A headless system is one whose consumer is a program
On May 21, 2026 we scored 447 software vendors on how well a machine can operate them. The number that mattered was not the average. It was this: an automated collector could fully assess only 220 of the 447. For the other 227, the machine surface could not be found, fetched, or resolved well enough to grade all five dimensions.
That is not a defect in the scorer. It is the definition of a headless system arriving as a measurement. A headless system is one whose primary consumer is a program, not a person. Under that definition, a system whose machine surface a program cannot even discover is not headless, whatever its architecture diagram claims.
This page is the definitional reference for the term this publication is named after, and it does three things. It defines a headless system by its consumer rather than by which layers were split apart, and argues why that beats the incumbent architectural definition. It separates the three unrelated things the word headless now means, because all three arrive on this domain from search and only one is the subject here. Then it uses The Headless Index to show which layers of the stack are headless systems today and which are not, with the measured evidence for each.
What a headless system actually is
Here is the definition, stated so it can be argued with.
A headless system is one whose primary consumer is a program, not a person. Its contract is a machine interface: a specification, a schema, an event stream, an authorization model. Any human interface it has is a downstream client of that contract, not the contract itself.
Most published answers to what is a headless system define headless architecture instead, and the incumbent definition is worth quoting exactly rather than paraphrasing. One of the category’s own vendors puts it this way: headless architecture “separates the front end (user interface) from the backend (content management and data storage)”, and the term refers to removing the “head” (the presentation layer) from the body of your content management system.
That was the right definition for the problem it solved: one content body, many display channels. The question it answers is which layers were decoupled from which, and it answers that well.
It has stopped discriminating. It classifies a headless CMS serving a marketing site read only by humans as headless, and a monolith with a complete, versioned, well-authorized internal API as not headless. Every practitioner knows the second system is easier for a program to operate. A definition that gets that backwards is describing the diagram, not the software.
Defining a category by its consumer is not a trick invented for this page, it is what specifications in this space already do. The A2A protocol documentation draws its own boundary against MCP that way, saying the distinction “depends on what an agent interacts with”. Tools and resources are “primitives with well-defined, structured inputs and outputs”, agents are “more autonomous systems” that “reason, plan, use multiple tools”, and the protocol you reach for depends on which one is on the other end. Same move: the consumer determines the category.
The consumer definition has a cost, and it should be stated up front. Under it, several celebrated headless CMS products are not headless systems. The index numbers later in this piece are the receipt for that claim.
Three different things are called headless
Three unrelated meanings share this search term, and our own Search Console data for headless-systems.com shows impressions on all three in the same 28-day window. Only the first describes a headless system in the sense defined above, and only the first is the subject of this page.
Software whose consumer is a program.
A headless system in the sense used here, and the sense the term is heading toward.
This is the meaning people are reaching for when they search for headless software, and it carries our highest-volume queries: headless system (279 impressions) and headless software (115).
A machine with no attached display.
Headless browsers and headless servers.
Chrome’s own documentation says that with Chrome Headless mode “you can run the browser in an unattended environment, without any visible UI”.
Queries landing here: headless machine, headlessly, and headless devices.
If this is what you came for, the Chrome documentation is the right page and this one is not.
Unstyled component libraries.
Headless UI kits and headless design systems ship behavior and accessibility without markup or styling.
Headless UI describes itself in one line: “Completely unstyled, fully accessible UI components, designed to integrate beautifully with Tailwind CSS.”
Query: headless design system (24 impressions).
The consumer of these components is still a human looking at a screen.
The real distinction is not which layer got removed. In meanings two and three, the head was removed for convenience: no display to allocate, no opinions to override. In meaning one, the consumer changed. That is a different kind of fact, and it is the one this publication tracks.
More than half the index could not be fully read by a machine
The Headless Index scores vendors on five dimensions, each out of 20: API-first design, headless operation, MCP posture, schema and observability, and webhooks and events. The snapshot discussed here was scored on 2026-05-21 under methodology version THI-v1.0, and covers 447 vendors across 13 categories. The headline average is 43.3 out of 100, median 45.
Ignore that average, because the interesting number is the denominator.
All five dimensions were measurable for 220 of 447 vendors. For the remaining 227, at least one dimension came back unresolved, and the score was computed against a reduced denominator: 93 vendors scored out of 40, 95 out of 60, 37 out of 80, and two out of 20. The mean count of unmeasured criteria per vendor across the index is 1.15.
Split the population and the industry looks like two different industries. Among the 220 vendors a collector could fully assess, the mean is 60.6 and the median 60.0. Among the 227 it could not, the mean is 26.5. The 43.3 headline is a blend of those two groups, not a distribution around a center, and reporting the blend hides the finding.
The same structure explains the band split, which is 177 F, 129 C, 113 B, 21 D, and 7 A. It is tempting to read 177 F grades as 177 vendors hostile to machine consumption, but the data does not support that. Of the 177 F-band vendors, 164 (93%) were graded on only two or three of five dimensions, and 87 were graded on exactly two. Their recorded rationales cluster on two values, a Headless Index score of 13 or of 25, and for 173 of those 174 vendors the score is an artifact of a partial denominator rather than a verdict on the 40% of the index that sits in the F band.
So the F band is a discoverability finding, not a hostility finding, and that distinction is the whole point. If an automated collector following documentation links cannot establish whether a vendor has webhooks, an agent operating that vendor in production cannot either. Undiscoverable and absent are the same fact from the consumer’s side, and the consumer is what the definition turns on. A headless system whose seams no program can locate is not a headless system in production.
One more structural number, which corrects an earlier internal reading of our own data. 117 of 447 vendors (26.2%) publish a machine-reachable public OpenAPI specification. The other 330 record the same reason string: no public specification discovered during collection. Spec publication is bimodal rather than scarce. Among the 116 whose specification could be scored, the mean JAIRF AI-readiness score is 75.4 out of 100, median 77.3, and the foundational compliance dimension that most directly reflects specification quality averages 83.4. Either you have a spec and it is decent, or you have nothing at all.
That threshold is load-bearing in the rubric. All seven A-band vendors publish a machine-reachable spec, and the band rules make the A band unreachable without one: no discoverable specification means no JAIRF score, and an unscored JAIRF caps a vendor at B by construction. The seven are Stripe (100), Grafana (78), Honeycomb (78), Razorpay (76), Zapier (76), Metronome (75), and Polar (75).
Which layers of the stack are headless systems now
This is the part no competing definition page can write, because it requires having measured something.
Rank the 13 layers by mean score and by the cleanest binary signal in the dataset, whether a vendor publishes a machine-reachable specification.
| Layer | Vendors | Mean thesis-fit | Machine-reachable OpenAPI spec |
|---|---|---|---|
| Feature flags and config | 15 | 55.3 | 5 of 15 (33%) |
| Analytics and events | 16 | 49.5 | 3 of 16 (19%) |
| Commerce | 18 | 47.6 | 6 of 18 (33%) |
| Object and file storage | 21 | 47.5 | 4 of 21 (19%) |
| Search and vector databases | 20 | 47.1 | 8 of 20 (40%) |
| Observability | 45 | 46.8 | 11 of 45 (24%) |
| Auth and identity | 41 | 45.7 | 13 of 41 (32%) |
| Payments | 54 | 45.3 | 20 of 54 (37%) |
| Project and task management | 30 | 43.0 | 8 of 30 (27%) |
| Content management | 22 | 42.0 | 1 of 22 (5%) |
| Communications | 35 | 41.2 | 9 of 35 (26%) |
| Workflow and automation | 56 | 38.7 | 11 of 56 (20%) |
| AI platforms | 74 | 36.2 | 18 of 74 (24%) |
Specification publication varies more than eightfold across layers, from 40% in search and vector databases down to 1 of 22 (4.5%) in content management, the table above rounding each layer to whole percent. Content management is the layer that invented the word headless, and it is last on the one signal that most directly determines whether a program can operate a system it has never seen before. That is the cost of the consumer definition, paid in public.
Now the pattern, stated as a position: the layers that became headless systems are the layers whose consumers were already programs.
Feature flags exist to be read by a running process, and they lead the index at 55.3. Analytics and events exist to be written by instrumentation. Search and vector databases exist to be queried by application code. Observability exists to be scraped and alerted on. None of those layers adopted API-first architecture as a strategy. They were never anything else, and the index is measuring how much of that native machine surface is documented and reachable.
An independent survey converges on the same shape from the demand side. Stacklok’s State of MCP in Software 2026, fielded in December 2025 across 100 senior technical leaders at software companies, reports that “Software developers are the primary users (80%), followed by data analysts and scientists (68%)”, with top use cases of test generation (68%), code review and QA automation (67%), and debugging production issues (56%). Nearly half of those companies report MCP in production, with 19% describing it as broad production, though the same report’s familiarity question puts hands-on production experience lower still. Every leading use case in that survey is a developer tool.
Put the supply-side ranking and the demand-side survey together and the honest reading is narrower than the slogan. Machine consumption is real, measurable, and concentrated in the developer-adjacent layers. For payments, commerce, content, and the back office, “the consumer is a program” describes where those layers are going, not where they are. That is a prediction, and it should be read as one. The escape-hatch dynamic in headless ERP is what the transition looks like from inside a layer that has not made it yet.
The vendors selling agents ship the least agent-consumable APIs
AI platforms is the largest category in the index at 74 vendors, and it finishes last of 13.
Mean 36.2, median 25.0, a gap of more than eleven points and the signature of a category carried by a handful of outliers. MCP posture averages 4.07 out of 20, with 44 of 74 vendors scoring zero. Webhooks and events averages 2.46 out of 20, the worst of any layer, with 50 of 74 scoring zero. 40 of the 74 land in the F band.
The outliers are real: OpenAI at 86, Replicate at 76, Vapi at 75, Azure OpenAI and Hugging Face at 74. Behind them sits a long tail of companies selling agent capability through interfaces other programs cannot drive, subscribe to, or discover.
Payments produces the same shape from a different direction. It has the worst MCP posture of any layer at 3.15 out of 20, with 29 of 54 vendors at zero and only three of 54 scoring 15 or better. That was measured months after UCP, ACP, and AP2 had all shipped to let agents buy things: ACP and AP2 in 2025, UCP in January 2026. The protocol layer moved. The layer underneath it did not.
There is a general lesson here about how to evaluate any claimed headless system or headless solution. Marketing an agent is cheap. Publishing a specification, a webhook catalog, and a scoped authorization model is expensive, verifiable, and rare. Gartner’s own estimate that of the thousands of vendors claiming agentic solutions only around 130 offer genuine agentic features is the same finding measured on the other side of the API.
What these numbers can and cannot tell you
A permanent definitional page that buries its caveats deserves to be attacked on them, so here they are.
The snapshot is dated: all 447 vendor files carry a scoring timestamp of 2026-05-21. Every count and percentage on this page is the published state of that one scoring batch as it stood on publication day, 2026-08-21, and it is left frozen there on purpose. The Headless Index keeps adding vendors, so the live index will report a population larger than 447 and ratios that have moved with it. Read the figures here as a single dated measurement, and the index itself as the current one. This is a single-day measurement, not a rolling one, and it says nothing about vendors that shipped a specification or an MCP server afterwards.
The denominator varies, and unmeasured criteria are recorded as a score of zero in the sub-score while being excluded from the denominator. That makes three of the five dimension averages partly measurement artifacts: webhooks and events was assessable for 222 of 447, headless operation for 259, and API-first for 355. The effect is large enough to change what a number means. Webhooks and events averages 4.95 out of 20 across all 447, which looks like the weakest dimension in the index. But webhooks could only be assessed for 222 vendors, and among those 222 the mean is 9.96 with no genuine zeros. So 4.95 measures undiscoverable webhook documentation, not absent webhooks, and anyone reporting it as the latter is overclaiming.
MCP posture survives that test and should carry the argument instead. It was assessable for 439 of 447 vendors, and 197 of its 205 zeros are genuine. The average across those 439 is 5.59 out of 20, and only 60 vendors (13.4%) score 15 or better, eighteen months after the protocol shipped. On the all-447 basis MCP posture averages 5.49, and for reference on that same basis schema and observability leads at 10.50, API-first sits at 9.42, and headless operation at 7.57. Schema and observability also has the widest coverage in the index, assessable for 446 of 447.
The deeper limitation is what the rubric measures at all. It measures whether a machine surface exists, is reachable, and is described. That is a necessary condition for machine consumption, not a sufficient one. Kai Pan’s Agent-First Tool API paper argues that the tool interfaces agents consume “remain rooted in human-oriented CRUD paradigms”, and names five mismatches between conventional APIs and agent requirements: exact-identifier dependence, rendering-oriented responses, single-shot interaction assumptions, user-equivalent authorization, and opaque error semantics. On 50 real operational tasks it measures 88% end-to-end task success for agent-first interfaces against 64% for optimized CRUD baselines, and positions the paradigm as a semantic layer above transport standards like MCP rather than a replacement for them. A vendor can score well here and still be undriveable by an agent that has to infer preconditions, success criteria, and safe retry behavior from endpoint names.
Which direction does the error run? The index can call a vendor headless when an agent still cannot operate it. It cannot call a vendor headless when there is no machine surface to find. A headless system measured this way is one a program can locate and address, which is less than one a program can reason about. So these numbers are a floor on the industry’s problem, not a ceiling.
MCP shipped a UI, and the definition holds
Now the hardest case for the consumer definition, and it comes from the protocol built for machine consumption.
MCP shipped an official UI layer. MCP Apps lets a tool return interactive UI components rendered inside the conversation, and the maintainers’ own framing is that tools alone were not enough: MCP is good at connecting models to data and letting them act, but there is a context gap between what tools can do and what users can see. If the canonical protocol for machine consumption concluded it needed HTML, a definition built on the absence of a human consumer looks like it is defining away something the ecosystem just added back.
The objection is real and it does not get hedged. It gets answered by asking who holds the contract.
With MCP Apps, the agent is still the client.
The interface is a resource the agent fetches over a ui:// URI and hands to a renderer, chosen and scoped by the agent, addressed by the same machine contract as any other resource.
The human sees the result, but the human is downstream of the machine rather than the party the interface is built for.
That is the inversion this definition is actually claiming.
Headless never meant no pixels. It meant the pixels stopped being the interface. Under the consumer definition, a service that returns a rendered component to a program that asked for one is more of a headless system than one whose only surface is a dashboard a human logs into.
The rest of the same specification release supports that reading rather than undercutting it.
The 2026-07-28 MCP specification removed the initialize and initialized handshake and the Mcp-Session-Id header, replacing a bidirectional stateful design with a stateless protocol core in which every request is self-describing.
Sessions are a human affordance.
Removing them is a protocol admitting that its consumer is a program that does not hold a connection open.
We took this tension apart in more detail when the specification landed: MCP shipped a UI and the thesis held.
Scale is not the argument here, but it is not nothing. The protocol’s lead maintainers report Tier 1 SDK downloads close to half a billion a month, with the TypeScript and Python SDKs each past a billion cumulative, and in the same post they cite Honeycomb reporting that nearly 20% of its monthly interactive queries are now made by agents rather than people. That figure is Honeycomb’s own, reported secondhand, and scoped to interactive queries at one vendor, but it is the closest available measurement of a consumer flipping from human to program, and Honeycomb is one of only seven vendors in the A band.
Where the headless system definition runs out
Two more objections, both stronger than the usual ones.
The first is that the demand may not arrive. Gartner forecasts that 40% of enterprise applications will feature task-specific AI agents by the end of 2026, up from less than 5% in 2025. The same firm, and the same analyst, Anushree Verma, also forecasts that over 40% of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value, and inadequate risk controls. A January 2025 Gartner poll of 3,412 webinar attendees found only 19% had made significant agentic AI investments, against 42% investing conservatively and 31% still waiting to see. Verma describes most agentic projects today as early-stage experiments driven by hype and often misapplied. Both forecasts are on the record, the first for a deadline this piece publishes inside with no verification that it landed, and the tension is unresolved.
The case for building headless never depended on either forecast. A headless system can be consumed by whatever shows up, agents included. A system without a machine surface cannot, and retrofitting one later is the expensive direction. That argument holds whether the agent projects succeed or get canceled, which is why it is the one worth building on.
The second objection is that computer-use agents route around the problem. If an agent can operate a human interface, software that never went headless is still reachable, and the case for building a machine surface weakens. This is the objection that got stronger after we scored the index, and it should be conceded rather than argued with. A critical survey of GUI agents records the OSWorld human baseline at 72.36%, a number agents were nowhere near in 2024, when a GPT-4V baseline scored about 12%. Agents passed that baseline in late 2025, and by June 2026 the top OSWorld-Verified results stood at 85.4%. The same survey still finds agents “nowhere near autonomous operation on open-ended desktop tasks”, but on this benchmark the human gap has closed and inverted.
But performance is the wrong axis, because performance moves and definitions should not. The definitional point is about the contract.
scraped system agent renders /orders, finds a button labelled "Refund"
breaks on: redesign, A/B test, locale change, new modal
contract: pixels, unversioned, unannounced, unsigned
headless system agent calls POST /v1/refunds {charge, amount}
breaks on: a versioned, documented, announced change
contract: a specification someone published on purposeA browser agent does not turn a product into a headless system. It makes it scraped. The contract is pixels, nobody signed it, and it changes without notice. From the agent’s side a scraped monolith and a headless service can look identical for exactly one release, and then they diverge. That is the distinction the consumer definition exists to draw, and it is precisely why defining by consumer beats defining by diagram.
The consumer is the definition
A headless system is one whose primary consumer is a program, not a person. It does more work than the architectural definition because it survives the cases that one fumbles. It separates a programmable monolith from a decoupled CMS only humans read, and it reads an agent fetching a UI resource as machine consumption rather than a retreat to the dashboard.
It also turns a semantic argument into an empirical one, which is the part that matters. If the consumer defines the category, then the question of how much software is headless becomes measurable, and the answer as of the 2026-05-21 snapshot is uncomfortable. More than half of a 447-vendor index could not be fully read by a machine following its own documentation. 26.2% publish a specification a program can fetch. 13.4% have a credible MCP posture. The layers that clear the bar are the layers whose consumers were programs all along, and the layer selling agents finishes last.
For anyone evaluating a vendor, this converts a marketing question into three checks. Is there a published specification a program can fetch without a sales call? Can the system tell you something changed without being polled? Can you scope what a non-human caller is allowed to do? A headless solution that fails those checks is a dashboard with an API attached, and the index says most of them are.
That is the gap this publication was started to measure, and it is why the protocol layer matters more than the architecture diagram. The word headless will eventually stop being a label, the way client-server did, because a machine contract will just be what software has. The interesting question is which layers get there on purpose and which get scraped in the meantime.