
Kurt FischmanFounder, Marshal
Kurt is the CEO of Marshal, the Managed AI Ops company that designs, deploys, and operates AI agents as critical infrastructure for founder-led businesses.

An AI agent maturity model for business teams ranks how far a business can safely let agents own work, scored by supervision capacity rather than agent capability. For a small or mid-sized business, the binding limit is how many agent decisions one accountable person can genuinely review each week. That number, not the vendor demo, sets the real ceiling on how far the agents can go.
An AI agent maturity model for a business team is a staged framework that ranks how much work a business can safely hand to agents, and where along that path it currently sits. Most models split the journey into levels, from ad hoc experiments to autonomous operation, and attach a next step to each level so a team knows what to fix before it advances.
The framing matters because agents are not bought once and forgotten. An agent is an amount of work you have delegated, and delegation only holds if someone can still check the result. A maturity model is supposed to tell a business how much delegation it has earned. The published ones answer a slightly different question, and the gap between the two is where most small businesses get hurt.
We build AI Agent Systems for founder-led businesses, and the wall we hit first is never the agent's ceiling; it is the owner's calendar. That single observation reorders the whole framework. A maturity model that never asks about the calendar is measuring the wrong thing, confidently, in five neat stages.
Every agent maturity model on the first page of search measures the same thing: how autonomously the agent can act. A 2026-07-27 scan of the models ranking for this topic found the same axis under different labels. One vendor model runs five stages from Experimentation to Autonomous, scoring cross-system orchestration and multi-agent coordination. A major CRM vendor's model runs four levels from fixed-rule chatbots to any-to-any multi-agent operability. The stages differ in wording. The measuring stick does not: more autonomy, more systems touched, more agents talking to each other.
Each of those models is also calibrated for an organization that does not look like a small business. They assume a Center of Excellence, a security review board, role-based access at scale, a RACI chart with named owners for intake and incident response, and a change-management function to drive adoption. Those are reasonable things for a company of five thousand. They are science fiction for a company of fifteen, where the person setting agent policy is also the person answering the phone.
The consensus sells maturity as more autonomy. The unspoken truth is that autonomy you cannot supervise is not maturity, it is unmonitored liability with a nicer dashboard. An enterprise can afford to separate the agent's capability from the question of who watches it, because it has a department for watching. A small business cannot. For it, capability and oversight are welded together by a single scarce resource, and no chart on the market names that resource.
Supervision capacity is the number of agent decisions one accountable person can genuinely review each week before review turns into rubber-stamping. It is a budget, measured in attention, and it is far smaller than most owners guess. Reading an agent's output, deciding whether it was right, and correcting it when it was wrong is real work, and it competes for the same hours as everything else the owner does.
Here is the arithmetic that no capability model runs. An agent that qualifies inbound leads might produce forty decisions a week at first, each one a small judgment the founder can skim and approve in a minute. Scale the same agent to four hundred decisions and the founder now needs roughly seven hours a week just to review, assuming a minute each, which is optimistic. A founder who can review forty qualified-lead decisions a week cannot review four hundred, and no maturity chart on the market bothers to ask which number is theirs. The agent got more mature by every published measure. The business did not.
An SMB's agent maturity is capped not by what its agents can do but by how many agent decisions one accountable human can still review per week before review becomes rubber-stamping, and that ceiling, the supervision cliff, sits far lower than the capability the vendor demo shows. Buy the cleverest agent you can afford and you will still stall at the same place: the moment its output arrives faster than one person can read it. This is why agent readiness is a question about the reviewer, not the model. The model is almost never the bottleneck. The demo never shows the part where someone has to read what the agent did.
The four supervision levels sort a business by how it reviews agent work, not by how clever the agents are. Level one is manual with drafting: the agent proposes, and a person approves every action before it lands. Level two is exception review: the agent acts on the routine cases and routes only the ambiguous ones to a human queue. Level three is sampled audit: the agent owns the workflow, a person spot-checks a defined sample and reads the audit trail. Level four is governed autonomy: the agent runs unattended inside hard limits, with a kill switch and a periodic review of aggregate outcomes.
The cliff sits between whatever level a business is actually staffed for and the level its agent's capability invites it to attempt. An agent capable of level-four autonomy handed to a team that can only staff level-one review does not deliver level-four maturity. It delivers level-four volume with level-one oversight, which is the exact shape of the failures that make the news. Maturity is the lower of the two numbers, always.
The table below sets the supervision-scored levels next to the capability-scored levels the vendors publish, so the difference in what each one measures is visible in one place.
Comparison of a supervision-scored agent maturity model against the capability-scored models published for the enterprise, across five dimensions that decide whether a small business advances safely.
| Dimension | Supervision-scored model | Capability-scored model |
|---|---|---|
| What it measures | Reviewable agent output one owner can absorb per week | How autonomously the agent can act across systems |
| Binding constraint | The accountable person's review budget | The agent's technical ceiling |
| Who it fits | Founder-led and mid-market teams with thin headcount | Enterprises with a governance department |
| Signal you have stalled | Approvals become rubber-stamps under volume | Agents cannot reach the next autonomy tier |
| Next-step it prescribes | Widen the review aperture before adding volume | Buy more autonomy and more integrations |
The supervision-scored model is the only one that ties the next step to the resource an SMB actually runs out of first, which is a person's attention, not an agent's capability.
Finding your own supervision number takes an afternoon and a calendar, not a consultant. Start by picking one agent workflow you already run or plan to run, and write down a single figure: how many decisions it will generate in a normal week. A lead-qualification agent generates one decision per inbound lead. A support-triage agent generates one per ticket. An invoice-matching agent generates one per invoice. The unit is the decision a human would otherwise have made, because that is the unit someone now has to check.
Next, be honest about how long one review actually takes. Not the best case, where the agent was obviously right and the reviewer nodded. The average case, which includes the ones where the reviewer had to open the record, reconstruct the context, disagree with the agent, and fix it. For most judgment work that number lands between one and five minutes, and it climbs the moment the stakes rise. Multiply the weekly decision count by that per-review minute figure, and you have the weekly review load in hours.
Now put that number against the hours the accountable person actually has. A founder wearing four hats does not have seven spare hours a week to read agent output, and pretending otherwise is how the cliff gets crossed on paper before it gets crossed in production. If the review load exceeds the available hours, the business is not ready for that agent at that volume, no matter what the capability chart says. The gap between review load and available hours is the single most useful number in the whole exercise, and it is the number every published model leaves out.
The exercise also exposes the cheap fix hiding in plain sight. If review takes four minutes because the reviewer has to reconstruct context every time, then attaching the agent's reasoning and the source record to each decision might cut that to ninety seconds, which nearly triples capacity without touching the agent at all. Supervision capacity is not fixed. It is engineered, and the arithmetic tells you exactly which lever moves it.
The cost of crossing the supervision cliff is not one big failure. It is a slow drift into unmonitored operation that looks, from the inside, exactly like success. Volume goes up. The dashboard stays green. The owner stops opening individual records because there are too many, and starts trusting the aggregate. Nothing announces the moment oversight ended, which is precisely why it is dangerous.
The first real cost is the mistake that ships to a customer because no one read it. An agent misqualifies a lead and routes a serious buyer to a dead queue, or approves a refund that should have been flagged, or sends a confidently wrong answer under the company's name. In a supervised workflow a human catches these before they land. Past the cliff, the catch never happens, and the business learns about the mistake from the customer instead of the audit trail.
The second cost is subtler and worse: the erosion of the owner's own judgment about what the agents are doing. Once review becomes rubber-stamping, the owner loses the running feel for where the agent is weak, which cases it fumbles, and when its behavior drifts. That feel is what makes it possible to improve the system. Lose it, and the business is flying on instruments it has stopped reading. Recovering costs far more than the review time that was skipped, because trust, once it has quietly failed, has to be rebuilt decision by decision.
The point of scoring maturity by supervision is to make the cliff visible before the drift starts. A model that shows an owner exactly how much review a given agent volume demands turns an invisible risk into a line item on a calendar, and a line item can be planned around. That is the whole value of measuring the right thing: it converts a failure that arrives without warning into a constraint you can see coming and manage.
Moving up a supervision level means widening the review aperture before adding agent volume, in that order. A business that reverses the order buys the volume first and discovers the aperture too late, usually when a customer finds the mistake the owner never had time to catch. The sequence is not optional, and it is cheaper than the alternative.
Widening the aperture is concrete work, not a mindset. It means building an exception queue so the agent stops sending routine cases for review and escalates only the ambiguous ones. It means risk tiering, so a low-reversibility action gets a light touch and a high-impact one gets a hard gate. It means an audit trail the owner can sample instead of reading every record, and it means naming one accountable person rather than diffusing responsibility until no one owns the failure. These are the mechanics of agent governance, and each one buys back review hours without dropping oversight.
Only after the aperture is wider does adding agent volume make sense, because now the same person can absorb more decisions per week without the review degrading. This is the actual shape of the SMB agentification roadmap: not a calendar of features to ship, but a sequence of supervision upgrades that each raise the ceiling before the next batch of work is delegated. A team that treats maturity this way never falls off the cliff, because it moves the cliff before it moves the volume.
The supervision maturity model breaks in two places worth naming before you adopt it. The first is genuinely low-stakes, fully reversible work, where a wrong agent action costs nothing and no one needs to review it at all. Sorting internal notes, drafting throwaway first passes, tagging records that a human will re-check downstream anyway: for these, supervision capacity is not the constraint, and forcing a review budget onto them just adds friction. Score those by throughput and move on.
The second place it breaks is at real enterprise scale, where supervision is a staffed function rather than one person's spare attention. A company that can hire reviewers, stand up a monitoring team, and run a governance board has decoupled capability from oversight in exactly the way a small business cannot, and the capability-scored models were built for that world. The supervision model is a small-business and mid-market instrument. Businesses large enough to buy their way past the human bottleneck should ignore it and use the enterprise frameworks as intended. Everyone still living inside a single owner's calendar should score maturity the way that calendar actually works.
An AI agent maturity model for business teams is a staged framework that ranks how much work a business can safely let agents own, and identifies what to fix before advancing. The useful version for a small business scores each stage by supervision capacity, meaning how much agent output one accountable person can review, rather than by how autonomous the agents are.
The enterprise versions score maturity by agent capability: autonomy, cross-system orchestration, and multi-agent coordination. This model scores it by supervision capacity, because a small business runs out of review hours long before it runs out of agent capability. The enterprise models assume a governance department exists to watch the agents, which is the assumption that fails for a founder-led team.
A small business finds its stage by counting how it reviews agent work today, not by rating the agents. If a person approves every action, it is at manual review. If the agent handles routine cases and escalates exceptions, it is at exception review. The honest test is whether the reviewer still reads the output or has quietly started rubber-stamping it.
The supervision cliff is the point where agent volume outruns the review budget of the accountable person, and oversight silently degrades into rubber-stamping. Maturity models built around agent capability never show the cliff, because it lives in the human's calendar, not the agent's specification. Crossing it feels like progress until the first unreviewed mistake reaches a customer.
Moving up a level means widening the review aperture before adding agent volume, through exception queues, risk tiering, audit trails, and one named owner. Each of those buys back review hours so the same person can absorb more decisions without oversight slipping. Adding volume first, then scrambling for oversight, is the sequence that drives teams off the cliff.
Buying a more capable agent raises a business only on a capability chart, and does nothing for real maturity unless someone also gained the hours to supervise the extra output. A cleverer agent that produces more decisions per week without more review capacity lowers effective maturity, because it pushes the team closer to the supervision cliff rather than further from it.
Join hundreds of small businesses operating at machine speed with agents on the job.