Agentic Engineering maturity model

From one task to an adaptive product system

The seven levels describe the largest unit that can be delegated reliably. They do not measure model intelligence or the number of installed tools. Each level needs its own evidence, boundaries, and human responsibility.

The short version

Autonomy does not grow through a claim. It grows when a smaller loop works repeatedly, fails visibly, and can continue or stop safely.

The progression

Each level expands the loop that can close reliably.

Higher levels build on proven lower ones. Each decision domain is assessed separately, not the team as a whole. Where authority, recovery, or evidence is weaker, that dimension limits the safe level.

  1. 01

    Level 1: directly supervised task

    System
    An agent supports one clearly bounded task in one session.
    Human
    A person frames the task, provides context, and reviews the result directly.
    Minimum evidence
    A representative task succeeds under direct supervision.
  2. 02

    Level 2: repeatable procedure

    System
    Versioned instructions or skills stabilize a recurring method.
    Human
    A person selects the procedure, supervises execution, and evaluates repeated results.
    Minimum evidence
    Repeated runs outperform unstructured prompting on relevant examples.
  3. 03

    Level 3: living repository

    System
    Context, decisions, checks, and learning survive across sessions.
    Human
    People own intent and material decisions. Agents maintain bounded repository artifacts.
    Minimum evidence
    Context, owners, fast checks, full gates, and learning are discoverable.
  4. 04

    Level 4: grounded system

    System
    Sources and tools provide attributable evidence and controlled external actions.
    Human
    People set access and authority boundaries and decide consequential external effects.
    Minimum evidence
    Sources, permissions, currentness, failure paths, and audit evidence are explicit.
  5. 05

    Level 5: stateful workflow

    System
    A bounded end-to-end workflow keeps state, waits deliberately, retries, recovers, or stops.
    Human
    People design boundaries and handle exceptions instead of scheduling every next step.
    Minimum evidence
    Repeated runs are observable, recoverable, reconciled after missed progress, idempotent where needed, safely stoppable, and evaluated against outcomes.
  6. 06

    Level 6: governed value stream

    System
    The system carries signals within explicit goals and risk limits through production evidence and the next decision.
    Human
    People govern goals and risk, retain veto and incident authority, and intervene on exceptions.
    Minimum evidence
    Isolation, policy and quality gates, rollback, audit, incident ownership, production feedback, and effective oversight are proven.
  7. 07

    Level 7: adaptive product system

    System
    Within a bounded decision domain, the system selects problems, experiments, and investments from trusted signals and learns from outcomes.
    Human
    People own strategy, budgets, decision domains, kill criteria, and stop authority.
    Minimum evidence
    Level 6 controls, trusted product signals, data and experiment boundaries, budgets, kill criteria, and human-governed investment decisions are proven.

Not a status model

Not every piece of work should reach Level 7.

A clearly bounded task belongs at Level 1. A recurring workflow may be complete at Level 5. The highest number is not the goal. Delegation, risk, and evidence need to fit the work.

More than technology

From Level 5 onward, autonomy changes the working system.

Responsibility gets redrawn.

People execute fewer individual steps. They shape goals, rules, decision authority, exceptions, and escalation.

Evidence replaces approval theatre.

Tests, policies, observability, rollback, and outcome signals carry routine decisions. People review new meaning and material risk.

Product, delivery, and operations become one loop.

Signals are triaged, tested through small experiments, observed in production, and used for the next investment decision.

Evidence and target

The boundary between today and tomorrow stays visible.

My public Agentic Engineering Harness proves reusable methods, skills, rules, and checks. The private implementation proves operation across real repositories. Value Pipeline is the commercial Level 7 target and remains in development.

View the agentic working model

Which loop can your system genuinely delegate today?

We do not start at Level 7. We start with a recurring constraint, an explicit boundary, and the smallest level that creates measurable value.