Agentic Product & Software Development

From signal to outcome without a human clock.

I use agents inside the product and in the work on the product. The target is autonomous signal-to-outcome loops that triage, clarify, build, verify, release, observe, and keep learning without waiting for a person at every step. Current evidence proves the building blocks; the adaptive product system is the target, not a finished product.

How I see it

The real step change is not more generated code. It is agents carrying signals autonomously through to a verifiable outcome without people triggering every next step.

Two changes, one system

Agents become part of the product and part of the work on the product.

A technical foundation for sound AI and ML judgment

My path into AI builds on my Cognitive Informatics studies: machine learning, robotics, AI, agent and multi-agent systems, neural networks, and the question of how systems perceive, decide, and fail. That gives me judgment around architecture limits, model limits, and realistic use cases.

Agents work on the product and inside it

I have used LLMs productively every day since February 2023. Since late 2025, most of my implementation has been produced through agent workflows: I frame outcomes, context, and constraints, orchestrate agents, review code and evidence, and remain accountable for product and technical decisions. I use the same approach for product clarification, research, architecture, tests, review, documentation, and operations, as well as for products where agentic capabilities become part of the user value.

Engineering method instead of prompt magic

User value and domain language lead. DDD with bounded contexts keeps responsibility legible, while TDD protects behavior without exceptions. Disposable spikes, prototypes, and grilling sessions burn down risk before assumptions become architecture.

Rules, skills, tools, and artifact learning

Agentic work becomes reproducible when context is cut cleanly, rules are explicit, permissions are clear, and reusable skills and tools exist. Corrections from tests, reviews, and operations return as reviewable changes to rules, tests, documentation, or runbooks.

Small increments with durable continuation

Every slice should deliver a useful end-to-end path. Durable run state, checkpoints, retries, observability, and explicit wake conditions keep token limits, session changes, or one failure from stopping the entire loop.

People remain valuable

Not as schedulers, but where different capabilities matter.

People provide intent, domain evidence, and policy, enable new tools, safeguard sensitive credentials and secrets, correct through feedback or veto, and handle break-glass cases. The system should request those contributions deliberately instead of making routine work wait for approval.

User value and domain boundaries before tool or model choice
DDD with explicit bounded contexts and TDD without exceptions
Disposable spikes, prototypes, and grilling to burn down risk early
Smallest useful end-to-end slice instead of a big bang
Durable run state, observability, rollback, and human veto

The agentic value stream

From real signals to outcomes and back again.

Signal & triage

People and systems provide customer, support, market, and operational signals. Agents deduplicate, prioritize by user value, and start the next permitted step.

Problem brief & bet

DDD and bounded contexts keep language and responsibility clean. Disposable spikes, prototypes, and grilling burn down the largest uncertainty early.

Build, validate, operate

TDD without exceptions, DevSecOps, and observability move with small, useful end-to-end increments. A passing slice becomes a candidate for safe integration and operation only with rollout, rollback, and policy evidence.

Outcome review & learning memory

Durable run state, checkpoints, retries, and wake conditions survive session and token limits. Evidence returns as reviewable tests, rules, skills, and runbooks.

From task to outcome

Seven maturity levels, one evidence-based path.

The public maturity model moves from directly supervised tasks through durable workflows and a governed value stream to Level 7: an adaptive product system. There, agents select bounded problems and experiments within explicit goals, budgets, and kill criteria. People retain strategy, accountability, and stop authority.

The public method, private implementation, and commercial target remain deliberately separate. This keeps current evidence visible without claiming product maturity that is still being built.

  1. 01 · Public Harness

    Agentic Engineering Harness

    Open-source skills, contracts, maturity models, and checks that others can use to build their own harness.

  2. 02 · Private Harness

    Concrete multi-repository operations

    My Gitea-first implementation puts context routing, worktree leases, reviews, quality gates, and integration into practice across real repositories.

  3. 03 · Value Pipeline

    Level 7: adaptive product system

    The target carries trusted signals through bounded product decisions and experiments to measured outcomes and new investment decisions. Value Pipeline is in development and is being proven first on my own system.

Technical target system

Provider-neutral, cost-aware, and visible to people.

Gitea-first control plane

Git, issues, changes, reviews, and CI remain legible inside the private system. Public artifacts can continue to live on GitHub.

Provider-neutral model routing

Cursor Headless can be a subscription-backed worker adapter, but it is not the platform. Models and providers should be selected by difficulty, risk, data access, tool needs, evidence, and cost. Premium speed tiers such as Fast Mode stay off by default.

Durable resume

The target gives runs state, claims, checkpoints, retry rules, and reconcilers. Pausing without a signal is healthy; a forgotten human trigger is not an operating model.

Human-readable projection

The target makes signals, decisions, evidence, cost, risk, escalation, and outcomes visible to people so they can add context, feedback, or veto at any time.

Practice and direction

I am building this out of real product work.

At Livable Places, I connect multi-agent research, local agent skills, and plugin-style workflows with product clarification, data and ML evidence, code, tests, delivery, and platform operations. Corrections do not disappear into chat. They become reviewable artifacts and better project context.

My public Agentic Engineering Harness makes this approach inspectable as an open-source project: portable Skills, repository and multi-repository rules, independent reviews, versioned learning artifacts, and explicit escalation boundaries. The goal is maximum evidence-backed autonomy - not human approval for every commit.

Where this creates value

For products and teams that take agents seriously from user value through operations.

We can start with a concrete product problem, a constraint in the delivery flow, or the next sensible step in autonomy. What matters is shaping user value, accountability, quality, and operations together.