Skip to content
← Insights

Recursive self-improvement · · 8 min read

RSI Agent: The Compounding Intelligence Layer for Every Company

The first generation of AI agents completes work. The RSI Agent improves the system doing the work — turning every task, failure, and correction into a compounding operational advantage.

An amber lattice of connected nodes throwing off sparks in the dark, its structure extending outward to points beyond itself.

The first generation of AI agents sells labour. The RSI Agent creates an asset that becomes more valuable every time the work gets done.

Every company investing in AI agents is about to confront the same uncomfortable truth: deploying an agent is easy; making it reliably better is hard.

Today’s agents can write code, answer customers, analyse contracts, run campaigns, and coordinate operations. But most are effectively reset after every task. They repeat avoidable mistakes. They depend on teams to inspect failures, rewrite prompts, adjust tools, add tests, and redeploy the system by hand.

That is not autonomous work. It is a new form of software maintenance.

The RSI Agent closes the loop.

RSI stands for recursive self-improvement. An RSI Agent observes how work gets done, finds the gap between the result and the goal, proposes a better approach, tests it against real evidence, and deploys only the changes that prove superior. Each cycle improves the system that runs the next cycle.

The distinction is fundamental:

A conventional agent completes a task. An RSI Agent improves the capability that completes every future task.

That turns AI from a rented capability into a compounding asset.

The biggest opportunity in AI is moving up the stack

Foundation models are becoming more capable, more available, and increasingly interchangeable for many business workloads. That is extraordinary for builders — but it also means durable value will not come simply from wrapping a model in a new interface.

Value moves to the layer that understands the customer’s goals, observes real outcomes, and continuously improves performance across models and tools.

The RSI Agent is that layer.

It sits above the underlying models and inside the operating flow of the business. It learns which instructions work, which tools fail, which sequences produce the best result, which exceptions require human judgement, and which changes improve performance without creating new risk.

The model may be rented. The improvement loop belongs to the customer — and to the platform that makes it possible.

This creates a powerful strategic position. The RSI Agent does not need to outspend frontier labs on model training. It benefits every time those labs release a better model, while preserving the proprietary evaluation data, workflows, and improvement history created in the customer environment.

In other words: foundation-model progress is not a threat to the RSI Agent. It is free leverage.

From system of action to system of improvement

The enterprise software stack has systems of record, systems of engagement, and — more recently — systems of action. The RSI Agent introduces the next category: the system of improvement.

Its operating loop is simple to understand and difficult to reproduce well:

  1. Observe the work. Capture agent traces, outputs, tool calls, costs, errors, approvals, and real-world outcomes.
  2. Identify the constraint. Determine whether failure came from reasoning, context, memory, instructions, tool selection, workflow design, or the underlying model.
  3. Generate improvements. Propose changes to prompts, policies, tools, tests, memory, routing, or agent architecture.
  4. Prove the gain. Run candidates against trusted evaluations, historical cases, simulations, and adversarial tests.
  5. Deploy safely. Promote only verified changes, with versioning, audit trails, staged rollout, and rollback.
  6. Compound. Make the improved system the starting point for the next cycle.

The magic is not that the agent can edit itself. Any model can rewrite a prompt. The magic is that an RSI Agent can distinguish a real improvement from a plausible-sounding change.

That makes evaluation — not generation — the heart of the product.

A flywheel built from proprietary outcomes

Every enduring technology company has a compounding advantage. Search engines improve from queries. Marketplaces improve from transactions. Networks improve as participants join.

The RSI Agent improves from work.

Every task produces a trace. Every human correction reveals a preference. Every failure becomes a regression test. Every successful exception becomes a reusable pattern. Over time, the RSI Agent builds an improvement graph connecting goals, contexts, actions, outcomes, and verified changes.

That graph becomes increasingly specific to how an organisation operates. It captures knowledge that rarely exists in documentation: the difference between the official process and the process that actually produces a great result.

This is the moat.

Competitors may access the same foundation models. They will not have the same history of attempted work, corrected failures, validated procedures, and outcome-linked evaluations. The longer an RSI Agent operates inside a company, the more expensive it becomes to replace — not because the customer is locked into static data, but because the system has accumulated tested judgement.

The resulting flywheel is unusually attractive:

More work → more outcome data → better evaluations → safer improvements → higher reliability → more work.

This is not a data flywheel built on passive collection. It is a performance flywheel built on verified learning.

The wedge is reliability. The destination is the AI control plane.

The immediate pain in enterprise AI is not access to intelligence. It is reliability.

Companies can already prototype impressive agents. What stops them from handing those agents consequential workflows is the distance between “usually works” and “can be trusted.” Closing that gap requires continuous evaluation, failure analysis, controlled experimentation, and governance — the exact capabilities an RSI Agent provides.

That creates a focused entry point: help companies improve one high-value agent in one measurable workflow. Demonstrate fewer escalations, higher task success, lower cost, or faster completion. Then expand.

Once the RSI Agent becomes the trusted improvement layer for one workflow, it can extend across:

  • additional agents and departments;
  • multiple foundation-model providers;
  • tool and workflow optimisation;
  • shared evaluation infrastructure;
  • permissions, approvals, and policy enforcement; and
  • organisation-wide intelligence about where automation works and where humans add the most value.

The wedge is an agent optimiser. The destination is the control plane through which a company’s entire AI workforce learns.

That supports a business model with both recurring platform revenue and usage-based expansion. As customers delegate more work, the RSI Agent processes more traces, runs more evaluations, manages more deployments, and creates more measurable value. Revenue grows with adoption; retention grows with accumulated intelligence.

Why now

Recursive self-improvement has existed as an idea for decades. It is becoming a company now because five curves have converged.

Agents can act. Models can use tools, write and execute code, inspect environments, and operate across long workflows.

Work is measurable. Agent traces, business outcomes, simulations, and human corrections provide the evidence needed for disciplined improvement.

Models can perform AI R&D. Google DeepMind’s AlphaEvolve has used language models and automated evaluation to discover and optimise algorithms. The Darwin Gödel Machine demonstrated agent designs that generate, evaluate, and select improved descendants.

AI is entering the development loop. OpenAI’s GPT-Red uses automated adversarial iteration to find weaknesses and improve the robustness of later systems. Anthropic reports that AI now contributes substantially to its own engineering work and is taking on increasingly open-ended development tasks.

Enterprises need a neutral layer. No serious company wants its operational learning trapped inside a single model vendor. A model-independent improvement system lets customers use the best available intelligence while retaining their workflows, evaluations, governance, and accumulated learning.

These are the prerequisites for the RSI Agent — and they have arrived at the same time.

Safety is not a constraint on the product. It is the product.

The phrase “self-improving AI” can reasonably trigger concern. An unrestricted agent changing production systems would be dangerous and commercially unusable.

The RSI Agent should be positioned around controlled compounding, not uncontrolled autonomy.

It operates within explicit boundaries. Candidate changes are tested before deployment. High-impact modifications require approval. Independent evaluators challenge the system. Permissions follow least-privilege principles. Every change is versioned and attributable. Rollouts are staged, monitored, and reversible.

Most importantly, the agent proposing a change is not the sole judge of whether that change is good.

This architecture turns safety into a market advantage. The enterprise buyer does not merely want a more capable agent. The buyer wants evidence that the agent is improving for the right reasons, within policy, without silently degrading elsewhere.

The RSI Agent makes that evidence visible.

A new category — and a much larger ambition

The first wave of AI companies made intelligence accessible. The second made it actionable. The RSI Agent makes it cumulative.

That is a far larger idea than automating individual tasks. It changes the nature of enterprise software from something a company installs into something the company cultivates. The product does not merely execute the organisation’s playbook; it helps the organisation discover a better playbook.

In the near term, that means agents that stop repeating the same mistakes. It means failed runs automatically become tests. It means teams spend less time tuning prompts and more time deciding what outcomes matter.

Over time, it means every company can develop a proprietary intelligence layer shaped by its customers, operations, standards, and judgement — a layer that becomes more capable as the company uses it.

The defining question of the next AI cycle will not be, “Which company has access to the smartest model?”

Every company will.

The defining question will be, “Which company has the fastest, safest, and most defensible improvement loop?”

The RSI Agent is built to own that loop.

Sources and further reading