Skip to main content

Does AI Memory Actually Work for Coding Agents?

We benchmarked it three times and it lost the first round. Here is everything, including that.

Memory does not make a model smarter — it stops it re-learning what it already worked out. That single idea explains every number below, including the ones that did not go our way.

Markus Sandelin — Independent Research, February 2026

The shape of the answer

Three benchmarks, increasing in scale, and the result flips as they grow.

On a toy codebase, memory lost outright. There was nothing worth remembering, so every recall was overhead with no payoff. We are told this is the sort of result you quietly drop; we put it first, because it is the clearest evidence of what memory is actually for.

On a real production codebase — 4,895 lines, 158 files — quality did not move at all. Both conditions landed in the same band. What changed was how much work it took to get there: fewer turns, less exploration, less re-reading of files the agent had already read in a previous session.

On a six-agent swarm, the lead agent stopped exploring the codebase altogether. It recalled the architecture and went straight to delegation.

And at the bottom of the model range, the result stops being about efficiency. A model that scored zero on the task scored 39 out of 40 with memory available. Nothing about the model changed.

The pattern: memory is worth exactly as much as the context you would otherwise lose. On a small problem, that is nothing.

Four Key Findings

Four things came out of this, and only one of them is the sort of thing a vendor puts on a homepage.

1

Memory doesn't improve code quality

Quality ceiling is a property of the model (84-96% across all conditions).

2

Memory reduces exploration overhead

28-40% fewer turns, 22-32% lower cost on complex tasks.

3

There's a complexity threshold

On trivial codebases, memory is pure overhead.

4

Memory enables smaller models

Haiku 4.5 scored 0/40 without memory, 39/40 with it. Memory isn't just efficiency — it's capability for smaller models.

What Everyone Else Claims

Industry claims vs. what they actually measure.

SystemClaimed SavingActually MeasuresBenchmark
Mem090% tokensMemory compressionLOCOMO
A-Mem85-93% tokensPer-operation costDialogue QA
MemMachine80% tokensRecall accuracyLOCOMO
Zep94.8% accuracyRetrieval precisionCustom
LettaN/A (honest)Agent capabilityTerminal-Bench
Stompy15-28%Task efficiencyCoding tasks

Methodology

Worth being specific about what was measured, because “we benchmarked it” is the least falsifiable sentence in software.

  • System: MCP-based, PostgreSQL, VoyageAI embeddings
  • Codebase: 4,895 lines Python/FastAPI, 158 source files
  • Three conditions: stompy (MCP recall), file (static CONTEXT.md), nomemory (cold start)
  • Three tasks of increasing complexity
  • Scoring: 25-point rubric (5 criteria x 5 points)
  • All runs: Claude Opus 4.6, identical codebase snapshot

Results

Two tables. The first is per-task; the second is the one that matters, because it shows the result moving with complexity rather than sitting still.

Table 4 — Per-Task Results

TaskStompyFileNoMemory
Task 1 (Moderate)23/25 $1.90 35t23/25 $1.80 29t24/25 $1.17 19t
Task 2 (High)21/25 $1.33 31t22/25 $3.22 51t21/25 $3.51 58t
Task 3 (Very High)23/25 $3.52 47t23/25 $3.16 44t22/25 $3.18 54t

Table 5 — Aggregate

ConditionQualityCostTurnsCost/Point
Stompy67/75$6.75113$0.101/pt
File68/75$8.18124$0.120/pt
NoMemory67/75$7.86131$0.117/pt

Table 6 — Complexity Gradient

TaskWinnerTurn SavingsCost Savings
Task 1 (Moderate)nomemory-35% turns-20% cost
Task 2 (High)stompy+28% fewer turns+22% cost savings
Task 3 (Very High)stompy+18% fewer turns+22% cost savings, +32% time savings

Table 7 — Phase 3: Multi-Agent Swarm Results (6 agents, full-stack booking feature)

ModelConditionScoreCostTurnsTime
Sonnet 4.6stompy40/40$3.9826.5m
Sonnet 4.6nomemory40/40$7.0449.6m
Opus 4.6stompy40/40$4.34299.6m
Opus 4.6nomemory40/40$7.657010.0m
Haiku 4.5stompy39/40$4.9527.5m
Haiku 4.5nomemory0/40$3.9735.8m

Table 8 — Phase 3: Cost Savings Summary

ModelWith MemoryWithoutSavings
Sonnet 4.6$3.98$7.0443%
Opus 4.6$4.34$7.6543%
Haiku 4.5$4.95 (39/40)$3.97 (0/40)Memory enables capability

When Memory Hurts

This section exists because the first benchmark failed, and a page that only showed the other two would be a worse page.

Phase 1: Toy Codebase Results

nomemory won: 70.3% quality vs 59.5% for stompy.

We're showing this because cherrypicking is dishonest.

There is a complexity threshold below which memory is pure overhead. On a trivial 800-line Express/TypeScript codebase, the model can hold the entire context in its window. Memory retrieval adds latency and noise without reducing exploration — because there is nothing to explore. The breakeven point appears to be around 2,000-3,000 lines of meaningful code with non-obvious architecture.

Use Cases

Multi-Agent Swarms

Problem: 6 agents independently explore codebase. 6x redundant exploration.

How Stompy helps: Lead locks architecture, workers recall. 43% cost savings (Sonnet & Opus). Haiku: 0/40 → 39/40 with memory.

# Lead agent stores architecture
lock_context("service_layer: PostgreSQLAdapter pattern, execute_query for reads, execute_update for writes...")
# Worker agents recall before coding
recall_context("my_app/service_layer")  # deeplink syntax

Ticketing & Project Management

Problem: Re-explain ticket schema every session.

How Stompy helps: Lock conventions once. All ticket tools benefit.

lock_context("ticketing_workflow: states=[open,in_progress,review,done], priorities=[critical,high,medium,low]...")
ticket(action="create", title="Fix auth bug", type="bug")
ticket_board(status="in_progress")

Admin & Operations

Problem: Infra decisions in Slack/heads. 3am incidents.

How Stompy helps: Lock operational knowledge.

lock_context("deployment: DO App Platform, NYC region, auto-deploy on main...")
recall_context("my_app/deployment")
db_query("SELECT * FROM mcp_global.mcp_sessions WHERE status = 'active'")

Cross-Session Development

Problem: Monday morning, explain project for 14th time.

How Stompy helps: Stompy remembers across sessions.

lock_context("api_conventions: REST endpoints at /api/v1/, Pydantic models...")
recall_context("my_app/api_conventions")
context_search("how do we handle auth")

Codebase Onboarding

Problem: New dev asks same questions previous dev already answered.

How Stompy helps: Previous dev's knowledge persists.

project_brief()
recall_batch(topics=["my_app/service_layer", "my_app/database_conventions", "my_app/test_patterns"])

Limitations

N=1 per cell in the swarm phase. That is a promising signal, not statistical proof, and we would rather say so here than have you work it out yourself.

  • Three models (Opus 4.6, Sonnet 4.6, Haiku 4.5), single codebase, N=1 per cell
  • Our own system on our own codebase
  • Pilot study, not statistical proof

What's Next

The honest gap is scale: these are controlled runs, not a fleet. The numbers we would most like to publish are the ones from people who are not us.

Completed

  • Phase 3: Multi-agent swarm — 43% cost savings across Opus & Sonnet, memory enables Haiku (0/40 → 39/40)

Upcoming

  • Multi-model validation (GPT-5-Codex, Gemini 2.5 Pro)
  • TOON serialization format efficiency
  • Longitudinal study: 27 sessions on sustained development
We built a memory system. We tested it honestly. The results were modest. We think modesty, grounded in controlled measurement, is what this field needs.