You’ve seen it. An AI agent starts strong, pulls the right data, answers cleanly. Then ten turns later it’s hallucinating yesterday’s stock price, forgetting the user’s name, and recommending a product that was discontinued last week. The root cause isn’t the model. It’s the infrastructure. Or more precisely, the lack of it.
Most teams treat context as a RAG problem. Stuff a vector database, retrieve chunks, call it done. That works for a single Q&A. For an agent that plans, executes, loops, and adapts over hours or days? It collapses. The agent doesn’t fail because the retriever is bad. It fails because there is no infrastructure-level context loop — no system that continuously updates, propagates, and reconciles context across the stack. That’s what we need to build.
The RAG Fallacy: Why Static Retrieval Fails Dynamic Agents
RAG is a snapshot. You embed a corpus, index it, and at inference time you retrieve the top-K chunks that match the query. For a chatbot answering “What is the refund policy?” that’s fine. For an agent that books a flight, monitors the price, cancels the old booking, and rebooks when the price drops? RAG can’t hold the thread.
The problem is threefold.
First, agents generate new context with every action. Every API call, every tool output, every decision creates state that the next step depends on. RAG doesn’t ingest that. It points at a static corpus. If your agent calls a weather API, the result lives in the LLM’s context window — until it rolls off. The next loop has no memory of that temperature reading unless you explicitly write it somewhere and retrieve it again. That’s not a loop. That’s a batch job.
Second, agents need to reconcile conflicting context. A user says “I want the red one.” The inventory system says “red is out of stock.” The agent needs to hold both facts, resolve the conflict, and inform the user — without forgetting the original preference. RAG can’t do that. It retrieves vectors. It doesn’t merge state.
Third, agents degrade over time because the context window is a finite resource. Every turn appends tokens. Eventually you hit the limit, truncate, and lose the early decisions that shaped the current goal. The agent doesn’t know why it started a workflow. It just keeps executing.
My take: RAG is a brilliant solution for information retrieval. It is a terrible solution for agent state management. Teams that try to force RAG into an agent loop end up with brittle systems that work in demos and fall apart in production after 20 minutes. You need a different primitive.
What Is a Context Loop? A New Infrastructure Primitive
A context loop is not a retrieval pattern. It’s an infrastructure primitive — a system that continuously ingests, merges, and propagates context across the agent stack in real time. Think of it as a live state stream that every component reads from and writes to.
Here’s what it looks like in practice.
The agent runtime emits context events: tool call results, user messages, intermediate reasoning, environment changes. Those events flow into a context store — not a vector database, but a time-ordered, key-value store with conflict resolution. The context store merges incoming data with existing state, resolves contradictions using priority rules (user intent > system data > model inference), and publishes the updated context back to the runtime.
The runtime doesn’t query for context. It subscribes. Every loop starts with the latest merged context, not a fresh retrieval. The context store acts as a persistent, evolving memory that outlives any single LLM call.
Key components:
- Context store: A durable, ordered log of state changes. Not a vector index. Supports merge semantics (last-write-wins, CRDTs, or custom conflict resolvers).
- Subscription layer: The runtime subscribes to context updates, not polls. Reduces latency and prevents stale reads.
- Propagation hooks: Every tool, every model call, every user interaction writes back to the context store. Nothing is ephemeral.
This is different from “memory” in the LLM sense. Memory is a prompt hack. Context loops are infrastructure. They sit between the agent and the world, ensuring that every step sees the full, reconciled picture.
How Context Loops Solve Model Churn and Agent Degradation
Model churn is the silent killer of production agents. A new model release comes out — say, GPT-5.2 or Claude 4.5 Opus — and suddenly your agent behaves differently. The old prompt engineering tricks break. The chain-of-thought format shifts. The agent starts ignoring context it used to respect.
Why? Because you hardcoded context management into the prompt. “Remember the following facts: …” That works until the model’s attention mechanism changes. Every model release is a regression risk for your agent’s memory.
Context loops decouple memory from the model. The context store holds the state. The model receives a compressed, structured summary of that state — not raw history. When the model changes, you update the compression logic, not the entire agent. The context loop stays the same.
Agent degradation is the second problem. Over long sessions, agents lose track of the original goal. They start making decisions based on recent context while ignoring the initial plan. This is the “drift” problem. Context loops solve it by maintaining a goal tree — a structured representation of the original objective, sub-tasks, and completion status. Every loop checks the goal tree before acting. If the agent starts drifting, the context store provides the correction.
A real example from our work: A customer support agent that handles refunds. The user says “I want a refund for order 12345.” The agent checks the order, finds it’s past the refund window, and offers store credit. The user agrees. Ten minutes later, the same agent is handling a different issue for the same user and tries to process a refund again — because the earlier decision fell out of context. With a context loop, the “refund denied, credit offered, accepted” state persists. The agent never re-enters that path.
Building Context Loops: Key Architectural Decisions
You can’t buy a context loop off the shelf — not yet. You have to build it. Here are the decisions that matter.
1. Choose a context store with merge semantics, not just append.
Most teams start with Redis or Postgres. Those work for simple key-value storage. But they don’t handle concurrent writes from multiple agent threads. If two sub-agents update the same user’s context simultaneously, you get a race condition. Use a store that supports CRDTs (conflict-free replicated data types) or a custom merge function. Crescendo’s latest AI news coverage highlighted a startup using FoundationDB for exactly this reason — strong consistency with merge logic.
2. Context compression is not summarization.
Don’t ask the LLM to summarize the context store every loop. That’s expensive and introduces hallucination. Instead, use a deterministic compression layer: extract key-value pairs, timestamps, and status flags. Send only the delta to the model. The full history lives in the store. The model gets a structured snapshot.
3. Make the loop model-agnostic from day one.
You will switch models. Guaranteed. The LLM leaderboard on BenchLM.ai shows 290 models as of July 2026. Some will be better for reasoning, some for speed. Design your context loop to output a generic context object — JSON, protobuf, whatever — that any model can consume. Don’t tie it to OpenAI’s message format or Anthropic’s system prompt syntax.
4. Instrument everything.
You can’t debug a context loop without observability. Log every context write, every merge conflict, every subscription event. If an agent makes a bad decision, you need to trace it back to the exact context state it saw. This is non-negotiable.
What about cost? A context store is cheaper than re-embedding every turn. You’re storing JSON, not vectors. The real cost is the subscription layer and the merge logic. Budget for that.
What about latency? The subscription model is faster than RAG because you’re not doing a vector search every loop. You’re reading from a local cache or a low-latency store. The tradeoff is complexity. You need a streaming infrastructure (Kafka, Redis Streams, or similar). Worth it.
Real-World Example: Nvidia’s Agentic AI Stack and Context Loops
Nvidia gets this. In July 2026, they announced an agentic AI infrastructure stack that explicitly includes a “context orchestration layer.” Their announcement on CIO.com describes a system where context is not a prompt artifact but a first-class infrastructure component.
The Nvidia stack uses NIM microservices for model inference, but the key innovation is the context bus — a real-time event stream that connects agents, tools, and memory. Every NIM service publishes context changes. The bus merges and redistributes. Agents subscribe to the bus, not to individual services.
This is exactly the pattern I described. Nvidia recognized that agent failures aren’t model failures — they’re context failures. By building a context bus, they allow agents to maintain coherent state across heterogeneous models and tools. You can swap a Llama-4 agent for a Gemini agent without rewriting the context logic.
The lesson: The infrastructure giants are moving toward context loops. If you’re building agents today, you should too. Don’t wait for a managed service. The primitives are available — event streams, merge stores, subscription layers. Assemble them now.
What’s the biggest mistake teams make when building context loops? They try to make the LLM manage its own context. They put the merge logic in the prompt. That’s fragile. The context loop should be deterministic, not probabilistic. Let the LLM reason. Let the infrastructure remember.
How does a context loop handle privacy and data retention? Same way any stateful system does. You set TTLs on context entries. You purge after session end. You encrypt at rest and in transit. The advantage is that you control the data, not the LLM provider. Your context never enters a model training set.
My Take
I’ve built three agent systems. The first used RAG for everything. It worked for two weeks, then broke under real traffic. The second used a hand-rolled memory buffer in Redis. Better, but race conditions killed it. The third used a context loop with CRDTs and a subscription layer. That one worked.
Here’s what I believe: The next wave of AI infrastructure will be about state management, not model performance. Model benchmarks are flattening. The Zapier LLM comparison shows diminishing returns between top models. The real differentiation will be how well your agent holds context. The company that builds the best context loop will win the agent platform war.
Most teams are optimizing the wrong thing. They chase lower latency, higher accuracy on eval sets, fancier prompts. Meanwhile their agents forget the user’s name after three turns. That’s not a model problem. That’s an infrastructure problem. Fix the infrastructure.
Don’t build another RAG pipeline. Build a context loop.
Key Takeaways
- RAG is static retrieval. Agents need dynamic, continuous context updates that persist across turns and tools.
- A context loop is an infrastructure primitive — a live state stream with merge semantics, subscription, and propagation.
- Context loops decouple memory from the model, solving model churn and long-session degradation.
- Build with a merge-capable store (CRDTs or similar), deterministic compression, and model-agnostic output.
- Nvidia’s agentic stack validates the pattern: context as a bus, not a prompt hack.
- The biggest mistake is putting merge logic in the prompt. Keep the loop deterministic.
- Instrument everything. You can’t debug what you don’t log.