Context Engineering for AI Agents: The Production Guide

By Sachin Shinde · October 2026 · 18 min read
Context engineering for AI agents: layers from instructions through retrieval, memory, tools, and a production control loop

Context engineering is the discipline of designing what an AI agent should know, retrieve, remember, and be permitted to do at each inference step. Prompt engineering is one layer within it. The unit of design is not the instruction—it is the complete context state the model operates on at each decision point: instructions, retrieved evidence, memory, tool definitions, and execution state.

A 2025 survey of over 1,400 research papers organized the field into three technical pillars: context retrieval and generation, context processing, and context management—all of which appear across RAG systems, memory architectures, tool-integrated reasoning, and multi-agent pipelines (Mei et al., arXiv:2507.13334, July 2025). Anthropic's engineering team describes the production objective as curating and maintaining the optimal information set at inference time, where the system prompt, external data, tools, MCP servers, and message history together form the context state—not the instruction alone.

This guide gives engineering leaders a production decision framework that maps each context layer to its failure mode, control owner, and evaluation signal, including a decision rule for when to preload context, retrieve information just in time, or use a hybrid approach.

Questions This Article Answers

What Is Context Engineering for AI Agents?

Context engineering for AI agents is the practice of curating and maintaining the optimal information set for a model at each inference step. It addresses what the agent should know at the current step, what it should retrieve from external sources, what it should remember from prior steps, what tools it can invoke, and what actions it is authorized to take. The unit of design is not the prompt; it is the full context state at each decision point in the agent loop.

Anthropic's engineering team defines context engineering as maintaining the optimal information at inference time and identifies the system prompt, external data, tools, Model Context Protocol servers, and message history as components of the context state—not supplementary elements around a prompt. A survey of over 1,400 research papers organized the field into context retrieval and generation, context processing, and context management. These three layers appear across production AI systems that go beyond single-turn chat: RAG systems, agentic architectures, multi-step reasoning pipelines, and multi-agent coordination.

No single taxonomy is yet standard. The framework in this article assigns each context layer a failure mode, a control owner, and an evaluation signal—a structure designed for engineering leaders who need to assign responsibility and measure results, not only understand the concept.

How Does Context Engineering Differ from Prompt Engineering?

Prompt engineering designs the wording and structure of a single instruction to guide model output. Context engineering designs the entire information environment the model operates in at each step of an agent loop. Prompt engineering is a craft; context engineering is a system. The two are not in competition—a well-engineered prompt is part of a well-engineered context.

DimensionPrompt EngineeringContext Engineering
Unit of designA single instruction or messageThe entire context window at each inference step
ScopeWording and structure of the requestInstructions, retrieved evidence, memory, tools, execution state
AssemblyStatic or templatedDynamically assembled per request and per step
Primary skillPrompt design and iterationRetrieval, memory, tool, permission, and evaluation design
Failure modeThe model misunderstands or ignores the instructionWrong evidence, drifted memory, excessive tool noise, or unauthorized content

The distinction has operational consequences. A better-worded prompt cannot compensate for retrieved evidence that is stale, a memory layer that has drifted from the source of truth, or a tool set that grants more authority than the task requires. Each context layer has its own failure mode and its own remediation path.

What Are the Layers Inside an Agent's Context?

A production agent's context has six active layers and a governing control plane. Each layer contributes a different type of information and carries a different failure risk. The table below maps each layer to its typical contents, its main failure mode, and the control that addresses it.

Minimum sufficient context production loop: scope the decision, retrieve evidence, filter and label, assemble context, authorize action, then evaluate and update
LayerTypical ContentsMain FailureControl to DesignEvaluation Signal
InstructionsRole, goal, constraints, output contractAmbiguity or brittle prompt logicVersioned instructions, hierarchy, regression testsPass rate on instruction-adherence eval set
Working requestCurrent user intent, task identifiers, unresolved fieldsWrong scope or stale assumptionsNormalize the request; record unresolved fields explicitlyTask-completion rate on representative inputs
Retrieved evidenceDocuments, records, search results, code chunksIrrelevant, stale, or decontextualized chunksHybrid retrieval, metadata filters, reranking, provenance fieldsRetrieval failure rate; downstream answer grounding rate
MemoryPreferences, prior decisions, durable workflow lessonsSummary drift, poisoning, privacy leakageTyped memory, write policy, expiry rules, audit trailMemory accuracy on recalled-fact test set
Tools and resourcesAPIs, files, MCP resources, executable functionsToo many tools, ambiguous choice, excessive authorityMinimal tool surface, explicit schemas, permission boundariesTool-selection accuracy; unauthorized-action rate
Execution stateTool results, intermediate artifacts, checkpointsContext overflow or hidden stateStructured state records, compaction threshold, reviewed artifactsState-consistency rate across multi-step tasks
Control planeEvals, approvals, monitoring, rollback capabilityGood-looking output paired with unsafe or incorrect actionRisk gates, human ownership of consequential actions, traces, regression setIncident rate; time-to-detect on regression

The layers do not require separate products. They require separate responsibilities. A single platform can host all of them, but the team should be able to identify which layer produced a failure, who owns the control, and what evaluation signal confirms the fix.

How Should Retrieval and RAG Supply an Agent's Context?

Retrieval is the mechanism that brings external knowledge into an agent's working context at request time. The quality of retrieved evidence affects whether the model can reason correctly—not the model's parameters alone. Retrieval engineering therefore belongs within context engineering, not alongside it.

The evidence unit matters before the retrieval algorithm. A chunk that loses its document heading, owner, source date, and permission attributes during splitting is harder to rerank correctly and harder to cite accurately. Anthropic's contextual retrieval approach prepends a brief document-level context to each chunk before embedding it.

In Anthropic's reported experiment across selected domains and embedding configurations, contextual embeddings reduced top-20 retrieval failure from 5.7% to 3.7%—a 35% improvement over the baseline. Combining contextual embeddings with contextual BM25 reduced failure to 2.9%, a 49% improvement. These results are vendor-reported and configuration-dependent; they support testing contextual retrieval in your own setup rather than treating the results as a universal guarantee.

Hybrid retrieval combines dense semantic search with keyword search. Dense retrieval finds conceptually similar content; BM25 keyword retrieval finds exact product names, codes, policy identifiers, and proprietary terms that semantic search can miss. A reranker trained on task-relevant labeled pairs can improve precision over the merged candidate set before the context is assembled.

Query ShapeUseful Retrieval PathWhat It Catches
Natural-language conceptDense semantic searchParaphrases, related ideas, implicit references
Exact identifier, code, or nameBM25 keyword searchProduct codes, policy numbers, proprietary terms
Mixed enterprise contentHybrid: both paths merged and rerankedBoth concept and identifier types in one pass
Recent or time-sensitive factFiltered hybrid with freshness metadataCurrent records; excludes superseded content

When Should an Agent Preload Context, Retrieve Just in Time, or Use a Hybrid?

The preload-versus-just-in-time decision depends on how stable the information is, how large the candidate set is, and whether the agent can know at session start which records it will need. Loading everything upfront is not a strategy—it consumes tokens budget, increases latency and can reduce reasoning quality when irrelevant information is present.

Preload stable context that changes infrequently and applies to every call in the session: role definitions, operating constraints, output format contracts, and system policies. This avoids retrieval latency for information the agent will always need.

Retrieve just in time when the agent operates over a large corpus, needs user-specific records, or cannot predict at startup which documents are relevant. Anthropic describes the pattern as holding lightweight identifiers (such as a conversation ID or a document key) and fetching actual content through a tool call when the agent determines it is needed, rather than loading every candidate into the context window at session start (Anthropic, September 2025).

Context TypePreferred StrategyTrigger for the Choice
Role instructions, constraints, output formatPreload at session startStable; applies to every call; small token footprint
Large document corpus or database recordsJust-in-time retrieval via tool callCannot predict at the start which records are needed
User-specific preferences or historyJust-in-time from a memory storeUser identity is known; the corpus is per-user and large
Structured task with stable role and dynamic evidenceHybrid: preload role, retrieve evidence per stepInstructions are fixed; evidence changes with each subtask
Time-sensitive or frequently updated contentJust-in-time with a freshness filterThe preloaded version would be stale before the session ends

Use a freshness boundary to invalidate just-in-time evidence that was retrieved earlier in the same session but may have changed during a long-running task. The agent should not assume that a record fetched at step one is still current at step twelve.

How Should Teams Manage Memory and Conversation History?

Memory management is the part of context engineering that decides what persists between model calls and for how long. Unmanaged conversation history grows the context window at each step, increases latency and cost, and eventually can exceed the model's context limit. Three main policies address this problem.

PolicyMechanismAdvantageRisk
TrimmingRemove older turns from the context windowDeterministic, inexpensive, and introduces no new errorPermanently discards information the model cannot recover
SummarizationReplace a history block with a compressed versionPreserves longer-range continuityIntroduces information loss, summary bias, latency, and summary-poisoning risk
CompactionServer-managed prune pass at a defined token thresholdTransparent to the application; preserves active stateThe compacted state is opaque; treat source artifacts as the truth

OpenAI's SDK cookbook recommends logging summaries and treating the reviewed source artifact as the authoritative record, rather than relying on the model's compressed memory of earlier turns. OpenAI's API guide documents server-managed compaction: when the rendered token count crosses a defined threshold, the service runs a compaction pass, emits a compaction item, and prunes context before continuing inference (accessed October 2026).

Long-term memory (preferences, prior decisions, durable workflow lessons) requires a separate typed schema, a write policy, expiry rules, and an audit trail. A preference record and a factual record should not share the same memory store without a type field, because they have different expiry, confidence, and update policies.

How Should Tools and MCP Resources Be Exposed to an Agent?

The Model Context Protocol separates the context surface into three primitives with different control owners: prompts are user-controlled, resources are application-controlled contextual data, and tools are model-controlled executable functions (current MCP specification, accessed October 2026). This distinction matters because information access and action authority have different risk profiles. A model reading a resource and a model calling a tool are not equivalent operations.

MCP PrimitiveControl OwnerRisk TypeDesign Rule
PromptsUserInstruction injection from user-supplied templatesValidate and scope user-supplied prompt content before it reaches the model
ResourcesApplicationStale, unauthorized, or poisoned contextual dataApply access control, freshness checks, and provenance labeling before exposure
ToolsModel (invocation-gated by application)Incorrect tool selection, excessive authority, consequential side effectsMinimal tool surface; explicit schemas; confirmation gates for consequential actions

Anthropic's engineering guidance warns that bloated tool sets create ambiguous decision points: when many tools overlap in scope, the model must choose between plausible options, and incorrect tool selection can compound into multi-step errors (Anthropic, September 2025). The design principle is the smallest tool set that can complete the next step, with clear parameter schemas and permission boundaries enforced by the application layer rather than by the model alone.

Tool calls should be logged with their source context so the team can trace which retrieved evidence or memory record led the model to invoke a given tool. This trace is a primary input for debugging agentic failures and for building the injection-resistance test suite described in the next section.

How Do Permissions and Prompt Injection Affect Context Design?

Permissions and prompt injection are not security concerns added after context engineering is complete. They are constraints that context architecture must satisfy from the start, because the context window is an attack surface for injection and the retrieval layer determines which content reaches it.

Prompt injection occurs when untrusted external content in an agent's context carries instructions that override the agent's intended behavior. The risk increases with the agent's action authority, a read-only agent that receives an injection may return a misleading answer; an agent with write access to a database or messaging system could potentially be manipulated into consequential operations.

OpenAI's engineering guidance on prompt injection frames the problem as source-and-sink control. The application should constrain what sinks a successfully injected instruction can reach, rather than relying solely on detecting or blocking every injection attempt.

Practical context-engineering controls include:

  • Labelling retrieved external content as untrusted and keeping it structurally separate from trusted system instructions
  • Restricting what actions an agent can take without a human confirmation step
  • Logging every tool call with its source context, so an injection's path can be traced after the fact
  • Including injection-resistance cases in the pre-release evaluation set

Context architecture decisions (how sources are labelled, which tools are available, and what confirmations are required) determine the potential impact of a successful injection. These decisions belong in context design, not only in a post-hoc security review. See also: AI agent security.

Security boundary for AI agents: untrusted content is source-labelled, checked by sink control and human approval, then either denied or allowed to reach an authorized action with trace and revoke

Why Can More Context Make an AI Agent Less Accurate?

Adding more information to an agent's context does not necessarily improve its answers. Position, relevance, and volume interact in ways that can reduce reasoning quality, which is why context engineering is a selection problem, not a maximization problem.

A 2023 study by Liu et al. examined how language models use information in long contexts across multi-document question-answering tasks. The researchers found a U-shaped performance curve: models used information best when it appeared at the beginning or end of the context, and performance degraded when the relevant passage was in the middle of a long sequence.

Increasing retrieved documents from 20 to 50 produced only marginal accuracy gains in the reported experiments with GPT-3.5-Turbo and Claude 1.3 (Liu et al., “Lost in the Middle,” arXiv:2307.03172, 2023).

Three production implications follow:

  • Position Matters Place the most relevant evidence early or at the end of the assembled context. Evidence buried in the middle of a long retrieved block may be less likely to be used correctly.
  • Volume has Diminishing Returns A small set of high-relevance chunks can outperform a large set of mixed-relevance results. Increasing retrieval depth without improving precision adds noise rather than useful signal.
  • A Larger Context Window is not a Substitute for Selection A bigger token budget raises the ceiling but does not remove the need to select, order, and label evidence deliberately before assembly.

How Should Teams Evaluate Context and Retrieval Quality?

Standard retrieval metrics (recall at k, mean reciprocal rank, normalized discounted cumulative gain) measure whether a relevant document appears in the retrieved set. They do not measure whether that document actually improved the model's answer on the downstream task.

A 2024 study tested this gap directly. The researchers found that conventional query-document relevance labels correlate only weakly with downstream RAG performance.

Their proposed eRAG method evaluates each retrieved document by passing it through the downstream generation task and using the resulting output to estimate document-level usefulness rather than label-level relevance (Salemi and Zamani, “Evaluating Retrieval Quality in Retrieval-Augmented Generation,” arXiv:2404.13781, April 2024).

Retrieval quality flow from query and candidate evidence through reranking and filtering into a context packet and downstream answer, showing that relevance is not the same as usefulness

A practical context quality scorecard covers six dimensions:

DimensionDefinitionEvaluation Method
RelevanceDoes each context chunk support the model's specific task at this step?Labeled retrieval test set; downstream answer grounding check
SufficiencyDoes the assembled context contain enough information to complete the task without guessing?Task-completion rate on representative failure cases
FreshnessDoes the context reflect the current state of the source, or has it been superseded?Source timestamp comparison; stale-document test set
ProvenanceCan each claim in the output be traced to a specific source, owner, and retrieval timestamp?Citation accuracy rate; source-tracing audit
ControllabilityCan the team update, expire, or remove any context layer without rebuilding the whole system?Layer-change test: modify one layer, verify no unintended side effects
CostDoes the assembled context fit within token and latency budgets for the use case?Token count per call; p95 latency under production load

Run this scorecard per layer, not only at the output level. A well-designed instruction layer can be undermined by a retrieval layer that returns irrelevant chunks. A well-tuned retrieval layer can be undermined by a memory layer that has drifted. Each layer requires its own test set and its own owner. See also: AI agent evaluation.

What Does a Production Context-Engineering Control Loop Look Like?

A production context-engineering system is a loop, not a pipeline. After each model call, the system validates the output, updates state, compacts stale context, and carries forward only the durable artifact needed for the next step. The loop includes the evaluation layer from the start, not as an afterthought.

The nine steps in order:

Scope the task and identify the decision the next model call must make.

Load stable instructions and only the policy relevant to that decision.

Retrieve candidate evidence using exact-match and semantic signals.

Filter by permissions, freshness, source authority, and task relevance.

Rerank, compress, order, and label evidence so provenance survives into generation.

Expose the smallest tool set that can complete the next step.

Run the model, validate the output against a grounding check, and gate any consequential action.

Persist only the durable state or artifact needed for the next step.

Compact or invalidate stale context, then measure the run against an evaluation set.

Realisier's internal AI platform, Pulse, runs this loop in production across RAG-based question answering, multi-provider model routing, and consent-based adoption analytics. The governing design principle is minimum sufficient context: the smallest set of current, attributable, and governable information needed for the next decision. See also AI agent observability and LLMOps.

When Is a Simple Prompt Enough?

Context engineering adds complexity. The overhead is justified only when the task requires it. Adding retrieval, memory, or tool layers to a task that does not need them raises cost, increases the failure surface, and slows iteration.

A single, well-engineered prompt is sufficient when the task is single-turn and the required information fits entirely within the prompt; when the model's parametric knowledge covers the domain and recency is not a concern; when there is no external knowledge corpus, no conversation history to maintain, and no tool execution required; and when the failure cost is low enough that no additional grounding or governance layer is required.

Use case examples: Summarizing a document supplied in full within the prompt, classifying content against a fixed vocabulary, and generating code from a complete specification included in the prompt.

Add a context layer when a simple prompt fails on a recurring, identifiable pattern: hallucinated facts the model could not have known, outdated information, incorrect decisions from missing history, or tool execution the model cannot simulate.

Context engineering addresses specific information-supply problems. Build the smallest system that passes its evaluation set, then add layers only when a demonstrated failure pattern requires them. See also: human-in-the-loop AI for governance patterns when context complexity grows.

What Does This Mean in Practice?

The useful production question is not how large the context window can be. It is what context the agent actually needs for the next decision, where that context comes from, and what controls govern its selection, expiry, and evaluation.

  • Start with instructions. Establish a clear, versioned, tested role definition before adding any other layer. Ambiguous instructions can compound at retrieval.
  • Add retrieval when the model's knowledge is insufficient for the task or when the information must be current and citable. Choose the retrieval path by query shape, not convention.
  • Add memory when the application requires continuity across turns that the model cannot hold in short-term context. Use typed memory with explicit write policies.
  • Add tools when the agent must take actions that retrieval alone cannot provide, and design the tool surface to match only the next step.
  • Evaluate each layer separately. A well-designed prompt can fail because retrieval is wrong. Good retrieval can be undermined by a memory layer that has drifted. Treat each layer as an independent engineering responsibility with its own test set and owner.

Building an AI Agent in Production?

Realisier Labs designs and builds production AI systems for US companies—including retrieval architecture, memory policy, tool boundaries, and the evaluation layer. Talk to Sachin about your context-engineering challenge.

Frequently Asked Questions

What Is a Context Engineer?

A context engineer designs the information environment an AI agent uses at each inference step. The role covers retrieval architecture, memory policy, tool selection, permission boundaries, and context evaluation—disciplines that span data engineering, MLOps, and product engineering. The title is recent; the responsibilities exist in teams that build production AI agents, regardless of what the role is called.

Is Context Engineering the Same as Prompt Engineering?

No. Prompt engineering designs the wording of a single instruction inside the context window. Context engineering designs the entire context window: what evidence the agent retrieves, what it remembers, what tools it can call, and what it is authorized to do. A prompt is one input to context engineering, not the whole discipline. A better prompt cannot fix a retrieval layer that returns stale or unauthorized content.

What Are the Main Components of Context Engineering?

Production context engineering has six active layers: role instructions, working request and task state, retrieved evidence, memory (short-term and long-term), tools and resources, and an execution-state record. A governing control plane (covering evaluation, approvals, monitoring, and rollback) sits across all six. Each layer has its own failure mode, control owner, and evaluation signal.

How Does Token Budget Management Fit into Context Engineering?

Token budget management is the cost dimension of context engineering. The design goal is not to fill the context window—it is to assemble the minimum sufficient, current, attributable, and governable set of tokens for the next decision. Excess tokens raise inference cost, increase latency, and can reduce accuracy when irrelevant content displaces useful evidence in the model's attention. Budget decisions belong in context selection, not only in post-hoc cost monitoring.

How Does Model Context Protocol Fit into Context Engineering?

MCP structures the context surface into three primitives: prompts (user-controlled), resources (application-controlled contextual data), and tools (model-controlled executable functions). This separation makes control boundaries explicit: reading information through a resource and executing an action through a tool carry different confirmation requirements, audit obligations, and potential impact. Context engineering should design the permission and logging model for each primitive type independently.

How Should a Team Start with Context Engineering?

Start with the instruction layer. Define the agent's role, operating constraints, and output contract clearly, version the definition, and write a small labeled evaluation set that tests adherence to those constraints. Add retrieval when the first retrieval-specific failures appear in evaluation—hallucinated facts, outdated answers, or scope errors the instructions alone cannot prevent. Add memory and tools only when the task requires continuity or external actions. Build the evaluation layer in parallel with every addition, not after deployment.