Claude Opus 5.5’s 1M Context Window Is Not Agent Memory

A large context helps one run, while durable memory requires files, records, retrieval, and explicit handoff.

Buda Team
← Back to Blog
Claude Opus 5.5’s 1M Context Window Is Not Agent Memory

Claude Opus 5.5’s 1M Context Window Is Not Agent Memory

Claude Opus 5.5 has a 1 million token context window and a maximum output of 128,000 tokens. That is enough capacity to inspect large repositories, long conversations, and substantial document sets in one run. It is not permanent Agent memory.

Context is what the model can attend to during a request or an active conversation. Memory is what a system can preserve, retrieve, update, and audit after the window changes, the session ends, or another Agent takes over. Confusing the two creates brittle long-running workflows: they look coherent while the original transcript remains available, then lose decisions when compaction, routing, or a new session changes what the model sees.

A large window delays forgetting; it does not solve persistence

A 1M window lets an Agent carry more evidence at once. It does not decide which facts deserve to survive, whether a source has become stale, which result was approved, or how another Agent should discover the decision tomorrow. Putting every historical artifact into every prompt also raises cost and makes the important constraint harder to find.

Durable memory needs an external structure: files for working material, records for accepted facts, retrieval for relevant context, and an explicit handoff that says what changed and what remains uncertain. The model reads that structure; it is not a substitute for it.

Context capacity and durable Agent memory solve different problems

Preserved thinking is a conversation contract, not a memory store

Opus 5.5 uses adaptive thinking that cannot be disabled. Its thinking blocks are tied to the model and conversation. Anthropic’s migration guide says tool-use integrations should keep conversations append-only and return thinking blocks unmodified with tool results. Editing earlier messages, tools, or the system prompt can invalidate those blocks for newer accounts.

This protects conversation integrity and makes some forms of reasoning extraction harder. It does not turn thinking blocks into reusable organizational knowledge. Most are not suitable records, and compatibility changes when a router moves the conversation to another model. Fable 5.1 and Mythos 5.1 can read Opus 5.5 thinking blocks on the Claude API; other fallback paths may continue without them.

Four kinds of state should stay separate

Working context holds the instructions and evidence needed for the current step. It should be focused and replaceable.

Progress state records what the Agent tried, which tool calls succeeded, and what is blocked. It allows a person or another Agent to resume without replaying the entire transcript.

Durable knowledge contains accepted facts, decisions, and source links. It needs ownership, update rules, and timestamps.

Audit history preserves who changed what and which result was approved. It should not be silently rewritten when the model summarizes a session.

A long task should move from context to records through an explicit handoff

Design long-running work around checkpoints

At the start, give the Agent a narrow objective, source set, permissions, and acceptance criteria. During execution, write progress to artifacts that survive the current response. At meaningful checkpoints, summarize decisions with links to evidence rather than copying every token. Before handoff, validate files and tests, mark unresolved questions, and store the accepted result in the appropriate system of record.

Retrieval should select what the next step needs. A useful memory layer can answer: what was decided, why, by whom, from which evidence, and when it should be reviewed again. A raw transcript usually cannot answer those questions reliably without another expensive interpretation pass.

What this means for Buda

Buda’s role is the workspace around the model: files, browser and terminal sessions, visible tool calls, Skills, and deliverable artifacts. Those surfaces can persist beyond one model call and make a handoff inspectable.

The practical test is not how many tokens fit. Give an Agent a task that lasts longer than one context-management cycle. Interrupt it, resume it, route one step elsewhere, and ask a reviewer to reconstruct the decision. If the work survives because the state is explicit, the system has memory. If it survives only because the transcript is still open, it has context.

A handoff test that can fail honestly

Take an article workflow with source notes, five language drafts and a reviewer decision. End the original session after the first language passes review. Ask a new Agent to continue without pasting the entire chat. It should find the approved version, the unverified claims, the remaining languages and the review decision from explicit files or records. If it cannot, the workflow has not preserved memory, regardless of the size of the context window.

A useful checkpoint says what changed, where to inspect it, which sources support it, which checks passed and which decisions still belong to a person. Date decisions that can expire. Do not overwrite the review trail when updating a summary; preserve the previous decision and the reason for its replacement. That is what lets a third person reconstruct a handoff without guessing what the model meant by “done.”

Using Opus 5.5 in Buda

Claude Opus 5.5 is now available in Buda. Select it in the Agent’s model picker while retaining the same workspace and files. The larger model context does not turn conversation history into permanent memory; continue saving decisions, evidence, and handoff checkpoints in files or trusted records.

Read the Buda launch announcement · Buda Credits

Sources