Durable State Should Outlive the Agent That Produced It
I reopened a repository after the agent session that had been working on it was gone. The terminal history was incomplete, the conversational context was no longer available, and the model that started the work did not matter anymore. What mattered was whether the repository could answer a few basic questions: What was the goal? What was not allowed to change? Which steps had been completed? What evidence supported those claims? Where should the next operator resume?
A chat transcript can sometimes answer those questions, but it is a poor place to keep operational truth. It mixes decisions with speculation, commands with explanations, and current facts with statements that became obsolete ten messages later. It is optimized for conversation, not for resumption.
That experience reinforced a design rule I now consider fundamental: durable state should outlive the agent that produced it.
The session is not the system
Agent tools naturally make the session feel like the center of the work. The session contains the prompt, tool calls, intermediate reasoning, command output, and final response. While the session is active, that can look like a complete operating environment. The illusion breaks as soon as the process stops, the context window is compacted, a credential expires, a different model takes over, or a human needs to inspect the work without replaying the conversation.
The underlying project has a longer life than any of those components. Its intent may last for months. Its constraints may reflect product or safety decisions that must remain stable across many runs. Its evidence may need to be checked by CI, another agent, or a human reviewer. Its handoff should remain readable even if the original runtime no longer exists.
This suggests a clean separation. The agent session is an execution context. Project state is a durable contract. The first can be replaced. The second must remain inspectable.
I have been applying that separation through SMALL Protocol. SMALL defines five machine-readable artifacts: human-owned intent and constraints, agent-owned plan and progress, and a system-owned handoff. The artifacts live with the project as plain files rather than inside a model provider's conversation store. They are schema-validated, versioned, and designed to make the current state resumable.
The important idea is not the filename convention. It is the ownership boundary. Humans state what the work is and what must not change. Agents propose and record execution. The system produces the resume point. Those responsibilities should not collapse into one long narrative written by whichever model happened to be active.
State and execution are different jobs
Once state is explicit, a second distinction becomes visible: describing the work is not the same as driving it.
A state layer can say that a task is pending, that a constraint protects a path, that progress evidence exists, or that a handoff points to the next action. It should not also have to schedule every command, enforce every iteration limit, or decide when a test has converged. Those are execution responsibilities.
This is why I keep SMALL separate from loopexec. SMALL is a governance and continuity layer. It is explicitly not an agent framework or workflow engine. loopexec is a deterministic runtime for bounded work loops. It runs an execution command, evaluates an external check, records what happened, and stops for a computed reason. One describes durable truth about the project. The other drives a bounded attempt to change that truth.
A simplified loop can look like this:
loopexec run --json \
--exec "make fix" \
--check "go test ./..." \
--max-iterations 10
The runtime is allowed to execute and observe. It is not allowed to redefine the project's intent because a loop became inconvenient. It cannot silently remove a human constraint to obtain a green result. A passing check does not erase the requirement to preserve evidence and produce a valid handoff.
That separation creates useful failure boundaries. If the plan is invalid, the state layer should reject it. If the external check is flaky, the execution runtime should refuse to treat it as a stable oracle. If two agents attempt to write state concurrently, orchestration should fail rather than invent an automatic merge. If a loop reaches its iteration cap, the project state should still be available for the next operator.
Without these boundaries, every failure turns into the same vague diagnosis: the agent did something wrong. With them, the system can identify whether the problem belongs to intent, state integrity, execution, verification, concurrency, or resumption.
Plain files are an architectural choice
Keeping state in plain files can look unsophisticated next to a large orchestration database. I think that simplicity is an advantage when the goal is durable operational truth.
A file can be read by a human, validated by a CLI, diffed in Git, inspected by CI, and handed to a different runtime. It does not require the original session service to reconstruct meaning. It can travel with a branch and participate in the same review process as the code it describes.
This does not mean arbitrary text is enough. Durable state becomes useful when it has a schema and invariants. SMALL requires completed tasks to have corresponding evidence. Progress is append-only. Handoffs contain a replay identifier. Strict validation checks that task references are known and that the canonical state layout has not accumulated unrelated files. The protocol is single-writer by design because silent conflict resolution would make state less trustworthy, not more.
These constraints make the files less convenient to improvise and more valuable to resume. That is the right trade. The purpose of operational state is not to preserve every thought an agent had. It is to preserve the minimum set of facts another operator needs to continue safely.
Resumption is the real portability test
Teams often describe an agent system as model-independent because it can send prompts to more than one provider. That is only transport portability. A stronger test is whether the work can survive a change of operator.
Can a different model enter the repository, read the current intent, see the constraints, identify pending work, inspect evidence, and continue without reconstructing the project from chat history? Can a human audit why a task was marked complete? Can CI validate the state without calling the model? Can the original runtime disappear without taking the project's memory with it?
If the answer is no, the system still belongs to the session. The provider may be swappable, but the operating state is not portable.
A durable handoff changes that relationship. It turns resumption into an explicit system behavior rather than a hopeful prompt. The next operator does not need to imitate the previous operator's reasoning style. It needs to honor the same intent, constraints, evidence rules, and verification boundary.
This also makes model upgrades less dramatic. A new model can be evaluated as an executor against the same state contract. A local model can handle a bounded step, a hosted model can take a more demanding one, and a human can intervene when judgment is required. The project does not have to migrate its truth every time the execution layer changes.
Preserve verdicts, not mythology
There is a related lesson in execution receipts. A live language-model trajectory is not reproducible in the way a deterministic build is. Even with the same prompt and nominal settings, the exact sequence of choices may differ. Pretending otherwise turns audit language into mythology.
What can be preserved is the verdict boundary: the state that was tested, the check that ran, the exit result, the model identity and context used, the stop reason, and the receipt connecting them. loopexec makes an explicit distinction between replaying a verdict and re-executing a live agent loop. Replay verifies the recorded outcome without calling the agent. Re-execution runs the non-deterministic process again and can only report a statistical relationship to the earlier run.
That distinction follows the same architecture. The durable layer stores what can be verified. The execution layer remains replaceable and honestly non-deterministic.
Design for the operator who comes next
The most useful question for an agent workflow is not whether the current model appears productive. It is whether the next operator can understand and continue the work after the current model is gone.
That next operator may be another model, a scheduled runtime, a teammate, or me several weeks later. It should not need privileged access to an old conversation or faith in a polished completion summary. It should find explicit intent, stable constraints, evidence-backed progress, a valid resume point, and an external way to verify the result.
Agents will keep changing. Model routes, context limits, tool APIs, and runtimes will change with them. Durable work needs a slower-moving layer underneath.
The strongest agent systems will not be the ones that preserve the longest conversations. They will be the ones whose state remains clear after the conversation is gone.
