Agent Harness vs. Framework vs. Runtime vs. Control Plane
When Claude Code and Codex share work, which layer stops a risky command, which resumes a broken session, and which still knows who owns the task?
A Claude Code session tries to force-push a branch. A Codex run reaches for a file outside its workspace. A session ends halfway through a task whose next step belongs to someone else. When something goes wrong, which component can stop it, which can recover it, and which still knows whose job it was?
Four kinds of software claim a piece of that answer:
- An agent framework gives you the building blocks for agent logic: model and tool abstractions, agent loops, and workflow primitives.
- An agent harness runs that loop as a working agent. It drives model and tool calls, manages context and state, applies approval policies, and keeps a multistep task moving.
- A coding runtime such as Claude Code or Codex is a harness packaged for software work, with its own tools, permission system, sandbox, and session store.
- A multi-agent control plane sits above the runtimes and governs the work itself: which tasks exist, who owns each one, what it depends on, what it may spend, and who must review it.
These are responsibilities, not exclusive product categories. LangChain’s documentation calls LangGraph an orchestration runtime, and many teams also use it as a framework. AWS runs its managed harness inside AgentCore Runtime. Claude Code is a harness, a runtime, and a product at once. This guide maps the responsibilities, then walks one handoff through all of them.
The four layers at a glance
In short: harnesses and runtimes stop tool calls, and each runtime resumes its own sessions. In this setup, Paperclip retains the shared task-owner record after a session ends. Its experimental tool gateway, off by default, can also gate calls to tools connected through Paperclip, but not the runtime’s own shell and file tools.
Scroll horizontally to see all columns
| Layer | Can block a tool call? | Can resume a run? | Keeps task ownership after the session ends? |
|---|---|---|---|
| Framework | Only where you build it in | Yes, if you built on its persistence | Not inherently; requires a work-record layer |
| Harness | Yes, through its approval policy | Yes, inside its own session state | Not inherently; requires a work-record layer |
| Coding runtime | Yes: permission rules and sandbox | Yes: its own sessions | Not inherently; requires a work-record layer |
| Paperclip organizational-work layer | Native shell and file tools: no. Connected MCP and HTTP tools: yes, if the experimental tool gateway is on | It can run the owner again. It does not replay another runtime’s state | Yes |
What an agent framework owns
LangChain describes itself as “the agent framework: abstractions and integrations for models, tools, and agent loops,” and describes LangGraph as “the orchestration runtime: durable execution, streaming, human-in-the-loop, and persistence.” If you are building your own agent, this is where you decide its shape.
A framework can give you real controls. You can interrupt a graph before a risky step, or checkpoint its state so a failed run resumes where it stopped. But those controls exist because you wired them into your own graph. They do nothing for an agent you didn’t build with that framework, such as a Claude Code session your colleague started.
What an agent harness owns
Microsoft’s Agent Framework documentation defines a harness as “the runtime scaffolding that turns a language model into an agent that can perform work. It drives model and tool calls, manages conversation state and context, applies approval policies, and can keep the agent progressing through a multistep task.” Its Harness composes existing framework components rather than defining a separate runtime. Standing tool approvals and auto-approval rules are on by default. In Python, background agents, file access, and looping are still experimental.
AWS uses the word more broadly: the orchestration loop plus the infrastructure under it, including compute, sandbox, tool connections, filesystem, memory, identity, and observability. Its managed harness runs inside AgentCore Runtime, with an isolated microVM per session. That isolation is a property of AWS’s product, not of harnesses in general. LangChain’s Deep Agents is a third shape: “planning, subagents, filesystem tools, and context management on top of LangGraph.”
Across all three, the harness is where per-step decisions happen. It decides whether a tool call runs, what context the model sees, and how a session continues.
What a coding runtime owns
Claude Code and Codex are harnesses packaged for software work. They ship with a terminal or IDE surface, local tools, a permission system, a sandbox, and a session store. Two of those responsibilities matter most in this comparison.
Permissions. Claude Code evaluates permission rules “in order: deny, then ask, then allow,” and its documentation states that rules “are enforced by Claude Code, not by the model.” Deny rules block in every mode, including bypassPermissions. A Bash rule matches the command text Claude writes, so it “isn’t a security boundary around the program.” For enforcement that doesn’t depend on command text, Claude Code has an OS-level sandbox that limits which paths a command can write and which domains it can reach. Codex offers --sandbox modes (read-only, workspace-write, danger-full-access) and an --ask-for-approval policy. Its --dangerously-bypass-approvals-and-sandbox flag runs “every command without approvals or sandboxing,” and Codex advises using it only inside an externally hardened environment.
Sessions. claude --resume <session> and claude --continue reopen a Claude Code conversation. codex resume and codex exec resume --last do the same for Codex. Both resume the runtime’s own conversation from that runtime’s own records.
What a multi-agent control plane owns
Paperclip describes itself as “the layer above all of that — the org chart, the task board, the budget controller, the audit trail,” and is direct about the boundary: “It’s not the AI.” Agents run on the runtimes you connect through adapters, including Claude Code and Codex. Roles and runtimes are separate, so the same job can move from one runtime to another.
What Paperclip governs is the task. A Paperclip issue has one owner, a status, and blocked by links. While a blocker is unresolved, agents treat the issue as not ready to start. An execution policy can intercept an agent’s attempt to mark work done and route it to a reviewer or approver first. If the reviewer requests changes, the issue returns to the original executor. Budgets pause an agent at its limit.
Two Paperclip controls can stop something, and they are easy to confuse. The execution policy gates issue transitions, such as marking work done. It doesn’t see the commands an agent runs. The tool gateway brokers calls to tools you connect through Paperclip, such as MCP servers and internal HTTP services. It can allow, block, require approval for, rate-limit, or defer each call. It is experimental and off by default until an operator turns on enableApps. Neither control reaches the runtime’s native shell and file tools, which run inside the runtime and whatever sandbox it uses. Paperclip doesn’t build your agents, doesn’t itself enforce a sandbox around the commands they run, and doesn’t replay another runtime’s internal state.
A worked scenario: one blocked command and one handoff
This is an illustrative scenario built from documented behavior, not a recorded incident. The configuration is stated so you can check each step against your own setup.
Setup
- Paperclip runs two agents. The API Engineer uses the
claude_localadapter. The Integration Engineer usescodex_local. Both run in git worktrees. - Issue A, “Add pagination to the runs endpoint,” is owned by the API Engineer and has a review stage.
- Issue B, “Update the client SDK for paginated runs,” is owned by the Integration Engineer and is blocked by A.
- The Claude Code settings the session loads include the deny rule
Bash(git push --force *). Paperclip’s Claude Code adapter leavesdangerouslySkipPermissionsat its default oftrue, because it runs Claude headless. - The Integration Engineer’s Codex adapter keeps its default engine,
auto, which runs Codex through ACP.
1. The agent tries a forbidden command
Midway through A, Claude Code proposes git push --force origin feature/pagination. Claude Code blocks it. The deny rule matches, and deny rules hold even in the bypass mode Paperclip uses for unattended runs. Paperclip is not consulted. It sees only what the run outputs and what the agent reports afterward.
That block is narrower than it looks. Three separate boundaries could stop a force push, and each checks something different:
- The permission deny rule matches the command text Claude writes. The same push wrapped in a script may not match.
- The OS-level sandbox limits which files a command can write and which hosts it can reach. Your Git host has to be an allowed host for ordinary pushes to work, so the sandbox doesn’t stop a force push to it.
- Branch protection on your Git host decides whether the remote accepts a force push at all, however the command was written. If it isn’t configured, nothing on the remote side refuses the push.
2. The session ends before the work is done
The run ends with A half finished. Claude Code resumes its own conversation, and Paperclip decides when that happens. Paperclip’s adapter stores the Claude Code session ID and resumes it on the next heartbeat if the working directory still matches. If the adapter can’t resume, it falls back to a fresh session.
A fresh session doesn’t have the old conversation. It has what survived outside it: the issue, its comments and documents, and the branch. That is the practical design rule. Write progress into the task, not only into the transcript.
3. The work changes hands
The API Engineer marks A done. Paperclip’s execution policy intercepts the change. A moves to review instead of done, and the reviewer becomes responsible. When the reviewer approves, A is done, B’s blocker is resolved, and B becomes ready for the Integration Engineer’s next run in Codex.
The Claude Code session for A no longer matters to the handoff. Ownership is in the record: who held A, who reviewed it, what they decided, and who owns B now. If the reviewer had requested changes, Paperclip would have returned A to the API Engineer, regardless of which sessions were still open.
4. Codex reaches its own boundary
While working on B, Codex tries to edit a configuration file outside its workspace. The Codex sandbox blocks the write. On the default engine (auto, which runs ACP), Paperclip’s Codex adapter keeps Codex in its writable workspace sandbox, with network access on for each turn.
The classic Codex CLI lane (engine: "cli") works differently. There the adapter adds --dangerously-bypass-approvals-and-sandbox by default. If you set dangerouslyBypassApprovalsAndSandbox yourself, that setting decides: false keeps the sandbox and true bypasses it. If you leave it unset, the adapter skips the bypass, and Codex runs in its workspace sandbox or in the mode you chose, when any of these is true:
extraArgschoose a sandbox mode or profile, such as--sandbox,--profile,--full-auto, or asandbox_mode=override.extraArgsset anapproval_policy=override, or turn network access off withsandbox_workspace_write.network_access=false.- The execution target denies network access.
With the bypass in effect, nothing in this stack would have blocked that write. Full auto “is about approval prompts, not about what an agent is allowed to do in your company.” Company permissions, budgets, and approval requirements still apply, but none of them sandboxes a shell command.
When you need a control plane
If one developer runs one runtime and can see every session, the runtime’s own tools cover the job. Claude Code’s Agent View, claude --resume, and codex resume answer “what is running and how do I get back to it?” A control plane would add overhead without adding an owner who isn’t already watching. Our guide to managing Claude Code and Codex sessions compares those options in detail.
The calculation changes when responsibility outlives the session. These are the usual signals:
- More than one runtime contributes to the same deliverable.
- The next owner of a piece of work is not in the session that produced it.
- Someone other than the executor has to review or approve the result.
- A spending limit has to stop work, not only display a number.
- Someone needs to answer “whose job was this?” after every session has ended.
That is the layer Paperclip provides. For how organizational work control compares with workflow, gateway, runtime-action, and identity control planes, see What is an AI agent control plane?
Assign each responsibility before you need it
Before the first unattended run, answer these for your own stack:
- Which component blocks a risky tool call, and is it still active in unattended mode? Check your adapter defaults, including which engine each adapter runs.
- Is that block a text match, an OS-level sandbox, or a rule on the remote service, such as branch protection?
- Where does a session resume from, and what makes it start fresh instead?
- What would a fresh session need to continue? Put that in the task.
- Which record names the owner after the session ends?
- Which system is authoritative for merge and deployment? Usually that is your Git host and CI, not a task status.
Frequently asked questions
Is Claude Code an agent harness?
Yes. Claude Code is a harness built for coding, packaged with its own runtime, permission rules, sandbox, and session store. So is Codex. “Coding runtime” is a narrower label for the same kind of product.
Is LangGraph a framework or a runtime?
Both. LangChain’s documentation calls LangGraph an orchestration runtime with durable execution and persistence, and developers also use it as a framework for defining agent workflows. Ask which responsibility you need it to own.
Can Paperclip block or sandbox the commands its agents run?
Not the runtime’s own shell and file commands. That boundary belongs to the coding runtime or the environment it runs in, so check each adapter’s defaults. Paperclip gates issue transitions through its execution policy, and its experimental tool gateway, off by default, can gate calls to connected MCP and HTTP tools.
Sources
Reviewed October 7, 2026. Paperclip adapter and tool gateway pages and Claude Code sandboxing rechecked October 8, 2026.
- Microsoft, Agent Harness
- AWS, AgentCore harness and AgentCore harness vs. Runtime
- LangChain, LangGraph overview
- Anthropic, Claude Code permissions, permission modes, sandboxing, and CLI reference
- OpenAI, Codex CLI commands
- Paperclip, What is Paperclip, agent adapters, Claude Code adapter, Codex adapter, tool gateway, issues, and execution policy