Latest release v2026.1005.0 is live (October 5, 2026).

See what's new
← All articles

Agent Harness vs. Framework vs. Runtime vs. Control Plane

When Claude Code and Codex share work, which layer stops a risky command, which resumes a broken session, and which still knows who owns the task?

A Claude Code session tries to force-push a branch. A Codex run reaches for a file outside its workspace. A session ends halfway through a task whose next step belongs to someone else. When something goes wrong, which component can stop it, which can recover it, and which still knows whose job it was?

Four kinds of software claim a piece of that answer:

  • An agent framework gives you the building blocks for agent logic: model and tool abstractions, agent loops, and workflow primitives.
  • An agent harness runs that loop as a working agent. It drives model and tool calls, manages context and state, applies approval policies, and keeps a multistep task moving.
  • A coding runtime such as Claude Code or Codex is a harness packaged for software work, with its own tools, permission system, sandbox, and session store.
  • A multi-agent control plane sits above the runtimes and governs the work itself: which tasks exist, who owns each one, what it depends on, what it may spend, and who must review it.

These are responsibilities, not exclusive product categories. LangChain’s documentation calls LangGraph an orchestration runtime, and many teams also use it as a framework. AWS runs its managed harness inside AgentCore Runtime. Claude Code is a harness, a runtime, and a product at once. This guide maps the responsibilities, then walks one handoff through all of them.

The four layers at a glance

In short: harnesses and runtimes stop tool calls, and each runtime resumes its own sessions. In this setup, Paperclip retains the shared task-owner record after a session ends. Its experimental tool gateway, off by default, can also gate calls to tools connected through Paperclip, but not the runtime’s own shell and file tools.

Scroll horizontally to see all columns

LayerCan block a tool call?Can resume a run?Keeps task ownership after the session ends?
FrameworkOnly where you build it inYes, if you built on its persistenceNot inherently; requires a work-record layer
HarnessYes, through its approval policyYes, inside its own session stateNot inherently; requires a work-record layer
Coding runtimeYes: permission rules and sandboxYes: its own sessionsNot inherently; requires a work-record layer
Paperclip organizational-work layerNative shell and file tools: no. Connected MCP and HTTP tools: yes, if the experimental tool gateway is onIt can run the owner again. It does not replay another runtime’s stateYes
Examples used in this guide: LangChain and LangGraph (frameworks); Microsoft Agent Framework Harness, LangChain Deep Agents, and the AgentCore harness (harnesses); Claude Code and Codex CLI (coding runtimes); and Paperclip (control plane). The table reflects first-party documentation reviewed October 7–8, 2026. It describes documented responsibilities, not a security certification of any product.

What an agent framework owns

LangChain describes itself as “the agent framework: abstractions and integrations for models, tools, and agent loops,” and describes LangGraph as “the orchestration runtime: durable execution, streaming, human-in-the-loop, and persistence.” If you are building your own agent, this is where you decide its shape.

A framework can give you real controls. You can interrupt a graph before a risky step, or checkpoint its state so a failed run resumes where it stopped. But those controls exist because you wired them into your own graph. They do nothing for an agent you didn’t build with that framework, such as a Claude Code session your colleague started.

What an agent harness owns

Microsoft’s Agent Framework documentation defines a harness as “the runtime scaffolding that turns a language model into an agent that can perform work. It drives model and tool calls, manages conversation state and context, applies approval policies, and can keep the agent progressing through a multistep task.” Its Harness composes existing framework components rather than defining a separate runtime. Standing tool approvals and auto-approval rules are on by default. In Python, background agents, file access, and looping are still experimental.

AWS uses the word more broadly: the orchestration loop plus the infrastructure under it, including compute, sandbox, tool connections, filesystem, memory, identity, and observability. Its managed harness runs inside AgentCore Runtime, with an isolated microVM per session. That isolation is a property of AWS’s product, not of harnesses in general. LangChain’s Deep Agents is a third shape: “planning, subagents, filesystem tools, and context management on top of LangGraph.”

Across all three, the harness is where per-step decisions happen. It decides whether a tool call runs, what context the model sees, and how a session continues.

What a coding runtime owns

Claude Code and Codex are harnesses packaged for software work. They ship with a terminal or IDE surface, local tools, a permission system, a sandbox, and a session store. Two of those responsibilities matter most in this comparison.

Permissions. Claude Code evaluates permission rules “in order: deny, then ask, then allow,” and its documentation states that rules “are enforced by Claude Code, not by the model.” Deny rules block in every mode, including bypassPermissions. A Bash rule matches the command text Claude writes, so it “isn’t a security boundary around the program.” For enforcement that doesn’t depend on command text, Claude Code has an OS-level sandbox that limits which paths a command can write and which domains it can reach. Codex offers --sandbox modes (read-only, workspace-write, danger-full-access) and an --ask-for-approval policy. Its --dangerously-bypass-approvals-and-sandbox flag runs “every command without approvals or sandboxing,” and Codex advises using it only inside an externally hardened environment.

Sessions. claude --resume <session> and claude --continue reopen a Claude Code conversation. codex resume and codex exec resume --last do the same for Codex. Both resume the runtime’s own conversation from that runtime’s own records.

What a multi-agent control plane owns

Paperclip describes itself as “the layer above all of that — the org chart, the task board, the budget controller, the audit trail,” and is direct about the boundary: “It’s not the AI.” Agents run on the runtimes you connect through adapters, including Claude Code and Codex. Roles and runtimes are separate, so the same job can move from one runtime to another.

What Paperclip governs is the task. A Paperclip issue has one owner, a status, and blocked by links. While a blocker is unresolved, agents treat the issue as not ready to start. An execution policy can intercept an agent’s attempt to mark work done and route it to a reviewer or approver first. If the reviewer requests changes, the issue returns to the original executor. Budgets pause an agent at its limit.

Two Paperclip controls can stop something, and they are easy to confuse. The execution policy gates issue transitions, such as marking work done. It doesn’t see the commands an agent runs. The tool gateway brokers calls to tools you connect through Paperclip, such as MCP servers and internal HTTP services. It can allow, block, require approval for, rate-limit, or defer each call. It is experimental and off by default until an operator turns on enableApps. Neither control reaches the runtime’s native shell and file tools, which run inside the runtime and whatever sandbox it uses. Paperclip doesn’t build your agents, doesn’t itself enforce a sandbox around the commands they run, and doesn’t replay another runtime’s internal state.

A worked scenario: one blocked command and one handoff

This is an illustrative scenario built from documented behavior, not a recorded incident. The configuration is stated so you can check each step against your own setup.

Setup

  • Paperclip runs two agents. The API Engineer uses the claude_local adapter. The Integration Engineer uses codex_local. Both run in git worktrees.
  • Issue A, “Add pagination to the runs endpoint,” is owned by the API Engineer and has a review stage.
  • Issue B, “Update the client SDK for paginated runs,” is owned by the Integration Engineer and is blocked by A.
  • The Claude Code settings the session loads include the deny rule Bash(git push --force *). Paperclip’s Claude Code adapter leaves dangerouslySkipPermissions at its default of true, because it runs Claude headless.
  • The Integration Engineer’s Codex adapter keeps its default engine, auto, which runs Codex through ACP.
Swimlane diagram of the worked scenario. Lanes: Claude Code (API Engineer, issue A), Codex (Integration Engineer, issue B), Paperclip (issues A and B), and the Git host if configured. Moment 1, force push attempted: Claude Code decides, because its deny rule matches the command text git push --force and blocks it; Paperclip only records the run output and is not consulted; branch protection on the Git host is a backstop that refuses a force push however it is spelled. Moment 2, session ends mid-task: Claude Code resumes its own session if the working directory still matches, otherwise starts fresh; Paperclip decides when, because the next heartbeat restarts the owner, and records the owner, status, comments, and branch. Moment 3, A handed to B: Paperclip decides, because its review stage intercepts done and approval unblocks B, and it records the reviewer's decision and B's owner; the Claude Code session no longer matters and B is ready for Codex's next run. Moment 4, write outside the workspace: Codex decides, because the workspace sandbox on the default ACP engine blocks the write, though on the CLI lane with the bypass in effect nothing blocks it; Paperclip only records the run output and is not in the path.
The four moments below, side by side. Illustrative scenario built from documented behavior, not a recorded incident.

1. The agent tries a forbidden command

Midway through A, Claude Code proposes git push --force origin feature/pagination. Claude Code blocks it. The deny rule matches, and deny rules hold even in the bypass mode Paperclip uses for unattended runs. Paperclip is not consulted. It sees only what the run outputs and what the agent reports afterward.

That block is narrower than it looks. Three separate boundaries could stop a force push, and each checks something different:

  • The permission deny rule matches the command text Claude writes. The same push wrapped in a script may not match.
  • The OS-level sandbox limits which files a command can write and which hosts it can reach. Your Git host has to be an allowed host for ordinary pushes to work, so the sandbox doesn’t stop a force push to it.
  • Branch protection on your Git host decides whether the remote accepts a force push at all, however the command was written. If it isn’t configured, nothing on the remote side refuses the push.

2. The session ends before the work is done

The run ends with A half finished. Claude Code resumes its own conversation, and Paperclip decides when that happens. Paperclip’s adapter stores the Claude Code session ID and resumes it on the next heartbeat if the working directory still matches. If the adapter can’t resume, it falls back to a fresh session.

A fresh session doesn’t have the old conversation. It has what survived outside it: the issue, its comments and documents, and the branch. That is the practical design rule. Write progress into the task, not only into the transcript.

3. The work changes hands

The API Engineer marks A done. Paperclip’s execution policy intercepts the change. A moves to review instead of done, and the reviewer becomes responsible. When the reviewer approves, A is done, B’s blocker is resolved, and B becomes ready for the Integration Engineer’s next run in Codex.

The Claude Code session for A no longer matters to the handoff. Ownership is in the record: who held A, who reviewed it, what they decided, and who owns B now. If the reviewer had requested changes, Paperclip would have returned A to the API Engineer, regardless of which sessions were still open.

4. Codex reaches its own boundary

While working on B, Codex tries to edit a configuration file outside its workspace. The Codex sandbox blocks the write. On the default engine (auto, which runs ACP), Paperclip’s Codex adapter keeps Codex in its writable workspace sandbox, with network access on for each turn.

The classic Codex CLI lane (engine: "cli") works differently. There the adapter adds --dangerously-bypass-approvals-and-sandbox by default. If you set dangerouslyBypassApprovalsAndSandbox yourself, that setting decides: false keeps the sandbox and true bypasses it. If you leave it unset, the adapter skips the bypass, and Codex runs in its workspace sandbox or in the mode you chose, when any of these is true:

  • extraArgs choose a sandbox mode or profile, such as --sandbox, --profile, --full-auto, or a sandbox_mode= override.
  • extraArgs set an approval_policy= override, or turn network access off with sandbox_workspace_write.network_access=false.
  • The execution target denies network access.

With the bypass in effect, nothing in this stack would have blocked that write. Full auto “is about approval prompts, not about what an agent is allowed to do in your company.” Company permissions, budgets, and approval requirements still apply, but none of them sandboxes a shell command.

When you need a control plane

If one developer runs one runtime and can see every session, the runtime’s own tools cover the job. Claude Code’s Agent View, claude --resume, and codex resume answer “what is running and how do I get back to it?” A control plane would add overhead without adding an owner who isn’t already watching. Our guide to managing Claude Code and Codex sessions compares those options in detail.

The calculation changes when responsibility outlives the session. These are the usual signals:

  • More than one runtime contributes to the same deliverable.
  • The next owner of a piece of work is not in the session that produced it.
  • Someone other than the executor has to review or approve the result.
  • A spending limit has to stop work, not only display a number.
  • Someone needs to answer “whose job was this?” after every session has ended.

That is the layer Paperclip provides. For how organizational work control compares with workflow, gateway, runtime-action, and identity control planes, see What is an AI agent control plane?

Assign each responsibility before you need it

Before the first unattended run, answer these for your own stack:

  1. Which component blocks a risky tool call, and is it still active in unattended mode? Check your adapter defaults, including which engine each adapter runs.
  2. Is that block a text match, an OS-level sandbox, or a rule on the remote service, such as branch protection?
  3. Where does a session resume from, and what makes it start fresh instead?
  4. What would a fresh session need to continue? Put that in the task.
  5. Which record names the owner after the session ends?
  6. Which system is authoritative for merge and deployment? Usually that is your Git host and CI, not a task status.

Frequently asked questions

Is Claude Code an agent harness?

Yes. Claude Code is a harness built for coding, packaged with its own runtime, permission rules, sandbox, and session store. So is Codex. “Coding runtime” is a narrower label for the same kind of product.

Is LangGraph a framework or a runtime?

Both. LangChain’s documentation calls LangGraph an orchestration runtime with durable execution and persistence, and developers also use it as a framework for defining agent workflows. Ask which responsibility you need it to own.

Can Paperclip block or sandbox the commands its agents run?

Not the runtime’s own shell and file commands. That boundary belongs to the coding runtime or the environment it runs in, so check each adapter’s defaults. Paperclip gates issue transitions through its execution policy, and its experimental tool gateway, off by default, can gate calls to connected MCP and HTTP tools.

Sources

Reviewed October 7, 2026. Paperclip adapter and tool gateway pages and Claude Code sandboxing rechecked October 8, 2026.

A team of agents
for every person.

Join the Paperclip Cloud beta waitlist,
or run Paperclip on your own infrastructure.