Latest release v2026.1005.0 is live (October 5, 2026).

See what's new
← All articles

Self-Hosted AI Agent Stacks: What You Run and What You Still Own

Paperclip, LangGraph, CrewAI, AutoGen, LlamaIndex, and OpenHands all publish MIT-licensed code. What each asks you to run, where model calls still go, and what your team maintains.

An engineering lead wants agents to prepare code changes and draft weekly reports from internal documents, and wants that work to run on the company’s own servers. This is an illustrative requirement, not a customer account. The decision comes down to three things: what will actually run on our machines, what will still leave them, and who keeps it working?

Every project in this guide publishes its code under the MIT license: Paperclip, LangGraph, CrewAI, AutoGen, LlamaIndex, and OpenHands. The license tells you what you may do with the code. It doesn’t tell you which services you have to operate, where prompts and file contents go when an agent calls a model, which layers are separate commercial products, or who backs up the state. Those answers differ far more than the licenses do.

So compare a self-hosted stack on four questions:

  1. What do you install and run? Processes, databases, ports, and the runtime versions they depend on.
  2. Where do model calls go? A self-hosted orchestrator can still send every prompt to a hosted model API.
  3. Which layer is commercial or hosted? Tracing, deployment, and document parsing are often separate products from the open-source core.
  4. What state do you keep, and who maintains it? Run state, indexes, credentials, upgrades, and access control become your team’s work.

The six projects do different jobs

Before comparing deployments, separate the roles. LangGraph, CrewAI, AutoGen, and LlamaIndex are frameworks: libraries you use to build an agent application. Paperclip and OpenHands Agent Canvas are applications you deploy: services that run or coordinate agents, including coding agents such as Claude Code and Codex.

Those two applications overlap. Both coordinate several coding agents. Agent Canvas centers on starting agent conversations and running automations. Paperclip centers on the work itself: which task exists, who owns it, what it may spend, and who has to approve it. A framework application can sit underneath either one.

That is why “which project is best?” is the wrong question for the opening scenario. The lead may need one of each, and every component adds its own row of obligations. For the management controls themselves, see 7 Platforms for Managing AI Agent Teams. For where different control layers enforce decisions, see What Is an AI Agent Control Plane?

The deployment matrix

Scroll horizontally to see all columns

ProjectWhat you run and maintainWhere model calls goSeparate commercial product
PaperclipNode.js 24+ server with embedded PostgreSQL and the agent CLIs you choose; backups of the database, stored files and master key; upgrades; TLSTo the provider each runtime is set up for: Anthropic or AWS Bedrock for Claude Code; OpenAI or a compatible endpoint you configure for Codex. How the runtime signs in is a separate settingNone in the documentation cited here
LangGraphYour app, the library, and the checkpointer database you chooseWhichever models your code callsLangSmith; self-hosting it is an Enterprise add-on
CrewAIA Python 3.10–3.13 project (uv, crewai CLI), its API keys, and Flow stateSet per agent; the docs include Ollama for local modelsCrewAI AMP (hosted) or CrewAI Factory (self-hosted, including on-premises)
AutoGenExisting Python 3.10+ code, and a plan to migrate itThrough model-client extensions such as OpenAI'sNone documented in the repository. Its status is separate: maintenance mode
LlamaIndexPython packages, your index and vector store, and the parsing pipelineEmbeddings default to OpenAI; a tutorial runs fully local with OllamaLlamaParse for parsing and managed indexes
OpenHands Agent CanvasFrontend, agent server, and automation backend (Node.js 24+, uv); host hardening, API key, TLSBring your own model; can also drive Claude Code, Codex, or GeminiOpenHands Cloud or Enterprise backends
An illustrative stack for the opening scenario. Inside a solid frame labeled on your servers: the Paperclip service, which keeps tasks, owners, budgets, and approvals in embedded PostgreSQL with secrets encrypted by a master key, owned by the platform team; a Claude Code or Codex CLI runtime started by a Paperclip adapter, owned by the platform team; an optional report service you build on LangGraph, CrewAI, or LlamaIndex, which Paperclip's HTTP adapter wakes and which fetches the task and posts results through the Paperclip API, owned by app developers; and tool connections, owned by tool owners. Task records stay in the database. Arrows cross out of the frame: model calls from the runtime to the provider it is configured for, shown here as Anthropic or OpenAI; model and parsing calls from the report service to hosted model or parsing APIs unless routed to servers you run; and tool calls to the Git host and connected apps.
In this illustrated configuration, the task records and the agent runtime stay on your server. The model calls from Claude Code or Codex do not.

Paperclip coordinates the work; it doesn’t run the model

What you run. Paperclip installs with one command, npx paperclipai onboard, on Node.js 24 or later. The default setup starts an embedded PostgreSQL database, uses local disk storage, and creates a secrets master key that encrypts every API key stored in Paperclip. The documentation tells you to back that key up: lose it and you lose access to those secrets. For a team server, the documented path adds an authenticated deployment mode, a dedicated service user, systemd, and Nginx with HTTPS in front of a server bound to 127.0.0.1:3100.

What still leaves. Paperclip does not run the model loop. Each agent uses an adapter that launches a runtime on the machine, such as Claude Code, Codex, Gemini CLI, or OpenCode, and that runtime calls its model provider. How a runtime signs in and where its calls go are separate questions. The Claude Code adapter authenticates with an Anthropic API key, AWS Bedrock settings, or a Claude subscription login. The Codex adapter inherits the host’s ChatGPT login by default, uses an OpenAI API key when you set one, and can send calls to an OpenAI-compatible provider or gateway you define. In the configuration this guide illustrates, Claude Code calling Anthropic or Codex calling OpenAI, the task context, the files the agent reads, and the tool results go to that hosted provider. Self-hosting Paperclip keeps the task record on your server. It doesn’t make the model call local, and we did not test a configuration that keeps every model call inside your network.

What you still own. Backups of the database, of the agent files and uploads the database doesn’t contain, and of the master key; upgrades; sign-in for your team; TLS; the runtime CLIs; and each runtime’s credentials. Two defaults deserve a deliberate decision before unattended work:

  • Approval prompts. The coding adapters run in full-auto mode, so the runtime uses its tools without stopping to ask. Company permissions, budgets, and approval requirements still apply to the work in Paperclip. For Codex, the default engine (auto, which runs ACP) keeps Codex in its writable workspace sandbox. The classic CLI lane (engine: "cli") bypasses approvals and the sandbox by default unless you turn that off.
  • Connected tools. The Tool Gateway can allow, block, or require a person’s approval for calls to connected MCP servers and HTTP services. It is experimental and stays off until an operator turns on enableApps. It governs those connected tools, not a runtime’s own shell and file tools.

LangGraph: you own the application and its state

LangGraph’s documentation calls it a low-level orchestration framework and runtime for long-running, stateful agents, with durable execution, streaming, and human-in-the-loop steps. It runs inside the application you write, and persistence is something you configure. The persistence guide shows an in-memory checkpointer, SQLite for local development, and PostgreSQL for production. A checkpoint lets a graph resume where it stopped. It doesn’t tell anyone who owns a failed task or who should look at it.

The commercial layer is LangSmith, which covers tracing, evaluation, and a deployment platform. Tracing is opt-in: you set LANGSMITH_TRACING=true and an API key. Self-hosted LangSmith is an add-on to the Enterprise plan that needs a license key. It runs as a set of services, including a frontend, backend, queue, platform backend, and code-execution backend, over ClickHouse, PostgreSQL, Redis, and optional blob storage, installed on Kubernetes. LangGraph’s MIT license doesn’t extend to that platform.

LangGraph fits the report half of the scenario when you have developers who want precise control over one workflow and are willing to run its database.

CrewAI: Flows and Crews on your own Python install

CrewAI organizes work as Flows, event-driven workflows that manage state, and Crews, groups of role-based agents that a Flow calls for a specific task. The framework installs locally: Python 3.10 through 3.13, the uv package manager, and the crewai CLI, with API keys kept in a .env file. Its LLM documentation includes Ollama for local models, so a Crew can point at a model server you run if you have hardware for the model you choose.

CrewAI also offers enterprise deployment options: CrewAI AMP, a hosted platform, and CrewAI Factory, a containerized deployment for your own infrastructure that the docs say supports on-premises installs. Treat either as a separate product decision from the MIT framework.

AutoGen: an existing installation, not a new build

Microsoft’s AutoGen repository says the project is in maintenance mode. It will not receive new features or enhancements and is now community-managed; contributions are limited to bug fixes, security patches, and documentation. Microsoft directs new users to Microsoft Agent Framework and publishes a migration guide. The code is MIT-licensed; the documentation is CC BY 4.0. The repository documents no hosted or commercial layer for AutoGen itself, and that is a separate fact from its maintenance status.

AutoGen is in this guide because teams with existing AutoGen code still have to decide what to do with it. If yours runs AutoGen, it keeps running where you put it, and bug and security fixes can still be contributed. For the opening scenario, a new build, the deciding fact is that new features will not arrive: anything the scenario needs beyond today’s code, you would build and maintain yourself.

LlamaIndex: the data stays local only if every step does

The LlamaIndex framework is a Python toolkit for agents over your own data: connectors, indexes, query engines, agents, and workflows that can be deployed as services. That makes it a natural candidate for the report half of the scenario. Its default path is not local, though. The documentation notes that an index embeds pages with OpenAI by default. The local-models tutorial swaps in a Hugging Face embedding model and Llama 3.1 8B served by Ollama, and warns that you will need a machine with roughly 32 GB of RAM.

Parsing is the other boundary. LlamaParse, the hosted platform formerly called LlamaCloud, handles scanned PDFs, tables, and structured extraction as an API. LiteParse is the open-source option for parsing on your own machine. Decide which one each pipeline uses before you call the stack local.

OpenHands Agent Canvas: self-hosted, with host and network access

The current OpenHands repository leads with Agent Canvas, which it describes as a self-hosted developer control center for coding agents and automations. The project is labeled beta. It runs the open-source OpenHands agent and can also drive Claude Code, Codex, Gemini, or any Agent Client Protocol agent, across local, Docker, VM, and cloud backends, with the model of your choice.

Its self-hosting guide is direct about what you take on. On a VM, Agent Canvas runs a static frontend, an agent server, and an automation backend behind an ingress proxy, protected by an API key you generate. The guide warns that the agent can read and write the host’s filesystem, run shell commands, and reach the network, and that anyone who can reach the agent server can do the same. It tells you to treat the VM like a machine that holds production credentials and to block everything except SSH before the first start. The quickstart without a sandbox gives the agent full access to the installing machine’s filesystem; the Docker option mounts a projects directory you choose.

Putting the scenario together

Start from the requirement rather than the project list. Each choice below is bounded by what the documentation establishes.

  • If you need to know who owns each agent task, what it may spend, and who approved it, run Paperclip on your server with Claude Code or Codex as the runtime. Paperclip keeps those records in its own database. In the illustrated configuration, model calls still go to Anthropic or OpenAI, and your team owns the server, the backups, and the runtime credentials.
  • If developers mainly want to start, watch, and steer coding agents interactively, Agent Canvas centers on that. Isolate its VM before the first start, as its guide instructs.
  • If report documents must stay inside your network, build the report service on LlamaIndex or CrewAI with a local embedding model, a local model server such as Ollama, and local parsing, and size the machine for the model. A hosted model or parsing API anywhere in that pipeline breaks the requirement.
  • If you need more than one of these, compose them. Paperclip’s HTTP adapter can wake a service you control: each run sends one JSON request with the run ID, the agent ID, and a small context object such as the task ID and wake reason. It does not package the task, its documents, or its artifacts. Your service fetches what it needs and posts results back through the Paperclip API with its own key. If the endpoint is on a private or loopback address, the Paperclip host must list its exact origin (scheme, host, and port) in PAPERCLIP_HTTP_ADAPTER_PRIVATE_ENDPOINT_ALLOWLIST, or the request is refused before it connects. The adapter is configured through the API rather than the agent form today. A report service built on LangGraph, CrewAI, or LlamaIndex could sit behind a Paperclip agent this way, but it is a generic mechanism, not a packaged integration: you would build, secure, and test the endpoint yourself.

None of these components makes a stack air-gapped on its own. Isolation depends on the model route, the parsing route, the tool connections, and updates. Both the Paperclip and Agent Canvas quick-start commands download the software when they run, so an offline environment needs its own plan for installation and upgrades.

Before you commit

  • Trace one real task end to end. List every network destination it touches: model API, parser, Git host, package registry.
  • Name an owner for each stateful store. Paperclip’s database and master key, a LangGraph checkpointer, a vector index.
  • Find the commercial layer before you budget. Self-hosted LangSmith, CrewAI AMP or Factory, LlamaParse, OpenHands Cloud or Enterprise.
  • Read each runtime’s permission default. Paperclip’s coding adapters run full auto; Agent Canvas without a sandbox reaches the host’s filesystem.
  • Plan upgrades. AutoGen shows what happens when a dependency stops getting new features.

Acceptance checks for the deployment

We propose these checks; we did not run them for this guide. Each one has a result you can observe before go-live.

  1. Restart and recover. Stop the host in the middle of a task, then start it again. Pass: the services come back on their own, and an operator can see where the interrupted work stands and resume or reassign it.
  2. Restore from backup. Restore the backed-up database, stored files, and encryption key onto a clean machine. For Paperclip, that key is the secrets master key. Pass: stored secrets decrypt, and the task history and documents are intact.
  3. Record outbound destinations. Run one representative task with egress logging on. Pass: every destination in the log appears in your matrix, and nothing in the log is a surprise.
  4. Watch an unauthorized action fail. Attempt one action the setup should block, such as a tool call outside policy or a write outside the agent’s workspace. Pass: it is refused, and the refusal is recorded where an operator will see it.

An open-source license gives you permission to run the software on your own machines; actually deploying and operating it there is a separate job. The license doesn’t remove the outside services your agents depend on; it makes you responsible for knowing which ones they are. Pick each component for the job it does, then write down every row it adds to your matrix.

Sources

Checked October 7, 2026. These sources establish documented behavior, not results of our own tests.

A team of agents
for every person.

Join the Paperclip Cloud beta waitlist,
or run Paperclip on your own infrastructure.