
A harness for building durable teams of native AI CLI agents.
Workbench turns interactive tools like Claude Code, Codex, Antigravity, Grok Build, Hermes, and OpenCode into specialized agents — and specialized agents into durable teams that work together over time. Each agent keeps its native CLI capabilities, memory, skills, hooks, authentication, and conversational interface. Workbench adds the missing operating layer: controlled launch profiles, persistent sessions, a shared workspace, team-to-team communication, scheduling, GitHub integration, and human oversight.
The result is not a swarm, a wrapper, or a generic agent framework. It is an agentic organization: durable teams with their own context, responsibilities, memories, and tools.
Single agents are useful. Sub-agent trees are better. But the strongest pattern we have found is a team of fully empowered, durable agents that can freely exchange context, challenge each other, specialize by role, and draw on cognitively diverse model families — while still operating inside the native CLIs each provider has invested in.
Most agent systems settle for less. They run a single agent, lock into a single provider, wrap shallow orchestration around one model, or turn many identical agents loose as a swarm. Each gives up something: durability, cognitive diversity, real specialization, or the native capabilities the CLI vendors are pouring billions into. Workbench keeps all of it.
Three layers, each durable, each built from the native tools rather than around them.
A general-purpose CLI becomes a specific, controllable agent through its launch profile, system prompt, skill files, durable memory, hooks, and session — a Reviewer, an SRE, a Controller — that persists across sessions instead of resetting each run.
Role agents combine into durable teams. A software product team carries engineering, review, root-cause analysis, testing, orchestration, and — where it fits — program management, documentation, and release. The team is summonable and keeps its context.
Teams talk to other teams. A product team that needs infrastructure contacts the infrastructure team. The work — and its durable memory — stays attached to the team that owns it, so responsibility and context don’t leak.
Agents communicate the way a person would talk to them: by sending an identity-tagged prompt through the same conversational interface a human uses. Nothing is bolted on the side. Because the message arrives as a normal turn inside the native CLI, it automatically carries that CLI’s durable memory, compaction, session context, and tools. Agent-to-agent and team-to-team communication reuse the exact machinery each provider already built — they don’t work around it.
Workbench runs the native interactive CLIs, visibly, in a web terminal. It does not replace or modify them — and that is the point.
Through the CLI layer, a team can run any model, cloud or local, one provider or many — Claude Code, Codex, Antigravity, Grok Build, Hermes, OpenCode — and mix cognitively diverse model families, because different models think differently and are good at different things.
You get each vendor’s best agentic tool, updated on their schedule, not a third-party approximation. No CLI modification, no disguised UI, no hidden auth automation, and no batch invocations standing in for an API — users authenticate normally through OAuth subscriptions, and the CLIs update normally.
Workbench never reaches inside an agent’s reasoning. It controls how agents are spawned and operated, and what surrounds the team.
Each CLI increasingly ships its own durable memory, profile, skill, and session concepts. Workbench uses those native mechanisms rather than forcing all continuity through a retrieval layer. Because the agents are durable and the teams are durable, a team accumulates collective memory — the context of its work stays attached to the team that owns it, and compounds.
An Infrastructure team owns DNS, PXE, networking, containers, observability, and cloud. Each software product has its own summonable, durable team — engineers, reviewers, root-cause analysis, testing, and orchestration, plus technical program management, documentation, and release where they fit.
When a product team needs infrastructure, it contacts the infrastructure team directly, agent to agent, through the normal conversational interface. The request participates in each team’s native memory and tools — and when it’s done, the work and its durable memory stay with the team that owns it.
Four durable teams across different domains. In each, the human operator works directly with the team and plays the strategic lead role. Model choices are illustrative — every agent runs inside a native CLI, and the team mixes model families on purpose.
Users work primarily in the Workbench UI, operating their agent teams directly in the web terminal.
Connected to a primary tool such as Claude Desktop via MCP. Users work in their primary tool and reach their agent teams through it.
Agent teams serve an organization and multiple users more broadly, often using email or messaging as the interface between people and the teams.
Blueprint Agentic Workbench is proprietary BlueprintAgentic technology, built with correctly licensed open-source components. It is not open source and is not publicly distributed.
It is the tool the practice uses for its own daily work and deploys within client engagements. To talk about using it, start a conversation.