← Work
Blueprint Agentic technology

Blueprint Agentic Workbench

A harness for building durable teams of native AI CLI agents.

Workbench turns interactive tools like Claude Code, Codex, Antigravity, Grok Build, Hermes, and OpenCode into specialized agents — and specialized agents into durable teams that work together over time. Each agent keeps its native CLI capabilities, memory, skills, hooks, authentication, and conversational interface. Workbench adds the missing operating layer: controlled launch profiles, persistent sessions, a shared workspace, team-to-team communication, scheduling, GitHub integration, and human oversight.

The result is not a swarm, a wrapper, or a generic agent framework. It is an agentic organization: durable teams with their own context, responsibilities, memories, and tools.

Why teams.

Single agents are useful. Sub-agent trees are better. But the strongest pattern we have found is a team of fully empowered, durable agents that can freely exchange context, challenge each other, specialize by role, and draw on cognitively diverse model families — while still operating inside the native CLIs each provider has invested in.

Most agent systems settle for less. They run a single agent, lock into a single provider, wrap shallow orchestration around one model, or turn many identical agents loose as a swarm. Each gives up something: durability, cognitive diversity, real specialization, or the native capabilities the CLI vendors are pouring billions into. Workbench keeps all of it.

From agents to an organization.

Three layers, each durable, each built from the native tools rather than around them.

Agents

Durable role agents.

A general-purpose CLI becomes a specific, controllable agent through its launch profile, system prompt, skill files, durable memory, hooks, and session — a Reviewer, an SRE, a Controller — that persists across sessions instead of resetting each run.

Teams

Specialized durable teams.

Role agents combine into durable teams. A software product team carries engineering, review, root-cause analysis, testing, orchestration, and — where it fits — program management, documentation, and release. The team is summonable and keeps its context.

Organization

Teams that call teams.

Teams talk to other teams. A product team that needs infrastructure contacts the infrastructure team. The work — and its durable memory — stays attached to the team that owns it, so responsibility and context don’t leak.

How agents talk.

Agents communicate the way a person would talk to them: by sending an identity-tagged prompt through the same conversational interface a human uses. Nothing is bolted on the side. Because the message arrives as a normal turn inside the native CLI, it automatically carries that CLI’s durable memory, compaction, session context, and tools. Agent-to-agent and team-to-team communication reuse the exact machinery each provider already built — they don’t work around it.

Why native CLIs.

Workbench runs the native interactive CLIs, visibly, in a web terminal. It does not replace or modify them — and that is the point.

Any model, any location.

Through the CLI layer, a team can run any model, cloud or local, one provider or many — Claude Code, Codex, Antigravity, Grok Build, Hermes, OpenCode — and mix cognitively diverse model families, because different models think differently and are good at different things.

Provider investment, preserved.

You get each vendor’s best agentic tool, updated on their schedule, not a third-party approximation. No CLI modification, no disguised UI, no hidden auth automation, and no batch invocations standing in for an API — users authenticate normally through OAuth subscriptions, and the CLIs update normally.

What Workbench controls.

Workbench never reaches inside an agent’s reasoning. It controls how agents are spawned and operated, and what surrounds the team.

Collective memory.

Each CLI increasingly ships its own durable memory, profile, skill, and session concepts. Workbench uses those native mechanisms rather than forcing all continuity through a retrieval layer. Because the agents are durable and the teams are durable, a team accumulates collective memory — the context of its work stays attached to the team that owns it, and compounds.

Teams that call teams — an example.

An Infrastructure team owns DNS, PXE, networking, containers, observability, and cloud. Each software product has its own summonable, durable team — engineers, reviewers, root-cause analysis, testing, and orchestration, plus technical program management, documentation, and release where they fit.

When a product team needs infrastructure, it contacts the infrastructure team directly, agent to agent, through the normal conversational interface. The request participates in each team’s native memory and tools — and when it’s done, the work and its durable memory stay with the team that owns it.

Agentic team examples.

Four durable teams across different domains. In each, the human operator works directly with the team and plays the strategic lead role. Model choices are illustrative — every agent runs inside a native CLI, and the team mixes model families on purpose.

Operations

Finance.

  • Controller — Claude Opus 4.7
    Senior oversight, multi-report reasoning across Finance functions
  • Analyst — GPT-5.5
    Heavy data analysis; top-tier reasoner
  • Reporting — Claude Sonnet 4.6
    Structured periodic generation; no reasoning premium needed
  • Compliance — Claude Opus 4.7
    Careful regulatory interpretation, low error tolerance
  • Taxes — Gemini 3.1 Pro
    Rules and interpretation; long-context reasoning
  • Process Manager — Claude Haiku 4.5
    Process orchestration and routing
    • Email Sub-Agent — Claude Haiku 4.5
      Checks team email and routes to the appropriate team member
Audit quorumClaude Haiku 4.5 · GPT-5.4 Mini · Gemini 3.5 Flash
Hard rules to facts; strict instruction-followers, not reasoners.
Creative

Marketing.

  • Program Manager — Claude Sonnet 4.6
    Process and coordination; balance of reasoning and rule following
  • Copywriter — GPT-5.4 Mini
    GPT family leads direct-response copy and headlines
  • Image Creator — Gemini 3.1 Pro
    Nano Banana 2 leads image-gen leaderboards
  • Video Creator — Gemini 3.1 Pro
    Veo 3.1 leads cinematic quality, 4K, native audio
  • Audio Creator — GPT-5.5
    Leads audio reasoning benchmarks
  • Social Media Manager — Gemini 3.5 Flash
    Cheapest and fastest for high-volume posting
  • Performance Marketing — Claude Sonnet 4.6
    Analytical reasoning, structured campaign briefs
Analyst quorumClaude Opus 4.7 · GPT-5.5 · Gemini 3.1 Pro
Top-tier strategy peers, one per provider.
Research

Ocean Geology.

  • Research Coordinator — Claude Sonnet 4.6
    Process and scheduling; balance of reasoning and rule following
    • Email Sub-Agent — Claude Haiku 4.5
      Manages communications between agents and team members
    • Document Manager Sub-Agent — Claude Haiku 4.5
      Keeps research materials organized and up to date
  • Marine Geophysicist — Gemini 3.1 Pro
    Leads grad-level science reasoning; 1M-token context for survey data
  • Sedimentologist — GPT-5.5
    Balanced reasoner; strong descriptive synthesis and writing
  • Geochemist — Claude Opus 4.7
    Leads tool-using scientific workflows; multi-step isotope reasoning
  • Paleoceanographer — Gemini 3.1 Pro
    Same grad-level lead; long context for deep-time synthesis
Research Assistant quorumClaude Opus 4.7 · GPT-5.5 · Gemini 3.1 Pro
Expertise-grade peers, one per provider.
Engineering

Software Development.

  • Program Manager — Claude Opus 4.7
    Senior architectural oversight, cross-system coordination
  • Product Engineer — Codex
    Code synthesis and direct execution
  • Test Engineer — Claude Opus 4.7
    Cognitive diversity vs. Product Engineer; deep test design and edge-case reasoning
  • Tech Writer — Claude Opus 4.7
    Long-form technical documentation; precise writing
  • SDLC Process Manager — Claude Haiku 4.5
    Strict process operation; rule-follower over reasoner
    • Code Test Runner Sub-Agent — Claude Haiku 4.5
      Runs code-based tests such as mocks and integration
    • UI Test Runner Sub-Agent — Claude Sonnet 4.6
      Uses a runbook to test the UI via browsers or desktop control, as a user would
    • CICD Sub-Agent — Claude Haiku 4.5
      Deterministic pipeline orchestration
    • Git Records Sub-Agent — Claude Sonnet 4.6
      Maintains issues, PRs, milestones, and related Git records
Reviewers quorumClaude Sonnet 4.6 · Codex · Gemini 3.1 Pro · Claude Haiku 4.5
Sonnet, Codex, and Gemini Pro for coding expertise. Haiku for compliance with policy, standards, and requirements.

Deployment types.

Direct.

Users work primarily in the Workbench UI, operating their agent teams directly in the web terminal.

Best for power users

Connected.

Connected to a primary tool such as Claude Desktop via MCP. Users work in their primary tool and reach their agent teams through it.

Best for business users

Server.

Agent teams serve an organization and multiple users more broadly, often using email or messaging as the interface between people and the teams.

Organization-wide

Availability.

Blueprint Agentic Workbench is proprietary BlueprintAgentic technology, built with correctly licensed open-source components. It is not open source and is not publicly distributed.

It is the tool the practice uses for its own daily work and deploys within client engagements. To talk about using it, start a conversation.