← BlueprintAgentic

Blueprint Agentic Workbench.

An agent that can call sub-agents is powerful; teams of agents collaborating are more powerful. Workbench is not a coding tool — it's an agent-teaming harness, built on 3 key facts.

  1. Teams of agents are far more powerful than agents with sub-agents.
  2. There are massive advantages to sticking with provider native agentic tools.
    • Native provider tools outperform third-party tools due to integrations with their models.
    • Third-party developers can't match the multi-billion-dollar investments major providers pour into their own agentic tools.
    • Provider subscriptions with native tools cost a fraction of pay-per-use options.
    • Global standards and concepts now allow for interoperability between provider native agentic tools.
  3. Using a variety of models is crucial because different models both think differently and are good at different things. This is true of models both within a given provider and between providers.

Workbench provides an organizing harness that allows native agentic tools like Claude Code or ChatGPT's Codex to team together without sacrificing any part of the native tools.

A community of agents.

Today's provider-native tools provide a powerful construct: full agents with sub-agents. If a full agent is an employee, sub-agents are interns. The Workbench enables higher-order agentics — full agents communicating and coordinating as teams, and teams in turn coordinating as a community.

The structure mimics human organizational hierarchies for essentially the same reasons.

Parallel work.

Each agent focuses on its own tasks while the other agents focus on theirs, each with its own automation and triggers.

Coordination.

Each agent can coordinate its work and share information with other agents. Your Financial Controller agent can direct your Financial Analyst agent. Your agentic Marketing team can work with your agentic Finance team.

Specialization.

Each agent can have a specific focus.

  • Use the model best suited to its domain.
  • Maintain chat history and context specific to its work.
  • Utilize skills, tools, rules and access levels specific to its role.
  • Achieve cost savings using lesser models when it makes sense.

Checks and balances.

Complementary team members assist each other with better decision-making.

  • High reasoning models often fail to follow rules, lower reasoning models are often the right choice for rule-following circumstances.
  • Quorums of cognitively diverse models prevent hallucination, surface diverse perspectives, and find edge cases.

Agentic Team Examples

Four teams across different domains. In each, the human operator works directly with the team and plays the strategic lead role.

Operations

Finance.

  • Controller — Claude Opus 4.7
    Senior oversight, multi-report reasoning across Finance functions
  • Analyst — GPT-5.5
    Heavy data analysis; top-tier reasoner
  • Reporting — Claude Sonnet 4.6
    Structured periodic generation; no reasoning premium needed
  • Compliance — Claude Opus 4.7
    Careful regulatory interpretation, low error tolerance
  • Taxes — Gemini 3.1 Pro
    Rules and interpretation; long-context reasoning
  • Process Manager — Claude Haiku 4.5
    Process orchestration and routing
    • Email Sub-Agent — Claude Haiku 4.5
      Checks team email and routes to the appropriate team member
Audit quorumClaude Haiku 4.5 · GPT-5.4 Mini · Gemini 3.5 Flash
Hard rules to facts; strict instruction-followers, not reasoners.
Creative

Marketing.

  • Program Manager — Claude Sonnet 4.6
    Process and coordination; balance of reasoning and rule following
  • Copywriter — GPT-5.4 Mini
    GPT family leads direct-response copy and headlines
  • Image Creator — Gemini 3.1 Pro
    Nano Banana 2 leads image-gen leaderboards
  • Video Creator — Gemini 3.1 Pro
    Veo 3.1 leads cinematic quality, 4K, native audio
  • Audio Creator — GPT-5.5
    Leads audio reasoning benchmarks
  • Social Media Manager — Gemini 3.5 Flash
    Cheapest and fastest for high-volume posting
  • Performance Marketing — Claude Sonnet 4.6
    Analytical reasoning, structured campaign briefs
Analyst quorumClaude Opus 4.7 · GPT-5.5 · Gemini 3.1 Pro
Top-tier strategy peers, one per provider.
Research

Ocean Geology.

  • Research Coordinator — Claude Sonnet 4.6
    Process and scheduling; balance of reasoning and rule following
    • Email Sub-Agent — Claude Haiku 4.5
      Manages communications between agents and team members
    • Document Manager Sub-Agent — Claude Haiku 4.5
      Keeps research materials organized and up to date
  • Marine Geophysicist — Gemini 3.1 Pro
    Leads grad-level science reasoning; 1M-token context for survey data
  • Sedimentologist — GPT-5.5
    Balanced reasoner; strong descriptive synthesis and writing
  • Geochemist — Claude Opus 4.7
    Leads tool-using scientific workflows; multi-step isotope reasoning
  • Paleoceanographer — Gemini 3.1 Pro
    Same grad-level lead; long context for deep-time synthesis
Research Assistant quorumClaude Opus 4.7 · GPT-5.5 · Gemini 3.1 Pro
Expertise-grade peers, one per provider.
Engineering

Software Development.

  • Program Manager — Claude Opus 4.7
    Senior architectural oversight, cross-system coordination
  • Product Engineer — GPT-5.5
    Code synthesis and direct execution
  • Test Engineer — Claude Opus 4.7
    Cognitive diversity vs. Product Engineer; deep test design and edge-case reasoning
  • Tech Writer — Claude Opus 4.7
    Long-form technical documentation; precise writing
  • SDLC Process Manager — Claude Haiku 4.5
    Strict process operation; rule-follower over reasoner
    • Code Test Runner Sub-Agent — Claude Haiku 4.5
      Runs code-based tests such as mocks and integration
    • UI Test Runner Sub-Agent — Claude Sonnet 4.6
      Uses a runbook to test the UI via browsers or desktop control, as a user would
    • CICD Sub-Agent — Claude Haiku 4.5
      Deterministic pipeline orchestration
    • Git Records Sub-Agent — Claude Sonnet 4.6
      Maintains issues, PRs, milestones, and related Git records
Reviewers quorumClaude Sonnet 4.6 · Codex · Gemini 3.1 Pro · Claude Haiku 4.5
Sonnet, Codex, and Gemini Pro for coding expertise. Haiku for compliance with policy, standards, and requirements.

Deployment Types

Direct.

Users interact primarily with the Workbench UI for their daily work, utilizing their agent teams directly.

Best for power users

Connected.

Connected to a primary tool such as Claude Desktop via MCP. Users interact with their primary tool for daily work, utilizing their agent teams through it.

Best for business users

Server.

Agent teams serve organizations and multiple users more broadly. Often uses email or messaging apps/services as the primary interface between users and the agents.

Custom deployment · pay-per-use accounts

Workbench is currently in Beta.

Research and Development

Workbench's principles came out of building multiple agentic ecosystems before it — agent-managed services working in team formats. Components of those fully custom-built ecosystems include Context Brokerage, Inference Brokerage, Conversation Brokerage, Agent Hosts, Meta-Programming Engineering Teams, Digital Content Generators, Policy-Based File Servers, Vector and Graph Knowledge Management, Agentic Monitoring and Observability, and Agentic CI/CD. Several case studies and concept papers document the journey: the possibilities of meta-coding through agentic teams, the efficacy of cross-model quorums, and the principles of ecosystems built from teams of full agents.

Concepts applied directly in Workbench:

Concepts that emerged from the same work but are not yet applied in Workbench — generally applicable to large agentic systems: