AIGIS Workbench

Governed autonomy you can actually watch.

AIGIS Workbench is the governed core of AIGIS software engineering — a local-first workbench where your engineering standards run as machine-enforced gates on autonomous coding agents, not prompts they are politely asked to honor. Two provider families run every role, and can be pitted against each other to reason and to QC each other's work.

>_ in plain termsNew to AI coding tools? Think of it as a safe control room: the AI writes the code, and Workbench checks it against your rules before anything is kept.

>_ Bring your own provider — Codex/GPT and Claude. AIGIS sells the engineering control layer, not model credits.

workbench_posture
>_posture........machine_enforced_gates
>_providers......codex_gpt + claude
>_autonomy.......unattended_or_one_decision
>_honesty........no_fabricated_metrics
Scroll
aigis_workbench — governed_apply · live

Work flows in from two providers, the gate decides, and only clean work stacks into history — held work never slips through. Illustrative loop.

Why it is different

Standards as enforced gates — not polite requests.

Most AI coding tools hand the agent a style guide and hope. AIGIS Workbench compiles your standards into gates that run against the work, block a bad apply, and stay separated from advisory rules. That is the whole thesis: governed autonomy, honest by construction.

shared_architecture

>_foundation......shared_with_flagship_adsp1 >_role............stands_alone_as_a_product >_second_role.....proves_architecture_for_flagships

Workbench is built on the same architecture that powers our flagship ADSP1. It stands on its own as a product — and it is where AIGIS exercises and validates the engineering quality that flows into everything we ship.

>_ in plain termsYou set the rules once, and the workbench actually enforces them — instead of hoping the AI remembers. It is the same engine behind our flagship product, offered on its own.

>_

Machine-enforced gates

Enforced rules run as deterministic, zero-cost checks that hold a patch back — not text an agent can quietly skip. Advisory rules stay clearly labeled as advice.

>_

Two provider families

Claude and Codex/GPT run every role. Route each role independently, or pit them against each other to reason and to review each other's work.

>_

Honest by construction

Reviewers are read-only, cost is real or labeled an estimate, and no score, benchmark, or metric on any surface is fabricated. If a number appears, the product measured it.

The differentiator

Two provider families. One artifact. Structural disagreement.

One primitive, two surfaces: two different provider families reason independently on the same artifact, their disagreement is surfaced structurally, and a neutral lighter model synthesizes the result — never a fabricated agreement score.

>_ in plain termsTwo different AIs answer the same thing on their own, then a third neutral one writes up where they agreed and where they still disagree — so you are never trusting a single opinion.

Debate · artifact = a prompt

Answer independently, then rebut.

One prompt. Claude and Codex each answer first with no peeking, then run configurable structured rebuttal rounds. A separate neutral model — the Documenter role — organizes one combined answer plus an explicit "still contested" list. Watch divergence, rebuttal, and synthesis live, side by side, over SSE.

Measured: rounds, tokens & cost per provider, latency, converged vs forced.
Derived: resolved / unresolved points, provenance.
Self-reported: confidence. No invented quality score.

Review · artifact = a finished patch

The opposite — and the same — provider QC it.

Click Review on a patch or lane and the opposite provider and the same provider each QC the work, read-only, run as a task with a brief. Same-provider review catches execution slips; opposite-provider catches blind spots. When the two reviewers disagree with each other, that disagreement is elevated — not averaged away.

Read-only reviewers. Disagreement is surfaced structurally.
Neutral synthesis combines the result — no fabricated agreement metric.

>_

Debate and Review are committed day-one launch capabilities, currently being finalized. We describe what they do — we do not fake it with invented transcripts, scores, or benchmarks.

The workspace

Three modes, four control surfaces, one governed core.

Register your repositories, then work the way the task wants — unattended, hands-on, or conversational — over one governed core. Runs and apply are scoped per repo. Premium light and dark themes.

>_ in plain termsPick how hands-on you want to be: let it run on its own, guide it step by step, or just chat with it — all in one place.

>_batch

Batch Process

The unattended governed pipeline: compose, lint, queue, run, review, apply.

>_manual

Manual

Ad-hoc cockpit: generate a grounded prompt, run one lane, gate the apply.

>_repo

Repo

Fluid single-agent conversation with git as the safety net.

>_backlog

Backlog

Drop a list, or let an agent propose ranked, evidence-cited work.

>_memory

Memory

Token-budgeted context retrieval and operator-approved memory cards.

>_standards

Standards

Enforced gates and advisory rules, honestly separated, with a drift readout.

>_settings

Settings

Per-role model routing, concurrency and burn caps, default review policy.

>_registry

Multi-repo registry

Register N repos with provider and scan depth; everything scopes per repo.

operational_status_bar

>_repo...........scoped_per_registered_repo >_providers......codex + claude live_status >_state..........applied / ready / idle >_infra..........poll_cadence · dedup · sse · in_flight

Batch Process

Runs unattended — or hands you exactly one decision.

Author a task, lint it against a hard gate, and route it into the pipeline. One task is a trajectory; several become a multi-task trajectory, with automatic lane routing and grounding pulled from the repo. Then the scheduler runs it — and every action publishes to one live governance stream.

>_ in plain termsLine up the work, press go, and walk away. If something needs you, it asks one clear question instead of quietly getting stuck.

Scheduler controls

>_Run all >_Pause >_Resume >_Stop >_Archive lanes >_Reset metrics
aigis_workbench — batch_process · agent_patch_scheduler
AIGIS Workbench — Batch Process running an unattended governed patch pipeline: active trajectories, tasks in queue, burn ledger, and applied-safe versus held-for-QC counts.

Real AIGIS Workbench UI — Batch Process running the Agent Patch Scheduler · scroll to explore

01Compose & lintAuthor a task; lint it against a hard gate.
02Route to laneAutomatic lane routing; grounded from repo.
03Agent runCodex or Claude produces a patch.
04AssessScope, self-mod, side-effects, integration.
05GateZero-cost gates run; failures are held.
06ApplyAtomic, operator-reviewed, clean-or-revert.
07Git commitCategorized into clean history.
>_

Task Composer & lint gate

The front door: author a task and lint it against a hard gate before it can enter the pipeline. Lint-green in, or it does not run.

>_

Burn Ledger

Live concurrency and burn-safety caps, a burn-fuse, repairs per session, rework ratio, and a session cap — with real token capture: cost per patch and burn rate in actual dollars, projected against the cap.

>_

Per-lane observability

A live strip per lane (task / phase / elapsed / policy), and a per-task drawer with the assessment ribbon, file-by-file colorized diff, clickable changed-file paths, and the exact agent prompt sent.

>_

Governance overlay

Every action publishes to one live SSE stream: the deterministic zero-cost gates, operator decisions, and inline live validation screenshots — one place to watch it all.

>_

Granular failure recovery

Distinguishes connection-lost, provider-error, and usage-limit, then offers the matched action — Continue, Rerun-after-reset, or Resolve — so unattended runs stay trustworthy.

>_

One-decision invariant

A skipped or orphaned gate surfaces exactly one operator decision — never a silent hang. Copy-to-clipboard on every brief and debrief.

Deterministic, zero-cost gates

proof-collision keep-green boundary-crossing atomic clean-or-revert apply
>_

Real dollar token capture in the Burn Ledger is a committed day-one launch feature, being finalized. Where a figure is an estimate, it is labeled an estimate. No burn number on this site is fabricated.

Hands-on modes

When you want to drive it yourself.

Not every task belongs in an unattended queue. Two more surfaces put you in the seat — with the same gates underneath.

Manual · the ad-hoc cockpit

One prompt, one lane, one clean apply.

A Prompt Generator turns a goal into a grounded prompt and runs a single lane. Lane placeholders show live state — Active or Ready — with Prepare, Run, and Remove. The Apply Queue uses tri-state "main clean / dirty" gating: apply is blocked while the working tree is dirty. Batch submissions and job queue are one glance away, per repo.

Apply is gated on a clean tree — never a surprise write.

Repo · git as the safety net

A fluid single-agent session.

A repositories registry with typed repo maps — index, architecture, source-of-truth, changelog — compiles into an AIGIS.md with a managed pointer block. Open a live conversational session: choose Codex or Claude, model and reasoning level, read-only or workspace-write. Work on a branch with checkpoint, discard, and change-set review. Both providers share one durable transcript, with live token streaming and honest stream diagnostics.

Durable cross-provider transcript — one shared history.

Context & queue

Feed agents the right context — and only grounded work.

Memory

Structure-aware, operator-approved.

A context compiler previews a token-budgeted retrieval with a why-reason per slice — non-destructive, showing exactly what an agent would receive. Memory cards cover user preference, project, continuity, rule, and decision; pinned cards are always included, and secrets are stripped on save. The capture queue is operator-approved and never auto-promotes.

Non-destructive preview. Nothing is remembered without approval.

Backlog

Drop a list — or let it propose.

The Backlog Composer turns each line of a list into an independent lint-green task in its own lane, auto-queued, with a persistent "N tasks queued → open Batch" readout that survives refresh. The Backlog Proposer reads the repo, maps, and memory and proposes ranked, evidence-cited work; approve routes it through the same pipeline, dismiss drops it.

Honesty floor: only grounded, cited proposals — nothing fabricated.

Governance controls

You set the rules. The workbench enforces them.

Standards

Enforced and advisory, honestly separated.

Two standard sets — General good engineering and AIGIS engineering — as a per-rule corpus with toggles; toggling a rule off omits it from the compiled AIGIS.md. Enforced gates surface live and stay separated from advisory rules, with a drift readout for what has changed since the maps were written. The Trust Ledger tracks per-lane trust level and record (applies / holds / fails), proposes graduation to the review policy, and supports drop-hard-on-failure and pin-to-enforce.

Enforced ≠ advisory — the difference is always visible.

Settings

Route every role, cap every run.

Per-role model routing across Tier 1/2/3, Bugfix, Reviewer, Prompt generator, Documenter, and Backlog Proposer — each Codex or Claude, each with its own reasoning and effort. Set concurrency and burn-safety caps, and a default review policy: enforce, agent, or bypass.

Provider abstraction across every role, including a read-only cross-provider reviewer.

What makes it different

Not another chat wrapper.

AIGIS Workbench does not replace Codex, Claude, or your editor. It governs the engineering process around them — and shows its work.

Raw AI coding
  • Standards the agent can ignore
  • One provider, one opinion
  • Silent stalls mid-run
  • Cost you find out about later
  • Confident, unverifiable metrics
AIGIS Workbench
  • Machine-enforced gates
  • Two provider families, adjudicated
  • One decision, never a silent hang
  • Real cost, or labeled an estimate
  • No fabricated numbers, anywhere
  • Read-only reviewers

Bring your own provider

Bring your own AI provider.

AIGIS Workbench does not include third-party model usage. Use your own Codex, Claude, OpenAI, Anthropic, or supported provider access where available. AIGIS sells the engineering control layer, not model credits.

>_

Provider availability, API access, subscriptions, and usage costs depend on each third-party provider. AIGIS Workbench pricing does not include those costs unless explicitly stated in the future.

Coming next

On the roadmap — not shipping day one.

Directions we are building toward. These are not part of the launch feature set and are not implied as available.

>_Preliminary pricing. Not final. Planned tool pricing — shown to communicate the intended model.

Planned tool pricing

A pro-consumer model — own it or subscribe.

AIGIS Tools pricing is still being shaped. The goal is a light free version, paid yearly editions users can keep, and optional subscriptions for people who want continuous updates.

Planned · Free

Workbench Lite

For trying the AIGIS engineering workflow

Free
planned free tier
  • 1 local repo
  • 1 active feature lane
  • Manual lane workflow
  • Basic patch receipts & apply queue
  • Bring your own provider
Join Early Access
Preliminary

Workbench Studio

Power users, technical founders & agencies

$499/yr edition
planned · or $79/mo
  • More lanes · multi-repo workflows
  • Standards compiler
  • Reusable lane templates
  • Advanced validation profiles
  • Searchable receipt archive · audit exports
Contact AIGIS
Future · Preliminary

Workbench Team

Small teams that want shared AI-dev standards

Planned
team pricing
  • Shared lanes & standards
  • Reviewer roles
  • Shared apply governance
  • Admin controls & policy packs
  • Audit visibility
Ask About Teams
Tool pricing is not final and may change before public launch. Early users may receive founder pricing. AIGIS Tools do not include Codex, Claude, OpenAI, Anthropic, or third-party model usage — you are responsible for any provider subscriptions, API usage, limits, terms, and costs from the AI provider you choose.