Skip to content

Architecture

0 is an open cybersecurity harness built on one rule: reproduce before trusting.

For practical setup, start with Scan Workflows, Console, or Research Workflows. This page describes the engine’s structure.

The harness has several workflow-specific evidence paths:

  • 0 orchestrates source and live-target assessments, package audits, AI/agent evaluations and research workflows. Parallel discovery, verification and fix generation depend on the selected command and configuration.
  • 0verse is the separate in-repo compiled-program evidence producer. Its explicit integration is opt-in, not a stage of every scan.

The diagram below describes the research-adapter contract, not the execution order of all CLI commands.

In that research contract, discover and verify are mandatory stages. Unsupported optional stages are marked skipped; failed or unavailable proof stays inconclusive. A mandatory verification stage means a result is evaluated, not that it necessarily reproduces. Discovery output cannot promote itself.

Research evidence envelopes track independent dimensions:

DimensionMeaning
Proof gradecandidate → reachable → observed → reproduced → impact-proven
Noveltyunchecked, novel, duplicate, or inconclusive
Execution privilegeLinux zero-cap, Windows windows-restricted, privileged, or unknown — with an evidence basis
ProvenanceTarget/build/config identity, producer, model, attempt, run
Native evidenceProtocol attempts, sanitizer crashes, hunt records, VM proofs

A reproduced result still has to clear more gates before it can be disclosed. Disclosure requires reproduced evidence plus a real novelty receipt; scope, publishability, redaction, and operator policy are all separate downstream gates. A proof grade never implies attacker privilege.

Privilege claims carry their own attestation gates, and all of them fail closed:

  • Linux zero-cap — runtime attestation of non-root real and effective UIDs, an all-zero effective capability set, no_new_privs (so a later exec can’t regain privilege), and a digest-bound artifact.
  • Windows LPE — a separate token-transition gate; Linux UID/cap facts are never treated as Windows proof. Needs a Windows context, a retained token-attestation artifact and receipt, exact Canary build/campaign/worker/ manifest binding, ≥2 target trials and ≥2 clean controls with distinct captures, a low/non-elevated start token, and a distinct high-integrity admin or LocalSystem finish token. Benchmark rows and automatic disclosure fail closed; human review stays mandatory.
  • Linux kernel N-of-K — every reproduced boot supplies a bound receipt, and the aggregate manifest is the evidence artifact. It binds each dmesg digest and rejects mixed kernels or repeated boot IDs.

These are trusted-orchestrator bindings (VM kernel, guest init, launcher), not hardware-backed attestation of a hostile worker. Human review remains mandatory. See Research Workflows for the research contracts. The deterministic replay JSON in Verification Results is a separate contract, not the schema for every research envelope.

AdapterNative proofStatus
HTTP protocol conformanceConcrete request/response + deterministic RFC oracleConnected
Userspace memory safetyFuzz loop, sanitizer/Miri crash, saved input, primitive classificationConnected
Agentic huntBest-of-N records, judge, skeptic/prover, native novelty resultConnected
Linux kernel reproducerFresh-boot N-of-K signature gate with dmesg bindingConnected
Linux external boot matrix importVersioned vulnerable/patched manifest, unique boot markers, clean-control gate, hashed logsConnected; explicitly external provenance
Windows Hyper-V evidence importBuild/campaign/worker-bound 0verse receipt, clean controls, repeated crash signature, retained dumpsConnected for crash reproduction; LPE disclosure fail-closed until token attestation
Mobile static intakeTyped candidates + scoped downstream handoff; no passive promotionConnected
XNU IOKitSelector discovery, reachability hints, deterministic programs; panic promotion offPartial, fail-closed
Unified web/AI/source/package/on-chain pipelineNative findings wrapped without rerunning the pipelineConnected

0verse is an in-repo Python evidence producer (the 0verse/ directory), not an @0/* package. It handles compiled-program evidence with its own Ghidra, angr, AFL++, PoV, and notary contracts. 0 consumes only explicit, versioned interfaces — the opt-in 0verse binary/NDJSON contract and verified external receipts. It never bundles 0verse into @0/*, schedules it as a generic scan worker, or promotes a hypothesis without the matching proof gate.

A shared differential runner can run identical input against two versions, builds, configs, or implementations; a failed side is inconclusive, never divergent. Novelty providers are pluggable per ecosystem, and zero checked records can never produce a novel verdict.

Research CLI paths emit evidence envelopes. Envelope availability depends on the producing path; do not assume a legacy finding or a different command’s output has the same schema. Managed storage and access are outside this repository.

Two import paths handle kernel proofs the generic VM runner can’t safely rebuild:

  • 0 research linux-matrix imports externally executed boots. The versioned manifest binds build IDs, literal crash/completion oracles, per-boot markers, thresholds, and log paths; 0 hashes the manifest, every log, and its verdict. The envelope says executionOrigin: external and never claims 0 ran the boots.
  • 0 research linux runs natively, bound to a required literal crash oracle (--expected-signature). A different KASAN/oops/GPF is recorded but can’t satisfy the N-boot gate. Each boot contributes its own hashed dmesg artifact, so a 2-of-3 claim carries the full three-boot audit trail.

Execution ownership and engine connections

Section titled “Execution ownership and engine connections”

CLI, browser, schedules, and MCP use the shared workflow runner. Assessment shortcuts retain a one-step run; fix, verification, research, and deep review use their own typed executors. An individual target tool still executes through the tool executor. Calling an HTTP tool does not start a full security assessment.

A workflow defines steps. A run captures the definition, bound inputs, execution status, and returned evidence. The core runner executes connected steps sequentially, applies workflow and step deadlines, and shares provider spend. Finding severity and proof status are independent of whether execution completed. The control database retains workflow snapshots and combined results; filesystem artifacts remain on the engine that produced them.

An engine connection chooses the process that owns sessions, runs, approvals, workspaces, provider connections, and execution environments. Model selection chooses a provider route within that engine. Workspace paths belong to that engine’s filesystem. Connecting to a server does not copy a laptop repository or send the laptop’s model credentials to it.

A remote engine runs the same services on another machine and owns that machine’s local or SmolVM execution boundaries. A remote executor would let one engine outsource tool execution while retaining ownership of the run. These are separate capabilities. The engine connection protocol does not provide a generic remote shell, arbitrary executor registration, or qualification of another sandbox. The existing runner and SmolVM contracts still apply on the selected engine.

The browser uses a trusted same-origin connection proxy. Registered descriptors contain capability and connection metadata; remote bearer secrets remain in the local connection service, and provider credentials remain on the execution engine. Resources and approvals are bound to their selected engine. A connection failure is a disconnected state, not proof that the engine stopped its work. Reconnecting reads the owning engine’s retained state.

See Engine Connections for registration, SSH tunnel setup, remote workflow commands, and the current lifecycle boundaries.

For web pentesting the agent is shell-first: bash (curl, python3, sqlmap, …) is the primary tool, not a fixed set of HTTP tools. LLM and code targets get specialized tools like send_prompt and read_file. The agentic scan path applies triage and a separate verification pass; source, template, replay and research paths have different evidence contracts. Reports can retain candidates without independent runtime proof.

The six labels below are a conceptual workflow, not a promise of six separate agents or exactly two sessions. Native agentic verification uses one bounded session per candidate; optional discovery fan-out adds other sessions.

Plan -> Discover -> Attack -> Triage -> Verify -> Report

1. Research agent (Plan + Discover + Attack + PoC)

Section titled “1. Research agent (Plan + Discover + Attack + PoC)”

The research session is instructed to:

  1. Plan — choose likely vulnerability classes and prioritize vectors.
  2. Discover — map endpoints, detect models, fingerprint technology and inspect source when provided.
  3. Attack — test hypotheses with the target-specific tools.
  4. Record evidence and PoCs — save the observed request/response or source path and, where applicable, executable reproduction steps.

These are model tasks, not a guarantee that every target gets comprehensive coverage or every saved finding includes a working PoC. Deterministic recon and specialized pipeline stages can also contribute findings.

When a scope document or challenge description exists, it’s passed to the agent as context — the same way a real pentester receives a brief.

Tool set depends on target type:

  • Web: shell-first instructions favor bash, save_finding and done. Role tool sets can also include crawl, submit_form, http_request and other gated capabilities; browser depends on browser availability.
  • LLM: send_prompt plus the applicable network-role tools.
  • Scoped source/npm review: read_file, literal search_files, list_files and finding tools. Broader execution tools depend on the caller’s tool profile; a scoped read-only review is not the full console tool set.

Budget-aware reflection. As the turn budget is consumed, the loop can inject continue prompts so the agent does not spend every turn on one dead end. The budget depends on the workflow, role, and depth; it is not one universal 40-turn setting. See Agent Loop and Budget Management.

The agentic scanner applies default-on holding-it-wrong and evidence-completeness filters, category oracles, and optional source/PoV/consensus gates. Routing, tooling and category determine which checks run; several checks make live requests or model calls. Guarded heuristic rejection can hold protected findings for verification. See Finding Triage for actual defaults, evidence limits and per-finding provenance.

EGATS caveat. The 2026-04-11 ablation found egatsTreeSearch regresses solve rate on hard challenges at ~10× the cost of the next-worst layer. It’s removed from the default moat aliases and opt-in only (0#116). Results varied by slice. npm-bench attribution needs repeated runs; routing research remains separate.

Native agentic verification receives one finding with request/response excerpts, original analysis, target/authentication and optional review-memory context, with five turns per finding. It is separate from the original full conversation, but not blind to all reasoning. Structured consensus is a tool-free model assessment; template-scan fallback can confirm without replaying. See Blind Verification for these distinct contracts.

Reports retain findings and their evidence/verification state; consumers must not assume every row has the same proof grade. Formats vary by command and include terminal, HTML, PDF, SARIF, Markdown, and JSON. A local report is not automatically published or assigned a public share URL.

Different models can serve discovery, review, refutation and fix roles when the workflow and configured route support them. Multiple sessions do not imply different providers or statistically independent errors. API child runtimes keep the parent provider, credential and base URL; a role-model name must work on that route rather than triggering automatic cross-provider credential discovery.

Opt-in Jev assistance handles bounded browser navigation, memory relevance, duplicate assessment and red-team attempt feedback. It does not authorize tools or verify an exploit. Browser decisions return to the main model on ambiguity; memory context cannot reject a finding; dedupe preserves original evidence; red-team break detection remains with its judge/oracle. See Features for explicit data-egress consent, provider settings and per-evaluator budgets.

Executable plugins provide versioned guest execution, composition, next-call activation, source evolution and rollback. The wired research-preview live-harness path adds agent.driver and ui.view replacement during a task, with console/tool and TUI host integration. Its local measurements cover specific lifecycle paths, not universal security improvement or hosted end-to-end qualification.

A generation graph declares providers and dependencies. The runtime owns preparation, migration, activation, and resource disposal. Browser, TUI, and desktop consume its shared catalog.

  • Sandboxed: each activate, driver, view, or dispose phase invokes the pinned executable source in a fresh guest through the authorized broker. LiveHarnessHost holds JSON state between phases.
  • Workspace-trusted: a separate canonical-workspace grant permits in-process ESM with Cordis-owned providers and effects. This grants host privileges; self-extension alone grants no host trust.

After an engine restart, source/spec artifacts and trust files remain. Active providers, JSON state, resources, and in-flight work require explicit reconstruction. Guests are fresh per phase.

The lifecycle is inspired by Cordis’s reversible effects and dependency management and DSH’s service composition. Component-owned effects define the cleanup boundary. External requests already sent remain effective; quality improvements require separate measurement.

See Improvement Plane for the exact current/planned distinction, Python support, autonomy settings, and long-horizon recovery requirements.

Every UI and output surface consumes a renderer-neutral document or event rather than another renderer’s terminal text. The versioned contract is 0.presentation/v1: reports retain their existing schemas, while interactive sessions use typed transcript entries and live producers emit ordered semantic events with a local source, sequence, timestamp, type, payload, and optional scan or session correlation.

The native OpenTUI console, plain terminal output, report formatters, and the browser dashboard are adapters over that contract. The dashboard exposes its same-origin live feed at GET /api/v1/presentation/events as Server-Sent Events; Last-Event-ID resumes only persisted events after the supplied timestamp/ID cursor. A producer’s sequence is monotonic only within that producer, so consumers must not infer global ordering or exactly-once delivery across processes.

The terminal remains authoritative for an interactive console. While OpenTUI owns stdout, direct process writes are captured as semantic records instead of being allowed to corrupt the renderer frame; the original stream is restored when the console exits.

Desktop is a development-only alpha, excluded from CLI releases. The native application shell uses the same local control plane — see Roadmap for current status. Use the CLI Console for the released terminal interface.

ModeTargetWhat it does
deepLLM API URLMulti-turn prompt injection, jailbreak, tool poisoning, and exfiltration investigation
probeLLM API URLLightweight surface scan of an LLM API
webWeb app URLCORS, headers, exposed files, SSRF, XSS, path traversal, fingerprinting
mcpMCP serverTool poisoning, schema abuse, permission escalation
http_auditAuthenticated HTTP targetWorker-oriented scoped web assessment using ZERO_TARGET_* configuration

Package audits and source reviews use separate audit and review commands, not scan --mode audit or scan --mode review. Mode inference and explicit options are documented in Configuration.

0 decouples the pipeline from the LLM backend. Each adapter implements one interface over a different provider:

AdapterBackendHow
LlmApiRuntimeConfigured API providerDirect provider requests and native tool calls
ProcessRuntimeClaude, Codex or Gemini CLISubprocess adapter; distinct from a tool sandbox
CliNativeRuntimeSupported installed coding CLICLI-native session execution
OpenRouterRuntimeOpenRouterSeparate adapter with ensemble support
OllamaRuntimeOllamaLocal model inference

createRuntime(config) constructs API, process or Ollama runtimes from config.type; registry helpers and CLI entry points perform their own selection. auto is selection policy, not an AutoRuntime class. MCP target/client/server integration is separate from the model-runtime factory. See Configuration and API Keys for path-specific resolution and fallback.

The LLM adapter above is not an execution sandbox. Keep three separate choices:

LayerResponsibilityCurrent choice
ControllerProvider authentication, scope, budgets, approvals, evidence and version selection0’s TypeScript harness
Toolbox artifactFilesystem containing tools, runtimes and dependenciesOCI image, provisioned before execution
Execution engineHost/guest boundary, mounts, networking, resource limits and teardownNative Apple Silicon SmolVM workbench; Docker or Linux SmolVM for separately configured offline workers

OCI does not mean Docker execution. It is the open image format that both Docker and smolvm consume. Smolvm boots the image in a microVM with a separate guest kernel; Docker containers share their execution host’s kernel. A local archive avoids a registry or Docker daemon during smolvm execution.

Upstream smolvm also supports unpacked root filesystems, Smolfiles and packed .smolmachine artifacts. Those are viable upstream provisioning options, not operator image inputs accepted by 0’s workbench. Workbench setup approves a local archive by SHA-256 and stores it in private operator state; execution stages an independently verified private copy. A hand-maintained mutable VM or a model-selected registry tag is not a substitute.

  • Small runtime image: useful for bounded Node source-evolution fixtures. It is not the pentest toolbox.
  • Toolbox image: the Dockerfile’s toolbox target contains the declared static-analysis, web-testing and identity tools without the 0 application.
  • Distribution image: the runtime target adds the bundled CLI to that same toolbox. This remains the default Dockerfile output.

The toolbox is an explicit inventory, not a promise to contain every security tool. Optional wordlists, browsers, privileged networking and specialized kernel/binary environments need their own provisioning and qualification. Installation of a network tool does not authorize or enable target access. See Improvement Plane for local provisioning and execution checks.

The smolvm execution profile starts the whole 0 CLI inside an online Linux microVM, including its shells, browser, agents and GitHub tools. This is an explicit operator profile, configured through workbench setup/settings, not a replacement of only the bash tool. Neither launch failure nor missing image, runtime or provider prerequisites falls back to host or Docker execution.

On Apple Silicon, setup checksum-verifies the complete pinned SmolVM 1.14.6 release, verifies its existing code signature and Hypervisor entitlement, and preserves the bundled rootfs, libraries and disk templates. It does not re-sign the binary or change shell startup files. Docker may build/export OCI artifacts; no Docker daemon, Colima VM or nested KVM is required to execute this profile.

The main guest receives only the explicitly selected writable workspace, private guest HOME/state, immutable admission marker and selected integration grants. Its non-root numeric UID/GID matches the operator’s host ownership so virtiofs writes do not require changing workspace permissions. Universal host HOME, SSH agents and Docker sockets are not exposed. The configured online workbench is a trusted engagement environment: commands inside it can access its deliberately granted credentials. A VM boundary does not isolate those commands from one another or authorize targets; the existing scope/approval controls still apply.

Generated executable plugins and evolution programs instead use fresh sibling microVMs through a bounded host-supervised handoff broker. Their source/control mounts are read-only; their writable workspace is guest-owned; networking and workbench credentials are absent. The separate HTTP replay profile accepts only the broker’s validated public HTTP request contract and explicit network grant. Action images require an operator-approved immutable catalog entry; an unknown image is refused rather than silently replaced. Admission authority is the fixed read-only virtiofs marker and dedicated broker mount, not a forgeable environment flag or a file in the checkout.

The Darwin supervisor uses kernel process information, exact per-run ownership capabilities, process-start identities and kqueue lifetime events. It never signals process-name guesses or an unverified PID/process group. A kernel-held admission lock covers reservation, execution, teardown and lease removal. Cancellation and abrupt controller death stop the complete owned VM family; teardown proof includes sibling controllers and VMMs. Unconfirmed teardown keeps the admission lease and recovery state instead of claiming capacity was freed. TTY execution inherits terminal stdio and restores terminal settings at teardown; non-TTY execution preserves argv, stdin and separate stdout/stderr.

The complete Kali/Bun/0/Chromium archive was qualified natively with four vCPUs, 4096 MiB RAM and 20 GiB guest storage. Its imported root filesystem uses about 6.4 GiB; image import needs additional temporary headroom. An 8 GiB guest import exhausted its deadline, so small-runtime fixture budgets are not evidence that the complete workbench image fits. Sibling disk budgets are selected by the trusted host profile, never by a guest request.

E2B can remain a separate cloud placement option. Local qualification does not establish equivalent cloud configuration, isolation or operational behavior.

Codex did migrate from TypeScript to Rust. OpenAI’s May 2025 announcement names four motivations: installation without a Node prerequisite, native security bindings, lower memory consumption without runtime garbage collection, and an extensible wire protocol. These are engineering goals, not evidence that changing languages improves security findings or model reasoning.

The current TypeScript Codex SDK still provides a TypeScript integration surface by spawning the native CLI and exchanging JSONL events. The useful lesson for 0 is a stable protocol boundary between clients, orchestration and execution—not that every component must be rewritten together.

Decision: retain the TypeScript harness; evaluate a small native execution supervisor before considering a full rewrite. A separate process with a versioned protocol is preferable to placing new native failure modes directly inside the credential-bearing controller. Candidate responsibilities are process ownership, PTYs, OS confinement, resource accounting and confirmed teardown. Rust does not itself provide any of those isolation guarantees.

Before committing to a port, measure representative CLI startup, peak memory, controller CPU, worker preparation, cancellation latency and sustained output handling. Separate provider wait time and image import from controller overhead. Require behavior parity for scope, credentials, sessions, evidence and rollback, then demonstrate a measured improvement or a concrete OS capability the current implementation lacks. No 0-versus-Rust performance benchmark has established that a full harness rewrite is currently warranted.

0 integrates with MCP in three roles:

  • As a client — selected agent/console paths connect to external MCP servers to expose their tools. The model still uses its configured inference runtime.
  • As a server — 0 mcp-server exposes a scoped subset of tools over stdio to an external host. --tools is an allowlist, not a capability grant: every exposed tool still runs through 0’s execution and engagement guards, and 0 keeps ownership of scope, rate limiting, persistence, and verifier state. External hosts (DSH, Codex, Claude Code) are optional clients; they don’t replace the native scan loop. See Improvement Plane for the separate future-worker promotion boundary.
  • As a scan target — --mode mcp probes servers for tool poisoning, schema abuse, and permission escalation.
Terminal window
0 mcp-server \
--target https://example.com \
--scan-id engagement-001 \
--scope ./scope.json \
--tools http_request,crawl,send_prompt,submit_form

Two execution surfaces, one public documentation home:

  • 0 CLI — local runs, CI, replay, exports, and console.
  • Managed control plane — a separately operated engagement layer (not in this repo) that is still in development. See Roadmap for status.

Fresh CLI scan/review/audit workflows normally allocate a run-local ~/.0/runs/<scan-id>/state.db, journal and report; explicit database paths, resume, console and SDK callers have different storage choices. The dashboard can inspect a selected database via --db-path, not every worker database at once. Managed persistence and organization ownership are separate service concerns.

The shell-first web workflow uses bash, save_finding, and done. The agent can compose commands such as curl -c cookies.txt … | jq. Structured tools remain available.

See Research for the rationale and Benchmark for results.

ToolUsed inPurpose
bashWeb, LLM, VerifyShell commands subject to tool, scope, and engagement restrictions; host execution by default.
browserWebPlaywright headless browser for XSS and JS-rendered pages.
save_findingFinding-producing rolesRecord a candidate and evidence; does not independently verify it.
doneAllSignal completion.
send_promptLLMSend prompts to AI/LLM apps.
read_fileSource, npmRead source for code review.
run_commandProfiles that expose local executionRun an allowlisted command on the host (not a sandbox); not part of every scoped read-only review.
list_filesSource, npmEnumerate a directory.
search_filesScoped source, npmLiteral search in regular text files, subject to scope and size limits.
crawl / submit_form / http_requestNetwork rolesStructured HTTP alongside shell-first workflows.