In development Community beta target Q1 2027 View roadmap
Architecture specification · not released product behavior

Monitoring & Cost Management

Operational signals, usage ledger, budgets, resource reservation and honest measurement of concurrency.

Architecture baseline · subject to change

Signals worth observing

The proposed operational view includes worker heartbeat, CPU/RAM/disk pressure, queue delay, task failure rates, lease expiry, API latency, budget usage, orphan environments and browser blocked events. Thresholds in the architecture are starting targets and need tuning against measured baselines.

Usage ledger

Usage can be associated with project, task attempt, agent definition and worker: CPU-seconds, RAM-seconds, disk time, network egress, browser time, model tokens/API calls and artifact storage. Estimated and actual cost should be kept distinct.

Reserve, meter and enforce

Reserve estimated cost and resources before scheduling, meter actual usage, settle completed attempts and return unused reservation. Soft limits may warn or throttle; hard limits block new assignments. Only safe-to-stop workloads should be automatically cancelled; external mutations may need controlled pause.

Operational status, not decorative metrics

A future operations view should show state, reason code, age and data freshness for fleet, queue, attempt, costs, artifacts, approvals and audit events. A stale worker or disconnected signal must not be presented as current health.

100-agent evidence

Separate deterministic mock work, lightweight model calls, mixed coding/research/browser workloads and external-API workloads. Record source/image digests, fleet inventory, workload seed, configuration, timestamps, raw metrics, queued count and true RUNNING count. The target requires repeated runs and failure injection.

Measurement status: No performance, success-rate, recovery-time or cost-efficiency result is asserted by this architecture website.

All documentation · Full architecture page · Roadmap