In development Community beta target Q1 2027 View roadmap

Architecture · technical baseline

A distributed worker system with a durable central coordination layer.

The architecture baseline separates control, placement and durable state from worker-side agent execution. Components and protocols described here are proposed unless explicitly labeled as source-verified.

Evidence boundary: Repository source for two upstream foundations has been audited, but ANDIP implementation phases remain planned. Diagrams below are architecture illustrations, not a live deployment.
Full platform view

Control-plane decisions; worker-side execution.

API, workflow, policy and scheduler services coordinate registered VPS workers. Durable state and artifacts support observable, recoverable attempts.

The proposed ANDIP platform includes a central control plane, several worker VPS nodes, durable task state, artifact storage and audit/monitoring signals.
Full platform architecture. The system keeps execution distributed while centralizing policy, scheduling and workflow state.

Why a central control plane?

It provides one place to validate requests, apply project policy, rank eligible workers and track task state. The control plane coordinates; it does not turn the platform into a single agent process or live machine monitor.

Why distributed workers?

Agent processes need compute, storage, browser and network resources near the selected execution environment. Workers report capacity and execute bounded attempts while the control plane retains the durable record.

Control & execution

Separate decisions from local runtime duties.

Workers use an outbound connection pattern in the MVP design. A lease identifies the authorized attempt and fencing generation.

Control plane policy and scheduler issue a fenced lease to a VPS worker, which prepares an isolated runtime, emits heartbeat events and uploads artifacts.
Control plane and worker communication. The control plane stores execution truth; worker-side processes must reconcile after disconnects.

Central control plane

Gateway/authentication, project and task management, workflow orchestration, policy, budget reservation, scheduler, VPS registry, state and event interfaces.

Worker runtime

Enrollment identity, resource probes, task claim, sandbox preparation, adapter launch, heartbeat, checkpoint/artifact reporting and idempotent cleanup.

Proposed state backbone

PostgreSQL is the planned source of truth for task state, leases and outbox events. Artifact payloads use separate S3-compatible storage. Redis or NATS should be introduced only if queue/outbox benchmarks justify them.

Transport

HTTPS long-poll plus heartbeat is the initial protocol direction. Workers make outbound TLS connections; public inbound worker ports are not required for this MVP approach.

Lifecycle & placement

Each state transition leaves an explicit trace.

Requests become durable tasks, capacity-aware assignments, isolated attempts and checked artifacts.

Seven-stage task lifecycle: request, plan, schedule, execute, monitor, verify and complete, with policy and budget checks before assignment.
Agent execution lifecycle. A failed attempt may enter controlled recovery; terminal state and independent verification govern completion.
Scheduler filters candidate workers, reserves capacity and issues a fenced lease. Coding, research and browser workloads may land on different VPS nodes; unmatched work remains queued.
Workload distribution across VPS. Placement considers compatibility, capacity, policy, locality, browser slots and budget rather than equally dividing tasks.
Failure handling

Recover only when the outcome is safe to recover.

Heartbeat gaps make a worker suspect; lease expiry moves its attempt into reconciliation. A newer fencing generation prevents a late worker from overwriting current state.

Pure or idempotent tasks may be retried. Compatible checkpoints require version and checksum validation. External mutations with uncertain outcomes are held for provider reconciliation or human review—not blindly replayed.

Exactly-once external side effects are not guaranteed unless the external system supplies idempotency or reconciliation support.

Heartbeat or lease loss enters reconciliation; safe idempotent work can resume, while uncertain external side effects are held for human review.
Failure detection and recovery. Retry policy distinguishes idempotent work from uncertain external mutation.
Source boundaries

Reusable components are not the distributed platform.

ANDIP preserves upstream projects as independent foundations and adds adapter contracts plus newly engineered orchestration infrastructure.

Agent Orchestrator and Jev Ultrafast are reusable upstream foundations. ANDIP integration interfaces connect them to new distributed scheduling and worker infrastructure.
Open-source component integration. Agent Orchestrator supports local session and coding patterns; Jev is an optional browser adapter.

Verified source

Agent Orchestrator source includes local lifecycle, adapter, runtime, worktree and HTTP/event patterns. Jev Ultrafast source includes a browser agent loop, DOM snapshot and guarded actions; its audit reported 31 tests and lint passing.

Verified source

Planned ANDIP components

Distributed scheduler, VPS registry, durable lease queue, multi-worker runtime, scoped secrets, centralized policy/budget, artifact contract, recovery and fleet monitoring require additional engineering.

Planned

Engineering assumptions

PostgreSQL lease/outbox design, rootless container execution, S3-compatible artifact storage and HTTPS long-poll worker protocol are architecture choices subject to implementation benchmarks and security validation.

Assumption

Features awaiting validation

Cross-VPS capacity placement, durable recovery, isolation, browser profiles, cost metering and the 100-agent workload target need reproducible tests. No ANDIP implementation completion is inferred from upstream source.

Under validation
Technical decisions

Keep the first system understandable—and measurable.

Why Kubernetes is not required for the initial MVP

The initial design registers already-provisioned Linux VPS and runs a worker service. At this stage, Kubernetes would introduce cluster operations and another control plane without demonstrated need for autoscaling or managed provisioning. Revisit only when test evidence and deployment needs justify it.

Agent supervision is not distributed orchestration

A local harness can create sessions, launch an agent and observe its process. Distributed orchestration must also own worker identity, fleet capacity, leases, queue durability, reservation, policy, artifact persistence and safe recovery across servers.

Architecture baseline source: AGENT_NATIVE_INFRASTRUCTURE_MASTER_WORKFLOW.md, audited 2026-10-08. All operational interfaces remain subject to engineering implementation and change.

Upstream repositories

Attribution and licensing remain with their projects.