Agent state machines¶
An agent state machine is a declarative, human-editable, machine-parseable program whose building blocks are agent6 runs, sandboxed tool calls, timed waits, and branches. It lets an operator compose small deterministic agents that run for a long time, and agent6 is the runner.
This document specifies the format and its runtime.
The runtime lives under src/agent6/machine/ (the engine and the format), with the lifecycles around it in src/agent6/app/machine*.
The agent6 machine subcommands drive it: list, create, check, test, graph, run, status, poke, stop, and replay (CLI surface).
It changes neither the security model nor the tool surface AGENTS.md binds; Security considerations records how each invariant holds.
1. Motivation¶
run and review are single-shot.
A machine expresses the long-running shape: timed polling, branches on agent output, side-effecting steps, and terminal states.
The operator authors the flow as a static graph, and the LLM works inside one state at a time.
The same inputs take the same path, and a crash loses no state: the snapshot-and-replay posture run has internally, one layer up.
2. Non-goals¶
- Not a general programming language
- the branch/predicate grammar is non-Turing-complete: no loops inside a predicate, no arbitrary code
- loops exist only as graph edges
- Not a distributed scheduler
- one supervisor process per machine (systemd / cron-friendly), each
agentstate a child process; restartable - one active state at a time; compose by running independent machines
- one supervisor process per machine (systemd / cron-friendly), each
- Not a new network surface
- anything that talks to the outside world is a tool, gated by the existing audit rules
- Not run unreviewed
machine createmay draft a machine; in a repositorymachine runrefuses the draft until the operator commits it
3. Design principles¶
- Control flow is static and operator-owned; work is dynamic and model-owned
- the graph of states/edges is fixed at author time
- inside an
agentstate runs the usual agent6 loop
- Authored in a text editor
- TOML: diff-friendly, commentable
- a state can be an agent6 run, so mini-agents are wired together rather than written in Python
- Everything nondeterministic is journaled as a fact
- wall-clock reads, tool stdout, agent outputs: appended to an immutable event log once observed and reduced
- the engine is a pure reducer over
(machine, blackboard, event) → blackboard' - replay reads the journal instead of re-observing the world: a run backtests offline
- Fail loudly (repo convention)
- one file parses to one validated machine or a precise error
- a missing transition target, an unreachable state, a blackboard type mismatch, an unknown key: load-time errors
- No implicit defaults (mirrors
Config:extra="forbid", frozen=True)- every variable declares a type and an explicit initial value (
valuefor[vars.operator],defaultfor mutable[vars.code]/[vars.agent]) - every state declares every outcome edge it can produce
- every variable declares a type and an explicit initial value (
4. The format¶
A machine is a single TOML file, suffix .asm.toml ("agent6 state machine").
TOML because the project already standardizes on it and tomllib parses it.
The parsed document is validated by a pydantic v2 model at the trust boundary (extra="forbid", frozen=True), like Config.
Naming.
load_machineaccepts any path; shell completion globs*.asm.toml.
4.1 Top-level shape¶
machine = "item-classifier" # stable id; names the instance dir
version = 1 # schema version; bumped on changes
initial = "poll" # name of the entry state
[budget] # required; max_transitions always binds
max_usd = 25.0 # optional cap on metered spend (Spend bounds, section 9)
max_transitions = 100000 # hard stop on total edges taken (runaway guard)
# The blackboard is three subtables, named by who may write each variable.
# The subtable header is the owner; there is no per-entry discriminator.
[vars.operator] # written at author time; immutable at runtime
inbox_dir = { type = "str", value = "/srv/inbox" }
poll_secs = { type = "int", value = 300 }
[vars.code] # written deterministically by a tool state's capture
pending = { type = "list[str]", default = [] }
cursor = { type = "str", default = "" }
[vars.agent] # written by an agent state's validated finish_session
verdict = { type = "classification", default = {} } # a [schemas.*] record type
[schemas.<name>] # named record types; see 4.6
...
[states.<name>] # one table per state; see 4.3
...
4.2 The blackboard: three owners¶
The key/value store is split into three subtables, named by who may write each variable. The subtable header carries the owner, so who may write a value is checkable at load time.
| subtable | written by | mutability | declared with | example |
|---|---|---|---|---|
[vars.operator] |
the human, at author time | immutable at runtime | value |
inbox_dir, poll_secs, thresholds, an API base |
[vars.code] |
a tool state's capture |
mutable (deterministic) | default |
pending, cursor |
[vars.agent] |
an agent state's validated finish_session payload |
mutable (LLM) | default |
verdict (a [schemas.*] record) |
Only tool states (into [vars.code]) and agent states (into [vars.agent]) ever mutate the blackboard; branch/wait/terminal only route, sleep, or end.
[vars.operator]: the machine's parameters, set once at author/commit time, never written by any state- declared with a concrete
value(not adefault) - a
capture/settargeting an operator var is a load-time error - any JSON-serializable value; the names above are illustrative
- declared with a concrete
[vars.code]: change only as a pure function of journaled tool output (what keeps the path deterministic and replayable)[vars.agent]: change only through the single validated structured output of oneagentstate (the LLM's one sanctioned channel into the blackboard)
machine check enforces the ownership wall.
- a
toolcapture targets only[vars.code]; anagentcapture only[vars.agent];[vars.operator]is read-only to every state - a
toolcannot smuggle a write into an LLM-owned variable; an agent cannot overwrite a deterministic one
Allowed types (all three subtables): str, int, float, bool, list[<scalar>], json, and any named record type declared in [schemas.*] (see Record schemas).
The two structured types differ on one axis, navigability:
jsonis an opaque blob: read or written wholesale only- passable to a tool/agent (
{{ x | json }}) or captured whole, never dotted (x.keyonjsonis a load-time error) - use it only when the machine never inspects the value's internals
- passable to a tool/agent (
- a record type (e.g.
classification) is navigable- every
.fieldread in a predicate or template checks against the schema atmachine checktime - a misspelled field is a load error
- every
Declared types make branch predicates statically type-checkable: scalars by their declared type, record fields by their schema, and json forbidden from being dotted at all.
The blackboard (all three subtables) is the only state that flows between states.
The whole blackboard is snapshotted to disk after every transition; [vars.operator] is fixed for the life of the machine.
4.3 State kinds¶
Every state has a kind.
| kind | what it does | outcome labels (edges) |
|---|---|---|
agent |
runs one agent6 loop (a Workflow) on a prompt |
ok · failed · budget_exhausted · timeout |
tool |
one sandboxed command via run_in_jail |
ok · nonzero · timeout |
wait |
sleeps until a wall-clock tick or an external signal | tick · signal |
branch |
pure predicate over the blackboard → next state | (chooses a goto directly) |
terminal |
ends the machine | (none; absorbing) |
- outcome labels are a fixed enum per kind, produced deterministically by the state executor
- a non-terminal, non-branch state declares
on = { ... }mapping every label its kind can emit to a target; an omitted label is a load error - the edge taken is a pure function of the closed label set
agent¶
[states.classify]
kind = "agent"
model = "inherit" # the configured worker model, or pin one
prompt = """
Classify the item at path {{ cursor }}.
Call finish_session with JSON {label, confidence}.
"""
output_schema = "classification" # a [schemas.*] entry; validates the payload
capture = { finish_json = "verdict" } # payload -> blackboard var `verdict`
timeout_secs = 600
on = { ok = "route", failed = "poll", timeout = "poll", budget_exhausted = "halt" }
# mode = "agent" # "agent" (default, read-only) | "run"
# Optional per-state overrides (inherit the effective config when unset):
# provider = "anthropic" # which [providers.*] entry backs this call
# effort = "high" # off | low | medium | high | xhigh | max
# temperature = 0.2
# max_usd = 1.5 # this agent slice's metered-spend cap
# max_tokens_fallback = 100000 # ...and its unmetered-token cap (-1/0/>0)
An agent state spins up a normal agent6 run: its own snapshot dir, transcript, budget slice, jail.
- its only control-flow signal is the outcome label
- its structured product is the
finish_sessionpayload, validated againstoutput_schema, captured into the blackboard- the state's task states the contract (the schema rendered field by field)
- the loop refuses a non-conforming
finish_sessionwith the problems, so the model retries in-run; the engine's own validation of the recorded payload stays the authority
- the LLM cannot pick the next state; it populates variables a downstream
branchreads
mode chooses the tool surface.
"agent"(default): a read-only, structured-output loop: the read and navigation tools plusfinish_session- withheld:
apply_edit,apply_patch,run_command,run_verify_command,run_metric_command,read_session,fetch,read_backgroundandstop_background
- withheld:
"run": real coding work (edit + verify + commit tools), asagent6 runhas
Where a run state's work lands:
- each run state executes in a fresh clone (the
--parallellane mechanism,[parallel].workdircache) checked out at the machine chain's tip - its commits land per state on the visible
agent6/machine-<id>branch; your checkout is never touched - merge the branch when you want the work (the run's ending names it and the
git mergeline) - a state's clone is removed as it lands, so nothing is left for
sessions pruneto sweep - the chain ref (
refs/agent6/machine-<id>/head) outlives the instance dir- a fresh instance over a leftover chain starts from HEAD once the chain's tip is in HEAD, dropping the ref
- while the tip is not in HEAD it refuses, naming the merge and delete remedies
- a branch left without its ref is an ordinary branch, and the instance starts from HEAD
- a machine with run states works its own tree everywhere:
tooland read-only agent states also run in fresh clones at the chain tip- an edit-then-check loop sees the committed work with no plumbing
- a
toolstate's tree writes are scratch, discarded with its clone; durable output goes to the blackboard or$AGENT6_MACHINE_DATA_DIR - a machine with no run states runs its tool states in your checkout, unchanged
- states are sequential continuations of the branch (each starts from the previous state's tree); lanes are parallel alternatives cut at base
- a
mode = "run"state still returns only its outcome label andfinish_sessionpayload machine runresolves a git commit identity up front ([git.commit]or the repo's git config) so the confined agent's commits succeed
run_command in a mode = "run" state is gated by sandbox.run_commands:
- under the default
ask, an unattended machine auto-denies every call (machine runwarns up front when amode = "run"state would hit this) - a machine spawned from the web or TUI hub parks each approval and question for the front-end (the spawn carries the detached
waitaway-mode)- the answer never depends on when the viewer attached
- grant per invocation with
--auto-approveorAGENT6_AUTO_APPROVE=1(ask upgrades to yes; a withheldnostays no), or setsandbox.run_commands = "yes"in the repo config --no-commandsorAGENT6_NO_COMMANDS=1withholds every command tool- agent states alone read the two env names
- a machine
[config]overlay cannot grant it (sandbox policy is operator-only) - edits and the auto-commit need no approval;
run_verify_commandand theverify_whencertification share the same gate: prefertoolstates over shelling out
The optional per-state knobs tune how that loop runs.
provider/effort/temperatureselect and tune the model;max_usd/max_tokens_fallbackbound this one agent slice- each falls back through the effective config when omitted (Machine config overlay)
- secrets are never expressed here, only a
providername that must exist in the effective config
tool¶
[states.scan]
kind = "tool"
command = ["scan-inbox", "--dir", "{{ inbox_dir }}", "--since", "{{ cursor }}"]
output_schema = "scan_result" # types `result` so fields are navigable
capture = { set = { pending = "{{ result.pending }}", cursor = "{{ result.cursor }}" } }
timeout_secs = 60
on = { ok = "have_items", nonzero = "poll", timeout = "poll" }
A single command, argv-style (never a shell string), through the existing run_in_jail.
nonzerois any non-zero exit- stdout parses as JSON, bound to the capture-scope name
result(Names, references, and namespaces) - a capture binds only on
ok:nonzeroandtimeoutleave the blackboard as it was- a branch reading a captured var on those edges reads the previous iteration's value
-
capture has two modes; a state uses at most one:
- Opaque whole-capture:
capture = { stdout_json = "<var>" }binds the entire parsed stdout to one variable- no
output_schemaneeded;resultis opaque and may not be dotted
- no
- Typed field-capture:
output_schema = "<record>"typesresult; pull fields withset = { <var> = "{{ result.<field> }}" }- every
result.<field>is statically checked, mirroring how anagentstate validatesfinish_session
- every
- Opaque whole-capture:
-
a
list-typed variable spliced as a bare argv element ("{{ pending }}") expands to one argument per element, and an empty list contributes no argument at all (Templating and list-splicing) scan-inboxis an illustrative stand-in: atoolstate runs whatever audited command the operator names
Network (opt-in, host network off by default).
- a
tool'snetworkis"auto","none"or"host""auto"(default): its own network where the host can give one, degrading to the host's with a warning onhardened"none": its own network, required; refuses onhardened, which cannot isolate a single tool
- only
network = "host"reaches the host network - the engine is a host-netns supervisor (each
agentstate is its own subprocess; Security considerations), so one opt-intoolcan be networked while every other jailed command stays offline - a
toolcommand is fixed and operator-reviewed: not the free exfiltration channel a networkedrun_commandwould be - honoring the opt-in is the operator's call via
sandbox.network(global/repo config, never the machine overlay):
sandbox.network |
jailed commands | tool w/ network="host" |
|---|---|---|
auto (def) |
no host network on strict |
⛔ refuse to run |
session |
the same (refuses on hardened) |
⛔ refuse to run |
only_explicit_states |
no host network | host network |
host |
host network | host network (and run_command) |
- the headline setup (offline commands + one networked reviewed tool):
sandbox.network = "only_explicit_states"+network = "host"on that state only_explicit_statesandsessionneedstrict; an unhonorable tool-network config refuses at startup, naming the conflictingsandbox.networkvalue and the fix- offline tool states each get their own network (separate launchers; no run-wide session network to share)
Secrets (opt-in, none by default).
- a
tool'spass_env: environment variable names its jailed command receives from the operator's environment, e.g.pass_env = ["X_TOKEN"] - only names the operator lists in
[machine].pass_envreach a jail (global/repo config, never the machine overlay) - a state naming one not listed refuses the run at startup, naming the variable
- a provider's
api_key_envis never allowed there, as for an MCP server'spass_env machine checknames every variable a state declares
Script bundles. A machine is a bundle: the .asm.toml plus an optional sibling scripts/ of operator-reviewed helpers (the kind machine create may draft).
- a
toolreferences one by a relative path startingscripts/(command = ["bash", "scripts/fetch.sh"]), resolved against the jail's mounted cwd: keep the bundle at or under the directory you runagent6from - a bare binary in
command[0]resolves against the jail PATH (the setmachine checkprobes andrun_commanduses), never the hostPATH; absolute paths for anything elsewhere machine checkvalidates the bundle: everyscripts/entry and static command reference resolves inside it (escaping symlinks rejected)strict: the bundle is RO-bound in every jail; a tool or agent cannot rewrite its own machine logic mid-runhardened: Landlock carves the same protection, granting read on the workspace and write on each top-level entry outside the bundle- new top-level entries are denied, so a tool writes to
$AGENT6_MACHINE_DATA_DIR
- new top-level entries are denied, so a tool writes to
Cross-iteration persistence: $AGENT6_MACHINE_DATA_DIR.
- a per-machine writable dir under the per-repo state dir (
<state-dir>/<repo-id>/machines/<id>/data/, out of the workspace), RW in every tool jail - under
hardenedthe workspace's existing entries are writable too, never a new top-level entry; the data dir is the durable home either way, and the journal records every transition
wait¶
[states.poll]
kind = "wait"
every_secs = "{{ poll_secs }}" # at most one of: every_secs | until
on = { tick = "scan", signal = "scan" }
A wait keeps a machine running for a long time without spending CPU or tokens.
- at most one of
every_secsoruntil(an absolute ISO-8601 instant); both is a load error - on entry the engine persists the absolute next-wake instant (
wait.json) before sleeping and journals it on theWaitFactafter the wake: replay re-reads it and never sleeps - without
--exit-on-waitthe engine blocks in-process until the instant or an externalsignal(a file/IPC poke) - the wake being journaled absolutely lets the
--exit-on-waitpersisted-wake driver (Reliability) run the identical file, no format change
Wait-forever (no timer). Declare zero timers to park indefinitely until an operator signal poke:
- a no-timer wait can never
tick: it declares onlysignal(atickedge is a load error, unreachable) - under
--exit-on-waitthe engine persists a signal-only pending wait (no wake instant) and resumes when poked
Poke payloads. agent6 machine poke <id> [--data <json> | --message <text>] carries an optional payload to the waking wait.
- one signal pending at a time: a second poke replaces the first, payload included
- the payload is journaled on the
signalWaitFact(replay-safe) and materialized to$AGENT6_MACHINE_DATA_DIR/poke.jsonfor the nexttool - no capture on
wait: the payload flows through the existing tool -> capture -> branch pattern - on replay the journaled payload reproduces the identical input
branch¶
[states.route]
kind = "branch"
when = [
{ if = "verdict.label == 'urgent' and verdict.confidence >= 0.7", goto = "record" },
{ else = true, goto = "poll" },
]
whenis ordered; the first matchingifwins; a finalelse = trueis required (total function, no stuck state)- the predicate grammar is restricted and non-Turing-complete (Execution semantics): comparisons,
and/or/not, membership,len(),has(), literals, blackboard references has()tests presence: the guard anoptionalrecord field needs before a read (has(out.score) and out.score > 0is safe;andshort-circuits)- no function calls beyond the fixed allow-list, no Python attribute access, no
eval - dotted references (
verdict.confidence) are data navigation by agent6's own evaluator, never Python attribute resolution - a hard security boundary: a
.asm.tomlmust never execute arbitrary code
terminal¶
[states.halt]
kind = "terminal"
status = "failed" # "ok" | "failed"
reason = "machine budget exhausted"
- absorbing: emits
machine.endand returns control to the CLI - a machine may have many terminal states (success and failure variants)
notify (any state)¶
Any state may carry an optional notify, a templated message emitted on entry:
[states.escalate]
kind = "wait"
notify = "needs a human: {{ reason }}" # or a table with a level
on = { signal = "resume" }
[states.done]
kind = "terminal"
notify = { message = "run finished", level = "info" } # info | warn | error
status = "ok"
reason = "done"
- entering the state journals a
machine.notifyevent (message + level) and fires the operator notify hook - presentation only: no edge, no control-flow effect, no blackboard write
machine.endis also a notify trigger, so a terminal need not setnotifyto be surfaced- the message is a blackboard template, checked at
machine check - a
waitstate emits once per park: the armed wake record is the machine's memory that it already entered- a resume that re-enters mid-park does not page the operator again
- every other state kind re-emits on a resume that re-enters it
Two channels render it; agent6 owns no push infrastructure:
- device-present front-ends (
agent6 web, the TUI Machines page,agent6 attach): an ephemeral notification - out-of-band:
[machine.notify].on_event(config.md), an operator argv on the host, outside the jail, on everymachine.notifyandmachine.end(fan out to ntfy/Pushover/email/Telegram)
4.4 Templating and list-splicing¶
Strings may contain {{ ... }} interpolations.
- an interpolation is one reference plus at most one filter, nothing more
- no arbitrary expressions, no chained filters, no method calls; anything richer belongs in a
branchpredicate (itself restricted) - this keeps validation and replay simple, and the format from quietly becoming a scripting language
There are two filters, both zero-argument:
| filter | applies to | result |
|---|---|---|
len |
str, list, or a json/record container |
the integer length |
json |
any value | compact JSON, object keys sorted (deterministic) |
- no
joinfilter: a delimited string a command must re-split is fragile and injection-prone; lists reach argv by splicing (below) - an interpolation renders to a string, except a lone filter-less
{{ ref }}incapture.set, which assigns the referenced value with its own type (it must match the target's declared type) - elsewhere a bare
{{ x }}is legal only for a scalar (str/int/float/bool); a barelist/json/record reference is a load error (applyjson, or splice a list in argv)
List-splicing (argv only).
- a
commandelement that is exactly"{{ listvar }}"(lone list reference, no filter, no surrounding text) expands to one argv element per item - an empty list contributes no argument, so the command runs one element shorter: guard it with a
branchonlen(x)where that changes the command's meaning - the only way a list crosses into a command; injection-safe (each element stays a distinct argument, never shell-re-parsed)
- two load errors guard it: splicing a non-list, and embedding
{{ listvar }}inside a larger string ("--x={{ items }}") - filter and reference grammar are validated at
machine check
4.5 Names, references, and namespaces¶
The rules for naming, writing, and reading variables.
agent6 machine check enforces each one, machine run re-checks them, and a violation is a load-time error.
Identifier grammar. A variable name and a state name each match ^[a-z][a-z0-9_]*$ (ASCII snake_case).
TOML quoted/dotted keys that would smuggle other characters ("last-seen", "a.b") are a load error.
The restriction exists because variable names appear as bare Name tokens in predicates (parsed by ast.parse); a non-identifier could not be one.
Three owners, one flat reference namespace. The [vars.operator], [vars.code], and [vars.agent] subtables decide who may write a variable.
They do not create three separate read namespaces.
Every variable is referenced everywhere (templates and predicates alike) by its bare name only: positions, never vars.code.positions and never code.positions.
The owner prefix never appears in a reference.
Three consequences, each a machine check error:
- Global uniqueness across owners. A name is declared in one of the three subtables only.
Declaring
positionsin both[vars.code]and[vars.agent]is rejected: "variable 'positions' declared in both[vars.code]and[vars.agent]; the three owner subtables share one read namespace". A bare reference is forbidden rather than resolved by precedence. - No bare top-level vars. Every variable must live under one of the three owner subtables.
A key written directly under
[vars](i.e.vars.positions) has no declared owner and is rejected: "vars.positionshas no owner subtable; put it in[vars.operator],[vars.code], or[vars.agent]". It is never silently ignored. - Reserved names. The bare names
vars,operator,code,agent, andresultmay not be used as variable names.resultis reserved for capture scope (below); the rest are reserved so a reference can never be read as an owner path.
Reference grammar (one grammar, used identically in predicates and templates).
ref := name ("." key)*
name := an identifier declared in one [vars.*] subtable
key := an identifier (a declared field of a record type)
- the first segment is a declared variable; the validator checks it exists
- further
.keysegments are ordered dictionary lookups by agent6's own evaluator: never Python attribute access, nevergetattr - a
.keyis legal only into a record type (Record schemas); each segment checks against the schema at load (a misspelled field is a load error) - dotting an opaque
jsonor a scalar is a load error:jsonis wholesale-only, which keeps every navigable path statically checkable
Capture scope and result.
- inside a state's
capturetable, the reservedresultdenotes the structured output the state just produced, visible only there - not a declarable blackboard variable; invisible outside the capturing state
- dottable only when typed by an
output_schemarecord (mandatory foragentstates; optional fortoolstates, which are otherwise whole-capture only)
A capture has two forms of target:
- a fixed source key (
stdout_jsonfortool,finish_jsonforagent) naming one blackboard variable to receive the whole output; -
a
set = { <var> = "<template>" }table assigning rendered templates (which may readresult/result.<field>) to blackboard variables. -
what a capture may write is the ownership wall:
tool->[vars.code]only,agent->[vars.agent]only; a[vars.operator]or undeclared target is a load error - the captured value's runtime type must match the target's declared type, or the machine halts loudly
State-name namespace.
- state names are a separate namespace from variables: referenced only by
initial,goto, andon, never in predicates or templates (a state and a variable may share a name) - every
goto/ontarget names a declared state; every declared state is reachable frominitial(each a load error otherwise)
4.6 Record schemas ([schemas.*])¶
A record type is a named, field-typed structure declared once under [schemas.<name>].
- used as a variable's
type(navigability) and as anagentstate'soutput_schema(payload validation at the trust boundary) - one mechanism for both
- the schema language is inline TOML
- each entry is
field = "<type>"orfield = { type = "<type>", ... }:
[schemas.classification]
label = { type = "str", enum = ["urgent", "normal", "spam"] }
confidence = "float"
note = { type = "str", optional = true }
Rules (all enforced at machine check):
| Rule | Behavior |
|---|---|
| Field types | str, int, float, bool, list[<scalar>], another schema name (recursive; cycles are a load error), or json (opaque escape hatch; itself not dottable, The blackboard) |
| Required by default | every field must be present in a validated payload unless optional = true (mirrors Config's extra="forbid"); unknown fields are rejected. An absent optional field is absent, not null: reading it unguarded is a runtime halt, has() is its predicate guard, and machine test exercises absence (dry-run synthesizes required fields only) |
enum |
string fields only; constrains a str to a fixed literal list, checked at the finish_session/capture boundary (earlier than a branch would re-check it) |
| Dotting | a .field in a predicate/template is type-checked against the schema (field must exist); a list/json/non-record field may not be dotted further |
4.7 Machine config overlay ([config])¶
A machine file may carry an optional top-level [config] table: an agent6 config fragment layered on the effective config for the run.
- the stack: defaults < global < repo <
--config FILE< the machine overlay (highest) - most knobs
agent6 config showlists are valid inside it; the refusals are below
[config.workflow]
verify_command = ["uv", "run", "pytest", "-q"]
[config.review]
trigger = "on_verify_fail"
[config.budget]
max_usd = 50.0
Unset keys read straight through to the lower layers, so a machine only states what it wants to change. Two hard rules:
- No connections/secrets, no sandbox policy, no presets, no MCP servers, no host hooks
[config.providers.*],[config.sandbox.*],[config.presets.*],[config.mcp.*], a top-levelpreset,git.run_repo_hooks,git.run_repo_filters,machine.notify,machine.pass_env,notify.on_complete,prompt.system_prompt_file: each a load-time error- endpoints, key-env names, and secrets live in the global config / secrets store; sandbox policy, presets, MCP servers, and host-argv hooks are operator decisions in the global/repo config
- a machine file may be LLM-drafted or shared: it must not widen its own egress, weaken its jail, or run host code through the overlay
- a top-level
presetis refused for the same reason: the operator's selection would resolve it into those same knobs - the overlay only routes to a provider name that already exists, and sets benign knobs (commit identity)
- Per-
agent-state knobs (State kinds) override the overlay for that one state- agent-loop precedence: per-state knob > machine
[config]> repo > global > built-in default
- agent-loop precedence: per-state knob > machine
5. Execution semantics¶
5.1 The engine as a pure reducer¶
load(file) -> Machine # pydantic, extra=forbid, frozen
blackboard = Machine.initial_vars()
state = Machine.initial
loop:
event = execute(state, blackboard) # the only impure step
blackboard = reduce(blackboard, event) # pure; a fact that cannot reduce halts
journal.append(event) # append-only, fsync
state = next_state(Machine, state, event, blackboard) # pure
snapshot(state, blackboard) # atomic temp+rename
if state is terminal: break
executeis the only place the world is touched (run an agent, run a tool, read the clock)- a fact is journaled only after it reduces;
reduceandnext_stateare pure - replaying the journal reproduces the exact path, branches included (the outputs a branch reads are in the journal)
5.2 Determinism guarantees and the predicate evaluator¶
- Branch edges are pure functions of the blackboard; the blackboard is a pure function of journaled events; no branch depends on un-logged state
- The predicate evaluator is a hand-written recursive walk over a small AST
ast.parse(..., mode="eval"), then a strict node allow-list:Compare,BoolOp(and,or),UnaryOp(not,-,+),Name,Constant,ListandTupleliterals of constants, and a fixed-nameCalllistAttributeis reinterpreted as record data navigation, never Python attribute access- anything outside the allow-list raises at
machine check - it parses but never calls
eval,exec, orgetattr; anAttributechain walks the blackboard dict, aNamemust be declared, any other free name is a load error
- Wall-clock, randomness, and external reads are captured as facts
agent6 machine replay <machine-id>feeds the recorded facts instead of touching the world: a completed run replays to the identical path offline
5.3 Persistence layout¶
Mirrors the existing per-run layout under the per-repo state dir, out of the workspace:
<state-dir>/<repo-id>/machines/<machine-id>/
machine.asm.toml # the exact source the run started from (replay, status)
journal.jsonl # append-only, fsync'd, one event per line
snapshots/<n>.json # blackboard + current state, atomic temp+rename
agent_transcripts/<utc-iso>-<seq>.json # one lossless request/response per file
states/<seq>-<state>/logs.jsonl # per-execution event stream (role.*/tool.*),
# the watchable live view; pruned to recent;
# <seq> is four digits: 0003-classify
states/<seq>-<state>/approvals/, questions/ # that execution's answer bridge
# (`<id>.answer` from a front-end), steer files beside them
data/ # writable scratch ($AGENT6_MACHINE_DATA_DIR)
machine.lock # single-writer guard (one process per machine)
worker.pid # the live worker; absent or stale once it exits
wait.json # the armed wait (state, seq, next wake), written before a
# wait sleeps or parks, cleared when it fires
signal # a poke awaiting the wait's next check (payload inside)
signal.consuming # a claimed poke, until the wake's step is acked
stop # a stop request awaiting the next transition boundary
frontends/<pid> # live front-end claims (a TUI or web watcher)
approvals/away.mode # a hub-spawned instance's away mode ("wait")
- each
agentstate execution emits alogs.jsonlstream understates/<seq>-<state>/(the samerole.*_delta/tool.*events a run emits): a running machine follows live like a run - the heavy per-state logs prune to the most recent
state_log_keep(default 50); the journal stays the complete transition history
Sizing for long-running machines:
- the journal grows one line per transition; a 10-minute-interval machine makes ~150k transitions a year (3 per idle tick)
- a
waitorbranchline is ~200 B; atoolline carries the command's full stdout and stderr; anagentline carries its payload
- a
- snapshots keep only the most recent
[machine] snapshot_keep(default 5,0= all)machine statusreads the newest readable one, falling back through the retained tail when the newest is corrupt- replay and recovery fold the whole journal and never read a snapshot
- per-state reasoning logs grow with agent-state executions only, and self-prune
- the journal has no rotation: archive or delete an instance dir when replay no longer needs it;
[budget] max_transitionsis the primary runaway guard
5.4 Idempotency and crash recovery¶
- a state runs, then one fsync'd
StepEventrecords its outcome and captured fact: the commit point - the capture validates before the StepEvent writes, so the journal never holds a fact a later
reducecould not replay- a tool's malformed stdout halts the machine loudly
- an agent's non-conforming
finish_sessionis refused in-run, so the model retries; a leg that never conforms lands outcomefailedand routes on that edge
- on restart the engine folds the journal and continues from the last StepEvent
- replay holds the journal to its facts: a recorded label or goto its fact does not imply, a branch the replayed blackboard would not take, an agent payload outside its schema, a journal without its begin event, events after its end, or an end that disagrees with the replayed position refuses the instance and names the remedy (archive the directory)
- the crash window is side-effect-done to StepEvent-on-disk: a kill there loses the fact and the step re-runs on resume
- the posture is at-least-once: a
toolwith an external side effect must be idempotent (the examples move a file or write$AGENT6_MACHINE_DATA_DIR, so a re-run is a no-op) - the journal is crash-tolerant: a torn final line drops on read and heals on the next append
6. Reliability for 24/7 operation¶
- Restartable
- a
waitblocks in-process or persists the next wake and exits 0, re-armed by a systemd timer / cron - the journal is the source of truth either way: a reboot loses nothing
- a
- Runaway guards
[budget]USD and[budget].max_transitionsstop the machine when crossed- a no-wait, no-spend loop is still bounded by
max_transitions
- Single writer:
machine.lock(flock) guarantees one process per machine id; a second invocation refuses - Health/visibility
machine status <id>: current state, blackboard, last N events, spend, next wakeagent6 attach <id>: the unified watcher, live (state overview, each transition, the active agent state's reasoning)- the TUI Machines page wraps the same view (Run opens it; Watch
wattaches) machine graph <file>: mermaid or Graphviz DOT (--format)
7. CLI surface¶
| command | effect |
|---|---|
agent6 machine create <task> [-o <file>] [--max-attempts N] |
LLM-drafted bundle: .asm.toml + every scripts/... file + a mock test per script (external seam), written into a drafting workspace of its own; per-draft gate: machine check, ruff, ty, mock tests in a no-network jail, dry-run; failures hand the problems back (--max-attempts, default 3); output: a draft for operator review + commit (Security considerations) |
agent6 machine check <file> |
validate: the [config] overlay against the config schema (and its refusals); parse; type-check vars; every edge target exists; every state reachable; every branch total; names unique across owners, each owned by a subtable; every reference declared; every capture inside the ownership wall; len() args and wait timings well-typed; the script bundle contained; script health (ruff with its own config discovery, the nearest .ruff.toml, ruff.toml or pyproject.toml [tool.ruff] above the bundle; ty on a private temp copy with no config); no execution, no network |
agent6 machine test <file> [--blackboard FIXTURE.toml] |
everything check does; the bundle's scripts/**/*_test.py mock tests in a no-network jail (strict only: elsewhere they count as skipped on the verdict line); a pure dry-run (no provider, no clock): per state, synthesize the success fact, push through the real reduce, confirm capture binds and the label routes; per branch, evaluate each when against defaults + --blackboard, print the winning goto; the full offline simulation, every seam mocked |
agent6 machine graph <file> [--format mermaid\|dot] |
emit the machine as a diagram. mermaid (default) prints stateDiagram-v2; dot prints Graphviz DOT for dot -Tsvg/dot -Tpng and the broader Graphviz/xdot ecosystem. Reachability is already computed at load, so both are pure renders of the same validated graph. |
agent6 machine run <file> [--exit-on-wait] [--auto-approve\|--no-commands] [--dangerously-disable-sandbox] |
start (or resume) a machine. Acquires the lock, drives the loop. With --exit-on-wait, persist the next wake and exit 0 (status waiting) at the first not-ready wait, for an external scheduler (systemd timer / cron) to resume. The approval and sandbox flags are the run flags, for the same reasons and with the same refusals. |
agent6 machine status <id> |
current state, blackboard, spend, next wake. Read-only. |
agent6 machine (machine list) |
this repo's machines: each instance's status and current state joined with the authored .asm.toml that declares it, then the authored files no instance has run (spec validity per file). Read-only. |
agent6 attach <id> |
follow a running instance live (the unified watcher; the same command follows a run): state overview + current state, each transition as it lands, and the active agent state's reasoning (its per-state logs.jsonl). Read-only; Ctrl-C to stop. agent6 attach --tui <id> opens the machine screen, where a running agent state takes a steer (s), as the web machine page does. |
agent6 machine poke <id> [--data <json>\|--message <text>] |
signal a waiting instance to wake on its next check; an optional payload reaches the next tool at $AGENT6_MACHINE_DATA_DIR/poke.json (journaled, replay-safe). |
agent6 machine stop <id> |
park at the next transition boundary through a durable marker (the state in flight finishes and journals its fact); wakes a sleeping wait, leaving it armed; no MachineEnd journaled (resumes with machine run); ended/not-running answered with the note and exit 0, no marker; a journal it cannot read refused; the same answers on the web machine page and the TUI machine screen (x) |
agent6 machine replay <id> |
deterministic replay from the journal (no world I/O); backtesting. |
agent6 config show/get/set/unset/add/remove/fix --machine-file FILE |
read and write the machine's own [config] overlay through the config surfaces, with the same refusals as hand-editing it. |
machine check names each problem with its state and rule (state 'act': branch is not total (no final `else`)).
It warns when a tool state's binary is not on the jail's PATH.
7.1 machine create¶
Describe a loop in plain language and get a first-cut bundle back.
It is an ordinary agent6 run handed this document's grammar, working in a drafting workspace of its own.
The model writes the .asm.toml and every scripts/... file there with apply_edit, one file at a time, and finishes when the bundle is complete.
No new tool.
The leg has the edit tools; run_commands = "no" withholds run_command, run_verify_command, run_metric_command and stop_background, and the operator's [workflow].metric is dropped.
No host is pre-allowed, so a headless fetch denies.
It never sees the operator's checkout, and its writes are bounded by the workspace the way any run's are by its repo.
- the workspace is an empty git repo under
[parallel].workdir(where lane clones and fork worktrees live)- each iteration commits, so the draft survives a failure for the operator to read
- every draft is gated:
machine check, ruff (the destination's ruff config), ty, themachine testdry-run, and mock tests in a no-network jail- the mock tests run under
strictonly: elsewhere they are counted as skipped - agent6 runs those validators itself between attempts (they need agent6, which no jailed command can reach) and hands the problems back
- up to
--max-attempts(default 3); the agent patches the files it wrote
- the mock tests run under
- the result is a draft:
-o <file>overwrites freely, else<name>.asm.tomlin the cwd, never clobbered (a collision prints to stdout, exits non-zero)- scripts land in
scripts/
- scripts land in
- each attempt is watchable: a draft dir under the state dir carries the prompt, the transcript, and a
logs.jsonlthe TUI/web follow live- the CLI streams in the foreground; the TUI and web start detached and follow the dir
- the Security considerations invariant holds:
createdrafts into the working tree only; the operator reviews and commits;machine runrefuses an uncommitted bundle
8. Module boundaries¶
The layering is ui → app → workflows → tools → sandbox, with agent6.machine a top-level package beside them, and workflows never import each other.
An agent state needs to invoke the loop workflow, so the engine cannot itself be a workflow without breaking that rule.
The engine does not import the workflow stack.
engine.drive takes a World; the live one, LiveWorld, runs an agent state through its agent_runner callable (Callable[[AgentRequest, Path | None], AgentExecResult]).
The second argument is the per-state event-log path (<instance>/states/<seq>-<state>/logs.jsonl) each agent-state execution streams to.
app/, which depends on both agent6.machine and agent6.workflows, builds that runner (build_machine_agent_runner in app/machine_agent.py) and wires it into the LiveWorld in app/machine/run.py.
The orchestration around machine create/run lives in app/machine/, and ui/cli adapts argv and renders.
So agent6.machine never gains an edge into agent6.workflows, and the tach graph stays acyclic.
Files (all from __future__ import annotations, strict pyright, pydantic only at the parse boundary, @dataclass(frozen=True, slots=True) for the internal value types):
machine/model.py: pydanticMachineSpec/state/var specs (the parse boundary).machine/_semantics.py: semantic validation andfinish_sessionpayload validation.machine/dryrun.py: the pure, no-I/O dry-run behindagent6 machine test.machine/predicate.py: the allow-list AST predicate evaluator.machine/template.py: the single interpolation/splicing engine shared by the validator and the runtime.machine/graph.py: the mermaid/DOT renderers.machine/journal.py: append-only event log, snapshots, locking, and persisted-wake state.machine/engine.py: the deterministic reducer loop.machine/authoring.py: the per-attempt prompt builder formachine create, around the grammar guide inagent6.prompts.machine; it imports nothing from the workflow stack.
No new runtime dependency (tomllib + pydantic + stdlib ast).
9. Security considerations¶
- No new LLM tool surface
- the fixed set in
tools/schema.pyis unchanged; machines orchestrate existing capabilities machine createis no exception: the drafting agent has the same edit tools any run has, pointed at a drafting workspace of its own- every command tool is withheld, and its
[workflow].metricandfetchreach are removed
- the fixed set in
- No arbitrary code execution from a file
- predicates and templates are parsed-then-walked against an allow-list; never
eval/exec, nevergetattr - dotted references are agent6-interpreted data navigation
- predicates and templates are parsed-then-walked against an allow-list; never
- All side effects stay jailed
toolstates go throughrun_in_jail; eachagentstate is an ordinary run in its own subprocess, commands jailed like any run'smode = "run"machines never touch the operator's checkout: fresh clones per state, commits onagent6/machine-<id>, tool-state tree writes discarded with the clone- per-state network and refusals: security.md, Network
- Spend bounds
[budget].max_transitionsis required and always bindsmax_usd(optional) caps cumulative metered spend; an unpriced model is bounded per state by the agent6 config's[budget].max_tokens_fallback(0refuses unmetered models outright)- a supervisor crash mid-state cannot re-grant its slice: the resume books the orphaned per-state totals as an
attempt.spendjournal event, counted everywhere
- Machines are operator artifacts
- the threat model assumes the file is operator-written and reviewed like code; an LLM may propose (
machine createdrafts), running requires operator review + commit createwrites into its own workspace, publishes one reviewed bundle into the working tree, and never auto-runsrunoperates on a committed bundle, records it under the instance dir at first run, and refuses a continuation whose bundle drifted from the recorded bytes- a live instance runs the logic it recorded; an edit takes effect on a new instance
- drafting is assistance; authorization stays human
- the threat model assumes the file is operator-written and reviewed like code; an LLM may propose (
- External-world tools remain out of scope
- adding a tool that reaches the network or an external service is a separate change: the
tools/schema.pysecurity-review note plus a network/jail audit - the examples here use illustrative stand-in tools only
- adding a tool that reaches the network or an external service is a separate change: the
10. Worked example¶
# item-classifier.asm.toml (illustrative). scan-inbox/archive-item are
# stand-in audited tools, not part of agent6; they only show the *shape*.
machine = "item-classifier"
version = 1
initial = "poll"
[budget]
max_usd = 25.0
max_transitions = 100000
[vars.operator] # operator inputs, fixed for the machine's life
inbox_dir = { type = "str", value = "/srv/inbox" }
poll_secs = { type = "int", value = 300 }
[vars.code] # set deterministically by a tool capture
pending = { type = "list[str]", default = [] } # set by the scan tool
cursor = { type = "str", default = "" } # set by the scan tool
[vars.agent] # set by an agent state's finish_session
verdict = { type = "classification", default = {} } # set by classify
[schemas.classification] # validates the agent's finish_session payload
label = { type = "str", enum = ["urgent", "normal", "spam"] }
confidence = "float"
[schemas.scan_result] # types the scan tool's stdout
pending = "list[str]"
cursor = "str"
[states.poll]
kind = "wait"
every_secs = "{{ poll_secs }}" # at most one of every_secs | until
on = { tick = "scan", signal = "scan" }
[states.scan]
kind = "tool"
command = ["scan-inbox", "--dir", "{{ inbox_dir }}", "--since", "{{ cursor }}"]
output_schema = "scan_result"
capture = { set = { pending = "{{ result.pending }}", cursor = "{{ result.cursor }}" } }
timeout_secs = 60
on = { ok = "have_items", nonzero = "poll", timeout = "poll" }
[states.have_items]
kind = "branch"
when = [
{ if = "len(pending) == 0", goto = "poll" },
{ else = true, goto = "classify" },
]
[states.classify]
kind = "agent"
prompt = """
Classify these pending items: {{ pending | json }}
Call finish_session with JSON {label:"urgent"|"normal"|"spam", confidence:0..1}.
"""
output_schema = "classification"
capture = { finish_json = "verdict" }
timeout_secs = 600
on = { ok = "route", failed = "poll", timeout = "poll", budget_exhausted = "halt" }
[states.route]
kind = "branch"
when = [
{ if = "verdict.label == 'urgent' and verdict.confidence >= 0.7", goto = "record" },
{ else = true, goto = "poll" },
]
[states.record]
kind = "tool"
# `{{ pending }}` is a lone list reference -> spliced to one argv element per item
command = ["archive-item", "--label", "{{ verdict.label }}", "{{ pending }}"]
timeout_secs = 30
on = { ok = "poll", nonzero = "poll", timeout = "poll" }
[states.halt]
kind = "terminal"
status = "failed"
reason = "machine budget exhausted"
agent6 machine graph prints the control flow, one edge per transition, labelled with the state's on key or branch predicate as written, whitespace collapsed:
stateDiagram-v2
[*] --> poll
poll --> scan: tick
poll --> scan: signal
scan --> have_items: ok
scan --> poll: nonzero
scan --> poll: timeout
have_items --> poll: len(pending) == 0
have_items --> classify: else
classify --> route: ok
classify --> poll: failed
classify --> poll: timeout
classify --> halt: budget_exhausted
route --> record: verdict.label == 'urgent' and verdict.confidence >= 0.7
route --> poll: else
record --> poll: ok
record --> poll: nonzero
record --> poll: timeout
halt --> [*]