stephen wolff

Resilient agent harness: a cheat sheet

working note engineering

The developer notes behind Vehicles on a harness.

Distilled from Petar Ivanov, “5 Resilience Patterns for AI Agents” (24 Sep 2026), re-aimed at a Braitenberg vehicle simulator and the harness that runs it.


The one idea

A request is a letter: if it’s lost, post it again. An agent run is a journey: if the train breaks down at stop 14, you don’t go back to the start and buy every ticket again. You resume from stop 14.

Failure maths: per-step success p over n steps gives pⁿ clean runs. 0.98²⁰ ≈ 0.67 (1 run in 3 fails) · 0.99²⁰ ≈ 0.82 (about 1 in 5). Failure is the main path, not the edge case.

Three properties break the classic request playbook: runs are long (a deploy lands mid-run), stateful (step 14 depends on 1–13), and act on the world (replaying a step repeats its side effects).


The five patterns

#PatternRule of thumbAnalogyGotcha
1Retry the step, not the runSort failures into three piles: time fixes it (HTTP 429 Too Many Requests, timeouts, 5xx server errors, 529 overloaded) → retry later; the model fixes it (bad ID) → return the error as a tool result; nothing fixes it (revoked key) → fail nowA musician fluffs bar 14: repeat bar 14, not the whole symphonyRetries multiply across layers. SDK (software development kit) retries × step retries = 9 calls. One layer owns retries
2Make every write safe to call twiceIdempotency key on every tool that writes, built from stable IDs (run, tool-call, tick), never generated inside the stepA lift button: pressing it five times still summons one liftA key minted inside the step is new on every retry, so it protects nothing
3Checkpoint so a crash means resumeSave state after every step (history, results, step counter). Something external must notice the dead run and re-enterA bookmark, not a photocopy of the whole bookCode outside a checkpointed step reruns on resume. Version saved state like an API, because runs outlive deploys
4Fall back across failure domainsCircuit-break, then switch to the same model on another cloud before a different model. Only switch to things your evals passed. Switch at a step boundaryIf the M5 is shut, take the A38, not the M5 southboundA sibling model at the same provider usually shares the outage
5Bound the loop, then hand to a humanCap steps, tokens and wall-clock time; set per-call timeouts. A stopped run = saved state + reason, ready for a humanA kitchen timer, not waiting for the smoke alarmThe worst failure throws no error: a loop where every call returns 200 forever. Default SDK timeouts are 10 min × 3 attempts

Approval-wait rule (from pattern 5): the saved decision in the database is the source of truth; the approval event is only a wake-up call. Poll on a short timeout so a missed event costs minutes, not a whole expiry.


Mapping onto vehicles and a harness

Article conceptIn the simulator
RunOne experiment: a world, some vehicles, some light sources, a seed
StepOne tick of the world (all vehicles sense → wire → move)
Model/tool callstep(state) -> state: a pure function, safe to replay
Side-effecting toolAnything leaving the process: sonifying a vehicle (a note-on), a real motor, a log sent to a server, publishing a frame
Idempotency keyf"{run_id}:{vehicle_id}:{tick}"
CheckpointFrozen WorldState saved after every tick (or every k ticks if cheap to recompute, as long as step is deterministic)
EvalsBehavioural tests: 2a must turn away, 2b must turn towards. These double as the TDD (test-driven development) suite
FallbackSwapping a vehicle’s “brain” (plain wiring → CTRNN (continuous-time recurrent neural network) → LLM (large language model) controller) only at a tick boundary, and only if it passes the behavioural tests
BudgetMax ticks, wall-clock limit, plus a stuck detector (vehicle orbits forever, or Vehicle 3a parks at the light)
Human hand-offOutcome("parked", state, reason) you can inspect, tweak and resume

The design point: make the core deterministic and pure (seeded RNG, random number generator; no I/O in step). Then patterns 1–3 are almost free: a replayed tick gives the same answer, and only the edges need idempotency keys.


Wiring crib (Braitenberg’s Vehicles, 1984)

VehicleWiringSignBehaviour
2a “Fear”same sideexcitatoryspeeds up and turns away
2b “Aggression”crossedexcitatoryspeeds up and turns towards, rams it
3a “Love”same sideinhibitoryturns towards, slows, stops facing it
3b “Explorer”crossedinhibitoryapproaches, then slows and wanders off

Differential drive: forward speed = (L + R) / 2, turn rate = (R − L) / axle. Left wheel faster → turns right.


Code sketch (Python)

# vehicles.py: pure domain, no I/O
from dataclasses import dataclass, replace
import math

@dataclass(frozen=True)
class Vehicle:
    id: str
    x: float
    y: float
    heading: float          # radians, anticlockwise from +x
    wiring: str             # "2a" | "2b" | "3a" | "3b"

@dataclass(frozen=True)
class WorldState:
    tick: int
    vehicles: tuple[Vehicle, ...]
    lights: tuple[tuple[float, float], ...]
    schema_version: int = 1  # version saved state like an API

SENSOR_ANGLE, SENSOR_DIST, AXLE = 0.5, 0.5, 1.0

def _intensity(px, py, lights):
    return sum(1 / (1 + (px - lx) ** 2 + (py - ly) ** 2) for lx, ly in lights)

def sense(v, lights):
    def at(offset):
        a = v.heading + offset
        return _intensity(v.x + SENSOR_DIST * math.cos(a),
                          v.y + SENSOR_DIST * math.sin(a), lights)
    return at(+SENSOR_ANGLE), at(-SENSOR_ANGLE)   # (left, right)

WIRING = {
    "2a": lambda l, r: (l, r),
    "2b": lambda l, r: (r, l),
    "3a": lambda l, r: (1 - l, 1 - r),
    "3b": lambda l, r: (1 - r, 1 - l),
}

def move(v, lights, dt=0.1):
    left_motor, right_motor = WIRING[v.wiring](*sense(v, lights))
    speed = (left_motor + right_motor) / 2
    heading = v.heading + (right_motor - left_motor) / AXLE * dt
    return replace(v, heading=heading,
                   x=v.x + speed * math.cos(heading) * dt,
                   y=v.y + speed * math.sin(heading) * dt)

def step(state: WorldState) -> WorldState:
    return replace(state, tick=state.tick + 1,
                   vehicles=tuple(move(v, state.lights) for v in state.vehicles))
# harness.py: the only place with state, time and side effects
import time
from dataclasses import dataclass

@dataclass(frozen=True)
class Budget:
    max_ticks: int = 10_000
    max_seconds: float = 60.0

    def exceeded(self, state, started):
        if state.tick >= self.max_ticks:
            return "tick_budget"
        if time.monotonic() - started > self.max_seconds:
            return "time_budget"
        return None

@dataclass(frozen=True)
class Outcome:
    status: str             # "done" | "parked"
    state: object
    reason: str | None = None

class Harness:
    def __init__(self, store, budget, step_fn, sinks=(), done=lambda s: False):
        self.store, self.budget, self.step = store, budget, step_fn
        self.sinks, self.done = sinks, done

    def run(self, run_id, initial):
        state = self.store.load(run_id) or initial     # resume, don't restart
        started = time.monotonic()
        while not self.done(state):
            if reason := self.budget.exceeded(state, started):
                self.store.save(run_id, state)
                return Outcome("parked", state, reason)  # hand to a human
            state = self.step(state)
            for sink in self.sinks:                      # side effects, keyed
                sink.emit(run_id, state)
            self.store.save(run_id, state)               # checkpoint every tick
        return Outcome("done", state)

class IdempotentSink:
    """Wraps anything that acts on the world (audio, motors, network)."""
    def __init__(self, act):
        self.act, self.seen = act, set()   # use a persistent store in real life

    def emit(self, run_id, state):
        for v in state.vehicles:
            key = f"{run_id}:{v.id}:{state.tick}"   # stable IDs only
            if key not in self.seen:
                self.act(key, v)
                self.seen.add(key)

Key tests (pytest)

import math
from vehicles import Vehicle, WorldState, move, step
from harness import Harness, Budget, IdempotentSink

LIGHT_AHEAD_LEFT = ((5.0, 2.0),)

def vehicle(wiring):
    return Vehicle("v1", 0.0, 0.0, 0.0, wiring)

# Behavioural "evals": the wiring does what Braitenberg says
def test_fear_turns_away_from_light():
    assert move(vehicle("2a"), LIGHT_AHEAD_LEFT).heading < 0

def test_aggression_turns_towards_light():
    assert move(vehicle("2b"), LIGHT_AHEAD_LEFT).heading > 0

def test_love_turns_towards_light():
    assert move(vehicle("3a"), LIGHT_AHEAD_LEFT).heading > 0

# Pattern 3: a crash then resume ends where an uninterrupted run ends
class MemoryStore:
    def __init__(self): self.data = {}
    def load(self, k): return self.data.get(k)
    def save(self, k, s): self.data[k] = s

def world():
    return WorldState(0, (vehicle("2b"),), LIGHT_AHEAD_LEFT)

def test_resume_matches_uninterrupted_run():
    clean = Harness(MemoryStore(), Budget(max_ticks=50), step).run("a", world())

    store = MemoryStore()
    Harness(store, Budget(max_ticks=20), step).run("b", world())   # "crash" at 20
    resumed = Harness(store, Budget(max_ticks=50), step).run("b", world())

    assert resumed.state == clean.state

# Pattern 2: replaying a tick doesn't repeat side effects
def test_sink_acts_once_per_vehicle_per_tick():
    calls = []
    sink = IdempotentSink(lambda key, v: calls.append(key))
    s = step(world())
    sink.emit("run", s)
    sink.emit("run", s)          # retry of the same tick
    assert calls == ["run:v1:1"]

# Pattern 5: the loop is bounded and parks with a reason
def test_budget_parks_the_run():
    out = Harness(MemoryStore(), Budget(max_ticks=5), step).run("c", world())
    assert (out.status, out.reason, out.state.tick) == ("parked", "tick_budget", 5)

Build order (working backwards)

  1. End goal: a watchable (and listenable) world where 2a/2b/3a/3b behave recognisably, runs survive a restart, and nothing sounds twice.
  2. Behavioural tests for the four wirings → pure step.
  3. Harness with budget + in-memory checkpoint store → resume test.
  4. Idempotent sinks (renderer, sonification) → replay test.
  5. Swap the store for SQLite or files; add schema_version migration.
  6. Stuck detector (low displacement over k ticks) as a second budget reason.
  7. Only then: pluggable brains (CTRNN, LLM), each gated by the same behavioural tests (pattern 4).

TL;DR

  • Retry the tick, not the run; one layer owns retries.
  • Keys from stable IDs on everything that leaves the process.
  • Checkpoint every tick; keep step pure so replays are free.
  • Swap brains only at tick boundaries, only if they pass the behavioural tests.
  • Cap ticks, time and stuckness; a parked run is state + reason, not a loss.