Macrae

Updated 2026-10-09 · final state · all on Cloudflare, agents on Modal, voice by ElevenLabs

A research agent for Pavel Jungwirth's group at the IOCB (Institute of Organic Chemistry and Biochemistry of the Czech Academy of Sciences, Prague). You chat with Jarvis, by voice or by typing, and it answers like a capable assistant, citing the group's papers whenever they are relevant. When you start a task, it plans the run, does real computational chemistry in cloud sandboxes, writes a short manuscript, and shows you every paper it read, every calculation it ran and what it cost. It learns from its own traces: it notices a capability it lacks, creates it, tests it, installs it and uses it again, while its authority stays fixed.

1 · The pieces

Everything public sits behind one Cloudflare Worker (a small program on Cloudflare's network) at one address. The Worker serves the page and hands the browser a short-lived voice session. It forwards /api calls to the backend, adding a secret header the browser never sees. The backend is a Cloudflare Container, owned by a Durable Object: a named Cloudflare object that keeps exactly one instance and decides when it sleeps. The backend answers paper questions from its own index and runs tasks as agent-runner flows. Harbor starts each agent step in a sandbox on Modal and keeps the full trace; the sandbox also streams every step back while it works. Runs, traces, the capability registry and the paper index are mirrored to R2, Cloudflare's object storage. The typed chat, the planner and the lesson distiller call Claude through the Anthropic API from the container.

Abbreviations in the diagram: API = application programming interface · LLM = large language model · RAG = retrieval-augmented generation (search the papers first, then answer with what was found, citing it) · BM25 = "Best Match 25", a keyword-ranking formula · vCPU = virtual processor core · GiB = gibibyte · CPU / GPU = central / graphics processing unit · PySCF = Python-based Simulations of Chemistry Framework · ATIF = Agent Trajectory Interchange Format (Harbor's trace file).

YOU · BROWSER ELEVENLABS CLOUDFLARE · ONE WORKER, ONE ADDRESS MODAL · STARTED BY HARBOR Web pageweb/ · chat with Jarvis • chat + mic button • [n] → source cards • panel tabs Tasks & runs · Evolutioncapabilities · manuscript • icons per event 📄 🔎 🧮 ✍️ ✅ ⚠️ • Dusk → Dawn table still loads when thebackend is asleep Voice agent"macrae" · voice "Jarvis" LLM: Claude Sonnet 5.5 server tools: search_papersstart_task run_status page tools: show_citationsopen_run cites author · year · page Worker "macrae"cloudflare/ · wrangler deploy • serves web/ (static files) • /api/* → backendadds the secret header • /voice/signed-urlElevenLabs key stays here • rate limits per visitortask starts · voice sessions • secrets (names only)ElevenLabs · Anthropic Modal · tool + admin secret never sent to the browser Durable ObjectMacraeBackend · binding BACKEND • one instance: macrae-backend • sleeps after 2 h idlenever while a run is going Container (standard-1)deploy/Dockerfile · ½ vCPU · 4 GiB • API (FastAPI) server/chat · runs · planner · costs • RAG rag/: BM25 + vectors[n] citations with page • agent-runner flowsHarbor: 1 sandbox per step • evolve/: lessons, registrycapabilities · dusk → dawn disk temporary, mirrored to R2 on restart: waits up to 14 min R2 bucket "macrae-data" object storage · Worker binding DATA container reaches it via the Worker index/paper index (chunks, vectors) state/runs/flow state.json, step logs state/jobs/Harbor trial folders = traces state/macrae/ · state/evolve/event logs · lessons · capabilities Sandbox per agent stepHarbor starts one per step Claude CodeClaude Opus 5.5 · Sonnet 5.5 • reads paper excerpts • runs PySCF (DFT) • writes manuscript.mdmath · figures · [n] citations • uses installed tools CPU (up to 32 cores) or GPU streams every step live (≤ 1 s) Harbor trial = the trace agent/trajectory.json ATIF: every message, tool calland its output agent/claude-code.txtlive stream while it runs + reward from the check Images (cached by Modal) agent: Claude Code + live wrapper learned: calc-image, pinned PySCF · numpy · matplotlib voice audio tool webhooks /api/tools/* /api/* sync: every step + 60 s
what the visitor touches voice backend cloud compute stored data / traces a task run a voice session

2 · Status

Checked on the live site on 2026-10-09 after the last deploy. "Live" means on the public site and checked there; "tested" means merged and covered by tests (Harbor, Modal and Claude faked), not re-run live; "not built" means planned only.

statewhatwave
liveSite on one Worker: the page, the /api proxy and /voice/signed-url. The backend is one Cloudflare Container (standard-1: ½ vCPU, 4 GiB) that sleeps after 2 h unless a run is going and mirrors its state to R2.v1
liveChat with Jarvis over 97 group papers (1210 chunks): answers cite [n] with author, year and page and say which parts don't come from the papers, with the cost of each answer (about $0.01).w4
liveVoice: ElevenLabs agent "macrae", voice "Jarvis", Claude Sonnet 5.5 through ElevenLabs, server tools as webhooks to the Worker, page tools in the browser.v1
liveReal agent runs on Modal through Harbor: the ion–water binding task. The first verified run found −25.5 kcal/mol for Na+–water at 2.2 Å (DFT, density functional theory, in PySCF, B3LYP/def2-TZVP with counterpoise correction). The same task started from the cloud container passed twice; those first runs reinstalled PySCF every time and hit a virtual environment without pip: the seed of what Macrae learned to fix.v1
liveLive traces (the sandbox streams every step within about 1 s), costs in money (section 5), a planner that decides hardware, stages and budget before each run, and lessons distilled from every trace (12 from the first runs).v2
liveAuthority holds: the first forge built a pinned environment, but its tests failed in a fresh sandbox, so the installer rejected it and recorded why.v3
liveDusk → dawn, the Frankenstein loop end to end. Dusk (20261008-225728-dusk-f086, nothing learned): both runs passed, reward 1, 0 errors; both traces showed the environment gap (no pip / ensurepip, PySCF reinstalled every run). Macrae started two forges by itself; each wrote a pinned environment and its tests (H2 RHF/FCI against textbook values, H atom −0.5 Ha), passed them in a fresh sandbox and was installed (calc-image, small-calc-env). Dawn (20261008-232837-dawn-0ac3, the image + 8 lessons): both passed again, same answers (Na+–water −25.46, K+–water −17.92 kcal/mol). Wall time 1035 → 690 s (−33 %), setup 505 → 236 s (−53 %), installs 43 → 0 s, cost $1.44 → $1.04 (−27 %). The first dawn attempt failed at the image build (the planner had set the image variable); the planner may no longer touch it. One run per ion per side: a demonstration, not a statistic.final
livePublic without login: rate limits, a daily spend cap, a kill switch, inputs that never reach a shell; the benchmark, forge and kill-switch routes need the operator secret. Downloadable proof: a trace zip and an offline HTML report for every run.w4
liveFrankenstein on the page: the capability ledger, installed capabilities and the fixed Authority card, the dusk → dawn table (Evolution tab), and the manuscript a finished run wrote.final
testedMethods card task: end-to-end tests with a fake Harbor, not re-run live after the last deploy.v1
not builtThe two BFF research tasks (section 6): scouted and approved, not built.—
ownerPaper PDFs stay private: they are indexed on the server and never published or committed.always

3 · What happens when you…

…ask a question
  1. Typed: the page sends the message, the anonymous session id and the recent history to POST /api/chat. The backend searches the papers for every message and sends back a Claude answer with its cost (at most 20 answers a minute per address).
  2. By voice: the mic button asks the Worker for a signed URL (a short-lived link; URL = uniform resource locator) and the ElevenLabs session starts in the browser. The voice agent calls search_papers, a webhook that goes to the Worker, then through the Durable Object to the container.
  3. RAG returns the best passages, numbered [1] [2], each with title, authors, page and DOI (Digital Object Identifier).
  4. Jarvis chats normally and uses the passages when they help, naming the source ("Košťál 2026, page 3"). It says which parts come from the papers and which from general knowledge, and never invents a citation.
  5. Numbered source chips appear under the answer (show_citations during a call). Clicking [n] opens that source's card.
…start a task (click Run, or ask)
  1. The page shows an estimate from past runs. POST /api/tasks/{id}/start returns a run_id at once, or "try again" when a rate limit, the cap on active runs or the daily spend cap is reached.
  2. The planner (one Claude call, with the task, its inputs, the lessons and the installed capabilities) decides the stages, the hardware and a budget. The page shows that decision as the first card.
  3. Script steps pull paper passages from RAG. Agent steps run Claude Code in Modal sandboxes through Harbor, with the installed capabilities mounted and the lessons in their instruction.
  4. The sandbox streams each step back; the page polls /api/runs/{id}/events every 1–2 s, and each tool call becomes a timeline row (📄 read, 🔎 search, 🧮 calc, ✍️ write, ✅ result). The cost meter grows with it.
  5. A check: verifies the output (numbers in result.json, a manuscript with math, a figure and citations) and gives a reward; a failed check feeds a retry. The state goes to R2 at every step.
  6. When the run ends, Jarvis posts the outcome, its cost and the papers it cited, the page shows the manuscript the agent wrote, and evolve distills lessons from the trace for the next run; a gap in the trace starts a forge.

4 · Where the traces are

A trace is everything one run did: the plan, every message, tool call and result, every file written, every check, and what it cost. Anyone can read it on the page and download it; the owner also has the raw copy in R2. The full guide is docs/TRACES.md in the repository.

wherewhat you findhow to get it
1 · The page timeline One row per event (plan, status, read, search, calc, write, think, cite, result, error, and the capability events gap, create, test, install, use, rejected), with a short human title, details, cost and source cards. The server builds it from the run's state.json and step logs, the lines the sandbox streams while a step runs, and each Harbor trial when it ends, without repeating anything seen live. The Manuscript it wrote box shows the result when the run ends; /manuscript also serves the full edit stream. Run view; GET /api/runs/{id}/events, /manuscript
2 · Zip and report A zip of every file behind the run: the flow's state, plan, costs, step logs, live stream lines, the Harbor job folders with trajectory.json, the results and the manuscript. And a self-contained HTML report that works offline: timeline, tool calls with their outputs, costs, manuscript, citations and checks. Download trace and Open report on each run; GET /api/runs/{id}/trace.zip
3 · R2 bucket macrae-data state/runs/<run_id>/: state.json, step logs, live lines, plan, costs, outputs.
state/jobs/<run_id>/…: Harbor trial folders, i.e. the full agent traces (agent/trajectory.json in ATIF, agent/claude-code.txt, the reward, the files the agent changed).
state/macrae/: the server's event logs and run → task records.
state/evolve/: what Macrae learned: lessons, the capability registry and ledger, benchmarks.
index/: the paper index.
The container pushes changes at every step and every 60 s, and restores them when it starts. At most about 60 s can be lost if a host dies without warning.
Cloudflare dashboard or wrangler r2 (owner only)

How the timeline fills in live: a wrapper around Claude Code inside the sandbox forwards each line of its output to POST /api/live/{run}/{step}, at most about 1 s late, with a per-run token (never the shared secret). The server stores these lines in <run>/live/<step>.jsonl and de-duplicates them against the final trajectory.json. Traces also feed learning: lessons for the next run, and python -m evolve export turns passing trajectories into a fine-tuning dataset.

5 · Costs

Every chat answer, run and step shows what it cost in money, split into LLM (large language model) dollars and compute dollars, and how long it took. A run's meter grows live while it runs; the header keeps a total for your session, with a small "what this cost" breakdown; and before you start a task the page estimates its cost and time from past runs. The price and rate tables in server/costs.py cite their sources and dates.

Tokens per step input · output · cache read / write per model, planner calls too from Claude Code's stream Modal sandbox seconds CPU core-seconds · memory GPU type and seconds per step, from Harbor LLM $ tokens × price table per model (server/costs.py); Claude Code's own total_cost_usd wins if given Compute $ seconds × Modal rate table (one dict, with its source) CPU, memory, GPU type Live cost meter on the page LLM $ + compute $ = total $ tokens · per step: model, hardware time per phase: plan · setup · work · check · and wall clock GET /api/runs/{id} → "costs"

Typical numbers:

6 · Research tasks

Two tasks are on the site, each a flow of script steps on the backend and a Claude Code agent in a Modal sandbox:

live

Ion–water binding, computed live

Pick an ion (Li+, Na+, K+, Mg2+, Ca2+). The agent runs a B3LYP/def2-SVP distance scan in PySCF and a counterpoise-corrected def2-TZVP energy at the minimum, and compares it with full and ECC-scaled (×0.75) point charges. It keeps a lab notebook (NOTES_TO_SELF.md) and writes a short manuscript through Write/Edit only: equations, a figure, and numbered citations of the group's papers, drafted early and revised as numbers arrive. About 5–10 min and $0.25–0.75.

tested

Methods card for a group paper

Give a DOI. The backend pulls that paper's passages from the index, and the agent writes a methods summary in which every claim cites a passage [n]; a check rejects uncited claims.

Planned, not built: two tasks on BFF (BayesicForceFields, the group's code for Bayesian learning of partial charges; Košťál, Shanks, Jungwirth, Martinez-Seara, JCTC 2026). Agents scouted the group's papers and the BFF code (research/scout/) and the owner approved two ideas: "Is 0.8 the right charge-scaling factor? A Bayesian answer with BFF" (acetate, ECC 0.70–0.90, a Bayes factor) and "Ca2+–acetate binding with error bars" (umbrella-sampling free energies from the BFF charge posterior, compared with Raman data). The wave that was to build them ran out of Claude session time.

7 · Frankenstein: an agent that builds itself

The hackathon brief: "By dawn, show a creature that learned to do things it could not do at dusk." Macrae notices a capability it is missing, creates it, tests it, installs it, and uses it again in later runs. Its capabilities may evolve; its authority may not. (SHA-256 = Secure Hash Algorithm, 256-bit: a fingerprint that changes if a single byte of the tool changes.)

1 · Gap the agent needs a toolthat isn't there and says "CAPABILITY_GAP: …" 2 · Create a sandbox step writesthe tool and its tests (Harbor on Modal) 3 · Test runs the tests;the step's check passes only if they all pass 4 · Install server checks SHA-256,tests and policy, then adds it to the registry 5 · Use again later runs get it in theirsandbox and instruction; no re-creating evolve: lessons + capability ledger feed the next run 🔒 Authority: fixed (server/policy.py, constants that no capability can change) • tools run only inside the Harbor sandbox • no secrets (environment allowlist) • no network hosts beyond a fixed list • no write access to the registry • a maximum runtime • asking for more → rejected The installer refuses any manifest that asks for more, and records a "rejected" event with the reason.

8 · Code map

folderwhat's in it
data/, scripts/479 group publications (titles, authors, DOIs) from the IOCB site.
web/Chat with Jarvis: a thread with mic and composer, answers with [n] citations opening source cards, a session cost total. Right panel with two tabs. Tasks & runs: tasks with cost estimates, the current run (plan card, cost meter, live timeline 📄🔎🧮✍️✅⚠️, Download trace / Open report, the manuscript it wrote with math and figure). Evolution: lessons run over run, the capability ledger with the Authority card, Dusk → Dawn. Offline state, dark mode, a bottom sheet on phones. Plain HTML/CSS/JS, no build step.
cloudflare/Worker macrae (index.js + worker.js): serves the page, sends /api through the Durable Object to the container, ElevenLabs signed URL, rate limits, operator-only routes (benchmark, forge, kill switch, admin status/restart), the live-trace ingest route, and the container's R2 endpoint. wrangler.toml, deploy.sh, a local dev server and a mock backend.
server/FastAPI (a Python web framework): tasks, runs, trace events (tool calls → read / calc / write…), live ingest, chat, search, ElevenLabs tools, costs, the planner, capabilities and the fixed policy, manuscripts, trace downloads, rate limits, spend cap and kill switch, draining on restart.
rag/PDF → clean pages (headers, hyphenation and reference lists removed) → chunks; BM25 + bge-small embeddings (BAAI General Embedding, by the Beijing Academy of Artificial Intelligence), fused ranking; [n] citations with page.
tasks/The two tasks and their flows (methods card, ion–water DFT), preparation and check scripts, the research protocol (notebook + manuscript), the calc-image capability seed, and tasks/common/: the Modal agent image with the live wrapper.
evolve/Lessons distilled from traces, the capability registry, the dusk → dawn benchmark, run-over-run metrics, a fine-tuning dataset export.
agent_runner/Runs flows of agent and script steps through Harbor (Docker or Modal), with checks, retries and several Claude accounts. An attempt that errored (for example on a Claude rate limit) never counts as passed.
voice/ElevenLabs agent "macrae" (voice "Jarvis", Claude Sonnet 5.5 through ElevenLabs): server tools, page tools and a setup script.
deploy/Backend Dockerfile (Harbor uses Modal, so there is no Docker inside), start.py (supervises the server, drains runs on restart), r2sync.py (R2 mirror, including what Macrae learned), the secret checker, docs.
proof/Saved live payloads (runs, lessons, capability ledger, dusk → dawn report) behind the numbers on this page and in the README.
tests/Cross-module tests: RAG → flow → API integration, an end-to-end run with a fake Claude and Harbor, and checks of these docs.
research/Scouting notes on the group's papers and the BFF code; a real run kept as a replayable bundle (demo/small-calc-na/).
docs/This page, the presenter's demo script and the traces guide.

9 · Where it runs: all on Cloudflare

The page, the Worker and the backend live on one Cloudflare account and one address, deployed with cloudflare/deploy.sh (which runs wrangler deploy). The backend is a Cloudflare Container, built from the same Docker image, and the Durable Object MacraeBackend owns it. It sleeps after 2 hours without use, but never while a run is going. Its disk is temporary, so the paper index is baked into the image, and everything the backend writes is mirrored to R2. The container reaches R2 through the Worker at a private address (r2.macrae), so it needs no S3 (Simple Storage Service) keys. On a restart, task starts pause and running flows get up to 14 minutes to finish, but a deploy that changes the image makes Cloudflare replace the container without waiting, so deploys happen when no run is active. Edits to the page alone don't restart the backend. The heavy work (agents, calculations) runs in Modal sandboxes. AWS (Amazon Web Services) was the first plan and is no longer used.

Open to everyone, no login. Each visitor gets an anonymous session id. Task starts are rate-limited per IP (Internet Protocol) address in the Worker and per address, per session, for the voice agent and globally in the backend, with a cap on runs active on Modal at once (4); chat answers are limited to 20 a minute per address. A daily spend cap ($20 by default) stops new paid work when reached, and the operator has a secret-protected kill switch. Every route validates its input, task inputs come from allowlists and never reach a shell, and paper text in tool outputs is handled as data, not as instructions. No secret reaches the browser, and the sandbox gets only what its step needs.

10 · How it was built

With agent-runner itself. A written contract (CONTRACT.md) fixed every interface. Then six Claude Opus 5.5 agents (high effort) built their modules at the same time across three Claude accounts, each in its own container. All six passed their checks on the first attempt.

moduletimecost (API-equivalent)
web + Worker18.5 min$5.24
server14.4 min$3.86
rag15.6 min$3.41
tasks + Modal16.5 min$4.95
voice10.2 min$2.99
deploy15.7 min$3.44
total~19 min wall clock$23.90, on Claude subscriptions

The later waves used the same method (Claude Opus 5.5 through agent-runner and Harbor, one git branch or folder per agent, then an integrator agent that merges, fixes the seams and tests end to end):

11 · What the owner provides

itemwhere it lives
ElevenLabs API key and agent idWorker secrets; never sent to the browser
Claude logins for the agents, and an Anthropic API key for chat, planner and distillerWorker secrets, passed to the container; rotate with wrangler secret put and a restart
Modal tokenWorker secret, passed to the container
Operator secret (benchmark, forge, kill switch)Worker secret MACRAE_ADMIN_SECRET, kept locally in deploy/.env; never given to the voice agent
Paper PDFsindexed on the server, never published or committed
Research questionsapproved by the owner: the two BFF tasks (not built yet)
HostingCloudflare (Workers Paid), Modal (per second), ElevenLabs

Glossary

RAGRetrieval-augmented generation: find the relevant passages first, give them to the model, and cite them in the answer.
WorkerSmall program on Cloudflare's network that serves the site and forwards API calls.
Durable ObjectA named Cloudflare object with its own storage. Here it owns the single backend container and decides when it sleeps.
ContainerThe backend's Docker image, run by Cloudflare on demand.
R2Cloudflare's object storage (files in a bucket). Holds the index, runs and traces.
HarborOpen-source harness that runs an agent in a sandbox and saves its trajectory, reward and files.
ModalCloud that starts sandboxes/containers on demand, CPU or GPU, billed per second.
Trace / trajectoryEvery message, tool call and result of one agent run (ATIF format in trajectory.json).
Signed URLA short-lived link (URL = uniform resource locator) that lets the browser open a voice session without seeing the API key.
BFFBayesicForceFields: Bayesian learning of force-field partial charges from reference simulations.
ECCElectronic continuum correction: scaling ionic charges (about 0.75–0.8) to mimic electronic polarisation.
CapabilityA tool the agent wrote, tested and installed for itself. Authority (what tools may do) stays fixed.
Dusk / dawnTwo states of the same benchmark: dusk with nothing learned, dawn with every installed capability and lesson.
LessonA short do/avoid/setting/tool rule distilled from a trace, with links to the evidence, given to the next run of the task.
PlannerOne Claude call before each run that picks the stages, the hardware and a budget, and says why.
ManuscriptThe short paper a research run writes as it goes (results/manuscript.md), only through Write and Edit, so the trace holds every revision; the page shows the final text.