Macrae

Updated 2026-10-08 21:10 CEST (Central European Summer Time) · status: live on Cloudflare · live traces, costs and "Frankenstein" in progress

A voice agent for the Jungwirth group at the IOCB (Institute of Organic Chemistry and Biochemistry of the Czech Academy of Sciences, Prague). You chat with it, by voice or by typing, and it answers from the group's papers with citations. When you start a task, it does real computational chemistry on cloud hardware and shows you every paper it read and every calculation it ran.

1 · The pieces

Everything public sits behind one Cloudflare Worker (a small program on Cloudflare's network) at one address. The Worker serves the page and hands the browser a short-lived voice session. It forwards /api calls to the backend, adding a secret header the browser never sees. The backend is a Cloudflare Container, owned by a Durable Object: a named Cloudflare object that keeps exactly one instance and decides when it sleeps. The backend answers paper questions from its own index and runs tasks as agent-runner flows. Harbor starts each agent step in a sandbox on Modal and keeps the full trace. Runs, traces and the paper index are mirrored to R2, Cloudflare's object storage.

Abbreviations in the diagram: API = application programming interface · LLM = large language model · RAG = retrieval-augmented generation (search the papers first, answer only from what was found) · BM25 = "Best Match 25", a keyword-ranking formula · vCPU = virtual processor core · GiB = gibibyte · CPU / GPU = central / graphics processing unit · PySCF = Python-based Simulations of Chemistry Framework · BFF = BayesicForceFields (Bayesian force fields, the group's code) · GROMACS = GROningen MAchine for Chemical Simulations · ATIF = Agent Trajectory Interchange Format (Harbor's trace file).

YOU · BROWSER ELEVENLABS CLOUDFLARE · ONE WORKER, ONE ADDRESS MODAL · STARTED BY HARBOR Web pageweb/ · chat with Jarvis • chat + mic button • [n] → source cards • panel "Tasks & runs" tasks · live run tracerecent runs • icons per event 📄 🔎 🧮 ✍️ ✅ ⚠️ still loads when thebackend is asleep Voice agent"macrae" · voice "Jarvis" LLM: Claude Sonnet 5.5 server tools: search_papersstart_task · run_status page tools: show_citationsopen_run answers only from [n] Worker "macrae"cloudflare/ · wrangler deploy • serves web/ (static files) • /api/* → backendadds the secret header • /voice/signed-urlElevenLabs key stays here • rate limits per visitor6 task starts, 10 voice / min • secrets (names only)ElevenLabs · Anthropic Modal token · tool secret never sent to the browser Durable ObjectMacraeBackend · binding BACKEND • one instance: macrae-backend • sleeps after 2 h idlenever while a run is going Container (standard-1)deploy/Dockerfile · ½ vCPU · 4 GiB • API (FastAPI) server/tasks · runs · events · tools • RAG rag/: BM25 + vectors[n] citations with page • agent-runner flowsHarbor: 1 sandbox per step • disk: runs/ jobs/ index/temporary, mirrored to R2 on restart: waits up to 14 min R2 bucket "macrae-data" object storage · Worker binding DATA container reaches it via the Worker index/paper index (chunks, vectors) state/runs/flow state.json, step logs state/jobs/Harbor trial folders = traces state/macrae/page event logs, run → task Sandbox per agent stepHarbor starts one per step Claude CodeClaude Opus 5.5 · Sonnet 5.5 • reads paper excerpts • runs PySCF, BFF, GROMACS • writes result.json+ a cited explanation CPU (up to 32 cores) or GPU v2: streams each step live(in progress) Harbor trial = the trace agent/trajectory.json ATIF: every message, tool calland its output agent/claude-code.txtlive stream while it runs + reward from the check Images (cached by Modal) agent: Claude Code + tools BFF: BFF + GROMACS (in progress) built once, then reused voice audio tool webhooks /api/tools/* /api/* sync: every step + 60 s
what the visitor touches voice backend cloud compute stored data / traces a task run a voice session

2 · Status

As of 2026-10-08, 21:10 CEST.

statewhatwhen
liveSite on the Worker "macrae": the page, the /api proxy and /voice/signed-url.2026-10-08
livePage: dark/light chat with Jarvis, the macrae orbit logo, and a right panel "Tasks & runs" (tasks, live run trace, recent runs).2026-10-08
liveBackend: one Cloudflare Container (standard-1: ½ vCPU, 4 GiB). It sleeps after 2 h unless a run is going, and syncs its state (runs, Harbor traces, paper index) to the R2 bucket "macrae-data". Health: ok. Modal: connected.2026-10-08
liveVoice: ElevenLabs agent "macrae", voice "Jarvis", Claude Sonnet 5.5 through ElevenLabs. Server tools search_papers, start_task and run_status are webhooks to the Worker. Page tools: show_citations, open_run.2026-10-08
liveAWS (Amazon Web Services) is no longer used: the Cloudflare Container replaced it.2026-10-08
verifiedA real Claude Code agent run on Modal through Harbor (local test). It computed the Na+–water binding energy with DFT (density functional theory) in PySCF: B3LYP functional (Becke 3-parameter, Lee–Yang–Parr), def2-TZVP basis set (triple-zeta valence with polarisation), counterpoise correction. Result: −25.5 kcal/mol at 2.2 Å, reward 1, 34 trace events (search, cite, read, calc, write, result).
It took 9 min: 170 s for the first image build on Modal, the agent installing PySCF itself, and a 4-min sleep. That is to be optimised.
2026-10-08, 20:18
verifiedTests: 277 Python tests and 57 Node tests pass. They cover the Worker → container path, the R2 sync, and draining runs on restart.2026-10-08
in progressFirst cloud run (container → Modal): the K+ small calculation.started 21:02
in progressTwo BFF research tasks, approved by the owner, being turned into clickable tasks (section 6).started 20:34
in progressv2: live traces (the sandbox streams each step), costs in money (section 5), a planner that decides what to run and where (shown first in the trace), and "Evolution": lessons distilled from each trace feed the next run, plus a tools library and a fine-tuning dataset export.started 20:43
in progressFrankenstein (v3): a capability lifecycle (create → test → install → use again) with fixed authority, and raw Harbor traces downloadable from the page (section 7). Queued after v2.queued 21:08
in progressVideo: a 90 s film celebrating Pavel Jungwirth's research, plus a 20–25 s accelerated demo. ElevenLabs music and narration.started 21:07
waiting on ownerPaper PDFs (Portable Document Format). Until they arrive the RAG index is empty, and Jarvis says the papers here don't cover the question.open
waiting on ownerA fresh Claude API key to replace the current one.open

3 · What happens when you…

…ask a question
  1. By voice: the mic button asks the Worker for a signed URL (a short-lived link; URL = uniform resource locator), and the ElevenLabs session starts in the browser. Typed, outside a call: the text goes straight to POST /api/search.
  2. The voice agent calls search_papers. That webhook goes to the Worker, then through the Durable Object to the container.
  3. RAG returns the best passages, numbered [1] [2], each with title, authors, page and DOI (Digital Object Identifier).
  4. Jarvis answers only from those passages and names the source ("Košťál 2026, page 3"). With no PDFs indexed yet, it says the papers don't cover the question.
  5. show_citations adds numbered source chips under the answer. Clicking [n] opens that source's card.
…start a task (click Run, or ask)
  1. POST /api/tasks/{id}/start returns a run_id. Jarvis posts a run chip in the chat, and the panel opens the live trace.
  2. v2 A planner first decides what to run and on which hardware, and the page shows that decision as the first card.
  3. Script steps pull paper passages from RAG. Agent steps run in Modal sandboxes through Harbor.
  4. The page polls /api/runs/{id}/events every 1–2 s. Each tool call becomes a timeline row (📄 read, 🔎 search, 🧮 calc, ✍️ write, ✅ result).
  5. A check: verifies the output (e.g. result.json has numbers) and gives a reward. The container pushes the run's state to R2 at every step.
  6. When the run ends, Jarvis posts the outcome with the papers it cited.

4 · Where the traces are

A trace is everything one run did: every message, tool call and result. The same trace appears in three places, from the most readable to the most raw.

wherewhat you findhow to get it
1 · The page timeline One row per event (status, read, search, calc, write, think, cite, result, error), with a short human title, details, and source cards. The server builds it from the run's state.json, the step logs and each Harbor trial. While a step is still running it reads the live claude-code.txt instead. Panel "Tasks & runs" → a run; or GET /api/runs/{id}/events
2 · R2 bucket macrae-data state/runs/<run_id>/: the flow's state.json, step logs, outputs.
state/jobs/<run_id>/…: Harbor trial folders, i.e. the full agent traces.
state/macrae/: the server's event logs and run → task records.
index/: the paper index.
The container pushes changes at every step and every 60 s, and restores them when it starts. At most about 60 s can be lost if a host dies without warning.
Cloudflare dashboard (owner only)
3 · Harbor trajectory.json Inside each trial folder: agent/trajectory.json in ATIF (steps with source, message, tool calls and their outputs), agent/claude-code.txt (Claude Code's raw stream), the reward, and the files the agent changed. In R2 under state/jobs/. v3 "Download trace" (zip) and "Open raw trajectory" on each run.

v2, in progress A wrapper around Claude Code inside the sandbox forwards each step to POST /api/live/{run}/{step}, at most 1 s late, with a per-run token. The server stores these lines in <run>/live/<step>.jsonl, so the timeline fills in while the agent works instead of when the step ends. Lines are de-duplicated against the final trajectory.json.

5 · Costs in progress (v2)

The page shows what each run costs in money, growing live while it runs. This is the design from the v2 addendum of the build contract. It is being built and is not on the live site yet.

Tokens per step input · output · cache read / write per model, planner calls too from Claude Code's stream Modal sandbox seconds CPU core-seconds · memory GPU type and seconds per step, from Harbor LLM $ tokens × price table per model (server/costs.py); Claude Code's own total_cost_usd wins if given Compute $ seconds × Modal rate table (one dict, with its source) CPU, memory, GPU type Live cost meter on the page LLM $ + compute $ = total $ tokens · per step: model, hardware time per phase: plan · setup · work · check · and wall clock GET /api/runs/{id} → "costs"

Numbers known so far:

6 · Research tasks with BFF in progress

BFF (BayesicForceFields) is the group's open-source code for Bayesian learning of partial charges (Košťál, Shanks, Jungwirth, Martinez-Seara, JCTC (Journal of Chemical Theory and Computation) 2026). It fits force-field charges to ab initio molecular dynamics (AIMD) references through a Gaussian-process surrogate and MCMC (Markov chain Monte Carlo) sampling. The result is a probability distribution of charges, not a single best guess. Agents scouted the group's papers and the BFF code, proposed ideas, and had them critiqued by a simulated chemist and a compute engineer. The owner approved two, which are now being turned into clickable tasks. Each one runs in Harbor on Modal, so every stage is traced.

task A

Is 0.8 the right charge-scaling factor? A Bayesian answer with BFF

The group scales ionic charges by about 0.8 with the ECC (electronic continuum correction), to account for electronic polarisation. Here BFF learns acetate's charges at several ECC factors between 0.70 and 0.90. It then compares 0.80 against 0.75 with a Bayes factor: how much better the reference data supports one than the other.

task B

Ca2+–acetate binding with error bars: from the BFF charge posterior to experiment

BFF gives a range of plausible charges (the posterior), not one set. This task carries that uncertainty through umbrella-sampling PMFs (potentials of mean force: the free-energy profile of calcium binding to acetate). That gives a binding free energy with error bars, which it compares with Raman spectroscopy data.

Both: about 10–14 min and about $0.40 per run on a 32-CPU Modal box, with BFF and GROMACS preinstalled in the image, so the agent never spends time installing. The agent is told to poll, not sleep. Until they ship, the panel shows the two current tasks: "Methods card for a group paper" and "Ion–water binding, computed live" (DFT on Modal).

7 · Frankenstein: an agent that builds itself in progress (v3)

The hackathon brief: "By dawn, show a creature that learned to do things it could not do at dusk." Macrae notices a capability it is missing, creates it, tests it, installs it, and uses it again in later runs. Its capabilities may evolve; its authority may not. This is designed and queued behind v2. Nothing below is live yet. (SHA-256 = Secure Hash Algorithm, 256-bit: a fingerprint that changes if a single byte of the tool changes.)

1 · Gap the agent needs a toolthat isn't there and says CAPABILITY_GAP: name: why 2 · Create a sandbox step writesthe tool and its tests (Harbor on Modal) 3 · Test runs the tests;the step's check passes only if they all pass 4 · Install server checks SHA-256,tests and policy, then adds it to the registry 5 · Use again later runs get it in theirsandbox and instruction; no re-creating evolve: lessons + capability ledger feed the next run 🔒 Authority: fixed (server/policy.py, constants that no capability can change) • tools run only inside the Harbor sandbox • no secrets (environment allowlist) • no network hosts beyond a fixed list • no write access to the registry • a maximum runtime • asking for more → rejected The installer refuses any manifest that asks for more, and records a "rejected" event with the reason.

8 · Code map

folderwhat's in itstatus
data/, scripts/479 group publications (titles, authors, DOIs) from the IOCB sitedone
web/Chat with Jarvis: a thread with mic and composer, [n] citations opening source cards, typed questions that search the papers even without voice. Right panel "Tasks & runs": task cards, the live run timeline (📄🔎🧮✍️✅⚠️), counters, cited papers, recent runs. Offline state, light/dark, a bottom sheet on phones.live v2/v3 views
cloudflare/Worker macrae (index.js + worker.js): serves the page, sends /api through the Durable Object to the container, ElevenLabs signed URL, rate limits, admin status/restart, and the container's R2 endpoint. wrangler.toml, deploy.sh, a local dev server and a mock backend.live
server/FastAPI (a Python web framework): tasks, runs, trace events (tool calls → read / calc / write…), search, ElevenLabs tools, draining on restart.live v2: live, costs, planner
rag/PDF → clean pages (headers, hyphenation and reference lists removed) → chunks; BM25 + bge-small embeddings (BAAI General Embedding, by the Beijing Academy of Artificial Intelligence), fused ranking; [n] citations with pagebuilt needs PDFs
tasks/"Methods card for a group paper" and "Ion–water binding, computed live" (DFT with PySCF on Modal), plus Modal support in agent-runner. The two BFF tasks are being added.live BFF tasks
voice/ElevenLabs agent "macrae" (voice "Jarvis", Claude Sonnet 5.5 through ElevenLabs, cites every claim): three server tools and a setup scriptlive
deploy/Backend Dockerfile (Harbor uses Modal, so there is no Docker inside), start.py (supervises the server, drains runs on restart), r2sync.py (R2 mirror), the secret checker, docslive
evolve/Lessons distilled from traces, a tools library, a fine-tuning dataset export, then the capability registryv2 → v3
research/Scouting notes on the group's papers and the BFF code, behind the two research tasksdone

9 · Where it runs: all on Cloudflare

The page, the Worker and the backend live on one Cloudflare account and one address, deployed with cloudflare/deploy.sh (which runs wrangler deploy). The backend is a Cloudflare Container, built from the same Docker image, and the Durable Object MacraeBackend owns it. It sleeps after 2 hours without use, but never while a run is going. Its disk is temporary, so the paper index is baked into the image, and everything the backend writes is mirrored to R2. The container reaches R2 through the Worker at a private address (r2.macrae), so it needs no S3 (Simple Storage Service) keys. On a restart, task starts pause and running flows get up to 14 minutes to finish. Edits to the page alone don't restart the backend. The heavy work (agents, calculations) runs in Modal sandboxes. AWS was the first plan and is no longer used.

10 · How it was built

With agent-runner itself. A written contract (CONTRACT.md) fixed every interface. Then six Claude Opus 5.5 agents (high effort) built their modules at the same time across three Claude accounts, each in its own container. All six passed their checks on the first attempt.

moduletimecost (API-equivalent)
web + Worker18.5 min$5.24
server14.4 min$3.86
rag15.6 min$3.41
tasks + Modal16.5 min$4.95
voice10.2 min$2.99
deploy15.7 min$3.44
total~19 min wall clock$23.90, on Claude subscriptions

Later flows on 2026-10-08 used the same method (Claude Opus 5.5, high effort, through agent-runner and Harbor): an integrator that fixed the seams and tested end to end; the move of the backend into a Cloudflare Container with R2; the chat redesign; the BFF research ideas; and now v2, Frankenstein (v3) and the video.

11 · What the owner provides

itemstatus
ElevenLabs API keyprovided live
Claude login for the agentsprovided waiting a fresh key to replace it
Modal token (for the container)provided connected
Paper PDFswaiting indexed on the server, never published
Research tasksapproved two BFF ideas, being built
Hostinglive Cloudflare

Glossary

RAGRetrieval-augmented generation: find the relevant passages first, answer only from them, cite them.
WorkerSmall program on Cloudflare's network that serves the site and forwards API calls.
Durable ObjectA named Cloudflare object with its own storage. Here it owns the single backend container and decides when it sleeps.
ContainerThe backend's Docker image, run by Cloudflare on demand.
R2Cloudflare's object storage (files in a bucket). Holds the index, runs and traces.
HarborOpen-source harness that runs an agent in a sandbox and saves its trajectory, reward and files.
ModalCloud that starts sandboxes/containers on demand, CPU or GPU, billed per second.
Trace / trajectoryEvery message, tool call and result of one agent run (ATIF format in trajectory.json).
Signed URLA short-lived link (URL = uniform resource locator) that lets the browser open a voice session without seeing the API key.
BFFBayesicForceFields: Bayesian learning of force-field partial charges from reference simulations.
ECCElectronic continuum correction: scaling ionic charges (about 0.75–0.8) to mimic electronic polarisation.
CapabilityA tool the agent wrote, tested and installed for itself. Authority (what tools may do) stays fixed.