Macrae
Updated 2026-10-09 · final state · all on Cloudflare, agents on Modal, voice by ElevenLabs
A research agent for Pavel Jungwirth's group at the IOCB (Institute of Organic Chemistry and Biochemistry of the Czech Academy of Sciences, Prague). You chat with Jarvis, by voice or by typing, and it answers like a capable assistant, citing the group's papers whenever they are relevant. When you start a task, it plans the run, does real computational chemistry in cloud sandboxes, writes a short manuscript, and shows you every paper it read, every calculation it ran and what it cost. It learns from its own traces: it notices a capability it lacks, creates it, tests it, installs it and uses it again, while its authority stays fixed.
1 · The pieces
Everything public sits behind one Cloudflare Worker (a small program on Cloudflare's network) at one
address. The Worker serves the page and hands the browser a short-lived voice session. It forwards
/api calls to the backend, adding a secret header the browser never sees. The backend is a
Cloudflare Container, owned by a Durable Object: a named Cloudflare object that keeps exactly
one instance and decides when it sleeps. The backend answers paper questions from its own index and runs tasks
as agent-runner flows. Harbor starts each agent step in a sandbox on Modal and keeps the full
trace; the sandbox also streams every step back while it works. Runs, traces, the capability registry and the
paper index are mirrored to R2, Cloudflare's object storage. The typed chat, the planner and the lesson
distiller call Claude through the Anthropic API from the container.
Abbreviations in the diagram: API = application programming interface · LLM = large language model · RAG = retrieval-augmented generation (search the papers first, then answer with what was found, citing it) · BM25 = "Best Match 25", a keyword-ranking formula · vCPU = virtual processor core · GiB = gibibyte · CPU / GPU = central / graphics processing unit · PySCF = Python-based Simulations of Chemistry Framework · ATIF = Agent Trajectory Interchange Format (Harbor's trace file).
2 · Status
Checked on the live site on 2026-10-09 after the last deploy. "Live" means on the public site and checked there; "tested" means merged and covered by tests (Harbor, Modal and Claude faked), not re-run live; "not built" means planned only.
| state | what | wave |
|---|---|---|
| live | Site on one Worker: the page, the /api proxy and
/voice/signed-url. The backend is one Cloudflare Container (standard-1: ½ vCPU, 4 GiB) that sleeps
after 2 h unless a run is going and mirrors its state to R2. | v1 |
| live | Chat with Jarvis over 97 group papers (1210 chunks): answers
cite [n] with author, year and page and say which parts don't come from the papers, with the cost of
each answer (about $0.01). | w4 |
| live | Voice: ElevenLabs agent "macrae", voice "Jarvis", Claude Sonnet 5.5 through ElevenLabs, server tools as webhooks to the Worker, page tools in the browser. | v1 |
| live | Real agent runs on Modal through Harbor: the ion–water binding task. The first verified run found −25.5 kcal/mol for Na+–water at 2.2 Å (DFT, density functional theory, in PySCF, B3LYP/def2-TZVP with counterpoise correction). The same task started from the cloud container passed twice; those first runs reinstalled PySCF every time and hit a virtual environment without pip: the seed of what Macrae learned to fix. | v1 |
| live | Live traces (the sandbox streams every step within about 1 s), costs in money (section 5), a planner that decides hardware, stages and budget before each run, and lessons distilled from every trace (12 from the first runs). | v2 |
| live | Authority holds: the first forge built a pinned environment, but its tests failed in a fresh sandbox, so the installer rejected it and recorded why. | v3 |
| live | Dusk → dawn, the Frankenstein loop end to end. Dusk
(20261008-225728-dusk-f086, nothing learned): both runs passed, reward 1, 0 errors; both traces showed
the environment gap (no pip / ensurepip, PySCF reinstalled every run). Macrae started two forges by itself; each
wrote a pinned environment and its tests (H2 RHF/FCI against textbook values, H atom −0.5 Ha), passed
them in a fresh sandbox and was installed (calc-image, small-calc-env). Dawn
(20261008-232837-dawn-0ac3, the image + 8 lessons): both passed again, same answers
(Na+–water −25.46, K+–water −17.92 kcal/mol). Wall time 1035 → 690 s
(−33 %), setup 505 → 236 s (−53 %), installs 43 → 0 s, cost $1.44 → $1.04 (−27 %). The first dawn
attempt failed at the image build (the planner had set the image variable); the planner may no longer touch it.
One run per ion per side: a demonstration, not a statistic. | final |
| live | Public without login: rate limits, a daily spend cap, a kill switch, inputs that never reach a shell; the benchmark, forge and kill-switch routes need the operator secret. Downloadable proof: a trace zip and an offline HTML report for every run. | w4 |
| live | Frankenstein on the page: the capability ledger, installed capabilities and the fixed Authority card, the dusk → dawn table (Evolution tab), and the manuscript a finished run wrote. | final |
| tested | Methods card task: end-to-end tests with a fake Harbor, not re-run live after the last deploy. | v1 |
| not built | The two BFF research tasks (section 6): scouted and approved, not built. | — |
| owner | Paper PDFs stay private: they are indexed on the server and never published or committed. | always |
3 · What happens when you…
- Typed: the page sends the message, the anonymous session id and the recent history to
POST /api/chat. The backend searches the papers for every message and sends back a Claude answer with its cost (at most 20 answers a minute per address). - By voice: the mic button asks the Worker for a signed URL (a short-lived link; URL = uniform resource
locator) and the ElevenLabs session starts in the browser. The voice agent calls
search_papers, a webhook that goes to the Worker, then through the Durable Object to the container. - RAG returns the best passages, numbered
[1] [2], each with title, authors, page and DOI (Digital Object Identifier). - Jarvis chats normally and uses the passages when they help, naming the source ("Košťál 2026, page 3"). It says which parts come from the papers and which from general knowledge, and never invents a citation.
- Numbered source chips appear under the answer (
show_citationsduring a call). Clicking[n]opens that source's card.
- The page shows an estimate from past runs.
POST /api/tasks/{id}/startreturns arun_idat once, or "try again" when a rate limit, the cap on active runs or the daily spend cap is reached. - The planner (one Claude call, with the task, its inputs, the lessons and the installed capabilities) decides the stages, the hardware and a budget. The page shows that decision as the first card.
- Script steps pull paper passages from RAG. Agent steps run Claude Code in Modal sandboxes through Harbor, with the installed capabilities mounted and the lessons in their instruction.
- The sandbox streams each step back; the page polls
/api/runs/{id}/eventsevery 1–2 s, and each tool call becomes a timeline row (📄 read, 🔎 search, 🧮 calc, ✍️ write, ✅ result). The cost meter grows with it. - A
check:verifies the output (numbers inresult.json, a manuscript with math, a figure and citations) and gives a reward; a failed check feeds a retry. The state goes to R2 at every step. - When the run ends, Jarvis posts the outcome, its cost and the papers it cited, the page shows the
manuscript the agent wrote, and
evolvedistills lessons from the trace for the next run; a gap in the trace starts a forge.
4 · Where the traces are
A trace is everything one run did: the plan, every message, tool call and result, every file written,
every check, and what it cost. Anyone can read it on the page and download it; the owner also has the raw copy
in R2. The full guide is docs/TRACES.md in the repository.
| where | what you find | how to get it |
|---|---|---|
| 1 · The page timeline | One row per event (plan, status, read, search,
calc, write, think, cite, result,
error, and the capability events gap, create, test,
install, use, rejected), with a short human title, details, cost and
source cards. The server builds it from the run's state.json and step logs, the lines the sandbox
streams while a step runs, and each Harbor trial when it ends, without repeating anything seen live. The
Manuscript it wrote box shows the result when the run ends; /manuscript also serves the full
edit stream. |
Run view; GET /api/runs/{id}/events, /manuscript |
| 2 · Zip and report | A zip of every file behind the run: the flow's state, plan, costs, step logs, live stream lines, the Harbor
job folders with trajectory.json, the results and the manuscript. And a self-contained HTML report
that works offline: timeline, tool calls with their outputs, costs, manuscript, citations and checks. |
Download trace and Open report on each run; GET /api/runs/{id}/trace.zip |
3 · R2 bucket macrae-data |
state/runs/<run_id>/: state.json, step logs, live lines, plan, costs, outputs.state/jobs/<run_id>/…: Harbor trial folders, i.e. the full agent traces
(agent/trajectory.json in ATIF, agent/claude-code.txt, the reward, the files the
agent changed).state/macrae/: the server's event logs and run → task records.state/evolve/: what Macrae learned: lessons, the capability registry and ledger, benchmarks.index/: the paper index.The container pushes changes at every step and every 60 s, and restores them when it starts. At most about 60 s can be lost if a host dies without warning. |
Cloudflare dashboard or wrangler r2 (owner only) |
How the timeline fills in live: a wrapper around Claude Code inside the sandbox forwards each line of
its output to POST /api/live/{run}/{step}, at most about 1 s late, with a per-run token (never the
shared secret). The server stores these lines in <run>/live/<step>.jsonl and de-duplicates
them against the final trajectory.json. Traces also feed learning: lessons for the next run, and
python -m evolve export turns passing trajectories into a fine-tuning dataset.
5 · Costs
Every chat answer, run and step shows what it cost in money, split into LLM (large language model) dollars and
compute dollars, and how long it took. A run's meter grows live while it runs; the header keeps a total for your
session, with a small "what this cost" breakdown; and before you start a task the page estimates its cost and time
from past runs. The price and rate tables in server/costs.py cite their sources and dates.
Typical numbers:
- A chat answer: a fraction of a cent to a few cents, shown under the answer.
- The ion–water calculation: about 5–10 min of an agent on Modal and about $0.25–0.75, almost all of it LLM; with the lab notebook and manuscript it is at the top of that range.
- Guard rails: a global daily spend cap (LLM + Modal, $20 by default) with a clear message when it is reached, a cap on runs active at once, rate limits, and a budget per run set by the planner.
- Hosting: standard-1 container about $29/month if it never sleeps, about $4–5/month at about 4 h/day. That is on top of the $5 Workers Paid plan. R2 stays in the free tier. An open, visible browser tab checks health every 20 s and keeps the container awake.
- Building Macrae itself: see section 10 ($23.90 API-equivalent for the first six modules).
- With a Claude subscription login, LLM dollars are the API-equivalent cost of the same tokens, not a bill.
6 · Research tasks
Two tasks are on the site, each a flow of script steps on the backend and a Claude Code agent in a Modal sandbox:
Ion–water binding, computed live
Pick an ion (Li+, Na+, K+, Mg2+, Ca2+). The agent runs a
B3LYP/def2-SVP distance scan in PySCF and a counterpoise-corrected def2-TZVP energy at the minimum, and compares
it with full and ECC-scaled (×0.75) point charges. It keeps a lab notebook (NOTES_TO_SELF.md) and
writes a short manuscript through Write/Edit only: equations, a figure, and numbered citations of the group's
papers, drafted early and revised as numbers arrive. About 5–10 min and $0.25–0.75.
Methods card for a group paper
Give a DOI. The backend pulls that paper's passages from the index, and the agent writes a methods summary in
which every claim cites a passage [n]; a check rejects uncited claims.
Planned, not built: two tasks on BFF (BayesicForceFields, the group's code for Bayesian learning
of partial charges; Košťál, Shanks, Jungwirth, Martinez-Seara, JCTC 2026). Agents scouted the group's papers and the
BFF code (research/scout/) and the owner approved two ideas: "Is 0.8 the right charge-scaling factor? A
Bayesian answer with BFF" (acetate, ECC 0.70–0.90, a Bayes factor) and "Ca2+–acetate binding with error
bars" (umbrella-sampling free energies from the BFF charge posterior, compared with Raman data). The wave that was
to build them ran out of Claude session time.
7 · Frankenstein: an agent that builds itself
The hackathon brief: "By dawn, show a creature that learned to do things it could not do at dusk." Macrae notices a capability it is missing, creates it, tests it, installs it, and uses it again in later runs. Its capabilities may evolve; its authority may not. (SHA-256 = Secure Hash Algorithm, 256-bit: a fingerprint that changes if a single byte of the tool changes.)
- Capabilities are versioned, tested tools with a manifest: name, purpose, inputs/outputs, tests,
SHA-256 fingerprint, the run that created it, created/tested/installed/last-used times and a use count. The
registry lives under
evolve/capabilities/in the backend's state. - Three kinds. Tools (scripts, e.g. a radial distribution function or a block-averaged error bar from a molecular dynamics trajectory); environment capabilities (a Dockerfile fragment with pinned packages, created when the traces show slow installs or version breakage, tested by building it on Modal and running a smoke command, then used as the task's image, with setup time recorded before and after); and writing capabilities (style rules, templates).
- Notes to self. During a run the agent keeps
NOTES_TO_SELF.md: what it tried, what was slow, what to do differently.evolvereads it with the rest of the trace and writes lessons with evidence links. The Evolution tab shows them, run over run per task (setup time, total time, cost, errors). - The ledger:
GET /api/capabilitieslists the capabilities and every event (gap, create, test, install, use, rejected) between dusk and dawn. The Capabilities section of the Evolution tab shows it as a timeline, with the installed capabilities and the fixed "Authority" card. - Dusk → dawn benchmark, the proof. A fixed suite (the ion–water calculation for Na+ and
K+) runs twice: at dusk, with nothing learned injected or
mounted, and at dawn, with every installed capability and lesson. Per task it records success, reward,
wall time, setup time, LLM $, compute $, errors and the capabilities used; a registry snapshot taken at dusk
keeps the report reproducible.
GET /api/dawn-reportholds the comparison, and the Dusk → Dawn section of the Evolution tab shows it as a table with the deltas. Only the operator starts a benchmark:POST /api/benchmark/{dusk|dawn}with theX-Macrae-Adminheader (the Worker secretMACRAE_ADMIN_SECRET, which the voice agent never has); the same holds for a hand-started forge and the kill switch. - Raw traces: each run has "Download trace" (a zip) and "Open report" (offline HTML).
8 · Code map
| folder | what's in it |
|---|---|
data/, scripts/ | 479 group publications (titles, authors, DOIs) from the IOCB site. |
web/ | Chat with Jarvis: a thread with mic and composer, answers with
[n] citations opening source cards, a session cost total. Right panel with two tabs. Tasks &
runs: tasks with cost estimates, the current run (plan card, cost meter, live timeline 📄🔎🧮✍️✅⚠️, Download
trace / Open report, the manuscript it wrote with math and figure). Evolution: lessons run over run, the
capability ledger with the Authority card, Dusk → Dawn. Offline state, dark mode, a bottom sheet on phones. Plain
HTML/CSS/JS, no build step. |
cloudflare/ | Worker macrae (index.js + worker.js):
serves the page, sends /api through the Durable Object to the container, ElevenLabs signed URL, rate
limits, operator-only routes (benchmark, forge, kill switch, admin status/restart), the live-trace ingest route,
and the container's R2 endpoint.
wrangler.toml, deploy.sh, a local dev server and a mock backend. |
server/ | FastAPI (a Python web framework): tasks, runs, trace events (tool calls → read / calc / write…), live ingest, chat, search, ElevenLabs tools, costs, the planner, capabilities and the fixed policy, manuscripts, trace downloads, rate limits, spend cap and kill switch, draining on restart. |
rag/ | PDF → clean pages (headers, hyphenation and reference lists removed) → chunks;
BM25 + bge-small embeddings (BAAI General Embedding, by the Beijing Academy of Artificial Intelligence), fused
ranking; [n] citations with page. |
tasks/ | The two tasks and their flows (methods card, ion–water DFT), preparation and
check scripts, the research protocol (notebook + manuscript), the calc-image capability seed, and
tasks/common/: the Modal agent image with the live wrapper. |
evolve/ | Lessons distilled from traces, the capability registry, the dusk → dawn benchmark, run-over-run metrics, a fine-tuning dataset export. |
agent_runner/ | Runs flows of agent and script steps through Harbor (Docker or Modal), with checks, retries and several Claude accounts. An attempt that errored (for example on a Claude rate limit) never counts as passed. |
voice/ | ElevenLabs agent "macrae" (voice "Jarvis", Claude Sonnet 5.5 through ElevenLabs): server tools, page tools and a setup script. |
deploy/ | Backend Dockerfile (Harbor uses Modal, so there is no Docker inside),
start.py (supervises the server, drains runs on restart), r2sync.py (R2 mirror,
including what Macrae learned), the secret checker, docs. |
proof/ | Saved live payloads (runs, lessons, capability ledger, dusk → dawn report) behind the numbers on this page and in the README. |
tests/ | Cross-module tests: RAG → flow → API integration, an end-to-end run with a fake Claude and Harbor, and checks of these docs. |
research/ | Scouting notes on the group's papers and the BFF code; a real run kept as a
replayable bundle (demo/small-calc-na/). |
docs/ | This page, the presenter's demo script and the traces guide. |
9 · Where it runs: all on Cloudflare
The page, the Worker and the backend live on one Cloudflare account and one address, deployed with
cloudflare/deploy.sh (which runs wrangler deploy). The backend is a Cloudflare
Container, built from the same Docker image, and the Durable Object MacraeBackend owns it. It
sleeps after 2 hours without use, but never while a run is going. Its disk is temporary, so the paper index is
baked into the image, and everything the backend writes is mirrored to R2. The container reaches R2 through
the Worker at a private address (r2.macrae), so it needs no S3 (Simple Storage Service) keys. On a
restart, task starts pause and running flows get up to 14 minutes to finish, but a deploy that changes the image
makes Cloudflare replace the container without waiting, so deploys happen when no run is active. Edits to the page
alone don't restart the backend. The heavy work (agents, calculations) runs in Modal sandboxes. AWS (Amazon Web Services) was the first
plan and is no longer used.
Open to everyone, no login. Each visitor gets an anonymous session id. Task starts are rate-limited per IP (Internet Protocol) address in the Worker and per address, per session, for the voice agent and globally in the backend, with a cap on runs active on Modal at once (4); chat answers are limited to 20 a minute per address. A daily spend cap ($20 by default) stops new paid work when reached, and the operator has a secret-protected kill switch. Every route validates its input, task inputs come from allowlists and never reach a shell, and paper text in tool outputs is handled as data, not as instructions. No secret reaches the browser, and the sandbox gets only what its step needs.
10 · How it was built
With agent-runner itself. A written contract (CONTRACT.md) fixed every interface. Then six Claude
Opus 5.5 agents (high effort) built their modules at the same time across three Claude accounts, each in its own
container. All six passed their checks on the first attempt.
| module | time | cost (API-equivalent) |
|---|---|---|
| web + Worker | 18.5 min | $5.24 |
| server | 14.4 min | $3.86 |
| rag | 15.6 min | $3.41 |
| tasks + Modal | 16.5 min | $4.95 |
| voice | 10.2 min | $2.99 |
| deploy | 15.7 min | $3.44 |
| total | ~19 min wall clock | $23.90, on Claude subscriptions |
The later waves used the same method (Claude Opus 5.5 through agent-runner and Harbor, one git branch or folder per agent, then an integrator agent that merges, fixes the seams and tests end to end):
- Integration and Cloudflare: the first integrator; the backend moved into a Cloudflare Container with R2; the chat redesign.
- Research: agents scouted the group's papers and the BFF code, proposed ideas, a simulated chemist and a compute engineer critiqued them, and the owner approved two.
- v2: live traces, costs, the planner, evolution. v3: Frankenstein (capabilities, fixed authority, dusk → dawn, the research protocol and manuscript), the RAG on the group's papers.
- Wave 4: ten agents in parallel on their own branches (chat answers, browser end-to-end tests, concurrency, costs for users, trace downloads, voice, public safety, the proof pack, design polish, docs). Many of them, and the integrator, hit the Claude session limit.
- Final integration (2026-10-09, night, by hand with Claude Code). The v3 apply had dropped the tasks
module (
checks.py,research.py,calc-image) while keeping the flows that call it, so every live task would have failed its check, and it had reverted the typed chat (/api/chat); both were restored. Four finished w4 branches (docs, costs, trace download, public safety) were merged; the other w4 jobs left nothing. agent-runner'suntil:had let rate-limited attempts count as passed; fixed. Three safety gaps were closed: the operator routes were open to any visitor, the chat had no limits, and what Macrae learned was not synced to R2. The v3 page (capabilities, dusk → dawn, manuscript) was built then, and the dusk → dawn benchmark ran on the live site.
11 · What the owner provides
| item | where it lives |
|---|---|
| ElevenLabs API key and agent id | Worker secrets; never sent to the browser |
| Claude logins for the agents, and an Anthropic API key for chat, planner and distiller | Worker
secrets, passed to the container; rotate with wrangler secret put and a restart |
| Modal token | Worker secret, passed to the container |
| Operator secret (benchmark, forge, kill switch) | Worker secret MACRAE_ADMIN_SECRET, kept
locally in deploy/.env; never given to the voice agent |
| Paper PDFs | indexed on the server, never published or committed |
| Research questions | approved by the owner: the two BFF tasks (not built yet) |
| Hosting | Cloudflare (Workers Paid), Modal (per second), ElevenLabs |
Glossary
| RAG | Retrieval-augmented generation: find the relevant passages first, give them to the model, and cite them in the answer. |
| Worker | Small program on Cloudflare's network that serves the site and forwards API calls. |
| Durable Object | A named Cloudflare object with its own storage. Here it owns the single backend container and decides when it sleeps. |
| Container | The backend's Docker image, run by Cloudflare on demand. |
| R2 | Cloudflare's object storage (files in a bucket). Holds the index, runs and traces. |
| Harbor | Open-source harness that runs an agent in a sandbox and saves its trajectory, reward and files. |
| Modal | Cloud that starts sandboxes/containers on demand, CPU or GPU, billed per second. |
| Trace / trajectory | Every message, tool call and result of one agent run (ATIF format in trajectory.json). |
| Signed URL | A short-lived link (URL = uniform resource locator) that lets the browser open a voice session without seeing the API key. |
| BFF | BayesicForceFields: Bayesian learning of force-field partial charges from reference simulations. |
| ECC | Electronic continuum correction: scaling ionic charges (about 0.75–0.8) to mimic electronic polarisation. |
| Capability | A tool the agent wrote, tested and installed for itself. Authority (what tools may do) stays fixed. |
| Dusk / dawn | Two states of the same benchmark: dusk with nothing learned, dawn with every installed capability and lesson. |
| Lesson | A short do/avoid/setting/tool rule distilled from a trace, with links to the evidence, given to the next run of the task. |
| Planner | One Claude call before each run that picks the stages, the hardware and a budget, and says why. |
| Manuscript | The short paper a research run writes as it goes (results/manuscript.md), only through Write and Edit, so the trace holds every revision; the page shows the final text. |