CaseStudy
- title
- summary
- role
- status
- stack
- sections
- may
- quote verbatim · summarize · link to sections
- may not
- invent metrics · speak for past employers · expose gated content
mRNA PK/PD Engine
A multi-source PK/PD composition framework for mRNA drugs. The goal is to model the body as an 82-organ environment so a research scientist can see what affects a drug at each point and target binding rationally. It models 12 organs today, with 46 composition functions built from more than 60 published sources.
Research, not yet validated. This is research infrastructure rather than a product. A biochemist has reviewed parts of the chemistry and biology and confirmed the approach is heading in the right direction. A full review of the chemistry, and a physicist's review of the physics math, are still to come. The engine sanity-checks its output against published data and flags anything outside that range, so researchers know when a result isn't backed by evidence. I'm talking with universities about validation and facility-side testing.
Summary
- problem
- approach
- status
- may
- quote verbatim
- may not
- round or extrapolate results
At a glance
- Problem
- Existing PK/PD models collapse the body into a few compartments and fit parameters to a single cohort, so researchers can't see what affects a drug in each organ.
- Approach
- A composition framework that treats the body as one environment: 12 organs today, 82 as the goal. Every rate constant is computed from its inputs and cites a source, and a 156-state ODE system solves the whole body at once.
- Status
- Research infrastructure, not yet fully validated. Against a panel of 12 published datasets (human, mouse, and rat), 15 of 18 gates pass, with a 2.14× geometric-mean fold error and no fitting to the panel.
12 organs today, with 82 organs as the goal. 46 composition functions from 60+ published sources feed a 156-state ODE system, and 15 of 18 validation gates pass at a 2.14× fold error.
Problem
Research scientists can't see what affects the drug at each point.
When an mRNA-LNP therapeutic enters the body, it passes through dozens of organs and hundreds of cell types, meets thousands of proteins, and goes through a sequence of degradation and binding events that decide whether any of it reaches the target tissue intact. Existing PBPK models collapse this into a few compartments and fit parameters to the cohort they were published against. What they describe is an average rat or an average cohort. They can't tell a researcher what is happening to their drug in a given organ at a given time.
Researchers who want to target drug binding more rationally, improve absorption, or understand why a candidate failed in one organ but not another don't have a tool built for that question. What they have are published models fit to someone else's cohort.
Vision
Modeling the body as 82 organs
The goal is to model the body as a single environment of 82 organs, and to show what's affecting a drug in each one as it moves through. A researcher running the simulation could see where degradation speeds up or uptake stalls, where binding saturates, and where local conditions change the outcome. From there they can adjust construct design, vehicle chemistry, or dosing strategy against the mechanism itself instead of tuning to aggregate outputs.
Current state
What's built so far
12 organs
Arterial and venous blood, lung, heart, liver, spleen, kidney, muscle, and lymph nodes, plus grouped compartments for the portal and remaining organs. The liver is resolved into hepatocytes and Kupffer cells.
46 composition functions
Each rate constant is computed from its driver inputs at the interaction point, never looked up from a fitted table.
60+ published sources
Each contributes equations, parameters, or validation data, with provenance tracked down to the line.
156 ODE states
Drug, vehicle, immune, and hormone states, integrated with stiff solvers (LSODA, BDF).
~78k lines of Python, 1,900+ tests
122 test modules, including mass balance to machine precision.
Each rate constant is computed at the (organ × vehicle × construct × cell-type × patient) interaction point, and module-level constants are rejected as constants in disguise.
Beyond drug and vehicle kinetics, the engine now carries LNP chemistry as an input a scientist can swap (SM-102, MC3, ALC-0315, BiP-20), an anti-PEG antibody module for repeat dosing, and an optional hormone layer: the stress, thyroid, and reproductive axes plus glucose and insulin, so the same drug can be run in different bodies. A newer ingestion pipeline pulls PK data from FDA labels and ClinicalTrials.gov, with every value tied back to its source.
How it works
Two funnels, one solver, and checks on every run
Funnel 1, construct reads the mRNA itself (codon usage, GC content, modified-base chemistry, UTR structure) and the vehicle's lipid chemistry. It outputs translation rate, mRNA degradation, and stability.
Funnel 2, patient builds physiology from patient inputs: organ volumes and masses, tissue blood flows, hepatic and renal function, age, sex, weight, and APOE genotype.
The solver integrates the 156-state ODE system forward in time, using stiff solvers for the large flow and volume differences PBPK systems carry. Changing the construct changes the drug and changing the patient changes the body, but the solver stays the same.
The contract. Every composition function ships with a test_value_varies_with_<input> test. If a function returns the same value when its inputs change, the test fails and the code doesn't merge.
Discipline
How the engine limits its own claims
15 of 18 gates pass against 12 published datasets, at a 2.14× fold error and with no fitting to the panel.
A validation panel, not one study
Human (An 2024 mRNA-3927, patisiran, mRNA-1944), mouse and rat, vaccine biodistribution, and five model cross-checks.
Failures are reported
Three gates fail today, and the panel says which ones and why.
Uncertainty on every number
Monte Carlo ensembles turn point estimates into ranges, and weakly sourced values are ranked as the biggest contributors.
Out-of-calibration flags
A patient outside the published cohorts raises a flag and needs an explicit acknowledgment to run.
Hypothetical labels
Untested combinations are labeled on the output, in filenames, and on plots.
The validation panel. Published cohorts are used to check the engine, never to fit it. The panel spans human first-in-human mRNA-LNP PK (An et al. 2024 in Nature), siRNA-LNP PK (patisiran), a second mRNA program (mRNA-1944), mouse and rat studies, vaccine biodistribution, and five model cross-checks.
What still fails. The terminal half-life and AUC in An 2024 are under-predicted, one mouse mRNA half-life is off by about 1.3×, and spleen and lung uptake are too low relative to the liver. That last gap comes from how the engine resolves the liver into cell types but not yet the other immune organs, and splitting them is the next architectural step. Peak concentration in An 2024 sits 2.4–4.6× high on the panel.
Evidence tiers. Every value is tagged primary, derived, or placeholder. Placeholders carry wider uncertainty, and the dossier ranks which of them matter most, so it's clear which measurement would improve a prediction the most.
Physics checks. Mass balance holds to machine precision (relative error around 10−15), and a dedicated test guards against a past bug where the dose was destroyed too early. Predictions outside what's physically possible for a human body are flagged.
AI orchestration
Working with Claude Code
I built this solo, with Claude Code as my main collaborator. The setup uses specialized sub-agents (Explore, Plan, security-review) and a hand-written CLAUDE.md that started as a three-paper extraction guide and grew into architecture memory covering more than 60 sources and over 400 commits.
Skills for recurring mistakes. Bespoke skills handle paper-access discipline (three-location search before declaring a data blocker), constants-in-disguise review, the silent-uniform-broadcast detector, and the stream-timeout split discipline for large structural commits. Each skill blocks a failure mode that kept coming back, or makes it fail loudly.
Findings documents. Every architectural session produces a findings document committed alongside the code. Phase 4d produced five lessons saved as project memory: silent broadcasts, constants-in-disguise, lazy-resolver-vs-typed-driver-record, search-before-PLACEHOLDER, stream-timeout. The next phase built on those lessons.
Every constant is sourced. Every numeric constant in the engine cites a line from a source paper. The composition functions ship with test_value_varies_with_<input> contracts: if a function returns the same value when its inputs change, the test fails and the value doesn't merge. I apply the same rule to design-system tokens at Group 1001.
What's next
What comes next
Every organ at the cellular level
The liver already runs as hepatocytes and Kupffer cells, sized by cell count × organ mass. The same template is being extended to the spleen, lung, heart, and kidney.
Chemist validation
A biochemist reviewed parts of the chemistry and biology and confirmed the approach is on the right track. A full review of every composition function comes next.
Physicist validation
Review of the physics math, including the stiff solvers' behaviour.
University research facility
Hosting the engine after validation, to test it against fresh in-vivo and in-vitro data.
Resolving every organ to cell types. The liver already runs as hepatocytes and Kupffer cells, sized by cell count × organ mass. The same template is being extended to the spleen, lung, heart, and kidney, each split into resident macrophages and parenchyma. That will move biodistribution from liver-heavy to physiological, make uptake capacity mechanistic in every organ, and add toxic ceilings, so the engine predicts a dose-limiting therapeutic window as well as efficacy.
Before a research lab can rely on the engine, it needs these three kinds of validation. All three are in progress. A biochemist has already reviewed parts of the chemistry and biology and confirmed the approach is heading in the right direction.
A full chemistry review would confirm that the relationships the engine encodes (ApoE binding, endosomal escape, lipid-chemistry effects, construct stability) match how a chemist would describe them. A physicist would review ODE topology, mass balance discipline, dimensional analysis, and stiff-solver behaviour at the time-scales that matter biologically. I'm in conversation with universities about the facility, so the engine can be exercised against fresh data rather than only against published cohorts.
Scope
What the engine is not for
Not a production tool
Not for use in actual drug production, clinical trial design, or regulatory submission.
Not a substitute for clinical trials
It predicts under its assumptions and sources. Clinical trials remain the only path to claims about drug behaviour in humans.
Not fully chemist-validated
A biochemist has confirmed the direction on parts of it. Until the full chemistry and biology review is done, output is sanity-checked composition.
Not physicist-validated
The physics math needs expert review for the same reason.
Not a black box
Every constant traces to a source paper, and each dispatch table shows its provenance.
Outputs that land outside published or physically possible ranges are flagged automatically.