What did the Dyna-1 paper (Nature 2026, Scripps Research) sho…
What did the Dyna-1 paper (Nature 2026, Scripps Research) show about predicting millisecond protein dynamics from the residues that are missing in NMR spectra? Explain the idea for a general scientific reader, how it was validated against relaxation-dispersion experiments, how well it did on functional regions such as KRAS, and what the honest limitations are. Cite the paper and primary sources.
Research session
What did the Dyna-1 paper (Nature 2026, Scripps Research) show about predicting millisecond protein dynamics from the residues that are missing in NMR spectra? Explain the idea for a general scientific reader, how it was validated against relaxation-dispersion experiments, how well it did on functional regions such as KRAS, and what the honest limitations are. Cite the paper and primary sources.
Dyna-1 turns a longstanding nuisance in protein NMR — residues that simply never show up in the spectrum — into a usable, learnable signal for where a protein is moving on the microsecond-to-millisecond timescale, and a deep-learning model reproduces that signal well enough to flag functionally important motions like the switch regions of K-Ras WEB pmc.ncbi.nlm.nih.gov-aHR0cHM6Ly9wbWMubmNiaS5u. The work is from Dorothee Kern's lab at Scripps Research (Howard Hughes Medical Institute) with Sergey Ovchinnikov (MIT) and Hannah Wayment-Steele; the version I read is the bioRxiv preprint (v3), which notes a peer-reviewed Nature version at doi 10.1038/s41586-026-10989-4 WEB pmc.ncbi.nlm.nih.gov-aHR0cHM6Ly9wbWMubmNiaS5u.
A note on sourcing: all quantitative and methodological statements below are traceable to the full text of the bioRxiv v3 preprint on PMC WEB pmc.ncbi.nlm.nih.gov-aHR0cHM6Ly9wbWMubmNiaS5u. I attempted to open the peer-reviewed Nature article directly to confirm final published numbers but the publisher link could not be read in a bounded way, so figures cited here are from the preprint the Nature version supersedes; treat small numeric details as preprint-stage until checked against the final paper.
The core idea, for a general reader
Many proteins do their jobs by flickering between shapes — an enzyme opening and closing over a substrate, a signaling protein toggling between "on" and "off." These interconversions on the microsecond-to-millisecond (µs–ms) range are central to catalysis, allostery, and ligand binding, but they have been hard to predict because there was never a large, standardized dataset of them to learn from — unlike the Protein Data Bank that made AlphaFold possible WEB pmc.ncbi.nlm.nih.gov-aHR0cHM6Ly9wbWMubmNiaS5u.
The paper's key insight is that such a dataset already existed, hidden in plain sight. In NMR, a residue exchanging between two chemical environments on the µs–ms timescale can have its signal broadened so much that it disappears — the peak never gets "assigned." Across the ~10,000 proteins with deposited chemical shifts in the Biological Magnetic Resonance Data Bank (BMRB), many residues are simply missing from the assignment lists. The authors made the deliberately "bold assumption" that residues missing an assignment are largely missing because they are exchange-broadened by µs–ms motion, and trained deep-learning models to predict which residues are missing WEB pmc.ncbi.nlm.nih.gov-aHR0cHM6Ly9wbWMubmNiaS5u. The striking result is that a model trained purely on "is this residue assigned or not?" also predicts exchange that was independently measured by dedicated relaxation experiments — i.e., it learned real dynamics, not just an assignment artifact WEB pmc.ncbi.nlm.nih.gov-aHR0cHM6Ly9wbWMubmNiaS5u.
The best model, Dyna-1, is built on an intermediate layer (layer 22) of the multimodal protein language model ESM-3, taking both sequence and structure as input WEB pmc.ncbi.nlm.nih.gov-aHR0cHM6Ly9wbWMubmNiaS5u. It outputs, per residue, a probability of being exchange-broadened/missing — p(missing) — which is used as a probability of µs–ms exchange, p(exchange) WEB pmc.ncbi.nlm.nih.gov-aHR0cHM6Ly9wbWMubmNiaS5u.
A supporting observation ties this to biology: residues with µs–ms exchange are more evolutionarily conserved than average, and — tellingly — residues with missing assignments are statistically indistinguishable in conservation from residues with measured exchange, supporting the assumption that the "missing" signal really is dynamics WEB pmc.ncbi.nlm.nih.gov-aHR0cHM6Ly9wbWMubmNiaS5u.
How it was built and validated
The authors assembled two datasets WEB pmc.ncbi.nlm.nih.gov-aHR0cHM6Ly9wbWMubmNiaS5u:
- RelaxDB: 133 curated single-domain ¹⁵N backbone relaxation datasets (from an original 163), with a fitting framework designed to separate true µs–ms exchange from anisotropic-tumbling artifacts WEB pmc.ncbi.nlm.nih.gov-aHR0cHM6Ly9wbWMubmNiaS5u.
- mBMRB: 9,381 proteins from the BMRB labeled simply by which backbone amides are missing an assignment — roughly two orders of magnitude more data than RelaxDB, used for training WEB pmc.ncbi.nlm.nih.gov-aHR0cHM6Ly9wbWMubmNiaS5u.
Validation was layered:
-
Predicting missing assignments (the training task) reached an AUROC of ~0.77 on the validation set for the ESM-3 layer-22 model; a last-layer ESM-2 model (~0.75) and an AlphaFold2 pair-representation model using only structure (~0.71) were close behind, indicating both sequence and structure information help WEB pmc.ncbi.nlm.nih.gov-aHR0cHM6Ly9wbWMubmNiaS5u.
-
Transfer to real exchange. The crucial test: hold out the missing-assignment residues and ask whether the model predicts exchange in residues that were assigned but show elevated R₂ (Rₑₓ) from relaxation data. Performance drops relative to the missing-assignment task but stays well above controls — the model generalizes from "invisible" residues to measurable exchange WEB pmc.ncbi.nlm.nih.gov-aHR0cHM6Ly9wbWMubmNiaS5u. Overall AUROC on RelaxDB was 0.63, rising to 0.66 after removing proteins whose apparent exchange was a phosphate-buffer artifact WEB pmc.ncbi.nlm.nih.gov-aHR0cHM6Ly9wbWMubmNiaS5u.
-
CPMG relaxation dispersion, the most direct experiment for µs–ms exchange. On six proteins with complete ¹⁵N CPMG data (cyclophilin A, TEM-1 β-lactamase, adenylate kinase, biliverdin reductase B, K-Ras, and DUSP3), Dyna-1's predictions agreed with the measured dispersion, and adding labels for "unsuppressed Rₑₓ" improved AUROC for 5 of the 6 proteins WEB pmc.ncbi.nlm.nih.gov-aHR0cHM6Ly9wbWMubmNiaS5u. In cyclophilin A specifically, Dyna-1 flagged Arg148, which conventional CPMG processing missed because its exchange is too fast to be suppressed by the pulse train; re-analysis of the raw data supported the prediction — a case where the model corrected the standard analysis rather than merely matching it WEB pmc.ncbi.nlm.nih.gov-aHR0cHM6Ly9wbWMubmNiaS5u. It also recovered the known catalytically essential β-sheet network in CypA (Ser99, Phe113, Met61, Arg55) WEB pmc.ncbi.nlm.nih.gov-aHR0cHM6Ly9wbWMubmNiaS5u.
-
A genuinely prospective experiment. After the model was publicly released (and after initial manuscript review), the authors ran new ¹⁵N CPMG on two proteins chosen from the test set — Chitinase 19 (Chi19, a 205-residue chitinase, predicted high exchange) and the ~70-residue hypothetical protein yjbJ (PDB 1RYK, predicted low). The chitinase showed exchange concentrated in its active site and chitin-binding loops (AUROC 0.62); yjbJ showed dispersion only at two N-terminal residues (AUROC 0.72) WEB pmc.ncbi.nlm.nih.gov-aHR0cHM6Ly9wbWMubmNiaS5u. This is the strongest form of validation because the predictions were fixed before the data existed.
Functional regions and K-Ras
The paper's biological headline is that dynamics linked to function — catalysis and ligand binding — are predicted especially well, mirroring the conservation trend WEB pmc.ncbi.nlm.nih.gov-aHR0cHM6Ly9wbWMubmNiaS5u. Concretely:
- In K-Ras, Dyna-1 predicts concerted millisecond dynamics in the switch I and switch II regions — the segments that toggle K-Ras between signaling states and are central to its oncogenic biology WEB pmc.ncbi.nlm.nih.gov-aHR0cHM6Ly9wbWMubmNiaS5u. Importantly, this is a retrospective match to previously published CPMG data, not a blind prediction. The measured exchange it was tested against comes from Hansen et al. (Nat. Struct. Mol. Biol., 2023), who characterized active GTP-bound wild-type K-Ras plus the oncogenic G12D and G12C P-loop mutants across picosecond-to-millisecond timescales, resolving the Switch I/II regions that had been largely unobservable by X-ray and prior NMR PUBMED 37640864.
- In BPTI, Dyna-1 reached AUROC 0.92, matching the known disulfide-isomerization dynamics and the distinct kinetic substates seen in long molecular-dynamics simulations WEB pmc.ncbi.nlm.nih.gov-aHR0cHM6Ly9wbWMubmNiaS5u.
- Broken down by structural context, Dyna-1 is strongest in core helices and surface loops, and stronger for more conserved residues — i.e., exactly the regions where µs–ms motion tends to be functional WEB pmc.ncbi.nlm.nih.gov-aHR0cHM6Ly9wbWMubmNiaS5u.
The blind prospective tests were the chitinase and yjbJ, not K-Ras; the K-Ras and other well-known cases are retrospective validations against existing literature data WEB pmc.ncbi.nlm.nih.gov-aHR0cHM6Ly9wbWMubmNiaS5u.
Honest limitations
The authors are candid about several WEB pmc.ncbi.nlm.nih.gov-aHR0cHM6Ly9wbWMubmNiaS5u:
- Noisy labels. The training "negatives" are not clean: many residues that are assignable also experience exchange, so labeling them "no dynamics" is often wrong. The signal is learnable partly because label noise can act as regularization, but it caps how good the ground truth is.
- Modest absolute performance. On the real relaxation benchmark the AUROC is 0.63–0.66 — clearly better than chance and biologically meaningful, but far from the near-saturated accuracy AlphaFold reached for static structure. Some individual proteins even score below 0.5.
- Buffer and artifact confounds. Apparent exchange from non-specific phosphate-buffer binding at charged surface residues degraded performance until those proteins were flagged; Dyna-1 correctly did not predict such artefactual "exchange," but it shows the experimental ground truth itself is contaminated.
- Missing assignments are only a partial proxy. Only a subset of exchanging residues broaden beyond detection; using RelaxDB counts as a rough guide, missing assignments capture roughly half of the residues with exchange, so the primary training label is incomplete by construction.
- Scope of the model. Dyna-1 was trained on proteins under 400 residues, so larger proteins are outside its validated range; ESM-3, its backbone, was trained without side-chain information, which the authors expect limits performance. More data alone gave only marginal gains — the stringent 30%-identity/0.5-TM-score split performed nearly as well as the lenient 80%-identity/1.0-TM split — suggesting further progress needs architectural innovation, not just scale.
- What it does and doesn't predict. Dyna-1 classifies whether a residue undergoes µs–ms exchange; it does not by itself give rates, excited-state populations, or the structures of the alternative conformations. External conditions (temperature, pH, cofactors) change rates and populations but not the underlying process the model flags.
- ¹⁵N-only evaluation. The CPMG benchmark used ¹⁵N data, so exchange visible only in ¹H is not counted — meaning true performance may be somewhat underestimated.
Bottom line
Dyna-1 is best understood as an early, more modest "dynamics" analogue to what AlphaFold did for structure: it defines a task (classify µs–ms exchange), mines an overlooked large-scale dataset (missing NMR assignments), and shows a language-model-based predictor can recover functionally important motions — validated retrospectively against relaxation dispersion and, importantly, in a genuine prospective NMR experiment WEB pmc.ncbi.nlm.nih.gov-aHR0cHM6Ly9wbWMubmNiaS5u. Its predictions for functional hotspots like the K-Ras switch regions are compelling, but the honest headline is a useful classifier (AUROC roughly 0.6–0.9 depending on the system) that flags where slow dynamics live, not a solved problem or a generator of full conformational ensembles with rates and populations WEB pmc.ncbi.nlm.nih.gov-aHR0cHM6Ly9wbWMubmNiaS5u.
Sources
- Wayment-Steele, El Nesr, Hettiarachchi, Ojoawo, Kariyawasam, Ovchinnikov, Kern. "Learning millisecond protein dynamics from what is missing in NMR spectra." bioRxiv (v3, 2026), doi 10.1101/2025.03.19.642801; peer-reviewed version Nature (2026), doi 10.1038/s41586-026-10989-4 WEB pmc.ncbi.nlm.nih.gov-aHR0cHM6Ly9wbWMubmNiaS5u
- Hansen, Xiang, Yuan, Bruschweiler-Li, Brüschweiler. "Excited-state observation of active K-Ras reveals differential structural dynamics of wild-type versus oncogenic G12D and G12C mutants." Nat. Struct. Mol. Biol. 30, 1446–1455 (2023) — the K-Ras relaxation-dispersion dataset Dyna-1 was tested against PUBMED 37640864
This is research context, not a substitute for reading the primary methods; if you need the exact final-published AUROC values or the supplementary methods, I can try a different route to the Nature version's full text on request.