Multi-Omics: Plasma Proteins Predict 17 Disease Risks
- Aug 5
- 4 min read
Risk calculators built from age, blood pressure, and a lipid panel still miss many people who develop disease within a decade. A multi-omics analysis of 23,776 UK Biobank participants tested how far plasma profiling closes that gap, measuring 159 NMR metabolites and 2,923 Olink protein targets in the same individuals across 17 incident diseases. Adding omics improved prediction for every endpoint. The more useful result is which layer did the work.
Key Takeaways
Plasma proteomics added more predictive value than NMR metabolomics for 16 of the 17 incident diseases in this 23,776-person UK Biobank cohort.
Multi-omics models built from proteins and metabolites together did not beat proteomics alone, likely because the extra features cost degrees of freedom without adding signal.
Metabolite signatures converged on shared lipid pathways, while protein signatures separated cleanly by disease and organ system.
For biomarker programs, the practical read is to budget the protein assay first, then add metabolomics where it answers a mechanistic question.
What Multi-Omics Adds Beyond Clinical Risk Factors
Three nested baselines framed the comparison: age plus sex; ASCVD, the standard cardiovascular risk variables; and PANEL, a wider set adding demographics, lifestyle, and routine laboratory measurements. Cox models were fitted on each baseline first, and the martingale residuals became the prediction target for the omics data, so the molecular layers were credited only with what the clinical variables left unexplained. That design choice is what makes the comparison worth reading. Training ran through mixOmics with five repeats of five-fold stratified cross-validation, scored by Harrell's C-index. Mean gains reached 0.063, 0.038, and 0.030 over the three baselines (p below 1E-6).
Key Findings
Proteomics carried most of the signal: under the PANEL baseline, protein profiling beat metabolite profiling for 16 of 17 endpoints, glaucoma the lone exception. Prostate cancer moved furthest, C-index 0.66 to 0.75.
Combining both layers gained almost nothing: combined models averaged a C-index decrement of 0.002 against proteomics-only under PANEL, though they did help for peripheral artery disease, COPD, renal disease, and fractures.
A novel skin cancer protein surfaced. Mendelian randomization put PRG3 at an inverse-variance weighted OR of 0.91 (95% CI 0.89 to 0.93, p = 1.04E-16), higher levels tracking lower risk.
The two layers describe different biology: top metabolites returned lipid-centric KEGG pathways shared across traits, while top proteins split into disease-specific clusters (MACE and CHD correlating at 0.587).

Figure 1. Study overview for the multi-omics prediction framework. Panel A lays out the three nested baseline predictor sets: age plus sex, the ASCVD cardiovascular variables, and the wider PANEL set spanning demographics, lifestyle, and laboratory measurements. Panel B traces the workflow across 23,776 UK Biobank participants, from Cox residualization through repeated cross-validation with mixOmics to C-index evaluation. Adapted from Du et al. (2026), Nature Communications.
Why Proteins Outperformed Metabolites
The Nightingale NMR panel returns 159 features, most of them tightly correlated lipoprotein and fatty acid measures, and the PANEL baseline already carried clinical lipids. So much of the metabolomic signal was spent before modeling began. Olink coverage of 2,923 targets spans more independent biology. Carrasco-Zanini and colleagues (2024) showed that sparse proteomic signatures beat clinical models across a wide disease range; this work puts both layers on identical footing, same participants, same folds, same scoring.
What the Feature Lists Say About Mechanism
Protein importance scores recovered known biology and went past it. Respiratory endpoints shared SCGB1A1 as protective and PRR4 as risk-associated, while renal and liver disease both leaned on FAP and ITGAM. Peripheral artery disease stood apart, marked by BOC and CDON alongside NT-proBNP. Every measurement came from plasma at one baseline draw, so brain or lung processes appear only as whatever leaks into circulation, a limit the authors state plainly.
In Practice: When Both Layers Share One Injection
The combined-omics penalty here reads as a statistics problem more than a biology one: two platforms, two batch structures, and a few thousand extra degrees of freedom against roughly 700 to 3,500 cases per endpoint. When both layers come off the same injection, as they do in the Omni-MS workflow we run at Dalton, metabolite and protein measurements share an aliquot and a batch, which removes one noise source before modeling begins. That matters for teams planning multi-omics biomarker discovery at cohort scale, where sample volume is usually the binding constraint, and it is why our multi-omics analytical services start from a single aliquot instead of parallel platforms.
Frequently Asked Questions
What is multi-omics integration used for?
Multi-omics integration combines two or more molecular layers, such as proteomics and metabolomics, measured on the same samples to explain or predict a phenotype. Clinical research uses it for risk prediction, patient stratification, and mechanism work where one layer leaves gaps. Here it covered 17 incident diseases in a prospective cohort.
Does multi-omics prediction beat proteomics alone for disease risk?
Not reliably at this sample size. Adding 159 NMR metabolites to 2,923 plasma proteins cost an average 0.002 of C-index against the proteomics-only model. Combined models did win for peripheral artery disease, COPD, renal disease, and fractures, so it depends on the endpoint.
How much plasma does a combined proteomics and metabolomics study need?
That depends on whether the two assays run on separate platforms or one. Split workflows need distinct aliquots for the protein and metabolite arms, doubling the volume ask and adding a second batch structure. A single-injection workflow measures both from one aliquot.
Conclusion
What holds up is the ranking: for predicting who develops disease over years, a broad plasma proteome beats a targeted NMR metabolite panel, and it does so after honest adjustment for clinical risk factors. What isn't settled is whether that ranking survives untargeted metabolomics, a different biobank, or repeated sampling. Teams sizing a discovery cohort can treat proteomics as the anchor layer and justify metabolomics on mechanistic grounds rather than incremental C-index.
Related Reading
See how we run these analyses in one lab: Dalton's multi-omics CRO services.
Citation
Du, J., Zhou, M., Wang, H., Wang, J., Raffield, L. M., Zhou, R., Li, Y., Chen, C., & Sun, Q. (2026). Multi-omics integration predicts the incidence of 17 diseases in the UK Biobank. Nature Communications, 17(1), 6271. https://doi.org/10.1038/s41467-026-73017-z
Note
This blog post summarizes findings from the above-cited research. Figures are adapted from the original publication. For full details, please refer to the source article.
By Seungjun Yeo, CEO at Dalton Bioanalytics. Specializing in multi-omics mass spectrometry for drug discovery and biomarker research.
