Multi-Omics: Machine Learning for Early Parkinson's
- Aug 2
- 4 min read
Parkinson's disease is usually caught only after most dopaminergic neurons are gone, leaving a narrow window for treatments that might slow it. A January 2026 study in PLOS ONE asked a sharper question: could a multi-omics fingerprint from blood and cerebrospinal fluid flag the disease earlier? Wei Liu and colleagues pulled DNA methylation, gene expression, and cerebrospinal fluid proteomics from 305 participants in the Parkinson's Progression Markers Initiative, then stacked network biology and machine learning on top. The payoff was a diagnostic model built from signals any single assay tends to miss.
Key Takeaways
Multi-omics integration beat every single-omics layer for classifying early Parkinson's, reaching an AUC of 0.72 on held-out patients.
Cerebrospinal fluid proteomics was the strongest single layer, a reminder that protein readouts often carry more diagnostic signal than upstream nucleic acid measurements.
The model rediscovered known Parkinson's genes such as DNAJC6, SORL1, and NEDD4L and pointed to neuroinflammation and autophagy pathways, supporting its biological plausibility.
External validation held up only modestly at AUC 0.57, so this panel needs larger replication before any clinical role.
How Multi-Omics Integration Sharpened the Signal
The team split the PPMI data into a 244-sample training set and a 61-sample test set. Feature selection ran per layer with sparse partial least squares discriminant analysis, then DIABLO, the multi-omics integration engine from Singh and colleagues (2019), tied the layers into one latent structure and retained 56 CpG sites, 61 genes, and 70 proteins. Rather than treating those as a flat list, the authors mapped them onto a STRING protein-protein interaction network and used hidden-node and propagation algorithms to surface 59 regulatory hubs.
Key Findings
Multi-omics model performance: the integrated model hit test AUC 0.72 (95% CI 0.51 to 0.85), accuracy 0.74, and F1 0.83.
Selected panel: DIABLO retained 56 CpG sites, 61 genes, and 70 CSF proteins, and network topology flagged 59 regulatory nodes.
Proteomics led the single-omics comparison, outperforming gene expression profiling and DNA methylation.
External check: on 26 GEO samples with only two omics layers, the model reached AUC 0.57, a real drop the authors report plainly.

Figure 1. Study workflow. The three-omics PPMI dataset splits into a training set (n=244) and a test set (n=61). Single-omics analysis feeds functional enrichment and multi-omics integration, which feeds network topological analysis. The trained model is checked by internal validation and by external validation on an independent two-omics cohort (n=26). Adapted from Liu et al. (2026), PLOS ONE.
Why Network Topology Carried Weight
A protein can look unremarkable alone yet sit at the center of disease-relevant wiring. That is the gap the network step closes. By propagating signal across STRING interactions and scoring hidden nodes, the authors ranked features by connectivity rather than fold-change, then fed those rankings into XGBoost classifiers trained with stratified five-fold cross-validation over 1,000 iterations. The topological model trailed the full integration, which says integration and topology are complementary, not interchangeable.
From PPMI Cohort to Clinical Readiness
An AUC near 0.72 across cerebrospinal fluid and blood is respectable for a screening-adjacent tool, though far from a standalone test. The honest read is that the external AUC of 0.57 shows how fragile these panels stay when the platform and available layers shift between cohorts. A validation set missing the proteomic layer probably cost more accuracy than the discussion admits.
One Sample, Three Omics: Making Integration Pay Off
Pulling methylation, transcripts, and proteins from the same participants is what gives this model its edge, since no single layer captured early disease on its own. When those layers come from separate platforms and separate aliquots, batch effects and sample-volume limits creep in, one reason single-injection workflows matter for multi-omics biomarker discovery. The multi-omics analysis we run at Dalton is built around that shared-sample principle, so cross-omics comparisons rest on one aliquot rather than three reconciled ones.
Frequently Asked Questions
What is multi-omics analysis?
Multi-omics analysis measures several molecular layers, such as DNA methylation, gene expression, and proteins, from the same samples and integrates them into one model. Methods like DIABLO align the layers so shared disease signal that any single assay would miss becomes visible.
Can multi-omics improve early Parkinson's disease diagnosis?
In this study, multi-omics integration reached AUC 0.72 on held-out patients, better than methylation or transcriptomics alone, with cerebrospinal fluid proteomics as the strongest single layer. External validation was weaker, so the approach is promising but not yet clinic-ready.
What samples does a multi-omics diagnostic study need?
This work used 305 PPMI participants with paired blood and cerebrospinal fluid, plus a 26-sample external cohort. Paired specimens across omics layers matter more than raw numbers, and underpowered external sets often cost accuracy.
Conclusion
The believable part is that combining three molecular layers with network-aware machine learning tops any single layer for early Parkinson's classification. What stays unproven is generalizability, since the external AUC slid to 0.57 with one layer missing. For teams designing biomarker studies, the practical lesson is to lock down paired, single-source sampling before scaling the cohort.
Related Reading
See how we run these analyses in one lab: Dalton's multi-omics CRO services.
Citation
Liu, W., Xu, L., Wang, X., and Wang, J. (2026). Integrative multi-omics and network-based machine learning for early diagnosis of Parkinson's disease. PLOS ONE, 21(1), e0329980. https://doi.org/10.1371/journal.pone.0329980
Note
This blog post summarizes findings from the above-cited research. Figures are adapted from the original publication. For full details, please refer to the source article.
By Seungjun Yeo, CEO at Dalton Bioanalytics. Specializing in multi-omics mass spectrometry for drug discovery and biomarker research.
