Pd-mci markers, screening methods and uses thereof

By integrating plasma transcriptomics and cerebrospinal fluid proteomics analysis, combined with Mendelian randomization and single-cell nuclear RNA sequencing, PD-MCI biomarkers were screened, solving the problem that existing technologies cannot accurately predict the development of mild cognitive impairment in Parkinson's disease patients, and enabling early risk prediction and precision treatment.

CN122487677APending Publication Date: 2026-07-31SHANGHAI EAST HOSPITAL EAST HOSPITAL TONGJI UNIV SCHOOL OF MEDICINE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI EAST HOSPITAL EAST HOSPITAL TONGJI UNIV SCHOOL OF MEDICINE
Filing Date
2026-04-30
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Current technologies cannot accurately predict whether Parkinson's disease patients will develop mild cognitive impairment (PD-MCI), and there is a lack of effective disease-modifying therapies. Longitudinal multi-omics analyses have failed to capture the temporal dynamics of disease progression and the interactions between molecular levels.

Method used

We employed integrated multi-omics analysis (plasma transcriptomics and cerebrospinal fluid proteomics), combined with Mendelian randomization and single-cell nuclear RNA sequencing, to validate the expression of biomarkers in brain cell types. We obtained treatment strategies by computational drug repositioning and screened biomarkers such as LRPPRC, WDR1, EIF3H, PLEKHH2, SLC35G2, UTS2, MMP12, MSMB, CA1, and POSTN.

Benefits of technology

The study identified the core molecular pathways driving PD-related cognitive decline, provided an early PD-MCI risk prediction model, and offered a basis for target screening of early intervention drugs, thereby improving the accuracy of PD-MCI diagnosis and treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122487677A_ABST
    Figure CN122487677A_ABST
Patent Text Reader

Abstract

This invention provides a PD-MCI biomarker, its screening method, and its application. The PD-MCI biomarker is PLEKHH2, CA1, and WDR1. This invention opens new avenues for therapeutic development by identifying the "neuroconnectopathy" mechanism, while drug repositioning candidates establish a clear next step for preclinical validation and clinical trials.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biomedicine, specifically to PD-MCI biomarkers, their screening methods, and applications. Background Technology

[0002] Parkinson's disease (PD) is the second most common neurodegenerative disease worldwide, characterized by both key motor symptoms and numerous non-motor complications. Among these, PD with mild cognitive impairment (PD-MCI) is the most devastating and burdensome, affecting up to 80% of patients within 15 years of onset. The emergence of PD-MCI signifies a severe decline in the quality of life for both patients and caregivers, and there are currently no effective disease-modifying therapies.

[0003] Neuropathological markers of PD-MCI include extensive α-synuclein aggregation (Lewy bodies), often coexisting with Alzheimer's disease-related pathologies (amyloid-β plaques, neurofibrillary tau tangles). This indicates a complex and multifactorial etiology. This complexity poses a significant challenge: the inability to accurately predict which PD patients will develop mild cognitive impairment or dementia, and the incomplete understanding of the underlying molecular mechanisms, hinder targeted prevention.

[0004] In recent years, high-throughput omics technologies have offered unprecedented opportunities to reveal the molecular complexity of neurodegenerative diseases. Genome-wide association studies (GWAS) have identified numerous risk sites for PD, but their ability to explain the risk of PD-MCI remains limited. In contrast, transcriptomic and proteomic analyses of readily available bodily fluids (such as plasma and cerebrospinal fluid [CSF]) provide dynamic snapshots of pathological processes and hold promise for discovering diagnostic / prognostic biomarkers. However, most studies to date have been cross-sectional or focused only on a single omics level, failing to capture the temporal dynamics of disease progression and the interactions between them at the molecular level. Therefore, longitudinal, multi-omics approaches are crucial for mapping the temporal molecular events leading to PD-MCI.

[0005] Furthermore, traditional association studies using whole tissues or biological fluids cannot determine the specific cellular origin of dysregulated molecules; while omics analyses reveal correlations, they cannot establish causal relationships. Summary of the Invention

[0006] This invention aims to overcome the aforementioned shortcomings by designing a comprehensive translational study using the Parkinson's Disease Progression Biomarkers Initiative (PPMI) cohort, which employs a large-scale, in-depth phenotypic approach. We believe that integrated multi-omics analysis (plasma transcriptomics + cerebrospinal fluid proteomics) can identify core molecules / pathways driving PD-related cognitive decline. Based on this, we propose integrating these findings into a PD-MCI risk prediction model, using MR to test causality, and validating it through snRNA-seq in specific brain cell types. Therefore, we aim to translate this into more precise treatment strategies through computational drug repositioning.

[0007] This invention provides a method for screening PD-MCI biomarkers, characterized by comprising the following steps:

[0008] S1. A well-defined PD-MCI cohort was established using the Movement Disorders Association (MDS) Level II diagnostic criteria and longitudinal cognitive assessment;

[0009] S2. Integrating multi-omics analysis of plasma transcriptomics and cerebrospinal fluid proteomics to identify molecular signatures of PD-MCI progression;

[0010] S3. Apply systems biology approaches to understand the underlying biological mechanisms of PD-MCI;

[0011] S4. Systematically evaluate several machine learning models to select the best algorithm for PD-MCI risk prediction and obtain potential biomarkers;

[0012] S5. Mendelian randomization was used to explore potential causal relationships, and single-cell nuclear RNA sequencing was used to verify the expression of biomarkers in brain cell types.

[0013] S6. Based on the finally selected biomarkers, therapeutic candidate drugs are obtained by calculating drug repositioning.

[0014] Furthermore, the method for screening PD-MCI biomarkers provided by this invention is further characterized in that:

[0015] The PD-MCI markers include LRPPRC, WDR1, EIF3H, PLEKHH2, SLC35G2, UTS2, MMP12, MSMB, CA1, and POSTN.

[0016] Furthermore, the method for screening PD-MCI biomarkers provided by this invention is further characterized in that:

[0017] In step S3, the systems biology approach includes: functional enrichment, PPI network analysis, and longitudinal proteomic trajectory analysis.

[0018] Furthermore, the method for screening PD-MCI biomarkers provided by this invention is further characterized in that:

[0019] Based on the S2-S3 steps, predict the preclinical window period.

[0020] Furthermore, the method for screening PD-MCI biomarkers provided by this invention is further characterized in that:

[0021] Predict drugs for early intervention based on the preclinical window period.

[0022] Furthermore, the method for screening PD-MCI biomarkers provided by this invention is further characterized in that:

[0023] In S4, the optimal algorithm is the svmRadialCost model.

[0024] In addition, the present invention also provides the application of at least one biomarker selected from PLEKHH2, CA1, WDR1, MSMB, and POSTN in the preparation of PD cognitive impairment diagnostic / screening products.

[0025] In addition, the present invention also provides the application of at least one biomarker selected from PLEKHH2, CA1, WDR1, MSMB, and POSTN in the preparation of PD cognitive impairment typing products.

[0026] In addition, the present invention also provides the application of at least one of the biomarkers PLEKHH2, CA1, WDR1, MSMB, and POSTN as intervention targets in the preparation of drugs for the prevention or treatment of PD-MCI. Attached Figure Description

[0027] Figure 1 A flowchart and comprehensive analytical framework for the PPMI cohort study to identify / validate proteins and transcriptomic biomarkers associated with cognitive decline in PD;

[0028] The flowchart for (A) PPMI cohort inclusion is as follows. From the total PPMI cohort (n=4115), we first identified PD patients (n=1454). The final analysis cohort consisted of 322 PD patients with complete baseline data (including cerebrospinal fluid proteomics, plasma RNA-seq) and 4-year longitudinal cognitive assessments. HC: healthy controls; SWEDD: scans showed no evidence of dopaminergic deficiency;

[0029] (B) Schematic diagram of the multimodal bioinformatics and statistical analysis pipeline. The analysis begins with standard preprocessing of proteomics and transcriptomics data, followed by differential expression analysis to identify candidate biomarkers. These candidates undergo time-series analysis (exploring dynamic changes), pathway enrichment analysis (deciphering biological processes), and LASSO regression (feature selection in predictive models). These features are used to build machine learning models, improve interpretability through SHAP analysis, and establish clinical relevance through clinical feature correlation analysis. Mendelian randomization (MR) is used to infer causality, and single-cell sequencing analysis is used for cell localization; these steps collectively guide the prioritization of hypothetical drug targets.

[0030] Figure 2 Integrating multi-omics analysis to identify key molecules and pathways associated with cognitive impairment in PD;

[0031] Among them, (AB) volcano plots were used to compare differential expression between PD-MCI patients and cognitively normal patients within 4 years. (A) Plasma RNA-seq analysis. (B) Cerebrospinal fluid proteomics analysis;

[0032] (CD) Gene Ontology (GO) enrichment analysis of biological processes. (C) GO terms significantly enriched in significant RNA. (D) GO terms significantly enriched in significant proteins;

[0033] (EF) Protein-protein interaction (PPI) network of core differentially expressed proteins.

[0034] Figure 3 Longitudinal proteomic trajectories reveal dynamic cerebrospinal fluid proteomic biomarkers in patients with PD-CI prior to cognitive impairment.

[0035] Among them, (A) Temporal expression patterns of individual proteins (top 10 selected from LASSO) (Z-score normalized levels). These trajectories illustrate heterogeneous temporal dynamics, with some proteins showing gradual increases (e.g., pro-inflammatory markers) while others show early decreases (e.g., synaptic proteins);

[0036] (B) Average protein expression trajectory for each cluster. The line represents the average Z-score of all proteins in the cluster, and the shading represents the standard error of the mean;

[0037] (C) A heatmap showing all significantly differentially expressed protein patterns;

[0038] (D) Pie chart illustrates the proportion of disordered proteins in the five trajectory clusters.

[0039] Figure 4 . Construct and compare machine learning models for predicting cognitive impairment in Parkinson's disease;

[0040] Among them, (AB) LASSO regression selected the most important biomarkers from the DEG results. (A) Plasma transcripts; (B) Cerebrospinal fluid proteomics;

[0041] (CE) Comparative performance of various machine learning algorithms. Box plots show the ROC-AUC score distributions of random forest, XGBoost, and logistic regression models evaluated by 5-fold cross-validation. (C) Model based on plasma transcriptome + demographics; (D) Model based on cerebrospinal fluid proteomics + demographics; (E) Model based on plasma transcriptome + cerebrospinal fluid proteomics + demographics;

[0042] (F) Overall model performance evaluation on the independent test set. The ROC curve illustrates the discriminative power of the model based on plasma transcriptomics + cerebrospinal fluid proteomics + demographics;

[0043] (G) Performance evaluation of the optimal svmRadialCost model on the independent test set.

[0044] Figure 5 SHAP interpretation of machine learning models used to predict cognitive impairment in Parkinson's disease;

[0045] Among them, (AC) is the SHAP analysis of model interpretability. This figure ranks the top features based on the mean absolute SHAP value, representing their overall importance in the optimal model prediction. (A) Cerebrospinal fluid proteomics + Demographics; (B) Plasma transcripts + Demographics; (C) Cerebrospinal fluid proteomics + Plasma transcripts + Demographics.

[0046] Figure 6 Cross-sectional associations between candidate protein biomarkers and domain-specific cognitive scores;

[0047] The heatmap shows the standardized beta coefficients from the multiple linear regression model. Each cell represents the association between a feature (X-axis) and a specific cognitive test score (Y-axis). The color scale follows the beta value, and p-values ​​are labeled as: p < 0.05, *: p < 0.01.

[0048] SDM TOTAL: Total score for semantic memory test; DVTTOTALRECALL: Total recall score for delayed visual test; DVTSFTANIM: Semantic fluency test in delayed visual test (animal category); DVTSDM: Semantic memory in delayed visual test; DVTRECOGRETENTION: Recognition retention in delayed visual test; DVTRECOGDISCINDEX: Recognition discrimination index; DVTDELAYED_RECALL: Delayed recall; DVSDSDM: Semantic memory in delayed visual recognition test; DVSSFTANIM: Semantic fluency test in visuospatial test (animal category); DVS_LNS: Alphanumeric sorting in visuospatial test.

[0049] Figure 7 Sensitivity, heterogeneity, and pleiotropic effects of WDR1 in Mendelian randomization analysis for Parkinson's disease and dementia;

[0050] Among them, (A) scatter plot. (B) forest plot. (C) funnel plot. (D) leave-one-out sensitivity analysis.

[0051] Figure 8 Single-cell RNA sequencing of human brain tissue reveals altered cellular community structure in PDD;

[0052] Among them, (A) integrated UMAP visualization of all cores in the PDD and control (CTRL) queues;

[0053] (B) UMAP map colored according to clinical diagnosis;

[0054] (C) Heatmap of predicted differential gene expression in different cell types between PDD and CTRL;

[0055] (D) CellChat analysis of PDD and CTRL.

[0056] Figure 9. Gene marker annotation of single-nuclear RNA sequencing data;

[0057] The dot plot illustrates the expression of selected marker genes in each cell type. The size of each dot represents the percentage of cells in a cluster that express a specific gene, and the color of the dot represents the average expression level of that gene after logarithmic normalization.

[0058] Figure 10. Quality control and annotation of single-nuclear RNA sequencing data;

[0059] Among them, (AC) are the unified manifold approximation and projection visualization results. (A) Each number represents a cell subpopulation, reflecting more refined cellular heterogeneity. (B) Displayed by sample group. (C) Displayed by cell type specificity. (D) A pie chart shows in detail the proportion of cell nuclei allocated to each cell type.

[0060] Figure 11. Full phenotypic association study;

[0061] The Manhattan plot (A–E) illustrates the associations between single nucleotide polymorphisms (SNPs) used as instrumental variables and a wide range of human phenotypes. A key finding is that these SNPs are significantly associated with target gene levels (belonging to the "biomarker" category) but not with other neuropsychiatric, metabolic, or immune phenotypes (no points are above the red horizontal line). This strongly demonstrates the specificity of the instrumental variables and significantly reduces the likelihood of causal estimation bias due to pleiotropy. (A) CA1; (B) MMP12; (C) MSMB; (D) POSTN; (E) PLEKHH2. The horizontal axis from left to right is: Infectious, Neoplasms, Blood / Immune, Endocrine / Metabolic, Mental, Nervous, Eye, Ear, Cardiovascular, Respiratory, Digestive, Skin, Musculoskeletal, Urinary / Renal, Pregnancy / Congenital, Lab findings, Health services, Special; the vertical axis from bottom to top is: 0, 2, 4, 6, 8, 10, 12, 14, 16, 18, 20. Detailed Implementation

[0062] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0063] In this embodiment, the following is adopted: Figure 1 The overall analytical framework shown in B follows a multi-stage, hypothesis-driven approach. It integrates differential expression analysis, machine learning-based prediction, genetic causal inference, single-cell validation, and drug repositioning, establishing a translational research pipeline from biomarker discovery to therapeutic hypotheses. This framework includes the following steps:

[0064] S1. A well-defined PD-MCI cohort was established using the Movement Disorder Society (MDS) Level II diagnostic criteria and longitudinal cognitive assessment.

[0065] S2. Integrated multi-omics analysis (plasma transcriptomics and cerebrospinal fluid proteomics) to identify molecular signatures of PD-MCI progression;

[0066] S3. Apply systems biology methods (functional enrichment, PPI network analysis, longitudinal proteomic trajectory) to understand the underlying biological mechanisms of PD-MCI;

[0067] S4. The system evaluated 27 machine learning models to select the best algorithm for PD-MCI risk prediction;

[0068] S5. Mendelian randomization was used to explore potential causal relationships, and single-cell nuclear RNA sequencing was used to verify the expression of biomarkers in brain cell types.

[0069] S6. Propose therapeutic candidates by calculating drug repositioning.

[0070] This integrated framework aims to translate multi-omics findings into clinically actionable insights for the early diagnosis and targeted intervention of PD-MCI.

[0071] The specific process is as follows:

[0072] In S1, a well-defined PD-MCI cohort was established using the Movement Disorder Society (MDS) Level II diagnostic criteria and longitudinal cognitive assessment.

[0073] The data used in this study were obtained from the Parkinson's Disease Progression Biomarkers Initiative (PPMI) database (www.ppmi-info.org / access-data-specimens / download-data, RRID:SCR_006431). PPMI is an ongoing international, multicenter, prospective observational study designed to identify biomarkers of PD progression. For this analysis, we included de-identified PD patient data. Written informed consent was provided by all participants, and the study was approved by the institutional review committees at each participating site.

[0074] The process of filtering target data is as follows Figure 1 As shown in Figure A, from the initial 4115 PPMI participants, we focused on 1454 PD patients. After applying inclusion criteria (matched multi-omics data and availability of 4-year cognitive follow-up), 322 patients constituted the final analysis cohort. That is, the primary cohort analyzed in this embodiment consists of PD patients with baseline plasma RNA-seq data, baseline cerebrospinal fluid proteomics data, and a complete 4-year longitudinal cognitive assessment.

[0075] In this embodiment, for the 322 patients, the Montreal Cognitive Assessment (MoCA) used a cutoff value of 26 as the primary cognitive test to define cognitive progression and diagnose PD-MCI according to the Movement Disorders Association Grade II criteria. Other cognitive tests were used to validate the stability of the biomarkers.

[0076] The diagnostic criteria for PD-MCI are as follows (Movement Disorders Association Class II):

[0077] 1. Cognitive impairment: Based on neuropsychological testing, there is objective evidence of impairment in at least one cognitive domain (attention, executive function, visuospatial function, memory, language), manifested as performance below 1.5 standard deviations of the appropriate norm.

[0078] 2. Functional status: Based on clinical assessment (MDS-UPDRS Part 2 score ≤ 1), daily living activities are generally normal.

[0079] 3. Time points: PD-MCI diagnosis was determined at any follow-up visit during the 4-year observation period (baseline, annual visits in years 1–4). Patients who developed PDD (dementia) were excluded from the PD-MCI analysis but included in the PDD subgroup analysis.

[0080] 4. Cognitive Assessment Area: Neuropsychological tests include: MoCA (Montreal Cognitive Assessment) for holistic cognition; HVLT-R (Hopkins Verbal Learning Test-Revised) for memory; LNS (Alphabetic Number Sorting) for attention / working memory; SDMT (Symbolic Digit Modal Test) for processing speed; Line Orientation Judgment for visuospatial function; and Semantic Fluency for language.

[0081] PD-CN (Parkinson's Disease with Normal Cognition) definition: If a patient maintains cognitive integrity during a 4-year follow-up period, meets all five cognitive domain criteria (including MoCA, HVLT-R, LNS, SDMT, line orientation judgment, and semantic fluency) at all time points, and has no significant impairment in comprehensive cognitive assessment, then the patient is defined as PD-CN. Although a total MoCA score ≥ 26 is an important reference indicator, it is not the only criterion. Considering the performance in all five cognitive domains, some patients with MoCA < 26 perform well in other cognitive domains and are therefore classified as PD-CN.

[0082] The table below shows the baseline characteristics of participants from the Parkinson's Disease Progression Biomarkers Initiative (PPMI) cohort, stratified by PD-MCI results.

[0083]

[0084] †: Chi-square test, ¶: Wilcoxon rank-sum test.

[0085] In step S2, multi-omics analysis integrating plasma transcriptomics and cerebrospinal fluid proteomics is performed to identify molecular features of PD-MCI progression, including the following steps:

[0086] S2.1. Multi-omics data acquisition and preprocessing

[0087] S2.1.1. Plasma RNA-seq

[0088] Total RNA was extracted from plasma samples according to the PPMI sequencing protocol. Library preparation was performed on the Illumina platform. Raw sequencing reads (FASTQ files) were quality controlled using FastQC (v0.11.9) and trimmed using Trimmomatic (v0.39). Reads were aligned to the human reference genome (GRCh38) using STAR (v2.7.10a), and gene counts were quantified using featureCounts (v2.0.1). Counts were normalized and transformed using the varianceStabilizingTransformation function in the DESeq2 R package (v1.34.0).

[0089] S2.1.2. Cerebrospinal Fluid Proteomics

[0090] Cerebrospinal fluid samples were analyzed using SomaLogic's SOMAscan platform, targeting protein biomarkers based on slow dissociation rate modified aptamer (SOMAmer) technology. SOMAscan data were processed according to SomaLogic's standard workflow "HybNormPlateScaleCal," which includes four core steps: hybridization normalization, plate calibration, intra-plate median signal normalization, and calibration. Data processed through this workflow were used as analyzable raw data. Protein intensities were log2 transformed and quantile normalized to correct for inter-sample variability. Batch effects were adjusted using the ComBat function from the sva R package.

[0091] S2.2. Differential Expression Analysis

[0092] Differential analysis of RNA-seq and proteomics data between patients who progressed and those who did not was performed using linear models in the limma R package (v3.50.3). The Benjamini-Hochberg method was used to control for false discovery rate (FDR). Significance was set at |log2(fold change)| > 0.3 and p-values ​​were adjusted to < 0.1.

[0093] The results are as follows:

[0094] Differential expression analysis was performed between PD patients who developed mild cognitive impairment (MCI) and those who did not, such as... Figure 2 As shown in A and 2B, 19 differentially expressed transcripts (5 upregulated) and 44 differentially expressed proteins (2 upregulated) were identified.

[0095] In step S3, systems biology methods (functional enrichment, PPI network analysis, longitudinal proteomic trajectory) are applied to understand the underlying biological mechanisms of PD-MCI, including the following steps:

[0096] S3.1. Functional Enrichment Analysis

[0097] Gene Ontology (GO) enrichment analysis was performed on significantly dysregulated genes and proteins using the clusterProfiler R package (v4.2.2). Terms with an FDR-adjusted p-value < 0.05 were considered significant.

[0098] The results showed that functional enrichment analyses of proteomics and RNA-seq revealed similar patterns, with significant pathways highly concentrated in biological processes crucial for the integrity of neural circuits. Figure 2 As shown in C and 2D, pathways such as axonal guidance, neuronal projection development, and synaptic organization are consistently and significantly enriched at both molecular levels. This striking convergence at the transcriptomic and proteomic levels provides compelling evidence for a "neuroconnectopathy" in PD-MCI pathology. The consistent dysregulation of axonal guidance (SLC38A11), synapsis (PLEKHH2), and extracellular matrix remodeling (COL6A1, ITGAV) at the molecular level suggests that PD-MCI involves progressive disruption of neural circuit integrity, rather than isolated protein dysregulation.

[0099] S3.2. Interaction Network Analysis

[0100] Gene-gene interaction networks were constructed using the Genemania database, and protein-protein interaction (PPI) networks were constructed using the STRING database (v11.5) for interaction network analysis.

[0101] The results revealed that gene-gene interaction network and protein-protein interaction (PPI) network analysis further refined these findings, such as... Figure 2 As shown in E and 2F, it distills the complex molecular landscape into core functional modules and identifies key hub genes, including SLC38A11, PLEKHH2, and protein network hubs such as COL6A1 and ITGAV, which act as central regulators of this disordered network.

[0102] S3.3. Longitudinal proteomic trajectory analysis

[0103] To characterize the temporal dynamics of protein expression during the progression of Parkinson's disease with mild cognitive impairment (PD-MCI), trajectory modeling and unsupervised clustering were performed. Protein expression levels were first standardized to Z-scores relative to matched controls to mitigate inter-individual variability. The longitudinal trajectories of these Z-scores over the years leading to diagnosis were then modeled using Local Estimated Scatter Smoothing (LOESS) regression to provide a continuous view of protein fluctuations.

[0104] Unsupervised hierarchical clustering was performed based on the Euclidean distance between LOESS predicted trajectories to group proteins with similar temporal patterns. The Ward.D2 method was used as the clustering criterion to minimize intra-cluster variance. This analysis robustly classified the proteins into five distinct clusters. Proteins were defined as misaligned if their absolute Z-score exceeded a threshold of 0.3, as shown in the trajectory heatmap.

[0105] To map the temporal sequence of molecular events leading to cognitive impairment, this embodiment tracked the expression of significantly differentially expressed proteins (identified in baseline analysis) prior to diagnosis (up to 4 years before PD-MCI). Figure 3 As shown in Figure B, unsupervised k-means clustering of these longitudinal profiles reveals that the PD-MCI proteomic landscape is not static. It evolves through distinct, coordinated dynamic patterns. Figure 3 As shown in C, we identified five main trajectory clusters (cluster 1 to cluster 5), each with unique temporal characteristics.

[0106] like Figure 3 As shown in Figure D, most differentially expressed proteins (65.9%) belonged to clusters defined by early and intermediate changes (clusters 3 and 4), indicating that key pathological processes begin 3–4 years before the clinical diagnosis of PD-MCI. The protein trajectory showed a turning point approximately 3 years before MCI diagnosis, which may be related to the temporal characteristics of PD patients enrolled in the PPMI cohort and the selection of time points during follow-up. Proteins in clusters 2 and 5 showed a relatively sustained decrease or increase before MCI diagnosis; these proteins are closely related to mechanisms such as extracellular matrix remodeling, metabolic homeostasis, mitochondrial integrity, oxidative stress, and inflammation. Figure 3 As shown in Figure A, the representative trajectories of individual proteins further highlight this heterogeneity: some markers show gradual increases or decreases, while others show curved patterns.

[0107] In summary, these findings suggest that the molecular pathophysiology of PD-MCI begins long before the onset of clinical symptoms and follows a multi-stage time pattern, providing a critical window for early intervention. This extended preclinical window offers unique opportunities for early therapeutic intervention, particularly with drugs targeting neuroinflammation (cluster 2) and synaptic dysfunction (cluster 4).

[0108] In step S4, the system evaluated 27 machine learning models to select the best algorithm for PD-MCI risk prediction. The specific process is as follows: Machine learning modeling and validation strategy:

[0109] S4.1. To address potential multicollinearity among features and improve robustness, we employ two complementary regularization methods: LASSO regression (L1 penalty) and elastic networks (combining L1 and L2 penalties).

[0110] LASSO regression was applied to identify sparse sets of differentially expressed transcripts and proteins that best distinguished between PD-MCI and PD-CN. The LASSO algorithm was implemented using the glmnet R package (v4.1-6) and 10-fold cross-validation was used to determine the optimal lambda parameter (min.lambda.1se).

[0111] The Elastic Network, incorporating L1 and L2 penalties, serves as a complementary approach to address the potential limitations of LASSO in the presence of highly correlated features. The Elastic Network is implemented using the caret R package, with the alpha parameter ranging from 0 to 1 (where 0 = Ridge Regression, 1 = LASSO). Optimal alpha and lambda are selected using 5-fold cross-validation based on minimum classification error. The Elastic Network method identified 8 transcripts and 9 proteins, which showed moderate overlap (75-80%) with the features selected by LASSO, demonstrating the robustness of this feature selection process.

[0112] S4.2. In order to fully identify the best prediction algorithm and ensure the generalization ability of the model, 29 machine learning models were systematically evaluated using the caret R package (v6.0-94).

[0113] A. The algorithm suite includes:

[0114] Linear models: glm (Generalized Linear Model), glmnet (Penalized Logistic Regression with Elastic Network), bayesglm (Bayesian GLM), spls (Sparse Partial Least Squares), pls (Partial Least Squares), lda, lda2, pda, pda2, kernelpls, widekernelpls, simpls;

[0115] Support Vector Machines: svmLinear (linear kernel), svmRadial (radial basis kernel), svmRadialSigma (radial kernel with sigma optimization), svmRadialCost (radial kernel with cost optimization).

[0116] Tree-based methods: rf (random forest), cforest (conditional random forest), RRF (regularized random forest), rpart (recursive partitioning), C5.0 (decision tree), ctree (conditional inference tree), C5.0Tree;

[0117] Neural networks: nnet (single-layer neural network), pcaNNet (neural network with PCA preprocessing);

[0118] Ensemble methods: xgbLinear, xgbTree, xgbDART (XGBoost variant), LogitBoost (boosting logistic regression), glmboost;

[0119] All models were evaluated using repeated (n=5) 5-fold cross-validation and stratified sampling to ensure a balanced class distribution. Model performance was evaluated using seven metrics: accuracy, sensitivity, specificity, PPV (positive predictive value), NPV (negative predictive value), F1 score, and AUC.

[0120] B. Training set / independent test set partitioning

[0121] To ensure rigorous model validation and avoid overfitting, a two-stage validation approach was used: (1) internal cross-validation during model development, and (2) independent test set for final performance evaluation. The dataset (n=322) was first divided into a training set (70%, n=225) and an independent test set (30%, n=97) using stratified random sampling to maintain the PD-MCI / PD-CN ratio (approximately 1:2) between the two sets.

[0122] The training set is used for model development, with hyperparameter tuning performed through repeated (n=5) 5-fold cross-validation and stratified sampling to ensure a balanced class distribution. The independent test set is used solely to evaluate the performance of the best-performing model (svmRadialCost in this example), providing an unbiased estimate of the model's generalization ability to new patients. The independent test set is not involved in the model development and parameter tuning process at all.

[0123] C. Result

[0124] A comprehensive assessment shows that, Figure 4As shown in F and the table below, a comprehensive evaluation of 29 machine learning algorithms determined svmRadialCost (a support vector machine with radial basis function and cost optimization) to be the best model. Although pda2 had a slightly higher AUC (0.7667), svmRadialCost performed better in terms of accuracy (0.75), specificity (0.667), and F1 score (0.778), with an AUC of 0.7598, second only to pda2. Considering all performance metrics, model interpretability, and clinical applicability, we selected svmRadialCost as the final model (AUC = 0.7598, sensitivity = 82.4%, specificity = 66.7%).

[0125] Complete performance metrics table for 29 algorithms

[0126]

[0127] like Figure 5 As shown, the SHAP analysis of the optimal svmRadialCost model elucidates its decision-making process. PLEKHH2 (plasma transcript) and CA1 (cerebrospinal fluid protein) emerged as top biological predictors, while age at PD diagnosis and years of education remained key demographic factors. This interpretable framework links the model's predictive power to specific biological insights derived from multi-omics analysis, particularly... Figure 2 The "neuroconnectopathy" pathway identified in the study.

[0128] In step S5, Mendelian randomization is used to explore potential causal relationships, and single-cell nuclear RNA sequencing is used to verify the expression of biomarkers in brain cell types. The specific process is as follows:

[0129] S5.1. Clinical relevance analysis to identify clinically significant molecules.

[0130] To assess the clinical relevance of the identified multi-omics traits, their associations with key neuropsychological assessments and demographic factors were evaluated. Multiple linear regression analysis was performed with cognitive scores from a range of tests as the dependent variable. Independent variables included baseline levels of candidate proteins and transcripts identified in differential expression analysis and LASSO regression. All models were adjusted for key covariates such as age and years of education to control for potential confounding effects. Results are presented as standardized Beta coefficients, quantifying the strength and direction of associations. Positive Beta indicates that higher levels of the omics trait are associated with better cognitive performance, while negative Beta indicates an association with poorer cognitive function.

[0131] like Figure 6As shown, the standardized effect size is visualized in a heatmap using a multiple linear regression model adjusted for age and years of education. This analysis reveals distinct and coherent patterns of association between specific proteins and cognitive function. For genes and proteins upregulated in the PD-MCI group identified in the differential expression analysis, they are generally and consistently negatively correlated with cognitive-related scores, and vice versa. The heatmap shows that PD-MCI molecular signatures are not uniform but rather correlated differently with impairments in specific cognitive domains. This domain-specific mapping enhances the biological plausibility of our candidate biomarkers and provides a more nuanced understanding of how different molecular pathways may contribute to the heterogeneous cognitive profile of PD-MCI.

[0132] S5.2. Mendelian randomization (MR) to verify causality

[0133] Protein quantitative trait loci (pQTLs) and expression quantitative trait loci (eQTLs) were obtained from the UKB-PPP and eQTLGen consortia. Summary statistics of results (PDD) were obtained from the FinnGen consortium (R9 version, finn-b-PD_DEMENTIA). MR analysis was performed using the TwoSampleMR R package (v0.5.7). Inverse variance weighting (IVW) was the primary method, supplemented by MR-Egger, weighted median, and weighted modal methods. Sensitivity analyses included Cochran's Q test for heterogeneity, MR-Egger intercept test for level pleiotropy, and leave-one-out analysis.

[0134] In this embodiment, pQTLs and eQTLs were used as instrumental variables. Among the features selected by LASSO regression, LRPPRC, WDR1, EIF3H, PLEKHH2, SLC35G2, and UTS2 have eQTL-based tools, while MMP12 and MSMB have pQTL-based tools. Notably, WDR1 showed an association trend with increased PDD risk (see table below: Mendelian randomization analysis assesses the causal effect of feature levels on PDD risk).

[0135]

[0136] Directionality was consistent across different MR methods (IVW, MR-Egger, weighted median, weighted modality), and sensitivity analysis showed minimum-level pleiotropy (MR-Egger intercept p > 0.05) and heterogeneity (Cochran's Q p > 0.05). Leave-one-out analysis confirmed robustness by indicating no single genetic variant driving the estimate. Figure 7As shown, the sensitivity, heterogeneity, and pleiotropic analysis of WDR1 in PPD under Mendelian randomization showed that elevated WDR1 gene expression / level is a causal risk factor for PPD, and the results are highly robust and reliable.

[0137] S5.3. Single-cell nuclear RNA sequencing (snRNA-seq) analysis

[0138] SnRNA-seq data from GEO303823 were processed using the Cell Ranger workflow (v7.1.0). Downstream analyses, including quality control, normalization, integration, clustering, and cell type annotation, were performed in Seurat (v4.3.0). Differential expression and visualization of candidate genes were then performed.

[0139] For each cell type (containing at least 3 cells in each group) and gene combination (among selected genes with expression records), log2 fold change was calculated, and the significance of differences in gene expression between groups was assessed using a nonparametric Wilcoxon rank-sum test. The raw p-values ​​were adjusted for false discovery rate (FDR) using the Benjamini-Hochberg method. Differential expression results were visualized using heatmaps. Cell-cell communication analysis was performed using CellChat and CellChatDB.human as references to characterize cell-cell interactions. Results are as follows: Figure 8 As shown.

[0140] After identifying peripheral predictive biomarkers, we aimed to localize their cellular origin within the pathological brain environment. Due to the lack of publicly available snRNA-seq datasets specifically for PD-MCI (mild cognitive impairment), we utilized snRNA-seq data from PDD (Parkinson's disease dementia) patients and healthy controls from GEO30382. While PDD represents a more severe stage of cognitive impairment than our primary outcome, PD-MCI, this approach provides valuable insights into the cellular ecology of late PD-related cognitive decline and validates whether the biomarkers we identified show dysregulation in the brain parenchyma of patients with clinically significant cognitive impairment. This approach is reasonable given the progressive nature of cognitive decline in PD, where PD-MCI represents an intermediate stage along the continuum towards PDD. Biomarkers dysregulated in PDD may also show alterations in PD-MCI, although perhaps at an earlier stage or with varying magnitudes.

[0141] Postmortem sequencing of single-cell nuclear RNA from the prefrontal cortex provided spatially resolved validation. For example... Figure 8 AB Figure 9 and Figure 10As shown, integrated transcriptomic analysis revealed a unique global shift in the PDD cellular landscape. We successfully resolved all major brain cell types, including excitatory neurons (35%), inhibitory neurons (22%), oligodendrocytes (26%), astrocytes (9%), microglia (3%), OPCs (4%), and endothelial cells (1%).

[0142] In addition to global transcriptomic changes, we quantified significant alterations in cell proportions, revealing a profound dysregulation of brain cellular ecology in PDD. Most notably, we observed a simultaneous significant reduction in oligodendrocytes and a significant increase in the proportion of microglia. This finding suggests that the pathophysiology of PDD involves simultaneous mechanisms of demyelination / loss of white matter integrity and persistent neuroinflammation.

[0143] Crucially, we then attempted to anchor the top causal candidates from previous analyses to their specific cellular contexts, confirming their altered expression across multiple cell types in PDD. Figure 8 C). Furthermore, CellChat further depicted cell-cell interactions within the PDD group ( Figure 8 D). This cell type-specific validation directly links our peripheral and genetic findings to the relevant cellular matrix in the brain, pinpointing dysfunction within excitatory and inhibitory neurons as a key event in the PDD pathological cascade, potentially leading to the observed "connection disorder".

[0144] In step S6, the specific method for proposing therapeutic candidate drugs by calculating drug repositioning is as follows:

[0145] The DrugBank database (v5.1.9) was queried to identify known and experimental drugs targeting preferred pivot proteins. Additionally, the DGIdb database was used for comprehensive drug-gene interaction mining.

[0146] In this embodiment, eight genes selected based on LASSO regression—LRPPRC, WDR1, EIF3H, PLEKHH2, SLC35G2, UTS2, MMP12, and MSMB—along with genes CA1 and POSTN, which showed significant and potential therapeutic relevance in differential expression analysis, were included in the drug relocation analysis to search for corresponding targeted drugs or regulatory molecules in drug databases. We found that five genes (CA1, MMP12, PLEKHH2, MSMB, and POSTN) found corresponding targeted drugs or regulatory molecules in drug databases (see table below). Although there are currently no known targeted drugs for LRPPRC, EIF3H, SLC35G2, UTS2, and WDR1 in the current database, we believe that these genes without known targeted drugs may provide new opportunities for future drug development.

[0147]

[0148] PheWAS analysis further depicted the gene-related profiles of candidate drug targets such as CA1, MMP12, and PLEKHH2, etc. Figure 11 As shown in the analysis, these candidate genes exhibit broad phenotypic associations, linking to phenotypes in multiple physiological systems, including infection, tumors, immunity, metabolism, the nervous system, and the cardiovascular system, indicating that these genes possess pleiotropic characteristics. However, all gene-phenotype associations were weak (p-value > 1 × 10⁻⁶). -6 The results did not reach genome-wide significance, which is consistent with the characteristics of "minor polygenicity". This analysis directly translates our mechanistic insights into tangible therapeutic hypotheses, providing a shortcut to clinical testing.

[0149] In this embodiment, a multi-omics approach is used to identify predictive factors for Parkinson's disease with mild cognitive impairment (PD-MCI), integrating transcriptomics and proteomics data to test models based on relevant biomarkers and gain insights.

[0150] The machine learning model built after screening differential features using LASSO and elastic network regression, while improving predictive performance compared to models using only basic demographic factors (sex, age, years of education), showed limited improvement. Transcriptomics features contributed more to the performance improvement than proteomics features. Interestingly, integrating transcriptomics and proteomics features did not produce superior predictive results, but rather a trade-off between the two. We propose several potential reasons for this. First, the core reason for the weaker proteomics performance may lie in the dynamic nature and detection characteristics of proteins. The underlying pathology of PD-MCI conversion may initially manifest as gene activation or repression at the transcriptomic level, with corresponding protein expression potentially lagging by weeks or even months. Furthermore, the longer half-life of proteins may prevent them from fully reflecting the risk of PD-MCI conversion at baseline. Second, the narrower dynamic range and lower diversity of proteomics detection may mask low-abundance key molecules, resulting in relatively limited detection capabilities compared to transcriptomics. In addition, the relatively limited sample size in this study further restricts the effective quantification of low-abundance molecules by proteomics, where stability and data noise issues may be more prominent. The trade-off observed when directly integrating two omics datasets likely stems primarily from signal conflict and data heterogeneity. The temporal scale differences between the two omics layers, coupled with the potential risk of overfitting in transcriptome data, may introduce potential interference and cancellation, preventing the expected complementary synergy.

[0151] SHAP analysis showed that age at PD diagnosis is a key predictor of MCI, consistent with previous research, highlighting the importance of timely intervention, especially in high-risk patients. Interestingly, longer duration of education was a protective factor for cognition in our study, possibly because higher cognitive reserve provides a greater buffer against decline. Furthermore, our study confirmed that sex is a meaningful predictor, with males being identified as MCI patients more frequently, indicating a sex-specific difference in cognitive impairment.

[0152] Among transcriptomic genes, PLEKHH2 showed the highest predictive weight. PLEKHH2 was significantly correlated with scores on multiple tests in visuospatial and delayed visual memory. Single-cell RNA data from patients with PDD showed that its RNA expression was elevated in oligodendrocytes, microglia, and excitatory neurons, but downregulated in inhibitory neurons. These findings collectively suggest that PLEKHH2 may play a potential role in extracellular matrix and cell dynamics-related processes within the brain.

[0153] Among proteins, carbonic anhydrase 1 (CA1) showed the highest predictive weight. The closely related CA3 was also among the proteins with significantly differential expression. As a popular drug target, CA1 already has several classic clinically available drugs. Two FDA-approved, retargeted carbonic anhydrase inhibitors—acetazolamide and metronidazole—were found to correct proteomic abnormalities, effectively prevent early molecular pathology associated with cerebral amyloid angiopathy (CAA) in Tg-SwDI mice carrying pathogenic mutations and exhibiting Alzheimer's disease (AD)-like cognitive deficits and severe CAA, maintain synaptic stability, and suppress neuroinflammation.

[0154] Pathway enrichment analysis of features showing significant intergroup differences in proteomics and transcriptomics points to alterations in key functions such as axonal guidance and synaptic transmission. Notably, the gene sets corresponding to these significant features from the two omics layers do not overlap. Trajectory clustering analysis of proteomic data and correlation analysis with clinical indicators further suggest the potential functional impact of significant upregulation or downregulation of related proteins. Due to the lack of GWAS data for PD-MCI and single-cell sequencing data specifically for PD-MCI, we utilized relevant data from PDD patients and healthy controls to further validate machine learning-derived features, revealing causal genetic associations and enrichment patterns across different cell types. Finally, using the screened genes, we preliminarily identified potential drug candidates that may provide insights for future therapeutic interventions for PD-MCI or PDD.

[0155] Machine learning models have demonstrated depth in identifying predictive features that go beyond traditional statistical methods, particularly in capturing complex nonlinear relationships between variables. Based on baseline omics data, this study identified several robust predictors through multivariate association analysis. These biomarkers have potential applications in precision medicine and routine clinical practice. Our findings may provide clinicians with novel decision support tools to help identify high-risk patients early and facilitate the clinical translation of timely intervention and personalized management strategies.

[0156] The svmRadialCost model demonstrates robust performance through rigorous internal validation (repeated 5-fold cross-validation and independent test sets).

[0157] In summary, through a tightly integrated, multi-stage analysis, we constructed a preliminary diagnostic model for PD-MCI (from PD patients) and elucidated its presumed pathological mechanisms using multiple approaches. We provide strong genetic evidence for a causal role of WDR1 in PD-related cognitive decline and identify high-priority therapeutic development targets. The svmRadialCost model (AUC = 0.7598, sensitivity = 82.4%, specificity = 66.7%) provides a promising tool for early risk stratification, requiring external validation and prospective clinical evaluation in independent cohorts prior to clinical deployment. Ultimately, this work provides a comprehensive resource guiding the transition from understanding the mechanisms of PD-MCI to targeted interventions for PD patients at risk of progressing to PD-MCI. The "neuroconnectopathy" mechanism identified here opens new avenues for therapeutic development, while drug repositioning candidates (MEDRONIC ACID, ZILEUTON, ATENOLOL) establish clear next steps for preclinical validation and clinical trials.

Claims

1. A method for screening PD-MCI biomarkers, characterized in that, It includes the following steps: S1. A well-defined PD-MCI cohort was established using the Movement Disorders Association (MDS) Level II diagnostic criteria and longitudinal cognitive assessment; S2. Integrating multi-omics analysis of plasma transcriptomics and cerebrospinal fluid proteomics to identify molecular signatures of PD-MCI progression; S3. Apply systems biology approaches to understand the underlying biological mechanisms of PD-MCI; S4. Systematically evaluate several machine learning models to select the best algorithm for PD-MCI risk prediction and obtain potential biomarkers; S5. Mendelian randomization was used to explore potential causal relationships, and single-cell nuclear RNA sequencing was used to verify the expression of biomarkers in brain cell types.

2. The method for screening PD-MCI biomarkers as described in claim 1, characterized in that: The PD-MCI markers include LRPPRC, WDR1, EIF3H, PLEKHH2, SLC35G2, UTS2, MMP12, MSMB, CA1, and POSTN.

3. The method for screening PD-MCI biomarkers as described in claim 1, characterized in that: Based on the finally selected biomarkers, therapeutic candidate drugs are obtained by calculating drug repositioning.

4. The method for screening PD-MCI biomarkers as described in claim 1, characterized in that: In step S3, the systems biology approach includes: functional enrichment, PPI network analysis, and longitudinal proteomic trajectory analysis.

5. The method for screening PD-MCI biomarkers as described in claim 1, characterized in that: Based on the S2-S3 steps, predict the preclinical window period.

6. The method for screening PD-MCI biomarkers as described in claim 4, characterized in that: Predict drugs for early intervention based on the preclinical window period.

7. The method for screening PD-MCI biomarkers as described in claim 1, characterized in that: In S4, the optimal algorithm is the svmRadialCost model.

8. The application of at least one of the biomarkers PLEKHH2, CA1, WDR1, MSMB, and POSTN in the preparation of PD cognitive impairment diagnostic / screening products.

9. Application of at least one biomarker among PLEKHH2, CA1, WDR1, MSMB, and POSTN in the preparation of PD cognitive impairment classification products.

10. The application of at least one of the biomarkers PLEKHH2, CA1, WDR1, MSMB, and POSTN as intervention targets in the preparation of drugs for the prevention or treatment of PD-MCI.