Parkinson's disease early-stage biomarker auxiliary screening method based on multi-omics data
By integrating multi-omics data and screening biomarker combinations using machine learning algorithms, the problem of insufficient sensitivity and specificity in the early diagnosis of Parkinson's disease has been solved, enabling early identification of high-risk individuals and buying time for early intervention and treatment of the disease.
Patent Information
- Application Number
- CN202511711053.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-20
- Publication Date
- 2026-02-13
AI Technical Summary
Current technologies lack objective biomarkers that can reliably diagnose Parkinson's disease in the absence of motor symptoms. Single biomarkers are insufficient to meet the sensitivity and specificity requirements for early diagnosis, and existing methods such as the SAA technique are highly invasive and cannot be applied on a large scale.
Using a multi-omics data integration approach, proteomic and metabolomic analyses are performed on blood, plasma, serum, or cerebrospinal fluid samples. Machine learning algorithms such as LASSO logistic regression, random forest, and genetic algorithms are combined to screen for the optimal combination of biomarkers and establish a classification model for early diagnosis.
It significantly improves the accuracy and reliability of early diagnosis of Parkinson's disease, can identify high-risk individuals before symptoms appear, supports early intervention, and the detection method is convenient and efficient.
Smart Images

Figure CN121528295A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of Parkinson's disease technology, and more particularly to a method for assisting in the screening of early biomarkers for Parkinson's disease based on multi-omics data. Background Technology
[0002] Diagnosis of Parkinson's disease typically relies on clinical symptoms, and the appearance of significant motor symptoms often indicates an advanced stage of the disease. Currently, there is a lack of objective biomarkers that can reliably diagnose Parkinson's disease in the asymptomatic stage. This significantly limits early intervention and clinical trials, as by the time patients have noticeable symptoms, more than 60% of the dopamine neurons in the substantia nigra of the midbrain may have already been lost. Therefore, identifying biomarkers that change in the preclinical or very early stages of Parkinson's disease is crucial for early detection and timely intervention.
[0003] Potential early biomarkers for Parkinson's disease include various types: imaging findings (such as mild abnormalities in dopamine transporter scans), proteins or metabolites in body fluids, and specific gene mutations / expression changes. Among these, body fluid biomarkers (blood, cerebrospinal fluid, etc.) are considered ideal screening methods due to their relatively convenient acquisition and repeatable testing. However, to date, no single biomarker has been proven to have sufficient sensitivity and specificity for the early diagnosis of Parkinson's disease. Biomarkers such as α-synuclein and tau protein levels in cerebrospinal fluid and uric acid levels in blood have some correlation but are insufficient as independent diagnostic tools. Recently, a technique called α-synuclein seed amplification assay (SAA) has been developed to detect abnormal α-synuclein in body fluid samples, showing promise as a highly specific pathological biomarker for Parkinson's disease. However, SAA requires lumbar puncture to obtain cerebrospinal fluid, which is invasive and not yet suitable for large-scale application. Furthermore, while neurofilament light chains (NfL) in the blood have been found to increase with the progression of Parkinson's disease, NfL also rises in other neurodegenerative diseases (such as Alzheimer's disease), thus lacking sufficient disease specificity. Therefore, a single biomarker is often insufficient, and a combination of biomarkers is needed to improve diagnostic efficacy.
[0004] The development of multi-omics research has provided new avenues for discovering such combined biomarkers. Multi-omics refers to the simultaneous acquisition of data from multiple levels, including the genome, transcriptome, proteome, and metabolome, to comprehensively analyze disease-related biological changes. For Parkinson's disease, changes at different omics levels may be interconnected: for example, mutations in certain pathogenic genes can lead to abnormal expression of specific proteins, resulting in changes in metabolites. If the products of these linked changes can be detected early in the disease's development, it may be possible to identify high-risk individuals before symptoms appear. Existing research has attempted to use machine learning to find diagnostic features of Parkinson's in multimodal data. For example, a 2024 proteomics study identified a combination of eight proteins in plasma that could predict Parkinson's disease up to seven years before the onset of motor symptoms. These eight proteins (such as the granulin precursor Granulin, MBL-2, adhesion molecule ICAM-1, and complement C3) are considered to reflect some key pathways in the early stages of the disease, such as inflammatory responses and complement activation. The model, combining the expression levels of these proteins, correctly identified 79% of preclinical Parkinson's individuals (primarily RBD patients). Furthermore, an increasing number of metabolomics studies report significant differences in the blood metabolic profiles of Parkinson's patients compared to healthy individuals, such as abnormal levels of uridine and sphingolipids. However, these differences exhibit considerable individual variability and require integration with other indicators to enhance their diagnostic value. Summary of the Invention
[0005] To address the shortcomings of the aforementioned technologies in identifying the most valuable set of biomarkers for early identification of Parkinson's disease from a massive amount of multi-omics candidate features and demonstrating that the sensitivity and specificity of this set are sufficient to meet the requirements of early diagnosis in clinical screening and diagnosis, this invention provides an auxiliary screening method for early Parkinson's disease biomarkers based on multi-omics data.
[0006] To achieve the above objectives, this invention provides a method for assisting in the screening of early biomarkers for Parkinson's disease based on multi-omics data, comprising the following steps:
[0007] S1, collect biological samples from subjects, including Parkinson's disease patients, people at high risk of Parkinson's disease, and healthy controls;
[0008] S2, Perform multi-omics detection on the biological sample to obtain proteomic data and metabolomic data;
[0009] S3, preprocess the multi-omics data;
[0010] S4. Univariate analysis was performed on the preprocessed multi-omics data to initially screen out candidate biomarkers.
[0011] S5. Input the candidate markers into the machine learning model for multivariate screening, and use machine learning algorithms to select the optimal feature combination.
[0012] S6, a classification model is built based on a selected combination of markers to distinguish early stages of Parkinson's disease;
[0013] S7. Evaluate model performance on an independent validation set, and calculate sensitivity, specificity, and AUC values.
[0014] S8, develops a marker detection panel to transform the selected marker combinations into an operable detection method.
[0015] As an improvement of the present invention, in S1, the biological sample includes blood, plasma, serum or cerebrospinal fluid.
[0016] As an improvement of the present invention, in S2, the protein concentration of the proteomic data is determined by multiplex ELISA or mass spectrometry quantitative detection method, and the metabolomic data is determined by chromatography-mass spectrometry targeted assay method.
[0017] As an improvement of the present invention, the proteomic data is quantitatively detected using a mass spectrometry method comprising the following formula:
[0018]
[0019] ,in Indicates the first Corrected concentration of each protein, Its original peak area, This represents the total peak area of all detected proteins in the sample. This is the preset scaling factor;
[0020] The metabolomics data were obtained by detecting metabolite concentrations using a Targeted Assay based on chromatography-mass spectrometry, including the following formulas:
[0021]
[0022] in This represents the peak area ratio of the metabolite to the internal standard. The concentration to be measured, and The slope and intercept are obtained by fitting a standard with a known concentration.
[0023] As an improvement of the present invention, in S3, the pretreatment includes removing undetectable proteins and metabolites from the sample in proteomics and metabolomics, and using missing value imputation and batch effect correction.
[0024] As an improvement of the present invention, in S4, the multi-omics indicators of the candidate biomarkers are combined to construct a candidate biomarker dataset containing multiple features.
[0025] As an improvement of the present invention, in S5, the machine learning algorithm adopts the LASSO logistic regression algorithm, random forest algorithm or genetic algorithm. The machine learning model is trained based on the candidate biomarker dataset, and the algorithm with the best performance is selected through cross-validation for feature screening to determine the biomarker combination with the most predictive value for the early diagnosis of Parkinson's disease.
[0026] As an improvement to the present invention, the LASSO logistic regression objective function includes the following formula:
[0027]
[0028] in For the sample size, This represents the total number of candidate biomarkers (features). For the first The true class labels of each sample Model predicts probability, These are the characteristic regression coefficients. is the hyperparameter for L1 regularization intensity.
[0029] As an improvement to the present invention, the random forest algorithm includes the following formula:
[0030] For a random forest model containing T decision trees, the j-th feature The importance is defined as the average reduction in Gini impurity resulting from splitting this feature across all trees:
[0031]
[0032] in Using features in the t-th tree The set of all nodes to be split; for any split node Its Gini reduction is:
[0033]
[0034] in The total number of samples in the training set. For nodes The number of samples in , The number of samples in the left and right child nodes after the split. For nodes Medium category The proportion. According to... Sort in descending order and select the first few. The markers constitute the candidate combination.
[0035] As an improved embodiment of the present invention, the genetic algorithm includes the following formula:
[0036]
[0037] in Describes a subset of candidate features of length . binary vectors ( (Total number of candidate biomarkers) = 1 indicates selecting the first A landmark, = 0 indicates removal.
[0038]
[0039] in Indicates to retain only The first feature of the non-zero position correspondence One sample, For sample labels, To pass The classification performance metrics obtained from cross-validation, genetic algorithms Guided by the principle of selection, crossover, and mutation operations, the algorithm searches for the element with the highest fitness. The non-zero components are the optimal combination of markers.
[0040] The beneficial effects of this invention are as follows: Compared with the prior art, this invention provides an auxiliary screening method for early biomarkers of Parkinson's disease based on multi-omics data. By integrating multi-omics data, this invention overcomes the limitations of traditional single-omics research and can comprehensively capture the early biological changes of Parkinson's disease from multiple levels such as genes, transcription, proteins, and metabolism. This significantly improves the accuracy and reliability of biomarker screening and provides data support for the early auxiliary diagnosis of Parkinson's disease. By adopting a biomarker combination strategy, compared with traditional single biomarker detection, the accuracy, sensitivity, and specificity of auxiliary diagnosis are significantly enhanced. By combining multiple biomarkers with predictive value, high-risk individuals can be identified more accurately before symptoms appear, gaining valuable time for early intervention and treatment of Parkinson's disease. Attached Figure Description
[0041] Figure 1 This is a flowchart of the method of the present invention;
[0042] Figure 2 This is a module interaction diagram of the present invention;
[0043] Figure 3 This is a schematic diagram illustrating the discriminative effect of the combination of early Parkinson's disease biomarkers of the present invention. Detailed Implementation
[0044] To more clearly illustrate the present invention, the invention will be further described below with reference to the accompanying drawings.
[0045] In the following description, specific examples are given to provide a more in-depth understanding of the invention. It is obvious that the described embodiments are merely some, not all, of the embodiments of the invention. It should be understood that the specific embodiments described are for illustrative purposes only and are not intended to limit the scope of the invention.
[0046] It should be understood that when the terms “comprising” and / or “including” are used in this specification, they indicate the presence of the said feature, integral, step, operation, element, or component, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, or combinations thereof.
[0047] Please see Figures 1-3 The present invention provides a method for assisting in the screening of early biomarkers for Parkinson's disease based on multi-omics data, comprising the following steps:
[0048] S1, collect biological samples from subjects, including Parkinson's disease patients, people at high risk of Parkinson's disease, and healthy controls;
[0049] S2 performs multi-omics testing on biological samples to obtain proteomic and metabolomic data.
[0050] S3, preprocesses multi-omics data;
[0051] S4. Univariate analysis was performed on the preprocessed multi-omics data to initially screen out candidate biomarkers.
[0052] S5. Input the candidate markers into the machine learning model for multivariate screening, and use machine learning algorithms to select the optimal feature combination.
[0053] S6, a classification model is built based on a selected combination of markers to distinguish early stages of Parkinson's disease;
[0054] S7. Evaluate model performance on an independent validation set, and calculate sensitivity, specificity, and AUC values.
[0055] S8, develops a marker detection panel to transform the selected marker combinations into an operable detection method.
[0056] First, this invention uses biological samples from different subjects, including Parkinson's disease patients, high-risk pre-Parkinson's disease individuals, and healthy controls, for comparative analysis. Through this comparison, the unique biological changes characteristic of the early stage and high-risk pre-Parkinson's disease stage can be captured more accurately. After collecting the biological samples, multi-omics detection technology is used to obtain comprehensive and accurate proteomic and metabolomic data.
[0057] Secondly, the preprocessing of the collected multi-omics data is crucial. Preprocessing can remove undetectable proteins and metabolites from the samples, effectively reduce noise interference in the data, improve data quality, and further ensure the integrity and consistency of the data, making the data between different samples comparable and creating conditions for accurate screening of biomarkers in the future.
[0058] After preprocessing the multi-omics data, a detailed analysis is performed on the preprocessed multi-omics data to initially screen out candidate biomarkers. This is similar to sifting out potentially valuable information from a vast ocean of data, that is, extracting key and useful feature information to narrow down the detection range, eliminate useless information, and improve detection efficiency and accuracy, so as to more accurately identify the risk of Parkinson's disease in subsequent analysis and judgment.
[0059] The candidate biomarkers are then input into a machine learning model, where multivariate screening is performed using machine learning algorithms. The best-performing algorithm is selected from among many for feature selection, thereby determining the biomarker combination with the most predictive value for the early diagnosis of Parkinson's disease, thus improving the accuracy and efficiency of the screening. Subsequently, a classification model is built based on the selected biomarker combination, which can accurately distinguish the early stages of Parkinson's disease, providing a strong basis for the early auxiliary diagnosis of the disease. The model performance is evaluated in an independent validation set, and the reliability and effectiveness of the model are comprehensively evaluated by calculating indicators such as sensitivity, specificity, and AUC value, ensuring that the model can accurately play its role in practical applications.
[0060] Finally, a biomarker detection panel was developed, transforming the selected biomarker combinations into an operational detection method. This innovative achievement makes the detection of early biomarkers for Parkinson's disease more convenient and efficient, and is expected to be widely used in clinical screening and diagnosis, providing important support for the early detection and timely intervention of Parkinson's disease, with significant clinical application value and social significance.
[0061] This invention integrates multi-omics data to more comprehensively capture early biological changes in Parkinson's disease, thereby improving the accuracy and reliability of biomarker screening. Compared with traditional single biomarker detection, the biomarker combination strategy adopted in this method significantly enhances the sensitivity and specificity of diagnosis, helping to identify high-risk individuals before symptoms appear. In addition, this method, combined with machine learning algorithms, can automatically mine the most valuable feature combinations from massive biomarker data, further improving screening efficiency and diagnostic performance. Evaluation through independent validation sets has demonstrated its practicality and effectiveness in clinical screening and auxiliary diagnosis, providing a new and powerful auxiliary tool for the early intervention and treatment of Parkinson's disease.
[0062] In this embodiment, biological samples include blood, plasma, serum, or cerebrospinal fluid. The selection of these biological samples is based on scientific evidence and has clinical applicability. The collection of blood, plasma, and serum samples is not only simpler than other samples, but also causes less trauma to the patient. Furthermore, they can reflect the overall metabolism and protein expression of the body. Although the collection of cerebrospinal fluid is relatively complex, it is closer to the central nervous system and is more accurate in detecting neurological markers closely related to Parkinson's disease. By collecting these different types of biological samples, biological information related to Parkinson's disease can be obtained from multiple perspectives, and early biomarkers can be accurately screened.
[0063] In this embodiment, protein concentrations were determined using multiplex ELISA or mass spectrometry quantitative detection methods, while metabolomics concentrations were determined using chromatography-mass spectrometry (GC-MS) targeted assay methods. Multiplex ELISA methods offer high sensitivity and specificity, accurately measuring the concentrations of multiple proteins in a sample. Their high throughput allows for the processing of large numbers of samples in a short time, enabling large-scale screening. Mass spectrometry, with its high precision and high resolution, plays a crucial role in proteomics research. It accurately measures the molecular weight of proteins, thereby identifying their types and concentrations and quantifying them precisely. GC-MS targeted assay methods in metabolomics detection can accurately detect specific metabolites, effectively eliminating interference from other substances and improving the accuracy of metabolite concentration measurements. These detection technologies enable the acquisition of high-quality proteomics and metabolomics data, ensuring the accuracy of subsequent detection.
[0064] In this embodiment, the proteomic data were quantitatively detected using mass spectrometry, including the following formula:
[0065]
[0066] ,in Indicates the first Corrected concentration of each protein, Its original peak area, This represents the total peak area of all detected proteins in the sample. This is the preset scaling factor;
[0067] Metabolomics data, obtained through a chromatography-mass spectrometry targeted assay to determine metabolite concentrations, include the following formulas:
[0068]
[0069] in This represents the peak area ratio of the metabolite to the internal standard. The concentration to be measured, and The slope and intercept obtained by fitting with known concentration standards are used to infer the actual concentration of metabolites in the sample. This quantitative result constitutes the input matrix for subsequent preprocessing and is used in conjunction with proteomics data to enter a unified missing value imputation and batch effect correction process. This ensures the consistency of multi-omics data in terms of numerical scale and systematic error, and supports the joint modeling of highly reliable biomarkers.
[0070] In this embodiment, the preprocessing method for multi-omics data involves removing undetectable proteins and metabolites from samples in proteomics and metabolomics, and employing missing value imputation and batch effect correction. Missing value imputation is used to handle unobserved values caused by insufficient detection sensitivity or concentrations below the instrument's detection limit. Specific operations can employ strategies based on data distribution characteristics: for missing low-abundance substances, they are often replaced with a certain proportion (e.g., 1 / 2 or 1 / 5) of the smallest non-zero value of that variable in all samples; or the K-nearest neighbor (KNN) algorithm can be used to weighted estimate missing entries based on the expression patterns of similar samples; multiple imputation or base... Machine learning-based matrix completion methods reconstruct missing data while preserving covariant relationships between variables, generating a complete quantitative matrix for subsequent analysis. After missing value imputation, batch effect correction is necessary to further eliminate systematic biases introduced by non-biological factors such as experimental batches, operation time, or instrument status. Typical operations include using empirical Bayesian methods such as ComBat to model known batch labels as covariates, correcting for mean shift and variance variability in the data, while preserving the influence of key biological variables such as disease grouping. Alternatively, linear mixed-effects models or SVA (Surrogate Variable Analysis) strategies can be combined to identify and remove implicit technical variation factors. Through these two collaborative steps, the consistency and reliability of multi-batch omics data are significantly improved, providing high-quality input for high-precision biomarker discovery or diagnostic model construction.
[0071] In this embodiment, multi-omics indicators of candidate biomarkers are combined to construct a candidate biomarker dataset containing multiple features. Combining indicators from different omics levels to construct a dataset can fully leverage the advantages of multi-omics data and comprehensively reflect the biological changes in the early stages of Parkinson's disease. This is because single-omics data often only provides information on one aspect, while Parkinson's disease is a complex disease whose pathogenesis involves multiple biological processes and pathways. By combining multi-omics indicators, comprehensive analysis can be performed from multiple levels such as genes, transcription, proteins, and metabolism, more accurately capturing key features related to the early stages of the disease. Furthermore, the constructed candidate biomarker dataset containing multiple features provides a rich data foundation for training machine learning models, which helps to learn and screen biomarker combinations with greater predictive value.
[0072] In this embodiment, the machine learning algorithm employs LASSO logistic regression, random forest, or genetic algorithms. The machine learning model is trained on a candidate biomarker dataset, and cross-validation is used to select the best-performing algorithm for feature selection, thereby determining the biomarker combination with the most predictive value for the early diagnosis of Parkinson's disease. LASSO logistic regression, by introducing an L1 regularization term, automatically compresses the coefficients of some irrelevant features to zero during feature selection, achieving feature sparsity and selecting biomarkers that play a crucial role in the early diagnosis of Parkinson's disease, effectively avoiding overfitting and improving the model's generalization ability. Random forest algorithms construct multiple decision trees and perform ensemble learning, using each decision tree to vote on the classification results of samples, ultimately determining the sample's classification. Random forest algorithms can handle high-dimensional data and have good robustness to noise and outliers, capable of mining biomarker combinations with stable predictive performance from complex candidate biomarker datasets. Genetic algorithms simulate selection, crossover, and mutation operations in the natural biological evolution process, continuously iterating and optimizing the candidate biomarker dataset. The optimal feature combination for early diagnosis of Parkinson's disease is determined by a genetic algorithm that searches the entire feature space for a global search capability, avoiding getting trapped in local optima. This identifies the most predictive combination of biomarkers for early diagnosis. Further, cross-validation is used to prevent overfitting during model training. In this process, data from different omics can be treated equally or given different weights; for example, metabolites may have higher stability and measurability in blood, requiring comprehensive consideration. The goal is to find a relatively simple biomarker combination that best distinguishes early Parkinson's disease states on the training set. Assuming m biomarker combinations are selected, which can be mixtures from different omics (e.g., 3 proteins + 2 metabolites), the candidate biomarker dataset is divided into multiple subsets. One subset is used as the validation set, and the remaining subsets are used as the training set for model training and evaluation. The performance of different algorithms on the validation set is compared, using metrics such as accuracy, recall, and F1 score. The algorithm with the best performance is selected for feature selection, ultimately ensuring that the determined biomarker combination has the highest predictive value in the early auxiliary diagnosis of Parkinson's disease.
[0073] Please see Figure 3 , Figure 3 This study demonstrates the discriminative performance of different machine learning models (KNN, Lasso, RandomForest, SVM, and XGBoost) in distinguishing early-stage Parkinson's disease patients from healthy controls based on specific combinations of biomarkers. The results show that these models have high accuracy and reliability for early diagnosis of Parkinson's disease, demonstrating the potential clinical value of the selected biomarker combinations in early screening.
[0074] In this embodiment, the LASSO logistic regression objective function includes the following formula:
[0075]
[0076] in For the sample size, This represents the total number of candidate biomarkers (features). For the first The true class labels of each sample Model predicts probability, These are the characteristic regression coefficients. is the hyperparameter for L1 regularization intensity.
[0077] In this embodiment, the random forest algorithm includes the following formula:
[0078] For a random forest model containing T decision trees, the j-th feature The importance is defined as the average reduction in Gini impurity resulting from splitting this feature across all trees:
[0079]
[0080] in Using features in the t-th tree The set of all nodes to be split; for any split node Its Gini reduction is:
[0081]
[0082] in The total number of samples in the training set. For nodes The number of samples in , The number of samples in the left and right child nodes after the split. For nodes Medium category The proportion. According to... Sort in descending order and select the first few. The markers constitute the candidate combination.
[0083] In this embodiment, the genetic algorithm includes the following formula:
[0084]
[0085] in Describes a subset of candidate features of length . binary vectors ( (Total number of candidate biomarkers) = 1 indicates selecting the first A landmark, = 0 indicates removal.
[0086]
[0087] in Indicates to retain only The first feature of the non-zero position correspondence One sample, For sample labels, To pass The classification performance metrics obtained from cross-validation, genetic algorithms Guided by the principle of selection, crossover, and mutation operations, the algorithm searches for the element with the highest fitness. The non-zero components are the optimal combination of markers.
[0088] This invention also provides an early biomarker-assisted screening system for Parkinson's disease based on multi-omics data, comprising:
[0089] The data acquisition module is used to acquire biological samples from subjects and perform multi-omics testing on the biological samples to obtain proteomic and metabolomic data.
[0090] The data preprocessing module is used to preprocess proteomic and metabolomic data.
[0091] The data filtering module is used to perform univariate analysis on the preprocessed multi-omics data, initially screen out candidate biomarkers, and input the candidate biomarkers into the machine learning model for multivariate filtering, using machine learning algorithms to select the optimal feature combination.
[0092] The model processing module is used to build a classification model based on a selected combination of markers to distinguish early stages of Parkinson's disease.
[0093] The performance evaluation module is used to evaluate model performance on an independent validation set and calculate sensitivity, specificity, and AUC values.
[0094] The detection panel development module is used to develop marker detection panels, transforming the selected marker combinations into operable detection methods.
[0095] The present invention also provides an electronic device, including a memory and a processor, and computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform steps such as in a method for assisting screening of early biomarkers for Parkinson's disease based on multi-omics data.
[0096] The present invention also provides a computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform steps as described in a method for assisting in the screening of early biomarkers for Parkinson's disease based on multi-omics data.
[0097] The following are embodiments of the present invention in practical operation:
[0098] Example 1: Joint screening of biomarkers using proteomics and metabolomics
[0099] For example, a prospective study included 50 participants diagnosed with Parkinson's disease in the later stages of follow-up (all without motor symptoms at baseline but exhibiting prodromal manifestations such as RBD), and 50 age- and sex-matched healthy controls. Plasma samples were collected from all participants at baseline. Twenty patients diagnosed with early-stage Parkinson's disease (with mild motor symptoms, Hoehn-Yahr stage I) were also included as a reference. The abundance of 500 proteins and 200 metabolites in plasma was simultaneously measured using a targeted mass spectrometry platform. Differential analysis showed that in the prodromal group and control group, 15 proteins and 10 metabolites showed p < 0.01 and fold differences > 1.5. These 25 items were then used in subsequent machine learning. LASSO logistic regression was used for feature selection (aimed at distinguishing between prodromal Parkinson's disease and controls). LASSO results retained eight of the most discriminative features, including five proteins (such as granulin precursors Granulin, MBL2, ICAM-1, complement C3, and ER chaperone BiP) and three metabolites (such as uric acid, specific glycosphingolipids, and melatonin metabolites). A logistic regression model based on these eight features achieved an AUC of 0.93 on the training set. Using an independent validation set (another 20 RBD individuals, including 8 newly diagnosed Parkinson's patients during follow-up), the model achieved an AUC of 0.85, with a sensitivity of 75% and a specificity of 80%. The combined detection of these eight biomarkers was used to create a multiplex immunoassay panel, batch-testing samples in the validation set and calculating a risk score for each individual. The results showed that newly diagnosed RBD individuals generally had higher risk scores, while those without the disease had lower scores, with a significant difference between the two groups. This example demonstrates that a multi-omics approach combining proteins and metabolites can screen for promising combinations of early biomarkers, with superior discriminative performance compared to single-class biomarkers. The selection of some biomarkers overlaps with recent independent studies reporting early Parkinson's markers, such as elevated levels of Granulin and complement C3 in the blood of Parkinson's patients, which enhances the credibility of the findings.
[0100] Example 2: Prospective Clinical Validation of Biomarker Panels
[0101] This embodiment prospectively validates the combination of 8 biomarkers selected in Example 1. In an independent cohort, 100 patients with iRBD (idiopathic RBD), all without Parkinson's symptoms at baseline, were included. The levels of 8 biomarkers in their plasma were measured using a multiplex assay, and a comprehensive risk score was calculated. Individuals were divided into high-risk (top 30%) and low-risk (bottom 30%) groups based on their scores. These 100 individuals were then followed up annually with neurological examinations. After 3 years, 12 individuals were diagnosed with Parkinson's disease or DLB, with 10 from the high-risk group and 2 from the low-risk group. The cumulative incidence rate in the high-risk group was significantly higher than that in the low-risk group (p=0.01). This indicates that the biomarker panel has some predictive value for conversion in the coming years. Simultaneously, the panel was tested in 20 healthy controls, and the risk scores of the healthy controls were all very low, with no false positives for high risk, indicating good specificity. Based on these results, the researchers recommend using this biomarker panel for the clinical assessment of RBD patients to screen for those who may have progressed to Parkinson's disease earlier, facilitating intervention. This panel can also be used as part of clinical trial recruitment criteria, such as prioritizing the recruitment of high-risk RBD patients into neuroprotective agent trials to increase the probability of observing efficacy.
[0102] Example 3: Expansion and Application of Marker Combinations
[0103] Given the diversity of Parkinson's disease, this embodiment explores optimized biomarker combinations for different subtypes. For example, patients with predominantly bradykinesia and significant cognitive impairment may exhibit different biological pathways. Therefore, this invention compares Parkinson's patients with controls based on clinical subtypes (tremor-dominant vs. rigidity-dominant) using multi-omics screening. Results showed that some inflammation-related proteins (such as IL-6) were more prominent in the tremor subtype, while certain metabolites (such as sphingomyelin) showed more significant differences in the rigidity subtype. Therefore, custom panels can be developed; for example, an early screening panel for suspected tremor patients can focus on covering inflammatory markers. Nevertheless, this invention found that most key biomarkers showed consistent trends across subtypes; therefore, the final panel primarily uses a universal combination, considering subtype factors only in interpretation. In clinical application, this invention integrates biomarker testing into annual health check-up packages, providing testing services for individuals with a family history of Parkinson's or older individuals with mild symptoms. If an abnormal risk score is detected, further neurological examination is recommended. This model has performed well in a small-scale pilot program, identifying a 55-year-old individual with olfactory loss who tested positive for biomarkers. A year of follow-up revealed mild hand tremors, prompt intervention with dopaminergic and exercise therapy, resulting in slow symptom progression. Currently, this invention is collaborating with testing reagent manufacturers to apply for registration of eight biomarker kits as in vitro diagnostic products for use in the auxiliary diagnosis of Parkinson's disease. It is believed that with the accumulation of more data, the method of this invention can be continuously updated with new biomarker combinations to improve accuracy, and ultimately be used for early screening, risk assessment, and personalized medicine for Parkinson's disease.
[0104] The advantages of this invention are:
[0105] 1. This invention, by integrating multi-omics data, overcomes the limitations of traditional single-omics research. It can comprehensively capture the early biological changes of Parkinson's disease from multiple levels such as genes, transcription, proteins and metabolism, significantly improving the accuracy and reliability of biomarker screening and providing data support for the early auxiliary diagnosis of Parkinson's disease.
[0106] 2. The biomarker combination strategy adopted in this invention significantly enhances the sensitivity and specificity of diagnosis compared with traditional single biomarker detection. By combining multiple biomarkers with predictive value, high-risk individuals can be identified more accurately before symptoms appear, thus gaining valuable time for early intervention and treatment of Parkinson's disease.
[0107] 3. This invention establishes a machine learning model based on machine learning algorithms, automatically mining the most valuable feature combinations from massive biomarker data. The automated and intelligent screening process not only improves screening efficiency but also further enhances the accuracy of assisted diagnosis.
[0108] The above-disclosed embodiments are merely a few specific examples of the present invention, but the present invention is not limited thereto. Any variations that can be conceived by those skilled in the art should fall within the protection scope of the present invention.
Claims
1. A method for assisting in the screening of early biomarkers for Parkinson's disease based on multi-omics data, characterized in that, Includes the following steps: S1, collect biological samples from subjects, including Parkinson's disease patients, people at high risk of Parkinson's disease, and healthy controls; S2, Perform multi-omics detection on the biological sample to obtain proteomic data and metabolomic data; S3, preprocess the multi-omics data; S4. Univariate analysis was performed on the preprocessed multi-omics data to initially screen out candidate biomarkers. S5. Input the candidate markers into the machine learning model for multivariate screening, and use machine learning algorithms to select the optimal feature combination. S6, A classification model is built based on the selected combination of biomarkers to distinguish the risk of a subject having early Parkinson's disease; S7. Evaluate the performance of the classification model on an independent validation set, and calculate the sensitivity, specificity, and AUC value. S8, develops a marker detection panel to transform the selected marker combinations into an operable detection method.
2. The method for assisting in the screening of early biomarkers for Parkinson's disease based on multi-omics data according to claim 1, characterized in that, In S1, the biological sample includes blood, plasma, serum, or cerebrospinal fluid.
3. The method for assisting in the screening of early biomarkers for Parkinson's disease based on multi-omics data according to claim 1, characterized in that, In S2, the proteomic data are obtained by determining protein concentration using multiplex ELISA or mass spectrometry quantitative detection methods, and the metabolomic data are obtained by determining metabolite concentration using chromatography-mass spectrometry targeted assay methods.
4. The method for assisting in the screening of early biomarkers for Parkinson's disease based on multi-omics data according to claim 3, characterized in that, The proteomic data were detected using mass spectrometry quantitative detection method. Including the following formulas: ,in Indicates the first Corrected concentration of each protein, Its original peak area, This represents the total peak area of all detected proteins in the sample. The scaling factor is preset; the metabolomics data were analyzed using a Targeted Assay for metabolite concentrations, which included the following formulas; ,in This represents the peak area ratio of the metabolite to the internal standard. The concentration to be measured, and The slope and intercept are obtained by fitting a standard with a known concentration.
5. The method for assisting in the screening of early biomarkers for Parkinson's disease based on multi-omics data according to claim 1, characterized in that, In S3, the preprocessing includes removing undetectable proteins and metabolites from the sample in proteomics and metabolomics, and using missing value imputation and batch effect correction.
6. The method for assisting in the screening of early biomarkers for Parkinson's disease based on multi-omics data according to claim 1, characterized in that, In S4, the multi-omics indicators of the candidate biomarkers are combined to construct a candidate biomarker dataset containing multiple features.
7. The method for assisting in the screening of early biomarkers for Parkinson's disease based on multi-omics data according to claim 5, characterized in that, In S5, the machine learning algorithm adopts LASSO logistic regression algorithm, random forest algorithm or genetic algorithm. The machine learning model is trained based on the candidate biomarker dataset. The algorithm with the best performance is selected through cross-validation for feature screening to determine the combination of biomarkers with the most predictive value for the early diagnosis of Parkinson's disease.
8. The method for assisting in the screening of early biomarkers for Parkinson's disease based on multi-omics data according to claim 7, characterized in that, The LASSO logistic regression objective function includes the following formula: ,in For the sample size, This represents the total number of candidate biomarkers (features). For the first The true class labels of each sample Model predicts probability, These are the characteristic regression coefficients. is the hyperparameter for L1 regularization intensity.
9. The method for assisting in the screening of early biomarkers for Parkinson's disease based on multi-omics data according to claim 7, characterized in that, in, The random forest algorithm includes the following formula: For a random forest model containing T decision trees, the j-th feature The importance is defined as the average reduction in Gini impurity resulting from splitting this feature across all trees: ,in Using features in the t-th tree The set of all nodes to be split; for any split node Its Gini reduction is: ,in The total number of samples in the training set. For nodes The number of samples in , The number of samples in the left and right child nodes after the split. For nodes Medium category The proportion. According to... Sort in descending order and select the first few. The markers constitute the candidate combination.
10. The method for assisting in the screening of early biomarkers for Parkinson's disease based on multi-omics data according to claim 7, characterized in that, in, The genetic algorithm includes the following formula: ,in Describes a subset of candidate features of length . binary vectors ( (Total number of candidate biomarkers) = 1 indicates selecting the first A landmark, = 0 indicates rejection; ,in Indicates to retain only The first feature of the non-zero position correspondence One sample, For sample labels, To pass The classification performance metrics obtained from cross-validation, genetic algorithms Guided by the principle of selection, crossover, and mutation operations, the algorithm searches for the element with the highest fitness. The non-zero components are the optimal combination of markers.