Protein combination marker related to Parkinson's disease, product and application
By screening 20 protein combination biomarkers and using multi-model machine learning, the problems of subjectivity and multi-stage prediction in early screening of Parkinson's disease have been solved, achieving high-precision and interpretable multi-stage prediction, which is suitable for multi-center promotion.
Patent Information
- Application Number
- CN202511206653.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2025-12-12
AI Technical Summary
Existing early screening methods for Parkinson's disease are highly subjective, have poor repeatability, are difficult to predict multiple stages, and most rely on invasive samples, thus lacking clinical applicability.
By employing plasma differential proteomics and multi-model machine learning, 20 protein combination biomarkers, including KITM, DKK3, PHC1, and MP2K3, were screened out. Combined with multi-class support vector machine, logistic regression, and random forest models, multi-stage prediction of Parkinson's disease was achieved.
It achieves non-invasive, high-precision, and highly interpretable multi-stage prediction of Parkinson's disease, supports individual staging and disease monitoring, is suitable for multi-center promotion, and has standardized procedures and modeling stability.
Smart Images

Figure CN121114448A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of biomedicine, research or analytical materials and machine learning, and specifically to protein combinatorial biomarkers, products and applications related to Parkinson's disease (PD). Background Technology
[0002] Parkinson's disease is a neurodegenerative disease with insidious onset and slow progression. Its clinical staging often relies on the Hoehn-Yahr scale or the UPDRS (Unified Parkinson's Disease Rating Scale). However, these assessment methods have drawbacks such as high subjectivity, poor repeatability, and low sensitivity in the early stages, especially in the early and middle stages of the disease and in determining the stage boundaries, where their accuracy is insufficient and may delay the opportunity for treatment.
[0003] Some existing studies have attempted to use plasma or cerebrospinal fluid proteomics for early screening of PD, but the following technical problems still exist: 1. It only distinguishes between early-stage and control samples and cannot predict multi-stage clinical progression. 2. It mainly depends on changes in a single protein or peak value, resulting in poor stability and reproducibility; 3. Most samples use invasive methods (such as cerebrospinal fluid), which limits their clinical applicability; 4. Lack of explanatory mechanisms hinders clinical trust and validation.
[0004] Therefore, there is an urgent need to develop a non-invasive, high-throughput, stable, reliable, and interpretable intelligent diagnostic method and system that supports multi-phase prediction of PD, in order to address the shortcomings of existing methods in terms of accuracy, scalability, and practicality. Summary of the Invention
[0005] This invention aims to provide a method and system for predicting the staging of Parkinson's disease based on plasma differential proteomics and multi-model machine learning. It overcomes the problems of strong subjectivity, poor accuracy, and difficulty in multi-stage prediction in the existing technology, and realizes automatic staging prediction of Parkinson's disease from early to late stage. It has the advantages of being non-invasive, highly accurate, highly interpretable, and clinically adaptable.
[0006] Firstly, a protein combinatorial marker is provided, comprising: KITM, DKK3, PHC1, MP2K3, PERE, VGFR1, IREB2, PPP5, HEPC, H2A2B, IKKA, FBRL, FAK1, EDIL3, NCK1, DRC11, GCYA1, TTC27, CD63 and LHPL2.
[0007] Secondly, the application of a reagent for detecting a combination of protein biomarkers in the preparation of products for diagnosing Parkinson's disease is provided, wherein the combination of protein biomarkers includes: KITM, DKK3, PHC1, MP2K3, PERE, VGFR1, IRB2, PPP5, HEPC, H2A2B, IKKA, FBRL, FAK1, EDIL3, NCK1, DRC11, GCYA1, TTC27, CD63, and LHPL2.
[0008] Thirdly, a kit is provided that includes a method for detecting the protein combination markers described in the first aspect.
[0009] Fourthly, the kit described in the third aspect is provided for use in the preparation of products for detecting Parkinson's disease.
[0010] Fifthly, a product for detecting Parkinson's disease is provided, the product comprising primers, probes, antibodies, aptamers or chips that are specific to the protein combination markers described in the first aspect.
[0011] Sixthly, a computer program product related to Parkinson's disease is provided, the computer program product being used to diagnose the risk of a subject having Parkinson's disease, including the following steps: The relative abundance of each protein biomarker in the plasma sample of the subject to be tested is obtained; the protein biomarkers include KITM, DKK3, PHC1, MP2K3, PERE, VGFR1, IREB2, PPP5, HEPC, H2A2B, IKKA, FBRL, FAK1, EDIL3, NCK1, DRC11, GCYA1, TTC27, CD63, and LHPL2; The relative abundance of each protein biomarker is input into the prediction model for calculation, and the stage classification results are output. Based on the staging and classification results, the diagnosis or prediction can be made regarding whether the subject has Parkinson's disease, is at risk of having Parkinson's disease, and the extent of disease progression.
[0012] In one possible implementation, the prediction model includes a multi-class support vector machine, a logistic regression model, or a random forest model.
[0013] In one possible implementation, the method for obtaining the prediction model includes: Obtain blood samples from Parkinson's disease patients and healthy individuals; The protein concentration in blood samples was detected to obtain a protein dataset; Differential protein screening based on protein datasets yields candidate datasets; The first machine learning model is used to filter features in the candidate dataset to identify protein biomarkers for a single model. The intersection of the protein biomarkers selected by the first machine learning model is used to obtain the protein combination biomarkers. The relative abundance of each protein biomarker in the candidate dataset is used as input, and the stage classification results are used as labels to train the second machine learning model. The model performance is then evaluated to obtain the prediction model.
[0014] Furthermore, the first machine learning model includes random forest, support vector machine, and LASSO regression.
[0015] Furthermore, the second machine learning model includes a multi-class support vector machine, a logistic regression model, or a random forest model.
[0016] Compared with the prior art, the present invention has the following beneficial effects: This invention provides a protein biomarker comprising 20 proteins, which, along with related computational methods, can be used to diagnose whether a subject has Parkinson's disease.
[0017] The protein biomarkers provided by this invention enable multi-stage prediction. This is the first time that proteome-based prediction of early, middle, and late-stage Parkinson's disease has been achieved.
[0018] The method provided by this invention has high interpretability and outputs important biomarkers when combined with SHAP analysis; The method provided by this invention has good clinical application value, supports individual staging, disease monitoring and early intervention, and is suitable for multi-center promotion. It also features standardized procedures, stable modeling, and translational potential. Attached Figure Description
[0019] Figure 1 The degree of influence of 20 proteins on model predictions; Figure 2 A comparison chart of the prediction performance of different machine learning models (SVM, LR, RF).
[0020] Figure 3 A model diagram for predicting the staging of Parkinson's disease patients. Detailed Implementation
[0021] Currently, existing screening methods for Parkinson's disease suffer from problems such as high subjectivity, low sensitivity, and insufficient interpretability.
[0022] In view of this, the present invention provides protein combinatorial biomarkers, products and applications related to Parkinson's disease.
[0023] The present invention will be further described below with reference to specific embodiments, but the content of the present invention is not limited thereto.
[0024] Example 1 This embodiment proposes a method for screening combinatorial protein biomarkers associated with Parkinson's disease based on plasma differential proteomics, including: S1. Plasma Sample Collection and Preprocessing Specifically, plasma samples were collected from 200 PD patients and 101 healthy controls, and uniform reduction, alkylation, enzymatic digestion, reaction termination, desalting and preservation operations were performed to ensure the consistency and reliability of sample processing.
[0025] S2, quantitative proteomics analysis Specifically, Astral DIA (Data-Independent Acquisition) high-throughput mass spectrometry technology, combined with the Vanquish Neo liquid chromatography system, was used to achieve highly sensitive quantitative detection of proteins in plasma.
[0026] S3. Data Preprocessing and Batch Effect Correction Specifically, Protein Discoverer software was used to identify and quantify proteins in the raw mass spectrometry data. The ComBat (sva package) method was then used to correct batch-wise systematic errors during the experiment, ensuring data comparability.
[0027] S4. Differential Protein Screening Based on the Hoehn-Yahr (HY) scoring system, 200 PD patients were reclassified into four stage subgroups for subsequent differential protein analysis. The specific definitions are as follows: HY score <1, defined as stage "no disease"; HY score between 1.0 and 2.5, defined as stage "early"; HY score between 3.0 and 4.0, defined as stage "intermediate"; and HY score = 5, defined as stage "late".
[0028] It should be noted that the Hoehn-Yahr classification is a clinical staging system used to assess the severity of Parkinson's disease (PD), proposed in 1967 by American neurologists Melvin Y. Hoehn and Raymond D. Yahr. This system divides the disease into five stages based on the impact of a patient's motor symptoms (such as tremor, rigidity, and bradykinesia) on their daily life, helping doctors determine disease progression, develop treatment plans, and assess prognosis.
[0029] Subsequently, the `limma` and `edgeR` packages in R were used to perform differential protein expression data analysis between groups under both large and small sample conditions. The differential expression screening criteria were: Fold Change > 1.5 or < 0.585, and P-value < 0.05. A preliminary set of differentially expressed proteins was obtained as a candidate dataset for subsequent modeling features. The differentially expressed protein set is the foundation for model construction and a necessary process for screening the final model and proteins.
[0030] S5. Feature Selection and Protein Combination Optimization S51. Using the caret package in R, call the createDataPartition function to divide the candidate dataset into training and test sets in a 7:3 ratio.
[0031] S52, Random Forest Feature Filtering The randomForest package was used to build a model, and biomarkers associated with Parkinson's disease stage and UPDRS3 score were screened based on the %IncMSE index.
[0032] S53, Support Vector Machine (SVM-RFE) Feature Filtering We use caret::train to build a support vector machine model, and combine the variable importance output by varImp to extract the protein features with the highest importance index.
[0033] S54, LASSO regression feature selection Using the glmnet package, the cv.glmnet function is used to perform cross-validation to select the optimal λ, and the coef function is used to extract the protein features corresponding to non-zero coefficients.
[0034] S55, Feature Fusion and Model Training The intersection of the screening results of S52, S53, and S54 was used to determine the combination of 20 core proteins as biomarkers, including: KITM, DKK3, PHC1, MP2K3, PERE, VGFR1, IREB2, PPP5, HEPC, H2A2B, IKKA, FBRL, FAK1, EDIL3, NCK1, DRC11, GCYA1, TTC27, CD63, LHPL2.
[0035] The descriptions of each protein are as follows: KITM protein, or thymidine kinase 2 (mitochondrial), is a thymidine kinase in mitochondria that participates in the synthesis of mitochondrial DNA precursors (thymidine nucleotides) and mitochondrial nucleotide metabolism. It is widely used as a target for antiviral and chemotherapeutic drugs.
[0036] DKK3 protein, or Dickkopf-associated protein 3, is a member of the Dickkopf family. It consists of 350 amino acids and has a molecular weight of approximately 38,390 Da. Its encoding gene is located on chromosome 11p15.3, and the secreted protein contains two cysteine-rich regions.
[0037] PHC1, or polyhomeotic homolog 1, is a protein closely related to chromatin regulation and cell cycle control. Gene location: The PHC1 gene is located on human chromosome 12p13.31, and the protein it encodes contains multiple domains, some of which are involved in protein-protein interactions.
[0038] The protein MP2K3, also known as MAP2K3, stands for dual specificity mitogen-activated protein kinase kinase 3, and is a member of the mitogen-activated protein kinase kinase (MAPKK) family. The human MAP2K3 protein consists of 338 amino acids and has a molecular weight of approximately 38.3 kDa. Its encoding gene contains multiple alternative splicing transcripts that can encode different isoforms.
[0039] PERE, also known as EPX, stands for eosinophil peroxidase. As a peroxidase released by eosinophils, it participates in host defense and inflammatory responses. It can mediate the tyrosine nitration of specific granule proteins in mature, resting eosinophils.
[0040] VGFR1 should be read as "VEGFR1," which stands for Vascular Endothelial Growth Factor Receptor 1. It is an important receptor of the vascular endothelial growth factor (VEGF) family, playing a crucial role in angiogenesis and hematopoietic regulation. Gene and Structure: VEGFR1 is encoded by the gene FLT1 (fms-like tyrosine kinase 1), located on human chromosome 13q12, and belongs to the receptor tyrosine kinase (RTK) family.
[0041] IREB2 (Iron-Responsive Element Binding Protein 2) is a protein closely related to the regulation of cellular iron metabolism and plays an important role in maintaining intracellular iron homeostasis. Encoded by the IREB2 gene, located on human chromosome 15q25.1, IREB2's protein structure contains multiple functional domains, with an RNA-binding domain at its core that specifically recognizes and binds to iron response elements (IREs) on mRNA—a conserved stem-loop structure.
[0042] Protein PPP5 (Phosphine 5) is a serine / threonine protein phosphatase belonging to the PPP (Phosphine 5) family. It plays an important regulatory role in cellular stress responses, signal transduction, and protein folding. Gene and structure: PPP5 is encoded by the PPP5C gene, located on human chromosome 19p13.3.
[0043] HEPC, or hepcidin, is an antimicrobial polypeptide rich in cysteine synthesized and secreted by the liver.
[0044] H2A2B is histone H2A type 2-B, a core component of the nucleosome. Histones play a central role in transcriptional regulation, DNA repair, DNA replication, and chromosome stability.
[0045] IKKA, or protein IKKα (IκB kinase α), is a key regulator in the nuclear factor κB (NF-κB) signaling pathway and is involved in a variety of cellular physiological and pathological processes.
[0046] FBRL protein is nucleolar fibroin, a nucleolar-localized rRNA 2′-O-methyltransferase fibrillarin that participates in the methylation and early processing of precursor ribosomal RNA, the assembly of ribosomal precursors, and the maintenance of nucleolar structure.
[0047] FAK1 (Focal Adhesion Kinase 1) is a non-receptor tyrosine kinase that plays a core regulatory role in cell-extracellular matrix (ECM) interactions, cell migration, proliferation, and survival.
[0048] EDIL3 (Epidermal Growth Factor-like and Latent Transforming Growth Factor β Binding Protein-like Domain-containing Protein 3) is a secreted glycoprotein that plays an important role in extracellular matrix (ECM) remodeling, angiogenesis, wound repair, and inflammation regulation.
[0049] NCK1 (Non-Catalytic Region of Tyrosine Kinase 1) is an important adaptor protein that acts as a "molecular bridge" in cell signal transduction by mediating protein-protein interactions, and participates in the regulation of various biological processes such as cell proliferation, migration, and differentiation.
[0050] DRC11 is a member of the dynein regulatory complex (DRC). The DRC is an important complex in cilia and flagella that is associated with axonemes and is mainly involved in regulating the activity of dynein, thereby affecting the motor function of cilia / flagellates (such as the frequency and direction of wagging).
[0051] GCYA1, also known as GUCY1A1, is a protein whose full name is soluble guanylate cyclase α1 subunit. It forms a NO-sensitive soluble guanylate cyclase with the β subunit, catalyzing the conversion of guanosine triphosphate (GTP) to cyclic guanosine monophosphate (cGMP), which is involved in vasodilation and cell signaling.
[0052] The protein TTC27 (Tetratricopeptide Repeat Domain 27) is a protein containing tetrapeptide repeat sequences (TPRs), and its function is closely related to the structural and functional regulation of cilia.
[0053] CD63 is a transmembrane protein widely distributed in cells, belonging to the tetraspanin superfamily, and is also known as lysosome-associated membrane protein 3 (LAMP-3). LHPL2, also known as LHFPL2, is officially named LHFPLtetraspan subfamily member 2 protein. It belongs to the LHFP protein family (tetraspan / membrane protein family) and is related to reproductive tract development / reproductive function. Its specific biological functions are still being explored.
[0054] Example 2 Predictive Model Construction and Validation 1. Based on the above protein combination markers, three multi-classification models were constructed. The relative abundance of each protein was used as input, and the final output was different stages. Healthy individuals were classified as "no disease", and PD patients were classified as "early", "intermediate", and "late".
[0055] Classification models include: multi-class support vector machine (SVM), multinomial logistic regression model, and random forest model.
[0056] 2. Model performance verification 2.1 The model performance was evaluated using 5-fold cross-validation and Bootstrap resampling.
[0057] 2.2 Model Interpretation and SHAP Analysis To enhance model interpretability, the fastshap package is introduced to calculate the SHAP (Shapley Additive Explanations) values of each feature protein in the random forest model. SHAP bar charts are plotted to quantify the contribution of each protein to the model's predictions, and the results are visualized using bar charts and other methods.
[0058] The following detailed examples illustrate this. Protein combinatorial biomarker screening and prediction model construction: 1. During the sample processing stage, plasma samples were collected from 200 Parkinson's disease patients and 101 healthy controls. Samples were processed according to standard operating procedures, specifically as follows: Trypsin was used for enzymatic digestion at a mass ratio of 1:50 (enzyme:protein), and the reaction was carried out at 37°C for 12 hours. The reaction was terminated by adding 1% trifluoroacetic acid (TFA) and allowing it to react at room temperature for 5 minutes. Subsequently, desalting was performed using a C18 solid-phase extraction column, and the processed samples were stored at -80°C for analysis.
[0059] During the mass spectrometry detection phase, a Vanquish Neo liquid chromatography system was used, with a flow rate set at 1.8 μL / min and a gradient elution time of 14 min. Mass spectrometry detection employed Astral DIA (data-independent acquisition) mode, with a primary mass spectrometry scan range of m / z 380 to 980 and a secondary mass spectrometry scan range of m / z 150 to 2000, to achieve high-resolution, high-throughput quantitative analysis of proteins in plasma.
[0060] 2. Raw mass spectrometry data were used for protein identification and quantification analysis using DIA-NN software (version 1.8). During data processing, R language (version 4.2.3) was used for differential expression analysis, modeling, and screening.
[0061] Key analytical packages include: sva (3.46.0) for batch effect correction; and limma (3.54.0) for screening differentially expressed proteins between the control and treatment groups.
[0062] Protein screening: Random Forest (version 4.7.1.2), SVM (caret (7.0.1) + kernlab (0.9.33)), and Lasso regression model (glmnet (version 4.1.8)) were used for protein feature screening.
[0063] Regarding model parameters, the number of trees (ntree) was set to 1000 in the random forest model; the radial basis function (svmRadial) was used in the support vector machine, and the penalty parameter C was set to 1.0; in the LASSO regression model, the search range of λ was set to 0.001 to 1, and the optimal value was selected through cross-validation (cv.glmnet). These steps were used to evaluate and screen 20 differentially expressed proteins closely related to Parkinson's disease stages.
[0064] To demonstrate representative quantitative information for some proteins, the expression data (in terms of relative abundance) of several key proteins in some samples are listed below:
[0065] 3. Predictive model construction and model performance evaluation 3.1 Prediction Model Construction Based on the aforementioned protein biomarkers, three classification models were constructed. Using the relative abundance of each protein as input, the final output is different stages: healthy individuals are classified as "no disease," and PD patients are classified as "early," "intermediate," and "late stage."
[0066] In one possible implementation, the method for constructing the prediction model includes: 3.1a. Different expression patterns were obtained after cluster analysis of protein expression in different samples; each expression pattern has similar protein expression characteristics.
[0067] Furthermore, the clustering analysis includes t-SNE clustering analysis.
[0068] In this embodiment, six expression patterns were obtained through t-SNE clustering analysis.
[0069] 3.1b. By combining sample data with existing clinical information, protein expression patterns are mapped to various stage classifications, including no disease, early stage, mid-stage, and late stage.
[0070] For example, if Chuster2 / 5 has a HY stage of 0, then it is classified as a disease-free stage.
[0071] Understandably, the purpose of this step is to establish a mapping relationship between protein expression patterns and stages, which facilitates label classification.
[0072] 3.1c. Using the relative abundance of each protein biomarker in the sample data as input and the stage classification results as labels, train multiple classification models to obtain a prediction model.
[0073] Furthermore, the classification models include: multi-class support vector machine (SVM), multinomial logistic regression model, and random forest model.
[0074] 3.2 Model Performance Evaluation After the model was trained, performance was compared by plotting ROC curves, and the contribution of each protein to stage prediction was shown using the SHAP method. The SHAP contribution is shown in Figure 1. Figure 1 The vertical axis represents the official Uniprot database protein IDs of the 20 proteins, and the horizontal axis represents the SHAP contribution value. It can be seen that the top five proteins that contribute the most to the model among the 20 proteins are LHPL2, CD63, GCYA1, KITM, and PHC1.
[0075] The model evaluation results are shown in Figure 2. All three prediction models can predict patient staging, with the random forest model showing the best predictive performance. In Figure 2, the left horizontal axis represents the AUC value, the vertical axis represents the model name, and the right horizontal axis represents the standard deviation. Figure 2 shows the AUC values (0.73–0.89) and standard deviations (0.023–0.057) of the three models. The SVM+kernlab model has an AUC value of 0.73 and a standard deviation of 0.057; the Lasso regression model has an AUC value of 0.81 and a standard deviation of 0.024; and the random forest model has an AUC value of 0.89 and a standard deviation of 0.056, indicating that this model has the best predictive performance.
[0076] Figure 3 A model diagram for predicting the staging of Parkinson's disease patients.
[0077] Example 3 Based on the foregoing embodiments, this embodiment provides a computer program product for diagnosing the risk of a subject having Parkinson's disease, including the following steps: S1. Obtain the relative abundance of each protein biomarker in the plasma sample of the subject to be tested; the protein biomarkers include KITM, DKK3, PHC1, MP2K3, PERE, VGFR1, IREB2, PPP5, HEPC, H2A2B, IKKA, FBRL, FAK1, EDIL3, NCK1, DRC11, GCYA1, TTC27, CD63, and LHPL2; S2. Input the relative abundance of each protein biomarker into the prediction model for calculation, and output the staging and classification results; S3. Based on the staging and classification results, diagnose or predict whether the subject has Parkinson's disease or is at risk of having Parkinson's disease, as well as the degree of disease progression.
[0078] In one possible implementation, the prediction model includes a multi-class support vector machine, a logistic regression model, or a random forest model.
[0079] Preferably, the prediction model is a random forest model.
[0080] In one possible implementation, the prediction model is obtained by the method described in the foregoing embodiments.
[0081] In practice, when the output result is "no disease", it means that the probability of the subject having Parkinson's disease is low or that the subject is healthy. When the output result is a stage of "no disease", "early", "middle" or "late", it means that the subject may have Parkinson's disease. The stage can also predict the severity of the subject's condition. In this case, further testing with other medical means is required.
[0082] Example 4 This embodiment provides a detection reagent.
[0083] Based on the description of Embodiment 1 or 2 above, it can be seen that the predictive effect of the biomarkers selected in this application is good. Medical personnel can combine the protein combination biomarkers together as biomarkers to detect and diagnose the sample to be tested, so as to diagnose whether the sample to be tested has PD.
[0084] Therefore, this embodiment also provides a reagent for detecting protein combination markers in blood, which can be used in the preparation of diagnostic PD products to diagnose whether a sample to be tested has PD.
[0085] Example 5 This embodiment also provides a kit that may contain the detection reagent described in Example 1 to diagnose whether a sample to be tested has PD. The limitations and technical solutions of this application regarding the kit can be found in the description of Example 1 above, and will not be repeated here. Similarly, the above kit can also be used in the preparation of PD detection products, and will not be repeated here.
[0086] Example 6 This application also provides a product for diagnosing PD, to diagnose whether a sample to be tested has PD; the product is specific for one or more of the 20 PD-related proteins discovered in this application, and the product includes primers, probes, antibodies, aptamers or chips.
[0087] As can be seen from the description of Example 1 and from conventional methods in the art, when using newly discovered protein combination markers as markers for blood samples, it should be feasible for those skilled in the art to produce corresponding, specific products (primers, probes, antibodies, aptamers, or chips, etc.), which will not be elaborated here.
[0088] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions, and improvements made by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the invention.
Claims
1. A protein combinatorial marker, characterized in that, include: KITM, DKK3, PHC1, MP2K3, PERE, VGFR1, IREB2, PPP5, HEPC, H2A2B, IKKA, FBRL, FAK1, EDIL3, NCK1, DRC11, GCYA1, TTC27, CD63 and LHPL2.
2. The application of a reagent for detecting a protein composite biomarker in the preparation of products for diagnosing Parkinson's disease, wherein the protein composite biomarker comprises: KITM, DKK3, PHC1, MP2K3, PERE, VGFR1, IREB2, PPP5, HEPC, H2A2B, IKKA, FBRL, FAK1, EDIL3, NCK1, DRC11, GCYA1, TTC27, CD63 and LHPL2.
3. A reagent kit, characterized in that, The kit includes detection reagents for detecting the protein combination markers of claim 1.
4. The use of the kit according to claim 3 in the preparation of products for detecting Parkinson's disease.
5. A product for detecting Parkinson's disease, characterized in that, The product includes primers, probes, antibodies, aptamers, or chips that are specific to the protein combinatorial markers of claim 1.
6. A computer program product related to Parkinson's disease, characterized in that, The computer program product is used to diagnose the risk of a subject having Parkinson's disease, including the following steps: The relative abundance of each protein biomarker in the plasma sample of the subject to be tested is obtained; the protein biomarkers include KITM, DKK3, PHC1, MP2K3, PERE, VGFR1, IREB2, PPP5, HEPC, H2A2B, IKKA, FBRL, FAK1, EDIL3, NCK1, DRC11, GCYA1, TTC27, CD63, and LHPL2; The relative abundance of each protein biomarker is input into the prediction model for calculation, and the stage classification results are output. Based on the staging and classification results, the diagnosis or prediction can be made regarding whether the subject has Parkinson's disease, is at risk of having Parkinson's disease, and the extent of disease progression.
7. The computer program product according to claim 6, characterized in that, The prediction model includes multi-class support vector machine, logistic regression model or random forest model.
8. The computer program product according to claim 6, characterized in that, The method for obtaining the prediction model includes: Obtain blood samples from Parkinson's disease patients and healthy individuals; The protein concentration in blood samples was detected to obtain a protein dataset; Differential protein screening based on protein datasets yields candidate datasets; The first machine learning model is used to filter features in the candidate dataset to identify protein biomarkers for a single model. The intersection of the protein biomarkers selected by the first machine learning model is used to obtain the protein combination biomarkers. The relative abundance of each protein biomarker in the candidate dataset is used as input, and the stage classification results are used as labels to train the second machine learning model. The model performance is then evaluated to obtain the prediction model.
9. The computer program product according to claim 8, characterized in that, The first machine learning model includes random forest, support vector machine and LASSO regression.
10. The computer program product according to claim 8, characterized in that, The second machine learning model includes multi-class support vector machine, logistic regression model or random forest model.