Gene composition for screening or diagnosing Parkinson's syndrome and application thereof
By combining brain tissue, cerebrospinal fluid, and blood gene compositions with blood DNA methylation site compositions, and using a maximum logic intelligent classifier, the problem of accurate diagnosis and classification of Parkinson's syndrome has been solved, achieving highly accurate diagnosis and personalized treatment, and promoting new drug development and disease management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 张正军
- Filing Date
- 2026-01-04
- Publication Date
- 2026-04-21
AI Technical Summary
The current technology lacks effective biomarkers for the accurate diagnosis of Parkinson's syndrome, and single biomarkers are inconsistent in different populations and sample types, leading to limitations in diagnosis and treatment.
By using a combination of gene samples from brain tissue, cerebrospinal fluid, and blood, combined with a combination of blood DNA methylation sites, and through analysis using a maximum logic intelligent classifier, a multi-sample gene biomarker database is constructed to achieve accurate diagnosis and classification of Parkinson's syndrome.
It improves the accuracy and specificity of diagnosis, reduces the rate of misdiagnosis and missed diagnosis, supports the whole-cycle management of diseases, promotes personalized treatment and new drug development, and reduces the cost of clinical diagnosis.
Smart Images

Figure CN121896339A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of biotechnology, specifically relating to a gene composition for screening or diagnosing Parkinson's syndrome and its application. Background Technology
[0002] Parkinson's disease (PD) is a progressive neurodegenerative disorder characterized by motor and nonmotor symptoms, affecting millions worldwide. Despite extensive research, the exact molecular mechanisms remain unclear, and there are currently no disease-modifying treatments. The identification of reliable biomarkers has been a long-standing challenge, with many studies focusing on known candidate genes such as SNCA, PRKN, LRRK2, FAM171A2, PINK1, GBA, and α-synuclein levels in cerebrospinal fluid (CSF). While these biomarkers have contributed to understanding Parkinson's disease, they have not led to practical diagnostic or therapeutic breakthroughs. The heterogeneity of Parkinson's disease and limitations in study design and validation across different populations hinder the application of biomarkers.
[0003] The most extensively studied biomarkers in brain tissue and cerebrospinal fluid (CSF) include α-synuclein, neurofilament light chains (NfL), and uric acid. α-synuclein, a key component of Lewy bodies, plays a central role in Parkinson's disease pathology, making it a primary target for early detection and analysis. The development of the highly validated α-synuclein seed amplification assay (SAA) exemplifies this. Similarly, neurofilament light chains have been explored as markers of neurodegeneration, while uric acid has been investigated for its potential neuroprotective properties. Meanwhile, considerable effort has been devoted to identifying Parkinson's disease biomarkers in peripheral blood, aiming to develop less invasive and more readily available diagnostic tools. Despite these efforts, many proposed biomarkers lack robustness in clinical diagnosis and prognosis due to their inability to be replicated in independent cohorts. Summary of the Invention
[0004] The purpose of this invention is to provide a gene composition for screening or diagnosing Parkinson's syndrome, which overcomes tissue heterogeneity and instability, and provides a new means for the accurate diagnosis and targeted treatment of Parkinson's syndrome.
[0005] To achieve the above objectives, the present invention provides the following solution: In a first aspect, the present invention provides a biological composition for screening or diagnosing Parkinson's syndrome, comprising at least one of the following compositions: a brain tissue gene composition, a cerebrospinal fluid gene composition, a blood gene composition, and a blood DNA methylation site composition; The brain tissue gene composition includes genes CLSTN1, EIF4G3, NEBL, TSPOAP1, PRKRA, CEP15, and PBLD; The cerebrospinal fluid gene composition includes genes PRL, TNFRSF18, RAB22A, ROBO3, HLA-DQA2, CEL, ARPP21, and DCTN2; The blood gene composition includes genes PRL, TNFRSF18, RAB22A, ROBO3, HLA-DQA2, CEL, ARPP21, DCTN2, FAM171A2, and SNCA; The blood DNA methylation site composition includes sites cg13114166, cg09304357, cg04358214, cg17173442, cg04195527, cg15037004, cg15824323, and cg19226593.
[0006] Secondly, the present invention provides a reagent combination for screening or diagnosing Parkinson's syndrome that requires the minimum number of genes and has the lowest cost without requiring whole-genome sequencing, including reagents for detecting gene expression levels or methylation levels of the biological composition.
[0007] Thirdly, the present invention provides a system for screening or diagnosing Parkinson's syndrome, comprising the following connected functional modules: The data acquisition module is used to acquire at least one of the following data from the user sample: genotype, expression level, or methylation level of the biological composition described in the first aspect; The data analysis module is used to analyze the data acquired by the data acquisition module to obtain at least one of the following results: risk of Parkinson's syndrome, diagnostic results, and identification of disease subtypes; The data output module is used to output at least one of the results of the data analysis module, namely, the risk of Parkinson's syndrome patients, the diagnosis results, and the identification of disease subtypes, to the display terminal.
[0008] Optionally, in analyzing the data acquired by the data acquisition module to obtain at least one of the following results: risk of Parkinson's syndrome, diagnostic result, and identification of disease subtype, the data analysis module is configured to: input the data acquired by the data acquisition module into a risk assessment model based on a maximum logic intelligent classifier, and analyze to obtain at least one of the following results: risk of Parkinson's syndrome, diagnostic result, and identification of disease subtype.
[0009] Optionally, in obtaining the risk of developing Parkinson's syndrome, the data analysis module is used for: The risk assessment model based on the maximum logic intelligent classifier includes a gene subset optimization module and a risk assessment module; The gene subset optimization module is used to optimize and solve the data acquired by the data acquisition module using a maximum logic intelligent classifier to obtain a target gene subset for characterizing the risk of Parkinson's syndrome; the target gene subset corresponds to a combination of several genes of any of the biological compositions described in the first aspect; The risk assessment module is used to calculate the risk of Parkinson's syndrome for a user sample based on a subset of target genes.
[0010] Optionally, the maximum logical intelligent classifier is: ; ; ; in, To adopt the first The first biological sampling method obtained The first user sample The expression vector corresponding to the gene group or the expression vector corresponding to the methylated CpG site Value vector; They are respectively using the first The first biological sampling method obtained The first user sample The first in the group 2. The expression value of a gene or the corresponding methylated CpG site value; The total number of groups of genes or methylated CpG sites; The number of biological sampling methods; For the first Risk of Parkinson's syndrome in a sample of users; This indicates taking the maximum value; The intercept; Let be a coefficient vector, used to represent the first... The contribution of each gene or methylated CpG site in the group to the risk of Parkinson's syndrome; This represents the optimized value of the coefficient vector; For the target gene subset; The number of combinations of genes or methylated CpG sites in the target gene subset; for The set of indices; and For penalty parameters; yes{ The union of}; for The number of elements in the middle; The number of samples; It is an indicator function; To set a threshold; Indicates the first One user sample was identified as having Parkinson's syndrome; Indicates the first None of the user samples had Parkinson's syndrome.
[0011] Optionally, the data analysis module includes a Parkinson's syndrome screening unit, which is used to determine the Parkinson's syndrome diagnosis result of the user sample based on the user sample's risk of developing Parkinson's syndrome and a set threshold.
[0012] Optionally, the Parkinson's syndrome subtype is determined by the number of combinations in the target gene subset.
[0013] Optionally, the threshold can be set to 0.5.
[0014] According to specific embodiments provided by the present invention, the present invention has the following technical effects: This invention provides a gene composition for screening or diagnosing Parkinson's syndrome. The invention develops corresponding gene compositions for different types of samples—brain tissue, cerebrospinal fluid, and blood—which improves diagnostic accuracy and specificity, and reduces misdiagnosis and missed diagnosis rates. Since different types of samples carry different molecular information at different stages of disease development, analyzing gene biomarkers from multiple sample types allows for individual and synergistic verification of local lesion characteristics, systemic molecular signals, and system-specific indicators, significantly improving the correlation between gene biomarkers and the disease. The gene composition effectively avoids false positive / false negative problems caused by individual differences, sample contamination, or local microenvironment interference, significantly improving diagnostic accuracy and specificity and providing a reliable basis for precise disease identification. Furthermore, the gene composition of this invention can expand the applicable diagnostic scenarios, covering the entire disease lifecycle. Because the accessibility of diagnostic samples varies significantly at different stages of Parkinson's syndrome and among different patient groups, summarizing gene biomarkers from multiple sample types allows for the construction of diversified diagnostic solutions for different scenarios, addressing clinical pain points that cannot be met by a single sample type, and achieving full-cycle diagnosis and management coverage of the disease. Furthermore, the gene composition facilitates precise subtyping and personalized treatment of Parkinson's syndrome, improving clinical treatment outcomes. Clinicians can develop targeted, individualized treatment plans (such as targeted therapy and immunotherapy) based on the patient's gene marker type, avoiding ineffective treatment or adverse reactions caused by the traditional "one-size-fits-all" treatment model. Finally, the gene composition can also drive innovation in diagnostic technology, reducing clinical diagnostic costs and burdens. In addition, the gene composition provides core support for disease etiology research and new drug development, promoting progress in medical research. Abnormal gene expression is an important carrier for revealing the etiology and pathogenesis of diseases. By integrating gene marker information from different samples, key molecular pathways and regulatory networks in the occurrence and development of diseases can be systematically identified, clarifying the core pathogenic genes and driving factors, providing direct molecular evidence for disease etiology research. The gene markers can serve as targets for new drug development, providing precise directions for the development of targeted drugs and immunotherapies, shortening the new drug development cycle, and improving the success rate of research and development. Furthermore, the construction of a multi-sample gene biomarker database can provide rich molecular data resources for medical research, promote multi-center, large-sample clinical research, facilitate interdisciplinary integration (such as medicine, biology, and big data science), and contribute to the overall progress of medical research.
[0015] This invention provides a system for screening or diagnosing Parkinson's syndrome, comprising the following connected functional modules: a data acquisition module for acquiring at least one data point from a user sample, including genotype, expression level, or methylation level of the biological composition described in the first aspect; a data analysis module for analyzing the data acquired by the data acquisition module to obtain at least one result among Parkinson's syndrome risk, diagnostic result, and disease subtype identification; and a data output module for outputting at least one result among Parkinson's syndrome patient risk, diagnostic result, and disease subtype identification from the data analysis module to a display terminal. Simultaneously, the data analysis module performs analysis based on a maximum logical intelligent classifier (S4 classifier). The maximum logical intelligent classifier (S4 classifier) used in the system can achieve an overall accuracy of up to 91.63% with a very small gene set, and even reaches 100% in some cohorts. The coefficients in the model have clear biological interpretations, indicating the directional influence of gene expression changes on disease risk. The system can also reveal Parkinson's syndrome subtypes and guide precision medicine; the classifier of this invention can classify Parkinson's patients into different molecular subtypes. This demonstrates the high heterogeneity of Parkinson's syndrome and implies that a "one-size-fits-all" treatment strategy is inefficient, laying a solid foundation for developing personalized therapies targeting specific subtypes. Furthermore, the system does not rely on p-values and multiple test corrections; instead, it builds its model by finding the minimum, optimal gene combination that perfectly distinguishes cases from controls, fundamentally avoiding overfitting and insufficient statistical power. Attached Figure Description
[0016] Figure 1 Paired scatter plots of FAM171A2, SNCA, DCTN2, and CLSTN1 based on brain tissue sample (GSE40396); Figure 2 Paired scatter plots of FAM171A2, SNCA, DCTN2, and CLSTN1 based on whole blood sample (GSE99039); Figure 3 Paired scatter plots of FAM171A2, SNCA, DCTN2, and CLSTN1 based on cerebrospinal fluid samples (PPMI); Figure 4 Venn diagram of Parkinson's disease subtypes in the GSE20295 dataset; Figure 5 A diagram of a sparse, directed gene regulatory network based on GSE20295 (prefrontal cortex); Figure 6 A sparse, directed gene regulatory network diagram based on cerebrospinal fluid and whole blood data; Figure 7 A clinical diagnostic view based on tissue data; Figure 8A clinical diagnostic view based on blood DNA methylation data; Figure 9 A clinical diagnostic view based on cerebrospinal fluid data; Figure 10 This is a schematic diagram of the functional modules of a system used for screening or diagnosing Parkinson's syndrome. Detailed Implementation
[0017] This invention provides a biological composition for screening or diagnosing Parkinson's syndrome, comprising at least one of the following compositions: a brain tissue gene composition, a cerebrospinal fluid gene composition, a blood gene composition, and a blood DNA methylation site composition; wherein the brain tissue gene composition comprises genes CLSTN1, EIF4G3, NEBL, TSPOAP1, PRKRA, CEP15, and PBLD; and the cerebrospinal fluid gene composition comprises genes PRL, TNFRSF18, RAB22A, ROBO3, HLA-DQA2, CEL, and ARP. P21 and DCTN2; the blood gene composition includes genes PRL, TNFRSF18, RAB22A, ROBO3, HLA-DQA2, CEL, ARPP21, DCTN2, FAM171A2 and SNCA; the blood DNA methylation site composition includes cg13114166, cg09304357, cg04358214, cg17173442, cg04195527, cg15037004, cg15824323 and cg19226593.
[0018] In this invention, the brain tissue gene composition, cerebrospinal fluid gene composition, and blood gene composition are used to screen for or diagnose Parkinson's syndrome based on the measured gene expression level or genotype, and the blood DNA methylation site composition is used to screen for or diagnose Parkinson's syndrome based on the DNA methylation level at specific sites in the blood.
[0019] This invention provides a reagent for screening or diagnosing Parkinson's syndrome that requires the minimum number of genes and the lowest cost, without the need for whole-genome sequencing. The reagent includes a combination of reagents that detect at least one of the transcriptome or methylation sites, expression levels, or methylation levels corresponding to the genotype of the biological composition.
[0020] In this invention, the reagent preferably includes primers for detecting at least one of the brain tissue gene composition, cerebrospinal fluid gene composition, and blood gene composition, or a combination of primers and qPCR amplification premix, or a combination of primers, qPCR amplification premix, and reverse transcription reagent; or the reagent preferably includes primers for detecting the genotype of at least one of the brain tissue gene composition, cerebrospinal fluid gene composition, and blood gene composition, or a combination of primers and PCR amplification premix, or a combination of primers, PCR amplification premix, and sequencing reagent.
[0021] In this embodiment of the invention, reagents for detecting the genotype, expression level, or methylation level of the biological composition can be found on publicly available data links from the U.S. National Institutes of Health (NIH). For example, the data source for the GSE20295 cohort is located at https: / / www.ncbi.nlm.nih.gov / geo / query / acc.cgi?acc=GSE20295. Other cohorts can be adapted by replacing the cohort number with the above URL.
[0022] like Figure 10 As shown, the present invention provides a system for screening or diagnosing Parkinson's syndrome, comprising a data acquisition module, a data analysis module, and a data output module connected as follows.
[0023] The data acquisition module is used to acquire at least one of the following data from user samples: genotype, expression level, or methylation level of the biological composition. The biological composition is at least one of the following: brain tissue gene composition, cerebrospinal fluid gene composition, blood gene composition, and blood DNA methylation site composition.
[0024] First, a biological sample (such as blood, cerebrospinal fluid, or brain tissue) is obtained from the user, and the genotype, expression level, or methylation level of the biological composition is detected. The generated detection data is the input to the data acquisition module.
[0025] The data analysis module is used to analyze the data acquired by the data acquisition module to obtain at least one of the following results: risk of Parkinson's syndrome, diagnostic results, and identification of disease subtypes.
[0026] The data output module is used to output at least one of the results of the data analysis module, namely, the risk of Parkinson's syndrome patients, the diagnosis results, and the identification of disease subtypes, to the display terminal.
[0027] In terms of analyzing the data acquired by the data acquisition module to obtain at least one of the following results: risk of Parkinson's syndrome, diagnostic result, and identification of disease subtype, the data analysis module is configured to: input the data acquired by the data acquisition module into a risk assessment model based on a max-logistic intelligence (S4) classifier, and analyze to obtain at least one of the following results: risk of Parkinson's syndrome, diagnostic result, and identification of disease subtype.
[0028] In a specific example, regarding the determination of the risk of developing Parkinson's syndrome, the data analysis module includes a risk assessment model based on a maximum logic intelligent classifier, comprising a gene subset optimization module and a risk assessment module.
[0029] The gene subset optimization module is used to optimize and solve the data acquired by the data acquisition module using a maximum logic intelligent classifier to obtain a target gene subset for characterizing the risk of Parkinson's syndrome; the target gene subset corresponds to a combination of several genes of the above-mentioned biological composition. The risk assessment module is used to calculate the risk of Parkinson's syndrome for a user sample based on a subset of target genes.
[0030] The maximum logic intelligent classifier is a deterministic mathematical model whose core lies in finding a subset of target genes that can perfectly distinguish between disease and health states through optimization algorithms. (Minimum subset of genes, different subsets for different data sources) and minimum number of classifier signatures .
[0031] Assumption It is the first Parkinson's syndrome status of individual users ( It indicates good health. (Indicates that a diagnosis has been made) It is the expression value corresponding to the gene or the methylated CpG site. Value. Typically, there are over 20,000 genes and... CpG site, here represent The first of three different biological sampling methods The expression value or the corresponding methylated CpG site Value. Set based on user experience. Therefore, superscript It can be removed. (This is for simplifying the symbols.) Using the logit join function (or any monotonic join function), you can perform a join on the first... Parkinson's syndrome risk (probability of disease) for a single user sample. Calculate using the following relationship: (1); in, Equivalent to the risk of developing Parkinson's syndrome ; It is the intercept term; For the first The size of each user sample is The observation vector; It is the size of The coefficient vector is used to characterize the relationship between the predictor variable (gene or CpG site) and the risk of disease.
[0032] The above formula (1) can be expressed as: (2).
[0033] The maximum logical intelligent classifier is constructed by finding a subset of target genes that can distinguish between Parkinson's syndrome and healthy controls with maximum accuracy. The maximum logical intelligent classifier is represented as follows: Given the complexity of Parkinson's syndrome and its multiple symptoms (subtypes), it is natural to assume that the epigenetic structure of all subtypes may differ. It is hypothesized that all subtypes may be associated with G group genes or CpG loci: (3); in, To adopt the first The first biological sampling method obtained The first user sample The expression vector corresponding to the gene group or the expression vector corresponding to the methylated CpG site A value vector is a The observation vector; They are respectively using the first The first biological sampling method obtained The first user sample The first in the group 2. The expression value corresponding to each gene or the methylated CpG site. value; For the first The number of genes or methylated CpG sites in the group; The total number of groups of genes or methylated CpG sites; This represents the number of biological sampling methods.
[0034] The competitive (risk) factor classifier is defined as follows: (4); in, For the first Risk of Parkinson's syndrome in a sample of users; This indicates taking the maximum value; The intercept; For the first The coefficient vector of the genotype or methylated CpG site is used to characterize the first gene. The contribution of each gene or methylated CpG site in the group to the risk of Parkinson's syndrome is a The coefficient vector.
[0035] In formula (4), the user's risk of developing Parkinson's syndrome Mainly composed of the largest item The decision is made by comparing the values of all factors; the winner is selected, and the genome with the highest value is adopted. The calculated probability value is taken as the user's risk of developing Parkinson's syndrome.
[0036] when When formula (4) is simplified to classical logistic regression, classical logistic regression is a special case of the new classifier. Compared with black-box machine learning methods such as random forest and deep learning, the model in formula (4) provides clear and interpretable features, selects genes and CpG sites, and builds a bridge between linear models and advanced machine learning. It retains key characteristics such as interpretability, computability, predictability and stability.
[0037] Extend the S4 classifier from K=1 to any K: (5); in, The optimized value of the coefficient vector is the optimized value of gene expression or methylation CpG sites in the target gene subset. For the target gene subset; The number of combinations of genes or methylated CpG sites in the target gene subset; and It is the set of target genes or CpG sites selected by the final classifier and the number of its final discriminators (similar to handwriting). The discriminator is the classifier. ={1, 2, ..., p} is the set of indices for all genes or CpG sites; the value of p is based on the number of genes or sites. For formula (3) The set of indices; and For penalty parameters; yes{ The union of}; for The number of elements in the middle; The number of samples; It is an indicator function; To set a threshold; Indicates the first One user sample was identified as having Parkinson's syndrome; Indicates the first None of the user samples had Parkinson's syndrome.
[0038] In practical applications, a threshold probability needs to be selected to determine the patient's category label. Based on user experience, a threshold of 0.5 is set. Therefore, if , Individuals in the sample were classified as disease-free; otherwise, they were classified as having Parkinson's syndrome.
[0039] When the S4 classifier reaches 100% accuracy, it establishes a bioequivalence and unique RNA-seq, DNA methylation geometric space, which is referred to as maximum logical intelligence.
[0040] It should be noted that adjusting the threshold in formula (5) between 0 and 1 will affect the intercept in formula (4), but will not affect the coefficient vector. Therefore, it will not change the clustering and classification results, which makes calculating AUC unnecessary.
[0041] Based on a risk assessment model using a maximum logic intelligent classifier, the system determines whether a user has Parkinson's syndrome and its specific subtype. The data analysis module includes a Parkinson's syndrome screening unit and a Parkinson's syndrome subtype determination unit.
[0042] The Parkinson's syndrome screening unit is used to determine the diagnosis of Parkinson's syndrome in a user sample based on the user sample's risk of developing Parkinson's syndrome and a set threshold.
[0043] The Parkinson's syndrome subtype determination unit is used to determine the user's Parkinson's syndrome subtype, wherein the Parkinson's syndrome subtype is determined by the number of combinations in the target gene subset.
[0044] The following detailed description, in conjunction with embodiments, illustrates a gene composition for screening or diagnosing Parkinson's syndrome provided by the present invention and its application therein; however, these descriptions should not be construed as limiting the scope of protection of the present invention.
[0045] Example 1 Methods for screening or diagnosing Parkinson's syndrome based on the expression values of 7 genes in brain tissue transcriptomes 1. Data source description: Five independent brain tissue cohorts (GSE20295, GSE7621, GSE20141, GSE8397 and GSE49036) from the U.S. National Institutes of Health Clinical Trials Open GEO Database were used for training and validation.
[0046] Table 1. Explanation of the source of brain tissue cohort data
[0047] For each target brain tissue cohort, data from the remaining brain tissue cohorts are used as the sample set for that target brain tissue cohort. For example, for the GSE20295 brain tissue cohort, data from the GSE7621, GSE20141, GSE8397, and GSE49036 brain tissue cohorts are used as the validation sample set for the GSE20295 brain tissue cohort. The sample set includes detection data from several user samples and the Parkinson's syndrome disease label corresponding to each user sample. The detection data are the expression level data of the biological composition of the user samples. The Parkinson's syndrome disease label is either diseased or healthy.
[0048] 2. Optimization Analysis of the Maximum Logical Intelligent Classifier (S4 Classifier) The quenching algorithm is used for optimization, and multiple competitive classifiers are obtained for each data. In the optimization analysis of this embodiment, three competitive classifiers (CF1, CF2, CF3) can be set, that is, the total number of groups G in formula (4) is 3.
[0049] The sample set is divided into a training set and a validation set. Using the detection data of user samples in the training set as input and the Parkinson's syndrome disease label corresponding to the user samples as output, the risk assessment model based on the maximum logic intelligent classifier is trained to obtain the trained risk assessment model corresponding to each target brain tissue cohort.
[0050] The trained risk assessment model was validated using a validation set, and the model that met the validation criteria was used as the final risk assessment model. The final risk assessment model was then used to analyze the detection data of user samples in the target brain tissue cohort to obtain the risk of Parkinson's syndrome. This process was described above and will not be repeated here.
[0051] Taking the GSE20295 queue as an example, its classifier formula is: CF1: -8.0203 + 1.2978×CLSTN1 + 8.5063× EIF4G3 - 9.5181×CEP15; CF2: 3.7333 + 0.5648×NEBL + 6.2586× PRKRA - 7.0323×CEP15; CF3: -7.1155 - 3.1464× NEBL + 8.6751× TSPOAP1 - 6.7211× PBLD; The final classification result is determined by CFmax = max(CF1, CF2, CF3), where CFmax is the value in formula (4). .
[0052] Taking the GSE20295 queue as an example, based on the classification of users by CF1, CF2, and CF3 classifiers, they can be divided into at least 7 different molecular subtypes. The Venn diagram of Parkinson's disease subtypes in the GSE20295 dataset is shown below. Figure 4 As shown, the overlapping part of the circles representing different classifiers indicates that the corresponding user sample is assessed as having Parkinson's syndrome based on both classifiers. For example, the number 6 in the overlapping part of the CF1 circle and the CF2 circle indicates that the CF1 genome and the CF2 genome of these 6 user samples can both be used to assess the user as having Parkinson's syndrome.
[0053] 3. Diagnostic Results The diagnostic results are shown in Table 2. The overall classification accuracy of the 7 gene combinations in the five brain tissue cohorts reached 91.63%, with a sensitivity of 97.27% and a specificity of 84.95%, and the accuracy of some cohorts reached 100%.
[0054] Table 2. Results of classifying PD and non-PD using independent classifiers and their joint S4 classifier.
[0055] Overall results: composite indicators of all participants in the clinical trials (5 cohorts).
[0056] Example 2 Screening and diagnosis of PD based on 8 gene combinations from cerebrospinal fluid 1. Data source: Cerebrospinal fluid samples from the PPMI cohort.
[0057] 2. Cerebrospinal fluid 8-gene combination: PRL, TNFRSF18, RAB22A, ROBO3, HLA-DQA2, CEL, ARPP21, DCTN2, and the expression level of each gene was detected.
[0058] 3. The S4 classifier was used to build a model to distinguish between PD and the control group. The specific method is the same as described above.
[0059] Table 3. Performance of individual classifiers and combined maximum competition classifiers for classifying Parkinson's disease patients and non-Parkinson's disease patients into corresponding groups using cerebrospinal fluid sample data.
[0060] As shown in Table 3, the accuracy, sensitivity, and specificity of the 8-gene set in cerebrospinal fluid samples in distinguishing between PD and non-PD were 0.7721, 0.7912, and 0.7000, respectively.
[0061] Example 3 1. Data source: Whole blood samples from the GSE99039 cohort.
[0062] 2. Blood 10-gene set: PRL, TNFRSF18, RAB22A, ROBO3, HLA-DQA2, CEL, ARPP21, FAM171A2, SNCA and DCTN2, and the expression level of each gene was detected.
[0063] 3. The S4 classifier was used to build a model to distinguish between PD and the control group. The specific method is the same as described above.
[0064] Table 4. Performance of individual classifiers and combined maximum competition classifiers for classifying Parkinson's disease patients and non-Parkinson's disease patients into corresponding groups using whole blood RNA sample data GSE99039.
[0065] As shown in Table 4, the accuracy of the 10-gene set in whole blood samples in distinguishing between PD and non-PD was 0.8219, the sensitivity was 0.7951, and the specificity was 0.8455.
[0066] Example 4 Screening based on blood DNA methylation 1. Data source: GSE72774 methylation dataset.
[0067] 2. Sites and Genes: The methylation levels of eight CpG sites were detected: cg13114166 (RPA2), cg09304357, cg04358214 (C16orf70), cg17173442 (RFXANK), cg04195527 (INSIG2), cg15037004 (ZNF366), cg15824323, and cg19226593 (DCTN2).
[0068] 3. Analytical Methods The S4 classifier was used to build a model to distinguish between PD and the control group, and the specific method was the same as described above.
[0069] 4. Results DCTN2 Hypomethylation at the cg19226593 site on the gene was associated with a reduced risk of PD. Using a classifier that included these sites, a classification accuracy of 79.35% was achieved in blood samples (see Table 5).
[0070] Table 5 shows the performance results of Parkinson's disease classification using an independent classifier.
[0071] Example 5 1. Analysis of key differentially expressed genes in different sample types A comprehensive reanalysis using the maximum logical intelligence method described in Examples 1-4 revealed several key differences that may influence the conclusions of the literature. Specifically, functional shifts in multiple genes were observed between blood and cerebrospinal fluid, with only a few genes (such as ROBO3 and DCTN2) maintaining stable expression in different regions. These differences are summarized below: In cerebrospinal fluid, higher expression of PRL, TNFRSF18, and ARPP21 is associated with a non-Parkinson's disease (healthy control-like) state; in blood, lower expression is associated with a healthy control-like state. These genes may have opposing regulatory roles in the systemic circulation and the central nervous system, meaning that targeting these genes may require different strategies in blood and cerebrospinal fluid.
[0072] HLA-DQA2 and CEL are expressed at higher levels in blood, similar to healthy controls; and at lower levels in cerebrospinal fluid, also similar to healthy controls. These may be immune genes associated with the blood-brain barrier (BBB), playing different roles in the peripheral and central nervous systems.
[0073] High expression of ROBO3 in blood and cerebrospinal fluid is beneficial, as it can act as a stable protective factor and may be involved in maintaining neuronal connectivity or the integrity of the blood-brain barrier.
[0074] DCTN2 was expressed at low levels in both blood and cerebrospinal fluid, similar to healthy controls. Its consistent expression across different regions suggests it may serve as a stable regulator of the blood-cerebrospinal fluid interface. This makes it a high-priority candidate gene for further research into Parkinson's disease progression and treatment.
[0075] RAB22A, FAM171A2, and SNCA have no effect in cerebrospinal fluid, but their expression in blood is correlated with that in healthy controls. They may be peripheral regulatory factors, which means that their role in Parkinson's disease is systemic rather than central nervous system specific.
[0076] 2. Paired scatter plots and difference indicators Paired scatter plots of FAM171A2, SNCA, DCTN2, and CLSTN1 were plotted from brain tissue sample GSE40396, whole blood sample GSE99039, and cerebrospinal fluid data, see [link to scatter plot]. Figures 1-3 It can be seen that the distribution of Parkinson's disease cases and healthy controls differs greatly and overlaps. Figure 1-3Taking the top-left panel (FAM171A2 and SNCA) as an example, the ranges for Parkinson's disease and healthy control status for each gene highly overlap. This observation strongly suggests that these genes cannot serve as individual biomarkers; precise biomarkers must be combinations of multiple genes. Therefore, the combinations in Table 2 represent the best-performing biomarkers.
[0077] Figure 2 The study showed a positive correlation between FAM171A2 and SNCA, although the correlation was low. Analysis confirmed that FAM171A2 plays a crucial role in whole blood; however, its function in cerebrospinal fluid appears to be negligible. Notably, the range of cerebrospinal fluid FAM171A2 levels presented in the *Science* paper was truncated, much narrower than that observed in this invention. Figure 3 The distribution shown in the figure. This difference suggests that reported cerebrospinal fluid levels may not fully capture the biological variability of FAM171A2. Therefore, relying on cerebrospinal fluid FAM171A2 measurements to assess Parkinson's disease may have limited utility and could even lead to misleading conclusions regarding its diagnostic or prognostic value.
[0078] 3. Network analysis of whole blood and cerebrospinal fluid samples Network analysis of gene interactions in whole blood and cerebrospinal fluid (CSF) samples revealed distinct connectivity patterns, indicating functional differences between these two biological regions. Network analysis was performed based on Tables 2–4 using generalized correlation metrics. Results are as follows: Figure 5 and Figure 6 As shown.
[0079] about Figure 5 In brain tissue samples from [the study], CLSTN1 is the receptor among seven new key genes in brain tissue, making it the most promising target gene. Parkinson's disease research has long focused on genes such as SNCA and PRKN, but our findings suggest that CLSTN1 plays a more central role in the neuropathology of Parkinson's disease. SNCA (α-synuclein) is known to play a role in Lewy body formation, but its importance in brain tissue structural networks is weaker compared to CLSTN1. Unlike SNCA, which lacks importance in the Parkinson's disease brain network, CLSTN1 becomes a key hub regulating dopaminergic neuron survival, mitochondrial function, and neuroinflammation. Our research indicates that CLSTN1, previously classified as a receptor for Alzheimer's disease and amyotrophic lateral sclerosis (ALS), may be a key target receptor in Parkinson's disease, paving the way for new therapeutic strategies.
[0080] exist Figure 6In the lower left whole blood (Parkinson's disease) network, among the identified genes, ARPP21 and HLA-DQA2 showed the highest elude in whole blood, indicating that they act as major regulatory hubs, influencing multiple downstream targets. Conversely, DCTN2 showed the lowest elude in whole blood, suggesting that its direct regulatory role in this region is relatively limited.
[0081] However, in the cerebrospinal fluid (Parkinson's disease) network, the connectivity of HLA-DQA2 changed significantly, with an outflow degree of zero, suggesting that it may act as a terminal node rather than a regulatory hub in this environment. Despite losing outward connectivity, ARPP21 maintained the highest outflow degree, highlighting its crucial role in both blood and cerebrospinal fluid regions. Interestingly, DCTN2, which has the lowest outflow degree in whole blood, exhibits the second highest outflow degree in cerebrospinal fluid, indicating that its function is environment-dependent.
[0082] These findings suggest fundamental functional changes in gene interactions between whole blood and cerebrospinal fluid, highlighting the dynamic nature of molecular regulation in neurodegenerative diseases such as Parkinson's disease. The crucial role of HLA-DQA2 in both regions underscores its potential significance as a robust biomarker for Parkinson's disease. Furthermore, DCTN2 is the only gene identified by DNA methylation analysis within the cerebrospinal fluid gene network, further reinforcing its importance. Given its strong presence in both whole blood and cerebrospinal fluid, DCTN2 may serve as a molecular bridge connecting peripheral and central regions, potentially linking systemic immune responses to neuroinflammatory or neurodegenerative processes.
[0083] This dual presence suggests that DCTN2 may be a highly relevant druggable target, as therapeutic interventions aimed at modulating its function could influence disease mechanisms in both the peripheral and central nervous systems. Targeting DCTN2 could provide a novel approach to mitigating the progression of Parkinson's disease, making it an attractive candidate target for further drug development and precision medicine strategies.
[0084] Further investigation revealed the interaction between DCTN2 and genes identified in brain tissue (namely CLSTN1, EIF4G3, NEBL, TSPOAP1, PRKRA, CEP15, and PBLD). We found that DCTN2 is an upstream gene of all these key genes.
[0085] 4. Clinical View To illustrate the clinical utility of the identified genetic signatures, visual diagnostic diagrams simulating real-world clinical workflows are provided. Key genes identified in previous sections are visualized. These diagrams simulate how such biomarkers function in a clinical diagnostic workflow, supporting early screening via blood tests followed by cerebrospinal fluid confirmation or brain imaging. Figures 7-9Characteristic patterns from tissues, blood, and cerebrospinal fluid were displayed.
[0086] The clinical view plots further demonstrate the robustness of the findings. In each dataset, the classifier clearly distinguishes between the Parkinson's disease group and the control group using only a few genes. These plots simulate the role of such biomarkers in clinical diagnosis, supporting early screening via blood tests, followed by cerebrospinal fluid confirmation or brain imaging.
[0087] Furthermore, the axes identified here can be used for patient stratification or subtype classification, offering potential applications for monitoring disease progression or treatment response. Its minimal genetic footprint allows for the translation into methods based on quantitative polymerase chain reaction (qPCR) or digital assays.
[0088] The above description is merely a preferred embodiment of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention. Personalized treatment using existing repurposed drugs based on the principles of the present invention, as well as the development of new targeted drugs based on the present invention, should also be considered within the scope of protection of the present invention.
Claims
1. A biological composition for screening or diagnosing Parkinson's syndrome, comprising at least one of the following compositions: a brain tissue gene composition, a cerebrospinal fluid gene composition, a blood gene composition, and a blood DNA methylation site composition; The brain tissue gene composition includes genes CLSTN1, EIF4G3, NEBL, TSPOAP1, PRKRA, CEP15, and PBLD; The cerebrospinal fluid gene composition includes genes PRL, TNFRSF18, RAB22A, ROBO3, HLA-DQA2, CEL, ARPP21, and DCTN2; The blood gene composition includes genes PRL, TNFRSF18, RAB22A, ROBO3, HLA-DQA2, CEL, ARPP21, DCTN2, FAM171A2, and SNCA; The blood DNA methylation site composition includes cg13114166, cg09304357, cg04358214, cg17173442, cg04195527, cg15037004, cg15824323 and cg19226593.
2. A reagent for screening or diagnosing Parkinson's syndrome, characterized in that, Includes reagents for detecting at least one of the genotype, expression level, or methylation level of the biological composition.
3. A system for screening or diagnosing Parkinson's syndrome, characterized in that, Includes the following functional modules connected in sequence: The data acquisition module is used to acquire at least one of the following data from the user sample: genotype, expression level, or methylation level of the biological composition of claim 1; The data analysis module is used to analyze the data acquired by the data acquisition module to obtain at least one of the following results: risk of Parkinson's syndrome, diagnostic results, and identification of disease subtypes; The data output module is used to output at least one of the results of the data analysis module, namely, the risk of Parkinson's syndrome patients, the diagnosis results, and the identification of disease subtypes, to the display terminal.
4. The system for screening or diagnosing Parkinson's syndrome according to claim 3, characterized in that, In terms of analyzing the data acquired by the data acquisition module to obtain at least one of the following results: risk of Parkinson's syndrome, diagnostic result, and identification of disease subtype, the data analysis module is configured to: input the data acquired by the data acquisition module into a risk assessment model based on a maximum logic intelligent classifier, and analyze to obtain at least one of the following results: risk of Parkinson's syndrome, diagnostic result, and identification of disease subtype.
5. The system for screening or diagnosing Parkinson's syndrome according to claim 4, characterized in that, Regarding the determination of the risk of developing Parkinson's syndrome, the data analysis module is used for: The risk assessment model based on the maximum logic intelligent classifier includes a gene subset optimization module and a risk assessment module; The gene subset optimization module is used to optimize and solve the data acquired by the data acquisition module using a maximum logic intelligent classifier to obtain a target gene subset for characterizing the risk of Parkinson's syndrome; the target gene subset corresponds to a combination of several genes of any of the biological compositions of claim 1; The risk assessment module is used to calculate the risk of Parkinson's syndrome for a user sample based on a subset of target genes.
6. The system for screening or diagnosing Parkinson's syndrome according to claim 5, characterized in that, The maximum logical intelligent classifier is: ; ; ; in, To adopt the first The first biological sampling method obtained The first user sample The expression vector corresponding to the gene group or the expression vector corresponding to the methylated CpG site Value vector; They are respectively using the first The first biological sampling method obtained The first user sample The first in the group 2. The expression value corresponding to each gene or the methylated CpG site. value; The total number of groups of genes or methylated CpG sites; The number of biological sampling methods; For the first Risk of Parkinson's syndrome in a sample of users; This indicates taking the maximum value; The intercept; Let be a coefficient vector, used to represent the first... The contribution of each gene or methylated CpG site in the group to the risk of Parkinson's syndrome; This refers to the optimized values for gene expression or methylated CpG sites within a subset of the target genes. For the target gene subset; The number of combinations of genes or methylated CpG sites in the target gene subset; for The set of indices; and For penalty parameters; yes{ The union of}; for The number of elements in the middle; The number of samples; It is an indicator function; To set a threshold; Indicates the first One user sample was identified as having Parkinson's syndrome; Indicates the first None of the user samples had Parkinson's syndrome.
7. The system for screening or diagnosing Parkinson's syndrome according to claim 6, characterized in that, The data analysis module includes a Parkinson's syndrome screening unit; The Parkinson's syndrome screening unit is used to determine the diagnosis result of Parkinson's syndrome in a user sample based on the user sample's risk of developing Parkinson's syndrome and a set threshold.
8. The system for screening or diagnosing Parkinson's syndrome according to claim 6, characterized in that, Parkinson's syndrome subtypes are determined by the number of combinations within a subset of target genes.
9. The system for screening or diagnosing Parkinson's syndrome according to claim 7, characterized in that, Set the threshold to 0.5.