Computer device for noninvasive recognition of Parkinson's disease based on intestinal flora and computer readable storage medium
By constructing a Parkinson's disease prediction model based on the gut microbiota and using machine learning algorithms to analyze the abundance of 11 key bacterial genera in fecal samples, the accuracy of Parkinson's disease identification in existing technologies has been addressed, enabling non-invasive and low-cost early diagnosis and supporting personalized medicine.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PEKING UNIVERSITY SIXTH HOSPITAL
- Filing Date
- 2024-10-21
- Publication Date
- 2026-04-21
AI Technical Summary
Existing technologies for identifying Parkinson's disease based on gut microbiota suffer from insufficient accuracy, difficulty in integrating different research data, and a lack of non-invasive, highly sensitive, and specific early identification methods.
A computer device and model were constructed to analyze the relative abundance of 11 key bacterial genera in the fecal samples of subjects, and a Parkinson's disease prediction model was established using machine learning algorithms. The model outputs the probability value of the subject's disease and determines whether the subject has Parkinson's disease.
This method provides a non-invasive, accurate, and low-cost method for the early identification of Parkinson's disease, improving the accuracy and repeatability of identification, supporting personalized medicine, reducing diagnostic costs, and possessing efficient early screening capabilities.
Smart Images

Figure CN121905461A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of healthcare informatics and relates to a computer device and computer-readable storage medium for non-invasive identification of Parkinson's disease based on gut microbiota. Background Technology
[0002] Parkinson's disease (PD) is a common neurodegenerative disorder, and its diagnosis typically relies on clinical presentation, neurological examination, and response to dopamine replacement therapy such as levodopa. However, the accuracy of these methods depends on the physician's experience and judgment, especially in the early stages of the disease, which can lead to misdiagnosis or missed diagnosis. Furthermore, while imaging examinations and genetic testing can aid in diagnosis in some cases, their application is limited by factors such as cost, technical complexity, and test sensitivity. Therefore, there is an urgent need for a rapid, accurate, and non-invasive early identification method for the clinical diagnosis of PD.
[0003] In recent years, with the deepening research into the pathogenesis of Parkinson's disease (PD), increasing evidence suggests that the gut microbiota plays a crucial role in the occurrence and development of PD. Studies have found that the composition of the gut microbiota in PD patients differs significantly from that in healthy controls, particularly in microbial groups related to gut health and immune regulation. For example, the abundance of known butyrate-producing bacteria such as Faecalibacterium, Roseburia, and Coprococcus_2 is significantly reduced in PD patients, potentially leading to increased intestinal inflammation; while the relative abundance of genera such as Akkermansia and Bilophila is significantly increased in PD patients. These findings suggest that the gut microbiota may be an important biomarker for PD identification.
[0004] While existing research has provided preliminary evidence of a link between PD and the gut microbiota, a stable and reproducible genus of gut microbiota for PD identification has not yet been established. Heterogeneity among different research findings, along with the diversity of microbiota research methods, makes reliable comparisons and validations across different populations and regions difficult. Furthermore, existing studies often focus on single-cohort data analysis, failing to fully utilize large-scale, multi-center datasets for integrated analysis, thus failing to effectively overcome batch effects and other biases in research. These shortcomings limit the potential of the gut microbiota in the early identification of PD.
[0005] Therefore, existing technologies still have many shortcomings in PD identification based on the gut microbiota. There is a need for an identification and prediction model that can integrate large amounts of real-world research data and eliminate batch effects to improve the accuracy and reproducibility of PD identification, and to provide a non-invasive identification method with high sensitivity and specificity for clinical practice. Summary of the Invention
[0006] The technical problem solved by this invention is to provide a computer device, model and application capable of predicting Parkinson's disease.
[0007] To address the aforementioned technical problems, in a first aspect, the present invention provides a computer device comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to perform the following steps:
[0008] S1. Receiving data: Receiving sample data, which is the relative abundance of each of the 11 genera in the ex vivo fecal sample of the subject;
[0009] The 11 genera are as follows: Alistipes, Blautia, Butyricicoccus, Erysipelatoclostridium, Faecalibacterium, Family XIII AD3011 group, Family XIIIUCG-001, Lachnoclostridium, Lachnospiraceae FCS020 group, Lachnospiraceae UCG-004 and Ruminococcaceae UCG-002;
[0010] S2. Input data: Input the sample data into the Parkinson's disease prediction model;
[0011] The Parkinson's disease prediction model is constructed according to the following steps: the relative abundance of each of the 11 bacterial genera and the corresponding Parkinson's disease status in the ex vivo fecal samples of known healthy subjects and known Parkinson's disease patients are used as training samples to train the Parkinson's disease prediction model.
[0012] The Parkinson's disease status refers to either not having Parkinson's disease or having Parkinson's disease;
[0013] S3. Output results: Output the probability value of the subject's Parkinson's disease status through the Parkinson's disease prediction model; and then calculate the subject's Parkinson's disease status based on the probability value.
[0014] In the computer device described above, the relative abundance of each bacterial genera is equal to (the number of 16S rRNA sequences in each of the 11 bacterial genera / the total number of 16S rRNA sequences in the sample)%.
[0015] The 16S rRNA sequence number can be obtained by detecting any one or more of the following methods: metagenomic sequencing, 16S-rRNA sequencing, ITS sequencing, qRT-PCR, DNA blotting, and in situ hybridization.
[0016] The determination of the subject's Parkinson's disease status based on the probability value is as follows: if the subject's probability value is greater than the threshold (specifically 0.526), then the subject is or is a candidate for having Parkinson's disease; if the subject's probability value is less than or equal to 0.526, then the subject is or is a candidate for not having Parkinson's disease.
[0017] In the computer device described above, the Parkinson's disease condition can specifically be either not having Parkinson's disease or having Parkinson's disease.
[0018] In a second aspect, the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the steps in the computer device described in the first aspect.
[0019] Thirdly, the present invention provides a computer-readable storage medium storing a computer program that causes a computer to perform the steps in the computer device described in the first aspect.
[0020] Fourthly, the present invention provides a device for predicting Parkinson's disease in patients, the device comprising:
[0021] S1, Data receiving module: used to receive sample data, which is the relative abundance of each of the 11 genera in the ex vivo fecal sample of the subject;
[0022] The 11 genera are as follows: Alistipes, Blautia, Butyricicoccus, Erysipelatoclostridium, Faecalibacterium, Family XIII AD3011 group, Family XIIIUCG-001, Lachnoclostridium, Lachnospiraceae FCS020 group, Lachnospiraceae UCG-004 and Ruminococcaceae UCG-002;
[0023] S2, Data Input Module: Used to input the sample data into the Parkinson's Disease Prediction Model;
[0024] The Parkinson's disease prediction model is constructed according to the following steps: the relative abundance of each of the 11 bacterial genera and the corresponding Parkinson's disease status in the ex vivo fecal samples of known healthy subjects and known Parkinson's disease patients are used as training samples to train the Parkinson's disease prediction model.
[0025] The Parkinson's disease status refers to either not having Parkinson's disease or having Parkinson's disease;
[0026] S3. Result Output Module: Used to output the probability value of the subject's Parkinson's disease status through the Parkinson's disease prediction model; and then to calculate the subject's Parkinson's disease status based on the probability value.
[0027] In the above text, the Parkinson's disease status can specifically refer to not having Parkinson's disease or having Parkinson's disease.
[0028] Fifthly, the present invention provides a method for predicting or assisting in the prediction of Parkinson's disease status in a subject, the method comprising the following steps:
[0029] S1. Obtain the relative abundance of each of the 11 bacterial genera in the ex vivo fecal samples of the subjects;
[0030] The 11 genera are as follows: Alistipes, Blautia, Butyricicoccus, Erysipelatoclostridium, Faecalibacterium, Family XIII AD3011 group, Family XIIIUCG-001, Lachnoclostridium, Lachnospiraceae FCS020 group, Lachnospiraceae UCG-004 and Ruminococcaceae UCG-002;
[0031] S2. Input the sample data into the Parkinson's disease prediction model;
[0032] The Parkinson's disease prediction model is constructed according to the following steps: the relative abundance of each of the 11 bacterial genera and the corresponding Parkinson's disease status in the ex vivo fecal samples of known healthy subjects and known Parkinson's disease patients are used as training samples to train the Parkinson's disease prediction model.
[0033] The Parkinson's disease status refers to either not having Parkinson's disease or having Parkinson's disease;
[0034] S3. Output the probability value of the subject's Parkinson's disease status through the Parkinson's disease prediction model; and then determine the subject's Parkinson's disease status based on the probability value.
[0035] In a sixth aspect, the present invention provides a method for constructing a Parkinson's disease prediction model, the method comprising the following steps: the Parkinson's disease prediction model is constructed according to the following steps: the relative abundance of each of the 11 genera in the ex vivo fecal samples of known healthy subjects and known Parkinson's disease patients and the corresponding Parkinson's disease status are used as training samples to train the Parkinson's disease prediction model.
[0036] The Parkinson's disease status refers to either not having Parkinson's disease or having Parkinson's disease.
[0037] Advantages and benefits of this invention: This invention constructs a non-invasive PD prediction method and its predictive model based on gut microbiota, which has significant advantages: By analyzing the gut microbiota characteristics in fecal samples, this method provides a non-invasive means of PD prediction, avoiding the pain of traditional diagnosis and improving patient acceptance. This invention integrates data from multiple independent studies, identifies and optimizes 11 key microbial genera significantly associated with PD, and constructs a high-performance predictive model that demonstrates high accuracy and AUC on the test set, effectively distinguishing PD patients from healthy controls. This method supports the development of personalized medicine; by analyzing the gut microbiota characteristics of patients, customized PD prediction methods and treatment plans can be developed, providing patients with more precise medical services. Furthermore, the key microbial genera identified and the predictive model constructed in this invention provide a new perspective for further research on the etiology and pathogenesis of PD, helping to reveal potential pathogenic mechanisms and providing a theoretical basis for future treatment strategies. Compared to expensive and complex neuroimaging or gene testing, this method is lower in cost, has efficient early screening and identification capabilities, and is expected to be widely used in clinical practice. In summary, this invention demonstrates significant advantages in non-invasive prediction of Parkinson's disease (PD), personalized medical support, and cost-effectiveness, possessing important clinical application value and research prospects. It can guide clinicians in providing preventative or treatment plans for patients. With its advantages of being non-invasive, low-cost, and highly accurate, it can better intervene in the early progression of PD. Attached Figure Description
[0038] Figure 1 A flowchart for predicting PD prevalence using a PD prediction model. Detailed Implementation
[0039] The present invention will now be described in further detail with reference to specific embodiments. The given embodiments are merely illustrative of the invention and not intended to limit its scope. The embodiments provided below can serve as a guide for further improvements by those skilled in the art and do not constitute a limitation on the invention in any way.
[0040] Unless otherwise specified, the experimental methods used in the following examples are conventional methods, performed according to the techniques or conditions described in the literature in this field or according to the product instructions. Unless otherwise specified, the materials and reagents used in the following examples are commercially available.
[0041] Unless otherwise specified, the quantitative experiments in the following examples are all repeated three times, and the results are averaged.
[0042] Example 1: Construction and Validation of a Parkinson's Disease (PD) Prediction Model
[0043] 1. Data Collection and Integration
[0044] This invention employs a comprehensive search of the PubMed database to screen research literature related to "PD" and "gut microbiota," up to July 31, 2023. All ultimately included studies provided raw 16S rRNA gene sequencing data (FASTQ format) and patient information, resulting in six 16S rRNA gene sequencing datasets from five independent studies, totaling 550 PD patients and 456 healthy controls (HC). Data were downloaded from the ENA database using the SRA tool, including PRJEB55464, PRJNA391524, DRA009229, PRJNA381395, and PRJEB27564. Basic information on the subjects in the included datasets is shown in Table 1.
[0045] Table 1 shows the basic information of the subjects included in the gene sequencing dataset.
[0046]
[0047]
[0048] BMI, Body mass index; HC, healthy control; MDS-UPDRSIII, MDS Unified-Parkinson Disease Rating Scale Part 3; PD, Parkinson's disease.
[0049] The 550 PD patients and 456 healthy controls enrolled were divided into a training set (90% of the original dataset, including 495 PD patients and 410 healthy controls) and a validation set (10% of the original dataset, including 55 PD patients and 46 healthy controls).
[0050] 2. Analysis of 16S rRNA data of gut microbiota
[0051] Raw 16S rRNA sequencing data underwent standardized quality control and preprocessing. Fastp (V0.14.1) was used to overlap 200bp paired-end reads to generate longer tags. OTU clustering analysis was performed using Usearch (V10.0.240) and the UPARSE algorithm. Taxonomic annotation utilized the Silva V132 reference database, calculating the relative abundance of different microbial taxa in each sample. Population differences were compared using the Wilcoxon rank-sum test and the Kruskal-Wallis test. Linear discriminant analysis (LDA) was performed using the LEfSe tool.
[0052] 3. Acquisition of PD-related gut microbiota
[0053] The analysis revealed significant differences in microbial groups between PD patients and healthy controls. The relative abundance of microorganisms from different datasets and groups of subjects was input separately, and the NetMoss R package (Version 2) was used to further screen key microbial genera associated with PD. A NetMoss score exceeding 0.6 was set as a threshold to identify potential microorganisms associated with PD, resulting in 35 genera.
[0054] These 35 bacterial genera include: Faecalibacterium, Blautia, Coprococcus-2, Sutterella, Negativibacillus, Odoribacter, Rumiclostridium-1, dgA-11gut group, Alistipes, Butyricicoccus, Elusimicrobium, Lachnospiraceae UCG-004, Coprococcus-1, Ruminococcus-2, Romboutsia, Erysipelatoclostridium, Megasphaera, Parasutterella, Lachnospiraceae UCG-008, Methanobrevibacter, Prevotella-7, LachnospiraceaeFCS020 group, Ruminococcaceae UCG-002, Rikenellaceae RC9 gut group, Lachnoclostridium, Family XIII AD3011 group, Turicibacter, Coprobacter, Ruminococcaceae UCG-014, CAG-873, Desulfovibrio, Prevotella-2, Ruminococcus-1, Ruminococcaceae UCG-003, Family XIII UCG-001.
[0055] 4. Obtaining the target bacterial genus
[0056] The random forest algorithm was trained using the relative abundance and corresponding grouping (HC group, PD group) of the above 35 bacterial genera in each sample. The training process used ten-fold cross-validation with ten repetitions: First, the MeanDecreaseGini score of the random forest model constructed from these 35 bacterial genera was calculated and sorted from highest to lowest score; then, the top 5-35 bacterial genera were selected for model construction using the recursive feature elimination (RFE) method. It was found that the model with more than 10 features tended to be stable, and the model constructed from 11 bacterial genera had the best effect.
[0057] The 11 bacterial genera are as follows: Alistipes, Blautia, Butyricicoccus, Erysipelatoclostridium, Faecalibacterium, Family XIII AD3011 group, Family XIII UCG-001, Lachnoclostridium, Lachnospiraceae FCS020 group, Lachnospiraceae UCG-004 and Ruminococcaceae UCG-002.
[0058] 5. Construction of a PD prediction model based on non-invasive detection of gut microbiome
[0059] 1) Obtaining the relative abundance of each of the 11 genera.
[0060] 16S rRNA of microbial genera from fecal samples of 495 PD patients and 410 healthy controls in the training set mentioned above was obtained. The relative abundance of each genera in 11 genera in each sample was calculated according to the gut microbiota 16S rRNA data analysis in the previous two sections.
[0061] The relative abundance of each of the 11 genera is calculated as follows: (Number of 16S rRNA sequences in each of the 11 genera / Total number of 16S rRNA sequences in the sample)%
[0062] The 11 genera are as follows: Alistipes, Blautia, Butyricicoccus, Erysipelatoclostridium, Faecalibacterium, Family XIII AD3011 group, Family XIIIUCG-001, Lachnoclostridium, Lachnospiraceae FCS020 group, Lachnospiraceae UCG-004, and Ruminococcaceae UCG-002.
[0063] 2) Construction of a PD prediction model based on non-invasive detection of gut microbiome
[0064] The relative abundance of the 11 bacterial genera in each sample included in 1) above and the corresponding PD disease grouping (HC group, PD group) were converted into a data matrix, resulting in a total of 1006 samples with relative abundance of the 11 bacterial genera, including 550 PD patients and 456 HC control samples.
[0065] The relative abundance of 11 bacterial genera and PD prevalence (PD or HC) in the training set of 495 PD patients and 410 HC control samples were used as input data for modeling. The random forest algorithm was trained using the R package 'caret' to establish a PD prediction model. The model training process used ten iterations of tenfold cross-validation. The optimal parameters for this model are as follows: 500 trees and a maximum of 3 features considered at each split.
[0066] Based on the probability value of Parkinson's disease status for each sample in the training set output by the PD prediction model, different thresholds are used to convert the probability into the subject's predicted Parkinson's disease status. ROC curves are plotted and AUC=1 is calculated. At the same time, Precision and Recall under each threshold are calculated, and Precision-Recall curves are plotted. The reasonable threshold for the model is determined to be 0.526.
[0067] If the probability value of a subject is greater than 0.526, the subject is judged to have Parkinson's disease (PD); if the probability value of a subject is less than or equal to 0.526, the subject is judged to not have Parkinson's disease (healthy, HC).
[0068] In the training set, the model has 100% specificity and 100% sensitivity.
[0069] 6. Evaluate the overall performance of the model using validation set ROC curves and AUC values.
[0070] The remaining data matrix from the data matrix obtained in step 2) of section 5 above (55 PD patients and 46 HC patients) was used as the validation set. The validation set data was input into the PD prediction model to obtain the model probability value and predicted category. The model output the probability value and predicted category. Based on the actual clinical subjects' HC and PD groupings, ROC curves were plotted and the validation set AUC was calculated to be 0.864.
[0071] The relative abundance of each of the 11 genera in the above validation set is shown in Table 2-3 below.
[0072] The model's accuracy, specificity, and sensitivity were further evaluated using a confusion matrix. The confusion matrix categorizes true labels and model predictions into four classes: True Positive (TP), False Positive (FP), True Negative (TN), and False Negative (FN). Accuracy, specificity, and sensitivity were calculated based on the confusion matrix results. In the validation set, the model achieved an accuracy of 80.2%, a specificity of 87.0%, and a sensitivity of 78.2%.
[0073] The above results demonstrate that the model of the present invention can accurately predict PD patients.
[0074] Table 2 shows the relative abundance of some genera among the 11 bacterial genera in each sample of the validation set.
[0075]
[0076]
[0077]
[0078]
[0079] Note: Each column represents the relative abundance of 11 enterobacterial genera. The group indicates the grouping information. HC is the healthy control group, and PD is the Parkinson's disease group. S1-S101 represent 101 samples.
[0080] Table 3 shows the relative abundance of some genera among the 11 genera in each sample of the validation set.
[0081]
[0082]
[0083]
[0084] Note: Each column represents the relative abundance of 11 enterobacterial genera. The group indicates the grouping information. HC is the healthy control group, and PD is the Parkinson's disease group. S1-S101 represent 101 samples.
[0085] Example 2: PD prediction model for clinical samples
[0086] 1. Clinical specimen collection
[0087] This invention enrolled 50 confirmed PD patients and 50 healthy controls from Peking University Sixth Hospital to construct an independent validation cohort for model validation. All clinical information and biological samples of the patients were obtained from confirmed PD patients hospitalized at Peking University Sixth Hospital. Healthy controls were obtained from the patients' spouses, caregivers, or recruited volunteers with similar lifestyles from the same region. Both the PD patients and the healthy controls were from China.
[0088] 1) The inclusion and exclusion criteria for PD patients are as follows:
[0089] Inclusion criteria:
[0090] ① Patients clinically diagnosed with PD;
[0091] ②Age > 18 years old.
[0092] Exclusion criteria:
[0093] ① Patients who could not be diagnosed with PD;
[0094] ② Comorbid with other neurological or psychiatric disorders;
[0095] ③ Comorbid acute or chronic infectious diseases;
[0096] ④ Comorbid inflammatory bowel disease;
[0097] ⑤ Patients who have taken antibiotics within the past 3 months;
[0098] ⑥ Patients who have received radiotherapy or chemotherapy for malignant tumors within the past 6 months.
[0099] 2) The inclusion and exclusion criteria for healthy controls are as follows:
[0100] Inclusion criteria:
[0101] ① Clinically diagnosed as PD;
[0102] ② Recruit volunteers such as the patient's spouse, caregiver, or those with similar lifestyles in the same area;
[0103] ③ Age > 18 years old.
[0104] Exclusion criteria:
[0105] ① Suffering from lung diseases, infectious diseases, or other obvious systemic diseases;
[0106] ② Suffering from acute or chronic infectious diseases;
[0107] ③ Suffering from inflammatory bowel disease;
[0108] ④ Patients who have taken antibiotics within the past 3 months;
[0109] ⑤ Patients who have received radiotherapy or chemotherapy for malignant tumors within the past 6 months.
[0110] 3) Collect basic clinical information of the subjects, including: age, gender, duration of disease, BMI and other general information; medical history such as PD medication; MDS-Unified Parkinson Disease Rating Scale (MDS-UPDRS) part III score and PD gastrointestinal dysfunction scale score used to evaluate PD motor function.
[0111] 4) Collect fecal samples: Collect fecal samples from PD patients and healthy controls. In a clean bench, use a sterile fecal retrieval device to scoop out the middle section of fresh feces and place it in a sterile centrifuge tube. After being flash-frozen in liquid nitrogen, store at -80°C.
[0112] 2. Metagenomics sequencing and analysis
[0113] This invention utilizes metagenomics technology to detect the gut microbiota of subjects. First, fecal DNA was extracted from each subject's stool sample using a fecal genomic DNA extraction kit (MagPure Stool DNA KF Kit B). The genomic DNA was sonicated to fragments of 350 bp, and after end repair, DNBSEQ sequencing adapters were added using the MGIEAsy library preparation kit. PCR amplification and purification were performed, and the library was quality controlled using an Agilent 2100 analyzer. A circularization reaction system was prepared to obtain single-stranded circular products, which were then sequenced using a DNBSEQ-2000 sequencer.
[0114] Low-quality sequences were removed, and the samples were assembled de novo using the MEGAHIT assembly software. MetaGeneMark software was used for prediction, and CD-HIT software was used for redundancy removal. Clustering was performed based on sequence similarity. Salmon software was used for quantification to obtain standardized gene abundance values. Kraken2 was used for species annotation. At the same time, the species-level abundance of metagenomic samples was estimated using Bayesian algorithm and Kraken classification results through Bracken.
[0115] 3. Evaluate the overall performance of the model using ROC curves and AUC values from independent validation cohorts.
[0116] 1) Obtaining the relative abundance of each of the 11 genera.
[0117] Metagenomic sequencing data of microbial genera in fecal samples from 50 PD patients and 50 healthy controls in the independent validation cohort described in section 1 above were obtained. The relative abundance of each genera in the 11 genera in each sample was calculated according to the metagenomic data analysis described in section 2 above (Tables 4-5), forming a matrix of relative abundance data.
[0118] The relative abundance of each of the 11 genera is calculated as follows: (Number of sequences annotated to each of the 11 genera / Total number of sequences in the ex vivo fecal sample)%
[0119] The 11 genera are as follows: Alistipes, Blautia, Butyricicoccus, Erysipelatoclostridium, Faecalibacterium, Family XIII AD3011 group, Family XIIIUCG-001, Lachnoclostridium, Lachnospiraceae FCS020 group, Lachnospiraceae UCG-004, and Ruminococcaceae UCG-002.
[0120] 2) The data matrix of 50 PD patients and 50 HC patients obtained in 1) above was used as an independent validation cohort. The data from the independent validation cohort was input into the PD prediction model to obtain the model's predicted probability value and predicted category. The model output the predicted probability value and predicted category. Based on the actual clinical subjects' HC and PD groupings, ROC curves were plotted and the validation set AUC was calculated to be 0.779.
[0121] The model's accuracy, specificity, and sensitivity were further evaluated using a confusion matrix. The confusion matrix categorized true labels and model predictions into four classes: True Positive (TP), False Positive (FP), True Negative (TN), and False Negative (FN). Accuracy, specificity, and sensitivity were calculated based on the confusion matrix results. In the independent validation cohort, the model achieved an accuracy of 80.0%, a specificity of 74.0%, and a sensitivity of 88.0%.
[0122] The above results demonstrate that the model of the present invention can accurately predict PD patients.
[0123] Table 4 shows the relative abundance of some genera among the 11 bacterial genera in each sample of the independent validation cohort.
[0124]
[0125]
[0126]
[0127]
[0128] Note: Each column represents the relative abundance of each enterobacterial genera, and the group indicates the grouping information. HC is the healthy control group, and PD is the Parkinson's disease group. V1-V100 represent 100 samples.
[0129] Table 5 shows the relative abundance of some genera among the 11 bacterial genera in each sample of the independent validation cohort.
[0130]
[0131]
[0132]
[0133] Note: Each column represents the relative abundance of each enterobacterial genera, and the group indicates the grouping information. HC is the healthy control group, and PD is the Parkinson's disease group. V1-V100 represent 100 samples.
[0134] The present invention has been described in detail above. For those skilled in the art, the invention can be practiced in a wide range of ways with equivalent parameters, concentrations, and conditions without departing from its spirit and scope, and without requiring unnecessary experiments. Although specific embodiments have been given, it should be understood that further modifications can be made to the invention. In summary, according to the principles of the invention, this application is intended to include any changes, uses, or improvements to the invention, including changes made using conventional techniques known in the art that depart from the scope disclosed herein. Some of the essential features can be applied within the scope of the following appended claims.
Claims
1. A computer device comprising a memory, a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to perform the following steps: S1. Receiving data: Receiving sample data, which is the relative abundance of each of the 11 genera in the ex vivo fecal sample of the subject; The 11 genera are as follows: Alistipes, Blautia, Butyricicoccus, Erysipelatoclostridium, Faecalibacterium, Family XIII AD3011 group, Family XIIIUCG-001, Lachnoclostridium, Lachnospiraceae FCS020 group, Lachnospiraceae UCG-004 and Ruminococcaceae UCG-002; S2. Input data: Input the sample data into the Parkinson's disease prediction model; The Parkinson's disease prediction model is constructed according to the following steps: the relative abundance of each of the 11 bacterial genera and the corresponding Parkinson's disease status in the ex vivo fecal samples of known healthy subjects and known Parkinson's disease patients are used as training samples to train the Parkinson's disease prediction model. The Parkinson's disease status refers to either not having Parkinson's disease or having Parkinson's disease; S3. Output results: Output the probability value of the subject's Parkinson's disease status through the Parkinson's disease prediction model; and then calculate the subject's Parkinson's disease status based on the probability value.
2. A computer program product, including a computer program, characterized in that: When executed by a processor, the computer program performs the steps described in the computer apparatus of claim 1.
3. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program that causes a computer to perform the steps in the computer device of claim 1.
4. A device for predicting Parkinson's disease in patients, the device comprising: S1, Data receiving module: used to receive sample data, which is the relative abundance of each of the 11 genera in the ex vivo fecal sample of the subject; The 11 genera are as follows: Alistipes, Blautia, Butyricicoccus, Erysipelatoclostridium, Faecalibacterium, Family XIII AD3011 group, Family XIIIUCG-001, Lachnoclostridium, Lachnospiraceae FCS020 group, Lachnospiraceae UCG-004 and Ruminococcaceae UCG-002; S2, Data Input Module: Used to input the sample data into the Parkinson's Disease Prediction Model; The Parkinson's disease prediction model is constructed according to the following steps: the relative abundance of each of the 11 bacterial genera and the corresponding Parkinson's disease status in the ex vivo fecal samples of known healthy subjects and known Parkinson's disease patients are used as training samples to train the Parkinson's disease prediction model. The Parkinson's disease status refers to either not having Parkinson's disease or having Parkinson's disease; S3. Result Output Module: Used to output the probability value of the subject's Parkinson's disease status through the Parkinson's disease prediction model; and then to calculate the subject's Parkinson's disease status based on the probability value.
5. A method for predicting or assisting in the prediction of a subject's Parkinson's disease status, characterized in that: The method includes the following steps: S1. Obtain the relative abundance of each of the 11 bacterial genera in the ex vivo fecal samples of the subjects; The 11 genera are as follows: Alistipes, Blautia, Butyricicoccus, Erysipelatoclostridium, Faecalibacterium, Family XIII AD3011 group, Family XIIIUCG-001, Lachnoclostridium, Lachnospiraceae FCS020 group, Lachnospiraceae UCG-004 and Ruminococcaceae UCG-002; S2. Input the sample data into the Parkinson's disease prediction model; The Parkinson's disease prediction model is constructed according to the following steps: the relative abundance of each of the 11 bacterial genera and the corresponding Parkinson's disease status in the ex vivo fecal samples of known healthy subjects and known Parkinson's disease patients are used as training samples to train the Parkinson's disease prediction model. The Parkinson's disease status refers to either not having Parkinson's disease or having Parkinson's disease; S3. Output the probability value of the subject's Parkinson's disease status through the Parkinson's disease prediction model; and then determine the subject's Parkinson's disease status based on the probability value.
6. A method for constructing a Parkinson's disease prediction model, characterized in that: The method includes the following steps: The Parkinson's disease prediction model is constructed according to the following steps: The relative abundance of each of the 11 bacterial genera and the corresponding Parkinson's disease status in the ex vivo fecal samples of known healthy subjects and known Parkinson's disease patients are used as training samples to train the Parkinson's disease prediction model. The Parkinson's disease status refers to either not having Parkinson's disease or having Parkinson's disease.