A method and system for diagnosing or predicting a disease based on a combination of gut microbial markers
By screening causally related gut microbiota biomarker combinations through Mendelian randomization analysis and combining them with machine learning models, the problem of poor correlation of feature values in existing autism prediction models has been solved, achieving efficient autism diagnosis and risk prediction, and has significant clinical application potential.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HANGZHOU DEPU MEDICAL LAB CO LTD
- Filing Date
- 2024-06-14
- Publication Date
- 2026-04-28
AI Technical Summary
Existing models for predicting autism based on gut microbiota markers have poor correlations with the disease, resulting in low causal relationships and an inability to effectively explain the relationship with autism, thus lacking accurate diagnostic or predictive methods.
Gut microbial biomarkers with a causal relationship to autism were screened using Mendelian randomization analysis. A combination of microbial biomarkers, including Ruminiclostridium, Sutterella, and Eggerthella, was constructed. Abundance data were obtained using 16S RNA sequencing in conjunction with a machine learning model to build a disease prediction model.
It achieves accurate diagnosis and risk prediction of autism, with an AUC of 0.9623, 0.8409 on the internal test set, and 0.75 on the external test set. It enables timely medical intervention and has great clinical application value.
Smart Images

Figure CN118957049B_ABST
Abstract
Description
[0001] Related patents
[0002] This application is a divisional application of Chinese Patent Application No. 2024107640214, filed on June 14, 2024, entitled "A combination, system and application of microbial biomarkers for diagnosing or predicting autism". Technical Field
[0003] This invention belongs to the field of microbial biomarker technology, specifically, it relates to a method and system for diagnosing or predicting diseases based on a combination of gut microbial biomarkers, particularly, the disease being autism. Background Technology
[0004] Autism spectrum disorders (ASD) are severe neurobehavioral developmental disorders that begin in infancy. Children with ASD typically develop symptoms between 6 and 24 months of age, although some may have normal development initially, gradually exhibiting regressive changes such as loss of language and social skills between 24 and 36 months. In recent years, the incidence of ASD has been increasing annually.
[0005] Clinical evidence demonstrates that gut microbiota dysbiosis and altered metabolic products are closely related to the development of ASD. The gut is the largest digestive, immune, and endocrine organ in the human body, and gut microbiota are considered a "second genome" influencing human physiological and psychological well-being. Compared to healthy children, children with autism have significantly increased numbers of Clostridium species in their feces, and autism significantly improves after treatment with the antibiotic vancomycin. Clostridium not only produces enterotoxins that cause gastrointestinal diseases but also neurotoxins that contribute to autism. Researchers used pyrosequencing to study the gut microbiota of children with autism and found that severely autistic children had significantly increased levels of Bacteroidetes and Actinobacteria. Therefore, metabolic disorders caused by abnormal gut microbiota may be one of the pathogenic mechanisms of ASD.
[0006] Therefore, screening gut microbiota biomarkers and using machine learning models to predict the risk of autism can help humans better predict autism and thus enable early intervention. However, there is currently a severe lack of models for predicting autism based on gut microbiota biomarkers; some predictive models have poor correlations with autism and low causal relationships, meaning that these biomarkers often cannot explain the relationship with autism. Summary of the Invention
[0007] To solve at least one of the above-mentioned technical problems, the inventors obtained gut microbiota biomarkers that are causally related to autism through Mendelian randomization analysis, and further determined the causal relationship between the biomarkers and autism. All the obtained biomarkers are gut microbiota that are causally associated with autism, which can be used for better diagnosis or prediction.
[0008] The first aspect of this invention provides a combination of microbial biomarkers for diagnosing or predicting the risk of autism, including Ruminiclostridium , Sutterella , Eggerthella , Holdemanella , Odoribacter , Oscillibacter , Turicibacter , Terrisporobacter , Blautia and Adlercreutzia .
[0009] Furthermore, the microbial biomarker combination also includes Ruminococcus , Slackia , Eubacterium , Desulfovibrio , Holdemania , Gordonibacter , Dorea , Collinsella and Anaerotruncus At least one of them.
[0010] In some preferred embodiments of the present invention, the microbial biomarker combination includes Ruminiclostridium , Sutterella , Eggerthella , Holdemanella , Odoribacter , Oscillibacter , Turicibacter , Terrisporobacter , Blautia , Adlercreutzia .
[0011] In some preferred embodiments of the present invention, the microbial biomarker combination includes Ruminiclostridium , Sutterella , Slackia , Eggerthella , Holdemanella , Odoribacter , Oscillibacter , Turicibacter , Terrisporobacter , Blautia , Adlercreutzia .
[0012] In some preferred embodiments of the present invention, the microbial biomarker combination includes Ruminiclostridium , Sutterella , Slackia , Eggerthella , Holdemanella , Odoribacter , Oscillibacter , Turicibacter , Terrisporobacter , Collinsella , Blautia , Adlercreutzia .
[0013] In some preferred embodiments of the present invention, the microbial biomarker combination includes Ruminococcus , Ruminiclostridium , Sutterella , Slackia , Eggerthella , Holdemanella , Odoribacter , Oscillibacter , Turicibacter , Terrisporobacter , Collinsella , Blautia , Adlercreutzia .
[0014] In some preferred embodiments of the present invention, the microbial biomarker combination includes Ruminococcus , Ruminiclostridium , Sutterella , Slackia , Eggerthella , Holdemanella , Odoribacter , Oscillibacter , Turicibacter , Terrisporobacter , Dorea , Collinsella , Blautia , Adlercreutzia .
[0015] In some preferred embodiments of the present invention, the microbial biomarker combination includes Ruminococcus , Ruminiclostridium , Sutterella , Slackia , Eggerthella , Holdemanella , Holdemania , Odoribacter , Oscillibacter , Turicibacter , Terrisporobacter , Dorea , Collinsella , Blautia , Adlercreutzia .
[0016] In some preferred embodiments of the present invention, the microbial biomarker combination includes Ruminococcus , Ruminiclostridium , Sutterella , Slackia , Eubacterium , Eggerthella ,Holdemanella , Holdemania , Odoribacter , Oscillibacter , Turicibacter , Terrisporobacter , Dorea , Collinsella , Blautia , Adlercreutzia .
[0017] In some preferred embodiments of the present invention, the microbial biomarker combination includes Ruminococcus , Ruminiclostridium , Sutterella , Slackia , Eubacterium , Desulfovibrio , Eggerthella , Holdemanella , Holdemania , Odoribacter , Oscillibacter , Turicibacter , Terrisporobacter , Dorea , Collinsella , Blautia , Adlercreutzia .
[0018] In some preferred embodiments of the present invention, the microbial biomarker combination includes Ruminococcus , Ruminiclostridium , Sutterella , Slackia , Eubacterium , Desulfovibrio , Eggerthella , Holdemanella , Holdemania , Odoribacter , Oscillibacter , Turicibacter , Terrisporobacter , Dorea , Collinsella , Blautia , Anaerotruncus , Adlercreutzia .
[0019] In some preferred embodiments of the present invention, the microbial biomarker combination includes Ruminococcus , Ruminiclostridium , Sutterella , Slackia , Eubacterium , Desulfovibrio , Eggerthella , Holdemanella , Holdemania , Odoribacter , Oscillibacter , Gordonibacter , Turicibacter , Terrisporobacter , Dorea ,Collinsella , Blautia , Anaerotruncus and Adlercreutzia .
[0020] By detecting the above-mentioned microbial biomarkers and obtaining their abundance data, models can be constructed to diagnose whether subjects have autism or predict whether subjects are at risk of developing autism.
[0021] A second aspect of the present invention provides a system for diagnosing or predicting the risk of autism, comprising the following modules:
[0022] The data input module is used to input the abundance data of any combination of microbial biomarkers described in the first aspect of the present invention in the obtained biological samples of the subjects;
[0023] A database storage module is used to store abundance data of the microbial biomarker combinations in biological samples of a population, the population samples including samples from non-autistic subjects and samples from autistic patients;
[0024] The disease prediction module is connected to the data input module and the database storage module, respectively, and is used to construct a prediction model using the abundance data of the combination of microbial markers in the biological samples of the population, and to diagnose whether the subject has autism or predict whether the subject is at risk of having autism based on the abundance data of the combination of microbial markers of the subject obtained from the data input module.
[0025] In this invention, the term "abundance data of a combination of microbial markers" includes the abundance data of each microbial marker in the combination of microbial markers.
[0026] In some embodiments of the present invention, the abundance data of each microbial biomarker are obtained based on qPCR, 16S RNA sequencing or metagenomic sequencing methods.
[0027] In some specific embodiments of the present invention, it refers to obtaining it using 16S RNA sequencing, specifically including:
[0028] Genomic DNA was extracted from the biological sample and 16S RNA was sequenced to obtain sequencing data.
[0029] The raw reads of the sequencing data are preprocessed to filter out high-quality reads, which are then compared with the 16S RNA gene reference database. Chimeric sequences are removed. Finally, the filtered sequences are clustered according to a certain method (including but not limited to NanoCLUST) to obtain multiple sequence clustering operation taxonomic units (OTUs). Taxonomy annotation is performed on the OTUs to obtain the abundance data of each microorganism.
[0030] In some embodiments of the present invention, the abundance is relative abundance.
[0031] In this invention, the biological sample includes, but is not limited to, feces, intestinal lavage fluid, and anal swab samples, preferably, a feces sample.
[0032] In some embodiments of the present invention, the disease prediction module, in which the prediction model is constructed using abundance data of the combination of microbial biomarkers in the biological samples of the population, includes the following steps:
[0033] S21, the abundance data of the combination of microbial markers in the biological samples of the population are randomly divided into two groups, one as a training set and the other as a test set. Each group includes the abundance data of the microbial markers in samples from non-autistic subjects and samples from autistic patients.
[0034] S22, using training set data, constructs a disease prediction model based on machine learning algorithms and performs multi-fold cross-validation;
[0035] S23. Validate the obtained prediction model on the test set.
[0036] In some embodiments of the present invention, stratified random sampling is used for grouping. The grouping ratio can be 7:3, 4:1, etc.
[0037] In some embodiments of the present invention, the machine learning algorithm is selected from any of the following algorithms: logistic regression algorithm, linear regression algorithm, random forest algorithm, neural network algorithm, support vector machine algorithm, Bayesian classification algorithm, gradient boosting algorithm, K-nearest neighbor algorithm, and decision tree algorithm.
[0038] In some specific embodiments of the present invention, the machine learning algorithm is a logistic regression algorithm, and the prediction model diagnoses whether the subject has autism or predicts whether the subject has a risk of having autism and the level of that risk based on the score obtained by the logistic regression algorithm.
[0039] In some preferred embodiments of the present invention, the microbial biomarkers include all 19 microbial biomarkers mentioned above, and the regression coefficients of each microbial biomarker in the prediction model are as follows:
[0040]
[0041] In some preferred embodiments of the present invention, when the score is less than 0.5, the subject does not have autism or has a low risk of having autism; when the score is greater than 0.75, the subject has autism or has a high risk of having autism; otherwise, the subject has a medium risk of having autism.
[0042] The third aspect of the present invention provides the use of an abundance detection reagent of any combination of microbial biomarkers described in the first aspect of the present invention in the preparation of a kit for diagnosing or predicting autism.
[0043] In some preferred embodiments of the present invention, the abundance detection reagent refers to a high-throughput sequencing reagent, including nucleic acid extraction, amplification and / or purification reagents.
[0044] In some embodiments of the present invention, the abundance detection reagent includes primers and / or probes. Further, the probes are fabricated into a chip.
[0045] Beneficial effects of the present invention
[0046] Compared with the prior art, the present invention achieves the following beneficial effects:
[0047] Using the microbial biomarkers of this invention, a machine learning model was established, achieving an AUC of 0.9623 on the training set, 0.8409 on the internal test set, and 0.75 on the external test set. This indicates that the microbial biomarkers of this invention can accurately diagnose whether a subject has autism or predict whether a subject is at risk of developing autism. Furthermore, based on a specific machine learning model, it is also possible to achieve precise predictions based on the predicted risk level of a subject developing autism, enabling timely medical intervention and possessing significant clinical application value. Attached Figure Description
[0048] Figure 1 The results of analyzing ebi-a-GCST90016596 using the inverse variance weighting method in Embodiment 1 of the present invention are shown.
[0049] Figure 2 The results of analyzing ebi-a-GCST90016599 using the inverse variance weighting method in Embodiment 1 of the present invention are shown.
[0050] Figure 3 The results of analyzing ebi-a-GCST90016602 using the inverse variance weighting method in Embodiment 1 of the present invention are shown.
[0051] Figure 4 The results of analyzing ebi-a-GCST90016603 using the inverse variance weighting method in Embodiment 1 of the present invention are shown.
[0052] Figure 5 The results of analyzing ebi-a-GCST90016606 using the inverse variance weighting method in Embodiment 1 of the present invention are shown.
[0053] Figure 6The results of analyzing ebi-a-GCST90016614 using the inverse variance weighting method in Embodiment 1 of the present invention are shown.
[0054] Figure 7 The results of analyzing ebi-a-GCST90016620 using the inverse variance weighting method in Embodiment 1 of the present invention are shown.
[0055] Figure 8 The results of analyzing finn-b-F5_AUTISM using the inverse variance weighting method in Embodiment 1 of the present invention are shown.
[0056] Figure 9 The results of analyzing finn-b-F5_PERVASIVE using the inverse variance weighting method in Embodiment 1 of the present invention are shown.
[0057] The results of analyzing finn-b-KRA_PSY_AUTISM using the inverse variance weighting method in Embodiment 1 of the present invention are shown.
[0058] Figure 10 The results of analyzing finn-b-KRA_PSY_AUTISM_EXMORE using the inverse variance weighting method in Embodiment 1 of the present invention are shown.
[0059] Figure 11 The results of analyzing IEU-A-802 using the inverse variance weighting method in Embodiment 1 of the present invention are shown.
[0060] Figure 12 The results of analyzing IEU-A-806 using the inverse variance weighting method in Embodiment 1 of the present invention are shown.
[0061] Figure 13 The results of analyzing ieu-a-1184 using the inverse variance weighting method in Embodiment 1 of the present invention are shown.
[0062] Figure 14 The results of analyzing IEU-A-1185 using the inverse variance weighting method in Embodiment 1 of the present invention are shown.
[0063] Figure 15 The performance evaluation of the machine learning model in Embodiment 3 of the present invention on the training set and the test set is shown.
[0064] Figure 16 The results of the machine learning model in Embodiment 3 of the present invention on the determination of the risk threshold for autism are shown.
[0065] Figure 17 The performance evaluation of the machine learning model in Embodiment 3 of the present invention on an external validation set is shown. Detailed Implementation
[0066] Unless otherwise stated, implied from the context, or as is customary in the art, all parts and percentages in this application are based on weight, and all testing and characterization methods used are concurrent with the filing date of this application. Where applicable, any patent, patent application, or disclosure relating to this application is incorporated herein by reference in its entirety, and its equivalent patent families are also incorporated herein by reference, in particular the definitions of relevant terms in the art disclosed in such documents. If any definition of a specific term disclosed in the prior art is inconsistent with any definition provided in this application, the definition provided in this application shall prevail.
[0067] To make the technical problems solved by the present invention, the technical solutions and the beneficial effects of the present invention clearer, the present invention will be further described in detail below with reference to the embodiments.
[0068] The following examples are used to illustrate preferred embodiments of the invention. Those skilled in the art will understand that the techniques disclosed in the examples represent techniques discovered by the inventors that can be used to implement the invention, and therefore can be considered preferred embodiments for implementing the invention. However, those skilled in the art should understand from this specification that many modifications can be made to the specific embodiments disclosed herein, still yielding the same or similar results, without departing from the spirit or scope of the invention.
[0069] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains, and all materials publicly cited herein and referenced by them are incorporated herein by reference.
[0070] Those skilled in the art will recognize, or can learn through routine experimentation, many equivalents of the specific embodiments of the invention described herein. These equivalents will be included in the claims.
[0071] Unless otherwise specified, the experimental methods used in the following examples are conventional methods. Unless otherwise specified, the instruments and equipment used in the following examples are all conventional laboratory instruments and equipment; unless otherwise specified, the experimental materials used in the following examples were all purchased from conventional biochemical reagent stores.
[0072] Example 1: Screening of gut microbiota biomarkers for autism
[0073] This embodiment screens for gut microbiota biomarkers associated with autism. The specific screening steps are as follows:
[0074] 1. Obtain the dataset
[0075] The genetic variation data of gut microbiota came from the genome-wide association study (GWAS) published by the Microbial Genome Consortium in 2021. This study analyzed the host genotypes and 16S rRNA metagenomic data of 18,340 participants, including 122,110 single nucleotide polymorphism (SNP) variant sites, representing 211 bacterial taxa, including 15 unknown bacterial taxa.
[0076] The genetic variation data for autism came from the IEU OpenGWAS project and included 15 datasets: ebi-a-GCST90016620, ebi-a-GCST90016614, ieu-a-1185, ebi-a-GCST90016606, ebi-a-GCST90016596, ebi-a-GCST90016599, ieu-a-802, ebi-a-GCST90016603, ebi-a-GCST90016602, ieu-a-1184, ieu-a-806, finn-b-F5_AUTISM, finn-b-KRA_PSY_AUTISM, finn-b-KRA_PSY_AUTISM_EXMORE, and finn-b-F5_PERVASIVE.
[0077] 2. Select instrumental variables
[0078] Inclusion criteria for instrumental variables: The significance threshold for SNPs was set at P < 1.0 × 10⁻⁶. -5 The criterion for chain imbalance is r. 2 <0.001, genetic distance 10000kb, remove highly correlated SNPs, retain the SNP with the lowest P value.
[0079] The formula for calculating the F-value is:
[0080] F=[R 2 / (R 2 -1)]×[(NK-1) / K]
[0081] Where N represents the sample size of the gut microbiota GWAS study, K represents the number of instrumental variables, and R0 2 R represents the degree of exposure explained by SNP. 2 The calculation formula is: R 2 =2×β 2 ×(1-MAF)×MAF, where MAF is the minor allele frequency and β is the effect size of the SNP on exposure.
[0082] An F>10 value is considered to effectively avoid bias from weak instrumental variables, therefore SNPs with an F<10 value are removed. A total of 1519 SNPs were obtained as instrumental variables.
[0083] 3. Statistical Analysis
[0084] Mendelian randomization analysis was performed using the R TwoSampleMR algorithm package. MR analysis methods included inverse-variance weighted analysis (IVW), weighted median, MR-Egger, simple mode, and weighted mode. Eigenvalue selection was mainly performed using the IVW method, with P < 0.05 considered statistically significant.
[0085] 4. Sensitivity and pleiotropic effects analysis
[0086] Regarding sensitivity, this embodiment calculates Cochran's Q statistic using IVW and MR-Egger regression, respectively. A p-value > 0.05 indicates no significant heterogeneity. Simultaneously, the inventors also used a leave-one-out method to systematically remove included SNPs, observing their impact on the analysis results and constructing a forest plot. If removing a certain SNP results in a p-value > 0.05, the SNP is considered to have no significant impact on the results.
[0087] Regarding pleiotropy, the inventors simultaneously used the intercept term of MR-Egger regression and the Mendelian randomization pleiotropy residual sum and outlier (MR-PRESSO) to test for level pleiotropy in the included SNPs. In MR-Egger regression, if the intercept tends to 0, then level pleiotropy can be considered absent. Level pleiotropy screening retains eigenvalues with p > 0.05.
[0088] 5. Summary and Analysis
[0089] IVW is considered the standard method for summarizing MR data. This method uses the Wald ratio method to calculate the causal effect estimate for each included instrumental SNP, and then performs a weighted summarization analysis.
[0090] The inverse variance weighted analysis results (Mendelian randomized forest plot) of the above 15 data sets in this embodiment are as follows: Figure 18 As shown.
[0091] Depend on Figures 1 - 15 It can be seen that in ebi-a-GCST90016596,Figure 1 (Heliotrophic NK4A136 group) Lachnospiraceae NK4A136 group Lachnospiraceae NC2004 group (Heliotrophic NC2004 group) and Butyricicoccus The causal relationship between *Butyricum* and autism reached a statistically significant level (P<0.05), among which, Lachnospiraceae NK4A136 group and Butyricicoccus It was negatively correlated with the occurrence of autism, suggesting it is a protective factor against autism, with effect sizes (Odd Ratio, OR) of 1.014 (0.95% CI: 1.001-1.027) and 0.984 (0.95% CI: 0.971-0.998), respectively. Lachnospiraceae NC2004 group It was positively correlated with the occurrence of autism, suggesting it is a risk factor for autism, with an odds ratio of 0.984 (0.95% CI: 0.970-0.997).
[0092] Depend on Figure 2 It can be seen that in ebi-a-GCST90016599, Christensenellaceae R-7 group (Christensen bacteria) Dorea ( Dorea (Fungi genus) Escherichia.Shigella (Shigella) Paraprevotella (Palapuleria) and Eubacterium coprostanoligenes group The causal relationship between (coprosterol-producing eubacteria) and autism reached a statistically significant level (P<0.05), among which, Christensenellaceae R-7group , Dorea and Escherichia.Shigella It was positively correlated with the occurrence of autism, suggesting it is a risk factor for autism, with OR values of 1.025 (0.95% CI: 1.002-1.048), 1.030 (0.95% CI: 1.003-1.057), and 1.032 (0.95% CI: 1.008-1.056), respectively. Paraprevotella and Eubacterium coprostanoligenes group It was negatively correlated with the occurrence of autism, suggesting that it is a protective factor against autism, with OR values of 0.985 (0.95% CI: 0.970-0.999) and 0.975 (0.95% CI: 0.951-0.999), respectively.
[0093] Depend on Figure 3 It can be seen that in ebi-a-GCST90016602, Marvinbryantia (Marvinburyian) Ruminococcaceae UCG 005 (Ruminococci family UCG 005) Parasutterella (Parasartella spp.) Sutterella (Sartreus genus) Lachnospiraceae NC2004 group (Clorizobacterium NC2004 group) Lachnospiraceae UCG 001(UCG 001, family Trichophyceae) Lachnospira (Syngonium) and Olsenella The causal relationship between *Eurotium tumefaciens* and autism reached a statistically significant level (P<0.05). Marvinbryantia , Parasutterella , Lachnospiraceae NC2004 group , Lachnospiraceae UCG 001 and Olsenella It was positively correlated with the occurrence of autism, suggesting it is a risk factor for autism, with OR values of 1.023 (0.95% CI: 1.002-1.045), 1.027 (0.95% CI: 1.010-1.045), 1.020 (0.95% CI: 1.003-1.038), 1.023 (0.95% CI: 1.003-1.043), and 1.015 (0.95% CI: 1.002-1.028), respectively. Ruminococcaceae UCG 005 , Sutterella and Lachnospira It was negatively correlated with the occurrence of autism, suggesting it was a protective factor against autism, with OR values of 0.979 (0.95% CI: 0.962-0.996), 0.965 (0.95% CI: 0.932-0.999), and 0.956 (0.95% CI: 0.917-0.997), respectively.
[0094] Depend on Figure 4 It can be seen that in ebi-a-GCST90016603, Marvinbryantia (Marvinburyian) Ruminococcus 1 (Ruminococcus 1) Butyrivibrio (Vibrio butyricum) Dialister (Alisteria oryzae) Sutterella (Sartreus genus) Dorea ( Dorea (Fungi genus) Ruminococcaceae UCG 005 (Ruminococci family UCG 005) Anaerofilum (Sulfate-reducing bacteria) and Lachnospiraceae UCG 001 The causal relationship between (UCG001, family Trichophytonceae) and autism reached a statistically significant level (P<0.05), among which, Marvinbryantia , Butyrivibrio and Lachnospiraceae UCG 001 It was positively correlated with the occurrence of autism, suggesting it was a risk factor for autism, with OR values of 1.027 (0.95% CI: 1.006-1.047), 1.014 (0.95% CI: 1.003-1.024), and 1.032 (0.95% CI: 1.009-1.056), respectively. Ruminococcus 1 , Dialister , Sutterella , Dorea ,Ruminococcaceae UCG 005 and Anaerofilum It was negatively correlated with the occurrence of autism, suggesting it was a protective factor against autism, with OR values of 0.974 (0.95% CI: 0.954-0.994), 0.978 (0.95% CI: 0.957-1.000), 0.956 (0.95% CI: 0.921-0.991), 0.969 (0.95% CI: 0.942-0.996), 0.971 (0.95% CI: 0.950-0.992), and 0.980 (0.95% CI: 0.965-0.995), respectively.
[0095] Depend on Figure 5 It can be seen that in ebi-a-GCST90016606, Candidatus Soleaferrea (Bacillus phloem) Gordonibacter (Goldenbacterium parvum) Roseburia (Roseidon genus) Turicibacter (Spirogenic enteric bacteria) and Lachnospiraceae UCG 008 The causal relationship between (UCG 008, family Trichophyceae) and autism reached a statistically significant level (P<0.05), among which, Roseburia and Lachnospiraceae UCG 008 It was positively correlated with the occurrence of autism, suggesting it was a risk factor for autism, with OR values of 1.029 (0.95% CI: 1.009-1.050) and 1.020 (0.95% CI: 1.003-1.037), respectively. Candidatus Soleaferrea , Gordonibacter and Turicibacter It was negatively correlated with the occurrence of autism, suggesting it was a protective factor against autism, with OR values of 0.970 (0.95% CI: 0.948-0.992), 0.988 (0.95% CI: 0.977-0.999), and 0.981 (0.95% CI: 0.964-0.999), respectively.
[0096] Depend on Figure 6 It can be seen that in ebi-a-GCST90016614, Ruminococcaceae UCG 005 (Ruminococci family UCG 005) Ruminococcus 1 (Ruminococcus 1) Parabacteroides (Pseudomonas) Ruminococcaceae UCG 013 (Ruminococci family UCG 013) Anaerofilum (Sulfate-reducing bacteria) Fusicatenibacter (Fusobacterium) and Sutterella The causal relationship between *Sartella* and autism reached a statistically significant level (P<0.05), where R... uminococcaceae UCG 005 , Ruminococcus 1 , Anaerofilum ,Fusicatenibacter and Sutterella It was positively correlated with the occurrence of autism, suggesting it is a risk factor for autism, with OR values of 1.018 (0.95% CI: 1.005-1.030), 1.014 (0.95% CI: 1.001-1.028), 1.010 (0.95% CI: 1.000-1.021), 1.015 (0.95% CI: 1.001-1.028), and 1.025 (0.95% CI: 1.008-1.044), respectively. Parabacteroides and Ruminococcaceae UCG 013 It was negatively correlated with the occurrence of autism, suggesting it was a protective factor against autism, with OR values of 0.981 (0.95% CI: 0.967-0.995) and 0.987 (0.95% CI: 0.974-1.000), respectively.
[0097] Depend on Figure 7 It can be seen that in ebi-a-GCST90016620, Adlercreutzia (Adlerkcroitz) Anaerofilum (Sulfate-reducing bacteria) Desulfovibrio (Desulfovibrio) and Gordonibacter The causal relationship between *Goldenella parviflora* and autism reached a statistically significant level (P<0.05), among which... Adlercreutzia and Anaerofilum It was positively correlated with the occurrence of autism, suggesting it was a risk factor for autism, with OR values of 1.026 (0.95% CI: 1.002-1.050) and 1.035 (0.95% CI: 1.015-1.056), respectively. Desulfovibrio and Gordonibacter It was negatively correlated with the occurrence of autism, suggesting it was a protective factor against autism, with OR values of 0.963 (0.95% CI: 0.937-0.989) and 0.986 (0.95% CI: 0.976-0.996), respectively.
[0098] Depend on Figure 8 It can be seen that in finn-b-F5_AUTISM, Terrisporobacter (Tropicoides) and Collinsella The causal relationship between *Collinus* spp. and autism reached a statistically significant level (P<0.05), among which, Terrisporobacter It was positively correlated with the occurrence of autism, suggesting it is a risk factor for autism, with an odds ratio (OR) of 2.987 (0.95% CI: 1.174-7.598), which is highly significant. Collinsella It was negatively correlated with the occurrence of autism, suggesting it is a protective factor against autism, with an OR of 0.316 (0.95% CI: 0.119-0.838), which is highly significant.
[0099] Depend on Figure 9 It can be seen that in finn-b-F5_PERVASIVE, Blautia (Broutella) Eubacterium oxidoreducens group (Redox Eubacteria) Adlercreutzia (Adlerkcroitz) Gordonibacter (Goldenbacterium parvum) Sellimonas The causal relationship between *Cephalosporium* and autism was statistically significant (P<0.05). Gordonibacter and Sellimonas It was positively correlated with the occurrence of autism, suggesting it is a risk factor for autism, with OR values of 1.890 (0.95% CI: 1.057-3.380) and 1.950 (0.95% CI: 1.101-3.452), which were highly significant. Blautia , Eubacterium oxidoreducens group and Adlercreutzia It was negatively correlated with the occurrence of autism, suggesting it is a protective factor against autism, with OR values of 0.142 (0.95% CI: 0.039-0.516), 0.335 (0.95% CI: 0.125-0.898), and 0.389 (0.95% CI: 0.161-0.944), which were highly significant.
[0100] Depend on Figure 10 It can be seen that in finn-b-KRA_PSY_AUTISM, Collinsella (Collinsella spp.) Odoribacter (Porphyromaceae family) and Terrisporobacter The causal relationship between *Teriphylloxera* and autism reached a statistically significant level (P<0.05), among which, Odoribacter and Terrisporobacter A positive correlation between the two factors suggests that they are risk factors for autism, with OR values of 3.006 (95% CI: 1.028-8.793) and 3.108 (95% CI: 1.200-8.049), which are highly significant. Collinsella It was negatively correlated with the occurrence of autism, suggesting it is a protective factor against autism, with an OR of 0.336 (95% CI: 0.125-0.905), which is highly significant.
[0101] Depend on Figure 11 It can be seen that in finn-b-KRA_PSY_AUTISM_EXMORE, Oscillibacter (Octospora spp.) and Blautia The causal relationship between *Brutella* spp. and autism reached a statistically significant level (P<0.05), among which, OscillibacterIt was positively correlated with the occurrence of autism, suggesting it is a risk factor for autism, with an OR value of 1.996 (95% CI: 1.030-3.865), which is highly significant; Blautia It was negatively correlated with the occurrence of autism, suggesting it is a protective factor against autism, with an OR of 0.281 (95% CI: 0.095-0.830), which is highly significant.
[0102] Depend on Figure 12 It can be seen that in IEU-A-802, Lachnospiraceae NC2004 group (Clorizobacterium NC2004 group) Howardella (Hodgkin's bacterium) Slackia (Slack bacteria) Eubacterium hallii group (Hodgkin's group of bacteria) Turicibacter (sporogenic intestinal bacteria) Desulfovibrio (Desulfovibrio) Eggerthella (Aegyptiella spp.) Adlercreutzia (Adlerkcroitz) Anaerotruncus (Anaerobic Entomologous bacteria) and Ruminiclostridium 6 The causal relationship between (Clostridium rumenii 6) and autism reached a statistically significant level (P<0.05), among which, Lachnospiraceae NC2004 group , Slackia , Turicibacter and Eggerthella It was positively correlated with the occurrence of autism, suggesting it is a risk factor for autism, with OR values of 1.367 (95% CI: 1.017-1.838), 1.386 (95% CI: 1.031-1.864), 1.372 (95% CI: 1.031-1.825), and 1.436 (95% CI: 1.054-1.955), which were highly significant. Howardella , Eubacterium hallii group , Desulfovibrio , Adlercreutzia , Anaerotruncus and Ruminiclostridium 6 It was negatively correlated with the occurrence of autism, suggesting it is a protective factor against autism. The odds ratios (ORs) were 0.789 (95% CI: 0.630-0.990), 0.575 (95% CI: 0.390-0.848), 0.683 (95% CI: 0.486-0.960), 0.695 (95% CI: 0.516-0.936, P = 0.017), 0.603 (95% CI: 0.397-0.916), and 0.733 (95% CI: 0.545-0.987), which were highly significant.
[0103] Depend on Figure 13 It can be seen that in IEU-A-806, Dorea ( Dorea (Fungi genus) Lachnospiraceae NC2004group (Clorizobacterium NC2004 group) Anaerotruncus (Anaerobic clumps) Ruminiclostridium 9 (Clostridium rumenii 9) Holdemania (Holdman's bacterium) and Eisenbergiella The causal relationship between Eisenberger's bacteria and autism reached a statistically significant level (P<0.05). Lachnospiraceae NC2004 group , Ruminiclostridium 9 and Holdemania It was positively correlated with the occurrence of autism, suggesting it is a risk factor for autism, with OR values of 1.255 (95% CI: 1.002-1.573), 1.556 (95% CI: 1.168-2.071), and 1.196 (95% CI: 1.002-1.429), which were highly significant. Dorea , Anaerotruncus and Eisenbergiella It was negatively correlated with the occurrence of autism, suggesting it is a protective factor against autism, with OR values of 0.721 (95% CI: 0.532-0.979), 0.699 (95% CI: 0.525-0.931), and 0.783 (95% CI: 0.652-0.939), which were highly significant.
[0104] Depend on Figure 14 It can be seen that in ieu-a-1184, Turicibacter (sporogenic intestinal bacteria) Lachnospiraceae NC2004 group (Clorizobacterium NC2004 group) Dorea ( Dorea (Fungi genus) Holdemanella (Holdman's bacterium) Ruminiclostridium 9 (Clostridium rumenii 9) Anaerotruncus (Anaerobic Entomologous bacteria) and Eisenbergiella The causal relationship between Eisenberger's bacteria and autism reached a statistically significant level (P<0.05). Turicibacter , Lachnospiraceae NC2004 group , Holdemanella and Ruminiclostridium 9 It was positively correlated with the occurrence of autism, suggesting it is a risk factor for autism, with OR values of 1.372 (95% CI: 1.056-1.782), 1.308 (95% CI: 1.029-1.662), 1.322 (95% CI: 1.030-1.696), and 1.486 (95% CI: 1.036-2.131), which were highly significant. Dorea , Anaerotruncus and EisenbergiellaIt was negatively correlated with the occurrence of autism, suggesting it is a protective factor against autism, with OR values of 0.705 (95% CI: 0.508-0.978), 0.696 (95% CI: 0.519-0.935), and 0.785 (95% CI: 0.649-0.949), which were highly significant.
[0105] Depend on Figure 15 It can be seen that in IEU-A-1185, Ruminococcus 1 (Ruminococcus 1) Turicibacter (sporogenic intestinal bacteria) Ruminiclostridium 5 (Clostridium rumenii 5) Sutterella (Sartreus genus) Dorea ( Dorea (Fungi genus) Ruminococcaceae UCG 005 The causal relationship between (Ruminococcus family UCG 005) and autism reached a statistically significant level (P<0.05), among which, Turicibacter It was positively correlated with the occurrence of autism, suggesting it is a risk factor for autism, with an OR of 1.372 (95% CI: 1.056-1.782), which is highly significant; Ruminococcus 1 , Ruminiclostridium 5 , Sutterella , Dorea and Ruminococcaceae UCG 005 It was negatively correlated with the occurrence of autism, suggesting it is a protective factor against autism. The odds ratios (ORs) were 0.831 (95% CI: 0.705-0.981), 0.812 (95% CI: 0.687-0.961), 0.821 (95% CI: 0.684-0.987, P = 0.036), 0.811 (95% CI: 0.686-0.959), and 0.776 (95% CI: 0.670-0.898), which were highly significant.
[0106] After filtering out eigenvalues with small effect values (OR values), the final eigenvalues (biomarkers) include 19 genera of gut microbiota, as shown in Table 1:
[0107] Table 1 19 microbial biomarkers
[0108]
[0109] Example 2: Autism Prediction Machine Learning Model
[0110] This embodiment uses the LogisticRegression function from the Python sklearn.linear_model module to perform logistic regression modeling on the data. Logistic regression maps the results of multiple linear regression analysis to the logit function z=1 / (1+exp(y)), and binarizes the data according to a threshold to predict binary variables.
[0111] This embodiment constructs a predictive model for calculating the risk of autism in a host based on the abundance values of gut microbiota biomarkers screened in Example 1. The specific steps are as follows:
[0112] S1, Data Acquisition: The datasets for both the control group and the case group were obtained from the microphenodb database.
[0113] S2, the dataset is split into a training set and an internal validation set. In this embodiment, 25% of the samples are used as the internal validation set.
[0114] S3. A logistic regression model was established using cross-validation. The model performed best when the regularization strength coefficient Cs was 50. Using the best regularization strength coefficient, a new logistic regression model was trained using the training set. The regression coefficients of each microbial biomarker are shown in Table 2. The constant term = 13.6968.
[0115] Table 2 Regression coefficients of microbial biomarkers
[0116]
[0117] The above model was used for analysis on the training set, and the AUC was 0.9623.
[0118] S4 uses the trained model to make predictions on the internal validation set.
[0119] The ROC curve of the internal validation set is as follows: Figure 16 As shown, AUC = 0.8409, accuracy = 78%.
[0120] S5. Based on the predicted probability values of the dataset, the samples are divided into low-risk, medium-risk, and high-risk groups. The prediction score for low-risk is <0.5, and the prediction score for high-risk is >0.75.
[0121] The total sample size was 324 cases, of which 54.3% (176 people) were low-risk, 20.1% (65 people) were medium-risk, and 25.6% (83 people) were high-risk. Figure 17 As shown in the figure. Therefore, it can be seen that the risk threshold selected in this embodiment can successfully stratify the population according to different risks of developing autism.
[0122] Example 3 Model Performance Evaluation
[0123] The inventors applied the autism prediction model constructed in Example 2 to an external validation set.
[0124] The external validation set was primarily derived from real physical examination sample data. The sample type was feces, totaling 19 samples, including 8 cases of autism and 11 healthy controls. Nanopore 16S RNA sequencing analysis was performed. The specific analysis steps are as follows:
[0125] (1) Sequencing
[0126] DNA extraction: DNA was extracted using a fecal sample bacterial genome extraction kit, and the concentration and purity of the extracted DNA were detected using Nanodrop;
[0127] 16S rRNA amplification: The V4 region of the bacterial 16S rRNA gene was targeted and amplified using a 16S sequencing library preparation kit. The reaction was repeated 30 times under the following conditions: 95℃ for 1 min, 94℃ for 30 s, 62℃ for 30 s, 62℃ for 2 min, and 62℃ for 5 min; then held at 4℃.
[0128] Recovery of PCR products: The target fragment was purified using AMPure XP beads and quantified using Qubit;
[0129] Adapter ligation: Mix the above barcoded libraries according to their concentration ratio to a 10 μL system, ensuring the total DNA volume is 200 ng. Add 1 μL of RAP to the mixed sample, gently mix with a pipette tip, and incubate at room temperature for 5 min.
[0130] Finally, the constructed library was sequenced using the GridION nanopore sequencer.
[0131] (2) Bioinformatics analysis
[0132] First, NanoStat software was used for quality control of the raw data, filtering and removing sequencing adapters to obtain high-quality reads. Minimap2 global alignment was used to identify chimeras, and then yacrd was used to remove chimeric sequences. Filtlong software was then used to filter data of specific lengths and Q values. Finally, the filtered sequences were clustered according to a specific method (NanoCLUST in this example) to obtain multiple sequence clustering operational taxonomic units (OTUs), and taxonomic annotations were performed on the OTUs to obtain species abundance information.
[0133] (3) Model prediction
[0134] Abundance information of gut microbiota markers in the case group and the control group was collected, and then model prediction and risk level classification were performed.
[0135] The predictive analysis of the external validation data was performed using the model constructed in Example 2, and the results are as follows: Figure 18 As shown, AUC=0.75, accuracy 62%. These results indicate that the model constructed using Example 2 can distinguish between autistic patients and healthy individuals, and can be used clinically to assist in the diagnosis of whether a subject has autism.
[0136] Example 4: Further Evaluation of Model Performance
[0137] To further verify that the model constructed in Example 2 can be used to predict the risk of developing autism, the inventors collected 110 samples. Using the predictive model, 16 cases were classified as high-risk, 65 as low-risk, and the remaining 29 as medium-risk.
[0138] Further follow-up revealed that 7 high-risk cases showed autistic tendencies or were diagnosed with autism, while 64 low-risk cases did not develop autism. This demonstrates the high accuracy of the predictive model.
[0139] High-risk groups will continue to be followed up annually, while medium-risk groups will receive early intervention through family education and other means.
[0140] Example 5: Further screening of markers
[0141] The inventors used a stepwise regression approach to further screen the microbial biomarkers obtained in Example 2, and the results are shown in Table 3:
[0142] Table 3 Results of further screening of microbial biomarkers
[0143]
[0144] Note: In Table 3, R 2 The coefficient of determination is represented by AIC, which stands for Akaike Information Criterion. AIC is a standard for evaluating the complexity of a statistical model and measuring its goodness of fit. AIC = 2k - 2ln(L), where k is the number of model parameters and L is the likelihood function. BIC stands for Bayesian Information Criterion, used to evaluate the goodness of a model. The smaller the BIC, the better the model. BIC = kln(n) - ln(L), where k is the number of model parameters, n is the sample size, and L is the likelihood function. MSE stands for Mean Squared Error, which is a measure of the difference between the model's predicted values and the actual values.
[0145] As shown in Table 3, the biomarkers selected through stepwise regression were removed sequentially according to their numbers, based on the combination of microbial biomarkers in Table 1. Gordonibacter , Anaerotruncus , Desulfovibrio , Eubacterium , Holdemania , Dorea , Ruminococcus , Collinsella and Slackia .
[0146] The model prediction was performed using the combination of biomarkers selected by the stepwise regression method described above, and the results are shown in Table 4.
[0147] Table 4. Prediction results of the model using the combinations of biomarkers obtained from stepwise regression.
[0148]
[0149] As shown in Table 4, when only microbial markers are present... Ruminiclostridium , Sutterella , Eggerthella , Holdemanella , Odoribacter , Oscillibacter , Turicibacter , Terrisporobacter , Blautia and Adlercreutzia At that time, the AUC was high on both the training set and the internal validation set, and reached 0.6932 on the external validation set. When further gradually increased... Slackia , Collinsella , Ruminococcus and Dorea After that, AUC did not improve significantly, but it gradually increased further. Holdemania , Eubacterium , Desulfovibrio , Anaerotruncus and Gordonibacter Afterwards, the AUC on the external validation set increased significantly, reaching around 0.75.
[0150] All documents mentioned in this invention are incorporated herein by reference as if each document were individually incorporated by reference. Furthermore, it should be understood that after reading the foregoing teachings of this invention, those skilled in the art can make various alterations or modifications to this invention, and these equivalent forms also fall within the scope defined by the appended claims.
Claims
1. A method for constructing a disease prediction model based on a combination of gut microbial biomarkers, characterized in that, The disease prediction model is used to diagnose or predict autism, and the combination of gut microbiota biomarkers includes: Ruminococcus , Ruminiclostridium , Sutterella , Slackia , Eubacterium , Desulfovibrio , Eggerthella , Holdemanella , Holdemania , Odoribacter , Oscillibacter , Gordonibacter , Turicibacter , Terrisporobacter , Dorea , Collinsella , Blautia , Anaerotruncus and Adlercreutzia The method includes the following steps: S21, the abundance data of the gut microbiota marker combination in the biological samples of the population are randomly divided into two groups, one is a training set and the other is a test set. Each group includes the abundance data of the gut microbiota markers in samples from non-autistic subjects and samples from autistic patients. S22, using training set data, constructs a disease prediction model based on machine learning algorithms and performs multi-fold cross-validation; S23. Validate the obtained disease prediction model on the test set. The machine learning algorithm used is logistic regression, with a regularization strength coefficient Cs of 50. The regression coefficients of each gut microbiota marker are as follows: Constants = 13.6968 The prediction model diagnoses whether a subject has autism or predicts whether a subject is at risk of having autism based on a score obtained from a logistic regression algorithm. When the score is less than 0.5, the subject does not have autism or has a low risk of having autism; when the score is greater than 0.75, the subject has autism or has a high risk of having autism; otherwise, the subject has a medium risk of having autism.
2. The method according to claim 1, characterized in that, Abundance data for the gut microbiota biomarker ensemble were obtained based on qPCR, 16S RNA sequencing, or metagenomic sequencing methods.
3. A system for diagnosing or predicting autism, characterized in that, Includes the following modules: The data input module is used to input the abundance data of the gut microbiota biomarker combination in the obtained subject biological samples, wherein the gut microbiota biomarker combination includes Ruminococcus , Ruminiclostridium , Sutterella , Slackia , Eubacterium , Desulfovibrio , Eggerthella , Holdemanella , Holdemania , Odoribacter , Oscillibacter , Gordonibacter , Turicibacter , Terrisporobacter , Dorea , Collinsella , Blautia , Anaerotruncus and Adlercreutzia ; A database storage module is used to store abundance data of the gut microbiota marker combinations in biological samples of a population, including samples from non-autistic subjects and samples from autistic patients. The disease prediction module, connected to both the data input module and the database storage module, is used to construct a prediction model using abundance data of the gut microbiota marker combinations in the biological samples of the population, and to diagnose whether a subject has autism or predict whether a subject is at risk of developing autism based on the abundance data of the gut microbiota marker combinations of the subjects obtained from the data input module. In the disease prediction module, the step of constructing a prediction model using abundance data of the gut microbiota marker combinations in the biological samples of the population includes the following steps: S21, the abundance data of the gut microbiota marker combination in the biological samples of the population are randomly divided into two groups, one as a training set and the other as a test set. Each group includes the abundance data of the microbiota markers in samples from non-autistic subjects and samples from autistic patients. S22, using training set data, constructs a disease prediction model based on machine learning algorithms and performs multi-fold cross-validation; S23. Validate the obtained prediction model on the test set. The machine learning algorithm used is logistic regression, with a regularization coefficient Cs of 50. The prediction model uses the scores obtained from the logistic regression algorithm to diagnose whether a subject has autism or to predict whether a subject is at risk of developing autism. The regression coefficients of each gut microbiota marker are as follows: Constants = 13.6968 When the score is less than 0.5, the subject does not have autism or has a low risk of having autism; when the score is greater than 0.75, the subject has autism or has a high risk of having autism; otherwise, the subject has a medium risk of having autism.
4. The system according to claim 3, characterized in that, The abundance data of the gut microbiota biomarker combinations obtained in the data input module are based on qPCR, 16S RNA sequencing or metagenomic sequencing methods.
Citation Information
Patent Citations
Oral flora microorganisms for children autism assessment
CN111455076A
Colorectal adenoma intestinal microbial marker and application thereof
CN117965715A