Training random forest models for assessing autism spectrum disorders using fecal biomarkers
By detecting multi-kingdom microbial biomarkers in fecal samples and analyzing them using a random forest model, this study addresses the inaccuracy of existing ASD diagnostic techniques, providing a highly sensitive and specific diagnostic method to support the clinical management and prevention of ASD.
Patent Information
- Application Number
- CN202480042812.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-02-26
- Filing Date
- 2024-08-21
- Publication Date
- 2026-02-13
AI Technical Summary
The current understanding of the association between the gut microbiome and autism spectrum disorder (ASD) is not comprehensive enough, especially in terms of the reproducibility of biomarkers and diagnostic accuracy.
By using fecal microbiota characteristics and combining them with random forest model analysis, we can detect multi-kingdom microbial biomarkers (including archaea, bacteria, fungi, viruses, microbial functional pathways, and metagenomic assembly genomes) to assess the presence or risk of ASD. We can also use machine learning methods to train a random forest classifier to improve diagnostic accuracy.
It achieves highly sensitive and specific ASD diagnosis in different age groups and populations, reduces the false positive rate, provides a non-invasive diagnostic tool, and supports clinical decision-making and management.
Smart Images

Figure CN121532831A_ABST
Abstract
Description
Background of the Invention
[0002] Autism spectrum disorder (ASD) is a neurodevelopmental disorder characterized by persistent challenges in limited and repetitive patterns of social interaction, communication, and behavior, interests, or activities. ASD affects an individual throughout their life and is typically diagnosed in early childhood. The disorder encompasses a wide range of symptoms and severity levels, leading to the use of the term "spectrum."
[0003] In recent decades, ASD has received significant attention due to its increasing prevalence and its impact on individuals, families, and society as a whole. According to the Centers for Disease Control and Prevention (CDC), approximately one in 36 children in the United States is diagnosed with ASD. While improved diagnostic criteria and increased awareness may have contributed to this, the prevalence has steadily risen over the years. The exact causes of ASD are still under investigation, but it is believed to result from a combination of genetic, environmental, and developmental factors. ASD affects individuals of all races, ethnicities, and socioeconomic backgrounds, and it is more prevalent in males than in females. Invention Overview
[0005] Although studies have explored the relationship between the gut microbiome and autism spectrum disorder (ASD), the conventional understanding of the association between the microbiome and ASD at the genomic level remains incomplete. For example, questions remain regarding the reproducibility of biomarkers across cohorts and ages. This application's specification explores the potential of using fecal microbiome signatures to aid in the diagnosis of ASD. Accumulated evidence from three independent studies suggests alterations in the gut microbiome in children with ASD.
[0006] In the first study, metagenomic sequencing was performed on fecal samples from 1017 phenotypically normal children, yielding 6.24 terabytes of sequence data. Metagenetic analysis revealed alterations in 19 archaea, 126 bacteria, 6 fungi, 24 viruses, 909 microbial genes, and 104 microbial pathways in children with ASD compared to neurotypical children. Machine learning analysis using a single-kingdom panel showed an area under the operator curve (AUROC) ranging from 0.61 to 0.85 in distinguishing children with ASD from neurotypical children. A multi-kingdom and microbial functional biomarker, consisting of a panel of 27 biomarkers (2 archaea, 2 fungi, 2 viruses, 8 bacteria, 6 microbial genes, and 7 pathways), demonstrated excellent diagnostic accuracy for ASD, with an AUC of 0.88 at 89.7% specificity and 90.5% sensitivity. The model maintained an AUC of 0.81 in two independent validation cohorts of different ages and showed a low false positive rate in two unrelated disease cohorts.
[0007] In a public dataset of 195 samples from different populations, the diagnostic value of the 27 biomarkers across multiple boundaries remained significant (p = 9.4e-07). Functional analysis revealed that the model's accuracy was primarily driven by the thiamine diphosphate biosynthesis pathway, which is less abundant in children with ASD, indicating a relationship between microbial thiamine metabolism and ASD. Overall, these findings highlight the potential application of multiple boundaries and functional gut microbiota biomarkers as non-invasive diagnostic tools for ASD.
[0008] In the second study, metagenomic sequencing was performed on 598 stool samples to investigate the potential of microbial taxonomy and metagenomically assembled genome (MAG) biomarkers to distinguish children with ASD from age- and sex-matched neurotypical children. For ASD detection, the combined microbial taxa and microbial MAG biomarkers showed better classification performance than either the taxa alone or the MAG biomarkers alone. A machine learning model incorporating 5 bacterial taxa and 44 microbial MAG biomarkers (2 viral MAGs and 42 bacterial MAGs) achieved an AUROC of 0.886 in the discovery cohort and 0.734 in the independent validation cohort. A trade-off was made to stratify the model's output into high-confidence and median-confidence regions by analogy to authorized behavioral diagnostic procedures and to improve model efficiency. In the validation cohort, the model's high-confidence region showed an AUROC of 0.906 with 85.7% sensitivity and 93.4% specificity in distinguishing children with ASD from neurotypical children.
[0009] Furthermore, a significant positive correlation was observed between higher ASD risk scores and more severe social impairment symptoms as measured by the Social Responsiveness Scale (SRS). The microbiome panel demonstrated superior classification performance in children aged 6 years and older (AUROC 0.845) compared to older children (>6 years), and the model was broadly applicable to subjects of different sexes, with or without gastrointestinal symptoms (constipation and diarrhea), and with or without comorbid mental illnesses (ADHD and anxiety). This study highlights the potential clinical effectiveness of fecal microbiome in aiding ASD diagnosis.
[0010] In the third study, a total of 1,627 children (aged 1–13 years, 24.4% female) from five independent cohorts were recruited. Extensive phenotypic data (236 factors) were collected, including age, sex, body mass index (BMI), diet, medication use, comorbidities, associated mental disorders, gastrointestinal (GI) symptoms (including stool consistency assessed by the Bristol Stool Morphology Score (BSFS)), family characteristics, and technical factors related to sample collection, storage, and processing. All stool samples were processed using the same standardized protocol to reduce batch effects caused by technical factors. The findings of the third study were validated in a public dataset of 237 fecal metagenomics. In the association of the gut microbiome in children with ASD, 31 fecal microbiome biomarkers were identified, demonstrating good and reproducible diagnostic performance for ASD across different ages and cohorts.
[0011] The following provides an overview of various embodiments of the invention by way of a list of examples. As used below, any reference to a series of examples should be understood to refer to each of those examples individually (e.g., "examples 1-4" should be understood as "examples 1, 2, 3 or 4").
[0012] Example 1 is a computer-implemented method for training a random forest model to predict the presence of autism spectrum disorder (ASD) in subjects, the method comprising: obtaining a training dataset comprising abundance measures of a set of microbial markers (or biomarkers) for a first cohort of subjects with ASD and abundance measures of a set of microbial markers (or biomarkers) for a second cohort of subjects without ASD; dividing the training dataset into k subsets; for each of the k subsets, training k candidate random forest models by: generating a validation dataset comprising one of the k subsets and a training subset comprising the remaining k-1 subsets of the k subsets; training candidate random forest models among the k candidate random forest models using the training subsets and evaluating the candidate random forest models using the validation dataset by calculating a performance metric for the candidate random forest models; evaluating the k candidate random forest models by calculating the area under the curve (AUC) value of each of the k candidate random forest models; and deploying one of the k candidate random forest models as a trained random forest model based on the one of the k candidate random forest models with the largest AUC value.
[0013] Example 2 is a computer implementation of the method of Example 1, wherein training a candidate random forest model using training subsets includes: generating n training subsets by sampling (i) objects from a first queue of objects and a second queue of objects and (ii) microbial markers (or biomarkers) from a set of microbial markers (or biomarkers); and constructing n decision trees using the n training subsets, the n decision trees forming the candidate random forest model.
[0014] Example 3 is a computer-implemented method of Examples 1-2, wherein objects from a first queue of objects and a second queue of objects are sampled with replacement, and wherein microbial markers from a set of microbial markers are sampled with replacement.
[0015] Example 4 is a computer-implemented method of Example 2, wherein objects from a first queue of objects and a second queue of objects are randomly sampled, and wherein microbial markers from a set of microbial markers are randomly sampled.
[0016] Example 5 is a computer-implemented method of Examples 1-4, wherein the performance metrics include one or more of the following: total number of false positives, total number of true positives, total number of false negatives, and total number of true negatives.
[0017] Example 6 is a computer implementation of the methods in Examples 1-5, wherein the k candidate random forest models with the largest AUC values have a first set of hyperparameters. The computer implementation method further includes: training k second candidate random forest models with a second set of hyperparameters different from the first set; and evaluating the k second candidate random forest models by calculating their AUC values; wherein the k candidate random forest models or the k second candidate random forest models are deployed as trained random forest models based on the k candidate random forest models with the largest AUC values or the k second candidate random forest models with the largest AUC values.
[0018] Example 7 is a computer implementation of the method in Example 6, which further includes: training k third candidate random forest models with a third set of hyperparameters that are different from the first set of hyperparameters and the second set of hyperparameters; and evaluating the k third candidate random forest models by calculating their AUC values; wherein the k candidate random forest models, k second candidate random forest models, or k third candidate random forest models are selected as the random forest models to be trained based on the k candidate random forest models with the largest AUC value, the k second candidate random forest models with the largest AUC value, or the k third candidate random forest models with the largest AUC value.
[0019] Example 8 is a computer implementation of the method of Example 6, wherein the first set of hyperparameters and the second set of hyperparameters differ in at least one of the following aspects: the number of decision trees in their respective candidate random forest models; the maximum number of nodes in the decision trees in their respective candidate random forest models; the maximum or minimum number of sampled objects from the first and second queues of objects used to construct the decision trees in their respective candidate random forest models; or the maximum or minimum number of sampled microbial markers from the set of microbial markers used to construct the decision trees in their respective candidate random forest models.
[0020] Example 9 is a computer implementation of the methods in Examples 1-5, wherein k candidate random forest models have a first set of hyperparameters, and the computer implementation further includes: training k second candidate random forest models with a second set of hyperparameters different from the first set; evaluating the k second candidate random forest models by calculating the AUC value of each of the k second candidate random forest models; and deploying one of the k candidate random forest models or one of the k second candidate random forest models as a trained random forest model based on one of the k candidate random forest models with the largest AUC value or one of the k second candidate random forest models with the largest AUC value.
[0021] Example 10 is a computer implementation of the method in Example 9, wherein a set of hyperparameters corresponding to one of the k candidate random forest models with the largest AUC value or one of the k second candidate random forest models with the largest AUC value is used to retrain the random forest model using the entire reference dataset and is deployed as the trained random forest model.
[0022] Example 11 is a computer implementation of the method of Example 9, which further includes: training k third candidate random forest models with a third set of hyperparameters different from the first set of hyperparameters and the second set of hyperparameters; and evaluating the k third candidate random forest models by calculating the AUC value of each of the k third candidate random forest models; wherein one of the k candidate random forest models, one of the k second candidate random forest models, or one of the k third candidate random forest models is selected as the random forest model to be trained based on one of the k candidate random forest models with the largest AUC value, one of the k second candidate random forest models with the largest AUC value, or one of the k third candidate random forest models with the largest AUC value.
[0023] Example 12 is a computer implementation of the method of Example 9, wherein the first set of hyperparameters and the second set of hyperparameters differ in at least one of the following aspects: the number of decision trees in their respective candidate random forest models; the maximum number of nodes in the decision trees in their respective candidate random forest models; the maximum or minimum number of sampled objects from the first and second queues of objects used to construct decision trees in their respective candidate random forest models; or the maximum or minimum number of sampled microbial markers from the set of microbial markers used to construct decision trees in their respective candidate random forest models.
[0024] Example 13 is a computer implementation of the method in Examples 1-12, which further includes: comparing the AUC values of each of the k candidate random forest models to identify the maximum AUC value.
[0025] Example 14 is a computer implementation of the method of Examples 1-13, wherein obtaining training data includes: obtaining a reference dataset; and splitting the reference dataset into a training dataset and a test dataset, wherein the test dataset is used to evaluate k candidate random forest models.
[0026] Example 15 is a computer-implemented method of Examples 1-14, wherein the abundance measurements of the first and second queue objects are relative abundance measurements.
[0027] Example 16 is a method for assessing the risk or presence of autism spectrum disorder (ASD) in subjects, comprising: (a) generating a classifier algorithm from a reference dataset, wherein the reference dataset is obtained from subjects in a cohort of subjects with a known classification of ASD and subjects in another cohort of subjects without ASD by measuring the relative abundance of one or more microbiome-derived biomarkers selected from Tables 1, 2, 3 and / or 4 in fecal samples taken from the subjects; (b) detecting the relative abundance of one or more microbiome-derived biomarkers selected from Tables 1, 2, 3 and / or 4 in fecal samples taken from the subjects; (c) comparing the relative abundance of one or more microbiome-derived biomarkers selected from Tables 1, 2, 3 and / or 4 obtained from (b) with the reference dataset by applying the classification algorithm generated from (a); and (d) determining the probability of ASD in the subjects, wherein if the probability is higher than a determined threshold, the subjects are determined to be at high risk of ASD.
[0028] Example 17 is a method for assessing the risk or presence of autism spectrum disorder (ASD) in subjects, comprising: (a) generating a classifier algorithm from a reference dataset, wherein the reference dataset is obtained from a cohort of subjects with a known classification of ASD and another cohort of subjects without ASD by measuring the relative abundance of all functional genes detected in the reference dataset in fecal samples of each subject in the cohort, and wherein the mean abundance and prevalence are greater than or equal to 0.15% and 5%, respectively; (b) detecting the relative abundance of all functional genes detected in the reference dataset in fecal samples taken from the subjects; (c) comparing the relative abundance of all functional genes detected in the reference dataset obtained from (b) with the reference dataset by applying the classification algorithm generated from (a); and (d) determining the probability of ASD in the subject, wherein if the probability is higher than a determined threshold, the subject is determined to be at high risk of having ASD.
[0029] Example 18 is a method for assessing the risk or presence of autism spectrum disorder (ASD) in subjects, comprising: (a) generating a classifier algorithm from a reference dataset, wherein the reference dataset is obtained from a cohort of subjects with a known classification of ASD and another cohort of subjects without ASD by measuring the relative abundance of all microbial functional pathways detected in the reference dataset in fecal samples of each subject in the cohort, and wherein the mean abundance and prevalence are greater than or equal to 0.15% and 5%, respectively; (b) detecting the relative abundance of all microbial functional pathways detected in the reference dataset in fecal samples taken from the subjects; (c) comparing the relative abundance of all microbial functional pathways detected in the reference dataset obtained from (b) with the reference dataset by applying the classification algorithm generated from (a); and (d) determining the probability of ASD in the subject, wherein if the probability is higher than a determined threshold, the subject is determined to be at high risk of having ASD.
[0030] Example 19 is a method for assessing the risk or presence of autism spectrum disorder (ASD) in subjects, comprising: (a) generating a classifier algorithm from a reference dataset, wherein the reference dataset is obtained from a cohort of subjects with a known classification of ASD and another cohort of subjects without ASD by measuring the relative abundance of all combinations of microbial functional pathways, bacteria, functional genes, archaea, fungi, and viruses detected in the reference dataset in fecal samples of each subject in the cohort, and wherein the mean abundance and prevalence are greater than or equal to 0.15% and 5%, respectively; (b) detecting the relative abundance of all combinations of microbial functional pathways, bacteria, functional genes, archaea, fungi, and viruses detected in the reference dataset in fecal samples taken from the subjects; (c) comparing the relative abundance of all combinations of microbial functional pathways, bacteria, functional genes, archaea, fungi, and viruses detected in the reference dataset obtained from (b) with the reference dataset by applying the classification algorithm generated from (a); and (d) determining the probability of ASD in the subject, wherein if the probability is higher than a determined threshold, the subject is determined to be at high risk of ASD.
[0031] Example 20 is a method of any of Examples 16-19, wherein the subject has not used probiotics, antidepressants, antiepileptic drugs, or been exposed to antibiotics in the month prior to the test.
[0032] Example 21 is a method of any one of Examples 16-20, where the object is male.
[0033] Example 22 is a method of any one of Examples 16-20, where the object is female.
[0034] Example 23 is a method of any of Examples 16-20, wherein a set of queue objects with a known classification of ASD and another set of queue objects without ASD are both male.
[0035] Example 24 is a method of any of Examples 16-20, wherein a set of queue objects with a known classification of ASD and another set of queue objects without ASD are both female.
[0036] Example 25 is a method of any of Examples 16-20, wherein a set of queue objects with a known classification of having ASD and another set of queue objects without having ASD include male and female objects.
[0037] Example 26 is a computer-readable medium comprising instructions that, when executed by one or more processors, cause one or more processors to perform the method of any of the preceding claims.
[0038] Example 27 is a system comprising: one or more processors; and a computer-readable medium comprising instructions that, when executed by the one or more processors, cause the one or more processors to perform the method of any one of the preceding claims. Brief description of the attached diagram
[0040] Figure 1A and Figure 1B Example steps of a method for training a random forest classifier to predict the presence of ASD in objects are shown.
[0041] Figure 2 Example steps of a method for training a random forest classifier to predict the presence of ASD in objects are shown.
[0042] Figure 3 Example steps of a method for applying a trained random forest model to predict the presence of ASD in objects are shown.
[0043] Figure 4 The different stages of the research are shown.
[0044] Figure 5 This demonstrates the variation in multi-kingdom microbiome composition explained by the phenotypic genome in a multivariate PERMANOVA analysis.
[0045] Figure 6 The study shows alpha diversity in the multi-kingdom microbiome as measured by the Shannon index in children with ASD and neurotypical children.
[0046] Figure 7 A volcano plot is shown, which illustrates the association between multi-kingdom species and ASD calculated by MaAsLin2 after adjusting for significant confounding factors.
[0047] Figure 8 This demonstrates the variation in microbiome function explained by the phenome in a multivariate PERMANOVA analysis.
[0048] Figure 9 A volcano plot is shown, which illustrates the association between microbiome function and ASD calculated by MaAsLin2 after adjusting for significant confounding factors.
[0049] Figure 10 A schematic diagram of the development of the random forest model is shown.
[0050] Figure 11 A box plot is shown, which illustrates the distribution of AUC scores obtained by random forest classification and evaluated on the test dataset.
[0051] Figure 12A graph showing the association between the top 27 markers assessed by MaAsLin2 and ASD is shown.
[0052] Figure 13 The differential abundance of 27 biomarkers between children with ASD and neurotypical children was shown.
[0053] Figure 14 The p-value and average accuracy of each of the 27 markers using the random forest model are shown to decrease.
[0054] Figure 15 The differences in the relative abundance of 27 biomarkers between children with ASD and neurotypical children in an independent hospital cohort, as assessed by a two-sided Wilcoxon rank-sum test, are shown.
[0055] Figure 16 The area under the curve for random forest models employing different features in the validation queue is shown.
[0056] Figure 17 The performance metrics details of a trained random forest model that classifies ASD using different features in the validation queue are shown.
[0057] Figure 18 Violin plots comparing microbial richness, Shannon diversity, and Pielou evenness at the species and MAG levels between the two groups are shown, with statistical significance determined by Welch's t-test.
[0058] Figure 19 A comparison of four different models (random forest classifier, logistic regression, support vector machine, and gradient boosting classifier) for detecting ASD using different taxonomic ranks of gut microbiota and microbial MAG is shown.
[0059] Figure 20 A framework for dataset partitioning, model training, and independent validation is shown.
[0060] Figure 21 The fecal microbiome-based model performance in the discovery cohort of a random forest model that uses 49 microbial biomarkers (5 microbial taxa and 44 microbial MAG biomarkers) to predict ASD events is shown.
[0061] Figure 22 A heatmap showing the correlation between host phenotype and identified microbial characteristics is presented, where the correlation is calculated by Spearman rank correlation analysis.
[0062] Figure 23The performance of the fecal microbiome-based model in the validation cohort of the best random forest model for predicting ASD events using 49 microbial biomarkers (5 microbial taxa and 44 microbial MAG biomarkers) is shown.
[0063] Figure 24 The distribution of predicted risk scores (ASD probabilities) for all samples in the validation cohort is shown, where the area under the right line represents the ASD probability distribution for children with ASD in the validation cohort, and the area under the left line represents the ASD probability distribution for healthy controls.
[0064] Figure 25 A heatmap showing the correlation between microbial biomarkers and different autism symptom scores is presented, where the correlations were calculated by Spearman rank correlation analysis.
[0065] Figure 26 The correlation between the single core symptom index and the SRS-T score and the predicted probability of ASD is shown (SRS-T: total T score; SRS-RRB: restricted interests and repetitive behaviors; SRS-SCI: social communication and interaction).
[0066] Figure 27 The AUROC of the best model tested in the four subgroups of the validation queue is shown.
[0067] Figure 28 The different stages of the research are shown.
[0068] Figure 29 This demonstrates the variation in the composition of multi-kingdom (archaea, bacteria, fungi, and viruses) microbiomes explained by phenotypic analysis in a multivariate PERMANOVA analysis.
[0069] Figure 30 The Shannon Index is shown to illustrate the alpha diversity in multi-kingdom (archaea, bacteria, fungi, and viruses) microbiomes as measured by the Shannon Index in children with ASD (n = 709) and NT children (n = 374).
[0070] Figure 31 A volcano plot is shown, which illustrates the associations between multi-kingdom (archaea, bacteria, fungi, and viruses) species and ASD calculated by MaAsLin2 after adjusting for significant confounding factors.
[0071] Figure 32 This demonstrates the variation in microbiome functions (pathways and genes) explained by the phenome in a multivariate PERMANOVA analysis.
[0072] Figure 33A volcano plot is shown, which illustrates the association between microbiome functions (pathways and KO genes) and ASD, calculated by MaAsLin2 after adjusting for significant confounding factors.
[0073] Figure 34 A schematic diagram of the development of the random forest model is shown.
[0074] Figure 35 A box plot is shown, which illustrates the distribution of AUC scores obtained by random forest classification and evaluated on the test dataset.
[0075] Figure 36 A graph showing the association between the top 31 markers assessed by MaAsLin2 and ASD is shown.
[0076] Figure 37 The differential abundance and p-values of 31 biomarkers between children with ASD and those with neurotypical ASD are shown.
[0077] Figure 38 The average accuracy decreases for each of the 31 markers used in the random forest model.
[0078] Figure 39 The AUC (95% CI) of the random forest model with different features in the validation queue is shown.
[0079] Figure 40 The performance metrics details of a trained random forest model that classifies ASD using different features in the validation queue are shown.
[0080] Figure 41 The association between ASD and 31 identified fecal microbiome markers was shown in five cohorts.
[0081] Figure 42 The AUC of the models tested using 31 biomarkers in independent cohorts of ADHD and atopic dermatitis is shown.
[0082] Figure 43 An example computer system including various hardware components is shown. Invention Details
[0084] Autism spectrum disorder (ASD) is a neurodevelopmental disorder of unknown etiology. Emerging evidence suggests a crucial role for the gut microbiota in the pathogenesis of ASD via the gut-brain axis and highlights the diagnostic potential of certain bacterial species for early risk prediction. Current risk prediction tests using the microbiome rely solely on bacterial data. By utilizing multi-kingdom microbiome data, the disclosed techniques can provide a cost-effective and accurate method to support clinical decision-making for ASD, thus contributing to improved ASD prevention and management. The inventors of this application have discovered a significant association between ASD and microbial taxa and metagenomically assembled genomes (MAGs). In some embodiments, associations have been found between ASD and multi-kingdom microbial biomarkers, including archaea, bacteria, fungi, viruses, genes, and microbial pathways. In this regard, the present invention provides a novel method for determining the presence or risk of ASD in a subject by detecting a set of microbial biomarkers (e.g., including microbial taxa and MAGs) in a fecal sample and analyzing them using a random forest model.
[0085] As used herein, the terms “microbial biomarker,” “microbial biomarker,” or “microbial marker” refer to a collection of multi-kingdom microbial species (including archaea, bacteria, fungi, and viruses), functional pathways, functional genes, microbial taxa, and MAGs. The relative abundance of deoxyribonucleic acid (DNA) fragments of each archaea, bacterium, fungus, and virus species in a fecal sample from each subject can be measured. The relative abundance of functional pathways and functional genes can be inferred from the relative abundance of the aforementioned DNA fragments, for example, by annotation using the Kyoto Encyclopedia of Genetics and Genomes (KEGG) Orthology (KO). The term “microbial taxa” refers to an operational taxonomic unit at various taxonomic levels (e.g., genus, species, and strain). Higher-level taxa cover all members belonging to that taxa; for example, “Lactococcus genus (…)…” Lactococcus ")" refers to all species belonging to the genus *Lactococcus*. The term "metagenomic assembly genome" refers to a partial or complete microbial genome. A MAG is constructed by first assembling sequencing reads into contigs, and then classifying these contigs into a MAG based on covariations in nucleotide frequencies, abundance, and / or abundance in a set of samples representing a set of sequences from genome assemblies with similar characteristics.
[0086] As used herein, the term "reference dataset" refers to a dataset generated from a cohort of subjects with a definitive diagnosis of ASD and another cohort of subjects who have not yet been diagnosed with ASD and show no signs of ASD. It can be generated by measuring the relative abundance of each of the "microbial biomarkers" for each subject in the cohort and then aggregating such data generated from each subject. To properly construct a "reference dataset," a sufficient number of individuals should be included (e.g., at least 10, 12, 15, 20, 24, 50, 100, or more individuals). The diagnosis of subjects in the cohort can be determined using standard methods. For example, it can be made by a post-specialist qualified child psychiatrist based on the Diagnostic and Statistical Manual of Mental Disorders, Fifth Edition (DSM-5).
[0087] As used in this article, the term "training dataset" refers to a subset of the "reference dataset." It includes the known health classifications (ASD or NT) of the objects and the feature vectors of the objects, which can be used to train machine learning models.
[0088] This invention provides a set of microbial biomarkers for determining the presence of ASD in an object or assessing the risk of ASD in an object. The set of microbial biomarkers may include (1) all microbial biomarkers listed in Tables 1, 2, 3 and / or 4, (2) all functional genes detected in a reference dataset, (3) all microbial functional pathways detected in a reference dataset, and / or (4) combinations of all microbial functional pathways, bacteria, functional genes, archaea, fungi and viruses detected in a reference dataset, wherein the mean abundance and prevalence of (2), (3) and (4) are greater than or equal to 0.15% and 5%, respectively.
[0089] Microbial biomarkers identified as ASD-related in this study can be detected and measured in fecal samples taken from subjects using known methods to assess the presence of ASD or the risk of developing ASD in the subjects. For example, various methods for determining the level of bacterial species in samples have been reported in the literature, such as amplification (e.g., by polymerase chain reaction) and sequencing of bacterial polynucleotide sequences using sequence similarity in commonly shared 16S rRNA bacterial sequences. Nucleic acid-based analytical methods can generally be used for the detection and quantification of any genetic biomarker. Furthermore, metagenomic next-generation sequencing is a recently developed technology for rapidly determining the presence and level of microbial species such as archaea, bacteria, fungi, and viruses in samples using the unique genomic sequence of each species. For biopathway biomarkers, analytical methods can be found in the literature and described herein for assessing their presence and level, for example, by analyzing relevant metabolites or byproducts of such pathways.
[0090] Table 1 below provides an example list of biomarkers (species, genes, and pathways) used to determine the presence or risk of ASD in an object, ordered by importance. (Regarding...) Figures 4 to 17 During the study (“the first study”), biomarkers for the microbial origins listed in Table 1 were identified.
[0091] Table 1
[0092]
[0093] Tables 2 and 3 below provide an example list of microbial biomarkers used to determine the presence or risk of ASD, with Table 2 providing an example list of microbial taxa and Table 3 providing an example list of microbial genetic sequences. The microbial biomarkers in Tables 2 and 3 are ordered by importance for determining the presence or risk of ASD in an individual. (Regarding...) Figures 18 to 27 The microbial biomarkers listed in Tables 2 and 3 were identified during the study (“Second Study”).
[0094] Table 2
[0095]
[0096] Table 3
[0097]
[0098]
[0099] Table 4 below provides an example list of biomarkers (species, genes, and pathways) used to determine the presence or risk of ASD in an object, ordered by importance. (Regarding...) Figures 28 to 42 During the study (“Third Study”), biomarkers for the microbial origins listed in Table 4 were identified.
[0100] Table 4
[0101]
[0102] Some implementations provide methods for determining the presence or risk of ASD in an object, the methods including generating a machine learning classifier trained on a random forest model. In some implementations, the object is a male object. In some implementations, the object is a female object. In some implementations, the object can be either male or female. In some implementations, the object is a child under the age of 18, such as a child under 14, 13, 12, 11, 10, 9, 8, 7, 6, or 5 years old, for example, a child aged 1 to 15, 2 to 14, 3 to 13, 4 to 12, or 5 to 11 years old, for example, a child aged 4 to 10, 5 to 10, or 6 to 10 years old. In some implementations, the object is of Asian descent, such as Chinese. In some implementations, the object can be of any race or ethnicity.
[0103] A classifier can be a machine learning model that outputs a probability or score for a given entity in a category. The probability can be a value from 0.0 to 1.0, and a value called a discrimination cutoff (or threshold) can be the boundary between categories. For example, if the discrimination cutoff is 0.6, entities with probabilities between 0.0 and 0.6 can be classified into the first category (e.g., NT), and entities with probabilities between 0.6 and 1.0 can be classified into the second category (e.g., ASD). In some cases, when a higher level of accuracy is required, uncertain categories can be introduced by using two discrimination cutoffs as boundaries between categories. The probability falling between the two cutoffs is called the median confidence interval and represents the uncertain category. Probabilities below the lower cutoff and above the higher cutoff are called the high confidence interval.
[0104] In one instance, if the two cutoff values are 0.389 and 0.670, entities with probabilities between 0.0 and 0.389 can be classified into Class I (e.g., NT), and entities with probabilities between 0.670 and 1 can be classified into Class II (e.g., ASD). If an entity's probability is between 0.389 and 0.670, it can be classified into an indeterminate category. The optimal cutoff values for cases and controls in the high confidence interval can be determined by the maximum proportions of the true positive / false positive ratio and the true negative / false negative ratio, respectively. In some instances, the optimal cutoff value for the median confidence interval can be determined using ROC analysis that maximizes the Youden index.
[0105] Random forests are supervised learning methods used in machine learning for classification and regression. They average the results of many decision trees applied to different subsets of a dataset to improve the projection accuracy of the dataset. Nested cross-validation can be applied to compute in-cohort accuracy by randomly splitting feature vectors into training and test datasets. The training dataset can be randomly split into k subsets. These subsets can be used to create k folds, where k-1 subsets are used as training subsets and the remaining subsets are used as validation subsets. The training subsets can be used to train the model and the validation dataset to evaluate the model's performance during training. The test dataset, which may not be used for training the model, can be used to evaluate the trained model.
[0106] In some instances, the validation process can be 20 (or 10) iterations of k-fold hierarchical cross-validation (balanced class proportions across folds), where k = 5. For each of the 5 splits in these instances, a random forest model can be trained on a training subset, which can then be used to predict classifications on the validation dataset. The model's performance can be evaluated using appropriate performance metrics such as accuracy, precision, recall, area under the curve (AUC), etc. In some instances, after all iterations of cross-validation are completed, the performance metrics obtained at each fold can be aggregated. This can be done by calculating the mean, median, or any other appropriate aggregation method. This process can be repeated 20 (or 10) times while tuning the model's hyperparameters. The cross-validation process is repeated with the tuned hyperparameters of the model, and the performance of different combinations of hyperparameter values is evaluated. This allows identification of the set of optimal hyperparameters that produce the best performance.
[0107] After this process has been repeated 20 (or 10) times, candidate random forest models can be evaluated on test data. In some instances, a λ hyperparameter can be chosen for each model to maximize the AUC of the receiver operating curve (ROC) under the constraint that the model contains at least five non-zero coefficients. In some instances, the λ hyperparameters of the model can be varied until the maximum AUC-ROC of the model is found, and those hyperparameters are chosen as λ hyperparameters.
[0108] In some instances, the training and tuning process is incremental, involving k-fold cross-validation followed by hyperparameter tuning. For example, in a five-fold hierarchical cross-validation (balanced class proportions across folds) where k = 5, a random forest model can be trained on a training subset for each of the five splits in this instance, and then the random forest model can be used to predict classifications on the validation dataset. The model's performance can be evaluated using appropriate performance metrics (e.g., accuracy, precision, recall, AUC, etc.). In some instances, after all iterations of cross-validation are completed, the fold with the best performance metric (e.g., maximum AUC) is selected and used to tune the model's hyperparameters by evaluating the model's performance metric on the test data. λ hyperparameters can be chosen for each model to maximize the AUC of the receiver operating curve (ROC) under the constraint that the model contains at least five non-zero coefficients. In some instances, the model's hyperparameters can be varied until the maximum AUC-ROC of the model is found, and those hyperparameters are selected as λ hyperparameters.
[0109] In some instances, after selecting the λ hyperparameter, candidate random forest models can be evaluated by resplitting the entire reference dataset into a 70 / 30% training-test sample split. This process can be repeated 20 (or 10) times to obtain 20 (or 10) performance metrics for each candidate random forest model. After completing all iterations of 20 (or 10) repetitions, the performance metrics obtained in each repetition can be aggregated. This can be done by calculating the mean, median, or any other suitable aggregation method. This allows for the identification of the best candidate random forest model that produces optimal performance.
[0110] The area under the receiver operating characteristic curve (AUC) can be used to compare the performance of models with different methods and features. AUC is a metric that considers the trade-off between sensitivity and specificity across all possible thresholds for comparing the performance of various classifiers. For random classifiers, a baseline AUC value can be 0.5.
[0111] The area under the precision-recall curve (AUPR) can be provided as a complementary assessment, taking into account the trade-off between precision (or positive predictive value) and recall (or sensitivity), where the baseline equals the proportion of positive disease cases in all samples. Precision is the score of relevant instances among retrieved instances, and recall is the total proportion of relevant instances retrieved. The precision-recall curve is a graph of precision (x-axis) vs. recall (y-axis). An AUPR value of 1.0 indicates high precision and recall, while a value of 0.0 indicates low precision and recall.
[0112] Figure 1A and Figure 1BExample steps of method 100 for training a random forest classifier to predict the presence of ASD in objects are shown. One or more steps of method 100 may be omitted during execution of method 100, and the steps of method 100 may be executed in any order and / or in parallel. One or more steps of method 100 may be executed by one or more processors. Method 100 may be implemented as a computer-readable medium or computer program product containing instructions that, when executed by one or more computers, cause the one or more computers to perform one or more steps of method 100.
[0113] Step 102 includes obtaining a reference dataset from a first queue of objects diagnosed with ASD and a second queue of objects not diagnosed with ASD and not showing any signs of ASD. Step 102 may include steps 104 and 106.
[0114] Step 104 includes identifying a set of microbial biomarkers. The set of microbial biomarkers may include one or more biomarkers of microbial origin from Tables 1, 2, 3 and / or 4.
[0115] Step 106 includes determining, for each individual in the first cohort and the second cohort, the relative abundance of each of the biomarkers for the microbial origin identified in step 104.
[0116] Step 108 involves splitting the reference dataset into a training dataset and a test dataset. The training dataset is used to train the model, and the performance of the trained model can be evaluated using the reserved test dataset.
[0117] Step 110 involves dividing the training dataset into k subsets. The k subsets may be of equal or different sizes. For example, each of the k subsets may include approximately equal numbers of abundance measurements of objects. In some instances, objects may be randomly placed into each of the k subsets such that each subset may include approximately equal distributions of objects diagnosed with ASD and objects not diagnosed with ASD.
[0118] Step 112 includes training k candidate random forest models for each of the k subsets by the following steps: in step 114, generating a validation dataset including one of the k subsets and a training subset including the remaining k-1 subsets of the k subsets; in step 116, training a candidate random forest model among the k candidate random forest models using the training subset; and in step 118, evaluating the candidate random forest models using the validation dataset by computing a performance metric for the candidate random forest models.
[0119] In some instances, training a candidate random forest model using training subsets involves generating n training subsets by sampling (i) objects from a first cohort of objects and a second cohort of objects, and (ii) microbial biomarkers from a set of microbial biomarkers, and constructing n decision trees using the n training subsets. The n decision trees can form the candidate random forest model. In some instances, objects from the first cohort of objects and the second cohort of objects are sampled with replacement. In some instances, microbial biomarkers from the set of microbial biomarkers are sampled with replacement. For example, a first portion of a particular group within the n groups may include abundance measurements of a first object against the first and second microbial biomarkers, a second portion of the particular group may include abundance measurements of a second object against the first and second microbial biomarkers, and a third portion of the particular group may again include abundance measurements of the first object against the first and second microbial biomarkers.
[0120] Each of n decision trees can be constructed (trained) using training data from the corresponding groups, the training data including abundance measurements and ASD indicators (e.g., 0 or 1, NT or ASD, etc.), the ASD indicators indicating whether an object is diagnosed with ASD. In some instances, the training data can be prepared as a set of feature vectors (abundance measurements) and their corresponding target labels (ASD indicators). During the training process, the decision tree algorithm recursively partitions the data into subsets based on the abundance measurements of the microbial biomarkers to minimize impurities or errors in each partition. This process continues until a stopping criterion is met (e.g., reaching the maximum tree depth, or the number of data points in a leaf node falling below a certain threshold, etc.).
[0121] Each non-leaf node in the trained decision tree examines the input abundance measure of a specific microbial biomarker and, based on whether the input abundance measure of that biomarker is less than or greater than a threshold associated with that node, allows the decision tree to traverse further along a first or second path. Each leaf node in the trained decision tree corresponds to one of the target labels (ASD index, i.e., 0 or 1). Once all decision trees have been trained, the random forest model combines (e.g., aggregates) their predictions to make a final decision. In some instances, all predictions are averaged, and the average is compared to a threshold to produce the final output. In other instances, each tree “votes” for the class (0 or 1) of its prediction, and the class with the most votes becomes the final prediction.
[0122] In some instances, each of the n groups may correspond to a specific microbial taxonomy, MAG, biological entity, or kingdom. For example, the first group of n groups may include abundance measurements for multiple objects that are microbial taxonomy markers or one or more microbial markers; the second group may include per million transcript measurements for multiple objects that are microbial genetic sequences or one or more microbial markers; the third group may include abundance measurements for multiple objects that are bacteria or one or more microbial markers; the fourth group may include abundance measurements for multiple objects that are fungi or one or more microbial markers; and the fifth group may include abundance measurements for multiple objects that are viruses or one or more microbial markers. According to this example, the first decision tree of the n decision trees may be specific to a microbial taxonomy, the second decision tree may be specific to a microbial genetic sequence, the third decision tree may be specific to bacteria, the fourth decision tree may be specific to fungi, and the fifth decision tree may be specific to viruses. In another instance, a specific group in the n groups may include abundance measurements of multiple objects of microbial markers for various types of biological entities or kingdoms (e.g., taxa, MAGs, bacteria, fungi, viruses, etc.), and the corresponding decision tree can be constructed using abundance or per million transcripts measurements of various types of biological entities.
[0123] In some instances, performance metrics include one or more of the following: total number of false positives, total number of true positives, total number of false negatives, and total number of true negatives. For example, after training a candidate random forest model using training data from k-1 subsets, validation data for each object from the remaining subsets is fed into n decision trees of the candidate random forest model to generate a set of outputs, which are aggregated (e.g., averaged) and compared with a discrimination cutoff to produce a final prediction for each object. The prediction for each object is compared to the corresponding ASD metric, and the prediction is considered a true positive (the candidate model predicts a positive and is correct), a true negative (the candidate model predicts a negative and is correct), a false positive (the candidate model predicts a positive but is incorrect), or a false negative (the candidate model predicts a negative but is incorrect).
[0124] Step 120 involves evaluating the k candidate random forest models by calculating the AUC value for each of the k candidate random forest models. In some instances, the performance metric calculated in step 118 can be used to plot an ROC curve by changing the discriminant value from 0 to 1, while tracking the total number of true positives and false positives at each discriminant value. The points representing the true positive rate (on the y-axis) versus the false positive rate (on the x-axis) form the ROC curve. Once the ROC curve is generated, the AUC is calculated by finding the area under the ROC curve. This is achieved by numerically integrating the curve using various methods such as the trapezoidal rule. In various instances, a validation dataset and / or a test dataset can be used for step 120.
[0125] Step 122 involves comparing the AUC values of each of the k candidate random forest models to identify the one with the largest AUC value. The candidate model with the largest AUC can be selected.
[0126] Step 124 includes using data from the test dataset to validate the candidate model with the highest AUC to obtain the model's AUC (e.g., the probability that the model can correctly distinguish between randomly selected correctly classified samples and incorrectly classified samples).
[0127] Step 126 involves tuning the model if the AUC from step 124 is unsatisfactory, either by (a) increasing the number of objects in the training dataset or (b) changing the model's hyperparameters. In some instances, one or more hyperparameters (e.g., n_estimators, mtry, ntree, nodesize, maxnodes) for each model candidate can be tuned until the maximum AUC-ROC for that model is found. In one instance, the parameter n_estimators can be confidently set to n_estimators = 2000. If the AUC value is below a performance threshold, it may be unsatisfactory.
[0128] Step 128 involves repeating steps 102 through 126 until the AUC from step 124 is satisfactory and a final model is obtained. Whether an AUC is acceptable can vary depending on the intended use. Lower AUC scores may indicate lower confidence levels, and higher AUC scores may indicate higher confidence levels. Acceptable performance thresholds may be AUC values greater than or equal to 0.5, 0.6, 0.7, 0.8, 0.9, or 1.0. The final model may include a decision tree using a subset of the set of microbial-derived biomarkers identified in step 104. This subset of microbial-derived biomarkers may be collectively referred to as a biomarker group. In some cases, step 128 includes identifying a biomarker group comprising a subset of the set of microbial-derived biomarkers used in the decision tree of the final model.
[0129] Step 130 includes calculating a provisional cutoff value for the final model of ASD prediction based on the Youden index. The Youden index can integrate sensitivity and specificity information, and by using Youden index analysis, an optimal cutoff value can be found, for example, providing a value that provides the best trade-off between sensitivity and specificity. The Youden index can be calculated using the following formula:
[0130]
[0131] Step 132 involves deploying the final model as a trained random forest model based on one of the k candidate random forest models with the largest AUC value exceeding the performance threshold. In some instances, deploying the trained random forest model may include loading model-related weights onto a dedicated hardware accelerator or other computer hardware.
[0132] Figure 2 Example steps of a method for training a random forest classifier to predict the presence of ASD in objects are shown. Figure 2 The example steps can be used in conjunction with Method 100. In some instances, when building the model, the different sets of hyperparameters are not compared with k-fold validation. Instead, the best model is selected from k-fold cross-validation. Then, the different sets of hyperparameters are applied to the best model using the entire training set, and the AUC is evaluated on 30% of the test dataset.
[0133] Figure 3 Example steps of method 300, which applies a trained random forest model to predict the presence of ASD in an object, are shown. Step 302 includes determining the relative abundance of biomarkers (i.e., a group of biomarkers) of the same microbial origin used in the decision tree of the final random forest model from a fecal sample of an object to be determined to have or be at risk of ASD. Biomarkers may include (1) biomarkers of microbial origin listed in Tables 1, 2, 3 and / or 4; (2) all functional genes detected in the reference dataset; (3) all microbial functional pathways detected in the reference dataset; and / or (4) combinations of microbial functional pathways, bacteria, functional genes, archaea, fungi and viruses detected in the reference dataset, wherein the mean abundance and prevalence of each biomarker in (2), (3) and (4) detected in the reference dataset are greater than or equal to 0.15% and 5%, respectively.
[0134] Step 304 includes comparing the relative abundance of the biomarkers of microbial origin obtained from the object according to step 302 with a threshold in a decision tree obtained from the final random forest classifier of method 100, wherein the decision tree is generated by training a random forest from reference data, and wherein the relative abundance of each of the biomarkers of microbial origin obtained from the object in step 302 runs down the decision tree to generate a risk score (or probability).
[0135] Step 306 includes determining that the subject has an increased risk of ASD when the risk score is greater than the discrimination cutoff, or determining that the subject does not have an increased risk of ASD when the risk score is not greater than the discrimination cutoff. Optionally, the risk score may be interpreted as a continuous scale, wherein a higher score indicates a higher risk of the subject having ASD.
[0136] Various techniques, such as shotgun metagenomic sequencing, can be used to characterize the microbial taxa, MAGs, species, genes, and functional pathways of a sample. In shotgun metagenomic sequencing, DNA is obtained from a heterologous sample of cells and fragmented into DNA fragments that can be compared with the genomes of various microbial taxa, MAGs, species, genes, and functional pathways to identify the microbial taxa, MAGs, species, genes, and functional pathways in the sample.
[0137] The reference dataset can be a collection of metagenomic data obtained by sequencing fecal samples collected from one cohort of individuals diagnosed with ASD and another cohort of individuals who were never diagnosed with ASD and showed no signs of ASD (non-ASD or NT). In various implementations, the number of individuals can be 10, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 1250, 1500, 1750, 2000, 2500, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 15000, 20000, 50000, 100000, 500000, or 1,000,000 individuals. In some implementations, comorbidities such as attention deficit hyperactivity disorder (ADHD) are used to stratify the reference dataset into subgroups to train subgroup-specific machine learning classifiers. In some implementations, confounding factors such as gastrointestinal (GI) symptoms, diet, and age are used to pair equal numbers of ASD and non-ASD subjects. In various instances, the ASD and non-ASD subjects are all male, all female, or include both male and female subjects. In some implementations, the ASD and non-ASD subjects do not have intellectual disability, neurological disorders, psychosis, depressive disorders, or other serious medical conditions. In some implementations, it is recommended that ASD and non-ASD subjects discontinue the use of probiotics, antidepressants, antiepileptic drugs, and antibiotic exposure for one month prior to stool sample collection. In some implementations, subjects participating in dietary intervention programs or with specific dietary habits (i.e., low-fat diets, metabolic diets, vegetarian diets, and food allergies) are excluded from the reference dataset.
[0138] In some instances, the same standardized protocol was used to generate sequencing data in metagenomic datasets, for example, steps including fecal DNA extraction, sequencing, raw data quality filtering, host read decontamination, and microbiome annotation were standardized. Metagenomic data in the reference dataset can be obtained by: (1) prospectively recruiting a cohort of subjects diagnosed with ASD and another cohort of subjects not diagnosed with ASD and not showing any signs of ASD and sequencing their fecal samples, or (2) downloading from one or more public databases, or (3) a combination of (1) and (2).
[0139] Stool samples are obtained from the individuals to be tested. Stool samples can be easily collected from individuals in clinics or at their homes. An appropriate amount of stool is collected and preserved according to standard procedures before further preparation. In some embodiments, the stool samples are frozen at -80°C. Optionally, the stool samples are frozen with a preservative to enhance the stability of nucleotides in the stool samples. The stool samples can then be used to generate metagenomic data by using the same standardized protocol used to generate the reference dataset, including steps such as stool DNA extraction, sequencing, raw data quality filtering, host read decontamination, and microbiome annotation.
[0140] Figures 4 to 17 This study presents the results of the first multi-border analysis exploring gut archaea, bacteria, fungi, viruses, and their genes and functions. This article presents a metagenomic analysis of 1017 phenotypically well-developed and neurotypical children with ASD, and validation of these findings in a public dataset of 195 fecal metagenomics. Twenty-seven fecal microbiome biomarkers were identified among 1,188 significant multi-border associations in the gut microbiome of children with ASD, demonstrating good and reproducible diagnostic performance for ASD across different age groups and cohorts.
[0141] exist Figures 4 to 17 In the example shown, the random forest model is first trained on a randomly selected training dataset (70%, 5-fold cross-validation) and then applied to a reserved test dataset (30%) to obtain performance. Hyperparameters (e.g., mtry, ntree, nodesize, maxnodes) are then tuned based on the model performance on the test dataset to avoid overfitting. Finally, using the optimal combination of hyperparameters, the queue is randomly split 20 times to obtain the distribution of the random forest prediction evaluations on the test dataset, and the average AUC is calculated accordingly for visualization of the results. The highest-ranking and most frequently selected microbial features are considered as predictive features for further annotation. The predictive performance of each feature is retrieved using the same training dataset. The trained models are then tested in independent validation queues to evaluate their robustness.
[0142] Figures 4 to 7 This study demonstrates the association between ASD and the composition of the fecal microbiome across multiple kingdoms. Figure 4The different phases of the first study are illustrated, including discovery phase 402, machine learning phase 404, deployment phase 406, and validation phase 408. The study included 1017 children (aged 1–13 years, 84.3% male) from six independent cohorts. Extensive phenotypic data (n=227) were collected, including age, sex, body mass index (BMI), diet, medication use, comorbidities, associated mental disorders, GI symptoms (including fecal consistency), and family characteristics. All fecal samples were processed using the same standardized protocol. For each metagenomic, a total of 6.24 terabytes (TB) of sequence data was obtained at an average depth of 6.28 gigabytes (GB). In the discovery cohort, metagenomic sequencing was performed on fecal samples from 331 children with ASD and 116 neurotypical children (aged 5–11 years, all male).
[0143] The dataset was processed using Kraken2 and HUMAnN3 and filtered for 15% prevalence. The dataset contains 8,915 species-level taxa (382 archaea, 8,145 bacteria, 80 fungi, and 308 viruses), 499 functional pathways, and 5,128 KO microbial genes. A total of 172 stool samples from 82 children with ASD and 90 neurotypical children (aged 4–11 years, all male) in an independent hospital cohort were sequenced, and these data were used for external validation. A community cohort of younger children (84 children with ASD and 40 neurotypical children (aged 1–8 years, all male) was included to validate findings across different age groups. To test potential sex-related differences in gut microbiome biomarkers in ASD, an independent female cohort of 39 children with ASD and 39 neurotypical children (aged 4–12 years, all female) was included. Two additional cohorts of non-ASD children with attention deficit hyperactivity disorder (ADHD) (n = 118) and atopic dermatitis (n = 78) were used to evaluate the specificity of these findings. Furthermore, 195 fecal metagenomic samples (aged 2–13 years, all male) from previously published studies were also included in the analysis as external validation.
[0144] Since the composition of the gut microbiota is primarily influenced by environmental and host factors, the effects of 227 host factors on gut microbiome composition were analyzed. These host factors included age, ASD, body mass index (BMI), comorbidities (n = 7), diet (n = 201), family characteristics (n = 5), GI parameters (n = 4), medication use (n = 4), and mental disorders (n = 3). In the discovery cohort (aged 5–11 years, n = 447, all male), these host factors combined explained 13.0%, 15.8%, 8.86%, and 10.5% of the inter-individual microbiome variation in archaea, bacteria, fungi, and viruses, respectively.
[0145] Figure 5 The variation in multi-kingdom (archaea, bacteria, fungi, and viruses) microbiome composition explained by phenotype is shown in a multivariate PERMANOVA analysis. The presence of GI symptoms and a diagnosis of ASD were the top two factors contributing to the greatest variation in microbiome composition. A total of 26 factors showed significant effects on all four kingdoms of gut microbiome composition across all host and dietary factors studied, and these factors were adjusted for in all subsequent association analyses. Next, changes in gut microbiome diversity in children with ASD and neurotypical children were assessed.
[0146] Figure 6 The alpha diversity in the multi-kingdom (archaea, bacteria, fungi, and viruses) microbiome, as measured by Shannon indices, is shown in children with ASD (n = 331) and neurotypical (NT) children (n = 116). p-values were calculated using MMUPHin (two-tailed test). Data are shown via interquartile range (IQR), with the median at the horizontal line and extension to the most extreme points within 1.5 × IQR. Outliers are represented as points. Children with ASD showed reduced Shannon diversity indices for archaea (p = 0.0081), bacteria (p = 0.034), and viruses (p = 0.0091) compared to neurotypical children. The abundance of archaea, bacteria, and viruses was also significantly lower in ASD children than in neurotypical children. However, there was no significant difference in fungal diversity (p = 0.77) or abundance (p = 0.45) between ASD and neurotypical children.
[0147] Figure 7 A volcano plot is shown, illustrating associations between multi-kingdom (archaea, bacteria, fungi, and viruses) species and ASD, calculated by MaAsLin2 after adjusting for significant confounding factors. Associations with an adjusted p-value (FDR) less than 0.05 are considered significant and are marked with solid circles. Figure 7ASD-enriched associations were found in the upper right quadrant of each plot, while ASD-reduced associations were found in the upper left quadrant of each plot. Top-ranked species are labeled.
[0148] To identify microbial biomarkers for ASD, the association between multi-kingdom species and ASD was examined. A total of 19 archaea, 126 bacteria, 6 fungi, and 24 viruses showed varying abundances between children with ASD and neurotypical children (FDR < 0.05). Compared to neurotypical children, the relative abundance of 161 of the 175 identified microbial species was significantly reduced in the gut of children with ASD. This finding was most significant for bacterial communities, with 122 bacterial species reduced in the gut of children with ASD, while only four bacterial species were enriched in ASD. The altered bacterial species in children with ASD were caused by *Westernella fusionis*, *Westernella sinensis*, and other short-chain fatty acid-producing bacteria (e.g., *Bacteroides pHL2737*). Lawsonibacter asaccharolyticus Enterobacter faecalis ) This is driven by a reduction in [something]. Furthermore, several opportunistic pathogens, such as Pseudomonas BT422 and Oligotrophoblast G4, are enriched in the gut of children with ASD.
[0149] Figure 8 and Figure 9 This study demonstrated the association between ASD and fecal microbiome function. Figure 8 This study illustrates the changes in microbiome function (pathways and genes) explained by the phenotype in a multivariate PERMANOVA analysis. The influence of host phenotypic factors on microbial function was explored based on Kyoto Encyclopedia of Genes and Genomes (KEGG) Orthology (KO genes, n = 5,128) and metabolic pathways (n = 499). Figure 8 As shown, host phenotypic factors explained 17.7% and 14.1% of the variations in microbiome functional pathways and KO genes, respectively. A diagnosis of ASD was listed as the primary factor, accounting for 4.1% and 2.8% of the variations in microbiome functional pathways and KO genes, respectively. A total of 20 host and dietary factors showed significant effects on gut microbiome function. After adjusting for these 20 host factors, 909 differentially expressed microbiome genes were identified: 849 genes were reduced and 60 genes were increased in children with ASD compared to neurotypical children.
[0150] Figure 9 A volcano plot is shown, illustrating the associations between microbiome functions (pathways and KO genes) and ASD, calculated by MaAsLin2 after adjusting for significant confounding factors. Associations with an adjusted p-value (FDR) less than 0.05 were considered significant and are marked with solid circles. Figure 9ASD-enriched associations were found in the upper right quadrant of each graph, and ASD-reduced associations were found in the upper left quadrant of each graph. The top-ranked pathways are labeled with their names.
[0151] At the pathway level, among 104 differentially expressed pathways (FDR < 0.05), 88 pathways showed a negative correlation with ASD, and 16 pathways showed a positive correlation. Compared to neurotypical children, several functional pathways involved in amino acid biosynthesis, particularly L-isoleucine, were enriched in children with ASD, while several fatty acid metabolic pathways, such as the oleic acid β-oxidation pathway and the fatty acid salvage pathway, showed a negative correlation with ASD. Thiamine diphosphate biosynthesis was found to be reduced in children with ASD compared to neurotypical children. Impaired thiamine diphosphate synthesis has been associated with ASD and other mental disorders in animal and human studies. A negative correlation was observed between ASD and the 4-aminobutyric acid (GABA) bypass pathway. GABA is a major inhibitory neurotransmitter in the mammalian central nervous system and has been associated with ASD in previous studies.
[0152] Figures 10 to 14 This demonstrates how a random forest classifier trained on multi-kingdom microbiome composition and function data can predict the absence of ASD. Figure 10 A schematic diagram of the random forest model development is shown. A balanced reference dataset 1016 was constructed from the original discovery cohort (n = 447) to compare ASD (n = 116) children with NT (n = 116) using a one-to-one pairing algorithm that considers all significant confounding factors (n = 26). In some instances, the reference dataset 1016 includes feature vectors of six components: archaea, bacteria, fungi, viruses, microbial pathways, and KO genes.
[0153] A 70 / 30% training-test sample split is performed on the reference dataset 1016 to generate a training dataset 1018 and a test dataset 1020. The training dataset 1018 is used to train an ensemble of decision trees 1022, which is then aggregated to form a random forest model 1024. In the illustrated example, each of the decision trees 1022 corresponds to a specific biological entity or kingdom. In various instances, the random forest model 1024 may be formed using decision trees 1022 representing a single kingdom or biological entity (e.g., a bacteria-only decision tree) or decision trees 1022 representing multiple kingdoms or biological entities (e.g., bacteria, fungi, and viruses). After training the random forest model 1024, the accuracy of the random forest model 1024 is estimated by calculating the model's AUC using the test dataset 1020. This training and testing process is repeated 20 times to obtain the distribution of the random forest prediction evaluation on the test dataset 1020.
[0154] Previously, the performance of archaea, fungi, viruses, KO genes, or functional pathways in ASD had not been explored. To avoid discrimination bias due to sample size imbalance and residual confounding factors, a matching sub-cohort was constructed based on the discovery cohort (n = 447) using a one-to-one pairing algorithm to select children with ASD (n = 116) and neurotypical children. A total of 26 confounding factors identified above were used in the pairing algorithm, including age, BMI, BSFS, GI symptoms (n = 2), and dietary factors (n = 21). After matching, there were no differences in these 26 confounding factors between children with ASD and neurotypical children in the matched sub-cohorts, which were then used for the development of a random forest model.
[0155] Figure 11 Box plots are shown, illustrating the distribution of AUC scores obtained from random forest classification on the test dataset. Differences between evaluation groups are analyzed using a two-sided Wilcoxon rank-sum test. Figure 11 In the diagram, an asterisk indicates that the FDR is less than 0.05. The prevalence and relative abundance of the top 27 biomarkers used by the random forest model are also shown in the balanced datasets of children with ASD (n = 116) and neurotypical children (n = 116).
[0156] Single-boundary biomarkers were used to test the accuracy of models in distinguishing children with ASD from neurotypical children. Among all single-boundary biomarkers, the microbial pathway model demonstrated the strongest predictive ability for ASD detection, with an average score of 0.85 AUC at 81.5% specificity and 84.9% sensitivity. Models using bacterial features had the second highest scores, with an average AUC of 0.81, followed by KO gene-based models (average AUC 0.79), archaea-based models (average AUC 0.69), fungal-based models (average AUC 0.67), and virus-based models (average AUC 0.61). Overall, these findings suggest that biomarker features from different boundaries offer promising predictive capabilities for the diagnosis of ASD.
[0157] Since all single-border features have demonstrated diagnostic potential for children with ASD, the performance of a model combining individual multi-border features was investigated. Data from all four borders (including their functional and genetic aspects) were combined. The combined model showed superior performance in diagnosing ASD with 88.6% specificity and 85.8% sensitivity (mean AUC 0.88) compared to the single-border feature-based model (all adjusted p < 0.05). Furthermore, the model based on the combined multi-border features achieved an accuracy of 87.1%, with a positive predictive value of 96.3% and a negative predictive value of 86.8%. These results confirm that the multi-border fecal microbiome biomarker group has higher diagnostic performance for ASD than the single-border group.
[0158] Figure 12 A graph showing the association between the top 27 biomarkers assessed by MaAsLin2 and ASD is presented. To identify the minimum number of microbiome biomarkers that achieve the highest accuracy, identified biomarkers were sequentially included in the model according to their ranking. Finally, for the diagnosis of ASD, a total of 27 microbiome features (2 archaea, 2 fungi, 2 viruses, 8 bacteria, 6 KO genes, and 7 pathways) showed an AUC of 0.88, with a sensitivity of 89.7%, a specificity of 90.5%, and an accuracy of 90.1%. The prevalence and relative abundance of biomarkers also differed significantly between children with ASD and neurotypical children.
[0159] Figure 13 The differential abundance of 27 biomarkers between children with ASD and neurotypical children was shown. It was observed that 19 biomarkers were significantly reduced and 8 biomarkers were significantly enriched in the gut of children with ASD (all FDR < 0.01).
[0160] Figure 14 The p-value and average accuracy reduction for each of the 27 markers used in the random forest model are shown (where... means FDR < 0.05, means FDR < 0.01, This means FDR < 0.001. The importance of these 27 features was re-analyzed in the final integrated model, and it was observed that the model's accuracy was primarily driven by the thiamine diphosphate biosynthesis pathway, supporting a potential role for thiamine diphosphate in the pathogenesis of ASD. Furthermore, the reduction of several bacteria was also among the top-ranking microbial features contributing to diagnostic accuracy, including *Westernella fusionis*, Lawsonibacter asaccharolyticus*Westernella esculenta* and *Bacteroides* PHL2737 were also identified. The palmitoleic acid biosynthesis pathway I was also found to be among the top 27 features. Palmitoleic acid has previously been associated with constipation; therefore, this enriched biosynthetic pathway may explain the prevalence of constipation in children with ASD. Overall, the analysis showed that 27 fecal microbiome group markers derived from archaea, bacteria, fungi, viruses, KO genes, and microbial functional pathways represent a potentially promising non-invasive tool for the diagnosis of ASD.
[0161] Figures 15 to 17 The random forest model is demonstrated on an independent hospital cohort. Figure 15 The relative abundance of 27 biomarkers between children with ASD (n = 82) and neurotypical children (n = 90) in an independent hospital cohort, as evaluated by a two-sided Wilcoxon rank-sum test, is shown. To externally validate the diagnostic value and avoid overly optimistic reporting of diagnostic accuracy, the 27 biomarkers were tested in both group and random forest models in an independent hospital cohort (82 boys with ASD and 90 neurotypical boys, aged 4–11 years). The relative abundance of 26 biomarkers remained significantly different between children with ASD and neurotypical children.
[0162] Figure 16 The AUC (95% CI) of the random forest model with different features in the validation queue is shown. The AUC of the model was found to remain in the range of 0.59 to 0.85.
[0163] Figure 17 The performance metrics of the trained random forest models classifying ASD using different features in the validation cohort are shown in detail. Sensitivity, specificity, and accuracy were calculated based on the Yoden index. The ensemble model using 27 biomarkers (AUC 0.81) ranked first with an accuracy of 82.1%, with a sensitivity of 83.8% and a specificity of 80.8%. Furthermore, in a subset of younger children in the validation cohort (n = 31, 6 years or younger), the model's AUC ranged from 0.58 to 0.92, with the ensemble model using 27 biomarkers achieving the highest accuracy. To test whether this group could be applied to predicting the risk of ASD in younger children, the ensemble model was tested with 27 biomarkers in another cohort of younger children (84 boys with ASD and 40 neurotypical boys, 1–8 years old), and after adjusting for imbalanced sample sizes, the model achieved an AUC of 0.81, with a sensitivity of 65.48% and a specificity of 90%.
[0164] When the age range was reduced to six years or younger (31 neurotypical children and 72 children with ASD) or even four years or younger (18 neurotypical children and 34 children with ASD), the integrated model showed promising diagnostic value across age groups, with AUCs of 0.81 and 0.84, respectively. Next, associations between these 27 biomarkers were tested with ASD in this younger, age-stratified cohort, confirming that most of these associations remained reproducible. The model with the 27 biomarkers was then investigated to determine if it would work well in women. The model was tested in an all-female cohort, achieving an AUC of 0.59. Despite a decrease in diagnostic accuracy in female children, the difference in the relative abundance of four key biomarkers between children with ASD and neurotypical children remained significant in women. These biomarkers included the thiamine diphosphate biosynthesis pathway, *Westernella fusionis*, and... Natribaculum Luteum And Bacteroides phage B40_8. In summary, these results demonstrate the robustness of the model and the group of 27 biomarkers across age and cohorts.
[0165] Given the shared alterations in the gut microbiota across different diseases, validating the disease specificity of the identified microbial biomarker group is important to ensure a low false-positive rate for ASD diagnosis. To this end, the model was evaluated in two non-ASD cohorts of children with attention deficit hyperactivity disorder (ADHD, n = 118) or atopic dermatitis (n = 78). ADHD and atopic dermatitis have previously been reported to be associated with alterations in the gut microbiota. The AUC values of the biomarker group were significantly lower in children with atopic dermatitis or ADHD (AUC 0.525; p = 0.704) and atopic dermatitis (AUC 0.547; p = 0.396) compared to an independent cohort of children with ASD. Based on thresholds calculated from the discovery cohort, a total of 20 subjects from these tested groups were predicted to have ASD, reflecting a false-positive rate of 10.2%. Overall, these results support the idea that a multi-boundary group of 27 biomarkers is highly specific for ASD.
[0166] To further test the reproducibility of the multi-border group of 27 biomarkers, 195 shotgun fecal metagenomic datasets from six public datasets representing Asians, Europeans, and Americans were integrated. The biomarker group demonstrated significant diagnostic value in distinguishing children with ASD from those with neurotypical disorder at an AUC of 0.704 (p = 9.4e-07, 95% CI = 0.629–0.779). This performance from public datasets further confirms the robustness and generalizability of the multi-border group of 27 biomarkers across different populations and geographic locations.
[0167] Due to the vast inter-individual heterogeneity of the gut microbiota, it is possible for different species or strains in different individuals to trigger similar pathologies or phenotypes by expressing common pathways. Therefore, targeting its broader metagenomic functions, rather than specific taxa, may represent an effective approach to studying the microbiome-mediated pathogenesis of ASD. Decreased plasma concentrations of thiamine (vitamin B1) and its associated metabolites (e.g., thiamine diphosphate) have been associated with ASD, but the underlying causes remain unclear. A consistent decrease in the relative abundance of the thiamine diphosphate biosynthesis pathway was found in four cohorts of ASD compared to neurotypical children (see [link to study]). Figure 9 and Figure 15 The microbial functional group also showed the highest performance in differentiating ASD from neurotypical children (see [link]). Figure 14 A total of ten enzymes are involved in the thiamine diphosphate biosynthesis pathway. In the discovery cohort (n = 447), all ten enzymes were reduced in children with ASD, and eight of them were significant after multiple comparison adjustments (FDR < 0.05).
[0168] These associations were further evaluated in an independent hospital cohort (n = 172). The relative abundance of all 10 enzymes was significantly reduced in ASD compared to neurotypical children. In summary, these findings highlight that reduced abundance of thiamine diphosphate biosynthesis genes in the gut microbiota appears to be closely associated with ASD. The above data support the idea that microbiota-mediated functions are altered in ASD, and this is also related to multi-kingdom genes.
[0169] In this study, a comprehensive analysis of multi-kingdom and functional microbiomes was performed using over 1000 metagenomic datasets from six different independent cohorts of children. A range of bacterial and non-bacterial biomarkers were identified for the first time, and their performance in children with ASD in the detection cohorts was evaluated. 1188 significant ASD-associated multi-kingdom microbial taxa, functional genes, and pathways were identified. Fungal, archaea, viral species, and functional microbiome pathways were shown to differentiate children with ASD from neurotypical children across different age groups. A random forest model based on a 27-feature group achieved high predictive value for ASD diagnosis (mean AUC = 0.88), including in children under six years of age (AUC = 0.81). The reproducible performance of the models across cohorts and age groups demonstrates their potential as a promising diagnostic and predictive tool for ASD.
[0170] A series of novel bacterial and non-bacterial biomarkers were discovered, and their associations with ASD were profiled. Several promising beneficial bacteria were observed, such as *Westernella fusionis*, *Westernella esculenta*, and... Lawsonibacter asaccharolyticusBacteroides faecalis and other bacteria showed a significant negative correlation with ASD. Furthermore, ASD was also positively correlated with opportunistic pathogens such as Pseudomonas and Oligotrophosporidia. Overall, this imbalanced bacterial ecosystem may play an important role in the pathophysiology of ASD. In addition, the discovery that specific microbial functions may contribute to ASD pathogenesis through dysregulation of thiamine diphosphate biosynthesis may explain the reduced levels of thiamine-related metabolites in children with ASD. Thiamine-related metabolites play crucial roles in mental health and neural signal transduction. These findings provide further evidence that thiamine diphosphate biosynthesis in the gut microbiome could also be used as a novel therapeutic target in the future.
[0171] Whether ASD-related gut microbiome dysbiosis is solely driven by dietary preferences has been controversial. The results presented in this paper show that diet influences the gut microbiome in children with ASD. However, ASD-related microbiome alterations, including microbial diversity and composition, persisted after adjusting for dietary factors and GI symptoms in the model, suggesting that diet may not significantly affect the accuracy of the biomarker panel. Furthermore, analyses were conducted on two common childhood disorders known to be associated with gut microbiome alterations, atopic dermatitis, and ADHD. The panel of 27 biomarkers demonstrated its continued specificity for the diagnosis of ASD.
[0172] As an additional aspect of this disclosure, test subjects identified by the methods of the present invention as having ASD or at an elevated risk of developing ASD may receive appropriate treatment as a therapeutic or preventative measure to address persistent symptoms or the potentially higher risk. For example, medications such as antipsychotics and / or antidepressants may be given to children, or treatments may be given to children, such as those specifically designed to address behavioral problems and / or improve language, communication, or social skills.
[0173] Analyzing cohorts of children across diverse lifestyles, ethnicities, and locations offers a unique opportunity to study the microbiome associated with ASD. By combining multiple small-hospital-based and community cohorts with potentially low universality, better representativeness of the ASD case spectrum and controls is achieved. Appropriate methodologies avoid the artificial findings resulting from batch effects observed in any single dataset. Using large, diverse training datasets also enables the development of more accurate diagnostic models, and the availability of independent validation datasets allows for more realistic estimations of this accuracy.
[0174] Figures 18 to 27Results from a second study were presented, which involved shotgun metagenomic sequencing and supervised machine learning of the gut microbiota in 598 children with ASD and neurotypical children to profile the microbial characterization of the fecal microbiota at different levels for ASD. Statistical and machine learning methods were used to compare the gut metagenomics in the discovery cohort and evaluate their performance in an independent validation cohort. A small set of microbial biomarkers that could aid in the diagnosis of ASD were identified.
[0175] In the second study, a total of 598 children aged 3 to 11 years were recruited (median age: 7 years; interquartile range (IQR) 6–9 years, 80% male), including 264 children with clinician-confirmed ASD and 334 unrelated neurotypical children. The discovery cohort (n = 273) included 129 children with ASD and 144 neurotypical children matched by age and sex, while the independent validation cohort (n = 325) consisted of 135 children with ASD and 190 neurotypical children. Baseline characteristics of the participants are shown in Table 5.
[0176] Table 5
[0177]
[0178] To characterize the gut microbiota distribution in children with ASD, the diversity and composition of the gut microbiome were compared between children with ASD and neurotypical children at both the species and high-resolution MAG levels. Compared with neurotypical children, children with ASD had lower microbial richness at both the species and MAG levels (p-value). 物种 =0.0001, p-value MAG = 0.105). For example... Figure 18 As shown, in children with ASD, microbial diversity (Shannon) and evenness (Pielou) at MAG levels were significantly reduced. The gut microbiota composition in children with ASD differed significantly from that in neurotypical children at both species and MAG levels (p-value). 物种 = 0.001, p-value MAG = 0.001 (based on Bray-Curtis dissimilarity). However, as shown in Table 6, at the MAG level, a larger proportion of microbiome variation was associated with the diagnosis of ASD compared to the species level (R0). 2 物种 = 0.54%, R 2 MAG = 1.83%), indicating that the microbial MAG provides higher resolution and more microbial information in the potential characterization of different microbial variations associated with ASD.
[0179] Table 6
[0180]
[0181] In the discovery cohort, the potential of the gut microbiome for ASD detection was examined. For example... Figure 19 As shown, microbial taxonomy and microbial MAG were used individually and in combination to measure model performance. First, a random forest (RF) classifier was used to analyze all microbial features, and their performance was compared with three other methods, including logistic regression (LG), support vector machine (SVM), and gradient boosting classifier (XGBoost). Among the four different classifiers, the general trend was that microbial taxonomy had better diagnostic performance at higher resolutions. When distinguishing between children with ASD and neurotypical children, microbial species and microbial MAG achieved area under the receiver operating characteristic (AUROC) curves of 0.809 and 0.827, respectively, in the RF classifier. The RF model outperformed the LG model in terms of diagnostic performance. 物种 : 0.737; AUROC MAG : 0.809), SVM (AUROC) 物种 : 0.762; AUROC MAG (0.809) and XGBoost classifier (AUROC) 物种 : 0.821; AUROC MAG The diagnostic performance of the model was 0.808. Furthermore, the combination of microbial taxonomy and microbial MAG achieved better diagnostic performance than either microbial taxonomy or microbial MAG alone, with average AUROCs of 0.858, 0.858, 0.811, and 0.816 for RF, LG, SVM, and XGBoost classifiers, respectively.
[0182] To reduce interference from redundant microbiomes in biomarker identification, pre-selection of microbial biomarkers associated with ASD was performed by analyzing the differential abundance of taxa and MAGs in the microbiome composition (ANCOM) test within the cohort. After adjusting for age, sex, and body mass index (BMI), 54 microbial taxa and 703 microbial MAGs showed significant differential abundance between children with ASD and neurotypical children at a threshold >0.6. Subsequently, the pre-selected microbial features used for biomarker identification and model training were combined in the ASD diagnosis using RF, LG, SVM, and XGBoost classifiers.
[0183] Figure 20The workflow for developing a fecal microbiome-based diagnostic model for ASD is shown. For children with ASD and neurotypical children, samples are randomly divided into a training set (70% of samples) for the discovery cohort and a test set (the remaining 30%) for evaluation, based on the prediction target. The split itself is performed randomly 10 times to assess sampling variation. Within the discovery cohort, the predictive model is trained and tested via cross-validation, and the final performance of the best model is evaluated in a separate validation cohort.
[0184] Figure 21 Results using the RF classifier are shown, where the RF classifier, through cross-validation, achieved the highest predictive power among 5 bacterial taxa and 44 microbial MAGs (2 viral MAGs and 42 bacterial MAGs), with an AUROC of 0.886 (95% confidence interval: 0.868–0.892; sensitivity: 0.773, specificity: 0.834 at the optimal cutoff). Of the 49 microbial biomarkers, 2 bacterial taxa and 9 microbial MAGs were enriched in children with ASD, while 3 bacterial taxa and 35 microbial MAGs showed lower abundance in children with ASD compared to neurotypical children, suggesting that ASD-specific microbiome alterations are primarily attributable to the gut microbiota deficient in ASD. Furthermore, the associations of these biomarkers with host factors were explored in the discovery cohort. Figure 22 As shown, negligible direct associations were found between these microbial biomarkers and sex, BMI, and GI symptoms, suggesting that these selected biomarkers should not be affected by these host factors. Figure 23 Results are shown using the optimal diagnostic model applied to an independent validation cohort of 135 children with ASD and 190 neurotypical children. The optimal model showed an AUROC of 0.734 (95% confidence interval: 0.682, 0.782; sensitivity = 80.0%, specificity = 71.6%) when distinguishing between children with ASD and neurotypical children. Microbial biomarkers for the same cohort tested in the LG classifier showed results comparable to the RF classifier (AUROC = 0.742), but superior to XGBoost (AUROC = 0.720) and SVM (AUROC = 0.709). The model output was stratified into high-confidence and median-confidence segments by applying the optimal cutoff values and categorizing the model with the currently authorized behavioral diagnostic device (Cognoa) and improving model utilization efficiency.
[0185] Figure 24The distribution of predicted risk scores (ASD probabilities) for all samples in the validation cohort is shown, where the area under line 2402 represents the ASD probability distribution for children with ASD in the validation cohort, and the area under line 2404 represents the ASD probability distribution for healthy controls. According to clinical settings, model accuracy can be improved by applying two cutoff values for the ASD probability to the classification algorithm, where the region between the two cutoff values is called the median confidence interval and is classified as uncertain. ASD probabilities falling outside the median confidence interval are called the high confidence interval.
[0186] like Figure 24 As shown, for ASD detection, the performance using the high-confidence segment achieved an AUROC of 0.906 with a sensitivity of 85.7% and a specificity of 93.4%. Compared to the overall model performance, the AUROC, sensitivity, and specificity in the high-confidence segment showed superior performance, with an increase of 23% in AUROC, 4.1% in sensitivity, and 38.6% in specificity. In the validation cohort, the distribution of ASD risk scores (ASD probabilities) for ASD cases peaked at 0.62, while the distribution of ASD risk scores (ASD probabilities) for neurotypical children peaked at 0.36, demonstrating the robust classification ability of the microbiome model.
[0187] In some cases, the Social Responsiveness Scale, Version 2 (SRS-2) is used to assess the severity of social impairment within the autism spectrum. Two subscales of the SRS-2, namely Social Communication and Interaction (SRS-SCI) and Repetitive and Restrictive Behaviors (SRS-RRB), were analyzed, corresponding to the two core symptom domains of ASD and conforming to the DSM-5 criteria for ASD. Figure 25 The study showed that 49 identified microbial biomarkers exhibited significant correlations between single core symptom indices (Social Communication and Interaction (SRS-SCI) and Repetitive and Restrictive Behaviors (SRS-RRB)) and total severity scores (SRS-T) in the validation cohort.
[0188] Higher abundance of Clostridium species enriched in ASD and microbial MAGs annotated with active rumenococci were positively correlated with more severe symptoms in ASD, including core symptom indices (social communication and interaction, and repetitive and restrictive behaviors) and total severity scores. Conversely, lower abundance of microbial MAGs lacking in ASD, annotated with bacteria of the Trichophyceae family, *Pseudomonas previae*, and *Desulfovibrio lavans*, was also associated with more severe symptoms in ASD.
[0189] Next, the association between the predicted ASD risk score (ASD probability) and these ASD clinical scores was explored. Figure 26The predicted ASD risk score (ASD probability) showed a consistent positive correlation with social communication and interaction scores, repetitive and restrictive behavior scores, and total severity scores. Based on the model, it was shown that children with more severe ASD symptoms, as defined by higher SRS scores, tended to have higher ASD probability scores. These results indicate that microbial biomarkers and microbiome-based models are not only associated with the presence of ASD but also reflect the severity of ASD symptoms.
[0190] To further determine the generalizability of the model across different child groups, subjects were stratified into subgroups in the validation cohort based on their age, sex, presence of GI symptoms, and presence of comorbid mental illnesses (ADHD and anxiety). Figure 27 As shown, the subgroups included children aged 6 years and under (n = 71; compared to children older than 6 years, n = 254), males (n = 224; compared to females, n = 101), children without GI symptoms (n = 309; compared to children with GI symptoms (including constipation and diarrhea), n = 16), children without comorbid mental illness (n = 251; compared to children with ADHD, n = 26), and children without anxiety (n = 255; compared to children with anxiety, n = 50). The model's performance was found to be higher in younger children (≤ 6 years) than in older children (> 6 years), with AUROC increasing from 0.734 to 0.845. However, its performance in other models remained largely unchanged in other subgroups (AUROC). 男性 0.745, AUROC 女性 0.708, AUROC 无gi症状 0.733, AUROC gi症状 0.720, AUROC 无ADHD 0.729, AUROC ADHD : 0.730; AUROC 无焦虑 0.740, AUROC 焦虑 (0.728). These observations suggest that the model has higher accuracy in younger children, but its diagnostic performance is not affected by gender or the presence of GI symptoms or comorbid mental illnesses (ADHD and anxiety).
[0191] Currently, the gold standard for clinical diagnosis of ASD is physician assessment based on clinical observation of developmental history and behavior, supplemented by several psychological tests. However, the dynamic nature and heterogeneous presentation of children's developmental trajectories pose challenges to early detection of the condition. Therefore, robust biomarkers may be helpful for the clinical application of early detection. The gut microbiota is increasingly recognized as a regulator of brain development and behavior, but its role as a potential biomarker in ASD detection has not been fully utilized. Using shotgun metagenomic sequencing and machine learning, ASD-related microbial taxa and MAGs were identified, and subsequently, microbiome-based models were developed to aid in the diagnosis of ASD.
[0192] The study focuses on childhood, a critical window for human growth and health. In fact, the model is more sensitive in children aged 6 and younger. Furthermore, metagenomic assembly is integrated with a reference-based microbial distribution approach to characterize novel microbial features in ASD. An internal gene catalog was established by providing more information (nucleotide sequences of multi-kingdom microbes present in the samples) and de novo metagenomic assembly of the high-resolution microbial community. As the findings suggest, higher predictive performance was observed at higher microbial resolution levels for ASD, indicating that the association strength of higher-resolution gut microbial features may exceed that at lower resolutions, and that they can more accurately capture microbial variations in ASD.
[0193] Based on the model, the predicted risk score (probability of ASD) showed a significant positive correlation with the severity of social skills (total T-score) and single core symptoms (social communication and interaction, and repetitive and restrictive behaviors) on the autism spectrum, indicating that children with ASD exhibiting more severe symptoms tend to have a more typical "ASD-type" gut microbiome distribution. The microbiome model was more sensitive in children aged 6 years or younger, suggesting that ASD-associated microbes can serve as a substitute for risk stratification of ASD in early childhood, which could help children access services with potentially long-term impacts on developmental outcomes earlier. Clinically, ASD accompanied by other mental and physical health conditions may contribute to reduced certainty by making differential diagnosis more challenging. ADHD and anxiety are the most common comorbidities in diagnosis, accounting for 40% to 70% of individuals with ASD. 40,41 However, microbiome-based models showed good applicability in different subsets of the validation cohort where comorbidity was present or absent.
[0194] Considering the dynamic nature of ASD during childhood development, a single cutoff / threshold output by the algorithm is unsuitable for ASD detection. By drawing an analogy with the current behavioral diagnostic device (Cognoa) and providing a more applicable model, high-confidence and median confidence intervals are set for the predicted risk score (ASD probability). When the predicted ASD probability ranges from 0.389 to 0.670, the model returns a median confidence result. The model outperforms the behavioral Cognoa device, as demonstrated by higher specificity and more uniform sensitivity and specificity, indicating greater reliability in identifying individuals without ASD.
[0195] Notably, beneficial symbionts such as *Bacteroides monomorpha*, *Trichophyton*, and *Brutella westermani* were observed to be reduced in cohorts of patients with ASD. In mouse studies, *Bacteroides monomorpha* and *Trichophyton* A2 have been reported to be associated with regulatory behavior and restoration of ASD-like phenotypes by modulating intestinal amino acid transport and serum glutamine levels, as well as by modulating peripheral immune regulation. *Brutella westermani* has also been reported to have anti-inflammatory effects on peripheral blood and to be significantly reduced in other psychotic conditions. Furthermore, children with ASD have been observed to have reduced species of tailed bacteriophages, which have been reported to enhance the ability to recognize and remember novel objects by upregulating brain genes involved in memory. Children with ASD are also characterized by enrichment in *Clostridium* and active *Ruminococcus*. *Clostridium* has been reported to be associated with brain tissue damage and neurological disorders by producing clostridial toxins that may distort dendritic spine complexity during childhood development. Tryptophan decarboxylase from active rumenococci can regulate the formation of the neurotransmitter tryptophan, suggesting a possible direct mechanism by which the gut microbiota influences host behavior.
[0196] Figure 28 The different phases of the third study are illustrated, including discovery phase 2802, machine learning phase 2804, deployment phase 2806, and validation phase 2808. A total of 1,627 children (aged 1–13 years, 24.4% female) were recruited from five independent cohorts in this study. In total, over 10 terabytes of sequence data were obtained for each metagenomic sample at an average depth of 6.34 gigabytes. In the discovery cohort, metagenomic sequencing was performed on fecal samples from 709 children with ASD and 374 neurotypical children (aged 3–12 years, 24.3% female). The dataset was processed using Kraken2 and HUMAnN3, filtered at a 15% prevalence, and contains KEGG Orthology of 8,972 species-level taxa (393 archaea, 8,178 bacteria, 78 fungi, and 323 viruses), 500 functional pathways, and 5,225 microbial gene families.
[0197] In a separate hospital cohort, a total of 172 stool samples from 82 children with ASD and 90 neurotypical children (aged 4–11 years, all male) were sequenced, and these data were used for external validation. A community cohort of younger children, consisting of 116 children with ASD and 60 neurotypical children (aged 1–8 years, 29.5% female), was also included to validate the findings across different age groups. In the external validation analysis, 237 stool metagenomics from published datasets (aged 2–13 years, 17.7% female) were included. Two additional cohorts of non-ASD children with attention deficit hyperactivity disorder (ADHD, n = 118) and atopic dermatitis (n = 78) were used to evaluate the specificity of the findings.
[0198] Extensive phenotypic data (236 factors) were collected, including age, sex, BMI, diet, medication use, comorbidities, associated mental disorders, GI symptoms (including stool consistency assessed by the Bristol Stool Morphology Score (BSFS)), family characteristics, and technical factors related to sample collection, preservation, and processing. All stool samples were processed using the same standardized protocol to reduce batch effects caused by technical factors. In the discovery cohort, metagenomic sequencing was performed on stool samples from 709 children with ASD and 374 neurotypical children (aged 3–12 years, 24.3% female). The dataset was processed by Kraken2 and HUMAnN3, filtered at a 15% prevalence, and included KEGGOrthology (KO genes) of 8,972 species-level taxa (393 archaea, 8,178 bacteria, 78 fungi, and 323 viruses), 500 functional pathways, and 5,225 microbial gene families.
[0199] In a separate hospital cohort, a total of 172 stool samples from 82 children with ASD and 90 neurotypical children (aged 4–11 years, all male) were sequenced, and this data was used for external validation. A community cohort of younger children, consisting of 116 children with ASD and 60 neurotypical children (aged 1–8 years, 29.5% female), was also included to validate the findings across different age groups. In the external validation analysis, 237 stool metagenomics from published datasets (aged 2–13 years, 17.7% female) were included. Two additional cohorts of non-ASD children with attention deficit hyperactivity disorder (ADHD, n = 118) and atopic dermatitis (n = 78) were used to evaluate the specificity of the findings.
[0200] Since the composition of the gut microbiome is primarily influenced by environmental and host factors, the effects of 236 host factors on gut microbiome composition were analyzed. These host factors included ASD, age, sex, BMI, comorbidity (n = 7), diet (n = 201), family characteristics (n = 5), gastrointestinal (GI) parameters / symptoms (n = 4), medication use (n = 4), comorbid mental disorders (n = 3), and technical factors (n = 7). In the cohort (aged 3–12 years, n = 1083, 24.3% female), these host factors combined explained 12.5%, 15.1%, 10.7%, and 11.7% of the inter-individual microbiome variation in archaea, bacteria, fungi, and viruses, respectively.
[0201] Figure 29 The variation in multi-kingdom (archaea, bacteria, fungi, and viruses) microbiome composition explained by phenotypic factors is shown in a multivariate PERMANOVA analysis. GI parameters and the diagnosis of ASD were the top two factors contributing to the greatest variation in microbiome composition. A total of 21 factors showed significant effects on all four kingdoms of gut microbiome composition across all host and dietary factors studied; these factors included ASD, age, sex, BMI, three GI parameters, fifteen dietary factors, and sequencing batch, and were therefore adjusted for in all subsequent association analyses. Next, changes in gut microbiome diversity were assessed in children with ASD and neurotypical children.
[0202] Figure 30 The alpha diversity in the multi-kingdom (archaea, bacteria, fungi, and viruses) microbiome, as measured by Shannon indices, is shown in children with ASD (n = 709) and NT (n = 374). Children with ASD showed reduced Shannon diversity indices for archaea, bacteria, and viruses compared to neurotypical children. However, there was no significant difference in fungal diversity between children with ASD and neurotypical children. The abundance of archaea, bacteria, fungi, and viruses was also significantly lower in children with ASD than in neurotypical children.
[0203] Figure 31 A volcano plot is shown, illustrating associations between multi-kingdom (archaea, bacteria, fungi, and viruses) species and ASD, calculated by MaAsLin2 after adjusting for significant confounding factors. Associations with an adjusted p-value (FDR) less than 0.05 are considered significant and are marked with solid circles. Figure 31 ASD-enriched associations were found in the upper right quadrant of each plot, while ASD-reduced associations were found in the upper left quadrant of each plot. Top-ranked species are labeled.
[0204] To identify microbial biomarkers for ASD, the association between multi-kingdom species and ASD was examined. A total of 14 archaea, 51 bacteria, 7 fungi, and 18 viruses showed varying abundances between children with ASD and neurotypical children (FDR < 0.05). Compared to neurotypical children, the relative abundance of 80 of the 90 identified microbial species was significantly reduced in the gut of children with ASD. This finding was most significant for the bacterial community, where 50 bacterial species were reduced in the gut of children with ASD, while only one bacterial species was enriched in ASD. The variation in bacterial species in children with ASD was driven by a reduction in Streptococcus thermophilus and short-chain fatty acid-producing bacteria (e.g., Bacteroides pHL2737, Lawsonia intracellularis, and Bacteroides species).
[0205] Figure 32 and Figure 33 The association between ASD and fecal microbiome function was shown. Figure 32 The variation in microbiome function (pathways and genes) explained by the phenotype in a multivariate PERMANOVA analysis is shown. The influence of host phenotypic factors on microbial function was explored based on kegorthology (KO genes, n = 5,225) and metabolic pathways (n = 500). Figure 32 As shown, host phenotypic factors explained 17.1% and 15.7% of the variation in microbiome functional pathways and KO genes, respectively. A diagnosis of ASD was listed as the primary factor, accounting for 4.3% and 3.6% of the variation in microbiome functional pathways and KO genes, respectively. A total of 19 host and dietary factors showed significant effects on gut microbiome function, including ASD, age, sex, BMI, two GI parameters, twelve dietary factors, and sequencing batch. After adjusting for these confounding factors, 27 differentially expressed KO genes were identified: 23 differentially expressed KO genes were reduced and 4 were increased in children with ASD compared to neurotypical children (FDR < 0.01).
[0206] Figure 33 A volcano plot is shown, illustrating the associations between microbiome functions (pathways and KO genes) and ASD calculated by MaAsLin 2 after adjusting for significant confounding factors. Associations with an adjusted p-value (FDR) less than 0.01 were considered significant and are marked with solid circles. Figure 33 In each graph, enriched associations in ASD were found in the upper right quadrant, and reduced associations in ASD were found in the upper left quadrant. The top-ranked pathways are labeled with their names.
[0207] Figures 34 to 38This demonstrates how a random forest classifier trained on multi-kingdom microbiome composition and function data can predict the presence or absence of ASD. Figure 34 A schematic diagram of the random forest model development is shown. A balanced reference dataset 3416 was constructed from the original discovery cohort (n = 1083) to compare ASD children (n = 301) with NT children (n = 301) using a one-to-one pairing algorithm that considers all significant confounding factors (n = 24). In some instances, the reference dataset 3416 includes feature vectors of six components: archaea, bacteria, fungi, viruses, microbial pathways, and KO genes.
[0208] A 70 / 30% training-test sample split is performed on the reference dataset 3416 to generate a training dataset 3418 and a test dataset 3420. The training dataset 3418 is used to train an ensemble of decision trees 3422, which is then aggregated to form a random forest model 3424. In the example shown, each of the decision trees 3422 corresponds to a specific biological entity or kingdom. In various instances, the random forest model 3424 may be formed using decision trees 3422 of a single kingdom or biological entity (e.g., a bacteria-only decision tree) or decision trees 3422 of multiple kingdoms or biological entities (e.g., bacteria, fungi, and viruses). After training the random forest model 3424, the accuracy of the random forest model 3424 is estimated by calculating the model's AUC using the test dataset 3420. This training and testing process is repeated 20 times to obtain the distribution of the random forest prediction evaluation on the test dataset 3420.
[0209] Previously, the performance of archaea, fungi, viruses, KO genes, or functional pathways in ASD had not been explored. To avoid discrimination bias caused by sample size imbalance and residual confounding factors, a matching sub-cohort was constructed based on the discovery cohort (n = 447) using a one-to-one matching algorithm for children with ASD (n = 301, 95 girls and 206 boys) and neurotypical children (n = 301, 95 girls and 206 boys). A total of 24 confounding factors identified above were used in the matching algorithm, including age, sex, BMI, three GI parameters, sequencing batch, and 17 dietary factors. After matching, there were no differences in these 24 confounding factors between children with ASD and neurotypical children in the matching sub-cohorts, which were then used for the development of a random forest model.
[0210] Specifically, the distributions of age (p = 0.56), sex (p = 1), BMI (p = 0.64), sequencing batch (p = 0.79), and BSFS (by chi-square test, p = 0.14) were comparable between children with ASD and neurotypical children. The prevalence of functional constipation (9.3% vs 8.9%; p = 0.5) or functional defecation disorder (10.3% vs 7.9%; p = 0.2) was also comparable between children with ASD and neurotypical children. Furthermore, there were no differences in the frequency of consumption of 17 dietary factors between children with ASD and neurotypical children (all p > 0.05 by chi-square test). These dietary factors included beverages (Yakult), dairy products (cheese and whole milk), eggs (boiled and scrambled / fried eggs), fruits (apples / pears), grains (wheat noodles / udon noodles), meat (grilled pork / lean pork, salmon), cooking oils (peanut oil, olive oil, corn oil, vegetable oil, and canola oil), snacks (sweet and ice cream), and vegetables (choy sum). These results suggest that children with ASD and neurotypical children share a similar environmental and host background, therefore the model is unlikely to be influenced by these important factors.
[0211] Figure 35 Box plots are shown, illustrating the distribution of AUC scores obtained from random forest classification on the test dataset. Differences between evaluation groups are analyzed using a two-sided Wilcoxon rank-sum test. Figure 35 In the diagram, an asterisk indicates that the FDR is less than 0.05. The prevalence and relative abundance of the top 31 biomarkers used by the random forest model are also shown in the balanced datasets of children with ASD (n = 301) and neurotypical children (n = 301).
[0212] Single-boundary biomarkers were used to test the accuracy of models in distinguishing children with ASD from neurotypical children. Among all single-boundary biomarkers, the microbial pathway model demonstrated the strongest predictive ability for ASD detection, with an average AUC score of 0.87 at 86.4% specificity and 89.7% sensitivity. Models using microbial genes had the second highest scores, with an average AUC of 0.86, followed by bacterial-based models (average AUC 0.85), archaea-based models (average AUC 0.76), fungal-based models (average AUC 0.74), and virus-based models (average AUC 0.68). Overall, these findings suggest that biomarker features from different boundaries offer promising predictive capabilities for the diagnosis of ASD.
[0213] Since all single-border features have demonstrated diagnostic potential for children with ASD, the performance of a model combining individual multi-border features was investigated. Data from all four borders (including their functional and genetic aspects) were combined. The combined model showed superior performance in diagnosing ASD at 95.1% specificity and 89.5% sensitivity (mean AUC 0.91) compared to single-border feature-based models (all adjusted p < 0.05). Furthermore, the model based on combined multi-border features achieved an accuracy of 90.5%, with a positive predictive value of 97.8% and a negative predictive value of 89.1%. These results confirm that the multi-border fecal microbiome biomarker group has higher diagnostic performance for ASD than the single-border group.
[0214] Figure 36 A graph showing the association between the top 31 biomarkers assessed by MaAsLin2 and ASD is presented. To identify the minimum number of microbiome biomarkers that achieve the highest accuracy, identified biomarkers were sequentially included in the model according to their ranking. Finally, for the diagnosis of ASD, a total of 31 microbiome features (2 archaea, 2 fungi, 3 viruses, 9 bacteria, 5 microbial gene families, and 10 pathways) showed an AUC of 0.91, with a specificity of 92.5% and a sensitivity of 94.1%. The prevalence and relative abundance of biomarkers also differed significantly between children with ASD and neurotypical children.
[0215] Figure 37 The differential abundance and p-values of 31 biomarkers between children with ASD and neurotypical children are shown. It can be observed that 21 biomarkers were significantly reduced and 10 biomarkers were significantly enriched in the gut of children with ASD (all FDR < 0.05).
[0216] Figure 38 The average accuracy reduction for each of the 31 markers used in the random forest model is shown (where... means FDR < 0.05, means FDR < 0.01, This means FDR < 0.001. The importance of these 31 features was re-analyzed in the final integrated model, and it was observed that the model's accuracy was primarily driven by the panthenol-7 biosynthetic pathway, GTPases, and thiamine diphosphate biosynthetic pathway, supporting a potential role for thiamine diphosphate in the pathogenesis of ASD. Furthermore, reductions in several bacteria were among the top-ranking microbial features contributing to diagnostic accuracy; these included *Streptococcus thermophilus*, *Lawsonia intracellularis*, *Westernella fusionis*, *Westernella sinensis*, and *Bacteroides* PHL2737. The palmitoleic acid biosynthetic pathway I was also found to be among these top 31 features. Palmitoleic acid has previously been associated with constipation; therefore, this enriched biosynthetic pathway may also explain the prevalence of constipation in children with ASD. Overall, the analysis shows that 31 fecal microbiome group markers derived from archaea, bacteria, fungi, viruses, KO genes, and microbial functional pathways represent a potentially promising non-invasive tool for the diagnosis of ASD.
[0217] Figures 39 to 42 The random forest model is validated on an independent hospital cohort. To externally validate its diagnostic value and avoid overly optimistic reports of diagnostic accuracy, the trained model and a cohort of 31 biomarkers were tested in an independent hospital cohort (82 boys with ASD and 90 neurotypical boys, aged 4–11 years).
[0218] Figure 39 The AUC (95% CI) of the random forest model with different features in the validation cohort is shown. Sensitivity, specificity, and accuracy were calculated based on the Youden index. The AUC was calculated after adjusting for technical factors and available covariates including age, sex, BMI, BSBF, functional constipation, and defecation disorders. Figure 40 The performance metrics of the trained random forest models classifying ASD using different features in the validation queue are shown in detail. The AUC of the trained models ranged from 0.55 to 0.87. Among them, the ensemble model using 31 biomarkers (adjusted AUC 0.87) ranked first with 82% accuracy, with a sensitivity of 91% and a specificity of 73%. The relative abundance of 28 of the 31 biomarkers remained significantly different between children with ASD and neurotypical children.
[0219] Furthermore, the AUCs of the trained models in the younger subset of children in the validation cohort (n = 31, 6 years or younger) ranged from 0.61 to 0.89, with the ensemble model using 31 biomarkers again achieving the highest accuracy. To test whether this group could be applied to predicting the risk of ASD in younger children, the model trained with 31 biomarkers was tested in another younger cohort (116 ASD subjects and 60 neurotypical subjects, 1–8 years old, 29.5% female). The model achieved an AUC of 0.89, a sensitivity of 75.9%, and a specificity of 86.7%, with relatively balanced performance for both males (AUC 0.88) and females (AUC 0.92). When the age range was reduced to 6 years or younger (42 neurotypical children; 88 subjects with ASD) and 4 years or younger (27 neurotypical children; 46 subjects with ASD), the model showed accuracy of 88.6% and 86.3%, respectively.
[0220] Figure 41 The associations between ASD and 31 identified fecal microbiome markers are shown in the five cohorts of this study. Coefficient values for each association were only labeled if the corresponding FDR was less than 0.05. Figure 42 The AUC of the model tested using 31 biomarkers in independent cohorts of ADHD and atopic dermatitis is shown. ADHD and atopic dermatitis have been reported to be associated with alterations in the gut microbiota. The biomarker cohorts had lower AUC values in children with either atopic dermatitis or ADHD (AUC 0.513, p = 0.809) or atopic dermatitis (AUC 0.575, p = 0.157). In summary, these results demonstrate the robustness of our trained model and the 31 biomarker cohorts across age, sex, and cohorts.
[0221] To further test the reproducibility of the multi-segment group of 31 biomarkers, 237 shotgun fecal metagenomic datasets from six public datasets representing Asians, Europeans, and Americans were integrated. The trained model showed an AUC of 0.78 (p = 9.17e-14, 95% CI = 0.72–0.84, sensitivity 65.30%, specificity 72.40%) in distinguishing between children with ASD and those with neurotypicality. More importantly, the model showed comparable performance for both males (AUC 0.77) and females (AUC 0.82), confirming its applicability to both sexes. Overall, this multi-segment group of 31 biomarkers may be relevant across different populations and geographic locations.
[0222] Due to the vast inter-individual heterogeneity of the gut microbiota, it is possible for different species or strains in different individuals to trigger similar pathologies or phenotypes by expressing common pathways. Therefore, targeting its broader metagenomic functions, rather than specific taxa, may represent an effective approach to studying the microbiome-mediated pathogenesis of ASD. Previous studies have shown that panthenol improves symptoms in children with ASD. Decreased plasma concentrations of thiamine (vitamin B1) and its associated metabolites (e.g., thiamine diphosphate) have been associated with ASD. However, the underlying causes of these observations remain unclear. We found a consistent decrease in the relative abundance of the panthenol-7 biosynthetic pathway and the thiamine diphosphate biosynthetic pathway in ASD across all three cohorts compared to neurotypical children, and this primarily drove the accuracy of our diagnostic model. Figure 41 The microbial functional group also showed the highest diagnostic value in differentiating children with ASD from neurotypical children. Figure 35 A total of 17 enzymes are involved in these two pathways, and most of them were reduced in different cohorts of children with ASD. In summary, these findings highlight that reduced abundance of genes involved in the biosynthesis of panthenol-7 and thiamine diphosphate in the gut microbiota appears to be closely associated with ASD. The above data support the idea that microbiota-mediated functions and microbial genes are altered in ASD.
[0223] In this study, a comprehensive analysis of multi-kingdom and functional microbiomes was performed using over 1600 metagenomic datasets from five different independent cohorts of children. A range of bacterial and non-bacterial biomarkers were identified for the first time, and their performance in detecting children with ASD in the cohorts was evaluated. 1188 significant ASD-associated multi-kingdom microbial taxa, functional genes, and pathways were identified. Fungal, archaea, viral species, and functional microbiome pathways were shown to differentiate children with ASD from neurotypical children across different age groups. A random forest model based on a 31-feature cohort achieved high predictive value for ASD diagnosis (mean AUC = 0.91), including in children under 4 years of age (AUC = 0.91). The reproducible performance of the models across cohorts, sex, age, and public datasets demonstrates their potential as a promising diagnostic and predictive tool for ASD.
[0224] A series of novel bacterial and non-bacterial biomarkers were discovered and their associations with ASD were outlined. Several promising beneficial bacteria, such as *Streptococcus thermophilus*, *Westernella fusionis*, and *Westernella sinensis*, were observed to exhibit a significant negative correlation with ASD. Furthermore, it was found that specific microbial functions can contribute to ASD pathogenesis through dysregulation of panthenol and thiamine diphosphate biosynthesis. Panthenol and thiamine-related metabolites play crucial roles in mental health and neural signal transduction. These findings provide further evidence that thiamine diphosphate biosynthesis in the gut microbiome could also serve as a novel therapeutic target in the future.
[0225] Figure 43 Example computer system 4300, comprising various hardware elements, is shown according to some embodiments of this disclosure. Computer system 4300 may be incorporated into or integrated with the apparatus described herein, and / or may be configured to perform some or all of the steps of the methods provided by the various embodiments. For example, in the various embodiments, computer system 4300 may be configured to perform either method 100 or 300. It should be noted that... Figure 43 This is merely intended to provide a general description of the various components; any one or all of the components may be used appropriately. Therefore, Figure 43 It extensively demonstrates how individual system components can be implemented in a relatively discrete or relatively more integrated manner.
[0226] In the illustrated example, computer system 4300 includes communication medium 4302, one or more processors 4304, one or more input devices 4306, one or more output devices 4308, communication subsystem 4310, and one or more memory devices 4312. Computer system 4300 can be implemented using various hardware implementations and embedded system technologies. For example, one or more components of computer system 4300 can be implemented in integrated circuits (ICs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), field-programmable gate arrays (FPGAs) (such as those commercially available through XILINX®, INTEL®, or LATTICESEMICONDUCTOR®), system-on-a-chip (SoCs), microcontrollers, printed circuit boards (PCBs), and / or hybrid devices (such as SoCs and FPGAs).
[0227] Various hardware components of computer system 4300 can be communicatively connected via communication medium 4302. Although communication medium 4302 is shown as a single connection for clarity, it should be understood that communication medium 4302 may include various numbers and types of communication media for transmitting data between hardware components. For example, communication medium 4302 may include one or more conductors (e.g., conductive traces, lines or leads on a PCB or integrated circuit (IC), microstrip, stripline, coaxial cable), one or more optical waveguides (e.g., optical fiber, strip waveguide), and / or one or more wireless connections or links (e.g., infrared wireless communication, radio communication, microwave wireless communication), etc.
[0228] In some implementations, communication medium 4302 may include one or more buses connecting pins of hardware components of computer system 4300. For example, communication medium 4302 may include a bus, called the system bus, connecting processor 4304 to main memory 4314, and a bus, called the expansion bus, connecting main memory 4314 to input device 4306 or output device 4308. The system bus itself may consist of several buses, including an address bus, a data bus, and a control bus. The address bus can transfer memory addresses from processor 4304 to address bus circuitry associated with main memory 4314, so that the data bus can access data contained at the memory address and transfer it back to processor 4304. The control bus can transmit commands from processor 4304 and return status signals from main memory 4314. Each bus may include multiple wires for carrying multi-bit information, and each bus may support serial or parallel data transmission.
[0229] Processor 4304 may include one or more central processing units (CPUs), graphics processing units (GPUs), neural network processors or accelerators, digital signal processors (DSPs), and / or other general-purpose or special-purpose processors capable of executing instructions. The CPU may take the form of a microprocessor, which may be fabricated on a single IC chip with a metal-oxide-semiconductor field-effect transistor (MOSFET) structure. Processor 4304 may include one or more multi-core processors, in which each core can simultaneously read and execute program instructions with other cores, improving the speed of multi-threaded programs.
[0230] Input device 4306 may include one or more of various user input devices (e.g., mouse, keyboard, microphone) and various sensor input devices, such as image capture devices, temperature sensors (e.g., thermometer, thermocouple, thermistor), pressure sensors (e.g., barometer, tactile sensor), motion sensors (e.g., accelerometer, gyroscope, tilt sensor), and light sensors (e.g., photodiode, photodetector, charge-coupled device). Input device 4306 may also include means for reading and / or receiving removable storage devices or other removable media. Such removable media may include optical discs (e.g., Blu-ray discs, DVDs, CDs), memory cards (e.g., CompactFlash cards, Secure Digital (SD) cards, Memory Sticks), floppy disks, Universal Serial Bus (USB) flash drives, external hard disk drives (HDDs) or solid-state drives (SSDs), etc.
[0231] Output device 4308 may include one or more of a variety of devices for converting information into a human-readable form, such as, but not limited to, display devices, speakers, printers, tactile or sensory devices, etc. Output device 4308 may also include means for writing to removable storage devices or other removable media, such as those described in reference input device 4306. Output device 4308 may also include various actuators for causing physical movement of one or more components. Such actuators may be hydraulic, pneumatic, or electric, and may be controlled using control signals generated by computer system 4300.
[0232] The communication subsystem 4310 may include hardware components for connecting the computer system 4300 to a system or device located outside the computer system 4300, such as via a computer network. In various embodiments, the communication subsystem 4310 may include wired communication devices (e.g., Universal Asynchronous Receiver-Transmitter (UART)), optical communication devices (e.g., optical modems), infrared communication devices, radio communication devices (e.g., wireless network interface controllers, BLUETOOTH® devices, IEEE 802.11 devices, Wi-Fi devices, Wi-Max devices, cellular devices), etc., connected to one or more input / output ports.
[0233] Memory device 4312 may include various data storage devices of computer system 4300. For example, memory device 4312 may include various types of computer memory with varying response times and capacities, ranging from faster response times and lower capacity memory (e.g., processor registers and caches (e.g., L0, L1, L2)) to medium response times and medium capacity memory (e.g., random access memory (RAM)) to lower response times and lower capacity memory (e.g., solid-state drives and hard disks). Although processor 4304 and memory device 4312 are shown as separate elements, it should be understood that processor 4304 may include different levels of on-processor memory, such as processor registers and caches that may be used by a single processor or shared among multiple processors.
[0234] The memory device 4312 may include a main memory 4314, which can be directly accessed by the processor 4304 via the address and data bus of the communication medium 4302. For example, the processor 4304 can continuously read and execute instructions stored in the main memory 4314. Thus, various software elements can be loaded into the main memory 4314 for reading and execution by the processor 4304, such as... Figure 43 As shown. Typically, main memory 4314 is volatile memory, which loses all data when power is lost, and therefore requires power to retain the stored data. Main memory 4314 may also include a small portion of non-volatile memory containing software (e.g., firmware, such as BIOS) for reading other software stored in memory device 4312 into main memory 4314. In some embodiments, the volatile memory of main memory 4314 is implemented as RAM (e.g., dynamic random access memory (DRAM)), and the non-volatile memory of main memory 4314 is implemented as read-only memory (ROM) (e.g., flash memory, erasable programmable read-only memory (EPROM), or electrically erasable programmable read-only memory (EEPROM)).
[0235] Computer system 4300 may include software elements shown as currently residing within main memory 4314. These software elements may include an operating system, device drivers, firmware, compilers, and / or other code, such as one or more application programs, which may include computer programs provided by various embodiments of this disclosure. By way of example only, one or more steps described with respect to any of the methods discussed above may be implemented as instructions 4316, which may be executed by computer system 4300. In one example, such instructions 4316 may be received by computer system 4300 using communication subsystem 4310 (e.g., via a wireless or wired signal carrying instructions 4316), transmitted by communication medium 4302 to memory device 4312, stored in memory device 4312, read into main memory 4314, and executed by processor 4304 to perform one or more steps of the method. In another instance, instruction 4316 may be received by computer system 4300 using input device 4306 (e.g., via a reader for removable media), transmitted by communication medium 4302 to memory device 4312, stored in memory device 4312, read into main memory 4314, and executed by processor 4304 to perform one or more steps of the method.
[0236] In some embodiments of this disclosure, instruction 4316 is stored on a computer-readable storage medium (or simply a computer-readable medium). This computer-readable medium may be non-transitory and therefore may be referred to as a non-transitory computer-readable medium. In some cases, the non-transitory computer-readable medium may be incorporated into the computer system 4300. For example, the non-transitory computer-readable medium may be one of the memory devices 4312 (such as…). Figure 43 (As shown). In some cases, a non-transitory computer-readable medium can be separated from the computer system 4300. In one instance, the non-transitory computer-readable medium may be provided to the input device 4306 (such as...). Figure 43 A removable medium (e.g., those described with reference to input device 4306) is used, wherein instructions 4316 are read into computer system 4300 by input device 4306. In another example, a non-transitory computer-readable medium may be a component of a remote electronic device (e.g., a mobile phone) that can wirelessly transmit data signals carrying instructions 4316 to computer system 4300 and be communicated by communication subsystem 4310 (e.g., ...). Figure 43 (As shown) Receive.
[0237] Instruction 4316 may take any suitable form to be read and / or executed by computer system 4300. For example, instruction 4316 may be source code (written in a human-readable programming language such as Java, C, C++, C#, Python), object code, assembly language, machine code, microcode, executable code, etc. In one instance, instruction 4316 is provided to computer system 4300 in the form of source code, and a compiler is used to translate instruction 4316 from source code into machine code, which can then be read into main memory 4314 for execution by processor 4304. As another example, instruction 4316 is provided to computer system 4300 in the form of an executable file with machine code, which can be immediately read into main memory 4314 for execution by processor 4304. In various instances, instruction 4316 may be provided to computer system 4300 in encrypted or unencrypted form, compressed or uncompressed form, as an installation package or for initialization for broader software deployment, etc.
[0238] In one aspect of this disclosure, a system (e.g., computer system 4300) is provided to perform methods according to various embodiments of this disclosure. For example, some embodiments may include a system comprising one or more processors (e.g., processor 4304) communicatively connected to a non-transitory computer-readable medium (e.g., memory device 4312 or main memory 4314). The non-transitory computer-readable medium may have instructions stored therein (e.g., instruction 4316) that, when executed by the one or more processors, cause the one or more processors to perform the methods described in the various embodiments.
[0239] In another aspect of this disclosure, a computer program product including instructions (e.g., instruction 4316) is provided to perform methods according to various embodiments of this disclosure. The computer program product may be tangibly contained in a non-transitory computer-readable medium (e.g., memory device 4312 or main memory 4314). The instructions may be configured to cause one or more processors (e.g., one or more processors 4304) to perform the methods described in the various embodiments.
[0240] In another aspect of this disclosure, a non-transitory computer-readable medium (e.g., memory device 4312 or main memory 4314) is provided. The non-transitory computer-readable medium may have instructions stored therein (e.g., instruction 4316) that, when executed by one or more processors (e.g., processor 4304), cause the one or more processors to perform the methods described in the various embodiments.
[0241] The methods, systems, and apparatuses described above are examples. Various configurations can omit, replace, or add various procedures or components as needed. For example, in alternative configurations, methods can be performed in a different order than described, and / or stages can be added, omitted, and / or combined. Furthermore, the features described in relation to certain configurations can be combined in various other configurations. Different aspects and elements of the configurations can be combined in a similar manner. Moreover, technology is evolving; therefore, many elements are examples and do not limit the scope of this disclosure or the claims.
[0242] Specific details are provided in the specification to offer a thorough understanding of the exemplary configurations, including implementations. However, the configurations may be implemented without these specific details. For example, well-known circuits, processes, algorithms, structures, and techniques have been shown without unnecessary detail to avoid obscuring the configurations. The specification of this application provides only exemplary configurations and does not limit the scope, applicability, or configuration of the claims. Rather, the preceding description of the configurations will provide those skilled in the art with a feasible description of the techniques for implementing the descriptions. Various changes may be made to the function and arrangement of the elements without departing from the spirit or scope of this disclosure.
[0243] Several example configurations have been described, and various modifications, alternative constructions, and equivalents may be used without departing from the spirit of this disclosure. For example, the aforementioned elements may be components of a larger system, where other rules may take precedence over or modify the application of the technology in other ways. Furthermore, multiple steps may be taken before, during, or after considering the aforementioned elements. Therefore, the above description does not limit the scope of the claims.
[0244] As used herein and in the appended claims, the singular forms “a,” “an,” and “the” include plural references unless the context clearly specifies otherwise. Thus, for example, reference to “a user” includes reference to one or more such users, and reference to “a processor” includes reference to one or more processors and their equivalents known to those skilled in the art, and so on.
[0245] Furthermore, when the terms “comprise,” “comprising,” “contains,” “containing,” “include,” “including,” and “includes” are used in this specification and the appended claims, they are intended to specifically describe the presence of the stated feature, unit, component, or step, but do not preclude the presence or addition of one or more other features, units, components, steps, actions, or groups.
[0246] It should also be understood that the embodiments and implementations described herein are for illustrative purposes only, and various modifications or changes made based on the embodiments and implementations will inspire those skilled in the art, and such modifications or changes will be included within the spirit and scope of this application and the scope of the appended claims.
[0247] All patents, patent applications and other publications (including GenBank accession numbers or similar serial numbers) cited in this application are incorporated herein by reference in their entirety for all purposes.
Claims
1. A computer-implemented method for training a random forest model to predict the presence of autism spectrum disorder (ASD) in subjects, the method comprising: Obtain the training dataset, which includes: A collection of identified microbial biomarkers; Abundance measurements of the set of microbial biomarkers of the first cohort of subjects suffering from ASD; and Abundance measurements of the set of microbial biomarkers of the second cohort of subjects who did not have ASD; The training dataset is divided into k subsets; For each of the k subsets, train k candidate random forest models in the following manner: Generate a validation dataset that includes one of the k subsets and a training subset that includes the remaining k-1 subsets from the k subsets; Train the candidate random forest model from the k candidate random forest models using the training subset; and The candidate random forest model is evaluated using the validation dataset by calculating its performance metrics. The k candidate random forest models are evaluated by calculating the area under the curve (AUC) value of each of the k candidate random forest models; and One of the k candidate random forest models with the largest AUC value is deployed as the trained random forest model.
2. The computer-implemented method according to claim 1, wherein training the candidate random forest model using the training subset comprises: n training subsets are generated by sampling (i) objects from the first queue objects and the second queue objects and (ii) microbial markers from the set of microbial markers from the training subsets; as well as The n sets of training subsets are used to construct n decision trees, and the n decision trees form the candidate random forest model.
3. The computer-implemented method of claim 2, wherein objects from the first queue object and the second queue object are sampled with replacement, and wherein microbial markers from the set of microbial markers are sampled with replacement.
4. The computer-implemented method of claim 2, wherein the objects from the first queue object and the second queue object are randomly sampled, and wherein the microbial markers from the set of microbial markers are randomly sampled.
5. The computer-implemented method according to claim 1, wherein the performance metric includes one or more of the following: total number of false positives, total number of true positives, total number of false negatives, and total number of true negatives.
6. The computer-implemented method according to claim 1, wherein the k candidate random forest models have a first set of hyperparameters, and wherein the computer-implemented method further comprises: Train k second-candidate random forest models with a second set of hyperparameters that are different from the first set of hyperparameters; as well as The k second candidate random forest models are evaluated by calculating the AUC value of each of the k second candidate random forest models; One of the k candidate random forest models or one of the k second candidate random forest models with the largest AUC value is deployed as the trained random forest model.
7. The computer-implemented method of claim 6, wherein a set of hyperparameters corresponding to one of the k candidate random forest models with the largest AUC value or one of the k second candidate random forest models with the largest AUC value is used to retrain the random forest model using the entire reference dataset and is deployed as the trained random forest model.
8. The computer-implemented method according to claim 6, further comprising: Train k third candidate random forest models with a third set of hyperparameters that are different from the first set of hyperparameters and the second set of hyperparameters; as well as The k third-candidate random forest models are evaluated by calculating the AUC value of each of the k third-candidate random forest models. The k candidate random forest models, the k second candidate random forest models, or the k third candidate random forest models with the largest AUC value are selected as the trained random forest model.
9. The computer-implemented method of claim 6, wherein the first set of hyperparameters and the second set of hyperparameters differ in at least one of the following aspects: The number of decision trees in their respective candidate random forest models; The maximum number of nodes in the decision tree in each of the respective candidate random forest models; The maximum or minimum amount of sampled objects from the first and second queue objects used to construct the decision tree in the respective candidate random forest models; or The maximum or minimum number of sampled microbial biomarkers from the set of microbial biomarkers used to construct the decision tree in the respective candidate random forest models.
10. The computer-implemented method according to claim 1, further comprising: The AUC values of each of the k candidate random forest models are compared to identify the maximum AUC value.
11. The computer-implemented method of claim 1, wherein obtaining the training data comprises: Obtain a reference dataset; as well as The reference dataset is divided into a training dataset and a test dataset, wherein the test dataset is used to evaluate the k candidate random forest models.
12. The computer-implemented method according to claim 1, wherein the abundance measurements of the first queue object and the second queue object are relative abundance measurements.
13. The computer-implemented method of claim 1, wherein the k candidate random forest models have a set of hyperparameters, comprising: 10 to 500 decision trees; One to five nodes in the decision tree; 70% to 80% of the objects from the first queue and the second queue are used to construct the decision tree in the respective candidate random forest models; and Five to 50 sampled microbial biomarkers from the set of microbial biomarkers are used to construct the decision tree in the respective candidate random forest models.
14. A non-transitory computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to train a random forest model to predict the presence of autism spectrum disorder (ASD) in a subject, the method comprising: Obtain the training dataset, which includes: A collection of identified microbial biomarkers; Abundance measurements of the set of microbial biomarkers of the first cohort of subjects suffering from ASD; and Abundance measurements of the set of microbial biomarkers of the second cohort of subjects who did not have ASD; The training dataset is divided into k subsets; For each of the k subsets, train k candidate random forest models in the following manner: Generate a validation dataset that includes one of the k subsets and a training subset that includes the remaining k-1 subsets from the k subsets; Train the candidate random forest model from the k candidate random forest models using the training subset; and The candidate random forest model is evaluated using the validation dataset by calculating its performance metrics. The k candidate random forest models are evaluated by calculating the area under the curve (AUC) value of each of the k candidate random forest models; and One of the k candidate random forest models with the largest AUC value is deployed as the trained random forest model.
15. The non-transient computer-readable medium of claim 14, wherein training the candidate random forest model using the training subset comprises: n training subsets are generated by sampling (i) objects from the first queue objects and the second queue objects and (ii) microbial markers from the set of microbial markers from the training subsets; as well as The n sets of training subsets are used to construct n decision trees, and the n decision trees form the candidate random forest model.
16. The non-transient computer-readable medium of claim 15, wherein objects from the first queue objects and the second queue objects are sampled with replacement, and wherein microbial markers from the set of said microbial markers are sampled with replacement.
17. The non-transitory computer-readable medium of claim 15, wherein the objects from the first queue objects and the second queue objects are randomly sampled, and wherein the microbial markers from the set of said microbial markers are randomly sampled.
18. The non-transient computer-readable medium of claim 14, wherein the performance metric includes one or more of the following: total number of false positives, total number of true positives, total number of false negatives, and total number of true negatives.
19. A system comprising: One or more processors; and A computer-readable medium including instructions that, when executed by one or more processors, cause the one or more processors to perform a method of training a random forest model to predict the presence of autism spectrum disorder (ASD) in a subject, the method comprising: Obtain the training dataset, which includes: A collection of identified microbial biomarkers; Abundance measurements of the set of microbial biomarkers of the first cohort of subjects suffering from ASD; and Abundance measurements of the set of microbial biomarkers of the second cohort of subjects who did not have ASD; The training dataset is divided into k subsets; For each of the k subsets, train k candidate random forest models in the following manner: Generate a validation dataset that includes one of the k subsets and a training subset that includes the remaining k-1 subsets from the k subsets; Train the candidate random forest model from the k candidate random forest models using the training subset; and The candidate random forest model is evaluated using the validation dataset by calculating its performance metrics. The k candidate random forest models are evaluated by calculating the area under the curve (AUC) value of each of the k candidate random forest models; and One of the k candidate random forest models with the largest AUC value is deployed as the trained random forest model.
20. The system of claim 19, wherein training the candidate random forest model using the training subset comprises: n training subsets are generated by sampling (i) objects from the first queue objects and the second queue objects and (ii) microbial markers from the set of microbial markers from the training subsets; as well as The n sets of training subsets are used to construct n decision trees, and the n decision trees form the candidate random forest model.