Autism prediction model based on blood microorganisms and application thereof

By constructing an ASD prediction model based on blood microbiota, 12 biomarkers were screened using blood metagenomic data from 1,946 ASD standard four-family pedigrees. Using dual-pathway differential analysis and multi-dimensional statistical control, the accuracy and stability issues of early ASD diagnosis were resolved, enabling early risk assessment and stratification of ASD, and providing a sensitive and accurate objective assessment tool.

CN121839140BActive Publication Date: 2026-06-26XIANGYA HOSPITAL CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIANGYA HOSPITAL CENT SOUTH UNIV
Filing Date
2026-03-06
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

In existing technologies, the diagnosis of autism spectrum disorder (ASD) relies on behavioral assessment, which has low accuracy, high subjectivity, and difficulty in early objective quantitative assessment. Furthermore, the analysis of gut microbiota characteristics based on fecal samples lacks stability and standardization across individuals and scenarios. Blood microbiota studies suffer from problems such as unknown sources, interference from low biomass samples, and insufficient sample size, making it difficult to achieve early and accurate assessment of ASD.

Method used

Based on blood metagenomic data from 1,946 standard four-family ASD pedigrees, we screened blood microbial biomarkers associated with ASD, constructed an autism prediction model, used machine learning methods for disease prediction, and employed dual-pathway differential analysis and multi-dimensional statistical control to screen out 12 biomarkers with high specificity and high correlation. Risk assessment was conducted based on the detection status and relative abundance changes of blood microorganisms.

Benefits of technology

It enables early risk assessment and stratification of ASD, providing a sensitive and accurate objective assessment tool that overcomes the subjective limitations of traditional diagnostic methods and has broad prospects for clinical translation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121839140B_ABST
    Figure CN121839140B_ABST
Patent Text Reader

Abstract

The application provides an autism prediction model based on blood microorganisms and application thereof. The prediction model takes blood microorganisms as markers, and predicts the risk of autism of a blood sample through the decrease of a blood microorganism one detection rate or relative abundance, or the increase of a blood microorganism two detection rate or relative abundance. The blood microorganism one comprises Xanthobacteraceae bacterium, Pseudomonas sp.3J6, Rhodococcus equi, Parabacteroides distasonis, Xylella fastidiosa, Enterorhabdus hofmannii, Wolbachia, Yersinia enterocolitica, Borrelia parkeri and Elizabethkingia sp. The blood microorganism two comprises Klebsiella michiganensis and Bdellovibrio sp. The application constructs a model based on blood metagenomic data of 1946 standard four-generation families of autism, fully develops the advantages of blood microorganisms, sensitively and accurately realizes the capture of disease signals, and can be applied to the prediction of autism in the clinic, and realizes the early risk assessment and stratification of diseases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of early disease risk assessment and auxiliary diagnosis technology, specifically involving autism spectrum disorder risk prediction technology based on blood microbial biomarkers, and particularly involving an autism prediction model based on blood microorganisms and its application. Background Technology

[0002] Autism Spectrum Disorder (ASD), also known as autism spectrum disorder, is a neurodevelopmental disorder characterized by impairments in social interaction, communication, and repetitive, stereotyped behaviors. Its onset typically occurs in early childhood. In recent years, the global incidence of ASD has shown a continuous upward trend. This disease not only severely impairs children's physical and mental health and development but also places a heavy burden on affected families and consumes significant social resources, making it a crucial public health issue that urgently needs attention and resolution.

[0003] Currently, the clinical diagnosis of ASD still mainly relies on behavioral assessment methods, including clinical behavioral observation, collection of the child's developmental history, and standardized scale assessment. However, ASD is characterized by early age of onset and high clinical heterogeneity. Its phenotypic spectrum is complex, and the symptom differences between individuals are extremely significant. Furthermore, it overlaps with symptoms of various developmental disorders and mental health conditions, making the results of these behavioral assessments susceptible to interference from factors such as the child's age, language ability, and comorbidities. In addition, behavioral assessment results are highly dependent on the clinician's professional experience and the cooperation of the child and their family. Different clinicians may arrive at different diagnoses for the same individual with ASD, making it difficult to objectively, quantitatively, and reproducibly assess the risk of disease in its early stages. This inevitably introduces a high degree of subjectivity into behavioral assessments, ultimately resulting in a low accuracy rate in the clinical diagnosis of ASD and failing to meet the clinical needs for early screening and intervention.

[0004] Microorganisms, as important environmental influencing factors, are closely related to the occurrence and development of ASD. Among them, the correlation between the gut microbiome and ASD has been reported by many studies and has received widespread attention. However, most studies on gut microbiota use fecal samples for microbial characterization. The process of obtaining fecal samples and the stability of the microbial characteristics in the samples are easily disturbed by external factors such as diet, clinical medication, and living environment. Furthermore, the consistency and standardization of gut microbial characteristics obtained from fecal samples are insufficient across individuals and testing scenarios. These limitations seriously restrict the translation and application of gut microbial characteristics into clinically stable ASD early screening tools.

[0005] Blood, as the core medium for the transport and exchange of substances throughout the human body, contains a microbial composition that can reflect, to some extent, the interactions between the host and the microbial communities in different parts of the body. Various blood microbial-related signals collectively constitute a unique and dynamic microbial genetic information database within the human circulatory system. Numerous studies in clinical trials of different diseases have found that specific blood microbial profiles have good disease differentiation capabilities. Therefore, screening blood microbial biomarkers and constructing ASD risk prediction models based on them has potential clinical value as an objective auxiliary assessment tool for ASD.

[0006] In the study of blood microbiota, blood has been widely considered a sterile environment for centuries. However, with the rapid development and innovation of microbial detection technologies and research methods, numerous studies have clearly confirmed the objective fact that microorganisms exist in blood, even in the blood of healthy individuals, completely overturning the traditional understanding of blood asepticity. Currently, the known sources of microorganisms in blood are complex, including direct invasion of the bloodstream, residual microbial DNA after infection, microbial translocation in tissues such as the intestines, and delivery systems via extracellular vesicles. However, a clear and unified conclusion has not yet been reached regarding the specific sources of blood microorganisms. Furthermore, blood is a typical low-biomass sample, easily affected by background factors such as contamination control and batch effects during microbial detection and analysis, making it difficult to guarantee the stability and reliability of test results. In addition, current research on the correlation between the microbiome and disease generally suffers from small sample sizes and a lack of large-scale clinical sample data to support universally applicable conclusions on microbial characteristics.

[0007] In summary, compared to microorganisms in single tissues such as the gut, blood microbiota possesses the systemic advantage of reflecting the interactions of microbiota throughout the human body. Constructing ASD risk prediction models based on blood microbial biomarkers is an important research direction for achieving objective, early, and accurate ASD assessment. However, current research urgently needs to overcome technical limitations such as unclear sources, interference from low-biomass samples, and insufficient sample size in blood microbial studies. This will allow for the discovery of stable and transferable blood microbial characteristics that can be applied to early risk assessment and prediction of ASD. Therefore, developing a method for screening blood microbial biomarkers and constructing ASD risk prediction models that can effectively address the aforementioned technical problems has significant clinical translational potential and research value, and represents a pressing technical challenge in this field. Summary of the Invention

[0008] The technical problem this invention aims to solve is to overcome the shortcomings of existing technologies and provide a blood microbiome-based autism prediction model and its applications. This invention is based on blood metagenomic data from 1946 standard four-person families with ASD (each family includes two neurodevelopmentally normal parents, one child with ASD, and one neurodevelopmentally normal sibling), providing a large sample size and comprehensive phenotypic data coverage. The study aims to systematically screen blood microbiome biomarkers associated with ASD and further utilize machine learning methods to construct a disease prediction model. This model fully leverages the advantages of blood microbiota, sensitively and accurately capturing disease signals, and can be applied clinically to predict ASD, enabling early risk assessment and stratification.

[0009] To address the aforementioned technical problems, this invention provides an autism biomarker based on blood microorganisms, wherein the autism biomarker includes: *Pseudomonas aeruginosa* (…). Ochrobactrum quorumnocens ), Pseudomonas aeruginosa sp. 3J6 ( Pseudoalteromonas sp.3J6 ), Rhodococcus equi ( Horsetail ), Parabacterium dilatatum ( Parabacteroides distasonis ), Fastidious Trichoderma ( Xylella fastidiosa ), Enterobacter hopterii ( Enterobacter hormaechei ), Wolbachia ( Wolbachia pipientis Yersinia enterocolitica (Yersinia enterocolitica) Yersinia enterocolitica ), pale spirochetes ( Treponema pallidum ), Elizabethany filiis ( Elizabethkingia anopheles ), Klebsiella Michigani ( Klebsiella michiganensis ), Bacteriophage Bdellovibrio ( Bdellovibrio bacteriovorus One or more of the following.

[0010] Based on a general technical concept, the present invention also provides an autism prediction model based on blood microorganisms. The prediction model uses microorganisms in an isolated blood sample as markers and outputs an autism risk score corresponding to the isolated blood sample based on the detection status or relative abundance of the microorganisms.

[0011] The autism risk scoring criteria are as follows: a decrease in the detection rate or relative abundance of blood microbe I leads to an increase in the autism risk score; and / or an increase in the detection rate or relative abundance of blood microbe II leads to an increase in the autism risk score.

[0012] The blood microorganisms include one or more of the following microorganisms: Paleobacterium ( Ochrobacterium who is innocent ), Pseudomonas aeruginosa sp. 3J6 ( Pseudoalteromonas sp.3J6 ), Rhodococcus equi ( Horsetail ), Parabacterium dilatatum ( Parabacteroides distasonis ), Fastidious Trichoderma ( Xylella fastidiosa), Enterobacter hopterii ( Enterobacter hormaechei ), Wolbachia ( Wolbachia pipientis Yersinia enterocolitica (Yersinia enterocolitica) Yersinia enterocolitica ), pale spirochetes ( Treponema pallidum ), Elizabethany filiis ( Elizabethkingia anopheles );

[0013] The blood microorganisms include Klebsiella Michigani ( Klebsiella michiganensis ) and / or Bdellovibrio phage ( Bdellovibrio bacteriovorus ).

[0014] The autism prediction model described above, further, includes a method for constructing the autism prediction model as follows:

[0015] (1) Modeling sample and input feature set: Blood metagenomic data of 1946 standard four-family families were used as modeling sample. The case group consisted of children diagnosed with ASD, and the control group consisted of siblings with normal neurodevelopment in the same family. The input feature set consisted of the characteristic data of blood microbial species, which were the detection status or relative abundance of the blood microorganisms.

[0016] (2) Dividing the modeling samples: The modeling samples are randomly divided into a training set and a test set; the training set is used for model training, parameter learning and feature subset performance evaluation, and the test set is used for independent performance verification of the model and does not participate in the model training process.

[0017] (3) Feature importance assessment: The training set is used for feature training. A recursive feature elimination method with random forest as the base classifier is adopted. The contribution of the blood microorganisms to the ASD classification model is quantitatively assessed by using feature importance score as the quantitative index. Candidate blood microorganisms are included in the model one by one in order of importance from high to low. The model performance is re-evaluated by using 10-fold cross-validation, and a feature subset with stable performance index is selected.

[0018] (4) Model evaluation: The AUC curve of the test set samples is used as the discrimination index. The AUC value ranges from 0 to 1. The closer the AUC value is to 1, the higher the risk level of the corresponding sample is.

[0019] Furthermore, in the aforementioned autism prediction model, the ratio of the training set to the test set is 8:2.

[0020] In the aforementioned autism prediction model, the blood microorganisms, ranked from highest to lowest importance, are: Treponema pallidum (…). Treponema pallidum ), Klebsiella Michigani ( Klebsiella michiganensis ), Pseudomonas alterniflora ( Pseudoalteromonas sp. . 3J6 ), Rhodococcus equi ( HorsetailYersinia enterocolitica (Yersinia enterocolitica) Yersinia enterocolitica ), Parabacterium dilatatum ( Parabacteroides distasonis The importance of ) is second only to Enterobacter hopterii ( Enterobacter hormaechei ), Elizabethany filiis ( Elizabethkingia anopheles ), Fastidious Trichoderma ( Xylella fastidiosa ), quorum sensing Paleobacterium ( Ochrobacterium who is innocent ), Bacteriophage Bdellovibrio ( Bdellovibrio bacteriovorus ), Wolbachia ( Wolbachia of the pipistrelle ).

[0021] In the aforementioned autism prediction model, further, during the process of successively incorporating candidate blood microorganisms into the model in descending order of importance, before each new species is included, its Pearson correlation coefficient with the species already included in the model is calculated; only when the correlation coefficient between the species and any species already included in the model is less than 0.7 is it allowed to be added to the model.

[0022] The autism prediction model described above is further improved by immediately re-evaluating the model performance using 10-fold cross-validation on the training set for each newly added species feature that meets the relevance requirements, and obtaining the stability performance index corresponding to the feature subset. The complete process of randomly dividing the training / test set → adding features step by step → 10-fold cross-validation is repeated 100 times, and the performance data corresponding to each feature combination in each round is summarized.

[0023] Based on a general technical concept, the present invention provides an application of the autism prediction model in the preparation of autism risk assessment products.

[0024] The above application, further, the method of the application includes:

[0025] S1. Extract microbial metagenomic DNA from isolated peripheral venous blood samples and perform high-throughput sequencing to obtain raw blood metagenomic sequencing data.

[0026] S2. The raw blood metagenomic sequencing data is preprocessed to obtain highly reliable microbial sequence data;

[0027] S3. Based on the preprocessed sequencing data, analyze the detection status of blood microorganisms, where detected = 1 and undetected = 0; calculate the relative abundance of each taxa.

[0028] S4. After confirming that there is no redundancy in the features according to the feature selection rules during model construction, organize them into a feature matrix that the model can recognize.

[0029] S5. Using the feature matrix from S4 as input, substitute it into the autism prediction model for prediction. The closer the AUC value of the sample is to 1, the higher the risk score of autism in the corresponding ex vivo peripheral venous blood sample.

[0030] In the above application, the preprocessing further includes:

[0031] S2-1. Perform quality control cleaning and dehumanization treatment on non-human sequences in the sequencing data to obtain high-quality microbial sequences;

[0032] S2-2. Perform species classification annotation and abundance quantification on the high-quality microbial sequences to obtain a species abundance matrix;

[0033] S2-3. The species abundance matrix is ​​sequentially subjected to taxonomic limitation, noise filtering, and multi-dimensional progressive removal of pollutants to construct a highly reliable blood microbial database.

[0034] Further, in the above application, S2-1 includes: using samtools to extract unaligned sequences from the sequencing data that have not been aligned to the human reference genome; using BBduk to perform quality control cleaning on the unaligned sequences, cutting low-quality bases with Q<20 at the end of the sequences, removing sequences with an average quality score lower than 20, and eliminating low-complexity sequences with an average entropy <0.6; then aligning them again to the human reference genome using bowtie2 to eliminate residual human-derived sequences; and combining this with FASTP to complete quality assessment and obtain high-quality microbial sequences.

[0035] Further, in the above application, S2-2 includes: using Kraken2 to perform species classification annotation on the high-quality microbial sequences, combining Bracken to perform probability redistribution of the sequences on the classification tree to estimate species abundance, and generating a species-level sequence count and relative abundance matrix.

[0036] Further, in the above application, S2-3 includes: sequentially performing microbial group range limitation, noise control, and four-step progressive pollutant filtering on the matrix, retaining only bacterial groups, setting the sequence count of species with a relative abundance of less than 0.005 or less than 10 assigned sequence pairs to zero, then sequentially retaining species detected in at least two sequencing batches, species with at least 100 sequences in at least one sample, removing potential pollutants reported in previous studies and species in the same batch with a correlation coefficient > 0.8 with any pollutant, removing species with a detection rate greater than 20%, and finally obtaining a blood microbial database containing highly reliable microbial species.

[0037] Compared with the prior art, the advantages of the present invention are as follows:

[0038] (1) This invention provides an autism prediction model based on blood microbiota. The study relies on a large sample of blood metagenomic data from 1946 standard four-person families (each family includes two neurodevelopmentally normal parents, one child with ASD, and one neurodevelopmentally normal sibling). The sufficient sample size and comprehensive phenotypic data ensure the statistical validity and representativeness of the research results from the data source. Simultaneously, the core research design employs a "autism-neurodevelopmentally normal sibling" pairing within the family, which maximizes the control of confounding factors such as genetic background and early shared environment at the family level, significantly improving the comparability between cases and controls. This makes the selected differential signals more likely to reflect the relevant changes of autism itself, providing a high-quality data foundation that aligns with the essence of the disease for subsequent biomarker screening and model construction. Furthermore, machine learning methods are used to construct a disease prediction model based on this. This model fully leverages the advantages of blood microbiota, sensitively and accurately capturing disease signals, and can be applied clinically to predict ASD, achieving early risk assessment and stratification of the disease.

[0039] (2) This invention provides an autism prediction model based on blood microbiota, which adopts a marker screening method with dual-path complementarity and multi-dimensional statistical control, which has both comprehensiveness and stability: On the one hand, by analyzing and integrating the results of dual-path difference analysis of detection status and relative abundance, it can simultaneously capture two types of disease-related information: whether microorganisms are detected and the changes in abundance after detection, effectively reducing the instability of results caused by single-path analysis; on the other hand, by using conditional logistic regression, including random effects / stratified variables, etc., it strictly controls covariates such as family ID, gender, age, BMI, and ancestry, and obtains a significant marker set after FDR correction. The 12 autism-related blood microbiota markers finally determined have the characteristics of high specificity and high correlation, providing high-quality core input features for model construction.

[0040] (3) This invention provides an autism prediction model based on blood microorganisms, which is the first ASD risk prediction model based on blood microorganisms. Its modeling process has been scientifically designed through multiple stages, which fundamentally ensures the stability, accuracy and cross-scenario generalizability of the model: 12 precisely selected biomarkers are used as core inputs, and feature importance is ranked through recursive feature elimination, and feature redundancy and collinearity are controlled through correlation thresholds; the evaluation strategy of training set / test set partitioning, cross-validation and repeated sampling is adopted to effectively reduce randomness and overfitting risk; the final model is determined by combining AUC as the core indicator with accuracy, and the model is validated in an independent Chinese cohort, which fully verifies the generalizability of the model; at the same time, this invention protects the complete technical route and model implementation method, rather than a single algorithm, and has stronger application flexibility and technical scalability.

[0041] (4) This invention provides an application of an autism prediction model in the preparation of a test kit. The prediction model fully utilizes the signal advantages of blood microorganisms themselves, and can sensitively and accurately capture autism-related signals, realizing an objective quantitative assessment of ASD, breaking through the subjective limitations of traditional autism assessment methods. The model can be directly applied to the prediction of ASD in clinical practice, realizing early risk assessment and stratification of the disease, providing a reliable objective detection method for early detection and early intervention of autism, and has important practical application value for clinical screening and disease management of autism, with broad prospects for clinical translation. Attached Figure Description

[0042] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0043] Figure 1 This is a flowchart for noise and pollution control.

[0044] Figure 2 The results of blood microbial screening are shown in Figure A, which represents blood microorganisms screened using the detection status method; and Figure B, which represents blood microorganisms screened using the relative abundance of species.

[0045] Figure 3 The figure shows the correlation analysis results between blood microbial signals and various clinical phenotypic domains in the human body. A and B in the figure are the correlation pairs corresponding to p-value < 0.05 and q-value < 0.05 after FDR correction, respectively.

[0046] Figure 4 The importance of 12 blood microbial biomarkers is statistically analyzed.

[0047] Figure 5 This is the result of sensitivity verification. Detailed Implementation

[0048] The present invention will be further described below with reference to specific preferred embodiments, but this does not limit the scope of protection of the present invention.

[0049] The materials, reagents, and instruments used in the following examples are all commercially available. Unless otherwise specified, the experimental methods used in the following examples are conventional methods in the art.

[0050] This invention provides a disease prediction model based on large-scale metagenomic data to identify ASD-related blood microbial biomarkers and construct a model to assist in ASD risk assessment and stratification. This invention systematically identifies ASD-related blood microbial biomarkers from metagenomic data and provides data support for the ASD prediction model. The blood microbial biomarkers are obtained through the following screening method:

[0051] (1) Data preparation:

[0052] Discovery Cohort: This study used peripheral blood samples from 1946 standard four-person families. Each family included two neurodevelopmentally normal parents, one child with ASD, and one neurodevelopmentally normal sibling. The child with ASD within the same family was defined as the case group, and their neurodevelopmentally normal sibling as the control group. A pedigree-matched study design was employed to control for confounding factors such as genetic background and early shared environment at the family level, thereby reducing the interference of uncontrollable inter-individual differences on the comparison results.

[0053] Validation cohort: A cohort consisting of 743 ASD patients and 729 neurodevelopmentally normal children was included as an independent validation set to examine the stability and generalizability of blood microbial signals.

[0054] Both cohorts collected peripheral blood samples and extracted DNA. Whole-genome sequencing was performed based on the DNA from peripheral blood to obtain the raw data required for metagenomic analysis. Sample consistency verification and sequencing quality assessment were completed before entering the microbial signal mining stage to ensure the reliability of subsequent microbial feature identification and comparative analysis.

[0055] (2) Data Preprocessing: To ensure consistency in data processing between the discovery and validation cohorts, minimize systemic bias caused by differences in processing procedures, and thereby improve the comparability of data from different sources and enhance the robustness and reproducibility of conclusions, a unified standardized preprocessing procedure was adopted for whole-genome sequencing data of ASD patients and normal children. The specific steps are as follows:

[0056] 2.1 Use samtools to extract non-mapping reads from whole-genome sequencing data as potential microbial sequence inputs.

[0057] 2.2 The unaligned sequences were quality controlled and cleaned using BBduk, including cutting low-quality bases (Q<20) at the ends of the sequences, removing sequences with an average quality score of less than 20, and removing low-complexity sequences (average entropy <0.6, sliding window length of 50, k-mer length of 5) to reduce the influence of sequencing noise and artifact signals.

[0058] 2.3. The sequences that have passed quality control are further processed to remove human sequences. This involves re-aligning the sequences to the human reference genome using bowtie2 to remove as many residual human sequences as possible, and using FASTP to assess the quality of the retained sequences to ensure that the obtained sequences are high-quality microbial sequence data.

[0059] 2.4. A standardized species annotation and quantification process was employed to analyze the microbial composition. Specifically, Kraken2 was used for classification annotation, and Bracken was combined to perform probability redistribution of sequences on the taxonomic tree to re-estimate species abundance, thereby generating species-level sequence counts and relative abundance matrices for subsequent statistical analysis and predictive model construction. Species classification annotation was performed on the raw blood metagenomic sequencing data to generate initial species-level sequence counts and relative abundance matrices. This resulted in an initial species annotation dataset containing 12,146 microbial species.

[0060] (3) Noise and contamination control: Given that peripheral blood is a low-biomass sample with low microbial genetic material content, sparse sequence signals, and is easily affected by background noise from reagents, environment, and batches, occasional detections and misclassifications are likely to occur, thus affecting the authenticity and reproducibility of species annotation results. To ensure that subsequent biomarker screening and model construction mainly reflect reliable disease-related signals, after obtaining the species-level sequence count and relative abundance matrix, further denoising and decontamination processing is implemented to reduce annotation errors and remove potential contamination signals as much as possible.

[0061] For detailed steps, please refer to [link / reference]. Figure 1 :

[0062] 3.1 Limitation of Microbial Groups: Based on the characteristics of the microbial composition of blood samples, only bacteria were retained, while archaea, viruses, and eukaryotes were removed, reducing the number of microorganisms to 10,841. This reduced the influence of irrelevant taxa and groups susceptible to exogenous contamination.

[0063] 3.2 Noise Control: Species with a relative abundance below 0.005 or fewer than 10 assigned sequence pairs were classified as undetectable, and their sequence counts were set to zero, reducing the number of microorganisms to 595. This was to reduce noise interference caused by random misassignment of low-abundance species.

[0064] 3.3 Four-Step Progressive Pollutant Filtration: Based on the above operations, perform the following pollutant filtration operations sequentially to eliminate polluted signals and invalid background signals:

[0065] 3.3.1 Batch reproducibility filtering: Only species detected in at least two sequencing batches are retained, while species appearing only in a single batch are removed, ensuring reproducibility across batches. The number of microorganisms was reduced to 453.

[0066] 3.3.2 Library background noise filtering: Only species with at least 100 assigned sequences in at least one sample are retained, excluding artifact signals related to library background noise levels. The number of microorganisms was reduced to 366.

[0067] 3.3.3 Filtering of potential pollutants and associated species: Potential pollutants were removed based on the list of pollutants reported in previous studies, reducing the number of microorganisms to 197. Species highly correlated with any pollutant (correlation coefficient > 0.8) within the same batch were also removed to reduce interference from common sources within the same batch. The number of microorganisms was reduced to 117.

[0068] 3.3.4 High Detection Rate Background Signal Filtering: Combining the low detection rate characteristics of species in blood metagenomic data, species with a detection rate greater than 20% are removed to avoid the influence of background signals with high detection rates that lack discriminative power. The final dataset contains 100 microorganisms.

[0069] 3.4 Screening of High-Reliability Strains: After the above noise reduction, taxonomic limitation and four-step decontamination process, a highly reliable blood microbial database is obtained for subsequent biomarker screening and downstream inference analysis, thereby improving the reliability and robustness of the results.

[0070] (4) Screening of blood microbial markers related to ASD.

[0071] 4.1 In order to screen blood microbial biomarkers related to ASD, this invention conducts differential analysis from two complementary pathways: "detection status" and "relative abundance", in order to simultaneously capture information on "detection status" and "change in abundance after detection" in low biomass blood metagenomic data.

[0072] Path 1: Difference Analysis Based on Species Detection Status: Using species detection status (whether or not it is detected) as the outcome variable, a conditional logistic regression model stratified by pedigree number is used to conduct inter-group comparisons. Covariates such as sex, age, body mass index, and ancestry are adjusted to capture inter-group differences in species "detection status." A conditional logistic regression model stratified by pedigree ID is used, with adjustments made for covariates such as sex, age, BMI, and ancestry. The model results are then corrected using FDR (Free Decision Linear Regression) to identify species with a significance threshold q < 0.05.

[0073] See the screening results Figure 2 In the figure, A represents blood microorganisms screened using the detection status method, identifying a total of 12 significantly different species: 10 species (respectively...) Ochrobactrum quorumnocens , Pseudoalteromonas sp. 3J6 , Horsetail , Parabacteroides distasonis , Xylella fastidiosa , Enterobacter Hormaechei, Wolbachia pipientis , Yersinia enterocolitica , Treponema pallidum and Elizabethkingia anopheles The detection rate was significantly lower in children with ASD (blue dots in the figure, with negative coefficients, -log). 10 (The higher the q-value, the stronger the significance). Two types (respectively...) Klebsiella michiganensis , Bdellovibrio bacteriovorous The detection rate of this substance is significantly higher in children with ASD (orange dots in the figure, with positive coefficients).

[0074] Pathway Two: Difference Analysis Based on Relative Abundance of Species: Using relative abundance of species as the outcome variable, a multivariate association analysis model is employed for inter-group comparisons. After simultaneously adjusting for covariates such as sex, age, body mass index, and ancestry, pedigree ID is included as a random effect term to effectively control for intra-family correlations and capture inter-group differences in species abundance changes after detection. The MaAsLin2 multivariate association model is used, adjusting for covariates such as sex, age, BMI, and ancestry, and incorporating pedigree ID into the random effect term to control for intra-family correlations. The model results are then corrected using FDR (Free Decision Mapping) to identify species with a significance threshold q < 0.05 after correction.

[0075] See the screening results Figure 2 In the figure, B represents blood microorganisms screened using relative abundance of species. Among them, 9 species showed significantly reduced abundance in ASD, namely... Pseudoalteromonas sp. 3J6 , Horsetail , Parabacteroides distasonis , Xylella fastidiosa , Enterobacter hormaechei , Wolbachia pipientis , Yersinia enterocolitica , Treponema pallidum and Elizabethkingia anopheles Two species showed significantly increased abundance in ASD, and both were... Klebsiella Michigan and Bdellovibrio bacteriovorus .

[0076] 4.2 Multiple Validation Correction and Candidate Species Identification: Multiple validation correction was applied to the results of both analysis paths to control the false positive rate. Based on the corrected significance threshold, candidate differentially abundant species were identified for each path. Comparison of the results from the two paths revealed a high degree of consistency in the differentially abundant species, and the 12 species identified by the "Detection Status Path" completely covered the 11 species identified by the "Relative Abundance Path."

[0077] 4.3 Integration of candidate species to form a biomarker set: The candidate differential species obtained from the two analysis paths are integrated (preferably the union set is selected) to finally determine 12 blood microbial biomarkers that are significantly associated with ASD (i.e., all 12 species identified by the "detection status path"), forming a blood microbial biomarker set associated with autism spectrum disorder (ASD). This set will be used for subsequent feature construction and prediction model training to improve the robustness, consistency and transferability of biomarker screening across batches.

[0078] Experiment 1: Assessment of the association between blood microbial species and clinical phenotypes of ASD.

[0079] To further assess the association between the identified blood microbial species and the clinical phenotype of ASD, we systematically categorized the phenotypic information of ASD samples and conducted correlation analysis. The specific steps were as follows:

[0080] 1. Systematic Classification of ASD Clinical Phenotypic Characteristics: The phenotypic information of ASD samples is divided into eight major categories, each containing 3 to 9 sub-phenotypes, constructing a systematic phenotypic analysis system. These categories are: host basic characteristics, pregnancy or perinatal related characteristics, diet-related characteristics, disease and medication-related characteristics, adaptive function, cognitive function, behavioral characteristics, and social communication characteristics.

[0081] 2. Determine the dual-characteristic association analysis strategy: In view of the low biomass and sparse distribution characteristics of blood metagenomic data, in order to improve the robustness of association analysis results, microbial-phenotype association analysis was carried out based on two characterization methods: microbial detection status and microbial relative abundance. The results of the two analysis strategies were compared and analyzed to reduce the risk of false positives introduced by low abundance fluctuations, zero expansion and batch noise.

[0082] 3. Select the appropriate model based on phenotypic type and control for confounding factors: Select the corresponding statistical model based on the phenotypic data type, while including covariates and controlling for family stratification effects, specifically:

[0083] (1) Continuous phenotype: For scores of adaptive function, cognitive function, etc., a linear regression model under the linear model framework is used to estimate the association between "microbe-phenotypic characteristics".

[0084] (2) Binary phenotype: For gender and whether or not exposed to a certain type of drug / disease factor, the Logistic regression model is used to conduct association analysis; if there is a significant class imbalance in the binary phenotype, which leads to estimation bias or separation problems in the conventional Logistic regression, Firth correction is introduced to improve the stability and interpretability of parameter estimation.

[0085] (3) Confounding control: all models included sex, age and ancestry as covariates, and family lineage as a stratified variable to control for interference from genetic background and common environment.

[0086] 4. Comprehensive association test and multiple test correction: The above-mentioned microbial-phenotype association test is carried out one by one between the pre-processed microbial species set and all candidate phenotypes; the P-values ​​obtained from the test are corrected for multiple tests using the FDR method to control the overall false positive rate.

[0087] 5. Analysis and summarization of significant association results: The significant association results of microorganism-phenotype obtained after correction are visualized and enrichment domains are summarized to identify key phenotypic dimensions and their corresponding microbial species combinations that are significantly associated with blood microbial signals in different phenotypic domains. This provides data support and theoretical basis for subsequent ASD risk assessment based on phenotypic stratification, expansion of prediction models, and interpretation of clinical significance.

[0088] Figure 3 The association analysis results between blood microbial signals and various clinical phenotypic domains of the human body are divided into two dimensions: significant associations with adaptive and cognitive functional domains and nominal significant associations with other phenotypic domains. The robustness of the association signals is verified through a dual characterization strategy. Figure 3 The above shows the original p-values ​​< 0.05. Figure 3 The following is the q value after FDR correction, which is <0.05.

[0089] Figure A shows a significant association between the adaptive and cognitive functional domains (q < 0.05). This dimension represents the core association region between blood microbial signals and ASD clinical phenotypes; no association signals of this level were observed in other phenotypic domains. All associations are negative; increased blood microbial signals (more detectable microorganisms, higher relative abundance) correlate with lower levels of adaptive and cognitive function. The cognitive functional domain shows the most prominent association, with a significantly higher association signal than the adaptive functional domain. Among these, IQ-related indicators (verbal IQ, nonverbal IQ, and full-scale IQ) exhibit the strongest and most consistent association signals under both microbial detection status and relative abundance analysis strategies, with no contradictory results. Several blood microorganisms show significant negative correlations with IQ indicators; typical species include... Pseudoalteromonas sp. 3J6 , Wolbachia of pipientia , Yersinia enterocolitica , Elizabethkingia anopheles , Klebsiella Michigan That is, the enhanced signals of these species all correspond to a decrease in the IQ index.

[0090] Figure B shows that only a few nominally significant associations (p<0.05) exist for the remaining phenotypic domains. Analysis of all phenotypic domains, excluding adaptive / cognitive functions, including host basic characteristics, perinatal pregnancy characteristics, and dietary characteristics, revealed no significant associations with q<0.05, and only a few nominally significant associations at the p<0.05 level. This indicates that blood microbial signals have a weak influence on these phenotypic domains and are not a core influencing factor. Representative associations are: Parabacteroides distasonis The signal was negatively correlated with the total score of the Social Communication Questionnaire (SCQ), and the enhanced signal in this species corresponded to a poorer performance of the social communication phenotype. Yersinia enterocolitica The signal enhancement was positively correlated with current casein and gluten intake, with higher levels of casein and gluten intake corresponding to enhanced signaling in this species.

[0091] The differential and correlation analyses obtained from two characterization strategies—microbial detection status and relative abundance—showed a high degree of consistency in overall trends. For example, the IQ index exhibited the strongest association under both strategies, and the association direction of core species remained unchanged. These results confirm that the association between blood microbial signals and the ASD clinical phenotype domain is not random; the signals are stable and reproducible, providing a reliable data analysis foundation for subsequent research into the biological mechanisms underlying this association.

[0092] Experiment 2: Verify the contribution of 12 blood microbial markers.

[0093] A random forest (RF) model constructed using these 12 taxa was used to distinguish ASD from its neurologically normal siblings. (See attached results.) Figure 4 .

[0094] Figure 4 The importance statistics for 12 blood microbial biomarkers are presented. Feature importance analysis using random forest shows that these 12 candidate biomarkers have significantly different contributions to the model's discrimination, among which... Treponema pallidum and Klebsiella michiganensis The most important is [the first], followed by [the second]. Pseudoalteromonas sp. 3J6 , Horsetail , Yersinia enterocolitica , Parabacteroides distasonis This suggests that key species have a higher weight in model decision-making.

[0095] Example 1

[0096] A blood microbiome-based autism prediction model uses blood microbes as markers to predict the risk of autism in blood samples by a decrease in the detection rate or relative abundance of blood microbe I and an increase in the detection rate or relative abundance of blood microbe II.

[0097] The blood microbiome includes one or more of the following microorganisms: Ochrobacterium who is innocent, Pseudoalteromonas sp.3J6 , Horsetail , Parabacteroids discordant , Xylella fastidiosa , Enterobacter hormaechei , Wolbachia pipientis , Yersinia enterocolitica , Treponema pallidum , Elizabethkingia anopheles ;

[0098] The blood microorganisms include Klebsiella michiganensis and / or Bdellovibrio bacteriovorous .

[0099] Specific model building methods include:

[0100] (1) Modeling sample source: Blood metagenomic data of 1946 standard four-family families. The case group consisted of children diagnosed with ASD, and the control group consisted of siblings with normal neurodevelopment in the same family. Blood microbiota 1 and blood microbiota 2 were input into the feature set.

[0101] (2) The modeling samples are randomly divided into a training set (80%) and a test set (20%): the training set is used for model training, parameter learning and feature subset performance evaluation; the test set is used for independent performance verification of the model and does not participate in the model training process.

[0102] (3) Feature Importance Assessment: The training set was used for feature training. The recursive feature elimination (RFE) method with random forest as the base classifier was used to quantitatively assess the contribution of 12 candidate species to the ASD classification model and generate a species importance ranking. Species were included in the model one by one from high to low importance. Before each new species was included, its Pearson correlation coefficient with the species already included in the model was calculated. Only when the correlation coefficient between the species and any species already included in the model was less than 0.7 was it allowed to be included in the model, so as to avoid duplicate information input and model instability caused by highly correlated features.

[0103] During the feature inclusion process, the following steps are performed to ensure the stability of performance estimation: For each new species feature that meets the relevance requirements, the model performance is immediately re-evaluated using 10-fold cross-validation in the training set to obtain the stability performance index corresponding to the feature subset; To reduce the randomness caused by random partitioning and control the risk of overfitting, the complete process of "randomly partitioning the training / test set → progressively adding features → 10-fold cross-validation" is repeated 100 times, and the performance data corresponding to each feature combination in each round is summarized.

[0104] (4) Model evaluation: The area under the receiver operating characteristic curve (AUC) is used as the main discriminant. The AUC value ranges from 0 to 1. The closer the AUC value is to 1, the stronger the discriminant power of the model for ASD-related target risks, that is, the more accurately the model can identify high-risk samples, and the higher the target risk level of the corresponding samples.

[0105] Example 2

[0106] An application of the autism prediction model of Example 1 in autism prediction. It was applied to an independent Chinese cohort. The above analysis was performed using packages such as randomForest (v4.6-147) and pROC (v1.15.3) in R software.

[0107] Its application methods include:

[0108] (1) Microbial metagenomic DNA was extracted from peripheral venous blood samples of the subjects and high-throughput sequencing was performed to obtain raw blood metagenomic sequencing data.

[0109] (2) According to the aforementioned data preprocessing, noise and contamination control process, the raw data is preprocessed: low abundance / low sequence number noise is filtered out (relative abundance <0.005 or sequence log number <10 is judged as not detected); only bacterial groups are retained and non-bacterial groups are removed; quality control steps such as batch repeatability, library background noise, and contaminant filtration are completed to obtain highly reliable microbial sequence data.

[0110] (3) Based on the preprocessed sequencing data, the 12 ASD-related blood microbial taxa defined by the quantitative analysis model were analyzed. Ochrobactrum quorumnocens , Pseudoalteromonas sp. 3J6 , Horsetail , Parabacteroides distasonis , Xylella fastidiosa , Enterobacter hormaechei , Wolbachia pipientis , Yersinia enterocolitica , Treponema pallidum , Elizabethkingia anopheles , Klebsiella michiganensis , Bdellovibrio bacteriovorous ): Extract the detection status of each taxonomic group (detected = 1, not detected = 0); calculate the relative abundance of each taxonomic group; according to the feature selection rules during model construction (Pearson correlation coefficient < 0.7), after confirming that there is no feature redundancy, organize them into a feature matrix that the model can recognize.

[0111] (4) Open the R software and load the pre-trained random forest model file (built based on the randomForest v4.6-147 package); take the feature matrix prepared in step 2 as input and substitute it into the model for prediction.

[0112] The AUC value predicted for this sample was calculated using the pROC v1.15.3 package. The AUC value ranges from 0 to 1. The closer the AUC value is to 1, the stronger the model's ability to discriminate the target risk related to ASD. In other words, the model can more accurately identify high-risk samples, and the higher the target risk level of the corresponding samples.

[0113] Experiment 3: Verify the sensitivity of the prediction model.

[0114] Figure 5 The results show the sensitivity verification. The figure shows that the model's AUC for the training set samples is 0.95, indicating that the model can accurately distinguish between ASD and normal individuals in the modeled population; the AUC on the internal test set is 0.88, proving that the model maintains high discriminative ability in independent samples within the internal set; and the AUC on the independent Chinese cohort is 0.83, indicating good cross-population generalization and stable prediction in non-modeled populations. This demonstrates that the predictive model based on blood microbial biomarkers can accurately determine whether a person has autism or predict the risk of autism in a subject, and possesses certain cross-population generalization and application stability, making it suitable for assisting clinical risk assessment, early screening, and risk stratification of autism.

[0115] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make many possible variations and modifications to the technical solutions of the present invention using the methods and techniques disclosed above, or modify them into equivalent embodiments with equivalent changes, without departing from the spirit and technical essence of the present invention. Therefore, any simple modifications, equivalent substitutions, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solutions of the present invention shall still fall within the protection scope of the technical solutions of the present invention.

Claims

1. An autism prediction model based on blood microbiota, characterized in that, The prediction model uses blood microorganisms in an ex vivo blood sample as markers and outputs an autism risk score corresponding to the ex vivo blood sample based on the detection status or relative abundance of the blood microorganisms. The autism risk scoring criteria are as follows: a decrease in the detection rate or relative abundance of blood microbe I and an increase in the detection rate or relative abundance of blood microbe II result in an increased autism risk score. The blood microorganisms include one or more of the following: Ochrobactrum quorumnocens, Pseudoalteromonas sp.3J6, Prescottella equi, Parabacteroides distasonis, Xylella fastidiosa, Enterobacter hormaechei, Wolbachia pipientis, Yersinia enterocolitica, Treponema pallidum, and Elizabethkingia anophelis; The blood microorganisms include Klebsiella michiganensis and / or Bdellovibrio bacteriovorus.

2. The autism prediction model according to claim 1, characterized in that, The method for constructing the autism prediction model includes: (1) Modeling sample and input feature set: Blood metagenomic data of a standard four-family autism pedigree were used as the modeling sample. The case group consisted of children diagnosed with autism, and the control group consisted of siblings with normal neurodevelopment in the same family. The input feature set was the detection status or relative abundance of the blood microorganisms. (2) Dividing the modeling samples: The modeling samples are randomly divided into a training set and a test set; (3) Feature importance assessment: The training set is used for feature training. A recursive feature elimination method with random forest as the base classifier is adopted. The feature importance score is used as a quantitative indicator to quantitatively assess the contribution of the blood microorganisms to the autism classification model. Candidate blood microorganisms are included in the model one by one in order of importance from high to low. The model performance is re-evaluated by 10-fold cross-validation, and a feature subset with stable performance indicators is selected. (4) Model evaluation: The AUC value of the test set samples is used as the discrimination index. The AUC value ranges from 0 to 1. The closer the AUC value is to 1, the higher the risk score of the corresponding sample is for autism.

3. The autism prediction model according to claim 2, characterized in that, The ratio of the training set to the test set is 8:

2.

4. The autism prediction model according to claim 2, characterized in that, The blood microorganisms, ranked from most important to least important, are: Treponema pallidum, Klebsiella michiganensis, Pseudoalteromonas sp. 3J6, Prescottella equi, Yersinia enterocolitica, Parabacteroides distasonis, Enterobacter hormaechei, Elizabethkingia anophelis, Xylella fastidiosa, Ochrobactrum quorumnocens, Bdellovibrio bacteriovorus, and Wolbachia apipatitidis.

5. The autism prediction model according to claim 2, characterized in that, In the process of incorporating candidate blood microorganisms into the model one by one from high to low importance, before each new species is included, its Pearson correlation coefficient with the species already included in the model is calculated; only when the correlation coefficient between the species and any species already included in the model is less than 0.7 is it allowed to be added to the model.

6. The autism prediction model according to claim 5, characterized in that, For each new species feature that meets the relevance requirements, the model performance is immediately re-evaluated using 10-fold cross-validation in the training set to obtain the stability performance index corresponding to the feature subset. The complete process of randomly dividing the training / test set → adding features step by step → 10-fold cross-validation is repeated 100 times, and the performance data corresponding to each feature combination in each round is summarized.

7. The application of the autism prediction model according to any one of claims 1 to 6 in the preparation of an autism risk assessment product.

8. The application according to claim 7, characterized in that, The application method includes: S1. Extract microbial metagenomic DNA from isolated peripheral venous blood samples and perform high-throughput sequencing to obtain raw blood metagenomic sequencing data. S2. The raw blood metagenomic sequencing data is preprocessed to obtain highly reliable blood microbial sequence data; S3. Based on the preprocessed sequencing data, analyze the detection status of blood microorganisms, where detected = 1 and not detected = 0; calculate the relative abundance of each taxa. S4. After confirming that there is no redundancy in the features according to the feature selection rules during model construction, organize them into a feature matrix that the model can recognize. S5. Using the feature matrix from S4 as input, substitute it into the autism prediction model for prediction. The closer the AUC value of the sample is to 1, the higher the risk score of autism in the corresponding ex vivo peripheral venous blood sample.

9. The application according to claim 8, characterized in that, The preprocessing includes: S2-1. Perform quality control cleaning and dehumanization treatment on non-human sequences in the sequencing data to obtain high-quality microbial sequences; S2-2. Perform species classification annotation and abundance quantification on the high-quality microbial sequences to obtain a species abundance matrix; S2-3. The species abundance matrix is ​​sequentially subjected to taxonomic limitation, noise filtering, and multi-dimensional progressive removal of pollutants to construct a highly reliable blood microbial sequence database.

10. The application according to claim 9, characterized in that, S2-1 includes: using samtools to extract unaligned sequences from the sequencing data that have not been aligned to the human reference genome; using BBduk to perform quality control cleaning on the unaligned sequences; cutting low-quality bases with Q<20 at the end of the sequences; removing sequences with an average quality score of less than 20 and eliminating low-complexity sequences with an average entropy <0.6; then aligning them again to the human reference genome using bowtie2 to eliminate residual human-derived sequences; and combining this with FASTP to complete quality assessment and obtain high-quality microbial sequences. S2-2 includes: using Kraken2 to perform species classification annotation on the high-quality microbial sequences, combining Bracken to perform probability redistribution of the sequences on the classification tree to estimate species abundance, and generating a species-level sequence count and relative abundance matrix. S2-3 includes: sequentially performing microbial group range limitation, noise control, and four-step progressive pollutant filtering on the matrix, retaining only bacterial groups, setting the sequence count of species with relative abundance below 0.005 or less than 10 assigned sequence pairs to zero, then sequentially retaining species detected in at least two sequencing batches, species with at least 100 sequences in at least one sample, removing potential pollutants reported in previous studies and species in the same batch with a correlation coefficient > 0.8 with any pollutant, removing species with a detection rate greater than 20%, and finally obtaining a blood microbial database containing highly reliable microbial species.

Citation Information

Patent Citations

  • CN105652016A

  • CN120967020A