Two mutually competing flora as core microbiome signature for human disease

Through genome-centric microbiome-wide correlation analysis and assembly of high-quality draft genomes, identifying and organizing them into robust bacterial flora, the limitations of identifying disease-related microbiome characteristics in the prior art are solved, and effective prediction and risk management of a variety of chronic diseases are achieved.

CN119948567APending Publication Date: 2025-05-06RUTGERS THE STATE UNIV +1

Patent Information

Application Number
CN202380050273.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-04-25
Filing Date
2023-04-25
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing metagenomic data analysis methods have limitations in identifying disease-related microbiome characteristics, especially the fact that the function of bacterial genes is limited by their vector ecological behavior.

Method used

Genome-centric microbiome-wide association analysis (MWAS) was used to assemble high-quality draft genomes (MAGs) to identify basic components and important microbiome characteristics of the intestinal ecosystem and organize these characteristics to be robust bacterial flora for analysis.

Benefits of technology

The microbiome characteristics associated with a variety of chronic diseases were effectively identified, especially through balanced regulation of seesaw network-like bacterial flora, disease risk management is achieved, and good predictive classification capabilities are shown in multiple independent metagenomic datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948567A_ABST
    Figure CN119948567A_ABST
Patent Text Reader

Abstract

Methods and systems are disclosed for determining a disease state by acquiring a first plurality of nucleic acid sequences of genomic DNA from a sample of the intestinal tract of a subject. A first plurality of genomic abundance values for a first plurality of intestinal bacteria and a second plurality of genomic abundance values for a second plurality of at least 20 intestinal bacteria species are determined from the nucleic acid sequence. A model is applied to at least the first plurality of genomic abundance values and the second plurality of genomic abundance values, or one or more combinations thereof, thereby determining the disease state of the subject as an output of the model.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. Provisional Patent Application Serial No. 63 / 334,503, filed on April 25, 2022, which is hereby incorporated by reference in its entirety.

[0003] Brief Description of Sequence Listing

[0004] Accompanying this submission is a "Sequence Listing XML" file named ST26_126146_5001_WO.XML containing SEQ ID NOs: 1-99534, created on April 22, 2023, and 2,491,699 kilobytes in size, in compliance with 37 CFR §§1.831 to 37 CFR §§1.835, submitted by mail as an XML file on a read-only compact disc (DVD). 37 CFR §1.835(a)(1). The Sequence Listing XML is hereby incorporated by reference in its entirety. Background Art

[0005] After millions of years of co-evolution, humans have formed a robust symbiotic relationship with the gut microbiome [6, 7]. The gut microbiome serves as an essential organ that supports host homeostasis, including metabolism, immunity, development, and behavior [8]. The reduction or loss of these health-related functions when the gut microbiome is dysbiotic has been identified as a risk factor for many chronic diseases, including type 2 diabetes mellitus (T2DM) [9-11]. A series of microbiome-wide association studies (MWAS) have attempted to identify microbiome features (including features such as genes, pathways, and taxa) associated with disease phenotypes as biomarkers in metagenomic datasets [12-15].

[0006] However, the gene-centric metagenomic data analysis methods currently used in most MWAS projects have significant limitations. They rely on existing databases to classify and functionally annotate individual genes and exclude novel genes from downstream analyses [12, 14, 15]. More importantly, these methods treat individual genes as independent units, ignoring the fact that the function of bacterial genes is constrained by the ecological behavior of their vectors. For example, two competing bacterial strains may encode the same functional gene. However, the two copies of the same gene will contribute differently to the community-wide expression of the gene's function because their vectors follow opposite growth trajectories within the intestinal habitat

[16] . Grouping the same functional genes from different bacterial strains into a single unit (such as a pathway or species) will mask or neutralize these opposing changes and lead to spurious correlations [5]. Summary of the Invention

[0007] Given the above background, a genome-centric MWAS was employed, in which high-quality draft genomes assembled from metagenomic datasets (metagenomic assembled genomes, MAGs) were used as the basic building blocks of the gut ecosystem and the most important microbiome features for disease phenotype association analysis. MAGs are not independent microbiome features. They engage in ecological interactions such as competition or cooperation with each other and organize themselves into higher-level structures called “communities” [5]. Each community may be a functional unit in the gut ecosystem, and its members may have widely different taxonomic backgrounds but exhibit co-abundance behavior. Community has been shown to be positively or negatively correlated with disease phenotypes

[17] . Therefore, MAGs and their community-level clustering are ecologically meaningful features that can be used to identify microbiome features associated with human diseases.

[0008] Gut microbiome dysbiosis is associated with an increased risk of multiple human diseases [1, 2]. To date, many efforts have focused on identifying gene- or taxonomic-based microbial signatures as disease biomarkers. However, such signatures remain controversial [3, 4] and overlook the fact that gut bacterial strains are not independent but form coherent functional groups (also known as "microbiota") that interact with each other and influence host health [5]. Therefore, the examples may suggest the search for strain-level microbiome signatures in the form of robust microbiota, through which the gut microbiome provides stable health-related functions to the host. The examples herein may show that two competing bacterial microbiota are organized as two ends of a robust, stable seesaw-like network, and their abundance is associated with multiple chronic diseases. During a 3-month high-fiber intervention and 1-year follow-up in patients with type 2 diabetes mellitus (T2DM), 141 of a total of 1,845 metagenomic assembled genomes (MAGs) formed two competing microbiotas because they had a stable ecological relationship while the gut microbiome underwent profound structural changes. Fifty genomes in cluster 1 contained more genes for plant polysaccharide degradation and butyrate production, while 91 genomes in cluster 2 included nearly all virulence or antibiotic resistance gene carriers predicted from 1,845 MAGs. A random forest regression model revealed that the abundance distribution of 141 genomes correlated with 41 of 43 bioclinical parameters. Using these 141 MAGs as reference genomes, this seesaw network was not only detectable but also facilitated machine learning models to predict and classify cases and controls for nine diseases in 12 independent metagenomic datasets from 1,874 participants across ethnic and geographic regions. These diseases include type 2 diabetes, atherosclerosis, hypertension, cirrhosis, inflammatory bowel disease, colorectal cancer, ankylosing spondylitis, schizophrenia, and Parkinson's disease. These two seesaw-like microbiota may serve as core microbiomes, and their balance can be modulated for disease risk management.

[0009] Therefore, one aspect of the present disclosure provides a method for determining one of a plurality of disease states of a subject and a system for performing the disclosed method. The method includes obtaining, in electronic form, a first plurality (e.g., at least 100,000) of nucleic acid sequences of a first genomic DNA of a first biological sample from the intestine of the subject at a computer system having at least one processor and a memory storing one or more programs executed by the one or more processors. The method also includes determining a first plurality of genomic abundance values ​​from the first plurality of nucleic acid sequences, the first plurality of genomic abundance values ​​including, for each corresponding intestinal bacterial species in the first plurality (e.g., at least 20) intestinal bacterial species, a first corresponding abundance value of the genome of the corresponding intestinal bacterial species in the first plurality of intestinal bacterial species in the first biological sample; and a second plurality of genomic abundance values ​​including, for each corresponding intestinal bacterial species in the second plurality (e.g., at least 20) intestinal bacterial species, a first corresponding abundance value of the genome of the corresponding intestinal bacterial species in the second plurality of intestinal bacterial species in the first biological sample. The method also includes applying, by at least one processor, the model to at least the first plurality of genomic abundance values ​​and the second plurality of genomic abundance values, or one or more combinations thereof, to thereby determine a disease state of the subject as an output of the model.

[0010] As disclosed herein, any embodiment disclosed herein may be applied to any other aspect, where applicable.

[0011] For those skilled in the art, further aspects and advantages of the present disclosure will become apparent from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be appreciated, the present disclosure is capable of other and different embodiments, and its several details are capable of modification in various obvious respects, all without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature and not restrictive.

[0012] Thus, one aspect of the present invention provides a method of identifying a panel of intestinal microorganisms at a computer system having one or more processors and a memory storing one or more programs executed by the one or more processors.

[0013] In some embodiments, the method includes electronically obtaining, for each respective subject in a first plurality of subjects having a first state having a biological characteristic, a corresponding plurality of genomic abundance values, the plurality of genomic abundance values ​​including, for each respective gut microbe in a plurality of gut microbes, a corresponding value for the abundance of a genome of the respective gut microbe in a biological sample from the gut of the respective subject.

[0014] In some such embodiments, for each respective subject in the first plurality of subjects, the biological sample from the intestine of the respective subject is a stool sample.

[0015] In some such embodiments, the method includes, for each corresponding subject in the first plurality of subjects, sequencing genomic DNA from a corresponding biological sample from the intestine of the corresponding subject, thereby obtaining a corresponding first plurality of at least 100,000 nucleic acid sequences.

[0016] In some such embodiments, the method includes electronically obtaining, for each respective subject in the first plurality of subjects, a corresponding first plurality of at least 100,000 nucleic acid sequences of genomic DNA from a corresponding biological sample from the intestine of the corresponding subject, and determining, for each respective subject in the first plurality of subjects, a corresponding genomic abundance value for each respective gut microbe in the plurality of gut microbes from the corresponding first plurality of at least 100,000 nucleic acid sequences.

[0017] In some such embodiments, the method comprises, for each respective subject in the first plurality of subjects, assembling a corresponding first plurality of gut microbial genomes from the corresponding first plurality of at least 100,000 nucleic acid sequences by metagenomic de novo sequence assembly, and, for each respective gut microbial genome in the corresponding first plurality of gut microbial genomes, calculating a corresponding genomic abundance of the corresponding gut microbial genome.

[0018] In some such embodiments, the method includes, for each respective subject in the first plurality of subjects, assigning each respective nucleic acid sequence in the corresponding first plurality of at least 100,000 sequences to a respective gut microbe in a plurality of gut microbes, thereby generating, for each respective gut microbe in the plurality of gut microbes, a corresponding count of the respective nucleic acid sequence in the corresponding first plurality of nucleic acid sequences assigned to the respective gut microbe, and determining, for each respective gut microbe in the plurality of gut microbes, a corresponding genomic abundance value for the respective gut microbe based on the corresponding count of the respective nucleic acid sequence assigned to the respective gut microbe.

[0019] In some embodiments, the method includes electronically obtaining, for each respective subject in a second plurality of subjects having a second state of the biological characteristic, a corresponding plurality of genomic abundance values, the plurality of genomic abundance values ​​including, for each respective gut microbe in a plurality of gut microbes, a corresponding value for the abundance of a genome of the respective gut microbe in a biological sample from the gut of the respective subject.

[0020] In some such embodiments, for each respective subject in the second plurality of subjects, the biological sample from the intestine of the respective subject is a stool sample.

[0021] In some such embodiments, the method comprises, for each respective subject in the second plurality of subjects, sequencing genomic DNA from a respective biological sample from the intestine of the respective subject, thereby obtaining a corresponding second plurality of at least 100,000 nucleic acid sequences.

[0022] In some such embodiments, the method includes electronically obtaining, for each respective subject in the second plurality of subjects, a corresponding second plurality of at least 100,000 nucleic acid sequences of genomic DNA from a corresponding biological sample from the intestine of the corresponding subject, and determining, for each respective subject in the second plurality of subjects, a corresponding genomic abundance value for each respective gut microbe in the plurality of gut microbes from the corresponding second plurality of at least 100,000 nucleic acid sequences.

[0023] In some such embodiments, the method comprises, for each respective subject in the first plurality of subjects, assembling a corresponding second plurality of gut microbial genomes from a corresponding second plurality of at least 100,000 nucleic acid sequences by metagenomic de novo sequence assembly, and, for each respective gut microbial genome in the corresponding second plurality of gut microbial genomes, calculating a corresponding genomic abundance of the corresponding gut microbial genome.

[0024] In some such embodiments, the method includes, for each respective subject in the first plurality of subjects, assigning each respective nucleic acid sequence in the corresponding second plurality of at least 100,000 sequences to a respective gut microbe in the plurality of gut microbes, thereby generating, for each respective gut microbe in the plurality of gut microbes, a corresponding count of the respective nucleic acid sequence in the corresponding second plurality of nucleic acid sequences assigned to the respective gut microbe, and determining, for each respective gut microbe in the plurality of gut microbes, a corresponding genomic abundance value for the respective gut microbe based on the corresponding count of the respective nucleic acid sequence assigned to the respective gut microbe.

[0025] In some such embodiments, the first state of the biological signature is the absence of a disease or condition and the second state of the biological signature is the presence of a disease or condition, the first state of the biological signature is a first severity of the disease or condition and the second state of the biological signature is a second severity of the disease or condition, the first state of the biological signature is an untreated disease or condition and the second state of the biological signature is a treated disease or condition, the first state of the biological signature is a disease or condition treated with a first therapy and the second state of the biological signature is a disease or condition treated with a second therapy, the first state of the biological signature is a first level of a nutrient in the diet and the second state of the biological signature is a second level of a nutrient in the diet, or the first state of the biological signature is a first age and the second state of the biological signature is a second age.

[0026] In some such embodiments, the plurality of intestinal microorganisms comprises a member selected from Table 1, Table 2, or Figures 42A to 42XX At least 20 species of intestinal microorganisms.

[0027] In some embodiments, the method includes calculating a first plurality of similarity metrics from a corresponding plurality of genomic abundance values ​​across a first plurality of subjects, wherein the first plurality of similarity metrics includes a first corresponding similarity metric for each unique gut microbe pair, and the first corresponding similarity metric quantifies the similarity between: (i) corresponding first vectors formed from corresponding genomic abundance values ​​for first microbes in unique gut microbe pairs across the first plurality of subjects; and (ii) corresponding second vectors formed from corresponding genomic abundance values ​​for second microbes in unique gut microbe pairs across the first plurality of subjects.

[0028] In the following embodiments, the method includes calculating a second plurality of similarity metrics using corresponding genomic abundance values ​​for a second plurality of subjects, wherein the second plurality of similarity metrics includes a second corresponding similarity metric for each unique gut microbe pair in the plurality of gut microbes, and the second corresponding similarity metric quantifies the similarity between: (i) a corresponding second vector formed from corresponding genomic abundance values ​​for the first microbe in the unique gut microbe pairs across the second plurality of subjects; and (ii) a corresponding second vector formed from corresponding genomic abundance values ​​for the second microbe in the unique gut microbe pairs across the second plurality of subjects.

[0029] In some embodiments, the method includes determining, for each corresponding unique gut microbe pair in a set of unique gut microbe pairs, a set of unique gut microbe pairs in a plurality of gut microbes based on a first plurality of similarity metrics and a second plurality of similarity metrics, wherein both the first corresponding similarity metric and the second corresponding similarity metric indicate a statistically significant positive correlation between the abundance of the first gut microbe and the abundance of the second gut microbe in the corresponding unique gut microbe pair, or both the first corresponding similarity metric and the second corresponding similarity metric indicate a statistically significant negative correlation between the abundance of the first gut microbe and the abundance of the second gut microbe in the corresponding unique gut microbe pair.

[0030] In some such embodiments, both the first corresponding similarity measure and the second similarity measure are Pearson correlation coefficients, intraclass correlation coefficients, or rank correlation coefficients.

[0031] In some such embodiments, a statistically significant positive correlation has a P-value less than 0.001.

[0032] In some embodiments, the method includes identifying a set of gut microbes comprising corresponding gut microbes represented in the set of unique gut microbe pairs.

[0033] In some such embodiments, the method includes clustering the corresponding gut microbes represented in the set of unique gut microbe pairs into one or more networks. Each corresponding connected network includes a corresponding plurality of nodes and a corresponding set of one or more edges.

[0034] In some such embodiments, each respective node in the corresponding plurality of nodes represents a unique gut microbe represented in the set of unique gut microbe pairs.

[0035] In some such embodiments, each respective edge in the corresponding set of one or more edges connects two nodes representing a respective unique gut microbe pair in the set of unique gut microbe pairs.

[0036] In some such embodiments, each respective node in the corresponding plurality of nodes is connected to at least one other respective node in the plurality of nodes via a respective edge in the corresponding set of one or more edges.

[0037] In some such embodiments, the method includes identifying a respective network of the one or more networks that contains the most nodes, thereby identifying the set of gut microbes represented by the corresponding plurality of nodes in the respective network.

[0038] In some such embodiments, the set of gut microbes includes all corresponding gut microbes represented in the set of unique gut microbe pairs.

[0039] In some such embodiments, the panel of intestinal microorganisms comprises a member selected from Table 1, Table 2, or Figures 42A to 42XX At least 20 species of intestinal microorganisms.

[0040] Another aspect of the present disclosure provides a method of training a model for assessing human health status at a computer system having one or more processors and a memory storing one or more programs executed by the one or more processors.

[0041] In some such embodiments, the method includes electronically obtaining, for each respective training subject in a plurality of training subjects: (i) a plurality of corresponding genomic abundance values, the plurality of corresponding genomic abundance values ​​comprising, for each respective gut microbe in a plurality of gut microbes, a corresponding value of the abundance of a genome of the respective gut microbe in a corresponding biological sample from the gut of the respective training subject; and (ii) a corresponding state of a biological characteristic of the respective training subject.

[0042] In some such embodiments, for each respective subject in a plurality of training subjects, genomic DNA from a respective biological sample from the intestine of the respective training subject is sequenced to obtain a corresponding plurality of at least 100,000 nucleic acid sequences.

[0043] In some such embodiments, the biological sample from the intestine of the respective subject is a stool sample from the respective training subject.

[0044] In some such embodiments, the plurality of intestinal microorganisms comprises a member selected from Table 1, Table 2, or Figures 42A to 42XX At least 20 species of intestinal microorganisms.

[0045] In some such embodiments, the plurality of intestinal microorganisms comprises at least 20 microorganisms selected from Table 1, Table 2, or Figures 42A to 42XX Those microorganisms with a connectivity of at least 2.

[0046] In some such embodiments, the method includes, for each respective training subject in a plurality of training subjects, obtaining in electronic form a corresponding plurality of at least 100,000 nucleic acid sequences of genomic DNA from a corresponding biological sample of the intestine of the respective training subject, and, for each respective gut microorganism in a plurality of gut microorganisms, determining a corresponding value of the abundance of the genome of the respective gut microorganism from the corresponding first plurality of at least 100,000 nucleic acid sequences.

[0047] In some such embodiments, the method includes, for each respective training subject in the plurality of training subjects, electronically assembling a corresponding plurality of gut microbial genomes from a corresponding plurality of at least 100,000 nucleic acid sequences by de novo metagenomic assembly, and, for each respective gut microbe in the plurality of gut microbes, calculating a corresponding value for the abundance of the genome of the corresponding gut microbe based on the prevalence of the corresponding nucleic acid sequence in the plurality of at least 100,000 nucleic acid sequences that were used to assemble the corresponding gut microbial genome corresponding to the corresponding gut microbe in the plurality of gut microbial genomes.

[0048] In some such embodiments, the method includes assigning, for each respective subject in a plurality of training subjects, each respective nucleic acid sequence in a corresponding plurality of at least 100,000 sequences to a respective gut microbe in a plurality of gut microbes, thereby generating, for each respective gut microbe in the plurality of gut microbes, a corresponding count of the respective nucleic acid sequence in the corresponding plurality of nucleic acid sequences assigned to the respective gut microbe, and determining, for each respective gut microbe in the plurality of gut microbes, a corresponding genomic abundance value for the respective gut microbe based on the corresponding counts assigned to the respective nucleic acid sequence.

[0049] In some such embodiments, the biological characteristic is a disease or condition, a therapy administered to a subject, or a subject's diet.

[0050] In some such embodiments, the disease or condition is selected from the group consisting of type 2 diabetes, hypertension, schizophrenia, atherosclerotic cardiovascular disease (ACVD), liver cirrhosis (LC), inflammatory bowel disease (IBD), colorectal cancer (CRC), ankylosing spondylitis (AS), and Parkinson's disease (PD).

[0051] In some such embodiments, the disease or condition is cancer.

[0052] In some embodiments, the method includes, for each corresponding training subject in a plurality of training subjects, inputting information about the corresponding training subject into a model comprising a plurality of parameters. The model applies the plurality of parameters to the information through at least 10,000 calculations to obtain a corresponding output from the model for the corresponding training subject. The corresponding output includes an indication of the corresponding state of the biological characteristic of the corresponding training subject, and the information about the corresponding training subject includes a corresponding genomic abundance value for each corresponding intestinal microorganism in the plurality of intestinal microorganisms. The plurality of intestinal microorganisms is selected from Table 1, Table 2, or Figures 42A to 42XX .

[0053] In some such embodiments, the indication of the corresponding state of the biological characteristic is a categorical output of a respective state among a plurality of possible states of the biological characteristic.

[0054] In some such embodiments, the indication of the corresponding state of the biological characteristic is a probabilistic output of the corresponding state of the biological characteristic.

[0055] In some such embodiments, the model is a neural network algorithm, a support vector machine algorithm, a naive Bayes algorithm, a nearest neighbor algorithm, a boosted tree algorithm, a random forest algorithm, a convolutional neural network algorithm, a decision tree algorithm, a regression algorithm, or a clustering algorithm.

[0056] In some such embodiments, the plurality of parameters is at least 1000, at least 10,000, at least 15,000, at least 50,000, at least 100,000, at least 250,000, at least 500,000, or at least 1,000,000 parameters.

[0057] In some such embodiments, the model applies the plurality of parameters to the information by at least 25,000, at least 50,000, at least 100,000, at least 250,000, at least 500,000, or at least 1,000,000 computations to obtain a corresponding output from the model for the corresponding training subject.

[0058] In some embodiments, the method includes adjusting the plurality of parameters based on one or more differences between, for each respective training subject in the first plurality of training subjects, (i) the corresponding output from the model and (ii) the corresponding state of the biological characteristic for the respective training subject.

[0059] Another aspect of the present disclosure provides a method for assessing a health condition of a subject at a computer system having one or more processors and a memory storing one or more programs executed by the one or more processors.

[0060] In some such embodiments, the method comprises electronically obtaining a plurality of genomic abundance values ​​comprising for a genome selected from Table 1, Table 2, or Figures 42A to 42XX For each corresponding intestinal microorganism in the plurality of at least 20 intestinal microorganisms, a corresponding abundance value of the genome of the corresponding intestinal bacterial species in the plurality of at least 20 intestinal microorganisms in the biological sample from the subject.

[0061] In some such embodiments, the method comprises sequencing genomic DNA from the biological sample from the intestine of the subject, thereby obtaining the plurality of at least 100,000 nucleic acid sequences.

[0062] In some such embodiments, the biological sample from the subject's intestine is a stool sample.

[0063] In some such embodiments, the plurality of intestinal microorganisms comprises at least 20 microorganisms selected from Table 1, Table 2, or Figures 42A to 42XX Those microorganisms with a connectivity of at least 2.

[0064] In some such embodiments, the method comprises obtaining in electronic form a plurality of at least 100,000 nucleic acid sequences of genomic DNA from the biological sample of the intestine of the subject; and for each corresponding gut microorganism in the plurality of gut microorganisms, determining a corresponding value of the abundance of the genome of the corresponding gut microorganism from the plurality of at least 100,000 nucleic acid sequences.

[0065] In some such embodiments, the method includes electronically assembling a corresponding plurality of gut microbial genomes from the plurality of at least 100,000 nucleic acid sequences by de novo metagenomic assembly, and calculating, for each corresponding gut microbe in the plurality of gut microbes, a corresponding value for the abundance of the genome of the corresponding gut microbe based on the prevalence of the corresponding nucleic acid sequence in the plurality of at least 100,000 nucleic acid sequences used to assemble the corresponding gut microbial genome corresponding to the corresponding gut microbe in the plurality of gut microbial genomes.

[0066] In some such embodiments, the method includes assigning each corresponding nucleic acid sequence in the plurality of at least 100,000 sequences to a corresponding gut microbe in the plurality of gut microbes, thereby generating, for each corresponding gut microbe in the plurality of gut microbes, a corresponding count of the corresponding nucleic acid sequence in the plurality of nucleic acid sequences assigned to the corresponding gut microbe, and determining, for each corresponding gut microbe in the plurality of gut microbes, a corresponding genomic abundance value for the corresponding gut microbe based on the corresponding count of the corresponding nucleic acid sequence assigned to the corresponding gut microbe.

[0067] In some embodiments, the method includes inputting the plurality of genomic abundance values ​​into a model comprising a plurality of parameters, wherein the model applies the plurality of parameters to the plurality of genomic abundance values ​​through at least 10,000 calculations to generate an indication of the health status of the subject as an output from the model.

[0068] In some such embodiments, the indication of the subject's health status is indicative of a biological characteristic. The biological characteristic is a disease or condition, a therapy administered to the subject, or the subject's diet.

[0069] In some such embodiments, the disease or condition is selected from the group consisting of type 2 diabetes, hypertension, schizophrenia, atherosclerotic cardiovascular disease (ACVD), liver cirrhosis (LC), inflammatory bowel disease (IBD), colorectal cancer (CRC), ankylosing spondylitis (AS), and Parkinson's disease (PD).

[0070] In some such embodiments, the disease or condition is cancer.

[0071] In some such embodiments, the indication of the subject's health condition is a categorical output of a respective state among a plurality of possible states of the subject's health condition.

[0072] In some such embodiments, the indication of the subject's health condition is a probabilistic output of a corresponding state of the subject's health condition.

[0073] In some such embodiments, the model is a neural network algorithm, a support vector machine algorithm, a naive Bayes algorithm, a nearest neighbor algorithm, a boosted tree algorithm, a random forest algorithm, a convolutional neural network algorithm, a decision tree algorithm, a regression algorithm, or a clustering algorithm.

[0074] In some such embodiments, the plurality of parameters is at least 1000, at least 10,000, at least 15,000, at least 50,000, at least 100,000, at least 250,000, at least 500,000, or at least 1,000,000 parameters.

[0075] In some such embodiments, the model applies the plurality of parameters to the information by at least 25,000, at least 50,000, at least 100,000, at least 250,000, at least 500,000, or at least 1,000,000 computations to obtain a corresponding output from the model for the corresponding training subject.

[0076] Another aspect of the present disclosure provides a computer system comprising one or more processors and a non-transitory computer-readable medium comprising computer-executable instructions that, when executed by the one or more processors, cause the processors to perform the methods described herein.

[0077] Another aspect of the present disclosure provides a non-transitory computer-readable storage medium storing instructions that, when executed by a computer system, cause the computer system to perform any of the methods described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] Figure 1 A block diagram illustrating an exemplary computing device according to some embodiments of the present disclosure is illustrated.

[0079] Figure 2A 、 2B , 2C, 2D, 2E, 2F and 2G collectively provide a flow chart of processes and characteristics for identifying a group of intestinal microorganisms according to some embodiments of the present disclosure.

[0080] Figure 3A 、 3B , 3C and 3D together provide a flowchart of the process and features for training a model for assessing human health status according to some embodiments of the present disclosure.

[0081] Figure 4A 、 4B 4C together provide a flow chart of the process and characteristics of assessing the health status of a subject according to some embodiments of the present disclosure.

[0082] Figure 5A 、 5B, 5C, 5D, 5E, 5F, 5G, and 5H together illustrate that reversible changes in the gut microbiota are associated with reversible changes in the metabolic phenotype in patients with T2DM. (A) Study design. Before the run-in period, written informed consent was signed, a personal information questionnaire was completed, and HbA1c was measured at screening. After the run-in period, physical examinations and sample collection were performed at baseline (M0), three months after high-fiber intervention or conventional diet (M3), and one year after stopping high-fiber intervention (M15). (B) Changes in fiber intake. (C) Overall changes in the gut microbiome as shown by principal coordinate analysis based on Bray-Curtis distance for 1845 genomes, and (D) average Bray-Curtis distance between groups. A PERMANOVA test (9,999 permutations) was performed to compare the groups. *P<0.05 and ***P<0.001. The color of the squares shows the magnitude of the average Bray-Curtis distance. (E) Changes in HbA1c, (F) percentage of participants with adequate glycemic control, (G) fasting blood glucose, and (H) glucose area under the curve (AUC) in the meal tolerance test (MTT). For (E), (G), and (H), data are shown as percentage change relative to baseline (± SEM). Comparisons in the same group were performed using the Friedman test followed by the Nemenyi post hoc test, and close letters reflect significance (P < 0.05). n = 67 in the W group and n = 28 in the U group. Comparisons between W and U at the same time point were performed using the Mann-Whitney test (two-sided), *P < 0.05, **P < 0.01, and ***P < 0.001. In W n=74(M0) (for Figure H, n=72), in W n=74(M3), in W n=67(M15), in U n=36(M0), in U n=36(M3) and in U n=28(M15).

[0083] Figure 6A 、 6B1, 6B2 and 6B3 together illustrate that although the introduction and interruption of the high-fiber intervention induced profound overall changes in the intestinal microbial ecosystem, the two competing bacterial communities formed a robust seesaw network. (A) Distribution of different types of genome pair correlations during the experiment. The three letters show the genome pair correlations at M0, M3 and M15, respectively. The stable correlations NNN and PPP are highlighted. (B) UPGMA clustering of 141 nodes based on robust positive and negative correlations of nodes shows two clusters (green range and purple range). The bar chart shows the changes in abundance of each node throughout the experiment, and the changes in abundance are expressed as median abundance by Z-score conversion. The differences in changes in each node over time were tested using the Friedman test followed by the Nemenyi post hoc test. P < 0.05 was considered significant. For (B), the color of the node represents the membership of the two bacterial communities: bacterial community 1 is green and bacterial community 2 is purple.

[0084] Figure 7A 、 7B, 7C1, and 7C2 together illustrate that the balance between two competing bacterial communities in a seesaw network is associated with metabolic health in patients with type 2 diabetes. (A) Changes in the total abundance of community 1, the total abundance of community 2, and their ratio during the W group experiment. Differences between time points were analyzed using the Friedman test followed by the Nemenyi test. The tightness of the letters reflects significance at P < 0.05. (B) Associations between 141 gene sets and clinical parameters were explored using random forest regression with leave-one-out cross-validation. The bar chart shows the Pearson correlation coefficient between the predicted and measured values. The asterisk before the parameter name indicates the significance of the Pearson correlation. P values ​​were adjusted by the Benjamini and Hochberg method. *Adjusted P < 0.05, **Adjusted P < 0.01, and ***Adjusted P < 0.001. BMI, body mass index; SBP, systolic blood pressure; DBP, diastolic blood pressure; WC, waist circumference; HP, hip circumference; TNF-α, tumor necrosis factor-α; WBC, white blood cell count; CRP, C-reactive protein; LBP, lipopolysaccharide-binding protein; TC, total cholesterol; TG, triglycerides; Lpa, lipoprotein a; HDL, high-density lipoprotein; APOA, apolipoprotein A; LDL, low-density lipoprotein; APOB, apolipoprotein B; GFR(MDRR), glomerular filtration rate; CysC, cystatin C; ACR, urine microalbumin-to-creatinine ratio; I MT, intima-media thickness; DAN, diabetic autonomic neuropathy score; MHR, mean heart rate; SDNN, standard deviation of NN intervals; SDANN, standard deviation of the mean NN intervals calculated over 5 minutes; SDNNIndex, mean of the standard deviations of NN intervals over 5-minute periods; rMSSD, root mean square of the differences between consecutive NN intervals; pNN50, percentage of consecutive NN intervals with interval differences greater than 50 ms; TP, total power; VLF, very low frequency power; LF, low frequency power; HF, high frequency power; DPN, diabetic peripheral neuropathy score. (C) Heritability differences for carbohydrate substrate utilization (CAZy), short-chain fatty acid production (SCFA), number of antibiotic resistance genes (ARGs), and number of virulence factor genes (VFs). (C) Heatmap showing the proportion of each category (CAZy) or gene copy number (SCFA, ARGs, and VFs) in each genome. For carbohydrate substrate utilization, CAZy genes were predicted in each genome. The proportion of CAZy genes for a specific substrate was calculated as the number of CAZy genes involved in its utilization divided by the total number of CAZy genes.Arabinoxylan-related CAZy family: CE1, CE2, CE4, CE6, CE7, GH10, GH11, GH115, GH43, GH51, GH67, GH3 and GH5; cellulose-related: GH1, GH44, GH48, GH8, GH9, GH3 and GH5; inulin-related: GH32 and GH91; mucin-related family: GH1, GH2, GH3, GH4, GH18, GH19, GH20, GH29, GH33, GH38, GH58, GH79, GH84, GH85, GH88, GH89, GH92, GH95, GH98, GH99, GH101, GH105, GH109, ​​GH110, GH113, PL6, PL8, PL12, PL13 and PL21; pectin-related: CE12, CE8, GH28, PL1 and PL9; starch-related: GH13, GH31 and GH97. For short-chain fatty acid production, FTHFS: formate-tetrahydrofolate ligase for acetate production; ScpC: propionyl-CoA succinate-CoA transferase and Pet: propionate-CoA transferase for propionate production; But: butyryl-CoA (butyryl-CoA): acetate CoA transferase, Buk: butyrate kinase, 4Hbt: butyryl-CoA: 4-hydroxybutyrate CoA transferase, Ato: butyryl-CoA: acetoacetate CoA transferase (AtoA: α subunit, AtoD: β subunit) for butyrate production. The Mann-Whitney test (two-sided) was used to analyze the differences between colonies 1 and 2. #P < 0.1, *P < 0.05, **P < 0.01, and ***P < 0.001. Colony 1 (green bars): n = 50, colony 2 (purple bars): n = 91.

[0085] Figure 8A1 、 8A2 , 8A3, 8A4, and 8B collectively demonstrate that the seesaw network-like microbiome signature exists in other independent human cohorts and supports classification models for different diseases. (A) Microbiome signatures support classification models for four different diseases. The area under the ROC curve (AUC) of the random forest classifier based on 141 gene groups in the microbiome signature to classify controls and patients in each dataset. Leave-one-out cross-validation was applied. Type 2 diabetes mellitus (T2DM): controls n=136, T2D n=136; atherosclerotic cardiovascular disease (ACVD): controls n=171, and ACVD n=214; cirrhosis (LC): controls n=83, and LC n=84; ankylosing spondylitis: controls n=83, and AS n=97. (B) Microbiome signatures are associated with key T2D phenotypes.

[0086] Figure 9A flow chart illustrating the participants in the trial described in Example 1.

[0087] Figure 10A 、 10B , 10C, and 10D together illustrate violin plots of energy and macronutrient intake during the trial period in the W and U groups. Groups were compared using the Friedman test followed by the Nemenyi post hoc test. Comparisons between W and U at the same time point were performed using the Mann-Whitney test (two-sided). *P < 0.05, **P < 0.01, and ***P < 0.001. Boxes show the median and interquartile range (IQR), whiskers indicate the lowest and highest values ​​within 1.5 times the IQR of the first and third quartiles, and outliers are shown as individual points.

[0088] Figure 11A 、 11B , 11C, and 11D together illustrate violin plots of changes in gut microbiome alpha diversity during the trial in the W and U groups. (A) Shannon index; (B) Simpson index; (C) Observed gene sets; (D) Chao 1 index. Comparisons within the same group were performed using the Friedman test followed by the Nemenyi post hoc test. Comparisons between W and U at the same time point were performed using the Mann-Whitney test (two-sided). *P < 0.05, **P < 0.01, and ***P < 0.001. Boxes show the median and interquartile range (IQR), whiskers indicate the lowest and highest values ​​within 1.5 times the IQR of the first and third quartiles, and outliers are shown as individual points.

[0089] Figure 12A 、 12B Together with 12C, we show that the co-abundance network of the prevalent genome in the experiment is scale-free, and the degree distribution fits well with a power-law model.

[0090] Figure 13 It shows that the introduction and interruption of high-fiber intervention significantly changes the network degree distribution.

[0091] Figure 14A 、 14B14C, 14D, 14E, and 14F collectively demonstrate that the 141 gene sets contribute the majority of interactions in the network. (A) Stacked bar chart shows the distribution of positive and negative edges within the 141 gene sets, between the 141 gene sets and other nodes, and within other nodes. The 141 gene sets have significantly higher degree (B), betweenness centrality (C), eigenvector centrality (D), closeness centrality (E), and pressure centrality (F) than other nodes in the network. A two-sided Mann-Whitey test was performed. ***P < 0.0001.

[0092] Figure 15 The figure illustrates 141 nodes that are widely shared among patients in group W. The histogram shows the distribution of the genome shared by the 74 patients with varying prevalence.

[0093] Figure 16A 、 16B and 16C together illustrate that similar β-diversity patterns were found based on 141 genomes compared to the β-diversity patterns based on all 1845 genomes. (A) Overall changes in the gut microbiome as shown by principal coordinate analysis based on Bray-Curtis distance and abundance of 141 genomes. (B) Average Bray-Curtis distance between groups (B). A PERMANOVA test (9,999 permutations) was performed to compare the groups. *P < 0.05 and ***P < 0.001. The color of the squares shows the magnitude of the average Bray-Curtis distance. (C) Procrustes analysis, which combines principal coordinate analysis for 1845 genomes and 141 genomes based on Bray-Curtis distance.

[0094] Figure 17A 、 17B Together with 17C, the 141 nodes organize themselves into two clusters, each with robust co-occurrence behavior, and can be identified as potential ecological communities. (A-C) Stacked bar charts show the number of positive and negative edges within or between communities. Red: within community 1; blue: within community 2; green: between communities.

[0095] Figure 18A and 18B Together, these results demonstrate that the genomes in cluster 1 have much lower genetic capacity for pathogenicity and antibiotic resistance. (A) Bar graph shows the number of genes encoding virulence factors (VFs) and VF classes. (B) Bar graph shows the number of ARGs and corresponding antibiotic resistance types.

[0096] Figure 19 A workflow for validating microbiome signatures in other datasets according to some embodiments of the present disclosure is described.

[0097] Figure 20 Correlations between microbiome signatures and host phenotypes in a cirrhosis dataset were demonstrated (Qin et al. 2014). Random forest regression with leave-one-out cross-validation was used to explore associations between microbiome signatures and clinical parameters. Bar graphs show the Pearson correlation coefficients between predicted and measured values. Asterisks preceding parameter names indicate the significance of the Pearson correlation. P values ​​were adjusted by the Benjamini and Hochberg method. **Adjusted P < 0.01 and ***P < 0.001. TB: total bilirubin, Crea: creatinine level, Alb: albumin level, BMI: body mass index. N = 167.

[0098] Figure 21A and 21B Together illustrate receiver operating characteristic (ROC) curves for the performance of a random forest classifier trained to predict human disease based on genomic abundance values ​​of 141 gut microbes in diseased and healthy subjects in a multi-disease study, as described in Example 1.

[0099] Figure 22A and 22B Clinical parameters during the intervention period in the W and U groups are collectively illustrated.

[0100] Figure 23 Characterization of the co-abundance network of prevalent gene sets in the W group at M0, M3, and M15 during the experiment is illustrated, denoted as GM0, GM3, and GM15.

[0101] Figure 24 This study describes the design of a high-fiber intervention clinical trial in Chinese patients with T2D. Patients with type 2 diabetes were randomly assigned to a treatment group (W group) that received a WTP diet for 3 months and a one-year follow-up after discontinuation of the WTP diet, or a control group (U group) that received usual care and a one-year follow-up.

[0102] Figure 25 This study describes the genome-resolved metagenomic analysis used in a clinical study of a high-fiber intervention in patients with type 2 diabetes. In this study, shotgun metagenomics was applied to explore the gut microbiome. On average, each sample had 91.5 million raw reads and 86.5 million high-quality reads. In short, after de novo assembly, binning, quality control, and deduplication, 1,845 high-quality, non-redundant genomes were obtained for further genome-level analysis. These genomes accounted for >70% of the total reads in our metagenomic dataset.

[0103] Figure 26A 、 26B, 26C, 26D, 26E, 26F, 26G, 26H, 26I, 26J, 26K, 26L and 26M collectively illustrate the classification performance of eight types of diseases using different numbers of gene sets selected using degree-based reverse selection.

[0104] Figure 27A 、 27B , 27C, 27D, 27E, 27F, 27G, 27H, 27I, 27J, 27K, 27L, and 27M collectively illustrate the random forest classification performance for eight types of diseases using different numbers of randomly selected genomes.

[0105] Figure 28A 、 28B , 28C, 28D, 28E, 28F, 28G, and 28H collectively illustrate the classification power of two competing bacterial groups identified from QD and various types of diseases. Microbiome signatures containing genomes of two competing bacterial groups were obtained from multiple diseases; T2D ( Figure 28A )、LC( Figure 28B )、SCZ( Figure 28C )、IBD( Figure 28D )、AS( Figure 28E )、ACVD( Figure 28F )、CRC( Figure 28G ) and QD( Figure 28H The microbiome signatures identified for each condition were used to classify controls versus patients in each dataset using a random forest classifier. Figures 28A to 28H All microbiome signatures were shown to have the ability to classify cases and controls across studies.

[0106] Figure 29A and 29B The classification ranking of microbiome signatures for QD and other types of diseases is collectively illustrated. Eight groups of microbiome signatures obtained from QD and multiple disease cases (T2D, LC, SCZ, IBD, AS, ACVD, CRC) were ranked based on their performance in classifying cases from controls across 11 datasets. For each dataset, the best performer with the highest AUC value had the lowest ranking number, while the worst performer with the lowest AUC value had the highest ranking number. All ranking numbers assigned to each group of signature microbiome signatures are plotted on Figure 29A middle. Figure 29B The sum of the ranking numbers for each group of microbiome signatures is shown. Microbiome signatures obtained from QD had the best performance in classifying healthy subjects from patients in different datasets.

[0107] Figure 30A and 30BTogether, we demonstrated the ability of the combined pool to classify cases and controls from different studies. Eight signature microbiomes obtained from QD and multiple disease cases (T2D, LC, SCZ, IBD, AS, ACVD, CRC) were pooled together as a combined microbiome signature. Figure 30A Comparison of the classification performance of the combined pool with each individual signature microbiome based on AUC values ​​is shown. Figure 30B The significance of intra-group comparisons is shown. Friedman test followed by Dunn's post hoc test was performed for analysis (#BH adjusted P < 0.1, *BH adjusted P < 0.05).

[0108] Figure 31A 、 31B and 31C together illustrate the ranking of the classification performance of microbiome signatures. Nine microbiome signatures obtained from combined pools, QDs, or cases of multiple diseases (T2D, LC, SCZ, IBD, AS, ACVD, CRC) were ranked according to their performance in classifying cases from controls in 11 datasets. All ranking numbers assigned to each signature microbiome were plotted on Figure 31A middle. Figure 31B The significance of within-group comparisons is shown. Figure 31C The sum of ranks for each group of microbiome signatures is shown. Kruskal-Wallis test followed by Dunn's post hoc test was performed for analysis (#BH adjusted P < 0.1, *BH adjusted P < 0.05). The microbiome signatures obtained from the combined pool had the best performance in classifying healthy subjects from patients in different datasets.

[0109] Figure 32 The selection of the combined core pool is described. A random forest classification based on the combined 788 genomes was performed on each data set. Each of the 788 genomes was ranked according to its importance. The total ranking was obtained by adding the ranking values ​​of the 11 data sets, and all 788 genomes were ranked again according to the sum value. The most important genomes in the 11 data sets obtained the lowest total ranking value. Starting from the least important genome, each genome was removed one by one from each data set based on the order of importance. The classification performance (AUC) of the genomes with the remaining numbers after each removal was calculated by the random forest model, and all genome numbers were ranked based on the AUC value. The ranking values ​​of each genome number in the 11 data sets were summed. The sum of the rankings of each genome number in the 11 data sets was plotted. 302 genomes obtained the lowest AUC sum ranking. After removing 18 genomes that showed inconsistent C1A and C1B assignments, 284 genomes were retained as the combined core pool.

[0110] Figure 33A、 33B , 33C, 33D, 33E, 33F, 33G, 33H, 33I, and 33J collectively illustrate the classification power of two competing bacterial groups identified from QD, multiple disease types, combined pools, and combined core pools. Microbiome signatures of genomes containing two competing bacterial groups from multiple diseases: T2D ( Figure 33A )、LC( Figure 33B )、AS( Figure 33C )、CRC( Figure 33D )、IBD( Figure 33E )、QD( Figure 33F )、AVCD( Figure 33G )、SCZ( Figure 33H ), combined collection ( Figure 33I ) and combined core pools ( Figure 33J The microbiome signatures identified for each condition were used to classify controls and patients in each dataset using a random forest classifier. Figure 31 shows that all microbiome signatures had the ability to classify cases and controls across the different studies.

[0111] Figure 34A and 34B Together, the ability of the combined core pool to classify cases and controls across studies is demonstrated. Figure 34A Comparison of classification performance based on AUC of the combined core pool and signature microbiome obtained from the combined pool, QD, and multiple disease cases (T2D, LC, SCZ, IBD, AS, ACVD, CRC) is shown. Figure 34B The significance of intra-group comparisons is shown. Friedman test followed by Dunn's post hoc test was performed for analysis (#BH adjusted P < 0.1, *BH adjusted P < 0.05, **BH adjusted P < 0.01).

[0112] Figure 35A 、 35B Together with 35C, we demonstrate the rank order of classification performance of microbiome signatures. Ten microbiome signatures derived from the combined core pool, combined pool, QD, or multiple disease cases (T2D, LC, SCZ, IBD, AS, ACVD, CRC) were ranked based on their performance in classifying cases from controls across 11 datasets. Figure 35A All rank values ​​assigned to each set of signature microbiome are plotted. Figure 35B The significance of within-group comparisons is shown. Figure 35CThe sum of ranks for each group of microbiome signatures is shown. Kruskal-Wallis test followed by Dunn's post hoc test was performed for analysis (#BH adjusted P < 0.1, *BH adjusted P < 0.05, **BH adjusted P < 0.01). The microbiome signatures obtained from the combined core pool had the best performance in classifying healthy subjects from patients in different datasets.

[0113] Figure 36 The pipeline for identifying microbiome signatures from case and control cohorts is described.

[0114] Figure 37 Combined case and control samples from 25 datasets corresponding to 15 different diseases (type 2 diabetes (T2D), hypertension (HT), schizophrenia (SCZ), atherosclerotic cardiovascular disease (ACVD), liver cirrhosis (LC), inflammatory bowel disease (IBD), colorectal cancer (CRC), ankylosing spondylitis (AS), Parkinson's disease (PD), multiple sclerosis (MS), Gaucher disease type II (GDII), COVID-19 (COV), Behçet's disease (BD), autism spectrum disorder (ASD), and pancreatic cancer (PC)) are described.

[0115] Figure 38A1 、 38A2 , 38A3, 38B1, 38B2, and 38B3 collectively illustrate a universal random forest classification model for cases and controls based on the abundance of 284 core gene sets. A: 80% training: controls, n = 1285; cases, n = 1424, 10-fold CV (A1: area under the ROC curve (AUC) of the random forest classifier; A2: score density of cases and controls; A3: probability score of cases and controls); B: 20% testing: controls, n = 319; cases, n = 356 (B1: area under the ROC curve (AUC) of the random forest classifier; B2: score density of cases and controls; B3: probability score of cases and controls).

[0116] Figure 39A and 39B Figure 1 illustrates repeated training of a generalized random forest classification model for cases and controls using a randomly selected number of gene sets. (A) Each data point represents the mean AUC of a random forest model trained ten times using a different set of randomly selected gene sets, with the total number of gene sets, X, determined for the training set (indicated on the x-axis). (B) Each data point represents the mean AUC of a random forest model trained ten times using a different set of randomly selected gene sets, with the total number of gene sets, X, determined for the test set (indicated on the x-axis).

[0117] Figure 40A and40B Pairwise ANI comparisons of gene compositions are collectively illustrated. Figure 40A All pairwise ANI comparisons of genes in the combined pool of 788 genomes are depicted. Figure 40B Pairwise ANI comparisons between the genomes of colony 1 and colony 2 are depicted.

[0118] Figure 41A 、 41B , 41C, 41D, 41E, 41F, 41G, 41H, and 41I collectively illustrate the corresponding contigs obtained for each of the 788 genomes (referenced by SEQ ID).

[0119] Figure 42A 、 42B , 42C, 42D, 42E, 42F, 42G, 42H, 42I, 42J, 42K, 42L, 42M, 42N, 42O, 42P, 42Q, 42R, 42S, 42T, 42U, 42V, 42W, 42X, 42Y, 42Z, 42AA, 42BB, 42CC, 42DD, 42EE, 42FF, 42GG, 42HH, 42II, 42JJ, 42KK, 42LL, 42MM, 42NN, 42OO, 42PP, 42QQ, 42RR, 42SS, 42TT, 42UU, 42VV, 42WW, and 42XX collectively illustrate the taxonomic assignments of 788 combined microbiomes.

[0120] Like reference numerals designate corresponding parts throughout the several views of the drawings. DETAILED DESCRIPTION

[0121] The methods and systems described herein facilitate determining a subject's disease state across a variety of disease states based on the composition of the subject's microbiome.

[0122] definition.

[0123] The terms used in this disclosure are only used to describe the purpose of specific embodiments, and are not intended to limit the present invention. As used in the description of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular "a / an" and "the" are intended to include plural forms as well. It should also be understood that, as used herein, the term "and / or" refers to and encompasses any and all possible combinations of one or more of the associated listed items. It should be further understood that, when used in this specification, the term "includes," "comprising" or any of its variations specify the existence of stated features, integers, steps, operations, elements and / or components, but does not exclude the existence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. In addition, in the case of the terms "including / includes," "having / has / with," or its variations used in specific embodiments and / or claims, such terms are intended to have inclusiveness in a manner similar to the term "comprising."

[0124] As used herein, the term "if" may be interpreted to mean "when" or "upon" or "in response to determining" or "in response to detecting," as appropriate to the context. Similarly, the phrase "if it is determined" or "if [stated condition or event] is detected" may be interpreted to mean "upon determining" or "in response to determining" or "upon detecting [stated condition or event]" or "in response to detecting [stated condition or event]," as appropriate to the context.

[0125] It should also be understood that although the terms first, second, etc. can be used to describe multiple elements in this article, these elements should not be limited to these terms. These terms are only used to distinguish an element from another element. For example, without departing from the scope of this disclosure, the first subject can be referred to as the second subject, and similarly, the second subject can be referred to as the first subject. Although the first subject and the second subject are subjects, these subjects are not the same subject. The terms "subject", "user" and "patient" are used interchangeably in this article.

[0126] As used herein, the term "measure of central tendency" refers to the central or representative value of a distribution of values. Non-limiting examples of measures of central tendency include the arithmetic mean, weighted mean, midrange, median, triple means, geometric mean, geometric median, winsorized mean, median, and mode of a distribution of values.

[0127] As used herein, the term "subject" refers to any living or non-living organism, including but not limited to humans (e.g., male humans, female humans, fetuses, pregnant women, children, etc.), non-human mammals or non-human animals. Any human or non-human animal can serve as a subject, including but not limited to mammals, reptiles, birds, amphibians, fish, ungulates, ruminants, bovines (e.g., cattle), equines (e.g., horses), caprines and ovines (e.g., sheep, goats), porcines (e.g., pigs), camelids (e.g., camels, llamas, alpacas), monkeys, apes (e.g., gorillas, chimpanzees), ursids (e.g., bears), poultry, dogs, cats, mice, rats, fish, dolphins, whales, and sharks. In some embodiments, the subject is a male or female (e.g., man, woman, or child) of any age.

[0128] As used herein, the terms "cancer," "cancerous tissue," or "tumor" refer to an abnormal mass of tissue, wherein the growth of the mass exceeds and is not coordinated with the growth of normal tissue, including solid masses (e.g., as in solid tumors) or liquid masses (e.g., as in blood cancers). Cancer or tumors can be defined as "benign" or "malignant" based on the following characteristics: degree of cell differentiation (including morphology and function), growth rate, local invasion, and metastasis. "Benign" tumors can be well differentiated, characterized by slower growth than malignant tumors, and remain located at the original site. In addition, in some cases, benign tumors do not have the ability to infiltrate, invade, or metastasize to distant sites. "Malignant" tumors may be poorly differentiated (anaplastic), characterized by rapid growth, accompanied by progressive infiltration, invasion, and destruction of surrounding tissues. In addition, malignant tumors may have the ability to metastasize to distant sites. Therefore, cancer cells are cells found in abnormal tissue masses, whose growth is not coordinated with the growth of normal tissue. Therefore, a "tumor sample" refers to a biological sample of a tumor obtained or derived from a subject, as described herein.

[0129] Non-limiting examples of cancer types include ovarian cancer, cervical cancer, uveal melanoma, colorectal cancer, chromophobe renal cell carcinoma, liver cancer, endocrine tumors, oropharyngeal cancer, retinoblastoma, bile duct cancer, adrenal cancer, neural cancer, neuroblastoma, basal cell carcinoma, brain cancer, breast cancer, non-clear cell renal cell carcinoma, glioblastoma, glioma, kidney cancer, gastrointestinal stromal tumor, medulloblastoma, bladder cancer, stomach cancer, bone cancer, non-small cell lung cancer, Thymoma, prostate cancer, clear cell renal cell carcinoma, skin cancer, thyroid cancer, sarcoma, testicular cancer, head and neck cancer (e.g., head and neck squamous cell carcinoma), meningioma, peritoneal cancer, endometrial cancer, pancreatic cancer, mesothelioma, esophageal cancer, small cell lung cancer, Her2-negative breast cancer, ovarian serous carcinoma, HR+ breast cancer, uterine serous carcinoma, uterine corpus endometrial carcinoma, gastroesophageal junction adenocarcinoma, gallbladder cancer, chordoma, and papillary renal cell carcinoma.

[0130] As used herein, the term "cancer status" or "cancer status" refers to characteristics of a cancer patient's condition, such as diagnostic status, cancer type, cancer location, primary origin of the cancer, cancer stage, cancer prognosis, and / or one or more additional characteristics of the cancer (e.g., tumor characteristics, such as morphology, heterogeneity, size, etc.). In some embodiments, one or more additional personal characteristics of the subject are used to further describe the subject's cancer status or cancer status, such as age, sex, weight, race, personal habits (e.g., smoking, alcohol consumption, diet), other relevant medical conditions (e.g., high blood pressure, dry skin, other diseases), current medications, allergies, relevant medical history, side effects of current cancer treatment and other medications, etc.

[0131] As used herein, the term "genomic abundance value" refers to the absolute amount or relative amount of a microbial genome in a biological sample from the intestinal tract of a subject. Genomic abundance values ​​can be expressed in different units, including copy number, molar concentration, mass (e.g., mass standardized for the size of a genome), unique sequence reads (e.g., unique sequence reads standardized for the size of a genome), the percentage of the total amount of any previous measurement relative to the measurement of all genomes in a sample, the percentage of the total amount of any previous measurement relative to the measurement of multiple genomes in a sample, etc. In some embodiments, the genomic abundance value is normalized for the total genomic abundance in the sample. In some embodiments, the genomic abundance value is normalized for the genomic abundance value of the control genome in the sample. In some embodiments, the value of multiple genomic abundance values ​​in the sample is standardized, normalized, and / or scaled. Examples of methods for normalizing genomic abundance values ​​are described in, for example, Lin, H., Peddada, SD, Analysis of microbial compositions: a review of normalization and differential abundance analysis, Biofilms Microbiomes, 6(60)(2020) and Lutz KC et al., A Survey of Statistical Methods for Microbiome Data Analysis, Frontiers in Applied Mathematics and Statistics, 8(2022), the contents of which are incorporated herein by reference in their entirety. Methods for measuring genomic abundance values ​​are known in the art. For example, metagenomic sequencing can be used to reconstruct microbial genomes to a large extent by performing next generation sequencing on genomic DNA in biological samples (such as biological samples from the intestine of a subject). For a review of metagenomic sequences, see, for example, Quince C et al., Shotgun metagenomics, from sampling to analysis, Nat Biotechnol, 35(9): 833-44(2017), the contents of which are incorporated herein by reference in their entirety. Genomic abundance can also be determined by quantifying the copy number of ribosomal genes, such as the 16S rRNA gene.Examples of rRNA quantification are described in: Manzari C. et al., Accurate quantification of bacterial abundance in metagenomic DNAs accounting for variable DNA integrity levels, Microb Genom., 6(10): mgen000417 (2020) and Barlow, JT et al., A quantitative sequencing framework for absolute abundance measurements of mucosal and lumenal microbial communities, Nat Commun., 11: 2590 (2020), the contents of which are incorporated herein by reference in their entirety.

[0132] As used herein, the term "relative abundance" refers to the ratio of a first amount of a compound (e.g., a genome of a first microorganism) measured in a sample to a second amount of the compound measured in a second sample. In some embodiments, relative abundance refers to the ratio of the amount of a compound (e.g., a genome of a first microorganism) to the total amount of the compound (e.g., the total amount of a microorganism genome or the total amount of multiple genomes) in the same sample. In other embodiments, relative abundance refers to the ratio of the amount of a compound (e.g., a genome of a first microorganism) in a first sample to the amount of the compound in a second sample. For example, the ratio of the normalized amount of the genome of a first microorganism in a first sample to the normalized amount of the genome of the first microorganism in a second sample and / or a reference sample.

[0133] As used herein, the terms "sequencing," "sequencing," and the like refer to any biochemical method that can be used to determine the order of a biological macromolecule such as a nucleic acid or protein. For example, sequencing data can include all or part of the nucleotide bases in a nucleic acid molecule such as an mRNA transcript or a genomic locus.

[0134] As used herein, the term "sequence read" or "read" refers to a nucleotide sequence generated by any nucleic acid sequencing method described herein or known in the art. A read can be generated from one end of a nucleic acid fragment (a "single-end read") or from both ends of a nucleic acid fragment (e.g., paired-end reads, double-ended reads). The length of a sequence read is generally associated with a particular sequencing technology. For example, high-throughput methods can provide sequence reads that can vary in size from tens to hundreds of base pairs (bp). In some embodiments, the mean, median, or average length of the sequence reads is between about 15 bp and 900 bp long (e.g., about 20 bp, about 25 bp, about 30 bp, about 35 bp, about 40 bp, about 45 bp, about 50 bp, about 55 bp, about 60 bp, about 65 bp, about 70 bp, about 75 bp, about 80 bp, about 85 bp, about 90 bp, about 95 bp, about 100 bp, about 110 bp, about 120 bp, about 130, about 140 bp, about 150 bp, about 200 bp, about 250 bp, about 300 bp, about 350 bp, about 400 bp, about 450 bp, or about 500 bp). In some embodiments, the mean, median, or average length of the sequence reads is about 1000 bp, 2000 bp, 5000 bp, 10,000 bp, or 50,000 bp or greater. For example, Sequencing can provide sequence reads ranging in size from tens to hundreds to thousands of base pairs. For example, Parallel sequencing can provide sequence reads that do not vary much, for example, most sequence reads can be less than 200bp. A sequence read (or sequencing read) can refer to sequence information corresponding to a nucleic acid molecule (e.g., a string of nucleotides). For example, a sequence read can correspond to a string of nucleotides (e.g., about 20 to about 150) from a portion of a nucleic acid fragment, can correspond to a string of nucleotides at one or both ends of a nucleic acid fragment, or can correspond to nucleotides of the entire nucleic acid fragment. Sequence reads can be obtained in a variety of ways, for example, using sequencing technology or using probes, such as in hybridization arrays or capture probes; or amplification techniques, such as polymerase chain reaction (PCR) or linear amplification using a single primer, or isothermal amplification.

[0135] As used herein, the term "read fragment" refers to any form of nucleotide sequence read, including raw sequence reads directly obtained by nucleic acid sequencing technology, or sequences derived from raw sequence reads, such as aligned sequence reads, folded sequence reads, or spliced ​​sequence reads.

[0136] As used herein, the term "read count" refers to the total number of nucleic acid reads generated during a nucleic acid sequencing reaction, which may or may not be equal to the number of nucleic acid molecules generated.

[0137] As used in this article, the term "read depth", "sequencing depth" or "depth" can refer to the total number of unique nucleic acid fragments covering a specific locus or region of a microorganism genome obtained through sequencing in a specific sequencing reaction. Sequencing depth can be expressed as "Yx", such as 50x, 100x, etc., where "Y" refers to the quantity of unique nucleic acid fragments covering a specific locus obtained through sequencing in a sequencing reaction. In this case, Y must be an integer because it represents the actual sequencing depth of a specific locus. Alternatively, read depth, sequencing depth or depth can refer to a measure (such as mean or mode) of the central tendency of the number of unique nucleic acid fragments of one of the multiple loci or regions of a microorganism genome obtained through sequencing in a specific sequencing reaction. For example, in some embodiments, sequencing depth refers to the average depth of each locus in a targeted sequencing group (panel), exome or whole genome of a microorganism. In this case, Y can be expressed as a fraction or decimal because it refers to the average coverage of multiple loci. When narrating the depth average, the actual depth of any particular site may be different from the depth narrated overall. A metric for the sequencing depth range in which the total number of loci providing a defined percentage falls can be determined. For example, 90% or 95% or 99% of the loci fall within a sequencing depth range. As will be appreciated by those skilled in the art, different sequencing technologies provide different sequencing depths. For example, low-pass whole genome sequencing can refer to a technology that provides a sequencing depth of less than 5x, less than 4x, less than 3x or less than 2x (e.g., about 0.5x to about 3x).

[0138] As used herein, the term "sequencing breadth" refers to which portion of a particular microorganism's genome has been sequenced. Sequencing width can be expressed as a fraction, decimal, or percentage, and is typically calculated as (the total number of loci analyzed / the total number of loci in the genome). The denominator of this fraction can be a repeat-masked genome, and therefore 100% can correspond to all reference genomes minus the masked portion. A repeat-masked genome can refer to a genome in which sequence repeats are masked (e.g., sequence reads are aligned with the unmasked portion of the genome). In some embodiments, any portion of a genome can be masked, and therefore the sequencing width of any desired portion of a genome can be assessed.

[0139] As used herein, the terms "sequence ratio" and "coverage ratio" interchangeably refer to any measurement of the number of units of a genomic sequence in a first one or more biological samples (e.g., a test sample and / or a tumor sample) compared to the number of units of a corresponding genomic sequence in a second one or more biological samples (e.g., a reference sample and / or a control sample). In some embodiments, the sequence ratio is a copy ratio, a log2-converted copy ratio (e.g., a log2 copy ratio), a coverage ratio, a base fraction, an allele fraction (e.g., a variant allele fraction), and / or a tumor ploidy. In some embodiments, the sequence ratio is a log2-converted copy ratio. N The scaled copy ratio, where N is any real number greater than 1.

[0140] As used herein, the term "sequencing probe" refers to a molecule that binds nucleic acid with an affinity based on the expected nucleotide sequence of the RNA or DNA present at that locus.

[0141] As used herein, the term "targeted panel" or "targeted gene set" refers to a combination of probes used to sequence (e.g., by next-generation sequencing) nucleic acids present in a biological sample (e.g., a tumor sample, a liquid biopsy sample, a germline tissue sample, a white blood cell sample, or a tumor or tissue organoid sample) from a subject that are selected to locate one or more loci of interest in the genome.

[0142] As used herein, the term "sensitivity" or "true positive rate" (TPR) refers to the number of true positives divided by the sum of the number of true positives and false negatives. Sensitivity can characterize the ability of an assay or method to correctly identify the proportion of a population that actually has a certain condition. For example, sensitivity can characterize the ability of a method to correctly identify the number of subjects within a population that have a specific biological characteristic.

[0143] As used herein, the term "specificity" or "true negative rate" (TNR) refers to the number of true negatives divided by the sum of the number of true negatives and false positives. Specificity can characterize the ability of an assay or method to correctly identify the proportion of a population that does not actually have a certain condition. For example, specificity can characterize the ability of a method to correctly identify the number of subjects within a population that do not have a particular biological characteristic.

[0144] As used interchangeably herein, the terms "classifier" or "model" refer to a machine learning model or algorithm.

[0145] In some embodiments, model includes unsupervised learning algorithm. An example of unsupervised learning algorithm is cluster analysis. In some embodiments, model includes supervised machine learning. Non-limiting examples of supervised learning algorithm include but are not limited to logistic regression, neural network, support vector machine, naive Bayes algorithm, nearest neighbor algorithm, random forest algorithm, decision tree algorithm, boosting tree algorithm, multinomial logistic regression algorithm, linear model, linear regression, gradient boosting, hybrid model, hidden Markov (Markov) model, Gaussian NB algorithm, linear discriminant analysis, diffusion model or any combination thereof. In some embodiments, model is a polynomial classifier algorithm. In some embodiments, model is a 2-stage stochastic gradient descent (SGD) model. In some embodiments, model is a deep neural network (e.g., a deep and wide sample-level model).

[0146] Neural Networks. In some embodiments, the model is a neural network (e.g., a convolutional neural network and / or a residual neural network). Neural network algorithms (also known as artificial neural networks (ANNs)) include convolutional neural network algorithms and / or residual neural network algorithms (deep learning algorithms). In some embodiments, a neural network is a machine learning algorithm trained to map an input data set to an output data set, wherein the neural network includes a group of interconnected nodes organized into multiple layers of nodes. For example, in some embodiments, a neural network architecture includes at least an input layer, one or more hidden layers, and an output layer. In some embodiments, a neural network includes any total number of layers and any number of hidden layers, wherein the hidden layers serve as trainable feature extractors that allow a set of input data to be mapped to one or a set of output values. In some embodiments, a deep learning algorithm is a neural network that includes multiple hidden layers (e.g., two or more hidden layers). In some cases, each layer of a neural network includes multiple nodes (or "neurons"). In some embodiments, a node receives input directly from the input data or from the output of a node in a previous layer and performs a specific operation, such as a summation operation. In some embodiments, the connection from the input to the node is associated with a parameter (e.g., a weight and / or a weighting factor). In some embodiments, a node performs a summation of all input pairs x. iand the product of its associated parameters. In some embodiments, the weighted sum offset has a bias b. In some embodiments, a threshold or activation function f is used to gate the output of a node or neuron, and f is, in some cases, a linear or nonlinear function. In some embodiments, the activation function is, for example, a rectified linear unit (ReLU) activation function, a leaky ReLU activation function, or other functions such as a saturated hyperbolic tangent function, an identity function, a binary step function, a logistic function, an arcTan function, a softsign function, a parametric rectified linear unit function, an exponential linear unit function, a softPlus function, a curved identity function, a softExponential function, a sinusoid function, a sine function, a Gaussian function, or a sigmoid function, or any combination thereof.

[0147] In some embodiments, one or more sets of training data are used during a training phase to "teach" or "learn" weighting factors, bias values, thresholds, or other computational parameters of a neural network. For example, in some embodiments, input data from a training dataset and a gradient descent or backpropagation method are used to train the parameters so that the output values ​​calculated by the ANN are consistent with the examples included in the training dataset. In some embodiments, the parameters are obtained from a backpropagation neural network training process.

[0148] Any of a variety of neural networks are suitable for use in accordance with the present disclosure. Examples include, but are not limited to, feedforward neural networks, radial basis function networks, recurrent neural networks, residual neural networks, convolutional neural networks, residual convolutional neural networks, and the like, or any combination thereof. In some embodiments, machine learning utilizes pre-trained and / or transfer-learned ANNs or deep learning architectures. In some embodiments, convolutional neural networks and / or residual neural networks are used in accordance with the present disclosure.

[0149] For example, a deep neural network model includes an input layer, a plurality of individually parameterized (e.g., weighted) convolutional layers, and an output scorer. Each of the convolutional layers and the parameters (e.g., weights) of the input layer contribute to a plurality of parameters (e.g., weights) associated with the deep neural network model. In some embodiments, at least 50 parameters, at least 100 parameters, at least 1000 parameters, at least 2000 parameters, or at least 5000 parameters are associated with the deep neural network model. Therefore, deep neural network models require the use of computers because they cannot be solved with brain power. In other words, given the model input, in such embodiments the model output needs to be determined using a computer rather than with brain power. See, e.g., Krizhevsky et al., 2012, “Imagenet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems 2, Pereira, Burges, Bottou, Weinberger, eds., pp. 1097-1105, Curran Associates, Inc.; Zeiler, 2012 “ADADELTA: anadaptive learning rate method,” CoRR, vol. abs / 1212.5701; and Rumelhart et al., 1988, “Neurocomputing: Foundations of research,” chapter Learning Representations by Back-propagating Errors, pp. 696-699, Cambridge, MA, USA: MIT Press, each of which is hereby incorporated by reference.

[0150] Neural network algorithms (including convolutional neural network algorithms) suitable for use as models are disclosed in, for example, Vincent et al., 2010, “Stacked denoising autoencoders: Learning useful representations in adeep network with a local denoising criterion,” J Mach Learn Res 11, pp. 3371-3408; Larochelle et al., 2009, “Exploring strategies for training deep neural networks,” J Mach Learn Res 10, pp. 1-40; and Hassoun, 1995, Fundamentals of Artificial Neural Networks, Massachusetts Institute of Technology, each of which is hereby incorporated by reference. Additional example neural networks suitable for use as models are disclosed in Duda et al., 2001, Pattern Classification, 2nd ed., John Wiley & Sons, Inc., New York; and Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, each of which is hereby incorporated by reference in its entirety. Other example neural networks suitable for use as models are also described in Draghici, 2003, Data Analysis Tools for DNA Microarrays, Chapman & Hall / CRC; and Mount, 2001, Bioinformatics: sequence and genome analysis, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York, each of which is hereby incorporated by reference in its entirety.

[0151] Support Vector Machine. In some embodiments, the model is a support vector machine (SVM). An SVM algorithm suitable for use as a model is described, for example, in Cristianini and Shawe-Taylor, 2000, “An Introduction to Support Vector Machines,” Cambridge University Press, Cambridge; Boser et al., 1992, “A training algorithm for optimal margin classifiers,” in Proceedings of the 5th Annual ACM Workshop on Computational Learning Theory, ACM Press, Pittsburgh, Pa., pp. 142-152; Vapnik, 1998, Statistical Learning Theory, Wiley, New York; Mount, 2001, Bioinformatics: sequence and genome analysis, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY; Duda, Pattern Classification, 2nd ed., 2001, John Wiley & Sons, Inc., pp. 259, 262-265; and Hastie, 2001, The Elements of Statistical Learning, Springer, New York; and Furey et al., 2000, Bioinformatics 16, 906-914, the entire contents of each of these documents are hereby incorporated by reference. When used for classification, SVM separates a given binary labeled data set with a hyperplane that is maximally away from the labeled data. For certain situations where linear separation is impossible, SVM operates in conjunction with a 'kernel' technique that automatically implements a nonlinear mapping of the feature space. The hyperplane found by the SVM in the feature space corresponds to a nonlinear decision boundary in the input space in some cases. In some embodiments, a plurality of parameters (e.g., weights) associated with the SVM define the hyperplane. In some embodiments, the hyperplane is defined by at least 10, at least 20, at least 50, or at least 100 parameters, and the SVM model requires a computer to calculate because it cannot be solved with brain power.

[0152] Naive Bayes Algorithm. In some embodiments, the model is a Naive Bayes algorithm. Naive Bayes models suitable for use as models are disclosed, for example, in Ng et al., 2002, "On discriminative vs. generative classifiers: A comparison of logistic regression and naive Bayes," Advances in Neural Information Processing Systems, 14, which is hereby incorporated by reference. Naive Bayes models are any model in the family of "probabilistic models" based on applying Bayes' theorem and strong (naive) independence assumptions between features. In some embodiments, they are combined with kernel density estimation. See, for example, Hastie et al., 2001, The elements of statistical learning: data mining, inference, and prediction, Tibshirani and Friedman, eds., Springer, New York, which is hereby incorporated by reference.

[0153] Nearest neighbor algorithm. In some embodiments, the model is a nearest neighbor algorithm. In some embodiments, the nearest neighbor model is memory-based and does not include a model to be fitted. For the nearest neighbor, given a query point x0 (test object), identify k training points x (r) , r, ..., the k nearest neighbors to x0 (here the training subjects), and then the k nearest neighbors are used to classify the point x0. In some embodiments, the distance is determined as d using the Euclidean distance in the feature space. (i) =||x (i) -x (o) ||. Typically, when using the nearest neighbor algorithm, the abundance data used to calculate the linear discriminant is normalized to have a mean of 0 and a variance of 1. In some embodiments, the nearest neighbor rule is refined to address issues of unequal class priors, differential misclassification costs, and feature selection. Many of these refinements involve some form of weighted voting on neighbors. For more information on nearest neighbor analysis, see Duda, Pattern Classification, Second Edition, 2001, John Wiley & Sons, Inc; and Hastie, 2001, The Elements of Statistical Learning, Springer, New York, each of which is incorporated herein by reference.

[0154] The k-nearest neighbor model is a non-parametric machine learning method in which the input consists of the k closest training instances in the feature space. The output is category membership. An object is classified by a majority vote of its neighbors, where the object is assigned to the most common category among its k nearest neighbors (k is a positive integer, usually very small). If k=1, the object will simply be assigned to the category of the single nearest neighbor. See, Duda et al., 2001, Pattern Classification, 2nd Edition, John Wiley & Sons, which is incorporated herein by reference. In some embodiments, the number of distance calculations required to solve the k-nearest neighbor model makes it possible to use a computer to solve the model for a given input because it cannot be performed mentally.

[0155] Random forest, decision tree and boosted tree algorithms. In some embodiments, the model is a decision tree. The following reference generally describes decision trees suitable for use as models: Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, pp. 395-396, which is hereby incorporated by reference. Tree-based methods divide the feature space into a set of rectangles and then fit a model (such as a constant) in each rectangle. In some embodiments, the decision tree is a random forest regression. For example, a specific algorithm is classification and regression tree (CART). Other specific decision tree algorithms include, but are not limited to, ID3, C4.5, MART and random forest. CART, ID3 and C4.5 are described in the following reference: Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, pp. 396-408 and 411-412, which is hereby incorporated by reference. CART, MART, and C4.5 are described in Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, Chapter 9, which is hereby incorporated by reference in its entirety. Random forests are described in Breiman, 1999, "Random Forests—Random Features," Technical Report 567, Statistics Department, UC Berkeley, September 1999, which is hereby incorporated by reference in its entirety. In some embodiments, the decision tree model includes at least 10, at least 20, at least 50, or at least 100 parameters (e.g., weights and / or decisions) and requires a computer to calculate because it cannot be solved mentally.

[0156] Regression. In some embodiments, the model uses a regression algorithm. In some embodiments, the regression algorithm is any type of regression. For example, in some embodiments, the regression algorithm is logistic regression. In some embodiments, the regression algorithm is logistic regression with lasso regularization, L2 regularization, or elastic net regularization. In some embodiments, consideration is given to pruning (removing) extracted features whose corresponding regression coefficients fail to meet a threshold. In some embodiments, a generalized logistic regression model that handles multi-category responses is used as the model. The logistic regression algorithm is disclosed in Agresti, An Introduction to Categorical Data Analysis, 1996, Chapter 5, pages 103-144, John Wiley & Son, New York, which is incorporated herein by reference. In some embodiments, the model utilizes the regression model disclosed in the following reference: Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York. In some embodiments, the logistic regression model includes at least 10, at least 20, at least 50, at least 100, or at least 1000 parameters (e.g., weights) and requires a computer to calculate because it cannot be solved mentally.

[0157] Linear Discriminant Analysis Algorithm. In some embodiments, linear discriminant analysis (LDA), normal discriminant analysis (NDA), or discriminant function analysis is a generalized Fisher linear discriminant method, which is a method used in statistics, pattern recognition, and machine learning to find linear feature combinations that characterize or distinguish between two or more classes of objects or events. In some embodiments, the resulting combination is used as a model (linear model) in some embodiments of the present disclosure.

[0158] Hybrid Models and Hidden Markov Models. In some embodiments, the model is a hybrid model, such as that described in McLachlan et al., Bioinformatics 18(3):413-422, 2002. In some embodiments, particularly those that include a temporal component, the model is a hidden Markov model, such as that described in Schliep et al., 2003, Bioinformatics 19(1):i255-i263.

[0159] Clustering. In some embodiments, the model is an unsupervised clustering model. In some embodiments, the model is a supervised clustering model. Clustering algorithms suitable for use as models are described in, for example, the following literature: Duda and Hart, Pattern Classification and Scene Analysis, pp. 211-256, 1973, John Wiley & Sons, Inc., New York, (hereinafter referred to as "Duda 1973"), which is incorporated herein by reference in its entirety. As an illustrative example, in some embodiments, the clustering problem is described as finding one of natural groupings in a data set. In order to identify natural groupings, two problems are solved. First, determine the way to measure the similarity (or dissimilarity) between two samples. This metric (e.g., similarity metric) is used to ensure that the samples in a cluster are more similar to each other than they are to the samples in other clusters. Second, determine the mechanism for partitioning data into clusters using a similarity metric. One method for starting a clustering study is to define a distance function and calculate the distance matrix between all pairs of samples in the training set. If distance is a good measure of similarity, the distance between reference entities in the same cluster is significantly smaller than the distance between reference entities in different clusters. However, in some embodiments, clustering does not use a distance metric. For example, in some embodiments, a non-metric similarity function s(x, x') can be used to compare two vectors x and x'. In some such embodiments, s(x, x') is a symmetric function whose value is larger when x and x' are "similar" in some way. Once a method for measuring the "similarity" or "dissimilarity" between points in a data set is selected, clustering uses a standard function that measures the clustering quality of any partition of the data. Data are clustered using a data set partition that extremes the standard function. Specific exemplary clustering techniques contemplated for use in the present disclosure include, but are not limited to, hierarchical clustering (clustering using a nearest neighbor algorithm, a farthest neighbor algorithm, an average connection algorithm, a centroid algorithm, or a sum of squares algorithm), k-means clustering, a fuzzy k-means clustering algorithm, and Jarvis-Patrick clustering. In some embodiments, clustering includes unsupervised clustering (e.g., without a preconceived number of clusters and / or without predetermining cluster allocation).

[0160] Integration and promotion of models. In some embodiments, an integration of models (two or more) is used. In some embodiments, promotion techniques such as AdaBoost are used in combination with many other types of learning algorithms to improve the performance of the model. In this method, the output of any model disclosed herein or its equivalent is combined into a weighted sum representing the final output of the promoted model. In some embodiments, any measure of central tendency known in the art (including but not limited to mean, median, mode, weighted mean, weighted median, weighted mode, etc.) is used to combine multiple outputs from the model. In some embodiments, a voting method is used to combine multiple outputs. In some embodiments, the corresponding models in the model integration are weighted or unweighted.

[0161] As used herein, the term "parameter" refers to any coefficient or similar value of an internal or external element (e.g., weight and / or hyperparameter) in an algorithm, model, regressor, and / or classifier that can affect (e.g., modify, customize, and / or adjust) one or more inputs, outputs, and / or functions in the algorithm, model, regressor, and / or classifier. For example, in some embodiments, a parameter refers to any coefficient, weight, and / or hyperparameter that can be used to control, modify, customize, and / or adjust the behavior, learning, and / or performance of an algorithm, model, regressor, and / or classifier. In some cases, a parameter is used to increase or decrease the effect of an input (e.g., a feature) on an algorithm, model, regressor, and / or classifier. As a non-limiting example, in some embodiments, a parameter is used to increase or decrease the effect of a node (e.g., a node of a neural network), where the node includes one or more activation functions. The assignment of parameters to a particular input, output, and / or function is not limited to any one paradigm for a given algorithm, model, regressor, and / or classifier, but can be used for any suitable algorithm, model, regressor, and / or classifier architecture to obtain the desired performance. In some embodiments, the parameter has a fixed value. In some embodiments, the values ​​of the parameters are adjusted manually and / or automatically. In some embodiments, the values ​​of the parameters are modified by a validation and / or training process (e.g., by error minimization and / or backpropagation methods) of the algorithm, model, regressor, and / or classifier. In some embodiments, the algorithms, models, regressors, and / or classifiers of the present disclosure include multiple parameters. In some embodiments, the plurality of parameters is n parameters, wherein: n≥2; n≥5; n≥10; n≥25; n≥40; n≥50; n≥75; n≥100; n≥125; n≥150; n≥200; n≥225; n≥250; n≥350; n≥500; n≥600; n≥750; n≥1,000; n≥2,000; n≥4,000; n≥5,000; n≥7,500; n≥10,000; n≥20,000; n≥40,000; n≥75,000; n≥100,000; n≥200,000; n≥500,000, n≥1 x 10 6 , n ≥ 5 x 10 6 or n ≥ 1 x 10 7 Therefore, the algorithms, models, regressors, and / or classifiers of the present disclosure cannot be performed using mental power. In some embodiments, n is between 10,000 and 1 x 10 7 Between 100,000 and 5x 10 6 between 500,000 and 1 x 10 6In some embodiments, the algorithms, models, regressors, and / or classifiers of the present disclosure operate in a k-dimensional space, where k is a positive integer of 5 or greater (e.g., 5, 6, 7, 8, 9, 10, etc.). Therefore, the algorithms, models, regressors, and / or classifiers of the present disclosure cannot be performed mentally.

[0162] As used herein, the term "untrained model" (e.g., "untrained classifier" and / or "untrained neural network") refers to a machine learning model or algorithm (such as a classifier or neural network) that has not been trained for a target dataset. In some embodiments, "training a model" (e.g., "training a neural network") refers to the process of training an untrained or partially trained model (e.g., "untrained or partially trained neural network"). In addition, it should be understood that the term "untrained model" does not exclude the possibility of using transfer learning techniques in such training of an untrained or partially trained model. For example, Fernandes et al., 2017, "Transfer Learning with PartialObservability Applied to Cervical Cancer Screening," Pattern Recognition and Image Analysis: 8th Iberian Conference Proceedings, 243-250, which is incorporated herein by reference, provides a non-limiting example of such transfer learning. When transfer learning is used, additional data is provided to the untrained model in addition to the main training dataset. Typically, this additional data is in the form of parameters (e.g., coefficients, weights, and / or hyperparameters) learned from another auxiliary training dataset. In addition, although a description of a single auxiliary training dataset has been disclosed, it should be understood that there is no limit to the number of auxiliary training datasets that can be used to supplement the primary training dataset when training an untrained model in the present disclosure. For example, in some embodiments, two or more auxiliary training datasets, three or more auxiliary training datasets, four or more auxiliary training datasets, or five or more auxiliary training datasets are used to supplement the primary training dataset through transfer learning, wherein each such auxiliary training dataset is different from the primary training dataset. In some such embodiments, any form of transfer learning is used. For example, consider the case where there is a first auxiliary training dataset and a second auxiliary training dataset in addition to the primary training dataset. In this case, a transfer learning technique (e.g., a second model that is the same or different from the first model) is used to apply the parameters learned from the first auxiliary training dataset (by applying the first model to the first auxiliary training dataset) to the second auxiliary training dataset, thereby generating a trained intermediate model, the parameters of which are then applied to the primary training dataset and applied to the untrained model in combination with the primary training dataset itself.Alternatively, in another exemplary embodiment, a first set of parameters learned from a first auxiliary training dataset (by applying a first model to the first auxiliary training dataset) and a second set of parameters learned from a second auxiliary training dataset (by applying a second model that is the same as or different from the first model to the second auxiliary training dataset) are each separately applied to a separate instance of the main training dataset (e.g., by separate independent matrix multiplications), and these two parameter applications to separate instances of the main training dataset are then combined with the main training dataset itself (or some simplified form of the main training dataset, such as principal components or regression coefficients learned from the main training set) to an untrained model in order to train the untrained model.

[0163] As used herein, the term "instruction" refers to a command given to a computer processor by a computer program. On a digital computer, each instruction is a sequence of 0s and 1s that describes the physical operation that the computer is to perform. Such instructions may include data transfer instructions and data manipulation instructions. In some embodiments, each instruction is an instruction type identified by a specific processor type for executing instructions in an instruction set. Examples of instruction sets include, but are not limited to, reduced instruction set computers (RISC), complex instruction set computers (CISC), minimal instruction set computers (MISC), very long instruction words (VLIW), explicitly parallel instruction computing (EPIC), and single instruction set computers (OISC).

[0164] Several aspects are described below with reference to the example applications used for illustration. It should be understood that many specific details, relationships, and methods are set forth to provide a complete understanding of the features described herein. However, those of ordinary skill in the relevant art will readily recognize that the features described herein can be practiced without one or more of the specific details or with other methods. The features described herein are not limited to the illustrated ordering of actions or events, as some actions can occur in different orders and / or simultaneously with other actions or events. In addition, implementation according to the methods of the features described herein does not require all illustrated actions or events.

[0165] Reference will now be made in detail to the embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent that one of ordinary skill in the art can practice the present disclosure without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail to avoid unnecessarily obscuring aspects of the embodiments.

[0166] Exemplary system embodiments.

[0167] Now that an overview of some aspects of the disclosure and some definitions used in the disclosure have been provided, Figure 1 Details of an exemplary system are described. Figure 1 1 is a block diagram illustrating a system 100 according to some embodiments. In some embodiments, the system 100 includes one or more processing units CPU 102 (also referred to as processors), one or more network interfaces 104, a user interface 106 including (optionally) a display 108 and an input system 110, non-persistent memory 111, persistent memory 112, and one or more communication buses 114 for interconnecting these components. The one or more communication buses 114 optionally include circuits (sometimes referred to as chipsets) that interconnect and control communications between system components. The non-persistent memory 111 typically includes high-speed random access memory such as DRAM, SRAM, DDR RAM, ROM, EEPROM, flash memory, while the persistent memory 112 typically includes a CD-ROM, digital versatile disk (DVD) or other optical storage, a cassette, a magnetic tape, a magnetic disk storage or other magnetic storage device, a magnetic disk storage device, an optical disk storage device, a flash memory device or other non-volatile solid-state storage device. The persistent memory 112 optionally includes one or more storage devices remote from the CPU 102. Persistent memory 112 and non-volatile storage within non-persistent memory 112 include non-transitory computer-readable storage media. In some implementations, non-persistent memory 111 or alternatively non-transitory computer-readable storage media sometimes in conjunction with persistent memory 112 store the following programs, modules, and data structures, or a subset thereof:

[0168] • an optional operating system 116, which includes procedures for handling various basic system services and for performing hardware-related tasks;

[0169] An optional network communication module (or instructions) 118 for connecting the system 100 to other devices and / or the communication network 104;

[0170] A microbiome assessment module 140 for determining one of a plurality of disease states in a subject based on the composition of the subject's microbiome; and

[0171] A data store 140 of subject information based on microbiome sequencing results 150 , including abundance values ​​152 of microorganisms in each of bacterial groups 152 -A and 152 -B as described herein.

[0172] In various embodiments, one or more of the above elements are stored in one or more previously mentioned memory devices, and correspond to the instruction set for performing the above functions. The modules, data or programs (e.g., instruction sets) identified above do not need to be implemented as independent software programs, processes, data sets or modules, and therefore the various subsets of these modules and data can be combined or otherwise rearranged in various implementations. In some implementations, non-persistent memory 111 optionally stores the modules and data structures identified above. In addition, in some embodiments, memory stores additional modules and data structures not described above. In certain embodiments, one or more of the above-identified elements are stored in a computer system other than visualization system 100, which can be addressed by visualization system 100 so that visualization system 100 can retrieve all or part of these data when needed.

[0173] although Figure 1 A "system 100" is depicted, but the diagram is intended more as a functional description of various features that may be present in a computer system than as a schematic diagram of the architecture of the embodiments described herein. In practice, and as will be appreciated by one of ordinary skill in the art, items shown separately may be combined and some items may be separated. Furthermore, although Figure 1 Certain data and modules are described in non-persistent storage 111 , but some or all of these data and modules may be stored in persistent storage 112 .

[0174] 1. Methods for identifying a panel of gut microbes

[0175] FIG2 is a schematic diagram of a method for identifying a group of intestinal microorganisms as discussed below. Figure 1 The present invention may be implemented by the computer system 100 shown and described herein.

[0176] Referring to block 200, in some embodiments, the method includes electronically obtaining, for each respective subject in a first plurality of subjects having a first state of a biological characteristic, a corresponding plurality of genomic abundance values, the plurality of genomic abundance values ​​comprising, for each respective gut microbe in a plurality of gut microbes, a corresponding value for the abundance of a genome of the respective gut microbe in a biological sample from the gut of the respective subject. In some embodiments, the first plurality of subjects includes at least 50, at least 100, at least 200, at least 500, at least 1000, at least 5000, at least 10,000, at least 20,000, at least 50,000, at least 100,000, at least 500,000, or at least 1,000,000 subjects. In some embodiments, the first plurality of subjects comprises no more than 1,000,000, no more than 500,000, no more than 100,000, no more than 50,000, no more than 20,000, no more than 10,000, no more than 1000 subjects, no more than 500 subjects, no more than 100 subjects, or no more than 50 subjects. In some embodiments, the first plurality of subjects consists of 50 to 100, 50 to 200, 50 to 500, 100 to 500, 200 to 500, 200 to 1000, 500 to 1000, 200 to 5,000, 1000 to 10,000, 5000 to 200,000, 10,000 to 50,000, 20,000 to 100,000, or 500,000 to 1,000,000. In some embodiments, the first plurality of subjects falls within another range that begins at no less than 50 subjects and ends at no more than 10,000,000 subjects. In some embodiments, the first plurality of subjects have similar demographic characteristics (such as age, sex, race). In some embodiments, the first plurality of subjects have similar physical characteristics (such as weight, height, BMI values). In some embodiments, the first plurality of subjects have similar health conditions (such as physical or mental conditions, medical history, gene carriers, or drug use). In some embodiments, the first plurality of subjects have similar behavior and lifestyle preferences (such as diet, physical exercise, or substance use).

[0177] In certain embodiments, the corresponding value of the abundance of a genome is a value representing the absolute abundance of a microorganism genome. In certain embodiments, the corresponding value of the abundance of a genome is a value representing a normalized abundance value or a relative abundance value (for example, the abundance of a microorganism is normalized relative to the abundance of a total microbial group of interest). In certain embodiments, the corresponding value of the abundance of a genome is a value representing an average abundance value (for example, the mean value of the abundance obtained at different time points or from different biological samples from a patient, or the mean value of the abundance obtained using different probes, etc.), or a combination of any one of the above. The corresponding value of the abundance of a genome is measured by any technology known in the art. In certain embodiments, the genomic abundance value of a genome is measured by quantitative PCR (qPCR) such as bacterial 16S rRNA qPCR, RT-PCR or qRT-PCR, for quantitative genome abundance in a region of interest, for example, as described in U.S. Patent number 11,427,865, the disclosure of which is incorporated herein by reference in its entirety. In some embodiments, the genomic abundance value is measured as follows: using targeted sequencing (e.g., 16S rRNA sequencing or any other suitable biomarker), partial genome sequencing, or whole genome sequencing, thereby quantifying the number of reads of the targeted region in the microbial genome to determine the abundance of the genome, for example, as disclosed in U.S. Patent Application Publication No. 2021 / 0403986 or U.S. Patent No. 11,332,783, the disclosures of which are hereby incorporated by reference in their entirety. In some embodiments, deep sequencing is used to determine the abundance of the targeted sequence, for example, as disclosed in U.S. Patent Application Publication No. 2018 / 0237863, the disclosures of which are hereby incorporated by reference in their entirety. In some embodiments, the sequencing depth is at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 150, at least 200, at least 300, at least 400, at least 500, at least 750, at least 1000 or more.In some embodiments, shotgun metagenomic sequencing is used to provide sequence reads of the genome in a sample, for example, as described in U.S. Patent No. 11,028,449, the contents of which are incorporated herein by reference in their entirety.

[0178] Referring to block 202, in some embodiments, for each corresponding subject in the first plurality of subjects, the biological sample from the intestine of the corresponding subject is a stool sample. In some embodiments, the sample is a tissue biopsy sample, an intestinal sample, or a mucosal sample. See, e.g., Tang Q, J et al., Current Sampling Methods for Gut Microbiota: A Call for More Precise Devices, Front Cell Infect Microbio1., 10: 151 (2020), the contents of which are incorporated herein by reference in their entirety.

[0179] Referring to block 204, in some embodiments, the method includes, for each corresponding subject in the first plurality of subjects, sequencing genomic DNA from a corresponding biological sample from the intestinal tract of the corresponding subject, thereby obtaining a corresponding first plurality (e.g., at least 100,000) of nucleic acid sequences. In some embodiments, the first plurality of nucleic acid sequences comprises at least 100,000, at least 250,000, at least 500,000, at least 1,000,000, at least 2,500,000, at least 5,000,000, at least 10,000,000, or at least 50,000,000 nucleic acid sequences. In some embodiments, the first plurality of nucleic acid sequences comprises no more than 250,000,000, no more than 100,000,000, no more than 50,000,000, no more than 25,000,000, no more than 10,000,000, no more than 5,000,000, no more than 1,000,000, or no more than 100,000 nucleic acid sequences. In some embodiments, the first plurality of nucleic acid sequences consists of 100,000 to 1,000,000, 200,000 to 5,000,000, 500,000 to 10,000,000, 1,000,000 to 20,000,000, 5,000,000 to 50,000,000, 10,000,000 to 100,000,000, or 50,000,000 to 250,000,000 nucleic acid sequences. In some embodiments, the first plurality of nucleic acid sequences falls within another range that begins with no less than 100,000 nucleic acid sequences and ends with no more than 250,000,000 nucleic acid sequences.

[0180] In some embodiments, a first plurality (e.g., at least 100,000) of nucleic acid sequences are obtained by metagenomic sequencing, for example, as disclosed in U.S. Patent Application Publication No. 2016 / 0239602 or U.S. Patent No. 11,495,326, the contents of which are incorporated herein by reference in their entirety. In some embodiments, metagenomic sequencing further includes generating a plurality of metagenomic fragment reads. In some embodiments, metagenomic sequencing further includes fragmenting the microbial genome into random fragments of a target size. The size of the resulting fragments may vary. In one embodiment, a fragment of approximately 500 nucleotides may be obtained. In some embodiments, a fragment of 100-2000 nucleotides (e.g., 200-800, 100-900, 100-1000, 300-800, 400-900 nucleotides) may be obtained. In some embodiments, the method may further include extracting metagenomic fragments from corresponding biological samples. In some embodiments, metagenomic sequencing further includes sequencing the fragments using a high-throughput sequencing method to generate a plurality of sequencing reads.

[0181] In some embodiments, the first plurality (e.g., at least 100,000) of nucleic acid sequences are obtained by targeted panel sequencing, e.g., as described in U.S. Patent Application Publication No. 2019 / 0316209. In some embodiments, targeted panel sequencing comprises hybridizing genomic DNA isolated from a biological sample from the intestinal tract of a subject with a set of probes that are specific for each microorganism being quantified (e.g., Table 1, Table 2, and / or Table 3). Figures 42A to 42XX In some embodiments, the probe set comprises one or more probes that hybridize to unique sequences in the genome of each of the plurality of microorganisms listed in (e.g., each of the plurality of microorganisms listed in); the recovered nucleic acid is then sequenced. In some embodiments, a combination of semi-unique sequences (e.g., sequences found in a small number of microbial genomes) can be used to deconvolute genomic abundance values ​​using an algorithm (e.g., a system of equations). In some embodiments, the probe set comprises at least one probe that hybridizes to a unique sequence in each microbial genome to be detected. In some embodiments, the probe set comprises at least 2, at least 3, at least 4, at least 5, at least 10, at least 25, at least 50 or more probes that hybridize to unique different sequences in each microbial genome to be detected. In some embodiments, the probe sets include at least 20, at least 30, at least 40, at least 50, at least 75, at least 100, at least 125, at least 150, at least 200, at least 150, at least 300, at least 400, at least 500, at least 750, at least 1000, at least 1250, at least 1500, at least 2000, at least 2500, at least 3000, at least 4000, at least 5000, at least 7500, at least 10,000 or more unique probes.

[0182] In some embodiments, the sequenced genomic DNA from the corresponding biological sample comprises at its end a partial or complete sequencing platform adapter sequence that can be used for sequencing using a sequencing platform of interest. Sequencing platforms of interest include, but are not limited to, HiSeq TM , MiSeq TM and Genome Analyzer TM Sequencing system; from Ion Torrent TM Ion PGM TM and Ion Proton TM Sequencing system: PACBIO RS II Sequel system from Pacific Biosciences, from Life Technologies TM SOLiD sequencing system from Roche, 454 GS FLX+ and GS Junior sequencing systems from Roche, MinION from Oxford Nanopore TM system or any other sequencing platform of interest.

[0183] In some embodiments, multiple genomic abundance values ​​are determined using a microarray comprising probe sequences capable of detecting a unique genomic sequence for each corresponding genome of a plurality of intestinal microorganisms. In some embodiments, the probe set on the microarray comprises at least one probe that hybridizes to a unique sequence for each genome of the microorganism to be detected. In some embodiments, the probe set comprises at least two, at least three, at least four, at least five, at least ten, at least twenty-five, at least fifty, or more probes that hybridize to a unique, different sequence for each genome of the microorganism to be detected. In some embodiments, the probe sets include at least 20, at least 30, at least 40, at least 50, at least 75, at least 100, at least 125, at least 150, at least 200, at least 150, at least 300, at least 400, at least 500, at least 750, at least 1000, at least 1250, at least 1500, at least 2000, at least 2500, at least 3000, at least 4000, at least 5000, at least 7500, at least 10,000 or more unique probes.

[0184] Referring to block 206, in some embodiments, the method includes electronically obtaining, for each respective subject in the first plurality of subjects, a corresponding first plurality of nucleic acid sequences (e.g., at least 100,000 nucleic acid sequences) of genomic DNA from a corresponding biological sample from the intestine of the respective subject, and determining, for each respective subject in the first plurality of subjects, a corresponding genomic abundance value for each respective gut microbe in the plurality of gut microbes from the corresponding first plurality (e.g., at least 100,000) nucleic acid sequences. In some embodiments, the genomic abundance values ​​determined for each respective subject in the first plurality of subjects include at least 20, at least 25, at least 50, at least 100, at least 250, at least 500, at least 1,000, at least 5,000, or at least 10,000 genomic abundance values. In some embodiments, the genomic abundance values ​​determined for each respective subject in the first plurality of subjects comprise no more than 250,000, no more than 100,000, no more than 50,000, no more than 25,000, no more than 10,000, no more than 5,000, no more than 1,000, no more than 100, no more than 50, no more than 30, or no more than 20 genomic abundance values. In some embodiments, the genomic abundance values ​​determined for each respective subject in the first plurality of subjects consist of 10 to 40, 20 to 50, 30 to 80, 40 to 100, 50 to 150, 60 to 200, 80 to 300, 90 to 500, 100 to 1000, 500 to 2,000, or 1,000 to 5,000 genomic abundance values. In some embodiments, the genomic abundance values ​​determined for each respective subject in the first plurality of subjects fall within another range starting at no less than 10 genomic abundance values ​​and ending at no more than 250,000 genomic abundance values.

[0185] With reference to box 208, in some embodiments, the method includes, for each corresponding subject in the first plurality of subjects, assembling the corresponding first plurality of intestinal microorganism genomes from the corresponding first plurality of (e.g., at least 100,000) nucleic acid sequences by metagenome de novo sequence assembly, and for each corresponding intestinal microorganism genome in the corresponding first plurality of intestinal microorganism genomes, calculating the corresponding genome abundance of the corresponding intestinal microorganism genome. In some embodiments, the metagenome de novo sequence assembly further includes generating overlapping groups based on sequencing reads generated by shotgun sequencing technology, for example, as described in U.S. Patent No. 10,529,443, the contents of which are incorporated herein by reference in their entirety. In some embodiments, the first plurality of (e.g., at least 100,000) nucleic acid sequences can be assembled into the complete genomes of a variety of intestinal microorganisms. In some embodiments, the first plurality of (e.g., at least 100,000) nucleic acid sequences can be assembled into partial genomes of a variety of intestinal microorganisms.

[0186] With reference to box 210, in some embodiments, the method includes, for each corresponding subject in the first plurality of subjects, assigning each corresponding nucleic acid sequence in the corresponding first plurality (e.g., at least 100,000) sequences to a corresponding intestinal microbe in a plurality of intestinal microbes, thereby generating a corresponding count of the corresponding nucleic acid sequence in the corresponding first plurality of nucleic acid sequences assigned to the corresponding intestinal microbe for each corresponding intestinal microbe in the plurality of intestinal microbes, and for each corresponding intestinal microbe in the plurality of intestinal microbes, determining a corresponding genome abundance value for the corresponding intestinal microbe based on the corresponding count of the corresponding nucleic acid sequence assigned to the corresponding intestinal microbe. In some embodiments, assigning each corresponding nucleic acid to the corresponding intestinal microbe includes mapping the nucleic acid to a reference nucleic acid, e.g., a contig listed in Figure 41. In some embodiments, assigning each corresponding nucleic acid to the corresponding intestinal microbe includes annotating genomic information based on an existing database. In some embodiments, the nucleic acid sequences are analyzed, and the annotation is to define taxonomic assignments using sequence similarity and phylogenetic placement methods or a combination of the two strategies.

[0187] The method based on sequence similarity includes those methods familiar to those skilled in the art, includes but is not limited to BLAST, BLASTx, tBLASTn, tBLASTx, RDP sorter, DNAclust and the various embodiments (such as Qiime or Mothur) of these algorithms.These methods rely on sequence reading and mapping to reference database and selecting the matching with best score and e value.In certain embodiments, phylogenetic method is used in combination with sequence similarity method to improve the calling accuracy of annotation or classification distribution.Common databases include but are not limited to GT-DBTK, National Center for Biotechnology Information (NCBI) Genbank, European Bioinformatics Institute-European Nucleotide Archive (European Bioinformatics Institute-European Nucleotide Archive; EBI-ENA), U.S. Department of Energy National Institute of Genetics (USDOE) integrated microbial genome (Integrated Microbial Genomes) & Microbiomes; IMG / M) and other available databases in this area.

[0188] Referring to block 212, in some embodiments, the first state of the biological characteristic is the absence of a disease or condition, such as type 2 diabetes (T2D), hypertension (HT), schizophrenia (SCZ), atherosclerotic cardiovascular disease (ACVD), cirrhosis (LC), inflammatory bowel disease (IBD), colorectal cancer (CRC), ankylosing spondylitis (AS), Parkinson's disease (PD), multiple sclerosis (MS), Gaucher disease type II (GDII), COVID-19 (COVID-19), Behçet's disease (BD), autism spectrum disorder (ASD), or pancreatic cancer (PC). In some embodiments, the disease or condition is cancer, Alzheimer's disease, cardiovascular disease, autoimmune disease, mental health disorder, infectious disease, or genetic disorder.

[0189] In some embodiments, the first state of a biological characteristic is the first severity of a disease or condition. In some embodiments, the severity of the disease is classified according to the type, frequency, or intensity of the subject's experience. In some embodiments, the severity of the disease is classified according to the progression or prognosis of the disease or condition (e.g., different stages of cancer). In some embodiments, the first state of a biological characteristic is an untreated disease or condition. In some embodiments, the first state of a biological characteristic is a disease or condition treated with a first therapy (e.g., surgery, radiotherapy, chemotherapy, targeted therapy, gene therapy, immunotherapy, drug therapy, dietary changes, lifestyle changes).

[0190] In some embodiments, the first state of the biological characteristic is a first level of a nutrient in the diet (such as carbohydrates, protein, fat, vitamins, fiber). In some embodiments, the first state of the biological characteristic is a first age. In some embodiments, a threshold value (e.g., a level of a biomarker, a diagnostic cutoff value, or a threshold nutrient intake level) is provided for determining the first state of the biological characteristic.

[0191] Referring to block 214, in some embodiments, the plurality of intestinal microorganisms comprises a member selected from Table 1, Table 2, or Figures 42A to 42XX In some embodiments, at least about 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more intestinal microorganisms are selected from Table 1, Table 2 or Table 3. Figures 42A to 42XX .

[0192] Table 1 - Taxonomic assignment of 141 non-redundant genomes identified in two competing bacterial communities

[0193]

[0194]

[0195]

[0196]

[0197]

[0198]

[0199]

[0200]

[0201]

[0202]

[0203]

[0204]

[0205]

[0206]

[0207]

[0208]

[0209]

[0210]

[0211]

[0212]

[0213]

[0214]

[0215]

[0216]

[0217]

[0218]

[0219]

[0220]

[0221]

[0222]

[0223]

[0224]

[0225]

[0226] Table 1, Table 2 and Figures 42A to 42XX The bacterial species listed in were identified by metagenomic sequencing of genomic DNA isolated from human fecal samples and were determined to be part of two competing microbial communities with respect to at least one biological characteristic, as described in the Examples. Briefly, genomic DNA isolated from each fecal sample was sequenced by next generation sequencing, and contigs of the microbial genome sequences were constructed de novo. In general, the contigs identified for each microorganism were expected to represent more than 95% of the entire genome of that microorganism. Genomic constructs that differed less than 1% from each other in sequence were combined and defined as being from the same microorganism. Tables 1, 2, and Figures 42A to 42XX The genomic contigs for each microorganism listed in are provided in the sequence listings submitted with this application. The taxonomic assignments for each microorganism are given in Table 1, Table 2, or Figures 42A to 42XX FIG41 provides a correspondence between the sequence identifier assigned to each contig and the microorganism to which it belongs. For example, the contig provided as SEQ ID NO: 1-68 corresponds to the genomic sequence of microorganism 1U001.8 (e.g. Figure 41A ), microorganism 1U001.8 is a microorganism classified as follows: Bacteria domain, Proteobacteria phylum, Gammaproteobacteria class, Enterobacteriales order, Enterobacteriaceae family, Escherichia genus and Escherichia coli species, and belongs to group 2 among the 141 core microorganisms identified in Table 1.

[0227] Thus, in some embodiments of the methods described herein, a genome identified in a metagenomic analysis is classified as corresponding to Table 1, Table 2, and / or Table 3 if the identified genomic construct has at least 97% sequence identity when compared to the microbial contigs provided in the sequence listing. Figures 42A to 42XX 41. In some embodiments, a genome identified in a metagenomic analysis is classified as corresponding to Table 1, Table 2, and / or Table 3 if the identified genomic construct has at least 98% sequence identity when compared to the microbial contigs provided in the sequence listing. Figures 42A to 42XX 41. In some embodiments, a genome identified in a metagenomic analysis is classified as corresponding to Table 1, Table 2, and / or Table 3 if the identified genomic construct has at least 99% sequence identity when compared to the microbial contigs provided in the sequence listing. Figures 42A to 42XX 41. In some embodiments, a genome identified in a metagenomic analysis is classified as corresponding to Table 1, Table 2, and / or Table 3 if the identified genomic construct has at least 99.5% sequence identity when compared to the microbial contigs provided in the sequence listing. Figures 42A to 42XX 41. In some embodiments, a genome identified in the metagenomic analysis is classified as corresponding to Table 1, Table 2, and / or Table 3 if the identified genomic construct has at least 97%, at least 97.5%, at least 98%, at least 98.5%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or more sequence identity when compared to a contig of a microorganism provided in the sequence listing. Figures 42A to 42XX The microorganisms listed in are shown in Figure 41.

[0228] Referring to block 216, in some embodiments, the method includes electronically obtaining, for each respective subject in a second plurality of subjects having a second state of the biological characteristic, a corresponding plurality of genomic abundance values, the plurality of genomic abundance values ​​comprising, for each respective gut microbe in the plurality of gut microbes, a corresponding value for the abundance of a genome of the respective gut microbe in a biological sample from the gut of the respective subject. In some embodiments, the second plurality of subjects includes at least 50, at least 100, at least 200, at least 500, at least 1000, at least 5000, at least 10,000, at least 20,000, at least 50,000, at least 100,000, at least 500,000, or at least 1,000,000 subjects. In some embodiments, the second plurality of subjects comprises no more than 1,000,000, no more than 500,000, no more than 100,000, no more than 50,000, no more than 20,000, no more than 10,000, no more than 1000 subjects, no more than 500 subjects, no more than 100 subjects, or no more than 50 subjects. In some embodiments, the second plurality of subjects consists of 50 to 100, 50 to 200, 50 to 500, 100 to 500, 200 to 500, 200 to 1000, 500 to 1000, 200 to 5,000, 1000 to 10,000, 5000 to 200,000, 10,000 to 50,000, 20,000 to 100,000, or 500,000 to 1,000,000. In some embodiments, the second plurality of subjects falls within another range that begins at no less than 50 subjects and ends at no more than 10,000,000 subjects. In some embodiments, the second plurality of subjects have similar demographic characteristics (such as age, sex, race). In some embodiments, the second plurality of subjects have similar physical characteristics (such as weight, height, BMI values). In some embodiments, the second plurality of subjects have similar health conditions (such as physical or mental conditions, medical history, gene carriers, or drug use). In some embodiments, the second plurality of subjects have similar behavior and lifestyle preferences (such as diet, physical exercise, or substance use).

[0229] In certain embodiments, the corresponding value of the abundance of a genome is a value representing the absolute abundance of a microorganism genome. In certain embodiments, the corresponding value of the abundance of a genome is a value representing a normalized abundance value or a relative abundance value (for example, the abundance of a microorganism is normalized relative to the abundance of a total microbial group of interest). In certain embodiments, the corresponding value of the abundance of a genome is a value representing an average abundance value (for example, the mean value of the abundance obtained at different time points or from different biological samples from a patient, or the mean value of the abundance obtained using different probes, etc.), or a combination of any one of the above. The corresponding value of the abundance of a genome is measured by any technology known in the art. In certain embodiments, the genomic abundance value of a genome is measured by quantitative PCR (qPCR) such as bacterial 16S rRNA qPCR, RT-PCR or qRT-PCR, for quantitative genome abundance in a region of interest, for example, as described in U.S. Patent number 11,427,865, the disclosure of which is incorporated herein by reference in its entirety. In some embodiments, the genomic abundance value is measured as follows: using targeted sequencing (e.g., 16S rRNA sequencing or any other suitable biomarker), partial genome sequencing, or whole genome sequencing, thereby quantifying the number of reads of the targeted region in the microbial genome to determine the abundance of the genome, for example, as disclosed in U.S. Patent Application Publication No. 2021 / 0403986 or U.S. Patent No. 11,332,783, the disclosures of which are hereby incorporated by reference in their entirety. In some embodiments, deep sequencing is used to determine the abundance of the targeted sequence, for example, as disclosed in U.S. Patent Application Publication No. 2018 / 0237863, the disclosures of which are hereby incorporated by reference in their entirety. In some embodiments, the sequencing depth is at least about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 70, 80, 90, 100, 110, 120, 130, 150, 200, 300, 500, 500, 700, 1000 or more. In some embodiments, shotgun metagenomic sequencing is used to provide sequence reads of the genome in a sample, for example, as described in U.S. Patent No. 11,028,449, the contents of which are incorporated herein by reference in their entirety.

[0230] Referring to block 218, in some embodiments, for each respective subject in the second plurality of subjects, the biological sample from the intestine of the respective subject is a stool sample. In some embodiments, the biological sample is selected from a tissue biopsy sample, an intestinal sample, or a mucosal sample.

[0231] Referring to block 220, in some embodiments, the method includes, for each corresponding subject in the second plurality of subjects, sequencing genomic DNA from a corresponding biological sample from the intestinal tract of the corresponding subject, thereby obtaining a corresponding second plurality (e.g., at least 100,000) of nucleic acid sequences. In some embodiments, the second plurality of nucleic acid sequences comprises at least 100,000, at least 250,000, at least 500,000, at least 1,000,000, at least 2,500,000, at least 5,000,000, at least 10,000,000, or at least 50,000,000 nucleic acid sequences. In some embodiments, the second plurality of nucleic acid sequences comprises no more than 250,000,000, no more than 100,000,000, no more than 50,000,000, no more than 25,000,000, no more than 10,000,000, no more than 5,000,000, no more than 1,000,000, or no more than 100,000 nucleic acid sequences. In some embodiments, the second plurality of nucleic acid sequences consists of 100,000 to 1,000,000, 200,000 to 5,000,000, 500,000 to 10,000,000, 1,000,000 to 20,000,000, 5,000,000 to 50,000,000, 10,000,000 to 100,000,000, or 50,000,000 to 250,000,000 nucleic acid sequences. In some embodiments, the second plurality of nucleic acid sequences falls within another range starting at no less than 100,000 nucleic acid sequences and ending at no more than 250,000,000 nucleic acid sequences.

[0232] In some embodiments, a second plurality (e.g., at least 100,000) of nucleic acid sequences are obtained by metagenomic sequencing, for example, as disclosed in U.S. Patent Application Publication No. 2016 / 0239602 or U.S. Patent No. 11,495,326, the contents of which are incorporated herein by reference in their entirety. In some embodiments, metagenomic sequencing further includes generating a plurality of metagenomic fragment reads. In some embodiments, metagenomic sequencing further includes fragmenting the microbial genome into random fragments of a target size. The size of the resulting fragments may vary. In one embodiment, a fragment of approximately 500 nucleotides may be obtained. In some embodiments, a fragment of 100-2000 nucleotides (e.g., 200-800, 100-900, 100-1000, 300-800, 400-900 nucleotides) may be obtained. In some embodiments, the method may further include extracting metagenomic fragments from corresponding biological samples. In some embodiments, metagenomic sequencing further includes sequencing the fragments using a high-throughput sequencing method to generate a plurality of sequencing reads.

[0233] In some embodiments, the second plurality (e.g., at least 100,000) of nucleic acid sequences are obtained by targeted panel sequencing, e.g., as described in U.S. Patent Application Publication No. 2019 / 0316209. In some embodiments, targeted panel sequencing comprises hybridizing genomic DNA isolated from a biological sample from the intestinal tract of a subject with a set of probes that are specific for each microorganism being quantified (e.g., Table 1, Table 2, and / or Table 3). Figures 42A to 42XX In some embodiments, the probe set comprises one or more probes that hybridize to unique sequences in the genome of each of the plurality of microorganisms listed in (e.g., each of the plurality of microorganisms listed in); the recovered nucleic acid is then sequenced. In some embodiments, a combination of semi-unique sequences (e.g., sequences found in a small number of microbial genomes) can be used to deconvolute genomic abundance values ​​using an algorithm (e.g., a system of equations). In some embodiments, the probe set comprises at least one probe that hybridizes to a unique sequence in each microbial genome to be detected. In some embodiments, the probe set comprises at least 2, at least 3, at least 4, at least 5, at least 10, at least 25, at least 50 or more probes that hybridize to unique different sequences in each microbial genome to be detected. In some embodiments, the probe sets include at least 20, at least 30, at least 40, at least 50, at least 75, at least 100, at least 125, at least 150, at least 200, at least 150, at least 300, at least 400, at least 500, at least 750, at least 1000, at least 1250, at least 1500, at least 2000, at least 2500, at least 3000, at least 4000, at least 5000, at least 7500, at least 10,000 or more unique probes.

[0234] In some embodiments, the sequenced genomic DNA from the corresponding biological sample may contain at its end a partial or complete sequencing platform adapter sequence that can be used for sequencing using a sequencing platform of interest. Sequencing platforms of interest include, but are not limited to, HiSeq TM , MiSeq TM and Genome Analyzer TM Sequencing system; from IonTorrent TM Ion PGM TM and Ion Proton TM Sequencing system: PACBIO RSII Sequel system from Pacific Biosciences, from Life Technologies TM SOLiD sequencing system from Roche, 454 GS FLX+ and GS Junior sequencing systems from Roche, MinION from Oxford Nanopore TM system or any other sequencing platform of interest.

[0235] Referring to block 222, in some embodiments, the method includes electronically obtaining, for each respective subject in the second plurality of subjects, a corresponding second plurality (e.g., at least 100,000) of nucleic acid sequences of genomic DNA from a corresponding biological sample from the intestine of the respective subject, and determining, for each respective subject in the second plurality of subjects, a corresponding genomic abundance value for each respective gut microbe in the plurality of gut microbes from the corresponding second plurality (e.g., at least 100,000) of nucleic acid sequences. In some embodiments, the genomic abundance values ​​determined for each respective subject in the second plurality of subjects include at least 20, at least 25, at least 50, at least 100, at least 250, at least 500, at least 1,000, at least 5,000, or at least 10,000 genomic abundance values. In some embodiments, the genomic abundance values ​​determined for each respective subject in the second plurality of subjects comprise no more than 250,000, no more than 100,000, no more than 50,000, no more than 25,000, no more than 10,000, no more than 5,000, no more than 1,000, no more than 100, no more than 50, no more than 30, or no more than 20 genomic abundance values. In some embodiments, the genomic abundance values ​​determined for each respective subject in the second plurality of subjects consist of 10 to 40, 20 to 50, 30 to 80, 40 to 100, 50 to 150, 60 to 200, 80 to 300, 90 to 500, 100 to 1000, 500 to 2,000, or 1,000 to 5,000 genomic abundance values. In some embodiments, the genomic abundance values ​​determined for each respective subject in the second plurality of subjects fall within another range starting at no less than 20 genomic abundance values ​​and ending at no more than 250,000 genomic abundance values.

[0236] Referring to box 224, in some embodiments, the method includes, for each corresponding subject in the second plurality of subjects, assembling a corresponding second plurality of intestinal microbial genomes from a corresponding second plurality (e.g., at least 100,000) nucleic acid sequences by metagenome de novo sequence assembly, and for each corresponding intestinal microbial genome in the corresponding second plurality of intestinal microbial genomes, calculating the corresponding genome abundance of the corresponding intestinal microbial genome. In some embodiments, the metagenome de novo sequence assembly further includes generating overlapping groups based on sequencing reads generated by shotgun sequencing technology, for example, as described in U.S. Patent No. 10,529,443. In some embodiments, the second plurality (e.g., at least 100,000) nucleic acid sequences can be assembled into complete genomes of a plurality of intestinal microorganisms. In some embodiments, the second plurality (e.g., at least 100,000) nucleic acid sequences can be assembled into partial genomes of a plurality of intestinal microorganisms.

[0237] Referring to box 226, in some embodiments, the method includes, for each corresponding subject in a second plurality of subjects, assigning each corresponding nucleic acid sequence in a corresponding second plurality (e.g., at least 100,000) of sequences to a corresponding intestinal microbe in a plurality of intestinal microbes, thereby generating a corresponding count of the corresponding nucleic acid sequence in the corresponding second plurality of nucleic acid sequences assigned to the corresponding intestinal microbe for each corresponding intestinal microbe in the plurality of intestinal microbes; and determining a corresponding genome abundance value for the corresponding intestinal microbe based on the corresponding count of the corresponding nucleic acid sequence assigned to the corresponding intestinal microbe for each corresponding intestinal microbe in the plurality of intestinal microbes. In some embodiments, assigning each corresponding nucleic acid to the corresponding intestinal microbe comprises mapping the nucleic acid to a reference nucleic acid, e.g., a contig listed in Figure 41. In some embodiments, assigning each corresponding nucleic acid to the corresponding intestinal microbe comprises annotating genomic information based on an existing database. In some embodiments, the nucleic acid sequences are analyzed, and the annotations are defined using sequence similarity and phylogenetic placement methods or a combination of the two strategies to define taxonomic assignments.

[0238] The method based on sequence similarity includes those methods familiar to those skilled in the art, includes but is not limited to BLAST, BLASTx, tBLASTn, tBLASTx, RDP sorter, DNAclust and the various embodiments (such as Qiime or Mothur) of these algorithms.These methods rely on sequence reading and mapping to reference database and selecting the matching with best score and e value.In certain embodiments, phylogenetic method is used in combination with sequence similarity method to improve the calling accuracy of annotation or classification distribution.Common databases include but are not limited to GT-DBTK, National Center for Biotechnology Information (NCBI) Genbank, European Bioinformatics Institute-European Nucleotide Archive (European Bioinformatics Institute-European Nucleotide Archive; EBI-ENA), U.S. Department of Energy National Institute of Genetics (USDOE) integrated microbial genome (Integrated Microbial Genomes) & Microbiomes; IMG / M) and other available databases in this area.

[0239] Referring to block 228, in some embodiments, the second state of the bio-signature is the presence of a disease or condition, such as type 2 diabetes (T2D), hypertension (HT), schizophrenia (SCZ), atherosclerotic cardiovascular disease (ACVD), cirrhosis (LC), inflammatory bowel disease (IBD), colorectal cancer (CRC), ankylosing spondylitis (AS), Parkinson's disease (PD), multiple sclerosis (MS), Gaucher disease type II (GDII), COVID-19 (COVID-19), Behçet's disease (BD), autism spectrum disorder (ASD), or pancreatic cancer (PC). In some embodiments, the disease or condition is cancer, Alzheimer's disease, cardiovascular disease, autoimmune disease, mental health disorder, infectious disease, or genetic disorder.

[0240] In some embodiments, the second state of the biological characteristic is the second severity of the disease or condition. In some embodiments, the severity of the disease is classified according to the type, frequency or intensity of the subject's experience. In some embodiments, the severity of the disease is classified according to the progression or prognosis (e.g., different stages of cancer) of the disease or condition. In some embodiments, the second state of the biological characteristic is a treated disease or condition, and the second state of the biological characteristic is a disease or condition treated with a second therapy (e.g., surgery, radiotherapy, chemotherapy, targeted therapy, gene therapy, immunotherapy, drug therapy, dietary changes, lifestyle changes). In some embodiments, the second state of the biological characteristic is the second level of nutrients (such as carbohydrates, protein, fat, vitamins, fiber) in the diet. In some embodiments, the second state of the biological characteristic is a second age. In some embodiments, a threshold value (e.g., the level of a biomarker, a diagnostic cutoff value, or a threshold nutrient intake level) is provided for determining the second state of the biological characteristic.

[0241] Referring to block 230, in some embodiments, the plurality of intestinal microorganisms comprises a member selected from Table 1, Table 2, or Figures 42A to 42XX In some embodiments, the plurality of intestinal microorganisms comprises at least 20 species of intestinal microorganisms selected from Table 1, Table 2 or Figures 42A to 42XX In some embodiments, the plurality of intestinal microorganisms comprises at least 25 intestinal microorganisms selected from Table 1, Table 2 or Figures 42A to 42XX In some embodiments, the plurality of intestinal microorganisms comprises at least 30 intestinal microorganisms selected from Table 1, Table 2 or Figures 42A to 42XX In some embodiments, the plurality of intestinal microorganisms comprises at least 40 intestinal microorganisms selected from Table 1, Table 2 or Figures 42A to 42XX In some embodiments, the plurality of intestinal microorganisms comprises at least 25 intestinal microorganisms selected from Table 1, Table 2 or Figures 42A to 42XXIn some embodiments, the plurality of intestinal microorganisms is selected from Table 1, Table 2, or Table 3. Figures 42A to 42XX In some embodiments, the plurality of intestinal microorganisms is a combination of at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 125, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, or all of the intestinal microorganisms listed in Table 1. In some embodiments, the plurality of intestinal microorganisms is a combination of at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 125, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, or all of the intestinal microorganisms listed in Table 2. In some embodiments, the plurality of intestinal microorganisms is a combination of at least 20, at least 25, at least 30, at least 400, at least 500, at least 600, at least 700, at least 800, or all of the intestinal microorganisms listed in Table 2. Figures 42A to 42XX All intestinal microorganisms listed in.

[0242] Referring to block 232, in some embodiments, the method includes calculating a first plurality of similarity metrics from a corresponding plurality of genomic abundance values ​​across a first plurality of subjects, wherein the first plurality of similarity metrics includes a first corresponding similarity metric for each unique gut microbe pair, and the first corresponding similarity metric quantifies the similarity between: (i) a corresponding first vector formed from corresponding genomic abundance values ​​for the first microbe in the unique gut microbe pairs across the first plurality of subjects; and (ii) a corresponding second vector formed from corresponding genomic abundance values ​​for the second microbe in the unique gut microbe pairs across the first plurality of subjects. In some embodiments, the corresponding first vector is a set of values, each value representing a genomic abundance value for the first microbe in one of the first plurality of subjects. In some embodiments, the corresponding second vector is a set of values, each value representing a genomic abundance value for the second microbe in one of the first plurality of subjects. In some embodiments, a unique genomic pair is formed by any two genomes detected in the first plurality of subjects. In some embodiments, the number of unique pairs of a total number of N genomes can be calculated as N(N-1) / 2, where N represents the number of non-repeated genomes detected in the first plurality of subjects. As the number of gut microbes in the group increases, the number of computations required to determine the set of all similarity measures increases as a second-order function of the number of microbes.

[0243] Referring to block 234, in some embodiments, the method includes calculating a second plurality of similarity metrics using the corresponding genomic abundance values ​​for the second plurality of subjects, wherein the second plurality of similarity metrics includes second corresponding similarity metrics for unique gut microbe pairs in the plurality of gut microbes, and the second corresponding similarity metrics quantify the similarity between: (i) a corresponding second vector formed from the corresponding genomic abundance values ​​for the first microbe in the gut microbes in the second plurality of subjects; and (ii) a corresponding second vector formed from the corresponding genomic abundance values ​​for the second microbe in the unique gut microbe pairs across the second plurality of subjects. In some embodiments, the corresponding second vector is a set of values, each value representing a genomic abundance value for the second microbe in one of the second plurality of subjects. In some embodiments, the corresponding second vector is a set of values, each value representing a genomic abundance value for the second microbe in one of the second plurality of subjects. In some embodiments, a unique genomic pair is formed by any two genomes detected in the second plurality of subjects. In some embodiments, the number of unique pairs of a total number of N genomes can be calculated as N(N-1) / 2, where N represents the number of non-repeated genomes detected in the first plurality of subjects. As the number of gut microbes in the group increases, the number of computations required to determine the set of all similarity measures increases as a second-order function of the number of microbes.

[0244] Referring to block 236, in some embodiments, the method includes determining a set of unique gut microbe pairs based on the first plurality of similarity metrics and the second plurality of similarity metrics, wherein for each corresponding unique gut microbe pair in the set of unique gut microbe pairs, both the first corresponding similarity metric and the second corresponding similarity metric indicate a statistically significant positive correlation between the abundance of the first gut microbe and the abundance of the second gut microbe in the corresponding unique gut microbe pair, or both the first corresponding similarity metric and the second corresponding similarity metric indicate a statistically significant negative correlation between the abundance of the first gut microbe and the abundance of the second gut microbe in the corresponding unique gut microbe pair. In some embodiments, the similarity metric is a correlation coefficient. In some embodiments, a threshold or cutoff value for a statistically significant positive correlation is defined. In some embodiments, a threshold or cutoff value for a statistically significant negative correlation is defined.

[0245] Referring to box 238, in some embodiments, one or both of the first corresponding similarity measure and the second similarity measure can be a Pearson correlation coefficient, an intraclass correlation coefficient, or a rank correlation coefficient. In some embodiments, the similarity measure is a Spearman's correlation coefficient or a maximum information coefficient (MIC). In some embodiments, the similarity measure is a Kendall tau rank correlation coefficient, also known as Kendall tau, which is used to measure the association between two measurements. In some embodiments, the similarity measure is calculated by any algorithm based on sparse correlations for component data (SparCC) or an algorithm based on sparse inverse covariance estimation for ecological association inference (SPIEC-EASI). In some embodiments, the algorithm based on SparCC is FastSpar.

[0246] Referring to block 240, in some embodiments, a statistically significant positive correlation has a P-value of less than 0.001. In some embodiments, a statistically significant positive correlation has a P-value of less than 0.05. In some embodiments, a statistically significant positive correlation has a P-value of less than 0.01. In some embodiments, a statistically significant positive correlation has a P-value of less than 0.001, less than 0.005, less than 0.01, less than 0.025, less than 0.05, or less than 0.075.

[0247] Referring to block 242, in some embodiments, the method includes identifying a set of gut microbes comprising corresponding gut microbes represented in the set of unique gut microbe pairs. In some embodiments, a microbiome network of the corresponding gut microbes represented in the set of unique gut microbe pairs is constructed. In some embodiments, the network is visualized using bioinformatics software (e.g., Cystoscape).

[0248] Referring to block 244 , in some embodiments, the method includes clustering the respective gut microbes represented in the unique gut microbe pair into one or more networks, each respective connected network including a corresponding plurality of nodes and a corresponding set of one or more edges.

[0249] Referring to block 246, in some embodiments, each corresponding node in the plurality of nodes represents a unique gut microbe represented in the set of unique gut microbe pairs. In some embodiments, the node size represents the average abundance of the genome. In some embodiments, the connecting rod between the nodes is considered to be a metal spring attached to the node pair. In some embodiments, a similarity metric is used to determine the repulsive and attractive forces of the spring. In some embodiments, the value of the similarity metric is used to determine the weight of the connecting rod.

[0250] Referring to block 248, in some embodiments, each corresponding edge in the corresponding set of one or more edges connects two nodes representing corresponding unique gut microbe pairs in the set of unique gut microbe pairs. In some embodiments, the positively correlated unique genome pairs are different from the negatively correlated unique genome pairs.

[0251] Referring to block 250 , in some embodiments, each respective node in the corresponding plurality of nodes is connected to at least one other respective node in the plurality of nodes via a respective edge in the corresponding set of one or more edges.

[0252] Referring to block 252, in some embodiments, the method includes identifying a corresponding network of the one or more networks that contains the most nodes, thereby identifying the set of gut microbes represented by the corresponding plurality of nodes in the corresponding network. In some embodiments, nodes that are not connected to the one or more networks are removed.

[0253] Referring to block 254, in some embodiments, the set of identified gut microbes includes all corresponding gut microbes represented in the set of unique gut microbe pairs. In some embodiments, the set of identified gut microbes includes all corresponding gut microbes represented by nodes of one or more microbiome networks.

[0254] Referring to block 256, in some embodiments, the panel of identified gut microbes comprises a member selected from Table 1, Table 2, or Figures 42A to 42XXIn some embodiments, the plurality of intestinal microorganisms comprises at least 20 species of intestinal microorganisms selected from Table 1, Table 2 or Figures 42A to 42XX In some embodiments, the plurality of intestinal microorganisms comprises at least 25 intestinal microorganisms selected from Table 1, Table 2 or Figures 42A to 42XX In some embodiments, the plurality of intestinal microorganisms comprises at least 30 intestinal microorganisms selected from Table 1, Table 2 or Figures 42A to 42XX In some embodiments, the plurality of intestinal microorganisms comprises at least 40 intestinal microorganisms selected from Table 1, Table 2 or Figures 42A to 42XX In some embodiments, the plurality of intestinal microorganisms comprises at least 25 intestinal microorganisms selected from Table 1, Table 2 or Figures 42A to 42XX In some embodiments, the plurality of intestinal microorganisms is selected from Table 1, Table 2, or Table 3. Figures 42A to 42XX In some embodiments, the plurality of intestinal microorganisms is a combination of at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 125, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, or all of the intestinal microorganisms listed in Table 1. In some embodiments, the plurality of intestinal microorganisms is a combination of at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 125, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, or all of the intestinal microorganisms listed in Table 2. In some embodiments, the plurality of intestinal microorganisms is a combination of at least 20, at least 25, at least 30, at least 400, at least 500, at least 600, at least 700, at least 800, or all of the intestinal microorganisms listed in Table 2. Figures 42A to 42XX All intestinal microorganisms listed in.

[0255] In some embodiments, the set of identified gut microbes is selected from those microbes listed in Table 5 having a connectivity of at least 2, at least 3, at least 4, at least 5, at least 10, at least 15, at least 20, at least 25, or more. Referring to block 308, in some embodiments, the plurality of gut microbes comprises at least 20 microbes selected from those microbes listed in Table 5 having a connectivity of at least 2. In some embodiments, the plurality of gut microbes comprises at least 20 microbes selected from those microbes listed in Table 5 having a connectivity of at least 3, at least 4, at least 5, at least 10, at least 15, at least 20, at least 25, or more.

[0256] 2. Methods for training models for assessing human health status

[0257] FIG3 is a diagram of a method for training a model for assessing human health as discussed below. The method 300 may be performed using a computer system (e.g., as described above with reference to FIG3 ). Figure 1 The computer system 100 shown and described above is implemented.

[0258] Referring to box 300, in some embodiments, the method includes, for each corresponding training subject in a plurality of training subjects, obtaining in electronic form: (i) a corresponding plurality of genome abundance values, for each corresponding gut microorganism in a plurality of gut microorganisms, the plurality of genome abundance values ​​including a corresponding value of the abundance of a genome of the corresponding gut microorganism in a corresponding biological sample from the gut of the corresponding training subject; and (ii) a corresponding state of a biological characteristic of the corresponding training subject.

[0259] In some embodiments, the plurality of training subjects comprises at least 50, at least 100, at least 200, at least 500, at least 1000, at least 5000, at least 10,000, at least 20,000, at least 50,000, at least 100,000, at least 500,000, or at least 1,000,000 subjects. In some embodiments, the plurality of subjects comprises no more than 1,000,000, no more than 500,000, no more than 100,000, no more than 50,000, no more than 20,000, no more than 10,000, no more than 1000 subjects, no more than 50 subjects, no more than 100 subjects, or no more than 50 subjects. In certain embodiments, a plurality of training subjects is by 50 to 100, 50 to 200, 50 to 500, 100 to 500, 200 to 500, 200 to 1000, 500 to 1000, 200 to 5,000, 1000 to 10,000, 5000 to 200,000, 10,000 to 50,000, 20,000 to 100,000 or 500,000 to 1,000,000 composition.In certain embodiments, a plurality of training subjects falls into and starts from being not less than 50 subjects and finally is not higher than another scope of 10,000,000 subjects.In certain embodiments, a plurality of training subjects has similar demographic characteristics (such as age, sex, race).In certain embodiments, a plurality of training subjects has similar physical characteristics (such as weight, height, BMI value).In certain embodiments, a plurality of training subjects has similar health status (such as physical or mental condition, medical history, gene carrier or drug use). In some embodiments, multiple subjects share or have similar behavioral and lifestyle preferences (such as diet, physical exercise, or substance use).

[0260] In certain embodiments, the corresponding value of the abundance of a genome is a value representing the absolute abundance of a microorganism genome. In certain embodiments, the corresponding value of the abundance of a genome is a value representing a normalized abundance value or a relative abundance value (for example, the abundance of a microorganism is normalized relative to the abundance of a total microbial group of interest). In certain embodiments, the corresponding value of the abundance of a genome is a value representing an average abundance value (for example, the mean value of the abundance obtained at different time points or from different biological samples from a patient, or the mean value of the abundance obtained using different probes, etc.), or a combination of any one of the above. The corresponding value of the abundance of a genome is measured by any technology known in the art. In certain embodiments, the genomic abundance value of a genome is measured by quantitative PCR (qPCR) such as bacterial 16S rRNA qPCR, RT-PCR or qRT-PCR, for quantitative genome abundance in a region of interest, for example, as described in U.S. Patent number 11,427,865, the disclosure of which is incorporated herein by reference in its entirety. In some embodiments, the genomic abundance value is measured as follows: using targeted sequencing (e.g., 16S rRNA sequencing or any other suitable biomarker), partial genome sequencing, or whole genome sequencing, thereby quantifying the number of reads of the targeted region in the microbial genome to determine the abundance of the genome, for example, as disclosed in U.S. Patent Application Publication No. 2021 / 0403986 or U.S. Patent No. 11,332,783, the disclosures of which are hereby incorporated by reference in their entirety. In some embodiments, deep sequencing is used to determine the abundance of the targeted sequence, for example, as disclosed in U.S. Patent Application Publication No. 2018 / 0237863, the disclosures of which are hereby incorporated by reference in their entirety. In some embodiments, the sequencing depth is at least about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 70, 80, 90, 100, 110, 120, 130, 150, 200, 300, 500, 500, 700, 1000 or more. In some embodiments, shotgun metagenomic sequencing is used to provide sequence reads of the genome in a sample, for example, as described in U.S. Patent No. 11,028,449, the contents of which are incorporated herein by reference in their entirety.

[0261] Referring to block 302, in some embodiments, the method includes, for each corresponding subject in the plurality of training subjects, sequencing genomic DNA from a corresponding biological sample from the intestinal tract of the corresponding training subject, thereby obtaining a corresponding plurality (e.g., at least 100,000) of nucleic acid sequences. In some embodiments, the plurality of nucleic acid sequences comprises at least 100,000, at least 250,000, at least 500,000, at least 1,000,000, at least 2,500,000, at least 5,000,000, at least 10,000,000, or at least 50,000,000 nucleic acid sequences. In some embodiments, the plurality of nucleic acid sequences comprises no more than 250,000,000, no more than 100,000,000, no more than 50,000,000, no more than 25,000,000, no more than 10,000,000, no more than 5,000,000, no more than 1,000,000, no more than 100,000 nucleic acid sequences. In some embodiments, the plurality of nucleic acid sequences consists of 100,000 to 1,000,000, 200,000 to 5,000,000, 500,000 to 10,000,000, 1,000,000 to 20,000,000, 5,000,000 to 50,000,000, 10,000,000 to 100,000,000, or 50,000,000 to 250,000,000 nucleic acid sequences. In some embodiments, the plurality of nucleic acid sequences falls within another range that begins with no less than 100,000 nucleic acid sequences and ends with no more than 250,000,000 nucleic acid sequences.

[0262] In some embodiments, multiple (e.g., at least 100,000) nucleic acid sequences are obtained by metagenomic sequencing, for example, as disclosed in U.S. Patent Application Publication No. 2016 / 0239602 or U.S. Patent No. 11,495,326, the contents of which are incorporated herein by reference in their entirety. In some embodiments, metagenomic sequencing further includes generating multiple metagenomic fragment reads. In some embodiments, metagenomic sequencing further includes fragmenting the microbial genome into random fragments of target size. The size of the resulting fragments may vary. In one embodiment, a fragment of approximately 500 nucleotides can be obtained. In some embodiments, a fragment of 100-2000 nucleotides (e.g., 200-800, 100-900, 100-1000, 300-800, 400-900 nucleotides) can be obtained. In some embodiments, the method may further include extracting metagenomic fragments from corresponding biological samples. In some embodiments, metagenomic sequencing further includes sequencing the fragments using a high-throughput sequencing method to generate multiple sequencing reads.

[0263] In some embodiments, a plurality (e.g., at least 100,000) of nucleic acid sequences are obtained by targeted panel sequencing, e.g., as described in U.S. Patent Application Publication No. 2019 / 0316209. In some embodiments, targeted panel sequencing comprises hybridizing genomic DNA isolated from a biological sample of the intestinal tract of a subject with a set of probes that are specific for each microorganism being quantified (e.g., Table 1, Table 2, and / or Table 3). Figures 42A to 42XX In some embodiments, the probe set comprises one or more probes that hybridize to unique sequences in the genome of each of the plurality of microorganisms listed in (e.g., each of the plurality of microorganisms listed in); the recovered nucleic acid is then sequenced. In some embodiments, a combination of semi-unique sequences (e.g., sequences found in a small number of microbial genomes) can be used to deconvolute genomic abundance values ​​using an algorithm (e.g., a system of equations). In some embodiments, the probe set comprises at least one probe that hybridizes to a unique sequence in each microbial genome to be detected. In some embodiments, the probe set comprises at least 2, at least 3, at least 4, at least 5, at least 10, at least 25, at least 50 or more probes that hybridize to unique different sequences in each microbial genome to be detected. In some embodiments, the probe sets include at least 20, at least 30, at least 40, at least 50, at least 75, at least 100, at least 125, at least 150, at least 200, at least 150, at least 300, at least 400, at least 500, at least 750, at least 1000, at least 1250, at least 1500, at least 2000, at least 2500, at least 3000, at least 4000, at least 5000, at least 7500, at least 10,000 or more unique probes.

[0264] In some embodiments, the sequenced genomic DNA from the corresponding biological sample comprises at its end a partial or complete sequencing platform adapter sequence that can be used for sequencing using a sequencing platform of interest. Sequencing platforms of interest include, but are not limited to, HiSeq TM , MiSeq TM and Genome Analyzer TM Sequencing system; from Ion Torrent TM Ion PGM TM and Ion Proton TM Sequencing system: PACBIO RS II Sequel system from Pacific Biosciences, from Life Technologies TM SOLiD sequencing system from Roche, 454 GS FLX+ and GS Junior sequencing systems from Roche, MinION from Oxford NanoporeTM system or any other sequencing platform of interest.

[0265] Referring to block 304, in some embodiments, the biological sample from the intestine of the corresponding subject is a stool sample from the corresponding training subject. In some embodiments, the sample is a tissue biopsy sample, an intestinal sample, or a mucosal sample. See, for example, Tang Q, J et al., Current Sampling Methods for Gut Microbiota: A Call for More Precise Devices, Front Cell Infect Microbiol., 10: 151 (2020), the contents of which are incorporated herein by reference in their entirety.

[0266] Referring to block 306, in some embodiments, the plurality of intestinal microorganisms comprises a member selected from Table 1, Table 2, or Figures 42A to 42XX In some embodiments, the plurality of intestinal microorganisms comprises at least 20 species of intestinal microorganisms selected from Table 1, Table 2 or Figures 42A to 42XX In some embodiments, the plurality of intestinal microorganisms comprises at least 25 intestinal microorganisms selected from Table 1, Table 2 or Figures 42A to 42XX In some embodiments, the plurality of intestinal microorganisms comprises at least 30 intestinal microorganisms selected from Table 1, Table 2 or Figures 42A to 42XX In some embodiments, the plurality of intestinal microorganisms comprises at least 40 intestinal microorganisms selected from Table 1, Table 2 or Figures 42A to 42XX In some embodiments, the plurality of intestinal microorganisms comprises at least 25 intestinal microorganisms selected from Table 1, Table 2 or Figures 42A to 42XX In some embodiments, the plurality of intestinal microorganisms is selected from Table 1, Table 2, or Table 3. Figures 42A to 42XXIn some embodiments, the plurality of intestinal microorganisms is a combination of at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 125, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, or all of the intestinal microorganisms listed in Table 1. In some embodiments, the plurality of intestinal microorganisms is a combination of at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 125, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, or all of the intestinal microorganisms listed in Table 2. In some embodiments, the plurality of intestinal microorganisms is a combination of at least 20, at least 25, at least 30, at least 400, at least 500, at least 600, at least 700, at least 800, or all of the intestinal microorganisms listed in Table 2. Figures 42A to 42XX All intestinal microorganisms listed in.

[0267] In some embodiments of the methods described herein, a genome identified in the metagenomic analysis is classified as corresponding to Table 1, Table 2, and / or Table 3 if the identified genomic construct has at least 97% sequence identity when compared to the microbial contigs provided in the sequence listing. Figures 42A to 42XX 41. In some embodiments, a genome identified in a metagenomic analysis is classified as corresponding to Table 1, Table 2, and / or Table 3 if the identified genomic construct has at least 98% sequence identity when compared to the microbial contigs provided in the sequence listing. Figures 42A to 42XX 41. In some embodiments, a genome identified in a metagenomic analysis is classified as corresponding to Table 1, Table 2, and / or Table 3 if the identified genomic construct has at least 99% sequence identity when compared to the microbial contigs provided in the sequence listing. Figures 42A to 42XX 41. In some embodiments, a genome identified in a metagenomic analysis is classified as corresponding to Table 1, Table 2, and / or Table 3 if the identified genomic construct has at least 99.5% sequence identity when compared to the microbial contigs provided in the sequence listing. Figures 42A to 42XX 41. In some embodiments, a genome identified in the metagenomic analysis is classified as corresponding to Table 1, Table 2, and / or Table 3 if the identified genomic construct has at least 97%, at least 97.5%, at least 98%, at least 98.5%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or more sequence identity when compared to a contig of a microorganism provided in the sequence listing. Figures 42A to 42XX The microorganisms listed in are shown in Figure 41.

[0268] In some embodiments, the plurality of intestinal microorganisms is selected from those microorganisms listed in Table 5 having a connectivity of at least 2, at least 3, at least 4, at least 5, at least 10, at least 15, at least 20, at least 25, or more. Referring to block 308, in some embodiments, the plurality of intestinal microorganisms includes at least 20 microorganisms selected from those microorganisms listed in Table 5 having a connectivity of at least 2. In some embodiments, the plurality of intestinal microorganisms includes at least 20 microorganisms selected from those microorganisms listed in Table 5 having a connectivity of at least 3, at least 4, at least 5, at least 10, at least 15, at least 20, at least 25, or more.

[0269] Referring to block 310, in some embodiments, the method includes, for each respective training subject in a plurality of training subjects, electronically obtaining a corresponding plurality (e.g., at least 100,000) of nucleic acid sequences of genomic DNA from a corresponding biological sample from the intestine of the respective training subject, and, for each respective intestinal microorganism in a plurality of intestinal microorganisms, determining a corresponding value for the abundance of the genome of the respective intestinal microorganism from the corresponding first plurality (e.g., at least 100,000) of nucleic acid sequences. In some embodiments, the genomic abundance values ​​determined for each respective subject in the plurality of training subjects include at least 20, at least 25, at least 50, at least 100, at least 250, at least 500, at least 1,000, at least 5,000, or at least 10,000 genomic abundance values. In some embodiments, the genomic abundance values ​​determined for each respective subject in the plurality of training subjects comprise no more than 250,000, no more than 100,000, no more than 50,000, no more than 25,000, no more than 10,000, no more than 5,000, no more than 1,000, no more than 100, no more than 50, no more than 30, or no more than 20 genomic abundance values. In some embodiments, the genomic abundance values ​​determined for each respective subject in the plurality of training subjects consist of 10 to 40, 20 to 50, 30 to 80, 40 to 100, 50 to 150, 60 to 200, 80 to 300, 90 to 500, 100 to 1000, 500 to 2,000, or 1,000 to 5,000 genomic abundance values. In some embodiments, the genome abundance values ​​determined for each respective subject in the plurality of training subjects fall within another range that begins at no less than 20 genome abundance values ​​and ends at no more than 250,000 genome abundance values.

[0270] Referring to box 312, in some embodiments, the method includes, for each corresponding training subject in a plurality of training subjects, assembling a corresponding plurality of intestinal microbial genomes in electronic form from a corresponding plurality of (e.g., at least 100,000) nucleic acid sequences by metagenome de novo assembly, and for each corresponding intestinal microbe in a plurality of intestinal microbes, calculating a corresponding value of the abundance of the genome of the corresponding intestinal microbe based on the prevalence of the corresponding nucleic acid sequence in a plurality of (e.g., at least 100,000) nucleic acid sequences, the corresponding nucleic acid sequence being used to assemble the corresponding intestinal microbe genome corresponding to the corresponding intestinal microbe in the plurality of intestinal microbe genomes. In some embodiments, the metagenome de novo sequence assembly further includes generating contigs based on sequencing reads generated by shotgun sequencing technology, as described in U.S. Patent No. 10,529,443, the contents of which are incorporated herein by reference in their entirety. In some embodiments, the first plurality (e.g., at least 100,000) nucleic acid sequences can be assembled into complete genomes of a plurality of intestinal microbes. In some embodiments, the first plurality (e.g., at least 100,000) nucleic acid sequences can be assembled into partial genomes of a plurality of intestinal microbes.

[0271] With reference to box 314, in some embodiments, the method includes, for each corresponding subject in a plurality of training subjects, assigning each corresponding nucleic acid sequence in a corresponding plurality (e.g., at least 100,000) sequences to a corresponding intestinal microbe in a plurality of intestinal microbes, thereby generating a corresponding count of the corresponding nucleic acid sequence in the corresponding plurality of nucleic acid sequences assigned to the corresponding intestinal microbe for each corresponding intestinal microbe in the plurality of intestinal microbes, and for each corresponding intestinal microbe in the plurality of intestinal microbes, determining the corresponding genome abundance value of the corresponding intestinal microbe based on the corresponding count of the corresponding nucleic acid sequence assigned to the corresponding intestinal microbe. In some embodiments, assigning each corresponding nucleic acid to the corresponding intestinal microbe includes mapping the nucleic acid to a reference nucleic acid, e.g., the contigs listed in Figure 41. In some embodiments, assigning each corresponding nucleic acid to the corresponding intestinal microbe includes annotating genomic information based on an existing database. In some embodiments, the nucleic acid sequences are analyzed, and the annotations are defined using sequence similarity and phylogenetic placement methods or a combination of the two strategies to define taxonomic assignments.

[0272] The method based on sequence similarity includes those methods familiar to those skilled in the art, includes but is not limited to BLAST, BLASTx, tBLASTn, tBLASTx, RDP sorter, DNAclust and the various embodiments (such as Qiime or Mothur) of these algorithms.These methods rely on sequence reading and mapping to reference database and selecting the matching with best score and e value.In certain embodiments, phylogenetic method is used in combination with sequence similarity method to improve the calling accuracy of annotation or classification distribution.Common databases include but are not limited to GT-DBTK, National Center for Biotechnology Information (NCBI) Genbank, European Bioinformatics Institute-European Nucleotide Archive (European Bioinformatics Institute-European Nucleotide Archive; EBI-ENA), U.S. Department of Energy National Institute of Genetics (USDOE) integrated microbial genome (Integrated Microbial Genomes) & Microbiomes; IMG / M) and other available databases in this area.

[0273] Referring to box 316, in some embodiments, the biological characteristic is a disease or condition, a therapy administered to the subject (e.g., surgery, radiation therapy, chemotherapy, targeted therapy, gene therapy, immunotherapy, medication use, dietary changes, lifestyle changes), or the subject's diet (such as a diet rich in or poor in carbohydrates, protein, fat, vitamins, or fiber).

[0274] Referring to box 318, in some embodiments, the disease or condition is selected from the group consisting of type 2 diabetes (T2D), hypertension (HT), schizophrenia (SCZ), atherosclerotic cardiovascular disease (ACVD), cirrhosis (LC), inflammatory bowel disease (IBD), colorectal cancer (CRC), ankylosing spondylitis (AS), and Parkinson's disease (PD), multiple sclerosis (MS), Gaucher disease type II (GDII), COVID-19 (COV), Behçet's disease (BD), autism spectrum disorder (ASD), or pancreatic cancer (PC). In some embodiments, the disease or condition is cancer, Alzheimer's disease, cardiovascular disease, autoimmune disease, mental health disease, infectious disease, or genetic disorder.

[0275] In some embodiments, the model is trained on a dataset collected across multiple conditions, and the model is trained to distinguish between healthy and unhealthy states. For example, as described in Example 5, a random forest classifier was trained on a dataset from 26 different studies that collectively studied the microbiome in 15 different conditions. As shown in FIG38 , the generated model is able to predict healthy or unhealthy condition states regardless of the condition. Thus, in some embodiments, the biological feature is any one of a plurality of diseases and / or conditions, wherein the first state is the presence of any one of the diseases or conditions, and the second state is the absence of any one of the diseases or conditions.

[0276] Referring to block 320, in some embodiments, the disease or condition is cancer.

[0277] Referring to block 322, in some embodiments, the method includes, for each respective training subject in a plurality of training subjects, inputting information about the respective training subject into a model comprising a plurality of parameters. The model applies the plurality of parameters to the information through at least 10,000 calculations to obtain a corresponding output from the model for the respective training subject. The corresponding output includes an indication of a corresponding state of a biological characteristic of the respective training subject. The information about the respective training subject includes a corresponding genomic abundance value for each respective intestinal microorganism in a plurality of intestinal microorganisms, and the plurality of intestinal microorganisms are selected from Table 1, Table 2, or Figures 42A to 42XX .

[0278] Referring to box 322, in some embodiments, the indication of the corresponding state of the biological feature is a classification output of the corresponding state among a plurality of possible states of the biological feature. In some embodiments, the possible state is a state from a healthy subject. In some embodiments, the possible state is a state from a patient. In some embodiments, the state from the patient is classified according to the type, frequency, or intensity of the patient's experience. In some embodiments, the state from the patient is classified according to the progression or prognosis of the disease or condition (e.g., different stages of cancer). In some embodiments, a threshold value (such as the level of a biomarker, a diagnostic cutoff value, or a threshold nutrient intake level) is provided for determining the state of a healthy subject or patient.

[0279] Referring to block 324, in some embodiments, the indication of the corresponding state of the biological feature is a probabilistic output of the corresponding state of the biological feature. In some embodiments, the corresponding state is a state from a healthy subject. In some embodiments, the corresponding state is a state from a patient. In some embodiments, the states from the patient are classified by the type, frequency, or intensity of the patient's experience. In some embodiments, the states from the patient are classified by the progression or prognosis of the disease or condition (e.g., different stages of cancer). In some embodiments, a threshold value (such as a biomarker level, a diagnostic cutoff value, or a threshold nutrient intake level) is provided for determining the state of a healthy subject or patient.

[0280] Referring to box 326, in some embodiments, the model is a neural network algorithm, a support vector machine algorithm, a naive Bayes algorithm, a nearest neighbor algorithm, a boosted tree algorithm, a random forest algorithm, a convolutional neural network algorithm, a decision tree algorithm, a regression algorithm, or a clustering algorithm.

[0281] Referring to box 328, in some embodiments, the plurality of parameters is at least 1000, at least 10,000, at least 15,000, at least 50,000, at least 100,000, at least 250,000, at least 500,000, at least 1,000,000 parameters, at least 2,500,000 parameters, at least 5,000,000 parameters, at least 10,000,000 parameters, or more parameters.

[0282] Referring to box 330, in some embodiments, the model applies multiple parameters to the information through at least 1000 calculations, at least 5000 calculations, at least 10,000 calculations, at least 25,000 calculations, at least 50,000 calculations, at least 100,000 calculations, at least 250,000 calculations, at least 500,000 calculations, at least 1,000,000 calculations, at least 2,500,000 calculations, at least 5,000,000 calculations, at least 10,000,000 calculations, or more calculations to obtain corresponding outputs from the model for corresponding training subjects.

[0283] Referring to box 330, in some embodiments, the method includes adjusting a plurality of parameters based on one or more differences between (i) a corresponding output from the model and (ii) a corresponding state of the biological characteristic of the corresponding training subject for each corresponding training subject in the first plurality of training subjects.

[0284] In some embodiments where deep learning techniques utilize a neural network as described above, training the neural network to improve the accuracy of its predictions involves modifying one or more parameters, including but not limited to weights in filters in convolutional layers and biases in network layers. In some embodiments, the weights and biases are further constrained by various forms of regularization (such as L1, L2, weight decay, and dropout).

[0285] For example, in some embodiments, the neural network or any model disclosed herein optionally has its parameters (e.g., weights) adjusted (adjusted to potentially minimize the error between the system's predicted indications and the measured indications of the training data) when the training data is labeled (e.g., labeled with an indication of the state of a biological feature). Various methods for minimizing the error function, such as gradient descent, include but are not limited to logarithmic loss, sum of squared errors, hinge loss methods. In some embodiments, these methods also include second-order methods or approximations such as momentum, Hessian-free estimation, Nesterov accelerated gradient, adagrad, etc. In some embodiments, these methods also combine unlabeled generative pre-training with labeled discriminative training.

[0286] Therefore, in some embodiments, the training of the neural network includes adjusting one or more parameters of a plurality of parameters by back propagation of a loss function. In some embodiments, the loss function is a regression task and / or a classification task. Non-limiting examples of loss functions suitable for regression tasks include, but are not limited to, mean squared error loss function, mean absolute error loss function, Huber loss function, Log-Cosh loss function, or quantile loss function. See, Wang et al., 2020, “A Comprehensive Survey of Loss Functions in Machine Learning,” Annals of DataScience, doi.org / 10.1007 / s40745-020-00253-5, last accessed on September 15, 2021, which is hereby incorporated by reference in its entirety. Non-limiting examples of loss functions suitable for classification tasks include, but are not limited to, binary cross entropy loss function, hinge loss function, or squared hinge loss function. In some embodiments, the loss function is any suitable regression task loss function or classification task loss function.

[0287] Other suitable methods for training neural networks contemplated for use with the present disclosure are further described herein (see, e.g., definition of: untrained model, above).

[0288] In some embodiments, the parameters of the neural network are randomly initialized prior to training.

[0289] In some embodiments, the neural network includes discarding regularization parameters. For example, in some embodiments, regularization is performed by adding a penalty to the loss function, where the penalty is proportional to the parameter value in the trained or untrained model. Generally speaking, regularization reduces the complexity of the model by adding a penalty to one or more parameters to reduce the importance of the corresponding hidden neurons associated with those parameters. This practice can produce more generalized models and reduce overfitting of the data. In some embodiments, regularization includes an L1 or L2 penalty.

[0290] In some embodiments, training a neural network includes an optimizer. In some embodiments, the optimizer can use a loss function to update parameters of a neural network or other model via back propagation. In some embodiments, training a neural network includes a learning rate.

[0291] In some embodiments, the learning rate is at least 0.0001, at least 0.0005, at least 0.001, at least 0.005, at least 0.01, at least 0.05, at least 0.1, at least 0.2, at least 0.3, at least 0.4, at least 0.5, at least 0.6, at least 0.7, at least 0.8, at least 0.9, or at least 1. In some embodiments, the learning rate is no more than 1, no more than 0.9, no more than 0.8, no more than 0.7, no more than 0.6, no more than 0.5, no more than 0.4, no more than 0.3, no more than 0.2, no more than 0.1, no more than 0.05, no more than 0.01, or less. In some embodiments, the learning rate is 0.0001 to 0.01, 0.001 to 0.5, 0.001 to 0.01, 0.005 to 0.8, or 0.005 to 1. In some embodiments, the learning rate falls within another range that starts at no less than 0.0001 and ends at no more than 1.

[0292] In some embodiments, the learning rate further comprises a learning rate decay (e.g., a reduction in the learning rate over one or more epochs). For example, the learning rate decay can be a learning rate reduction of 0.5 or 0.1. In some embodiments, the learning rate is a differential learning rate. In some embodiments, the training of the neural network further utilizes a scheduler that conditionally applies a learning rate decay based on an evaluation of a performance metric over a threshold number of training epochs (e.g., applying a learning rate decay when a performance metric fails to meet a threshold performance value after at least a threshold number of training epochs).

[0293] In some embodiments, the performance of the neural network is measured at one or more time points using a performance metric, including but not limited to a training loss metric, a validation loss metric, and / or a mean absolute error. In some embodiments, the performance metric is the area under the receiver operating characteristic (AUROC) and / or the area under the precision-recall curve (AUPRC).

[0294] For example, in some embodiments, the performance of the neural network is measured by validating the model using a validation (e.g., development) dataset. In some such embodiments, the neural network is trained to form a trained neural network when the neural network meets minimum performance requirements based on the validation.

[0295] In some embodiments, any suitable validation method may be used, including but not limited to K-fold cross-validation, advanced cross-validation, random cross-validation, grouped cross-validation (e.g., K-fold grouped cross-validation), bootstrap bias-corrected cross-validation, random search, and / or Bayesian hyperparameter optimization.

[0296] In some embodiments, a method is provided for training a model comprising multiple parameters by the following procedure, the procedure comprising (i) inputting, for each respective training subject in a plurality of training subjects, a corresponding genomic abundance value for each respective gut microbe in a plurality of gut microbes, thereby obtaining, for each respective training subject in the plurality of training subjects, a corresponding predicted state of a biological feature as an output of the model, and (ii) for each respective training subject in the plurality of training subjects, refining multiple model parameters based on a difference between the corresponding state of the biological feature of the respective training subject and the corresponding predicted state of the biological feature.

[0297] 3. Methods for assessing the health status of subjects

[0298] FIG4 is a schematic diagram of a method for training a model for assessing human health as discussed below. The method may be performed using a computer system (e.g., as described above with reference to FIG4). Figure 1 The present invention may be implemented by the computer system 100 shown and described herein.

[0299] Referring to block 400, in some embodiments, a plurality of genomic abundance values ​​are obtained electronically for a genome selected from Table 1, Table 2, or Figures 42A to 42XX For each respective intestinal microorganism in a plurality (e.g., at least 20) intestinal microorganisms of the subject, the genomic abundance value comprises the corresponding abundance value of the genome of the corresponding intestinal bacterial species in the plurality (e.g., at least 20) intestinal microorganisms in the biological sample from the subject. In some embodiments, the plurality of intestinal microorganisms comprises a selected from Table 1, Table 2, or Figures 42A to 42XX In some embodiments, the plurality of intestinal microorganisms comprises at least 25 intestinal microorganisms selected from Table 1, Table 2 or Figures 42A to 42XX In some embodiments, the plurality of intestinal microorganisms comprises at least 30 intestinal microorganisms selected from Table 1, Table 2 or Figures 42A to 42XX In some embodiments, the plurality of intestinal microorganisms comprises at least 40 intestinal microorganisms selected from Table 1, Table 2 or Figures 42A to 42XXIn some embodiments, the plurality of intestinal microorganisms comprises at least 25 intestinal microorganisms selected from Table 1, Table 2 or Figures 42A to 42XX In some embodiments, the plurality of intestinal microorganisms is selected from Table 1, Table 2, or Table 3. Figures 42A to 42XX In some embodiments, the plurality of intestinal microorganisms is a combination of at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 125, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, or all of the intestinal microorganisms listed in Table 1. In some embodiments, the plurality of intestinal microorganisms is a combination of at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 125, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, or all of the intestinal microorganisms listed in Table 2. In some embodiments, the plurality of intestinal microorganisms is a combination of at least 20, at least 25, at least 30, at least 400, at least 500, at least 600, at least 700, at least 800, or all of the intestinal microorganisms listed in Table 2. Figures 42A to 42XX All intestinal microorganisms listed in.

[0300] In certain embodiments, the corresponding value of the abundance of a genome is a value representing the absolute abundance of a microorganism genome. In certain embodiments, the corresponding value of the abundance of a genome is a value representing a normalized abundance value or a relative abundance value (for example, the abundance of a microorganism is normalized relative to the abundance of a total microbial group of interest). In certain embodiments, the corresponding value of the abundance of a genome is a value representing an average abundance value (for example, the mean value of the abundance obtained at different time points or from different biological samples from a patient, or the mean value of the abundance obtained using different probes, etc.), or a combination of any one of the above. The corresponding value of the abundance of a genome is measured by any technology known in the art. In certain embodiments, the genomic abundance value of a genome is measured by quantitative PCR (qPCR) such as bacterial 16S rRNA qPCR, RT-PCR or qRT-PCR, for quantitative genome abundance in a region of interest, for example, as described in U.S. Patent number 11,427,865, the disclosure of which is incorporated herein by reference in its entirety. In some embodiments, the genomic abundance value is measured as follows: using targeted sequencing (e.g., 16S rRNA sequencing or any other suitable biomarker), partial genome sequencing, or whole genome sequencing, thereby quantifying the number of reads of the targeted region in the microbial genome to determine the abundance of the genome, for example, as disclosed in U.S. Patent Application Publication No. 2021 / 0403986 or U.S. Patent No. 11,332,783, the disclosures of which are hereby incorporated by reference in their entirety. In some embodiments, deep sequencing is used to determine the abundance of the targeted sequence, for example, as disclosed in U.S. Patent Application Publication No. 2018 / 0237863, the disclosures of which are hereby incorporated by reference in their entirety. In some embodiments, the sequencing depth is at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 150, at least 200, at least 300, at least 400, at least 500, at least 750, at least 1000 or more.In some embodiments, shotgun metagenomic sequencing is used to provide sequence reads of the genome in a sample, for example, as described in U.S. Patent No. 11,028,449, the contents of which are incorporated herein by reference in their entirety.

[0301] In some embodiments of the methods described herein, a genome identified in the metagenomic analysis is classified as corresponding to Table 1, Table 2, and / or Table 3 if the identified genomic construct has at least 97% sequence identity when compared to the microbial contigs provided in the sequence listing. Figures 42A to 42XX 41. In some embodiments, a genome identified in a metagenomic analysis is classified as corresponding to Table 1, Table 2, and / or Table 3 if the identified genomic construct has at least 98% sequence identity when compared to the microbial contigs provided in the sequence listing. Figures 42A to 42XX 41. In some embodiments, a genome identified in a metagenomic analysis is classified as corresponding to Table 1, Table 2, and / or Table 3 if the identified genomic construct has at least 99% sequence identity when compared to the microbial contigs provided in the sequence listing. Figures 42A to 42XX 41. In some embodiments, a genome identified in a metagenomic analysis is classified as corresponding to Table 1, Table 2, and / or Table 3 if the identified genomic construct has at least 99.5% sequence identity when compared to the microbial contigs provided in the sequence listing. Figures 42A to 42XX 41. In some embodiments, a genome identified in the metagenomic analysis is classified as corresponding to Table 1, Table 2, and / or Table 3 if the identified genomic construct has at least 97%, at least 97.5%, at least 98%, at least 98.5%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or more sequence identity when compared to a contig of a microorganism provided in the sequence listing. Figures 42A to 42XX The microorganisms listed in are shown in Figure 41.

[0302] Referring to block 402, in some embodiments, the method includes sequencing genomic DNA from a biological sample from the intestine of a subject to obtain a plurality (e.g., at least 100,000) of nucleic acid sequences. In some embodiments, the plurality of nucleic acid sequences comprises at least 100,000, at least 250,000, at least 500,000, at least 1,000,000, at least 2,500,000, at least 5,000,000, at least 10,000,000, or at least 50,000,000 nucleic acid sequences. In some embodiments, the plurality of nucleic acid sequences comprises no more than 250,000,000, no more than 100,000,000, no more than 50,000,000, no more than 25,000,000, no more than 10,000,000, no more than 5,000,000, no more than 1,000,000, or no more than 100,000 nucleic acid sequences. In some embodiments, the plurality of nucleic acid sequences consists of 100,000 to 1,000,000, 200,000 to 5,000,000, 500,000 to 10,000,000, 1,000,000 to 20,000,000, 5,000,000 to 50,000,000, 10,000,000 to 100,000,000, or 50,000,000 to 250,000,000 nucleic acid sequences. In some embodiments, the plurality of nucleic acid sequences falls within another range that begins with no less than 100,000 nucleic acid sequences and ends with no more than 250,000,000 nucleic acid sequences.

[0303] In some embodiments, multiple (e.g., at least 100,000) nucleic acid sequences are obtained by metagenomic sequencing, for example, as disclosed in U.S. Patent Application Publication No. 2016 / 0239602 or U.S. Patent No. 11,495,326, the contents of which are incorporated herein by reference in their entirety. In some embodiments, metagenomic sequencing further includes generating multiple metagenomic fragment reads. In some embodiments, metagenomic sequencing further includes fragmenting the microbial genome into random fragments of target size. The size of the resulting fragments may vary. In one embodiment, a fragment of approximately 500 nucleotides can be obtained. In some embodiments, a fragment of 100-2000 nucleotides (e.g., 200-800, 100-900, 100-1000, 300-800, 400-900 nucleotides) can be obtained. In some embodiments, the method may further include extracting metagenomic fragments from corresponding biological samples. In some embodiments, metagenomic sequencing further includes sequencing the fragments using a high-throughput sequencing method to generate multiple sequencing reads.

[0304] In some embodiments, the first plurality (e.g., at least 100,000) of nucleic acid sequences are obtained by targeted panel sequencing, e.g., as described in U.S. Patent Application Publication No. 2019 / 0316209. In some embodiments, targeted panel sequencing comprises hybridizing genomic DNA isolated from a biological sample from the intestinal tract of a subject with a set of probes that are specific for each microorganism being quantified (e.g., Table 1, Table 2, and / or Table 3). Figures 42A to 42XX In some embodiments, the probe set comprises one or more probes that hybridize to unique sequences in the genome of each of the plurality of microorganisms listed in (e.g., each of the plurality of microorganisms listed in); the recovered nucleic acid is then sequenced. In some embodiments, a combination of semi-unique sequences (e.g., sequences found in a small number of microbial genomes) can be used to deconvolute genomic abundance values ​​using an algorithm (e.g., a system of equations). In some embodiments, the probe set comprises at least one probe that hybridizes to a unique sequence in each microbial genome to be detected. In some embodiments, the probe set comprises at least 2, at least 3, at least 4, at least 5, at least 10, at least 25, at least 50 or more probes that hybridize to unique different sequences in each microbial genome to be detected. In some embodiments, the probe sets include at least 20, at least 30, at least 40, at least 50, at least 75, at least 100, at least 125, at least 150, at least 200, at least 150, at least 300, at least 400, at least 500, at least 750, at least 1000, at least 1250, at least 1500, at least 2000, at least 2500, at least 3000, at least 4000, at least 5000, at least 7500, at least 10,000 or more unique probes.

[0305] In some embodiments, the sequenced genomic DNA from the corresponding biological sample comprises at its end a partial or complete sequencing platform adapter sequence that can be used for sequencing using a sequencing platform of interest. Sequencing platforms of interest include, but are not limited to, HiSeq TM , MiSeq TM and Genome Analyzer TM Sequencing system; from Ion Torrent TM Ion PGM TM and Ion Proton TM Sequencing system: PACBIO RS II Sequel system from Pacific Biosciences, from Life Technologies TM SOLiD sequencing system from Roche, 454 GS FLX+ and GS Junior sequencing systems from Roche, MinION from Oxford NanoporeTM system or any other sequencing platform of interest.

[0306] Referring to block 404, in some embodiments, the biological sample from the intestine of the respective subject is a stool sample. In some embodiments, the sample is a tissue biopsy sample, an intestinal sample, or a mucosal sample. See, e.g., Tang Q, J et al., Current Sampling Methods for Gut Microbiota: A Call for More Precise Devices, Front Cell Infect Microbiol., 10: 151 (2020), the contents of which are incorporated herein by reference in their entirety.

[0307] Referring to box 406, in some embodiments, the plurality of intestinal microorganisms is selected from those microorganisms listed in Table 5 having a connectivity of at least 2, at least 3, at least 4, at least 5, at least 10, at least 15, at least 20, at least 25, or more. Referring to box 308, in some embodiments, the plurality of intestinal microorganisms includes at least 20 microorganisms selected from those microorganisms listed in Table 5 having a connectivity of at least 2. In some embodiments, the plurality of intestinal microorganisms includes at least 20 microorganisms selected from those microorganisms listed in Table 5 having a connectivity of at least 3, at least 4, at least 5, at least 10, at least 15, at least 20, at least 25, or more.

[0308] Referring to block 408, in some embodiments, the method includes electronically obtaining a plurality (e.g., at least 100,000) of nucleic acid sequences of genomic DNA from a biological sample from the intestine of the subject; and for each corresponding intestinal microorganism in the plurality of intestinal microorganisms, determining a corresponding value for the abundance of the genome of the corresponding intestinal microorganism from the plurality of at least 100,000 nucleic acid sequences. In some embodiments, the genomic abundance values ​​determined for the subject include no more than 250,000, no more than 100,000, no more than 50,000, no more than 25,000, no more than 10,000, no more than 5,000, no more than 1,000, no more than 100, no more than 50, no more than 30, or no more than 20 genomic abundance values. In some embodiments, the genomic abundance values ​​determined for a subject consist of 10 to 40, 20 to 50, 30 to 80, 40 to 100, 50 to 150, 60 to 200, 80 to 300, 90 to 500, 100 to 1000, 500 to 2,000, or 1,000 to 5,000 genomic abundance values. In some embodiments, the genomic abundance values ​​determined for a subject fall within another range starting at no less than 20 genomic abundance values ​​and ending at no more than 250,000 genomic abundance values.

[0309] With reference to box 410, in some embodiments, the method includes assembling corresponding multiple intestinal microorganism genomes in electronic form by a plurality of (e.g., at least 100,000) nucleic acid sequences from the beginning by metagenome assembly, and for each corresponding intestinal microorganism in the multiple intestinal microorganisms, a corresponding value of the abundance of the genome of the corresponding intestinal microorganism is calculated based on the prevalence of the corresponding nucleic acid sequence in the plurality of (e.g., at least 100,000) nucleic acid sequences, and the corresponding nucleic acid sequence is used to assemble the corresponding intestinal microorganism genome corresponding to the corresponding intestinal microorganism in the multiple intestinal microorganism genomes. In some embodiments, the metagenome de novo sequence assembly further includes generating contigs based on sequencing reads generated by shotgun sequencing technology, as described in U.S. Patent No. 10,529,443, the content of which is incorporated herein by reference in its entirety. In some embodiments, a plurality of (e.g., at least 100,000) nucleic acid sequences can be assembled into the complete genomes of the multiple intestinal microorganisms. In some embodiments, a plurality of (e.g., at least 100,000) nucleic acid sequences can be assembled into partial genomes of the multiple intestinal microorganisms.

[0310] With reference to box 412, in some embodiments, these methods include assigning each corresponding nucleic acid sequence in a plurality of (e.g., at least 100,000) sequences to a corresponding intestinal microorganism in a plurality of intestinal microorganisms, thereby for each corresponding intestinal microorganism in a plurality of intestinal microorganisms, generating a corresponding count of the corresponding nucleic acid sequence in a plurality of nucleic acid sequences assigned to the corresponding intestinal microorganism, and for each corresponding intestinal microorganism in a plurality of intestinal microorganisms, determining the corresponding genome abundance value of the corresponding intestinal microorganism based on the corresponding count of the corresponding nucleic acid sequence assigned to the corresponding intestinal microorganism. In some embodiments, assigning each corresponding nucleic acid to the corresponding intestinal microorganism includes mapping the nucleic acid to a reference nucleic acid. In some embodiments, assigning each corresponding nucleic acid to the corresponding intestinal microorganism includes annotating genome information based on an existing database. In some embodiments, the nucleic acid sequence is analyzed, and annotation is to define classification assignment using sequence similarity and phylogeny placement methods or a combination of the two strategies.

[0311] The method based on sequence similarity includes those methods familiar to those skilled in the art, includes but is not limited to BLAST, BLASTx, tBLASTn, tBLASTx, RDP sorter, DNAclust and the various embodiments (such as Qiime or Mothur) of these algorithms.These methods rely on sequence reading and mapping to reference database and selecting the matching with best score and e value.In certain embodiments, phylogenetic method is used in combination with sequence similarity method to improve the calling accuracy of annotation or classification distribution.Common databases include but are not limited to GT-DBTK, National Center for Biotechnology Information (NCBI) Genbank, European Bioinformatics Institute-European Nucleotide Archive (European Bioinformatics Institute-European Nucleotide Archive; EBI-ENA), U.S. Department of Energy National Institute of Genetics (USDOE) integrated microbial genome (Integrated Microbial Genomes) & Microbiomes; IMG / M) and other available databases in this area.

[0312] Referring to box 414, in some embodiments, the method includes inputting the multiple genomic abundance values ​​into a model including multiple parameters, wherein the model applies the multiple parameters to the multiple genomic abundance values ​​through multiple calculations (e.g., at least 10,000 times) to generate an indication of the health status of the subject as an output from the model.

[0313] Referring to box 416, in some embodiments, the indication of the subject's health status is indicative of a biological characteristic, where the biological characteristic is a disease or condition, a therapy administered to the subject (e.g., surgery, radiation therapy, chemotherapy, targeted therapy, gene therapy, immunotherapy, medication use, dietary changes, lifestyle changes), or the subject's diet (such as a diet rich in or poor in carbohydrates, protein, fat, vitamins, fiber).

[0314] Referring to box 418, in some embodiments, the disease or condition is selected from the group consisting of type 2 diabetes (T2D), hypertension (HT), schizophrenia (SCZ), atherosclerotic cardiovascular disease (ACVD), cirrhosis (LC), inflammatory bowel disease (IBD), colorectal cancer (CRC), ankylosing spondylitis (AS), and Parkinson's disease (PD), multiple sclerosis (MS), Gaucher disease type II (GDII), COVID-19 (COV), Behçet's disease (BD), autism spectrum disorder (ASD), or pancreatic cancer (PC). In some embodiments, the disease or condition is cancer, Alzheimer's disease, cardiovascular disease, autoimmune disease, mental health disease, infectious disease, or genetic disorder.

[0315] Referring to block 420, in some embodiments, the disease or condition is cancer.

[0316] In some embodiments, the model has been trained on a dataset collected across a variety of conditions, and the model is trained to distinguish between healthy and unhealthy states. For example, as described in Example 5, a random forest classifier was trained on a dataset from 26 different studies that collectively studied the microbiome in 15 different conditions. As shown in FIG38 , the resulting model is able to predict healthy or unhealthy condition states regardless of the condition. Thus, in some embodiments, the biological feature is any one of a variety of diseases and / or conditions, wherein the first state is the presence of any one of the diseases or conditions, and the second state is the absence of any one of the diseases or conditions.

[0317] With reference to box 422, in some embodiments, the indication of the subject's health status is a category output of a corresponding state in a plurality of possible states of the subject's health status. In some embodiments, the corresponding state of the subject's health status is with reference to the severity of the disease or illness. In some embodiments, the severity of the disease is classified according to the progression or prognosis (e.g., different stages of cancer) of the disease or illness. In some embodiments, a threshold value (such as the level of a biomarker, a diagnostic cutoff value, or a threshold nutrient intake level) is provided for determining the state of the subject's health status. In some embodiments, the corresponding state of the subject's health status is the absence or presence of a disease or illness.

[0318] With reference to box 424, in some embodiments, the indication of the subject's health status is a probability output of a corresponding state of the subject's health status. In some embodiments, the corresponding state of the subject's health status is with reference to the severity of the disease or illness. In some embodiments, the severity of the disease is classified according to the progression or prognosis of the disease or illness (e.g., different stages of cancer). In some embodiments, a threshold value (such as the level of a biomarker, a diagnostic cutoff value, or a threshold nutrient intake level) is provided for determining the subject's state. In some embodiments, the corresponding state of the subject's health status is the absence or presence of a disease or illness.

[0319] Referring to box 426, in some embodiments, the model is a neural network algorithm, a support vector machine algorithm, a naive Bayes algorithm, a nearest neighbor algorithm, a boosted tree algorithm, a random forest algorithm, a convolutional neural network algorithm, a decision tree algorithm, a regression algorithm, or a clustering algorithm.

[0320] Referring to box 428, in some embodiments, the plurality of parameters is at least 1000, at least 10,000, at least 15,000, at least 50,000, at least 100,000, at least 250,000, at least 500,000, at least 1,000,000 parameters, at least 2,500,000 parameters, at least 5,000,000 parameters, at least 10,000,000 parameters, or more parameters.

[0321] Referring to box 430, in some embodiments, the model applies multiple parameters to the information through at least 1000 calculations, at least 5000 calculations, at least 10,000 calculations, at least 25,000 calculations, at least 50,000 calculations, at least 100,000 calculations, at least 250,000 calculations, at least 500,000 calculations, at least 1,000,000 calculations, at least 2,500,000 calculations, at least 5,000,000 calculations, at least 10,000,000 calculations, or more calculations to obtain corresponding outputs from the model for corresponding training subjects.

[0322] Examples

[0323] Example 1: Seesaw-like microbiome networks as a common microbiome signature of human disease

[0324] It is hypothesized that microorganisms required to provide essential health-related functions to the host [7] should maintain stable ecological interactions with each other to achieve structural and functional stability [18,19]. To identify microbiome signatures based on stable interactions between MAGs, patients with T2DM were randomized at baseline (M0) to receive a 3-month (M3) high-fiber intervention (W group; n = 74) or standard care (U group; n = 36) and then followed up for one year (M15) in an open-label controlled trial ( Figure 3A and Figure 7). A high-fiber intervention was used to impose a positive environmental perturbation that significantly and reversibly altered the abundance of gut microbiome members [16, 17]. Co-abundance network analysis performed at each of the three time points allowed us to identify pairs of MAGs whose correlations remained unchanged despite significant changes in the abundance of the entire community caused by the perturbation. We found that these genomic pairs originated from 141 MAGs and formed two bacterial communities that were organized at opposite ends of a robust, stable seesaw-like network. Together, these seesaw network-like gene panels enabled machine learning models for predicting responses of multiple metabolic phenotypes to dietary intervention in a T2DM cohort and for predictive classification of cases versus controls from 12 independent metagenomic datasets of 1,874 subjects across diverse cohorts and chronic diseases, including T2DM, atherosclerotic cardiovascular disease (ACVD), hypertension, liver cirrhosis (LC), inflammatory bowel disease (IBD), colorectal cancer (CRC), ankylosing spondylitis (AS), schizophrenia, and Parkinson's disease (PD), suggesting the identification of common microbiome signatures across diverse human diseases.

[0325] Reversible changes in the gut microbiota are associated with reversible changes in the host metabolic phenotype

[0326] Throughout the study, the dietary fiber intake of the U group remained unchanged, while the dietary fiber intake of the W group showed a substantial increase from M0 to M3 and a decrease from M3 to M15, but was still higher than that of M0 ( Figure 5B Compared with the U group, the W group had significantly higher fiber intake at M3 and M15, but similar energy and macronutrient expenditure (Figure 10).

[0327] To investigate structural changes in the gut microbiota in response to the introduction and withdrawal of a high-fiber intervention, shotgun metagenomic sequencing was performed on 315 stool samples collected from 110 patients in the W and U groups, of whom 95 patients provided samples at all three time points, while 15 patients provided samples only at M0 and M3 ( Figure 9). To achieve strain and subspecies level resolution, 1,845 non-redundant high-quality draft genomes were reconstructed from the metagenomic dataset (if the average nucleotide identity ANI between two genomes was > 99%, the two genomes were superimposed into one MAG), and these MAGs accounted for more than 70% of the total reads. In the context of β-diversity based on Bray-Curtis distance, the overall structure of the intestinal flora in the W group changed significantly from M0 to M3 (PERMANOVA test, P < 0.001) and returned to the structure of M0 at M15, while there was no difference in the U group at the three time points ( Figure 5C , D). Similar changes were observed in alpha diversity based on the Shannon index and Simpson index (Figure 11). These results show that, as previously reported, high-fiber intervention triggered significant structural changes in the gut microbiota

[16] , however, after withdrawal of the intervention, the gut microbiota returned to baseline, indicating that the community structure had high resilience.

[0328] To determine whether the host metabolic phenotype would also show reversible changes similar to those in the gut microbiota, we examined 43 bioclinical parameters at three time points. Hemoglobin A1c (HbA1c) in the U group did not show any changes throughout the trial. The high-fiber intervention reduced the HbA1c level in the W group by an average of 15.22% ± 9.82% (mean ± sd) from M0 to M3, and this reduction was significantly greater than that in the U group. At one-year follow-up, HbA1c in the W group was significantly higher than that in M3, but still lower than that in M0 ( Figure 5E At M3, the proportion of patients achieving adequate glycemic control (HbA1c ≤ 7%) was also significantly higher in the W group (61.6% vs. 33.3% in the U group), but no difference was shown between the two groups at M15 ( Figure 5F Fasting and postprandial blood glucose levels in the meal tolerance test showed a similar trend to HbA1c ( Figure 5G , H). The W group also showed a reduction in inflammation, hyperlipidemia, obesity, and T2DM complications from M0 to M3, but rebounded at one-year follow-up (Figure 22). Mantel tests based on the Manhattan distance of all 43 bioclinical parameters and the Bray-Curtis distance of the gut microbiota showed that clinical outcomes were significantly correlated with gut microbial structure (R2 = 0.09, P = 2 × 10-4). These results suggest that changes in host metabolic phenotype are associated with reversible changes in the gut microbiota in response to the presence / absence of high-fiber intervention.

[0329] Stably interacting genome pairs form a seesaw network with two competing bacterial populations

[0330] To facilitate the identification of genome pairs that maintain stable ecological interactions during the experiment, a co-abundance network was constructed for each time point based on the abundance matrix of MAGs representing prevalent microorganisms. A total of 477 MAGs were selected for network construction because they were detected in more than 75% of the samples at each time point in the W group. They were also dominant because they accounted for approximately 60% of the total abundance of 1,845 MAGs. Pairwise correlations were calculated for all 113,526 possible genome pairs in these 477 prevalent MAGs, and three co-abundance networks were constructed, one for each time point (GM0, GM3, and GM15) ( Figure 23 ). The co-abundance network of the prevalent gene sets in the W group at M0, M3 and M15 during the trial is represented as GM0 (442; 4231), GM3 (421; 2587) and GM15 (429; 4592). The numbers in brackets are the order and size of the network. The correlation between gene sets was calculated using FastSpar, n = 67 patients. All significant correlations with P ≤ 0.001 were included. The co-abundance network was visualized. The edges between nodes represent the correlation. Red and blue indicate positive and negative correlations, respectively. The node size indicates the average abundance of the gene set. The layout of nodes and edges was determined by edge-weighted Spring embedded layout with correlation coefficients as weights. The three networks have similar order S, i.e., the total number of nodes (MAG), SM0 (442), SM3 (421) and SM15 (429), but their size L, i.e., the total number of edges (correlations), LM0 (4231), LM3 (2587) and LM15 (4592) are very different. L decreased to 61.14% of GM0 in GM3 and rebounded to 108.53% of GM0 in GM15. This was confirmed by the change in connectivity, which is defined as the ratio of realized ecological interactions among potential ecological interactions (in undirected networks, connectivity = L / (S(S-1) / 2), whose values ​​are in the interval [0, 1])20. Connectivity decreased from 0.043 in GM0 to 0.029 in GM3 and rebounded to 0.050 in GM15. The high fiber intervention greatly reduced the interactions between common gene sets in the network. In addition, it was found that the distribution of degree (i.e., the number of edges a node has) was well fitted by the power law model (Figure 12, R2 values ​​ranged from 0.79 to 0.82), indicating that there are a small number of nodes with high degrees (a characteristic of scale-free networks) that can resist random errors or decay

[21] . Here, a hub is defined as a node that is connected to more than one-fifth of the total nodes in the network ( Figure 13Of the 24 hubs, 10 were in GM0 and 20 were in GM15, but none were in GM3. This indicates that the overall structure of the gut microbiome may have undergone profound changes during the trial, especially as the high-fiber intervention led to a loss of interactions between pairs of genomes.

[0331] If a genome pair maintained the same ecological interaction across all three time points, their ecological relationship was considered robust and stable. Of the 113,526 possible genome pairs, 92.39% had no correlation at any of the three time points, indicating that establishing an ecological relationship between two genomes is a rare event ( Figure 6A ). Among the 477 prevalent gene sets, 184 gene sets had 517 positive correlations and 118 negative correlations across all three time points. The co-abundance network of the 184 gene sets whose correlations remained unchanged in GM0, GM3, and GM15 was visualized. The correlations between gene sets were calculated using FastSpar. All significant correlations with P ≤ 0.001 were included. The node size indicates the average abundance of the gene set. The lines between the nodes represent the correlations, and red and blue indicate positive and negative correlations, respectively. The edge-weighted Spring embedded layout was applied to layout the network with the correlation coefficient as the weight (-1 or 1). Of these 184 gene sets, 43 were excluded from subsequent analysis because they had no interactions with the remaining 141 nodes. The remaining 141 gene sets (which included 586 gene set pairs with stable correlations throughout the experiment) became candidates for microbiome signatures due to their robust interactions. The researchers then explored how these 141 gene sets were connected to each other and to the remaining nodes ( Figure 14A ). 141 genomes have significantly higher degree, betweenness centrality, eigenvector centrality, closeness centrality, and pressure centrality than the rest of the genomes in the network ( Figure 14B -F). This indicates that the 141 gene groups exert relatively strong control over the interactions with other nodes (reflected by betweenness centrality and eigenvector centrality) and the information flow in the network (reflected by proximity centrality and pressure centrality). Deleting these 141 nodes causes the network to collapse, as an average of 86.08% of the total edges disappear. This suggests that the 141 gene groups can be considered core nodes of the network, as they are highly connected not only internally but also with other nodes.

[0332] These 141 gene sets were also highly prevalent among the participants, as 140 of them were present in >90% of the 74 individuals in the W group, and 104 of them were present in 100% of the individuals ( Figure 15). These 141 gene sets were also mostly dominant members of the gut microbiota, as 111 of them had a higher abundance than the median of the 1,845 MAGs. Based on the Bray-Curtis distance, β-diversity analysis showed a significant correlation between the profiles of the 141 MAGs and the total 1,845 MAGs, as demonstrated by the Mantel test (R2 = 0.62, P = 0.001) and the Procrustes analysis (P = 0.001) (Figure 16, Figure 5C , D). These indicated that changes in 141 MAGs led to major shifts in the entire gut microbiome at the three time points.

[0333] Bacteria that are positively correlated with each other and exhibit robust co-occurrence behavior can be considered as ecological communities [5]. The 141 genomes organized themselves into two communities, and the genomes of each community were highly interconnected and positively correlated. There were 50 genomes in community 1 and 91 genomes in community 2 (Figure 6B). The correlations between the 141 genomes were calculated using FastSpar, n = 67 patients. All significant correlations with P ≤ 0.001 were included. The co-abundance network within the 141 genomes was visualized. The edges between the nodes represent the correlations. Red and blue indicate positive and negative correlations, respectively. The color of the nodes represents the membership of the two communities: green for community 1 and purple for community 2. All genomes in community 1 were from the Firmicutes phylum, while all genomes in community 2 were from five different phyla, including Firmicutes, Bacteroidetes, Proteobacteria, Actinobacteria, and Fusobacteria. There were only negative edges between the two communities, indicating a competitive relationship. The abundance of members of cluster 1 increased from M0 to M3 and then decreased from M3 to M15, while members of cluster 2 showed the opposite pattern of change (Figure 6B). Therefore, there is a robust cooperative relationship between members within each cluster, while there is a competitive relationship between the two clusters. The seesaw network with 141 nodes is in two polarized clusters. In this network, the edges between nodes represent correlations. Red and blue indicate positive and negative correlations, respectively. Our data showed that in the W group, before and after the high-fiber intervention and at the one-year follow-up, the two clusters of 141 genomes formed a stable seesaw network that existed in all three ecological networks. The co-abundance network in the U group was constructed based on the 141 genomes, and the seesaw network was also observed at each time point. FastSpar was used to calculate the correlation between the genomes at each time point in the U group. All significant correlations with P≤0.001 were included. The network was visualized as follows: the lines between the nodes represent correlations, and red and blue indicate positive and negative correlations, respectively. Node size indicates the average abundance of the genome in patient samples with data at three time points in the U group (n = 28). The color of the node represents the members of the two bacterial groups: bacterial group 1 is green and bacterial group 2 is purple. The percentage of correlation follows the microbiome characteristics of the seesaw network pattern (i.e., positive edges within each bacterial group and negative edges between the two bacterial groups) in yellow, and the ratio of negative correlation within each bacterial group to positive correlation between bacterial groups is represented by black in the 100% stacked bar chart.

[0334] This suggests that the detection of a stable seesaw network-like genome is not dependent on high-fiber intervention but may be an inherent feature of the human gut microbiome. Metagenomic functions of two competing bacterial communities regulate host metabolic phenotypes.

[0335] We then determined whether the balance between the two competing microbiota could be modulated by dietary fiber and how this affected host metabolic phenotypes. From M0 to M3, the total abundance of community 1 increased significantly, while the total abundance of community 2 decreased significantly. At M15, community 1 decreased to a level similar to M0, and community 2 rebounded but remained lower than M0. Subsequently, high-fiber intervention significantly increased the ratio of community 1 to community 2 from M0 to M3. At the one-year follow-up period, the ratio decreased significantly and was not different from baseline ( Figure 7A ). Overall, the mantel test based on the Manhattan distance of all 43 bioclinical parameters and the Bray-Curtis distance based on the 141 genomes showed that clinical outcomes were significantly associated with the variation of the seesaw network-like genome (R2=0.11, P=1×10-4). The association between the 141 genomes and each host bioclinical parameter was explored using a machine learning algorithm. Random forest regression based on the 141 genomes with leave-one-out cross-validation predicted 41 of the 43 bioclinical parameters, with significant Pearson correlation coefficients between the predicted and measured values ​​between 0.11 and 0.44 ( Figure 7B These results revealed that the 141 genomes function as two competing bacterial groups in a seesaw network, constituting an important microbiome signature for T2DM and associated metabolic phenotypes.

[0336] To explore the genetic basis for the association between the dynamic changes in the seesaw network-like microbiome signature and the response to host metabolic phenotypes, a genome-centric analysis of the metagenomes of the two competing bacterial communities was performed. Because dietary fiber can alter the balance between the two bacterial communities, genes encoding carbohydrate-active enzymes (CAZy) and genes encoding key enzymes in short-chain fatty acid (SCFA) production were identified to compare the genetic capacity for carbohydrate utilization between the two communities. Compared with the genomes in community 2, the genomes in community 1 had a higher proportion of CAZy genes targeting arabinoxylan (P < 0.001) and cellulose (P < 0.01), and a lower proportion of CAZy genes targeting inulin utilization (P < 0.01) (Figure 7C). There were no differences in genes targeting starch, pectin, and mucin utilization between the two communities. Our previous study showed that the gut microbiome benefits patients with T2DM by producing acetate and butyrate through carbohydrate fermentation

[16] . Among the terminal genes in the butyrate biosynthesis pathway, both from carbohydrates (i.e., but and buk) and proteins (i.e., atoA / D and 4Hbt), the copy number of but was significantly higher in colony 1, while there were no differences in the other terminal genes between the two colonies (Figure 7C). More than one-third of the genome in colony 1 contained the but gene, while less than 5% of the genome in colony 2 contained this gene (Fisher's exact test P < 0.001). Compared with colony 2, colony 1 also showed a trend towards higher heritability for acetate production (P = 0.06), but lower heritability for propionate production (P < 0.05) (Figure 7C). These results indicate that colony 1 has a significantly higher heritability for utilizing complex plant polysaccharides and producing acetate and butyrate than colony 2.

[0337] From a pathogenicity perspective, 21 of the 1845 MAGs encoded 750 virulence factor (VF) genes. Of the 21 VF-encoding genomes, 3 belonged to cluster 1, while 18 belonged to cluster 2. Three of the 50 genomes in cluster 1 harbored a VF gene involved in antiphagocytosis. In cluster 2, 18 of the 91 genomes encoded 747 VF genes across 15 different VF categories (i.e., acid resistance, adhesion, antiphagocytosis, biofilm formation, efflux pumps, endotoxins, invasion, iron uptake, manganese uptake, motility, nutritional factors, proteases, regulation, secretion systems, and toxins) (Figures 7C and 18A). Notably, 98.53% of all VF genes in cluster 2 were hidden in 8 genomes (1 in Enterobacter kobei, 2 in Escherichia flexneri, 3 in Escherichia coli, and 2 in Klebsiella). The high enrichment of virulence factor genes in the genomes of cluster 2 (P < 2.2 × 10-16, Fisher's exact test) suggests that this cluster may play an important role in aggravating metabolic disease phenotypes. In terms of antibiotic resistance genes (ARGs), in cluster 1, only 1 genome (accounting for 2.00% of the genomes of this cluster) contained a copy of an ARG related to phenol ( Figure 5C , S18B). In cluster 2, 17 genomes (accounting for 18.68% of the genomes in this cluster) encode 40 ARGs that confer resistance to seven different antibiotic classes (i.e., aminoglycosides, β-lactams, fosfomycin, glycopeptides, quinolones, macrolides, and tetracyclines). Therefore, cluster 2 may serve as a reservoir of ARGs for horizontal transfer to opportunistic pathogens.

[0338] In summary, our data suggest that two competing microbiotas have different genetic capabilities, with microbiota 1 being potentially beneficial and microbiota 2 being detrimental. Boosting microbiota 1 with dietary fiber shifts the balance between the two microbiotas, resulting in more favorable metabolic outcomes.

[0339] Stable network-like microbiome signatures exist across ethnic and geographic cohorts

[0340] It was then asked whether these 141 genomes, organized as two competing groups in a stable seesaw network, could be a common microbiome signature across different diseases in other independent metagenomic study cohorts. To answer this question, we used the 141 identified genomes as reference genomes to retrieve their abundance from an independent T2DM study

[22] ( Figure 19). In this validation dataset, the reference genome accounted for 32.93% of the total abundance, and 135 reference genomes were constructed into a co-abundance network, in which 99.21% of all edges followed a seesaw network pattern of microbiome characteristics (i.e., positive edges within each bacterial community and negative edges between two bacterial communities). This further supports the existence of a similar seesaw network in T2DM patients. In addition, in the 136 healthy controls of the same study

[22] , the reference genome accounted for 35.29% of the total abundance, and 128 genomes were constructed into a co-abundance network, in which 98.60% of all edges of the network conformed to the seesaw model. In the context of β diversity based on Bray-Curtis distance, the microbiome characteristics between T2DM patients and healthy controls showed significant differences based on the abundance matrix of the reference genome. In the principal coordinate analysis plot based on Bray-Curtis distance, the composition of the microbiome characteristics between controls and patients was different. 95% confidence ellipses are projected for controls and patients respectively. The p-value of the PERMANOVA test is indicated. Furthermore, a random forest regression model was trained using the abundance matrix of genomes in the microbiome signature and phenotypic data, and the predicted values ​​of BMI, fasting insulin, and HbA1c from the model were found to be significantly correlated with the measured values ​​( Figure 20 A random forest model was trained to see if it could discriminate between patients and controls. Receiver operating characteristic curve analysis demonstrated moderate predictive ability, with an area under the curve (AUC) of 0.70 after leave-one-out cross-validation. Thus, the seesaw network-like microbiome signature not only persists but also maintains a similar relationship with host metabolic phenotypes in an independent T2DM study.

[0341] We further hypothesized that the seesaw network-like microbiome signature represents an intrinsic feature of the human gut microbiome, the disruption of which may be associated with diseases other than T2DM. First, the same validation analysis was performed on metagenomic datasets from case-control studies of three different types of diseases, including ACVD23 (a chronic metabolic disease), LC24 (a liver disease), and AS25 (an autoimmune disease). Members of the two competing bacterial groups in the seesaw network-like microbiome signature showed similar ecological interactions in four independent human gut metagenomic datasets. Correlations between genomes were calculated using FastSpar. All significant correlations (P≤0.001) belonged to the seesaw model (positive correlations within bacterial groups and negative correlations between bacterial groups). The network was visualized as follows: the lines between nodes represent correlations, and red and blue indicate positive and negative correlations, respectively. The color of the node represents the membership of the two seesaw groups: green for bacterial group 1 and purple for bacterial group 2. The percentage of correlations that follow a seesaw network-like microbiome signature (i.e., positive edges within each microbiome, negative edges between two microbiome signatures) is shown in yellow, and the ratio of negative correlations within each microbiome to positive correlations between microbiome signatures is a 100% stacked bar in black. In ACVD patients and their controls, reference genomes from microbiome signatures accounted for 32.73% and 36.22% of the total abundance, respectively, and 139 genomes from patients and 137 genomes from controls were constructed into a co-abundance network, of which 94.33% and 98.49% of all edges were consistent with the seesaw model, respectively. The reference genomes of the microbiome signature accounted for 33.84%, 35.83%, and 41.02% of the total abundance in the metagenomic datasets of healthy controls (the same control cohort was used in the LC and AS studies), LC, and AS patients, respectively. 117, 125, and 123 reference genomes were constructed into co-abundance networks, of which 100%, 98.68%, and 88.54% of all total edges were consistent with the seesaw network model in the metagenomic datasets of healthy controls, LC, and AS patients, respectively. In the PCoA plot based on Bray-Curtis distance, the microbiome signatures in all three datasets showed significant differences between controls and patients. For the LC study, a random forest model was trained using the abundance matrix of the reference genome and phenotypic data, and the predicted values ​​of total bilirubin, albumin levels, and BMI based on the model were found to be significantly correlated with the measured values ​​(Figure 21). Compared with the T2DM dataset

[22] , the random forest classifier based on microbiome signatures showed better predictive ability in distinguishing cases from controls of ACVD (AUC = 0.81), LC (AUC = 0.91), and AS (AUC = 0.98) (Figure 8A).To further confirm the relevance of microbiome signatures to human disease, we tested genomes derived from microbiome signatures in a wider range of disease types and across diverse ethnicities and geographic regions. These datasets included hypertension (Chinese cohort), IBD (US cohort and Dutch cohort), CRC (Chinese cohort and Australian cohort), schizophrenia (Chinese cohort), and PD (Chinese cohort). On average, the reference genomes accounted for 31.82% ± 4.05% (mean ± sd) of the total abundance of the entire microbial community in the dataset. The microbiome signature has been shown to perform predictive classification of cases and controls in metagenomic datasets from studies for hypertension

[26] (AUC = 0.74), IBD (AUC = 0.70 for IBD dataset 127, AUC = 0.90 for IBD dataset 228, and AUC = 0.83 for IBD dataset 328), CRC (AUC = 0.73 for CRC dataset 129 and AUC = 0.74 for CRC dataset 230), schizophrenia (AUC = 0.69), and PD31 (AUC = 0.76) (Figure 22). These results show that our microbiome signature is present in healthy controls from independent studies and in a variety of patient populations across different ethnicities and regions. The association between the 141 gene sets and host phenotypes and the ability of these gene sets to discriminate between controls and patients with various types of diseases as biomarkers suggest that these seesaw network-like gene sets organized into two groups represent common microbiome signatures associated with a wide variety of human disease phenotypes.

[0342] The classification performance of eight diseases was verified using different numbers of gene sets selected by degree-based reverse selection. Figure 26A )、ACVD( Figure 26B )、LC( Figure 26C )、AS( Figure 26D )、PD( Figure 26E )、SCZ( Figure 26F ), CRC-1, CRC-2, CRC-3( Figures 26G-26I )、IBD-1、IBD-2、IBD-3( Figures 26J-26L ),hypertension( Figure 26M ) were used to build random forest regression models for eight different types of diseases using 13 datasets obtained from the National Natural Science Foundation of China (NSFC). Abundant reads associated with 141 identified genomes were recruited to classify healthy subjects from patients. For each classification model, all selected genomes were ranked by their degree (connectivity). Genomes with relatively low degrees were gradually removed. After removing each genome, the predictive ability corresponding to the area under the curve (AUC) was calculated. When fewer than about 30 higher-degree genomes were used to classify healthy subjects from patients, the predictive ability began to decline.

[0343] We further validated the classification performance of eight types of diseases using different numbers of randomly selected gene sets. Figure 27A )、ACVD( Figure 27B )、LC( Figure 27C )、AS( Figure 27D )、PD( Figure 27E )、SCZ( Figure 27F ), CRC-1, CRC-2, CRC-3( Figures 27G-27I )、IBD-1、IBD-2、IBD-3( Figures 27J-27L ),hypertension( Figure 27M ) were used to construct random forest regression models for eight different disease types using 13 datasets obtained from a random forest regression model. Abundant reads associated with 141 identified gene sets were recruited to classify healthy subjects from patients. For each classification model, classification performance was predicted using a varying number of randomly selected gene sets. Predictive power, expressed as the area under the curve (AUC), was calculated for each set of randomly selected gene sets. Predictive power began to decline when fewer than approximately 30 randomly selected gene sets were used to classify healthy subjects from patients.

[0344] discuss

[0345] As described here, a genome-based, reference-free, and ecological interaction-focused approach led to the identification of a stable seesaw network of genomes from two competing microbiota, alterations of which were associated with a wide range of host phenotypes in patients with type 2 diabetes. Furthermore, random forest models based on these genomes predicted the classification of cases and controls for multiple diseases, suggesting that these genomes may form a common microbiome signature across a wide range of ethnicities, geographic regions, and disease states.

[0346] The genomes in this common microbiome signature are organized into a seesaw-like network that combines cooperative and competitive interactions. Although cooperative ecological networks may be effective, they create dependencies and the potential for mutual collapse, which may have destabilizing effects on the human gut microbiome. This destabilizing effect of cooperation can be mitigated by introducing ecological competition into the network

[32] . Therefore, a seesaw-like network with both cooperative and competitive interactions may represent a stable microbiome structure

[32] . Interestingly, although the seesaw-like network is stable, the weights at the two ends (i.e., the abundance of community 1 and community 2) are modifiable, and such changes are associated with the health of the host. When large amounts of complex fiber become available, community 1 and community 2 do not change in membership or the nature of their interactions with each other, but instead experience large changes in community-level abundance in a competitive manner. Members of community 1 have a higher genetic capacity for degrading complex plant polysaccharides and produce beneficial metabolites (including SCFAs) that may suppress pathogenic organisms in community 216. Members of commensal 2 need to be kept at low levels, as their overgrowth may compromise host health through, for example, increased inflammation

[33] . However, pathogenic organisms in commensal 2 cannot be eliminated; for example, they may serve as essential agents for training our immune system from early in life [34, 35]. Therefore, the balance between commensal 1 and commensal 2 becomes crucial in determining whether the gut microbiome maintains health or exacerbates disease. This seesaw-like network between commensal 1 and commensal 2 allows the genomes in our common microbiome signature to readily respond to changes in external energy inputs to the gut microbial ecosystem and regulate the effects of external energy inputs on host health while maintaining their structural integrity. This structural integrity may be key to ensuring the long-term ecological stability of the gut microbiome and its ability to provide essential health-related functions to the host.

[0347] This seesaw network structure may have been stabilized by natural selection during the long-term coevolution between the microbiome and its host [18, 36]. This selection pressure may be exerted by dietary fiber, which only directly interacts with the intestinal microorganisms as an external energy source [37, 38]. Studies of coprolites show that ancient humans had much higher dietary fiber intakes and have only significantly decreased in the past 150 years [39, 40] (the intake of plant fiber in prehistoric diets was 130g / d

[41] , while the median intake in the modern American diet is 12-14g / d

[42] ). Such high fiber intake in evolutionary history may have favored beneficial bacteria in group 1 because their higher genetic ability to use plant polysaccharides as an external energy supply enables them to gain a competitive advantage over pathogenic organisms in group 2 in the intestinal microbial ecosystem

[43] . Similar to tall trees that are the foundation species of closed forests, group 1 can serve as a "foundation bacteria" to stabilize the healthy intestinal microbiome and prevent pathogens

[44] . The health benefits of dietary fiber in preventing and alleviating a variety of chronic diseases have been demonstrated epidemiologically and clinically, suggesting that the dominance of commensal microbiota 1 over commensal microbiota 2 can improve host health [16, 38, 45, 46].

[0348] Furthermore, the genomes in our seesaw network-like common microbiome signature cluster can be considered as part of the human core gut microbiome [47, 48]. This is because: 1) they are common across different ethnic and geographical populations; 2) they show temporal stability not only in membership but also in interactions with each other and with the host; 3) they account for approximately 10% of the gut microbiome but are disproportionately important in shaping the ecological community; 4) they provide essential health-related functions to the host; and 5) this core microbiome organized into a seesaw network may have been established during long-term coevolution and has become the ecological basis for regulating host health.

[0349] In fact, this seesaw network can be detected in other independent metagenomic datasets and has been shown to be associated with different diseases, suggesting that this evolutionarily conserved ecological structure may be critical for the restoration and maintenance of human health. In addition, the seesaw network structure showed stable relationships both internally within the network and externally with multiple host clinical markers, suggesting that genome-based microbiota can serve as robust disease biomarkers. Within the seesaw network, the imbalance between the two competing microbiota may serve as a common biological basis for many human diseases. Targeting this core intestinal microbiome, restoring and maintaining the dominance of beneficial bacteria over harmful bacteria may help reduce disease risk or alleviate symptoms, thereby opening up new avenues for the management and prevention of chronic diseases.

[0350] Materials and Methods

[0351] Clinical trials

[0352] Study design

[16] : This clinical trial was conducted at Qidong People's Hospital (Jiangsu, China) to investigate the effects of a high-fiber diet under free-living conditions in a cohort of individuals with clinically diagnosed T2DM (Qidong). The study protocol was approved by the Ethics Committee of Shanghai General Hospital (2014KY104), and the study was conducted in accordance with the principles of the Declaration of Helsinki. All participants provided written informed consent. The trial was registered with the Chinese Clinical Trial Registry (ChiCTR-IPC-14005346). Study design and participation procedures are as follows: Figure 9 shown.

[0353] This study recruited Chinese Han patients with T2DM (age: 37-70 years; HbA1c: 6.5%-12.0%). A more detailed description of the inclusion and exclusion criteria is shown in the Chinese Clinical Trial Registry (chictr.org.cn).

[0354] Patients received a high-fiber diet (WTP diet) as a treatment group (W group) or usual care (normal diet) as a control group (U group) for 3 months. Total calorie and macronutrient prescriptions were based on the age-specific Chinese Dietary Reference Intakes (Chinese Nutrition Society, 2013). The WTP diet was based on whole grains, traditional Chinese medicinal foods, and prebiotics, and included three ready-to-eat prepared foods

[16] . The daily diet, including standard dietary and exercise recommendations, was formulated according to the Chinese Diabetes Society T2DM Guidelines

[49] . Patients in the W group were given the WTP diet as a self-administered intervention at home for 3 months, while patients in the U group received usual care. The WTP diet intervention was discontinued in the W group at the end of the third month (at M3). Then, W and U continued to be followed up for one year (M15). Nutrient intake was calculated according to the Chinese Food Composition 200950 using a meal-based food frequency questionnaire and 24-h dietary recall. Both groups of patients continued to take antidiabetic medications as prescribed by their doctors.

[0355] Figure 22A and 22BClinical parameters during the intervention period in the W and U groups are collectively described. Data are shown as mean ± SEM (N). Intra-group comparisons were performed using the Friedman test followed by the Nemenyi post hoc test, with identical letters (a, b, or c) indicating no significant difference and different letters indicating significant difference (P < 0.05). Comparisons between W and U at the same time point were performed using the Mann-Whitney test (two-sided). FBG, fasting blood glucose; MTT glucose AUC, area under the glucose curve (AUC) in a meal tolerance test; MTT C-peptide AUC, area under the C-peptide curve (AUC) during meal tolerance test; HOMA-IR = 1.5 + FBG * fasting C-peptide / 2800; HOMA-β = 0.27 * fasting C-peptide / (FBG-3.5); BMI, body mass index; BD, body weight; SBP, systolic blood pressure; DBP, diastolic blood pressure; WC, waist circumference; HP, hip circumference; WHR, waist-to-hip ratio; TNF-α, tumor necrosis factor-α; WBC, white blood cell count; CRP, C-reactive protein; LBP, lipopolysaccharide binding protein; TC, total cholesterol; TG, triglycerides; Lpa, lipoprotein a; HDL, high-density lipoprotein; APOA, apolipoprotein A; LDL, low-density lipoprotein; APOB, apolipoprotein B; Lipoprotein B, GFR, glomerular filtration rate, CysC, cystatin C, ACR, urinary microalbumin-to-creatinine ratio, IMT, intima-media thickness, DAN, diabetic autonomic neuropathy score, MHR, mean heart rate, SDNN, standard deviation of the NN interval, SDANN, standard deviation of the mean NN interval calculated over a 5-minute period, SDNNIndex, mean of the standard deviations of the NN intervals over a 5-minute period, rMSSD, root mean square of the differences between consecutive NN intervals, pNN50, percentage of consecutive NN interval differences greater than 50 ms, TP, total power, VLF, very low-frequency power, LF, low-frequency power, HF, high-frequency power, and DPN, diabetic peripheral neuropathy score.

[0356] Before the 2-week run-in period, all participants attended a lecture on diabetes intervention and improvement and received diabetes education and metabolic assessment. Based on the inclusion and exclusion criteria, 119 eligible individuals were enrolled and divided into two groups (n = 79 in the W group and n = 40 in the U group) in a 2:1 ratio using SAS software.

[0357] Physical examinations were performed at Qidong People's Hospital (Jiangsu, China) at M0, M3, and M15. Instructions for sample collection were provided to participants the day before. Participants provided stool and first morning urine as required. After a fasting venous blood sample was collected, a 3-h meal tolerance test (Chinese steamed buns containing 75 g of effective carbohydrate; MTT test was performed) was performed and venous blood samples were collected 30, 60, 120, and 180 min after the meal. All blood samples were kept at room temperature for 30 min to obtain serum, which was centrifuged at 3000 rpm for 20 min at 4°C. The fasting serum was divided into two parts, one for hospital testing and the other for laboratory testing. Stool, urine, and serum samples were immediately stored on dry ice and then transported to the laboratory and frozen at -80°C. Subsequently, anthropometric markers and the diabetes complication index were measured. Ewing's test

[51] and 24-h Holter monitoring were performed to evaluate diabetic autonomic neuropathy (DAN). B-mode carotid ultrasound was performed to evaluate atherosclerosis. The Michigan Neuropathy Screening Scale

[52] was administered to assess diabetic peripheral neuropathy (DPN). In addition, a meal-based food frequency questionnaire and a 24-h dietary recall were recorded for nutrient intake calculation. In addition, medication use was self-reported.

[0358] Fasting venous blood was used to measure HbA1c, fasting blood glucose, fasting insulin, fasting C-peptide, C-reactive protein (CRP), a complete blood test, blood biochemistry tests, and five thyroid analytes. Postprandial blood glucose, insulin, and C-peptide were measured using venous blood samples collected at 30, 60, 120, and 180 minutes after the MTT. Urinalysis and measurement of the urine microalbumin-to-creatinine ratio were performed using early morning fasting urine. The above measurements were performed at Qidong People's Hospital. TNF-α (R&D Systems, MN, USA), lipopolysaccharide binding protein (Hycult Biotech, PA, USA), leptin (P&C, PCDBH0287, China), and adiponectin (P&C, PCDBH0016, China) were quantified by enzyme-linked immunosorbent assay (ELISA) using fasting venous blood at Shanghai Jiao Tong University.

[0359] Homeostasis model assessment of insulin resistance (HOMA-IR) and pancreatic β-cell function (HOMA-β) were calculated based on fasting blood glucose (mmol / L) and fasting C-peptide (pmol / L): HOMA-IR = 1.5 + FBG * fasting C-peptide / 2800;

[0360] HOMA-β = 0.27 * fasting C-peptide / (FBG-3.5). Glomerular filtration rate was estimated by the formula GFR (ml / min per 1.73 m2) = 186 * Scr - 1.154 * age - 0.203 * 0.742 (if female) * 1.233 (if Chinese) 54, where Scr (serum creatinine) is in mg / dl and age is in years.

[0361] Gut microbiome analysis

[0362] Metagenomic sequencing. DNA was extracted from fecal samples using a method previously described

[17] . Metagenomic sequencing was performed using an Illumina Hiseq 3000 from GENEWIZ (Beijing, China). Cluster generation of sequencing primers, template hybridization, isothermal amplification, linearization, and blocking denaturation and hybridization were performed according to the workflow specified by the service provider. The constructed libraries had an insert size of approximately 500 bp and were then subjected to high-throughput sequencing to obtain paired-end reads of 150 bp each in the forward and reverse directions.

[0363] Data quality control. Prinseq

[55] was used to: 1) trim reads from the 3′ end until the first nucleotide was reached with a quality threshold of 20; 2) remove read pairs when the read was <60 bp or contained “N” bases; and 3) dedupe the reads. Reads that could be aligned to the human genome (Homo sapiens, UCSC hg19) were removed (aligned with Bowtie2

[56] using --reorder --no-hd --no-contain --dovetail).

[0364] De novo assembly, abundance calculation, and taxonomic assignment of genomes. De novo assembly was performed for each sample using IDBA_UD

[57] (-step 20 --mink20 --maxk 100 --min_contig 500 --pre_correction). The assembled contigs were further binned using MetaBAT

[58] (--minContig 1500 --superspecific-B 20). The quality of the bins was assessed using CheckM

[59] . Bins with completeness > 95%, contamination < 5%, and strain heterogeneity < 5% were retained as high-quality draft genomes. The assembled high-quality draft genomes were further deduplicated using dRep

[60] . The abundance of the genome in each sample was calculated using DiTASiC

[61] , and the estimated counts with a P value < 0.05 were removed, and all samples were reduced to 36 million reads (a sample with a read mapping ratio < 25% was not fully represented by the high-quality genome and was removed from further analysis). The taxonomic assignment of genomes was performed using GTDB-Tk

[62] .

[0365] Functional analysis of the gut microbiome. Genomes were annotated using Prokka

[63] . KEGG ortholog (KO) IDs were assigned to predicted protein sequences in each genome using KofamKOALA

[64] via HMMSEARCH for Kofam. Antibiotic resistance genes were predicted using ResFinder

[65] with default parameters. Virulence factors were identified based on the core set of the Virulence Factor Database of Pathogens (VFDB

[66] , downloaded in July 2020). Predicted protein sequences were aligned with reference sequences in VFDB using BLASTP (best hit with E value < le-5, identity > 80% and query coverage > 70%). Genes encoding carbohydrate-active enzymes (CAZy) were identified using dbCAN (version 6.0)

[67] , and the best hit alignment was retained. Antibiotic resistance genes were identified using ResFinder

[68] . Genes encoding formate-tetrahydrofolate ligase, propionyl-CoA:succinate-CoA transferase, propionate CoA transferase, 4Hbt, AtoA, AtoD, Buk and But were identified as described previously

[16] .

[0366] Gut microbiome network construction and analysis. Fastspar

[69] was used to calculate the correlation between genomes with 1,000 permutations, and correlations with P ≤ 0.001 were retained for further analysis. In the W group, the common genomes shared by more than 75% of the samples at each time point were used to construct the co-abundance network at each time point. The network was visualized using Cystoscape v3.8.1

[70] . The layout of nodes and edges was determined by edge-weighted Spring embedded layout using correlation coefficients as weights. The links between nodes were regarded as metal springs attached to the node pairs. The correlation coefficient was used as a function to determine the repulsive and attractive forces of the springs

[70] . The layout algorithm set the positions of the nodes to minimize the sum of the forces in the network. Robust stable edges were defined as the same positive / negative correlations between the same two genomes in all three networks. Stable genome pairs were clustered based on robust positive edges (set to 1) and negative edges (set to -1) using the UPGMA clustering method. iTOL71 was used to integrate and visualize clustering trees, taxonomic information, and abundance changes of 141 genomes.

[0367] Validation in other independent cohorts. Independent metagenomic datasets from four case-control studies were included to validate the generalizability of the Seesaw network core. These datasets included 136 controls and 136 individuals with T2DM from Qin et al., 2012; 171 controls and 214 individuals with atherosclerotic cardiovascular disease from Jie et al., 2017; 83 controls and 84 individuals with cirrhosis from Qin et al., 2014; and 83 controls and 97 individuals with ankylosing spondylitis from Wen et al., 2017 (Table S8). DiTASiC was used to calculate the abundance of 141 genomes in each sample, and estimated counts with P values ​​< 0.05 were removed and further converted to relative abundance divided by the total number of reads. Fastspar was used to calculate correlations between genomes with 1,000 permutations, and correlations with P ≤ 0.001 were retained for network construction. A 30-fold repeated 5-fold cross-validation was used, and correlations shared by more than 95% of the 150 networks constructed by the cross-validation process were retained in the final network.

[0368] Statistical analysis.

[0369] Statistical analysis was performed in R (R version 3.6.1). Intra-group comparisons were performed using the Friedman test followed by the Nemenyi post hoc test. Comparisons between W and U at the same time point were performed using the Mann-Whitney test. Pearson chi-square tests were performed to compare differences in categorical data between groups or time points. The PERMANOVA test (9,999 permutations) was used to compare the gut microbiota structure of each group. A P value of less than 0.05 was considered statistically significant.

[0370] The Mann-Whitney test and Fisher's exact test were used to compare the functions of cluster 1 and cluster 2. Random forest regression and classification analyses based on microbiome signatures and clinical parameters / groups were performed using leave-one-out cross-validation.

[0371] To test whether the genomic abundance values ​​of the 141 microbiome groups identified above can be used to predict human health, random forest classifiers were trained on microbiome datasets obtained from patients and healthy controls in at least one study each of Parkinson's disease, schizophrenia, inflammatory bowel disease, colorectal cancer, and hypertension. Figure 21A and 21BTogether, we demonstrate the ability of microbiome signatures to discriminate between healthy subjects and patients as biomarkers in a broader dataset of diseases across ethnicities and geographic regions. The microbiome signatures supported predictive classification models for eight additional independent datasets. The area under the receiver operating characteristic (AUC) curve (ROC) of a random forest classifier for classifying controls and patients in each dataset based on 141 gene sets from the microbiome signatures was calculated. Leave-one-out cross-validation was used. Parkinson's disease (1): control n=40, Parkinson's disease n=39; Schizophrenia (2): control n=81, schizophrenia n=90; Colorectal cancer (CRC) dataset 1(3): control n=54, CRC n=74; Colorectal cancer (CRC) dataset 2(4): control n=63, CRC n=46; Inflammatory bowel disease (IBD) dataset 1(5): control n=26, IBD n=80; Inflammatory bowel disease (IBD) dataset 2(6): control n=34, IBD n=121; Inflammatory bowel disease (IBD) dataset 3(7): control n=22, IBD n=43. Hypertension (8): Control n=41, Hypertension n=99. 1=Qian, YW. et al. Gut metagenomics-derived genes as potential biomarkers of Parkinson's disease. Brain 143, 2474-2489 (2020). 2=Zhu, F. et al. Metagenome-wide association of gut microbiome features for schizophrenia. Nat Commun 11, 1612, doi: 10.1038 / s41467-020-15457-9 (2020). 3=Yu, J. et al. Metagenomic analysis of faecal microbiome as a tool towards targeted non-invasive biomarkers for colorectal cancer. Gut 66, 70-+, doi: DOI 10.1136 / gutjnl-2015-309800 (2017). 4=Feng, Q. et al. Gutmicrobiome development along the colorectal adenoma-carcinoma sequence. Nature communications 6, 6528, doi: 10.1038 / ncomms7528 (2015).5 = Lloyd-Price, J. et al. Multiomics of the gut microbial ecosystem in inflammatory bowel diseases. Nature 569, 655-662, doi: 10.1038 / s41586-019-1237-9 (2019). 6 = Franzosa, EA et al. Gut microbiome structure and metabolic activity in inflammatory bowel disease. Nature microbiology 4, 293-305, doi: 10.1038 / s41564-018-0306-4 (2019). 7 = Li, J. et al. Gut microbiota dysbiosis contributes to the development of hypertension. Microbiome 5, 14, doi: 10.1186 / s40168-016-0222-x (2017).

[0372] Because the corresponding microbiota communities are highly correlated, it was hypothesized that information about relative abundance values ​​would be highly correlated, such that fewer than the full set of genome abundance values ​​would provide adequate classification power. As a first test, multiple random forest classifiers were trained on microbiota datasets obtained from patients and healthy controls in at least one study each of type 2 diabetes (T2D), atherosclerotic cardiovascular disease (ACVD), liver cirrhosis (LC), ankylosing spondylitis (AS), Parkinson's disease (PD), schizophrenia (SCZ), colorectal cancer (CRC), inflammatory bowel disease (IBD), and hypertension. The first random forest classifier trained for each condition used genomic data from all 141 genomes listed in Table 1. Each subsequent classifier was trained on at least one genome, with the order of genome discarding determined by the degree of connectivity within the microbiota (the number of connections a genome has with other genomes in the microbiota) in ascending order, as shown in Table 5.

[0373] Table 5. Genome discarding order.

[0374]

[0375]

[0376]

[0377] The performance of each model was then determined as the AUC of the ROC curve and plotted as shown in Figure 26. As shown in Figure 26, the number of gene sets required to adequately characterize a clinical model of the disease state is less than the total 141 gene sets. In fact, in most (if not all) cases, models trained using only the 10-15 most relevant gene sets were adequate for clinical use (e.g., AUC of 0.65 or higher).

[0378] Next, determine how many randomly selected genomes from the 141 identified genomes are sufficient to meet the model with clinical utility. In short, multiple random forest classifiers were trained based on the microbiome data sets obtained for patients and healthy controls in at least one study of type 2 diabetes (T2D), atherosclerotic cardiovascular disease (ACVD), cirrhosis (LC), ankylosing spondylitis (AS), Parkinson's disease (PD), schizophrenia (SCZ), colorectal cancer (CRC), inflammatory bowel disease (IBD) and hypertension. Specifically, for each data set, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130 and 140 genomes randomly selected from the 141 genomes identified in Table 1 were used to train 10 classifiers (a total of 150 models for each data set). The average AUC of the ROC curve for each group of x randomly selected genomes was determined and plotted in Figure 27. As shown in Figure 27, the number of gene sets required to adequately characterize a clinical model of a disease state is less than the full set of 141. In fact, in most, if not all, cases, models trained using only 15-20 randomly selected gene sets are sufficient for clinical use (e.g., AUC of 0.65 or higher).

[0379] Example 2 - Identification of Human Disease Microbiome Signatures

[0380] To investigate the microbiome signatures of seven different diseases (type 2 diabetes (T2D), schizophrenia (SCZ), atherosclerotic cardiovascular disease (ACVD), liver cirrhosis (LC), inflammatory bowel disease (IBD), colorectal cancer (CRC), and ankylosing spondylitis (AS)), 11 metagenomic datasets were obtained from publications (Qin, J. et al. A metagenome-wide association study of gut microbiota in type 2 diabetes. Nature 490, 55-60, doi: 10.1038 / nature11450 (2012). Jie, Z. et al. The gut microbiome in atherosclerotic cardiovascular disease. Nat Commun 8, 845, doi: 10.1038 / s41467-017-00900-1 (2017). Qin, N. et al. Alterations of the human gut microbiome in liver cirrhosis. Nature 513, 59-64, doi: 10.1038 / nature 13568 (2014). Wen, C. et al. Quantitative metagenomics reveals unique gut microbiome biomarkers in ankylosing spondylitis. GenomeBiol 18, 142, doi: 10.1186 / s13059-017-1271-6 (2017). Zhu, F. et al. Metagenome-wide association of gut microbiome features for schizophrenia. Nat Commun 11, 1612, doi: 10.1038 / s41467-020-15457-9 (2020). Yu, J. et al.Metagenomic analysis of faecalmicrobiome as a tool towards targeted non-invasive biomarkers for colorectalcancer.Gut 66, 70-+, doi: DOI 10.1136 / gutjnl-2015-309800(2017).Feng, Q. et al. Gutmicrobiome development along the colorectal adenoma-carcinoma sequence. Nature communications 6, 6528, doi: 10.1038 / ncomms7528 (2015). Lloyd-Price, J. et al. Multi-omics of the gut microbial ecosystem in inflammatory bowel diseases. Nature 569, 655-662, doi: 10.1038 / s41586-019-1237-9 (2019). Franzosa, EA et al. Gutmicrobiome structure and metabolic activity in inflammatory bowel disease. Nature microbiology 4, 293-305, doi: 10.1038 / s41564-018-0306-4 (2019)). To achieve genome-level resolution, non-redundant high-quality draft genomes (HQMAGs) were reconstructed from metagenomic datasets. If the average nucleotide identity (ANI) between two HQMAGs was > 99%, two HQMAGs were superimposed into one.

[0381] To facilitate the identification of genome pairs that maintain stable ecological interactions between case and control groups, a co-abundance network was constructed for each case based on the HQMAG abundance matrix representing prevalent microorganisms. Co-abundance networks are a data-driven approach to studying ecological interactions between microorganisms across habitats. Prevalent HQMAGs in biological samples for each indication were selected for network construction. In each case cohort and its corresponding control cohort (e.g., type 2 diabetes group and healthy subject cohort), pairwise correlations of all possible genome pairs were calculated based on the abundance of these prevalent HQMAGs, and seven co-abundance networks were constructed. The network is represented by the order S (i.e., the total number of nodes (HQMAGs)) and the size L (i.e., the total number of edges (correlations)). Fastspar (a fast and scalable correlation estimation tool for microbiome studies) was used to calculate the correlation between 1,000 permuted genomes at each time point based on the abundance of genomes across patients, and correlations with P ≤ 0.001 were retained for further analysis. The network was visualized using Cystoscape v3.8.176. Edge-weighted Spring Embedded Layout uses correlation coefficients as weights to determine the layout of nodes and edges. The links between nodes are treated as metal springs attached to pairs of nodes. The correlation coefficients are used to determine the repulsive and attractive forces of the springs. The layout algorithm positions nodes to minimize the sum of the forces in the network. Differences in the co-abundance networks between the case and control cohorts were observed.

[0382] Genomes were considered to have robust and stable ecological relationships if the pairs maintained the same ecological interactions between the case and control cohorts. Robust stable edges were defined as unchanged positive / negative correlations between two identical genomes between the case and control cohorts. Stable genome pairs were clustered using the average clustering method based on robust positive edges (set to 1) and negative edges (set to -1). Various trees were displayed, manipulated, and annotated using iTOL77, an online tool, to integrate and visualize cluster trees, taxonomic information, and genome abundance changes.

[0383] Genome pairs with no correlation or no stable ecological interactions were removed from the analysis. Prevalent HQMAGs with positive and negative correlations were selected for further analysis. For each case cohort and its corresponding control cohort, the maximally interconnected HQMAG group was identified from all available interconnected HQMAG groups (C1, C2, ...), and HQMAGs that did not interact with the maximally interconnected HQMAG group were removed from further analysis.

[0384] The remaining HQMAGs (which included pairs of genomes with stable correlations) were further defined as genomes with stable ecological interactions (GSEIs) and became candidates for our microbiome signatures. GSEIs had significantly higher degrees, betweenness centrality, eigenvector centrality, closeness centrality, and stress centrality than the remaining genomes in the network. This suggests that GSEIs can be considered core nodes of the network because they are highly connected not only within themselves but also with other nodes.

[0385] Bacteria that are positively correlated and exhibit robust co-occurrence behavior can be considered ecological communities. The 141 GSEIs organized themselves into two competing communities, with the genomes of each community highly interconnected and positively correlated. Robust cooperation exists between members of each community, while competition exists between the two communities.

[0386] The process of identifying microbiome signatures from case and control cohorts is as follows Figure 36 As described.

[0387] Members of two competing bacterial communities identified from QD and various types of diseases (including type 2 diabetes (T2D), schizophrenia (SCZ), atherosclerotic cardiovascular disease (ACVD), liver cirrhosis (LC), inflammatory bowel disease (IBD), colorectal cancer (CRC), ankylosing spondylitis (AS)) were used to predict their ability to classify cases from controls by random forest classifiers in different studies. The AUC values ​​of the members from T2D ( Figure 28A )、LC( Figure 28B )、SCZ( Figure 28C )、IBD( Figure 28D )、AS( Figure 28E )、ACVD( Figure 28F )、CRC( Figure 28G ) and QD( Figure 28H Figure 28 shows that all microbiome signatures have the ability to classify cases and controls across studies. The precise classification ability of microbiome signatures varied between studies and between groups.

[0388] The classification ability of eight microbiome signatures was ranked based on their performance on 11 datasets. The eight microbiome signatures obtained from QD and multiple disease cases (T2D, LC, SCZ, IBD, AS, ACVD, CRC) were ranked based on their performance in classifying cases and controls in each dataset. The rank values ​​assigned to each signature microbiome were plotted on Figure 29A middle. Figure 29BThe sum of the rankings for each group of microbiome features is shown. The lower the rank, the better the classification performance. The microbiome features derived from QD had the best performance in classifying healthy subjects from patients in 11 datasets.

[0389] Example 3 - Identification of combinatorial core microbiome signatures from genome assembly pools.

[0390] 1. Genome Assembly Pools

[0391] 921 genomes (two competing bacterial groups found in QD and seven types of diseases including T2D, LC, SCZ, IBD, AS, ACVD, and CRC) were further deduplicated using dRep. If the average nucleotide identity ANI between two genomes was > 99%, they were superimposed into one. 788 non-redundant genomes were obtained. Pairwise ANI comparisons were performed on 310,078 genome pairs among the 788 genomes. The ANI distribution is shown in Figure 2. Figure 40A As shown. We further investigated the AM comparison between genomes assigned to two competing bacterial groups (Cluster 1 and Cluster 2). After removing genomes with inconsistent bacterial group assignments from the 788 genomes, Cluster 1 had 311 genomes and Cluster 2 had 440 genomes. The genome pairs between Cluster 1 and Cluster 2 were calculated by multiplying the total number of genomes in Cluster 1 (440) by the total number of genomes in Cluster 2 (331). The ANI distribution for the 136,840 genome pairs is shown in Figure 2. Figure 40B shown.

[0392] Genome abundance in each sample was calculated using DiTASiC, which applies kallisto for pseudo-alignment and uses a generalized linear model to account for shared reads between genomes, removing estimated counts with P values ​​> 0.05.

[0393] A machine learning classifier based on a random forest algorithm was trained to compare the ability of the combined 788 gene sets to individual microbiome signatures obtained from QD and multiple disease cases (including T2D, LC, SCZ, IBD, AS, ACVD, CRC) in classifying patients and controls. Figure 30A The area under the ROC curve (AUC) of the random forest classifier for classifying controls and patients in each dataset based on the combined pool or individual microbiome features is shown in Figure . Figure 30B Significance of intra-group comparisons is shown. Friedman test followed by Dunn post hoc test was performed for analysis (#BH adjusted P < 0.1, *BH adjusted P < 0.05). Overall, the combined pool had the best ability to classify cases and controls across studies.

[0394] The classification performance of each model was further ranked. The nine groups of microbiome signatures were ranked based on their performance in classifying cases from controls across 11 datasets. Figure 31A The rank values ​​assigned to each set of characteristic microbiome are plotted. Figure 31B The significance of within-group comparisons is shown. Figure 31C The sum of the ranking values ​​for each group of microbiome features is shown. Kruskal-Wallis test followed by Dunn's post hoc test was performed for analysis (#BH adjusted P < 0.1, *BH adjusted P < 0.05). The results confirmed that the microbiome features obtained from the combined pool had the best performance in classifying healthy subjects from patients in 11 datasets.

[0395] 2. Combined Core Pools of Genomes

[0396] The combined core pool of genomes was selected from the 788 genomes of combination by the step stated below. Each data set was subjected to a random forest classification based on the 788 genomes of combination. Each of the 788 genomes was ranked according to its importance for each data set. Overall ranking was obtained by adding the ranking values ​​in the 11 data sets, and all 788 genomes were ranked again based on the sum total. The most important genomes in the 11 data sets obtained the lowest overall ranking value (Table 3).

[0397] Table 3 - Ranking of genome importance

[0398] Genome Dataset 1 Dataset 2 Dataset 3 … Dataset 11 Overall ranking 1 D1R1 D2R1 D3R1 … D11R1 SR1 2 D1R2 D2R2 D3R2 … D11R2 SR2 3 D1R3 D2R3 D3R3 … D11R3 SR3 … … … … … … … 788 D1R788 D2R788 D3R788 … D11R788 SR788

[0399] Starting with the least important genome, each genome was removed from each dataset one by one according to the order of importance. The classification performance (AUC) of the remaining genome numbers after each round of removal was calculated using a random forest model, and all genome numbers were ranked based on the AUC value. The ranking values ​​for each genome number in the 11 datasets were summed (Table 4).

[0400] Table 4 - Ranking of genome numbers based on AUC

[0401]

[0402] exist Figure 32 The sum of the ranking values ​​for each genome number across the 11 datasets is plotted in Figure 3. 302 genomes achieved the lowest summed AUC ranking. After removing 18 genomes that exhibited inconsistent C1A and C1B assignments, 284 genomes remained as the combined core pool genomes.

[0403] The taxonomic abilities of two competing bacterial groups identified from T2D ( Figure 33A )、LC( Figure 33B )、AS( Figure 33C )、CRC( Figure 33D )、IBD( Figure 33E )、QD( Figure 33F )、AVCD( Figure 33G )、SCZ( Figure 33H ), combined collection ( Figure 33I ) and combined core pools ( Figure 33J The microbiome signatures identified for each condition were used to classify controls and patients in each dataset using a random forest classifier. Figure 31 shows that all microbiome signatures had the ability to classify cases and controls across the different studies.

[0404] As illustrated in Figure 34, the capacity of the combined core pool had the best capacity to classify cases and controls across the different studies. Figure 34A The areas under the ROC curves (AUCs) of random forest classifiers for classifying controls and patients in each dataset based on the combined core pool, combined pool, or individual microbiome signatures were compared. Figure 34B The significance of intra-group comparisons is shown. Friedman test followed by Dunn's post hoc test was performed for analysis (#BH adjusted P < 0.1, *BH adjusted P < 0.05, **BH adjusted P < 0.01).

[0405] The classification performance of microbiome signatures was ranked based on AUC values ​​in 11 datasets. Figure 35A All rank values ​​assigned to each set of signature microbiome are plotted. Figure 35B The significance of within-group comparisons is shown. Figure 35C The sum of the ranking values ​​for each group of microbiome features is shown. Kruskal-Wallis test followed by Dunn's post hoc test was performed for analysis (#BH adjusted P < 0.1, *BH adjusted P < 0.05, **BH adjusted P < 0.01). Those results confirmed that the microbiome features obtained from the combined core pool had the best performance in classifying healthy subjects from patients in different datasets.

[0406] Example 4 - A universal random forest classification model based on 284 core genomes from two competing bacterial communities in a seesaw network.

[0407] We used 25 metagenomic datasets from case-control studies of 15 different diseases (type 2 diabetes (T2D), hypertension (HT), schizophrenia (SCZ), atherosclerotic cardiovascular disease (ACVD), liver cirrhosis (LC), inflammatory bowel disease (IBD), colorectal cancer (CRC), ankylosing spondylitis (AS), Parkinson's disease (PD), multiple sclerosis (MS), Gaucher disease type II (GDII), COVID-19 (COVID-19), Behçet's disease (BD), autism spectrum disorder (ASD), and pancreatic cancer (PC)) to train random forest classification models for each dataset using the abundance of the same 284 core gene sets present in our seesaw network of two competing microbiota (basal and pathogenic). These models enabled us to distinguish cases from controls in the 25 metagenomic datasets using leave-one-out cross-validation, with most having an AUC above 0.7.

[0408] To classify cases and controls, 25 datasets corresponding to 15 different diseases were used. Figure 37 ) were combined with case and control samples, and patients with any disease were considered cases. All samples in the combined dataset were divided into two groups: 80% for training and 20% for testing. 80% of the samples were used as a training set to build a random forest classification model based on the abundance of the 284 combined core gene sets (controls, n = 1285; cases, n = 1424; 10-fold cross validation). 20% of the samples were used as a test set to obtain the probability score of the random forest classification model based on the abundance of the 284 combined core gene sets (controls, n = 319; cases, n = 356).

[0409] like Figure 38A1 As shown, the training set produced an AUC of 0.74 for classifying cases from controls. The optimal cutoff value was 0.5028, the specificity value was 0.7275, and the sensitivity value was 0.6374. Figure 38B1 As shown, the test set produced an AUC of 0.76 for classifying cases from controls. The optimal cutoff value was 0.531, the specificity value was 0.6489, and the sensitivity value was 0.7492. The model produced a significantly higher probability score for cases than controls, which was not observed in the training set ( Figure 38A2 、 Figure 38A3 ) and the test set ( Figure 38B2 、 Figure 38B3 ). Therefore, the identified microbial genomes described herein can be used to train a general model that distinguishes disease from controls.

[0410] The success of our random forest model based on 284 core genomes from two competing bacterial communities in a seesaw network suggests that biological signals associated with these genomes can be robustly detected despite variations introduced by various confounding factors, ranging from biological to technical. Further refinement and testing of our general model will make a significant contribution to translational metagenomics.

[0411] Example 5 - Repeated training of a universal random forest classification model based on 284 core genomes of two competing bacterial communities in a seesaw network.

[0412] We used 25 metagenomic datasets from case-control studies covering 15 different diseases to construct random forest classification models using a randomly selected number of genomes from the 284 core genomes.

[0413] In brief, multiple random forest classifiers were trained on microbiome datasets obtained from patients and healthy controls in at least one study each of type 2 diabetes (T2D), atherosclerotic cardiovascular disease (ACVD), liver cirrhosis (LC), ankylosing spondylitis (AS), Parkinson's disease (PD), schizophrenia (SCZ), colorectal cancer (CRC), inflammatory bowel disease (IBD), and hypertension. Specifically, the datasets were randomly split into 80% for training the RF model and 20% for testing. For each data set, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270 and 280 groups of genomes were randomly selected from the 284 genomes identified in Table 2 to train 10 classifiers (a total of 290 models for each data set). The average AUC of the ROC curve of each group of x randomly selected genomes is determined and plotted in Figure 39. As shown in Figure 39, the number of genomes required for the clinical model that fully meets the disease state is less than all 284 genomes. In fact, in most (if not all) cases, the model trained using only 15-20 randomly selected genomes is sufficient to meet clinical use (e.g., AUC is 0.65 or higher).

[0414] Cited References and Alternative Examples

[0415] All references cited herein are incorporated by reference in their entirety and for all purposes to the ...

Claims

1. A method for identifying a group of intestinal microorganisms, the method comprising: At a computer system having one or more processors and a memory storing one or more programs executed by the one or more processors: A) for each respective subject in a first plurality of subjects having a first state of a biological characteristic, obtaining in electronic form a corresponding plurality of genome abundance values, the corresponding plurality of genome abundance values ​​comprising, for each respective gut microbe in a plurality of gut microbes, a corresponding value of the abundance of a genome of the respective gut microbe in a biological sample from the gut of the respective subject; B) for each corresponding subject in a second plurality of subjects having a second state of the biological characteristic, obtaining in electronic form a corresponding plurality of genome abundance values, the corresponding plurality of genome abundance values ​​comprising, for each corresponding intestinal microorganism in the plurality of intestinal microorganisms, a corresponding value of the abundance of a genome of the corresponding intestinal microorganism in a biological sample from the intestinal tract of the corresponding subject; C) calculating a first plurality of similarity metrics from a corresponding plurality of genomic abundance values ​​across said first plurality of subjects, wherein: The first plurality of similarity measures includes a first corresponding similarity measure for each unique pair of gut microbes in the plurality of gut microbes, and The first corresponding similarity metric quantifies the similarity between: (i) a corresponding first vector formed by corresponding genomic abundance values ​​of a first microbe in the unique gut microbe pairs across the first plurality of subjects; and (ii) a corresponding second vector formed by corresponding genomic abundance values ​​of a second microbe in the unique gut microbe pairs across the first plurality of subjects; D) calculating a second plurality of similarity metrics using the corresponding genomic abundance values ​​of the second plurality of subjects, wherein: The second plurality of similarity measures comprises a second corresponding similarity measure for each unique pair of gut microbes in the plurality of gut microbes, and The second corresponding similarity metric quantifies the similarity between: (i) a corresponding second vector formed by corresponding genomic abundance values ​​of a first microbe in the unique gut microbe pairs across the second plurality of subjects; and (ii) a corresponding second vector formed by corresponding genomic abundance values ​​of a second microbe in the unique gut microbe pairs across the second plurality of subjects; E) determining a set of unique gut microbe pairs from the plurality of gut microbes based on the first plurality of similarity metrics and the second plurality of similarity metrics, wherein for each respective unique gut microbe pair in the set of unique gut microbe pairs: The first corresponding similarity measure and the second corresponding similarity measure both indicate a statistically significant positive correlation between the abundance of the first gut microbe and the abundance of the second gut microbe in the corresponding unique gut microbe pair, or The first corresponding similarity measure and the second corresponding similarity measure both indicate a statistically significant negative correlation between the abundance of a first gut microbe and the abundance of a second gut microbe in the respective unique gut microbe pair; and F) identifying a set of gut microorganisms comprising the corresponding gut microorganisms represented in the set of unique gut microorganism pairs.

2. The method according to claim 1, wherein: The obtaining A) comprises: (i) for each respective subject in the first plurality of subjects, obtaining in electronic form a corresponding first plurality of at least 100,000 nucleic acid sequences of genomic DNA from a corresponding biological sample from the intestine of the respective subject, and (ii) for each respective subject in the first plurality of subjects, determining a corresponding genomic abundance value for each respective gut microbe in the plurality of gut microbes from the corresponding first plurality of at least 100,000 nucleic acid sequences; and The obtaining B) comprises: (i) for each respective subject in the second plurality of subjects, obtaining in electronic form a corresponding second plurality of at least 100,000 nucleic acid sequences of genomic DNA from a corresponding biological sample from the intestine of the respective subject, and (ii) for each respective subject in the second plurality of subjects, determining a corresponding genomic abundance value for each respective gut microbe in the plurality of gut microbes from the corresponding second plurality of at least 100,000 nucleic acid sequences.

3. The method according to claim 2, wherein: The determining A)(ii) comprises, for each respective subject in the first plurality of subjects: assembling a corresponding first plurality of gut microbial genomes from said corresponding first plurality of at least 100,000 nucleic acid sequences by metagenomic de novo sequence assembly, and For each corresponding gut microbial genome in the corresponding first plurality of gut microbial genomes, calculating a corresponding genome abundance of the corresponding gut microbial genome; and The determining B)(ii) comprises, for each respective subject in the second plurality of subjects: assembling a corresponding second plurality of gut microbial genomes from the corresponding second plurality of at least 100,000 nucleic acid sequences by de novo metagenomic sequence assembly, and For each corresponding gut microbial genome in the corresponding second plurality of gut microbial genomes, a corresponding genome abundance of the corresponding gut microbial genome is calculated.

4. The method according to claim 2, wherein: The determining A)(ii) comprises, for each respective subject in the first plurality of subjects: assigning each corresponding nucleic acid sequence in the corresponding first plurality of at least 100,000 sequences to a corresponding gut microorganism in the plurality of gut microorganisms, thereby generating, for each corresponding gut microorganism in the plurality of gut microorganisms, a corresponding count of the corresponding nucleic acid sequence in the corresponding first plurality of nucleic acid sequences assigned to the corresponding gut microorganism, and For each respective gut microbe in the plurality of gut microbes, determining the respective genomic abundance value for the respective gut microbe based on the respective counts of respective nucleic acid sequences assigned to the respective gut microbe; and The determining B)(ii) comprises, for each respective subject in the first plurality of subjects: assigning each corresponding nucleic acid sequence in the corresponding second plurality of at least 100,000 sequences to a corresponding gut microorganism in the plurality of gut microorganisms, thereby generating, for each corresponding gut microorganism in the plurality of gut microorganisms, a corresponding count of the corresponding nucleic acid sequence in the corresponding second plurality of nucleic acid sequences assigned to the corresponding gut microorganism, and For each respective gut microorganism of the plurality of gut microorganisms, the respective genomic abundance value of the respective gut microorganism is determined based on the respective count of the respective nucleic acid sequence assigned to the respective gut microorganism.

5. The method according to any one of claims 2 to 4, further comprising: for each corresponding subject in the first plurality of subjects, sequencing genomic DNA from the corresponding biological sample from the intestine of the corresponding subject, thereby obtaining the corresponding first plurality of at least 100,000 nucleic acid sequences; as well as For each corresponding subject in the second plurality of subjects, genomic DNA from the corresponding biological sample from the intestine of the corresponding subject is sequenced to obtain the corresponding second plurality of at least 100,000 nucleic acid sequences.

6. The method according to any one of claims 1 to 5, wherein: the first state of the biological characteristic is the absence of a disease or condition, and the second state of the biological characteristic is the presence of the disease or condition; The first state of the biological characteristic is a first severity of a disease or condition, and the second state of the biological characteristic is a second severity of the disease or condition: the first state of the bio-signature is an untreated disease or condition, and the second state of the bio-signature is a treated disease or condition; The first state of the bio-signature is a disease or condition treated with a first therapy, and the second state of the bio-signature is a disease or condition treated with a second therapy; The first state of the biological signature is a first level of a nutrient in the diet, and the second state of the biological signature is a second level of a nutrient in the diet; or The first state of the biometric is a first age, and the second state of the biometric is a second age.

7. The method of any one of claims 1 to 6, wherein the plurality of intestinal microorganisms comprises at least 20 intestinal microorganisms selected from Table 1, Table 2, or Figures 42A to 42XX.

8. The method according to any one of claims 1 to 7, wherein: for each respective subject in the first plurality of subjects, the biological sample from the intestinal tract of the respective subject is a stool sample; as well as For each respective subject in the second plurality of subjects, the biological sample from the intestine of the respective subject is a stool sample.

9. The method according to any one of claims 1 to 8, wherein both the first corresponding similarity measure and the second similarity measure are Pearson correlation coefficients, intraclass correlation coefficients or rank correlation coefficients.

10. The method of any one of claims 1 to 9, wherein a statistically significant positive correlation has a P value of less than 0.

001.

11. The method according to any one of claims 1 to 10, wherein the set of gut microbes comprises all corresponding gut microbes represented in the set of unique gut microbe pairs.

12. The method according to any one of claims 1 to 10, wherein said identifying (F) comprises: The corresponding gut microbes represented in the set of unique gut microbe pairs are clustered into one or more networks, wherein each corresponding connected network comprises a corresponding plurality of nodes, and a corresponding set of one or more edges, wherein: Each corresponding node in the corresponding plurality of nodes represents a unique gut microbe represented in the set of unique gut microbe pairs, Each corresponding edge in the corresponding set of one or more edges connects two nodes representing a corresponding unique gut microbe pair in the set of unique gut microbe pairs, and each respective node in the corresponding plurality of nodes is connected to at least one other respective node in the plurality of nodes via a respective edge in the corresponding set of one or more edges; and A corresponding network of the one or more networks containing the most nodes is identified, thereby identifying a group of intestinal microorganisms represented by the corresponding plurality of nodes in the corresponding network.

13. The method of any one of claims 1 to 12, wherein the panel of intestinal microorganisms comprises at least 20 intestinal microorganisms selected from Table 1, Table 2, or Figures 42A to 42XX.

14. A method for training a model for assessing human health status, the method comprising: At a computer system having one or more processors and a memory storing one or more programs executed by the one or more processors: A) for each respective training subject in a plurality of training subjects, obtaining in electronic form: (i) a plurality of corresponding genome abundance values, the plurality of corresponding genome abundance values ​​comprising, for each corresponding intestinal microorganism in a plurality of intestinal microorganisms, a corresponding value of the abundance of the genome of the corresponding intestinal microorganism in the corresponding biological sample from the intestinal tract of the corresponding training subject, and (ii) the corresponding state of the biological characteristic of the corresponding training subject; B) for each respective training subject in the plurality of training subjects, inputting information about the respective training subject into a model comprising a plurality of parameters, wherein the model applies the plurality of parameters to the information by at least 10,000 calculations to obtain a corresponding output from the model for the respective training subject, wherein: the corresponding output comprises an indication of the corresponding state of the biological characteristic of the corresponding training subject, and The information about the corresponding training subject includes the corresponding genome abundance value of each corresponding intestinal microorganism in the plurality of intestinal microorganisms, and The plurality of intestinal microorganisms are selected from Table 1, Table 2, or Figures 42A to 42XX; and C) adjusting the plurality of parameters based on, for each respective training subject in the first plurality of training subjects, one or more differences between (i) the corresponding output from the model and (ii) the corresponding state of the biological characteristic of the respective training subject.

15. The method of claim 14, wherein the acquiring A) comprises, for each respective training subject in the plurality of training subjects: (i) obtaining in electronic form a corresponding plurality of at least 100,000 nucleic acid sequences of genomic DNA from a corresponding biological sample of the intestine of said corresponding training subject; and (ii) for each respective gut microorganism in the plurality of gut microorganisms, determining the corresponding value of the abundance of the genome of the respective gut microorganism from the corresponding first plurality of at least 100,000 nucleic acid sequences.

16. The method of claim 15, wherein the determining A)(ii) comprises, for each respective training subject in the plurality of training subjects: Assembling in electronic form a corresponding plurality of intestinal microbial genomes from the corresponding plurality of at least 100,000 nucleic acid sequences by de novo metagenomic sequence assembly, and For each corresponding gut microorganism in the plurality of gut microorganisms, the corresponding value of the abundance of the genome of the corresponding gut microorganism is calculated based on the prevalence of the corresponding nucleic acid sequence in the plurality of at least 100,000 nucleic acid sequences, and the corresponding nucleic acid sequence is used to assemble a corresponding gut microorganism genome corresponding to the corresponding gut microorganism in the plurality of gut microorganism genomes.

17. The method of claim 15, wherein the determining A)(ii) comprises, for each respective subject in the plurality of training subjects: assigning each corresponding nucleic acid sequence in the corresponding plurality of at least 100,000 sequences to a corresponding gut microorganism in the plurality of gut microorganisms, thereby generating, for each corresponding gut microorganism in the plurality of gut microorganisms, a corresponding count of the corresponding nucleic acid sequence in the corresponding plurality of nucleic acid sequences assigned to the corresponding gut microorganism, and For each respective gut microorganism of the plurality of gut microorganisms, the respective genomic abundance value of the respective gut microorganism is determined based on the respective count of the respective nucleic acid sequence assigned to the respective gut microorganism.

18. The method according to any one of claims 15 to 17, further comprising sequencing the genomic DNA of the corresponding biological sample from the intestine of each corresponding subject among the multiple training subjects, thereby obtaining the corresponding multiple of at least 100,000 nucleic acid sequences.

19. The method of any one of claims 14 to 18, wherein the plurality of intestinal microorganisms comprises at least 20 intestinal microorganisms selected from Table 1, Table 2, or Figures 42A to 42XX.

20. The method of any one of claims 14 to 19, wherein the plurality of intestinal microorganisms comprises at least 20 microorganisms selected from those microorganisms having a connectivity of at least 2 in Table 1, Table 2, or Figures 42A to 42XX.

21. The method according to any one of claims 14 to 20, wherein for each respective subject of the plurality of training subjects, the biological sample from the intestinal tract of the respective subject is a stool sample from the respective training subject.

22. The method of any one of claims 14 to 21, wherein the biological characteristic is a disease or condition, a therapy administered to the subject, or the subject's diet.

23. The method of claim 22, wherein the disease or condition is selected from the group consisting of type 2 diabetes, hypertension, schizophrenia, atherosclerotic cardiovascular disease (ACVD), liver cirrhosis (LC), inflammatory bowel disease (IBD), colorectal cancer (CRC), ankylosing spondylitis (AS) and Parkinson's disease (PD).

24. The method of claim 22, wherein the disease or condition is cancer.

25. The method of any one of claims 14 to 24, wherein the indication of the corresponding state of the biological characteristic is a categorical output of a respective state among a plurality of possible states of the biological characteristic.

26. The method of any one of claims 14 to 24, wherein the indication of the corresponding state of the biological feature is a probabilistic output for the corresponding state of the biological feature.

27. The method according to any one of claims 14 to 26, wherein the model is a neural network algorithm, a support vector machine algorithm, a naive Bayes algorithm, a nearest neighbor algorithm, a boosted tree algorithm, a random forest algorithm, a convolutional neural network algorithm, a decision tree algorithm, a regression algorithm or a clustering algorithm.

28. The method of any one of claims 14 to 27, wherein the plurality of parameters is at least 1000, at least 10,000, at least 15,000, at least 50,000, at least 100,000, at least 250,000, at least 500,000, or at least 1,000,000 parameters.

29. The method of any one of claims 14 to 28, wherein the model applies the plurality of parameters to the information by calculating at least 25,000 times, at least 50,000 times, at least 100,000 times, at least 250,000 times, at least 500,000 times, or at least 1,000,000 times to obtain a corresponding output from the model for the corresponding training subject.

30. A method for assessing the health status of a subject, the method comprising: At a computer system having one or more processors and a memory storing one or more programs executed by the one or more processors: A) obtaining in electronic form a plurality of genomic abundance values, the plurality of genomic abundance values ​​comprising, for each corresponding intestinal microorganism in a plurality of at least 20 intestinal microorganisms selected from Table 1, Table 2, or Figures 42A to 42XX, corresponding abundance values ​​of genomes of corresponding intestinal bacterial species in the plurality of at least 20 intestinal microorganisms in a biological sample from the subject; as well as B) inputting the multiple genomic abundance values ​​into a model comprising multiple parameters, wherein the model applies the multiple parameters to the multiple genomic abundance values ​​through at least 10,000 calculations to generate an indication of the health status of the subject {e.g., a category output or a probability output} as an output from the model.

31. The method according to claim 30, wherein the obtaining A) comprises: (i) obtaining in electronic form a plurality of at least 100,000 nucleic acid sequences of genomic DNA from said biological sample from the intestine of said subject {including support for the minimum number, maximum number and range of NA sequences in the specification}; and (ii) for each respective gut microorganism in the plurality of gut microorganisms, determining the corresponding value of the abundance of the genome of the respective gut microorganism from the plurality of at least 100,000 nucleic acid sequences.

32. The method of claim 31 , wherein the determining A)(ii) comprises: Assembling a corresponding plurality of intestinal microbial genomes in electronic form from the plurality of at least 100,000 nucleic acid sequences by de novo metagenomic sequence assembly, and For each corresponding gut microorganism in the plurality of gut microorganisms, the corresponding value of the abundance of the genome of the corresponding gut microorganism is calculated based on the prevalence of the corresponding nucleic acid sequence in the plurality of at least 100,000 nucleic acid sequences, and the corresponding nucleic acid sequence is used to assemble a corresponding gut microorganism genome corresponding to the corresponding gut microorganism in the plurality of gut microorganism genomes.

33. The method of claim 31 , wherein the determining A)(ii) comprises: assigning each corresponding nucleic acid sequence in the plurality of at least 100,000 sequences to a corresponding gut microbe in the plurality of gut microbes, thereby generating, for each corresponding gut microbe in the plurality of gut microbes, a corresponding count of the corresponding nucleic acid sequence in the plurality of nucleic acid sequences assigned to the corresponding gut microbe, and For each respective gut microorganism of the plurality of gut microorganisms, the respective genomic abundance value of the respective gut microorganism is determined based on the respective count of the respective nucleic acid sequence assigned to the respective gut microorganism.

34. The method according to any one of claims 31 to 33, further comprising sequencing genomic DNA from the biological sample from the intestine of the subject, thereby obtaining the plurality of at least 100,000 nucleic acid sequences.

35. The method of any one of claims 30 to 34, wherein the plurality of intestinal microorganisms comprises at least 20 microorganisms selected from those microorganisms having a connectivity of at least 2 in Table 1, Table 2, or Figures 42A to 42XX.

36. The method according to any one of claims 30 to 35, wherein the biological sample from the intestinal tract of the subject is a stool sample.

37. The method of any one of claims 30 to 36, wherein the indication of the subject's health status is an indication of a biological characteristic, wherein the biological characteristic is a disease or condition, a therapy administered to the subject, or the subject's diet.

38. The method of claim 37, wherein the disease or condition is selected from the group consisting of type 2 diabetes, hypertension, schizophrenia, atherosclerotic cardiovascular disease (ACVD), liver cirrhosis (LC), inflammatory bowel disease (IBD), colorectal cancer (CRC), ankylosing spondylitis (AS) and Parkinson's disease (PD).

39. The method of claim 37, wherein the disease or condition is cancer.

40. The method of any one of claims 30 to 39, wherein the indication of the subject's health condition is a categorical output of a respective state among a plurality of possible states of the subject's health condition.

41. The method of any one of claims 30 to 39, wherein the indication of the subject's health condition is a probabilistic output of the corresponding state of the subject's health condition.

42. The method according to any one of claims 30 to 41, wherein the model is a neural network algorithm, a support vector machine algorithm, a naive Bayes algorithm, a nearest neighbor algorithm, a boosted tree algorithm, a random forest algorithm, a convolutional neural network algorithm, a decision tree algorithm, a regression algorithm or a clustering algorithm.

43. The method of any one of claims 30 to 42, wherein the plurality of parameters is at least 1000, at least 10,000, at least 15,000, at least 50,000, at least 100,000, at least 250,000, at least 500,000, or at least 1,000,000 parameters.

44. A method according to any one of claims 30 to 43, wherein the model applies the multiple parameters to the information by calculating at least 25,000 times, at least 50,000 times, at least 100,000 times, at least 250,000 times, at least 500,000 times, or at least 1,000,000 times to obtain a corresponding output from the model for the corresponding training subject.

45. A computer system, the computer system comprising: one or more processors; as well as A non-transitory computer readable medium comprising computer executable instructions which, when executed by the one or more processors, cause the processors to perform the method according to any one of claims 1 to 44.

46. ​​A non-transitory computer-readable storage medium having program code instructions stored thereon, which, when executed by a processor, cause the processor to perform the method of any one of claims 1 to 44.

Citation Information

Patent Citations

  • Methods for genome assembly and haplotype phasing

    US10529443B2

  • Microbiome based systems, apparatus and methods for monitoring and controlling industrial processes and systems

    US11028449B2

  • Sample analysis, presence determination of a target sequence

    US11332783B2

  • Absolute quantification of nucleic acids and related methods and systems

    US11427865B2

  • Metagenomic library and natural product discovery platform

    US11495326B2

Cited By

  • Type 2 diabetes mellitus prediction method and system based on microorganism co-abundance characteristics

    CN121790004A

  • A method and system for predicting type 2 diabetes based on microbial co-abundance features

    CN121790004B