Two competing guilds as core microbiome signatures of human diseases

JP2025517828A5Pending Publication Date: 2026-05-01RUTGERS THE STATE UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
RUTGERS THE STATE UNIV
Filing Date
2023-04-25
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Current gene-centered approaches for metagenomic data analysis in microbiome-wide association studies (MWAS) have limitations, as they rely on existing databases and treat individual genes independently, ignoring ecological interactions and leading to spurious correlations.

Method used

A genome-centered MWAS approach using high-quality draft genomes (metagenome-assembled genomes, MAGs) as features for correlation analysis, which considers ecological interactions and organizes genomes into guilds for functional analysis.

Benefits of technology

This approach allows for the identification of ecologically meaningful microbiome signatures associated with human diseases, as demonstrated by the correlation of guild abundances with chronic diseases and the ability to predict disease states using machine learning models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method and system for determining a disease state by obtaining a first plurality of nucleic acid sequences from genomic DNA from a sample derived from the intestine of a subject. From the nucleic acid sequences, a first plurality of genomic abundance values for a first plurality of gut bacteria and a second plurality of genomic abundance values for a second plurality of at least 20 gut bacteria are determined. A model is applied to at least the first plurality of genomic abundance values and the second plurality of genomic abundance values, or one or more combinations thereof, thereby determining the disease state of the subject as the output of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 63 / 334,503, filed on April 25, 2022, which is hereby incorporated by reference in its entirety.

[0002] Brief Description of the Sequence Listing This submission is accompanied by a "Sequence Listing XML" file named ST26_126146_5001_WO.XML, created on April 22, 2023, having a size of 2,491,699 kilobytes and containing sequence numbers 1 - 99534, and is submitted by mail as an XML file on a read - only optical disk (DVD) in accordance with 37 CFR §§ 1.831 - 1.835. 37 CFR § 1.835(a)(1). The Sequence Listing XML is hereby incorporated by reference in its entirety.

Background Art

[0003] Over the course of long - term co - evolution, humans have developed a robust symbiotic relationship with the gut microbiota [6, 7]. The gut microbiota functions as an essential organ to support host homeostasis in metabolism, immunity, development, and behavior, etc. [8]. The attenuation or loss of such health - related functions in the gut microbiota with dysregulated gut - bacterial symbiosis has been identified as a risk factor for many chronic diseases, including type 2 diabetes (T2DM) [9 - 11]. A series of microbiota - wide association studies (MWAS) have attempted to identify microbiota signatures (including features such as genes, pathways, taxa, etc.) associated with disease phenotypes as biomarkers in metagenomic datasets [12 - 15].

[0004] However, current gene-centered approaches for metagenomic data analysis in the majority of MWAS projects have significant limitations. They rely on existing databases for the taxonomic and functional annotation of individual genes and exclude novel genes from downstream analysis [12, 14, 15]. More importantly, such approaches treat individual genes as independent units, ignoring the fact that the function of bacterial genes is constrained by the ecological behavior of their carriers. For example, two competing bacterial strains may encode the same functional gene. However, these two copies of the same gene will contribute differently to the expression of the entire community of the gene's function due to the opposite growth trajectories of their carriers within the intestinal habitat

[16] . By grouping the same functional genes from different bacterial strains into one unit, such as a pathway or a species, these opposite changes will be masked or neutralized, resulting in spurious correlations [5].

Summary of the Invention

[0005] In view of the above background, a genome-centered MWAS was adopted that uses high-quality draft genomes assembled from metagenomic datasets (metagenome-assembled genomes, MAGs) as the most important microbiome features for correlation analysis with the basic components of the intestinal ecosystem and disease phenotypes. MAGs are also not independent microbiome features. They have ecological interactions such as competition or cooperation with each other and are organized into higher-level structures called "guilds" [5]. Each guild is potentially a functional unit of the intestinal ecosystem, and its members may have a wide variety of taxonomic backgrounds but show coexistence behavior. Guilds have been shown to be positively or negatively correlated with disease phenotypes

[17] . Therefore, MAGs and their guild-level aggregation are ecologically meaningful features for identifying microbiome signatures associated with human diseases.

[0006] Dysbiosis of the gut microbiota has been associated with an increased risk of a wide range of human diseases [1, 2]. So far, many attempts have focused on identifying gene-based or taxon-based microbial signatures as disease biomarkers. However, there is still room for debate about such signatures [3, 4], and the fact that gut bacterial strains do not act independently but rather form coherent functional groups (aka "guilds") that interact with each other to affect host health has been overlooked [5]. Accordingly, embodiments may propose to look for strain-level microbiome signatures in the form of robust guilds through which the gut microbiota provides stable health-related functions to the host. Embodiments may show that two competing bacterial guilds are organized as the two ends of a robustly stable seesaw-like network, and their abundances correlate with a wide range of chronic diseases. Of a total of 1,845 metagenome-assembled genomes (MAGs), 141 experienced severe structural changes in the gut microbiota during a 3-month high-fiber intervention and 1-year follow-up of type 2 diabetes (T2DM) patients, while forming two competing guilds considering stable ecological relationships. The 50 genomes of guild 1 contained more genes for plant polysaccharide degradation and butyrate production, while the 91 genomes of guild 2 contained almost all pathogenic or antibiotic resistance gene carriers predicted from the 1,845 MAGs. A random forest regression model showed that the abundance distribution of the 141 genomes was associated with 41 of 43 bioclinical parameters. Using these 141 MAGs as reference genomes, such a seesaw network was not only detectable but also facilitated a machine learning model for predictive classification between cases and controls of nine diseases including T2DM, atherosclerosis, hypertension, cirrhosis, inflammatory bowel disease, colorectal cancer, ankylosing spondylitis, schizophrenia, and Parkinson's disease in 12 independent metagenomic datasets from 1,874 participants across ethnic and geographical groups. The two seesaw network guilds function as a core microbiome, and their balance can be regulated for disease risk management.

[0007] Accordingly, one aspect of the present disclosure provides a method and a system for implementing the disclosed method to determine a disease state among a plurality of disease states of a subject. The method includes obtaining, in a computer system having at least one processor and a memory storing one or more programs for execution by the one or more processors, in electronic form, a first plurality (e.g., at least 100,000) of nucleic acid sequences for a first genomic DNA from a first biological sample derived from the intestine of the subject. The method also includes determining, from the first plurality of nucleic acid sequences, a first plurality of genomic abundance values including, for each respective species of a first plurality (e.g., at least 20) of intestinal bacterial species, a first corresponding abundance value for the genome of each respective species of the first plurality of intestinal bacterial species in the first biological sample, and a second plurality of genomic abundance values including, for each respective species of a second plurality (e.g., at least 20) of intestinal bacterial species, a first corresponding abundance value for the genome of each respective species of the second plurality of intestinal bacterial species in the first biological sample. The method also includes applying, by the at least one processor, a model to at least the first plurality of genomic abundance values and the second plurality of genomic abundance values, or one or more combinations thereof, thereby determining, as an output of the model, the disease state of the subject.

[0008] As disclosed herein, where applicable, any embodiment disclosed herein may be applied to any other aspect.

[0009] Further aspects and advantages of the present disclosure will become readily apparent to those skilled in the art from the following detailed description. In which only exemplary embodiments of the present disclosure are shown and described. As will be understood, the present disclosure is capable of other different embodiments and some of the details thereof are capable of modification in various obvious respects without departing from the present disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature and not as restrictive.

[0010] Accordingly, one aspect of the present invention provides a method for identifying a set of gut microbiota in a computer system having one or more processors and a memory storing one or more programs for execution by the one or more processors.

[0011] In some embodiments, the method includes obtaining, in electronic form, for each respective subject of a first plurality of subjects having a first state of a biological characteristic, a corresponding plurality of genomic abundance values including, for each respective gut microbiota of a plurality of gut microbiota, a corresponding value for the abundance of the genome of the respective gut microbiota in a biological sample derived from the gut of the respective subject.

[0012] In some such embodiments, for each respective subject of the first plurality of subjects, the biological sample derived from the gut of the respective subject is a fecal sample.

[0013] In some such embodiments, the method includes sequencing genomic DNA from the corresponding biological sample derived from the gut of each respective subject of the first plurality of subjects, thereby obtaining a corresponding first plurality of at least 100,000 nucleic acid sequences.

[0014] In some such embodiments, the method includes obtaining, in electronic form, a corresponding first plurality of at least 100,000 nucleic acid sequences for genomic DNA from the corresponding biological sample derived from the gut of each respective subject of the first plurality of subjects, and determining, for each respective subject of the first plurality of subjects, corresponding genomic abundance values for each respective gut microbiota of a plurality of gut microbiota from the corresponding first plurality of at least 100,000 nucleic acid sequences.

[0015] In some such embodiments, the method comprises, for each respective subject among a first plurality of subjects, assembling a corresponding first plurality of gut microbiome genomes by metagenomic de novo sequence assembly from a corresponding first plurality of at least 100,000 nucleic acid sequences, and for each respective gut microbiome genome among the corresponding first plurality of gut microbiome genomes, calculating a corresponding genomic abundance of the respective gut microbiome genome.

[0016] In some such embodiments, the method comprises, for each respective subject among a first plurality of subjects, assigning each respective nucleic acid sequence among a corresponding first plurality of at least 100,000 sequences to a respective gut microbiome among a plurality of gut microbiomes, thereby generating a corresponding count of each respective nucleic acid sequence among the corresponding first plurality of nucleic acid sequences assigned to the respective gut microbiome for each respective gut microbiome among the plurality of gut microbiomes, and for each respective gut microbiome among the plurality of gut microbiomes, determining a corresponding genomic abundance value for the respective gut microbiome based on the corresponding count of each respective nucleic acid sequence assigned to the respective gut microbiome.

[0017] In some embodiments, the method comprises obtaining, for each respective subject among a second plurality of subjects having a second state of a biological characteristic, in electronic form, a corresponding plurality of genomic abundance values including a corresponding value for the abundance of the genome of each respective gut microbiome in a gut-derived biological sample of the respective subject for each respective gut microbiome among a plurality of gut microbiomes.

[0018] In some such embodiments, for each respective subject among the second plurality of subjects, the gut-derived biological sample of the respective subject is a fecal sample.

[0019] In some such embodiments, the method includes, for each respective subject of the second plurality of subjects, sequencing genomic DNA from a corresponding biological sample derived from the intestine of the respective subject, thereby obtaining a corresponding second plurality of at least 100,000 nucleic acid sequences.

[0020] In some such embodiments, the method includes, in electronic form, for each respective subject of the second plurality of subjects, obtaining a corresponding second plurality of at least 100,000 nucleic acid sequences for genomic DNA from a corresponding biological sample derived from the intestine of the respective subject, and for each respective subject of the second plurality of subjects, determining a corresponding genomic abundance value for each respective gut microbe among the plurality of gut microbes from the corresponding second plurality of at least 100,000 nucleic acid sequences.

[0021] In some such embodiments, the method includes, for each respective subject of the first plurality of subjects, assembling a corresponding second plurality of gut microbe genomes by metagenomic de novo sequence assembly from the corresponding second plurality of at least 100,000 nucleic acid sequences, and for each respective gut microbe genome among the corresponding second plurality of gut microbe genomes, calculating the corresponding genomic abundance of the respective gut microbe genome.

[0022] In some such embodiments, the method includes, for each respective subject of the first plurality of subjects, assigning each respective nucleic acid sequence among the corresponding second plurality of at least 100,000 sequences to a respective gut microbe among the plurality of gut microbes, thereby generating a corresponding count of each respective nucleic acid sequence among the corresponding second plurality of nucleic acid sequences assigned to the respective gut microbe for each respective gut microbe among the plurality of gut microbes, and for each respective gut microbe among the plurality of gut microbes, determining a corresponding genomic abundance value for the respective gut microbe based on the corresponding count of the respective nucleic acid sequences assigned to the respective gut microbe.

[0023] In some such embodiments, the first state of the biological characteristic is the absence of a disease or disorder, the second state of the biological characteristic is the presence of a disease or disorder, or the first state of the biological characteristic is the first severity of a disease or disorder, the second state of the biological characteristic is the second severity of the disease or disorder, or the first state of the biological characteristic is an untreated disease or disorder, the second state of the biological characteristic is a treated disease or disorder, or the first state of the biological characteristic is a disease or disorder treated with a first therapy, the second state of the biological characteristic is a disease or disorder treated with a second therapy, or the first state of the biological characteristic is a first level of nutrients in the diet, the second state of the biological characteristic is a second level of nutrients in the diet, or the first state of the biological characteristic is a first age, the second state of the biological characteristic is a second age.

[0024] In some such embodiments, the plurality of gut microbiota includes at least 20 gut microbiota selected from Table 1, Table 2, or FIGS. 42A - 42XX.

[0025] In some embodiments, the method comprises calculating a first plurality of similarity metrics from corresponding plurality of genomic abundance values across a first plurality of subjects, the first plurality of similarity metrics including a first corresponding similarity metric for each unique pair of gut microbiota among the plurality of gut microbiota, the first corresponding similarity metric quantifying the similarity between (i) a corresponding first vector formed by the corresponding genomic abundance values of a first microorganism in a unique pair of gut microbiota across the first plurality of subjects and (ii) a corresponding second vector formed by the corresponding genomic abundance values of a second microorganism in the unique pair of gut microbiota across the first plurality of subjects.

[0026] In some embodiments, the method comprises calculating a second plurality of similarity metrics using corresponding genomic abundance values for a second plurality of subjects, the second plurality of similarity metrics including a second corresponding similarity metric for each unique pair of gut microbiota among the plurality of gut microbiota, the second corresponding similarity metric quantifying a similarity between (i) a corresponding second vector formed by the corresponding genomic abundance value of a first microbiota in a unique pair of gut microbiota across the second plurality of subjects and (ii) a corresponding second vector formed by the corresponding genomic abundance value of a second microbiota in the unique pair of gut microbiota across the second plurality of subjects.

[0027] In some embodiments, the method comprises determining a set of unique pairs of gut microbiota among the plurality of gut microbiota based on the first plurality of similarity metrics and the second plurality of similarity metrics, wherein both the first corresponding similarity metric and the second corresponding similarity metric exhibit a statistically significant positive correlation between the abundance of the first gut microbiota and the abundance of the second gut microbiota in each unique pair of gut microbiota, or both the first corresponding similarity metric and the second corresponding similarity metric exhibit a statistically significant negative correlation between the abundance of the first gut microbiota and the abundance of the second gut microbiota in each unique pair of gut microbiota.

[0028] In some such embodiments, both the first corresponding similarity metric and the second similarity metric are a Pearson correlation coefficient, an intraclass correlation coefficient, or a rank correlation coefficient.

[0029] In some such embodiments, the statistically significant positive correlation has a P-value of less than 0.001.

[0030] In some embodiments, the method comprises identifying a set of gut microbiota including each gut microbiota represented by the set of unique pairs of gut microbiota.

[0031] In some such embodiments, the method includes clustering each gut microbe, represented by a unique pair set of gut microbes, into one of a plurality of networks. Each respective connected network includes a corresponding plurality of nodes and a corresponding set of one or more edges.

[0032] In some such embodiments, each respective node among the corresponding plurality of nodes represents a unique gut microbe represented by a unique pair set of gut microbes.

[0033] In some such embodiments, each respective edge in the corresponding set of one or more edges connects two nodes that represent unique pairs of gut microbes in the unique pair set of gut microbes.

[0034] In some such embodiments, each respective node among the corresponding plurality of nodes is connected to at least one other respective node among the plurality of nodes via each respective edge in the corresponding set of one or more edges.

[0035] In some such embodiments, the method includes identifying each network in one or more networks that contains the most nodes, thereby identifying the set of gut microbes represented by the corresponding plurality of nodes in each network.

[0036] In some such embodiments, the set of gut microbes includes all respective gut microbes represented by a unique pair set of gut microbes.

[0037] In some such embodiments, the set of gut microbes includes at least 20 gut microbes selected from Table 1, Table 2, or FIGS. 42A - 42XX.

[0038] Another aspect of the present disclosure provides a method for training a model for evaluating human health in a computer system having one or more processors and a memory storing one or more programs for execution by the one or more processors.

[0039] In some embodiments, the method comprises, for each respective training subject among a plurality of training subjects in electronic form, (i) a corresponding plurality of genomic abundance values including corresponding values of the genomes of each respective gut microorganism among a plurality of gut microorganisms with respect to the abundance of the genomes of each respective gut microorganism in a corresponding biological sample derived from the gut of the respective training subject, and (ii) the corresponding state of the biological characteristics of each respective training subject.

[0040] In some such embodiments, the method comprises, for each respective subject among a plurality of training subjects, sequencing genomic DNA from a corresponding biological sample derived from the gut of the respective training subject, thereby obtaining a corresponding plurality of at least 100,000 nucleic acid sequences.

[0041] In some such embodiments, the biological sample derived from the gut of each respective training subject is a fecal sample from the respective training subject.

[0042] In some such embodiments, the plurality of gut microorganisms includes at least 20 gut microorganisms selected from Table 1, Table 2, or FIGS. 42A - 42XX.

[0043] In some such embodiments, the plurality of gut microorganisms includes at least 20 microorganisms selected from those microorganisms in Table 1, Table 2, or FIGS. 42A - 42XX having at least 2 binding properties.

[0044] In some such embodiments, the method comprises, for each respective training subject among a plurality of training subjects, electronically obtaining, from the corresponding biological sample derived from the intestine of each respective training subject, a corresponding plurality of at least 100,000 nucleic acid sequences for genomic DNA, and for each respective gut microorganism among a plurality of gut microorganisms, determining, from the corresponding first plurality of at least 100,000 nucleic acid sequences, a corresponding value for the abundance of the genome of each respective gut microorganism.

[0045] In some such embodiments, the method comprises, for each respective training subject among a plurality of training subjects, electronically assembling, by metagenomic de novo sequence assembly from the corresponding plurality of at least 100,000 nucleic acid sequences, the corresponding plurality of gut microorganism genomes, and for each respective gut microorganism among a plurality of gut microorganisms, calculating, based on the morbidity rate of each respective nucleic acid sequence among the corresponding plurality of at least 100,000 nucleic acid sequences used to assemble each respective gut microorganism genome among the plurality of gut microorganism genomes corresponding to each respective gut microorganism, a corresponding value for the abundance of the genome of each respective gut microorganism.

[0046] In some such embodiments, the method comprises, for each respective subject among a plurality of training subjects, assigning each respective nucleic acid sequence among the corresponding plurality of at least 100,000 sequences to each respective gut microorganism among a plurality of gut microorganisms, thereby generating, for each respective gut microorganism among a plurality of gut microorganisms, a corresponding count of each respective nucleic acid sequence among the corresponding plurality of nucleic acid sequences assigned to each respective gut microorganism, and for each respective gut microorganism among a plurality of gut microorganisms, determining, based on the corresponding count of each respective nucleic acid sequence assigned to each respective gut microorganism, a corresponding genomic abundance value for each respective gut microorganism.

[0047] In some such embodiments, the biological characteristic is a disease or disorder, a therapy administered to the subject, or the subject's diet.

[0048] In some such embodiments, the disease or disorder is selected from the group consisting of type 2 diabetes, hypertension, schizophrenia, atherosclerotic cardiovascular disease (ACVD), liver cirrhosis (LC), inflammatory bowel disease (IBD), colorectal cancer (CRC), ankylosing spondylitis (AS), and Parkinson's disease (PD).

[0049] In some embodiments, the disease or disorder is cancer.

[0050] In some embodiments, the method includes, for each respective training subject among a plurality of training subjects, inputting information regarding each respective training subject into a model that includes a plurality of parameters. The model applies the plurality of parameters to the information via at least 10,000 calculations in order to obtain a corresponding output for each respective training subject from the model. The corresponding output includes an indicator of a corresponding state of a biological characteristic of each respective training subject, and the information regarding each respective training subject includes corresponding genomic abundance values for each respective gut microorganism among a plurality of gut microorganisms. The plurality of gut microorganisms is selected from Table 1, Table 2, or FIGS. 42A - 42XX.

[0051] In some such embodiments, the indicator of the corresponding state of the biological characteristic is a class output of each state among a plurality of possible states of the biological characteristic.

[0052] In some such embodiments, the indicator of the corresponding state of the biological characteristic is a probability output of the corresponding state of the biological characteristic.

[0053] In some such embodiments, the model is a neural network algorithm, a support vector machine algorithm, a naive Bayes algorithm, a nearest neighbor algorithm, a boosting tree algorithm, a random forest algorithm, a convolutional neural network algorithm, a decision tree algorithm, a regression algorithm, or a clustering algorithm.

[0054] In some such embodiments, the plurality of parameters are at least 1,000, at least 10,000, at least 15,000, at least 50,000, at least 100,000, at least 250,000, at least 500,000, or at least 1,000,000 parameters.

[0055] In some such embodiments, the model applies the plurality of parameters to the information via at least 25,000, at least 50,000, at least 100,000, at least 250,000, at least 500,000, or at least 1,000,000 calculations to obtain corresponding outputs for each training subject from the model.

[0056] In some embodiments, the method includes adjusting a plurality of parameters based on one or more differences between (i) a corresponding output from the model and (ii) a corresponding state of the biological characteristics of each respective training subject among a first plurality of training subjects.

[0057] Another aspect of the present disclosure provides a method for evaluating the health of a subject in a computer system having one or more processors and a memory storing one or more programs for execution by the one or more processors.

[0058] In some embodiments, the method includes obtaining, for each respective gut microbe among a plurality of at least 20 gut microbes selected from Table 1, Table 2, or FIGS. 42A - 42XX in electronic form, a plurality of genomic abundance values including corresponding abundance values for the genome of each species of gut bacteria among the plurality of at least 20 gut microbes in a biological sample from the subject.

[0059] In some such embodiments, the method includes sequencing genomic DNA from a biological sample derived from the gut of the subject, thereby obtaining a plurality of at least 100,000 nucleic acid sequences.

[0060] In some such embodiments, the biological sample derived from the intestine of the subject is a fecal sample.

[0061] In some such embodiments, the plurality of gut microbiota comprises at least 20 microorganisms selected from those microorganisms in Table 1, Table 2, or Figures 42A - 42XX that have at least 2 binding properties.

[0062] In some such embodiments, the method comprises, in electronic form, obtaining a plurality of at least 100,000 nucleic acid sequences for genomic DNA from a biological sample derived from the intestine of the subject, and determining, for each respective gut microbiota among the plurality of gut microbiota, a corresponding value for the abundance of the genome of the respective gut microbiota relative to the plurality of at least 100,000 nucleic acid sequences.

[0063] In some such embodiments, the method comprises, in electronic form, assembling a corresponding plurality of gut microbiota genomes by metagenomic de novo sequence assembly from the plurality of at least 100,000 nucleic acid sequences, and calculating, for each respective gut microbiota among the plurality of gut microbiota, a corresponding value for the abundance of the genome of the respective gut microbiota based on the prevalence of each nucleic acid sequence among the plurality of at least 100,000 nucleic acid sequences used to assemble the respective gut microbiota genome corresponding to the respective gut microbiota in the plurality of gut microbiota genomes.

[0064] In some such embodiments, the method assigns each respective nucleic acid sequence among a plurality of at least 100,000 sequences to a respective gut microorganism among a plurality of gut microorganisms, thereby generating, for each respective gut microorganism among the plurality of gut microorganisms, a corresponding count of each respective nucleic acid sequence among the plurality of nucleic acid sequences assigned to the respective gut microorganism, and determining, for each respective gut microorganism among the plurality of gut microorganisms, a corresponding genomic abundance value for the respective gut microorganism based on the corresponding count of each respective nucleic acid sequence assigned to the respective gut microorganism.

[0065] In some embodiments, the method includes inputting a plurality of genomic abundance values into a model that includes a plurality of parameters. The model applies the plurality of parameters to the plurality of genomic abundance values via at least 10,000 calculations to generate, as an output from the model, an indicator of the health of the subject.

[0066] In some such embodiments, the indicator of the health of the subject is an indicator of a biological characteristic. The biological characteristic is a disease or disorder, a therapy administered to the subject, or the diet of the subject.

[0067] In some such embodiments, the disease or disorder is selected from the group consisting of type 2 diabetes, hypertension, schizophrenia, atherosclerotic cardiovascular disease (ACVD), liver cirrhosis (LC), inflammatory bowel disease (IBD), colorectal cancer (CRC), ankylosing spondylitis (AS), and Parkinson's disease (PD).

[0068] In some embodiments, the disease or disorder is cancer.

[0069] In some such embodiments, the indicator of the health of the subject is a class output for each respective state among a plurality of possible states of the health of the subject.

[0070] In some such embodiments, the indicator of the health of the subject is a probability output for the corresponding state of the health of the subject.

[0071] In some such embodiments, the model is a neural network algorithm, a support vector machine algorithm, a naive Bayes algorithm, a nearest neighbor algorithm, a boosting tree algorithm, a random forest algorithm, a convolutional neural network algorithm, a decision tree algorithm, a regression algorithm, or a clustering algorithm.

[0072] In some such embodiments, the plurality of parameters are at least 1000, at least 10,000, at least 15,000, at least 50,000, at least 100,000, at least 250,000, at least 500,000, or at least 1,000,000 parameters.

[0073] In some such embodiments, the model applies the plurality of parameters to the information via at least 25,000, at least 50,000, at least 100,000, at least 250,000, at least 500,000, or at least 1,000,000 calculations to obtain corresponding outputs for each training target from the model.

[0074] Another aspect of the present disclosure provides a computer system. The computer system includes one or more processors and a non-transitory computer-readable medium including computer-executable instructions that, when executed by the one or more processors, cause the processors to perform the methods described herein.

[0075] Another aspect of the present disclosure provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores instructions that, when executed by a computer system, cause the computer system to perform any of the methods described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0076]

Figure 1

Figure 2A

Figure 2B

Figure 2C

Figure 2D

Figure 2E

Figure 2F

Figure 2G

Figure 3A

Figure 3B

Figure 3C

Figure 3D

Figure 4A

Figure 4B

Figure 4C

Figure 5A

Figure 5B

Figure 5C

Figure 5D

Figure 5E

Figure 5F

Figure 5G

Figure 5H

Figure 6A

Figure 6B1

Figure 6B2

Figure 6B3

Figure 7A

Figure 7B

Figure 7C1

Figure 7C2

Figure 8A1

Figure 8A2

Figure 8A3

Figure 8A4

Figure 8B

Figure 9

Figure 10A

Figure 10B

Figure 10C

Figure 10D

Figure 11A

Figure 11B

Figure 11C

Figure 11D

Figure 12A

Figure 12B

Figure 12C

Figure 13

Figure 14A

Figure 14B

Figure 14C

Figure 14D

Figure 14E

Figure 14F

Figure 15

Figure 16A

Figure 16B

Figure 16C

Figure 17A

Figure 17B

Figure 17C

Figure 18A

Figure 18B

Figure 19

Figure 20

Figure 21A

Figure 21B

Figure 22A

Figure 22B

Figure 23

Figure 24

Figure 25

Figure 26A

Figure 26B

Figure 26C

Figure 26D

Figure 26E

Figure 26F

Figure 26G

Figure 26H

Figure 26I

Figure 26J

Figure 26K

Figure 26L

Figure 26M

Figure 27A

Figure 27B

Figure 27C

Figure 27D

Figure 27E

Figure 27F

Figure 27G

Figure 27H

Figure 27I

Figure 27J

Figure 27K

Figure 27L

Figure 27M

Figure 28A

Figure 28B

Figure 28C

Figure 28D

Figure 28E

Figure 28F

Figure 28G

Figure 28H

Figure 29A

Figure 29B

Figure 30A

Figure 30B

Figure 31A

Figure 31B

Figure 31C

Figure 32

Figure 33A

Figure 33B

Figure 33C

Figure 33D

Figure 33E

Figure 33F

Figure 33G

Figure 33H

Figure 33I

Figure 33J

Figure 34A

Figure 34B

Figure 35A

Figure 35B

Figure 35C

Figure 36

Figure 37

Figure 38A1

Figure 38A2

Figure 38A3

Figure 38B1

Figure 38B2

Figure 38B3

Figure 39A

Figure 39B

Figure 40A

Figure 40B

Figure 41A

Figure 41B

Figure 41C

Figure 41D

Figure 41E

Figure 41F

Figure 41G

Figure 41H

Figure 41I

Figure 42A

Figure 42B

Figure 42C

Figure 42D

Figure 42E

Figure 42F

Figure 42G

Figure 42H

Figure 42I

Figure 42J

Figure 42K

Figure 42L

Figure 42M

Figure 42N

Figure 42O

Figure 42P

Figure 42Q

Figure 42R

Figure 42S

Figure 42T

Figure 42U

Figure 42V

Figure 42W

Figure 42X

Figure 42Y

Figure 42Z

Figure 42AA

Figure 42BB

Figure 42CC

Figure 42DD

Figure 42EE

Figure 42FF

Figure 42GG

Figure 42HH

Figure 42II

Figure 42JJ

Figure 42KK

Figure 42LL

Figure 42MM

Figure 42NN

Figure 42OO

Figure 42PP

Figure 42QQ

Figure 42RR

Figure 42SS

Figure 42TT

Figure 42UU

Figure 42VV

Figure 42WW

Figure 42XX

[0077] Like reference numerals refer to corresponding parts throughout the drawings.

Best Mode for Carrying Out the Invention

[0078] The methods and systems described herein facilitate determination of a subject's disease state among a plurality of disease states based on the composition of the subject's microbiome.

[0079] Definitions. The terms used in this disclosure are for the purpose of describing particular embodiments only and are not intended to limit the invention. As used in the description of the invention and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. The term “and / or” as used herein refers to any and all possible combinations of one or more of the associated listed items and is to be understood to be inclusive. The terms “includes,” “comprising,” or any variation thereof as used herein specify the presence of the stated feature, integer, step, operation, element, and / or component, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Further, as long as the terms “including,” “include,” “having,” “has,” “with” or variations thereof are used in either the detailed description or the claims, such terms are intended to be inclusive in the same manner as the term “comprising.”

[0080] As used herein, the term "if" can be interpreted to mean "when", "upon", "in response to a determination that", "in accordance with a determination that", or "as appropriate in the context", depending on the context, that the preceding condition described is true. Similarly, the phrases "when determined" or "when [a described condition or event] is detected" can be interpreted to mean "at the time of determination", "in response to the determination", "at the time of detection of [a described condition or event]", or "when [a described condition or event] is detected", depending on the context.

[0081] Also, terms such as first, second, etc. may be used herein to describe various elements, but it should also be understood that these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first object may be referred to as a second object without departing from the scope of the present invention, and similarly, a second object may be referred to as a first object. The first object and the second object are both objects, but not the same object. The terms "object", "user", and "patient" are used interchangeably herein.

[0082] As used herein, the term "measure of central tendency" refers to the center or representative value of a distribution of values. Non-limiting examples of measures of central tendency include the arithmetic mean, weighted mean, midrange, midhinge, trimmed mean, geometric mean, geometric median, winsorized mean, median, and mode of a distribution of values.

[0083] As used herein, the term "subject" refers to any living or non-living entity, including but not limited to a human (e.g., male human, female human, fetus, pregnant female, child, etc.), non-human mammal, or non-human animal. Any human or non-human animal, including but not limited to mammals, reptiles, birds, amphibians, fish, ungulates, ruminants, genus Bos (e.g., cattle), family Equidae (e.g., horses), goats and sheep (e.g., sheep, goats), family Suidae (e.g., pigs), family Camelidae (e.g., camels, llamas, alpacas), monkeys, apes (e.g., gorillas, chimpanzees), family Ursidae plantigrade carnivores (e.g., bears), poultry, dogs, cats, mice, rats, fish, dolphins, whales, sharks, can function as a subject. In some embodiments, the subject is a male or female of any age (e.g., male, female, or child).

[0084] As used herein, the terms "cancer", "cancer tissue", or "tumor" refer to an abnormal mass of tissue whose growth exceeds and is uncoordinated with the growth of normal tissue, including both solid masses (e.g., such as solid tumors) or fluid masses (e.g., such as blood cancers). Cancers or tumors can be defined as "benign" or "malignant" according to the following characteristics: degree of cell differentiation, including morphology and functionality, growth rate, local invasion, and metastasis. "Benign" tumors are well-differentiated, have a characteristically slower growth than malignant tumors, and may remain localized to the site of origin. Additionally, in some cases, benign tumors do not have the ability to infiltrate, invade, or metastasize to distant sites. "Malignant" tumors are poorly differentiated (anaplastic), are associated with progressive infiltration, invasion, and destruction of surrounding tissues, and characteristically grow rapidly. Furthermore, malignant tumors may have the ability to metastasize to distant sites. Thus, cancer cells are cells found within an abnormal mass of tissue whose growth is uncoordinated with the growth of normal tissue. Thus, a "tumor sample" refers to a biological sample obtained from or derived from a subject's tumor, as described herein.

[0085] Non-limiting examples of cancer types include ovarian cancer, cervical cancer, uveal melanoma, colorectal cancer, chromophobe renal cell carcinoma, liver cancer, endocrine tumors, oropharyngeal cancer, retinoblastoma, biliary tract cancer, adrenal cancer, nerve cancer, neuroblastoma, basal cell carcinoma, brain tumors, breast cancer, non-clear cell renal cell carcinoma, glioblastoma, glioma, kidney cancer, gastrointestinal stromal tumor, medulloblastoma, bladder cancer, stomach cancer, bone cancer, non-small cell lung cancer, thymoma, prostate cancer, clear cell renal cell carcinoma, skin cancer, thyroid cancer, sarcoma, testicular cancer, head and neck cancer (e.g., head and neck squamous cell carcinoma), meningioma, peritoneal cancer, endometrial cancer, pancreatic cancer, mesothelioma, esophageal cancer, small cell lung cancer, Her2-negative breast cancer, ovarian serous cancer, HR+ breast cancer, uterine serous cancer, endometrial cancer of the uterine body, gastroesophageal junction adenocarcinoma, gallbladder cancer, chordoma, and papillary renal cell carcinoma.

[0086] As used herein, the terms "cancer condition" or "cancer symptom" refer to characteristics of the symptoms of a cancer patient, e.g., the diagnostic state, cancer type, cancer location, cancer origin, cancer stage, cancer prognosis, and / or one or more additional characteristics of the cancer (e.g., tumor characteristics such as morphology, heterogeneity, size, etc.). In some embodiments, one or more additional personal characteristics of the subject are used to further describe the subject's cancer condition or cancer symptom, e.g., age, gender, weight, race, personal habits (e.g., smoking, drinking, diet), other relevant health conditions (e.g., hypertension, dry skin, other diseases), current drug therapy, allergies, relevant medical history, current side effects of cancer treatment, and other drug therapies.

[0087] As used herein, the term "genomic abundance value" refers to the absolute or relative amount of the genome of a microorganism in a biological sample from the intestine of a subject. Genomic abundance values can be expressed in different units, including copy number, molar concentration, mass (e.g., normalized to the size of the genome), unique sequence reads (e.g., normalized to the size of the genome), percentage of any of the previous metrics relative to the total amount of the metric across all genomes in the sample, percentage of any of the previous metrics relative to the total amount of the metric across multiple genomes in the sample, etc. In some embodiments, the genomic abundance value is normalized relative to the total genomic abundance in the sample. In some embodiments, the genomic abundance value is normalized relative to the genomic abundance value relative to a control genome in the sample. In some embodiments, the values for multiple genomic abundance values in the sample are standardized, normalized, and / or scaled. Examples of methods for normalizing genomic abundance values are described, for example, in Lin, H., Peddada, S.D., Analysis of microbial compositions: a review of normalization and differential abundance analysis, Biofilms Microbiomes, 6(60)(2020) and Lutz K.C., et al., A Survey of Statistical Methods for Microbiome Data Analysis, Frontiers in Applied Mathematics and Statistics, 8(2022), the contents of which are incorporated herein by reference in their entirety. Methods for measuring genomic abundance values are known in the art. For example, metagenomic sequencing can be used to substantially reconstruct microbial genomes from next-generation sequencing of genomic DNA in biological samples such as biological samples from the intestine of a subject.For a review of metagenomic sequences, see, for example, Quince C, et al., Shotgun metagenomics, from sampling to analysis, Nat Biotechnol, 35(9):833-44(2017), the contents of which are incorporated herein by reference in their entirety. Genome abundance may also be determined by quantification of the copy number of ribosomal genes, such as the 16S rRNA gene. As described in Manzari C., et al., Accurate quantification of bacterial abundance in metagenomic DNAs accounting for variable DNA integrity levels, Microb Genom., 6(10):mgen000417(2020) and Barlow, J.T., et al., A quantitative sequencing framework for absolute abundance measurements of mucosal and lumenal microbial communities, Nat Commun., 11:2590(2020), the contents of which are incorporated herein by reference in their entirety.

[0088] As used herein, the term "relative abundance" refers to the ratio of a first amount of a compound measured in a sample, e.g., the genome of a first microorganism, to a second amount of a compound measured in a second sample. In some embodiments, relative abundance refers to the ratio of the amount of a compound, e.g., the amount of the genome of a first microorganism, to the total amount of compounds in the same sample, e.g., the total amount of microbial genomes or the total amount of a plurality of genomes. In other embodiments, relative abundance refers to the ratio of the amount of a compound in a first sample, e.g., the amount of the genome of a first microorganism, to the amount of a compound in a second sample. For example, it is the ratio of the normalized amount of the genome for a first microorganism in a first sample to the normalized amount of the genome for the first microorganism in a second sample and / or a reference sample.

[0089] As used herein, the terms "sequencing," "sequence determination," and the like refer to any biochemical process that can be used to determine the order of a biomacromolecule such as a nucleic acid or a protein. For example, sequencing data can include all or a portion of the nucleotide bases in a nucleic acid molecule such as an mRNA transcript or a genomic locus.

[0090] As used herein, the terms "sequence read" or "read" refer to a nucleotide sequence that is described herein or generated by any nucleic acid sequencing process known in the art. Reads can be generated from one end of a nucleic acid fragment ("single-end read") or from both ends of a nucleic acid fragment (e.g., paired-end reads, double-end reads). The length of a sequence read is often related to a particular sequencing technology. For example, high-throughput methods provide sequence reads that can vary in size from tens to hundreds of base pairs (bp). In some embodiments, the sequence read is of an average, median, or mean length of about 15 bp to 900 bp (e.g., about 20 bp, about 25 bp, about 30 bp, about 35 bp, about 40 bp, about 45 bp, about 50 bp, about 55 bp, about 60 bp, about 65 bp, about 70 bp, about 75 bp, about 80 bp, about 85 bp, about 90 bp, about 95 bp, about 100 bp, about 110 bp, about 120 bp, about 130, about 140 bp, about 150 bp, about 200 bp, about 250 bp, about 300 bp, about 350 bp, about 400 bp, about 450 bp, or about 500 bp). In some embodiments, the sequence read is of an average, median, or mean length of about 1000 bp, 2000 bp, 5000 bp, 10,000 bp, or 50,000 bp, or more. For example, Nanopore® sequencing can provide sequence reads that can vary in size from tens to hundreds to thousands of base pairs. Illumina® parallel sequencing can provide, for example, sequence reads that do not vary as much, such that most sequence reads can be less than 200 bp. A sequence read (or sequencing read) can refer to the sequence information corresponding to a nucleic acid molecule (e.g., a string of nucleotides). For example, a sequence read can correspond to a string of nucleotides (e.g., about 20 to about 150) derived from a portion of a nucleic acid fragment, to a string of nucleotides at one or both ends of a nucleic acid fragment, or to the nucleotides of an entire nucleic acid fragment.Array reads can be obtained in a variety of ways, for example, using sequencing techniques or amplification techniques such as hybridization arrays or capture probes, or polymerase chain reaction (PCR) or linear amplification using a single primer or isothermal amplification.

[0091] As used herein, the term "read segment" refers to any form of nucleotide sequence read that includes raw sequence reads obtained directly from nucleic acid sequencing techniques or sequences derived therefrom, such as aligned sequence reads, folded sequence reads, or stitched sequence reads.

[0092] As used herein, the term "read count" refers to the total number of nucleic acid reads generated during a nucleic acid sequencing reaction, which may or may not be equivalent to the number of nucleic acid molecules generated.

[0093] As used herein, the terms "read depth", "sequencing depth", or "depth" may refer to the total number of unique nucleic acid fragments encompassing a particular locus or region of a microbial genome that are sequenced in a particular sequencing reaction. The sequencing depth can be expressed as "Y-fold", for example, 50-fold, 100-fold, etc., where "Y" refers to the number of unique nucleic acid fragments encompassing a particular locus that are sequenced in the sequencing reaction. In such cases, Y is necessarily an integer as it represents the actual sequencing depth for a particular locus. Alternatively, read depth, sequencing depth, or depth may refer to a measure of the central tendency (e.g., mean or mode) of the number of unique nucleic acid fragments encompassing one of a plurality of loci or regions of a microbial genome that are sequenced in a particular sequencing reaction. For example, in some embodiments, the sequencing depth refers to the average depth across all loci in a targeted sequencing panel, exome, or entire genome of a microbe. In such cases, Y can be expressed as a fraction or decimal as it refers to the average coverage across multiple loci. When the average depth is enumerated, the actual depth for any particular locus may differ from the overall depth that is enumerated. A metric can be determined that provides a range of sequencing depths that includes a defined percentage of the total number of loci. For example, a range of sequencing depths that includes 90%, 95%, or 99% of the loci. As will be understood by those skilled in the art, different sequencing technologies provide different sequencing depths. For example, low-pass whole genome sequencing may refer to a technology that provides a sequencing depth of less than 5-fold, less than 4-fold, less than 3-fold, or less than 2-fold, for example, from about 0.5-fold to about 3-fold.

[0094] As used herein, the term "sequencing breadth" refers to what proportion of a particular microbial genome has been sequenced. Sequencing breadth can be expressed as a fraction, decimal, or percentage, and is generally calculated as (number of loci analyzed / total number of loci in the genome). The denominator of the percentage can be the repeat-masked genome, and thus 100% can correspond to all of the reference genome excluding the masked portions. A repeat-masked genome can refer to a genome in which sequence repeats are masked (e.g., sequence reads align to unmasked portions of the genome). In some embodiments, any portion of the genome can be masked, and thus the breadth of sequencing can be evaluated for any desired portion of the genome.

[0095] As used herein, the terms "sequence ratio" and "coverage ratio" interchangeably refer to any measure of the number of units of genomic sequence in a first one or more biological samples (e.g., test and / or tumor samples) compared to the number of units of the respective genomic sequences in a second one or more biological samples (e.g., reference sample and / or control sample). In some embodiments, the sequence ratio is a copy ratio, log 2 -transformed copy ratio (e.g., log 2 copy ratio), coverage ratio, base fraction, allele fraction (e.g., variant allele fraction), and / or tumor ploidy. In some embodiments, the sequence ratio is a logarithmic N -transformed copy ratio, where N is any real number greater than 1.

[0096] As used herein, the term "sequencing probe" refers to a molecule that binds to a nucleic acid with an affinity based on the expected nucleotide sequence of the RNA or DNA present at that locus.

[0097] As used herein, the term "targeted panel" or "targeted gene panel" refers to a combination of probes for sequencing nucleic acids (e.g., by next-generation sequencing) present in a biological sample from a subject (e.g., a tumor sample, a liquid biopsy sample, a germ cell tissue sample, a white blood cell sample, or a tumor or tissue organoid sample) that are selected to map to one or more loci of interest within the genome.

[0098] As used herein, the term "sensitivity" or "true positive rate" (TPR) refers to the number of true positives divided by the sum of the number of true positives and false negatives. Sensitivity can characterize the ability of an assay or method to accurately identify the proportion of a population that truly has a symptom. For example, sensitivity can characterize the ability of a method to accurately identify the number of subjects within a population that have a particular biological characteristic.

[0099] As used herein, the term "specificity" or "true negative rate" (TNR) refers to the number of true negatives divided by the sum of the number of true negatives and false positives. Specificity can characterize the ability of an assay or method to accurately identify the proportion of a population that truly does not have a symptom. For example, specificity can characterize the ability of a method to correctly identify the number of subjects within a population that do not have a particular biological characteristic.

[0100] As used interchangeably herein, the term "classifier" or "model" refers to a machine learning model or algorithm.

[0101] In some embodiments, the model includes an unsupervised learning algorithm. An example of an unsupervised learning algorithm is cluster analysis. In some embodiments, the model includes supervised machine learning. Non-limiting examples of supervised learning algorithms include, but are not limited to, logistic regression, neural networks, support vector machines, naive Bayes algorithms, nearest neighbor algorithms, random forest algorithms, decision tree algorithms, boosted tree algorithms, polynomial logistic regression algorithms, linear models, linear regression, gradient boosting, hybrid models, hidden Markov models, Gaussian NB algorithms, linear discriminant analysis, diffusion models, or any combination thereof. In some embodiments, the model is a multinomial classifier algorithm. In some embodiments, the model is a two-stage stochastic gradient descent (SGD) model. In some embodiments, the model is a deep neural network (e.g., a deep and wide sample-level model).

[0102] Neural network. In some embodiments, the model is a neural network (e.g., a convolutional neural network and / or a residual neural network). A neural network algorithm, also known as an artificial neural network (ANN), includes convolutional and / or residual neural network algorithms (deep learning models). In some embodiments, a neural network is a machine learning algorithm trained to map an input data set to an output data set, and the neural network includes a group of interconnected nodes organized into multiple layers of nodes. For example, in some embodiments, a neural network architecture may include at least an input layer, one or more hidden layers, and an output layer. In some embodiments, a neural network may include any total number of layers and any number of hidden layers, and the hidden layers function as trainable feature extractors that enable mapping a set of input data to an output value or a set of output values. In some embodiments, a deep learning algorithm is a neural network that includes multiple hidden layers, e.g., two or more hidden layers. In some cases, each layer of a neural network includes a number of nodes (or "neurons"). In some embodiments, a node receives an input coming directly from either input data or the output of nodes in the previous layer and performs a specific operation, e.g., a summation operation. In some embodiments, the connection from the input to the node is associated with parameters (e.g., weights and / or weight coefficients). In some embodiments, a node has an input, x iSum the products of all pairs of the weights and their associated parameters. In some embodiments, the weighted sum is offset by a bias b. In some embodiments, the output of a node or neuron is gated using a threshold or activation function f, which in some cases is a linear or non-linear function. In some embodiments, the activation function can be, for example, a rectified linear unit (ReLU) activation function, a leaky ReLU activation function, or a saturated hyperbolic tangent, identity, binary step, logistic, arcTan, soft sign, parametric rectified linear unit, exponential linear unit, softPlus, bent identity, softExponential, sinusoid, sine, Gaussian, or sigmoid function, or any combination thereof.

[0103] In some embodiments, the weighting coefficients, bias values, and thresholds, or other computational parameters of the neural network are "taught" or "learned" during a training phase using one or more sets of training data. For example, in some embodiments, the parameters are trained using input data from a training data set and gradient descent or backpropagation such that the output value(s) calculated by the ANN match the examples included in the training data set. In some embodiments, the parameters are obtained from an error backpropagation neural network training process.

[0104] Any of a variety of neural networks are suitable for use according to the present disclosure. Examples include, but are not limited to, feedforward neural networks, radial basis function networks, recurrent neural networks, residual neural networks, convolutional neural networks, residual convolutional neural networks, or any combination thereof. In some embodiments, machine learning utilizes a pre-trained and / or transfer-learned ANN or deep learning architecture. In some embodiments, convolutional and / or residual neural networks are used according to the present disclosure.

[0105] For example, a deep neural network model includes an input layer, a plurality of individually parameterized (e.g., weighted) convolutional layers, and an output scorer. Each parameter (e.g., weight) of the convolutional layers as well as the input layer contributes to a plurality of parameters (e.g., weights) associated with the deep neural network model. In some embodiments, at least 50 parameters, at least 100 parameters, at least 1000 parameters, at least 2000 parameters, or at least 5000 parameters are associated with the deep neural network model. Thus, since the deep neural network model cannot be solved mentally, it is necessary to use a computer. In other words, when an input to the model is provided, in such embodiments, the model output needs to be determined using a computer rather than mentally. For example, see Krizhevsky et al., 2012, “Imagenet classification with deep convolutional neural networks,” in Advances in Neural Information Processing Systems 2, Pereira, Burges, Bottou, Weinberger, eds., pp. 1097-1105, Curran Associates, Inc., Zeiler, 2012 “ADADELTA: an adaptive learning rate method,” ‘CoRR, vol. abs / 1212.5701, and Rumelhart et al., 1988, “Neurocomputing: Foundations of research,” ch. Learning Representations by Back-propagating Errors, pp. 696-699, Cambridge, MA, USA: MIT Press, each of which is incorporated herein by reference.

[0106] Neural network algorithms, including convolutional neural network models suitable for use as models, are disclosed, for example, in Vincent et al., 2010, “Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion,” J Mach Learn Res 11, pp. 3371 - 3408, Larochelle et al., 2009, “Exploring strategies for training deep neural networks,” J Mach Learn Res 10, pp. 1 - 40, and Hassoun, 1995, Fundamentals of Artificial Neural Networks, Massachusetts Institute of Technology, each of which is incorporated herein by reference. Further exemplary neural networks suitable for use as models are disclosed in Duda et al., 2001, Pattern Classification, Second Edition, John Wiley & Sons, Inc., New York, and Hastie et al., 2001, The Elements of Statistical Learning, Springer - Verlag, New York, each of which is incorporated herein by reference in its entirety. Further exemplary neural networks suitable for use as models are also described in Draghici, 2003, Data Analysis Tools for DNA Microarrays, Chapman & Hall / CRC, and Mount, 2001, Bioinformatics: sequence and genome analysis, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York, each of which is incorporated herein by reference in its entirety.

[0107] Support vector machine. In some embodiments, the model is a support vector machine (SVM). SVM algorithms suitable for use as a model are described, for example, in Cristianini and Shawe-Taylor, 2000, “An Introduction to Support Vector Machines,” Cambridge University Press, Cambridge; Boser et al., 1992, “A training algorithm for optimal margin classifiers,” in Proceedings of the 5th Annual ACM Workshop on Computational Learning Theory, ACM Press, Pittsburgh, Pa., pp. 142-152; Vapnik, 1998, Statistical Learning Theory, Wiley, New York; Mount, 2001, Bioinformatics: sequence and genome analysis, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y.; Duda, Pattern Classification, Second Edition, 2001, John Wiley & Sons, Inc., pp. 259, 262-265; and Hastie, 2001, The Elements of Statistical Learning, Springer, New York; and Furey et al., 2000, Bioinformatics 16, 906-914, each of which is hereby incorporated by reference in its entirety. When used for classification, an SVM separates a given set of binary-labeled data from labeled data using a hyperplane that is maximally distant. In certain cases where linear separation is not possible, the SVM functions in combination with a “kernel” technique that automatically realizes a non-linear mapping into the feature space. The hyperplane found by the SVM in the feature space corresponds, in some cases, to a non-linear decision boundary in the input space.In some embodiments, a plurality of parameters (e.g., weights) associated with the SVM define a hyperplane. In some embodiments, the hyperplane is defined by at least 10, at least 20, at least 50, or at least 100 parameters, and the SVM model requires a computer to compute because it cannot be solved mentally.

[0108] Naive Bayes algorithm. In some embodiments, the model is a Naive Bayes model. A Naive Bayes model suitable for use as a model is disclosed, for example, in Ng et al., 2002, “On discriminative vs. generative classifiers: A comparison of logistic regression and naive Bayes,” Advances in Neural Information Processing Systems, 14, which is incorporated herein by reference. A Naive Bayes model is any model within the family of “probabilistic models” based on applying Bayes' theorem with a strong (naive) independence assumption between features. In some embodiments, they are combined with kernel density estimation. See, for example, Hastie et al., 2001, The elements of statistical learning: data mining, inference, and prediction, eds. Tibshirani and Friedman, Springer, New York, which is incorporated herein by reference.

[0109] Nearest neighbor algorithm. In some embodiments, the model is a nearest neighbor algorithm. In some implementations, the nearest neighbor model is memory-based and does not include a model to be fitted. For the nearest neighbor, when a query point x 0 (test subject) is given, k training points x (r) , r,... x 0The k (here, the training target) that is closest to the distance is identified, and then the point x 0 is classified using the k nearest neighbors. In some embodiments, the Euclidean distance in the feature space is used to determine the distance d (i) = ||x (i)- x (o) ||. Typically, when a nearest neighbor algorithm is used, the abundance data used to calculate the linear discriminant is standardized to have a mean of zero and a variance of one. In some embodiments, the nearest neighbor rule is improved to address issues of unequal class priors, differential misclassification costs, and feature selection. Many of these improvements involve some form of weighted voting for adjacency. For further information on nearest neighbor analysis, see Duda, Pattern Classification, Second Edition, 2001, John Wiley & Sons, Inc, and Hastie, 2001, The Elements of Statistical Learning, Springer, New York, each of which is incorporated herein by reference.

[0110] The k-nearest neighbor model is a non-parametric machine learning method in which the input consists of the k nearest training examples in the feature space. The output is class membership. The object is classified by a plurality of votes from adjacent objects, and the object is assigned to the most common class among its k nearest neighbors (k is typically a small positive integer). In the case of k = 1, the object is simply assigned to the class of its single nearest neighbor. See Duda et al., 2001, Pattern Classification, Second Edition, John Wiley & Sons, which is incorporated herein by reference. In some embodiments, the number of distance calculations required to solve the k-nearest neighbor model is such that a computer is used to solve the model for a given input because it cannot be performed mentally.

[0111] Random forest, decision tree, and boosting tree algorithms. In some embodiments, the model is a decision tree. Decision trees suitable for use as a model are outlined in Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, pp. 395-396, which is incorporated herein by reference. Decision tree-based methods partition the feature space into a set of rectangles and fit a model (such as a constant) to each. In some embodiments, the decision tree is a random forest regression. For example, one particular algorithm is the Classification and Regression Trees (CART) method. Other particular decision tree algorithms include, but are not limited to, ID3, C4.5, MART, and random forest. CART, ID3, and C4.5 are described in Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, pp. 396-408 and pp. 411-412, which is incorporated herein by reference. CART, MART, and C4.5 are described in Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, Chapter 9, which is incorporated herein by reference in its entirety. Random forest is described in Breiman, 1999, “Random Forests--Random Features,” Technical Report 567, Statistics Department, U.C. Berkeley, September 1999, which is incorporated herein by reference. In some embodiments, the decision tree model includes at least 10, at least 20, at least 50, or at least 100 parameters (e.g., weights and / or decisions), which cannot be mentally resolved and thus require a computer to compute.

[0112] Regression. In some embodiments, the model uses a regression algorithm. In some embodiments, the regression algorithm is any type of regression. For example, in some embodiments, the regression algorithm is logistic regression. In some embodiments, the regression algorithm is logistic regression with Lasso, L2, or elastic net regularization. In some embodiments, these extracted features having corresponding regression coefficients that fail to meet a threshold are removed (deleted) from consideration. In some embodiments, a generalization of the logistic regression model for handling multi-categorical responses is used as the model. The logistic regression algorithm is disclosed in Agresti, An Introduction to Categorical Data Analysis, 1996, Chapter 5, pp. 103 - 144, John Wiley & Son, New York, which is incorporated herein by reference. In some embodiments, the model utilizes a regression model disclosed in Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York. In some embodiments, the logistic regression model includes at least 10, at least 20, at least 50, at least 100, or at least 1000 parameters (e.g., weights), and since it cannot be mentally solved, a computer is required to compute it.

[0113] Linear discriminant analysis algorithm. In some embodiments, linear discriminant analysis (LDA), also known as normal discriminant analysis (NDA), or discriminant function analysis, is a generalization of Fisher's linear discriminant, a method used in statistics, pattern recognition, and machine learning to find a linear combination of features that characterize or separate two or more classes of objects or events. In some embodiments, the resulting combination is used as a model (linear model) in some embodiments of the present disclosure.

[0114] Mixture models and hidden Markov models. In some embodiments, the model is a mixture model as described in McLachlan et al., Bioinformatics 18(3):413-422, 2002. In some embodiments, particularly those embodiments that include a time component, the model is a hidden Markov model as described in Schliep et al., 2003, Bioinformatics 19(1):i255-i263.

[0115] Clustering. In some embodiments, the model is an unsupervised clustering model. In some embodiments, the model is a supervised clustering model. Clustering suitable for use as a clustering algorithm is described, for example, on pages 211-256 of Duda and Hart, Pattern Classification and Scene Analysis, 1973, John Wiley & Sons, Inc., New York (hereinafter in this specification, "Duda 1973"), which is hereby incorporated by reference in its entirety. As an illustrative example, in some embodiments, the clustering problem is described as one of finding natural groupings in a dataset. To identify natural groupings, two problems are addressed. First, determine a method for measuring the similarity (or dissimilarity) between two samples. This metric (e.g., similarity measure) is used to ensure that samples within one cluster are more similar to each other than samples within the other cluster. Second, determine a mechanism for dividing the data into clusters using the similarity measure. One way to start a clustering investigation is to define a distance function and compute a matrix of distances between all pairs of samples in the training set. If distance is a good measure of similarity, the distance between reference entities within the same cluster is significantly smaller than the distance between reference entities in different clusters. However, in some implementations, clustering does not use a distance metric. For example, in some embodiments, a non-metric similarity function s(x, x') is used to compare two vectors x and x'. In some such embodiments, s(x, x') is a symmetric function that has a large value when x and x' are "similar" in some way. Once a method for measuring "similarity" or "dissimilarity" between points in the dataset is selected, clustering uses a criterion function that measures the clustering quality of any partition of the data. The data is clustered using a partition of the dataset that extremizes the criterion function.Certain exemplary clustering techniques contemplated for use in the present disclosure include, but are not limited to, hierarchical clustering (agglomerative clustering using the nearest neighbor algorithm, farthest neighbor algorithm, average linkage algorithm, centroid algorithm, or sum of squares algorithm), k-means clustering, fuzzy k-means clustering, and Jarvis-Patrick clustering. In some embodiments, clustering includes unsupervised clustering (e.g., without a pre-conceived number of clusters and / or prior determination of cluster assignments).

[0116] Ensembles of models and boosting. In some embodiments, an ensemble (two or more) of models is used. In some embodiments, boosting techniques such as AdaBoost are used in combination with many other types of learning algorithms to improve the performance of the model. In this approach, the output of any of the models disclosed herein, or their equivalents, is combined into a weighted sum that represents the final output of the boosted model. In some embodiments, multiple outputs from the models are combined using any measure of central tendency known in the art, including, but not limited to, mean, median, mode, weighted mean, weighted median, weighted mode, etc. In some embodiments, the multiple outputs are combined using a voting method. In some embodiments, each model within the ensemble of models is weighted or unweighted.

[0117] As used herein, the term "parameter" refers to any coefficient of an internal or external element (e.g., weight and / or hyperparameter) in an algorithm, model, regressor, and / or classifier that can affect (e.g., modify, adapt, and / or adjust) one or more inputs, outputs, and / or functions in the algorithm, model, regressor, and / or classifier, or similarly any value. For example, in some embodiments, a parameter refers to any coefficient, weight, and / or hyperparameter that can be used to control, modify, adapt, and / or adjust the behavior, learning, and / or performance of an algorithm, model, regressor, and / or classifier. In some cases, a parameter is used to increase or decrease the effect of an input (e.g., feature) on an algorithm, model, regressor, and / or classifier. As a non-limiting example, in some embodiments, a parameter is used to increase or decrease the effect of a node (e.g., a neural network) that includes one or more activation functions. The assignment of a parameter to a particular input, output, and / or function is not limited to any one paradigm for a given algorithm, model, regressor, and / or classifier, and can be used in any suitable algorithm, model, regressor, and / or classifier architecture for a desired performance. In some embodiments, a parameter has a fixed value. In some embodiments, the value of a parameter is manually and / or automatically adjustable. In some embodiments, the value of a parameter is modified by a validation and / or training process for an algorithm, model, regressor, and / or classifier (e.g., by an error minimization and / or backpropagation method). In some embodiments, the algorithms, models, regressors, and / or classifiers of the present disclosure include a plurality of parameters.In some embodiments, the plurality of parameters are n parameters, where n ≥ 2; n ≥ 5; n ≥ 10; n ≥ 25; n ≥ 40; n ≥ 50; n ≥ 75; n ≥ 100; n ≥ 125; n ≥ 150; n ≥ 200; n ≥ 225; n ≥ 250; n ≥ 350; n ≥ 500; n ≥ 600; n ≥ 750; n ≥ 1,000; n ≥ 2,000; n ≥ 4,000; n ≥ 5,000; n ≥ 7,500; n ≥ 10,000; n ≥ 20,000; n ≥ 40,000; n ≥ 75,000; n ≥ 100,000; n ≥ 200,000; n ≥ 500,000, n ≥ 1 × 10. 6 , n ≥ 5 × 10 6 , or n ≥ 1 × 10 7 . Thus, the algorithms, models, regressors, and / or classifiers of the present disclosure cannot be implemented mentally. In some embodiments, n is from 10,000 to 1 × 10 7 , from 100,000 to 5 × 10 6 , or from 500,000 to 1 × 10 6 . In some embodiments, the algorithms, models, regressors, and / or classifiers of the present disclosure operate in a k-dimensional space, where k is a positive integer of 5 or greater (e.g., 5, 6, 7, 8, 9, 10, etc.). Thus, the algorithms, models, regressors, and / or classifiers of the present disclosure cannot be implemented mentally.

[0118] As used herein, the term "untrained model" (e.g., "untrained classifier" and / or "untrained neural network") refers to a machine learning model or algorithm such as a classifier or neural network that has not been trained on a target dataset. In some embodiments, "training of a model" (e.g., "training of a neural network") refers to the process of training an untrained or partially trained model (e.g., "untrained or partially trained neural network"). Further, it should be understood that the term "untrained model" does not exclude the possibility that transfer learning techniques may be used for such training of an untrained or partially trained model. For example, Fernandes et al., 2017, "Transfer Learning with Partial Observability Applied to Cervical Cancer Screening," Pattern Recognition and Image Analysis: 8th Iberian Conference Proceedings, 243-250, which is incorporated herein by reference, provides non-limiting examples of such transfer learning. In examples where transfer learning is used, the above-mentioned untrained model is provided with additional data more than that of the primary training dataset. Typically, this additional data is in the form of parameters (e.g., coefficients, weights, and / or hyperparameters) learned from another auxiliary training dataset. Further, although the description of a single auxiliary training dataset is disclosed, it should be understood that there is no limit to the number of auxiliary training datasets that can be used to complement the primary training dataset when training an untrained model in the present disclosure. For example, in some embodiments, two or more auxiliary training datasets, three or more auxiliary training datasets, four or more auxiliary training datasets, or five or more auxiliary training datasets are used to complement the primary training dataset through transfer learning, and each such auxiliary dataset is different from the primary training dataset. In some such embodiments, any method of transfer learning is used.For example, consider the case where, in addition to a primary training dataset, there are a first auxiliary training dataset and a second auxiliary training dataset. In such a case, the parameters learned from the first auxiliary training dataset (by application of a first model to the first auxiliary training dataset) are applied to the second auxiliary training dataset using transfer learning techniques (e.g., a second model that is the same as or different from the first model), which then results in a trained intermediate model whose coefficients are applied to the primary training dataset, which, together with the primary training dataset itself, is applied to an untrained model. Alternatively, in another exemplary embodiment, a first set of parameters learned from the first auxiliary training dataset (by application of a first model to the first auxiliary training dataset) and a second set of parameters learned from the second auxiliary training dataset (by application of a second model that is the same as or different from the first model to the second auxiliary training dataset) are each separately applied to separate instances of the primary training dataset (e.g., by separate independent matrix multiplications), and both such applications of the parameters to separate the instances of the primary training dataset, together with the primary training dataset itself (or some reduced form of the primary training dataset such as the principal components or regression coefficients learned from the primary training dataset), are then applied to the untrained model to train the untrained model.

[0119] As used herein, the term "instruction" refers to an instruction given to a computer processor by a computer program. In a digital computer, each instruction is a sequence of 0s and 1s that describes a physical operation to be performed by the computer. Such instructions can include data transfer instructions and data manipulation instructions. In some embodiments, each instruction is a type of instruction within an instruction set recognized by a particular processor type used to execute the instruction. Examples of instruction sets include, but are not limited to, Reduced Instruction Set Computer (RISC), Complex Instruction Set Computer (CISC), Minimal Instruction Set Computer (MISC), Very Long Instruction Word (VLIW), Explicitly Parallel Instruction Computing (EPIC), and One Instruction Set Computer (OISC).

[0120] Some aspects are described below with reference to exemplary applications for illustration. It should be understood that numerous specific details, relationships, and methods are set forth in order to provide a complete understanding of the features described herein. However, one of ordinary skill in the art will readily recognize that the features described herein can be practiced without one or more of the specific details, or in other ways. The features described herein are not limited by the order of the acts or events recited, since some acts may occur in a different order and / or concurrently with other acts or events. Further, not all acts or events illustrated are required to implement a methodology in accordance with the features described herein.

[0121] Reference is now made in detail to embodiments, examples of which are illustrated in the accompanying drawings. In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to one of ordinary skill in the art that the present disclosure may be practiced without these specific details. In other instances, well-known methods, procedures, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.

[0122] Exemplary System Embodiments Since an overview of some aspects of the present disclosure and some definitions used in the present disclosure have been provided, the details of the exemplary system will be described in conjunction with FIG. 1. FIG. 1 is a block diagram showing a system 100 according to some implementation manners. In some implementation manners, the system 100 includes one or more processing units CPU 102 (also referred to as processors), one or more network interfaces 104, a user interface 106 including a display 108 and an input system 110 (optionally), a non-persistent memory 111, a persistent memory 112, and one or more communication buses 114 for interconnecting these components. One or more communication buses 114 optionally include a circuit (sometimes referred to as a chipset) for interconnecting and controlling communications between system components. The non-persistent memory 111 typically includes high-speed random access memories such as DRAM, SRAM, DDR RAM, ROM, EEPROM, flash memory, etc., while the persistent memory 112 typically includes CD-ROM, digital versatile disk (DVD), or other optical storage, magnetic cassette, magnetic tape, magnetic disk storage, or other magnetic storage devices, magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid state storage devices. The persistent memory 112 optionally includes one or more storage devices located remotely from the CPU(s) 102. The persistent memory 112 and the non-volatile memory device(s) within the non-persistent memory 112 include non-transitory computer-readable storage media. In some embodiments, the non-persistent memory 111 or alternatively, the non-transitory computer-readable storage media stores the following programs, modules, and data structures, or subsets thereof, optionally in conjunction with the persistent memory 112: ● An optional operating system 116 including procedures for processing various basic system services and executing hardware-dependent tasks; ● An optional network communication module (or instructions) 118 for connecting the system 100 to other devices and / or communication network 104; ● A microbiome evaluation module 140 for determining a subject's disease state among a plurality of disease states based on the composition of the subject's microbiome; and ● A data store of subject information 140 based on microbiome sequencing results 150, including the presence values 152 of microorganisms in each of guilds 152-A and 152-B, as described herein.

[0123] In various embodiments, one or more of the identified elements described above are stored in one or more of the aforementioned memory devices and correspond to a set of instructions for performing the aforementioned functions. The identified modules, data, or programs (e.g., sets of instructions) described above need not be implemented as separate software programs, procedures, data sets, or modules, and thus, various subsets of these modules and data may be combined in various implementations or rearranged otherwise. In some embodiments, the non-persistent memory 111 optionally stores a subset of the modules and data structures identified above. Further, in some embodiments, the memory stores additional modules and data structures not described above. In some embodiments, one or more of the identified elements described above are stored in a computer system other than that of the visualization system 100 and are addressable by the visualization system 100, and thus the visualization system 100 can retrieve all or part of such data when needed.

[0124] FIG. 1 depicts the "System 100", which is not a structural schematic of the implementation described herein, but is rather intended as a functional description of various features that may exist in a computer system. In practice, as will be recognized by those skilled in the art, the separately shown items may be combined and some items may be separated. Further, FIG. 1 depicts certain data and modules within the non-persistent memory 111, but some or all of these data and modules may alternatively be stored in the persistent memory 112.

[0125] 1. Method for identifying a set of gut microbiota FIG. 2 is a schematic diagram of a method for identifying a set of gut microbiota, as described below. The method can be implemented using a computer system (e.g., the computer system 100 shown and described above with reference to FIG. 1).

[0126] Referring to block 200, in some embodiments, the method includes, for each respective subject in a first plurality of subjects having a first state of a biological characteristic in electronic form, obtaining corresponding plural genomic abundance values including, for each respective gut microorganism among a plurality of gut microorganisms, a corresponding value for the abundance of the genome of each respective gut microorganism in a biological sample derived from the gut of the respective subject. In some embodiments, the first plurality of subjects includes at least 50, at least 100, at least 200, at least 500, at least 1000, at least 5000, at least 10,000, at least 20,000, at least 50,000, at least 100,000, at least 500,000, or at least 1,000,000 subjects. In some embodiments, the first plurality of subjects includes up to 1,000,000, up to 500,000, up to 100,000, up to 50,000, up to 20,000, up to 10,000, up to 1000, up to 500, up to 100, or up to 50 subjects. In some embodiments, the first plurality of subjects consists of 50 - 100, 50 - 200, 50 - 500, 100 - 500, 200 - 500, 200 - 1000, 500 - 1000, 200 - 5,000, 1000 - 10,000, 5000 - 200,000, 10,000 - 50,000, 20,000 - 100,000, or 500,000 - 1,000,000 subjects. In some embodiments, the first plurality of subjects falls within another range starting from at least 50 subjects and ending with up to 10,000,000 subjects. In some embodiments, the first plurality of subjects share similar demographic characteristics (such as age, gender, ethnicity, etc.). In some embodiments, the first plurality of subjects share similar physical characteristics (such as weight, height, BMI value, etc.). In some embodiments, the first plurality of subjects share a similar health status (such as physical or mental state, medical history, genetic carrier, or drug use, etc.). In some embodiments, the first plurality of subjects share or are similar in terms of behavior and lifestyle preferences (such as diet, physical activity, or drug use).

[0127] In some embodiments, the corresponding value for the genomic abundance is a value representing the absolute abundance of the microbial genome. In some embodiments, the corresponding value for the genomic abundance is a value representing a normalized abundance value or a relative abundance value (e.g., the abundance of one microorganism normalized to the abundance of the total microbiome of interest). In some embodiments, the corresponding value for the genomic abundance is a value representing an averaged abundance value (e.g., the average of abundances obtained at different time points or from different biological samples from a patient, or the average of abundances obtained using different probes), or any combination of the above. The corresponding value for the genomic abundance is measured by any technique known in the art. In some embodiments, the genomic abundance value of the genome is measured by quantitative PCR (qPCR), such as bacterial 16S rRNA qPCR, RT-PCR, or qRT-PCR, to quantify the abundance of a region of interest in the genome, as described in, for example, U.S. Patent No. 11,427,865, the entire disclosure of which is incorporated herein by reference. In some embodiments, the genomic abundance value is measured by targeted sequencing (e.g., 16S rRNA sequencing or any other suitable biomarker), partial genomic sequencing, or whole-genome sequencing, thereby quantifying the number of reads of a targeted region within the microbial genome to determine the genomic abundance, as disclosed in, for example, U.S. Patent Application Publication No. 2021 / 0403986 or U.S. Patent No. 11,332,783, the entire disclosures of which are incorporated herein by reference. In some embodiments, deep sequencing is used to determine the abundance of a targeted sequence, as disclosed in, for example, U.S. Patent Application Publication No. 2018 / 0237863, the entire disclosure of which is incorporated herein by reference.In some embodiments, the depth of sequencing is at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 150, at least 200, at least 300, at least 400, at least 500, at least 750, at least 1000, or more. In some embodiments, for example, shotgun metagenome sequencing is used to provide sequence reads of the genome in a sample, as described in U.S. Patent No. 11,028,449, the content of which is hereby incorporated by reference in its entirety.

[0128] Referring to block 202, in some embodiments, for each respective one of the first plurality of subjects, the biological sample derived from the intestine of each respective subject is a fecal sample. In some embodiments, the sample is a tissue biopsy, an intestinal, or a mucosal sample. For example, see Tang Q, Jet al., Current Sampling Methods for Gut Microbiota: A Call for More Precise Devices, Front Cell Infect Microbiol., 10:151 (2020), the content of which is hereby incorporated by reference in its entirety.

[0129] Referring to block 204, in some embodiments, the method comprises sequencing genomic DNA from a corresponding biological sample derived from the intestine of each respective target among a first plurality of targets, thereby obtaining a corresponding first plurality of at least 100,000 nucleic acid sequences. In some embodiments, the first plurality of nucleic acid sequences comprises at least 100,000, at least 250,000, at least 500,000, at least 1,000,000, at least 2,500,000, at least 5,000,000, at least 10,000,000, or at least 50,000,000 nucleic acid sequences. In some embodiments, the first plurality of nucleic acid sequences comprises 250,000,000 or fewer, 100,000,000 or fewer, 50,000,000 or fewer, 25,000,000 or fewer, 10,000,000 or fewer, 5,000,000 or fewer, 1,000,000 or fewer, 100,000 or fewer nucleic acid sequences. In some embodiments, the first plurality of nucleic acid sequences consists of 100,000 to 1,000,000, 200,000 to 5,000,000, 500,000 to 10,000,000, 1,000,000 to 20,000,000, 5,000,000 to 50,000,000, 10,000,000 to 100,000,000, or 50,000,000 to 250,000,000 nucleic acid sequences. In some embodiments, the first plurality of nucleic acid sequences begins with at least 100,000 nucleic acid sequences and falls within another range ending with 250,000,000 or fewer nucleic acid sequences.

[0130] In some embodiments, the first plurality (e.g., at least 100,000) of nucleic acid sequences are obtained by metagenomic sequencing, such as disclosed in U.S. Patent Application Publication No. 2016 / 0239602 or U.S. Patent No. 11,495,326, the contents of which are hereby incorporated by reference in their entirety. In some embodiments, metagenomic sequencing further includes generating a plurality of metagenomic fragment reads. In some embodiments, metagenomic sequencing further includes fragmenting the microbial genome into randomly sized target fragments. The resulting fragments can vary in size. In one embodiment, fragments of approximately 500 nucleotides can be obtained. In some embodiments, fragments of 100 to 2000 nucleotides, such as 200 to 800 nucleotides, 100 to 900 nucleotides, 100 to 1000 nucleotides, 300 to 800 nucleotides, 400 to 900 nucleotides can be obtained. In some embodiments, the method can further include extracting metagenomic fragments from a corresponding biological sample. In some embodiments, metagenomic sequencing further includes sequencing the fragments using a high-throughput sequencing method to generate a plurality of sequencing reads.

[0131] In some embodiments, the first plurality (e.g., at least 100,000) of nucleic acid sequences are obtained by targeted panel sequencing, as described, for example, in U.S. Patent Application Publication No. 2019 / 0316209. In some embodiments, targeted panel sequencing hybridizes genomic DNA isolated from a biological sample derived from the gut of a subject to a panel of probes that includes one or more probes that hybridize to unique sequences within the genome of each of a plurality of microorganisms, such as each of the microorganisms listed in Tables 1, 2, and / or FIGS. 42A - 42XX, prior to sequencing the recovered nucleic acids. In some embodiments, a combination of quasi-unique sequences (e.g., sequences found in a small number of microbial genomes) can be used to back-calculate genomic abundance values using an algorithm, such as a system of equations. In some embodiments, the panel of probes includes at least one probe that hybridizes to a sequence unique to each detected microbial genome. In some embodiments, the panel of probes includes at least 2, at least 3, at least 4, at least 5, at least 10, at least 25, at least 50, or more probes that hybridize to different sequences unique to each detected microbial genome. In some embodiments, the panel of probes includes at least 20, at least 30, at least 40, at least 50, at least 75, at least 100, at least 125, at least 150, at least 200, at least 150, at least 300, at least 400, at least 500, at least 750, at least 1000, at least 1250, at least 1500, at least 2000, at least 2500, at least 3000, at least 4000, at least 5000, at least 7500, at least 10,000, or more unique probes.

[0132] In some embodiments, the sequenced genomic DNA from the corresponding biological sample includes at its ends partial or complete sequencing platform adapter sequences useful for sequencing using the targeted sequencing platform. Targeted sequencing platforms include, but are not limited to, the Illumina® HiSeq®, MiSeq®, and Genome Analyzer® sequencing systems, the Ion Torrent® Ion PGM® and Ion Proton® sequencing systems, the Pacific Biosciences PACBIO RS II Sequel system, the Life Technologies® SOLiD sequencing system, the Roche 454GS FLX+ and GS Junior sequencing systems, the Oxford Nanopore MinION® system, or any other targeted sequencing platform.

[0133] In some embodiments, the plurality of genomic abundance values are determined using a microarray that includes probe sequences capable of detecting the unique genomic sequences of each of the plurality of gut microorganisms. In some embodiments, the panel of probes on the microarray includes at least one probe that hybridizes to a sequence unique to each microbial genome to be detected. In some embodiments, the panel of probes includes at least two, at least three, at least four, at least five, at least ten, at least twenty-five, at least fifty, or more probes that hybridize to different sequences unique to each microbial genome to be detected. In some embodiments, the panel of probes includes at least twenty, at least thirty, at least forty, at least fifty, at least seventy-five, at least one hundred, at least one hundred and twenty-five, at least one hundred and fifty, at least two hundred, at least one hundred and fifty, at least three hundred, at least four hundred, at least five hundred, at least seven hundred and fifty, at least one thousand, at least one thousand two hundred and fifty, at least one thousand five hundred, at least two thousand, at least two thousand five hundred, at least three thousand, at least four thousand, at least five thousand, at least seven thousand five hundred, at least ten thousand, or more unique probes.

[0134] Referring to block 206, in some embodiments, the method includes, for each respective target in a first plurality of targets, obtaining, in electronic form, a corresponding first plurality of nucleic acid sequences (e.g., at least 100,000) for genomic DNA from a corresponding biological sample derived from the intestine of the respective target, and for each respective target in the first plurality of targets, determining, from the corresponding first plurality of at least 100,000 nucleic acid sequences, a corresponding genomic abundance value for each respective gut microbe among a plurality of gut microbes. In some embodiments, the genomic abundance values determined for each respective target in the first plurality of targets include genomic abundance values of at least 20, at least 25, at least 50, at least 100, at least 250, at least 500, at least 1,000, at least 5,000, or at least 10,000. In some embodiments, the genomic abundance values determined for each respective target in the first plurality of targets include genomic abundance values of 250,000 or less, 100,000 or less, 50,000 or less, 25,000 or less, 10,000 or less, 5,000 or less, 1,000 or less, 100 or less, 50 or less, 30 or less, or 20 or less. In some embodiments, the genomic abundance values determined for each respective target in the first plurality of targets consist of genomic abundance values of 10 to 40, 20 to 50, 30 to 80, 40 to 100, 50 to 150, 60 to 200, 80 to 300, 90 to 500, 100 to 1,000, 500 to 2,000, or 1,000 to 5,000. In some embodiments, the genomic abundance values determined for each respective target in the first plurality of targets fall within another range that starts with a genomic abundance value of 10 or less and ends with a genomic abundance value of 250,000 or less.

[0135] Referring to block 208, in some embodiments, the method comprises, for each respective target in the first plurality of targets, assembling a corresponding first plurality of gut microbial genomes by metagenomic de novo sequence assembly from a corresponding first plurality of at least 100,000 nucleic acid sequences, and for each respective gut microbial genome in the corresponding first plurality of gut microbial genomes, calculating a corresponding genomic abundance of the respective gut microbial genome. In some embodiments, the metagenomic de novo sequence assembly further comprises generating contigs based on sequencing reads generated by shotgun sequencing techniques, such as those described in U.S. Patent No. 10,529,443, the contents of which are hereby incorporated by reference in their entirety. In some embodiments, the first plurality of (e.g., at least 100,000) nucleic acid sequences can be assembled into the entire genomes of a plurality of gut microbes. In some embodiments, the first plurality of (e.g., at least 100,000) nucleic acid sequences can be assembled into partial genomes of a plurality of gut microbes.

[0136] Referring to block 210, in some embodiments, the method comprises, for each respective target in the first plurality of targets, assigning each respective nucleic acid sequence in the corresponding first plurality of at least 100,000 sequences to each respective gut microbe among the plurality of gut microbes, thereby generating, for each respective gut microbe among the plurality of gut microbes, a corresponding count of each respective nucleic acid sequence assigned to the respective gut microbe; and determining, for each respective gut microbe among the plurality of gut microbes, a corresponding genomic abundance value for the respective gut microbe based on the corresponding count of each respective nucleic acid sequence assigned to the respective gut microbe. In some embodiments, assigning each respective nucleic acid to each respective gut microbe comprises mapping the nucleic acid to a reference nucleic acid, e.g., a contig listed in FIG. 41. In some embodiments, assigning each respective nucleic acid to each respective gut microbe comprises annotating genomic information based on an existing database. In some embodiments, the nucleic acid sequences are analyzed and the annotation defines a taxonomic assignment using sequence similarity methods and phylogenetic placement methods, or a combination of the two strategies.

[0137] Array similarity-based methods include, but are not limited to, those well-known to those skilled in the art, including various implementations of these algorithms such as BLAST, BLASTx, tBLASTn, tBLASTx, RDP classifier, DNAclust, and Qiime or Mothur. These methods rely on mapping sequence reads to a reference database and selecting the matches with the best scores and e-values. In some embodiments, phylogenetic methods are used in combination with sequence similarity methods to improve the call accuracy of annotation or taxonomic assignment. Common databases include, but are not limited to, GT-DBTK, National Center for Biotechnology Information (NCBI) Genbank, European Bioinformatics Institute-European Nucleotide Archive (European Bioinformatics Institute-European Nucleotide Archive; EBI-ENA), National Institute of Genetics, U.S. Department of ENERGY (USDOE) Integrated Microbial Genomes (Integrated Microbial Genomes) & Microbiomes; IMG / M), and other databases available in the art.

[0138] Referring to block 212, in some embodiments, the first state of the biological characteristic is the absence of a disease or disorder, such as type 2 diabetes (T2D), hypertension (HT), schizophrenia (SCZ), atherosclerotic cardiovascular disease (ACVD), cirrhosis (LC), inflammatory bowel disease (IBD), colorectal cancer (CRC), ankylosing spondylitis (AS), and Parkinson's disease (PD), multiple sclerosis (MS), type 2 Gaucher disease (GDII), COVID-19 (COV), Behçet's disease (BD), autism spectrum disorder (ASD), or pancreatic cancer (PC). In some embodiments, the disease or disorder is cancer, Alzheimer's disease, cardiovascular disease, autoimmune disease, mental health disease, infectious disease, or genetic disorder.

[0139] In some embodiments, the first state of the biological characteristic is the first severity of a disease or disorder. In some embodiments, the severity of a disease is classified by the type, frequency or intensity experienced by the subject. In some embodiments, the severity of a disease is classified by the progression or prognosis of the disease or disorder, such as different stages of cancer. In some embodiments, the first state of the biological characteristic is an untreated disease or disorder. In some embodiments, the first state of the biological characteristic is a disease or disorder treated with a first therapy, such as surgery, radiation therapy, chemotherapy, targeted therapy, gene therapy, immunotherapy, drug therapy, dietary change, lifestyle change.

[0140] In some embodiments, the first state of the biological characteristic is the first level of nutrients in the diet, such as carbohydrates, proteins, fats, vitamins, fiber. In some embodiments, the first state of the biological characteristic is the first age. In some embodiments, a threshold is provided for determining the first state of the biological characteristic, such as the level of a biomarker, a diagnostic cut-off value, or a threshold nutrient intake level.

[0141] Referring to block 214, in some embodiments, the plurality of gut microbiota includes at least 20 gut microbiota selected from Table 1, Table 2, or FIGS. 42A - 42XX. In some embodiments, at least about 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more gut microbiota are selected from Table 1, Table 2, or FIGS. 42A - 42XX.

Table 1 - 1

Table 1 - 2

Table 1 - 3

Table 1 - 4

Table 1 - 5

Table 1 - 6

Table 1 - 7

Table 2 - 1

Table 2 - 2

Table 2 - 3

Table 2 - 4

Table 2-5

Table 2-6

Table 2-7

Table 2-8

Table 2-9

Table 2-10

Table 2-11

Table 2-12

Table 2-13

[0142] The bacterial species listed in Table 1, Table 2, and Figures 42A - 42XX were identified by metagenomic sequencing of genomic DNA isolated from human fecal samples as described in the Examples and were determined to be part of two competing microbiome guilds for at least one biological property. Briefly, genomic DNA was isolated from each fecal sample, sequenced by next-generation sequencing, and contigs for microbial genomic sequences were constructed de novo. Generally, the contigs identified for each microorganism are predicted to represent more than 95% of the entire genome for the microorganism. Genomic constructs with less than 1% sequence divergence from each other were combined and defined as being from the same microorganism. The genomic contigs for each microorganism listed in Table 1, Table 2, and Figures 42A - 42XX are provided in the sequence listing submitted with this application. The taxonomic assignment for each microorganism is shown in Table 1, Table 2, or Figures 42A - 42XX. A correspondence between the sequence identifier assigned to each contig and the microorganism to which it belongs is provided in Figure 41. For example, the contigs provided as SEQ ID NOs: 1 - 68 are classified as Bacteria domain, Proteobacteria phylum, Gammaproteobacteria class, Enterobacterales order, Enterobacteriaceae family, Escherichia genus, and Escherichia coli species and correspond to the genomic sequence of microorganism 1U001.8 (as shown in Figure 41A) that is in guild 2 of the 141 core microorganisms identified in Table 1.

[0143] Thus, in some embodiments of the methods described herein, as shown in FIG. 41, when the identified genomic construct has at least 97% sequence identity compared to the contigs of the microorganisms provided in the sequence listing, the genome identified by metagenomic analysis is classified as corresponding to the microorganisms listed in Table 1, Table 2, and / or FIGS. 42A - 42XX. In some embodiments, as shown in FIG. 41, when the identified genomic construct has at least 98% sequence identity compared to the contigs of the microorganisms provided in the sequence listing, the genome identified by metagenomic analysis is classified as corresponding to the microorganisms listed in Table 1, Table 2, and / or FIGS. 42A - 42XX. In some embodiments, as shown in FIG. 41, when the identified genomic construct has at least 99% sequence identity compared to the contigs of the microorganisms provided in the sequence listing, the genome identified by metagenomic analysis is classified as corresponding to the microorganisms listed in Table 1, Table 2, and / or FIGS. 42A - 42XX. In some embodiments, as shown in FIG. 41, when the identified genomic construct has at least 99.5% sequence identity compared to the contigs of the microorganisms provided in the sequence listing, the genome identified by metagenomic analysis is classified as corresponding to the microorganisms listed in Table 1, Table 2, and / or FIGS. 42A - 42XX. In some embodiments, the genome identified by metagenomic analysis, as shown in FIG. 41, when the identified genomic construct has at least 97%, at least 97.5%, at least 98%, at least 98.5%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or more sequence identity compared to the contigs of the microorganisms provided in the sequence listing, is classified as corresponding to the microorganisms listed in Table 1, Table 2, and / or FIGS. 42A - 42XX.

[0144] Referring to block 216, in some embodiments, the method includes, for each respective subject in a second plurality of subjects having a second state of a biological characteristic in electronic form, obtaining corresponding plural genomic abundance values including, for each respective gut microorganism among a plurality of gut microorganisms, a corresponding value for the abundance of the genome of each respective gut microorganism in a biological sample derived from the gut of the respective subject. In some embodiments, the second plurality of subjects includes at least 50, at least 100, at least 200, at least 500, at least 1000, at least 5000, at least 10,000, at least 20,000, at least 50,000, at least 100,000, at least 500,000, or at least 1,000,000 subjects. In some embodiments, the second plurality of subjects includes subjects of 1,000,000 or less, 500,000 or less, 100,000 or less, 50,000 or less, 20,000 or less, 10,000 or less, 1000 or less, 500 or less, 100 or less, or 50 or less. In some embodiments, the second plurality of subjects consists of 50 to 100, 50 to 200, 50 to 500, 100 to 500, 200 to 500, 200 to 1000, 500 to 1000, 200 to 5,000, 1000 to 10,000, 5000 to 200,000, 10,000 to 50,000, 20,000 to 100,000, or 500,000 to 1,000,000. In some embodiments, the second plurality of subjects falls within another range starting from subjects of 50 or more and ending with subjects of 10,000,000 or less. In some embodiments, the second plurality of subjects share similar demographic characteristics (such as age, gender, ethnicity, etc.). In some embodiments, the second plurality of subjects share similar physical characteristics (such as weight, height, BMI value, etc.). In some embodiments, the second plurality of subjects share a similar health status (such as physical or mental state, medical history, genetic carrier, or drug use, etc.). In some embodiments, the second plurality of subjects share or are similar in behaviors and lifestyle preferences (such as diet, physical activity, or drug use).

[0145] In some embodiments, the corresponding value to the genomic abundance is a value representing the absolute abundance of the microbial genome. In some embodiments, the corresponding value to the genomic abundance is a value representing a normalized abundance value, or a relative abundance value (e.g., the abundance of one microorganism normalized to the abundance of the total microbiome of interest). In some embodiments, the corresponding value to the genomic abundance is an averaged abundance value (e.g., the average of abundances obtained at different time points, or from different biological samples from a patient, or the average of abundances obtained using different probes, etc.), or a value representing any combination of the above. The corresponding value to the genomic abundance is measured by any technique known in the art. In some embodiments, the genomic abundance value of the genome is measured by quantitative PCR (qPCR), such as bacterial 16S rRNA qPCR, RT-PCR, or qRT-PCR, to quantify the abundance of the region of interest in the genome, as described in, for example, U.S. Patent No. 11,427,865, the entire disclosure of which is incorporated herein by reference. In some embodiments, the genomic abundance value is measured by targeted sequencing (e.g., 16S rRNA sequencing, or any other suitable biomarker), partial genomic sequencing, or whole-genomic sequencing, thereby quantifying the number of reads of the targeted region within the microbial genome to determine the genomic abundance, as disclosed in, for example, U.S. Patent Application Publication No. 2021 / 0403986 or U.S. Patent No. 11,332,783, the entire disclosures of which are incorporated herein by reference. In some embodiments, deep sequencing is used to determine the abundance of the targeted sequence, as disclosed in, for example, U.S. Patent Application Publication No. 2018 / 0237863, the disclosure of which is incorporated herein by reference in its entirety.In some embodiments, the sequencing depth is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 70, 80, 90, 100, 110, 120, 130, 150, 200, 300, 500, 500, 700, 1000, or more. In some embodiments, shotgun metagenomic sequencing is used to provide sequence reads of the genome in a sample, such as as described in U.S. Patent No. 11,028,449, the contents of which are hereby incorporated by reference in their entirety.

[0146] Referring to block 218, in some embodiments, for each respective object in the second plurality of objects, the biological sample derived from the intestine of each respective object is a fecal sample. In some embodiments, the biological sample is selected from a tissue biopsy, intestine, or mucosal sample.

[0147] Referring to block 220, in some embodiments, the method comprises sequencing genomic DNA from a corresponding biological sample derived from the intestine of each respective target among a second plurality of targets, thereby obtaining a corresponding second plurality of (e.g., at least 100,000) nucleic acid sequences. In some embodiments, the second plurality of nucleic acid sequences comprises at least 100,000, at least 250,000, at least 500,000, at least 1,000,000, at least 2,500,000, at least 5,000,000, at least 10,000,000, or at least 50,000,000 nucleic acid sequences. In some embodiments, the second plurality of nucleic acid sequences comprises nucleic acid sequences of targets that are 250,000,000 or fewer, 100,000,000 or fewer, 50,000,000 or fewer, 25,000,000 or fewer, 10,000,000 or fewer, 5,000,000 or fewer, 1,000,000 or fewer, 100,000 or fewer. In some embodiments, the second plurality of nucleic acid sequences consists of 100,000 to 1,000,000, 200,000 to 5,000,000, 500,000 to 10,000,000, 1,000,000 to 20,000,000, 5,000,000 to 50,000,000, 10,000,000 to 100,000,000, or 50,000,000 to 250,000,000 nucleic acid sequences. In some embodiments, the second plurality of nucleic acid sequences falls within another range that starts with at least 100,000 nucleic acid sequences and ends with 250,000,000 or fewer nucleic acid sequences.

[0148] In some embodiments, the second plurality (e.g., at least 100,000) of nucleic acid sequences are obtained by metagenomic sequencing, such as disclosed in, for example, U.S. Patent Application Publication No. 2016 / 0239602 or U.S. Patent No. 11,495,326, the contents of which are hereby incorporated by reference in their entirety. In some embodiments, metagenomic sequencing further includes generating a plurality of metagenomic fragment reads. In some embodiments, metagenomic sequencing further includes fragmenting the microbial genome into randomly sized fragments. The resulting fragments can vary in size. In one embodiment, fragments of approximately 500 nucleotides can be obtained. In some embodiments, fragments of 100 to 2000 nucleotides, such as 200 to 800 nucleotides, 100 to 900 nucleotides, 100 to 1000 nucleotides, 300 to 800 nucleotides, 400 to 900 nucleotides can be obtained. In some embodiments, the method can further include extracting metagenomic fragments from a corresponding biological sample. In some embodiments, metagenomic sequencing further includes sequencing the fragments using a high-throughput sequencing method to generate a plurality of sequencing reads.

[0149] In some embodiments, the second plurality (e.g., at least 100,000) of nucleic acid sequences are obtained by targeted panel sequencing, as described, for example, in U.S. Patent Application Publication No. 2019 / 0316209. In some embodiments, targeted panel sequencing comprises hybridizing genomic DNA isolated from a biological sample derived from the gut of a subject to a panel of probes that hybridize to unique sequences within the genome of each of a plurality of microorganisms, such as each of the microorganisms listed in Tables 1, 2, and / or FIGS. 42A - 42XX, prior to sequencing the recovered nucleic acids. In some embodiments, a combination of quasi-unique sequences (e.g., sequences found in a small number of microbial genomes) can be used to back-calculate genomic abundance values using an algorithm, such as a system of equations. In some embodiments, the panel of probes comprises at least one probe that hybridizes to a unique sequence for each detected microbial genome. In some embodiments, the panel of probes comprises at least 2, at least 3, at least 4, at least 5, at least 10, at least 25, at least 50, or more probes that hybridize to different unique sequences for each detected microbial genome. In some embodiments, the panel of probes comprises at least 20, at least 30, at least 40, at least 50, at least 75, at least 100, at least 125, at least 150, at least 200, at least 150, at least 300, at least 400, at least 500, at least 750, at least 1000, at least 1250, at least 1500, at least 2000, at least 2500, at least 3000, at least 4000, at least 5000, at least 7500, at least 10,000, or more unique probes.

[0150] In some embodiments, the genomic DNA for sequencing from a corresponding biological sample may include at its ends partial or complete sequencing platform adapter sequences useful for sequencing using the targeted sequencing platform. Examples of the targeted sequencing platform include, but are not limited to, the Illumina® HiSeq®, MiSeq® and Genome Analyzer® sequencing systems, the Ion Torrent® Ion PGM® and Ion Proton® sequencing systems, the Pacific Biosciences PACBIO RS II Sequel system, the Life Technologies® SOLiD sequencing system, the Roche 454GS FLX+ and GS Junior sequencing systems, the Oxford Nanopore MinION® system, or any other sequencing platform of interest.

[0151] Referring to block 222, in some embodiments, the method comprises, for each respective object in the second plurality of objects, in electronic form, obtaining a corresponding second plurality of (e.g., at least 100,000) nucleic acid sequences for genomic DNA from a corresponding biological sample derived from the intestine of the respective object; and for each respective object in the second plurality of objects, determining a corresponding genomic abundance value for each respective gut microbe among the plurality of gut microbes from the corresponding second plurality of (e.g., at least 100,000) nucleic acid sequences. In some embodiments, the genomic abundance values determined for each respective object in the second plurality of objects include genomic abundance values of at least 20, at least 25, at least 50, at least 100, at least 250, at least 500, at least 1,000, at least 5,000, or at least 10,000. In some embodiments, the genomic abundance values determined for each respective control in the second plurality of objects include genomic abundance values of 250,000 or less, 100,000 or less, 50,000 or less, 25,000 or less, 10,000 or less, 5,000 or less, 1,000 or less, 100 or less, 50 or less, 30 or less, or 20 or less. In some embodiments, the genomic abundance values determined for each respective object in the second plurality of objects consist of genomic abundance values of 10-40, 20-50, 30-80, 40-100, 50-150, 60-200, 80-300, 90-500, 100-1000, 500-2,000, or 1,000-5,000. In some embodiments, the genomic abundance values determined for each respective object in the second plurality of objects fall within another range starting with a genomic abundance value of 20 or less and ending with a genomic abundance value of 250,000 or less.

[0152] Referring to block 224, in some embodiments, the method includes, for each respective target in the second plurality of targets, assembling a corresponding second plurality of gut microbial genomes by metagenomic de novo sequence assembly from a corresponding second plurality of at least 100,000 nucleic acid sequences, and for each respective gut microbial genome in the corresponding second plurality of gut microbial genomes, calculating a corresponding genomic abundance of each respective gut microbial genome. In some embodiments, the metagenomic de novo sequence assembly further includes generating contigs based on sequencing reads generated by shotgun sequencing technology, such as that described in U.S. Patent No. 10,529,443. In some embodiments, the second plurality of (e.g., at least 100,000) nucleic acid sequences can be assembled into the whole genomes of a plurality of gut microbes. In some embodiments, the second plurality of (e.g., at least 100,000) nucleic acid sequences can be assembled into partial genomes of a plurality of gut microbes.

[0153] Referring to block 226, in some embodiments, the method comprises, for each respective target among the second plurality of targets, assigning each respective nucleic acid sequence among the corresponding second plurality of at least 100,000 sequences to a respective gut microorganism among the plurality of gut microorganisms, thereby generating, for each respective gut microorganism among the plurality of gut microorganisms, a corresponding count of each respective nucleic acid sequence among the corresponding second plurality of nucleic acid sequences assigned to the respective gut microorganism; and determining, for each respective gut microorganism among the plurality of gut microorganisms, a corresponding genomic abundance value for the respective gut microorganism based on the corresponding count of each respective nucleic acid sequence assigned to the respective gut microorganism. In some embodiments, assigning each respective nucleic acid to a respective gut microorganism comprises mapping the nucleic acid to a reference nucleic acid, e.g., a contig listed in FIG. 41. In some embodiments, assigning each respective nucleic acid to a respective gut microorganism comprises annotating genomic information based on an existing database. In some embodiments, the nucleic acid sequences are analyzed and the annotation defines taxonomic assignments using sequence similarity methods and phylogenetic placement methods, or a combination of the two strategies.

[0154] Array similarity-based methods include, but are not limited to, those well known to those skilled in the art, including various implementations of these algorithms such as BLAST, BLASTx, tBLASTn, tBLASTx, RDP classifier, DNAclust, and Qiime or Mothur. These methods rely on mapping sequence reads to a reference database and selecting the matches with the best scores and e-values. In some embodiments, phylogenetic methods are used in combination with sequence similarity methods to improve the call accuracy of annotation or taxonomic assignment. Common databases include, but are not limited to, GT-DBTK, National Center for Biotechnology Information (NCBI) Genbank, European Bioinformatics Institute - European Nucleotide Archive (European Bioinformatics Institute - European Nucleotide Archive; EBI-ENA), National Institute of Genetics, U.S. Department of ENERGY (USDOE) Integrated Microbial Genomes (Integrated Microbial Genomes) & Microbiomes; IMG / M), and other databases available in the art.

[0155] Referring to block 228, in some embodiments, the second state of the biological characteristic is a disease or disorder, such as type 2 diabetes (T2D), hypertension (HT), schizophrenia (SCZ), atherosclerotic cardiovascular disease (ACVD), cirrhosis (LC), inflammatory bowel disease (IBD), colorectal cancer (CRC), ankylosing spondylitis (AS), and Parkinson's disease (PD), multiple sclerosis (MS), type 2 Gaucher disease (GDII), COVID-19 (COV), Behçet's disease (BD), autism spectrum disorder (ASD), or pancreatic cancer (PC). In some embodiments, the disease or disorder is cancer, Alzheimer's disease, cardiovascular disease, autoimmune disease, mental health disease, infectious disease, or genetic disorder.

[0156] In some embodiments, the second state of the biological characteristic is the second severity of a disease or disorder. In some embodiments, the severity of a disease is classified by the type, frequency or intensity experienced by the subject. In some embodiments, the severity of a disease is classified by the progression or prognosis of the disease or disorder, such as the different stages of cancer. In some embodiments, the second state of the biological characteristic is a treated disease or disorder, and the second state of the biological characteristic is a disease or disorder treated with a second therapy, such as surgery, radiation therapy, chemotherapy, targeted therapy, gene therapy, immunotherapy, drug therapy, dietary modification, lifestyle modification. In some embodiments, the second state of the biological characteristic is the second level of nutrients in the diet, such as carbohydrates, proteins, fats, vitamins, fiber. In some embodiments, the second state of the biological characteristic is the second age. In some embodiments, a threshold is provided for determining the second state of the biological characteristic, such as the level of a biomarker, a diagnostic cut-off value, or a threshold nutrient intake level.

[0157] Referring to block 230, in some embodiments, the plurality of gut microbiota includes at least 20 gut microbiota selected from Table 1, Table 2, or FIGS. 42A - 42XX. In some embodiments, the plurality of gut microbiota includes at least 25 gut microbiota selected from Table 1, Table 2, or FIGS. 42A - 42XX. In some embodiments, the plurality of gut microbiota includes at least 30 gut microbiota selected from Table 1, Table 2, or FIGS. 42A - 42XX. In some embodiments, the plurality of gut microbiota includes at least 40 gut microbiota selected from Table 1, Table 2, or FIGS. 42A - 42XX. In some embodiments, the plurality of gut microbiota includes at least 25 gut microbiota selected from Table 1, Table 2, or FIGS. 42A - 42XX. In some embodiments, the plurality of gut microbiota includes at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 80, at least 85, at least 90, at least 95, at least 100, at least 125, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, or all of the gut microbiota selected from Table 1, Table 2, or FIGS. 42A - 42XX. In some embodiments, the plurality of gut microbiota is at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 125, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, or all of the gut microbiota selected from Table 1, Table 2, or FIGS. 42A - 42XX.In some embodiments, the plurality of gut microbes are all the gut microbes listed in Table 1. In some embodiments, the plurality of gut microbes are all the gut microbes listed in Table 2. In some embodiments, the plurality of gut microbes are all the gut microbes listed in FIGS. 42A-42XX.

[0158] Referring to block 232, in some embodiments, the method is to calculate a first plurality of similarity metrics from corresponding plurality of genomic abundance values across a first plurality of subjects, the first plurality of similarity metrics including a first corresponding similarity metric for each unique pair of gut microbes among the plurality of gut microbes, the first corresponding similarity metric quantifying the similarity between (i) a corresponding first vector formed by the corresponding genomic abundance values of the first microbe in a unique pair of gut microbes across the first plurality of subjects and (ii) a corresponding second vector formed by the corresponding genomic abundance values of the second microbe in the unique pair of gut microbes across the first plurality of subjects. In some embodiments, the corresponding first vector is a set of values where each value represents the genomic abundance value of the first microbe for one subject among the first plurality of subjects. In some embodiments, the corresponding second vector is a set of values where each value represents the genomic abundance value of the second microbe for one subject among the first plurality of subjects. In some embodiments, a unique genomic pair is formed by any two genomes detected across the first plurality of subjects. In some embodiments, the number of unique pairs relative to the total number of N genomes can be calculated by N(N - 1) / 2, where N represents the non - repeating number of genomes detected across the first plurality of subjects. As the number of gut microbes in the set increases, the number of calculations required to determine the set of all similarity metrics increases as a quadratic function of the number of microbes.

[0159] Referring to block 234, in some embodiments, the method comprises calculating a second plurality of similarity metrics using corresponding genomic abundance values for a second plurality of subjects, the second plurality of similarity metrics including a second corresponding similarity metric for each unique pair of gut microbiota among the plurality of gut microbiota, the second corresponding similarity metric quantifying the similarity between (i) a corresponding second vector formed by the corresponding genomic abundance value of a first microbiota in a unique pair of gut microbiota across the second plurality of subjects and (ii) a corresponding second vector formed by the corresponding genomic abundance value of a second microbiota in the unique pair of gut microbiota across the second plurality of subjects. In some embodiments, the corresponding second vector is a set of values, each value representing the genomic abundance value of the second microbiota for one subject among the second plurality of subjects. In some embodiments, the corresponding second vector is a set of values, each value representing the genomic abundance value of the second microbiota for one subject among the second plurality of subjects. In some embodiments, a unique genomic pair is formed by any two genomes detected across the second plurality of subjects. In some embodiments, the number of unique pairs relative to the total number of N genomes can be calculated by N(N - 1) / 2, where N represents the non - repeating number of genomes detected across the first plurality of subjects. As the number of gut microbiota in the set increases, the number of calculations required to determine the set of all similarity metrics increases as a quadratic function of the number of microbiota.

[0160] Referring to block 236, in some embodiments, determining a set of unique pairs of gut microbiota among a plurality of gut microbiota based on a first plurality of similarity metrics and a second plurality of similarity metrics, wherein both the first corresponding similarity metric and the second corresponding similarity metric show a statistically significant positive correlation between the abundance of the first gut microbiota and the abundance of the second gut microbiota in the unique pair of each gut microbiota, or both the first corresponding similarity metric and the second corresponding similarity metric show a statistically significant negative correlation between the abundance of the first gut microbiota and the abundance of the second gut microbiota in the unique pair of each gut microbiota. In some embodiments, the similarity metric is a correlation coefficient. In some embodiments, a threshold or cut-off value for a statistically significant positive correlation is defined. In some embodiments, a threshold or cut-off value for a statistically significant negative correlation is defined.

[0161] Referring to block 238, in some embodiments, one or both of the first corresponding similarity metric and the second similarity metric can be a Pearson correlation coefficient, an intraclass correlation coefficient, or a rank correlation coefficient. In some embodiments, the similarity metric is a Spearman's correlation coefficient or a maximum information coefficient (MIC). In some embodiments, the similarity metric is a rank correlation coefficient also known as Kendall's tau used to measure the association between two scales. In some embodiments, the similarity metric is calculated by a SparCC-based algorithm for compositional data or a SPIEC-EASI-based algorithm for ecological association inference. In some embodiments, the SparCC-based algorithm is FastSpar.

[0162] Referring to block 240, in some embodiments, a statistically significant positive correlation has a P-value of less than 0.001. In some embodiments, a statistically significant positive correlation has a P-value of less than 0.05. In some embodiments, a statistically significant positive correlation has a P-value of less than 0.01. In some embodiments, a statistically significant positive correlation has a P-value of less than 0.001, less than 0.005, less than 0.01, less than 0.025, less than 0.05, or less than 0.075.

[0163] Referring to block 242, in some embodiments, the method includes identifying a set of gut microbiota that includes each gut microbiota represented by a unique pair of gut microbiota. In some embodiments, a microbiome network is constructed that includes each gut microbiota represented by a unique pair of gut microbiota. In some embodiments, the network is visualized by bioinformatics software, such as Cytoscape.

[0164] Referring to block 244, in some embodiments, the method is to cluster each gut microbiota represented by a unique pair of gut microbiota into one of a plurality of networks, where each connected network includes a corresponding set of a plurality of nodes and one or more corresponding edges.

[0165] Referring to block 246, in some embodiments, each node among the corresponding plurality of nodes represents a unique gut microbiota represented by a unique pair of gut microbiota. In some embodiments, the node size represents the average abundance of the genome. In some embodiments, the link between nodes is treated as a metal spring attached to a pair of nodes. In some embodiments, a similarity metric is used to determine the repulsive and attractive forces of the spring. In some embodiments, the value of the similarity matrix is used to determine the weight of the link.

[0166] Referring to block 248, in some embodiments, each respective edge in a corresponding set of one or more edges connects two nodes representing each respective unique pair of gut microbes in a set of unique pairs of gut microbes. In some embodiments, a unique pair of genomes that are positively correlated is differentiated from a unique pair of genomes that are negatively correlated.

[0167] Referring to block 250, in some embodiments, each respective node in a corresponding plurality of nodes is connected to at least one other respective node in the plurality of nodes via each respective edge in a corresponding set of one or more edges.

[0168] Referring to block 252, in some embodiments, the method includes identifying each respective network in one or more networks that contains the most nodes, thereby identifying a set of gut microbes represented by the corresponding plurality of nodes in each respective network. In some embodiments, nodes that are not connected to one or more networks are removed.

[0169] Referring to block 254, in some embodiments, the identified set of gut microbes includes all respective gut microbes represented by a set of unique pairs of gut microbes. In some embodiments, the identified set of gut microbes includes all respective gut microbes represented by nodes of one or more microbiome networks.

[0170] Referring to block 256, in some such embodiments, the identified set of gut microbiota comprises at least 20 gut microbiota from Table 1, Table 2, or Figures 42A - 42XX. In some embodiments, the plurality of gut microbiota comprises at least 25 gut microbiota selected from Table 1, Table 2, or Figures 42A - 42XX. In some embodiments, the plurality of gut microbiota comprises at least 30 gut microbiota selected from Table 1, Table 2, or Figures 42A - 42XX. In some embodiments, the plurality of gut microbiota comprises at least 40 gut microbiota selected from Table 1, Table 2, or Figures 42A - 42XX. In some embodiments, the plurality of gut microbiota comprises at least 25 gut microbiota selected from Table 1, Table 2, or Figures 42A - 42XX. In some embodiments, the plurality of gut microbiota comprises at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 125, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, or all of the gut microbiota selected from Table 1, Table 2, or Figures 42A - 42XX. In some embodiments, the plurality of gut microbiota is at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 125, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, or all of the gut microbiota selected from Table 1, Table 2, or Figures 42A - 42XX.In some embodiments, the plurality of gut microorganisms are all the gut microorganisms listed in Table 1. In some embodiments, the plurality of gut microorganisms are all the gut microorganisms listed in Table 2. In some embodiments, the plurality of gut microorganisms are all the gut microorganisms listed in FIGS. 42A-42XX.

[0171] In some embodiments, the identified set of gut microorganisms is selected from the microorganisms in Table 5 having at least 2, at least 3, at least 4, at least 5, at least 10, at least 15, at least 20, at least 25, or more binding properties. Referring to block 308, in some embodiments, the plurality of gut microorganisms includes at least 20 microorganisms selected from the microorganisms listed in Table 5 as having at least 2 binding properties. In some embodiments, the plurality of gut microorganisms includes at least 20 microorganisms selected from those microorganisms listed in Table 5 as having at least 3, at least 4, at least 5, at least 10, at least 15, at least 20, at least 25, or more binding properties.

[0172] 2. Method for training a model for evaluating human health FIG. 3 is a schematic diagram of a method for training a model for evaluating human health, as described below. Method 300 can be implemented using a computer system (e.g., computer system 100 shown and described above with reference to FIG. 1).

[0173] Referring to block 300, in some embodiments, the method includes, for each respective training subject among a plurality of training subjects in electronic form, (i) for each respective gut microorganism among the plurality of gut microorganisms, a corresponding plurality of genomic abundance values including a corresponding value for the abundance of the genome of the respective gut microorganism in a corresponding biological sample derived from the gut of the respective training subject, and (ii) the corresponding state of the biological characteristics of the respective training subject.

[0174] In some embodiments, the plurality of training subjects includes at least 50, at least 100, at least 200, at least 500, at least 1000, at least 5000, at least 10,000, at least 20,000, at least 50,000, at least 100,000, at least 500,000, or at least 1,000,000 training subjects. In some embodiments, the plurality of subjects includes subjects of 1,000,000 or less, 500,000 or less, 100,000 or less, 50,000 or less, 20,000 or less, 10,000 or less, 1000 or less, 500 or less, 100 or less, or 50 or less. In some embodiments, the plurality of training subjects consists of 50 to 100, 50 to 200, 50 to 500, 100 to 500, 200 to 500, 200 to 1000, 500 to 1000, 200 to 5,000, 1000 to 10,000, 5000 to 200,000, 10,000 to 50,000, 20,000 to 100,000, or 500,000 to 1,000,000. In some embodiments, the plurality of training subjects falls within another range starting from 50 subjects or more and ending at 10,000,000 subjects or less. In some embodiments, the plurality of training subjects share similar demographic characteristics (such as age, gender, ethnicity, etc.). In some embodiments, the plurality of subjects share similar physical characteristics (such as weight, height, BMI value, etc.). In some embodiments, the plurality of training subjects share similar health conditions (such as physical or mental state, medical history, genetic carrier, or drug use, etc.). In some embodiments, the plurality of subjects share or have similar behaviors and lifestyle preferences (such as diet, physical exercise, or drug use).

[0175] In some embodiments, the corresponding value to the genomic abundance is a value representing the absolute abundance of the microbial genome. In some embodiments, the corresponding value to the genomic abundance is a value representing a normalized abundance value or a relative abundance value (e.g., the abundance of one microorganism normalized to the abundance of the total microbiome of interest). In some embodiments, the corresponding value to the genomic abundance is a value representing an averaged abundance value (e.g., the average of abundances obtained at different time points or from different biological samples from a patient, or the average of abundances obtained using different probes), or any combination of the above. The corresponding value to the genomic abundance is measured by any technique known in the art. In some embodiments, the value of the genomic abundance of the genome is measured by quantitative PCR (qPCR), such as bacterial 16S rRNA qPCR, RT-PCR, or qRT-PCR, to quantify the abundance of the region of interest in the genome, as described, for example, in U.S. Patent No. 11,427,865, the entire disclosure of which is incorporated herein by reference. In some embodiments, the genomic abundance value is measured by targeted sequencing (e.g., 16S rRNA sequencing or any other suitable biomarker), partial genome sequencing, or whole genome sequencing, thereby quantifying the number of reads of the targeted region within the microbial genome to determine the genomic abundance, as disclosed, for example, in U.S. Patent Application Publication No. 2021 / 0403986 or U.S. Patent No. 11,332,783, the entire disclosures of which are incorporated herein by reference. In some embodiments, deep sequencing is used to determine the abundance of the targeted sequence, as disclosed, for example, in U.S. Patent Application Publication No. 2018 / 0237863, the disclosure of which is incorporated herein by reference in its entirety.In some embodiments, the sequencing depth is at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 70, 80, 90, 100, 110, 120, 130, 150, 200, 300, 500, 500, 700, 1000, or more. In some embodiments, shotgun metagenome sequencing is used to provide genomic sequence reads in a sample, such as described in U.S. Patent No. 11,028,449, the contents of which are incorporated herein by reference in their entirety.

[0176] Referring to block 302, in some embodiments, the method comprises sequencing genomic DNA from a corresponding biological sample derived from the intestine of each respective one of a plurality of training subjects, thereby obtaining a corresponding plurality of (e.g., at least 100,000) nucleic acid sequences. In some embodiments, the plurality of nucleic acid sequences comprises at least 100,000, at least 250,000, at least 500,000, at least 1,000,000, at least 2,500,000, at least 5,000,000, at least 10,000,000, or at least 50,000,000 nucleic acid sequences. In some embodiments, the plurality of nucleic acid sequences comprises 250,000,000 or fewer, 100,000,000 or fewer, 50,000,000 or fewer, 25,000,000 or fewer, 10,000,000 or fewer, 5,000,000 or fewer, 1,000,000 or fewer, 100,000 or fewer nucleic acid sequences. In some embodiments, the plurality of nucleic acid sequences consists of 100,000 to 1,000,000, 200,000 to 5,000,000, 500,000 to 10,000,000, 1,000,000 to 20,000,000, 5,000,000 to 50,000,000, 10,000,000 to 100,000,000, or 50,000,000 to 250,000,000 nucleic acid sequences. In some embodiments, the plurality of nucleic acid sequences falls within another range starting with at least 100,000 nucleic acid sequences and ending with 250,000,000 or fewer nucleic acid sequences.

[0177] In some embodiments, a plurality of (e.g., at least 100,000) nucleic acid sequences are obtained by metagenomic sequencing, such as disclosed in U.S. Patent Application Publication No. 2016 / 0239602 or U.S. Patent No. 11,495,326, the contents of which are hereby incorporated by reference in their entirety. In some embodiments, metagenomic sequencing further includes generating a plurality of metagenomic fragment reads. In some embodiments, metagenomic sequencing further includes fragmenting the microbial genome into randomly sized fragments. The resulting fragments can vary in size. In one embodiment, fragments of approximately 500 nucleotides can be obtained. In some embodiments, fragments of 100 to 2000 nucleotides, such as 200 to 800 nucleotides, 100 to 900 nucleotides, 100 to 1000 nucleotides, 300 to 800 nucleotides, 400 to 900 nucleotides can be obtained. In some embodiments, the method can further include extracting metagenomic fragments from a corresponding biological sample. In some embodiments, metagenomic sequencing further includes sequencing the fragments using a high-throughput sequencing method to generate a plurality of sequencing reads.

[0178] In some embodiments, a plurality of (e.g., at least 100,000) nucleic acid sequences are obtained by targeted panel sequencing, as described, for example, in U.S. Patent Application Publication No. 2019 / 0316209. In some embodiments, targeted panel sequencing hybridizes genomic DNA isolated from a biological sample derived from the intestine of a subject to a panel of probes that includes one or more probes that hybridize to unique sequences within the genome of each of a plurality of microorganisms, such as each of the microorganisms listed in Tables 1, 2, and / or FIGS. 42A-42XX, prior to sequencing the recovered nucleic acids. In some embodiments, a combination of quasi-unique sequences (e.g., sequences found in a small number of microbial genomes) can be used to back-calculate genomic abundance values using an algorithm, such as a system of equations. In some embodiments, the panel of probes includes at least one probe that hybridizes to a sequence unique to each detected microbial genome. In some embodiments, the panel of probes includes at least 2, at least 3, at least 4, at least 5, at least 10, at least 25, at least 50, or more probes that hybridize to different sequences unique to each detected microbial genome. In some embodiments, the panel of probes includes at least 20, at least 30, at least 40, at least 50, at least 75, at least 100, at least 125, at least 150, at least 200, at least 150, at least 300, at least 400, at least 500, at least 750, at least 1000, at least 1250, at least 1500, at least 2000, at least 2500, at least 3000, at least 4000, at least 5000, at least 7500, at least 10,000, or more unique probes.

[0179] In some embodiments, the genomic DNA sequenced from the corresponding biological sample includes at its ends partial or complete sequencing platform adapter sequences useful for sequencing using the targeted sequencing platform. Examples of targeted sequencing platforms include, but are not limited to, the Illumina® HiSeq®, MiSeq®, and Genome Analyzer® sequencing systems, the Ion Torrent® Ion PGM® and Ion Proton® sequencing systems, the Pacific Biosciences PACBIO RS II Sequel system, the Life Technologies® SOLiD sequencing system, the Roche 454GS FLX+ and GS Junior sequencing systems, the Oxford Nanopore MinION® system, or any other targeted sequencing platform.

[0180] Referring to block 304, in some embodiments, the biological sample derived from the intestine of each subject to be trained is a fecal sample from each subject to be trained. In some embodiments, the sample is a tissue biopsy, intestine, or mucosal sample. See Tang Q, et al., Current Sampling Methods for Gut Microbiota: A Call for More Precise Devices, Front Cell Infect Microbiol., 10:151 (2020), the contents of which are incorporated herein by reference in their entirety.

[0181] Referring to block 306, in some embodiments, the plurality of gut microbiota includes at least 20 gut microbiota selected from Table 1, Table 2, or FIGS. 42A - 42XX. In some embodiments, the plurality of gut microbiota includes at least 25 gut microbiota selected from Table 1, Table 2, or FIGS. 42A - 42XX. In some embodiments, the plurality of gut microbiota includes at least 30 gut microbiota selected from Table 1, Table 2, or FIGS. 42A - 42XX. In some embodiments, the plurality of gut microbiota includes at least 40 gut microbiota selected from Table 1, Table 2, or FIGS. 42A - 42XX. In some embodiments, the plurality of gut microbiota includes at least 25 gut microbiota selected from Table 1, Table 2, or FIGS. 42A - 42XX. In some embodiments, the plurality of gut microbiota includes at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 125, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, or all of the gut microbiota selected from Table 1, Table 2, or FIGS. 42A - 42XX. In some embodiments, the plurality of gut microbiota is at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 125, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, or all of the gut microbiota selected from Table 1, Table 2, or FIGS. 42A - 42XX.In some embodiments, the plurality of gut microbiota are all the gut microbiota listed in Table 1. In some embodiments, the plurality of gut microbiota are all the gut microbiota listed in Table 2. In some embodiments, the plurality of gut microbiota are all the gut microbiota listed in FIGS. 42A to 42XX.

[0182] In some embodiments of the methods described herein, as shown in FIG. 41, if the identified genomic construct has at least 97% sequence identity compared to the contigs of the microorganisms provided in the sequence listing, the genome identified by metagenomic analysis is classified as corresponding to the microorganisms listed in Table 1, Table 2, and / or FIGS. 42A - 42XX. In some embodiments, as shown in FIG. 41, if the identified genomic construct has at least 98% sequence identity compared to the contigs of the microorganisms provided in the sequence listing, the genome identified by metagenomic analysis is classified as corresponding to the microorganisms listed in Table 1, Table 2, and / or FIGS. 42A - 42XX. In some embodiments, as shown in FIG. 41, if the identified genomic construct has at least 99% sequence identity compared to the contigs of the microorganisms provided in the sequence listing, the genome identified by metagenomic analysis is classified as corresponding to the microorganisms listed in Table 1, Table 2, and / or FIGS. 42A - 42XX. In some embodiments, as shown in FIG. 41, if the identified genomic construct has at least 99.5% sequence identity compared to the contigs of the microorganisms provided in the sequence listing, the genome identified by metagenomic analysis is classified as corresponding to the microorganisms listed in Table 1, Table 2, and / or FIGS. 42A - 42XX. In some embodiments, the genome identified by metagenomic analysis, as shown in FIG. 41, is such that the identified genomic construct has at least 97%, at least 97.5%, at least 98%, at least 98.5%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or more sequence identity compared to the contigs of the microorganisms provided in the sequence listing, and is classified as corresponding to the microorganisms listed in Table 1, Table 2, and / or FIGS. 42A - 42XX.

[0183] In some embodiments, the plurality of gut microorganisms are selected from the microorganisms in Table 5 having at least 2, at least 3, at least 4, at least 5, at least 10, at least 15, at least 20, at least 25, or more binding properties. Referring to block 308, in some embodiments, the plurality of gut microorganisms includes at least 20 microorganisms selected from the microorganisms listed in Table 5 as having at least 2 binding properties. In some embodiments, the plurality of gut microorganisms includes at least 20 microorganisms selected from those microorganisms listed in Table 5 as having at least 3, at least 4, at least 5, at least 10, at least 15, at least 20, at least 25, or more binding properties.

[0184] Referring to block 310, in some embodiments, the method includes, for each respective training subject among a plurality of training subjects, obtaining, in electronic form, a corresponding plurality of at least 100,000 nucleic acid sequences for genomic DNA from a corresponding biological sample derived from the intestine of each respective training subject; and for each respective intestinal microorganism among a plurality of intestinal microorganisms, determining, from the corresponding first plurality of (e.g., at least 100,000) nucleic acid sequences, a corresponding value for the abundance of the genome of each respective intestinal microorganism. In some embodiments, the genomic abundance values determined for each respective subject among the plurality of training subjects include genomic abundance values of at least 20, at least 25, at least 50, at least 100, at least 250, at least 500, at least 1,000, at least 5,000, or at least 10,000. In some embodiments, the genomic abundance values determined for each respective subject among the plurality of training subjects include genomic abundance values of 250,000 or less, 100,000 or less, 50,000 or less, 25,000 or less, 10,000 or less, 5,000 or less, 1,000 or less, 100 or less, 50 or less, 30 or less, or 20 or less. In some embodiments, the genomic abundance values determined for each respective subject among the plurality of training subjects consist of genomic abundance values of 10 - 40, 20 - 50, 30 - 80, 40 - 100, 50 - 150, 60 - 200, 80 - 300, 90 - 500, 100 - 1000, 500 - 2,000, or 1,000 - 5,000. In some embodiments, the genomic abundance values determined for each respective subject among the plurality of training subjects fall within another range that starts with a genomic abundance value of 20 or less and ends with a genomic abundance value of 250,000 or less.

[0185] Referring to block 312, in some embodiments, the method, for each respective training subject among a plurality of training subjects, electronically assembles a corresponding plurality of gut microbial genomes by metagenomic de novo sequence assembly from a corresponding plurality of (e.g., at least 100,000) nucleic acid sequences, and for each respective gut microbe among the plurality of gut microbes, calculates a corresponding value for the abundance of the genome of each respective gut microbe based on the morbidity rate of each respective nucleic acid sequence among the plurality of at least 100,000 nucleic acid sequences used to assemble each respective gut microbial genome among the corresponding plurality of gut microbial genomes. In some embodiments, the metagenomic de novo sequence assembly further includes generating contigs based on sequencing reads generated by shotgun sequencing technology, as described in U.S. Patent No. 10,529,443, the contents of which are hereby incorporated by reference in their entirety. In some embodiments, the first plurality of (e.g., at least 100,000) nucleic acid sequences can be assembled into the entire genomes of the plurality of gut microbes. In some embodiments, the first plurality of (e.g., at least 100,000) nucleic acid sequences can be assembled into partial genomes of the plurality of gut microbes.

[0186] Referring to block 314, in some embodiments, the method includes, for each respective target among a plurality of training targets, assigning each respective nucleic acid sequence among a corresponding plurality of (e.g., at least 100,000) sequences to each respective gut microorganism among a plurality of gut microorganisms, thereby generating, for each respective gut microorganism among the plurality of gut microorganisms, a corresponding count of each respective nucleic acid sequence among the corresponding plurality of nucleic acid sequences assigned to the respective gut microorganism; and determining, for each respective gut microorganism among the plurality of gut microorganisms, a corresponding genomic abundance value for the respective gut microorganism based on the corresponding count of each respective nucleic acid sequence assigned to the respective gut microorganism. In some embodiments, assigning each respective nucleic acid to each respective gut microorganism includes mapping the nucleic acid to a reference nucleic acid, e.g., a contig listed in FIG. 41. In some embodiments, assigning each respective nucleic acid to each respective gut microorganism includes annotating genomic information based on an existing database. In some embodiments, the nucleic acid sequences are analyzed and the annotation defines taxonomic assignments using sequence similarity methods and phylogenetic placement methods, or a combination of the two strategies.

[0187] Array similarity-based methods include, but are not limited to, those well-known to those skilled in the art, including various implementations of these algorithms such as BLAST, BLASTx, tBLASTn, tBLASTx, RDP classifier, DNAclust, and Qiime or Mothur. These methods rely on mapping sequence reads to a reference database and selecting the best-scoring and e-value matches. In some embodiments, phylogenetic methods are used in combination with sequence similarity methods to improve the call accuracy of annotation or taxonomic assignment. Common databases include, but are not limited to, GT-DBTK, National Center for Biotechnology Information (NCBI) Genbank, European Bioinformatics Institute - European Nucleotide Archive (European Bioinformatics Institute - European Nucleotide Archive; EBI-ENA), National Institute of Genetics, U.S. Department of ENERGY (USDOE) Integrated Microbial Genomes (Integrated Microbial Genomes) & Microbiomes; IMG / M), and other databases available in the art.

[0188] Referring to block 316, in some embodiments, the biological characteristics are a disease or disorder, a therapy administered to the subject, such as surgery, radiation therapy, chemotherapy, targeted therapy, gene therapy, immunotherapy, drug therapy, dietary changes, lifestyle changes, or the subject's diet, such as a diet rich or poor in carbohydrates, proteins, fats, vitamins, or fiber.

[0189] Referring to block 318, in some embodiments, the disease or disorder is selected from the group consisting of type 2 diabetes (T2D), hypertension (HT), schizophrenia (SCZ), atherosclerotic cardiovascular disease (ACVD), cirrhosis (LC), inflammatory bowel disease (IBD), colorectal cancer (CRC), ankylosing spondylitis (AS), and Parkinson's disease (PD), multiple sclerosis (MS), type 2 Gaucher disease (GDII), COVID-19 (COV), Behçet's disease (BD), autism spectrum disorder (ASD), or pancreatic cancer (PC). In some embodiments, the disease or disorder is cancer, Alzheimer's disease, cardiovascular disease, autoimmune disease, mental health disease, infectious disease, or genetic disorder.

[0190] In some embodiments, the model is trained on a dataset collected across multiple disorders, and the model is trained to distinguish between a healthy state and an unhealthy state. For example, as described in Example 5, the random forest classifier was trained on a dataset from 26 different studies that collectively looked at the microbiomes in 15 different disorders. As shown in FIG. 38, the resulting model was exponentiated to predict a healthy or unhealthy disorder state regardless of the disorder. Thus, in some embodiments, the biological characteristic is any one of a plurality of diseases and / or disorders, the first state is the presence of any one of the diseases or disorders, and the second state is the absence of any of the diseases or disorders.

[0191] Referring to block 320, in some embodiments, the disease or disorder is cancer.

[0192] Referring to block 322, in some embodiments, the method includes, for each respective training subject among a plurality of training subjects, inputting information regarding each respective training subject into a model that includes a plurality of parameters. The model applies the plurality of parameters to the information via at least 10,000 calculations to obtain a corresponding output for each respective training subject from the model. The corresponding output includes an indicator of a corresponding state of a biological characteristic of each respective training subject. The information regarding each respective training subject includes corresponding genomic abundance values for each respective gut microbe among a plurality of gut microbes, and the plurality of gut microbes are selected from Table 1, Table 2, or FIGS. 42A - 42XX.

[0193] Referring to block 322, in some embodiments, the indicator of the corresponding state of the biological characteristic is a class output of each state among a plurality of possible states of the biological characteristic. In some embodiments, the possible state is a state from a healthy subject. In some embodiments, the possible state is a state from a patient. In some embodiments, the state from a patient is classified by the type, frequency, or intensity experienced by the patient. In some embodiments, the state from a patient is classified by the progression or prognosis of a disease or disorder, e.g., different stages of cancer. In some embodiments, a threshold for determining the state of a healthy subject or patient is provided, such as a biomarker level, a diagnostic cut - off value, or a threshold nutrient intake level.

[0194] Referring to block 324, in some embodiments, the indicator of the corresponding state of the biometric characteristic is a probability output of the corresponding state of the biometric characteristic. In some embodiments, the corresponding state is the state from a healthy subject. In some embodiments, the corresponding state is the state from a patient. In some embodiments, the state from the patient is classified by the type, frequency, or intensity experienced by the patient. In some embodiments, the state from the patient is classified by the progression or prognosis of a disease or disorder, for example, different stages of cancer. In some embodiments, a threshold is provided for determining the state of a healthy subject or patient, such as a biomarker level, a diagnostic cut-off value, or a threshold nutrient intake level.

[0195] Referring to block 326, in some embodiments, the model is a neural network algorithm, a support vector machine algorithm, a naive Bayes algorithm, a nearest neighbor algorithm, a boosted tree algorithm, a random forest algorithm, a convolutional neural network algorithm, a decision tree algorithm, a regression algorithm, or a clustering algorithm.

[0196] Referring to block 328, in some embodiments, the plurality of parameters is at least 1000, at least 10,000, at least 15,000, at least 50,000, at least 100,000, at least 250,000, at least 500,000, at least 1,000,000, at least 2,500,000, at least 5,000,000, at least 10,000,000, or more.

[0197] Referring to block 330, in some embodiments, the model applies a plurality of parameters to information via at least 1,000 calculations, at least 5,000 calculations, at least 10,000 calculations, at least 25,000 calculations, at least 50,000 calculations, at least 100,000 calculations, at least 250,000 calculations, at least 500,000 calculations, at least 1,000,000 calculations, at least 2,500,000 calculations, at least 5,000,000 calculations, at least 10,000,000 calculations, or more, and obtains corresponding outputs for each training target from the model.

[0198] Referring to block 330, in some embodiments, the method includes adjusting a plurality of parameters based on one or more differences between (i) the corresponding output from the model and (ii) the corresponding state of the biometric characteristics of each respective training target among the first plurality of training targets.

[0199] In some embodiments where deep learning techniques utilize neural networks as described above, training of the neural network to improve its prediction accuracy involves modifying one or more parameters including, but not limited to, the weights within the filters in the convolutional layer and the biases within the network layers. In some embodiments, the weights and biases are further constrained in various forms of regularization such as L1, L2, weight decay, and dropout.

[0200] For example, in some embodiments, if the training data is labeled (e.g., with an indicator of the state of a living body characteristic), either the neural network or the model disclosed herein has its parameters (e.g., weights) adjusted (adjusted to potentially minimize the error between the predicted indicator of the system and the measured indicator of the training data). Various methods are used to minimize an error function such as gradient descent, which includes, but is not limited to, log loss, sum of squared errors, hinge loss methods. In some embodiments, these methods further include second-order methods or approximations such as the momentum method, Hessian-free estimation, Nesterov's accelerated gradient method, AdaGrad. In some embodiments, the method also combines unlabeled generative pre-training and labeled discriminative training.

[0201] Accordingly, in some embodiments, training of the neural network includes adjusting one or more of a plurality of parameters by backpropagation of error through a loss function. In some embodiments, the loss function is a regression task and / or a classification task. Non-limiting examples of loss functions suitable for regression tasks include, but are not limited to, mean squared error loss function, mean absolute error loss function, Huber loss function, Log-Cosh loss function, or quantile loss function. See Wang et al., 2020, “A Comprehensive Survey of Loss Functions in Machine Learning,” Annals of Data Science, doi.org / 10.1007 / s40745-020-00253-5, last accessed September 15, 2021, which is incorporated herein by reference in its entirety. Non-limiting examples of loss functions suitable for classification tasks include, but are not limited to, binary cross-entropy loss function, hinge loss function, or squared hinge loss function. In some embodiments, the loss function is any suitable regression task loss function or classification task loss function.

[0202] Other suitable methods for training a neural network contemplated for use in this disclosure are further described herein (e.g., the untrained model described above).

[0203] In some embodiments, the parameters of the neural network are randomly initialized prior to training.

[0204] In some embodiments, the neural network includes dropout normalization parameters. For example, in some embodiments, normalization is performed by adding a penalty to the loss function, where the penalty is proportional to the value of the parameters in the trained or untrained model. Generally, normalization reduces the complexity of the model by adding a penalty to one or more parameters and reduces the importance of each hidden neuron associated with those parameters. Such practices can result in a more generalized model and can reduce overfitting of the data. In some embodiments, normalization includes an L1 or L2 penalty.

[0205] In some embodiments, training the neural network includes an optimizer. In some embodiments, the optimizer may use the loss function to update the parameters of the neural network or other model via backpropagation of errors. In some embodiments, training the neural network includes a learning rate.

[0206] In some embodiments, the learning rate is at least 0.0001, at least 0.0005, at least 0.001, at least 0.005, at least 0.01, at least 0.05, at least 0.1, at least 0.2, at least 0.3, at least 0.4, at least 0.5, at least 0.6, at least 0.7, at least 0.8, at least 0.9, or at least 1. In some embodiments, the learning rate is 1 or less, 0.9 or less, 0.8 or less, 0.7 or less, 0.6 or less, 0.5 or less, 0.4 or less, 0.3 or less, 0.2 or less, 0.1 or less, 0.05 or less, 0.01 or less, or less than that. In some embodiments, the learning rate is 0.0001 - 0.01, 0.001 - 0.5, 0.001 - 0.01, 0.005 - 0.8, or 0.005 - 1. In some embodiments, the learning rate falls within another range starting from 0.0001 or more and ending at 1 or less.

[0207] In some embodiments, the learning rate further includes learning rate decay (e.g., a decrease in the learning rate over one or more epochs). For example, the learning decay rate can be a 0.5 or 0.1 decrease in the learning rate. In some embodiments, the learning rate is a differential learning rate. In some embodiments, training the neural network further uses a scheduler that conditionally applies learning rate decay based on the evaluation of a performance metric over a threshold number of training epochs (e.g., learning rate decay is applied when the performance metric fails to meet a threshold performance value for at least a threshold number of training epochs).

[0208] In some embodiments, the performance of the neural network is measured at one or more time points using performance metrics including, but not limited to, a training loss metric, a validation loss metric, and / or a mean absolute error. In some embodiments, the performance metric is an area under the receiver operating characteristic curve (AUROC) and / or an area under the precision - recall curve (AUPRC).

[0209] For example, in some embodiments, the performance of a neural network is measured by validating the model using a validation (e.g., development) dataset. In some such embodiments, by training a neural network, a trained neural network is formed when the neural network meets minimum performance requirements based on the validation.

[0210] In some embodiments, any suitable method for validation can be used, including but not limited to K-fold cross-validation, advanced cross-validation, random cross-validation, grouped cross-validation (e.g., K-fold grouped cross-validation), bootstrap bias-corrected cross-validation, random search, and / or Bayesian hyperparameter optimization.

[0211] In some embodiments, a method for training a model including a plurality of parameters is provided by a procedure including: (i) inputting corresponding genomic abundance values for each of a plurality of gut microbiota for each of a plurality of training subjects, thereby obtaining, as an output from the model, corresponding predicted states of biological characteristics for each of the plurality of training subjects; and (ii) improving a plurality of model parameters based on the difference between the corresponding states of biological characteristics for each of the plurality of training subjects and the corresponding predicted states of biological characteristics for each of the plurality of training subjects.

[0212] 3. Method for Evaluating the Health of a Subject FIG. 4 is a schematic diagram of a method for training a model for evaluating human health, as described below. The method can be implemented using a computer system (e.g., the computer system 100 shown and described above with reference to FIG. 1).

[0213] Referring to block 400, in some embodiments, the method comprises obtaining, for each respective gut microbe of a plurality of (e.g., at least 20) gut microbes selected from Table 1, Table 2, or FIGS. 42A - 42XX in electronic form, a plurality of genomic abundance values including corresponding abundance values for the genome of each species of gut bacteria among the plurality of at least 20 gut microbes in a biological sample from a subject. In some embodiments, the plurality of gut microbes includes at least 25 gut microbes selected from Table 1, Table 2, or FIGS. 42A - 42XX. In some embodiments, the plurality of gut microbes includes at least 30 gut microbes selected from Table 1, Table 2, or FIGS. 42A - 42XX. In some embodiments, the plurality of gut microbes includes at least 40 gut microbes selected from Table 1, Table 2, or FIGS. 42A - 42XX. In some embodiments, the plurality of gut microbes includes at least 25 gut microbes selected from Table 1, Table 2, or FIGS. 42A - 42XX. In some embodiments, the plurality of gut microbes includes at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 125, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, or all of the gut microbes selected from Table 1, Table 2, or FIGS. 42A - 42XX.In some embodiments, the plurality of gut microbiota are at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 125, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, or all of the gut microbiota selected from Table 1, Table 2, or FIGS. 42A - 42XX. In some embodiments, the plurality of gut microbiota are all of the gut microbiota listed in Table 1. In some embodiments, the plurality of gut microbiota are all of the gut microbiota listed in Table 2. In some embodiments, the plurality of gut microbiota are all of the gut microbiota listed in FIGS. 42A - 42XX.

[0214] In some embodiments, the corresponding value to the genomic abundance is a value representing the absolute abundance of the microbial genome. In some embodiments, the corresponding value to the genomic abundance is a value representing a normalized abundance value or a relative abundance value (e.g., the abundance of one microorganism normalized to the abundance of the total microbiome of interest). In some embodiments, the corresponding value to the genomic abundance is a value representing an averaged abundance value (e.g., the average of abundances obtained at different time points, or from different biological samples from a patient, or the average of abundances obtained using different probes), or any combination of the above. The corresponding value to the genomic abundance is measured by any technique known in the art. In some embodiments, the value of the genomic abundance of the genome is measured by quantitative PCR (qPCR), such as bacterial 16S rRNA qPCR, RT-PCR, or qRT-PCR, to quantify the abundance of regions of interest in the genome, as described, for example, in U.S. Patent No. 11,427,865, the entire disclosure of which is incorporated herein by reference. In some embodiments, the genomic abundance value is measured by targeted sequencing (e.g., 16S rRNA sequencing, or any other suitable biomarker), partial genome sequencing, or whole genome sequencing, thereby quantifying the number of reads of a targeted region within the microbial genome to determine the abundance of the genome, as disclosed, for example, in U.S. Patent Application Publication No. 2021 / 0403986 or U.S. Patent No. 11,332,783, the entire disclosures of which are incorporated herein by reference. In some embodiments, deep sequencing is used to determine the abundance of a targeted sequence, as disclosed, for example, in U.S. Patent Application Publication No. 2018 / 0237863, the disclosure of which is incorporated herein by reference in its entirety.In some embodiments, the depth of sequencing is at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 70, at least 80, at least 90, at least 100, at least 110, at least 120, at least 130, at least 150, at least 200, at least 300, at least 400, at least 500, at least 750, at least 1000, or more. In some embodiments, for example, shotgun metagenomic sequencing is used to provide sequence reads of the genome in a sample, as described in U.S. Patent No. 11,028,449, the contents of which are hereby incorporated by reference in their entirety.

[0215] In some embodiments of the methods described herein, as shown in FIG. 41, when the identified genomic construct has at least 97% sequence identity compared to the contigs of the microorganisms provided in the sequence listing, the genome identified by metagenomic analysis is classified as corresponding to the microorganisms listed in Table 1, Table 2, and / or FIGS. 42A - 42XX. In some embodiments, as shown in FIG. 41, when the identified genomic construct has at least 98% sequence identity compared to the contigs of the microorganisms provided in the sequence listing, the genome identified by metagenomic analysis is classified as corresponding to the microorganisms listed in Table 1, Table 2, and / or FIGS. 42A - 42XX. In some embodiments, as shown in FIG. 41, when the identified genomic construct has at least 99% sequence identity compared to the contigs of the microorganisms provided in the sequence listing, the genome identified by metagenomic analysis is classified as corresponding to the microorganisms listed in Table 1, Table 2, and / or FIGS. 42A - 42XX. In some embodiments, as shown in FIG. 41, when the identified genomic construct has at least 99.5% sequence identity compared to the contigs of the microorganisms provided in the sequence listing, the genome identified by metagenomic analysis is classified as corresponding to the microorganisms listed in Table 1, Table 2, and / or FIGS. 42A - 42XX. In some embodiments, the genome identified by metagenomic analysis, as shown in FIG. 41, when the identified genomic construct has at least 97%, at least 97.5%, at least 98%, at least 98.5%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9%, or more sequence identity compared to the contigs of the microorganisms provided in the sequence listing, is classified as corresponding to the microorganisms listed in Table 1, Table 2, and / or FIGS. 42A - 42XX.

[0216] Referring to block 402, in some embodiments, the method includes sequencing genomic DNA from a biological sample derived from the intestine of a subject, thereby obtaining a plurality of (e.g., at least 100,000) nucleic acid sequences. In some embodiments, the plurality of nucleic acid sequences includes at least 100,000, at least 250,000, at least 500,000, at least 1,000,000, at least 2,500,000, at least 5,000,000, at least 10,000,000, or at least 50,000,000 nucleic acid sequences. In some embodiments, the plurality of nucleic acid sequences includes 250,000,000 or fewer, 100,000,000 or fewer, 50,000,000 or fewer, 25,000,000 or fewer, 10,000,000 or fewer, 5,000,000 or fewer, 1,000,000 or fewer, 100,000 or fewer nucleic acid sequences. In some embodiments, the plurality of nucleic acid sequences consists of from 100,000 to 1,000,000, from 200,000 to 5,000,000, from 500,000 to 10,000,000, from 1,000,000 to 20,000,000, from 5,000,000 to 50,000,000, from 10,000,000 to 100,000,000, or from 50,000,000 to 250,000,000 nucleic acid sequences. In some embodiments, the plurality of nucleic acid sequences falls within another range that starts with at least 100,000 nucleic acid sequences and ends with 250,000,000 or fewer nucleic acid sequences.

[0217] In some embodiments, a plurality of (e.g., at least 100,000) nucleic acid sequences are obtained by metagenomic sequencing, such as disclosed in U.S. Patent Application Publication No. 2016 / 0239602 or U.S. Patent No. 11,495,326, the contents of which are hereby incorporated by reference in their entirety. In some embodiments, metagenomic sequencing further includes generating a plurality of metagenomic fragment reads. In some embodiments, metagenomic sequencing further includes fragmenting a microbial genome into randomly sized fragments of a target size. The resulting fragments can vary in size. In one embodiment, fragments of approximately 500 nucleotides can be obtained. In some embodiments, fragments of 100 to 2000 nucleotides, such as 200 to 800, 100 to 900, 100 to 1000, 300 to 800, 400 to 900 nucleotides can be obtained. In some embodiments, the method can further include extracting metagenomic fragments from a corresponding biological sample. In some embodiments, metagenomic sequencing further includes sequencing the fragments using a high-throughput sequencing method to generate a plurality of sequencing reads.

[0218] In some embodiments, the first plurality (e.g., at least 100,000) of nucleic acid sequences are obtained by targeted panel sequencing, as described, for example, in U.S. Patent Application Publication No. 2019 / 0316209. In some embodiments, targeted panel sequencing comprises hybridizing genomic DNA isolated from a biological sample from the subject's gut to a panel of probes that hybridize to unique sequences within the genome of each of the quantified microorganisms, e.g., each of the plurality of microorganisms listed in Tables 1, 2, and / or FIGS. 42A - 42XX, prior to sequencing the recovered nucleic acids. In some embodiments, a combination of quasi-unique sequences (e.g., sequences found in a small number of microbial genomes) can be used to back-calculate genomic abundance values using an algorithm, e.g., a system of equations. In some embodiments, the panel of probes comprises at least one probe that hybridizes to a sequence unique to each detected microbial genome. In some embodiments, the panel of probes comprises at least 2, at least 3, at least 4, at least 5, at least 10, at least 25, at least 50, or more probes that hybridize to different sequences unique to each detected microbial genome. In some embodiments, the panel of probes comprises at least 20, at least 30, at least 40, at least 50, at least 75, at least 100, at least 125, at least 150, at least 200, at least 150, at least 300, at least 400, at least 500, at least 750, at least 1000, at least 1250, at least 1500, at least 2000, at least 2500, at least 3000, at least 4000, at least 5000, at least 7500, at least 10,000, or more unique probes.

[0219] In some embodiments, the genomic DNA sequenced from the corresponding biological sample contains at its ends partial or complete sequencing platform adapter sequences useful for sequencing using the targeted sequencing platform. Examples of targeted sequencing platforms include, but are not limited to, the Illumina® HiSeq®, MiSeq®, and Genome Analyzer® sequencing systems, the Ion Torrent® Ion PGM® and Ion Proton® sequencing systems, the Pacific Biosciences PACBIO RS II Sequel system, the Life Technologies® SOLiD sequencing system, the Roche 454GS FLX+ and GS Junior sequencing systems, the Oxford Nanopore MinION® system, or any other targeted sequencing platform.

[0220] Referring to block 404, in some embodiments, the biological sample derived from the intestine of each subject is a fecal sample. In some embodiments, the sample is a tissue biopsy, intestine, or mucosal sample. See, e.g., Tang Q, Jet al., Current Sampling Methods for Gut Microbiota: A Call for More Precise Devices, Front Cell Infect Microbiol., 10:151 (2020), the content of which is hereby incorporated by reference in its entirety.

[0221] Referring to block 406, in some embodiments, the plurality of gut microbes are selected from the microbes in Table 5 having at least 2, at least 3, at least 4, at least 5, at least 10, at least 15, at least 20, at least 25, or more binding properties. Referring to block 308, in some embodiments, the plurality of gut microbes includes at least 20 microbes selected from the microbes listed in Table 5 as having at least 2 binding properties. In some embodiments, the plurality of gut microbes includes at least 20 microbes selected from those microbes listed in Table 5 as having at least 3, at least 4, at least 5, at least 10, at least 15, at least 20, at least 25, or more binding properties.

[0222] Referring to block 408, in some embodiments, the method includes obtaining, in electronic form, a plurality of (e.g., at least 100,000) nucleic acid sequences for genomic DNA from a biological sample derived from the gut of a subject, and for each respective gut microbe among the plurality of gut microbes, determining a corresponding value for the abundance of the genome of the respective gut microbe from the plurality of at least 100,000 nucleic acid sequences. In some embodiments, the genomic abundance value determined for a subject includes a genomic abundance value of 250,000 or less, 100,000 or less, 50,000 or less, 25,000 or less, 10,000 or less, 5,000 or less, 1,000 or less, 100 or less, 50 or less, 30 or less, or 20 or less. In some embodiments, the genomic abundance value determined for a subject consists of a genomic abundance value of 10 - 40, 20 - 50, 30 - 80, 40 - 100, 50 - 150, 60 - 200, 80 - 300, 90 - 500, 100 - 1000, 500 - 2,000, or 1,000 - 5,000. In some embodiments, the genomic abundance value determined for a subject falls within another range starting with a genomic abundance value of 20 or less and ending with a genomic abundance value of 250,000 or less.

[0223] Referring to block 410, in some embodiments, assembling, in electronic form, a plurality of (e.g., at least 100,000) gut microbial genomes by metagenomic de novo sequence assembly from a plurality of nucleic acid sequences; and for each respective gut microbe among the plurality of gut microbes, calculating a corresponding value for the abundance of the genome of each respective gut microbe based on the prevalence of each respective nucleic acid sequence among the plurality of at least 100,000 nucleic acid sequences used to assemble the respective gut microbial genome corresponding to each respective gut microbe. In some embodiments, the metagenomic de novo sequence assembly further includes generating contigs based on sequencing reads generated by shotgun sequencing techniques as described in U.S. Patent No. 10,529,443, the contents of which are incorporated herein by reference in their entirety. In some embodiments, the plurality of (e.g., at least 100,000) nucleic acid sequences can be assembled into partial genomes of the plurality of gut microbes. In some embodiments, the plurality of (e.g., at least 100,000) nucleic acid sequences can be assembled into whole genomes of the plurality of gut microbes.

[0224] Referring to block 412, in some embodiments, the method assigns each respective nucleic acid sequence in a plurality of (e.g., at least 100,000) arrays to a respective gut microorganism among a plurality of gut microorganisms, thereby generating, for each respective gut microorganism among the plurality of gut microorganisms, a corresponding count of each respective nucleic acid sequence among the plurality of nucleic acid sequences assigned to the respective gut microorganism, and determining, for each respective gut microorganism among the plurality of gut microorganisms, a corresponding genomic abundance value for the respective gut microorganism based on the corresponding count of each respective nucleic acid sequence assigned to the respective gut microorganism. In some embodiments, assigning each nucleic acid to a respective gut microorganism includes mapping the nucleic acid to a reference nucleic acid. In some embodiments, assigning a respective gut microorganism to each respective nucleic acid includes annotating genomic information based on an existing database. In some embodiments, nucleic acid sequences are analyzed and the annotation defines a taxonomic assignment using sequence similarity methods and phylogenetic placement methods, or a combination of the two strategies.

[0225] Array similarity-based methods include, but are not limited to, those well known to those skilled in the art, including various implementations of these algorithms such as BLAST, BLASTx, tBLASTn, tBLASTx, RDP classifier, DNAclust, and Qiime or Mothur. These methods rely on mapping sequence reads to a reference database and selecting the matches with the best scores and e-values. In some embodiments, phylogenetic methods are used in combination with sequence similarity methods to improve the call accuracy of annotations or taxonomic assignments. Common databases include, but are not limited to, GT-DBTK, National Center for Biotechnology Information (NCBI) Genbank, European Bioinformatics Institute - European Nucleotide Archive (European Bioinformatics Institute - European Nucleotide Archive; EBI-ENA), National Institute of Genetics, U.S. Department of ENERGY (USDOE) Integrated Microbial Genomes (Integrated Microbial Genomes) & Microbiomes; IMG / M), and other databases available in the art.

[0226] Referring to block 414, in some embodiments, the method is to input a plurality of genomic abundance values into a model that includes a plurality of parameters, and the model applies the plurality of parameters to the plurality of genomic abundance values through a plurality of (e.g., at least 10,000) calculations to generate an indicator of the health of the subject as an output from the model, including inputting.

[0227] Referring to block 416, in some embodiments, the health indicator of interest is an indicator of a biological characteristic, and the biological characteristic is a disease or disorder, a therapy administered to the subject, such as surgery, radiation therapy, chemotherapy, targeted therapy, gene therapy, immunotherapy, drug therapy, dietary changes, lifestyle changes, or the subject's diet such as a diet rich or poor in carbohydrates, proteins, fats, vitamins, or fiber.

[0228] Referring to block 418, in some embodiments, the disease or disorder is selected from the group consisting of type 2 diabetes (T2D), hypertension (HT), schizophrenia (SCZ), atherosclerotic cardiovascular disease (ACVD), cirrhosis (LC), inflammatory bowel disease (IBD), colorectal cancer (CRC), ankylosing spondylitis (AS), and Parkinson's disease (PD), multiple sclerosis (MS), type 2 Gaucher disease (GDII), COVID-19 (COV), Behçet's disease (BD), autism spectrum disorder (ASD), or pancreatic cancer (PC). In some embodiments, the disease or disorder is cancer, Alzheimer's disease, cardiovascular disease, autoimmune disease, mental health disease, infectious disease, or genetic disorder.

[0229] Referring to block 420, in some embodiments, the disease or disorder is cancer.

[0230] In some embodiments, the model is trained on a dataset collected across multiple disorders, and the model is trained to distinguish between a healthy state and an unhealthy state. For example, as described in Example 5, the random forest classifier was trained on a dataset from 26 different studies that collectively looked at the microbiomes in 15 different disorders. As shown in FIG. 38, the resulting model was exponentiated to predict a healthy or unhealthy disorder state regardless of the disorder. Thus, in some embodiments, the biological characteristic is any one of a plurality of diseases and / or disorders, the first state is the presence of any one of the diseases or disorders, and the second state is the absence of any of the diseases or disorders.

[0231] Referring to block 422, in some embodiments, an indicator of a subject's health is a class output for each state in a plurality of possible states of the subject's health. In some embodiments, each state of the subject's health is referred to by the severity of a disease or disorder. In some embodiments, the severity of a disease is classified by the progression or prognosis of the disease or disorder, e.g., different stages of cancer. In some embodiments, a threshold for determining the subject's health state, such as a biomarker level, a diagnostic cut-off value, or a threshold nutrient intake level, is provided. In some embodiments, each state of the subject's health is the absence or presence of a disease or disorder.

[0232] Referring to block 424, in some embodiments, an indicator of a subject's health is a probability output for a corresponding state of the subject's health. In some embodiments, the corresponding state of the subject's health is referred to by the severity of a disease or disorder. In some embodiments, the severity of a disease is classified by the progression or prognosis of the disease or disorder, e.g., different stages of cancer. In some embodiments, a threshold for determining the subject's state, such as a biomarker level, a diagnostic cut-off value, or a threshold nutrient intake level, is provided. In some embodiments, the corresponding state of the subject's health is the absence or presence of a disease or disorder.

[0233] Referring to block 426, in some embodiments, the model is a neural network algorithm, a support vector machine algorithm, a naive Bayes algorithm, a nearest neighbor algorithm, a boosted tree algorithm, a random forest algorithm, a convolutional neural network algorithm, a decision tree algorithm, a regression algorithm, or a clustering algorithm.

[0234] Referring to block 428, in some embodiments, the plurality of parameters are at least 1000, at least 10,000, at least 15,000, at least 50,000, at least 100,000, at least 250,000, at least 500,000, at least 1,000,000, at least 2,500,000, at least 5,000,000, at least 10,000,000, or more.

[0235] Referring to block 430, in some embodiments, the model applies the plurality of parameters to the information via at least 1000 calculations, at least 5000 calculations, at least 10,000 calculations, at least 25,000 calculations, at least 50,000 calculations, at least 100,000 calculations, at least 250,000 calculations, at least 500,000 calculations, at least 1,000,000 calculations, at least 2,500,000 calculations, at least 5,000,000 calculations, at least 10,000,000 calculations, or more, and obtains the corresponding output for each training target from the model.

Example

[0236] Example 1 - Seesaw network guild as a common microbiome signature for human diseases The microorganisms necessary to provide health-related functions essential to the host [7] were hypothesized to maintain stable ecological interactions with each other for structural and functional stability [18, 19]. To identify the microbiome signature based on stable interactions between MAGs, patients with T2DM were randomized at baseline (M0) and received either a 3-month (M3) high-fiber intervention (W group; n = 74) or standard treatment (U group; n = 36), followed by a 1-year follow-up (M15) in an open-label control trial (Figure 3A and Figure 7). Positive environmental perturbations were exerted using the high-fiber intervention to dramatically and reversibly change the abundance of members of the gut microbiome [16, 17]. Co-occurrence network analysis at each of the three time points made it possible to identify MAG pairs that could maintain their correlation despite significant changes in the abundance of the entire community due to perturbation. These genomic pairs were from 141 MAGs, and they were found to form two guilds organized as the two ends of a robustly stable seesaw-like network. Also, these seesaw-networked genomes supported a machine learning model for the predictive classification of cases and controls in 12 independent metagenomic datasets from 1,874 subjects across different cohorts and various chronic diseases, including T2DM, atherosclerotic cardiovascular disease (ACVD), hypertension, liver cirrhosis (LC), inflammatory bowel disease (IBD), colorectal cancer (CRC), ankylosing spondylitis (AS), schizophrenia, and Parkinson's disease (PD), suggesting the identification of a common microbiome signature across different human diseases.

[0237] Reversible changes in the gut microbiome associated with reversible changes in host metabolic phenotypes The dietary fiber intake of the U group did not change throughout the study, while the W group had a significantly increased dietary fiber intake from M0 to M3 and a decreased dietary fiber intake from M3 to M15, but remained higher than M0 (Figure 5B). Compared with the U group, the W group had significantly higher fiber intakes at M3 and M15, but similar energy and macronutrient consumption (Figure 10).

[0238] To investigate the structural changes of the gut microbiota in response to the introduction and discontinuation of high-fiber intervention, shotgun metagenomic sequencing was performed on 315 fecal samples collected from 110 patients in the W group and the U group. Among them, 95 patients provided samples at all three time points, and 15 patients provided samples only at M0 and M3 (Figure 9). To achieve resolution at the strain and subspecies levels, 1,845 non-redundant high-quality draft genomes were reconstructed from the metagenomic dataset (when the average nucleotide identity (ANI) between them was >99%, two genomes collapsed into one MAG), and these MAGs accounted for more than 70% of all reads. In the context of beta diversity based on the Bray-Curtis distance, the overall structure of the gut microbiota in the W group changed significantly from M0 to M3 (PERMANOVA test, P<0.001), returned to the M0 structure at M15, but there was no difference in the U group over the three time points (Figure 5C, D). Similar changes in alpha diversity based on the Shannon index and Simpson index were also observed (Figure 11). These results showed that, as previously reported, high-fiber intervention induced significant structural changes in the gut microbiota

[16] , but after the intervention was discontinued, the gut microbiota returned to the baseline, indicating high elasticity of the community structure.

[0239] To determine whether the host metabolic phenotype also shows reversible changes similar to the gut microbiota, 43 clinical parameters were examined over three time points. Hemoglobin A1c (HbA1c) in the U group did not change throughout the study. With the high-fiber intervention, the HbA1c level in the W group decreased by an average of 15.22% ± 9.82% (mean ± s.d) from M0 to M3, and such a decrease was significantly greater than the decrease in the U group. During the one-year follow-up, HbA1c increased significantly from M3, but remained lower than M0 in the W group (Figure 5E). The proportion of patients who achieved appropriate blood glucose control (HbA1c ≤ 7%) at M3 was also significantly higher in the W group (61.6% vs 33.3% in the U group), but there was no difference between the two groups at M15 (Figure 5F). The levels of fasting and postprandial blood glucose in the oral glucose tolerance test followed a similar trend to HbA1c (Figure 5G, H). The W group also showed remission of inflammation, hyperlipidemia, obesity, and T2DM complications from M0 to M3, but recovered during the one-year follow-up (Figure 22). The Mantel test using the Manhattan distance based on all 43 clinical parameters and the Bray-Curtis distance of the gut microbiota showed that the clinical outcome was significantly correlated with the gut bacterial structure (R2 = 0.09, P = 2×10-4). These results indicate that changes in the host metabolic phenotype were associated with reversible changes in the gut microbiota in response to the presence / absence of high-fiber intervention.

[0240] Genome pairs with stable interactions form a seesaw network with two competing guilds To facilitate the identification of genomic pairs that can maintain their ecological interactions stable during the experiment, co-occurrence networks were constructed for each time point based on the abundance matrix of MAGs representing epidemic microorganisms. Since more than 75% of the samples were detectable at each time point in the W group, a total of 477 MAGs were selected for network construction. They were also dominant as they accounted for approximately 60% of the total abundance of 1,845 MAGs. Pairwise correlations were calculated for all 113,526 possible genomic pairs among these 477 epidemic MAGs, and one of three co-occurrence networks was constructed for each time point (GM0, GM3, and GM15) (, Figure 23). During the experiment, the co-occurrence networks of epidemic genomes in M0, M3, and M15 of the W group are shown as GM0 (442; 4231), GM3 (421; 2587), and GM15 (429; 4592). The numbers in parentheses are the order and size of the network. The correlations between genomes were calculated using FastSpar (n = 67 patients). All significant correlations with P ≤ 0.001 were included. The co-occurrence networks were visualized. The edges between nodes represent correlations. Red and blue indicate positive and negative correlations, respectively. The node size indicates the average abundance of the genome. The layout of nodes and edges was determined by an edge-weighted spring embedding layout that weights the correlation efficiency. The three networks had a similar order S, i.e., the total number of nodes (MAG), SM0 (442), SM3 (421), and SM15 (429), but they varied considerably in their size L, i.e., the total number of edges (correlations), LM0 (4231), LM3 (2587), and LM15 (4592). At GM3, L decreased to 61.14% of L at GM0, and at GM15, it recovered to 108.53% of L at GM0. This was confirmed by the change in the degree of connectivity, defined as the proportion of realized ecological interactions among potential ones (in an undirected network, degree of connectivity = L / (S(S - 1) / 2), and the value is in the [0,1] interval)20. The degree of connectivity decreased from 0.043 at GM0 to 0.029 at GM3 and recovered to 0.050 at GM15. High fiber intervention dramatically reduced the interactions between epidemic genomes within the network.In addition, the degree distribution, i.e., the number of edges a node has, was found to fit well the power-law model (Figure 12, R2 value of 0.79 - 0.82) that should show the characteristics of a scale-free network resistant to the presence of nodes with high degrees, random errors or decay

[21] . In this specification, a hub as a node was defined as a node connecting to more than 1 / 5 of all nodes in the network (Figure 13). Among 24 hubs, 10 were in GM0 and 20 were in GM15, but none were in GM3. This indicates that the overall structure of the gut microbiota may have undergone severe changes during the test, and in particular, the high-fiber intervention resulted in a loss of interaction between genomic pairs.

[0241] When genomic pairs maintained the same ecological interactions across all three time points, their ecological relationships were considered robust a...

Claims

1. A method for identifying a set of intestinal microorganisms, In a computer system having one or more processors and memory for storing one or more programs to be executed by the one or more processors, A) To obtain, in electronic format, a plurality of corresponding genome abundance values ​​for each of the first plurality of objects having a first state of biological characteristics, including a corresponding value for the abundance of each intestinal microorganism among the plurality of intestinal microorganisms, with respect to the abundance of the genome of each intestinal microorganism in a biological sample derived from the intestines of each of the respective objects, B) To obtain, in electronic format, for each of the second of multiple subjects having a second state of biological properties, a plurality of corresponding genome abundance values, including the corresponding values ​​for each of the intestinal microorganisms in the plurality of intestinal microorganisms, with respect to the abundance of the genome of each intestinal microorganism in a biological sample derived from the intestine of each of the subjects, C) Calculating a first set of similarity metrics from the corresponding set of genome abundance values ​​across the first set of subjects, The first plurality of similarity metrics include a first corresponding similarity metric for each unique pair of intestinal microorganisms in the plurality of intestinal microorganisms, The first corresponding similarity metric involves calculating a first set of similarity metrics that quantify the similarity between (i) a corresponding first vector formed by the corresponding genome abundance values ​​of a first microorganism in a specific pair of the intestinal microorganisms across the first set of subjects, and (ii) a corresponding second vector formed by the corresponding genome abundance values ​​of a second microorganism in a specific pair of the intestinal microorganisms across the first set of subjects, D) Calculating a second set of similarity metrics using the corresponding genome abundance values ​​for the second set of subjects, The second plurality of similarity metrics includes a second corresponding similarity metric for each unique pair of intestinal microorganisms among the plurality of intestinal microorganisms, The second corresponding similarity metric involves calculating a second set of similarity metrics that quantify the similarity between (i) a corresponding second vector formed by the corresponding genomic abundance values ​​of the first microorganism in a specific pair of the intestinal microorganisms across the second set of subjects, and (ii) a corresponding second vector formed by the corresponding genomic abundance values ​​of the second microorganism in a specific pair of the intestinal microorganisms across the second set of subjects, E) Determining a set of unique pairs of intestinal microorganisms in the plurality of intestinal microorganisms based on the first set of similarity metrics and the second set of similarity metrics, for each unique pair of intestinal microorganisms in the set of unique pairs of intestinal microorganisms, Both the first corresponding similarity metric and the second corresponding similarity metric show a statistically significant positive correlation between the abundance of the first intestinal microorganism and the abundance of the second intestinal microorganism in each intrinsic pair of the intestinal microorganisms, or The determination that both the first corresponding similarity metric and the second corresponding similarity metric show a statistically significant negative correlation between the abundance of the first intestinal microorganism and the abundance of the second intestinal microorganism in each intrinsic pair of the intestinal microorganisms, F) A method comprising identifying a set of intestinal microorganisms, each intestinal microorganism represented by a unique pair set of intestinal microorganisms.

2. The acquisition A) is, (i) In electronic format, for each of the first plurality of targets, to obtain a first plurality of at least 100,000 nucleic acid sequences corresponding to the genomic DNA from the corresponding biological sample derived from the intestine of each of the first plurality of targets, (ii) For each of the first multiple objects, the corresponding first multiple This includes determining the corresponding genome abundance value for each of the multiple intestinal microorganisms from at least 100,000 nucleic acid sequences, The acquisition B) mentioned above is (i) In electronic format, for each of the second plurality of subjects, to obtain a second plurality of at least 100,000 nucleic acid sequences corresponding to the genomic DNA from the corresponding biological sample derived from the intestine of each of the second plurality of subjects, (ii) The method according to claim 1, comprising determining the corresponding genome abundance value for each of the intestinal microorganisms in the plurality of intestinal microorganisms from the corresponding second plurality of at least 100,000 nucleic acid sequences for each of the second plurality of targets.

3. The above determination A)(ii) applies to each of the first multiple objects, The process involves assembling a plurality of corresponding first intestinal microbial genomes by metagenomic de novo sequence assembly from the aforementioned plurality of corresponding first nucleic acid sequences, This includes calculating the corresponding genome abundance of each of the intestinal microbial genomes in the aforementioned first plurality of corresponding intestinal microbial genomes, The above determination B) (ii) applies to each of the second set of objects, The process involves assembling a second group of corresponding gut microbiota genomes by a metagenomic de novo sequence assembly from the aforementioned second group of at least 100,000 nucleic acid sequences, The method according to claim 2, comprising calculating the corresponding genome abundance of each of the intestinal microbial genomes in the corresponding second plurality of intestinal microbial genomes.

4. The above determination A)(ii) applies to each of the first multiple objects, Assigning each nucleic acid sequence in the corresponding first plurality of at least 100,000 sequences to each of the plurality of intestinal microorganisms, thereby generating a corresponding count for each of the corresponding nucleic acid sequences in the corresponding first plurality of nucleic acid sequences assigned to each of the plurality of intestinal microorganisms, This includes determining the corresponding genome abundance value for each of the multiple intestinal microorganisms based on the corresponding count of the nucleic acid sequence assigned to each of the intestinal microorganisms, The above determination B) (ii) is made with respect to each of the first multiple objects, Assigning each nucleic acid sequence in the corresponding second plurality of at least 100,000 sequences to each of the plurality of intestinal microorganisms, thereby generating a corresponding count for each of the corresponding nucleic acid sequences in the corresponding second plurality of nucleic acid sequences assigned to each of the plurality of intestinal microorganisms, The method according to claim 2, comprising determining the corresponding genome abundance value for each of the intestinal microorganisms in the plurality of intestinal microorganisms based on the corresponding count of the nucleic acid sequence assigned to each of the intestinal microorganisms.

5. For each of the first plurality of targets, the genomic DNA from the corresponding biological sample derived from the intestine of each target is sequenced, thereby obtaining the corresponding first plurality of at least 100,000 nucleic acid sequences. The method according to any one of claims 2 to 4, further comprising sequencing the genomic DNA from the corresponding biological sample derived from the intestine of each of the second plurality of targets, thereby obtaining the corresponding second plurality of at least 100,000 nucleic acid sequences.

6. The first state of the biological characteristics is the absence of disease or disability, and the second state of the biological characteristics is the presence of disease or disability. The first state of the biological characteristics is the first severity of the disease or disorder, and the second state of the biological characteristics is the second severity of the disease or disorder. The first state of the biological characteristics is an untreated disease or disorder, and the second state of the biological characteristics is a treated disease or disorder. Whether the first state of the biological characteristics is a disease or disorder treated by the first therapy, and whether the second state of the biological characteristics is a disease or disorder treated by the second therapy, The first state of the biological properties is a first level of nutrient in a diet, and the second state of the biological properties is a second level of nutrient in a diet, or The method according to any one of claims 1 to 4, wherein the first state of the biological characteristics is a first age, and the second state of the biological characteristics is a second age.

7. The method according to any one of claims 1 to 4, wherein the plurality of intestinal microorganisms comprises at least 20 intestinal microorganisms selected from Table 1, Table 2, or Figures 42A to 42XX.

8. For each of the first plurality of subjects, the biological sample derived from the intestine of each subject is a fecal sample. The method according to any one of claims 1 to 4, wherein, for each of the second plurality of targets, the biological sample derived from the intestine of each of the targets is a fecal sample.

9. The method according to any one of claims 1 to 4, wherein both the first corresponding similarity metric and the second similarity metric are Pearson correlation coefficients, intraclass correlation coefficients, or rank correlation coefficients.

10. The method according to any one of claims 1 to 4, wherein a statistically significant positive correlation has a p-value of less than 0.

001.

11. The method according to any one of claims 1 to 4, wherein the set of intestinal microorganisms includes all of the intestinal microorganisms represented by a unique set of pairs of intestinal microorganisms.

12. The above identification F) Clustering each of the aforementioned intestinal microorganisms, represented by a unique set of pairs of intestinal microorganisms, into one of more networks, wherein each connected network includes a corresponding set of multiple nodes and one or more corresponding edges. Each of the corresponding nodes represents a unique set of intestinal microorganisms, Each of the corresponding edges in the set of one or more edges connects two nodes representing each unique pair of intestinal microorganisms in the unique set of intestinal microorganisms, Each node in the corresponding plurality of nodes connects to at least one other node in the plurality of nodes via each edge in the corresponding set of one or more edges. Each node is connected to a cluster, The method according to any one of claims 1 to 4, comprising: identifying each network in one or more networks containing the most nodes; thereby identifying the set of intestinal microorganisms represented by the corresponding number of nodes in each of the networks.

13. The method according to any one of claims 1 to 4, wherein the set of intestinal microorganisms comprises at least 20 intestinal microorganisms selected from Table 1, Table 2, or Figures 42A to 42XX.

14. A method for training a model to assess human health, In a computer system having one or more processors and memory for storing one or more programs to be executed by the one or more processors, A) In electronic format, for each of the multiple training subjects, (i) A plurality of corresponding genome abundance values, including a corresponding value for each of the intestinal microorganisms in the plurality of intestinal microorganisms, relative to the genome abundance of each intestinal microorganism in the corresponding biological sample derived from the intestine of the respective training subject, and (ii) To acquire the corresponding state of the biological characteristics of each of the training subjects, B) For each of the multiple training targets, inputting information about each training target into a model that includes multiple parameters, wherein the model applies the multiple parameters to the information through at least 10,000 calculations in order to obtain a corresponding output from the model for each training target. The corresponding output includes an index of the corresponding state of the biological characteristics of each of the training subjects, The information relating to each of the training targets includes the corresponding genome abundance values ​​for each of the multiple intestinal microorganisms. The aforementioned multiple intestinal microorganisms are selected from Table 1, Table 2, or Figures 42A to 42XX, and input accordingly. C) A method comprising, for each of the first plurality of training subjects, adjusting the plurality of parameters based on one or more differences between (i) the corresponding output from the model and (ii) the corresponding state of the biological characteristics of each of the training subjects.

15. The acquisition A) described above applies to each of the training targets among the multiple training targets, (i) Obtaining, in electronic format, a plurality of at least 100,000 corresponding nucleic acid sequences for the genomic DNA from the corresponding biological sample derived from the intestine of each of the training subjects, (ii) The method according to claim 14, comprising determining the corresponding value for the abundance of the genome of each of the plurality of intestinal microorganisms from a first plurality of corresponding nucleic acid sequences of at least 100,000 nucleic acids.

16. The above decision A)(ii) is made with respect to each of the training subjects among the plurality of training subjects, The assembly of a plurality of corresponding intestinal microbial genomes by a metagenomic de novo sequence assembly from the plurality of corresponding nucleic acid sequences, For each of the plurality of intestinal microorganisms, the abundance of the genome of each of the intestinal microorganisms is determined based on the prevalence of each nucleic acid sequence in the plurality of at least 100,000 nucleic acid sequences used to assemble each intestinal microorganism genome in the plurality of intestinal microorganism genomes corresponding to each of the plurality of intestinal microorganisms. The method according to claim 15, comprising calculating the corresponding value.

17. The above decision A)(ii) applies to each of the multiple training targets, Assigning each nucleic acid sequence in the aforementioned plurality of at least 100,000 corresponding sequences to each of the plurality of intestinal microorganisms, thereby generating a corresponding count for each of the corresponding nucleic acid sequences in the plurality of corresponding nucleic acid sequences assigned to each of the plurality of intestinal microorganisms, The method according to claim 15, comprising determining the corresponding genome abundance value for each of the intestinal microorganisms in the plurality of intestinal microorganisms based on the corresponding count of the nucleic acid sequence assigned to each of the intestinal microorganisms.

18. The method according to any one of claims 15 to 17, further comprising sequencing the genomic DNA from the corresponding biological sample derived from the intestine of each of the plurality of training subjects, thereby obtaining the corresponding plurality of at least 100,000 nucleic acid sequences.

19. The method according to any one of claims 14 to 17, wherein the plurality of intestinal microorganisms comprises at least 20 intestinal microorganisms selected from Table 1, Table 2, or Figures 42A to 42XX.

20. The method according to any one of claims 14 to 17, wherein the plurality of intestinal microorganisms comprises at least 20 microorganisms selected from those microorganisms in Table 1, Table 2, or Figures 42A to 42XX, each having at least two binding capabilities.

21. The method according to any one of claims 14 to 17, wherein for each of the subjects among the plurality of training subjects, the biological sample derived from the intestines of each subject is a fecal sample derived from each of the training subjects.

22. The method according to any one of claims 14 to 17, wherein the biological characteristic is a disease or disorder, a therapy administered to the subject, or a diet for the subject.

23. The method according to claim 22, wherein the disease or disorder is selected from the group consisting of type 2 diabetes mellitus, hypertension, schizophrenia, atherosclerotic cardiovascular disease (ACCVD), cirrhosis of the liver (LC), inflammatory bowel disease (IBD), colorectal cancer (CRC), ankylosing spondylitis (AS), and Parkinson's disease (PD).

24. The method according to claim 22, wherein the disease or disorder is cancer.

25. The method according to any one of claims 14 to 17, wherein the index of the corresponding state of the biological characteristic is a class output of each state in a plurality of possible states of the biological characteristic.

26. The method according to any one of claims 14 to 17, wherein the index of the corresponding state of the biological characteristics is a probability output of the corresponding state of the biological characteristics.

27. The aforementioned model includes neural network algorithms, support vector machine algorithms, naive Bayes algorithms, nearest neighbor algorithms, boost tree algorithms, random forest algorithms, and convolutional neural network algorithms. The method according to any one of claims 14 to 17, wherein the method is a decision tree algorithm, a regression algorithm, or a clustering algorithm.

28. The method according to any one of claims 14 to 17, wherein the plurality of parameters are at least 1,000, at least 10,000, at least 15,000, at least 50,000, at least 100,000, at least 250,000, at least 500,000, or at least 1,000,000 parameters.

29. The method according to any one of claims 14 to 17, wherein the model applies the plurality of parameters to the information through at least 25,000, at least 50,000, at least 100,000, at least 250,000, at least 500,000, or at least 1,000,000 calculations in order to obtain corresponding outputs from the model for each of the training targets.

30. A method for evaluating the health of a subject, In a computer system having one or more processors and memory for storing one or more programs to be executed by the one or more processors, A) To obtain, in electronic format, multiple genome abundance values ​​for each of the at least 20 intestinal microorganisms selected from Table 1, Table 2, or Figures 42A to 42XX, including the corresponding abundance values ​​for each species of intestinal bacteria in the biological sample derived from the subject, B) A method comprising inputting the plurality of genomic abundance values ​​into a model that includes a plurality of parameters, wherein the model applies the plurality of parameters to the plurality of genomic abundance values ​​through at least 10,000 calculations in order to generate an indicator of the health of the subject {e.g., a class output or a probabilistic output} as an output from the model.

31. The acquisition A) is, (i) Obtaining, in electronic format, at least 100,000 nucleic acid sequences of the genomic DNA from the biological sample derived from the intestine of the subject, (ii) The method according to claim 30, comprising determining, for each of the plurality of intestinal microorganisms, a corresponding value for the abundance of the genome of each intestinal microorganism from the plurality of at least 100,000 nucleic acid sequences.

32. The above decision A) (ii) is, The assembly of a plurality of corresponding intestinal microbial genomes by a metagenomic de novo sequence assembly from the plurality of at least 100,000 nucleic acid sequences in electronic form, The method according to claim 31, comprising: for each of the plurality of intestinal microorganisms, calculating the corresponding value of the genome of each intestinal microorganism to the abundance of each intestinal microorganism based on the prevalence of each nucleic acid sequence in the plurality of at least 100,000 nucleic acid sequences used to assemble each intestinal microorganism genome in the plurality of intestinal microorganism genomes corresponding to each intestinal microorganism.

33. The above decision A) (ii) is, Assigning each nucleic acid sequence in the plurality of at least 100,000 sequences to each of the plurality of intestinal microorganisms, thereby generating a corresponding count for each of the plurality of nucleic acid sequences assigned to each of the plurality of intestinal microorganisms. For each of the intestinal microorganisms among the plurality of intestinal microorganisms, The method according to claim 31, comprising determining the corresponding genome abundance value for each of the intestinal microorganisms based on the corresponding count of each nucleic acid sequence assigned to the substance.

34. The method according to any one of claims 31 to 33, further comprising sequencing the genomic DNA from the biological sample derived from the intestine of the subject, thereby obtaining the plurality of at least 100,000 nucleic acid sequences.

35. The method according to any one of claims 30 to 33, wherein the plurality of intestinal microorganisms comprises at least 20 microorganisms selected from those microorganisms in Table 1, Table 2, or Figures 42A to 42XX, each having at least two binding affinity.

36. The method according to any one of claims 30 to 33, wherein the biological sample derived from the intestine of the subject is a fecal sample.

37. The method according to any one of claims 30 to 33, wherein the indicator of health of the subject is an indicator of biological characteristics, and the biological characteristics are a disease or disorder, a therapy administered to the subject, or a diet for the subject.

38. The method according to claim 37, wherein the disease or disorder is selected from the group consisting of type 2 diabetes mellitus, hypertension, schizophrenia, atherosclerotic cardiovascular disease (ACCVD), cirrhosis of the liver (LC), inflammatory bowel disease (IBD), colorectal cancer (CRC), ankylosing spondylitis (AS), and Parkinson's disease (PD).

39. The method according to claim 37, wherein the disease or disorder is cancer.

40. The method according to any one of claims 30 to 33, wherein the indicator of the health of the subject is a class output of each of the multiple possible states of the health of the subject.

41. The method according to any one of claims 30 to 33, wherein the indicator of the health of the subject is a probability output of the corresponding state of the health of the subject.

42. The method according to any one of claims 30 to 33, wherein the model is a neural network algorithm, a support vector machine algorithm, a naive Bayes algorithm, a nearest neighbor algorithm, a boosted tree algorithm, a random forest algorithm, a convolutional neural network algorithm, a decision tree algorithm, a regression algorithm, or a clustering algorithm.

43. The method according to any one of claims 30 to 33, wherein the plurality of parameters are at least 1,000, at least 10,000, at least 15,000, at least 50,000, at least 100,000, at least 250,000, at least 500,000, or at least 1,000,000 parameters.

44. The method according to any one of claims 30 to 33, wherein the model applies the plurality of parameters to the information through at least 25,000, at least 50,000, at least 100,000, at least 250,000, at least 500,000, or at least 1,000,000 calculations in order to obtain corresponding outputs from the model for each training target.

45. A computer system, One or more processors, A computer system comprising: a non-temporary computer-readable medium containing computer-executable instructions that, when executed by one or more of the aforementioned processors, cause the processors to carry out the method according to any one of claims 1 to 4, 14 to 17, or 30 to 33.

46. A non-temporary computer-readable storage medium having a program code instruction stored in it that, when executed by the processor, causes the processor to perform the method according to any one of claims 1 to 4, 14 to 17, or 30 to 33.