Methods for predicting response to therapy for a condition by core microbiome bacterial populations

CN122804060APending Publication Date: 2026-09-22RUTGERS THE STATE UNIV +1
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
CN202480083303.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-01
Filing Date
2024-10-31
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

虽然这些方法无疑为微生物组的结构配置和潜在功能特征提供了重要的见解,但它们可能不足以代表强调这一复杂系统的稳定性和复原力的重要生态相互作用

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

Methods and systems for predicting a subject's response to a therapy by obtaining a first plurality of nucleic acid sequences for genomic DNA from a sample of a subject's gut. A plurality of genomic abundance values for a plurality of gut bacteria is determined from the nucleic acid sequences. A model is applied to the plurality of genomic abundance values, obtaining as an output of the model a prediction of the subject's response to a therapy.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-citation of related applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 595,189, filed November 1, 2023, the contents of which are hereby incorporated herein by reference in their entirety for all purposes.

[0003] Brief description of sequence lists

[0004] This submission incorporates by reference the “Sequence List XML” file named ST26_126146_5002_PR.XML, containing SEQ ID NO: 1 to 99534, created on April 22, 2023, and measuring 2,491,699 kilobytes, conforming to 37 CFR § 1.831 to 1.835, which was submitted by mail as an XML file on a read-only optical disc (DVD) for use in U.S. Provisional Application Serial No. 63 / 498,177. 37 CFR § 1.835(a)(1). The entire Sequence List XML is incorporated herein by reference.

[0005] This submission incorporates by reference the “Sequence List XML” file named ST26_126146_5001_WO.XML containing SEQ ID NO: 1 to 99534, created on April 24, 2024, and measuring 2,491,699 kilobytes in size, conforming to 37 CFR §§ 1.831 to 1.835, which was submitted by mail as an XML file on a read-only optical disc (DVD) on April 25, 2024, for use in PCT application serial number PCT / US24 / 26282. 37 CFR § 1.835(a)(1). The Sequence List XML is incorporated herein by reference in its entirety. Background Technology

[0006] The human gut microbiome, a hallmark of complex adaptive systems (CAS), boasts trillions of microorganisms and exhibits rich phylogenetic diversity. This complex ecosystem not only maintains positive interactions with its host environment but also displays dynamic adaptability, playing a crucial role in maintaining health and modulating disease susceptibility. Among the astonishing microbial diversity present in the human gut, the concept of the "core microbiome" has garnered considerable attention. This core is hypothesized to combine microorganisms ubiquitous in healthy individuals, thus making significant contributions to maintaining homeostasis in nutrition, metabolism, immunity, and behavior. The indispensable role of this core microbiome, analogous to that of vital organs, underscores its importance in overall health management.

[0007] Historically, the classification of the core microbiome has primarily relied on assessments of the presence or absence of specific taxa or genes / pathways in cohorts of healthy individuals, supplemented by quantifications of their abundance or prevalence. While these approaches have undoubtedly provided important insights into the structural configuration and potential functional characteristics of the microbiome, they may be insufficient to represent the crucial ecological interactions that underscore the stability and resilience of this complex system. This oversight is particularly significant when considering the critical roles these interactions play in the onset, progression, and remission of various disease states.

[0008] As a CAS (Consciousness-Oriented System), the microbiome follows a modular design principle. The components of the CAS are organized into modules that interconnect to form a network. In the gut ecosystem, individual microbes are integrated into modular structures called microbiomes. Although each microbiome contains microbes with different taxonomic backgrounds, they act as coherent functional units or modules within the CAS of the microbiome. Members of the microbiome exhibit cooperative behavior through co-abundance, and different microbiomes may cooperate or compete to form ecological networks. Therefore, characterizing the core microbiome, in terms of the microbiome, becomes a promising and intriguing approach. Summary of the Invention

[0009] Throughout co-evolution, the gut microbiota has played a crucial role in maintaining human health. However, identifying core microbial components that reliably deliver fundamental health benefits remains a significant challenge. These core members are believed to maintain their ecological interactions, whether cooperative or competitive, despite changing environmental conditions. Based on a high-fiber intervention trial in patients with type 2 diabetes and 26 different case-control datasets, 284 high-quality metagenomic assemblies were identified, consistently forming stable genomic pairs among individuals regardless of dietary changes or disease progression. These genomes correspond to two microbiota, including the most resilient and highly interconnected bacteria, collectively associated with a wide range of health conditions. One microbiota is rich in genes related to plant polysaccharide degradation and butyrate production, while the other is characterized by high prevalence of genes associated with virulence and antibiotic resistance. Using these genomes as references, a random forest model skillfully distinguished between cases and controls across 15 different diseases and predicted patient responses to immunotherapy. Therefore, this core microbiome profile holds the potential to serve as a unified therapeutic target for enhancing health.

[0010] Individual microbial cells are considered fundamental components or reagents of CAS, representing the primary ecologically significant structural and functional units in the gut ecosystem. In some embodiments, high-quality metagenomically assembled genomes (HQMAGs) are used as an alternative to profiling these microbial cells, thus providing a more comprehensive and realistic description of the microbiome compared to approaches centered on genes, pathways, or taxonomic units. This perspective encompasses the full genetic potential and ecological characteristics of microorganisms, reinforcing the fundamental ecological axiom that organisms (or more precisely, cells) interact with and with their environment, rather than with genes / pathways or taxa.

[0011] To identify the core components of the gut microbiome, a genome-centric, reference-free approach was employed, emphasizing the stability of ecological interactions. This method involves detecting stable relationships between HQMAGs under different conditions, introducing environmental disturbances into the gut ecosystem via dietary interventions or disease progression. These stable relationships can reveal the core members of the microbiome. This aligns with a fundamental principle of systems biology that stable relationships often imply key system components. In the context of the gut microbiome, these core components likely perform essential functions that contribute to system resilience and host health, requiring their persistence and predictable interaction patterns. Therefore, revealing these stable relationships can illuminate these key microbial components, potentially revealing the pillars of conserved ecological networks across individuals, communities, or health states within the gut microbiome.

[0012] A robust seesaw-like network comprising two competing bacterial communities was identified. This was demonstrated through a high-fiber intervention pre- and post-intervention (QD test). Figure 1 A) The network is identified by searching for stable genomic pairs in the co-abundance networks between individuals or between healthy and diseased cohorts. This seesaw-like network embodies interactions of cooperation and competition and may indicate key features of a stable microbiome structure. HQMAGs identified in this novel core microbiome have demonstrated correlations with various clinical parameters in patients with type 2 diabetes (T2DM) receiving high-fiber interventions. Furthermore, general machine learning models based on these HQMAGs in the seesaw-like core microbiome successfully distinguished cases from controls in 26 independent datasets spanning 15 different diseases. Moreover, these HQMAGs support machine learning models for predicting personalized treatment responses to immunotherapy in patients with cancer or autoimmune diseases. This disclosure introduces novel concepts and analytical paradigms for studying the core gut microbiome. This paradigm provides enhanced health maintenance strategies and disease management, enabling personalized interventions to adapt to the complex interactions of microbial relationships within the gut ecosystem.

[0013] Therefore, one aspect of this disclosure provides a method and system for training a model for predicting a subject's response to a therapy. The method includes, at a computer system having at least one processor and a memory storing one or more programs executed by the one or more processors, acquiring, electronically for each of a plurality of training subjects, wherein each of the plurality of training subjects has received a therapy for a condition: (i) corresponding plurality of genomic abundance values ​​for the corresponding training subject at a time prior to receiving the therapy, wherein the corresponding plurality of genomic abundance values ​​include, for each of a plurality of gut microbiota, a corresponding value for the abundance of the genome of the corresponding gut microbiota in a corresponding biological sample from the gut of the corresponding training subject, and (ii) an indication of the corresponding training subject's response to the therapy. The method further includes, for each of a plurality of training subjects, inputting information about the corresponding training subject into a model comprising multiple parameters, wherein the model applies the multiple parameters to the information through at least 10,000 calculations to obtain a corresponding output from the model for the corresponding training subject, wherein the corresponding output includes a prediction of the corresponding training subject's response to the therapy, and the information about the corresponding training subject includes the corresponding genomic abundance value of each of a plurality of gut microbiota, and the plurality of gut microbiota are selected from Table 1, Table 2, or Figures 13A to 13XX The method also includes adjusting multiple parameters based on one or more differences between (i) the corresponding output from the model and (ii) the corresponding indication of the corresponding training subject's response to the therapy for each of the first plurality of training subjects.

[0014] Therefore, another aspect of this disclosure provides methods and systems for using models to predict a subject's response to a therapy. The method includes acquiring, in electronic form, a plurality of genomic abundance values ​​selected from Table 1, Table 2, or Table 3 at a computer system having one or more processors and memory storing one or more programs executed by the processors. Figures 13A to 13XX The method includes the abundance value of the genome of each corresponding gut microbiome from a variety of gut microbiota in a biological sample from the subject. The method also includes inputting multiple genome abundance values ​​into a model comprising multiple parameters, wherein the model applies the multiple parameters to the multiple genome abundance values ​​through at least 10,000 calculations to generate a prediction of the subject's response to the therapy as an output from the model.

[0015] As disclosed herein, any embodiments disclosed herein may be applied to any other aspect where applicable.

[0016] Further aspects and advantages of this disclosure will become apparent to those skilled in the art from the following detailed description, in which only illustrative embodiments of the disclosure are shown and described. As will be appreciated, this disclosure is capable of other and different embodiments, and certain details thereof can be modified in various obvious ways without departing from this disclosure. Therefore, the drawings and description are to be regarded in an illustrative rather than restrictive manner.

[0017] Therefore, one aspect of the present invention provides a method for training a model for predicting a subject's response to a therapy at a computer system having one or more processors and memory storing one or more programs to be executed by the one or more processors.

[0018] In some embodiments, the method includes, for each of a plurality of training subjects, acquiring electronically the following information, wherein each of the plurality of training subjects has received a therapy for a condition: (i) corresponding plurality of genome abundance values ​​for the corresponding training subject prior to receiving the therapy, wherein the corresponding plurality of genome abundance values ​​include, for each of a plurality of gut microbes, a corresponding value for the abundance of the genome of the corresponding gut microbe in a corresponding biological sample from the gut of the corresponding training subject, and (ii) an indication of the corresponding training subject’s response to the therapy.

[0019] In some such embodiments, the method includes sequencing the genomic DNA of a corresponding biological sample from the gut of each of a plurality of training subjects, thereby obtaining a plurality of at least 100,000 nucleic acid sequences.

[0020] In some such embodiments, the method includes, for each of a plurality of training subjects, electronically acquiring a corresponding plurality of at least 100,000 nucleic acid sequences for genomic DNA from a corresponding biological sample from the gut of the corresponding training subject.

[0021] In some such embodiments, the method includes determining, for each of a plurality of gut microbes, a corresponding value for the abundance of the genome of the corresponding gut microbe from a plurality of at least 100,000 nucleic acid sequences.

[0022] In some such embodiments, the method includes, for each of the plurality of training subjects, assembling, electronically, a plurality of gut microbial genomes corresponding to a plurality of at least 100,000 nucleic acid sequences from a metagenomic de novo sequence assembly, and for each of the plurality of gut microbes, calculating a corresponding value for the abundance of the genome of the corresponding gut microbe based on the universality of the corresponding nucleic acid sequence among the plurality of at least 100,000 nucleic acid sequences, the corresponding nucleic acid sequence being used to assemble the corresponding gut microbial genome of the plurality of gut microbial genomes corresponding to the corresponding gut microbe.

[0023] In some such embodiments, the method includes, for each of a plurality of training subjects, assigning each corresponding nucleic acid sequence from a plurality of at least 100,000 sequences to a corresponding gut microbe from a plurality of gut microbes, thereby generating a corresponding count of the corresponding nucleic acid sequence from the plurality of nucleic acid sequences assigned to the corresponding gut microbe for each corresponding gut microbe from the plurality of gut microbes, and determining a corresponding genomic abundance value for the corresponding gut microbe based on the corresponding count of the corresponding nucleic acid sequence assigned to the corresponding gut microbe for each corresponding gut microbe from the plurality of gut microbes.

[0024] In some such embodiments, a variety of gut microbiota includes those selected from Table 1, Table 2, or... Figures 13A to 13XX At least 20 types of gut microbiota.

[0025] In some such embodiments, the multiple gut microbiota comprises at least 20 microorganisms selected from Table 1, Table 2, or... Figures 13A to 13XX Microorganisms with a connectivity of at least 2.

[0026] In some such embodiments, the biological sample from the gut of the corresponding subject is a fecal sample from the corresponding training subject.

[0027] In some such embodiments, the therapy is a biological therapy, immunotherapy, chemotherapy, radiotherapy, gene therapy, hormone therapy, photodynamic therapy, targeted therapy, small molecule, antibody, polynucleotide, natural compound, immunomodulator, bone marrow therapy, stem cell therapy, surgical therapy, induction therapy, maintenance therapy, or a combination thereof.

[0028] In some such embodiments, the condition is selected from the group consisting of: type 2 diabetes, hypertension, schizophrenia, atherosclerotic cardiovascular disease (ACVD), cirrhosis (LC), inflammatory bowel disease (IBD), colorectal cancer (CRC), ankylosing spondylitis (AS) and Parkinson's disease (PD), inflammatory bowel disease (IBD), rheumatoid arthritis (RA), or advanced melanoma and B-cell lymphoma.

[0029] In some such embodiments, the condition is cancer.

[0030] In some embodiments, the method includes, for each of a plurality of training subjects, inputting information about the respective training subject into a model comprising a plurality of parameters, wherein the model applies the plurality of parameters to the information through at least 10,000 computations to obtain a corresponding output from the model for the respective training subject, wherein the corresponding output includes a prediction of the respective training subject's response to the therapy, and the information about the respective training subject includes a corresponding genomic abundance value for each of a plurality of gut microbiota, wherein the plurality of gut microbiota is selected from Table 1, Table 2, or Figures 13A to 13XX .

[0031] In some such embodiments, the prediction of the response of the corresponding training subject is the category output of the corresponding response among a plurality of possible responses of the corresponding training subject.

[0032] In some such embodiments, the prediction of the response of the corresponding training subject is a probability output for the response of the corresponding training subject.

[0033] In some such embodiments, the model is a neural network algorithm, support vector machine algorithm, naive Bayes algorithm, nearest neighbor algorithm, boosting tree algorithm, random forest algorithm, convolutional neural network algorithm, decision tree algorithm, regression algorithm, or clustering algorithm.

[0034] In some such embodiments, the plurality of parameters is at least 1,000, at least 10,000, at least 15,000, at least 50,000, at least 100,000, at least 250,000, at least 500,000, or at least 1,000,000 parameters.

[0035] In some such embodiments, the model applies the plurality of parameters to the information through at least 25,000, at least 50,000, at least 100,000, at least 250,000, at least 500,000, or at least 1,000,000 calculations to obtain a corresponding output from the model for the respective training subject.

[0036] In some embodiments, the method includes adjusting a plurality of parameters based on one or more differences between (i) a corresponding output from the model and (ii) a corresponding indication of the corresponding training subject’s response to the therapy for each of a plurality of training subjects.

[0037] Another aspect of this disclosure provides a method for using a model at a computer system to predict a subject’s response to a therapy for a condition, the computer system having one or more processors and memory storing one or more programs executed by the one or more processors.

[0038] In some such embodiments, the method includes acquiring a plurality of genomic abundance values ​​in electronic form, the plurality of genomic abundance values ​​including those selected from Table 1, Table 2, or Figures 13A to 13XX The abundance value of the genome of each corresponding gut microbe in the multiple gut microbiota from the subject's biological sample.

[0039] In some such embodiments, the method includes sequencing the genomic DNA of the biological sample from the subject's gut to obtain the plurality of at least 100,000 nucleic acid sequences.

[0040] In some such embodiments, the method includes electronically acquiring at least 100,000 nucleic acid sequences targeting genomic DNA from the biological sample of the subject’s gut.

[0041] In some such embodiments, the method includes determining a corresponding value for the abundance of the genome of a given gut microbe from a plurality of at least 100,000 nucleic acid sequences for each given gut microbe among a plurality of gut microbes.

[0042] In some such embodiments, the method includes electronically assembling a plurality of gut microbial genomes from a plurality of at least 100,000 nucleic acid sequences via de novo metagenomic sequencing, and for each corresponding gut microbial among the plurality of gut microbes, calculating a corresponding value for the abundance of the genome of the corresponding gut microbe based on the universality of the corresponding nucleic acid sequence among the plurality of at least 100,000 nucleic acid sequences, the corresponding nucleic acid sequence being used to assemble the corresponding gut microbial genome from the plurality of gut microbial genomes corresponding to the corresponding gut microbe.

[0043] In some such embodiments, the method includes assigning each corresponding nucleic acid sequence from the plurality of at least 100,000 sequences to a corresponding gut microbe from the plurality of gut microbes, thereby generating a corresponding count of the corresponding nucleic acid sequence from the plurality of nucleic acid sequences assigned to the corresponding gut microbe for each corresponding gut microbe from the plurality of gut microbes, and determining a corresponding genomic abundance value for the corresponding gut microbe based on the corresponding count of the corresponding nucleic acid sequence assigned to the corresponding gut microbe for each corresponding gut microbe from the plurality of gut microbes.

[0044] In some such embodiments, the multiple gut microbiota comprises at least 20 microorganisms selected from Table 1, Table 2, or... Figures 13A to 13XX Microorganisms with a connectivity of at least 2.

[0045] In some such embodiments, the biological sample from the subject's intestines is a fecal sample.

[0046] In some such embodiments, the therapy is a biological therapy, immunotherapy, chemotherapy, radiotherapy, gene therapy, hormone therapy, photodynamic therapy, targeted therapy, small molecule, antibody, polynucleotide, natural compound, immunomodulator, bone marrow therapy, stem cell therapy, surgical therapy, induction therapy, maintenance therapy, or a combination thereof.

[0047] In some such embodiments, the condition is selected from the group consisting of: type 2 diabetes, hypertension, schizophrenia, atherosclerotic cardiovascular disease (ACVD), cirrhosis (LC), inflammatory bowel disease (IBD), colorectal cancer (CRC), ankylosing spondylitis (AS) and Parkinson's disease (PD), inflammatory bowel disease (IBD), rheumatoid arthritis (RA), advanced melanoma and B-cell lymphoma.

[0048] In some such embodiments, the condition is cancer.

[0049] In some embodiments, the method includes inputting multiple genomic abundance values ​​into a model that includes multiple parameters, wherein the model applies the multiple parameters to the multiple genomic abundance values ​​through at least 10,000 calculations to generate a prediction of the subject’s response to the therapy as an output from the model.

[0050] In some such embodiments, the prediction of a subject's response is the output of the category of the corresponding response among a plurality of possible responses of the subject.

[0051] In some such embodiments, the prediction of a subject's response is a probability output for the corresponding subject's response.

[0052] In some such embodiments, the model is a neural network algorithm, support vector machine algorithm, naive Bayes algorithm, nearest neighbor algorithm, boosting tree algorithm, random forest algorithm, convolutional neural network algorithm, decision tree algorithm, regression algorithm, or clustering algorithm.

[0053] In some such embodiments, the plurality of parameters is at least 1,000, at least 10,000, at least 15,000, at least 50,000, at least 100,000, at least 250,000, at least 500,000, or at least 1,000,000 parameters.

[0054] In some such embodiments, the model applies the multiple parameters to the information by calculating at least 25,000, at least 50,000, at least 100,000, at least 250,000, at least 500,000, or at least 1,000,000 times to obtain a corresponding output from the model for the respective subject.

[0055] In some embodiments, the method includes treating the subject by: i) administering the therapy to the subject when the predicted response to the therapy meets a threshold probability that the subject will have a favorable response to the therapy; ii) administering one or more of a variety of gut microbiota to the subject when the predicted response to the therapy does not meet a threshold probability that the subject will have a favorable response to the therapy.

[0056] Another aspect of this disclosure provides a computer system. The computer system includes one or more processors and a non-transitory computer-readable medium including computer-executable instructions that, when executed by the one or more processors, cause the processors to perform the methods described herein.

[0057] Another aspect of this disclosure provides a non-transitory computer-readable storage medium. This non-transitory computer-readable storage medium stores instructions that, when executed by a computer system, cause the computer system to perform any of the methods described herein. Attached Figure Description

[0058] Figure 1 A block diagram illustrating an exemplary computing device according to some embodiments of the present disclosure is provided.

[0059] Figure 2A , Figure 2B , Figure 2C and Figure 2D Together, flowcharts are provided of processes and features for training a model for predicting a subject’s response to a therapy for a disease, according to some embodiments of this disclosure.

[0060] Figure 3A , Figure 3B and Figure 3C Together, flowcharts are provided for processes and characteristics of predicting a subject’s response to a therapy for a condition, according to some embodiments of this disclosure.

[0061] Figure 4A , Figure 4B , Figure 4C , Figure 4D , Figure 4E , Figure 4F , Figure 4G and Figure 4HThis study collectively illustrates the association between reversible alterations in the gut microbiota induced by a high-fiber diet and corresponding changes in the metabolic phenotype of patients with type 2 diabetes mellitus (T2DM). (A) Study design of the QD trial. Written informed consent, personal information questionnaires, and HbA1c-based screening were conducted during induction. After induction, physical examinations and sample collection were performed at baseline (M0), three months after high-fiber intervention (W) or a regular diet (U) (M3), and one year after cessation of high-fiber intervention (M15). (B) Changes in fiber intake. (C) Overall changes in the gut microbiota as shown by principal coordinate analysis based on Bray-Curtis distances of 1845 HQMAGs, and (D) Mean Bray-Curtis distances between groups. PERMANOVA (9,999 permutations) was performed to compare groups. *P < 0.05 and ***P < 0.001. The color intensity of the squares indicates the magnitude of the mean Bray-Curtis distance. (E) Changes in HbA1c, (F) Percentage of participants achieving adequate glycemic control, (G) Fasting blood glucose, and (H) Area under the glucose curve (AUC) in the meal tolerance test (MTT). For (E), (G), and (H), data are presented as percentage changes relative to baseline (±SEM). Comparisons within the same group were performed using the Friedman test followed by the Nemenyi post-hoc test, with compact letters reflecting significance (P<0.05). n=67 in group W and n=28 in group U. Comparisons between W and U at the same time point were performed using the Mann-Whitney test (two-tailed), *P<0.05, **P<0.01, and ***P<0.001. In W, n=74 (M0) (for graph H, n=72), in W, n=74 (M3), in W, n=67 (M15), in U, n=36 (M0), in U, n=36 (M3) and in U, n=28 (M15).

[0062] Figure 5A and Figure 5B The combined illustration shows that despite substantial global changes in the gut microbiota induced by high-fiber intervention, two competing bacterial communities associated with HbA1c levels formed a robust seesaw-like network within the ecosystem. (A) Distribution of correlations between different types of genome pairs during the trial. Three letters indicate the correlations of genome pairs subsequently at M0, M3, and M15 (N for negative correlation, P for positive correlation, and U for no correlation). Stable correlations, NNN and PPP, are highlighted. (B) Correlation between genome clusters and HbA1c using a linear mixed-effects model with the MaAslin2 package. Abundance was log-transformed. Subjects were used as random effects. N=67. *BH-adjusted P<0.05, ***BH-adjusted P<0.001.

[0063] Figure 6A , Figure 6B , Figure 6C1 , Figure 6C2 , Figure 6D , Figure 6E1 , Figure 6E2 , Figure 6E3 , Figure 6E4 , Figure 6E5 , Figure 6E6 , Figure 6E7 , Figure 6E8 , Figure 6E9 , Figure 6E10 and Figure 6E11This study illustrates how the genomes of two competing microbiota predicted metabolic health outcomes in patients with type 2 diabetes mellitus (T2DM) in a QD trial, and distinguished cases from controls across seven diseases in eleven independent case-control metagenomic datasets (Case-Control Dataset Set I). (A) Changes in mean range-scale robust central log ratio (rclr) transformation abundance and its proportion in the W group trial for microbiota 1 and 2. Differences between time points were analyzed using the Friedman test followed by the Nemenyi test. The compactness of the letters reflects significance of P < 0.05. (B) Prediction of clinical parameters using the genomes of two competing microbiota. A linear mixed-effects model was trained with subjects as random effects, based on the mean range-scale rclr abundance of microbiota 1 and 2 and clinical parameters at M0 and M3. Clinical parameters were predicted based on the trained model using the mean range-scale rclr abundance of microbiota 1 and 2 at M15. The bar graph shows the Pearson correlation coefficient between the predicted clinical parameters and the measured values ​​at M15. An asterisk preceding the parameter name indicates the significance of the Pearson correlation. P values ​​were adjusted using the Benjamini and Hochberg method. *Adjusted P < 0.05, **Adjusted P < 0.01, and ***Adjusted P < 0.001. BMI, Body Mass Index; SBP, Systolic Blood Pressure; DBP, Diastolic Blood Pressure; WC, Waist Circumference; HP, Hip Circumference; TNF-α, Tumor Necrosis Factor-α; WBC, White Blood Cell Count; CRP, C-Reactive Protein; LBP, Lipopolysaccharide-Binding Protein; TC, Total Cholesterol; TG, Triglycerides; Lpa, Lipoprotein A; HDL, High-Density Lipoprotein; APOA, Apolipoprotein A; LDL, Low-Density Lipoprotein; APOB, Apolipoprotein B; GFR (MDRR), Glomerular Filtration Rate; CysC, Cystatin C; ACR, Urinary Microalbumin to Creatinine Ratio; IM T, intima-media thickness; DAN, diabetic autonomic neuropathy score; MHR, mean heart rate; SDNN, standard deviation of NN intervals; SDANN, standard deviation of the mean NN interval calculated over 5 minutes; SDNNIndex, mean of the standard deviations of NN intervals in 5-minute segments; rMSSD, root mean square of the difference between consecutive NN intervals; pNN50, percentage of the difference between consecutive NN intervals greater than 50 ms; TP, total power; VLF, very low frequency power; LF, low frequency power; HF, high frequency power; DPN, diabetic peripheral neuropathy score. (C) Differences in the heritability of carbohydrate substrate utilization (CAZy), short-chain fatty acid production (SCFA), antibiotic resistance genes (ARG), and virulence factor genes (VF). Heatmaps show the proportion (CAZy) or gene copy number (SCFA, ARG, and VF) of each class in each genome. For carbohydrate substrate utilization, CAZy genes were predicted in each genome.The proportion of CAZy genes for a specific substrate is calculated by dividing the number of CAZy genes involved in its utilization by the total number of CAZy genes. The CAZy families associated with arabinoxylan are: CE1, CE2, CE4, CE6, CE7, GH10, GH11, GH115, GH43, GH51, GH67, GH3, and GH5; those associated with cellulose are: GH1, GH44, GH48, GH8, GH9, GH3, and GH5; those associated with inulin are: GH32 and GH91; and those associated with mucins are: GH1, GH2, GH3, GH4, GH18, GH19, and GH20. GH29, GH33, GH38, GH58, GH79, GH84, GH85, GH88, GH89, GH92, GH95, GH98, GH99, GH101, GH105, GH109, ​​GH110, GH113, PL6, PL8, PL12, PL13 and PL21; pectin-related: CE12, CE8, GH28, PL1 and PL9; starch-related: GH13, GH31 and GH97. For short-chain fatty acid production, FTHFS (formate-tetrahydrofolate ligase) was used for acetate production; ScpC (propionyl-CoA succinate-CoA transferase) and Pct (propionate-CoA transferase) were used for propionate production; But (butyryl-CoA): butyrate-CoA transferase, Buk (butyrate kinase), 4Hbt (butyryl-CoA): 4-hydroxybutyrate-CoA transferase, and Ato (butyryl-CoA): acetoacetate-CoA transferase (AtoA: α-subunit, AtoD: β-subunit) were used for butyrate production. The Mann-Whitney test (two-tailed) was used to analyze the differences between community 1 and community 2. #P < 0.1, *P < 0.05, **P < 0.01, and ***P < 0.001. Number of genomes in community 1 (green bar): n = 50; number of genomes in community 2 (purple bar): n = 91. (D) Datasets from 11 different datasets across 7 diseases, including type 2 diabetes (T2D), cirrhosis (LC), ankylosing spondylitis (AS), atherosclerotic cardiovascular disease (ACVD), schizophrenia (SCZ), colorectal cancer (CRC), and inflammatory bowel disease (IBD), were collected as case-control dataset set I. The sample size for each case-control dataset is shown in red and green numbers in the left figure, respectively. For each dataset in the set, metagenomic reads were recruited into 141 genomes from two competing bacterial communities in QD to estimate the genomic abundance in each sample. (E) Based on the abundance matrix of the 141 genomes, a random forest classification model with leave-one-out cross-validation was trained to classify case and control subjects in each dataset. The ROC curves and area under the curve (AUC) are shown here.

[0064] Figure 7A and Figure 7B The genomes forming two competing microbial communities, as identified from case-control datasets specific to a single disease, are jointly illustrated, demonstrating significant effectiveness in classifying cases and controls across independent datasets of different diseases within Case-Control Dataset Set I. (A) Identification of two competing microbial communities in a seesaw network in Case-Control Dataset Set I. Sample sizes for each case-control study are shown in red and green numbers in the left figure, respectively. Case-Control Dataset Set I comprises 11 published metagenomic case-control datasets covering seven diseases, including type 2 diabetes (T2D), cirrhosis (LC), ankylosing spondylitis (AS), atherosclerotic cardiovascular disease (ACVD), schizophrenia (SCZ), colorectal cancer (CRC), and inflammatory bowel disease (IBD). Datasets from three studies were combined for CRC analysis. Datasets from two studies were combined for IBD analysis. The correlation percentages of two competing microbial communities following a seesaw network pattern (i.e., positive edges within each community and negative edges between the two communities) are shown in yellow, and the ratio of negative correlation within each community to positive correlation between the communities is shown as stacked black bars representing 100%. (B) Two competing microbial communities identified in one dataset were used as predictors to classify cases and controls in the same dataset and all other 10 datasets. A random forest classification model with leave-one-out cross-validation was applied to each dataset for each of the two competing microbial communities. The area under the ROC curve (AUC) values ​​are shown in the heatmap.

[0065] Figure 8A , Figure 8B1 , Figure 8B2 , Figure 8B3 , Figure 8B4 , Figure 8B5 , Figure 8B6 , Figure 8B7 , Figure 8B8 , Figure 8B9 , Figure 8B10 , Figure 8B11 , Figure 8B12 , Figure 8B13 , Figure 8B13 , Figure 8B14 , Figure 8B15 , Figure 8B16 , Figure 8C1 and Figure 8C2A combined core genome from all identified competing microbiota is illustrated, demonstrating its ability to effectively distinguish cases from controls across a wider range of diseases and predict treatment outcomes in independent datasets. (A) Combined genomes and combined core genomes were identified from eight groups of two competing microbiota generated from the QD dataset and the case-control dataset set I. All HQMAGs in each group of the two competing microbiota were deduplicated based on a cutoff of 99% average nucleotide identity (ANI) between the two genomes. 788 non-redundant HQMAGs were obtained as the combined genomes for all eight groups of the two competing microbiota. Random forest classification models with leave-one-out cross-validation were constructed based on the 788 HQMAGs in each dataset. HQMAGs were ranked based on their importance across all models. One HQMAG was subsequently removed from the least important HQMAG (ranked highest in importance) to perform the random forest classification model in each dataset. In each dataset, the number of HQMAGs was ranked according to the area under the ROC curve (AUC) value. A scatter plot shows the relationship between the number of HQMAGs and model performance. The y-axis represents the sum of rankings based on AUC values ​​(smaller values ​​indicate better performance). 302 HQMAGs achieved optimal performance. After excluding 18 HQMAGs that exhibited inconsistent C1A and C1B assignments across datasets, a total of 284 HQMAGs were retained from the 302 HQMAGs as the combined core genome for all 8 groups from both competing bacterial populations. (B) Using the combined core as a predictor in Case-Control Dataset Collection II, which contains 15 published metagenomic case-control datasets for 10 diseases, including one dataset for ankylosing spondylitis (AS#2): cases n=85, controls n=55; one dataset for autism spectrum disorder (ASD): cases n=64, controls n=64; one dataset for Behcet's disease (BD): cases n=24, controls n=52; one dataset for COVID-19: cases n=47, controls n=19; and three datasets for colorectal cancer (CRC): CRC#4 cases n=40, controls n=40; ​​CRC#5 cases n =61 cases, control n=52; CRC#6 cases n=52, control n=52; 1 case on Graves' disease (GD), case n=88, control n=62; 2 cases on hypertension (HT): HT#1 cases n=60, control n=56, HT#2 cases n=99, control n=41; 1 case on multiple sclerosis (MS): case n=24, control n=24; 3 cases on pancreatic cancer (PC): PC#1 cases n=43, control n=235, PC#2 cases n=57, control n=50, PC#3 cases n=44, control n=32; and 1 case on Parkinson's disease (PD): case n=39, control n=40. A random forest classification model with leave-one-out cross-validation is applied to each dataset.(C) The combined core genome was used as a predictor in the treatment dataset set to predict responders (R) and non-responders (NR) under treatment. For inflammatory bowel disease (IBD), 14-week remission was used to determine R and NR for patients with IBD receiving anti-cytokine or anti-antigen therapy. IBD_anti-cytokine: Rn=29, NRn=18; IBD_anti-integrin #1: Rn=27, NRn=40; ​​IBD_anti-integrin #2: Rn=29, NRn=53. For rheumatoid arthritis (RA), responders to methotrexate (MTX) were a priori defined as any newly diagnosed RA patient with a disease activity score (DAS28) of ≥1.8 in 28 joints at month 4 after initiation of MTX monotherapy. Rn=19, NRn=28. For advanced melanoma, progression-free survival was used to determine R and NR for immune checkpoint inhibitor (ICI) therapy. AM_ICI#1: Rn=4, NRn=7; AM_ICI#2: Rn=10, NRn=8; AM_ICI#3: Rn=12, NRn=13; AM_ICI#4: Rn=25, NRn=30; AM_ICI#5: Rn=26, NRn=28. For B-cell lymphoma, the treating physician categorized the tumor response to CAR-T cell immunotherapy as complete or incomplete remission (partial remission, stable disease, disease progression, or death) 180 days after CAR-T cell infusion. The model was trained in a German cohort and validated in a US cohort. Germany: Rn=21, NRn=29; USA: Rn=21, NRn=24.

[0066] Figure 9A , Figure 9B and Figure 9CThe combined core genome from all eight groups of two competing microbiota is illustrated in both the case-control datasets I and II, demonstrating its discriminatory power in classifying healthy individuals versus patients for the colorectal cancer (CRC), inflammatory bowel disease (IBD), and pancreatic cancer (PC) datasets. Prediction matrices are shown for the combined core genome from all eight groups of two competing microbiota within each dataset (diagonal values), across dataset pairs (one dataset used for model training and the other for testing), and in a setting where one dataset is retained (the model is trained on all datasets except one and tested on the retained dataset). A random forest classification model with leave-one-out cross-validation is applied. The area under the ROC curve (AUC) values ​​are shown in the matrix. (A) CRC. #1: Case n=74, Control n=54; #2: Case n=46, Control n=63; #3: Case n=22, Control n=60; #4: Case n=40, Control n=40; ​​#5: Case n=61, Control n=52; #6: Case n=52, Control n=52. (B) IBD. #1: Case n=80, Control n=26; #2: Case n=121, Control n=34; #3: Case n=43, Control n=22; (C) PC. #1: Case n=43, Control n=235; #2: Case n=57, Control n=50; #3: Case n=44, Control n=32.

[0067] Figure 10A1 , Figure 10A2 , Figure 10B1 , Figure 10B2 , Figure 10C1 , Figure 10C2 , Figure 10D1 and Figure 10D2The combined core genomes of two competing microbiota support the prediction of treatment outcomes in a ensemble of treatment datasets for inflammatory bowel disease, rheumatoid arthritis, advanced melanoma, and B-cell lymphoma. The abundance of the combined core genomes (284 HQMAGs) in the pre-treatment samples was used as a predictor in a random forest classification model to predict responders (R) and non-responders (NR) under treatment. The area under the ROC curve (AUC) and AUC values ​​are shown in the figure. (A) R and NR were determined using 14-week remission. IBD_anti-cytokine, Rn=29, NRn=18; IBD_anti-integrin #1, Rn=27, NRn=40; ​​IBD_anti-integrin #2, Rn=29, NRn=53. (B) Responders to MTX were a priori defined as any new-onset RA patient with a disease activity score (DAS28) (25) improvement ≥1.8 at month 4 after initiation of MTX monotherapy. Rn=19, NRn=28. (C) Overall response rate (ORR, left matrix) and progression-free survival (PFS12, right matrix) were used to determine R and NR, respectively. Prediction matrices based on microbiome-based response predictions were assessed via ORR (left matrix) and PFS12 (right matrix), in each cohort (values ​​on the diagonal), across cohort pairs (one cohort used for model training and the other for testing), and leave-one-cohort settings (model trained on all but one cohort and tested on the leave-one cohort). ORR: Rn=94, NRn=71; PFS12: Rn=77, NRn=86. (D) Therapists classified tumor response to CAR-T cell immunotherapy at 180 days post-CAR-T cell infusion as complete or incomplete response (partial response, stable disease, disease progression, or death). The model was trained in #1 (German cohort) and validated in #2 (US cohort). #1: Rn=21, NRn=29; #2: Rn=21, NRn=24.

[0068] Figure 11A Figure 11A2 Figure 11B1 , Figure 11B2 , Figure 11C1 , Figure 11C2 , Figure 11D1 and Figure 11D2A combined core genome of two competing bacterial communities was illustrated, providing a general model for distinguishing cases and controls across various diseases (Case-Control Datasets I and II). (A) All control and case samples from Case-Control Datasets I and II (comprising a total of 26 datasets for 15 different diseases) were combined and randomly assigned, with 80% used for training the random forest classification model and 20% for testing. (B) ROC curves and area under the ROC curve (AUC). (C) Density plot of probability scores between cases and controls. The probability scores were generated by the random forest classification model and show the probability that a sample is predicted as a case. (D) Box plot of probability scores between control and case samples. The Mann-Whitney test was applied. ***P<0.001. Training: Controls n=1,285, Cases n=1424; Testing: Controls n=319, Cases n=356.

[0069] Figure 12A , Figure 12B , Figure 12C , Figure 12D , Figure 12E , Figure 12F , Figure 12G , Figure 12H and Figure 12I Together they illustrate the corresponding contigs obtained for each of the 788 genomes (referenced by SEQ ID).

[0070] Figure 13A , Figure 13B , Figure 13C , Figure 13D , Figure 13E , Figure 13F , Figure 13G , Figure 13H , Figure 13I , Figure 13J , Figure 13K , Figure 13L , Figure 13M , Figure 13N , Figure 13O , Figure 13P , Figure 13Q , Figure 13R , Figure 13S , Figure 13T , Figure 13U , Figure 13V , Figure 13W , Figure 13X , Figure 13Y , Figure 13Z , Figure 13AA , Figure 13BB , Figure 13CC , Figure 13DD , Figure 13EE , Figure 13FF , Figure 13GG , Figure 13HH , Figure 13II , Figure 13JJ , Figure 13KK , Figure 13LL , Figure 13MM , Figure 13NN , Figure 1300 , Figure 13PP , Figure 13 QQ , Figure 13RR , Figure 13SS , Figure 13TT , Figure 13UU , Figure 13VV , Figure 13WW and Figure 13XX Together they illustrate the taxonomic allocation of 788 genomes in the composite microbiome.

[0071] Figure 14A and Figure 14B A joint illustration shows how CC-TCG predicts treatment outcomes in independent datasets. CC-TCG was used as a predictor in the treatment dataset ensemble to predict whether responders (R) and non-responders (NR) were receiving treatment. For inflammatory bowel disease (IBD), 14-week remission was used to determine R and NR for patients with IBD receiving anti-cytokine or anti-antigen therapy. For rheumatoid arthritis (RA), responders to methotrexate (MTX) were a priori defined as any newly diagnosed RA patient whose disease activity score (DAS28) for 28 joints improved to R1.8 at four months after initiation of MTX monotherapy. For advanced melanoma (AM), progression-free survival was used to determine R and NR for immune checkpoint inhibitor (ICI) therapy. For B-cell lymphoma, the treating physician categorized tumor response to CAR-T cell immunotherapy as complete or incomplete (partial remission, stable disease, disease progression, or death) at 180 days after CAR-T cell infusion. The model was trained in a German cohort and validated in a US cohort. Sample sizes for R and NR are listed.

[0072] Figure 15A , Figure 15B , Figure 15C and Figure 15DThe combined core of TCG is illustrated to predict treatment outcomes in a ensemble of treatment datasets for inflammatory bowel disease, rheumatoid arthritis, advanced melanoma, and B-cell lymphoma. The abundance of the combined core genome (284 HQMAGs) in the pre-treatment samples was used as a predictor in a random forest classification model to predict responders (R) and non-responders (NR) under treatment. The area under the ROC curve (AUC) and AUC values ​​are shown in the figure. (15A) R and NR were determined using 14-week remission. IBD_anticytokine, n=29 for R, n=18 for NR; IBD_antiintegrin #1, Rn=27, NRn=40; ​​and IBD_antiintegrin #2, Rn=29, NRn=53. (15B) Responders to MTX were defined a priori as any new-onset RA patient whose disease activity score (DAS28) (25) improved to R1.8 at month 4 after initiation of MTX monotherapy. Rn=19, NRn=28. (15C) Overall response rate (ORR, left matrix) and progression-free survival (PFS12, right matrix) were used to determine R and NR, respectively. Prediction matrices of microbiome-based response predictions were assessed via ORR (left matrix) and PFS12 (right matrix), in each cohort (values ​​on the diagonal), across cohort pairs (one cohort used to train the model and the other for testing), and leave-one-cohort settings (the model was trained on all but one cohort and tested on the leave-one cohort). ORR: Rn=94, NRn=71; PFS12: Rn=77, NRn=86. (15D) Therapeutic physicians classified tumor response to CAR-T cell immunotherapy as complete or incomplete response (partial response, stable disease, disease progression, or death) at 180 days after CAR-T cell infusion. The model was trained in #1 (German cohort) and validated in #2 (US cohort). #,1: Rn=21, NRn=29; #,2: Rn=21, NRn=24.

[0073] Figure 16A and Figure 16B The combined illustrations demonstrate the functional concordance at the community level among the eight TCGs identified in the QD test and CCDC-I. (16A) The PCoA plot based on Euclidean distance shows a significant separation between C1A and C1B in the eight TCG groups. (16B) Ward association based on Euclidean distance shows the separation between C1A and C1B in the eight TCG groups. The Euclidean distance was calculated based on the average copy number of the KEGG ortholog in each community.

[0074] Figure 17A , Figure 17B , Figure 17C , Figure 17D , Figure 17E and Figure 17F The combined illustration shows the differences in genes associated with carbohydrate utilization and SCFA production between C1A and C1B. Bar plots (17A, 17C, and 17E) show the proportion of CAZy genes for a specific substrate, calculated as the number of CAZy genes involved in its utilization divided by the total number of CAZy genes. The CAZy families associated with arabinoxylan are: CE1, CE2, CE4, CE6, CE7, GH10, GH11, GH115, GH43, GH51, GH67, GH3, and GH5; those associated with cellulose are: GH1, GH44, GH48, GH8, GH9, GH3, and GH5; those associated with inulin are: GH32 and GH91; and those associated with mucin are: GH1, GH2, GH3, GH4, GH18, GH19, GH20, and GH5. 29, GH33, GH38, GH58, GH79, GH84, GH85, GH88, GH89, GH92, GH95, GH98, GH99, GH101, GH105, GH109, ​​GH110, GH113, PL6, PL8, PL12, PL13 and PL21; pectin-related: CE12, CE8, GH28, PL1 and PL9; and starch-related: GH13, GH31 and GH97. Bar plots (17B, 17D, and 17F) show the copy numbers of genes encoding FTHFS: formate-tetrahydrofolate ligase for acetate production; ScpC: propionyl-CoA (CoA) succinate-CoA transferase and Pct: propionate-CoA transferase for propionate production; But, butyryl-CoA (butyryl-CoA): acetate-CoA transferase; Buk, butyrate kinase; 4Hbt, butyryl-CoA: 4-hydroxybutyrate-CoA transferase; and Ato, butyryl-CoA: acetoacetate-CoA transferase (AtoA: α subunit, AtoD: β subunit) for butyrate production. Differences between C1A1 and C1B2 were analyzed using the Mann-Whitney test (two-tailed). #p<0.1, *p<0.05, **p<0.01, ***p<0.001.

[0075] In the various views of the accompanying drawings, the same reference numerals refer to the corresponding parts. Detailed Implementation

[0076] The methods and systems described in this article help predict a subject's response to treatments for a condition based on the composition of the subject's microbiome.

[0077] definition.

[0078] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in the description of the invention and the appended claims, unless the context clearly indicates otherwise, the singular forms “a / an” and “the” are intended to equally include the plural forms. It should also be understood that the term “and / or” as used herein refers to and covers any and all possible combinations of one or more of the associated listed items. It should be further understood that, when used in this specification, the terms “comprising,” “including,” or any variations thereof specifically describe the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. Furthermore, in the context of the specific embodiments and / or claims, the terms “including / include,” “having / has / with,” or variations thereof are intended to be inclusive in a manner similar to the term “comprising.”

[0079] As used herein, the term "if" can be interpreted as "when," "at," "in response to determination," or "in response to detection," depending on the context. Similarly, the phrase "if determination" or "if [the stated condition or event] is detected" can be interpreted as "when determination," "in response to determination," "when [the stated condition or event] is detected," or "in response to the detection of [the stated condition or event]," depending on the context.

[0080] It should also be understood that although the terms first, second, etc., may be used herein to describe multiple elements, these elements should not be limited to these terms. These terms are used only to distinguish one element from another. For example, without departing from the scope of this disclosure, a first subject may be referred to as a second subject, and similarly, a second subject may be referred to as a first subject. Although both the first subject and the second subject are subjects, they are not the same subject. The terms “subject,” “user,” and “patient” are used interchangeably herein.

[0081] As used in this article, a "measure of central tendency" refers to the central or representative value of a distribution of values. Non-restricted examples of measures of central tendency include the arithmetic mean, weighted mean, median, midline, three-mean, geometric mean, geometric median, shortened mean, median, and mode of a distribution of values.

[0082] As used herein, the term "subject" means any living or non-living organism, including but not limited to humans (e.g., human males, human females, fetuses, pregnant women, children, etc.), non-human mammals, or non-human animals. Any human or non-human animal may serve as a subject, including but not limited to mammals, reptiles, birds, amphibians, fish, ungulates, ruminants, bovids (e.g., cattle), equines (e.g., horses), caprines and sheep (e.g., sheep, goats), suidae (e.g., pigs), camelids (e.g., camels, llamas, alpacas), monkeys, apes (e.g., gorillas, chimpanzees), bears (e.g., bears), poultry, dogs, cats, mice, rats, fish, dolphins, whales, and sharks. In some embodiments, the subject is a male or female of any age (e.g., a man, a woman, or a child).

[0083] As used herein, the term "application" in relation to the methods of the present invention means a method for therapeutic or preventative prevention, treatment, or improvement of a syndrome, symptom, or disease as described herein. Such methods include the simultaneous application of an effective amount of the therapeutic agent at different times during a treatment process or in combination. The methods of the present invention should be understood to include all known therapeutic treatment regimens.

[0084] As used herein, the terms “cancer,” “cancerous tissue,” or “tumor” refer to an abnormal mass of tissue in which the growth of the mass exceeds that of normal tissue and is not in harmony with the growth of normal tissue, including both solid masses (e.g., as in solid tumors) and fluid masses (e.g., as in hematologic malignancies). Cancer or tumor can be defined as “benign” or “malignant” based on the following characteristics: degree of cellular differentiation (including morphology and function), growth rate, local invasion, and metastasis. A “benign” tumor is well differentiated and is characterized by slower growth than a malignant tumor and remains in its original site. Furthermore, in some cases, benign tumors do not have the ability to infiltrate, invade, or metastasize to distant sites. A “malignant” tumor may be poorly differentiated (anaplastic) and is characterized by rapid growth with progressive infiltration, invasion, and destruction of surrounding tissues. Furthermore, malignant tumors may have the ability to metastasize to distant sites. Thus, cancer cells are cells found in abnormal masses of tissue whose growth is not in harmony with the growth of normal tissue. Thus, a “tumor sample” refers to a biological sample of a tumor obtained from or derived from a subject, as described herein.

[0085] Non-limiting examples of cancer types include ovarian cancer, cervical cancer, uveal melanoma, colorectal cancer, colorophobic renal cell carcinoma, liver cancer, endocrine tumors, oropharyngeal cancer, retinoblastoma, cholangiocarcinoma, adrenal cancer, neurocarcinoma, neuroblastoma, basal cell carcinoma, brain cancer, breast cancer, non-clear cell renal cell carcinoma, glioblastoma, glioma, kidney cancer, gastrointestinal stromal tumor, medulloblastoma, bladder cancer, stomach cancer, bone cancer, and non-small cell lung cancer. Thymoma, prostate cancer, clear cell renal cell carcinoma, skin cancer, thyroid cancer, sarcoma, testicular cancer, head and neck cancer (e.g., head and neck squamous cell carcinoma), meningioma, peritoneal cancer, endometrial cancer, pancreatic cancer, mesothelioma, esophageal cancer, small cell lung cancer, HER2-negative breast cancer, ovarian serous carcinoma, HR+ breast cancer, uterine serous carcinoma, endometrial carcinoma of the uterine body, adenocarcinoma of the gastroesophageal junction, gallbladder cancer, chordoma, and papillary renal cell carcinoma.

[0086] As used herein, the term “cancer status” or “cancer condition” refers to the characteristics of a cancer patient’s condition, such as diagnostic status, cancer type, cancer location, primary origin of cancer, cancer stage, cancer prognosis, and / or one or more additional characteristics of the cancer (e.g., tumor characteristics, such as morphology, heterogeneity, size, etc.). In some embodiments, one or more additional personal characteristics of the subject are used to further describe the subject’s cancer status or cancer condition, such as age, sex, weight, race, personal habits (e.g., smoking, alcohol consumption, diet), other relevant medical conditions (e.g., hypertension, dry skin, other diseases), current medications, allergies, relevant medical history, current cancer treatments, and side effects of other medications.

[0087] As used herein, the terms “treat,” “treating,” “treatment,” or “therapy” refer to both therapeutic treatment and preventative or protective measures aimed at preventing or alleviating a target pathological condition or symptom. Those requiring treatment include those diagnosed with the condition, those predisposed to it (e.g., a genetic predisposition), and those seeking to prevent it. The terms “prevent,” “preventing,” and “prevention” refer to reducing the likelihood of the onset (or recurrence) of a disease, symptom, condition, or related symptoms. This term implies obtaining a beneficial or desired outcome, such as a clinical outcome. Beneficial or desired outcomes may include, but are not limited to, the reduction of one or more symptoms. As used herein, the term “reduction,” for example, relating to symptoms of a condition, means reducing at least one of the frequency and magnitude of symptoms of a patient’s condition.

[0088] As used herein, “response” means the response of a subject with a pathology treatable with a biological, chemical, or physical therapy to said biological, chemical, or physical therapy. Normative criteria may vary depending on the disease.

[0089] As used herein, “immunotherapy” refers to any therapy that directly or indirectly alters a patient’s immune response or immune system. For immunotherapy strategies, a strong immune response detected at the tumor site has been found to be a reliable biomarker for many cancers, such as colon and rectal cancer, and a pre-existing immune response is assumed to be associated with better treatment outcomes. Immune response encompasses any form of immune response in the patient through direct or indirect action on the cancer or tumor site, or both. Immune response refers to the host cancer patient’s immune response to the tumor and encompasses the presence, quantity, or activity of cells and associated signaling molecules involved in the host immune response, including all cytokines, chemokines, growth factors, and stem cell growth factors. In some embodiments, immune response encompasses a variety of different cell subtypes, such as T cell lineages, B cell lineages, natural killer cells, macrophages, dendritic cells, myeloid-derived suppressor cells, lytic dendritic cells, fibroblasts, endothelial cells, and a large number of signaling molecules (cytokines, chemokines, and other signaling molecules).

[0090] As used herein, "immunotherapy agent" refers to a compound, composition, or treatment that indirectly or directly enhances, stimulates, or strengthens the body's immune response against cancer cells and / or reduces the side effects of other anticancer therapies. Therefore, immunotherapy is a therapy that directly or indirectly stimulates or strengthens the immune system's response to cancer cells and / or reduces side effects that may be caused by other anticancer agents. Immunotherapy is also referred to in the art as immunologic therapy, biological therapy, biological response modifier therapy, and biotherapy. Examples of common immunotherapy agents known in the art include, but are not limited to, cytokines, cancer vaccines, monoclonal antibodies, and non-cytokine adjuvants. Alternatively, immunotherapy may include administering a dose of immune cells (T cells, NK cells, dendritic cells, B cells, etc.) to a patient.

[0091] Immunotherapy agents can be non-specific, meaning they generally enhance the immune system, making it more effective in fighting the growth and / or spread of cancer cells, or they can be specific, meaning they target the cancer cells themselves. Immunotherapy regimens can combine non-specific and specific immunotherapy agents. Non-specific immunotherapy agents are substances that stimulate or indirectly enhance the immune system. Non-specific immunotherapy agents have been used alone as a primary therapy for cancer, as well as as a supplement to primary therapy, in which case they act as adjuvants to enhance the effectiveness of other therapies, such as cancer vaccines. Non-specific immunotherapy agents can also work in the latter case to reduce the side effects of other therapies, such as bone marrow suppression induced by certain chemotherapy agents. Non-specific immunotherapy agents can act on key immune system cells and elicit secondary responses, such as increased production of cytokines and immunoglobulins. Alternatively, the agent itself may contain cytokines. Non-specific immunotherapy agents are generally classified as cytokine or non-cytokine adjuvants.

[0092] Many cytokines have been found to have applications in cancer treatment, either as general non-specific immunotherapy to enhance the immune system or as adjuvants provided in conjunction with other therapies. Suitable cytokines include, but are not limited to, interferons, interleukins, and colony-stimulating factors.

[0093] The interferons (IFNs) considered in this invention include common types of IFNs, IFN-α (IFN-a), IFN-β (IFN-β), and IFN-γ (IFN-y). IFNs can act directly on cancer cells, for example, by slowing their growth, promoting their development into cells with more normal behavior, and / or increasing their antigen production, thereby making cancer cells more easily recognized and destroyed by the immune system. IFNs can also act indirectly on cancer cells, for example, by slowing angiogenesis, enhancing the immune system, and / or stimulating natural killer (NK) cells, T cells, and macrophages. Recombinant IFN-α is commercially available from Roferon (Roche Pharmaceuticals) and Intron A (Schering Corporation). The use of IFN-α alone or in combination with other immunotherapeutic agents or with chemotherapy agents has shown efficacy in treating a variety of cancers, including melanoma (including metastatic melanoma), kidney cancer (including metastatic kidney cancer), breast cancer, prostate cancer, and cervical cancer (including metastatic cervical cancer).

[0094] Interleukins considered in this invention include IL-2, IL-4, IL-11, and IL-12. Examples of commercially available recombinant interleukins include Proleukin® (IL-2; Chiron Corporation) and Neumega® (IL-12; Wyeth Pharmaceuticals). Zymogenetics, Inc. (Seattle, Washington) is currently testing a recombinant form of IL-21, which is also being considered for use in combinations of these inventions. Interleukins, alone or in combination with other immunotherapeutic agents or chemotherapeutic agents, have shown efficacy in treating a variety of cancers, including kidney cancer (including metastatic kidney cancer), melanoma (including metastatic melanoma), ovarian cancer (including recurrent ovarian cancer), cervical cancer (including metastatic cervical cancer), breast cancer, colorectal cancer, lung cancer, brain cancer, and prostate cancer. The combination of interleukin and IFN-α has also shown good activity in the treatment of various cancers (Negrier et al., Ann Oncol. 2002 13(9):1460-8; Touranietal, JClin Oncol. 2003 21(21):398794).

[0095] The colony-stimulating factors (CSFs) considered in this invention include granulocyte colony-stimulating factor (G-CSF or filgrastim), granulocyte-macrophage colony-stimulating factor (GM-CSF or sargramostim), and erythropoietin (epoetin alfa, darbepoietin). Treatment with one or more growth factors helps stimulate the production of new blood cells in patients receiving conventional chemotherapy. Therefore, CSF treatment can help reduce chemotherapy-related side effects and allow for the use of higher doses of chemotherapy agents. Various recombinant colony-stimulating factors are commercially available, such as Neupogen® (G-CSF; Amgen), Neulasta (pelfilgrastim; Amgen), Leukine (GM-CSF; Berlex), Procrit (erythropoietin; OrthoBiotech), Epogen (erythropoietin; Amgen), and Arnesp (erythropoietin). Colony-stimulating factors have shown efficacy in the treatment of cancers, including melanoma, colorectal cancer (including metastatic colorectal cancer), and lung cancer.

[0096] Non-cytokine adjuvants suitable for combinations applicable to this invention include, but are not limited to, levamisole, alum, BCG, Freund's incomplete adjuvant (IFA), QS-21, DETOX, keyhole hemocyanin (KLH), and dinitrophenyl (DNP). Combinations of non-cytokine adjuvants with other immunotherapeutic and / or chemotherapeutic agents have demonstrated efficacy against various cancers, including, for example, colon and colorectal cancer (levamisole); melanoma (BCG and QS-21); and kidney and bladder cancer (BCG).

[0097] In addition to having specific or non-specific targets, immunotherapeutic agents can be active, i.e., stimulating the body's own immune response, or they can be passive, i.e., contained in immune system components generated outside the body.

[0098] Passive specific immunotherapy typically involves the use of one or more monoclonal antibodies that are specific to specific antigens found on the surface of cancer cells or to specific cell growth factors. Monoclonal antibodies can be used to treat cancer in a variety of ways, such as enhancing a subject's immune response to a specific type of cancer, interfering with cancer cell growth by targeting specific cell growth factors (such as those involved in angiogenesis), or interfering with cancer cell growth by enhancing the delivery of other anticancer agents to cancer cells when linked or conjugated to agents such as chemotherapy agents, radioactive particles, or toxins.

[0099] Currently used as cancer immunotherapy agents, monoclonal antibodies suitable for inclusion in the combinations of this invention include, but are not limited to, rituximab (Rituxan®), trastuzumab (Herceptin®), ibritumomab tiuxetan (Zevalin®), tositumomab (Bexxar®), cetuximab (C-225, Erbitux®), bevacizumab (Avastin®), gemtuzumab ozogamicin (Mylotarg®), alemtuzumab (Campath®), and BL22. Monoclonal antibodies are used to treat a variety of cancers, including breast cancer (including advanced metastatic breast cancer), colorectal cancer (including advanced and / or metastatic colorectal cancer), ovarian cancer, lung cancer, prostate cancer, cervical cancer, melanoma, and brain tumors.

[0100] Other examples include antibodies against specific co-stimulatory molecules. Co-stimulatory molecules include, for example, B7-1 / CD80, CD28, B7-2 / CD86, CTLA-4, B7-H1 / PD-L1, Gi24 / Dies 1 / VISTA, B7-H2, ICOS, B7-H3, PD-1, B7-H4, PD-L2 / B7-DC, B7-H6, PDCD6, BTLA, 4-1BB / TNFRSF9 / CD137, CD40 ligand / TNFSF5, and 4-1BB ligand / TNFSF9. GITR / TNFRSF18, HVEM / TNFRSF14, CD27 / TNFRSF7, LIGHT / TNFSF14, CD27 ligand / TNFSF7, OX40 / TNFRSF4, CD30 / TNFRSF8, OX40 ligand / TNFSF4, CD30 ligand / TNFSF8, TACI / TNFRSF13B, CD40 / TNFRSF5, 2B4 / CD244 / SLAMF4, CD84 / SLAMF5, BLAME / SLAMF8, CD229 / SLAMF3, CD2CRACC / SLAMF7, CD2F-10 / SLAMF9, NTB-A / SLAMF6, CD48 / SLAMF2, SLAM / CD150, CD58 / LFA-3, CD2 Ikaros, CD53 integrin α4 / CD49d, CD82 / Kai-1 integrin α4β1, CD90 / Thyl integrin α4β7 / LPAM-1, CD96 LAG-3, CD160 LMIR1 / CD300A, CRTAM TCL1A, DAP 12 TCL1B, Dectin-1 / CLEC7A TIM-1 / KIM-1 / HAVCR, DPPIV / CD26 TIM-4, EphB6 TSLP, HLA class I TSLP R, HLA-DR. Specifically, antibodies are selected from the group consisting of: anti-CTLA4 antibodies (e.g., ipilimumab), anti-PD1 antibodies, anti-PDL1 antibodies, anti-TIMP3 antibodies, anti-LAG3 antibodies, anti-B7H3 antibodies, anti-B7H4 antibodies, anti-TREM antibodies, anti-BTLA antibodies, anti-light antibodies, or anti-B7H6 antibodies.

[0101] Monoclonal antibodies can be used alone or in combination with other immunotherapeutic agents or chemotherapy agents.

[0102] Active specific immunotherapy often involves the use of cancer vaccines. Cancer vaccines have been developed that contain intact cancer cells, parts of cancer cells, or one or more antigens derived from cancer cells. Cancer vaccines, alone or in combination with one or more immunotherapeutic or chemotherapeutic agents, are being investigated for the treatment of several types of cancer, including melanoma, kidney cancer, ovarian cancer, breast cancer, colorectal cancer, and lung cancer. Combining nonspecific immunotherapeutic agents with cancer vaccines is useful in order to enhance the body's immune response.

[0103] Immunotherapy can consist of adoptive immunotherapy, as described by Nicholas P. Restifo, Mark E. Dudley, and Steven A. Rosenberg, "Adoptive immunotherapy for cancer: harnessing the T cell response," Nature Reviews Immunology, Vol. 12, April 2012. In adoptive immunotherapy, the patient's circulating lymphocytes or tumor-infiltrating lymphocytes are isolated in vitro, activated by lymphokines such as IL-2, or exfiltrated by tumor necrosis genes, and then re-administered (Rosenberg et al., 1988; 1989). The activated lymphocytes are preferably the patient's own cells, which have been isolated earlier from blood or tumor samples and activated (or "expanded") in vitro. This form of immunotherapy has resulted in several cases of regression in melanoma and renal cell carcinoma.

[0104] As used herein, the term "genomic abundance value" refers to the absolute or relative amount of microbial genomes in a biological sample from the gut of a subject. Genomic abundance values ​​can be expressed in various units, including copy number, molar concentration, mass (e.g., mass normalized to genome size), unique sequence reads (e.g., unique sequence reads normalized to genome size), a percentage of any prior measure relative to the total amount of all genomes in the sample, a percentage of any prior measure relative to the total amount of multiple genomes in the sample, etc. In some embodiments, the genomic abundance value is normalized relative to the total genomic abundance in the sample. In some embodiments, the genomic abundance value is normalized relative to the genomic abundance value of a control genome in the sample. In some embodiments, the values ​​of multiple genomic abundance values ​​in the sample are standardized, normalized, and / or scaled. Examples of methods for normalizing genome abundance values ​​are described, for example, in Lin, H., Peddada, SD, Analysis of microbial compositions: a review of normalization and differential abundance analysis, Biofilms Microbiomes, 6(60) (2020) and Lutz KC et al., A Survey of Statistical Methods for Microbiome Data Analysis, Frontiers in Applied Mathematics and Statistics, 8 (2022), the contents of which are incorporated herein by reference in their entirety. Methods for measuring genome abundance values ​​are known in the art. For example, metagenomic sequencing can be used to reconstruct microbial genomes to a large extent by next-generation sequencing of genomic DNA in biological samples, such as biological samples from the gut of a subject. For a review of metagenomic sequences, see, for example, Quince C et al., Shotgun metagenomics, from sampling to analysis, Nat Biotechnol, 35(9):833-44 (2017), the contents of which are incorporated herein by reference in their entirety. Genomic abundance can also be determined by quantifying the copy number of ribosomal genes, such as the 16S rRNA gene.Examples of rRNA quantification are described in the following: Manzari C. et al., Accurate quantification of bacterial abundance in metagenomic DNAs accounting for variable DNA integrity levels, Microb Genom., 6(10):mgen000417 (2020) and Barlow, JT et al., A quantitative sequencing framework for absolute abundance measurements of mucosal and lumenal microbial communities, Nat Commun., 11:2590 (2020), the contents of which are incorporated herein by reference in their entirety.

[0105] As used herein, the term "relative abundance" refers to the ratio of a first amount of a compound (e.g., the genome of a first microorganism) measured in a sample to a second amount of a compound measured in a second sample. In some embodiments, relative abundance refers to the ratio of the amount of a compound (e.g., the genome of a first microorganism) in the same sample to the total amount of compounds (e.g., the total amount of a microbial genome or the total amount of multiple genomes). In other embodiments, relative abundance refers to the ratio of the amount of a compound (e.g., the genome of a first microorganism) in a first sample to the amount of that compound in a second sample. For example, the ratio of the normalized amount of the genome of a first microorganism in a first sample to the normalized amount of that genome of a first microorganism in a second sample and / or a reference sample.

[0106] As used herein, the terms “sequencing,” “sequencing determination,” etc., refer to any biochemical method that can be used to determine the order of biological macromolecules such as nucleic acids or proteins. For example, sequencing data may include all or part of the nucleotide bases in nucleic acid molecules such as mRNA transcripts or genomic loci.

[0107] As used herein, the term "sequence read" or "read" refers to a nucleotide sequence produced by any nucleic acid sequencing method described herein or known in the art. Reads can be generated from one end of a nucleic acid fragment ("single-end read") or from both ends of a nucleic acid fragment (e.g., paired-end read, double-ended read). The length of a sequence read is typically related to the specific sequencing technology. For example, high-throughput methods can provide sequence read sizes ranging from tens to hundreds of base pairs (bp). In some embodiments, the average, median, or average length of the sequence reads is approximately 15 bp to 900 bp (e.g., approximately 20 bp, approximately 25 bp, approximately 30 bp, approximately 35 bp, approximately 40 bp, approximately 45 bp, approximately 50 bp, approximately 55 bp, approximately 60 bp, approximately 65 bp, approximately 70 bp, approximately 75 bp, approximately 80 bp, approximately 85 bp, approximately 90 bp, approximately 95 bp, approximately 100 bp, approximately 110 bp, approximately 120 bp, approximately 130 bp, approximately 140 bp, approximately 150 bp, approximately 200 bp, approximately 250 bp, approximately 300 bp, approximately 350 bp, approximately 400 bp, approximately 450 bp, or approximately 500 bp). In some embodiments, the average, median, or average length of the sequence reads is approximately 1000 bp, 2000 bp, 5000 bp, 10,000 bp, or 50,000 bp or greater. For example, Nanopore® sequencing can provide sequence reads ranging in size from tens to hundreds to thousands of base pairs. Illumina® parallel sequencing, for example, can provide sequence reads with minimal variation; most reads are less than 200 bp. A sequence read (or sequencing read) can refer to the sequence information corresponding to a nucleic acid molecule (e.g., a string of nucleotides). For example, a sequence read can correspond to a string of nucleotides from a portion of a nucleic acid fragment (e.g., about 20 to about 150), a string of nucleotides at one or both ends of a nucleic acid fragment, or the nucleotides of the entire nucleic acid fragment. Sequence reads can be obtained in various ways, such as using sequencing technologies or using probes, such as in hybridization arrays or capture probes; or amplification techniques, such as polymerase chain reaction (PCR) or linear amplification using a single primer, or isothermal amplification.

[0108] As used herein, the term "read segment" refers to any form of nucleotide sequence read, including raw sequence reads obtained directly by nucleic acid sequencing technology, or sequences derived from raw sequence reads, such as aligned, folded, or spliced ​​sequence reads.

[0109] As used in this article, the term "read count" refers to the total number of nucleic acid reads generated during a nucleic acid sequencing reaction, which may or may not be equal to the number of nucleic acid molecules generated.

[0110] As used herein, the terms “read depth,” “sequencing depth,” or “depth” can refer to the total number of unique nucleic acid fragments covering a specific locus or region of a microbial genome sequenced in a particular sequencing reaction. Sequencing depth can be expressed as “Yx,” such as 50x, 100x, etc., where “Y” refers to the number of unique nucleic acid fragments covering a specific locus sequenced in a sequencing reaction. In this case, Y must be an integer, as it represents the actual sequencing depth of the specific locus. Alternatively, read depth, sequencing depth, or depth can refer to a measure of central tendency (e.g., mean or mode) of the number of unique nucleic acid fragments covering one of multiple loci or regions of a microbial genome sequenced in a particular sequencing reaction. For example, in some embodiments, sequencing depth refers to the average depth of each locus in a targeted sequencing panel, exome, or the entire genome of a microorganism. In this case, Y can be expressed as a fraction or decimal, as it refers to the average coverage of multiple loci. When describing depth averages, the actual depth of any particular site may differ from the overall described depth. A measure can be determined to provide a defined percentage of the total number of loci falling within a range of sequencing depths. For example, sequencing depths ranging from 90% to 95% or 99% of gene loci. As those skilled in the art will understand, different sequencing technologies offer different sequencing depths. For example, low-pass whole-genome sequencing can refer to technologies that provide sequencing depths of less than 5x, less than 4x, less than 3x, or less than 2x (e.g., about 0.5x to about 3x).

[0111] As used herein, the term "sequencing breadth" refers to which portion of a particular microbial genome has been sequenced. Sequencing width can be expressed as a fraction, decimal, or percentage, and is typically calculated as (number of loci analyzed / total number of loci in the genome). The denominator of this fraction can be a duplicate-masked genome, and therefore 100% can correspond to all reference genomes minus the masked portions. A duplicate-masked genome can refer to a genome in which sequence repetitions are masked (e.g., sequence reads aligned with unmasked portions of the genome). In some embodiments, any portion of the genome can be masked, and therefore the sequencing width of any desired portion of the genome can be assessed.

[0112] As used herein, the terms “sequence ratio” and “coverage ratio” are interchangeable and refer to any measure of the number of units of genomic sequence in a first or more biological samples (e.g., test samples and / or tumor samples) compared to the number of units of the corresponding genomic sequence in a second or more biological samples (e.g., reference samples and / or control samples). In some embodiments, the sequence ratio is the copy ratio, the log2-converted copy ratio (e.g., log2 copy ratio), the coverage ratio, the base fraction, the allele fraction (e.g., the variant allele fraction), and / or the tumor ploidy. In some embodiments, the sequence ratio is the log2-converted copy ratio (e.g., log2 copy ratio), the coverage ratio, the base fraction, the allele fraction (e.g., the variant allele fraction), and / or the tumor ploidy. N The converted copy ratio, where N is any real number greater than 1.

[0113] As used herein, the term “sequencing probe” refers to a molecule that binds to nucleic acids with an affinity based on the expected nucleotide sequence of RNA or DNA present at the locus.

[0114] As used herein, the term “target group” or “targeted genome” refers to a combination of probes used to sequence (e.g., by next-generation sequencing) nucleic acids present in a biological sample (e.g., a tumor sample, a liquid biopsy sample, a germ tissue sample, a leukocyte sample, or a tumor or tissue organoid sample) from a subject, which are selected to locate one or more loci of interest in the genome.

[0115] As used herein, the term “sensitivity” or “true positive rate” (TPR) refers to the number of true positives divided by the sum of the number of true positives and false negatives. Sensitivity characterizes the ability of a assay or method to correctly identify the proportion of a population that truly possesses a certain condition. For example, sensitivity characterizes the ability of a method to correctly identify the number of subjects within a population who have a specific biological characteristic.

[0116] As used herein, the term "specificity" or "true negative rate" (TNR) refers to the number of true negatives divided by the sum of the number of true negatives and false positives. Specificity characterizes the ability of a assay or method to correctly identify the proportion of a population that does not actually have a certain condition. For example, specificity can characterize the ability of a method to correctly identify the number of subjects within a population who do not possess a specific biological characteristic.

[0117] As used interchangeably in this article, the terms "classifier" or "model" refer to machine learning models or algorithms.

[0118] In some embodiments, the model includes an unsupervised learning algorithm. An example of an unsupervised learning algorithm is cluster analysis. In some embodiments, the model includes supervised machine learning. Non-limiting examples of supervised learning algorithms include, but are not limited to, logistic regression, neural networks, support vector machines, Naive Bayes, nearest neighbor algorithms, random forest algorithms, decision tree algorithms, boosting tree algorithms, multinomial logistic regression algorithms, linear models, linear regression, gradient boosting, mixture models, hidden Markov models, Gaussian NB algorithms, linear discriminant analysis, diffusion models, or any combination thereof. In some embodiments, the model is a multinomial classifier algorithm. In some embodiments, the model is a 2-stage stochastic gradient descent (SGD) model. In some embodiments, the model is a deep neural network (e.g., a depth-width sample-level model).

[0119] Neural network. In some embodiments, the model is a neural network (e.g., a convolutional neural network and / or a residual neural network). Neural network algorithms, also known as artificial neural networks (ANNs), include convolutional and / or residual neural network algorithms (deep learning algorithms). In some embodiments, a neural network is a machine learning algorithm trained to map an input dataset to an output dataset, wherein the neural network includes interconnected groups of nodes organized into multiple layers of nodes. For example, in some embodiments, the neural network architecture includes at least an input layer, one or more hidden layers, and an output layer. In some embodiments, the neural network includes any total number of layers and any number of hidden layers, wherein the hidden layers serve as trainable feature extractors that allow mapping a set of input data to one or a set of output values. In some embodiments, the deep learning algorithm is a neural network including multiple hidden layers (e.g., two or more hidden layers). In some cases, each layer of the neural network includes multiple nodes (or "neurons"). In some embodiments, a node receives input directly from the input data or from the output of a node in a previous layer and performs a specific operation, such as a summation operation. In some embodiments, the connection from the input to the node is associated with parameters (e.g., weights and / or weighting factors). In some embodiments, the node is associated with all input pairs x. iThe summation of the product of the parameter and its associated parameters. In some embodiments, the weighted summation is offset by the bias b. In some embodiments, a threshold or activation function f may be used to gate the output of a node or neuron, and in some cases, f is a linear or nonlinear function. In some embodiments, the activation function is, for example, a modified linear unit (ReLU) activation function, a leaky ReLU activation function, or other functions such as a saturated hyperbolic tangent function, an identity function, a binary step function, a logic function, an arcTan function, a softsign function, a parametric modified linear unit function, an exponential linear unit function, a softPlus function, a curved identity function, a softExponential function, a Sinusoid function, a Sine function, a Gaussian function, or a sigmoid function, or any combination thereof.

[0120] In some implementations, one or more sets of training data are used during the training phase to "teach" or "learn" the weighting factors, bias values, thresholds, or other computational parameters of the neural network. For example, in some implementations, input data from the training dataset and gradient descent or backpropagation are used to train the parameters such that the output values ​​computed by the ANN are consistent with the instances included in the training dataset. In some embodiments, the parameters are obtained from the backpropagation neural network training process.

[0121] Any of a variety of neural networks is suitable for use according to this disclosure. Examples include, but are not limited to, feedforward neural networks, radial basis function networks, recurrent neural networks, residual neural networks, convolutional neural networks, residual convolutional neural networks, and any combination thereof. In some embodiments, machine learning utilizes ANN or deep learning architectures that employ pre-trained and / or transfer learning. In some implementations, convolutional neural networks and / or residual neural networks are used according to this disclosure.

[0122] For example, a deep neural network model includes an input layer, multiple individually parameterized (e.g., weighted) convolutional layers, and an output scorer. The parameters (e.g., weights) of each convolutional layer and the input layer contribute to multiple parameters (e.g., weights) associated with the deep neural network model. In some embodiments, at least 50 parameters, at least 100 parameters, at least 1000 parameters, at least 2000 parameters, or at least 5000 parameters are associated with the deep neural network model. Therefore, deep neural network models require the use of a computer because they cannot be solved by mental effort. In other words, given a model input, the model output in such embodiments needs to be determined using a computer, not by mental effort. See, for example, Krizhevsky et al., 2012, “Imagenet classification with deep convolutional neural networks”, in Advances in Neural Information Processing Systems 2, edited by Pereira, Burges, Bottou, and Weinberger, pp. 1097–1105, Curran Associates, Inc.; Zeiler, 2012, “ADADELTA: an adaptive learning rate method”, CoRR, Vol. abs / 1212.5701; and Rumelhart et al., 1988, “Neurocomputing: Foundations of research”, Chapter Learning Representations by Back-propagating Errors, pp. 696–699, Cambridge, MA, USA: MIT Press, each of which is hereby incorporated by reference.

[0123] Neural network algorithms (including convolutional neural network algorithms) suitable for use as models are disclosed, for example, Vincent et al., 2010, “Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion,” J Mach Learn Res 11, pp. 3371–3408; Larochelle et al., 2009, “Exploring strategies for training deep neural networks,” J Mach Learn Res 10, pp. 1–40; and Hassoun, 1995, Fundamentals of Artificial Neural Networks, Massachusetts Institute of Technology, each of which is hereby incorporated by reference. Additional example neural networks suitable for use as models are disclosed in Duda et al., 2001, Pattern Classification, 2nd ed., John Wiley & Sons, Inc., New York; and Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, each of which is hereby incorporated by reference in its entirety. Other example neural networks suitable for use as models are also described in the following: Draghici, 2003, Data Analysis Tools for DNA Microarrays, Chapman & Hall / CRC; and Mount, 2001, Bioinformatics: sequence and genome analysis, ColdSpring Harbor Laboratory Press, Cold Spring Harbor, New York, each of which is hereby incorporated in full by reference.

[0124] Support Vector Machine. In some embodiments, the model is a Support Vector Machine (SVM). Suitable SVM algorithms for use as models are described, for example, in Cristianini and Shaw-Taylor, 2000, “An Introduction to Support Vector Machines,” Cambridge University Press, Cambridge; Boser et al., 1992, “A training algorithm for optimal margin classifiers,” in Proceedings of the 5th Annual ACM Workshop on Computational Learning Theory, ACM Press, Pittsburgh, Pa., pp. 142–152; Vapnik, 1998, Statistical Learning Theory, Wiley, New York; Mount, 2001, Bioinformatics: sequence and genome analysis, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY; Duda, Pattern Classification, 2nd Edition, 2001, John Wiley & Sons, Inc., pp. 259, 262–265; and Hastie, 2001, The Elements of Statistical Learning, Springer, New York. York; and Furey et al., 2000, Bioinformatics 16, 906-914, each of which is hereby incorporated by reference in full. When used for classification, SVM separates a given dataset of binary labels from a hyperplane that is maximally farthest from the labeled data. For certain cases where linear separation is impossible, SVM operates in conjunction with a 'kernel' technique that automatically implements a nonlinear mapping to the feature space. The hyperplane discovered by SVM in the feature space corresponds in some cases to a nonlinear decision boundary in the input space. In some embodiments, multiple parameters (e.g., weights) associated with SVM define the hyperplane. In some embodiments, the hyperplane is defined by at least 10, at least 20, at least 50, or at least 100 parameters, and the SVM model requires computation by a computer because it cannot be solved mentally.

[0125] Naive Bayes algorithm. In some embodiments, the model is a Naive Bayes algorithm. Suitable Naive Bayes models for use as models are disclosed, for example, in: Ng et al., 2002, “On discriminative vs. generative classifiers: A comparison of logistic regression and naive Bayes,” Advances in Neural Information Processing Systems, 14, which is hereby incorporated by reference. Naive Bayes models are any models in the “probabilistic models” family that are based on the application of Bayes’ theorem and the assumption of strong (naive) independence between features. In some embodiments, they are combined with kernel density estimation. See, for example, Hastie et al., 2001, Theelements of statistical learning: data mining, inference, and prediction, eds. Tibshirani and Friedman, Springer, New York, which is hereby incorporated by reference.

[0126] The nearest neighbor algorithm. In some embodiments, the model is a nearest neighbor algorithm. In some implementations, the nearest neighbor model is memory-based and does not include the model to be fitted. For the nearest neighbor, given a query point x0 (test subject), k training points x (r) Let r, ..., identify the k nearest neighbors to x0 (here, the training subjects), and then classify point x0 using the k nearest neighbors. In some embodiments, the distance is determined using Euclidean distance in the feature space. Typically, when using the nearest neighbor algorithm, the abundance data used to compute the linear discriminant is standardized to have a mean of 0 and a variance of 1. In some embodiments, the nearest neighbor rule is refined to address issues of unequal class priors, differential misclassification costs, and feature selection. Many of these refinements involve some form of weighted voting on neighbors. For more information on nearest neighbor analysis, see Duda, Pattern Classification, 2nd Edition, 2001, John Wiley & Sons, Inc; and Hastie, 2001, The Elements of Statistical Learning, Springer, New York, each of which is hereby incorporated by reference.

[0127] The k-nearest neighbor model is a nonparametric machine learning method where the input consists of the k nearest training samples in the feature space. The output is a class membership. An object is classified by the majority vote of its neighbors, where the object is assigned to the class most common among its k nearest neighbors (k is a positive integer, typically small). If k=1, the object is assigned to the class of that single nearest neighbor. See, Duda et al., 2001, Pattern Classification, 2nd Edition, John Wiley & Sons, which are hereby incorporated by reference. In some embodiments, the amount of distance computation required to solve the k-nearest neighbor model makes it necessary to use a computer to solve the model for a given input, as the model cannot be mentally generated.

[0128] Random forest, decision tree, and boosting tree algorithms. In some embodiments, the model is a decision tree. Decision trees suitable for use as models are generally described in: Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, pp. 395-396, which are hereby incorporated by reference. Tree-based methods divide the feature space into a set of rectangles and then fit a model (e.g., a constant) into each rectangle. In some embodiments, the decision tree is a random forest regression. For example, a particular algorithm is Classification and Regression Tree (CART). Other specific decision tree algorithms include, but are not limited to, ID3, C4.5, MART, and random forest. CART, ID3, and C4.5 are described in the following reference: Duda, 2001, Pattern Classification, John Wiley & Sons, Inc., New York, pp. 396-408 and 411-412, which are hereby incorporated by reference. CART, MART, and C4.5 are described in: Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York, Chapter 9, which is hereby incorporated by reference in its entirety. Random forests are described in: Breiman, 1999, “Random Forests--RandomFeatures,” Technical Report 567, Statistics Department, UC Berkeley, September 1999, which is hereby incorporated by reference in its entirety. In some embodiments, decision tree models include at least 10, at least 20, at least 50, or at least 100 parameters (e.g., weights and / or decisions) and require computation by a computer because they cannot be solved mentally.

[0129] Regression. In some embodiments, the model uses a regression algorithm. In some embodiments, the regression algorithm is any type of regression. For example, in some embodiments, the regression algorithm is logistic regression. In some embodiments, the regression algorithm is logistic regression with lasso regularization, L2 regularization, or elastic network regularization. In some embodiments, those extracted features with corresponding regression coefficients that fail to meet a threshold are pruned from consideration. In some embodiments, a generalized logistic regression model that handles multi-class responses is used as the model. The logistic regression algorithm is disclosed in Agresti, An Introduction to Categorical Data Analysis, 1996, Chapter 5, pp. 103-144, John Wiley & Son, New York, which is incorporated herein by reference. In some embodiments, the model employs a regression model disclosed in Hastie et al., 2001, The Elements of Statistical Learning, Springer-Verlag, New York. In some embodiments, the logistic regression model includes at least 10, at least 20, at least 50, at least 100, or at least 1000 parameters (e.g., weights) and requires computation by a computer because it cannot be solved mentally.

[0130] Linear discriminant analysis algorithm. In some embodiments, linear discriminant analysis (LDA), normal discriminant analysis (NDA), or discriminant function analysis is a generalized Fisher linear discriminant method, which is a method used in statistics, pattern recognition, and machine learning to find linear combinations of features that characterize or distinguish two or more classes of objects or events. In some embodiments, the resulting combinations are used as models (linear models) in some embodiments of this disclosure.

[0131] Hybrid models and Hidden Markov models. In some embodiments, the model is a hybrid model, such as that described in: McLachlan et al., Bioinformatics 18(3):413-422, 2002. In some embodiments, particularly those including a time component, the model is a Hidden Markov model, such as that described in: Schliep et al., 2003, Bioinformatics 19(1):i255-i263.

[0132] Clustering. In some embodiments, the model is an unsupervised clustering model. In some embodiments, the model is a supervised clustering model. Clustering algorithms suitable for use as models are described, for example, in the following literature: Duda and Hart, Pattern Classification and Scene Analysis, pp. 211-256, 1973, John Wiley & Sons, Inc., New York, (hereinafter referred to as "Duda 1973"), which is incorporated herein by reference in its entirety. As an illustrative example, in some embodiments, the clustering problem is described as finding one of the natural groupings in a dataset. To identify natural groupings, two problems are addressed. First, a way to measure the similarity (or dissimilarity) between two samples is determined. This measure (e.g., a similarity measure) is used to ensure that samples in one cluster are more similar to each other than they are similar to samples in other clusters. Second, a mechanism for partitioning the data into clusters using the similarity measure is determined. One approach to begin clustering research is to define a distance function and compute a distance matrix between all pairs of samples in the training set. If distance is a good measure of similarity, then the distance between reference entities in the same cluster is significantly smaller than the distance between reference entities in different clusters. However, in some implementations, clustering does not use distance metrics. For example, in some embodiments, a non-metric similarity function s(x,x') can be used to compare two vectors x and x'. In some such embodiments, s(x,x') is a symmetric function, with a larger value when x and x' are "similar" in some way. Once a method for measuring the "similarity" or "dissimilarity" between points in the dataset is selected, clustering uses a standard function that measures the clustering quality of any partition of the data. Data partitions that extremes the standard function are used to cluster the data. Specific exemplary clustering techniques contemplated for use in this disclosure include, but are not limited to, hierarchical clustering (aggregate clustering using nearest neighbor, farthest neighbor, average connectivity, centroid, or sum of squares algorithms), k-means clustering, fuzzy k-means clustering, and Jarvis-Patrick clustering. In some embodiments, clustering includes unsupervised clustering (e.g., without a pre-defined number of clusters and / or without a pre-determined cluster assignment).

[0133] Model ensemble and boosting. In some embodiments, model ensemble (two or more) is used. In some embodiments, boosting techniques such as AdaBoost are combined with many other types of learning algorithms to improve model performance. In this method, the outputs of any of the models disclosed herein, or their equivalents, are combined into a weighted sum representing the final output of the boosted model. In some embodiments, multiple outputs from the models are combined using any measure of central tendency known in the art, including but not limited to mean, median, mode, weighted mean, weighted median, weighted mode, etc. In some embodiments, a voting method is used to combine multiple outputs. In some embodiments, the corresponding models in the model ensemble are weighted or unweighted.

[0134] As used herein, the term "parameter" refers to any coefficient or similar value of an internal or external element (e.g., weights and / or hyperparameters) in an algorithm, model, regressor, and / or classifier that can affect (e.g., modify, customize, and / or adjust) one or more inputs, outputs, and / or functions in the algorithm, model, regressor, and / or classifier. For example, in some embodiments, a parameter refers to any coefficient, weight, and / or hyperparameter that can be used to control, modify, customize, and / or adjust the behavior, learning, and / or performance of an algorithm, model, regressor, and / or classifier. In some cases, parameters are used to increase or decrease the influence of inputs (e.g., features) on an algorithm, model, regressor, and / or classifier. As a non-limiting example, in some embodiments, parameters are used to increase or decrease the influence of nodes (e.g., nodes in a neural network), where nodes include one or more activation functions. The assignment of parameters to a particular input, output, and / or function is not limited to any particular paradigm of a given algorithm, model, regressor, and / or classifier, but can be used for any suitable algorithm, model, regressor, and / or classifier architecture to achieve the desired performance. In some embodiments, parameters have fixed values. In some embodiments, the values ​​of the parameters may be adjusted manually and / or automatically. In some embodiments, the values ​​of the parameters are modified through validation and / or training processes of the algorithm, model, regressor, and / or classifier (e.g., through error minimization and / or backpropagation methods). In some embodiments, the algorithms, models, regressors, and / or classifiers of this disclosure include multiple parameters. In some embodiments, the plurality of parameters is n parameters, wherein: n≥2; n≥5; n≥10; n≥25; n≥40; n≥50; n≥75; n≥100; n≥125; n≥150; n≥200; n≥225; n≥250; n≥350; n≥500; n≥600; n≥750; n≥1,000; n≥2,000; n≥4,000; n≥5,000; n≥7,500; n≥10,000; n≥20,000; n≥40,000; n≥75,000; n≥100,000; n≥200,000; n≥500,000; n≥1×10 6 n≥5×10 6 or n≥1×10 7 Therefore, the algorithms, models, regressors, and / or classifiers of this disclosure cannot be performed mentally. In some embodiments, n is 10,000 with 1 × 10⁻⁶. 7 Between 100,000 and 5×10 6 Between 500,000 and 1×10 6In some embodiments, the algorithms, models, regressors, and / or classifiers of this disclosure operate in a k-dimensional space, where k is a positive integer of 5 or greater (e.g., 5, 6, 7, 8, 9, 10, etc.). Therefore, the algorithms, models, regressors, and / or classifiers of this disclosure cannot be executed mentally.

[0135] As used herein, the term "untrained model" (e.g., "untrained classifier" and / or "untrained neural network") refers to a machine learning model or algorithm (such as a classifier or neural network) that has not yet been trained on a target dataset. In some embodiments, "training a model" (e.g., "training a neural network") refers to the process of training an untrained or partially trained model (e.g., "untrained or partially trained neural network"). Furthermore, it should be understood that the term "untrained model" does not preclude the possibility of using transfer learning techniques in such training of an untrained or partially trained model. For example, a non-limiting example of such transfer learning is provided below: Fernandes et al., 2017, "Transfer Learning with Partial Observability Applied to Cervical Cancer Screening," Pattern Recognition and Image Analysis: Iberian Conference Proceedings, 8th Edition, pp. 243-250, which is hereby incorporated by reference. In the case of using transfer learning, supplementary data is provided to the untrained model described above, in addition to the data from the primary training dataset. Typically, this supplementary data is in the form of parameters (e.g., coefficients, weights, and / or hyperparameters) learned from another auxiliary training dataset. Furthermore, although a description of a single auxiliary training dataset has been disclosed, it should be understood that there is no limitation on the number of auxiliary training datasets that can be used to supplement the primary training dataset when training an untrained model in this disclosure. For example, in some embodiments, two or more, three or more, four or more, or five or more auxiliary training datasets are used to supplement the primary training dataset by transfer learning, wherein each such auxiliary training dataset is different from the primary training dataset. In some such embodiments, transfer learning of any kind is used. For example, consider the case where a first auxiliary training dataset and a second auxiliary training dataset exist in addition to the primary training dataset. In this scenario, transfer learning techniques (e.g., a second model, the same as or different from the first model) are used to apply parameters learned from the first auxiliary training dataset (by applying the first model to the first auxiliary training dataset) to the second auxiliary training dataset, thereby generating a trained intermediate model. The parameters of this intermediate model are then applied to the main training dataset and combined with the main training dataset itself to train the untrained model. Alternatively, in another exemplary embodiment, a first set of parameters learned from the first auxiliary training dataset (by applying the first model to the first auxiliary training dataset) and a second set of parameters learned from the second auxiliary training dataset (by applying a second model, the same as or different from the first model, to the second auxiliary training dataset) are each applied separately to individual instances of the main training dataset (e.g., through separate independent matrix multiplications). These two applications of parameters to individual instances of the main training dataset are then combined with the main training dataset itself (or some simplified form of the main training dataset, such as principal components or regression coefficients learned from the main training set) and applied to the untrained model to train the untrained model.

[0136] As used herein, the term "AUC" refers to, for example, the area under the ROC curve. This value assesses the merit of a test on a given sample population, with a value of 1 indicating a good test down to 0.5, meaning the test provides a random response when classifying the test subject. Since the AUC ranges only from 0.5 to 1.0, small changes in AUC are more significant than similar changes in measures ranging from 0% to 1% or 0% to 100%. When giving a percentage change in AUC, it is calculated based on the fact that the measure ranges from 0.5 to 1.0 across the entire range. Various statistical packages can calculate the AUC of an ROC curve. AUC can be used to compare the accuracy of classification algorithms across the entire data range. By definition, a classification algorithm with a larger AUC has a greater ability to correctly classify unknowns between two groups of interest (disease and no disease, responders and non-responders).

[0137] As used herein, the term "instruction" refers to a command given to a computer processor by a computer program. On a digital computer, each instruction is a sequence of 0s and 1s that describes the physical operation the computer is to perform. Such instructions can include data transfer instructions and data manipulation instructions. In some embodiments, each instruction is a type of instruction in an instruction set that is identified by the specific processor type used to execute the instruction. Examples of instruction sets include, but are not limited to, Reduced Instruction Set Computers (RISC), Complex Instruction Set Computers (CISC), Minimal Instruction Set Computers (MISC), Very Long Instruction Words (VLIW), Explicit Parallel Instruction Computation (EPIC), and Single Instruction Set Computers (OISC).

[0138] The following examples illustrate several aspects. It should be understood that many specific details, relationships, and methods are presented to provide a complete understanding of the features described herein. However, those skilled in the art will readily recognize that the features described herein can be practiced without one or more of these specific details or using other methods. The features described herein are not limited to the illustrated order of actions or events, as some actions may occur in a different order and / or simultaneously with other actions or events. Furthermore, implementing the methods according to the features described herein does not require all the actions or events shown.

[0139] Select abbreviation***

[0140] HQMAG – High-Quality Metagenomic Assembly of Genomes

[0141] CCDC-I – Case-Control Dataset Collection I

[0142] CCDC-II – Case-Control Dataset Collection II

[0143] TDC – Collection of Treatment Datasets

[0144] TCG – Two competing bacterial groups

[0145] QD-TCG – Two competing bacterial groups (141 HQMAGs) identified from QD studies

[0146] Two competing bacterial groups (788 HQMAGs) in the C-TCG-combination

[0147] CC-TCG – A core set of two competing bacterial communities (284 HQMAGs)

[0148] ANI – Average Nucleotide Identity

[0149] Detailed reference will now be made to embodiments, examples of which are illustrated in the accompanying drawings. Numerous specific details are set forth in the following detailed description to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure can be practiced without these specific details. In other instances, well-known methods, processes, components, circuits, and networks have not been described in detail so as not to unnecessarily obscure aspects of the embodiments.

[0150] Exemplary system implementation.

[0151] Since an overview of some aspects of this disclosure and some definitions used in this disclosure have already been provided, in conjunction with Figure 1 The details of the exemplary system are described. Figure 1This is a block diagram illustrating a system 100 according to some embodiments. In some embodiments, system 100 includes one or more processing units, a CPU 102 (also referred to as a processor), one or more network interfaces 104, a user interface 106 including (optionally) a display 108 and an input system 110, non-persistent memory 111, persistent memory 112, and one or more communication buses 114 for interconnecting these components. The one or more communication buses 114 may optionally include circuitry (sometimes referred to as a chipset) that interconnects and controls communication between system components. Non-persistent memory 111 typically includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, ROM, EEPROM, flash memory, while persistent memory 112 typically includes CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic tape cassette, magnetic tape, disk storage or other magnetic storage devices, disk storage devices, optical disc storage devices, flash memory devices or other non-volatile solid-state storage devices. Persistent memory 112 may optionally include one or more storage devices located remotely from CPU 102. The non-volatile storage devices within persistent memory 112 and non-persistent memory 112 include non-transitory computer-readable storage media. In some implementations, non-persistent memory 111 or alternatively, non-transitory computer-readable storage media may be combined with persistent memory 112 to store programs, modules, and data structures or subsets thereof:

[0152] • An optional operating system 116, which includes procedures for handling various basic system services and for performing hardware-related tasks;

[0153] • Optional network communication module (or instruction) 118 for connecting system 100 to other devices and / or communication network 104;

[0154] • Microbiome assessment module 140, used to determine one of multiple disease states in a subject based on the composition of the subject's microbiome; and

[0155] • Data storage 140 of subject information based on microbiome sequencing results 150, including abundance values ​​152 of microorganisms in each of microbial communities 152-A and 152-B as described herein.

[0156] In various implementations, one or more of the aforementioned elements are stored in one or more of the previously mentioned memory devices and correspond to an instruction set for performing the functions described above. The modules, data, or programs (e.g., instruction sets) identified above do not need to be implemented as separate software programs, processes, datasets, or modules, and therefore various subsets of these modules and data can be combined or otherwise rearranged in various implementations. In some implementations, non-persistent memory 111 optionally stores a subset of the modules and data structures identified above. Furthermore, in some embodiments, the memory stores additional modules and data structures not described above. In some embodiments, one or more of the identified elements are stored in a computer system other than the visualization system 100, which is addressable by the visualization system 100, such that the visualization system 100 can retrieve all or part of this data when needed.

[0157] although Figure 1 A "System 100" is depicted, but this diagram is intended more as a functional description of the various features that may exist in a computer system than as a structural diagram of the implementation described herein. In practice, and as those skilled in the art will recognize, items shown individually can be combined and some items can be separated. Furthermore, although Figure 1 Some data and modules in non-persistent memory 111 are described, but some or all of these data and modules may be stored in persistent memory 112.

[0158] 1. Methods for training models to predict subjects' responses to treatments for a disease.

[0159] Figure 2 is a schematic diagram of the method for training a model to predict a subject's response to a treatment for a condition, as described below. This method can utilize a computer system (e.g., as described above). Figure 1 This is achieved through a computer system 100 that is displayed and described in real time.

[0160] Referring to block 200, in some embodiments, the method includes, for each of a plurality of training subjects, electronically acquiring: (i) corresponding plurality of genomic abundance values ​​for the corresponding training subject at a time prior to receiving therapy, wherein the corresponding plurality of genomic abundance values ​​include, for each of a plurality of gut microbiota, a corresponding value for the abundance of the genome of the corresponding gut microbe in a corresponding biological sample from the gut of the corresponding training subject; and (ii) an indication of the corresponding training subject's response to the therapy. Each of the plurality of training subjects has received therapy for the condition.

[0161] In some embodiments, the plurality of training subjects includes at least 50, at least 100, at least 200, at least 500, at least 1000, at least 5000, at least 10,000, at least 20,000, at least 50,000, at least 100,000, at least 500,000, or at least 1,000,000 subjects. In some embodiments, the plurality of training subjects includes no more than 1,000,000, no more than 500,000, no more than 100,000, no more than 50,000, no more than 20,000, no more than 10,000, no more than 1,000 subjects, no more than 500 subjects, no more than 100 subjects, or no more than 50 subjects. In some embodiments, the number of training subjects may consist of 50 to 100, 50 to 200, 50 to 500, 100 to 500, 200 to 500, 200 to 1000, 500 to 1000, 200 to 5,000, 1000 to 10,000, 5,000 to 200,000, 10,000 to 50,000, 20,000 to 100,000, or 500,000 to 1,000,000. In some embodiments, the number of training subjects may fall into another range, starting with no fewer than 50 subjects and ending with no more than 100,000,000 subjects. In some embodiments, the number of subjects may share similar health conditions (such as physical or mental condition, medical history, gene carrier status, or drug use).

[0162] In some embodiments, a corresponding biological sample is collected from the gut of the respective training subject prior to treatment or therapy. In some embodiments, the biological sample is collected no more than 15 minutes, 30 minutes, 1 hour, 2 hours, 3 hours, 4 hours, 5 hours, 6 hours, 12 hours, or 24 hours prior to treatment or therapy. In some embodiments, the biological sample is collected 1 day, 2 days, 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 3 weeks, 4 weeks, or longer prior to treatment or therapy. In some embodiments, the biological sample is collected approximately 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, or longer prior to treatment or therapy.

[0163] In some embodiments, sample data (including plasma and stool samples) and corresponding clinical information (including sex / age / body fat count / underlying diseases / histopathological features, etc.) are collected for each training subject prior to treatment. A whole-microbiome analysis is performed on the individual biological samples. In some embodiments, the samples are tissue biopsy samples, intestinal samples, or mucosal samples. See, for example, Tang Q, Jet et al., Current Sampling Methods for GutMicrobiota: A Call for More Precise Devices, Front Cell Infect Microbiol., 10:151 (2020), the contents of which are incorporated herein by reference in their entirety. In some embodiments, the biological sample from the gut of the respective subject is a stool sample from the respective training subject.

[0164] In some embodiments, the corresponding value for genome abundance is a value representing the absolute abundance of the microbial genome. In some embodiments, the corresponding value for genome abundance is a value representing a normalized abundance value or a relative abundance value (e.g., the abundance of a microorganism normalized relative to the abundance of the total microbiome of interest). In some embodiments, the corresponding value for genome abundance is a value representing an average abundance value (e.g., the average abundance obtained at different time points or from different biological samples from a patient, or the average abundance obtained using different probes, etc.), or a combination of the above. The corresponding value for genome abundance is measured using any technique known in the art. In some embodiments, the genome abundance value is measured by quantitative PCR (qPCR) such as bacterial 16S rRNA qPCR, RT-PCR, or qRT-PCR to quantify the abundance of regions of interest in the genome, for example, as described in U.S. Patent No. 11,427,865, the disclosure of which is incorporated herein by reference in its entirety. In some embodiments, genome abundance values ​​are determined by using targeted sequencing (e.g., 16S rRNA sequencing or any other suitable biomarker), partial genome sequencing, or whole genome sequencing, thereby quantifying the number of reads of the target region in the microbial genome to determine genome abundance, for example as disclosed in U.S. Patent Application Publication No. 2021 / 0403986 or U.S. Patent No. 11,332,783, the disclosures of which are incorporated herein by reference in their entirety. In some embodiments, deep sequencing is used to determine the abundance of the target sequence, for example as disclosed in U.S. Patent Application Publication No. 2018 / 0237863, the disclosure of which is incorporated herein by reference in its entirety.In some embodiments, the sequencing depth is at least 2X, at least 3X, at least 4X, at least 5X, at least 6X, at least 7X, at least 8X, at least 9X, at least 10X, at least 11X, at least 12X, at least 13X, at least 14X, at least 15X, at least 16X, at least 17X, at least 18X, at least 19X, at least 20X, at least 21X, at least 22X, at least 23X, at least 24X, at least 25X, at least 26X, at least 27X, at least 28X, at least 29X, at least 30X, at least 31X, at least 32X, at least 33X, at least 34X, at least 35X, at least 36X, at least 37X, at least 38X, or at least 39X, at least 40X, at least 41X, at least 42X, at least 43X, at least 44X, at least 45X, at least 46X, at least 47X, at least 48X, at least 49X, at least 50X, at least 51X, at least 52X, at least 53X, at least 54X, at least 55X, at least 56X, at least 57X, at least 58X, at least 59X, at least 60X, at least 70X, at least 80X, at least 90X, at least 100X, at least 110X, at least 120X, at least 130X, at least 150X, at least 200X, at least 300X, at least 400X, at least 500X, at least 750X, at least 1000X or more. In some embodiments, shotgun metagenomic sequencing is used to provide sequence reads of the genome in the sample, for example, as described in U.S. Patent No. 11,028,449, the contents of which are incorporated herein by reference in their entirety.

[0165] In some embodiments, subject response is characterized by clinical outcome measures, including but not limited to complete response, partial response, no response, survival, development of adverse events, or any combination thereof. In some embodiments, a responder has a complete response to treatment, while a non-responder has no response or partial response. In some embodiments, training subjects undergo routine clinical examination, laboratory analysis, and computed tomography. Tumor response is assessed using RECIST criteria. In some embodiments, a complete response is defined as the complete radiographic disappearance of the minimum radiographically detectable disease or stable disease; a partial response is defined as a measurable reduction of at least 50% in the longest dimension of the disease; disease stability is defined as a reduction of less than 25% in the longest dimension; and progressive disease is defined as tumor growth of more than 25% in the longest dimension or the development of new lesions. In some embodiments, the overall response rate is defined as the sum of the complete and partial response rates, and the tumor control rate is defined as the sum of the overall response rate and the disease stability rate.

[0166] In some embodiments, the indication of subject response is characterized by the actual therapeutic efficacy of the therapy, including progression-free survival (PFS), duration of progression-free survival during treatment, overall survival (OS), response to the therapy (RT), overall response rate (ORR), sustained clinical effect (DCB), disease activity score, or any combination thereof, or any other method known in the art for assessing the progression or prognosis of a disease or condition.

[0167] In some embodiments, “progression-free survival” (PFS) has the meaning as understood in the art, referring to the length of time a patient has had a disease (such as cancer) during and after treatment for that disease without it becoming progressive. In some embodiments, measuring progression-free survival is used as an assessment of the effectiveness of a new treatment. In some embodiments, PFS is determined in a randomized clinical trial; in some such embodiments, PFS refers to the time from randomization until objective tumor progression and / or death.

[0168] In some embodiments, ORR may be defined as the proportion of patients who identify a partial (PR) or complete (CR) response as the best overall response (BOR) based on some metrics, such as the Response Evaluation Criteria in Solid Tumors (RECIST 1.1). Stable disease (SD) is categorized as non-responsive along with progressive disease (PD). In some embodiments, ORR has its meaning as understood in the art as the proportion of patients whose tumor size decreases by a predetermined amount and for a minimum duration. In some embodiments, the duration of response is typically measured from the initial response time up to the recorded tumor progression. In some embodiments, ORR involves the sum of a partial response and a complete response.

[0169] In some embodiments, "clinical effect" refers to clinical benefit. In some embodiments, such clinical benefit is or includes a reduction in tumor size, an increase in progression-free survival, an increase in overall survival, a reduction in total tumor burden, and a reduction in symptoms caused by tumor growth, such as pain, organ failure, bleeding, skeletal damage, and other related sequelae of metastatic cancer, and combinations thereof. In some embodiments, clinical effect is a "continuous clinical effect" (DCB) maintained over a relevant time period. In some embodiments, the relevant time period is at least 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, 7 months, 8 months, 9 months, 10 months, 11 months, 1 year, 2 years, 3 years, 4 years, 5 years, or longer.

[0170] In some embodiments, subject responses are measured using a Disease Activity Score (DAS) (see, for example, Vander Heijde DM et al., J Rheumatol, 1993, 20(3): 579-81; Prevoo ML et al., Arthritis Rheum, 1995, 38: 44-8). The DAS system represents the current state and changes in disease activity. The DAS scoring system uses a weighted mathematical formula derived from clinical trials of RA. For example, DAS 28 is 0.56(T28) + 0.28(SW28) + 0.70(Ln ESR) + 0.014 GH, where T represents the number of tender joints, SW represents the number of swollen joints, ESR represents the erythrocyte sedimentation rate, and GH represents overall health status. Various DAS values ​​represent high or low disease activity and remission, and changes and endpoint scores lead to patient classification by degree of response (none, moderate, good).

[0171] In some embodiments, the level of an immune response or immune parameter in a cancer patient induced by immunotherapy is used to measure an indicator of the subject's response. In some embodiments, the immune response or immune parameter is characterized by the expression levels of various biomarkers of the host immune response in conjunction with the occurrence of cancer at a given stage of cancer development (i.e., treatment efficacy). In some embodiments, the expression levels of the biomarker are compared to, and, when necessary, to, a reference value. Therefore, the reference value for the same biomarker is predetermined and is known to indicate a reference value associated with distinguishing between low and high levels of immune response to the biomarker in cancer patients. The predetermined reference value for the biomarker is associated with cancer patients who respond to treatment, or conversely, with cancer patients who do not respond to treatment.

[0172] In some embodiments, the variation in combinations of biomarkers is quantified. In some embodiments, combinations of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more different biomarkers are quantified.

[0173] In some embodiments, immunohistochemistry is used to quantify biomarkers. Example biomarkers include 18S, ACE, ACTB, AGTR1, AGTR2, APC, APOA1, ARF1, AXIN1, BAX, BCL2, BCL2L1, CXCR5, BMP2, BRCA1, BTLA, C3, CASP3, CASp9, CCL1, CCL11, CCL13, CCL16, CCL17, CCL18, CCL19, CCL2, CCL20, CCL21, CCL22, CCL23, CCL24, CCL25, CCL26, CCL27, CCL28, CCL3, CCL5, CCL7, CCL8, CCNB1, CCND1, and C. CNE1, CCR1, CCR10, CCR2, CCR3, CCR4, CCR5, CCR6, CCR7, CCR8, CCR9, CCRL2, CD154, CD19, CD1a, CD2, CD226, CD244, PDCD1LG1, CD28, CD34, CD36, CD3 8. CD3E, CD3G, CD3Z, CD4, CD40LG, CD5, CD54, CD6, CD68, CD69, CLIP, CD80, CD83, SLAMF5, CD86, CD8A, CDH1, CDH7, CDK2, CDK4, CDKN1A, CDKN1B, CDKN 2A, CDKN2B, CEACAM1, COL4A5, CREBBP, CRLF2, CSF1, CSF2, CSF3, CTLA4, CTNN81, CTSC, CX3CL1, CX3CRI, CXCL1, CXCL10, CXCL11, CXCL12, CXCL13, CX CL14, CXCL16, CXCL2, CXCL3, CXCL5, CXCL6, CXCL9, CXCR3, CXCR4, CXCR6, CYP1A2, CYP7A1, DCC, DCN, DEFA6, DICER1, DKK1, Dok-1, Dok-2, DOK6, DVL1 , E2F4, EBI3, ECE1, ECGF1, EDN1, EGF, EGFR, EIF4E, CD105, ENPEP, ERBB2, EREG, FCGR3A, CGR3B, FN1, FOXP3, FYN, FZD1, GAPD, GLI2, GNLY, GOLPH4, GR B2, GSK3B, GSTP1, GUSB, GZMA, GZMH, GZMK, HLA-B, HLA-C, HLA-, MA, HLA-DMB, HLA-DOA, HLA-DOB, HLA-DPA1, HLA-DQA2, HLA-DRA, HLX1, HMOX1, HRAS,HSPB3, HUWE1, ICAM1, ICAM-2, ICOS, ID1, ifna1, ifna17, ifna2, ifna5, ifna6, ifna8, IFNAR1, IFNAR2, IFNG, IFNGR1, IFNG R2, IGF1, IHH, IKBKB, IL10, IL12A, IL12B, IL12RB1, IL12RB2, IL13, IL13RA2, IL15, IL15RA, IL17, IL17R, IL17RB, IL18, IL1 A, IL1B, IL1RI, IL2, IL21, IL21R, IL23A, IL23R, IL24, IL27, IL2RA, IL2RB, IL2RG, IL3, IL31RA, IL4, IL4RA, IL5, IL6, IL7, IL7RA, IL8, CXCR1, CXCR2, IL9, IL9R, IRF1, ISGF3G, ITGA4, ITGA7, integrin, αE (antigen CD103, human mucosal lymphocytes, antigen 1; α polypeptide), hCG33203 gene, ITGB3 , JAK2, JAK3, KLRB1, KLRC4, KLRF1, KLRG1, KRAS, LAG3, LAIR2, LEF1, LGALS9, LILRB3, LRP2, LTA, SLAMF3, MADCAM1, MADH3, M ADH7, MAF, MAP2K1, MDM2, MICA, MICB, MKI67, MMP12, MMP9, MTA1, MTSS1, MYC, MYD88, MYH6, NCAM1, NFATC1, NKG7, NLK, NOS2A, P2X7, PDCD1, PECAM-, CXCL4, PGK1, PIAS1, PIAS2, PIAS3, PIAS4, PLAT, PML, PP1A, CXCL7, PPP2CA, PRF1, PROM1, PSMB5, ​​PTCH, PTGS2, PTP4A3, PTPN6, PTPRC, RAB23, RAC / RHO, RAC2, RAF, RB1, RBL1, REN, Drosha, SELE, SELL, SELP, SERPINE1, SFRP1, SIRP β1, SKI, SLAMF1, SLAMF6, SLAMF7, SLAMF8, SMAD2, SMAD4, SMO, SMOH, SMURF1, SOCS1, SOCS2, SOCS3, SOCS4, SOCS5 , SOCS6, SOCS7, SOD1, SOD2, SOD3, SOS1, SOX17, CD43, ST14, STAM, STAT1, STAT2, STAT3, STAT4, STAT5A, STAT5B,STAT6, STK36, TAP1, TAP2, TBX21, TCF7, TERT, TFRC, TGFA, TGFB1, TGFBR1, TGFBR2, TIMP3, TLR1, TLRO1, TLR2, TLR3, TLR4, TLR5, TLR6, TLR7, TLR8, TLR9, TNF, TNFRSF10A, TNFRSF11A, TNFRSF18, TNFRSF1A, TNFRSF1B, OX-40, TNFRSF5, TNFRSF6, TNFRSF7, TNFRSF8, TNFRSF9, TNFSF10, TNFSF6, TOB1, TP53, TSLP, VCAM1, VEGF, WIF1, WNT1, WNT4, XCL1, XCR1, ZAP70, and ZIC2. ,

[0174] In some embodiments, the treatment regimen or therapy may be administered via any common route, as long as the target tissue or cells are available through that route. This includes, but is not limited to, intravenous, catheter insertion, in situ, intradermal, subcutaneous, intramuscular, intraperitoneal intratumoral, oral, nasal, buccal, rectal, vaginal, or local application. The choice of therapeutic agent and dosage regimen may depend on a variety of factors, such as the combination of drugs used, the specific disease being treated, and the patient's condition and medical history.

[0175] Referring to block 202, in some embodiments, the method includes sequencing the genomic DNA of a corresponding biological sample from the gut of each of the plurality of training subjects to obtain a plurality of corresponding nucleic acid sequences. In some embodiments, the plurality of corresponding nucleic acid sequences comprises at least 10,000, at least 100,000, at least 250,000, at least 500,000, at least 1,000,000, at least 2,500,000, at least 5,000,000, at least 10,000,000, or at least 50,000,000 nucleic acid sequences. In some embodiments, the number of nucleic acid sequences corresponding to multiple nucleic acid sequences may be no more than 250,000,000, no more than 100,000,000, no more than 50,000,000, no more than 25,000,000, no more than 10,000,000, no more than 5,000,000, no more than 1,000,000, or no more than 100,000 nucleic acid sequences. In some embodiments, the corresponding plurality of nucleic acid sequences comprises 100,000 to 1,000,000, 200,000 to 5,000,000, 500,000 to 10,000,000, 1,000,000 to 20,000,000, 5,000,000 to 50,000,000, 10,000,000 to 100,000,000, or 50,000,000 to 250,000,000 nucleic acid sequences. In some embodiments, the corresponding plurality of nucleic acid sequences falls into another range starting at no less than 1,000 nucleic acid sequences and ending at no more than 250,000,000 nucleic acid sequences.

[0176] In some embodiments, multiple nucleic acid sequences are obtained via metagenomic sequencing, for example, as disclosed in U.S. Patent Application Publication No. 2016 / 0239602 or U.S. Patent No. 11,495,326, the contents of which are incorporated herein by reference in their entirety. In some embodiments, metagenomic sequencing further includes generating multiple metagenomic reads. In some embodiments, metagenomic sequencing further includes fragmenting the microbial genome into random fragments of a target size. The size of the resulting fragments may vary. In one embodiment, a fragment of approximately 500 nucleotides may be obtained. In some embodiments, fragments of 100-2000 nucleotides (e.g., 200-800, 100-900, 100-1000, 300-800, 400-900 nucleotides) may be obtained. In some embodiments, the method may further include extracting metagenomic fragments from the corresponding biological sample. In some embodiments, metagenomic sequencing further includes sequencing the fragments using high-throughput sequencing methods to generate multiple sequencing reads.

[0177] In some embodiments, targeted genome sequencing is used to obtain multiple corresponding nucleic acid sequences. Examples of targeted genome sequencing are described in U.S. Patent Application Publication No. 2019 / 0316209. In some embodiments, targeted genome sequencing includes hybridizing genomic DNA isolated from a biological sample from the gut of a subject with a set of probes prior to sequencing the recovered nucleic acids. This set of probes includes one or more probes that hybridize to unique sequences in the genome of each quantified microorganism. In some embodiments, the microorganisms include those listed in Tables 1 and 2 and / or... Figures 13A to 13XX The list includes a variety of microorganisms. In some embodiments, combinations of semi-unique sequences (e.g., sequences found in a small number of microbial genomes) can be used to deconvolve genome abundance values ​​using algorithms (e.g., systems of equations). In some embodiments, the probe set includes at least one probe that hybridizes to a unique sequence in each microbial genome to be detected. In some embodiments, the probe set includes at least two, at least three, at least four, at least five, at least ten, at least twenty-five, at least fifty, or more probes that hybridize to unique and distinct sequences in each microbial genome to be detected. In some embodiments, the probe set includes at least 20, at least 30, at least 40, at least 50, at least 75, at least 100, at least 125, at least 150, at least 200, at least 150, at least 300, at least 400, at least 500, at least 750, at least 1000, at least 1250, at least 1500, at least 2000, at least 2500, at least 3000, at least 4000, at least 5000, at least 7500, at least 10,000 or more unique probes.

[0178] In some embodiments, the sequenced genomic DNA from the corresponding biological sample contains at its ends partial or complete sequencing platform adaptor sequences that can be used for sequencing using the sequencing platform of interest. Sequencing platforms of interest include, but are not limited to, HiSeq™, MiSeq™, and Genome Analyzer™ sequencing systems from Illumina®; Ion PGM™ and Ion Proton™ sequencing systems from IonTorrent™; the PACBIO RSII Sequel system from Pacific Biosciences; the SOLiD sequencing system from Life Technologies™; the 454 GS FLX+ and GS Junior sequencing systems from Roche; the MinION™ system from Oxford Nanopore; or any other sequencing platform of interest.

[0179] Referring to block 204, in some embodiments, the method includes, for each of a plurality of training subjects, electronically acquiring a plurality of corresponding nucleic acid sequences for genomic DNA from a corresponding biological sample of the gut of the corresponding training subject.

[0180] Referring to block 206, in some embodiments, the method includes determining, for each of a plurality of gut microbes, a corresponding value for the abundance of the genome of the corresponding gut microbe from a plurality of corresponding nucleic acid sequences.

[0181] In some embodiments, the genomic abundance values ​​determined for each corresponding subject among a plurality of training subjects include at least 20, at least 25, at least 50, at least 75, at least 100, at least 125, at least 150, at least 175, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 700, at least 800, at least 900, at least 1000, at least 1500, at least 2000, at least 25000, at least 5000, or at least 10000 genomic abundance values, wherein each genomic abundance value corresponds to a different gut microbiota. In some embodiments, the genome abundance values ​​determined for each corresponding subject among a plurality of training subjects include genome abundance values ​​of no more than 250,000, no more than 100,000, no more than 50,000, no more than 25,000, no more than 10,000, no more than 5,000, no more than 2,500, no more than 1,000, no more than 750, no more than 500, or fewer. In some embodiments, the genome abundance values ​​determined for each corresponding subject among a plurality of training subjects consist of the following: 10 to 40, 20 to 50, 30 to 80, 40 to 100, 50 to 150, 60 to 200, 80 to 300, 90 to 500, 100 to 1,000, 500 to 2,000, or 1,000 to 5,000 genome abundance values. In some embodiments, the genomic abundance values ​​determined for each corresponding subject among a plurality of training subjects fall into another range that begins with no less than 10 genomic abundance values ​​and ends with no more than 250,000 genomic abundance values.

[0182] Referring to block 208, in some embodiments, for each corresponding training subject among a plurality of training subjects, the method includes assembling a plurality of corresponding gut microbial genomes electronically from corresponding plurality of nucleic acid sequences via metagenomic de novo sequencing, and for each corresponding gut microbial among the plurality of gut microbes, calculating a corresponding value for the abundance of the genome of the corresponding gut microbe based on the universality of the corresponding nucleic acid sequence in the plurality of nucleic acid sequences used to assemble the corresponding gut microbial genome corresponding to the corresponding gut microbe in the plurality of gut microbial genomes. In some embodiments, metagenomic de novo sequencing further includes generating contigs based on sequencing reads generated by shotgun sequencing technology. This technique is described, for example, in U.S. Patent No. 10,529,443, the contents of which are incorporated herein by reference in their entirety. In some embodiments, a first plurality of nucleic acid sequences are assembled into a complete genome of a plurality of gut microbes. In some embodiments, a plurality of nucleic acid sequences are assembled into a partial genome of a plurality of gut microbes.

[0183] Referring to block 210, in some embodiments, for each corresponding subject among a plurality of training subjects, the method includes assigning each corresponding nucleic acid sequence among a plurality of nucleic acid sequences to a corresponding gut microbe among a plurality of gut microbes, thereby generating a corresponding count of the corresponding nucleic acid sequence among the plurality of nucleic acid sequences assigned to the corresponding gut microbe for each corresponding gut microbe among the plurality of gut microbes, and determining a corresponding genomic abundance value for the corresponding gut microbe based on the corresponding count of the corresponding nucleic acid sequence assigned to the corresponding gut microbe for each corresponding gut microbe among the plurality of gut microbes. In some embodiments, assigning each corresponding nucleic acid to the corresponding gut microbe includes mapping the nucleic acid to a reference nucleic acid (e.g., the contigs listed in Figure 12). In some embodiments, assigning each corresponding nucleic acid to the corresponding gut microbe includes annotating genomic information based on existing databases. In some embodiments, nucleic acid sequences are analyzed, and the annotation is defined by using sequence similarity and phylogenetic placement methods or a combination of both strategies to define the classification assignment.

[0184] Sequence similarity-based methods for assigning each nucleic acid sequence to a corresponding gut microbe include those familiar to those skilled in the art, including but not limited to BLAST, BLASTx, tBLASTn, tBLASTx, RDP classifiers, DNAclust, and various implementations of these algorithms (such as Qiime or Mothur). These methods rely on mapping sequence reads to a reference database and selecting matches with the best score and e-value. In some embodiments, phylogenetic methods are combined with sequence similarity methods to improve the accuracy of annotation or classification assignments. Common databases include, but are not limited to, GT-DBTK, NCBI Genbank, the European Bioinformatics Institute-European Nucleotide Archive (EBI-ENA), the US Department of Energy National Institute of Genetics (USDOE) Integrated Microbial Genomes & Microbiomes (IMG / M), and other databases available in the art.

[0185] In some embodiments, a microarray comprising probe sequences capable of detecting unique genomic sequences of each corresponding genome of a variety of gut microbes is used to determine multiple genomic abundance values. In some embodiments, the probe set on the microarray comprises at least one probe that hybridizes to a unique sequence of each microbial genome to be detected. In some embodiments, the probe set comprises at least two, at least three, at least four, at least five, at least ten, at least 25, at least 50 or more probes that hybridize to unique and distinct sequences of each microbial genome to be detected. In some embodiments, the probe set includes at least 20, at least 30, at least 40, at least 50, at least 75, at least 100, at least 125, at least 150, at least 200, at least 150, at least 300, at least 400, at least 500, at least 750, at least 1000, at least 1250, at least 1500, at least 2000, at least 2500, at least 3000, at least 4000, at least 5000, at least 7500, at least 10,000 or more unique probes.

[0186] Referring to box 212, in some embodiments, the variety of gut microbiota includes those selected from Table 1, Table 2, or... Figures 13A to 13XXAt least 20 types of gut microbiota. In some embodiments, at least about 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100 or more types of gut microbiota are selected from Table 1, Table 2 or Figures 13A to 13XX In some embodiments, the variety of gut microbiota includes those selected from Table 1, Table 2, or... Figures 13A to 13XX At least 25 gut microbiota. In some embodiments, the gut microbiota comprises those selected from Table 1, Table 2, or... Figures 13A to 13XX At least 30 gut microbiota. In some embodiments, the gut microbiota comprises those selected from Table 1, Table 2, or... Figures 13A to 13XX At least 40 gut microbiota. In some embodiments, the gut microbiota comprises those selected from Table 1, Table 2, or... Figures 13A to 13XX At least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 125, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, or all of the gut microbiota. In some embodiments, the multiple gut microbiota are selected from Table 1, Table 2, or... Figures 13A to 13XX At least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 125, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, or all gut microbiota. In some embodiments, multiple gut microbiota refers to all gut microbiota listed in Table 1. In some embodiments, multiple gut microbiota refers to all gut microbiota listed in Table 2. In some embodiments, multiple gut microbiota is... Figures 13A to 13XX All gut microbes listed in the table.

[0187] Table 1 - Taxonomic allocation of 141 non-redundant genomes identified in two competing bacterial groups (QD-TCG)

[0188]

[0189] Table 2 - Classification and allocation of 284 core microbiomes (CC-TCG)

[0190]

[0191] Table 1, Table 2 and Figures 13A to 13XX The bacterial species listed were identified by metagenomic sequencing of genomic DNA isolated from human fecal samples and by identifying them as part of two competing microbial communities relative to at least one biological characteristic, as described in the examples. In short, genomic DNA isolated from each fecal sample was sequenced using next-generation sequencing, and contiguous groups of microbial genome sequences were constructed de novo. Generally, the contiguous groups identified for each microorganism are expected to represent more than 95% of the entire genome of that microorganism. Genome constructs with less than 1% sequence difference from each other were grouped and defined as originating from the same microorganism. Tables 1, 2, and... Figures 13A to 13XX The genomic contiguous groups of each microorganism listed are provided in the sequence listing submitted with this application. The taxonomic allocation for each microorganism is shown in Table 1, Table 2, or... Figures 13A to 13XXThe sequence identifiers assigned to each contiguous group and their corresponding microorganisms are given in Figure 12. For example, the contiguous groups provided as SEQ ID NO:1 to 68 correspond to the genome sequence of microorganism 1U001.8 (e.g., ...). Figure 12A As indicated in the table), microorganism 1U001.8 is a microorganism classified as follows: domain Bacteria, phylum Proteobacteria, class Gamma Proteobacteria, order Enterobacteriaceae, family Enterobacteriaceae, genus Escherichiae, and species Escherichia coli, and belongs to group 2 of the 141 core microorganisms identified in Table 1.

[0192] Therefore, in some embodiments of the methods described herein, if the identified genome construct has at least 97% sequence identity when compared with the microbial contigs provided in the sequence listing, the genomes identified in the metagenomic analysis will be classified as corresponding to Tables 1, 2, and / or Figures 13A to 13XX The microorganisms listed are shown in Figure 12. In some embodiments, if the identified genome construct has at least 98% sequence identity when compared with the microbial contiguous groups provided in the sequence listing, the genomes identified in the metagenomic analysis will be classified as corresponding to Tables 1, 2, and / or Figures 13A to 13XX The microorganisms listed are shown in Figure 12. In some embodiments, if the identified genome construct has at least 99% sequence identity when compared with the microbial contiguous groups provided in the sequence listing, the genomes identified in the metagenomic analysis will be classified as corresponding to Tables 1, 2 and / or Figures 13A to 13XX The microorganisms listed are shown in Figure 12. In some embodiments, if the identified genome construct has at least 99.5% sequence identity when compared with the microbial contigs provided in the sequence listing, the genomes identified in the metagenomic analysis will be classified as corresponding to Tables 1, 2, and / or Figures 13A to 13XX The microorganisms listed are shown in Figure 12. In some embodiments, if the identified genome construct has at least 97%, at least 97.5%, at least 98%, at least 98.5%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or higher sequence identity when compared with the contiguous group of microorganisms provided in the sequence listing, then the genomes identified in the metagenomic analysis will be classified as corresponding to Tables 1, 2 and / or Figures 13A to 13XX The microorganisms listed are shown in Figure 12.

[0193] Referring to box 214, in some embodiments, the multiple gut microbiota includes at least 20 microorganisms selected from Table 1, Table 2, or... Figures 13A to 13XXThe group of identified gut microbes is selected from those microbes having a connectivity of at least 2. In some embodiments, the group of identified gut microbes is selected from those microbes having a connectivity of at least 2, at least 3, at least 4, at least 5, at least 10, at least 15, at least 20, at least 25 or greater.

[0194] Referring to box 216, in some embodiments, the biological sample from the intestine of the corresponding subject is a fecal sample from the corresponding training subject. In some embodiments, the biological sample is a sample obtained from the small intestine or large intestine, preferably the colon or rectum, more preferably in the form of a fecal sample or rectal swab or a biopsy specimen of the gastrointestinal mucosa.

[0195] Referring to block 218, in some embodiments, the therapy is a biological therapy, immunotherapy, chemotherapy, radiotherapy, gene therapy, hormone therapy, photodynamic therapy, targeted therapy, small molecule, antibody, polynucleotide, natural compound, immunomodulator, bone marrow therapy, stem cell therapy, surgical therapy, induction therapy, maintenance therapy, or a combination thereof.

[0196] Referring to box 220, in some embodiments, the condition is selected from the group consisting of: type 2 diabetes, hypertension, schizophrenia, atherosclerotic cardiovascular disease (ACVD), cirrhosis (LC), inflammatory bowel disease (IBD), colorectal cancer (CRC), ankylosing spondylitis (AS) and Parkinson's disease (PD), inflammatory bowel disease (IBD), rheumatoid arthritis (RA), or advanced melanoma and B-cell lymphoma. In some embodiments, the condition is, for example, hypertension (HT), schizophrenia (SCZ), multiple sclerosis (MS), type II Gaucher disease (GDII), COVID-19 (COV), Behcet's disease (BD), autism spectrum disorder (ASD), or pancreatic cancer (PC). In some embodiments, the condition is cancer, Alzheimer's disease, cardiovascular disease, autoimmune disease, mental health disease, infectious disease, or genetic disease.

[0197] In some embodiments, a condition is categorized by any indicator of a patient's biological state, function, structure, process, response, or condition. Such indicators include any of a number of variables (parameters) commonly measured in medicine to assess a patient for purposes such as diagnosis, prognosis, and / or treatment. Generally, indicators of interest herein are those whose values ​​(which may be quantitative or qualitative) reflect, characterize, or relate to the function or structure of organs and organ systems, and / or whose values ​​reflect, characterize, or relate to the presence or severity of a condition. In some embodiments, a disease or symptom is categorized by its progression or prognosis, such as different stages of cancer, the type, frequency, or severity of a condition that can be objectively measured or experienced by the subject. In some embodiments, the condition may be obtained through a medical device that can be used to analyze the state of a body part, wound, or lesion, or the state of a physical matrix, such as a test strip, test strip, filter, or other matrix, which provides a means for detecting or measuring the presence of substances in a biological sample obtained from the subject. In some embodiments, it is intended to detect antibodies or biomarkers against pathogens (e.g., viruses, bacteria, fungi), abnormal tissues (e.g., tumor sites) in biological samples, and / or detect their presence in biological samples from patients, for purposes such as diagnosing the presence of symptoms or diseases.

[0198] Referring to box 222, in some embodiments, the condition is cancer.

[0199] Referring to block 224, in some embodiments, the method includes, for each of a plurality of training subjects, inputting information about the respective training subject into a model including multiple parameters. The model, for example, applies the multiple parameters to the information through at least 10,000 calculations to obtain a corresponding output from the model for that respective training subject. The corresponding output includes a prediction of the respective training subject's response to the therapy, and the information about the respective training subject includes a corresponding genomic abundance value for each of a plurality of gut microbiota, and the plurality of gut microbiota is selected from Table 1, Table 2, or... Figures 13A to 13XX .

[0200] In some embodiments, the model is trained on a dataset collected from multiple therapies for a condition, and the model is trained to distinguish between response and non-response states. In some embodiments, the model includes a learned statistical classifier system. In some embodiments, the learned statistical classifier system is a random forest, a classification and regression tree, a boosting tree, or a neural network. For example, as described in Example 3, the random forest classifier was trained on a dataset from 11 different studies that collectively investigated the microbiome in four different conditions. As shown in Figure 8C, the resulting model was able to predict whether a patient would respond to or not respond to anti-cytokine or anti-integrin therapy, methotrexate treatment for new-onset rheumatoid arthritis, immune checkpoint inhibitor (ICI) therapy for advanced melanoma, and CD19-CAR-T immunotherapy for B-cell lymphoma.

[0201] Referring to block 226, in some embodiments, the prediction of a corresponding training subject's response is a category output of the corresponding response among a plurality of possible responses of the corresponding training subject. This method allows setting a single "cutoff" value, allowing differentiation between responders and non-responders to treatment. In some embodiments, the prediction of a corresponding training subject's response includes a prediction of the objective response rate of the human subject to the treatment or therapy, and wherein the prediction of the objective response rate includes an indication or classification of the amount of a complete or partial response to the treatment.

[0202] Referring to block 228, in some embodiments, the prediction of the response of the corresponding training subject is a probability output for the response of the corresponding training subject. As disclosed above, the method allows setting a single “cutoff” value, allowing differentiation between responders and non-responders to treatment. In some embodiments, the method includes using the model to calculate a probability value for the subject; comparing the probability value to a threshold derived from a cohort of responders / non-responders to determine whether the probability value is above / below the threshold; and classifying the subject as a responder / non-responder if the probability value is above / below the threshold. In embodiments, the threshold may be a probability value of at least 50%, 55%, 50%, 65%, 70%, 75%, or about 80% or more. In other embodiments, the probability value is a positive predictive value as measured by the area under the curve (AUC) of a receiver operating characteristic (ROC) curve. In some embodiments, a multivariate logistic regression model, a neural network model, a random forest model, or a decision tree model is used to calculate the probability value.

[0203] Referring to box 230, in some embodiments, the model is a neural network algorithm, support vector machine algorithm, naive Bayes algorithm, nearest neighbor algorithm, boosting tree algorithm, random forest algorithm, convolutional neural network algorithm, decision tree algorithm, regression algorithm, or clustering algorithm.

[0204] Referring to block 232, in some embodiments, the plurality of parameters are at least 1,000, at least 10,000, at least 15,000, at least 50,000, at least 100,000, at least 250,000, at least 500,000, or at least 1,000,000, at least 2,500,000, at least 5,000,000, at least 10,000,000, or more parameters.

[0205] Referring to box 234, in some embodiments, the model applies multiple parameters to the information through at least 1,000 calculations, at least 5,000 calculations, at least 10,000 calculations, at least 25,000 calculations, at least 50,000 calculations, at least 100,000 calculations, at least 250,000 calculations, at least 500,000 calculations, at least 1,000,000 calculations, at least 2,500,000 calculations, at least 5,000,000 calculations, at least 10,000,000 calculations or more, in order to obtain the corresponding output of the corresponding training subject from the model.

[0206] Referring to block 236, in some embodiments, the method includes adjusting a plurality of parameters based on, for each of a plurality of training subjects, one or more differences between (i) a corresponding output from the model and (ii) a corresponding indication of the training subject’s response to the therapy.

[0207] In some embodiments of deep learning techniques that utilize neural networks as described above, training the neural network to improve the accuracy of its predictions involves modifying one or more parameters, including but not limited to weights in filters within convolutional layers and biases within network layers. In some embodiments, the weights and biases are further constrained by various forms of regularization (such as L1, L2, weight decay, and dropout).

[0208] For example, in some embodiments, the neural network or any model disclosed herein optionally has its parameters (e.g., weights) adjusted (to potentially minimize the error between the system's predicted indication and the measurement indication of the training data) after the training data is labeled (e.g., labeled with indicators of the state of biological characteristics). Various methods for minimizing the error function, such as gradient descent, include, but are not limited to, logarithmic loss, sum of squared errors, and hinged loss methods. In some embodiments, these methods further include second-order methods or approximations, such as momentum methods, Hessian-free estimation, Nesterov's accelerated gradient, adagrad, etc. In some embodiments, the method also combines label-free generative pre-training and labeled discriminative training.

[0209] Therefore, in some embodiments, training the neural network involves adjusting one or more parameters among a plurality of parameters through backpropagation of the loss function. In some embodiments, the loss function is a regression task and / or a classification task. Non-limiting examples of loss functions suitable for regression tasks include, but are not limited to, the mean squared error loss function, the mean absolute error loss function, the Huber loss function, the Log-Cosh loss function, or the quantile loss function. See, Wang et al., 2020, “A Comprehensive Survey of Loss Functions in Machine Learning,” Annals of DataScience, doi.org / 10.1007 / s40745-020-00253-5, last accessed on September 15, 2021, which is incorporated herein by reference in its entirety. Non-limiting examples of loss functions suitable for classification tasks include, but are not limited to, the binary cross-entropy loss function, the hinge loss function, or the squared hinge loss function. In some embodiments, the loss function is any suitable regression task loss function or classification task loss function.

[0210] This paper further describes other suitable methods for training neural networks that are intended to be used in this disclosure (see, for example, the definition above: untrained model).

[0211] In some embodiments, the parameters of the neural network are randomly initialized before training.

[0212] In some embodiments, the neural network includes discarding regularization parameters. For example, in some embodiments, regularization is performed by adding a penalty to the loss function, where the penalty is proportional to the parameter values ​​in the trained or untrained model. Generally, regularization reduces model complexity by adding a penalty to one or more parameters to reduce the importance of the corresponding hidden neurons associated with those parameters. Such practices can produce more general models and reduce overfitting of data. In some embodiments, regularization includes L1 or L2 penalties.

[0213] In some embodiments, training a neural network includes an optimizer. In some embodiments, the optimizer may use a loss function to update the parameters of the neural network or other model via backpropagation. In some embodiments, training a neural network includes a learning rate.

[0214] In some embodiments, the learning rate is at least 0.0001, at least 0.0005, at least 0.001, at least 0.005, at least 0.01, at least 0.05, at least 0.1, at least 0.2, at least 0.3, at least 0.4, at least 0.5, at least 0.6, at least 0.7, at least 0.8, at least 0.9, or at least 1. In some embodiments, the learning rate is no more than 1, no more than 0.9, no more than 0.8, no more than 0.7, no more than 0.6, no more than 0.5, no more than 0.4, no more than 0.3, no more than 0.2, no more than 0.1, no more than 0.05, no more than 0.01, or less. In some embodiments, the learning rate is from 0.0001 to 0.01, 0.001 to 0.5, 0.001 to 0.01, 0.005 to 0.8, or 0.005 to 1. In some embodiments, the learning rate falls into another range that begins at a value not less than 0.0001 and ends at a value not greater than 1.

[0215] In some embodiments, the learning rate further includes learning rate decay (e.g., a reduction in the learning rate over one or more epochs). For example, the learning rate decay rate may be a reduction of 0.5 or 0.1. In some embodiments, the learning rate is a differential learning rate. In some embodiments, training the neural network also uses a scheduler that conditionally applies learning rate decay based on the evaluation of the performance metric within a threshold number of training epochs (e.g., applying learning rate decay when the performance metric fails to meet a threshold performance value after at least a threshold number of training epochs).

[0216] In some embodiments, performance metrics are used to measure the performance of the neural network at one or more time points. These performance metrics include, but are not limited to, training loss metrics, validation loss metrics, and / or mean absolute error. In some embodiments, the performance metrics are the area under the receiving operating features (AUROC) and / or the area under the precision-recall curve (AUPRC).

[0217] For example, in some embodiments, the performance of a neural network is measured by validating the model using a validation dataset (e.g., development). In some such embodiments, the neural network is trained to form a trained neural network when the neural network meets minimum performance requirements based on the validation.

[0218] In some embodiments, any suitable validation method may be used, including but not limited to K-fold cross-validation, advanced cross-validation, random cross-validation, grouped cross-validation (e.g., K-fold grouped cross-validation), bootstrap bias-corrected cross-validation, random search, and / or Bayesian hyperparameter optimization.

[0219] In some embodiments, a method is provided for training a model comprising multiple parameters by means of the following procedure, the procedure comprising (i) for each corresponding training subject among a plurality of training subjects, inputting a corresponding genomic abundance value for each corresponding gut microbe among a plurality of training subjects, thereby obtaining a corresponding prediction of the training subject’s response to a therapy as the output of the model for each corresponding training subject among a plurality of training subjects, and (ii) refining the multiple model parameters based on the difference between the corresponding actual response to the therapy for the training subject and the corresponding predicted response to the therapy for the training subject.

[0220] 2. Methods for using models to predict subjects' responses to treatments for a disease.

[0221] Figure 3 is a schematic diagram of a method for predicting a subject's response to a treatment for a condition, as described below. Method 300 can use a computer system (e.g., see above). Figure 1 The computer system 100 shown and described is used to implement this.

[0222] Referring to block 300, in some embodiments, the method includes acquiring a plurality of genomic abundance values ​​in electronic form, the plurality of genomic abundance values ​​including those selected from Table 1, Table 2, or Figures 13A to 13XX The abundance value of the genome of each of the multiple gut microbiota in the subject's biological sample.

[0223] In some embodiments, a corresponding biological sample from the gut of the subject is collected prior to treatment or therapy. In some embodiments, the biological sample is collected no more than 15 minutes, 30 minutes, 1 hour, 2 hours, 3 hours, 4 hours, 5 hours, 6 hours, 12 hours, or 24 hours prior to treatment or therapy. In some embodiments, the biological sample is collected 1 day, 2 days, 3 days, 4 days, 5 days, 6 days, 1 week, 2 weeks, 3 weeks, 4 weeks, or longer prior to treatment or therapy. In some embodiments, the biological sample is collected approximately 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, or longer prior to treatment or therapy.

[0224] In some embodiments, sample data (including plasma and stool samples) and corresponding clinical information (including sex / age / body fat count / underlying diseases / histopathological features, etc.) are collected for each subject prior to treatment. A whole-microbiome analysis is performed on the individual biological samples. In some embodiments, the samples are tissue biopsy samples, intestinal samples, or mucosal samples. See, for example, Tang Q, Jet et al., Current Sampling Methods for Gut Microbiota: ACall for More Precise Devices, Front Cell Infect Microbiol., 10:151 (2020), the contents of which are incorporated herein by reference in their entirety. In some embodiments, the biological sample from the gut of the respective subject is a stool sample from the respective subject.

[0225] In some embodiments, the multiple gut microbiota include those selected from Table 1, Table 2, or... Figures 13A to 13XX At least 25 gut microbiota. In some embodiments, the gut microbiota comprises those selected from Table 1, Table 2, or... Figures 13A to 13XX At least 30 gut microbiota. In some embodiments, the gut microbiota comprises those selected from Table 1, Table 2, or... Figures 13A to 13XX At least 40 gut microbiota. In some embodiments, the gut microbiota comprises those selected from Table 1, Table 2, or... Figures 13A to 13XX At least 25 gut microbiota. In some embodiments, the gut microbiota comprises those selected from Table 1, Table 2, or... Figures 13A to 13XXAt least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 125, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, or all of the gut microbiota. In some embodiments, the multiple gut microbiota are selected from Table 1, Table 2, or... Figures 13A to 13XX At least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 125, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, or all gut microbiota. In some embodiments, multiple gut microbiota refers to all gut microbiota listed in Table 1. In some embodiments, multiple gut microbiota refers to all gut microbiota listed in Table 2. In some embodiments, multiple gut microbiota is... Figures 13A to 13XX All gut microbes listed in the table.

[0226] In some embodiments, the corresponding value for genome abundance is a value representing the absolute abundance of the microbial genome. In some embodiments, the corresponding value for genome abundance is a value representing a normalized abundance value or a relative abundance value (e.g., the abundance of a microorganism normalized relative to the abundance of the total microbiome of interest). In some embodiments, the corresponding value for genome abundance is a value representing an average abundance value (e.g., the average abundance obtained at different time points or from different biological samples from a patient, or the average abundance obtained using different probes, etc.), or a combination of the above. The corresponding value for genome abundance is measured using any technique known in the art. In some embodiments, the genome abundance value is measured by quantitative PCR (qPCR) such as bacterial 16S rRNA qPCR, RT-PCR, or qRT-PCR to quantify the abundance of regions of interest in the genome, for example, as described in U.S. Patent No. 11,427,865, the disclosure of which is incorporated herein by reference in its entirety. In some embodiments, genome abundance values ​​are determined by using targeted sequencing (e.g., 16S rRNA sequencing or any other suitable biomarker), partial genome sequencing, or whole genome sequencing, thereby quantifying the number of reads of the target region in the microbial genome to determine genome abundance, for example as disclosed in U.S. Patent Application Publication No. 2021 / 0403986 or U.S. Patent No. 11,332,783, the disclosures of which are incorporated herein by reference in their entirety. In some embodiments, deep sequencing is used to determine the abundance of the target sequence, for example as disclosed in U.S. Patent Application Publication No. 2018 / 0237863, the disclosure of which is incorporated herein by reference in its entirety.In some embodiments, the sequencing depth is at least 2X, at least 3X, at least 4X, at least 5X, at least 6X, at least 7X, at least 8X, at least 9X, at least 10X, at least 11X, at least 12X, at least 13X, at least 14X, at least 15X, at least 16X, at least 17X, at least 18X, at least 19X, at least 20X, at least 21X, at least 22X, at least 23X, at least 24X, at least 25X, at least 26X, at least 27X, at least 28X, at least 29X, at least 30X, at least 31X, at least 32X, at least 33X, at least 34X, at least 35X, at least 36X, at least 37X, at least 38X, or at least 39X, at least 40X, at least 41X, at least 42X, at least 43X, at least 44X, at least 45X, at least 46X, at least 47X, at least 48X, at least 49X, at least 50X, at least 51X, at least 52X, at least 53X, at least 54X, at least 55X, at least 56X, at least 57X, at least 58X, at least 59X, at least 60X, at least 70X, at least 80X, at least 90X, at least 100X, at least 110X, at least 120X, at least 130X, at least 150X, at least 200X, at least 300X, at least 400X, at least 500X, at least 750X, at least 1000X or more. In some embodiments, shotgun metagenomic sequencing is used to provide sequence reads of the genome in the sample, for example, as described in U.S. Patent No. 11,028,449, the contents of which are incorporated herein by reference in their entirety.

[0227] In some embodiments of the methods described herein, if the identified genome construct has at least 97% sequence identity when compared with the microbial contigs provided in the sequence listing, the genomes identified in the metagenomic analysis will be classified as corresponding to Tables 1, 2, and / or Figures 13A to 13XX The microorganisms listed are shown in Figure 12. In some embodiments, if the identified genome construct has at least 98% sequence identity when compared with the microbial contiguous groups provided in the sequence listing, the genomes identified in the metagenomic analysis will be classified as corresponding to Tables 1, 2, and / or Figures 13A to 13XX The microorganisms listed are shown in Figure 12. In some embodiments, if the identified genome construct has at least 99% sequence identity when compared with the microbial contiguous groups provided in the sequence listing, the genomes identified in the metagenomic analysis will be classified as corresponding to Tables 1, 2 and / or Figures 13A to 13XXThe microorganisms listed are shown in Figure 12. In some embodiments, if the identified genome construct has at least 99.5% sequence identity when compared with the microbial contiguous groups provided in the sequence listing, the genomes identified in the metagenomic analysis will be classified as corresponding to the microorganisms listed in Tables 1, 2, and / or Figure AXX, as shown in Figure 12. In some embodiments, if the identified genome construct has at least 97%, at least 97.5%, at least 98%, at least 98.5%, at least 99%, at least 99.1%, at least 99.2%, at least 99.3%, at least 99.4%, at least 99.5%, at least 99.6%, at least 99.7%, at least 99.8%, at least 99.9% or higher sequence identity when compared with the microbial contiguous groups provided in the sequence listing, the genomes identified in the metagenomic analysis will be classified as corresponding to the microorganisms listed in Tables 1, 2, and / or Figure AXX, as shown in Figure 12. Figures 13A to 13XX The microorganisms listed are shown in Figure 12.

[0228] Referring to block 302, in some embodiments, the method includes sequencing genomic DNA from a biological sample of the gut of a subject to obtain a plurality of nucleic acid sequences. In some embodiments, the plurality of nucleic acid sequences comprises at least 10,000, at least 100,000, at least 250,000, at least 500,000, at least 1,000,000, at least 2,500,000, at least 5,000,000, at least 10,000,000, or at least 50,000,000 nucleic acid sequences. In some embodiments, the plurality of nucleic acid sequences comprises no more than 250,000,000, no more than 100,000,000, no more than 50,000,000, no more than 25,000,000, no more than 10,000,000, no more than 5,000,000, no more than 1,000,000, or no more than 100,000 nucleic acid sequences. In some embodiments, the plurality of nucleic acid sequences comprises 100,000 to 1,000,000, 200,000 to 5,000,000, 500,000 to 10,000,000, 1,000,000 to 20,000,000, 5,000,000 to 50,000,000, 10,000,000 to 100,000,000, or 50,000,000 to 250,000,000 nucleic acid sequences. In some embodiments, the plurality of nucleic acid sequences falls into another range starting at no less than 1,000 nucleic acid sequences and ending at no more than 250,000,000 nucleic acid sequences.

[0229] In some embodiments, multiple nucleic acid sequences are obtained via metagenomic sequencing, for example, as disclosed in U.S. Patent Application Publication No. 2016 / 0239602 or U.S. Patent No. 11,495,326, the contents of which are incorporated herein by reference in their entirety. In some embodiments, metagenomic sequencing further includes generating multiple metagenomic reads. In some embodiments, metagenomic sequencing further includes fragmenting the microbial genome into random fragments of a target size. The size of the resulting fragments may vary. In one embodiment, a fragment of approximately 500 nucleotides may be obtained. In some embodiments, fragments of 100-2000 nucleotides (e.g., 200-800, 100-900, 100-1000, 300-800, 400-900 nucleotides) may be obtained. In some embodiments, the method may further include extracting metagenomic fragments from the corresponding biological sample. In some embodiments, metagenomic sequencing further includes sequencing the fragments using high-throughput sequencing methods to generate multiple sequencing reads.

[0230] In some embodiments, targeted genome sequencing is used to obtain multiple corresponding nucleic acid sequences. Examples of targeted genome sequencing are described in U.S. Patent Application Publication No. 2019 / 0316209. In some embodiments, targeted genome sequencing includes hybridizing genomic DNA isolated from a biological sample from the gut of a subject with a set of probes prior to sequencing the recovered nucleic acids. This set of probes includes one or more probes that hybridize to unique sequences in the genome of each quantified microorganism. In some embodiments, the microorganisms include those listed in Tables 1 and 2 and / or... Figures 13A to 13XX The list includes a variety of microorganisms. In some embodiments, combinations of semi-unique sequences (e.g., sequences found in a small number of microbial genomes) can be used to deconvolve genome abundance values ​​using algorithms (e.g., systems of equations). In some embodiments, the probe set includes at least one probe that hybridizes to a unique sequence in each microbial genome to be detected. In some embodiments, the probe set includes at least two, at least three, at least four, at least five, at least ten, at least twenty-five, at least fifty, or more probes that hybridize to unique and distinct sequences in each microbial genome to be detected. In some embodiments, the probe set includes at least 20, at least 30, at least 40, at least 50, at least 75, at least 100, at least 125, at least 150, at least 200, at least 150, at least 300, at least 400, at least 500, at least 750, at least 1000, at least 1250, at least 1500, at least 2000, at least 2500, at least 3000, at least 4000, at least 5000, at least 7500, at least 10,000 or more unique probes.

[0231] In some embodiments, the sequenced genomic DNA from the corresponding biological sample contains at its ends partial or complete sequencing platform adaptor sequences that can be used for sequencing using the sequencing platform of interest. Sequencing platforms of interest include, but are not limited to, HiSeq™, MiSeq™, and Genome Analyzer™ sequencing systems from Illumina®; Ion PGM™ and Ion Proton™ sequencing systems from IonTorrent™; the PACBIO RSII Sequel system from Pacific Biosciences; the SOLiD sequencing system from Life Technologies™; the 454 GS FLX+ and GS Junior sequencing systems from Roche; the MinION™ system from Oxford Nanopore; or any other sequencing platform of interest.

[0232] Referring to block 304, in some embodiments, the method includes acquiring, in electronic form, multiple nucleic acid sequences targeting genomic DNA from a biological sample of the subject's gut.

[0233] Referring to block 306, in some embodiments, the method includes determining a corresponding value for the abundance of the genome of a given gut microbe from a plurality of nucleic acid sequences for each given gut microbe among a plurality of gut microbes. In some embodiments, the genome abundance values ​​determined for a subject include at least 20, at least 25, at least 50, at least 75, at least 100, at least 125, at least 150, at least 175, at least 200, at least 250, at least 300, at least 350, at least 400, at least 500, at least 600, at least 700, at least 700, at least 800, at least 900, at least 1000, at least 1500, at least 2000, at least 25000, at least 5000, or at least 10000 genome abundance values, wherein each genome abundance value corresponds to a different gut microbe. In some embodiments, genome abundance values ​​include genome abundance values ​​of no more than 250,000, no more than 100,000, no more than 50,000, no more than 25,000, no more than 10,000, no more than 5,000, no more than 2,500, no more than 1,000, no more than 750, no more than 500, or fewer. In some embodiments, genome abundance values ​​consist of 10 to 40, 20 to 50, 30 to 80, 40 to 100, 50 to 150, 60 to 200, 80 to 300, 90 to 500, 100 to 1,000, 500 to 2,000, or 1,000 to 5,000 genome abundance values. In some embodiments, the number of genome abundance values ​​falls into another range starting at no less than 10 genome abundance values ​​and ending at no more than 250,000 genome abundance values.

[0234] Referring to block 308, in some embodiments, the method includes assembling a plurality of gut microbial genomes in electronic form from a plurality of nucleic acid sequences via metagenomic de novo sequencing, and for each corresponding gut microbial in the plurality of gut microbes, calculating a corresponding value for the abundance of the genome of the corresponding gut microbe based on the universality of the corresponding nucleic acid sequence in the plurality of nucleic acid sequences used to assemble the corresponding gut microbial genome in the plurality of gut microbial genomes corresponding to the corresponding gut microbe. In some embodiments, metagenomic de novo sequencing further includes generating contigs based on sequencing reads generated by shotgun sequencing technology. This technique is described, for example, in U.S. Patent No. 10,529,443, the contents of which are incorporated herein by reference in their entirety. In some embodiments, the plurality of nucleic acid sequences may be assembled into a complete genome of a plurality of gut microbes. In some embodiments, the plurality of nucleic acid sequences may be assembled into a partial genome of a plurality of gut microbes.

[0235] Referring to block 310, in some embodiments, these methods include assigning each corresponding nucleic acid sequence from a plurality of nucleic acid sequences to a corresponding gut microbe from a plurality of gut microbes, thereby generating a corresponding count of the corresponding nucleic acid sequence from the plurality of nucleic acid sequences assigned to the corresponding gut microbe for each corresponding gut microbe from the plurality of gut microbes, and determining a corresponding genomic abundance value for the corresponding gut microbe based on the corresponding count of the corresponding nucleic acid sequence assigned to the corresponding gut microbe for each corresponding gut microbe from the plurality of gut microbes. In some embodiments, assigning each corresponding nucleic acid to a corresponding gut microbe includes mapping that nucleic acid to a reference nucleic acid. In some embodiments, assigning each corresponding nucleic acid to a corresponding gut microbe includes annotating genomic information based on an existing database. In some embodiments, nucleic acid sequences are analyzed, and the annotation is defined by using sequence similarity and phylogenetic placement methods or a combination of both strategies to define the classification assignment.

[0236] Sequence similarity-based methods for assigning each corresponding nucleic acid sequence in the corresponding gut microbiota include those familiar to those skilled in the art, including but not limited to BLAST, BLASTx, tBLASTn, tBLASTx, RDP classifiers, DNAclust, and various implementations of these algorithms (such as Qiime or Mothur). These methods rely on mapping sequence reads to a reference database and selecting matches with the best score and e-value. In some embodiments, phylogenetic methods are combined with sequence similarity methods to improve the accuracy of annotation or classification assignments. Common databases include, but are not limited to, GT-DBTK, NCBI Genbank, the European Bioinformatics Institute-European Nucleotide Archive (EBI-ENA), the US Department of Energy National Institute of Genetics (USDOE) Integrated Microbial Genomes & Microbiomes (IMG / M), and other databases available in the art.

[0237] Referring to box 312, in some embodiments, the multiple gut microbiota includes at least 20 microorganisms selected from Table 1, Table 2, or... Figures 13A to 13XX The gut microbiota includes those microorganisms having a connectivity of at least 2, at least 3, at least 4, at least 5, at least 10, at least 15, at least 20, at least 25, or greater. In some embodiments, the multiple gut microbiota comprises at least 20 microorganisms selected from Table 1, Table 2, or... Figures 13A to 13XXThe list includes those microorganisms with a connectivity degree of at least 2. In some embodiments, the multiple gut microbiota comprises at least 20 microorganisms selected from Table 1, Table 2, or Table 3. Figures 13A to 13XX The microorganisms listed in the document have a connectivity of at least 3, at least 4, at least 5, at least 10, at least 15, at least 20, at least 25 or greater.

[0238] Referring to block 314, in some embodiments, the biological sample from the subject's intestine is a fecal sample. In some embodiments, the sample is a tissue biopsy sample, an intestinal sample, or a mucosal sample. In some embodiments, the biological sample is a sample obtained from the small intestine or large intestine, preferably the colon or rectum, more preferably in the form of a fecal sample or rectal swab or a biopsy specimen of the gastrointestinal mucosa.

[0239] Referring to block 316, in some embodiments, the therapy is a biological therapy, immunotherapy, chemotherapy, radiotherapy, gene therapy, hormone therapy, photodynamic therapy, targeted therapy, small molecule, antibody, polynucleotide, natural compound, immunomodulator, bone marrow therapy, stem cell therapy, surgical therapy, induction therapy, maintenance therapy, or a combination thereof.

[0240] Referring to box 318, in some embodiments, the condition is selected from the group consisting of: type 2 diabetes, hypertension, schizophrenia, atherosclerotic cardiovascular disease (ACVD), cirrhosis (LC), inflammatory bowel disease (IBD), colorectal cancer (CRC), ankylosing spondylitis (AS) and Parkinson's disease (PD), inflammatory bowel disease (IBD), rheumatoid arthritis (RA), advanced melanoma, and B-cell lymphoma. In some embodiments, the condition is, for example, hypertension (HT), schizophrenia (SCZ), multiple sclerosis (MS), type II Gaucher disease (GDII), COVID-19 (COV), Behcet's disease (BD), autism spectrum disorder (ASD), or pancreatic cancer (PC). In some embodiments, the condition is cancer, Alzheimer's disease, cardiovascular disease, autoimmune disease, mental health disease, infectious disease, or genetic disease.

[0241] In some embodiments, a condition is categorized by any indicator of a patient's biological state, function, structure, process, response, or condition. Such indicators include any of a number of variables (parameters) commonly measured in medicine to assess a patient for purposes such as diagnosis, prognosis, and / or treatment. Generally, indicators of interest herein are those whose values ​​(which may be quantitative or qualitative) reflect, characterize, or relate to the function or structure of organs and organ systems, and / or whose values ​​reflect, characterize, or relate to the presence or severity of a condition. In some embodiments, a disease or symptom is categorized by its progression or prognosis, such as different stages of cancer, the type, frequency, or severity of a condition that can be objectively measured or experienced by the subject. In some embodiments, the condition may be obtained through a medical device that can be used to analyze the state of a body part, wound, or lesion, or the state of a physical matrix, such as a test strip, test strip, filter, or other matrix, which provides a means for detecting or measuring the presence of substances in a biological sample obtained from the subject. In some embodiments, it is intended to detect antibodies or biomarkers against pathogens (e.g., viruses, bacteria, fungi), abnormal tissues (e.g., tumor sites) in biological samples, and / or detect their presence in biological samples from patients, for purposes such as diagnosing the presence of symptoms or diseases.

[0242] Referring to block 320, in some embodiments, the condition is cancer.

[0243] In some implementations, the condition is inflammatory bowel disease and the treatment includes anti-cytokine or anti-integrin therapy. For example, as reported in Example 5, for each of three datasets containing microbiome metagenomic data collected before and after treatment with anti-cytokine or anti-integrin therapy, a random forest classifier was constructed based on the abundance of 284 HQMAGs in CC-TCG to predict treatment response. The model predicted treatment responses with ROC AUC values ​​ranging from 0.64 to 0.69.

[0244] The method according to any one of claims 19 to 29, wherein the condition is rheumatoid arthritis and the therapy comprises methotrexate treatment. For example, as reported in Example 5, for a dataset with microbiome metagenomic data collected before and after methotrexate treatment, a random forest classifier was constructed based on the abundance of 284 HQMAGs in CC-TCG to predict treatment response. The model predicted a treatment response with an ROC AUC value of 0.69.

[0245] The method according to any one of claims 19 to 29, wherein the condition is B-cell lymphoma and the therapy comprises CAR-T cell immunotherapy. In some embodiments, the CAR-T cell immunotherapy is CD19-CAR-T immunotherapy. For example, as reported in Example 5, for each of five datasets containing microbiome metagenomic data collected before and after treatment with CD19-CAR-T immunotherapy, a random forest classifier was constructed based on the abundance of 284 HQMAGs in CC-TCG to predict treatment response. The model predicted treatment responses with ROC AUC values ​​ranging from 0.64 to 0.69.

[0246] The method according to any one of claims 19 to 29, wherein the condition is melanoma and the therapy comprises immune checkpoint inhibitor treatment. For example, as reported in Example 5, for one of two datasets containing microbiome metagenomic data collected before and after treatment with an immune checkpoint inhibitor (ICI), a random forest classifier was constructed based on the abundance of 284 HQMAGs in CC-TCG to predict treatment response. The model predicted treatment outcomes on the training dataset with an AUC of 0.66, and on the other dataset not used for training, the AUC was 0.64.

[0247] Referring to box 322, in some embodiments, the method includes inputting multiple genomic abundance values ​​into a model that includes multiple parameters. The model applies the multiple parameters to the multiple genomic abundance values ​​through, for example, at least 10,000 calculations to generate a prediction of a subject's response to a therapy as output from the model.

[0248] In some embodiments, the model is trained on a dataset collected from multiple therapies for a condition, and the model is trained to distinguish between response and non-response states. In some embodiments, the model includes a learned statistical classifier system. In some embodiments, the learned statistical classifier system is a random forest, a classification and regression tree, a boosting tree, or a neural network. For example, as described in Example 3, the random forest classifier was trained on a dataset from 11 different studies that collectively investigated the microbiome in four different conditions. As shown in Figure 8C, the resulting model was able to predict whether a patient would respond to or not respond to anti-cytokine or anti-integrin therapy, methotrexate treatment for new-onset rheumatoid arthritis, immune checkpoint inhibitor (ICI) therapy for advanced melanoma, and CD19-CAR-T immunotherapy for B-cell lymphoma.

[0249] In some embodiments, subject response is characterized by clinical outcome measures, including but not limited to complete response, partial response, no response, survival, development of adverse events, or any combination thereof. In some embodiments, a responder has a complete response to treatment, while a non-responder has no response or partial response. In some embodiments, patients undergo routine clinical examination, laboratory analysis, and computed tomography. Tumor response is assessed using RECIST criteria. In some embodiments, a complete response is defined as the complete radiographic disappearance of the minimum radiographically detectable disease or stable disease; a partial response is defined as a measurable reduction of at least 50% in the longest dimension of the disease; disease stability is defined as a reduction of less than 25% in the longest dimension; and progressive disease is defined as tumor growth of more than 25% in the longest dimension or the development of new lesions. In some embodiments, the overall response rate is defined as the sum of the complete and partial response rates, and the tumor control rate is defined as the sum of the overall response rate and the disease stability rate.

[0250] In some embodiments, the indication of subject response is characterized by the actual therapeutic efficacy of the therapy, including progression-free survival (PFS), duration of progression-free survival during treatment, overall survival (OS), response to the therapy (RT), overall response rate (ORR), sustained clinical effect (DCB), disease activity score, or any combination thereof, or any other method known in the art for assessing the progression or prognosis of a disease or condition.

[0251] In some embodiments, “progression-free survival” (PFS) has the meaning as understood in the art, referring to the length of time a patient has had a disease (such as cancer) during and after treatment for that disease without it becoming progressive. In some embodiments, measuring progression-free survival is used as an assessment of the effectiveness of a new treatment. In some embodiments, PFS is determined in a randomized clinical trial; in some such embodiments, PFS refers to the time from randomization until objective tumor progression and / or death.

[0252] In some embodiments, ORR may be defined as the proportion of patients who identify a partial (PR) or complete (CR) response as the best overall response (BOR) based on some metrics, such as the Response Evaluation Criteria in Solid Tumors (RECIST 1.1). Stable disease (SD) is categorized as non-responsive along with progressive disease (PD). In some embodiments, ORR has its meaning as understood in the art as the proportion of patients whose tumor size decreases by a predetermined amount and for a minimum duration. In some embodiments, the duration of response is typically measured from the initial response time up to the recorded tumor progression. In some embodiments, ORR involves the sum of a partial response and a complete response.

[0253] In some embodiments, "clinical effect" refers to clinical benefit. In some embodiments, such clinical benefit is or includes a reduction in tumor size, an increase in progression-free survival, an increase in overall survival, a reduction in total tumor burden, and a reduction in symptoms caused by tumor growth, such as pain, organ failure, bleeding, skeletal damage, and other related sequelae of metastatic cancer, and combinations thereof. In some embodiments, clinical effect is a "continuous clinical effect" (DCB) maintained over a relevant time period. In some embodiments, the relevant time period is at least 1 month, 2 months, 3 months, 4 months, 5 months, 6 months, 7 months, 8 months, 9 months, 10 months, 11 months, 1 year, 2 years, 3 years, 4 years, 5 years, or longer.

[0254] In some embodiments, subject responses are measured using a Disease Activity Score (DAS) (see, for example, Vander Heijde DM et al., J Rheumatol, 1993, 20(3): 579-81; Prevoo ML et al., Arthritis Rheum, 1995, 38: 44-8). The DAS system represents the current state and changes in disease activity. The DAS scoring system uses a weighted mathematical formula derived from clinical trials of RA. For example, DAS 28 is 0.56(T28) + 0.28(SW28) + 0.70(Ln ESR) + 0.014 GH, where T represents the number of tender joints, SW represents the number of swollen joints, ESR represents the erythrocyte sedimentation rate, and GH represents overall health status. Various DAS values ​​represent high or low disease activity and remission, and changes and endpoint scores lead to patient classification by degree of response (none, moderate, good).

[0255] In some embodiments, the level of an immune response or immune parameter in a cancer patient induced by immunotherapy is used to measure an indicator of the subject's response. In some embodiments, the immune response or immune parameter is characterized by the expression levels of various biomarkers of the host immune response in conjunction with the occurrence of cancer at a given stage of cancer development (i.e., treatment efficacy). In some embodiments, the expression levels of the biomarker are compared to, and, when necessary, to, a reference value. Thus, the reference value for the same biomarker is predetermined and is known to indicate a reference value associated with distinguishing between low and high levels of immune response to the biomarker in cancer patients. The predetermined reference value for the biomarker is associated with cancer patients who respond to treatment, or conversely, with cancer patients who do not respond to treatment.

[0256] In some embodiments, the variation in combinations of biomarkers is quantified. In some embodiments, combinations of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50 or more different biomarkers are quantified.

[0257] In some embodiments, immunohistochemistry is used to quantify biomarkers. Example biomarkers include 18S, ACE, ACTB, AGTR1, AGTR2, APC, APOA1, ARF1, AXIN1, BAX, BCL2, BCL2L1, CXCR5, BMP2, BRCA1, BTLA, C3, CASP3, CASp9, CCL1, CCL11, CCL13, CCL16, CCL17, CCL18, CCL19, CCL2, CCL20, CCL21, CCL22, CCL23, CCL24, CCL25, CCL26, CCL27, CCL28, CCL3, CCL5, CCL7, CCL8, CCNB1, CCND1, and C. CNE1, CCR1, CCR10, CCR2, CCR3, CCR4, CCR5, CCR6, CCR7, CCR8, CCR9, CCRL2, CD154, CD19, CD1a, CD2, CD226, CD244, PDCD1LG1, CD28, CD34, CD36, CD3 8. CD3E, CD3G, CD3Z, CD4, CD40LG, CD5, CD54, CD6, CD68, CD69, CLIP, CD80, CD83, SLAMF5, CD86, CD8A, CDH1, CDH7, CDK2, CDK4, CDKN1A, CDKN1B, CDKN 2A, CDKN2B, CEACAM1, COL4A5, CREBBP, CRLF2, CSF1, CSF2, CSF3, CTLA4, CTNN81, CTSC, CX3CL1, CX3CRI, CXCL1, CXCL10, CXCL11, CXCL12, CXCL13, CX CL14, CXCL16, CXCL2, CXCL3, CXCL5, CXCL6, CXCL9, CXCR3, CXCR4, CXCR6, CYP1A2, CYP7A1, DCC, DCN, DEFA6, DICER1, DKK1, Dok-1, Dok-2, DOK6, DVL1 , E2F4, EBI3, ECE1, ECGF1, EDN1, EGF, EGFR, EIF4E, CD105, ENPEP, ERBB2, EREG, FCGR3A, CGR3B, FN1, FOXP3, FYN, FZD1, GAPD, GLI2, GNLY, GOLPH4, GR B2, GSK3B, GSTP1, GUSB, GZMA, GZMH, GZMK, HLA-B, HLA-C, HLA-, MA, HLA-DMB, HLA-DOA, HLA-DOB, HLA-DPA1, HLA-DQA2, HLA-DRA, HLX1, HMOX1, HRAS,HSPB3, HUWE1, ICAM1, ICAM-2, ICOS, ID1, ifna1, ifna17, ifna2, ifna5, ifna6, ifna8, IFNAR1, IFNAR2, IFNG, IFNGR1, IFNG R2, IGF1, IHH, IKBKB, IL10, IL12A, IL12B, IL12RB1, IL12RB2, IL13, IL13RA2, IL15, IL15RA, IL17, IL17R, IL17RB, IL18, IL1 A, IL1B, IL1RI, IL2, IL21, IL21R, IL23A, IL23R, IL24, IL27, IL2RA, IL2RB, IL2RG, IL3, IL31RA, IL4, IL4RA, IL5, IL6, IL7, IL7RA, IL8, CXCR1, CXCR2, IL9, IL9R, IRF1, ISGF3G, ITGA4, ITGA7, integrin, αE (antigen CD103, human mucosal lymphocytes, antigen 1; α polypeptide), hCG33203 gene, ITGB3 , JAK2, JAK3, KLRB1, KLRC4, KLRF1, KLRG1, KRAS, LAG3, LAIR2, LEF1, LGALS9, LILRB3, LRP2, LTA, SLAMF3, MADCAM1, MADH3, M ADH7, MAF, MAP2K1, MDM2, MICA, MICB, MKI67, MMP12, MMP9, MTA1, MTSS1, MYC, MYD88, MYH6, NCAM1, NFATC1, NKG7, NLK, NOS2A, P2X7, PDCD1, PECAM-, CXCL4, PGK1, PIAS1, PIAS2, PIAS3, PIAS4, PLAT, PML, PP1A, CXCL7, PPP2CA, PRF1, PROM1, PSMB5, ​​PTCH, PTGS2, PTP4A3, PTPN6, PTPRC, RAB23, RAC / RHO, RAC2, RAF, RB1, RBL1, REN, Drosha, SELE, SELL, SELP, SERPINE1, SFRP1, SIRP β1, SKI, SLAMF1, SLAMF6, SLAMF7, SLAMF8, SMAD2, SMAD4, SMO, SMOH, SMURF1, SOCS1, SOCS2, SOCS3, SOCS4, SOCS5 , SOCS6, SOCS7, SOD1, SOD2, SOD3, SOS1, SOX17, CD43, ST14, STAM, STAT1, STAT2, STAT3, STAT4, STAT5A, STAT5B,STAT6, STK36, TAP1, TAP2, TBX21, TCF7, TERT, TFRC, TGFA, TGFB1, TGFBR1, TGFBR2, TIMP3, TLR1, TLRO1, TLR2, TLR3, TLR4, TLR5, TLR6, TLR7, TLR8, TLR9, TNF, TNFRSF10A, TNFRSF11A, TNFRSF18, TNFRSF1A, TNFRSF1B, OX-40, TNFRSF5, TNFRSF6, TNFRSF7, TNFRSF8, TNFRSF9, TNFSF10, TNFSF6, TOB1, TP53, TSLP, VCAM1, VEGF, WIF1, WNT1, WNT4, XCL1, XCR1, ZAP70, and ZIC2. ,

[0258] Referring to block 324, in some embodiments, the prediction of a subject's response is a category output of the corresponding response among a plurality of possible responses for the respective subject. This method allows for setting a single "cutoff" value, allowing for differentiation between responders and non-responders to treatment. In some embodiments, the prediction of a subject's response includes a prediction of the objective response rate of the human subject to the treatment or therapy, and wherein the prediction of the objective response rate includes an indication or classification of the amount of a complete or partial response to the treatment.

[0259] Referring to block 326, in some embodiments, the prediction of a subject's response is a probability output for the response of the corresponding subject. As disclosed above, the method allows setting a single "cutoff" value, allowing differentiation between responders and non-responders to treatment. In some embodiments, the method includes using the model to calculate a probability value for the subject; comparing the probability value to a threshold derived from a cohort of responders / non-responders to determine whether the probability value is above or below the threshold; and classifying the subject as a responder / non-responder if the probability value is above or below the threshold. In embodiments, the threshold may be a probability value of at least 50%, 55%, 50%, 65%, 70%, 75%, or about 80% or more. In other embodiments, the probability value is a positive predictive value as measured by the area under the curve (AUC) of a receiver operating characteristic (ROC) curve. In some embodiments, a multivariate logistic regression model, a neural network model, a random forest model, or a decision tree model is used to calculate the probability value.

[0260] Referring to box 328, in some embodiments, the model is a neural network algorithm, support vector machine algorithm, naive Bayes algorithm, nearest neighbor algorithm, boosting tree algorithm, random forest algorithm, convolutional neural network algorithm, decision tree algorithm, regression algorithm, or clustering algorithm.

[0261] Referring to block 330, in some embodiments, the plurality of parameters are at least 1,000, at least 10,000, at least 15,000, at least 50,000, at least 100,000, at least 250,000, at least 500,000, at least 1,000,000, at least 2,500,000, at least 5,000,000, at least 10,000,000, or more parameters.

[0262] Referring to block 332, in some embodiments, the model applies multiple parameters to the information through at least 1,000 calculations, at least 5,000 calculations, at least 10,000 calculations, at least 25,000 calculations, at least 50,000 calculations, at least 100,000 calculations, at least 250,000 calculations, at least 500,000 calculations, at least 1,000,000 calculations, at least 2,500,000 calculations, at least 5,000,000 calculations, at least 10,000,000 calculations or more, to obtain the corresponding output of the subject from the model.

[0263] Referring to block 334, in some embodiments, the method further includes treating the subject by: i) administering the therapy to the subject when the predicted response to the therapy meets a threshold probability that the subject will have a favorable response to the therapy; ii) administering one or more of a variety of gut microbiota to the subject when the predicted response to the therapy does not meet a threshold probability that the subject will have a favorable response to the therapy.

[0264] In some embodiments, the administration includes identifying one or more of a plurality of gut microbes that are underrepresented in the subject, for example, as determined based on corresponding genomic abundance values ​​of the microbes, and administering the identified one or more gut microbes to the subject. In some embodiments, the identification includes determining whether the abundance of the gut microbes (e.g., as determined based on corresponding genomic abundance values ​​of the microbes) meets a corresponding threshold amount. When the abundance of the microbes does not meet the corresponding threshold amount, the identified microbes are used for administration. In some embodiments, the corresponding threshold amount is a relative abundance. In some embodiments, the corresponding threshold amount is a quantity relative to the abundance of one or more different gut microbes in the subject. In some embodiments, the corresponding threshold amount is a quantity relative to the total abundance of a plurality of gut microbes in the subject.

[0265] In some embodiments, the application includes applying a predefined group of microorganisms. In some embodiments, the predefined microbiome includes at least five microorganisms selected from Table 1, Table 2, or... Figures 13A to 13XXThe gut microbiota. In some embodiments, the predefined microbiome includes at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 55, and 60 species. 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 175, 180, 190, 200, 225, 250, 275, 300, 350, 400, 450, 500, 600, 700 or more varieties are selected from Table 1, Table 2 or Figures 13A to 13XX Gut microbiota.

[0266] In some embodiments, the predefined microbiome includes only those selected from Table 1, Table 2, or... Figures 13A to 13XX The gut microbiota assigned to community 1. That is, the predefined microbiome does not include those selected from Table 1, Table 2, or... Figures 13A to 13XX The microorganisms assigned to community 2. In some embodiments, the predefined microbiome includes at least 5 species selected from Table 1, Table 2, or... Figures 13A to 13XX The gut microbiota assigned to flora 1. In some embodiments, the predefined microbiome includes at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, and 45 species. 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 175, 180, 190, 200, 225, 250, 275, 300, 350, 400 or more varieties are selected from Table 1, Table 2 or... Figures 13A to 13XX The gut microbiota allocated to flora 1.

[0267] In some embodiments, the predefined microbiome includes only those selected from Table 1, Table 2, or... Figures 13A to 13XX The gut microbiota assigned to community 2. That is, the predefined microbiome does not include those selected from Table 1, Table 2, or... Figures 13A to 13XX The microorganisms assigned to community 1. In some embodiments, the predefined microbiome includes at least 5 species selected from Table 1, Table 2, or... Figures 13A to 13XXThe gut microbiota assigned to community 2. In some embodiments, the predefined microbiome includes at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, and 45 species. 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 175, 180, 190, 200, 225, 250, 275, 300, 350, 400 or more varieties are selected from Table 1, Table 2 or... Figures 13A to 13XX The gut microbiota allocated to Group 2.

[0268] In some embodiments, when the predicted response to the therapy does not meet a threshold probability that the subject will have a favorable response to the therapy, the method further includes administering the therapy to the subject. In some embodiments, the therapy is administered to the subject at approximately the same time as the administration of one or more of a plurality of gut microbes. In some embodiments, the therapy is administered to the subject after the administration of one or more of a plurality of gut microbes. In some embodiments, the therapy is administered to the subject for at least 1 day, at least 2 days, at least 3 days, at least 4 days, at least 5 days, at least 6 days, at least 7 days, at least 1 week, at least 2 weeks, at least 3 weeks, at least 4 weeks, at least 5 weeks, at least 6 weeks, at least 7 weeks, at least 8 weeks, or longer after the administration of one or more of a plurality of gut microbes. In some embodiments, the therapy is administered to the subject for no more than 3 months, no more than 2 months, no more than 1 month, no more than 4 weeks, no more than 3 weeks, no more than 2 weeks, no more than 1 week, no more than 6 days, no more than 5 days, no more than 4 days, no more than 3 days, or no more than 2 days after the administration of one or more of a plurality of gut microbes. In some embodiments, the treatment is administered to the subject 1 day to 2 months, 1 day to 1 month, 1 day to 3 weeks, 1 day to 2 weeks, 1 day to 1 week, 1 day to 3 days, 2 days to 2 months, 2 days to 1 month, 2 days to 3 weeks, 2 days to 2 weeks, 2 days to 1 week, 2 days to 3 days, 3 days to 2 months, 3 days to 1 month, 3 days to 3 weeks, 3 days to 2 weeks, 3 days to 1 week, 1 week to 2 months, 1 week to 1 month, 1 week to 3 weeks, or 1 week to 2 weeks after administration of one or more of a variety of gut microbes.

[0269] In some embodiments, if a subject is classified as a predicted nonresponder prior to treatment, the clinician may treat that subject differently than a subject classified as a predicted responder. Classifying a subject as a predicted nonresponder or a predicted responder may allow for the use of a more patient-specific or alternative treatment regimen.

[0270] In some embodiments, the treatment regimen or therapy may be administered via any common route of administration, provided that the target tissue or cells are available via that route. This includes, but is not limited to, intravenous, catheter-directed, in situ, intradermal, subcutaneous, intramuscular, intraperitoneal intratumoral administration, oral, nasal, buccal, rectal, vaginal, or local application. The choice of therapeutic agent and dosage regimen may depend on various factors, such as the combination of drugs used, the specific disease being treated, and the patient's condition and medical history.

[0271] In some embodiments, one or more of a variety of gut microbiota may be administered to a non-responsive individual via, but not limited to, oral administration or colonoscopy. Gut microbiota therapeutic compositions as described herein can be prepared and administered using methods known in the art. Typically, the compositions are formulated for oral, colonoscopy, or nasogastric delivery, although any suitable method may be used.

[0272] In some embodiments, non-responders receive fecal microbiota transplantation from a responder population using methods disclosed, for example, in US 20230109343, US 20200147151, or US 2021036172. In some embodiments, non-responders receive an effective amount of preselected isolated gut microbiota from fecal material of responders. In some embodiments, non-responders receive an effective amount of fecal microbiota from Table 1, Table 2, or Table 2. Figures 13A to 13XX A pre-selected isolated gut microbiota. In some embodiments, one or more of the various gut microbiota administered to non-responders comprise a therapeutically effective or sufficient amount selected from Tables 1, 2, or 3. Figures 13A to 13XXAn isolated or purified gut microbial community of at least 1, at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80, at least 85, at least 90, at least 95, at least 100, at least 125, at least 150, at least 200, at least 250, at least 300, at least 400, at least 500, at least 600, at least 700, at least 800, or all of the gut microbiota. In some embodiments, one or more of the gut microbiota administered to a non-responder comprises at least about 1 × 10⁻⁶. 3 One live colony-forming unit (CFU) of bacteria or at least about 1 × 10⁻⁶ 4 1×10 5 1×10 6 1×10 7 1×10 8 1×10 9 1×10 10 1×10 11 1×10 12 1×10 13 1×10 14 1×10 15 One live CFU (or any derived range thereof). In some embodiments, a single dose will contain a quantity of gut microbiota (such as a specific bacterium or species, genus or family described herein), wherein the live CFU of the specific bacterium is at least, at most or exactly 1 × 10⁻⁶. 4 1×10 5 1×10 6 1×10 7 1×10 8 1×10 9 1×10 10 1×10 11 1×10 12 1×10 13 1×10 14 1×10 15 or greater than 1×10 15 (or any derivable range thereof). In some embodiments, a single dose will contain at least, at most, or exactly 1 × 10⁻⁶. 4 1×10 5 1×10 6 1×10 7 1×10 8 1×10 9 1×10 10 1×10 111×10 12 1×10 13 1×10 14 1×10 15 or greater than 1×10 15 Total gut microbiota of live CFU (or any of the derived ranges thereof).

[0273] In some embodiments, multiple gut microbiota are administered simultaneously or sequentially with one or more therapies targeting a disease or condition. In some embodiments, some, most, or substantially all of the colon, gut, or gut microbiota of the subject are removed prior to administration of the composition.

[0274] In some embodiments, multiple gut microbiota are administered more than once. In some respects, the composition is administered daily, weekly, or monthly. In some embodiments, multiple gut microbiota are administered for two, three, or four months to induce and / or maintain an appropriate microbiome in the GI tract of unresponsive individuals.

[0275] Example

[0276] Example 1 – Two competing bacterial groups identified in the QD trial (QD-TCG) distinguished cases from controls in 10 independent case-control metagenomic datasets for 6 different diseases.

[0277] 1.1 Reversible changes in the gut microbiota are correlated with reversible changes in the host metabolic phenotype.

[0278] QD is an open-label, randomized, controlled intervention trial in which patients with type 2 diabetes mellitus (T2DM) were randomized at baseline (M0) to receive either a 3-month (M3) high-fiber intervention (W group; n=74) or standard of care (U group; n=36), followed by a one-year follow-up (M15). Figure 4A Throughout the trial, dietary fiber intake remained constant in group U, while dietary fiber intake in group W showed a significant increase from M0 to M3, and a decrease from M3 to M15. Figure 4B Compared to group U, group W had significantly higher fiber intake at M3 and M15. Figure 4B However, energy and protein consumption were similar between the two groups throughout the study period. Although a decreasing trend was observed in the W group, there was no significant difference in fat intake between the two groups. Compared to the U group, the W group had similar carbohydrate consumption at M0 and M3, but higher consumption at M15.

[0279] To investigate the gut microbiota response to the introduction and withdrawal of a high-fiber intervention, we performed shotgun metagenomic sequencing on 315 stool samples collected from 110 patients in groups W and U. 95 patients provided samples from all three time points, while 15 patients provided samples only from M0 and M3. For genome-level resolution, we reconstructed 1,845 non-redundant, high-quality draft genomes (HQMAGs, where two HQMAGs were superimposed if the average nucleotide identity (ANI) between them was >99%) from the metagenomic dataset. These HQMAGs comprised over 70% of the total reads. In the context of β-diversity based on Bray-Curtis distance, the overall structure of the gut microbiota in group W changed significantly from M0 to M3 (PERMANOVA test, P < 0.001), reverting to the structure of M0 at M15; in group U, there was no difference across the three time points. Figures 4C to 4D Similar changes in abundance-weighted α-diversity based on Shannon and Simpson indices were also observed. Regarding the unweighted abundance index, both the abundance and Chao1 in groups W and U increased from M0 to M3 and remained at M15. These results indicate that the high-fiber intervention induced significant structural and abundance changes in the gut microbiota; however, upon withdrawal from the intervention, the gut microbiota structure returned to baseline, demonstrating high resilience of the community structure, although there may be a delay in the increase in abundance.

[0280] To determine whether host metabolic phenotype would show reversible changes similar to those in the gut microbiota, we examined 43 bioclinical parameters across 9 categories at 3 time points. Hemoglobin A1c (HbA1c) in group U showed no significant change throughout the trial. The high-fiber intervention significantly reduced HbA1c levels in group W from M0 to M3 by 15.22% ± 9.82% (mean ± SD). During the one-year follow-up in group W, HbA1c significantly increased from M3 but remained lower than at M0. Figure 4E At M3, the proportion of patients achieving adequate glycemic control (HbA1c ≤ 7%) was significantly higher in the W group (61.6% vs. 33.3% in the U group), but no difference was observed between the two groups at M15. Figure 4F In meal tolerance tests, fasting blood glucose and postprandial blood glucose levels showed similar trends to HbA1c. Figure 4G , Figure 4H Of the remaining 40 bioclinical parameters, 14 also showed remission from M0 to M3, but rebounded during the one-year follow-up in the W group. These results suggest that changes in host metabolic phenotype are associated with reversible changes in the gut microbiota in response to the introduction or withdrawal of a high-fiber intervention.

[0281] 1.2 Genome pairs with stable interactions form a seesaw-like network of two competing bacterial communities.

[0282] To facilitate the identification of genomic pairs that maintained stable ecological interactions during the trial, particularly in the W group with profound microbiota and host phenotypic changes, we constructed co-abundance networks for each time point based on the abundance matrix of HQMAGs representing prevalent microbes. Co-abundance networks are a data-driven approach to studying ecological interactions between microbes across habitats. A total of 477 HQMAGs were selected for network construction because they were detectable in over 75% of the samples at each time point in the W group. These 477 HQMAGs also account for approximately 60% of the total abundance of 1,845 HQMAGs. In the W group, we calculated the pairwise correlations of all 113,526 possible genomic pairs in these 477 prevalent HQMAGs based on the abundance between patients at each time point and constructed three co-abundance networks (G) using Fastspar. M0 G M3 and G M15 Fastspar is a fast and scalable tool for estimating the correlation of constituent data. These three networks have similar orders S, i.e., the total number of nodes (HQMAG), S... M0 (442) S M3 (421) and S M15 (429), but they are in size L, i.e., the total number of edges (correlation) L M0 (4,231), L M3 (2,587) and L M15 There is a significant difference between (4,592) and G. M3 L in the middle drops to G M0 61.14% of the total, while G M15 L rises to G M0 108.53% of the total. The change in connectivity confirms this pattern; connectivity is defined as the proportion of ecological interactions realized between potential ecological interactions (in undirected networks, connectivity = ...). (Range: [0,1]). Connectivity from G M0 The value decreased from 0.043 to G M3 0.029, and rebounded to G M15 The changes in L and connectivity indicate that high-fiber intervention significantly reduces the correlation between generalized genomes in the network. Furthermore, we found that the distribution of degree, i.e., the number of edges a node has, closely matches the power-law model (R0). 2 Value G M0 0.79, G M3 0.82, G M15 (0.79), indicating a network hub21 The existence of hubs. Defining a hub as a node connecting more than one-fifth of the total nodes in the network, we found 24 hubs, 10 of which are in G. M0 20 in G M15 But none of them are in G M3 These results indicate that the overall structure of the gut microbiome underwent profound changes during the experiment, particularly the loss of interactions between genome pairs due to the high-fiber intervention.

[0283] We believe that if two genomes maintain the same type of correlation at all three time points, they have a strong and stable ecological relationship. Of the 113,526 possible genome pairs, 92.39% showed no correlation at any of the three time points, indicating that even brief ecological relationships between two genomes are rare. Figure 5A Interestingly, at all three time points, 517 genome pairs showed positive correlations, while 118 showed negative correlations. All 635 stable correlations involved 184 HQMAGs. Based on connected component clustering analysis, these HQMAGs were divided into 17 disconnected clusters. Clusters C2-C16 contained only 2 to 9 HQMAGs, which were only positively correlated. Cluster C1 was more complex, as it had 141 HQMAGs, which showed both negative and positive correlations. To identify the existence of sub-clusters in C1, we further constructed a clustering tree based on negative and positive correlations using the average linkage method and applied WGCNA analysis. The 141 genomes in C1 were further divided into two sub-clusters, C1A and C1B. Interestingly, among C1A, C1B, and C2-C16, only C1A and C1B were significantly correlated with our primary result, HbA1c. Figure 5B Therefore, we will focus on C1A and C1B for further analysis.

[0284] C1A and C1B can be considered as microbial communities because the HQMAGs in each cluster are highly correlated, both robust and transient, showing only positive correlations. Figure 5B These two bacterial communities are connected only by negative edges, indicating a competitive relationship that forms a seesaw-like network. This type of network characteristic is called Two Competing Groups (TCGs). Members of the TCG have significantly higher degree, betweenness centrality, eigenvector centrality, proximity centrality, and pressure centrality than other genomes in the network. This finding suggests that these two communities exert relatively strong control over the interactions (reflected by betweenness centrality and eigenvector centrality) and information flow (reflected by proximity centrality and pressure centrality) of other nodes in the network. Deleting either community would cause the network to collapse, as an average of 86.08% of the total edges would be lost. This suggests that the genomes of the two communities can be considered as three large networks (G...M0 G M3 and G M15 These are the core nodes of the gut microbiota network, because they are not only the most stable nodes in the gut microbiota network, but also the nodes with the highest degree of connectivity.

[0285] Members of both microbial communities were also very prevalent among participants, with 137 members present in >90% of individuals in both groups W and U, and 95 members present in 100% of individuals. Furthermore, the majority of these 141 HQMAGs were also major members of the gut microbiota, with 111 having an abundance higher than the median of 1,845 HQMAGs and representing 20.78% of the total sequencing reads. Beta diversity analysis based on Bray-Curtis distance showed significant correlations between the members of both microbial communities and the profile of all 1,845 HQMAGs, as demonstrated by the Mantel test (R²). 2 =0.62, P=0.001) and Procrustes analysis (P=0.001) prove that ( Figure 4C , Figure 4D This indicates that changes in these two bacterial communities led to major changes in the entire gut microbiota at three time points.

[0286] Our data indicate that 141 HQMAG TCGs exist in all three ecosystem networks G of group W. M0 G M3 and G M15 Furthermore, the presence of TCG in the W group at M0 indicates that this microbial organization exists regardless of whether a high-fiber intervention was performed in our study. Given the similarity of the overall gut microbiota structure between the W and U groups at M0 and the U group at all three time points ( Figure 4C , Figure 4D We hypothesized that TCGs could also be observed in Group U throughout the experiment. Therefore, we constructed an abundance network based on the abundance of 141 HQMAGs in individuals within Group U at each time point. In the co-abundance network, 99.8%, 99.51%, and 99.74% of the total edges were consistent with our TCGs, indicating a positive correlation within the microbiota and a negative correlation between microbiota. This suggests that TCG detection is independent of the high-fiber intervention, indicating that this pattern may be an inherent structure of the gut microbiome in our study.

[0287] 1.3 The dynamics of the two competing bacterial communities are related to changes in the host's metabolic phenotype.

[0288] We sought to determine whether TCG balance could be modulated by dietary fiber and to describe how TCG affects the host's metabolic phenotype. In the W group, from M0 to M3, the abundance of microbiota 1 significantly increased, while the abundance of microbiota 2 significantly decreased. Then at M15, microbiota 1 decreased to a level similar to M0, while microbiota 2 increased, but not significantly different from M3. Subsequently, from M0 to M3, the high-fiber intervention significantly increased the ratio of microbiota 1 to microbiota 2. During the one-year follow-up, this ratio significantly decreased ( Figure 6A Throughout the trial, the abundance and proportion of the two bacterial communities in group U remained unchanged. These results indicate that changes in the balance between the two communities are accompanied by changes in dietary fiber intake, overall gut microbiota, and host phenotype. To further explore the importance of the two communities to host health, we applied a linear mixed-effects model to identify the associations between the two communities and each host bioclinical parameter. The model was trained using the abundance of the two communities and the bioclinical parameters at M0 and M3. Subsequently, based on the community abundance at M15, the model was used to generate predicted values ​​for each bioclinical parameter. These predicted values ​​were then correlated with the bioclinical parameters measured at M15. Forty-two of the 43 bioclinical parameters showed significant Pearson correlation coefficients between predicted and measured values, ranging from 0.14 to 0.88 (adjusted p-value < 0.05). Figure 6B These results indicate that TCG constitutes an important microbiome feature of T2DM and related metabolic phenotypes.

[0289] Next, we performed a genome-centric analysis of 141 HQMAGs in the TCG to explore the genetic basis of the association between dynamic changes in seesaw-like microbiome traits and host metabolic phenotypes. Since dietary fiber can alter the balance between the two communities, we first attempted to identify genes encoding carbohydrate-active enzymes (CAZy) and genes encoding key enzymes in short-chain fatty acid (SCFA) production to compare the genetic capacity for carbohydrate utilization between the two communities. Compared to the genome in community 2 (C1B), the genome in community 1 (C1A) was richer in CAZy genes targeting arabinoxylan (P<0.001) and cellulose (P<0.01), and had a lower proportion of CAZy genes targeting inulin utilization (P<0.01) (Figure 6C). There were no differences in genes targeting starch, pectin, and mucin utilization between the two communities. Our previous research has shown that the gut microbiota benefits patients with type 2 diabetes mellitus (T2DM) through the production of acetic acid and butyric acid via carbohydrate fermentation. 11Among the terminal genes of the butyrate biosynthesis pathway from both carbohydrates (i.e., but and buk) and proteins (i.e., atoA / D and 4Hbt), the copy number of but was significantly higher in group 1, while there were no differences in other terminal genes between the two groups (Fig. 6C). More than one-third of the genome in group 1 contained the but gene, compared to less than 5% of the genome in group 2 (Fisher exact test, P < 0.001). Compared to group 2, group 1 also showed a higher trend toward heritability for acetate production (P = 0.06), but a lower heritability for propionate production (P < 0.05) (Fig. 6C). These results indicate that group 1 has a significantly higher heritability for utilizing complex plant polysaccharides and producing acetate and butyrate compared to group 2.

[0290] From a pathogenicity perspective, 21 of the 1,845 HQMAGs encode 750 virulence factor (VF) genes. Of these 21 VF-encoding HQMAGs, 3 belong to group 1 and 18 to group 2. Three of the 50 HQMAGs in group 1 have a VF gene involved in antiphagocytosis. In group 2, 18 of the 91 HQMAGs encode 747 VF genes across 15 different VF categories (i.e., acid tolerance, adhesion, antiphagocytosis, biofilm formation, efflux pumps, endotoxins, invasion, iron uptake, manganese uptake, motility, trophic factors, proteases, regulation, secretion systems, and toxins) (Figure 6C). Notably, 98.53% of all VF genes in bacterial community 2 were contained in eight HQMAGs (one in *Enterobacter kobei*, two in *Escherichia flexneri*, three in *Escherichia coli*, and two in *Klebsiella*). The HQMAGs of bacterial community 2 showed a high enrichment of virulence factor genes (Fisher exact test, P < 2.2 × 10⁻⁶). -16 This suggests that this bacterial community may play an important role in exacerbating metabolic disease phenotypes.

[0291] Regarding antibiotic resistance genes (ARGs), in communal 1, only one HQAMG (representing 2.00% of the communal genome) contains a copy of a phenol-associated ARG (Figure 6C). In communal 2, 17 HQMAGs (representing 18.68% of the communal genome) encode 40 ARGs that are resistant to seven different classes of antibiotics (i.e., aminoglycosides, β-lactams, fosfomycin, glycopeptides, quinolones, macrolides, and tetracyclines). Therefore, communal 2 could serve as a reservoir of ARGs for horizontal transfer to opportunistic pathogens. In conclusion, our data show that the two competing communes have different genetic capabilities, with communal 1 potentially beneficial and communal 2 potentially harmful.

[0292] 1.4 Two competing bacterial groups identified in the QD trial (QD-TCG) distinguished cases from controls in 10 independent case-control metagenomic datasets for 6 different diseases.

[0293] Since two competing microbial communities identified in the QD trial (QD-TCG) were found to be important microbiome signatures in response to T2DM dietary interventions, we investigated whether they could serve as biomarkers to distinguish between T2DM and controls. We addressed this in a separate T2DM metagenomic dataset comprising 136 T2DM cases and 136 controls, using all 141 HQMAGs from the QD-TCG as reference genomes. These were used for read recruitment analysis, a widely used method for estimating the abundance of reference genomes within metagenomics. On average, 35.28% and 32.92% of reads were recruited in case and control samples, respectively. Following this, we developed a machine learning classifier based on a random forest algorithm, utilizing the abundance of the 141 HQMAGs to determine whether we could distinguish between cases and controls. Receiver operating characteristic (ROC) curve analysis showed moderate diagnostic ability, with an area under the curve (AUC) of 0.70, determined by leave-one-out cross-validation.

[0294] We then hypothesized that QD-TCG might represent an inherent pattern in the human microbiome, independent of disease type. To assess this, we collected 10 independent case and control metagenomic datasets covering six different diseases: cirrhosis (LC), ankylosing spondylitis (AS), atherosclerotic cardiovascular disease (ACVD), schizophrenia (SCZ), colorectal cancer (CRC), and inflammatory bowel disease (IBD). We termed this collection of 10 datasets, along with the aforementioned T2DM dataset, Case-Control Dataset Set I (CCDC-I). On average, from these 10 metagenomic datasets, 32.12% and 31.34% of reads were recruited to the two microbiota in case and control samples, respectively. Figure 6DWithin each dataset, a random forest classifier using 141 HQMAGs demonstrated diagnostic ability to distinguish between cases and controls, with AUCs ranging from 0.68 for SCZ to 0.98 for AS#1 (Figure 6E).

[0295] These findings highlight the discriminative power of QD-TCG in distinguishing between cases and controls across various disease types. This demonstrates common microbiome characteristics associated with a broad spectrum of human diseases.

[0296] Materials and Methods

[0297] Clinical trials

[0298] Study Design: The QD trial was conducted at Qidong People's Hospital (Jiangsu Province, China) to examine the effects of a high-fiber diet under free-living conditions in a cohort of individuals clinically diagnosed with type 2 diabetes mellitus (T2DM) in Qidong. The study protocol was approved by the Ethics Committee of Shanghai General Hospital (2014KY104), and the study was conducted in accordance with the principles of the Declaration of Helsinki. All participants provided written informed consent. This trial is registered with the Chinese Clinical Trial Registry (ChiCTR-IPC-14005346).

[0299] This study recruits Han Chinese patients with type 2 diabetes mellitus (T2DM) (age: 37 to 70 years; HbA1c: 6.5% to 12.0%). A more detailed description of the inclusion and exclusion criteria is available at the Chinese Clinical Trial Registry, at www.chictr.org.cn.

[0300] Patients received either a high-fiber diet (WTP diet) as the treatment group (W group) or standard care (standard diet) as the control group (U group) for 3 months. Total calories and macronutrient prescriptions were based on the Chinese Dietary Reference Intakes for specific ages (Chinese Nutrition Society, 2013). The WTP diet consisted primarily of whole grains, traditional Chinese medicinal foods, and prebiotics, and included three ready-to-eat prepared foods. Standard care included standard dietary and exercise recommendations based on the Chinese Diabetes Association's T2DM guidelines. Patients in the W group received the WTP diet for a 3-month home-administered intervention, while patients in the U group received standard care. The WTP diet intervention was discontinued at the end of the third month (M3). Both W and U groups were then followed up for one year (M15). Nutrient intake was calculated according to the Chinese Food Composition 2009 using a meal-based food frequency questionnaire and 24-hour dietary recall. Both groups continued to take their prescribed antidiabetic medications.

[0301] Prior to the 2-week run-inperiod, all participants attended a lecture on diabetes intervention and improvement, and received diabetes education and metabolic assessment. 119 eligible individuals were enrolled based on inclusion and exclusion criteria and randomly assigned to two groups in a 2:1 ratio using SAS software (n=79 in the W group and n=40 in the U group).

[0302] Physical examinations were conducted at Qidong People's Hospital (Jiangsu Province, China) at times M0, M3, and M15. Sample collection instructions were provided to participants the day before. Participants provided stool and first morning urine as required. After collecting fasting venous blood samples, a 3-hour meal tolerance test (containing 75g of a Chinese steamed bun with available carbohydrates; MTT test) was performed, and postprandial venous blood samples were collected at 30, 60, 120, and 180 minutes. All blood samples were left at room temperature for 30 minutes to obtain serum, which was then centrifuged at 3000 rpm for 20 minutes at 4°C. The fasting serum was aliquoted, one for hospital testing and the other for laboratory testing. Stool, urine, and serum samples were immediately stored on dry ice and then transported to the laboratory and frozen at -80°C. Subsequently, anthropometric markers and the diabetic complication index were measured. Ewing tests and 24-hour Holter monitoring were performed to assess diabetic autonomic neuropathy (DAN). B-mode carotid ultrasound was performed to assess atherosclerosis. The Michigan Neuropathy Screening Scale was used to assess diabetic peripheral neuropathy (DPN). Additionally, meal-based food frequency questionnaires and 24-hour dietary reviews were recorded for nutrient intake calculations.

[0303] Fasting venous blood was used to measure HbA1c, fasting blood glucose, fasting insulin, fasting C-peptide, C-reactive protein (CRP), complete blood count, blood biochemistry, and five thyroid analytes. Postprandial blood glucose, insulin, and C-peptide were measured using venous blood samples taken at MTT 30, 60, 120, and 180 minutes. Morning fasting urine was used for urinalysis and to measure the urinary microalbumin-to-creatinine ratio. The above measurements were performed at Qidong People's Hospital. At Shanghai Jiao Tong University, fasting venous blood was used to quantify TNF-α (R&D Systems, MN, USA), lipopolysaccharide-binding protein (Hycult Biotech, PA, USA), leptin (P&C, PCDBH0287, China), and adiponectin (P&C, PCDBH0016, China) using enzyme-linked immunosorbent assay (ELISA).

[0304] Homeostasis model assessments of insulin resistance (HOMA-IR) and pancreatic β-cell function (HOMA-β) were calculated based on fasting blood glucose (mmol / L) and fasting C-peptide (pmol / L): HOMA-IR = 1.5 + FBG * fasting C-peptide / 2800; HOMA-β = 0.27 * fasting C-peptide / (FBG - 3.5). GFR (ml / min per 1.73 mcg) was calculated using the formula... 2 =186*Scr -1.154 *age -0.203 The glomerular filtration rate is estimated using *0.742 (for women) *1.233 (for Chinese), where Scr (serum creatinine) is expressed in mg / dl and age is expressed in years.

[0305] Gut microbiome analysis

[0306] Metagenomic sequencing. DNA was extracted from fecal samples using methods described previously. Metagenomic sequencing was performed using an Illumina HiSeq 3000 from Genewiz (Beijing, China). Primer clustering, template hybridization, isothermal amplification, linearization, blocking denaturation, and hybridization were all performed according to the service provider's specified workflow. The constructed library had an insert size of approximately 500 bp, which was then subjected to high-throughput sequencing to obtain 150 bp paired-end reads in both the forward and reverse directions.

[0307] Data quality control. Prinseq is used to: 1) trim reads from the 3' end until the first nucleotide of a quality threshold of 20 is reached; 2) delete read pairs when reads are <60 bp or contain "N" bases; and 3) remove duplicate reads. Reads that can be aligned to the human genome (Homo sapiens, UCSC hg19) are removed (using --reorder --no-hd --no-contain --dovetail for Bowtie2 alignment).

[0308] De novo genome assembly, abundance calculation, and classification assignment. De novo assembly was performed on each sample using IDBA_UD (--step 20 --mink20 --maxk 100 --min_contig 500 --pre_correction). The assembled contigs were further binned using MetaBAT (--minContig 1500 --superspecific -B 20). Bin quality was assessed using CheckM. Bins with >95% integrity, <5% contamination, and <5% strain heterogeneity were retained as high-quality draft genomes. The assembled high-quality draft genomes were further deduplicated using dRep. DiTASiC (which uses kallisto for pseudo-alignment and resolves shared reads between genomes using a generalized linear model) was used to calculate genome abundance in each sample, removing estimated counts with p-values ​​>0.05, and reducing all samples to 36 million reads (a sample with a read mapping ratio <25% was not well represented by a high-quality genome and was removed in downstream analysis). Genome classification and assignment are performed using GTDB-Tk with default parameters.

[0309] Gut microbiome network construction and analysis. In the W group, a co-abundance network was constructed for each time point using a universal genome shared by more than 75% of the samples at each time point. Fastspar (a fast and scalable correlation estimation tool for microbiome research) was used to calculate the correlations between genomes with 1,000 permutations at each time point based on the abundance across patient genomes, and correlations with p ≤ 0.001 were reserved for further analysis. The network was visualized using Cytoscape v3.8.1. A weighted Spring-embedded layout was used to determine the placement of nodes and edges using correlation coefficients as weights. Links between nodes were considered as metal springs attached to node pairs. Correlation coefficients were used to determine the repulsive and attractive forces of the springs. The layout algorithm set the positions of nodes to minimize the sum of forces in the network. Robust stable edges were defined as invariant positive / negative correlations between the same two genomes in all three networks at M0, M3, and M15. Nodes in the stable network were clustered using connected component clustering analysis in cystoscopy. To identify the presence of sub-clusters within the C1 cluster, the average linkage method was used, based on cluster trees with negative correlation (set to -1) and positive correlation (set to 1), followed by WGCNA analysis.

[0310] Case-Control Dataset Collection I (CCDC-I). Eleven independent metagenomic datasets for T2DM, LC, AS, ACVD, SCZ, CRC, and IBD were downloaded from the SRA or ENA databases. Group information, i.e., cases or controls, was collected from the corresponding papers or curated MetagenomicData tables. Quality control of raw reads was performed by KneadData (see URLbitbucket.org / biobakery / kneaddata). DiTASiC was used to recruit reads and estimate the abundance of 141 HQMAGs in the QD-CTG for each sample, removing estimated counts with p-values ​​> 0.05, and further converting to relative abundance divided by the total number of reads. Based on the estimated abundance of HQMAGs in each dataset, a random forest classification model was constructed for classifying cases and controls using leave-one-out cross-validation. High-quality reads for each sample were assembled de novo using MEGAHIT (-min-contig-len 500, -presets meta-large). The assembled contigs were binned using MetaBAT 2 and MaxBin 2. Bin quality was assessed using CheckM. Bins with >95% integrity, <5% contamination, and <5% strain heterogeneity were retained as HQMAGs and further deduplicated using dRep. In both the case and control groups, co-abundance correlations were calculated using Fastspar with 1,000 permutations based on HQMAGs shared by more than 75% of samples in both groups. Correlations with P ≤ 0.001 were retained for further analysis. Robust stable edges were defined as invariant positive / negative correlations between two identical genomes in both the case and control groups. The same clustering analyses performed in the QD trials were performed on stable networks. The abundance of HQMAGs in each TCG per sample was estimated using DiTASiC. A random forest classification model was constructed using leave-one-out cross-validation to classify cases and controls to test each TCG.

[0311] Case-Control Dataset Collection II (CCDC-II). Fifteen independent metagenomic datasets for AS, ASD, BD, COVID-19, CRC, GD, HT, MS, PC, and PD were downloaded from the SRA or ENA databases. Group information, i.e., cases or controls, was collected from corresponding papers or curatedMetagenomicData tables. Raw read quality control was performed using KneadData. DiTASiC was used to recruit reads and estimate the abundance of HQMAG in C-TCG and CC-TCG. Based on the estimated abundance of HQMAG in each dataset, a random forest classification model was constructed for classifying cases and controls using leave-one-out cross-validation.

[0312] Treatment Dataset Collection (TDC). Eleven independent metagenomic datasets of pre-treatment samples associated with four diseases (IBD, rheumatoid arthritis (RA), advanced melanoma, and B-cell lymphoma) were downloaded from the SRA or ENA databases. Respondent and non-responder categories for each sample were collected from the corresponding papers. Raw read quality control was performed using KneadData. DiTASiC was used to recruit reads and estimate the abundance of HQMAG in CC-TCG. Based on the estimated abundance of HQMAG in each dataset, a random forest classification model for predicting responders and non-responders was constructed using leave-one-out cross-validation.

[0313] Gut microbiome functional analysis. HQMAGs were annotated using Prokka. KEGG ortholog (KO) IDs were assigned to predicted protein sequences in each HQMAG using KofamKOALA with default parameters via HMMSEARCH against Kofam. KOs were further assigned to KEGG modules. Antibiotic resistance genes were predicted using ResFinder with default parameters. Virulence factor identification was based on the core set of the Virulence Factor Database (VFDB). Predicted protein sequences were aligned to reference sequences in the VFDB using BLASTP (best hits were E<1e-5, identity >80%, and query coverage >70%). Genes encoding carbohydrate-active enzymes (CAZy) were identified using dbCAN (version 6.0), with best hit alignments retained. As previously described, genes encoding formate-tetrahydrofolate ligase, propionyl-CoA:succinate-CoA transferase, propionate-CoA transferase, 4Hbt, AtoA, AtoD, Buk, and But were identified.

[0314] Statistical analysis.

[0315] Statistical analyses were performed in R (R version 3.6.1). Within-group comparisons were performed using the Friedman test, followed by the Nemenyi post-hoc test for repeated measures. The Mann-Whitney test (two-tailed) was used to compare W and U at the same time point. The Pearson chi-square test was performed to compare differences in categorical data between groups or between time points. The PERMANOVA test (9,999 permutations) was used to compare the gut microbiota structure of each group.

[0316] The Mann-Whitney test (two-tailed) and Fisher's exact test (two-tailed) were used to compare the target functions between community 1 and community 2. Hierarchical clustering analysis based on Jaccard distance on the KO map was performed to compare HQMAG in CC-TCG.

[0317] A linear mixed-effects model with subject ID as the random effect was applied to explore the association between bacterial abundance and clinical parameters in QD-TCG. For each HQMAG belonging to a bacterial community, the robust abundance after clr transformation in each sample was first range-scaled. Subsequently, the bacterial abundance was obtained as the average range-scaled abundance of HQMAGs belonging to that bacterial community. Time points M0 and M3 were used to train the linear mixed-effects model, and M15 was used for testing.

[0318] Classification analysis was performed using a random forest with leave-one-out cross-validation based on the HQMAG of the TCG in each dataset of CCDC-I, CCDC-II, and TDC. A general model was also performed using a random forest with 10x cross-validation based on the HQMAG in CC-TCG.

[0319] Example 2 – Combined core genomes (CC-TCG) from all two competing bacterial groups showed better performance in case-control classification across diseases.

[0320] 2.1 Disease, as an ecosystem disturbance, led to the identification of more groups of two competing bacterial communities with stable correlations in the independent case-control dataset.

[0321] To expand our understanding of two competing gut microbiota, QD-TCG, which are identified based on revealing stable genomic pairs despite dietary intervention, we focused our study on precisely locating stable genomic pairs in both cases and controls in a cross-sectional dataset, where disease progression is seen as a perturbation of the gut ecosystem. To this end, we first reconstructed a non-redundant HQMAG (QD-TCG) from the CCDC-I metagenomic dataset. Figure 7A In the case and control datasets for T2D, LC, AS, ACVD, SCZ, CRC, and IBD, we obtained a total of 682, 384, 564, 1308, 812, 1644, and 774 HQMAGs, respectively. For each disease, we constructed a co-abundance network for both the case and control groups based on the prevalence of HQMAGs shared by more than 75% of participants in both groups. We identified stable HQMAG pairs that showed positive / negative correlations in both the case and control groups and further grouped them using component clustering analysis. Notably, in the studies of T2D, LC, AS, ACVD, SCZ, CRC, and IBD, the prevalent HQMAGs with stable correlations were clustered into 2, 6, 10, 4, 10, 8, and 7 clusters.

[0322] In each disease type, we observed a dominant cluster that possessed the majority of the genomes in the stable network and exhibited complex interactions of both positive and negative correlations. To determine whether C1 contained subclusters, we constructed cluster trees using negative and positive correlations from the mean linkage method and applied WGCNA analysis. Each cluster C1 was then divided into two subclusters, C1A and C1B. Interestingly, each pair of C1A and C1B followed the pattern of two competing microbiomes we observed in QD-TCG. Most stable correlations between HQMAGs belonging to C1A and C1B showed positive correlations within each cluster and negative correlations between the two clusters. In the studies of T2D, LC, AS, ACVD, SCZ, CRC, and IBD, these correlations accounted for 76.88%, 95.31%, 100%, 96.23%, 96.43%, 96.72%, and 97.10% of the stable correlations within or between C1A and C1B, respectively. These results highlight the potential for detecting TCG in case-control datasets and demonstrate the presence of TCG in the gut microbiome of people across various diseases, ethnicities, and geographic regions.

[0323] Next, we investigated whether TCGs identified from CCDC-I could serve as biomarkers to distinguish between cases and controls. To this end, we used the corresponding HQMAG as a reference genome and performed read recruitment analysis on each TCG in the CCDC-I metagenomic dataset. On average, 54.58%, 28.32%, 26.93%, 63.62%, 44.54%, 55.72%, and 34.21% of reads from CCDC-I were recruited to TCGs identified in studies of T2D, LC, AS, ACVD, SCZ, CRC, and IBD, respectively. For each CCDC-I dataset, we then employed a random forest classifier, using the abundance of HQMAG in each TCG to distinguish between cases and controls. In the 11 datasets of CCDC-I, TCG from studies on T2D, LC, AS, ACVD, SCZ, CRC, and IBD achieved moderate to excellent diagnostic capabilities, classifying cases and controls in datasets 8, 9, 9, 9, 8, 7, and 7, with mean AUCs of 0.79, 0.78, 0.78, 0.77, 0.77, 0.77, and 0.74, respectively. Figure 7B These findings highlight the versatility of TCGs in classifying cases and controls across a wide range of disease types, suggesting that TCGs with stable correlations constitute common microbiome features associated with a broad spectrum of human diseases.

[0324] 2.2 The combined core genome (CC-TCG) from all two competing bacterial groups showed better performance in case and control classification across diseases.

[0325] In summary, we collected 925 HQMAGs from QD trials and CCDC-I, involving 8 different TCGs. After redundancy removal analysis based on a genomic ANI cutoff of >99%, these were merged into a pool of 788 non-redundant HQMAGs. This pool of 788 non-redundant HQMAGs is referred to here as the combined genome of two competing communities (C-TCGs), representing the confluence of the 8 TCG sets. Within the C-TCGs, 701 HQMAGs were specific to one of the 8 sets, while 87 were shared by multiple sets. Of the specific ones, 301 belonged to C1A, and 400 to C1B. Of the shared ones, 10 and 40 consistently belonged to C1A and C1B, respectively, while 37 showed inconsistent distributions across different TCGs. Figure 8A To determine whether C-TCG improves case-control classification across diseases, we performed read recruitment analysis on the CCDC-I metagenomic dataset using the corresponding HQMAG as the reference genome. C-TCG accounted for an average of 84.54% of the total abundance. Subsequently, a random forest classifier was trained on each CCDC-I dataset using the abundance of HQMAG in C-TCG. Overall, C-TCG demonstrated superior case-control classification ability in CCDC-I compared to TCG alone, with significantly higher AUC values ​​than classifiers trained on TCGs from T2D, AS, IBD, SCZ, and LC studies.

[0326] Next, we aimed to identify the C-TCG members most relevant to classification performance as the core genome. We ranked the 788 HQMAGs in the C-TCGs based on feature importance in the random forest model built for each CCDC-I dataset. Starting with the least important HQMAG, we sequentially removed one HQMAG and trained a new random forest model on each dataset, evaluating the AUC value used for classifying cases and controls. This yielded 788 different classifiers, each using a variable number of HQMAGs from 1 to 788 for each CCDC-I dataset. We assigned rankings based on the AUC value of their corresponding models using HQMAG numbers, with lower rankings indicating higher AUC values. Notably, the classifier trained on the first 302 HQMAGs showed the best classification performance in CCDC-I, as evidenced by the minimum cumulative ranking ( Figure 8AOf the 302 HQMAGs, 103 were specific to C1A, 181 were specific to C1B, and 18 showed inconsistent C1A and C1B assignments across different TCGs. After discarding the 18 inconsistent HQMAGs, we obtained a set of 284 HQMAGs that were not only most relevant to classification performance but also consistently assigned to the two competing microbiota. We termed these HQMAGs the combined core set of the two competing microbiota (CC-TCG). Overall, the random forest classifier based on CC-TCGs demonstrated superior performance in classifying cases and controls compared to both C-TCGs and individual TCGs from the QD trial and CCDC-I, with significantly higher AUC values ​​than classifiers trained on TCGs from the CRC, T2D, AS, IBD, SCZ, and LC studies.

[0327] To decipher the genetic basis of the association between these genomes and host health, we performed a genome-centric analysis of HQMAGs belonging to the CC-TCG. We first conducted targeted functional analyses and compared HQMAGs assigned to C1A and C1B. Similar to the findings with QD-TCG, C1A had a significantly higher copy number of genes involved in butyrate biosynthesis and a lower copy number of genes involved in propionate production. Regarding carbohydrate degradation genes, HQMAGs in C1A were enriched with CAZy genes for arabinoxylan and cellulose utilization. These findings suggest that HGMAGs from C1A have a higher genetic capacity for utilizing complex plant polysaccharides and producing butyrate compared to C1B. From the perspective of antibiotic resistance and pathogenicity, C1A had fewer ARGs and VFs than C1B.

[0328] Furthermore, we performed untargeted functional analysis based on the assignment of KEGG orthologs (KOs) to all predicted genes from 284 core HQMAGs. In total, we identified 3,553 and 5,495 KOs in C1A and C1B, respectively. Hierarchical clustering analysis based on KO profiles demonstrated significant functional differences between C1A and C1B (PERMANOVA, P = 0.001). KOs from C1A and C1B were further mapped to 253 and 291 KEGG modules, respectively. There were 250 shared modules between the two groups. C1A had 3 unique modules for acarbose biosynthesis, benzoate degradation, and staphylococcal ferritin B biosynthesis. C1B had 41 unique modules, including those for multidrug resistance, KDO2-lipid A modification, pathogenicity traits, and γ-aminobutyric acid (GABA) production. In conclusion, these results suggest that CC-TCGs possess different heritability, with C1A potentially beneficial and C1B potentially harmful.

[0329] 2.3 The combined core genome of two competing bacterial communities (CC-TCG) distinguishes cases from controls in other datasets.

[0330] To further validate CC-TCG's ability to distinguish between cases and controls across a range of diseases, we compiled 15 additional independent metagenomic datasets. These datasets include case and control data from 10 different diseases, including ankylosing spondylitis (AS), autism spectrum disorder (ASD), Behcet's disease (BD), COVID-19, colorectal cancer (CRC), Graves' disease (GD), hypertension (HT), multiple sclerosis (MS), pancreatic cancer (PC), and Parkinson's disease (PD). These 15 datasets are collectively referred to as Case-Control Dataset Collection II (CCDC-II). Using 284 HQMAGs from CC-TCG as reference genomes, we performed read recruitment analysis on the CCDC-II metagenomic datasets. On average, we found that 34.41% and 34.76% of reads were recruited in case and control samples, respectively. For each dataset in CCDC-II, we then used the abundance of the 284 HQMAGs from CC-TCG to train a random forest classifier to distinguish cases from controls. CC-TCG demonstrated moderate to excellent diagnostic capabilities in 10 out of 15 datasets, particularly those related to AS, ASD, COVID-19, CRC, GD, HT, MS, and PC, although it only achieved an AUC of 0.58 for HT#2, and AUCs between 0.6 and 0.7 for BD, PD, CRC#4, and CRC#5 datasets (Figure 8B).

[0331] For diseases such as CRC, IBD, and PC, which have more than two distinct datasets, we employed cross-dataset analysis (one dataset for model training and another for testing) and leave-one-out-of-dos (LODO) analysis to assess the general applicability of CC-TCG in diagnosing these diseases. For CRC, we found insufficient model portability from one dataset to another, similar to the microbiome characterization report in CRC by Thomas et al. However, in the cases of IBD and PC, classification models trained on a single CC-TCG dataset showed moderate to excellent performance when classifying cases and controls on other datasets. As done in the LODO analysis, pooling the training dataset improved the performance of CC-TCG-based models in classifying CRC cases and controls, with AUC values ​​ranging from 0.66 to 0.79. Figure 9A The LODO analysis yielded AUC values ​​ranging from 0.75 to 0.91 for IBD, while the AUC values ​​for PC ranged from 0.71 to 0.72. Figures 9B to 9CThese results demonstrate CC-TCG's classification capabilities in completely independent datasets.

[0332] Materials and Methods

[0333] Gut microbiome analysis

[0334] Metagenomic sequencing. DNA was extracted from fecal samples using methods described previously. Metagenomic sequencing was performed using an Illumina Hiseq 3000 from GENEBITDA (Beijing, China). Primer clustering, template hybridization, isothermal amplification, linearization, blocking denaturation, and hybridization were all performed according to the service provider's specified workflow. The constructed library had an insert size of approximately 500 bp, which was then subjected to high-throughput sequencing to obtain 150 bp paired-end reads in both the forward and reverse directions.

[0335] Data quality control. Prinseq is used to: 1) trim reads from the 3' end until the first nucleotide of a quality threshold of 20 is reached; 2) delete read pairs when reads are <60 bp or contain "N" bases; and 3) remove duplicate reads. Reads that can be aligned to the human genome (Homo sapiens, UCSC hg19) are removed (using --reorder --no-hd --no-contain --dovetail for Bowtie2 alignment).

[0336] De novo genome assembly, abundance calculation, and classification assignment. De novo assembly was performed on each sample using IDBA_UD (--step 20 --mink20 --maxk 100 --min_contig 500 --pre_correction). The assembled contigs were further binned using MetaBAT (--minContig 1500 --superspecific -B 20). Bin quality was assessed using CheckM. Bins with >95% integrity, <5% contamination, and <5% strain heterogeneity were retained as high-quality draft genomes. The assembled high-quality draft genomes were further deduplicated using dRep. DiTASiC (which applies kallisto for pseudo-alignment and resolves shared reads between genomes using a generalized linear model) was used to calculate genome abundance in each sample, removing estimated counts with p-values ​​>0.05, and reducing all samples to 36 million reads (a sample with a read mapping ratio <25% was not well represented by a high-quality genome and was removed in downstream analysis). Genome classification and assignment are performed using GTDB-Tk with default parameters.

[0337] Gut microbiome network construction and analysis. In the W group, a co-abundance network was constructed for each time point using a universal genome shared by more than 75% of the samples at each time point. Fastspar (a fast and scalable correlation estimation tool for microbiome research) was used to calculate the correlations between genomes with 1,000 permutations at each time point based on the abundance across patient genomes, and correlations with p ≤ 0.001 were reserved for further analysis. The network was visualized using Cytoscape v3.8.1. A weighted Spring-embedded layout was used to determine the placement of nodes and edges using correlation coefficients as weights. Links between nodes were considered as metal springs attached to node pairs. Correlation coefficients were used to determine the repulsive and attractive forces of the springs. The layout algorithm set the positions of nodes to minimize the sum of forces in the network. Robust stable edges were defined as invariant positive / negative correlations between the same two genomes in all three networks at M0, M3, and M15. Nodes in the stable network were clustered using connected component clustering analysis in cystoscopy. To identify the presence of sub-clusters within the C1 cluster, the average linkage method was used, based on cluster trees with negative correlation (set to -1) and positive correlation (set to 1), followed by WGCNA analysis.

[0338] Case-Control Dataset Collection I (CCDC-I). Eleven independent metagenomic datasets for T2DM, LC, AS, ACVD, SCZ, CRC, and IBD were downloaded from the SRA or ENA databases. Group information (cases or controls) was collected from corresponding papers or curated MetagenomicData tables. KneadData (…) https: / / bitbucket.org / biobakery / kneaddataQuality control of raw reads was performed. DiTASiC was used to recruit reads and estimate the abundance of 141 HQMAGs in the QD-CTG for each sample. Estimated counts with p-values ​​> 0.05 were removed, and the results were further converted to relative abundance divided by the total number of reads. Based on the estimated abundance of HQMAGs in each dataset, a random forest classification model for classifying cases and controls was constructed using leave-one-out cross-validation. High-quality reads for each sample were assembled de novo using MEGAHIT ((-min-contig-len 500, -presets meta-large)). The assembled contigs were binned using MetaBAT 2 and MaxBin 2. Bin quality was assessed using CheckM. Bins with >95% integrity, <5% contamination, and <5% strain heterogeneity were retained as HQMAGs and further deduplicated using dRep (Table 3). Co-abundance correlations were calculated in both case and control groups using Fastspar with 1,000 permutations based on HQMAGs shared by more than 75% of samples in both groups. Correlations with P ≤ 0.001 were retained for further analysis. Robust stable edges were defined as invariant positive / negative correlations between two identical genomes in both the case and control groups. The same clustering analyses performed in the QD trials were performed on stable networks. The abundance of HQMAGs in each TCG within each sample was estimated using DiTASiC. A random forest classification model was constructed using leave-one-out cross-validation to classify cases and controls to test each TCG.

[0339] Case-Control Dataset Collection II (CCDC-II). Fifteen independent metagenomic datasets for AS, ASD, BD, COVID-19, CRC, GD, HT, MS, PC, and PD were downloaded from the SRA or ENA databases. Group information, i.e., cases or controls, was collected from corresponding papers or compiled metagenomic data. Raw read quality control was performed using KneadData. DiTASiC was used to recruit reads and estimate the abundance of HQMAG in C-TCG and CC-TCG. Based on the estimated abundance of HQMAG in each dataset, a random forest classification model was constructed for classifying cases and controls using leave-one-out cross-validation.

[0340] Treatment Dataset Collection (TDC). Eleven independent metagenomic datasets of pre-treatment samples associated with four diseases (IBD, rheumatoid arthritis (RA), advanced melanoma, and B-cell lymphoma) were downloaded from the SRA or ENA databases. Respondent and non-responder categories for each sample were collected from the corresponding papers. Raw read quality control was performed using KneadData. DiTASiC was used to recruit reads and estimate the abundance of HQMAG in CC-TCG. Based on the estimated abundance of HQMAG in each dataset, a random forest classification model for predicting responders and non-responders was constructed using leave-one-out cross-validation.

[0341] Gut microbiome functional analysis. HQMAGs were annotated using Prokka. KEGG ortholog (KO) IDs were assigned to predicted protein sequences in each HQMAG using KofamKOALA with default parameters via HMMSEARCH against Kofam. KOs were further assigned to KEGG modules. Antibiotic resistance genes were predicted using ResFinder with default parameters. Virulence factor identification was based on the core set of the Virulence Factor Database (VFDB). Predicted protein sequences were aligned to reference sequences in the VFDB using BLASTP (best hits were E<1e-5, identity >80%, and query coverage >70%). Genes encoding carbohydrate-active enzymes (CAZy) were identified using dbCAN (version 6.0), with best hit alignments retained. As previously described, genes encoding formate-tetrahydrofolate ligase, propionyl-CoA:succinate-CoA transferase, propionate-CoA transferase, 4Hbt, AtoA, AtoD, Buk, and But were identified.

[0342] Statistical analysis.

[0343] Statistical analyses were performed in R (R version 3.6.1). Within-group comparisons were performed using the Friedman test, followed by the Nemenyi post-hoc test for repeated measures. The Mann-Whitney test (two-tailed) was used to compare W and U at the same time point. The Pearson chi-square test was performed to compare differences in categorical data between groups or between time points. The PERMANOVA test (9,999 permutations) was used to compare the gut microbiota structure of each group.

[0344] The Mann-Whitney test (two-tailed) and Fisher's exact test (two-tailed) were used to compare the target functions between community 1 and community 2. Hierarchical clustering analysis based on Jaccard distance on the KO map was performed to compare HQMAG in CC-TCG.

[0345] A linear mixed-effects model with subject ID as the random effect was applied to explore the association between bacterial abundance and clinical parameters in QD-TCG. For each HQMAG belonging to a bacterial community, the robust abundance after clr transformation in each sample was first range-scaled. Subsequently, the bacterial abundance was obtained as the average range-scaled abundance of HQMAGs belonging to that bacterial community. Time points M0 and M3 were used to train the linear mixed-effects model, and M15 was used for testing.

[0346] Classification analysis was performed using a random forest with leave-one-out cross-validation based on the HQMAG of the TCG in each dataset of CCDC-I, CCDC-II, and TDC. A general model was also performed using a random forest with 10x cross-validation based on the HQMAG in CC-TCG.

[0347] Example 3 – Combined Core Genome of Two Competing Microbes (CC-TCG) Predicts Immunotherapy Outcomes Across Various Independent Datasets Covering Multiple Diseases.

[0348] Previous studies have linked the gut microbiota to the efficacy of biotherapies in many diseases. We hypothesize that pre-treatment changes in the CC-TCG may be a predictor of clinical success in disease. To test our hypothesis, we compiled 11 pre-treatment metagenomic datasets associated with four diseases: IBD, rheumatoid arthritis (RA), advanced melanoma, and B-cell lymphoma, along with their corresponding categories indicating responders and non-responders to therapy. These 11 datasets are referred to as the Treatment Dataset Collection (TDC). We performed read recruitment analysis on the metagenomic datasets of the TDC using HQMAG from CC-TCG as the reference genome. On average, we found that 16.22% and 13.06% of reads were recruited in the responder and non-responder samples, respectively. For each dataset in the TDC, we constructed a random forest classifier based on the abundance of HQMAG in CC-TCG to predict treatment response (Figure 8C).

[0349] In three datasets for IBD treatment, patients were categorized as responders or non-responders to anti-cytokine or anti-integrin therapy based on 14-week remission. CC-TCG predicted remission at week 14 with an AUC of 0.68 for IBD_anti-cytokine, 0.64 for IBD_anti-integrin #1, and 0.69 for IBD_anti-integrin #2 (Figures 8C and 10A). | In the methotrexate treatment dataset for new-onset RA, responders to methotrexate were defined as any new-onset RA patient with improved disease activity scores across 28 joints. CC-TCG predicted improvement under methotrexate treatment with an AUC of 0.69 (Figures 8C and 10B). | A cross-cohort dataset of immune checkpoint inhibitor (ICI) therapy for advanced melanoma was collected in TDC by Lee et al. These cohorts were drawn from Barcelona (AM_ICI#1), Leeds (AM_ICI#2), Manchester (AM_ICI#3), PRIMM-UK (AM_ICI#4), and PRIMM-NL (AM_ICI#5). Responders and responders were defined based on overall response rate (ORR) or progression-free survival (PFS12) at 12 months. On average, CC-TCG predicted AUCs of 0.61 and 0.7 for ORR and PFS12 in each cohort, respectively (Figures 8C and 10C). | Insufficient portability of the predictive model was found from one cohort to another. However, when the training dataset was pooled, as done in the LODO analysis, the predictive performance improved, with mean AUCs of 0.66 and 0.67 for ORR and FFS12, respectively. | In the CD19-CAR-T immunotherapy dataset for B-cell lymphoma, responses to therapy were categorized as complete or incomplete remission at 180 days after CAR-T cell infusion. We trained the predictive model in a cohort from Germany (BCL_CD19-CAR-T#1) and validated it in a cohort from the United States (BCL_CD19-CAR-T#2). CC-TCG predicted response to CD19-CAR-T immunotherapy with an AUC of 0.66 in the German cohort, and the model was fully transferable to predict the US cohort with an AUC of 0.64. These results confirm the association between the pre-treatment gut microbiome and treatment efficacy and highlight the potential of using CC-TCG to predict treatment outcomes across a range of diseases.

[0350] Materials and Methods

[0351] Gut microbiome analysis

[0352] Metagenomic sequencing. DNA was extracted from fecal samples using methods described previously. Metagenomic sequencing was performed using an Illumina Hiseq 3000 from GENEBITDA (Beijing, China). Primer clustering, template hybridization, isothermal amplification, linearization, blocking denaturation, and hybridization were all performed according to the service provider's specified workflow. The constructed library had an insert size of approximately 500 bp, which was then subjected to high-throughput sequencing to obtain 150 bp paired-end reads in both the forward and reverse directions.

[0353] Data quality control. Prinseq is used to: 1) trim reads from the 3' end until the first nucleotide of a quality threshold of 20 is reached; 2) delete read pairs when reads are <60 bp or contain "N" bases; and 3) remove duplicate reads. Reads that can be aligned to the human genome (Homo sapiens, UCSC hg19) are removed (using --reorder --no-hd --no-contain --dovetail for Bowtie2 alignment).

[0354] De novo genome assembly, abundance calculation, and classification assignment. De novo assembly was performed on each sample using IDBA_UD (--step 20 --mink20 --maxk 100 --min_contig 500 --pre_correction). The assembled contigs were further binned using MetaBAT (--minContig 1500 --superspecific -B 20). Bin quality was assessed using CheckM. Bins with >95% integrity, <5% contamination, and <5% strain heterogeneity were retained as high-quality draft genomes. The assembled high-quality draft genomes were further deduplicated using dRep. DiTASiC (which uses kallisto for pseudo-alignment and resolves shared reads between genomes using a generalized linear model) was used to calculate genome abundance in each sample, removing estimated counts with p-values ​​>0.05, and reducing all samples to 36 million reads (a sample with a read mapping ratio <25% was not well represented by a high-quality genome and was removed in downstream analysis). Genome classification and assignment are performed using GTDB-Tk with default parameters.

[0355] Gut microbiome network construction and analysis. In the W group, a co-abundance network was constructed for each time point using a universal genome shared by more than 75% of the samples at each time point. Fastspar (a fast and scalable correlation estimation tool for microbiome research) was used to calculate the correlations between genomes with 1,000 permutations at each time point based on the abundance across patient genomes, and correlations with p ≤ 0.001 were reserved for further analysis. The network was visualized using Cytoscape v3.8.1. A weighted Spring-embedded layout was used to determine the placement of nodes and edges using correlation coefficients as weights. Links between nodes were considered as metal springs attached to node pairs. Correlation coefficients were used to determine the repulsive and attractive forces of the springs. The layout algorithm set the positions of nodes to minimize the sum of forces in the network. Robust stable edges were defined as invariant positive / negative correlations between the same two genomes in all three networks at M0, M3, and M15. Nodes in the stable network were clustered using connected component clustering analysis in cystoscopy. To identify the presence of sub-clusters within the C1 cluster, the average linkage method was used, based on cluster trees with negative correlation (set to -1) and positive correlation (set to 1), followed by WGCNA analysis.

[0356] Case-Control Dataset Collection I (CCDC-I). Eleven independent metagenomic datasets for T2DM, LC, AS, ACVD, SCZ, CRC, and IBD were downloaded from the SRA or ENA databases. Group information, i.e., cases or controls, was collected from corresponding papers or curatedMetagenomicData tables. Quality control of raw reads was performed using KneadData. DiTASiC was used to recruit reads and estimate the abundance of 141 HQMAGs in the QD-CTG in each sample, removing estimated counts with p-values ​​> 0.05, and further converting to relative abundance divided by the total number of reads. Based on the estimated abundance of HQMAGs in each dataset, a random forest classification model was constructed for classifying cases and controls using leave-one-out cross-validation. High-quality reads for each sample were assembled de novo using MEGAHIT ((-min-contig-len 500, -presets meta-large)). The assembled contigs were binned using MetaBAT2 and MaxBin 2. Bin quality was assessed using CheckM. Bins with >95% integrity, <5% contamination, and <5% strain heterogeneity were retained as HQMAGs and further deduplicated using dRep (Table 3). Co-abundance correlations were calculated in both the case and control groups using Fastspar with 1,000 permutations based on HQMAGs shared by more than 75% of samples in both groups. Correlations with P ≤ 0.001 were retained for further analysis. Robust stable edges were defined as invariant positive / negative correlations between two identical genomes in both the case and control groups. The same clustering analyses performed in the QD trials were performed on stable networks. The abundance of HQMAGs in each TCG within each sample was estimated using DiTASiC. A random forest classification model was constructed using leave-one-out cross-validation to classify cases and controls to test each TCG.

[0357] Case-Control Dataset Collection II (CCDC-II). Fifteen independent metagenomic datasets (Table 4) on AS, ASD, BD, COVID-19, CRC, GD, HT, MS, PC, and PD were downloaded from the SRA or ENA databases. Group information, i.e., cases or controls, was collected from corresponding papers or curatedMetagenomicData tables. Quality control of raw reads was performed using KneadData. DiTASiC was used to recruit reads and estimate the abundance of HQMAG in C-TCG and CC-TCG. Based on the estimated abundance of HQMAG in each dataset, a random forest classification model was constructed for classifying cases and controls using leave-one-out cross-validation.

[0358] Treatment Dataset Collection (TDC). Eleven independent metagenomic datasets of pre-treatment samples associated with four diseases (IBD, rheumatoid arthritis (RA), advanced melanoma, and B-cell lymphoma) were downloaded from the SRA or ENA databases. Respondent and non-responder categories for each sample were collected from the corresponding papers. Raw read quality control was performed using KneadData. DiTASiC was used to recruit reads and estimate the abundance of HQMAG in CC-TCG. Based on the estimated abundance of HQMAG in each dataset, a random forest classification model for predicting responders and non-responders was constructed using leave-one-out cross-validation.

[0359] Gut microbiome functional analysis. HQMAGs were annotated using Prokka. KEGG ortholog (KO) IDs were assigned to predicted protein sequences in each HQMAG using KofamKOALA with default parameters via HMMSEARCH against Kofam. KOs were further assigned to KEGG modules. Antibiotic resistance genes were predicted using ResFinder with default parameters. Virulence factor identification was based on the core set of the Virulence Factor Database (VFDB). Predicted protein sequences were aligned to reference sequences in the VFDB using BLASTP (best hits were E<1e-5, identity >80%, and query coverage >70%). Genes encoding carbohydrate-active enzymes (CAZy) were identified using dbCAN (version 6.0), with best hit alignments retained. As previously described, genes encoding formate-tetrahydrofolate ligase, propionyl-CoA:succinate-CoA transferase, propionate-CoA transferase, 4Hbt, AtoA, AtoD, Buk, and But were identified.

[0360] Statistical analysis.

[0361] Statistical analyses were performed in R (R version 3.6.1). Within-group comparisons were performed using the Friedman test, followed by the Nemenyi post-hoc test for repeated measures. The Mann-Whitney test (two-tailed) was used to compare W and U at the same time point. The Pearson chi-square test was performed to compare differences in categorical data between groups or between time points. The PERMANOVA test (9,999 permutations) was used to compare the gut microbiota structure of each group.

[0362] The Mann-Whitney test (two-tailed) and Fisher's exact test (two-tailed) were used to compare the target functions between community 1 and community 2. Hierarchical clustering analysis based on Jaccard distance on the KO map was performed to compare HQMAG in CC-TCG.

[0363] A linear mixed-effects model with subject ID as the random effect was applied to explore the association between bacterial abundance and clinical parameters in QD-TCG. For each HQMAG belonging to a bacterial community, the robust abundance after clr transformation in each sample was first range-scaled. Subsequently, the bacterial abundance was obtained as the average range-scaled abundance of HQMAGs belonging to that bacterial community. Time points M0 and M3 were used to train the linear mixed-effects model, and M15 was used for testing.

[0364] Classification analysis was performed using a random forest with leave-one-out cross-validation based on the HQMAG of the TCG in each dataset of CCDC-I, CCDC-II, and TDC. A general model was also performed using a random forest with 10x cross-validation based on the HQMAG in CC-TCG.

[0365] Example 4 – A general model based on a combined core genome of two competing bacterial communities to distinguish between cases and controls of different diseases.

[0366] We demonstrate the power of CC-TCG in distinguishing cases from controls and predicting treatment response. This is confirmed by intermediate to excellent random forest models built for each dataset in CCDC-I, CCDC-II, and TDC. Further validation of the robustness and portability of the CC-TCG-trained models was achieved through cross-dataset analysis of diseases with multiple datasets and LODO analysis.

[0367] Our findings prompted us to explore whether a general CC-TCG-based model could distinguish cases from controls, regardless of disease type. To this end, we merged all 26 datasets from CCDC-I and CCDC-II, covering 1,780 cases and 1,604 controls across 15 disease types. HQMAG was very prevalent in CC-TCG, with 210 cases present in over 80% of the samples and 101 cases present in over 90% of the samples.

[0368] A consistent CC-TCG trend was observed in 20 out of 26 datasets, with control samples exhibiting greater diversity and a higher C1A to C1B ratio in CC-TCG compared to case samples. We randomly assigned all case and control samples, with 80% used to train a CC-TCG-based random forest classifier and the remaining 20% ​​reserved for testing. Figure 11A ).

[0369] By employing 10-fold cross-validation, we found that the general classifier exhibited a moderate ability, with an AUC of 0.73, to distinguish between cases and controls (Figure 11B). The model generated probability scores indicating the likelihood of a sample being classified as a case, which showed a unimodal distribution in both case and control samples, with clear separation between the peaks (Figure 11C). The probability scores in controls were significantly lower than those in cases (Figure 11D). When the general classifier was applied to the test data, an AUC of 0.76 was achieved in distinguishing between cases and controls (Figure 11B). A similar distribution of probability scores for cases and controls was observed in the test data, parallel to the training data (Figure 11C). Similarly, the probability scores for controls were significantly lower than those for cases in the test data. For both the training and test data, a probability score of 0.52 provided a classification specificity of 0.7 and a sensitivity of 0.7 to distinguish between cases and controls. These results demonstrate the feasibility of using CC-TCG with a general model to distinguish between cases and controls, which is agnostic to disease type. In other words, changes in CC-TCG can serve as a general indicator of health recovery and maintenance.

[0370] discuss

[0371] Our research reveals the existence of a core gut microbiome structure comprised of two antagonistic bacterial communities, manifesting as a seesaw-like network in humans. This discovery is attributed to a unique combination of methodologies: genome-centric, reference-free, and interaction-focused approaches. These findings reveal strong associations between this core microbiome signature and diverse host phenotypes, particularly in individuals with type 2 diabetes mellitus (T2DM). Furthermore, our random forest model demonstrates that these bacterial genomes can help differentiate between cases and controls across multiple diseases and predict immunotherapy outcomes, highlighting the relevance of this core microbiome signature across different ethnicities, geographic locations, and disease states.

[0372] Several pieces of evidence support characterizing these two interconnected microbial communities as core gut microbiome features. First, their presence is consistent across populations regardless of race and geographic location. Second, they exhibit remarkable temporal stability in terms of membership and interaction patterns. Third, although they comprise approximately 10% of the gut microbiome, they exert a significant influence on the ecological community due to their highly interconnected and stable role within the gut ecosystem. Finally, the structure of these communities may be the product of natural selection over a long history of co-evolution between the microbes and their hosts, and may be regulated by dietary fiber, a direct external energy source for the gut ecosystem.

[0373] Until about 150 years ago, historical dietary trends in humans favored high fiber intake, supporting this view. High dietary fiber may provide an evolutionary advantage to beneficial bacteria in Group 1, as they have a stronger ability to degrade plant polysaccharides, thus giving them an advantage over pathogenic bacteria in Group 2. It is noteworthy that this seesaw-like network is not solely associated with T2DM, nor is it entirely dependent on a high-fiber diet. It can be detected in other independent metagenomic datasets, showing associations with a variety of diseases, and may be a fundamental property of the human microbiome.

[0374] Members of Community 1 possess an exceptional ability to degrade complex plant polysaccharides, producing beneficial metabolites such as short-chain fatty acids (SCFAs). These metabolites may inhibit the overgrowth of pathogenic bacteria in Community 2, whose unregulated proliferation can harm host health through mechanisms such as inflammation. However, maintaining a certain number of these pathogenic bacteria is essential because they play a crucial role in activating our immune system early in life. Community 1, analogous to the tall trees that form the foundational species in a dense forest, can be considered the "basal microbiota," constructing and stabilizing the gut environment unfavorable to Community 2 (pathogenic bacteria). Therefore, maintaining the delicate balance between the basal and pathogenic microbiota is crucial for determining whether the gut microbiome promotes health or induces disease.

[0375] The identified networks are characterized by cooperative and competitive interactions. While cooperation improves overall metabolic efficiency, it can lead to dependence and instability, which can be mitigated by introducing competition into the network. Interestingly, this seesaw-like network remains stable, but the relative abundance of basal and pathogenic microbiota may fluctuate, suggesting the role of external energy inputs, particularly dietary fiber. Understanding how external energy inputs affect the balance between order and disorder in complex adaptive systems is crucial for comprehending the ecological dynamics within the gut microbiome. “Order” (stable and predictable interactions) can be achieved through sufficient dietary fiber input by competing microbiota as the backbone of the gut ecosystem, thereby enhancing the dominance of basal microbiota relative to pathogenic microbiota. Conversely, “disorder” (instability and potential systemic disruption) can occur if energy input into the gut ecosystem shifts from external dietary fiber to host-produced mucins. These dynamics are permanently offset by energy inputs, as demonstrated by the identified seesaw-like network.

[0376] Our analysis builds on a relationship-centric approach, arguing that stable relationships within the microbiome help identify its core components. This approach involves studying the interactions and symbiotic patterns among various microbial members and interpreting these relationships as indicators of ecological function. Using this method, the seesaw-like network of two competing bacterial communities becomes a powerful feature of the gut microbiome, demonstrating stability across different ethnicities, geographic locations, and disease states. This finding highlights the importance of relationship-based analysis in microbiome research, showcasing the significance of cooperative and competitive dynamics within the gut microbiome. It also describes how external energy inputs, such as dietary fiber, can modulate interaction dynamics to maintain a balance between order (stability) and disorder (instability) within the microbiome.

[0377] In summary, our findings suggest that a seesaw-like network may be a fundamental characteristic of the human gut microbiome. We emphasize the crucial role of dietary fiber as an external energy input, essential for maintaining gut microbiome order to benefit host health. This characteristic may have been established through natural selection during a long co-evolutionary history between the microbiome and its host in high-fiber diets, highlighting the complex interactions of these principles within the human gut microbiome.

[0378] Our relationship-centric approach, which enhances our understanding of gut microbiome dynamics, can guide the development of interventions targeting this core feature. The ultimate goal is to restore and maintain the dominance of the basal microbiota over pathogenic microbiota, thereby promoting and protecting human health. Our research underscores the importance of recognizing stable relationships as key indicators of microbial components and their ecological roles, providing a promising framework for future microbiome research. Further research is needed to fully understand these complex dynamics and unlock the enormous potential of microbiome-centric therapies.

[0379] Example 5 – Integrated TCG prediction of immunotherapy response in independent treatment datasets.

[0380] Previous reports have linked the gut microbiota to the efficacy of biotherapies in a variety of diseases. It is hypothesized that pre-treatment changes in CC-TCG can predict the clinical success of these treatments. To test this hypothesis, 11 pre-treatment metagenomic datasets associated with four diseases were compiled: IBD (Lee, JWJ et al., Cell Host Microbe 29:1294-1304.e4 (2021) and Ananthakrishnan, AN et al., Cell Host Microbe 21:603-10.e3 (2017)), rheumatoid arthritis (RA) (Artacho, A. et al., Arthritis Rheumatol.73:931-42 (2021)), advanced melanoma (Lee, KA et al., Nat. Med.28, 535-44 (2022)), and B-cell lymphoma (Stein-Thoeringer, CK et al., Nat. Med.29:906-16 (2023)). These datasets included responder and non-responder categories and were collectively referred to as the Treatment Dataset Collection (TDC) (Figure 14). Read recruitment analysis was performed on the TDC metagenomic dataset using 284 HQMAGs from CC-TCG as reference genomes. On average, 16.22% and 13.06% of reads were recruited in the responder and non-responder samples, respectively. For each dataset in TDC, a random forest classifier was constructed based on the abundance of the 284 HQMAGs from CC-TCG to predict treatment response (Figure 14).

[0381] In three IBD treatment datasets, CC-TCG predicted remission at week 14, with an AUC of 0.68 for IBD_anti-cytokine, 0.64 for IBD_anti-integrin #1, and 0.69 for IBD_anti-integrin #2. Figure 14B and Figure 15A In the methotrexate treatment dataset for newly diagnosed RA, CC-TCG predicted improvement, with an AUC of 0.69 (). Figure 14B and Figure 15B ).

[0382] The immune checkpoint inhibitor (ICI) treatment dataset for advanced melanoma includes multiple cohorts, including Barcelona (AM_ICI#1), Leeds (AM_ICI#2), Manchester (AM_ICI#3), PRIMM-UK (AM_ICI#4), and PRIMM-NL (AM_ICI#5). On average, CC-TCG predicted overall response rate (ORR) and 12-month progression-free survival (FPS12), with AUC values ​​of 0.61 and 0.7 in each cohort, respectively. Figure 14B and Figure 15C The predictive model lacks portability from one cohort to another. However, predictive performance improves when the training dataset is pooled, as in the LODO analysis, with mean AUC values ​​of 0.66 for ORR and 0.67 for FPS12. In the CD19-CAR-T immunotherapy dataset for B-cell lymphoma, the predictive model was trained for a cohort from Germany (BCL_CD19-CAR-T#1) and validated for a cohort from the United States (BCL_CD19-CAR-T#2). CC-TCG predicted response to CD19-CAR-T immunotherapy in the German cohort with an AUC of 0.66, and the model was sufficiently portable to predict response in the US cohort with an AUC of 0.64. Figure 14A and Figure 15D These results highlight the potential of using CC-TCG to predict treatment outcomes across a wide range of diseases and therapies.

[0383] Example 6 – A dichotomy of conserved health-related functions of the genome in the TCG.

[0384] Using HQMAGs allows for the annotation and understanding of health-related functionality between two bacterial communities within a TCG. With a genomic ANI cutoff of 99%, HQMAGs from eight different TCGs identified in QD and CCDC-I showed limited overlap, with 701 out of 788 genomes in the C-TCG being specific to only one TCG. Simultaneously, only 87 appeared in more than one TCG. While this reveals differences in genomic composition among HQMAGs from different TCGs derived from various study datasets, we observed similar functional profiles within C1A and C1B, as well as functional dichotomies between the two communities across the eight TCGs, through non-targeted functional annotation of KEGG orthologs (KOs) in each community. Pearson correlation analysis showed a strong correlation of KO profiles within C1A (r = 0.89 ± 0.071, mean ± SD) and C1B (r = 0.96 ± 0.023). The correlation between C1A and C1B was lower (r = 0.76 ± 0.093). Principal coordinate analysis and cluster analysis of Euclidean distances calculated from the KO profiles further illustrated the significant differences between C1A and C1B from different TCGs (Figure 16), as well as a significant PERMANOVA test (p=0.002). Of the 5,870 identified KOs, 2,473 showed significant copy number differences between the two communities, with 354 higher in C1A and 2,119 higher in C1B. Therefore, although HQMAGs from different TCGs have very little overlap, the functions of each community within different TCGs are highly consistent.

[0385] To understand why C1A in QD-TCG is promoted by fiber and negatively correlated with HbA1c, genes encoding carbohydrate-active enzymes (CAZy) and key enzymes in short-chain fatty acid (SCFA) production were identified to compare the genetic capacity for carbohydrate utilization between the two populations. In QD-TCG, compared to the genomes in C1B, the genomes in C1A were richer in CAZy genes for arabinoxylan (p<0.001) and cellulose (p<0.001), and the proportion of CAZy genes utilized by inulin was lower (p<0.01). Figure 17A The two communities showed no differences in genes related to starch, pectin, and mucin utilization. Previous work has shown that gut microbiota benefit patients with type 2 disease by producing acetate and butyrate via carbohydrate fermentation. In the QD-TCG, among the terminal genes of butyrate biosynthesis pathways from both carbohydrates (e.g., but and buk) and proteins (e.g., atoA / D and 4Hbt), the copy number of but was significantly higher in C1A, while there were no differences in other terminal genes between the two communities. Figure...

Claims

1. A method for training a model to predict a subject's response to a therapy for a condition, comprising: In a computer system having one or more processors and memory storing one or more programs to be executed by the one or more processors: A) For each of the multiple training subjects, obtain electronically the following information, wherein each of the multiple training subjects has received treatment for the condition: (i) Prior to receiving the therapy, for the corresponding training subject, a plurality of corresponding genome abundance values, wherein the plurality of corresponding genome abundance values ​​include, for each of a plurality of gut microbiota, a corresponding value for the abundance of the genome of the corresponding gut microbiota in a corresponding biological sample from the gut of the corresponding training subject, and (ii) An indication of the response of the corresponding training subject to the therapy of the corresponding training subject; B) For each of the plurality of training subjects, information about the corresponding training subject is input into a model comprising multiple parameters, wherein the model applies the multiple parameters to the information through at least 10,000 calculations to obtain a corresponding output from the model for the corresponding training subject, wherein: The corresponding output includes a prediction of the corresponding training subject's response to the therapy. The information regarding the corresponding training subjects includes the corresponding genomic abundance value for each of the multiple gut microbiota, and The various gut microbiota mentioned are selected from Table 1, Table 2 or Figure 13A to Figure 13XX; as well as C) Adjust the plurality of parameters based on one or more differences between (i) the corresponding output from the model and (ii) the corresponding indication of the corresponding training subject’s response to the therapy for each of the plurality of training subjects.

2. The method of claim 1, wherein obtaining A) comprises, for each corresponding training subject among the plurality of training subjects: (i) electronically acquiring at least 100,000 corresponding nucleic acid sequences for genomic DNA from the corresponding biological sample of the gut of the corresponding training subject; and (ii) For each of the multiple gut microbes, the corresponding value of the abundance of the genome of the corresponding gut microbe is determined from the corresponding plurality of at least 100,000 nucleic acid sequences.

3. The method of claim 2, wherein determining A)(ii) comprises, for each corresponding training subject among the plurality of training subjects: Multiple gut microbial genomes are assembled electronically from at least 100,000 corresponding nucleic acid sequences using metagenomic de novo sequencing. For each of the plurality of gut microbes, the corresponding value of the abundance of the genome of the corresponding gut microbe is calculated based on the prevalence of the corresponding nucleic acid sequence among the plurality of at least 100,000 nucleic acid sequences used to assemble the corresponding gut microbe genome among the plurality of gut microbe genomes corresponding to the corresponding gut microbe.

4. The method of claim 2, wherein determining A)(ii) comprises, for each corresponding subject among the plurality of training subjects: Each corresponding nucleic acid sequence from the plurality of at least 100,000 sequences is assigned to a corresponding gut microbe among the plurality of gut microbes, thereby generating a corresponding count of the corresponding nucleic acid sequence among the plurality of nucleic acid sequences assigned to the corresponding gut microbe for each corresponding gut microbe among the plurality of gut microbes, and For each of the multiple gut microbes, the corresponding genomic abundance value of the corresponding gut microbe is determined based on the corresponding count of the corresponding nucleic acid sequence assigned to the corresponding gut microbe.

5. The method according to any one of claims 2 to 4, further comprising, for each of the plurality of training subjects, sequencing the genomic DNA of the corresponding biological sample from the gut of the corresponding training subject to obtain at least 100,000 nucleic acid sequences of the corresponding plurality.

6. The method according to any one of claims 1 to 5, wherein the plurality of gut microbiota comprises at least 20 gut microbiota selected from Table 1, Table 2 or Figures 13A to 13XX.

7. The method according to any one of claims 1 to 5, wherein the plurality of gut microbiota comprises at least 50, at least 100, at least 150, at least 200, or at least 250 gut microbiota selected from Table 2.

8. The method according to any one of claims 1 to 5, wherein the plurality of gut microbiota comprises all gut microbiota from Table 2.

9. The method according to any one of claims 1 to 8, wherein the plurality of gut microbiota comprises at least 20 microorganisms, the at least 20 microorganisms being selected from those microorganisms having a connectivity of at least 2 in Table 1, Table 2 or Figures 13A to 13XX.

10. The method according to any one of claims 1 to 9, wherein for each of the plurality of training subjects, the biological sample from the intestine of the respective subject is a fecal sample from the respective training subject.

11. The method according to any one of claims 1 to 10, wherein the therapy is a biological therapy, immunotherapy, chemotherapy, radiotherapy, gene therapy, hormone therapy, photodynamic therapy, targeted therapy, small molecule, antibody, polynucleotide, natural compound, immunomodulator, bone marrow therapy, stem cell therapy, surgical therapy, induction therapy, maintenance therapy, or a combination thereof.

12. The method of claim 1, wherein the condition is selected from the group consisting of: type 2 diabetes, hypertension, schizophrenia, atherosclerotic cardiovascular disease (ACVD), cirrhosis (LC), inflammatory bowel disease (IBD), colorectal cancer (CRC), ankylosing spondylitis (AS) and Parkinson's disease (PD), inflammatory bowel disease (IBD), rheumatoid arthritis (RA) or advanced melanoma and B-cell lymphoma.

13. The method of claim 1, wherein the disease is cancer.

14. The method according to any one of claims 1 to 13, wherein the prediction of the response of the corresponding training subject is a category output of the corresponding response among a plurality of possible responses of the corresponding training subject.

15. The method according to any one of claims 1 to 13, wherein the prediction of the response of the corresponding training subject is a probability output for the response of the corresponding training subject.

16. The method according to any one of claims 1 to 15, wherein the model is a neural network algorithm, a support vector machine algorithm, a Naive Bayes algorithm, a nearest neighbor algorithm, a boosting tree algorithm, a random forest algorithm, a convolutional neural network algorithm, a decision tree algorithm, a regression algorithm, or a clustering algorithm.

17. The method according to any one of claims 1 to 16, wherein the plurality of parameters is at least 1,000, at least 10,000, at least 15,000, at least 50,000, at least 100,000, at least 250,000, at least 500,000, or at least 1,000,000 parameters.

18. The method according to any one of claims 1 to 17, wherein the model applies the plurality of parameters to the information through at least 25,000, at least 50,000, at least 100,000, at least 250,000, at least 500,000, or at least 1,000,000 calculations to obtain corresponding outputs from the model for the respective training subjects.

19. A method for predicting a subject's response to a therapy for a condition, comprising: In a computer system having one or more processors and memory storing one or more programs to be executed by the one or more processors: A) Obtain multiple genome abundance values ​​in electronic form, the multiple genome abundance values ​​including the corresponding genome abundance value of the corresponding gut bacterial species in the multiple gut microbiota selected from Table 1, Table 2 or Figure 13A to Figure 13XX in the biological sample of the subject. as well as B) Input the plurality of genomic abundance values ​​into a model that includes a plurality of parameters, wherein the model applies the plurality of parameters to the plurality of genomic abundance values ​​through at least 10,000 calculations to generate a prediction of the subject’s response to the therapy as an output from the model.

20. The method of claim 19, wherein obtaining A) comprises: (i) Obtain in electronic form at least 100,000 nucleic acid sequences targeting genomic DNA from a biological sample of the subject’s gut; as well as (ii) For each of the plurality of gut microbes, a corresponding value for the abundance of the genome of the corresponding gut microbe is determined from the plurality of at least 100,000 nucleic acid sequences.

21. The method of claim 20, wherein determining A)(ii) comprises: Multiple gut microbial genomes were assembled electronically from at least 100,000 nucleic acid sequences using metagenomic de novo sequencing. For each of the plurality of gut microbes, the corresponding value of the abundance of the genome of the corresponding gut microbe is calculated based on the prevalence of the corresponding nucleic acid sequence among the plurality of at least 100,000 nucleic acid sequences used to assemble the corresponding gut microbe genome among the plurality of gut microbe genomes corresponding to the corresponding gut microbe.

22. The method of claim 20, wherein determining A)(ii) comprises: Each corresponding nucleic acid sequence from the plurality of at least 100,000 sequences is assigned to a corresponding gut microbe from the plurality of gut microbes, thereby generating a corresponding count of the corresponding nucleic acid sequence from the plurality of nucleic acid sequences assigned to the corresponding gut microbe for each corresponding gut microbe from the plurality of gut microbes, and For each of the multiple gut microbes, the corresponding genomic abundance value of the corresponding gut microbe is determined based on the corresponding count of the corresponding nucleic acid sequence assigned to the corresponding gut microbe.

23. The method according to any one of claims 20 to 22, further comprising sequencing the genomic DNA of the biological sample from the intestine of the subject to obtain the plurality of at least 100,000 nucleic acid sequences.

24. The method according to any one of claims 19 to 23, further comprising treating the subject by: The therapy is administered to the subject when the predicted response to the therapy meets a threshold probability that the subject will have a favorable response to the therapy; and When the predicted response of the subject to the therapy does not meet the threshold probability that the subject will have a favorable response to the therapy, one or more of the multiple gut microbiota are administered to the subject.

25. The method according to any one of claims 19 to 24, wherein the plurality of gut microbiota comprises at least 20 gut microbiota selected from Table 1, Table 2 or Figures 13A to 13XX.

26. The method according to any one of claims 19 to 24, wherein the plurality of gut microbiota comprises at least 50, at least 100, at least 150, at least 200, or at least 250 gut microbiota selected from Table 2.

27. The method according to any one of claims 19 to 24, wherein the plurality of gut microbiota comprises all gut microbiota from Table 2.

28. The method according to any one of claims 19 to 27, wherein the plurality of gut microbiota comprises at least 20 microorganisms selected from those microorganisms having a connectivity of at least 2 in Table 1, Table 2 or Figures 13A to 13XX.

29. The method according to any one of claims 19 to 28, wherein the biological sample from the intestine of the subject is a fecal sample.

30. The method according to any one of claims 19 to 29, wherein the therapy is a biological therapy, immunotherapy, chemotherapy, radiotherapy, gene therapy, hormone therapy, photodynamic therapy, targeted therapy, small molecule, antibody, polynucleotide, natural compound, immunomodulator, bone marrow therapy, stem cell therapy, surgical therapy, induction therapy, maintenance therapy, or a combination thereof.

31. The method according to claims 19 to 30, wherein the condition is selected from the group consisting of: type 2 diabetes, hypertension, schizophrenia, atherosclerotic cardiovascular disease (ACVD), cirrhosis (LC), inflammatory bowel disease (IBD), colorectal cancer (CRC), ankylosing spondylitis (AS) and Parkinson's disease (PD), inflammatory bowel disease (IBD), rheumatoid arthritis (RA), advanced melanoma and B-cell lymphoma.

32. The method according to claims 19 to 30, wherein the condition is cancer.

33. The method according to any one of claims 19 to 29, wherein the condition is inflammatory bowel disease and the therapy comprises anti-cytokine or anti-integrin treatment.

34. The method according to any one of claims 19 to 29, wherein the condition is rheumatoid arthritis and the treatment includes methotrexate treatment.

35. The method according to any one of claims 19 to 29, wherein the condition is B-cell lymphoma and the therapy comprises CAR-T cell immunotherapy.

36. The method according to any one of claims 19 to 29, wherein the condition is melanoma and the therapy comprises immune checkpoint inhibitor treatment.

37. The method of any one of claims 19 to 36, wherein the prediction of the subject's response is a category output of the corresponding response among a plurality of possible responses of the corresponding subject.

38. The method according to any one of claims 19 to 36, wherein the prediction of the subject's response is a probability output for the corresponding subject's response.

39. The method according to any one of claims 19 to 38, wherein the model is a neural network algorithm, a support vector machine algorithm, a naive Bayes algorithm, a nearest neighbor algorithm, a boosting tree algorithm, a random forest algorithm, a convolutional neural network algorithm, a decision tree algorithm, a regression algorithm, or a clustering algorithm.

40. The method according to any one of claims 19 to 39, wherein the plurality of parameters is at least 1,000, at least 10,000, at least 15,000, at least 50,000, at least 100,000, at least 250,000, at least 500,000, or at least 1,000,000 parameters.

41. The method according to any one of claims 19 to 40, wherein the model applies the plurality of parameters to the information through at least 25,000, at least 50,000, at least 100,000, at least 250,000, at least 500,000, or at least 1,000,000 calculations to obtain a corresponding output from the model for the respective subject.

42. A computer system comprising: One or more processors; as well as A non-transitory computer-readable medium comprising computer-executable instructions, which, when executed by the one or more processors, cause the processors to perform the method according to any one of claims 1 to 41.

43. A non-transitory computer-readable storage medium having program code instructions stored thereon, the program code instructions, when executed by a processor, causing the processor to perform the method according to any one of claims 1 to 41.

Citation Information

Patent Citations

  • Methods for genome assembly and haplotype phasing

    US10529443B2

  • Microbiome based systems, apparatus and methods for monitoring and controlling industrial processes and systems

    US11028449B2

  • Sample analysis, presence determination of a target sequence

    US11332783B2

  • Absolute quantification of nucleic acids and related methods and systems

    US11427865B2

  • Metagenomic library and natural product discovery platform

    US11495326B2