Method and system for microbiota subspecies characterization, identification, diagnostics, and therapy

The method and system for microbiota subspecies analysis using machine learning and deep learning provide a subspecies-level catalog of human gut microbiota, addressing the limitations of species-level methods by enhancing disease diagnosis and therapeutic prediction through improved strain-level understanding.

WO2026039715A1PCT designated stage Publication Date: 2026-02-19RES DEVMENT FOUND
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/042133
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-15
Filing Date
2025-08-15
Publication Date
2026-02-19

AI Technical Summary

Technical Problem

Current microbiome analysis methods are limited to species level, obscuring intraspecies variations and hindering the establishment of causal links between the microbiome and health due to strain-level differences, which are host-specific and can differ by a single nucleotide, making it difficult to link strain identities to strain-specific functions and impeding differential abundance testing.

Method used

A method and system for microbiota subspecies analysis using machine learning and deep learning to create a consistently annotated catalog of human gut microbiota at a subspecies level, enabling the identification of subspecies with underappreciated roles in human health and disease, and utilizing a panhashome approach for rapid subspecies quantification and identification of genes driving intraspecies variations.

Benefits of technology

The subspecies-level analysis improves disease diagnosis and prediction of therapeutic responsiveness by identifying disease-associated subspecies and their functional characteristics, outperforming species-level methods in predictive accuracy and mechanistic understanding of microbiome-phenotype interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025042133_19022026_PF_FP_ABST
    Figure US2025042133_19022026_PF_FP_ABST
Patent Text Reader

Abstract

Methods and systems related to microbiota bacteria subspecies classification and identification are provided. The methods and systems provided herein can be used, e.g., to identify bacteria subspecies that can confer increased or decreased risk to a disease, such as for example a gastrointestinal cancer or inflammatory disease. The methods and systems may include implementation of machine learning, deep learning, and annotated datasets for microbiota bacteria subspecies classification and identification.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD AND SYSTEM FOR MICROBIOTA SUBSPECIES CHARACTERIZATION, IDENTIFICATION, DIAGNOSTICS, AND THERAPY BACKGROUND FIELD

[0001] The present disclosure relates generally to the field of software engineering, molecular biology, and medicine. More particularly, it concerns methods and systems for microbiota subspecies analysis involving machine learning approaches. DESCRIPTION OF RELATED ART

[0002] Although advances have been made for using machine learning for microbiota analysis, significant limitations still exist. The development of advanced computational tools has been used to quantify microbe abundances primarily by using custom references1,2, with microbes being annotated to species level. Strains within the same species can notably differ in their gene content3, and these intraspecies variations can drive distinct functional characteristics between strains belonging to the same species. Conflating different strains can obscure associations with a condition or a treatment, thus hindering deductive research. One of the main challenges in establishing causal links between the microbiome and health is the limitation of current analysis methods to species level due to these intraspecies variations. The bacterial strains, as the highest resolution, could differ at only one nucleotide and are mostly host- specific4. Systematically quantifying the microbiota on a strain level is therefore unsuitable, and most existing algorithms only focus on the dominant strain within a sample. Accordingly, linking strain identities to strain-specific functions is difficult, contributing to challenges in investigating the underlying mechanisms of microbiome-disease interactions. These issues in microbiome research suggest that both species, and strain-level microbiome profiling are suboptimal for differential abundance testing across samples and studies, highlighting the need to enhance the resolution5–10. Some work has been performed to try to address the level between species and strains, or subspecies resolution, in one11,12or a limited range of selected species13, where functional, phenotype-specific differences between sibling subspecies were observed. Although it is possible to differentiate strains based on single nucleotide variations (SNVs), current tools have several drawbacks. Detangling species variability based on calling SNVs directly from sequencing data4,14,15defines new subspecies groups for every analyzed cohort, thus disabling formulation of generally applicable conclusions, albeit acknowledging the intraspecies variability. In light of the above challenges, there is a strong need for a consistently annotated catalog of the human gut microbiota on a subspecies level in microbiome research, which could help or enablethe identification of subspecies with underappreciated roles in human health and disease. Clearly, there is a need for new methods and systems for microbiota analysis. SUMMARY

[0003] The present disclosure overcomes limitations in the art by providing methods and systems for microbiota subspecies analysis. Systems and methods provided herein may utilize machine learning or deep learning (also referred to as artificial intelligence or AI methodologies) and may be used for the diagnosis of diseases, including but not limited to cancers, gastrointestinal diseases, inflammatory disorders (e.g., Chron’s disease, inflammatory bowel disorder (IBD), etc.), and predicting responders and nonresponders to therapy or treatment. It has been observed that utilization of subspecies analysis surprisingly and unexpectedly improved analysis or microbiome data, e.g., for the detection of diseases such as colorectal cancer or IBD. A consistently annotated catalog of human microbiota at a subspecies level is also provided herein, as are methods for generation and classification of microbiota at a subspecies level.

[0004] Systems and methods for subspecies of microbiome bacteria as provided herein can allow for significant advantages for diagnosis of disease and / or identification of therapeutically helpful or disease-causing bacteria, and / or predicting responsiveness to therapy against a disease. For example, subspecies analysis improved prediction of cancer responses to immunotherapies (e.g., FIGS. 13A-D showing data from melanoma patients). Subspecies microbiome can be performed on a subject such as a human patient before and / or during treatment with a therapy (e.g., an anticancer immunotherapy, chemotherapy, therapeutic to treat an inflammatory or gastrointestinal disease, etc.) to assess the likelihood of therapeutic responsiveness. In some aspects, subspecies analysis is used to diagnose a disease. As shown in the Examples and figures, data is provided herein showing that subspecies bacteria analysis can improve microbiome analysis and identify effects (e.g., improved diagnosis of a disease, or improved predicted responsiveness to a therapeutic) would are not observed at either the higher (bacterial species) level or the lower (bacterial strain) level. Microbial strains from the same species can have distinct functional characteristics owing to their different gene content. As the highest resolution, strains are mainly host-specific, thus obscuring unbiased associations, and hindering deductive research. In the instant disclosure, human gut microbiota are comprehensively defined at consistently-annotated subspecies resolution in an unbiased, cohort-independent manner. The disclosed embodiments describe the development of panhashome, a sketching-based method for rapid subspecies quantification and identification of genes that drive the intraspecies variations and showed that subspecies carry implicit information that can be undetectable at species level. By meta-analysis of colorectal cancer (CRC) datasets, disease-associated subspecieswere identified whose sibling subspecies or parental species are not. Subspecies-based machine- learning CRC diagnostic algorithm outperformed species-level methods by leveraging this unique subspecies-level information. The subspecies catalogue provided herein can allow for the identification of genes that can drive the functional differences between subspecies as fundamental step in mechanistically understanding microbiome-phenotype interactions. Systems and related methods for subspecies classification and analysis of microbiome data are also provided herein.

[0005] As shown in the below examples, a comprehensive human gut microbiota subspecies (HuMSub) catalogue was generated by defining the human microbiota at subspecies resolution in an unbiased and cohort-independent manner. It is demonstrated that at this resolution, it is possible to generalize across different and distinct populations worldwide while still maintaining specificity, and subspecies can carry novel implicit information not present at the species level. This concept was empirically demonstrated using meta-analysis of all available colorectal cancer patient datasets, where subspecies were identified as associated with the disease whose sibling subspecies or species are not. Since sibling subspecies share a large portion of their genome, with this work a unique setup was established for fast and easy subspecies quantification, as well as identification of specific genes that drive the subspecies variations and link them to the host phenotype. Additionally, subspecies-based machine-learning approaches were created and demonstrated that the novel layer of information provided on the subspecies level was consistently observed to outperform machine-learning algorithms that rely on species resolution. The HuMSub catalogue contains a comprehensive collection of consistently annotated subspecies of the human microbiome to date, setting the ground for the analysis of new and reanalysis of existing datasets at an unprecedented depth. Identifying specific genes and, hence, functions is fundamental in understanding the underlying mechanism of microbiome-disease interactions. This catalogue, and the workflows as described herein can be used to identify specific genes that can drive differences in subspecies abundance and function, thus providing basis for mechanistic explanations for the associations, which have been lacking in routine microbiome research.

[0006] A preferred form of the present application presents a computer- implemented method for creating an annotated catalog or dataset of a number of microbiome bacteria subspecies to identify a characteristic specific to one or more microbiome bacteria subspecies where the characteristic is associated with a disease or a therapy. For example, a characteristic could encompass a genome or function or subspecies specific phenotype associated with a human disease or therapy. This form includes building a dataset of subspecies as a collection of hash values specific to said subspecies, by selecting hash values that appear in a range of at least 10-95% of the genomes of said subspecies, but in range of less than 1-9% of the genomes fromdifferent subspecies. In a preferred form, hash values are selected that appear in at least about 20% of the genomes of said subspecies, but in less than about 5% of the genomes from different subspecies.

[0007] For diagnostic purposes, a machine learning model based on the annotated catalog of microbiome bacteria subspecies is used. In this example, a machine learning classifier model is trained using one or more studies containing microbiome samples quantified at the subspecies level and associated with a disease or therapy. During training, the classifiers leverage the information present at the subspecies level to successfully extract non-linear combination of subspecies abundances that can serve as markers of the said disease. The classifiers are used to operate the trained machine learning model by applying the classifiers to a microbiome sample to identify a characteristic as associated with a disease, a disease resistance, or a therapy. For example, the characteristic may be one or more genomes or functions associated with a cancer, such as a gastrointestinal tract cancer, including colorectal cancer.

[0008] In another preferred form, a computer-implemented method is provided for creating a catalog or dataset of microbiome bacteria subspecies to identify a characteristic in microbiome bacteria in a human associated with a disease or therapy. For diagnostic purposes, this form may include training a machine learning model using studies associating a characteristic with a disease or therapy to identify classifiers. The form uses the dataset of subspecies and applies the classifiers to a microbiota-containing sample from a human to identify the presence of the characteristic.

[0009] In another preferred form a system is provided to analyze genome data for a microbiome bacteria sample obtained from a human (or possibly another living organism). The microbiome bacteria sample may be obtained from the human body and then analyzed outside the body (such as in an analytical tool as described further). In certain embodiments, the genome data for the microbiome bacteria sample (e.g., a biological sample comprising microbiome bacteria, stool sample, etc.) is obtained by in vitro sequencing analysis or other similar analytic techniques operated on the sample outside of the living body (e.g., by a sequencing analytic tool, DNA sequencing, RNA sequencing, next generation sequencing (NGS), third-generation sequencing, high throughput DNA sequencing). For example, the sequencing analysis may comprise whole metagenome shotgun sequencing (WGS) to obtain genome data for the microbiome bacteria. WGS generally involves obtaining more comprehensive genome information based on sequencing the DNA in the microbiome bacteria sample. A variety of next-generation DNA sequencing options are commercially available, including for example Illumina platforms, and third- generation sequencing approaches such as PacBio SMRT (Menlo Park, CA) or total genomesequencing from and Oxford Nanopore Technologies (ONT; Oxford, UK). The sequencing analysis may be a metagenomic analysis or shotgun sequencing approach comprising extracting and analyzes total DNA of the bacterial community present in the sample; metagenomics sequencing can utilize culture-independent techniques to obtain collective genome information from all bacteria or microorganisms present Using one or more computers, a dataset of subspecies is built as a collection of hash values specific to a subspecies, by selecting hash values that appear in a range of at least 10-95% (e.g., at least 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, or at least 95%) of the genomes of said subspecies, but in range of less than 1-9% (e.g., less than 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, or less than 1% ) of the genomes from different subspecies. In a preferred form, hashes that appear in at least about 20% of the genomes of said subspecies, but in less than about 5% of the genomes from different subspecies are selected. One or more studies which associate a characteristic, such as a gene, pathway or function, with a disease or therapy are processed with the dataset using a machine learning model to build classifiers. In a preferred form such a machine learning model may be based on a gradient boosting approach, such as a Light Gradient Boosting Machine Inputting a microbiome bacteria sample from a human into the machine learning model based on the subspecies dataset identifies said characteristic in the sample.

[0010] One or more studies which associate a characteristic with a disease or therapy can be used to train a machine learning model based on the subspecies dataset to build classifiers to identify a characteristic associated with a disease or therapy. The classifiers are applied to a microbiome bacteria sample from a human to identify the presence of the characteristic associated with a disease or therapy in the sample.

[0011] An aspect of the present disclosure relates to a computer-implemented method for creating an annotated catalog or dataset of a number of microbiome bacteria subspecies comprising: building a dataset of subspecies as a collection of hash values specific to a subspecies, including-sketching genomes of microbiota species producing a set of hash values for a genome; selecting hash values specific to a subspecies that appear in a range of at least 10-95% of the genomes of said subspecies, but in range of less than 1-9% of the genomes from different subspecies; and including selected hash values specific to said subspecies in said dataset. The dataset of subspecies may be built by selecting hash values that appear in at least about 20% of the genomes of said subspecies, but in less than about 5% of the genomes from different subspecies. Said dataset of subspecies may be used to identify a characteristic of one or more microbiome bacteria subspecies where the characteristic is associated with a disease, a disease resistance, or a therapy. Said characteristic may be associated with cancer, inflammatory disease, or any diseasewhere changes in microbiota on a subspecies level can be detected. The inflammatory disease may be inflammatory bowel disease (IBD), Chron’s disease, or ulcerative colitis. The cancer may be a gastrointestinal cancer, a melanoma, or an adenoma. Said characteristic may be associated with a gastrointestinal tract cancer, including colorectal cancer, and said association of said characteristic is an increased incidence of the cancer. The characteristic may be associated with resistance to a disease, wherein the disease is cancer or any disease where changes in microbiota on a subspecies level can be detected. The cancer may be a gastrointestinal cancer, colorectal cancer, adenoma, or any cancer where changes in microbiota on a subspecies level can be detected. The method may comprise identifying a disease, disease resistance, or response to therapy, comprising training a machine learning model using a one or more studies applied to the subspecies dataset which associate a characteristic with a disease or therapy to build classifiers; operating said trained machine learning model by applying said classifiers to a sample to identify said characteristic. The method may include inputting a microbiota-containing sample from a human into the trained machine learning model and identifying if said characteristic is present in the sample. The machine learning model may include gradient boosting, and / or may comprise a Light Gradient Boosting Machine. The machine learning model may comprise an Artificial Neural Network. The method may comprise training the machine learning model with additional data such as risk factors, including one or more of age, lifestyle factors, family and / or personal history. The method may comprise training the machine learning model with data from one or more colorectal cancer studies. The method may comprise training the machine learning model with additional data such as fecal occult blood test. The method may comprise training the machine learning model with top ranked subspecies associated with minimal microbial signature. The method may include a hash function building said hash values, wherein the hash function comprises a murmur function used for sketching. Building the subspecies dataset may include Sourmash employing a hash function to sketch genomes of microbiota species. Genomes of microbiota species may be clustered into species-level OTUs (species-level operational taxonomic units) having at least 95% average nucleotide identity prior to said building the dataset of subspecies. Said building the dataset of subspecies may comprise clustering species in an unsupervised manner. The clustering species in an unsupervised manner may comprise using a Prediction Strength metric with a cutoff between 0.5 and 0.95. The Prediction Strength metric may include a cutoff of about 0.8. Said building the dataset of subspecies may comprise removing contaminated or incomplete genomes. GUNC and / or BUSCO may be re used to remove contaminated or incomplete genomes. A murmur hash function may be used to sketch genomes of microbiota species. The subspecies dataset may include an identification of one or more genesdescribing the functional variation between subspecies. The method may comprise building the subspecies dataset by obtaining sketches of genomes annotated with subspecies information. The sketches of genomes may include or consist essentially of sketches of coding sequences predicted in microbial genomes.

[0012] Another aspect of the present disclosure relates to a computer- implemented method for creating a catalog or dataset of microbiome bacteria subspecies to identify a characteristic in microbiome bacteria in a human associated with a disease, a disease resistance, or a therapy, comprising: building a dataset of subspecies where the dataset is identified as quantified subspecies abundances using hash values specific to said subspecies by obtaining sketches of genomes annotated with subspecies information; training a machine learning model using said dataset of subspecies and using studies associating a characteristic with a disease, disease resistance, or therapy to identify classifiers; applying the classifiers to a microbiota- containing sample; and identifying in the microbiota-containing sample and using said classifiers, said characteristic associated with the disease, disease resistance, or therapy. The method may further comprise identifying a function specific to a subspecies using hash values specific to said subspecies by obtaining sketches of genomes annotated with subspecies information; where said subspecies is associated with a disease and sibling subspecies and / or parental species are not associated with said disease. The dataset of subspecies may be built by selecting hash values specific to said subspecies that appear in at least about 10- 95% of the genomes of said subspecies, but in less than about 1-9% of the genomes from different subspecies. The dataset of subspecies may be built by selecting hash values that appear in at least about 20% of the genomes of said subspecies, but in less than about 5% of the genomes from different subspecies. Said characteristic may be associated with cancer. The therapy may be an anti-cancer therapy, immunotherapy, chemotherapy, biologicals, hormone therapy, hyper-and hypothermia, photodynamic therapy, radiation therapy, stem cell transplant, checkpoint therapy, or anti-PD1 therapy. The characteristic may be predicted response to the therapy or response of the human (e.g., therapeutic response, an adverse side effect, or pharmacokinetics) to the therapy.

[0013] Yet another aspect of the present disclosure relates to a system to classify a microbiome bacteria sample from a human, the system comprising: one or more computers; and one or more non-transitory computer readable media couple to the one or more computers and having instructions stored thereon which, when executed by the one or more computers causes the one or more computers to perform operations comprising- building a dataset of subspecies as a collection of hash values specific to a subspecies, by selecting hash values that appear in at least about 10- 95% of the genomes of said subspecies, but in less than about 1-9% of the genomes fromdifferent subspecies; obtaining one or more studies which associate a characteristic with a disease or therapy; processing the studies and dataset of subspecies using a machine learning model to build classifiers to identify a characteristic associated with a disease or therapy; inputting microbiome bacteria sample from a human into the machine learning model, applying said classifiers and identifying said characteristic associated with a disease or therapy in the sample. Said characteristic may be associated with a gastrointestinal tract cancer, including colorectal cancer and said association of said characteristic may be an increased incidence of cancer. The system may include a hash function building, wherein said hash values comprises a murmur function used for sketching.

[0014] Another aspect of the present disclosure relates to a method of detecting or identifying risk of a disease in a mammalian subject comprising: (i) obtaining a microbiome sample from the subject; (ii) sequencing genome data from the microbiome sample; (iii) comparing the genome data from the microbiome sample to a catalog or dataset of microbiome bacteria subspecies generated by any one of claims 30-35 or the annotated catalog or dataset of microbiome bacteria subspecies of any one of claims 1-29; wherein if microbiome subspecies associated with the disease are identified in the microbiome sample, then the subject has an increased risk of or has the disease. The disease may be a cancer (e.g., a gastrointestinal cancer, a colorectal cancer, an adenoma, or a melanoma). The disease may be an inflammatory disease (e.g., Chron’s disease, inflammatory bowel disorder (IBD), or ulcerative colitis). The subject may be a human.

[0015] As used herein, the term “subspecies” refers to groups of bacteria within a species of bacteria. Subspecies group together at least two or more strains of bacteria. Multiple or many strains of bacteria within a bacterial species can be grouped together as a subspecies based on similarity estimated using hashing approaches. Preferably, subspecies share a highly similar or essentially identical phenotype.

[0016] As used herein the specification, “a” or “an” may mean one or more. As used herein in the claim(s), when used in conjunction with the word “comprising,” the words “a” or “an” may mean one or more than one.

[0017] The use of the term “or” in the claims is used to mean “and / or” unless explicitly indicated to refer to alternatives only or the alternatives are mutually exclusive, although the disclosure supports a definition that refers to only alternatives and “and / or.” As used herein “another” may mean at least a second or more.

[0018] Throughout this application, the term “about” is used to indicate that a value includes the inherent variation of error for the device, the inherent variation in the methodbeing employed to determine the value, the variation that exists among the study subjects, or a value that is within 10% of a stated value.

[0019] As used herein, “essentially free,” in terms of a specified component, is used herein to mean that none of the specified component has been purposefully formulated into a composition and / or is present only as a contaminant or in trace amounts. The total amount of the specified component resulting from any unintended contamination of a composition is therefore well below 0.05%, preferably below 0.01%. Most preferred is a composition in which no amount of the specified component can be detected with standard analytical methods.

[0020] As used in this specification and claim(s), the words “comprising” (and any form of comprising, such as “comprise” and “comprises”), “having” (and any form of having, such as “have” and “has”), “including” (and any form of including, such as “includes” and “include”) or “containing” (and any form of containing, such as “contains” and “contain”) are inclusive or open-ended and do not exclude additional, unrecited elements or method steps.

[0021] The terms “subject,” “host,” “patient,” and “individual” are used interchangeably herein to refer to any mammalian subject for whom therapy is desired, particularly humans. Other subjects may include cattle, dogs, cats, guinea pigs, rabbits, rats, mice, horses, and so on. Preferably the mammalian subject is a human subject or patient.

[0022] “Treatment” and “treating” refer to administration or application of a therapeutic agent to a subject or performance of a procedure or modality on a subject for the purpose of obtaining a therapeutic benefit of a disease or health-related condition.

[0023] “Annotated” refers to describing the structure and function of the components of a genome, by predicting, analyzing, and interpreting them in order to extract their biological significance and understand the biological processes in which they participate. For example, annotations of a genome may include, but are not limited to Kegg modules and pathways, BioCyc pathway / genome collection and others.

[0024] Other objects, features and advantages of the present disclosure will become apparent from the following detailed description. It should be understood, however, that the detailed description and the specific examples, while indicating preferred embodiments, are given by way of illustration only, since various changes and modifications within the spirit and scope of the disclosed embodiments will become apparent to those skilled in the art from this detailed description.BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The patent or application file contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.

[0026] The following drawings form part of the present specification and are included to further demonstrate certain aspects of the present disclosure. The disclosed embodiments may be better understood by reference to one or more of these drawings in combination with the detailed description of specific embodiments presented herein.

[0027] FIGS. 1A-D. Subspecies from the HuMSub catalogue are ubiquitous across the bacterial domain and are directly quantifiable. (A) Phylogenetic tree of 3483 species. The inner color ring represents phyla attribution, and the height of the red bars shows the number of detected subspecies. (B) The number of different taxonomic categories denoting the five largest phyla. (C) Boxplot of F1-score distribution on species and subspecies level. (D) Boxplot of L2- distance distribution on species and subspecies level. MetaPhlAn4 performance is included only in L2-distance calculation since it defaults to lower taxonomic levels if a higher one can’t be determined, which would result in an artificially low F1-score in this benchmark setting.

[0028] FIGS. 2A-E. Global biogeographical subspecies distribution. (A) Clustered heatmap showing the prevalence of subspecies from the six biggest phyla. (B) UpSet plot showing the number of detected subspecies shared by different continents. (C) Histogram showing the distribution of the Geographical Enrichment Score (GES) for all analyzed subspecies. The vertical dashed line, at GES=0.4, denotes the threshold for geographically restricted subspecies. (D) Barplot showing the percentage of geographically restricted subspecies (GES > 0.4) for the top five analyzed phyla. (E) Stacked barplot of GES values of the top ten species where at least one subspecies is geographically restricted (GES > 0.4).

[0029] FIGS. 3A-E. Subspecies contain implicit information not available at species level. (A) Vulcano plot showing the Random Effect Size and adjusted p-value at the subspecies level. Blue dots represent significantly associated subspecies whose sibling subspecies aren’t associated. Red dots represent significantly associated subspecies whose parental species aren’t associated. (B) Forest plot of effect sizes and random effect sizes for Fusobacterium animalis at species and subspecies levels. (C) Forest plot of effect sizes and the random effect sizes of the top ten subspecies positively and negatively associated with CRC, while their parental species isn’t (Red dots from panel A). (D and E). Forest plots of effect sizes and random effect sizes for Ruthenibacterium lactatiformans (D) and Faecalibacterium prausnitzii (E) species and their subspecies. In B-E color dots represent effect sizes estimated for each study individuallyusing standardized mean difference, while black diamonds represent random effect sizes. Species and subspecies are significantly associated if abs(random effect size) > 0.2 and adjusted p-value < 0.1.

[0030] FIGS. 4A-E. Subspecies are better predictors of colorectal cancer. (A) Boxplots of the AUROC at the species and subspecies level obtained through the LODO approach initialized ten times with different random states. The study used for testing is denoted on the x- axis, while all others were used for training. Difference in performance between species and subspecies level models was tested using Mann-Whitney U-test. (B) Line plot of median AUROC performances between different numbers of studies used for training on species and subspecies level. Each point represents a median AUROC for all combinations of that number of studies used for training, averaged for all test studies. Each model was initialized ten times. (C) Boxplots of the AUROC at species and subspecies level, with and without the result of FOBT used to weigh the model output. The red dashed line indicates the AUROC performance of FOBT alone. (D) Clustered heatmap showing the importance ranking of subspecies found as the top 10 most important ones for at least one study. Marked in red are subspecies whose sibling subspecies exist in the dataset but weren’t marked in the top 50. (E) Line plot of AUROC performance obtained through LODO approach, where only the most important subspecies are included during training. To prevent overfitting, subspecies were ranked using models not trained on the dataset used for testing.

[0031] FIGS. 5A-H. Sibling subspecies can differently affect host physiology. (A) Boxplot showing number of genes fully specific for a certain subspecies and number of genes with high-impact variants specific for a subspecies of top ten species associated with CRC. (B) Barplot denoting KEGG modules to which of the genes shown in A) they belong. (C) Barplot showing KEGG modules to which genes with high-impact variants of Ruthenibacterium lactatiformans belong. Aerobic and anaerobic cobalamin biosynthesis are grouped together. (D- H) MSA output of genes sirC (D), cobI (E), argH (F), hemH (G) and gltX (H) found in the corresponding species. Red color marks variants that are predicted to lower the stability of the protein (abs(^^G)>1). Sequences were dereplicated at the full length on the subspecies level. MSA visualizations are shortened to the region of interest.

[0032] FIGS.6A-E. Summary of the subspecies delineation, related to FIGS.1A- D. (A) Workflow of subspecies delineation. Completeness estimation was done using BUSCO, while contamination estimation was done using both GUNC and BUSCO. (B) Jointplot of checkM and BUSCO completeness assessments. checkM completeness was obtained from the HumGut catalog. (C) Histogram showing the distribution of subspecies number per species. (D) Stackedbarplot showing the percentage of species and subspecies grouped on the phylum level. (E) Stacked barplot describing the distribution of subspecies cluster sizes.

[0033] FIGS. 7A-E. The subspecies-specific panhashome approach is efficient and precise in quantifying subspecies, related to FIGS. 1A-D. (A) Heatmap showing pairwise- calculated containment of subspecies-specific hashes across all subspecies. Labels are omitted for visualization purposes. (B) Boxplot denoting runtime for MetaPhlAn4 run on 8 cores using default options and sourmash-based subspecies quantification. (C) Boxplot denoting memory required for MetaPhlAn4 using default options on 8 cores and sourmash-based subspecies quantification. (D) Boxplot of F1-score distribution on species and subspecies level calculated on samples simulated using genomes outside of the HumGut catalog. (E) Boxplot of L2-distance distribution on species and subspecies level calculated on samples simulated using genomes outside of the HumGut catalog.

[0034] FIGS. 8A-D. Geographical Enrichment Score (GES) correctly estimates subspecies’ geographical restriction, related to FIGS.2A-E. (A) UpSet plot showing the number of detected subspecies shared by different African countries. (B) Jointplot denoting the correlation between GES and median adjusted p-value obtained through pairwise country comparisons for each subspecies. (C) Histogram showing the distribution of median differences in prevalence between countries for each subspecies. In red, subspecies marked as “low specific” (GES < 0.4); in blue, subspecies marked as “high specific” (GES > 0.4). (D) Barplot showing a number of geographically restricted (GES > 0.4) subspecies grouped into phyla per NCBI taxonomy.

[0035] FIGS. 9A-D. Sibling subspecies are differentially abundant between healthy and CRC samples, related to FIGS.3A-E. (A) Forest plot of effect sizes and the random effect sizes of the top ten subspecies positively and negatively associated with CRC, while at least one of their sibling subspecies isn’t (Blue dots from Figure 3A). (B and C). Forest plots of effect sizes and random effect sizes for Porphyromonas asaccharolytica (B) and UBA1691 sp900544375 (C) species and their subspecies. Color dots represent effect sizes estimated for each study individually using standardized mean difference, while the black diamonds represent random effect sizes. (D) Clustered heatmap showing significantly associated species using univariate statistics on each study separately. On the subspecies level, if any of the subspecies are associated, the species as marked as associated. Species and subspecies are significantly associated if abs(log2FC) > 0.5 and FDR < 0.2.

[0036] FIGS. 10A-D. Subspecies enable improved prognostics of CRC, related to FIGS.4A-E. (A) Boxplots of the AUROC at the species and subspecies level obtained through the ten-fold cross-validation initialized thirty times with different random states. (B) Heatmapshowing median AUROC of species- and subspecies-level models when trained and tested only on one dataset. (C) Boxplot showing the negative predictive value (NPV) of species- and subspecies-level models obtained through ten-fold cross-validation initialized thirty times with different random states. (D) Line plots of median AUROC performances between different numbers of studies used for training on species and subspecies level. Each point represents a median AUROC for all combinations of that number of studies used for training, where each model was initialized ten times. Indicated study refers to a study used for testing.

[0037] FIGS. 11A-G. Subspecies-specific genes and variants, related to FIGS. 5A-H. (A) Histogram showing the frequency of all effects of variants defined as subspecies- specific in the top ten subspecies associated with CRC. The vertical dashed line at abs(^^G)=1 denotes the threshold for high-impact variants. (B) Barplot denoting KEGG modules to which genes that contain high-impact subspecies-specific variants belong. (C) Barplot denoting KEGG modules to which genes that are detected as fully subspecies-specific belong. (D and E) MSA output of genes cobM (D) and cbiG (E) from Ruthenibacterium lactatiformans. Red color marks variants that are predicted to lower the stability of the protein (abs(^^G)>1). All protein sequences were dereplicated on the subspecies level. MSA visualizations are shortened to the region of interest. (F and G) Forest plots of effect sizes and random effect sizes for Enterocloster sp005845215 (F) and Porphyromonas sp900548415 (G) species and their subspecies. Color dots represent effect sizes estimated for each study individually using standardized mean difference, while black diamonds represent random effect sizes.

[0038] FIG. 12: Subspecies improve detection of adenoma. Boxplot of the AUROC of adenoma predictors at species and subspecies level, obtained through five-fold cross- validation initialized ten times with different random states. Difference in performance was tested using the Mann-Whitney U-test.

[0039] FIGS. 13A-F: Subspecies provide deeper understanding of the microbiome's impact on response to anticancer immunotherapy. (A) PCA plot of centered log- transformed abundances of responders and non-responders on the species (top) and the subspecies level (bottom). Non-transparent points denote medioids of the responder (blue) and non-responder (red) groups. (B) Volcano plot showing the Random Effect Size and adjusted values at the subspecies level between responders and non-responders. Red dots represent significantly associated subspecies whose parental species aren’t associated. Blue dots represent significantly associated subspecies whose sibling subspecies aren’t associated. (C and D) Forest plots of effect sizes and random effect sizes for Bilophila wadsworthia (C) and Ruthenibacterium lactatiformans (D) species and their subspecies. Color dots represent the effect sizes estimated for each studyindividually using the standardized mean difference, while black diamonds represent the random effect sizes. Species and subspecies are significantly associated if abs(random effect size) > 0.2 and adjusted Pvalue < 0.2. (E) Boxplot of the AUROC at the species, subspecies, and strain level obtained through 5-fold cross-validation repeated 10 times on the anti-PD1 treated, combined treated, and all samples together. (F) Line plot of median AUROC performances between different numbers of studies used for training at species and subspecies level. Each point represents median AUROC for all combinations of the specified number of studies used for training, averaged across all test studies. Each model was initialized ten times.

[0040] FIG. 14: Subspecies increase the performance of inflammatory bowel disease (IBD) prediction models. Boxplot of the AUROC of IBD predictors at species and subspecies level, obtained through ten-fold cross-validation initialized thirty times with different random states. Difference in performance was tested using the Mann-Whitney U-test.

[0041] FIG.15 is a flow diagram of a method for creating a dataset of subspecies, according to some embodiments.

[0042] FIG.16 is a flow diagram of a method for analyzing a microbiome bacteria sample according to subspecies, according to some embodiments.

[0043] FIG.17 is a block diagram of one embodiment of a computer system that may be used to implement techniques disclosed herein. DESCRIPTION OF ILLUSTRATIVE EMBODIMENTS I. Machine Learning and Deep Learning Systems and methods

[0044] As noted above, the methods and systems described herein broadly build one or more datasets of bacteria subspecies and then use machine learning to identify subspecies that correlate to the trait or characteristic of interest. For example, machine learning can be used to identify bacterial subspecies in a human gut microbiome that correlate with or cause an increased or decreased risk of a cancer or other diseases in humans (e.g., an inflammatory gastrointestinal disease). Machine learning can be also used to identify bacterial subspecies in a human that correlate with good health or resistance to a cancer or disease in the human. Various techniques for building such datasets of bacteria subspecies and machine learning approaches to identify characteristics are described generally and specific techniques used in preferred embodiments are outlined herein.

[0045] Artificial Intelligence is often thought of as encompassing Machine Learning techniques which in turn encompasses Deep learning. As used in the present application, a machine learning model is a type of mathematical model which, after being "trained" on a given dataset, can be used to make predictions or classifications on new data. Machine learningtechniques that use gradient boosting are particularly useful with high-dimensional sparse datasets and in a preferred embodiment hereof a machine learning model using the Light Gradient Boosting Machine is described. Artificial Neural Networks (ANNs), decision trees and support vector machines are known machine learning examples. Training methods used can be either supervised, semi-supervised (sometimes known as reinforcement) or unsupervised.

[0046] Machine Learning is a subset of artificial intelligence (AI) methods, which leverage large datasets to recognize, classify, and predict patterns. Machine Learning approaches may include Prediction Strength metrics or various algorithms. For example, specific machine learning approaches that can be used in various aspects of the present disclosure include: GUNC, BUSCO, Sourmash, Prediction Strength Metric, MetaPhlAn4, LEMMI, LGBM, LODO, and Mastiff. LGBM, LODO.

[0047] A number of types of ANNs exist and are inspired by biological neural networks. Typically, ANNs include an input layer and an output layer, with multiple hidden layers in between. Feedforward ANNs are common and Convolutional neural networks (CNN) are a type of feedforward ANNs. Other types include regulatory feedback, radial basis function, deep belief networks, recurrent neural networks, modular neural networks, physical neural networks, dynamic neural networks, and memory networks. CNN can assimilate both isolated topographies as well as entire images and classify these images according to their unique features . In addition to CNN, other neural network algorithms that may be used in aspects of the present disclosure including Radial Function Networks, Multi-Layer Perceptrons (MLP), Self Organizing Maps, and Recurrent Neural Networks (RNN).

[0048] A hash function may be used during generation of a subspecies listing. Hash functions and associated hash tables are useful in storing and retrieving data records and are computationally and storage efficient. In computer science and data mining, multiple hashing- based approaches could be used, such as MinHash (the min-wise independent permutations locality sensitive hashing scheme) or FracMinHash. Among its use cases, those techniques can quickly estimate how similar two sets are. The Jaccard similarity coefficient is a commonly used indicator of the similarity between two sets. While different hash functions may be employed with the methods and systems described herein, MurmurHash is believed suitable along with its known variants, including FarmHash, Jenkins Hash, CityHash, and / or DJB2.

[0049] Some embodiments use FracMinHash (such as used in SourMash)-based sketching to map and store subspecies genomic information, it being understood that other hash functions exist. In particular, Sourmash software can be preferably used with different hash functions employed. Another preferred form uses a FracMinHash sketching technique in theSourmash software. Generally speaking, a sketch is a set (or collection) of selected hash values from a genome.

[0050] Quantifying subspecies was previously challenging due to the high genetic similarity and large genomic segments that are shared among strains within the same species. While methods utilizing SNVs can quantify strains and subspecies, they only offer an indirect calculation of their abundance based on the overall species abundance. To overcome these issues, a method is provided herein for direct quantification of the previously defined subspecies from metagenomic sequencing, for which a custom-defined concept of subspecies-specific “panhashome” is developed, as a collection of hash values specific to each subspecies. All genomes of microbiota species may be sketched using Sourmash21. Then, one can choose or select the hashes that appeared in at least 20% of the genomes of a particular subspecies, but in less than 5% of the genomes from different subspecies (e.g., as shown in Figure 2A). A subspecies quantification benchmark can be constructed and was shown to validate the newly established quantification approach. Metagenomic samples can be simulated, and the Euclidian (L2) distance can be assessed, which measures the difference between ground truth and quantified abundances, and the F1-score as a rate of false positives and false negatives, on both species and subspecies resolution. Computational resources required for the quantification can be tracked. To help interpret the performance metrics, which could be influenced by the complexity of simulated microbiome samples, etaPhlAn41can be included as a computational tool for species-level profiling commonly used in the field. Since there is no possibility to include custom references in MetaPhlAn4, it may be used only on the species level. The median F1-score and L2 distances observed in the below examples were 0.794 and 0.144, respectively, suggesting a very high performance of our approach, slightly outperforming MetaPhlAn4 (Figures 1C and 1D) for the species-level profiling. Both runtime and total memory use were markedly reduced when using this sourmash-based approach compared to MetaPhlAn4, showing that the quantification of subspecies is both faster and less memory intensive compared to previous state-of-the-art species- level quantification (Figures 2B and 2C). These results demonstrate surprising and unexpected advantages for using the subspecies analysis methods provided herein. The high performance of the sourmash-based approach is in agreement with the LEMMI independent benchmark23. II. Subspecies Catalogue Datasets

[0051] In some aspects, subspecies catalogues that can be used in machine learning or deep learning algorithms are provided herein. The HuMSub catalogue that comprehensively defines human microbiota at subspecies resolution by establishing a panhashome-based method is provided, and can enable identification of specific genes that driveinter-subspecies phenotypic differences. Meta-analysis of colorectal cancer datasets reveals disease-associated subspecies that can be used to train machine-learning diagnostic algorithms that were observed to consistently outperform species-level methods. Subspecies-level datasets and analysis can enable or be used for improving understanding of microbiome-condition interactions.

[0052] The HuMSub listing provided in the Examples below contains 5361 subspecies. The HuMSub catalogue can be generated or defined irrespective of what, if any, machine learning method(s) or disease studies are subsequently performed using the HuMSub listing. The subspecies catalogue may increase when new microbiome bacteria are identified or as genetic drift occurs within bacteria. The subspecies abundancies, obtained using the catalogue and the panhashome approach can be further used for subsequent analysis as described herein.

[0053] Previous research, pioneered by efforts in a limited range of selected species, focused on advancing from the classical species resolution, pointing to the need of generation of a uniform and comprehensive reference system on the subspecies level that can be optimally compared across studies. To generate the HuMSub catalogue, a comprehensive genome catalog16in which genomes originate from diverse populations is used, and then focused on predicting the coding sequences. In contrast to approaches using genome-wide SNVs, methods described herein enabled detecting consistent variations that cause actual phenotypic diversities. The sketching approach for subspecies delineation, as implemented in Sourmash21, enables defining subspecies groups as collections of genomes with known sequences, which can render identifying the subspecies-specific functions relatively straightforward. With the custom subspecies-specific panhashome concept, optimal computational efficiency of sketching was utilized to generate a highly compact resource, taking up 127 MB of memory space, enabling it to be used without requiring extensive resources, and allowing an easy sharing compared to classical reference-based tools.

[0054] Even though higher taxonomic resolution could be followed by a loss of replicability across populations, the worldwide population analysis showed that HuMSub catalogue generalizes across highly diverse metagenomic samples, enabling its use in meta- analyses and other means of cross-cohort research. Moreover, HuMSub enables identification of functional microbiota differences relevant to the condition or the host phenotype based on the differential sibling-subspecies associations. Such approaches advance from the species-level methods currently used, since they allow accessing the functional changes and direct contribution of the driver subspecies critical to the phenotype in a direct comparison to their sibling, non- associated subspecies.

[0055] Microbiota species-level machine learning approaches are promising as a new non-invasive tool for the detection and diagnosis of certain diseases, including colorectal cancer (CRC). However, their direct applicability is still limited mainly by the insufficient prognostic rates embedded in the intrinsic level of information available at the species level. By using HuMSub, 204 CRC-associated subspecies were found, out of which, 95 (46.5%) were instances where at least one subspecies was significantly associated with CRC and at least one sibling subspecies was not, and 26 subspecies associations were uncovered, not detectable at the corresponding parental species level. Based on this new layer of associations between the microbiota and CRC, it was found that, without any prior knowledge, subspecies-based classifiers outperformed the corresponding species-level models in distinguishing samples from healthy and diseased subjects. The identification of subspecies interactions that are not detectable at a species level provides additional support regarding the unexpected advantages of the subspecies methodologies and systems provided herein.

[0056] Based on the superior prognostic performance (median LODO AUROC=0.842, reaching over 0.890 for certain datasets), training the subspecies-based machine learning models by including additional studies, optionally in combination with other easy-to-use simple pre-diagnostic tests, as well as by modeling for the different risk factors, including age, lifestyle factors, family and / or personal history, can enable development of diagnostic machine-, or deep-learning based approaches for accurate detection of various diseases, including CRC and other cancers, as well as predicting responders vs. non-responders to certain treatments, based on single microbiota sequencing. HuMSub increases inter-study microbiota-analysis reproducibility and reduces variability between metagenomic studies, thereby enabling establishment of a consistent microbial signature for a host phenotype. Dwelling on the subspecies differences, HuMSub catalogue enables discovery of new mechanistic insights into the interplay between the microbiota and host phenotype, and sets the ground for analyses of new and reanalyses of existing datasets at an unprecedented depth. III. Diseases

[0057] Subspecies datasets, machine learning, and / or deep learning approaches can be used to identify subspecies of bacteria that correlate with or cause an increased or decreased risk of a disease, such as for example, a gastrointestinal disease, an inflammatory disease (e.g., IBD, Chron’s disease, etc.), a cancer, or a gastrointestinal disease. The methods and systems provided herein can be used to determine or predict how a disease (e.g., a disease already diagnosed in a mammalian subject such as a human) may respond to a therapeutic (e.g., therapeutic efficacy of a drug, possible adverse side effects from the drug, pharmacokinetics of a drug, etc.).For example, the systems can be used to predict responsiveness of a patient to an anticancer drug, such as for example the effectiveness of an immunotherapy against melanoma.

[0058] The disease may be a gastrointestinal disease. For example, the disease may be irritable bowel syndrome (IBD), gastroesophageal reflux disease (GERD), ulcerative colitis, celiac disease, diverticulitis, gallstones, peptic ulcers, gastroenteritis, pancreatitis, dyspepsia, colorectal polyp, gastritis, gastroparesis, or a disease of the esophagus, stomach, small intestine, large intestine or, rectum. It is anticipated that methods and systems provided herein can be applied to essentially any disease characterized by alterations in microbiota. For example, the disease may be cachexia, metabolic disease including obesity, dyslipidemia, diabetes, atherosclerosis, autoimmune disease, neuroinflammatory disease, bone disease, neurodegenerative disease, or an inflammatory disease.

[0059] Subspecies datasets, machine learning, and / or deep learning approaches can be used to identify subspecies of bacteria that can be used as a therapeutic (e.g., to treat or reduce the risk of a disease) or that correlate with or cause an increased or decreased effectiveness of a therapeutic (e.g., a drug) in a mammalian subject, preferably a human. In some preferred aspects, the therapeutic is a subspecies of bacteria that can be useful for treating or reducing the risk of a gastrointestinal disease or cancer, for example as described above.

[0060] Subspecies datasets, machine learning, and / or deep learning approaches can be used to identify subspecies of bacteria that correlate with resistance to the disease in a mammalian subject, preferably a human. The disease may be a gastrointestinal disease. For example, the disease may be irritable bowel syndrome (IBD), gastroesophageal reflux disease (GERD), ulcerative colitis, celiac disease, diverticulitis, gallstones, peptic ulcers, gastroenteritis, pancreatitis, dyspepsia, colorectal polyp, gastritis, gastroparesis, or a disease of the esophagus, stomach, small intestine, large intestine or rectum. The disease may be any disease where microbiota is changed. For example, the disease may be cachexia, metabolic disease incliding obesity, dylipidaemia, diabetes, atherosclerosis, autommune disease, neuroinflammatory disease, bone disease, neurodegenerative disease, or inflammatory disease. A. Cancers

[0061] Subspecies datasets, machine learning, and / or deep learning approaches can be used to identify subspecies of bacteria that correlate with or cause an increased or decreased risk of a cancer in a mammalian subject, preferably a human. In some preferred aspects, the cancer is a cancer of the gastrointestinal tract, such as a cancer of the colon, intestine, stomach, anus, pancreas, liver, or esophagus. Nonetheless, it is anticipated that the methods, datasets, and systems provided herein can be used to identify subspecies bacteria that increase or decrease the risk of orare involved with causing or ameliorating a wide variety of cancers. The systems and methods provided herein can also be used to determine the effectiveness of an anticancer drug (e.g., checkpoint inhibitor) against the cancer. The cancer may be a gastrointestinal cancer or a non- gastrointestinal cancer or a that is not present in the gastrointestinal system.

[0062] The cancer may be, e.g., a cancer from the bladder, blood, bone, bone marrow, brain, breast, colon, esophagus, gastrointestine, gum, head, kidney, liver, lung, nasopharynx, neck, ovary, prostate, skin, stomach, testis, tongue, or uterus. The cancer may specifically be of the following histological type, though it is not limited to these: neoplasm, malignant; carcinoma; carcinoma, undifferentiated; giant and spindle cell carcinoma; small cell carcinoma; papillary carcinoma; squamous cell carcinoma; lymphoepithelial carcinoma; basal cell carcinoma; pilomatrix carcinoma; transitional cell carcinoma; papillary transitional cell carcinoma; adenocarcinoma; gastrinoma, malignant; cholangiocarcinoma; hepatocellular carcinoma; combined hepatocellular carcinoma and cholangiocarcinoma; trabecular adenocarcinoma; adenoid cystic carcinoma; adenocarcinoma in adenomatous polyp; adenocarcinoma, familial polyposis coli; solid carcinoma; carcinoid tumor, malignant; branchiolo- alveolar adenocarcinoma; papillary adenocarcinoma; chromophobe carcinoma; acidophil carcinoma; oxyphilic adenocarcinoma; basophil carcinoma; clear cell adenocarcinoma; granular cell carcinoma; follicular adenocarcinoma; papillary and follicular adenocarcinoma; nonencapsulating sclerosing carcinoma; adrenal cortical carcinoma; endometroid carcinoma; skin appendage carcinoma; apocrine adenocarcinoma; sebaceous adenocarcinoma; ceruminous adenocarcinoma; mucoepidermoid carcinoma; cystadenocarcinoma; papillary cystadenocarcinoma; papillary serous cystadenocarcinoma; mucinous cystadenocarcinoma; mucinous adenocarcinoma; signet ring cell carcinoma; infiltrating duct carcinoma; medullary carcinoma; lobular carcinoma; inflammatory carcinoma; paget's disease, mammary; acinar cell carcinoma; adenosquamous carcinoma; adenocarcinoma w / squamous metaplasia; thymoma, malignant; ovarian stromal tumor, malignant; thecoma, malignant; granulosa cell tumor, malignant; androblastoma, malignant; sertoli cell carcinoma; leydig cell tumor, malignant; lipid cell tumor, malignant; paraganglioma, malignant; extra-mammary paraganglioma, malignant; pheochromocytoma; glomangiosarcoma; malignant melanoma; amelanotic melanoma; superficial spreading melanoma; melanoma in giant pigmented nevus; epithelioid cell melanoma; blue nevus, malignant; sarcoma; fibrosarcoma; fibrous histiocytoma, malignant; myxosarcoma; liposarcoma; leiomyosarcoma; rhabdomyosarcoma; embryonal rhabdomyosarcoma; alveolar rhabdomyosarcoma; stromal sarcoma; mixed tumor, malignant; mullerian mixed tumor; nephroblastoma; hepatoblastoma; carcinosarcoma; mesenchymoma, malignant; brenner tumor,malignant; phyllodes tumor, malignant; synovial sarcoma; mesothelioma, malignant; dysgerminoma; embryonal carcinoma; teratoma, malignant; struma ovarii, malignant; choriocarcinoma; mesonephroma, malignant; hemangiosarcoma; hemangioendothelioma, malignant; kaposi's sarcoma; hemangiopericytoma, malignant; lymphangiosarcoma; osteosarcoma; juxtacortical osteosarcoma; chondrosarcoma; chondroblastoma, malignant; mesenchymal chondrosarcoma; giant cell tumor of bone; ewing's sarcoma; odontogenic tumor, malignant; ameloblastic odontosarcoma; ameloblastoma, malignant; ameloblastic fibrosarcoma; pinealoma, malignant; chordoma; glioma, malignant; ependymoma; astrocytoma; protoplasmic astrocytoma; fibrillary astrocytoma; astroblastoma; glioblastoma; oligodendroglioma; oligodendroblastoma; primitive neuroectodermal; cerebellar sarcoma; ganglioneuroblastoma; neuroblastoma; retinoblastoma; olfactory neurogenic tumor; meningioma, malignant; neurofibrosarcoma; neurilemmoma, malignant; granular cell tumor, malignant; malignant lymphoma; hodgkin's disease; hodgkin's; paragranuloma; malignant lymphoma, small lymphocytic; malignant lymphoma, large cell, diffuse; malignant lymphoma, follicular; mycosis fungoides; other specified non-hodgkin's lymphomas; malignant histiocytosis; multiple myeloma; mast cell sarcoma; immunoproliferative small intestinal disease; leukemia; lymphoid leukemia; plasma cell leukemia; erythroleukemia; lymphosarcoma cell leukemia; myeloid leukemia; basophilic leukemia; eosinophilic leukemia; monocytic leukemia; mast cell leukemia; megakaryoblastic leukemia; myeloid sarcoma; and hairy cell leukemia. IV. Examples

[0063] The following examples are included to demonstrate preferred embodiments of the present disclosure. It should be appreciated by those of skill in the art that the techniques disclosed in the examples which follow represent techniques discovered by to function well in the practice of the disclosed embodiments, and thus can be considered to constitute preferred modes for its practice. However, those of skill in the art should, in light of the present disclosure, appreciate that many changes can be made in the specific embodiments which are disclosed and still obtain a like or similar result without departing from the spirit and scope of the disclosed embodiments. Example 1 – Subspecies from the HuMSub catalogue are ubiquitous across the bacterial domain and are directly quantifiable

[0064] The genomes from HumGut16were used, the most comprehensive catalog of bacterial genomes from the human gut at the time of starting the project. HumGut was built using the UHGG catalog17, with additional assembled and reference genomes, all clustered into species-level OTUs at 95% average nucleotide identity. Since some species could containrelatively diverse genomes, those groups were used instead of direct GTDB and NCBI species for further analysis. Although the genomes were already satisfying the medium quality standard18, more aggressive quality filtering was applied based on GUNC19and BUSCO20to remove more of the contaminated and incomplete genomes (Figures 1A and 1B). For subspecies clustering, sketches were created of coding sequences in their amino acid form, predicted in all genomes using Sourmash21. This was done with the expectation that variants in coding sequences would lead to distinct subspecies phenotypes instead of hard-to-interpret intergenic variation. For species with more than three genomes, the pairwise distances were then calculated and clustered the genomes in an unsupervised manner using the Prediction Strength metric22with a cutoff of 0.8 to select the optimal number of clusters. In some cases, a prediction strength cutoff of 0.5-0.95 is used for clustering groups, but a cutoff of 0.7-0.9 is believed to be preferable.

[0065] After quality filtering and clustering, 225918 genomes were kept represented across 3483 species. Using the above-defined approach, the HuMSub catalogue was generated (Figure 1A) where a total of 5361 subspecies belonging to 977 species were identified, revealing that 28% of species contain previously routinely neglected subspecies variations (Figures 1B and 1C). Notably, the clustering protocol didn’t induce a taxonomic bias, as ratios of phyla remained similar between species and subspecies levels (FIG. 6D). 70% of subspecies contained more than one reference genome (FIG. 6E), suggesting that the clustering approach effectively generalizes across the bacterial genomes.

[0066] Quantifying subspecies was previously challenging due to the high genetic similarity and large genomic segments that are shared among strains within the same species. While methods utilizing SNVs can quantify strains and subspecies, they only offer an indirect calculation of their abundance based on the overall species abundance. To overcome these issues, a method for direct quantification of the previously defined subspecies from metagenomic sequencing has been established, for which a custom-defined concept of subspecies-specific “panhashome” was developed, as a collection of hash values specific to each subspecies. Specifically, all genomes in the species were first sketched using Sourmash21. Then, the hashes that appeared in at least 20% of the genomes of a particular subspecies, but in less than 5% of the genomes from different subspecies (FIG. 7A) were chosen. A subspecies quantification benchmark was constructed to validate the newly established quantification approach. Ten metagenomic samples were simulated and the Euclidian (L2) distance was assessed, which measures the difference between ground truth and quantified abundances, and the F1-score as a rate of false positives and false negatives, on both species and subspecies resolution. Additionally, the computational resources required for the quantification were tracked. To help interpret theperformance metrics, which could be influenced by the complexity of simulated microbiome samples, MetaPhlAn41was included as a computational tool for species-level profiling commonly used in the field. Since there is no possibility to include custom references in MetaPhlAn4, it was used only on the species level. The median F1-score and L2 distance were 0.794 and 0.144, respectively, suggesting a very high performance of this approach, slightly outperforming MetaPhlAn4 (Figures 1C and 1D) for the species-level profiling. Both runtime and total memory use were markedly reduced when using the sourmash-based approach compared to MetaPhlAn4, showing that the quantification of subspecies is both faster and less memory intensive compared to previous state-of-the-art species-level quantification (Figures 2B and 2C). The high performance of the sourmash-based approach is in agreement with the LEMMI independent benchmark23.

[0067] To further validate the quantification approach described herein on novel genomes, a new set of ten metagenomic samples was simulated using genomes uploaded to RefSeq after the publication of the HumGut catalogue. The ones whose NCBI taxonomic IDs were originally contained in the catalogue were kept and further quality-filtered. To annotate them with subspecies information, the similarity of their sketches was used to the subspecies-specific ones and kept genomes with detected similarity > 0.7 (discussed further herein). This benchmark enabled assessing the quantification performance in real-world use. Similar results were obtained, with a median F1-score of 0.777 and L2 distance of 0.151, further indicating high performance on real-world data (Figures 2D and 2E). Example 2 – Global biogeographical subspecies distribution

[0068] The human gut microbiome varies across different populations, even among cohorts otherwise similar in phenotype24–26. The variability increases by increasing the taxonomic resolution, to the extent that most strains are not only population, but even host- specific4. This hinders deductive research and prevents strain level from becoming widely used in metagenomic quantification.

[0069] Given that subspecies offer an increased resolution compared to species and might face similar limitations, its potential utility for meaningful global-scale microbiome research was explored. 5272 publicly available human metagenomic samples with country information were searched for the presence of subspecies using Mastiff27. By grouping the samples to a continent level, it was found that 62% of the quantified subspecies are found in at least three of the four analyzed continents (Figure 2A). The African continent showed the highest number of specific subspecies (Figure 2B), which could be explained by the different host lifestyles. The three African countries: Madagascar10, Cameroon28, and Tanzania29–31, indeed show previouslyneglected Africa-enriched microbiome, further confirmed by a recent study32, with subspecies that were shared across those countries (FIG.8A).

[0070] To better summarize the biogeographical distribution of the subspecies, a single metric was defined that helped survey their enrichment. The prevalence of subspecies between all countries was compared and the median difference was weighted with adjusted p- values to get a Geographical Enrichment Score (GES ^ [0,1]) for each subspecies (see Methods for details). Calculated as such, GES represents the level of geographical restriction of a subspecies, where GES=0 shows no restriction, and GES=1 demonstrates high country-level specificity. To validate the metric, its correlation with median adjusted p-values (FIG. 8B) was confirmed for each subspecies obtained through cross-country pairwise comparisons. From there, it was concluded that GES=0.4 effectively distinguishes country-specific versus wide-spread subspecies. This threshold was used to split all subspecies into groups of high and low specificity. By looking at the distribution of median differences in prevalence obtained through the same comparisons as P-values, it was found that subspecies characterized with low specificity showed minimal median differences in prevalence compared to highly specific ones with a much broader distribution (FIG. 8C). Based on this threshold, 1869 subspecies were shared for many of the analyzed countries, and 338 subspecies were found specific to a certain region (FIG.2C). These country-specific subspecies accounted for a high percentage of Actinobacteria, Bacteroidota, and Firmicutes phyla (FIG. 2D). Firmicutes phylum contains the highest number of restricted subspecies (FIG. 8D), in agreement with a previous study13. By looking at the species whose subspecies showed the highest difference in GES, examples were identified where one subspecies is geographically restricted while others aren’t (FIG.2E). This represents novel information that would be missed when analyzing on both species and strain resolution. The highest difference showed Anaerostipes hadrus, where one of its subspecies was highly prevalent in Great Britain, Denmark, and Sweden, Korea, and the Netherlands, while the others were lowly but equally present in almost all analyzed countries. Collectively, these data suggest that unlike species, the subspecies within the HuMSub catalogue generally across different and distinct populations worldwide while maintaining specificity. Example 3 – Subspecies contain implicit information not available at species resolution

[0071] Colorectal cancer (CRC) is the third most common cancer responsible for the second-most cancer-related deaths worldwide33. To assess the practical use of the HuMSub catalogue, the inventors used microbiome samples from seven CRC studies34-40, totaling to 555 CRC samples and 530 healthy controls, and quantified them at the subspecies level. Out of the2800 identified subspecies, 218 were found significantly associated (abs(random effect size) > 0.2; FDR < 0.1). Out of those, 104 were instances where at least one subspecies was significantly associated with CRC, while at least one sibling subspecies was not (Figures 3A, marked in blue, 4A and Table 1). As a proof of principle, Fusobacterium animalis, previously associated with CRC development (doi.org / 10.1158 / 1940-6207.CAPR-16-0178), with two sibling subspecies, where only 001002 is found significantly increased in cancer samples (FIG. 3B). A recent study (doi.org / 10.1038 / s41586-024-07182-w) focusing entirely on this species experimentally demonstrated an existence of two clades, where only one populates the tumor niche and contributes to intestinal adenomas. Similarly, Porphyromonas asaccharolytica, also described as a contributor to CRC43showed subspecies 001002 as the highest CRC-associated subspecies, while its sibling subspecies 001001 didn’t show association with CRC (FIG.9B).jda_eulav_p_gol59retsulcjda_eulav_peulav_preppu_icrew1o0096722448627335440263 1 1 8 1 4 3 3 1 8l6301241 629 682 555 34 0 3 7 9 4 5 3 3 24774 0 0 2 7 8 0 0 7_ic 15.10.10..400- 1..400- 0.. .5. 3 0614.915.4.3.4.4.1100-0-0-.0.0.00-.00-0-0-0-0-.0ezi7s56 841 7 37 4_86995 1347 63 932 736381 56 62 8397985781 74t6 54997806565 004921 300074 05 20 7357889343 79c68f1 5 8 02 1 01 3 7878 7 2 5 66e9f 67 7 7 1 96 2 35 7 4658 2 3 2 06_57101 19 770334 054854 779163 06 55 2282462828 0e2n.3.3. 2. 2..02. 2.3.3. 4.2.3. 2. 3. 3.3.2.2.2.42a0 0 00- 0 - 00-0-0- 0 0 00- 00-0-0-0-0-.0emsei2c01 2 1 2 2e00000000 0401020202010302020401040302010p1s0100101001001001001001001001001001001001001001001001001001090 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0b2 6 7 5 3 1 9 1 8 1 5 4 9 1 3 6 0 8 5 2us 211 3 9 9 4 0 0 0 0 1 8 7 0 0 3 0 7 3 1 216110000020600052108320100073000010001071160045.21570268870.0507700.029650.0-25357104457011984086134299 91 13 52 1135784504238039483396486108654435549 21 93 3. 3505244 0 473.0.01.1..01.1..0.0.01.0. .0 1. 0- 1..0.4.5. 03.- - 0 0 - 0 0 - - - 0 0 0 0 -0-0-0-.00-3113 2113 61 4233 27449653 39 17 34 26444 3119 01 863 124 3 097 416 2 123 28 0 8 4116997280682706 91 4896382829235961 125 321563 313974 57 2114 42602267 2547 4270 2 8 4 54306 4 13620.30.5 921-0-.3 89 7 817922727 4 3 6 3. 13. 007023. 5 91 700.02.02-.20.04.20.2. 0 2-0-0-.0 0 3.02.02-.03.0-0-2.20. 2-0-.02.0-30201020102020201040101 2 3 2 4 4 2 4 2 1 60 0 0 0 0 0 0 0 0 00 0 0 0 0 0 0 0 0 0 01 1 1 1 1 1 1 10 0 0 0 0 0 0 0 0 0 0 00 0 0 0 0 0 01 1 1 1 1 1 1 1 1 1 1 1 1 10 0 0 0 0 00 0 0 0 0 0 0 0 0 0 0 0 0 0 01 1 1 4 60 0 0 0 0 0 0 0 0 0 0 0 0 0 0 05 2 34 0 2 8 5 5 5 8 1 1 0 7 1 1 3 5 07 19 7 9 3 2 7 1 4 8 8 5 5 1 6 7 9 1 6 70 00214004110411030908060407010505140106410759105.20400529180.0421800.079860.0-4529 4 9 1 5 3 5 0 3 0 6 7 0 0 6 2 6 2 9 9362 9 3 8 0 7 1 7 7 3 5 9 7 5 1 1 4 1 0 2741.426845379842345437314948393806348743 24 3.0.-0.-0.-0.-0- 0..00- 1..00.-0.-0.-0.-0.-0.-0- 1..00- 1..4. .000-0-0- -3449466047 44 25 68 22627641377678 48 16 61 1289 3 4 2 54157 8 1 2 42 39767 7 23 6 8 9 1 3 3 9 6 6 1028324155 15 6 9 5 8 828 279 9 2 495696 4 87 3 6 7 1 3 2 9 6833405953 2 8 15 8 1 8 9 310 277 98 79207 0 3 26 3 9 0 0 9 29094148 9 1 5 7 0 22 2 3 3 2213 2 1 522 0432. 6 5 5. . . .2.2.3.2.2 2 3 2 2 2 2 2 20 0 .0.0 . . . . ..0 ..00- . . .- -0-0-0-0-0-0-0-0-0-0-0-0-0-0-0-1040802010204010301 2 2 4 2 2 2 1 1 3 2 3 20 0 0 0 0 0 0 00 0 0 0 0 0 0 0 0 0 0 0 01 1 1 1 1 1 10 0 0 0 0 0 0 0 0 0 0 0 0 00 0 0 0 0 01 1 1 1 1 1 1 1 1 1 1 1 1 1 10 0 0 0 00 0 0 0 0 0 0 0 0 0 0 0 0 0 0 00 9 6 80 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 04 8 02 6 6 5 3 8 6 3 6 1 2 7 0 0 2 3 7 30 22 2 1 0 6 1 7 9 3 7 8 4 1 3 9 2 7 0 80 40010402100000081206170623061303280003131680924.53330783400.0901000.046202.0-648717788639 27 77765089818496878063041 6 4 962 .6 6 4 0 80.3745673 5. 97059423576714442907 0543 01160.-0.-0- 1..00-0- 0..00.-0.-0- 0.10..00.-0- 0..00- .4-0.-0- 25.. 0.00- 0532260 67 7037 78 501463 8633 79 2 59 88347192 332984 399 264 2 62 32 68 006 0614257391 631 60134 114 580 0047089 562 46 3 37694 010048565 03218639 4 93 1501608 7708 3558 80 123880 1675 998359 689 3 41 51 424.20.2.133-0-0-. 0232922.12 5 8 733 23202.3.-0-.03.20.2.-0-0-.0.02.0-0-.04.20.2.2.-0-0-0-.03.0-.020102020202010201010201 3 2 3 3 2 1 3 1 1 10 0 0 0 0 0 0 0 0 00 0 0 0 0 0 0 0 0 0 01 1 1 1 1 1 1 10 0 0 0 0 0 0 0 0 0 0 00 0 0 0 0 0 01 1 1 1 1 1 1 1 1 1 1 1 1 10 0 0 0 0 00 0 0 0 0 0 0 0 0 0 0 0 0 0 03 9 1 8 40 0 0 0 0 0 0 0 0 0 0 0 0 0 0 03 0 29 4 6 1 2 3 1 5 4 2 6 2 3 2 2 8 33 36 0 2 1 4 5 5 8 0 5 1 8 9 2 1 2 7 3 90 06121006291800341106100710220800041020071741239.495711125288875579 868 3 1 42 8 19 646 2 6 7 8707 062597391762 68178942050700000103000 0. 4 1 4 2. . . .0.0 0 0 0 0 0 0 00 0 0.0. . .0.0.0.081 471 5927363797432144 05332 642 0-3 5 9 1200 00 0.60E 080 080001 0417 7 8 10000003 0 3 100000000.0 0.0.5.0.0.0.0.0.0.00.0.017 10 5 471.3 57 2815 64283 3 74 9 3.0 247 8 8838 46836100-7 0.3001 3. 0- - . 3.001148 5330.. 8 7 800-840.0.00-.0.0.00-0-874391498147627 7 3 5 9 88 3 6 8 22 4 4 4 3 1437 5 6 45252484523595882.60- 04..3.3.909 1 8 4.2 3.3.00-0-0-.00.10.00.00- 1.00-0-34 28 19 58 1 3 1 3 189 7 58 5 413 2 1057177634 79710869 4 28083614 222383 54 4 07 8 92 8 7791 61 5674 24 6 1 10 906994231150522 85 902 02. 2. 2.2. .-02 2 2 22. 32.2.0- 00-0-.0.0.0.00-.00-0-301030201020101 2 1 2 1 40 0 0 0 00 0 0 0 0 01 1 10 0 0 0 0 0 0 00 0 010101010101010101010 0 0 0 0 0 0 0 0 009 1 0 3 5 1 8 70 0 05 0 2 9 5 14 9 0 3 07 5 0 05 3 0 2 5 7 41 1 0 0033434121221020080

[0072] On the other hand, 28 subspecies were associated with CRC where the parental species didn’t show association with disease, representing 26.9% of previously discussed associations (Figures 3A, marked in red, 3C and Table 2). Among these, the subspecies Ruthenibacterium lactatiformans 001003 was increased, while subspecies 001001 and 001002, together with the species level, were not significantly different in the samples from CRC subjects compared to those from healthy controls. (Figure 3D). Similarly, the subspecies denoted as UBA1691 sp900544375001002 was increased in CRC, while its sibling subspecies 001004 and the species level of this bacterium were not (FIG.9C). Contrary to this, some subspecies were less abundant in samples from CRC subjects. For example, Faecalibacterium prausnitzii is a species described as beneficial in ameliorating CRC. In the meta-analysis, two of its subspecies were decreased in samples from subjects with CRC, while the species level itself wasn’t (Figure 3E). This indicates that not only does subspecies level provide new insights into statistical associations with a disease, but that they could also improve reproducibility and to some extent explain discrepant results between studies. To further examine this possibility, univariate statistics were used for differential abundance testing within each study individually, both on the species and subspecies level, and found 31 species whose subspecies were significantly associated with CRC in more studies compared to the species level alone (FIG. 9D). This demonstrates that subspecies contain a layer of information fully neglected or missed when analyzing the microbiome on a species level and demonstrates that the subspecies are better suited for cross-study comparisons, providing associations with improved replicability between cohorts.jda_eulav_p_gol59retsulcjda_eulav_peulav_preppu_icre 0 3 7 6 9 7 6 8 6 9w 7 1 5 1 6 1 4 8 81 2 8 5o 1 2 3 1 1 2 0 9 069109 4 6l 0676470453514881139457 5 253 0 3_ic318.07.06.17.39.43.10.14.39.025 4 4.30.50 7.5 00 0 0 0 0 0 0 0 0 0 0 0.0.0ez6 6 4 2is 8 6 0 7 1930 1131 76 94 644816 23_ 6 5 0 48 04 2927 2 12t 4 0 3 71 18 8094 2 45ce9f6937946 2563 3914 22 75 023 1.3 3 33f5e2421233 120_.0.0.0.0 .3 50. 272 22 120- 0392 12-0-.0.0 .0-.0 .0.-0 .0n -aemsei2c0401030302010204010204 2 1e0p10s010101010101010100 0 01010 0 0009001005004001001001000005001 1 1500100 070 0b2 0 8 7 5 2 3 3 1 8 5 6315us1160201070100210308070501606499988955.2982442313770.010144700.07908360.0--9091-47-90-55-89-8235487 2 9 3 33 6 5 4 0 9 5 8 2959 8 6 92 5 3 3 3 2 0 8 9 8785 9 89 0 0 3 0 2 8 1 6 7 0812 82 7 7 7 5 1 9 5 1 4 2 6801149 4 4 9 5 2 7 4 4 9 5 665.00.50.30.30.30.30.00..00.-0- 0.40.00.40.093 464 273 9 4 4738983026234 261924 61 516 82900567 8 4 815 23 299955489363 8 98 317 9 1 1159 5 642 73 2 0 5228098123852 223 2 2 2 322.. 32 6.0 .0 .0.0.0.20.20 .00- .002.20 .0 .- - - - - - - -0-4 2 1 1 30 0301 2 1 1 2 3 2 30 0 00 0 0000 0 0 0 0 0 00 0 01 1 1 1010 0 0 0 0 01 1 10 0 0 0 0101 1 1 1 10 0 00 0 0 0 0 0000 0 0 00 0 09 6 3 8 2 3 2030 0 03 3 08 1 1 7 4 9 5 8554 21 9 22 2 0 8 3 8 4 1 01780 7 04 1 0 1 0 1 1 0 0 102 0 1 0Example 4 – Subspecies are better predictors of colorectal cancer

[0073] To further investigate the usability and practical importance of the findings that subspecies carry more informative gut microbiome profile stratification than species, machine learning was used to train classifiers for distinguishing samples from CRC and healthy subjects. A Light Gradient Boosting Machine (LGBM)46was used, an ensemble-type model designed to work well with sparse and high-dimensional data. To directly compare the information content of species vs. subspecies inputs, same-structured models were trained on two different levels: MetaPhlAn-based species-level genome bins (SGBs) abundances and subspecies-level profiles quantified using the HuMSub catalogue. To assess their performance, either all samples were pooled together (cross- validation) or the Leave-One-Dataset-Out (LODO) approach47was used. The former allows the model to train on some samples from all studies, thus enabling potential per-study adjustments. On the other hand, with the LODO approach, the model is trained on all studies except one, which is then used entirely for testing.

[0074] Samples from all seven up-to-date reported studies were used as in the meta- analysis, totaling 532 CRC samples and 501 controls. Importantly, species-level models achieved median performance similar to the current literature in this kind of setting (AUROC=0.79)39, establishing a credible baseline. During the cross-validation, subspecies-level models mildly outperformed species-level ones (FIG. 10A). Strikingly, by employing the LODO approach that simulates real-world scenarios, it was found that subspecies-level models (median AUROC=0.838, exceeding 0.890 for select datasets) consistently outperformed species-level ones (median AUROC=0.785) for all 6 studies (FIG. 4A), except when testing on VogtmannE_201640, which similarly resulted in lowest prediction rates as previously reported at species level39. These results suggest that subspecies resolution inherently contains novel information, which is more transferable between studies, explaining the increase in performance. Similarly, by training the models on only one of the studies and testing on all others, we observed that in two-thirds of study-to-study comparisons, subspecies-level classifiers outperformed the species-level ones (FIG. 10B). Together with an increase in negative predictive value (NPV), an important metric for screening-type tools (FIG. 10C), the increase in the performance of subspecies-level models suggests their superior predictive power, making them more suitable for clinical use.

[0075] Next, it was investigated if the performance of the classifiers would increase with more data being available for training. For that, the model was trained on a different number of datasets using the LODO approach and found that the subspecies level didn’t reach the plateau, whilethe species level started decreasing in performance by increasing the used datasets for training (FIG. 4B). Additionally, the difference between subspecies- and species-level performance persisted along the different training sizes, and a similar trend was observed when investigating per-study performance (FIG. 10D). To address the possibility of further increasing the predictive power by incorporating additional parameters, results of a fecal occult blood test (FOBT) available for study ZellerG_2014 were included36. Since this information isn’t available in other datasets, it was used to weigh the final model output (see STAR Methods for details). There, higher performance (median AUROC=0.893) was achieved compared to all other input combinations (Figure 4C), suggesting that inclusion of additional parameters can further improve the predictive performance.

[0076] To gain further insights into the improved performance of classifiers relying on subspecies- compared to the species-level ones, the importance of the individual subspecies learned by the trained models during the LODO approach were ranked and compared between the studies (Figure 4D).14 out of 47 total top-ranked subspecies (29.78%, marked in red) didn’t have their sibling subspecies ranked in the top 50 of any of the studies. This suggests that the differential abundance of sibling subspecies, not quantifiable at the species level, enables classifiers to extract a CRC-related microbial signature with improved specificity. Therefore, new models were trained using only the top- ranked subspecies to define the minimal microbial signature required for successful training. It was found that, for almost all studies, there was a plateau in performance at 64 top subspecies (Figure 4E), which can render classifiers more suitable and easier to interpret for use in diagnostics in a clinical setting. Example 5 – Sibling subspecies can differently affect host physiology

[0077] The existence of closely related groups of bacterial genomes, some of which are associated with a certain host phenotype while others are not, presents a unique opportunity to investigate potential functional mechanisms of the bacteria in host physiology. To investigate whether indeed the subspecies resolution could provide insights into the underlying mechanistic origins of the subspecies-phenotype associations, the annotated gene differences between the subspecies was looked at. Both potential gain and loss of function were considered, either by presence or absence of a complete gene, or by a destabilizing high-impact variant in one of the subspecies, effectively rendering that gene incapable of performing the specific function. Since subspecies groups in the HuMSub catalogue were initially defined based on variations in their gene sequences, this information was extracted by profiling the top ten subspecies positively associated with CRC (Figure 3C) together with their non-associated sibling subspecies and defining their functional differences. In accordance with preserving their evolutionary fitness, most of the detected SNVs didn’t have a significant predictedeffect on protein stability (FIG. 11A). However, there were many more cases of differences by high- impact (abs(^^G) > 1) gene variants, than by the subspecies-specific accessory genes (Figure 5A). These high-impact variants have been attributed to genes involved in biosynthetic pathways of nucleotides and amino acids (Figure 5B). Interestingly, differences in the type of the modules inactivated by the two mechanisms were uncovered: nucleotide and amino acid biosynthesis were mainly affected by high-impact SNVs (FIG. 11B), while gene specificity was responsible for differences in the synthesis of other types of molecules, such as metabolites and vitamins (FIG.11C).

[0078] The possible role of several specific subspecies in the development of CRC was then looked into. Ruthenibacterium lactatiformans 001003 was first looked at as the top associated subspecies where its parental species is not associated with CRC (Figure 3C and 3D). Between the CRC-associated 001003 and its sibling subspecies 001001 and 001002, 183 high-impact variants were detected of which 31 were located in 24 genes with annotated KEGG modules (Table 3). Five of those variants were found in four genes (K03394, K24866, K05936 and K02189) belonging to modules for aerobic (M00925) and anaerobic (M00924) cobalamin (vitamin B12) biosynthesis (Figure 5C). A highly destabilizing variant (abs(^^G) = 2.4) was found located at position 38 of gene sirC (K24866) in all protein sequences of 001002 and some 001001 subspecies, which was absent in 80% of 001003 sequences (Figure 5D). Similarly, another high-impact destabilizing variant (abs(^^G) = 2.7) was located at position 224 of gene cobI (K03394) in all sequences of non-associated 001001 and 001002 subspecies, not present in over 80% of 001003 sequences (Figure 5E). An additional disruptive variant (abs(^^G) = 1.65) at position 5 of gene cobM (K05936) was found in all sequences of the subspecies 001001 (FIG.11D), and finally, two destabilizing variants were found at positions 195 (abs(^^G) = 1.03) and 252 (abs(^^G) = 1.0) of gene cbiG (K02189) of several 001001 sequences (FIG.11E). These data indicate that unlike the CRC-associated 001003, the non-associated subspecies 001001 and 001002 have lost the ability to produce vitamin B12 in aerobic and anaerobic conditions due to these highly disruptive variants. Interestingly, higher concentrations of B12are measured in the serum of CRC patients, and these B12levels are positively associated with the stage of the cancer. Higher B12levels are associated with lower methylation levels of retrotransposable long interspersed nuclear elements (LINEs), both in tumor tissue and surrounding peripheral blood mononuclear cells49. LINEs are preventive biomarkers of CRC50. Supplementation of this vitamin increases the risk of CRC51, suggesting a potential causative effect of higher B12 levels on CRC progression.Pe3- ni ,erec niirebroItacIInirehtao_r eS G K A R cnelrelelfaA P E V K er_ecn905 7eeE r n efe G05 2_E3R1. 7V0 3E 09eg _ _ 5_ G19 E E1 G143T M5 1 L_ 187Q_11._ _ 3T4 1_E_ 5T3 1r_eU O4G N6_1Z94N52Z0 6M0 _M5 _1N K G00U O G N75U O G N15ec0101010 0nepr _uo003 0 01010003030303efr 4747 047 047 047 0er g1010101010itis7on 102 648po 3 2 622 3P3P n e e-ie,dyyahwehdtlaaprefcfyolrgod>u =oPD-6r-eenstoncuElg0000 8M sanotcalonoculgohps- o6 h p40470K 8 -883.2N P E M T K S S A V A R X0 382_1E324 E32_ E016 E016X00 3R. 7 GE6_GE5 GE4 GE4J0 _L1 7 _8 9_ _30 _ _ 4 _ _ 4_1 1_8 4T M_T M1T M3 2T M3 2Z0 .Z9 2U O3U O6 1U O6 _U6 _N K10N51G N40G N69G N86O G N86010101010 00 0 01 104 270 04 370 04 3 070 04 3 0004 3 0004 30117 7 70 01010101043 1 4 987162 57224sisehtnysoibm A NcuF- L-P D U2900M -4-luxeh-onibara-L-ateb86091K 9253.1VI A R V YI00 20 1E5284E528 2_1V 00 35U001 _0 1.2G_E_GT M9 5_E4_R. 29 5 L1 7Q01 _1 7V 01L5U_T M_ _8O4U O4Z992 _Z0._K1Z0JG00G N08G N08N51N G00N N010010101010104 3 003 003 003 00007 047 040404 304 101 171717 70 0 0 0101014309 5203 39121 341D R S D V G K A N 222 _ R_1.E9E62_1 G05 26_ L1 5 _E3_G1_ 551_E4_R1. 61 6 L1 3367 _8972T M4 _T M6 2 _89729Z N51U O G N61U O G N65Z N5101010103 00004 30010007 04 304 304 301717 70 01010543 3 8462862421ene e-it tnoihtem,sisehtnysoibenietsyC0600Mtnedneped26471K 8260.1YD A R K A K 00 80 1E91 _0 1.3G06 _ E3E9_G063E7141V 00 11X00 56T M851_E_T M85G1_E_T M8 6 Q01 _1 0XJ 00 _L4U_ _O8 _U8 _U0Z0.L3 _Z101.7H00G N81O G N81O G N92N H00N K1010101010 0004 3 003 003 010010007 047 0404 304 304 201 171717 70 0 0 010103159255 3 91 1 2835191-non>,=yaPw6h et sa op tcetuarfh,pessoahhppeesvoittnaediPxo0000Marefsesnaatrtcl ]ue y 0d1er ]sa ]sl 2o .2.e8o .d1.bir 4t.an5.1.la 2. o22et h:pCoru 1.sn: aarCtosEt1:tE[ro[ aohpes gaC tE[612661407 0200 04K0K0K 4479231.44185.722211 1.1VPL G L P 024 E2 1_8 42H_X01X0 56M6 1 J.8 20J_00 _1O0 _ _ 9Z10 .7N76Z N64N K1010 0010104 370 04 1 070 04 201170 01003 140103

[0079] It was next investigated whether such mechanistic explanations are also observable with subspecies that show lower levels of association with CRC. The functional potential of the Enterocloster sp005845215 subspecies 001002 that was increased in CRC samples was compared relative to its sibling non-associated subspecies 001001 (FIG. 11F) and 108 gene differences were found, from which 11 could be linked to 10 KEGG modules (Table 4). Among those genes, argH (KO K01755), which is a part of the biosynthetic pathway of the amino acid arginine (M00845), showed a highly destabilizing variant (abs(^^G) = 1.3) at position 447 in the non- associated subspecies 001001 (Figure 5F). Interestingly, both asymmetric and symmetric dimethylarginines are higher in patients with CRC52. Moreover, increased dietary intake of arginine is associated with the increased incidence of CRC53, suggesting a causal relation between the increased arginine levels and CRC. Indeed, these previous findings led to development of arginine-degrading enzymes as a new class of anti-CRC drugs54. As another example, subspecies Porphyromonas sp900548415 001002 that was increased in CRC was compared to its CRC non-associated sibling subspecies 001001 (FIG.11G), and 130 different genes were uncovered, 38 of which being annotated to 21 modules (Table 5). Among these, the gene hemH (KO K01772) that belongs to one of the pathways for the biosynthesis of heme (M00926), had two destabilizing variants in positions 173 (abs(^^G) = 1.4) and 227 (abs(^^G) = 2.0) within the non-associated subspecies 001001 (Figure 5G). Additionally, the gene gltX (KO K01885) was similarly destabilized (abs(^^G) = 1.2) by a variant at position 74 (Figure 6H) in the same subspecies, which is part of another pathway of heme synthesis (M00121). Increased dietary intake, as well as the microbiota-produced heme, is linked to gut epithelial hyperproliferation55and chronic intestinal inflammation that critically contribute to the development of CRC56, suggesting that the increased abundance of subspecies 001002 could be mechanistically linked, through heme production, to favorable conditions for CRC development. Together, these data indicate that at the subspecies resolution, microbiome profiling reveals an additional layer of information necessary for in-depth functional analysis, thus providing new insights into the underlying origins of microbiota-phenotype associations.csed_eludom ludom oitcnufehto_r eS K R L cnelrelelfaQ R S H er_eegO2O3O6O6_ N_c9_ _ _e N E8E58N E19N58neernG_06654_ G0 1_ G 76 0E 3G 4026_efT50 _T79 _1_ _70e U0E U0T E U1T E U0E r G M G M G M G M ecp01010 0neruo0 01010r010 010 0101efeg616 6060r _31131 113131itis n4300528o o 10p3 1 1,sisehtnysoibenisPyA L D 500M esaniketatrapsa-VIS Y O2_Y.8O1O2N E1 1N_ _9M04N9G_76 2_D0063E2G03 2E8G0604T11 9AJ 00_1 __T51 7 _T5 _2U E_Z10U E U0E G M N D G M G M 010 0 001 1 10000006 1106 1106 210 06 10311313131936 01222239,sisehtnysoib800M etaniccusoninigra o o a hC aylm mesmedED [ UpeE[55075372 oit487-cnlK R TerelelfaN er_eO1_O2_.6eg 1N E0N M0E3 _O_1 501N0eN0G_501 5 1D043_4 G_ 0 _A01c01neeE3nG 40T U2E6T3U23E0J_ 0_Z11 r _222M G M0efT E_G N DerU G M78010010 ecp0106 2 0102 002ne uo0011061 061 05refrg49 0313131elb e_1ar 4T 6 i 35 3ti59392son 6po 3-esoculg-D-ahpla> =P1L RIN R O4O3O2O2O2N_72N_72N_9_9N1_9N E E E E E19G92G 92G 8 G 8 G 8_9_9_ 34_ 68_ 6T0 3T0T20T0T80U E_U E7_6U E1G M_5U E0G M2_U E0G M6G M G M2_0101010 0000 010104 290 04 290 04 202029049 0401414141941414833104 221 482stnalp,nsiivsealhftonbyisoRib100M / esaniknivalfoNbirM F -R S M V D O1_O8_O1O4O4N E74N4_4 E0N4N_8N_86E0E5E5G6 6 6_09 G_ 7G_ 7G_ 7G_ 7T02T21T211 3131U E_U E3U E_T E7T E7G M71G M_8G M13U G M_6U G M_601010 0 00 01010104 20 0 0009 04 104 104 204 20191919 94 4 4141401942 1 53 2130292-esoculg>=negocylgP6R A S O1_O1O3N E05N_8N_45E56E0G T9 G G 6_0 7 701_31 _21U E_T34U E_T 4 U E7G M2G M7G M_7010010104 2 002 009 0404 1019194 41422030327S S 3_O340N_46E072 G 671 _2E7T1_ U E7M7G M_70100104 1 001904019414504191>=D P ATF / GN,aiMreF / tcnia vb ald fn oa bir-5 / -es 5a (-ni 6-moaneidmaN V O3 7N_E4O_0N76E2G7G 9_2_ 2T190U E7T _ U E3G M7G M2_0100104 1 009 04 2019414473231A A S L T V R 7_7O1_O1O32 N5N_5N_9 E3 38423E E 9G0_ 68 G 3 5_6G_ 9E3T282U E3T281U E3T0U E7M_G M_G M1_G M_0010 02 01 10 004 200904 20 04 2019 941414827 32953132>=ePtaTloGf,osridsyehhtarntyesoTib100M62etaoretpor esdayhhtindys69700K 35-2103.1RG M T E E A Q O2N_O1O1O6O6O4E5_3N0N_8N_6N_ _3 E5E4E262N196G 5E EG0G 59G 62G 62G 8_T8_ _ 1 _ _ _ 629U E2T08_E8T01 9_ T09 82 T0 2T0314U G M_9U E G M60U E U E U E G M G M_7G M_7G M_6010 001 1010101000000 0 04 2904 2904 20 04 20 04 20 04 2019 9 9 94141414141435 46 0 1 31 761729471,sis-elyhtmn ay ts uloi gb ,eairmeetcHab900M oc / nirynhiprryohpportooproprp-K C R M T S S O4 7 1N_O O O41_9N7_2N8_5N E48E9E E0GG G 6 G 8_6 _ 2 _ 7 5T80E31T9039 T15_8_ T2U U E U E U E G M_6G M1_G M66G M9_010 0 001010104 290 04 290 04 2 090 04 201194 41414720 5 72319292eedintiodeilcmiurnyoPbir000M esahtnysP T C -R D D G A A Q N O1_O1_O1_O1_O1N74N40N40N4N_E E E E0E40G40G 85G 85G 85G 8_ T9 6_ _ _ _ 508_ T21 8_ T21T821T821U E9U E U E_U E_U E_G M1G M79G M79G M79G M7901010101004 2 010 02 002 00009 049 0404 204 2011919194 4 4 4149462854 51 2 60757Example 6 – Methods

[0080] The following methods were used in Examples 1-5.

[0081] Genome filtering. The HumGut collection was quality filtered using a custom workflow. Briefly, genes from all genomes were called with prodigal57, which were used as input for a two-fold filtering approach. The first fold was the detection of contaminated genomes, where GUNC19was used with GTDB reference and kept only genomes marked as ‘passed’. The second fold was the completeness and contamination assessment using BUSCO20. To best utilize the BUSCO approach, a provisional run was done on a set of genomes containing one representative per species-level cluster in order to get their BUSCO taxonomic clade. Then, that information was used to run BUSCO on all genomes, now forcing the run on the previously selected taxonomies. This was done to prevent incomplete genomes from being evaluated on the lowest taxonomic clade, thus potentially obtaining an artificially high completeness score. The quality of analyzed genomes was assessed using the following formulas: ^^^^^^^^^^^^^^^^^^^^^^^^ ൌ ^^^^^^^^^^^^^^^^^ ^ 0.5 ∗ ^^^^^^^^^^^^^^^^^^^^^ / 100^^^^^^^^^^^^^^^^^^^^^^^^^^ ൌ ^^^^^^^^^^^^^^^^^^^^ / 100^^^^^^^^^^^^^^^^^^^^^^^^ ൌ ^^^^^^^^^^^^^^^^^^^^^^^^ െ 5 ∗ ^^^^^^^^^^^^^^^^^^^^^^^^^^ െ ^^^^^^^^^^^^^^^^^^^^^ / 100^and only genomes with QualityScore > 0.7 were kept for further use.

[0082] Clustering. Genes called during the quality-filtering step were relied on as the input for sketching using Sourmash21with the “sketch” command and parameters “-p protein,k=51,scaled=1”. The output of the sketch command for all genomes that passed QC and belonged to one HumGut_95 cluster was further processed together using the sourmash compare command with “--ksize 51 –protein --ani” flags. This command produces a similarity matrix, where values represent ANI estimates from protein sketches, which differs from the standard use of ANI. This similarity matrix was then transformed into a distance matrix (1 – similarity_matrix) and used for hierarchical clustering, where the input genomes were randomly attributed to training and testing sets. We then performed clustering on these sets using the single linkage method, generating linkage matrices that describe the cluster formations. For possible cluster numbers up to 15, we assigned cluster labels and calculated prediction strengths by comparing labels between training and testing sets. The optimal number of clusters was determined based on the mean Prediction Strength across 15 iterations, selecting the highest cluster number that exceeds a predefined cutoff of 0.8, as suggested in the original publication by Tibshirani and Walther(doi.org / 10.1198 / 106186005X59243). The maximal number of clusters was data-driven, where out of 977 clustering events, there is a median of 2, a mean of 2.93 and a maximum of 12 subspecies. Finally, hierarchical clustering was applied to the entire distance matrix. We named subspecies as follows: [cluster95]001[subspecies_number], where cluster95 is the original species-level cluster ID from the HumGut catalogue and subspecies_number was attributed during clustering. If no subspecies were detected within the input species-level cluster, all genomes were labeled as [cluster95]001001.

[0083] Definition of Subspecies-specific hashes. Sourmash21combined with custom scripts was relied on to generate subspecies-specific panhashomes. Briefly, “sourmash sketch -p dna,k=51,scaled=1000” were used on genomes annotated with subspecies information to obtain their sketches. Then, using a custom Rust script, all sketches were parsed per species- level cluster and found all hash values specific to each subspecies. To be selected in this embodiment, the hash values needed to be found in more than 20% of the respective subspecies’ genomes while not present in more than 5% of all others. However, it is believed that in some applications the hash values needed to be found in more than 10-95% of the respective subspecies’ genomes while not present in more than 1-9% of all others Finally, the specific hash values were compiled into one signature per subspecies and then indexed with sourmash index.

[0084] A custom workflow was generated to leverage information about specific hashes. There, sequencing reads from input metagenomic samples are sketched using the same parameters used to generate subspecies-specific panhashomes. Then, the gather58algorithm from sourmash was relied on to quantify subspecies-specific hash values from the index. The default sourmash output is then summarized to produce the final table of relative abundances.

[0085] Quantification benchmark. CAMISIM59was used to simulate ten metagenomic samples through an Art simulator60with default error profiles. Each sample was simulated using the mean fragment size of 350 and SD=30, for a total of 3 GBp per sample. For samples simulated using HumGut genomes as a source, 15 genomes were randomly selected per subspecies or all genomes from a subspecies if it is represented with less than that. Then, CAMISIM was used to further randomly sample 1500 genomes from the selection to represent subspecies with a maximum of three genomes. Log distribution was used for the abundance profiles, with a mean value of 1 and SD=2.75.

[0086] To further validate the quantification approach for real-world use, a new set of metagenomic samples was simulated, now using genomes uploaded to RefSeq since 01.01.2022, after the publication of HumGut. The list of all genomes was obtained from ftp.ncbi.nlm.nih.gov / genomes / refseq / assembly_summary_refseq.txt, as of 16.02.2024. Further, the list was filtered to keep only genomes with NCBI taxonomic IDs found in HumGut. Then, thequality filtering workflow described herein was run and genomes that passed GUNC combined with BUSCO QualityScore > 0.8 were kept, which included 2432 genomes. Those genomes were annotated with subspecies information. There, the similarity was calculated between input genomes and subspecies-specific hashes in a pairwise manner, calling the final subspecies assignment if the similarity score is higher than 0.8. With this, 324 genomes ended up used for metagenome simulation. Finally, sequencing samples were simulated using the same parameters as for HumGut genomes, now including all input genomes.

[0087] In order to further give context to evaluation metrics, MetaPhlAn41with the vOct22_CHOCOPhlAnSGB_202212.1 database was run on the same simulated samples. To match the abundance estimations with the ground truth, all abundances estimated at the species level were looked at and their NCBI taxonomic IDs extracted. That was used to match them to species-level taxonomic IDs from the input genomes while ignoring abundances estimated at levels higher than species. Abundances were summed from the quantification to obtain species- level abundances across subspecies from the same species.

[0088] For evaluation metrics, F1-score and Euclidian (L2) distance were used. For all quantifications, an abundance threshold at 0.001 (0.1%) was used, thus ignoring all taxa simulated at a level lower than that. The F1-score is calculated as a harmonic mean of the precision (the ratio of true positive calls and total positive calls) and recall (the ratio of true negative calls and total negative calls). On the other hand, L2 distance was used as a measure of discordance between ground truth and estimated abundances using the scipy.spatial.distance.euclidean function in Python. The tracking of used computational resources was done through the benchmark directive within the SnakeMake workflow. There, each sample was analyzed with MetaPhlAn within a single rule, while sourmash quantification included two – one for sample sketching and the other for the ‘gather’ command. The total resources for sourmash quantification were summed between those rules. The computational time was found in the ‘s’ field, while the total memory required was in ‘max_uss’.

[0089] Subspecies biogeographical distribution. Mastiff27, a sourmash-based tool, was used to analyze the geographic distribution of subspecies. It enables querying the pre- generated sourmash index including public human metagenomic samples. As a result, a table including containment estimates for each input query is returned.0.2 was used as the containment threshold for inferring subspecies presence / absence and subsequent prevalence calculations, meaning that a metagenomic sample must contain at least 20% of subspecies-specific hashes to be declared positive for that subspecies. Since the index was generated with the parameters “k=21,scaled=1000”, the subspecies-specific panhashomes were regenerated to match those requirements. Since the resolution of sketching was effectively lowered, specific hashes for 3038subspecies were able to be found. CuratedMetagenomicData61was used to obtain metadata of metagenomic samples. Here, filtered samples were filtered and kept only ones originating from fecal samples from healthy adult individuals. Further, to reduce the noise from countries with fewer samples, only those with more than 30 samples in total were kept. With this, 5272 metagenomic samples from Europe, Asia, North America, and Africa were ended up with. For prevalence calculations at the country level, each country subspecies was called positive if at least 10% of its samples were positive.

[0090] Public metagenomic datasets. Sequencing data was downloaded from seven datasets34–40from seven countries in total. The samples were filtered and only ones labeled as “CRC” or “control” in the “study_condition” column of metadata tables obtained from curatedMetagenomicData61were kept.

[0091] Sequencing data preprocessing. All downloaded data was preprocessed using the ‘QC’ module from the ATLAS v262with default parameters. In short, using tools from the BBmap suite v37.78, sequencing reads were quality-trimmed, and contaminations from the human genome were filtered out.

[0092] Machine learning. The Light Gradient Boosting Machine (LGBM) classifier46was used, implemented in the lightgbm Python library. The models were constructed with 500 estimators, using ‘gbdt’ as the boosting type, ‘binary_logloss’ as the loss function, 0.01 learning rate, and 11 leaves.

[0093] Here, the same subspecies-level quantifications were used as for meta- analysis. However, species-level abundances were obtained using MetaPhlAn4 and the vOct22_CHOCOPhlAnSGB_202212.1 database with default parameters. Before training, all features found at the abundance <0.00001 in more than 85% of samples were filtered out, to ensure that we only keep the highly prevalent and not very low-abundant ones. This filtering approach enabled us to reduce the “Course of dimensionality” and eliminate the irreproducible features found only in a subset of samples. This means that for the LODO approach, we filtered features separately for each combination of studies included within the training. With these criteria, on average, 63.5% of all detected subspecies were filtered out.

[0094] Two approaches to estimate the performance of the models were used: on all datasets together and the Leave-One-Dataset-out (LODO) approach. For the first one, all datasets were combined together and tenfold cross-validation was used. For the latter, the LODO approach was used to simulate the real-world use case - four datasets were used for training, while the fifth one was used for testing, which was repeated five times. For all of the approaches, the model was reinitiated and, where applicable, the cross-validation was reinitiated ten times with different random states. The importance of features was assessed using the build-it‘feature_importances_’ method of the LBGM classifier. To include information from FOBT, the model was trained using the LODO approach but increased the output value by 0.2 if the test sample was positive in FOBT.

[0095] Identification of subspecies-specific genes. To guarantee that genes that finally resulted in the differential abundance between subspecies were found, subspecies-specific panhashomes were used as a starting point. First, the minimal set of input genomes was found where those hash values could be found. Then, those genomes were combined and sourmash signature k-mers -k 51 were used to translate them into k-mers. Next, prodigal-derived gene calls obtained during the initial quality-filtering step were used to find the corresponding complete sequences from which the specific k-mers were extracted. As described herein, k-mers may include between 8 and 20 nucleotides or amino acids, between 8 and 15 nucleotides or amino acids, between 8 and 12 nucleotides or amino acids, or preferably between 8 and 10 nucleotides or amino acids. Then, MMseqs263was used to cluster the extracted genes based on their sequence similarity with the threshold at 0.9. Further, those representatives were annotated using DRAM64. If two or more genes were annotated with the same function, those clusters were merged and considered together. There could be two mutually exclusive scenarios: a gene as a whole could be defined as present in only one of the subspecies, or it could be present in multiple subspecies but with subspecies-specific variants. If the first possibility was true, then it was reported the whole gene as subspecies-specific. Alternatively, the multiple sequence alignments (MSA) with muscle65were performed and defined the position of variants found in more than 80% of certain subspecies while not found in more than 20% of all other subspecies. Then, PROSTATA66, a transformer- based tool, was used to predict the effect of a variant on the protein stability. All variants with abs(^^G) > 1 were considered as high-impact variants. For MSA visualization, pyMSAviz Python library was used. Quantification and statistical analysis:

[0096] Subspecies biogeographical distribution. To calculate if a subspecies is significantly differently prevalent between countries, a robust approach based on the Chi-Square Test of Independence (implemented as scipy.stats.chi2_contingency in SciPy67Python library) was used, followed by the Benjamini-Hochberg procedure for adjusting p-values, effectively controlling the false discovery rate. If the adjusted p-value for a species was lower than 0.05, the post-hoc Tukey's Honestly Significant Difference (HSD) test was used to identify which pairs of countries were indeed different in prevalence. To summarize the information obtained through statistical testing, it was sought to define a single metric, geographical enrichment score (GES) that would enable them to distinguish geographically restricted subspecies. It was defined asGESsubsp=[(abs(median_difference) + log(1 / (median(adjusted p-value) + 0.0001))) / 2]. Finally, the GES was scaled across subspecies to [0,1].

[0097] Meta-analysis. Relative abundances obtained through the quantification workflow were used and arcsine-square root-transformed them. To adjust the transformed abundances for BMI and age, linear models were created and fitted to the transformed abundances of all samples within each study individually, adjusted for disease presence, age, and BMI (abundance ~ study_condition + age + BMI). Then, the inventors used them to predict adjusted mean transformed abundances per study, which further served as input. For statistics, the metafor68R package was used. The escalc function was used to estimate the effect sizes for each species and subspecies separately by calculating the standardized mean difference (Cohen’s d). Then, those estimates were used with the rma function to calculate random effect sizes and the corresponding p-values. Finally, those p-values were corrected for multiple testing using the Benjamini-Hochberg procedure with the cutoff at 0.1.

[0098] Machine learning. On each taxonomic level, models were generated using different random states to minimize the effect of the random model initialization on the performance estimation. To compare model performances, the Mann-Whitney U test was used with the Benjamini-Hochberg procedure to correct for multiple testing where applicable. Example 7 – Subspecies Analysis Improves Prediction of Adenoma Status

[0099] Metagenomic sequencing data from patients diagnosed with adenoma was analyzed, which can generally be considered an intermediary step toward colorectal cancer (CRC). These data represent a subset of data obtained but not included for the analysis of CRC. In total, 136 samples positive for adenoma and 512 healthy controls were analyzed, both on the species and subspecies levels. Subspecies-level abundances are obtained through the use of the HuMSub catalog, while species-level ones were independently estimated using the established MetaPhlAn tool that is available through the curatedMetagenomicData package.

[0100] LGBM models with the same architecture as for CRC prediction described above (1000 predictors using ‘gbdt’ as the boosting type, ‘binary_logloss’ as the loss function, 0.01 learning rate, and 11 leaves) were used.

[0101] Machine learning (ML) was used to build models to directly predict adenoma status. To evaluate these models, we used five-fold cross-validation (CV) initiated ten times. There, the subspecies-level models (median AUROC 0.702) outperformed species-level ones (median AUROC 0.662) (FIG. 12). These results demonstrate that additional useful information exists at the subspecies level that is not available at the current alternative resolutions. These results support the idea that subspecies level analysis in combination with machine learning can be used to provide improved adenoma diagnoses.Example 8 – Subspecies Analysis Improves Prediction of Responses to Immunotherapy in melanoma patients

[0102] Microbiota may affect response of melanoma cancer patients to immunotherapy. To analyze its contribution on the subspecies level, six publicly available datasets containing metagenomic sequencing data from patients diagnosed with cancer and treated with immune checkpoint blockage therapy were analyzed. The patients were grouped into “responder” and “non-responder” groups based on the progression of the cancer after the induction of the therapy. Since species-level abundances for all studies are not available through the curatedMetagenomicData61package, these were quantified using the MetaPHlAn4 tool with default settings. The HuMSub catalog described above was used to obtain subspecies-level abundancies. In total, 339 responder and 202 non-responder samples were analyzed. Within the PCA space on center-log-transformed abundances, the subspecies level achieved better separation of responders vs. non-responder samples, with the normalized Euclidean distance between medioids of 6.94, compared to species resolution’s distance of 4.50 (FIG.13A). Using the meta- analysis approach with adjusting for gender, five subspecies were identified that were associated with the response but whose parental species are not associated with response (FIG.13B). Thus, the subspecies analysis identified useful data for predicting responses that were not observed at a species level. For example, only Bilophila wadsworthia 001001 was increased in responder patients, while its sibling subspecies 001002, as well as its parental species, were not (FIG.13C). Conversely, subspecies Ruthenibacterium lactatiformans 001003, previously detected as increased in CRC patients (FIG. 3D) is now increased in non-responders, while its sibling subspecies and parental species aren’t (FIG. 13D). These associations indicated that subspecies provide a more useful resolution to analyse the microbiome, which was further exploited using ML. For this, the proposed subspecies resolution was compared to the current state-of-the-art in the field described by Gunjur et al (doi.org / 10.1038 / s41591-024-02823-z). There, it is proposed that the strain level, together with the stratification of samples based on the immune therapy applied (anti-PD1 and combined anti-PD1 / anti-CTLA4), increases the predictive power of ML models in discriminating responders vs. non-responders. However, even though the proposed strain reference was delineated using samples from the study included in the analysis, which can artificially increase its performance, the subspecies level consistently outperformed it (FIG.13E). Similarly, subspecies also outperformed the species level which was used as control (FIG.13E). Finally, the performance of the models was investigated by increasing the number of studies included during training. Species-level analysis didn’t increase the performance with additional studies, owing to the lower reproducibility between studies observed at the species level as additional studies were incorporated into the analysis. In complete contrast to the species-levelanalysis, when utilizing the subspecies-level analysis, additional studies continued to increase the performance (FIG.13F). These results demonstrate that subspecies analysis surprisingly resulted in the exact opposite (improvement instead of worsening) predictive power as compared to species-level analysis, as data from additional studies are incorporated into the analysis. As described above, strain-level analysis can result in serious problems and complications with microbiome analysis. These results demonstrate that subspecies-level analysis can improve the analytical and predictive power of microbiome analysis. These results also confirm, using a newly identified phenotype that was generated using an independent set of datasets (data sets from multiple independent studies), that the subspecies resolution provides surprising and unexpected advantages for microbiome analysis, including the use of microbiome analysis in diagnostics and therapy.

[0103] Meta-analysis: Relative abundances obtained through the quantification workflow were used and arcsine-square root-transformed them. To adjust the transformed abundances for gender, linear models were created that were fitted to the transformed abundances of all samples within each study individually, adjusted for disease presence, age, and BMI (abundance ~ response + gender). Then, they were used to predict adjusted mean transformed abundances per study, which further served as input. For statistics, the metafor R package was used. The escalc function was used to estimate the effect sizes for each species and subspecies separately by calculating standardized mean difference. Then, the inventors used those estimates with the rma function to calculate random effect sizes and the corresponding p-values. Finally, those p-values were corrected for multiple testing using the Benjamini-Hochberg procedure.

[0104] Machine Language (ML): The ML approach described above was used with the following modifications. The LightGBM classifier hyperparameters were optimized using Optuna with a five-fold stratified cross-validation scheme. The search space included num_leaves, learning_rate, feature_fraction, reg_lambda, and n_estimators. During each fold, the training set was oversampled using the RandomOverSampler from imblearn to address class imbalance. Missing values were imputed with zeros. The optimization objective was the average area under the receiver operating characteristic curve (AUROC) across the validation folds. The best hyperparameter set identified by Optuna was subsequently evaluated using stratified 10-fold cross-validation repeated across 10 different random states. For each fold, the training set was again oversampled, and missing values were imputed. Models were trained using LightGBM with the tuned parameters and evaluated on held-out test data.

[0105] For melanoma datasets, the same subspecies-level quantifications were used as for meta-analysis. However, since not all species-level abundances were available through curated MetagenomicData, we quantified them directly using MetaPhlAn4 with defaultparameters. Before training, all features found at the abundance <0.00001 in more than 85% of samples were filtered out. Example 9 - Subspecies Analysis Improves Diagnosis and Prediction of Inflammatory Bowel Disease (IBD)

[0106] Metagenomic sequencing data from 81 samples positive for inflammatory bowel disease (IBD) and 221 controls were obtained79. Using a subspecies analysis approach as described above, the species-level abundances were obtained through curated MetagenomicData, while we quantified the subspecies-level ones using the HuMSub catalogue.

[0107] Because of IBD's high effect on microbiome signature, it was hypothesized that ML-based prediction models trained on subspecies-level information could be used to predict the data from metagenomics sequencing data. Similar to the above, five-fold cross- validation (CV) initiated ten times was used and demonstrated that subspecies-based models (median AUROC 0.920) outperformed species-level ones (median AUROC 0.856) (FIG. 14). These results again demonstrate that subspecies level analysis can provide superior power for microbiota analysis as compared to a species level analysis, and this subspecies analysis avoids problems associated with strain level analysis.

[0108] Methods used for subspecies analysis were the same the methods described above for CRC and adenoma analysis, with the following modifications: : LGBM models with the same architecture as for CRC and adenoma prediction (1000 predictors using ‘gbdt’ as the boosting type, ‘binary_logloss’ as the loss function, 0.01 learning rate, and 11 leaves) were used.

[0109] FIG.15 is a flow diagram of a method for creating a dataset of subspecies, according to some embodiments. Method 1500 may be performed by any suitable computer system, such as that described below with respect to FIG.17. Method 1500 has many variations, including those described below.

[0110] Method 1500 includes, at 1510, accessing a catalog comprising coding sequences from a plurality of microbiota species. Examples of catalogs are described herein. In some embodiments, the coding sequences include one or more k-mers. The k-mers may include, for example, between 8 and 20 nucleotides or amino acids, between 8 and 15 nucleotides or amino acids, between 8 and 12 nucleotides or amino acids, or preferably between 8 and 10 nucleotides or amino acids.

[0111] Method 1500 continues, at 1520, with generating sets of hash values corresponding to coding sequences of genomes from each of the microbiota species. In some embodiments, generating the sets of hash values includes employing a hash function to sketch thegenomes of the microbiota species. In some embodiments, the catalog is filtered to remove contaminated or incomplete genomes before generating the sets of hash values.

[0112] Method 1500 continues, at 1530, with clustering the hash values to identify one or more subspecies in the plurality of microbiota species. In some embodiments, clustering the hash values includes clustering in an unsupervised manner. Clustering in the unsupervised manner may include using a Prediction Strength metric with a cutoff between 0.5 and 0.95.

[0113] Method 1500 continues, at 1540, with assigning groups of the hash values to a particular subspecies by selecting the hash values that appear in a range of at least 10-95% of the genomes of said particular subspecies, but in range of less than 1-9% of the genomes from different subspecies. In some embodiments, assigning groups of the hash values to the particular subspecies includes selecting the hash values that appear in at least about 20% of the genomes of said particular subspecies, but in range of less than about 5% of the genomes from different subspecies.

[0114] Method 1500 continues, at 1550, with generating a dataset describing the subspecies and corresponding assigned groups of hash values to the subspecies. In certain embodiments, the subspecies includes multiple strains of microbiota. In some embodiments, the dataset describing the subspecies includes an identification of one or more genes describing functional variation between subspecies.

[0115] FIG.16 is a flow diagram of a method for analyzing a microbiome bacteria sample according to subspecies, according to some embodiments. Method 1600 may be performed by any suitable computer system, such as that described below with respect to FIG.17. Method 1600 has many variations, including those described below. In some embodiments, one or more aspects of method 1600 may be combined with one or more aspects of method 1500, described above.

[0116] Method 1600 includes, at 1610, receiving, at a computer system, genome data from in vitro sequencing analysis on a microbiome bacteria sample. Method 1600 continues, at 1620, with analyzing, by the computer system, the genome data according to a library of microbiome bacteria subspecies to determine a presence of microbiome bacteria subspecies in the genome data, wherein the library of microbiome bacteria subspecies includes a listing of microbiome bacteria subspecies, a particular subspecies being a grouping of hash values defined as including the hash values that appear in a range of at least 10-95% of the genomes of said particular subspecies, but in range of less than 1-9% of the genomes from different subspecies, wherein the hash values are generated from coding sequences of a plurality of microbiota species, and wherein the hash values correspond to coding sequences of genomes from each of the microbiota species. In some embodiments, determining the presence of the microbiome bacteriasubspecies in the genome data includes determining an amount of the microbiome bacteria subspecies in the genome data.

[0117] In certain embodiments, method 1600 further includes identifying a characteristic of at least one of the microbiome bacteria subspecies in the sample based on the analyzing. The characteristic may be associated with a disease, a disease resistance, or a therapy. In some embodiments, the characteristic is associated with cancer, inflammatory disease, or any disease where changes in microbiota on a subspecies level are detectable. The inflammatory disease may be inflammatory bowel disease (IBD), Chron’s disease, or ulcerative colitis. The cancer may be a gastrointestinal cancer, a melanoma, or an adenoma. In some embodiments, the characteristic is associated with a gastrointestinal tract cancer, including colorectal cancer, and said association of said characteristic is an increased incidence of the cancer.

[0118] In certain embodiments, the analyzing is completed by a trained machine learning algorithm implemented on the computer system, where the trained machine learning algorithm has been trained using the library of microbiome bacteria subspecies. In some embodiments, the trained the machine learning model is trained with additional data such as risk factors, including one or more of age, lifestyle factors, family and / or personal history.

[0119] Various techniques described herein, may be performed by one or more computer programs. The term "program" is to be construed broadly to cover a sequence of instructions in a programming language that a computing device can execute or interpret. These programs may be written in any suitable computer language, including lower-level languages such as assembly and higher-level languages such as Python.

[0120] Program instructions may be stored on a "non-transitory, computer- readable storage medium" or a "non-transitory, computer-readable medium." The storage of program instructions on such media permits execution of the program instructions by a computer system. These are broad terms intended to cover any type of computer memory or storage device that is capable of storing program instructions. The term "non-transitory," as is understood, refers to a tangible medium. Note that the program instructions may be stored on the medium in various formats (source code, compiled code, etc.).

[0121] The phrases "computer-readable storage medium" and "computer- readable medium" are intended to refer to both a storage medium within a computer system as well as a removable medium such as a CD-ROM, memory stick, or portable hard drive. The phrases cover any type of volatile memory within a computer system including DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc., as well as non-volatile memory such as magnetic media, e.g., a hard drive, or optical storage. The phrases are explicitly intended to cover the memory of a server that facilitates downloading of program instructions, the memories within anyintermediate computer system involved in the download, as well as the memories of all destination computing devices. Still further, the phrases are intended to cover combinations of different types of memories.

[0122] In addition, a computer-readable medium or storage medium may be located in a first set of one or more computer systems in which the programs are executed, as well as in a second set of one or more computer systems which connect to the first set over a network. In the latter instance, the second set of computer systems may provide program instructions to the first set of computer systems for execution. In short, the phrases "computer-readable storage medium" and "computer-readable medium" may include two or more media that may reside in different locations, e.g., in different computers that are connected over a network.

[0123] Note that in some cases, program instructions may be stored on a storage medium but not enabled to execute in a particular computing environment. For example, a particular computing environment (e.g., a first computer system) may have a parameter set that disables program instructions that are nonetheless resident on a storage medium of the first computer system. The recitation that these stored program instructions are "capable" of being executed is intended to account for and cover this possibility. Stated another way, program instructions stored on a computer-readable medium can be said to "executable" to perform certain functionality, whether or not current software configuration parameters permit such execution. Executability means that when and if the instructions are executed, they perform the functionality in question.

[0124] Similarly, systems that implement the methods described with respect to any of the disclosed techniques are also contemplated. One such system is described with reference to FIG 17. FIG 17 is a block diagram of another embodiment of such a computer system. Computer system 1700 includes a processor subsystem 1780 that is coupled to a system memory 1720 and I / O interfaces(s) 1740 via an interconnect 1760 (e.g., a system bus). I / O interface(s) 1740 is coupled to one or more I / O devices 1750. Computer system 1700 may be any of various types of devices, including, but not limited to, a server system, personal computer system, desktop computer, laptop or notebook computer, mainframe computer system, tablet computer, handheld computer, workstation, network computer, a consumer device such as a mobile phone, music player, or personal data assistant (PDA). Although a single computer system 1700 is shown in FIG.17 for convenience, system 1700 may also be implemented as two or more computer systems operating together.

[0125] Processor subsystem 1780 may include one or more processors or processing units. In various embodiments of computer system 1700, multiple instances of processor subsystem 1780 may be coupled to interconnect 1760. In various embodiments,processor subsystem 1780 (or each processor unit within 1780) may contain a cache or other form of on-board memory.

[0126] System memory 1720 is usable to store program instructions executable by processor subsystem 1780 to cause system 1700 perform various operations described herein. System memory 1720 may be implemented using different physical memory media, such as hard disk storage, floppy disk storage, removable disk storage, flash memory, random access memory (RAM-SRAM, EDO RAM, SDRAM, DDR SDRAM, RAMBUS RAM, etc.), read only memory (PROM, EEPROM, etc.), and so on. Memory in computer system 1700 is not limited to primary storage such as memory 1720. Rather, computer system 1700 may also include other forms of storage such as cache memory in processor subsystem 1780 and secondary storage on I / O Devices 1750 (e.g., a hard drive, storage array, etc.). In some embodiments, these other forms of storage may also store program instructions executable by processor subsystem 1780.

[0127] I / O interfaces 1740 may be any of various types of interfaces configured to couple to and communicate with other devices, according to various embodiments. In one embodiment, I / O interface 1740 is a bridge chip (e.g., Southbridge) from a front-side to one or more back-side buses. I / O interfaces 1740 may be coupled to one or more I / O devices 1750 via one or more corresponding buses or other interfaces. Examples of I / O devices 1750 include storage devices (hard drive, optical drive, removable flash drive, storage array, SAN, or their associated controller), network interface devices (e.g., to a local or wide-area network), or other devices (e.g., graphics, user interface devices, etc.). In one embodiment, computer system 1700 is coupled to a network via a network interface device 1750 (e.g., configured to communicate over WiFi, Bluetooth, Ethernet, etc.).

[0128] Memory 1720 may include a non-transitory computer-readable storage medium storing program instructions 1722 in various embodiments. Program instructions 1722 may include instructions that are executable to perform methods 1500-1600, for example.

[0129] All of the methods disclosed and claimed herein can be made and executed without undue experimentation in light of the present disclosure. While the compositions and methods of the disclosed embodiments have been described in terms of preferred embodiments, it will be apparent to those of skill in the art that variations may be applied to the methods and in the steps or in the sequence of steps of the method described herein without departing from the concept, spirit and scope of the disclosed embodiments. More specifically, it will be apparent that certain agents which are both chemically and physiologically related may be substituted for the agents described herein while the same or similar results would be achieved. All such similar substitutes and modifications apparent to those skilled in the art are deemed to be within the spirit, scope and concept of the disclosed embodiments as defined by the appended claims.

[0130] The present disclosure includes references to "embodiments," which are non-limiting implementations of the disclosed concepts. References to "an embodiment," "one embodiment," "a particular embodiment," "some embodiments," "various embodiments," and the like do not necessarily refer to the same embodiment. A large number of possible embodiments are contemplated, including specific embodiments described in detail, as well as modifications or alternatives that fall within the spirit or scope of the disclosure. Not all embodiments will necessarily manifest any or all of the potential advantages described herein.

[0131] This disclosure may discuss potential advantages that may arise from the disclosed embodiments. Not all implementations of these embodiments will necessarily manifest any or all of the potential advantages. Whether an advantage is realized for a particular implementation depends on many factors, some of which are outside the scope of this disclosure. In fact, there are a number of reasons why an implementation that falls within the scope of the claims might not exhibit some or all of any disclosed advantages. For example, a particular implementation might include other circuitry outside the scope of the disclosure that, in conjunction with one of the disclosed embodiments, negates or diminishes one or more the disclosed advantages. Furthermore, suboptimal design execution of a particular implementation (e.g., implementation techniques or tools) could also negate or diminish disclosed advantages. Even assuming a skilled implementation, realization of advantages may still depend upon other factors such as the environmental circumstances in which the implementation is deployed. For example, inputs supplied to a particular implementation may prevent one or more problems addressed in this disclosure from arising on a particular occasion, with the result that the benefit of its solution may not be realized. Given the existence of possible factors external to this disclosure, it is expressly intended that any potential advantages described herein are not to be construed as claim limitations that must be met to demonstrate infringement. Rather, identification of such potential advantages is intended to illustrate the type(s) of improvement available to designers having the benefit of this disclosure. That such advantages are described permissively (e.g., stating that a particular advantage "may arise") is not intended to convey doubt about whether such advantages can in fact be realized, but rather to recognize the technical reality that realization of such advantages often depends on additional factors.

[0132] Unless stated otherwise, embodiments are non-limiting. That is, the disclosed embodiments are not intended to limit the scope of claims that are drafted based on this disclosure, even where only a single example is described with respect to a particular feature. The disclosed embodiments are intended to be illustrative rather than restrictive, absent any statements in the disclosure to the contrary. The application is thus intended to permit claims coveringdisclosed embodiments, as well as such alternatives, modifications, and equivalents that would be apparent to a person skilled in the art having the benefit of this disclosure.

[0133] For example, features in this application may be combined in any suitable manner. Accordingly, new claims may be formulated during prosecution of this application (or an application claiming priority thereto) to any such combination of features. In particular, with reference to the appended claims, features from dependent claims may be combined with those of other dependent claims where appropriate, including claims that depend from other independent claims. Similarly, features from respective independent claims may be combined where appropriate.

[0134] Accordingly, while the appended dependent claims may be drafted such that each depends on a single other claim, additional dependencies are also contemplated. Any combinations of features in the dependent that are consistent with this disclosure are contemplated and may be claimed in this or another application. In short, combinations are not limited to those specifically enumerated in the appended claims.

[0135] Where appropriate, it is also contemplated that claims drafted in one format or statutory type (e.g., apparatus) are intended to support corresponding claims of another format or statutory type (e.g., method).

[0136] Because this disclosure is a legal document, various terms and phrases may be subject to administrative and judicial interpretation. Public notice is hereby given that the following paragraphs, as well as definitions provided throughout the disclosure, are to be used in determining how to interpret claims that are drafted based on this disclosure.

[0137] References to a singular form of an item (i.e., a noun or noun phrase preceded by "a," "an," or "the") are, unless context clearly dictates otherwise, intended to mean "one or more." Reference to "an item" in a claim thus does not, without accompanying context, preclude additional instances of the item. A "plurality" of items refers to a set of two or more of the items.

[0138] The word "may" be used herein in a permissive sense (i.e., having the potential to, being able to) and not in a mandatory sense (i.e., must).

[0139] The terms "comprising" and "including," and forms thereof, are open- ended and mean "including, but not limited to."

[0140] When the term "or" is used in this disclosure with respect to a list of options, it will generally be understood to be used in the inclusive sense unless the context provides otherwise. Thus, a recitation of "x or y" is equivalent to "x or y, or both," and thus covers 1) x but not y, 2) y but not x, and 3) both x and y. On the other hand, a phrase such as "either x or y, but not both" makes clear that "or" is being used in the exclusive sense.

[0141] A recitation of "w, x, y, or z, or any combination thereof" or "at least one of … w, x, y, and z" is intended to cover all possibilities involving a single element up to the total number of elements in the set. For example, given the set [w, x, y, z], these phrasings cover any single element of the set (e.g., w but not x, y, or z), any two elements (e.g., w and x, but not y or z), any three elements (e.g., w, x, and y, but not z), and all four elements. The phrase "at least one of … w, x, y, and z" thus refers to at least one element of the set [w, x, y, z], thereby covering all possible combinations in this list of elements. This phrase is not to be interpreted to require that there is at least one instance of w, at least one instance of x, at least one instance of y, and at least one instance of z.

[0142] Various "labels" may precede nouns or noun phrases in this disclosure. Unless context provides otherwise, different labels used for a feature (e.g., "first circuit," "second circuit," "particular circuit," "given circuit," etc.) refer to different instances of the feature. Additionally, the labels "first," "second," and "third" when applied to a feature do not imply any type of ordering (e.g., spatial, temporal, logical, etc.), unless stated otherwise.

[0143] The phrase "based on" or is used to describe one or more factors that affect a determination. This term does not foreclose the possibility that additional factors may affect the determination. That is, a determination may be solely based on specified factors or based on the specified factors as well as other, unspecified factors. Consider the phrase "determine A based on B." This phrase specifies that B is a factor that is used to determine A or that affects the determination of A. This phrase does not foreclose that the determination of A may also be based on some other factor, such as C. This phrase is also intended to cover an embodiment in which A is determined based solely on B. As used herein, the phrase "based on" is synonymous with the phrase "based at least in part on."

[0144] The phrases "in response to" and "responsive to" describe one or more factors that trigger an effect. This phrase does not foreclose the possibility that additional factors may affect or otherwise trigger the effect, either jointly with the specified factors or independent from the specified factors. That is, an effect may be solely in response to those factors, or may be in response to the specified factors as well as other, unspecified factors. Consider the phrase "perform A in response to B." This phrase specifies that B is a factor that triggers the performance of A, or that triggers a particular result for A. This phrase does not foreclose that performing A may also be in response to some other factor, such as C. This phrase also does not foreclose that performing A may be jointly in response to B and C. This phrase is also intended to cover an embodiment in which A is performed solely in response to B. As used herein, the phrase "responsive to" is synonymous with the phrase "responsive at least in part to." Similarly, the phrase "in response to" is synonymous with the phrase "at least in part in response to."

[0145] Within this disclosure, different entities (which may variously be referred to as "units," "circuits," other components, etc.) may be described or claimed as "configured" to perform one or more tasks or operations. This formulation-[entity] configured to [perform one or more tasks]-is used herein to refer to structure (i.e., something physical). More specifically, this formulation is used to indicate that this structure is arranged to perform the one or more tasks during operation. A structure can be said to be "configured to" perform some tasks even if the structure is not currently being operated. Thus, an entity described or recited as being "configured to" perform some tasks refers to something physical, such as a device, circuit, a system having a processor unit and a memory storing program instructions executable to implement the task, etc. This phrase is not used herein to refer to something intangible.

[0146] In some cases, various units / circuits / components may be described herein as performing a set of task or operations. It is understood that those entities are "configured to" perform those tasks / operations, even if not specifically noted.

[0147] The term "configured to" is not intended to mean "configurable to." An unprogrammed FPGA, for example, would not be considered to be "configured to" perform a particular function. This unprogrammed FPGA may be "configurable to" perform that function, however. After appropriate programming, the FPGA may then be said to be "configured to" perform the particular function.

[0148] For purposes of United States patent applications based on this disclosure, reciting in a claim that a structure is "configured to" perform one or more tasks is expressly intended not to invoke 35 U.S.C. § 112(f) for that claim element. Should Applicant wish to invoke Section 112(f) during prosecution of a United States patent application based on this disclosure, it will recite claim elements using the "means for" [performing a function] construct.

[0149] The following numbered clauses set out various non-limiting embodiments disclosed herein: A1. A computer-implemented method for creating an annotated catalog or dataset of a number of microbiome bacteria subspecies comprising: building a dataset of subspecies as a collection of hash values specific to a subspecies, including- sketching genomes of microbiota species producing a set of hash values for a genome; selecting hash values specific to a subspecies that appear in a range of at least 10-95% of the genomes of said subspecies, but in range of less than 1-9% of the genomes from different subspecies; and including selected hash values specific to said subspecies in said dataset.A2. The method of any previous clause within set A, the dataset of subspecies is built by selecting hash values that appear in at least about 20% of the genomes of said subspecies, but in less than about 5% of the genomes from different subspecies. A3. The method of any previous clause within set A, said dataset of subspecies being used to identify a characteristic of one or more microbiome bacteria subspecies where the characteristic is associated with a disease, a disease resistance, or a therapy. A4. The method of any previous clause within set A, said characteristic being associated with cancer, inflammatory disease, or any disease where changes in microbiota on a subspecies level can be detected. A5. The method of any previous clause within set A, wherein the inflammatory disease is inflammatory bowel disease (IBD), Chron’s disease, or ulcerative colitis. A6. The method of any previous clause within set A, wherein the cancer is a gastrointestinal cancer, a melanoma, or an adenoma. A7. The method of any previous clause within set A, said characteristic being associated with a gastrointestinal tract cancer, including colorectal cancer, and said association of said characteristic is an increased incidence of the cancer. A8. The method of any previous clause within set A, wherein the characteristic is associated with resistance to a disease, wherein the disease is cancer or any disease where changes in microbiota on a subspecies level can be detected. A9. The method of any previous clause within set A, wherein the cancer is a gastrointestinal cancer, colorectal cancer, adenoma, or any cancer where changes in microbiota on a subspecies level can be detected. A10. The method of any previous clause within set A, identifying a disease, disease resistance, or response to therapy, comprising training a machine learning model using a one or more studies applied to the subspecies dataset which associate a characteristic with a disease or therapy to build classifiers; operating said trained machine learning model by applying said classifiers to a sample to identify said characteristic. A11. The method of any previous clause within set A, wherein the method includes inputting a microbiota-containing sample from a human into the trained machine learning model and identifying if said characteristic is present in the sample. A12. The method of any previous clause within set A, wherein the machine learning model includes gradient boosting, and / or comprises a Light Gradient Boosting Machine. A13. The method of any previous clause within set A, wherein the machine learning model comprises an Artificial Neural Network.A14. The method of any previous clause within set A, wherein the method comprises training the machine learning model with additional data such as risk factors, including one or more of age, lifestyle factors, family and / or personal history. A15. The method of any previous clause within set A, wherein the method comprises training the machine learning model with data from one or more colorectal cancer studies. A16. The method of any previous clause within set A, wherein the method comprises training the machine learning model with additional data such as fecal occult blood test. A17. The method of any previous clause within set A, wherein the method comprises training the machine learning model with top ranked subspecies associated with minimal microbial signature. A18. The method of any previous clause within set A, including a hash function building said hash values, the hash function comprising a murmur function used for sketching. A19. The method of any previous clause within set A, wherein building the subspecies dataset includes Sourmash employing a hash function to sketch genomes of microbiota species. A20. The method of any previous clause within set A, wherein the genomes of microbiota species are clustered into species-level OTUs having at least 95% average nucleotide identity prior to said building the dataset of subspecies. A21. The method of any previous clause within set A, wherein said building the dataset of subspecies comprises clustering species in an unsupervised manner. A22. The method of any previous clause within set A, wherein the clustering species in an unsupervised manner comprises using a Prediction Strength metric with a cutoff between 0.5 and 0.95. A23. The method of any previous clause within set A, wherein the Prediction Strength metric includes a cutoff of about 0.8. A24. The method of any previous clause within set A, wherein said building the dataset of subspecies comprises removing contaminated or incomplete genomes. A25. The method of any previous clause within set A, wherein GUNC and / or BUSCO are used to remove contaminated or incomplete genomes. A26. The method of any previous clause within set A, wherein a murmur hash function is used to sketch genomes of microbiota species. A27. The method of any previous clause within set A, wherein the subspecies dataset includes an identification of one or more genes describing the functional variation between subspecies. A28. The method of any previous clause within set A, wherein the method comprises building the subspecies dataset by obtaining sketches of genomes annotated with subspecies information.A29. The method of any previous clause within set A, wherein the sketches of genomes comprise or consist essentially of sketches of coding sequences predicted in microbial genomes. Set B B1. A computer-implemented method for creating a catalog or dataset of microbiome bacteria subspecies to identify a characteristic in microbiome bacteria in a human associated with a disease, a disease resistance, or a therapy, comprising: building a dataset of subspecies where the dataset is identified as quantified subspecies abundances using hash values specific to said subspecies by obtaining sketches of genomes annotated with subspecies information; training a machine learning model using said dataset of subspecies and using studies associating a characteristic with a disease, disease resistance, or therapy to identify classifiers; applying the classifiers to a microbiota-containing sample; and identifying in the microbiota-containing sample and using said classifiers, said characteristic associated with the disease, disease resistance, or therapy. B2. The method of any previous clause within set B, wherein the method further comprises identifying a function specific to a subspecies using hash values specific to said subspecies by obtaining sketches of genomes annotated with subspecies information; where said subspecies is associated with a disease and sibling subspecies and / or parental species are not associated with said disease. B3. The method of any previous clause within set B, wherein the dataset of subspecies is built by selecting hash values specific to said subspecies that appear in at least about 10- 95% of the genomes of said subspecies, but in less than about 1-9% of the genomes from different subspecies. B4. The method of any previous clause within set B, wherein the dataset of subspecies is built by selecting hash values that appear in at least about 20% of the genomes of said subspecies, but in less than about 5% of the genomes from different subspecies. B5. The method of any previous clause within set B, said characteristic being associated with cancer. B6. The method of any previous clause within set B, wherein the therapy is an anti-cancer therapy, immunotherapy, chemotherapy, biologicals, hormone therapy, hyper-and hypothermia, photodynamic therapy, radiation therapy, stem cell transplant, checkpoint therapy, or anti-PD1 therapy. B7. The method of any previous clause within set B, wherein the characteristic is predicted response of the disease or the human to the therapy. Set CC1. A system to classify a microbiome bacteria sample from a human, the system comprising: one or more computers; and one or more non-transitory computer readable media couple to the one or more computers and having instructions stored thereon which, when executed by the one or more computers causes the one or more computers to perform operations comprising- building a dataset of subspecies as a collection of hash values specific to a subspecies, by selecting hash values that appear in at least about 10- 95% of the genomes of said subspecies, but in less than about 1-9% of the genomes from different subspecies; obtaining one or more studies which associate a characteristic with a disease or therapy; processing the studies and dataset of subspecies using a machine learning model to build classifiers to identify a characteristic associated with a disease or therapy; inputting microbiome bacteria sample from a human into the machine learning model, applying said classifiers and identifying said characteristic associated with a disease or therapy in the sample. C2. The system of any previous clause within set C, said characteristic being associated with a gastrointestinal tract cancer, including colorectal cancer and said association of said characteristic is an increased incidence of cancer. C3. The system of any previous clause within set C, including a hash function building said hash values comprising a murmur function used for sketching. Set D D1. A method of detecting or identifying risk of a disease in a mammalian subject comprising: (i) obtaining a microbiome sample from the subject; (ii) sequencing genome data from the microbiome sample; (iii) comparing the genome data from the microbiome sample to a catalog or dataset of microbiome bacteria subspecies generated by any one of claims 30-35 or the annotated catalog or dataset of microbiome bacteria subspecies of any one of claims 1-29; wherein if microbiome subspecies associated with the disease are identified in the microbiome sample, then the subject has an increased risk of or has the disease. D2. The method of any previous clause within set D, wherein the disease is a cancer. D3. The method of any previous clause within set D, wherein the cancer is a gastrointestinal cancer, a colorectal cancer, an adenoma, or a melanoma. D4. The method of any previous clause within set D, wherein the disease is an inflammatory disease.D5. The method of any previous clause within set D, wherein the inflammatory disease is Chron’s disease, inflammatory bowel disorder (IBD), or ulcerative colitis. D6. The method of any previous clause within set D, wherein the subject is a human. REFERENCES The following references, to the extent that they provide exemplary procedural or other details supplementary to those set forth herein, are specifically incorporated herein by reference. 1. Blanco-Míguez, A., Beghini, F., Cumbo, F., McIver, L.J., Thompson, K.N., Zolfo, M., Manghi, P., Dubois, L., Huang, K.D., Thomas, A.M., et al. (2023). Extending and improving metagenomic taxonomic profiling with uncharacterized species using MetaPhlAn 4. Nature Biotechnology, 1–12. 2. Ruscheweyh, H.-J., Milanese, A., Paoli, L., Karcher, N., Clayssen, Q., Keller, M.I., Wirbel, J., Bork, P., Mende, D.R., Zeller, G., et al. (2022). Cultivation-independent genomes greatly expand taxonomic-profiling capabilities of mOTUs across various environments. Microbiome 10, 1–12. 3. Van Rossum, T., Ferretti, P., Maistrenko, O.M., and Bork, P. (2020). Diversity within species: interpreting strains in microbiomes. Nat Rev Microbiol 18, 491–506. 10.1038 / s41579- 020-0368-1. 4. Truong, D.T., Tett, A., Pasolli, E., Huttenhower, C., and Segata, N. (2017). Microbial strain-level population structure and genetic diversity from metagenomes. Genome Res. 27, 626– 638. 10.1101 / gr.216242.116. 5. Davenport, R., and Smith, C. (1952). Panophthalmitis due to an Organism of the Bacillus Subtilis Group. British Journal of Ophthalmology 36, 389–392. 10.1136 / bjo.36.7.389. 6. Gordienko, E.N., Kazanov, M.D., and Gelfand, M.S. (2013). Evolution of Pan-Genomes of Escherichia coli, Shigella spp., and Salmonella enterica. J Bacteriol 195, 2786–2792. 10.1128 / JB.02285-12. 7. Schloissnig, S., Arumugam, M., Sunagawa, S., Mitreva, M., Tap, J., Zhu, A., Waller, A., Mende, D.R., Kultima, J.R., Martin, J., et al. (2013). Genomic variation landscape of the human gut microbiome. Nature 493, 45–50. 8. Greenblum, S., Carr, R., and Borenstein, E. (2015). Extensive strain-level copy-number variation across human gut microbiome species. Cell 160, 583–594. 9. Scholz, M., Ward, D.V., Pasolli, E., Tolio, T., Zolfo, M., Asnicar, F., Truong, D.T., Tett, A., Morrow, A.L., and Segata, N. (2016). Strain-level microbial epidemiology and population genomics from shotgun metagenomics. Nature methods 13, 435–438. 10. Pasolli, E., Asnicar, F., Manara, S., Zolfo, M., Karcher, N., Armanini, F., Beghini, F., Manghi, P., Tett, A., Ghensi, P., et al. (2019). Extensive unexplored human microbiomediversity revealed by over 150,000 genomes from metagenomes spanning age, geography, and lifestyle. Cell 176, 649–662. 11. Karcher, N., Nigro, E., Punčochář, M., Blanco-Míguez, A., Ciciani, M., Manghi, P., Zolfo, M., Cumbo, F., Manara, S., Golzato, D., et al. (2021). Genomic diversity and ecology of human-associated Akkermansia species in the gut microbiome revealed by extensive metagenomic assembly. Genome Biol 22, 209.10.1186 / s13059-021-02427-7. 12. Lawley, B., Munro, K., Hughes, A., Hodgkinson, A.J., Prosser, C.G., Lowry, D., Zhou, S.J., Makrides, M., Gibson, R.A., Lay, C., et al. (2017). Differentiation of Bifidobacterium longum subspecies longum and infantis by quantitative PCR using functional gene targets. PeerJ 5, e3375. 13. Costea, P.I., Coelho, L.P., Sunagawa, S., Munch, R., Huerta‐Cepas, J., Forslund, K., Hildebrand, F., Kushugulova, A., Zeller, G., and Bork, P. (2017). Subspecies in the global human gut microbiome. Mol Syst Biol 13, 960.10.15252 / msb.20177589. 14. Van Rossum, T., Costea, P.I., Paoli, L., Alves, R., Thielemann, R., Sunagawa, S., and Bork, P. (2022). metaSNV v2: detection of SNVs and subspecies in prokaryotic metagenomes. Bioinformatics 38, 1162–1164. 15. Tan, C., Cui, W., Cui, X., and Ning, K. (2019). Strain-GeMS: optimized subspecies identification from microbiome data based on accurate variant modeling. Bioinformatics 35, 1789–1791. 16. Hiseni, P., Rudi, K., Wilson, R.C., Hegge, F.T., and Snipen, L. (2021). HumGut: a comprehensive human gut prokaryotic genomes collection filtered by metagenome data. Microbiome 9, 1–12. 17. Almeida, A., Nayfach, S., Boland, M., Strozzi, F., Beracochea, M., Shi, Z.J., Pollard, K.S., Sakharova, E., Parks, D.H., Hugenholtz, P., et al. (2021). A unified catalog of 204,938 reference genomes from the human gut microbiome. Nat Biotechnol 39, 105–114. 10.1038 / s41587-020-0603-3. 18. Bowers, R.M., Kyrpides, N.C., Stepanauskas, R., Harmon-Smith, M., Doud, D., Reddy, T.B.K., Schulz, F., Jarett, J., Rivers, A.R., Eloe-Fadrosh, E.A., et al. (2017). Minimum information about a single amplified genome (MISAG) and a metagenome-assembled genome (MIMAG) of bacteria and archaea. Nat Biotechnol 35, 725–731.10.1038 / nbt.3893. 19. Orakov, A., Fullam, A., Coelho, L.P., Khedkar, S., Szklarczyk, D., Mende, D.R., Schmidt, T.S., and Bork, P. (2021). GUNC: detection of chimerism and contamination in prokaryotic genomes. Genome biology 22, 1–19. 20. Manni, M., Berkeley, M.R., Seppey, M., Simão, F.A., and Zdobnov, E.M. (2021). BUSCO Update: Novel and Streamlined Workflows along with Broader and DeeperPhylogenetic Coverage for Scoring of Eukaryotic, Prokaryotic, and Viral Genomes. Molecular Biology and Evolution 38, 4647–4654.10.1093 / molbev / msab199. 21. Brown, C.T., and Irber, L. (2016). sourmash: a library for MinHash sketching of DNA. The Journal of Open Source Software 1, 27.10.21105 / joss.00027. 22. Tibshirani, R., muand Walther, G. (2005). Cluster Validation by Prediction Strength. Journal of Computational and Graphical Statistics 14, 511–528.10.1198 / 106186005X59243. 23. Seppey, M., Manni, M., and Zdobnov, E.M. (2020). LEMMI: a continuous benchmarking platform for metagenomics classifiers. Genome Res 30, 1208–1216.10.1101 / gr.260398.119. 24. Rothschild, D., Weissbrod, O., Barkan, E., Kurilshikov, A., Korem, T., Zeevi, D., Costea, P.I., Godneva, A., Kalka, I.N., Bar, N., et al. (2018). Environment dominates over host genetics in shaping human gut microbiota. Nature 555, 210–215.10.1038 / nature25973. 25. Falony, G., Joossens, M., Vieira-Silva, S., Wang, J., Darzi, Y., Faust, K., Kurilshikov, A., Bonder, M.J., Valles-Colomer, M., Vandeputte, D., et al. (2016). Population-level analysis of gut microbiome variation. Science 352, 560–564.10.1126 / science.aad3503. 26. Zhernakova, A., Kurilshikov, A., Bonder, M.J., Tigchelaar, E.F., Schirmer, M., Vatanen, T., Mujagic, Z., Vila, A.V., Falony, G., Vieira-Silva, S., et al. (2016). Population-based metagenomics analysis reveals markers for gut microbiome composition and diversity. Science 352, 565–569.10.1126 / science.aad3369. 27. Irber, L., and Brown, C.T. (2022). Mastiff. Version 0.0.3. 28. Lokmer, A., Cian, A., Froment, A., Gantois, N., Viscogliosi, E., Chabé, M., and Ségurel, L. (2019). Use of shotgun metagenomics for the identification of protozoa in the gut microbiota of healthy individuals from worldwide populations with various industrialization levels. PloS one 14, e0211139. 29. Smits, S.A., Leach, J., Sonnenburg, E.D., Gonzalez, C.G., Lichtman, J.S., Reid, G., Knight, R., Manjurano, A., Changalucha, J., Elias, J.E., et al. (2017). Seasonal cycling in the gut microbiome of the Hadza hunter-gatherers of Tanzania. Science 357, 802–806. 30. Rampelli, S., Schnorr, S.L., Consolandi, C., Turroni, S., Severgnini, M., Peano, C., Brigidi, P., Crittenden, A.N., Henry, A.G., and Candela, M. (2015). Metagenome sequencing of the Hadza hunter-gatherer gut microbiota. Current Biology 25, 1682–1693. 31. Tett, A., Huang, K.D., Asnicar, F., Fehlner-Peach, H., Pasolli, E., Karcher, N., Armanini, F., Manghi, P., Bonham, K., Zolfo, M., et al. (2019). The Prevotella copri Complex Comprises Four Distinct Clades Underrepresented in Westernized Populations. Cell Host & Microbe 26, 666-679.e7.10.1016 / j.chom.2019.08.018.32. Carter, M.M., Olm, M.R., Merrill, B., Dahan, D., Tripathi, S., Spencer, S., Yu, F., Jain, S., Neff, N., Jha, A., et al. (2022). Ultra-deep sequencing of Hadza hunter-gatherers recovers vanishing gut microbes. Science. 33. Sung, H., Ferlay, J., Siegel, R.L., Laversanne, M., Soerjomataram, I., Jemal, A., and Bray, F. (2021). Global Cancer Statistics 2020: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries. CA Cancer J Clin 71, 209–249. 10.3322 / caac.21660. 34. Feng, Q., Liang, S., Jia, H., Stadlmayr, A., Tang, L., Lan, Z., Zhang, D., Xia, H., Xu, X., Jie, Z., et al. (2015). Gut microbiome development along the colorectal adenoma–carcinoma sequence. Nat Commun 6, 6528.10.1038 / ncomms7528. 35. Wirbel, J., Pyl, P.T., Kartal, E., Zych, K., Kashani, A., Milanese, A., Fleck, J.S., Voigt, A.Y., Palleja, A., Ponnudurai, R., et al. (2019). Meta-analysis of fecal metagenomes reveals global microbial signatures that are specific for colorectal cancer. Nat Med 25, 679–689. 10.1038 / s41591-019-0406-6. 36. Zeller, G., Tap, J., Voigt, A.Y., Sunagawa, S., Kultima, J.R., Costea, P.I., Amiot, A., Böhm, J., Brunetti, F., Habermann, N., et al. (2014). Potential of fecal microbiota for early-stage detection of colorectal cancer. Molecular Systems Biology 10, 766.10.15252 / msb.20145645. 37. Yu, J., Feng, Q., Wong, S.H., Zhang, D., Liang, Q.Y., Qin, Y., Tang, L., Zhao, H., Stenvang, J., Li, Y., et al. (2017). Metagenomic analysis of faecal microbiome as a tool towards targeted non-invasive biomarkers for colorectal cancer. Gut 66, 70–78.10.1136 / gutjnl-2015- 309800. 38. Yachida, S., Mizutani, S., Shiroma, H., Shiba, S., Nakajima, T., Sakamoto, T., Watanabe, H., Masuda, K., Nishimoto, Y., Kubo, M., et al. (2019). Metagenomic and metabolomic analyses reveal distinct stage-specific phenotypes of the gut microbiota in colorectal cancer. Nat Med 25, 968–976.10.1038 / s41591-019-0458-7. 39. Thomas, A.M., Manghi, P., Asnicar, F., Pasolli, E., Armanini, F., Zolfo, M., Beghini, F., Manara, S., Karcher, N., Pozzi, C., et al. (2019). Metagenomic analysis of colorectal cancer datasets identifies cross-cohort microbial diagnostic signatures and a link with choline degradation. Nat Med 25, 667–678.10.1038 / s41591-019-0405-7. 40. Vogtmann, E., Hua, X., Zeller, G., Sunagawa, S., Voigt, A.Y., Hercog, R., Goedert, J.J., Shi, J., Bork, P., and Sinha, R. (2016). Colorectal cancer and the human gut microbiome: reproducibility with whole-genome shotgun sequencing. PloS one 11, e0155362. 41. Gupta, A., Dhakan, D.B., Maji, A., Saxena, R., PK, V.P., Mahajan, S., Pulikkan, J., Kurian, J., Gomez, A.M., Scaria, J., et al. (2019). Association of Flavonifractor plautii, aflavonoid-degrading bacterium, with the gut microbiome of colorectal cancer patients in India. MSystems 4, 10–1128. 42. Yang, Y., Du, L., Shi, D., Kong, C., Liu, J., Liu, G., Li, X., and Ma, Y. (2021). Dysbiosis of human gut microbiome in young-onset colorectal cancer. Nature communications 12, 6757. 43. Okumura, S., Konishi, Y., Narukawa, M., Sugiura, Y., Yoshimoto, S., Arai, Y., Sato, S., Yoshida, Y., Tsuji, S., Uemura, K., et al. (2021). Gut bacteria identified in colorectal cancer patients promote tumourigenesis via butyrate secretion. Nat Commun 12, 5674.10.1038 / s41467- 021-25965-x. 44. Dikeocha, I.J., Al-Kabsi, A.M., Chiu, H.-T., and Alshawsh, M.A. (2022). Faecalibacterium prausnitzii ameliorates colorectal tumorigenesis and suppresses proliferation of HCT116 colorectal cancer cells. Biomedicines 10, 1128. 45. Ferreira-Halder, C.V., de Sousa Faria, A.V., and Andrade, S.S. (2017). Action and function of Faecalibacterium prausnitzii in health and disease. Best Practice & Research Clinical Gastroenterology 31, 643–648. 46. Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., and Liu, T.-Y. (2017). LightGBM: A Highly Efficient Gradient Boosting Decision Tree. In Advances in Neural Information Processing Systems (Curran Associates, Inc.). 47. Riester, M., Wei, W., Waldron, L., Culhane, A.C., Trippa, L., Oliva, E., Kim, S., Michor, F., Huttenhower, C., Parmigiani, G., et al. (2014). Risk Prediction for Late-Stage Ovarian Cancer by Meta-analysis of 1525 Patient Samples. JNCI: Journal of the National Cancer Institute 106, dju048.10.1093 / jnci / dju048. 48. Haghighat, G., Khajeh-Mehrizi, A., and Ranjbar, H. (2023). Evaluation of Serum Vitamin B12 Levels in Patients with Colon and Breast Cancer: A Case-Control Study. Int J Hematol Oncol Stem Cell Res 17, 240–244.10.18502 / ijhoscr.v17i4.13914. 49. Boughanem, H., Hernandez-Alonso, P., Tinahones, A., Babio, N., Salas-Salvadó, J., Tinahones, F.J., and Macias-Gonzalez, M. (2020). Association between Serum Vitamin B12 and Global DNA Methylation in Colorectal Cancer Patients. Nutrients 12, 3567. 10.3390 / nu12113567. 50. Goel, A., Xicola, R.M., Nguyen, T., Doyle, B.J., Sohn, V.R., Bandipalliam, P., Rozek, L.S., Reyes, J., Cordero, C., Balaguer, F., et al. (2010). Aberrant DNA Methylation in Hereditary Nonpolyposis Colorectal Cancer Without Mismatch Repair Deficiency. Gastroenterology 138, 1854-1862.e1.10.1053 / j.gastro.2010.01.035. 51. Oliai Araghi, S., Kiefte-de Jong, J.C., van Dijk, S.C., Swart, K.M.A., van Laarhoven, H.W., van Schoor, N.M., de Groot, L.C.P.G.M., Lemmens, V., Stricker, B.H., Uitterlinden, A.G., et al. (2019). Folic Acid and Vitamin B12 Supplementation and the Risk of Cancer: Long-term Follow-up of the B Vitamins for the Prevention of Osteoporotic Fractures (B-PROOF) Trial. Cancer Epidemiol Biomarkers Prev 28, 275–282.10.1158 / 1055-9965.EPI-17-1198. 52. Bednarz-Misa, I., Fleszar, M.G., Zawadzki, M., Kapturkiewicz, B., Kubiak, A., Neubauer, K., Witkiewicz, W., and Krzystek-Korpacka, M. (2020). L-Arginine / NO Pathway Metabolites in Colorectal Cancer: Relevance as Disease Biomarkers and Predictors of Adverse Clinical Outcomes Following Surgery. Journal of Clinical Medicine 9, 1782. 10.3390 / jcm9061782. 53. Alexandrou, C., Al-Aqbi, S.S., Higgins, J.A., Boyle, W., Karmokar, A., Andreadi, C., Luo, J.-L., Moore, D.A., Viskaduraki, M., Blades, M., et al. (2018). Sensitivity of Colorectal Cancer to Arginine Deprivation Therapy is Shaped by Differential Expression of Urea Cycle Enzymes. Sci Rep 8, 12096.10.1038 / s41598-018-30591-7. 54. Phillips, M.M., Sheaff, M.T., and Szlosarek, P.W. (2013). Targeting Arginine-Dependent Cancers with Arginine-Degrading Enzymes: Opportunities and Challenges. Cancer Res Treat 45, 251–262.10.4143 / crt.2013.45.4.251. 55. Ijssennagger, N., Belzer, C., Hooiveld, G.J., Dekker, J., van Mil, S.W.C., Müller, M., Kleerebezem, M., and van der Meer, R. (2015). Gut microbiota facilitates dietary heme-induced epithelial hyperproliferation by opening the mucus barrier in colon. Proc Natl Acad Sci U S A 112, 10038–10043.10.1073 / pnas.1507645112. 56. Seiwert, N., Adam, J., Steinberg, P., Wirtz, S., Schwerdtle, T., Adams-Quack, P., Hövelmeyer, N., Kaina, B., Foersch, S., and Fahrer, J. (2021). Chronic intestinal inflammation drives colorectal tumor formation triggered by dietary heme iron in vivo. Arch Toxicol 95, 2507–2522.10.1007 / s00204-021-03064-6. 57. Hyatt, D., Chen, G.-L., LoCascio, P.F., Land, M.L., Larimer, F.W., and Hauser, L.J. (2010). Prodigal: prokaryotic gene recognition and translation initiation site identification. BMC Bioinformatics 11, 119.10.1186 / 1471-2105-11-119. 58. Irber, L., Brooks, P.T., Reiter, T., Pierce-Ward, N.T., Hera, M.R., Koslicki, D., and Brown, C.T. (2022). Lightweight compositional analysis of metagenomes with FracMinHash and minimum metagenome covers. bioRxiv.10.1101 / 2022.01.11.475838. 59. Fritz, A., Hofmann, P., Majda, S., Dahms, E., Dröge, J., Fiedler, J., Lesker, T.R., Belmann, P., DeMaere, M.Z., Darling, A.E., et al. (2019). CAMISIM: simulating metagenomes and microbial communities. Microbiome 7, 1–12. 60. Huang, W., Li, L., Myers, J.R., and Marth, G.T. (2012). ART: a next-generation sequencing read simulator. Bioinformatics 28, 593–594.61. Pasolli, E., Schiffer, L., Manghi, P., Renson, A., Obenchain, V., Truong, D.T., Beghini, F., Malik, F., Ramos, M., Dowd, J.B., et al. (2017). Accessible, curated metagenomic data through ExperimentHub. Nat Methods 14, 1023–1024.10.1038 / nmeth.4468. 62. Kieser, S., Brown, J., Zdobnov, E.M., Trajkovski, M., and McCue, L.A. (2020). ATLAS: a Snakemake workflow for assembly, annotation, and genomic binning of metagenome sequence data. BMC bioinformatics 21, 1–8. 63. Steinegger, M., and Söding, J. (2017). MMseqs2 enables sensitive protein sequence searching for the analysis of massive data sets. Nature biotechnology 35, 1026–1028. 64. Shaffer, M., Borton, M.A., McGivern, B.B., Zayed, A.A., La Rosa, S.L., Solden, L.M., Liu, P., Narrowe, A.B., Rodríguez-Ramos, J., Bolduc, B., et al. (2020). DRAM for distilling microbial metabolism to automate the curation of microbiome function. Nucleic Acids Research 48, 8883–8900.10.1093 / nar / gkaa621. 65. Edgar, R.C. (2004). MUSCLE: multiple sequence alignment with high accuracy and high throughput. Nucleic acids research 32, 1792–1797. 66. Umerenkov, D., Nikolaev, F., Shashkova, T.I., Strashnov, P.V., Sindeeva, M., Shevtsov, A., Ivanisenko, N.V., and Kardymon, O.L. (2023). PROSTATA: a framework for protein stability assessment using transformers. Bioinformatics 39, btad671. 10.1093 / bioinformatics / btad671. 67. Virtanen, P., Gommers, R., Oliphant, T.E., Haberland, M., Reddy, T., Cournapeau, D., Burovski, E., Peterson, P., Weckesser, W., Bright, J., et al. (2020). SciPy 1.0: fundamental algorithms for scientific computing in Python. Nat Methods 17, 261–272.10.1038 / s41592-019- 0686-2. 68. Viechtbauer, W. (2010). Conducting Meta-Analyses in R with the metafor Package. Journal of Statistical Software 36, 1–48.10.18637 / jss.v036.i03. 69. Dildar M, et al. (2021) “A Review Using Deep Learning Techniques.” Int J Environ Res Public Health. May 20;18(10):5479. 70. LeCun, Yann; Bengio, Yoshua; Hinton, Geoffrey (2015). "Deep Learning" (PDF). Nature.521 (7553): 436–444. PMID 26017442. 71. Gopalakrishnan, V., Spencer, C.N., Nezi, L., Reuben, A., Andrews, M.C., Karpinets, T.V., Prieto, P.A., Vicente, D., Hoffman, K., Wei, S.C., et al. (2018). Gut microbiome modulates response to anti–PD-1 immunotherapy in melanoma patients. Science 359, 97–103. 72. McCulloch, J.A., Davar, D., Rodrigues, R.R., Badger, J.H., Fang, J.R., Cole, A.M., Balaji, A.K., Vetizou, M., Prescott, S.M., Fernandes, M.R., et al. (2022). Intestinal microbiota signatures of clinical response and immune-related adverse events in melanoma patients treated with anti-PD-1. Nat Med 28, 545–556.73. Routy, B., Le Chatelier, E., Derosa, L., Duong, C.P.M., Alou, M.T., Daillère, R., Fluckiger, A., Messaoudene, M., Rauber, C., Roberti, M.P., et al. (2018). Gut microbiome influences efficacy of PD-1–based immunotherapy against epithelial tumors. Science 359, 91– 97. 74. Lee, K.A., Thomas, A.M., Bolte, L.A., Björk, J.R., de Ruijter, L.K., Armanini, F., Asnicar, F., Blanco-Miguez, A., Board, R., Calbet-Llopart, N., et al. (2022). Cross-cohort gut microbiome associations with immune checkpoint inhibitor response in advanced melanoma. Nat Med 28, 535–544. 75. Frankel, A.E., Coughlin, L.A., Kim, J., Froehlich, T.W., Xie, Y., Frenkel, E.P., and Koh, A.Y. (2017). Metagenomic Shotgun Sequencing and Unbiased Metabolomic Profiling Identify Specific Human Gut Microbiota and Metabolites Associated with Immune Checkpoint Therapy Efficacy in Melanoma Patients. Neoplasia 19, 848–855. 76. Pasolli, E., Schiffer, L., Manghi, P., Renson, A., Obenchain, V., Truong, D.T., Beghini, F., Malik, F., Ramos, M., Dowd, J.B., et al. (2017). Accessible, curated metagenomic data through ExperimentHub. Nat Methods 14, 1023–1024. 77. Blanco-Míguez, A., Beghini, F., Cumbo, F., McIver, L.J., Thompson, K.N., Zolfo, M., Manghi, P., Dubois, L., Huang, K.D., Thomas, A.M., et al. (2023). Extending and improving metagenomic taxonomic profiling with uncharacterized species using MetaPhlAn 4. Nature Biotechnology, 1–12. 78. Viechtbauer, W. (2010). Conducting Meta-Analyses in R with the metafor Package. Journal of Statistical Software 36, 1–48. 79. Nielsen, H.B., Almeida, M., Juncker, A.S., Rasmussen, S., Li, J., Sunagawa, S., Plichta, D.R., Gautier, L., Pedersen, A.G., Le Chatelier, E., et al. (2014). Identification and assembly of genomes and genetic elements in complex metagenomic samples without using reference genomes. Nat Biotechnol 32, 822–828. 80. Luiz Irber, Phillip T. Brooks, Taylor Reiter, N. Tessa Pierce-Ward, Mahmudur Rahman Hera, David Koslicki, C. Titus Brown; “Lightweight compositional analysis of metagenomes with FracMinHash and minimum metagenome covers,” bioRxiv .01.11.475838, 2022.

Claims

WHAT IS CLAIMED IS:

1. A method, comprising: accessing a catalog comprising coding sequences from a plurality of microbiota species; generating sets of hash values corresponding to coding sequences of genomes from each of the microbiota species; clustering the hash values to identify one or more subspecies in the plurality of microbiota species; assigning groups of the hash values to a particular subspecies by selecting the hash values that appear in a range of at least 10-95% of the genomes of said particular subspecies, but in range of less than 1-9% of the genomes from different subspecies; and generating a dataset describing the subspecies and corresponding assigned groups of hash values to the subspecies.

2. The method according to claim 1, wherein the generated dataset is implemented in a further method to: analyze genome data acquired by in vitro sequencing analysis of a microbiome bacteria sample; and conduct a bacteria related diagnosis based on the analyzed genome data.

3. The method according to any one of claims 1 or 2, wherein the coding sequences comprise one or more k-mers.

4. The method according to of claim 3, wherein the k-mers comprise between 8 and 20 nucleotides or amino acids, between 8 and 15 nucleotides or amino acids, between 8 and 12 nucleotides or amino acids, or preferably between 8 and 10 nucleotides or amino acids.

5. The method according to any one of claims 1-4, wherein assigning groups of the hash values to the particular subspecies includes selecting the hash values that appear in at least about 20% of the genomes of said particular subspecies, but in range of less than about 5% of the genomes from different subspecies.

6. The method according to any one of claims 1-5, wherein the subspecies includes multiple strains of microbiota.

7. The method according to any one of claims 1-6, wherein the generating the sets of hash values includes employing a hash function to sketch the genomes of the microbiota species.

8. The method according to any one of claims 1-7, wherein clustering the hash values includes clustering in an unsupervised manner.

9. The method according to claim 8, wherein the clustering in the unsupervised manner comprises using a Prediction Strength metric with a cutoff between 0.5 and 0.95; the cutoff may preferably be between 0.7 and 0.9, or may be 0.

8.

10. The method according to any one of claims 1-9, further comprising filtering the catalog to remove contaminated or incomplete genomes before generating the sets of hash values.

11. The method according to any one of claims 1-10, wherein the dataset describing the subspecies includes an identification of one or more genes describing functional variation between subspecies.

12. A method, comprising: receiving, at a computer system, genome data from in vitro sequencing analysis on a microbiome bacteria sample; and analyzing, by the computer system, the genome data according to a library of microbiome bacteria subspecies to determine a presence of microbiome bacteria subspecies in the genome data, wherein the library of microbiome bacteria subspecies includes a listing of microbiome bacteria subspecies, a particular subspecies being a grouping of hash values defined as including the hash values that appear in a range of at least 10-95% of the genomes of said particular subspecies, but in range of less than 1- 9% of the genomes from different subspecies, wherein the hash values are generated from coding sequences of a plurality of microbiota species, and wherein the hash values correspond to coding sequences of genomes from each of the microbiota species.

13. The method according to claim 12, further comprising identifying a characteristic of at least one of the microbiome bacteria subspecies in the sample based on the analyzing.

14. The method according to claim 13, wherein the characteristic is associated with a disease, a disease resistance, or a therapy.

15. The method according to any one of claim 13 or 14, wherein the characteristic is associated with cancer, inflammatory disease, or any disease where changes in microbiota on a subspecies level are detectable.

16. The method according to claim 15, wherein the inflammatory disease is inflammatory bowel disease (IBD), Chron’s disease, or ulcerative colitis.

17. The method according to any one of claim 15 or 16, wherein the cancer is a gastrointestinal cancer, a melanoma, or an adenoma.

18. The method according to any one of claims 13-17, wherein the characteristic is associated with a gastrointestinal tract cancer, including colorectal cancer, and said association of said characteristic is an increased incidence of the cancer.

19. The method according to any one of claims 12-18, wherein the analyzing is completed by a trained machine learning algorithm implemented on the computer system, and wherein the trained machine learning algorithm has been trained using the library of microbiome bacteria subspecies.

20. The method according to claim 19, wherein the trained the machine learning model is trained with additional data such as risk factors, including one or more of age, lifestyle factors, family and / or personal history.

21. The method according to any one of claims 12-20, wherein determining the presence of the microbiome bacteria subspecies in the genome data includes determining an amount of the microbiome bacteria subspecies in the genome data.

22. The method according to any one of claims 12-21, further comprising conducting a bacteria related diagnosis based on determining the presence of microbiome bacteria subspecies in the analyzed genome data.

23. Predicting a response to a treatment regime by implementing the method according to any one of claims 1-22.

24. Selecting a treatment regime by implementing the method according to any one of claims 1-22, the treatment regime optionally including treatment using a particular drug.

25. Predicting a likelihood of a subject contracting a disease through an analysis implementing the method according to any one of claims 1-22.

26. A diagnostic test to acquire genome data from a microbiome bacteria sample and determine a presence of microbiome bacteria subspecies in the genome data that implements the method according to any one of claims 1-22.

27. A method, comprising: accessing a catalog comprising coding sequences from a plurality of microbiota species; generating sets of hash values corresponding to coding sequences of genomes from each of the microbiota species; clustering the hash values to identify one or more subspecies in the plurality of microbiota species; assigning groups of the hash values to a particular subspecies by selecting the hash values that appear in a range of at least 10-95% of the genomes of said particular subspecies, but in range of less than 1-9% of the genomes from different subspecies; generating a dataset describing the subspecies and corresponding assigned groups of hash values to the subspecies; and determining a presence of microbiome bacteria subspecies in a microbiome bacteria sample by analyzing the microbiome bacteria sample according to the dataset describing the subspecies.