Metagenomic and multi-OMIC biomarker discovery and diagnostics

Metagenomic and multi-omic analysis with machine learning models for detecting CRC, CRA, and CRAA using microbial biomarkers in stool samples addresses the limitations of current methods, enhancing detection and compliance by offering a non-invasive alternative to colonoscopy.

WO2026076249A1PCT designated stage Publication Date: 2026-04-09PRESCIENT METABIOMICS JV LLC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-10-02
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Current non-invasive methods for detecting colorectal neoplasia, such as colorectal cancer (CRC), colorectal adenoma (CRA), and colorectal advanced adenoma (CRAA), are limited in sensitivity and specificity, particularly for early and advanced adenomas, creating a gap in healthcare that could be addressed by developing sensitive and accurate biomarkers derived from microbial species in stool and other samples.

Method used

The method involves metagenomic and multi-omic analysis of biological samples to identify taxonomic and functional biomarkers, using machine learning models to discriminate between healthy and diseased states, enabling the detection of CRC, CRA, and CRAA through fecal or other sample analysis, thereby providing a non-invasive screening alternative to colonoscopy.

Benefits of technology

This approach enhances the detection of colorectal neoplasia, improving compliance and reducing the need for invasive procedures by providing a more accurate and efficient screening method that can identify early and advanced adenomas, thus potentially reducing the frequency of colonoscopies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025049242_09042026_PF_FP_ABST
    Figure US2025049242_09042026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure in various aspects and embodiments provides systems and methods for evaluating or screening subjects for the presence or absence of colorectal neoplasia, such as colorectal cancer (CRC), colorectal adenoma (CRA), and / or colorectal advanced adenoma (CRAA), by metagenomic or multi-omic analysis of microbiome in biological samples, including fecal samples. In aspects, the present disclosure provides methods for generating machine learning models or signatures based on metagenomic or multi-omic analysis of the microbiome in biological samples, including fecal samples, to evaluate or screen subjects for the presence or absence of disorders, such as but not limited to CRC, CRA, and CRAA.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] MBI-004PC / 108458-5004

[0002] METAGENOMIC AND MULTI-OMIC BIOMARKER DISCOVERY AND DIAGNOSTICS

[0003] PRIORITY

[0004] This Applications claims the benefit of, and claims priority to, U.S. Provisional Application No. 63 / 702,319 filed October 2, 2024, which is hereby incorporated by reference in its entirety.

[0005] BACKGROUND

[0006] Colorectal cancers are among the most prevalent cancers worldwide with an estimated 1.8 million new colon cancer cases and over 700,000 rectal cancer cases reported in 2018. Bray et al., Global cancer statistics 2018: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries, CA Cancer J Clin. 2018; 68(6):394- 424. Further, colon cancer is now the leading cause of cancer-related deaths in the U.S. in men over 20 and under 50. Siegel RL, et al. Cancer Statistics, 2024, CA: A Cancer Journal for Clinicians (2024). Despite the strong evidence demonstrating that screening of individuals with average CRC risk reduces mortality, compliance amongst individuals is limited due to the invasiveness, discomfort and fear associated with colonoscopy. Lauby- Secretan et al., The IARC Perspective on Colorectal Cancer Screening. N Engl J Med. 2018; 378(18):1734-1740. This has created a significant gap in the health care system and CRC prevention in particular, emphasizing the need for sensitive, accurate, non-invasive diagnostics to detect colon carcinomas and adenomas. The present disclosure fills this gap by providing biomarkers derived from microbial species present in stool and other samples.

[0007] BRIEF DESCRIPTION OF DRAWINGS

[0008] FIG. 1 is a schematic diagram illustrating the workflow for identifying and evaluating taxonomic and functional biomarkers (including multi-omic markers) for generating machine learning models to discriminate between conditions, such as (but not limited to) colorectal neoplasia and healthy controls.

[0009] DBl / 162792885.1 1 MBI-004PC / 108458-5004

[0010] DETAILED DESCRIPTION

[0011] The present disclosure in various aspects and embodiments provides methods and systems for evaluating or screening subjects for the presence or absence of colorectal neoplasia, such as colorectal cancer (CRC), colorectal adenoma (CRA), and / or colorectal advanced adenoma (CRAA), by metagenomic and / or multi-omic analysis of biological samples such as fecal, blood, serum, plasma, urine, saliva, biopsy tissues, mucosa tissue sample or swab, intestinal lavage or aspirant, and other biofluids and cell samples (referred to herein as “biological samples”) containing microbiome DNA, RNA, Proteins, and other molecules for molecular analysis. As used herein, the term “multi-omic” can include analysis of human genomic markers (in addition to microbial markers), transcriptomic markers (microbial and / or human), metabolomic / metabolic markers, proteomic markers, and immune system biomarkers. In other aspects, the present disclosure provides methods and systems for generating machine learning models or “signatures” (biomarker profiles or patterns) based on metagenomic and / or multi-omic analysis of biological samples, including but not limited to fecal samples, to evaluate (e.g., screen) subjects for the presence or absence of disorders, such as but not limited to colon disorders such as colorectal neoplasia (e.g., CRC, CRA, and CRAA).

[0012] CRC is one of the most common and deadly cancers worldwide. The majority of CRCs develop through a multistep process known as the adenoma-carcinoma sequence, in which normal colonic epithelium transforms into adenomatous polyps and eventually invasive adenocarcinoma. This process typically occurs over 10-15 years and involves the accumulation of genetic and epigenetic alterations that drive the initiation and progression of neoplasia. By far, interest has focused on CRC, where a number of studies have tried to delineate the relationships between the gut microbiome and disease progression. Indeed, diet and other environmental factors may play an important role in the initiation and progression of colorectal carcinomas. The gut microbiota is believed to play an essential role as an intermediary between our diet and the types of nutrients and toxic byproducts that may be produced.

[0013] CRC is a heterogeneous disease, the majority of which are considered sporadic without underlying heritable features. Frank et al., Concordant and discordant familial

[0014] DBl / 162792885.1 2 MBI-004PC / 108458-5004 cancer: Familial risks, proportions and population impact, Int J Cancer 2017; 140(7) : 1510- 1516. A wide variety of environmental factors including a western diet, obesity, cigarette smoking, alcohol consumption and lack of exercise are known CRC risk factors. Chief amongst these risk factors is diet where an estimated -38% of incipient CRC cases were linked. Additional evidence for environmental influence of CRC is based on findings that the incidence of CRC is influenced by emigration, wherein a subject’s risk of CRC development is altered based on the diet and lifestyle of the recipient country. Each of the above-mentioned CRC risk modifiers is also known to modulate the composition of the gut microbiota. Greathouse et al., Gut microbiome meta-analysis reveals dysbiosis is independent of body mass index in predicting risk of obesity -associated CRC, BMJ Open Gastroenterol 2019; 6(l):e000247; Lee c / a / .. Association between Cigarette Smoking Status and Composition of Gut Microbiota: Population-Based Cross-Sectional Study, J Clin Med 2018; 7(9):282; Rodriguez-Gonzalez et al., Microbiota and Alcohol Use Disorder: Are Psychobiotics a Novel Therapeutic Strategy?. Curr Pharm Des 2020; 26(20):2426-2437; Allen et al., Exercise Alters Gut Microbiota Composition and Function in Lean and Obese Humans, Med Sci Sports Exerc 2018; 50(4):747-757. This association has drawn substantial attention to the gut microbiota as a potential mediator of CRC initiation and / or progression. The large number of species and genes encoded in the gut microbiome represents a source of potential biomarkers for diagnostics and prognostics of early, premalignant adenomas, advanced adenomas, and CRC.

[0015] Microbial dysbiosis is a term used to describe microbiota that differentially represent microbes not typically observed in healthy individuals. In disease states, tens of microbial species may be differentially represented, however, in any single individual afflicted by disease, some but generally not all differentially represented species may be apparent. Such dysbiosis can include microbes present in CRC stool samples that are typically absent in healthy individuals, whereas more subtle signatures involve shifts in microbes present in healthy subjects that are on average over- or under-represented in CRC. The identification and characterization of CRC-specific dysbiosis (as well as CRA- and CRAA-specific dysbiosis) is important for the development of an effective screening tool.

[0016] DBl / 162792885.1 3 MBI-004PC / 108458-5004

[0017] In various aspects and embodiments, the present disclosure enables detection of colorectal cancer (CRC), as well as colorectal adenoma (CRA), and colorectal advanced adenoma (CRAA) (as well as other disorders that can be discriminated based on microbiome composition in samples) based on the composition of a subject’s microbiome (e.g., as present in fecal or other biological samples). In various embodiments, the present disclosure provides taxonomic and gene features (and methods and systems for generating the same) to distinguish healthy subjects from those with CRC or early or advanced adenomas. In some aspects, the present disclosure provides machine learning (ML) models (and methods and systems for making the same) that avoid a variety of pitfalls associated with metagenomic and multi-omic data (e.g., data heterogeneity, noise, overfitting, etc.), including metagenomic and multi-omic data collected using heterogeneous methods and analysis procedures.

[0018] The methods disclosed herein leverage informative biomarkers for each disease class. For example, in some embodiments the methods identify and / or provide features of high importance for distinguishing CRC, CRA, or CRAA from healthy controls.

[0019] In one aspect, the present disclosure provides a method for evaluating a biological subject for the presence of a colorectal neoplasm. In this aspect, the disclosure provides a method for screening subjects as an alternative to invasive procedures such as colonoscopy, to thereby increase screening compliance, and enable early detection (and removal and / or treatment) of neoplasms. This aspect further provides a method for treating a population of subjects for any existing colorectal neoplasia. In various embodiments the method comprises quantifying genetic elements from a biological sample from the subject(s), such as (but not limited to) a fecal sample. Other biological samples (e.g., mucosal tissue samples, blood, exosome, circulating cell-free DNA, urine, saliva) that allow for sampling of the microbiome, including the gut microbiome, and / or human gene mutations and / or gene expression find use according to this disclosure.

[0020] In embodiments, the genetic elements are associated with (e.g., are indicative of) colorectal cancer (CRC), colorectal adenoma (CRA), or colorectal advanced adenoma (CRAA), and such features can be identified and selected using metagenomic sequencing

[0021] DBl / 162792885.1 4 MBI-004PC / 108458-5004 and machine learning models as described herein. The genetic elements comprise elements associated with microbial taxonomic classification and elements associated with one or more microbial gene classifications or functions. In this manner, in embodiments, the process prepares an abundance profile of the genetic elements, and the abundance profile is evaluated for a signature indicating the presence or absence of CRC, CRA, and / or CRAA in the subject. The subject can therefore be identified as likely to have (or not have) CRC, CRA, and / or CRAA. The process can provide a binary classification (i.e., presence of absence) or a statistical output indicating the likelihood that the subject has CRC, CRA, or CRAA. In embodiments, the abundance profile of genetic elements is evaluated for signatures classifying the profile as either (1) CRC and / or CRAA, or (2) CRA or control. In various embodiments, the method provides for improved detection of adenomas (CRA and / or CRAA) over known detection tests. Subjects identified as having CRC, CRAA, and / or CRA, can be treated as described herein.

[0022] Among the non-invasive CRC detection tests that have been described is the fecal immunochemical test (FIT) that is associated with limited sensitivity of 79% for detecting CRC and a poor sensitivity (-25%) for the detection of advanced adenomas. Lee et al., Accuracy of fecal immunochemical tests for colorectal cancer: systematic review and metaanalysis, Ann Intern Med 2014; 160(3): 171; Hundt et al., Comparative evaluation of immunochemical fecal occult blood tests for colorectal adenoma detection, Ann Intern Med 2009; 150(3): 162-9. A multi-target stool assay quantitatively examines KRAS mutations, aberrant NDRG4 and BMP3 methylation, along with -actin and hemoglobin immunoassays. This assay performs better than FIT, detecting CRC cases (-92% compared to 74%) with greater sensitivity, whereas advanced premalignant lesions were still poorly detected by both assays (-42% and -24% respectively). Imperiale et al., Multitarget stool DNA testing for colorectal-cancer screening, N Engl J Med. 2014; 370(14):1287-97. These outcomes highlight another important gap in the healthcare system: the relatively poor ability of existing non-invasive methods to detect early and advanced adenomas. The development of a diagnostic that addresses this gap could significantly improve the detection of premalignant lesions and reduce the number of colonoscopies required for average risk subjects.

[0023] DBl / 162792885.1 5 MBI-004PC / 108458-5004

[0024] In some embodiments, the subject is at low risk for CRC or colorectal polyps such as CRA or CRAA. In such embodiments, low risk individuals screened according to the present disclosure can avoid or delay more invasive colonoscopy procedures (or have such procedures less frequently). That is, the method can be performed as a screening process as an alternative to colonoscopy. According to these embodiments, low risk subjects can be screened at lower cost and at higher efficiency to the healthcare system, and subjects thereby identified where colonoscopy or other treatments are more warranted. “Low risk subjects” (as understood in the art) are subjects with no previous incidence of colorectal cancer or polyps (e.g., a colonoscopy was previously performed on the subject without detecting colorectal cancer or polyps, such as CRA or CRAA), and do not have a family history of colorectal cancer or colorectal polyps. In some embodiments, subjects at low risk do not have inflammatory bowel disease such as Crohn’s disease or ulcerative colitis. In various embodiments, the subject at low risk is at least 45 years of age, or at least 50 years of age, or at least 55 years of age, or at least 60 years of age. In some embodiments, the subject at low risk is less than 75 years of age or less than 70 years of age or less than 65 years of age. In still other embodiments, the subject is less than 45 years of age or less than 50 years of age. In embodiments, the method is performed at a determined frequency, such as at least about annually (e.g., about annually), or at least about every other year (e.g., about once every two years), or once every three to five years.

[0025] In other embodiments, the subject is high or medium risk for CRC or colorectal polyps (as understood in the art). In such embodiments, these subjects can be more frequently monitored for development of colorectal neoplasia, enabling early detection and treatment without frequent colonoscopies. For example, subjects at high or medium risk include those with prior incidence of CRC or colorectal polyps (e.g., CRA or CRAA), and / or family history of CRC or colorectal polyps. In some embodiments, subjects at high or medium risk have inflammatory bowel disease such as Crohn's disease or ulcerative colitis. In various embodiments, the method is performed at a determined frequency, such as at least about annually (e.g., about annually) or at least about every other year (e.g., once every two years), or once every three to five years. In some embodiments, the method is performed more frequently such as about twice per year. In various embodiments, the subject of high or medium risk is at least 45 years of age, or at least 50 years of age, or at least 55 years of

[0026] DBl / 162792885.1 6 MBI-004PC / 108458-5004 age, or at least 60 years of age. In some embodiments, the subject at high or medium risk is at least 65 years of age, or at least 70 years of age, or at least 75 years of age. In still other embodiments, the subject is less than 45 years of age or less than 50 years of age.

[0027] In embodiments, once a baseline biomarker profile according to this disclosure is established for the subject (either low-risk, medium-risk, or high-risk subject), subsequent testing can be scored relative to this baseline, allowing any movement of the score towards CRA, CRAA, or CRC signatures to be identified.

[0028] In various embodiments, the genetic elements from a biological sample (such as but not limited to a fecal sample) are quantified by nucleic acid sequencing, which can include genomic sequencing and / or RNA sequencing (e g., cDNA sequencing or small RNA sequencing). In various embodiments, the nucleic acid sequencing comprises shotgun metagenomic sequencing, targeted amplicon sequencing, and / or hybridization capture probe sequencing, among any other sequencing technique. Profiling genetic elements with nucleic acid sequencing can be used for training and test cohorts (e.g., to train and validate machine learning models and gene signatures), as well as for evaluating subjects for the presence or absence of the signatures.

[0029] Several studies have examined gut microbiota using either 16S rDNA, shotgun metagenomic, targeted amplicon-based or hybridization capture probe sequencing. These studies have explored fecal and mucosal-associated populations and different stages along the adenoma-carcinoma progression. A meta-analysis of fecal or other microbiota samples datasets resulted in the identification of several bacterial species enriched in CRC: Bacteroides fragilis, Fusobacterium nucleatum, Parvimonas micra, Porphyromonas assacharolytica, Prevotella intermedia, Alistipes finegoldii, and Thermoanaeroovibrio acidaminovorans. Dai, el al., Multi-cohort analysis of colorectal cancer metagenome identified altered bacteria across populations and universal bacterial markers. Microbiome 2018; 6(l):70. A separate pair of meta-analyses identified an expanded set of twenty-nine species enriched over eight distinct geographical regions. Thomas et al., Metagenomic analysis of colorectal cancer datasets identifies cross-cohort microbial diagnostic signatures and a link with choline degradation, Nat Med 2019; 25(4):667-678; Wirbel et al., Meta-

[0030] DBl / 162792885.1 7 MBI-004PC / 108458-5004 analysis of fecal metagenomes reveals global microbial signatures that are specific for colorectal cancer, Nat Med 2019; 25(4):679-689. A number of studies have analyzed the human gut microbiota associated with colonic tumors and normal adjacent tissue leading to the identification of dysbiotic signatures associated with CRC. While specific taxa vary from study to study, some common themes include the frequent identification of elevated relative abundance of E. colt, Fusobacterium nucleatum, and enterotoxin-producing Bacteroides fragilis (ETBF) strain (Arthur et al., Intestinal inflammation targets cancer-inducing activity of the microbiota. Science 2012; 338(6103): 120-3; Kostic el al., Fusobacterium nucleatum potentiates intestinal tumorigenesis and modulates the tumor-immune microenvironment. Cell Host Microbe 2013; 14(2):207-15; Wu et al., A human colonic commensal promotes colon tumorigenesis via activation of T helper type 17 T cell responses, Nat Med 2009; 15(9): 1016-22. Additional taxa associated with CRC have also been identified, but they are less uniformly observed across studies.

[0031] An important observation established by these studies is that the magnitude of difference in relative abundance derived from tissue samples is substantially greater compared to stool samples, where differentially abundant taxa are more subtle. A number of characteristics of fecal samples (and other biological samples comprising microbiota) present specific challenges in the identification of diagnostic biomarkers for the early detection of adenomas and carcinomas, including high dimensionality, data sparsity, and generalizability. Despite the massive quantity of DNA sequence data generated in shotgun metagenomic sequence analysis of stool samples, for example, the best performing biomarkers obtained in such analyses are detected in only a relatively small number of samples. This defines the problem of data sparsity and dictates that a high-performance diagnostic based on next-generation sequencing (NGS) data requires multiple independent biomarkers to compensate for low prevalence of any single biomarker in the human population.

[0032] Thus, in various embodiments, the metagenomic sequencing is deep sequencing of genomic DNA isolated from the fecal sample or other biological sample. In various embodiments, the nucleic acid sequencing involves sequencing at least about 20,000,000 reads (i.e., raw reads per sample). In various embodiments, the nucleic acid sequencing

[0033] DBl / 162792885.1 8 MBI-004PC / 108458-5004 involves sequencing at least about 25,000,000 reads, or at least about 30,000,000 reads, or at least about 40,000,000 reads, or at least about 50,000,000 reads, or at least about 60,000,000 reads, or at least about 75,000,000 reads, or at least about 100,000,000 reads per sample. Well known quality control metrics can be utilized to remove low quality reads, which are generally less than about 15%, or less than about 10% of the raw reads. Generally, the reads will have less than about 5%, or less than about 4%, or less than about 3%, or less than about 2% human reads. In various embodiments, human reads are removed from the analysis of microbial DNA.

[0034] In various embodiments, the nucleic acid sequencing comprises one or more of shotgun metagenomic sequencing, rDNA sequencing, and targeted nucleic acid sequencing (e.g., targeted amplicon sequencing or hybridization capture probe sequencing). In embodiments, the nucleic acid sequencing includes multiple workflows, for example, may comprise rDNA sequencing, and one or more of shotgun sequencing, and targeted nucleic acid sequencing. Sequencing can be conducted using any known library preparation protocol, including by employing sample tags for a multiplex workflow. See for example, U.S. Patent No. 8,603,749 and U.S. Patent No. 9,453,262, which are hereby incorporated by reference in their entireties. Profiling of genetic elements can determine either relative or quantitative abundance.

[0035] In some embodiments, library preparation from DNA samples for sequencing employs total DNA isolated from fecal or other biological samples (e.g., GI mucosal samples). Numerous kits for making sequencing libraries from DNA are available commercially. In some embodiments, library preparation comprises fragmentation of the DNA, end-repair, addition of sequencing adapters (e.g., by ligation or amplification), and amplification to enrich for products that have adapters ligated to both ends. For example, DNA can be fragmented such that the mean fragment size is in the range of 100 base pairs to about 5000 base pairs, such as in the range of about 250 bps to about 4000 bps, or the range of about 500 bps to about 3000 bps, or in the range of about 1000 bps to about 3000 bps. In some embodiments, the mean fragment size is less than 1000 bps, such as in the range of 200 to 1000 bps (e.g., 200 to 500 bps). To facilitate multiplexing, different barcoded adapters can be used with different biological samples (e.g., from different subjects). In some

[0036] DBl / 162792885.1 9 MBI-004PC / 108458-5004 embodiments, barcodes can be introduced at the PCR amplification step by using different barcoded PCR primers to amplify different biological samples. The library may be subject to shotgun metagenomic sequencing in some embodiments.

[0037] In some embodiments, the nucleic acid sequencing focuses on one or more genomic loci to allow for taxonomic analysis. Taxonomic analysis can include profiling at various levels, e.g., Kingdom, Phylum, Class, Order, Family, Genus, and Species. In embodiments, taxonomic analysis includes strain level analysis. In embodiments, taxonomic classification includes analysis of bacteria, fungi, archaea, and in embodiments viruses (e.g., components of the virome including bacteriophages and eukaryotic viruses, as known in the art).

[0038] In embodiments, taxonomic analysis includes rDNA analysis. Analysis of rDNA (genes encoding rRNA) can include 16S rDNA, 18S rDNA, and internal transcribed spacer (ITS) sequencing. 16S and ITS sequence analysis allows for taxonomic analysis of bacteria and archaea, and 18S and ITS sequence analysis allows for taxonomic analysis of eukaryotes (e.g., fungi). The 16S rRNA gene comprises nine variable regions interspersed throughout the highly conserved 16S sequence. In some embodiments, sub-regions of the gene are amplified by targeted PCR for sequencing, ranging from single variable regions, such as V4 or V6, to three variable regions, such as VI to V3 or V3 to V5. Similarly, 18S rRNA genes comprise variable regions (VI to V9) which can be used to discriminate at the family, order, genus, and species (and sub-species) levels as is known in the art. In some embodiments, sub-regions of the gene are amplified by targeted PCR for sequencing, ranging from single variable regions to a plurality of variable regions. The ITS lies between the large and small rRNA subunit gene loci and can be species specific. This polymorphism is due to the presence of tRNA genes. The ITS region can be amplified by targeted PCR for sequencing and taxonomic analysis.

[0039] In some embodiments, 16S / 18S / ITS sequences are clustered based on similarity to generate operational taxonomic units (OTUs). Representative OTU sequences can be compared with reference databases to determine taxonomy. In some embodiments, sequences of >95% identity are considered to represent the same genus, whereas sequences of >97% identity are considered to represent the same species. Methods of determining

[0040] DBl / 162792885.1 10 MBI-004PC / 108458-5004

[0041] OTUs are known in the art. In some embodiments, strain or subspecies are further distinguished based on analysis of polymorphisms. Taxonomic analysis of 16S, 18S, and ITS DNA sequences is well known in the art. See Ze-Gang Wei et al., Comparison of Methods for Picking the Operational Taxonomic Units from Amplicon Sequences, Front. Microbiol., 24 March 2021.

[0042] Sequence reads other than rDNA can also be analyzed to infer taxonomy by comparison to reference microbial genomes. Helene LCF, et al., New Insights into the Taxonomy of Bacteria in the Genomic Era and a Case Study with Rhizobia, International Journal of Microbiology Vol. 2022.

[0043] In various embodiments, targeted genomic fragments are captured from a metagenomic library, optionally followed by amplification. For example, nucleic acid capture probes can be used that hybridize to conserved regions of rDNA or conserved regions of functional orthologs. Sequence capture allows targeted enrichment of informative DNA. In concert with NGS, capture provides an efficient strategy for high-throughput screening of regions of interest. In various embodiments, a capture strategy reduces the required sequencing depth to less than about 25,000,000 reads, or less than about 20,000,000 reads, or less than about 15,000,000 reads, or less than about 10,000,000 reads, or less than about 5,000,000 reads, or less than about 2,000,000 reads. An exemplary sequence capture protocol comprises: fragmentation of input DNA (e.g., by shearing or with use of enzymes); addition of sequencing adapters (e.g., by ligation or amplification using fusion primers) to form library molecules; incubating the library with pools of capturable oligonucleotide probes designed to target (and hybridize to) specific regions of interest within the DNA fragment library. An exemplary capturable moiety is biotin, which can be conjugated to probe oligonucleotides. Probe / target hybrids are then captured from the library (e.g., using streptavidin-coated magnetic beads). The result is a sequencing-ready library that is highly enriched for the targeted DNA.

[0044] In embodiments, and particularly for evaluating samples for the presence of absence of signatures, genetic elements can be quantified by PCR (qPCR) or microarray according to known processes. For example, genus-specific or species-specific sequences can be

[0045] DBl / 162792885.1 11 MBI-004PC / 108458-5004 quantitatively amplified and detected (e g., from rDNA in the sample) as well as indicative sequences (which may be conserved sequences) in gene function elements. In this manner, abundance profiles of genetic elements (e.g., informative features) can be constructed without a sequencing workflow. See WO 2024 / 206741, which is hereby incorporated by reference in its entirety.

[0046] In embodiments, the genetic elements can be profiled using quantitative microbiome profiling (QMP) or Relative Microbiome Profiling (RMP). By quantifying microbial loads using techniques like quantitative PCR, QMP can provide an accurate picture of microbial communities. In embodiments, known quantities of DNA or microbial cells are added to samples to provide a reference for absolute quantification.

[0047] In various embodiments, metagenomic sequence reads are annotated for microbial taxonomy and / or gene classification or function. Such annotation can be conducted using any of the publicly available tools and databases (including those available from Huttenhower Lab, Harvard T.H Chan School of Public Health).

[0048] In embodiments, metagenomic sequence reads are annotated to identify microbial communities. Exemplary tools for profiling microbial communities from metagenomic shotgun sequencing data include MetaPhlAn and StrainPhlAn (Huttenhower Lab). MetaPhlAn is a computational tool for profiling the composition of microbial communities (Bacteria, Archaea and Eukaryotes) including at species-level. With StrainPhlAn and similar tools, it is possible to perform strain-level microbial profiling. MetaPhlAn relies on unique clade-specific marker genes identified from ~1M microbial genomes. Bioinformatics tools for strain level analysis may include one or more of StrainScan, Krakenuniq, StrainSeeker, and StrainPhlAn.

[0049] In embodiments, sequence reads are annotated for predicted function, using any number of available tools. For example, HUMAnN3 (Huttenhower Lab) can be used for functional annotation, providing insights into the metabolic pathways and gene functions represented in the microbiome. HUMAnN can be used for profiling the abundance of microbial metabolic pathways and other molecular functions from metagenomic or metatranscriptomic sequencing data. Other available tools include MelonnPan (for

[0050] DBl / 162792885.1 12 MBI-004PC / 108458-5004 predicting metabolite composition from microbiome sequencing data), ShortBRED (for profiling protein families in shotgun sequencing data), MetaWIBELE (to identify potentially bioactive gene products in microbial communities), FUGAsseM (to predict function of uncharacterized gene products in microbial communities), and MACARRoN (to identify potentially bioactive novel metabolites in microbial communities). Other available tools known in the art include DIAMOND, EIMMER, MetaCyc, MGnify, CAZy, and the NCBI Metagenome Database, and others.

[0051] In embodiments, metagenomic sequence reads are annotated to identify viruses, including eukaryotic viruses and bacteriophages, and which may include bacteriophages specific for Fusobacterium and Porphyromonas species, or butyrate-producing species. Bioinformatics tools for analyzing the virome include ViroProfiler, Hecatomb, VIGA, ViromeFlowX, and VIROME.

[0052] In exemplary embodiments, metagenomic sequence reads are annotated for Enzyme Commission (EC) number. EC number is a numerical classification scheme for enzymes, based on the chemical reactions they catalyze. Table 8 shows EC number features statistically associated with the presence of CRC as compared to controls, as determined from metagenomic sequencing data.

[0053] In exemplary embodiments, sequence reads are annotated for gene orthology relationships. For example, eggNOG (evolutionary genealogy of genes: Non-supervised Orthologous Groups) is a public resource to establish orthology relationships between genes identified from the metagenomic sequence reads. Table 9 shows features from the eggNOG resource (based on metagenomic sequence data) that are indicative of CRC.

[0054] In exemplary embodiments, metagenomic sequence reads are annotated based on gene homology. Homology can be determined at various levels of identity to seed sequences, such as 50% or greater, 70% or greater, 80% or greater, or 90% or greater. Table 10 shows UniRef90 features (at least 90% identity to microbial seed sequences in the database) that are indicative of the presence of CRC as compared to controls, as determined from metagenomic sequencing data.

[0055] DBl / 162792885.1 13 MBI-004PC / 108458-5004

[0056] In exemplary embodiments, metagenomic sequence reads are annotated for gene ontology (GO) (gene product function). Table 11 shows GO features associated with the presence of CRC as compared to controls, as determined from metagenomic sequencing data.

[0057] In exemplary embodiments, the metagenomic sequence reads are annotated according to the KEGG Orthology database, or similar database. The KEGG Orthology (KO) database is a database of molecular functions represented in terms of functional orthologs. A functional ortholog is manually defined in the context of KEGG molecular networks, namely, KEGG pathway maps, BRITE hierarchies and KEGG modules. Each node of the network, such as a box in the KEGG pathway map, is given a KO identifier (called K number) as a functional ortholog defined from experimentally characterized genes and proteins in specific organisms, which are then used to assign orthologous genes in other organisms based on sequence similarity. The resulting KO grouping may correspond to a group of highly similar sequences within a limited organism group or it may be a more divergent group. Table 12 shows KO features associated with the presence of CRC as compared to controls, as determined from metagenomic sequencing data.

[0058] In exemplary embodiments, the metagenomic sequence reads are annotated according to pathways (e.g., KEGG pathway database, or similar database). Table 13 shows pathway features associated with the presence of CRC as compared to controls, as determined from metagenomic sequencing data.

[0059] In exemplary embodiments, the metagenomic sequence reads are annotated according to protein families and domains (e.g., Pfam database, or similar database). Table 14 shows Pfam database features associated with the presence of CRC as compared to controls, as determined from metagenomic sequencing data.

[0060] In exemplary embodiments, the metagenomic sequence reads are annotated according to gene product enzymatic reactions. Table 15 shows reaction database features associated with the presence of CRC as compared to controls, as determined from metagenomic sequencing data.

[0061] DBl / 162792885.1 14 MBI-004PC / 108458-5004

[0062] In exemplary embodiments, the metagenomic sequence reads are annotated according to microbial taxonomy, including at various taxonomic levels (including optionally genus, species, and / or strain). Table 16 shows taxonomic features associated with the presence of CRC as compared to controls, as determined from metagenomic sequencing data.

[0063] In embodiments, the taxonomic and functional annotated data are batch corrected before feature selection or analysis of the genetic elements profile. Batch effect correction is a statistical approach used to remove systematic variations in data that arise from differences in experimental conditions, sample processing, or equipment across different batches of data. This correction ensures that the observed biological differences in data are tally reflective of the underlying biological variation rather than artifacts introduced by technical discrepancies. Batch effects in microbiome data can arise from sampling and processing differences in studies especially when data is collected from multiple metagenomic studies from around the world or from different collection sites in a single country. Available resources for batch correction of metagenomic data include but are not limited to MMUPHin, DeBias-M, and ConQur. Meta-analysis datasets can be evaluated before and after batch correction, including with use of publicly available tools, including Principal Component Analysis (PCA), ADONIS, and Isolation Forest Algorithms. Meta- analysis datasets can be evaluated before and after batch correction including with Leave-1 - Out Modeling, Anomaly Analysis (box plots of study vs Log2 cross validation), and others. Using these packages and tools, an optimum meta-analysis dataset can be prepared for feature selection, including by use of logistic regression in some embodiments.

[0064] In embodiments, logistic regression can be performed across all features (taxonomic, functional, KEGG pathways, etc.) while controlling for study ID as a covariate. This helps in identifying features significantly associated with the condition or outcome of interest. In embodiments, logistic regression is used to identify statistically significant features (e.g., for feature reduction before constructing full machine learning models), as described more fully herein. Selected features may have a p-value of 0.05 or less, or 0.01 or less, or 0.005 or less, or 0.001 or less, in their association with the disorder of interest (e g., CRC, CRA, or CRAA).

[0065] DBl / 162792885.1 15 MBI-004PC / 108458-5004

[0066] For the detection of the disorder (e.g., CRC, CRAA, and / or CRA), the number of genetic elements quantified and evaluated in a test sample will be sufficient to provide for a high performance test (e.g., by allowing for the analysis of numerous informative features). For example, in various embodiments, the genetic elements can be analyzed (e.g., with respect to each relevant model) for the presence of at least about 50 features, or at least about 100 features, at least about 200 features, at least about 500 features, or at least about 800 features, or at least about 1000 features. Exemplary features for detecting CRC are displayed in Tables 8 to 16 (with an ensemble feature set shown in Table 17). Exemplary features for detecting CRA are displayed in Table 18. Exemplary features for detecting CRAA are displayed in Table 19. The number of features for each test need not be the same for each model. For example, in some embodiments the model or “signature” for detecting CRC, CRA, or CRAA may include at least about 50 features, or at least about 100 features, or at least about 200 features, or at least about 500 features or at least about 750 features, or at least about 1000 features. In some embodiments, the models or signatures for detecting CRC, CRA, or CRAA include less than about 500 features, such as less than about 250 features, or less than about 200 features, or less than about 150 features, or less than about 100 features. In embodiments, the number of features is in the range of 50 to 200 features. Models with more or less features can nevertheless be constructed according to the present disclosure.

[0067] In various embodiments, a portion of the genetic elements are associated with (indicative of) colorectal cancer (CRC). For example, the genetic elements can comprise one or more taxonomic or gene function features listed in Tables 8 to 16. In various embodiments, the genetic elements comprise at least five taxonomic or gene function features listed in Tables 8 to 16. In some embodiments, the genetic elements comprise at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or gene function features listed in Tables 8 to 16. In various embodiments, the genetic elements comprise at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 taxonomic features listed in Tables 8 to 16; and at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 gene function features listed in Tables 8 to 16. In some embodiments, at least five, or at least ten, or at least 20 genetic elements for detecting CRC correspond to bacterial species that generally reside in the oral

[0068] DBl / 162792885.1 16 MBI-004PC / 108458-5004 cavity. In some embodiments, the difference in relative abundance between disease and nondisease samples (or vice versa) is at least about 1.1 fold, or at least about 1.2 fold, or at least about 1.3 fold, or at least about 1.4 fold, or at least about 1.5 fold, or at least about 2 fold. In some embodiments, the difference in prevalence between disease and non-disease samples (or vice versa) is at least about 1.1 fold, or at least about 1.2 fold, or at least about 1.3 fold, or at least about 1.4 fold, or at least about 1.5 fold, or at least about 2 fold.

[0069] In embodiments, the genetic elements associated with (indicative of) CRC include features of microbial taxonomy (e.g., sequences annotated for taxonomy, for example at genus, species, and / or strain level), along with at least 2, at least 3, at least 4, or at least 5 selected from: (1) features defining an enzyme or reaction classification (e.g., sequences annotated by EC number or enzymatic reaction), (2) features defining gene orthology or ontology (e.g., sequences annotated with gene orthology or gene ontology database), (3) features defining gene homology (e.g., sequences annotated with UniRef90 database), (4) features defining pathways (e.g., sequences annotated with KEGG pathway database), and (5) features defining protein families and / or domains (e.g., sequences annotated with a protein family or protein domain database).

[0070] For example, the genetic elements can comprise one or more taxonomic or gene function features listed in Table 17. In various embodiments, the genetic elements comprise at least five taxonomic or gene function features listed in Table 17. In some embodiments, the genetic elements comprise at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or gene function features listed in Table 17. In various embodiments, the genetic elements comprise at least one, at least two, at least five, at least about 10, at least about 20 taxonomic features listed in Table 17; and at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 gene function features listed in Table 17.

[0071] In various embodiments, at least a portion of the genetic elements analyzed are associated with (indicative of) colorectal adenoma (CRA). For example, genetic elements can comprise one or more taxonomic or gene function features listed in Table 18, including annotations of genetic elements as described. In various embodiments, the genetic elements

[0072] DBl / 162792885.117 MBI-004PC / 108458-5004 comprise at least five taxonomic or gene function features listed in Table 18. In some embodiments, the genetic elements comprise at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or gene function features (e.g., which may be listed in Table 18). In various embodiments, the genetic elements comprise at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 taxonomic features; and at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 gene function features (e.g., as may be listed in Table 18). Certain genetic elements may have differential abundance in samples (e g., fecal samples or other biological samples) from CRA subjects (as compared to controls), and other genetic elements may have differential prevalence in fecal or other biological samples from CRA subjects (as compared to control subjects). In some embodiments, the genetic elements include a plurality of those having differential abundance in CRA, and a plurality of those having differential prevalence in CRA. In some embodiments, the difference in relative abundance between disease and nondisease samples (or vice versa) is at least about 1.1 fold, or at least about 1.2 fold, or at least about 1.3 fold, or at least about 1.4 fold, or at least about 1.5 fold, or at least about 2 fold. In some embodiments, the difference in prevalence between disease and non-disease samples (or vice versa) is at least about 1.1 fold, or at least about 1.2 fold, or at least about 1.3 fold, or at least about 1.4 fold, or at least about 1.5 fold, or at least about 2 fold. Other exemplary features associated with CRA which may be employed according to this paragraph are disclosed in WO 2024 / 206741, which is hereby incorporated by reference in its entirety.

[0073] In embodiments, the genetic elements associated with (indicative of) CRA include features of microbial taxonomy (e.g., sequences annotated for taxonomy, for example at genus, species, and / or strain level), along with at least 2, at least 3, at least 4, or at least 5 selected from: (1) features defining an enzyme or reaction classification (e.g., sequences annotated by EC number or enzymatic reaction), (2) features defining gene orthology or ontology (e g., sequences annotated with gene orthology or gene ontology database), (3) features defining gene homology (e.g., sequences annotated with UniRef90 database), (4) features defining pathways (e.g., sequences annotated with KEGG pathway database), and (5) features defining protein families and / or domains (e.g., sequences annotated with a protein family or protein domain database).

[0074] DBl / 162792885.1 18 MBI-004PC / 108458-5004

[0075] In various embodiments, the genetic elements are associated with (indicative of) colorectal advanced adenoma (CRAA). For example, the genetic elements can comprise one or more taxonomic or gene function features listed in Table 19, and the genetic elements may be annotated as described. In various embodiments, the genetic elements comprise at least five taxonomic or gene function features listed in Table 19. In some embodiments, the genetic elements comprise at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or gene function features (e.g., as may be listed in Table 19). In various embodiments, the genetic elements comprise at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 taxonomic features; and at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 gene function features (e.g., as may be listed in Table 19). Certain genetic elements may have differential abundance in fecal or other biological samples from CRAA subjects (as compared to controls), and other genetic elements may have differential prevalence in fecal or other biological samples from CRAA subjects (as compared to control subjects). In some embodiments, the genetic elements include a plurality of those having differential abundance in CRAA, and a plurality of those having differential prevalence in CRAA. In some embodiments, the difference in relative abundance between disease and non-disease biological samples (or vice versa) is at least about 1.1 fold, or at least about 1 .2 fold, or at least about 1.3 fold, or at least about 1.4 fold, or at least about 1.5 fold, or at least about 2 fold. In some embodiments, the difference in prevalence between disease and non-disease biological samples (or vice versa) is at least about 1.1 fold, or at least about 1.2 fold, or at least about 1.3 fold, or at least about 1.4 fold, or at least about 1.5 fold, or at least about 2 fold. Other exemplary features associated with CRA are disclosed in WO 2024 / 206741, which is hereby incorporated by reference in its entirety.

[0076] In embodiments, the genetic elements associated with (indicative of) CRAA include features of microbial taxonomy (e.g., sequences annotated for taxonomy, for example at genus, species, and / or strain level), along with at least 2, at least 3, at least 4, or at least 5 selected from: (1) features defining an enzyme or reaction classification (e.g., sequences annotated by EC number or enzymatic reaction), (2) features defining gene orthology or

[0077] DBl / 162792885.1 19 MBI-004PC / 108458-5004 ontology (e.g., sequences annotated with gene orthology or gene ontology database), (3) features defining gene homology (e.g., sequences annotated with UniRef90 database), (4) features defining pathways (e.g., sequences annotated with KEGG pathway database), and (5) features defining protein families and / or domains (e.g., sequences annotated with a protein family or protein domain database).

[0078] In various embodiments, the profile of genetic elements (which is optionally an abundance profile) is evaluated for signatures indicating the presence or absence of each of CRC, CRA, and CRAA. As disclosed herein, the microbiome profile for CRC, CRA, and CRAA do not exhibit linear relationship with one another, and therefore each can be evaluated using separate models or signatures.

[0079] In various embodiments, the signature(s) are generated from a training set (e.g., of metagenomic data) by machine learning (ML). For example, the signature indicating the presence or absence of CRA is trained with fecal or other biological samples from a CRA cohort and biological samples from a control cohort. The signature indicating the presence or absence of CRAA is trained with fecal or other biological samples from a CRAA cohort and biological samples from a control cohort. The signature indicating the presence or absence of CRC is trained with fecal or other biological samples from a CRC cohort and biological samples from a control cohort. In each instance, the control cohort is considered a healthy cohort, that is, defined by the absence of CRC, CRA, and CRAA. In embodiments, the signature indicates the presence or absence of diagnostic target variables classified as Aneo (CRC + CRAA; positive) and Naneo (CRA + Control; negative), in a binary classification context, which is trained with samples from CRC / CRAA and CRA / Control cohorts by machine learning. In some embodiments, samples of the control cohort are not from subjects having significant gastrointestinal ailments, such as Crohn’s disease or ulcerative colitis.

[0080] According to the various aspects and embodiments of this disclosure, the training set comprises at least about 50 samples, or at least about 100 samples, or at least about 150 samples, or at least about 200 samples, or at least about 500 samples, or at least about 1000 samples that are positive for CRA, CRAA, or CRC. In some embodiments, the training set

[0081] DBl / 162792885.1 20 MBI-004PC / 108458-5004 comprises at least about 25 non-disease or healthy controls, or at least about 50 non-disease or healthy controls, or at least about 100 non-disease or healthy controls, or at least about 500 non-disease or healthy controls, or at least about 1000 non-disease or healthy controls. One of skill in the art will be able to assemble training sets representing disease and control samples in a manner that results in adequate statistical powering. The training set need not be sourced from a single study or geographic area. In some embodiments, biological samples are sourced and / or processed at different geographies (e.g., at least two different countries or continents). In these embodiments, the separate procurement, processing, or sequencing provides added diversity of research protocols, and may also provide subject genetic, ethnic, age-related, and / or environmental variation (including variation in diet).

[0082] Signatures can be trained using one or a plurality of machine learning algorithms. In some embodiments, machine learning algorithms are trained with selected statistically significant features, such as those selected using logistic regression as described. In some embodiments, at least one of the machine learning algorithms utilized is a supervised machine learning algorithm. In these or other embodiments, the machine learning algorithms comprise one or more of unsupervised or semi-supervised machine learning. Various machine learning algorithms are known and can be used according to the present disclosure, including but not limited to one or more of parametric / non-parametric distance measures, logistic regression, support vector machines, decision trees, random forests, neural networks, probit regression, Fisher's linear discriminant, Naive Bayes classifier, perceptron, quadratic classifiers, kernel estimation, k-nearest neighbor, learning vector quantization, and principal components analysis. In embodiments, the machine learning employs an Al-enabled, massively parallel computational and automated machine learning platform and workflow (computer program) for comparative machine learning modeling, optimization, testing, evaluation, and ranking of models, such as one or more of deep learning, gradient boosted, neural networks, ensemble, or blender modeling algorithms, such as but not limited to Gradient Boosted Trees Classifier, extreme Gradient Boosted Trees Classifiers, Light Gradient Boosted Trees Classifiers, Light Gradient Boosting on Elastic Net Predictions, Keras Slim Residual Neural Network Classifiers, Generalized Additive Models, Elastic Net Classifiers, Random Forest Classifiers, Deep Forest Classifiers, Average Blender

[0083] DBl / 162792885.1 21 MBI-004PC / 108458-5004

[0084] Classifiers, TensorFlow Multilayer Perceptron Classifiers, TensorFlow Neural Network Classifiers, and Rule-Fit Classifiers.

[0085] Feature reduction prior to full model construction can use various tools. In some embodiments, features are selected from training cohorts by ensemble ranking of feature importance, which ranks features according to their importance in a predictive model. Alternatively or in addition, features are selected according to their statistical significance (individually) for predicting CRA, CRAA, or CRC. For example, individual features can be selected whose abundance or prevalence is predictive of the presence or absence of CRC, CRA, or CRAA with a p-value less than or equal to 0.05 in the training group, or a p-value less than or equal to 0.01, or a p-value less than or equal to 0.005, or a p-value less than or equal to 0.001 in the training group (or other selected statistical threshold). In embodiments, such statistically significant features are identified using logistic regression.

[0086] In embodiments, diagnostic models or signatures are prepared by machine learning analyses by dividing data into groups based on feature type (e.g., functional annotation, taxonomic composition, KEGG pathways, etc.), and models can be selected based on performance metrics such as accuracy, AUC, logloss, and other relevant criteria as known in the art. In embodiments, backward feature elimination is performed to refine the models and select those with the fewest features that maintain strong performance metrics. In embodiments, model construction comprises feature merging, that is, combining features from top-performing models across different analysis groups to create a final, comprehensive model, and optionally further applying backward feature elimination. In embodiments, the model construction further comprises Permutation Testing, which involves conducting permutation tests on a holdout set (HO) to validate the model's generalization capability and ensure that the performance holds up across different data permutations. Multifold External Holdout Testing involves conducting permutation tests on external holdout sets (XHO) that are reserved and are not used for model training purposes.

[0087] In embodiments, independent signatures are created for different geographic regions (e g., continents, or countries), to allow the signatures to be further tailored for subpopulations of subjects. For example, such signatures can involve different microbial

[0088] DBl / 162792885.1 22 MBI-004PC / 108458-5004 taxonomic features and / or different gene function features for different geographic regions. Geographic regions can be defined by different continents or different countries or different regions. In embodiments, signatures are created that are generalizable across different geographic regions, including globally, by selecting features that are informative independent of geography and / or ethnicity (or subject age) for example.

[0089] Accuracy of a model can be assessed using the standard Receiving Operator Characteristics (ROC) curve analysis to calculate true positive, false positive, true negative and false negative rates, overall accuracy and area under the curve. The term “ROC” or “ROC curve,” refers to a Receiver Operator Characteristic curve. A ROC curve can be a graphical representation of the performance of a binary classifier system. For any given method, a curve can be generated by plotting the sensitivity against the specificity at various threshold settings. Furthermore, provided at least one of three parameters (e.g., sensitivity, specificity, and the threshold setting), a ROC curve can determine the value or expected value for any unknown parameter. The unknown parameter can be determined using a curve fitted to a ROC curve. For example, provided the presence / absence or abundance of one or more features, the expected sensitivity and / or specificity of a test can be determined. The term “AUC” or “ROC-AUC” can refer to the area under a receiver operator characteristic curve. This metric can provide a measure of diagnostic utility of a method, considering both the sensitivity and specificity of the method. A ROC-AUC can range from 0.5 to 1 .0, where a value closer to 0.5 can indicate a method has limited diagnostic utility (e.g., lower sensitivity and / or specificity) and a value closer to 1.0 indicates that the method has greater diagnostic utility (e.g., higher sensitivity and / or specificity).

[0090] In various embodiments, the signature(s) have a sensitivity for classifying samples for the presence or absence of CRC, CRA, or CRAA of at least about 0.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about 0.95. In embodiments, the signature(s) have a specificity for classifying samples for the presence or absence of CRC, CRA, and / or CRAA of at least about 0.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about 0.95. For example, the signatures can classify each of CRC, CRA, and / or CRAA with a sensitivity of at least 0.75, and a specificity of at least 0.75. For example, the signatures may have an area under the curve (AUC) for classifying samples for

[0091] DBl / 162792885.1 23 MBI-004PC / 108458-5004 the presence or absence of CRC, CRA, and / or CRAA of at least about 0.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about 0.95.

[0092] In various embodiments, if the subject is not identified according to the process described herein as likely to have CRA, CRAA, or CRC, no further procedure is conducted. That is, the subject is not scheduled for a colonoscopy or other evaluation for colorectal cancer or adenoma. The test can be repeated at some frequency, such as about once every year to about once every five years. In embodiments, once a baseline biomarker profile is established for the subject, subsequent testing can be relatively scored against this baseline, allowing any movement of the score towards CRA, CRAA, or CRC signatures to be identified. This time-interval testing can further increase the accuracy of the test. Where the subject is identified as likely having one or more of CRA, CRAA, or CRC according to the processes described herein, a tailored diagnostic or treatment plan is initiated. For example, the subject can undergo a procedure that involves imaging of the colon, such as colonoscopy or CT coIonography (or other scan or imaging technique) to confirm the result, which can also involve removal of one or more polyps and / or obtaining a biopsy of growths suspected of involving colorectal cancer. Where colorectal cancer is confirmed, the subject is treated for CRC. For example, the subject can undergo one or more of surgery (e.g., cancer resection, including partial colectomy in some embodiments), chemotherapy, radiation therapy, and immunotherapy for colorectal cancer. Exemplary chemotherapy or immunotherapy for colorectal cancer may include one or more of 5-fluorouracil (5-FU), capecitabine (XELODA) (which is metabolized by the tumor to 5-FU), irinotecan, leucovorin, oxaliplatin, cetuximab, panitumumab, regorafenib, bevacizumab, aflibercept, and ramucirumab. Exemplary combination therapies further comprise FOLFOX (5-FU, leucovorin, and oxaliplatin), FOLFIRI (leucovorin, 5-FU, and irinotecan), CAPEOX (capecitabine and oxaliplatin), FOLFOXIRI (leucovorin, 5-FU, oxaliplatin, and irinotecan), 5-FU with leucovorin or capecitabine alone, and trifluridine and tipiracil combination (LONSURF). In some embodiments, the subject receives an immune checkpoint inhibitor, such as an antibody or other molecule that inhibits PD-1, PD-L1, PD-L2, or cytotoxic T- lymphocyte-associated protein 4 (CTLA-4). In embodiments, selection of treatments for the subject can be, in part, facilitated by the testing, which can include testing for biomarker signatures correlated with drug response or toxicity (as already described).

[0093] DBl / 162792885.1 24 MBI-004PC / 108458-5004

[0094] Further, radiation therapy can be used in conjunction with resection, chemotherapy, immunotherapy, or alone. Types of radiation therapy include External-Beam Radiation Therapy (EBRT), Internal Radiation Therapy (brachytherapy), Endocavitary radiation therapy, Interstitial brachytherapy, and Radioembolization.

[0095] In aspects, the present disclosure provides a method and systems for preparing a genetic signature of genetic elements (i.e., informative features) indicative of the presence of a disorder, such as colorectal neoplasm or other disorder (e.g., including but not limited to a colon disorder). The method comprises providing a training cohort of fecal or other samples from subj ects confirmed to have the disorder of interest (e.g., CRA, CRAA, or CRC, or other disorder of interest) (or providing RNA or DNA isolated therefrom), and conducting genomic nucleic acid sequencing of DNA isolated from the fecal or other biological samples as already described. In various embodiments, the disorder of interest is a colon disorder selected from Crohn’s disease, ulcerative colitis, irritable bowel syndrome (IBS), diverticulitis, CRA, CRAA, and CRC. Other disorders include any of those in which the microbiome composition from a patient sample can be indicative of the disorder, and numerous such conditions are known in the art, including various inflammatory and autoimmune disorders. In embodiments, a genetic signature is created to identify and / or characterize (e.g., classify) dysbiosis (e.g., gut health), identify presence of cancers of other (non-intestinal) tissues (e.g., stomach, esophageal, liver, kidney, pancreas, lung, CNS), characterize (e.g., classify) liver and kidney dysfunction, autoimmune or inflammatory disease (e.g., Rheumatoid Arthritis, multiple sclerosis, fibromyalgia, fibrosis), cardiovascular disease, metabolic syndrome (e.g., obesity and insulin resistance), and CNS disorders (e.g., Alzheimer’s disease, Parkinson’s disease, or Lewy body dementia), including identifying and classifying for subtypes thereof. In embodiments, the signature is used to predict drug response / toxicity predictions, including for chemotherapy or immunotherapy (e.g., immune checkpoint inhibitor therapy, such as PD-1 / PD-L1 blockade or CTLA-4 inhibitor).

[0096] In embodiments, the microbial metagenomic markers are combined with other markers, creating a “multi-omic” test, which in embodiments includes analysis of human genomic markers, transcriptomic markers (microbial and / or human), metabolic markers,

[0097] DBl / 162792885.1 25 MBI-004PC / 108458-5004 proteomic markers, and immune system markers. Human genomic markers in embodiments include genetic and epigenetic markers. Genetic markers can include tumor mutational markers and epigenetic markers, of which numerous are known in the art, including mutations in KRAS, NRAS, TP53, BRAF, PIK3CA, APC, AKT1, ERBB2, FBXW7, KIT, MLH1, MSH2, MSH6, PMS2, MUTYH, SMARCB1, SMO, and STK11. See, Youssef O., et al. Gene mutations in stool from gastric and colorectal neoplasia patients by nextgeneration sequencing. World J Gastroenterol . 2017 Dec 21;23 (47): 8291-8299; Armaghany T. et al. Genetic alterations in colorectal cancer. Gastrointest Cancer Res. 2012 Jan;5(l):19- 27; Ciepiela, I., et al. Tumor location matters, next generation sequencing mutation profiling of left-sided, rectal, and right-sided colorectal tumors in 552 patients. Sci Rep 14, 4619 (2024). Exemplary epigenetic markers include markers of aberrant methylation (such as of NDRG4, BMP3, MLH1, SFRP1, and SEPT9, among others), and altered histone modifications, such as methylation, acetylation, and phosphorylation (e.g., H3K79me2). Metabolic markers (including for use in fecal samples) include levels of volatile organic compounds, short-chain fatty acids (e.g., lower levels of butyrate, acetate, and propionate), bile acids (e.g., altered levels of allo-cholic acid, allo-DCA, and / or allo-LCA), polyamines (e.g., elevated putrescine, cadaverine, spermidine, and modified derivatives) and certain amino acids (e.g., elevated levels of branched-chain amino acids, leucine, isoleucine, valine, proline, isovalerate, isobutyrate, valerate, glutamate, and phenylacetate), which can reflect the metabolic and microbial changes that accompany the development of colorectal cancer (CRC) and its precursors. Other metabolites include levels of Kreb cycle metabolites, such as succinate. Exemplary proteomic and / or inflammatory markers include elevated levels of one or more of fecal calprotectin (FC), cathepsin (CAT), lactoferrin (LTF), matrix metalloproteinase-9 (MMP9), retinol-binding protein 4 (RBP4), serpin peptidase inhibitor clade A member 3 (SERPINA3), and S100A6. Further, systemic inflammatory markers that can be measured in samples include CRP, IL-6, TNF-a, IL-ip, and zonulin. Other markers include Trimethylamine N-oxide (TMAO), Phenylacetylglutamine (PAG), uremic toxins, and Secondary bile acids (SB As), which can be used as markers of increased cardiovascular risk among other things. Metabolic markers can be quantified in samples using known methods, such as LC-MS. Proteomic markers can also be quantified using known methods,

[0098] DBl / 162792885.1 26 MBI-004PC / 108458-5004 such as ELISA. Zhgun ES and Ilina EN. Fecal Metabolites As Non-Invasive Biomarkers of Gut Diseases. Acta Naturae. 2020 Apr-Jun; 12(2):4-14.

[0099] In embodiments, the microbial markers described herein are used along with a fecal occult blood test (FIT).

[0100] The metagenomic sequence data is annotated as described, for taxonomic and gene information, and which can be subject to batch correction as already described. Features statistically associated with the disorder (such as but not limited to CRC, CRA, and CRAA) can be identified by processes such as logistic regression and / or large language models to provide for feature reduction prior to machine learning model construction. A gene signature is then trained using selected statistically significant features by machine learning processes (as described), and which classifies samples for the presence or absence of the disorder (e.g., CRA, CRAA, or CRC), and which in embodiments identifies subtypes of complex disorders.

[0101] For example, the signature is generated from a training set (e.g., of metagenomic data) by machine learning (e.g., supervised machine learning). For example, the signature indicating the presence or absence of the disorder is trained with fecal and / or other biological samples from a cohort having the disorder of interest, and biological samples from a control cohort. The control cohort is considered a healthy cohort, that is, defined by the absence of the disorder. In embodiments, the training set comprises at least about 50 samples, or at least about 100 samples, or at least about 150 samples, or at least about 200 samples, or at least about 500 samples, or at least about 1000 samples that are positive for the disorder of interest. In some embodiments, the training set comprises at least about 25 non-disease or healthy controls, or at least about 50 non-disease or healthy controls, or at least about 100 non-disease or healthy controls, or at least about 500 non-disease or healthy controls, or at least about 1000 non-disease or healthy controls. One of skill in the art will be able to assemble training sets representing disease and control samples in a manner that results in adequate statistical powering. The training set need not be sourced from a single study or geographic area. In some embodiments, biological samples are sourced and / or processed at different geographies. In these embodiments, the separate procurement, processing, or

[0102] DBl / 162792885.1 T1 MBI-004PC / 108458-5004 sequencing provides added diversity of research protocols, and may also provide subject genetic, ethnic, and / or environmental variation (including variation in diet).

[0103] Signatures can be trained using one or a plurality of machine learning algorithms as already described. Feature reduction prior to full model construction can use various tools. In some embodiments, features are selected from training cohorts by ensemble ranking of feature importance, which ranks features according to their importance in a predictive model. Alternatively or in addition, features are selected according to their statistical significance (individually) for predicting the disorder of interest (e.g., using logistic regression or other tool). For example, individual features can be selected whose abundance or prevalence is predictive of the presence or absence of the disorder with a p-value less than or equal to 0.05 in the training group, or a p-value less than or equal to 0.01, or a p-value less than or equal to 0.005, or a p-value less than or equal to 0.001 in the training group (or other selected statistical threshold).

[0104] In embodiments, diagnostic models or signatures are prepared by machine learning analyses by dividing data into groups based on feature type (e.g., functional annotation, taxonomic composition, KEGG pathways, etc. as already described), and models can be selected based on performance metrics such as accuracy, AUC, logloss, and other relevant criteria as known in the art. In embodiments, backward feature elimination is performed to refine the models and select those with the fewest features that maintain strong performance metrics. In embodiments, model construction comprises feature merging (combine features from top-performing models across different analysis groups to create a final, comprehensive model), and optionally further applying backward feature elimination. In embodiments, the model construction further comprises Permutation Testing, which involves conducting permutation tests on a holdout set (HO) to validate the model's generalization capability and ensure that the performance holds up across different data permutations. Multifold External Holdout Testing involves conducting permutation tests on external holdout sets (XHO) that are reserved and are not used for model training purposes.

[0105] In embodiments, the genetic elements associated with (indicative of) of the disorder that are selected for inclusion in the signature or model, include features of microbial

[0106] DBl / 162792885.1 28 MBI-004PC / 108458-5004 taxonomy (e.g., sequences annotated for taxonomy, for example at genus, species, and / or strain level), along with at least 2, at least 3, at least 4, or at least 5 selected from: (1) features defining an enzyme or reaction classification (e.g., sequences annotated by EC number or enzymatic reaction), (2) features defining gene orthology or ontology (e.g., sequences annotated with gene orthology or gene ontology database), (3) features defining gene homology (e.g., sequences annotated with UniRef90 database), (4) features defining pathways (e.g., sequences annotated with KEGG pathway database), and (5) features defining protein families and / or domains (e.g., sequences annotated with a protein family or protein domain database).

[0107] In this aspect, the gene signature for CRC can comprise microbial taxonomic classification features and microbial gene function features as described above and as exemplified in Table 8 to 17. The gene signature for CRA and CRAA can comprise microbial taxonomic classification features and microbial gene function features as described above and as exemplified in Tables 18 and 19, respectively. In this aspect, the method can employ any sample suitable for evaluating the microbiome of the cohort, including fecal samples as well as other biological samples, including human biofluids (e.g., blood, serum, plasma, urine, saliva), tissues, mucosa, and cell samples.

[0108] As described, the nucleic acid sequencing for model construction may comprise one or more of shotgun metagenomic sequencing, rDNA sequencing, and targeted nucleic acid sequencing, including targeted amplicon sequencing and hybridization capture probe sequencing. Any sequencing technique can be employed. The genetic elements in the samples are assigned to a reference genome for taxonomic classification (which can include rDNA analysis), and / or genetic elements are assigned to a gene function (as already described). Taxonomic and gene function features can also be analyzed at the protein level using known methods.

[0109] For example, microbial taxonomic classification features and microbial gene function features are selected that have a differential abundance or differential prevalence in fecal or other biological samples from subjects with the disorder of interest, as compared to control subjects. In some embodiments, the features comprise at least five taxonomic and / or

[0110] DBl / 162792885.1 29 MBI-004PC / 108458-5004 gene function features. The features may comprise at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or gene function features. In some embodiments, the features comprise at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 taxonomic features; and at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 gene function features.

[0111] In some embodiments, microbial taxonomic classification features and microbial gene function features are selected that have a differential abundance or differential prevalence in fecal or other biological samples from CRC subjects, as compared to control subjects. In some embodiments, the features comprise at least five taxonomic and / or gene function features, and which are optionally listed in Tables 8 to 17. The features may comprise at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or gene function features, which are optionally listed in Tables 8 to 17. In some embodiments, the features comprise at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 taxonomic features that are optionally listed in Tables 8 to 17; and at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 gene function features that are optionally listed in Tables 8 to 17.

[0112] In some embodiments, microbial taxonomic classification features and microbial gene function features are selected that have a differential abundance or differential prevalence in fecal or other biological samples from CRA subjects, as compared to control subjects. In some embodiments, the features comprise at least five taxonomic and / or gene function features, and which are optionally listed in Table 18. The features may comprise at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or gene function features, which are optionally listed in Table 18. In some embodiments, the features comprise at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 taxonomic features that are optionally listed in Table 18; and at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 gene function features that are optionally listed in Table 18.

[0113] In some embodiments, microbial taxonomic classification features and microbial gene function features are selected that have a differential abundance or differential

[0114] DBl / 162792885.1 30 MBI-004PC / 108458-5004 prevalence in fecal or other biological samples from CRAA subjects, as compared to control subjects. In some embodiments, the features comprise at least five taxonomic and / or gene function features, and which are optionally listed in Table 19. The features may comprise at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or gene function features, which are optionally listed in Table 19. In some embodiments, the features comprise at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 taxonomic features that are optionally listed in Table 19; and at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 gene function features that are optionally listed in Table 19.

[0115] In various embodiments regarding models for detection of colorectal neoplasia, at least three gene signatures are trained that: classify samples for the presence or absence of CRA, classify samples for the presence or absence of CRAA, and classify samples for the presence or absence of CRC. For example, the signature classifying samples for the presence or absence of CRA is trained with fecal or other biological samples from a CRA cohort and samples from a control cohort by machine learning. The signature classifying samples for the presence or absence of CRC is trained with fecal or other biological samples from a CRC cohort and samples from a control cohort by machine learning. The signature classifying samples for the presence or absence of CRAA is trained with fecal or other biological samples from a CRAA cohort and samples from a control cohort by machine learning. The machine learning can be as already described, and can include supervised machine learning, unsupervised machine learning, or semi-supervised machine learning, or a combination thereof.

[0116] In embodiments, feature importance tables (e.g., for CRA, CRAA, or CRC, or other disorder) are produced through the final stages of a discovery pipeline, e.g., after data acquisition, taxonomic and functional annotation, and / or quality analysis and batch correction, as illustrated in FIG. 1. The methods and systems provide for feature reduction protocols, which can employ a “Microbiome Explorer Workbench ML Modeling Analysis and Optimization” module as shown. In embodiments, a first feature reduction step is performed using logistic regression, wherein features are evaluated according to p-values and strength of association. The system can be further configured to visualize the full set of

[0117] DBl / 162792885.1 31 MBI-004PC / 108458-5004 results within the Workbench. From this interface, a user may invoke automated machine learning (AutoML) operations to iteratively reduce the feature space, producing the final feature importance tables.

[0118] In embodiments, large language models (LLM) integration can be used for feature interpretation and / or reduction. In embodiments, the method or system integrates a LLM trained on a corpus comprising more than 3,000 published peer-reviewed papers relating to the disorder of interest, such as colon cancer progression and biology. The LLM is accessible via a chatbot interface and is configured to provide contextual explanations of the features identified during the pipeline analysis. The chatbot is further configured to deliver summaries of relevant published findings, thereby enabling the user to interpret the biological and clinical significance of specific features in light of peer-reviewed literature.

[0119] In embodiments, the methods and systems combine the pipeline-based feature reduction and the LLM interpretive component into a unified framework. The methods and systems are configured to generate feature importance tables through logistic regression and AutoML operations as described above, and further to provide real-time interpretive outputs regarding those features by means of the LLM chatbot. This integrated approach enables users not only to identify and prioritize relevant features, but also to immediately access context from the scientific literature regarding those features, thereby accelerating discovery and decision-making.

[0120] In various embodiments, the signature(s) created have a sensitivity for classifying samples for a disorder (such as for the presence or absence of CRA, CRAA, or CRC) of at least about 0.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about 0.95. In various embodiments, the signature(s) created have a specificity for classifying samples for the presence or absence of the disorder (such as CRA, CRAA, or CRC) of at least about 0.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about 0.95. For example, the signatures classify each of CRA, CRAA, and CRC with a sensitivity of at least 0.75. For example, the signatures created may have an area under the curve (AUC) for classifying samples for the presence or absence of CRA, CRAA,

[0121] DBl / 162792885.1 32 MBI-004PC / 108458-5004 or CRC of at least about 0.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about 0.95.

[0122] As used herein, the term “about”, unless the context requires otherwise, means ±10% of an associated numerical value.

[0123] Other aspects and embodiments of the invention will be apparent from the following non-limiting examples.

[0124] EXAMPLES

[0125] Example 1: Microbiome Analysis Pipeline

[0126] FIG. 1 illustrates diagrammatically a microbiome analysis pipeline for determining informative features and preparing machine learning models according to the present disclosure.

[0127] First, data collection is conducted, which can involve clinical protocol optimization and generative Al. For example, generative Al may be applied to refine clinical protocols, optimizing for better alignment with study goals, regulatory requirements, and patient outcomes. The Al can suggest protocol adjustments based on historical data and predictive modeling. Further, statistical and machine learning analysis plans can be defined to implement clinical protocols for Metabiome discovery and predictive modeling. Objectives are defined including working hypotheses and systems biology framework, which may be based on Al-enabled reviews of applicable life science and medical literatures.

[0128] To optimize data collection, generative Al can be applied to simulate and optimize data collection protocols, ensuring that the most relevant and high-quality data is gathered. This step includes identifying optimal sampling strategies, sequencing methods, and conditions to enhance data quality and consistency across studies. Optimized data collection protocols are applied to collect datasets required and specified for the Metabiome discovery and predictive modeling pipeline.

[0129] For data acquisition and initial processing, NGS shotgun sequencing dataset(s) are acquired for the Metabiome system of interest. The input is raw FASTQ sequencing files, which are subjected to quality control and human read removal. An exemplary software for

[0130] DBl / 162792885.1 33 MBI-004PC / 108458-5004 this purpose is KneadData, which can be employed for initial quality filtering and removal of human reads. This step ensures that the dataset is clean and contains only the relevant microbial sequences for downstream analysis.

[0131] The shotgun sequencing data after QC, is used for taxonomic and functional annotation, which can employ various available databases, and specifically to identify metagenomic biomarkers followed by inferential prediction of molecular interactions and functions from reference databases, and / or to identify and assemble genes present in the FASTQ dataset followed by functional annotation of genes from reference databases. Commercial and open-source software code and programs are often available in R or Python code libraries that may be useful for intermediate steps in the Metabiome pipeline. For instance, programs developed by the Harvard Huttenhower Microbiome Laboratory include:

[0132] MetaPhlAn: Used for taxonomic profiling, identifying the composition of microbial communities in the samples;

[0133] HUMAnN3: Used for functional annotation, providing insights into the metabolic pathways and gene functions represented in the microbiome;

[0134] StrainPhlAn: Strain-level resolution on single nucleotide polymorphisms within conserved and unique species marker genes,

[0135] MelonnPan: a computational method for predicting metabolite composition from microbiome sequencing data;

[0136] PPANINI: A computational pipeline to prioritize microbial genes based on their metagenomic properties (e.g. prevalence and abundance);

[0137] ShortBRED: A system for profiling protein families of interest at veiy high specificity in shotgun sequencing data;

[0138] MetaWIBELE: Identify and prioritize potentially bioactive (and often uncharacterized) gene products in microbial communities;

[0139] FUGAsseM Predict function of uncharacterized gene products in microbial communities;

[0140] DBl / 162792885.1 34 MBI-004PC / 108458-5004

[0141] MACARRoN: Identify potentially bioactive novel metabolites in microbial communities.

[0142] The taxonomic and functional annotated data are batch corrected. Batch effect correction is a statistical approach used to remove systematic variations in data that arise from differences in experimental conditions, sample processing, or equipment across different batches of data. This correction ensures that the observed biological differences in data are truly reflective of the underlying biological variation rather than artifacts introduced by technical discrepancies. Batch effects in microbiome data arise from sampling and processing differences in studies especially when data is collecting from multiple different metagenomic studies from around the world or from different collection sites in a single country. Examples of open-source batch correction programs include:

[0143] MMUPHin: a Bioconductor package implementing meta-analysis methods for microbial community profiles. It has interfaces for: a) covariate-controlled batch and study effect adjustment, b) meta-analytic differential abundance testing, and meta-analytic discovery of c) discrete (cluster-based) or d) continuous unsupervised population structure. MMUPHin enables the normalization and combination of multiple microbial community studies, and can help in identifying microbes, genes, or pathways that are differential with respect to combined phenotypes. Finally, the package can find clusters or gradients of sample types that reproduce consistently among studies.

[0144] DeBias-M is a method for inference and correction of processing bias in microbiome data within a phenotype prediction framework, designed to operate in the context of multiple processing batches or studies (collectively termed “batches”). DEBIAS-M learns biascorrection factors for each microbe in each batch that simultaneously minimize batch effects and maximize cross-study associations with phenotypes. DEBIAS-M can improve modeling of microbiome data and identification of interpretable signals that are reproducible across studies.

[0145] ConQur is designed to remove microbiome batch effects using a two-part quantile regression model to minimize any study-specific biases that may have been introduced due to differences in sample preparation, sequencing platforms, or other study-related variables. ConQuR accommodates the complex distributions of microbial read counts by non-

[0146] DBl / 162792885.1 35 MBI-004PC / 108458-5004 parametric modeling, and it generates batch removed zero-inflated read counts that can be used in subsequent analyses.

[0147] Tools for evaluating meta-analysis datasets before and after batch correction include:

[0148] Principal Component Analysis (PC A): Conduct PC A before and after batch correction to visually assess whether the study bias has been effectively removed;

[0149] ADONIS: Permutational Multivariate Analysis of Variance to statistically test whether batch effects have been reduced;

[0150] Isolation Forest Algorithms: Identify outliers and ensure that the study bias has been minimized.

[0151] Tools for evaluating each study for fitness include: LOSO (Leave-l-Out Modeling), Anomaly Analysis (box plots of study vs Log2 cross validation), comparing Sensitivity, Specificity, Wrongly Prediction Score of each study; allowance for study, country, and subject case diversity in meta-analysis.

[0152] Using these packages and tools, an optimum meta-analysis dataset is selected for feature selection and logistic regression. Logistic regression is performed across all features (taxonomic, functional, KEGG pathways, etc.) while controlling for study ID as a covariate. This helps in identifying features significantly associated with the condition or outcome of interest.

[0153] Machine Learning (ML) analysis and construction of diagnostic algorithms and models can employ the following tools or methods. Holistic ML approach can be employed with AutoML including: (1) Conduct machine learning analyses by dividing data into groups based on feature type (e.g., functional annotation, taxonomic composition, KEGG pathways); Choose models based on performance metrics such as accuracy, AUC, logloss, and other relevant criteria; Perform backward feature elimination to refine the models and select those with the fewest features that maintain strong performance metrics.

[0154] Final model construction can involve: Feature Merging (combine features from all top-performing models across different analysis groups to create a final, comprehensive model), Final Model Optimization (applying backward feature elimination to ensure that it

[0155] DBl / 162792885.1 36 MBI-004PC / 108458-5004 is both parsimonious and robust, with a focus on maintaining performance while minimizing feature count), and Generalization Testing. For example, Permutation Testing can involve conducting permutation tests on the holdout set (HO) to validate the model's generalization capability and ensure that the performance holds up across different data permutations. Multifold External Holdout Testing involves conducting permutation tests on an external holdout set (XHO) that are reserved and are not used for model training purposes.

[0156] Methods can further involve interactive data visualization and exploration. Specifically, an interactive dashboard (MICROBIOME EXPLORER) can be implemented, allowing users to engage with data tables from different analysis groups (e.g., taxonomy, KEGG, functional annotations). The user can dynamically plot data and apply filters to explore specific subsets of features. This allows for more in-depth exploration and better understanding of the results. The user can interactively select features within the dashboard to send to AutoML for further analysis. The selected features will be processed in the backend, and users will receive visualizations of the final model results directly within the dashboard. The dashboard provides real-time updates and visual feedback, enabling users to refine selections and analysis iteratively.

[0157] Users may further conduct generative Al exploration and interpretation of biomarker annotations and models. This can include chatbot interaction, for example, to integrate a chatbot specifically trained on microbiome literature, allowing users to interactively explore the final features derived from machine learning models. Users can query the chatbot to understand the biological relevance of specific features, receive literature references, and explore potential pathways, gene functions, and microbial interactions related to the selected features. The chatbot can offer contextual assistance by linking features to known clinical outcomes, potential therapeutic targets, and known associations in microbiome literature, providing users with a deeper understanding of the data. Users can customize their queries to focus on specific aspects, such as functional relevance, taxonomic insights, or associations with particular diseases, enabling targeted exploration of the data. The chatbot can further provide concise summaries of relevant scientific literature, highlighting key findings and suggesting further reading, thus enriching the data exploration process with valuable references. As users interact with the chatbot, it can learn from the interactions, refining its

[0158] DBl / 162792885.1 37 MBI-004PC / 108458-5004 responses and recommendations, making the tool progressively more tailored to the user’s needs and preferences.

[0159] Example 2: CRC-CTR Taxonomical Analysis From NGS Data Acquisition to Metagenome Explorer Database Assembly Table 1 summarizes patient demographics and study inclusion. Samples are stool samples collected prior to colonoscopy.

[0160] DBl / 162792885.1 38 MBI-004PC / 108458-5004

[0161] After data acquisition and initial processing with KneadData taxonomic annotation is conducted as described in Example 1. In particular, taxonomic annotation with MetaPhlAn resulted in 10,054 taxonomic features, which were further selected for downstream analysis.

[0162] DBl / 162792885.1 39 MBI-004PC / 108458-5004

[0163] The data is evaluated for batch correction. This analysis assesses the impact of batch effects on the microbiome taxa dataset, using dimensionality reduction and statistical methods to evaluate the structure and variability across different groups. To evaluate the presence of batch effects in the raw dataset, two distinct dimensionality reduction techniques, Principal Component Analysis (PCA) (linear) and Uniform Manifold Approximation and Projection (UMAP) (non-linear), were applied. These methods allow exploration of whether the dataset exhibits clustering or separation due to non-biological factors, such as batch processing, across three key categorical variables: target, Country, and ProjectNam eDate. ADONIS (PERMANOVA) was also employed to statistically quantify the impact of these variables on overall dataset variation.

[0164] PCA is a linear technique that highlights the main directions of variance in the data, enabling the visualization of batch effects that may dominate the variance structure. The first two principal components (PCI and PC2) were used to explore whether the dataset clusters were based on batch or experimental conditions. Target-Based PCA was used to visualize the separation of samples based on the target variable, assessing whether target groups (e g., CRC vs Control) explain the variance in the data. Country-Based PCA was used to demonstrate the extent to which geographical origins of samples influence the dataset. ProjectNameDate-Based PCA was used to investigate whether the project-specific timeframes affect the data clustering.

[0165] UMAP is a non-linear technique that captures both global and local structures in the data, to uncover more subtle patterns or complex batch effects that PCA might miss. Target- Based UMAP examines if the dataset clusters based on biological target groupings (e g., disease vs. control). Country-Based UMAP investigates the influence of geographical origin on the dataset structure. ProjectNameDate-Based UMAP evaluates how project timing and batches may contribute to sample separation.

[0166] To statistically evaluate the influence of batch effects, we conducted ADONIS (also known as PERMANOVA) tests on the Bray-Curtis dissimilarity matrix. This non-parametric test quantifies the amount of variance explained by target, Country, and ProjectNameDate. Results are summarized in Table 2:

[0167] DBl / 162792885.1 40 MBI-004PC / 108458-5004

[0168] Target: R2of 0.0042 indicates that only 0.42% of the total variance in the data is explained by the target variable (e.g., CRC vs Control). This is a very small effect size.

[0169] Country: R2of 0.1173 indicates that 11.73% of the total variance in the data is explained by the Country variable. This is a moderate effect size, suggesting that geographical location significantly impacts the microbiome composition.

[0170] ProjectNameDate: R2: 0.1522 indicates that 15.22% of the total variance is explained by the ProjectNameDate variable. This is the largest effect size among the three variables, suggesting that batch effects related to the timing or project-specific factors are strongly influencing the dataset.

[0171] The P-value of 0.001 for all variables shows that these effects are statistically significant, meaning there is some detectable difference in the microbiome composition between the groups defined by these variables. The R2values represent the effect sizes and indicate how much of the overall variance is explained by each variable.

[0172] To further explore potential batch effects and outliers, the raw dataset was processed through an AutoML platform, focusing on anomaly detection models. AutoML (Automated Machine Learning) automates the machine learning process, from data preprocessing to model deployment. The goal was to identify outliers in the microbiome taxa data, which could indicate samples that deviate significantly due to batch effects or other non-biological factors. Four anomaly detection models were outputted by the AutoML platform, and their performance was evaluated using synthetic AUCs across validation, cross-validation, and holdout datasets, as shown in Table 3:

[0173] DBl / 162792885.1 41 MBI-004PC / 108458-5004

[0174] Isolation Forest Anomaly Detection with Calibration emerged as the best-performing model with the highest AUC across all datasets (Validation AUC: 0.8203, Holdout AUC: 0.8309). This model provided the most reliable outlier detection and demonstrated consistent performance in identifying potential anomalies in the microbiome data. Local Outlier Factor (LOF) also performed well but had slightly lower AUC scores (Validation AUC: 0.7991, Holdout AUC: 0.7839) compared to Isolation Forest. Double Median Absolute Deviation showed moderate performance. Mahalanobis Distance model was not suitable for this task, with AUC scores of 0.5 (indicating random performance). Accordingly, the Isolation Forest model was selected as the most appropriate for detecting anomalies in the microbiome dataset, based on its strong and consistent AUC performance. This model can therefore be used to identify samples that may be outliers due to batch effects or other non-biological factors, helping refine the dataset for further analysis.

[0175] To make the anomaly scores more interpretable, a percentile rank transformation was applied. This transformation helps normalize the distribution of anomaly scores, particularly since the raw scores were concentrated between 0 and 1. Percentile ranks allow comparison of the relative standing of each sample within the dataset, where a percentile rank of 0 represents the lowest anomaly score and 1 represents the highest. This transformed anomaly score helps in visualizing the distribution of outliers across the dataset based on various categorical variables such as target, Country, and ProjectNam eDate.

[0176] To better understand how the anomaly scores were distributed across different categories, box plots were generated. These plots visualize the percentile-ranked anomaly scores for: (1) Target (CRC vs Control), showing whether the samples labeled as CRC or Control exhibit different outlier tendencies; (2) Country, showing how geographical location may influence the distribution of anomaly scores; and (3) ProjectNam eDate, assessing whether different projects or time frames contribute to outlier detection, possibly indicating batch effects. The analysis is summarized in Table 4:

[0177] DBl / 162792885.1 42 MBI-004PC / 108458-5004

[0178] The ANOVA test was run to relate the impact of Target, Country, and Project on the anomaly scores in the microbiome dataset. The results reveal that all three variables — Target, Country, and Project — have statistically significant effects on the anomaly scores. However, the magnitude of their contributions varies.

[0179] To identify potential outliers in the anomaly scores, the Interquartile Range (IQR) method was applied, a robust statistical technique that defines outliers as points that fall significantly outside the central 50% of the data. This method is particularly effective for skewed or non-normally distributed data, such as microbiome taxonomic data.

[0180] The first quartile (QI, 25th percentile) and third quartile (Q3, 75th percentile) of the anomaly scores were calculated: QI = 0.006; Q3 = 0.015.

[0181] The Interquartile Range (IQR), which measures the spread of the middle 50% of the data, was computed as: IQR = Q3 - QI = 0.015 - 0.006 = 0.009.

[0182] Outliers were defined as points that fall outside 1.5 times the IQR below QI or above Q3: Lower bound = QI - 1.5 * IQR = 0.006 - 1.5 * 0.009 = -0.0075 (Since anomaly scores cannot be negative, this lower bound is not relevant for positive scores); Upper bound = Q3 + 1.5 * IQR = 0.015 + 1.5 * 0.009 = 0.0285. Any anomaly scores greater than 0.0285 were flagged as potential outliers. Based on the IQR method, 211 outliers were identified from 1,726 samples. These samples had anomaly scores exceeding the upper threshold of 0.0285. The remaining 1,515 samples were classified as non-outliers.

[0183] Box plots were generated to visualize the distribution of the anomaly scores, with the outliers clearly indicated as points falling outside the whiskers of the plot (extending to QI - 1.5 * IQR and Q3 + 1.5 * IQR).

[0184] DBl / 162792885.1 43 MBI-004PC / 108458-5004

[0185] Outliers only by Target: CONTROL (95); CRC (116).

[0186] Outliers only by Country (Table 5):

[0187] Outliers only by Project (Table 6):

[0188] To assess the potential impact of outliers removal on batch effects, we performed an analysis both before and after the removal of outliers. This was done using an ADONIS (PERMANOVA) test, by the variables Target, Country, and ProjectNameDate. The test was applied to both the full dataset and the dataset after removing outliers. Below, we compare the results of the ADONIS test for the two datasets (Table 7):

[0189] The proportion of variance explained by Target (R2) remained relatively unchanged before and after outlier removal, suggesting that the outliers had minimal impact on how well Target explains the variation in the dataset. In both cases, the R2value was low (0.0042

[0190] DBl / 162792885.1 44 MBI-004PC / 108458-5004 before and 0.0041 after), indicating that Target contributes only a small amount to the overall variation in the data. However, the p-value remained significant (0.001), implying that the effect of Target, although small, is statistically significant.

[0191] The R2value for Country slightly decreased after outlier removal, from 0.1173 to 0.1135, but the reduction was relatively small. This suggests that Country still explains a substantial portion of the variance in the dataset, even after removing the outliers.

[0192] The most noticeable change occurred with the ProjectNam eDate variable. Before outlier removal, Proj ectNameDate explained 15.22% of the variance (R2= 0.1522), but after outlier removal, this dropped significantly to 3.699% (R2= 0.03699).

[0193] The analysis indicates that the outliers had a significant impact on the variation explained by Proj ectNameDate, suggesting that batch effects related to project-specific factors were largely driven by these outliers. After outlier removal, the influence of Proj ectNameDate diminished, while the effects of Country and Target remained stable. This highlights the importance of addressing outliers in batch effect analysis to ensure more accurate interpretations of the data.

[0194] After removing outliers from the dataset, two complementary statistical methods were applied — Logistic Regression and Kruskal -Wallis tests by country — to further reduce the feature set. This two-pronged approach identifies predictive and robust features, accounting for both statistical significance across countries and predictive power while controlling for confounding variables like Country and Project ID.

[0195] For logistic regression, the target variable was binary (e.g., CRC vs. Control), with CRC samples coded as 1 and control samples as 0. Both Country and Project ID were included as covariates to control confounding factors and ensure the predictive power of the features was not influenced by these variables. Each feature (e.g., taxonomic) was tested in a logistic regression model of the form: logit (T’(Target)) = Po + Pi • Feature + P2 ‘ Country + P3 • Project ID

[0196] DBl / 162792885.1 45 MBI-004PC / 108458-5004 where: P(Target) is the probability of the target outcome (CRC), Feature represents each individual feature, and Country and Project ID were included as control variables.

[0197] Each feature was evaluated based on:

[0198] P-value: A measure of statistical significance. Features with a low p-value (< 0.05) were considered predictive;

[0199] Coefficient (Strength): The magnitude of the logistic regression coefficient, indicating the strength of association between the feature and the target (CRC).

[0200] Directionality: Positive or negative, based on the sign of the coefficient.

[0201] In parallel with the logistic regression, the Kruskal-Wallis test was applied across individual countries to identify features that showed significant differences between the CRC and control groups within each country. This helped in evaluating the country-specific performance of features and detecting those that were consistently significant across multiple geographic regions. For each feature, the Kruskal-Wallis test was conducted within each country, and the p-value (FRD corrected) was computed to determine the statistical significance of the feature. A median FDR-corrected p-value was calculated across all countries, providing a summary measure of the feature's significance across geographic regions.

[0202] The combination of both Logistic Regression and Kruskal-Wallis tests allowed for a robust feature selection process. Features that were statistically significant (p-value < 0.05) in both tests — Logistic Regression and Kruskal-Wallis (median p-value across countries) — were prioritized. These features are not only predictive but also consistent across countries, making them more reliable for downstream machine learning models. In cases where a broader set of features was needed, we also considered features that were significant in either the logistic regression or the Kruskal -Wallis test, as some features might perform better in specific countries while others might be more predictive across all samples.

[0203] DBl / 162792885.1 46 MBI-004PC / 108458-5004

[0204] Features with significant p-values across multiple countries were selected based on their median FDR-corrected p-values (Kruskal -Wallis). These features were robust across different geographic regions. Further, features with strong coefficients and low p-values were selected (logistic regression), indicating that they are predictive of the target outcome (CRC vs control) even after accounting for the potential effects of Country and Project ID.

[0205] After applying both tests and filtering based on significance, a reduced set of N predictive features was identified. These features were selected for their: (1) Predictive power (as indicated by logistic regression), (2) Consistency across countries (as indicated by the Kruskal-Wallis test), and (3) Robustness to confounding effects (as both tests controlled for relevant covariates such as Country and Project ID). This dual approach of using both logistic regression and country-specific Kruskal-Wallis tests ensured that the final feature set is both statistically significant and reliable across different geographic regions, making it highly suited for downstream machine learning applications.

[0206] Example 3: Logistic Regression Feature Selection Process for CRC, CRA, CRAA, and Controls

[0207] Following the identification of features with significant p-values across multiple countries using Kruskal-Wallis tests, logistic regression analysis is applied to further refine the selection for different cohorts, including CRC, CRA, CRAA, and Controls. This process evaluates features based on three key metrics — Strength, Directionality, and Strength- Power — to assess their predictive power and robustness across these groups.

[0208] Strength is calculated directly from the logistic regression output, where the coefficient ( ) of each feature quantifies its influence on the target variable (e.g., distinguishing between CRC, CRA, CRAA, and Controls). The coefficient measures how much the log-odds of the target (e.g., CRC vs. Control) change with a one-unit increase in the feature. A larger absolute value of the coefficient indicates a stronger influence, whether positive or negative, on the target outcome across these cohorts.

[0209] Directionality is determined by the sign of the coefficient, indicating whether the feature's influence is positive or negative. A positive coefficient implies that an increase in

[0210] DBl / 162792885.1 47 MBI-004PC / 108458-5004 the feature increases the likelihood of the target being CRC, CRA, or CRAA, relative to Controls. The directionality for these features is classified as "Positive". A negative coefficient means that as the feature increases, the likelihood of the target being CRC, CRA, or CRAA decreases, making it more likely that the sample belongs to the Control group. These features are classified as having "Negative" directionality.

[0211] Strength-Power categorizes the Strength (coefficient) into quartiles, providing a clearer interpretation of the effect size for each feature. After calculating the coefficients, they are divided into quartiles based on their magnitude, with features being assigned to one of the following categories: (1) Very Strong Negative are features with a very strong negative influence on the likelihood of the target being CRC, CRA, or CRAA, relative to Controls, typically falling in the lowest quartile; (2) Weak Positive are features with a weak positive influence, typically found in the second quartile; (3) Moderate Positive are features with a moderate positive influence, categorized in the third quartile; and (4) Strong Positive are features with a strong positive influence on the likelihood of the target being CRC, CRA, or CRAA, generally in the highest quartile. This categorization allows for a more practical interpretation of logistic regression outputs across the different groups, helping to identify the most influential features in predicting CRC, CRA, and CRAA compared to Controls. By selecting features with strong coefficients, consistent directionality, and reliable Strength- Power classifications, the final set of predictive features is both statistically significant and robust across geographic regions and cohorts, making it highly suitable for downstream machine learning applications.

[0212] Example 3: CRC-CTR Taxonomic and Functional Gene Analysis From NGS Data

[0213] Using the workflow generally described in Example 2, exemplary features that are indicative of CRC were determined, and which are provided in the following tables.

[0214] Table 8 shows annotation for Enzyme Commission Number. EC No. classifies enzymes based on the chemical reactions they catalyze. Each number represents a specific enzyme and its function. Annotating the microbiome genetic elements for EC No. produces statistically significant features for discriminating CRC from controls. Further, the

[0215] DBl / 162792885.1 48 MBI-004PC / 108458-5004 enzymatic activities within the microbiome shed light on metabolic pathways and biochemical transformations facilitated by microbial enzymes.

[0216] Table 9 shows annotation using the eggNOG database, which is a database that groups proteins into orthologous groups across species, represents evolutionary relationships and functional similarities between proteins. Annotating the microbiome genetic elements for gene orthology produces statistically significant features for discriminating CRC from controls. Further, the information helps to understand shared functions between microbial species based on their protein content.

[0217] Table 10 shows annotation using UniRef90 database, which represents gene functions based on UniRef90 cluster. Annotating the microbiome genetic elements using UniRef90 produces statistically significant features for discriminating CRC from controls. Further, this information highlights the functional potential of the microbiome by showing the abundance of gene clusters involved in various biological activities.

[0218] Table 11 shows annotation using Gene Ontology (GO) database, which defines gene functions across three domains: Biological Process, Molecular Function, and Cellular Component. Annotating the microbiome genetic elements using GO database produces statistically significant features for discriminating CRC from controls. Further, this information helps to inform gene function at the molecular level (e.g., binding, catalysis), their role in biological processes (e.g., metabolism), and where the proteins operate in the cell.

[0219] Table 12 shows annotation using Kegg Orthology (KO) database, which groups genes into functional units, providing insight into their roles in metabolic pathways. Annotating the microbiome genetic elements using KO database produces statistically significant features for discriminating CRC from controls. This information links genes to their biological roles in various processes, including metabolism and regulation, and helps to map genes to specific pathways.

[0220] Table 13 shows annotation by Pathway (e.g., using KEGG pathway database), which provides insights into the abundance of metabolic pathways within the microbiome.

[0221] DBl / 162792885.1 49 MBI-004PC / 108458-5004

[0222] Annotating the microbiome genetic elements by pathway produces statistically significant features for discriminating CRC from controls. Further, this information helps to understand the functional potential of microbial communities by showing which pathways are prevalent and how they contribute to overall metabolism.

[0223] Table 14 shows annotation by Protein families, using the Pfam database, which categorizes proteins into families based on sequence similarity. Annotating the microbiome genetic elements by protein families produces statistically significant features for discriminating CRC from controls. This information further provides functional domains of proteins, helping to understand protein functions within the microbiome.

[0224] Table 15 shows annotation by Reactions (RXN), which represent sets of genes or proteins that catalyze specific biochemical reactions in metabolic pathways. Annotating the microbiome genetic elements by reaction produces statistically significant features for discriminating CRC from controls. This information further highlights how microbial communities contribute to metabolic processes, linking genes and enzymes to their specific roles in biochemical reactions.

[0225] Table 16 shows annotation by taxonomic classifications of microbiome samples (e.g., Kingdom, Phylum, Genus), which produces statistically significant features for discriminating CRC from controls. Further, this information allows one to explore the composition and diversity of microbial communities at different taxonomic levels, which is crucial for understanding the structure and ecological roles of microbes in various environments or conditions.

[0226] Once the individual tables are processed in the pipeline, features are first reduced through Kruskal-Wallis tests and logistic regression. Subsequently, further reduction occurs via the AutoML process. For each table, the top-performing model identified in AutoML is used to extract the most impactful features. These features are then merged to create an ensembled dataset, combining the key features from each top model. This final merged dataset undergoes another AutoML run to further reduce the feature set, ultimately

[0227] DBl / 162792885.1 50 MBI-004PC / 108458-5004 identifying the minimal number of mixed features (from different tables) that provide the best overall performance).

[0228] Table 17: Machine learning ensemble features for discriminating CRC

[0229] DBl / 162792885.1 51 MBI-004PC / 108458-5004

[0230] DB1 / 162792885.1 52 MBI-004PC / 108458-5004

[0231] DB1 / 162792885.1 53 MBI-004PC / 108458-5004

[0232] Table 17 presents ensemble machine learning features for discriminating colorectal cancer from control samples. These features were selected based on their relative importance in the predictive models developed using supervised machine learning algorithms, such as random forests and gradient boosting machines. The selection process involved ranking features by their contribution to model accuracy and refining the list through cross-validation to prevent overfitting.

[0233] Validation testing results of the AutoML modeling of the ensembled dataset (Table 17) show cross validation (CV) AUC of 0.87 and log loss of 0.46. Results show HO (holdout) AUC of 0.84 and HO log loss of 0.48. Using the process described above, Machine learning ensemble features for discriminating CRA (Table 18) and CRAA (Table 19) were also determined using CRA and CRAA cohorts.

[0234] Table 18: Machine learning ensemble features for discriminating CRA

[0235] DBl / 162792885.1 54 MBI-004PC / 108458-5004

[0236] DB1 / 162792885.1 55 MBI-004PC / 108458-5004

[0237] DB1 / 162792885.1 56 MBI-004PC / 108458-5004

[0238] DB1 / 162792885.1 57 MBI-004PC / 108458-5004

[0239] DB1 / 162792885.1 58 MBI-004PC / 108458-5004

[0240] DB1 / 162792885.1 59 MBI-004PC / 108458-5004

[0241] DB1 / 162792885.1 60 MBI-004PC / 108458-5004

[0242] DB1 / 162792885.1 61 MBI-004PC / 108458-5004

[0243] DB1 / 162792885.1 62 MBI-004PC / 108458-5004

[0244] DB1 / 162792885.1 63 MBI-004PC / 108458-5004

[0245] DB1 / 162792885.1 64 MBI-004PC / 108458-5004

[0246] DB1 / 162792885.1 65 MBI-004PC / 108458-5004

[0247] Table 19: Machine learning ensemble features for discriminating CRAA

[0248] DBl / 162792885.1 66 MBI-004PC / 108458-5004

[0249] DB1 / 162792885.1 67 MBI-004PC / 108458-5004

[0250] DB1 / 162792885.1 68 MBI-004PC / 108458-5004

[0251] DB1 / 162792885.1 69 MBI-004PC / 108458-5004

[0252] DB1 / 162792885.1 70 MBI-004PC / 108458-5004

[0253] DB1 / 162792885.1 71 MBI-004PC / 108458-5004

[0254] DB1 / 162792885.1 72 MBI-004PC / 108458-5004

[0255] DB1 / 162792885.1 73 MBI-004PC / 108458-5004

[0256] DB1 / 162792885.1 74 MBI-004PC / 108458-5004

[0257] In summary, metagenomic data is annotated for microbial taxonomic and gene function features, which undergo feature reduction through a dual-stage methodology. Initially, features can be filtered using the Kruskal-Wallis test and logistic regression, focusing on statistical significance and predictive capacity. Subsequently, these features are further refined, which can employ automated machine learning techniques. The topperforming model from each individual dataset (e.g., table) is selected based on its ability to highlight the most impactful features. These feature sets are then merged into a unified ensembled dataset, comprising the optimal features from each top model across all tables.

[0258] DBl / 162792885.1 75 MBI-004PC / 108458-5004

[0259] This merged dataset undergoes an additional round of processing (e.g., within the AutoML framework) to further reduce the feature set, aiming to identify the minimal yet most effective combination of features drawn from multiple tables, thereby maximizing model performance and predictive accuracy.

[0260] To ensure robust model evaluation, 10-fold cross-validation is employed throughout the process. This method provides a comprehensive assessment of the model's generalization ability by training on 80% of the data while testing on 20%, iterating across multiple folds to mitigate overfitting and ensure a well-rounded model validation.

[0261] REFERENCES

[0262] [1] Fearon ER, Vogelstein B. A genetic model for colorectal tumorigenesis. Cell. 1990;61(5):759-767.

[0263] [2] Clevers H, Nusse R. Wnt / p-catenin signaling and disease. Cell. 2012;149(6): 1192-1205.

[0264] [3] Vogelstein B, et al. Genetic alterations during colorectal-tumor development. N Engl J Med. 1988;319(9):525-532.

[0265] [4] Baker SJ, et al. p53 gene mutations occur in combination with 17p allelic deletions as late events in colorectal tumorigenesis. Cancer Res. 1990; 50(23): 7717-7722.

[0266] [5] Lao VV, Grady WM. Epigenetics and colorectal cancer. Nat Rev Gastroenterol Hepatol. 2011;8(12):686-700.

[0267] [6] Fraga MF, et al. Loss of acetylation at Lysl6 and trimethylation at Lys20 of histone H4 is a common hallmark of human cancer. Nat Genet. 2005;37(4):391-400.

[0268] [7] Lengauer C, et al. Genetic instability in colorectal cancers. Nature. 1997;386(6625):623- 627.

[0269] [8] Boland CR, Goel A. Microsatellite instability in colorectal cancer. Gastroenterology. 2010;138(6):2073-2087.e3.

[0270] [9] Til H, et al. The intestinal microbiota in colorectal cancer. Cancer Cell. 2018;33(6):954- 964.

[0271] DBl / 162792885.1 76 MBI-004PC / 108458-5004

[0272]

[0010] Fridman WH, et al. The immune contexture in human tumours: impact on clinical outcome. Nat Rev Cancer. 2012;12(4):298-306.

[0273]

[0011] Vander Heiden MG, et al. Understanding the Warburg effect: the metabolic requirements of cell proliferation. Science. 2009;324(5930): 1029-1033.

[0274] DBl / 162792885.1 77 MBI-004PC / 108458-5004

[0275] TABLE 8: Enzyme Commission (EC)

[0276] DB1 / 151327673.3 93

[0277] MBI-004PC / 108458-5004

[0278] DB1 / 151327673.3 94

[0279] MBI-004PC / 108458-5004

[0280] DB1 / 151327673.3 95

[0281] MBI-004PC / 108458-5004

[0282] DB1 / 151327673.3 96

[0283] MBI-004PC / 108458-5004

[0284] DB1 / 151327673.3 97

[0285] MBI-004PC / 108458-5004

[0286] DB1 / 151327673.3 98

[0287] MBI-004PC / 108458-5004

[0288] Table 9: Eggnog

[0289] DB1 / 162798841.4 99

[0290] MBI-004PC / 108458-5004

[0291] DB1 / 162798841.4 100

[0292] MBI-004PC / 108458-5004

[0293] DB1 / 162798841.4 101

[0294] MBI-004PC / 108458-5004

[0295] DB1 / 162798841.4 102

[0296] MBI-004PC / 108458-5004

[0297] DB1 / 162798841.4 103

[0298] MBI-004PC / 108458-5004

[0299] DB1 / 162798841.4 104

[0300] MBI-004PC / 108458-5004

[0301] DB1 / 162798841.4 105

[0302] MBI-004PC / 108458-5004

[0303] DB1 / 162798841.4 106

[0304] MBI-004PC / 108458-5004

[0305] DB1 / 162798841.4 107

[0306] MBI-004PC / 108458-5004

[0307] DB1 / 162798841.4 108

[0308] MBI-004PC / 108458-5004

[0309] DB1 / 162798841.4 109

[0310] MBI-004PC / 108458-5004

[0311] DB1 / 162798841.4 110

[0312] MBI-004PC / 108458-5004

[0313] DB1 / 162798841.4 111

[0314] MBI-004PC / 108458-5004

[0315] DB1 / 162798841.4 112

[0316] MBI-004PC / 108458-5004

[0317] DB1 / 162798841.4 113

[0318] MBI-004PC / 108458-5004

[0319] DB1 / 162798841.4 114

[0320] MBI-004PC / 108458-5004

[0321] DB1 / 162798841.4 115

[0322] MBI-004PC / 108458-5004

[0323] DB1 / 162798841.4 116

[0324] MBI-004PC / 108458-5004

[0325] DB1 / 162798841.4 117

[0326] MBI-004PC / 108458-5004

[0327] DB1 / 162798841.4 118

[0328] MBI-004PC / 108458-5004

[0329] DB1 / 162798841.4 119

[0330] MBI-004PC / 108458-5004

[0331] DB1 / 162798841.4 120

[0332] MBI-004PC / 108458-5004

[0333] DB1 / 162798841.4 121

[0334] MBI-004PC / 108458-5004

[0335] DB1 / 162798841.4 122

[0336] MBI-004PC / 108458-5004

[0337] DB1 / 162798841.4 123

[0338] MBI-004PC / 108458-5004

[0339] DB1 / 162798841.4 124

[0340] MBI-004PC / 108458-5004

[0341] DB1 / 162798841.4 125

[0342] MBI-004PC / 108458-5004

[0343] DB1 / 162798841.4 126

[0344] MBI-004PC / 108458-5004

[0345] DB1 / 162798841.4 127

[0346] MBI-004PC / 108458-5004

[0347] DB1 / 162798841.4 128

[0348] MBI-004PC / 108458-5004

[0349] DB1 / 162798841.4 129

[0350] MBI-004PC / 108458-5004

[0351] DB1 / 162798841.4 130

[0352] MBI-004PC / 108458-5004

[0353] DB1 / 162798841.4 131

[0354] MBI-004PC / 108458-5004

[0355] DB1 / 162798841.4 132

[0356] MBI-004PC / 108458-5004

[0357] DB1 / 162798841.4 133

[0358] MBI-004PC / 108458-5004

[0359] DB1 / 162798841.4 134

[0360] MBI-004PC / 108458-5004

[0361] DB1 / 162798841.4 135

[0362] MBI-004PC / 108458-5004

[0363] DB1 / 162798841.4 136

[0364] MBI-004PC / 108458-5004

[0365] DB1 / 162798841.4 137

[0366] MBI-004PC / 108458-5004

[0367] DB1 / 162798841.4 138

[0368] MBI-004PC / 108458-5004

[0369] DB1 / 162798841.4 139

[0370] MBI-004PC / 108458-5004

[0371] DB1 / 162798841.4 140

[0372] MBI-004PC / 108458-5004

[0373] DB1 / 162798841.4 141

[0374] MBI-004PC / 108458-5004

[0375] DB1 / 162798841.4 142

[0376] MBI-004PC / 108458-5004

[0377] DB1 / 162798841.4 143

[0378] MBI-004PC / 108458-5004

[0379] DB1 / 162798841.4 144

[0380] MBI-004PC / 108458-5004

[0381] DB1 / 162798841.4 145

[0382] MBI-004PC / 108458-5004

[0383] DB1 / 162798841.4 146

[0384] MBI-004PC / 108458-5004

[0385] DB1 / 162798841.4 147

[0386] MBI-004PC / 108458-5004

[0387] DB1 / 162798841.4 148

[0388] MBI-004PC / 108458-5004

[0389] DB1 / 162798841.4 149

[0390] MBI-004PC / 108458-5004

[0391] DB1 / 162798841.4 150

[0392] MBI-004PC / 108458-5004

[0393] DB1 / 162798841.4 151

[0394] MBI-004PC / 108458-5004

[0395] DB1 / 162798841.4 152

[0396] MBI-004PC / 108458-5004

[0397] DB1 / 162798841.4 153

[0398] MBI-004PC / 108458-5004

[0399] DB1 / 162798841.4 154

[0400] MBI-004PC / 108458-5004

[0401] DB1 / 162798841.4 155

[0402] MBI-004PC / 108458-5004

[0403] DB1 / 162798841.4 156

[0404] MBI-004PC / 108458-5004

[0405] DB1 / 162798841.4 157

[0406] MBI-004PC / 108458-5004

[0407] DB1 / 162798841.4 158

[0408] MBI-004PC / 108458-5004

[0409] DB1 / 162798841.4 159

[0410] MBI-004PC / 108458-5004

[0411] DB1 / 162798841.4 160

[0412] MBI-004PC / 108458-5004

[0413] DB1 / 162798841.4 161

[0414] MBI-004PC / 108458-5004

[0415] DB1 / 162798841.4 162

[0416] MBI-004PC / 108458-5004

[0417] DB1 / 162798841.4 163

[0418] MBI-004PC / 108458-5004

[0419] DB1 / 162798841.4 164

[0420] MBI-004PC / 108458-5004

[0421] DB1 / 162798841.4 165

[0422] MBI-004PC / 108458-5004

[0423] DB1 / 162798841.4 166

[0424] MBI-004PC / 108458-5004

[0425] DB1 / 162798841.4 167

[0426] MBI-004PC / 108458-5004

[0427] DB1 / 162798841.4 168

[0428] MBI-004PC / 108458-5004

[0429] DB1 / 162798841.4 169

[0430] MBI-004PC / 108458-5004

[0431] DB1 / 162798841.4 170

[0432] MBI-004PC / 108458-5004

[0433] DB1 / 162798841.4 171

[0434] MBI-004PC / 108458-5004

[0435] DB1 / 162798841.4 172

[0436] MBI-004PC / 108458-5004

[0437] DB1 / 162798841.4 173

[0438] MBI-004PC / 108458-5004

[0439] DB1 / 162798841.4 174

[0440] MBI-004PC / 108458-5004

[0441] DB1 / 162798841.4 175

[0442] MBI-004PC / 108458-5004

[0443] DB1 / 162798841.4 176

[0444] MBI-004PC / 108458-5004

[0445] DB1 / 162798841.4 177

[0446] MBI-004PC / 108458-5004

[0447] DB1 / 162798841.4 178

[0448] MBI-004PC / 108458-5004

[0449] DB1 / 162798841.4 179

[0450] MBI-004PC / 108458-5004

[0451] DB1 / 162798841.4 180

[0452] MBI-004PC / 108458-5004

[0453] DB1 / 162798841.4 181

[0454] MBI-004PC / 108458-5004

[0455] DB1 / 162798841.4 182

[0456] MBI-004PC / 108458-5004

[0457] DB1 / 162798841.4 183

[0458] MBI-004PC / 108458-5004

[0459] DB1 / 162798841.4 184

[0460] MBI-004PC / 108458-5004

[0461] DB1 / 162798841.4 185

[0462] MBI-004PC / 108458-5004

[0463] DB1 / 162798841.4 186

[0464] MBI-004PC / 108458-5004

[0465] DB1 / 162798841.4 187

[0466] MBI-004PC / 108458-5004

[0467] DB1 / 162798841.4 188

[0468] MBI-004PC / 108458-5004

[0469] DB1 / 162798841.4 189

[0470] MBI-004PC / 108458-5004

[0471] DB1 / 162798841.4 190

[0472] MBI-004PC / 108458-5004

[0473] TABLE 10: Genes (UniRetPO)

[0474] DB1 / 162798840.4

[0475] MBI-004PC / 108458-5004

[0476] DB1 / 162798840.4

[0477] MBI-004PC / 108458-5004

[0478] DB1 / 162798840.4

[0479] MBI-004PC / 108458-5004

[0480] DB1 / 162798840.4 194

[0481] MBI-004PC / 108458-5004

[0482] DB1 / 162798840.4

[0483] MBI-004PC / 108458-5004

[0484] DB1 / 162798840.4 196

[0485] MBI-004PC / 108458-5004

[0486] DB1 / 162798840.4 197

[0487] MBI-004PC / 108458-5004

[0488] DB1 / 162798840.4 198

[0489] MBI-004PC / 108458-5004

[0490] DB1 / 162798840.4

[0491] MBI-004PC / 108458-5004

[0492] DB1 / 162798840.4

[0493] MBI-004PC / 108458-5004

[0494] DB1 / 162798840.4 201

[0495] MBI-004PC / 108458-5004

[0496] DB1 / 162798840.4 202

[0497] MBI-004PC / 108458-5004

[0498] DB1 / 162798840.4 203

[0499] MBI-004PC / 108458-5004

[0500] DB1 / 162798840.4 204

[0501] MBI-004PC / 108458-5004

[0502] DB1 / 162798840.4 205

[0503] MBI-004PC / 108458-5004

[0504] DB1 / 162798840.4 206

[0505] MBI-004PC / 108458-5004

[0506] DB1 / 162798840.4 207

[0507] MBI-004PC / 108458-5004

[0508] DB1 / 162798840.4 208

[0509] MBI-004PC / 108458-5004

[0510] DB1 / 162798840.4

[0511] MBI-004PC / 108458-5004

[0512] DB1 / 162798840.4 210

[0513] MBI-004PC / 108458-5004

[0514] DB1 / 162798840.4 211

[0515] MBI-004PC / 108458-5004

[0516] DB1 / 162798840.4 212

[0517] MBI-004PC / 108458-5004

[0518] DB1 / 162798840.4 213

[0519] MBI-004PC / 108458-5004

[0520] DB1 / 162798840.4 214

[0521] MBI-004PC / 108458-5004

[0522] DB1 / 162798840.4

[0523] MBI-004PC / 108458-5004

[0524] DB1 / 162798840.4 216

[0525] MBI-004PC / 108458-5004

[0526] DB1 / 162798840.4 217

[0527] MBI-004PC / 108458-5004

[0528] DB1 / 162798840.4 218

[0529] MBI-004PC / 108458-5004

[0530] DB1 / 162798840.4 219

[0531] MBI-004PC / 108458-5004

[0532] DB1 / 162798840.4 220

[0533] MBI-004PC / 108458-5004

[0534] DB1 / 162798840.4

[0535] MBI-004PC / 108458-5004

[0536] DB1 / 162798840.4

[0537] MBI-004PC / 108458-5004

[0538] DB1 / 162798840.4

[0539] MBI-004PC / 108458-5004

[0540] DB1 / 162798840.4 224

[0541] MBI-004PC / 108458-5004

[0542] DB1 / 162798840.4

[0543] MBI-004PC / 108458-5004

[0544] DB1 / 162798840.4 226

[0545] MBI-004PC / 108458-5004

[0546] DB1 / 162798840.4 227

[0547] MBI-004PC / 108458-5004

[0548] DB1 / 162798840.4 228

[0549] MBI-004PC / 108458-5004

[0550] DB1 / 162798840.4 229

[0551] MBI-004PC / 108458-5004

[0552] DB1 / 162798840.4 230

[0553] MBI-004PC / 108458-5004

[0554] DB1 / 162798840.4

[0555] MBI-004PC / 108458-5004

[0556] DB1 / 162798840.4

[0557] MBI-004PC / 108458-5004

[0558] DB1 / 162798840.4

[0559] MBI-004PC / 108458-5004

[0560] DB1 / 162798840.4 234

[0561] MBI-004PC / 108458-5004

[0562] DB1 / 162798840.4

[0563] MBI-004PC / 108458-5004

[0564] DB1 / 162798840.4 236

[0565] MBI-004PC / 108458-5004

[0566] DB1 / 162798840.4 237

[0567] MBI-004PC / 108458-5004

[0568] DB1 / 162798840.4

[0569] MBI-004PC / 108458-5004

[0570] DB1 / 162798840.4

[0571] MBI-004PC / 108458-5004

[0572] DB1 / 162798840.4

[0573] MBI-004PC / 108458-5004

[0574] DB1 / 162798840.4 241

[0575] MBI-004PC / 108458-5004

[0576] DB1 / 162798840.4

[0577] MBI-004PC / 108458-5004

[0578] DB1 / 162798840.4 243

[0579] MBI-004PC / 108458-5004

[0580] DB1 / 162798840.4 244

[0581] MBI-004PC / 108458-5004

[0582] DB1 / 162798840.4 245

[0583] MBI-004PC / 108458-5004

[0584] DB1 / 162798840.4 246

[0585] MBI-004PC / 108458-5004

[0586] DB1 / 162798840.4 247

[0587] MBI-004PC / 108458-5004

[0588] DB1 / 162798840.4

[0589] MBI-004PC / 108458-5004

[0590] DB1 / 162798840.4 249

[0591] MBI-004PC / 108458-5004

[0592] DB1 / 162798840.4

[0593] MBI-004PC / 108458-5004

[0594] DB1 / 162798840.4

[0595] MBI-004PC / 108458-5004

[0596] DB1 / 162798840.4

[0597] MBI-004PC / 108458-5004

[0598] DB1 / 162798840.4

[0599] MBI-004PC / 108458-5004

[0600] DB1 / 162798840.4 254

[0601] MBI-004PC / 108458-5004

[0602] DB1 / 162798840.4

[0603] MBI-004PC / 108458-5004

[0604] DB1 / 162798840.4

[0605] MBI-004PC / 108458-5004

[0606] DB1 / 162798840.4

[0607] MBI-004PC / 108458-5004

[0608] DB1 / 162798840.4

[0609] MBI-004PC / 108458-5004

[0610] DB1 / 162798840.4

[0611] MBI-004PC / 108458-5004

[0612] DB1 / 162798840.4

[0613] MBI-004PC / 108458-5004

[0614] DB1 / 162798840.4 261

[0615] MBI-004PC / 108458-5004

[0616] DB1 / 162798840.4

[0617] MBI-004PC / 108458-5004

[0618] DB1 / 162798840.4 263

[0619] MBI-004PC / 108458-5004

[0620] DB1 / 162798840.4 264

[0621] MBI-004PC / 108458-5004

[0622] DB1 / 162798840.4 265

[0623] MBI-004PC / 108458-5004

[0624] DB1 / 162798840.4 266

[0625] MBI-004PC / 108458-5004

[0626] DB1 / 162798840.4 267

[0627] MBI-004PC / 108458-5004

[0628] DB1 / 162798840.4 268

[0629] MBI-004PC / 108458-5004

[0630] DB1 / 162798840.4 269

[0631] MBI-004PC / 108458-5004

[0632] DB1 / 162798840.4 270

[0633] MBI-004PC / 108458-5004

[0634] DB1 / 162798840.4 271

[0635] MBI-004PC / 108458-5004

[0636] DB1 / 162798840.4 272

[0637] MBI-004PC / 108458-5004

[0638] DB1 / 162798840.4 273

[0639] MBI-004PC / 108458-5004

[0640] DB1 / 162798840.4 274

[0641] MBI-004PC / 108458-5004

[0642] DB1 / 162798840.4 275

[0643] MBI-004PC / 108458-5004

[0644] DB1 / 162798840.4 276

[0645] MBI-004PC / 108458-5004

[0646] DB1 / 162798840.4 zn

[0647] MBI-004PC / 108458-5004

[0648] DB1 / 162798840.4 278

[0649] MBI-004PC / 108458-5004

[0650] DB1 / 162798840.4 279

[0651] MBI-004PC / 108458-5004

[0652] DB1 / 162798840.4

[0653] MBI-004PC / 108458-5004

[0654] DB1 / 162798840.4 281

[0655] MBI-004PC / 108458-5004

[0656] DB1 / 162798840.4

[0657] MBI-004PC / 108458-5004

[0658] DB1 / 162798840.4

[0659] MBI-004PC / 108458-5004

[0660] DB1 / 162798840.4 284

[0661] MBI-004PC / 108458-5004

[0662] DB1 / 162798840.4

[0663] MBI-004PC / 108458-5004

[0664] DB1 / 162798840.4 286

[0665] MBI-004PC / 108458-5004

[0666] DB1 / 162798840.4 287

[0667] MBI-004PC / 108458-5004

[0668] DB1 / 162798840.4

[0669] MBI-004PC / 108458-5004

[0670] DB1 / 162798840.4

[0671] MBI-004PC / 108458-5004

[0672] DB1 / 162798840.4 290

[0673] MBI-004PC / 108458-5004

[0674] DB1 / 162798840.4

[0675] MBI-004PC / 108458-5004

[0676] Table 11 : Gene Ontology (GO)

[0677] DB1 / 151328179.4 292

[0678] MBI-004PC / 108458-5004

[0679] DB1 / 151328179.4 293

[0680] MBI-004PC / 108458-5004

[0681] DB1 / 151328179.4 294

[0682] MBI-004PC / 108458-5004

[0683] DB1 / 151328179.4 295

[0684] MBI-004PC / 108458-5004

[0685] DB1 / 151328179.4 296

[0686] MBI-004PC / 108458-5004

[0687] DB1 / 151328179.4 297

[0688] MBI-004PC / 108458-5004

[0689] DB1 / 151328179.4 298

[0690] MBI-004PC / 108458-5004

[0691] DB1 / 151328179.4 299

[0692] MBI-004PC / 108458-5004

[0693] DB1 / 151328179.4 300

[0694] MBI-004PC / 108458-5004

[0695] DB1 / 151328179.4 301

[0696] MBI-004PC / 108458-5004

[0697] DB1 / 151328179.4 302

[0698] MBI-004PC / 108458-5004

[0699] DB1 / 151328179.4 303

[0700] MBI-004PC / 108458-5004

[0701] DB1 / 151328179.4 304

[0702] MBI-004PC / 108458-5004

[0703] DB1 / 151328179.4 305

[0704] MBI-004PC / 108458-5004

[0705] DB1 / 151328179.4 306

[0706] MBI-004PC / 108458-5004

[0707] DB1 / 151328179.4 307

[0708] MBI-004PC / 108458-5004

[0709] DB1 / 151328179.4 308

[0710] MBI-004PC / 108458-5004

[0711] DB1 / 151328179.4 309

[0712] MBI-004PC / 108458-5004

[0713] DB1 / 151328179.4 310

[0714] MBI-004PC / 108458-5004

[0715] DB1 / 151328179.4 311

[0716] MBI-004PC / 108458-5004

[0717] DB1 / 151328179.4 312

[0718] MBI-004PC / 108458-5004

[0719] DB1 / 151328179.4 313

[0720] MBI-004PC / 108458-5004

[0721] DB1 / 151328179.4 314

[0722] MBI-004PC / 108458-5004

[0723] DB1 / 151328179.4 315

[0724] MBI-004PC / 108458-5004

[0725] DB1 / 151328179.4 316

[0726] MBI-004PC / 108458-5004

[0727] DB1 / 151328179.4 317

[0728] MBI-004PC / 108458-5004

[0729] DB1 / 151328179.4 318

[0730] MBI-004PC / 108458-5004

[0731] Table 12: Kegg Orthology (KO)

[0732] DB1 / 151328280.3 319

[0733] MBI-004PC / 108458-5004

[0734] DB1 / 151328280.3 320

[0735] MBI-004PC / 108458-5004

[0736] DB1 / 151328280.3 321

[0737] MBI-004PC / 108458-5004

[0738] DB1 / 151328280.3 322

[0739] MBI-004PC / 108458-5004

[0740] DB1 / 151328280.3 323

[0741] MBI-004PC / 108458-5004

[0742] DB1 / 151328280.3 324

[0743] MBI-004PC / 108458-5004

[0744] DB1 / 151328280.3 325

[0745] MBI-004PC / 108458-5004

[0746] DB1 / 151328280.3 326

[0747] MBI-004PC / 108458-5004

[0748] DB1 / 151328280.3 327

[0749] MBI-004PC / 108458-5004

[0750] DB1 / 151328280.3 328

[0751] MBI-004PC / 108458-5004

[0752] Table 13: Pathways

[0753] DB1 / 151328414.3 329

[0754] MBI-004PC / 108458-5004

[0755] DB1 / 151328414.3 330

[0756] MBI-004PC / 108458-5004

[0757] DB1 / 151328414.3 331

[0758] MBI-004PC / 108458-5004

[0759] DB1 / 151328414.3 332

[0760] MBI-004PC / 108458-5004

[0761] DB1 / 151328414.3 333

[0762] MBI-004PR / 108458-5004

[0763] Table 14: Protein Families (Pfam)

[0764] DB1 / 162798839.3 334

[0765] MBI-004PR / 108458-5004

[0766] DB1 / 162798839.3 335

[0767] MBI-004PR / 108458-5004

[0768] DB1 / 162798839.3 336

[0769] MBI-004PR / 108458-5004

[0770] DB1 / 162798839.3 337

[0771] MBI-004PR / 108458-5004

[0772] DB1 / 162798839.3 338

[0773] MBI-004PR / 108458-5004

[0774] DB1 / 162798839.3 339

[0775] MBI-004PR / 108458-5004

[0776] DB1 / 162798839.3 340

[0777] MBI-004PR / 108458-5004

[0778] DB1 / 162798839.3 341

[0779] MBI-004PR / 108458-5004

[0780] DB1 / 162798839.3 342

[0781] MBI-004PR / 108458-5004

[0782] DB1 / 162798839.3 343

[0783] MBI-004PR / 108458-5004

[0784] DB1 / 162798839.3 344

[0785] MBI-004PR / 108458-5004

[0786] DB1 / 162798839.3 345

[0787] MBI-004PR / 108458-5004

[0788] DB1 / 162798839.3 346

[0789] MBI-004PR / 108458-5004

[0790] DB1 / 162798839.3 347

[0791] MBI-004PR / 108458-5004

[0792] DB1 / 162798839.3 348

[0793] MBI-004PR / 108458-5004

[0794] DB1 / 162798839.3 349

[0795] MBI-004PR / 108458-5004

[0796] DB1 / 162798839.3 350

[0797] MBI-004PR / 108458-5004

[0798] DB1 / 162798839.3 351

[0799] MBI-004PR / 108458-5004

[0800] DB1 / 162798839.3 352

[0801] MBI-004PR / 108458-5004

[0802] DB1 / 162798839.3 353

[0803] MBI-004PR / 108458-5004

[0804] DB1 / 162798839.3 354

[0805] MBI-004PR / 108458-5004

[0806] DB1 / 162798839.3 355

[0807] MBI-004PR / 108458-5004

[0808] DB1 / 162798839.3 356

[0809] MBI-004PR / 108458-5004

[0810] DB1 / 162798839.3 357

[0811] MBI-004PR / 108458-5004

[0812] DB1 / 162798839.3 358

[0813] MBI-004PR / 108458-5004

[0814] DB1 / 162798839.3 359

[0815] MBI-004PR / 108458-5004

[0816] DB1 / 162798839.3 360

[0817] MBI-004PR / 108458-5004

[0818] DB1 / 162798839.3 361

[0819] MBI-004PC / 108458-5004

[0820] Table 15: Reactions (RXN)

[0821] DB1 / 151328712.4 362

[0822] MBI-004PC / 108458-5004

[0823] DB1 / 151328712.4 363

[0824] MBI-004PC / 108458-5004

[0825] DB1 / 151328712.4 364

[0826] MBI-004PC / 108458-5004

[0827] DB1 / 151328712.4 365

[0828] MBI-004PC / 108458-5004

[0829] DB1 / 151328712.4 366

[0830] MBI-004PC / 108458-5004

[0831] DB1 / 151328712.4 367

[0832] MBI-004PC / 108458-5004

[0833] DB1 / 151328712.4 368

[0834] MBI-004PC / 108458-5004

[0835] DB1 / 151328712.4 369

[0836] MBI-004PC / 108458-5004

[0837] DB1 / 151328712.4 370

[0838] MBI-004PC / 108458-5004

[0839] DB1 / 151328712.4 371

[0840] MBI-004PC / 108458-5004

[0841] DB1 / 151328712.4 372

[0842] MBI-004PC / 108458-5004

[0843] DB1 / 151328712.4 373

[0844] MBI-004PC / 108458-5004

[0845] DB1 / 151328712.4 374

[0846] MBI-004PC / 108458-5004

[0847] DB1 / 151328712.4 375

[0848] MBI-004PC / 108458-5004

[0849] DB1 / 151328712.4 376

[0850] MBI-004PC / 108458-5004

[0851] Table 16: Microbial taxa

[0852] DB1 / 151330358.2 377

[0853] MBI-004PC / 108458-5004

[0854] DB1 / 151330358.2 378

[0855] MBI-004PC / 108458-5004

[0856] DB1 / 151330358.2 379

[0857] MBI-004PC / 108458-5004

[0858] DB1 / 151330358.2 380

[0859] MBI-004PC / 108458-5004

[0860] DB1 / 151330358.2 381

[0861] MBI-004PC / 108458-5004

[0862] DB1 / 151330358.2 382

[0863] MBI-004PC / 108458-5004

[0864] DB1 / 151330358.2 383

[0865] MBI-004PC / 108458-5004

[0866] DB1 / 151330358.2 384

[0867] MBI-004PC / 108458-5004

[0868] DB1 / 151330358.2 385

[0869] MBI-004PC / 108458-5004

[0870] DB1 / 151330358.2 386

[0871] MBI-004PC / 108458-5004

[0872] DB1 / 151330358.2 387

[0873] MBI-004PC / 108458-5004

[0874] DB1 / 151330358.2 388

[0875] MBI-004PC / 108458-5004

[0876] DB1 / 151330358.2 389

[0877] MBI-004PC / 108458-5004

[0878] DB1 / 151330358.2 390

[0879] MBI-004PC / 108458-5004

[0880] DB1 / 151330358.2 391

[0881] MBI-004PC / 108458-5004

[0882] DB1 / 151330358.2 392

[0883] MBI-004PC / 108458-5004

[0884] DB1 / 151330358.2 393

[0885] MBI-004PC / 108458-5004

[0886] DB1 / 151330358.2 394

[0887] MBI-004PC / 108458-5004

[0888] DB1 / 151330358.2 395

[0889] MBI-004PC / 108458-5004

[0890] DB1 / 151330358.2 396

[0891] MBI-004PC / 108458-5004

[0892] DB1 / 151330358.2 397

Claims

MBI-004PC / 108458-5004CLAIMS1. A method for evaluating a subject for the presence of a colorectal neoplasm, comprising: quantifying genetic elements from a biological sample from the subject, wherein the abundance or prevalence of the genetic elements is associated with colorectal cancer (CRC), colorectal adenoma (CRA), or colorectal advanced adenoma (CRAA) to thereby prepare an abundance profile of the genetic elements, and wherein the genetic elements comprise elements associated with microbial taxonomic classification and genetic elements associated with one or more microbial gene functions; evaluating the abundance profile for a signature indicating the presence of CRC, CRA, and / or CRAA in the subject, and determining whether the subject is likely to have CRC, CRA, or CRAA.

2. The method of claim 1, wherein the subject is at low risk for CRC, CRA, CRAA, or colorectal polyps.

3. The method of claim 2, wherein the subj ect has no previous incidence of CRC, CRA, or CRAA,4. The method of claim 2 or 3, wherein the method is performed as an alternative to colonoscopy.

5. The method of claim 1, wherein the subject is at high or medium risk for CRC, CRA, CRAA, or colorectal polyps.

6. The method of claim 5, wherein the subject has prior incidence of CRC, CRA, or CRAA, and / or family history of CRC.

7. The method of claim 5 or 6, wherein the method is performed at least once annually or at least every other year.DBl / 162792885.1 78MBI-004PC / 108458-50048. The method of any one of claims 1 to 7, wherein the subject is at least 45, or at least 50 years of age, or at least 55 years of age, or at least 60 years of age.

9. The method of any one of claims 1 to 7, wherein the subject is less than 45 years of age, or less than 50 years of age.

10. The method of any one of claims 1 to 9, wherein the biological sample is a fecal, blood, serum, intestinal mucosa, mucosal swab, colonoscopy aspirant, lavage, or biopsy tissue sample.

11. The method of claim 10, wherein the biological samples is a fecal sample or intestinal mucosa sample.

12. The method of any one of claims 1 to 11, wherein the genetic elements are quantified by a method comprising nucleic acid sequencing, PCR, qPCR, or microarray.

13. The method of claim 12, wherein the genetic elements are quantified by nucleic acid sequencing, and which involves sequencing at least about 20,000,000 reads.

14. The method of claim 13, wherein the nucleic acid sequencing involves sequencing at least about 40,000,000 reads.

15. The method of any one of claims 12 to 14, wherein the nucleic acid sequencing comprises one or more of shotgun metagenomic sequencing, rDNA sequencing, targeted amplicon nucleic acid sequencing, or hybridization capture probe sequencing.

16. The method of claim 15, wherein the nucleic acid sequencing comprises 16S rDNA, 18S rDNA, or ITS amplicon sequencing; and comprises targeted amplicon nucleic acid sequencing or hybridization capture probe sequencing.DBl / 162792885.1 79MBI-004PC / 108458-500417. The method of claim 15 or 16, wherein one or more genetic elements are quantified by capturing from a sequencing library, and optionally amplified by PCR, followed by sequencing.

18. The method of any one of claims 1 to 17, wherein raw data is determined from a plurality of independent studies, and is batch corrected.

19. The method of any one of claims 1 to 18, wherein at least a portion of the genetic elements are indicative of colorectal cancer (CRC).

20. The method of claim 19, wherein the genetic elements that are indicative of CRC comprise elements annotated for a plurality of Enzyme Commission (EC) number, protein family, pathway, reaction, gene orthology, gene homology, gene ontology, and taxonomy.

21. The method of claim 20, wherein the genetic elements that are indicative of CRC include features of microbial taxonomy, along with at least 2, at least 3, at least 4, or at least 5 selected from: (1) features defining an enzyme or reaction classification, (2) features defining gene orthology or ontology, (3) features defining gene homology, (4) features defining pathways, and (5) features defining protein families and / or domains.

22. The method of claim 20 or 21, wherein the genetic elements that are indicative of CRC comprise at least five taxonomic or gene function features listed in Tables 8 to 17.

23. The method of claim 22, wherein the genetic elements that are indicative of CRC comprise at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or gene function features listed in Tables 8 to 17.

24. The method of claim 23, wherein the genetic elements that are indicative of CRC comprise at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 taxonomic features listed in Table 16; and at least one, at least two, at least five, atDBl / 162792885.1 80MBI-004PC / 108458-5004 least about 10, at least about 20, or at least about 50 gene function features listed in Tables 8 to 15.

25. The method of any one of claims 19 to 24, wherein the genetic elements that are indicative of CRC have differential abundance or differential prevalence in samples from CRC subjects, as compared to control subjects.

26. The method of any one of claims 1 to 25, wherein at least a portion of the genetic elements are indicative of colorectal adenoma (CRA).

27. The method of claim 26, wherein the genetic elements that are indicative of CRA comprise elements annotated for a plurality of: Enzyme Commission (EC) number, protein family, pathway, reaction, gene orthology, gene homology, gene ontology, and taxonomy.

28. The method of claim 27, wherein the genetic elements that are indicative of CRA include features of microbial taxonomy, along with at least 2, at least 3, at least 4, or at least 5 selected from: (1) features defining an enzyme or reaction classification, (2) features defining gene orthology or ontology, (3) features defining gene homology, (4) features defining pathways, and (5) features defining protein families and / or domains.

29. The method of claim 28, wherein the genetic elements that are indicative of CRA comprise at least five taxonomic or gene function features listed in Table 18.

30. The method of claim 29, wherein the genetic elements that are indicative of CRA comprise at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or gene function features listed in Table 18.

31. The method of any one of claims 27 to 30, wherein the genetic elements that are indicative of CRA comprise at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 taxonomic features listed in Table 18; and at least one, at leastDBl / 162792885.1 81MBI-004PC / 108458-5004 two, at least five, at least about 10, at least about 20, or at least about 50 gene function features listed in Table 18.

32. The method of any one of claims 27 to 31, wherein the genetic elements that are indicative of CRA have differential abundance or differential prevalence in samples from CRA subjects, as compared to control subjects.

33. The method of any one of claims 1 to 32, wherein at least a portion of the genetic elements are indicative of colorectal advanced adenoma (CRAA).

34. The method of claim 33, wherein the genetic elements that are indicative of CRAA comprise elements annotated for a plurality of: Enzyme Commission (EC) number, protein family, pathway, reaction, gene orthology, gene homology, gene ontology, and taxonomy.

35. The method of claim 34, wherein the genetic elements that are indicative of CRAA include features of microbial taxonomy, along with at least 2, at least 3, at least 4, or at least 5 selected from: (1) features defining an enzyme or reaction classification, (2) features defining gene orthology or ontology, (3) features defining gene homology, (4) features defining pathways, and (5) features defining protein families and / or domains.

36. The method of claim 34 or 35, wherein the genetic elements that are indicative of CRAA comprise at least five taxonomic or gene function features listed in Table 19.

37. The method of claim 36, wherein the genetic elements that are indicative of CRAA comprise at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or gene function features listed in Table 19.

38. The method of any one of claims 35 to 37, wherein the genetic elements that are indicative of CRAA comprise at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 taxonomic features listed in Table 19; and at least one, at leastDBl / 162792885.1 82MBI-004PC / 108458-5004 two, at least five, at least about 10, at least about 20, or at least about 50 gene function features listed in Table 19.

39. The method of any one of claims 36 to 38, wherein the genetic elements indicative of CRAA have differential abundance or differential prevalence in samples from CRAA subjects, as compared to control subjects.

40. The method of any one of claims 1 to 39, wherein the abundance profile is evaluated for signatures indicating the presence or absence of each of CRC, CRA, and CRAA.

41. The method of any one of claims 1 to 40, wherein the abundance profile is evaluated for signatures classifying the profile as either (1) CRC and / or CRAA, or (2) CRA or control.

42. The method of any one of claims 1 to 41, wherein:(a) the signature indicating the presence or absence of CRC is trained with samples from a CRC cohort and samples from a control cohort by machine learning;(b) the signature indicating the presence or absence of CRA is trained with samples from a CRA cohort and samples from a control cohort by machine learning; and(c) the signature indicating the presence or absence of CRAA is trained with samples from a CRAA cohort and samples from a control cohort by machine learning; and(d) the signature indicating the presence or absence of diagnostic target variables classified as Aneo (CRC + CRAA; positive) and Naneo (CRA + Control; negative), in a binary classification context, is trained with samples from CRC / CRAA and CRA / Control cohorts by machine learning.

43. The method of claim 42, wherein the signature of (a), (b), (c), and / or (d) is trained with genetic elements having an abundance or prevalence that is associated with CRC, CRA, or CRAA, respectively, and which is statistically significant, optionally wherein the genetic elements are identified by logistic regression prior to training the signatures.DBl / 162792885.1 83MBI-004PC / 108458-500444. The method of claim 43, wherein the statistical significance is an FDR-corrected p- value of 0.05 or less, or a p-value of 0.01 or less.

45. The method of any one of claims 42 to 44, wherein the signature(s) are trained using a plurality of machine learning algorithms.

46. The method of claim 45, wherein at least one machine learning algorithm is supervised machine learning.

47. The method of claim 46, wherein the machine learning algorithms further comprise one or more of unsupervised and semi-supervised machine learning.

48. The method of any one of claims 42 to 47, wherein the machine learning comprises one or more of parametric / non-parametric distance measures, logistic regression, support vector machines, decision trees, random forests, neural networks, probit regression, Fisher's linear discriminant, Naive Bayes classifier, perceptron, quadratic classifiers, kernel estimation, k-nearest neighbor, learning vector quantization, and principal components analysis.

49. The method of any one of claims 42 to 48, wherein the machine learning comprises comparative machine learning modeling, optimization, testing, evaluation, and ranking of models, including with one or more of deep learning, gradient boosted, neural networks, ensemble, or blender modeling algorithms, including one or more of Gradient Boosted Trees Classifier, extreme Gradient Boosted Trees Classifiers, Light Gradient Boosted Trees Classifiers, Light Gradient Boosting on Elastic Net Predictions, Keras Slim Residual Neural Network Classifiers, Generalized Additive Models, Elastic Net Classifiers, Random Forest Classifiers, Deep Forest Classifiers, Average Blender Classifiers, TensorFlow Multilayer Perceptron Classifiers, TensorFlow Neural Network Classifiers, and Rule-Fit Classifiers.

50. The method of any one of claims 42 to 49, wherein the signature(s) have a sensitivity for classifying samples for the presence or absence of CRA, CRAA, or CRC of at least aboutDBl / 162792885.1 84MBI-004PC / 108458-50040.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about0.95.

51. The method of any one of claims 42 to 50, wherein the signature(s) have a specificity for classifying samples for the presence or absence of CRA, CRAA, or CRC of at least about 0.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about 0.95.

52. The method of any one of claims 1 to 51, wherein:(a) if the subject is not identified as likely to have CRA, CRAA, or CRC, no further procedure is conducted, and(b) if the subject is identified as likely having one or more of CRA, CRAA, or CRC, a further procedure or treatment is initiated.

53. The method of claim 52, wherein the further procedure comprises imaging of the colon.

54. The method of claim 53, wherein the procedure is a colonoscopy, which optionally involves removal of one or more polyps and / or biopsy of growths suspected of comprising CRC55. The method of any one of claims 52 to 54, wherein if the subject is confirmed to have CRC, the subject is treated for CRC by one or more of surgery, chemotherapy, radiation therapy, and immunotherapy.

56. A method for preparing a signature of genetic elements indicative of the presence of a colorectal neoplasm, the method comprising: providing a training cohort of biological samples from subjects confirmed to have CRC, CRA, or CRAA, or RNA or DNA isolated therefrom; conducting genomic nucleic acid sequencing of DNA isolated from the samples to prepare a set of genetic elements for each sample;DBl / 162792885.1 85MBI-004PC / 108458-5004 annotating the genetic elements for microbial taxonomy and microbial gene functions; selecting genetic elements whose abundance or prevalence has a statistically significant association with CRC, CRA, or CRAA; and training a gene signature with the selected genetic elements that classifies samples for the presence or absence of CRA, CRAA, or CRC, wherein the gene signature comprises microbial taxonomic classification features and microbial gene function features.

57. The method of claim 56, wherein the samples are selected from fecal, blood, serum, plasma, urine, saliva, biopsy tissues, mucosa tissue sample or swab, and intestinal lavage or aspirant.

58. The method of claim 57, wherein samples are fecal samples or mucosa tissue samples.

59. The method of any one of claims 56 to 58, wherein the nucleic acid sequencing involves sequencing at least about 20,000,000 reads per sample.

60. The method of claim 59, wherein the nucleic acid sequencing involves sequencing at least about 40,000,000 reads.

61. The method of any one of claims 56 to 60, wherein the nucleic acid sequencing comprises one or more of shotgun metagenomic sequencing, rDNA sequencing, and targeted nucleic acid sequencing.

62. The method of claim 61, wherein the nucleic acid sequencing comprises 16S rDNA sequencing and / or 18S rDNA sequencing and / or ITS sequencing; and one or more of shotgun sequencing and targeted nucleic acid sequencing.

63. The method of any one of claims 56 to 62, wherein genetic elements within the samples are amplified, optionally by PCR.DBl / 162792885.1 86MBI-004PC / 108458-500464. The method of any one of claims 56 to 63, wherein genetic elements are assigned to a reference genome for taxonomic classification, and / or genetic elements are assigned to a gene function.

65. The method of claim 64, wherein the genetic elements are annotated for a plurality of: Enzyme Commission (EC) number, protein family, pathway, reaction, gene orthology, gene homology, gene ontology, and microbial taxonomy.

66. The method of claim 64 or 65, wherein raw data is determined from a plurality of independent studies, and is batch corrected.

67. The method of any one of claims 56 to 66, wherein microbial taxonomic classification features and microbial gene function features are selected that have a statistically significant differential abundance or differential prevalence in samples from CRC, CRA, or CRAA subjects, as compared to control subjects.

68. The method of claim 67, wherein the microbial gene function features include at least 2, at least 3, at least 4, or at least 5 selected from: (1) features defining an enzyme or reaction classification, (2) features defining gene orthology or ontology, (3) features defining gene homology, (4) features defining pathways, and (5) features defining protein families and / or domains.

69. The method of claim 67 or 68, wherein the features comprise at least five taxonomic and / or gene function features, and which are optionally listed in Tables 8 to 19.

70. The method of claim 69, wherein the features comprise at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or gene function features, which are optionally listed in Tables 8 to 19.DBl / 162792885.1 87MBI-004PC / 108458-500471 . The method of claim 69 or 70, wherein the features comprise at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 taxonomic features that are optionally listed in Table 16; and at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 gene function features that are optionally listed in Tables 8 to 19.

72. The method of any one of claims 56 to 71, wherein at least three gene signatures are trained that:(1) classify samples for the presence or absence of CRA,(2) classify samples for the presence or absence of CRAA, and(3) classify samples for the presence or absence of CRC.

73. The method of claim 72, wherein:(a) the signature classifying samples for the presence or absence of CRC is trained with samples from a CRC cohort and samples from a control cohort by machine learning;(b) the signature classifying samples for the presence or absence of CRA is trained with samples from a CRA cohort and samples from a control cohort by machine learning; and(c) the signature classifying samples for the presence or absence of CRAA is trained with samples from a CRAA cohort and samples from a control cohort by machine learning.

74. The method of claim 73, wherein the signature(s) are trained using a plurality of machine learning algorithms.

75. The method of claim 73 or 74, wherein at least one machine learning algorithm is supervised machine learning.

76. The method of claim 75, wherein the machine learning algorithms further comprise one or more of unsupervised or semi-supervised machine learning.DBl / 162792885.1 88MBI-004PC / 108458-500477. The method of any one of claims 73 to 76, wherein the machine learning comprises one or more of parametric / non-parametric distance measures, logistic regression, support vector machines, decision trees, random forests, neural networks, probit regression, Fisher's linear discriminant, Naive Bayes classifier, perceptron, quadratic classifiers, kernel estimation, k-nearest neighbor, learning vector quantization, and principal components analysis.

78. The method of any one of claims 73 to 77, wherein the machine learning comprises comparative machine learning modeling, optimization, testing, evaluation, and ranking of models, including with one or more of deep learning, gradient boosted, neural networks, ensemble, or blender modeling algorithms, including one or more of Gradient Boosted Trees Classifier, extreme Gradient Boosted Trees Classifiers, Light Gradient Boosted Trees Classifiers, Light Gradient Boosting on Elastic Net Predictions, Keras Slim Residual Neural Network Classifiers, Generalized Additive Models, Elastic Net Classifiers, Random Forest Classifiers, Deep Forest Classifiers, Average Blender Classifiers, TensorFlow Multilayer Perceptron Classifiers, TensorFlow Neural Network Classifiers, and Rule-Fit Classifiers.

79. The method of any one of claims 73 to 78, wherein the signature(s) have a sensitivity for classifying samples for the presence or absence of CRC, CRA, or CRAA of at least about 0.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about 0.95.

80. The method of any one of claims 73 to 79, wherein the signature(s) have a specificity for classifying samples for the presence or absence of CRC, CRA, or CRAA of at least about 0.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about 0.95.

81. A method for preparing a genetic signature of genetic elements indicative of the presence of a disorder, the method comprising: providing a training cohort of biological samples from subjects confirmed to have the disorder, or RNA or DNA isolated therefrom;DBl / 162792885.1 89MBI-004PC / 108458-5004 conducting genomic nucleic acid sequencing of DNA isolated from the samples to prepare a set of genetic elements for each sample; annotating the genetic elements for microbial taxonomy and microbial gene functions; selecting genetic elements whose abundance or prevalence has a statistically significant association with the disorder; and training a gene signature with the selected genetic elements that classifies samples for the presence or absence of the disorder, wherein the gene signature comprises microbial taxonomic classification features and microbial gene function features.DBl / 162792885.1 90

Citation Information

Patent Citations

  • Method for diagnosing adenomas and / or colorectal cancer (CRC) based on analyzing the gut microbiome

    EP2955232A1

  • Improved method for the screening, diagnosis and / or monitoring of colorectal advanced neoplasia, advanced adenoma and / or colorectal cancer

    US20220145400A1