Discovery, functional analysis, and diagnosis of colorectal adenoma and cancer biomarkers.

Biomarkers from microbial species in fecal samples, combined with machine learning, improve colorectal cancer screening by enhancing detection sensitivity and specificity, reducing the need for invasive procedures.

JP2026514442APending Publication Date: 2026-05-11PRESCIENT METABIOMICS JV LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
PRESCIENT METABIOMICS JV LLC
Filing Date
2024-03-29
Publication Date
2026-05-11

AI Technical Summary

Technical Problem

Current colorectal cancer screening methods, such as colonoscopy, are invasive and have low compliance due to discomfort, while non-invasive tests like fecal immunochemical tests (FIT) and multi-target fecal assays have limited sensitivity for detecting colorectal adenomas and cancers, highlighting a gap in healthcare systems for early and accurate detection.

Method used

Development of biomarkers derived from microbial species in fecal and other samples, combined with machine learning models, for detecting colorectal adenomas and cancers through metagenomic and multi-omics analysis of biological samples, providing improved detection and reducing the need for invasive procedures.

Benefits of technology

Enhances the detection of colorectal adenomas and cancers with higher sensitivity and specificity, enabling early detection and reducing the number of invasive colonoscopies required, particularly for low-risk individuals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026514442000001_ABST
    Figure 2026514442000001_ABST
Patent Text Reader

Abstract

In various aspects and embodiments, the Disclosure provides methods for evaluating subjects for the presence or absence of colorectal neoplasms, such as colorectal cancer (CRC), colorectal adenoma (CRA), and progressive colorectal adenoma (CRAA), by metagenomic and multi-omics analysis of feces or other biological samples. In other aspects, the Disclosure provides methods for generating machine learning models or “signatures” based on metagenomic and multi-omics analysis of feces or other biological samples to evaluate subjects for the presence or absence of colorectal diseases, including but not limited to CRC, CRA, and CRAA.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Priority This application claims priority and the benefit thereof to U.S. Provisional Application No. 63 / 455,698, filed on Mar. 30, 2023, the entire disclosure of which is incorporated herein by reference.

[0002] Sequence Listing This application includes a sequence listing submitted in XML format via EFS-Web. The content of the XML copy named "MBI-003PC_108458-5003_Sequence_Listing" was created on Mar. 29, 2024, is 73,728 bytes in size, and the entire content thereof is incorporated herein by reference.

Background Art

[0003] Colorectal cancer is one of the most common cancers worldwide, with approximately 1.8 million new cases and over 700,000 cases of rectal cancer reported in 2018. (Bray et al., Global cancer statistics 2018: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries, CA Cancer J Clin. 2018;68(6):394-424). Furthermore, colorectal cancer is currently the leading cause of death for men under 50 years of age. (Siegel RL, et al. Cancer Statistics, 2024, CA: A Cancer Journal for Clinicians (2024)). Despite strong evidence showing that screening individuals with average CRC risk reduces mortality, compliance among individuals is limited due to the invasiveness, discomfort, and fear associated with colonoscopy. Lauby-Secretan et al., The IARC Perspective on Colorectal Cancer Screening, N Engl J Med. 2018;378(18):1734-1740. This creates a significant gap in healthcare systems and, in particular, in the prevention of colorectal cancer (CRC), highlighting the need for highly sensitive, accurate, and non-invasive diagnostics for detecting colorectal adenomas and cancers. This disclosure addresses this gap by providing biomarkers derived from microbial species present in fecal and other samples. [Brief explanation of the drawing]

[0004] [Figure 1] A is a principal coordinate analysis (PCoA) plot of the microbiome profiles derived from samples within the analyzed trials. B is a PCoA plot of the microbiome profiles derived from samples for each disease class. C is a PCoA plot of the microbiome profiles derived from samples within the analyzed trials after supervised normalization. D is a PCoA plot of the microbiome profiles derived from samples for each disease class after supervised normalization. [Figure 2] This is a cross-correlation plot. A CRC model was trained using samples from the trials listed at the top. These models were then used to predict the sample (test set) from each trial. [Figure 3] This workflow describes two different feature selection methods, Feature Importance Rank Ensembling (FIRE) and Statistical Estimation of Associations between Microbial Communities and Host Phenotypes (SIAMCAT), either separately or in combination. [Figure 4] This figure shows the independent feature selection method SIAMCAT. Features are ranked according to their significance score. Box plots displaying the relative abundance of the sample are shown to visualize the contribution of features to the difference representation, scaling change, prevalence shift, and area under the curve (AUC). [Figure 5] This figure shows performance evaluations obtained using mean AUC based on a combination of taxonomic and genetic features (KEGG ortholog (KO) group, or taxonomic groups associated with colorectal adenoma (CRA), progressive colorectal adenoma (CRAA), or colorectal cancer (CRC)). Here, the feature selection methods FIRE and SIAMCAT are applied separately and together. [Figure 6] These are Venn diagrams showing the top 800, 500, 200, 100, 50, and 20 features generated from FIRE and SIAMCAT combinations for CRA, CRAA, and CRC. The total number and percentage of overlapping features across disease classes are shown. The number of features corresponding to taxonomic (T) and genetic (K) features are shown separately. The number of analyzed features is shown in descending order from left to right and top to bottom (800, 500, 200, 100, 50, and 20 features). [Figure 7] This is a Venn diagram showing the direction of change in characteristics. The central plot shows Venn diagrams of the number of overlapping characteristics among 800 taxonomic and genetic characteristics generated from the FIRE and SIAMCAT combinations for CRA, CRAA, and CRC. The surrounding diagrams consider the direction of change in a pairwise manner to show similarities and differences between characteristics common to the disease classes. (FC) = magnification change. [Figure 8] This figure shows the differential representation of bacterial taxonomic classes between disease classes. The number of features at the class level was summed as either an overrepresentation (positive value) or underrepresentation (negative value) compared to the control sample. [Figure 9] This figure shows the differential representation of bacterial families between disease classes. The number of features at the family level was summed as either an overrepresentation (positive value) or underrepresentation (negative value) compared to the control sample. The family distinguishes CRAA from the control and other disease classes. [Figure 10] This cladogram shows important taxonomic features that indicate the differential representation of CRA, CRAA, and CRC. [Figure 11] This figure shows the overexpression of pathogenicity-determining factors, biofilm formation, invasin, pathogenicity factor (VF) regulators, LPS, secretory system, and pathogenicity effectors in CRA, CRAA, and CRC compared to healthy controls. [Figure 12] A and B are Euclidean distance plots showing the distance from the centroid calculated for all taxonomic ranks based on relative abundance (A) and supervised normalization (B) for all samples in each test. [Figure 13]A to N are graphs showing the quantitative detection of CRC taxonomic features by sequencing and PCR: (A) Fusobacterium nucleatum, (B) Streptococcus salivarius, (C) Parvimonas micra, (D) Roseburia intestinalis, (E) Eubacterium ventriosum, (F) Clostridium symbiosum, (G) Gemella morbillorum, (H) Cloacibacillus evryensis, (I) Bacteroides stercoris, (J) Butyricimonas virosa, (K) Collinsella stercoris, (L) Fecalibacterium prausnitzii, (M) Intestimonas butyriciproducens, and (N) Vaillonella parvula. [Figure 14] A to Q are graphs showing the quantitative detection of CRA taxonomic features by sequencing and PCR: (A) Bacteroides salyersiae, (B) Dorea formicigenerans, (C) Ruminococcus bicirculans, (D) Clostridium spiroforme, (E) Alistipes shahii, (F) Dorea longicatena, (G) Gemella sanguinis, (H) Streptococcus thermophilus, (I) Bifidobacterium animalis, (J) Bifidobacterium pseudocatenulatum, (K) Escherichia coli, (L) Gordonibacter pamelaeae, (M) Parabacteroides goldsteinii, (N) Bifidobacterium adolescentis, (O) Bacteroides cellulosilyticus, (P) Bacteroides caccae, and (Q) Bacteroides nordii. [Figure 15]A-E are graphs showing the quantitative detection of CRAA taxonomic features by sequencing and PCR: (A) Intestinibacter bartletti, (B) Bacteroides xylanisolvens, (C) Bacteroides thetaiotaomicron, (D) Flavinofactor plautii, and (E) Mogibacterium diversum. [Modes for carrying out the invention]

[0005] In various aspects and embodiments, the Disclosure provides methods for evaluating a subject for the presence or absence of colorectal neoplasms, such as colorectal cancer (CRC), colorectal adenoma (CRA), and progressive colorectal adenoma (CRAA), by metagenomic and multi-omics analysis of biological samples (hereinafter referred to as “biological samples”), such as feces, blood, serum, plasma, urine, saliva, biopsy tissue, mucosal tissue samples or swabs, bowel cleansing fluids or aspirates, and other biological fluids and cell samples, containing human and microbiome DNA, RNA, proteins, and other molecules for molecular analysis. In other aspects, the Disclosure provides methods for generating machine learning models or “signatures” (biomarker profiles or patterns) based on metagenomic and multi-omics analysis of biological samples, including fecal samples, for evaluating a subject for the presence or absence of colorectal diseases, including but not limited to CRC, CRA, and CRAA.

[0006] CRC is a heterogeneous disease, and the majority of cases are thought to be sporadic without underlying genetic features. (Frank et al., Concordant and discordant familial cancer: Familiar risks, proportions and population impact, Int J Cancer 2017;140(7):1510-1516). A wide variety of environmental factors, including Western diet, obesity, smoking, alcohol consumption, and lack of exercise, are known CRC risk factors. Of these risk factors, diet is the most important, with an estimated 38% of initiation CRC cases being related to diet. Further evidence of the environmental impact of CRC is based on the finding that the incidence of CRC is affected by migration and that the risk of developing CRC in a subject changes based on the diet and lifestyle of the receiving country. It is also known that each of the above CRC risk modifiers modulates the composition of the gut microbiota.Greathouse et al.,Gut microbiome meta-analysis reveals dysbiosis is independent of body mass index in predicting risk of obesity-associated CRC,BMJ Open Gastroenterol 2019;6(1):e000247;Lee et al.,Association between Cigarette Smoking Status and Composition of Gut Microbiota:Population-Based Cross-Sectional Study,J Clin Med 2018;7(9):282 ;Rodriguez-Gonzalez et al.,Microbiota and Alcohol Use Disorder: Are Psychobiotics a Novel Therapeutic Strategy?,Curr Pharm Des 2020;26(20):2426-2437;Allen et al.,Exercise Alters Gut Microbiota Composition and Function in Lean and Obese Humans,Med Sci Sports Exerc 2018;50(4):747-757. This relevance has drawn substantial attention to the gut microbiota as a potential mediator of CRC initiation and / or progression. The large number of species and genes encoded by the gut microbiota represents a potential source of biomarkers for the diagnosis and prognosis of early, pre-malignant, progressive adenomas, and CRC.

[0007] In various embodiments, the Disclosure enables the detection of colorectal adenoma (CRA), progressive colorectal adenoma (CRAA), and / or colorectal cancer (CRC) (and other colon disorders) based on the composition of the target microbiome (e.g., present in other biological samples such as fecal or mucosal tissue samples) and / or other multi-omics molecular analytes. In various embodiments, the Disclosure provides taxonomic and genetic features to distinguish healthy subjects from those with early and advanced adenomas, as well as those with carcinomas. In some embodiments, the Disclosure provides machine learning (ML) models that avoid various pitfalls associated with metagenomic and multi-omics data (e.g., data heterogeneity, noise, overfitting, etc.), including metagenomic and multi-omics data collected using heterogeneous methods and analytical procedures.

[0008] The methods disclosed herein utilize the most informative biomarkers for each disease class. For example, in some embodiments, the methods provide features that are crucial for distinguishing CRC, CRA, or CRAA from healthy controls. While the features of CRC were heavily reliant on taxonomic features, the features of CRA and CRAA were more balanced in their representation of genetic and taxonomic features. As disclosed herein, the optimal features for each disease class show little overlap and indicate that the progression from adenoma to carcinoma reflects a unique selective environment for the microbiome that does not follow a simple linear relationship.

[0009] In one embodiment, the Disclosure provides a method for evaluating a biological subject for the presence of a colorectal neoplasm. In this embodiment, the Disclosure provides a method for screening a subject as an alternative to invasive procedures such as colonoscopy, thereby increasing screening compliance and enabling early detection of neoplasms. In various embodiments, the method includes quantifying genetic elements from a biological sample from the subject, such as a fecal sample. Other biological samples that allow sampling of the microbiome, including the gut microbiota (e.g., mucosal tissue samples, blood, saliva), may also be used. The genetic elements are related to colorectal cancer (CRC), colorectal adenoma (CRA), or progressive colorectal adenoma (CRAA) and can be selected using a machine learning model described herein. The genetic elements include elements related to microbial taxonomic classification and elements related to the function of one or more microbial genes. Thus, the process creates an abundance profile of the genetic elements, which is evaluated for signatures indicating the presence or absence of CRC, CRA, and / or CRAA in the subject. Therefore, subjects can be identified as likely to have (or not have) CRC, CRA, and / or CRAA. This process may provide a binary classification (i.e., presence or absence) or a statistical output indicating the likelihood that a subject has CRC, CRA, or CRAA. In various embodiments, the method provides improved detection of adenomas (CRA and / or CRAA) compared to known detection tests.

[0010] Among the non-invasive CRC detection tests described, there is the fecal immunochemical test (FIT), which has a limited sensitivity of 79% for detecting CRC and a low sensitivity (approximately 25%) for detecting progressive adenoma. (Lee et al., Accuracy of fecal immunochemical tests for colorectal cancer: systematic review and meta-analysis, Ann Intern Med 2014;160(3):171; Hundt et al., Comparative evaluation of immunochemical fecal occult blood tests for colorectal adenoma detection, Ann Intern Med 2009;150(3):162-9). Multi-target fecal assays quantitatively examine KRAS mutations, NDRG4 and BMP3 methylation abnormalities, along with β-actin and hemoglobin immunoassays. This assay performed better than FIT, detecting CRC cases with higher sensitivity (approximately 92% compared to 74%), although advanced premalignant lesions were still poorly detected by both assays (approximately 42% and 24%, respectively). Imperiale et al., Multitarget stool DNA testing for colorectal-cancer screening, N Engl J Med. 2014;370(14):1287-97. These results highlight another significant gap in the healthcare system: the relatively inadequate ability of existing non-invasive methods to detect early and advanced adenomas. Developing diagnostic methods to address this gap could significantly improve the detection of premalignant lesions and reduce the number of colonoscopies required for subjects at average risk.

[0011] In some embodiments, subjects are at low risk of colorectal polyps such as CRC, CRA, or CRAA. In such embodiments, low-risk individuals screened by this disclosure can avoid or delay more invasive colonoscopy. That is, this method can be implemented as a screening process as an alternative to colonoscopy. According to these embodiments, low-risk subjects can be screened at a lower cost and with greater efficiency for the healthcare system, thereby identifying subjects who are more likely to require colonoscopy or other treatment. A “low-risk subject” (as understood in the art) is a subject who has not had a past occurrence of colorectal cancer or polyps (e.g., a subject has previously undergone a colonoscopy but no colorectal cancer or polyps, e.g., CRA or CRAA, were detected) and does not have a family history of colorectal cancer or colorectal polyps. In some embodiments, low-risk subjects do not have inflammatory bowel disease such as Crohn's disease or ulcerative colitis. In various embodiments, the low-risk group is at least 45 years old, or at least 50 years old, or at least 55 years old, or at least 60 years old. In some embodiments, the low-risk group is under 75 years old, or under 70 years old, or under 65 years old. In yet another embodiment, the group is under 45 years old or under 50 years old.

[0012] In other embodiments, subjects are at high or moderate risk of CRC or colorectal polyps (as understood in the art). In such embodiments, these subjects may be monitored more frequently for the development of colorectal neoplasms, enabling early detection and treatment without frequent colonoscopy. For example, high-risk or moderate-risk subjects may have a history of CRC or colorectal polyps (e.g., CRA or CRAA) and / or a family history of CRC or colorectal polyps. In some embodiments, high-risk or moderate-risk subjects do not have inflammatory bowel disease such as Crohn's disease or ulcerative colitis. In various embodiments, the method is performed at a predetermined frequency, such as at least once a year or at least every year. In some embodiments, the method is performed at least twice a year. In various embodiments, high-risk or moderate-risk subjects are at least 45 years old, or at least 50 years old, or at least 55 years old, or at least 60 years old. In some embodiments, high-risk or medium-risk individuals are at least 65 years old, or at least 70 years old, or at least 75 years old. In yet other embodiments, the target group is under 45 years old or under 50 years old.

[0013] In various embodiments, genetic elements derived from biological samples (such as, but not limited to, fecal samples) are quantified by nucleic acid sequencing, which may include genome sequencing and / or RNA sequencing (e.g., cDNA sequencing). In various embodiments, nucleic acid sequencing may include shotgun metagenomic sequencing, targeted amplicon sequencing, and / or hybridization capture probe sequencing, among any other sequencing techniques.

[0014] Several studies have examined the gut microbiota using either 16S rDNA, shotgun metagenomics, targeted amplicon-based, or hybridization capture probe sequencing. These studies investigate fecal and mucosal-related populations, as well as different stages along the progression of adenoma and cancer. Meta-analysis of datasets from fecal or other microbiota samples identified seven CRC-rich bacterial species: Bacteroides fragilis, Fusobacterium nucleatum, Parvimonas micra, Porphyromonas assacharolytica, Prevotella intermedia, Alistipes finegoldii, and Thermoanaeroovibrio acidaminovorans. (Dai, et al., Multi-cohort analysis of colorectal cancer metagenome identified altered bacteria across populations and universal bacterial markers, Microbiome 2018;6(1):70). A separate pair of meta-analyses identified an expanded set of 29 species enriched across eight different geographical regions. Thomas et al., Metagenomic analysis of colorectal cancer datasets identifies cross-cohort microbial diagnostic signatures and a link with choline degradation, Nat Med 2019;25(4):667-678; Wirbel et al., Meta-analysis of fecal metagenomes reveals global microbial signatures that are specific for colorectal cancer, Nat Med 2019;25(4):679-689. Many studies have analyzed the human gut microbiota associated with colorectal tumors and normal adjacent tissues, leading to the identification of characteristics of enterotoxosis associated with CRC.While specific taxa vary from study to study, some common themes include the frequent identification of increased relative abundances of E. coli, Fusobacterium nucleatum, and enterotoxin-producing Bacteroides fragilis (ETBF) strains (Arthur et al., Intestinal inflammation targets cancer-inducing activity of the microbiota, Science 2012;338(6103):120-3; Kostic et al., Fusobacterium nucleatum potentiates intestinal tumorigenesis and modulates the tumor-immune microenvironment, Cell Host Microbe 2013;14(2):207-15; Wu et al., A human colonic commensal promotes colon tumorigenesis via activation of T helper type 17 T cell responses, Nat Med). 2009;15(9):1016-22. Additional taxa related to CRC have also been identified, but they are not very uniform across the study.

[0015] An important observation established by these tests is that the magnitude of the difference in relative abundance derived from tissue samples is significantly greater compared to fecal samples, and fecal samples have more complex differentially abundant taxa and often rely on AI methods to interpret. Many characteristics of feces (and other biological samples including the microbiota) present unique challenges in the identification of diagnostic biomarkers for the early detection of adenomas and carcinomas, such as high dimensionality, data sparsity, and generalizability. Despite the large amount of DNA sequence data generated in shotgun metagenomic sequencing of fecal samples, for example, the best-performing biomarkers obtained from such analyses are detected only in relatively few samples. This defines the problem of data sparsity and indicates that in high-performance diagnostics based on next-generation sequencing (NGS) data, multiple independent biomarkers are required to compensate for the low prevalence of any single biomarker within the human population.

[0016] Thus, in various embodiments, metagenomic sequencing is deep sequencing of genomic DNA isolated from fecal samples or other biological samples. In various embodiments, nucleic acid sequencing includes sequencing at least about 20,000,000 reads (i.e., raw reads per fecal sample). In various embodiments, nucleic acid sequencing includes sequencing at least about 25,000,000 reads, or at least about 30,000,000 reads, or at least about 40,000,000 reads, or at least about 50,000,000 reads, or at least about 60,000,000 reads, or at least about 75,000,000 reads, or at least about 100,000,000 reads per sample. Well-known quality control metrics can be utilized to exclude low-quality reads, which are generally less than about 15%, or less than about 10% of the raw reads. Generally, reads have less than about 5%, or less than about 4%, or less than about 3%, or less than about 2% human reads. In various embodiments, human reads are excluded from the analysis.

[0017] In various embodiments, nucleic acid sequencing includes one or more shotgun metagenomic sequencing, rDNA sequencing, and targeted nucleic acid sequencing (e.g., targeted amplicon sequencing or hybridization capture probe sequencing). In embodiments, nucleic acid sequencing may include multiple workflows, for example, rDNA sequencing, as well as one or more of shotgun sequencing and targeted nucleic acid sequencing. Sequencing can be performed using any known library preparation protocol, including the use of sample tags in a multiplex workflow. See, for example, U.S. Patents 8,603,749 and 9,453,262. These patents are incorporated herein by reference in their entirety.

[0018] In some embodiments, library creation from a DNA sample for sequencing uses total DNA isolated from feces or other biological samples (e.g., GI mucosal samples). Many kits for creating sequencing libraries from DNA are commercially available. In some embodiments, library creation includes fragmentation of the DNA, end repair, addition of sequencing adapters (e.g., by ligation or amplification), and amplification to enrich for products having adapters ligated to both ends. For example, the DNA can be fragmented such that the average fragment size ranges from 100 base pairs to about 5000 base pairs, e.g., in the range of about 250 bps to about 4000 bps, or in the range of about 500 bps to about 3000 bps, or in the range of about 1000 bps to about 3000 bps. In some embodiments, the average fragment size is less than 1000 bps, e.g., in the range of 200 - 1000 bps (e.g., 200 - 500 bps). To facilitate multiplexing, different barcoded adapters may be used with different biological samples (e.g., from different subjects). In some embodiments, the barcode can be introduced in the PCR amplification step by using different barcoded PCR primers to amplify different biological samples. The library can, in some embodiments, be subject to shotgun metagenomic sequencing.

[0019] In some embodiments, nucleic acid sequencing focuses on one or more genomic loci to enable taxonomic analysis, including rDNA analysis. Analysis of rDNA (genes encoding rRNA) may include 16S rDNA, 18S rDNA, and internal transcription spacer (ITS) sequencing. 16S and ITS sequence analysis enables taxonomic analysis of bacteria and archaea, while 18S and ITS sequence analysis enables taxonomic analysis of eukaryotes (e.g., fungi). The 16S rRNA gene contains nine variable regions scattered throughout the highly conserved 16S sequence. In some embodiments, subregions of the gene are amplified by targeted PCR for sequencing, ranging from a single variable region such as V4 or V6 to three variable regions such as V1–V3 or V3–V5. Similarly, the 18S rRNA gene contains variable regions (V1–V9) that can be used to identify at the family, order, genus, and species (and subspecies) levels, as is well known in the art. In some embodiments, gene subregions are amplified by targeted PCR for sequencing, ranging from a single variable region to multiple variable regions. ITSs are located between large and small rRNA subunit loci and may be species-specific. This polymorphism is due to the presence of tRNA genes. ITS regions can be amplified by targeted PCR for sequencing and taxonomic analysis.

[0020] In some embodiments, 16S / 18S / ITS sequences are clustered based on similarity to generate operational taxonomic units (OTUs). Representative OTU sequences can be compared to a reference database to determine the classification. In some embodiments, sequences with more than 95% identity are considered to represent the same genus, while sequences with more than 97% identity are considered to represent the same species. Methods for determining OTUs are known in the art. In some embodiments, strains or subspecies are further distinguished based on polymorphism analysis. Taxonomic analysis of 16S, 18S, and ITS DNA sequences is well known in the art. See Ze-Gang Wei et al., Comparison of Methods for Picking the Operational Taxonomic Units From Amplicon Sequences, Front. Microbiol., 24 March 2021.

[0021] By analyzing sequence reads other than rDNA and comparing them with reference microbial genomes, it is possible to infer the most likely classification methods. (Helene LCF, et al., New Insights into the Taxonomy of Bacteria in the Genomic Era and a Case Study with Rhizobia, International Journal of Microbiology Vol.2022.)

[0022] In various embodiments, sequence reads are analyzed to determine the abundance of gene function. Sequence reads may be analyzed according to the KEGG Orthology database or a similar database. The KEGG Orthology (KO) database is a database of molecular functions represented in relation to functional orthologues. Functional orthologues are manually defined in relation to the KEGG molecular network, i.e., the KEGG pathway map, BRITE hierarchy, and KEGG modules. Each node in the network, such as a box in the KEGG pathway map, is assigned a KO identifier (called a K number) as a functional orthologue defined from experimentally characterized genes and proteins in a particular organism, and these are then used to assign orthologue genes from other organisms based on sequence similarity. The resulting KO grouping may correspond to a group of highly similar sequences within a limited population of organisms, or it may be a more diverse group. Thus, in various embodiments, gene function is assigned to sequence reads (e.g., according to the KO database), and the abundance of gene function is determined for a biological sample.

[0023] In various embodiments, a target genome fragment is captured from a metagenomic library and optionally followed by amplification. For example, a nucleic acid capture probe that hybridizes to a conserved region of rDNA or a conserved region of a functional ortholog can be used. Sequence capture enables targeted enrichment of informational DNA. In conjunction with NGS, capture provides an efficient strategy for high-throughput screening of the target region. In various embodiments, the capture strategy reduces the required sequencing depth to less than approximately 25,000,000, or less than approximately 20,000,000, or less than approximately 15,000,000, or less than approximately 10,000,000, or less than approximately 5,000,000, or about 2,000,000 reads. An exemplary sequence capture protocol involves fragmenting the input DNA (e.g., by cleavage or the use of enzymes), adding a sequencing adapter to form a library molecule (e.g., by ligation or amplification using fusion primers), and incubating the library with a pool of captureable oligonucleotide probes designed to target (and hybridize to) a specific region of interest within the DNA fragment library. An exemplary captureable portion is biotin, which can be conjugated with the probe oligonucleotide. The probe / target hybrid is then captured from the library (e.g., using streptavidin-coated magnetic beads). The result is a sequenceable library with a high concentration of the targeted DNA.

[0024] In further embodiments, genetic elements can be quantified by PCR (qPCR) according to known processes. For example, genus-specific or species-specific sequences can be quantitatively amplified and detected, as can conserved sequences in functional gene elements (e.g., from rDNA in a sample). In this way, abundance profiles of genetic elements (e.g., informational features) can be constructed without a sequencing workflow.

[0025] For the detection of CRA, CRAA, and / or CRC, the number of quantified genetic elements is sufficient to provide a high-performance test (e.g., by enabling the analysis of a large number of informative features). For example, in various embodiments, the genetic elements can be analyzed for the presence of at least about 50 features, or at least about 100 features, at least about 200 features, at least about 500 features, or at least about 800 features, or at least about 1000 features (for each model). Exemplary features for detecting CRA, CRAA, and CRC are shown in Tables 3, 4, and 5, respectively. The number of features in each test does not have to be the same for each model. For example, in some embodiments, a model or "signature" for detecting CRC may include at least about 500 features, or at least about 750 features, or at least about 1000 features. In some embodiments, a model or signature for detecting CRAA and CRA may include significantly fewer, less than about 500, for example less than about 250 features (e.g., in the range of 50 to 200 features). Nevertheless, in accordance with this disclosure, it is possible to construct models with more or fewer features.

[0026] In various embodiments, the genetic elements analyzed are related to colorectal adenoma (CRA). For example, a genetic element may include one or more taxonomic or gene function features listed in Table 3. In various embodiments, a genetic element includes at least five taxonomic or gene function features listed in Table 3. In some embodiments, a genetic element includes at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or gene function features listed in Table 3. In various embodiments, a genetic element includes at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 taxonomic features listed in Table 3, and at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 gene function features listed in Table 3. As shown in Table 3, certain genetic elements exhibit different abundances (compared to controls) in samples from CRA subjects (such as fecal samples or other biological samples), and other genetic elements exhibit different prevalences (compared to controls) in fecal samples or other biological samples from CRA subjects. In some embodiments, the genetic elements include multiple genetic elements having different abundances in CRA, and multiple genetic elements having different prevalences in CRA. In some embodiments, the difference in relative abundance between diseased and non-disease samples (or vice versa) is at least about 1.1 times, or at least about 1.2 times, or at least about 1.3 times, or at least about 1.4 times, or at least about 1.5 times, or at least about 2 times. In some embodiments, the difference in prevalence between diseased and non-disease samples (or vice versa) is at least about 1.1 times, or at least about 1.2 times, or at least about 1.3 times, or at least about 1.4 times, or at least about 1.5 times, or at least about 2 times.

[0027] In various embodiments, the genetic element is related to colorectal progressive adenoma (CRAA). For example, the genetic element may include one or more taxonomic or genetic functional features listed in Table 4. In various embodiments, the genetic element includes at least five taxonomic or genetic functional features listed in Table 4. In some embodiments, the genetic element includes at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or genetic functional features listed in Table 4. In various embodiments, the genetic element includes at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 taxonomic features listed in Table 4, and at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 genetic functional features listed in Table 4. As shown in Table 4, certain genetic elements exhibit different abundances (compared to controls) in fecal samples from CRAA subjects or other biological samples, and other genetic elements exhibit different prevalences (compared to control subjects) in fecal samples from CRAA subjects or other biological samples. In some embodiments, the genetic elements include multiple genetic elements having different abundances in CRAA, and multiple genetic elements having different prevalences in CRAA. In some embodiments, the difference in relative abundance between diseased and non-disease biological samples (or vice versa) is at least about 1.1 times, or at least about 1.2 times, or at least about 1.3 times, or at least about 1.4 times, or at least about 1.5 times, or at least about 2 times. In some embodiments, the difference in prevalence between diseased and non-disease biological samples (or vice versa) is at least about 1.1 times, or at least about 1.2 times, or at least about 1.3 times, or at least about 1.4 times, or at least about 1.5 times, or at least about 2 times.

[0028] In various embodiments, the genetic element is related to colorectal cancer (CRC). For example, the genetic element may include one or more taxonomic or gene function features listed in Table 5. In various embodiments, the genetic element includes at least five taxonomic or gene function features listed in Table 5. In some embodiments, the genetic element includes at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or gene function features listed in Table 5. In various embodiments, the genetic element includes at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 taxonomic features listed in Table 5, and at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 gene function features listed in Table 5. As shown in Table 5, certain genetic elements show different abundances (compared to controls) in fecal samples or other biological samples from CRC subjects, and other genetic elements show different prevalences (compared to controls) in fecal samples or other biological samples from CRC subjects. In some embodiments, the genetic elements include multiple genetic elements having different abundances in CRC, and multiple genetic elements having different prevalences in CRC. In some embodiments, at least five, or at least ten, or at least twenty genetic elements for detecting CRC correspond to bacterial species commonly found in the oral cavity. In some embodiments, the difference in relative abundance between diseased and non-disease samples (or vice versa) is at least about 1.1 times, or at least about 1.2 times, or at least about 1.3 times, or at least about 1.4 times, or at least about 1.5 times, or at least about 2 times. In some embodiments, the difference in prevalence between diseased and non-disease samples (or vice versa) is at least about 1.1 times, or at least about 1.2 times, or at least about 1.3 times, or at least about 1.4 times, or at least about 1.5 times, or at least about 2 times.

[0029] In various embodiments, the abundance profiles of genetic elements are evaluated for signatures indicating the presence or absence of CRC, CRA, and CRAA, respectively. As disclosed herein, the microbiome profiles of CRC, CRA, and CRAA do not exhibit a linear relationship with one another, and therefore each is best evaluated using a separate model or signature.

[0030] In various embodiments, signatures are generated from a training set using a machine learning (ML) model. For example, a signature indicating the presence or absence of CRA is trained using fecal or other biological samples from a CRA cohort and biological samples from a control cohort. A signature indicating the presence or absence of CRAA is trained using fecal or other biological samples from a CRAA cohort and biological samples from a control cohort. A signature indicating the presence or absence of CRC is trained using fecal or other biological samples from a CRC cohort and biological samples from a control cohort. In all cases, the control cohort is considered to be a healthy cohort, i.e., defined by the absence of CRC, CRA, and CRAA. In some embodiments, the control cohort samples do not come from subjects with serious gastrointestinal diseases such as Crohn's disease or ulcerative colitis.

[0031] According to various aspects and embodiments of this disclosure, a training set includes at least about 50 samples, or at least about 100 samples, or at least about 150 samples, or at least about 200 samples, or at least about 500 samples, or at least about 1000 samples that are positive for CRA, CRAA, or CRC. In some embodiments, a training set includes at least about 25 non-disease or healthy controls, or at least about 50 non-disease or healthy controls, or at least about 100 non-disease or healthy controls, or at least about 500 non-disease or healthy controls, or at least about 1000 non-disease or healthy controls. Those skilled in the art will be able to assemble training sets representing disease and control samples in a manner that yields appropriate statistical power. The training set does not need to be supplied from a single test or geographical area. In some embodiments, the biological samples are supplied and / or processed in different geographical locations (e.g., at least two different countries or continents). In these embodiments, separate acquisition, processing, or sequencing may provide additional diversity in the research protocol, as well as genetic, racial, and / or environmental variability (including dietary variability) of the subjects.

[0032] Signatures can be trained using one or more machine learning algorithms. In some embodiments, at least one of the machine learning algorithms used is a supervised machine learning algorithm. In these or other embodiments, the machine learning algorithm includes one or more of unsupervised machine learning or semi-supervised machine learning. A variety of machine learning algorithms are known and can be used in accordance with this disclosure. These algorithms include, but are not limited to, one or more of the following: parametric / nonparametric distance measurement, logistic regression, support vector machines, decision trees, random forests, neural networks, probit regression, Fisher's linear discriminant, naive Bayesian classifiers, perceptrons, quadratic classifiers, kernel estimation, k-nearest neighbors, learning vector quantization, and principal component analysis. In embodiments, machine learning uses AI-enabled massively parallel computing and automated machine learning platforms and workflows (computer programs) to perform comparative machine learning modeling, optimization, testing, evaluation, and model ranking. This includes, but is not limited to, one or more of the following algorithms: deep learning, gradient boosting, neural networks, ensembles, or Blender modeling algorithms, such as gradient boosted tree classifiers, eXtreme gradient boosted tree classifiers, optical gradient boosted tree classifiers, optical gradient boosting on elastic network predictions, Keras Slim residual neural network classifiers, generalized additive models, elastic network classifiers, random forest classifiers, deep forest classifiers, mean Blender classifiers, TensorFlow multilayer perceptron classifiers, TensorFlow neural network classifiers, and Rule-Fit classifiers.

[0033] Features can be selected using any known process. In some embodiments, the signature(s) include features selected from the training cohort by ensemble ranking of feature importance, which ranks the features according to their importance in the predictive model. Alternatively or additionally, features are selected (individually) according to their statistical significance to predict CRA, CRAA, or CRC. For example, individual features with abundance or prevalence that can predict the presence or absence of CRC, CRA, or CRAA can be selected, having a p-value of ≤0.05 in the training group, or a p-value of ≤0.01, or ≤0.005, or ≤0.001 (or other selected statistical thresholds) in the training group. In some embodiments, the signature(s) include features selected from the training cohort by feature importance rank unassembly (FIRE) and by statistical estimation of association between microbial communities and phenotypes (SIAMCAT). For example, overlapping features from both processes can be selected. An iterative process for feature selection is illustrated in Figure 3.

[0034] The accuracy of a model can be assessed using standard receiver operating characteristic (ROC) curve analysis to calculate true positive, false positive, true negative, and false negative rates, overall accuracy, and area under the curve. The term "ROC" or "ROC curve" refers to the receiver operating characteristic curve. An ROC curve can be a graphical representation of the performance of a binary classifier system. In any case, an ROC curve can be constructed by plotting sensitivity against specificity at various threshold settings. Furthermore, given at least one of three parameters (e.g., sensitivity, specificity, and threshold setting), an ROC curve can determine the value or expected value of any unknown parameter. The unknown parameter can be determined using a curve fitted to the ROC curve. For example, the presence or absence or abundance of one or more features can determine the expected sensitivity and / or specificity of a test. The term "AUC" or "ROC-AUC" refers to the area under the receiver operating characteristic curve. This metric can provide a measure of the diagnostic utility of a method, taking into account both its sensitivity and specificity. The ROC-AUC can range from 0.5 to 1.0. A value close to 0.5 may indicate that the diagnostic utility of the method is limited (e.g., lower sensitivity and / or specificity), while a value close to 1.0 indicates that the diagnostic utility of the method is higher (e.g., higher sensitivity and / or specificity).

[0035] In various embodiments, the signature(s) have a sensitivity of at least about 0.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about 0.95, for classifying a sample for the presence or absence of CRA, CRAA, or CRC. In embodiments, the signature(s) have a specificity of at least about 0.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about 0.95, for classifying a sample for the presence or absence of CRA, CRAA, or CRC. For example, the signature can provide sensitivity for classifying CRA, CRAA, and CRC, respectively, having a sensitivity of at least 0.75 and a sensitivity of at least 0.75. For example, the signature is an area under the curve (AUC) for classifying a sample for the presence or absence of CRA, CRAA, or CRC, which is at least about 0.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about 0.95.

[0036] In various embodiments, if the subject is not identified as likely to have CRA, CRAA, or CRC according to the processes described herein, no further action is taken. That is, the subject is not scheduled to undergo colonoscopy or other evaluation for colorectal cancer or adenoma. If the subject is identified as likely to have one or more of CRA, CRAA, or CRC according to the processes described herein, a diagnostic or treatment plan tailored to the purpose is initiated. For example, the subject may undergo a procedure involving imaging of the colon, such as colonoscopy or CT colonography (or other scanning or imaging technique), to confirm the results, which may include removal of one or more polyps and / or taking biopsies of growths suspected to be involved in colorectal cancer. If colorectal cancer is confirmed, the subject is treated for CRC. For example, the subject may undergo one or more of the following for colorectal cancer: surgery for colorectal cancer (e.g., cancer resection including partial colectomy in some embodiments), chemotherapy, radiotherapy, and immunotherapy. Exemplary chemotherapy or immunotherapy for colorectal cancer may include one or more of the following: 5-fluorouracil (5-FU), capecitabine (XELODA) (metabolized to 5-FU by the tumor), irinotecan, leucovorin, oxaliplatin, cetuximab, panitumumab, regorafenib, bevacizumab, aflibercept, and ramucirumab. Furthermore, exemplary combination therapies include FOLFOX (5-FU, leucovorin, and oxaliplatin), FOLFIRI (leucovorin, 5-FU, and irinotecan), CAPEOX (capecitabine and oxaliplatin), FOLFOXIRI (leucovorin, 5-FU, oxaliplatin, and irinotecan), 5-FU and leucovorin or capecitabine alone, and the combination of trifluridine and tipiracil (LONSURF). In some embodiments, the subject receives an immune checkpoint inhibitor, such as an antibody or other molecule that inhibits PD-1, PD-L1, PD-L2, or cytotoxic T lymphocyte-associated protein 4 (CTLA-4).

[0037] Furthermore, radiotherapy can be used in combination with or alone with resection, chemotherapy, and immunotherapy. Types of radiotherapy include external beam radiation therapy (EBRT), internal beam radiation therapy (broad-irradiation therapy), intracavitary radiotherapy, interstitial brachytherapy, and radiation embolism.

[0038] In other embodiments, the present disclosure provides a method for constructing a genetic signature of a genetic element (i.e., an informational feature) indicating the presence of a colorectal neoplasm. The method comprises providing a training cohort of fecal or other samples from subjects confirmed to have CRA, CRAA, or CRC (or providing RNA or DNA isolated therefrom), and performing genomic nucleic acid sequencing of the DNA isolated from the fecal or other biological sample, as already described. A genetic signature is then trained to classify the samples in the presence or absence of CRA, CRAA, or CRC. In this embodiment, the genetic signature includes a microbial taxonomic classification feature and a microbial gene functional feature, as described above and illustrated in Tables 3-5. In this embodiment, the method can use any sample suitable for evaluating the microbiome of a cohort, including not only fecal samples but also other biological samples such as human biological fluids (e.g., blood, serum, plasma, urine, saliva), tissues, mucosa, and cell samples.

[0039] As described, nucleic acid sequencing may include one or more shotgun metagenomic sequencing, rDNA sequencing, and targeted nucleic acid sequencing, including targeted amplicon sequencing and hybridization capture probe sequencing. Any sequencing technique may be used. Genetic elements in the sample are assigned to a reference genome for taxonomic classification (which may include rDNA analysis), and / or genetic elements are assigned to gene function (as already described). Taxonomic and gene function features may also be analyzed at the protein level using known methods.

[0040] For example, in fecal or other biological samples from CRA subjects, microbial taxonomic features and microbial gene function features are selected that have different abundances or different prevalences compared to control subjects. In some embodiments, the features include at least five taxonomic features and / or gene function features, optionally listed in Table 3. The features include at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic features and / or gene function features, optionally listed in Table 3. In some embodiments, the features include at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 taxonomic features, optionally listed in Table 3, and at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 gene function features, optionally listed in Table 3.

[0041] In some embodiments, microbial taxonomic features and microbial gene function features are selected from CRAA subjects in feces or other biological samples that have different abundances or different prevalences compared to control subjects. In some embodiments, the features include at least five taxonomic features and / or gene function features, optionally listed in Table 4. The features include at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic features and / or gene function features, optionally listed in Table 4. In some embodiments, the features include at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 taxonomic features, optionally listed in Table 4, and at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 gene function features, optionally listed in Table 4.

[0042] In some embodiments, microbial taxonomic features and microbial gene function features are selected from fecal or other biological samples from CRC subjects that have different abundances or different prevalences compared to control subjects. In some embodiments, the features include at least five taxonomic features and / or gene function features, optionally listed in Table 5. The features include at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic features and / or gene function features, optionally listed in Table 5. In some embodiments, the features include at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 taxonomic features, optionally listed in Table 5, and at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 gene function features, optionally listed in Table 5.

[0043] In various embodiments, at least three gene signatures are trained to classify samples based on the presence or absence of CRA, the presence or absence of CRAA, and the presence or absence of CRC. For example, the signature for classifying samples based on the presence or absence of CRA is trained by machine learning using fecal or other biological samples from a CRA cohort and samples from a control cohort. The signature for classifying samples based on the presence or absence of CRC is trained by machine learning using fecal or other biological samples from a CRC cohort and samples from a control cohort. The signature for classifying samples based on the presence or absence of CRAA is trained by machine learning using fecal or other biological samples from a CRAA cohort and samples from a control cohort. The machine learning may be as already described and may include supervised machine learning, unsupervised machine learning, semi-supervised machine learning, or a combination thereof.

[0044] As already explained, features can be selected. For example, a signature(s) may include features selected from a training cohort by ensemble ranking of feature importance. Alternatively, or additionally, features are selected (individually) according to their statistical significance to predict CRA, CRAA, or CRC. For example, individual features may be selected based on their abundance or prevalence that significantly predicts the presence or absence of CRC, CRA, or CRAA, as demonstrated by a p-value of ≤0.05 in the training group, or ≤0.01 in the training group, or ≤0.005, or ≤0.001 (or any selected statistical threshold). In some embodiments, a signature includes features selected from a training cohort by feature importance rank unassembly (FIRE) and by statistical estimation of associations between microbial communities and phenotypes (SIAMCAT).

[0045] In various embodiments, the created signature(s) have a sensitivity of at least about 0.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about 0.95, for classifying a sample for the presence or absence of CRA, CRAA, or CRC. In various embodiments, the created signature(s) have a specificity of at least about 0.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about 0.95, for classifying a sample for the presence or absence of CRA, CRAA, or CRC. For example, a signature can provide sensitivity for classifying CRA, CRAA, and CRC, respectively, with a sensitivity of at least 0.75. For example, the created signature has an area under the curve (AUC) of at least approximately 0.70, at least approximately 0.75, at least approximately 0.80, at least approximately 0.90, or at least approximately 0.95 for classifying a sample for the presence or absence of CRA, CRAA, or CRC.

[0046] In other embodiments, the Disclosure provides a method for creating genetic signatures of fecal (or other biological sample) genetic elements indicating colon disorders (including, but not limited to, CRA, CRAA, and CRC). In various embodiments, the method includes providing a training cohort (or DNA isolated therefrom) of fecal or other biological samples from subjects and control subjects confirmed to have colon disorders, and performing genomic nucleic acid sequencing of the DNA isolated from the samples (as already described). Genetic signatures are then trained to classify the samples for the presence of colon disorders and for the absence of colon disorders. According to this embodiment, the genetic signatures include features selected from the training cohort by ensemble ranking of feature importance and by the statistical significance of individual features.

[0047] For example, features may include features selected from the training cohort by ensemble ranking of feature importance and according to their statistical significance (individually) for predicting colonic disorders. For example, individual features with abundance or prevalence that can predict the presence or absence of colonic disorders can be selected, having a p-value of 0.05 or less in the training group, or a p-value of 0.01 or less, or a p-value of 0.005 or less, or a p-value of 0.001 or less (or other selected statistical thresholds) in the training group. In some embodiments, the signature includes features selected from the training cohort by feature importance rank unassembly (FIRE) and by statistical estimation of associations between microbial communities and phenotypes (SIAMCAT).

[0048] In various embodiments, the colon disorder is selected from Crohn's disease, ulcerative colitis, irritable bowel syndrome (IBS), bifurcationitis, colorectal adenoma (CRA), progressive colorectal adenoma (CRAA), and colorectal cancer (CRC).

[0049] In various embodiments, the feature includes the above-described taxonomic and microbial gene function features. Exemplary taxonomic and gene function features are shown in Tables 3, 4, and 5 for CRA, CRAA, and CRC, respectively. For example, in stool or other samples from colonic disorder subjects, taxonomic and microbial gene function features are selected that have different abundances or different prevalences compared to control subjects. In some embodiments, the feature includes at least five taxonomic and / or gene function features. In some embodiments, the feature includes at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or gene function features. In embodiments, the feature includes at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 taxonomic features, and at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 gene function features.

[0050] In various embodiments, the signature(s) are trained using one or more machine learning algorithms, including supervised machine learning, unsupervised machine learning, semi-supervised machine learning, or a combination thereof, as already described.

[0051] In various embodiments, the generated signature(s) have a sensitivity of at least about 0.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about 0.95 for classifying a sample for the presence or absence of colonic dysfunction. In various embodiments, the generated signature(s) have a specificity of at least about 0.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about 0.95 for classifying a sample for the presence or absence of colonic dysfunction. For example, the created signature(s) have an area under the curve (AUC) of at least about 0.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about 0.95 for classifying a sample for the presence or absence of colonic dysfunction.

[0052] As used herein, the term "approximately" means ±10% of the associated number unless otherwise specified in the context.

[0053] Other aspects and embodiments of the present invention are evident from the following non-limiting examples. [Examples]

[0054] Example 1: Materials and Method Data normalization and visualization using PCoA. Individual taxonomic numbers were normalized using the M-value weighted trimmed mean (TMM) with the edgeR package, and then converted to log-CPM using Voom, implemented in the "limma" package in R version 4.2.1. (Robinson et al., edgeR: a Bioconductor package for differential expression analysis of digital gene expression data, Bioinformatics 2010;26(1):139-40; Law et al., voom: Precision weights unlock linear model analysis tools for RNA-seq read counts, Genome Biol 2014;15(2):R29). Next, the data were logarithmically transformed and further normalized using supervised normalization (SNM) to remove significant batch effects between projects while preserving biological differences between disease classes. Poore et al., Microbiome analyses of blood and tissues suggest cancer diagnostic approach, Nature 2020;579(7800):567-574. Supervised normalization (SNM) of microarrays was implemented in the "snm" package from R. Mecham et al., Supervised normalization of microarrays, Bioinformatics 2010;26(10):1308-15. The effect of supervised normalization was visualized using principal coordinate analysis (PCoA). PCoA was performed using Euclidean distance for both relative abundance values ​​and SNM-converted count tables, including taxa of all ranks (from kingdom to species). Differences in microbial composition between projects and disease classes were separately evaluated using the Adonis function from the vegan package (available on the World Wide Web at CRAN.R-project.org / kgage=vegan).The mean distance of samples from the centroid in the PCoA plot was compared across projects using the Kruskal-Wallis test.

[0055] Machine learning. Models were created using an automated machine learning platform called DataRobot (DR, available at www.dacerob.com on the World Wide Web). A Python script was developed for automated data submission to DR, enabling the development of models for multiple datasets. For prediction of an external test dataset, the best model among all developed models was selected based on the largest Area Under Curve (AUC) value. "Blender models," obtained using several machine learning algorithms (combining the predictions of two or more models), were not used here. For classification purposes for each target, the sample set was randomly split into a training set (80% of the samples) and a test set (20% of the samples). The training set was used to develop a set of high-performance predictive models in DR (using more than 10 different machine learning algorithms for classification, including the eXtreme gradient boosted tree classifier, the Keras Slim residual neural network classifier with training schedules, the elastic network classifier, and the optical gradient boosted tree classifier with early stopping capabilities). Using the developed models, disease status was predicted for the remaining 20% ​​of samples (data not used in training) to define the top-performing model (best external test AUC). The model's performance against test set parameters is ultimately complemented by the sensitivity, specificity, and accuracy of external tests.

[0056] DataRobot feature lists. Feature lists control the subset of features that DataRobot uses to build models. DataRobot automatically creates several feature lists per project, including two main lists: informational features and DataRobot (DR) reduced features. Informational features are all features (usually all features) that provide potentially useful information for modeling. DR reduced features are a subset of features selected based on the feature impact calculation of the best model. The DR reduced features list consists of features that provide 95% of the model's accumulated impact. While computational analysis does not require limiting the number of features, practical considerations for further laboratory analysis (by qPCR) recommend using models with fewer features. Typically, the informational features list contains almost all features in the dataset (1000-2000 for classification annotations, 5000-10000 for feature annotations), and the DR reduced features list contains 100 or fewer features; therefore, for comparison purposes, we used a model built with DR reduced features.

[0057] FIRE feature selection. The feature reduction and selection method "Feature Importance Rank Ensembling" (FIRE, available on the World Wide Web at www.datarobot.com / blog / using-feature-importance-rank-ensembling-fire-for-advanced-feature-selection / ) was used. In this method, features are derived from multiple diverse predictive models built by DR. By default, DR sorts the models by a selected criterion, e.g., external test AUC. The median rank of each feature is calculated by aggregating the respective ranks of several top models (the number of top models to consider was empirically chosen to be 5).

[0058] Feature importance is calculated in DR using an algorithm that measures the informational content of the variables. This calculation is performed individually for each feature in the dataset. In some embodiments, the FIRE procedure includes the following steps: (a) calculating the feature importance of the top models (e.g., 3, 4, 5, 6, 7, 8, 9, 10) (determined by an external test AUC), (b) obtaining the feature ranking, (c) calculating the median rank of each feature, (d) sorting the aggregated list by the calculated median rank, (e) defining a threshold number of features to select, (f) defining a feature list based on the newly selected features, and (g) removing redundant features selected by two or more models. By sorting the aggregated list by median rank, a list of ranked feature importance is derived.

[0059] Since the optimal number of features was unknown, the first loop iterated through and tested several thresholds in large increments (800, 700, 600, 500, 400, 300, 200, 100, 90, 80, 70, 60, 50, 40, 30), and the second loop iterated through and tested several thresholds in small increments around the first threshold found (e.g., if the best threshold from the first loop was equal to 80, then 85, 84, 83, 82, 81, 79, 78, 77, 76, 75). Finally, the threshold that provided the best external test AUC was obtained. In some analyzed datasets, the total number of features was between 800 and 900, so the maximum number of FIRE features was considered to be 800. The inventors implemented FIRE in a multi-stage process with the aim of improving the predictive accuracy of disease classes and, furthermore, leveraging the ensemble nature of FIRE, establishing feature selections whose basis is not tied to a single model and its associated biases, thereby improving generalizability.

[0060] SIAMCAT feature selection. Another feature selection method is based on statistical estimation of associations between microbial communities and phenotypes, referred to as Statistical Estimation of Associations between Microbiota and Host Phenotypes (SIAMCAT) version 2.1.0. Wirbel et al., Microbiome meta-analysis and cross-disease comparison enabled by the SIAMCAT machine learning toolbox, Genome Biol 2021;22(1):93. SIAMCAT is part of a set of computational microbiome analysis tools developed at EMBL. SIAMCAT provides sign for feature abundances (using p-values ​​and adjusted p-values), thereby indicating the sign of changes in abundance between normal and diseased states. Furthermore, this analysis used an adjusted p-value (adj_pval) of 0.001. Raw data were preprocessed, and taxonomic profiles were created using bioBakery 3 pipeline version 3.0.0a7.

[0061] Combining FIRE and SIAMCAT. Since FIRE and SIAMCAT feature lists are selected based on different criteria (ensemble ranking for the former, statistical significance for the latter), we also decided to test lists that combine FIRE and SIAMCAT lists to create a new FIRE_SIAMCAT feature list.

[0062] Example 2: Analysis of publicly available datasets To evaluate the accuracy and feasibility of developing a non-invasive diagnostic fecal test for early and advanced precancerous adenomas and late-stage carcinomas, the inventors incorporated data generated from 13 trials conducted by laboratories in eight countries and analyzed fecal samples using shotgun metagenomic sequencing. Wirbel et al.,Meta-analysis of fecal metagenomes reveals global microbial signatures that are specific for colorectal cancer,Nat Med 2019;25(4):679-689 Thomas et al.,of colorectal cancer datasets identifies cross-cohort microbial diagnostic signatures and a link with choline degradation,Nat Med 2019;25(4):667-678;Yu et al. al.,Metagenomic analysis of faecal microbiome as a tool towards targeted non-invasive biomarkers for colorectal cancer,Gut 2017;66(1):70-78;Feng et al.,Gut microbiome development along the colorectal adenoma-carcinoma sequence,Nat Commun 2015;6:6528;Zeller et al.,Potential of fecal microbiota for early-stage detection of colorectal cancer,Mol Syst Biol 2014;10(11):766;Yachida et al.,Metagenomic and metabolomic analyzes reveal distinct stage-specific phenotypes of the gut microbiota in colorectal cancer,Nat Med 2019;25(6):968-976;Gao et al.,Alterations,Interactions,and Diagnostic Potential of Gut Bacteria and Viruses in Colorectal Cancer,Front Cell Infect Microbiol 2021;11:657867;Gupta et al.,Association of Flavonifractor plautii,a Flavonoid-Degrading Bacterium,with the Gut Microbiome of Colorectal Cancer Patients in India,mSystems 2019 Nov 12;4(6):e00438-19;Hale et al.,Distinct microbes,metabolites,and ecologies define the microbiome in deficient and proficient mismatch repair colorectal cancers,Genome Med 2018;10(1):78;Vogtmann et al.,Colorectal Cancer and the Human Gut Microbiome:Reproducibility with Whole-Genome Shotgun Sequencing,PLoS One 2016;11(5):e0155362;Spanogiannopoulos et al.,Host and gut bacteria share metabolic pathways for anti-cancer drug metabolism,Nat Microbiol 2022;7(10):1605-1620;Hannigan et al.Diagnostic Potential and Interactive Dynamics of the Colorectal Cancer Virome, mBio 2018;9(6):e02248-18. In total, our analysis included sequence data generated from 1,705 subjects, including 703 healthy controls, 196 cases of precancerous colorectal adenoma (CRA), 48 cases of advanced precancerous colorectal adenoma (CRAA), and 765 cases of colorectal cancer (CRC) confirmed by colonoscopy. Descriptive statistics for the cohorts from each study are shown in Table 1.

[0063] To analyze the publicly available datasets, we first performed data normalization using the weighted trimmed mean (TMM) and logarithm per million (TMM-Voom data) of the M-values, and then further normalized using supervised normalization (SNM) as described in Mecham et al., Supervised normalization of microarrays, Bioinformatics 2010;26(10):1308-15. The effect of supervised normalization was visualized using principal coordinate analysis (PCoA). PCoA was performed using Euclidean distance for both TMM-Voom and TMM-Voom-SNM conversion count tables, which include taxa of all ranks (from kingdom to species). Differences in microbial composition between projects and disease classes were evaluated separately.

[0064] PCoA plots show variations that can be explained by supervised normalization, but are not necessarily by the original project, in R 2The variability was significantly reduced from 10.3% to 0.0975% (Figures 1A-B). Only 1.5% of the variability was explained by disease class using TMM-Voom normalized data (Figure 1C). After supervised normalization, this variability decreased to 0.896%, but the differences between disease types remained significant (Figure 1D). To identify projects and samples representing potential outliers, the distance to the centroid within each project was determined. These distances were similar across projects for both TMM-Voom and TMM-Voom-SNM data (Figures 12A, B). Although not ideal, project-specific variability was fully expected. Rather than attempting to remove trials based on specific criteria, the inventors chose to face the challenge of performance optimization of these datasets. Furthermore, it is difficult to distinguish between trial effects and country-specific effects. Therefore, the possibility that biomarkers associated with adenomas and carcinomas are susceptible to country or region-specific influences remains unresolved. These results highlight important challenges associated with meta-analysis of gut microbiota data. These analyses, in particular, showed that while the β-variances between studies were similar, there was, as expected, significant inter-sample variability within each study.

[0065] To further investigate features associated with meta-analysis, namely trial variability and population variability, models were trained using data from individual trials representing different population cohorts, and the extent to which each model predicted all other trials was determined (Figure 2). In most cases, models trained on data from specific trials performed well on their own, although not necessarily yielding the best results. Some trials used for training generated relatively high AUCs across all or most trials in the trial set. This result may be a way to measure the generalizability of features derived from a particular trial population. Other trials used as the training set predicted one or a few trials with high AUCs, but showed greater variability overall. Finally, some trials did not perform relatively well across most or all trials. The AUC variability observed in this analysis highlights one of the major challenges associated with meta-analysis of microbiome profiles. Numerous reasons can explain the variability between trials, including sampling differences, methodological variability, geographical influences, and cohort demographics. This analysis, among other things, showed that some trials generated high predictive power for CRC samples in their own studies and several additional trials. No single trial provided a good prediction of CRC across all trials. It should be noted that in such analyses, even the highest-quality trials yielded results comparable to the lowest-quality trials. Factors contributing to trial quality are variable and difficult to define.

[0066] Example 3: Feature Generation In developing a fecal microbiome diagnostic analysis pipeline, the inventors focused on evaluating two relevant data features. First, the relative abundance of taxonomic features enumerated through the bioBakery pipeline. See, for example, McIver et al., bioBakery: a meta-omic analysis environment. Bioinformatics. 2018;34(7):1235-1237. Second, the inclusion of gene features obtained from shotgun metagenomic sequencing analysis was investigated. This report evaluates annotations of KO gene functions. The KEGG Ortholog (KO) group is a database of molecular functions represented in relation to functional orthologues. See Kanehisa et al., KEGG: new perspectives on genomes, pathways, diseases and drugs, Nucleic Acids Res 2017;45(D1):D353-D361. Importantly, gene features often have high relative abundances because each represents the sum of orthogonal genes within the entire community. This characteristic is considered potentially beneficial compared to taxonomic features, which often suffer from the problem of sparseness. While we do not wish to be constrained by theory, the inventors hypothesized that genetic features could positively contribute to predictive performance and function as complementary features to those derived from classification methods, and they tested this hypothesis.

[0067] Example 4: Feature Processing and Selection We implemented a strategy to evaluate a wide variety of feature reduction methods and compare their overall impact on predictive accuracy. Each of these has its own specific advantages and disadvantages. Feature selection schemes based on filtering data to remove low-prevalence features are risky from a fecal microbiome perspective, as many of the best diagnostic features (species) are often low in abundance and prevalence. This knowledge means that feature selection must be performed in a way that preserves their informational features. This fact also indicates that the best-performing models are likely to require more features for optimal accuracy. Here, we evaluate a feature reduction method called Feature Importance Rank Unassembly (FIRE). Specific to this method, the highest-ranked features are derived from multiple diverse models. We chose to identify the five best-performing models using our ensemble procedure (by external test AUC). To define a novel workflow, we added another statistically-based feature selection method called SIAMCAT (described below) (Figure 3). One key finding based on the FIRE implementation is that the best-performing model differs depending on the disease class, highlighting that there is no single model or set of models that is optimal for distinguishing between health and disease (CRA, CRAA, CRC). Therefore, FIRE is run separately for CRA, CRAA, and CRC samples, respectively, using feature selection and training to achieve target-specific model optimization, which then achieves optimal disease-class-specific diagnostic performance.

[0068] As an additional layer to the ensemble, the inventors process taxonomic and genetic features using a tool known as SIAMCAT. This allows for visualization of different abundances, prevalences, and AUCs of features, and features are ranked based on statistical significance. SIAMCAT-based feature selection allows for setting arbitrary significance cutoffs. In this study, features with a corrected p-value p<0.001 were used. Unsurprisingly, features generated by FIRE and SIAMCAT partially overlap. The inventors' pipeline combines a relatively large number of features generated by FIRE and SIAMCAT, where redundancy is eliminated. The unique features generated after combining FIRE and SIAMCAT represent a new list of features used to classify samples into healthy or diseased classes. In practice, any number of features can be selected, but the optimal number of features must be determined empirically (see below). In this embodiment, FIRE is run on a mixture of taxonomic and genetic features, and the results of the top 5 models are shown.

[0069] Example 5: FIRE Feature Selection To establish the optimal number of features, FIRE was iteratively run, starting with a large number of features, such as 800-1000, to establish baseline performance based on external test AUC. These analyses were performed separately for each disease target, as well as for taxonomic and genetic features (Table 2). When taxonomic and functional genetic features were evaluated separately, no patterns were observed across disease classes. For CRC, 800 genetic features and 400 taxonomic features provided the best performance. This was strongly contrasted by CRAA and CRA analyses. For CRAA, the optimal number of genetic features was significantly less (40), as was the number of taxonomic features (100). For CRA, 40 genetic features and 70 taxonomic features were observed to be optimal for performance.

[0070] These results surprisingly showed that the number of optimal features for CRC, both taxonomic and genetic, was significantly greater than the number determined for CRAA and CRA (Table 2). The reason(s) for this are unclear, but it may reflect that colorectal tumors have the greatest impact on the colonic microbiota and their encoded functions, thereby generating a broader range of discriminative biomarkers. For both CRAA and CRA, a relatively small number of genetic features (40) were determined to be optimal. While we do not wish to be bound by theory, we speculate that this may reflect a relative lack of discriminative genes and taxonomic biomarkers in the early stages of the disease.

[0071] Table 3 shows a list of the top 300 characteristics of colorectal adenomas (CRAs) that indicate the classification of the identified genera. Table 3 also shows the ratio of change in relative abundance compared to CRA-negative samples, the prevalence shift, and the rank indicating the weight or importance of the change. The prevalence shift value between the two classes is positive when the prevalence of CRA is high and negative when the prevalence of the control group is high. In some cases, the ratio change is zero, and therefore the prevalence shift column indicates the prevalence in CRA.

[0072] Table 4 shows a list of the top 300 characteristics of colorectal progressive adenoma (CRAA) indicating the classification of the identified genera. Table 4 also shows the ratio of change in relative abundance compared to CRAA-negative samples, the prevalence shift, and the rank indicating the weight or importance of the change. The prevalence shift value between the two classes is positive when the prevalence of CRAA is high and negative when the prevalence of the control group is high. In some cases, the ratio change is zero, and therefore the prevalence shift column indicates the prevalence in CRAA.

[0073] Table 5 shows a list of the top 300 characteristics of colorectal cancer (CRC) indicating the classification of identified genera. Table 5 also shows the ratio of change in relative abundance compared to CRC-negative samples, the prevalence shift, and the rank indicating the weight or importance of the change. The prevalence shift value between the two classes is positive when the prevalence of CRC is high and negative when the prevalence of the control group is high. In some cases, the ratio change is zero, and therefore the prevalence shift column indicates the prevalence in CRC.

[0074] The taxonomic characteristics of CRCs are unique in that they are highly concentrated in diseases and species that are normally present in the oral cavity and are often overexpressed. While most studies investigating these microorganisms focus on their behavior in the oral cavity rather than the gut, accumulating evidence suggests that these taxa are pathogenic organisms that can cause or contribute to disease in a variety of contexts.

[0075] Several interesting observations can be made from this output. Perhaps most importantly, the top 24 features reflect bacterial species commonly found in the oral cavity. While the results clearly indicate the presence of these organisms in a healthy gut microbiota, each of these taxa shows an increase in relative abundance (multiplier change) and prevalence (frequency of non-zero measurements) in the CRC compared to controls. All taxonomic features except three show an increase in relative abundance in the CRC compared to healthy control subjects. Two of these features, Roseburia intestinalis and members of the genus Anaerostipes, are part of a normal symbiotic gut microbiota. The specific reasons for their decreased relative abundance and increased prevalence are unknown. The third example is unexpected, relating to the known oral bacterium Streptococcus salivalus. The possible reasons for its decreased relative abundance and increased prevalence are unknown.

[0076] The overrepresentation of oral microorganisms in CRC fecal samples is consistent with the idea that the tumor microenvironment co-selects these oral species through unknown adaptive advantages lacking in healthy individuals and / or defense mechanisms rendered ineffective in CRC. While the factors promoting this adaptive advantage may be complex, one factor that could explain these outcomes is thought to be a metabolic shift occurring in colon cancer epithelium associated with the transition from health and adenoma to carcinoma. Specifically, intestinal oxygen consumption resulting from the oxidative metabolism of butyrate for energy is replaced by the fermentation of lactic acid, which does not consume oxygen. One key consequence of this metabolic shift is an increase in oxygen tension in the tumor microenvironment. This increase in oxygen content may be sufficient, or at least one contributing factor, to positively select the observed aerobic oral species.

[0077] Example 6: SIAMCAT Feature Selection An independent method for feature selection called SIAMCAT was investigated. This method calculates and displays the relative abundance of each feature in all analyzed samples, the statistical significance of features represented differentially within the dataset, the observed multiplier change between healthy controls and each disease class, the change in prevalence, and the feature AUC (Figure 4). In this example, feature importance is related to CRC using only taxonomic features.

[0078] While the output from SIAMCAT reveals only slight scaling changes for some taxonomic features, significant shifts in prevalence are observed within these same species. Unsurprisingly, many of the top features generated by FIRE and SIAMCAT overlap, although some features are unique to one method or the other. This provides a reasonable basis for evaluating the relative impact of combining non-redundant features on classification performance and assessing the relative strength of our approach's performance. We evaluated possible stepwise improvements in our approach by generating AUC using the best models from our machine learning algorithms to establish a baseline comparison with FIRE and SIAMCAT individually and in combination (Figure 5). While FIRE feature selection generally outperformed SIAMCAT, the SIAMCAT values ​​are evident in instances where combining FIRE and SIAMCAT analysis features and results yielded the best performance (Figure 5). Furthermore, in all analyses including non-redundant features derived from FIRE and SIAMCAT, SIAMCAT features were consistently among the most important features positively contributing to the external test AUC. Applying SIAMCAT to control and CRC samples generated a list of taxonomic feature importance that closely matched taxa reported by several studies.

[0079] It was observed that the best performance depended on the disease target. For CRA, the best performance was achieved when combining taxonomic and KO features selected by combining FIRE-SIAMCAT. This approach increased the external AUC by approximately 8% (baseline AUC = 0.80 vs. 0.87). The analysis of CRAA performance was somewhat more complex. All feature selection strategies performed best when using FIRE, and there was little difference in performance when using taxonomic features alone or in combination with genetic features. The results of CRC showed a similar trend to CRAA. We observed that the combination of taxonomic and genetic features performed better than taxonomic features alone, and furthermore, taxonomic features alone performed better than genetic features alone. Features selected by FIRE resulted in a 3% increase in external AUC (baseline AUC = 0.94 vs. 0.97). The results of CRC showed that the combination of taxonomic and KO features performed better than taxonomic features alone, and furthermore, taxonomic features alone performed better than KO features alone. The best performance was observed from the features selected by FIRE, resulting in a slight 2% increase in the external AUC (baseline AUC = 0.80 vs. 0.82). These results highlight the challenges associated with optimizing model performance. Since no single model optimally performs for all three disease classes, the microbiome and microbial community associated with each disease class requires the definition of separate computational workflows.

[0080] To gain biological insights into the features that most significantly contribute to distinguishing disease class predictions, or potential common features shared across disease classes, we assessed feature overlap across disease classes (generated from a combination of FIRE and SIAMCAT) (Figure 6). It is noteworthy that when comparing the top 20 key features for each disease class (Venn diagram, lower right), there was no overlap in either taxonomic or genetic features. Another interesting difference in this comparison is that 60% of CRC features were taxonomic, significantly more than observed in CRA (40%) and CRAA (25%). When examining the top 50 features (Venn diagram, lower center), these differences persist but dissipate. Within the top 50 features, slight overlap in features across disease classes begins to be observed. A comparison of the top 100 features shows that the proportion of taxonomic features becomes perfectly even across disease classes. As more features are compared, the proportion of genetic features compared to taxonomic features continues to increase, and an increase in overlap between disease classes is observed. Examining 800 features reveals potentially interesting biological aspects of the microbiome in association with colon neoplasms. It is noteworthy that, among overlapping features, genetic features are more dominant than taxonomic features. Particularly interesting, while there were no overlapping taxonomic features across the three disease classes, 48 ​​genetic features (2.4%) were shared. This imbalance is also evident in all pairwise comparisons of overlapping features, with taxonomic features accounting for approximately 9–14% of the overlapping features. Considering the similarity in the frequency of common taxonomic features across disease classes, the number of common genetic features is significantly higher between CRC and CRA compared to any other pairwise relationship.

[0081] As a next step, the inventors refined these comparisons to determine the true biological similarity of common taxonomic and genetic features by considering the direction of change in both feature types, i.e., increase or decrease in the disease (Figure 7). When the top 800 features of each disease class are considered binary (top left; all increase > 0; or all decrease < 0 compared to healthy controls), most overlapping features occur between CRA and CRC samples, but some overlap remains between CRAA and CRC samples. Further analysis of these relationships in different permutations (top right) shows the relationship between CRA and CRAA samples, highlighting that there are few common features showing the same direction of change. These figures also highlight the large number of overlapping features shared between CRC and CRAA, and that the direction of change is opposite. The comparison between CRA and CRC samples is visualized by comparing the Venn diagrams in the bottom left and bottom right. While the number of common features considering the direction of change remains large, the remaining figures (bottom left CRC FC<0, CRA FC>0, CRAA FC<0, and bottom right CRC FC>0, CRA FC<0, CRAA FC>0) show that most of the common features of CRA and CRAA represent examples of change in the opposite direction.

[0082] Thus, many common features were found between CRA and CRC, and the direction of their changes was also frequently the same (Figure 7). This result is the most surprising and difficult to explain, but it suggests that the intestinal microenvironment and selective pressure are similar in CRA and CRC, but different in CRAA. Although not bound by theory, we hypothesize that the high proportion of common genetic features compared to taxonomic features may reflect functional redundancy between related and even distantly related taxa, which are essentially interchangeable within the community thanks to common genes encoded in their respective genomes. The high proportion of genetic features in the expanded list of feature importance suggests the need for a detailed analysis of functional attributes that are overexpressed and underexpressed in health and disease.

[0083] Example 7: Model validation: Analysis of taxonomic features. To assess whether the taxonomic features among the top 800 in each class behave consistently and / or are biased towards specific phylogenetic groups, we analyzed key features at the class and family levels. Of the top 800 features, 66 represented taxa that were over- or under-expressed in the CRA, 86 in the CRAA, and 79 in the CRC. In total, the feature importance list included taxa from 12 classes (Figure 8). Note that these features were considered discriminative based on the AI ​​model used, although they did not necessarily achieve statistical significance in the comparison. Clostridium had the most features that were under-expressed in the CRA. This class was strongly over-expressed in both the CRAA and CRC.

[0084] The next most dominant class among the top features is Bacteroidium. Of the 15 taxa of the top features in CRA, 12 showed a positive scaling change, while 9 taxa of CRC showed a positive scaling change. In contrast, few taxa in this class were distinctive in CRAA and were mainly underrepresented. Two classes (Tissierellia and Fusobacteriia) were overrepresented and were exclusively included in the high importance list of CRC, but not present in CRA or CRAA. The class that best represented CRAA was Methanobacteria, which was uniquely overrepresented in CRAA but not in CRA or CRC. Two additional classes, Actinobacteria and Coriobacteria, were very overrepresented in CRAA but were either absent or underrepresented in the feature importance lists of CRA and CRC. The consistency of these results is extremely remarkable considering the vast phylogenetic space and evolutionary distance within bacterial classes.

[0085] Next, these results were examined at a higher resolution at the family level. The top features of all disease classes were attributed to 42 different bacteriological families. As observed during class analysis, the distribution of important features for each disease class is biased towards specific families. Examination of families belonging to the three classes Methanobacteria, Actinobacteria, and Coriobacteria shows that the underlying families provide significant strength in distinguishing CRAA from other disease classes and healthy subjects (Figure 9). For CRA, two species within Barnesiellaceae were uniquely underexpressed, while single species in Selenomonadaceae, Sutterellaceae, and Neiseriaceae were all overexpressed. In the case of CRAA, two species within Acidaminococceae were uniquely underexpressed in CRAA. The overexpression of two species in both Methanobactericeae and Propionibacteriaceae was exclusive to the list of importance of CRAA features. Two species within Tannerellaceae uniquely overexpressed in CRA and underexpressed in CRAA. Species within other families, including Actinomycetaceae and Eggerthellaceae, overexpressed in CRAA and underexpressed in CRA. Taxonomy belonging to Lachnospiraceae, which is not exclusively overexpressed in CRAA, includes 10 underlying taxa that distinguish CRAA from other disease classes. The overexpression of two families (Peptoniphilaceae and Fusobacteriaceae), each possessing two species, is specific to CRC samples. Two families, including Clostridiaceae (3 taxa) and Erysipelotrichaceae (2 taxa), distinguish CRC from other disease classes and healthy controls.

[0086] Among the most important taxonomic features of CRA, differences in the expression of six Bacteroides spp. were observed. Several genera, including Prevotella spp., Parabacteroides spp., and Veillonella spp., were represented by two or more species, all of which were overexpressed compared to the healthy control sample and Eubacterium spp. and Roseburia spp. Both Eubacterium spp. and Roseburia spp. were underexpressed. All other differentially rich genera were represented by a single species. Key features of CRAA included overexpression of genera compared to the healthy control sample, including four species of Actinomyces, three species of Collinsiella, two species of Enorma, two species of Lactobacillus, two species of Dorea, and two species of Coprococcus. Two species of Bacteroides were underexpressed compared to the healthy control sample. Two Alistipes spp. showed different representations compared to the control sample. Important features of CRC included two species of Actinomyces, three species of Porphorymonas, four species of Prevotella, two species of Peptostreptococcus, and three species of Fusobacterium. Three species of Bacteroides and two species of Veillonella were overexpressed in CRC compared to healthy control samples. Some of these features did not show a significant metric change compared to the control, but they did show a significant change in prevalence. All other genera were represented by a single species. Interestingly, the taxa in the importance list of CRC features did not show underexpression.

[0087] As shown in Figure 10, the following taxonomic features increased in CRA: Actinomyces odontolyticus, Bifidobacterium adolescentis, Bifidobacterium longum, Enterorhabdus caecimuris, Gordonibacter pamelaeae, Bacteroides eggerthii, Bacteroides intestinalis, Bacteroides nordii, Bacteroides plebeius, Bacteroides salyersiae, Bacteroides stercoris, Barnesiella intestinihomonis, Butyricimonas virosa, Prevotella copri, Prevotella stercorea, Alistipes shahii, Parabacteroides gordonii, Parabacteroides goldsteinii, Gemella sanguinis, Streptococcus thermophilus, Eubacterium ventriosum, Anaerostipes hadrus, Blautia obeum, Blautia wexlerae, Dorea formicigenerans, Dorea longicatena, Fusicatenibacter saccharivorans, Eubacterium rectale, Roseburia sp. CAG309, Roseburia sp. CAG431, Oscillibacter sp. CAG 241, Faecalibacterium prausnitzii, Clostridium leptum, Ruminococcaceae bacterium D16, Ruminococcus lactaris, Clostridium spiroforme, Firmicutes bacterium CAG 110, Veillonella atypica, Veillonella tobetsuensis, Neisseria flavescens, Klebsiella pneumoniae, and Klebsiella variicola.Among the most important taxonomic features of CRA, differences in the expression of six Bacteroides spp. were observed. Several genera, including Prevotella spp., Parabacteroides spp., and Veillonella spp., were represented by two different species, all of which were overexpressed compared to healthy control samples and Eubacterium spp. and Roseburia spp. Both Eubacterium spp. and Roseburia spp. were underexpressed. All other genera were represented by a single species.

[0088] As shown in Figure 10, the following taxonomic features increased in CRAA: Methanobrevibacter smithii, Bifidobacterium longum, Propionibacterium freudenreichii, Olsenella scatoligenes, Collinsella aerofaciens, Collinsella intestinalis, Collinsella stercoris, Enorma massiliensis, Adlercreutzia equolifaciens, Asaccharobacter celatus, Gordonibacter pamelaeae, Slackia isoflavoniconvertens, Bacteroides ovatus, Bacteroides thetaiotaomicron, Bacteroides uniformis, Bacteroides vulgatus, Bacteroides xylanisolvens, Alistripes inops, Alistripes putredinis, Parabacteroides distasonis, Streptococcus mitis, Clostridium Sp.CAG 167, Eubacterium hallii, Anaerostipes hadrus, Blautia wexlerae, Ruminococcus torquea, Coprococcus catus, Coprococcus comes, Dorea formicigenerans, Dorea longicatena, Fusicatenibacter saccharivorans, Clostridium bolteae, Roseburia faecis, Oscillibacter sp.CAG 241, Intestinibacter bartlettii, Firmicutes bacterium CAG 170, Firmicutes bacterium CAG 238, Firmicutes bacterium CAG 94, Phascolarctobacterium faecium, and Haemophilus parainfluenzii.Key features of CRAAs included overexpression of genera compared to healthy control samples, including four species of Actinomyces, three species of Collinsiella, two species of Enorma, two species of Lactobacillus, two species of Dorea, and two species of Coprococcus. Two species of Bacteroides were underexpressed compared to healthy control samples. Two species of Alistipes spp. showed different representations compared to control samples.

[0089] As shown in Figure 10, the following taxonomic features increased in CRC: Actinomyces turicensis, Bifidobacterium catenulatum, Collinsella aerofaciens, Slackia exigua, Bacteroides fragilis, Bacteroides nordii, Bacteroides plebeius, Butyricimonas virosa, Porphyromonas asaccharolytica, Porphyromonas endodontalis, Porphyromonas uenonis, Alloprevotella tannerae, Prevotella intermedia, Prevotella nigrescens, Prevotella sp CAG 520, Prevotellastercorea, Gemella morbillorum, Streptococcus pasteurianus, Streptococcus salivarius, Clostridium sp CAG 58, Hungatella hathewayi, Mogibacterium diversum, Eubactenum eligens, Eubacterium ramulus, Eubacterium ventriosum, Anaerostipes hadrus, Coprococcus catus, Eisenbergiella tayi, Clostridtum symbiosum, Roseburia intestinalis, Roseburia sp CAG 303, Peptostreptococcus anaerobius, Peptostreptococcus stomatis, Faecalibacterium prausnltzii, Ruminococcaceae bacterium D16, Ruthenibacterium lactatiformans, Solobacterium moorei, Firmicutes bacterium CAG 94, Dialister pneumosintes, Veillonella parvula, Veillonella sp T110116, ParvomonasThe species represented included *Fusobacterium micra*, *Fusobacterium naviforme*, *Fusobacterium nucleatum*, *Fusobacterium sp. oral taxon 370*, *Eikenella corrodens*, *Escherichia coli*, and *Morganella morganii*. Key features of CRC included two *Actinomyces* spp. species, three *Porphorymonas* spp. species, four *Prevotella* spp. species, two *Peptostreptococcus* spp. species, and three *Fusobacterium* spp. species. Some of these features did not show significant ptasticity changes compared to controls, but they did show a significant increase in prevalence. Three *Bacteroides* spp. species and two *Veillonella* spp. species were overrepresented in CRC compared to healthy control samples. All other genera were represented as single species.

[0090] The inventors conducted a detailed meta-analysis of publicly available microbiome shotgun sequencing data derived from fecal samples of healthy control donors and donors diagnosed with CRA, CRAA, and CRC by colonoscopy. The characteristics of the 13 studies analyzed varied considerably, including eight different countries, various disease states, cohort sizes, and the number of reads meeting quality control criteria (Table 1). Most studies attempted to balance sex and age within their respective cohorts, but generally there were more males than females. Additional factors such as DNA preparation and sequencing methods varied between studies, and importantly, some studies collected fecal or other samples after colonoscopy. These factors may lead to variability in study results. Despite these confounding factors, the taxonomic features identified in individual studies, while variable, define a consensus finding, at least with respect to CRC. The fact that similar taxa are identified, given the known inter-individual variability of the microbiome and the different dietary habits in each participating country, strongly suggests that the selective forces acting on CRC are superior to diet and other known selective pressures.

[0091] These findings led to two surprising conclusions regarding the contribution of the selective microenvironment generated by colonic lesions and / or the gut microbiota to disease onset and progression. First, there are significant differences between the microbiota of CRA and CRAA, raising the question of whether progressive adenoma is simply a larger form of adenoma. Our results suggest that this assumption requires further investigation, and instead indicate that the microbiota adenoma / progressive adenoma interaction represents a very different process. Second, and perhaps even more surprising, given the strong opposing behavior of the CRA and CRAA microbiota in terms of genetic expression, the CRA and CRC microbiota are functionally substantially synonymous. In fact, the genetic features indicating the consistent direction of change in the CRA and CRC microbiota are no different from those observed in CRAA. Of the 408 features that were expressed differently in all three disease classes compared to healthy controls, 389 (95%) showed a pattern of directional agreement between CRA and CRC, and directional disagreement between CRAA and the other disease classes. In this regard, the same genetic features that were commonly increased in both CRA and CRC samples were decreased in CRAA compared to healthy control samples, and vice versa. However, since almost all of the progressive adenoma samples were obtained from a single study, these results are considered provisional.

[0092] Example 8: Verification of high-impact features using external data validation tests Following the completion of formal analyses of FIRE and SIAMCAT, data from a new study focusing on the Spanish cohort became available (Front.Microbiol.,2024;11(14):1-17). This data was incorporated and used for additional external validation. The new dataset consisted of 30 CRC samples and 30 CONTROL samples. These data included 30 additional polyp samples that were not analyzed due to the lack of specifications to distinguish between early adenomas (CRA) and late adenomas (CRAA). BB3 was used to establish taxonomic assignments. The new external dataset was scored using a model developed for taxonomic annotation (eXtreme gradient boosted tree classifier) ​​with informational features and thresholds corresponding to those that maximized the F1 score (0.629). The resulting model's predictive performance yielded the following results: (AUC=0.8089, Sensitivity=0.7333, Specificity=0.8333, and Precision=0.7833). Depending on the metrics evaluated, these scores were better than or equivalent to those achieved in external HO data analysis using data from the original metagenomic data. These results indicate that the modeling effort was largely successful, reflecting that the CRC vs. CONTROL classification model demonstrated good generalizability and produced satisfactory performance on new data.

[0093] Example 9: Verification of high-impact features via qPCR Designing primers specific to a target taxa presents a multifaceted challenge. The greatest challenge lies in the large proportion of known sequence space occupied by the target taxa compared to the unknown sequence space existing on Earth. In this regard, the quality of any primer design must be deemed acceptable unless otherwise demonstrated. Targeting unique gene sequences present in the target taxa but not in its vicinity is the most direct method for performing specific qPCR. However, target genes may be widely present in sequenced isolates but not actually present in uncharacterized samples, potentially leading to underestimated abundances in qPCR reactions. Conversely, mapping sequence reads from shotgun metagenomic sequencing of fecal samples is incomplete and limited in accuracy based on the availability of known sequences. In this respect, relative abundance measurements generated by sequence enumeration may be incomplete and therefore may differ from those generated by qPCR. Many of these nuances can be directly assessed in candidate primer design by determining the sequences of PCR products generated from tens or hundreds of reactions and evaluating the purity of the generated product sequences. Nonspecific priming or amplification of neighboring sequences can and should be quantified after a comprehensive initial assessment and before any commercial trials are deployed.

[0094] Based on a list of 300 taxa generated by applying FIRE feature selection (from each category CRC, CRA, and CRAA) and ensemble modeling, top target species were used for primer design. Multiple primer sets were identified for each target based on gene sequences identified as specific to the target. Primer pairs targeting all bacteria using the 16s rRNA gene were used as internal housekeeping controls to normalize results between samples: allbacteria_16S Fw GCAGGCCTAACACATGCAAGTC (SEQ ID NO: 1), allbacteria_16S Rv CTGCTGCCTCCCGTAGGAGT (SEQ ID NO: 2, product size 120 base pairs). A list of all working primers for each category is shown in Table 6.

[0095] Each primer was tested with 10 to 48 samples, approximately 50% from CRC / CRA / CRAA subjects and approximately 50% from CTR (control subjects). Theoretical and results-based evaluations of primer design considered several metrics, including Tm, melting curve, presence or absence of primer dimers, presence or absence of harpins, Ct amplification number, base number, product size, primer couple specificity (based on blast alignment of primer 3), and reproducibility of sequencing data.

[0096] Shotgun sequencing and qPCR data were compared for each specific target taxa to evaluate the quantitative relative abundance of each sample. The results are summarized in Figures 13A-N (CRC), 14A-Q (CRA), and 15A-E (CRAA). The scatter plots evaluate the quantitative relative abundance of each sample based on shotgun sequencing (left) and qPCR (right) for each specific target taxa. Each dot represents one sample. Sequencing samples are ordered from lowest to highest value, and qPCR samples are sorted according to the order of the sequencing samples. The y-axis of the sequencing graph shows the raw data values, and the y-axis of the qPCR graph shows the relative abundance values ​​calculated by Method 2 (-Delata Delta C(T)). The bar graph shows the mean values ​​of the sequencing and qPCR data for the tested samples. Error bars represent the standard error values.

[0097] According to these data, 16 primer pairs reproduce the sequencing data for CRC, 18 for CRA, and 6 for CRAA. Certain primer pairs (not shown) showed high specificity for the target, but the abundance of the taxa was very low. These include Peptostreptococcus stomatis and Dialister pneumosintes for CRC, Caprococcus catus for CRA, and Actinomyces graeviitzii for CRAA. [Table 1-1] Table 1-2 Table 2 Table 3-1 Table 3-2 Table 3-3 Table 3-4 Table 3-5 Table 3-6 Table 3-7 Table 3-8 Table 3-9 Table 3-10 Table 3-11 Table 3-12 Table 3-13 Table 3-14 Table 3-15 Table 3-16 Table 3-17 Table 3-18 Table 3-19 Table 3-20 Table 3-21 Table 3-22 Table 3-23 Table 3-24 Table 3-25 Table 3-26 Table 3-27 Table 3-28 Table 3-29 Table 3-30 Table 3-31 Table 3-32 Table 3-33 Table 3-34 Table 3-35 Table 3-36 Table 3-37 Table 3-38 Table 3-39 Table 3-40 Table 4-1 Table 4-2 Table 4-3 Table 4-4 Table 4-5 Table 4-6 Table 4-7 Table 4-8 Table 4-9 Table 4-10 Table 4-11 Table 4-12 Table 4-13 Table 4-14 Table 4-15 Table 4-16 Table 4-17 Table 4-18 Table 4-19 Table 4-20 Table 4-21 Table 4-22 Table 4-23 Table 4-24 Table 4-25 Table 4-26 Table 4-27 Table 4-28 Table 4-29 Table 4-30 Table 4-31 Table 4-32 Table 4-33 Table 4-34 Table 4-35 Table 4-36 Table 4-37 Table 4-38 Table 4-39 Table 5-1 Table 5-2 Table 5-3 Table 5-4 Table 5-5 Table 5-6 Table 5-7 Table 5-8 Table 5-9 Table 5-10 Table 5-11 Table 5-12 Table 5-13 Table 5-14 Table 5-15 Table 5-16 Table 5-17 Table 5-18 Table 5-19 Table 5-20 Table 5-21 Table 5-22 Table 5-23 Table 5-24 Table 5-25 Table 5-26 Table 5-27 Table 5-28 Table 5-29 Table 5-30 Table 5-31 Table 5-32 Table 5-33 Table 5-34 Table 5-35 Table 5-36 Table 5-37 Table 5-38 Table 5-39 Table 5-40 Table 6-1 Table 6-2 Table 6-3 Table 6-4

Claims

1. A method for evaluating the presence of colorectal neoplasms in a subject, The quantification of genetic elements from biological samples from the subject, wherein the abundance or prevalence of the genetic elements is associated with colorectal cancer (CRC), colorectal adenoma (CRA), or progressive colorectal adenoma (CRAA), thereby creating an abundance profile of the genetic elements, and the quantification includes elements related to the taxonomic classification of microorganisms and / or genetic elements related to the gene function of one or more microorganisms. To evaluate the abundance profiles of signatures indicating the presence of CRC, CRA, and / or CRAA in the aforementioned subject, and The method comprising determining whether the subject is likely to have CRC, CRA, or CRAA.

2. The method according to claim 1, wherein the subject has a low risk of CRC, CRA, CRAA, or colorectal polyps.

3. The method according to claim 1 or 2, wherein the subject does not have a history of CRC, CRA, or CRAA.

4. The method according to claim 2 or 3, wherein the method is performed as an alternative to colonoscopy.

5. The method according to claim 1, wherein the subject is at high or moderate risk of CRC, CRA, CRAA, or colorectal polyps.

6. The method according to claim 5, wherein the subject has a history of CRC, CRA, or CRAA, and / or a family history of CRC.

7. The method according to claim 5 or 6, wherein the method is performed at least once a year or at least every other year.

8. The method according to any one of claims 1 to 7, wherein the subject is at least 45 years old, or at least 50 years old, or at least 55 years old, or at least 60 years old.

9. The method according to any one of claims 1 to 7, wherein the subject is under 45 years of age or under 50 years of age.

10. The method according to any one of claims 1 to 9, wherein the biological sample is feces, blood, serum, intestinal mucosa, mucosal swab, colonoscopy aspirate, lavage solution, biopsy tissue sample, or other biological sample.

11. The method according to any one of claims 1 to 10, wherein the genetic elements are quantified by a procedure including nucleic acid sequencing, PCR, qPCR, or microarray.

12. The method according to claim 11, wherein the genetic element is quantified by nucleic acid sequencing, and the nucleic acid sequencing comprises sequencing of at least about 20,000,000 reads.

13. The method according to claim 12, wherein the nucleic acid sequencing comprises sequencing of at least about 40,000,000 reads.

14. The method according to any one of claims 11 to 13, wherein the nucleic acid sequencing comprises one or more shotgun metagenomic sequencing, rDNA sequencing, targeted amplicon nucleic acid sequencing, or hybridization capture probe sequencing.

15. The method according to claim 14, wherein the nucleic acid sequencing comprises 16S rDNA, 18S rDNA, or ITS amplicon nucleic acid sequencing, and further comprises targeted amplicon nucleic acid sequencing or hybridization capture probe sequencing.

16. The method according to claim 14 or 15, wherein one or more genetic elements are quantified by capturing them from a sequencing library, optionally amplified by PCR, and subsequently sequenced.

17. The method according to any one of claims 1 to 16, wherein the genetic element represents or is associated with colorectal adenoma (CRA).

18. The method according to claim 17, wherein the genetic element includes at least five taxonomic or genetic functional features listed in Table 3.

19. The method according to claim 18, wherein the genetic element includes at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or genetic functional features as listed in Table 3.

20. The method according to claim 19, wherein the genetic element comprises at least one, at least two, at least five, at least about ten, at least about twenty, or at least about fifty taxonomic features listed in Table 3, and at least one, at least two, at least five, at least about ten, at least about twenty, or at least about fifty gene functional features listed in Table 3.

21. The method according to any one of claims 18 to 20, wherein the genetic element has a different abundance or a different prevalence in a sample from a CRA subject compared to a control subject.

22. The method according to any one of claims 1 to 16, wherein the genetic element is related to progressive colorectal adenoma (CRAA).

23. The method according to claim 22, wherein the genetic element includes at least five taxonomic features or gene functional features listed in Table 4.

24. The method according to claim 23, wherein the genetic element includes at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or genetic functional features as listed in Table 4.

25. The method according to claim 24, wherein the genetic element comprises at least one, at least two, at least five, at least about ten, at least about twenty, or at least about fifty taxonomic features listed in Table 4, and at least one, at least two, at least five, at least about ten, at least about twenty, or at least about fifty gene functional features listed in Table 4.

26. The method according to any one of claims 23 to 25, wherein the genetic element has a different abundance or a different prevalence in a sample from a CRAA subject compared to a control subject.

27. The method according to any one of claims 1 to 16, wherein the genetic element is related to colorectal cancer (CRC).

28. The method according to claim 27, wherein the genetic element includes at least five taxonomic features or genetic functional features listed in Table 5.

29. The method according to claim 28, wherein the genetic element includes at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or genetic functional features as listed in Table 5.

30. The method according to claim 28 or 29, wherein the genetic element comprises at least one, at least two, at least five, at least about ten, at least about twenty, or at least about fifty taxonomic features listed in Table 5, and at least one, at least two, at least five, at least about ten, at least about twenty, or at least about fifty gene functional features listed in Table 5.

31. The method according to any one of claims 27 to 30, wherein the genetic element has a different abundance or a different prevalence in a sample from a CRC subject compared to a control subject.

32. The method according to any one of claims 1 to 31, wherein the abundance profile is evaluated for signatures indicating the presence or absence of CRC, CRA, and CRAA, respectively.

33. (a) The signature indicating the presence or absence of CRA is trained by machine learning using samples from the CRA cohort and samples from the control cohort. (b) The signature indicating the presence or absence of CRC is trained by machine learning using samples from the CRC cohort and samples from the control cohort, (c) The method according to any one of claims 1 to 32, wherein the signature indicating the presence or absence of CRAA is trained by machine learning using samples from a CRAA cohort and samples from a control cohort.

34. The method according to claim 33, wherein the signature(s) are trained using a plurality of machine learning algorithms.

35. The method according to claim 33 or 34, wherein at least one machine learning algorithm is supervised machine learning.

36. The method according to claim 35, wherein the machine learning algorithm further comprises one or more of unsupervised machine learning and semi-supervised machine learning.

37. The method according to any one of claims 33 to 36, wherein the machine learning method includes one or more of the following: parametric / nonparametric distance measurement, logistic regression, support vector machine, decision tree, random forest, neural network, probit regression, Fisher's linear discriminant, simple Bayes classifier, perceptron, quadratic classifier, kernel estimation, k-nearest neighbors, learning vector quantization, and principal component analysis.

38. The method according to any one of claims 33 to 36, wherein the machine learning includes comparative machine learning modeling, optimization, testing, evaluation, and ranking of models, the machine learning includes one or more of the following algorithms: deep learning, gradient boosting, neural networks, ensembles, or blender modeling, the machine learning includes one or more of the following algorithms: gradient boosted tree classifier, eXtreme gradient boosted tree classifier, optical gradient boosted tree classifier, optical gradient boosting on elastic network prediction, Keras Slim residual neural network classifier, generalized additive model, elastic network classifier, random forest classifier, deep forest classifier, mean blender classifier, TensorFlow multilayer perceptron classifier, TensorFlow neural network classifier, and Rule-Fit classifier.

39. The method according to any one of claims 33 to 38, wherein the signature(s) include features selected from a training cohort by ensemble ranking of feature importance and by the statistical significance of individual features.

40. The method according to claim 39, wherein the signature(s) include features selected from a training cohort by feature importance rank assembling (FIRE) and by statistical estimation of associations between microbial communities and phenotypes (SIAMCAT).

41. The method according to any one of claims 33 to 40, wherein the signature(s) have a sensitivity of at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.90, or at least about 0.95 for classifying a sample for the presence or absence of CRA, CRAA, or CRC.

42. The method according to any one of claims 33 to 41, wherein the signature(s) have a specificity of at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.90, or at least about 0.95 for classifying a sample for the presence or absence of CRA, CRAA, or CRC.

43. (a) If the subject is not identified as likely to have CRA, CRAA, or CRC, no further action will be taken, and (b) If the subject is identified as likely to have one or more of CRA, CRAA, or CRC, further treatment or therapy is initiated, the method according to any one of claims 1 to 42.

44. The method according to claim 43, wherein the further procedure includes imaging of the colon.

45. The method according to claim 44, wherein the procedure is a colonoscopy and optionally includes the removal of one or more polyps and / or biopsy of growths suspected to contain CRCs.

46. The method according to claim 44 or 45, wherein, if it is confirmed that the subject has CRCs, the subject is treated for the CRCs by one or more of the following: surgery, chemotherapy, radiotherapy, and immunotherapy.

47. A method for creating a genetic signature of a genetic element indicating the presence of a colorectal neoplasm, wherein the method is To provide a training cohort of biological samples from subjects confirmed to possess CRA, CRAA, or CRC, or RNA or DNA isolated from said biological samples. Perform genomic nucleic acid sequencing of the DNA isolated from the aforementioned sample. The method comprising training a gene signature for classifying samples based on the presence or absence of CRA, CRAA, or CRC, wherein the gene signature includes microbial taxonomic classification features and microbial gene functional features.

48. The method according to claim 47, wherein the sample is selected from feces, blood, serum, plasma, urine, saliva, biopsy tissue, mucosal tissue sample or swab, and bowel cleansing solution or aspirate.

49. The method according to claim 48, wherein the sample is a fecal sample or a mucosal tissue sample.

50. The method according to any one of claims 47 to 49, wherein the nucleic acid sequencing comprises sequencing of at least about 20,000,000 reads per sample.

51. The method according to claim 50, wherein the nucleic acid sequencing comprises sequencing of at least about 40,000,000 reads.

52. The method according to any one of claims 47 to 51, wherein the nucleic acid sequencing comprises one or more of shotgun metagenomic sequencing, rDNA sequencing, and targeted nucleic acid sequencing.

53. The method according to claim 52, wherein the nucleic acid sequencing comprises one or more of 16S rDNA sequencing and / or 18S rDNA sequencing and / or ITS sequencing, as well as shotgun sequencing and targeted nucleic acid sequencing.

54. The method according to any one of claims 47 to 53, wherein the genetic elements in the sample are optionally amplified by PCR.

55. The method according to any one of claims 47 to 54, wherein the genetic element is assigned to a reference genome for taxonomic classification, and / or the genetic element is assigned to a gene function.

56. The method according to any one of claims 47 to 55, wherein a sample from a CRA subject is selected to have microbial taxonomic characteristics and microbial gene functional characteristics having different abundances or different prevalences compared to a control subject.

57. The method according to claim 56, wherein the aforementioned features include, optionally, at least five taxonomic features and / or genetic functional features listed in Table 3.

58. The method according to claim 57, wherein the aforementioned features optionally include at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or genetic functional features as listed in Table 3.

59. The method according to claim 57 or 58, wherein the features include, optionally, at least one, at least two, at least five, at least about ten, at least about twenty, or at least about fifty taxonomic features listed in Table 3, and optionally, at least one, at least two, at least five, at least about ten, at least about twenty, or at least about fifty genetic functional features listed in Table 3.

60. The method according to any one of claims 47 to 55, wherein a sample from a CRAA subject is selected to have microbial taxonomic characteristics and microbial gene functional characteristics having different abundances or different prevalences compared to a control subject.

61. The method according to claim 60, wherein the aforementioned features include, optionally, at least five taxonomic features and / or genetic functional features listed in Table 4.

62. The method according to claim 61, wherein the aforementioned features optionally include at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or genetic functional features listed in Table 4.

63. The method according to claim 61 or 62, wherein the features include, optionally, at least one, at least two, at least five, at least about ten, at least about twenty, or at least about fifty taxonomic features listed in Table 4, and optionally, at least one, at least two, at least five, at least about ten, at least about twenty, or at least about fifty genetic functional features listed in Table 4.

64. The method according to any one of claims 47 to 55, wherein a sample from a CRC subject is selected to have microbial taxonomic characteristics and microbial gene functional characteristics having different abundances or different prevalences compared to a control subject.

65. The method according to claim 64, wherein the aforementioned features include, optionally, at least five taxonomic features and / or genetic functional features listed in Table 5.

66. The method according to claim 65, wherein the aforementioned features optionally include at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or genetic functional features listed in Table 5.

67. The method according to claim 65 or 66, wherein the features include, optionally, at least one, at least two, at least five, at least about ten, at least about twenty, or at least about fifty taxonomic features listed in Table 5, and optionally, at least one, at least two, at least five, at least about ten, at least about twenty, or at least about fifty genetic functional features listed in Table 5.

68. At least three gene signatures are trained, (1) Classify the samples based on the presence or absence of CRA, (2) Classify the samples based on the presence or absence of CRAA, and (3) The method according to any one of claims 47 to 67, for classifying samples based on the presence or absence of CRC.

69. (a) The signature for classifying samples based on the presence or absence of CRA is trained by machine learning using samples from the CRA cohort and samples from the control cohort. (b) The signature for classifying samples based on the presence or absence of CRC is trained by machine learning using samples from the CRC cohort and samples from the control cohort, (c) The method according to claim 68, wherein the signature for classifying samples based on the presence or absence of CRAA is trained by machine learning using samples from a CRAA cohort and samples from a control cohort.

70. The method according to claim 69, wherein the signature(s) are trained using a plurality of machine learning algorithms.

71. The method according to claim 69 or 70, wherein at least one machine learning algorithm is supervised machine learning.

72. The method according to claim 71, wherein the machine learning algorithm further comprises one or more of unsupervised machine learning or semi-supervised machine learning.

73. The method according to any one of claims 69 to 72, wherein the machine learning method includes one or more of the following: parametric / nonparametric distance measurement, logistic regression, support vector machine, decision tree, random forest, neural network, probit regression, Fisher's linear discriminant, simple Bayes classifier, perceptron, quadratic classifier, kernel estimation, k-nearest neighbors, learning vector quantization, and principal component analysis.

74. The method according to any one of claims 71 to 73, wherein the machine learning includes one or more of the following: deep learning, gradient boosting, neural networks, ensembles, or Blender modeling algorithms, and comprises comparative machine learning modeling, optimization, testing, evaluation, and ranking of models, which includes one or more of the following: gradient boosted tree classifier, eXtreme gradient boosted tree classifier, optical gradient boosted tree classifier, optical gradient boosting on elastic network prediction, Keras Slim residual neural network classifier, generalized additive model, elastic network classifier, random forest classifier, deep forest classifier, mean Blender classifier, TensorFlow multilayer perceptron classifier, TensorFlow neural network classifier, and Rule-Fit classifier.

75. The method according to any one of claims 69 to 74, wherein the signature(s) include features selected from a training cohort by ensemble ranking of feature importance and by the statistical significance of individual features.

76. The method according to claim 75, wherein the signature(s) include features selected from a training cohort by feature importance rank assembling (FIRE) and by statistical estimation of associations between microbial communities and phenotypes (SIAMCAT).

77. The method according to any one of claims 47 to 76, wherein the signature(s) have a sensitivity of at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.90, or at least about 0.95 for classifying a sample for the presence or absence of CRA, CRAA, or CRC.

78. The method according to any one of claims 47 to 77, wherein the signature(s) have a specificity of at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.90, or at least about 0.95 for classifying a sample for the presence or absence of CRA, CRAA, or CRC.

79. A method for creating a gene signature for a genetic element that exhibits colon disorders, wherein the method is To provide a training cohort of biological samples from subjects and control subjects confirmed to have colon disorders, or DNA isolated from said biological samples. Perform genomic nucleic acid sequencing of the DNA isolated from the aforementioned sample. The method comprising training a gene signature for classifying samples based on (1) the presence or absence of the colon disorder, wherein the gene signature includes features selected from a training cohort by ensemble ranking of feature importance and by the statistical significance of individual features.

80. The method according to claim 79, wherein the biological sample is selected from feces, blood, serum, plasma, urine, saliva, biopsy tissue, mucosal tissue sample or swab, and bowel cleansing solution or aspirate.

81. The method according to claim 80, wherein the sample is a fecal sample or a mucosal tissue sample.

82. The method according to any one of claims 79 to 81, wherein the signature(s) include features selected from a training cohort by feature importance rank assembling (FIRE) and by statistical estimation of associations between microbial communities and phenotypes (SIAMCAT).

83. The method according to any one of claims 79 to 82, wherein the colon disorder is selected from Crohn's disease, ulcerative colitis, irritable bowel syndrome (IBS), bifurcationitis, colorectal adenoma (CRA), progressive colorectal adenoma (CRAA), and colorectal cancer (CRC).

84. The method according to claims 79 to 83, wherein the aforementioned features include microbial taxonomic classification features and microbial gene function features.

85. The method according to any one of claims 79 to 84, wherein the nucleic acid sequencing comprises sequencing of at least about 20,000,000 reads per sample.

86. The method according to claim 85, wherein the nucleic acid sequencing comprises sequencing at least about 40,000,000 reads per sample, or at least about 50,000,000 reads, or at least about 60,000,000 reads, or at least about 75,000,000 reads, or at least about 100,000,000 reads, or at least about 150,000,000 reads.

87. The method according to any one of claims 79 to 86, wherein the nucleic acid sequencing comprises one or more of shotgun metagenomic sequencing, 16S rDNA sequencing, and targeted nucleic acid sequencing.

88. The method according to claim 87, wherein the nucleic acid sequencing comprises one or more of 16S rDNA sequencing, shotgun sequencing, and targeted nucleic acid sequencing.

89. The method according to any one of claims 79 to 88, wherein the genetic elements in the sample are optionally amplified by PCR.

90. The method according to any one of claims 79 to 89, wherein the genetic element is assigned to a reference genome for taxonomic classification, and / or the genetic element is assigned to a gene function.

91. The method according to any one of claims 79 to 90, wherein in a sample from a subject with colon disorders, microbial taxonomic characteristics and microbial gene functional characteristics having different abundances or different prevalences compared to a control subject are selected.

92. The method according to claim 91, wherein the features include at least five taxonomic features and / or genetic functional features.

93. The method according to claim 92, wherein the features include at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or genetic functional features.

94. The method according to claim 92 or 93, wherein the features include at least one, at least two, at least five, at least about ten, at least about twenty, or at least about fifty taxonomic features, and at least one, at least two, at least five, at least about ten, at least about twenty, or at least about fifty genetic functional features.

95. The method according to any one of claims 79 to 94, wherein the signature(s) are trained using a plurality of machine learning algorithms.

96. The method according to claim 95, wherein at least one machine learning algorithm is supervised machine learning.

97. The method according to claim 96, wherein the machine learning algorithm further comprises one or more of unsupervised machine learning or semi-supervised machine learning.

98. The method according to any one of claims 95 to 97, wherein the machine learning method includes one or more of the following: parametric / nonparametric distance measurement, logistic regression, support vector machine, decision tree, random forest, neural network, probit regression, Fisher's linear discriminant, simple Bayes classifier, perceptron, quadratic classifier, kernel estimation, k-nearest neighbors, learning vector quantization, and principal component analysis.

99. The method according to any one of claims 96 to 98, wherein the machine learning includes comparative machine learning modeling, optimization, testing, evaluation, and ranking of models, the machine learning includes one or more of the following algorithms: deep learning, gradient boosting, neural networks, ensembles, or blender modeling, the machine learning includes one or more of the following algorithms: gradient boosted tree classifier, eXtreme gradient boosted tree classifier, optical gradient boosted tree classifier, optical gradient boosting on elastic network prediction, Keras Slim residual neural network classifier, generalized additive model, elastic network classifier, random forest classifier, deep forest classifier, mean blender classifier, TensorFlow multilayer perceptron classifier, TensorFlow neural network classifier, and Rule-Fit classifier.

100. The method according to any one of claims 79 to 99, wherein the signature(s) have a sensitivity of at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.90, or at least about 0.95 for classifying a sample for the presence or absence of the colonic disorder.

101. The method according to any one of claims 79 to 100, wherein the signature(s) have a specificity of at least about 0.70, at least about 0.75, at least about 0.80, at least about 0.90, or at least about 0.95 for classifying a sample for the presence or absence of colonic disorder.