A method for screening for colorectal cancer using fecal microbiome profiling

JP2025519728A5Pending Publication Date: 2026-03-10BARCELONA SUPERCOMPUTING CENT CENT NAT DE SUPERCOMPUTACION +3
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-06-16
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Current colorectal cancer (CRC) screening methods, particularly the fecal immunochemical test (FIT), suffer from high false-positive rates, leading to unnecessary colonoscopies and increased healthcare costs.

Method used

A method combining fecal microbiome profiling with a two-stage AI-based classification algorithm to identify specific bacterial taxa and metabolic pathways associated with CRC development, thereby reducing unnecessary colonoscopies and improving diagnostic accuracy.

Benefits of technology

The method achieves a sensitivity close to 100% for CRC detection while significantly reducing the false-positive rate, thereby optimizing resource allocation and improving patient outcomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The present invention relates to a two-step method for screening for colorectal cancer (CRC) using fecal microbiome profiling. The method includes determining the levels of two or more bacterial taxa in a fecal sample isolated from a subject, classifying CRC samples versus non-CRC samples by a computer algorithm as a first step, and classifying samples classified as non-CRC in the first step into CR samples and non-CR samples using two or more bacterial taxa specifically enriched in clinically relevant (CR) samples compared to non-CR samples by a computer algorithm as a second step. The present invention also relates to a kit including reagents and a computer program for performing the method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of medicine. More particularly, the present invention relates to a method for screening for colorectal cancer using fecal microbiome profiling.

Background Art

[0002] Colorectal cancer (CRC) is the third most common type of cancer and the second leading cause of cancer-related deaths worldwide (1), and is the main cause of nearly 900,000 deaths per year. CRC exhibits various molecular phenotypes and significant resistance to therapies. It has been suggested that this malignant disease progresses from the pathological transformation of normal colonic epithelium to adenomatous polyps, which ultimately lead to invasive cancer. This process is slow and involves the accumulation of genetic and / or epigenetic mutations (2). Non-environmental risk factors in CRC include age and genetic susceptibility (3).

[0003] The incidence of CRC has increased with economic development and the westernization of eating and lifestyle habits, suggesting a significant effect of environmental and lifestyle factors that may act in combination with genetic predisposition (4). Regarding this matter, more and more evidence is associated with changes in the gastrointestinal microbiota and CRC development (5).

[0004] Earlier studies have shown that changes in the gut microbiota may affect colonic tumorigenesis via chronic inflammation or the production of carcinogenic compounds (7) (6). Comparing paired tumor and normal tissues, or fecal samples from CRC patients and healthy subjects, reveals differences in the relative abundance of some microbial species or genera (8, 9). The diagnosis of CRC is difficult and involves a complex process that usually begins with the detection of initial symptoms by the patient, followed by clinical diagnostic procedures mainly based on colonoscopy.

[0005] Prevention and early diagnosis of CRC can save many lives (10, 11), and regular screening of populations over a certain age is carried out in many countries. Current CRC screening consists of a two-step procedure with a non-invasive test (most commonly, quantification of fecal immunochemical test (FIT) for occult hemoglobin in stool), followed by colonoscopy if the test is positive (FIT-positive, assigned threshold hemoglobin concentration) (12, 13). This approach is effective, but it results in a high false-positive rate in the first step and many unnecessary colonoscopies (only about 20 - 30% of the colonoscopies performed in FIT-positive individuals exhibit clinically relevant characteristics, and CRC is only 3 - 5%) (14).

[0006] Colonoscopy is an invasive, expensive, and time-consuming procedure. Therefore, additional biomarkers that can better stratify individuals at higher risk for CRC and pre-cancerous lesions associated with the risk who should undergo colonoscopy should significantly reduce healthcare costs.

[0007] Many current studies have focused on finding additional criteria, such as risk factors and other biomarkers, that are considered by the decision algorithms used to individualize positive FIT tests for colonoscopy. These include the examination of molecular biomarkers related to the processes underlying colorectal carcinogenesis, from circulating tumor cells (15), extracellular DNA (16), microRNA (17), and also metabolites from plasma (18) samples, and germline risk gene variants from blood DNA (19). As evidence for the existence of microbiome changes associated with CRC and the possible involvement of the microbiota in cancer origin and progression is increasing (5, 21), microbial markers have emerged in recent years as promising additional factors to be considered in early screening.

[0008] In addition, better understanding the role of gut metabolism, microbiota, and microbiota-host interactions at the initiation stage of CRC can help establish preventive measures such as dietary changes or the use of probiotics or prebiotics.

[0009] Overall, there is a need for early diagnostic non-invasive techniques to diagnose this malignant disease and enable the best prognosis and quality of life for patients.

Summary of the Invention

[0010] Brief Description of the Invention The present invention discloses an innovative approach for the early detection of colorectal cancer (CRC) that combines the microbiome profiling of samples with a two-stage AI-based classification algorithm designed for the early detection of clinically relevant cases to reduce the number of unnecessary colonoscopies and to provide a better prognosis for CRC patients.

[0011] To explore potential predictive biomarkers present in FIT and other types of fecal samples and to shed light on the potential role of the gut microbiome in CRC development, targeted sequencing of the V3-V4 region of the 16S rRNA gene from DNA directly extracted from FIT tubes collected within a population screening program conducted in Catalonia, Spain, was used to perform microbiome profiling (22).

[0012] A total of 2,889 FIT-positive samples and 246 FIT-negative samples were analyzed. Their microbial composition and metabolic capabilities were determined and studied together with various colonoscopy results to see how much they vary between samples.

[0013] Significant differences in specific taxa and metabolic pathways were observed between related stages in the CRC development along the path from healthy tissue to cancer. Using the diagnostic evaluation of colonoscopy, changes in the composition, taxa co-occurrence, and metabolic characteristics of the microbiota associated with clinically relevant traits such as the presence of polyps or definite pre-cancerous lesions were reconstructed, suggesting the potential role of microbiota in the origin and progression of CRC.

[0014] Finally, using machine learning algorithms, a two-step classification mechanism was developed and validated that combines information from a high-sensitivity bacterial signature, gender, age, and hemoglobin to limit unnecessary colonoscopies while minimizing the false negative rate (Figure 1). This classification mechanism achieved a sensitivity close to 100% for CRC while significantly reducing the current false positive rate.

[0015] The present invention relates to the method described in the claims. In a first embodiment, the present disclosure is a method for diagnosing a subject as to whether they have colorectal cancer (CRC) or for classifying a subject in a patient cohort as to whether they have a higher risk of developing CRC, comprising: (i) determining the levels of three or more bacterial taxa in a fecal sample isolated from the subject; (ii) as a first step, classifying CRC samples versus non-CRC samples using a computer algorithm with two or more bacterial taxa specifically enriched in CRC samples compared to non-CRC samples, the hemoglobin content of the sample, and the age and gender of the donor; (iii) as a second step, classifying the samples classified as non-CRC in the first step into CR samples and non-CR samples using a computer algorithm with two or more bacterial taxa specifically enriched in clinically relevant (CR) samples compared to non-CR samples, the hemoglobin content of the sample, and the age and gender of the donor, where CR includes intermediate risk lesions, high risk lesions, carcinoma in situ (CIS), and colorectal cancer (CRC). comprising: Three or more bacterial taxa in step (i) are species of the genus Hungatella, species of the genus Colinsella, species of the genus Tyzzerella, Phascolarctobacterium succinatutens, species of the genus Lactobacillus, species of the genus Akkermansia, Akkermansia muciniphila, O. Mollicutes_RF39.UCF, species of the family Ruminococcaceae_UCG.002, species of the family Ruminococcaceae_UCG.0010, species of the genus Odoribacter, O. Rhodospirillales.UCF, species of the genus Victivallis, species of the family Ruminococcaceae_UCG.005, species of the genus Negativibacillus, species of the family Christensenellaceae_R.7_group, species of the genus Oxalobacter, species of the genus Butyrivibrio, species of the family XIII_UCG.001, species of the genus Gemella, species of the genus Peptostreptococcus, species of the genus Pediococcus) Lactobacillus vaginalis, Enorma massiliensis, Megamonas funiformis, Peptostreptococcus anaerobius, Peptoniphilus lacrimalis, Lactobacillus oris, Alloscardovia omnicolens, Allisonella histaminiformans, Acidaminococcus fermatans, Collinsella bouchesdurhonensis, Corynebacterium spp., Veillonella dispar, Ezakiella spp., O. Chloroplast.UCF, Sphingomonas spp., Dialister succinatiphilus, Finegoldia magna, Bacteroides coprophilus, Eggerthella spp., Acidaminococcus spp., Enterococcus spp.) Sutterella wadsworthensis, Bacteroides fragilis, Bacteroides plebeius, Bacteroides coprocola, Bifidobacterium longum, Bilofila spp., Parabacteroides merdae, DTU08 spp., Oscillibacter spp., Parabacteroides goldsteinii, Parabacteroides spp., Bacteroides spp., Coprobacter secundus, Prevotella timonensis, Streptococcus parasanguinis, Peptostreptococcus anaerobius, Streptococcus sobrinus, Lachnospiraceae_FCS020_group bacteria, Bifidobacterium dentium, Porphyromonas spp., Lachnospiraceae_UCC.008 spp., Enterobacter spp., Hungatella hathewayi, Ezakiella spp., Leukonostoc spp.) Parabacteroides johnsonii, Bacteroides finegoldii, species of Eisenbergiella spp., Alistipes finegoldii, F. Erysipelotrichaceae.UCG, Dorea formicigenerans, Bacteroides caccae, Fusobacterium unclassified.S106, Peptostreptococcus unclassified.S87, Erysipelotrichaceae_UCG.003 unclassified.S297, Alistipes putredinis, Prevotella unclassified.S33, and Coprococcus comes, selected from the group consisting of. refers to the method.

[0016] In a second embodiment, the present disclosure (a) A reagent for implementing a method for determining the presence or abundance of bacteria in a fecal sample for determining the level of two or more bacterial taxa in step (i) of the method of the first embodiment, and (b) A computer program stored on a computer-readable data carrier or chip, which, when executed by a computer, includes instructions for causing the computer to perform steps (ii) and (iii) of the method of the first embodiment refers to a kit comprising.

[0017] The following figures are merely illustrative of the invention and should in no way be construed as limiting the scope of the invention as set forth in the appended claims.

Brief Description of the Drawings

[0018]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

[0019] DA_taxa: All specifically abundant taxa intersecting between CRIPREV and the validation dataset were used as features.

[0020] 4-4 taxa panel: The 4-taxa panel for each stage.

[0021] 4 - 4 classification group panel, adjW: 4 - 4 classification group panel for each stage, and the penalty for CR samples in the second stage is smaller.

[0022] FIT_filter_4 - 4 classification group panel: Samples with a FIT value (μg hemoglobin / g feces) exceeding 954 were sent for colonoscopy, and the remaining samples were subjected to a classification mechanism.

[0023] FIT_filter_4 - 4 classification group panel_adjW: Samples with a FIT value (μg hemoglobin / g feces) exceeding 954 were sent for colonoscopy, and the remaining samples were subjected to a classification mechanism. The penalty for CR samples in the second stage is smaller.

Modes for Carrying Out the Invention

[0024] Definitions The present invention will be described in more detail below with reference to the figures. However, the specific embodiments, examples, or results described herein are for illustrative purposes only and should in no way be construed as limiting the scope of the present invention as defined by the appended claims.

[0025] It should be understood that the specific methodologies, protocols, and reagents described herein may vary, and thus the present invention is not limited thereto. It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of the present invention, which is limited only by the appended claims. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art.

[0026] Each document cited in this specification (including all patents, patent applications, scientific publications, manufacturer's specifications, instructions, etc.), whether supra or infra, is hereby incorporated by reference in its entirety. In the event of a conflict between the definitions or teachings of such incorporated references and the definitions or teachings set forth herein, the text of this specification shall control.

[0027] As used herein, the term "comprising" or variations thereof, e.g., "comprises" (particularly in the context of the claims), shall be construed as open-ended terms or non-exclusive inclusion (i.e., meaning "including but not limited to") unless otherwise noted. The term "comprising" is intended to include the more limiting terms "consisting essentially of" or "substantially comprising", and "consisting of".

[0028] In the case of a chemical compound or composition, the term "consisting essentially of" or "substantially comprising" means that certain additional components, i.e., components that do not substantially affect the essential characteristics of the compound or composition, may be present, e.g., unavoidable impurities.

[0029] The terms "a", "an", and "the", as used herein in the context of describing the invention (particularly in the context of the claims), shall each be read and construed to include at least one element or component, and shall be construed to include both the singular and the plural unless otherwise indicated herein or clearly contradicted by the context.

[0030] In addition, unless expressly stated to the contrary, the term "or" refers to an inclusive "or" and not an exclusive "or" (i.e., meaning "and / or").

[0031] The phrase "selected from the group consisting of" means that one or more members (plural possible) of the group may be used in any combination (plural possible).

[0032] All numerical values are assumed, in this specification, whether or not explicitly indicated, to be modified by the term "about". The recitation of a range of values is intended, in this specification, merely as a concise way of referring individually to each separate value falling within the range and each separate value is incorporated into this specification as if it were individually recited herein.

[0033] The use of the terms "for example", "e.g.", "such as", or variations thereof is intended, unless otherwise claimed, merely to clarify the invention and does not impose a limitation on the scope of the invention. These terms should be construed to mean "but not limited to".

[0034] The term "species of" means an unclassified bacterial species from the same bacterial genus. For example, Akkermansis spp. means unclassified Akkermansia species. Instead, the term "unclassified" is used herein to indicate an unclassified bacterial species from the same bacterial genus. Thus, "species of" and "unclassified" have the same meaning and both terms are used interchangeably herein.

[0035] The term "bacterial level(s)" means the relative abundance of a given bacterial taxon relative to other bacterial taxa present in the same sample.

[0036] The term "bacterial profile" means a series of relative abundances of bacterial taxa for a given sample.

[0037] The term "taxon" means a member of a taxonomic rank and includes, for example, a bacterial family, genus, or species.

[0038] The term "FIT value" means hemoglobin content, i.e., μg hemoglobin / g feces. The term "fecal immunochemical test" or "FIT" means any fecal test for determining occult hemoglobin in feces by immunochemistry, such as a fecal immunochemical tab (FIT) or immunochemical fecal occult blood (iFOB).

[0039] The term "clinically relevant" (CR) means a grouping defined for risk stages in the development of CRC that includes intermediate risk lesions (IRL), high risk lesions (HRL), carcinoma in situ (CIS), and colorectal cancer (CRC), but does not include negative / healthy (N), lesions not related to risk (LNAR), and low risk lesions (LRL).

[0040] The term "CRIPREV" means the research project in the Catalan CRC Screening Program from which samples for the present invention were received. CriPrev: Prevention of Colorectal Cancer in the Average Risk Population Using Genomic Biomarkers and Microbiomics. PERIS, grant from the Government of Catalonia (reference: SLT002 / 16 / 00398).

[0041] The term "method for determining the presence or abundance of bacteria" means any method or protocol used to determine the presence or abundance of bacteria, including sequencing of PCR from gene amplicons such as the 16S rRNA gene, whole shotgun sequencing, cell-based methods such as flow cytometry, quantitative PCR (qPCR), proteomics, and antibody-based detection methods.

[0042] In this specification, there is no term that should be construed as indicating any element not claimed as an essential element for the practice of the present invention.

[0043] Embodiments In a first aspect, the present disclosure is a method for diagnosing a subject as to whether they have colorectal cancer (CRC) or for classifying a subject in a patient cohort as to whether they have a higher risk of developing CRC, comprising: (i) determining the levels of two or more bacterial taxa in a fecal sample isolated from a subject in a patient cohort; (ii) as a first step, classifying CRC samples versus non-CRC samples using a computer algorithm with two or more bacterial taxa specifically enriched in CRC samples compared to non-CRC samples, the hemoglobin content of the sample, and the age and sex of the donor; (iii) as a second step, classifying the samples classified as non-CRC in the first step into CR samples and non-CR samples using a computer algorithm with two or more bacterial taxa specifically enriched in clinically relevant (CR) samples compared to non-CR samples, the hemoglobin content of the sample, and the age and sex of the donor, where CR includes intermediate-risk lesions, high-risk lesions, carcinoma in situ (CIS), and CRC; comprising: Two or more bacterial taxa are the species of the genus Hungateella, the species of the genus Collinsella, the species of the genus Tyzzerella, Faecalibacterium succinatutens, the species of the genus Lactobacillus, the species of the genus Akkermansia, Akkermansia muciniphila, O. mollicutes_RF39.UCF, the species of the family Lachnospiraceae_UCG.002, the species of the family Lachnospiraceae_UCG.0010, the species of the genus Odoribacter, O. rhodospirillales.UCF, the species of the genus Victivallis, the species of the family Lachnospiraceae_UCG.005, the species of the genus Negativibacillus, the species of the family Christensenellaceae_R.7_group, the species of the genus Oxalobacter, the species of the genus Butyrivibrio, the species of the family XIII_UCG.001, the species of the genus Gemella, the species of the genus Peptostreptococcus, the species of the genus Pediococcus, Lactobacillus vaginalis, Eubacterium massiliense, Megamonas funiformis, Peptostreptococcus anaerobius, Peptoniphilus lacrimalis, Lactobacillus oris, Alloscardovia omnicolens, Allisonella histaminiformans, Acidaminococcus fermentans, Collinsella bowersoxdurrhonensis, the species of the genus Corynebacterium, Veillonella dispar, the species of the genus Escherichia, O. chloroplast.UCF, the species of the genus Sphingomonas, Dialister succinatiphilus, Finegoldia magna, Bacteroides coprophilus, the species of the genus Eggerthella, the species of the genus Acidaminococcus, the species of the genus Enterococcus, Shuttleworthia wautersii, Bacteroides fragilis, Bacteroides prevotii, Bacteroides coprocola, Bifidobacterium longum, the species of the genus Biophila, Parabacteroides merdae, the species of DTU08, the species of the genus Oscillibacter, Parabacteroides goldsteinii, the species of the genus Parabacteroides, the species of the genus Bacteroides, Coprobacter secundus, Prevotella timonensis, Streptococcus parasanguinis, Peptostreptococcus anaerobius, Streptococcus sobrinus, the bacteria of the family Lachnospiraceae_FCS020_group, Bifidobacterium dentium, the species of the genus Porphyromonas, the family Lachnospiraceae_UCC.Selected from the group consisting of species of 008, species of Enterobacter, Hungateella hathewayi, species of Ezakiella, species of Leuconostoc, Parabacteroides johnsonii, Bacteroides finegoldii, species of Eisenbergiella, Alistipes finegoldii, F. erythroperotrichaceae.UCG, Dorea formicigenerans, and Bacteroides caccae. Relates to a method.

[0044] In another aspect, the present disclosure is a method for diagnosing a subject as having colorectal cancer (CRC) or for classifying a subject as having a higher risk of developing CRC in a patient cohort, comprising: (i) determining the levels of three or more bacterial taxa in a fecal sample isolated from a subject in a patient cohort; (ii) as a first step, classifying CRC samples versus non-CRC samples using a computer algorithm with two or more bacterial taxa specifically enriched in CRC samples compared to non-CRC samples, the hemoglobin content of the sample, and the age and gender of the donor; (iii) as a second step, classifying the samples classified as non-CRC in the first step into CR samples and non-CR samples using a computer algorithm with two or more bacterial taxa specifically enriched in clinically relevant (CR) samples compared to non-CR samples, the hemoglobin content of the sample, and the age and gender of the donor, where CR includes intermediate risk lesions, high risk lesions, carcinoma in situ (CIS), and CRC; comprising the steps of Two or more bacterial taxa are species of the genus Hungateella, species of the genus Collinsella, species of the genus Tyzzerella, Faecalibacterium succinatutens, species of the genus Lactobacillus, species of the genus Akkermansia, Akkermansia muciniphila, O. mollicutes_RF39.UCF, species of the family Lachnospiraceae_UCG.002, species of the family Lachnospiraceae_UCG.0010, species of the genus Odoribacter, O. rhodospirillales.UCF, species of the genus Victivallis, species of the family Lachnospiraceae_UCG.005, species of the genus Negativibacillus, species of the family Christensenellaceae_R.7_group, species of the genus Oxalobacter, species of the genus Butyrivibrio, species of the family XIII_UCG.001, species of the genus Gemella, species of the genus Peptostreptococcus, species of the genus Pediococcus, Lactobacillus baguinalis, Eubacterium massiliense, Megamonas funiformis, Peptostreptococcus anaerobius, Peptoniphilus lacrimalis, Lactobacillus oris, Alloscardovia omnicolens, Allisonella histaminiformans, Acidaminococcus fermentans, Collinsella bowersdorferiensis, species of the genus Corynebacterium, Veillonella dispar, species of the genus Escherichia, O. chloroplast.UCF, species of the genus Sphingomonas, Dialister succinatiphilus, Finegoldia magna, Bacteroides coprophilus, species of the genus Eggerthella, species of the genus Acidaminococcus, species of the genus Enterococcus, Suttrella wadsworthia, Bacteroides fragilis, Bacteroides prevotii, Bacteroides coprocola, Bifidobacterium longum, species of the genus Virgibacillus, Parabacteroides merdae, species of DTU08, species of the genus Oscillibacter, Parabacteroides goldsteinii, species of the genus Parabacteroides, species of the genus Bacteroides, Coprobacter secundus, Prevotella timonensis, Streptococcus parasanguinis, Peptostreptococcus anaerobius, Streptococcus sobrinus, Lachnospiraceae_FCS020_group bacteria, Bifidobacterium dentium, species of the genus Porphyromonas, Lachnospiraceae_UCC.A group selected from the species of 008, species of the genus Enterobacter, Hungateella hathewayi, species of the genus Ezakiella, species of the genus Leuconostoc, Parabacteroides johnsonii, Bacteroides finegoldii, species of the genus Eisenbergiella, Alistipes finegoldii, F. erythroperotrichaceae_UCG, Dorea formicigenerans, Bacteroides caccae, Fusobacterium genus unclassified.S106, Peptostreptococcus genus unclassified.S87, Erythroperotrichaceae_UCG.003 unclassified.S297, Alistipes putredinis, Prevotella genus unclassified.S33, and Coprococcus comes. Relating to a method.

[0045] In a more preferred embodiment, the taxon in either step ii) or iii) is selected from any of the following: Bacteroides coprocola, Bifidobacterium longum, Porphyromonas genus unclassified.S30, Eisenbergiella genus unclassified.S226, Peptostreptococcus genus unclassified.S87, Negativibacillus genus unclassified.S269, unclassified.unclassified.S306, Acidaminococcus genus unclassified.S307, Bacteroides coprocola, Bifidobacterium longum, Odoribacter genus unclassified.S27, Porphyromonas genus unclassified.S30, Christensenellaceae_R.7_group unclassified.S209, Eisenbergiella genus unclassified.S226, Peptostreptococcus genus unclassified.S87, Lachnospiraceae_UCG.005 unclassified.S92, and Akkermansia genus unclassified.S361.

[0046] One skilled in the art will recognize that the phrase "classifying patients having a risk of developing colorectal cancer" includes the diagnosis of non-CRC, as well as various stages of CRC development, such as negative (N), lesions not associated with risk (LNAR), low-risk lesions (LRL), intermediate-risk lesions (IRL), high-risk lesions (HRL) and carcinoma in situ (CIS), and the diagnosis of colorectal cancer (CRC). In order to achieve maximum sensitivity (error 0), CRC that is misclassified may also be included in the second stage, so CRC is considered in both stages of the method. This provides a second opportunity for the sample to be classified as clinically relevant in the model.

[0047] In a preferred embodiment, the fecal sample is a fecal immunochemical test (FIT) sample. The fecal sample of the patient is preferably a sample used in the fecal immunochemical test (FIT). In a preferred embodiment, since no additional fecal sample needs to be collected and stored from the patient for analysis, the fecal sample is a FIT-positive sample (i.e., having a hemoglobin content exceeding 20 μg hemoglobin / g feces). The method of the present invention significantly reduces the current false positive rate of FIT. Of course, any fecal sample can be used in the method of the present invention, and the method of the present invention is not limited to FIT samples. In another preferred embodiment, the fecal sample is a FIT-negative sample (i.e., having a hemoglobin content of 20 μg hemoglobin / g feces or less).

[0048] According to the present invention, the method includes determining the levels of two or more bacterial taxa, preferably three or more bacterial taxa, in steps (ii) and (iii). This should not be understood as a limiting characteristic, that is, in the present invention, the levels of combinations of 4, 5, 6, 7 or more taxa may be determined in each step if this is suitable or desirable. It is understood that one of two or more bacterial taxa may be present simultaneously in each step.

[0049] Examples of combinations of bacteria for which a level is determined are combinations of bacteria selected from the group consisting of (the meanings of the terms "taxadown", "taxatop", and "taxarandom" are explained in the section "Combinations of taxa" below). In a more preferred embodiment, when the sample is FIT positive, the taxa are selected from any of the following combinations:

[0050] [Table 1] TIFF2025519728000004.tif194149 TIFF2025519728000005.tif194149 TIFF2025519728000006.tif194149 TIFF2025519728000007.tif194149 TIFF2025519728000008.tif194149 TIFF2025519728000009.tif194149 TIFF2025519728000010.tif209149 TIFF2025519728000011.tif221149 TIFF2025519728000012.tif221149 TIFF2025519728000013.tif221149 TIFF2025519728000014.tif221149 TIFF2025519728000015.tif70149

[0051] In the above group, the first half of the taxa are the taxa determined in the first stage, and the next half are the bacterial taxa determined in the second stage.

[0052] In a more preferred embodiment, the taxon is selected from the group consisting of species of the genus Akkermansia, Akkermansia muciniphila, Bacteroides fragilis, Bacteroides prevotii, species of the genus Negativibacillus, Bacteroides coprocola, Bacteroides caccae, and Dorea formicigenerans.

[0053] In an even more preferred embodiment, in the first stage of the method, the levels of species of the genus Akkermansia, Akkermansia muciniphila, Bacteroides fragilis and Bacteroides prevotii are determined and the subject is classified as to whether it has CRC, and in the second stage, the levels of species of the genus Negativibacillus, Bacteroides coprocola, Bacteroides caccae and Dorea formicigenerans are determined and the subject is classified as to whether it has a risk of developing CRC. Preferably, in the first stage, higher levels of species of the genus Akkermansia and / or Akkermansia muciniphila, and lower levels of Bacteroides fragilis and / or Bacteroides prevotii are associated with CRC, and in the second stage, higher levels of species of the genus Negativibacillus and / or Bacteroides coprocola, and / or lower levels of Bacteroides caccae and / or Dorea formicigenerans are associated with the risk of developing CRC.

[0054] In the most preferred embodiment, the combination of bacteria whose levels are determined is Akkermansia unclassified.S361 and Akkermansia muciniphila for step (ii) (stage 1), and Bacteroides coprocola and Dorea formicigenerans for step (iii) (stage 2). In a preferred embodiment, in the first and second stages, a first ratio including the centered log-ratio (clr) of the following taxon:

[0055]

Number

[0056]

Number

[0057] In another embodiment of the first aspect of the present invention, when the sample is FIT-negative, the bacterial taxa are: Allistipes putredinis, Anaerostipes hadrus, Bacteroides coprocola, Bacteroides eggerthii, Bifidobacterium animalis, Bifidobacterium bifidum, Bifidobacterium longum, Blautia massiliensis, Blautia obeum, Coprococcus comes, Coprococcus eutactus, Dorea longicatena, Fusobacterium necrophorum, Parvimonas micra, Peptostreptococcus stomatis, Solobacterium moorei, Bifidobacterium_unclassified.S5, Adlercreutzia_unclassified.S168, Porphyromonas_unclassified.S30, Paraprevotella_unclassified.S182, Prevotella_unclassified.S33, Parvimonas_unclassified.S67, Coprococcus_unclassified.S223, Dorea_unclassified.S225, Eisenbergiella_unclassified.S226, Lachnoclostridium_unclassified.S77, Peptococcus_unclassified.S249, Peptostreptococcus_unclassified.S87, Flavonifractor_unclassified.S265, GCA.900066225_unclassified.S267, Negativibacillus_unclassified.S269, Oscillospira_unclassified.S271, Lachnospiraceae_UCG.008_unclassified.S281, Erysipelotrichaceae_UCG.003_unclassified.S297, unclassified Faecalitalea. S300, unclassified unclassified. S306; unclassified Acidaminococcus. S307, unclassified Fusobacterium. S106, unclassified Desulfovibrio. S323, Blautia stercoris, Butyrivibrio crossotus, Parabacteroides distasonis, Roseburia inulinivorans, Sellimonas intestinalis, unclassified Olsenella. S24, unclassified Odoribacter. S27, unclassified Weissella. S204, unclassified Streptococcus. S55, unclassified Christensenellaceae_R.7_group. S209, unclassified Ruminospiraceae_UCG.010. S242, unclassified Marvinbryantia. S244, unclassified Intestinibacter. S818, unclassified Lachnospiraceae_NK4A214_group. S277, unclassified Lachnospiraceae_UCG.005. S92, unclassified Lachnospiraceae_UCG.014. S94, unclassified Veillonella. S104 and unclassified Akkermansia. S361, and is selected from the group consisting of.

[0058] In a preferred embodiment, the taxon is selected from the group consisting of unclassified Fusobacterium. S106, unclassified Peptostreptococcus. S87, unclassified Erysipelotrichaceae_UCG.003. S297, Allistipes putredinis, unclassified Prevotella. S33, unclassified Akkermansia. S361, Coprococcus comes, Bifidobacterium longum. Preferably, any combination of these taxa is used to classify a subject as to whether they have a risk of developing CRC.

[0059] In a more preferred embodiment, in the first stage, the levels of Fusobacterium unclassified.S106, Peptostreptococcus unclassified.S87, Erysipelotrichaceae_UCG.003 unclassified.S297, and Allistipes putredinis are determined, and in the second stage, the levels of Prevotella unclassified.S33, Akkermansia unclassified.S361, Coprococcus comes, and Bifidobacterium longum are determined to classify the subject as to whether they have a risk of developing CRC.

[0060] In another preferred embodiment of the first aspect of the present invention, in step (iii), it is considered that subjects classified into the cohort of subjects having a risk of developing CRC require a colonoscopy, and in step (iii), it is not considered that subjects not classified into the cohort of subjects having a risk of developing CRC require a colonoscopy.

[0061] The computer algorithm in step (iii) of the method of the present disclosure is selected from the group consisting of an artificial intelligence algorithm, a machine learning algorithm, and a trained neural network algorithm. Preferably, the computer algorithm is a trained neural network algorithm.

[0062] In a second aspect, the present invention relates to (a) a reagent for carrying out a method for determining the presence or abundance of bacteria in a fecal sample for determining the levels of two or more bacterial taxa in step (i) of the method of the previous embodiment, and (b) a computer program comprising instructions that cause a computer to perform steps (ii) and (iii) of the method of the present invention when the program is executed by the computer and relates to a kit comprising the same.

[0063] In a preferred embodiment, in step (a), the levels of three or more taxa in step (i) of the method are determined. In a preferred embodiment, the kit (a) Reagents for carrying out a method for determining the presence or abundance of bacteria in a fecal sample for determining the levels of two or more bacterial taxa in step (i) of the method of the second aspect of the present invention, and (b) A computer program stored on a computer-readable data carrier or chip, which, when executed by a computer, includes instructions for causing the computer to perform steps (ii) and (iii) of the method of the present invention comprising.

[0064] In a preferred embodiment, in step (a), the levels of three or more taxa in step (i) of the method are determined.

[0065] In a more preferred embodiment, the reagent is for performing 16S rRNA gene sequencing.

Example

[0066] The examples given below are for illustrative purposes only and do not limit the present invention described above in any way. Example 1: Sample collection and subjects A total of 2,889 FIT-positive (>20 μg hemoglobin / g feces) and 246 FIT-negative (<20 μg hemoglobin / g feces) samples from the Catalan CRC Screening Program were analyzed. A summary of the distribution of FIT-positive samples across several characteristics is shown in Table 1.

[0067]

Table 2

[0068] The collected metadata included six different clinical variables for each sample, including the diagnosis after colonoscopy evaluation (Table 2), the number of polyps, the FIT value (μg hemoglobin / g feces), the hospital where the sample was collected, and the gender and age of the donor. The colonoscopy diagnoses examined were: negative (N), colorectal cancer (CRC), and various lesions that may be associated with the development of colorectal cancer: carcinoma in situ (CIS), high-risk lesions (HRL), intermediate-risk lesions (IRL), low-risk lesions (LRL), and lesions not associated with risk (LNAR) (23). In addition, the samples were classified into two groups according to the clinical relevance of the colonoscopy-based diagnosis (24). CRC, CIS, HRL, and IRL were considered clinically relevant colonoscopies (CR), and N, LNAR, and LRL were considered clinically irrelevant colonoscopies (non-CR).

[0069]

Table 3

[0070] Example 2: DNA Extraction and 16S rRNA Sequencing In the following, 16S rRNA gene sequencing was used as a method for the identification, classification, and quantification of bacterial taxa in complex biological mixtures such as fecal samples. However, those skilled in the art will recognize that other analytical methods, such as polymerase chain reaction (PCR), PCR multiplexing, "next-generation sequencing" (NGS), RNA panels, proteomics, gas chromatography / mass spectrometry, and liquid chromatography / mass spectrometry, may also be used if suitable or desirable.

[0071] 500 μl of the FIT sample was prepared in a test tube and stored at -80 °C until further processing. DNA was extracted using the DNeasy PowerLyzer PowerSoil Kit (Qiagen, ref. QIA12855) according to the manufacturer's instructions. The extraction tubes were agitated twice in a 96-well plate at 30 Hz / s for 5 minutes using a Tissue lyser II (Qiagen). 4 μl of each DNA sample was used to amplify the V3-V4 region of the bacterial 16S ribosomal RNA gene using the following universal primers in limited cycle PCR. V3-V4-Forward (5'-TCGTCGGCAGCGTCAGATGTGTATAAGAGACAGCCTACGGGNGGCWGCAG-3') and V3-V4-Reverse (5'-GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAGGACTACHVGGGTATCTAATCC-3')

[0072] To prevent uneven base composition in further MiSeq sequencing, the sequencing phase was shifted by adding a variable number of bases (0 - 3) as spacers to both the forward and reverse primers (a total of 4 forward and 4 reverse primers were used). PCR was performed in a 10 μl volume reaction using Kapa HiFi HotStart Ready Mix (Roche, ref. KK2602) at a primer concentration of 0.2 μM. The cycle conditions were an initial denaturation at 95 °C for 3 minutes, followed by 25 cycles of 95 °C for 30 s, 55 °C for 30 s, and 72 °C for 30 s, and finally a final extension step at 72 °C for 5 minutes.

[0073] After the first PCR step, water was added to a total volume of 50 μl, and the reaction was purified at a 0.9X ratio using AMPure XP beads (Beckman Coulter) according to the manufacturer's instructions. The PCR product was eluted from the magnetic beads with 32 μl of buffer EB (Qiagen), and 30 μl of the eluate was transferred to a new 96-well plate. The primers used in the first PCR contained overhangs that allowed the addition of full-length Nextera adapters with barcodes for multiplex sequencing in the second PCR step, resulting in a sequencing-ready library. To do so, 5 μl of the first amplification product was used as a template for a second PCR with a final volume of 50 μl of Nextera XT v2 adapter primers using the same PCR mix and temperature profile as the first PCR, but for only 8 cycles. After the second PCR, 25 μl of the final product was used for purification and normalization according to the manufacturer's protocol with the SequalPrep normalization kit (Invitrogen). The library was eluted in 20 μl and pooled for sequencing.

[0074] In an ABI 7900HT real-time cycler (Applied Biosystems), the final pool was quantified by qPCR using the Kapa library quantification kit for Illumina Platforms (Kapa Biosystems). Sequencing was performed on an Illumina MiSeq with 2 × 300 bp reads, using v3 chemistry at a loading concentration of 18 pM. To increase sequence diversity, 10% of the PhIX control library was added.

[0075] Two bacterial mock communities were obtained from BEI Resources of the Human Microbiome Project (HM-276D and HM-277D), each containing genomic DNA of ribosomal operons from 20 bacterial species (25). Mock DNA was amplified and sequenced in the same manner as all other FIT samples. Negative controls for the DNA extraction and PCR amplification steps were also included in parallel, using the same conditions and reagents. These negative controls did not exhibit visible bands or quantifiable amounts of DNA by Bioanalyzer, while all of our samples clearly showed visible bands after 25 cycles.

[0076] For the FIT-positive group, an average of 56,219.03 selected reads per sample were obtained, which included a total of 376 assigned taxa. The phyla Bacteroidetes and Firmicutes were the most representative phyla, and the 10 most abundant genera, in this order, were: Bacteroides, Faecalibacterium, Prevotella, Blautia, F. Lachnospiraceae_UCG, Ruminococcus, Agathobacter, Bifidobacterium, Alistipes, and Akkermansia (Figure 3). These results are consistent with previous studies using fecal samples (49 - 53). The similarity of the microbiome profiles obtained from FIT and fecal samples was also confirmed by comparing the data from the 5 individuals included in this study, for which whole-genome shotgun Illumina data and Ion-Torrent V2 - 4, V6 - 8 16S profiling data of feces were available (35) (Figure 4).

[0077] Example 3. Microbiome Analysis For each of the separate sequencing runs, amplicon sequence variant (ASV) tables were obtained using the dada2 (v.1.10.1) pipeline (27). The plotQualityProfile function of dada2 was used to examine the quality profiles of the forward and reverse sequencing reads, and according to these plots, the filterAndTrim function was used to select and trim low-quality sequencing reads. A matrix with learned error rates was obtained by the learnErrors dada2 function.

[0078] Dereplication (combining identical sequencing reads into unique sequences) and sample inference (from the matrix of estimated learning error rates) were performed, paired reads were merged to obtain complete denoised sequences, and chimeric sequences were excluded from these. Classification was assigned to the ASVs by mapping to the SILVA 16s rRNA database (v.132) (28). Negative controls (non-template samples) and positive controls (mock microbial communities containing a mixture of 20 strains at known ratios) were sequenced and analyzed in each of the runs to determine the potential contamination background and to evaluate the accuracy of the pipeline. ASV and classification tables were obtained separately for each run and then the results were merged. Samples and controls without metadata information were discarded in further analysis.

[0079] The phylogenetic tree was reconstructed by using phangorn (v.2.5.5) (29) and the Decipher R package (v2.10.2) (30), which was integrated with the merged ASV and taxonomy table, and their assigned metadata created a phyloseq (v.1.26.1) object (31). The estimate_richness function of the phyloseq package was used to characterize alpha diversity metrics, including Observed index, Shannon, Simpson, InvSimpson, PD Chao1, ACE, and standard error measures such as se.Chao1 and se.ACE. Faith's phylogenetic diversity, an alpha diversity metric that incorporates the branch lengths of the phylogenetic tree, was calculated using the picante package (v.1.8.1).

[0080] In addition, various distance metrics were calculated based on differences in taxonomic composition between samples using the Phyloseq and Vegan (v.2.5 - 6) packages (Oksanen et al. 2019, Vegan: Community Ecology Package. https: / / CRAN.R-project.org / package=vegan). These metrics include Jensen-Shannon divergence (JSD), weighted Unifrac, unweighted unifrac, Bray-Curtis dissimilarity, Jaccard, and Canberra. The Aitchison distance between samples was also calculated using the cmultRepl and codaSeq.clr functions of the CodaSeq (v.0.99.6) (32) and zCompositions (v.1.3.4) (33) packages. Normalization was performed by transforming the counts to centered log-ratios (clr) (34). Centered log-ratios are a transformation of raw counts to make samples comparable, taking into account the compositional nature of microbiome data. This is the application of logarithms to the ratio of observed frequencies and their geometric mean. Prior to this transformation, multiplicative simple zero replacement implemented in the cmultRepl function of the zCompositions package (display method = "CZM") was performed. Clr can result in both positive and negative values. Samples with less than 1000 reads and taxa that did not appear in most samples and had low abundances were selected and removed. Finally, taxa at each taxonomic rank were aggregated to study trends at various taxonomic depths.

[0081] Example 4. Statistical Analysis Using the seven distance metrics mentioned above, we performed Permutational Multivariate Analysis of Variance (PERMANOVA) using the adonis function of the Vegan R package (v. 2.5 - 6) to determine the association between clinical variables and the overall microbial composition of the samples. Diagnosis, gender, and age variables were considered as covariates. We also applied the ANOSIM (Analysis of Similarities) test using the anosim function of the Vegan R package to determine the differences between and within groups.

[0082] Using the linear model implemented in the R package lme4 (v. 1.1 - 21)(41), we performed a specific abundance analysis using the clr data for various taxonomic classes across various clinical variables. A linear model was constructed with diagnosis (Dx), gender, age, number of polyps, and hospital and FIT value (only for FIT-positive samples) as fixed effects, and sequencing run as a random effect that could be a factor for batch effects. Considering all diagnoses, and further, by changing all other diagnoses to "others", CRC vs. non-CRC samples were compared to evaluate this linear model. To determine the differences between samples with CR or non-CR colonoscopy defined previously (Table 2), a second linear model was applied considering a variable called risk as a fixed effect instead of diagnosis.

[0083] Analysis of Variance (ANOVA) was applied to determine the significance for each of the fixed effects included in the model using the Car R package (v. 3.0 - 6)(42). To determine specific differences between groups, multiple comparisons were performed on the results obtained from the linear model using the Tukey test in the glht function of the multcomp R package (v. 1.4 - 12)(43). Bonferroni was applied for multiple testing correction, and statistical significance was defined as a p-value less than 0.05. In addition, the selbal package (v. 0.1.0)(44) was used to study the groups of taxa (balances) with potential predictive power for CRC status in FIT-positive samples.

[0084] Example 5. Machine learning classification A further aspect of the present invention relates to a novel two-step classification mechanism that enhances the inclusion of colorectal cancer and clinically relevant cases and prioritizes the reduction of false negatives over false positives. Feature selection is based on differential analysis: the combination of the centered log-ratio (Clr) of the selected taxa with clinical variables (gender, age, and hemoglobin content).

[0085] Briefly, a prediction model based on a two-step classification (Figure 2) was developed using the neural network (NN) algorithm implemented in the caret package (v.6.0 - 85) (47). At each step, 75% of the data was randomly trained by 10-fold cross-validation and tested on the remaining samples. To avoid "lucky" splits and evaluate the variability of the prediction performance, the process was repeated 100 times. Feature selection was performed based on specific abundance results incorporating the hemoglobin content, age, and gender variables, including taxa found to have significant differences in abundance in our invention. Samples with missing values for the metadata under consideration were excluded. Taxa abundances were included as clr. The two-step classification mechanism proceeds as follows: In the first step, the method classifies CRC versus non-CRC samples. Samples classified as non-CRC in the first step are subjected to a second model that classifies CR versus non-CR samples, including misclassified CRCs, to improve sensitivity. At the end of the two-step classification, the average percentage of misclassified CRC and CR samples was calculated to evaluate the performance of the model.

[0086] To validate this strategy, a model trained with all CRIPREV samples was built and tested in two independent datasets: a cohort from the USA (48) and 100 pre-samples from the same Catalan screening. For the USA cohort, Catalan hemoglobin thresholds (hemoglobin > 20 μg / g feces) were applied to select FIT-positive samples for inclusion in the validation. Their raw data was processed according to exactly the same methodology disclosed in this document (see microbiome analysis, materials and methods). Perhaps because that study used only the V4 region of the 16S rRNA gene as compared to V3-V4 in the present invention, Bacteroides fragilis was unfortunately not assigned. The design and construction of the classification mechanism is described in more detail below.

[0087] Example 6. Design and construction of the classification mechanism for FIT-positive samples The two-step classification mechanism proceeds as follows: In the first step, the method classifies colorectal cancer (CRC) versus non-CRC samples. Samples classified as non-CRC in the first step are subjected to a second model that classifies clinically relevant (CR) versus non-CR samples. Clinically relevant refers to the grouping of colonoscopy diagnoses ranging from intermediate-risk lesions to high-risk lesions and CRC that require clinical follow-up.

[0088] The inputs used by the model are three data related to FIT (gender, age, and FIT value), and a normalized and filtered amplicon sequence variant (ASV) table (obtained from sequencing data as described herein). ASV is limited to a selected set of taxa (the optimal model for inclusion of CR cases included the four taxa described herein, but other combinations from the related taxa identified herein could also have been used).

[0089] The model was trained using approximately 2,800 sample data from the CRIPREV project. In the first-stage classification, 10-fold cross-validation was performed, and the best model was the one used to predict the independent test set. Some of the model's specificities: Method: Implemented in the caret package by using the nnet, train function (v.6.0 - 85).

[0090] MaxNWts (maximum allowable number of weights): 2000.

[0091] Weights: We varied the weights to penalize more heavily against the expected minority classes: 0.75 for CRC and 0.25 for others.

[0092] After prediction, a confusion matrix was constructed, and the samples classified as others were subjected to a second classification detailed below. If the model classified all samples as CRC (AUC: 0.5, null permissibility for classification), all samples were subjected to the second classification. In the second stage, the same training set was used, but CRC samples were excluded from this training set, and intermediate-risk and high-risk lesions were labeled as clinically relevant. The model was trained to recognize clinically relevant samples. For this, 10-fold cross-validation was performed, and the best model was the one used to predict the independent test set. Some of the model's specificities: Method: Implemented in the caret package (v.6.0 - 85) using nnet.

[0093] MaxNWts (maximum allowable number of weights): 2000.

[0094] Weights: We varied the weights to penalize more heavily against the expected minority classes: 0.60 for clinically relevant samples and 0.40 for clinically irrelevant samples.

[0095] Performance Evaluation To evaluate the strategy, three independent strategies were applied. 1. CRIPREV Sample At each stage, a model was constructed that trained on a random 75% of the data by 10-fold cross-validation and tested on the remaining samples. To avoid "lucky" splits and to evaluate the variability of the prediction performance, the process was repeated 100 times. Feature selection was performed based on specific abundance results that included taxa found to have abundances that were significantly different in our invention and incorporated FIT-values, age, and gender variables. Samples with missing values for the metadata under consideration were excluded.

[0096] 2. Independent Research The model was trained on 100% of the CRIPREV data and tested for performance on an independent dataset cohort of 135 samples from the USA from previously published studies. In the immediately preceding study, the Catalan hemoglobin threshold (>20 μg hemoglobin / g feces) was applied to select FIT-positive samples for inclusion in the validation. Their raw data was processed according to exactly the same methodology described in this document. Unfortunately, probably because that study used only the V4 region of the 16S rRNA gene as compared to V3-V4 in the present invention, Bacteroides fragilis was not assigned.

[0097] 3. Samples newly obtained from the Catalan Screening program Using the model trained on 100% of the CRIPREV data, the performance on an independent dataset of 100 additional FIT-positive samples from Catalan CRC screening was tested. Non-limiting examples of thresholds useful for diagnosing CRC using the method of the present invention are described below.

[0098] For various thresholds: (I) For each of the stages, the ratio calculated from all dysregulated taxa (overrepresented taxa / underrepresented taxa) (II) For each stage, the ratio calculated from a 4-taxa panel (overrepresented taxa / underrepresented taxa) (III) The average of the key species in clinically relevant groups It was determined considering

[0099] The best results so far have been considering a second option, a filter based on thresholds using two different ratios, each with a 4 - taxon panel for each step. Two ratios were calculated from the amplicon sequence variant table normalized by centered log - ratio (clr): The first ratio:

[0100] [Number] The second ratio:

[0101] [Number] A filter with the condition of having a first ratio greater than - 0.5512273 (based on the mean of the first ratio in CRC patients) or a second ratio greater than 0 was applied. The results obtained were: Using the CRIPREV dataset: Percentage of clinically relevant samples detected: 85.41 Percentage of CRC samples: 86.57 Percentage of colonoscopies skipped: 14.92 Using the validation dataset: Percentage of clinically relevant samples detected: 81.25 Percentage of CRC samples: 88 Percentage of colonoscopies skipped: 14 It was.

[0102] Alpha and Beta Diversity By calculating alpha and beta diversity metrics, the overall diversity of the microbiota in the samples was quantified. When considering all diagnoses, significant differences (P<0.05) were observed in the observed indicator alpha diversity metric (measuring the number of species per sample), but not when comparing CR vs non-CR samples in particular (Figure 5). For the Shannon and Simpson indices considering differences in abundance, when considering all diagnoses, significant differences were only observed in the Simpson index (assigning a greater weight to dominant species).

[0103] Using distances (beta diversity) between the microbial profiles of the samples, such as Aitchison distance, MDS plots were generated (Figure 5). Clear clustering of samples with the same diagnosis or risk (CR vs non-CR) was not observed. However, using the adonis test and Aitchison distance, considering gender and age as covariates and the sequencing run as a potential source of batch effect, a significant effect of diagnosis was detected (P=0.001). The ANOSIM test also supported a significant but subtle difference between diagnosis groups and greater similarity within groups (R: 0.07463, p-value: 0.001). Generally, this suggests the existence of a significant but subtle difference in the overall microbiota composition among FIT-positive samples with various colonoscopy outcomes.

[0104] Example 7: Microbiota in FIT-positive samples Using comparative analysis, significant differences in the relative abundances of several taxa were detected by various fixed-effect variables (FIT-positive samples). These analyses identified, for example, 34 species whose abundances changed significantly among colonoscopy diagnoses (Table 3 and Figure 7).

[0105]

Table 4

[0106] Based on the observation that CRC was the most definitive diagnosis (Figure 7), CRC vs. non-CRC samples were specifically compared, and 41 species that were specifically abundant were identified (Figure 8A). These included the overrepresentation of Akkermansia muciniphila and species of the genus Akkermansia in CRC compared to non-CRC samples, as well as the underrepresentation of Bacteroides prevotii and Bacteroides fragilis. In addition, using the selbal package for the same comparison (CRC vs. non-CRC), it was determined that the ratio between the species (balances) most associated with the CRC state was given by a decrease (compared to non-CRC samples) in the ratio of a group of taxa including B. fragilis (G1: species of the genus Bifidobacterium, Bacteroides fragilis, Subdoligranulum variabile, and species of the genus Eggerthella) to a second group of taxa including species of the genus Akkermansia (G2: species of the genus Akkermansia, species of the genus Gemella, Peptostreptococcus stomatis, species of the genus Adlercreutzia, and species of the genus Butyrivibrio). Finally, applying the same linear model to the comparison of CR vs. non-CR samples identified 34 species that were specifically abundant (Figure 8B).

[0107] Colorectal polyp (54), a benign tumor that protrudes into the colonic mucus and bulges into the intestinal lumen, has long been identified as a potential precursor to CRC. This disclosure includes 66.82% of samples in which colonoscopy detected the presence of polyps and the number of polyps ranged from 1 to 22. Some CRC samples did not have polyps, while some negative samples had 1 to 3 polyps, and it was observed that some lesions not related to clinically relevant colonoscopy had a significant amount of polyps (1 to 11 polyps). Species whose abundance was significantly correlated with the number of polyps were detected (Table 4).

[0108]

Table 5

[0109] 7.1 Taxonomic Combinations Next, as potential candidates for classification, to determine the possible combinations of taxa included in the list that were found to be specifically enriched by diagnosis (41) and clinical relevance (34), we used our validation set (100 pilot samples from the Catalan screening for colorectal cancer). A total of 27 taxa that intersect between the CRIPREV project and these pilot samples were identified, and these are the taxa included in the results presented here.

[0110] We determined the various combinations of taxa considering the effect sizes observed in our statistical tests (the statistical tests presented here, where they were detected as dysregulated by the variable of interest). We defined the top and bottom taxa from the list and determined subsets of taxa as follows: 4 taxa from the top of the list (50 random combinations) 4 taxa from the bottom of the list (50 random combinations) 4 random taxa (50 random combinations) 2 taxa from the top of the list (all possible combinations) 2 taxa from the bottom of the list (all possible combinations) 1 taxon from the top of the list (all possible combinations) 1 taxon from the bottom of the list (all possible combinations)

[0111] By using 100 preliminary samples from the same local screening, a potential subset of taxa with classification ability (i.e., specifically enriched in the differential analysis test of the present invention) was determined. Considering the effect sizes observed in the statistical tests of the present invention, various combinations of taxa were determined. At each step, the top (with high size effect) and bottom (with low size effect) taxa were defined from the list, and the determination of subsets of taxa was performed as follows: 4 taxa from the top of the list (50 random combinations), 4 taxa from the bottom of the list (50 random combinations), 4 random taxa (50 random combinations), 2 taxa from the top of the list (all possible combinations), 2 taxa from the bottom of the list (all possible combinations), 1 taxon from the top of the list (all possible combinations), and 1 taxon from the bottom of the list (all possible combinations). From Figure 8, it can be seen that both the species of the genus Akkermansia and Akkermansia muciniphila have the largest effect sizes in the group of specifically enriched taxa that are overrepresented in CRC.

[0112] Using the validation set, a total of 948 models were tested. Based on several metrics (AUC1 >= 0.55, specificity > 0.2, AUC2 > 0.5, and specificity2 > 0) for selecting 13.5% (128 / 948) of the models, the models were screened. The strategy that selected more models was the strategy that included the subset of 4 taxa with the largest effect size (Figure 9). The selected models were divided into three grades considering their prediction performance values: Grade 1 included 8 models that had a sensitivity of 100% for CRC, a sensitivity of 96% or more for clinically relevant individuals, and 12% or more of discarded unnecessary colonoscopies. The list of Grade 1 combinations is:

[0113]

Table 6

[0114] Grades 2 and 3 each included the selected combinations of 50 and 70.

[0115] Furthermore, the potential of various 27 taxa was explored by evaluating in how many models each of them appeared (Figure 10, Table 5), and the species of the genus Akkermansia was the taxon that appeared in most of the selected models.

[0116]

Table 7

[0117] Of the 128 selected models, 124 included at least one of the 8 taxa (4 taxa per stage): Akkermansia muciniphila, species of the genus Akkermansia, Bacteroides fragilis and Bacteroides prevotii, Bacteroides coprocola, species of the genus Negativibacillus, Dorea formicigenans or Bacteroides caccae, which are included in the selected models of the present application.

[0118] 7.2 Development of a two-stage machine learning classification mechanism. Evaluation and verification of strategies in FIT positive samples.

[0119] Since samples with different diagnoses presented significant differences in the abundance of various bacterial taxa, a machine learning approach was explored to develop a sample classification mechanism capable of distinguishing samples that are more likely to benefit from colonoscopy procedures (i.e., samples with clinically relevant diagnoses).

[0120] In this regard, false negatives (i.e., individuals with clinically relevant lesions who do not proceed to colonoscopy) are of greater medical interest compared to false positives (individuals without lesions who undergo colonoscopy), so emphasis was placed on achieving high sensitivity as opposed to high precision.

[0121] To derive this prediction mechanism, we explored the effects of using various machine learning algorithms and the use of feature selection to limit the parameters set for all bacterial taxa that showed a significant difference, or only a few of them (see section "Materials and Methods"). When more taxa were included, better AUC and specificity were observed (Table 6). This can be interpreted as a better reduction of the false positive rate. On the other hand, when restricted to only a panel of taxa, better recall and sensitivity for CRC and CR samples were obtained, but insufficient AUC and specificity were also obtained (Table 7). However, in the context of current screening, there still exists a satisfactory reduction of the false positive rate with a good prioritization of relevant cases. The best results were achieved by a two-stage classification mechanism trained to classify CRC samples in the first stage and any CR samples in the second stage. This final classification mechanism takes into account information on the FIT values available from gender, age, and FIT test results, and the abundances from two different subsets of two out of four taxa (first stage: species of the genus Akkermansia, Akkermansia muciniphila, Bacteroides fragilis, and Bacteroides prevotii, and second stage: species of the genus Negativibacillus, Bacteroides coprocola, Bacteroides caccae, and Dorea formicigenerans). This classification mechanism obtained 98.98% sensitivity for CRC samples and 97.98% for clinically relevant samples.

[0122]

Table 8

[0123]

Table 9

[0124] A model with all samples (excluding *Bacteroides fragilis*) was constructed and this strategy was validated by testing it in an independent cohort of 135 FIT-positive samples from the USA. The results of this model, adjusted in the USA cohort, yielded 100% sensitivity for CRC and 98.46% for CR lesions, reducing 20% of unnecessary colonoscopies (A). Validation was also performed by an independent dataset consisting of 100 preliminary samples from the same Catalan Screening, detecting 100% of all CRC samples and 96% of CR samples, with a 12% reduction in false positives (Table 8).

[0125]

Table 10

[0126] Using 100 unbiased preliminary samples from the Catalan Screening, we explored how changes in some parameters of the classification mechanism affect sensitivity and the number of colonoscopies saved. For example, by penalizing the minority class (CR) less in the second stage, a better reduction (26%) of unnecessary colonoscopies was obtained, but at the cost of including fewer CR samples (90%). Similarly, applying a FIT value threshold above which a benefit in colonoscopy is assumed can reduce the number of samples to be tested. For such a threshold passed by 18% of our samples, applying the value 954 μg hemoglobin / g feces (third quartile in CR samples) saves 14% of unnecessary colonoscopies at the end of the process. When we combined both methods, we were able to reach 30% of colonoscopies saved at the cost of a reduction in CR detection (87%). However, in all cases mentioned, 100% of CRC samples were detected. This indicates that the algorithm can be fine-tuned to optimize cost-effectiveness (Figure 11).

[0127] Although certain representative embodiments and details have been shown to illustrate the present invention, it will be apparent to those skilled in the art that various changes and modifications may be made and that the invention may be practiced otherwise than as particularly described and claimed, within the scope of the appended claims.

[0128] Example 8. Microbiome in FIT-negative samples. Further validation of the two-step classification mechanism.

[0129] Differential analysis yielded 39 taxa with significant results (P-value < 0.05) when CRC was compared to others, and in the second step, 42 taxa specifically enriched by CR vs. non-CR.

[0130] Using this dataset, the machine learning classification mechanism strategy was also evaluated according to the same method as the previously shown FIT-positive samples, but in this case considering the taxa found to be specifically enriched. For the first step, it included "unclassified Fusobacterium.S106", "unclassified Peptostreptococcus.S87", "unclassified UCG.003 of Erysipelotrichaceae.S297", "Allistipes putredinis", gender, and age. For the second step, it included "unclassified Prevotella.S33", "unclassified Akkermansia.S361", "Coprococcus comes", "Bifidobacterium longum", gender, and age.

[0131] The best strategy for the two-step classification mechanism (including 4 taxa and clinical variables per step) was applied. The strategy was evaluated by training and testing 100 models (to avoid lucky splits when creating the training and test sets). The obtained results are shown in the following table.

[0132]

Table 11

[0133] The sensitivities for CRC and CR at the end of the procedure were: Sensitivity CRC at the end of the 2-step procedure: 98.38 Sensitivity CR at the end of the 2-step procedure: 95.73913 It was so.

[0134] This indicates that the method of the present invention is predictable for FIT-negative samples. References 1. Bray, F. et al. Global cancer statistics 2018: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J. Clin. 68, 394-424 (2018). 2. Hong, S. N. Genetic and epigenetic alterations of colorectal cancer. Intest Res 16, 327-337 (2018). 3. Valle, L. et al. Update on genetic predisposition to colorectal cancer and polyposis. Mol. Aspects Med. 69, 10-26 (2019). 4. Murphy, N. et al. Lifestyle and dietary environmental factors in colorectal cancer susceptibility. Mol. Aspects Med. 69, 2-9 (2019). 5. Saus, E., Iraola-Guzman, S., Willis, J. R., Brunet-Vega, A. & Gabaldon, T. Microbiome and colorectal cancer: Roles in carcinogenesis and clinical potential. Mol. Aspects Med. 69, 93-106 (2019). 6. Zou, S., Fang, L.& Lee, M.-H. Dysbiosis of gut microbiota in promoting the development of colorectal cancer. Gastroenterol. Rep. 6, 1-12 (2018). 7. Zackular, J. P., Rogers, M. A. M., Ruffin, M. T., 4th & Schloss, P. D. The human gut microbiome as a screening tool for colorectal cancer. Cancer Prev. Res. 7, 1112-1121(2014). 8. Sheng, Q.-S. et al. Comparison of Gut Microbiome in Human Colorectal Cancer in Paired Tumor and Adjacent Normal Tissues. Onco. Targets. Ther. 13, 635-646 (2020). 9. Yu, J. et al. Metagenomic analysis of faecal microbiome as a tool towards targeted non-invasive biomarkers for colorectal cancer. Gut 66, 70-78 (2017). 10. Winawer, S. J. The history of colorectal cancer screening: a personal perspective. Dig. Dis. Sci. 60, 596-608(2015). 11. Young, G. P., Rabeneck, L.& Winawer, S. J. The Global Paradigm Shift in Screening for Colorectal Cancer. Gastroenterology 156, 843-851.e2 (2019). 12. Zou, S., Fang, L. & Lee, M.-H. Dysbiosis of gut microbiota in promoting the development of colorectal cancer. Gastroenterol. Rep. 6, 1-12 (2018). 13. Vega, P., Valentin, F. & Cubiella, J. Colorectal cancer diagnosis: Pitfalls and opportunities. World J. Gastrointest. Oncol. 7, 422-433 (2015). 14. Inici. http: / / www.prevenciocolonbcn.org / ca / . 15. Alix-Panabieres, C. & Pantel, K. Circulating tumor cells: liquid biopsy of cancer. Clin. Chem. 59, 110-118 (2013). 16. Bettegowda, C. et al. Detection of circulating tumor DNA in early- and late-stage human malignancies. Sci. Transl. Med. 6, 224ra24 (2014). 17. Duran-Sanchon, S. et al. Identification and Validation of MicroRNA Profiles in Fecal Samples for Detection of Colorectal Cancer. Gastroenterology 158, 947-957.e4 (2020). 18. Nannini, G., Meoni, G., Amedei, A. & Tenori, L. Metabolomics profile in gastrointestinal cancers: Update and future perspectives. World J. Gastroenterol. 26, 2514 - 2532 (2020). 19. Thomas, M. et al. Genome-wide Modeling of Polygenic Risk Score in Colorectal Cancer Risk. Am. J. Hum. Genet. 107, 432 - 444 (2020). 20. Janney, A., Powrie, F. & Mann, E. H. Host-microbiota maladaptation in colorectal cancer. Nature 585, 509 - 517 (2020). 21. Sepich-Poore, G. D. et al. The microbiome and human cancer. Science 371, eabc4552 (2021). 22. Quintero, E. et al. Colonoscopy versus fecal immunochemical testing in colorectal-cancer screening. N. Engl. J. Med. 366, 697 - 706 (2012). 23. Atkin, W. S. et al. European guidelines for quality assurance in colorectal cancer screening and diagnosis. First Edition--Colonoscopic surveillance following adenoma removal. Endoscopy 44 Suppl 3, SE151 - 63 (2012). 24. Click, B., Pinsky, P. F., Hickey, T., Doroudi, M. & Schoen, R. E. Association of Colonoscopy Adenoma Findings With Long-term Colorectal Cancer Incidence. JAMA 319, 2021-2031(2018). 25. Willis, J. R. et al. Citizen science charts two major ‘stomatotypes’ in the oral microbiome of adolescents and reveals links with habits and drinking water composition. Microbiome 6, 218 (2018). 26. Willis, J. R. et al. Oral microbiome in down syndrome and its implications on oral health. J. Oral Microbiol. 13, 1865690 (2020). 27. Callahan, B. J. et al. DADA2: High resolution sample inference from amplicon data. doi:10.1101 / 024034. 28. Quast, C. et al. The SILVA ribosomal RNA gene database project: improved data processing and web-based tools. Nucleic Acids Res. 41, D590-6 (2013). 29. Schliep, K. P. phangorn: phylogenetic analysis in R. Bioinformatics vol. 27 592-593 (2011). 30. Wright, E., Erik & Wright, S. Using DECIPHER v2.0 to Analyze Big Biological Sequence Data in R. The R Journal vol. 8 352 (2016). 31. McMurdie, P. J. & Holmes, S. phyloseq: an R package for reproducible interactive analysis and graphics of microbiome census data. PLoS One 8, e61217 (2013). 32. Gloor, G. B. & Reid, G. Compositional analysis: a valid approach to analyze microbiome high-throughput sequencing data. Can. J. Microbiol. 62, 692 - 703 (2016). 33. Palarea-Albaladejo, J. & Martin-Fernandez, J. A. zCompositions - R package for multivariate imputation of left-censored data under a compositional approach. Chemometrics and Intelligent Laboratory Systems vol. 143 85 - 96 (2015). 34. Gloor, G. B., Macklaim, J. M., Pawlowsky-Glahn, V. & Egozcue, J. J. Microbiome Datasets Are Compositional: And This Is Not Optional. Frontiers in Microbiology vol. 8 (2017). 35. Mas-Lloret, J. et al. Gut microbiome diversity detected by high-coverage 16S and shotgun sequencing of paired stool and colon sample. Sci Data 7, 92 (2020). 36. Babraham Bioinformatics - FastQC A Quality Control tool for High Throughput Sequence Data. https: / / www.bioinformatics.babraham.ac.uk / projects / fastqc / . 37. Bolger, A. M., Lohse, M. & Usadel, B. Trimmomatic: a flexible trimmer for Illumina sequence data. Bioinformatics vol. 30 2114 - 2120 (2014). 38. Wood, D. E., Lu, J. & Langmead, B. Improved metagenomic analysis with Kraken 2. Genome Biol. 20, 257(2019). 39. Lu, J., Breitwieser, F. P., Thielen, P. & Salzberg, S. L. Bracken: estimating species abundance in metagenomics data. PeerJ Computer Science vol. 3 e104 (2017). 40. Hmisc: Harrell Miscellaneous. https: / / CRAN.R-project.org / package = Hmisc. 41. Bates, D., Machler, M., Bolker, B. & Walker, S. Fitting Linear Mixed-Effects Models Using lme4. J. Stat. Softw. 67, (2015). 42. Fox, J., Friendly, M. & Weisberg, S. Hypothesis Tests for Multivariate Linear Models Using the car Package. The R Journal vol. 5 39 (2013). 43. Hothorn, T., Bretz, F. & Westfall, P. Simultaneous inference in general parametric models. Biom. J. 50, 346 - 363 (2008). 44. Rivera - Pinto, J. et al. Balances: a New Perspective for Microbiome Analysis. mSystems 3, (2018). 45. Kurtz, Z. D. et al. Sparse and Compositionally Robust Inference of Microbial Ecological Networks. PLOS Computational Biology vol. 11 e1004226 (2015). 46. Woloszynek, S. et al. Exploring thematic structure and predicted functionality of 16S rRNA amplicon data. PLOS ONE vol. 14 e0219235 (2019). 47. Kuhn, M. Building Predictive Models in R Using the caret Package. J. Stat. Softw. 28, (2008). 48. Baxter, N. T., Koumpouras, C. C., Rogers, M. A. M., Ruffin, M. T., 4th & Schloss, P. D. DNA from fecal immunochemical test can replace stool for detection of colonic lesions using a microbiota-based model. Microbiome 4, 59 (2016). 49. Abrahamson, M., Hooker, E., Ajami, N. J., Petrosino, J. F. & Orwoll, E. S. Successful collection of stool samples for microbiome analyses from a large community-based population of elderly men. Contemp Clin Trials Commun 7, 158-162 (2017). 50. Feng, Y. et al. An examination of data from the American Gut Project reveals that the dominance of the genus Bifidobacterium is associated with the diversity and robustness of the gut microbiota. Microbiologyopen 8, e939 (2019). 51. Yang, T.-W. et al. Enterotype-based Analysis of Gut Microbiota along the Conventional Adenoma-Carcinoma Colorectal Cancer Pathway. Sci. Rep. 9, 1-13 (2019). 52. Sweeney, T. E. & Morton, J. M. The human gut microbiome: a review of the effect of obesity and surgically induced weight loss. JAMA Surg. 148, 563-569 (2013). 53. Rinninella, E. et al. What is the Healthy Gut Microbiota Composition? A Changing Ecosystem across Age, Environment, Diet, and Diseases. Microorganisms 7, (2019). 54. Shussman, N. & Wexner, S. D. Colorectal polyps and polyposis syndromes. Gastroenterol. Rep. 2, 1-15 (2014).

Claims

1. 1. A method for diagnosing a subject as suffering from colorectal cancer (CRC) or for classifying a subject in a patient cohort as having a higher risk of developing CRC, comprising: (i) determining the levels of three or more bacterial taxa in a fecal sample isolated from a subject; (ii) as a first step, classifying CRC versus non-CRC samples using a computer algorithm two or more bacterial taxa that are differentially abundant in CRC samples compared to non-CRC samples, the hemoglobin content of the samples, and the age and sex of the donor; (iii) as a second step, using a computer algorithm to classify the samples classified as non-CRC in the first step into CR samples and non-CR samples using two or more bacterial taxa that are differentially abundant in clinically relevant (CR) samples compared to non-CR samples, the hemoglobin content of the samples, and the age and sex of the donor, where CR includes intermediate-risk lesions, high-risk lesions, carcinoma in situ (CIS), and colorectal cancer (CRC). Including, The three or more bacterial taxa in step (i) are selected from the group consisting of Hungatellae spp., Collinsella spp., Tyzzerella spp., Phascolarctobacterium succinatutens, Lactobacillus spp., Akkermansia spp., Akkermansia muciniphila, O. mollicutes_RF39.UCF, Ruminococcaceae_UCG.UCF, and Ruminococcaceae_UCG.UCF. Ruminococceae_UCG.002 spp., Ruminococceae_UCG.0010 spp., Odoribacter spp., O. Rhodospirillares.UCF, Victivallis spp., Ruminococceae_UCG. 005 spp. (Ruminococcaceae_UCG.005 spp.), Negativibacillus spp., Christensenellaceae_R.7_group spp., Oxalobacter spp., Butyrivibrio spp., Family_XIII_UCG. 001 species, Gemella spp., Peptostreptococcus spp., Pediococcus spp., Lactobacillus vaginalis, Enorma massiliensis, Megamonas funiformis, Peptostreptococcus anaerobius, Peptoniphilus lacrimalislacrimalis, Lactobacillus oris, Alloscardovia omnicolens, Allisonella histaminiformans, Acidaminococcus fermatans, Collinsella bouchesdurhonensis, Corynebacterium spp., Veillonella dispar, Ezakiella spp. spp., O. Chloroplast. UCF, Sphingomonas spp., Dialister succinatiphilus, Finegoldia magna, Bacteroides coprofilus, Eggerthella spp., Acidaminococcus spp., Enterococcus spp., Sutterella watswartensis wadsworthensis, Bacteroides fragilis, Bacteroides plebeius, Bacteroides coprocola, Bifidobacterium longum, Bilophila spp., Parabacteroides merdae, DTU08 spp., Oscillibacter spp., Parabacteroides goltzeinii, goldsteinii), Parabacteroides spp., Bacteroides spp.spp.), Coprobacter secundus, Prevotella timonensis, Streptococcus parasanguinis, Peptostreptococcus anaerobius, Streptococcus sobrinus, Lachnospiraceae_FCS020 group bacteria, Bifidobacterium dentium dentium), Porphyromonas spp., Lachnospiraceae UCC. Lachnospiraceae_UCC.008 spp., Enterobacter spp., Hungatell a hathewayi, Ezakiella spp., Leuconostoc spp., Parabacteroides johnsonii, Bacteroides finegoldii, Eisenbergiella spp., Alistipes finegoldii, finegoldii), F. Erysipelotrichaceae UCG, Dorea formicigenerans, Bacteroides caccae, Fusobacterium unclassified S106, Peptostreptococcus unclassified S87, Erysipelotrichaceae UCG.003 unclassified. S297, Alistipes putredinis, Prevotella unclassified S33 and Coprococcus comes, method.

2. 2. The method of claim 1, wherein the fecal sample is a fecal immunochemical test (FIT) sample.

3. If the sample is FIT positive, the bacterial taxon is selected from the group consisting of Akkermansia spp., Akkermansia muciniphila, Bacteroides fragilis, Bacteroides plebeius, Negativibacillus spp., Bacteroides coprocola, Bacteroides caccae, and Dorea formisigenerans.

3. The method of claim 2, wherein the bacterium is selected from the group consisting of:

4. In the first step of the method, levels of Akkermansia spp., Akkermansia muciniphila, Bacteroides fragilis, and Bacteroides plebeius are determined to classify the subject as having CRC, and in the second step, levels of Negativibacillus spp., Bacteroides coprocola, Bacteroides caccae, and Dorea formisigenerans are determined to classify the subject as having CRC.

4. The method of claim 3, wherein the level of CRC (C. formicigenerans) is determined to classify the subject as being at risk for developing CRC.

5. In the first stage, higher levels of Akkermansia spp. and / or Akkermansia muciniphila and lower levels of Bacteroides fragilis and / or Bacteroides plebeius are associated with CRC, and in the second stage, higher levels of Negativibacillus spp. and / or Bacteroides coprocola and / or lower levels of Bacteroides caccae are associated with CRC.

5. The method of claim 4, wherein P. caccae and / or Dorea formicigenerans are associated with the risk of developing CRC.

6. In the first and second stages, A first ratio comprising the central logarithmic ratios (clr) of the following taxa: [Equation 1] is greater than -0.5512273, Second Ratio [Equation 2] If is greater than 0, the subject is diagnosed as being at risk for developing CRC; The method of claim 5.

7. If the sample is FIT negative, the bacterial taxon is Fusobacterium unclassified S106, Peptostreptococcus unclassified S87, Erysipelotrichaceae UCG.003 unclassified S297, Alistipes putredinis, Prevotella unclassified S33, Akkermansia unclassified S40, Akkermansia unclassified S50, Akkermansia unclassified S60, Akkermansia unclassified S70, Akkermansia unclassified S80, Akkermansia unclassified S90, Akkermansia unclassified S1 ...

7. The method of claim 6, wherein the bacterial strain selected from the group consisting of S361, Coprococcus comes, and Bifidobacterium longum is selected and determined to classify the subject as having a risk of developing CRC.

8. In the first stage, higher levels of Fusobacterium unclassified S106, Peptostreptococcus unclassified S87, Erysipelotrichaceae UCG.003 unclassified S297, and Alistipes putredinis were detected, and in the second stage, higher levels of Prevotella unclassified S33, Akkermansia unclassified S34, and Akkermansia unclassified S35 were detected.

8. The method of claim 7, wherein the expression of S361, Coprococcus.comes, and Bifidobacterium.longum is determined to classify the subject as being at risk for developing CRC.

9. 9. The method of any one of claims 1 to 8, wherein in step (iii) subjects classified into the cohort of subjects at risk of developing CRC are considered to require colonoscopy, and wherein in step (iii) subjects not classified into the cohort of subjects at risk of developing CRC are not considered to require colonoscopy.

10. 10. The method of claim 1, wherein the computer algorithm is selected from the group consisting of an artificial intelligence algorithm, a machine learning algorithm, and a trained neural network algorithm.

11. The method of claim 10 , wherein the computer algorithm is a trained neural network algorithm.

12. (a) reagents for carrying out a method for determining the presence or abundance of bacteria in a fecal sample, for determining the level of two or more bacterial taxa in step (i) of the method of claim 1; and (b) a computer program stored on a computer-readable data carrier or chip, comprising instructions that, when said program is executed by a computer, cause said computer to carry out steps (ii) and (iii) of the method of claim 1. Kit including:

13. 13. The kit of claim 12, wherein the reagents are for performing 16S rRNA gene sequencing.