Bacterial biomarkers for colorectal cancer diagnosis

A consortium of bacterial biomarkers, including Fusobacterium, Parvimonas, Bacteroides, and Faecalibacterium, identified via 16S rRNA gene sequencing, enhances CRC detection accuracy from 0.65 to 0.8, offering a non-invasive early detection method.

WO2025158075A1PCT designated stage Publication Date: 2025-07-31SERVIZO GALEGO DE SAUDE +2
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/051982
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-26
Filing Date
2025-01-27
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Current colorectal cancer (CRC) screening methods, such as fecal occult blood tests, are non-specific and invasive, often detecting the disease at advanced stages, leading to low survival rates and high mortality. There is a need for more specific, non-invasive biomarkers to enhance early detection of CRC.

Method used

The use of a consortium of bacterial biomarkers, including Fusobacterium, Parvimonas, Bacteroides, and Faecalibacterium, identified through 16S rRNA gene sequencing, particularly using Oxford Nanopore Technologies' long-read sequencing, to analyze fecal samples for early detection of CRC.

Benefits of technology

This approach improves the diagnostic accuracy for CRC from an AUC of 0.65 to 0.8, providing a non-invasive method for early detection of colorectal neoplasia, potentially reducing mortality by identifying precancerous stages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000042_0001
    Figure IMGF000042_0001
  • Figure 00000073_0000
    Figure 00000073_0000
  • Figure 00000074_0000
    Figure 00000074_0000
Patent Text Reader

Abstract

The present invention relates to a bacterial biomarker or to a combination of bacterial biomarkers in intestinal or stool samples which on one hand help in the diagnosis of colorectal cancer, and on the other hand predict the progression of colorectal cancer, as well as the response to treatment of a subject suffering from this disease.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Bacterial Biomarkers for colorectal cancer diagnosis

[0002] Technical Field of the Invention

[0003] The present invention relates to a bacterial biomarker or to a combination of bacterial biomarkers in intestinal or stool samples which on one hand help in the diagnosis of colorectal cancer, and on the other hand predict the progression of colorectal cancer, as well as the response to treatment of a subject suffering from this disease.

[0004] State of the Art

[0005] Colorectal cancer (CRC) is one of the most common types of cancer worldwide, after breast, prostate and lung cancers (considering both sexes and all ages). Nevertheless, it is the second type of cancer that causes the most deaths worldwide, with 935173 deaths in 2020 [1], In Spain, CRC is a malignant neoplasm with the highest incidence [2]; specifically, the province of A Coruna (Galicia, NW of Spain) is an area where the incidence and mortality of CRC are increasing every year [2], The usual precursors of CRC are colorectal polyps and the progression from benign polyps to carcinomas, which occurs very slowly and sporadically, usually taking approximately 10 to 15 years if there is no specific genetic syndrome (which occurs only in ~5% of cases) [3], Most common symptoms of CRC are unspecific (e.g., presence of mucus or blood in stool, abdominal or pelvic pain), and most of the time patients do not show any signs until illness worsens, so colorectal tumors are frequently diagnosed at advanced stages [4-6], Later CRC detection decreases the survival rate and increases the morbidity of patients, with a 5-year survival rate of 75-90% if cancer is detected in early stages (e.g., stages I-II) and only -15% when CRC is diagnosed in more serious and advanced stages (e.g., stage IV) [6], Stool tests are easy and inexpensive methods for selecting CRC high-risk individuals. In Spain, with the aim of increasing the detection of cases, CRC follow-up programs have been implemented for individuals aged 50 years and above. These programs utilize fecal occult blood tests (FOBT), which promote a downward trend in CRC mortality [7], Nevertheless, worldwide epidemiological studies have reported that the number of individuals who debuted before 50 years old has increased [8-10] and moreover, the specificity of FOBT is quite low, being significantly higher only for advanced CRC phases (stages III-IV)

[0011] , In addition, if FOBT results are positive, colonoscopy is recommended, which is unnecessary in most cases where other disorders may cause rectal bleeding, involving a significant workload in hospitals.

[0006] Therefore, there is an urgent need for new, more specific, and non-invasive biomarkers to enhance early detection of colorectal neoplasia.

[0007] On the other hand, prior to the advent of the Next Generation Sequencing (NGS) and Third Generation Sequencing (TGS) platforms, the analysis of microbial compositions relied on traditional bacteriological procedures. However, the high proportion of viable but non-culturable or difficult-to-cultivate microorganisms in the biosphere complicates the analysis of microbiome compositions. As a matter of fact, nowadays more than 99% of the potentially 1011-1012microbial species on Earth remain undiscovered or uncultured. Nevertheless, high-throughput sequencing methods enabled, for the first time in history, the possibility to identify, without culturing, practically ~100% of the genetic material of all microorganisms in a given sample. For this reason, nowadays 16S rRNA gene sequencing represents the gold-standard method for microbiome studies, standing as the fundamental driver behind the increasing volume of studies shaping environmental and animal microbial communities.

[0008] Oxford Nanopore Technologies (ONT) and its innovative long-read sequencing method, which can generate >10 Kb sequences directly from native DNA, has positioned itself as a compelling alternative to short-read technologies such as Illumina or other long-read expensive technologies like PacBio for 16S rRNA analysis. In comparison to Illumina, which is restricted to sequence small areas of the 16S rRNA such as V3-V4 (~400 nt) and identification mostly at the genus level, ONT can obtain the full region (~1500 nt, V1-V9), which allows identifying reads at the species level more consistently. The main limitation is the relatively higher error ONT reads achieve in comparison to other technologies, although its chemistry and basecallers, such as Dorado, are constantly being improved and will presumably get to a similar quality eventually, having recently achieved above Q20 in a small proportion of reads (in contrast to Illumina and PacBio’ s consistent +Q30). Putting things into perspective, Q20 indicates an error rate of 1% (Q15 = 5% error rate), which is the threshold needed to confidently assign an OTU (Operational Taxonomic Unit) to a specific species in full length 16S rRNA. This limitation has led to different bioinformatic approximations used for each technology. While Illumina and PacBio reads’ quality allows for the creation of precise ASVs (Amplicon Sequence Variant), typically through DADA2, this method is not prepared for the current quality profile of ONT reads, which has caused the elaboration of various tools such as Emu or NanoClust. Nonetheless, future advancements in ONT’s technology might allow for the use of procedures typically reserved for Illumina or PacBio.

[0009] To assess the capabilities and applications of ONT’s long reads to generate disease- related bacterial biomarkers in comparison to Illumina, colorectal cancer (CRC) was used as an example in this study, which stands out as one of the most diagnosed tumors all over the world.

[0010] Particularly in Spain (Europe), CRC has the highest incidence among tumors, with over 40,000 new cases throughout 2023. Moreover, CRC is the second leading cause of cancer mortality, with more than 15,000 deaths reported, according to the Spanish Association Against Cancer latest annual report. The CRC assessment programs implemented worldwide predominantly utilize a guaiac-based fecal occult blood test (gFOBT) or a fecal immunochemical test (FIT), which are conducted biennially, given that colon cancer development is a slow and gradual process. As a result, the implementation of these screenings allows to: (a) increase the number of annual diagnostics, (b) improve the CRC survival statistics and (c) achieve lower mortality rates. However, it is important to highlight that blood detection in fecal matter does not necessarily indicate the presence of a carcinoma, and that a colonoscopy (an invasive and expensive medical technique) is required to confirm or deny the FOBT or FIT result. Therefore, there is an urgent need to develop novel screening methods, more specific, non-invasive and cheaper, to enhance the detection of intestinal lesions even at precancerous stages.

[0011] Interestingly, as already indicated previously, several gut microbes could develop harmful effects on colonocytes’ integrity and homeostasis. Multiple in vitro and in vivo experiments demonstrated over the last years that specific bacteria, such as Escherichia coli pks+, enterotoxigenic strains of Bacteroides fragilis (ETBF), Parvimonas micra or Fusobacterium nucleatum are extremely related with the colorectal tumorigenesis process. More precisely, microorganisms have the potential to directly and indirectly: (a) affect the epithelial permeability (e.g., by modulating the expression of tight junction proteins), (b) promote chronic tissue inflammation (e.g., through the secretion of toxins, enhancing bacterial adherence to epithelial cells), (c) deregulate host anti-tumoral immune activities (e.g., F. nucleatum through the production of certain adhesins such as Fap2, affecting the function of natural killer T cells), (d) trigger chromosomic instability and DNA mutagenesis (e.g., ETBF via the BFT toxin promotes the hyperproduction of reactive oxygen species or E. coli pks+, a pathogen which induces a characteristic mutational signature via the colibactin molecule), (e) alter eukaryotic DNA methylation patterns (e.g., P. micra promotes hypermethylation in genes related to cytoskeleton regulation and tumor suppression) and (d) modulate several cell-signaling pathways (e.g.; E-cadherin / p-catenin, TLR4 / MYD88 / NF-KB or SMO / RAS / p38 MAPK). In turn, these deleterious activities lead to hyperproliferation, senescence, tumor growth and invasiveness, additionally inducing the initial stages of the metastatic process. Consequently, tumor progression and effectiveness of anti-cancer therapies could be directly correlated with the gut bacteriome established in each patient. Accordingly, multiple studies have put forth specific microbiome-derived biomarkers for a non- invasive CRC diagnostic method during the preceding decade. These investigations have demonstrated the potential to differentiate individuals with malignant dysplasia from those without lesions through a simple analysis of their fecal material, utilizing 16S rRNA gene sequencing or quantitative PCR (qPCR) procedures.

[0012] Thus, the present invention further evaluates the potential of ONT to improve upon Illumina’s gold standard 16S rRNA analysis using a considerable cohort of subjects (n=123), examined previously with Illumina and composed of colorectal cancer patients and their healthy counterparts, in order to obtain more precise CRC biomarkers. Two different sequencing systems and approaches, Illumina-V3V4 and 0NT-V1V9 (RIO.4.1 chemistry), three Dorado basecalling models (fast, hac, sup; v4.1.0) and two databases (SILVA and Emu’s Default database) were compared.

[0013] Brief Description of the Drawings

[0014] Figure 1. Schematic representation of the workflow during the present study. All colorectal cancer (CRC) patients involved in this study were diagnosed and treated at the University Hospital of A Coruna (CHUAC, Spain). CRC diagnosis was based on: (1) positive colonoscopy for colorectal neoplasia confirmed by histopathological analysis of biopsied tissues or (2) through CRC screening programs consisting of a positive fecal occult blood test (FOBT) followed by a positive colonoscopy validated by histopathological analysis. The cohort for the study consisted of 93 patients diagnosed with CRC and 30 healthy individuals without any digestive disorders (non-CRC). All the participants completed a questionnaire and collected saliva (S) and fecal (F) samples at home. Tissue samples such as normal colorectal mucosa tissue (NM) and adenocarcinoma tissue (Ac) of CRC patients were collected by the surgery team after colon laparoscopic resection at CHUAC. Besides, gingival crevicular fluid samples (GCF) were collected at Pardinas Medical Dental Clinic (A Coruna, Spain) during an oral examination. Finally, all different-nature samples were sent to the microbiology laboratory at CHUAC where they were processed, sequenced and analyzed using different bioinformatic tools.

[0015] Figures 2. Bacteriome landscape of the different-nature samples obtained from colorectal cancer patients (CRC) and individuals without any digestive disorders (non-CRC) analyzed by 16S rRNA Illumina sequencing. Barplots show the relative abundance (RA) at the family level (A), genera level (B) and species level (C), with a prevalence filter of 30%. The panel D shows a Venn diagram indicating the number of bacterial genera detected in each of the samples, as well as the number of genera common to all the samples with a list of oral bacteria. Samples analyzed: saliva (S) and feces (F) from both non-CRC and CRC individuals; gingival crevicular fluid (GCF), adenocarcinomas (Ac), and normal colorectal mucosa tissues (NM) from CRC patients.

[0016] Figure 3. Differential abundance analysis of the bacteriomes of fecal samples (F) from CRC and non-CRC individuals made at the genus (above), species (middle) and Amplicon Sequencing Variants (ASV, down) levels. Effect size (log fold change), standard error and adjusted p-values for each entry were obtained using the ANCOM-BC method, with a prevalence filter of 30%. Only effect sizes with adjusted p-values <0.05 are shown: (***: <0.001, **: <0.01, *: <0.05).

[0017] Figure 4. Landscape of oral related bacteria among different-nature samples of CRC) and non-CRC individuals. A) Alluvial barplot of oral related bacterial and their relative abundance (RA) among the different samples: saliva (S) from non-CRC and CRC individuals, gingival crevicular fluid (GCF) from CRC patients, adenocarcinomas (Ac) from CRC patients, normal colorectal mucosa tissues (NM) from CRC patients and feces (F) from both non-CRC and CRC participants. B) Differential abundance analysis of groups of ASVs from typical oral related genera obtained from feces (F) from CRC and non-CRC individuals. Effect size (log fold change), standard error and adjusted p-values for each entry were obtained using the ANCOM-BC method (***: <0.001, **: <0.01, >0.1). No prevalence cut was used for this analysis to show unsignificant entries belonging to oral related bacteria.

[0018] Figure 5. Heat map showing the associations found among bacteria in the oral cavity of CRC patients. A) Gingival crevicular fluid samples (GCF) of CRC patients. B) Saliva samples (S) of CRC patients. Species belonging to the Socransky complexes (red, orange, green, yellow, blue and purple) are marked with rectangles of the corresponding color.

[0019] Figure 6. Networks among oral bacteria present in colon tissue of CRC patients. A) Non- neoplastic colon tissue (NM). B) Colorectal adenocarcinoma tissue (Ac). Edge color corresponds to correlation strength, shown in the color key. Color of nodes corresponds to Socransky complexes color and, additionally, light blue was used for other oral related species and gray for gut associated species. Size of nodes was related to mean abundance in tissue.

[0020] Figure 7. Heat map of associations between oral-related bacteria and all the bacteria present in intestinal samples of CRC patients. A) Non-neoplastic colon tissue (NM). B) Colorectal adenocarcinoma tissue (Ac). Species belonging to the Socransky complexes were marked with rectangles of the corresponding color.

[0021] Figure 8. Heat map of associations between bacteria in the oral cavity and the tumor of CRC patients. A) Gingival plaque (GCF) and adenocarcinoma (Ac) paired samples. B) Saliva (S) and Ac paired samples.

[0022] Figure 9. Tissue enterotypes analysis in colorectal cancer (CRC) patients. A) Dirichlet Multinomial Mixtures (DMM) was used to infer the optimal number of community types in NM samples. Model fit was measured by Akaike's Information Criterion (aic) (dotted line), Bayesian Information Criterion (bic) (dashed line) and Laplace approximation (Iplc) (solid line). B) NM samples from both enterotypes (enterotype 1 : blue, enterotype 2: yellow) and their corresponding AC sample (enterotype 1 : red, enterotype 2: green) were represented with a CCA plot.

[0023] Figure 10. Prediction performance of bacterial biomarkers detected in fecal samples. Prediction performance is indicated as AUC values obtained from ROC curves of a leave- one-out cross validation method based on models with 1, 2, 3 or 4 bacterial genera.

[0024] Figure 11. Comparison of ONT -VI V9 basecalling models. A) Distribution of the average quality of reads per base-calling model. B) P-diversity analysis using SILVA database and MDS+JSD. The same samples with different models are connected by lines. C) a- diversity analysis using SILVA and Emu’s Default database. Significance levels are given through Wilcoxon rank sum tests and comparisons between groups are not shown.

[0025] Fig. 12. Comparison of Illumina-V3V4 and ONT-V1V9 approaches using SILVA. A) a- diversity indexes at the genus level, comparing the two volunteer groups (cancer vs. control). B) P-diversity analysis at the genus level. C) P-diversity analysis at the species level. D) Mean Centered Log Ratio (CLR) abundance correlation between approaches and for each group (cancer vs. control). Each point represents a different genus, and color indicates if that taxa appear, on average, on both, none or only one of the approaches. Three relevant genera, which contain multiple species related to colorectal cancer are highlighted (Fusobacterium, Parvimonas and Peptostreptococcus) .

[0026] Fig. 13 Comparison of Illumina-V3V4 and ONT-V1V9 approaches with both databases for specific genera in each subject group. Three important genera, which contain multiple species associated with colorectal cancer are highlighted (Fusobacterium, Parvimonas and P e piastre ptococcus). A) Centered Log Ratio (CLR) abundance of each genus per group using Wilcoxon rank sum tests to assess significance. B) Percentage of samples where each genus is present per group.

[0027] Fig. 14 Colorectal cancer biomarkers obtained with ONT-V1V9 and both databases. A) Differential abundance analysis (DAA) through ANCOM-BC using ONT-V1V9 and SILVA. Control subjects are the reference group, meaning a higher Log Fold Change (LFC) indicates higher abundance of a taxon in the cancer group. B) DAA through ANCOM-BC using ONT-V1V9 and Emu’s Default database. Control subjects are the reference group, meaning a higher Log Fold Change (LFC) indicates higher abundance of a taxon in the cancer group. C) Centered Log Ratio (CLR) abundance of species with significant differences between cancer and control groups, indicated through Wilcoxon rank sum tests. D) AUC of two models (Illumina- V3V4 or ONT-V1V9) using manually selected features based on read identity to reference, ANCOM-BC and CLR abundance. Sample size (n) refers to the number of features included in each model.

[0028] Detailed Description of the Invention

[0029] In the present invention, we performed an in-depth analysis of the oral and intestinal microbiome of samples obtained from a cohort of CRC patients (gingival crevicular fluid, saliva, non-neoplastic tissues, adenocarcinoma tissues, and feces), to establish correlations between bacteria within and between niches and to provide new insights about the oral-gut connection. Correlation analyses were performed to identify the bacterial consortia present in the oral and gut environments, including the tumor tissues. The comparison between microbiomes from CRC patients and those of healthy people allowed us to describe a specific combination of bacteria that could serve as a powerful noninvasive biomarker for CRC diagnosis.

[0030] Current CRC follow-up screenings programs in Spain are based on FOBT tests, and if the result is positive, a colonoscopy is recommended [7], Despite this, in many cases, the CRC symptoms are non-specific, and the disease is only apparent when there is rectal bleeding or acute abdominal pain, which often corresponds to an advanced tumor stage, with a higher likelihood of distant metastasis [8, 69, 5], At this critical phase, an individual’s response to chemotherapy agents and their survival could also be compromised. Therefore, there is an urgent need for new, specific, and minimally invasive biomarkers to enhance the early detection of colorectal cancer.

[0031] Over the last few years, multiple studies have demonstrated that the gut microbiome of patients with CRC is significantly unbalanced. Colorectal dysbiosis is characterized by the loss of mutualistic species, significant reduction in microbial biodiversity, and strong decline in healthy microbiota functions [30, 70, 36, 31, 25, 50], In example 1 of the current study, we identified a clear over-representation of common well-known periodontal pathogens in all intestinal samples obtained from CRC diagnosed individuals. Interestingly, this finding suggests that some of these pathobionts could potentially serve as non-invasive fecal biomarkers for CRC diagnosis (specifically, Fusobacterium and Parvimonas). Most of the oral microorganisms detected in the samples were strict anaerobes (e.g., Fusobacterium, Parvimonas, Prevotella, Peptostreptococcus and Porphyromonas), which are probably translocated from the subgingival cavity to the gut. These oral bacteria, especially Fusobacterium and Parvimonas, were also over- represented in samples from other CRC studied cohorts. Therefore, we suggest that the co-occurrence of these oral microorganisms in the gut, which are typical drivers of chronic oral inflammatory diseases, may play an important role in the development of colorectal tumors.

[0032] In example 1 of the current study, we found that the abundance of Fusobacterium increased from the normal mucosa to tumors. In addition, fecal samples from CRC patients showed to be significantly more enriched in Fusobacterium when compared to feces of non-CRC individuals. Furthermore, according to our results, another oral pathogen that overgrows in the colon of patients with CRC, is Parvimonas. Furthermore, Peptostreptococcus was another enriched genus in our CRC patient cohort. Moreover, this invention further shows that a consistent depletion of Faecalibacterium and Blautia genera was observed in CRC individuals when compared to the non-CRC group.

[0033] The present invention supports that there is no clear overlap in microbiome composition between oral and fecal samples obtained from non-CRC individuals. In Spain, a significant proportion of the adult population between the ages of 50 and 80 years, estimated at approximately 35%, is affected by PD

[0104] , In addition, recent studies revealed that patients diagnosed with PD have an increased risk of CRC development by -44% when compared to individuals with good oral health, suggesting a positive correlation between oral disorders and CRC [37, 53], Micro-communities of various oral pathobionts, naturally found in saliva or gingival fluids

[0105] , have the potential to migrate from the oral cavity to the gut [106, 107], This migration can contribute to the colonization of new colorectal microenvironments, including abnormal gut structures, especially during the progression of the PD, when these pathobionts experience overgrowth [37, 53], The similarities between the colorectal epithelium and the subgingival cavity, such as similar pH and low percentage of oxygen levels, create favorable conditions for the adaptation and overgrowth of these harmful periodontal organisms [108, 109], In particular, the increased mucosal tissue mass, as occurs in dysplasia, provides an advantage for the growth of oral microbes due to the increased supply of nutrients. Usually, the formation of multi-species consortia enhances the viability of individual bacteria, facilitates their adhesion and invasion of host cells, disrupts cell adhesive contacts, and further increases collective virulence as well as host vascular permeability [110, 111], This could explain why inflammatory gum infections were triggered mostly by various microbes, suggesting that the microbial cluster detected in our invention, could promote a strong pro-inflammatory effect in the gut, in a similar way that occurs in the oral cavity, altering eukaryotic cell signaling pathways [76, 77, 90, 91, 82], However, it remains unknown how CRC-associated bacteria interact in this dysbiotic environment. In this invention we observed that oral associated bacteria in tumors correlated with each other in a similar way as they did in the oral niche. Different bacterial complexes were established in the subgingival cavity by Socransky in the year 1998

[0112] , Bacteria from orange and red complexes (proteolytic and strict anaerobes) were associated with PD whereas yellow and purple ones (facultative anaerobes and saccharolytic) were established as earlier colonizers and health associated. Nevertheless, it was shown that in supragingival liquid, members of both types of bacteria coexist in the biofilms

[0113] and bacteria from the “early colonizers” group (e.g., Streptococcus) protect strict anaerobes (e.g., Fusobacterium) from oxidative stress

[0114] , This synergism could explain the positive correlation among “early colonizers” and aerobes with strict anaerobes in the saliva fluid and normal colorectal tissues, observed in the current invention. It is worth mentioning that although Fusobacterium and Parvimonas correlated positively in gingival and tissue samples, they did not correlate in saliva, suggesting that the synergy between these two pathogens is not as favorable or required in this oral environment. The lower correlation between strict and facultative anaerobic bacteria in carcinomas could indicate reduced oxygen irrigation in this microenvironment. In relation to this, Galeano et al.

[0109] conducted a study on the spatial distribution of intratumoral bacteria in CRC tumors. Their findings demonstrated that bacterial communities tend to populate microniches that are less vascularized and, therefore, have lower oxygen pressure. Although we were able to observe that the co-occurrence patterns of oral bacteria in carcinomas were like those in the oral cavity, it is interesting to highlight that not all PD related pathogens can reach the gut environment. For example, although Fusobacterium, Parvimonas and Peptostreptoccocus are present in adenocarcinomas, Tannerella and Treponema are absent in most patients. Furthermore, when we correlated the presence of bacteria in the saliva or subgingival fluid with their presence in colorectal tumors, no clear correlations were found. Therefore, our data indicate that the cluster of oral pathobionts detected in carcinomas represents a subset of the more complex ones present in the oral cavity.

[0034] Based on our bacteriome co-occurrence data, we consider that a preferred approach for diagnosing CRC in fecal samples should not focus solely on the detection of one bacteria, but on a consortium of oral and intestinal pathogens. Therefore, we investigated the use of oral over-abundant genera in the gut of patients with CRC as potential biomarkers in fecal samples. Initially, we tested Fusobacterium as a single diagnostic biomarker, obtaining a modest prediction performance (AUC 0.65). Specifically, in our invention, we expanded this bacteriome biomarker model by integrating other over-abundant bacteria in fecal samples from CRC patients (Parvimonas and Bacteroides) and another enriched genus in non-CRC fecal samples (Faecalibacterium), resulting in an improved performance, up to an AUC of 0.8. The genera included in our model were: Fusobacterium, Parvimonas, Bacteroides and Faecalibacterium. Our analysis revealed that the combination of bacteria, including intestinal related genera (Bacteroides and Faecalibacterium), improves the discriminatory power between non-CRC and CRC individuals. Moreover, it is important to remark that, most of these bacterial genera (Fusobacterium, Parvimonas and Bacteroides) were also found to be over-represented in adenomatous polyps, which are recognized as typical CRC precursor lesions [39, 42, 32], The data obtained in example 1 of the current study thus revealed that a specific consortium of oral and gut bacteria (Fusobacterium, Parvimonas, Bacteroides and Faecalibacterium) could be used as a predictive model for the detection of malignant neoplasms in the colon, even at early stages that are not detected by colonoscopy.

[0035] It must be noted that in the context example 1 of the current study, the genus Fusobacterium is found at higher levels in biological intestinal or stool samples in patients with CRC compared to the general population. Furthermore, it must be noted that the following sequences NR_028933.1, NR_027588.1, NR_117734.1, NR_042365.1, NR_113384.1, NR_074412.1, NR_026084.1, NR_026085.1, NR_044689.2 and NR_044820.1 (NCBI reference sequence ID) are sequences representative of this genus and are therefore useful for the identification of bacteria belonging to this genus in said samples through a PCR method or by means of any sequencing technique. One skilled in the art will have knowledge of other sequences representative of this genus which will be used for identifying those bacteria belonging to this genus in a stool or intestinal sample.

[0036] Moreover, it must be noted that in the context of example 1 of the current study, the genus Parvimonas is found at higher levels in biological intestinal or stool samples in patients with CRC compared to the general population. Furthermore, it must be noted that the following sequences NR_036934.1, NR_114675.1 and NR_114338.1 (NCBI reference sequence ID) are sequences representative of this genus and are therefore useful for the identification of bacteria belonging to this genus in said samples through a PCR method or by means of any sequencing technique.

[0037] Furthermore, it must be noted that in the context of example 1 of the current study, the genus Bacteroides is found at higher levels in biological intestinal or stool samples in patients with CRC with respect to the general population. Furthermore, it must be noted that the following sequences NR_074784.2, NR_119164.1, NR_112936.1 and NR_112141.1 (NCBI reference sequence ID) are sequences representative of this genus and are therefore useful for the identification of bacteria belonging to this genus in said samples through a PCR method or by means of any sequencing technique.

[0038] One skilled in the art will have knowledge of other sequences representative of this genus which will be used for identifying those bacteria belonging to this genus in a stool or intestinal sample.

[0039] Finally, it must be noted that in the context of at least example 1 of the current study, the genus Faecalibacterium is found at lower levels in biological intestinal or stool samples in patients with CRC with respect to the general population. Furthermore, it must be noted that the following sequences NR_028961.1, NR_178280.1, NR_178311.1, NR_178971.1, NR_178972.1 and NR_179388.1 (NCBI reference sequence ID) are sequences representative of this genus and are therefore useful for the identification of bacteria belonging to this genus in said samples through a PCR method or by means of any sequencing technique.

[0040] One skilled in the art will have knowledge of other sequences representative of this genus which will be used for identifying those bacteria belonging to this genus in a stool or intestinal sample. Therefore, a first aspect of the present invention relates to the in vitro use of the level or the concentration in an intestinal or stool sample of bacteria belonging to the genus Fusobacterium, Parvimonas, Bacteroides and / or Faecalibacterium, or any combination thereof, for the diagnosis / of CRC or colorectal polyps (as an early sign or marker of CRC) in a patient, or for obtaining useful data which allows said diagnosis.

[0041] An alternative embodiment of the first aspect of the invention relates to a method for the in vitro diagnosis or for obtaining useful data which helps in said diagnosis of a subject suspected of possibly suffering from CRC or colorectal polyps (as an early sign or marker of CRC), which method comprises using, as an indicator in an intestinal or stool sample obtained from said subject, the level or the concentration in said sample of bacteria belonging to the genus Fusobacterium, Parvimonas, Bacteroides and / or Faecalibacterium, or any combination thereof, wherein if the level or the concentration in said intestinal or stool sample of bacteria belonging to Fusobacterium, Parvimonas, Bacteroides and / or Faecalibacterium, or any combination thereof differs or varies from that present in an intestinal or stool sample obtained from a healthy subject or with respect to a reference value, it is indicative of said subject having CRC or colorectal polyps.

[0042] In a preferred embodiment of the first aspect of the invention, the method comprises using, as an indicator in an intestinal or stool sample obtained from said subject, the level or the concentration in said sample of bacteria belonging to the genus Fusobacterium, wherein if the level or the concentration in said intestinal or stool sample of bacteria belonging to said genus is increased with respect to that present in an intestinal or stool sample obtained from a healthy subject or with respect to a reference value, it is indicative of said subject having CRC or colorectal polyps.

[0043] In another preferred embodiment of the first aspect of the invention, the method comprises using, as an indicator in an intestinal or stool sample obtained from said subject, of the level or the concentration in said sample of bacteria belonging to the genus Parvimonas, wherein if the level or the concentration in said intestinal or stool sample of bacteria belonging to said genus is increased with respect to that present in an intestinal or stool sample obtained from a healthy subject or with respect to a reference value, it is indicative of said subject having CRC or colorectal polyps. In another preferred embodiment of the first aspect of the invention, the method comprises using, as an indicator in an intestinal or stool sample obtained from said subject, the level or the concentration in said sample of bacteria belonging to the genus Bacteroides, wherein if the level or the concentration in said intestinal or stool sample of bacteria belonging to said genus is increased with respect to that present in an intestinal or stool sample obtained from a healthy subject or with respect to a reference value, it is indicative of said subject having CRC or colorectal polyps.

[0044] In another preferred embodiment of the first aspect of the invention, the method comprises using, as an indicator in an intestinal or stool sample obtained from said subject, the level or the concentration in said sample of bacteria belonging to the genus Faecalibacterium, wherein if the level or the concentration in said intestinal or stool sample of bacteria belonging to said genus is decreased with respect to that present in an intestinal or stool sample obtained from a healthy subject or with respect to a reference value, it is indicative of said subject having CRC or colorectal polyps.

[0045] In another preferred embodiment of the first aspect of the invention, the method comprises using, as an indicator in an intestinal or stool sample obtained from said subject, the level or the concentration in said sample of bacteria belonging to the genus Fusobacterium and Parvimonas, wherein if the level or the concentration in said intestinal or stool sample of bacteria belonging to said genus is increased with respect to that present in an intestinal or stool sample obtained from a healthy subject or with respect to one or more reference values, it is indicative of said subject having CRC or colorectal polyps.

[0046] In yet another preferred embodiment of the first aspect of the invention, the method comprises using, as an indicator in an intestinal or stool sample obtained from said subject, the level or the concentration in said sample of bacteria belonging to the genera Fusobacterium, Parvimonas, Bacteroides and / or Faecalibacterium, wherein if the level or the concentration in said intestinal or stool sample of bacteria belonging to the genera Faecalibacterium is decreased with respect to that present in an intestinal or stool sample obtained from a healthy subject or with respect to one or more reference values, and wherein if the level or the concentration in said intestinal or stool sample of bacteria belonging to the genera Fusobacterium, Parvimonas, and / or Bacteroides is increased with respect to that present in an intestinal or stool sample obtained from a healthy subject or with respect to a reference value, it is indicative of said subject having CRC or colorectal polyps.

[0047] In the context of the present invention, CRC refers to the abnormal and uncontrolled cell growth of the cells that form the large intestine, leading to a alteration in its original structure. Cancer cells distinguish themselves from the adjacent normal cells by their morphology and the acquisition of new cellular characteristics. In the context of the present invention, colorectal polyps are abnormal, non-malignant formations that originate in the mucosal layer of the large intestine. It is estimated that 95% of colorectal carcinomas develop from adenomatous polyps.

[0048] It must be noted that the present invention sufficiently describes the following sequences (NCBI reference sequence ID): NR_028933.1, NR_027588.1, NR_117734.1, NR_042365.1, NRJ 13384.1, NR_074412.1, NR_026084.1, NR_026085.1,

[0049] NR_044689.2, NR_044820.1, NR_036934.1, NRJ 14675.1, NRJ 14338.1,

[0050] NR_074784.2, NRJ 19164.1, NRJ 12936.1, NRJ 12141.1, NR_028961.1,

[0051] NR_178280.1, NR_178311.1, NRJ78971.1, NRJ78972.1 and NRJ79388.1, which allow identifying the presence as well as the concentration of each of the genera herein mentioned.

[0052] A second aspect of the present invention relates to in vitro method for determining the risk that a subject has colorectal cancer or colorectal polyps, comprising the following steps: a. determining the levels or the concentration of bacteria belonging to the genus Fusobacterium, Parvimonas, Bacteroides and / or Faecalibacterium, or any combination thereof, in an intestinal or stool sample isolated from said subject; and b. identifying the subject as a subject at risk of developing colorectal cancer by a predictive model which correlates at least one or more of the levels identified in step (a) with representative levels of the same from samples obtained or isolated from subjects previously identified as suffering from colorectal cancer or colorectal polyps, said predictive model having been generated by training a computer with a plurality of the concentration levels of the bacteria belonging to the genus identified in step (a) from previously identified subjects having colorectal cancer or colorectal polyps, by machine learning on said plurality of levels so as to obtain representative score profiles associated with colorectal cancer.

[0053] A preferred embodiment of the present invention refers to the in vitro method of the second aspect of the invention, wherein step a) comprises the determination of the levels or the concentration of the bacteria belonging to the following genus: Fusobacterium, and Parvimonas, in an intestinal or stool sample isolated from said subject.

[0054] Another preferred embodiment of the present invention refers to the in vitro method of the second aspect of the invention, wherein step a) comprises the determination of the levels or the concentration of the bacteria belonging to the following genus: Fusobacterium, Parvimonas, and Bacteroides, in an intestinal or stool sample isolated from said subject.

[0055] Another preferred embodiment of the present invention refers to the in vitro method of the second aspect of the invention, wherein step a) comprises the determination of the levels or the concentration of all of the bacteria belonging to the following genus: Fusobacterium, Parvimonas, Bacteroides and Faecalibacterium, in an intestinal or stool sample isolated from said subject.

[0056] A third aspect of the invention refers to an in vitro method for diagnosing colorectal cancer in a subject in need thereof, comprising the following steps: a. determining the levels or the concentration of bacteria belonging to the genus Fusobacterium, Parvimonas, Bacteroides and / or Faecalibacterium, or any combination thereof, in an intestinal or stool sample isolated from said subject; and b . identifying the subj ect as a subj ect having colorectal cancer by a predictive model which correlates at least one or more of the levels identified in step (a) with representative levels of the same from samples obtained or isolated from subjects previously identified as suffering from colorectal cancer, said predictive model having been generated by training a computer with a plurality of the concentration levels of the bacteria belonging to the genus identified in step (a) from previously identified subjects having colorectal cancer, by machine learning on said plurality of levels so as to obtain representative score profiles associated with colorectal cancer.

[0057] A preferred embodiment of the present invention refers to the in vitro method of the third aspect of the invention, wherein step a) comprises the determination of the levels or the concentration of the bacteria belonging to the following genus: Fusobacterium, and Parvimonas, in an intestinal or stool sample isolated from said subject.

[0058] Another preferred embodiment of the present invention refers to the in vitro method of the third aspect of the invention, wherein step a) comprises the determination of the levels or the concentration of the bacteria belonging to the following genus: Fusobacterium, Parvimonas, and Bacteroides, in an intestinal or stool sample isolated from said subject. Another preferred embodiment of the present invention refers to the in vitro method of the third aspect of the invention, wherein step a) comprises the determination of the levels or the concentration of all of the bacteria belonging to the following genus: Fusobacterium, Parvimonas, Bacteroides and Faecalibacterium, in an intestinal or stool sample isolated from said subject.

[0059] In a preferred embodiment of any of the previous aspects of the invention, the amplification reaction is carried out by means of a real-time polymerase chain reaction.

[0060] In the context of the present invention, reference value is preferably understood to be the result of the data of a mathematical algorithm which uses the concentration and the total amount of each bacterium belonging to one or more of the genera proposed in the present invention in the general population or in a healthy subject. The best sensitivity and specificity value will automatically be proposed using the algorithm. This algorithm will provide values along with the proposed value, with the changing sensitivity and specificity of the values in order to provide researchers with more information and to allow them to decide on the best test and cut-off value for each patient or specific situation.

[0061] In the context of the previous embodiments of the invention, the bacteria belonging to any of the genera proposed in an intestinal or stool sample isolated from said subject can be determined, in a non-limiting manner, by means of massive sequencing of the genome of the stool such that the total number of sequences of that bacterium existing in the stool is obtained together with the total number of other bacteria present in that stool sample. Preferably, said determination is performed by means of Next Generation Sequencing methods.

[0062] It will be appreciated that the term "machine learning" generally refers to algorithms that give a computer the ability to learn without being explicitly programmed, including algorithms that learn from and make predictions about data. Machine learning algorithms employed by the embodiments disclosed herein may include, but are not limited to, random forest ("RF"), least absolute shrinkage and selection operator ("LASSO") logistic regression, regularized logistic regression, XGBoost, decision tree learning, artificial neural networks ("ANN"), deep neural networks ("DNN"), support vector machines, rulebased machine learning, and / or others.

[0063] For clarity, algorithms such as linear regression or logistic regression can be used as part of a machine learning process. However, it will be understood that using linear regression or another algorithm as part of a machine learning process is distinct from performing a statistical analysis such as regression with a spreadsheet program. Whereas statistical modeling relies on finding relationships between variables (e.g., mathematical equations) to predict an outcome, a machine learning process may continually update model parameters and adjust a classifier as new data becomes available, without relying on explicit or rules-based programming.

[0064] In a particular embodiment, the second step of the third aspect of the invention is performed by a classification method, which results in identifying the subject as a subject at risk of suffering from colorectal polyps or adenomas and / or colorectal cancer.

[0065] In one embodiment, step (b) is carried out by a classification method; preferably selected from logistic regression, random forest, gradient boosting (GB), adaptive boosting (AB), extreme Gradient Boosting (XGB) k-nearest neighbors (kNN), artificial neural network (ANN), support vector machine (SVM), and combinations thereof.

[0066] In an embodiment, the predictive model is generated by training the computer with a plurality of levels or concentrations, as indicated in the second or third aspect of the invention from previously identified samples from subjects having colorectal adenomas and / or colorectal cancer by machine learning on said plurality of levels so as to obtain representative multivariable data sets associated with this group of patients; wherein the training comprises the following steps:

[0067] (i) training data, from a plurality of score profiles, is randomly stratified into:

[0068] - a calibration dataset (for example in a percentage of about 75%), and

[0069] - a validation dataset (for example in a percentage of about 25%);

[0070] (ii) the predictive model is seeded on the calibration dataset (particularly is developed by applying a machine learning method selected from a regression method, a classification method or a combination thereof on the calibration dataset);

[0071] (iii) the predictive model is optimized by an internal cross validation; preferably by a k- fold cross validation, wherein each of the k cases of the k-fold cross validation is used for testing only once and one at a time; and

[0072] (iv) the predictive model is further validated by predicting new samples using the validation dataset.

[0073] In some embodiments, the second step is performed by a classification method wherein the patients are assigned a probability of belonging to given category such as patients having colorectal adenomas and / or colorectal cancer. In some embodiments, the classification method is carried out by a method selected from gradient boosting, support vector machine (SVM), decision trees, K nearest neighbors, Naive Bayes or neural networks. In a preferred embodiment, the classification method is carried out by a Gradient Boosting. As used herein, Gradient Boosting is a machine learning algorithm that uses a gradient boosting framework. Gradient Boosting trees, a decision-tree-based ensemble model, differ fundamentally from conventional statistical techniques that aim to fit a single model using the entire dataset. Such ensemble approach improves performance by combining strengths of models that learn the data by recursive binary splits, such as trees, and of "boosting", an adaptive method for combining several simple (base) models. At each iteration of the gradient boosting algorithm, a subsample of the training data is selected at random (without replacement) from the entire training data set, and then a simple base learner is fitted on each subsample. The final boosted trees model is an additive tree model, constructed by sequentially fitting such base learners on different subsamples. This procedure incorporates randomization, which is known to substantially improve the predictor accuracy and also increase robustness. Additionally, boosted trees can fit complex nonlinear relationships, and automatically handle interaction effects between predictors as addition to other advantages of tree-based methods, such as handling features of different types and accommodating missing data. Hence, in many cases their predictive performance is superior to most traditional modelling methods. In a particular embodiment, the second step is performed by a regression method; preferably selected from multiple linear regression (MLR), principal component regression (PCR), partial least squares regression (PLSR), artificial neural network (ANN), support vector machine (SVM), random forest (RF), lassor regression, ridge regression and combinations thereof.

[0074] In a particular embodiment, the second step is performed by a classification method, more in particular, by gradient boosting, which includes the value of one or more variables of the gene expression profile collected in step (i) and which contribute to the identification of the patients having colorectal adenomas and / or colorectal cancer.

[0075] In a particular embodiment, the second step is performed by a regression method which includes the value of one or more variables of the score profile collected in step (i) and which contribute to the identification of the patients having colorectal adenomas and / or colorectal cancer.

[0076] In another preferred embodiment of the first, second or third aspect of the invention or of any of the preferred embodiments of the invention, said levels or concentration of bacteria refer to the total amount of the bacteria belonging to the genus with respect to the total bacteria present in said sample.

[0077] In another preferred embodiment of the first, second or third aspect of the invention or of any of the preferred embodiments of the invention, the levels or the concentration of bacteria belonging to the genera Fusobacterium, Parvimonas, Bacteroides and / or Faecalibacterium, are determined by performing an amplification reaction from a nucleic acid preparation derived from said sample using a pair of primers capable of amplifying one or more representative regions of said genera.

[0078] In yet another preferred embodiment of the first, second or third aspect of the invention or of any of the preferred embodiments of the invention, the amplification reaction is carried out by means of a real-time polymerase chain reaction. Preferably, the detection of the amplification product is carried out by means of a fluorescent intercalating agent. Even more preferably, the detection of the amplification product(s) is carried out by means of a labeled probe, wherein the probe preferably comprises at its 5' end a reporter pigment and at its 3' end a “quencher” pigment or silencer or buffering agent.

[0079] In a fourth aspect of the invention, the method of the first, second or third aspect of the invention or of any of the embodiments thereof further comprises storing the results of the method in a data carrier, wherein said data carrier is preferably a computer readable medium.

[0080] In a fifth aspect of the invention, the method of the first, second or third aspect of the invention or of any of the embodiments thereof comprises at least the implementation of the comparative step and optionally the provision of a result as a consequence of said comparison using a computer program.

[0081] A sixth aspect of the invention relates to a kit comprising one or more pairs of primers capable of amplifying bacteria belonging to the genera Fusobacterium, Parvimonas, Bacteroides and / or Faecalibacterium. Preferably and as mentioned, a sixth aspect of the invention relates to a method for the detection of bacteria from a stool or intestinal sample, comprising the following steps: i) contacting the sample to be analyzed with a reaction mixture containing specific primers capable of amplifying bacteria belonging to the genera Fusobacterium, Parvimonas, Bacteroides and / or Faecalibacterium preferably for performing multiplex PCR, ii) performing amplification by means of polymerase chain reaction, iii) identifying the formation of the products of the preceding step, said formation being indicative of the levels or the concentration of bacteria belonging to one or more of the genera Fusobacterium, Parvimonas, Bacteroides and / or Faecalibacterium .

[0082] In relation to this sixth aspect of the invention, it preferably provides a method for simultaneously detecting Fusobacterium, Parvimonas, Bacteroides and / or Faecalibacterium . In a particular embodiment of this sixth aspect of the invention, DNA fragments included or comprised in sequences 1 to 12 are amplified.

[0083] In another embodiment of this sixth aspect of the invention, the amplification products which allow identifying the different bacterial species and groups are detected by means of using probes. In a more preferred embodiment, these probes have a length between 15 and 25 nucleotides. The primers can be designed by means of multiple alignment with programs such as CLUSTAL X, which allow identifying highly conserved regions that serve as a template.

[0084] Given the great abundance of PCR inhibitors, such as humic and fulvic acids, heavy metals, heparin, etc., which may give rise to false negatives, and despite the existence of methods which reduce the concentration of molecules of this type, it is advisable (cf. J. Hoorfar et al., “Making internal amplification control mandatory for diagnostic PCR” J. of Clinical Microbiology, Dec. 2003, pp. 5835) for PCR assays to contain an internal amplification control (IAC). This IAC is simply a DNA fragment which is amplified simultaneously with the target sample, such that its absence at the end of the assays is indicative of the presence of factors which have led to an unwanted PCR development.

[0085] On the other hand, as indicated in Example 2, full 16S rRNA VI V9 sequencing through Oxford Nanopore and its new R10.4.1 chemistry achieved accurate species-level bacterial identification, facilitating the discovery of more precise disease-related biomarkers and increasing the taxonomic fidelity of future microbiome analyses.

[0086] Therefore, a seventh aspect of the invention refers to an in vitro method for determining the risk that a subject has colorectal cancer, comprising the following steps: a. determining the levels or the concentration of bacteria belonging to any one of the genus and / or species belonging to the list consisting of: the genus Fusobacterium, preferably the species Fusobacterium nucleatum, the genus Parvimonas, preferably the species Parvimonas micra, the genus Bacteroides, preferably the species Bacteroides fragilis, the genus Peptostreptococcus, preferably the species Peptostreptococcus stomatis and / or Peptostreptococcus anaerobius, the genus Blautia, preferably the species Blautia luti, the genus Agathobaculum, preferably the species Agathobaculum butyriciproducens, the genus Gemella, preferably the species Gemella morbilloriim. the genus Clostridium, preferably the species Clostridium perfringens, the genus Sullerella, preferably the species Sutterella wadsworthensis, the genus Dialister, preferably the species Dialister pneumosinles, the genus Romboutsia, preferably the species Romboutsia ilealis, the genus Paraprevotella, preferably the species Paraprevotella clara, the genus Longibaculum, preferably the species Longibaculum sp. KGMB06250, the genus Raoultibacter, preferably the species Raoultibacter massiliensis, and / or the genus Faecalibacterium, preferably the species Faecalibacterium prausnilzii, or any combination thereof, in an intestinal or stool sample isolated from said subject; and b. comparing the levels or the concentration of said bacteria in said intestinal or stool sample with one or more reference values, wherein an increase in the number of sequences of any one of the following: the genus Fusobacterium, preferably the species Fusobacterium nucleatum, the genus Parvimonas, preferably the species Parvimonas micro, the genus Bacteroides, preferably the species Bacteroides fragilis, the genus Peptostreptococcus, preferably the species Peptostreptococcus stomatis and / or Peptostreptococcus anaerobius, the genus B lamia, preferably the species Blautia hili, the genus Agathobaculum, preferably the species Agathobaculum butyriciproducens, the genus Gemella, preferably the species Gemella morbillorum, the genus Clostridium, preferably the species Clostridium perfringens, the genus Sutterella, preferably the species Sutterella wadsworthensis, the genus Dialister, preferably the species Dialister pneumosintes, the genus Romboutsia, preferably the species Romboutsia ilealis, the genus Paraprevotella, preferably the species Paraprevotella clara, the genus Longibaculum, preferably the species Longibaculum sp. KGMB06250, the genus Raoultibacter, preferably the species Raoultibacter massiliensis,' and / or a decreased in the number of sequences of Faecalibacterium, preferably the species Faecalibacterium prausnitzii, in said sample with respect to said one or more reference values is indicative of an increase in the risk that the subject has colorectal cancer; wherein the said levels or the concentration of bacteria of step a) are determined by an Oxford Nanopore sequencing platform for accurate species-level bacterial identification, the method comprising:

[0087] • isolating and preparing nucleic acids from the intestinal or stool sample isolated from said subject;

[0088] • performing an Oxford Nanopore sequencing analysis on said nucleic acids, thereby generating sequencing data;

[0089] • identifying and quantifying bacteria at the genus and / or species level in the sample based on said sequencing data; and

[0090] • executing step b); wherein, preferably, the one or more reference values of step b) are understood to refer to the concentration and / or the total amount of each bacterium belonging to one or more of the genera proposed, and / or any scores obtained by combining such bacterial levels, in the general population or in a healthy subject.

[0091] As used herein, the term “Oxford Nanopore sequencing” (also referred to as “ONT sequencing”) refers to a method of nucleic acid sequencing in which individual polynucleotide strands pass through, or in proximity to, protein nanopores embedded in a membrane, and an electrical potential is applied across said membrane. The passage of the polynucleotide through or near the nanopore results in characteristic changes in ionic current, which are measured in real-time and correlated to specific nucleotide sequences. This platform enables long-read sequencing of nucleic acids and real-time data analysis without the need for fluorescent labels or reversible terminators. Furthermore, as used herein, “Oxford Nanopore sequencing carried out using primers targeting the V1-V9 regions of bacterial rRNA genes (0NT-V1V9)” refers to a specialized application of the aforesaid ONT sequencing methodology. In this application, the bacterial 16S rRNA gene is selectively amplified or enriched using primers that anneal to conserved sites flanking the VI through V9 hypervariable regions of said gene, thereby generating an amplicon comprising substantially the entire 16S rRNA gene. The resulting amplicons are then processed for ONT library preparation, optionally comprising end-repair, adapter ligation, or barcoding steps, and subsequently introduced into an Oxford Nanopore flow cell. Sequencing is achieved by translocating or threading the prepared 16S rRNA amplicons through the nanopores under an applied voltage; changes in ionic current corresponding to nucleotide composition are continuously measured and recorded. The output signal is computationally translated into a nucleotide sequence read spanning substantially the entire V1-V9 region of the bacterial 16S rRNA gene.

[0092] This ONT-V1V9 approach facilitates long-read, high-resolution profiling of bacterial communities, enabling the characterization, identification, or phylogenetic analysis of bacterial taxa in a sample by interrogating nearly the full length of the 16S rRNA locus within a single sequencing read.

[0093] In an embodiment of the seventh aspect of the invention, the Oxford Nanopore sequencing is carried out using primers targeting the V1-V9 regions of bacterial rRNA genes (ONT-V1V9), thereby facilitating the discovery of more precise disease-related biomarkers and increasing the taxonomic fidelity of the resulting microbiome analysis.

[0094] In another embodiment of the seventh aspect of the invention, preferably in combination with any previous or subsequent embodiment of the seventh aspect of the invention, the determined levels or concentration of bacteria at least belong to the genus and / or species selected from the genus Fusobacterium, preferably the species Fusobacterium nucleatum, the genus Parvimonas, preferably the species Parvimonas micra and / or the genus Faecalibacterium, preferably the species Faecalibacterium prausnitzii.

[0095] In another embodiment of the seventh aspect of the invention, preferably in combination with any previous or subsequent embodiment of the seventh aspect of the invention, the determined levels or concentration of bacteria further belong to the genus and / or species: genus Bacteroides, preferably the species Bacteroides fragilis.

[0096] In another embodiment of the seventh aspect of the invention, preferably in combination with any previous or subsequent embodiment of the seventh aspect of the invention, the determined levels or concentration of bacteria further belong to the genus and / or species: genus Agathobaculum, preferably the species Agathobaculum butyriciproducens. In another embodiment of the seventh aspect of the invention, preferably in combination with any previous or subsequent embodiment of the seventh aspect of the invention, the determined levels or concentration of bacteria further belong to the genus and / or species: genus Peptostreptococcus, preferably the species Peptostreptococcus stomatis and Peptostreptococcus anaerobius, genus Gemella, preferably the species Gemella morbillorum, genus Dialister, preferably the species Dialister pneumosintes, genus Sutterella, preferably the species Sutterella wadsworthensis, genus Clostridium, preferably the species Clostridium perfringens, and genus Romboutsia, preferably species Romboutsia ilealis.

[0097] In another embodiment of the seventh aspect of the invention, preferably in combination with any previous or subsequent embodiment of the seventh aspect of the invention, the determined levels or concentration of bacteria further belong to the genus and / or species: genus Paraprevotella, preferably the species Paraprevotella clara.

[0098] In another embodiment of the seventh aspect of the invention, preferably in combination with any previous or subsequent embodiment of the seventh aspect of the invention, the determined levels or concentration of bacteria further belong to the genus and / or species: genus Longibaculum, preferably the species Longibaculum sp. KGMB06250.

[0099] In another embodiment of the seventh aspect of the invention, preferably in combination with any previous or subsequent embodiment of the seventh aspect of the invention, the determined levels or concentration of bacteria further belong to the genus and / or species: genus Raoultibacter, preferably the species Raoultibacter massiliensis.

[0100] It is noted that figure 14 clearly indicates in figure 14D the specific combinations indicated above.

[0101] An eighth aspect of the invention refers to an in vitro method for diagnosing colorectal cancer in a subject in need thereof, comprising the following steps: a) determining the levels or the concentration, in an intestinal or stool sample isolated from the subject, of the bacteria according to step a) of the seventh aspect of the invention, or according to any preferred embodiment of the seventh aspect of the invention, wherein the said levels or the 1 concentration of bacteria of step a) are determined by an Oxford Nanopore sequencing platform in accordance with the seventh aspect of the invention, or according to any preferred embodiment of the seventh aspect of the invention; and b) comparing the levels or the concentration of said bacteria in said intestinal or stool sample with one or more reference values, wherein an increase or a decrease in said sample with respect to said one or more reference values, as defined in step b) of the seventh aspect of the invention, or in any preferred embodiment of the seventh aspect of the invention, is indicative that the subject has colorectal cancer; wherein, preferably, the one or more reference values of step b) are understood to refer to the concentration and / or the total amount of each bacterium belonging to one or more of the genera proposed, and / or any scores obtained by combining such bacterial levels, in the general population or in a healthy subject.

[0102] A ninth aspect of the invention refers to in vitro method for determining the risk that a subject has colorectal cancer, wherein the method comprises: a) determining the levels or the concentration, in an intestinal or stool sample isolated from the subject, of the bacteria according to step a) of the seventh aspect of the invention, or according to any preferred embodiment of the seventh aspect of the invention, wherein the said levels or the concentration of bacteria of step a) are determined by an Oxford Nanopore sequencing platform in accordance with the seventh aspect of the invention, or in accordance with any preferred embodiment of the seventh aspect of the invention; and b) identifying the subject as a subject at risk of developing colorectal cancer by a predictive model which correlates at least one or more of the levels identified in step (a) with representative levels of the same from samples obtained or isolated from subjects previously identified as suffering from colorectal cancer, said predictive model having been generated by training a computer with a plurality of the concentration levels of the bacteria belonging to the genus identified in step (a) from previously identified subjects having colorectal cancer, by machine learning on said plurality of levels so as to obtain representative score profiles associated with colorectal cancer.

[0103] In a tenth aspect of the invention, the method of the seventh to ninth aspects of the invention further comprises storing the results of the method in a data carrier, preferably wherein said data carrier is a computer readable medium.

[0104] An eleventh aspect of the invention, refers to a computer-implemented method for determining the risk that a subject has colorectal cancer or is at risk of having colorectal cancer, wherein said method comprises at least receiving the levels or concentrations of the bacteria (step a)) and carrying out the comparative step (step (b)) and optionally the provision of a result as a consequence of said comparison, as these steps are defined in the method according to the seventh to ninth aspects of the invention.

[0105] Throughout the description, the term “specific” means that the primers comprise a nucleotide sequence that is completely complementary to the genes or gene fragments used by the present invention.

[0106] A preferred embodiment of the sixth aspect of the invention relates to the use of the kit or of the methodology therein described for implementing the methodology according to any of the first, second, or third aspects of the present invention.

[0107] The following examples merely illustrate the present invention and must not be understood as limiting same.

[0108] Examples of the Invention

[0109] Example 1

[0110] MATERIALS AND METHODS

[0111] Patient’s recruitment

[0112] A total of 159 patients with CRC were recruited. To select CRC patients for this study, several exclusion criteria were established: (1) no antibiotic intake in less than one month, (2) no infectious disease in the last three months, (3) no chemotherapy and / or radiotherapy treatments prior to sample collection, (4) no genomic predisposition to develop CRC (other cases of CRC in first-degree relatives or Lynch syndrome among others), (5) no diagnosis of other gut disorders (such as inflammatory bowel disease), (6) no immunological diseases and (7) no transplants and / or any immunosuppressor treatment. Based on these criteria only 93 of the initial 159 patients were included. Likewise, CRC patients’ companions / couples were asked to participate in the study and only 34 agreed to participate. Individuals included in this non-CRC control cohort followed the same exclusion criteria as CRC patients but were not diagnosed with any type of cancer. The rigorous selection process performed in the study ensured that the control group consisted of individuals who were free from cancer, making it a suitable comparison cohort for studying the microbiome in the context of CRC. Only 30 of the 34 individuals met the requirements. Informed consent was obtained from all the participants before the sample collection phase.

[0113] Sample collection

[0114] A total of 93 samples of ca. 20 mL of feces (F) and 93 samples of 5 mL of unstimulated saliva (S) were collected at home by CRC patients included in the study, before starting the low-residue diet required for laparoscopy. Thirty healthy individuals were recruited from the same samples: F (n = 30) and S (n = 28). A previous interview was conducted with all participants to recover data related to age, sex, weight, height, and lifestyle habits, such as diet or sporting activity, oral diseases, antibiotic, intake, and previous surgeries. Furthermore, the F samples were kept in the presence of 10 mL of RNAlater reagent (Thermo Fisher Scientific, Waltham, Massachussets, USA). Gingival crevicular fluid (GCF) samples (n = 19) from CRC patients were collected by inserting sterile paper points (ISO 30, Henry Schein, USA) in the subgingival sulcus of different teeth for 30 s. Eight paper points per patient were obtained and placed in an Eppendorf tube containing 500 pL of RNAlater (Thermo Fisher Scientific, Waltham, Massachussets, USA). A total of 56 adenocarcinoma tissue samples from the colon (Ac) and 58 non-neoplastic colon tissue samples from the surrounding areas (NM) were collected via sterile dissection during surgical resection. A total of 2 g of amoxicillin / clavulanic acid was administered to the patients at the beginning of the laparoscopic and postoperatory stages (a total of 3 doses every 8 h). All tissue samples were immediately stored in the presence of 500 pL of RNAlater reagent (Thermo Fisher Scientific, Waltham, Massachussets, USA) for nucleic acid extraction and sequencing. All samples were stored at -80 °C until further analysis. Oral and dental check-up

[0115] As cited above, GCF samples from CRC patients who agreed to undergo an oral examination were collected by a periodontist at the Pardinas Medical Dental Clinic (A Coruna, Galicia, Spain). Only 19 of 93 went to the dental check-ups. Periodontal assessment for each patient was done according to the 2017 World Workshop on the Classification of Periodontal and Peri -Implant Diseases and Conditions

[0054] , The gingival condition of each patient was recorded using the Silness-Loe gingival index

[0055] , The Decayed, Missing, and Filled Teeth (DMFT) index was assessed for each patient using intraoral examination and Cone-beam Computed Tomography (CBCT; Carestream Dental LLC, Atlanta, USA). The presence or absence of periapical lesions was ruled out by CBCT analysis in each patient.

[0116] Bacterial DNA extraction

[0117] Stool and saliva samples

[0118] First, F and S samples of CRC and non-CRC individuals were defrosted and then centrifuged (2 min at 5000 rpm and 4 °C). Next, 2 mL of each supernatant were centrifuged again (10 min at 13000 rpm at 4°C). Final pellets were resuspended in 100 pL of nuclease-free water. To induce cellular wall lysis of all bacteria, samples were incubated (37°C, at 400 rpm, 1 h) in presence of 5 pL of an enzymatic cocktail (EC) containing 20 mg / mL of lysozyme, 1.25 KU / mL of lysostaphin and 0.625 KU / mL of mutanolysin (all enzymes from Sigma-Aldrich, St. Louis, Missouri, USA). Bacterial DNA from F and S samples was extracted using the MasterPureTM Complete DNA and RNA Purification Kit (Epicentre, Madison, Wisconsin, USA).

[0119] Tissue and gingival crevicular fluid samples

[0120] Bacterial DNA from tissue samples was extracted from 20 mg of colon mucosa (non- neoplastic and adenocarcinoma tissues) using the AllPrep® DNA / RNA Mini kit (Qiagen, Hilden, Germany). Tissue homogenization was performed using Lysing Matrix E tubes (MP Biomedicals, St. Ana, California, USA) and a 1600 MiniG system (Thermo Fisher Scientific, Waltham, Massachussets, USA) at 1500 rpm for 10 min. A total of 30 pL of EC was added to samples after homogenization. In the GCF samples, only a 2 min vortex at high speed was used to detach bacteria from endodontic paper points. The paper points were carefully discarded, and 500 pL of sterile phosphate buffered saline was added. A new centrifugation at 13000 rpm for 30 min at 4 °C was carried out and the bacterial pellet was used to extract bacterial DNA following the AllPrep® DNA / RNA Mini kit manufacturer’s instructions with an additional enzymatic lysis (30 pL of EC). For all samples, the final DNA was eluted in EB buffer (Qiagen, Hilden, Germany) and stored at -20 °C until library preparation. Negative controls were used for each extraction were done to avoid contamination.

[0121] 16S rRNA metabarcoding

[0122] For microbial identification, two hypervariable regions of the 16S rRNA gene (V3-V4) were amplified by PCR using 5 ng / pL of DNA and the following oligonucleotides: 5' TCGTCGGCAGCGTCAGATGTGTATAAGAGACAGCCTACGGGNGGCWGCAG as the forward primer and 5' GTCTCGTGGGCTCGGAGATGT GTATAAGAGACAGGACTACHVGGGTATCTAATCC as the reverse primer

[0056] , All libraries were prepared following the Illumina 16S Metagenomic Sequencing Library Preparation protocol (Illumina, San Diego, California, USA). The pooled final libraries were diluted to a final concentration of 10 pM, and and 20% of 10 pM PhiX (Illumina, USA) was added for sequencing using paired end Illumina MiSeq v3 reagent kits 2x300 (Illumina, San Diego, California, USA). DNA quantification was performed using the Qubit dsDNA HS Assay Kit (Invitrogen, Waltham, Massachusetts, USA), and the library size was checked using a 2100 Bioanalyzer Instrument (Agilent Technologies, St. Clara, California, USA). Negative controls were included in all cases to avoid contamination.

[0123] Bioinformatic analysis

[0124] The quality of all FASTQ files generated from 16S rRNA gene sequencing was checked using FastQC

[0057] , Afterwards, sequences were analyzed using QIIME2 (version 2021.11)

[0058] on a per sequencing run basis, utilizing DADA2 to trim, denoise, correct sequencing errors and remove chimeras, resulting in several tables of Amplicon Sequence Variants (ASVs)

[0059] , The resulting features from each sequencing run were then merged and collapsed into a single feature table and classified using the SILVA 138 99% reference database

[0060] through QIIME2. Afterwards, R (version 4.1)

[0061] and Phyloseq (version 1.36.0)

[0062] were used to create a Phyloseq object from QIIME 2 results, to process it and to clean it. Mainly, filtering ASVs to those from the bacterial kingdom, subtracting the raw count of ASVs that appeared in control samples from the rest on a per sequencing run basis and removing ASVs in specific genera that are typically involved in reagent contamination

[0063] , Additionally, taxonomy levels were propagated (e.g., An ASV that was classified at the genus level but not classified at the species level had their last known level propagated, filling the species level with Genus NA). To calculate the mean relative abundance for each type of sample the cleaned and ratified Phyloseq was used. Filters composed of a minimum abundance of 0.01% and minimum prevalence of 30% were applied, firstly in a per biosample basis (each biosample can be sequenced multiple times) and then in a per group basis (e.g., feces of CRC patients), obtaining the mean relative abundance in multiple taxonomy levels. Alpha and beta diversity analyses were also performed through R and Phyloseq. Differential abundance analysis (DAA) was performed using R package ANCOM-BC (version 2.0.1)

[0064] on the cleaned, non- rarified Phyloseq object, adjusting the p-values by the Holm-Bonferroni method

[0065] and using a prevalence cut of 30%, except in the test of agglomerations of oral genera, in which prevalence filter was removed to show every genus of interest.

[0125] To perform the co-occurrence study raw counts were normalized with ANCOM-BC

[0064] and species with less than 0.01% of mean abundance or in less than 30% of the samples were filtered out. Paired samples (S, GCF, NM, and Ac from the same CRC individuals) were used to study the intra- and inter-niche correlations at the bacterial level. To assess the correlations among the bacteria in the oral samples, unsupervised sPCA from mixOmics R package was performed. To elucidate the associations of oral bacteria in tumors a multivariate analysis (sPLS-canonical) from the mixOmics R package was applied using the normalized dataset of all bacterial counts from the tissue sample and a subset of the oral -associated bacteria present in the samples

[0066] ,

[0126] Normalized data from 47 MN tissues were used as input for DMM algorithm

[0067] to identify the optimal number of clusters (tissue enterotypes) based on Laplace approximation. Subsequently, Bray Curtis distances were calculated and for graphical visualization, an ordination technique was performed with non-metric multidimensional scaling (NMDS) method

[0068] , Taxonomic differences among enterotypes were assessed using Wilcoxon rank sum paired tests and canonical correlation analysis (CCA) was used to plot these differences using R

[0061] ,

[0127] Logarithmic regression models were used to evaluate the discriminatory capacity of bacteria as biomarkers in feces. Accuracy was evaluated using Receiver Operating Characteristic (ROC) curves and the area under the curve (AUC) was calculated and validated using the leave -one-out cross-validation (LOOCV) method. Additionally, the bootstrapping algorithm implemented in the Boruta R library (Kursa and Rudnicki, 2010) was used to blind the selection of additional biomarkers. The analysis was performed at the genus level.

[0128] RESULTS

[0129] CRC patients and healthy controls characteristics

[0130] A total of 93 patients with CRC and 30 healthy individuals (non-CRC) were included in the present study. CRC diagnosis was confirmed by colonoscopy and histopathological analysis in all cases. The characteristics of both groups are summarized in Tables 1 and

[0131] 2. Figure 1 summarizes the workflow of this study.

[0132] Table 1. Characteristics of the colorectal cancer patients’ cohort (CRC, n = 93) and the control group (non-CRC; n = 30).

[0133] Colorectal Control Total p- cancer value

[0134] (CRC) (non-CRC) (CRC + non-CRC)

[0135] Sex Females 36 (38.71) 23 (76.67) 59 (47.97) 0.020 n (%) Males 57 (61.29) 7 (23.33) 64 (52.03)

[0136] Age (years) Females 65 63 65 0.107

[0137] Median Males 70 64 67 0.053

[0138] Height Females 170 160 160 0.984

[0139] (cm)

[0140] Median Males 169 167.5 169 0.281

[0141] Weight Females 71 67 68 0.491

[0142] (kg)

[0143] Median Males 77 82 77.5 0.635

[0144] Location Cecum 6 (6.45) - 6 (6.45) n (%) Ascending 29 (31.18) 29 (31.18) colon

[0145] Hepatic 1 (1.07) - 1 (1.07) flexure

[0146] Transvers 3 (3.22) - 3 (3.22) e colon

[0147] Splenic 2 (2.15) - 2 (2.15) flexure

[0148] Descendin 22 (23.65) - 22 (23.65) g colon

[0149] Sigmoid 18 (19.35) - 18 (19.35) colon

[0150] Rectum 10 (10.75) - 10 (10.75)

[0151] Undetermi 2 (2.15) - 2 (2.15) ned

[0152] CRC T [1-4] T1 (17.20); T2 T1 (17.20); T2 Stage (12.90); (12.90); (TNM) T3 (62.36); T4 T3 (62.36); T4 0 / (5.38) (5.38) / o

[0153] N [0-5] NO (64.51); N1 NO (64.51); N1 (17.20); N2 (17.20); N2 (7.53); N3 (7.53); N3 (2.15);

[0154] (2.15); N4 N4 (2.15); N5

[0155] (2.15); N5 (2.15)

[0156] (2.15)

[0157] MO (92.47); MO (92.47); Ml

[0158] Ml (2.15); (2.15);

[0159] M [0-3]

[0160] M2 (2.15); M3 M2 (2.15); M3 (1-07) (1-07)

[0161] Undetermi 2.06 2.06 ned

[0162] Bristol Type 1 0 (0) 0 (0) 0 (0) 0.399

[0163] Stool Scale n (%) Type 2 8 (8.60) 3 (10) 11 (8.94)

[0164] Type 3 10 (10.76) 1 (3.33) 11 (8.94)

[0165] Type 4 38 (40.86) 19 (63.34) 57 (46.35) Type 5 8 (8.60) 3 (10) 11 (8.94)

[0166] Type 6 10 (10.75) 1 (3.33) 11 (8.94)

[0167] Type 7 5 (5.38) 1 (3.33) 6 (4.88)

[0168] Undetermi 14 (15.05) 2 (6.67) 16 (13.01) ned

[0169] Periodonta No 33 (35.48) 14 (46.67) 47 (38.21) 0.535

[0170] 1 Disease

[0171] (PD) n (%) Yes 54 (58.07) 15 (50) 69 (56.10)

[0172] Undetermi 6 (6.45) 1 (3.33) 7 (5.69) ned

[0173] Physical No 44 (47.31) 21 (70) 65 (52.85) 0.083

[0174] Activity n (%) Yes 40 (43.01) 7 (23.33) 47 (38.21)

[0175] Undetermi 9 (9.68) 2 (6.67) 11 (8.94) ned

[0176] Alcohol No 54 (58.06) 21 (70) 75 (60.98) 0.560

[0177] Consumpti on n (%) Yes 30 (32.26) 7 (23.33) 37 (30.08)

[0178] Undetermi 9 (9.68) 2 (6.67) 11 (8.94) ned

[0179] Tobacco No 70 (75.27) 22 (73.33) 92 (74.8) 0.788

[0180] Consumpti on n (%) Yes 14 (15.05) 6 (20) 20 (16.26)

[0181] Undetermi 9 (9.68) 2 (6.67) 11 (8.94) ned Caffeine No 66 (70.97) 19 (63.93) 85 (69.11) 0.511

[0182] Consumpti on n (%) Yes 18 (19.35) 9 (29.40) 27 (21.95)

[0183] Undetermi 9 (9.68) 2 (6.67) 11 (8.94) ned

[0184] Sleep No 58 (62.37) 23 (76.67) 81 (65.85) 0.280

[0185] Disorders n (%) Yes 24 (25.81) 6 (20) 30 (24.39)

[0186] Undetermi 11 (11.82) 1 (3.33) 12 (9.76) ned

[0187] Environm City 31 (33.33) 18 (60) 49 (39.84) ent

[0188] 0.551

[0189] (Most Countrysi 45 (48.39) 8 (26.67) 53 (43.09) habitual) de n (%) Both 8 (8.60) 3 (10) 11 (8.94)

[0190] Undetermi 9 (9.68) 1 (3.33) 10 (8.13) ned

[0191] Table 2. Oral health characteristics of the colorectal cancer (CRC) patients who attended the oral and dental check-up.

[0192] Colorectal Cancer (CRC) n %

[0193] Periodontal Disease IA 1 5.26

[0194] Gingivitis

[0195] (PD)6IB 1 5.26

[0196] Initial IIA 3 15.79

[0197] Periodontitis IIB 3 15.79

[0198] Mild IIIA 4 21.06

[0199] Periodontitis IIIB 1 5.26

[0200] Progressive IVA 0 0 Periodontitis IVB 3 15.79

[0201] No PD 2 10.53

[0202] No teeth11 5.26

[0203] Gingival Index20.1 - 1 7 36.84

[0204] 1.1 - 2 10 52.64

[0205] 2.1 - 3 1 5.26

[0206] No teeth11 5.26

[0207] Tooth Loss 1 - 5 9 47.36

[0208] (Number of missing teeth) 6 - 10 4 21.06

[0209] 11-15 4 21.06

[0210] > 15 1 5.26

[0211] No teeth11 5.26

[0212] Caries Index 1 - 10 4 21.06

[0213] (DMFT Index: decayed, 11 - 20 12 63.15 missing, filled) > 20 2 10.53

[0214] No teeth11 5.26

[0215] Periapical lesion No 14 73.68

[0216] Yes 5 26.32

[0217] 'No teeth: colorectal cancer (CRC) patient with completely missing teeth due to acute periodontitis (dental implants only).2Gingival Index: Level 0.1 - 1: Mild degree of gingival inflammation, slight gingival color change, no bleeding. Level 1.1 - 2: Moderate degree of inflammation, gingival reddening and swelling, gingival bleeding on probing and pressure. Level 2.1 - 3: Strong level of inflammation, intense gingival reddening and swelling, plentiful bleeding, and the possibility of ulceration.

[0218] A total of 377 samples were sequenced and analyzed using different bioinformatic procedures. In the CRC group, 93 fecal samples (F), 93 saliva samples (S), 19 gingival crevicular fluid samples (GCF), 58 normal colorectal mucosa samples (NM) and 56 adenocarcinoma samples (Ac) were obtained. For the non-CRC group, 30 F and 28 S samples were processed and analysed. The colon distribution of the different Ac samples studied in the CRC patients revealed that 31.18% were located in the ascending colon, 23.65% in the descending colon and 19.35% in the sigmoid colon (Table 1). Most of the CRC patients of this study were diagnosed in advanced stages, being the 62.36% of them in the T3 stage and 5.38% in the T4 stage (Table 1). Nevertheless, 64.51% of the CCR patients had no lymph node invasion (NO) and 92.47% of the patients had no distant metastasis (Table 1). Interestingly, a high percentage of patients with CRC (58.07%) were admitted having oral disorders to the questionnaire (e.g., halitosis, gingivitis, periodontitis, and tartar on teeth or caries, among others). In contrast, only 23.33% of the individuals in the non-CRC control group presented dental or gum disorders (Table 1). Specifically, among CRC patients who agreed to attend the oral check-up, 89.47% had periodontal disease in early or advanced stages (Table 2). Further characteristics of the studied cohorts (CRC and non-CRC) are summarized in Tables 1 and 2.

[0219] Microbiome composition of CRC patients

[0220] Microbiome analysis of different samples obtained from a cohort of patients with CRC was performed (Figure 1). To achieve this, the S and F samples from CRC patients were compared with samples from non-CRC individuals. The median number of reads per sample across sequencing runs was 49737 (21 million reads in total), decreasing to a median of 32104 (13 million reads in total) after quality control and DADA2 processing. This resulted in 15213 Amplicon Variant Sequences (ASVs) with a median length of 418 bp. Regarding the oral cavity microbiome of CRC patients, important periodontal pathobionts were identified specifically in GCF samples (Figure 2), being the most represented: Fusobacterium sp. (15%), Porphyromonas gingivalis (9.48%), Prevotella intermedia (3.20%), Prevotella nigrescens (2.10%), Tannerella forsythia (2.03%), Alloprevotella tannerae (1.72%), Treponema denticola (1.65%) and the emerging periodontal pathogen Filifactor alocis (1.58%) (Figure 2C).

[0221] Regarding the gut microbiome, and focusing on F samples, Ruminococcaceae, Lachnospiraceae and Bacteroidaceae were the most abundant families in the two groups (CRC and non-CRC) as shown in Figure 2A. However, Prevotellaceae, Enterob acteriaceae and Rikenellaceae were more abundant in the CRC group than in the non-CRC group (Figure 2A). Interestingly, the Fusobacteriaceae family was significantly over-enriched in the CRC samples (adjusted p-value <0.001). In contrast, in the F samples from the healthy group, the families Oscillospiraceae, Streptococcaceae and Bifidobacteriaceae were more represented compared to F samples in CRC patients (Figure 2A). Consequently, Fusobacterium and Parvimonas genera in fecal samples of CRC patients were significantly over-abundant (adjusted p-values <0.001 in both cases) compared to F samples of the non-CRC group (Figure 3). Notably, when taking sex into account these differences were still visible, being more abundant in both males and females diagnosed with CRC. More specifically, at the species level, Parvimonas sp., an important periodontal pathogen, and Bacteroides fragilis were significantly enriched (adjusted p-values <0.001 in both cases) in F samples from CRC patients when compared to F samples from non-CRC individuals (Figure 3). In contrast, focusing on ASVs level analysis, Blautia sp. and Faecalibacterium sp. were significantly more abundant in F samples from the non-CRC control group (adjusted p-values of <0.05 and <0.01, respectively) (Figure 3).

[0222] When the analysis was focused only on oral related microorganisms, the genera Fusobacterium and Parvimonas appeared across all types of samples except in F of non- CRC individuals (Figure 4A), showing more abundance in GCF, Ac, NM and F of CRC patients (Figure 4A). Differential abundance analysis of specific combinations of ASVs integrating typical oral related genera was conducted using data obtained from F samples from CRC and non-CRC individuals. The data revealed that the group formed by Parvimonas, Fusobacterium, Prevotella, Peptostreptococcus and Porphyromonas was over-abundant in CRC compared to non-CRC fecal samples (Figure 4B). However, only Fusobacterium, Parvimonas and Peptostreptococcus showed significant differences when used alone (Figure 3).

[0223] In addition, the microbiome differences between the Ac and NM samples obtained from the same individuals with CRC were analyzed (Figure 2). At the genus level, Fusobacterium and Prevotella were over-represented in Ac samples compared to NM samples from the same individual. In contrast, the Ruminococcus gnavus and torques groups, Anaerostipes and Blautia were more abundant in NM than in Ac (Figure 2B). At the species level, the results demonstrated that Bacteroides fragilis was enriched in Ac tissue (8.7%) when compared to NM samples (5%) of CRC patients, in accordance with results observed in F samples of CRC patients when compared to samples from non-CRC individuals (Figure 2C).

[0224] No significant differences in alpha diversity among the CRC and non-CRC groups for F and S samples were detected. When quantifying the influence of covariables in betadiversity measures, F appeared to be influenced by patient group (CRC, non-CRC; p- value <0.01) and age (p-value <0.05), but not sex. Analysis of Ac samples revealed significant influences of age (p-value <0.05) and sex (p-value <0.05), but not of the location (right vs. left colon), tumor size, metastasis, or lymph node affectation. Additionally, S samples were not influenced by patient group, age, or sex. Bacterial co-ocurrence

[0225] To corroborate the hypothesis that oral microbes translocate in complex multispecies clusters, a correlation study was conducted. Paired S, GCF, NM, and Ac samples from the same CRC individuals were used to study intra- and interniche bacterial correlations. In GCF samples from CRC patients, members of the red and orange Socransky complexes, well-known as late colonizers and periodontopathogens, including Fusobacterium, Parvimonas and Tannerella forsythia, were positively correlated with each other. Members of the green and purple complexes, known as early colonizers and health-associated, also clustered together (Figure 5A). Interestingly, in S samples of CRC patients, Fusobacterium periodonticum correlated positively with members or the red complex (i.e. Porphyromonas) as well as with other facultative anaerobes (i.e., Gemella) and to a lesser extent with aerobes (i.e., Rothia aeria) (Figure 5B).

[0226] When studying the correlations of oral species in NM samples from CRC patients, common patterns with S samples were detected. In this niche, proteolytic species from the orange complex (Fusobacterium, Parvimonas and Peptostreptococcus) and facultative anaerobes (Gemella, Granulicatella and Streptococcus) were positively correlated (Figure 6A, 7A). In the Ac samples, the same proteolytic species (Fusobacterium, Parvimonas and Peptostreptococcus) clustered together and had little or no correlation with facultative anaerobes (apart from Gemella), resembling the subgingival niche. It is worth mentioning that in tumors, the proteolytic oral cluster also correlated with intestinal microorganisms, such as unclassified Hungatella and Bacteroides fragilis (positively) or Agathobacter and Faecalibacterium (negatively) (Figure 6B, 7B). Therefore, oral proteolytic bacteria have a correlation pattern in the gut, which seems to be a pattern found in oral niches.

[0227] Finally, the correlation between bacteria in saliva or subgingival fluids and their presence in tumor tissue was analysed, and no strong correlations were found. In fact, only Gemella and Veillonella correlated positively with itself when comparing the Ac and S samples (Figure 8). Therefore, these results suggest that the levels of oral bacteria in the gut are not related to their corresponding levels in the oral cavity.

[0228] Gut enterotypes and colonization of oral bacteria

[0229] To evaluate whether certain bacterial communities in the colon are more likely to interact or facilitate colonization by oral bacteria, we analyzed the distinct enterotypes in NM samples. The DMM algorithm identified the optimal number of enterotypes in the NM tissues as two based on the Laplace approximation and NMD ordination (Figure 9A). Although some overlapping was observed between the clusters, the differences among them were significant. Indeed, the canonical correlation analysis (CCA) performed for the two possible enterotypes in the control / affected tissue pairs showed a higher distance among the enterotypes than among the control / affected tissue pairs (Figure 9B). Patients classified in enterotype 1 or 2 did not differ in the clinical parameters evaluated (age, sex, tumor status, or location) (Table 3).

[0230] Table 3. Clinical parameters of patient classified by tissue cluster (enterotype).

[0231] Cluster-1 Cluster-2 -value

[0232] Right 14 11

[0233] Localization* , „ 0.15

[0234] Left 9 13

[0235] T1 2 4

[0236] T2 4 2

[0237] Tumor size 0.19

[0238] T3 16 17

[0239] T4 1 0

[0240] Spread to NO 16 12 lymph nodes „T„

[0241] JN > 0 7 12

[0242] MO 21 23

[0243] Metastasis ,, „ „ , 0.45

[0244] M > 0 2 1 Male 17 16

[0245] Sexn0.5

[0246] Female 6 8

[0247] Age Mean 68.78 66.29 0.15

[0248] *The right colon refers to: the cecum, ascending colon, hepatic flexure and transverse colon. The left colon refers to: the splenic flexure, descending colon, sigmoid colon and rectum.

[0249] Differences in the bacterial composition were evaluated according to the enterotypes assigned to the tumor tissues. A total of 39 species were found at significantly different levels between tumors of enterotype 1 and 2. Among them, only an uncultured Fusobacterium could potentially have an oral origin, and is not one of the most prevalent or abundant.

[0250] Determination of the best bacterial consortium for CRC diagnostics

[0251] The role of Fusobacterium as a potential biomarker in fecal samples was tested in our cohort. ROC curves were utilized to assess its discriminatory power, revealing an AUC of 0.65 when only Fusobacterium was included in the model (Figure 10). Interestingly, the inclusion in the model of other over-abundant bacteria in CRC feces like Parvimonas (oral) and Bacteroides (intestinal), improved the discriminatory power between CRC patients and non-CRC individuals, obtaining an AUC of 0.75 (Figure 10). However, the addition of other over-abundant bacteria in CRC feces like Peptostreptoccocus did not increase the efficiency of the model, whereas the addition of healthy associated bacteria like Blautia or Faecalibacterium increased the AUC value up to 0.77 and 0.8, respectively. Feature selection using the Boruta algorithm confirmed the efficiency of Fusobacterium and Parvimonas as biomarkers for CRC. However, this algorithm suggested a combination of six other genera, namely, Incertae sedis, Odoribacter, Faecalitalea, UCG- 010, Slackia and Parvimonas, which achieved an AUC of 0.86, although the relative proportions of most of these poorly characterized bacteria were very low in the samples. Therefore, the present invention provides new evidence that oral pathobionts, probably forming synergistic consortia, colonize the colorectal mucosa, and can be detected in fecal samples. The periodontal bacterial associations reported in the present work may enhance the development of a pro-inflammatory microenvironment, collaborating with the onset of tumorigenesis. In the present invention, we propose that the cluster formed by Fusobacterium, Parvimonas, Bacteroides and Faecalibacterium can be used as an excellent biomarker for early diagnosis of CRC. We also aimed to highlight that PD is a risk factor for CRC initiation. Therefore, oral treatments could contribute to reducing the incidence and prevalence of cancer; however, CRC screening for patients diagnosed with PD may be key to improving the early diagnosis of CRC.

[0252] REFERENCES

[0253] 1. International Agency for Research on Cancer (IARC), World Health Organization (WHO). Global Cancer Observatory (GLOBOCAN) 2020. Lyon (France): IARC; 2020 [accessed 2022 Dec 12], https: / / gco.iarc.fr / .

[0254] 2. Asociacion Espanola Contra el Cancer. Informe dinamico: Cancer de Colon. Madrid (Spain): Asociacion Espanola Contra el Cancer; 2023 Jan 01 [accessed 2022 Sep 26], https: / / observatorio.contraelcancer.es / informes / informe-dinamico-cancer-de-colon.

[0255] 3. Kong F & Cai Y (2019) Study Insights into Gastrointestinal Cancer through the Gut Microbiota. BioMed Research International 2019, 8721503, doi: 10.1155 / 2019 / 8721503.

[0256] 4. Louis P, Hold GL & Flint HJ (2014) The gut microbiota, bacterial metabolites and colorectal cancer. Nature Reviews Microbiology 12, 661-672, doi: 10.1038 / nrmicro3344. 5. Salehiniya H, Pouyesh V, Tarazoj AA, Mohammadian-Hafshejani A, Aghajani M, Yousefi SM & Gandomani HS (2017) Colorectal cancer in the world: incidence, mortality and risk factors. Biomedical Research and Therapy 4, doi: 10.15419 / bmrat.v4il0.372.

[0257] 6. Sawicki T, Ruszkowska M, Danielewicz A, Niedzwiedzka E, Arlukowicz T & Przybylowicz KE (2021) A review of colorectal cancer in terms of epidemiology, risk factors, development, symptoms and diagnosis. Cancers 13, 2025, doi: 10.3390 / cancersl3092025.

[0258] 7. Servizo Galego de Saude (SERGAS). Programa de detection precoz do cancro colorrectal. Galicia (Spain): SERGAS;2013 [accessed 2022 Sep 28], https: / / www.sergas.es / Saude-publica / Programa-de-detecci%C3%B3n-precoz-do- cancro-col orrectal ?i di oma=es .

[0259] 8. Burnett-Hartman AN, Lee JK, Demb J & Gupta S (2021) An update on the epidemiology, molecular characterization, diagnosis, and screening strategies for early- onset colorectal cancer. Gastroenterology 160, 1041-1049, doi: 10.1053 / j.gastro.2020.12.068.

[0260] 9. Siegel RL, Torre LA, Soerjomataram I, Hayes RB, Bray F, Weber TK & Jemal A (2019) Global patterns and trends in colorectal cancer incidence in young adults. Gut 68, 2179-2185, doi: 10.1136 / gutjnl-2019-319511.

[0261] 10. Vuik FE, Nieuwenburg SA, Bardou M, Lansdorp-Vogelaar I, Dinis-Ribeiro M, Bento MJ, Zadnik V, Pellise M, Esteban L, Kaminski MF, Suchanek S, Ngo O, Majek O, Leja M, Kuipers EJ & Spaander MC (2019) Increasing incidence of colorectal cancer in young adults in Europe over the last 25 years. Gut 68, 1820-1826, doi: 10.1136 / gutjnl- 2018-317592.

[0262] 11. Elsafi SH, Alqahtani NI, Zakary NY & Al Zahrani EM (2015) The sensitivity, specificity, predictive values, and likelihood ratios of fecal occult blood test for the detection of colorectal cancer in hospital settings. Clinical experimental gastroenterology, 279-284.

[0263] 12. Watson AJ & Collins PD (2011) Colon cancer: a civilization disorder. Digestive Diseases 29, 222-228, doi: 10.1159 / 000323926.

[0264] 13. Ahmad AF, Dwivedi G, O’Gara F, Caparros-Martin J & Ward NC (2019) The gut microbiome and cardiovascular disease: current knowledge and clinical potential. Am J Physiol Heart Circ Physiol 317, H923-H938, doi: 10.1152 / ajpheart.00376.2019.

[0265] 14. Chok KC, Ng KY, Koh RY & Chye SM (2021) Role of the gut microbiome in Alzheimer’s disease. Reviews in the Neurosciences 32, 767-789, doi: 10.1515 / revneuro- 2020-0122.

[0266] 15. Cryan JF, O'Riordan KJ, Sandhu K, Peterson V & Dinan TG (2020) The gut microbiome in neurological disorders. The Lancet Neurology 19, 179-194, doi:

[0267] 10.1016 / S 1474-4422( 19)30356-4.

[0268] 16. du Teil Espina M, Gabarrini G, Harmsen HJM, Westra J, van Winkelhoff AJ & van Dijl JM (2018) Talk to your gut: the oral-gut microbiome axis and its immunomodulatory role in the etiology of rheumatoid arthritis. FEMS microbiology reviews 43, 1-18, doi: 10.1093 / femsre / fuy035

[0269] 17. Durack J & Lynch SV (2018) The gut microbiome: Relationships with disease and opportunities for therapy. Journal of Experimental Medicine 216, 20-40, doi : 10.1084 / jem.20180448

[0270] 18. Gasmi Benahmed A, Gasmi A, Do§a A, Chirumbolo S, Mujawdiya PK, Aaseth J, Dadar M & Bjorklund G (2021) Association between the gut and oral microbiome with obesity. Anaerobe 70, 102248, doi: 10.1016 / j. anaerobe.2020.102248.

[0271] 19. Peirce JM & Alvina K (2019) The role of inflammation and the gut microbiome in depression and anxiety. Journal of Neuroscience Research 97, 1223-1241, doi: 10.1002 / jnr.24476.

[0272] 20. Stokholm J, Blaser MJ, Thorsen J, Rasmussen MA, Waage J, Vinding RK, Schoos A-MM, Kunoe A, Fink NR, Chawes BL, Bonnelykke K, Brejnrod AD, Mortensen MS, Al-Soud WA, Sorensen SJ & Bisgaard H (2018) Maturation of the gut microbiome and risk of asthma in childhood. Nature Communications 9, 141, doi: 10.1038 / s41467-017- 955 02573-2.

[0273] 21. Mohammadi M, Mirzaei H & Motallebi M (2021) The role of anaerobic bacteria in the development and prevention of colorectal cancer: A review study. Anaerobe, 102501, doi: 10.1016 / j. anaerobe.2021.102501.

[0274] 22. Nejman D, Livyatan I, Fuks G, Gavert N, Zwang Y, Geller LT, Rotter-Maskowitz A, Weiser R, Mallei G, Gigi E, Meltser A, Douglas GM, Kamer I, Gopalakrishnan V, Dadosh T, Levin-Zaidman S, Avnet S, Atlan T, Cooper ZA, Arora R, Cogdill AP, Khan MAW, Ologun G, Bussi Y, Weinberger A, Lotan-Pompan M, Golani O, Perry G, Rokah M, Bahar-Shany K, Rozeman EA, Blank CU, Ronai A, Shaoul R, Amit A, Dorfman T, Kremer R, Cohen ZR, Harnof S, Siegal T, Yehuda-Shnaidman E, Gal -Yam EN, Shapira H, Baldini N, Langille MGI, Ben-Nun A, Kaufman B, Nissan A, Golan T, Dadiani M, Levanon K, Bar J, Yust-Katz S, Barshack I, Peeper DS, Raz DJ, Segal E, Wargo JA, Sandbank J, Shental N & Straussman R (2020) The human tumor microbiome is composed of tumor type-specific intracellular bacteria. Science (New York, NY) 368, 973-980, doi: 10.1126 / science.aay9189.

[0275] 23. Osman MA, Neoh HM, Ab Mutalib NS, Chin SF, Mazlan L, Raja Ali RA, Zakaria AD, Ngiu CS, Ang MY & Jamal R (2021) Parvimonas micra, Peptostreptococcus stomatis, Fusobacterium nucleatum and Akkermansia muciniphila as a four-bacteria biomarker panel of colorectal cancer. Scientific Reports 11, 2925, doi: 10.1038 / s41598- 021-82465-0.

[0276] 24. Parhi L, Alon-Maimon T, Sol A, Nejman D, Shhadeh A, Fainsod-Levi T, Yajuk O, Isaacson B, Abed J, Maalouf N, Nissan A, Sandbank J, Yehuda-Shnaidman E, Ponath F, Vogel J, Mandelboim O, Granot Z, Straussman R & Bachrach G (2020) Breast cancer colonization by Fusobacterium nucleatum accelerates tumor growth and metastatic progression. Nature Communications 11, 3259, doi: 10.1038 / s41467-020-16967-2.

[0277] 25. Purcell RV, Visnovska M, Biggs PJ, Schmeier S & Frizelle FA (2017) Distinct gut microbiome patterns associate with consensus molecular subtypes of colorectal cancer. Scientific Reports 7, 11590, doi: 10.1038 / s41598-017-11237-6.

[0278] 26. Wong SH & Yu J (2019) Gut microbiota in colorectal cancer: mechanisms of action and clinical applications. Nature Reviews Gastroenterology & Hepatology 16, 690-, doi: 10.1038 / s41575-019-0209-8.

[0279] 27. Schmidt TSB, Raes J & Bork P (2018) The Human Gut Microbiome: From Association to Modulation. Cell 172, 1198-1215, doi: 10.1016 / j cell.2018.02.044.

[0280] 28. Wong SH & Yu J (2019) Gut microbiota in colorectal cancer: mechanisms of action and clinical applications. Nature Reviews Gastroenterology & Hepatology 16, 690-704, doi: 10.1038 / s41575-019-0209-8.

[0281] 29. Yang Y, Nguyen M, Khetrapal V, Sonnert ND, Martin AL, Chen H, Kriegel MA & Palm NW (2022) Within-host evolution of a gut pathobiont facilitates liver translocation. Nature, doi: 10.1038 / s41586-022-04949-x.

[0282] 30. Bertocchi A, Carloni S, Ravenda PS, Bertalot G, Spadoni I, Lo Cascio A, Gandini S, Lizier M, Braga D, Asnicar F, Segata N, Klaver C, Brescia P, Rossi E, Anselmo A, Guglietta S, Maroli A, Spaggiari P, Tarazona N, Cervantes A, Marsoni S, Lazzari L, Jodice MG, Luise C, Erreni M, Pece S, Di Fiore PP, Viale G, Spinelli A, Pozzi C, Penna G & Rescigno M (2021) Gut vascular barrier impairment leads to intestinal bacteria dissemination and colorectal cancer metastasis to liver. Cancer Cell 39, 708-724 e711, doi: 10.1016 / j. ccell.2021.03.004. 31. Koliarakis I, Messaritakis I, Nikolouzakis TK, Hamilos G, Souglakos J & Tsiaoussis J (2019) Oral bacteria and intestinal dysbiosis in colorectal cancer. International Journal of Molecular Sciences 20, 4146.

[0283] 32. Xu J, Yang M, Wang D, Zhang S, Yan S, Zhu Y & Chen W (2020) Alteration of the abundance of Parvimonas micra in the gut along the adenoma-carcinoma sequence. Oncology Letters 20, 106, doi: 10.3892 / ol.2020.11967.

[0284] 33. Akimoto N, Ugai T, Zhong R, Hamada T, Fujiyoshi K, Giannakis M, Wu K, Cao Y, Ng K & Ogino S (2021) Rising incidence of early-onset colorectal cancer — a call to action. Nature Reviews Clinical Oncology 18, 230-243, doi: 10.1038 / s41571-020-00445- 1.

[0285] 34. Hofseth LJ, Hebert JR, Chanda A, Chen H, Love BL, Pena MM, Murphy EA, Sajish M, Sheth A, Buckhaults PJ & Berger FG (2020) Early-onset colorectal cancer: initial clues and current views. Nature Reviews Gastroenterology & Hepatology 17, 352-364, doi: 10.1038 / s41575-019-0253-4.

[0286] 35. Abed J, Maalouf N, Manson AL, Earl AM, Parhi L, Emgard JEM, Klutstein M, Tayeb S, Almogy G, Atlan KA, Chaushu S, Israeli E, Mandelboim O, Garrett WS & Bachrach G (2020) Colon cancer-associated Fusobacterium nucleatum may originate from the oral cavity and reach colon tumors via the circulatory system. Frontiers in Cellular and Infection Microbiology 10, 400, doi: 10.3389 / fcimb.2020.00400.

[0287] 36. Flemer B, Warren RD, Barrett MP, Cisek K, Das A, Jeffery IB, Hurley E, O'Riordain M, Shanahan F & O'Toole PW (2018) The oral microbiota in colorectal cancer is distinctive and predictive. Gut 67, 1454-1463, doi: 10.1136 / gutjnl-2017-314814.

[0288] 37. Mesa F, Mesa-Lopez MJ, Egea-Valenzuela J, Benavides-Reyes C, Nibali L, Ide M, Mainas G, Rizzo M & Magan-Fernandez A (2022) A new comorbidity in periodontitis: Fusobacterium nucleatum and colorectal cancer. Medicine 58, doi: 10.3390 / medicina58040546.

[0289] 38. Wang N & Fang JY (2022) Fusobacterium nucleatum, a key pathogenic factor and microbial biomarker for colorectal cancer. Trends in microbiology, doi: 10.1016 / j.tim.2022.08.010.

[0290] 39. Brentar M, Kuijvenhoven J, Stokkers P & Struys E (2022) The potential of fecal microbiota and amino acids to detect and monitor patients with adenoma. Gut microbes 14, 2038863. 40. Clos-Garcia M, Garcia K, Alonso C, Iruarrizaga-Lejarreta M, D’ Amato M, Crespo A, Iglesias A, Cubiella J, Bujanda L & Falcon-Perez JM (2020) Integrative analysis of fecal metagenomics and metabolomics in colorectal cancer. Cancers 12, 1142.

[0291] 41. Eklbf V, Lbfgren-Burstrbm A, Zingmark C, Edin S, Larsson P, Karling P,

[0292] Alexeyev O, Rutegard J, Wikberg ML & Palmqvist R (2017) Cancer-associated fecal microbial markers in colorectal cancer detection. International journal of cancer 141, 2528-2536.

[0293] 42. Goedert JJ, Gong Y, Hua X, Zhong H, He Y, Peng P, Yu G, Wang W, Ravel J & Shi J (2015) Fecal microbiota characteristics of patients with colorectal adenoma detected by screening: a population-based study. EBioMedicine 2, 597-603.

[0294] 43. Tarallo S, Ferrero G, Gallo G, Francavilla A, Clerico G, Realis Luc A, Manghi P, Thomas AM, Vineis P & Segata N (2019) Altered fecal small RNA profiles in colorectal cancer reflect gut microbiome composition in stool samples. Msystems 4, e00289-00219.

[0295] 44. Yachida S, Mizutani S, Shiroma H, Shiba S, Nakajima T, Sakamoto T, Watanabe H, Masuda K, Nishimoto Y, Kubo M, Hosoda F, Rokutan H, Matsumoto M, Takamaru H, Yamada M, Matsuda T, Iwasaki M, Yamaji T, Yachida T, Soga T, Kurokawa K, Toyoda A, Ogura Y, Hayashi T, Hatakeyama M, Nakagama H, Saito Y, Fukuda S, Shibata T & Yamada T (2019) Metagenomic and metabolomic analyses reveal distinct stage-specific phenotypes of the gut microbiota in colorectal cancer. Nature Medicine 25, 968-976, doi: 10.1038 / s41591-019-0458-7.

[0296] 45. Young C, Wood HM, Fuentes Balaguer A, Bottomley D, Gallop N, Wilkinson L, Benton SC, Brealey M, John C & Burtonwood C (2021) Microbiome analysis of more than 2,000 nhs bowel cancer screening programme samples shows the potential to improve screening accuracymicrobiome analysis improves crc screening accuracy. Clinical Cancer Research 27, 2246-2254.

[0297] 46. Zeller G, Tap J, Voigt AY, Sunagawa S, Kultima JR, Costea PI, Amiot A, Bohm J, Brunetti F & Habermann N (2014) Potential of fecal microbiota for early-stage detection of colorectal cancer. Molecular Systems Biology 10, 766.

[0298] 47. Drewes JL, White JR, Dejea CM, Fathi P, lyadorai T, Vadivelu J, Roslani AC, Wick EC, Mongodin EF, Loke MF, Thulasi K, Gan HM, Goh KL, Chong HY, Kumar S, Wanyiri JW & Sears CL (2017) High-resolution bacterial 16S rRNA gene profile metaanalysis and biofilm status reveal common colorectal cancer consortia. NPJ Biofilms Microbiomes 3, 34, doi: 10.1038 / s41522-017-0040-3. 48. Mangifesta M, Mancabelli L, Milani C, Gaiani F, de’ Angelis N, de’ Angelis GL, van Sinderen D, Ventura M & Turroni F (2018) Mucosal microbiota of intestinal polyps reveals putative biomarkers of colorectal cancer. Scientific Reports 8, 13974, doi: 10.1038 / s41598-018-32413-2.

[0299] 49. Wang Y, Zhang Y, Qian Y, Xie Y-H, Jiang S-S, Kang Z-R, Chen Y-X, Chen Z-F & Fang J-Y (2021) Alterations in the oral and gut microbiome of colorectal cancer patients and association with host clinical factors. 149, 925-935, doi: 10.1002 / ijc.33596.

[0300] 50. Zhang M, Lv Y, Hou S, Liu Y, Wang Y & Wan X (2021) Differential mucosal microbiome profiles across stages of human colorectal cancer. Life 11, 831.

[0301] 51. Lbwenmark T, Lbfgren-Burstrbm A, Zingmark C, Eklbf V, Dahlberg M, Wai S, Larsson P, Ljuslinder I, Edin S & Palmqvist R (2020) Parvimonas micra as a putative non-invasive faecal biomarker for colorectal cancer. Scientific Reports 10, 15250, doi: 10.1038 / s41598-020-72132-l.

[0302] 52. Komiya Y, Shimomura Y, Higurashi T, Sugi Y, Arimoto J, Umezawa S, Uchiyama S, Matsumoto M & Nakajima A (2019) Patients with colorectal cancer have identical strains of Fusobacterium nucleatum in their colorectal cancer and oral cavity. Gut 68, 1335-1337, doi: 10.1136 / gutjnl-2018-316661.

[0303] 53. Xuan K, Jha AR, Zhao T, Uy JP & Sun C (2021) Is periodontal disease associated with increased risk of colorectal cancer? A meta-analysis. International Journal of Dental Hygiene 19, 50-61, doi: 10.1111 / idh.12483.

[0304] 54. Papapanou PN, Sanz M, Buduneli N, Dietrich T, Feres M, Fine DH, Flemmig TF, Garcia R, Giannobile WV, Graziani F, Greenwell H, Herrera D, Kao RT, Kebschull M, Kinane DF, Kirkwood KL, Kocher T, Komman KS, Kumar PS, Loos BG, Machtei E, Meng H, Mombelli A, Needleman I, Offenbacher S, Seymour GJ, Teles R & Tonetti MS (2018) Periodontitis: Consensus report of workgroup 2 of the 2017 World Workshop on the Classification of Periodontal and Peri-Implant Diseases and Conditions. Journal of clinical periodontology 45 Suppl 20, S162-sl70, doi: 10.1111 / jcpe.12946.

[0305] 55. Loe H (1967) The Gingival Index, the Plaque Index and the Retention Index Systems. Journal of periodontology 38, Suppl :610-616, doi: 10.1902 / jop.1967.38.6.610.

[0306] 56. Klindworth A, Pruesse E, Schweer T, Peplies J, Quast C, Hom M & Glbckner FO (2013) Evaluation of general 16S ribosomal RNA gene PCR primers for classical and next-generation sequencing-based diversity studies. Nucleic acids research 41, el, doi: 10.1093 / nar / gks808. 57. Andrews S. FastQC. California (USA): GitHub Inc, 2023 Apr 25. [accessed 2023 Jun 10] . https: / / github . com / s-andrews / FastQC .

[0307] 58. Bolyen E, Rideout JR, Dillon MR, Bokulich NA, Abnet CC, Al-Ghalith GA,

[0308] 1103 Alexander H, Alm EJ, Arumugam M, Asnicar F, Bai Y, Bisanz JE, Bittinger K, Brejnrod A, Brislawn CJ, Brown CT, Callahan BJ, Caraballo-Rodriguez AM, Chase J, Cope EK, Da Silva R, Diener C, Dorrestein PC, Douglas GM, Durall DM, Duvallet C, Edwardson CF, Ernst M, Estaki M, Fouquier J, Gauglitz JM, Gibbons SM, Gibson DL, Gonzalez A, Gorlick K, Guo J, Hillmann B, Holmes S, Holste H, Huttenhower C, Huttley GA, Janssen S, Jarmusch AK, Jiang L, Kaehler BD, Kang KB, Keefe CR, Keim P, Kelley ST, Knights D, Koester I, Kosciolek T, Kreps J, Langille MGI, Lee J, Ley R, Liu Y-X, Loftfield E, Lozupone C, Maher M, Marotz C, Martin BD, McDonald D, McIver LJ, Melnik AV, Metcalf JL, Morgan SC, Morton JT, Naimey AT, Navas-Molina JA, Nothias LF, Orchanian SB, Pearson T, Peoples SL, Petras D, Preuss ML, Pruesse E, Rasmussen LB, Rivers A, Robeson MS, Rosenthal P, Segata N, Shaffer M, Shiffer A, Sinha R, Song SJ, Spear JR, Swafford AD, Thompson LR, Torres PJ, Trinh P, Tripathi A, Turnbaugh PJ, Ul-Hasan S, van der Hooft JJJ, Vargas F, Vazquez-Baeza Y, Vogtmann E, von Hippel M, Walters W, Wan Y, Wang M, Warren J, Weber KC, Williamson CHD, Willis AD, Xu ZZ, Zaneveld JR, Zhang Y, Zhu Q, Knight R & Caporaso JG (2019) Reproducible, interactive, scalable and extensible microbiome data science using QIIME 2. Nature Biotechnology 37, 852-857, doi: 10.1038 / s41587-019-0209-9.

[0309] 59. Callahan BJ, McMurdie PJ, Rosen MJ, Han AW, Johnson AJA & Holmes SP (2016) DADA2: High-resolution sample inference from Illumina amplicon data. Nature Methods 13, 581-583, doi: 10.1038 / nmeth.3869.

[0310] 60. Quast C, Pruesse E, Yilmaz P, Gerken J, Schweer T, Yarza P, Peplies J & Glockner FO (2013) The SILVA ribosomal RNA gene database project: improved data processing and web-based tools. Nucleic acids research 41, D590-596, doi: 10.1093 / nar / gksl219.

[0311] 61. Ihaka R & Gentleman R (1996) R: A Language for Data Analysis and Graphics.

[0312] Journal of Computational and Graphical Statistics 5, 299-314, doi:

[0313] 10.1080 / 10618600.1996.10474713.

[0314] 62. McMurdie PJ & Holmes S (2013) phyloseq: an R package for reproducible interactive analysis and graphics of microbiome census data. PloS one 8, e61217, doi: 10.1371 / journal. pone.0061217.

[0315] 63. Salter SJ, Cox MJ, Turek EM, Calus ST, Cookson WO, Moffatt MF, Turner P, Parkhill J, Loman NJ & Walker AW (2014) Reagent and laboratory contamination can critically impact sequence-based microbiome analyses. BMC Biology 12, 87, doi: 10.1186 / sl2915-014-0087-z.

[0316] 64. Lin H & Peddada SD (2020) Analysis of compositions of microbiomes with bias correction. Nature Communications 11, 3514, doi: 10.1038 / s41467-020-17041-7.

[0317] 65. Holm S (1979) A simple sequentially rejective multiple test procedure. Scandinavian Journal of Statistics 6, 65-70.

[0318] 66. Rohart F, Gautier B, Singh A & Le Cao K-A (2017) mixOmics: An R package for ‘omics feature selection and multiple data integration. PLoS computational biology 13, el005752, doi: 0.1371 / joumal.pcbi.1005752.

[0319] 67. Holmes I, Harris K & Quince C (2012) Dirichlet multinomial mixtures: generative models for microbial metagenomics. PloS one 7, e30126.

[0320] 68. Oksanen J SG, Blanchet F, Kindt R, Legendre P, Minchin P, O'Hara R, Solymos P, Stevens M, Szoecs E, Wagner H, Barbour M, Bedward M, Bolker B, Borcard D, Carvalho G, Chirico M, De Caceres M, Durand S, Evangelista H, FitzJohn R, Friendly M, Furneaux B, Hannigan G, Hill M, Lahti L, McGlinn D, Ouellette M, Ribeiro Cunha E, Smith T, Stier A, Ter Braak C, Weedon J (2023) vegan: Community Ecology Package.

[0321] 69. Mattiuzzi C, Sanchis-Gomar F & Lippi G (2019) Concise update on colorectal cancer epidemiology. Annals of translational medicine 7, 609, doi: 10.21037 / atm.2019.07.91.

[0322] 70. Fan X, Jin Y, Chen G, Ma X & Zhang L (2021) Gut microbiota dysbiosis drives the development of colorectal cancer. Digestion 102, 508-515, doi: 10.1159 / 000508328.

[0323] 71. Conde-Perez K, Buetas E, Aja-Macaya P, Arribas E, Iglesias-Corras I, Trigo-Tasende N, Nasser-Ali M, Estevez LS, Rumbo-Feal S & Otero-Alen B (2022) Parvimonas micra can translocate from the subgingival sulcus of the human oral cavity to colorectal adenocarcinoma. Molecular Oncology, accepted author manuscript, doi: 10.1002 / 1878- 0261.13506.

[0324] 72. Jin S, Wetzel D & Schirmer M (2022) Deciphering mechanisms and implications of bacterial translocation in human health and disease. Current Opinion in Microbiology 67, 102147.

[0325] 73. Yu TC, Zhou YL, Fang JYJJoG & Hepatology (2022) Oral pathogen in the pathogenesis of colorectal cancer. 37, 273-279.

[0326] 74. Shen X, Li J, Li J, Zhang Y, Li X, Cui Y, Gao Q, Chen X, Chen Y & Fang JY (2021) Fecal enterotoxigenic Bacteroides fragilis-Peptostreptococcus stomatis-Parvimonas micra biomarker for noninvasive diagnosis and prognosis of colorectal laterally spreading tumor. Frontiers in Oncology 11, 661048, doi: 10.3389 / fonc.2021.661048. 75. Ang MY, Dy mock D, Tan JL, Thong MH, Tan QK, Wong GJ, Paterson IC & Choo SW (2013) Genome Sequence of Parvimonas micra Strain A293, Isolated from an Abdominal Abscess from a Patient in the United Kingdom. 1, e01025-01013, doi: doi: 10.1128 / genomeA.01025-13.

[0327] 76. Hu L, Liu Y, Kong X, Wu R, Peng Q, Zhang Y, Zhou L & Duan L (2021) Fusobacterium nucleatum facilitates M2 macrophage polarization and colorectal carcinoma progression by activating TLR4 / NF-KB / S 100A9 cascade. Front Immunol 12, 658681, doi: 10.3389 / fimmu.2021.658681.

[0328] 77. Long X, Wong CC, Tong L, Chu ESH, Ho Szeto C, Go MYY, Coker OO, Chan AWH, Chan FKL, Sung JJY & Yu J (2019) Peptostreptococcus anaerobius promotes colorectal carcinogenesis and modulates tumour immunity. Nature Microbiology 4, 2319- 2330, doi: 10.1038 / s41564-019-0541-3.

[0329] 78. Lowenmark T, Lofgren-Burstrom A, Zingmark C, Ljuslinder I, Dahlberg M, Edin S & Palmqvist R (2022) Tumour colonisation of Parvimonas micra is associated with decreased survival in colorectal cancer patients. Cancers 14, doi: 10.3390 / cancersl4235937.

[0330] 79. Mu W, Jia Y, Chen X, Li H, Wang Z, Cheng BJFiC & Microbiology I (2020) Intracellular Porphyromonas gingivalis promotes the proliferation of colorectal cancer cells via the MAPK / ERK signaling pathway. 10, 584798.

[0331] 80. Van Dalen PJ, Van Deutekom-Mulder EC, De Graaff J & Van Steenbergen TJ (1998) Pathogenicity of Peptostreptococcus micros morphotypes and Prevotella species in pure and mixed culture. Journal of medical microbiology 47, 135-140, doi: 10.1099 / 00222615- 47-2-135.

[0332] 81. Wang X, Jia Y, Wen L, Mu W, Wu X, Liu T, Liu X, Fang J, Luan Y, Chen P, Gao J, Nguyen K-A, Cui J, Zeng G, Lan P, Chen Q, Cheng B & Wang Z (2021) Porphyromonas gingivalis promotes colorectal carcinoma by activating the hematopoietic NLRP3 inflammasome. Cancer research 81, 2745-2759, doi: 10.1158 / 0008-5472. CAN-20-3827 %J Cancer Research.

[0333] 82. Zhao L, Zhang X, Zhou Y, Fu K, Lau HC-H, Chun TW-Y, Cheung AH-K, Coker OO, Wei H, Wu WK-K, Wong SH, Sung JJ-Y, To KF & Yu J (2022) Parvimonas micra promotes colorectal tumorigenesis and is associated with prognosis of colorectal cancer patients. Oncogene, doi: 10.1038 / s41388-022-02395-7.

[0334] 83. Castellarin M, Warren RL, Freeman JD, Dreolini L, Krzywinski M, Strauss J, Barnes R, Watson P, Allen-Vercoe E, Moore RA & Holt RA (2012) Fusobacterium nucleatum infection is prevalent in human colorectal carcinoma. Genome Res 22, 299-306, doi: 10.1101 / gr.126516.111.

[0335] 84. Hertel J, Heinken A, Martinelli F & Thiele I (2021) Integration of constraint-based modeling with fecal metabolomics reveals large deleterious effects of Fusobacterium spp. on community butyrate production. Gut microbes 13, 1-23, doi:

[0336] 10.1080 / 19490976.2021.1915673.

[0337] 85. Kostic AD, Gevers D, Pedamallu CS, Michaud M, Duke F, Earl AM, Ojesina Al, Jung J, Bass AJ, Tabernero J, Baselga J, Liu C, Shivdasani RA, Ogino S, Birren BW, Huttenhower C, Garrett WS & Meyerson M (2012) Genomic analysis identifies association of Fusobacterium with colorectal carcinoma. Genome Res 22, 292-298, doiilO.l 101 / gr.l26573.111.

[0338] 86. Mima K, Nishihara R, Qian ZR, Cao Y, Sukawa Y, Nowak JA, Yang J, Dou R, Masugi Y & Song M (2016) Fusobacterium nucleatum in colorectal carcinoma tissue and patient prognosis. Gut 65, 1973-1980.

[0339] 87. Rubinstein MR, Wang X, Liu W, Hao Y, Cai G & Han YW (2013) Fusobacterium nucleatum promotes colorectal carcinogenesis by modulating E-cadherin / p-catenin signaling via its FadA adhesin. Cell host & microbe 14, 195-206, doi: 10.1016 / j.chom.2013.07.012.

[0340] 88. Serna G, Ruiz-Pace F, Hernando J, Alonso L, Fasani R, Landolfi S, Comas R, Jimenez J, Elez E, Bullman S, Tabernero J, Capdevila J, Dienstmann R & Nuciforo P (2020) Fusobacterium nucleatum persistence and risk of recurrence after preoperative treatment in locally advanced rectal cancer. Annals of oncology: official journal of the European Society for Medical Oncology 31, 1366-1375, doi: 10.1016 / j.annonc.2020.06.003.

[0341] 89. Tahara T, Yamamoto E, Suzuki H, Maruyama R, Chung W, Garriga J, Jelinek J, Yamano HO, Sugai T, An B, Shureiqi I, Toyota M, Kondo Y, Estecio MR & Issa JP (2014) Fusobacterium in colonic flora and molecular features of colorectal carcinoma. Cancer research 74, 1311-1318, doi: 10.1158 / 0008-5472.Can-13-1865.

[0342] 90. Xu C, Fan L, Lin Y, Shen W, Qi Y, Zhang Y, Chen Z, Wang L, Long Y, Hou T, Si J & Chen S (2021) Fusobacterium nucleatum promotes colorectal cancer metastasis through miR-1322 / CCL20 axis and M2 polarization. Gut microbes 13, 1980347, doi: 10.1080 / 19490976.2021.1980347.

[0343] 91. Yu T, Guo F, Yu Y, Sun T, Ma D, Han J, Qian Y, Kryczek I, Sun D, Nagarsheth N, Chen Y, Chen H, Hong J, Zou W & Fang JY (2017) Fusobacterium nucleatum Promotes Chemoresistance to Colorectal Cancer by Modulating Autophagy. Cell 170, 548- 563.e516, doi: 10.1016 / j.cell.2017.07.008.

[0344] 92. Nakatsu G, Li X, Zhou H, Sheng J, Wong SH, Wu WK, Ng SC, Tsoi H, Dong Y, Zhang N, He Y, Kang Q, Cao L, Wang K, Zhang J, Liang Q, Yu J & Sung JJ (2015) Gut mucosal microbiome across stages of colorectal carcinogenesis. Nat Commun 6, 8727, doi: 10.1038 / ncomms9727.

[0345] 93. Yu J, Feng Q, Wong SH, Zhang D, Liang Qy, Qin Y, Tang L, Zhao H, Stenvang J, Li Y, Wang X, Xu X, Chen N, Wu WKK, Al-Aama J, Nielsen HJ, Kiilerich P, Jensen BAH, Yau TO, Lan Z, Jia H, Li J, Xiao L, Lam TYT, Ng SC, Cheng AS-L, Wong VW-S, Chan FKL, Xu X, Yang H, Madsen L, Datz C, Tilg H, Wang J, Brunner N, Kristiansen K, Arumugam M, Sung JJ-Y & Wang J (2017) Metagenomic analysis of faecal microbiome as a tool towards targeted non-invasive biomarkers for colorectal cancer. Gut 66, 70-78, doi: 10.1136 / gutjnl-2015-309800 %J Gut.

[0346] 94. Boleij A, Hechenbleikner EM, Goodwin AC, Badani R, Stein EM, Lazarev MG, Ellis B, Carroll KC, Albesiano E, Wick EC, Platz EA, Pardoll DM & Sears CL (2015)The Bacteroides fragilis toxin gene is prevalent in the colon mucosa of colorectal cancer patients. Clinical infectious diseases : an official publication of the Infectious Diseases Society of America 60, 208-215, doi: 10.1093 / cid / ciu787.

[0347] 95. Bachem A, Makhlouf C, Binger KJ, de Souza DP, Tull D, Hochheiser K, Whitney PG, Fernandez-Ruiz D, Dahling S, Kastenmuller W, Jonsson J, Gressier E, Lew AM, Perdomo C, Kupz A, Figgett W, Mackay F, Oleshansky M, Russ BE, Parish IA, Kallies A, McConville MJ, Turner SJ, Gebhardt T & Bedoui S (2019) Microbiota-derived shortchain fatty acids promote the memory potential of antigen-activated CD8+ T cells. Immunity 51, 285-297. e285, doi: 10.1016 / j.immuni.2019.06.002.

[0348] 96. Liu X, Mao B, Gu J, Wu J, Cui S, Wang G, Zhao J, Zhang H & Chen W (2021) Blautia — a new functional genus with potential probiotic properties? Gut microbes 13, 1875796, doi: 10.1080 / 19490976.2021.1875796.

[0349] 97. Luu K, Ye JY, Lagishetty V, Liang F, Hauer M, Sedighian F, Kwaan MR, Kazanjian KK, Hecht JR, Lin AY & Jacobs JP (2023) Fecal and tissue microbiota are associated with tumor T-cell infiltration and mesenteric lymph node involvement in colorectal cancer. Nutrients 15, 316.

[0350] 98. Luu M, Riester Z, Baldrich A, Reichardt N, Yuille S, Busetti A, Klein M, Wempe A, Leister H, Raifer H, Picard F, Muhammad K, Ohl K, Romero R, Fischer F, Bauer CA, Huber M, Gress TM, Lauth M, Danhof S, Bopp T, Nerreter T, Mulder IE, Steinhoff U, Hudecek M & Visekruna A (2021) Microbial short-chain fatty acids modulate CD8+ T cell responses and improve adoptive immunotherapy for cancer. Nature Communications 12, 4077, doi: 10.1038 / s41467-021-24331-l.

[0351] 99. Chen L, Wang W, Zhou R, Ng SC, Li J, Huang M, Zhou F, Wang X, Shen B, M AK, Wu K & Xia B (2014) Characteristics of fecal and mucosa-associated microbiota inChinese patients with inflammatory bowel disease. Medicine 93, e51, doi : 10.1097 / md.0000000000000051.

[0352] 100. Kim CH, Park J & Kim M (2014) Gut microbiota-derived short-chain Fatty acids, T cells, and inflammation. Immune network 14, 277-288, doi: 10.4110 / in.2014.14.6.277.

[0353] 101. Maioli TU, Borras-Nogues E, Torres L, Barbosa SC, Martins VD, Langella P, Azevedo VA & Chatel JM (2021) Possible benefits of faecalibacterium prausnitzii for obesity-associated gut disorders. Frontiers in pharmacology 12, 740636, doi: 10.3389 / fphar.2021.740636.

[0354] 102. Miquel S, Martin R, Rossi O, Bermudez-Humaran LG, Chatel JM, Sokol H, Thomas M, Wells JM & Langella P (2013) Faecalibacterium prausnitzii and human intestinal health. Current Opinion in Microbiology 16, 255-261, doi: 10.1016 / j.mib.2013.06.003.

[0355] 103. Zhang J, He Y, Xia L, Yi J, Wang Z, Zhao Y, Song X, Li J, Liu H & Liang X (2022) Expansion of colorectal cancer biomarkers based on gut bacteria and viruses. Cancers 14, 4662, doi: 10.3390 / cancers 14194662.

[0356] 104. Espana CGdCdDd, 2022. Atlas de Salud Bucodental en Espana. Una Hamada a la accion., in: Comunicacion Gid (Ed.).

[0357] 105. Simon-Soro A, Ren Z, Krom BP, Hoogenkamp MA, Cabello-Yeves PJ, Daniel SG, Bittinger K, Tomas I, Koo H & Mira A (2022) Polymicrobial aggregates in human saliva build the oral biofilm. mBio 13, e0013122, doi: 10.1128 / mbio.00131-22.

[0358] 106. Kolenbrander PE, Palmer Jr RJ, Rickard AH, Jakubovics NS, Chalmers NI & Diaz PI (2006) Bacterial interactions and successions during plaque development. Periodontology 2000 42, 47-79, doi: 10.1111 / j.l600-0757.2006.00187.x.

[0359] 107. Kolenbrander PE, Palmer RJ, Periasamy S & Jakubovics NS (2010) Oral multispecies biofilm development and the key role of cell-cell distance. Nature Reviews Microbiology 8, 471-480, doi: 10.1038 / nrmicro2381.

[0360] 108. Flynn KJ, Baxter NT & Schloss PD (2016) Metabolic and community synergy of oral bacteria in colorectal cancer. mSphere 1, doi: 10.1128 / mSphere.00102-16. 109. Galeano Nino JL, Wu H, LaCourse KD, Kempchinsky AG, Baryiames A, Barber B, Futran N, Houlton J, Sather C & Sicinska E (2022) Effect of the intratumoral microbiota on spatial and cellular heterogeneity in cancer. Nature 611, 810-817.

[0361] 110. El-Awady A, de Sousa Rabelo M, Meghil MM, Rajendran M, Elashiry M, Stadler AF, Foz AM, Susin C, Romito GA, Arce RM & Cutler CW (2019) Polymicrobial synergy within oral biofilm promotes invasion of dendritic cells and survival of consortia members, npj Biofilms and Microbiomes 5, 11, doi: 10.1038 / s41522-019-0084-7.

[0362] 111. Farrugia C, Stafford GP, Gains AF, Cutts AR & Murdoch C (2022) Fusobacterium nucleatum mediates endothelial damage and increased permeability following single species and polymicrobial infection. 93, 1421-1433, doi: 10.1002 / JPER.21-0671.

[0363] 112. Socransky S, Haffajee A, Cugini M, Smith C & Kent Jr R (1998) Microbial complexes in subgingival plaque. Journal of clinical periodontology 25, 134-144, doi: 10.1111 / j. l600-051x.l998.tb02419.x.

[0364] 113. Mark Welch JL, Rossetti BJ, Rieken CW, Dewhirst FE & Borisy GG (2016) Biogeography of a human oral microbiome at the micron scale. Proceedings of the National Academy of Sciences 113, E791-E800, doi: 10.1073 / pnas.1522149113.

[0365] 114. Bradshaw DJ, Marsh PD, Allison C & Schilling KM (1996) Effect of oxygen, inoculum composition and flow rate on development of mixed-culture oral biofilms. Microbiology 142, 623-629.

[0366] 115. Rawat PS, Li Y, Zhang W, Meng X & Liu W (2022) Hungatella hathewayi, an efficient glycosaminoglycan-degrading Firmicutes from human gut and Its chondroitin ABC exolyase with high activity and broad substrate specificity. Applied Environmental Microbiology 88, e01546-01522, doi: 10.1128 / aem.01546-22.

[0367] 116. Mori M, Ponce-de-Leon M, Pereto J & Montero F (2016) Metabolic complementation in bacterial communities: necessary conditions and optimality. Frontiers in microbiology 7, 1553, doi: 10.3389 / fmicb.2016.01553.

[0368] 117. Sears CL & Pardoll DM (2011) Perspective: alpha-bugs, their microbial partners, and the link to colon cancer. Journal of Infectious Diseases 203, 306-311, doi : 10.1093 / jinfdis / jiq061.

[0369] 118. Mirzaei R, Mirzaei H, Alikhani MY, Sholeh M, Arabestani MR, Sai dijam M,

[0370] Karampoor S, Ahmadyousefi Y, Moghadam MS, Irajian GR, Hasanvand H & Yousefimashouf R (2020) Bacterial biofilm in colorectal cancer: What is the real mechanism of action? Microbial pathogenesis 142, 104052, doi: 10.1016 / j.micpath.2020.104052. 119. Parsaei M, Sarafraz N, Moaddab SY & Ebrahimzadeh Leylabadlo H (2021) The importance of Faecalibacterium prausnitzii in human health and diseases. New microbes and new infections 43, 100928, doi: 10.1016 / j.nmni.2021.100928.

[0371] 120. Rosero JA, Killer J, Sechovcova H, Mrazek J, Benada O, Fliegerova K, Havlik J & Kopecny J (2016) Reclassification of Eubacterium rectale (Hauduroy et al. 1937) Prevot 1938 in a new genus Agathobacter gen. Nov. As agathobacter rectalis comb. Nov., and description of Agathobacter ruminis sp. Nov., isolated from the rumen contents of sheep and cows. International journal of systematic and evolutionary microbiology 66, 768-773, doi: 10.1099 / ijsem.0.000788.

[0372] 121. Mirzaei R, Afaghi A, Babakhani S, Sohrabi MR, Hosseini-Fard SR, Babolhavaeji K, Khani Ali Akbari S, Yousefimashouf R & Karampoor S (2021) Role of microbiota- derived short-chain fatty acids in cancer development and prevention. Biomedicine & Pharmacotherapy 139, 111619, doi: 10.1016 / j .biopha.2021.111619.

[0373] 122. Lv M, Zhang J, Deng J, Hu J, Zhong Q, Su M, Lin D, Xu T, Bai X, Li J & Guo X (2023) Analysis of the relationship between the gut microbiota enterotypes and colorectal adenoma. Frontiers in microbiology 14, Original Research, doi: 10.3389 / fmicb.2023.1097892.

[0374] 123. Zhao R, Xia D, Chen Y, Kai Z, Ruan F, Xia C, Gong J, Wu J & Wang X (2023) Improved diagnosis of colorectal cancer using combined biomarkers including Fusobacterium nucleatum, fecal occult blood, transferrin, CEA, CAI 9-9, gender, and age. Cancer medicine 12, 14636-14645, doi: 10.1002 / cam4.6067.

[0375] Example 2.

[0376] Methods

[0377] 1.1 Recruitment of participants and fecal sampling

[0378] All volunteers provided informed signed consent prior to the initiation of the sample collection phase, which was performed in the University Hospital of A Coruna (HU AC; Galicia, Spain). Strict adherence to clinical guidelines and regulations was maintained throughout the recruitment period (Research Ethical Committee of Galicia, Spain: code CEIm-G 2018 / 609). A total of 93 CRC diagnosed subjects were enrolled in the project between 2019 and 2022. Certain inclusion criteria were followed as previously described

[0044] : (a) no antibiotic treatment within the last month, (b) no infectious disease, (c) no chemotherapy and / or radiotherapy treatments prior to sample collection, (d) no genetic predisposition to CRC development, (e) no intestinal inflammatory disorders, (f) no immunological diseases, (g) no medical history of transplantation, (h) not currently undergoing immunosuppressive treatment. Moreover, a total of 30 CRC cancer-free patients’ companions / couples were asked to participate in the study, meeting the same inclusion requirements as the CRC group. Samples (n=123), each containing approximately 20 mL of fecal material, were self-collected by each patient at home and preserved in 10 mL of RNAlater reagent (Thermo Fisher Scientific, Waltham, MA, USA). A preliminary interview with each volunteer was conducted to gather individual data (e.g.: age, sex, weight or height, among others) and lifestyle habits (e.g.: dietary patterns or physical activity, among others). Fecal samples were stored at -80°C until DNA extraction.

[0379] 1.2 DNA extraction

[0380] As previously outlined fecal samples underwent a brief pre-processing protocol prior to the DNA extraction procedure. The MasterPure™ Complete DNA / RNA Purification Kit (Epicentre, USA) was used in accordance with the manufacturer’s instructions.

[0381] 1.3 16S rRNA metabarcoding sequencing

[0382] 1.3.1 Illumina (MiSeq™) 16S rRNA V3-V4

[0383] Two highly variable regions within the 16S rRNA gene (V3-V4) were selectively amplified through PCR, employing 5’

[0384] TCGTCGGCAGCGTCAGATGTGTATAAGAGACAGCCTACGGGNGGCWGCAG as forward primer and 5’

[0385] GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAGGACTACHVGGGTATCTA ATCC as reverse primer. Nuclease-free water was included as negative control in each PCR reaction to avoid bacterial contaminations. Subsequently, libraries were constructed by following the Illumina 16S Metagenomic Sequencing Library Preparation protocol (Illumina, San Diego, CA, USA). Library concentration was measured using a Qubit dsDNA HS Assay Kit (Invitrogen, USA) and a Qubit 2.0 fluorometer (Invitrogen, USA). Libraries were pooled and diluted to a final concentration of 10 pM and mixed with 20% of 10 pM PhiX control (Illumina, USA). Samples were finally sequenced using a MiSeq Reagent Kit v3 (600 cycles) (Illumina, USA) and a MiSeq platform (Illumina, USA).

[0386] 1.3.2 Oxford Nanopore (MinlON™) 16S rRNA V1-V9. Complete bacterial 16S rRNA hypervariable regions (VI -V9) were amplified, using the following forward and reverse primers: 5’ AGMGTTYGATYMYGGCTCAG and 5’ TACGGYTACCTTGTTACGACTT, respectively. Negative controls (nuclease-free water) were included to avoid contaminations. For each PCR reaction, 200 fmol of fecal DNA was used. Afterwards, DNA Oxford Nanopore libraries were constructed by following the Native Barcoding Kit 96 (SQK-NBD114.96) protocol (Oxford Nanopore, Oxford, UK). In the same way as Illumina libraries, DNA concentration was assessed by using a Qubit 2.0 fluorometer with the corresponding Qubit dsDNA HS Assay Kit (Invitrogen, USA). Pooled libraries were loaded onto R10.4.1 Flow Cells (FLO-MINI 14) and sequenced for 72 h following the manufacturer’s instructions, detailed in the Native Barcoding Kit 96 protocol (Oxford Nanopore, Oxford, UK).

[0387] 1.4 Bioinformatic analysis

[0388] Oxford Nanopore Technologies (ONT) reads were first basecalled using Dorado duplex (v. 0.5.3) with three models: fast, hac and sup (dna, rlO.4.1, e8.2, 400bps, v4.1.0). The resulting reads were then separated into simplex, which were demultiplexed using Dorado (kit SQK-NBD 114-96), and duplex reads (“consensus” of two parental simplex reads). Afterwards, duplex reads in which parental reads had different barcodes or which lengths differed more than 50% were removed. The resulting reads were quality controlled with chopper (v. 0.5.0), trimming 20 nt from the front and back and selecting reads between 1200-1900 nt with a specific minimum average quality score, depending on the basecalling model (fast: Q7; hac and sup: Q12). Additionally, Duplex Tools (v. 0.2.9)

[0048] was used to detect and remove reads with mid-strand adapters. Host contamination was assessed with Kraken2 using its Standard 64Gb database

[0049] , Lastly, reads were identified using Emu (v. 3.4.5)

[0013] with its Default database (rrnDB v. 5.6 combined with NCBI 16S and SILVA (v. 138.1).

[0389] Illumina reads were analyzed through QIIME 2 (v. 2021.11), where DADA2 was used to trim, denoise, correct sequencing errors and remove chimeras on a per sequencing run basis, producing Amplicon Sequence Variants (ASVs), which were then classified using a feature classifier created with RESCRIPt (v. 2021.11.0) and SILVA (v. 138.1). Additionally, contamination in these ASVs was also assessed with Kraken2’s 64Gb database. Results from both ONT and Illumina were merged and analyzed in R (v. 4.2.0), mainly through Phyloseq (v. 1.42.0)

[0058] for data management, ANCOM-BC (v. 2.0.1) for differential abundance analysis (prevalence cut of 10%, adjusting significance by Holm- Bonferroni

[0060] ) and microbiome (v. 1.20.0)

[0061] for centered log-ratio abundance normalization (CLR). In order to asses [3-diversity differences, a PERMANOVA analysis through adonis2, using a multi-dimensional scaling (MDS) and the Jensen-Shannon distance (JSD), was performed. Additionally, pairwise comparisons were conducted using Wilcoxon rank-sum tests, adjusting significance for multiple comparisons using Holm-Bonferroni. Significance values across analyses are represented as * (p < 0.05), ** (p < 0.01) or *** (p < 0.001).

[0390] For biomarker identification, automated feature selection was performed using the Boruta algorithm, with two prevalences cuts of 10% and 30%. Additionally, combinations of manually selected features were tested using a Random Forest machine learning algorithm and evaluated using the leave-one-out cross-validation method. The performance of the model was expressed through the area under the receiver operating characteristic curve (AUC) value.

[0391] 2. Results

[0392] Quality control of Illumina reads resulted in a median of 32104 reads per sample with 90% of those being Q30 and 95% being Q20, generating ASVs with a median length of 418 nt. Meanwhile, quality control of ONT reads provided a total median of ~109k reads per sample and median length of 1480 nt. Median average quality varied across basecalling models, with the sup model achieving QI 8 and some reads above Q30 (Fig. 11 A). Duplex rate for sup was on average 8±2.73 %, out of which 3.76±1.61 % were filtered. Rarefaction curves were closed at the species level on both ONT and Illumina.

[0393] Comparison of ONT-V1V9 samples regarding [3-diversity analysis (MDS+JSD) revealed no differences between basecalling models (Fig. 11B, p-value>0.8), but did show significant differences between databases (p-value<0.001). Analysis of a-diversity was significantly different for the fast model, showing higher values of observed features (Fig. 11C, p-value<0.05). Additionally, Emu’s Default database also resulted in significantly higher observed features when compared to SILVA (Fig. 11C, p-value<0.05). Moreover, significant differences were observed in the percentage of reads identified differently in respect to the sup model at each taxonomic level (p-value<0.001). For example, at species level, using the SILVA database, the identification of fast reads differed a median of 9.3% from sup, while hac differed 3.47%. Differences at higher levels such as family were smaller, having a median of 3.67% and 1.20% for fast and hac, respectively. Interestingly, Emu’s Default database had higher standard deviation compared to SILVA in this measurement. Consequently, the sup model (QI 8) was chosen for the following analyses.

[0394] When comparing Illumina-V3V4 and ONT-V1V9 with the sup basecalling model, both using the SILVA database, a-diversity at the genus level was not influenced, as shown in Fig. 12A. Regarding P-diversity, samples overlapped at the genus level and not at the species level (Fig. 12B and Fig. 12C), although PERMANOVA analysis indicated significantly different results in both cases (p-value<0.001). Normalized CLR abundance at the genus level correlated well on average between Illumina-V3V4 and ONT-V1V9 (Pearson correlation: >0.8, Fig. 2D), showing few taxa present in only one of the two approaches, which in most cases was due to slight differences in the taxonomy caused by the use of different classifiers. The percentage of feature counts identified as a known species (e.g. at species level and not as uncultured, unidentified, unclassified, Taxon NA...) varied significantly (p-value<0.001), obtaining a median of 16.75% with Illumina- V3V4, 27.74% with ONT-V1V9 (SILVA) and, interestingly, 100% with ONT-V1V9 (Default).

[0395] In Figure 13 A the abundance of three important genera, potential indicators of the presence of colorectal tumors along the large bowel, is shown across technologies (Illumina-V3V4 and ONT-V1V9) and databases (Default Emu database and SILVA). In all cases these genera are significantly more abundant in fecal samples from CRC patients than in the control group (p-value<0.001) and their abundances between approaches are similar. The percentage of samples with these taxa can be seen in Figure 13B, with very similar proportions between methods except in the case of ONT-V1V9 (SILVA) for Parvimonas, where false positives were observed in controls (these occurrences were confirmed manually, identified as Parvimonas sp., not P. micra, at the species level). When studying differential abundance analysis through ANCOM-BC (Figure 14A and Figure 14B) and differences in CLR abundance (Figure 14C) using ONT-V1V9, both databases concur in multiple CRC biomarkers, such as Parvimonas micra, Fusobacterium nucleatum, Peptostreptococcus stomatis, Peptostreptococcus anaerobius, Gemella morbillorum, Sutterella wadsworthensis, Clostridium perfringens, Bacteroides fragilis and Dialister pneumosintes. In contrast, biomarkers for healthy controls are not shared between databases, with Emu’s Default database indicating more, such as: Agathobaculum butyriciproducens, Romboutsia ilealis, Anaerostipes rhamnosivorans and Anaerocolumna cellulosilytica.

[0396] In order to identify useful combinations of CRC microbial biomarkers, multiple machine learning models were created, using both automatic and manual feature selection. Models using automatically selected features are shown in Table 4, with the Default Emu database and two different prevalence levels (10% or 30%). All of them obtained at least 0.9 AUC (Area Under Curve) with just 10 features. For example, using the Default database at 10% prevalence the most important features according to the Boruta algorithm were P. micra, A. cellulosilytica, A. rhamnosivorans, P. stomatis, A. butyriciproducens, P. anaerobius, Prevotella stercorea, Candidatus Saccharibacteria bacterium oral taxon 957 (also named Candidatus Nanosynbacter featherlites (TM7) in NCBI), Olsenella timonensis and Raoultibacter timonensis (Table 4).

[0397] However, further evaluation of these taxa at the read level showed low identity to their reference. Specifically, A. cellulosilytica and A. rhamnosivorans, which were big contributors to these models, had a maximum of 90-92% identity, which should be classified as unknown Lachnospiraceae. For this reason, manual combinations of taxa with high read identity and coverage, significant differences between groups (using ANCOM-BC and CLR abundance) or identified as relevant by the Boruta algorithm were tested. These combinations and their AUC are shown in Figure 14D (and compared to previously described combinations with Illumina-V3V4

[0044] , The usage of P. micra and F. nucleatum provides an AUC of 0.71, increasing to 0.76 by adding B. fragilis, to 0.82 by adding A. butyriciproducens and, lastly, obtaining a maximum AUC of 0.87 with a total of 14 features. All comparable combinations to Illumina-V3V4 obtained slightly higher AUC (e.g. P. micra+F. nucleatum vs. Parvimonas+Fusobacterium). Selection Features Feature AUC count

[0398] Boruta Top P. micra, A. butyriciproducens, A. cellulosilytica, O. timonensis, A. bacterium, 10 0.92

[0399] 10 (30%) R. timonensis. Streptococcus sp. A12, B. luti, Clostridium sp. BNL1100, S. vari- abile

[0400] Boruta Top P. micra, A. cellulosilytica, A. rhamnosivorans, P. stomatis, A. butyricipro- 10 0.91

[0401] 10 (10%) ducens, P. anaerobius, P. stercorea, Candidates Saccharibacteria bacterium oral taxon 957, O. timonensis, R. timonensis

[0402] Manual F. nucleatum., P. micra, B. fragilis, A. butyriciproducens 4 0.82

[0403] Top 4

[0404] Manual F. nucleatum, P. micra, B. fragilis, A. butyriciproducens, P. stomatis, P. anaer- 14 0.87

[0405] Top 14 obius, G. morbillorum, D. pneumosintes, S. wadsworthensis, C. perfringens, R. ilealis, P. clara, Longibaculum sp. KGMB06250, R. massiliensis

[0406] Table 4. Area Under the Curve (AU C) for the prediction of cancer / control with different machine learning models using 0NT-V1V9 and Emu’s Default database, based on if feature selection is automatic (Boruta, with two prevalence thresholds of 10% and 30%) or manual

[0407] Conclusion example 2

[0408] Bacterial abundance between Illumina-V3V4 and ONT-V1V9 at the genus level correlated well (R2>0.8). Nanopore sequencing identified more specific bacterial biomarkers for colorectal cancer than those obtained with Illumina, such as Parvimonas micra, Fusobacterium nucleatum, Peptostreptococcus stomatis, Peptostreptococcus anaerobius, Gemella morbillorum, Clostridium perfringens, Bacteroides fragilis and Sutterella wadsworthensis. Prediction of colorectal cancer through manual feature selection and machine learning resulted in an AUC of 0.87 with 14 species or 0.82 with just 4 species (P. micra, F. nucleatum, B. fragilis and Agathobaculum butyriciproducens).

[0409] Conclusion

[0410] Full 16S rRNA V1V9 sequencing through Oxford Nanopore and its new RIO.4.1 chemistry achieved accurate species-level bacterial identification, facilitating the discovery of more precise disease-related biomarkers and increasing the taxonomic fidelity of future microbiome analyses. CLAUSES:

[0411] 1. An in vitro method for determining the risk that a subject has colorectal cancer, comprising the following steps: a. determining the levels or the concentration of bacteria belonging to the genus Fusobacterium, Parvimonas, Bacteroides and / or Faecalibacterium, or any combination thereof, in an intestinal or stool sample isolated from said subject; and b. comparing the levels or the concentration of said bacteria in said intestinal or stool sample with one or more reference values, wherein an increase in the number of sequences of Fusobacterium, Parvimonas, and / or Bacteroides, and a decreased in the number of sequences of Faecalibacterium, in said sample with respect to said one or more reference values is indicative of an increase in the risk that the subject has colorectal cancer; wherein the levels or the concentration of bacteria belonging to the genus Fusobacterium, Parvimonas, Bacteroides and / or Faecalibacterium, are determined by performing an amplification reaction from a nucleic acid preparation derived from said sample using at least one pair of primers capable of amplifying at least one representative region of said genus, and detecting the amplification product, wherein the one or more reference values are understood to refer to the concentration and / or the total amount of each bacterium belonging to one or more of the genera proposed, in the general population or in a healthy subject.

[0412] 2. The in vitro method for determining the risk that a subject has colorectal cancer according to clause 1, wherein the method comprises: a. determining the levels or the concentration of at least the bacteria belonging to the genus Fusobacterium, Parvimonas, Bacteroides and Faecalibacterium, in an intestinal or stool sample isolated from said subject; and b. comparing the levels or the concentration of said bacteria in said intestinal or stool sample with a reference value, wherein an increase in the number of sequences of Fusobacterium, Parvimonas, and / or Bacteroides, and a decreased in the number of sequences of Faecalibacterium, in said sample with respect to said one or more reference values is indicative of an increase in the risk that the subject has colorectal cancer. 3. An in vitro method for diagnosing colorectal cancer in a subject in need thereof, comprising the following steps: a) determining the levels or the concentration of bacteria belonging to the genus Fusobacterium, Parvimonas, Bacteroides and / or Faecalibacterium, or any combination thereof, in an intestinal or stool sample isolated from said subject; and b) comparing the levels or the concentration of said bacteria in said intestinal or stool sample with one or more reference values, wherein an increase in the number of sequences of Fusobacterium, Parvimonas, and / or Bacteroides, and a decreased in the number of sequences of Faecalibacterium, in said sample with respect to said one or more reference values is indicative that the subject has colorectal cancer; wherein the levels or the concentration of bacteria belonging to the genus Fusobacterium, Parvimonas, Bacteroides and / or Faecalibacterium, are determined by performing an amplification reaction from a nucleic acid preparation derived from said sample using at least one pair of primers capable of amplifying at least one representative region of said genus, and detecting the amplification product, wherein the one or more reference values are understood to refer to the concentration and / or the total amount of each bacterium belonging to one or more of the genera proposed, in the general population or in a healthy subject.

[0413] 4. The in vitro method for diagnosing colorectal cancer in a subject in need thereof according to clause 3, wherein the method comprises the following steps: a) determining the levels or the concentration of the bacteria belonging to the genus Fusobacterium, Parvimonas, Bacteroides and Faecalibacterium, in an intestinal or stool sample isolated from said subject; and b) comparing the levels or the concentration of said bacteria in said intestinal or stool sample with a reference value, wherein an increase in the number of sequences of Fusobacterium, Parvimonas, and / or Bacteroides, and a decreased in the number of sequences of Faecalibacterium, in said sample with respect to said one or more reference values is indicative that the subject has colorectal cancer. An in vitro method for determining the risk that a subject has colorectal cancer, comprising the following steps: a. determining the levels or the concentration of bacteria belonging to the genus Fusobacterium, Parvimonas, Bacteroides and / or Faecalibacterium, or any combination thereof, in an intestinal or stool sample isolated from said subject; and b. identifying the subject as a subject at risk of developing colorectal cancer by a predictive model which correlates at least one or more of the levels identified in step (a) with representative levels of the same from samples obtained or isolated from subjects previously identified as suffering from colorectal cancer, said predictive model having been generated by training a computer with a plurality of the concentration levels of the bacteria belonging to the genus identified in step (a) from previously identified subjects having colorectal cancer, by machine learning on said plurality of levels so as to obtain representative score profiles associated with colorectal cancer. An in vitro method for determining the risk that a subject has colorectal cancer according to clause 5, wherein the method comprises the following steps: a) determining the levels or the concentration of bacteria belonging to at least all of the following genus: Fusobacterium, Parvimonas, Bacteroides and Faecalibacterium, in an intestinal or stool sample isolated from said subject; and b) identifying the subject as a subject at risk of developing colorectal cancer by a predictive model which correlates at least one or more of the levels identified in step (a) with representative levels of the same from samples obtained or isolated from subjects previously identified as suffering from colorectal cancer, said predictive model having been generated by training a computer with a plurality of the concentration levels of the bacteria belonging to the genus identified in step (a) from previously identified subjects having colorectal cancer, by machine learning on said plurality of levels so as to obtain representative score profiles associated with colorectal cancer. An in vitro method for diagnosing colorectal cancer in a subject in need thereof, comprising the following steps: a. determining the levels or the concentration of bacteria belonging to the genus Fusobacterium, Parvimonas, Bacteroides and / or Faecalibacterium, or any combination thereof, in an intestinal or stool sample isolated from said subject; and b . identifying the subj ect as a subj ect having colorectal cancer by a predictive model which correlates at least one or more of the levels identified in step (a) with representative levels of the same from samples obtained or isolated from subjects previously identified as suffering from colorectal cancer, said predictive model having been generated by training a computer with a plurality of the concentration levels of the bacteria belonging to the genus identified in step (a) from previously identified subjects having colorectal cancer, by machine learning on said plurality of levels so as to obtain representative score profiles associated with colorectal cancer. An in vitro method for diagnosing colorectal cancer in a subject in need thereof according to clause 7, wherein the method comprises the following steps: a) determining the levels or the concentration of bacteria belonging to at least all of the following genus: Fusobacterium, Parvimonas, Bacteroides and Faecalibacterium, in an intestinal or stool sample isolated from said subject; and b) identifying the subj ect as a subj ect having colorectal cancer by a predictive model which correlates at least one or more of the levels identified in step (a) with representative levels of the same from samples obtained or isolated from subjects previously identified as suffering from colorectal cancer, said predictive model having been generated by training a computer with a plurality of the concentration levels of the bacteria belonging to the genus identified in step (a) from previously identified subjects having colorectal cancer, by machine learning on said plurality of levels so as to obtain representative score profiles associated with colorectal cancer. The method according to any one of clauses 1 to 4, wherein the amplification reaction is carried out by means of a real-time polymerase chain reaction. The method according to any one of clauses 5 or 8, wherein the levels or the concentration of bacteria belonging to the genus Fusobacterium, Parvimonas, Bacteroides and / or Faecalibacterium as identified in step (a), are determined by performing an amplification reaction from a nucleic acid preparation derived from said sample using at least one pair of primers capable of amplifying at least one representative region of said genus and detecting the amplification product.

[0414] 11. The method according to any of clauses 1 to 10, wherein said method further comprises storing the results of the method in a data carrier, preferably wherein said data carrier is a computer readable medium.

[0415] 12. A computer-implemented method for determining the risk that a subject has colorectal cancer, wherein said method comprises at least the comparative step (step (b)) and optionally the provision of a result as a consequence of said comparison, as these steps are defined in the method according to any of clauses 1 to 4.

[0416] 13. A computer-implemented method for determining the risk that a subject has colorectal cancer, wherein said method comprises at least the comparative step (step (b)) and optionally the provision of a result as a consequence of said comparison, as these steps are defined in the method according to any of clauses 5 to 8.

[0417] 14. A kit capable of implementing the methodology described in any of clauses 1 to

[0418] 11.

Claims

CLAIMS1. An in vitro method for determining the risk that a subject has colorectal cancer, comprising the following steps: a. determining the levels or the concentration of bacteria belonging to any one of the genus and / or species selected from: genus Fusobacterium, preferably the species Fusobacterium nucleatum; genus Parvimonas, preferably the species Parvimonas micra; genus Bacteroides, preferably the species Bacteroides fragilis; genus Peptostreptococcus, preferably the species Peptostreptococcus stomatis and / or Peptostreptococcus anaerobius; genus Blautia, preferably the species Blautia luti; genus Agathobaculum, preferably the species Agathobaculum butyriciproducens; genus Gemella, preferably the species Gemella morbillorum; genus Clostridium, preferably the species Clostridium perfringens; genus Sutterella, preferably the species Sutterella wadsworthensis; genus Dialister, preferably the species Dialister pneumosintes; genus Romboutsia, preferably species Romboutsia ilealis; genus Paraprevotella, preferably the species Paraprevotella clara; genus Longibaculum, preferably the species Longibaculum sp. KGMB06250; genus Raoultibacter, preferably the species Raoultibacter massiliensis; and / or genus Faecalibacterium, preferably the species Faecalibacterium prausnitzii, or any combination thereof, in an intestinal or stool sample isolated from said subject; and b. comparing the levels or the concentration of said bacteria in said intestinal or stool sample with one or more reference values, wherein an increase in the number of sequences of genus Fusobacterium, preferably the species Fusobacterium nucleatum; genus Parvimonas, preferably the species Parvimonas micra; genus Bacteroides, preferably the species Bacteroides fragilis; genus Peptostreptococcus, preferably the speciesPeptostreptococcus stomatis and / or Peptostreptococcus anaerobius; genus Blautia, preferably the species Blautia luti; genus Agathobaculum, preferably the species Agathobaculum butyriciproducens; genus Gemella, preferably the species Gemella morbillorum; genus Clostridium, preferably the species Clostridium perfringens; genus Sutterella,preferably the species Sutterella wadsworthensis; genus Dialister, preferably the species Dialister pneumosintes; genus Romboutsia, preferably species Romboutsia ilealis; genus Paraprevotella, preferably the species Paraprevotella clara; genus Longibaculum, preferably the species Longibaculum sp. KGMB06250; genus Raoultibacter, preferably the species Raoultibacter massiliensis; and a decreased in the number of sequences of Faecalibacterium, preferably the species Faecalibacterium prausnitzii, in said sample with respect to said one or more reference values is indicative of an increase in the risk that the subject has colorectal cancer; wherein the said levels or the concentration of bacteria of step a) are determined by an Oxford Nanopore sequencing platform for accurate species-level bacterial identification, the method comprising:• isolating and preparing nucleic acids from the intestinal or stool sample isolated from said subject;• performing an Oxford Nanopore sequencing analysis on said nucleic acids, thereby generating sequencing data;• identifying and quantifying bacteria at the genus and / or species level in the sample based on said sequencing data; and• executing step b); wherein the one or more reference values of step b) are understood to refer to the concentration and / or the total amount of each bacterium belonging to one or more of the genera and / or species proposed, and / or any scores obtained by combining such bacterial levels, in the general population or in a healthy subject.

2. The method of claim 1, wherein the Oxford Nanopore sequencing is carried out using primers targeting the V1-V9 regions of bacterial rRNA genes (ONT-V1V9).

3. The method of any one of claims 1 or 2, wherein the determined levels or concentration of bacteria at least belong to the species selected from Fusobacterium nucleatum, Parvimonas micra and Faecalibacterium prausnitzii.

4. The method of claim 3, wherein the determined levels or concentration of bacteria further belong to the species Bacteroides fragilis.

5. The method of claim 4, wherein the determined levels or concentration of bacteria further belong to the species: Agathobaculum butyr iciproducens.

6. The method of claim 5, wherein the determined levels or concentration of bacteria further belong to the species: Peptostreptococcus stomatis, Peptostreptococcus anaerobius, Gemella morbillorum, Dialister pneumosintes, Sutterella wadsworthensis, Clostridium perfringens, and Romboutsia ilealis.

7. The method of claim 6, wherein the determined levels or concentration of bacteria further belong to the species Paraprevotella clara.

8. The method of claim 7, wherein the determined levels or concentration of bacteria further belong to the species Longibaculum sp. KGMB06250.

9. The method of claim 8, wherein the determined levels or concentration of bacteria further belong to the species Raoultibacter massiliensis.

10. An in vitro method for diagnosing colorectal cancer in a subject in need thereof, comprising the following steps: a) determining the levels or the concentration of the bacteria according to step a) of any one of claims 1 to 9 wherein the said levels or the concentration of bacteria of step a) are determined by an Oxford Nanopore sequencing platform in accordance with any one of claims 1 to 9; and b) comparing the levels or the concentration of said bacteria in said intestinal or stool sample with one or more reference values, wherein an increase or a decrease, as defined in step b) of any one of claims 1 to 9, in said sample with respect to said one or more reference values is indicative that the subject has colorectal cancer; wherein the one or more reference values of step b) are understood to refer to the concentration and / or the total amount of each bacterium belonging to one or more of thegenera proposed, and / or any scores obtained by combining such bacterial levels, in the general population or in a healthy subject.

11. An in vitro method for determining the risk that a subject has colorectal cancer, wherein the method comprises: a) determining the levels or the concentration of the bacteria according to any one of claims 1 to 9 wherein the said levels or the concentration of bacteria of step a) are determined by an Oxford Nanopore sequencing platform in accordance with any one of claims 1 to 9; and b) identifying the subject as a subject at risk of developing colorectal cancer by a predictive model which correlates at least one or more of the levels identified in step (a) with representative levels of the same from samples obtained or isolated from subjects previously identified as suffering from colorectal cancer, said predictive model having been generated by training a computer with a plurality of the concentration levels of the bacteria belonging to the genus identified in step (a) from previously identified subjects having colorectal cancer, by machine learning on said plurality of levels so as to obtain representative score profiles associated with colorectal cancer.

12. An in vitro method for diagnosing colorectal cancer in a subject in need thereof, comprising the following steps: a. determining the levels or the concentration of bacteria belonging to the genus Fusobacterium, Parvimonas, Bacteroides and Faecalibacterium in an intestinal or stool sample isolated from said subject; and b . identifying the subj ect as a subj ect having colorectal cancer by a predictive model which correlates at least one or more of the levels identified in step (a) with representative levels of the same from samples obtained or isolated from subjects previously identified as suffering from colorectal cancer, said predictive model having been generated by training a computer with a plurality of the concentration levels of the bacteria belonging to the genus identified in step (a) from previously identified subjects having colorectal cancer, by machine learning on said plurality oflevels so as to obtain representative score profiles associated with colorectal cancer.

13. The method according to any of claims 1 to 12, wherein said method further comprises storing the results of the method in a data carrier, preferably wherein said data carrier is a computer readable medium.

14. A computer-implemented method for determining the risk that a subject has colorectal cancer, wherein said method comprises at least receiving the levels or concentrations of the bacteria according to step a) and the comparative step (step (b)) and optionally the provision of a result as a consequence of said comparison, as these steps are defined in the method according to any of claims 1 to 13.

Citation Information

Patent Citations

  • Microbial marker of colorectal cancer and application of marker

    CN109943636A

  • Method for diagnosing colon tumor via bacterial metagenomic analysis

    EP3564390A1

  • Compositions and methods for diagnosing colorectal cancer

    US20230083456A1

  • Methods of determining colorectal cancer status in an individual

    WO2018109219A1

  • Compositions and methods for diagnosing colorectal cancer

    WO2023036266A1