Methods for detecting colorectal tumors

By measuring specific microorganisms and metabolites in fecal samples, the method addresses the lack of early-stage colorectal cancer detection, enabling non-invasive tumor detection and improving patient outcomes.

JP7854661B2Active Publication Date: 2026-05-07INSTITUTE OF SCIENCE TOKYO +2
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
INSTITUTE OF SCIENCE TOKYO
Filing Date
2021-06-07
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing methods fail to identify bacteria associated with early stages of colorectal cancer progression, such as multiple polyps and intramucosal carcinoma, limiting effective detection of colorectal tumors.

Method used

A method involving the measurement and comparison of specific microorganisms and metabolites in fecal samples, including Actinomyces odontolyticus, Phascolarctobacterium succinatutens, and other bacteria, along with amino acids, organic acids, and bile acids, to detect colorectal tumors through non-invasive stool examination.

Benefits of technology

Enables the detection of colorectal tumors like adenomas and intramucosal carcinomas through a non-invasive stool test, improving early detection and potentially reducing mortality rates by allowing for timely endoscopic removal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007854661000001
    Figure 0007854661000001
  • Figure 0007854661000002
    Figure 0007854661000002
  • Figure 0007854661000003
    Figure 0007854661000003
Patent Text Reader

Abstract

To identify bacteria associated with multiple polyps or an intra-mucosa cancer and establish a novel colorectal tumor detection means based on these bacteria, a method for detecting a colorectal tumor is provided, the method being characterized by including the following steps (1) to (3): (1) a step for measuring an amount of microorganisms such as Actinomyces odontolyticus in the feces of a subject; (2) a step for comparing the value measured in step (1) with a corresponding value in the feces of a normal person; and (3) a step for determining, on the basis of the result of the comparison in step (2), that the subject has colorectal tumor when the value measured in step (1) is higher or lower than the corresponding value in the feces of the normal person.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for detecting colorectal tumors such as adenomas and intramucosal cancers. If colorectal cancer can be detected at a stage before it becomes advanced cancer, the mortality rate of patients can be reduced, and since cancer can be removed by endoscopy, the quality of life of patients can also be improved.

Background Art

[0002] Colorectal cancer is the most common cancer in Japan, second only to gastric cancer. Although the Westernization of lifestyle such as diet is considered to be the cause, its mechanism is not clear.

[0003] The number of cells in one human is about 37 trillion, and the number of intestinal bacteria per person is approximately 40 trillion, with a weight of about 1 - 1.5 kg. It has recently been found that disturbances in these intestinal flora are related to various diseases such as inflammatory bowel disease. In 2012, it was reported that Fusobacterium nucleatum, which is known as a causative bacterium of periodontal disease in the oral cavity, is characteristically present in large numbers in the feces of patients with colorectal cancer, and this has been verified so far (Non-Patent Document 1, Non-Patent Document 2).

Prior Art Documents

Non-Patent Documents

[0004]

Non-Patent Document 1

Non-Patent Document 2

[0005] Colorectal cancer progresses from multiple polyps (adenomas) to intramucosal carcinoma and then to advanced cancer. While several bacteria associated with advanced colorectal cancer have been identified, no bacteria or metabolites associated with multiple polyps or intramucosal carcinoma in the stages before advanced cancer were known.

[0006] Against this background, the present invention aims to identify bacteria associated with multiple polyps and intramucosal carcinoma, and to provide a novel method for detecting colorectal tumors based on these bacteria. [Means for solving the problem]

[0007] As a result of diligent research to solve the above problems, the inventors of this invention discovered that bacteria such as Actinomyces odontolyticus are associated with the early stages of colorectal cancer development, and thus completed the present invention.

[0008] In other words, the present invention provides the following [1] to [7]. [1] A method for detecting colorectal tumors, characterized by comprising the following steps (1) to (3), (1) A step of measuring the amount of microorganisms in the feces of a subject, wherein the microorganisms are Actinomyces odontolyticus, Phascolarctobacterium succinatutens, Actinomyces viscosus, Desulfovibrio longreachensis, Solobacterium moorei, Porphyromonas uenonis, Colinsella aerofaciens, Desulfovibrio vietnamensis, and Peptostreptococcus stomatis. The process involves using at least one microorganism selected from the group consisting of Lactobacillus sanfranciscensis, Parvimonas micra, Gemella morbillorum, Selenomonas sputigena, Bifidobacterium longum subsp. longum, Eubacterium eligens, and Lachnospira multipara. (2) A step of comparing the value measured in step (1) with the corresponding value in the feces of a healthy person. (3) As a result of comparing the results of process (2), the values ​​measured in process (1) were found to be for Actinomyces odontolyticus, Phascolarctobacterium succinatutens, Actinomyces viscosus, Desulfovibrio longreachensis, Solobacterium moorei, Porphyromonas uenonis, Colinsella aerofaciens, Desulfovibrio vietnamensis, and Peptostreptococcus stomatis. If the amount is that of stomatis, Lactobacillus sanfranciscensis, Parvimonas micra, Gemella morbillorum, or Selenomonas sputigena, the subject is determined to have a colorectal tumor if the value is higher than the corresponding value in the stool of a healthy person. If the amount measured in step (1) is that of Bifidobacterium longum subsp. longum, Eubacterium eligens, or Lachnospira multipara, the subject is determined to have a colorectal tumor if the value is lower than the corresponding value in the stool of a healthy person.

[0009] [2] The method for detecting colorectal tumors according to [1], characterized in that in step (1), the microorganism is at least one microorganism selected from the group consisting of Actinomyces odontolyticus, Phascolarctobacterium succinatutens, Actinomyces viscosus, and Desulfovibrio longreachensis, and in step (3), the subject is determined to have a colorectal tumor when the value measured in step (1) is higher than the corresponding value in the feces of a healthy person.

[0010] [3] The method for detecting colorectal tumors according to [1], characterized in that in step (1), the microorganism is Actinomyces odontolyticus, and in step (3), the subject is determined to have a colorectal tumor when the value measured in step (1) is higher than the corresponding value in the feces of a healthy person.

[0011] [4] A method for detecting colorectal tumors according to any one of [1] to [3], further comprising the following steps (a-1) to (a-3), (a-1) A step of measuring the amount of amino acids in the feces of a subject, wherein the amino acid is at least one amino acid selected from the group consisting of isoleucine, leucine, valine, phenylalanine, tyrosine, glycine, and serine. (a-2) A step of comparing the value measured in step (a-1) with the corresponding value in the feces of a healthy person. (a-3) A step in which, as a result of the comparison in step (a-2), if the value measured in step (a-1) is higher than the corresponding value in the stool of a healthy person, the subject is determined to have a colorectal tumor.

[0012] [5] A method for detecting colorectal tumors according to any one of [1] to [3], further comprising the following steps (b-1) to (b-3), (b-1) A step of measuring the amount of organic acid in the feces of a subject, wherein the organic acid is at least one organic acid selected from the group consisting of succinic acid, fumaric acid, malic acid, and isovaleric acid. (b-2) A step of comparing the value measured in step (b-1) with the corresponding value in the feces of a healthy person. (b-3) A step in which, as a result of the comparison in step (b-2), if the value measured in step (b-1) is higher than the corresponding value in the stool of a healthy person, the subject is determined to have a colorectal tumor.

[0013] [6] A method for detecting colorectal tumors according to any one of [1] to [3], further comprising the following steps (c-1) to (c-3), (c-1) A step of measuring the amount of bile acid in the feces of a subject, wherein the bile acid is at least one bile acid selected from the group consisting of deoxycholic acid, glycocholic acid, and taurocholic acid. (c-2) A step of comparing the value measured in step (c-1) with the corresponding value in the feces of a healthy person. (c-3) A step in which, as a result of the comparison in step (c-2), if the value measured in step (c-1) is higher than the corresponding value in the stool of a healthy person, the subject is determined to have a colorectal tumor.

[0014] [7] A method for detecting a colorectal tumor according to any one of [1] to [6], characterized in that the colorectal tumor is an adenoma or an intramucosal carcinoma.

[0015] This specification includes the content described in the specification and / or drawings of the Japanese Patent Application, Japanese Patent Application No. 2020-098956, which forms the basis of the priority of this application. [Effects of the Invention]

[0016] This invention provides a novel method for detecting colorectal tumors. This method makes it possible to detect colorectal tumors such as adenomas and intramucosal carcinomas, which were previously only detectable by endoscopy, through a non-invasive stool examination. [Brief explanation of the drawing]

[0017] [Figure 1] Figure showing global metagenomic and metabolomic characteristics of fecal samples. [Figure 2] A diagram illustrating the different taxonomic and metabolomic signatures of cancer as it progresses through each stage. [Figure 3] Diagram showing CRC-related changes in microbial genes grouped into KO genes and KEGG pathway modules. [Figure 4] A diagram illustrating the dynamics of microorganisms and their diagnostic potential during the multi-stage progression of CRC. [Figure 5] A diagram outlining the research and the metagenomic analysis pipeline. [Figure 6] A diagram showing the clinical information of the subjects. [Figure 7] A diagram showing the microbial community structure and human genome content in fecal metagenomics. [Figure 8] A diagram showing the distribution of tumor sites and sexes within the overall structure of the metagenomic and metabolome. [Figure 9] A figure comparing the abundance of A. parvulum in metagenomic analysis and qPCR analysis. [Figure 10] A figure showing the replication rate estimated using GRiD. [Figure 11] A diagram showing changes in metabolites at different stages of colorectal cancer. [Figure 12] A diagram showing metabolome changes in the tricarboxylic acid (TCA) pathway and metagenomic changes in amino acid metabolism and other representative pathways. [Figure 13] Figure illustrating the potential of metagenomics and metabolome as diagnostic markers for early (S0) and advanced (SIII / IV) CRC. [Figure 14] A diagram showing the correlation between food intake and gut microbiota. [Figure 15]Enlarged view of the figures for Actinomyces odontoliticus, Phascolactobacterium succinatytens, Actinomyces viscodus, and Desulfovibrio longlichensis in Figure 2. The bar graphs in the figure show, from left to right, the abundance of each organism in healthy individuals, patients with multiple polyps, patients with stage 0 colorectal cancer, patients with stage I and II colorectal cancer, and patients with stage III and IV colorectal cancer. [Figure 16] This is an enlarged view of the figures for isoleucine, leucine, valine, phenylalanine, tyrosine, glycine, and serine in Figure 2. The bar graphs in the figure show, from left to right, the amounts present in healthy individuals, patients with multiple polyps, patients with stage 0 colorectal cancer, patients with stage I and II colorectal cancer, and patients with stage III and IV colorectal cancer. [Figure 17] This is an enlarged view of the figures for succinic acid, fumaric acid, malic acid, and isovaleric acid in Figure 12. The bar graphs in the figure show, from left to right, the amounts present in healthy individuals, patients with multiple polyps, patients with stage 0 colorectal cancer, patients with stage I and II colorectal cancer, and patients with stage III and IV colorectal cancer. [Figure 18] This is an enlarged view of the figures for deoxycholic acid (DCA), glycocholic acid, and taurocholic acid in Figure 2. The bar graphs in the figure show, from left to right, the levels present in healthy individuals, patients with multiple polyps, patients with stage 0 colorectal cancer, patients with stage I and II colorectal cancer, and patients with stage III and IV colorectal cancer. [Modes for carrying out the invention]

[0018] The present invention will be described in detail below. In this invention, "colorectal tumor" mainly refers to adenoma (multiple polyps) or intramucosal carcinoma (stage 0 colorectal cancer), but is not limited to these.

[0019] The present invention provides a method for detecting colorectal tumors, characterized by comprising the following steps (1) to (3). In step (1), the amount of microorganisms in the subject's feces is measured.

[0020] Examples of microorganisms include Actinomyces odontolitis, Phascolactobacterium succinatytens, Actinomyces viscosus, Desulfovibrio longlichensis (hereinafter, these four microorganisms may be referred to as "Group A microorganisms"), Salobacterium moorei, Porphyromonas uenonis, Corinzera aerofasiens, Desulfovibrio vientaensis, Peptostreptococcus stomatis, Lactobacillus sanfrancisensis, Parvimonas micra, Gemella morbilorum, Selenomonas sptigena (hereinafter, these nine microorganisms may be referred to as "Group B microorganisms"), Bifidobacterium longum subsp. longum, Eubacterium erigens, and Lachnospira multipara (hereinafter, these three microorganisms may be referred to as "Group C microorganisms").

[0021] In step (1), the amount of at least one microorganism selected from the group consisting of the above microorganisms is measured, preferably the amount of at least one microorganism selected from the group consisting of Actinomyces odontolitis, Phascolactobacterium succinatytens, Actinomyces viscosus, and Desulfovibrio longlichensis is measured, and more preferably the amount of Actinomyces odontolitis is measured.

[0022] The amount of microorganisms can be measured according to known methods, for example, by extracting total DNA from a fecal sample, reading its sequence, and basing it on that sequence and a 16S ribosomal RNA database (e.g., SILVA).

[0023] In step (2), the values ​​measured in step (1) are compared with the corresponding values ​​in the stool of healthy individuals. The corresponding values ​​in the stool of healthy individuals may be determined before measuring the index in the subject's stool, or they may be determined simultaneously with the measurement of the index in the subject's stool.

[0024] In step (3), if the value measured in step (1) is higher or lower than the corresponding value in the stool of a healthy person, based on the comparison results from step (2), the subject is determined to have a colorectal tumor.

[0025] Microorganisms in Group A specifically increase in adenomas or intramucosal carcinomas, and then show a decreasing trend thereafter (Figures 2b and 15). Therefore, when measuring the amount of Group A microorganisms, a subject is diagnosed with a colorectal tumor if the measured value is higher than the corresponding value in the stool of a healthy individual. Furthermore, measuring the amount of Group A microorganisms can specifically detect adenomas or intramucosal carcinomas among colorectal tumors.

[0026] Microorganisms in Group B increase in adenomas or intramucosal carcinomas and continue to increase thereafter (Figure 2b). Therefore, when measuring the amount of Group B microorganisms, a subject is diagnosed with a colorectal tumor if the measured value is higher than the corresponding value in the stool of a healthy individual. Furthermore, measuring the amount of Group B microorganisms can detect not only adenomas and intramucosal carcinomas but also more advanced colorectal cancers (stages I, II, III, and IV colorectal cancer).

[0027] Microorganisms in group C tend to decrease in adenomas (Figure 2b). Therefore, when measuring the amount of group C microorganisms, a subject is diagnosed with a colorectal tumor if the measured value is lower than the corresponding value in the stool of a healthy individual. Furthermore, measuring the amount of group C microorganisms can specifically detect adenomas among colorectal tumors.

[0028] The method for detecting colorectal tumors of the present invention, as described above, detects colorectal tumors by measuring the amount of microorganisms in the subject's stool. However, in order to improve the detection accuracy, the method may also be combined with measuring the amount of metabolites, such as amino acids, organic acids, and bile acids, in the subject's stool.

[0029] The detection of colorectal tumors by the amount of metabolites can be performed in the same way as the detection of colorectal tumors by the amount of microorganisms described above. The amount of metabolites can be measured according to known methods, for example, using CE-MS.

[0030] Examples of amino acids include isoleucine, leucine, valine, phenylalanine, tyrosine, glycine, and serine, and at least one of these amino acids can be selected as the target of measurement.

[0031] These amino acids are specifically increased in mucosal cancers (Figures 2c and 16), so when their measured values ​​are higher than the corresponding values ​​in the stool of healthy individuals, the subject is diagnosed with a colorectal tumor. Furthermore, by measuring the amount of these amino acids, mucosal cancers can be specifically detected among colorectal tumors.

[0032] Examples of organic acids include succinic acid, fumaric acid, malic acid, and isovaleric acid, and at least one selected from this group of organic acids can be used as the target of measurement.

[0033] Succinic acid, fumaric acid, and malic acid are specifically increased in mucosal cancer (Figures 12 and 17), and isovaleric acid is specifically increased in advanced colorectal cancer (Figures 2c and 17). Therefore, when the measured values ​​of these acids are higher than the corresponding values ​​in the stool of healthy individuals, the subject is determined to have a colorectal tumor. Furthermore, measuring the amounts of succinic acid, fumaric acid, and malic acid allows for the specific detection of mucosal cancer among colorectal tumors, and measuring the amount of isovaleric acid allows for the specific detection of advanced colorectal cancer among colorectal tumors.

[0034] Examples of bile acids include deoxycholic acid, glycocholic acid, and taurocholic acid, and at least one of these bile acids can be selected as the target of measurement.

[0035] Deoxycholic acid increases specifically in adenomas (Figure 2c, Figure 18), and glycolic acid and taurocholic acid increase specifically in intramucosal cancers (Figure 2c, Figure 18). Therefore, when the measured values are higher than the corresponding values in the feces of healthy subjects, the subject is determined to have a colorectal tumor. In addition, when measuring the amount of deoxycholic acid, adenomas can be specifically detected among colorectal tumors, and when measuring the amounts of glycolic acid and taurocholic acid, intramucosal cancers can be specifically detected among colorectal tumors.

Example

[0036] The present invention will be described in more detail below with reference to examples, but the present invention is not limited to these examples.

[0037] Colorectal cancer (CRC) affects more than 250,000 people worldwide every year. 9,10 , 6,7 , 8 Most sporadic CRCs develop through the formation of polypoid adenomas, preceded by intramucosal cancer (high-grade dysplastic adenomas), which may progress to the malignant form. 2 This process is known as the adenoma-carcinoma sequence and occurs through a multistage mechanism involving specific mutations. 2 Since it takes decades for the final malignant tumor to develop, 3 the early detection of cancer and endoscopic removal are priorities in cancer prevention.

[0038] Changes in the intestinal microbiota are thought to be involved in changes in human health and disease, including CRC. 4 Fecal metagenomics is a useful tool for quantifying the intestinal microbiome and holds promise for diagnosis. 5 In particular, F. nucleatum has been implicated in CRC, 6,7 and its mechanism is being elucidated in mice. 8 In addition, intestinal metabolites, including amino acids and bacterial metabolites (e.g., bile acids and short-chain fatty acids), are also said to be associated with the cancer state in the intestine. 9,10Therefore, comprehensive metagenomic and metabolome analysis may offer another approach to understanding the development of CRC associated with changes in the gut microbiome.

[0039] The inventors performed whole-genome shotgun metagenomics and capillary electrophoresis-time-of-flight mass spectrometry (CE-TOFMS) based metabolomics on fecal samples collected from patients with various stages of colorectal neoplasms, obtaining evidence of different stage-specific phenotypes of microorganisms and metabolites in the feces.

[0040] Metagenomic data were collected from 616 subjects and metabolome data from 406 subjects (Figure 5). Based on colonoscopy and histological findings, subjects were classified into nine groups: 1) Normal (no significant colonoscopy findings), (2) Few polyps (up to two small polyps less than 5 mm), (3) Multiple polypoid adenomas with low-grade dysplasia (MP, 3 or more, mostly 5 or more), (4) Intramucosal carcinoma (high-grade dysplasia polypoid adenoma), stage 0 / pTis CRC (S0), (5) Stage I CRC, (6) Stage II CRC, (7) Stage III CRC, and (8) Stage IV CRC based on the 8th International Union for Cancer Control (UICC) TNM classification. The remaining group was (9) Normal individuals with a history of colorectal surgery (HS). Polypoid adenomas in MP were limited to conventional colorectal adenomas with low-grade dysplasia that were pathologically proven (i.e., tubular adenoma, tubular cystoid adenoma, villous adenoma), and serrated adenomas were not included. Subjects in groups (1) and (2) were used as healthy controls. CRCs with stages I and II were grouped into stage I / II (SI / II), and CRCs with stages III and IV were grouped into stage III / IV (SIIII / IV) for analysis. The clinical characteristics of each group are shown in Figure 6. Smoking history (Brinkman index) differed between groups, with patients with more advanced stages tending to have lower Brinkman indexes (i.e., less smoking) compared to patients with earlier stages, and patients with HS showed even lower values. Metagenomic data were obtained from 28 CRC patients with stages I-III both preoperatively and postoperatively.

[0041] Figure 1 shows a summary of taxonomic data obtained from 616 subjects and metabolome data obtained from 406 subjects. Subjects rich in the genus Bacteroides were found to have fewer Prevotella species. Notably, the genus Megamonas was detected at a high frequency in 118 out of 616 patients (19.2%) across the entire group, but was rarely detected in the remaining patients, indicating an unusual distribution within the population. While Megamonas has not been reported as a dominant genus in gut microbiota studies targeting Caucasians, it has been detected in studies targeting Chinese individuals, suggesting that this genus may be characteristic of Asian populations. 11 The human genome content is calculated for each stage of the CRC (S0, P = 0.00175; SI / II, P = 8.19 × 10⁻⁶). -5 SIII / IV, P = 1.05 × 10 -5 In a one-sided Mann-Whitney U test, the levels were significantly higher than in healthy controls (Figure 7). Principal component analysis (PCA) identified the two most varied clusters, Bacteroides and Prevotella, in all subjects (Figure 1b) and 251 healthy controls (Figure 7b). These two clusters are defined as the major components of the human gut microbiota. 12 Dirichlet polynomial mixture model 13 When fitting was performed using the model, four community types were obtained for all 616 participants, one of which was comprised of Prevotella species (Figure 7e). Regarding known risk factors (obesity (P=0.8773), alcohol consumption (P=0.2989), and smoking history (P=0.3151)), no significant differences (P<0.005) were observed among the four community types. Furthermore, colonoscopy findings (healthy, MP, S0, SI / II, SIII / IV, P = 0.2864) and tumor location (left colon, right colon, rectum, P = 0.5231) were not associated with any of the community types. In terms of age, the Prevotella cluster was more frequently observed in men (79.1%) (P = 7.80 × 10⁻⁶). -10 ).

[0042] Propionic acid and butyrate are the main energy sources for the large intestine. 14 However, these two were ranked as the two most abundant metabolites. In PCA, in addition to propionic acid and butyrate, dihydrouracil and urea also showed considerable variability (Figure 1b). In PCA, stage, tumor site, and sex were not associated with variations in the metabolite profile (Figure 8).

[0043] Compared to healthy controls, microbiome shifts were observed not only in MP and S0 but also in SI / II and SIII / IV, and were found to be highly different between stages. In S0, SI / II, and SIII / IV samples, numerous species belonging to the Firmicutes, Fusobacteria, and Bacteroidetes lineages were predominantly elevated, increasing with increasing malignancy. Elevated Fusobacteria species were observed in at least two stages, while many elevated Firmicutes and Bacteroidetes species were stage-specific. In the Proteobacteria phylum, many species were elevated only in MP (Figure 2a). The genus Bifidobacterium was mainly decreased in S0. There were two patterns of significant (P < 0.005) species increases: the first was an increase from early to late stages, and the second was an increase only in the early stages. The former included F. nucleatum (e.g., F. nucleatum ssp. nucleatum (S0, P = 7.64 × 10)). -5 , q = 0.0492; SI / II, P = 1.47 × 10 -8 q = 5.63 × 10 -5 SIII / IV, P = 6.93 × 10 -11 q = 2.62 × 10 -7 )), Solobacterium moorei(S0, P = 1. 18 × 10 -5 , q = 0.0381; SI / II, P = 0.000601, q = 0.1355; SIII / IV, P = 5.91 × 10 -5, q = 0.0195), Peptostreptococcus stomatis (SI / II, P = 4.62 × 10). -6 , q = 3.22 × -3 ; SIII / IV, P = 1.98 × -10 , q = 3.75 × -7 )、Peptostreptococcus anaerobius(SI / II, P = 7.57 × 10). -5 , q = 2.90 × -3 ; SIII / IV, P = 1.83 × -10 , q = 3.75 × -7 )、Lactobacillus sanfranciscensis(S0, P = 0.000616, q = 0.118; SIII / IV, P = 0.00328, q = 0.101)、Parvimonas micra(SI / II, P = 1.89 × 10). -5 , q = 1.04 × -3 ; SIII / IV, P = 6.15 × -13 , q = 4.66 × 10-9) and Gemella morbillorum (S0, P = 0.000257, q = 0.0800; SI / II, P = 5.62 × 10). -7 , q = 6.15 × 10-4; SIII / IV, P = 1.79 × 10–9, q = 2.11 × -6 ) was significantly higher than that of Atopobium parvulum(MP, P = 0.00338, q = 0.153; S0, P = 7.03 × 10). -5The identified species were Actinomyces odontolyticus (S0, P = 0.000164, q = 0.0636), Desulfovibrio longreachensis (S0, P = 0.00164, q = 0.188), and Phascolarctobacterium succinatutens (S0, P = 0.00236, q = 0.242). The presence of A. parvulum was also confirmed by quantitative PCR (Figure 9).

[0044] Furthermore, we identified new species associated with CRC. Among these, Colinsella aerofaciens (P = 0.000840, q = 0.0544), Dorea longicatena (P = 0.000925, q = 0.0557), Porphyromonas uenonis (P = 0.000439, q = 0.0475), Selenomonas sputigena (P = 0.00369, q = 0.101), and Streptococcus anginosus (P = 0.00177, q = 0.0788) showed significant elevation in SIII / IV across all four analysis pipelines used (see "Methods"). Previous studies 5 Similarly, the butyrate-producing bacteria Lachnospira multipara (S0, P = 0.000596, q = 0.585; SI / II, P = 0.000801, q = 0.725; SIII / IV, P = 0.000116, q = 0.877) and Eubacterium eligens (S0, P = 0.00147, q = 0.698) showed a significant decrease in CRC stages. Furthermore, sulfide-producing bacteria such as Desulfovibrio vietnamensis (SIII / IV, P = 0.00109, q = 0.0565), D. longreachensis (S0, P = 0.00164, q = 0.188), and Bilophila wadsworthia (SIII / IV, P = 0.00408, q = 0.101) showed elevated levels.

[0045] Bacterial replication rate 15 is G. morbillorum(SI / II, P = 0.000436; SIII / IV, P = 5.09 × 10 -7 ), P. micra(SI / II, P = 0.000114; SIII / IV, P = 2.00 × 10 -6 ), P. stomatis(SI / II, P = 0.00424; SIII / IV, P = 8.07 × 10 -6 In each stage, the replication rates were significantly higher, with F. nucleatum ssp. nucleatum (P = 0.000386) and D. longicatena (P = 0.000246) being significantly higher in SIII / IV compared to healthy controls (Figure 10). The high replication rate can explain the high abundance of these species, suggesting that these bacteria may be metabolically activated. In particular, S. moorei (MP, P = 0.000299) and C. aerofaciens (S0, P = 0.000215) showed significantly higher replication rates in MP and S0, respectively, and their relative presence increased in subsequent stages. These results suggest that changes in many microbial species may be involved in the early progression of CRC.

[0046] A total of 65 metabolites showed significant differences (P < 0.005) compared to healthy individuals at at least one stage (Figure 11). Bile acids, short-chain fatty acids, amino acids, and key carbon metabolism components were of interest due to their close relationship with the gut microbiota. 16 (Figures 2c and 12).

[0047] Compared to healthy individuals, the concentrations of MP deoxycholate (DCA) (P = 0.000118, q = 0.0503), S0 glycocholate (P = 0.00100, q = 0.0508), and taurocholate (P = 0.00195, q = 0.0655) were found to be significantly increased. Furthermore, the concentrations of branched-chain amino acids (isoleucine (S0, P=0.00124, q=0.0508), leucine (S0, P=0.000371, q=0.0507; SIII / IV, P=0.00314, q=0.0931), valine (S0, P=0.000483, 0.0508)), as well as phenylalanine (S0, P = 0.000697, q = 0.0508), tyrosine (S0, P = 0.00136, q = 0.0508), glycine (S0, P = 0.00497, q = 0.120), and serine (SIII / IV, P = 0.00178, q = 0.0900) increased. Isoverate is a branched-chain fatty acid produced from leucine through bacterial fermentation. 17 The value gradually increased from S0 to SIII / IV (SIII / IV, P = 0.00188, q = 0.0900).

[0048] Since DCA levels were elevated in the MP stage, we searched for species that might correlate with this metabolite. B. wadsworthia was the only species significantly associated with DCA in the MP stage (Figure 2d). While the correlation coefficient was positive in other stages, it was not statistically significant. B. wadsworthia is a conjugated form of taurocorate, a precursor of DCA. 18 It is known to grow in culture media containing [the specified substance].

[0049] A. parvulum is significantly increased in MP and S0, and it has been reported that it constitutes a network hub of H2S-producing bacteria in inflammatory bowel disease patients through a high co-occurrence relationship with Streptococcus. 19We investigated the correlation between the abundance of this species and other species. The number of bacteria correlated with Atopobium increased significantly at both the genus and species levels during S0 (Figure 2e, f). There was a strong correlation between A. parvulum and A. odontolyticus, S. anginosus, S. moorei, and G. morbillorum, and the relative presence of these species increased during S0 and beyond. The increase in Atopobium even in the early stages of CRC suggests that it may have a strong influence on H2S-producing bacteria.

[0050] A total of 1,243 Kyoto Encyclopedia of Genes and Genomes (KEGG) orthology genes (KO genes) were significantly elevated (P < 0.005) in at least one stage and significantly decreased in 96 genes compared to healthy controls. Because the amino acid concentration in the feces changed significantly from S0 (Figure 2c), the gene load of microorganisms was examined to confirm the role of microorganisms in amino acid metabolism (Figure 3a).

[0051] Changes in gene quantities are shown in pathway representations (Figures 3b and 12b). Pathway modules illustrating the limits of bacterial biosynthesis and degradation are described in "Microbial Metabolism in Diverse Environments" (map01120) and literature. 20-23 I modified the KEGG pathway map I referenced and built it manually.

[0052] Among the pathways with the greatest differences in expression levels, aromatic amino acid metabolism and sulfide production pathways were found to be associated with CRC. Genes involved in the biosynthesis of phenylalanine and tyrosine were significantly elevated, and among them, pheC (P = 1.94 × 10) was particularly elevated. -5 q = 0.0297 in S0) was identified as the top score marker to distinguish S0 cases from healthy controls (Figures 4b and 13). In the catabolic pathway, toxic phenylacetate 24-26Genes involved in the degradation of phenylalanine through the production of [unclear] were elevated, mainly in MP. Genes involved in tryptophan biosynthesis were significantly decreased in SIII / IV (P < 0.005). Genotoxic hydrogen sulfide 27 Disimiratory sulfate reductase subunit A (dsrA), which is involved in the production of [the substance], was significantly elevated in SIII / IV (P = 0.00499, q = 0.0729) (Figure 3b). dsrA is a component of Desulfovibrio spp. 28 It has been found to be active in many sulfate-reducing bacteria, including Desulfovibrio piger (P = 0.0178, q = 0.226), for example, it showed substantially high values ​​in SIII / IV.

[0053] Most S0 lesions are curable with endoscopic approaches, and there are many opportunities for detection. 3 To investigate the potential of gut metagenomic and metabolome parameters as diagnostic markers, random forest and LASSO logistic regression classification were used to distinguish between S0 and SIII / IV cases and healthy controls. Four different models were constructed based on species only, KO gene only, metabolite only, or a combination of these three. A comparison of the classification capabilities of the four models showed that the combined models performed better than the individual models in both S0 and SIII / IV classifications (Figure 4b,c). The area under the ROC curve (AUC) of the random forest classifier detected S0 and SIII / IV CRC patients at 0.78 and 0.85, respectively. The results obtained using the LASSO logistic regression classifier are shown in Figure 13.

[0054] In the S0 classification, the high-ranking features were mainly KO genes, such as pheC (encoding cyclohexadienyl dehydratase). Other features included D. longreachensis, S. moorei, leucine, valine, phenylalanine, and succinic acid (Figure 4b), and these were detected to be distributed differentially (Figure 2b,c). The higher ranks of the SIII / IV classification included oral anaerobic bacteria such as P. micra, P. stomatis, F. nucleatum, and P. anaerobius, which had previously been identified as marker species for CRC. 5,29,30 (Figure 4b).

[0055] Metagenomic data were obtained from 28 CRC patients (SI / II and SIII / IV) before surgical treatment and for approximately one year after treatment. Of the 22 species described in Figure 2b, the relative abundances of P. stomatis, P. anaerobius, P. micra, P. uenonis, and D. longicatena decreased after tumor resection (Figure 4d,e). No significant differences were observed when comparing these five species with fecal samples from HS subjects (Figure 4f). The main results obtained in this study are shown in Figure 4a.

[0056] This study investigated the relationship between the gut ecosystem and multi-stage tumorigenesis, providing information on the microbial and microbial metabolite profiles in CRC. The results revealed shifts in microbial and metabolome activity in the MP and S0 stages, in addition to more advanced stages. These shifts differed significantly between stages. Two patterns of species increase were observed: one that continued from the early stages, and another that increased only in the early stages. The latter pattern was the primary focus of this study, as changes in the gut microbiota may predispose individuals to CRC development. In particular, increased abundance of F. nucleatum and S. moorei was observed in the S0 stage. These species are not only known to be associated with advanced CRC, but may also contribute to the early stages of tumorigenesis. 5,31 .

[0057] In particular, A. parvulum and A. odontolyticus showed significant increases only in MP or S0. Furthermore, network analysis revealed a strong correlation between A. parvulum and Streptococcus spp., known as H2S-producing bacteria. 32 A. parvulum has been shown to constitute a network hub of H2S-producing bacteria in patients with inflammatory bowel disease, through a high co-occurrence relationship with Streptococcus. 19 A. odontolyticus is often found in the oral cavity and digestive tract of healthy humans, and has been shown to be one of the dominant Actinomyces species that form biofilms on tooth surfaces in particular. 33 Previously, although with a small sample size, the presence of A. odontolyticus in fecal samples from adenocarcinoma patients has been reported. 34 Further research is needed to elucidate the precise mechanisms by which these bacteria contribute to tumor formation.

[0058] DCA levels were significantly elevated in MP patients. DCA is known to be associated with increased DNA damage and mutations. 35 In animal studies, administration of bile acids increased the incidence of tumors in the intestines. 36 B. wadsworthia's growth is promoted by bile. 18 However, in this study, it was the only one that showed a significant correlation with DCA (Figure 2d). Furthermore, the concentrations of conjugated bile acids (taurocholate and glycocholate) were also elevated in S0. The results of the questionnaire survey showed that in S0, B. wadsworthia was positively correlated with dietary protein (P = 0.00278) and meat (P = 0.00248) intake (Figure 14). B. wadsworthia, a close relative of the genus Desulfovibrio, is known to cause inflammation, but its association with carcinogenesis has not been extensively studied. 29 , often discussed in the context of dysbiosis 9,18These microorganisms are a normal part of the gut ecosystem, but an excess of these bacteria and their metabolites in the large intestine can cause inflammation and DNA damage.

[0059] The following limitations should be considered in this study. First, this study includes a large CRC cohort, but a validation cohort is needed in the future. Second, the total number of cells in the fecal samples was not determined in this study. Recent reports have shown that microbial burden is an important factor in the changes in the microbiome observed in several diseases. 37 Recent reports have shown that microbial burden is a key factor in the changes in the microbiome observed in several diseases. 37 Absolute amounts may be a better indicator of multi-stage CRC carcinogenesis than relative amounts. Thirdly, the inventor's method (fecal samples collected at the first bowel movement after initiation of oral administration of a bowel cleansing agent) appears suitable for collecting and cryopreserving material from subjects undergoing colonoscopy in a hospital, but there are no similar studies reported for comparison. Furthermore, the impact of bowel cleansing on metabolome data in CRC cases has not yet been investigated. Therefore, further research is needed to validate the inventor's sampling protocol.

[0060] This study was unable to clarify the possibility of microbial involvement in the pre-adenoma stage or to reveal a more detailed causal relationship between the microbiome / metabolomes and tumors. Therefore, the inventors are considering conducting longitudinal microbiome analysis by regularly collecting stool and tissue biopsies from individuals who have undergone colonoscopy. Furthermore, in order to understand the role of the microbiome in carcinogenesis of CRC, it is necessary to clarify the relationship between the gut microbiome and the molecular characteristics of tumors in individual CRC patients. In this study, cases with hereditary diseases or suspected hereditary diseases were excluded. Adenomas were limited to conventional colorectal adenomas that do not include serrated lesions. Recently, the possibility of tissue bacteria being involved in the latter has been reported. 38-40Metagenomic and metabolome data obtained from stool and tissue samples from patients with genetic disorders or suspected genetic disorders, or from patients with serrated lesions, may reveal other aspects of tumorigenesis in CRCs.

[0061] In conclusion, the inventors observed dynamic changes in microbial composition, gene load of the gut microbiota, and metabolites during the multi-stage progression of CRC. While it is unclear whether these species and metabolites directly cause tumorigenesis, structural changes in the gut microbiota may lead to changes in the oncogenic microenvironment. Furthermore, this study revealed that the progression of CRC may be influenced not only by the presence of cancer-related organisms but also by the metabolic output of the entire microbial community. The inventors believe that CRC is not only fundamentally a genetic disease but also a microbial disease.

[0062] Online content: All methods, additional references, Nature Research report summaries, source data, code and data availability statements, and relevant accession codes are available at https: / / doi.org / 10.1038 / s41591-019-0458-7.

[0063] 〔method〕 Research participants and sample collection This study involved subjects who underwent total colonoscopy at the National Cancer Center Hospital (Tokyo). The samples and clinical information used in this study were obtained with informed consent and approval from the institutional review committees of each participating institution (National Cancer Center: 2013-244, Tokyo Institute of Technology: 2014018). Stool samples were collected immediately after the first bowel movement following the initiation of oral administration of bowel cleansing agents at the hospital and frozen with dry ice. The inventors had previously compared this type of sample with a frozen stool sample (standard sample) collected one day prior to colonoscopy. 41The study demonstrated that there were no significant differences in the taxonomic profiles, the pairwise-Pearson correlation coefficients, or the taxonomic abundance of the 20 dominant genera between the two groups. Participants were given the same commercially available low-residue diet the day before their colonoscopy. Data on lifestyle, including diet, were collected from a prospective study based on the Japan Public Health Center. 42 Based on examples used in previous studies, a detailed questionnaire (475 questions, 25 pages) was used to obtain the data. Cases with hereditary diseases or suspected hereditary diseases (e.g., familial adenomatous polyposis, hereditary nonpolyposis colorectal cancer, high microsatellite instability), inflammatory bowel disease, a history of abdominal surgery, and cases for which stool samples were insufficient for data collection were excluded from this study. The intergroup distribution of body mass index (BMI), alcohol intake (grams per day), and smoking habit (Brinkman index) was analyzed using the Kruskal-Wallis rank-sum test. The intergroup distribution of sex was analyzed using Pearson's χ² test (df=5).

[0064] DNA extraction From frozen fecal samples, using the GNOME DNA Isolation Kit (MP Biomedicals), as previously described 43 DNA was extracted using the bead-beat method as described above. DNA quality was evaluated using an Agilent 4200 TapeStation (Agilent Technologies). After final precipitation, the DNA samples were resuspended in TE buffer and stored at -80°C before further analysis.

[0065] Whole-genome shotgun sequencing Sequencing libraries were generated using the Nextera XT DNA Sample Prep Kit (Illumina). Library quality was verified using an Agilent 4200 TapeStation. Whole-genome shotgun sequencing of fecal samples was performed on the HiSeq2500 platform (Illumina). All samples were paired-end sequenced with a read length of 150 bp, and the target data size was 5.0 Gb.

[0066] quality control For a total of 31,797,649,036 (average 49,375,231) paired-end reads covering 4,772,084,552,120 (average 7,410,069,180) base pairs, the following quality control was performed: Raw reads containing the letter "N" (base pair not identified) were discarded. Reads containing the DNA sequence of bacteriophage phiX were processed using Bowtie 2 (version 2.2.9). 44 Using the preset option "-fast-local", the read was identified and discarded by mapping it. (cutadapt (version 1.9.1)) 45Using the following options, reads were trimmed for adapter and primer sequences, with the following choices used: "-a CTGTCTTATACATCTCCGAGCCCACGAGAC -O 33 -q 17" for the forward primer sequence and "-a CTGTCTTATACATCTGACGCTGCCGACGA -O 32 -q 17" for the reverse primer sequence. Reads containing consecutive quality values ​​of 17 or less were tail-cut at the 3' end using the cutadapt program. Next, reads shorter than 50 base pairs were discarded. Reads with an average quality value of 25 or less were discarded. Then, using Bowtie2 (version 2.2.9), the reads were mapped to the human genome (24 gi numbers: 568336000~568336023, http: / / www.ncbi.nlm.nih.gov / nuccore / 568336023 / , GRCh38). The mapped reads were determined to be derived from the human genome and were discarded. Finally, unpaired reads were discarded. As a result, a total of 28,482,269,496 paired-end reads (average 44,227,127) with a total of 4,114,878,107,497 base pairs (average 6,389,562,279) were used for the following analysis.

[0067] Taxonomic profiling High-quality reeds are available in BLAST+ (version 2.2.30). 47 (Cutoff: E < 1 × 10) -8The reads were aligned with a pre-calculated Operational Taxonomic Unit (OTU) dataset stored in VITCOMIC246 using BLASTn (cutoff: E<1×10⁻¹). This filtered for bacterial and archaeal 16S rRNA sequences, while excluding tRNA, 23S rRNA, and internal transcriptional spacer (ITS) sequences. Several reference strategies exist for the taxonomic assignment of metagenomic data. The inventors used the All-Species Living Tree Project (LTP) database from the SILVA database for two reasons: Firstly, a single-locus database has the advantage of containing more taxa than those covering multiple loci or complete genomes. Secondly, the reference genes in this database consist only of isolated strains of specific types, allowing them to be cultured under specific conditions and relatively easily manipulated in animal experiments. The filtered reads were then processed using BLASTn (cutoff: E<1×10⁻¹). -8 (Array identity > 97%, alignment coverage > 80%, bit score > 70) 48,49 Using the SILVA database (version 123) 48 Aligned with the LTP. Only top hits were selected. As a result, a total of 8,367 species and 1,941 genera were identified. The relative abundance of a species was calculated per sample and defined as the number of reads assigned to that species divided by the total number of aligned reads in the sample. If a read was aligned to multiple taxonomic sequences in the database and their alignment scores were equal, these taxonomic sequences were given a value of 1 divided by the number of taxonomic sequences so that they could "share" the read. The relative abundance of a genus was calculated as the sum of all species belonging to that genus. Hereafter, the generated profiles will be referred to as "species profiles" and "genus profiles". To validate this species profiling pipeline, species-level profiles were obtained using three other pipelines. Using the high-quality reads described above, the mOTU profiler 50Species-level metagenomic OTUs (mOTUs) were obtained from taxonomic profiling using MetaPhlAn251 (version 2.6.0) with default parameters. MetaPhlAn251 is based on species-specific marker genes. The profiler's database is based on 10 universal single-copy marker genes. These marker genes are derived from both reference genomes and metagenomic genomes. The inventors used two databases: one containing OTUs derived from metagenomic data and one not. The generated profiles include 651 and 838 species. In addition, using high-quality reads, species-level taxonomic profiling was performed using MetaPhlAn251 (version 2.6.0) with default parameters, yielding 623 species. MetaPhlAn2 is based on species-specific marker genes.

[0068] Analysis of microbial community structure The overall community structure of 616 metagenomic data reads and 406 metabolome data reads was investigated using PCA. Furthermore, the community type of the metagenomic samples was analyzed using a Dirichlet polynomial mixture model based on the number of sequence reads. 13 The "DirichletMultinomial" package in R. 52 The following was used: Up to four clusters existed, of which the third cluster was predominantly Prevotella (Figure 7e). The other three clusters had more diverse structures and were predominantly Bacteroides. Cancer stage, tumor location, and sex were examined in different community types. The distribution of age, BMI, smoking habits (Brinkman index), and alcohol intake (grams per day) in the four community types was examined using ANOVA. In addition to sex, colonoscopy findings (healthy, MP, S0, SI / II, SIII / IV, HS), and tumor location (left colon, right colon, rectum, double / triple cancer), the distribution of the other three groups without cancer (healthy, MP, HS) was tested using Pearson's χ² test (sex, df=3, colonoscopy findings, df=15, tumor location, df=18).

[0069] Construction of a taxonomic tree GraPhlAn (version 0.9.7) 53 A taxonomic tree was constructed using [a specific method / tool]. Taxonomic hierarchical information was obtained from the SILVA database. Figure 2b uses seven levels: domain, phylum, family, genus, and species. Species are defined as having an average abundance of 10 -6 The above steps were used to filter the results so that only those with P < 0.005 (one-sided Mann-Whitney U test) at any stage are displayed.

[0070] Analysis of microbial replication rates The replication rate was calculated using the Growth Rate Index (GRiD) (version 1.2). This algorithm is based on estimating the coverage ratio of peaks (ori) and troughs (ter) of a reference bacterial genome using Tukey's biweight function M estimator. The GRiD value shows a positive correlation with the replication rate. Of the 22 species shown in Figure 2b, 20 (A. odontolyticus, Actinomyces viscosus, A. parvulum, Bifidobacterium longum ssp. longum, B. wadsworthia, C. aerofaciens, D. longicatena, E. eligens, F. nucleatum ssp. nucleatum, G. morbillorum, L. multip. morbillorum, L. multipara, L. sanfranciscensis, P. micra, P. anaerobius, P. stomatis, P. succinatutens, P. uenonis, S. sputigena, S. moorei, S. anginosus) were calculated using GRiD15, and a single reference genome was used in each calculation by setting the parameter to "single". Reference genomes were downloaded from the NCBI website, and the reference genome ID and the SILVA LTP ID were cross-referenced using the NCBI accession number. If the SILVA identifier could not be cross-referenced with the NCBI database, the "Representative Genome" was selected in the "RefSeq Category." Reference genomes for D. longreachensis and D. vietnamensis were not found in the NCBI database. Replication rates at each stage (MP, S0, SI / II, SIII / IV) were compared with healthy individuals, and statistical significance was assessed using a one-sided Mann-Whitney U test at an α level of 0.005. Based on GRiD coverage requirements, samples with coverage of 0.2 or less for each reference genome were excluded from statistical testing.For A. odontolyticus, A. viscosus, A. parvulum, L. multipara, S. anginosus, and S. sputigena, only a few samples showed high recall in this analysis.

[0071] Network analysis of genus and species Spearman's correlation coefficient was calculated using the relative abundance profiles of genus and species at each stage (MP, S0, SI / II, SIII / IV). The genus correlation network was constructed using species with a correlation coefficient of 0.4 or higher or 0.4 or lower with the genus Atopobium. The network was constructed using yEd Graph Editor (version 3.18.11) (https: / / www.yworks.com / products / yed).

[0072] Array assembly IDBA_UD (version 1.1.1) 54 Using the parameters -mink 20 -maxk 120 -step 10, high-quality reads were assembled sample by sample. A total of 81,002,850 scaffolds were generated (an average of 125,780.8 scaffolds per sample).

[0073] Functional profiling MetaGeneMark (version 3.26) 55 Using the parameter -g 11, open reading frames (ORFs) were predicted from the obtained scaffold. As a result, 156,163,520 ORFs with an amino acid length of 50 or more were identified using DIAMOND (version 0.9.10) in the KEGG GENES database (as of 2017). 56The reads were annotated (cutoff: sequence identity > 40, bit score > 70, coverage > 80), resulting in a total of 126,761,506 ORFs, or annotated "genes." The gene abundance was calculated as follows: High-quality reads were mapped to scaffolds using Bowtie 2 version 2.2.944. The read coverage of each ORF on each scaffold was evaluated and defined as the number of base pairs mapped to the corresponding scaffold region divided by the length of the ORF. If two or more ORFs matched a single gene, the abundance of each gene was calculated as the average of the score values. The total gene abundance is 7,242 KO genes, which are functional units as defined by KEGG. The profile obtained in this way will be called the "KO gene profile."

[0074] Functional characteristics of pathways We obtained amino acid-related KO genes with pathway information from the KEGG BRITE "ko00001.keg" (list of KO genes with pathway maps) in the "Amino acid metabolism" category. We also collected KEGG modules published in "Microbial metabolism in diverse environments" (map01120) to investigate known pathway modules in microorganisms. Figure 3a shows KO genes with a prevalence of 5% or more in all 576 samples, compared to healthy controls, based on Mann-Whitney U test analysis (P < 0.005) at any stage (MP, S0, SI / II, SIII / IV). The representative reaction pathways shown in Figures 3b and 12b were manually constructed by referring to literature and modifying KEGG pathway reference maps. Figures 3b and 12b show KO genes with P < 0.005 at any stage (MP, S0, SI / II, SIII / IV), omitting the remaining KO genes within the pathway. Genes and KEGG orthology genes are linked in KEGG. For each knockout gene, the abundance of microbial genes was summed so that each component of the knockout gene represents an organism. Organism names are stored in KEGG as three-letter codes.

[0075] Verification of A. parvulum by quantitative PCR. The abundance of A. parvulum was investigated using quantitative PCR (qPCR) in stool samples from 73 S0 CRC patients and 73 healthy controls. The copy number of the target region of the A. parvulum 16S rRNA gene in 1 μg of extracted DNA was estimated. The PCR products were sequenced using A. parvulum F-primer (5'-TGGATAATACCGAATACTTCGAGACT-3') and A. parvulum R-primer (5'-TGCAGGTACCGTCACTTTCG-3'), and qPCR was performed using Rotor-Gene Q (QIAGEN) with TB Green Premix Ex Taq II (Takara Bio).

[0076] Metabolome analysis Quantitative analysis of charged metabolites by CE-TOFMS has been previously described. 57 The procedure was carried out as described above. Metabolites in feces were extracted by vigorously shaking methanol containing 20 μM each of methionine sulfone and d-camphor-10-sulfonic acid as internal standards. 58 All CE-TOFMS experiments were performed using the Agilent CE system. CE-TOFMS metabolome data were obtained for 517 compounds. For analysis, concentrations below the detection limit were replaced with zero, and metabolites that were below the detection limit in all samples were excluded.

[0077] Identification of metagenomic and metabolomic markers for the detection of colorectal cancer To identify metagenomic and metabolome markers that distinguish samples from CRC patients with S0 (n = 27) and SIII / IV (n = 54) classifications and samples from healthy controls (n = 127), we constructed classification models based on species, KO gene, and metabolite profiles using two different methods: random forest and LASSO logistic regression. 59Model validation involved 10x stratified cross-validation (re-sampling the dataset partitions 10 times). Each test validated the model's accuracy using ROC and performed abundance filtering to remove features with low abundance by calculating the average relative abundance at each stage. Abundance thresholds were determined to optimize AUC. Features were then standardized (centered to mean 0 and divided by the standard deviation of each feature). Models were designed using one of the three feature types (species, KO gene, metabolite) individually or in combination of all three. For the random forest model, two steps were performed. The first step involved building models using each of the three profiles (species, KO gene, metabolite) independently. This model used all pre-filtered features to calculate "feature importance" (see explanation below) by running a random forest function with the specified parameters (500 trees, balanced class weights, maximum feature = square root of all features), followed by recursive feature elimination with parameter step=0.1 using five different random seeds. 60 The optimal number of features was determined using [method / tool ​​name]. In the second stage, a combinatorial model was constructed using the features determined for each model of species, KO gene, and metabolite. Furthermore, a recursive feature removal method was used, similar to the above. 60 Feature selection was performed using the Python package "scikit-learn". All analyses were performed using the Python package "scikit-learn". The contribution of features to the model was output as the number of feature imports. Figure 4 shows features whose importance is non-zero in at least 50% of the tests. In the LASSO logistic regression model, features were selected using L1 regularization. The contribution of features to the model was calculated using the absolute percentage of the regression coefficients. Figure 9 of the extended data shows features with non-zero coefficients in at least 50% of the tests. When the diagnostic potential of known CRC risk factors (age, sex, BMI, smoking, and alcohol consumption) was evaluated, the predictive accuracy was low (Figure 13c).

[0078] Correlation between species and metabolome For each stage (MP (n = 40), S0 (n = 27), SI / II (n = 69), SIII / IV (n = 54)), the pairwise correlation coefficients between species and metabolites were calculated using Spearman's correlation coefficient. Focusing on species-metabolite pairs that were present in higher abundance in MP or S0 samples compared to healthy control samples (P < 0.005; one-sided Mann-Whitney U test) and had a correlation coefficient of 0.6 or greater in MP or S0 (P < 0.005), 169 pairs were thus formed. Among these, when selecting pairs with a species abundance of 10 -4 or more, only one pair of B. wadsworthia and DCA remained in MP.

[0079] statistical analysis The abundance of each species, KO gene, metabolite, and bacterial replication rate was compared one-to-one with healthy controls using the one-sided Mann-Whitely U test to determine whether they were significantly increased or decreased at each stage (MP, S0, SI / II, SIII / SIV). A P < 0.005 was considered statistically significant. Also, the Benjamini-Hochberg false-discovery rate-corrected P value (q value) was estimated.

[0080] Summary of the report Further information on the study design is published in the Nature Research Reporting Summary linked to this paper.

[0081] How to obtain the data The raw sequence data reported in this paper are deposited in the DDBJ Sequence Read Archive (DRA) as DRA006684 and DRA008156.

[0082] 〔Explanation of Figures〕 Figure 1. Global metagenomic and metabolome characteristics of fecal samples. a. Relative abundance of the top 30 genera (top) and human genome fractions for 616 subjects in the healthy control group (normal and few polyps) (n = 251), MP group (n = 67), S0 group (n = 73), SI / II group (n = 111), SIII / IV group (n = 74), and HS group (n = 40). Percentage of metabolite concentrations (bottom) for 406 subjects in the healthy control group (n=149), MP group (n=45), S0 group (n=30), SI / II group (n=80), SIII / IV group (n=68), and HS group (n=34). b. PCA of genus names (n=616) (left) and PCA of metabolites (n=406) (right). PC, principal component.

[0083] Figure 2. Different taxonomic and metabolomic signatures for each stage as cancer progresses. a,b. In the four stages MP (n=67), S0 (n=73), SI / II (n=111), and SIII / IV (n=111), we evaluated whether the abundance of species was significantly increased or decreased compared to healthy controls (n=251) (P < 0.005; one-sided Mann-Whitney U test). a. Phylum distribution of the number of species increased or decreased in each of the four stages compared to healthy controls. For each phylum (Proteobacteria, Bacteroidetes, Fusobacteria, Firmicutes, Actinobacteria), we counted the number of species that significantly increased or decreased in each of the four stages compared to healthy controls. The number of species in which changes were observed at all stages (ubiquitous, green) was very small; most species showed changes specific to one stage (stage-specific, red) or changes common to another stage (common, blue). b. The phylogenetic tree shows 361 different species classified into the lineages of Proteobacteria, Bacteroidetes, Fusobacteria, Firmicutes, and Actinobacteria. Species showing a significant increase (orange) or decrease (green) (P < 0.005) are indicated within the outer circle. Species showing a particularly large increase (false detection rate adjusted P < 0.1; one-sided Mann-Whitney U test) are shown in red. The innermost circle shows the relative abundance of the species averaged across all samples. The average relative abundance is 1 × 10⁻⁶. -6Species with lower levels were excluded from the phylogenetic tree. The box plots show the relative abundances of cancer-related species (green circles), butyrate producers (yellow circles), hydrogen sulfide producers (pink circles), or species newly detected in all four metagenomic pipelines (see "Methods") that changed significantly (P < 0.005) compared to healthy controls (blue circles). The Y-axis of each box plot represents the relative abundance. Each plot shows the boxes for healthy controls (leftmost bar) and the four groups (MP, S0, SI / II, SIII / IV) from left to right. c. Metabolome analysis revealed significant increases or decreases in fecal concentrations of bile acids, branched-chain fatty acids (BCFAs), and amino acids (P < 0.005; one-sided Mann-Whitney U test). Significant increases or decreases were observed in each stage (MP (n=45), S0 (n=30), SI / II (n=80), SIII / IV (n=68)) compared to healthy controls (n=149). BCAAs are branched-chain amino acids, and AAAs are aromatic amino acids. d. Relationship between B. wadsworthia and DCA. B. wadsworthia had the highest Spearman correlation coefficient (CC = 0.63, P = 1.50 × 10) with DCA of MP. -5 (Calculated using the asymptotic t approximation method) e, Genus network analysis of A. parvulum. Genus correlation networks were constructed in the healthy control group (n=251), MP group (n=67), S0 group (n=73), SI / II group (n=111), and SIII / IV group (n=74). In total, there was a correlation of 1 × 10⁻¹⁶ between Atopodium and Atopodium at at least one stage. -4 We used 22 genera with abundances exceeding 1 × 10⁻¹⁰ and Spearman correlation coefficients exceeding 0.4. Node size is proportional to the richness of the genus. Edge width is proportional to the strength of the correlation. f, Species correlation coefficient with A. parvulum in each group. Abundance is 1 × 10⁻¹⁰. -5The above shows the species with a Spearman correlation coefficient of 0.5 or higher. Species shown in red are shown in section b. Significant changes (increases and decreases) are indicated as follows: +++, increases when P < 0.005; ++, increases when P < 0.01; +, increases when P < 0.05; ---, decreases when P < 0.005; --, decreases when P < 0.01; -, decreases when P < 0.05. The boxes represent the 25-75% lines, the black line represents the median, and the whiskers represent the maximum and minimum values ​​within 1.5 times the interquartile range.

[0084] Figure 3. Changes in CRC-related changes of KO genes and microbial genes grouped into KEGG pathway modules. a,b. For each of the four stages, MP (n=67), S0 (n=73), SI / II (n=111), and SIII / IV (n=74), the gene amount was evaluated to determine whether it was significantly increased or decreased compared to healthy controls (n=251) (P < 0.005; one-sided Mann-Whitney U test). a. The relative abundance of KO genes involved in amino acid metabolism and representative microbial metabolism (methane metabolism, sulfur metabolism, aromatic decomposition, etc.) that showed a significant difference in at least one of the four stages is shown in a heatmap. KO genes with a prevalence of 5% or more (KO genes detected in 5% or more of 576 individuals) are shown. Significant changes (increase, decrease) are shown as follows. ++ indicates an increase at P < 0.005, ++ indicates an increase at P < 0.01, + indicates an increase at P < 0.05, --- indicates a decrease at P < 0.005, -- indicates a decrease at P < 0.01, - indicates a decrease at P < 0.05. The representative KO genes shown in b and a are shown using pathway modules modified from the KEGG pathway map's "Valine, Leucine, Isoleucine Biosynthesis," "Lysine Degradation," "Tyrosine Metabolism," "Phenylalanine Metabolism," "Phenylalanine, Tyrosine, Tryptophan Biosynthesis," "Methane Metabolism," "Sulfur Metabolism," and "Benzoate Degradation." Each box within the pathway represents a KO gene, shown in red if it increased at any stage, and in blue if it decreased. The bar graph shows the average relative gene amount for each sample within the five groups (from left to right: healthy (H), MP, S0, SI / II, SIII / IV), and is color-coded according to the order of the values. Each KO gene is composed of genes from organisms represented by circles. The size and color of the circles are proportional to the relative abundance of the organism's genes. The organism's genes are grouped in a row and indicated by the organism's name. The three organisms most abundant in healthy controls are indicated by three-letter codes (e.g., pru for Prevotella ruminicola, bvu for Bacteroides vulgatus). See Figure 12 for other amino acids. The dots in each pathway represent intermediate metabolites.APS, adenosine 5'-phosphosulfate; PASP, 3'-phosphoadenosine 5'-phosphosulfate.

[0085] Figure 4. Microbial dynamics and their diagnostic potential in the multi-stage progression of CRC. a. Graphical representation of major microbial and metabolome changes in the multi-stage progression of CRC. b, c. Metagenomic and metabolome markers for detecting S0 (n=27) (left) and SIII / IV (n=54) (right) CRC patients identified from healthy controls (n=127) by a random forest classification method based on species (red), KO gene (blue), metabolite (black) individually, or a combination of the three features (green). In the S0 classification, each model used 29 species, 16 KO genes, and 24 metabolites as features. In the SIII / IV classification, each model used 55 species, 5 KO genes, and 62 metabolites as features. In the combined models, species, KO gene, and metabolite features were selected from individual models. GABA represents γ-aminobutyric acid, and KO represents KEGG orthology genes. The x-axis shows the contribution of each feature to the model in each test (see "Methods"). The boxes represent the 25-75% line, the black line represents the median, the whiskers extend to the maximum and minimum values ​​within 1.5 times the interquartile range, and the dots indicate outliers. Boxes are shown in red if they overstate S0 or SIII / IV compared to healthy controls, in light blue if they understate, and in gray if there is no significant change (P < 0.005; one-sided Mann-Whitney U test). The performance of the classifier using AUC was evaluated using 10 randomized cross-validations. d. Comparison of relative species abundance before and after surgery (n = 28). The x-axis shows the logarithmically transformed multiplier change in expression, and the y-axis shows the P-value analyzed using the one-sided Wilcoxon signed-rank test. The horizontal dashed line indicates a P-value of 0.005. The size of the circles shows the average abundance of each species in the pre- and post-operative states. Of the species shown in Figure 2, butyrate producers, H2S producers, cancer-related species reported in other cohorts, or species newly reported to be associated with increased or decreased CRC in this study are highlighted in red (increase) and blue (decrease).e. Relative abundances (P < 0.005; one-sided Wilcoxon signed-rank test) of five species that significantly decreased postoperatively compared to preoperatively in 28 patients with SI / II / III CRC: D. longicatena, P. micra, P. uenonis, P. anaerobius, and P. stomatis. The boxes represent the 25-75% line, the black line represents the median, the whiskers extend to the maximum and minimum values ​​within 1.5 times the interquartile range, and the dots indicate outliers. Increases are shown in red, and decreases in blue. f. The relative abundances of the same five species in HS samples (n = 40) for which preoperative fecal samples were not available are shown compared to healthy controls (n = 251).

[0086] Figure 5. Overview of the study and the metagenomic analysis pipeline. a. Overview of the study. Whole-genome shotgun sequencing data was collected using fecal samples from 616 subjects, and functional and taxonomic profiles were created. Metabolite profiles were created using CE-TOFMS analysis on fecal samples from 406 subjects. Samples from 347 subjects were available for both sequencing analysis and CE-TOFMS data analysis. KO represents KEGG orthology genes. b. Flowchart of the pipeline used for metagenomic analysis. The inventor's metagenomics pipeline consists of three parts: quality control, functional profiling, and taxonomic profiling. Raw reads first undergo quality control checks, then go through several analysis steps to finally generate functional and taxonomic profiles based on KEGG orthology genes.

[0087] Figure 6. Clinical information of subjects. a,b. Distribution of age, sex, BMI, Brinkman index, and alcohol consumption for 616 subjects with metagenomic data (a) and 406 subjects with metabolome data (b). The boxes represent the 25-75th percentile line, the black line represents the median, the whiskers extend to the maximum and minimum values ​​within 1.5 times the interquartile range, and the dots indicate outliers.

[0088] Figure 7. Microbial community structure and human genome content in fecal metagenomics. a. The proportion of human genome reads to the total number of raw reads changes with the progression of colorectal cancer. The proportion of human genomes (proportion of reads mapped to the human genome) in the feces of the MP group (n = 67), S0 group (n = 73), SI / II group (n = 111), and SIII / IV group (n = 74) was significantly increased compared to the healthy control group (H, n = 251) (P < 0.005; one-sided Mann-Whitney U test). The box represents the 25-75% line, the black line represents the median, the vertical line represents the maximum value within 1.5 times the interquartile range, and the dots represent outliers beyond 1.5 times the interquartile range. b,c. PCA of genera (n=251) (b) and metabolites (n=149) (c) in the healthy control group. d. Fecal metagenomics (n = 616) were optimally classified into four community types by fitting to a Dirichlet polynomial mixed model. e. Composition of the top 30 genera in each of the four community types. fi. Distribution of stage (healthy, MP, S0, SI / II, SIII / IV, HS) (f), tumor site (right colon, left colon, rectum, double or triple cancer) (g), sex (h), and age (i) in each of the four community types. The patient distribution by community type is n=191 for community type 1, n=172 for community type 2, n=129 for community type 3, and n=124 for community type 4. The box in i represents the 25-75% line, the black line represents the median, the whiskers extend to the maximum and minimum values ​​within 1.5 times the interquartile range, and the dots indicate outliers.

[0089] Figure 8. Distribution of tumor site and sex in the overall structure of the metagenomic and metabolome. a, b. PCA plots of genus profiles (n = 616) grouped by tumor site (a) and sex (b). c, d. PCA plots of metabolite profiles (n = 406) grouped by tumor site (c) and sex (d).

[0090] Figure 9. Comparison of A. parvulum abundance in metagenomic analysis and qPCR analysis. a,b. A. parvulum abundance estimated by whole-genome shotgun metagenomic sequencing data (a) and qPCR using 16S rRNA gene copy number (b) in samples of 73 S0 CRC patients (green) and 73 healthy controls (red). c. Spearman correlation coefficient of A. parvulum abundance between the two methods was calculated using asymptotic approximation. d. qPCR showed a statistically significant difference in the number of A. parvulum genes between healthy controls and S0 CRC patients (one-sided Mann-Whitney U test). The box represents the 25-75% line, the black line represents the median, the whiskers extend to the maximum and minimum values ​​within 1.5 times the interquartile range, and the dots indicate outliers.

[0091] Figure 10 shows the replication rate estimated using GRiD. The replication rates were plotted for the 20 species shown in Figure 2. The Y-axis (GRiD) is defined as the ratio of coverage at the peak (ori) and trough (ter) of the reference bacterial genome. Samples with sufficient coverage are plotted for mapping to the reference genome (coverage > 0.2). The number of samples varies depending on the dependent species and is indicated in parentheses. P-values ​​were calculated for each stage (MP, S0, SI / II, SIII / IV) using a one-sided Mann-Whitney U test and compared to healthy controls. Significant increases or decreases are indicated as follows: +++, increased at P < 0.005; ++, increased at P < 0.01; +, increased at P < 0.05; ---, decreased at P < 0.005; --, decreased at P < 0.01; -, decreased at P < 0.05. The box represents the 25-75% line, the black line represents the median, the whiskers extend to the maximum and minimum values ​​within 1.5 times the interquartile range, and the dots indicate outliers.

[0092] Figure 11. Changes in metabolites at different stages of colorectal cancer. Capillary electrophoresis time-of-flight mass spectrometry (CE-TOFMS) analysis showed that 65 metabolites at any stage (MP, n = 45; S0, n = 30; SI / II, n = 80; SIII / IV, n = 68) showed a statistically significant difference (P < 0.005; one-sided Mann-Whitney U test) compared to healthy controls (n = 149). Significant changes (increases and decreases) are indicated as follows: ++ indicates an increase at P<0.005, ++ indicates an increase at P<0.01, + indicates an increase at P<0.05, --- indicates a decrease at P<0.005, -- indicates a decrease at P<0.01, and - indicates a decrease at P<0.05. The box represents the 25-75% line, the black line represents the median, the whiskers extend to the maximum and minimum values ​​within 1.5 times the interquartile range, and the dots indicate outliers.

[0093] Figure 12. Metabolic changes in the tricarboxylic acid (TCA) pathway and metagenomic changes in amino acid metabolism and other representative pathways. a. Quantification levels of metabolites involved in the tricarboxylic acid (TCA) pathway. The levels of the three TCA metabolites succinate, fumarate, and malate were significantly higher in S0 samples (SIII / IV samples for fumarate) compared with healthy control samples (P < 0.005; one-sided Mann-Whitney U test) (++, P < 0.005; ++, P < 0.01; +, P < 0.05). The reason why succinate, fumarate, and malate accumulate in the feces of early colorectal cancer patients, despite extremely low concentrations of other TCA intermediates such as 2-oxoglutaric acid, is unknown. Some bacteria are known to synthesize ATP using the reverse reaction of succinate dehydrogenase, producing succinate as a byproduct. This is part of fumarate respiration, which uses fumarate as the electron acceptor rather than molecular oxygen. The box represents the 25-75% line, the black line represents the median, the whiskers extend to the maximum and minimum values ​​within 1.5 times the interquartile range, and the dots indicate outliers. Concentration is shown on the y axis (nmol g). -1The data is shown in ). Healthy individuals (n = 127), MP (n = 45), S0 (n = 30), SI / II (n = 80), SIII / IV (n = 68). ND represents "not detected and / or not determined". b. Pathway modules of metabolic types omitted from Figure 3b. The pathway modules are improvements on the KEGG pathway maps for "Alanine, Aspartate, and Glutamate Metabolism", "Cysteine ​​and Methionine Metabolism", "Methane Metabolism", and "Arginine and Proline Metabolism". "Leucine Degradation" was constructed based on the leucine metabolism of Clostridium difficile because a bacterial map does not exist in KEGG. For each KO gene, the bar graph shows, from left to right, the average abundance of the KO gene in each sample within five groups: healthy individuals (n=251), MP (n=67), S0 (n=73), SI / II (n=111), and SIII / IV (n=74), and is color-coded according to the order of the values. Each KO gene is composed of the genes of organisms represented by circles. The size and color of the circles are proportional to the relative abundance of the organism's genes. The organism's genes are grouped in a single column and indicated by the organism's name. The three organisms most abundant in healthy controls are indicated by three-letter codes (e.g., ova is Oscillibacter valericigenes, kpe is Klebsiella pneumoniae 342). The dots in each pathway represent intermediate metabolites. The color of the boxes for pathway components indicates whether there was a significant increase (P < 0.005; one-sided Mann-Whitney U test) in any stage (MP, S0, SI / II, SIII / IV) compared to healthy controls, with red indicating this.

[0094] Figure 13. Potential metagenomic and metabolome markers for early (S0) and advanced (SIII / IV) CRC. a. ROC curves as performance evaluations for LASSO logistic regression and random forest models, used to distinguish samples from S0 (left two panels) and SIII / IV (right two panels) CRC patients from healthy control samples. Models were designed based on species (red), KO gene (blue), metabolite (black), or combinations of these three features (green). For S0 classification, 29 species were used in species-based models, 16 in KO gene-based models, and 24 in metabolite-based models. For SIII / IV classification, 55 species were used in species-based models, 5 in KO gene-based models, and 62 in metabolite-based models. In combination models, species, KO gene, and metabolite features were selected from individual models. Classification accuracy was evaluated by AUC using 10 randomized cross-validation tests. In the LASSO logistic regression model, both individual and combined models were constructed using all features that met the abundance threshold. The discriminant features among all features are shown. b. Discriminant features identified from the LASSO logistic regression classifier and random forest classifier to distinguish cases of S0 (n = 27) and SIII / IV (n = 54) from healthy controls (n = 127). The box plot colors indicate significant increases (red) and decreases (light blue) in each group compared to the healthy control group (P < 0.005; one-sided Mann-Whitney U test). Dark gray boxes indicate features that are not statistically significant. The x-axis shows the contribution of each feature to the model in each test (see "Methods"). The boxes represent the 25-75% line, the black line the median, the whiskers extend to the maximum and minimum values ​​within 1.5× of the interquartile range, and the points indicate outliers. c. Analysis of confounding factors that may affect metagenomic and metabolome classification. We analyzed the AUC for factors such as age, sex, BMI, smoking, and alcohol exposure. Smoking and alcohol values ​​are shown as the Brinkman index and alcohol intake, respectively.Although patient gender and Brinkman index differed significantly between groups, neither the random forest model nor the logistic regression model achieved high accuracy.

[0095] Figure 14. Correlation between dietary intake and gut microbiota. For Fusobacterium spp., Akkermansia muciniphila, and sulfur-producing bacteria (B. wadsworthia, Pyramidobacter piscolens), whose relationship with dietary intake has been reported to date, Spearman's correlation coefficients were examined between dietary fiber (soluble dietary fiber, insoluble dietary fiber, total dietary fiber), dietary protein intake (protein, meat), dietary fat (lipids), dietary calcium (calcium), dairy product intake (milk), and energy intake (energy). +++ indicates correlation at P < 0.005; + indicates correlation at P < 0.05. Samples without dietary data were excluded from the calculation of correlation coefficients. Health, n = 242, MP, n = 67, S0, n = 72, SI / II, n = 109, SIII / IV, n = 71.

[0096] [References] 1. Brenner, H., Kloor, M. & Pox, CP Colorectal cancer. Lancet 383, 1490-1502 (2014). 2. Fearon, ER & Vogelstein, B. A genetic model for colorectal tumorigenesis. Cell 61, 759-767 (1990). 3. Jones, S. et al. Comparative lesion sequencing provides insights into tumor evolution. Proc. Natl Acad. Sci. USA 105, 4283-4288 (2008). 4. Ashktorab, H., Kupfer, S. S., Brim, H. & Carethers, J. M. Racial disparity in gastrointestinal cancer risk. Gastroenterology 153, 910-923 (2017). 5. Zeller, G. et al. Potential of fecal microbiota for early-stage detection of colorectal cancer. Mol. Syst. Biol. 10, 766 (2014). 6. Castellarin, M. et al. Fusobacterium nucleatum infection is prevalent in human colorectal carcinoma. Genome Res. 22, 299-306 (2012). 7. Kostic, A. D. et al. Genomic analysis identifies association of Fusobacterium with colorectal carcinoma. Genome Res. 22, 292-298 (2012). 8. Yang, Y. et al. Fusobacterium nucleatum increases proliferation of colorectal cancer cells and tumor development in mice by activating Toll-like receptor 4 signaling to nuclear factor-kB, and up-regulating expression of microRNA-21. Gastroenterology 152, 851-866 (2017). 9. Louis, P., Hold, G. L. & Flint, H. J. The gut microbiota, bacterial metabolites and colorectal cancer. Nat. Rev. Microbiol. 12, 661-672 (2014). 10. Hirayama, A. et al. Quantitative metabolome profiling of colon and stomach cancer microenvironment by capillary electrophoresis time-of-flight mass spectrometry. Cancer Res. 69, 4918-4925 (2009). 11. Liao, M. et al. Comparative analyses of fecal microbiota in Chinese isolated Yao population, minority Zhuang and rural Han by 16sRNA sequencing. Sci. Rep. 8, 1142 (2018). 12. Arumugam, M. et al. Enterotypes of the human gut microbiome. Nature 473, 174-180 (2011). 13. Ding, T. & Schloss, P. D. Dynamics and associations of microbial community types across the human body. Nature 509, 357-360 (2014). 14. Wong, J. M., de Souza, R., Kendall, C. W., Emam, A. & Jenkins, D. J. Colonic health: fermentation and short chain fatty acids. J. Clin. Gastroenterol. 40, 235-243 (2006). 15. Emiola, A. & Oh, J. High throughput in situ metagenomic measurement of bacterial replication at ultra-low sequencing coverage. Nat. Commun. 9, 4956 (2018). 16. Brestoff, J. R. & Artis, D. Commensal bacteria at the interface of host metabolism and the immune system. Nat. Immunol. 14, 676-684 (2013). 17. Zarling, E. J. & Ruchim, M. A. Protein origin of the volatile fatty acids isobutyrate and isovalerate in human stool. J. Lab. Clin. Med. 109, 566-570 (1987). 18. Devkota, S. et al. Dietary-fat-induced taurocholic acid promotes pathobiont expansion and colitis in Il10 - / - mice. Nature 487, 104-108 (2012). 19. Mottawea, W. et al. Altered intestinal microbiota-host mitochondria crosstalk in new onset Crohn’s disease. Nat. Commun. 7, 13419 (2016). 20. Xu, H. et al. Isoleucine biosynthesis in Leptospira interrogans serotype lai strain 56601 proceeds via a threonine-independent pathway. J. Bacteriol. 186, 5400-5409 (2004). 21. Bui, T. P. et al. Production of butyrate from lysine and the Amadori product fructoselysine by a human gut commensal. Nat. Commun. 6, 10062 (2015). 22. Prieto, M. A., Diaz, E. & Garcia, J. L. Molecular characterization of the 4-hydroxyphenylacetate catabolic pathway of Escherichia coli W: engineering a mobile aromatic degradative cluster. J. Bacteriol. 178, 111-120 (1996). 23. Teufel, R. et al. Bacterial phenylalanine and phenylacetate catabolic pathway revealed. Proc. Natl Acad. Sci. USA 107, 14390-14395 (2010). 24. Russell, W. R. et al. High-protein, reduced-carbohydrate weight-loss diets promote metabolite profiles likely to be detrimental to colonic health. Am. J. Clin. Nutr. 93, 1062-1072 (2011). 25. Russell, W. R. et al. Major phenylpropanoid-derived metabolites in the human gut can arise from microbial fermentation of protein. Mol. Nutr. Food Res. 57, 523-535 (2013). 26. Windey, K., De Preter, V. & Verbeke, K. Relevance of protein fermentation to gut health. Mol. Nutr. Food Res. 56, 184-196 (2012). 27. Attene-Ramos, M. S., Wagner, E. D., Plewa, M. J. & Gaskins, H. R. Evidence that hydrogen sulfide is a genotoxic agent. Mol. Cancer Res. 4, 9-14 (2006). 28. Loubinoux, J., Bisson-Boutelliez, C., Miller, N. & Le Faou, A. E. Isolation of the provisionally named Desulfovibrio fairfieldensis from human periodontal pockets. Oral Microbiol. Immunol. 17, 321-323 (2002). 29. Feng, Q. et al. Gut microbiome development along the colorectal adenoma- carcinoma sequence. Nat. Commun. 6, 6528 (2015). 30. Yu, J. et al. Metagenomic analysis of faecal microbiome as a tool towards targeted non-invasive biomarkers for colorectal cancer. Gut 66, 70-78 (2017). 31. Bullman, S. et al. Analysis of Fusobacterium persistence and antibiotic response in colorectal cancer. Science 358, 1443-1448 (2017). 32. Carbonero, F., Benefiel, A. C., Alizadeh-Ghamsari, A. H. & Gaskins, H. R. Microbial pathways in colonic sulfur metabolism and links with health and disease. Front. Physiol. 3, 448 (2012). 33. Kononen, E. & Wade, W. G. Actinomyces and related organisms in human infections. Clin. Microbiol. Rev. 28, 419-442 (2015). 34. Kasai, C. et al. Comparison of human gut microbiota in control subjects and patients with colorectal carcinoma in adenoma: terminal restriction fragment length polymorphism and next-generation sequencing analyses. Oncol. Rep. 35, 325-333 (2016). 35. Bernstein, H., Bernstein, C., Payne, C. M. & Dvorak, K. Bile acids as endogenous etiologic agents in gastrointestinal cancer. World J. Gastroenterol. 15, 3329-3340 (2009). 36. Suzuki, K. & Bruce, W. R. Increase by deoxycholic acid of the colonic nuclear damage induced by known carcinogens in C57BL / 6J mice. J. Natl Cancer Inst. 76, 1129-1132 (1986). 37. Vandeputte, D. et al. Quantitative microbiome profiling links gut community variation to microbial load. Nature 551, 507-511 (2017). 38. Tahara, T. et al. Fusobacterium in colonic flora and molecular features of colorectal carcinoma. Cancer Res. 74, 1311-1318 (2014). 39. Ito, M. et al. Association of Fusobacterium nucleatum with clinical and molecular features in colorectal serrated pathway. Int J. Cancer 137, 1258-1268 (2015). 40. Mima, K. et al. Fusobacterium nucleatum in colorectal carcinoma tissue and patient prognosis. Gut 65, 1973-1980 (2016). 41. Nishimoto, Y. et al. High stability of faecal microbiome composition in guanidine thiocyanate solution at room temperature and robustness during colonoscopy. Gut 65, 1574-1575 (2016). 42. Tsugane, S. & Sawada, N. The JPHC study: design and some findings on the typical Japanese diet. Jpn J. Clin. Oncol. 44, 777-782 (2014). 43. Furet, J. P. et al. Comparative assessment of human and farm animal faecal microbiota using real-time quantitative PCR. FEMS Microbiol. Ecol. 68, 351-362 (2009). 44. Langmead, B. & Salzberg, S. L. Fast gapped-read alignment with Bowtie 2. Nat. Methods 9, 357-359 (2012). 45. Martin, M. Cutadapt removes adapter sequences from high-throughput sequencing reads. EMBnet J. 17, 10 (2011). 46. Mori, H., Maruyama, T., Yano, M., Yamada, T. & Kurokawa, K. VITCOMIC2: visualization tool for the phylogenetic composition of microbial communities based on 16S rRNA gene amplicons and metagenomic shotgun sequencing. BMC Syst. Biol. 12, 30 (2018). 47. Camacho, C. et al. BLAST+: architecture and applications. BMC Bioinformatics 10, 421 (2009). 48. Yarza, P. et al. Update of the All-Species Living Tree Project based on 16S and 23S rRNA sequence analyses. Syst. Appl. Microbiol. 33, 291-299 (2010). 49. Yarza, P. et al. Uniting the classification of cultured and uncultured bacteria and archaea using 16S rRNA gene sequences. Nat. Rev. Microbiol. 12, 635-645 (2014). 50. Sunagawa, S. et al. Metagenomic species profiling using universal phylogenetic marker genes. Nat. Methods 10, 1196-1199 (2013). 51. Truong, D. T. et al. MetaPhlAn2 for enhanced metagenomic taxonomic profiling. Nat. Methods 12, 902-903 (2015). 52. Holmes, I., Harris, K. & Quince, C. Dirichlet multinomial mixtures: generative models for microbial metagenomics. PLoS ONE 7, e30126 (2012). 53. Asnicar, F., Weingart, G., Tickle, T. L., Huttenhower, C. & Segata, N. Compact graphical representation of phylogenetic data and metadata with GraPhlAn. PeerJ 3, e1029 (2015). 54. Peng, Y., Leung, H. C., Yiu, S. M. & Chin, F. Y. IDBA-UD: a de novo assembler for single-cell and metagenomic sequencing data with highly uneven depth. Bioinformatics 28, 1420-1428 (2012). 55. Besemer, J. & Borodovsky, M. Heuristic approach to deriving models for gene finding. Nucleic Acids Res. 27, 3911-3920 (1999). 56. Kanehisa, M. & Goto, S. KEGG: Kyoto Encyclopedia of Genes and Genomes. Nucleic Acids Res. 28, 27-30 (2000). 57. Soga, T. et al. Quantitative metabolome analysis using capillary electrophoresis mass spectrometry. J. Proteome Res. 2, 488-494 (2003). 58. Mishima, E. et al. Evaluation of the impact of gut microbiota on uremic solute accumulation by a CE-TOFMS-based metabolomics approach. Kidney Int. 92, 634-645 (2017). 59. Tibshirani, R. Regression shrinkage and selection via the LASSO. JR Stat. Soc. B 58, 267-288 (1996). 60. Pedregosa, F. et al. Scikit-learn: machine learning in Python. J. Mach. Learn. Res. 12, 2825-2830 (2011).

[0097] All publications, patents, and patent applications cited herein are incorporated herein by reference in their entirety. [Industrial applicability]

[0098] This invention is applicable to industrial fields related to the detection of colorectal tumors.

Claims

1. A method for detecting colorectal tumors, characterized by comprising the following steps (1) to (3): (1) A step of measuring the amount of microorganisms in the feces of a subject, wherein the microorganism is at least one microorganism selected from the group consisting of Actinomyces odontolyticus, Phascolarctobacterium succinatutens, Actinomyces viscosus, Desulfovibrio longreachensis, Solobacterium moorei, Gemella morbillorum, and Bifidobacterium longum subsp. longum. (2) A step of comparing the value measured in step (1) with the corresponding value in the feces of a healthy person. (3) As a result of the comparison in step (2), if the value measured in step (1) is the amount of Actinomyces odontolyticus, Phascolarctobacterium succinatutens, Actinomyces viscosus, Desulfovibrio longreachensis, Solobacterium moorei, or Gemella morbillorum, the subject is determined to have a colorectal tumor if the value is higher than the corresponding value in the stool of a healthy person; if the value measured in step (1) is the amount of Bifidobacterium longum subsp. longum, the subject is determined to have a colorectal tumor if the value is lower than the corresponding value in the stool of a healthy person.

2. The method for detecting colorectal tumors according to claim 1, characterized in that in step (1), the microorganism is at least one microorganism selected from the group consisting of Actinomyces odontolyticus, Phascolarctobacterium succinatutens, Actinomyces viscosus, and Desulfovibrio longreachensis, and in step (3), the subject is determined to have a colorectal tumor when the value measured in step (1) is higher than the corresponding value in the feces of a healthy person.

3. The method for detecting a colorectal tumor according to claim 1, characterized in that in step (1), the microorganism is Actinomyces odontolyticus, and in step (3), the subject is determined to have a colorectal tumor when the value measured in step (1) is higher than the corresponding value in the feces of a healthy person.

4. Furthermore, the method for detecting colorectal tumors according to any one of claims 1 to 3, characterized in that it further includes the following steps (a-1) to (a-3), (a-1) A step of measuring the amount of amino acids in the feces of a subject, wherein the amino acid is at least one amino acid selected from the group consisting of isoleucine, leucine, valine, phenylalanine, tyrosine, glycine, and serine. (a-2) A step of comparing the value measured in step (a-1) with the corresponding value in the feces of a healthy person. (a-3) A step in which, as a result of the comparison in step (a-2), if the value measured in step (a-1) is higher than the corresponding value in the stool of a healthy person, the subject is determined to have a colorectal tumor.

5. Furthermore, the method for detecting colorectal tumors according to any one of claims 1 to 3, characterized in that it further includes the following steps (b-1) to (b-3), (b-1) A step of measuring the amount of organic acid in the feces of a subject, wherein the organic acid is at least one organic acid selected from the group consisting of succinic acid, fumaric acid, malic acid, and isovaleric acid. (b-2) A step of comparing the value measured in step (b-1) with the corresponding value in the feces of a healthy person. (b-3) A step in which, as a result of the comparison in step (b-2), if the value measured in step (b-1) is higher than the corresponding value in the stool of a healthy person, the subject is determined to have a colorectal tumor.

6. Furthermore, the method for detecting colorectal tumors according to any one of claims 1 to 3, characterized in that it further includes the following steps (c-1) to (c-3), (c-1) A step of measuring the amount of bile acid in the feces of a subject, wherein the bile acid is at least one bile acid selected from the group consisting of deoxycholic acid, glycocholic acid, and taurocholic acid. (c-2) A step of comparing the value measured in step (c-1) with the corresponding value in the feces of a healthy person. (c-3) A step in which, as a result of the comparison in step (c-2), if the value measured in step (c-1) is higher than the corresponding value in the stool of a healthy person, the subject is determined to have a colorectal tumor.

7. A method for detecting a colorectal tumor according to any one of claims 1 to 6, characterized in that the colorectal tumor is an adenoma or an intramucosal carcinoma.

Citation Information

Patent Citations

  • Method for diagnosing adenomas and / or colorectal cancer (CRC) based on analyzing the gut microbiome

    EP2955232A1

  • Method for detecting early colon cancer

    WO2019151515A1