Discovery, functional analysis and diagnosis of colorectal adenoma and cancer biomarkers

By performing metagenomic and multi-omics analysis on fecal and other biological samples and using machine learning models to identify biomarkers in the microbial composition, a non-invasive method for diagnosing colorectal tumors is provided. This method solves the problems of low screening compliance and insufficient detection sensitivity in existing technologies, and enables early detection and efficient screening.

CN121152884APending Publication Date: 2025-12-16PRESCIENT METABIOMICS JV LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202480033383.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-03-30
Filing Date
2024-03-29
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

Current colonoscopy procedures suffer from low individual compliance due to their invasiveness and discomfort, and existing non-invasive methods have low sensitivity in detecting colonic adenomas and cancers, making it difficult to detect colorectal tumors at an early stage.

Method used

By performing metagenomic and multi-omics analyses on fecal and other biological samples, and using machine learning models to identify biomarkers in the microbial composition, a feature profile is constructed to distinguish healthy subjects from those with colorectal adenomas and cancers, providing a non-invasive diagnostic method.

Benefits of technology

It improves compliance with colorectal cancer screening, enables early detection of tumors, significantly improves the detection of precancerous lesions, reduces the frequency of colonoscopy, and improves the sensitivity and specificity of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121152884A_ABST
    Figure CN121152884A_ABST
Patent Text Reader

Abstract

The present disclosure provides, in various aspects and embodiments, methods for assessing the presence or absence of colorectal neoplasms, such as colorectal cancer (CRC), colorectal adenoma (CRA), and advanced colorectal adenoma (CRAA), in a subject by metagenomic and multiomic analysis of faeces or other biological samples. In other aspects, the disclosure provides methods for generating machine learning models or "profiles" based on metagenomic and multi-omics analysis of faeces or other biological samples to evaluate the presence or absence of colon conditions, including but not limited to CRC, CRA, and CRAA, in a subject.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] priority This application claims priority and benefit to U.S. Provisional Application No. 63 / 455,698, filed March 30, 2023, which is hereby incorporated by reference in its entirety.

[0002] sequence list This application contains a sequence list, which has been submitted in XML format via EFS-Web. An XML copy named “MBI-003PC_108458-5003_Sequence_Listing” was created on March 29, 2024, and is 73,728 bytes in size; its contents are incorporated herein by reference in their entirety. Background Technology

[0003] Colorectal cancer is one of the most prevalent cancers worldwide, with an estimated 1.8 million new cases reported in 2018, including over 700,000 cases of rectal cancer. (Bray) Bray F, Ferlay J, Soerjomataram I, et al ., Global cancer statistics 2018: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries , CA Cancer J Clin 2018; 68(6):394-424. Furthermore, colon cancer is now a leading cause of death in men under 50. Siegel RL et al. Cancer Statistics, 2024 , CA: A Cancer Journal for Clinicians (2024). Despite strong evidence that screening individuals with average CRC risk reduces mortality, individual adherence is limited due to the invasiveness, discomfort, and fear associated with colonoscopy. Lauby-Secretan Bray F, Ferlay J, Soerjomataram I, et al , The IARC Perspective on Colorectal Cancer Screening , N Engl J Med 2018; 378(18):1734-1740. This creates a significant gap in healthcare systems and CRC prevention, highlighting the need for sensitive, accurate, and non-invasive diagnostic methods for detecting colonic adenomas and cancer. This disclosure fills this gap by providing biomarkers derived from microbial species present in feces and other samples. Attached Figure Description

[0004] Figure 1 A is the principal coordinate analysis (PCoA) diagram of the microbial community spectrum derived from the analyzed samples. Figure 1 B is a PCoA diagram of the microbiome profile derived from samples of each disease category. Figure 1 C is the PCoA diagram of the microbial community profile obtained from the analyzed samples after supervised normalization. Figure 1D is a PCoA diagram of the microbiome profile derived from samples of each disease category after supervised normalization.

[0005] Figure 2 This is a cross-correlation plot. The samples from the studies listed above were used individually to train the CRC models. These models were then used to predict the samples (test set) from each study.

[0006] Figure 3 A workflow is shown that illustrates two different feature selection methods, namely Feature Importance Ranking Integration (FIRE) and Statistical Inference of Associations Between Microbial Communities and Host Phenotypes (SIAMCAT), used alone or in combination.

[0007] Figure 4 The independent feature selection method SIAMCAT is illustrated. Features are ranked according to significance scores. Box plots are shown, displaying the relative abundance of samples to visualize representation, fold change, prevalence shift, and feature contribution to the area under the curve (AUC).

[0008] Figure 5 The performance evaluation obtained using the mean AUC based on the following is shown: combinations of taxonomic and genetic features (KEGG ortholog (KO) groups, or taxonomic units associated with colorectal adenoma (CRA), advanced colorectal adenoma (CRAA), or colorectal cancer (CRC)), as well as the feature selection methods FIRE and SIAMCAT applied individually and together.

[0009] Figure 6 A Venn diagram is shown, displaying the top 800, 500, 200, 100, 50, and 20 features generated by combining FIRE and SIAMCAT for CRA, CRAA, and CRC. The number of overlapping features between disease categories and their percentage of the total are shown. The number of features corresponding to taxonomic features (T) and genetic features (K) are shown separately. The number of features analyzed is shown in descending order from left to right and top to bottom (800, 500, 200, 100, 50, and 20 features).

[0010] Figure 7 A Venn diagram is shown, illustrating the direction of feature variation. The Venn diagram in the central plot illustrates the number of overlapping features for 800 taxonomic and genetic features of CRA, CRAA, and CRC, generated by a combination of FIRE and SIAMCAT. Surrounding plots consider the direction of variation in pairs to illustrate the similarities and differences between common features across disease categories. (FC) = fold change.

[0011] Figure 8 The differences in bacterial taxonomy between different disease categories are shown. The number of class-level features compared to control samples is summarized as either overpresented (positive value) or underpresented (negative value).

[0012] Figure 9 The differences in bacterial families across different disease categories are illustrated. The number of features at the family level compared to control samples is summarized as either overpresented (positive value) or underpresented (negative value). These families distinguish CRAA from controls and other disease categories.

[0013] Figure 10 A branching diagram is shown, illustrating important taxonomic features and demonstrating the differences in CRA, CRAA, and CRC.

[0014] Figure 11 The diagram illustrates the overemphasis on virulence determinants—adhesion, biofilm formation, infiltrates, virulence factor (VF) regulators, LPS, secretion system, and virulence effectors—in CRA, CRAA, and CRC compared to healthy controls.

[0015] Figure 12 A and Figure 12 B is the Euclidean distance plot, which illustrates the distances to the centroid calculated based on relative abundance for all samples across all taxonomic ranks in each study. Figure 12 A) and the distance from the centroid after supervised normalization ( Figure 12 B).

[0016] Figures 13A to 13N are charts illustrating the taxonomic characteristics of CRC detected quantitatively by sequencing and PCR: (A) Fusobacterium nucleatum ( Fusobacterium nucleatum (B) Streptococcus salivarius ( Streptococcus salivarius (C) Micromonas ( Parvimonas micra (D) Enteric Roseola ( Roseburia intestinalis ), (E) Eubacterium truncatum ( Eubacterium ventriosum ), (F) symbiotic Clostridium ( Clostridium symbiosum (G) Measles twins ( Gemella morbillorum ), (H) Cryptosporidium effervescence ( Cloacibacillus evryensis (I) Bacteroides feces Bacteroides stercoris ), (J) Toxic butyric acid mononucleosis ( Butyricimonas virosa (K) fecal Collins bacteria ( Collinsella stercoris ), (L) Plasmodium falciparum ( Fecalibacterium prausnitzii ), (M) Enteromonas butyrate-producing bacteria ( Intestimonas butyriciproducens ), and (N) Varionella spp. ( Vaillonella parvula ).

[0017] Figures 14A to 14Q are charts illustrating the taxonomic characteristics of CRAs detected quantitatively by sequencing and PCR: (A) Bacteroides sarcopene ( Bacteroides salyersiae (B) Formic acid-producing Dorperella ( Dorea formicigenerans (C) Dual-circulation rumen cocci ( Ruminococcus bicirculans (D) Clostridium helicobacter ( Clostridium spiroforme (E) Salmonella sarcoptera ( Alistipes shahii (F) Long chain Dorsal bacteria ( Dorea longicatena (G) hematogenous twin cocci ( Gemella sanguinis (H) thermophilic streptococci ( Streptococcus thermophilus (I) Bifidobacterium animalis ( Bifidobacterium animalis ), (J) Bifidobacterium pseudosporidioides ( Bifidobacterium pseudocatenulatum (K) Escherichia coli ( Escherichia coli ), (L) Gordon's bacillus ( Gordonibacter pamelaeae ), (M) Parabacterium goeringii ( Parabacteroides goldsteinii ), (N) Bifidobacterium adolescentis ( Bifidobacterium adolescentis ), (O) Cellulolytic Bacteroides ( Bacteroides cellulosilyticus ), (P) fecal bacterium ( Bacteroides caccae ), and (Q) Bacteroides norovirus ( Bacteroides nordii ).

[0018] Figures 15A to 15E are charts illustrating the taxonomic characteristics of CRAA detected by sequencing and PCR: (A) Pasteurella multocida ( Intestinibacter bartletti (B) Bacteroides xylanolyticus Bacteroides xylanisolvens (C) Bacteroides polymorpha ( Bacteroides thetaiotaomicron (D) Prevotella flavonoidolytica ( Flavinofractor plautii ), and (E) Diplostomum meristem ( Mogibacterium diversum ). Detailed Implementation

[0019] This disclosure provides, in various aspects and embodiments, methods for evaluating the presence or absence of colorectal tumors (such as colorectal cancer (CRC), colorectal adenoma (CRA), and advanced colorectal adenoma (CRAA)) in a subject by performing metagenomic and multi-omics analysis on biological samples, such as feces, blood, serum, plasma, urine, saliva, biopsy tissue, mucosal tissue samples or swabs, intestinal lavage fluid or aspirates, and other biofluid and cellular samples containing human and microbiome DNA, RNA, proteins, and other molecules for molecular analysis (referred to herein as "biological samples"). In other aspects, this disclosure provides methods for generating machine learning models or "signatures" (biomarker profiles or patterns) based on metagenomic and multi-omics analysis of biological samples (including fecal samples) to evaluate the presence or absence of colonic conditions, such as, but not limited to, CRC, CRA, and CRAA, in a subject.

[0020] CRC is a heterogeneous disease, with most cases considered sporadic and lacking an underlying genetic component. Frank et al. , Concordant and discordant familial cancer: Familial risks, proportions and population impact , Int J Cancer 2017; 140(7):1510-1516. A variety of environmental factors, including a Western diet, obesity, smoking, alcohol consumption, and lack of exercise, are known risk factors for CRC. Of these, diet is the most important, estimated to be related to approximately 38% of new CRC cases. Additional evidence for the influence of the environment on CRC is based on the finding that CRC incidence is influenced by immigration, with the risk of developing CRC altered by the diet and lifestyle of the receiving country. It is also known that each of the aforementioned CRC risk modulators modulates the composition of the gut microbiota. Greathouse et al. , Gut microbiome meta-analysis reveals dysbiosis is independent of body mass index in predicting risk of obesity-associated CRC , BMJ Open Gastroenterol 2019; 6(1):e000247;Lee et al. , Association between Cigarette Smoking Status and Composition of Gut Microbiota: Population-Based Cross-Sectional Study , J Clin Med 2018; 7(9):282; Rodriguez-Gonzalez et al. , Microbiota and Alcohol Use Disorder: Are Psychobiotics a Novel Therapeutic Strategy? , Curr Pharm Des 2020;26(20):2426-2437;Allen et al. , Exercise Alters Gut Microbiota Composition and Function in Lean and Obese Humans , Med Sci Sports Exerc2018;50(4):747-757. This association has drawn widespread attention to the gut microbiome as a potential mediator of CRC initiation and / or progression. The large number of species and genes encoded in the gut microbiome represent a source of potential biomarkers for the diagnosis and prognosis of early-stage, precancerous adenomas, late-stage adenomas, and CRC.

[0021] In various aspects and embodiments, this disclosure enables the detection of colorectal adenomas (CRA), advanced colorectal adenomas (CRAA), and / or colorectal cancer (CRC) (and other colonic conditions) based on the composition of the subject's microbiome (e.g., as present in feces or other biological samples, such as mucosal tissue samples) and / or other multi-omics molecular analytes. In various embodiments, this disclosure provides taxonomic and genetic signatures to distinguish healthy subjects from those with early and late-stage adenomas and those with cancer. In some aspects, this disclosure provides machine learning (ML) models that avoid many pitfalls associated with metagenomic and multi-omics data (e.g., data heterogeneity, noise, overfitting, etc.), including metagenomic and multi-omics data collected using heterogeneous methods and analytical procedures.

[0022] The methods disclosed herein utilize the most informative biomarkers for each disease category. For example, in some embodiments, the methods provide features that are highly important for distinguishing CRC, CRA, or CRAA from healthy controls. While CRC features are overly reliant on taxonomic features, CRA and CRAA features are more balanced in their presentation of genetic and taxonomic characteristics. As disclosed herein, the optimal features for each disease category show little overlap, suggesting that the progression from adenoma to carcinoma reflects the unique selective environment of the microbiome and does not follow a simple linear relationship.

[0023] In one aspect, this disclosure provides a method for evaluating biological subjects for the presence of colorectal cancer. In this regard, the disclosure provides a method for screening subjects as an alternative to invasive procedures, such as colonoscopy, thereby improving screening compliance and enabling early detection of tumors. In various embodiments, the method includes quantifying genetic elements from biological samples (such as fecal samples) from the subject. Other biological samples (e.g., mucosal tissue samples, blood, saliva) that allow sampling of the microbiome, including the gut microbiome, may also be used. These genetic elements are associated with colorectal cancer (CRC), colorectal adenoma (CRA), or advanced colorectal adenoma (CRAA) and can be selected using machine learning models as described herein. Genetic elements include elements associated with microbial taxonomic classification and elements associated with the function of one or more microbial genes. In this manner, the process prepares an abundance profile of genetic elements and evaluates the abundance profile against a characteristic profile indicating the presence or absence of CRC, CRA, and / or CRAA in the subject. Thus, it is possible to identify whether a subject may have (or not have) CRC, CRA, and / or CRAA. The process can provide a binary classification (i.e., presence or absence) or a statistical output indicating the likelihood that a subject has CRC, CRA, or CRAA. In various embodiments, the method provides an improved detection of adenomas (CRA and / or CRAA) compared to known detection tests.

[0024] Among the non-invasive CRC detection tests described is the fecal immunochemical test (FIT), which has limited sensitivity of 79% for detecting CRC and poor sensitivity (approximately 25%) for detecting advanced adenomas. et al. , Accuracy of fecal immunochemical tests for colorectal cancer: systematic review and meta- , analysis 2014; 160(3):171;Hundt Ann Intern Med , et al. Comparative evaluation of , immunochemical fecal occult blood tests for colorectal adenoma detection 2009; 150(3):162-9. A multi-target stool assay quantitatively examines KRAS mutations, aberrant NDRG4 and BMP3 methylation, and β-actin and hemoglobin immunoassays. This assay outperformed FIT, detecting CRC cases with higher sensitivity (approximately 92% vs. 74%), while both assays remained poor in detecting advanced precancerous lesions (approximately 42% and 24%, respectively). Ann , Intern Med et al. , Multitarget stool DNA testing for colorectal-cancer screening N Engl J Med2014; 370(14):1287-97. These results highlight another important gap in the healthcare system: the relatively poor ability of existing non-invasive methods to detect early and late-stage adenomas. Developing a diagnostic method to fill this gap could significantly improve the detection of precancerous lesions and reduce the number of colonoscopies required for average-risk subjects.

[0025] In some embodiments, the subject has a low risk of developing CRC or colorectal polyps (such as CRA or CRAA). In such embodiments, low-risk individuals screened according to this disclosure can avoid or delay more invasive colonoscopy procedures. That is, the method can serve as an alternative screening process to colonoscopy. According to these embodiments, low-risk subjects can be screened for in the healthcare system at a lower cost and with greater efficiency, thereby identifying subjects who are more in need of colonoscopy or other treatment. A “low-risk subject” (as understood in the art) is a subject who has not previously had colorectal cancer or polyps (e.g., a subject who has previously undergone colonoscopy but no colorectal cancer or polyps, such as CRA or CRAA, were detected) and has no family history of colorectal cancer or polyps. In some embodiments, low-risk subjects do not have inflammatory bowel disease, such as Crohn's disease or ulcerative colitis. In various embodiments, low-risk subjects are at least 45 years old, or at least 50 years old, or at least 55 years old, or at least 60 years old. In some implementations, the low-risk subjects are less than 75 years of age, less than 70 years of age, or less than 65 years of age. In still other implementations, the subjects are less than 45 years of age or less than 50 years of age.

[0026] In other embodiments, subjects are at high or intermediate risk of developing CRC or colorectal polyps (as understood in the art). In such embodiments, the development of colorectal tumors in these subjects can be monitored more frequently, enabling early detection and treatment without frequent colonoscopies. For example, high- or intermediate-risk subjects include those who have previously had CRC or colorectal polyps (e.g., CRA or CRAA) and / or have a family history of CRC or colorectal polyps. In some embodiments, high- or intermediate-risk subjects have inflammatory bowel disease, such as Crohn's disease or ulcerative colitis. In various embodiments, the method is performed at a defined frequency, such as at least about once a year or at least about once every other year. In some embodiments, the method is performed at least twice a year. In various embodiments, high- or intermediate-risk subjects are at least 45 years old, or at least 50 years old, or at least 55 years old, or at least 60 years old. In some embodiments, high- or intermediate-risk subjects are at least 65 years old, or at least 70 years old, or at least 75 years old. In other implementation schemes, the subjects are less than 45 years of age or less than 50 years of age.

[0027] In various embodiments, genetic elements from biological samples (such as, but not limited to, fecal samples) are quantified by nucleic acid sequencing, which may include genome sequencing and / or RNA sequencing (e.g., cDNA sequencing). In various embodiments, nucleic acid sequencing includes shotgun metagenomic sequencing, targeted amplicon sequencing, and / or hybridization capture probe sequencing, as well as any other sequencing techniques.

[0028] Several studies have examined the gut microbiota using 16S rDNA sequencing, shotgun metagenomic sequencing, targeted amplicon sequencing, or hybridization capture probe sequencing. These studies explored fecal and mucosal-associated communities as well as different stages of adenoma and cancer progression. Meta-analyses of fecal or other microbiota sample datasets identified seven bacteria enriched in CRC: Bacteroides fragilis. (Bacteroides fragilis), Fusobacterium nucleatum (Fusobacterium nucleatum), Micromonas (Parvimonas micra), Porphyromonas unsaturates (Porphyromonas asaccharolytica), Prevotella intermedia (Prevotella intermedia), Fenrir albopictus (Alistipes finegoldii) Japanese amino acid thermophilic anaerobic vibrio (Thermoanaerovibrio acidaminovorans) Dai, et al. , Multi-cohort analysis of colorectal cancer metagenome identified altered bacteria across populations and universal bacterial markers , Microbiome 2018; 6(1):70. A pair of separate meta-analyses identified an extended set of twenty-nine species enriched in eight different geographic regions. Thomas et al. , Metagenomic analysis of colorectal cancer datasets identifies cross-cohortmicrobial diagnostic signatures and a link with choline degradation , Nat Med 2019; 25(4):667-678;Wirbel et al. , Meta-analysis of fecal metagenomes reveals global microbial signatures that are specific for colorectal cancer , Nat Med 2019; 25(4):679-689. Numerous studies have analyzed the human gut microbiota associated with colon cancer and normal adjacent tissues, thus identifying a spectrum of dysbiosis associated with CRC. While specific taxa vary from study to study, some common themes include the frequent identification of relatively increased abundance of *Escherichia coli* (E. coli). E. coli ), Fusobacterium nucleatum and enterotoxins Bacteroides fragilis (ETBF) strain (Arthur et al. , Intestinal inflammation targets cancer-inducing activity of the microbiota , Science 2012; 338(6103):120-3; Kostic et al. , Fusobacterium nucleatum potentiates intestinal tumorigenesis and modulates the tumor-immune microenvironment , Cell Host Microbe 2013; 14(2):207-15; Wu et al. , A human colonic commensal promotes colon tumorigenesis via activation of T helper type 17 T cell responses , Nat Med 2009; 15(9):1016-22. Additional taxonomic units related to CRC have also been identified, but low consistency has been observed across different studies.

[0029] A key observation identified in these studies is that the relative abundance variation from tissue samples is much greater than that from fecal samples, and the highly differentiated taxa in fecal samples are more nuanced, often requiring AI methods for deciphering. Many characteristics of feces (and other biological samples containing the microbiome) present unique challenges for identifying diagnostic biomarkers for the early detection of adenomas and cancers, including high dimensionality, data sparsity, and generalization. For example, despite the large amount of DNA sequence data generated in shotgun metagenomic sequencing of fecal samples, the best-performing biomarkers obtained from such analyses are detected only in a relatively small number of samples. This defines the data sparsity problem and prescribes that high-performance diagnostics based on next-generation sequencing (NGS) sequence data requires multiple independent biomarkers to compensate for the low detection rate of any single biomarker in the human population.

[0030] Therefore, in various embodiments, metagenomic sequencing is the deep sequencing of genomic DNA isolated from fecal samples or other biological samples. In various embodiments, nucleic acid sequencing involves sequencing at least about 20,000,000 reads (i.e., raw reads for each fecal sample). In various embodiments, nucleic acid sequencing involves sequencing at least about 25,000,000 reads, or at least about 30,000,000 reads, or at least about 40,000,000 reads, or at least about 50,000,000 reads, or at least about 60,000,000 reads, or at least about 75,000,000 reads, or at least about 100,000,000 reads for each sample. Well-known quality control metrics can be used to remove low-quality reads, which are typically less than about 15% or less than 10% of the original reads. Generally, the reads will have less than about 5%, or less than about 4%, or less than about 3%, or less than about 2% of human reads. In various implementation schemes, human readings are removed from the analysis.

[0031] In various embodiments, nucleic acid sequencing includes one or more of shotgun metagenomic sequencing, rDNA sequencing, and targeted nucleic acid sequencing (e.g., targeted amplicon sequencing or hybridization capture probe sequencing). In embodiments, nucleic acid sequencing includes multiple workflows, which may include, for example, rDNA sequencing, and one or more of shotgun sequencing and targeted nucleic acid sequencing. Sequencing can be performed using any known library preparation protocol, including multiplex detection workflows using sample tags. See, for example, U.S. Patent Nos. 8,603,749 and 9,453,262, which are hereby incorporated herein by reference in their entirety.

[0032] In some embodiments, the preparation of sequencing libraries from DNA samples utilizes total DNA isolated from feces or other biological samples (e.g., GI mucosal samples). Many kits for preparing sequencing libraries from DNA are commercially available. In some embodiments, library preparation includes: DNA fragmentation, end repair, addition of sequencing adapters (e.g., by ligation or amplification), and amplification to enrich products with adapters at both ends. For example, DNA may be fragmented such that the average fragment size is in the range of 100 to 5000 base pairs, such as in the range of about 250 bp to 4000 bp, or in the range of about 500 bp to 3000 bp, or in the range of about 1000 bp to 3000 bp. In some embodiments, the average fragment size is less than 1000 bp, such as in the range of 200 to 1000 bp (e.g., 200 to 500 bp). To facilitate multiplexing, different barcode adapters can be used for different biological samples (e.g., biological samples from different subjects). In some implementations, barcoding can be introduced into the PCR amplification step by using different barcoded PCR primers to amplify different biological samples. In some implementations, the library can be subjected to shotgun metagenomic sequencing.

[0033] In some embodiments, nucleic acid sequencing focuses on one or more genomic loci to allow for taxonomic analysis, including rDNA analysis. rDNA (gene encoding rRNA) analysis may include 16S rDNA, 18S rDNA, and internal transcribed spacer (ITS) sequencing. 16S and ITS sequence analysis allows for taxonomic analysis of bacteria and archaea, while 18S and ITS sequence analysis allows for taxonomic analysis of eukaryotes (e.g., fungi). The 16S rRNA gene contains nine variable regions scattered throughout the highly conserved 16S sequence. In some embodiments, sequencing is performed by targeting subregions of the gene amplified by PCR, ranging from a single variable region, such as V4 or V6, to three variable regions, such as V1 to V3 or V3 to V5. Similarly, the 18S rRNA gene contains variable regions (V1 to V9), which can be used for differentiation at the family, order, genus, and species (and subspecies) level, as is known in the art. In some embodiments, sequencing is performed by targeting subregions of the gene amplified by PCR, ranging from a single variable region to multiple variable regions. ITS regions are located between large and small rRNA subunit gene loci and can be species-specific. This polymorphism is due to the presence of tRNA genes. ITS regions can be sequenced and taxonomically analyzed by targeted PCR amplification.

[0034] In some implementations, 16S / 18S / ITS sequences are clustered based on similarity to generate operational taxonomic units (OTUs). Representative OTU sequences can be compared with a reference database to determine taxonomy. In some implementations, sequences with >95% identity are considered to represent the same genus, while sequences with >97% identity are considered to represent the same species. Methods for determining OTUs are known in the art. In some implementations, strains or subspecies are further distinguished based on polymorphism analysis. Taxonomic analysis of 16S, 18S, and ITS DNA sequences is well known in the art. See Ze-Gang Wei et al. Comparison of Methods for Picking the Operational Taxonomic Units From Amplicon Sequences , Front. Microbiol March 24, 2021.

[0035] It can also analyze sequence reads other than rDNA, and infer possible taxonomy by comparing them with a reference microbial genome. (Helene LCF, et al.) New Insights into the Taxonomy of Bacteria in the Genomic Era and a Case Study with Rhizobia, International Journal of Microbiology Volume 2022.

[0036] In various implementations, sequence reads are analyzed to determine the abundance of gene function. Sequence reads can be analyzed against the KEGGorthology database or a similar database. The KEGG Ortholog (KO) database is a molecular function database represented by functional orthologs. Functional orthologs are manually defined in the context of a KEGG molecular network (i.e., KEGG pathway maps, BRITE hierarchies, and KEGG modules). Each node in the network (such as a box in a KEGG pathway map) is assigned a KO identifier (called a K number) as a functional ortholog defined from genes and proteins experimentally characterized in a particular organism. These identifiers are then used to assign orthologous genes in other organisms based on sequence similarity. The resulting KO groupings may correspond to a highly similar set of sequences within a limited group of organisms, or they may be a more diverse group. Thus, in various implementations, sequence reads are assigned to gene functions (such as according to the KO database), and the abundance of gene function is determined for a biological sample.

[0037] In various implementations, target genomic fragments are captured from metagenomic libraries and subsequently amplified optionally. For example, nucleic acid capture probes that hybridize to conserved regions of rDNA or conserved regions of functional orthologs can be used. Sequence capture allows for targeted enrichment of informative DNA. Combined with NGS, capture provides an efficient strategy for high-throughput screening of regions of interest. In various implementations, capture strategies reduce the required sequencing depth to less than about 25,000,000 reads, or less than about 20,000,000 reads, or less than about 15,000,000 reads, or less than about 10,000,000 reads, or less than about 5,000,000 reads, or less than about 2,000,000 reads. An exemplary sequence capture protocol includes: fragmenting the input DNA (e.g., by cleavage or using an enzyme); adding sequencing adapters (e.g., by ligation or amplification using fusion primers) to form a library molecule; and incubating the library with a pool of captureable oligonucleotide probes designed to target (and hybridize) specific regions of interest within the DNA fragment library. An exemplary captureable component is biotin, which can be conjugated to probe oligonucleotides. The probe / target hybrid is then captured from the library (e.g., using streptavidin-coated magnetic beads). The result is a sequencing-ready library highly enriched with the target DNA.

[0038] In other embodiments, genetic elements can be quantified using PCR (qPCR) according to known procedures. For example, genus-specific or species-specific sequences, as well as conserved sequences in gene functional elements, can be quantitatively amplified and detected (e.g., from rDNA in a sample). In this way, abundance profiles of genetic elements (e.g., informative traits) can be constructed without a sequencing workflow.

[0039] For the detection of CRA, CRAA, and / or CRC, the quantified number of genetic elements will be sufficient to provide a high-performance test (e.g., by allowing the analysis of a large number of informative features). For example, in various embodiments, genetic elements can be analyzed for the presence of at least about 50 features, or at least about 100 features, at least about 200 features, at least about 500 features, or at least about 800 features, or at least about 1000 features (for each model). Tables 3, 4, and 5 illustrate exemplary features for the detection of CRA, CRAA, and CRC, respectively. The number of features per test for each model need not be the same. For example, in some embodiments, a model or “feature spectrum” for the detection of CRC may include at least about 500 features, or at least about 750 features, or at least about 1000 features. In some embodiments, the model or feature spectrum for the detection of CRAA and CRA is significantly less and may include fewer than about 500 features, such as fewer than about 250 features (e.g., in the range of 50 to 200 features). However, models with more or fewer features can be constructed according to this disclosure.

[0040] In various embodiments, the analyzed genetic element is associated with colorectal adenoma (CRA). For example, the genetic element may contain one or more taxonomic or gene functional features listed in Table 3. In various embodiments, the genetic element contains at least five taxonomic or gene functional features listed in Table 3. In some embodiments, the genetic element contains at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or gene functional features listed in Table 3. In various embodiments, the genetic element contains at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 taxonomic features listed in Table 3; and at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 gene functional features listed in Table 3. As shown in Table 3, some genetic elements have differential abundance (compared to controls) in samples from CRA subjects (e.g., fecal samples or other biological samples), while other genetic elements have differential detection rates (compared to control subjects) in fecal or other biological samples from CRA subjects. In some embodiments, the genetic elements include multiple genetic elements with differential abundance in the CRA and multiple genetic elements with differential detection rates in the CRA. In some embodiments, the relative abundance difference between disease and non-disease samples (or vice versa) is at least about 1.1-fold, or at least about 1.2-fold, or at least about 1.3-fold, or at least about 1.4-fold, or at least about 1.5-fold, or at least about 2-fold. In some embodiments, the detection rate difference between disease and non-disease samples (or vice versa) is at least about 1.1-fold, or at least about 1.2-fold, or at least about 1.3-fold, or at least about 1.4-fold, or at least about 1.5-fold, or at least about 2-fold.

[0041] In various embodiments, the genetic element is associated with advanced colorectal adenoma (CRAA). For example, the genetic element may include one or more taxonomic or gene functional features listed in Table 4. In various embodiments, the genetic element includes at least five taxonomic or gene functional features listed in Table 4. In some embodiments, the genetic element includes at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or gene functional features listed in Table 4. In various embodiments, the genetic element includes at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 taxonomic features listed in Table 4; and at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 gene functional features listed in Table 4. As shown in Table 4, some genetic elements have different abundances (compared to controls) in the feces or other biological samples of CRAA subjects, while other genetic elements have different detection rates (compared to control subjects) in the feces or other biological samples of CRAA subjects. In some embodiments, the genetic elements include multiple genetic elements with differential abundance in the CRAA and multiple genetic elements with differential detection rates in the CRAA. In some embodiments, the relative abundance difference between disease and non-disease biological samples (or vice versa) is at least about 1.1-fold, or at least about 1.2-fold, or at least about 1.3-fold, or at least about 1.4-fold, or at least about 1.5-fold, or at least about 2-fold. In some embodiments, the detection rate difference between disease and non-disease biological samples (or vice versa) is at least about 1.1-fold, or at least about 1.2-fold, or at least about 1.3-fold, or at least about 1.4-fold, or at least about 1.5-fold, or at least about 2-fold.

[0042] In various embodiments, the genetic element is associated with colorectal cancer (CRC). For example, the genetic element may include one or more taxonomic or gene functional features listed in Table 5. In various embodiments, the genetic element includes at least five taxonomic or gene functional features listed in Table 5. In some embodiments, the genetic element includes at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or gene functional features listed in Table 5. In various embodiments, the genetic element includes at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 taxonomic features listed in Table 5; and at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 gene functional features listed in Table 5. As shown in Table 5, some genetic elements have different abundances (compared to controls) in the feces or other biological samples of CRC subjects, while other genetic elements have different detection rates (compared to control subjects) in the feces or other biological samples of CRC subjects. In some embodiments, the genetic elements include multiple genetic elements with differential abundance in CRC and multiple genetic elements with differential detection rates in CRC. In some embodiments, at least five, at least ten, or at least 20 genetic elements used for detecting CRC correspond to bacterial species commonly found in the oral cavity. In some embodiments, the relative abundance difference between diseased and non-disease-related samples (or vice versa) is at least about 1.1-fold, or at least about 1.2-fold, or at least about 1.3-fold, or at least about 1.4-fold, or at least about 1.5-fold, or at least about 2-fold. In some embodiments, the detection rate difference between diseased and non-disease-related samples (or vice versa) is at least about 1.1-fold, or at least about 1.2-fold, or at least about 1.3-fold, or at least about 1.4-fold, or at least about 1.5-fold, or at least about 2-fold.

[0043] In various implementations, the abundance profiles of genetic elements are evaluated against the characteristic profiles indicating the presence or absence of each of the CRC, CRA, and CRAA. As disclosed herein, the microbiome profiles of the CRC, CRA, and CRAA do not exhibit linear relationships with each other, and therefore each microbiome profile can be optimally evaluated using a separate model or characteristic profile.

[0044] In various implementations, a machine learning (ML) model is used to generate feature profiles from a training set. For example, a feature profile indicating the presence or absence of CRA is trained using fecal or other biological samples from a CRA cohort and biological samples from a control cohort. A feature profile indicating the presence or absence of CRAA is trained using fecal or other biological samples from a CRAA cohort and biological samples from a control cohort. A feature profile indicating the presence or absence of CRC is trained using fecal or other biological samples from a CRC cohort and biological samples from a control cohort. In each case, the control cohort is considered a healthy cohort, defined as the absence of CRC, CRA, and CRAA. In some implementations, the samples in the control cohort are not from subjects with severe gastrointestinal diseases such as Crohn's disease or ulcerative colitis.

[0045] According to various aspects and embodiments of this disclosure, the training set comprises at least about 50 samples, or at least about 100 samples, or at least about 150 samples, or at least about 200 samples, or at least about 500 samples, or at least about 1000 samples that are positive for CRA, CRAA, or CRC. In some embodiments, the training set comprises at least about 25 non-disease or healthy controls, or at least about 50 non-disease or healthy controls, or at least about 100 non-disease or healthy controls, or at least about 500 non-disease or healthy controls, or at least about 1000 non-disease or healthy controls. Those skilled in the art will be able to assemble training sets representing disease and control samples in a manner that produces sufficient statistical power. The training set need not originate from a single study or geographic region. In some embodiments, biological samples originate from different geographic locations (e.g., at least two different countries or continents) and / or are processed in different geographic locations. In these embodiments, separate procurement, processing, or sequencing provides additional diversity to the study protocol and may also provide information on genetic, ethnic, and / or environmental variations (including dietary variations) of the subjects.

[0046] One or more machine learning algorithms can be used to train the feature spectrum. In some embodiments, at least one of the machine learning algorithms used is a supervised machine learning algorithm. In these or other embodiments, the machine learning algorithm includes one or more of unsupervised or semi-supervised machine learning. Various machine learning algorithms are known and can be used according to this disclosure, including, but not limited to, one or more of parametric / nonparametric distance measures, logistic regression, support vector machines, decision trees, random forests, neural networks, probabilistic regression, Fisher's linear discriminant analysis, Naive Bayes classifiers, perceptrons, quadratic classifiers, kernel estimation, k-nearest neighbors, learned vector quantization, and principal component analysis. In the implementation, machine learning employs AI-enabled massively parallel computing and automated machine learning platforms and workflows (computer programs) for comparing machine learning modeling, optimization, testing, evaluation, and model ranking, such as one or more of deep learning, gradient boosting, neural networks, ensemble or hybrid modeling algorithms, such as, but not limited to, gradient boosting tree classifiers, extreme gradient boosting tree classifiers, light gradient boosting tree classifiers, elastic network predictive light gradient boosting, Keras Slim residual neural network classifiers, generalized additive models, elastic network classifiers, random forest classifiers, deep forest classifiers, average mixer classifiers, TensorFlow multilayer perceptron classifiers, TensorFlow neural network classifiers, and rule fitting classifiers.

[0047] Features can be selected using any known process. In some implementations, the feature spectrum includes features selected from the training cohort through an ensemble ranking of feature importance, ranking features according to their importance in the predictive model. Alternatively or additionally, features are selected based on their statistical significance (individually) to predict CRA, CRAA, or CRC. For example, individual features may be selected whose abundance or detection rate predicts the presence or absence of CRC, CRA, or CRAA, where the p-value in the training group is less than or equal to 0.05, or less than or equal to 0.01, or less than or equal to 0.005, or less than or equal to 0.001 (or other selected statistical thresholds). In some implementations, the feature spectrum includes features selected from the training cohort through a feature importance ranking ensemble (FIRE) and statistical inference of associations between microbial communities and phenotypes (SIAMCAT). For example, features overlapping between the two processes may be selected. Figure 3 The iterative process of feature selection is illustrated in the figure.

[0048] The accuracy of a model can be evaluated using standard Receiver Operating Characteristic (ROC) curve analysis to calculate true positive, false positive, true negative, and false negative rates, overall accuracy, and area under the curve. The term "ROC" or "ROC curve" refers to a Receiver Operating Characteristic curve. An ROC curve can be a graphical representation of the performance of a binary classifier system. For any given method, an ROC curve can be generated by plotting sensitivity versus specificity at various threshold settings. Furthermore, an ROC curve can determine the value or expected value of any unknown parameter as long as at least one of three parameters (e.g., sensitivity, specificity, and threshold setting) is provided. Unknown parameters can be determined using a curve fitted to an ROC curve. For example, providing the presence / absence or abundance of one or more features can determine the expected sensitivity and / or specificity of the test. The term "AUC" or "ROC-AUC" can refer to the area under the Receiver Operating Characteristic curve. This metric provides a measure of the diagnostic utility of a method while taking into account both the method's sensitivity and specificity. ROC-AUC can range from 0.5 to 1.0, where a value close to 0.5 may indicate limited diagnostic utility of the method (e.g., low sensitivity and / or specificity), while a value close to 1.0 indicates greater diagnostic utility of the method (e.g., high sensitivity and / or specificity).

[0049] In various embodiments, the characteristic spectrum has a sensitivity for classifying a sample based on the presence or absence of CRA, CRAA, or CRC, said sensitivity being at least about 0.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about 0.95. In embodiments, the characteristic spectrum has a specificity for classifying a sample based on the presence or absence of CRA, CRAA, or CRC, said specificity being at least about 0.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about 0.95. For example, the characteristic spectrum can provide a sensitivity for classifying each of CRA, CRAA, and CRC, with a sensitivity of at least 0.75, and a sensitivity of at least 0.75. For example, the characteristic spectrum may have an area under the curve (AUC) for classifying the sample based on the presence or absence of CRA, CRAA, or CRC, wherein the area under the curve is at least about 0.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about 0.95.

[0050] In various implementation schemes, if the process described herein does not identify a subject as potentially having CRA, CRAA, or CRC, no further procedures are performed. That is, the subject is not scheduled for colonoscopy or other evaluations for colorectal cancer or adenomas. When the process described herein identifies a subject as potentially having one or more of CRA, CRAA, or CRC, a customized diagnostic or treatment plan is initiated. For example, the subject may undergo procedures involving colonic imaging, such as colonoscopy or CT colonography (or other scanning or imaging techniques) to confirm the results, which may also involve the removal of one or more polyps and / or obtaining biopsies of growths suspected of being related to colorectal cancer. When colorectal cancer is confirmed, the subject receives treatment for CRC. For example, the subject may receive one or more surgical procedures for colorectal cancer (e.g., cancer resection, including partial colectomy in some implementations), chemotherapy, radiation therapy, and immunotherapy. Exemplary chemotherapy or immunotherapy for colorectal cancer may include one or more of the following: 5-fluorouracil (5-FU), capecitabine (XELODA) (which is metabolized by the tumor to 5-FU), irinotecan, leucovorin, oxaliplatin, cetuximab, panitumumab, regorafenib, bevacizumab, aflibercept, and ramucirumab. Exemplary combination therapies also include FOLFOX (5-FU, leucovorin, and oxaliplatin), FOLFIRI (leucovorin, 5-FU, and irinotecan), CAPEOX (capecitabine and oxaliplatin), FOLFOXIRI (leucovorin, 5-FU, oxaliplatin, and irinotecan), 5-FU with leucovorin or capecitabine alone, and the combination of trifluuridine and tipiracil (LONSURF). In some implementations, the subject receives an immune checkpoint inhibitor, such as an antibody or other molecule that inhibits PD-1, PD-L1, PD-L2, or cytotoxic T-lymphocyte-associated protein 4 (CTLA-4).

[0051] In addition, radiotherapy can be used in combination with resection, chemotherapy, and immunotherapy, or it can be used alone. Types of radiotherapy include external beam radiotherapy (EBRT), internal brachytherapy, intracavitary radiotherapy, interstitial brachytherapy, and radioembolization.

[0052] In other respects, this disclosure provides a method for preparing a genetic signature profile of genetic elements (i.e., informative features) indicating the presence of colorectal tumors. The method includes providing a training cohort of fecal or other samples (or providing RNA or DNA isolated from) subjects confirmed to have CRA, CRAA, or CRC, and performing genomic nucleic acid sequencing on the DNA isolated from the fecal or other biological samples, as previously described. A genetic signature profile is then trained to classify the samples for the presence or absence of CRA, CRAA, or CRC. In this respect, the genetic signature profile includes microbial taxonomic features and microbial gene functional features as described above and exemplified in Tables 3 to 5. In this respect, the method can employ any sample suitable for evaluating the microbiome of the cohort, including fecal samples and other biological samples, including human biological fluids (e.g., blood, serum, plasma, urine, saliva), tissues, mucous membranes, and cell samples.

[0053] As described above, nucleic acid sequencing may include one or more of shotgun metagenomic sequencing, rDNA sequencing, and targeted nucleic acid sequencing, including targeted amplicon sequencing and hybridization capture probe sequencing. Any sequencing technology may be used. Genetic elements in the sample are assigned to a reference genome for taxonomic classification (which may include rDNA analysis), and / or genetic elements are assigned to gene functions (as previously described). Taxonomic and gene functional characterization can also be analyzed at the protein level using known methods.

[0054] For example, microbial taxonomic features and microbial gene function features that exhibit differential abundance or differential detection rate in fecal or other biological samples of CRA subjects compared to control subjects are selected. In some embodiments, the features include at least five taxonomic and / or gene function features, which are optionally listed in Table 3. The features may include at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or gene function features, which are optionally listed in Table 3. In some embodiments, the features include at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 taxonomic features, which are optionally listed in Table 3; and at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 gene function features, which are optionally listed in Table 3.

[0055] In some embodiments, microbial taxonomic features and microbial gene functional features that exhibit differential abundance or differential detection rate in fecal or other biological samples of CRAA subjects compared to control subjects are selected. In some embodiments, the features include at least five taxonomic and / or gene functional features, which are optionally listed in Table 4. The features may include at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or gene functional features, which are optionally listed in Table 4. In some embodiments, the features include at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 taxonomic features optionally listed in Table 4; and at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 gene functional features optionally listed in Table 4.

[0056] In some embodiments, microbial taxonomic features and microbial gene functional features that exhibit differential abundance or differential detection rate in the feces or other biological samples of CRC subjects compared with control subjects are selected. In some embodiments, the features include at least five taxonomic and / or gene functional features, which are optionally listed in Table 5. The features may include at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or gene functional features, which are optionally listed in Table 5. In some embodiments, the features include at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 taxonomic features optionally listed in Table 5; and at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 gene functional features optionally listed in Table 5.

[0057] In various implementations, at least three gene feature profiles are trained: one for classifying samples based on the presence or absence of CRA, one for classifying samples based on the presence or absence of CRAA, and one for classifying samples based on the presence or absence of CRC. For example, a feature profile for classifying samples based on the presence or absence of CRA can be trained using machine learning, using fecal or other biological samples from a CRA cohort and samples from a control cohort. Similarly, a feature profile for classifying samples based on the presence or absence of CRC can be trained using machine learning, using fecal or other biological samples from a CRAA cohort and samples from a control cohort. Machine learning can be as described above and can include supervised machine learning, unsupervised machine learning, or semi-supervised machine learning, or a combination thereof.

[0058] Features can be selected as described above. For example, a feature spectrum may include features selected from the training cohort by ensemble ranking of feature importance. Alternatively or additionally, features may be selected based on the statistical significance (individually) of features to predict CRA, CRAA, or CRC. For example, individual features may be selected based on the abundance or detection rate of an individual feature that significantly predicts the presence or absence of CRC, CRA, or CRAA, expressed as a p-value less than or equal to 0.05 in the training group, or a p-value less than or equal to 0.01 in the training group, or a p-value less than or equal to 0.005, or a p-value less than or equal to 0.001 (or any selected statistical threshold). In some embodiments, the feature spectrum includes features selected from the training cohort by Feature Importance Ranking Integration (FIRE) and Statistical Inference of Associations Between Microbial Communities and Phenotypes (SIAMCAT).

[0059] In various embodiments, the created characteristic spectrum has a sensitivity for classifying a sample based on the presence or absence of CRA, CRAA, or CRC, said sensitivity being at least about 0.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about 0.95. In various embodiments, the created characteristic spectrum has a specificity for classifying a sample based on the presence or absence of CRA, CRAA, or CRC, said specificity being at least about 0.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about 0.95. For example, the characteristic spectrum can provide a sensitivity of at least 0.75 for classifying each of CRA, CRAA, and CRC. For example, the created characteristic spectrum may have an area under the curve (AUC) for classifying samples for the presence or absence of CRA, CRAA, or CRC, wherein the area under the curve is at least about 0.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about 0.95.

[0060] In other aspects, this disclosure provides a method for preparing a genetic profile of fecal (or other biological samples) genetic elements indicative of colonic diseases (including, but not limited to, CRA, CRAA, and CRC). In various embodiments, the method includes providing a training cohort of fecal or other biological samples (or DNA isolated therefrom) from subjects confirmed to have colonic diseases and control subjects, and performing genomic nucleic acid sequencing on the DNA isolated from the samples (as previously described). A genetic profile is then trained to classify the samples for the presence of colonic diseases and for the absence of colonic diseases. According to this aspect, the genetic profile includes features selected from the training cohort by an integrated ranking of feature importance and the statistical significance of individual features.

[0061] For example, a feature spectrum may include features selected from the training cohort based on an integrated ranking of feature importance and (individually) the statistical significance of features predicting colonic disease. For instance, individual features may be selected whose abundance or detection rate predicts the presence or absence of colonic disease, wherein the p-value in the training group is less than or equal to 0.05, or less than or equal to 0.01, or less than or equal to 0.005, or less than or equal to 0.001 (or other selected statistical thresholds). In some embodiments, the feature spectrum includes features selected from the training cohort through Feature Importance Ranking Integration (FIRE) and Statistical Inference of Associations Between Microbial Communities and Phenotypes (SIAMCAT).

[0062] In various implementation schemes, the colonic disease is selected from Crohn's disease, ulcerative colitis, irritable bowel syndrome (IBS), diverticulitis, colorectal adenoma (CRA), advanced colorectal adenoma (CRAA), and colorectal cancer (CRC).

[0063] In various embodiments, the features include microbial taxonomic features and microbial gene function features as described above. Tables 3, 4, and 5 illustrate exemplary taxonomic and gene function features for CRA, CRAA, and CRC, respectively. For example, microbial taxonomic features and microbial gene function features that show differential abundance or differential detection rate in fecal or other biological samples from subjects with colonic diseases compared to control subjects are selected. In some embodiments, the features include at least five taxonomic and / or gene function features. In some embodiments, the features include at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or gene function features. In embodiments, the features include at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 taxonomic features; and at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 gene function features.

[0064] In various implementations, one or more machine learning algorithms are used to train the feature spectrum, as described above, including supervised machine learning, unsupervised machine learning, or semi-supervised machine learning, or a combination thereof.

[0065] In various embodiments, the generated feature spectrum has a sensitivity for classifying the sample based on the presence or absence of colonic disease, said sensitivity being at least about 0.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about 0.95. In various embodiments, the generated feature spectrum has a specificity for classifying the sample based on the presence or absence of colonic disease, said specificity being at least about 0.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about 0.95. For example, the created feature spectrum may have an area under the curve (AUC) for classifying the sample based on the presence or absence of colonic disease, said AUC being at least about 0.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about 0.95.

[0066] As used in this article, the term “about” means ±10% of the relevant value unless the context requires otherwise.

[0067] Other aspects and embodiments of the invention will become apparent from the following non-limiting examples.

[0068] Example

[0069] Example 1: Materials and Methods

[0070] Data normalization and visualization via PCoA. Using the edgeR package, the discrete taxonomic counts were normalized using the M-value weighted trimmed mean (TMM), and converted to log-CPM using Voom, implemented in the 'limma' package in R version 4.2.1. Robinson et al. , edgeR: a Bioconductor package for differential expression analysis of digital gene expression data , Bioinformatics 2010; 26(1):139-40;Law et al. , voom: Precision weights unlock linear model analysis tools for RNA-seq read counts Genome Biol 2014; 15(2):R29. The data were then log-transformed and further normalized using a supervised normalization method (SNM) to remove significant batch effects between items while preserving biological differences between disease categories. Poore et al. , Microbiome analyzes of blood and tissues suggest cancer diagnostic approach Nature 2020;579(7800):567-574. Supervised Normalization (SNM) Method for Microarrays in R. Mecham et al. , Supervised normalization of microarrays Bioinformatics Implemented in the 'snm' package of 2010; 26(10):1308-15. The effects of supervised normalization were visualized using principal coordinate analysis (PCoA). PCoA was performed using Euclidean distance on a table of relative abundance values ​​and SNM transformed counts from all taxa (kingdom to species). Differences in microbial composition between items and disease categories were assessed individually using the adonis function from the vegan package (available at CRAN.R-project.org / package=vegan on the World Wide Web). The mean distance between samples and centroids in the PCoA plots was compared between items using the Kruskal-Wallis test.

[0071] Machine learning.The models were created using an automated machine learning platform called DataRobot (DR; available at www.datarobot.com on the World Wide Web). A Python script was developed to automatically submit data to DR, allowing for the development of models for multiple datasets. The best model among all developed models was selected based on the maximum area under the curve (AUC) value predicted for the external test dataset. "Blender models" obtained by combining predictions from several machine learning algorithms (combining predictions from two or more models) were not used here. To classify each target, the sample set was randomly divided into a training set (80% of the samples) and a test set (20% of the samples). The training set was used to develop a set of high-performance predictive models in DR (using over ten different machine learning algorithms for classification, such as extreme gradient boosting tree classifiers, residual neural network classifiers using training plans, elastic network classifiers, and light gradient boosting tree classifiers with early stopping). The developed models were used to predict the disease state of the remaining 20% ​​of samples (data not used for training), and a top-performing model (with the highest external test AUC) was defined. Model performance on the test set parameters was ultimately supplemented by sensitivity, specificity, and accuracy of the external tests.

[0072] DataRobot feature list. Feature lists control the subset of features DataRobot uses to build its models. DataRobot automatically creates several feature lists for each project, including two main lists: informative features and DataRobot (DR)-Reduced features. Informative features are all features that provide potentially valuable information for modeling (typically all features). DR-Reduced features are a subset of features selected based on the feature influence calculation for the best model. The DR-Reduced feature list consists of features that provide 95% cumulative influence on the model. While computational analysis does not require limiting the number of features, practical considerations for further laboratory analysis (via qPCR) point to models with fewer features. Since informative feature lists typically contain almost all features in the dataset (1000–2000 in taxonomic annotations and 5000–10000 in feature annotations), while DR-Reduced feature lists contain no more than 100 features, we used a model built using the DR-Reduced feature list for comparison.

[0073] FIRE feature selection.The feature reduction and selection method "Feature Importance Rank Integration" (FIRE, available at www.datarobot.com / blog / using-feature-importance-rank-ensembling-fire-for-advanced-feature-selection / ) was used. In this method, features are derived from multiple different predictive models built by DataRobot (DR). By default, DR ranks the models according to a selected criterion, such as external test AUC. The median rank of each feature is calculated by aggregating the ranks of each of several top-ranking models (the number of top-ranking models to be considered is empirically chosen and equals five).

[0074] Feature importance in DR is calculated using an algorithm that measures the information content of variables—this calculation is performed independently for each feature in the dataset. In some implementations, the FIRE procedure includes the following steps: (a) calculating the feature importance of the top models (e.g., 3, 4, 5, 6, 7, 8, 9, 10) (determined by external test AUC), (b) obtaining the rank of the features, (c) calculating the median rank of each feature, (d) sorting the aggregated list by the calculated median rank, (e) defining a threshold number of features to select, and (f) defining a feature list based on the newly selected features, and (g) removing redundant features selected by two or more models. By sorting the aggregated list by median rank, we obtain the ranked list of feature importance.

[0075] Since the optimal number of features was unknown, we iteratively tested several thresholds in the first loop with large increments (800, 700, 600, 500, 400, 300, 200, 100, 90, 80, 70, 60, 50, 40, 30), and in the second loop with small increments around the first threshold found (e.g., 85, 84, 83, 82, 81, 79, 78, 77, 76, 75, if the optimal threshold obtained in the first loop equals 80). Finally, we adopted the threshold that provided the highest external test AUC. We considered the maximum number of FIRE features to be 800 because the total number of features was between 800 and 900 in some datasets analyzed. We implemented FIRE in a multi-step process to improve the accuracy of disease category prediction, and furthermore, to leverage the ensemble nature of FIRE to build feature selection, which improves generalization since its basis does not depend on a single model and its associated biases.

[0076] SIAMCAT feature selection.Another feature selection method is based on statistical inference of the association between microbial communities and phenotypes, and is called Statistical Inference of Associations Between Microbial Communities and Host Phenotypes (SIAMCAT) version 2.1.0. Wirbel et al. , Microbiome meta-analysis and cross-disease comparison enabled by the SIAMCAT machine learning toolbox , Genome Biol 2021;22(1):93. SIAMCAT is part of the computational microbiome analysis toolkit developed by EMBL. SIAMCAT provides signs for characteristic abundances (with p-values ​​and adjusted p-values), thus showing the sign of abundance changes between normal and disease states. Additionally, we used an adjusted p-value (adj_pval) of 0.001 in this analysis. Raw data were preprocessed and taxonomically analyzed using bioBakery 3 workflow version 3.0.0a7.

[0077] Combine FIRE and SIAMCAT. Since the FIRE and SIAMCAT feature lists were selected based on different criteria (the first criterion is ensemble ranking, and the second criterion is statistical significance), we decided to also test a list that combines the FIRE and SIAMCAT lists to create a new FIRE_SIAMCAT feature list.

[0078] Example 2: Analysis of publicly available datasets

[0079] To evaluate the accuracy and feasibility of developing a non-invasive diagnostic stool test for early and late-stage precancerous adenomas and advanced cancers, we imported data from 13 studies conducted by laboratories in 8 countries and analyzed stool samples using shotgun metagenomic sequencing. Wirbel et al. , Meta-analysis of fecal metagenomes reveals global microbial signatures that are specific for colorectal cancer , Nat Med 2019; 25(4):679-689 Thomas et al. , of colorectal cancer datasets identifies cross-cohort microbial diagnostic signatures and a link with choline degradation , Nat Med 2019; 25(4):667-678;Yu et al. , Metagenomic analysis of faecal microbiome as a tool towards targeted non-invasive biomarkers for colorectal cancer , Gut 2017; 66(1):70-78;Feng et al. , Gut microbiome development along the colorectal adenoma-carcinoma sequence , Nat Commun 2015; 6:6528; Zeller et al. , Potential of fecal microbiota for early-stage detection of colorectal cancer , Mol Syst Biol2014; 10(11):766;Yachida et al. , Metagenomic and metabolomic analyses reveal distinct stage-specific phenotypes of the gut microbiota in colorectal cancer , Nat Med 2019; 25(6):968-976;Gao et al. , Alterations, Interactions, and Diagnostic Potential of Gut Bacteria and Viruses in Colorectal Cancer , Front Cell Infect Microbiol 2021; 11:657867;Gupta et al. , Association of Flavonifractor plautii , a Flavonoid-Degrading Bacterium, with the Gut Microbiome of Colorectal Cancer Patients in India , mSystems 2019 Nov12;4(6):e00438-19; Hale et al. , Distinct microbes, metabolites, and ecologies define the microbiome in deficient and proficient mismatch repair colorectal cancers , Genome Med 2018; 10(1):78;Vogtmann et al. , Colorectal Cancer and the Human Gut Microbiome: Reproducibility with Whole-Genome Shotgun Sequencing , PLoS One 2016; 11(5):e0155362; Spanogiannopoulos et al. , Host and gut bacteria share metabolic pathways for anti-cancer drug metabolism , Nat Microbiol 2022;7(10):1605-1620; Hannigan et al. , Diagnostic Potential and Interactive Dynamics of the Colorectal Cancer Virome , mBio 2018; 9(6):e02248-18. In summary, our analysis included sequence data generated from 1,705 participants, comprising 703 healthy controls, 196 cases of precancerous colorectal adenoma (CRA), 48 cases of advanced precancerous colorectal adenoma (CRAA), and 765 cases of colorectal cancer (CRC), all confirmed by colonoscopy. Descriptive statistics for each study cohort are shown in Table 1.

[0080] To analyze publicly available datasets, we first normalize the data using the M-value weighted trimmed mean (TMM) and the number of log-counts per million (TMM-Voom data), and then further normalize using a supervised method (SNM), such as Mecham. wait people , Supervised normalization of microarrays BioinformaticsAs described in 2010; 26(10):1308-15. The effects of supervised normalization were visualized using principal coordinate analysis (PCoA). PCoA was performed on Euclidean distance pairs on TMM-Voom and TMM-Voom-SNM transformed count tables that included taxa from all ranks (kingdom to species). Differences in microbial composition between items and disease categories were assessed separately.

[0081] The PCoA plot shows that supervised normalization significantly reduced the variation that could be explained across different items, from R... 2 =10.3% reduced to 0.0975% ( Figure 1 A to Figure 1 B). Using TMM-Voom normalized data, only 1.5% of the variation can be explained by disease category ( Figure 1 C). After supervised normalization, this variation decreased to 0.896%, although the differences between disease types remained significant. Figure 1 D). To identify items and samples representing potential outliers, we determined the distance to the centroid within each item. For both TMM-Voom and TMM-Voom-SNM data, these distances were similar across different items. Figure 12 A, Figure 12 B). While not ideal, project-specific variation was entirely expected. We chose to confront the performance optimization challenges of these datasets rather than attempting to eliminate studies based on provisional criteria. Furthermore, it is difficult to distinguish between study-specific and country-specific effects. Therefore, the possibility that biomarkers associated with adenomas and cancer are susceptible to country- or region-specific effects remains unresolved. These results highlight the significant challenges associated with meta-analyses of gut microbiome data. These analyses exemplify, Among other things The β dispersion was similar across different studies, but as expected, there was significant variability among samples within each study.

[0082] To further examine the characteristics relevant to meta-analysis, namely study variability and population variability, we used data from individual studies representing different population cohorts to train the models and determined the predictive power of each model for all other studies. Figure 2In most cases, models trained using data from specific studies perform well, although not always optimally. Some studies used for training generated relatively high AUCs on the test sets of all or most studies. This result may be a way to measure the generalization of characteristics derived from a specific study population. Other studies used as training sets predicted one or more studies with high AUCs, but overall showed greater variability. Finally, some studies performed relatively poorly in most or all studies. The variability of AUCs generated in this analysis illustrates one of the major challenges associated with meta-analysis of microbiome profiles. The causes of cross-study variability can be numerous, including sampling differences, methodological variations, geographic effects, cohort demographics, etc. This analysis suggests that... Among other things Some studies demonstrated high predictive power for CRC samples used in their own studies and several additional studies. No single study was able to predict CRC well across all studies. It should be noted that even the highest quality studies in this analysis can only be considered relative to the weakest performing studies. The factors contributing to study quality are diverse and difficult to define.

[0083] Example 3: Feature Generation

[0084] In our efforts to develop a diagnostic analysis workflow for the fecal microbiome, we focused on evaluating two key data characteristics. First, the relative richness of taxonomic features enumerated through the bioBakery workflow. See, for example, McIver. et al. , bioBakery: a meta'omic analysis environment . Bioinformatics 2018;34(7):1235-1237. Secondly, we explored the inclusion of gene signatures derived from shotgun metagenomic sequence analysis. In this report, we evaluated KO gene functional annotation. The KEGG orthologs (KO) group is a molecular functional database represented by functional orthologs. Kanehisa et al. , KEGG: new perspectives on genomes, pathways, diseases and drugs , Nucleic Acids Res 2017; 45(D1):D353-D361. Importantly, genetic traits generally have higher relative abundance because each genetic trait represents the sum of orthologous genes across the entire community. This property is considered potentially beneficial compared to taxonomic traits, which often suffer from sparsity problems. Not wishing to be bound by theory, we hypothesize that genetic traits may positively contribute to predictive performance and act as complementary features to taxonomic traits, and we test this hypothesis.

[0085] Example 4: Feature Processing and Selection

[0086] We implemented several strategies to evaluate a variety of feature reduction methods to compare their overall impact on predictive accuracy. Each of these methods has its own specific advantages and disadvantages. Feature selection schemes based on filtering data to remove low-detection features are risky in the context of fecal microbiota, as many of the best diagnostic features (species) are low in abundance and often have low detection rates. This realization implies the need for feature selection in a way that preserves informative features. This fact also suggests that the best-performing model may require a larger number of features to achieve optimal accuracy. Here, we evaluate a feature reduction method called Feature Importance Ranking Integration (FIRE). The unique aspect of this method is that the highest-ranked features are derived from multiple different models. We select the top 5 models in our integration procedure (by external test AUC). We have added another statistically based feature selection method, called SIAMCAT (described below), to define a novel workflow. Figure 3 A key finding from FIRE implementation is that the best-performing model varies by disease category, highlighting that no single model or set of models can optimally distinguish between health and disease (CRA, CRAA, CRC). Therefore, we performed FIRE using independent feature selection and training on CRA, CRAA, and CRC samples to achieve target-specific model optimization, thereby achieving optimal disease-category-specific diagnostic performance.

[0087] As an additional layer of integration, we process taxonomic and genetic features using a tool called SIAMCAT, which allows visualization of differential abundance, detection rate, feature AUC, and ranking of features based on statistical significance. Feature selection based on SIAMCAT allows setting any significance cutoff value. In this study, we used features with corrected p-values ​​p < 0.001. Unsurprisingly, the features generated by FIRE and SIAMCAT partially overlap. Our workflow combines a relatively large number of features generated by FIRE and SIAMCAT, with any redundancy removed. The unique features generated by the combination of FIRE and SIAMCAT represent a new feature list used to classify samples into healthy or disease categories. In practice, any number of features can be selected, but the optimal number must be determined empirically (see below). In this embodiment, FIRE is run on a mixture of taxonomic and genetic features, and the results of the top 5 models are shown.

[0088] Example 5: FIRE Feature Selection

[0089] To determine the optimal number of features, FIRE is iteratively performed starting with a large number of features, such as 800-1000, to establish baseline performance based on external test AUC. We performed these analyses separately for each disease target, as well as taxonomic and genetic features (Table 2). When evaluating taxonomic and functional genetic features separately, we did not observe any common patterns across different disease categories. For CRC, 800 genetic features and 400 taxonomic features provided optimal performance. This contrasts sharply with the CRAA and CRA analyses. For CRAA, the optimal number of genetic features was significantly lower (40), and the number of taxonomic features was also lower (100). For CRA, we observed that 40 genetic features and 70 taxonomic features were optimal for performance.

[0090] These results surprisingly show that the optimal number of features for CRC (both taxonomic and genetic) is significantly higher than that identified for CRAA and CRA (Table 2). The reason for this is unclear, but it may reflect that colorectal tumors have the greatest impact on the colonic microbiota and its encoded functions, thus generating a wider range of discriminative biomarkers. For both CRAA and CRA, a relatively small number of genetic features (40) were identified as optimal. Not wanting to be bound by theory, we speculate that this may reflect the relative scarcity of discriminative genetic and taxonomic biomarkers in the early stages of the disease.

[0091] Table 3 presents a list of the top 300 characteristics of colorectal adenomas (CRAs), showing the taxonomy of the identified genera. Table 3 also shows the fold change in relative abundance compared to CRA-negative samples, the detection rate shift, and the ranking of the weights or importance indicating the change. The detection rate shift between the two categories is positive when the detection rate is high in CRAs, and negative when the detection rate is high in the control group. In some cases, the fold change is zero, therefore the detection rate shift column indicates the detection rate in CRAs.

[0092] Table 4 presents a list of the top 300 characteristics of advanced colorectal adenomas (CRAA), showing the taxonomy of identified genera. Table 4 also shows the fold change in relative abundance compared to CRAA-negative samples, the detection rate shift, and the ranking of the weights or importance indicating the change. The detection rate shift between the two categories is positive when the detection rate is high in CRAA, and negative when the detection rate is high in the control group. In some cases, the fold change is zero, so the detection rate shift column indicates the detection rate in CRAA.

[0093] Table 5 presents a list of the top 300 characteristics of colorectal cancer (CRC), showing the taxonomy of identified genera. Table 5 also shows the fold change in relative abundance compared to CRC-negative samples, the detection rate shift, and the ranking of the weights or importance indicating the change. The detection rate shift between the two categories is positive when the detection rate is high in CRC, and negative when the detection rate is high in the control group. In some cases, the fold change is zero, so the detection rate shift column indicates the detection rate in CRC.

[0094] The taxonomic characteristics of CRCs are unique because they are highly enriched with features that are overemphasized in diseases and species that typically colonize the oral cavity. Most studies examining these microbes have focused on their behavior in the oral cavity rather than the gut, but mounting evidence suggests that these taxa are pathogenic bacteria capable of causing or contributing to disease in a variety of conditions.

[0095] Several interesting observations can be drawn from this output. Perhaps the most important fact is that the first 24 features reflect bacterial species commonly found in the oral cavity. While the results clearly indicate the presence of these organisms in a healthy gut microbiota, each of these taxa shows an increase in relative abundance (fold change) and detection rate (non-zero measurement frequency) in the CRC relative to the control. Compared to healthy control subjects, the relative abundance of all but three taxonomic features in the CRC was increased. Two of these features, intestinal Rosellariasis and anaerobic Corynebacterium (…), were… Anaerostipes Members of this genus are part of the normal symbiotic gut microbiota. The specific reasons for their decreased relative abundance and increased detection rate are unclear. The third case was unexpected, involving a known oral bacterium, *Streptococcus salivarius*. The possible reasons for its decreased relative abundance and increased detection rate are also unclear.

[0096] The overemphasis of oral microbes in CRC fecal samples is consistent with the view that the tumor microenvironment selects these oral species through a combination of unknown adaptive advantages lacking in healthy individuals and / or ineffective defense mechanisms in CRC. While the factors driving this adaptive advantage may be complex, one factor that could explain these results is the metabolic shift accompanying the transformation of colon cancer epithelium from healthy and adenomatous to carcinoma, where aerobic processes in the gut, driven by butyrate oxidation for energy, are replaced by anaerobic lactic acid fermentation. A significant consequence of this metabolic shift is increased oxygen tension in the tumor microenvironment. This increased oxygen content may be sufficient to induce the observed aerobic oral species, or at least a contributing factor to the positive selection of observed aerobic oral species.

[0097] Example 6: SIAMCAT Feature Selection

[0098] We explored an independent method for feature selection called SIAMCAT. This method calculates and displays the relative abundance of each feature in all analyzed samples, the statistical significance of features showing central differences in the dataset, the fold change observed between healthy controls and each disease category, the change in detection rate, and the feature AUC (). Figure 4 In this embodiment, feature importance is for CRC that uses only taxonomic features.

[0099] The SIAMCAT output clearly shows that while the fold changes of several taxonomic features are negligible, these same species exhibit significant shifts in detection rates. Unsurprisingly, many top-level features generated by FIRE and SIAMCAT overlap, but several features are unique to one or the other. This provides a reasonable basis for combining non-redundant features to assess their relative impact on classification performance and to evaluate the relative performance advantage of our method. We evaluate the potential incremental improvements of our method by generating AUCs using the best model from a machine learning algorithm to establish a baseline for comparison with FIRE and SIAMCAT used alone or in combination. Figure 5 While FIRE feature selection generally outperforms SIAMCAT, the value of SIAMCAT is clearly evident in cases where combining FIRE and SIAMCAT to analyze features and results yields optimal performance. Figure 5 Furthermore, for all analyses involving non-redundant features derived from FIRE and SIAMCAT, SIAMCAT features consistently remain among the most important features contributing positively to the AUC of external tests. Applying SIAMCAT to control and CRC samples generated a list of taxonomic feature importance that is highly consistent with the taxonomic units reported in several studies.

[0100] We observed that optimal performance depended on the disease target. The best performance for CRA was achieved when combining taxonomic and KO features selected by FIRE-SIAMCAT. This approach increased the external AUC by nearly 8% (baseline AUC = 0.80 vs. 0.87). Analysis of CRAA performance was slightly more complex. All feature selection strategies performed best when using FIRE, while performance differences were minimal when using taxonomic features alone or in combination with genomic features. The results for CRC were similar to those for CRAA. We observed that taxonomic and genomic features outperformed taxonomic features alone, and taxonomic features alone outperformed genomic features alone. Features selected by FIRE generated a 3% gain in external AUC (baseline AUC = 0.94 vs. 0.97). The results for CRC showed that the combination of taxonomic and KO features outperformed taxonomic features alone, and taxonomic features alone outperformed KO features alone. The best performance was observed from features selected by FIRE, resulting in a modest 2% increase in external AUC (baseline AUC = 0.80 vs. 0.82). These results highlight the challenges in optimizing model performance. Different computational workflows need to be defined for the microbiota and microbiome associated with each disease category, because no single model can perform best in all three disease categories.

[0101] To obtain biological insights into the features that most significantly contribute to disease category prediction or potential common features shared between disease categories, we evaluated the overlap of features between different disease categories (generated by a combination of FIRE and SIAMCAT). Figure 6Notably, when comparing the top 20 key features for each disease category (bottom right Venn diagram), there was no overlap between taxonomic and genetic features. Another interesting difference in this comparison is that 60% of the features for CRC are taxonomic, significantly larger than those observed for CRA (40%) and CRAA (25%). These differences were preserved but weakened when examining the top 50 features (bottom middle Venn diagram). In the top 50 features, we begin to observe moderate overlap between different disease categories. Comparing the top 100 features shows that the proportion of taxonomic features is fairly balanced across different disease categories. As more features are compared, the proportion of genetic features relative to taxonomic features increases, and we observe increasing overlap between different disease categories. An examination of 800 features reveals potentially interesting biological aspects of the microbiome in the context of colorectal cancer. Notably, among the overlapping features, genetic features dominate relative to taxonomic features. Particularly interesting is that we observed no overlapping taxonomic features across all three disease categories, but 48 genetic features (2.4%) were shared. This imbalance is also evident in all pairwise comparisons of overlapping features, with taxonomic features accounting for approximately 9–14% of overlapping features. Although the frequency of shared taxonomic features is similar between different disease categories, it is noteworthy that the number of shared genetic features between CRC and CRA is significantly higher relative to any other pairwise relationship.

[0102] Next, we improved these comparisons by considering the direction of change in both trait types—that is, increase or decrease in disease—to determine the true biological similarity of shared taxonomic and genetic traits. Figure 7 Considering the top 800 features for each disease category in a binary manner (top left; all increases >0 or all decreases <0 relative to healthy controls), most overlapping features appear between CRA and CRC samples, although some overlap still exists between CRAA and CRC. Further analysis of these relationships in different arrangements (top right) shows the relationship between CRA and CRAA samples, highlighting a small number of shared features showing the same direction of change. These plots also highlight the large number of overlapping features shared between CRC and CRAA showing opposite directions of change. A visual comparison of CRA and CRC samples is possible by comparing the bottom left and bottom right Venn diagrams. Considering the still large number of shared features when considering the direction of change, the remaining plots (bottom left: CRC FC<0, CRA FC>0, CRAA FC<0 and bottom right: CRC FC>0, CRA FC<0, CRAA FC>0) illustrate that most shared features between CRA and CRAA represent changes in opposite directions.

[0103] Therefore, although CRA and CRC share many common features and their changes usually occur in the same direction ( Figure 7 This result is quite surprising and difficult to interpret, but it suggests that the gut microenvironment and selective pressures are similar in CRA and CRC, but different in CRAA. Without being bound by theory, we hypothesize that the higher proportion of shared genetic traits relative to taxonomic traits may reflect functional redundancy in closely related or even distant taxa that are essentially interchangeable within the community due to shared genes encoded in their respective genomes. The high proportion of genetic traits in the expanded trait importance list suggests the need for a detailed analysis of overpresented and underpresented functional attributes in both health and disease.

[0104] Example 7: Model Validation: Taxonomic Feature Analysis.

[0105] To assess whether the top 800 taxonomic features for each category showed consistency and / or bias toward a specific phylogenetic group, we analyzed important features at the class and family levels. Of the top 800 features, 66 represented overpresented or underpresented taxa in the CRA, 86 represented overpresented or underpresented taxa in the CRAA, and 79 represented overpresented or underpresented taxa in the CRC. The feature importance list contains a total of 12 class taxa. Figure 8 It should be noted that these features are not necessarily statistically significant in comparisons, but are considered discriminative based on the AI ​​model used. Clostridia had the most underpresented features in the CRA. This class was significantly overpresented in both the CRAA and CRC.

[0106] Among the top-level characteristics, the next most important class is Bacteroidetes ( Bacteroidia Of the 15 taxa in the top-level characteristics of the CRA, twelve showed positive fold changes, while nine taxa in the CRC showed positive fold changes. In contrast, the CRAA had fewer identifiable taxa in this class and exhibited a significant deficiency. (The two classes (Mucariaceae)...) Tissierellia ) and Fusobacteria ( Fusobacteriia This is overemphasized and unique in the CRC high importance list, but not in the CRA or CRAA. The class most indicative of CRAA is Methanobacteria ( Methanobacteria This is uniquely overpresented in CRAA, but not in CRA or CRC. Two additional classes, Actinobacteria (… Actinobacteria ) and Rhodotorula ( CoriobacteriaThis trait is significantly overpresented in the CRAA, but absent or underpresented in the CRA and CRC trait importance lists. Given the vast phylogenetic space and evolutionary distance within the class Bacteria, the consistency of these results is remarkable.

[0107] Next, we examined these results at a higher family resolution. The top features across all disease categories involved 42 different bacterial families. As we observed when analyzing classes, we found a preference for specific families in the distribution of important features for each disease category. Examining those belonging to the Methanobacteria class (…) Methanobacteria Actinomycetes ( Actinobacteria The families of the three classes (including) and the class Rhodotorulatae, indicating that these families have a significant ability to distinguish CRAA from other disease categories and healthy subjects. Figure 9 For CRA, Barnesaceae ( Barnesiellaceae Two species within the family ) uniquely exhibited deficiency, while those from the family Lunaedomycetes ( Selenomonadaceae ), Sartorius ( Sutterellaceae ) and Neisseriaceae ( Neiseriaceae Individual species within the family CRAA are all overpresented. Regarding CRAA, the family Aminococcidae (…) Acidaminococcaceae Two species within the family Methanobacteria are uniquely underrepresented in the CRAA. Methanobactericeae ) and Propionibacterium ( Propionibacteriaceae The overpresentation of two species within this family is unique to the CRAA trait importance list. Tanneraceae ( Tannerellaceae Two species within this family are uniquely overrepresented in the CRA but underrepresented in the CRAA. Other families, including Actinidiaceae (…), are also included. Actinomycetaceae ) and Igerzilidae ( Eggerthellaceae Species within the family Trichophyceae are overrepresented in the CRAA but underrepresented in the CRA. Lachnospiraceae While not uniquely overrepresented in the CRAA, the taxonomic unit of α-acidophilia is involved in 10 sub-taxonomic units that distinguish the CRAA from other disease categories. Two families, each with two species (Acidophilaceae, α-acidophilia ... Peptoniphilaceae ) and Fusobacteriaceae ( Fusobacteriaceae The over-presentation of )) is unique to CRC samples. Two families, including Clostridium ( Clostridiaceae (3 taxonomic units) and Erysipelothrix family ( Erysipelotrichaceae (2 taxonomic units) distinguish CRC from other disease categories and healthy controls.

[0108] Among the most important taxonomic features of the CRA, we observed 6 Bacteroides genera ( Bacteroides ) Species differences are presented. Several genera are represented by two or more species, including the genus Prevotella ( Prevotella ) species, genus *Pseudomonas* ( Parabacteroides ) species and Veillonella genus ( Veillonella ) species, all of which were overpresented compared to healthy control samples, as well as Eubacterium ( Eubacterium ) species and genus Roselle ( Roseburia Both species were underrepresented. All other diverse genera were represented by a single species. Key features of the CRAA compared to the healthy control samples showed overrepresentation of genera, including those belonging to the genus *Actinomyces* (…). Actinomyces Four species belonging to the genus Collins ( Collinsiella Three species belonging to the genus Ernomama ( Enorma Two species belonging to the genus *Lactobacillus*. Lactobacillus Two species belonging to the genus *Dolphobella* (…). Dorea Two species and belonging to the genus *Factococcus* ( Coprococcus Two species belonging to the genus *Bacteroides*. Compared to the healthy control sample, the samples were found to contain more *Bacteroides*. Bacteroides The presentation of two species of ) is insufficient. Two other fungal genera ( Alistipes The presentation of species differed significantly from the control samples. Key characteristics of CRC include two genera of actinomycetes (…). Actinomyces ) species, 3 porphyromonas genera ( Porphorymonas ) species, 4 Prevotella genera ( Prevotella ) species, 2 genera of Peptostreptococcus ( Peptostreptococcus ) species 、 3 Fusobacterium species ( Fusobacterium ) species, 3 Bacteroides genera ( Bacteroides ) species, 2 Veillonellae genus ( Veillonella Species that are overpresented in CRC relative to healthy controls. Several of these features do not show significant fold changes relative to controls, but do show significant changes in detection rates. All other genera are represented by a single species. Interestingly, no taxa in the CRC feature importance list show underpresentation.

[0109] like Figure 10 As shown, the following taxonomic characteristics have increased in the CRA: cariogenic actinomycetes ( Actinomyces odontolytic Bifidobacterium adolescentis () Bifidobacterium adolescentis ), Bifidobacterium longum ( Bifidobacterium longum ), mouse cecal bacillus ( Enterorhabdus caecimuris ), Gordon's bacillus ( Gordonibacter pamelaea ), Bacteroides igerzii ( Bacteroides eggerthii ), Enterobacterium ( Intestinal Bacteroides ), Bacteroides nori ( Bacteroides nordii ), Bacteroides ( Bacteroides plebeius ), Bacteroides salisii Bacteroides salyersiae ), fecal Bacteroides ( Bacteroides suffocatus ), human enterobacteria ( Barnesiella intestinihomonis ), Toxic butyric acid bacteria ( Butyricimonas virosa Prevotella Copri ( Prevotella copri ), fecal prednisone ( Prevotella scatcorea ), Sachetsia sahedron ( Alistipes shahii ), Parabacterium Gordonii ( Parabacteroides gordonii ), Parabacterium goeringii ( Parabacteroides goldsteinii ), hematogenous twin cocci ( Blood twin Streptococcus thermophilus () Streptococcus thermophilus ), Protuberance Eubacterium ( Eubacterium ventriosum ), strong anaerobic bacteria ( Anaerostipes hadrus ), Broutbacterium ovale ( I am dying. ), Blautia wexleri ( Blautia wexlerae ), Formic acid-producing Dorper ( Golden ant-producing ), Long-chain Dorperella ( Long-chained dorea ), Sugar-containing spindle bacteria ( Fusicatenibacter saccharivorans ), Proctobacterium rectum ( Eubacterium rectale ), Roseola species CAG 309 ( Roseburiasp. CAG 309), species of the genus *Roseidon* CAG 431 ( Roseburia sp. CAG431), CAG 241 (a species of the genus *Vibrio*) Oscillibacter sp. CAG 241), Plasminogen pyrenoidosa ( Faecalibacterium prausnitzii ), Clostridium perfringens ( Clostridium leptum ), Rumenaceae bacteria D16 ( Ruminococcaceae bacteria D16), Gastrococcus lactis ( Ruminococcus lactaris ), Clostridium helicobacter ( Clostridium spiroforme Firmicutes bacteria CAG 110 ( Firmicutes bacteria CAG 110), atypical Veillonella ( Veillonella atypical ), Veillonella Tobetsucho ( Veillonella Tobetsu ), Neisseria palea Neisseria flavuscens ), Klebsiella pneumoniae ( Klebsiella pneumonia ) and Klebsiella phytogenes ( Klebsiella variicola Among the most important taxonomic features of the CRA, we observed six Bacteroides genera ( Bacteroides The differences between species are presented. Several genera are represented by two different species, including the genus Prevotella (…). Prevotella ) species, genus *Pseudomonas* ( Parabacteroids ) species and Veillonella genus ( Veillonella ) species, all of which were overpresented compared to healthy control samples, as well as Eubacterium ( Eubacterium ) species and genus Roselle ( Rosemary Both of these species are insufficient. All other genera are represented by a single species.

[0110] For example Figure 10 As shown, the following taxonomic characteristics are added in the CRAA: *Methanobacterium spp.* Methanobrevibacter smithii ), Bifidobacterium longum ( Bifidobacterium longum ), Propionibacterium fischeri ( Propionibacterium freudenreichii ), Olsenella odorifera ( Olsenella scatoligenes ), Collins acetifera ( Collinsella aerofaciens ), Enteric Collins bacteria ( Collinsella intestinal fecal Collins bacteria ( Collinsella dung beetle ), Massey Ernomazic bacillus ( Huge from Marseille ), Estrol-producing Adlerkleinichthys ( Adlercreutzia equilifaciens ), hidden non-sugar bacteria ( Asaccharobacter celatus ), Gordon's bacillus ( Gordonibacter pamelaea ), Slack bacteria that convert isoflavones ( Slackia isoflavoniconvertens ), Bacteroides ovalis ( Bacteroides ovatus ), Bacteroides multiforme ( Bacteroides thetaiotaomicron ), Bacteroides monomorpha ( Bacteroides uniformis ), Bacteroides commonis ( Bacteroides vulgaris ), Bacteroides xylan Bacteroides xylanisolvens ), lack of other branches ( The helpless alistripes ), putrefactive bacteria ( Rot streaks ), Parabacterium dilatatum ( Parabacteroides distasonis ), Streptococcus suis ( Streptococcus milius Clostridium species CAG167 ( Clostridium sp. CAG 167), Hodgkin's E. ( Eubacterium hallii ), strong anaerobic bacteria ( Anaerostipes hadrus ), Brontë viridissima ( Blautia wexlerae ), Ruminococcus twistis ( Ruminococcus torciata ), Agkistrodon halys ( Coprococcus catus ), associated fecal cocci ( Coprococcus companion ), Formic acid-producing Dorper ( Dorea ant-producing ), Long-chain Dorperella ( Long-chained dorea ), Sugar-containing spindle bacteria ( Fusicatenibacter saccharivorans Clostridium botulinum ( Clostridium bolts ), fecal roseola ( Roseburia faeces ), Species CAG 241 of the genus *Vibrio* Oscillibacter sp. CAG 241), Pasteurella multocida ( Intestinibacter bartlettiiFirmicutes bacteria CAG 170 ( Firmicutes bacteria CAG 170), Firmicutes bacteria CAG 238 ( Firmicutes bacteria CAG238), Firmicutes bacteria CAG94 ( Firmicutes bacteria CAG 94), fecal colic ( Phascolarctobacterium faecium ) and Haemophilus parainfluenzae ( Haemophilus parainfluenzae Compared to healthy control samples, CRAA exhibited an overemphasis on key features of genera, including those belonging to the genus Actinomycetes (). Actinomyces Four species belonging to the genus Collins ( ) Collinsiella Three species belonging to the genus Ernomama ( Huge Two species belonging to the genus *Lactobacillus*. Lactobacillus Two species belonging to the genus *Dolphobella* (…). Golden Two species and belonging to the genus *Factococcus* ( Coprococcus) Two species. Compared to the healthy control sample, it belongs to the genus Bacteroides ( Bacteroides The presentation of two species of ) is insufficient. Two other fungal genera ( Common grebe The species presented differed significantly from the control sample.

[0111] For example Figure 10 As shown, the following taxonomic characteristics have increased in the CRC: Zurich Actinobacterium ( Actinomyces from Turkey ), Bifidobacterium chain Bifidobacterium catenulatum ), Collins acetifera ( Collinsella aerofaciens ), demanding Slack bacteria ( Slackia exigusa ), Bacteroides fragilis ( Bacteroides fragilis ), Bacteroides nori ( Bacteroides nordii ), Bacteroides ( Bacteroides plebeian ), Toxic butyric acid bacteria ( Butyricimonas virosa ), Porphyromonas unsaturates ( Porphyromonas asaccharolytic ), Porphyromonas pulpis ( Porphyromonas endodontalis ), Porphyromonas ursae ( Porphyromonas veononis ), Tannabepus ( Alloprevotella tannerae Prevotella intermedia ( Prevotella intermedia Prevotella melanogaster ( Prevotella nigrescens Prevotella species CAG 520 Prevotella sp. CAG 520), fecal Prevotella ( Prevotellastercorea ), Measles twins ( Twin measles ), Pasteurella multocida ( Streptococcus pasteurianus ), Streptococcus salivarius ( Streptococcus salivarius Clostridium species CAG 58 ( Clostridium sp.CAG 58), Hargartella hominis ( Hungatella hathawayi ), Bacillus difficile ( Mogibacterium diversum ), picky eubacterium ( Choosing Eubactenum ), Mycobacteria ( Eubacterium ramulus ), Protuberance Eubacterium ( Eubacterium ventrally ), strong anaerobic bacteria ( Anaerostipes hadrus ), Agkistrodon halys ( Coprococcus cat ), Eisenberger thyrizinski ( Eisenbergiella tayi ), symbiotic Clostridium ( Clostridium symbiosum ), Enterobacter rosenbergii ( Intestinal roseburia ), Roseola species CAG 303 ( Roseburia sp. CAG303), Anaerobic Streptococcus ( Peptostreptococcus anaerobic ), Peptostreptococcus stomatitis ( Peptostreptococcus stomatis ), Plasmodium falciparum ( Faecalibacterium prausnltzii ), Ruminaceae bacteria D16 ( Ruminococcaceae bacteria D16), Ruthenium lactis ( Ruthenibacterium lactate-forming ), Lactobacillus murineis ( Solobacterium moorei Firmicutes bacteria CAG94 ( Firmicutes bacteria CAG 94), Alistair pyogenes ( Dialister pneumosints ), Varroa minor cocci ( Veillonella small Veillonella species T110116 Veillonella sp T110116), Micromonas vaginalis ( Parvomonas micra ), Fusobacterium navicularis ( Fusobacterium ship-shaped ), Fusobacterium nucleatum ( Fusobacterium nucleatum ), Oral taxonomic unit 370 of the genus Fusobacterium ( Fusobacterium sp. oral taxon 370), E. coli ( Eikenella corrodens ), Escherichia coli ( Escherichia coli ) and Morganella morganii ( Morganella morganii Key characteristics of CRC include two genera of actinomycetes ( ). Actinomyces ) species, 3 porphyromonas genera ( Porphorymonas ) species, 4 Prevotella genera ( Prevotella ) species, 2 genera of Peptostreptococcus ( Peptostreptococcus ) species 、 3 Fusobacterium species ( Fusobacterium ) species. Several of these characteristics did not show a significant fold change relative to the control, but did show a significant increase in detection rate. Compared to healthy control samples, three Bacteroides species ( Bacteroides ) species, 2 Veillonella ( Veillonella Species are overrepresented in the CRC. All other genera are represented by a single species.

[0112] We conducted an in-depth meta-analysis of publicly available microbiome shotgun sequencing data from fecal samples of healthy control donors and those diagnosed with CRA, CRAA, and CRC via colonoscopy. The 13 studies analyzed were highly diverse in characteristics, including eight different countries, different disease states, cohort sizes, and the number of readings through quality control indicators (Table 1). Although most studies attempted to balance sex and age within their respective cohorts, males were generally more prevalent than females. Other factors, such as DNA preparation and sequencing methods, varied between studies, and importantly, some studies collected feces or other samples after colonoscopy. These factors can introduce variability into the results. Despite these confounding factors, the taxonomic features identified in individual studies, while differing, did define consensus findings, at least for CRC. Given the known inter-individual microbiome variability and the different dietary habits in each participating country, the fact that similar taxa were identified strongly suggests that selective forces at play in CRC are dominant relative to dietary and other known selective pressures.

[0113] These findings lead to two surprising conclusions regarding the contribution of the selective microenvironment and / or gut microbiota generated by colonic lesions to disease occurrence and progression. First, the significant differences between the CRA and CRAA microbiota raise questions about whether advanced adenomas are simply larger forms of adenomas. Our results not only suggest that this hypothesis requires further investigation but also indicate that the microbiota-adenoma / advanced adenoma interaction represents a highly distinct process. Second, and perhaps more surprisingly, despite the strongly opposite behaviors of the CRA and CRAA microbiota in gene presentation, they are functionally essentially equivalent. In fact, genetic features exhibiting a consistent direction of change in the CRA and CRC microbiota are not unrelated to those observed in the CRAA. Of the 408 features that differed across all three disease categories compared to healthy controls, 389 (95%) showed a consistent pattern between CRA and CRC, and an inconsistent pattern between CRAA and other disease categories. In this respect, the same genetic features that increased in both CRA and CRC samples were reduced in CRAA relative to healthy control samples. vice versa. However, these results are considered preliminary because almost all of the late-stage adenoma samples came from a single study.

[0114] Example 8: Verifying high-impact features through external data validation tests

[0115] Following the completion of the formal analysis of FIRE and SIAMCAT, data from a new study on the Spanish cohort were obtained. Front. Microbiol ., 2024; 11(14):1-17). This data was imported and used for additional external validation. The new dataset consisted of 30 CRC samples and 30 control samples. The data included an additional 30 polyp samples that were not analyzed due to the lack of a specification to distinguish between early-stage adenomas (CRA) and late-stage adenomas (CRAA). We used BB3 to establish taxonomic classification. A model developed for taxonomic annotation (Limited Gradient Boosting Tree Classifier) ​​was used to score the new external dataset using informative features and a threshold corresponding to maximizing the F1 score (0.629). The resulting predictive performance of the model yielded the following results: (AUC=0.8089, Sensitivity=0.7333, Specificity=0.8333 and Accuracy=0.7833). According to the evaluated metrics, these scores were either better or comparable to those achieved by external HO data analysis using the original metagenomic data. These results indicate that the modeling work was largely successful, reflecting that the CRC classification model exhibits good generalization and satisfactory performance on new data.

[0116] Example 9: Validation of high-impact characteristics via qPCR

[0117] Designing primers for target taxa of interest presents multiple challenges. The greatest challenge relates to the vast gap between the known sequence space occupied by the target taxa and the unknown sequence space existing on Earth. In this respect, the quality of any primer design must be deemed acceptable unless otherwise demonstrated. Targeting unique gene sequences present in the taxa of interest but absent in their nearest neighbors represents the most straightforward way to perform specific qPCR. However, while the target gene is ubiquitous in sequenced isolates, it may actually be absent in uncharacterized samples, thus generating underestimated abundance in qPCR reactions. Conversely, shotgun metagenomic sequencing of fecal samples provides imperfect sequence read mapping and limited accuracy based on the availability of known sequences. In this respect, relative abundance measurements generated by sequence enumeration may not be perfect and may therefore differ from those generated by qPCR. Many of these nuances can be evaluated directly in candidate primer design by sequencing the PCR products generated from dozens or hundreds of reactions to assess the purity of sequences in the generated products. After a comprehensive initial evaluation and before deploying any commercial testing, nonspecific primer binding or amplification of nearest-neighbor sequences can and should be quantified.

[0118] Primers were designed using a list of 300 taxonomic units generated by applying FIRE feature selection (from CRC, CRA, and CRAA for each category) and ensemble modeling, using top-level target species. Multiple primer sets were identified for each target based on gene sequences unique to the target. Primer pairs targeting total bacteria using the 16S rRNA gene were used as internal housekeeping controls to normalize results across different samples: total bacteria_16S Fw GCAGGCCTAACACATGCAAGTC (SEQ ID NO: 1), total bacteria_16S Rv CTGCTGCCTCCCGTAGGAGT (SEQ ID NO: 2), with a product size of 120 base pairs. A list of all working primers for each category is shown in Table 6.

[0119] Each primer was tested on 10–48 samples, with approximately 50% from CRC / CRA / CRAA subjects and approximately 50% from CTR (control subjects). The theoretical and outcome-based evaluations of primer design considered several metrics: Tm, melting curve, presence or absence of primer dimers, presence or absence of hairpins, Ct amplification number, base number, product size, primer pair specificity (based on Primer 3 blast alignment), and reproducibility of sequencing data.

[0120] We evaluated the quantitative relative abundance values ​​for each sample, comparing shotgun sequencing and qPCR data for each specific target taxonomic unit. The results are summarized in Figures 13A–13N (CRC), 14A–14Q (CRA), and 15A–15E (CRAA). Scatter plots show the quantitative relative abundance values ​​for each sample based on shotgun sequencing (left) and qPCR (right) for each specific target taxonomic unit. Each point represents one sample. Sequencing samples are arranged in order from lowest to highest value, and qPCR samples are ordered according to the sequence of the sequencing samples. The y-axis of the sequencing plot shows the raw data values, while the y-axis of the qPCR plot shows the relative abundance values ​​calculated using Method 2 (-Delata Delta C(T)). Bar charts show the mean values ​​of the sequencing and qPCR data for the test samples. Error bars represent the values ​​of the standard error.

[0121] Based on this data, 16 primer pairs from CRC, 18 primer pairs from CRA, and 6 primer pairs from CRAA reproduced the sequencing data. Some primer pairs (not shown) exhibited high target specificity but very low abundance of taxa. These included: *Peptostreptococcus stomatitidis* from CRC. Peptostreptococcus stomatis ) and Alisteria pulmonaryis ( Dialister pneumosintes ); CRA's fecal cocci ( Caprococcus catus); and Gravenz Actinomyces of CRAA ( Actinomyces graevenitzii ).

[0122] Table 1: Selected metagenomic items representing different subject populations used for modeling. Descriptive statistics for each item: sex, BMI, age, disease category, raw readings, post-control readings, and percentage of human readings. For continuous variables, the mean and standard deviation are shown, and for categorical variables, the number of samples within each category is shown.

[0123]

[0124]

[0125] Table 2: FIRE Feature Selection. This table reports the AUC values ​​for the external (20%) dataset. The bold underlined values ​​are the maximum external test AUC achieved for a specific annotation, target, and FIRE feature set size.

[0126] 1 Extreme gradient boosting tree classifier

[0127] 2 Keras Slim residual neural network classifier with training plan (1 layer: 64 units)

[0128] 3 Elastic network classifier (L2 / binomial bias)

[0129] 4 An early-stopping, lightweight gradient boosting tree classifier is used.

[0130]

[0131] Table 3: Feature list of colorectal adenomas (CRAs). This table provides fold changes in relative abundance, detection rate shifts, and the weights or importance of taxonomic features, changes, and shifts. Detection rate shifts between two categories are positive when the detection rate is high in CRAs, and negative when the detection rate is high in the control group.

[0132]

[0133]

[0134]

[0135]

[0136]

[0137]

[0138]

[0139]

[0140]

[0141]

[0142]

[0143]

[0144]

[0145]

[0146]

[0147]

[0148]

[0149]

[0150]

[0151]

[0152]

[0153]

[0154]

[0155]

[0156]

[0157]

[0158]

[0159]

[0160]

[0161]

[0162]

[0163]

[0164]

[0165]

[0166]

[0167]

[0168]

[0169] Table 4: Feature list of advanced colorectal adenomas (CRAA). This table provides fold changes in relative abundance, detection rate shifts, and the weights or importance of features, changes, and shifts. Detection rate shifts between two categories are positive when the detection rate is high in CRAA, and negative when the detection rate is high in the control group.

[0170]

[0171]

[0172]

[0173]

[0174]

[0175]

[0176]

[0177]

[0178]

[0179]

[0180]

[0181]

[0182]

[0183]

[0184]

[0185]

[0186]

[0187]

[0188]

[0189]

[0190]

[0191]

[0192]

[0193]

[0194]

[0195]

[0196]

[0197]

[0198]

[0199]

[0200] Table 5: Feature list of colorectal cancer (CRC). This table provides fold changes in relative abundance, detection rate shifts, and the weights or importance of features, changes, and shifts. Detection rate shifts between two categories are positive when the detection rate is high in CRC, and negative when the detection rate is high in the control group.

[0201]

[0202]

[0203]

[0204]

[0205]

[0206]

[0207]

[0208]

[0209]

[0210]

[0211]

[0212]

[0213]

[0214]

[0215]

[0216]

[0217]

[0218]

[0219]

[0220]

[0221]

[0222]

[0223]

[0224]

[0225]

[0226]

[0227]

[0228]

[0229]

[0230] Table 6: Examples of representative qPCR primers.

[0231]

[0232]

[0233]

Claims

1. A method for evaluating the presence of colorectal tumors in a biological subject, comprising: Quantify the genetic elements from biological samples from the subject, wherein the abundance or detection rate of the genetic elements is correlated with colorectal cancer (CRC), colorectal adenoma (CRA), or advanced colorectal adenoma (CRAA), thereby preparing an abundance profile of the genetic elements, and wherein the genetic elements include elements related to microbial taxonomic classification and / or genetic elements related to the function of one or more microbial genes; The abundance spectrum was evaluated based on the characteristic spectrum indicating the presence of CRC, CRA, and / or CRAA in the subjects, and Determine whether the subject may have CRC, CRA, or CRAA.

2. The method of claim 1, wherein the subject has a low risk of having CRC, CRA, CRAA, or colorectal polyps.

3. The method of claim 1 or 2, wherein the subject has not previously experienced CRC, CRA, or CRAA.

4. The method of claim 2 or 3, wherein the method is performed as an alternative to colonoscopy.

5. The method of claim 1, wherein the subject has a high or moderate risk of having CRC, CRA, CRAA, or colorectal polyps.

6. The method of claim 5, wherein the subject has a history of CRC, CRA, or CRAA, and / or has a family history of CRC.

7. The method of claim 5 or 6, wherein the method is performed at least once a year or at least once every other year.

8. The method of any one of claims 1 to 7, wherein the subject is at least 45 years old, or at least 50 years old, or at least 55 years old, or at least 60 years old.

9. The method of any one of claims 1 to 7, wherein the subject is less than 45 years of age or less than 50 years of age.

10. The method of any one of claims 1 to 9, wherein the biological sample is feces, blood, serum, intestinal mucosa, mucosal swabs, colonoscopy aspirate, irrigation fluid, or biopsy tissue sample or other biological sample.

11. The method of any one of claims 1 to 10, wherein the genetic element is quantified by a procedure including nucleic acid sequencing, PCR, qPCR, or microarray.

12. The method of claim 11, wherein the genetic element is quantified by nucleic acid sequencing, and the nucleic acid sequencing involves sequencing at least about 20,000,000 reads.

13. The method of claim 12, wherein the nucleic acid sequencing involves sequencing at least about 40,000,000 reads.

14. The method of any one of claims 11 to 13, wherein the nucleic acid sequencing comprises one or more of shotgun metagenomic sequencing, rDNA sequencing, targeted amplicon sequencing, or hybridization capture probe sequencing.

15. The method of claim 14, wherein the nucleic acid sequencing comprises 16S rDNA, 18S rDNA, or ITS amplicon sequencing; and comprises targeted amplicon nucleic acid sequencing or hybridization capture probe sequencing.

16. The method of claim 14 or 15, wherein one or more genetic elements are quantified by capturing from a sequencing library and optionally by PCR amplification followed by sequencing.

17. The method of any one of claims 1 to 16, wherein the genetic element indicates or is associated with colorectal adenoma (CRA).

18. The method of claim 17, wherein the genetic element comprises at least five taxonomic or gene functional features listed in Table 3.

19. The method of claim 18, wherein the genetic elements comprise at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or gene functional features listed in Table 3.

20. The method of claim 19, wherein the genetic element comprises at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 taxonomic features listed in Table 3; and at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 gene functional features listed in Table 3.

21. The method of any one of claims 18 to 20, wherein the genetic element has differential abundance or differential detection rate in the sample of the CRA subject compared with that of a control subject.

22. The method of any one of claims 1 to 16, wherein the genetic element is associated with advanced colorectal adenoma (CRAA).

23. The method of claim 22, wherein the genetic element comprises at least five taxonomic or gene functional features listed in Table 4.

24. The method of claim 23, wherein the genetic elements comprise at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or gene functional features listed in Table 4.

25. The method of claim 24, wherein the genetic element comprises at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 taxonomic features listed in Table 4; and at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 gene functional features listed in Table 4.

26. The method of any one of claims 23 to 25, wherein the genetic element has differential abundance or differential detection rate in the sample of the CRAA subject compared with that of the control subject.

27. The method of any one of claims 1 to 16, wherein the genetic element is associated with colorectal cancer (CRC).

28. The method of claim 27, wherein the genetic element comprises at least five taxonomic or gene functional features listed in Table 5.

29. The method of claim 28, wherein the genetic elements comprise at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or gene functional features listed in Table 5.

30. The method of claim 28 or 29, wherein the genetic element comprises at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 taxonomic features listed in Table 5; and at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 gene functional features listed in Table 5.

31. The method of any one of claims 27 to 30, wherein the genetic element has differential abundance or differential detection rate in the sample of the CRC subject compared with that of the control subject.

32. The method of any one of claims 1 to 31, wherein the abundance spectrum is evaluated for a feature spectrum indicating the presence or absence of each of CRC, CRA, and CRAA.

33. The method according to any one of claims 1 to 32, wherein: (a) Using machine learning, a feature spectrum indicating the presence or absence of CRA is trained using samples from the CRA cohort and samples from the control cohort; (b) Using machine learning, a feature spectrum indicating the presence or absence of CRC is trained using samples from the CRC cohort and samples from the control cohort; and (c) Using machine learning, a feature spectrum indicating the presence or absence of CRAA is trained using samples from the CRAA cohort and samples from the control cohort.

34. The method of claim 33, wherein multiple machine learning algorithms are used to train the feature spectrum.

35. The method of claim 33 or 34, wherein at least one machine learning algorithm is supervised machine learning.

36. The method of claim 35, wherein the machine learning algorithm further comprises one or more of unsupervised and semi-supervised machine learning.

37. The method of any one of claims 33 to 36, wherein the machine learning comprises one or more of parametric / nonparametric distance measurement, logistic regression, support vector machine, decision tree, random forest, neural network, probabilistic regression, Fisher linear discriminant, Naive Bayes classifier, perceptron, quadratic classifier, kernel estimation, k-nearest neighbor, learned vector quantization, and principal component analysis.

38. The method of any one of claims 33 to 36, wherein the machine learning includes comparative machine learning modeling, optimization, testing, evaluation, and model ranking, including one or more of deep learning, gradient boosting, neural networks, ensemble or hybrid modeling algorithms, including gradient boosting tree classifiers, extreme gradient boosting tree classifiers, light gradient boosting tree classifiers, elastic network predictive light gradient boosting, Keras Slim residual neural network classifiers, generalized additive models, elastic network classifiers, random forest classifiers, deep forest classifiers, average mixer classifiers, TensorFlow multilayer perceptron classifiers, TensorFlow neural network classifiers, and rule fitting classifiers.

39. The method of any one of claims 33 to 38, wherein the feature spectrum comprises features selected from the training queue by an integrated ranking of feature importance and the statistical significance of individual features.

40. The method of claim 39, wherein the feature spectrum comprises features selected from the training cohort via Feature Importance Ranking Integration (FIRE) and Statistical Inference of Associations Between Microbial Communities and Phenotypes (SIAMCAT).

41. The method of any one of claims 33 to 40, wherein the characteristic spectrum has a sensitivity for classifying a sample based on the presence or absence of CRA, CRAA, or CRC, said sensitivity being at least about 0.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about 0.

95.

42. The method of any one of claims 33 to 41, wherein the characteristic spectrum has specificity for classifying a sample based on the presence or absence of CRA, CRAA, or CRC, said specificity being at least about 0.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about 0.

95.

43. The method according to any one of claims 1 to 42, wherein: (a) If the subject is not identified as potentially having CRA, CRAA, or CRC, no further procedures will be performed, and (b) If the subject is identified as potentially having one or more of CRA, CRAA, or CRC, further procedures or treatments shall be initiated.

44. The method of claim 43, wherein the further procedure includes colon imaging.

45. The method of claim 44, wherein the procedure is a colonoscopy, the colonoscopy optionally involving the removal of one or more polyps and / or biopsy of growths suspected of including CRC.

46. ​​The method of claim 44 or 45, wherein if the subject is confirmed to have CRC, the subject's CRC is treated by one or more of surgery, chemotherapy, radiotherapy, and immunotherapy.

47. A method for preparing a genetic profile of genetic elements indicating the presence of colorectal tumors, the method comprising: Provide a training cohort of biological samples, or RNA or DNA isolated from, subjects confirmed to have CRA, CRAA, or CRC; The DNA isolated from the sample was subjected to genomic nucleic acid sequencing; The training involves classifying samples based on the presence or absence of CRA, CRAA, or CRC, wherein the gene signature profile includes microbial taxonomic features and microbial gene functional features.

48. The method of claim 47, wherein the sample is selected from feces, blood, serum, plasma, urine, saliva, biopsy tissue, mucosal tissue samples or swabs, and intestinal lavage fluid or aspirate.

49. The method of claim 48, wherein the sample is a fecal sample or a mucosal tissue sample.

50. The method of any one of claims 47 to 49, wherein the nucleic acid sequencing involves sequencing at least about 20,000,000 reads per sample.

51. The method of claim 50, wherein the nucleic acid sequencing involves sequencing at least about 40,000,000 reads.

52. The method of any one of claims 47 to 51, wherein the nucleic acid sequencing comprises one or more of shotgun metagenomic sequencing, rDNA sequencing, and targeted nucleic acid sequencing.

53. The method of claim 52, wherein the nucleic acid sequencing comprises 16S rDNA sequencing and / or 18S rDNA sequencing and / or ITS sequencing; and one or more of shotgun sequencing and targeted nucleic acid sequencing.

54. The method of any one of claims 47 to 53, wherein the genetic elements in the sample are amplified, optionally by PCR.

55. The method of any one of claims 47 to 54, wherein genetic elements are assigned to a reference genome for taxonomic classification, and / or genetic elements are assigned to gene functions.

56. The method of any one of claims 47 to 55, wherein microbial taxonomic features and microbial gene functional features that have differential abundance or differential detection rate in the samples of CRA subjects compared with control subjects are selected.

57. The method of claim 56, wherein the features comprise at least five taxonomic and / or gene functional features, and the taxonomic and / or gene functional features are optionally listed in Table 3.

58. The method of claim 57, wherein the features comprise at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or gene functional features, which are optionally listed in Table 3.

59. The method of claim 57 or 58, wherein the features include at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 taxonomic features optionally listed in Table 3; and at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 gene functional features optionally listed in Table 3.

60. The method of any one of claims 47 to 55, wherein microbial taxonomic features and microbial gene functional features that have differential abundance or differential detection rate in the samples of CRAA subjects compared with control subjects are selected.

61. The method of claim 60, wherein the features comprise at least five taxonomic and / or gene functional features, and the taxonomic and / or gene functional features are optionally listed in Table 4.

62. The method of claim 61, wherein the features comprise at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or gene functional features, and the taxonomic or gene functional features are optionally listed in Table 4.

63. The method of claim 61 or 62, wherein the features include at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 taxonomic features optionally listed in Table 4; and at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 gene functional features optionally listed in Table 4.

64. The method of any one of claims 47 to 55, wherein microbial taxonomic features and microbial gene functional features that have differential abundance or differential detection rate in the samples of CRC subjects compared with control subjects are selected.

65. The method of claim 64, wherein the features comprise at least five taxonomic and / or gene functional features, and the taxonomic and / or gene functional features are optionally listed in Table 5.

66. The method of claim 65, wherein the features comprise at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or gene functional features, and the taxonomic or gene functional features are optionally listed in Table 5.

67. The method of claim 65 or 66, wherein the features include at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 taxonomic features optionally listed in Table 5; and at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 gene functional features optionally listed in Table 5.

68. The method of any one of claims 47 to 67, wherein at least three gene signature profiles are trained, said gene signature profiles being: (1) Classify the samples based on the presence or absence of CRA. (2) Classify the samples based on the presence or absence of CRAA, and (3) Classify the samples based on the presence or absence of CRC.

69. The method of claim 68, wherein: (a) Using machine learning, a feature spectrum is trained to classify samples based on the presence or absence of CRA using samples from the CRA cohort and samples from the control cohort. (b) Using machine learning, a feature spectrum is trained to classify samples based on the presence or absence of CRC, using samples from the CRC cohort and samples from the control cohort. and (c) Using machine learning, a feature spectrum is trained to classify samples based on the presence or absence of CRAA, using samples from the CRAA cohort and samples from the control cohort.

70. The method of claim 69, wherein multiple machine learning algorithms are used to train the feature spectrum.

71. The method of claim 69 or 70, wherein at least one machine learning algorithm is supervised machine learning.

72. The method of claim 71, wherein the machine learning algorithm further comprises one or more of unsupervised or semi-supervised machine learning.

73. The method of any one of claims 69 to 72, wherein the machine learning comprises one or more of parametric / nonparametric distance measurement, logistic regression, support vector machine, decision tree, random forest, neural network, probabilistic regression, Fisher linear discriminant, Naive Bayes classifier, perceptron, quadratic classifier, kernel estimation, k-nearest neighbor, learned vector quantization, and principal component analysis.

74. The method of any one of claims 71 to 73, wherein the machine learning includes comparative machine learning modeling, optimization, testing, evaluation, and model ranking, including one or more of deep learning, gradient boosting, neural networks, ensemble or hybrid modeling algorithms, including gradient boosting tree classifiers, extreme gradient boosting tree classifiers, light gradient boosting tree classifiers, elastic network predictive light gradient boosting, Keras Slim residual neural network classifiers, generalized additive models, elastic network classifiers, random forest classifiers, deep forest classifiers, average mixer classifiers, TensorFlow multilayer perceptron classifiers, TensorFlow neural network classifiers, and rule fitting classifiers.

75. The method of any one of claims 69 to 74, wherein the feature spectrum comprises features selected from the training queue by an integrated ranking of feature importance and the statistical significance of individual features.

76. The method of claim 75, wherein the feature spectrum comprises features selected from the training cohort via Feature Importance Ranking Integration (FIRE) and Statistical Inference of Associations Between Microbial Communities and Phenotypes (SIAMCAT).

77. The method of any one of claims 47 to 76, wherein the characteristic spectrum has a sensitivity for classifying a sample based on the presence or absence of CRA, CRAA, or CRC, said sensitivity being at least about 0.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about 0.

95.

78. The method of any one of claims 47 to 77, wherein the characteristic spectrum has specificity for classifying a sample based on the presence or absence of CRA, CRAA, or CRC, said specificity being at least about 0.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about 0.

95.

79. A method for preparing a genetic profile of genetic elements indicative of colonic diseases, the method comprising: Provide training cohorts of biological samples, or DNA isolated from, subjects confirmed to have colonic disease and control subjects; The DNA isolated from the sample was subjected to genomic nucleic acid sequencing; The training classifies samples into gene feature profiles for the presence of (1) the colonic disease and (2) the absence of the colonic disease, wherein the gene feature profiles include features selected from the training queue by an integrated ranking of feature importance and the statistical significance of individual features.

80. The method of claim 79, wherein the biological sample is selected from feces, blood, serum, plasma, urine, saliva, biopsy tissue, mucosal tissue samples or swabs, and intestinal lavage fluid or aspirate.

81. The method of claim 80, wherein the sample is a fecal sample or a mucosal tissue sample.

82. The method of any one of claims 79 to 81, wherein the feature spectrum comprises features selected from the training cohort via Feature Importance Ranking Integration (FIRE) and Statistical Inference of Associations Between Microbial Communities and Phenotypes (SIAMCAT).

83. The method of any one of claims 79 to 82, wherein the colonic disease is selected from Crohn's disease, ulcerative colitis, irritable bowel syndrome (IBS), diverticulitis, colorectal adenoma (CRA), advanced colorectal adenoma (CRAA), and colorectal cancer (CRC).

84. The method of claims 79 to 83, wherein the features include microbial taxonomic features and microbial gene functional features.

85. The method of any one of claims 79 to 84, wherein the nucleic acid sequencing involves sequencing at least about 20,000,000 reads per sample.

86. The method of claim 85, wherein the nucleic acid sequencing involves sequencing at least about 40,000,000 reads, or at least about 50,000,000 reads, or at least about 60,000,000 reads, or at least about 75,000,000 reads, or at least about 100,000,000 reads, or at least about 150,000,000 reads per sample.

87. The method of any one of claims 79 to 86, wherein the nucleic acid sequencing comprises one or more of shotgun metagenomic sequencing, 16S rDNA sequencing, and targeted nucleic acid sequencing.

88. The method of claim 87, wherein the nucleic acid sequencing comprises one or more of 16S rDNA sequencing, shotgun sequencing, and targeted nucleic acid sequencing.

89. The method of any one of claims 79 to 88, wherein the genetic elements in the sample are amplified, optionally by PCR.

90. The method of any one of claims 79 to 89, wherein genetic elements are assigned to a reference genome for taxonomic classification, and / or genetic elements are assigned to gene functions.

91. The method of any one of claims 79 to 90, wherein microbial taxonomic features and microbial gene functional features that have differential abundance or differential detection rate in samples from subjects with colonic diseases compared with control subjects are selected.

92. The method of claim 91, wherein the features comprise at least five taxonomic and / or gene functional features.

93. The method of claim 92, wherein the features include at least about 10, at least about 25, at least about 50, or at least about 100 taxonomic or genetic functional features.

94. The method of claim 92 or 93, wherein the features include at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 taxonomic features; and at least one, at least two, at least five, at least about 10, at least about 20, or at least about 50 gene functional features.

95. The method of any one of claims 79 to 94, wherein multiple machine learning algorithms are used to train the feature spectrum.

96. The method of claim 95, wherein at least one machine learning algorithm is supervised machine learning.

97. The method of claim 96, wherein the machine learning algorithm further comprises one or more of unsupervised or semi-supervised machine learning.

98. The method of any one of claims 95 to 97, wherein the machine learning comprises one or more of parametric / nonparametric distance measurement, logistic regression, support vector machine, decision tree, random forest, neural network, probabilistic regression, Fisher linear discriminant, Naive Bayes classifier, perceptron, quadratic classifier, kernel estimation, k-nearest neighbor, learned vector quantization, and principal component analysis.

99. The method of any one of claims 96 to 98, wherein the machine learning includes comparative machine learning modeling, optimization, testing, evaluation, and model ranking, including one or more of deep learning, gradient boosting, neural networks, ensemble or hybrid modeling algorithms, including gradient boosting tree classifiers, extreme gradient boosting tree classifiers, light gradient boosting tree classifiers, elastic network predictive light gradient boosting, Keras Slim residual neural network classifiers, generalized additive models, elastic network classifiers, random forest classifiers, deep forest classifiers, average mixer classifiers, TensorFlow multilayer perceptron classifiers, TensorFlow neural network classifiers, and rule fitting classifiers.

100. The method of any one of claims 79 to 99, wherein the characteristic spectrum has a sensitivity for classifying the sample for the presence or absence of the colonic disease, the sensitivity being at least about 0.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about 0.

95.

101. The method of any one of claims 79 to 100, wherein the characteristic spectrum has a specificity for classifying the sample for the presence or absence of the colonic disease, the specificity being at least about 0.70, or at least about 0.75, or at least about 0.80, or at least about 0.90, or at least about 0.95.

Citation Information

Patent Citations

  • Multitag sequencing ecogenomics analysis-US

    US8603749B2

  • Multitag sequencing ecogenomics analysis-us

    US9453262B2