Stratification methods for assessing the progression and risk of advanced adenoma and colorectal cancer

WO2025083468A8PCT designated stage expired Publication Date: 2026-05-07MAINZ BIOMED NV +3
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
MAINZ BIOMED NV
Filing Date
2024-10-16
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Current screening methods for colorectal cancer (CRC) and advanced adenomas (AA) are either invasive and uncomfortable, such as colonoscopy, or have low sensitivity and specificity, such as fecal occult blood testing (FOBT). There is a need for non-invasive methods that can improve the detection of CRC and AA with increased sensitivity and specificity.

Method used

A computer-implemented method that transforms stool sample data into a diagnostic code by determining the expression levels of specific genes in the stool sample using a machine-learning model trained with labeled diagnostic data. This method includes providing a stool sample, determining the mRNA expression levels of distinct genes, and using a computing system to generate a diagnostic code based on the sample data.

Benefits of technology

The method provides a non-invasive means to assess the risk of CRC and AA with improved sensitivity and specificity, potentially leading to earlier detection and more effective management of these conditions.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The present disclosure concerns a computer-implemented method for transforming stool sample data into a diagnostic code, the method including: a) providing a stool sample from a subject; b) determining, as the stool sample data, expression levels of at least two distinct genes from a plurality of mRNA segments in the stool sample; and c) determining, by a computing system executing a machine-learning model and based on the stool sample data, a diagnostic code, wherein the machine-learning model was trained using stool diagnostic training data including a plurality of stool sample data labeled with at least one of two or more corresponding diagnostic codes.
Need to check novelty before this filing date? Find Prior Art

Description

STRATIFICATION METHODS FOR ASSESSING THE PROGRESSION AND RISK OF ADVANCED ADENOMA AND COLORECTAL CANCERBACKGROUND

[0001] Colorectal cancer (CRC) is one of the few cancer types for which screening has been proven to reduce cancer mortality in average-risk individuals. Indeed, the spread of the disease in terms of local invasion as well as to lymph nodes and distant organs at the time of the diagnosis is an important prognostic factor, with five-year survival rates of more than 90% for individuals with localized lesions but only 10% for those having their CRC metastasized to distal organs. Early detection is thus a key factor in reducing mortality from CRC. Advanced adenomas (AA) are also important to detect since they are considered to be the precursors of CRC while non-advanced adenomas (< 1 cm without advanced histology) may not be associated with increased colorectal cancer risk. Several screening regiments for CRC and AA are recommended such as fecal occult blood testing and colonoscopy. While colonoscopy remains the gold standard for the detection of colorectal lesions (up to 95% sensitivity for CRC and 76% for AA), compliance is not optimal owing to discomfort and unpleasant preparation procedures. The risk of complications, cost and access are other limitations of this procedure. On the other hand, the improved immunological version of fecal occult blood testing also referred to as the fecal immunochemical test (FIT), which detects human hemoglobin in stool samples, has been used for some time with some success but poor precursor lesion detection rates (66-80% sensitivity for CRC but only 10-28% for AA) albeit an excellent specificity (93-95%) limits its effectiveness. It is therefore imperative to explore alternate or complementary strategies with the potential to improve CRC screening performance, especially for the detection of cancers at their early stages and AA.

[0002] In this context, a number of initiatives have been undertaken over the last ten years, from stool testing as a non-invasive approach to the implementation of personalized CRC screening trying to meet with desirable features for a CRC screening test. Interestingly, many of the stool-based testing strategies are based on the high rate of tumor cell exfoliation into the colon-rectal lumen, a parameter that appears to be independent of blood release. An effective and well- documented strategy is the FDA-approved multitarget stool DNA test, an approach based on the detection of specific DNA aberrations from the CRC cells shed into the stool in combination with FIT, which results in an improvement of sensitivity for both CRC (92.3%) and AA (42.4%) detection compared toFIT alone, although achieved through a reduction of specificity to 87% thus generating almost three times more false positives. At first sight, the cost-benefit of such new methods for the medical system may temper screening recommendations but the high cost of CRC treatment, particularly for more advanced disease, is considered to improve the costeffectiveness of CRC screening. Furthermore, higher threshold costs for a biomarker test that could significantly increase the sensitivity of AA detection while maintaining reasonable specificity, would likely be cost-effective relative to currently available non- invasive tests.

[0003] Still based on the significant exfoliation of dysplastic cells from colorectal lesions into the lumen, host mRNA has also been investigated in the stools as a potential biomarker. While isolated from purified exfoliated colonocytes or directly extracted from the stools, host mRNA has been found to be a reliable source of biomarkers for detecting colorectal cancers. It was previously confirmed that the target mRNAs originated from the tumor or surrounding mucosa and that the number of detected gene transcripts was affected by the number of exfoliated tumor cells, exfoliation of inflammatory cells, tumor size and transcript expression level in the tumor but not primary vs distal location. More recently, it has been demonstrated that the inclusion of a multi-target RNA assay significantly strengthens both sensitivity and specificity for CRC detection. Droplet digital PCR was also evaluated as a potential alternative to qPCR for stool mRNA multiplex analysis. However, one important question that remains to be tested for the validation of a multi-target stool mRNA test pertains to AA detection since only ITGA6 has thus far been found to be overrepresented in stool samples of patients bearing AA. Another aspect that needs to be evaluated for potential clinical implementation is the robustness of the test under realistic preservation conditions, as mRNA are considered to be relatively susceptible to degradation in the stools.

[0004] It would be desirable to be provided with a non-invasive method to identify subjects who have an increased risk of having an advanced adenoma or a colorectal cancer with increased sensitivity and / or specificity.SUMMARY

[0005] The present disclosure relates to non-invasive methods for assessing the risk of a subject of having an advanced adenoma or a colorectal cancer based on stool sample data, including the mRNA levels of one or more genes present in the subject's stool sample.

[0006] The present disclosure provides a computer-implemented method for transforming stool sample data from a subject into a diagnostic code by executing a machine-learning model that is trained using stool diagnostic training data including a plurality of stool sample data labeled with at least one of two or more corresponding diagnostic codes.

[0007] According to a first aspect, the present disclosure provides a computer- implemented method for transforming stool sample data into a diagnostic code, the method including a) providing a stool sample from a subject; b) determining, as the stool sample data, expression levels of at least two distinct genes from a plurality of mRNA segments in the stool sample; and c) determining, by a computing system executing a machinelearning model and based on the stool sample data, a diagnostic code, wherein the machine-learning model was trained using stool diagnostic training data including a plurality of stool sample data labeled with at least one of two or more corresponding diagnostic codes.

[0008] In a second aspect, the disclosure provides a computer-implemented method for transforming stool sample data, the method including a) providing a stool sample from a subject, in which the stool sample includes a plurality of mRNA transcripts; b) determining, as the stool sample data, expression levels of at least two distinct genes from a plurality of mRNA transcripts in the stool sample; c) providing a representation of a graphical user interface for display on a computing system; d) receiving an input from the graphical user interface, the input comprising the stool sample data; e) determining, by the computing system executing a machine-learning model and based on the stool sample data, a diagnostic code, wherein the machine-learning model was trained using stool diagnostic training data including a plurality of stool sample data labeled with at least one of two or more corresponding diagnostic codes; and f) providing, for display on the graphical user interface, the diagnostic code.

[0009] In embodiments, the stool sample data and stool diagnostic training data further include subject age, subject gender, or both age and gender. In embodiments, the stool sample data and stool diagnostic training data further include the presence / absence (binary) or abundance (e.g., concentration) of hemoglobin in the stool (occult blood). In embodiments, the stool sample data and stool diagnostic training data include at least one colorectal epithelial cell, which, in embodiments, includes the plurality of mRNA transcripts.

[0010] Regarding the diagnostic code, in embodiments the diagnostic code includes normal (i.e., healthy; no indication of non-advanced adenomas (NA), advanced adenoma (AA), or colorectal cancer (CRC)), NA, AA, CRC, or AA / CRC, where AA / CRC indicates a subject can have either AA, CRC, or both. In embodiments, subsets of diagnostic codes are contemplated, e.g., a) normal, NA, AA, CRC, and AA / CRC; or b) normal, NA, AA, and CRC; or c) normal, NA, and AA / CRC; or d) normal, AA, CRC, and AA / CRC; or e) normal, AA, and CRC; or f) normal and AA / CRC.

[0011] Some embodiments of the aspects described above further include a step of prescribing a colonoscopy for the subject if the diagnostic code determined for the subject’s stool sample data is AA, CRC, or AA / CRC, and embodiments further include a step of referring, prescribing, or submitting the subject to a chemotherapy, a radiotherapy and / or a surgery if the diagnostic code determined for the subject’s stool sample data is AA, CRC, or AA / CRC.

[0012] According to a third aspect, the present disclosure provides a kit for transforming stool sample data, wherein the kit includes at least two reagents for determining a mRNA expression level of at least two distinct genes from the plurality of mRNA transcripts to obtain a diagnostic code in a stool sample from the subject. In some embodiments, the kit further comprises a container for storing a stool sample. In additional embodiments, the at least two reagents are for determining the mRNA expression level of the S100A4 gene, the PTGS2 gene, the GADD45B gene, the ITGA6 gene, the MYBL2 gene, the MYC gene, the CEACAM5 gene, and / or the MACC1 gene. In some additional embodiments, the kit further comprises at least one additional reagent for determining the mRNA expression level of the PTGS2 gene and / or of the ITGA6 gene. In yet additional embodiments, the at least two reagents are for determining the mRNA expression level of the CEACAM5 gene, the ITGA6 gene and / or the MACC1 gene. In still other embodiments, the at least two reagents are for determining the mRNA expression level of the S100A4 gene, the GADD45B gene, the ITGA6 gene, the MYBL2 gene, the MYC gene and / or the PTGS2 gene. In particular embodiments, at least two reagents for determining a mRNA expression level of at least two distinct genes from the plurality of mRNA transcripts are for determining the mRNA expression level of the CEACAM5, ITGA6, MACC1, PTGS2, and S100A4. In some embodiments, the kit further comprises a reverse-transcriptase. In still other embodiments, the kit can be used in combination or further means for determining the presence of hemoglobin in the stool sample. In still other embodiments, the kit can be used in combination or further includes a fecalimmunochemical test (FIT) to determine quantitatively the presence of hemoglobin in the stool sample. In yet further embodiments, the kit can be used in combination with, or further comprises, reagents for determining the presence of a DNA mutation and / or an aberrant DNA methylation pattern associated with a predisposition to a colorectal cancer in the colorectal epithelial cell of the subject. In some embodiments, the at least one DNA mutation is located in the KRAS gene, the BRAF gene, or both KRAS and BRAF genes. In additional embodiments, the aberrant DNA methylation pattern is located in the NDRG4 gene and / or the BMP3 gene. In still some further embodiments, the kit can be used in combination or further comprises a Cologuard™ and / or a Colo Alert™ assay to determine the presence of hemoglobin in the stool sample, the presence of DNA mutation and / or the presence of the abnormal DNA methylation pattern.

[0013] The present disclosure also provides a method for stratifying subjects with respect to their relative risk of having an advanced adenoma or a colorectal cancer. The method is based on the differential expression of certain genes of colorectal epithelial cells, which are present in the stool of the subject and measured by the concentrations of their transcripts. The method is also based on the relative stability of mRNA transcripts in the stool.

[0014] According to a fourth aspect, the present disclosure concerns a method of stratifying the risk of a subject of having an advanced adenoma or a colorectal cancer in a subject. The method comprises a) providing a stool sample from the subject, wherein the stool sample comprises a plurality of mRNA transcripts from the subject. The method also comprises b) determining the mRNA expression level of at least two distinct genes from the plurality of mRNA transcripts to obtain a test expression profile. The method further comprises c) comparing the test expression profile(s) with a control expression profile, wherein the control expression profile comprises the mRNA expression level of the at least two genes and is derived from a plurality of control mRNA transcripts (which can, in some embodiments, be derived from a control colorectal epithelial cell) from a control subj ect known to lack the advanced adenoma or the colorectal cancer. In some embodiments, the stool sample comprises at least one colorectal epithelial cell. In additional embodiments, the at least one colorectal epithelial cell comprises the plurality of mRNA transcripts. In yet additional embodiments, the test expression profile and the control expression profile comprise the mRNA expression level of at least two of the S100A4 gene, the PTGS2 gene, the GADD45B gene, the ITGA6 gene, the MYBL2 gene, the MYC gene, the CEACAM5 gene, and / or the MACC1 gene. In some specificembodiments, the present disclosure provides a method of stratifying the risk of a subject of having an advanced adenoma or a colorectal cancer in a subject, wherein the method comprises a) providing a stool sample from the subject, wherein the stool sample comprises at least one colorectal epithelial cell from the subject, b) determining the mRNA expression level of at least two distinct genes from the at least one colorectal epithelial cell to obtain a test expression profile, wherein the test expression profile comprises the mRNA expression level of at least two of the S100A4 gene, the PTGS2 gene, the GADD45B gene, the ITGA6 gene, the MYBL2 gene, the MYC gene, the CEACAM5 gene, and / or the MACC1 gene and c) comparing the test expression profile with a control expression profile, wherein the control expression profile comprises the mRNA expression level of the at least two genes and is derived from a control colorectal epithelial cell from a control subject known to lack the advanced adenoma or the colorectal cancer. In an embodiment, step b) comprises determining the mRNA expression level from at least one additional gene from plurality of mRNA transcripts of the colorectal epithelial cell, wherein the test expression profile and the control expression profile further comprises the expression level of the PTGS2 gene and / or of the ITGA6 gene. The method described herein can be used for stratifying the risk of the subj ect of having the advanced adenoma. In such embodiment, the test expression profile and the control expression profile can comprise the mRNA expression level of the CEACAM5 gene, the ITGA6 gene, the MACC1 gene, and / or the B2M gene. Alternatively, or in combination, the method described herein can be used for stratifying the risk of the subject of having the colorectal cancer. In such embodiments, the test expression profile and the control expression profile can comprise the mRNA expression level of two or more of the S100A4 gene, the GADD45B gene, the ITGA6 gene, the MYBL2 gene, the MYC gene and / or the PTGS2 gene. In some embodiments, the test expression profile for colorectal cancer can also include one or more of CEACAM5, MACC1 and ITGA6. In an embodiment, step b) comprises using a reversetranscriptase polymerase chain reaction (RT-PCR) to obtain the mRNA expression level of the at least two genes of the test expression profile and / or the control expression profile. In yet another embodiment, step b) comprises using a quantitative polymerase chain reaction (qPCR) to obtain the mRNA expression level of the at least two gens of the test expression profile and / or the control expression profile. In some embodiments, the further comprises, prior to step b), storing the stool sample. In additional embodiment, the method further comprises determining the presence of hemoglobin in the stool sample. In some specific embodiments, the method comprises using a fecal immunochemical test (FIT) to determine the quantitative presence of hemoglobin in the stool sample. In furtherembodiments, the method further comprises determining the presence of a DNA mutation and / or an aberrant DNA methylation pattern associated with a predisposition to a colorectal cancer in the colorectal epithelial cell of the subject. For example, the DNA mutation can be located in the KRAS gene, or the BRAF gene, or both KRAS and BRAF genes. In another example, the aberrant DNA methylation pattern can be located in the NDRG4 gene and / or the BMP3 gene. In some embodiments, the method comprising using the Cologuard™ or ColoAlert™ assay to determine the presence of hemoglobin in the stool sample, the presence of the DNA mutation and / or the presence of the aberrant DNA methylation pattern. In some embodiments, the method is for screening for subjects suitable for colonoscopy. In some embodiments, the method further comprises submitting the subject having been stratified as being at increased risk of developing the colorectal cancer to a chemotherapy, a radiotherapy and / or a surgery. In some embodiments, the colorectal cancer is a colon cancer or a rectal cancer. In some embodiments, the method further includes subject age and / or gender and / or the presence / absence (binary) or abundance (e.g., concentration) of hemoglobin in the stool (occult blood).

[0015] According to a fifth aspect, the present disclosure provides a kit for stratifying the risk of a subject of having an advanced adenoma or a colorectal cancer in a subject, wherein the kit comprises at least two reagents for determining the mRNA expression level of at least two distinct genes from the plurality of mRNA transcripts to obtain a test expression profile in a stool sample from the subject. In some embodiments, the kit further comprises a container for storing a stool sample. In additional embodiments, the at least two reagents are for determining the mRNA expression level of the S100A4 gene, the PTGS2 gene, the GADD45B gene, the ITGA6 gene, the MYBL2 gene, the MYC gene, the CEACAM5 gene, and / or the MACC1 gene. In some additional embodiments, the kit further comprises at least one additional reagent for determining the mRNA expression level of the PTGS2 gene and / or of the ITGA6 gene. In yet additional embodiments, the at least two reagents are for determining the mRNA expression level of the CEACAM5 gene, the ITGA6 gene and / or the MACC1 gene. In still other embodiments, the at least two reagents are for determining the mRNA expression level of the S100A4 gene, the GADD45B gene, the ITGA6 gene, the MYBL2 gene, the MYC gene and / or the PTGS2 gene. In particular embodiments, at least two reagents for determining a mRNA expression level of at least two distinct genes from the plurality of mRNA transcripts are for determining the mRNA expression level of the CEACAM5, ITGA6, MACC1, PTGS2, and S100A4. In some embodiments, the kit further comprises a reverse-transcriptase. Instill other embodiments, the kit can be used in combination or further means for determining the presence of hemoglobin in the stool sample. In still other embodiments, the kit can be used in combination or further comprises a fecal immunochemical test (FIT) to determine the presence of hemoglobin in the stool sample. In yet further embodiments, the kit can be used in combination or further comprises reagents for determining the presence of a DNA mutation and / or an aberrant DNA methylation pattern associated with a predisposition to a colorectal cancer in the colorectal epithelial cell of the subject. In some embodiments, the at least one DNA mutation is located in the KRAS gene, the BRAF gene, or both KRAS and BRAF genes. In additional embodiments, the aberrant DNA methylation pattern is located in the NDRG4 gene and / or the BMP3 gene. In still some further embodiments, the kit can be used in combination or further comprises a Cologuard™ and / or a ColoAlert™ assay to determine the presence of hemoglobin in the stool sample, the presence of DNA mutation and / or the presence of the abnormal DNA methylation pattern.

[0016] In embodiments of the various aspects described above and herein, mRNA expression level can be provided as an absolute amount or can be provided in a normalized amount, as described below.BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Having thus generally described the nature of the invention, reference will now be made to the accompanying drawings, showing by way of illustration, a preferred embodiment thereof, and in which:

[0018] FIG. 1A and FIG. IB illustrate the detection and analysis of selected mRNA targets found to be differentially in stool samples of patients with colorectal cancer (CRC) stages I-III or advanced adenomas (AA). Results in A and B are expressed as median (interquartile range) of copy number and score, respectively, relative to control patients.** P< 0.001 to*** P<0.0005 using the Kruskal- Wallis test. FIG. 1 A (left panel): For S100A4, a significant increase was observed in CRC stages I-III as compared to controls (Ctrl) or patients with AA as one of the six targets identified as being overrepresented in the stools of patients with CRC (right panel) For CEACAM5, a significant increase was observed in CRC stages I-III and AA as compared to controls (Ctrl) as one of the three targets identified as being overrepresented in the stools of patients with either AA or CRC. FIG. IB: Scores were calculated using an algorithm that combined all six targets for CRC (left) and the threetargets identifying AA and CRC (right) lesions relative to controls. FIG. 1C: Receiver operating characteristics (ROC) curve analysis showing the two groups of targets for CRC (left) and AA and CRC (right). Area under the curve (AUC) values are indicated.

[0019] FIG. 2A and FIG. 2B illustrate the ROC curve analysis of an optimized combination of five of the targets for the detection of patients with AA or CRC. AUC is indicated and sensitivity and specificity are provided in % (95% CI). FIG. 2A: ROC curve analysis of the combination of the three targets identified for detecting AA and CRC, CEAC AM5, ITGA6 and MACC 1 with the two stronger targets for detecting CRC, PTGS2 and S100A4, for AA and CRC. FIG. 2B: Same combination as in FIG. 2A but including the FIT component.

[0020] FIG. 3A and FIG. 3B provide the target stability analyses in stool samples over a 5-day period. Target stability was tested under various conditions of conservation and target detection was monitored throughout the 5 days in samples maintained at -20 °C with (FIT 5d -20) and without (5d -20 °C) a thaw cycle, at 4 °C (l-5d 4 °C) and at room temperature (l-5d RT). FIG. 3A: As illustrated with PTGS2, copy numbers remained relatively stable during the 5 days in both control stool samples (Ctrl) and samples obtained from CRC patients. FIG. 3B: Cumulative scores including the four tested targets PTGS2, CEACAM5, ITGA2 and ITGA6 showed that overall, the targets were relatively stable under cooled conditions and for 3 days at room temperature.

[0021] FIG. 4A to FIG. 4G provide the detection and analysis of selected mRNA targets found to be overrepresented in stool samples of patients with colorectal cancers (CRC) stages 1-111 or advanced adenomas (AA). As shown for S100A4 (FIG. 1A), a significant increase was observed for the five other targets GADD45B, ITGA2, MYBL2, MYC and PTGS2 in CRC stages 1111 as compared to controls (Ctrl) while for three of the targets, CEACAM5 (FIG. 1A) ITGA6 and MACC1, a significant increase was observed in samples from patients with CRC stages 1-111 or AA as compared to controls (Ctrl). Results are expressed as median (interquartile range) of copy number relative to control patients. * P< 0.05 to*** P<0.0005 using the Kruskal-Wallis test. FIG. 4A provides the copy number of GADD45B in function of the sample received. FIG. 4B provides the copy number of ITGA2 in function of the sample received. FIG. 4C provides the copy number of MYBL2 in function of the sample received. FIG. 4D provides the copy number of MYC in function of the sample received. FIG. 4E provides the copy number of PTGS2 in function of the sample received. FIG. 4F provides the copy number of ITGA6 in functionof the sample received. FIG. 4G provides the copy number of MACC1 in function of the sample received.

[0022] FIG. 5A to FIG. 5C provide additional information on target stability analyses in stool samples over a 5 -day period. Target stability was tested under various conditions of conservation as in FIG. 3 for CEACAM5, ITGA2 and ITGA6. Copy number (copy nb) were evaluated in the stool samples throughout the 5 days at -20 °C with (FIT 5d -20) and without (5d -20) a thaw cycle, at 4 °C (1-25 5d 4) and at room temperature (l-5d RT). FIG. 5A provides the copy number of ITGA6 in control (left panel) or CRC (right panel) samples in function of days in storage. FIG. 5B provides the copy number of CEACAM5 in control (left panel) or CRC (right panel) samples in function of days in storage. FIG. 5C provides the copy number of ITGA2 in control (left panel) or CRC (right panel) samples in function of days in storage.

[0023] FIG. 6A and FIG. 6B show amplification curves for a dilution series of known initial concentrations (standard curves; FIG. 6A) and for a set of patient samples (FIG. 6B).

[0024] FIG. 7A and FIG. 7B show amplification curves for a series of samples measured for mRNA marker CEACAM5 (FIG. 7A) and for a series of samples measured for housekeeping gene CTTN.

[0025] FIG. 8 shows a schematic illustrating the combined use of mRNA markers, housekeeping genes, absolute quantification and relative quantification as analytical steps in a method for identifying, diagnosing and / or stratifying the risk of a subject for having CRC or AA as described herein.

[0026] FIG. 9A to FIG. 9C show comparisons of the quantification of MACC1, PTGS2, and S100A4, respectively, both by absolute quantification and relative quantification as compared to the housekeeping gene GAPDH.

[0027] FIG. 10A and FIG. 10B are pie charts illustrating the distribution of pathological results for CRC (FIG. 10A) and AA (FIG. 10B).

[0028] FIG. 11 is an exemplary flow diagram of an artificial intelligence / machine learning algorithm developed to classify a (stool) sample as a predicted control (i.e., the subject is preliminarily predicted to not have AA or CRC) or predicted case (i.e., the subject ispreliminarily predicted to have AA or CRC).

[0029] FIG. 12 is an exemplary flow diagram of one output from an artificial intelligence / machine-leaming algorithm developed to sensitively and specifically detect and quantify colorectal cancer and advanced adenoma risk.

[0030] FIG. 13 shows a flowchart of a method for transforming stool sample data, according to an example implementation.DETAILED DESCRIPTION

[0031] The present disclosure provides a method for stratifying subjects with respect to their relative risk of having an advanced adenoma or a colorectal cancer by determining the expression levels of a plurality of genes in the subjects' stool sample. The subjects that can be stratified by the method can be mammals and, in some embodiments, humans. The subjects may or may not have been previously investigated for their predisposition to develop an advanced adenoma or a colorectal cancer. The subjects may or may not have been previously treated for an advanced adenoma or a colorectal cancer.

[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. In the case of conflict, the present document, including definitions will control.

[0033] As used herein, “around”, “about” or “approximately” shall generally mean within + / - 20 percent, and more preferably within + / -10 percent of a given value or range. Numerical quantities given herein are approximate, meaning that the term “around”, “about” or “approximately” can be inferred if not expressly stated.

[0034] The term “gene” as used herein refers to a locatable region of genomic sequence. In cells, a gene is a portion of DNA that contains both “coding” sequences that determine what the gene does, and “non-coding” sequences that determine when the gene is active (expressed). A gene is a union of genomic sequences encoding a coherent set of potentially overlapping functional products. The molecules resulting from gene expression, whether RNA or protein, are known as gene products.

[0035] The term “genetic marker” as used herein refers to alteration in DNA that may indicate the presence of, or an increased risk of developing, a specific disease or disorder, e.g.,colorectal cancer or advanced adenoma.

[0036] The term “gene expression” means the production of a protein or a functional RNA (e.g., mRNA) from its gene.

[0037] As used herein, a “housekeeping gene” refers to (typically) constitutive genes that are required for the maintenance of basal cellular functions that are essential for the existence of a cell, regardless of its specific role in the tissue or organism, or genes whose expression is otherwise consistent or substantially unaffected in the cell type interest and relevant conditions. Thus, they are generally expressed in all cells, or in at least the cell type of interest, of an organism under normal and pathophysiological conditions, irrespective of tissue type, developmental stage, cell cycle state, or external signal. As such, housekeeping genes can be used as internal controls for experimental studies. The reliability of any relative RT-PCR experiment can be improved by including an invariant endogenous control (reference gene) in the assay to correct for sample to sample variations in RT-PCR efficiency and errors in sample quantification. A biologically meaningful reporting of target mRNA expression level requires accurate and relevant normalization to some standard and is recommended in quantitative RT-PCR. Many housekeeping genes are known to those persons skilled in the art, such as for example, but not limited to, glyceraldehyde-3 -phosphate dehydrogenase (GAPDH; e.g., transcript variant 4, mRNA, NCBI Reference Sequence: NM_001289746.1); cortactin (CTTN; e.g., transcript variant 3, mRNA; NCBI Reference Sequence: NM_001184740.1); Ras-related protein Rab-7a (RAB7A; mRNA, NCBI Reference Sequence: NM_004637.5); 18S ribosomal RNA (RRN18S), RNA polymerase 2 subunit A (PolR2A); YWHAH (tyrosine 3- monooxygenase / tryptophan 5 -monooxygenase activation protein, eta polypeptide; NM- 003405); UBA3 (Ubiquitin-activating enzyme 3; NM-003968); RPS24 (ribosomal protein S24; NM-033022); RPL13 (ribosomal protein L13; NM-000977); and PGK1 (phosphoglycerate kinase 1; NM-000291).

[0038] The term “prognosis” means a forecasting of the probable course and outcome of a disease.

[0039] The term “primer” refers to a strand of nucleic acid complementary to a target nucleic acid sequence to be amplified, that serves as a starting point for DNA amplification.

[0040] The term “mRNA transcript” and “mRNA segment” are used interchangeablyherein and refer to messenger RNA, or fragment thereof, that has been transcribed from a DNA template.

[0041] The term “overexpressed” refers to a state wherein there exists any measurable increase over normal or baseline levels. For example, a molecule that is overexpressed in a disease is one that is manifested in a measurably higher level in the presence of the disease than in the absence of the disease.

[0042] The term “underexpressed” refers to a state wherein there exists any measurable decrease from normal or baseline levels. For example, a molecule that is underexpressed in a disease is one that is manifested in a measurably lower level in the presence of the disease than in the absence of the disease.

[0043] The term “alternatively expressed” refers to a state wherein there exists any measurable difference from normal or baseline levels. For example, a molecule that is alternatively expressed in a disease is one that is manifested in a measurably higher or lower level in the presence of the disease than in the absence of the disease. The term can be considered to encompass both underexpressed and overexpressed.

[0044] The term “normal sample” or “control” refers to a biological sample (e.g., stool) from a subject determined to be negative for colorectal cancer or advanced adenoma.

[0045] The term “biological sample” refers any sample obtained from a subject in which expression of a mRNA of interest can be detected. A biological sample can include blood, saliva, biopsy, skin tissue, liquid culture, stool and urine, but is not particularly limited thereto. In embodiments described herein, the biological sample is stool.

[0046] The term “iterative modeling process” or “iterative process” refers to a modeling method for creating one or more models to describe the relationship between predictors and outcomes in target datasets that includes a repeatable or loop-able subroutine or process (e.g. a run, a for loop, an epoch, a cycle).

[0047] The term ‘machine-learning model’ as used herein generally refers to a mathematical representation of a real-world process that is created by training a machinelearning algorithm using training data. Unless specified otherwise, the term machinelearning model as used herein may refer to a stool sample diagnostics machine-learning model.

[0048] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," and / or "including" when used in this specification, specify the presence of stated features, elements, and / or components, but do not preclude the presence or addition of one or more other features, elements, components, and / or groups thereof. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items. As used herein, phrases such as "between X and Y" and "between about X and Y" should be interpreted to include X and Y. As used herein, phrases such as "between about X and Y" mean "between about X and about Y." As used herein, phrases such as "from about X to Y" mean "from about X to about Y."

[0049] Moreover, the present disclosure also contemplates that in some embodiments, any feature or combination of features set forth herein can be excluded or omitted.

[0050] In the claims, as well as in the specification, all transitional phrases such as "comprising," "including," "carrying," "having," "containing," "involving," "holding," "composed of," or the like are to be understood to be open-ended, i.e., to mean including but not limited to. Only the transitional phrases "consisting of and "consisting essentially of shall be closed or semi-closed transitional phrases, respectively, as set forth in the United States Patent Office Manual of Patent Examining Procedures, Section 2111.03. "Consisting essentially of is to be interpreted as encompassing the recited materials or steps and those that do not materially affect the basic and novel character! stic(s) of the disclosure. The open-end phrases such as "comprising" include and encompass the close- ended phrases. Comprising may be amended to the more limiting phrases "consisting essentially of of "consisting of as needed.

[0051] All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The use of any and all examples, or exemplary language (e.g., "such as"), is intended for illustration and does not pose a limitation on the scope of the disclosure unless otherwise claimed. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the invention.

[0052] Furthermore, the disclosure encompasses all variations, combinations, andpermutations in which one or more limitations, elements, clauses, and descriptive terms from one or more of the listed claims are introduced into another claim. For example, any claim that is dependent on another claim can be modified to include one or more limitations found in any other claim that is dependent on the same base claim. Where elements are presented as lists, e.g., in Markush group format, each subgroup of the elements is also disclosed, and any element(s) can be removed from the group.

[0053] Broadly, the methods of the present disclosure allow for the analysis, transformation, and classification of data associated with a stool sample from a subject. The computer- implemented methods described herein utilizes a machine-learning model trained using “stool diagnostic training data” that is labeled with two or more corresponding diagnostic codes in order to determine a diagnostic code for the subject stool sample data, e.g., normal, NA, AA, CRC, or AA / CRC. Other methods herein allow for the stratification of subjects into at least two groups: a first group of subjects having an increased risk of having an AA or a CRC (e.g., high risk group) and a second group of subjects having a decreased risk of having an AA or a CRC (e.g., low risk group). In some embodiments of the stratification method, the method can also allow the stratification of the high risk group into two subgroups: a first subgroup of subjects having an increased risk of having AA (e.g., AA subgroup) and a second subgroup of subjects having an increased risk of having a CRC (e.g., CRC subgroup). The methods are based on the detection of increased mRNA transcript levels of at least two different genes present in the stool of the subjects. Subjects who have received a diagnostic code of AA, CRC, or AA / CRC or have been stratified in the high risk group, the AA subgroup or the CRC subgroup can receive tailored recommendations and treatments. For example, subj ects who have received a diagnostic code of AA, CRC, or AA / CRC, or who have been stratified in the high risk group, especially in the CRC subgroup, can receive a recommendation to perform a colonoscopy and / or be subject to a colonoscopy. In another example, subjects who have received a diagnostic code of AA, CRC, or AA / CRC, or who have been stratified in the high risk group, especially in the CRC subgroup, can receive a recommendation to receive a chemotherapy, a radiotherapy or to undergo surgery and / or receive the chemotherapy, the radiotherapy or be subject to surgery. Subjects who have received a diagnostic code of AA, CRC, or AA / CRC, or who have been stratified in the low risk group can receive tailored recommendations and treatments.

[0054] The methods described herein rely, at least in part, on assessing the expression level of a combination of genes in one or more cells from the subject and determining if such genes are overexpressed (or alternatively expressed) in the stool sample obtained from thesubj ects. The mRNA transcripts which are being submitted to these methods are present in a stool sample from the subject. It is understood that, in some embodiments, the mRNA transcripts can either be shed from cells of the colorectal epithelium and can be found in the stool sample in a cell-free manner. It also is understood that the mRNA transcripts can be present in one or more colorectal epithelial cell which is shed and present in the stool sample. The mRNA transcripts and / or the colorectal epithelial cell comprising same can be shed from an advanced adenoma(s) or a malignant epithelial tumor(s) that may be present in the subject. It has been surprisingly shown in the Example below that mRNA transcripts are stable in a stool sample and can conveniently be used to stratify the risk even though the stool sample had been previously stored.

[0055] In some embodiments of all methods described herein, in addition to assessing the expression level of a combination of genes in one or more cells from a subject (or in training with stool diagnostic training data), subject age, subject gender and / or the presence / absence (binary) or abundance (e.g., concentration) of hemoglobin in the stool (occult blood) of the subject is considered by the machine-learning model.

[0056] As a first step, the methods thus comprise providing a stool sample from the subj ect, wherein the stool sample comprises a plurality of mRNA transcripts from the colorectal epithelial cells from the subject. In one embodiment, the stool sample from the subject comprises at least one cell (or in some embodiments a plurality of cells) from the colorectal epithelium of the subject. In an embodiment, the cell is an epithelial cell. In still another embodiment, the cell is derived or shed from the colon's epithelium, e.g., the cell is a colon epithelial cell also referred to as a colonocyte. In yet another embodiment, the cell is derived or shed from the rectum's epithelium, e.g., the cell is a rectal epithelial cell. In still a further embodiment, the cell is derived or shed from the colon or the rectum, it is a colorectal epithelial cell. In an embodiment, the cell is an immune cell. In some embodiments, the method comprises obtaining the stool sample of the subject. In embodiments, the stool sample is about 5 grams, or about 4 grams, or about 3 grams, or about 2.5 grams, or about 2 grams, or about 1.5 grams, or about 1 grams, or about 750 mg, or about 500 mg, or about 250 mg, or about 100 mg, or about 75 mg or about 50 mg, or about 25 mg, or about 10 mg, or about 7.5 mg, or about 5 mg, or about 2.5 mg, or about 1 mg, or about 750 pg, or about 500 pg, or about 250 pg, or about 100 pg.

[0057] In some embodiments, the method can be performed directly on the stool sample which has been obtained from the subject. In other embodiments, the method can beperformed on a stool sample which has been processed. For example, the method can be performed on a stool sample which has been diluted with an appropriate solution (which can, in some embodiments, include RNase inhibitors) and / or filtered. As such, the method can include, in some embodiments, diluting and / or filtering the stool sample.

[0058] In yet another example, the stool sample or the processed stool sample can be stored prior to the next (e.g., determining) step. The stool sample or the processed stool sample can be stored at freezing temperatures (e.g., between -25 °C and -15 °C, in some embodiments at -18 °C, or between about -25 °C and -100 °C, or about -80 °C), at refrigerating temperatures (e.g., between 0 °C and 10 °C, in some embodiments at 4 °C) and / or at room temperatures (e.g., between 20 °C and 30 °C, in some embodiments at 23 °C). As such, the method can include storing the stool sample or the processed stool sample after it has been obtained or processed and before it is being further characterized. The stool sample or the processed stool sample can be stored for at least 1, 2, 3, 4, 5 days or more prior to the determination of the mRNA expression levels. In some embodiments, the method can include storing the stool sample or the processed stool sample prior to determining the mRNA expression levels.

[0059] Once the stool sample (which may have been processed and / or stored) has been obtained, the expression level of at least two distinct genes from the one or more cells present in the stool sample is determined. The expression level of the at least two distinct genes can be obtained by determining, e.g., the relative amount or the absolute amount of the mRNA being expressed from each gene. The determination of the expression level of the combination of genes can be made simultaneously (in a multiplex format) or subsequently.

[0060] The determination of the expression level of the combination of genes can include the reverse transcription of the mRNA transcripts associated with each gene, the amplification of the cDNA molecules associated with each gene of the combination and / or the hybridization of an oligonucleotide (which may be a primer or a probe) to the mRNA transcripts / cDNA molecules associated with each gene of the combination. In embodiments of the methods in which the mRNA transcripts are being reverse-transcribed and amplified, their (relative) abundance can be determined by detecting a signal associated with the amplified nucleic acid molecules. In some embodiments, the method can include performing a reverse-transcription step to convert the mRNA transcripts into cDNA molecules. In some additional embodiments, the method can include performing apolymerase chain reaction (PCR) step to amplify the number of cDNA molecules. In yet further embodiments, the method can include performing a quantitative polymerase chain reaction (qPCR) step to quantify the number of cDNA molecules. In yet further embodiments, the method can include performing a digital polymerase chain reaction (dPCR) step to quantify the number of cDNA molecules. In further embodiments, the method includes competitive RT-PCR, and in yet further embodiments, the method includes quantitative RT-PCR (also “qRT-PCR”, “RT-qPCR”, or “real-time RT-PCR”). While these methods can quickly and efficiently amplify small quantities of RNA in a relatively short period of time, e.g., one hour or less, one or more of several variables need to be controlled for in gene-expression analysis. Examples include the amount of starting material, whether reliable extraction of like amounts of non-degraded RNA from each sample is achieved; consistent reverse transcriptase efficiency resulting in equal amounts of cDNA; adequate primer specificity; and whether inhibitors are present in samples. In order to account for these effects, results can be normalized to a factor independent of the target gene(s) mRNA being quantified. For example, mRNAs can be normalized against a specified number or cells, against a comparator mRNA transcript from a gene distinct from the target gene(s), or combination thereof. Normalization against a comparator mRNA can be conducted against mRNA transcripts from one or more genes that are expressed at a relatively constant level across different experimental conditions and cell types and whose expression is known not to be modulated in advanced adenoma or colorectal cancer cells. In some embodiments, such comparator mRNAs are from one or more housekeeping genes.

[0061] Thus, in some embodiments, the mRNA expression level can be provided as an absolute amount or can be provided in a normalized amount. In some embodiments in which the method provides a normalized amount, the method can further include determining the number of cells in which the mRNA expression level has been determined, the number of copies of a mRNA or cDNA generated therefrom per cell or per unit volume of processed biological sample, and / or determining the mRNA expression level of one or more housekeeping gene in the stool sample or the processed stool sample. In some embodiments, the mRNA expression levels can be provided as ratios of one another. For the purpose of determining the relative amount, a ACP value is calculated between gene of interest and housekeeping gene as described below.

[0062] In some embodiments, normalization includes simultaneously measuring theexpression of target gene(s) of interest and one or more housekeeping genes in each sample. The crossing point “Cp” (also cycle threshold “Ct”; take-off point “TOP”; or quantification cycle “Cq”) values are determined for both the target gene and the housekeeping gene. The relative expression of the target gene can then be calculated by subtracting the Cp value of the housekeeping gene from the Cp value of the target gene (ACp; or Ct housekeeping gene from Ct target gene for ACt, etc.). When calculated in this manner, a lower / smaller ACp value indicates lower expression of the target gene (and higher / larger indicates higher expression of the target gene). An inverse calculation can also be performed, in which the Cp value of the target gene can be subtracted from the Cp value of the housekeeping gene. When calculated in this manner, a lower / smaller ACp value indicates higher expression of the target gene. A ACp value can be further used to calculate fold changes or relative expression levels between different samples or experimental conditions, or iteratively transformed by one or more operations by a machine-learning model to determine a diagnostic code.

[0063] The determining step, be it the first determining step in the computer-implemented methods or the sole determining step in the stratification method, in its most basic form, provides expression levels of at least two distinct genes from a plurality of mRNA segments in the stool sample, i.e., the mRNA expression levels of the at least two genes whose expression has been quantified (stool sample data). In the case of the computer- implemented methods, the stool sample data are passed to the next step. In the stratification methods, the determining step provides a test expression profile which comprises the mRNA expression level of the at least two genes whose expression has been quantified. The stool sample data (SSD), or for stratification methods the test expression profile (TEP), can include the mRNA expression level of the CEA adhesion molecule 5 (also referred to as CEACAM5, CD66e or CEA and having the Gene ID 1048). The SSD or TEP can include the mRNA expression level of the growth arrest and DNA damage inducible beta gene (also referred to as GADD45B, GADD45BETA or MYD118 and having the Gene ID: 4616). The SSD or TEP can include the mRNA expression level of the integrin subunit alpha 2 gene (also referred to as ITGA2, BR, CD49B, GPla, HPA-5, VLA-2 or VLAA2 and having the Gene ID: 3673). The SSD or TEP can include the mRNA expression level of the MET transcriptional regulator MACC1 (also referred to as MACC1, 7A5 or SH3BP4L and having the Gene ID: 346389). The SSD or TEP can include the mRNA expression level of the MYB proto-oncogene like 2 gene (also referred to as MYBL2, B- MYB or BMYB and having the Gene ID: 4605). The SSD or TEP can include the mRNAexpression level of the MYC proto-oncogene, bHLH transcription factor (also referred to as MYC, MRTL, MYCC, bHLHe39 or c-Myc and having the Gene ID: 4609). The SSD or TEP can include the mRNA expression level of the SI 00 calcium binding protein A4 (also referred to as S100A4, 18A2, 42 A, CAPL, FSP1, MTS1, P9KA or PEL98 and having Gene ID: 6275). In some embodiments, the SSD or TEP can include the mRNA expression level of prostaglandin-endoperoxide synthase 2 (also referred to as PTGS2, COX2, COX-2, PHS-2, PGG / HS, PGHS-2, hCox-2, or GRIPGHS, and having Gene ID: 5743). In some embodiments, the SSD or TEP can include the mRNA expression level of the integrin subunit alpha 6 gene (also referred to as ITGA6, JEB6, CD49f, VLA-6, ITGA6A, or ITGA6B, and having Gene ID: 3655). In some further embodiments, the SSD or TEP can include the mRNA expression level of the beta-2-microglobulin gene (also referred to as B2M or IMD43 and having the Gene ID 567). In some further optional embodiments, the SSD or TEP can include the mRNA expression level of the integrin subunit alpha 1 gene (also referred to as ITGA1, CD49a or VLA1 and having the Gene ID: 3672). In a specific embodiment, the SSD or TEP can include the mRNA expression profile of CEACAM5, ITGA6 and MACC1. In yet another specific embodiment, the SSD or TEP can include the mRNA expression profile of CEACAM5, ITGA6, MACC1 and B2M. In still yet another embodiment, the SSD or TEP can include the mRNA expression profile of PTGS2 and S100A4. In still yet another embodiment, the SSD or TEP can include the mRNA expression profile of CEACAM5, ITGA6, MACC1, PTGS2 and S100A4, optionally in combination with the mRNA expression profile of B2M.

[0064] In some optional embodiments, the SSD or TEP can include the mRNA expression level of ITGA6. In such embodiments, it is possible that the SSD or TEP can include the mRNA expression level of the alpha- and / or beta-isoform of ITGA6 gene transcript. In some optional embodiments, the SSD or TEP can include the mRNA expression level ofPTGS2.

[0065] In some specific embodiments, the SSD or TEP can include at least one, two, three, four or five level of mRNA expression the S100A4 gene, the GADD45B gene, the ITGA6 gene, the MYBL2 gene, the MYC gene and / or the PTGS2 gene.

[0066] In some specific embodiments, the SSD or TEP can include at least one, two, three, four or five level of mRNA expression the CEACAM5 gene, the ITGA6 gene, the MACC 1 gene, the PTGS2 gene and / or the S100A4 gene.

[0067] In some embodiments, the SSD or TEP can include the mRNA expression level of B2M and CEAC AM5, or B2M and GADD45B, or B2M and ITGA6 A, or B2m and ITGA6, or B2M and MACC1, or B2M and MYBL2, or B2M and MYC, or B2M and PTGS2, or B2M and S100A4, or CEACAM5 and GADD45B, or CEACAM5 and ITGA6A, or CEACAM5 and ITGA6, or CEACAM5 and MACC1, or CEACAM5 and MYBL2, or CEACAM5 and MYC, or CEACAM5 and PTGS2, or CEACAM5 and S100A4, or GADD45B and ITGA6A, or GADD45B and ITGA6, or GADD45B and MACC1, or GADD45B and MYBL2, or GADD45B and MYC, or GADD45B and PTGS2, or GADD45B and S100A4, or ITGA6A and ITGA6, or ITGA6A and MACC1, or ITGA6A and MYBL2, or ITGA6A and MYC, or ITGA6A and PTGS2, or ITGA6A and S100A4, or ITGA6 and MACC1, or ITGA6 and MYBL2, or ITGA6 and MYC, or ITGA6 and PTGS2, or ITGA6 and S100A4, or MACC1 and MYBL2, or MACC1 and MYC, or MACC1 and PTGS2, or MACC1 and S100A4, or MYBL2 and MYC, or MYBL2 and PTGS2, or MYBL2 and S100A4, or MYC and PTGS2, or MYC and S100A4, orPTGS2 and S100A4.

[0068] In some embodiments, the SSD or TEP can include the mRNA expression level of one marker detectable in stool samples and enriched in samples with CRC, or PTGS2, S100A4, GADD45B, ITGA6A, MYBL2 or MYC.

[0069] In some embodiments, the SSD or TEP can include the mRNA expression level of two markers detectable in stool samples and enriched in samples with CRC (i.e., PTGS2, S100A4, GADD45B, ITGA6A, MYBL2 or MYC). Such embodiments are specifically described above.

[0070] In some embodiments, the SSD or TEP can include the mRNA expression level of three markers detectable in stool samples and enriched in samples with CRC. Such embodiments include PTGS2, S100A4 and GADD45B, or PTGS2, S100A4 and ITGA6A, or orPTGS2, S100A4 and MYBL2, orPTGS2, S100A4 and MYC, orPTGS2, GADD45B and ITGA6A, or PTGS2, GADD45B and MYBL2, or PTGS2, GADD45B and MYC, or PTGS2, ITGA6A and MYBL2, or PTGS2, ITGA6A and MYC, or PTGS2, MYBL2 and MYC, or S100A4, GADD45B and ITGA6A, or S100A4, GADD45B and MYBL2, or S100A4, GADD45B and MYC, or S100A4, ITGA6A and MYBL2, or S100A4, ITGA6A and MYC, or S100A4, MYBL2 and MYC, or GADD45B, ITGA6A and MYBL2, or GADD45B, ITGA6A and MYC, or ITGA6A, MYBL2 and MYC.

[0071] In some embodiments, the SSD or TEP can include the mRNA expression level offour markers detectable in stool samples and enriched in samples with CRC, including PTGS2, S100A4, GADD45B, ITGA6A, MYBL2 or MYC. Such embodiments include PTGS2, S100A4, GADD45B and ITGA6A, orPTGS2, S100A4, GADD45B and MYBL2, or PTGS2, S100A4, GADD45B and MYC, or PTGS2, S100A4, ITGA6A and MYBL2, or PTGS2, S100A4, ITGA6A and MYC, or PTGS2, S100A4, MYBL2 and MYC, or PTGS2, S100A4, ITGA6A and MYC, or S100A4, GADD45B, ITGA6A and MYBL2, or S100A4, GADD45B, MYBL2 and MYC, or S100A4, GADD45B, ITGA6A and MYC, or PTGS2, GADD45B, ITGA6A and MYBL2, or PTGS2, GADD45B, MYBL2 and MYC, or PTGS2, GADD45B, ITGA6A and MYC, or PTGS2, ITGA6A, MYBL2 and MYC, or S100A4, GADD45B, ITGA6A and MYBL2, or S100A4, GADD45B, ITGA6A and MYC, or GADD45B, ITGA6A, MYBL2 and MYC.

[0072] In some embodiments, the SSD or TEP can include the mRNA expression level of five markers detectable in stool samples and enriched in samples with CRC. Such embodiments include PTGS2, S100A4, GADD45B, ITGA6A and MYBL2, or PTGS2, S100A4, GADD45B, ITGA6A and MYC, or S100A4, GADD45B, ITGA6A, MYBL2 and MYC, or PTGS2, GADD45B, ITGA6A, MYBL2 and MYC, or PTGS2, S100A4, ITGA6A, MYBL2 and MYC, or PTGS2, S100A4, GADD45B, MYBL2 and MYC.

[0073] In some embodiments, the SSD or TEP can include the mRNA expression level of six markers detectable in stool samples and enriched in samples with CRC, including PTGS2, S100A4, GADD45B, ITGA6A, MYBL2 and MYC.

[0074] In some embodiments, the SSD or TEP can include the mRNA expression level of one marker detectable in stool samples and enriched in samples with CRC and AA, including CEACAM5, ITGA6, MACC1 or B2M.

[0075] In some embodiments, the SSD or TEP can include the mRNA expression level of two markers detectable in stool samples and enriched in samples with CRC and AA. Such embodiments include CEACAM5 and ITGA6, or CEACAM5 and MACC1, or CEAC AM5 and B2M, or ITGA6 and MACC 1 , or ITGA6 and B2M, or MACC 1 and B2M.

[0076] In some embodiments, the SSD or TEP can include the mRNA expression level of three markers detectable in stool samples and enriched in samples with CRC and AA. Such embodiments include CEACAM5, ITGA6 and MACC1, or CEACAM5, ITGA6 and B2M, or ITGA6, MACC1 and B2M, or CEACAM5, MACC1 and B2M.

[0077] In some embodiments, the SSD or TEP can include the mRNA expression level of four markers detectable in stool samples and enriched in samples with CRC and AA, including CEACAM5, ITGA6, MACC1 and B2M.

[0078] In some embodiments, the SSD or TEP can include the mRNA expression level of one marker detectable in stool samples and enriched in samples with CRC (PTGS2, S100A4, GADD45B, ITGA6A, MYBL2 or MYC) and one marker detectable in stool samples and enriched in samples with CRC and AA (CEACAM5, ITGA6, MACC1 or B2M). Such embodiments are specifically described above.

[0079] In some embodiments, the SSD or TEP can include the mRNA expression level of two markers detectable in stool samples and enriched in samples with CRC and one marker detectable in stool samples and enriched in samples with CRC and AA. Such embodiments can include PTGS2, S100A4 and CEACAM5, or PTGS2, GADD45B and CEACAM5, or PTGS2, ITGA6A and CEACAM5, or PTGS2, MYBL2 and CEACAM5, or PTGS2, MYC and CEACAM5, or S100A4, GADD45B and CEACAM5, or S100A4, ITGA6A and CEACAM5, or S100A4, MYBL2 and CEACAM5, or S100A4, MYC and CEACAM5, or GADD45B, ITGA6A and CEACAM5, or GADD45B, MYBL2 and CEACAM5, or GADD45B, MYC and CEACAM5, or ITGA6A, MYBL2 and CEACAM5, or ITGA6A, MYC and CEACAM5, or MYBL2, MYC and CEACAM5, or PTGS2, S100A4 and ITGA6, or PTGS2, GADD45B and ITGA6, or PTGS2, ITGA6A and ITGA6, or PTGS2, MYBL2 and ITGA6, or PTGS2, MYC and ITGA6, or S100A4, GADD45B and ITGA6, or S100A4, ITGA6A and ITGA6, or S100A4, MYBL2 and ITGA6, or S100A4, MYC and ITGA6, or GADD45B, ITGA6A and ITGA6, or GADD45B, MYBL2 and ITGA6, or GADD45B, MYC and ITGA6, or ITGA6A, MYBL2 and ITGA6, or ITGA6A, MYC and ITGA6, or MYBL2, MYC and ITGA6, or PTGS2, S100A4 and MACC1, or PTGS2, GADD45B and MACC1, or PTGS2, ITGA6A and MACC1, or PTGS2, MYBL2 and MACC1, or PTGS2, MYC and MACC1, or S100A4, GADD45B and MACC1, or S100A4, ITGA6A and MACC1, or S100A4, MYBL2 and MACC1, or S100A4, MYC and MACC1, or GADD45B, ITGA6A and MACC1, or GADD45B, MYBL2 and MACC1, or GADD45B, MYC and MACC1, or ITGA6A, MYBL2 and MACC1, or ITGA6A, MYC and MACC1, or MYBL2, MYC and MACC1, or PTGS2, S100A4 and B2M, or PTGS2, GADD45B and B2M, or PTGS2, ITGA6A and B2M, or PTGS2, MYBL2 and B2M, or PTGS2, MYC and B2M, or S100A4, GADD45B and B2M, or S100A4, ITGA6A and B2M, or S100A4, MYBL2 and B2M, or S100A4, MYC and B2M, or GADD45B, ITGA6A and B2M, or GADD45B, MYBL2 and B2M, orGADD45B, MYC and B2M, or ITGA6A, MYBL2 and B2M, or ITGA6A, MYC and B2M, or MYBL2, MYC and B2M.

[0080] In some embodiments, the SSD or TEP can include the mRNA expression level of three markers detectable in stool samples and enriched in samples with CRC and one marker detectable in stool samples and enriched in samples with CRC and AA. Such embodiments can include PTGS2, S100A4, GADD45B and CEACAM5, or PTGS2, S100A4, ITGA6A and CEACAM5, or PTGS2, S100A4, MYBL2 and CEACAM5, or PTGS2, S100A4, MYC and CEACAM5, or PTGS2, GADD45B, ITGA6A and CEACAM5, or PTGS2, GADD45B, MYBL2 and CEACAM5, or PTGS2, GADD45B, MYC and CEACAM5, or PTGS2, ITGA6A, MYBL2 and CEACAM5, or PTGS2, ITGA6A, MYC and CEACAM5, orPTGS2, MYBL2, MYC and CEACAM5, or S100A4, GADD45B, ITGA6A and CEACAM5, or S100A4, GADD45B, MYBL2 and CEACAM5, or S100A4, GADD45B, MYC and CEACAM5, or S100A4, ITGA6A, MYBL2 and CEACAM5, or S100A4, ITGA6A, MYC and CEACAM5, or S100A4, MYBL2, MYC and CEACAM5, or GADD45B, ITGA6A, MYBL2 and CEACAM5, or GADD45B, ITGA6A, MYC and CEACAM5, or ITGA6A, MYBL2, MYC and CEACAM5, or PTGS2, S100A4, GADD45B and ITGA6, or PTGS2, S100A4, ITGA6A and ITGA6, or PTGS2, S100A4, MYBL2 and ITGA6, or PTGS2, S100A4, MYC and ITGA6, or PTGS2, GADD45B, ITGA6A and ITGA6, or PTGS2, GADD45B, MYBL2 and ITGA6, or PTGS2, GADD45B, MYC and ITGA6, or PTGS2, ITGA6A, MYBL2 and ITGA6, or PTGS2, ITGA6A, MYC and ITGA6, or PTGS2, MYBL2, MYC and ITGA6, or S100A4, GADD45B, ITGA6A and ITGA6, or S100A4, GADD45B, MYBL2 and ITGA6, or S100A4, GADD45B, MYC and ITGA6, or S100A4, ITGA6A, MYBL2 and ITGA6, or S100A4, ITGA6A, MYC and ITGA6, or S100A4, MYBL2, MYC and ITGA6, or GADD45B, ITGA6A, MYBL2 and ITGA6, or GADD45B, ITGA6A, MYC and ITGA6, or ITGA6A, MYBL2, MYC and ITGA6, or PTGS2, S100A4, GADD45B and MACC1, or PTGS2, S100A4, ITGA6A and MACC1, or PTGS2, S100A4, MYBL2 and MACC1, or PTGS2, S100A4, MYC and MACC1, or PTGS2, GADD45B, ITGA6A and MACC1, orPTGS2, GADD45B, MYBL2 and MACCl, orPTGS2, GADD45B, MYC andMACCl, or PTGS2, ITGA6A, MYBL2 and MACC1, or PTGS2, ITGA6A, MYC and MACC1, or PTGS2, MYBL2, MYC and MACC1, or S100A4, GADD45B, ITGA6A and MACC1, or S100A4, GADD45B, MYBL2 and MACC1, or S100A4, GADD45B, MYC and MACC1, or S100A4, ITGA6A, MYBL2 and MACC1, or S100A4, ITGA6A, MYC and MACC1, or S100A4, MYBL2, MYC and MACC1, or GADD45B, ITGA6A, MYBL2 and MACC1,or GADD45B, ITGA6A, MYC and MACC1, or ITGA6A, MYBL2, MYC and MACC1, or PTGS2, S100A4, GADD45B and B2M, or PTGS2, S100A4, ITGA6A and B2M, or PTGS2, S100A4, MYBL2 and B2M, or PTGS2, S100A4, MYC and B2M, or PTGS2, GADD45B, ITGA6A and B2M, or PTGS2, GADD45B, MYBL2 and B2M, or PTGS2, GADD45B, MYC and B2M, or PTGS2, ITGA6A, MYBL2 and B2M, or PTGS2, ITGA6A, MYC and B2M, or PTGS2, MYBL2, MYC and B2M, or S100A4, GADD45B, ITGA6A and B2M, or S100A4, GADD45B, MYBL2 and B2M, or S100A4, GADD45B, MYC and B2M, or S100A4, ITGA6A, MYBL2 and B2M, or S100A4, ITGA6A, MYC and B2M, or S100A4, MYBL2, MYC and B2M, or GADD45B, ITGA6A, MYBL2 and B2M, or GADD45B, ITGA6A, MYC and B2M, or ITGA6A, MYBL2, MYC and B2M.

[0081] In some embodiments, the SSD or TEP can include the mRNA expression level of four markers detectable in stool samples and enriched in samples with CRC and one marker detectable in stool samples and enriched in samples with CRC and AA. Such embodiments can include PTGS2, S100A4, GADD45B, ITGA6A and CEACAM5, or PTGS2, S100A4, GADD45B, MYBL2 and CEACAM5, or PTGS2, S100A4, GADD45B, MYC and CEACAM5, or PTGS2, S100A4, ITGA6A, MYBL2 and CEACAM5, or PTGS2, S100A4, ITGA6A, MYC and CEACAM5, or PTGS2, S100A4, MYBL2, MYC and CEACAM5, or PTGS2, S100A4, ITGA6A, MYC and CEACAM5, or S100A4, GADD45B, ITGA6A, MYBL2 and CEACAM5, or S100A4, GADD45B, MYBL2, MYC and CEACAM5, or S100A4, GADD45B, ITGA6A, MYC and CEACAM5, or PTGS2, GADD45B, ITGA6A, MYBL2 and CEACAM5, or PTGS2, GADD45B, MYBL2, MYC and CEACAM5, or PTGS2, GADD45B, ITGA6A, MYC and CEACAM5, or PTGS2, ITGA6A, MYBL2, MYC and CEACAM5, or S100A4, GADD45B, ITGA6A, MYBL2 and CEACAM5, or S100A4, GADD45B, ITGA6A, MYC and CEACAM5, or GADD45B, ITGA6A, MYBL2, MYC and CEACAM5, or PTGS2, S100A4, GADD45B, ITGA6A and ITGA6, or PTGS2, S100A4, GADD45B, MYBL2 and ITGA6, or PTGS2, S100A4, GADD45B, MYC and ITGA6, or PTGS2, S100A4, ITGA6A, MYBL2 and ITGA6, or PTGS2, S100A4, ITGA6A, MYC and ITGA6, or PTGS2, S100A4, MYBL2, MYC and ITGA6, or PTGS2, S100A4, ITGA6A, MYC and ITGA6, or S100A4, GADD45B, ITGA6A, MYBL2 and ITGA6, or S100A4, GADD45B, MYBL2, MYC and ITGA6, or S100A4, GADD45B, ITGA6A, MYC and ITGA6, or PTGS2, GADD45B, ITGA6A, MYBL2 and ITGA6, or PTGS2, GADD45B, MYBL2, MYC and ITGA6, or PTGS2, GADD45B, ITGA6A, MYC and ITGA6, or PTGS2, ITGA6A, MYBL2, MYC and ITGA6, or S100A4, GADD45B, ITGA6A, MYBL2 and ITGA6, or S100A4,GADD45B, ITGA6A, MYC and ITGA6, or GADD45B, ITGA6A, MYBL2, MYC and ITGA6, or PTGS2, S100A4, GADD45B, ITGA6A and MACC1, or PTGS2, S100A4, GADD45B, MYBL2 and MACC1, or PTGS2, S100A4, GADD45B, MYC and MACC1, or PTGS2, S100A4, ITGA6A, MYBL2 and MACC1, orPTGS2, S100A4, ITGA6A, MYC and MACC1, or PTGS2, S100A4, MYBL2, MYC and MACC1, or PTGS2, S100A4, ITGA6A, MYC and MACC1, or S100A4, GADD45B, ITGA6A, MYBL2 and MACC1, or S100A4, GADD45B, MYBL2, MYC and MACC1, or S100A4, GADD45B, ITGA6A, MYC and MACC1, or PTGS2, GADD45B, ITGA6A, MYBL2 and MACC1, or PTGS2, GADD45B, MYBL2, MYC and MACC1, or PTGS2, GADD45B, ITGA6A, MYC and MACC1, or PTGS2, ITGA6A, MYBL2, MYC and MACC1, or S100A4, GADD45B, ITGA6A, MYBL2 and MACC1, or S100A4, GADD45B, ITGA6A, MYC and MACC1, or GADD45B, ITGA6A, MYBL2, MYC and MACC1, or PTGS2, S100A4, GADD45B, ITGA6A and B2M, or PTGS2, S100A4, GADD45B, MYBL2 and B2M, or PTGS2, S100A4, GADD45B, MYC and B2M, or PTGS2, S100A4, ITGA6A, MYBL2 and B2M, or PTGS2, S100A4, ITGA6A, MYC and B2M, or PTGS2, S100A4, MYBL2, MYC and B2M, or PTGS2, S100A4, ITGA6A, MYC and B2M, or S100A4, GADD45B, ITGA6A, MYBL2 and B2M, or S100A4, GADD45B, MYBL2, MYC and B2M, or S100A4, GADD45B, ITGA6A, MYC and B2M, or PTGS2, GADD45B, ITGA6A, MYBL2 and B2M, orPTGS2, GADD45B, MYBL2, MYC and B2M, orPTGS2, GADD45B, ITGA6A, MYC and B2M, or PTGS2, ITGA6A, MYBL2, MYC and B2M, or S100A4, GADD45B, ITGA6A, MYBL2 and B2M, or S100A4, GADD45B, ITGA6A, MYC and B2M, or GADD45B, ITGA6A, MYBL2, MYC and B2M.

[0082] In some embodiments, the SSD or TEP can include the mRNA expression level of five markers detectable in stool samples and enriched in samples with CRC and one marker detectable in stool samples and enriched in samples with CRC and AA. Such embodiments can include PTGS2, S100A4, GADD45B, ITGA6A, MYBL2 and CEACAM5, or PTGS2, S100A4, GADD45B, ITGA6A, MYC and CEACAM5, or S100A4, GADD45B, ITGA6A, MYBL2, MYC and CEACAM5, or PTGS2, GADD45B, ITGA6A, MYBL2, MYC and CEACAM5, orPTGS2, S100A4, ITGA6A, MYBL2, MYC and CEACAM5, or PTGS2, S100A4, GADD45B, MYBL2, MYC and CEACAM5, or PTGS2, S100A4, GADD45B, ITGA6A, MYBL2 and ITGA6, or PTGS2, S100A4, GADD45B, ITGA6A, MYC and ITGA6, or S100A4, GADD45B, ITGA6A, MYBL2, MYC and ITGA6, or PTGS2, GADD45B, ITGA6A, MYBL2, MYC and ITGA6, or PTGS2, S100A4, ITGA6A, MYBL2, MYC and ITGA6, orPTGS2, S100A4, GADD45B,MYBL2, MYC and ITGA6, or PTGS2, S100A4, GADD45B, ITGA6A, MYBL2 and MACC1, or PTGS2, S100A4, GADD45B, ITGA6A, MYC and MACC1, or S100A4, GADD45B, ITGA6A, MYBL2, MYC and MACC1, or PTGS2, GADD45B, ITGA6A, MYBL2, MYC and MACC1, or PTGS2, S100A4, ITGA6A, MYBL2, MYC and MACC1, or PTGS2, S100A4, GADD45B, MYBL2, MYC and MACC1, or PTGS2, S100A4, GADD45B, ITGA6A, MYBL2 and B2M, or PTGS2, S100A4, GADD45B, ITGA6A, MYC and B2M, or S100A4, GADD45B, ITGA6A, MYBL2, MYC and B2M, or PTGS2, GADD45B, ITGA6A, MYBL2, MYC and B2M, or PTGS2, S100A4, ITGA6A, MYBL2, MYC and B2M, or PTGS2, S100A4, GADD45B, MYBL2, MYC and B2M.

[0083] In some embodiments, the SSD or TEP can include the mRNA expression level of six markers detectable in stool samples and enriched in samples with CRC and one marker detectable in stool samples and enriched in samples with CRC and AA. Such embodiments can include PTGS2, S100A4, GADD45B, ITGA6A, MYBL2, MYC and CEACAM5, or PTGS2, S100A4, GADD45B, ITGA6A, MYBL2, MYC and ITGA6, or PTGS2, S100A4, GADD45B, ITGA6A, MYBL2, MYC and MACC1, or PTGS2, S100A4, GADD45B, ITGA6A, MYBL2, MYC and B2M.

[0084] In some embodiments, the SSD or TEP can include the mRNA expression level of one marker detectable in stool samples and enriched in samples with CRC and two markers detectable in stool samples and enriched in samples with CRC and AA. Such embodiments can include PTGS2, CEACAM5 and ITGA6, or S100A4, CEACAM5 and ITGA6, or GADD45B, CEACAM5 and ITGA6, or ITGA6A, CEACAM5 and ITGA6, or MYBL2, CEACAM5 and ITGA6, or MYC, CEACAM5 and ITGA6, or PTGS2, CEACAM5 and MACC1, or S100A4, CEACAM5 and MACC1, or GADD45B, CEACAM5 andMACCl, orITGA6A, CEACAM5 and MACCl, orMYBL2, CEACAM5 and MACC1, or MYC, CEACAM5 and MACC1, or PTGS2, CEACAM5 and B2M, or S100A4, CEACAM5 and B2M, or GADD45B, CEACAM5 and B2M, or ITGA6A, CEACAM5 and B2M, orMYBL2, CEACAM5 andB2M, or MYC, CEACAM5 and B2M, or PTGS2, ITGA6 and MACC1, or S100A4, ITGA6 and MACC1, or GADD45B, ITGA6 and MACC1, or ITGA6A, ITGA6 and MACC1, or MYBL2, ITGA6 and MACC1, or MYC, ITGA6 and MACC1, or PTGS2, ITGA6 and B2M, or S100A4, ITGA6 and B2M, or GADD45B, ITGA6 and B2M, or ITGA6A, ITGA6 and B2M, or MYBL2, ITGA6 and B2M, or MYC, ITGA6 and B2M, orPTGS2, MACC1 and B2M, or S100A4, MACC1 and B2M, or GADD45B, MACC1 and B2M, or ITGA6A, MACC1 and B2M, or MYBL2, MACC1 and B2M, or MYC, MACC1 and B2M.

[0085] In some embodiments, the SSD or TEP can include the mRNA expression level of one marker detectable in stool samples and enriched in samples with CRC and three markers detectable in stool samples and enriched in samples with CRC and AA. Such embodiments can include PTGS2, CEACAM5, ITGA6 and MACC1, or S100A4, CEACAM5, ITGA6 and MACC1, or GADD45B, CEACAM5, ITGA6 and MACC1, or ITGA6A, CEACAM5, ITGA6 and MACC1, or MYBL2, CEACAM5, ITGA6 and MACC1, or MYC, CEACAM5, ITGA6 and MACC1, orPTGS2, CEACAM5, ITGA6 and B2M, or S100A4, CEACAM5, ITGA6 and B2M, or GADD45B, CEACAM5, ITGA6 and B2M, or ITGA6A, CEACAM5, ITGA6 and B2M, or MYBL2, CEACAM5, ITGA6 and B2M, or MYC, CEACAM5, ITGA6 and B2M, or PTGS2, ITGA6, MACC1 and B2M, or S100A4, ITGA6, MACC1 and B2M, or GADD45B, ITGA6, MACC1 and B2M, or ITGA6A, ITGA6, MACC1 and B2M, or MYBL2, ITGA6, MACC1 and B2M, or MYC, ITGA6, MACC1 and B2M, or PTGS2, CEACAM5, MACC1 and B2M, or S100A4, CEACAM5, MACC1 and B2M, or GADD45B, CEACAM5, MACC1 and B2M, or ITGA6A, CEACAM5, MACC1 and B2M, or MYBL2, CEACAM5, MACC1 and B2M, or MYC, CEACAM5, MACC1 and B2M.

[0086] In some embodiments, the SSD or TEP can include the mRNA expression level of one marker detectable in stool samples and enriched in samples with CRC and four markers detectable in stool samples and enriched in samples with CRC and AA. Such embodiments can include PTGS2, CEACAM5, ITGA6, MACC1 and B2M, or S100A4, CEACAM5, ITGA6, MACC1 and B2M, or GADD45B, CEACAM5, ITGA6, MACC1 and B2M, or ITGA6A, CEACAM5, ITGA6, MACC1 and B2M, orMYBL2, CEACAM5, ITGA6, MACC1 and B2M, or MYC, CEACAM5, ITGA6, MACC1 and B2M.

[0087] In some embodiments, the SSD or TEP can include the mRNA expression level of two markers detectable in stool samples and enriched in samples with CRC and two markers detectable in stool samples and enriched in samples with CRC and AA. Such embodiments can include PTGS2, S100A4 and CEACAM5 and ITGA6, or PTGS2, GADD45B and CEACAM5 and ITGA6, or PTGS2, ITGA6A and CEACAM5 and ITGA6, or PTGS2, MYBL2 and CEACAM5 and ITGA6, or PTGS2, MYC and CEACAM5 and ITGA6, or S100A4, GADD45B and CEACAM5 and ITGA6, or S100A4, ITGA6A and CEACAM5 and ITGA6, or S100A4, MYBL2 and CEACAM5 and ITGA6, or S100A4, MYC and CEACAM5 and ITGA6, or GADD45B, ITGA6A and CEACAM5 and ITGA6, or GADD45B, MYBL2 and CEACAM5 and ITGA6, or GADD45B, MYC and CEACAM5 and ITGA6, or ITGA6A, MYBL2 and CEACAM5 and ITGA6, orITGA6A, MYC and CEACAM5 and ITGA6, or MYBL2, MYC and CEACAM5 and ITGA6, or PTGS2, S100A4 and CEACAM5 and MACC1, or PTGS2, GADD45B and CEACAM5 and MACC1, or PTGS2, ITGA6A and CEACAM5 and MACC1, or PTGS2, MYBL2 and CEACAM5 and MACC1, or PTGS2, MYC and CEACAM5 and MACC1, or S100A4, GADD45B and CEACAM5 and MACC1, or S100A4, ITGA6A and CEACAM5 and MACC1, or S100A4, MYBL2 and CEACAM5 and MACC1, or S100A4, MYC and CEACAM5 and MACC1, or GADD45B, ITGA6A and CEACAM5 and MACC1, or GADD45B, MYBL2 and CEACAM5 and MACC1, or GADD45B, MYC and CEACAM5 and MACC1, or ITGA6A, MYBL2 and CEACAM5 and MACC1, or PTGS2, S100A4 and CEACAM5 and B2M, or PTGS2, GADD45B and CEACAM5 and B2M, or PTGS2, ITGA6A and CEACAM5 and B2M, or PTGS2, MYBL2 and CEACAM5 and B2M, or PTGS2, MYC and CEACAM5 and B2M, or S100A4, GADD45B and CEACAM5 and B2M, or S100A4, ITGA6A and CEACAM5 and B2M, or S100A4, MYBL2 and CEACAM5 and B2M, or S100A4, MYC and CEACAM5 and B2M, or GADD45B, ITGA6A and CEACAM5 and B2M, or GADD45B, MYBL2 and CEACAM5 and B2M, or GADD45B, MYC and CEACAM5 and B2M, or ITGA6A, MYBL2 and CEACAM5 and B2M, or ITGA6A, MYC and CEACAM5 and B2M, or MYBL2, MYC and CEACAM5 and B2M, or PTGS2, S100A4 and ITGA6 and MACC1, or PTGS2, GADD45B and ITGA6 and MACC1, or PTGS2, ITGA6A and ITGA6 and MACC1, or PTGS2, MYBL2 and ITGA6 and MACC1, or PTGS2, MYC and ITGA6 and MACC1, or S100A4, GADD45B and ITGA6 and MACC1, or S100A4, ITGA6A and ITGA6 and MACC1, or S100A4, MYBL2 and ITGA6 and MACC1, or S100A4, MYC and ITGA6 and MACC1, or GADD45B, ITGA6A and ITGA6 and MACC1, or GADD45B, MYBL2 and ITGA6 and MACC1, or GADD45B, MYC and ITGA6 and MACC1, or ITGA6A, MYBL2 and , ITGA6 and MACClor ITGA6A, MYC and , ITGA6 and MACClor MYBL2, MYC and ITGA6 and MACC1, or PTGS2, S100A4 and ITGA6 and B2M, or PTGS2, GADD45B and ITGA6 and B2M, or PTGS2, ITGA6A and ITGA6 and B2M, or PTGS2, MYBL2 and ITGA6 and B2M, or PTGS2, MYC and ITGA6 and B2M, or S100A4, GADD45B and ITGA6 and B2M, or S100A4, ITGA6A and ITGA6 and B2M, or S100A4, MYBL2 and ITGA6 and B2M, or S100A4, MYC and ITGA6 and B2M, or GADD45B, ITGA6A and ITGA6 and B2M, or GADD45B, MYBL2 and ITGA6 and B2M, or GADD45B, MYC and ITGA6 and B2M, or ITGA6A, MYBL2 and ITGA6 and B2M, or ITGA6A, MYC and ITGA6 and B2M, or MYBL2, MYC and ITGA6 and B2M, orPTGS2, S100A4 and MACCl and B2M, orPTGS2, GADD45B and MACCl and B2M, or PTGS2, ITGA6A and MACC1 and B2M, or PTGS2, MYBL2 and MACC1 and B2M,or PTGS2, MYC and MACC1 and B2M, or S100A4, GADD45B and MACC1 and B2M, or S100A4, ITGA6A and MACC1 and B2M, or S100A4, MYBL2 and MACC1 and B2M, or S100A4, MYC and MACC1 and B2M, or GADD45B, ITGA6A and MACC1 and B2M, or GADD45B, MYBL2 and MACC1 and B2M, or GADD45B, MYC and MACC1 and B2M, or ITGA6A, MYBL2 and MACC1 and B2M, or ITGA6A, MYC and MACC1 and B2M, or MYBL2, MYC and MACC1 and B2M, or ITGA6A, MYC and MACC1 and B2M, or MYBL2, MYC and MACC1 and B2M.

[0088] In some embodiments, the SSD or TEP can include the mRNA expression level of two markers detectable in stool samples and enriched in samples with CRC and three markers detectable in stool samples and enriched in samples with CRC and AA. Such embodiments can include PTGS2, S100A4 and CEACAM5, ITGA6 and MACC1, or PTGS2, GADD45B and CEACAM5, ITGA6 and MACC1, or PTGS2, ITGA6A and CEACAM5, ITGA6 and MACC1, or PTGS2, MYBL2 and CEACAM5, ITGA6 and MACC1, orPTGS2, MYC and CEACAM5, ITGA6 and MACC1, or S100A4, GADD45B and CEACAM5, ITGA6 and MACC1, or S100A4, ITGA6A and CEACAM5, ITGA6 and MACC1, or S100A4, MYBL2 and CEACAM5, ITGA6 and MACC1, or S100A4, MYC and CEACAM5, ITGA6 and MACC1, or GADD45B, ITGA6A and CEACAM5, ITGA6 and MACC1, or GADD45B, MYBL2 and CEACAM5, ITGA6 and MACC1, or GADD45B, MYC and CEACAM5, ITGA6 and MACC1, or ITGA6A, MYBL2 and CEACAM5, ITGA6 and MACC1, or ITGA6A, MYC and CEACAM5, ITGA6 and MACC1, or MYBL2, MYC and CEACAM5, ITGA6 and MACC1, or PTGS2, S100A4 and CEACAM5, ITGA6 and B2M, or PTGS2, GADD45B and CEACAM5, ITGA6 and B2M, or PTGS2, ITGA6A and CEACAM5, ITGA6 and B2M, or PTGS2, MYBL2 and CEACAM5, ITGA6 and B2M, or PTGS2, MYC and CEACAM5, ITGA6 and B2M, or S100A4, GADD45B and CEACAM5, ITGA6 and B2M, or S100A4, ITGA6A and CEACAM5, ITGA6 and B2M, or S100A4, MYBL2 and CEACAM5, ITGA6 and B2M, or S100A4, MYC and CEACAM5, ITGA6 and B2M, or GADD45B, ITGA6A and CEACAM5, ITGA6 and B2M, or GADD45B, MYBL2 and CEACAM5, ITGA6 and B2M, or GADD45B, MYC and CEACAM5, ITGA6 and B2M, or ITGA6A, MYBL2 and CEACAM5, ITGA6 and B2M, or ITGA6A, MYC and CEACAM5, ITGA6 and B2M, or MYBL2, MYC and CEACAM5, ITGA6 and B2M, or PTGS2, S100A4 and ITGA6, MACC1 and B2M, or PTGS2, GADD45B and ITGA6, MACC1 and B2M, or PTGS2, ITGA6A and ITGA6, MACC1 and B2M, or PTGS2, MYBL2 and ITGA6, MACC1 and B2M, or PTGS2, MYC and ITGA6, MACC1 and B2M, or S100A4, GADD45B andITGA6, MACC1 and B2M, or S100A4, ITGA6A and ITGA6, MACC1 and B2M, or S100A4, MYBL2 and ITGA6, MACC1 andB2M, or S100A4, MYC and ITGA6, MACC1 and B2M, or GADD45B, ITGA6A and ITGA6, MACC1 and B2M, or GADD45B, MYBL2 and ITGA6, MACC1 and B2M, or GADD45B, MYC and ITGA6, MACC1 and B2M, or ITGA6A, MYBL2 and ITGA6, MACC1 and B2M, or ITGA6A, MYC and ITGA6, MACC1 and B2M, or MYBL2, MYC and ITGA6, MACC1 and B2M, or PTGS2, S100A4 and CEACAM5, MACC1 and B2M, or PTGS2, GADD45B and CEACAM5, MACC1 and B2M, or PTGS2, ITGA6A and CEACAM5, MACC1 and B2M, or PTGS2, MYBL2 and CEACAM5, MACC1 and B2M, orPTGS2, MYC and CEACAM5, MACC1 and B2M, or S100A4, GADD45B and CEACAM5, MACC1 and B2M, or S100A4, ITGA6A and CEACAM5, MACC1 and B2M, or S100A4, MYBL2 and CEACAM5, MACC1 and B2M, or S100A4, MYC and CEACAM5, MACC1 and B2M, or GADD45B, ITGA6A and CEACAM5, MACC1 and B2M, or GADD45B, MYBL2 and CEACAM5, MACC1 and B2M, or GADD45B, MYC and CEACAM5, MACC1 and B2M, or ITGA6A, MYBL2 and CEACAM5, MACC1 and B2M, or ITGA6A, MYC and CEACAM5, MACC1 and B2M, or MYBL2, MYC and CEACAM5, MACC1 and B2M. In one particular embodiment, the two markers detectable in stool samples and enriched in samples with CRC are PTGS2 and S100A4, and the three markers detectable in stool samples and enriched in samples with CRC and AA are CEACAM5, ITGA6 and MACC1.

[0089] In some embodiments, the SSD or TEP can include the mRNA expression level of two markers detectable in stool samples and enriched in samples with CRC and four markers detectable in stool samples and enriched in samples with CRC and AA. Such embodiments can include PTGS2, S100A4 and CEACAM5, ITGA6, MACC1 and B2M, or PTGS2, GADD45B and CEACAM5, ITGA6, MACC1 and B2M, or PTGS2, ITGA6A and CEACAM5, ITGA6, MACC1 and B2M, or PTGS2, MYBL2 and CEACAM5, ITGA6, MACC1 andB2M, orPTGS2, MYC and CEACAM5, ITGA6, MACC1 andB2M, or S100A4, GADD45B and CEACAM5, ITGA6, MACC1 and B2M, or S100A4, ITGA6A and CEACAM5, ITGA6, MACC1 and B2M, or S100A4, MYBL2 and CEACAM5, ITGA6, MACC1 and B2M, or S100A4, MYC and CEACAM5, ITGA6, MACC1 and B2M, or GADD45B, ITGA6A and CEACAM5, ITGA6, MACC1 and B2M, or GADD45B, MYBL2 and CEACAM5, ITGA6, MACC1 and B2M, or GADD45B, MYC and CEACAM5, ITGA6, MACC1 andB2M, orITGA6A, MYBL2 and CEACAM5, ITGA6, MACC1 and B2M, or ITGA6A, MYC and CEACAM5, ITGA6, MACC1 and B2M, or MYBL2, MYC and CEACAM5, ITGA6, MACC1 and B2M.

[0090] In some embodiments, the SSD or TEP can include the mRNA expression level of three markers detectable in stool samples and enriched in samples with CRC and two markers detectable in stool samples and enriched in samples with CRC and AA. Such embodiments can include PTGS2, S100A4, GADD45B and CEACAM5 and ITGA6, or PTGS2, S100A4, ITGA6A and CEACAM5 and ITGA6, or PTGS2, S100A4, MYBL2 and CEACAM5 and ITGA6, or PTGS2, S100A4, MYC and CEACAM5 and ITGA6, or PTGS2, GADD45B, ITGA6A and CEACAM5 and ITGA6, or PTGS2, GADD45B, MYBL2 and CEACAM5 and ITGA6, or PTGS2, GADD45B, MYC and CEACAM5 and ITGA6, or PTGS2, ITGA6A, MYBL2 and CEACAM5 and ITGA6, or PTGS2, ITGA6A, MYC and CEACAM5 and ITGA6, or PTGS2, MYBL2, MYC and CEACAM5 and ITGA6, or S100A4, GADD45B, ITGA6A and CEACAM5 and ITGA6, or S100A4, GADD45B, MYBL2 and CEACAM5 and ITGA6, or S100A4, GADD45B, MYC and CEACAM5 and ITGA6, or S100A4, ITGA6A, MYBL2 and CEACAM5 and ITGA6, or S100A4, ITGA6A, MYC and CEACAM5 and ITGA6, or S100A4, MYBL2, MYC and CEACAM5 and ITGA6, or GADD45B, ITGA6A, MYBL2 and CEACAM5 and ITGA6, or GADD45B, ITGA6A, MYC and CEACAM5 and ITGA6, or ITGA6A, MYBL2, MYC and CEACAM5 and ITGA6, or PTGS2, S100A4, GADD45B and CEACAM5 and MACC1, or PTGS2, S100A4, ITGA6A and CEACAM5 and MACC1, or PTGS2, S100A4, MYBL2 and CEACAM5 and MACC1, or PTGS2, S100A4, MYC and CEACAM5 and MACC1, or PTGS2, GADD45B, ITGA6A and CEACAM5 and MACC1, or PTGS2, GADD45B, MYBL2 and CEACAM5 and MACC1, or PTGS2, GADD45B, MYC and CEACAM5 and MACC1, or PTGS2, ITGA6A, MYBL2 and CEACAM5 and MACC1, or PTGS2, ITGA6A, MYC and CEACAM5 and MACC1, or PTGS2, MYBL2, MYC and CEACAM5 and MACC1, or S100A4, GADD45B, ITGA6A and CEACAM5 and MACCl, or S100A4, GADD45B, MYBL2 and CEACAM5 and MACCl, or S100A4, GADD45B, MYC and CEACAM5 and MACC1, or S100A4, ITGA6A, MYBL2 and CEACAM5 and MACC1, or S100A4, ITGA6A, MYC and CEACAM5 and MACC1, or S100A4, MYBL2, MYC and CEACAM5 and MACC1, or GADD45B, ITGA6A, MYBL2 and CEACAM5 and MACC1, or GADD45B, ITGA6A, MYC and CEACAM5 and MACC1, or ITGA6A, MYBL2, MYC and CEACAM5 and MACC1, or PTGS2, S100A4, GADD45B and CEACAM5 and B2M, or PTGS2, S100A4, ITGA6A and CEACAM5 and B2M, or PTGS2, S100A4, MYBL2 and CEACAM5 and B2M, or PTGS2, S100A4, MYC and CEACAM5 and B2M, or PTGS2, GADD45B, ITGA6A and CEACAM5 and B2M, or PTGS2, GADD45B, MYBL2 and CEACAM5 and B2M, or PTGS2, GADD45B, MYC and CEACAM5 and B2M, or PTGS2, ITGA6A, MYBL2 and CEACAM5 and B2M, orPTGS2, ITGA6A, MYC and CEACAM5 and B2M, or PTGS2, MYBL2, MYC and CEACAM5 and B2M, or S100A4, GADD45B, ITGA6A and CEACAM5 and B2M, or S100A4, GADD45B, MYBL2 and CEACAM5 and B2M, or S100A4, GADD45B, MYC and CEACAM5 and B2M, or S100A4, ITGA6A, MYBL2 and CEACAM5 and B2M, or S100A4, ITGA6A, MYC and CEACAM5 and B2M, or S100A4, MYBL2, MYC and CEACAM5 and B2M, or GADD45B, ITGA6A, MYBL2 and CEACAM5 and B2M, or GADD45B, ITGA6A, MYC and CEACAM5 and B2M, or ITGA6A, MYBL2, MYC and CEACAM5 and B2M, or PTGS2, S100A4, GADD45B and ITGA6 and MACC1, or PTGS2, S100A4, ITGA6A and ITGA6 and MACC1, or PTGS2, S100A4, MYBL2 and ITGA6 and MACC1, or PTGS2, S100A4, MYC and ITGA6 and MACC1, or PTGS2, GADD45B, ITGA6A and ITGA6 and MACC1, or PTGS2, GADD45B, MYBL2 and ITGA6 and MACC1, or PTGS2, GADD45B, MYC and ITGA6 and MACC1, or PTGS2, ITGA6A, MYBL2 and ITGA6 and MACC1, or PTGS2, ITGA6A, MYC and ITGA6 and MACC1, or PTGS2, MYBL2, MYC and ITGA6 and MACC1, or S100A4, GADD45B, ITGA6A and ITGA6 and MACC1, or S100A4, GADD45B, MYBL2 and ITGA6 and MACC1, or S100A4, GADD45B, MYC and ITGA6 and MACC1, or S100A4, ITGA6A, MYBL2 and ITGA6 and MACC1, or S100A4, ITGA6A, MYC and ITGA6 and MACC1, or S100A4, MYBL2, MYC and ITGA6 and MACC1, or GADD45B, ITGA6A, MYBL2 and ITGA6 and MACC1, or GADD45B, ITGA6A, MYC and ITGA6 and MACC1, or ITGA6A, MYBL2, MYC and ITGA6 and MACC1, or PTGS2, S100A4, GADD45B and ITGA6 and B2M, or PTGS2, S100A4, ITGA6A and ITGA6 and B2M, or PTGS2, S100A4, MYBL2 and ITGA6 and B2M, or PTGS2, S100A4, MYC and ITGA6 and B2M, or PTGS2, GADD45B, ITGA6A and ITGA6 and B2M, or PTGS2, GADD45B, MYBL2 and ITGA6 and B2M, or PTGS2, GADD45B, MYC and ITGA6 and B2M, or PTGS2, ITGA6A, MYBL2 and ITGA6 and B2M, or PTGS2, ITGA6A, MYC and ITGA6 and B2M, or PTGS2, MYBL2, MYC and ITGA6 and B2M, or S100A4, GADD45B, ITGA6A and ITGA6 and B2M, or S100A4, GADD45B, MYBL2 and ITGA6 and B2M, or S100A4, GADD45B, MYC and ITGA6 and B2M, or S100A4, ITGA6A, MYBL2 and ITGA6 and B2M, or S100A4, ITGA6A, MYC and ITGA6 and B2M, or S100A4, MYBL2, MYC and ITGA6 and B2M, or GADD45B, ITGA6A, MYBL2 and ITGA6 and B2M, or GADD45B, ITGA6A, MYC and ITGA6 and B2M, or ITGA6A, MYBL2, MYC and ITGA6 and B2M, or PTGS2, S100A4, GADD45B and MACC1 and B2M, or PTGS2, S100A4, ITGA6A and MACC1 and B2M, or PTGS2, S100A4, MYBL2 and MACC1 and B2M, or PTGS2, S100A4, MYC and MACC1 and B2M, or PTGS2, GADD45B, ITGA6A and MACC1 and B2M, or PTGS2, GADD45B, MYBL2 and MACC1 and B2M, or PTGS2, GADD45B,MYC and MACC1 and B2M, or PTGS2, ITGA6A, MYBL2 and MACC1 and B2M, or PTGS2, ITGA6A, MYC and MACC1 and B2M, or PTGS2, MYBL2, MYC and MACC1 and B2M, or S100A4, GADD45B, ITGA6A and MACC1 and B2M, or S100A4, GADD45B, MYBL2 and MACC1 and B2M, or S100A4, GADD45B, MYC and MACC1 and B2M, or S100A4, ITGA6A, MYBL2 and MACC1 and B2M, or S100A4, ITGA6A, MYC and MACC1 and B2M, or S100A4, MYBL2, MYC and MACC1 and B2M, or GADD45B, ITGA6A, MYBL2 and MACC1 and B2M, or GADD45B, ITGA6A, MYC and MACC1 and B2M, or ITGA6A, MYBL2, MYC and MACC1 and B2M.

[0091] In some embodiments, the SSD or TEP can include the mRNA expression level of three markers detectable in stool samples and enriched in samples with CRC and three markers detectable in stool samples and enriched in samples with CRC and AA. Such embodiments can include PTGS2, S100A4, GADD45B and CEACAM5, ITGA6 and MACC1, orPTGS2, S100A4, ITGA6A and CEACAM5, ITGA6 and MACCl, orPTGS2, S100A4, MYBL2 and CEACAM5, ITGA6 and MACC1, or PTGS2, S100A4, MYC and CEACAM5, ITGA6 and MACC1, or PTGS2, GADD45B, ITGA6A and CEACAM5, ITGA6 and MACC1, or PTGS2, GADD45B, MYBL2 and CEACAM5, ITGA6 and MACC1, orPTGS2, GADD45B, MYC and CEACAM5, ITGA6 and MACC1, or PTGS2, ITGA6A, MYBL2 and CEACAM5, ITGA6 and MACC1, or PTGS2, ITGA6A, MYC and CEACAM5, ITGA6 and MACC1, or PTGS2, MYBL2, MYC and CEACAM5, ITGA6 and MACC1, or S100A4, GADD45B, ITGA6A and CEACAM5, ITGA6 and MACC1, or S100A4, GADD45B, MYBL2 and CEACAM5, ITGA6 and MACC1, or S100A4, GADD45B, MYC and CEACAM5, ITGA6 and MACC1, or S100A4, ITGA6A, MYBL2 and CEACAM5, ITGA6 and MACC1, or S100A4, ITGA6A, MYC and CEACAM5, ITGA6 and MACC1, or S100A4, MYBL2, MYC and CEACAM5, ITGA6 and MACC1, or GADD45B, ITGA6A, MYBL2 and CEACAM5, ITGA6 and MACC1, or GADD45B, ITGA6A, MYC and CEACAM5, ITGA6 and MACC1, or ITGA6A, MYBL2, MYC and CEACAM5, ITGA6 and MACC1, or PTGS2, S100A4, GADD45B and CEACAM5, ITGA6 and B2M, or PTGS2, S100A4, ITGA6A and CEACAM5, ITGA6 and B2M, or PTGS2, S100A4, MYBL2 and CEACAM5, ITGA6 and B2M, or PTGS2, S100A4, MYC and CEACAM5, ITGA6 and B2M, or PTGS2, GADD45B, ITGA6A and CEACAM5, ITGA6 and B2M, or PTGS2, GADD45B, MYBL2 and CEACAM5, ITGA6 and B2M, or PTGS2, GADD45B, MYC and CEACAM5, ITGA6 and B2M, or PTGS2, ITGA6A, MYBL2 and CEACAM5, ITGA6 and B2M, or PTGS2, ITGA6A, MYC and CEACAM5, ITGA6 and B2M, or PTGS2, MYBL2, MYC and CEACAM5, ITGA6 and B2M, orS100A4, GADD45B, ITGA6A and CEACAM5, ITGA6 and B2M, or S100A4, GADD45B, MYBL2 and CEACAM5, ITGA6 and B2M, or S100A4, GADD45B, MYC and CEACAM5, ITGA6 and B2M, or S100A4, ITGA6A, MYBL2 and CEACAM5, ITGA6 and B2M, or S100A4, ITGA6A, MYC and CEACAM5, ITGA6 and B2M, or S100A4, MYBL2, MYC and CEACAM5, ITGA6 and B2M, or GADD45B, ITGA6A, MYBL2 and CEACAM5, ITGA6 and B2M, or GADD45B, ITGA6A, MYC and CEACAM5, ITGA6 and B2M, or ITGA6A, MYBL2, MYC and CEACAM5, ITGA6 and B2M, or PTGS2, S100A4, GADD45B and ITGA6, MACC1 and B2M, or PTGS2, S100A4, ITGA6A and ITGA6, MACC1 and B2M, or PTGS2, S100A4, MYBL2 and ITGA6, MACC1 and B2M, or PTGS2, S100A4, MYC and ITGA6, MACC1 and B2M, or PTGS2, GADD45B, ITGA6A and ITGA6, MACC1 and B2M, or PTGS2, GADD45B, MYBL2 and ITGA6, MACC1 and B2M, or PTGS2, GADD45B, MYC and ITGA6, MACC1 and B2M, or PTGS2, ITGA6A, MYBL2 and ITGA6, MACC1 and B2M, or PTGS2, ITGA6A, MYC and ITGA6, MACC1 and B2M, or PTGS2, MYBL2, MYC and ITGA6, MACC1 and B2M, or S100A4, GADD45B, ITGA6A and ITGA6, MACC1 and B2M, or S100A4, GADD45B, MYBL2 and ITGA6, MACC1 and B2M, or S100A4, GADD45B, MYC and ITGA6, MACC1 and B2M, or S100A4, ITGA6A, MYBL2 and ITGA6, MACC1 and B2M, or S100A4, ITGA6A, MYC and ITGA6, MACC1 and B2M, or S100A4, MYBL2, MYC and ITGA6, MACC1 and B2M, or GADD45B, ITGA6A, MYBL2 and ITGA6, MACC1 and B2M, or GADD45B, ITGA6A, MYC and ITGA6, MACC1 and B2M, or ITGA6A, MYBL2, MYC and ITGA6, MACC1 and B2M, or PTGS2, S100A4, GADD45B and CEACAM5, MACC1 and B2M, or PTGS2, S100A4, ITGA6A and CEACAM5, MACC1 and B2M, or PTGS2, S100A4, MYBL2 and CEACAM5, MACC1 and B2M, orPTGS2, S100A4, MYC and CEACAM5, MACC1 and B2M, or PTGS2, GADD45B, ITGA6A and CEACAM5, MACC1 and B2M, or PTGS2, GADD45B, MYBL2 and CEACAM5, MACC1 and B2M, or PTGS2, GADD45B, MYC and CEACAM5, MACC1 and B2M, or PTGS2, ITGA6A, MYBL2 and CEACAM5, MACC1 and B2M, or PTGS2, ITGA6A, MYC and CEACAM5, MACC1 and B2M, or PTGS2, MYBL2, MYC and CEACAM5, MACC1 and B2M, or S100A4, GADD45B, ITGA6A and CEACAM5, MACC1 and B2M, or S100A4, GADD45B, MYBL2 and CEACAM5, MACC1 and B2M, or S100A4, GADD45B, MYC and CEACAM5, MACC1 and B2M, or S100A4, ITGA6A, MYBL2 and CEACAM5, MACC1 and B2M, or S100A4, ITGA6A, MYC and CEACAM5, MACC1 and B2M, or S100A4, MYBL2, MYC and CEACAM5, MACC1 and B2M, or GADD45B, ITGA6A, MYBL2 and CEACAM5, MACC1 and B2M, or GADD45B, ITGA6A, MYC and CEACAM5, MACC1 and B2M,or ITGA6A, MYBL2, MYC and CEACAM5, MACC1 and B2M.

[0092] In some embodiments, the SSD or TEP can include the mRNA expression level of three markers detectable in stool samples and enriched in samples with CRC and four markers detectable in stool samples and enriched in samples with CRC and AA. Such embodiments can include PTGS2, S100A4, GADD45B and CEACAM5, ITGA6, MACC1 and B2M, or PTGS2, S100A4, ITGA6A and CEACAM5, ITGA6, MACC1 and B2M, or PTGS2, S100A4, MYBL2 and CEACAM5, ITGA6, MACC1 and B2M, or PTGS2, S100A4, MYC and CEACAM5, ITGA6, MACC1 and B2M, or PTGS2, GADD45B, ITGA6A and CEACAM5, ITGA6, MACC1 and B2M, or PTGS2, GADD45B, MYBL2 and CEACAM5, ITGA6, MACC1 andB2M, orPTGS2, GADD45B, MYC and CEACAM5, ITGA6, MACC1 and B2M, or PTGS2, ITGA6A, MYBL2 and CEACAM5, ITGA6, MACC1 and B2M, or PTGS2, ITGA6A, MYC and CEACAM5, ITGA6, MACC1 and B2M, orPTGS2, MYBL2, MYC and CEACAM5, ITGA6, MACC1 and B2M, or S100A4, GADD45B, ITGA6A and CEACAM5, ITGA6, MACC1 and B2M, or S100A4, GADD45B, MYBL2 and CEACAM5, ITGA6, MACC1 and B2M, or S100A4, GADD45B, MYC and CEACAM5, ITGA6, MACC1 and B2M, or S100A4, ITGA6A, MYBL2 and CEACAM5, ITGA6, MACC1 and B2M, or S100A4, ITGA6A, MYC and CEACAM5, ITGA6, MACC1 and B2M, or S100A4, MYBL2, MYC and CEACAM5, ITGA6, MACC1 and B2M, or GADD45B, ITGA6A, MYBL2 and CEACAM5, ITGA6, MACC1 and B2M, or GADD45B, ITGA6A, MYC and CEACAM5, ITGA6, MACC1 and B2M, or ITGA6A, MYBL2, MYC and CEACAM5, ITGA6, MACC1 and B2M.

[0093] In some embodiments, the SSD or TEP can include the mRNA expression level of four markers detectable in stool samples and enriched in samples with CRC and two markers detectable in stool samples and enriched in samples with CRC and AA. Such embodiments can include PTGS2, S100A4, GADD45B, ITGA6A and CEACAM5 and ITGA6, or PTGS2, S100A4, GADD45B, MYBL2 and CEACAM5 and ITGA6, or PTGS2, S100A4, GADD45B, MYC and CEACAM5 and ITGA6, or PTGS2, S100A4, ITGA6A, MYBL2 and CEACAM5 and ITGA6, or PTGS2, S100A4, ITGA6A, MYC and CEACAM5 and ITGA6, or PTGS2, S100A4, MYBL2, MYC and CEACAM5 and ITGA6, or PTGS2, S100A4, ITGA6A, MYC and CEACAM5 and ITGA6, or S100A4, GADD45B, ITGA6A, MYBL2 and CEACAM5 and ITGA6, or S100A4, GADD45B, MYBL2, MYC and CEACAM5 and ITGA6, or S100A4, GADD45B, ITGA6A, MYC and CEACAM5 and ITGA6, or PTGS2, GADD45B, ITGA6A, MYBL2 and CEACAM5 andITGA6, or PTGS2, GADD45B, MYBL2, MYC and CEACAM5 and ITGA6, or PTGS2, GADD45B, ITGA6A, MYC and CEACAM5 and ITGA6, or PTGS2, ITGA6A, MYBL2, MYC and CEACAM5 and ITGA6, or S100A4, GADD45B, ITGA6A, MYBL2 and CEACAM5 and ITGA6, or S100A4, GADD45B, ITGA6A, MYC and CEACAM5 and ITGA6, or GADD45B, ITGA6A, MYBL2, MYC and CEACAM5 and ITGA6, or PTGS2, S100A4, GADD45B, ITGA6A and CEACAM5 and MACC1, or PTGS2, S100A4, GADD45B, MYBL2 and CEACAM5 and MACC1, or PTGS2, S100A4, GADD45B, MYC and CEACAM5 and MACC1, or PTGS2, S100A4, ITGA6A, MYBL2 and CEACAM5 and MACC1, or PTGS2, S100A4, ITGA6A, MYC and CEACAM5 and MACC1, or PTGS2, S100A4, MYBL2, MYC and CEACAM5 and MACC1, or PTGS2, S100A4, ITGA6A, MYC and CEACAM5 and MACC1, or S100A4, GADD45B, ITGA6A, MYBL2 and CEACAM5 and MACC1, or S100A4, GADD45B, MYBL2, MYC and CEACAM5 and MACC1, or S100A4, GADD45B, ITGA6A, MYC and CEACAM5 and MACC1, or PTGS2, GADD45B, ITGA6A, MYBL2 and CEACAM5 and MACC1, or PTGS2, GADD45B, MYBL2, MYC and CEACAM5 and MACC1, or PTGS2, GADD45B, ITGA6A, MYC and CEACAM5 andMACCl, orPTGS2, ITGA6A, MYBL2, MYC and CEACAM5 and MACC1, or S100A4, GADD45B, ITGA6A, MYBL2 and CEACAM5 and MACC1, or S100A4, GADD45B, ITGA6A, MYC and CEACAM5 and MACC1, or GADD45B, ITGA6A, MYBL2, MYC and CEACAM5 and MACC1, or PTGS2, S100A4, GADD45B, ITGA6A and CEACAM5 and B2M, or PTGS2, S100A4, GADD45B, MYBL2 and CEACAM5 and B2M, or PTGS2, S100A4, GADD45B, MYC and CEACAM5 and B2M, or PTGS2, S100A4, ITGA6A, MYBL2 and CEACAM5 and B2M, or PTGS2, S100A4, ITGA6A, MYC and CEACAM5 and B2M, or PTGS2, S100A4, MYBL2, MYC and CEACAM5 and B2M, or PTGS2, S100A4, ITGA6A, MYC and CEACAM5 and B2M, or S100A4, GADD45B, ITGA6A, MYBL2 and CEACAM5 and B2M, or S100A4, GADD45B, MYBL2, MYC and CEACAM5 and B2M, or S100A4, GADD45B, ITGA6A, MYC and CEACAM5 and B2M, orPTGS2, GADD45B, ITGA6A, MYBL2 and CEACAM5 and B2M, or PTGS2, GADD45B, MYBL2, MYC and CEACAM5 and B2M, or PTGS2, GADD45B, ITGA6A, MYC and CEACAM5 and B2M, or PTGS2, ITGA6A, MYBL2, MYC and CEACAM5 and B2M, or S100A4, GADD45B, ITGA6A, MYBL2 and CEACAM5 and B2M, or S100A4, GADD45B, ITGA6A, MYC and CEACAM5 and B2M, or GADD45B, ITGA6A, MYBL2, MYC and CEACAM5 and B2M, or PTGS2, S100A4, GADD45B, ITGA6A and ITGA6 and MACC1, or PTGS2, S100A4, GADD45B, MYBL2 and ITGA6 and MACC1, or PTGS2, S100A4, GADD45B, MYC and ITGA6 and MACC1, or PTGS2, S100A4, ITGA6A, MYBL2 and ITGA6 andMACC1, or PTGS2, S100A4, ITGA6A, MYC and ITGA6 and MACC1, or PTGS2, S100A4, MYBL2, MYC and ITGA6 and MACC1, or PTGS2, S100A4, ITGA6A, MYC and ITGA6 and MACC1, or S100A4, GADD45B, ITGA6A, MYBL2 and ITGA6 and MACC1, or S100A4, GADD45B, MYBL2, MYC and ITGA6 and MACC1, or S100A4, GADD45B, ITGA6A, MYC and ITGA6 and MACC1, or PTGS2, GADD45B, ITGA6A, MYBL2 and ITGA6 and MACC1, or PTGS2, GADD45B, MYBL2, MYC and ITGA6 and MACC1, or PTGS2, GADD45B, ITGA6A, MYC and ITGA6 and MACC1, or PTGS2, ITGA6A, MYBL2, MYC and ITGA6 and MACC1, or S100A4, GADD45B, ITGA6A, MYBL2 and ITGA6 and MACC1, or S100A4, GADD45B, ITGA6A, MYC and ITGA6 and MACC1, or GADD45B, ITGA6A, MYBL2, MYC and ITGA6 and MACC1, or PTGS2, S100A4, GADD45B, ITGA6A and ITGA6 and B2M, or PTGS2, S100A4, GADD45B, MYBL2 and ITGA6 and B2M, or PTGS2, S100A4, GADD45B, MYC and ITGA6 and B2M, or PTGS2, S100A4, ITGA6A, MYBL2 and ITGA6 and B2M, or PTGS2, S100A4, ITGA6A, MYC and ITGA6 and B2M, or PTGS2, S100A4, MYBL2, MYC and ITGA6 and B2M, or PTGS2, S100A4, ITGA6A, MYC and ITGA6 and B2M, or S100A4, GADD45B, ITGA6A, MYBL2 and ITGA6 and B2M, or S100A4, GADD45B, MYBL2, MYC and ITGA6 and B2M, or S100A4, GADD45B, ITGA6A, MYC and ITGA6 and B2M, or PTGS2, GADD45B, ITGA6A, MYBL2 and ITGA6 and B2M, or PTGS2, GADD45B, MYBL2, MYC and ITGA6 and B2M, or PTGS2, GADD45B, ITGA6A, MYC and ITGA6 and B2M, or PTGS2, ITGA6A, MYBL2, MYC and ITGA6 and B2M, or S100A4, GADD45B, ITGA6A, MYBL2 and ITGA6 and B2M, or S100A4, GADD45B, ITGA6A, MYC and ITGA6 and B2M, or GADD45B, ITGA6A, MYBL2, MYC and ITGA6 and B2M, or PTGS2, S100A4, GADD45B, ITGA6A and MACC1 and B2M, or PTGS2, S100A4, GADD45B, MYBL2 and MACC1 and B2M, or PTGS2, S100A4, GADD45B, MYC and MACC1 and B2M, or PTGS2, S100A4, ITGA6A, MYBL2 and MACC1 and B2M, or PTGS2, S100A4, ITGA6A, MYC and MACC1 and B2M, or PTGS2, S100A4, MYBL2, MYC and MACC1 and B2M, or PTGS2, S100A4, ITGA6A, MYC and MACC1 and B2M, or S100A4, GADD45B, ITGA6A, MYBL2 and MACC1 and B2M, or S100A4, GADD45B, MYBL2, MYC and MACC1 and B2M, or S100A4, GADD45B, ITGA6A, MYC and MACC1 and B2M, or PTGS2, GADD45B, ITGA6A, MYBL2 and MACC1 and B2M, or PTGS2, GADD45B, MYBL2, MYC and MACC1 and B2M, or PTGS2, GADD45B, ITGA6A, MYC and MACC1 and B2M, or PTGS2, ITGA6A, MYBL2, MYC and MACC1 and B2M, or S100A4, GADD45B, ITGA6A, MYBL2 and MACC1 and B2M, or S100A4, GADD45B, ITGA6A, MYC and MACC1 and B2M, or GADD45B, ITGA6A, MYBL2, MYC andMACC1 and B2M.

[0094] In some embodiments, the SSD or TEP can include the mRNA expression level of four markers detectable in stool samples and enriched in samples with CRC and three markers detectable in stool samples and enriched in samples with CRC and AA. Such embodiments can include PTGS2, S100A4, GADD45B, ITGA6A and CEACAM5, ITGA6 and MACCl, or PTGS2, S100A4, GADD45B, MYBL2 and CEACAM5, ITGA6 and MACC1, or PTGS2, S100A4, GADD45B, MYC and CEACAM5, ITGA6 and MACC1, or PTGS2, S100A4, ITGA6A, MYBL2 and CEACAM5, ITGA6 and MACC1, or PTGS2, S100A4, ITGA6A, MYC and CEACAM5, ITGA6 and MACC1, or PTGS2, S100A4, MYBL2, MYC and CEACAM5, ITGA6 and MACC1, or PTGS2, S100A4, ITGA6A, MYC and CEACAM5, ITGA6 and MACC1, or S100A4, GADD45B, ITGA6A, MYBL2 and CEACAM5, ITGA6 and MACC1, or S100A4, GADD45B, MYBL2, MYC and CEACAM5, ITGA6 and MACC1, or S100A4, GADD45B, ITGA6A, MYC and CEACAM5, ITGA6 and MACC1, or PTGS2, GADD45B, ITGA6A, MYBL2 and CEACAM5, ITGA6 and MACC1, or PTGS2, GADD45B, MYBL2, MYC and CEACAM5, ITGA6 and MACC1, or PTGS2, GADD45B, ITGA6A, MYC and CEACAM5, ITGA6 and MACC1, or PTGS2, ITGA6A, MYBL2, MYC and CEACAM5, ITGA6 andMACCl, or S100A4, GADD45B, ITGA6A, MYBL2 and CEACAM5, ITGA6 and MACC1, or S100A4, GADD45B, ITGA6A, MYC and , or GADD45B, ITGA6A, MYBL2, MYC and CEACAM5, ITGA6 and MACC1, or PTGS2, S100A4, GADD45B, ITGA6A and CEACAM5, ITGA6 and B2M, or PTGS2, S100A4, GADD45B, MYBL2 and CEACAM5, ITGA6 and B2M, or PTGS2, S100A4, GADD45B, MYC and CEACAM5, ITGA6 and B2M, or PTGS2, S100A4, ITGA6A, MYBL2 and CEACAM5, ITGA6 andB2M, orPTGS2, S100A4, ITGA6A, MYC and CEACAM5, ITGA6 andB2M, or PTGS2, S100A4, MYBL2, MYC and CEACAM5, ITGA6 and B2M, or PTGS2, S100A4, ITGA6A, MYC and CEACAM5, ITGA6 and B2M, or S100A4, GADD45B, ITGA6A, MYBL2 and CEACAM5, ITGA6 and B2M, or S100A4, GADD45B, MYBL2, MYC and CEACAM5, ITGA6 and B2M, or S100A4, GADD45B, ITGA6A, MYC and CEACAM5, ITGA6 and B2M, or PTGS2, GADD45B, ITGA6A, MYBL2 and CEACAM5, ITGA6 and B2M, or PTGS2, GADD45B, MYBL2, MYC and CEACAM5, ITGA6 and B2M, or PTGS2, GADD45B, ITGA6A, MYC and CEACAM5, ITGA6 and B2M, or PTGS2, ITGA6A, MYBL2, MYC and CEACAM5, ITGA6 and B2M, or S100A4, GADD45B, ITGA6A, MYBL2 and CEACAM5, ITGA6 and B2M, or S100A4, GADD45B, ITGA6A, MYC and CEACAM5, ITGA6 and B2M, or GADD45B, ITGA6A,MYBL2, MYC and CEACAM5, ITGA6 and B2M, or PTGS2, S100A4, GADD45B, ITGA6A and ITGA6, MACC1 and B2M, or PTGS2, S100A4, GADD45B, MYBL2 and ITGA6, MACC1 and B2M, or PTGS2, S100A4, GADD45B, MYC and ITGA6, MACC1 and B2M, or PTGS2, S100A4, ITGA6A, MYBL2 and ITGA6, MACC1 and B2M, or PTGS2, S100A4, ITGA6A, MYC and ITGA6, MACC1 and B2M, or PTGS2, S100A4, MYBL2, MYC and ITGA6, MACC1 and B2M, or PTGS2, S100A4, ITGA6A, MYC and ITGA6, MACC1 and B2M, or S100A4, GADD45B, ITGA6A, MYBL2 and ITGA6, MACC1 and B2M, or S100A4, GADD45B, MYBL2, MYC and ITGA6, MACC1 and B2M, or S100A4, GADD45B, ITGA6A, MYC and ITGA6, MACC1 andB2M, orPTGS2, GADD45B, ITGA6A, MYBL2 and ITGA6, MACC1 and B2M, or PTGS2, GADD45B, MYBL2, MYC and ITGA6, MACC1 and B2M, or PTGS2, GADD45B, ITGA6A, MYC and ITGA6, MACC1 and B2M, or PTGS2, ITGA6A, MYBL2, MYC and ITGA6, MACC1 and B2M, or S100A4, GADD45B, ITGA6A, MYBL2 and ITGA6, MACC1 and B2M, or S100A4, GADD45B, ITGA6A, MYC and ITGA6, MACC1 and B2M, or GADD45B, ITGA6A, MYBL2, MYC and ITGA6, MACC1 and B2M, or PTGS2, S100A4, GADD45B, ITGA6A and CEACAM5, MACC1 and B2M, or PTGS2, S100A4, GADD45B, MYBL2 and CEACAM5, MACC1 and B2M, or PTGS2, S100A4, GADD45B, MYC and CEACAM5, MACC1 and B2M, or PTGS2, S100A4, ITGA6A, MYBL2 and CEACAM5, MACC1 and B2M, or PTGS2, S100A4, ITGA6A, MYC and CEACAM5, MACC1 and B2M, or PTGS2, S100A4, MYBL2, MYC and CEACAM5, MACC1 and B2M, or PTGS2, S100A4, ITGA6A, MYC and CEACAM5, MACC1 and B2M, or S100A4, GADD45B, ITGA6A, MYBL2 and CEACAM5, MACC1 and B2M, or S100A4, GADD45B, MYBL2, MYC and CEACAM5, MACC1 and B2M, or S100A4, GADD45B, ITGA6A, MYC and CEACAM5, MACC1 and B2M, or PTGS2, GADD45B, ITGA6A, MYBL2 and CEACAM5, MACC1 and B2M, or PTGS2, GADD45B, MYBL2, MYC and CEACAM5, MACC1 and B2M, or PTGS2, GADD45B, ITGA6A, MYC and CEACAM5, MACC1 and B2M, or PTGS2, ITGA6A, MYBL2, MYC and CEACAM5, MACC1 and B2M, or S100A4, GADD45B, ITGA6A, MYBL2 and CEACAM5, MACC1 and B2M, or S100A4, GADD45B, ITGA6A, MYC and CEACAM5, MACC1 and B2M, or GADD45B, ITGA6A, MYBL2, MYC and CEACAM5, MACC1 and B2M.

[0095] In some embodiments, the SSD or TEP can include the mRNA expression level of four markers detectable in stool samples and enriched in samples with CRC and four markers detectable in stool samples and enriched in samples with CRC and AA. Such embodiments can include PTGS2, S100A4, GADD45B, ITGA6A and CEACAM5,ITGA6, MACC1 and B2M, or PTGS2, S100A4, GADD45B, MYBL2 and CEACAM5, ITGA6, MACC1 and B2M, or PTGS2, S100A4, GADD45B, MYC and CEACAM5, ITGA6, MACC1 and B2M, or PTGS2, S100A4, ITGA6A, MYBL2 and CEACAM5, ITGA6, MACC1 andB2M, orPTGS2, S100A4, ITGA6A, MYC and CEACAM5, ITGA6, MACC1 and B2M, or PTGS2, S100A4, MYBL2, MYC and CEACAM5, ITGA6, MACC1 and B2M, or PTGS2, S100A4, ITGA6A, MYC and CEACAM5, ITGA6, MACC1 and B2M, or S100A4, GADD45B, ITGA6A, MYBL2 and CEACAM5, ITGA6, MACC1 and B2M, or S100A4, GADD45B, MYBL2, MYC and CEACAM5, ITGA6, MACC1 and B2M, or S100A4, GADD45B, ITGA6A, MYC and CEACAM5, ITGA6, MACC1 and B2M, or PTGS2, GADD45B, ITGA6A, MYBL2 and CEACAM5, ITGA6, MACC1 and B2M, or PTGS2, GADD45B, MYBL2, MYC and CEACAM5, ITGA6, MACC1 and B2M, or PTGS2, GADD45B, ITGA6A, MYC and CEACAM5, ITGA6, MACC1 and B2M, or PTGS2, ITGA6A, MYBL2, MYC and CEACAM5, ITGA6, MACC1 and B2M, or S100A4, GADD45B, ITGA6A, MYBL2 and CEACAM5, ITGA6, MACC1 and B2M, or S100A4, GADD45B, ITGA6A, MYC and CEACAM5, ITGA6, MACC1 and B2M, or GADD45B, ITGA6A, MYBL2, MYC and CEACAM5, ITGA6, MACC1 and B2M, or.

[0096] In some embodiments, the SSD or TEP can include the mRNA expression level of five markers detectable in stool samples and enriched in samples with CRC and two markers detectable in stool samples and enriched in samples with CRC and AA. Such embodiments can include PTGS2, S100A4, GADD45B, ITGA6A, MYBL2 and CEACAM5 and ITGA6, or PTGS2, S100A4, GADD45B, ITGA6A, MYC andCEACAM5 and ITGA6, or S100A4, GADD45B, ITGA6A, MYBL2, MYC andCEACAM5 and ITGA6, or PTGS2, GADD45B, ITGA6A, MYBL2, MYC andCEACAM5 and ITGA6, or PTGS2, S100A4, ITGA6A, MYBL2, MYC and CEACAM5 and ITGA6, or PTGS2, S100A4, GADD45B, MYBL2, MYC and CEACAM5 and ITGA6, or PTGS2, S100A4, GADD45B, ITGA6A, MYBL2 and CEACAM5 and MACC1, or PTGS2, S100A4, GADD45B, ITGA6A, MYC and CEACAM5 and MACC1, or S100A4, GADD45B, ITGA6A, MYBL2, MYC and CEACAM5 and MACC1, or PTGS2, GADD45B, ITGA6A, MYBL2, MYC and CEACAM5 and MACC1, or PTGS2, S100A4, ITGA6A, MYBL2, MYC and CEACAM5 and MACC1, or PTGS2, S100A4, GADD45B, MYBL2, MYC and CEACAM5 and MACC1, or PTGS2, S100A4, GADD45B, ITGA6A, MYBL2 and CEACAM5 and B2M, or PTGS2, S100A4, GADD45B, ITGA6A, MYC and CEACAM5 and B2M, or S100A4, GADD45B,ITGA6A, MYBL2, MYC and CEACAM5 and B2M, or PTGS2, GADD45B, ITGA6A, MYBL2, MYC and CEACAM5 and B2M, or PTGS2, S100A4, ITGA6A, MYBL2, MYC and CEACAM5 and B2M, or PTGS2, S100A4, GADD45B, MYBL2, MYC and CEACAM5 and B2M, or PTGS2, S100A4, GADD45B, ITGA6A, MYBL2 and ITGA6 and MACC1, or PTGS2, S100A4, GADD45B, ITGA6A, MYC and ITGA6 and MACC1, or S100A4, GADD45B, ITGA6A, MYBL2, MYC and ITGA6 and MACC1, or PTGS2, GADD45B, ITGA6A, MYBL2, MYC and ITGA6 and MACC1, or PTGS2, S100A4, ITGA6A, MYBL2, MYC and ITGA6 and MACC1, or PTGS2, S100A4, GADD45B, MYBL2, MYC and ITGA6 and MACC1, or PTGS2, S100A4, GADD45B, ITGA6A, MYBL2 and ITGA6 and B2M, or PTGS2, S100A4, GADD45B, ITGA6A, MYC and ITGA6 and B2M, or S100A4, GADD45B, ITGA6A, MYBL2, MYC and ITGA6 and B2M, or PTGS2, GADD45B, ITGA6A, MYBL2, MYC and ITGA6 and B2M, or PTGS2, S100A4, ITGA6A, MYBL2, MYC and ITGA6 and B2M, or PTGS2, S100A4, GADD45B, MYBL2, MYC and ITGA6 and B2M, or PTGS2, S100A4, GADD45B, ITGA6A, MYBL2 and MACC1 and B2M, or PTGS2, S100A4, GADD45B, ITGA6A, MYC and MACC1 and B2M, or S100A4, GADD45B, ITGA6A, MYBL2, MYC and MACC1 and B2M, or PTGS2, GADD45B, ITGA6A, MYBL2, MYC and MACC1 and B2M, or PTGS2, S100A4, ITGA6A, MYBL2, MYC and MACC1 and B2M, or PTGS2, S100A4, GADD45B, MYBL2, MYC and MACC1 and B2M.

[0097] In some embodiments, the SSD or TEP can include the mRNA expression level of five markers detectable in stool samples and enriched in samples with CRC and three markers detectable in stool samples and enriched in samples with CRC and AA. Such embodiments can include PTGS2, S100A4, GADD45B, ITGA6A, MYBL2 and CEACAM5, ITGA6 and MACC1, or PTGS2, S100A4, GADD45B, ITGA6A, MYC and CEACAM5, ITGA6 and MACC1, or S100A4, GADD45B, ITGA6A, MYBL2, MYC and CEACAM5, ITGA6 and MACC1, or PTGS2, GADD45B, ITGA6A, MYBL2, MYC and CEACAM5, ITGA6 and MACC1, or PTGS2, S100A4, ITGA6A, MYBL2, MYC and CEACAM5, ITGA6 and MACC1, or PTGS2, S100A4, GADD45B, MYBL2, MYC and CEACAM5, ITGA6 and MACC1, or PTGS2, S100A4, GADD45B, ITGA6A, MYBL2 and CEACAM5, ITGA6 and B2M, or PTGS2, S100A4, GADD45B, ITGA6A, MYC and CEACAM5, ITGA6 and B2M, or S100A4, GADD45B, ITGA6A, MYBL2, MYC and CEACAM5, ITGA6 and B2M, or PTGS2, GADD45B, ITGA6A, MYBL2, MYC and CEACAM5, ITGA6 and B2M, or PTGS2, S100A4, ITGA6A, MYBL2, MYC and CEACAM5, ITGA6 and B2M, or PTGS2, S100A4, GADD45B, MYBL2, MYC andCEACAM5, ITGA6 and B2M, or PTGS2, S100A4, GADD45B, ITGA6A, MYBL2 and ITGA6, MACC1 and B2M, or PTGS2, S100A4, GADD45B, ITGA6A, MYC and ITGA6, MACC1 and B2M, or S100A4, GADD45B, ITGA6A, MYBL2, MYC and ITGA6, MACC1 and B2M, or PTGS2, GADD45B, ITGA6A, MYBL2, MYC and ITGA6, MACC1 and B2M, or PTGS2, S100A4, ITGA6A, MYBL2, MYC and ITGA6, MACC1 and B2M, or PTGS2, S100A4, GADD45B, MYBL2, MYC and ITGA6, MACC1 and B2M, or PTGS2, S100A4, GADD45B, ITGA6A, MYBL2 and CEACAM5, MACC1 and B2M, or PTGS2, S100A4, GADD45B, ITGA6A, MYC and CEACAM5, MACC1 and B2M, or S100A4, GADD45B, ITGA6A, MYBL2, MYC and CEACAM5, MACC1 and B2M, or PTGS2, GADD45B, ITGA6A, MYBL2, MYC and CEACAM5, MACC1 and B2M, or PTGS2, S100A4, ITGA6A, MYBL2, MYC and CEACAM5, MACC1 and B2M, or PTGS2, S100A4, GADD45B, MYBL2, MYC and CEACAM5, MACC1 and B2M.

[0098] In some embodiments, the SSD or TEP can include the mRNA expression level of five markers detectable in stool samples and enriched in samples with CRC and four markers detectable in stool samples and enriched in samples with CRC and AA. Such embodiments can include PTGS2, S100A4, GADD45B, ITGA6A, MYBL2 and CEACAM5, ITGA6, MACC1 and B2M, or PTGS2, S100A4, GADD45B, ITGA6A, MYC and CEACAM5, ITGA6, MACC1 and B2M, or S100A4, GADD45B, ITGA6A, MYBL2, MYC and CEACAM5, ITGA6, MACC1 and B2M, or PTGS2, GADD45B, ITGA6A, MYBL2, MYC and CEACAM5, ITGA6, MACC1 and B2M, or PTGS2, S100A4, ITGA6A, MYBL2, MYC and CEACAM5, ITGA6, MACC1 and B2M, or PTGS2, S100A4, GADD45B, MYBL2, MYC and CEACAM5, ITGA6, MACC1 and B2M.

[0099] In some embodiments, the SSD or TEP can include the mRNA expression level of six markers detectable in stool samples and enriched in samples with CRC and two markers detectable in stool samples and enriched in samples with CRC and AA. Such embodiments can include PTGS2, S100A4, GADD45B, ITGA6A, MYBL2, MYC and CEACAM5 and ITGA6, or PTGS2, S100A4, GADD45B, ITGA6A, MYBL2, MYC and CEACAM5 and MACC1, or PTGS2, S100A4, GADD45B, ITGA6A, MYBL2, MYC and CEACAM5 and B2M, or PTGS2, S100A4, GADD45B, ITGA6A, MYBL2, MYC and ITGA6 and MACC1, or PTGS2, S100A4, GADD45B, ITGA6A, MYBL2, MYC and ITGA6 and B2M, or PTGS2, S100A4, GADD45B, ITGA6A, MYBL2, MYC and MACC1 and B2M.

[0100] In some embodiments, the SSD or TEP can include the mRNA expression level of six markers detectable in stool samples and enriched in samples with CRC and three markers detectable in stool samples and enriched in samples with CRC and AA. Such embodiments can include PTGS2, S100A4, GADD45B, ITGA6A, MYBL2 and MYC and CEACAM5, ITGA6 and MACC1, or PTGS2, S100A4, GADD45B, ITGA6A, MYBL2 and MYC and CEACAM5, ITGA6 and B2M, or PTGS2, S100A4, GADD45B, ITGA6A, MYBL2 and MYC and ITGA6, MACC1 and B2M, or PTGS2, S100A4, GADD45B, ITGA6A, MYBL2 and MYC and CEACAM5, MACC1 and B2M.

[0101] In an embodiment, the SSD or TEP can include the mRNA expression level of six markers detectable in stool samples and enriched in samples with CRC and four markers detectable in stool samples and enriched in samples with CRC and AA. Such embodiments can include PTGS2, S100A4, GADD45B, ITGA6A, MYBL2 and MYC and CEACAM5, ITGA6, MACC1 and B2M.

[0102] In some embodiments, the SSD or TEP can include the mRNA expression level of at least one, two, three, four or five level of mRNA expression the CEACAM5 gene, the ITGA6 gene, the MACC1 gene, the PTGS2 gene and / or the S100A4 gene.

[0103] In some embodiments, the mRNA expression levels of the at least two distinct genes that comprise SSD or a TEP comprise a Ct value, a ACt value (compared to, e.g., expression of a housekeeping gene), an absolute amount of expression, or combination thereof. In some embodiments of the method, two or more of a CEACAM5 expression value, a PTGS2 expression value, a S100A4 expression value, a MACC1 expression value, and a ITGA6 expression value is utilized for the SSD or TEP. In embodiments, a B2M expression value is also utilized for the SSD or TEP. In some embodiments, at least one of CEACAM5, PTGS2, S100A4, MACC1, ITGA6, and B2M utilizes a ACt, a Ct value, or an absolute amount of expression value for the SSD or to determine the TEP. In some embodiments comprising PTGS2, ITGA6, or both, a ACt value is used to determine the SSD or TEP. In some embodiments comprising CEACAM5, a Cp value is used to determine the SSD or TEP.

[0104] In an embodiment, CEACAM5 utilizes a Cp expression value, PTGS2 utilizes a ACt expression value, S100A4 utilizes an absolute amount expression value (e.g., copies / pL), MACC1 utilizes an absolute amount expression value (e.g., copies / pL) expression value, and a ITGA6 utilizes a ACt expression value for the SSD or TEP. In some embodiments, morethan one expression value type can be used for a target mRNA for a SSD or TEP. For example, in embodiments, a ACt expression value and an absolute value can be utilized for PTGS2 for a SSD or TEP. In an embodiment, CEACAM5 utilizes a Cp expression value, PTGS2 utilizes a ACt expression value and an absolute amount expression value (e.g., copies / pL), S100A4 utilizes an absolute amount expression value (e.g., copies / pL), MACC1 utilizes an absolute amount expression value (e.g., copies / pL) expression value, and a ITGA6 utilizes a ACt expression value for the SSD or TEP.

[0105] The SSD or TEP comprises the mRNA expression level of a combination of at least two of any of the genes described herein. In some embodiments, the SSD or TEP comprises the mRNA expression level of a combination of at least three of any of the genes described herein. In some embodiments, the SSD or TEP comprises the mRNA expression level of a combination of at least four of any of the genes described herein. In some embodiments, the SSD or TEP comprises the mRNA expression level of a combination of at least five of any of the genes described herein. In some embodiments, the SSD or TEP comprises the mRNA expression level of a combination of at least six of any of the genes described herein. In some embodiments, the SSD or TEP comprises the mRNA expression level of a combination of at least seven of any of the genes described herein. In some embodiments, the SSD or TEP comprises the mRNA expression level of a combination of at least eight of any of the genes described herein. In some embodiments, the SSD or TEP comprises the mRNA expression level of a combination of at least nine of any of the genes described herein.

[0106] In some embodiments, the SSD or TEP comprises the mRNA expression level of a combination of at least ten of any of the genes described herein. In some embodiments, the SSD or TEP comprises the mRNA expression level of a combination of at least eleven of any of the genes described herein. In some embodiments, the SSD or TEP comprises the mRNA expression level of a combination of at least twelve of any of the genes described herein.

[0107] Once the TEP has been obtained, it is compared to a control expression profile (CEP). The control expression profile comprises the mRNA expression level of the at least two genes who are present on the test expression level. The control expression profile can be obtained or derived from one or more mRNA transcripts and / or cells (such as, for example, one or more epithelial cells and, in further embodiments, one or more colorectal epithelial cell) from a control subject which is known not to experience an advancedadenoma or a colorectal cancer (e.g., a healthy control subject). The control expression profile can be obtained or derived from one or more mRNA transcripts and / or cells from a control subject having a non-advanced adenoma and lacking an advanced adenoma or a colorectal cancer. The control expression profile can be obtained or derived from one or more RNA transcripts and / or one or more cells from a healthy (e.g., non cancerous) tissue from a subject which may, in some embodiments, have an advanced adenoma or a colorectal cancer. In some embodiments, the control subject can be aged- and gender- matched with the subject whose risk is being stratified. In some embodiments, the control expression profile is obtained or derived from a plurality of control subjects. In some further embodiments, the method can include determining the mRNA expression profile of at least two genes from the control subject to provide the control expression profile.

[0108] The control expression profile can include the mRNA expression level of the CEA adhesion molecule 5 (also referred to as CEACAM5, CD66e or CEA and having the Gene ID 1048). The control expression profile can include the mRNA expression level of the growth arrest and DNA damage inducible beta gene (also referred to as GADD45B, GADD45BETA or MYD118 and having the Gene ID: 4616). The control expression profile can include the mRNA expression level of the integrin subunit alpha 2 gene (also referred to as ITGA2, BR, CD49B, GPla, HPA-5, VLA-2 or VLAA2 and having the Gene ID: 3673). The control expression profile can include the mRNA expression level of the MET transcriptional regulator MACC 1 (also referred to as MACC 1 , 7 A5 or SH3BP4L and having the Gene ID: 346389). The control expression profile can include the mRNA expression level of the MYB proto-oncogene like 2 gene (also referred to as MYBL2, B- MYB or BMYB and having the Gene ID: 4605). The control expression profile can include the mRNA expression level of the MYC proto-oncogene, bHLH transcription factor (also referred to as MYC, MRTL, MYCC, bHLHe39 or c-Myc and having the Gene ID: 4609). The control expression profile can include the mRNA expression level of the S100 calcium binding protein A4 (also referred to as S100A4, 18A2, 42A, CAPL, FSP1, MTS1, P9KA or PEL98 and having Gene ID: 6275). In aspecific embodiment, the control expression profile can include the mRNA expression profile of CEACAM5, ITGA6 and MACC1. In yet another specific embodiment, the control expression profile can include the mRNA expression profile of CEACAM5, ITGA6, MACC1 and B2M. In still yet another embodiment, the control expression profile can include the mRNA expression profile of PTGS2 and S100A4. In still yet another embodiment, the control expression profile can include the mRNA expression profile of CEACAM5, ITGA6, MACC1, PTGS2and S100A4, optionally in combination with the mRNA expression profile of B2M.

[0109] In some optional embodiments, the control expression profile can include the mRNA expression level of the integrin subunit alpha 6 gene (also referred to as ITGA6, CD49f, ITGA6B or VLA-6 and having the Gene ID: 3655). In such embodiments, it is possible that the control expression profile can include the mRNA expression level of the alpha- and / or beta-isoform of ITGA6 gene transcript. In some optional embodiments, the control expression profile can include the mRNA expression level of the prostaglandinendoperoxide synthase 2 gene (also referred to as PTGS2, COX-2, COX2, GRIPGHS, PGG / HS, PGHS-2, PHS-2 or hCox-2 and having the Gene ID: 5743).

[0110] In some further optional embodiments, the control expression profile can include the mRNA expression levels of the beta-2-microglobulin gene (also referred to as B2M or IMD43 and having the Gene ID 567). In some further optional embodiments, the control expression profile can include the mRNA expression level of the integrin subunit alpha 1 gene (also referred to as ITGA1, CD49a or VLA1 and having the Gene ID: 3672).

[0111] In some specific embodiments, the control expression level can include at least one, two, three, four or five levels of mRNA expression the S 100A4 gene, the GADD45B gene, the ITGA6 gene, the MYBL2 gene, the MYC gene and / or the PTGS2 gene. The control expression profile comprises the mRNA expression level of the genes which are also reported on the TEP. In some embodiments, the control expression profile comprises the mRNA expression level of a combination of at least two of any of the genes described herein. In some embodiments, the control expression profile comprises the mRNA expression level of a combination of at least three of any of the genes described herein. In some embodiments, the control expression profile comprises the mRNA expression level of a combination of at least four of any of the genes described herein. In some embodiments, the control expression profile comprises the mRNA expression level of a combination of at least five of any of the genes described herein. In some embodiments, the control expression profile comprises the mRNA expression level of a combination of at least six of any of the genes described herein. In some embodiments, the control expression profile comprises the mRNA expression level of a combination of at least seven of any of the genes described herein. In some embodiments, the control expression profile comprises the mRNA expression level of a combination of at least eight of any of the genes described herein. In some embodiments, the control expression profile comprises the mRNA expression level of a combination of at least nine of any of the genesdescribed herein. In some embodiments, the control expression profile comprises the mRNA expression level of a combination of at least ten of any of the genes described herein. In some embodiments, the control expression profile comprises the mRNA expression level of a combination of at least eleven of any of the genes described herein. In some embodiments, the control expression profile comprises the mRNA expression level of a combination of at least twelve of any of the genes described herein.

[0112] For the stratification methods, a comparison is then conducted to determine if the mRNA expression levels present in the TEP are different than the corresponding mRNA expression levels present in the control expression profile. This comparison can be performed on a gene-by-gene basis. For example, if the TEP comprises the mRNA expression level of the CEACAM5 gene, such expression level is compared to the mRNA expression level of the CEACAM5 in the control expression profile. If it is determined that the mRNA expression levels of at least two genes in the TEP is higher than the mRNA expression levels of the same two genes in the control expression profile, this is indicative that the stratified subject has an increased risk of having an advanced adenoma or a colorectal cancer when compared to the control subject. If it is determined that the mRNA expression levels of at least two genes in the control expression profile is lower than the mRNA expression levels of the same two genes in the TEP, this is indicative that the stratified subject has an increased risk of having an advanced adenoma or a colorectal cancer when compared to the control subject.

[0113] For the computer-implemented methods, the second determining step includes determining, by a computing system executing a machine-learning model and based on the stool sample data, a diagnostic code, wherein the machine-learning model was trained using stool diagnostic training data including a plurality of stool sample data labeled with at least one of two or more corresponding diagnostic codes.

[0114] Referring to the machine-learning model, many types of functionality and neural networks can be employed to perform functions of the machine-learning model. The machinelearning model can operate according to machine-learning tasks as classified into several categories, including supervised learning, semi-supervised learning, and unsupervised learning. In supervised learning, the machine-learning model builds a mathematical model from a set of data that contains both the inputs and the desired outputs. The set of data is sample data known as the “stool diagnostic training data” or simply “training data”, in order to make predictions or decisions without being explicitly programmed to perform the task.

[0115] In semi-supervised learning, the machine-learning model develops mathematical models from incomplete training data, where a portion of the sample input does not have labels. A classification model can then be used when the outputs are restricted to a limited set of values.

[0116] In unsupervised learning, the machine-learning model builds a mathematical model from a set of data that contains only inputs and no desired output labels. Unsupervised learning models are used to find structure in related training data, such as grouping or clustering of data points. Unsupervised learning can discover patterns in data and can group the inputs into categories.

[0117] Alternative machine-learning models may be used to learn and classify types of data to consider for generating the diagnostic codes, such as deep learning methods such as Artificial Neural Networks (ANNs), Convolutional Neural Networks (CNNs), Recursive Neural Networks, Deep Boltzmann Machines (DBMs), Deep Belief Networks (DBNs), Stacked AutoEncoders, and other modeling techniques derived from a neural network framework, or generative models. Deep machine-learning may use neural networks to analyze prior test results through a collection of interconnected processing nodes. The connections between the nodes may be dynamically weighted. Neural networks learn relationships through repeated exposure to data and adjustment of internal weights. Neural networks may capture nonlinearity and interactions among independent variables without pre-specification. Whereas traditional regression analysis requires that nonlinearities and interactions be detected and specified manually, neural networks perform the tasks automatically.

[0118] Additionally, or alternatively, contemplated iterative modeling processes include Turing machines, evolutionary programming methods, including genetic models and genetic programming (e.g., tree-based genetic programming, stack-based genetic programming, linear (including machine code) genetic programming, grammatical evolution, extended compact genetic programming (ECGP), embedded Cartesian genetic programming (ECGP), probabilistic incremental program evolution (PIPE), and strongly typed genetic programming (STGP)). Other evolutionary programming methods include gene expression programming, evolution strategy, differential evolution, neuroevolution, learning classifier systems, or reinforcement learning systems, where solution is a set of classifiers (rules or conditions) that can be binary, real, neural net, or S-expression types. In the case of learning classifier systems, fitness may be determined with either a strength or accuracy based reinforcement learning orsupervised learning approach.

[0119] Additional or alternative contemplated iterative modeling processes may include Monte Carlo methods, Markov chains, stepwise linear and logistical regression, decision trees, Random Forests, Support Vector Machines, Bayesian modeling techniques, or Gradient- Boosting techniques, so long as the process includes a repeatable or loop-able subroutine or process (e.g. a run, a for loop, an epoch, a cycle).

[0120] Still other machine-learning models or functions can be implemented to generate the diagnostic codes, such as any number of classifiers that receives input parameters and outputs a classification (e.g., diagnostic code). Any semi-supervised, supervised, or unsupervised learning may be used, particularly iterative modeling processes, i.e., modeling method to describe the relationship between predictors and outcomes in target datasets that includes a repeatable or loop-able subroutine or process (e.g. a run, a for loop, an epoch, a cycle).

[0121] Stool diagnostic training data (SDTD) may be used in raw form (unadjusted, absolute), or normalized (e.g., to housekeeper gene expression) by the machine-learning model during its training or model building. The machine-learning model of the disclosure can also transform the SDTD, for example, with computational operators (e.g., logical statements like IF, AND, OR), mathematical operators (e.g., arithmetic operations like multiplication, division, subtraction, and addition; trigonometric operations; logistic functions; calculus operations; “floor” or “ceiling” operators; or any other mathematical operators), constants (e.g., a constant numerical value, including integers or values like pi), a predictor (e.g., observed or measured values or formulas), features (e.g., characteristics), variables, ternary operators (e.g., an operator that takes three arguments where the first argument is a comparison argument, the second is the result upon a true comparison, and the third is the result upon a false comparison), algorithms, formulas, literals, functions (e.g., unary functions, binary functions, etc.), binary operators (e.g., an operator that operates on two operands and manipulates them to return a result), weights and weight vectors, nodes and hidden nodes, gradient descent, sigmoidal activation functions, hyper-parameters, and biases, any or all of which may be used to train with the stool diagnostic training data labeled with two or more diagnostic codes to build the machine-learning model useful to determine an appropriate diagnostic code for stool sample data from a subject. SSD is transformed by any of the operations above according to the trained machine-learning model.

[0122] The machine-learning model may thus be considered an application of rules in combination with learning from prior data to identify appropriate outputs. Analyzing stool diagnostic training data allows the machine-learning model to learn patterns of stool sample data and appropriate outputs (diagnostic codes) to then apply to tool sample data from a subject.

[0123] In some embodiments, the computer-implemented methods further comprise providing a representation of a graphical user interface for display on a computing system; receiving an input from the graphical user interface, the input comprising the stool sample data; and after the diagnostic code is determined by the machine-learning model, providing, for display on the graphical user interface, the diagnostic code.

[0124] As indicated above, the methods described herein can also be used to stratify the risk of the subject of having an advanced adenoma. In such embodiments, the SSD, TEP, or CEP include the mRNA expression levels of the CEACAM5 gene, the ITGA6 gene and / or the MACC1 gene. In some additional embodiments, the SSD, TEP, or CEP include the mRNA expression levels of the CEACAM5 gene, the ITGA6 gene and the MACC1 gene. The method can include determining the mRNA expression level of any one of the CEAC AM5 gene, the ITGA6 gene and / or the MACC 1 gene in the subj ect who is to receive a diagnostic code or who is being stratified and / or the control subject (for a CEP). In some specific embodiments, the method can include determining the mRNA expression level of the CEACAM5 gene, the ITGA6 gene and the MACC1 gene in the subject who is to receive a diagnostic code or who is being stratified and / or the control subject. The method can include comparing the mRNA expression level of the CEACAM5 gene, the ITGA6 gene and / or the MACC1 gene between the test and the control expression profiles. In some embodiments, the method can include comparing the mRNA expression level of the CEACAM5 gene, the ITGA6 gene and the MACC1 gene between the test and the control expression profiles for the stratification methods, or include the mRNA expression level of the CEACAM5 gene, the ITGA6 gene and the MACC1 gene for the SSD for determination by a trained machine-learning model of a diagnostic code. If it has been determined that the mRNA expression level of at least one, at least two or all three genes (e.g., the CEACAM5 gene, the ITGA6 gene and / or the MACC1 gene) present in the test expression profile is increased with respect the corresponding mRNA expression level in the control expression profile, it is indicative that the stratified subject has an increased risk, with respect to the control subject, of having an advanced adenoma. If it has been determined that the mRNA expression level of at least one, at least two or all three genes(e.g., the CEACAM5 gene, the ITGA6 gene and / or the MACC1 gene) present in the control expression profile is decreased with respect the corresponding mRNA expression level in the test expression profile, it is indicative that the stratified subject has an increased risk, with respect to the control subject, of having an advanced adenoma. For an example machine-learning model-related method, if a subject’s SSD includes the mRNA expression level of at least one, at least two or all three genes (e.g., the CEACAM5 gene, the ITGA6 gene and / or the MACC1 gene) that are increased in comparison to the stool diagnostic training data, the subject is more likely to receive a diagnostic code of AA, and if decreased in comparison to the stool diagnostic training data, the subject is less likely to receive a diagnostic code of AA.

[0125] As indicated above, the methods described herein can also be used to stratify the risk of the subject of having a colorectal cancer. In such embodiments, the SSD, TEP, or CEP include the mRNA expression levels of the S100A4 gene, the GADD45B gene, the ITGA6 gene, the MYBL2 gene, the MYC gene and / or the PTGS2 gene. In some additional embodiments, the SSD, TEP, or CEP include the mRNA expression levels of the S100A4 gene, the GADD45B gene, the ITGA6 gene, the MYBL2 gene, the MYC gene and the PTGS2 gene. The method can include determining the mRNA expression level of any one of the S100A4 gene, the GADD45B gene, the ITGA6 gene, the MYBL2 gene, the MYC gene and / or the PTGS2 gene in the subject who is to receive a diagnostic code or who is being stratified and / or the control subject. In some specific embodiments, the method can include determining the mRNA expression level of the S100A4 gene, the GADD45B gene, the ITGA6 gene, the MYBL2 gene, the MYC gene and the PTGS2 gene in the subj ect who is to receive a diagnostic code or who is being stratified and / or the control subject. The method can include comparing the mRNA expression level of the S100A4 gene, the GADD45B gene, the ITGA6 gene, the MYBL2 gene, the MYC gene and / or the PTGS2 gene between the test and the control expression profiles for the stratification methods, or include the mRNA expression level of the S 100A4 gene, the GADD45B gene, the ITGA6 gene, the MYBL2 gene, the MYC gene and / or the PTGS2 gene for the SSD for determination by a trained machine-learning model of a diagnostic code. In some embodiments, the method can include comparing the mRNA expression level of the S100A4 gene, the GADD45B gene, the ITGA6 gene, the MYBL2 gene, the MYC gene and / or the PTGS2 gene between the test and the control expression profiles for the stratification methods, or include the same genes for the SSD for determination by a trained machine-learning model of a diagnostic code. If it has been determined that themRNA expression level of at least one, at least two, at least three, at least four or all five genes (e.g., the S100A4 gene, the GADD45B gene, the ITGA6 gene, the MYBL2 gene, the MYC gene and / or the PTGS2 gene) present in the test expression profile is increased with respect the corresponding mRNA expression level in the control expression profile, it is indicative that the stratified subject has an increased risk, with respect to the control subject, of having a colorectal cancer. If it has been determined that the mRNA expression level of at least one, at least two, at least three, at least four or all five genes (e.g., the S100A4 gene, the GADD45B gene, the ITGA6 gene, the MYBL2 gene, the MYC gene and / or the PTGS2 gene) present in the control expression profile is decreased with respect the corresponding mRNA expression level in the test expression profile, it is indicative that the stratified subject has an increased risk, with respect to the control subject, of having a colorectal cancer. For an example machine-learning model-related method, if a subject’s SSD includes the mRNA expression level of at least one, at least two, at least three, at least four or all five genes (e.g., the S100A4 gene, the GADD45B gene, the ITGA6 gene, the MYBL2 gene, the MYC gene and / or the PTGS2 gene) that are increased in comparison to the stool diagnostic training data, the subject is more likely to receive a diagnostic code of CRC, and if decreased in comparison to the stool diagnostic training data, the subject is less likely to receive a diagnostic code of CRC.

[0126] The methods described herein can be used in conjunction with other methods and assays used to aid the diagnosis of advanced adenoma or colorectal cancer. It is recognized in the art that the presence of hemoglobin, a component of blood, in the stool is present more frequently in subjects having an advanced adenoma or a colorectal cancer than in subjects having a non-advanced adenoma or healthy subjects. As such, the methods described herein can be used in combination with the determination of the presence of hemoglobin to increase the sensitivity of the methods. The methods described herein can thus include determining the presence of hemoglobin in the stool sample or the processed stool sample. In some embodiments, the method can include performing a fecal immunochemical test (FIT; e.g., SENTiFIT® FOB Gold® test, IDK® TurbiFIT™ or IDK® Hemoglobin ELISA) to determine the presence or absence of hemoglobin in the stool sample. In some further embodiments, the method can include performing a guaiac-based fecal occult blood test (gFOBT). In some additional embodiments, the method can include performing the Cologuard™ and / or the ColoAlert™ test. One of the strengths of combining the method of the present disclosure based on multitarget mRNA with other assays is the higher level of detection of advanced adenomas with mRNA targets ascompared to other non-invasive methods including FIT or gFOBT, Cologuard™ and / or Colo Alert™.

[0127] The subject who is to receive a diagnostic code or who is being stratified may have previously been determined with the presence of an advanced adenoma or a colorectal cancer. F or example, the presence of hemoglobin in the stool of the subj ect who is to receive a diagnostic code or who is being stratified may have previously been determined. In some embodiments, the subject may previously have been submitted to a FIT or a gFOBT tests to determine the presence of hemoglobin in a stool sample (which may be the same or different from the one used to determine the mRNA expression levels). In some embodiments, the subject may previously had been determined to have hemoglobin in his / her stool.

[0128] In some embodiments, the method includes a SSD, TEP, or CEP including CEACAM5 that utilizes a Cp expression value, PTGS2 that utilizes a ACt expression value, S100A4 that utilizes an absolute amount expression value (e.g., copies / pL), MACC1 that utilizes an absolute amount expression value (e.g., copies / pL) expression value, and ITGA6 that utilizes a ACt expression value; and further includes a FIT test to determine the presence of hemoglobin in the stool sample. In some embodiments, more than one expression value type can be used for a target mRNA for a S SD, TEP, or CEP. For example, in embodiments, a ACt expression value and an absolute value can be utilized for PTGS2 for a S SD, TEP, or CEP. In an embodiment, the method includes a S SD, TEP, or CEP including CEACAM5 that utilizes a Cp expression value, PTGS2 that utilizes a ACt expression value and an absolute amount expression value (e.g., copies / pL), S100A4 that utilizes an absolute amount expression value (e.g., copies / pL), MACC1 that utilizes an absolute amount expression value (e.g., copies / pL) expression value, and ITGA6 that utilizes a ACt expression value; and further includes a FIT test.

[0129] It is also recognized that some DNA mutations increase the predisposition of a subject to develop an advanced adenoma or a colorectal cancer. As such, the methods described herein can be used in combination with the determination of the presence or the absence of one or more DNA mutation in the genome of cells of the subjects to increase the sensitivity of the methods described herein. The methods described herein can thus include determining the presence or absence of at least one DNA mutation (which may be, for example, a deletion, an insertion and / or a duplication) in the genome of the subjects, wherein the at least one DNA mutation is associated with an increase in thepredisposition of developing an advanced adenoma or a colorectal cancer. For example, the DNA mutation can be located in the KRAS gene, and / or the BRAF gene. In some embodiments, the method can include performing the Cologuard™ or ColoAlert™ test to determine the presence of the at least one DNA mutation in the cells of the subject. The subject who is to receive a diagnostic code or who is being stratified may have previously been diagnosed with the presence of an advanced adenoma or a colorectal cancer. For example, the presence of at least one DNA mutation in the cell(s) of the subject may have previously been determined.

[0130] In some embodiments, the methods includes a S SD or TEP including CEACAM5 that utilizes a Cp expression value, PTGS2 that utilizes a ACt expression value, S100A4 that utilizes an absolute amount expression value (e.g., copies / pL), MACC1 that utilizes an absolute amount expression value (e.g., copies / pL) expression value, and ITGA6 that utilizes a ACt expression value; and further includes (1) a FIT test and (2) determining the presence of a DNA mutation (e.g., KRAS or BRAF). In some embodiments, more than one expression value type can be used for a target mRNA for a test expression profile. For example, in embodiments, both a ACt expression value and an absolute value can be utilized for PTGS2 for a test expression profile. In an embodiment, the method includes a test expression profile including CEACAM5 that utilizes a Cp expression value, PTGS2 that utilizes a ACt expression value and an absolute amount expression value (e.g., copies / pL), S100A4 that utilizes an absolute amount expression value (e.g., copies / pL), MACC1 that utilizes an absolute amount expression value (e.g., copies / pL) expression value, and ITGA6 that utilizes a ACt expression value; and further includes (1) a FIT test and (2) determining the presence of a DNA mutation (e.g., KRAS or BRAF). In embodiments, the method further includes determining aberrant DNA methylation pattern (e.g., NDRG4 gene and / or BMP3).

[0131] The subject who is to receive a diagnostic code or who is being stratified may have previously been diagnosed with the presence of an advanced adenoma or a colorectal cancer. For example, the presence of at least one DNA mutation in the cell(s) of the subject may have previously been determined. In some embodiments, the subject may previously have been submitted to a Cologuard™ or ColoAlert™ test to determine the presence of the DNA mutation in the cell from the subject.

[0132] It is also recognized that advanced adenoma and malignant colorectal tumors can be visualized in situ and help the physician in determining if a subject has an advancedadenoma or a colorectal cancer. As such, the methods described herein can be used in combination with the visualization of part of the subject's colorectal tract. The colorectal tract can be visualized using a colonoscopy, a flexible sigmoidoscopy and / or a CT coIonography. In some embodiments, the colorectal tract can be visualized using a colonoscopy. The methods described herein can thus include a step of visualizing part or the entire subject's colorectal tract to determine the presence of advanced adenoma(s) and / or malignant colorectal tumor(s).

[0133] In some embodiments, the subject may have previously been submitted to a visualization of part or all of its colorectal tract.

[0134] Alternatively, the methods described herein can be used prior to the visualization of the subject's colorectal tract. For example, the methods described herein can be used to prioritize subjects being stratified in the high risk group or the CRC subgroup of being submitted to imaging, such as a colonoscopy. Imaging techniques are uncomfortable for some subjects or can be of limited availability in some geographical areas. As such, there may be, under certain circumstances, a need to prioritize subjects which would benefit from such imaging analysis because they are in the high risk group or the CRC subgroup. In some embodiments, the methods include recommending to the subj ect who receives a diagnostic code for CRCor who has been stratified in the high risk group or the CRC subgroup of having an imaging analysis of their colorectal tract (such as a colonoscopy) performed to aid the physician in his / her diagnosis.

[0135] The methods described herein can also be used in the context of a clinical trial to include or exclude subjects from a clinical study or to attribute them to a treatment arm of the clinical study.

[0136] It is also recognized that advanced adenoma and malignant colorectal tumors can be detected in a pathology analysis (such as, for example, a histology analysis) and help the physician in determining if a subject has an advanced adenoma or a colorectal cancer. As such, the methods described herein can be used in combination with a pathological analysis of a tissue of a subject suspected of being an advanced adenoma or a malignant colorectal tumor. In some embodiments, the tissue of the subject may have previously been submitted to pathological analysis.

[0137] Having received a diagnostic code or having been stratified to a particular group or a particular subgroup, the subject may receive a tailored therapeutic regimen that is suitableto alleviate the symptoms or treat the condition that subject has been assigned to. As such, the methods described herein can be used in the treatment of an advanced adenoma or a colorectal cancer (such as a colon cancer or a rectal cancer) in subjects who have been stratified in the high risk group or who have received a CRC diagnostic code. The treatment can include submitting the subject to one or more rounds of chemotherapy (e.g., 5- fluorouracil, leucovorin, capeci tabine, irinotecan and / or oxaliplatin) optionally in combination with therapeutic antibodies. The treatment can include submitting the subject to one or more rounds of radiation therapy. The treatment can include submitted the subj ect to surgery (e.g., surgical resection of the advanced adenoma or the malignant colorectal tumor).

[0138] The methods described herein can be used to tailor the treatment regimen of a subject which has received at least one dose of chemotherapy, at least one dose or radiotherapy and / or has already been submitted to a surgery to remove one or more advanced adenoma or one or more colorectal malignant tumor. In such embodiments, the methods can be used to determine if the subject receives a AA or CRC diagnostic code, or if the subject is at risk of having pre-cancerous or cancerous colorectal cells, or if the treatment provided was sufficient to no longer receive a AA or CRC diagnostic code or to reduce the risk of having pre-cancerous or cancerous colorectal cells. As such, the methods described herein can be used after the subject has received at least one first therapy or surgery and before the subject received a further therapy or is submitted to a further surgery. The methods described herein can help the physician to determine if a more aggressive or a less aggressive therapeutic or surgical regimen is required.

[0139] The present disclosure also provides a kit for transforming stool sample data or a kit for performing the stratification methods. The kits comprise means (e.g., reagents) for determining the mRNA expression levels of the at least one, two, three, four, five or more genes present in the SSD or TEP. For example, the kits can comprise at least one, two, three, four, five or more pair of primers for amplifying the cDNA molecules corresponding to the mRNA molecules whose expression is being determined. The kits can include reagents for the detection of the mRNA expression level of the CEA adhesion molecule 5 (also referred to as CEACAM5, CD66e or CEA and having the Gene ID 1048). The kits can include reagents for the detection the mRNA expression level of the growth arrest and DNA damage inducible beta gene (also referred to as GADD45B, GADD45BETA or MYD118 and having the Gene ID: 4616). The kits can include reagents for the detection of the mRNA expression level of the integrin subunit alpha 2 gene (also referredto as ITGA2, BR, CD49B, GPla, HPA-5, VLA-2 or VLAA2 and having the Gene ID: 3673). The kits can include reagents for the detection of the mRNA expression level of the MET transcriptional regulator MAC Cl (also referred to as MACC1, 7A5 or SH3BP4L and having the Gene ID: 346389). The SSD or TEP can include the mRNA expression level of the MYB proto-oncogene like 2 gene (also referred to as MYBL2, B-MYB or BMYB and having the Gene ID: 4605). The kits can include reagents for the detection of the mRNA expression level of the MYC proto-oncogene, bHLH transcription factor (also referred to as MYC, MRTL, MYCC, bHLHe39 or c-Myc and having the Gene ID: 4609). The kits can include reagents for the detection of the mRNA expression level of the SI 00 calcium binding protein A4 (also referred to as S100A4, 18A2, 42 A, CAPL, FSP1, MTS1, P9KA or PEL98 and having Gene ID: 6275). In some further optional embodiments, the kits can include the reagents for the detection of the mRNA expression level of the beta- 2-microglobulin gene (also referred to as B2M or IMD43 and having the Gene ID 567). In some further optional embodiments, the kit can include reagents for the detection of the mRNA expression level of the integrin subunit alpha 1 gene (also referred to as ITGA1, CD49a or VLA1 and having the Gene ID: 3672). In a specific embodiment, the kits can include reagents for the detection of the mRNA expression profile of CEACAM5, ITGA6 and MACC 1. In yet another specific embodiment, the SSD or TEP can include the mRNA expression profile of CEACAM5, ITGA6, MACC1 and B2M. In still yet another embodiment, the kits can include reagents for the detection of the mRNA expression profile of PTGS2 and S100A4. In still yet another embodiment, the kits can include reagents for the detection of the mRNA expression profile of CEACAM5, ITGA6, MACC1, PTGS2 and S100A4, optionally in combination with the mRNA expression profile of B2M.

[0140] In some optional embodiments, the kits can include reagents for the detection of the mRNA expression level of the integrin subunit alpha 6 gene (also referred to as ITGA6, CD49f, ITGA6B or VLA-6 and having the Gene ID: 3655). In such embodiments, it is possible that the kits can include reagents for the detection of the expression of the mRNA expression level of the alpha- and / or beta-isoform of ITAGA6 gene transcript. In some optional embodiments, kit can include reagents for the detection of the mRNA expression level of the prostaglandin-endoperoxide synthase 2 gene (also referred to asPTGS2, COX- 2, COX2, GRIPGHS, PGG / HS, PGHS-2, PHS-2 or hCox-2 and having the Gene ID: 5743). In some specific embodiments, the kits can include at least one, two, three, four or five reagents for the detection of the level of expression the S100A4 gene, the GADD45Bgene, the ITGA6 gene, the MYBL2 gene, the MYC gene and / or the PTGS2 gene.

[0141] The kits can, in some embodiments, include a polymerase for performing a polymerase chain reaction. In some embodiments, the kits can also include primers for performing the reverse-transcription step. The kits can, in some additional embodiments, include a reverse transcriptase for performing the reverse transcription step. In some embodiments, the kits can include probes intended to be cleaved during the amplification step (e.g., Taqman® probes for example) to allow a quantitative PCR detection of the mRNA transcripts. The kits can also include a container for a stool sample from the subject and / or for storing the stool sample prior to the determination step. The kits also comprises instructions to use the means for determining the mRNA expression levels to obtain the test expression profile. The kits can also include instructions on how to stratify the subjects whose stool sample is being analyzed based on their risk of having an advanced adenoma or a colorectal cancer.

[0142] The present invention will be more readily understood by referring to the following examples which are given to illustrate the invention rather than to limit its scope.EXAMPLES

[0143] Patients and samples for initial experiments

[0144] Two sets of patient samples were used. The first set of samples was collected from patients and healthy controls from the Hamamatsu University School of Medicine with written informed consent. The study was approved by the institutional research ethics committee of the Hamamatsu University School of Medicine. Complete information about this set has been provided in previous studies (Herring et al., 2018; Beaulieu et al., 2016; Herring et al., 2017). Briefly, the study cohort used herein included 24 patients with AA defined as being 10 mm or larger at their greatest dimension and 78 patients with CRC (24 stage I, 32 stage II and 22 stage III) diagnosed by colonoscopy and histopathology as well as 32 healthy controls. For controls and AA, stool samples were collected before colonoscopy. The immunochemical fecal occult blood test (iFOBT) was performed on all patients and controls as described (Beaulieu et al., 2016).

[0145] The second set of samples was collected from three healthy controls and three patients diagnosed with CRC stage II or Ill by colonoscopy and histopathology from the Centre Hospitalier Universitaire de Sherbrooke (CHUS) with written informed consent.The study was approved by the institutional research ethics committee of the CHUS. This set of samples was used for mRNA target stability experiments. Each sample was split into 13 aliquots stored under various conditions for up to 5 days as follows: #1, 5 days at -80°C used as control; #2, 5 days at -20 °C; #3, 5 days at -20 °C with a thaw / freeze cycle; #4-8, 1-5 days at 4°C and #9-13, 1-5 days at 23°C.

[0146] Mainz Biomed conducted as a third clinical data set the Pilot Study for Colorectal Cancer and Advanced Adenoma Detection with the Mainz Biomed Colorectal Cancer Test (eAArly DETECT), IRB approved Pro00067506 by Advarra, 17. Nov. 2022. All subjects were enrolled between December 19, 2022, and August 31, 2023, across 21 US sites. The study was conducted for feasibility and test development and included two groups:

[0147] Group 1 : Subjects that are advised to have or are scheduled for a screening colonoscopy and are at average risk for colorectal cancer.

[0148] Group 2: Subjects that are suspected to have at least one advanced pre-cancerous lesion or colorectal cancer. Subjects in this group are those that have been pre-identified with imaging, a positive non -invasive screening test, and / or colonoscopy which requires additional intervention.

[0149] The study ended with 258 evaluable subjects. Subjects withdrew for the following reasons: subject withdrew consent, subject did not have a colonoscopy, subject did not follow study procedures (did not return sample, waited too long to return sample, or sample was collected after colonoscopy), or sample was not received at the lab in time.

[0150] The following results are based on data from 258 patients and demonstrate that the Mainz Biomed test achieved a CRC sensitivity of 97% (CI: 83.3-99.9), AA sensitivity of 88% (CI: 77.2-94.5), combined sensitivity for CRC, and AA of 91% (CI: 88.4-98.3) and a specificity of 93% (CI: 88.4-97.3).

[0151] Compared to Cologuard and ColoSense, the Mainz Biomed mt-sRNA test demonstrated both improved sensitivity and specificity for CRC and AA, with a substantially improved sensitivity for AA from 42.4-44.9% to 88% (See Table 4-1 for a comparison of test performance).

[0152] RNA isolation, reverse transcription, and PCR amplification.

[0153] RNA was isolated from fecal samples and reverse transcribed as described previously (Hamaya et al., 2010; Dydensborg etal., 2006). The quantitative analysis of the eight RNA markers, five (5) target genes (CEACAM5, ITGA6, MACC1, PTGS2, and S100A4) and three (3) housekeeping genes (RAB7A, CTTN, and GAPDH), from the stool sample was carried out using RT-PCR on the QuantstudioDx (Thermo Fisher Scientific™). Reverse transcription involves the use of the enzyme reverse transcriptase which reverse transcribes the RNA template into complementary DNA (cDNA). This cDNA can then be analyzed quantitatively by real-time PCR. The TaqMan assays contains two specific, unlabeled primers and a TaqMan probe designed by the manufacturer to span exons to prevent non-specific amplification of genomic DNA. The hydrolysis probe is labeled at the 5'-end with a fluorophore (reporter dye) and at the 3 '-end with a non-fluorescent quencher. The proximity of the reporter dye to the quencher inhibits the fluorescence of the reporter molecule. During amplification, the probe binds specifically to the cDNA fragment. The 5'-nuclease activity of the polymerase cleaves the hybridized probe, separating the reporter from the quencher and generating a fluorescence signal.

[0154] To analyze the expression level of the mRNA target genes, the housekeeping genes are used as references. For each sample, one or more of the housekeeping genes are analyzed in addition to the actual target. A housekeeping gene is a gene that is necessary for fundamental cellular processes, and therefore, each individual housekeeping gene is supposed to be expressed at a constant level in each cell. The levels of expression between different housekeeping genes are different. Consequently, the expression levels can be used as a reference to normalize the expression of a target gene on the input RNA. For normalization,the CP difference between the respective RNA marker and one of the housekeeping gene (ACP) is calculated.

[0155] Commercially available TaqMan primer and probe mixtures were used for the preamplification of the 27 preselected targets as described herein and detailed in Table 1. Quantitative polymerase chain reaction (qPCR) was performed using the TaqMan Gene Expression Assay with conditions described previously (Herring et al., 2018).

[0156] A list of specific targets tested is provided in Table 1. All primer and probe mixtures were first tested on a subset of stool samples including controls, AA and CRC to select those that were consistently detectable in the stools. Further analysis on the whole set of samples allowed the selection of those specifically enriched in CRC and AA or only CRC.

[0157] Table 1Table 2: Overview of the mt-sRNA Test Reagent Kit Components used for clinical data

[0158] Data presentation and statistical analysis for initial experiments.

[0159] Stool mRNA data were calculated as copy number per pl of reaction. For each transcript, a standard reference curve was generated using a serial fivefold dilution of a cDNA stock solution of the target sequence quantified on a NanoDrop 1000 Spectrophotometer (NanoDrop, Wilmington, DE, USA). Prism 8 was used for calculating statistics. Comparison mRNA expression (in copy number) in stool controls and patients with AA and CRC stage I-III lesions were expressed as median with interquartile range and analyzed by the Kruskal-Wallis test followed by Dunn's multiple comparison test. Area under the receiver operating characteristic (ROC) curves were calculated to establish sensitivities and specificities for each marker expressed in % with a 95% confidence interval. Scores were calculated for each marker on a scale of 0 to 3 on the basis of three cut-off values established from the ROC curve: (the lower cut-off corresponding to a sensitivity of 80%, medium cut-off corresponding to a specificity of 90% and higher cut-off corresponding to a specificity of 99%) as established previously. Statistical significance was defined as P < 0.05.

[0160] Twenty-seven (27) specific targets chosen on the basis of their reported over expression in colorectal cancerous lesions were screened. Preliminary evaluation of these using a subset of 30 samples (10 controls, 10 AA and 10 CRC) revealed that 14 were consistently detected in stools of patients bearing colorectal lesions (Table 1). Further testing with other primer and probe mixtures for poorly detected targets was tried but not further studied herein since 14 appeared to be enough to run the validation assay considering that for a clinical assay, the multiplex PCR capacity is limited to 4 to 5 targets depending on the manufacturer.

[0161] Further investigation of the 14 targets was performed on the set of 132 samples obtained from healthy controls (32) and patients bearing colorectal lesions (24 AA and 78 CRC). As shown in Table 1, a number of targets were found to be significantly over- represented in samples from patients with CRC while a few identified patients bearing AA or CRC. As illustrated with S100A4 (FIG. 1 A), the median copy number for the transcripts of the first group which also included GADD45B, ITGA2, MYBL2, MYC and PTGS2 (FIG. 4) were found to be significantly increased in the stools of patients with CRC ascompared with the controls while only three including CEACAM5 (FIG. 1 A), ITGA6 and MACC1 (FIG. 4) were found to be over-represented also in patients with AA.

[0162] Scores were then calculated for the two groups of markers. Because copy numbers varied considerably between the targets, from -200 for MYC to 40,000 for CEACAM5, individual scores were determined for all targets by attributing a value of Oto 3 for each patient sample based the on cut-off values of the targets, as described above. Then, an overall score for the each of the two groups of markers was determined for controls and patients with AA or CRC. As shown in FIG. IB, the overall score for the 6 markers of the first group significantly recognized the samples from CRC patients vs those of the controls while the overall scores of the three markers of the second group distinguished the samples from patients bearing CRC or AA from those of the controls. ROC curves for the two groups were determined (FIG. 1C). For the first group, the area under the curve (AUC) for CRC was 0.970 corresponding to a sensitivity of 89% for 95% specificity but AUC was only 0.825 for AA with a 58% sensitivity (for 95% specificity). In the second group, AUC was 0.914 for CRC and 0.917 for AA showing a sensitivity of 79% and 75%, respectively (for 95% specificity).

[0163] Considering that detecting 75% of the AA could be achieved using the three markers of the second group (i.e. CEACAM5, ITGA6 and MACC1), various combinations of markers belonging to the first group were included in order to improve CRC detection using a maximum of five targets (Table 2). Results showed that adding the two markers S100A4 and PTGS2 significantly improved the rate of CRC detection up to 89% (for 95% specificity) (FIG. 2A). Interestingly, considering the result of the FIT in combination with the multi-target score further increased CRC detection up to 95% (for a 97% specificity) but had no significant effect on AA detection (FIG. 2B).

[0164] A selection of target combinations is provided in Table 1. Sensitivities and Specificities were determined based on optimal cutoff values. AUC: Area under the curve, AA: Advanced adenoma, CRC: Colorectal cancers stage I, II and Ill.

[0165] Table 2

[0166] Considering that the ultimate goal would be to evaluate the feasibility of using the multi -target mRNA stool test in a clinical set-up, the stability of the mRNA targets in stool samples was evaluated subjected to various conditions of preservation that mimic the clinical reality. Stool samples were obtained from three controls and three patientsdiagnosed with CRC. Four of the identified targets in stools were selected for testing, including two for each group identified above: CEACAM5, ITGA6, ITGA2 and PTGS2. Conditions to be tested included conventional freezing at -20 °C with and without a thaw cycle, conservation at 4°C and conservation at room temperature (23°C), for a 5-day period. As shown for PTGS2 (FIG. 3A) as well as CEACAM5, ITGA6 and ITGA2 (FIG. 5), the mRNA targets were found to be very stable under all frozen and cooled conditions over the 5-day period while some variations were observed at room temperature for some markers such as PTGS2 (FIG. 3 A). Score compilation of the data confirmed the relative stability of the targets for all conditions including ambient temperature for at least 3 days (FIG. 3B).

[0167] In this example, a multitarget stool mRNA test was shown to represent a powerful assay for detecting patients with colorectal cancers and demonstrate its usefulness to also detect high risk adenomas. One interest of the procedure relies on its relative simplicity considering that high sensitivities and specificities can be obtained with a selection of only five targets, thus compatible with multiplex PCR in stool samples, an approach already in place in the clinic to investigate gastrointestinal infections.

[0168] One strength of the multitarget stool mRNA test presented herein is that transcripts are directly isolated from the stools by conventional extraction methods thus being compatible with automation rather than procedures that require enrichment protocols for exfoliated colorectal cells prior to RNA extraction and processing. Another strength is the relatively low number of targets required to optimize the assay. It is worth mentioning that an important part of this proof-of-concept study was finding specific targets to identify samples from patients with AA among others that appear to be overrepresented in CRC and then selecting the strongest combination to allow the detection of both AA and CRC.

[0169] It is interesting to contextualize the findings that this study, relying on the use of only five mRNA targets, allowed the detection of 75% of the samples obtained from patients with AA and 89% of the samples obtained from patients with CRC, using a specificity of 95%. It was chosen to express the data using this optimal specificity which generates less than 5% of false positives in order to allow a fair comparison to other tests. Incidentally, integration of the FIT component to the mRNA data increased CRC sensitivity up to 95%, consistent with the fact that the origins of exfoliated cells and blood in the stools are likely to be different. Overall, a multi-target stool mRNA-FIT test allowsthe detection of 75% of the AA and 95% of the CRC with less than 4% of false positives. These numbers compared advantageously to any other screening test for colorectal cancerous lesions. As shown with the inclusion of the FIT component, diversification of target types improves sensitivity.

[0170] Another finding is the possibility to include a factor for predicting AA vs CRC, which could provide pertinent information ahead of colonoscopy. Indeed, considered separately, the combination of the three targets CEACAM5, ITGA6 and MACC1 selected to predict AA provided 75% and 79% sensitivity (for 95% specificity) for AA and CRC respectively and the two targets S100A4 and PTGS2 selected to improve CRC detection provided 29% and 80% sensitivity (for 95% specificity) for AA and CRC prediction, suggesting that using distinct repertoires of targets for AA and CRC could be used to improve patient stratification for colonoscopy. Specific analysis of S100A4 and PTGS2 scores for patients identified as positive in the multi-target stool mRNA test could contribute to discriminating between patients carrying AA vs those with CRC considering that, for instance, a patient with a score of >4.5 for S100A4 and PTGS2 displays a 17% probability of having a AA vs 73% odds of having a CRC.

[0171] Finally, the assessment of target stability revealed that stool sample collection to perform the multitarget stool mRNA test does not require particular conditions, being relatively stable for at least 3 days, even at room temperature. Part of this relatively surprising observation may result from the possibility that mRNA degradation is prevented in exfoliated cells, which are the main source of host mRNA in the stools. Another part results from the procedure used for selecting the mRNA targets. Incidentally, it was not surprising that only half of the 27 selected targets were amplified in stool samples. The efficient amplification of these targets was also dependent on the use of the TaqMan Gene Expression Assay which was found to be more sensitive and specific than conventional qPCR for stool samples while requiring relatively short intact mRNA sequences.

[0172] The COLOFUTURE Study

[0173] COLOFUTURE is an independent, multi-center case cohort study. Data for the COLOFUTURE study were generated from stool samples collected at eight clinical sites in Germany and Norway. The cohort comprised 220 subjects, including 38 colorectal cancer (CRC), 55 advanced adenoma (AA), 10 with non-advanced adenomas (AD), and 117 normalcontrol subjects based on colonoscopy and pathology review. The cohort was comprised of 45% female and 55% male, with an average age of 62.3 years. The age distribution of the interim analysis and age-related stage distribution found in the cohorts is shown in Table 3.

[0174] Table 3

[0175] Results were compared to colonoscopy and pathology to determine the sensitivity and specificity for early detection of CRC and AA vs. control samples (non-advanced adenomas, AD and normal colons; FIG. 10A and FIG. 10B show pie charts illustrating the distribution of pathological results for CRC (FIG. 10A) and AA (FIG. 10B)). Nucleic acids were extracted from stabilized stool samples with a silica-bead-based extraction method. mRNA biomarkers were analyzed utilizing one-step qRT-PCR in combination with TaqMan probes; DNA markers were analyzed by qPCR; and human hemoglobin was quantified by FIT; and the Emerge Quantitative Evolution Al platform by Liquid Biosciences was used to develop classifiers capable of distinguishing CRC and AA from AD and normal samples.

[0176] The names and functions of the genetic biomarkers analyzed are shown in Table 4:

[0177] Table

[0178] Findings: The clinical performance of this approach in terms of sensitivity and specificity for CRC, AA, and a combined group of diseased subjects (with two-sided 95% Clopper-Pearson confidence intervals (CI)) is shown in Table 5:

[0179] Table 5

[0180] CRC sensitivity by stage was also assessed, with the results shown in Table 6.

[0181] Table 6

[0182] Where indicates the four CRC samples were not staged at the time of the analysis. Further, sensitivity for high grade dysplasia was 75% (9 / 12).

[0183] This innovative setup represents the inaugural instance of a multimodal analysis involving DNA, mRNA expression, FIT analysis, and AI / ML developed algorithm. Sensitivity for detection of CRC was found to be 94.4 %, while AAs were detected with 80 % sensitivity. Specificity was 97.5 % and 95.2 %, respectively (negative colonoscopies plus AD). Therefore, the results indicate an unexpectedly substantial and meaningful enhancement in the effectiveness of non-invasive CRC screening, particularly for detection of AA, where an increased sensitivity is urgently needed in order to reduce CRC incidence and mortality to a significant level.

[0184] Materials and methods

[0185] Subjects received a DNA / RNA Shield™ Fecal Collection tube (Zymo Research, Cat. No. R1101-E) to collect their stool sample. After bowel movement they transfer an amount of stool into the collection tube on the order of 500 pg to 1 g. DNA and RNA from the asubject' s stool were extracted by use of MagMAX Microbiome Ultra Nucleic Acid Isolation Kit (Applied Biosystems Cat. No. A42357). The DNA / RNA eluate was analyzed for the mRNA expression of the biomarkers B2M, CEACAM5, ITGA6, MACC1, PTGS2, S100A4 with absolute quantification and / or relative quantification by the use of one or more housekeeping gene(s).

[0186] For Real-Time One-Step RT-PCR (RT-qPCR) the TaqPath™ 1-Step Multiplex Master Mix (No ROX) from ThermoFisher Scientific (Order No. A28522) was used in combinationwith nuclease-free water.

[0187] The following TaqMan Gene Expression Assays from ThermoFisher were used for the detection of the mRNA markers: (1) CEACAM5 FAM-MGB, Assay-ID: Hs00944025_ml; (2) ITGA6 FAM-MGB, Assay-ID: Hs0104101 l_ml; (3) PTGS2 FAM-MGB, Assay-ID: Hs00153133_ml; (4) MACC1 FAM-MGB, Assay-ID: Hs00766186_ml; (5) S100A4 FAM- MGB, Assay-ID: Hs00243202_ml; and (6) B2M FAM-MGB, Assay-ID: Hs00984230_ml.

[0188] For relative quantification by use of housekeeping genes, the following TaqMan Gene Expression Assays from ThermoFisher were used: (1) CTTN VIC-MGB, Assay-ID: HsOl 124232 ml; (2) GAPDH VIC-MGB, Assay-ID: Hs99999905_ml; and (3) RAB7A VIC- MGB, Assay-ID: Hs01115139_ml.

[0189] For absolute quantification, as Positive Control plasmids from Eurofins Genomics containing the relevant sequences of the mRNA biomarkers and the binding sites for primers and probes of the mRNA-assays were used. For relative quantification, Universal Human Reference RNA from Applied Biosystems™ (Order No. QS0639) was used as Positive Control.

[0190] For each assay at least one no-template negative control (e.g., PCR-H2O instead of RNA, NTC) and a positive control (PC) was analyzed in parallel with test samples. For the evaluation of the amount of detected RNA (absolute quantification), a standard curve was prepared using the plasmid (Table 7).

[0191] Table 7

[0192] Where “ *” indicates the concentrations that are measured in the RT-qPCR as a standard curve. For all samples a master mix for each reaction was prepared according to Table 8, below.

[0193] Table 8

[0194] 15 pL of the master mix was added to each of a number of wells of a 96-well plate corresponding to the number of samples and controls needed for each experiment, followed by adding 5 pL of either a test-RNA or a control into the 15 pL of master mix of each well, followed by performing the following qRT-PCR protocol:

[0195] Table 9

[0196] The amplification curves (determination of the crossing points (Cp)) were evaluated with the “Absolute Quantification” method (Abs Quant / 2nd Derivative Max) on the LightCycler® 480II. The fluorescence signals for the different assays were detected on channels 465-510 nm (FAM) and 533-580 nm (VIC).

[0197] For relative quantification by use of housekeeping gene(s), a ACp value was calculated by subtraction of the Cp (i.e., Ct) values for one mRNA marker and one housekeeping gene (ACp = Cp mRNA marker - Cp HKG). The no-template-control (NTC) did not show an amplification curve, and the positive control demonstrated a similar Cp from run to run, both as expected.

[0198] Absolute quantification

[0199] Based on the calculated standard curve, the LightCycler® 480II “Absolute Quantification” method automatically calculated the copy number for each mRNA marker. FIG. 6A shows the amplification curves for a set of standards of known initial concentrations (i.e., standard curves) measured for mRNA marker CEACAM5, and FIG. 6B show amplification curves of experimental samples measured for CEACAM5 mRNA. The respective values are also summarized in Table 10.

[0200] Table 10

[0201] Relative quantification

[0202] For relative quantification analysis, mRNA expression levels of one or more housekeeping genes were analyzed in the same way as the mRNA markers, without the need for a standard curve. The relative expression level of each mRNA marker was calculated in relation to the housekeeping gene(s) by subtracting the measured Cp value of the housekeeping gene (HKG) from the Cp of each sample according to the formula: ACp = Cp mRNA marker - Cp HKG. FIG. 7A shows amplification curves of samples measured for mRNA markerCEACAM5, and FIG. 7B shows amplification curves of samples measured for housekeeping gene CTTN. The respective values for samples quantified for CEACAM5 using relative quantification are also summarized in Table 11.

[0203] Table 11

[0204] FIG. 8 illustrates the combined use of mRNA markers, housekeeping genes, absolute quantification and relative quantification as analytical steps in a method for identifying, diagnosing and / or stratifying the risk of a subject for having CRC or AA as described herein. FIG. 9A to FIG. 9C show comparisons of the quantification of MACC1, PTGS2, and S100A4, respectively, both by absolute quantification and relative quantification as compared to the housekeeping gene GAPDH.

[0205] Algorithm

[0206] FIG. 11 is an exemplary flow diagram of an artificial intelligence / machine learning (AI / ML) algorithm developed to classify a (stool) sample as a “predicted control” (i.e., the subject is preliminarily predicted to not have AA or CRC) or “predicted case” (i.e., the subject is preliminarily predicted to have AA or CRC), which is an optional step for the AI / ML described below.

[0207] FIG. 12 is an exemplary flow diagram of one output model from an AI / ML algorithmdeveloped herein to sensitively and specifically detect and quantify colorectal cancer and advanced adenoma risk. The described model is represented by boxes 1 through 5 (as noted in the lower right corner of each box). For Step 0 (box 1), if the CEACAM5 expression level is a Cp >5 and MACC1 expression level is a concentration of >5 Copies / pL, the yes condition obtains, and the result of Step 0 is assigned as the ACp for PTGS2 and housekeeper gene GAPDH divided by 2 (i.e., (Cp PTGS2 - Cp GAPDH) / 2). If measured CEACAM5 and MACC1 values do not meet these conditions, a no condition obtains, and the result of Step 0 is a “7”.

[0208] For Step 3 (box 2), if the MACC1 expression level in a sample is a concentration of >6 copies / pL and the MACC1 expression level in a sample is a concentration of >6 copies / pL, a yes condition obtains, and the result of Step 3 is a “15”. If the conditions are not met, the result of Step 3 is a “2”.

[0209] For Step 6 (box 3), if the MACC1 expression level is greater than the result of Step 3, a yes condition obtains, and the result of Step 6 is the square of the result of Step 0 (i.e., 72or 49 for a Step 0 “no” result and (Cp PTGS2 - Cp GAPDH) / 2)2for a Step 0 “yes” result). If the MACC1 expression level is not greater than the result value for Step 3, a no condition obtains, and the result of Step 6 is a “3”.

[0210] For Step 9 (box 4), if the FIT concentration (in this case, IDK Hemoglobin ELISA) does not equal the result of Step 6 and the S100A4 expression level concentration in copies / pL does not equal the result of Step 3, then a yes condition obtains, and the result of Step 9 is the result of Step 6 subtracted from the IDK FIT concentration. Otherwise, the result is a “0”. For the Classification step (box 5), if the result of Step 9 is less than 0, a yes condition obtains, and the result of the Classification step is a “0” (i.e., a sample that is classified as a negative for CRC). If the result of Step 9 is not less than 0, the result of the Classification step is a “1”, or a sample that is classified as positive for CRC.

[0211] The result of the measurement and quantitation of mRNA expression levels for the described markers and assessment in the algorithm as described resulted in heretofore unattainable levels of sensitivity and specificity for CRC and AA.

[0212] FIG. 13 shows a flowchart of a computer-implemented method 200 for transforming stool sample data according to an example implementation. Method 200 may include one or more operations, functions, or actions as illustrated by one or more of blocks 202-212.Although the blocks are illustrated in a sequential order, these blocks may also be performed in parallel, and / or in a different order than those described herein. Also, the various blocks may be combined into fewer blocks, divided into additional blocks, and / or removed based upon the desired implementation.

[0213] It should be understood that for this and other processes and methods disclosed herein, flowcharts show functionality and operation of one possible implementation of present examples. In this regard, each block or portions of each block may represent a module, a segment, or a portion of program code, which includes one or more instructions executable by a processor for implementing specific logical functions or steps in the process. The program code may be stored on any type of computer readable medium or data storage, for example, such as a storage device including a disk or hard drive. Further, the program code can be encoded on a computer-readable storage media in a machine-readable format, or on other non- transitory media or articles of manufacture. The computer readable medium may include non- transitory computer readable medium or memory, for example, such as computer-readable media that stores data for short periods of time like register memory, processor cache and Random Access Memory (RAM). The computer readable medium may also include non- transitory media, such as secondary or persistent long term storage, like read only memory (ROM), optical or magnetic disks, compact-disc read only memory (CD-ROM), for example. The computer readable media may also be any other volatile or non-volatile storage systems. The computer readable medium may be considered a tangible computer readable storage medium, for example.

[0214] In addition, each block or portions of each block in FIG. 13 , and within other processes and methods disclosed herein, may represent circuitry that is wired to perform the specific logical functions in the process. Alternative implementations are included within the scope of the examples of the present disclosure in which functions may be executed out of order from that shown or discussed, including substantially concurrent or in reverse order, depending on the functionality involved, as would be understood by those reasonably skilled in the art.

[0215] At block 202, the method 200 includes providing a stool sample from a subject, in which the stool sample includes a plurality of mRNA transcripts.

[0216] At block 204, the method 200 includes determining, as the stool sample data, expression levels of at least two distinct genes from a plurality of mRNA transcripts in thestool sample.

[0217] At block 206, the method 200 includes providing a representation of a graphical user interface for display on a computing system.

[0218] At block 208, the method 200 includes receiving an input from the graphical user interface, the input comprising the stool sample data. In one example, the stool sample data associated with the subject further comprises subject age, gender, or age and gender.

[0219] At block 210, the method 200 includes determining, by the computing system executing a machine-learning model and based on the stool sample data, a diagnostic code, wherein the machine-learning model was trained using stool diagnostic training data including a plurality of stool sample data labeled with at least one of two or more corresponding diagnostic codes.

[0220] At block 212, the method 200 includes providing, for display on the graphical user interface, the diagnostic code.

[0221] In conclusion, this example demonstrates the usefulness of host mRNAs as biomarkers to identify patients carrying curable colorectal cancers as well as precancerous lesions.

[0222] While the invention has been described in connection with specific embodiments thereof, it will be understood that the scope of the claims should not be limited by the preferred embodiments set forth in the examples, but should be given the broadest interpretation consistent with the description as a whole.

Claims

WHAT IS CLAIMED IS:

1. A computer-implemented method for transforming stool sample data, the method comprising: a) providing a stool sample from a subject; b) determining, as the stool sample data, expression levels of at least two distinct genes from a plurality of mRNA segments in the stool sample; and c) determining, by a computing system executing a machine-learning model and based on the stool sample data, a diagnostic code, wherein the machine-learning model was trained using stool diagnostic training data including a plurality of stool sample data labeled with at least one of two or more corresponding diagnostic codes.

2. A computer-implemented method for transforming stool sample data, the method comprising: a) providing a stool sample from a subject; b) determining, as the stool sample data, expression levels of at least two distinct genes from a plurality of mRNA transcripts in the stool sample; c) providing a representation of a graphical user interface for display on a computing system; d) receiving an input from the graphical user interface, the input comprising the stool sample data; e) determining, by the computing system executing a machine-learning model and based on the stool sample data, a diagnostic code, wherein the machine-learning model was trained using stool diagnostic training data including a plurality of stool sample data labeled with at least one of two or more corresponding diagnostic codes; f) providing, for display on the graphical user interface, the diagnostic code.

3. The method of claim 1 or claim 2, wherein the stool sample data and stool diagnostic training data further comprises subject age and gender.

4. The method of claim 1 or claim 2, wherein the stool sample data and stool diagnostic training data further comprises subject age or gender.

5. The method of claim 1 or claim 2, wherein the stool sample data and stool diagnostic training data further comprises the presence or abundance of hemoglobin (occult blood) in the stool sample.

6. The method of claim 5 comprising using a fecal immunochemical test (FIT) to determine the presence of hemoglobin in the stool sample7. The method of claim 1 or claim 2, wherein the stool sample comprises at least one colorectal epithelial cell.

8. The method of claim 7, wherein the at least one colorectal epithelial cell comprises the plurality of mRNA transcripts.

9. The method of any one of claims 1 to 8, wherein the diagnostic code comprises normal, non-advanced adenomas (NA), advanced adenoma (AA), colorectal cancer (CRC), or AA / CRC.

10. The method of any one of claims 1 to 8, wherein the diagnostic code is selected from the group consisting of: a) normal, NA, AA, CRC, and AA / CRC; or b) normal, NA, AA, and CRC; or c) normal, NA, and AA / CRC; or d) normal, AA, CRC, and AA / CRC; or e) normal, A A, and CRC; or f) normal and AA / CRC.

11. The method of any one of claims 1 to 10, further comprising a step of prescribing a colonoscopy for the subject if the diagnostic code determined for the subject’s stool sample data is AA, CRC, or AA / CRC.

12. The method of any one of claims 1 to 11, further comprising submitting the subject to a chemotherapy, a radiotherapy and / or a surgery if the diagnostic code determined for the subject’s stool sample data is AA, CRC, or AA / CRC.

13. The method of any one of claims 1 to 12, wherein the mRNA expression level of at least two distinct genes comprises at least two of the CEACAM5 gene, the MACC1 gene, the PTGS2 gene, the S100A4 gene, the ITGA6 gene, the GADD45B gene, the ITGA2 gene, the MYBL2 gene, and the MYC gene.

14. The method of any one of claims 1 to 13, wherein the mRNA expression level of at least two distinct genes comprises at least two of the CEACAM5 gene, the MACC1 gene, the PTGS2 gene, the S100A4 gene, and the ITGA6 gene.

15. The method of any one of claims 1 to 14, wherein the mRNA expression level of at least two distinct genes comprises the CEACAM5 gene, the ITGA6 gene and / or the MACC1 gene.

16. The method of any one of claims 1 to 15, wherein the mRNA expression level of at least two distinct genes comprises the CEACAM5 gene, the ITGA6 gene, the MACC1 gene, the PTGS2 gene, and the S100A4 gene.

17. The method of any one of claims 1 to 12, wherein the mRNA expression level of at least two distinct genes comprises the S100A4 gene, the GADD45B gene, the ITGA2 gene, the MYBL2 gene, the MYC gene and / or the PTGS2 gene.

18. The method of any one of claims 1 to 17, wherein step b) comprises using a reverse-transcriptase polymerase chain reaction (RT-PCR) to obtain the mRNA expression level.

19. The method of any one of claims 1 to 18, wherein step b) comprises using a quantitative polymerase chain reaction (qPCR) to obtain the mRNA expression level.

20. The method of any one of claim 1 to 19, wherein the mRNA expression level for each gene is expressed as an absolute amount or a normalized amount.

21. The method of any one of claim 1 to 20, wherein the mRNA expression level for each gene is expressed as one or more of a Cp value, a ACp value, or an absolute amount of mRNA or cDNA made from a mRNA.

22. The method of any one of claims 1 to 21, further comprising prior to step b), storing the stool sample.

23. The method of any one of claims 1 to 22, further comprising determining the mRNA expression level of one or more housekeeping genes.

24. the method of claim 23, wherein the housekeeping genes comprises one or more of GAPDH, CTTN or RAB7A.

25. The method of any one of claims 1 to 24, further comprising determining the presence of a DNA mutation and / or an aberrant DNA methylation pattern associated with a predisposition to a colorectal cancer in the colorectal epithelial cell of the subject.

26. The method of claim 25, wherein the at least one DNA mutation is located in the KRAS, the BRAF gene, or both KRAS and BRAF genes.

27. The method of claim 25, wherein the aberrant DNA methylation pattern is located in the NDRG4 gene and / or the BMP3 gene.

28. The method of any one of claims 23 to 27 comprising using the Cologuard™ or ColoAlert™ assay to determine the presence of hemoglobin in the stool sample, the presence of DNA mutation and / or the presence of the abnormal DNA methylation pattern.

29. The method of any one of claims 1 to 28, wherein the colorectal cancer is a colon cancer.

30. The method of any one of claims 1 to 29, wherein the colorectal cancer is a rectal cancer.

31. A kit for transforming stool sample data, wherein the kit comprises at least two reagents for determining a mRNA expression level of at least two distinct genes from the plurality of mRNA transcripts to obtain a diagnostic code in a stool sample from the subject.

32. The kit of claim 31, further comprising a container for storing a stool sample.

33. The kit of claim 32, wherein the stool sample comprises about 1 grams, or about 750 mg, or about 500 mg, or about 250 mg, or about 100 mg, or about 75 mg or about 50 mg, or about 25 mg, or about 10 mg, or about 7.5 mg, or about 5 mg, or about 2.5 mg, or about 1 mg, or about 750 pg, or about 500 pg, or about 250 pg, or about 100 pg of stool.

34. The kit of any one of claims 31 to 33, wherein the at least two reagents are for determining the mRNA expression level of the S100A4 gene, the GADD45B gene, the ITGA2 gene, the MYBL2 gene, the MYC gene, the PTGS2 gene, the ITGA6 gene, the CEACAM5 gene, the B2M gene, and / or the MACC1 gene.

35. The kit of any one of claims 31 to 33, wherein the at least two reagents are for determining the mRNA expression level of the S100A4 gene, the PTGS2 gene, the ITGA6 gene, the CEACAM5 gene, and / or the MACC1 gene.

36. The kit of any one of claims 31 to 35, wherein the at least two reagents are for determining the mRNA expression level of the CEACAM5 gene, the ITGA6 gene and / or the MAC Cl gene.

37. The kit of any one of claims 31 to 36, wherein the at least two reagents are for determiningthe mRNA expression level of the CEACAM5 gene, the ITGA6 gene, the MACC1 gene, the PTGS2 gene, and the S100A4 gene.

38. The kit of any one of claims 31 to 36, wherein the at least two reagents are for determining the mRNA expression level of the S100A4 gene, the GADD45B gene, the ITGA2 gene, the MYBL2 gene, the MYC gene and / or the PTGS2 gene.

39. The kit of any one of claims 31 to 38, further comprising reagents for determining an mRNA expression level of one or more housekeeping genes.

40. The kit of claim 39, wherein the housekeeping genes comprises one or more of GAPDH, CTTN or RAB7A.

41. The kit of any one of claims 31 to 40, further comprising a reverse-transcriptase.

42. The kit of any one of claims 31 to 41, further means for determining the presence of hemoglobin in the stool sample.

43. The kit of claim 42, further comprising using a fecal immunochemical test (FIT) to determine the presence of hemoglobin in the stool sample.

44. The kit of any one of claims 31 to 43, further comprising reagents for determining the presence of a DNA mutation and / or an aberrant DNA methylation pattern associated with a predisposition to a colorectal cancer in the colorectal epithelial cell of the subject.

45. The kit of claim 44, wherein the at least one DNA mutation is located in the KRAS gene, the BRAF gene, or both the KRAS and BRAF genes.

46. The kit of claim 44 or 45, wherein the aberrant DNA methylation pattern is located in the NDRG4 gene and / or the BMP3 gene.

47. The kit of any one of claims 44 to 46, further comprising a Cologuard™ and / or a ColoAlert™ assay to determine the presence of hemoglobin in the stool sample, the presence of DNA mutation and / or the presence of the abnormal DNA methylation pattern.

48. A method of stratifying the risk of a subject of having an advanced adenoma or a colorectal cancer in a subject, the method comprising: providing a stool sample from the subject, wherein the stool sample comprises a plurality of mRNA transcripts from the subject;determining the mRNA expression level of at least two distinct genes from the plurality of mRNA transcripts to obtain a test expression profile; and comparing the test expression profile with a control expression profile, wherein the control expression profile comprises the mRNA expression level of the at least two genes and is derived from a plurality of control mRNA transcripts from a control subject known to lack the advanced adenoma or the colorectal cancer; wherein if it is determined that the test expression profile of the subject comprises at least two genes whose expression are increased with the respect to the control expression profile, the subject is stratified as having an increased risk of having the advanced adenoma or the colorectal cancer, when compared to the control subject.

49. The method of claim 48, wherein the stool sample comprises at least one colorectal epithelial cell.

50. The method of claim 49, wherein the at least one colorectal epithelial cell comprises the plurality of mRNA transcripts.

51. The method of any one of claims 48 to 50, wherein the test expression profile and the control expression profile comprise the mRNA expression level of at least two of the CEACAM5 gene, the MACC1 gene, the PTGS2 gene, the S100A4 gene, the ITGA6 gene, the GADD45B gene, the ITGA2 gene, the MYBL2 gene, and the MYC gene.

52. The method of any one of claims 48 to 51, wherein the test expression profile and the control expression profile comprise the mRNA expression level of at least two of the CEACAM5 gene, the MACC1 gene, the PTGS2 gene, the S100A4 gene, and the ITGA6 gene.

53. The method of any one of claims 48 to 52, wherein the test expression profile and the control expression profile comprise the mRNA expression level of the CEACAM5 gene, the ITGA6 gene and / or the MACC1 gene.

54. The method of any one of claims 48 to 53, wherein the test expression profile and the control expression profile comprise the mRNA expression level of the CEACAM5 gene, the ITGA6 gene, the MACC1 gene, the PTGS2 gene, and the S100A4 gene.

55. The method of any one of claims 48 to 51, wherein the test expression profile and the control expression profile comprise the mRNA expression level of the S100A4 gene, theGADD45B gene, the ITGA2 gene, the MYBL2 gene, the MYC gene and / or the PTGS2 gene.

56. The method of any one of claims 48 to 55, wherein step b) comprises using a reversetranscriptase polymerase chain reaction (RT-PCR) to obtain the mRNA expression level of the at least two genes of the test expression profile and / or the control expression profile.

57. The method of any one of claims 48 to 56, wherein step b) comprises using a quantitative polymerase chain reaction (qPCR) to obtain the mRNA expression level of the at least two genes of the test expression profile and / or the control expression profile.

58. The method of any one of claim 48 to 57, wherein the mRNA expression level for each gene is expressed as an absolute amount or a normalized amount.

59. The method of any one of claim 48 to 57, wherein the mRNA expression level for each gene is expressed as one or more of a Cp value, a ACp value, or an absolute amount of mRNA or cDNA made from a mRNA.

60. The method of any one of claims 48 to 59, further comprising, prior to step b), storing the stool sample.

61. The method of any one of claims 48 to 60, further comprising determining the presence of hemoglobin in the stool sample.

62. The method of claim 61 comprising using a fecal immunochemical test (FIT) to determine the presence of hemoglobin in the stool sample.

63. The method of any one of claims 48 to 62, further comprising determining the mRNA expression level of one or more housekeeping genes.

64. the method of claim 63, wherein the housekeeping genes comprises one or more of GAPDH, CTTN or RAB7A.

65. The method of any one of claims 48 to 64, wherein the test expression profile and control expression profile comprise the CEACAM5 gene, the ITGA6 gene, the MACC1 gene, the PTGS2 gene, and the S100A4 gene; the mRNA expression level for the CEACAM5 gene is calculated as a Cp value;the mRNA expression level for the MACC1 gene is calculated as a concentration value; the mRNA expression level for the PTGS2 gene is calculated as a ACp value; the mRNA expression level for the S100A4 gene is calculated as a concentration value; and wherein the presence of hemoglobin in the stool sample is determined using a fecal immunochemical test resulting in a FIT concentration value; and wherein if the CEACAM5 expression level is a Cp >5 and MACC1 expression level is a concentration of >5 copies / pL then a first value is obtained, the first value being defined as ACp for PTGS2 and housekeeper gene GAPDH divided by 2; otherwise, the first value is a 7; if the MACC1 expression level in a sample is a concentration of >6 copies / pL, a second value is obtained, the second value being a 15; otherwise, the second value is a 2; if the MACC1 expression level is between 2 and 15, then a third value is obtained, the third value being a “3”, otherwise the third value is the square of the first value; if the FIT concentration value does not equal the third value and the S100A4 expression level concentration does not equal the second value, then a fourth value is obtained, the fourth value being the result of subtracting the third value from the FIT concentration value; otherwise, the fourth value is a “0”; and if the fourth value is less than 0, a fifth value is obtained, the fifth value being a “0”; otherwise the fifth value is a “1”, wherein a “0” indicates a sample is negative for CRC, and a “1” indicates a sample is positive for CRC.

66. The method of any one of claims 48 to 65, further comprising determining the presence of a DNA mutation and / or an aberrant DNA methylation pattern associated with a predisposition to a colorectal cancer in the colorectal epithelial cell of the subject.

67. The method of claim 66, wherein the at least one DNA mutation is located in the KRAS, the BRAF gene, or both KRAS and BRAF genes.

68. The method of claim 66, wherein the aberrant DNA methylation pattern is located in the NDRG4 gene and / or the BMP3 gene.

69. The method of any one of claims 61 to 68, comprising using the Cologuard™ or ColoAlert™ assay to determine the presence of hemoglobin in the stool sample, the presence of DNA mutation and / or the presence of the abnormal DNA methylation pattern.

70. The method of any one of claims 48 to 69 for screening for subjects suitable for colonoscopy.

71. The method of any one of claims 48 to 70, further comprising submitting the subject having been stratified as being at increased risk of developing the colorectal cancer to a chemotherapy, a radiotherapy and / or a surgery.

72. The method of any one of claims 48 to 71, wherein the colorectal cancer is a colon cancer.

73. The method of any one of claims 48 to 72, wherein the colorectal cancer is a rectal cancer.

74. A kit for stratifying the risk of a subject of having an advanced adenoma or a colorectal cancer in a subject, wherein the kit comprises at least two reagents for determining a mRNA expression level of at least two distinct genes from the plurality of mRNA transcripts to obtain a test expression profile in a stool sample from the subject.

75. The kit of claim 74, further comprising a container for storing a stool sample.

76. The kit of claim 75, wherein the stool sample comprises about 1 grams, or about 750 mg, or about 500 mg, or about 250 mg, or about 100 mg, or about 75 mg or about 50 mg, or about 25 mg, or about 10 mg, or about 7.5 mg, or about 5 mg, or about 2.5 mg, or about 1 mg, or about 750 pg, or about 500 pg, or about 250 pg, or about 100 pg of stool.

77. The kit of any one of claims 74 to 76, wherein the at least two reagents are for determining the mRNA expression level of the S100A4 gene, the GADD45B gene, the ITGA2 gene, the MYBL2 gene, the MYC gene, the PTGS2 gene, the ITGA6 gene, the CEACAM5 gene, the B2M gene, and / or the MACC1 gene.

78. The kit of any one of claims 74 to 76, wherein the at least two reagents are for determining the mRNA expression level of the S100A4 gene, the PTGS2 gene, the ITGA6 gene, the CEACAM5 gene, and / or the MACC1 gene.

79. The kit of any one of claims 74 to 78, wherein the at least two reagents are for determining the mRNA expression level of the CEACAM5 gene, the ITGA6 gene and / or the MAC Cl gene.

80. The kit of any one of claims 74 to 79, wherein the atleasttwo reagents are for determining the mRNA expression level of the CEACAM5 gene, the ITGA6 gene, the MACC1 gene, the PTGS2 gene, and the S100A4 gene.

81. The kit of any one of claims 74 to 80, wherein the at least two reagents are for determining the mRNA expression level of the S100A4 gene, the GADD45B gene, the ITGA2 gene, the MYBL2 gene, the MYC gene and / or the PTGS2 gene.

82. The kit of any one of claims 74 to 81, further comprising reagents for determining an mRNA expression level of one or more housekeeping genes.

83. The kit of claim 82, wherein the housekeeping genes comprises one or more of GAPDH, CTTN or RAB7A.

84. The kit of any one of claims 74 to 83, further comprising a reverse-transcriptase.

85. The kit of any one of claims 74 to 84, further means for determining the presence of hemoglobin in the stool sample.

86. The kit of claim 85, further comprising using a fecal immunochemical test (FIT) to determine the presence of hemoglobin in the stool sample.

87. The kit of any one of claims 74 to 86, further comprising reagents for determining the presence of a DNA mutation and / or an aberrant DNA methylation pattern associated with a predisposition to a colorectal cancer in the colorectal epithelial cell of the subject.

88. The kit of claim 87, wherein the at least one DNA mutation is located in the KRAS gene, the BRAF gene, or both the KRAS and BRAF genes.

89. The kit of claim 87 or claim 88, wherein the aberrant DNA methylation pattern is located in the NDRG4 gene and / or the BMP3 gene.

90. The kit of any one of claims 87 to 89, further comprising a Cologuard™ and / or a ColoAlert™ assay to determine the presence of hemoglobin in the stool sample, the presence of DNA mutation and / or the presence of the abnormal DNA methylation pattern.