Compositions and methods for diagnosing early-stage precancerous colorectal advanced adenomas or cancers

A biomarker panel measured in biological samples using advanced PCR techniques and machine learning analysis provides a non-invasive, cost-effective method for diagnosing colorectal cancer and precancerous advanced adenomas, addressing the limitations of current screening methods.

WO2025117915A1PCT designated stage expired Publication Date: 2025-06-05EL CAPITAN BIOSCIENCES INC

Patent Information

Application Number
PCT/US2024/057989
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-29
Filing Date
2024-11-29
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Current methods for screening colorectal cancer (CRC) are invasive, expensive, or have low sensitivity for early-stage and precancerous lesions, highlighting the need for a non-invasive, cost-effective, and high-performing screening method.

Method used

A method involving the measurement of a panel of biomarkers selected from genes such as PPBP, OLR1, CXCL8, MMP3, HBA, TIMP1, MMP7, TGFBI, MYC, SPP1, NKD2, IL11, ETV4, CLEC5A, S100A9, TCN1, MMP1, IGF2, GDF15, TUBB3, RPP40, EPHX4, CBX2, ASCL2, EVA1A, PRSS33, RAET1L, and ULBPl in biological samples like stool or blood to diagnose CRC or precancerous colorectal advanced adenomas using quantitative RT-PCR, Droplet Digital PCR, or nucleic acid sequencing, with machine learning classifiers for analysis.

Benefits of technology

This approach enables accurate diagnosis of CRC and precancerous advanced adenomas with high sensitivity and specificity, potentially reducing the need for invasive procedures and lowering screening costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000039_0001
    Figure IMGF000039_0001
  • Figure IMGF000040_0001
    Figure IMGF000040_0001
  • Figure IMGF000041_0001
    Figure IMGF000041_0001
Patent Text Reader

Abstract

The present disclosure provides methods and compositions, e.g., kits, for diagnosing colorectal cancer or precancerous colorectal advanced adenomas based on a subject's biomarker panels in a biological sample (e.g., a feces sample or a blood sample).
Need to check novelty before this filing date? Find Prior Art

Description

COMPOSITIONS AND METHODS FOR DIAGNOSING EARLY-STAGE PRECANCEROUS COLORECTAL ADVANCED ADENOMAS OR CANCERSSEQUENCE LISTING

[0001] The sequence listing that is contained in the file named “081996-8005W001_Seq”, which is 168,540 bytes and was created on November 28, 2024, is filed herewith by electronic submission and is incorporated by reference herein.CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to US provisional application 63 / 604,179, filed November 29, 2023, the disclosure of which is incorporated herein by reference.FIELD OF THE INVENTION

[0003] The present disclosure generally relates to diagnosis and treatment of cancers or precancerous polyps. In particular, the present disclosure relates to biomarker panels in a biological sample (e.g., a feces sample or a blood sample) for diagnosing and treating colorectal cancer or precancerous colorectal advanced adenomas.BACKGROUND

[0004] Colorectal cancer (CRC) is the third most commonly diagnosed cancer and the third most common cause of cancer-related death in both men and women (Bray et al. 2024). The adenoma-carcinoma sequence is a fundamental concept in colorectal cancer, which refers to the stepwise progression from normal epithelium to dysplastic epithelium to carcinoma, due to accumulation of mutations over a number of years leading to activation of oncogenes and loss of tumor suppressor genes (Fearon and Vogelstein 1990).

[0005] Any approved screening method, if used as recommended, significantly reduces CRC-specific mortality. In general, two types of screening are recommended for colorectal cancer: Visual screening procedures such as colonoscopy and stool-based tests. While decennial colonoscopy is likely to be the most effective screening method with a 73% reduction in mortality, the adherence to this method is an important factor in CRC-specific mortality (Zheng et al. 2023). Given the screening modality and the possibility of choosing, less invasive tests have a higher uptake rate (Heidenreich et al. 2022).

[0006] To prevent CRC, stool -based screening with a fecal immunochemical test (FIT) is most commonly used worldwide. FIT tests quantify the level of occult blood in stool. Patients with an abnormally high level of occult blood in stool are referred for a colonoscopy. The evaluation of CRC screening programs in Europe revealed that for a positive FIT, the reportedpositive predictive value for advanced adenomas ranged from 5% in Ireland to 30% in Italy (Navarro et al. 2017). More complex non-invasive options like multi-target stool DNA (MT- sDNA) test detect curable-stage CRC with high sensitivity 93% (95% CI, 84-98%) to 100% (69-100%) and outperforms FIT in the detection of both advanced adenomas and serrated precursors with sensitivity increasing in association with risk of progression to cancer (“Multitarget Stool DNA Testing for Colorectal -Cancer Screening | NEJM,” n.d.; Bosch et al. 2019).

[0007] Multitarget stool mRNA-FIT is an emerging platform for CRC screening with promising performance: the detection of 46% of the AA and 94% of the CRC. Test performance including specificity is highly correlated with the genes selected and their combination (Barnell et al. 2019; Herring et al. 2021).

[0008] While considerable progress has been made in developing CRC screening tests, all existing approaches have significant flaws. Colonoscopies are invasive and expensive, FIT tests have low sensitivity for early stage and precancerous lesions while MT-sDNA methods are expensive with mediocre specificity. Hence, there is a need to develop a high performing, non-invasive and low-cost method of CRC screening.SUMMARY OF INVENTION

[0009] The present disclosure in one aspect provides a method for diagnosing colorectal cancer or precancerous colorectal advanced adenomas in a subject. In one embodiment, said method comprises: measuring in a biological sample obtained from the subject levels of a panel of biomarkers comprising one or more genes (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10 or more genes) selected from Table 1, evaluating the measured levels of the panel of biomarkers, and determining that the subject is healthy or has precancerous colorectal advanced adenomas or colorectal cancer.

[0010] In some embodiments, the panel of biomarkers comprises one or more genes (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10 or more genes) selected from the group consisting of: PPBP, OLR1, CXCL8, MMP3, HBA, TIMP1, MMP7, TGFBI, MYC, SPP1, NKD2, IL11, ETV4, CLEC5A, S100A9, TCN1, MMP1, IGF2, GDF15, TUBB3, RPP40, EPHX4, CBX2, ASCL2, EVA1A, PRSS33, RAET1L, and ULBPl.

[0011] In some embodiments, the panel of biomarkers comprises a gene pair selected from the Table 2. In some embodiments, the panel of biomarkers comprises a gene pair selected from the Table 4.

[0012] In some embodiments, the method is for diagnosing colorectal cancer, and the panel of biomarkers comprises a gene pair selected from the group consisting of: (a) PPBP and TIMP1; (b) CXCL8 and PPBP; (c) PPBP and CBX2; (d) MMP7 and PPBP; (e) PPBP and MMP1; (f) HBA and MYC; (g) MYC and PPBP; (h) OLR1 and PPBP; (i) MMP3 and SPP1; (j) PPBP and ETV4; (k) OLR1 and TUBB3; (1) PPBP and SPP1; (m) CXCL8 and MMP3; (n) HBA and MMP3; (o) OLR1 and MMP3; (p) PPBP and CEACAM6; (q) PPBP and S100A9; (r) OLR1 and IFI6; (s) CXCL8 and TUBB3; and (t) PPBP and CLEC5A.

[0013] In some embodiments, the method is for diagnosing precancerous colorectal advanced adenomas, and the panel of biomarkers comprises a gene pair selected from the group consisting of: (a) MYC and CGREF1; (b) MYC and TCN1; (c) MYC and CEL; (d) HBA and MYC; (e) MYC and CPNE1; (f) MYC and TRIM29; (g) MYC and S100A9; (h) MYC and SPTBN2; (i) MYC and RHPN1; (j) HBA and CGREF1; (k) MYC and CEMIP; (1) HBA and TCN1; (m) MMP7 and CGREF1; (n) MYC and UBE2C; (o) EPHX4 and MYC; (p) MYC and CLEC5A; (q) HBA and PPBP; (r) MYC and PPBP; (s) MYC and CXCL3; and (t) TIMP1 and TCN1.

[0014] In certain embodiment, the subject is a human or non-human.

[0015] In certain embodiment, the biological sample is a stool sample, a blood sample, a urine sample, a lymph sample, a saliva sample or a tissue sample.

[0016] In certain embodiment, the biological sample is a stool sample.

[0017] In certain embodiment, the levels of a panel of biomarkers are measured by detecting mRNA expression levels or expressed protein levels.

[0018] In certain embodiment, the mRNA expression levels are detected using quantitative RT-PCR or Droplet Digital PCR.

[0019] In certain embodiment, the levels of the panel of biomarkers are determined by nucleic acid sequencing.

[0020] In certain embodiment, the evaluating step and / or the determining step comprises analyzing the levels of a panel of biomarkers by a machine learning classifier.

[0021] In certain embodiment, the machine learning classifier is random forests or generalized linear model.

[0022] In another aspect, the present disclosure provides a panel of biomarkers for use in diagnosing colorectal cancer or precancerous colorectal advanced adenomas in a subject comprising one or more genes (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10 or more genes) selected from Table 1, Table 2 or Table 4.

[0023] In some embodiments, the present disclosure provides a panel of biomarkers for use in diagnosing colorectal cancer in a subject comprising a gene pair selected from the group consisting of: (a) PPBP and TIMP1; (b) CXCL8 and PPBP; (c) PPBP and CBX2; (d) MMP7 and PPBP; (e) PPBP and MMP1 ; (f) HB A and MYC; (g) MYC and PPBP; (h) OLR1 and PPBP; (i) MMP3 and SPP1; (j) PPBP and ETV4; (k) OLR1 and TUBB3; (1) PPBP and SPP1; (m) CXCL8 and MMP3; (n) HBA and MMP3; (o) OLR1 and MMP3; (p) PPBP and CEACAM6; (q) PPBP and S100A9; (r) OLR1 and IFI6; (s) CXCL8 and TUBB3; and (t) PPBP and CLEC5A.

[0024] In some embodiments, the present disclosure provides a panel of biomarkers for use in diagnosing precancerous colorectal advanced adenomas in a subject comprising a gene pair selected from the group consisting of: (a) MYC and CGREF1; (b) MYC and TCN1; (c) MYC and CEL; (d) HBA and MYC; (e) MYC and CPNE1; (f) MYC and TRIM29; (g) MYC and S100A9; (h) MYC and SPTBN2; (i) MYC and RHPN1; (j) HBA and CGREF1; (k) MYC and CEMIP; (1) HBA and TCN1; (m) MMP7 and CGREF1; (n) MYC and UBE2C; (o) EPHX4 and MYC; (p) MYC and CLEC5A; (q) HBA and PPBP; (r) MYC and PPBP; (s) MYC and CXCL3; and (t) TIMP1 and TCN1.

[0025] In another aspect, the present disclosure provides a kit or an integrated system of diagnosing colorectal cancer or precancerous colorectal advanced adenomas in a subject, comprising an agent for detecting in a biological sample obtained from the subject levels of the panel of biomarkers comprising one or more genes (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10 or more genes) selected from Table 1, Table 2 or Table 4.

[0026] In certain embodiment, the present disclosure provides a kit or an integrated system of diagnosing colorectal cancer in a subject, comprising an agent for detecting in a biological sample obtained from the subject levels of the panel of biomarkers comprising a gene pair selected from the group consisting of: (a) PPBP and TEMPI; (b) CXCL8 and PPBP; (c) PPBP and CBX2; (d) MMP7 and PPBP; (e) PPBP and MMP1; (f) HBA and MYC; (g) MYC and PPBP; (h) OLR1 and PPBP; (i) MMP3 and SPP1; (j) PPBP and ETV4; (k) OLR1 and TUBB3; (1) PPBP and SPP1; (m) CXCL8 and MMP3; (n) HBA and MMP3; (o) OLR1 and MMP3; (p) PPBP and CEACAM6; (q) PPBP and S100A9; (r) OLR1 and IFI6; (s) CXCL8 and TUBB3; and (t) PPBP and CLEC5A.

[0027] In certain embodiment, the present disclosure provides a kit or an integrated system of diagnosing precancerous colorectal advanced adenomas in a subject, comprising an agent for detecting in a biological sample obtained from the subject levels of the panel of biomarkers comprising a gene pair selected from the group consisting of: (a) MYC and CGREF1; (b) MYCand TCN1; (c) MYC and CEL; (d) HBA and MYC; (e) MYC and CPNE1; (f) MYC and TRIM29; (g) MYC and S100A9; (h) MYC and SPTBN2; (i) MYC and RHPN1; (j) HBA and CGREF1; (k) MYC and CEMIP; (1) HBA and TCN1; (m) MMP7 and CGREF1; (n) MYC and UBE2C; (o) EPHX4 and MYC; (p) MYC and CLEC5A; (q) HBA and PPBP; (r) MYC and PPBP; (s) MYC and CXCL3; and (t) TIMP1 and TCN1.

[0028] In certain embodiment, the agent is selected from the group consisting of: primers, nucleic acids and oligonucleotides.

[0029] In certain embodiment, the levels of the panel of biomarkers are measured by detecting mRNA expression levels or expressed protein levels.

[0030] In certain embodiment, the mRNA expression levels are detected using quantitative RT-PCR or Droplet Digital PCR.

[0031] In certain embodiment, the levels of the panel of biomarkers are determined by nucleic acid sequencing.

[0032] In another aspect, the present disclosure provides use of a biomarker-specific reagent in the manufacture of a kit for diagnosing colorectal cancer or precancerous colorectal advanced adenomas in a subject, wherein the biomarker-specific reagent specifically binds to the panel of biomarkers comprising one or more genes (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10 or more genes) selected from Table 1, Table 2 or Table 4.

[0033] In some embodiments, the panel of biomarkers comprises one or more genes (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10 or more genes) selected from the group consisting of: PPBP, OLR1, CXCL8, MMP3, HBA, TIMP1, MMP7, TGFBI, MYC, SPP1, NKD2, IL11, ETV4, CLEC5A, S100A9, TCN1, MMP1, IGF2, GDF15, TUBB3, RPP40, EPHX4, CBX2, ASCL2, EVA1A, PRSS33, RAET1L, and ULBPl.

[0034] In certain embodiment, the present disclosure provides use of a biomarker-specific reagent in the manufacture of a kit for diagnosing colorectal cancer in a subject, wherein the biomarker-specific reagent specifically binds to the panel of biomarkers comprising a gene pair selected from the group consisting of: (a) PPBP and TEMPI; (b) CXCL8 and PPBP; (c) PPBP and CBX2; (d) MMP7 and PPBP; (e) PPBP and MMP1; (f) HBA and MYC; (g) MYC and PPBP; (h) OLR1 and PPBP; (i) MMP3 and SPP1; (j) PPBP and ETV4; (k) OLR1 and TUBB3; (1) PPBP and SPP1; (m) CXCL8 and MMP3; (n) HBA and MMP3; (o) OLR1 and MMP3; (p) PPBP and CEACAM6; (q) PPBP and S100A9; (r) OLR1 and IFI6; (s) CXCL8 and TUBB3; and (t) PPBP and CLEC5A.

[0035] In certain embodiment, the present disclosure provides use of a biomarker-specific reagent in the manufacture of a kit for diagnosing precancerous colorectal advanced adenomas in a subject, wherein the biomarker-specific reagent specifically binds to the panel of biomarkers comprising a gene pair selected from the group consisting of: (a) MYC and CGREF1; (b) MYC and TCN1; (c) MYC and CEL; (d) HBA and MYC; (e) MYC and CPNE1; (f) MYC and TRIM29; (g) MYC and S100A9; (h) MYC and SPTBN2; (i) MYC and RHPN1; (j) HBA and CGREF1; (k) MYC and CEMIP; (1) HBA and TCN1; (m) MMP7 and CGREF1; (n) MYC and UBE2C; (o) EPHX4 and MYC; (p) MYC and CLEC5A; (q) HBA and PPBP; (r) MYC and PPBP; (s) MYC and CXCL3; and (t) TIMP1 and TCN1.

[0036] In certain embodiment, the subject is determined to have colorectal cancer or precancerous colorectal advanced adenomas by a machine learning classifier based on levels of the panel of biomarkers.

[0037] In certain embodiment, the levels of the panel of biomarkers are measured by detecting mRNA expression levels or expressed protein levels.

[0038] In certain embodiment, the mRNA expression levels are detected using quantitative RT-PCR or Droplet Digital PCR.

[0039] In certain embodiment, the levels of the panel of biomarkers are determined by nucleic acid sequencing.

[0040] In certain embodiment, the machine learning classifier is random forest or generalized linear model.

[0041] In certain embodiment, the biomarker-specific reagent comprises primers, nucleic acids and / or oligonucleotides.

[0042] In another aspect, the present disclosure provides a method for treating colorectal cancer or precancerous colorectal advanced adenomas in a subject. In one embodiment, the method comprises: administering to the subject a therapeutically effective amount of a drug useful for treating colorectal cancer or precancerous colorectal advanced adenomas, wherein the subject has been determined to have colorectal cancer or precancerous colorectal advanced adenomas by a machine learning classifier based on levels of the panel of biomarkers comprising one or more genes (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10 or more genes) selected from Table1, Table 2 or Table 4.

[0043] In some embodiments, the panel of biomarkers comprises one or more genes (e.g.,2, 3, 4, 5, 6, 7, 8, 9, 10 or more genes) selected from the group consisting of: PPBP, OLR1, CXCL8, MMP3, HBA, TIMP1, MMP7, TGFBI, MYC, SPP1, NKD2, IL11, ETV4, CLEC5A,S100A9, TCN1, MMP1, IGF2, GDF15, TUBB3, RPP40, EPHX4, CBX2, ASCL2, EVA1A, PRSS33, RAET1L, and ULBPE

[0044] In certain embodiment, the present disclosure provides a method for treating colorectal cancer in a subject, wherein the subject has been determined to have colorectal cancer by a machine learning classifier based on levels of the panel of biomarkers comprising a gene pair selected from the group consisting of: (a) PPBP and TIMP1; (b) CXCL8 and PPBP; (c) PPBP and CBX2; (d) MMP7 and PPBP; (e) PPBP and MMP1; (f) HBA and MYC; (g) MYC and PPBP; (h) OLR1 and PPBP; (i) MMP3 and SPP1; (j) PPBP and ETV4; (k) OLR1 and TUBB3; (1) PPBP and SPP1; (m) CXCL8 and MMP3; (n) HBA and MMP3; (o) OLR1 and MMP3; (p) PPBP and CEACAM6; (q) PPBP and S100A9; (r) OLR1 and IFI6; (s) CXCL8 and TUBB3; and (t) PPBP and CLEC5A.

[0045] In certain embodiment, the present disclosure provides a method for treating precancerous colorectal advanced adenomas in a subject, wherein the subject has been determined to have precancerous colorectal advanced adenomas by a machine learning classifier based on levels of the panel of biomarkers comprising a gene pair selected from the group consisting of: (a) MYC and CGREF1; (b) MYC and TCN1; (c) MYC and CEL; (d) HBA and MYC; (e) MYC and CPNE1; (f) MYC and TRIM29; (g) MYC and S100A9; (h) MYC and SPTBN2; (i) MYC and RHPN1; (j) HBA and CGREF1; (k) MYC and CEMIP; (1) HBA and TCN1; (m) MMP7 and CGREF1; (n) MYC and UBE2C; (o) EPHX4 and MYC; (p) MYC and CLEC5A; (q) HBA and PPBP; (r) MYC and PPBP; (s) MYC and CXCL3; and (t) TIMP1 and TCN1.

[0046] In certain embodiment, the machine learning classifier is random forests and / or generalized linear model.

[0047] In certain embodiment, the levels of the panel of biomarkers are determined by nucleic acid sequencing or RT-PCR or Droplet Digital PCR.BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The following drawings form part of the present specification and are included to further demonstrate certain aspects of the present disclosure. The disclosure may be better understood by reference to one or more of these drawings in combination with the detailed description of specific embodiments presented herein.

[0049] FIG. 1 shows the distribution of AUCs for all possible 2 gene combinations of the genes listed in Table 1 obtained by testing on a held out test set using random forest and GLM models.DETAILED DESCRIPTION OF THE INVENTION

[0050] Before the present disclosure is described in greater detail, it is to be understood that this disclosure is not limited to particular embodiments described, and as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present disclosure will be limited only by the appended claims.

[0051] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present disclosure, the preferred methods and materials are now described.

[0052] All publications and patents cited in this specification are herein incorporated by reference as if each individual publication or patent were specifically and individually indicated to be incorporated by reference and are incorporated herein by reference to disclose and describe the methods and / or materials in connection with which the publications are cited. The citation of any publication is for its disclosure prior to the filing date and should not be construed as an admission that the present disclosure is not entitled to antedate such publication by virtue of prior disclosure. Further, the dates of publication provided could be different from the actual publication dates that may need to be independently confirmed.

[0053] As will be apparent to those of skill in the art upon reading this disclosure, each of the individual embodiments described and illustrated herein has discrete components and features which may be readily separated from or combined with the features of any of the other several embodiments without departing from the scope or spirit of the present disclosure. Any recited method can be carried out in the order of events recited or in any other order that is logically possible.

[0054] Definitions

[0055] The following definitions are provided to assist the reader. Unless otherwise defined, all terms of art, notations and other scientific or medical terms or terminology used herein are intended to have the meanings commonly understood by those of skill in the chemical and medical arts. In some cases, terms with commonly understood meanings are defined herein for clarity and / or for ready reference, and the inclusion of such definitions herein should not necessarily be construed to represent a substantial difference over the definition of the term as generally understood in the art.

[0056] As used herein, the singular forms “a”, “an” and “the” include plural references unless the context clearly dictates otherwise.

[0057] As used herein, the term “administering” means providing a pharmaceutical agent or composition to a subject, and includes, but is not limited to, administering by a medical professional and self-administering.

[0058] The term “amount” or “level” generally refers to the quantity of a substance of interest. In the context of a panel of biomarker, a level of a panel of biomarkers refers to the quantity of the polynucleotides (e.g., RNA or DNA) of interest or the polypeptides of interest present in a sample. Such quantity may be expressed in the absolute terms, i.e., the total quantity of the polynucleotides or polypeptides in the sample, or in the relative terms, i.e., the concentration of the polynucleotides or polypeptides in the sample.

[0059] As used herein, the term “biomarker” refers to a detectable organic biomolecule associated with a particular phenotype or risk of developing a particular phenotype, such as a polynucleotide (e.g., RNA (e.g., mRNA) or DNA (e.g., cDNA)) or a polypeptide, which is differentially present in a biological sample taken from a subject having a certain condition (e.g., having colorectal cancer or precancerous colorectal advanced adenomas) as compared to a comparable biological sample taken from a subject who does not have such condition, such as a healthy subject or a non-cancer patient. For example, a biomarker can be a polynucleotide, such as RNA (e.g., mRNA), which is present at an elevated or decreased level in a biological sample (e.g., a tissue sample, a feces sample, or a blood sample) of a colorectal cancer patient compared to a comparable sample (e.g., a tissue sample, a feces sample, or a blood sample) of a subject with a negative diagnosis (e.g., a healthy subject).

[0060] As used herein, the term “cancer” refers to any diseases involving an abnormal cell growth and include all stages and all forms of the disease that affects any tissue, organ or cell in the body. The term includes all known cancers and neoplastic conditions, whether characterized as malignant, benign, soft tissue, or solid, and cancers of all stages and grades including pre- and post-metastatic cancers. In general, cancers can be categorized according to the tissue or organ from which the cancer is located or originated and morphology of cancerous tissues and cells. As used herein, cancer types include, without limitation, acute lymphoblastic leukemia (ALL), acute myeloid leukemia, adrenocortical carcinoma, anal cancer, astrocytoma, childhood cerebellar or cerebral, basal-cell carcinoma, bile duct cancer, bladder cancer, bone tumor, brain cancer, cerebellar astrocytoma, cerebral astrocytoma / malignant glioma, ependymoma, medulloblastoma, supratentorial primitive neuroectodermal tumors,visual pathway and hypothalamic glioma, breast cancer, Burkitt's lymphoma, cervical cancer, chronic lymphocytic leukemia, chronic myelogenous leukemia, colorectal cancer, emphysema, endometrial cancer, ependymoma, esophageal cancer, Ewing's sarcoma, retinoblastoma, gastric (stomach) cancer, glioma, head and neck cancer, heart cancer, Hodgkin lymphoma, islet cell carcinoma (endocrine pancreas), Kaposi sarcoma, kidney cancer (renal cell cancer), laryngeal cancer, leukemia, liver cancer, lung cancer, neuroblastoma, non-Hodgkin lymphoma, ovarian cancer, pancreatic cancer, pharyngeal cancer, prostate cancer, rectal cancer, renal cell carcinoma (kidney cancer), retinoblastoma, Ewing family of tumors, skin cancer, stomach cancer, testicular cancer, throat cancer, thyroid cancer, vaginal cancer.

[0061] The term “colon cancer” used interchangeably with the term “colorectal cancer” or “rectal cancer” refers to any cancerous neoplasia of the colon (including the rectum and appendix). Many colorectal cancers arise from precancerous colorectal advanced adenomas or adenomatous polyps, which are usually benign, but some may develop into cancer over time. The diagnosis of localized colon cancer is often through colonoscopy. Once localized colon cancer is diagnosed, it is usually surgically removed and then treated with chemotherapy. Precancerous colorectal advanced adenomas or colorectal adenomatous polyps are a risk factor for colorectal cancer. The removal of colorectal adenomatous polyps at the time of colonoscopy would reduce the risk of having colorectal cancer. In addition, clinical data has shown that early detection and curative surgical resection of colorectal cancer will significantly improve survival rates.

[0062] It is noted that in this disclosure, terms such as “comprises”, “comprised”, “comprising”, “contains”, “containing” and the like have the meaning attributed in United States Patent law; they are inclusive or open-ended and do not exclude additional, un-recited elements or method steps. Terms such as “consisting essentially of’ and “consists essentially of’ have the meaning attributed in United States Patent law; they allow for the inclusion of additional ingredients or steps that do not materially affect the basic and novel characteristics of the claimed disclosure. The terms “consists of’ and “consisting of’ have the meaning ascribed to them in United States Patent law; namely that these terms are close ended.

[0063] The terms “assessing”, “assaying”, “measuring” and “detecting” can be used interchangeably and refer to both quantitative and semi-quantitative determinations. Where either a quantitative and semi-quantitative determination is intended, the phrase “measuring a level” of a polynucleotide or polypeptide of interest or “detecting” a polynucleotide or polypeptide of interest can be used.

[0064] The term “hybridizing” refers to the binding, duplexing, or hybridizing of a nucleic acid molecule preferentially to a particular nucleotide sequence under stringent conditions. The term “stringent conditions” refers to conditions under which a probe will hybridize preferentially to its target subsequence, and to a lesser extent to, or not at all to, other sequences in a mixed population (e.g., a cell lysate or DNA preparation from a tissue biopsy). A “stringent hybridization” and “stringent hybridization wash conditions” in the context of nucleic acid hybridization (e.g., as in array, microarray, Southern or northern hybridizations) are sequence dependent, and are different under different environmental parameters. An extensive guide to the hybridization of nucleic acids is found in, e.g., Tijssen Laboratory Techniques in Biochemistry and Molecular Bio logy — Hybridization with Nucleic Acid Probes part I, Ch. 2, “Overview of principles of hybridization and the strategy of nucleic acid probe assays,” (1993) Elsevier, N.Y. Generally, highly stringent hybridization and wash conditions are selected to be about 5° C lower than the thermal melting point (Tm) for the specific sequence at a defined ionic strength and pH. The Tm is the temperature (under defined ionic strength and pH) at which 50% of the target sequence hybridizes to a perfectly matched probe. Very stringent conditions are selected to be equal to the Tm for a particular probe. An example of stringent hybridization conditions for hybridization of complementary nucleic acids which have more than 100 complementary residues on an array or on a filter in a Southern or northern blot is 42° C using standard hybridization solutions (see, e.g., Sambrook and Russell Molecular Cloning: A Laboratory Manual (3rd ed.) Vol. 1-3 (2001) Cold Spring Harbor Laboratory, Cold Spring Harbor Press, NY). An example of highly stringent wash conditions is 0.15 M NaCl at 72° C for about 15 minutes. An example of stringent wash conditions is a 0.2* SSC wash at 65° C for 15 minutes. Often, a high stringency wash is preceded by a low stringency wash to remove background probe signal. An example medium stringency wash for a duplex of, e.g., more than 100 nucleotides, is lx SSC at 45° C for 15 minutes. An example of a low stringency wash for a duplex of, e.g., more than 100 nucleotides, is 4xSSC to 6xSSC at 40° C for 15 minutes.

[0065] The term “nucleic acid” and “polynucleotide” are used interchangeably and refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof. Polynucleotides may have any three-dimensional structure, and may perform any function, known or unknown. Non-limiting examples of polynucleotides include a gene, a gene fragment, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozymes, cDNA, shRNA, single-stranded short or long RNAs, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence,control regions, isolated RNA of any sequence, nucleic acid probes, and primers. The nucleic acid molecule may be linear or circular.

[0066] As used herein, the term “precancerous colorectal advanced adenoma” or “advanced adenoma” refers the highest risk stage of precancerous lesions. Typically, a precancerous colorectal advanced adenoma is a precancerous polyp in the colon or rectum that has at least one of the following characteristics: it’s at least 10 mm in size, it has a significant villous component, and it has high-grade dysplasia.

[0067] In general, a “protein” is a polypeptide (i.e., a string of at least two amino acids linked to one another by peptide bonds). Proteins may include moieties other than amino acids (e.g., may be glycoproteins) and / or may be otherwise processed or modified. Those of ordinary skill in the art will appreciate that a “protein” can be a complete polypeptide chain as produced by a cell (with or without a signal sequence), or can be a functional portion thereof. Those of ordinary skill will further appreciate that a protein can sometimes include more than one polypeptide chain, for example linked by one or more disulfide bonds or associated by other means.

[0068] As used herein, the term “subject” refers to a human or any non-human animal (e.g., mouse, rat, rabbit, dog, cat, cattle, swine, sheep, horse or primate). A human includes pre and post-natal forms. In many embodiments, a subject is a human being. A subject can be a patient, which refers to a human presenting to a medical provider for diagnosis or treatment of a disease. The term “subject” is used herein interchangeably with “individual” or “patient”. A subject can be afflicted with or is susceptible to a disease or disorder but may or may not display symptoms of the disease or disorder.

[0069] As used herein, the term “therapeutically effective amount” means the amount of agent that is sufficient to prevent, treat, reduce and / or ameliorate the symptoms and / or underlying causes of any disorder or disease, or the amount of an agent sufficient to produce a desired effect on a cell. In one embodiment, a “therapeutically effective amount” is an amount sufficient to reduce or eliminate a symptom of a disease. In another embodiment, a therapeutically effective amount is an amount sufficient to overcome the disease itself.

[0070] The term “treatment,” “treat,” or “treating” refers to a method of reducing the effects of a cancer (e.g., breast cancer, lung cancer, ovarian cancer or the like) or symptom of cancer. Thus, in the disclosed method, treatment can refer to a 10%, 20%, 30%, 40%, 50%, 60%, 70%), 80%), 90%), or 100% reduction in the severity of a cancer or symptom of the cancer. For example, a method of treating a disease is considered to be a treatment if there is a10% reduction in one or more symptoms of the disease in a subject as compared to a control. Thus, the reduction can be a 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100% or any percent reduction between 10 and 100% as compared to native or control levels. It is understood that treatment does not necessarily refer to a cure or complete ablation of the disease, condition, or symptoms of the disease or condition.

[0071] As used herein, the term “biological sample” refers to a sample obtained from a subject. The sample may comprise cells or can be cell free. The biological sample can be stool, blood, sputum, urine, saliva, serum cerebrospinal, tissue, cells, secretions or the like. A noninvasive sampling method, such as a stool sample (i.e., a feces sample) or a blood sample can be obtained and used in the methods of the present disclosure.

[0072] Panel of Biomarkers

[0073] Identification of specific and sensitive biomarkers suitable for the development of improved early cancer diagnosis is one of the major focuses of cancer research. While there may be numerous potential tumor biomarkers identified using various techniques such as DNA / RNA microarray analysis and deep sequencing, reliable biomarker panels for diagnosing early colorectal cancer or precancerous colorectal advanced adenomas are still lacking. The present disclosure provides a novel panel of biomarkers useful for diagnosing colorectal cancer or precancerous colorectal advanced adenomas. The methods and compositions described herein are based, in part, on the discovery of a panel of biomarkers correlated with colorectal cancer or precancerous colorectal advanced adenomas in a patient.

[0074] The methods provided herein comprise measuring a panel of biomarkers with the specificity and sensitivity required for managing and diagnosing subjects that have or may have a colon cancer or precancerous colorectal advanced adenomas. The methods provided herein can be used in an outpatient clinic or inpatient environment. Outpatient clinical diagnostics are useful to reduce costs of unnecessary, often invasive or painful, procedures.

[0075] In one aspect, the present disclosure provides a method of diagnosing colorectal cancer or precancerous colorectal advanced adenomas in a subject. In certain embodiments, the method comprises: measuring a biological sample obtained from the subject levels (e.g., expression levels, such as mRNA levels) of a panel of biomarkers, evaluating the measured levels of the panel of biomarkers, and determining that the subject is healthy or has precancerous colorectal advanced adenomas and / or colorectal cancer. In certain embodiments, the subject is a human or non-human.

[0076] In another aspect, the present disclosure provides a panel of biomarkers for use in diagnosing colorectal cancer or precancerous colorectal advanced adenomas in a subject comprising one or more genes (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10 or more genes) selected from Table 1. In some embodiments, the panel of biomarkers comprises one or more genes (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10 or more genes) selected from the group consisting of: PPBP, OLR1, CXCL8, MMP3, HBA, TIMP1, MMP7, TGFBI, MYC, SPP1, NKD2, IL11, ETV4, CLEC5A, S100A9, TCN1, MMP1, IGF2, GDF15, TUBB3, RPP40, EPHX4, CBX2, ASCL2, EVA1A, PRSS33, RAET1L, and ULBP1. In some embodiments, the panel of biomarkers comprises a gene pair selected from the Table 2. In some embodiments, the panel of biomarkers comprises a gene pair selected from the Table 4.

[0077] As used herein, the term “panel of biomarkers” refers to a selection of at least two biomarkers. The panel can comprise from 2 to 12 (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 or 12) biomarkers. In certain embodiments, the biomarkers include certain genes. In certain embodiments, the biomarkers are selected from Table 1. The mRNA and protein sequences are referred to Ensemble Number.

[0078] Measuring Biomarkers

[0079] The methods of the present disclosure involve detecting or measuring at least a subset of the predicting biomarkers disclosed herein (for example, in Table 1), in a biological sample obtained from a subject suspected of having or at risk of having cancers or precancerous polyps.

[0080] Sample Preparation

[0081] The biological sample can be a tissue sample, a urine sample, a lymph sample, a saliva sample, a blood sample or a stool sample. The biological samples often contain tumor cells or debris thereof (e.g., cancer cell apoptotic products). A noninvasive sampling method, such as a stool sample (i.e., a feces sample) or a blood sample can be obtained and used in the methods of the present disclosure.

[0082] In certain embodiments, a tissue sample can be processed to perform in situ hybridization. For example, the tissue sample can be paraffin-embedded before fixing on a glass microscope slide, and then deparaffinized with a solvent, typically xylene.

[0083] In certain embodiments, the biological sample is a stool sample.

[0084] In certain embodiments, the method further comprises isolating the nucleic acid, e.g., DNA or RNA from the sample. Various methods of extraction are suitable for isolating the DNA or RNA from cells or tissues, such as phenol and chloroform extraction, and variousother methods as described in, for example, Ausubel et al., Current Protocols of Molecular Biology (1997) John Wiley & Sons, and Sambrook and Russell, Molecular Cloning: A Laboratory Manual 3rded. (2001).

[0085] Commercially available kits can also be used to isolate DNA and / or RNA, including for example, the NucliSens extraction kit (Biomerieux, Marcy 1'Etoile, France), QIAamp™ mini blood kit, Agencourt Genfind™, Rneasy® mini columns (Qiagen), PureLink® RNA mini kit (Thermo Fisher Scientific), and Eppendorf Phase Lock Gels™. A skilled person can readily extract or isolate RNA or DNA following the manufacturer’s protocol.

[0086] Methods of Measuring Biomarkers

[0087] The biomarkers disclosed herein can be detected in the level of DNA (e.g., genomic DNA) or RNA (e.g., mRNA) using proper methods known in the art including, without limitation, amplification assay, hybridization assay, and sequencing assay. The gene expression level can be detected in the RNA (e.g., mRNA) level or protein level using proper methods known in the art. In certain embodiments, the levels of a panel of biomarkers are measured by detecting mRNA expression levels. In certain embodiments, the mRNA expression levels are detected using quantitative RT-PCR or Droplet Digital PCR. In certain embodiments, the levels of the panel of biomarkers are determined by nucleic acid sequencing. In certain embodiments, the levels of a panel of biomarkers are measured by detecting expressed protein levels. In certain embodiments, the expressed protein levels are detected using Immunoassays.

[0088] Sequencing methods

[0089] Sequencing methods useful in the measurement of the biomarkers involves sequencing of the target nucleic acid. Any sequencing known in the art can be used to detect the biomarkers of interest. In general, sequencing methods can be categorized to traditional or classical methods and high throughput sequencing (next generation sequencing). Traditional sequencing methods include Maxam-Gilbert sequencing (also known as chemical sequencing) and Sanger sequencing (also known as chain-termination methods).

[0090] High throughput sequencing, or next generation sequencing, by using methods distinguished from traditional methods, such as Sanger sequencing, is highly scalable and able to sequence the entire genome or transcriptome at once. High throughput sequencing involves sequencing-by-synthesis, sequencing-by-ligation, and ultra-deep sequencing (such as described in Marguiles et al., Nature 437 (7057): 376-80 (2005)). Sequence-by-synthesis involves synthesizing a complementary strand of the target nucleic acid by incorporating labeled nucleotide or nucleotide analog in a polymerase amplification. Immediately after orupon successful incorporation of a label nucleotide, a signal of the label is measured, and the identity of the nucleotide is recorded. The detectable label on the incorporated nucleotide is removed before the incorporation, detection and identification steps are repeated. Examples of sequence-by-synthesis methods are known in the art, and are described for example in U.S. Pat. No. 7,056,676, U.S. Pat. No. 8,802,368 and U.S. Pat. No. 7,169,560, the contents of which are incorporated herein by reference. Sequencing-by-synthesis may be performed on a solid surface (or a microarray or a chip) using fold-back PCR and anchored primers. Target nucleic acid fragments can be attached to the solid surface by hybridizing to the anchored primers, and bridge amplified. This technology is used, for example, in the Illumina® sequencing platform.

[0091] Pyrosequencing involves hybridizing the target nucleic acid regions to a primer and extending the new strand by sequentially incorporating deoxynucleotide triphosphates corresponding to the bases A, C, G, and T (U) in the presence of a polymerase. Each base incorporation is accompanied by release of pyrophosphate, converted to ATP by sulfurylase, which drives synthesis of oxyluciferin and the release of visible light. Since pyrophosphate release is equimolar with the number of incorporated bases, the light given off is proportional to the number of nucleotides adding in any one step. The process is repeated until the entire sequence is determined.

[0092] In certain embodiments, the biomarkers described herein are detected by whole transcriptome shotgun sequencing (RNA sequencing). The method of RNA sequencing has been described (see Wang Z, Gerstein M and Snyder M, Nature Review Genetics (2009) 10:57- 63; Maher CA et al., Nature (2009) 458:97-101; Kukurba K & Montgomery SB, Cold Spring Harbor Protocols (2015) 2015(11): 951-969).

[0093] Amplification assay

[0094] A nucleic acid amplification assay involves copying a target nucleic acid (e.g., DNA or RNA), thereby increasing the number of copies of the amplified nucleic acid sequence. Amplification may be exponential or linear. Exemplary nucleic acid amplification methods include, but are not limited to, amplification using the polymerase chain reaction ("PCR", see U.S. Patents 4,683,195 and 4,683,202; PCR Protocols: A Guide To Methods And Applications (Innis et al., eds, 1990)), reverse transcriptase polymerase chain reaction (RT-PCR), quantitative real-time PCR (qRT-PCR); quantitative PCR, such as TaqMan®, nested PCR, Digital PCR, such as Droplet Digital PCR (ddPCR), ligase chain reaction (See Abravaya, K., et al., Nucleic Acids Research, 23:675-682, (1995), branched DNA signal amplification (see, Urdea, M. S., et al., AIDS, 7 (suppl 2):S11-S14, (1993), amplifiable RNA reporters, Q-betareplication (see Lizardi et al., Biotechnology (1988) 6: 1197), transcription-based amplification (see, Kwoh et al., Proc. Natl. Acad. Sci. USA (1989) 86: 1173-1177), boomerang DNA amplification, strand displacement activation, cycling probe technology, self-sustained sequence replication (Guatelli et al., Proc. Natl. Acad. Sci. USA (1990) 87: 1874-1878), rolling circle replication (U.S. Patent No. 5,854,033), isothermal nucleic acid sequence based amplification (NASBA), and serial analysis of gene expression (SAGE). Droplet Digital PCR (ddPCR) is a method for performing digital PCR that is based on water-oil emulsion droplet technology. A sample is fractionated into many droplets, and PCR amplification of the template molecules occurs in each individual droplet. ddPCR technology uses reagents and workflows similar to those used for most standard TaqMan probe-based assays. The massive sample partitioning is a key aspect of the ddPCR technique.

[0095] In certain embodiments, the nucleic acid amplification assay is a PCR-based method. PCR is initiated with a pair of primers that hybridize to the target nucleic acid sequence to be amplified, followed by elongation of the primer by polymerase which synthesizes the new strand using the target nucleic acid sequence as a template and dNTPs as building blocks. Then the new strand and the target strand are denatured to allow primers to bind for the next cycle of extension and synthesis. After multiple amplification cycles, the total number of copies of the target nucleic acid sequence can increase exponentially.

[0096] In certain embodiments, intercalating agents that produce a signal when intercalated in double stranded DNA may be used. Exemplary agents include SYBR GREEN™ and SYBR GOLD™. Since these agents are not template-specific, it is assumed that the signal is generated based on template-specific amplification. This can be confirmed by monitoring signal as a function of temperature because melting point of template sequences will generally be much higher than, for example, primer-dimers, etc.

[0097] In certain embodiments, a detectably labeled primer or a detectably labeled probe can be used, to allow detection of the biomarkers corresponding to that primer or probe. In certain embodiments, multiple labeled primers or labeled probes with different detectable labels can be used to allow simultaneous detection of multiple biomarkers.

[0098] Hybridization assay

[0099] Nucleic acid hybridization assays use probes to hybridize to the target nucleic acid, thereby allowing detection of the target nucleic acid. Non-limiting examples of hybridization assay include Northern blotting, Southern blotting, in situ hybridization, microarray analysis, and multiplexed hybridization-based assays.

[0100] In certain embodiments, the probes for hybridization assay are detectably labeled. In certain embodiments, the nucleic acid-based probes for hybridization assay are unlabeled. Such unlabeled probes can be immobilized on a solid support such as a microarray and can hybridize to the target nucleic acid molecules which are detectably labeled.

[0101] In certain embodiments, hybridization assays can be performed by isolating the nucleic acids (e.g., RNA or DNA), separating the nucleic acids (e.g., by gel electrophoresis) followed by transfer of the separated nucleic acid on suitable membrane filters (e.g., nitrocellulose filters), where the probes hybridize to the target nucleic acids and allows detection. See, for example, Molecular Cloning: A Laboratory Manual, J. Sambrook et al., eds., 2nd edition, Cold Spring Harbor Laboratory Press, 1989, Chapter 7. The hybridization of the probe and the target nucleic acid can be detected or measured by methods known in the art. For example, autoradiographic detection of hybridization can be performed by exposing hybridized filters to photographic film.

[0102] In some embodiments, hybridization assays can be performed on microarrays. Microarrays provide a method for the simultaneous measurement of the levels of large numbers of target nucleic acid molecules. The target nucleic acids can be RNA, DNA, cDNA reverse transcribed from mRNA, or chromosomal DNA. The target nucleic acids can be allowed to hybridize to a microarray comprising a substrate having multiple immobilized nucleic acid probes arrayed at a density of up to several million probes per square centimeter of the substrate surface. The RNA or DNA in the sample is hybridized to complementary probes on the array and then detected by laser scanning. Hybridization intensities for each probe on the array are determined and converted to a quantitative value representing relative levels of the RNA or DNA. See, U.S. Patent Nos. 6,040,138, 5,800,992 and 6,020,135, 6,033,860, and 6,344,316.

[0103] Techniques for the synthesis of these arrays using mechanical synthesis methods are described in, e.g., U.S. Patent No. 5,384,261. Although a planar array surface is often employed the array may be fabricated on a surface of virtually any shape or even a multiplicity of surfaces. Arrays may be peptides or nucleic acids on beads, gels, polymeric surfaces, fibers such as fiber optics, glass or any other appropriate substrate, see U.S. Patent Nos. 5,770,358, 5,789,162, 5,708,153, 6,040,193 and 5,800,992. Arrays may be packaged in such a manner as to allow for diagnostics or other manipulation of an all-inclusive device. Useful microarrays are also commercially available, for example, microarrays from Affymetrix, from Nano String Technologies, QuantiGene 2.0 Multiplex Assay from Panomics.

[0104] In certain embodiments, hybridization assays can be in situ hybridization assay. In situ hybridization assay is useful to detect the presence of gene mutations. Probes useful for in situ hybridization assay can be mutation specific probes, which hybridize to a specific gene mutation to detect the presence or absence of the specific mutation of interest. Methods for use of unique sequence probes for in situ hybridization are described in U.S. Pat. No. 5,447,841, incorporated herein by reference. Probes can be viewed with a fluorescence microscope and an appropriate filter for each fluorophore, or by using dual or triple band-pass filter sets to observe multiple fhiorophores. See, e.g., U.S. Pat. No. 5,776,688 to Bittner, et al., which is incorporated herein by reference. Any suitable microscopic imaging method can be used to visualize the hybridized probes, including automated digital imaging systems. Alternatively, techniques such as flow cytometry can be used to examine the hybridization pattern of the probes.

[0105] Immunoassay

[0106] Immunoassays used herein typically involves using antibodies that specifically bind to biomarker protein. Such antibodies can be obtained using methods known in the art (see, e.g., Huse et al., Science (1989) 246: 1275-1281; Ward et al, Nature (1989) 341 :544-546), or can be obtained from commercial sources. Examples of immunoassays include, without limitation, Western blotting, enzyme-linked immunosorbent assay (ELISA), enzyme immunoassay (EIA), radioimmunoassay (RIA), immunoprecipitations, sandwich assays, competitive assays, immunofluorescent staining and imaging, immunohistochemistry (IHC), and fluorescent activating cell sorting (FACS). For a review of immunological and immunoassay procedures, see Basic and Clinical Immunology (Stites & Terr eds., 7thed. 1991). Moreover, the immunoassays can be performed in any of several configurations, which are reviewed extensively in Enzyme Immunoassay (Maggio, ed., 1980); and Harlow & Lane, supra. For a review of the general immunoassays, see also Methods in Cell Biology: Antibodies in Cell Biology, volume 37 (Asai, ed. 1993); Basic and Clinical Immunology (Stites & Terr, eds., 7thed. 1991).

[0107] Any of the assays and methods provided herein for the measurement of the gene expression level can be adapted or optimized for use in automated and semi-automated systems or point of care assay systems.

[0108] The gene expression level described herein can be normalized using a proper method known in the art. For example, the gene expression level can be normalized to a standard level of a standard marker, which can be predetermined, determined concurrently, ordetermined after a sample is obtained from the subject. The standard marker can be run in the same assay or can be a known standard marker from a previous assay. For another example, the gene expression level can be normalized to an internal control which can be an internal marker, or an average level or a total level of a plurality of internal markers.

[0109] The level of mRNA expression of each of the biomarkers described herein can be normalized to a reference level for a control gene. The control value can be predetermined, determined concurrently, or determined after a sample is obtained from the subject. The standard can be run in the same assay or can be a known standard from a previous assay. In the cases when the level of RNA expression is determined by RNA sequencing, the level of RNA expression of each of the biomarkers can be normalized to the total reads of the sequencing. The normalized levels of mRNA expression of the biomarker genes can be transformed into a score, e.g., using the methods and models described herein.

[0110] Methods for Predicting Colorectal Cancer or Precancerous Colorectal advanced adenomas

[0111] In some embodiments, the method disclosed herein comprises classifying the subject as 1) healthy or 2) having colorectal cancer or precancerous colorectal adenoma based on the measured levels of the biomarker panel. In some embodiments, the method comprises evaluating the measured levels of the biomarker panel by a machine learning classifier and determining that the subject is 1) healthy or 2) has colorectal cancer or precancerous colorectal adenoma.

[0112] In statistics, classification is the problem of identifying which of a set of categories an observation (or observations) belongs to. As used herein, the term “classification” used interchangeably with the term “classifying” refers to the identification of the subject as 1) being healthy or 2) having colorectal cancer or precancerous colorectal advanced adenomas based on the measured levels of the biomarker panel. A “classifier” refers to an algorithm that implements the classification.

[0113] As used herein, the term “machine learning” refers to a computer-implemented technique that gives computer systems the ability to progressively improve performance on a specific task with data, i.e., to learn from the data, without being explicitly programmed. Machine learning technique adopts algorithms that can learn from and make prediction on data through building a model, i.e., a description of a system using mathematical concepts, from sample inputs. A core objective of machine learning is to generalize from the experience, i.e., to perform accurately on new data after having experienced a learning data set.

[0114] Machine learning models can be categorized as either supervised or unsupervised. Supervised learning involves learning a function that maps an input to an output based on example input-output pairs. In the context of biomedical diagnosis or prognosis, machine learning techniques generally involves supervised learning process, in which the computer is presented with example inputs (e.g., signature of gene expression) and their desired outputs (e.g., likelihood of having colorectal cancer or precancerous colorectal advanced adenomas) to learn a general rule that maps inputs to outputs. Different models, i.e., hypothesis, can be employed in the generalization process. For the best performance in the generalization, the complexity of the hypothesis should match the complexity of the function underlying the data.

[0115] Supervised models include but not limited to logistic regression, support vector machine, decision trees, random forest, artificial neural network, linear regression, elastic net and naive bayes. In certain embodiments, the machine learning classifier is random forest or generalized linear model.

[0116] Logistic Regression

[0117] The simplest idea of linear regression is to find a line that best fits the data. Extensions of linear regression include multiple linear regression (e.g., finding a plane of best fit) and polynomial regression (e.g., finding a curve of best fit). Logistic regression is similar to linear regression but is used to model the probability of a finite number of outcomes, typically two.

[0118] Support Vector Machine

[0119] A Support Vector Machine (SVM) is a supervised classification technique that, at the most fundamental level, find a hyperplane or a boundary between two classes of data that maximizes the margin between the two classes. There are many planes that can separate the two classes, but only one plane can maximize the margin or distance between the classes.

[0120] Decision Tree

[0121] A decision tree is a decision support tool that uses a tree-like model of decisions and their possible consequences, including chance event outcomes, resource costs, and utility. Typically, a decision tree is a flowchart-like structure in which each internal node represents a “test” on an attribute, each branch represents the outcome of the test, and each leaf node represents a class label (decision taken after computing all attribute). The paths from root to leaf represent classification rules. Decision trees are intuitive and easy to build but fall short when it comes to accuracy.

[0122] Random Forest

[0123] Random forests are an ensemble learning technique that builds off of decision trees. Random forests involve creating multiple decision trees using bootstrapped datasets of the original data and randomly selecting a subset of variables at each step of the decision tree. The model then selects the mode of all of the predictions of each decision tree. By relying on a “majority wins” model, it reduces the risk of error from an individual tree.

[0124] Generalized Linear Model

[0125] A generalized linear model (GLM) is a flexible generalization of ordinary linear regression by allowing the linear model to be related to the response variable via a link function and by allowing the magnitude of the variance of each measurement to be a function of its predicted value. Ordinary linear regression predicts the expected value of a given unknown quantity (the response variable, a random variable) as a linear combination of a set of observed values (predictors). This implies that a constant change in a predictor leads to a constant change in the response variable (i.e. a linear-response model). This is appropriate when the response variable can vary, to a good approximation, indefinitely in either direction, or more generally for any quantity that only varies by a relatively small amount compared to the variation in the predictive variables, e.g. human heights. However, these assumptions are inappropriate for some types of response variables. For example, in cases where the response variable is expected to be always positive and varying over a wide range, constant input changes lead to geometrically (i.e. exponentially) varying, rather than constantly varying, output changes. Similarly, a model that predicts a probability of making a yes / no choice (a Bernoulli variable) is even less suitable as a linear-response model, since probabilities are bounded on both ends (they must be between 0 and 1). Generalized linear models cover all these situations by allowing for response variables that have arbitrary distributions (rather than simply normal distributions), and for an arbitrary function of the response variable (the link function) to vary linearly with the predictors (rather than assuming that the response itself must vary linearly).

[0126] Artificial Neural Network

[0127] Artificial neural networks (ANNs), usually simply called neural networks (NNs), are inspired by the biological neural networks that constitute animal brains. An ANN is based on a collection of connected units or nodes called artificial neurons, which loosely model the neurons in a biological brain. Each connection, like the synapses in a biological brain, can transmit a signal to other neurons. An artificial neuron receives a signal then processes it and can signal neurons connected to it. The “signal” at a connection is a real number, and the output of each neuron is computed by some non-linear function of the sum of its inputs. Theconnections are called edges. Neurons and edges typically have a weight that adjusts as learning proceeds. The weight increases or decreases the strength of the signal at a connection. Neurons may have a threshold such that a signal is sent only if the aggregate signal crosses that threshold. Typically, neurons are aggregated into layers. Different layers may perform different transformations on their inputs. Signals travel from the first layer (the input layer) to the last layer (the output layer), possibly after traversing the layers multiple times.

[0128] Elastic net

[0129] Elastic net is a regularized regression method that linearly combines the Li and L2 penalties of the LASSO (least absolute shrinkage and selection operator) and ridge methods. The LASSO method constructs a linear model, which penalizes the regression coefficients with an Li penalty, shrinking many of them to zero. Any features which have non-zero regression coefficients are “selected” by the lasso algorithm. As a result, the LASSO method performs both feature selection and regularization in order to enhance the prediction accuracy and interpretability of the resulting machine learning model. Elastic net is one of the improvements to the LASSO. Elastic net regularization combines the Li penalty of LASSO with the L2 penalty of ridge regression.

[0130] Unsupervised learning, on the other hand, is to draw inferences and find patterns from input data without references to labeled outcomes. Two main methods used in unsupervised learning include clustering and dimensionality reduction.

[0131] Clustering is an unsupervised technique that involves the grouping or clustering of data points. Common clustering algorithm include k-means clustering, hierarchical clustering, mean shift clustering, and density-based clustering.

[0132] Dimensionality reduction is a process of reducing the number of random variables to obtain a set of principle variables. Common dimensionality reduction algorithm include principal component analysis (PCA), regularized regression and Boruta.

[0133] The use of the machine learning classifier can classify the sample as 1) a healthy or 2) having colorectal cancer and / or precancerous lesions with a sensitivity, specificity, positive predictive value, negative predictive value, and / or overall accuracy of at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%.

[0134] In certain embodiments, the methods of the present disclosure further comprise sending the 1) a healthy or 2) having colorectal cancer and / or precancerous lesions classification results to a clinician. In some embodiments, the method of the present disclosureprovides a diagnosis or prognosis in the form of a probability that the subject 1) is healthy or 2) having colorectal cancer and / or precancerous lesions. For example, the subject may have about a 0%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95% or greater probability of 1) being healthy or 2) having colorectal cancer and / or precancerous lesions.

[0135] Computer-implemented Methods, Systems and Devices

[0136] Any of the methods described herein may be totally or partially performed with a computer system including one or more processors, which can be configured to perform the steps. Thus, embodiments are directed to computer systems configured to perform the steps of any of the methods described herein, potentially with different components performing a respective step or a respective group of steps. Although presented as numbered steps, steps of methods herein can be performed at a same time or in a different order. Additionally, portions of these steps may be used with portions of other steps from other methods. Also, all or portions of a step may be optional. Any of the steps of any of the methods can be performed with modules, circuits, or other means for performing these steps.

[0137] Any of the computer systems mentioned herein may utilize any suitable number of subsystems. In some embodiments, a computer system includes a single computer apparatus, where the subsystems can be the components of the computer apparatus. In other embodiments, a computer system can include multiple computer apparatuses, each being a subsystem, with internal components. The subsystems can be interconnected via a system bus. Additional subsystems include, for examples, a printer, keyboard, storage device(s), monitor, which is coupled to display adapter, and others. Peripherals and input / output (I / O) devices, which couple to I / O controller, can be connected to the computer system by any number of means known in the art, such as serial port. For example, serial port or external interface (e.g. Ethernet, Wi-Fi, etc.) can be used to connect computer system to a wide area network such as the Internet, a mouse input device, or a scanner. The interconnection via system bus allows the central processor to communicate with each subsystem and to control the execution of instructions from system memory or the storage device(s) (e.g., a fixed disk, such as a hard drive or optical disk), as well as the exchange of information between subsystems. The system memory and / or the storage device(s) may embody a computer readable medium. Any of the data mentioned herein can be output from one component to another component and can be output to the user.

[0138] A computer system can include a plurality of the same components or subsystems, e.g., connected together by external interface or by an internal interface. In some embodiments,computer systems, subsystem, or apparatuses can communicate over a network. In such instances, one computer can be considered a client and another computer a server, where each can be part of a same computer system. A client and a server can each include multiple systems, subsystems, or components.

[0139] It should be understood that any of the embodiments of the present disclosure can be implemented in the form of control logic using hardware (e.g., an application specific integrated circuit or field programmable gate array) and / or using computer software with a generally programmable processor in a modular or integrated manner. As used herein, a processor includes a multi-core processor on a same integrated chip, or multiple processing units on a single circuit board or networked. Based on the disclosure and teachings provided herein, a person of ordinary skill in the art will know and appreciate other ways and / or methods to implement embodiments of the present disclosure using hardware and a combination of hardware and software.

[0140] Any of the software components or functions described in this application may be implemented as software code to be executed by a processor using any suitable computer language such as, for example, Java, C++ or Perl using, for example, conventional or object- oriented techniques. The software code may be stored as a series of instructions or commands on a computer readable medium for storage and / or transmission, suitable media include random access memory (RAM), a read only memory (ROM), a magnetic medium such as a hard-drive or a floppy disk, or an optical medium such as a compact disk (CD) or DVD (digital versatile disk), flash memory, and the like. The computer readable medium may be any combination of such storage or transmission devices.

[0141] Such programs may also be encoded and transmitted using carrier signals adapted for transmission via wired, optical, and / or wireless networks conforming to a variety of protocols, including the Internet. As such, a computer readable medium according to an embodiment of the present invention may be created using a data signal encoded with such programs. Computer readable media encoded with the program code may be packaged with a compatible device or provided separately from other devices (e.g., via Internet download). Any such computer readable medium may reside on or within a single computer product (e.g. a hard drive, a CD, or an entire computer system), and may be present on or within different computer products within a system or network. A computer system may include a monitor, printer, or other suitable display for providing any of the results mentioned herein to a user.

[0142] Kits

[0143] In another aspect, the present disclosure provides kits or an integrated system for use in the methods described above. The kits may comprise any or all of the reagents to perform the methods described herein. In certain embodiments, the kit comprises primers for detecting the nucleic acids specific to the panel of biomarkers in a sample.

[0144] “Primer” as used herein refers to an oligonucleotide molecule with a length of 7-40 nucleotides, preferably 10-38 nucleotides, preferably 15-30 nucleotides, or 15-25 nucleotides, or 17-20 nucleotides. For example, the primer can an oligonucleotide having a length of 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 nucleotides. Primers are used in the amplification of a DNA sequence by polymerase chain reaction (PCR) as well known in the art. For a DNA template sequence to be amplified, a pair of primers can be designed at its 5’ upstream and its 3’ downstream sequence, i.e., 5’ primer and 3’ primer, each of which can specifically hybridize to a separate strand of the DNA double strand template. 5’ primer is complementary to the anti-sense strand of the DNA double strand template; and 3’ primer is complementary to the sense strand of the DNA template. As known in the art, the “sense strand” of a double stranded DNA template is the strand which contains the sequence identical to the mRNA sequence transcribed from the DNA template (except that “U” in RNA corresponds to “T” in the DNA) and encoding for a protein product. The complementary sequence of the sense strand is the “anti-sense strand.”

[0145] In certain embodiments, the kit further comprises an agent for amplifying the target nucleic acid using the primers. In addition, the kits may include instructional materials containing directions (i.e., protocols) for the practice of the methods provided herein. While the instructional materials typically comprise written or printed materials, they are not limited to such. Any medium capable of storing such instructions and communicating them to an end user is contemplated by this disclosure. Such media include but are not limited to electronic storage media (e.g., magnetic discs, tapes, cartridges, chips), optical media (e.g., CD ROM), and the like. Such media may include addresses to internet sites that provide such instructional materials.

[0146] In another aspect, the present disclosure provides oligonucleotide probes for detecting the nucleic acids specific to the panel of biomarkers in a sample. In certain embodiments, the probes are attached to a solid support, such as an array slide or chip, e.g., as described in Eds., Bowtell and Sambrook DNA Microarrays: A Molecular Cloning Manual (2003) Cold Spring Harbor Laboratory Press. Construction of such devices are well known in the art, for example as described in US Patents and Patent Publications U.S. Patent No.5,837,832; PCT application W095 / 11995; U.S. Patent No. 5,807,522; US Patent Nos. 7,157,229, 7,083,975, 6,444,175, 6,375,903, 6,315,958, 6,295,153, and 5,143,854, 2007 / 0037274, 2007 / 0140906, 2004 / 0126757, 2004 / 0110212, 2004 / 0110211, 2003 / 0143550, 2003 / 0003032, and 2002 / 0041420. Nucleic acid arrays are also reviewed in the following references: Biotechnol Annu Rev (2002) 8:85-101; Sosnowski et al. Psychiatr Genet (2002)12(4): 181-92; Heller, Annu Rev Biomed Eng (2002) 4: 129-53; Koi chinsky et al., Hum. Mutat (2002) 19(4):343-60; and McGail et al., Adv Biochem Eng Biotechnol (2002) 77:21-42.

[0147] A microarray can be composed of a large number of unique, single-stranded polynucleotides, usually either synthetic antisense polynucleotides or fragments of cDNAs, fixed to a solid support. Typical polynucleotides are preferably about 6-60 nucleotides in length, more preferably about 15-30 nucleotides in length, and most preferably about 18-25 nucleotides in length. For certain types of arrays or other detection kits / sy stems, it may be preferable to use oligonucleotides that are only about 7-20 nucleotides in length. In other types of arrays, such as arrays used in conjunction with chemiluminescent detection technology, preferred probe lengths can be, for example, about 15-80 nucleotides in length, preferably about 50-70 nucleotides in length, more preferably about 55-65 nucleotides in length, and most preferably about 60 nucleotides in length.

[0148] Techniques for the synthesis of these arrays using mechanical synthesis methods are described in, e.g., U.S. Patent No. 5,384,261. Although a planar array surface is often employed the array may be fabricated on a surface of virtually any shape or even a multiplicity of surfaces. Arrays may also be nucleic acids on beads, gels, polymeric surfaces, fibers such as fiber optics, glass or any other appropriate substrate, see U.S. Patent Nos. 5,770,358, 5,789,162, 5,708,153, 6,040,193 and 5,800,992. Arrays may be packaged in such a manner as to allow for diagnostics or other manipulation of an all-inclusive device.

[0149] The probes and primers necessary for practicing the present disclosure can be synthesized and labeled using well known techniques. Oligonucleotides used as probes and primers may be chemically synthesized according to the solid phase phosphoramidite triester method first described by Beaucage and Caruthers, Tetrahedron Letts. (1981) 22: 1859-1862, using an automated synthesizer, as described in Needham-Van Devanter et al, Nucleic Acids Res. (1984) 12:6159-6168.

[0150] Methods for Treating Cancer

[0151] In yet another aspect, the present disclosure provides a method for treating colorectal cancer or precancerous colorectal adenoma in a subject. In some embodiments, themethod comprises administering to the subject a therapeutically effective amount of a drug useful for treating colorectal cancer or precancerous colorectal adenoma, wherein the subject has been determined to have colorectal cancer or precancerous colorectal adenoma, for example, by a machine learning classifier, based on levels (e.g., expression levels) of at least two biomarkers measured in a feces sample isolated from the subject, wherein the panel of biomarkers comprises biomarkers selected from the group disclosed herein.

[0152] The drug that can be used in the method disclosed herein include, without limitation: alkylating agents or agents with an alkylating action, such as cyclophosphamide (CTX; e.g. cytoxan®), chlorambucil (CHL; e.g. leukeran®), cisplatin (CisP; e.g. platinol®) busulfan (e.g. myleran®), melphalan, carmustine (BCNU), streptozotocin, triethylenemelamine (TEM), mitomycin C, and the like; anti-metabolites, such as methotrexate (MTX), etoposide (VP 16; e.g. vepesid®), 6-mercaptopurine (6MP), 6-thiocguanine (6TG), cytarabine (Ara-C), 5- fluorouracil (5-FU), capecitabine (e.g.Xeloda®), dacarbazine (DTIC), and the like; antibiotics, such as actinomycin D, doxorubicin (DXR; e.g. adriamycin®), daunorubicin (daunomycin), bleomycin, mithramycin and the like; alkaloids, such as vinca alkaloids such as vincristine (VCR), vinblastine, and the like; and other antitumor agents, such as paclitaxel (e.g. taxol®) and pactitaxel derivatives, the cytostatic agents, glucocorticoids such as dexamethasone (DEX; e.g. decadron®) and corticosteroids such as prednisone, nucleoside enzyme inhibitors such as hydroxyurea, amino acid depleting enzymes such as asparaginase, leucovorin, folinic acid, raltitrexed, and other folic acid derivatives, and similar, diverse antitumor agents. The following agents may also be used as additional agents: arnifostine (e.g. ethyol®), dactinomycin, mechlorethamine (nitrogen mustard), streptozocin, cyclophosphamide, lornustine (CCNU), doxorubicin lipo (e.g. doxil®), gemcitabine (e.g. gemzar®), daunorubicin lipo (e.g. daunoxome®), procarbazine, mitomycin, docetaxel (e.g. taxotere®), aldesleukin, carboplatin, oxaliplatin, cladribine, camptothecin, CPT 11 (irinotecan), 10-hydroxy 7-ethyl- camptothecin (SN38), floxuridine, fludarabine, ifosfamide, idarubicin, mesna, interferon alpha, interferon beta, mitoxantrone, topotecan, leuprolide, megestrol, melphalan, mercaptopurine, plicamycin, mitotane, pegaspargase, pentostatin, pipobroman, plicamycin, teniposide, testolactone, thioguanine, thiotepa, uracil mustard, vinorelbine, and chlorambucil.

[0153] In some embodiment, the drug used in the method disclosed herein include, without limitation: Alymsys® (Bevacizumab), Avastin® (Bevacizumab), Camptosar® (Irinotecan Hydrochloride), Capecitabine, Cetuximab, Cyramza® (Ramucirumab), Eloxatin® (Oxaliplatin), Erbitux® (Cetuximab), 5-FU (Fluorouracil Injection), Fluorouracil Injection,Ipilimumab, Irinotecan Hydrochloride, Keytruda® (Pembrolizumab), Leucovorin Calcium, Lonsurf® (Trifluridine and Tipiracil Hydrochloride), Mvasi® (Bevacizumab), Opdivo® (Nivolumab), Oxaliplatin, Panitumumab, Pembrolizumab, Ramucirumab, Regorafenib, Stivarga® (Regorafenib), Trifluridine and Tipiracil Hydrochloride, Vectibix® (Panitumumab), Xeloda® (Capecitabine), Yervoy® (Ipilimumab), Zaltrap® (Ziv-Aflibercept), Zirabev® (Bevacizumab), Ziv-Aflibercept.

[0154] The drug described herein may be administered in any desired and effective manner: for oral ingestion, or as an ointment or drop for local administration to the eyes, or for parenteral or other administration in any appropriate manner such as intraperitoneal, subcutaneous, topical, intradermal, inhalation, intrapulmonary, rectal, vaginal, sublingual, intramuscular, intravenous, intraarterial, intrathecal, or intralymphatic. Further, the drug may be administered in conjunction with other treatments.

[0155] The following examples are provided to better illustrate the claimed disclosure and are not to be interpreted as limiting the scope of the disclosure. All specific compositions, materials, and methods described below, in whole or in part, fall within the scope of the present disclosure. These specific compositions, materials, and methods are not intended to limit the disclosure, but merely to illustrate specific embodiments falling within the scope of the disclosure. One skilled in the art may develop equivalent compositions, materials, and methods without the exercise of inventive capacity and without departing from the scope of the disclosure. It will be understood that many variations can be made in the procedures herein described while still remaining within the bounds of the present disclosure. It is the intention of the inventors that such variations are included within the scope of the disclosure.EXAMPLE 1

[0156] This example shows the identification of biomarker panels for colorectal cancer or precancerous colorectal advanced adenomas.

[0157] It is well established that cells derived from tumors or precancerous lesions are shed into the stool and can be detected (Berger and Ahlquist 2012). Existing FDA approved methods to detect the presence of tumor or adenomas / advanced adenoma cells in stool commonly depend on detecting the presence of tumor derived DNA (“Multi target Stool DNA Testing for Colorectal-Cancer Screening | NEJM,” n.d.).

[0158] Detecting tumor derived RNA instead of DNA is a promising approach to improve detection of shed abnormal cells in stool (Herring et al. 2021). RNA will frequently have morecopies in the tumor cell than DNA which would increase sensitivity in detecting abnormal growths. In addition numerous genes have been shown to have very little to no expression in normal colon tissue but high expression in some tumors (Scanlan et al. 2002) suggesting mRNA biomarkers can also be highly specific. And finally stool derived mRNA can cheaply be measured using very sensitive and specific RT-qPCR assays.

[0159] The choice of which mRNA biomarkers to use will have a dramatic impact on colorectal cancer screening performance (Herring et al. 2021). Here we describe a methodology to identify mRNA biomarkers that can be used to screen for the presence of colorectal cancer (CRC) advanced adenomas (AA) and adenomas (A) in stool.

[0160] CRC Transcriptomic datasets were downloaded from The Cancer Genome Atlas (TCGA) database (Weinstein et al. 2013) and healthy colon controls were downloaded from the Genotype-Tissue Expression (GTEx) database. The combined TCGA / GTE dataset consisted of 478 colon cancer tissue samples, 166 rectal cancer tissue samples and 702 normal colon / rectum tissue samples.

[0161] Batch correction on the downloaded data was performed using combat-seq (“ComBat-Seq: Batch Effect Adjustment for RNA-Seq Count Data | NAR Genomics and Bioinformatics | Oxford Academic,” n.d.). A differential expression analysis comparing gene expression in CRC tissue vs healthy colon tissue was performed using edgeR (Robinson, McCarthy, and Smyth 2010). A variety of metrics was also calculated for each gene. This includes the percentile ranking of each gene’s median expression level in CRC and healthy tissue and each gene’s area under the curve (AUC) which quantifies how well the gene’s expression does in predicting if a sample is derived from cancerous or normal tissue. The closer the AUC is to 1 the better the gene predicts if the tissue sample is cancerous or normal. An AUC of 1 means the gene's expression level can perfectly predict if the sample is cancerous or normal with a sensitivity and specificity of 100%.

[0162] Stool is a complex matrix with microbial RNA being the dominant source of RNA in stool. It is expected that RNA derived from cancer cells will constitute a very small fraction of the total RNA present in stool. Therefore, additional filters were applied to identify and rank genes in terms of their ability to detect the presence of abnormal cells in stool.

[0163] We determined that genes with strong differential expression comparing CRC samples to normal samples and that are ubiquitously differentially expressed across most / all tumors compared to healthy controls would be good candidates for stool-based biomarkers for CRC / AA / A screening. The following filter was used to obtain these genes: FDR<0.001 andAUC > 0.9 and a log base 2-fold change in expression between CRC and healthy tissue > 2. The gene list was then sorted by median expression in CRC samples from highest to lowest (Table 1). All genes selected have an individual AUC of 0.9 or higher so perform very well as single biomarkers in classifying cancer from normal samples.EXAMPLE 2

[0164] This example illustrates the validation of the performance of classification models using at least two biomarkers identified in Example 1 to distinguish subjects having colorectal cancer or precancerous colorectal advanced adenomas from healthy subjects.

[0165] To test combinations of our selected biomarkers in correctly classifying CRC and normal samples. We trained classification models on the TCGA / GTEx data using all possible combinations of two genes and tested how well they performed on an independent test dataset. The test data was obtained from the gene expression omnibus database with accession number GSE165255 and consisted of 50 paired CRC / normal patient samples. This test dataset was not used as part of biomarker selection so represents an independent test of how well the biomarkers perform. We tested multiple classification models, including random forests and generalized linear model (GLM). Diagnostic performance in distinguishing between cancerous and normal tissue was evaluated by calculating the AUC for all combinations of two genes.

[0166] The performance on the held out test dataset was exceptionally good with a median AUC of 0.95 for both the random forest and GLM classification models. In addition, 90% of the two gene combinations had an AUC > 0.86 using the GLM classifier and 90% of the 2 gene combinations had an AUC > 0.81 for the random forest classifier (FIG. 1). All 2 gene combinations and their performance are listed in Table 2.

[0167] Larger combinations of genes must perform at least as well as the two gene combinations because any combination of 3 or more genes must comprise at least one of the 2 gene combinations and will not perform any worse than the best performing 2 gene combination that is included in the larger combinations. Therefore, we have also demonstrated that most larger combinations are also expected to perform well.

[0168] The set of mRNA biomarkers we have identified perform very well individually or in combination in distinguishing between abnormal and healthy tissue and represent a promising approach to building a low cost, high performance colorectal cancer screening assay.EXAMPLE 3

[0169] This example shows the clinical validation of the biomarkers obtained in EXAMPLE 1.

[0170] Stool was retrospectively collected from 131 patients: 49 CRC, 29 AA, 16 adenoma (AD), 19 hyperplastic polyps (HP), 15 normal colonoscopy. Advanced precancerous lesions were defined using generally accepted criteria i.e. greater than 1 cm in size in any dimension, or a villous component greater than 25% or the presence of high-grade dysplasia. Fifty-eight genes were selected from the set of 134 to test on the clinical samples. We have previously filed a patent application (Application Number PCT / US24 / 50917, whose content is incorporated herein via reference) that describes a novel method to use blood associated mRNA markers for the detection of occult blood in stool. We selected HBA to test in combination with the 58 CRC / normal tissue derived genes for a total of 59 genes tested on clinical samples. For HBA the TaqMan assay was designed to recognize either the HBA1 or HBA2 hemoglobin subunit and henceforth will be referred to as HBA. In addition, we measured the levels of GAPDH as a reference.

[0171] A total of 60 genes were measured on the clinical stool samples. The purpose of this case-control study was to test if the strong differential expression observed in the CRC / normal tissue datasets would translate to clinical utility. The primary outcome of this study was the AUC of CRC detection for each biomarker where the HP, AD and Normal colonoscopy samples were all combined into a single control group and compared to the CRC samples (Table 3). Secondary outcomes were the AUC of AA detection for each individual biomarker calculating the ability to distinguish stool samples of AA patients from stool samples of the control group (HP, AD and Normal colonoscopy samples) (Table 3), and the CRC and AA AUC for all 2 gene combinations (Table 4).

[0172] The methodology used to extract RNA from stool is as follows. RNA was extracted from 0.25 g of stool using the Omega Bio-Tek E.Z.N.A.® Stool RNA Kit (R6828), which employs a phenol-chloroform extraction method. To remove potential PCR inhibitors, the Zymo Research OneStep™ PCR Inhibitor Removal Kit (D6035) was used following the RNA extraction. The extracted RNA was resuspended in 50 pL of DEPC-treated water per tube. RNA concentration and purity were measured using a NanoDrop spectrophotometer. The typical RNA yield was approximately 450 pg per sample. The RNA was then aliquoted and stored at -80°C until further use.

[0173] After RNA was extracted from stool, gene transcripts were quantified with the RT- qPCR TaqMan assay. The PCR protocol used was as follows. Takara One Step PrimeScript™ III RT-qPCR Mix with UNG (RR601 A) was utilized for qRT-PCR. Each reaction was set up with a total volume of 10 pL, including 2 pL of extracted RNA. The reactions were performedon an ABI 7500 Real-Time PCR System with the following thermal cycling conditions: reverse transcription at 50°C for 5 minutes, initial denaturation at 95°C for 10 seconds, followed by 40 cycles of 95°C for 10 seconds and 60°C for 34 seconds. For each gene the primers / probes were designed in house to be RNA specific either by splitting the primer / probe over a splice site. Or by placing primers on exons separated by a large intron (> 1000 bps). Table 5 shows the sequence for the primers and probes used.

[0174] The delta-delta Ct method (Livak and Schmittgen 2001) was used to calculate the relative fold gene expression. The reference samples used in the delta-delta Ct calculation was the mean Ct value for those cancer samples with detectable signal defined as Ct values < 40 cycles. The reference gene used was GAPDH. For those stool samples with no detectable copies the fold gene expression was set to 0.

[0175] Gene performance in correctly predicting case stool (CRC, AA) versus control stool (HP, AD, Normal) was obtained using random forests. The machine learning experiments were performed using the tidymodels package in R, default parameters were used for the random forest model. No hyperparameter optimization or feature selection was performed. Classification performance was estimated using an average of five-fold cross validations repeated ten times with random splits of the data for each cross-validation experiment. Fivefold cross validation consists of randomly splitting the data into 5 roughly equal groups, training the model on the combination of four of the groups and testing the model on the held- out group. This process is repeated until all groups have been used as the held-out test set. The input data to the machine learning experiments was the normalized relative fold-expression for each of genes and gene combinations. Sensitivity and specificity were calculated using a 0.5 class probability cutoff. If the reported class probability was > 0.5 the sample was predicted to be a case sample if the class probability was < 0.5 the sample was predicted to be a control sample.

[0176] In general, the clinical utility of the 59 genes tested was very high. The large majority of the genes (95%) had an AUC > 0.6 predicting the stool obtained from CRC patients versus controls. Many of the genes also had a significant ability to detect the signature of AAs in stool, 48% of the genes had an AUC > 0.6 in predicting stool obtained from AA patients compared to the controls (Table 3). A significant portion of the genes had very strong individual performance; the top five genes PPBP, OLR1, CXCL8, MMP3, HBA all had CRC AUCs > 0.91 and the top ten genes all had individual AUCS > 0.85. The top performing gene PPBP at 92.6% specificity detected 88% of CRC samples and 45% of AA samples. Thiscompares favorably to FIT testing which is the most common method used for CRC screening worldwide (Allison et al. 2014). FIT has -70-80% sensitivity for CRC detection and 20-30% AA detection at -95% specificity (“Colonoscopy versus FIT-Fecal DNA for Colon Cancer Screening,” n.d.). Many of our genes had high specificity with 12 of the genes with CRC AUC > 0.6 having >=95% specificity (Table 3).

[0177] In general, combining genes significantly improved performance over using single genes only. The average CRC AUC using single gene biomarkers was 0.75 and 0.59 AA AUC compared to 0.85 average CRC AUC and 0.64 AA AUC using two gene combinations. The top performing combination for CRC AUC (PPBP, TIMP1) was also better than any individual gene with a CRC AUC of 0.985 compared to the top performing individual gene (PPBP), CRC AUC of 0.95. This pattern also holds true for AA performance. The top performing AA gene was MYC with an AA, AUC of 0.76 the top performing two gene combination was MYC combined with CBREF1 with an AA, AUC of 0.85 considerably better than MYC by itself (Table 4). More than 99% of the two gene combinations had > 0.6 CRC AUC with 72% of gene combinations having an AA AUC > 0.6. Larger gene combinations also performed very well. We combined the top twenty genes as sorted by individual AUC plus any gene with a specificity > 95%. A total of 25 genes were tested as a single panel (PPBP, OLR1, CXCL8, MMP3, HBA, TIMP1, MMP7, TGFBI, MYC, SPP1, NKD2, IL11, ETV4, CLEC5A, S100A9, TCN1, MMP1, IGF2, GDF15, TUBB3, EPHX4, CBX2, ASCL2, EVA1A, ULBP1). This larger gene combination had an CRC AUC of 0.987, an AA AUC of 0.74 with a sensitivity of CRC and AA detection respectively of 96% and 46% with a specificity of 86.8%.

[0178] Of particular interest was the combination of HBA with other mRNA biomarkers. Detecting the presence of blood associated proteins in stool using FIT is the basis of the most widespread and successful form of CRC screening worldwide. However, FIT has lower sensitivity for early-stage tumors and precancerous lesions. Pairing FIT with nucleic acid tumor derived biomarkers can improve sensitivity of detection for these types of lesions and is the basis of a number of FDA approved screening tests (“Multitarget Stool DNA Testing for Colorectal-Cancer Screening | NEJM,” n.d.). Since FIT is a protein based test, pairing it with nucleic acid based biomarkers increases complexity of sample collection and cost. We previously showed using blood associated mRNA biomarkers can replace FIT for detection of occult blood and combine well with other nucleic acid-based biomarkers. Here we test combining HBA, one of our blood associated biomarkers with a larger set of nucleic acid biomarkers. HBA quantifies the amount of both HBA1 and HBA2 which are two hemoglobinsubunit genes with a high degree of sequence similarity. Similar to our previous study, HBA combines very well with other nucleic acid biomarkers with an average CRC AUC of 0.94 across all 2 gene combinations compared to an individual CRC AUC of 0.91 for HBA by itself (Table 4) The ability to detect occult blood using nucleic acid-based biomarkers greatly reduces the cost and complexity of pairing the occult blood signal with other nucleic acid-based stool biomarkers.

[0179] We tested 58 out of the 134 biomarkers on clinical samples that we predicted to be informative for CRC / AA diagnosis. Similar to the results observed in tissue, the large majority of genes had differential representation in the stool of CRC or AA patients compared to controls and showed strong potential clinical utility with 95% of individual biomarkers having 0.6 CRC AUC or larger. It is expected that the large majority of the rest of the 134 genes we have not tested as part of our set of 58 will also have differential representation in the stool of CRC and or AA patients compared to controls.

[0180] In addition, we have shown the large majority of two gene combinations also have some ability to discriminate CRC or AA derived stool samples from controls with 99% of two gene combinations having an AUC of 0.6 or larger. Larger combinations of genes must perform at least as well as the two gene combinations. Since any combination 3 or above must consist of at least one of the 2 gene combinations and will not perform any worse than the best performing 2 gene combination that is included in the larger combinations. Therefore, we have also demonstrated that most larger combinations that include the 2 gene combinations we list here are also expected to have clinical utility.ReferencesAllison, James E., Callum G. Fraser, Stephen P. Halloran, and Graeme P. Young. 2014.“Population Screening for Colorectal Cancer Means Getting FIT : The Past, Present, and Future of Colorectal Cancer Screening Using the Fecal Immunochemical Test for Hemoglobin (FIT).” Gut and Liver 8 (2): 117-30.Bamell, Erica K., Yiming Kang, Elizabeth M. Wurtzler, Malachi Griffith, Aadel A.Chaudhuri, Obi L. Griffith, Andrew R. Barnell, et al. 2019. “Noninvasive Detection of High-Risk Adenomas Using Stool -Derived Eukaryotic RNA Sequences as Biomarkers.” Gastroenterology 157 (3): 884-887. e3.Berger, Barry M., and David A. Ahlquist. 2012. “Stool DNA Screening for Colorectal Neoplasia: Biological and Technical Basis for High Detection Rates.” Pathology 44(2): 80-88.Bosch, L. J. W., V. Melotte, S. Mongera, K. L. J. Daenen, V. M. H. Coupe, S. T. van Turenhout, E. M. Stoop, et al. 2019. “Multitarget Stool DNA Test Performance in an Average-Risk Colorectal Cancer Screening Population.” The American Journal of Gastroenterology 114 (12): 1909-18.Bray, Freddie, Mathieu Laversanne, Hyuna Sung, Jacques Ferlay, Rebecca L. Siegel, Isabelle Soerjomataram, and Ahmedin Jemal. 2024. “Global Cancer Statistics 2022: GLOBOCAN Estimates of Incidence and Mortality Worldwide for 36 Cancers in 185 Countries.” CA: A Cancer Journal for Clinicians, April.“Colonoscopy versus FIT -Fecal DNA for Colon Cancer Screening.” n.d. ACS. Accessed November 5, 2024.“ComBat-Seq: Batch Effect Adjustment for RNA-Seq Count Data | NAR Genomics and Bioinformatics | Oxford Academic.” n.d. Accessed October 17, 2023.Fearon, Eric R., and Bert Vogelstein. 1990. “A Genetic Model for Colorectal Tumorigenesis.” Cell 61 (5): 759-67.Heidenreich, Sebastian, Lila J. Finney Rutten, Lesley-Ann Miller-Wilson, Cecilia Jimenez- Moreno, Gin Nie Chua, and Deborah A. Fisher. 2022. “Colorectal Cancer Screening Preferences among Physicians and Individuals at Average Risk: A Discrete Choice Experiment.” Cancer Medicine 11 (16): 3156-67.Herring, Elizabeth, Eric Tremblay, Nathalie McFadden, Shigeru Kanaoka, and Jean-Frangois Beaulieu. 2021. “Multitarget Stool mRNA Test for Detecting Colorectal Cancer Lesions Including Advanced Adenomas.” Cancers 13 (6): 1228.Livak, Kenneth J., and Thomas D. Schmittgen. 2001. “Analysis of Relative Gene Expression Data Using Real-Time Quantitative PCR and the 2-AACT Method.” Methods 25 (4): 402-8.“Multitarget Stool DNA Testing for Colorectal-Cancer Screening | NEJM.” n.d. Accessed September 6, 2022.Navarro, Mercedes, Andrea Nicolas, Angel Ferrandez, and Angel Lanas. 2017. “Colorectal Cancer Population Screening Programs Worldwide in 2016: An Update.” World Journal of Gastroenterology 23 (20): 3632.Robinson, Mark D., Davis J. McCarthy, and Gordon K. Smyth. 2010. “edgeR: A Bioconductor Package for Differential Expression Analysis of Digital Gene Expression Data.” Bioinformatics 26 (1): 139-40.Scanlan, Matthew J., Ali O. Gure, Achim A. Jungbluth, Lloyd J. Old, and Yao-Tseng Chen. 2002. “Cancer / Testis Antigens: An Expanding Family of Targets for Cancer Immunotherapy.” Imm unological Reviews 188 (October):22-32.Weinstein, John N., Eric A. Collisson, Gordon B. Mills, Kenna M. Shaw, Brad A. Ozenberger, Kyle Ellrott, Ilya Shmulevich, Chris Sander, and Joshua M. Stuart. 2013. “The Cancer Genome Atlas Pan-Cancer Analysis Project.” Nature Genetics 45 (10): 1113-20.Zheng, Senshuang, Jelle J. A. Schrijvers, Marcel J. W. Greuter, Giirsah Kats-Ugurlu, Wenli Lu, and Geertruida H. de Bock. 2023. “Effectiveness of Colorectal Cancer (CRC) Screening on All-Cause and CRC-Specific Mortality Reduction: A Systematic Review and Meta- Analysis.” Cancers 15 (7): 1948.Table 1: CRC / AA Screening genes.AUC: Area under the curve; logFC = log 2 fold change of the gene expression comparing cancer to healthy samples; FDR: False discovery rate, the expected proportion of times a gene with the given FDR is not truly differentially expressed between cancer and healthy samples; N_percentile: The percentile ranking of the genes median expression level in normal samples; T_percentile: The percentile ranking of the genes median expression level in tumor samples. This table was generated using the TCGA + GTEx samples.Table 2: Classification Results of Two Genes CombinationIllTable 3: Performance of Genes on Clinical Stool Samples.AUC: Area under the curve; CRC AUC: the AUC observed comparing CRC samples to control samples (HP+AD+Normal Colonoscopy); AA AUC: the AUC observed comparing A samples to control samples (HP+AD+Normal Colonoscopy); CRC and AA sensitivity: the sensitivity of detection for CRC and AA stool samples. The specificity is the proportion of control samples correctly classified as neither CRC or AA samples. The top twenty genes and genes with 90% or greater specificity have an near their name.Table 4: Two Genes Combination Results on Clinical SamplesTable 5: Primers and Probes UsedFor Primer: forward primerRev Primer: reverse primerFluo: fluorescent dye used on the TaqMan probeQuencher: fluorescent quencher was used on the TaqMan probe

Claims

WHAT IS CLAIMED IS:

1. A method of diagnosing colorectal cancer or precancerous colorectal advanced adenomas in a subject, said method comprising: measuring in a biological sample obtained from the subject levels of a panel of biomarkers comprising one or more genes selected from Table 1, evaluating the measured levels of the panel of biomarkers, and determining that the subject is 1) healthy or 2) has precancerous colorectal advanced adenomas or colorectal cancer.

2. The method of claim 1, wherein the panel of biomarkers comprises one or more genes selected from the group consisting of: PPBP, 0LR1, CXCL8, MMP3, HBA, TIMP1, MMP7, TGFBI, MYC, SPP1, NKD2, IL11, ETV4, CLEC5A, S100A9, TCN1, MMP1, IGF2, GDF15, TUBB3, RPP40, EPHX4, CBX2, ASCL2, EVA1A, PRSS33, RAET1L, ULBP1, GAD1 and SALL4.

3. The method of claim 1, wherein the panel of biomarkers comprises a gene pair selected from the Table 2 or Table 4.

4. The method of claim 1 , wherein the method is for diagnosing colorectal cancer, wherein the panel of biomarkers comprises a gene pair selected from the group consisting of: (a) PPBP and TIMP1; (b) CXCL8 and PPBP; (c) PPBP and CBX2; (d) MMP7 and PPBP; (e) PPBP and MMP1; (f) HBA and MYC; (g) MYC and PPBP; (h) 0LR1 and PPBP; (i) MMP3 and SPP1; (j) PPBP and ETV4; (k) 0LR1 and TUBB3; (1) PPBP and SPP1; (m) CXCL8 and MMP3; (n) HBA and MMP3; (o) 0LR1 and MMP3; (p) PPBP and CEACAM6; (q) PPBP and S100A9; (r) 0LR1 and IFI6; (s) CXCL8 and TUBB3; and (t) PPBP and CLEC5A.

5. The method of claim 1, wherein the method is for diagnosing precancerous colorectal advanced adenomas, wherein the panel of biomarkers comprises a gene pair selected from the group consisting of:(a) MYC and CGREF1; (b) MYC and TCN1; (c) MYC and CEL; (d) HBA and MYC; (e) MYC and CPNE1; (f) MYC and TRIM29; (g) MYC and S100A9; (h) MYC and SPTBN2; (i) MYC and RHPN1; (j) HBA and CGREF1; (k) MYC and CEMIP; (1) HBA and TCN1 ; (m) MMP7 and CGREF 1 ; (n) MYC and UBE2C; (o) EPHX4 and MYC; (p) MYC and CLEC5A; (q) HBA and PPBP; (r) MYC and PPBP; (s) MYC and CXCL3; and (t) TIMP1 and TCN1.

6. The method of claim 1, wherein the biological sample is a stool sample.

7. The method of claim 1, wherein the levels of a panel of biomarkers are measured by detecting mRNA expression levels or expressed protein levels; optionally, wherein the mRNA expression levels are detected using quantitative RT-PCR, Droplet Digital PCR or nucleic acid sequencing.

8. The method of claim 1, wherein the levels of the panel of biomarkers are determined by quantitative RT-PCR or Droplet Digital PCR or nucleic acid sequencing.

9. The method of any one of claims 1-8, wherein the evaluating step and / or the determining step comprises analyzing the levels of the panel of biomarkers by a machine learning classifier; optionally, wherein the machine learning classifier is random forests or generalized linear model.

10. The method of any one of claims 1-9, wherein the evaluating step comprises comparing the levels of the panel of biomarkers to reference levels of the panel of biomarkers.

11. A method of diagnosing colorectal cancer or precancerous colorectal advanced adenomas in a subject, said method comprising measuring in a stool sample obtained from the subject mRNA levels of a panel of biomarkers comprising one or more genes selected from the group consisting of: PPBP, OLR1, CXCL8, MMP3, HBA, TEMPI, MMP7, TGFBI, MYC, SPP1, NKD2, IL11, ETV4, CLEC5A, S100A9, TCN1, MMP1, IGF2, GDF15, TUBB3, RPP40, EPHX4, CBX2, ASCL2, EVA1A, PRSS33, RAET1L, ULBP1, GAD1 and SALL4.

12. A kit or an integrated system of diagnosing colorectal cancer or precancerous colorectal advanced adenomas in a subject, comprising an agent for detecting in a biological sample obtained from the subject levels of the panel of biomarkers comprising one or more genes selected from Table 1, Table 2 or Table 4.

13. The kit or the integrated system of claim 12, wherein the agent is selected from the group consisting of: a primer, a nucleic acid and an oligonucleotide.

14. Use of a biomarker-specific reagent in the manufacture of a kit for diagnosing colorectal cancer or precancerous colorectal advanced adenomas in a subject, wherein the biomarker-specific reagent specifically binds to the panel of biomarkers comprising one or more genes selected from Table 1, Table 2 or Table 4.

15. The use of claim 14, wherein the subject is determined to have colorectal cancer or precancerous colorectal advanced adenomas by a machine learning classifier based on levels of the panel of biomarkers; optionally, wherein the machine learning classifier is random forests or generalized linear model.

16. The use of claim 15, wherein the levels of the panel of biomarkers are measured by detecting mRNA expression levels or expressed protein levels; optionally, wherein the mRNA expression levels are detected using quantitative RT-PCR, Droplet Digital PCR or nucleic acid sequencing.

17. The use of any one of claims 14-16, wherein the biomarker-specific reagent comprises a primer, a nucleic acid and / or an oligonucleotide.

18. A method for treating colorectal cancer or precancerous colorectal advanced adenomas in a subject, the method comprising: administering to the subject a therapeutically effective amount of a drug useful for treating colorectal cancer or precancerous colorectal advanced adenomas, wherein the subject has been determined to have colorectal cancer or precancerous colorectal advanced adenomas by evaluating levels of the panel of biomarkers comprising one or more genes selected from Table 1, Table 2 or Table 4.

19. The method of claim 18, wherein the levels of the panel of biomarkers are evaluated using a machine learning classifier such as random forests and / or generalized linear model.

20. The method of claim 18 or 19, wherein the levels of the panel of biomarkers are determined by quantitative RT-PCR, Droplet Digital PCR or nucleic acid sequencing.

Citation Information

Patent Citations

  • Novel genes, compositions, kits, and methods for identification, assessment, prevention, and therapy of colon cancer

    US20030148410A1

  • Method and apparatus for determining a probability of colorectal cancer in a subject

    US20150141286A1

  • Protein biomarker profiles for detecting colorectal tumors

    US20170176441A1

Cited By

  • Application of CGREF1 inhibitor in preparation of medicine for preventing and treating colorectal cancer metastasis

    CN120960432A