Occult blood mRNA biomarkers paired with DNA / RNA mutations and DNA methylation

The integration of blood RNA biomarkers with DNA/RNA mutational markers in stool samples simplifies and cost-reduces colorectal cancer screening by using a single collection tube, enhancing sensitivity and accuracy.

WO2025217467A1PCT designated stage Publication Date: 2025-10-16EL CAPITAN BIOSCIENCES INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/024177
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-10
Filing Date
2025-04-10
Publication Date
2025-10-16

AI Technical Summary

Technical Problem

Current fecal occult blood tests (FOBTs) for colorectal cancer screening are non-specific, costly, and complex, requiring multiple sample collection tubes and assays, which complicates the patient experience and increases costs.

Method used

A method using blood RNA biomarkers in stool samples, combined with DNA or RNA mutational markers, to diagnose colorectal cancer (CRC) or advanced precancerous lesions (APL), reducing the need for multiple collection tubes and assays, and utilizing PCR for sensitive detection.

Benefits of technology

Simplifies the screening process, reduces costs, and enhances sensitivity by using a single collection tube for both blood and genetic markers, enabling accurate detection of CRC or APL.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000030_0001
    Figure IMGF000030_0001
  • Figure IMGF000012_0001
    Figure IMGF000012_0001
  • Figure IMGF000013_0001
    Figure IMGF000013_0001
Patent Text Reader

Abstract

The present disclosure provides methods and compositions, e.g., kits, for diagnosing colorectal cancer based on a mRNA level of a blood-specific gene and a mutation or DNA methylation.
Need to check novelty before this filing date? Find Prior Art

Description

OCCULT BLOOD MRNA BIOMARKERS PAIRED WITH DNA / RNA MUTATIONS AND DNA METHYLATIONSEQUENCE LISTING

[0001] The sequence listing that is contained in the file named “081996-8006W001_Seq”, which is 3,583 bytes and was created on April 10, 2025, is filed herewith by electronic submission and is incorporated by reference herein.CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to US provisional application 63 / 632,018, filed April 10, 2024, the disclosure of which is incorporated herein by reference.FIELD OF THE INVENTION

[0003] The present disclosure generally relates to diagnosis, screening and treatment of gastrointestinal symptoms or conditions. In particular, the present disclosure relates to compositions and methods for detecting blood-specific mRNA in a feces sample for diagnosing and / or treating colorectal cancer.BACKGROUND

[0004] The fecal occult blood test (FOBT) is a diagnostic test to assess the presence of blood in stool. In healthy subjects, the volume of blood leaked into the gastrointestinal tract is 0.5 to 1.5 mL per day2,5. This small amount of blood is not usually detected in the fecal occult blood test10. Although FOBTs are commonly used as a diagnostic test in clinical settings, including for iron deficiency anemia, ulcerative colitis (UC) and acute diarrhea, their performance in predicting presumptive causes is poor9. For Colorectal Cancer (CRC), FOBT is only used for screening11.

[0005] FOBTs commonly used in clinical laboratories are based on detecting heme, heme- derived porphyrins, or colorectal hemoglobin protein14. Heme can be detected by guaiac-based methods (gFOBT, guaiac based fecal occult blood test), which utilizes the pseudo-peroxidase activity of hemoglobin, wherein guaiac is oxidized by hydrogen peroxidase. Because this reaction takes place with any peroxidase present in stool, gFOBT tests are non-specific to human Hb, with interference by any foodstuffs with peroxidase content, by certain chemicals or even medications. Heme can also be detected by converting the non-fluorescent heme to fluorescent porphyrins18. Alternatively, human hemoglobin or transferrin can be detected by immunological methods (iFOBT, human hemoglobin immunochemical based FOBT, or fecalimmunochemical test, FIT)17The only randomized clinical trials (RCT) comparing gFOBT and iFOBT published so far from van Rossum et al. concluded that the performance of iFOBTs is clearly superior to that of gFOBTs in detecting any type of colorectal neoplasia21. Moreover, iFOBTs does not necessitate any dietary modifications, improving patient adherence19. While detecting occult blood in stool via blood proteins is well established, there are circumstances in which it would be ideal if different means to detect the presence of blood in stool could be applied. In particular, the field of colorectal cancer screening would benefit from alternative approaches for detecting occult blood.

[0006] Sporadic CRC is thought to arise in part from accumulation of somatic gene mutations. The most commonly mutated genes are TP53, APC, and KRAS, which have been associated with survival and proliferation of cancer cells22. Hereditary CRC has two well- documented forms: familial adenomatous polyposis and Lynch syndrome. In both cases, the familial predisposition is due to germline variants. Thus, for either sporadic or hereditary CRC, the presence of genetic mutational burden is associated with clinical implications including disease screening and selection of treatment. Indeed, an existing FDA approved screening test as part of their biomarker panel includes mutations associated with CRC13.

[0007] The multitarget approach to colorectal cancer screening is now a part of daily practice in colon cancer screening with the FDA-approved FIT-fecal DNA test13. In addition, a multitarget FIT-RNA screening test is under review at the FDA6. Current multitarget approaches to screen for colorectal cancer and precancerous lesions combine FIT signal with other biomarker types, commonly nucleic acid-based biomarkers (DNA or RNA). Due to multiple different biomarker types being measured, combining FIT with other biomarkers typically requires at least two separate sample collection tubes and preservation buffers, plus at least two different assays - the FIT assay and a separate assay (commonly PCR) to measure the level of the other biomarkers. This complicates the patient experience and adds to the cost of the colorectal screening test. There is a continuing need to develop new diagnostic methods with good sensitivity for detecting a gastrointestinal disease and fecal occult blood, which is less complex and lower cost.SUMMARY OF INVENTION

[0008] We have previously described a new approach to quantify the level of blood in stool that uses blood RNA biomarkers instead of blood protein biomarkers. There are a number of advantages to RNA based methods of detecting occult blood. For nucleic acid based multitarget CRC screening tests only a single collection tube would be needed using our novel RNAapproach for occult blood detection, which would reduce the cost and complexity of the test. In addition, PCR can be used to quantify the level of the blood RNA biomarkers in stool. PCR is a highly sensitive and specific assay whose limit of detection can be as low as 3-5 copies8.

[0009] The present disclosure provides a method of diagnosing CRC or advanced precancerous lesion (APL) by detecting blood markers and DNA or RNA mutational markers in stool samples. The combination of our blood markers with DNA or RNA mutational detection in stool would remove the necessity of two collection tubes and reduce the cost of a multitarget CRC screening test that pairs occult blood detection with DNA / RNA mutation detection. In one embodiment, the method comprises: obtaining a stool sample of the subject; measuring a messenger RNA (mRNA) level of at least one blood-specific gene; detecting a mutation at the DNA level or an abnormal expression at RNA level of a CRC / APL-mutated gene, or a DNA methylation of a CRC / APL-methylated gene; evaluating the measured mRNA level of the blood-specific gene, the mutation or abnormal expression of the CRC / APL-mutated gene or the DNA methylation of the CRC -methylated gene; and determining whether the subject is 1) healthy or 2) has CRC or APL.

[0010] In some embodiments, the at least one blood-specific gene is selected from Table 1 (blood group antigen genes), Table 2 (platelet genes) or Table 3 (hemoglobin subunit genes). In some embodiments, the at least one CRC / APL-mutated gene is selected from Table 4. In some embodiments, the at least one CRC / APL-methylated gene is selected from Table 5.

[0011] In some embodiments, wherein the subject is a human.

[0012] In some embodiments, the mRNA level is detected using quantitative RT-PCR orDigital PCR. In some embodiments, wherein the mRNA level is determined by nucleic acid sequencing.

[0013] In some embodiments, the DNA methylation is detected by methylation-specific polymerase chain reaction (MSP), nucleic acid sequencing, mass spectrometry, methylationspecific nuclease, mass-based separation or targeted capture.

[0014] In some embodiments, the evaluating step and / or the determining step comprises using a machine learning classifier.

[0015] In another aspect, the present disclosure provides a kit or an integrated system of diagnosing CRC / APL in a subject. In some embodiments, the kit or integrated system comprises (a) a first agent for detecting in a stool sample obtained from the subject a mRNA level of a blood-specific gene and (b) a second agent for detecting in the stool sample a mutation at the DNA level or abnormal expression at RNA level of a CRC / APL-mutated gene or a DNAmethylation of a CRC / APL-methylated gene. In some embodiments, the first or second agent is selected from the group consisting of: primers, nucleic acids and oligonucleotides.

[0016] In another aspect, the present disclosure provides use of a biomarker-specific reagent in the manufacture of a kit for diagnosing CRC / APL in a subject, wherein the biomarker-specific reagent specifically binds to a mRNA of a blood-specific gene.

[0017] In yet another aspect, the present disclosure provides a method for treating CRC in a subject. In some embodiments, the method comprises: administering to the subject a therapeutically effective amount of a drug useful for treating CRC, wherein the subject has been determined to have CRC based on (a) a mRNA level of a blood-specific gene and (b) a mutation at DNA level or abnormal expression at RNA level of a CRC-mutated gene or a DNA methylation of a CRC -methylated gene.DETAILED DESCRIPTION OF THE INVENTION

[0018] Before the present disclosure is described in greater detail, it is to be understood that this disclosure is not limited to particular embodiments described, and as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present disclosure will be limited only by the appended claims.

[0019] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present disclosure, the preferred methods and materials are now described.

[0020] All publications and patents cited in this specification are herein incorporated by reference as if each individual publication or patent were specifically and individually indicated to be incorporated by reference and are incorporated herein by reference to disclose and describe the methods and / or materials in connection with which the publications are cited. The citation of any publication is for its disclosure prior to the filing date and should not be construed as an admission that the present disclosure is not entitled to antedate such publication by virtue of prior disclosure. Further, the dates of publication provided could be different from the actual publication dates that may need to be independently confirmed.

[0021] As will be apparent to those of skill in the art upon reading this disclosure, each of the individual embodiments described and illustrated herein has discrete components and features which may be readily separated from or combined with the features of any of the otherseveral embodiments without departing from the scope or spirit of the present disclosure. Any recited method can be carried out in the order of events recited or in any other order that is logically possible.

[0022] Definitions

[0023] The following definitions are provided to assist the reader. Unless otherwise defined, all terms of art, notations and other scientific or medical terms or terminology used herein are intended to have the meanings commonly understood by those of skill in the chemical and medical arts. In some cases, terms with commonly understood meanings are defined herein for clarity and / or for ready reference, and the inclusion of such definitions herein should not necessarily be construed to represent a substantial difference over the definition of the term as generally understood in the art.

[0024] As used herein, the singular forms “a”, “an” and “the” include plural references unless the context clearly dictates otherwise.

[0025] As used herein, the term “or” is an inclusive “or” operator and is equivalent to the term “and / or”, unless the context clearly dictates otherwise. The term “based on” is not exclusive and allowed for being based on additional factors not described unless the context clearly dictates otherwise. In addition, the singular forms “a,” “an” and “the” include plural references unless the content clearly dictates otherwise. The meaning of “in . . .” includes “within . . .” and “on . . .”.

[0026] As used herein, the term “administering” means providing a pharmaceutical agent or composition to a subject, and includes, but is not limited to, administering by a medical professional and self-administering.

[0027] The term “amount” or “level” generally refers to the quantity of a substance of interest. In the context of a biomarker or of a panel of biomarkers, a level of a panel of biomarkers refers to the quantity of the polynucleotides (e.g., mRNA) of interest or the polypeptides of interest present in a sample. Such quantity may be expressed in the absolute terms, i.e., the total quantity of the polynucleotides or polypeptides in the sample, or in the relative terms, i.e., the concentration of the polynucleotides or polypeptides in the sample.

[0028] The terms “assessing”, “assaying”, “measuring” and “detecting” can be used interchangeably and refer to both quantitative and semi-quantitative determinations. Where either a quantitative and semi-quantitative determination is intended, the phrase “measuring a level” of a polynucleotide or polypeptide of interest or “detecting” a polynucleotide or polypeptide of interest can be used.

[0029] As used herein, the term “biomarker” refers to a detectable organic biomolecule associated with a particular phenotype or risk of developing a particular phenotype, such as a polynucleotide (e.g., RNA (e.g., mRNA) or DNA (e.g., cDNA)) or a polypeptide, which is differentially present in a biological sample taken from a subject having a certain condition (e.g., having a gastrointestinal disease) as compared to a comparable biological sample taken from a subject who does not have such condition, such as a healthy subject or a non-cancer patient. For example, a biomarker can be a polynucleotide, such as RNA (e.g., mRNA), which is present at an elevated or decreased level in a biological sample (e.g., a tissue sample, a feces sample, or a blood sample) of a gastrointestinal disease patient compared to a comparable sample (e.g., a feces sample) of a subject with a negative diagnosis (e.g., a healthy subject).

[0030] The term “colon cancer” used interchangeably with the term “colorectal cancer” or “rectal cancer” refers to any cancerous neoplasia of the colon (including the rectum and appendix). Many colorectal cancers arise from precancerous colorectal adenomas or adenomatous polyps, which are usually benign, but some may develop into cancer over time. The diagnosis of localized colon cancer is often through colonoscopy. Once localized colon cancer is diagnosed, it is usually surgically removed and then may be treated with chemotherapy. Precancerous colorectal adenomas or colorectal adenomatous polyps are a risk factor for colorectal cancer. The removal of colorectal adenomatous polyps at the time of colonoscopy would reduce the risk of having colorectal cancer. In addition, clinical data has shown that early detection and curative surgical resection of colorectal cancer will significantly improve survival rates.

[0031] As used herein, advanced precancerous lesion (APL) refers cells or tissues that are not yet cancerous, but show significant abnormal changes (often in structure, size, or organization) that strongly suggest a likelihood of becoming malignant over time. Typically, colorectal APL refers to advanced adenomas as defined by size >1 cm, with villous or tubulovillous histology and high-grade dysplasia.

[0032] It is noted that in this disclosure, terms such as “comprises”, “comprised”, “comprising”, “contains”, “containing” and the like have the meaning attributed in United States Patent law; they are inclusive or open-ended and do not exclude additional, un-recited elements or method steps. Terms such as “consisting essentially of’ and “consists essentially of’ have the meaning attributed in United States Patent law; they allow for the inclusion of additional ingredients or steps that do not materially affect the basic and novel characteristics of the claimed disclosure. The terms “consists of’ and “consisting of’ have the meaning ascribed to them in United States Patent law; namely that these terms are close ended.

[0033] As used herein, the term “complementary” and “complementarity” refer to nucleotides (e.g., a nucleotide) or a polynucleotide (e.g., a sequence of nucleotides) associated with base pairing rule. For Example, sequence 5’-A-G-T-3’ is complementary to sequence 3’- T-C-A-5’. Complementary can be “partial,” in which only some nucleic acid bases are matched according to the base pairing rule. Alternatively, there may be “completely” or “total” complementarity between nucleic acids. The complementary degree between nucleic acid chains affects the efficiency and strength of hybridization between nucleic acid chains. This is especially important in the amplification reactions and detection methods that depend upon binding between nucleic acids.

[0034] As used herein, the term “nucleic acid” and “polynucleotide” are used interchangeably and refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof. Polynucleotides may have any three-dimensional structure, and may perform any function, known or unknown. Non-limiting examples of polynucleotides include a gene, a gene fragment, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozymes, cDNA, shRNA, single-stranded short or long RNAs, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, control regions, isolated RNA of any sequence, nucleic acid probes, and primers. The nucleic acid molecule may be linear or circular.

[0035] As used herein, the term “polymerase chain reaction” or “PCR” is used to refer to a technique for amplifying a target sequence and is well recognized by one skilled in the art. The PCR generally consists of introducing a large excess of two oligonucleotide primers into the DNA mixture containing the desired target sequence, followed by a precise sequence of thermal cycling in the presence of a DNA polymerase. The two primers are complementary to their respective strands of the double stranded target sequence. For amplification, the mixture is denatured and then the primers annealed with its complementary sequence within the target molecule. After annealing, the primers were extended with a polymerase so as to form a new pair of complementary chains. The steps of denaturation, primer annealing, and polymerase extension can be repeated many times (that is, denaturation, annealing, and extension constitute one “cycle” and there can be numerous “cycles”) to obtain a high concentration of amplified fragments of the desired target sequence. The length of the amplified fragment of the target sequence is determined by the relative position of the primers with respect to each other, so the length is a controllable parameter. Since the desired amplified fragment of the target sequence becomes the predominant sequence (in terms of concentration) in the mixture, it is called “PCR amplified”, “PCR products” or “amplicons”.

[0036] The term “nucleic acid detection assay” refers to any method of determining the nucleotide composition of the target nucleic acid. Nucleic acid detection assay includes but is not limited to DNA sequencing methods and probe hybridization methods.

[0037] The term “amplifiable nucleic acid” refers to a nucleic acid that can be amplified by any amplification method. It is expected that “amplifiable nucleic acid” will normally comprise a “sample template”.

[0038] The term “sample template” refers to the nucleic acid originating from a sample that is for analysis of presence of the “target” (as defined below). In contrast, “background template” is used to refer to nucleic acids other than the sample template, which may be or may not be present in the sample. Background template is most often inadvertent, which may be the result of carryover, or due to the presence of nucleic acid contaminants sought to be purified from the sample. For Example, nucleic acids from an organism other than those to be tested may exist as a background for the test sample.

[0039] The term “primer” refers to an oligonucleotide, whether occurring naturally as in a purified restriction digestion or is produced synthetically, that is capable of acting as a point of initiation of synthesis when placed in the conditions for induction of synthesis of product extended from primers complementary to nucleic acid chain (e.g., in the presence of nucleotides and an inducing agent such as a DNA polymerase and under proper temperature and pH). The primer is typically single-stranded for maximum efficiency in amplification, but may alternatively be double-stranded. In the case of double chains, the primer is first treated to separate its strands before being used to prepare the extension product. In general, the primer is oligodeoxyribonucleotides. At minimum, the length of primer should be sufficient to initiate the synthesis of the extension product in the presence of the inducer. The exact length of the primer will depend on many factors, including temperature, source of the primer, and the use of the method.

[0040] The term “probe” refers to an oligonucleotide (e.g., a sequence of nucleotides) whether occurring naturally as in a purified restriction digest or produced synthetically, recombinantly, or by PCR amplification, that is capable of selectively hybridize to another sensitive target oligonucleotide. A probe can be single — stranded or double — stranded. Probes are useful in the detection, identification, and isolation of specific gene sequences (e.g., “capture probes”). It is anticipated that in some embodiments, any probe used in the present invention may be labeled with any “report molecule” to make it detectable in any detection system.

[0041] The term “methylation” refers to cytosine methylation at position C5 or N4 of cytosines, the N6 position of adenine, or other types of nucleic acid methylation. In vitro DNA amplified oligonucleotides are usually unmethylated because typical in vitro DNA amplification methods do not retain the methylation pattern of the amplified template. However, “unmethylated DNA” or “methylated DNA” can also refer to amplified DNA whose original template was unmethylated or methylated, respectively.

[0042] The term “methylated nucleotides” or “methylated nucleotide bases” refer to the presence of a methyl moiety on a nucleotide base, where the methyl moiety is not present in a recognized typical nucleotide bases. For Example, cytosine does not contain a methyl moiety in its pyrimidine ring, but 5-methyl cytosine contains a methyl moiety at 5-position of its pyrimidine ring. Therefore, cytosine is not a methylated nucleotide and 5-methylcytosine is a methylated nucleotide. In another example, the thymine contains a methyl moiety at 5-position of its pyrimidine ring; however, for the purposes herein, thymine is not considered a methylated nucleotide when present in DNA since thymine is a typical nucleotide base of DNA.

[0043] The methylation status can be optionally represented or indicated by a “methylation value” (e.g., representing methylation frequency, fraction, ration, percent, etc.). A methylation values can be generated, for example, by quantifying the amount of intact nucleic acid present following restriction digestion with a methylation dependent restriction enzyme, by comparing amplification profiles after the bisulfite reaction, or by comparing the sequences of bisulfite- treated and untreated nucleic acids. Thus, a value such as methylation value, represents methylation status and can be used as a quantitative indicator of methylation status across multiple copies of a locus.

[0044] As used herein, the term “bisulfite reagent” refers to a reagent comprising bisulphite, disulphite, hydrogen sulfite or a combination thereof. Cytosine nucleotides without methylation in the DNA treated with a bisulfite reagent will convert or translate into uracil, methylated cytosine while other bases remain unchanged. This allows differentiation between the methylated and unmethylated cytidine such as those in two nucleotide sequences of CpG.

[0045] The term “methylation assay” refers to any assay used to determine the methylation status of one or more CpG dinucleotide sequences within a nucleic acid sequence.

[0046] As used herein, the term “subject” refers to a human or any non-human animal (e.g., mouse, rat, rabbit, dog, cat, cattle, swine, sheep, horse or primate). A human includes pre and post-natal forms. In many embodiments, a subject is a human being. A subject can be a patient, which refers to a human presenting to a medical provider for diagnosis or treatment of a disease. The term “subject” is used herein interchangeably with “individual” or “patient”. A subject canbe afflicted with or is susceptible to a disease or disorder but may or may not display symptoms of the disease or disorder.

[0047] As used herein, the term “therapeutically effective amount” means the amount of agent that is sufficient to prevent, treat, reduce and / or ameliorate the symptoms and / or underlying causes of any disorder or disease, or the amount of an agent sufficient to produce a desired effect on a cell. In one embodiment, a “therapeutically effective amount” is an amount sufficient to reduce or eliminate a symptom of a disease. In another embodiment, a therapeutically effective amount is an amount sufficient to overcome the disease itself.

[0048] The term “treatment,” “treat,” or “treating” refers to a method of reducing the effects of a cancer (e.g., breast cancer, lung cancer, ovarian cancer or the like) or symptom of cancer. Thus, in the disclosed method, treatment can refer to a 10%, 20%, 30%, 40%, 50%, 60%, 70%), 80%), 90%), or 100% reduction in the severity of a cancer or symptom of the cancer. For example, a method of treating a disease is considered to be a treatment if there is a 10% reduction in one or more symptoms of the disease in a subject as compared to a control. Thus, the reduction can be a 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100% or any percent reduction between 10 and 100% as compared to native or control levels. It is understood that treatment does not necessarily refer to a cure or complete ablation of the disease, condition, or symptoms of the disease or condition.

[0049] Panel of Biomarkers

[0050] Fecal occult blood test is a test that checks for occult (hidden) blood in the stool. Blood in the stool may be a sign of colorectal cancer or other gastrointestinal conditions, such as polyps, ulcers, or hemorrhoids. Guaiac fecal occult blood test (gFOBT) and immunochemical fecal occult blood test (iFOBT) are two types of fecal occult blood tests. Guaiac fecal occult blood test uses a chemical substance called guaiac to check for blood in the stool. Immunochemical fecal occult blood test uses an antibody to check for blood in the stool.

[0051] The present disclosure in one aspect provides a panel of biomarkers that includes at least one blood associated mRNA biomarker. In particular, the disclosed blood associated biomarker is useful for combining with a CRC / APL-mutation biomarker or a CRC / APL-DNA- methylation biomarker for diagnosing CRC / APL. The methods utilizing the biomarker panel described herein would require one nucleic acid test assay instead of two different types of assays, including one protein based test, in order to diagnose the gastrointestinal disease, which simplifies the patient experience, reduces to chance of lab errors when combining the results of different assays, and reduces the cost of stool collection and assay reagents. The methods and composition to detect fecal occult blood is of high particular value in mRNA-basedmethods of colorectal cancer screening. The methods provided herein can be used in an outpatient clinic or inpatient environment.

[0052] As used herein, the term “panel of biomarkers” refers to a selection of at least two biomarkers. The panel can comprise from 2 to 20 (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20) or more biomarkers. In certain embodiments, the panel of biomarkers disclosed herein comprises at least one blood-specific mRNA biomarker and a CRC-specific biomarker.

[0053] In certain embodiments, the blood-specific biomarkers used in the method disclosed herein are listed in Table 1 (blood group antigen genes), Table 2 (platelet genes) and Table 3 (hemoglobin subunit genes). The term “blood-specific biomarker” as used herein refers to a gene that is specifically expressed in at least one type of blood cells. It is understood that some blood-specific biomarkers disclosed herein may not express exclusively in blood cells, but rather expressed in a highly enriched level in blood cells. In certain embodiments, the expression of a group of blood-specific biomarkers are used to detect fecal occult blood. In such embodiments, each of the biomarkers used in detecting fecal occult blood is a bloodspecific biomarker.

[0054] Table 1: Blood-specific Biomarkers from Blood Group Antigen Genes

[0055] Table 2: Blood-specific Biomarkers from Platelet Genes

[0056] Table 3: Hemoglobin Subunit Genes

[0057] In certain embodiments, the CRC-mutated biomarkers used in the method disclosed herein are genes commonly mutated in colorectal cancer or APL and could be paired with our fecal occult blood detection biomarkers that were obtained from the cosmic database4and are listed in Table 4.

[0058] Table 4: Genes commonly mutated in CRC / APL

[0059] In addition to somatic mutations, epigenetic alterations are thought to contribute substantially to the initiation and progression of colorectal cancer1, 22. DNA methylation is present in early stage and precancerous lesions and hence represent promising biomarkers tobuild methods of early-stage detection of CRC22. Indeed, assessing the presence of tumor associated methylated DNA is a well-established means to screen for the presence of colorectal cancer11. Pairing protein based FIT assay for occult blood detection with DNA methylation biomarkers is the basis of a number of published multitarget DNA screening methods17.

[0060] Compared with the protein based FIT assay, the blood biomarkers for occult blood detection paired with DNA methylation biomarkers would result in considerable cost and complexity reduction in the sample collection process. In certain embodiments, the DNA methylation biomarkers used in the method disclosed herein are listed in Table 5.

[0061] Table 5: Genes commonly methylated in CRC / APL

[0062] In addition to pairing well with other nucleic acid biomarkers, the blood biomarkers used in the method disclosed herein can be used singly or combined to boost signal and increase detection sensitivity. For example, HBA1 and HBA2 code for two different subunits of hemoglobin and are in the top ten highest expressed genes in whole blood. HBA1 and HBA2 have identical protein sequences and are 97% identical in DNA sequence (Turner et al., 2015). Given the strong sequence similarity between the two above-mentioned genes, a single PCR assay with primers that recognize both HBA1 and HBA2 can be designed. The single PCR assay is able to combine the signal from two of the highest expressing genes in blood in a single biomarker, and therefore boost the detection of the presence of trace amounts of bleeding in stool. Biologically, this approach is sound as functional hemoglobin requires equal protein expression of both protein subunits.

[0063] Measuring Biomarkers

[0064] The methods of the present disclosure involve detecting or measuring at least one blood-specific biomarker (for example, in Table 1-3) and at least one CRC-specific biomarkers disclosed herein (for example, in Table 4-5), in a stool sample obtained from a subject suspected of having or at risk of having CRC.

[0065] Sample Preparation

[0066] The methods disclosed herein use a stool sample (i.e., a feces sample), which may contain occult blood cells and tumor cells or debris thereof (e.g., cancer cell apoptotic products).

[0067] In certain embodiments, the method comprises a step of isolating nucleic acid, e.g., ribonucleic acid (RNA) from the sample. Various methods of extraction are suitable for isolating nucleic acid from cells or tissues, such as phenol and chloroform extraction, and various other methods as described in, for example, Ausubel et al., Current Protocols of Molecular Biology ( 1997) John Wiley & Sons, and Sambrook and Russell, Molecular Cloning: A Laboratory Manual 3rded. (2001). In some embodiments, nucleic acid isolation kit is used to obtain the nucleic acid for analysis. Other suitable methods for obtaining nucleic acid from a subject for testing is well known to one skilled in the art.

[0068] In certain embodiments, the method comprises a step of isolating deoxyribonucleic acid (DNA) from the sample. Various methods of extraction are suitable for isolating the DNA from cells or tissues, such as phenol and chloroform extraction, and various other methods as described in, for example, Machiels et al., Biotechniques, 28(2): 286-290 (2000). In some embodiments, a DNA extraction kit is used to obtain the DNA for analysis. Other suitable methods for obtaining genomic DNA and / or free DNA from a subject for testing for the presence of gastrointestinal diseases is well known to one skilled in the art.

[0069] Commercially available kits can also be used to isolate RNA, including for example, the NucleoSpin RNA Stool kit (Takara), Rneasy® mini columns (Qiagen), Stool Total RNA Purification kit (Norgen Biotek), NZY Stool RNA Isolation kit (NZYTech), and PureLink® RNA mini kit (Thermo Fisher Scientific). A skilled person can readily extract or isolate RNA or DNA following the manufacturer’s protocol.

[0070] Methods of Measuring Biomarkers

[0071] The biomarkers disclosed herein can be detected in the level of RNA (e.g., mRNA), mutation of DNA and methylation level of DNA using proper methods known in the art including, without limitation, amplification assay, hybridization assay, and sequencing assay. In certain embodiments, the mRNA expression levels are detected using quantitative RT-PCR or Digital PCR. In certain embodiments, the mRNA levels of the panel of biomarkers are determined by nucleic acid sequencing. In certain embodiments, the DNA mutations aredetermined by nucleic acid sequencing. In certain embodiments, the DNA methylation levels are determined by methylation specific PCR (MSP), DNA methylation-based chip, targeted DNA methylation sequencing, digital PCR and quantitative fluorescence PCR.

[0072] Sequencing methods

[0073] Sequencing methods useful in the measurement of the biomarkers involve sequencing of the target nucleic acid. Any sequencing known in the art can be used to detect the biomarkers of interest. In general, sequencing methods can be categorized to traditional or classical methods and high throughput sequencing (next generation sequencing). Traditional sequencing methods include Maxam-Gilbert sequencing (also known as chemical sequencing) and Sanger sequencing (also known as chain-termination methods).

[0074] High throughput sequencing, or next generation sequencing, by using methods distinguished from traditional methods, such as Sanger sequencing, is highly scalable and able to sequence the entire genome or transcriptome at once. High throughput sequencing involves sequencing-by-synthesis, sequencing-by-ligation, and ultra-deep sequencing (such as described in Marguiles et al., Nature 437 (7057): 376-80 (2005)). Sequence-by-synthesis involves synthesizing a complementary strand of the target nucleic acid by incorporating labeled nucleotide or nucleotide analog in a polymerase amplification. Immediately after or upon successful incorporation of a label nucleotide, a signal of the label is measured, and the identity of the nucleotide is recorded. The detectable label on the incorporated nucleotide is removed before the incorporation, detection and identification steps are repeated. Examples of sequence-by-synthesis methods are known in the art, and are described for example in U.S. Pat. No. 7,056,676, U.S. Pat. No. 8,802,368 and U.S. Pat. No. 7,169,560, the contents of which are incorporated herein by reference. Sequencing-by-synthesis may be performed on a solid surface (or a microarray or a chip) using fold-back PCR and anchored primers. Target nucleic acid fragments can be attached to the solid surface by hybridizing to the anchored primers, and bridge amplified. This technology is used, for example, in the Illumina® sequencing platform.

[0075] Pyrosequencing involves hybridizing the target nucleic acid regions to a primer and extending the new strand by sequentially incorporating deoxynucleotide triphosphates corresponding to the bases A, C, G, and T (U) in the presence of a polymerase. Each base incorporation is accompanied by release of pyrophosphate, converted to ATP by sulfurylase, which drives synthesis of oxyluciferin and the release of visible light. Since pyrophosphate release is equimolar with the number of incorporated bases, the light given off is proportional to the number of nucleotides adding in any one step. The process is repeated until the entire sequence is determined.

[0076] In certain embodiments, the biomarkers described herein are detected by whole transcriptome shotgun sequencing (RNA sequencing). The method of RNA sequencing has been described (see Wang Z, Gerstein M and Snyder M, Nature Review Genetics (2009) 10:57- 63; Maher CA et al., Nature (2009) 458:97-101; Kukurba K & Montgomery SB, Cold Spring Harbor Protocols (2015) 2015(11): 951-969).

[0077] Amplification assay

[0078] A nucleic acid amplification assay involves copying a target nucleic acid (e.g., DNA or RNA), thereby increasing the number of copies of the amplified nucleic acid sequence. Amplification may be exponential or linear. Exemplary nucleic acid amplification methods include, but are not limited to, amplification using the polymerase chain reaction ("PCR", see U.S. Patents 4,683,195 and 4,683,202; PCR Protocols: A Guide To Methods And Applications (Innis et al., eds, 1990)), reverse transcriptase polymerase chain reaction (RT-PCR), quantitative real-time PCR (qRT-PCR); quantitative PCR, such as TaqMan®, nested PCR, Digital PCR, such as Droplet Digital PCR (ddPCR), ligase chain reaction (See Abravaya, K., et al., Nucleic Acids Research, 23:675-682, (1995), branched DNA signal amplification (see, Urdea, M. S., et al., AIDS, 7 (suppl 2):S11-S14, (1993), amplifiable RNA reporters, Q-beta replication (see Lizardi et al., Biotechnology (1988) 6: 1197), transcription-based amplification (see, Kwoh et al., Proc. Natl. Acad. Sci. USA (1989) 86: 1173-1177), boomerang DNA amplification, strand displacement activation, cycling probe technology, self-sustained sequence replication (Guatelli et al., Proc. Natl. Acad. Sci. USA (1990) 87: 1874-1878), rolling circle replication (U.S. Patent No. 5,854,033), isothermal nucleic acid sequence based amplification (NASBA), and serial analysis of gene expression (SAGE). Droplet Digital PCR (ddPCR) is a method for performing digital PCR that is based on water-oil emulsion droplet technology. A sample is fractionated into many droplets, and PCR amplification of the template molecules occurs in each individual droplet. ddPCR technology uses reagents and workflows similar to those used for most standard TaqMan probe-based assays. The massive sample partitioning is a key aspect of the ddPCR technique.

[0079] In certain embodiments, the nucleic acid amplification assay is a PCR-based method. PCR is initiated with a pair of primers that hybridize to the target nucleic acid sequence to be amplified, followed by elongation of the primer by polymerase which synthesizes the new strand using the target nucleic acid sequence as a template and dNTPs as building blocks. Then the new strand and the target strand are denatured to allow primers to bind for the next cycle of extension and synthesis. After multiple amplification cycles, the total number of copies of the target nucleic acid sequence can increase exponentially.

[0080] In certain embodiments, intercalating agents that produce a signal when intercalated in double stranded DNA may be used. Exemplary agents include SYBR GREEN™ and SYBR GOLD™. Since these agents are not template-specific, it is assumed that the signal is generated based on template-specific amplification. This can be confirmed by monitoring signal as a function of temperature because the melting point of template sequences will generally be much higher than, for example, primer-dimers, etc.

[0081] In certain embodiments, a detectably labeled primer or a detectably labeled probe can be used, to allow detection of the biomarkers corresponding to that primer or probe. In certain embodiments, multiple labeled primers or labeled probes with different detectable labels can be used to allow simultaneous detection of multiple biomarkers.

[0082] Hybridization assay

[0083] Nucleic acid hybridization assays use probes to hybridize to the target nucleic acid, thereby allowing detection of the target nucleic acid. Non-limiting examples of hybridization assay include Northern blotting, Southern blotting, in situ hybridization, microarray analysis, and multiplexed hybridization-based assays.

[0084] In certain embodiments, the probes for hybridization assay are detectably labeled. In certain embodiments, the nucleic acid-based probes for hybridization assay are unlabeled. Such unlabeled probes can be immobilized on a solid support such as a microarray and can hybridize to the target nucleic acid molecules which are detectably labeled.

[0085] In certain embodiments, hybridization assays can be performed by isolating the nucleic acids (e.g., RNA or DNA), separating the nucleic acids (e.g., by gel electrophoresis) followed by transfer of the separated nucleic acid on suitable membrane filters (e.g., nitrocellulose filters), where the probes hybridize to the target nucleic acids and allows detection. See, for example, Molecular Cloning: A Laboratory Manual, J. Sambrook et al., eds., 2nd edition, Cold Spring Harbor Laboratory Press, 1989, Chapter 7. The hybridization of the probe and the target nucleic acid can be detected or measured by methods known in the art. For example, autoradiographic detection of hybridization can be performed by exposing hybridized filters to photographic film.

[0086] In some embodiments, hybridization assays can be performed on microarrays. Microarrays provide a method for the simultaneous measurement of the levels of large numbers of target nucleic acid molecules. The target nucleic acids can be RNA, DNA, cDNA reverse transcribed from mRNA, or chromosomal DNA. The target nucleic acids can be allowed to hybridize to a microarray comprising a substrate having multiple immobilized nucleic acid probes arrayed at a density of up to several million probes per square centimeter of the substratesurface. The RNA or DNA in the sample is hybridized to complementary probes on the array and then detected by laser scanning. Hybridization intensities for each probe on the array are determined and converted to a quantitative value representing relative levels of the RNA or DNA. See, U.S. Patent Nos. 6,040,138, 5,800,992 and 6,020,135, 6,033,860, and 6,344,316.

[0087] Techniques for the synthesis of these arrays using mechanical synthesis methods are described in, e.g., U.S. Patent No. 5,384,261. Although a planar array surface is often employed the array may be fabricated on a surface of virtually any shape or even a multiplicity of surfaces. Arrays may be peptides or nucleic acids on beads, gels, polymeric surfaces, fibers such as fiber optics, glass or any other appropriate substrate, see U.S. Patent Nos. 5,770,358, 5,789,162, 5,708,153, 6,040,193 and 5,800,992. Arrays may be packaged in such a manner as to allow for diagnostics or other manipulation of an all-inclusive device. Useful microarrays are also commercially available, for example, microarrays from Affymetrix, from Nano String Technologies, QuantiGene 2.0 Multiplex Assay from Panomics.

[0088] In certain embodiments, hybridization assays can be in situ hybridization assay. In situ hybridization assay is useful to detect the presence of gene mutations. Probes useful for in situ hybridization assay can be mutation specific probes, which hybridize to a specific gene mutation to detect the presence or absence of the specific mutation of interest. Methods for use of unique sequence probes for in situ hybridization are described in U.S. Pat. No. 5,447,841, incorporated herein by reference. Probes can be viewed with a fluorescence microscope and an appropriate filter for each fluorophore, or by using dual or triple band-pass filter sets to observe multiple fhiorophores. See, e.g., U.S. Pat. No. 5,776,688 to Bittner, et al., which is incorporated herein by reference. Any suitable microscopic imaging method can be used to visualize the hybridized probes, including automated digital imaging systems. Alternatively, techniques such as flow cytometry can be used to examine the hybridization pattern of the probes.

[0089] Any of the assays and methods provided herein for the measurement of the gene expression level can be adapted or optimized for use in automated and semi-automated systems or point of care assay systems.

[0090] The gene expression level described herein can be normalized using a proper method known in the art. For example, the gene expression level can be normalized to a standard level of a standard marker, which can be predetermined, determined concurrently, or determined after a sample is obtained from the subject. The standard marker can be run in the same assay or can be a known standard marker from a previous assay. For another example,the gene expression level can be normalized to an internal control which can be an internal marker, or an average level or a total level of a plurality of internal markers.

[0091] The level of mRNA expression of each of the biomarkers described herein can be normalized to a reference level for a control gene. The control value can be predetermined, determined concurrently, or determined after a sample is obtained from the subject. The standard can be run in the same assay or can be a known standard from a previous assay. In the cases when the level of RNA expression is determined by RNA sequencing, the level of RNA expression of each of the biomarkers can be normalized to the total reads of the sequencing. The normalized levels of mRNA expression of the biomarker genes can be transformed into a score, e.g., using the methods and models described herein.

[0092] DNA methylation detecting methods

[0093] Detection means for determining methylation level can include the use of methylation-specific polymerase chain reaction (MSP), nucleic acid sequencing, mass spectrometry, methylation-specific nuclease, mass-based separation or targeted capture.

[0094] In one embodiment, the detection of the methylated regions comprises performing bisulfite conversion of the DNA obtained from a biological sample from a subject; and detecting or determining methylation of multiple methylated regions of the bisulfite-converted DNA.

[0095] In one embodiment, the procedure of MSP detection method includes: 1) amplifying the methylated fragments in the selected target regions from the bisulfite-converted DNA using a pair of specific primers; 2) amplifying the non-methylated fragments in the selected target regions from the bisulfite-converted DNA using a pair of specific primers; 3) analyzing the amplified products resulting from the above 1) and 2) by an agar gel electrophoresis; and 4) determining the methylation levels of the selected target regions according to the presence or absence of the bands or density of the bands from the electrophoresis result.

[0096] In another embodiment, the MSP procedure includes: 1) quantifying the methylation level in the selected target regions for the bisulfite-converted DNA using a pair of primers and a probe for the selected target regions; 2) quantifying the non-methylation level in the selected target regions for the bisulfite-converted DNA using a specific primers and probes; and 3) calculating the methylation rate of each selected target region according to the non- methylation level and the quantification of the methylation level in each region.

[0097] In one particular embodiment, the process of DNA methylation-based chip detection includes: 1) amplifying the whole genome from the bisulfite-converted DNA; 2)contacting said bisulfite-converted DNA with a chip comprising methylated and nonmethylated capture probes to form a target-captured complex when a target region is present in said bisulfite-converted DNA; 3) performing a labelled single-base elongation reaction of said target-captured complex; and 4) amplifying and reading sequence signals captured according to the fluorescent staining reaction, and calculating the methylation levels in the target regions.

[0098] In one particular embodiment, the process of target DNA methylation sequencing process includes: 1) amplifying the whole genome from the bisulfite-converted DNA; 2) linking the amplified product in 1) with a linker; 3) targeted-capturing the library building product in 2); 4) sequencing of the capture products of 3); and 5) calculating the methylation levels in the selected target regions according to the sequencing result.

[0099] Methods for Diagnosing Colorectal Cancer

[0100] In some embodiments, the method disclosed herein comprises classifying the subject as 1) healthy or 2) having CRC or APL based on the measured biomarker panel. In some embodiments, the method comprises evaluating the measured biomarker panel by a machine learning classifier and determining that the subject is 1) healthy or 2) has CRC / APL.

[0101] In statistics, classification is the problem of identifying which of a set of categories an observation (or observations) belongs to. As used herein, the term “classification” used interchangeably with the term “classifying” refers to the identification of the subject as 1) being healthy or 2) having colorectal cancer or precancerous colorectal adenomas based on the measured levels of the biomarker panel. A “classifier” refers to an algorithm that implements the classification.

[0102] As used herein, the term “machine learning” refers to a computer-implemented technique that gives computer systems the ability to progressively improve performance on a specific task with data, i.e., to learn from the data, without being explicitly programmed. Machine learning technique adopts algorithms that can learn from and make prediction on data through building a model, i.e., a description of a system using mathematical concepts, from sample inputs. A core objective of machine learning is to generalize from the experience, i.e., to perform accurately on new data after having experienced a learning data set.

[0103] Machine learning models can be categorized as either supervised or unsupervised. Supervised learning involves learning a function that maps an input to an output based on example input-output pairs. In the context of biomedical diagnosis or prognosis, machine learning techniques generally involves supervised learning process, in which the computer is presented with example inputs (e.g., signature of gene expression) and their desired outputs(e.g., likelihood of having colorectal cancer or precancerous colorectal adenomas) to learn a general rule that maps inputs to outputs. Different models, i.e., hypothesis, can be employed in the generalization process. For the best performance in the generalization, the complexity of the hypothesis should match the complexity of the function underlying the data.

[0104] Supervised models include but not limited to logistic regression, support vector machine, decision trees, random forest, artificial neural network, linear regression, elastic net and naive bayes. In certain embodiments, the machine learning classifier is random forest or elastic net.

[0105] Logistic Regression

[0106] The simplest idea of linear regression is to find a line that best fits the data. Extensions of linear regression include multiple linear regression (e.g., finding a plane of best fit) and polynomial regression (e.g., finding a curve of best fit). Logistic regression is similar to linear regression but is used to model the probability of a finite number of outcomes, typically two.

[0107] Support Vector Machine

[0108] A Support Vector Machine (SVM) is a supervised classification technique that, at the most fundamental level, find a hyperplane or a boundary between two classes of data that maximizes the margin between the two classes. There are many planes that can separate the two classes, but only one plane can maximize the margin or distance between the classes.

[0109] Decision Tree

[0110] A decision tree is a decision support tool that uses a tree-like model of decisions and their possible consequences, including chance event outcomes, resource costs, and utility. Typically, a decision tree is a flowchart-like structure in which each internal node represents a “test” on an attribute, each branch represents the outcome of the test, and each leaf node represents a class label (decision taken after computing all attribute). The paths from root to leaf represent classification rules. Decision trees are intuitive and easy to build but fall short when it comes to accuracy.

[0111] Random Forest

[0112] Random forests are an ensemble learning technique that builds off of decision trees. Random forests involve creating multiple decision trees using bootstrapped datasets of the original data and randomly selecting a subset of variables at each step of the decision tree. The model then selects the mode of all of the predictions of each decision tree. By relying on a “majority wins” model, it reduces the risk of error from an individual tree.

[0113] Artificial Neural Network

[0114] Artificial neural networks (ANNs), usually simply called neural networks (NNs), are inspired by the biological neural networks that constitute animal brains. An ANN is based on a collection of connected units or nodes called artificial neurons, which loosely model the neurons in a biological brain. Each connection, like the synapses in a biological brain, can transmit a signal to other neurons. An artificial neuron receives a signal then processes it and can signal neurons connected to it. The “signal” at a connection is a real number, and the output of each neuron is computed by some non-linear function of the sum of its inputs. The connections are called edges. Neurons and edges typically have a weight that adjusts as learning proceeds. The weight increases or decreases the strength of the signal at a connection. Neurons may have a threshold such that a signal is sent only if the aggregate signal crosses that threshold. Typically, neurons are aggregated into layers. Different layers may perform different transformations on their inputs. Signals travel from the first layer (the input layer) to the last layer (the output layer), possibly after traversing the layers multiple times.

[0115] Elastic net

[0116] Elastic net is a regularized regression method that linearly combines the Li and L2 penalties of the LASSO (least absolute shrinkage and selection operator) and ridge methods. The LASSO method constructs a linear model, which penalizes the regression coefficients with an Li penalty, shrinking many of them to zero. Any features which have non-zero regression coefficients are “selected” by the lasso algorithm. As a result, the LASSO method performs both feature selection and regularization in order to enhance the prediction accuracy and interpretability of the resulting machine learning model. Elastic net is one of the improvements to the LASSO. Elastic net regularization combines the Li penalty of LASSO with the L2 penalty of ridge regression.

[0117] Unsupervised learning, on the other hand, is to draw inferences and find patterns from input data without references to labeled outcomes. Two main methods used in unsupervised learning include clustering and dimensionality reduction.

[0118] Clustering is an unsupervised technique that involves the grouping or clustering of data points. Common clustering algorithm include k-means clustering, hierarchical clustering, mean shift clustering, and density -based clustering.

[0119] Dimensionality reduction is a process of reducing the number of random variables to obtain a set of principle variables. Common dimensionality reduction algorithm include principal component analysis (PCA), regularized regression and Boruta.

[0120] The use of the machine learning classifier can classify the sample as 1) a healthy or 2) having CRC / APL with a sensitivity, specificity, positive predictive value, negativepredictive value, and / or overall accuracy of at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%.

[0121] In certain embodiments, the methods of the present disclosure further comprise sending the 1) a healthy or 2) having CRC classification results to a clinician. In some embodiments, the method of the present disclosure provides a diagnosis or prognosis in the form of a probability that the subject 1) is healthy or 2) having CRC. For example, the subject may have about a 0%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95% or greater probability of 1) being healthy or 2) having CRC.

[0122] Computer-implemented Methods, Systems and Devices

[0123] Any of the methods described herein may be totally or partially performed with a computer system including one or more processors, which can be configured to perform the steps. Thus, embodiments are directed to computer systems configured to perform the steps of any of the methods described herein, potentially with different components performing a respective step or a respective group of steps. Although presented as numbered steps, steps of methods herein can be performed at a same time or in a different order. Additionally, portions of these steps may be used with portions of other steps from other methods. Also, all or portions of a step may be optional. Any of the steps of any of the methods can be performed with modules, circuits, or other means for performing these steps.

[0124] Any of the computer systems mentioned herein may utilize any suitable number of subsystems. In some embodiments, a computer system includes a single computer apparatus, where the subsystems can be the components of the computer apparatus. In other embodiments, a computer system can include multiple computer apparatuses, each being a subsystem, with internal components. The subsystems can be interconnected via a system bus. Additional subsystems include, for examples, a printer, keyboard, storage device(s), monitor, which is coupled to display adapter, and others. Peripherals and input / output (I / O) devices, which couple to I / O controller, can be connected to the computer system by any number of means known in the art, such as serial port. For example, serial port or external interface (e.g. Ethernet, Wi-Fi, etc.) can be used to connect computer system to a wide area network such as the Internet, a mouse input device, or a scanner. The interconnection via system bus allows the central processor to communicate with each subsystem and to control the execution of instructions from system memory or the storage device(s) (e.g., a fixed disk, such as a hard drive or optical disk), as well as the exchange of information between subsystems. The system memory and / orthe storage device(s) may embody a computer readable medium. Any of the data mentioned herein can be output from one component to another component and can be output to the user.

[0125] A computer system can include a plurality of the same components or subsystems, e.g., connected together by external interface or by an internal interface. In some embodiments, computer systems, subsystem, or apparatuses can communicate over a network. In such instances, one computer can be considered a client and another computer a server, where each can be part of a same computer system. A client and a server can each include multiple systems, subsystems, or components.

[0126] It should be understood that any of the embodiments of the present disclosure can be implemented in the form of control logic using hardware (e.g., an application specific integrated circuit or field programmable gate array) and / or using computer software with a generally programmable processor in a modular or integrated manner. As used herein, a processor includes a multi-core processor on a same integrated chip, or multiple processing units on a single circuit board or networked. Based on the disclosure and teachings provided herein, a person of ordinary skill in the art will know and appreciate other ways and / or methods to implement embodiments of the present disclosure using hardware and a combination of hardware and software.

[0127] Any of the software components or functions described in this application may be implemented as software code to be executed by a processor using any suitable computer language such as, for example, Java, C++ or Perl using, for example, conventional or object- oriented techniques. The software code may be stored as a series of instructions or commands on a computer readable medium for storage and / or transmission, suitable media include random access memory (RAM), a read only memory (ROM), a magnetic medium such as a hard-drive or a floppy disk, or an optical medium such as a compact disk (CD) or DVD (digital versatile disk), flash memory, and the like. The computer readable medium may be any combination of such storage or transmission devices.

[0128] Such programs may also be encoded and transmitted using carrier signals adapted for transmission via wired, optical, and / or wireless networks conforming to a variety of protocols, including the Internet. As such, a computer readable medium according to an embodiment of the present invention may be created using a data signal encoded with such programs. Computer readable media encoded with the program code may be packaged with a compatible device or provided separately from other devices (e.g., via Internet download). Any such computer readable medium may reside on or within a single computer product (e.g. a hard drive, a CD, or an entire computer system), and may be present on or within different computerproducts within a system or network. A computer system may include a monitor, printer, or other suitable display for providing any of the results mentioned herein to a user.

[0129] Kits

[0130] In another aspect, the present disclosure provides kits or an integrated system for use in the methods described above. The kits may comprise any or all of the reagents to perform the methods described herein. In certain embodiments, the kit comprises primers for detecting the nucleic acids specific to the panel of biomarkers in a sample.

[0131] “Primer” as used herein refers to an oligonucleotide molecule with a length of 7-40 nucleotides, preferably 10-38 nucleotides, preferably 15-30 nucleotides, or 15-25 nucleotides, or 17-20 nucleotides. For example, the primer can an oligonucleotide having a length of 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 nucleotides. Primers are used in the amplification of a DNA sequence by polymerase chain reaction (PCR) as well known in the art. For a DNA template sequence to be amplified, a pair of primers can be designed at its 5’ upstream and its 3’ downstream sequence, i.e., 5’ primer and 3’ primer, each of which can specifically hybridize to a separate strand of the DNA double strand template. 5’ primer is complementary to the anti-sense strand of the DNA double strand template; and 3’ primer is complementary to the sense strand of the DNA template. As known in the art, the “sense strand” of a double stranded DNA template is the strand which contains the sequence identical to the mRNA sequence transcribed from the DNA template (except that “U” in RNA corresponds to “T” in the DNA) and encoding for a protein product. The complementary sequence of the sense strand is the “anti-sense strand.”

[0132] In certain embodiments, the kit further comprises an agent for amplifying the target nucleic acid using the primers. In addition, the kits may include instructional materials containing directions (i.e., protocols) for the practice of the methods provided herein. While the instructional materials typically comprise written or printed materials, they are not limited to such. Any medium capable of storing such instructions and communicating them to an end user is contemplated by this disclosure. Such media include but are not limited to electronic storage media (e.g., magnetic discs, tapes, cartridges, chips), optical media (e.g., CD ROM), and the like. Such media may include addresses to internet sites that provide such instructional materials.

[0133] In another aspect, the present disclosure provides oligonucleotide probes for detecting the nucleic acids specific to the panel of biomarkers in a sample. In certain embodiments, the probes are attached to a solid support, such as an array slide or chip, e.g., asdescribed in Eds., Bowtell and Sambrook DNA Microarrays: A Molecular Cloning Manual (2003) Cold Spring Harbor Laboratory Press. Construction of such devices are well known in the art, for example as described in US Patents and Patent Publications U.S. Patent No. 5,837,832; PCT application W095 / 11995; U.S. Patent No. 5,807,522; US Patent Nos. 7,157,229, 7,083,975, 6,444,175, 6,375,903, 6,315,958, 6,295,153, and 5,143,854, 2007 / 0037274, 2007 / 0140906, 2004 / 0126757, 2004 / 0110212, 2004 / 0110211, 2003 / 0143550, 2003 / 0003032, and 2002 / 0041420. Nucleic acid arrays are also reviewed in the following references: Biotechnol Annu Rev (2002) 8:85-101; Sosnowski et al. Psychiatr Genet (2002)12(4): 181-92; Heller, Annu Rev Biomed Eng (2002) 4: 129-53; Koi chinsky et al., Hum. Mutat (2002) 19(4):343-60; and McGail et al., Adv Biochem Eng Biotechnol (2002) 77:21-42.

[0134] A microarray can be composed of a large number of unique, single-stranded polynucleotides, usually either synthetic antisense polynucleotides or fragments of cDNAs, fixed to a solid support. Typical polynucleotides are preferably about 6-60 nucleotides in length, more preferably about 15-30 nucleotides in length, and most preferably about 18-25 nucleotides in length. For certain types of arrays or other detection kits / systems, it may be preferable to use oligonucleotides that are only about 7-20 nucleotides in length. In other types of arrays, such as arrays used in conjunction with chemiluminescent detection technology, preferred probe lengths can be, for example, about 15-80 nucleotides in length, preferably about 50-70 nucleotides in length, more preferably about 55-65 nucleotides in length, and most preferably about 60 nucleotides in length.

[0135] Techniques for the synthesis of these arrays using mechanical synthesis methods are described in, e.g., U.S. Patent No. 5,384,261. Although a planar array surface is often employed the array may be fabricated on a surface of virtually any shape or even a multiplicity of surfaces. Arrays may also be nucleic acids on beads, gels, polymeric surfaces, fibers such as fiber optics, glass or any other appropriate substrate, seeU.S. PatentNos. 5,770,358, 5,789,162, 5,708,153, 6,040,193 and 5,800,992. Arrays may be packaged in such a manner as to allow for diagnostics or other manipulation of an all-inclusive device.

[0111] The probes and primers necessary for practicing the present disclosure can be synthesized and labeled using well known techniques. Oligonucleotides used as probes and primers may be chemically synthesized according to the solid phase phosphoramidite triester method first described by Beaucage and Caruthers, Tetrahedron Letts. (1981) 22: 1859-1862, using an automated synthesizer, as described in Needham-Van Devanter et al, Nucleic Acids Res. (1984) 12:6159-6168.

[0136] Methods for Treating Colorectal Cancer

[0137] In yet another aspect, the present disclosure provides a method for treating CRC in a subject. In some embodiments, the method comprises administering to the subject a therapeutically effective amount of a drug useful for treating CRC, wherein the subject has been determined to have CRC based on the measured biomarkers comprising at least one bloodspecific biomarker and a CRC-specific biomarker.

[0138] In some embodiment, the drug used in the method of treating CRC as disclosed herein include, without limitation: Alymsys® (Bevacizumab), Avastin® (Bevacizumab), Camptosar® (Irinotecan Hydrochloride), Capecitabine, Cetuximab, Cyramza® (Ramucirumab), Eloxatin® (Oxaliplatin), Erbitux® (Cetuximab), 5-FU (Fluorouracil Injection), Fluorouracil Injection, Ipilimumab, Irinotecan Hydrochloride, Keytruda® (Pembrolizumab), Leucovorin Calcium, Lonsurf® (Trifluridine and Tipiracil Hydrochloride), Mvasi® (Bevacizumab), Opdivo® (Nivolumab), Oxaliplatin, Panitumumab, Pembrolizumab, Ramucirumab, Regorafenib, Stivarga® (Regorafenib), Trifluridine and Tipiracil Hydrochloride, Vectibix® (Panitumumab), Xeloda® (Capecitabine), Yervoy® (Ipilimumab), Zaltrap® (Ziv-Aflibercept), Zirabev® (Bevacizumab), Ziv-Aflibercept.

[0139] The drug described herein may be administered in any desired and effective manner: for oral ingestion, or as an ointment or drop for local administration to the eyes, or for parenteral or other administration in any appropriate manner such as intraperitoneal, subcutaneous, topical, intradermal, inhalation, intrapulmonary, rectal, vaginal, sublingual, intramuscular, intravenous, intraarterial, intrathecal, or intralymphatic. Further, the drug may be administered in conjunction with other treatments.

[0140] The following examples are provided to better illustrate the claimed disclosure and are not to be interpreted as limiting the scope of the disclosure. All specific compositions, materials, and methods described below, in whole or in part, fall within the scope of the present disclosure. These specific compositions, materials, and methods are not intended to limit the disclosure, but merely to illustrate specific embodiments falling within the scope of the disclosure. One skilled in the art may develop equivalent compositions, materials, and methods without the exercise of inventive capacity and without departing from the scope of the disclosure. It will be understood that many variations can be made in the procedures herein described while still remaining within the bounds of the present disclosure. It is the intention of the inventors that such variations are included within the scope of the disclosure.EXAMPLE 1

[0141] We have previously demonstrated that a number of blood specific biomarkers (HBA1, HBA2, HBB) had a similar ability to detect occult blood in stool as FIT and had similar diagnostic performance in identifying the presence of a colorectal cancer lesion (CRC) or advanced precancerous lesion (APL) (see PCT / US2024 / 050917). Here we show that pairing one of our blood specific biomarkers with other nucleic acid-based biomarkers has similar performance as pairing FIT with nucleic acid-based biomarkers.

[0142] We performed a study in which HBA1, HBA2, FIT and SDC2 methylated DNA were quantified in the stool of patients with CRC or APL as well as control patients who had a normal colonoscopy or who had benign hyperplastic polyps. The primary outcome of this study is the sensitivity and specificity of detection for CRC, and APL via pairing FIT with a methylated DNA CRC biomarker or pairing the same methylated DNA biomarker with our blood specific mRNA biomarkers (HBA1, HBA2). Methylated SDC2 is a well-established CRC / APL stool-based biomarker and is the basis of a number of existing or in development CRC / APL screening tests.

[0143] HBA1 and HBA2 were selected for testing on clinical samples because these genes had the highest expression in whole blood based on our analysis of RNA-seq whole blood samples and we previously showed these genes had similar performance in detecting CRC and APL lesions as does FIT. These genes also code for hemoglobin subunits and hemoglobin protein is the primary target for most existing FIT assays. HBA1 and HBA2 have very high sequence similarity so TaqMan primers and a TAqMan probe were designed to recognize either HBA1 or HBA2 sequence. Recognizing two blood associated genes with a single TaqMan assay increases the sensitivity of detection without increasing the cost or complexity.RNA was extracted from stool collected from 60 patients, 15 CRC, 15 APL and 30 Controls. Advanced precancerous lesions were defined using generally accepted criteria i.e. greater than 1 cm in size in any dimension, or a villous component greater than 25% or the presence of high-grade dysplasia.

[0144] E.Z.N.A Stool RNA Kit (OMEGA®, product number: R6828) can be used to extract RNA from stool samples. By placing up to 200 mg of stool sample in a 2 mL centrifuge tube, and following the protocol provided by the above kit, total stool RNA can be extracted.

[0145] Method of extracting and measuring methylated SDC2 DNA from stool sample

[0146] A commercial kit approved for CRC screening in China was used to measure the level of SDC2 methylation in stool (Changanxin® Human SDC2 Gene MethylationDetection Kit). The assay was performed per the kit instructions. Genomic DNA was extracted from 2.5g of human stool using the specified SDC2 Methylation Detection Kit (fluorescence PCR method). The protocol involved enzymatic lysis, SPE column purification, and bisulfite conversion using sodium nitrite to differentiate methylated DNA. Critical steps included sequential washing with ethanol-diluted buffers (e.g., Solution C at a 1 :4 ethanol ratio), magnetic bead-based purification, and thermal incubation (95 °C for denaturation). Converted DNA was eluted and stored at -20°C. Fluorescent PCR was conducted in 30 p.L reactions containing PCR Mix 1 (DNA polymerase, dNTPs), PCR Mix 2 (primers targeting methylated SDC2 and ACTB), and 10 p.L DNA template. Amplification involved 48 cycles (95° C / 15 s, 85 ° C / 90 s, 72 ° C / 30 s) on compatible thermal cyclers (e.g., Roche LightCycler 480 II). Methylation status was determined by Ct values: ACTB (Texas Red channel, Ct 36) validated sample integrity, while SDC2 (FAM channel, Ct <38) indicated methylation positivity. Invalid results (ACTB Ct36) required retesting. Strict contamination controls and biosafety protocols were enforced throughout.

[0147] Method to perform FIT

[0148] A fecal immunochemical test (FIT) was simultaneously performed for each sample as RNA extraction using the Luminex platform from Tellgen Corporation, China, employing flow cytometry to quantitatively detect hemoglobin (HB) and transferrin (TF) protein in human fecal sample suspensions.

[0149] Method to measure HBA1 and HBA2

[0150] After RNA was extracted from stool the amount of HBA1 and HBA2 transcripts were quantified using RT-qPCR with a TaqMan assay. Since the HBA1 and HBA2 TaqMan assay was designed to recognize either HBA1 or HBA2, henceforth we will refer to this assay as HBA collectively. The PCR protocol used was as follows: Takara One Step PrimeScript™ III RT-qPCR Mix with UNG (RR601A) was utilized for qRT-PCR. Each reaction was set up with a total volume of 10 pL, including 2 pL of extracted RNA. The reactions were performed on an ABI 7500 Real-Time PCR System with the following thermal cycling conditions: reverse transcription at 50°C for 5 minutes, initial denaturation at 95°C for 10 seconds, followed by 40 cycles of 95°C for 10 seconds and 60°C for 34 seconds.

[0151] The primers and probes used in the TaqMan assays were designed in-house to recognize spliced mRNA and not genomic DNA. This was done by having either a primer or a probe span a splice junction. The primer and probe sequences are given in Table 6. It is notedthat either HBA1 & HBA2 probe or HBB probe includes a FAM moiety at the 5’ end of the DNA sequence, and a MGB moiety at the 3’ end of the DNA sequence, respectively.

[0152] The diagnostic performance of the different biomarker combinations was obtained using 10-fold cross validation and random forests with default parameters.

[0153] Table 6: Sequence of Primers and Probes

[0154] FIT performance by itself was 86.7% sensitivity for CRC detection, 66% sensitivity for APL detection and 86.7% specificity. HBA had very similar performance as FIT with the same CRC sensitivity and specificity (Table 7). There was a strong degree of overlap in samples that were high for both FIT and the mRNA-based measure of occult blood (HBA). Of the 13 out of 15 CRC samples that were positive for FIT, all of them were also positive for HBA. The correlation between the levels of occult blood measured using FIT and our HBA biomarker was strongly statistically significant R= -0.55 (Pearson's Correlation Coefficient), p-value=6e-6 and in the correct direction.

[0155] Table 7: Performance of stool mRNA occult blood biomarkers paired with DNA methylation

[0156] For this dataset SDC2 methylation was the poorest performing biomarker with a CRC sensitivity of 67.7% and a AA sensitivity of 33.3% at a 80% specificity. Importantlypairing HBA with the SDC2 methylation biomarker had similar performance as pairing FIT with SDC2 methylation.

[0157] Our results demonstrate that using mRNA biomarkers strongly associated with blood; similar performance in detecting occult blood in stool and screening for cancerous and advanced precancerous lesions can be obtained as the current gold standard FIT protein based occult blood detection method. In addition we show, our mRNA blood associated biomarkers can be used as a replacement for FIT when paired with other nucleic acid biomarkers.

[0158] Current multi-target stool CRC screening methods pair nucleic acid-based assays with FIT, which increases cost and difficulty of use. With our nucleic acid-based approach to occult blood detection in stool, we can pair the signal from occult blood with other nucleic acid-based biomarkers without requiring two different assays, workflows, or methods of stool collection. This represents a considerable savings in cost and complexity and is a real breakthrough in the field. While we tested pairing our blood associated mRNA biomarker with one nucleic acid-based biomarker of a different type (SDC2 DNA methylation), it is expected our mRNA biomarkers can serve as a replacement for FIT when paired with other nucleic acidbased biomarkers. This is due to the close and strong overlap between HBA positive and FIT positive samples and the statistically significant correlation between the levels of HBA and FIT in patient stool samples.

[0159] While the disclosure has been particularly shown and described with reference to specific embodiments (some of which are preferred embodiments), it should be understood by those having skill in the art that various changes in form and detail may be made therein without departing from the spirit and scope of the present disclosure as disclosed herein.References1. Ashktorab, H., Brim, H., 2014. DNA Methylation and Colorectal Cancer. Curr. Colorectal Cancer Rep. 10, 425-430. https: / / doi.org / 10.1007 / sl l888-014-0245-22. Cameron, A. D. Gastro-intestinal blood loss measured by radioactive chromium. Gut 1, 177-182 (1960).3. Chen, J.-J., Wang, A.-Q., Chen, Q.-Q., 2017. DNA methylation assay for colorectal carcinoma. Cancer Biol. Med. 14, 42-49. https: / / doi.Org / 10.20892 / j.issn.2095-3941.2016.00824. COSMIC: the Catalogue Of Somatic Mutations In Cancer | Nucleic Acids Research |Oxford Academic [WWW Document], n.d. URL https: / / academic.oup.com / nar / article / 47 / Dl / D941 / 5146192 (accessed 1.24.24).5. Ebaugh, F. G., Clemens, T., Rodnan, G. & Peterson, R. E. Quantitative measurement of gastrointestinal blood loss. I. The use of radioactive Cr51 in patients with gastrointestinal hemorrhage. Am. J. Med. 25, 169-181 (1958).6. Geneoscopy Submits Premarket Approval Application to FDA for its Noninvasive Colorectal Cancer RNA Biomarker Screening Test - Geneoscopy - Transforming Gastrointestinal Health, 2023. URL https: / / www.geneoscopy.com / geneoscopy-submits- premarket-approval-application-to-fda-for-its-noninvasive-colorectal-cancer-ma-biomarker- screening-test / (accessed 1.24.24).7. He, C., Huang, Q., Zhong, S., Chen, L.S., Xiao, H., Li, L., 2023. Screening and identifying of biomarkers in early colorectal cancer and adenoma based on genome-wide methylation profiles. World J. Surg. Oncol. 21, 312. https: / / doi.org / 10.1186 / sl2957-023- 03189-18. Kralik, P., Ricchi, M., 2017. A Basic Guide to Real Time PCR in Microbial Diagnostics: Definitions, Parameters, and Everything. Front. Microbiol. 8.9. Lee, M.W., Pourmorady, J.S., Laine, L., 2020. Use of Fecal Occult Blood Testing as a Diagnostic Tool for Clinical Indications: A Systematic Review and Meta- Analysis. Am. J. Gastroenterol. 115, 662-670. https: / / doi.org / 10.14309 / ajg.000000000000049510. Medical Advisory Secretariat. Fecal occult blood test for colorectal cancer screening: an evidence-based analysis. Ont. Health TechnoL Assess. Ser. 9, 1-40 (2009).11. Medical Advisory Secretariat. Optical coherence tomography for age-related macular degeneration and diabetic macular edema: an evidence-based analysis. Ont. Health TechnoL Assess. Ser. 9, 1-22 (2009).12. Muller, D., Gyorffy, B., 2022. DNA methylation-based diagnostic, prognostic, and predictive biomarkers in colorectal cancer. Biochim. Biophys. Acta BBA - Rev. Cancer 1877, 188722. https: / / doi.Org / 10.1016 / j.bbcan.2022.18872213. Multitarget Stool DNA Testing for Colorectal-Cancer Screening | NEJM [WWW Document], n.d. URL https: / / www.nejm.org / doi / full / 10.1056 / nejmoal311194 (accessed 9.6.22).14. Navarro, M., Nicolas, A., Ferrandez, A. & Lanas, A. Colorectal cancer population screening programs worldwide in 2016: An update. World J. Gastroenterol. 23, 3632-3642 (2017).15. Niedermaier, T., Tikk, K., Gies, A., Bieck, S., Brenner, H., 2020. Sensitivity of FecalImmunochemical Test for Colorectal Cancer Detection Differs According to Stage and Location. Clin. Gastroenterol. Hepatol. 18, 2920-2928. e6. https : / / doi . org / 10.1016 / j . cgh.2020.01.02516. Rademakers, G., Massen, M., Koch, A., Draht, M.X., Buekers, N., Wouters, K.A.D., Vaes, N., De Meyer, T., Carvalho, B., Meijer, G.A., Herman, J.G., Smits, K.M., van Engeland, M., Melotte, V., 2021. Identification of DNA methylation markers for early detection of CRC indicates a role for nervous system-related genes in CRC. Clin. Epigenetics 13, 80. https: / / doi.org / 10.1186 / sl3148-021-01067-917. Stracci, F., Zorzi, M., Grazzini, G., 2014. Colorectal Cancer Screening: Tests, Strategies, and Perspectives. Front. Public Health 2.18. Stiirzlinger, H., Conrads-Frank, A., Eisenmann, A., Invansits, S., Jahn, B., Janzic, A., Jelenc, M., Kostnapfel, T., Mencej Bedrac, S., Miihlberger, N., Siebert, U., Sroczynski, G., 2023. Stool DNA testing for early detection of colorectal cancer: systematic review using the HTA Core Model® for Rapid Relative Effectiveness Assessment. GMS Ger. Med. Sci. 21, Doc06. https: / / doi.org / 10.3205 / 00032019. Tinmouth, J., Lansdorp-Vogelaar, I. & Allison, J. E. Faecal immunochemical tests versus guaiac faecal occult blood tests: what clinicians and colorectal cancer screening programme organisers need to know. Gut 64, 1327-1337 (2015).20. Turner, A., Sasse, J., Varadi, A., 2015. Development and validation of a high throughput, closed tube method for the determination of haemoglobin alpha gene (HBA1 and HBA2) numbers by gene ratio assay copy enumeration-PCR (GRACE-PCR). BMC Med. Genet. 16, 115. https: / / doi.org / 10.1186 / sl2881-015-0258-y21. van Rossum, L. G. et al. Random comparison of guaiac and immunochemical fecal occult blood tests for colorectal cancer in a screening population. Gastroenterology 135, 82- 90 (2008).22. Zhuang, Y., Wang, H., Jiang, D., Li, Y., Feng, L., Tian, C., Pu, M., Wang, X., Zhang, J., Hu, Y., Liu, P., 2021. Multi gene mutation signatures in colorectal cancer patients: predict for the diagnosis, pathological classification, staging and prognosis. BMC Cancer 21, 380. https: / / doi.org / 10.1186 / sl2885-021-08108-923. Zitt, Marion, Zitt, Matthias, Muller, H.M., 2007. DNA methylation in colorectal cancer- -impact on screening and therapy monitoring modalities? Dis. Markers 23, 51-71. https: / / doi.org / 10.1155 / 2007 / 891967

Claims

WHAT IS CLAIMED IS:

1. A method of diagnosing colorectal cancer (CRC) or advanced precancerous lesion (APL) in a subject, said method comprising: obtaining a stool sample of the subject; measuring in the stool sample a messenger RNA (mRNA) level of at least one blood-specific gene; detecting in the stool sample a mutation detected at the DNA or an abnormal expression at the RNA level of at least one CRC / APL-mutated gene or a DNA methylation of at least one CRC / APL methylated gene; evaluating the measured mRNA level of the blood-specific gene, the mutation or abnormal expression of the CRC / APL-mutated gene or the DNA methylation of the CRC / APL-methylated gene; and determining whether the subject is 1) healthy or 2) has CRC or APL.

2. The method of claim 1, wherein the at least one blood-specific gene is selected from Table 1 (blood group antigen genes), Table 2 (platelet genes) or Table 3 (hemoglobin subunit genes); optionally, wherein the at least one blood-specific gene is selected from the group consisting of: HBE1, HBG1, HBG2, HBM, HBQ1, HBZ, HBB, HBA2, HBD, HBA1, S100A9, CSF3R, IFITM2 and LCP1; preferably, wherein the at least one blood-specific gene is selected from the group consisting of: HBA1, HBA2 and HBB.

3. The method of claim 1 or 2, wherein the at least one CRC / APL-mutated gene is selected from Table 4; optionally, wherein at least one CRC / APL mutated gene is selected from the group consisting of: KRAS and BRAF.

4. The method of claim 1 or 2, wherein at least one CRC / APL-methylated gene is selected from Table 5;Optionally, wherein at least one CRC / APL-methylated gene is selected from the group consisting of: SDC2, NDRG4, BMP3, LASS4, LRRC4, PPP2R5C.

5. The method of any one of the preceding claims, wherein the mRNA level of either the blood-marker or CRC / APL-mutated gene is measured using quantitative RT-PCR or Droplet Digital PCR.

6. The method of any one of the preceding claims, wherein the mRNA level or mutation is measured by nucleic acid sequencing.

7. The method of any one of the preceding claims, wherein the DNA methylation is detected by methylation-specific polymerase chain reaction (MSP), nucleic acid sequencing, mass spectrometry, methylation-specific nuclease, mass-based separation or targeted capture.

8. The method of claim any one of the preceding claims, wherein the evaluating step and / or the determining step comprises using a machine learning classifier.

9. A kit or an integrated system of diagnosing CRC or APL in a subject, comprising (a) a first agent for detecting in a stool sample obtained from the subject a mRNA level of at least one blood-specific gene and (b) a second agent for detecting in the stool sample a mutation at DNA level or an abnormal expression at RNA level of at least one CRC / APL-mutated gene or a DNA methylation of at least one CRC / APL-methylated gene.

10. The kit or the integrated system of claim 9, wherein the at least one blood-specific gene is selected from Table 1 (blood group antigen genes), Table 2 (platelet genes) or Table 3 (hemoglobin subunit genes); optionally, wherein the at least one blood-specific gene is selected from the group consisting of: HBE1, HBG1, HBG2, HBM, HBQ1, HBZ, HBB, HBA2, HBD, HBA1, S100A9, CSF3R, IFITM2 and LCP1; preferably, wherein the at least one blood-specific gene is selected from the group consisting of HBA1, HBA2 and HBB.

11. The kit or the integrated system of claim 9 or 10, wherein the at least one CRC / APL- mutated gene is selected from Table 4; optionally, wherein at least one CRC / APL mutated gene is selected from the group consisting of KRAS and BRAF.

12. The kit or the integrated system of claim 9 or 10, wherein at least one CRC / APL- methylated gene is selected from Table 5; optionally, wherein at least one CRC / APL-methylated gene is selected from the group consisting of SDC2, NDRG4, BMP3, LASS4, LRRC4, PPP2R5C.

13. The kit or the integrated system of anyone of claims 9-12, wherein the first or the second agent is selected from the group consisting of primers, nucleic acids and oligonucleotides.

14. A method for treating CRC in a subject, the method comprising: administering to the subject a therapeutically effective amount of a drug useful for treating CRC,wherein the subject has been determined to have CRC based on (a) a mRNA level of at least one blood-specific gene in a stool sample of the subject, and (b) a mutation at DNA level or an abnormal expression at RNA level of at least one CRC -mutated gene or a DNA methylation of at least one CRC-methylated gene in the stool sample.

15. The method of claim 14, wherein the drug is selected from the group consisting of Bevacizumab, Irinotecan Hydrochloride, Capecitabine, Cetuximab, Ramucirumab, Oxaliplatin, Cetuximab, 5-FU, Ipilimumab, Irinotecan Hydrochloride, Pembrolizumab, Leucovorin Calcium, Trifluridine, Tipiracil Hydrochloride, Nivolumab, Panitumumab, Regorafenib, and Ziv-Aflibercept.

Citation Information

Patent Citations

  • K-Ras Oligonucleotide Microarray and Method for Detecting K-Ras Mutations Employing the Same

    US20070298419A1

  • Use of picoplatin and bevacizumab to treat colorectal cancer

    US20110052580A1

  • Method of diagnosing neoplasms - ii

    US20190024188A1