Compositions and methods for fecal occult blood test in diagnosing gastrointestinal disease

By measuring blood-specific and disease-specific mRNA levels in stool samples with a panel of biomarkers and a machine learning classifier, the method addresses the complexity and cost issues of current FOBTs, providing a more efficient and accurate diagnosis for gastrointestinal diseases.

WO2025080943A9PCT designated stage expired Publication Date: 2025-08-28EL CAPITAN BIOSCIENCES INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/050917
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-11
Filing Date
2024-10-11
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Current fecal occult blood tests (FOBTs) for gastrointestinal diseases, such as colorectal cancer, are non-specific and require multiple assays, leading to complexity and high costs, with existing methods like guaiac-based tests being interfered by foodstuffs and immunological tests needing dietary modifications.

Method used

A method involving the measurement of blood-specific and disease-specific mRNA levels in stool samples using quantitative RT-PCR or Droplet Digital PCR, combined with a machine learning classifier, to diagnose gastrointestinal diseases like colorectal cancer, utilizing a panel of biomarkers including HBB, HBA2, HBA1, S100A9, and other genes.

Benefits of technology

This approach simplifies the diagnostic process, enhances sensitivity, and reduces costs by using a single assay to detect gastrointestinal diseases with high accuracy, improving patient experience and reducing the complexity of current screening methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000010_0001
    Figure IMGF000010_0001
  • Figure IMGF000011_0001
    Figure IMGF000011_0001
  • Figure IMGF000012_0001
    Figure IMGF000012_0001
Patent Text Reader

Abstract

The present disclosure provides methods and compositions, e.g., kits, for diagnosing a gastrointestinal disease based on mRNA levels of at least one blood-specific gene in a feces sample from a subject. In some embodiments, the gastrointestinal disease is diagnosed based on mRNA levels of a panel of biomarkers comprising at least one blood-specific gene and a group of disease-specific genes.
Need to check novelty before this filing date? Find Prior Art

Description

COMPOSITIONS AND METHODS FOR FECAL OCCULT BLOOD TEST IN _ DIAGNOSING GASTROINTESTINAL DISEASE _SEQUENCE LISTING

[0001] The sequence listing that is contained in the file named “081996-8004W001”, which is 6,018 bytes (as measured in Microsoft Windows) and was created on October 03, 2024, is filed herewith by electronic submission and is incorporated by reference herein.CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to US provisional application 63 / 589,624, filed October 11, 2023, the disclosure of which is incorporated herein by reference.FIELD OF THE INVENTION

[0003] The present disclosure generally relates to diagnosis and treatment of gastrointestinal symptoms or conditions. In particular, the present disclosure relates to compositions and methods for detecting blood-specific mRNA in a feces sample for diagnosing and / or treating gastrointestinal diseases.BACKGROUND

[0004] The fecal occult blood test (FOBT) is a diagnostic test to assess the presence of blood in stool. In healthy subjects, the volume of blood leaked into the gastrointestinal tract is 0.5 to 1.5 mL per day1,2. This small amount of blood is not usually detected in the fecal occult blood test3. Although FOBTs are commonly used as a diagnostic test in clinical settings including iron deficiency anemia, ulcerative colitis (UC) and acute diarrhea, their performance in predicting presumptive causes is poor4. For Colorectal Cancer (CRC), FOBT is only used for screening5. The presence of occult blood in stool is well established as a biomarker to identify the potential presence of colorectal cancer or advanced precancerous lesions (APL).

[0005] FOBTs commonly used in clinical laboratories are based on detecting heme, heme- derived porphyrins, or colorectal hemoglobin protein6. Heme can be detected by guaiac-based methods (gFOBT, guaiac based fecal occult blood test), which utilizes the pseudo-peroxidase activity of hemoglobin, wherein guaiac is oxidized by hydrogen peroxidase. Because this reaction takes place with any peroxidase present in stool, gFOBT tests are non-specific to human Hb, with interference by any foodstuffs with peroxidase content, by certain chemicals or even medications. Heme can also be detected by converting the non-fluorescent heme to fluorescent porphyrins7. Human hemoglobin can be detected by immunological methods (fecal immunochemical test or FIT). The only randomized clinical trials (RCT) comparing gFOBTand FIT published so far from van Rossum et al. concluded that the performance of FIT is clearly superior to that of gFOBTs in detecting any type of colorectal neoplasia8. Moreover, FIT does not necessitate any dietary modifications, improving patient adherence7.

[0006] The genomic approach to the management of gastrointestinal diseases is now a part of practice in colon cancer screening with the FDA-approved FIT-fecal DNA test9, as well as in the diagnosis of infectious diarrhea10and Covid-1911. Currently, at least 2 multi-target stool mRNA tests are under development in CRC screening12. Gene expression mRNA biomarkers from tumor cells exfoliated in the intestinal lumen are exponentially amplified relative to DNA markers13which may provide additional detection sensitivity. Current multitarget nucleic acidbased approaches to screen for colorectal cancer and advanced precancerous lesions depend on combining the signal obtained from at least two assays: a PCR based assay and a protein-based assay (FIT). This complicates the patient experience and adds to the cost of the colorectal screening test. There is a continuing need to develop new diagnostic methods with good sensitivity for detecting gastrointestinal disease and fecal occult blood, which is less complex and of lower cost.SUMMARY OF INVENTION

[0007] The present disclosure in one aspect provides a method for diagnosing a gastrointestinal disease in a subject. In one embodiment, the method comprises: obtaining a ribonucleic acid (RNA) from a stool sample of the subject; measuring messenger RNA (mRNA) level of at least one blood-specific gene in the obtained RNA; evaluating the measured mRNA level of the at least one blood-specific gene; and determining whether the subject is 1) healthy or 2) has the gastrointestinal disease.

[0008] In certain embodiment, the at least one blood-specific gene is selected from Table 1. In certain embodiments, the at least one blood-specific gene is selected from the group consisting of: HBB, HBA2, HBD, HBA1, S100A9, CSF3R, IFITM2 and LCP1. In certain embodiments, the at least one blood-specific gene is selected from the group consisting of: HBA1, HBA2 and HBB. In certain embodiments, the at least one blood-specific gene is selected from the group consisting of: HBE1, HBG1, HBG2, HBM, HBQ1 and HBZ. In certain embodiments, the at least one blood-specific gene is a blood-specific biomarker (e.g., a blood antigen gene or a platelet-specific gene).

[0009] In some embodiments, the method for diagnosing a gastrointestinal disease as disclosed herein further comprises measuring mRNA levels of one or more genes selected from a group of disease-specific genes.

[0010] In some embodiments, the gastrointestinal disease is colorectal cancer (CRC). In some embodiments, the group of CRC-specific genes are selected from Tables 2-4. In some embodiments, the group of CRC-specific genes are selected from Tables 2 and 3. In some embodiments, the group of CRC-specific genes are selected from Table 2.

[0011] In some embodiments, wherein the subject is a human.

[0012] In some embodiments, the mRNA levels are detected using quantitative RT-PCR or Droplet Digital PCR or Partition PCR. In some embodiments, wherein the mRNA levels are determined by nucleic acid sequencing. In some embodiments, wherein the mRNA levels are determined by CRISPR-based assay, or nuclease protection assay.

[0013] In some embodiments, the evaluating step and / or the determining step comprises analyzing the levels of a panel of biomarkers by a machine learning classifier. In some embodiments, the machine learning classifier is random forests or elastic net.

[0014] In another aspect, the present disclosure provides a panel of mRNA biomarkers for use in diagnosing a gastrointestinal disease in a subject, wherein the panel of mRNA biomarkers comprises at least one blood-specific gene. In some embodiments, the panel of mRNA biomarkers further comprises a group of disease-specific genes.

[0015] In another aspect, the present disclosure provides a kit or an integrated system of diagnosing a gastrointestinal disease in a subject. In some embodiments, the kit or integrated system comprises an agent for detecting in a stool sample obtained from the subject mRNA levels of a panel of biomarkers comprising at least one blood-specific gene. In some embodiments, the panel of biomarkers further comprises a group of disease-specific genes. In some embodiments, the agent is selected from the group consisting of: primers, nucleic acids and oligonucleotides.

[0016] In another aspect, the present disclosure provides use of a biomarker-specific reagent in the manufacture of a kit for diagnosing a gastrointestinal disease in a subject, wherein the biomarker-specific reagent specifically binds to a panel of mRNA biomarkers comprising at least one blood-specific gene. In some embodiment, the panel of mRNA biomarkers further comprises a group of disease-specific genes.

[0017] In yet another aspect, the present disclosure provides a method for treating a gastrointestinal disease in a subject. In some embodiments, the method comprises: administering to the subject a therapeutically effective amount of a drug useful for treating the gastrointestinal disease, wherein the subject has been determined to have the gastrointestinal disease based on mRNA levels of a panel of biomarkers comprising at least one blood-specificgene. In some embodiments, the panel of mRNA biomarkers further comprises a group of disease-specific genes.BRIEF DESCRIPTION OF DRAWING

[0018] The following drawings form part of the present specification and are included to further demonstrate certain aspects of the present disclosure. The disclosure may be better understood by reference to one or more of these drawings in combination with the detailed description of specific embodiments presented herein.

[0019] FIG. 1 illustrates the enhanced expression of HBA and HBB in stool samples from CRC patients as compared to normal controls. Top panel is a scatter plot of the Ct values observed for those stool samples with detectable expression. Each bottom panel represents the percentage of samples for each sample type for which detectable expression could be measured, respectively. The TaqMan assay used was designed to recognize either HBA1 or HBA2 therefore the results of this assay are referred to as HBA collectively. Strong differences in measured Ct values are observed comparing stool from CRC patients compared to normal controls.

[0020] FIG. 2 illustrates the comparison of HBA and HBB expression across all 90 samples. Each dot represents a sample. The TaqMan assay used was designed to recognize either HBA1 or HBA2, therefore, the results of this assay are referred to as HBA. For those samples that expression could not be detected the Ct value was set to 40. A strong correlation between HBA and HBB is observed.DETAILED DESCRIPTION OF THE INVENTION

[0021] Before the present disclosure is described in greater detail, it is to be understood that this disclosure is not limited to particular embodiments described, and as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present disclosure will be limited only by the appended claims.

[0022] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present disclosure, the preferred methods and materials are now described.

[0023] All publications and patents cited in this specification are herein incorporated by reference as if each individual publication or patent were specifically and individually indicated to be incorporated by reference and are incorporated herein by reference to disclose and describe the methods and / or materials in connection with which the publications are cited. The citation of any publication is for its disclosure prior to the filing date and should not be construed as an admission that the present disclosure is not entitled to antedate such publication by virtue of prior disclosure. Further, the dates of publication provided could be different from the actual publication dates that may need to be independently confirmed.

[0024] As will be apparent to those of skill in the art upon reading this disclosure, each of the individual embodiments described and illustrated herein has discrete components and features which may be readily separated from or combined with the features of any of the other several embodiments without departing from the scope or spirit of the present disclosure. Any recited method can be carried out in the order of events recited or in any other order that is logically possible.

[0025] Definitions

[0026] The following definitions are provided to assist the reader. Unless otherwise defined, all terms of art, notations and other scientific or medical terms or terminology used herein are intended to have the meanings commonly understood by those of skill in the chemical and medical arts. In some cases, terms with commonly understood meanings are defined herein for clarity and / or for ready reference, and the inclusion of such definitions herein should not necessarily be construed to represent a substantial difference over the definition of the term as generally understood in the art.

[0027] As used herein, the singular forms “a”, “an” and “the” include plural references unless the context clearly dictates otherwise.

[0028] As used herein, the term “administering” means providing a pharmaceutical agent or composition to a subject, and includes, but is not limited to, administering by a medical professional and self-administering.

[0029] The term “level” generally refers to the quantity of a substance of interest. In the context of a biomarker, a level of a biomarker refers to the quantity of the polynucleotide (e.g., mRNA) of interest or the polypeptide of interest present in a sample. Such quantity may be expressed in the absolute terms, i.e., the total quantity of the polynucleotide or polypeptide in the sample, or in the relative term, i.e., the concentration of the polynucleotide or polypeptide in the sample.

[0030] The terms “assessing”, “assaying”, “measuring” and “detecting” can be used interchangeably and refer to both quantitative and semi-quantitative determinations. Where either a quantitative and semi-quantitative determination is intended, the phrase “measuring a level” of a polynucleotide or polypeptide of interest or “detecting” a polynucleotide or polypeptide of interest can be used.

[0031] As used herein, the term “biomarker” refers to a detectable organic biomolecule associated with a particular phenotype or risk of developing a particular phenotype, such as a polynucleotide (e.g., RNA (e.g., mRNA) or DNA (e.g., cDNA)) or a polypeptide, which is differentially present in a biological sample taken from a subject having a certain condition (e.g., having a gastrointestinal disease) as compared to a comparable biological sample taken from a subject who does not have such condition, such as a healthy subject or a non-cancer patient. For example, a biomarker can be a polynucleotide, such as RNA (e.g., mRNA), which is present at an elevated or decreased level in a biological sample (e.g., a tissue sample, a feces sample, or a blood sample) of a gastrointestinal disease patient compared to a comparable sample (e.g., a feces sample) of a subject with a negative diagnosis (e.g., a healthy subject).

[0032] As used herein, the term “gastrointestinal disease” refers to a broad category of medical conditions that affect the gastrointestinal (GI) tract, which includes the esophagus, stomach, intestines, and accessory organs such as the liver and pancreas. It includes at least both colorectal cancer (CRC) and advanced precancerous lesions (APL).

[0033] As used herein, the term “advanced precancerous lesions” refer to tissue abnormalities that have a higher risk of progressing to cancer if left untreated. These lesions are more developed or progressed compared to early-stage precancerous changes and show more significant cellular atypia (abnormalities in cell shape, size, and organization) and architectural disorganization.

[0034] The term “colon cancer” used interchangeably with the term “colorectal cancer” or “rectal cancer” refers to any cancerous neoplasia of the colon (including the rectum and appendix). Many colorectal cancers arise from precancerous colorectal adenomas or adenomatous polyps, which are usually benign, but some may develop into cancer over time. The diagnosis of localized colon cancer is often through colonoscopy. Once localized colon cancer is diagnosed, it is usually surgically removed and then treated with chemotherapy. Precancerous colorectal adenomas or colorectal adenomatous polyps are a risk factor for colorectal cancer. The removal of colorectal adenomatous polyps at the time of colonoscopy would reduce the risk of having colorectal cancer. In addition, clinical data has shown that earlydetection and curative surgical resection of colorectal cancer will significantly improve survival rates.

[0035] It is noted that in this disclosure, terms such as “comprises”, “comprised”, “comprising”, “contains”, “containing” and the like have the meaning attributed in United States Patent law; they are inclusive or open-ended and do not exclude additional, un-recited elements or method steps. Terms such as “consisting essentially of’ and “consists essentially of’ have the meaning attributed in United States Patent law; they allow for the inclusion of additional ingredients or steps that do not materially affect the basic and novel characteristics of the claimed disclosure. The terms “consists of’ and “consisting of’ have the meaning ascribed to them in United States Patent law; namely that these terms are close ended.

[0036] The term “nucleic acid” and “polynucleotide” are used interchangeably and refer to a polymeric form of nucleotides of any length, either deoxyribonucleotides or ribonucleotides, or analogs thereof. Polynucleotides may have any three-dimensional structure, and may perform any function, known or unknown. Non-limiting examples of polynucleotides include a gene, a gene fragment, exons, introns, messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozymes, cDNA, shRNA, single-stranded short or long RNAs, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of any sequence, control regions, isolated RNA of any sequence, nucleic acid probes, and primers. The nucleic acid molecule may be linear or circular.

[0037] As used herein, the term “subject” refers to a human or any non-human animal (e.g., mouse, rat, rabbit, dog, cat, cattle, swine, sheep, horse or primate). A human includes pre and post-natal forms. In many embodiments, a subject is a human being. A subject can be a patient, which refers to a human presenting to a medical provider for diagnosis or treatment of a disease. The term “subject” is used herein interchangeably with “individual” or “patient”. A subject can be afflicted with or is susceptible to a disease or disorder but may or may not display symptoms of the disease or disorder.

[0038] As used herein, the term “therapeutically effective amount” means the amount of agent that is sufficient to prevent, treat, reduce and / or ameliorate the symptoms and / or underlying causes of any disorder or disease, or the amount of an agent sufficient to produce a desired effect on a cell. In one embodiment, a “therapeutically effective amount” is an amount sufficient to reduce or eliminate a symptom of a disease. In another embodiment, a therapeutically effective amount is an amount sufficient to overcome the disease itself.

[0039] The term “treatment,” “treat,” or “treating” refers to a method of reducing the effects of a cancer (e.g., breast cancer, lung cancer, ovarian cancer or the like) or symptom ofcancer. Thus, in the disclosed method, treatment can refer to a 10%, 20%, 30%, 40%, 50%, 60%, 70%), 80%), 90%), or 100% reduction in the severity of a cancer or symptom of the cancer. For example, a method of treating a disease is considered to be a treatment if there is a 10% reduction in one or more symptoms of the disease in a subject as compared to a control. Thus, the reduction can be a 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100% or any percent reduction between 10 and 100% as compared to native or control levels. It is understood that treatment does not necessarily refer to a cure or complete ablation of the disease, condition, or symptoms of the disease or condition.

[0040] Biomarkers

[0041] Fecal occult blood test is a test that checks for occult (hidden) blood in the stool. Blood in the stool may be a sign of colorectal cancer or other gastrointestinal conditions, such as polyps, ulcers, or hemorrhoids. Guaiac fecal occult blood test (gFOBT) and immunochemical fecal occult blood test (FIT) are two types of fecal occult blood tests. Guaiac fecal occult blood test uses a chemical substance called guaiac to check for blood in the stool. Immunochemical fecal occult blood test uses an antibody to check for blood in the stool.

[0042] The present disclosure in one aspect provides an RNA based biomarker panel that includes at least one blood associated mRNA biomarker. In some embodiments, the disclosed biomarker panel combines blood associated mRNA biomarkers and disease-specific biomarkers useful for diagnosing a gastrointestinal disease. The methods utilizing the biomarker panel described herein would require one assay instead of two in order to diagnose the gastrointestinal disease, which simplifies the patient experience and reduces the cost of stool collection and assay reagents. The methods and composition to detect fecal occult blood is of high particular value in mRNA-based methods of colorectal cancer screening. The methods provided herein can be used in an outpatient clinic or inpatient environment.

[0043] In certain embodiments, the method for diagnosing a gastrointestinal disease in a subject disclosed herein comprises: obtaining a ribonucleic acid (RNA) from a stool sample of the subject; measuring messenger RNA (mRNA) levels of a panel of biomarkers in the obtained RNA, said panel of biomarkers comprising at least one blood-specific gene and a group of disease-specific genes; evaluating the measured mRNA levels of the panel of biomarkers; and determining whether the subject is 1) healthy or 2) has the gastrointestinal disease. In certain embodiments, the subject is a human or non-human.

[0044] As used herein, the term “panel of biomarkers” refers to a selection of at least two biomarkers. The panel can comprise from 2 to 20 (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20) or more biomarkers. In certain embodiments, the panel of biomarkersdisclosed herein comprises at least one blood-specific biomarker and a group of diseasespecific biomarkers.

[0045] In certain embodiments, the blood-specific biomarkers used in the method disclosed herein are listed in Table 1. The term “blood-specific biomarker” as used herein refers to a gene that is specifically expressed in at least one type of blood cells. It is understood that some blood-specific biomarkers disclosed herein may not express exclusively in blood cells, but rather expressed in a highly enriched level in blood cells. In certain embodiments, the expression of a group of blood-specific biomarkers are used to detect fecal occult blood. In such embodiments, each of the biomarkers used in detecting fecal occult blood is a bloodspecific biomarker.

[0046] Table 1: Information of Blood-specific Biomarkers

[0047] In certain embodiments, the disease-specific biomarkers are CRC-specific. In certain embodiments, the CRC-specific biomarkers are selected from the group consisting of MAGEA3, DKK4, NOTUM, KLK6, SIM2, TCN1, MYC, LCN2, CDH3, KRT6B, KRT80 and MMP7. The information of certain CRC-specific biomarkers can be found in Tables 2-4. These biomarkers were obtained from the literature and an inhouse computational analysis of publicly available CRC / Normal RNA-seq datasets. The mRNA sequences are referred to Ensemble Number.

[0048] Table 2: CRC-specific Biomarkers

[0049] Table 3: CRC-specific Biomarkers

[0050] Table 4: CRC-specific Biomarkers

[0051] In certain embodiments, the group of CRC-specific biomarkers comprises two or more (e.g., 2, 3, 4, 5, 6 or 7) genes selected from the group consisting of: KRT80, KLK6, MMP7, CDH3, NOTUM, MYC and SIM2.

[0052] In certain embodiments, the group of CRC-specific biomarkers consists of: NOTUM, KLK6, SIM2, CDH3, KRT80 and MMP7. In certain embodiments, the group of CRC-specific biomarkers consists of: MAGEA3, DKK4, NOTUM, KLK6, SIM2, TCN1, MYC, LCN2, CDH3, KRT6B, KRT80 and MMP7.

[0053] Measuring Biomarkers

[0054] The methods of the present disclosure involve detecting or measuring at least one blood-specific biomarker (for example, in Table 1) and a group of gastrointestinal diseasespecific biomarkers disclosed herein (for example, in Table 2-4), in a stool sample obtained from a subject suspected of having or at risk of having a gastrointestinal disease.

[0055] Sample Preparation

[0056] The methods disclosed herein use a stool sample (i.e., a feces sample), which may contain occult blood cells and tumor cells or debris thereof (e.g., cancer cell apoptotic products).

[0057] In certain embodiments, the method comprises a step of isolating ribonucleic acid (RNA) from the sample. Various methods of extraction are suitable for isolating the RNA from cells or tissues, such as phenol and chloroform extraction, and various other methods as described in, for example, Ausubel et al., Current Protocols of Molecular Biology (1997) John Wiley & Sons, and Sambrook and Russell, Molecular Cloning: A Laboratory Manual 3rded. (2001).

[0058] Commercially available kits can also be used to isolate RNA, including for example, the NucleoSpin RNA Stool kit (Takara), Rneasy® mini columns (Qiagen), Stool Total RNA Purification kit (Norgen Biotek), NZY Stool RNA Isolation kit (NZYTech), and PureLink® RNA mini kit (Thermo Fisher Scientific). A skilled person can readily extract or isolate RNA or DNA following the manufacturer’s protocol.

[0059] Methods of Measuring Biomarkers

[0060] The biomarkers disclosed herein can be detected in the level of RNA (e.g., mRNA) using proper methods known in the art including, without limitation, amplification assay, hybridization assay, and sequencing assay. In certain embodiments, the mRNA expression levels are detected using quantitative RT-PCR or Droplet Digital PCR or Partition PCR. In certain embodiments, the levels of the panel of biomarkers are determined by nucleic acid sequencing.

[0061] Sequencing methods

[0062] Sequencing methods useful in the measurement of the biomarkers involve sequencing of the target nucleic acid. Any sequencing known in the art can be used to detect the biomarkers of interest. In general, sequencing methods can be categorized to traditional or classical methods and high throughput sequencing (next generation sequencing). Traditional sequencing methods include Maxam-Gilbert sequencing (also known as chemical sequencing) and Sanger sequencing (also known as chain-termination methods).

[0063] High throughput sequencing, or next generation sequencing, by using methods distinguished from traditional methods, such as Sanger sequencing, is highly scalable and able to sequence the entire genome or transcriptome at once. High throughput sequencing involves sequencing-by-synthesis, sequencing-by-ligation, and ultra-deep sequencing (such as described in Marguiles et al., Nature 437 (7057): 376-80 (2005)). Sequence-by-synthesis involves synthesizing a complementary strand of the target nucleic acid by incorporating labeled nucleotide or nucleotide analog in a polymerase amplification. Immediately after or upon successful incorporation of a label nucleotide, a signal of the label is measured, and the identity of the nucleotide is recorded. The detectable label on the incorporated nucleotide is removed before the incorporation, detection and identification steps are repeated. Examples of sequence-by-synthesis methods are known in the art, and are described for example in U.S. Pat. No. 7,056,676, U.S. Pat. No. 8,802,368 and U.S. Pat. No. 7,169,560, the contents of which are incorporated herein by reference. Sequencing-by-synthesis may be performed on a solid surface (or a microarray or a chip) using fold-back PCR and anchored primers. Target nucleic acid fragments can be attached to the solid surface by hybridizing to the anchored primers, and bridge amplified. This technology is used, for example, in the Illumina® sequencing platform.

[0064] Pyrosequencing involves hybridizing the target nucleic acid regions to a primer and extending the new strand by sequentially incorporating deoxynucleotide triphosphates corresponding to the bases A, C, G, and T (U) in the presence of a polymerase. Each base incorporation is accompanied by release of pyrophosphate, converted to ATP by sulfurylase,which drives synthesis of oxyluciferin and the release of visible light. Since pyrophosphate release is equimolar with the number of incorporated bases, the light given off is proportional to the number of nucleotides adding in any one step. The process is repeated until the entire sequence is determined.

[0065] In certain embodiments, the biomarkers described herein are detected by whole transcriptome shotgun sequencing (RNA sequencing). The method of RNA sequencing has been described (see Wang Z, Gerstein M and Snyder M, Nature Review Genetics (2009) 10:57- 63; Maher CA et al., Nature (2009) 458:97-101; Kukurba K & Montgomery SB, Cold Spring Harbor Protocols (2015) 2015(11): 951-969).

[0066] Amplification assay

[0067] A nucleic acid amplification assay involves copying a target nucleic acid (e.g., DNA or RNA), thereby increasing the number of copies of the amplified nucleic acid sequence. Amplification may be exponential or linear. Exemplary nucleic acid amplification methods include, but are not limited to, amplification using the polymerase chain reaction ("PCR", see U.S. Patents 4,683,195 and 4,683,202; PCR Protocols: A Guide To Methods And Applications (Innis et al., eds, 1990)), reverse transcriptase polymerase chain reaction (RT-PCR), quantitative real-time PCR (qRT-PCR); quantitative PCR, such as TaqMan®, nested PCR, Digital PCR, such as Droplet Digital PCR (ddPCR), ligase chain reaction (See Abravaya, K., et al., Nucleic Acids Research, 23:675-682, (1995), branched DNA signal amplification (see, Urdea, M. S., et al., AIDS, 7 (suppl 2):S11-S14, (1993), amplifiable RNA reporters, Q-beta replication (see Lizardi et al., Biotechnology (1988) 6: 1197), transcription-based amplification (see, Kwoh et al., Proc. Natl. Acad. Sci. USA (1989) 86: 1173-1177), boomerang DNA amplification, strand displacement activation, cycling probe technology, self-sustained sequence replication (Guatelli et al., Proc. Natl. Acad. Sci. USA (1990) 87: 1874-1878), rolling circle replication (U.S. Patent No. 5,854,033), isothermal nucleic acid sequence based amplification (NASBA), and serial analysis of gene expression (SAGE). Droplet Digital PCR (ddPCR) is a method for performing digital PCR that is based on water-oil emulsion droplet technology. A sample is fractionated into many droplets, and PCR amplification of the template molecules occurs in each individual droplet. ddPCR technology uses reagents and workflows similar to those used for most standard TaqMan probe-based assays. The massive sample partitioning is a key aspect of the ddPCR technique.

[0068] In certain embodiments, the nucleic acid amplification assay is a PCR-based method. PCR is initiated with a pair of primers that hybridize to the target nucleic acid sequence to be amplified, followed by elongation of the primer by polymerase which synthesizes thenew strand using the target nucleic acid sequence as a template and dNTPs as building blocks. Then the new strand and the target strand are denatured to allow primers to bind for the next cycle of extension and synthesis. After multiple amplification cycles, the total number of copies of the target nucleic acid sequence can increase exponentially.

[0069] In certain embodiments, intercalating agents that produce a signal when intercalated in double stranded DNA may be used. Exemplary agents include SYBR GREEN™ and SYBR GOLD™. Since these agents are not template-specific, it is assumed that the signal is generated based on template-specific amplification. This can be confirmed by monitoring signal as a function of temperature because melting point of template sequences will generally be much higher than, for example, primer-dimers, etc.

[0070] In certain embodiments, a detectably labeled primer or a detectably labeled probe can be used, to allow detection of the biomarkers corresponding to that primer or probe. In certain embodiments, multiple labeled primers or labeled probes with different detectable labels can be used to allow simultaneous detection of multiple biomarkers.

[0071] Hybridization assay

[0072] Nucleic acid hybridization assays use probes to hybridize to the target nucleic acid, thereby allowing detection of the target nucleic acid. Non-limiting examples of hybridization assay include Northern blotting, Southern blotting, in situ hybridization, microarray analysis, and multiplexed hybridization-based assays.

[0073] In certain embodiments, the probes for hybridization assay are detectably labeled. In certain embodiments, the nucleic acid-based probes for hybridization assay are unlabeled. Such unlabeled probes can be immobilized on a solid support such as a microarray and can hybridize to the target nucleic acid molecules which are detectably labeled.

[0074] In certain embodiments, hybridization assays can be performed by isolating the nucleic acids (e.g., RNA or DNA), separating the nucleic acids (e.g., by gel electrophoresis) followed by transfer of the separated nucleic acid on suitable membrane filters (e.g., nitrocellulose filters), where the probes hybridize to the target nucleic acids and allows detection. See, for example, Molecular Cloning: A Laboratory Manual, J. Sambrook et al., eds., 2nd edition, Cold Spring Harbor Laboratory Press, 1989, Chapter 7. The hybridization of the probe and the target nucleic acid can be detected or measured by methods known in the art. For example, autoradiographic detection of hybridization can be performed by exposing hybridized filters to photographic film.

[0075] In some embodiments, hybridization assays can be performed on microarrays. Microarrays provide a method for the simultaneous measurement of the levels of large numbersof target nucleic acid molecules. The target nucleic acids can be RNA, DNA, cDNA reverse transcribed from mRNA, or chromosomal DNA. The target nucleic acids can be allowed to hybridize to a microarray comprising a substrate having multiple immobilized nucleic acid probes arrayed at a density of up to several million probes per square centimeter of the substrate surface. The RNA or DNA in the sample is hybridized to complementary probes on the array and then detected by laser scanning. Hybridization intensities for each probe on the array are determined and converted to a quantitative value representing relative levels of the RNA or DNA. See, U.S. Patent Nos. 6,040,138, 5,800,992 and 6,020,135, 6,033,860, and 6,344,316.

[0076] Techniques for the synthesis of these arrays using mechanical synthesis methods are described in, e.g., U.S. Patent No. 5,384,261. Although a planar array surface is often employed the array may be fabricated on a surface of virtually any shape or even a multiplicity of surfaces. Arrays may be peptides or nucleic acids on beads, gels, polymeric surfaces, fibers such as fiber optics, glass or any other appropriate substrate, see U.S. Patent Nos. 5,770,358, 5,789,162, 5,708,153, 6,040,193 and 5,800,992. Arrays may be packaged in such a manner as to allow for diagnostics or other manipulation of an all-inclusive device. Useful microarrays are also commercially available, for example, microarrays from Affymetrix, from Nano String Technologies, QuantiGene 2.0 Multiplex Assay from Panomics.

[0077] In certain embodiments, hybridization assays can be in situ hybridization assay. In situ hybridization assay is useful to detect the presence of gene mutations. Probes useful for in situ hybridization assay can be mutation specific probes, which hybridize to a specific gene mutation to detect the presence or absence of the specific mutation of interest. Methods for use of unique sequence probes for in situ hybridization are described in U.S. Pat. No. 5,447,841, incorporated herein by reference. Probes can be viewed with a fluorescence microscope and an appropriate filter for each fluorophore, or by using dual or triple band-pass filter sets to observe multiple fhiorophores. See, e.g., U.S. Pat. No. 5,776,688 to Bittner, et al., which is incorporated herein by reference. Any suitable microscopic imaging method can be used to visualize the hybridized probes, including automated digital imaging systems. Alternatively, techniques such as flow cytometry can be used to examine the hybridization pattern of the probes.

[0078] Any of the assays and methods provided herein for the measurement of the gene expression level can be adapted or optimized for use in automated and semi-automated systems or point of care assay systems.

[0079] The gene expression level described herein can be normalized using a proper method known in the art. For example, the gene expression level can be normalized to astandard level of a standard marker, which can be predetermined, determined concurrently, or determined after a sample is obtained from the subject. The standard marker can be run in the same assay or can be a known standard marker from a previous assay. For another example, the gene expression level can be normalized to an internal control which can be an internal marker, or an average level or a total level of a plurality of internal markers.

[0080] The level of mRNA expression of each of the biomarkers described herein can be normalized to a reference level for a control gene. The control value can be predetermined, determined concurrently, or determined after a sample is obtained from the subject. The standard can be run in the same assay or can be a known standard from a previous assay. In the cases when the level of RNA expression is determined by RNA sequencing, the level of RNA expression of each of the biomarkers can be normalized to the total reads of the sequencing. The normalized levels of mRNA expression of the biomarker genes can be transformed into a score, e.g., using the methods and models described herein.

[0081] Methods for Diagnosing Gastrointestinal Disease

[0082] In some embodiments, the method disclosed herein comprises classifying the subject as 1) healthy or 2) having a gastrointestinal disease based on the measured mRNA level of the blood-specific biomarker. In some specification, the method disclosed herein is based on the measured mRNA levels of the biomarker panel described herein. In some embodiments, the method comprises evaluating the measured levels of the biomarker panel by a machine learning classifier and determining that the subject is 1) healthy or 2) has the gastrointestinal disease.

[0083] In statistics, classification is the problem of identifying which of a set of categories an observation (or observations) belongs to. As used herein, the term “classification” used interchangeably with the term “classifying” refers to the identification of the subject as 1) being healthy or 2) having colorectal cancer or precancerous colorectal adenomas based on the measured levels of the biomarker panel. A “classifier” refers to an algorithm that implements the classification.

[0084] As used herein, the term “machine learning” refers to a computer-implemented technique that gives computer systems the ability to progressively improve performance on a specific task with data, i.e., to learn from the data, without being explicitly programmed. Machine learning technique adopts algorithms that can learn from and make prediction on data through building a model, i.e., a description of a system using mathematical concepts, from sample inputs. A core objective of machine learning is to generalize from the experience, i.e., to perform accurately on new data after having experienced a learning data set.

[0085] Machine learning models can be categorized as either supervised or unsupervised. Supervised learning involves learning a function that maps an input to an output based on example input-output pairs. In the context of biomedical diagnosis or prognosis, machine learning techniques generally involves supervised learning process, in which the computer is presented with example inputs (e.g., signature of gene expression) and their desired outputs (e.g., likelihood of having colorectal cancer or precancerous colorectal adenomas) to learn a general rule that maps inputs to outputs. Different models, i.e., hypothesis, can be employed in the generalization process. For the best performance in the generalization, the complexity of the hypothesis should match the complexity of the function underlying the data.

[0086] Supervised models include but not limited to logistic regression, support vector machine, decision trees, random forest, artificial neural network, linear regression, elastic net and naive bayes. In certain embodiments, the machine learning classifier is random forest or elastic net.

[0087] Logistic Regression

[0088] The simplest idea of linear regression is to find a line that best fits the data. Extensions of linear regression include multiple linear regression (e.g., finding a plane of best fit) and polynomial regression (e.g., finding a curve of best fit). Logistic regression is similar to linear regression but is used to model the probability of a finite number of outcomes, typically two.

[0089] Support Vector Machine

[0090] A Support Vector Machine (SVM) is a supervised classification technique that, at the most fundamental level, find a hyperplane or a boundary between two classes of data that maximizes the margin between the two classes. There are many planes that can separate the two classes, but only one plane can maximize the margin or distance between the classes.

[0091] Decision Tree

[0092] A decision tree is a decision support tool that uses a tree-like model of decisions and their possible consequences, including chance event outcomes, resource costs, and utility. Typically, a decision tree is a flowchart-like structure in which each internal node represents a “test” on an attribute, each branch represents the outcome of the test, and each leaf node represents a class label (decision taken after computing all attribute). The paths from root to leaf represent classification rules. Decision trees are intuitive and easy to build but fall short when it comes to accuracy.

[0093] Random Forest

[0094] Random forests are an ensemble learning technique that builds off of decision trees. Random forests involve creating multiple decision trees using bootstrapped datasets of the original data and randomly selecting a subset of variables at each step of the decision tree. The model then selects the mode of all of the predictions of each decision tree. By relying on a “majority wins” model, it reduces the risk of error from an individual tree.

[0095] Artificial Neural Network

[0096] Artificial neural networks (ANNs), usually simply called neural networks (NNs), are inspired by the biological neural networks that constitute animal brains. An ANN is based on a collection of connected units or nodes called artificial neurons, which loosely model the neurons in a biological brain. Each connection, like the synapses in a biological brain, can transmit a signal to other neurons. An artificial neuron receives a signal then processes it and can signal neurons connected to it. The “signal” at a connection is a real number, and the output of each neuron is computed by some non-linear function of the sum of its inputs. The connections are called edges. Neurons and edges typically have a weight that adjusts as learning proceeds. The weight increases or decreases the strength of the signal at a connection. Neurons may have a threshold such that a signal is sent only if the aggregate signal crosses that threshold. Typically, neurons are aggregated into layers. Different layers may perform different transformations on their inputs. Signals travel from the first layer (the input layer) to the last layer (the output layer), possibly after traversing the layers multiple times.

[0097] Elastic net

[0098] Elastic net is a regularized regression method that linearly combines the Li and L2penalties of the LASSO (least absolute shrinkage and selection operator) and ridge methods. The LASSO method constructs a linear model, which penalizes the regression coefficients with an Li penalty, shrinking many of them to zero. Any features which have non-zero regression coefficients are “selected” by the lasso algorithm. As a result, the LASSO method performs both feature selection and regularization in order to enhance the prediction accuracy and interpretability of the resulting machine learning model. Elastic net is one of the improvements to the LASSO. Elastic net regularization combines the Li penalty of LASSO with the L2 penalty of ridge regression.

[0099] Unsupervised learning, on the other hand, is to draw inferences and find patterns from input data without references to labeled outcomes. Two main methods used in unsupervised learning include clustering and dimensionality reduction.

[0100] Clustering is an unsupervised technique that involves the grouping or clustering of data points. Common clustering algorithm include k-means clustering, hierarchical clustering, mean shift clustering, and density -based clustering.

[0101] Dimensionality reduction is a process of reducing the number of random variables to obtain a set of principle variables. Common dimensionality reduction algorithm include principal component analysis (PCA), regularized regression and Boruta.

[0102] The use of the machine learning classifier can classify the sample as 1) a healthy or 2) having gastrointestinal disease with a sensitivity, specificity, positive predictive value, negative predictive value, and / or overall accuracy of at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%.

[0103] In certain embodiments, the methods of the present disclosure further comprise sending the 1) a healthy or 2) having gastrointestinal disease classification results to a clinician. In some embodiments, the method of the present disclosure provides a diagnosis or prognosis in the form of a probability that the subject 1) is healthy or 2) having gastrointestinal disease. For example, the subject may have about a 0%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95% or greater probability of 1) being healthy or 2) having gastrointestinal disease.

[0104] Computer-implemented Methods, Systems and Devices

[0105] Any of the methods described herein may be totally or partially performed with a computer system including one or more processors, which can be configured to perform the steps. Thus, embodiments are directed to computer systems configured to perform the steps of any of the methods described herein, potentially with different components performing a respective step or a respective group of steps. Although presented as numbered steps, steps of methods herein can be performed at a same time or in a different order. Additionally, portions of these steps may be used with portions of other steps from other methods. Also, all or portions of a step may be optional. Any of the steps of any of the methods can be performed with modules, circuits, or other means for performing these steps.

[0106] Any of the computer systems mentioned herein may utilize any suitable number of subsystems. In some embodiments, a computer system includes a single computer apparatus, where the subsystems can be the components of the computer apparatus. In other embodiments, a computer system can include multiple computer apparatuses, each being a subsystem, with internal components. The subsystems can be interconnected via a system bus. Additional subsystems include, for examples, a printer, keyboard, storage device(s), monitor, which iscoupled to display adapter, and others. Peripherals and input / output (I / O) devices, which couple to I / O controller, can be connected to the computer system by any number of means known in the art, such as serial port. For example, serial port or external interface (e.g. Ethernet, Wi-Fi, etc.) can be used to connect computer system to a wide area network such as the Internet, a mouse input device, or a scanner. The interconnection via system bus allows the central processor to communicate with each subsystem and to control the execution of instructions from system memory or the storage device(s) (e.g., a fixed disk, such as a hard drive or optical disk), as well as the exchange of information between subsystems. The system memory and / or the storage device(s) may embody a computer readable medium. Any of the data mentioned herein can be output from one component to another component and can be output to the user.

[0107] A computer system can include a plurality of the same components or subsystems, e.g., connected together by external interface or by an internal interface. In some embodiments, computer systems, subsystem, or apparatuses can communicate over a network. In such instances, one computer can be considered a client and another computer a server, where each can be part of a same computer system. A client and a server can each include multiple systems, subsystems, or components.

[0108] It should be understood that any of the embodiments of the present disclosure can be implemented in the form of control logic using hardware (e.g., an application specific integrated circuit or field programmable gate array) and / or using computer software with a generally programmable processor in a modular or integrated manner. As used herein, a processor includes a multi-core processor on a same integrated chip, or multiple processing units on a single circuit board or networked. Based on the disclosure and teachings provided herein, a person of ordinary skill in the art will know and appreciate other ways and / or methods to implement embodiments of the present disclosure using hardware and a combination of hardware and software.

[0109] Any of the software components or functions described in this application may be implemented as software code to be executed by a processor using any suitable computer language such as, for example, Java, C++ or Perl using, for example, conventional or object- oriented techniques. The software code may be stored as a series of instructions or commands on a computer readable medium for storage and / or transmission, suitable media include random access memory (RAM), a read only memory (ROM), a magnetic medium such as a hard-drive or a floppy disk, or an optical medium such as a compact disk (CD) or DVD (digital versatile disk), flash memory, and the like. The computer readable medium may be any combination of such storage or transmission devices.

[0110] Such programs may also be encoded and transmitted using carrier signals adapted for transmission via wired, optical, and / or wireless networks conforming to a variety of protocols, including the Internet. As such, a computer readable medium according to an embodiment of the present invention may be created using a data signal encoded with such programs. Computer readable media encoded with the program code may be packaged with a compatible device or provided separately from other devices (e.g., via Internet download). Any such computer readable medium may reside on or within a single computer product (e.g. a hard drive, a CD, or an entire computer system), and may be present on or within different computer products within a system or network. A computer system may include a monitor, printer, or other suitable display for providing any of the results mentioned herein to a user.

[0111] Kits

[0112] In another aspect, the present disclosure provides kits or an integrated system for use in the methods described above. The kits may comprise any or all of the reagents to perform the methods described herein. In certain embodiments, the kit comprises primers for detecting the nucleic acids specific to the panel of biomarkers in a sample.

[0113] “Primer” as used herein refers to an oligonucleotide molecule with a length of 7-40 nucleotides, preferably 10-38 nucleotides, preferably 15-30 nucleotides, or 15-25 nucleotides, or 17-20 nucleotides. For example, the primer can an oligonucleotide having a length of 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29 or 30 nucleotides. Primers are used in the amplification of a DNA sequence by polymerase chain reaction (PCR) as well known in the art. For a DNA template sequence to be amplified, a pair of primers can be designed at its 5’ upstream and its 3’ downstream sequence, i.e., 5’ primer and 3’ primer, each of which can specifically hybridize to a separate strand of the DNA double strand template. 5’ primer is complementary to the anti-sense strand of the DNA double strand template; and 3’ primer is complementary to the sense strand of the DNA template. As known in the art, the “sense strand” of a double stranded DNA template is the strand which contains the sequence identical to the mRNA sequence transcribed from the DNA template (except that “U” in RNA corresponds to “T” in the DNA) and encoding for a protein product. The complementary sequence of the sense strand is the “anti-sense strand.”

[0114] In certain embodiments, the kit further comprises an agent for amplifying the target nucleic acid using the primers. In addition, the kits may include instructional materials containing directions (i.e., protocols) for the practice of the methods provided herein. While the instructional materials typically comprise written or printed materials, they are not limitedto such. Any medium capable of storing such instructions and communicating them to an end user is contemplated by this disclosure. Such media include but are not limited to electronic storage media (e.g., magnetic discs, tapes, cartridges, chips), optical media (e.g., CD ROM), and the like. Such media may include addresses to internet sites that provide such instructional materials.

[0115] In another aspect, the present disclosure provides oligonucleotide probes for detecting the nucleic acids specific to the panel of biomarkers in a sample. In certain embodiments, the probes are attached to a solid support, such as an array slide or chip, e.g., as described in Eds., Bowtell and Sambrook DNA Microarrays: A Molecular Cloning Manual (2003) Cold Spring Harbor Laboratory Press. Construction of such devices are well known in the art, for example as described in US Patents and Patent Publications U.S. Patent No. 5,837,832; PCT application W095 / 11995; U.S. Patent No. 5,807,522; US Patent Nos. 7,157,229, 7,083,975, 6,444,175, 6,375,903, 6,315,958, 6,295,153, and 5,143,854, 2007 / 0037274, 2007 / 0140906, 2004 / 0126757, 2004 / 0110212, 2004 / 0110211, 2003 / 0143550, 2003 / 0003032, and 2002 / 0041420. Nucleic acid arrays are also reviewed in the following references: Biotechnol Annu Rev (2002) 8:85-101; Sosnowski et al. Psychiatr Genet (2002)12(4): 181-92; Heller, Annu Rev Biomed Eng (2002) 4: 129-53; Koi chinsky et al., Hum. Mutat (2002) 19(4):343-60; and McGail et al., Adv Biochem Eng Biotechnol (2002) 77:21-42.

[0116] A microarray can be composed of a large number of unique, single-stranded polynucleotides, usually either synthetic antisense polynucleotides or fragments of cDNAs, fixed to a solid support. Typical polynucleotides are preferably about 6-60 nucleotides in length, more preferably about 15-30 nucleotides in length, and most preferably about 18-25 nucleotides in length. For certain types of arrays or other detection kits / systems, it may be preferable to use oligonucleotides that are only about 7-20 nucleotides in length. In other types of arrays, such as arrays used in conjunction with chemiluminescent detection technology, preferred probe lengths can be, for example, about 15-80 nucleotides in length, preferably about 50-70 nucleotides in length, more preferably about 55-65 nucleotides in length, and most preferably about 60 nucleotides in length.

[0117] Techniques for the synthesis of these arrays using mechanical synthesis methods are described in, e.g., U.S. Patent No. 5,384,261. Although a planar array surface is often employed the array may be fabricated on a surface of virtually any shape or even a multiplicity of surfaces. Arrays may also be nucleic acids on beads, gels, polymeric surfaces, fibers such as fiber optics, glass or any other appropriate substrate, see U.S. Patent Nos. 5,770,358, 5,789,162,5,708,153, 6,040,193 and 5,800,992. Arrays may be packaged in such a manner as to allow for diagnostics or other manipulation of an all-inclusive device.

[0118] The probes and primers necessary for practicing the present disclosure can be synthesized and labeled using well known techniques. Oligonucleotides used as probes and primers may be chemically synthesized according to the solid phase phosphoramidite triester method first described by Beaucage and Caruthers, Tetrahedron Letts. (1981) 22: 1859-1862, using an automated synthesizer, as described in Needham-Van Devanter et al, Nucleic Acids Res. (1984) 12:6159-6168.

[0119] Methods for Treating Gastrointestinal Disease

[0120] In yet another aspect, the present disclosure provides a method for treating a gastrointestinal disease in a subject. In some embodiments, the method comprises administering to the subject a therapeutically effective amount of a drug useful for treating the gastrointestinal disease, wherein the subject has been determined to have the gastrointestinal disease by a machine learning classifier based on mRNA levels of a panel of biomarkers. In some embodiments, the panel of biomarkers comprises at least one blood-specific gene. In some embodiments, the panel of biomarkers comprises at least one blood-specific gene and a group of disease-specific genes.

[0121] In some embodiment, the drug used in the method of treating CRC as disclosed herein include, without limitation: Alymsys® (Bevacizumab), Avastin® (Bevacizumab), Camptosar® (Irinotecan Hydrochloride), Capecitabine, Cetuximab, Cyramza® (Ramucirumab), Eloxatin® (Oxaliplatin), Erbitux® (Cetuximab), 5-FU (Fluorouracil Injection), Fluorouracil Injection, Ipilimumab, Irinotecan Hydrochloride, Keytruda® (Pembrolizumab), Leucovorin Calcium, Lonsurf® (Trifluridine and Tipiracil Hydrochloride), Mvasi® (Bevacizumab), Opdivo® (Nivolumab), Oxaliplatin, Panitumumab, Pembrolizumab, Ramucirumab, Regorafenib, Stivarga® (Regorafenib), Trifluridine and Tipiracil Hydrochloride, Vectibix® (Panitumumab), Xeloda® (Capecitabine), Yervoy® (Ipilimumab), Zaltrap® (Ziv-Aflibercept), Zirabev® (Bevacizumab), Ziv-Aflibercept.

[0122] The drug described herein may be administered in any desired and effective manner: for oral ingestion, or as an ointment or drop for local administration to the eyes, or for parenteral or other administration in any appropriate manner such as intraperitoneal, subcutaneous, topical, intradermal, inhalation, intrapulmonary, rectal, vaginal, sublingual, intramuscular, intravenous, intraarterial, intrathecal, or intralymphatic. Further, the drug may be administered in conjunction with other treatments.

[0123] The following examples are provided to better illustrate the claimed disclosure and are not to be interpreted as limiting the scope of the disclosure. All specific compositions, materials, and methods described below, in whole or in part, fall within the scope of the present disclosure. These specific compositions, materials, and methods are not intended to limit the disclosure, but merely to illustrate specific embodiments falling within the scope of the disclosure. One skilled in the art may develop equivalent compositions, materials, and methods without the exercise of inventive capacity and without departing from the scope of the disclosure. It will be understood that many variations can be made in the procedures herein described while still remaining within the bounds of the present disclosure. It is the intention of the inventors that such variations are included within the scope of the disclosure.EXAMPLE 1

[0124] This example shows the identification and validation of biomarkers useful for a mRNA-based fecal occult blood test. First, putative biomarkers were identified using a combination of computational and biological analysis. Second, a selection of blood biomarkers was validated using RT-qPCR in RNA extracted from stool collected from patients diagnosed with CRC / APL or patients with a normal colonoscopy.

[0125] To identify genes that could be used in occult blood detection, a computational analysis was performed to identify highly expressed genes in blood. Whole blood tissue expression data was downloaded from the Genotype-Tissue Expression (GTEx) database and the read counts were converted to transcripts per million (TPM). The genes were then sorted by their expression level from highest to lowest expression.

[0126] Many of the highly expressed genes in blood are also expressed in colon tissue. Genes that are expressed in both colon tissue and whole blood are not specific biomarkers for bleeding in the gut. To filter out genes also expressed in colon tissue, GTEx expression data for colon tissue was analyzed to remove genes expressed in colon tissue. The top 8 list of genes expressed in blood after filtering out colon expressed genes are: HBB, HBA2, HBD, HBA1, S100A9, CSF3R, IFITM2 and LCP1. The top four most expressed genes in whole blood with little detectable expression in colon tissue are hemoglobin subunit genes.

[0127] In addition, we identified a list of likely red blood cell specific gene markers per established blood biology. We obtained a list of hemoglobin subunit genes by searching the HUGO Gene Nomenclature Committee (HGNC) database16using the search term “hemoglobin subunits” and identified HBE1, HBG1, HBG2, HBM, HBQ1 and HBZ hemoglobin subunits as additional biomarkers.

[0128] We also include genes that code for blood group antigens. These are cell surface markers found primarily but not exclusively on red blood cells. We identified these genes from the HGNC database using the search term “blood group antigen”. Finally, we obtained a list of mRNA transcripts present in platelets by searching the HUGO database using the key word “platelets”. The identified blood-specific genes are listed in Table 1.

[0129] An internal validation study was then performed testing the ability of HBA1, HBA2 and HBB to detect occult blood in stool and comparing the performance in CRC / APL detection with the current gold standard FIT method. The primary outcome of this study is the sensitivity and specificity of detection for cancerous, and advanced precancerous lesions, using HBA1 and HBA2 and HBB as biomarkers. The secondary outcome is comparing performance of the mRNA biomarkers with FIT.

[0130] HBA1, HBA2 and HBB were selected for testing on clinical samples because these genes had the highest expression in whole blood based on our analysis of RNA-seq whole blood samples. These genes also code for hemoglobin subunits and hemoglobin protein is the primary target for most existing FIT assays. HBA1 and HBA2 have very high sequence similarity so primers and a probe were designed to recognize either HBA1 or HBA2 sequence. Recognizing two blood associated genes with a single TaqMan assay increases the sensitivity of detection without increasing the cost or complexity.

[0131] RNA was extracted from stool collected from 90 patients, 30 CRC, 30 APL and 30 samples taken from patients with a normal colonoscopy. Advanced precancerous lesions were defined using generally accepted criteria i.e. greater than 1 cm in size in any dimension, or a villous component greater than 25% or the presence of high-grade dysplasia.

[0132] The methodology used to extract RNA from stool is as follows. RNA was extracted from 0.25 g of stool using the Omega Bio-Tek E.Z.N. A.® Stool RNA Kit (R6828), which employs a phenol-chloroform extraction method. To remove potential PCR inhibitors, the Zymo Research OneStep™ PCR Inhibitor Removal Kit (D6035) was used following the RNA extraction. The extracted RNA was resuspended in 50 pL of DEPC-treated water per tube. RNA concentration and purity were measured using a NanoDrop spectrophotometer. The typical RNA yield was approximately 450 pg per sample. The RNA was then aliquoted and stored at -80°C until further use.

[0133] After RNA was extracted from stool the amount of HBA1, HBA2 and HBB and the 13 CRC-specific genes, transcripts were quantified using RT-qPCR using TaqMan assays. Since the HBA1 and HBA2 TaqMan assay was designed to recognize either HBA1 or HBA2, henceforth we will refer to this assay as HBA collectively. The PCR protocol usedwas as follows: Takara One Step PrimeScript™ III RT-qPCR Mix with UNG (RR601 A) was utilized for qRT-PCR. Each reaction was set up with a total volume of 10 pL, including 2 pL of extracted RNA. The reactions were performed on an ABI 7500 Real-Time PCR System with the following thermal cycling conditions: reverse transcription at 50°C for 5 minutes, initial denaturation at 95°C for 10 seconds, followed by 40 cycles of 95°C for 10 seconds and 60°C for 34 seconds.

[0134] As a comparison, a fecal immunochemical test (FIT) was simultaneously performed for each sample as RNA extraction using the Luminex platform from Tellgen Corporation, China, employing flow cytometry to quantitatively detect hemoglobin (HB) and transferrin (TF) protein in human fecal sample suspensions.

[0135] The primers and probes used in the TaqMan assays were designed in-house to recognize spliced mRNA and not genomic DNA. This was done by having either a primer or a probe span a splice junction. The primer and probe sequences are given in Table 5. It is noted that either HBA1 & HBA2 probe or HBB probe includes a FAM moiety at the 5’ end of the DNA sequence, and a MGB moiety at the 3’ end of the DNA sequence, respectively.

[0136] Table 5 : Sequence of Primers and Probes

[0137] FIT performance was 76.7% sensitivity for CRC detection, 26.7% sensitivity for APL detection and 93.3% specificity; these results are in close agreement with published FIT performance3. To directly compare the performance of HBA / HBB with FIT, we calculated the sensitivity for CRC and AA detection at the same specificity observed for FIT. By setting the specificity to be the same as FIT for HBA and HBB we can directly and fairly compare the sensitivities across all biomarkers at the same specificity. Both HBA and HBB performed the same or better than FIT (Table 6). HBA had better CRC and AA sensitivity compared to FIT while HBB had the same CRC sensitivity but better AA sensitivity compared to FIT.

[0138] There was a strong degree of overlap in samples that were high for both FIT and the mRNA-based measures of occult blood (HBA, HBB). Of the 23 out of 30 CRC samplesthat were positive for FIT, 21 of them were also positive for HBA and 20 were positive for HBB assuming a specificity cutoff of 93.3%. Of the 8 out of 30 APL samples that were positive for FIT, 5 of them were positive for both HBA and HBB.

[0139] Table 6: Performance of stool mRNA occult blood biomarkers compared with FIT

[0140] Both HBA and HBB had higher levels in the stool of CRC and APL patients compared to normal controls. The mean Ct value for HBA in CRC stool samples was 29.3 compared to a mean Ct value in normal stool of 36.3, which represents -130 fold higher levels observed in CRC stool compared to controls. The mean Ct value for HBB in CRC stool samples was 29 compared to a mean Ct value in normal stool of 36.6, which represents -190 fold higher levels observed in CRC stool compared to controls (FIG. 1). The mean Ct value in APL samples for HBB was 33.9 for HBA, the mean Ct value in APL samples was 33.8, representing a fold enrichment in APL compared to normal stool of -6.5 and -5.7 respectively. The correlation between HBA and HBB across all samples was very strong with a Pearson correlation coefficient of 0.83, p-value le-16.

[0141] To test if our blood-specific mRNA biomarkers can also be used as a replacement for FIT when combined with a panel of CRC-specific biomarkers, we extracted RNA from 130 stool samples (48 CRC stool samples, 22 APL stool samples and 60 normal stool samples) and performed RT-qPCR to measure the abundance of 13 CRC-specific genes ETV4, SPTBN2, IFITM1, TESC, OLR1, IPO4, RHPN1, PPBP, MYC, TIMP1, CXCL8, TGFBI, MMP7 (as described in Table 2-4) and HBA mRNA. We also performed FIT tests on all these stool samples. We calculated the performance combining HBA with this panel of 13 CRC-specific biomarkers compared to the performance of FIT combined with the same panel of 13 biomarkers (Table 7). Performance was calculated using a random forests classifier and five-fold cross validation using scikit-learn module in Python.

[0142] The combination of HBA and CRC-associated biomarkers performed as well as the combination of FIT and CRC-associated biomarkers. At the same specificity of 91.7%, the HBA plus the CRC gene panel combination had a sensitivity of 93.8%, which is a slightlybetter compared to the FIT and CRC gene panel combination with a sensitivity of 91.7%. The AA sensitivity of HBA plus CRC gene panel had the same APL sensitivity as the FIT plus the CTC gene panel.

[0143] Table 7: Performance combining HBA or FIT with CRC gene panel

[0144] Our results demonstrate that using mRNA level of blood-specific biomarkers had similar performance in detecting occult blood in stool and screening for cancerous and advanced precancerous lesions as the current gold standard of FIT protein based occult blood detection, both individually and combined with other stool biomarkers. Current multi-target stool CRC screening methods pair nucleic acid-based assays with FIT, which increases cost and difficulty of use. With our nucleic acid-based approach to detect occult blood in stool, we can pair the signal from occult blood with other nucleic acid-based biomarkers without requiring two different assays, workflows, or methods of stool collection. This represents a considerable savings in cost and complexity and is a breakthrough in the field. While we tested 3 of our blood associated genes on clinical stool samples, it is expected the other blood associated genes provided in this disclosure will also perform similarly, as all blood associated genes were selected on the basis of the same criteria.

[0145] While the disclosure has been particularly shown and described with reference to specific embodiments (some of which are preferred embodiments), it should be understood by those having skill in the art that various changes in form and detail may be made therein without departing from the spirit and scope of the present disclosure as disclosed herein.References1. Cameron, A. D. Gastro-intestinal blood loss measured by radioactive chromium. Gut 1, 177-182 (1960).2. Ebaugh, F. G., Clemens, T., Rodnan, G. & Peterson, R. E. Quantitative measurement of gastrointestinal blood loss. I. The use of radioactive Cr51 in patients with gastrointestinal hemorrhage. Am. J. Med. 25, 169-181 (1958).3. Medical Advisory Secretariat. Fecal occult blood test for colorectal cancer screening: an evidence-based analysis. Ont. Health TechnoL Assess. Ser. 9, 1-40 (2009).4. Lee, M. W., Pourmorady, J. S. & Laine, L. Use of Fecal Occult Blood Testing as a Diagnostic Tool for Clinical Indications: A Systematic Review and Meta- Analysis. Am.J. Gastroenterol. 115, 662-670 (2020).5. Medical Advisory Secretariat. Optical coherence tomography for age-related macular degeneration and diabetic macular edema: an evidence-based analysis. Ont. Health TechnoL Assess. Ser. 9, 1-22 (2009).6. Navarro, M., Nicolas, A., Ferrandez, A. & Lanas, A. Colorectal cancer population screening programs worldwide in 2016: An update. World J. Gastroenterol. 23, 3632- 3642 (2017).7. Tinmouth, J., Lansdorp-Vogelaar, I. & Allison, J. E. Faecal immunochemical tests versus guaiac faecal occult blood tests: what clinicians and colorectal cancer screening programme organisers need to know. Gut 64, 1327-1337 (2015).8. van Rossum, L. G. et al. Random comparison of guaiac and immunochemical fecal occult blood tests for colorectal cancer in a screening population. Gastroenterology 135, 82-90 (2008).9. Imperiale, T. F. et al. Multitarget stool DNA testing for colorectal-cancer screening. N. Engl. J. Med. 370, 1287-1297 (2014).10. Wilber, E., Baker, J. M. & Rebolledo, P. A. Clinical Implications of Multiplex PathogenPanels for the Diagnosis of Acute Viral Gastroenteritis. J. Clin. Microbiol. 59, eOl 51319 (2021). Mesoraca, A. et al. Evaluation of SARS-CoV-2 viral RNA in fecal samples. Virol. J. 17, 86 (2020). Herring, E., Tremblay, E., McFadden, N., Kanaoka, S. & Beaulieu, J.-F. Multitarget Stool mRNA Test for Detecting Colorectal Cancer Lesions Including Advanced Adenomas. Cancers 13, 1228 (2021). Kanaoka, S., Yoshida, K.-L, Miura, N., Sugimura, H. & Kajimura, M. Potential usefulness of detecting cyclooxygenase 2 messenger RNA in feces for colorectal cancer screening. Gastroenterology 127, 422-427 (2004).

Claims

WHAT IS CLAIMED IS:

1. A method of diagnosing a gastrointestinal disease in a subject, said method comprising: obtaining a ribonucleic acid (RNA) from a stool sample of the subject; measuring messenger RNA (mRNA) level of at least one blood-specific gene; evaluating the measured mRNA level of the at least one blood-specific gene; and determining whether the subject is 1) healthy or 2) has the gastrointestinal disease.

2. The method of claim 1, wherein the at least one blood-specific gene is selected from Table 1.

3. The method of claim 1, wherein the at least one blood-specific gene is selected from the group consisting of: HBB, HBA2, HBD, HBA1, S100A9, CSF3R, IFITM2 and LCP1.

4. The method of claim 1, wherein the at least one blood-specific gene is selected from the group consisting of: HBA1, HBA2 and HBB.

5. The method of claim 1, wherein the at least one blood-specific gene is selected from the group consisting of: HBE1, HBG1, HBG2, HBM, HBQ1 and HBZ.

6. The method of claim 1, wherein the at least one blood-specific gene is a blood-specific antigen gene or a platelet-specific gene.

7. The method of claim 1, wherein the mRNA level of the at least one blood-specific gene is measured using a primer having a sequence of any one of SEQ ID NOs: 1, 2, 4 or 5, or a probe having a sequence of SEQ ID NO: 3 or 6.

8. The method according to any one of claims 1-7, further comprising measuring mRNA levels of a group of disease-specific genes.

9. The method of claim 8, wherein the gastrointestinal disease is colorectal cancer or an advanced precancerous lesion.

10. The method of claim 9, wherein the group of disease-specific genes are selected from Tables 2-4.

11. The method of claim 1, wherein the subject is a human.

12. The method of claim 1, wherein the mRNA levels are detected using quantitative RT- PCR or Droplet Digital PCR or Partition PCR.

13. The method of claim 1, wherein the mRNA levels are determined by nucleic acid sequencing.

14. The method of claim 1, wherein the evaluating step and / or the determining step comprises analyzing the levels of a panel of biomarkers by a machine learning classifier.

15. The method of claim 14, wherein the machine learning classifier is random forests or elastic net.

16. A panel of mRNA biomarkers for use in diagnosing a gastrointestinal disease in a subject, said panel of mRNA biomarkers comprising at least one blood-specific gene.

17. The panel of mRNA biomarkers of claim 16, further comprising a group of diseasespecific genes.

18. A kit or an integrated system of diagnosing a gastrointestinal disease in a subject, comprising an agent for detecting in a stool sample obtained from the subject mRNA level of at least one blood-specific gene.

19. The kit or integrated system of claim 18, further comprising a second agent for detecting mRNA levels of a group of disease-specific genes.

20. The kit or the integrated system of claim 18, wherein the agent is selected from the group consisting of: primers, nucleic acids and oligonucleotides.

21. The kit or integrated system of claim 18, wherein the mRNA levels are detected using quantitative RT-PCR or Droplet Digital PCR or Partition PCR.

22. The kit or integrated system of claim 18, wherein the mRNA levels are determined by nucleic acid sequencing.

23. Use of a biomarker-specific reagent in the manufacture of a kit for diagnosing a gastrointestinal disease in a subject, wherein the biomarker-specific reagent specifically binds to at least one blood-specific gene.

24. A method for treating a gastrointestinal disease in a subject, the method comprising: administering to the subject a therapeutically effective amount of a drug useful for treating the gastrointestinal disease, wherein the subject has been determined to have the gastrointestinal disease based on mRNA level of at least one blood-specific gene.

25. The method of claim 24, wherein the subject has been determined to have the gastrointestinal disease based on mRNA levels of a panel of biomarkers comprising the at least one blood-specific gene and a group of disease-specific genes.