Detecting gastric neoplasms
DNA methylation markers in gastric cancer detection address the challenge of late detection by providing sensitive and specific early detection methods, utilizing RRBS and methylation-specific assays in fecal and plasma samples.
Patent Information
- Application Number
- JP2023085323
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2015-08-31
- Filing Date
- 2023-05-24
- Publication Date
- 2025-11-17
- Estimated Expiration
- 2036-08-31
AI Technical Summary
Gastric cancer is difficult to treat in Western countries due to late detection, and existing methods for early detection, such as mutational markers, are cumbersome and lack sensitivity and specificity.
Utilizing DNA methylation markers, particularly in CpG islands, to develop high-throughput assays for early detection of gastric cancer, using techniques like reduced representation bisulfite sequencing (RRBS) to identify differentially methylated regions (DMRs) in fecal and plasma samples, employing markers like ARHGEF4, ABCB1, CLEC11A, and CYP26C1, and analyzing methylation status through methods like methylation-specific PCR and sequencing.
The method provides high sensitivity and specificity for early detection of gastric cancer, enabling refined screening and diagnostic tools using fecal or blood samples.
Smart Images

Figure 0007771126000016 
Figure 0007771126000017 
Figure 0007771126000018
Abstract
Description
[Technical Field]
[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims priority to and the benefit of U.S. Provisional Application No. 62 / 212,221, the entire contents of which are incorporated herein by reference.
[0002] Presented herein are techniques related to detecting neoplasms, particularly (but not limited to) methods, compositions and related uses for detecting pre-cancerous and malignant neoplasms (e.g., gastric cancer). [Background technology]
[0003] Gastric cancer is the third most common cause of cancer-related deaths worldwide (see, e.g., World Health Organization. Cancer: Fact Sheet No. 297. WHO) and remains difficult to treat in Western countries because most patients present with advanced disease. In the United States, gastric malignancies are currently the 15th most common cancer (see, e.g., Surveillance, Epidemiology, and End Results Program. SEER Stat FactSheets: Stomach Cancer. National Cancer Institute). In most developed countries, however, gastric cancer rates have declined dramatically over the past half century.
[0004] The decline in stomach cancer is due, in part, to the widespread use of refrigeration. Refrigeration has several beneficial effects: increased consumption of fresh fruits and vegetables; reduced intake of salt, used as a food preservative; and reduced contamination of food with carcinogenic compounds resulting from the spoilage of unfrozen meat products. Salt and salted foods can damage the gastric mucosa, leading to inflammation and an associated increase in DNA synthesis and cell proliferation. Other factors likely contributing to the decline in stomach cancer rates include improved hygiene and antibiotic use, as well as lower rates of chronic Helicobacter pylori infection, a result of increased screening in some countries (see, e.g., Global Cancer Facts & Figures, 3rd ed. American Cancer Society).
[0005] However, gastric cancer remains difficult to treat in Western countries, as most patients present with advanced disease, and even those in the most favorable condition who undergo curative surgical resection often die from recurrent disease.
[0006] Therefore, there is a need for improved methods and techniques for the early detection of gastric cancer. Summary of the Invention
[0007] Methylated DNA has been investigated as a potential biomarker in tissues of most tumor types. In many instances, DNA methyltransferases add methyl groups to DNA at cytosine-phosphate-guanine (CpG) island sites as an epigenetic control of gene expression. In a biologically intriguing mechanism, acquired methylation in the promoter regions of tumor suppressor genes is thought to silence their expression and thus contribute to carcinogenesis. DNA methylation may be a more chemically and biologically stable diagnostic tool than RNA or protein expression (Laird (2010) "Principles and challenges of genome-wide DNA methylation analysis" Nat Rev Genet 11: 191-203). Furthermore, in other cancers, such as sporadic colon cancer, methylation markers provide superior specificity and are more broadly informative and sensitive than individual DNA mutations (Zou et al (2007) “Highly methylated genes in colorectal neoplasia: implications for screening” Cancer Epidemiol Biomarkers Prev 16: 2686-96).
[0008] Analysis of CpG islands has provided important insights when used in animal models and human cell lines. For example, Zhang et al. found that amplicons from different parts of the same CpG island had different levels of methylation (Zhang et al. (2009) "DNA methylation analysis of chromosome 21 gene promoters at single base pair and single allele resolution" PLoS Genet 5:e1000438). Furthermore, the methylation levels were assigned to two modes: hypermethylated and unmethylated sequences, further supporting a dual-switching pattern of DNA methyltransferase activity (Zhang et al. (2009) "DNA methylation analysis of chromosome 21 gene promoters at single base pair and single allele resolution" PLoS Genet 5:e1000438). In vivo analysis of mouse tissues and in vitro analysis of cell lines demonstrated that only approximately 0.3% of high-CpG-density promoters (HCPs, defined as having >7% CpG sequences within a 300-base pair region) are methylated, whereas low-CpG-density sites (defined as having <5% CpG sequences within a 300-base pair region) are frequently methylated in a dynamic, tissue-specific manner (Meissner et al. (2008) "Genome-scale DNA methylation maps of pluripotent and differentiated cells" Nature 454: 766-70). HCPs include promoters for ubiquitous constitutive and developmental genes.Among the HCP sites that were >50% methylated were several established markers (e.g., Wnt 2, NDRG2, SFRP2, and BMP3) (Meissner et al. (2008) “Genome-scale DNA methylation maps of pluripotent and differentiated cells” Nature 454: 766-70).
[0009] Collectively, cancers of the gastrointestinal system account for a higher number of deaths than cancers of any other organ system. Annual deaths in the United States from upper GI cancer exceed 90,000, compared with approximately 50,000 from colorectal cancer. Gastric adenocarcinoma is the second leading cause of cancer death worldwide (i.e., 1 million deaths per year). It is most common in Pacific Asia, Russia, and parts of Latin America. Because most patients present with advanced disease, the prognosis is poor (5-year survival rate <5-15%).
[0010] The genomic mechanisms underlying gastric cancer are complex and incompletely understood. Chronic gastritis due to Helicobacter pylori infection, along with environmental exposures, contributes to pathogenesis. Chromosomal and microsatellite instability have been associated with gastric cancer and epigenetic alterations. CpG methylator phenotypes have also been described in 24–47% of cases. Commonly observed mutations include those in p53, APC, β-catenin, and k-ras, which are relatively rare and heterogeneous in nature.
[0011] Gastric cancer sheds cells and DNA into the digestive stream and is ultimately secreted in the feces. Highly sensitive assays have been used to detect mutant DNA in matched feces from gastric cancer patients whose excised tumors are known to contain the same sequences. However, some limitations with mutational markers relate to the underlying heterogeneity and cumbersome process of their detection; typically, each mutation site spanning multiple genes must be assayed separately to achieve high sensitivity.
[0012] Epigenetic DNA methylation at cytosine-phosphate-guanine (CpG) island sites has been investigated as a potential biomarker for most tumor types. In a biologically intriguing mechanism, acquired methylation events in the promoter regions of tumor suppressor genes silence their expression and contribute to carcinogenesis. DNA methylation may be a more chemically and biologically stable diagnostic tool than RNA or protein expression. Furthermore, in other cancers, such as sporadic colon cancer, aberrant methylation markers offer broader insights, higher sensitivity, and superior specificity than individual DNA mutations.
[0013] Several methods are available for the search for novel methylation markers. Microarrays based on CpG methylation interrogation are a suitable, high-throughput approach, but this strategy is primarily directed to known regions of proven tumor suppressor promoters of interest. Alternative methods for genome-wide analysis of DNA methylation have been developed in the last decade. There are three basic approaches. The first uses digestion of DNA with restriction enzymes that recognize specific methylated sites, followed by several analytical techniques in which primers are used to generate methylation data limited to the restriction enzyme recognition site or to amplify DNA in an evaluation step (e.g., methylation-specific PCR; MSP). The second approach uses antibodies against methylated cytosines or other methylation-specific binding domains to enrich for the methylated fraction of genomic DNA, followed by microarray analysis or sequencing to map the fragment relative to a reference genome. This approach does not offer single-nucleotide resolution of all methylated sites within the fragment. A third approach begins with bisulfite treatment to convert all unmethylated tyrosines to uracil, followed by restriction enzyme digestion and complete sequencing of all fragments after binding to an adaptor ligand. The choice of restriction enzyme can enrich the fragments for CpG-dense regions, reducing the number of overlapping sequences that can determine multiple gene locations during analysis. This latter approach, called reduced representation bisulfite sequencing (RRBS), is known but has not yet been used to study gastric cancer.
[0014] RRBS generates single-nucleotide resolution CpG methylation status data at moderate to high read coverage for 80-90% of all CpG islands and the majority of tumor suppressor promoters. Analysis of these reads results in the identification of differentially methylated regions (DMRs). In previous RRBS analysis of pancreatic cancer samples, hundreds of DMRs were discovered, many of which had never been associated with carcinogenesis and were unannotated. Further validation studies on independent tissue samples confirmed the marker CpGs in their performance with 100% sensitivity and specificity.
[0015] A clinical approach with highly discriminatory markers could have a major impact: for example, assaying such markers in separate media (such as feces or blood) could lead to refined screening or diagnostic tools for the detection of gastric neoplasms.
[0016] Experiments conducted in developing embodiments of this technology identified and described 123 novel DNA methylation markers generated from gastric cancer samples. Along with cancer, normal stomach tissue, normal colon tissue, and normal white blood cell DNA were sequenced. Markers were validated across multiple sample populations to identify and optimize the most robust candidates. Further experiments identified 10 optimal markers for detecting gastric cancer (ARHGEF4, ELMO1, ABCB1, CLEC11A, ST8SIA1, SFMBT2, CD1D, CYP26C1, ZNF569, and C13ORF18; see Example 2 and Tables 4, 5, and 6). Further experiments identified 30 optimal markers for detecting gastric cancer (ELMO1, ARHGEF4, EMX1, SP9, CLEC11A, ST8SIA1, BMP3, KCNA3, DMRTA2, KCNK12, CD1D, PRKCB, CYP26C1, ZNF568, ABCB1, ELOVL2, PKIA, SFMBT2 (893), PCBP3, MATK, GRN2D, NDRG4, DLX4, PPP2R5C, FGF14, ZNF132, CHST2 (7890), FLI1, c13orf18, or ZNF569; see Example 1 and Tables 1, 2, and 3). Further experiments identified 12 optimal markers for detecting gastric cancer (ELMO1, ZNF569, C13orf18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4 and ST8SIA1; see Example 3 and Tables 7 and 8).
[0017] Thus, presented herein are techniques for gastric cancer screening markers that result in high signal-to-noise ratios and low background levels when detected in samples taken from subjects (e.g., fecal samples; gastric tissue samples; plasma samples). The markers were identified in patient-control studies by comparing the methylation status of DNA markers from gastric tissue or plasma of subjects with gastric cancer with the methylation status of the same DNA markers from control subjects (e.g., DNA from normal gastric tissue, normal colon epithelium, normal plasma, and normal leukocytes; see Examples 1, 2, and 3, and Tables 1-8) (e.g., plasma tissue; see Example 3).
[0018] As described herein, the techniques yield many methylation markers and subsets thereof (e.g., a set of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more markers) with high discrimination for gastric cancer overall. Experiments used selection filters for candidate markers to identify markers that exhibit high signal-to-noise ratios and low background levels for high specificity purposes for screening or diagnosing gastric cancer.
[0019] In some embodiments, the techniques involve assessing the presence and methylation status of one or more of the markers identified herein in a biological sample (e.g., stomach tissue, plasma sample). These markers include one or more of the methylation variable regions (DMRs) described herein (e.g., Tables 1, 2, 4, and / or 7). Methylation status is assessed in embodiments of the techniques. Thus, the techniques described herein are not limited by the method by which the methylation status of a gene is measured. For example, in some embodiments, the methylation status is measured by genome scanning. For example, one method involves limited-target genome scanning (Kawai et al. (1994) Mol. Cell. Biol. 14: 7421-7427), and another example involves methylation-sensitive arbitrary-start PCR (Gonzalgo et al. (1997) Cancer Res. 57: 594-599). In some embodiments, changes in methylation patterns at specific CpGs are monitored by digestion of genomic DNA with methylation-sensitive restriction enzymes followed by Southern analysis of the desired region (digestion-Southern). In some embodiments, analyzing changes in methylation patterns involves a PCR-based process that involves digestion of genomic DNA with methylation-sensitive restriction enzymes prior to PCR amplification (Singer-Sam et al. (1990) Nucl. Acids Res. 18: 687). Additionally, other techniques have been reported that utilize bisulfite treatment of DNA as a starting point for methylation analysis. These include methylation-specific PCR (MSP) (Herman et al. (1992) Proc. Natl. Acad. Sci. USA 93: 9821-9826) and restriction enzyme digestion of PCR products amplified from bisulfite-converted DNA (Sadri and Hornsby (1996) Nucl. Acids Res. 24: 5058-5059; and Xiong and Laird (1997) Nucl. Acids Res. 25: 2532-2534).PCR techniques have been developed for the detection of genetic mutations (Kuppuswamy et al. (1991) Proc. Natl. Acad. Sci. USA 88: 1143-1147) and for the quantification of allele-specific expression (Szabo and Mann (1995) Genes Dev. 9: 3097-3108; and Singer-Sam et al. (1992) PCR Methods Appl. 1: 160-163). Such techniques use internal primers that anneal to PCR-generated templates and terminate 5' adjacent to the single nucleotide to be quantified. A method using the "quantitative Ms-SNuPE assay" described in U.S. Pat. No. 7,037,650 is used in some embodiments.
[0020] When assessing methylation status, the methylation status is often expressed as the fraction or percentage of an individual strand of DNA that is methylated at a specific site (e.g., a single nucleotide, a specific region or locus, or a longer desired sequence (e.g., a subsequence of DNA of ∼100 bp, 200 bp, 500 bp, 1000 bp or more) relative to the total population of DNA in a sample containing that site. Previously, the amount of unmethylated nucleic acid is determined by PCR using standards. A known amount of DNA is then bisulfite-treated, and the resulting methylation-specific sequences are determined using real-time PCR or other exponential amplification (e.g., as demonstrated by the QuARTS assay (e.g., U.S. Pat. No. 8,361,720; U.S. Patent Application Publication No. 2012 / 0122088; and U.S. Patent Application Publication No. 2012 / 0122106, both of which are incorporated by reference)).
[0021] For example, in some embodiments, the method includes generating a standard curve for the unmethylated target by using an external standard. The standard curve is composed of at least two points and relates the real-time Ct values for unmethylated DNA to the known quantitative standard. A second standard curve for the methylated target is then generated from the at least two points and the external standard. This second standard curve relates the Ct values for methylated DNA to the known quantitative standard. Test sample Ct values are then determined for the methylated and unmethylated populations, and the genome equivalents of DNA are calculated from the standard curve generated by the first two steps. The percentage of methylation at the desired site is calculated from the amount of methylated DNA relative to the total amount of DNA in the population (e.g., (number of methylated DNA) / (number of methylated DNA + number of unmethylated DNA) × 100).
[0022] Also disclosed herein are compositions and kits for carrying out the above-described methods. For example, in some embodiments, reagents (e.g., primers, probes) specific to one or more markers are provided, either individually or as a set (e.g., a set of primer pairs for amplifying multiple markers). Additional reagents for carrying out detection assays (e.g., enzymes, buffers, positive and negative controls for carrying out QuARTS, PCR, sequencing, bisulfite, or other assays) are also provided. Reaction mixtures containing the above-described reagents are also provided. Furthermore, master mix reagent sets are provided, which contain multiple reagents that can be added to each other and / or to a test sample to complete a reaction mixture.
[0023] In some embodiments, the technology described herein relates to a programmable machine designed to perform a series of arithmetic or logical operations, such as those provided by the methods described herein. For example, some embodiments of the technology relate to (e.g., are implemented in) computer software and / or computer hardware. In one aspect, the technology relates to a computer that includes a form of memory, elements for performing arithmetic and logical operations, and a processing element (e.g., a microprocessor) for executing a series of instructions (e.g., methods as described herein) for reading, manipulating, and storing data. In some embodiments, the microprocessor is part of a system for determining the methylation status (e.g., of one or more DMRs (e.g., DMRs 1-274 as shown in Tables 1, 2, 4, and 7)); comparing the methylation status (e.g., of one or more DMRs (e.g., DMRs 1-274 as shown in Tables 1, 2, 4, and 7)); generating a standard curve; determining Ct values; calculating the fraction, frequency, or percentage of methylation (e.g., of one or more DMRs (e.g., DMRs 1-274 as shown in Tables 1, 2, 4, and 7)); identifying CpG islands; determining the specificity and / or sensitivity of an assay or marker; calculating ROCkyokusenn and associated AUC; and sequence analysis (all as described herein or known in the art).
[0024] In some embodiments, the microprocessor or computer uses the methylation status data in an algorithm to predict the site of the cancer.
[0025] In some embodiments, a software or hardware element receives the results of multiple assays and determines a single numerical result for reporting to a user indicative of cancer risk based on the results (e.g., determining the methylation status of multiple DMRs (e.g., as shown in Tables 1, 2, 4, and 7)). Related embodiments calculate a risk factor based on a mathematical combination (e.g., weighted combination, linear combination) of the results of multiple assays (e.g., determining the methylation status of multiple markers (e.g., as shown in Tables 1, 2, 4, and 7)). In some embodiments, the methylation status of the DMRs defines one dimension and can have values in a multidimensional space, and the coordinate defined by the methylation status of the multiple DMRs is the result (e.g., for reporting to a user) related to cancer risk.
[0026] Some embodiments include a recording medium and a memory element (e.g., volatile and / or non-volatile memory) for use in storing instructions (e.g., an embodiment of a process as described herein) and / or data (e.g., artifacts (e.g., methylation measurements, sequencing, and associated statistical descriptions)). Some embodiments also relate to systems that include one or more of a CPU, a graphics card, and a user interface (e.g., including an output device (e.g., a display) and an input device (e.g., a keyboard)).
[0027] The programmable machines associated with the above technologies include conventional existing technologies and technologies under development or yet to be developed (eg, quantum computers, chemical computers, optical computers, spintronics-based computers, etc.).
[0028] In some embodiments, the techniques involve wired transmission media (e.g., metal cables, optical fibers) or wireless transmission media for data transmission. For example, some embodiments involve data transmission across a network (e.g., a local area network (LAN), a wide area network (WAN), a specialized network, the Internet, etc.). In some embodiments, the programmable machine has a client / server relationship.
[0029] In some embodiments, the data is stored on a computer-readable recording medium (eg, a hard disk, flash memory, optical media, floppy disk, etc.).
[0030] In some embodiments, the techniques presented herein involve multiple programmable devices operating in concert to implement the methods as described herein. For example, in some embodiments, multiple computers (e.g., connected to a network) may operate simultaneously to collect and process data (e.g., in a cluster or grid computing or some other computing implementation, relying on complete computers (having on-board CPUs, storage, power supplies, network interfaces, etc.) connected by conventional network interfaces (e.g., Ethernet, fiber optics) or wireless communication technologies).
[0031] For example, some embodiments provide a computer including a computer-readable medium. The embodiment includes a random access memory (RAM) coupled to a processor. The processor executes computer-executable program instructions stored in the memory. Such a processor may include a microprocessor, an ASIC, a state machine, or other processor, and may be any of a number of computer processors (e.g., processors offered by Intel Corporation of Santa Clara, California and Motorola Corporation of Schaumburg, Illinois). Such a processor may include or be coupled to a medium (e.g., a computer-readable medium) that stores instructions that, when executed by the processor, cause the processor to perform the steps described herein.
[0032] Embodiments of computer-readable media include, but are not limited to, electronic, optical, magnetic, or other storage or transmission devices capable of providing computer-readable instructions to a processor. Other examples of suitable media include, but are not limited to, floppy disks, CD-ROMs, magnetic disks, memory chips, ROM, RAM, ASICs, configured processors, all optical media, all magnetic tapes, or any other magnetic media from which a computer can read instructions. Also, various other forms of computer-readable media, including routers, private or public networks, or other wireless or wired transmission devices or paths, may transmit or convey instructions to a computer. The instructions may include code based on any suitable computer programming language, including, for example, C, C++, C#, Visual Basic, Java, Python, Perl, and JavaScript.
[0033] In some embodiments, the computer is connected to a network. The computer may also include many external or internal devices (e.g., a mouse, CD-ROM, DVD, keyboard, display, or other input or output devices). Examples of computers include personal computers, digital assistants, personal digital assistants, mobile phones, cell phones, smartphones, pagers, digital tablets, laptop computers, Internet appliances, and other processor-based devices. In general, computers related to aspects of the technology described herein may be any type of processor-based platform running any operating system capable of supporting one or more programs, including the technology described herein (e.g., Microsoft Windows, Linux, UNIX, Mac OS X, etc.). Some embodiments encompass personal computers running other application programs (e.g., multiple applications). The multiple applications may be stored in memory and may include, for example, a word processing application, a spreadsheet application, an email application, a quick messaging application, a presentation application, an Internet browser application, a calendar / organizer application, and any other application executable by a client device.
[0034] All such components, computers, and systems described herein as associated with the above technology may be logical or virtual.
[0035] Thus, provided herein are techniques for screening for neoplasia in a sample obtained from a subject. The methods include assessing the methylation status of a marker in a sample (e.g., gastric tissue) (e.g., a plasma sample) obtained from the subject; and identifying the subject as having a neoplasia when the methylation status of the marker differs from the methylation status of the marker assessed in a subject without the neoplasia. The marker comprises a base in a methylation variable region (DMR) selected from the group consisting of DMRs 1-274 as set forth in Tables 1, 2, 4, and / or 7. The techniques relate to identifying and differentiating between gastric cancers. In some embodiments, the neoplasia is a gastric neoplasia (e.g., gastric cancer). Some embodiments provide methods that involve assessing multiple markers (e.g., assessing 2-11-100 or 200 or 274 markers).
[0036] The techniques are not limited by the methylation state assessed. In some embodiments, assessing the methylation state of a marker in a sample comprises quantifying the methylation state of a single base. In some embodiments, quantifying the methylation state of a marker in a sample comprises determining the degree of methylation at multiple bases. Further, in some embodiments, the methylation state of a marker comprises an increased methylation state of the marker relative to the marker's normal methylation state. In some embodiments, the methylation state of a marker comprises a decreased methylation state of the marker relative to the marker's normal methylation state. In some embodiments, the methylation state of a marker comprises a pattern of methylation of the marker that differs from the marker's normal methylation state.
[0037] Furthermore, in some embodiments, the marker is a region of less than 100 bases, a region of less than 500 bases, a region of less than 1000 bases, a region of less than 5000 bases, or in some embodiments, the marker is 1 base. In some embodiments, the marker is in a promoter with high CpG density.
[0038] The techniques are not limited by the type of sample. For example, in some embodiments, the sample is a fecal sample, a tissue sample (e.g., a stomach tissue sample), a blood sample (e.g., plasma, serum, whole blood), stool, or a urine sample.
[0039] Furthermore, the techniques are not limited by the method used to determine methylation status. In some embodiments, quantifying involves using methylation-specific polymerase chain reaction, nucleic acid sequencing, mass spectrometry, methylation-specific nucleases, mass-based separation, or target capture. In some embodiments, quantifying involves the use of methylation-specific oligonucleotides. In some embodiments, the techniques utilize parallel large-scale sequencing (e.g., next-generation sequencing) (e.g., sequencing-by-synthesis, real-time (e.g., single-molecule) sequencing, bead emulsion sequencing, nanopore sequencing, etc.) to determine methylation status.
[0040] The above techniques provide reagents for detecting DMRs, for example, in some embodiments, a set of oligonucleotides is provided that includes the sequences set forth in SEQ ID NOS: 1 to 109. In some embodiments, oligonucleotides are provided that include a sequence complementary to a chromosomal region that contains a base in a DMR.
[0041] The techniques provide various panels of markers, for example, in some embodiments, the markers include chromosomal regions annotated as ARHGEF4, ELMO1, ABCB1, CLEC11A, ST8SIA1, SFMBT2, CD1D, CYP26C1, ZNF569, or C13ORF18, which contain the marker. Further, embodiments provide methods for analyzing DMRs according to Table 4, which are DMR numbers 49, 43, 196, 66, 1, 237, 249, 250, 251, or 252. Some embodiments provide for determining the methylation status of markers, wherein the chromosomal region has an annotation that is ELMO1, ARHGEF4, EMX1, SP9, CLEC11A, ST8SIA1, BMP3, KCNA3, DMRTA2, KCNK12, CD1D, PRKCB, CYP26C1, ZNF568, ABCB1, ELOVL2, PKIA, SFMBT2 (893), PCBP3, MATK, GRN2D, NDRG4, DLX4, PPP2R5C, FGF14, ZNF132, CHST2 (7890), FLI1, c13orf18, or ZNF569, including markers (see Table 2). Further, embodiments provide a method of analyzing a DMR based on Table 2 that is DMR number 253, 251, 254, 255, 256, 249, 257, 258, 259, 260, 261, 262, 250, 263, 1, 264, 265, 196, 266, 118, 267, 268, 269, 270, 271, 272, 46, 273, 252, or 237. Further, in such embodiments where the sample is a plasma sample, the method analyzes a DMR from Tables 1, 2, 4, and 7 that is DMR number 253, 251, 254, 255, 256, 249, 257, 258, 259, 260, 261, 262, 250, 263, 1, 264, 265, 196, 266, 118, 267, 268, 269, 270, 271, 272, 46, 273, 252, or 237. In some embodiments, the method comprises determining the methylation status of two markers (e.g., a pair of markers set forth in Tables 1, 2, 4, or 7).
[0042] Kit embodiments are provided, e.g., kits that include a bisulfite reagent; and a control nucleic acid that includes a sequence based on a DMR selected from the group consisting of DMR1-274 (based on Tables 1, 2, 4, or 7) and has a methylation status associated with subjects who do not have cancer. In some embodiments, the kit includes a bisulfite reagent and an oligonucleotide as described herein. In some embodiments, the kit includes a bisulfite reagent; and a control nucleic acid that includes a sequence based on a DMR selected from the group consisting of DMR1-274 (based on Tables 1, 2, 4, or 7) and has a methylation status associated with subjects who do not have cancer. Some kit embodiments include a sample collector for collecting a sample (e.g., a stool sample, a stomach tissue sample, a plasma sample) from a subject; reagents for isolating nucleic acids from the sample; and an oligonucleotide as described herein.
[0043] The technology relates to embodiments of compositions (e.g., reaction mixtures). In some embodiments, compositions are provided that include a nucleic acid containing a DMR and a bisulfite reagent. Some embodiments provide compositions that include a nucleic acid containing a DMR and an oligonucleotide as described herein. Some embodiments provide compositions that include a nucleic acid containing a DMR and a methylation-sensitive restriction enzyme. Some embodiments provide compositions that include a nucleic acid containing a DMR and a polymerase.
[0044] Further related method embodiments are provided for screening for neoplasms in a sample obtained from a subject (e.g., a stomach tissue sample, a plasma sample, a fecal sample). The method includes determining the methylation status of a marker in a sample containing a base in a DMR that is one or more of DMRs 1-274 (Tables 1, 2, 4, or 7); comparing the methylation status of the marker from the subject sample with the methylation status of the marker from a normal control sample from a subject without cancer; and determining a confidence interval and / or p-value for the difference in methylation status between the subject sample and the normal control sample. In some embodiments, the confidence interval is 90%, 95%, 97.5%, 98%, 99%, 99.5%, 99.9%, or 99.99%, and the p-value is 0.1, 0.05, 0.025, 0.02, 0.01, 0.005, 0.001, or 0001. Some embodiments of the method include reacting a nucleic acid containing a DMR with a bisulfite reagent to produce a bisulfite-reacted nucleic acid; sequencing the bisulfite-reacted nucleic acid to reveal the nucleotide sequence of the bisulfite-reacted nucleic acid; comparing the nucleotide sequence of the bisulfite-reacted nucleic acid to the nucleotide sequence of a nucleic acid containing a DMR from a subject without cancer to identify differences in the two sequences; and, if there are differences, identifying the subject as having a neoplasia.
[0045] A system for screening for neoplasms in a sample obtained from a subject is provided by the above-described technique. Exemplary embodiments of the system include, for example, a system for screening for neoplasms in a sample obtained from a subject (e.g., a gastric tissue sample, a plasma sample, a fecal sample). The system includes an analytical component configured to determine the methylation status of the sample, a software component for comparing the methylation status of the sample to the methylation status of a control sample or a reference sample recorded in a database, and an alert component configured to alert a user to a methylation status associated with cancer. The alert, in some embodiments, is determined by software receiving results from multiple assays (e.g., determining the methylation status of multiple markers (e.g., as shown in Tables 1, 2, or 7)) and calculating a value or result for reporting based on the multiple results. Some embodiments provide a database of weighting variables associated with each DMR described herein for use in calculating a value or result and / or alert for reporting to a user (e.g., a physician, nurse, clinician, etc.). In some embodiments, all results from multiple assays are recorded, and in some embodiments, one or more results are used to indicate a score, value, or result based on a combination of one or more results based on multiple assays that is indicative of cancer risk in the subject.
[0046] In some embodiments of the system, the sample includes a nucleic acid containing a DMR. In some embodiments, the system further includes a component for isolating the nucleic acid and a component for collecting the sample (e.g., a component for collecting a stool sample). In some embodiments, the system includes a nucleic acid sequence containing a DMR. In some embodiments, the database includes nucleic acid sequences from subjects who do not have cancer. Nucleic acids (e.g., a collection of multiple nucleic acids, each nucleic acid containing a DMR) are also provided. In some embodiments, a collection of multiple nucleic acids, each nucleic acid containing a sequence from a subject who does not have the eye. Related system embodiments include a collection of multiple nucleic acids as described and a database of multiple nucleic acid sequences associated with the collection of multiple nucleic acids. Some embodiments further include a bisulfite reagent. Some embodiments further include a nucleic acid sequencer.
[0047] In some embodiments, methods are provided for characterizing a sample (e.g., a gastric tissue sample, a plasma sample, a fecal sample) from a human patient. For example, in some embodiments, such embodiments include obtaining DNA from the human patient sample; assessing the methylation status of a DNA methylation marker comprising a base in a methylation variable region (DMR) selected from the group consisting of DMRs 1-274 according to Tables 1, 2, or 7; and comparing the assessed methylation status of the one or more DNA methylation markers to a reference methylation level for the one or more DNA methylation markers for human patients without gastric neoplasia.
[0048] Such methods are not limited to a particular type of sample from a human patient. In some embodiments, the sample is a gastric tissue sample. In some embodiments, the sample is a plasma sample. In some embodiments, the sample is a fecal sample, a tissue sample, a gastric tissue sample, a blood sample, or a urine sample.
[0049] In some embodiments, such methods include assessing a plurality of DNA methylation markers. In some embodiments, such methods include assessing 12 to 10 DNA methylation markers. In some embodiments, such methods include assessing the methylation state of one or more DNA methylation markers in a sample, including assessing the methylation state of a single base. In some embodiments, such methods include assessing the methylation state of one or more DNA methylation markers in a sample, including determining the degree of methylation state at a plurality of bases. In some embodiments, such methods include assessing the methylation state of a forward strand or assessing the methylation state of a reverse strand.
[0050] In some embodiments, the DNA methylation marker is a region of less than 100 bases. In some embodiments, the DNA methylation marker is a region of less than 500 bases. In some embodiments, the DNA methylation marker is a region of less than 1000 bases. In some embodiments, the DNA methylation marker is a region of less than 10,000 bases. In some embodiments, the DNA methylation marker is a single base. In some embodiments, the DNA methylation marker is a promoter with high CpG density.
[0051] In some embodiments, the evaluating comprises using methylation-specific polymerase chain reaction, nucleic acid sequencing, mass spectrometry, methylation-specific nucleases, mass-based separation, or target capture.
[0052] In some embodiments, the evaluating comprises the use of a methylation-specific oligonucleotide. In some embodiments, the methylation-specific oligonucleotide is selected from the group consisting of SEQ ID NOs: 1-109.
[0053] In some embodiments, the chromosomal region having an annotation selected from the group consisting of ARHGEF4, ELMO1, ABCB1, CLEC11A, ST8SIA1, SFMBT2, CD1D, CYP26C1, ZNF569, and C13ORF18 comprises a methylation marker. In some embodiments, the DMR is based on Table 4 and is selected from the group consisting of DMR numbers 49, 43, 196, 66, 1, 237, 249, 250, 251, and 252.
[0054] In some embodiments, the chromosomal region having an annotation selected from the group consisting of ELMO1, ARHGEF4, EMX1, SP9, CLEC11A, ST8SIA1, BMP3, KCNA3, DMRTA2, KCNK12, CD1D, PRKCB, CYP26C1, ZNF568, ABCB1, ELOVL2, PKIA, SFMBT2 (893), PCBP3, MATK, GRN2D, NDRG4, DLX4, PPP2R5C, FGF14, ZNF132, CHST2 (7890), FLI1, c13orf18, and ZNF569 comprises a DNA methylation marker.
[0055] In some embodiments, the DMR is based on Table 2 and is selected from the group consisting of DMR numbers 253, 251, 254, 255, 256, 249, 257, 258, 259, 260, 261, 262, 250, 263, 1, 264, 265, 196, 266, 118, 267, 268, 269, 270, 271, 272, 46, 273, 252, and 237.
[0056] In some embodiments where the obtained sample is a plasma sample, the markers include chromosomal regions annotated as including markers ELMO1, ZNF569, C13orf18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4, and ST8SIA1.
[0057] In some embodiments, the DMR is based on Table 1, 2, 4 or 7 and is selected from the group consisting of 253, 237, 252, 261, 251, 196, 250, 265, 256, 249 and 274.
[0058] In some embodiments, such methods comprise determining the methylation status of two DNA methylation markers, hi some embodiments, such methods comprise determining the methylation status of a pair of DNA methylation markers set forth in Tables 1, 2, 4, or 7.
[0059] In some embodiments, the techniques provide methods for characterizing samples obtained from human patients. In some embodiments, such methods include determining the methylation status of a DNA methylation marker in a sample containing a base in a DMR selected from the group consisting of DMRs 1-274 based on Tables 1, 2, 4, and 7; comparing the methylation status of the DNA methylation marker from the patient sample with the methylation status of the DNA methylation marker from a normal control sample from a human subject without cancer; and determining a confidence interval and p-value for the difference in methylation status between the human patient and the normal control sample. In some embodiments, the confidence interval is 90%, 95%, 97.5%, 98%, 99%, 99.5%, 99.9%, or 99.99%, and the p-value is 0.1, 0.05, 0.025, 0.02, 0.01, 0.005, 0.001, or 0.0001.
[0060] In some embodiments, the technology provides a method for characterizing a sample (e.g., a gastric tissue sample, a plasma sample, or a fecal sample) obtained from a human subject, comprising reacting a nucleic acid containing a DMR with a bisulfite reagent to produce a bisulfite-reacted nucleic acid; sequencing the bisulfite-reacted nucleic acid to provide a nucleotide sequence of the bisulfite-reacted nucleic acid; and comparing the nucleotide sequence of the bisulfite-reacted nucleic acid to the nucleotide sequence of a nucleic acid containing a DMR from a subject without gastric cancer.
[0061] In some embodiments, the technology provides a system for characterizing a sample (e.g., a gastric tissue sample, a plasma sample, or a fecal sample) obtained from a human subject. The system includes an analytical component configured to determine the methylation state of the sample, a software component configured to compare the methylation state of the sample with control or reference sample methylation states stored in a database, and an alerting component configured to determine a single value based on a combination of the methylation states and alert a user to a methylation state associated with gastric cancer. In some embodiments, the sample includes a nucleic acid comprising a DMR.
[0062] In some embodiments, such systems further include a component for isolating nucleic acids, hi some embodiments, such systems further include a component for collecting a sample.
[0063] In some embodiments, the sample is a fecal sample, a tissue sample, a stomach tissue sample, a blood sample, or a urine sample.
[0064] In some embodiments, the database includes nucleic acid sequences that include DMRs. In some embodiments, the database includes nucleic acid sequences from subjects that do not have gastric cancer.
[0065] Further embodiments will be apparent to those skilled in the relevant arts based on the teachings contained herein.
[0066] These and other features, aspects and advantages of the present technology will be better understood with reference to the following figures.
[0067] It is understood that the drawings are not necessarily to scale, and that objects in the drawings are not to scale relative to each other. The drawings are representations intended to clarify and facilitate understanding of various embodiments of the devices, systems, compositions, and methods disclosed herein. In all cases, the same reference numbers are used throughout the drawings to describe the same or similar parts. Furthermore, the drawings are not intended to limit the scope of the present teachings in any way. [Brief explanation of the drawings]
[0068] [Figure 1] 1 shows that ELMO1, as described in Example 2, exhibited high discrimination against gastric cancer. [Figure 2]
[0023] Figure 1 shows oligonucleotide sequences for FRET cassettes used to detect methylated DNA features by QuARTs (Quantitative Allele-Specific Real-Time Target Signal Amplification) assays. Each FRET sequence contains a fluorophore and quencher that can be multiplexed together in three separate assays. [Figure 3] Receiver operating characteristic curves are shown highlighting the performance in plasma of a three marker panel (ELMO1, ZNF569, and c13orf18) compared to the individual marker curves. At 100% specificity, the panel detected gastric cancer with a sensitivity of 86%. [Figure 4A] Log-scale absolute strand counts of methylated ELMO1 in plasma according to stage, illustrating the progression from normal mucosa to stage 4 gastric cancer. Quantitative marker levels increased with GC stage. [Figure 4B] Figure 1 shows a bar graph demonstrating the sensitivity (100% specificity) of gastric cancer in plasma according to stage for a three marker panel (ELMO1, ZNF569 and c13orf18). [Figure 5A] Performance of 12 gastric cancer markers (ELMO1, ZNF569, C13orf18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4, and ST8SIA1; ELMO1 was run twice, the second time in biplex format) is shown. 90% specificity (A), 95% specificity (B), and 100% specificity (C) are shown in matrix format for 74 plasma samples. Markers are arranged vertically and samples horizontally. Samples are aligned with normal on the left and cancer on the right. Positive hits are in light gray, and misses are in dark gray. Here, the top-performing marker (ELMO1) is listed first at 90% specificity, and the remaining panels are at 100% specificity. This plot allows markers to be evaluated in a combinatorial manner. [Figure 5B] Performance of 12 gastric cancer markers (ELMO1, ZNF569, C13orf18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4, and ST8SIA1; ELMO1 was run twice, the second time in biplex format) is shown. 90% specificity (A), 95% specificity (B), and 100% specificity (C) are shown in matrix format for 74 plasma samples. Markers are arranged vertically and samples horizontally. Samples are aligned with normal on the left and cancer on the right. Positive hits are in light gray, and misses are in dark gray. Here, the top-performing marker (ELMO1) is listed first at 90% specificity, and the remaining panels are at 100% specificity. This plot allows markers to be evaluated in a combinatorial manner. [Figure 5C]Performance of 12 gastric cancer markers (ELMO1, ZNF569, C13orf18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4, and ST8SIA1; ELMO1 was run twice, the second time in biplex format) is shown. 90% specificity (A), 95% specificity (B), and 100% specificity (C) are shown in matrix format for 74 plasma samples. Markers are arranged vertically and samples horizontally. Samples are aligned with normal on the left and cancer on the right. Positive hits are in light gray, and misses are in dark gray. Here, the top-performing marker (ELMO1) is listed first at 90% specificity, and the remaining panels are at 100% specificity. This plot allows markers to be evaluated in a combinatorial manner. [Figure 6] Shows the performance of three gastric cancer markers (first panel) in 74 plasma samples, with 100% specificity. Markers are arranged vertically, samples horizontally. Cancers are arranged according to stage. Positive hits are in light gray, misses are in dark gray (ELMO1, ZNF569, and c13orf18). [Figure 7A] Box plots (linear scale) of 12 gastric cancer markers (ELMO1, ZNF569, C13orf18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4, and ST8SIA1; ELMO1 was performed twice, the second time in biplex format) in plasma are shown. Control samples (N=38) are listed on the left, and gastric cancer cases (N=36) are listed on the right. The vertical axis is the % methylation normalized to β-actin strands. [Figure 7B]Box plots (linear scale) of 12 gastric cancer markers (ELMO1, ZNF569, C13orf18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4, and ST8SIA1; ELMO1 was performed twice, the second time in biplex format) in plasma are shown. Control samples (N=38) are listed on the left, and gastric cancer cases (N=36) are listed on the right. The vertical axis is the % methylation normalized to β-actin strands. [Figure 7C] Box plots (linear scale) of 12 gastric cancer markers (ELMO1, ZNF569, C13orf18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4, and ST8SIA1; ELMO1 was performed twice, the second time in biplex format) in plasma are shown. Control samples (N=38) are listed on the left, and gastric cancer cases (N=36) are listed on the right. The vertical axis is the % methylation normalized to β-actin strands. [Figure 7D] Box plots (linear scale) of 12 gastric cancer markers (ELMO1, ZNF569, C13orf18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4, and ST8SIA1; ELMO1 was performed twice, the second time in biplex format) in plasma are shown. Control samples (N=38) are listed on the left, and gastric cancer cases (N=36) are listed on the right. The vertical axis is the % methylation normalized to β-actin strands. [Figure 7E] Box plots (linear scale) of 12 gastric cancer markers (ELMO1, ZNF569, C13orf18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4, and ST8SIA1; ELMO1 was performed twice, the second time in biplex format) in plasma are shown. Control samples (N=38) are listed on the left, and gastric cancer cases (N=36) are listed on the right. The vertical axis is the % methylation normalized to β-actin strands. [Figure 7F] Box plots (linear scale) of 12 gastric cancer markers (ELMO1, ZNF569, C13orf18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4, and ST8SIA1; ELMO1 was performed twice, the second time in biplex format) in plasma are shown. Control samples (N=38) are listed on the left, and gastric cancer cases (N=36) are listed on the right. The vertical axis is the % methylation normalized to β-actin strands. [Figure 7G] Box plots (linear scale) of 12 gastric cancer markers (ELMO1, ZNF569, C13orf18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4, and ST8SIA1; ELMO1 was performed twice, the second time in biplex format) in plasma are shown. Control samples (N=38) are listed on the left, and gastric cancer cases (N=36) are listed on the right. The vertical axis is the % methylation normalized to β-actin strands. [Figure 7H] Box plots (linear scale) of 12 gastric cancer markers (ELMO1, ZNF569, C13orf18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4, and ST8SIA1; ELMO1 was performed twice, the second time in biplex format) in plasma are shown. Control samples (N=38) are listed on the left, and gastric cancer cases (N=36) are listed on the right. The vertical axis is the % methylation normalized to β-actin strands. [Figure 7I] Box plots (linear scale) of 12 gastric cancer markers (ELMO1, ZNF569, C13orf18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4, and ST8SIA1; ELMO1 was performed twice, the second time in biplex format) in plasma are shown. Control samples (N=38) are listed on the left, and gastric cancer cases (N=36) are listed on the right. The vertical axis is the % methylation normalized to β-actin strands. [Figure 7J] Box plots (linear scale) of 12 gastric cancer markers (ELMO1, ZNF569, C13orf18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4, and ST8SIA1; ELMO1 was performed twice, the second time in biplex format) in plasma are shown. Control samples (N=38) are listed on the left, and gastric cancer cases (N=36) are listed on the right. The vertical axis is the % methylation normalized to β-actin strands. [Figure 7K] Box plots (linear scale) of 12 gastric cancer markers (ELMO1, ZNF569, C13orf18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4, and ST8SIA1; ELMO1 was performed twice, the second time in biplex format) in plasma are shown. Control samples (N=38) are listed on the left, and gastric cancer cases (N=36) are listed on the right. The vertical axis is the % methylation normalized to β-actin strands. [Figure 7L] Box plots (linear scale) of 12 gastric cancer markers (ELMO1, ZNF569, C13orf18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4, and ST8SIA1; ELMO1 was performed twice, the second time in biplex format) in plasma are shown. Control samples (N=38) are listed on the left, and gastric cancer cases (N=36) are listed on the right. The vertical axis is the % methylation normalized to β-actin strands. [Figure 7M]Box plots (linear scale) of 12 gastric cancer markers (ELMO1, ZNF569, C13orf18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4, and ST8SIA1; ELMO1 was performed twice, the second time in biplex format) in plasma are shown. Control samples (N=38) are listed on the left, and gastric cancer cases (N=36) are listed on the right. The vertical axis is the % methylation normalized to β-actin strands. DETAILED DESCRIPTION OF THE INVENTION
[0069] Presented herein is technology relating to, but not limited to, methods, compositions, and related uses for detecting tumorigenesis, and more particularly, for detecting pre-cancerous and malignant tumors (e.g., gastric cancer). As the technology is described herein, the section headings used are for organizational purposes only and should not be construed as limiting the subject matter in any way.
[0070] In this detailed description of various embodiments, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the disclosed embodiments. Those skilled in the art will appreciate that these various embodiments may be practiced with or without these specific details. In other instances, structures and devices are shown in block diagram form. Furthermore, those skilled in the art will readily appreciate that the specific order in which various methods are presented and performed is illustrative, and that this order can be changed and still be within the spirit and scope of the various embodiments disclosed herein.
[0071] All literature and similar materials cited in this application, including but not limited to patents, patent applications, articles, books, papers, and Internet web pages, are expressly incorporated herein by reference for any purpose. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the various embodiments described herein belong. When it is clear that the definition of a term in an incorporated reference differs from the definition provided in the present teachings, the definition provided in the present teachings shall control.
[0072] [Definition] To facilitate understanding of this invention, a number of terms and phrases are defined below. Additional definitions are set forth throughout the detailed description.
[0073] Throughout the specification and claims, unless the context clearly dictates otherwise, the following terms have the meanings clearly associated therewith. As used herein, the phrase "in one embodiment" does not necessarily refer to the same embodiment, although it may. Furthermore, as used herein, the phrase "other embodiments" does not necessarily refer to different embodiments, although it may. Thus, as described below, various embodiments of the invention can be readily combined without departing from the scope or spirit of the invention.
[0074] Additionally, as used herein, the term "or" is an inclusive "or" modifier and is synonymous with the term "and / or" unless the context clearly dictates otherwise. The term "based on" is not exclusive and takes into account additional, unlisted elements unless the context clearly dictates otherwise. Additionally, as used herein, the meanings of "a," "an," and "the" include plural references. The meaning of "in" includes "in" and "on."
[0075] As used herein, "nucleic acid" or "nucleic acid molecule" generally refers to ribonucleic acid or deoxyribonucleic acid, which may be unmodified or modified DNA or RNA. "Nucleic acid" includes, but is not limited to, single-stranded and double-stranded nucleic acids. As used herein, the term "nucleic acid" also includes DNA containing one or more modified bases as described above. Thus, DNA with a backbone modified for stability or other reasons is a "nucleic acid." As used herein, the term "nucleic acid" encompasses chemical forms of DNA characteristic of viruses and cells (including, for example, single cells and complex cells), as well as chemically, enzymatically, or metabolically modified forms of nucleic acids.
[0076] The terms "oligonucleotide" or "polynucleotide" or "nucleotide" or "nucleic acid" refer to a molecule having two or more deoxyribonucleotides or ribonucleotides, preferably three or more, and usually more than 10. The exact size can depend on many factors, which in turn depend on the ultimate function or use of the oligonucleotide.
[0077] Oligonucleotides can be produced by any method, including chemical synthesis, DNA replication, reverse transcription, or a combination thereof. Typical deoxyribonucleotides of DNA are thymine, adenine, cytosine, and guanine. Typical ribonucleotides of RNA are uracil, adenine, cytosine, and guanine.
[0078] As used herein, a "locus" or "region" of a nucleic acid refers to a subregion of a nucleic acid, e.g., a gene on a chromosome, a single nucleotide, a CpG island, etc.
[0079] The terms "complementary" and "complementarity" refer to nucleotides (e.g., single nucleotides) or polynucleotides (e.g., consecutive nucleotides) related by the base-pairing rules. For example, the sequence "5'-AGT-3'" is complementary to the sequence "3'-TCA-5'." Complementarity may be "partial," in which only some of the nucleic acid bases match according to the base-pairing rules. Alternatively, there may be "complete" or "total" complementarity between nucleic acids. The degree of complementarity between nucleic acid strands affects the efficiency and strength of hybridization between nucleic acid strands. This is particularly important in amplification reactions and detection methods that rely on binding between nucleic acids.
[0080] The term "gene" refers to a nucleic acid (e.g., DNA or RNA) sequence that comprises coding sequences necessary for the production of an RNA, polypeptide, or precursor thereof. A functional polypeptide can be encoded by a full-length coding sequence or any portion of a coding sequence, so long as the desired activity or functional property of the polypeptide (e.g., enzymatic activity, ligand binding, signal transduction, etc.) is retained. The term "portion," when used in reference to a gene, refers to a fragment of that gene. Fragments can range in size from a few nucleotides to the entire gene minus one nucleotide. Thus, "nucleotides comprising at least a portion of a gene" can include multiple fragments of the gene or the entire gene.
[0081] The term "gene" also encompasses the coding region of a structural gene, as well as sequences adjacent to the coding region at both the 5' and 3' ends, e.g., sequences located adjacent to the coding region for a distance of about 1 kb on either end, where the gene corresponds to the length of the full-length mRNA (e.g., including coding, regulatory, structural, and other sequences). Sequences located 5' of the coding region and present on the mRNA are referred to as 5' non-translated or 5' untranslated sequences. Sequences located 3' or downstream of the coding region and present on the mRNA are referred to as 3' non-translated or 3' untranslated sequences. The term "gene" encompasses both cDNA and genomic forms of a gene. In some organisms (e.g., eukaryotes), genomic forms or clones of a gene contain coding regions interrupted by non-coding sequences called "introns" or "intervening regions" or "intervening sequences." Introns are segments of a gene that are transcribed into nuclear RNA (hnRNA); introns may contain regulatory elements such as enhancers. Introns are removed or "spliced out" from the nuclear or primary transcript; introns are therefore absent in the messenger RNA (mRNA) transcript. mRNA functions during translation to specify the sequence or order of amino acids in a nascent polypeptide.
[0082] In addition to containing introns, genomic forms of a gene may also include sequences located on both the 5' and 3' end of the sequences present on the RNA transcript. These sequences are referred to as "flanking" sequences or regions (these flanking sequences are located 5' or 3' to the untranslated sequences present on the mRNA transcript). The 5' flanking region may contain regulatory sequences, such as promoters and enhancers, that control or influence the transcription of the gene. The 3' flanking region may contain sequences that direct the termination of transcription, post-transcriptional cleavage, and polyadenylation.
[0083] The term "wild-type," when constructed with respect to a gene, refers to a gene that has the characteristics of a gene isolated from a naturally occurring source. The term "wild-type," when constructed with respect to a gene product, refers to a gene product that has the characteristics of a gene product isolated from a naturally occurring source. The term "naturally occurring," when applied to a subject, refers to the fact that the subject can be found in nature. For example, a polypeptide or polynucleotide sequence that is present in an organism (including viruses), that can be isolated from a natural source, and that has not been intentionally modified by man in a laboratory is naturally occurring. A wild-type gene is often the gene or allele that is most frequently observed in a population and is thus the randomly designed "normal" or "wild-type" form of that gene. In contrast, the terms "modified" or "mutant," when constructed with respect to a gene or gene product, refer to a gene or gene product that exhibits modifications (e.g., altered properties) in sequence and / or functional properties when compared to the wild-type gene or gene product, respectively. It should be noted that naturally occurring mutants can be isolated and identified by the fact that they have altered properties when compared to the wild-type gene or gene product.
[0084] The term "allele" refers to a variant of a gene, including, but not limited to, mutant forms and variants, polymorphic loci and single nucleotide polymorphic loci, frameshifts, and splice variants. Alleles can occur naturally in a population or can occur during the lifespan of a particular individual in any population.
[0085] Thus, the terms "mutant" and "variant," when used in reference to a nucleotide sequence, refer to a nucleotide sequence that differs from another, normally related, nucleotide sequence by one or more nucleotides. A "variant" is a difference between two different nucleotide sequences, typically one of the sequences being a reference sequence.
[0086] "Amplification" is specialized nucleic acid replication involving template specificity. This is in contrast to non-specific template replication (e.g., replication that is template-dependent but not dependent on a specific template). Template specificity is distinguished here from fidelity of replication (e.g., synthesis of the appropriate nucleotide sequence) and nucleotide (ribo- or deoxyribo-) specificity. Template specificity is frequently described in terms of "target" specificity. The target sequence is a "target" in the sense that it is sought to be sorted out from other nucleic acids. Amplification techniques are primarily designed for this sorting.
[0087] Nucleic acid amplification generally refers to the generation of multiple copies of a polynucleotide or portion of a polynucleotide, typically starting from small amounts of polynucleotide (e.g., a single polynucleotide molecule, 10-100 copies of a polynucleotide molecule, which may or may not be exactly identical), and the amplification product or amplicon is generally detectable. Polynucleotide amplification encompasses a variety of chemical and enzymatic processes. The generation of multiple DNA copies from one or a few copies of a target or template DNA molecule during polymerase chain reaction (PCR) or ligase chain reaction (LCR; see, e.g., U.S. Pat. No. 5,494,810, which is incorporated herein by reference in its entirety) is a form of amplification. Additional types of amplification include, but are not limited to, allele-specific PCR (see, e.g., U.S. Pat. No. 5,639,611, incorporated herein by reference in its entirety), assembly PCR (see, e.g., U.S. Pat. No. 5,965,408, incorporated herein by reference in its entirety), helicase-dependent amplification (see, e.g., U.S. Pat. No. 7,662,594, incorporated herein by reference in its entirety), hot-start PCR (see, e.g., U.S. Pat. Nos. 5,773,258 and 5,338,671, each incorporated herein by reference in its entirety), inter-sequence-specific PCR, inverse PCR (see, e.g., Triglia, et al. (1988) Nucleic Acids Res., 16:8186, incorporated herein by reference in its entirety), ligation-mediated PCR (see, e.g., Guilfoyle, R. et al., Nucleic Acids Res., 16:8186, incorporated herein by reference in its entirety), and ligation-mediated PCR (see, e.g., Guilfoyle, R. et al., Nucleic Acids Res., 16:8186, incorporated herein by reference in its entirety). Research, 25:1854-1858 (1997); U.S. Pat. No. 5,508,169, each of which is incorporated herein by reference in its entirety), methylation-specific PCR (see, e.g., Herman, et al., (1996) PNAS 93(13)9821-9826, each of which is incorporated herein by reference in its entirety), miniprimer PCR, multiplex ligation-dependent probe amplification (see, e.g., Schouten, et al., (2002) Nucleic Acids Research 30(12):e57, which is incorporated herein by reference in its entirety), multiplex PCR (see, e.g., Chamberlain, et al., (1988) Nucleic Acids Research 16(23)11141-11156; Ballabio, et al., (1990) Human Genetics 84(6)571-573; Hayden, et al., (2008) BMC Genetics 9:80, each of which is incorporated herein by reference in its entirety), nested PCR, overlap-extension PCR (see, e.g., Higuchi, et al., (1988) Nucleic Acids Research 16(15)7351-7367, which is incorporated herein by reference in its entirety), real-time PCR (see, e.g., Higuchi, et al., (1992) Biotechnology 10:413-417; Higuchi, et al., (1993) Biotechnology 11:1026-1030, each of which is incorporated herein by reference in its entirety), reverse transcription PCR (see, e.g., Bustin, SA (2000) J. Molecular Endocrinology 25:169-193, each of which is incorporated herein by reference in its entirety), solid-phase PCR, thermal asymmetric interlaced PCR, and touchdown PCR (see, e.g., Don, et al., Nucleic Acids Research (1991) 19(14) 4008; Roux, K. (1994) Biotechniques 16(5) 812-814; Hecker, et al. al., (1996) Biotechniques 20(3)478-485, each of which is incorporated herein by reference in its entirety. Polynucleotide amplification can also be achieved using digital PCR (see, e.g., Kalinina, et al., Nucleic Acids Research. 25; 1999-2004, (1997); Vogelstein and Kinzler, Proc Natl Acad Sci USA. 96; 9236-41, (1999); International Patent Application Publication No. 05023091A2; U.S. Patent Application Publication No. 20070202525, each of which is incorporated herein by reference in its entirety).
[0088] The term "polymerase chain reaction" ("PCR") refers to the method of K.B. Mullis, U.S. Pat. Nos. 4,683,195, 4,683,202, and 4,965,188, which describes a method for increasing the concentration of a segment of a target sequence in a mixture of genomic DNA without cloning or purification. This process for amplifying a target sequence consists of introducing a large excess of two oligonucleotide primers into a DNA mixture containing the desired target sequence, followed by a precise sequence of thermal cycling in the presence of a DNA polymerase. The two primers are complementary to each strand of the double-stranded target sequence. To effect amplification, the mixture is denatured, and the primers are then annealed to their complementary sequences within the target molecule. After annealing, the primers are extended with a polymerase to form a new set of complementary strands. The steps of denaturation, primer annealing, and polymerase extension can be repeated many times (i.e., denaturation, annealing, and extension constitute one "cycle," and there can be many "cycles") to obtain amplified segments of the desired target sequence at high concentrations. The length of the amplified segments of the desired target sequence is determined by the relative positions of the primers with respect to each other, and therefore, this length is a controllable parameter. Because of the repetitive aspect of the process, this method is referred to as "polymerase chain reaction" ("PCR"). Because the amplified segments of the desired target sequence become the predominant sequences (in terms of concentration) in the mixture, they are referred to as "PCR-amplified," and are "PCR products" and "amplicons."
[0089] Template specificity is achieved in most amplification techniques by the choice of enzyme. Amplification enzymes are enzymes that, under the conditions in which they are used, process only specific sequences of nucleic acid in a heterogeneous mixture of nucleic acids. For example, in the case of Q-beta replicase, MDV-1 RNA is the specific template for this replicase (Kacian et al., Proc. Natl. Acad. Sci. USA, 69:3038
[1972] ). Other nucleic acids are not replicated by this amplification enzyme. Similarly, in the case of T7 RNA polymerase, this amplification enzyme has stringent specificity for its own promoter (Chamberlin et al., Nature, 228:227
[1970] ). In the case of T4 DNA ligase, this enzyme will not ligate two oligonucleotides or polynucleotides where there is a mismatch between the oligonucleotide or polynucleotide substrate and the template at the ligation junction (Wu and Wallace (1989) Genomics 4:560). Finally, thermostable template-dependent DNA polymerases (e.g., Taq and Pfu DNA polymerases) are known to exhibit high specificity for the sequences bound to them, and thus defined by the primers, due to their ability to function at high temperatures, which create thermodynamic conditions that favor primer hybridization with the target sequence but not with non-target sequences (H.A. Erlich (ed.), PCR Technology, Stockton Press
[1989] ).
[0090] As used herein, the term "nucleic acid detection assay" refers to any method for determining the nucleotide composition of a nucleic acid of interest. Nucleic acid detection assays include, but are not limited to, DNA sequencing, probe hybridization, structure-specific cleavage assays (e.g., INVADER assay (Hologic, Inc.) and are described, e.g., in U.S. Pat. Nos. 5,846,717, 5,985,557, 5,994,069, 6,001,567, 6,090,543, and 6,872,816; Lyamichev et al., J. Am. Chem. Soc., 1999; 1998; 1999; 2000; 2001; 2002; 2003; 2004; 2005; 2006; 2007; 2008; 2009; 2010; 2011; 2012; 2013; 2014; 2015; 2016; 2017; 2018; 2019; 2020; 2021; 2022; 2023; 2024; 2025; 2026; 2027; 2028 ...9; 2030; 2031; 2032; 2033; 2034; 2035; 2036; 2037; 2040; 2041; 2042; 2043; 2 et al., Nat. Biotech., 17:292 (1999), Hall et al. al., PNAS, USA, 97:8272 (2000), and U.S. Patent Application Publication No. 2009 / 0253142, each of which is incorporated herein by reference in its entirety for all purposes), enzymatic mismatch cleavage methods (e.g., Variagenics, U.S. Pat. Nos. 6,110,684, 5,958,692, 5,851,770, which are incorporated herein by reference in their entireties), polymerase chain reaction, branched hybridization (e.g., Chiron, U.S. Pat. Nos. 5,849,481, 5,710,264, 5,124,246, and 5,624,802, which are incorporated herein by reference in their entireties), rolling circle replication (e.g., U.S. Pat. No. 6,210,884, Nos. 6,183,960, and 6,235,502, which are incorporated herein by reference in their entireties), NASBA (e.g., U.S. Pat. No. 5,409,818, which is incorporated herein by reference in its entirety), molecular beacon technology (e.g., U.S. Pat. No. 6,150,097, which is incorporated herein by reference in its entirety), E-sensor technology (Motorola, U.S. Pat. Nos. 6,248,229, 6,221,583, 6,013,170, and 6,063,573, which are incorporated herein by reference in their entireties), cycling probe technology (e.g., U.S. Pat. Nos. 5,403,711, 5,011,769, and 5,660,988, which are incorporated herein by reference in their entireties), Dade These include the Behring signal amplification method (e.g., U.S. Patent Nos. 6,121,001, 6,110,677, 5,914,230, 5,882,867, and 5,792,614, which are incorporated herein by reference in their entireties), the ligase chain reaction (e.g., Barnay Proc. Natl. Acad. Sci. USA 88,189-93 (1991)), and sandwich hybridization (e.g., U.S. Patent No. 5,288,609, which is incorporated herein by reference in its entirety).
[0091] The term "amplifiable nucleic acid" is used in reference to a nucleic acid that can be amplified by any amplification method. "Amplifiable nucleic acid" is generally intended to include a "sample template."
[0092] The term "sample template" refers to nucleic acid originating from a sample that is analyzed for the presence of a "target" (defined below). In contrast, "background template" is used to refer to nucleic acid other than the sample template, which may or may not be present in the sample. Background template is most often inadvertent. It may be the result of carryover or may be due to the presence of nucleic acid contaminants sought to be purified away from the sample. For example, nucleic acids from organisms other than those to be detected may be present as background in a test sample.
[0093] The term "primer" refers to an oligonucleotide, whether naturally occurring, as in a purified restriction digest, or synthetically produced, that can serve as a point of initiation for the synthesis of a primer extension product complementary to a nucleic acid strand when placed under conditions that induce the synthesis (e.g., in the presence of nucleotides and an inducing agent such as DNA polymerase, at a suitable temperature and pH). The primer is preferably single-stranded for maximum amplification efficiency, but may alternatively be double-stranded. If double-stranded, the primer is first treated to separate its strands before being used to prepare extension products. Preferably, the primer is an oligodeoxyribonucleotide. The primer must be sufficiently long to initiate the synthesis of extension products in the presence of an inducing agent. The exact length of the primer will depend on numerous factors, including temperature, source of primer, and the method used.
[0094] The term "probe" refers to an oligonucleotide (e.g., a nucleotide sequence) that is naturally occurring in a purified restriction digest product or that is produced synthetically, recombinantly, or by PCR amplification, and that is capable of hybridizing to another oligonucleotide of interest. Probes can be single-stranded or double-stranded. Probes are useful for the detection, identification, and isolation of specific gene sequences (e.g., "capture sequences"). It is contemplated that the probes used in the present invention will, in some embodiments, be labeled with a "reporter molecule" so as to be detectable by any detection system, including, but not limited to, enzymatic (e.g., ELISA, as well as enzyme-based histochemical assays), fluorescent, radioactive, and luminescent systems. It is not intended that the present invention be limited to a particular detection system or label.
[0095] As used herein, "methylation" refers to methylation of cytosine at the C5 or N4 position of cytosine, methylation of adenine, or methylation of the N6 position of other types of nucleic acids. Because typical in vitro DNA amplification methods do not preserve the methylation pattern of the amplification template, in vitro amplified DNA is usually unmethylated. However, "unmethylated DNA" or "methylated DNA" also refer to amplified DNA in which the original template is unmethylated or methylated, respectively.
[0096] Therefore, as used herein, "methylated nucleotide" or "methylated nucleotide base" refers to the presence of a methyl residue on a nucleotide base, and the methyl residue is not present in the recognized typical nucleotide base.For example, cytosine does not contain a methyl residue on its pyrimidine ring, but 5-methylcytosine contains a methyl residue at the 5th position of its pyrimidine ring.Therefore, cytosine is not a methylated nucleotide, but 5-methylcytosine is a methylated nucleotide.In another example, thymine contains a methyl residue at the 5th position of its pyrimidine ring, but for the purposes of this specification, thymine is not considered a methylated nucleotide when present in DNA, because thymine is a typical nucleotide base of DNA.
[0097] As used herein, a "methylated nucleic acid molecule" refers to a nucleic acid molecule that contains one or more methylated nucleotides.
[0098] As used herein, the "methylation state," "methylation profile," and "methylation status" of a nucleic acid molecule refer to the absence of one or more methylated nucleotide bases in a nucleic acid molecule. For example, a nucleic acid molecule containing a methylated cytosine is considered to be methylated (e.g., the methylation state of the nucleic acid molecule is methylated). A nucleic acid molecule that does not contain any methylated nucleotides is considered to be unmethylated.
[0099] The methylation state of a particular nucleic acid sequence (e.g., as described herein, a genetic marker or a DNA region) can indicate the methylation state of all bases in the sequence, or it can indicate the methylation state of a portion of the bases (e.g., one or more cytosines) within the sequence, or it can indicate information about the local methylation density within the sequence, with or without giving information about the exact location within the sequence where the methylation is present.
[0100] The methylation state of a nucleotide locus in a nucleic acid molecule refers to the presence or absence of a methylated nucleotide at a particular locus in the nucleic acid molecule. For example, when the nucleotide present at the seventh nucleotide in the nucleic acid molecule is 5-methylcytosine, the methylation state of the cytosine at the seventh nucleotide of the nucleic acid molecule is methylated. Similarly, when the nucleotide present at the seventh nucleotide in the nucleic acid molecule is cytosine (not 5-methylcytosine), the methylation state of the cytosine at the seventh nucleotide of the nucleic acid molecule is unmethylated.
[0101] Methylation status can optionally be expressed or indicated by a "methylation value" (e.g., representing a methylation frequency, fraction, ratio, percentage, etc.). Methylation values can be generated, for example, by quantifying the amount of intact nucleic acid following restriction digestion with a methylation-dependent restriction enzyme, or by comparing amplification profiles after a bisulfite reaction, or by comparing sequences of nucleic acid that have been treated with and without bisulfite. Thus, a value, e.g., a methylation value, represents a methylation status and can therefore be used as an indicator of the amount of methylation status across multiple copies of a locus. This is particularly useful when it is desirable to compare the methylation status of sequences in a sample to a threshold or reference value.
[0102] As used herein, "methylation frequency" or "percent methylation" refers to the number of instances where a molecule or locus is methylated compared to the number of instances where the molecule or locus is unmethylated.
[0103] Thus, a methylation state refers to the methylation state of a nucleic acid (e.g., a genomic sequence). Furthermore, a methylation state refers to the characteristics of a nucleic acid fragment at a specific genomic locus related to methylation. These characteristics include, but are not limited to, whether any cytosine (C) residues in the DNA sequence are methylated, the location of the methylated C residues, the frequency or percentage of methylated C residues in any particular region of the nucleic acid, and allelic differences in methylation factors, such as differences in allelic origin. The terms "methylation state," "methylation profile," and "methylation status" also refer to the relative concentration, absolute concentration, or pattern of methylated or unmethylated C residues in any particular region of a nucleic acid in a biological sample. For example, if a cytosine (C) sequence in a nucleic acid sequence is methylated, it can be referred to as having "hypermethylated" or "increased methylation," whereas if a cytosine (C) residue in a DNA sequence is unmethylated, it can be referred to as having "hypomethylated" or "decreased methylation." Similarly, if a cytosine (C) residue in a nucleic acid sequence is methylated when compared to other nucleic acid sequences (e.g., in a different region or different individual), the sequence is considered to be hypermethylated, or have increased methylation, compared to other nucleic acid sequences. Also, if a cytosine (C) residue in a DNA sequence is unmethylated when compared to other nucleic acid sequences (e.g., in a different region or different individual), the sequence is considered to be hypomethylated, or have decreased methylation, compared to other nucleic acid sequences. Furthermore, the term "methylation pattern," as used herein, refers to the collection of methylated and unmethylated nucleotides in a nucleic acid sequence. Two nucleic acids may have the same or similar methylation frequency or percentage, but may have different methylation patterns when the positions of the methylated and unmethylated nucleotides are different but the number of methylated and unmethylated nucleotides in a region is the same or similar.Sequences are referred to as "differentially methylated" or "having differences in methylation" or "different methylation states" when the extent, frequency, or pattern of methylation (some with increased or decreased methylation compared to others) differs. The term "differential methylation" refers to the difference in the level or pattern of nucleic acid methylation in a cancer-positive sample compared to the level or pattern of nucleic acid methylation in a cancer-negative sample. It can also refer to the difference in the level or pattern between patients whose cancer recurs after surgery and those who do not. Differential methylation and specific levels or patterns of DNA methylation are prognostic and predictive biomarkers, for example, and precise classification or predictive characteristics have been defined in the past.
[0104] Methylation state frequencies can be used to represent a group of individuals or a sample from a single individual. For example, a nucleotide locus with a 50% methylation state frequency is methylated in 50% of instances and unmethylated in 50% of instances. The frequencies can be used, for example, to represent the degree to which a nucleotide locus or nucleic acid region is methylated in a group of individuals or a collection of nucleic acids. Thus, when the methylation in a first group or pool of nucleic acid molecules differs from the methylation in a second group or pool of nucleic acid molecules, the methylation state frequency of the first group or pool can differ from the methylation state frequency of the second group or pool. The frequencies can also be used, for example, to represent the degree to which a nucleotide locus or nucleic acid region is methylated in a single individual. For example, the frequencies can be used to represent the degree to which a population of cells from a tissue sample is methylated or unmethylated at a nucleotide locus or nucleic acid region.
[0105] As used herein, "nucleotide locus" refers to the position of a nucleotide in a nucleic acid molecule. The nucleotide locus of a methylated nucleotide refers to the position of the methylated nucleotide in a nucleic acid molecule.
[0106] Typically, methylation of human DNA occurs on dinucleotide sequences containing adjacent guanine and cytosine, with the cytosine located 5' of the guanine (also referred to as CpG dinucleotide sequences). Most cytosines within CpG dinucleotides are methylated in the human genome, but some remain unmethylated in specific CpG dinucleotide-rich genomic regions known as CpG islands (see, e.g., Antequera et al. (1990) Cell 62:503-514).
[0107] As used herein, "CpG island" refers to a G:C-rich region of genomic DNA that contains an increased number of CpG dinucleotides compared to the total genomic DNA. A CpG island may be at least 100, 200, or more base pairs in length, the G:C content of the region is at least 50%, and the ratio of observed to expected CpG frequency is 0.6; in some instances, a CpG island may be at least 500 base pairs in length, the G:C content of the region is at least 55%, and the ratio of observed to expected CpG frequency is 0.65. The ratio of observed to expected CpG frequency can be calculated by the method set forth in Gardiner-Garden et al. (1987) J. Mol. Biol. 196: 261-281. For example, the observed CpG frequency relative to the expected frequency can be calculated by the formula R=(A×B) / (C×D), where R is the ratio of the observed CpG frequency to the expected frequency, A is the number of CpG dinucleotides in the sequence being analyzed, B is the total number of nucleotides in the sequence being analyzed, C is the total number of C nucleotides in the sequence being analyzed, and D is the total number of G nucleotides in the sequence being analyzed. Methylation state is typically determined in CpG islands, for example in promoter regions. Although other sequences in the human genome are susceptible to DNA methylation, such as CpA and CpT, they can be assessed (Ramsahoye (2000) Proc. Natl. Acad. Sci. USA 97:5237-5242; Salmon and Kaye (1970) Biochim. Biophys. Acta. 204:340-351; Grafstrom (1985) Nucleic Acids Res. 13:2827-2842; Nyce (1986) Nucleic Acids Acids Res. 14:4353-4367; see Woodcock (1987) Biochem. Biophys. Res. Commun. 145:888-894).
[0108] As used herein, a reagent that modifies the nucleotides of a nucleic acid molecule in correlation with the methylation state of the nucleic acid molecule, or a methylation-specific reagent, refers to a compound, composition, or other agent that can change the nucleotide sequence of a nucleic acid molecule in a manner that reflects the methylation state of the nucleic acid molecule. A method of treating a nucleic acid molecule with the reagent includes contacting the nucleic acid molecule with the reagent, optionally in combination with additional steps, to achieve the desired nucleotide sequence change. The change in the nucleotide sequence of the nucleic acid molecule can result in a nucleic acid molecule in which each methylated nucleotide is modified with a different nucleotide. The change in the nucleic acid nucleotide sequence can result in a nucleic acid molecule in which each unmethylated nucleotide is modified with a different nucleotide. The change in the nucleic acid nucleotide sequence can result in a nucleic acid molecule in which each selected unmethylated nucleotide (e.g., each unmethylated cytosine) is modified with a different nucleotide. The use of the reagent to alter a nucleic acid nucleotide sequence can result in a nucleic acid molecule in which each nucleotide that is a methylated nucleotide (e.g., each methylated cytosine) is modified with a different nucleotide. As used herein, the use of a reagent that modifies a selected nucleotide refers to a reagent that modifies one of the four nucleotides typically present in a nucleic acid molecule (C, G, T, and A for DNA, and C, G, U, and A for RNA), such as a reagent that modifies one nucleotide without modifying the other three nucleotides. In one exemplary embodiment, the reagent modifies an unmethylated selected nucleotide to generate a different nucleotide. In another exemplary embodiment, the reagent can deaminate an unmethylated cytosine nucleotide. An exemplary reagent is bisulfite.
[0109] As used herein, the term "bisulfite reagent" refers, in some embodiments, to a reagent comprising bisulfite, disulfite, hydrogen sulfite, or a combination thereof, for distinguishing between methylated and unmethylated cytidines, e.g., in CpG dinucleotide sequences.
[0110] The term "methylation assay" refers to an assay for determining the methylation state of one or more CpG dinucleotide sequences within a sequence of a nucleic acid.
[0111] The term "MS AP-PCR" (Methylation-Sensitive Arbitrarily-Primed Polymerase Chain Reaction) refers to an art-recognized technique that involves global scanning of the genome using CG-rich primers to align regions most likely to contain CpG dinucleotides, and is described in Gonzalgo et al. (1997) Cancer Research 57:594-599.
[0112] The term "MethyLight™" refers to the art-recognized fluorescence-based real-time PCR technology described in Eads et al. (1999) Cancer Res. 59:2302-2306.
[0113] The term "HeavyMethyl™" refers to an assay of methylation-specific blocking probes (also referred to herein as blockers) that cover CpG positions between or are covered by amplification primers that allow methylation-specific selective amplification of a nucleic acid sample.
[0114] The term "HeavyMethyl™ MethyLight™" assay refers to the HeavyMethyl™ MethyLight™ assay, which is a variant of the MethyLight™ assay that combines a methylation-specific blocking probe that covers the CpG positions between the amplification primers.
[0115] The term "Ms-SNuPE" (methylation-sensitive single nucleotide primer extension) refers to the art-recognized assay described in Gonzalgo & Jones (1997) Nucleic Acids Res. 25:2529-2531.
[0116] The term "MSP" (methylation-specific PCR) refers to the art-recognized methylation assay described in Herman et al. (1996) Proc. Natl. Acad. Sci. USA 93:9821-9826 and U.S. Pat. No. 5,786,146.
[0117] The term "COBRA" (Combined Bisulfate Restriction Analysis) refers to an art-recognized methylation assay described in Xiong & Laird (1997) Nucleic Acids Res. 25:2532-2534.
[0118] The term "MCA" (methylated CpG island amplification) refers to the methylation assay described in Toyota et al. (1999) Cancer Res. 59: 2307-12, and WO 00 / 26401A1.
[0119] As used herein, a "selected nucleotide" refers to one of the four nucleotides typically occurring in nucleic acid molecules (C, G, T, and A for DNA and C, G, U, and A for RNA), and can include methylated derivatives of a typically occurring nucleotide (e.g., when C is a selected nucleotide, both methylated and unmethylated C are included in the meaning of the selected nucleotide), with a methylated selected nucleotide specifically referring to a methylated typically occurring nucleotide, and an unmethylated selected nucleotide specifically referring to an unmethylated typically occurring nucleotide.
[0120] The term "methylation-specific restriction enzyme" or "methylation-selective restriction enzyme" refers to an enzyme that selectively digests nucleic acids depending on the methylation state of its recognition site. In the case of a restriction enzyme that specifically cleaves when the recognition site is unmethylated or hemimethylated, cleavage may not occur or may occur with significantly reduced efficiency when the recognition site is methylated. In the case of a restriction enzyme that specifically cleaves when the recognition site is methylated, cleavage may not occur or may occur with significantly reduced efficiency when the recognition site is unmethylated. Preferably, the methylation-specific restriction enzyme has a recognition sequence containing a CG dinucleotide (e.g., a recognition sequence such as CGCG or CCCGGG). In some preferred embodiments, the restriction enzyme does not cleave when the cytosine in the dinucleotide is methylated at the carbon atom C5.
[0121] As used herein, a "different nucleotide" refers to a nucleotide that is chemically different from the selected nucleotide; typically, the different nucleotide has different Watson-Crick base pairing properties than the selected nucleotide; and the typically occurring nucleotide that is complementary to the selected nucleotide is not identical to the typically occurring nucleotide that is complementary to the different nucleotide. For example, when C is the selected nucleotide, U or T can be the different nucleotide, as exemplified by the complementarity of C to G and U or T to A. As used herein, a nucleotide that is complementary to a selected nucleotide or a different nucleotide refers to a nucleotide that base pairs with the selected nucleotide or different nucleotide under high stringency conditions with a higher affinity than the base pairing of a complementary nucleotide with three of the four typically occurring nucleotides. One example of complementarity is Watson-Crick base pairing in DNA (e.g., AT and CG) and RNA (e.g., AU and CG). Thus, for example, under high stringency conditions, G base pairs with a higher affinity to C than it does to G, A, or T, and therefore when C is the selected nucleotide, G is the complementary nucleotide to the selected nucleotide.
[0122] As used herein, the "sensitivity" of a given marker refers to the proportion of samples reporting DNA methylation values above a threshold that distinguishes between neoplastic and non-neoplastic samples. In some embodiments, a positive is defined as a neoplasm identified in tissue reporting a DNA methylation value above the threshold (e.g., in a disease-associated range), and a false negative is defined as a neoplasm identified in tissue reporting a DNA methylation value below the threshold (e.g., in a non-disease-associated range). The sensitivity value therefore reflects the probability that a DNA methylation measurement of a given marker obtained from a known disease sample will fall within the disease range associated with the measurement. As defined herein, the clinical relevance of a calculated sensitivity value represents an estimate of the probability that a given marker, when used in symptomatic subjects, will detect the presence of a clinical condition.
[0123] As used herein, the "specificity" of a given marker refers to the proportion of non-neoplastic samples reporting DNA methylation values below a threshold that distinguishes between neoplastic and non-neoplastic samples. In some embodiments, a negative is defined as a tissue-ascertained non-neoplastic sample reporting a DNA methylation value below the threshold (e.g., a non-disease-associated range), and a false positive is defined as a tissue-ascertained non-neoplastic sample reporting a DNA methylation value above the threshold (e.g., a disease-associated range). The specificity value therefore reflects the probability that a DNA methylation measurement of a given marker obtained from a known non-disease sample will fall within the non-disease range associated with the measurement. As defined herein, the clinical relevance of a calculated specificity value represents an estimate of the probability that a given marker will detect the absence of clinical symptoms when used in asymptomatic patients.
[0124] The term "AUC" as used herein is an abbreviation for "area under a curve." In particular, AUC refers to the area under a receiver operating characteristic (ROC) curve. An ROC curve is a plot of the true positive rate against the false positive rate for different possible cut points of a diagnostic test. The ROC curve shows a trade-off between sensitivity and specificity depending on the cut point selected (any increase in sensitivity may be accompanied by a decrease in specificity). The area under the ROC curve (AUC) is a measure of the accuracy of a diagnostic test (the larger the area, the better; the best is 1; a random test may have a ROC curve lying on the diagonal with an area of 0.5; see: JP Egan. (1975) Signal Detection Theory and ROC Analysis, Academic Press, New York).
[0125] As used herein, the term "neoplasm" refers to an "abnormal mass or excessive growth of tissue, not in harmony with the growth of normal tissue." See, e.g., "The Spread of Tumors in the Human Body," London, Butterworth & Co, 1952.
[0126] As used herein, the term "adenoma" refers to a benign tumor of glandular origin. These growths are benign but, over time, can evolve into malignant tumors.
[0127] The terms "pre-cancerous" or "pre-neoplastic" and their cognates refer to any cell proliferative disorder undergoing malignant transformation.
[0128] A "site" of a neoplasm, adenoma, cancer, etc. is a tissue, organ, cell type, anatomical region, body part, etc. in a subject's body in which the neoplasm, adenoma, cancer, etc. is located.
[0129] As used herein, a "diagnostic" testing method includes detecting or identifying a disease state or condition in a subject, determining the likelihood that a subject will be infected with a given disease or condition, determining the likelihood that a subject with a disease or condition will respond to treatment, determining the prognosis (or likely progression or regression) of a subject with a disease or condition, and determining the effectiveness of treatment for a subject with a disease or condition. For example, diagnostics can be used to detect the presence or likelihood of a subject being infected with a neoplasm or the likelihood that such a subject will respond favorably to a compound (e.g., a pharmaceutical, e.g., a drug) or other treatment.
[0130] The term "marker," as used herein, refers to a substrate (e.g., a nucleic acid or region of a nucleic acid) that can diagnose cancer by distinguishing cancer cells from normal cells, for example, based on its methylation state.
[0131] The term "isolated," when used in reference to a nucleic acid, such as in an "isolated oligonucleotide," refers to a nucleic acid sequence that has been identified and separated from at least one contaminant nucleic acid with which it is normally associated in its natural source. An isolated nucleic acid exists in a form or context that is different from that in which it is found in nature. In contrast, non-isolated nucleic acids, such as DNA and RNA, are found in the state in which they occur in nature. Examples of non-isolated nucleic acids include a given DNA sequence (e.g., a gene) found on a host cell chromosome in close proximity to adjacent genes, and an RNA sequence, such as a particular mRNA sequence encoding a particular protein, found within a cell in a mixture with many other mRNAs encoding multiple proteins. However, an isolated nucleic acid encoding a particular protein includes, by way of example, a nucleic acid in a cell that normally expresses that protein, such that the nucleic acid is in a chromosomal location different from that of natural cells or is otherwise flanked by nucleic acid sequences different from those in which it is found in nature. An isolated nucleic acid or oligonucleotide can exist in single-stranded or double-stranded form. When an isolated nucleic acid or oligonucleotide is used to express a protein or oligonucleotide, the oligonucleotide will at least contain a sense or coding strand (i.e., the oligonucleotide is single-stranded), but may contain both a sense and an antisense strand (i.e., the oligonucleotide can be double-stranded). The isolated nucleic acid may be combined with other nucleic acids or molecules after being isolated from its natural or typical environment. For example, the isolated nucleic acid may be present in a host cell, for example, in heterologous expression.
[0132] The term "purified" refers to either a nucleic acid or amino acid sequence molecule that has been removed, isolated, or separated from its natural environment. An "isolated nucleic acid sequence" may therefore be a purified nucleic acid sequence. "Substantially purified" molecules are at least 60% free, preferably at least 75% free, and more preferably at least 90% free from other components with which they are naturally associated. As used herein, the terms "purified" or "to purify" also refer to the removal of contaminants from a sample. Removal of contaminating proteins results in an enrichment of the polypeptide or nucleic acid of interest in the sample. In another example, a recombinant polypeptide is expressed in a plant, bacterial, yeast, or mammalian host cell, and the polypeptide is purified by removal of host cell proteins, thereby enriching the recombinant polypeptide in the sample.
[0133] The term "composition comprising" a given polynucleotide sequence or polypeptide refers broadly to any composition that contains the given polynucleotide sequence or polypeptide. The composition may include an aqueous solution containing salts (e.g., NaCl), detergents (e.g., SDS), and other components (e.g., Denhardt's solution, dried milk, salmon sperm DNA, etc.).
[0134] The term "sample" is used in its broadest sense. In one sense, it may refer to animal cells or tissue. In another sense, it is meant to include specimens or cultures obtained from any source, as well as biological and environmental samples. Biological samples can be obtained from plants or animals (including humans) and can encompass fluids, solids, tissues, and gases. In some embodiments, the sample is a plasma sample. In some embodiments, the sample is a gastric tissue sample. In some embodiments, the sample is a stool sample. Environmental samples include environmental materials such as surface material, soil, water, and industrial samples. These examples should not be construed as limiting the sample types applicable to the present invention.
[0135] As used herein, the term "remote sample," when used in some contexts, refers to a sample that is indirectly collected from a site other than the source of the sample's cells, tissues, or organs. For example, when sample material of pancreatic origin is assessed in a stool sample (e.g., not from a sample taken directly from the pancreas), the sample is a remote sample.
[0136] As used herein, the term "patient" or "subject" refers to an organism undergoing various tests administered by a technique. The term "subject" includes animals, preferably mammals, including humans. In preferred embodiments, the subject is a primate. In even more preferred embodiments, the subject is a human.
[0137] As used herein, the term "kit" refers to any delivery system for delivering materials. In the context of a reaction assay, such a delivery system includes a system that allows for the storage, transport, or delivery of reaction reagents (e.g., oligonucleotides, enzymes, etc. in appropriate containers) and / or supporting materials (e.g., buffers, instructions for conducting the assay, etc.) from one location to another. For example, a kit includes one or more containers (e.g., boxes) containing the relevant reaction reagents and / or supporting materials. As used herein, the term "fragmented kit" refers to a delivery system that includes two or more separate containers, each containing a portion of the overall kit components. The containers can be delivered to the intended recipient together or separately. For example, a first container can contain an enzyme for use in an assay, while a second container contains oligonucleotides. The term "fragmented kit" is intended to encompass, but is not limited to, kits containing analyte-specific reagents (ASRs) regulated by Section 520(e) of the Federal Food, Drug, and Cosmetic Act. Indeed, any delivery system containing two or more separate containers, each containing a portion of the total kit components, is included in the term "fragmented kit." In contrast, a "combined kit" refers to a delivery system containing all components of a reaction assay in a single container (e.g., in a single container housing each of the desired components). The term "kit" includes both fragmented and composite kits.
[0138] [Embodiments of the present technology] Provided herein are techniques for detecting tumorigenesis, and in particular (but not limited to), methods, compositions, and related uses for detecting precancerous conditions and malignant tumors, such as gastric cancer. Markers were identified in a case-control study by comparing the methylation status of DNA markers from tumors of subjects with gastric cancer (e.g., gastric cancer) with the methylation status of the same DNA markers from control subjects (see Examples 1 and 2). Additional experiments identified 10 optimal markers for gastric cancer detection (ARHGEF4, ELMO1, ABCB1, CLEC11A, ST8SIA1, SFMBT2, CD1D, CYP26C1, ZNF569, C13ORF18; see Example 2 and Table 4). Further experiments identified 30 optional markers for gastric cancer detection (see Example 1 and Table 2). Further experiments identified 12 optimal markers for detecting gastric cancer (e.g., gastric cancer) in plasma samples (ELMO1, ZNF569, C13 or f18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4, and ST8SIAI; see Example 3, Tables 1, 2, 4, and 7).
[0139] Markers and / or panels of markers (e.g., chromosomal regions containing markers (see Table 4) with annotations selected from ARHGEF4, ELMO1, ABCB1, CLEC11A, ST8SIAI, SFMBT2, CD1D, CYP26C1, ZNF569, or 13ORF18) were identified in a case-control approach by comparing the methylation status of stomach tissue from subjects with gastric cancer (e.g., gastric cancer) with the methylation status of the same DNA markers from control subjects (see Example 2).
[0140] Markers and / or panels of markers (e.g., chromosomal regions containing markers having an annotation selected from ELMO1, ARHGEF4, EMX1, SP9, CLEC11A, ST8SIA1, BMP3, KCNA3, DMRTA2, LCNK12, SD1D, PRLCB, CYP26C1, ZNF568, ABCB1, ELOVL2, PKIA, SFMBT2(839), PCBP3, MATK, GRN2D, NDRG4, DLX4, PPP2R5C, FGF14, ZNF132, CHST2(7890), FL11, c13 or f18, or ZNF569 (see Table 2) were identified in a case-control approach by comparing the methylation status of stomach tissue from subjects with gastric cancer (e.g., gastric cancer) with the methylation status of the same DNA markers from control subjects (see Example 1).
[0141] Markers and / or panels of markers (e.g., chromosomal regions containing markers having annotations selected from ELMO1, ZNF569, C13 or f18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4, ST8SIA1 (see Tables 1, 2, 4, and 7) were identified in a case-control approach by comparing the methylation status of DNA markers (from gastric plasma of subjects with gastric cancer (e.g., gastric cancer)) with the methylation status of the same DNA markers from control subjects (see Example 3).
[0142] Additionally, the technology includes chromosomal regions that contain a panel of various markers, for example, in some embodiments, the markers are annotated as and include ARHGEF4, ELMO1, ABCB1, CLEC11A, ST8SIA1, SFMBT2, CD1D, CYP26C1, ZNF569, or C13ORF18 (see Table 4). Additionally, embodiments provide methods for analyzing DMRs in Table 4, including DMR numbers 49, 43, 196, 66, 1, 237, 249, 250, 251, or 252. In some embodiments, measurements of the methylation status of markers are provided, wherein a chromosomal region is annotated as and comprises the markers ELMO1, ARHGEF4, EMX1, SP9, CLEC11A, ST8SIA1, BMP3, KCNA3, DMRTA2, KCNK12, CD1D, PRKCB, CYP26C1, ZNF568, ABCB1, ELOVL2, PKIA, SFMBT2(893), PCBP3, MATK, GRN2D, NDRG4, DLX4, PPP2R5C, FGF14, ZNF132, CHST2(7890), FLI1, c13 or fl8, or ZNF569 (see Table 2). In further embodiments, specific methods of analyzing DMRs from Table 2 provide methods for analyzing DMRs in Table 2 having DMR numbers 253, 251, 254, 255, 256, 249, 257, 258, 259, 260, 261, 262, 250, 263, 1, 264, 265, 196, 266, 118, 267, 268, 269, 270, 271, 272, 46, 273, 252, or 237. In some embodiments, the obtained sample is a plasma sample, and the markers are annotated as ELMO1, ZNF569, C13C18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4, and ST8SIA1, and include chromosomal regions containing the markers.Further, in such embodiments, the sample is a plasma sample and the method analyzes a DMR of Tables 1, 2, 4, and 7 that is DMR number 253, 251, 254, 255, 256, 249, 257, 258, 259, 260, 261, 262, 250, 263, 1, 264, 265, 196, 266, 118, 267, 268, 269, 270, 271, 272, 46, 273, 252, or 237. In some embodiments, the method comprises measuring the methylation status of two markers (e.g., a pair of markers set forth in a column of Table 1, 2, 4, or 7).
[0143] Although the disclosure herein refers to certain illustrated embodiments, it will be understood that these embodiments are presented by way of example and not by way of limitation.
[0144] In certain aspects, the present technology provides compositions and methods for identifying, detecting, and / or classifying cancers, such as gastric cancer (e.g., gastric cancer). The methods involve detecting the methylation status of at least one methylation marker in a biological sample isolated from a subject (e.g., a stool sample, a gastric tissue sample, a plasma sample), where a change in the methylation status of the marker is indicative of the presence, classification, or location of gastric cancer (e.g., gastric cancer). Certain embodiments relate to markers containing a methylation variable region (DMR, e.g., DMR1-274, Tables 1, 2, 4, and 7) for use in the diagnosis (e.g., screening) of neoplastic cells in neoplastic cell proliferative disorders (e.g., cancer), including early detection during precancerous stages of the disease.
[0145] In addition to embodiments in which the methylation of at least one marker, region of a marker, or base of a marker comprising a DMR shown herein and listed in Tables 1, 2, 4, or 7 (e.g., DMR DMR1-274) is analyzed, the present technology also provides panels of markers comprising at least one marker, region of a marker, or base of a marker comprising a DMR that have utility in the detection of cancer, particularly gastric cancer.
[0146] Some embodiments of the technology are based on the analysis of the CpG methylation status of at least one marker, region of a marker, or base of a marker.
[0147] In some embodiments, the present technology provides for the use of bisulfite techniques in combination with one or more methylation assays that measure the methylation status of CpG dinucleotide sequences in at least one marker comprising a DMR (e.g., DMR1-274, see Tables 1, 2, 4, and 7). Genomic CpG dinucleotides can be methylated or demethylated (alternatively known as up- and down-methylation, respectively). However, the methods of the present invention are suitable for analyzing biological material of heterogeneous nature (e.g., low concentrations of tumor cells or biological material derived therefrom) in the background of a separate sample (e.g., blood, tissue excreta, or stool).
[0148] Thus, when analyzing the methylation status of CpG positions in such samples, quantitative assays can be used to detect the level (e.g., percent, fraction, ratio, population, degree) of methylation at a particular CpG position.
[0149] According to this technology, measuring the methylation status of CpG dinucleotide sequences in markers containing DMRs has utility in both the diagnosis and characterization of cancers such as gastric cancer.
[0150] [Marker combinations] In some embodiments, the technology relates to assessing the methylation status of a combination of markers including those in Table 1 (e.g., DMR numbers 1-248), or Table 4 (e.g., DMR numbers 49, 43, 196, 66, 1, 237, 249, 250, 251, 252), or Table 2 (e.g., DMR numbers 253, 251, 254, 255, 256, 249, 257, 258, 259, 260, 261, 262, 250, 263, 1, 264, 265, 196, 266, 118, 267, 268, 269, 270, 271, 272, 46, 273, 252, 237), or Table 7 (DMR number 274) or additional markers comprising a DMR. In some embodiments, assessing the methylation status of one or more markers improves the specificity and / or sensitivity of screening or diagnosing for a suspected tumor (e.g., gastric cancer) in a subject. In some embodiments, a marker or combination of markers distinguishes between tumor types and / or locations.
[0151] Different cancers are predicted by different combinations of markers (e.g., identified by statistical techniques for the specificity and sensitivity of prediction). The present technology provides methods for determining predictive and validated combinations for several cancers.
[0152] [Method for analyzing methylation state] The most commonly used method for analyzing nucleic acids for the presence of 5-methylcytosine is based on the bisulfite method described by Frommer et al. (1992) Proc. Natl. Acad. Sci. USA 89:1827-31, explicitly incorporated herein by reference in its entirety for all purposes), or a variant thereof, for the detection of 5-methylcytosine in DNA. The bisulfite method for detecting 5-methylcytosine is based on the observation that non-5-methylcytosine cytosines react with hydrogen sulfite ions (also known as bisulfite). The reaction typically involves the following steps: First, cytosine reacts with hydrogen sulfite to form sulfonated cytosine. Next, spontaneous deamination of the sulfonated intermediate results in sulfonated uracil. Finally, the sulfonated uracil is desulfonated under alkaline conditions to form uracil. Detection is possible because uracil base pairs with adenine (behaving like thymine), while 5-methylcytosine base pairs with guanine (behaving like cytosine), allowing for differentiation of methylated from unmethylated cytosine by, for example, bisulfite gene sequencing (Grigg G, & Clark S, Bioessays (1994) 16: 431-36; Grigg G, DNA Seq. (1996) 6: 189-98) or methylation-specific PCR (MSP), as disclosed, for example, in U.S. Pat. No. 5,786,146.
[0153] Some conventional techniques involve enclosing the DNA to be analyzed in an agarose membrane, thereby preventing DNA diffusion and renaturation (bisulfite reacts only with single-stranded DNA), and replacing the precipitation and purification steps with a rapid dialysis (Olek A, et al. (1996) "A modified and improved method for bisulfite-based cytosine methylation analysis" Nucleic Acids Res. 24: 5064-6). Thus, it is possible to analyze individual cells for methylation status, demonstrating the usefulness and sensitivity of the method. An overview of conventional methods for detecting 5-methylcytosine is given by Rein, T., et al. (1998) Nucleic Acids Res. 26: 2255.
[0154] Bisulfite techniques typically involve amplifying short, specific fragments of a known nucleic acid after bisulfite treatment and then analyzing the products by sequencing to analyze the positions of individual cytosines (Olek & Walter (1997) Nat. Genet. 17: 275-6), or primer extension reactions (Gonzalgo & Jones (1997) Nucleic Acids Res. 25: 2529-31; WO 95 / 00669; U.S. Pat. No. 6,251,594). Some methods use enzymatic digestion (Xiong & Laird (1997) Nucleic Acids Res. 25: 2532-4). Detection by hybridization has also been described in the art (Olek et al., WO 99 / 28498). Furthermore, the use of bisulfite technology for methylation detection on individual genes has been described (Grigg & Clark (1994) Bioessays 16: 431-6; Zeschnigk et al. (1997) Hum Mol Genet. 6: 387-95; Feil et al. (1994) Nucleic Acids Res. 22: 695; Martin et al. (1995) Gene 157: 261-4; WO 9746705; WO 9515373).
[0155] Various methylation analysis methods are known in the art and can be used in conjunction with bisulfite treatment according to the present technique. These analyses can determine the methylation state of one or more CpG dinucleotides (e.g., CpG islands) within a nucleic acid sequence. Such analyses include, among other techniques, sequencing of bisulfite-treated nucleic acids, PCR (for sequence-specific amplification), Southern blot analysis, and the use of methylation-sensitive restriction enzymes.
[0156] For example, gene sequencing has been amenable to analysis of methylation patterns, and 5-methylcytosine is distinguished by using bisulfite treatment (Frommer et al. (1992) Proc. Natl. Acad. Sci. USA 89: 1827-1831). Furthermore, restriction enzyme digestion of PCR products amplified from bisulfite-converted DNA finds use in assessing methylation state, as described, for example, by Sadri & Hornsby (1997) Nucl. Acids Res. 24: 5058-5059, or as performed in a method known as COBRA (combined bisulfate restriction analysis) (Xiong & Laird (1997) Nucleic Acids Res. 25: 2532-2534).
[0157] The COBRA™ assay is a quantitative methylation assay useful for determining DNA methylation levels at specific loci in small amounts of genomic DNA (Xiong & Laird, Nucleic Acids Res. 25:2532-2534, 1997). Briefly, restriction enzyme digestion is used to reveal methylation-dependent sequence differences in PCR products of sodium bisulfite-treated DNA. Methylation-dependent sequence differences are first introduced into genomic DNA by standard bisulfite treatment according to the method described by Frommer et al. (Proc. Natl. Acad. Sci. USA 89:1827-1831, 1992). PCR amplification of the bisulfite-converted DNA is then performed using specific primers for the CpG island of interest, followed by restriction endonuclease digestion, gel electrophoresis, and detection with a specific labeled hybridization probe. Methylation levels in the original DNA sample are represented by the relative amounts of digested and undigested PCR products in a linear quantification across a broad spectrum of DNA methylation levels. Furthermore, this technique can be reliably applied to DNA obtained from microdissected paraffin-embedded tissue samples.
[0158] Typical reagents for COBRA™ analysis (e.g., as might be found in a typical COBRA™-based kit) may include, but are not limited to, PCR primers for specific loci (e.g., specific genes, markers, DMRs, regions of genes, regions of markers, bisulfite-treated DNA sequences, CpG islands, etc.); restriction enzymes and appropriate buffers; gene hybridization oligonucleotides; control hybridization oligonucleotides; kinase labeling kits for oligonucleotide probes; and labeled nucleotides. Additionally, bisulfite conversion reagents may include DNA denaturing buffers, sulfonation buffers, DNA repair reagents or kits (e.g., precipitation, ultrafiltration, affinity columns); desulfonation buffers, and DNA repair compositions.
[0159] Preferably, assays such as "MethyLight™" (fluorescence-based real-time PCR technology) (Eads et al., Cancer Res. 59:2302-2306, 1999), Ms-SNuPE™ (methylation-sensitive single nucleotide primer extension) reactions (Gonzalgo & Jones, Nucleic Acids Res. 25:2529-2531, 1997), methylation-specific PCR ("MSP"; Herman et al., Proc. Natl. Acad. Sci. USA 93:9821-9826, 1996; US Pat. No. 5,786,146), and methylated CpG island amplification ("MCA"; Toyota et al., Cancer Res. 59:2307-12, 1999) are used alone or in combination with one or more of these methods.
[0160] The "HeavyMethyl™" assay technology is a quantitative method for assessing methylation differences based on methylation-specific amplification of bisulfite-treated DNA. Methylation-specific blocking probes ("blockers") covering the CpG positions between or covered by the amplification primers enable methylation-specific selective amplification of nucleic acid samples.
[0161] The term "HeavyMethyl™ MethyLight™" assay refers to the HeavyMethyl™ MethyLight™ assay, which is a variant of the MethyLight™ assay, in which the MethyLight™ assay combines a methylation-specific blocking probe that covers the CpG positions between the amplification primers. The HeavyMethyl™ assay may also be used in combination with methylation-specific amplification primers.
[0162] Typical reagents for HeavyMethyl™ analysis (e.g., as may be found in a typical MethyLight™-based kit) may include, but are not limited to, PCR primers for a specific locus (e.g., a specific gene, marker, DMR, region of a gene, region of a marker, bisulfite-treated DNA sequence, CpG island, or bisulfite-treated DNA sequence or CpG island, etc.); oligonucleotide blocking; optimized PCR buffer and deoxynucleotides; and Taq polymerase.
[0163] MSP (methylation-specific PCR) can assess the methylation status of virtually any group of CpG sites within a CpG island, without the need for methylation-sensitive restriction enzymes (Herman et al. Proc. Natl. Acad. Sci. USA 93:9821-9826, 1996; US Pat. No. 5,786,146). Briefly, DNA is modified with sodium bisulfite to convert unmethylated but not methylated cytosines to uracil, and the products are sequentially amplified with primers specific for methylated versus unmethylated DNA. MSP requires only small amounts of DNA, has a sensitivity of 0.1% of the methylated alleles at a given CpG island locus, and can be performed on DNA extracted from paraffin-embedded samples. Typical reagents for MSP analysis (e.g., as might be found in a typical MSP-based kit) may include, but are not limited to, methylated and unmethylated PCR primers for specific loci (e.g., specific genes, markers, DMRs, regions of genes, regions of markers, bisulfite-treated DNA sequences, CpG islands, etc.); optimized PCR buffers and deoxynucleotides, and specific probes.
[0164] The MethyLight™ assay is a high-throughput quantitative methylation assay that utilizes fluorescence-based real-time PCR (e.g., TaqMan®) without the need for further manipulation after the PCR step (Eads et al., Cancer Res. 59:2302-2306, 1999). Briefly, the MethyLight™ process begins with a mixed sample of genomic DNA, which is converted by standard methods into a mixed pool with methylation-dependent sequence differences in a sodium bisulfite reaction (the bisulfite step converts unmethylated cytosine residues to uracil). Fluorescence-based PCR is then performed in a "biased" reaction, for example, with PCR primers that overlap known CpG dinucleotides. Sequence discrimination occurs both at the level of the amplification step and the level of the fluorescence detection step.
[0165] The MethyLight™ assay is used as a quantitative test for methylation patterns in nucleic acids (e.g., genomic DNA samples), with sequence discrimination occurring at the level of probe hybridization. In the quantitative version, PCR reactions provide methylation-specific amplification in the presence of fluorescent probes that overlap specific putative methylation sites. An unbiased control for DNA input is provided by reactions in which neither the primers nor the probe overlap any CpG dinucleotides. Quantitative testing of genomic methylation can also be achieved by probing biased PCR pools with either control nucleotides that do not cover known methylation sites (e.g., fluorescent-based versions of HeavyMethyl™ and MSP technology) or oligonucleotides that cover potential methylation sites.
[0166] The MethyLight™ process can be used with any desired probe (e.g., TaqMan® probe, Lightcycler® probe, etc.). For example, in some applications, double-stranded genomic DNA is treated with sodium bisulfite and targeted to one of two sets of probes in a TaqMan® PCR reaction, e.g., with an MSP primer and / or a HeavyMethyl blocker oligonucleotide and a TaqMan® probe. TaqMan® probes are doubly labeled with fluorescent "reporter" and "quencher" molecules and are designed to be specific for regions of relatively high GC content, so they melt at a temperature approximately 10°C higher during PCR cycles than the forward or reverse primers. This allows the TaqMan® probe to remain fully hybridized during the PCR annealing / extension step. Taq polymerase can enzymatically synthesize new strands during PCR, ultimately resulting in the annealed TaqMan® probe. Taq polymerase 5'-3' endonuclease activity may then be displaced onto the TaqMan® probe by digesting it, releasing a fluorescent reporter molecule for quantitative detection of an unquenched signal using a real-time fluorescence detection system.
[0167] Typical reagents for MethyLight™ analysis (e.g., as may be found in a typical MethyLight™-based kit) may include, but are not limited to, PCR primers for specific loci (e.g., specific genes, markers, DMRs, regions of genes, regions of markers, bisulfite-treated DNA sequences, CpG islands, etc.); TaqMan™ or Lightcycer™ probes; optimized PCR buffer and deoxynucleotides; and Taq polymerase.
[0168] The QM™ (Quantitative Methylation) Assay is an alternative quantitative test for methylation patterns in genomic DNA samples, in which sequence discrimination occurs at the level of probe hybridization. In this quantitative version, PCR reactions provide unbiased amplification in the presence of fluorescent probes that overlap specific putative methylation sites. An unbiased control for DNA input is provided by reactions in which neither the primers nor the probe overlap any CpG dinucleotides. Quantitative testing of genomic methylation is also achieved by probing biased PCR pools with either control oligonucleotides that do not cover known methylation sites (fluorescence-based versions of HeavyMethyl™ and MSP technology) or oligonucleotides that cover potential methylation sites.
[0169] The QM™ process can use any suitable probe (e.g., TaqMan® probe, Lightcycler® probe) in the amplification step. For example, double-stranded genomic DNA is treated with sodium bisulfite and exposed to an unbiased primer and TaqMan® probe. The TaqMan® probe can remain fully hybridized during the PCR annealing / extension step. During PCR, Taq polymerase can eventually access the annealed TaqMan® probe as it enzymatically synthesizes a new strand. The Taq polymerase 5'-3' endonuclease activity can then displace the TaqMan® by digesting it, releasing a fluorescent reporter molecule for quantitative detection of an unquenched signal using a real-time fluorescence detection system. Typical reagents for QM™ analysis (e.g., as might be found in a typical QM™-based kit) may include, but are not limited to, PCR primers for specific loci (e.g., specific genes, markers, DMRs, regions of genes, regions of markers, bisulfite-treated DNA sequences, CpG islands, etc.); TaqMan® or Lightcycer® probes; optimized PCR buffer and deoxynucleotides; and Taq polymerase.
[0170] The Ms-SNuPE™ technology is a quantitative method for assessing methylation differences at specific CpG sites based on bisulfite treatment of DNA followed by single-nucleotide primer extension (Gonzalgo & Jones, Nucleic Acids Res. 25:2529-2531, 1997). Briefly, genomic DNA is reacted with sodium bisulfite to convert unmethylated cytosines to uracil, leaving 5-methylcytosines unchanged. Amplification of the desired target sequence is then performed using PCR specific to the bisulfite-converted DNA, and the resulting product is isolated and used as a template for methylation analysis at the CpG sites of interest. This method allows for analysis of small amounts of DNA (e.g., microdissected pathology sections) and avoids the use of restriction enzymes to determine the methylation status at CpG sites.
[0171] Typical reagents for Ms-SNuPE™ analysis (e.g., as found in a typical Ms-SNuPE™-based kit) may include, but are not limited to, PCR primers for specific loci (e.g., specific genes, markers, DMRs, regions of genes, regions of markers, bisulfite-treated DNA sequences, CpG islands, etc.); optimized PCR buffers and deoxynucleotides; gel extraction kits; positive control primers; Ms-SNuPE™ primers for specific loci; reaction buffers (Ms-SNuPE reactions); and labeled nucleotides. Additionally, bisulfite conversion reagents may include DNA denaturing buffers; sulfonation buffers; DNA repair reagents or kits (e.g., precipitation, ultrafiltration, affinity columns); desulfonation buffers; and DNA repair compositions.
[0172] Reduced bisulfite sequencing (RRBS) begins with bisulfite treatment of nucleic acids to convert all unmethylated cytosines to uracils, followed by restriction enzyme digestion (e.g., with an enzyme that recognizes sites containing CG sequences, such as MspI), and completes fragment sequencing after ligation with an adaptor ligand. The choice of restriction enzyme reduces the number of overlapping sequences that may map to multiple gene locations during analysis and enriches for fragments in CpG-dense regions. Thus, RRBS reduces the complexity of the nucleic acid sample by selecting a subset of restriction fragments for sequencing (e.g., by size selection using preparative gel electrophoresis). In contrast to whole-genome bisulfite sequencing, all fragments produced by restriction enzyme digestion contain DNA methylation information for at least one CpG dinucleotide. Thus, RRBS enriches the sample for promoters, CpG islands, and other genomic features with a high frequency of restriction enzyme cleavage sites in these regions, providing an assay for assessing the methylation state of one or more genomic loci.
[0173] A typical RRMS protocol involves digesting a nucleic acid sample with a restriction enzyme such as MspI, filling in overhangs and A-tailing, ligating adapters, bisulfite conversion, and PCR (see, e.g., Meissner et al. (2005) "Genome-scale DNA methylation mapping of clinical samples at single-nucleotide resolution" Nat Methods 7: 133-6; Meissner et al. (2005) "Reduced representation bisulfite sequencing for comparative high-resolution DNA methylation analysis" Nucleic Acids Res. 33: 5868-77).
[0174] In some embodiments, quantitative allele-specific real-time target and signal amplification (QuARTS) assays are used to assess methylation status. During each QuARTS assay, three reactions occur sequentially: amplification (reaction 1) and target probe cleavage (reaction 2) in the first reaction; and FRET cleavage and fluorescent signal generation (reaction 3) in the second reaction. When a target nucleic acid is amplified with specific primers, a specific detection probe with a flap sequence loosely binds to the amplification product. The presence of a specific invading oligonucleotide at the target binding site triggers cleavage to release the flap sequence by cleaving between the detection probe and the flap sequence. The flap sequence is complementary to the non-hairpin portion of the FRET cassette. Therefore, the flap sequence functions as an invading oligonucleotide on the FRET cassette, resulting in cleavage between the FRET cassette fluorophore and the quencher, generating a fluorescent signal. The cleavage reaction can cleave multiple probes per target, thus releasing multiple fluorophores per flap, resulting in exponentially increasing signal amplification. QuARTS can detect multiple targets in a single reaction by using FRET cassettes with different dyes (see, e.g., Zou et al. (2010) "Sensitive quantification of methylated markers with a novel methylation-specific technology" Clin Chem 56: A199; U.S. Patent Application Nos. 12 / 946,737, 12 / 946,745, 12 / 946,752, and 61 / 548,639).
[0175] The term "bisulfite reagent" refers to a reagent containing bisulfite, disulfite, hydrogen sulfite, or a combination thereof, which, as disclosed herein, serves to distinguish between methylated and unmethylated CpG dinucleotide sequences. Methods for such treatment are known in the art (e.g., PCT / EP2004 / 011715, incorporated herein by reference). Bisulfite treatment is preferably carried out in the presence of a denaturing solvent, such as, but not limited to, n-alkylene glycol or diethylene glycol dimethyl ether (DME), or in the presence of dioxane or a dioxane derivative. In some embodiments, the denaturing solvent is used at a concentration of 1% to 35% (v / v). In some embodiments, the bisulfite reaction is carried out in the presence of a scavenger, such as, but not limited to, a chroman derivative (e.g., 6-hydroxy-2,5,7,8-tetramethylchromium 2-carboxylate or trihydroxybenzoic acid and its derivatives, e.g., gallic acid) (see PCT / EP2004 / 011715, incorporated herein by reference). The bisulfite conversion is preferably carried out at a temperature of 30°C to 70°C, with the temperature being briefly raised above 80°C during the reaction (see PCT / EP2004 / 011715, incorporated herein by reference). The bisulfite-treated DNA is preferably purified prior to quantification. This may be done by any means known in the art, such as, but not limited to, ultrafiltration (e.g., by Microcon™ columns (manufactured by Millipore™)). Purification was performed using a modified manufacturing protocol (see, for example, PCT / EP2004 / 011715, incorporated herein by reference).
[0176] In some embodiments, fragments of the treated DNA are amplified using a set of primer oligonucleotides (see, e.g., Tables 3 and / or 5) and an amplification enzyme according to the present invention. Amplification of several DNA segments can be carried out sequentially in one and the same reaction vessel. Typically, amplification is carried out using the polymerase chain reaction (PCR). The amplified products are typically 100-2000 base pairs in length.
[0177] In another embodiment of the method, the methylation status of a CpG position within or near a marker containing a DMR (e.g., DMR1-274 as provided in Tables 1, 2, 4, and 7) can be detected by using a methylation-specific primer oligonucleotide. This technique (MSP) is described in U.S. Pat. No. 6,265,171 to Herman. The use of methylation-status-specific primers for the amplification of bisulfite-treated DNA allows for the discrimination between methylated and unmethylated nucleic acids. An MSP primer pair contains at least one primer that hybridizes to a bisulfite-treated CpG dinucleotide. Therefore, the sequence of the primer contains at least one CpG dinucleotide. MSP primers specific for unmethylated DNA contain a "T" at the C position in the CpG.
[0178] The fragments obtained by amplification can be directly or indirectly labeled for detection. In some embodiments, the label is a fluorescent label, a radionuclide, or a detectable molecular fragment with a typical mass that can be detected by mass spectrometry. The label is a mass label, and in some embodiments, the labeled amplification product has a single positive or negative net charge, which can improve detectability in mass spectrometry. Detection can be achieved and visualized by, for example, matrix-assisted laser desorption / ionization (MALDI) or electrospray ionization (ESI).
[0179] Methods for isolating DNA suitable for these assay techniques are known in the art. In particular, some embodiments involve isolating nucleic acids as described in U.S. Patent Application No. 13 / 470,251 ("Nucleic Acid Isolation"), which is incorporated herein by reference.
[0180] 〔method〕 In some embodiments of the present technology, a method is provided that includes the steps of: 1) contacting nucleic acid obtained from a subject (e.g., genomic DNA, isolated from a body fluid such as a stool sample, stomach tissue, or plasma sample) with at least one reagent or set of reagents that distinguish between methylated and unmethylated markers in at least one marker comprising a DMR (e.g., DMR1-274, e.g., as shown in Tables 1, 2, 4, and 7); 2) Detection of neoplasia or proliferative disorders (eg, with a sensitivity of greater than or equal to 80% and a specificity of greater than or equal to 80%).
[0181] In some embodiments of the present technology, a method is provided that includes the steps of: 1) contacting nucleic acid obtained from a subject (e.g., genomic DNA, e.g., isolated from a body fluid such as a stool sample or stomach tissue) with at least one reagent or set of reagents that distinguish between methylated and unmethylated CpG dinucleotides in at least one marker selected from the group consisting of ARHGEF4, ELMO1, ABCB1, CLEC11A, ST8SIA1, SFMBT2, CD1D, CYP26C1, ZNF569, and C13ORF18; 2) Detection of gastric cancer (e.g., with a sensitivity greater than or equal to 80% and a specificity greater than or equal to 80%).
[0182] In some embodiments of the present technology, a method is provided that includes the steps of: 1) contacting nucleic acid obtained from a subject (e.g., genomic DNA, e.g., isolated from a plasma sample) with at least one reagent or set of reagents that distinguish between methylated and unmethylated CpG dinucleotides in at least one marker selected from a chromosomal region having an annotation selected from the group consisting of ELMO1, ARHGEF4, EMX1, SP9, CLEC11A, ST8SIA1, BMP3, KCNA3, DMRTA2, KCNK12, CD1D, PRKCB, CYP26C1, ZNF568, ABCB1, ELOVL2, PKIA, SFMBT2(893), PCBP3, MATK, GRN2D, NDRG4, DLX4, PPP2R5C, FGF14, ZNF132, CHST2(7890), FLI1, c13 or fl8, or ZNF569; 2) Detection of gastric cancer (e.g., with a sensitivity of greater than or equal to 80% and a specificity of greater than or equal to 80%).
[0183] In some embodiments of the present technology, a method is provided that includes the steps of: 1) contacting nucleic acid obtained from a subject (e.g., genomic DNA, e.g., isolated from a body fluid such as a stool sample or stomach tissue) with at least one reagent or set of reagents that distinguish between methylated and unmethylated CpG dinucleotides in at least one marker selected from a chromosomal region annotated with ELMO1, ZNF569, C13orf18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4, and ST8SIA1; 2) Detection of gastric cancer (e.g., with a sensitivity of greater than or equal to 80% and a specificity of greater than or equal to 80%).
[0184] Preferably, the sensitivity is about 70% to about 100%, or about 80% to about 90%, or about 80% to about 85%. Preferably, the specificity is 70% to about 100%, or about 80% to about 90%, or about 80% to about 85%.
[0185] Genomic DNA can be isolated by any technique, including the use of commercially available kits. Briefly, the DNA of interest is encapsulated in a cell membrane, and the biological sample must be enzymatically, chemically, or mechanically disrupted and dissolved. The DNA solution is then purified of proteins and other contaminants, for example, by digestion with proteinase K. The genomic DNA is then eluted from the solution. This can be accomplished by a variety of techniques, including salting out, organic solvent extraction, and binding of DNA to a solid-phase support. The choice of method will be influenced by several factors, including time, cost, and the required quality of the DNA. All types of clinical samples, including those from tumorigenic or pre-tumorigenic events, are suitable for use in this method (e.g., cell lines, histological slides, biopsies, paraffin-embedded tissues, body fluids, stool, gastric tissue, colonic fluid, urine, plasma, serum, whole blood, isolated blood cells, cells isolated from blood, and combinations thereof).
[0186] The above techniques are not limited to methods used to prepare samples or nucleic acids for experiments, for example, DNA can be isolated from stool, blood, or plasma samples using direct gene capture, e.g., methods detailed in U.S. Patent Application No. 61 / 485,386 and related publications.
[0187] The genomic DNA sample is treated with at least one reagent or a series of reagents that distinguish between methylated and unmethylated CpG dinucleotides within at least one marker that includes a DMR (e.g., DMR1-274, e.g., based on Tables 1, 2, 4, and 7).
[0188] In some embodiments, the reagent converts the demethylated cytosine base at the 5' position to uracil, thymine, or other bases dissimilar to cytosine upon hybridization, although in some embodiments the reagent may be a methylation-sensitive restriction enzyme.
[0189] In some embodiments, the genomic DNA sample is treated with a method that converts demethylated cytosine bases at the 5' position to uracil, thymine, or other bases different from cytosine by a reaction. In some embodiments, this treatment is followed by alkaline hydrolysis with hydrogen sulfate (hydrogen, sulfuric, disulfate).
[0190] The processed nucleic acid is analyzed for the methylation status of the target gene sequence (at least one gene, genome sequence, or nucleotides from a marker comparing the DMR, e.g., at least one DMR selected from DMRs 1-274, e.g., based on Tables 1, 2, 4, or 7). The above methods of analysis are selected from known techniques, e.g., QuARTS and MSP, which are included in the present invention.
[0191] Aberrant methylation, more particularly hypermethylation of markers comparing DMRs (at least one DMR selected from DMR1-274, e.g., based on Tables 1, 2, 4, or 7), is associated with gastric cancer and more particularly with predicted tumor location.
[0192] The present invention relates to the analysis of several samples related to gastric cancer. For example, in some embodiments, the sample includes tissue or biological fluid from a patient. In some embodiments, the sample includes secretions. In some embodiments, the sample includes blood, serum, plasma, gastric secretions, pancreatic juice, a gastrointestinal biopsy sample, microdissected cells from a gastrointestinal biopsy, gastrointestinal cells discarded in the gastrointestinal lumen, or gastrointestinal cells recovered from feces. In some embodiments, the subject is a human. These samples can be derived from the upper or lower gastrointestinal tract, or include cells, tissues, and / or secretions from both the upper and lower gastrointestinal tract. The samples can be derived from the liver, bile duct, pancreas, stomach, colon, or rectum. In some embodiments, the sample includes cells, secretions, or tissue from the esophagus, small intestine, appendix, duodenum, polyps, gallbladder, anus, and / or peritoneum, cellular fluid, ascites, urine, feces, pancreatic juice, fluid obtained during endoscopy, blood, mucus, or saliva. In some embodiments, the sample is stool.
[0193] Such samples can be obtained by any of a number of methods known in the art (e.g., obvious to those skilled in the art).For example, urine samples and fecal samples can be easily obtained, and blood samples, ascites samples, plasma samples, or pancreatic juice samples can be obtained parenterally, for example, by using a needle and syringe.Acellular or substantially acellular samples can be obtained by subjecting the sample to various techniques known in the art, including, but not limited to, centrifugation and filtration.Although it is generally preferred to use non-invasive techniques to obtain the sample, it is also preferred to obtain samples such as tissue homogenates, tissue slices, and biopsy specimens.
[0194] In some embodiments, the technology relates to a method for treating a patient (e.g., a patient with gastric cancer (e.g., gastric cancer), a patient with early-stage gastric cancer, or a patient who may progress to gastric cancer), the method comprising determining the methylation status of a DMR as provided herein, and administering a treatment to the patient based on the results of determining the methylation status. The treatment may be the administration of a pharmaceutical compound, a vaccine, performing a surgical procedure, imaging the patient, or performing other tests. Preferably, the application is in a clinical screening method, a method for prognostic evaluation, a method for monitoring the outcome of a treatment, a method for identifying patients who are likely to respond to a particular therapeutic treatment, a method for imaging the patient or subject, or a method for drug screening or development.
[0195] In some embodiments of the present technology, a method for diagnosing gastric cancer (e.g., gastric cancer) is provided. The terms "diagnosing" and "diagnosis" as used herein refer to a method that allows a person skilled in the art to estimate and even determine whether a subject has a certain disease or condition, or whether the subject may progress to a certain disease or condition in the future. A person skilled in the art often makes a diagnosis based on one or more indicators of diagnosis (e.g., biomarkers (e.g., DMRs as disclosed herein), methylation status that represents the presence, severity, or absence of a condition, etc.).
[0196] Along with diagnosis, clinical cancer prognosis involves determining the pathogenicity and recurrence rate of cancer and planning the most effective treatment. If a more accurate diagnosis can be made or the likely risk of cancer progression can be assessed, appropriate treatment, and in some cases, less risky treatment, can be selected. Cancer biomarker findings (e.g., measurement of methylation status) are useful for distinguishing subjects with a good prognosis and / or low risk of cancer progression (who require no or little treatment) from subjects who are prone to cancer progression or cancer recurrence (who will benefit from more intensive treatment).
[0197] Such "diagnosing" or "diagnosis," as used herein, further includes determining the risk of cancer progression or determining a prognosis, such as selecting an appropriate (or effective) treatment based on the measurement of a diagnostic biomarker (e.g., DMR) reported herein, or predicting a clinical outcome (medical treatment or no medical treatment) by monitoring current treatment and possibly modifying treatment. Furthermore, in some embodiments of the present disclosure, the combined determination of biomarkers over time can facilitate diagnosis and / or prognosis. Temporal changes in biomarkers can predict clinical outcomes, monitor the progression of gastric cancer, or directly monitor the effectiveness of appropriate treatment against the cancer. In some such embodiments, for example, biological samples are monitored to predict changes in the methylation status of one or more biomarkers (e.g., DMR) identified herein (potentially monitoring one or more biomarkers) over time during the progression of effective treatment.
[0198] Further specific examples of patient status include methods for determining whether to initiate or continue cancer prevention or treatment in a patient. In some specific examples, the methods include a system of biological samples from a patient over a period of time; each of the biological samples includes measuring changes in the methylation status of at least one biomarker identified herein. Changes in the methylation status of biomarkers can be used to predict the risk of cancer progression, predict clinical outcome, determine whether to initiate or continue cancer prevention or treatment, and determine whether a current treatment is effective for cancer. For example, a first sample selected for initiation of treatment and a second sample selected some time after initiation of treatment. The methylation status can be measured in each of the samples obtained at different times or quantitatively from different characteristics. Changes in the methylation status of biomarker levels from different samples can correlate with the risk, progression, effective treatment, or prognosis of gastric cancer in a patient.
[0199] More particularly, the methods and properties of the invention are for the treatment or diagnosis of disease at an early stage, e.g., before symptoms of the disease appear. In some particular cases, the methods and properties of the invention are for the treatment or diagnosis of disease at a clinical stage.
[0200] More specifically, many diagnostic or prognostic determinations of one or more biomarkers involve changes over time in the markers used to make the diagnostic or prognostic determination. For example, diagnostic markers have been determined at early stage and again at secondary stage. Specifically, an increase in a marker from early stage to secondary stage can be diagnostic of a particular type or severity of cancer, or prognostic. Similarly, a decrease in a marker from early stage to secondary stage can indicate a particular type or severity of cancer, or prognostic. Furthermore, the degree of change in one or more markers correlates with cancer severity or future adverse events. A skilled practitioner will appreciate these and other specific measurements and comparisons can be made to generate similar biomarkers at multiple stages, and a biomarker can be measured at one stage and a second biomarker at a secondary stage, and the comparison of these markers can provide diagnostic information.
[0201] As used herein, the phrase "diagnostic determination" refers to a method by which a skilled artisan can predict the course or outcome of a patient's condition. A "predictive" term does not refer to the ability to predict the course or outcome of a condition with 100% accuracy, nor does it refer to a course or outcome that is more or less likely based on the methylation status of a biomarker (e.g., a DMR). Instead, a skilled artisan will understand that a "predictive" term refers to an increased probability of a certain course or outcome occurring; and that course or outcome is more likely to occur in patients who exhibit the condition compared to those individuals who do not exhibit the condition. For example, in individuals who do not exhibit the condition (e.g., have a normal methylation status of one or more DMRs), the likelihood of an outcome (e.g., suffering from gastric cancer) may be very low.
[0202] Specifically, statistical analysis is related to predictive indicators where the outcome is more likely to be progressive. For example, specifically, a methylation status different from a normal reference sample derived from a cancer-free patient can be shown by determining the level of statistical significance in patients with a similar level of methylation status in the reference sample than in patients suffering from cancer. In addition, the change in methylation status from the reference level (e.g., "normal") can be reflected in the patient's prognosis, and the degree of change in methylation status is related to the severity of the adverse outcome. Statistical significance is often determined by comparing two or more populations and determining the confidence difference or p-value. For example, Dowdy et al. and Wearden, Statistics for Research, John Wiley & Sons, New York, 1983, is incorporated by reference in its entirety. Typical confidence intervals for the current patient status are 90%, 95%, 97.5%, 98%, 99%, 99.5%, 99.9%, and 99.99%, with typical p-values of 0.1, 0.05, 0.025, 0.02, 0.01, 0.005, 0.001, and 0.0001.
[0203] In other specific embodiments, threshold levels of change in the methylation status of the biomarkers (e.g., DMRs) disclosed herein for prognosis or diagnosis can be established, and the level of methylation of the biomarker in a biological sample can be readily compared to the threshold level of methylation status. Preferred threshold changes in methylation status for the biomarkers disclosed herein are about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 50%, about 75%, about 100%, and about 150%. However, in other embodiments, "nomograms" can be established, whereby the methylation status of a given prognostic or diagnostic indicator (biomarker or combination of biomarkers) directly correlates with a propensity toward a given outcome. Those skilled in the art are familiar with the use of such nomograms, as they relate two numerical values, understanding that the measurement, like the marker concentration, is uncertain because it refers to an individual sample measurement rather than a population average.
[0204] In some embodiments, a reference sample is analyzed simultaneously with the biological sample, and the results obtained from the biological sample are compared to the results obtained from the reference sample. Additionally, a standard distribution curve is expected to be generated to compare the assay results for the biological sample. Such a standard distribution curve represents the methylation status of the biomarker as a function of assay unit, e.g., the intensity of the fluorescent signal when a fluorescent label is used. Using samples from multiple donors, the standard distribution curve defines a standard methylation status of one or more biomarkers in normal tissue, as well as "at-risk" levels of one or more biomarkers in tissue from donors with metaplasia and donors with gastric cancer. In one embodiment of the method, a patient is identified as having a metaplasia exhibiting an aberrant methylation status of one or more DMRs described herein in a biological sample collected from the patient. In yet another embodiment of the method, the patient is considered to have cancer upon detection of an aberrant methylation status of one or more such biomarkers in a biological sample from the patient.
[0205] The above analysis of markers can be performed separately or simultaneously with additional markers in a single test sample. For example, each marker can be tested simultaneously for efficient sample processing, further enabling broader diagnosis or accurate prognosis. Furthermore, certain skills in the art recognize the utility of testing multiple samples (e.g., serially over time) from the same patient. Such testing of serial samples has been found to identify changes in methylation markers over time. Changes in methylation status, as well as the absence of changes in methylation status, provide useful information about the disease state, including, but not limited to, the presence or amount of tissue recovered, the applicability of drug treatment, the effectiveness of various treatments, and patient outcomes, including risk of future conditions.
[0206] The above-mentioned biomarker analyses are carried out in various practical systems. For example, the use of microplates and full automation have facilitated the processing of large numbers of test samples. Alternatively, a single sample type has evolved to facilitate immediate treatment and diagnosis in a rapid manner, for example, for temporary transport or emergency room settings.
[0207] In some embodiments, when comparing the baseline methylation status, the patient is diagnosed as having gastric cancer, which occurs when there is a difference in the methylation status of at least one biomarker in the sample. Conversely, no change in methylation status is identified in the biological sample. The patient is identified as having no gastric cancer, no risk of cancer, or a low risk of cancer. In this regard, patients with cancer or at risk can be distinguished from patients who are generally cancer-free or at low risk. Those patients at risk for developing gastric cancer are placed on a more intensive or routine cleaning regimen, including endoscopic observation. Meanwhile, patients at generally no or low risk can avoid endoscopy until a future screening time, e.g., screening consistent with current technology, indicating a risk of gastric cancer in those patients.
[0208] As previously mentioned, specific developments of the current methods involve detecting changes in the methylation status of one or more biomarkers, which can be measured qualitatively or quantitatively. A reliable threshold measurement is thus indicated for determining whether a patient has a risk of progression to gastric cancer (e.g., gastric cancer) at the stage of diagnosis. For example, the methylation status of one or more biomarkers in a biological sample can vary from a predetermined baseline methylation status. In some embodiments of the method, the baseline methylation status is a biomarker with a detectable methylation status. In other embodiments of the method, if a reference sample is run simultaneously with the biological sample, the predetermined baseline methylation status is the methylation status in the reference sample. In other embodiments of the method, the predetermined methylation status is based on a later or standard distribution curve. In other embodiments of the method, the predetermined methylation status is a distinct state or a range of states. The predetermined methylation is selected based in part on the implementation of the method, such as the specificity desired, within clearly acceptable limits for those skilled in the art.
[0209] Beyond the diagnostic methods involved, preferred patients are vertebrate patients. Preferred vertebrates are warm-blooded; preferred warm-blooded vertebrates are mammals. Preferred mammals are mostly humans. As used herein, "patient" includes both human and animal patients. Accordingly, veterinary treatment methods are provided herein. Such current techniques are not only useful for the diagnosis of mammals such as humans, but also for endangered and valuable mammals, such as Siberian tigers; animals of economic importance, such as those farmed for human consumption; and animals of social importance to humans, such as pets and animals kept in zoos. Examples of such animals include, but are not limited to, carnivorous mammals such as cats and dogs; domestic pigs, piglets, boars, and wild boars; and ruminants or ungulates, such as domestic cattle, bulls, sheep, giraffes, deer, goats, bison, camels, and horses. Thus, the above diagnosis and treatment of livestock is also indicated, including, but not limited to, domesticated pigs, ruminants, ungulates, horses (including racehorses), and the like. The immediately indicated patient context further includes a system for diagnosing gastric cancer (e.g., gastric cancer) in a patient. The system is defined, for example, as a mass-produced kit that has been used to screen for gastric cancer risk or diagnosis in a patient from whom a biological sample has been collected. Illustrative systems are consistent with the present technology for measuring DMRs of methylation status according to Tables 1, 2, 4, or 7. [Example]
[0210] Example 1 - Identification of Markers for the Detection of Gastric Cancer Experiments conducted in the course of developing embodiments for this technology identified 123 DNA methylation markers for gastric cancer, corresponding to 248 DMRs (e.g., DMR numbers 1 to 248; Table 1). These DNA methylation markers were identified from data generated by CpG island enrichment accompanied by parallel large-scale sequencing of patient-control tissue sample sets. Controls included DNA obtained from normal gastric tissue, normal colonic epithelium, and normal leukocytes. The method utilized reduced representation bisulfite sequencing (RRBS). The third step included analyzing the data for coverage cutoffs, eliminating all non-informative sites, contrasting percent methylation between diagnostic subgroups using logistic regression, computationally generating CpG "islands" based on distinct groupings of closely related methylated sites, and receiver operating characteristic analysis. The fourth step was to filter the % methylation data (both individual CpGs and in silico clusters) to maximize signal-to-noise ratio, minimize background methylation, reveal tumor heterogeneity, and emphasize receiver operating characteristic (ROC) performance. Furthermore, this analysis ensured the identification of CpG hotspots throughout the defined DNA length for easy and optimal design of downstream marker assays (methylation-specific PCR, small-fragment deep sequencing).
[0211] Each sample yielded approximately 2–3 million high-quality CpGs, which, after analysis and filtering, resulted in fewer than 1,000 highly discriminative individual sites. These were concentrated in 123 localized regions of differential methylation, some spanning 30–40 bases and others spanning kilobases (see Table 1). All DMRs had an AUC of 0.8 or greater, with some demonstrating perfect discrimination of patients from controls (AUC of 1.0). In addition to the 0.8 AUC threshold, identified gastric cancer markers were required to exhibit methylation density in tumors greater than 20-fold that of normal gastric or colonic mucosa and <1.0% methylation in non-neoplastic gastric or colonic mucosa.
[0212] [Table 1]
[0213] [Table 2]
[0214] [Table 3]
[0215] [Table 4]
[0216] [Table 5]
[0217] [Table 6]
[0218] Of the 123 disclosed markers, 22 were selected based on AUC and fold change (cancer vs. normal stomach + normal colon) and incorporated into a larger tissue validation study along with 73 exploratory markers from other GI cancer sites (colon, pancreas, bile duct, and esophagus). The samples in these different cohorts ranged from 15 to 50 samples, representing both cancer and normal, and where applicable, adenoma. The validation platform was quantitative methylation-specific PCR using primers optimized by a methylation control population. All assays were performed on Roche 480 LightCyclers. Results were analyzed for receiver operating characteristic (ROC), fold change, and complementarity. The top 30 markers, assessed by sensitivity (for cancer and adenoma), are listed in Table 2. Table 3 provides primer information for the markers listed in Table 2. Of these, 10 were selected for further tissue validation in a collaborative study with Mayo Korean (see Example 2).
[0219] [Table 7]
[0220] [Table 8]
[0221] [Table 9]
[0222] Example 2 - Gastric Cancer Detection with Novel Methylated DNA Markers: Tissue Validation in Patient Cohorts from the United States and Korea Experiments conducted in the course of developing embodiments for this technology identified novel methylated DNA by whole methylome sequencing and then validated the best candidates in tissues from patient cohorts in the United States and South Korea (see Tables 4, 5, and 6).
[0223] After attempting to discover the entire methylome using reduced-representation bisulfite sequencing, 17 candidate methylated DNA markers were selected for blinded testing in gastric tissue collections from patients in the United States and South Korea. DNA from microdissected tissues was quantified by methylation-specific polymerase chain reaction, and marker levels were normalized by beta-actin (total human DNA). GC cases consisted of 35 (US) and 50 (South) patients with pathologically confirmed untreated adenocarcinoma. Controls included pathologically normal epithelium obtained at sites distant from the primary GC from 65 (US) and 50 (South) healthy patients demographically matched to the cases. For each marker, the area under the receiver operating characteristic curve (AUC) was calculated from nominal logistic regression.
[0224] Overall, the top 10 markers included ARHGEF4, ELMO1, ABCB1, CLEC11A, ST8SIA1, SFMBT2, CD1D, CYP26C1, ZNF569, and C13ORF18 (see Table 4). At 95% specificity, this panel of markers detected GC in 100% of the US cohort and 94% of the Korean cohort. The individual AUCs were 0.92, 0.91, 0.89, 0.88, 0.85, 0.82, 0.82, 0.81, 0.75, and 0.66 in the US cohort and 0.73, 0.90, 0.78, 0.75, 0.80, 0.94, 0.87, 0.79, 0.75, and 0.79 in the Korean cohort, respectively (see Table 6). Some markers, such as ELMO1, showed similarly high discrimination in each patient cohort, while others, such as ARHGEF4, showed more variable discrimination due to high background from normal controls (NL) across patient cohorts (see Figure 1). Appendix 5 shows the forward and reverse primer information for the 10 markers shown in Appendix 4.
[0225] [Table 10]
[0226] [Table 11]
[0227] [Table 12]
[0228] Example 3 - Detection of gastric cancer by evaluation of novel methylated DNA markers in plasma This example explored the use of selected methylated DNA marker (MDM) candidates from Examples I and II to evaluate plasma-based approaches for gastric adenocarcinoma (GC) detection. Indeed, this example demonstrates that a panel of novel MDMs evaluated in plasma accurately distinguishes GC cases from controls. Marker levels were shown to be negligible in controls, elevated in cases, and gradually increase with GC stage.
[0229] Archival plasma samples from a single, integrated institution that met the case inclusion criteria were included in the study. Cases included 37 patients with pathologically confirmed adenocarcinoma, spanning all stages, histological types, and gastric sites. Case plasma samples were collected before resection, chemotherapy, or radiation. Archival plasma samples from 38 age- and sex-matched healthy volunteers were selected as controls. DNA was extracted from 2 ml of plasma and bisulfite-treated. Selected MDMs (ELMO1, ZNF569, C13orf18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4, and ST8SIA1) (Table 7 shows the DMR information for LRRC4) were then quantified using a quantitative allele-specific real-time target and signal amplification method without optimization. These markers have demonstrated high performance in tissues as assessed by clinical sensitivity and specificity, fold change in methylation signal when comparing gastric cancer versus normal gastric tissue, and very low background methylation in DNA from blood cells.
[0230] To facilitate optimal analysis in plasma, these markers (previously examined in tissues using methylation-specific PCR assays) were redesigned using the original DMR sequences in combination with a quantitative allele-specific real-time target and signal amplification (QuARTs) assay. The QuARTs assay includes PCR primers (see Table 8), a detection probe (see Table 8), and an interstitial oligo (Integrated DNA Technologies), GoTaq DNA polymerase (Promega), Cleavase II (Hologic), and a fluorescence resonance energy transfer (FRET) reporter cassette containing FAM, HEX, and Quasar 670 dyes (Biosearch Technologies).
[0231] Figure 2 shows the oligonucleotide sequences for the FRET cassettes used to detect methylated DNA features by quantitative allele-specific real-time target and signal amplification (QuARTs) assays. Each FRET sequence contains a fluorophore and a quencher that can be multiplexed together in three separate assays.
[0232] Standards were derived from amplified and truncated sequences from plasmid controls. Absolute copy numbers of serially diluted standards were obtained from Poisson modeling.
[0233] DNA was purified from 2 mL of archived plasma samples and bisulfite converted using in-house methods and automated instrumentation. Treatment control genes were pre-spiked to all samples to correct for variability in target recovery over the course of the study.
[0234] Samples were preamplified with 12 primer sets, diluted, and subjected to QuARTs. The platform for the downstream step was the LightCycler 480 (Roche). Strand counts for each marker were normalized to the strands from β-actin included in the multiplex reaction.
[0235] The results were logistically analyzed and plotted in matrix format to assess complementarity.
[0236] Figure 3 shows receiver operating characteristic curves highlighting the performance in plasma of the three marker panel (ELMO1, ZNF569, and c13orf18) compared to the individual marker curves. At 100% specificity, the panel detected gastric cancer with a sensitivity of 86%.
[0237] Figure 4A shows the stage-dependent log-scale absolute strand counts of methylated ELMO1 in plasma, illustrating the progression from normal mucosa to stage 4 gastric cancer. Quantitative marker levels increased with GC stage.
[0238] FIG. 4B shows a bar graph demonstrating the sensitivity (100% specificity) of gastric cancer in plasma according to stage of a three-marker panel (ELMO1, ZNF569 and c13orf18).
[0239] Indeed, ELMO1 was the most discriminatory marker, alone with an AUC of 94% (95% CI 89-99%). At 100% specificity, a panel of three MDMs (ELMO1, ZNF569, and C13orf18) detected GC with a sensitivity of 86% (71-95%). According to GC stage (Figure 4A), the sensitivity of the panel at 100% specificity was 50%, 92%, 100%, and 100% for stages 1 (n = 8), 2 (n = 13), 3 (n = 4), and 4 (n = 11), respectively (p = 0.01 for trend). Quantitative MDM levels increased directly with GC stage. As shown by the distribution of methylated ELOMO1 (Fig. 4B), the median plasma strands / ml of plasma in controls was 0.4, and the strand counts in stages 1, 2, 3, and 4 of CG cases were 16, 111, 101, and 213, respectively (p for trend = 0.01). Neither age nor sex affected marker levels.
[0240] Figures 5A, B, and C show the performance of 12 gastric cancer markers (ELMO1, ZNF569, C13orf18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4, and ST8SIA1; ELMO1 was run twice, the second time in biplex format). 90% specificity (A), 95% specificity (B), and 100% specificity (C) are shown in matrix format for 74 plasma samples. Markers are arranged vertically and samples horizontally. Samples are aligned with normal on the left and cancer on the right. Positive hits are in light gray, and misses are in dark gray. Here, the top-performing marker (ELMO1) is listed first at 90% specificity, and the remaining panels are at 100% specificity. This plot allows markers to be evaluated in a combinatorial manner.
[0241] Figure 6 shows the performance of three gastric cancer markers (first panel) in 74 plasma samples, with 100% specificity. Markers are arranged vertically and samples horizontally. Cancers are arranged according to stage. Positive hits are in light gray, misses are in dark gray (ELMO1, ZNF569, and c13orf18).
[0242] Figure 7A-M show box plots (linear scale) of 12 gastric cancer markers (ELMO1, ZNF569, C13orf18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4, and ST8SIA1; ELMO1 was performed twice, the second time in a biplex format) in plasma. Control samples (N=38) are listed on the left, and gastric cancer cases (N=36) are listed on the right. The vertical axis is the % methylation normalized to β-actin strands.
[0243] [Table 13]
[0244] [Table 14]
[0245] [Table 15]
[0246] The publications and patents mentioned in the foregoing specification are incorporated herein by reference in their entirety for all purposes. Various modifications and variations of the described compositions, methods, and uses of the technology will be apparent to those skilled in the art without departing from the scope and spirit of the technology as described. Although the technology has been described in connection with illustrative specific embodiments, it will be understood that the invention as claimed should not be unduly limited to such specific embodiments. Indeed, various modifications of the described modes for carrying out the invention that are obvious to those skilled in pharmacology, biochemistry, medicine, or related fields are intended to be within the scope of the following claims.
Claims
1. 1. A method for assessing the methylation status of SFMBT2 and FER1L4 in a sample from a human subject having or suspected of having a gastric neoplasia, comprising: treating DNA from said sample with a reagent that modifies the DNA in a methylation-specific manner; amplifying the treated DNA using primers specific for methylation variable regions (DMRs) from at least SFMBT2 and FER1L4; and measuring the methylation status of the DMRs from at least SFMBT2 and FER1L4 using methylation-specific PCR, quantitative methylation-specific PCR, methylation-sensitive DNA restriction enzyme analysis, quantitative bisulfite pyrosequencing and / or bisulfite genomic sequencing PCR; The method, wherein the sample is a gastric tissue sample or a plasma sample.
2. The method of claim 1 further comprising the step of extracting DNA from the sample.
3. the DMR is associated with an area under the ROC curve (AUC) that is 0.5 or greater; the ROC curve distinguishes between DNA from the sample derived from a human subject having or suspected of having a gastric neoplasia and a control DNA sample; 2. The method of claim 1, wherein the control DNA sample is from a human subject who does not have a gastric neoplasia.
4. the DMR has a higher percentage of methylation compared to a control DNA sample; 2. The method of claim 1, wherein the control DNA sample is from a human subject who does not have a gastric neoplasia.
5. the DMR has a high frequency of hypermethylation compared to a control DNA sample; 2. The method of claim 1, wherein the control DNA sample is from a human subject who does not have a gastric neoplasia.
6. 2. The method of claim 1, comprising amplifying a DMR from at least one additional gene selected from the group consisting of NDRG4, ELMO1, ZNF569, C13orf18, CD1D, ARHGEF4, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4 and ST8SIA1.
7. 2. The method of claim 1, wherein the reagent that modifies DNA in a methylation-specific manner comprises one or more of a methylation-sensitive restriction enzyme, a methylation-dependent restriction enzyme, and / or a bisulfite reagent.
8. 2. The method of claim 1, wherein the DMR is in a coding region or a regulatory region.
9. The method of claim 1, wherein the primers specific to SFMBT2 comprise SEQ ID NOs: 5 and 6, or SEQ ID NOs: 83 and 84.