Detecting gastric neoplasm
Patent Information
- Application Number
- JP2025023534
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2015-08-31
- Filing Date
- 2025-02-17
- Publication Date
- 2026-02-04
AI Technical Summary
Current methods for detecting gastric cancer are inadequate, as most patients present with advanced disease, and existing markers are cumbersome and not highly sensitive.
The development of novel DNA methylation markers, specifically 123 markers identified from gastric cancer samples, which are confirmed across multiple sample populations to optimize robust candidates for detecting gastric cancer.
These methylation markers provide a high signal-to-noise ratio and low background levels, enabling accurate screening and diagnosis of gastric cancer with high discrimination.
Smart Images

Figure 00000064_0000 
Figure 00000064_0001 
Figure 00000064_0002
Abstract
Description
Technical Field
[0001] (Cross - reference to related applications) This application claims the benefit of, and claims priority to, U.S. Provisional Application No. 62 / 212,221, which is hereby incorporated by reference in its entirety.
[0002] Techniques related to detecting neoplasms (in particular, but not limited to, methods, compositions, and related uses for detecting pre - cancerous conditions and malignant neoplasms (such as gastric cancer)) are presented herein.
Background Art
[0003] Gastric cancer is the third most common cause of cancer - related death worldwide (see, e.g., World Health Organization. Cancer: Fact Sheet No 297. WHO), and is still difficult to treat in Western countries because most patients present with advanced disease. In the United States, malignant tumors of the stomach are currently the 15th most common cancer (see, e.g., Surveillance, Epidemiology, and End Results Program. SEER Stat FactSheets: Stomach Cancer. National Cancer Institute). However, in most developed countries, the rate of gastric cancer has decreased dramatically over the past half - century.
[0004] The decrease in gastric cancer is partly due to the widespread use of refrigeration. Refrigeration has several beneficial effects: increased consumption of fresh fruits and vegetables; reduced intake of salt, which is used as a food preservative; and reduced contamination of food by carcinogenic compounds resulting from the spoilage of unfrozen meat products. Salt and salted foods can damage the gastric mucosa, leading to inflammation and a related increase in DNA synthesis and cell proliferation. Other factors that may have contributed to the decline in the rate of gastric cancer include improved hygiene and the use of antibiotics, as well as a lower rate of chronic infection with Helicobacter pylori as a result of increased screening in some countries (see, for example, Global Cancer Facts & Figures, 3rd ed. American Cancer Society).
[0005] However, gastric cancer continues to be difficult to treat in Western countries because most patients present with advanced disease. Even patients who are in the most favorable condition and have undergone curative surgical resection often die due to recurrent disease.
[0006] Therefore, improved methods and techniques for the early detection of gastric cancer are needed. SUMMARY OF THE INVENTION
[0007] Methylated DNA has been studied as a promising group of biomarkers in most tumor types of tissue. In many instances, DNA methyltransferases add methyl groups to DNA at cytosine-phosphate-guanine (CpG) island sites as an epigenetic control of gene expression. In a biologically interesting mechanism, the acquired methylation phenomenon in the promoter region of tumor suppressor genes is thought to turn off expression and thus contribute to carcinogenesis. DNA methylation can be a diagnostic tool that is chemically and biologically more stable than RNA or protein expression (Laird (2010) “Principles and challenges of genome-wide DNAmethylation analysis” Nat Rev Genet 11: 191-203). Furthermore, in other cancers such as sporadic colorectal cancer, methylation markers provide excellent characteristics and are more widely beneficial and sensitive than individual DNA mutations (Zou et al (2007) “Highly methylated genes in colorectal neoplasia:implications for screening” Cancer Epidemiol Biomarkers Prev 16: 2686-96).
[0008] Analysis of CpG islands has provided important insights when used in animal models and human cell lines. For example, Zhang et al. discovered that amplicons from different parts of the same CpG island have different levels of methylation (Zhang et al. (2009) “DNA methylation analysis of chromosome 21 gene promoters at single base pair and single allele resolution” PLoS Genet 5:e1000438). Furthermore, methylation levels are assigned to two modes between hypermethylated and unmethylated sequences, further supporting a binary, switch-like pattern of DNA methyltransferase activity (Zhang et al. (2009) “DNA methylation analysis of chromosome 21 gene promoters at single base pair and single allele resolution” PLoS Genet 5:e1000438). In vivo analysis of mouse tissues and in vitro analysis of cell lines have demonstrated that only approximately 0.3% of promoters with a high CpG density (HCP, defined as having >7% CpG sequences within a 300 base pair region) are methylated, while sites with low CpG density (defined as having <5% CpG sequences within a 300 base pair region) are frequently methylated in a dynamic, tissue-specific manner (Meissner et al. (2008) “Genome-scale DNA methylation maps of pluripotent and differentiated cells” Nature 454: 766-70). HCPs contain promoters for ubiquitous and developmental genes.Among these, HCP sites methylated at >50% were several established markers (e.g., Wnt 2, NDRG2, SFRP2, and BMP3) (Meissner et al. (2008) “Genome-scale DNA methylation maps of pluripotent and differentiated cells” Nature 454: 766-70).
[0009] Overall, gastrointestinal cancers account for a higher number of deaths than cancers of any other organ system. The annual number of deaths in the United States due to upper GI cancers exceeds 90,000, compared to approximately 50,000 for colorectal cancers. Gastric adenocarcinoma is the second most common cause of cancer death worldwide (i.e., 1 million deaths per year). It is most common in Pacific Rim Asia, Russia, and parts of Latin America. The prognosis is poor (5-year survival rate <5-15%) because most patients present with advanced disease.
[0010] The genomic mechanisms underlying gastric cancer are complex and not fully understood. Chronic gastritis based on Helicobacter pylori infection contributes to the etiology, along with environmental exposure. Chromosomal instability and microsatellite instability have been associated with gastric cancer and epigenetic changes. The CpG methylator phenotype also accounts for 24-47% of cases. Commonly observed mutations include p53, APC, β-catenin, and k-ras, but these are relatively infrequent and heterogeneous in nature.
[0011] Gastric cancer cells and DNA flow through the digestive stream and are ultimately secreted into feces. Sensitive assays have been used to date to detect mutant DNA in the feces of gastric cancer patients who have been found to have tumors containing the same sequences. However, several limitations with mutant markers relate to the fundamental heterogeneity of their detection and cumbersome processing, typically requiring each mutation site spanning multiple genes to be assayed separately to achieve sensitivity.
[0012] Epigenetic methylation of DNA at cytosine-phosphate-guanine (CpG) island sites has been studied as a promising set of biomarkers in most tumor types of tissue. In a biologically interesting mechanism, acquired methylation in the promoter region of tumor suppressor genes halts expression and contributes to carcinogenesis. DNA methylation can be a more chemically and biologically stable diagnostic tool than RNA or protein expression. Furthermore, in other cancers such as sporadic colon cancer, aberrant methylation markers are more broadly useful and sensitive than individual DNA mutations, providing superior characteristics.
[0013] Several methods are available for the discovery of novel methylation markers. Microarrays based on interrogation of CpG methylation are a suitable high-throughput approach, but this strategy is mainly directed at known regions of desired tumor suppressor promoters that have been demonstrated. Alternative methods for genome-wide analysis of DNA methylation have been developed over the past decade. There are three basic approaches. The first uses digestion of DNA by restriction enzymes that recognize specific methylated sites (followed by several analytical techniques that yield or evaluate methylation data limited to restriction enzyme recognition sites, or in which primers are used to amplify DNA in an evaluation step (e.g., methylation-specific PCR; MSP)). The second approach involves enrichment of the methylated fraction of genomic DNA using antibodies against methylated cytosine or other methylation-specific binding domains (subsequent to microarray analysis or sequencing to map the fragments relative to a reference genome). This approach does not show single nucleotide resolution of all methylated sites within the fragments. The third approach begins with bisulfite treatment to convert all unmethylated thymines to uracil, followed by restriction enzyme treatment and complete sequencing of all fragments after ligation to adapter ligands. The choice of restriction enzyme can enrich the fragments for CpG-rich regions and reduce the number of repetitive sequences that can determine multiple gene positions during analysis. This latter approach (referred to as reduced representation bisulfite sequencing (RRBS)) is known but has not yet been used for studying gastric cancer.
[0014] RRBS generates single nucleotide resolution CpG methylation status data at moderate to high read coverage for 80 - 90% of all CpG islands and most of the tumor suppressor promoters. In cancer patient-control studies, Analysis of these reads results in the identification of methylated variable regions (DMRs). In previous RRBS analysis of pancreatic cancer samples, hundreds of DMRs were discovered, many of which were never associated with carcinogenesis and were unannotated. Further validation tests on independent tissue samples confirmed marker CpGs with 100% sensitivity and specificity in terms of their performance.
[0015] Clinical approaches to highly discriminable markers can have a major impact. For example, assays of such markers in remote media (such as feces or blood) can lead to more accurate screening or diagnostic tools for the detection of gastric neoplasms.
[0016] Experiments conducted in developing embodiments for this technology have identified and described 123 novel DNA methylation markers similarly generated from gastric cancer samples. Along with the cancer, normal gastric tissue, normal colon tissue, and normal leukocyte DNA were sequenced. The markers were confirmed on multiple sample populations to identify and optimize the most robust candidates. Further experiments identified 10 optimal markers for detecting gastric cancer (ARHGEF4, ELMO1, ABCB1, CLEC11A, ST8SIA1, SFMBT2, CD1D, CYP26C1, ZNF569 and C13ORF18; see Example 2, and Tables 4, 5 and 6). Further experiments identified 30 optimal markers for detecting gastric cancer (ELMO1, ARHGEF4, EMX1, SP9, CLEC11A, ST8SIA1, BMP3, KCNA3, DMRTA2, KCNK12, CD1D, PRKCB, CYP26C1, ZNF568, ABCB1, ELOVL2, PKIA, SFMBT2 (893), PCBP3, MATK, GRN2D, NDRG4, DLX4, PPP2R5C, FGF14, ZNF132, CHST2 (7890), FLI1, c13orf18 or ZNF569; see Example 1, and Tables 1, 2 and 3). Further experiments identified 12 optimal markers for detecting gastric cancer (ELMO1, ZNF569, C13orf18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4 and ST8SIA1; see Example 3, and Tables 7 and 8).
[0017] Accordingly, provided herein are techniques for gastric cancer screening markers that, when detected from a sample taken from a subject (e.g., a fecal sample; a gastric tissue sample; a plasma sample), result in a high signal-to-noise ratio and a low background level. The markers were identified in patient-control studies by comparing the methylation status of DNA markers from gastric tissue or plasma of subjects having gastric cancer to the methylation status of the same DNA markers from control subjects (e.g., DNA from normal gastric tissue, normal colonic epithelium, normal plasma and normal white blood cells; see Examples 1, 2 and 3, and Tables 1-8) (e.g., plasma tissue; see Example 3).
[0018] As described herein, the above techniques result in a number of methylation markers and subsets thereof (e.g., a set of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12 or more markers) with high discrimination for gastric cancer overall. Experiments used a selection filter for candidate markers to identify markers that exhibit a high signal-to-noise ratio and a low background level for high performance for the purpose of screening or diagnosing gastric cancer.
[0019] In some embodiments, the technology relates to assessing the presence and methylation status of one or more of the markers in a biological sample (e.g., gastric tissue, plasma sample) identified herein. These markers include one or more of the methylation variable regions (DMRs) as described herein (e.g., in Tables 1, 2, 4, and / or 7). The methylation status is evaluated in embodiments of the technology. Thus, the technology shown herein is not limited to methods by which the methylation status of a gene is measured. For example, in some embodiments, the methylation status is measured by a genome scanning method. For example, one method involves restriction target genome scanning (Kawai et al. (1994) Mol. Cell. Biol. 14: 7421-7427), and another example involves methylation-sensitive arbitrarily primed PCR (Gonzalgo et al. (1997) Cancer Res. 57: 594-599). In some embodiments, changes in the methylation pattern at a particular CpG are monitored by digestion of genomic DNA with a methylation-sensitive restriction enzyme (followed by Southern analysis of the desired region) (the digestion-Southern method). In some embodiments, analyzing changes in the methylation pattern involves a PCR-based process that involves digestion of genomic DNA with a methylation-sensitive restriction enzyme prior to PCR amplification (Singer-Sam et al. (1990) Nucl. Acids Res. 18: 687). Additionally, other techniques that utilize bisulfite treatment of DNA as a starting point for methylation analysis have been reported. These include methylation-specific PCR (MSP) (Herman et al. (1992) Proc. Natl. Acad. Sci. USA 93: 9821-9826) and restriction enzyme digestion of PCR products amplified from bisulfite-converted DNA (Sadri and Hornsby (1996) Nucl. Acids Res. 24: 5058-5059; and Xiong and Laird (1997) Nucl. Acids Res. 25: 2532-2534). PCR techniques for detecting gene mutations (Kuppuswamy et al. (1991) Proc. Natl. Acad. Sci. USA 88: 1143-1147), and PCR techniques for quantifying allele-specific expression (Szabo and Mann (1995) Genes Dev. 9: 3097-3108; and Singer-Sam et al. (1992) PCR Methods Appl. 1: 160-163) have been developed. Such techniques use internal primers (terminating at the 5' that anneals to the PCR-generated template and touches the single nucleotide to be quantified). Methods using the "quantitative Ms-SNuPE assay" described in U.S. Patent No. 7,037,650 are used in some embodiments States.
[0020] When assessing methylation status, the methylation status is often expressed as the fraction or percentage of individual strands of DNA that are methylated at a particular site (e.g., a single nucleotide, a particular region or locus, a longer desired sequence (e.g., a sub-sequence of DNA of ~100bp, 200bp, 500bp, 1000bp or more)) relative to the total population of DNA in a sample containing the particular site. Previously, the amount of unmethylated nucleic acid has been determined by PCR using a standard. Then, a known amount of DNA is bisulfite-treated and the resulting methylation-specific sequences are determined using real-time PCR or other exponential amplification (e.g., the QuARTS assay (e.g., as shown by U.S. Patent No. 8,361,720, which is incorporated herein by reference; U.S. Patent Application Publication No. 2012 / 0122088; and U.S. Patent Application Publication No. 2012 / 0122106)).
[0021] For example, in some embodiments, the method includes generating a standard curve for an unmethylated target by using an external standard. The standard curve is composed of at least two points and relates to the real-time Ct values for unmethylated DNA relative to a known quantitative standard. Then, a second standard curve for a methylated target is composed of at least two points and an external standard. This second standard curve relates to the Ct for methylated DNA relative to a known quantitative standard. Next, the test sample Ct values are determined for the methylated and unmethylated populations and the genomic equivalents of the DNA are calculated from the standard curves generated by the first two steps. The percentage of methylation at the desired site is calculated from the amount of methylated DNA (e.g., (number of methylated DNA) / (number of methylated DNA + number of unmethylated DNA) × 100) relative to the total amount of DNA in the population.
[0022] This specification also discloses compositions and kits for carrying out the above methods. For example, in some embodiments, reagents specific to one or more markers (e.g., primers, probes) are provided either singly or in sets (e.g., sets of primer pairs for amplifying multiple markers). Also provided are additional reagents for performing detection assays (e.g., enzymes, buffers, positive controls, and negative controls for carrying out QuARTS, PCR, sequencing, bisulfite, or other assays). Also provided is a reaction mixture containing the above reagents. Further disclosed is a master mix reagent set containing a plurality of reagents that can be added to each other and / or to a test sample to complete the reaction mixture.
[0023] In some embodiments, the techniques described herein are associated with a programmable machine designed to perform a series of arithmetic or logical operations such as those brought about by the methods described herein. For example, some embodiments of the above techniques are associated with computer software and / or computer hardware (e.g., implemented thereon). In one aspect, the above techniques relate to a computer having in the form of memory, elements for performing arithmetic and logical operations, and a processing element (e.g., a microprocessor) for executing a series of instructions (e.g., methods as shown herein) for reading, handling, and complementing data. In some embodiments, the microprocessor determines the methylation state (e.g., of one or more DMRs (such as DMR1 - 274 as shown in Tables 1, 2, 4, and 7)); compares the methylation states (e.g., of one or more DMRs (such as DMR1 - 274 as shown in Tables 1, 2, 4, and 7)); generates a standard curve; determines Ct values; calculates the portion, frequency, or percentage of methylation (e.g., of one or more DMRs (such as DMR1 - 274 as shown in Tables 1, 2, 4, and 7)); identifies CpG islands; determines the specificity and / or sensitivity of an assay or marker; ROCkyokusenn and calculate the associated AUC; it is part of a system for array analysis (all described herein or known in the art).
[0024] In some embodiments, the microprocessor or computer uses methylation state data in an algorithm for predicting the site of cancer.
[0025] In some embodiments, a software element or a hardware element receives the results of a plurality of assays and determines a single numerical result for reporting to a user indicative of cancer risk based on the results (e.g., determining the methylation state of a number of DMRs (such as those shown in Tables 1, 2, 4, and 7)). Related embodiments calculate a risk factor based on a mathematical combination (e.g., weighted combination, linear combination) of the results of a number of assays (e.g., determining the methylation state of a number of markers (such as a number of DMRs (such as shown in Tables 1, 2, 4, and 7))). In some embodiments, the methylation state of a DMR defines a one-dimensionality and can have a value in a multi-dimensional space, and the coordinates defined by the methylation states of a number of DMRs are the results regarding cancer risk (e.g., for reporting to a user).
[0026] Some embodiments include a recording medium and memory elements. The memory elements (e.g., volatile memory and / or non-volatile memory) are useful in the storage of instructions (e.g., an embodiment of a process as described herein) and / or data (e.g., artifacts (e.g., methylation measurements, sequence determinations, and associated statistical descriptions)). Also, some embodiments relate to a system comprising one or more of a CPU, a graphics card, and a user interface (e.g., including an output device (e.g., a display) and an input device (e.g., a keyboard)).
[0027] Programmable machines related to the above technology include conventional existing technologies and technologies under development or to be developed in the future (e.g., quantum computers, chemical computers, optical computers, computers based on spintronics, etc.).
[0028] In some embodiments, the above technology includes a wired transmission medium (e.g., metal cable, optical fiber) or a wireless transmission medium for data transmission. For example, some embodiments relate to data transmission across a network (e.g., local area network (LAN), wide area network (WAN), special network, Internet, etc.). In some embodiments, the programmable machine has a client / server relationship.
[0029] In some embodiments, the data is recorded on a computer-readable recording medium (e.g., hard disk, flash memory, optical medium, floppy disk, etc.).
[0030] In some embodiments, the technology described herein is associated with a plurality of programmable devices that operate in cooperation to implement the methods as described herein. For example, in some embodiments, a plurality of computers (e.g., connected to a network) can operate simultaneously to collect and process data in the implementation of cluster computing or grid computing or some other computer construct that relies on complete computers (having an on-board CPU, storage device, power supply, network interface, etc.) connected by (e.g., conventional network interfaces (e.g., Ethernet (registered trademark), optical fiber) or wireless communication technologies).
[0031] For example, some embodiments provide a computer including a computer-readable medium. Such embodiments include a random access memory (RAM) coupled to a processor. The processor executes computer-executable program instructions stored in the memory. Such a processor may include a microprocessor, an ASIC, a state machine, or other processors, and can be any of a number of computer processors (e.g., processors provided by Intel Corporation of Santa Clara, California and Motorola Corporation of Schaumburg, Illinois). Such a processor may include, or be connected to, a medium (e.g., a computer-readable medium) that stores instructions that, when executed by the processor, cause the processor to perform the steps described herein.
[0032] Embodiments of computer-readable media include, but are not limited to, electronic, optical, magnetic, or other storage or transmission devices capable of providing computer-readable instructions to a processor. Other examples of suitable media include, but are not limited to, floppy disks, CD-ROMs, magnetic disks, memory chips, ROMs, RAMs, ASICs, pre-programmed processors, all optical media, all magnetic tapes, or any other magnetic media readable by a computer. Also, various other forms (including routers, personal networks or public communication lines, or other wireless or wired transmission devices or channels) can transmit or convey instructions to a computer. The instructions can include code based on any suitable computer programming language (e.g., C, C++, C#, Visual Basic, Java®, Python, Perl, and JavaScript®).
[0033] The computer is connected to a network in some embodiments. Also, the computer may comprise many external or internal devices (e.g., a mouse, CD-ROM, DVD, keyboard, display, or other input or output devices). Examples of computers are personal computers, digital assistants, personal digital assistants, form phones, mobile phones, smartphones, portable small wireless paging devices, digital tablets, laptop computers, Internet appliances, and other processor-based devices. Generally, a computer related to multiple aspects of the above technologies is any platform based on any type of processor that can be operated by any operating system (e.g., Microsoft Windows, Linux®, UNIX®, Mac OS X, etc.) that can support one or more programs including the above technologies shown herein. Some embodiments include personal computers that execute other application programs (e.g., multiple applications). The multiple applications may be stored in memory, including, for example, word processing applications, spreadsheet applications, email applications, instant messaging applications, presentation applications, Internet browser applications, calendar / organizer applications, and any other application executable by a client device.
[0034] As described herein in connection with the above technologies, all such components, computers, and systems may be logical or virtual.
[0035] Accordingly, this specification discloses techniques related to a method for screening for neoplasms in a sample obtained from a subject. The method includes evaluating the methylation status of markers in a sample obtained from a subject (e.g., gastric tissue) (e.g., plasma sample); and identifying that the subject has a neoplasm when the methylation status of the markers is different from the methylation status of the markers evaluated in a subject without a neoplasm. Here, the markers include bases in methylation variable regions (DMRs) selected from the group consisting of DMR1 to 274 as shown in Tables 1, 2, 4, and / or 7. The above techniques relate to identifying and differentiating gastric cancer. In some embodiments, the neoplasm is a gastric neoplasm (e.g., gastric cancer). Some embodiments provide a method that includes evaluating a plurality of markers (e.g., evaluating 2 to 11 to 100 or 200 or 274 markers).
[0036] The above techniques are not limited to the methylation status to be evaluated. In some embodiments, evaluating the methylation status of markers in a sample includes quantifying the methylation status of a single base. In some embodiments, quantifying the methylation status of markers in a sample includes determining the degree of methylation in a plurality of bases. Further, in some embodiments, the methylation status of a marker includes an increased methylation status of the marker relative to the normal methylation status of the marker. In some embodiments, the methylation status of a marker includes a decreased methylation status of the marker relative to the normal methylation status of the marker. In some embodiments, the methylation status of a marker includes a methylation pattern of the marker that is different from the normal methylation status of the marker.
[0037] Furthermore, in some embodiments, the marker is a region of less than 100 bases, a region of less than 500 bases, a region of less than 1000 bases, a region of less than 5000 bases, or in some embodiments, the marker is 1 base. In some embodiments, the marker is in a promoter with a high CpG density.
[0038] The above technology is not limited to the type of sample. For example, in some embodiments, the sample is a fecal sample, a tissue sample (e.g., a gastric tissue sample), a blood sample (e.g., plasma, serum, whole blood), excreta, or a urine sample.
[0039] Furthermore, the above technology is not limited to the method used to determine the methylation state. In some embodiments, quantification includes using methylation-specific polymerase chain reaction, nucleic acid sequencing, mass spectrometry, methylation-specific nuclease, mass-based separation, or target capture. In some embodiments, quantification includes using methylation-specific oligonucleotides. In some embodiments, the above technology utilizes parallel large-scale sequencing (e.g., next-generation sequencing) (e.g., sequencing by synthesis, real-time (e.g., single-molecule) sequencing, bead emulsion sequencing, nanopore sequencing, etc.) to determine the methylation state.
[0040] The above technology provides reagents for detecting DMR. For example, in some embodiments, a set of oligonucleotides containing the sequences shown by SEQ ID NOs: 1 to 109 is provided. In some embodiments, oligonucleotides containing sequences complementary to chromosomal regions having bases in the DMR are provided.
[0041] The above technology provides various panels of markers. For example, in some embodiments, the marker includes a chromosomal region having an annotation that the marker is ARHGEF4, ELMO1, ABCB1, CLEC11A, ST8SIA1, SFMBT2, CD1D, CYP26C1, ZNF569, or C13ORF18. Further, embodiments provide a method for analyzing DMRs based on Table 4, where the DMR numbers are 49, 43, 196, 66, 1, 237, 249, 250, 251, or 252. Some embodiments provide for determining the methylation state of a marker, where the chromosomal region has an annotation that the marker is ELMO1, ARHGEF4, EMX1, SP9, CLEC11A, ST8SIA1, BMP3, KCNA3, DMRTA2, KCNK12, CD1D, PRKCB, CYP26C1, ZNF568, ABCB1, ELOVL2, PKIA, SFMBT2 (893), PCBP3, MATK, GRN2D, NDRG4, DLX4, PPP2R5C, FGF14, ZNF132, CHST2 (7890), FLI1, c13orf18, or ZNF569 (see Table 2). Further, embodiments provide a method for analyzing DMRs based on Table 2, where the DMR numbers are 253, 251, 254, 255, 256, 249, 257, 258, 259, 260, 261, 262, 250, 263, 1, 264, 265, 196, 266, 118, 267, 268, 269, 270, 271, 272, 46, 273, 252, or 237. Further, in such embodiments where the sample is a plasma sample, a method for analyzing DMRs based on Tables 1, 2, 4, and 7, where the DMR numbers are 253, 251, 254, 255, 256, 249, 257, 258, 259, 260, 261, 262, 250, 263, 1, 264, 265, 196, 266, 118, 267, 268, 269, 270, 271, 272, 46, 273, 252, or 237. In some embodiments, the above method includes determining the methylation state of two markers (e.g., a pair of markers shown in Tables 1, 2, 4, or 7).
[0042] Embodiments of kits (e.g., bisulfite reagents; and control nucleic acids comprising a sequence based on a DMR selected from the group consisting of DMR1-274 (based on Tables 1, 2, 4, or 7) and having a methylation state associated with a subject without cancer) are provided. In some embodiments, the kit comprises a bisulfite reagent and oligonucleotides as described herein. In some embodiments, the kit comprises a bisulfite reagent; and control nucleic acids comprising a sequence based on a DMR selected from the group consisting of DMR1-274 (based on Tables 1, 2, 4, or 7) and having a methylation state associated with a subject without cancer. Some kit embodiments include a sample collector for recovering a sample (e.g., a fecal sample, a gastric tissue sample, a plasma sample) from a subject; reagents for isolating nucleic acids from the sample; and oligonucleotides as described herein.
[0043] The above technology relates to embodiments of compositions (e.g., reaction mixtures). In some embodiments, compositions comprising a nucleic acid comprising a DMR and a bisulfite reagent are provided. Some embodiments provide compositions comprising a nucleic acid comprising a DMR and oligonucleotides as described herein. Some embodiments provide compositions comprising a nucleic acid comprising a DMR and a methylation-sensitive restriction enzyme. Some embodiments provide compositions comprising a nucleic acid comprising a DMR and a polymerase.
[0044] Embodiments of related further methods are provided for screening for neoplasms in a sample obtained from a subject (e.g., a gastric tissue sample, a plasma sample, a fecal sample). The method includes determining the methylation status of a marker in a sample containing a base in one or more DMRs of DMR1-274 (Tables 1, 2, 4, or 7); comparing the methylation status of the marker from the subject sample to the methylation status of the marker from a normal control sample from a subject without cancer; and determining a confidence interval and / or a p-value for the difference in methylation status between the subject sample and the normal control sample. In some embodiments, the confidence interval is 90%, 95%, 97.5%, 98%, 99%, 99.5%, 99.9%, or 99.99%, and the p-value is 0.1, 0.05, 0.025, 0.02, 0.01, 0.005, 0.001, or 0001. Some embodiments of the method include reacting a nucleic acid containing a DMR with a bisulfite reagent to generate a bisulfite-treated nucleic acid; sequencing the bisulfite-treated nucleic acid to show the nucleotide sequence of the bisulfite-treated nucleic acid; comparing the nucleotide sequence of the bisulfite-treated nucleic acid to the nucleotide sequence of a nucleic acid containing a DMR from a subject without cancer to identify a difference between the two sequences; and when there is a difference, identifying that the subject has a neoplasm.
[0045] A system for screening for neoplasms in a sample obtained from a subject is provided by the above technique. Exemplary embodiments of the system include, for example, a system for screening for neoplasms in a sample obtained from a subject (e.g., a gastric tissue sample, a plasma sample, a fecal sample). The system includes an analysis component configured to determine the methylation state of the sample, a software component for comparing the methylation state of the sample with the methylation state of a control sample or a reference sample recorded in a database, and a notification component configured to alert the user to a methylation state associated with cancer. In some embodiments, the notification is determined by software that receives results from a number of assays (e.g., determining the methylation state of a number of markers (such as those shown in Tables 1, 2, or 7)), and by calculating a value or result for reporting based on the number of results. Some embodiments provide a database of weighted variables associated with each DMR shown herein for use in calculating values or results and / or notifications for reporting to a user (e.g., a physician, a nurse, a clinician, etc.). In some embodiments, all results from a number of assays are recorded, and in some embodiments, one or more results are used to indicate a score, value, or result based on a composite of one or more results (representing the cancer risk in the subject) based on the number of assays.
[0046] In some embodiments of the system, the sample contains nucleic acids that include DMRs. In some embodiments, the system further includes components for isolating nucleic acids and components for collecting the sample (e.g., components for collecting fecal samples). In some embodiments, the system includes nucleic acid sequences that include DMRs. In some embodiments, the database includes nucleic acid sequences from subjects without cancer. Also, nucleic acids (e.g., a collection of multiple nucleic acids, each having a DMR) are provided. In some embodiments, a collection of multiple nucleic acids, each having a sequence from a subject without eyes. Related system embodiments include a collection of multiple nucleic acids as described and a database of multiple nucleic acid sequences associated with the collection of multiple nucleic acids. Some embodiments further include bisulfite reagents. Some embodiments further include a nucleic acid sequencer.
[0047] In some embodiments, a method for characterizing a sample from a human patient (e.g., a gastric tissue sample, a plasma sample, a fecal sample) is provided. For example, in some embodiments, such embodiments include obtaining DNA from a sample of a human patient; evaluating the methylation status of a DNA methylation marker that includes bases in a methylation variable region (DMR) selected from the group consisting of DMR1 - 274 based on Table 1, 2, or 7; and comparing the evaluated methylation status of one or more DNA methylation markers to methylation level criteria for one or more DNA methylation markers for human patients without gastric neoplasms.
[0048] Such methods are not limited to a particular type of sample from a human patient. In some embodiments, the sample is a gastric tissue sample. In some embodiments, the sample is a plasma sample. In some embodiments, the sample is a fecal sample, a tissue sample, a gastric tissue sample, a blood sample, or a urine sample.
[0049] In some embodiments, such a method includes evaluating a plurality of DNA methylation markers. In some embodiments, such a method includes evaluating from 12 to 107 DNA methylation markers. In some embodiments, such a method includes evaluating the methylation status of one or more DNA methylation markers in a sample, including evaluating the methylation status of a single nucleotide. In some embodiments, such a method includes evaluating the methylation status of one or more DNA methylation markers in a sample, including determining the degree of methylation status at multiple nucleotides. In some embodiments, such a method includes evaluating the methylation status of the forward strand or evaluating the methylation status of the reverse strand.
[0050] In some embodiments, the DNA methylation marker is a region of less than 100 bases. In some embodiments, the DNA methylation marker is a region of less than 500 bases. In some embodiments, the DNA methylation marker is a region of less than 1000 bases. In some embodiments, the DNA methylation marker is a region of less than 10,000 bases. In some embodiments, the DNA methylation marker is a single base. In some embodiments, the DNA methylation marker is a promoter with a high CpG density.
[0051] In some embodiments, evaluating includes using methylation-specific polymerase chain reaction, nucleic acid sequencing, mass spectrometry, methylation-specific nucleases, mass-based separation, or target capture.
[0052] In some embodiments, evaluating includes using methylation-specific oligonucleotides. In some embodiments, the methylation-specific oligonucleotides are selected from the group consisting of SEQ ID NOs: 1-109.
[0053] In some embodiments, the chromosomal region having an annotation selected from the group consisting of ARHGEF4, ELMO1, ABCB1, CLEC11A, ST8SIA1, SFMBT2, CD1D, CYP26C1, ZNF569 and C13ORF18 contains methylation markers. In some embodiments, the DMR is based on Table 4 and is selected from the group consisting of DMR numbers 49, 43, 196, 66, 1, 237, 249, 250, 251 and 252.
[0054] In some embodiments, the chromosomal region having an annotation selected from the group consisting of ELMO1, ARHGEF4, EMX1, SP9, CLEC11A, ST8SIA1, BMP3, KCNA3, DMRTA2, KCNK12, CD1D, PRKCB, CYP26C1, ZNF568, ABCB1, ELOVL2, PKIA, SFMBT2 (893), PCBP3, MATK, GRN2D, NDRG4, DLX4, PPP2R5C, FGF14, ZNF132, CHST2 (7890), FLI1, c13orf18 and ZNF569 contains DNA methylation markers.
[0055] In some embodiments, the DMR is based on Table 2 and is selected from the group consisting of DMR numbers 253, 251, 254, 255, 256, 249, 257, 258, 259, 260, 261, 262, 250, 263, 1, 264, 265, 196, 266, 118, 267, 268, 269, 270, 271, 272, 46, 273, 252 and 237.
[0056] In some embodiments where the obtained sample is a plasma sample, the marker contains a chromosomal region having an annotation of ELMO1, ZNF569, C13orf18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4 and ST8SIA1 that contains the marker.
[0057] In some embodiments, the DMR is based on Table 1, 2, 4 or 7 and is selected from the group consisting of 253, 237, 252, 261, 251, 196, 250, 265, 256, 249 and 274.
[0058] In some embodiments, such a method comprises determining the methylation status of two DNA methylation markers. In some embodiments, such a method comprises determining the methylation status of pairs of DNA methylation markers shown in Table 1, 2, 4 or 7.
[0059] In some embodiments, the above technique provides a method for characterizing a sample obtained from a human patient. In some embodiments, such a method comprises determining the methylation status of a DNA methylation marker in a sample comprising a base at a DMR selected from the group consisting of DMR1 - 274 based on Table 1, 2, 4 and 7; comparing the methylation status of the DNA methylation marker from the patient sample to the methylation status of the DNA methylation marker from a normal control sample from a human subject without cancer; and determining the confidence interval and p - value of the difference in methylation status between the human patient and the normal control sample. In some embodiments, the confidence interval is 90%, 95%, 97.5%, 98%, 99%, 99.5%, 99.9% or 99.99%, and the p - value is 0.1, 0.05, 0.025, 0.02, 0.01, 0.005, 0.001 or 0.0001.
[0060] In some embodiments, the technology provides a method for characterizing a sample (e.g., a gastric tissue sample, a plasma sample, a fecal sample) obtained from a human subject. The method includes reacting a nucleic acid containing a DMR with a bisulfite reagent to generate a bisulfite-treated nucleic acid; sequencing the bisulfite-treated nucleic acid to provide a nucleotide sequence of the bisulfite-treated nucleic acid; and comparing the nucleotide sequence of the bisulfite-treated nucleic acid with the nucleotide sequence of a nucleic acid containing a DMR from a subject without gastric cancer.
[0061] In some embodiments, the technology provides a system for characterizing a sample (e.g., a gastric tissue sample, a plasma sample, a fecal sample) obtained from a human subject. The system includes an analysis component configured to determine a methylation state of the sample, a software component configured to compare the methylation state of the sample with a control sample methylation state or a reference sample methylation state recorded in a database, and an alert component configured to determine a single value based on a combination of methylation states and alert a user of a methylation state associated with gastric cancer. In some embodiments, the sample includes a nucleic acid containing a DMR.
[0062] In some embodiments, such a system further includes a component for separating nucleic acids. In some embodiments, such a system further includes a component for collecting the sample.
[0063] In some embodiments, the sample is a fecal sample, a tissue sample, a gastric tissue sample, a blood sample, or a urine sample.
[0064] In some embodiments, the database includes nucleic acid sequences containing DMRs. In some embodiments, the database includes nucleic acid sequences from subjects without gastric cancer.
[0065] Further embodiments will be apparent to those of ordinary skill in the relevant art based on the teachings contained herein.
[0066] These and other features, aspects, and advantages of the present technology will be better understood with reference to the following figures.
[0067] It is understood that the drawings are not necessarily to scale, and that objects in the drawings are not scaled in relation to each other. The drawings are a depiction for the purpose of clarifying and understanding various embodiments of the apparatus, systems, compositions, and methods disclosed herein. In any case, the same reference numbers are used throughout the drawings to describe the same or similar parts. Further, the drawings are not at all intended to limit the scope of the present teachings.
Brief Description of the Drawings
[0068]
Figure 1
Figure 2
Figure 3
Figure 4A
Figure 4B
Figure 5A
Figure 5B
Figure 5C
Figure 6
Figure 7A
Figure 7B
Figure 7C
Figure 7D
Figure 7E
Figure 7F
Figure 7G
Figure 7H
Figure 7I
Figure 7J
Figure 7K
Figure 7L
Figure 7M
DETAILED DESCRIPTION OF THE INVENTION
[0069] Disclosed herein is technology related to methods, compositions, and related uses (among others) for detecting tumorigenesis and, more particularly, for detecting pre-cancerous tumors and malignancies (such as gastric cancer). As the technology is described herein, the headings of the sections used are for purposes of organization only and should not be construed in any way as limiting the subject matter.
[0070] In the detailed description of the various embodiments, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the disclosed embodiments. Those skilled in the art will correctly understand that these various embodiments can be practiced with or without these specific details. In other instances, structures and devices are shown in block diagram form. Further, those skilled in the art can readily and correctly understand that the particular order in which the various methods are presented and executed is an example and that this order can be changed and still be within the spirit and scope of the various embodiments disclosed herein.
[0071] Without limitation, all documents and similar materials cited in this application, including patents, patent applications, articles, books, papers, and Internet web pages, are expressly incorporated herein by reference for any purpose. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the various embodiments described herein belong. When the definition of a term in an incorporated reference is clearly different from the definition given in this teaching, the definition given in this teaching shall control.
[0072] 〔Definitions〕 To facilitate understanding of the present invention, a number of terms and phrases are defined below. Additional definitions will be explained in the detailed description.
[0073] In the specification and claims, unless the context clearly indicates otherwise, the following terms have the meanings expressly associated with them in this specification. As used herein, the phrase "in one embodiment" may, but need not necessarily, refer to the same embodiment. Further, as used herein, the phrase "in other embodiments" may, but need not necessarily, refer to different embodiments. Thus, as described below, the various embodiments of the present invention may be readily combined without departing from the scope or spirit of the present invention.
[0074] Furthermore, as used herein, the term "or" is an inclusive disjunctive term and is synonymous with the term "and / or" unless the context clearly dictates otherwise. The term "based on" is not exclusive and takes into account additional elements not recited unless the context clearly dictates otherwise. Further, in the specification, the meanings of "a", "an", and "the" include the plural forms. The meaning of "in" includes "in" and "on".
[0075] As used herein, "nucleic acid" or "nucleic acid molecule" generally refers to ribonucleic acid or deoxyribonucleic acid, which can be unmodified or modified DNA or RNA. "Nucleic acid" includes, without limitation, single-stranded and double-stranded nucleic acids. As used herein, the term "nucleic acid" also includes DNA containing one or more modified bases as described above. Thus, DNA having a modified backbone for stability or other reasons is a "nucleic acid". As used herein, the term "nucleic acid" encompasses not only the chemical form of DNA characteristic of viruses and cells (including, for example, single cells and complex-type cells), but also forms of nucleic acids that are chemically, enzymatically, or metabolically modified.
[0076] The terms "oligonucleotide" or "polynucleotide" or "nucleotide" or "nucleic acid" refer to a molecule having two or more deoxyribonucleotides or ribonucleotides, preferably three or more, and usually more than ten. The exact size can depend on a number of factors and thus depends on the ultimate function or use of the oligonucleotide.
[0077] Oligonucleotides can be made in any of a number of ways including chemical synthesis, DNA replication, reverse transcription, or combinations thereof. The typical deoxyribonucleotides of DNA are thymine, adenine, cytosine, and guanine. The typical ribonucleotides of RNA are uracil, adenine, cytosine, and guanine.
[0078] As used herein, a "locus" or "region" of a nucleic acid refers to a sub-region of the nucleic acid, such as, for example, a gene on a chromosome, a single nucleotide, a CpG island, etc.
[0079] The terms "complementary" and "complementarity" refer to nucleotides (e.g., a single nucleotide) or polynucleotides (e.g., contiguous nucleotides) that are associated by base pairing rules. For example, the sequence "5'-A-G-T-3'" is complementary to the sequence "3'-T-C-A-5'". Complementarity may be "partial", where only some of the nucleic acid bases conform to the base pairing rules. Alternatively, "complete" or "total" complementarity may exist between nucleic acids. The degree of complementarity between nucleic acid strands affects the efficiency and strength of hybridization between the nucleic acid strands. This is particularly important in amplification reactions and detection methods that rely on binding between nucleic acids.
[0080] The term "gene" refers to a nucleic acid (e.g., DNA or RNA) sequence that contains the coding sequence necessary for the production of an RNA, polypeptide, or precursor thereof. A functional polypeptide can be encoded by the full-length coding sequence or any portion of the coding sequence, as long as the desired activity or functional property of the polypeptide (e.g., enzymatic activity, ligand binding, signal transduction, etc.) is retained. The term "portion", when used in reference to a gene, refers to a fragment of that gene. A fragment can range in size from a few nucleotides to up to the full gene minus one nucleotide. Thus, "a nucleotide containing at least one portion of a gene" may include multiple fragments of the gene or may include the complete gene.
[0081] The term "gene" also encompasses the coding region of a structural gene, as well as the sequences flanking the coding region at both the 5' and 3' termini, for example, sequences located adjacent to the coding region over a distance of about 1 kb at either terminus, in which case the gene corresponds to the length of the full-length mRNA (including, for example, coding sequences, regulatory sequences, structural sequences, and other sequences). The sequence located 5' to the coding region and present on the mRNA is referred to as the 5' non-translated or untranslated sequence. The sequence located 3' or downstream of the coding region and present on the mRNA is referred to as the 3' non-translated or untranslated sequence. The term "gene" encompasses both cDNA and the genomic form of the gene. In some organisms (e.g., eukaryotes), the genomic form or clone of a gene contains coding regions interrupted by non-coding sequences referred to as "introns" or "intervening regions" or "intervening sequences". Introns are segments of a gene that are transcribed into nuclear RNA (hnRNA), and introns may contain regulatory elements such as enhancers. Introns are removed or "spliced out" from the nucleus or the primary transcript, and thus do not exist in messenger RNA (mRNA) transcripts. mRNA functions to specify the sequence or order of amino acids in the nascent polypeptide during translation.
[0082] In addition to containing introns, the genomic form of a gene may also include sequences located at both the 5' and 3' termini of the sequences present on the RNA transcript. These sequences are referred to as "flanking" sequences or regions (these flanking sequences are located 5' or 3' to the non-translated sequences present on the mRNA transcript). The 5' flanking region may contain regulatory sequences such as promoters and enhancers that control or influence the transcription of the gene. The 3' flanking region may contain sequences that direct transcription termination, post-transcriptional cleavage, and polyadenylation.
[0083] The term "wild type" refers to a gene that, when constituted with respect to a gene, has the characteristics of a gene isolated from a naturally occurring source. The term "wild type" refers to a gene product that, when constituted with respect to a gene product, has the characteristics of a gene product isolated from a naturally occurring source. The term "naturally occurring" refers to the fact that a subject can be found in nature when used with respect to a subject. For example, a polypeptide sequence or polynucleotide sequence that exists in an organism (including a virus), can be isolated from a natural source, and has not been intentionally modified by human hands in a laboratory is naturally occurring. A wild-type gene is often the gene or allele that is most frequently observed in a population, and thus is the randomly designed "normal" or "wild-type" form of that gene. In contrast, the terms "modified" or "variant" refer to a gene or gene product that, when constituted with respect to a gene or gene product, respectively, exhibits a modification (e.g., a changed characteristic) in sequence and / or functional characteristics when compared to the wild-type gene or gene product. Note that naturally occurring mutants can be isolated, and they are identified by the fact that they have characteristics that are changed when compared to the wild-type gene or gene product.
[0084] The term "allele" refers to a variant of a gene, which includes, but is not limited to, mutant forms and variants, polymorphic loci, as well as single nucleotide polymorphic loci, frameshifts, and splice variants. Alleles can occur naturally in a population or can occur during the lifetime of a particular individual in any population.
[0085] Thus, the terms "mutant form" and "variant" refer to a nucleotide sequence that differs by one or more nucleotides from another, usually related, nucleotide sequence when used with respect to a nucleotide sequence. A "variant" is the difference between two different nucleotide sequences, typically one of the sequences being a reference sequence.
[0086] "Amplification" is a special nucleic acid replication involving template specificity. This is in contrast to non-specific template replication (e.g., replication that is template-dependent but does not depend on a specific template). Template specificity is here distinguished from replication fidelity (e.g., synthesis of the appropriate nucleotide sequence) and nucleotide (ribo- or deoxyribo-) specificity. Template specificity is frequently described in terms of "target" specificity. The target sequence is "target" in the sense that it is required to be selected from other nucleic acids. Amplification techniques are mainly designed for this selection.
[0087] The amplification of nucleic acids generally refers to the generation of multiple copies of a polynucleotide or a portion of a polynucleotide, typically starting from a small amount of polynucleotides (e.g., a single polynucleotide molecule, 10 to 100 copies of a polynucleotide molecule, which may or may not be exactly identical), and the amplification product or amplicon is generally detectable. The amplification of polynucleotides encompasses various chemical and enzymatic processes. The generation of multiple DNA copies from one or a few copies of a target or template DNA molecule during polymerase chain reaction (PCR) or ligase chain reaction (LCR, see, e.g., U.S. Patent No. 5,494,810, which is hereby incorporated by reference in its entirety) is a form of amplification. Further types of amplification include, but are not limited to, allele-specific PCR (see, e.g., U.S. Patent No. 5,639,611, which is hereby incorporated by reference in its entirety), assembly PCR (see, e.g., U.S. Patent No. 5,965,408, which is hereby incorporated by reference in its entirety), helicase-dependent amplification (see, e.g., U.S. Patent No. 7,662,594, which is hereby incorporated by reference in its entirety), hot start PCR (see, e.g., U.S. Patents Nos. 5,773,258 and 5,338,671, each of which is hereby incorporated by reference in its entirety), intersequence-specific PCR, inverse PCR (see, e.g., Triglia, et al. (1988) Nucleic Acids Res., 16:8186, which is hereby incorporated by reference in its entirety), ligation-mediated PCR (see, e.g., Guilfoyle, R., et al., Nucleic Acids Research, 25:1854-1858 (1997), U.S. Patent No. 5,508,169, each of which is hereby incorporated by reference in its entirety), methylation-specific PCR (see, e.g., Herman, et al., (1996) PNAS 93(13)9821-9826, which is hereby incorporated by reference in its entirety), miniprimer PCR, multiplex ligation-dependent probe amplification (see, e.g., Schouten, et al., see (2002) Nucleic Acids Research 30(12):e57, which is incorporated herein by reference in its entirety), multiplex PCR (e.g., see Chamberlain, et al., (1988) Nucleic Acids Research 16(23)11141-11156, Ballabio, et al., (1990) Human Genetics 84(6)571-573, Hayden, et al., (2008) BMC Genetics 9:80, each of which is incorporated herein by reference in its entirety), nested PCR, overlap extension PCR (e.g., see Higuchi, et al., (1988) Nucleic Acids Research 16(15)7351-7367, which is incorporated herein by reference in its entirety), real-time PCR (e.g., see Higuchi, et al., (1992) Biotechnology 10:413-417, Higuchi, et al., (1993) Biotechnology 11:1026-1030, each of which is incorporated herein by reference in its entirety), reverse transcription PCR (e.g., see Bustin, S.A. (2000) J. Molecular Endocrinology 25:169-193, which is incorporated herein by reference in its entirety), solid-phase PCR, thermal asymmetric interlaced PCR, and touchdown PCR (e.g., see Don, et al., Nucleic Acids Research (1991) 19(14)4008, Roux, K. (1994) Biotechniques 16(5)812-814, Hecker, et. Reference is made to al., (1996) Biotechniques 20(3) 478-485, each of which is hereby incorporated by reference in its entirety), which is included. Polynucleotide amplification can also be achieved using digital PCR (see, for example, Kalinina, et al., Nucleic Acids Research. 25; 1999-2004, (1997), Vogelstein and Kinzler, Proc Natl Acad Sci USA. 96; 9236-41, (1999), International Patent Application Publication No. 05023091A2, US Patent Application Publication No. 20070202525, each of which is hereby incorporated by reference in its entirety).
[0088] The term "polymerase chain reaction" ("PCR") refers to the methods of U.S. Patent Nos. 4,683,195, 4,683,202, and 4,965,188 to K.B. Mullis that describe a method for increasing the concentration of a segment of a target sequence in a mixture of genomic DNA without cloning or purification. This process for amplifying a target sequence consists of introducing two highly excess oligonucleotide primers into a DNA mixture containing the desired target sequence, followed by performing thermal cycling in the correct order in the presence of a DNA polymerase. The two primers are complementary to each strand of the double-stranded target sequence. To effect amplification, the mixture is denatured and then the primers are annealed to their complementary sequences within the target molecule. After annealing, the primers are extended with a polymerase to form a new set of complementary strands. The steps of denaturation, primer annealing, and polymerase extension are repeated many times (i.e., denaturation, annealing, and extension constitute one "cycle" and there can be a number of "cycles") to obtain an amplified segment of the desired target sequence at a high concentration. The length of the amplified segment of the desired target sequence is determined by the relative positions of the primers with respect to each other and is thus a controllable parameter. Because of the repetitive nature of the process, this method is referred to as the "polymerase chain reaction" ("PCR"). Since the amplified segments of the desired target sequence become the dominant sequences (with respect to concentration) in the mixture, they are referred to as "amplified by PCR" and are "PCR products" and "amplicons".
[0089] Template specificity is achieved, in most amplification techniques, by enzyme selection. Amplification enzymes are enzymes that process only specific sequences of nucleic acids in a heterogeneous mixture of nucleic acids under the conditions in which they are used. For example, in the case of Q-beta replicase, MDV-1 RNA is the specific template for this replicase (Kacian et al., Proc. Natl. Acad. Sci. USA, 69:3038
[1972] ). Other nucleic acids are not replicated by this amplification enzyme. Similarly, in the case of T7 RNA polymerase, this amplification enzyme has stringent specificity for its own promoter (Chamberlin et al, Nature, 228:227
[1970] ). In the case of T4 DNA ligase, this enzyme does not ligate two oligonucleotides or polynucleotides where there is a mismatch at the ligation junction between this oligonucleotide or polynucleotide substrate and this template (Wu and Wallace (1989) Genomics 4:560). Finally, thermostable template-dependent DNA polymerases (e.g., Taq and Pfu DNA polymerases) exhibit high specificity for the bound sequence with respect to their ability to function at high temperatures, and as a result it has been found that they are defined by the primer, and the high temperature creates thermodynamic conditions that are favorable for primer hybridization to the target sequence but not for hybridization to non-target sequences (H. A. Erlich (ed.), PCR Technology, Stockton Press
[1989] ).
[0090] As used herein, the term "nucleic acid detection assay" refers to any method for determining the nucleotide composition of a nucleic acid of interest. Nucleic acid detection assays include, but are not limited to, DNA sequencing methods, probe hybridization methods, structure-specific cleavage assays (e.g., the INVADER assay (Hologic, Inc.)), and are described, for example, in U.S. Patent Nos. 5,846,717, 5,985,557, 5,994,069, 6,001,567, 6,090,543, and 6,872,816, Lyamichev et al., Nat. Biotech., 17:292 (1999), Hall et al., PNAS, USA, 97:8272 (2000), and U.S. Patent Application Publication No. 2009 / 0253142, each of which is hereby incorporated by reference in its entirety for all purposes), enzyme mismatch cleavage methods (e.g., Variagenics, U.S. Patent Nos. 6,110,684, 5,958,692, 5,851,770, which are hereby incorporated by reference in their entirety), polymerase chain reaction, branched hybridization formation methods (e.g., Chiron, U.S. Patent Nos. 5,849,481, 5,710,264, 5,124,246, and 5,624,802, which are hereby incorporated by reference in their entirety), rolling circle replication (e.g., U.S. Patent Nos. 6,210,884, 6,183,960, and 6,235,502, which are hereby incorporated by reference in their entirety), NASBA (e.g., U.S. Patent No. 5,409,818, which is hereby incorporated by reference in its entirety), molecular beacon technology (e.g., U.S. Patent No. 6,150,097, which is hereby incorporated by reference in its entirety), E sensor technology (Motorola, U.S. Patent Nos. 6,248,229, 6,221,583, 6,013,170, and 6,063,573, which are hereby incorporated by reference in their entirety), cycling probe technology (e.g., U.S. Patent Nos. 5,403,711, 5,011,769, and 5,660,988, which are hereby incorporated by reference in their entirety), Dade Behring signal amplification methods (e.g., U.S. Patent Nos. 6,121,001, 6,110,677, 5,914,230, 5,882,867, and 5,792,614, which are hereby incorporated by reference in their entirety), ligase chain reaction (e.g., Barnay Proc. Natl. Acad. Sci USA 88, 189-93 (1991)), and sandwich hybridization formation methods (e.g., U.S. Patent No. 5,288,609, which is hereby incorporated by reference in its entirety) are included.
[0091] The term "amplifiable nucleic acid" is used with respect to nucleic acids that can be amplified by any amplification method. "Amplifiable nucleic acids" are usually intended to include "sample templates".
[0092] The term "sample template" refers to nucleic acids originating from a sample that is analyzed for the presence of a "target" (defined below). In contrast, the term "background template" is used with respect to nucleic acids other than the sample template that may or may not be present in the sample. Background templates are most frequently inadvertent. This can be the result of carryover or can be due to the presence of nucleic acid contaminants that were required to be purified away from the sample. For example, nucleic acids from organisms other than those to be detected may be present as background in the test sample.
[0093] The term "primer" refers to an oligonucleotide that can function as an initiation point when placed under conditions (e.g., in the presence of inducers such as nucleotides and DNA polymerase, at a suitable temperature and pH) that induce the synthesis of a primer extension product complementary to a nucleic acid strand, whether it is naturally occurring or synthetically produced, such as in a purified restriction digest. The primer is preferably single-stranded for maximum amplification efficiency, but alternatively may be double-stranded. If double-stranded, the primer is first treated to separate the strands before being used in the preparation of the extension product. Preferably, the primer is an oligodeoxyribonucleotide. The primer must be of sufficient length to initiate the synthesis of the extension product in the presence of an inducer. The exact length of the primer will depend on a number of factors, including temperature, primer source, and the method used.
[0094] The term "probe" refers to an oligonucleotide (e.g., a nucleotide sequence) that occurs naturally in a purified restriction digest or is made by synthesis, recombination, or PCR amplification and is capable of hybridizing to another oligonucleotide of interest. The probe may be single-stranded or double-stranded. Probes are useful for the detection, identification, and isolation of specific gene sequences (e.g., "capture sequences"). The probes used in the present invention are, in some embodiments, intended to be labeled with a "reporter molecule" such that they are detectable with any detection system, including but not limited to enzyme (e.g., ELISA in addition to enzyme-based histochemical assays), fluorescence, radioactivity, and luminescence systems. The present invention is not intended to be limited to a particular detection system or label.
[0095] As used herein, "methylation" refers to the methylation of cytosine at the C5 or N4 position of cytosine, adenine, or the N6 position of other types of nucleic acids. Since typical in vitro DNA amplification methods do not preserve the methylation pattern of the amplification template, DNA amplified in vitro is usually not methylated. However, "unmethylated DNA" or "methylated DNA" also refers to amplified DNA whose original template was unmethylated or methylated, respectively.
[0096] Accordingly, as used herein, "methylated nucleotide" or "methylated nucleotide base" refers to the presence of a methyl residue on a nucleotide base, where the methyl residue is not present on the typical nucleotide bases that are recognized. For example, cytosine does not contain a methyl residue on its pyrimidine ring, but 5-methylcytosine contains a methyl residue at the 5-position of its pyrimidine ring. Thus, cytosine is not a methylated nucleotide, while 5-methylcytosine is a methylated nucleotide. In other examples, thymine contains a methyl residue at the 5-position of its pyrimidine ring, but for the purposes herein, since thymine is a typical DNA nucleotide base, thymine is not considered a methylated nucleotide when present in DNA.
[0097] As used herein, a "methylated nucleic acid molecule" refers to a nucleic acid molecule that contains one or more methylated nucleotides.
[0098] As used herein, the "methylation state", "methylation profile" and "methylation status" of a nucleic acid molecule refer to the presence or absence of one or more methylated nucleotide bases in the nucleic acid molecule. For example, a nucleic acid molecule containing methylated cytosine is considered to be methylated (e.g., the methylation state of the nucleic acid molecule is methylated). A nucleic acid molecule that does not contain any methylated nucleotides is considered to be unmethylated.
[0099] The methylation state of a particular nucleic acid sequence (e.g., when described herein, a gene marker or DNA region) can indicate the methylation state of all bases in the sequence, or can indicate the methylation state of a portion of the bases within the sequence (e.g., one or more cytosines), or can indicate information regarding the local methylation density within the sequence, with or without giving the exact location within the sequence where methylation is present.
[0100] The methylation state of a nucleotide locus in a nucleic acid molecule refers to the presence or absence of a methylated nucleotide at a particular locus in the nucleic acid molecule. For example, when the nucleotide present at the 7th nucleotide in a nucleic acid molecule is 5-methylcytosine, the methylation state of the cytosine at the 7th nucleotide of the nucleic acid molecule is methylated. Similarly, when the nucleotide present at the 7th nucleotide in a nucleic acid molecule is cytosine (not 5-methylcytosine), the methylation state of the cytosine at the 7th nucleotide of the nucleic acid molecule is not methylated.
[0101] The methylation status can optionally be represented or indicated by a "methylation value" (e.g., representing methylation frequency, fraction, ratio, percentage, etc.). The methylation value can be generated, for example, by quantifying the amount of the nucleic acid as it is after subjecting it to the following restriction digestion with a methylation-dependent restriction enzyme, or by comparing the amplification profiles after bisulfite reaction, or by comparing the sequences of bisulfite-treated and untreated nucleic acids. Thus, a value, such as a methylation value, represents the methylation status and can therefore be used as an indicator of the amount of the methylation status across multiple copies of a locus. This is particularly useful when it is desirable to compare the methylation status of sequences in a sample to a threshold or reference value.
[0102] As used herein, "methylation frequency" or "methylation percentage (%)" refers to the number of instances in which a molecule or locus is methylated compared to the number of instances in which the molecule or locus is not methylated.
[0103] Thus, the methylation state represents the methylation state of a nucleic acid (e.g., a genomic sequence). Further, the methylation state refers to the characteristics of a nucleic acid fragment at a specific genomic locus associated with methylation. The above characteristics include, but are not limited to, whether any cytosine (C) residue within this DNA sequence is methylated, the position of the methylated C residue, the frequency or ratio of methylated C in any specific region of the nucleic acid, and differences in alleles in methylation factors, such as differences in the origin of alleles. The terms "methylation state", "methylation profile", and "methylation status" also refer to the pattern of methylated or unmethylated C in any specific region of a nucleic acid in a biological sample in terms of relative concentration, absolute concentration. For example, if the cytosine (C) sequence within a nucleic acid sequence is methylated, it can be termed as having "hypermethylated" or "increased methylation", while if the cytosine (C) residue within a DNA sequence is not methylated, it can be termed as having "hypomethylated" or "decreased methylation". Similarly, when compared with other nucleic acid sequences (e.g., different regions or different individuals, etc.), if the cytosine (C) residue within a nucleic acid sequence is methylated, that sequence is considered to have hypermethylated or increased methylation compared to other nucleic acid sequences. Also, when compared with other nucleic acid sequences (e.g., different regions or different individuals, etc.), if the cytosine (C) residue within a DNA sequence is not methylated, that sequence is considered to have hypomethylated or decreased methylation compared to other nucleic acid sequences. Further, the term "methylation pattern", as used herein, refers to the set sites of methylated and unmethylated nucleotides with respect to the sequence of a nucleic acid. Two nucleic acids may have the same or similar methylation frequency or methylation ratio, but may have different methylation patterns when the positions of methylated and unmethylated nucleotides are different while the number of methylated and unmethylated nucleotides in the region is the same or similar.An array is said to have "differentially methylated" or "differences in methylation" or "different methylation states" when the range (having increased or decreased methylation compared to others), frequency, or pattern of methylation is different. The term "differential methylation" refers to the difference in the level or pattern of nucleic acid methylation in a cancer-positive sample compared to the level or pattern of nucleic acid methylation in a cancer-negative sample when compared. It can also be referred to as the difference in level or pattern between patients in whom cancer has recurred after surgery and those in whom it has not. Specific levels or patterns of differential methylation and DNA methylation are predictive and prognostic biomarkers, for example, and have previously had their accurate classification or predictive characteristics elucidated.
[0104] Methylation state frequency can be used to represent a sample from a group of multiple individuals or a single individual. For example, a nucleotide locus having a methylation state frequency of 50% is methylated in 50% of the examples and not methylated in 50% of the examples. The above frequency can be used, for example, to represent the degree to which a nucleotide locus or nucleic acid region is methylated in a group of individuals or a collection of nucleic acids. Thus, when the methylation in a first group or pool of nucleic acid molecules is different from the methylation in a second group or pool of nucleic acid molecules, the methylation state frequency of the first group or pool can be different from the methylation state frequency of the second group or pool. The above frequency can also be used, for example, to represent the degree to which a nucleotide locus or nucleic acid region is methylated in a single individual. For example, the above frequency can be used to represent the degree to which a population of cells from a tissue sample is methylated or not methylated at a nucleotide locus or nucleic acid region.
[0105] As used herein, "nucleotide locus" refers to the position of a nucleotide in a nucleic acid molecule. A nucleotide locus of a methylated nucleotide refers to the position of the methylated nucleotide in a nucleic acid molecule.
[0106] Typically, methylation of human DNA occurs on dinucleotide sequences containing adjacent guanine and cytosine, with the cytosine located 5' to the guanine (also referred to as CpG dinucleotide sequences). Most cytosines within CpG dinucleotides are methylated in the human genome, although some remain unmethylated in specific CpG dinucleotide-rich genomic regions known as CpG islands (see, for example, Antequera et al. (1990) Cell 62:503-514).
[0107] As used herein, "CpG island" refers to a G:C rich region of genomic DNA that contains an increased number of CpG dinucleotides compared to the genomic DNA as a whole. The CpG island may be at least 100, 200, or more base pairs in length, the G:C content of the region is at least 50%, the ratio of the observed CpG frequency to the predicted frequency is 0.6, and in some instances, the CpG island may be at least 500 base pairs in length, the G:C content of the region is at least 55%, and the ratio of the observed CpG frequency to the predicted frequency is 0.65. The ratio of the observed CpG frequency to the predicted frequency can be calculated by the method defined by Gardiner-Garden et al (1987) J. Mol. Biol. 196: 261-281. For example, the observed CpG frequency relative to the predicted frequency can be calculated by the formula R=(A×B) / (C×D), where R is the ratio of the observed CpG frequency to the predicted frequency, A is the number of CpG dinucleotides in the sequence being analyzed, B is the total number of nucleotides in the sequence being analyzed, C is the total number of C nucleotides in the sequence being analyzed, and D is the total number of G nucleotides in the sequence being analyzed. The methylation state is typically determined in CpG islands, for example, in the promoter region. Although other sequences in the human genome are prone to DNA methylation such as CpA and CpT, they can be evaluated (see Ramsahoye (2000) Proc. Natl. Acad. Sci. USA 97:5237-5242; Salmon and Kaye (1970) Biochim. Biophys. Acta. 204:340-351; Grafstrom (1985) Nucleic Acids Res. 13:2827-2842; Nyce (1986) Nucleic Acids Res. 14:4353-4367; Woodcock (1987) Biochem. Biophys. Res. Commun. 145:888-894).
[0108] As used herein, a reagent that modifies nucleotides of a nucleic acid molecule in correlation with the methylation state of the nucleic acid molecule, or a methylation-specific reagent, refers to a compound or composition or other agent that can change the nucleotide sequence of a nucleic acid molecule in a manner that reflects the methylation state of the nucleic acid molecule. A method of treating a nucleic acid molecule with the above reagent involves contacting the nucleic acid molecule with the reagent and, if necessary, combining with additional steps to effect a desired change in the nucleotide sequence. The above changes in the nucleotide sequence of the nucleic acid molecule can result in a nucleic acid molecule in which each methylated nucleotide is modified to a different nucleotide. The above changes in the nucleic acid nucleotide sequence can result in a nucleic acid molecule in which each unmethylated nucleotide is modified to a different nucleotide. The above changes in the nucleic acid nucleotide sequence can result in a nucleic acid molecule in which each selected unmethylated nucleotide (e.g., each unmethylated cytosine) is modified to a different nucleotide. The use of the above reagent to change the nucleic acid nucleotide sequence can result in a nucleic acid molecule in which each nucleotide that is a methylated nucleotide (e.g., each methylated cytosine) is modified to a different nucleotide. As used herein, the use of a reagent that modifies a selected nucleotide refers to a reagent that modifies one of the four nucleotides (C, G, T, and A in the case of DNA, and C, G, U, and A in the case of RNA) that are typically present in a nucleic acid molecule, such as a reagent that modifies one nucleotide without modifying the other three nucleotides. In one typical example, the above reagent results in a different nucleotide by modifying a selected unmethylated nucleotide. In another typical example, the above reagent can deaminate an unmethylated cytosine nucleotide. A typical reagent is bisulfite.
[0109] As used herein, the term "bisulfite reagent" refers to a reagent comprising bisulfite, disulfite, hydrogen sulfite, or combinations thereof, in some embodiments, for example, in a CpG dinucleotide sequence, to distinguish methylated cytosine from unmethylated cytosine.
[0110] The term "methylation assay" refers to an assay for determining the methylation state of one or more CpG dinucleotide sequences within a nucleic acid sequence.
[0111] The term "MS AP-PCR" (Methylation-Sensitive Arbitrarily-Primed Polymerase Chain Reaction) refers to a technique recognized in the art that takes into account a global scan of the genome using CG-rich primers to align regions most likely to contain CpG dinucleotides, as described in Gonzalgo et al. (1997) Cancer Research 57:594-599.
[0112] The term "MethyLight™" refers to a fluorescence-based real-time PCR technique recognized in the art as described in Eads et al. (1999) Cancer Res. 59:2302-2306.
[0113] The term "HeavyMethyl™" refers to an assay that covers CpG positions between amplification primers that enable methylation-specific selective amplification of a nucleic acid sample, or an assay of methylation-specific blocking probes (also referred to herein as blockers) covered by such primers.
[0114] The term "HeavyMethyl™ MethyLight™" assay refers to a variant of the MethyLight™ assay that combines methylation-specific blocking probes that cover CpG positions between amplification primers, i.e., the HeavyMethyl™ MethyLight™ assay.
[0115] The term "Ms-SNuPE" (methylation-sensitive single nucleotide primer extension) refers to an assay recognized in the art described in Gonzalgo & Jones (1997) Nucleic Acids Res. 25:2529-2531.
[0116] The term "MSP" (methylation-specific PCR) refers to a methylation assay recognized in the art described in Herman et al. (1996) Proc. Natl. Acad. Sci. USA 93:9821-9826 and U.S. Pat. No. 5,786,146.
[0117] The term "COBRA" (combined bisulfite restriction analysis) refers to a methylation assay recognized in the art described in Xiong & Laird (1997) Nucleic Acids Res. 25:2532-2534.
[0118] The term "MCA" (methylated CpG island amplification) refers to a methylation assay described in Toyota et al. (1999) Cancer Res. 59: 2307-12, and WO 00 / 26401 A1.
[0119] As used herein, "selected nucleotide" refers to one of the four nucleotides typically present in a nucleic acid molecule (C, G, T, and A in the case of DNA, and C, G, U, and A in the case of RNA), and may include methylated derivatives of the typically present nucleotides (e.g., when C is the selected nucleotide, both methylated and non-methylated C are included in the meaning of the selected nucleotide). Methylated selected nucleotides specifically refer to methylated typically present nucleotides, and non-methylated selected nucleotides specifically refer to non-methylated typically present nucleotides.
[0120] The term "methylation-specific restriction enzyme" or "methylation-selective restriction enzyme" refers to an enzyme that selectively digests nucleic acids depending on the methylation state of its recognition site. In the case of a restriction enzyme that specifically cleaves when the recognition site is not methylated or is hemimethylated, cleavage may not occur or may occur with a significantly reduced effect when the recognition site is methylated. In the case of a restriction enzyme that specifically cleaves when the recognition site is methylated, cleavage may not occur or may occur with a significantly reduced effect when the recognition site is not methylated. Preferably, it is a methylation-specific restriction enzyme with a recognition sequence containing a CG dinucleotide (such as a recognition sequence like CGCG or CCCGGG in the examples). Some more preferred embodiments are restriction enzymes that do not cleave when the cytosine in the above dinucleotide is methylated at carbon atom C5.
[0121] As used herein, "different nucleotide" refers to a nucleotide that is chemically different from a selected nucleotide, typically where the different nucleotide has Watson-Crick base pairing properties different from those of the selected nucleotide and the typically occurring nucleotide that is complementary to the selected nucleotide is not identical to the typically occurring nucleotide that is complementary to the different nucleotide. For example, when C is the selected nucleotide, U or T may be a different nucleotide, as exemplified by the complementarity of C to G and of U or T to A. As used herein, a nucleotide that is complementary to a selected nucleotide or to a different nucleotide forms a base pair with the selected nucleotide or the different nucleotide that has a higher affinity than base pairing of the complementary nucleotide with three of the four typically occurring nucleotides under high stringency conditions. An example of complementarity is Watson-Crick base pairing in DNA (e.g., A-T and C-G) and RNA (e.g., A-U and C-G). Thus, for example, under high stringency conditions, G forms a base pair with C with a higher affinity than with G, A, or T, such that when C is the selected nucleotide, G is the nucleotide complementary to the selected nucleotide.
[0122] As used herein, the "sensitivity" of a given marker refers to the proportion of samples that report a DNA methylation value above a threshold that distinguishes between samples of neoplasms and samples that are not neoplasms. In some embodiments, positive is defined as a neoplasm confirmed in a tissue that reports a DNA methylation value above the threshold (e.g., in a range associated with a disease), and false negative is defined as a neoplasm confirmed in a tissue that reports a DNA methylation value below the threshold (e.g., in a range associated with non-disease). The value of sensitivity thus reflects the probability that a DNA methylation measurement of a given marker obtained from a known disease sample can be within the range of the disease associated with the measurement. As defined herein, the clinical relevance of the calculated sensitivity value represents an estimate of the probability that a given marker can detect the presence of a clinical symptom when used in a symptomatic subject.
[0123] As used herein, the “specificity” of a given marker refers to the proportion of non-neoplastic samples that report DNA methylation values below the threshold that distinguishes neoplastic samples from non-neoplastic samples. In some embodiments, negative is defined as a non-neoplastic sample confirmed in a tissue that reports a DNA methylation value below the threshold (e.g., in a range associated with non-disease), and false positive is defined as a non-neoplastic sample confirmed in a tissue that reports a DNA methylation value above the threshold (e.g., in a range associated with disease). The value of specificity thus reflects the probability that a DNA methylation measurement of a given marker obtained from a known non-disease sample can be within the non-disease range associated with the measurement. As defined herein, the clinical relevance of the calculated specificity value, when used in asymptomatic patients, represents an estimate of the probability that a given marker can detect the absence of clinical symptoms.
[0124] The term “AUC” as used herein is an abbreviation for “area under a curve”. In particular, AUC refers to the area under the receiver operating characteristic (ROC) curve. The ROC curve is a plot of the true positive rate against the false positive rate for different possible cut-off points of a diagnostic test. The ROC curve shows the trade-off between sensitivity and specificity that depends on the selected cut-off point (any increase in sensitivity can be accompanied by a decrease in specificity). The area under the ROC curve (AUC) is a measure of the accuracy of a diagnostic test (the larger the area, the better; the maximum is 1; a random test can have a diagonal ROC curve with an area of 0.5; see: J. P. Egan. (1975) Signal Detection Theory and ROC Analysis, Academic Press, New York).
[0125] As used herein, the term "neoplasm" refers to "an abnormal mass of tissue, excessive growth, not in harmony with the growth of normal tissue". See, for example, "The Spread of Tumors in the Human Body", London, Butterworth & Co, 1952.
[0126] As used herein, the term "adenoma" refers to a benign tumor of glandular origin. Although these growths are benign, over time, they may evolve into malignancy.
[0127] The terms "preneoplastic" or "preneoplastic" and their synonyms refer to any cell proliferation disorder that has undergone malignant transformation.
[0128] The "site" of a neoplasm, adenoma, cancer, etc. is a region of tissue, organ, cell type, anatomical structure, part of the body, etc. in the body of the subject where the neoplasm, adenoma, cancer, etc. is located.
[0129] As used herein, a "diagnostic" test method includes detecting or identifying the disease state or symptoms of a subject, which determines the likelihood that the subject is infected with a given disease or condition, determines the likelihood that a subject with the disease or condition will respond to treatment, determines the prognosis of a subject with the disease or condition (or the possible progression or regression), and determines the effectiveness of treatment of a subject with the disease or condition. For example, the diagnosis can be used to detect the likelihood that a subject is infected with a neoplasm or the likelihood that such a subject will respond favorably to a compound (e.g., a medicament, e.g., a drug) or other treatment.
[0130] The term "marker", as used herein, refers to a substrate (e.g., a nucleic acid or a region of a nucleic acid) that can diagnose cancer, for example, by distinguishing cancer cells from normal cells based on its methylation state.
[0131] The term "isolated", when used in reference to a nucleic acid, as in "isolated oligonucleotide", refers to a nucleic acid sequence that has been identified and separated from at least one contaminating nucleic acid with which it is ordinarily associated in its natural source. An isolated nucleic acid exists in a form or context different from that in which it occurs in nature. In contrast, non-isolated nucleic acids such as DNA and RNA are found in their naturally occurring states. Examples of non-isolated nucleic acids include a given DNA sequence (e.g., a gene) that is found on a host cell chromosome in proximity to adjacent genes and an RNA sequence such as a specific mRNA sequence that encodes a particular protein and is found intracellularly as a mixture with many other mRNAs that encode a number of different proteins. However, an isolated nucleic acid that encodes a particular protein includes, for example, nucleic acids within a cell that normally express the protein where the nucleic acid is at a chromosomal location different from that of its natural cell or where, alternatively, nucleic acid sequences different from those normally found flank it. An isolated nucleic acid or oligonucleotide can exist in single-stranded or double-stranded form. When an isolated nucleic acid or oligonucleotide is utilized to express a protein or oligonucleotide, the oligonucleotide includes the minimal sense or coding strand (i.e., the oligonucleotide is single-stranded), but can also include both the sense and antisense strands (i.e., the oligonucleotide can be double-stranded). An isolated nucleic acid may be combined with other nucleic acids or molecules after it has been isolated from its natural or typical environment. For example, an isolated nucleic acid may be present in a host cell that has been placed, for example, in a heterologous expression context.
[0132] The term "purified" refers to a molecule that is either removed from its natural environment, isolated, or separated, whether it is a nucleic acid or an amino acid sequence. An "isolated nucleic acid sequence" may therefore be a purified nucleic acid sequence. A "substantially purified" molecule is at least 60% free, preferably at least 75% free, and more preferably at least 90% free from other components with which it is naturally associated. As used herein, the terms "purified" or "purifying" also refer to removing contaminants from a sample. By removing contaminating proteins, the proportion of the polypeptide or nucleic acid of interest in the sample is increased. In other examples, a recombinant polypeptide is expressed in a plant, bacterium, yeast, or mammalian host cell, the polypeptide is purified by removing host cell proteins, and the proportion of the recombinant polypeptide in the sample is thereby increased.
[0133] The term "composition comprising" a given polynucleotide sequence or polypeptide broadly refers to any composition containing the given polynucleotide sequence or polypeptide. The composition may include an aqueous solution containing salts (e.g., NaCl), surfactants (e.g., SDS), and other components (e.g., Denhardt's solution, dried milk, salmon sperm DNA, etc.).
[0134] The term "sample" is used in its broadest sense. In one sense, it may refer to animal cells or tissues. In another sense, it means including not only biological and environmental samples, but also specimens or cultures obtained from any source. Biological samples are obtained from plants or animals (including humans) and may include fluids, individuals, tissues, and gases. In some embodiments, the sample is a plasma sample. In some embodiments, the sample is a gastric tissue sample. In some embodiments, the sample is a fecal sample. Environmental samples include environmental substances such as surface materials, soil, water, and industrial samples. These examples should not be construed as limiting the types of samples applicable to the present invention.
[0135] As used herein, a "distant sample" refers to a sample that is indirectly recovered from a site that is not the source of the cells, tissue, or organ of the sample, when used in some contexts. For example, when a sample substance originating from the pancreas is evaluated in a fecal sample (e.g., not derived from a sample directly taken from the pancreas), the sample is a distant sample.
[0136] As used herein, the terms "patient" or "subject" refer to an organism that undergoes various tests provided by the technology. The term "subject" includes animals, preferably mammalian animals including humans. In a preferred embodiment, the subject is a primate animal. In an even more preferred embodiment, the subject is a human.
[0137] As used herein, the term "kit" refers to any delivery system for delivering materials. In the context of a reaction assay, such a delivery system includes a system that enables the storage, transport, or delivery of reaction reagents (e.g., oligonucleotides, enzymes, etc. in suitable containers) and / or support materials (e.g., buffers, instructions for performing the assay, etc.) from one location to another. For example, a kit includes one or more containers (e.g., boxes) containing the relevant reaction reagents and / or support materials. As used herein, the term "subdivided kit" refers to a delivery system that includes two or more separate containers, each containing a portion of the total kit components. The containers can be delivered together or separately to the intended recipient. For example, the first container may contain an enzyme for use in the assay, while the second container may contain oligonucleotides. The term "subdivided kit" is intended to include, but is not limited to, kits containing analyte-specific reagents (ASR) that are regulated by section 520(e) of the Federal Food, Drug, and Cosmetic Act. In practice, any delivery system that includes two or more separate containers, each containing a portion of the total kit components, is included in the term "subdivided kit". In contrast, a "composite kit" refers to a delivery system that includes all components of a reaction assay in a single container (e.g., within a single container that houses each of the desired components). The term "kit" includes both subdivided and composite kits.
[0138] 〔Embodiments of the present technology〕 Provided herein are technologies related to the detection of tumor formation, and in particular, but not limited to, methods, compositions, and related uses for detecting pre-cancerous conditions and malignant tumors such as gastric cancer. The methylation status of DNA markers derived from tumors of subjects with gastric cancer (e.g., stomach cancer) was identified by a case-control method by comparing it with the methylation status of the same DNA markers from control subjects (see Examples 1 and 2). In additional experiments, 10 optimal markers for gastric cancer detection (ARHGEF4, ELMO1, ABCB1, CLEC11A, ST8SIA1, SFMBT2, CD1D, CYP26C1, ZNF569, C13ORF18; see Example 2 and Table 4) were identified. By further additional experiments, 30 optional markers for gastric cancer detection were identified (see Example 1 and Table 2). By further additional experiments, 12 optimal markers for the detection of gastric cancer (e.g., stomach cancer) in plasma samples were identified (ELMO1, ZNF569, C13 or f18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4 and ST8SIAI; see Example 3, Tables 1, 2, 4 and 7).
[0139] Markers and / or marker panels (e.g., having an annotation selected from ARHGEF4, ELMO1, ABCB1, CLEC11A, ST8SIAI, SFMBT2, CD1D, CYP26C1, ZNF569 or 13ORF18, and chromosomal regions containing the markers (see Table 4)) were identified by a case-control method by comparing the methylation status of gastric tissue of subjects with gastric cancer (e.g., stomach cancer) with the methylation status of the same DNA markers from control subjects (see Example 2).
[0140] A marker and / or a panel of markers (e.g., having an annotation selected from ELMO1, ARHGEF4, EMX1, SP9, CLEC11A, ST8SIA1, BMP3, KCNA3, DMRTA2, LCNK12, SD1D, PRLCB, CYP26C1, ZNF568, ABCB1, ELOVL2, PKIA, SFMBT2(839), PCBP3, MATK, GRN2D, NDRG4, DLX4, PPP2R5C, FGF14, ZNF132, CHST2(7890), FL11, c13 or f18 or ZNF569, and a chromosomal region containing the marker (see Table 2)) were identified by a case-control method by comparing the methylation state of the marker in gastric tissue of a subject having gastric cancer (e.g., stomach cancer) with the methylation state of the same DNA marker from a control subject (see Example 1).
[0141] A marker and / or a panel of markers (e.g., having an annotation selected from ELMO1, ZNF569, C13 or f18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4, ST8SIA1, and a chromosomal region containing the marker (see Tables 1, 2, 4 and 7)) were identified by a case-control method by comparing the methylation state of the DNA marker (derived from plasma of the stomach of a subject having gastric cancer (e.g., stomach cancer)) with the methylation state of the same DNA marker from a control subject (see Example 3).
[0142] Furthermore, the present technology relates to a panel of various markers. For example, in some embodiments, the markers have the annotation of ARHGEF4, ELMO1, ABCB1, CLEC11A, ST8SIA1, SFMBT2, CD1D, CYP26C1, ZNF569 or C13ORF18, and include a chromosomal region containing the markers (see Table 4). Further, in an embodiment, it shows a method for analyzing the DMRs in Table 4, where the DMR numbers are 49, 43, 196, 66, 1, 237, 249, 250, 251 or 252. In some embodiments, it provides a measurement of the methylation state of the markers, where the chromosomal region has the annotation of ELMO1, ARHGEF4, EMX1, SP9, CLEC11A, ST8SIA1, BMP3, KCNA3, DMRTA2, KCNK12, CD1D, PRKCB, CYP26C1, ZNF568, ABCB1, ELOVL2, PKIA, SFMBT2(893), PCBP3, MATK, GRN2D, NDRG4, DLX4, PPP2R5C, FGF14, ZNF132, CHST2(7890), FLI1, c13 or f18 or ZNF569 and includes the markers (see Table 2). In a further embodiment, the specific method for DMR analysis from Table 2 provides a method for analyzing the DMRs in Table 2 where the DMR numbers are 253, 251, 254, 255, 256, 249, 257, 258, 259, 260, 261, 262, 250, 263, 1, 264, 265, 196, 266, 118, 267, 268, 269, 270, 271, 272, 46, 273, 252 or 237. In some embodiments, the obtained sample is a plasma sample, and the markers have the annotation of ELMO1, ZNF569, C13 or f18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4 and ST8SIA1, and include a chromosomal region containing the markers.Furthermore, in such embodiments, the sample is a plasma sample, and the method analyzes the DMRs of Tables 1, 2, 4, and 7 that are DMR number 253, 251, 254, 255, 256, 249, 257, 258, 259, 260, 261, 262, 250, 263, 1, 264, 265, 196, 266, 118, 267, 268, 269, 270, 271, 272, 46, 273, 252, or 237. In some embodiments, the method includes measuring the methylation status of two markers (e.g., a set of markers shown in the columns of Tables 1, 2, 4, or 7).
[0143] The disclosure herein refers to specific exemplary embodiments, but it will be understood that these embodiments are shown by way of illustration and not by way of limitation.
[0144] In certain aspects, the technology provides compositions and methods for the identification, detection, and / or classification of cancers such as gastric cancer (e.g., stomach cancer). The method includes detecting the methylation status of at least one methylated marker in a biological sample isolated from a subject (e.g., a fecal sample, a gastric tissue sample, a plasma sample), wherein a change in the methylation status of the marker is indicative of the presence, classification, or location of gastric cancer (e.g., stomach cancer). Specific embodiments relate to markers comprising methylation variable regions (DMRs, e.g., DMR1 - 274, Tables 1, 2, 4, and 7) that are used in the diagnosis (e.g., screening) of neoplastic cell proliferative diseases (e.g., cancer), including early detection during the pre - cancerous stage of the disease.
[0145] In addition to embodiments in which the methylation of at least one marker, a region of a marker, or the bases of a marker comprising a DMR (e.g., DMR, DMR1 - 274) shown herein and listed in Tables 1, 2, 4, or 7 is analyzed, the technology also provides a panel of markers comprising at least one marker, a region of a marker, or the bases of a marker that includes DMRs having utility in the detection of cancer, particularly gastric cancer.
[0146] Some embodiments of the present technology are based on the analysis of the CpG methylation status of at least one marker, a region of a marker, or a base of a marker.
[0147] In some embodiments, the present technology is provided for the use of the bisulfite method in combination with one or more methylation assays that measure the methylation status of CpG dinucleotide sequences in at least one marker that includes a DMR (see, e.g., DMR1-274, Tables 1, 2, 4, and 7). CpG dinucleotides in the genome can be methylated or demethylated (or are known as up- and down-methylated, respectively). However, the methods of the present invention are suitable for the analysis of biological materials of heterogeneous nature (such as low concentrations of tumor cells or biological materials derived therefrom) in the background of a remote sample (such as blood, tissue excretions, or feces).
[0148] Thus, when analyzing the methylation status of the position of CpG in such a sample, a quantitative assay can be used to detect the level of methylation (such as percentage, fraction, ratio, population, degree) at a specific CpG position.
[0149] According to the present technology, the measurement of the methylation status of CpG dinucleotide sequences in markers that include DMRs is useful for both the diagnosis and characterization of cancers such as gastric cancer.
[0150] 〔Combination of Markers〕 In some embodiments, the present technology relates to evaluating the methylation status of a combination of markers including Table 1 (e.g., DMR numbers 1-248), or Table 4 (e.g., DMR numbers 49, 43, 196, 66, 1, 237, 249, 250, 251, 252), or Table 2 (e.g., DMR numbers 253, 251, 254, 255, 256, 249, 257, 258, 259, 260, 261, 262, 250, 263, 1, 264, 265, 196, 266, 118, 267, 268, 269, 270, 271, 272, 46, 273, 252, 237), or Table 7 (DMR number 274) or additional markers including DMRs. In some embodiments, evaluating the methylation status of one or more markers improves specificity and / or the sensitivity of screening or diagnosis for the prediction of tumors (e.g., gastric cancer) in a subject. In some embodiments, the marker or combination of markers differentiates between tumor types and / or localizations.
[0151] Various cancers are predicted by various combinations of markers (e.g., identified by statistical techniques regarding prediction specificity and sensitivity). The present technology provides a method for determining predictive and validated combinations for several cancers.
[0152] Method for analyzing methylation state The most commonly used methods for analyzing nucleic acids for the presence of 5-methylcytosine are based on the bisulfite method described by Frommer et al. (Frommer et al. (1992) Proc. Natl. Acad. Sci. USA 89: 1827-31, which is hereby incorporated by reference in its entirety for all purposes) for the detection of 5-methylcytosine in DNA or variants thereof. The bisulfite method for finding 5-methylcytosine is based on the observation that cytosine that is not 5-methylcytosine reacts with hydrogen sulfite ions (also known as bisulfite). The reaction is typically carried out by the following steps. First, cytosine reacts with hydrogen sulfite to form sulfonated cytosine. Next, spontaneous deamination of the sulfonated reaction intermediate results in sulfonated uracil. Finally, sulfonated uracil is desulfonated under alkaline conditions to form uracil. Since uracil forms a base pair with adenine (behaves like thymine), while 5-methylcytosine forms a base pair with guanine (behaves like cytosine), detection is possible. This allows for the discrimination of methylated cytosine from unmethylated cytosine, for example, by bisulfite gene sequencing (Grigg G, & Clark S, Bioessays (1994) 16: 431-36; Grigg G, DNA Seq. (1996) 6: 189-98), or by methylation-specific PCR (MSP) as disclosed, for example, in U.S. Patent No. 5,786,146.
[0153] Some of the prior art relates to methods that involve surrounding DNA to be analyzed within an agarose membrane, thereby preventing diffusion and reannealing of the DNA (bisulfite reacts only with single-stranded DNA) and replacing the precipitation and purification steps with rapid dialysis (Olek A, et al. (1996) “A modified and improved method for bisulfite based cytosine methylation analysis” Nucleic Acids Res. 24: 5064-6). Thus, it is possible to analyze individual cells for their methylation status, which indicates the usefulness and sensitivity of the method. An overview of the prior methods for detecting 5-methylcytosine is given by Rein, T., et al. (1998) Nucleic Acids Res. 26: 2255.
[0154] Bisulfite technology typically involves amplifying short specific fragments of a known nucleic acid after bisulfite treatment in order to analyze the position of individual cytosines, and then analyzing the products by sequencing (Olek & Walter (1997) Nat. Genet. include primer extension reactions (Gonzalgo & Jones (1997) Nucleic Acids Res. 25: 2529-31; International Publication No. WO 95 / 00669; U.S. Patent No. 6,251,594). Some methods use enzymatic digestion (Xiong & Laird (1997) Nucleic Acids Res. 25: 2532-4). Detection by hybridization is also described in the art (Olek et al., International Publication No. WO 99 / 28498). Additionally, the use of bisulfite technology for methylation detection of individual genes is described (Grigg & Clark (1994) Bioessays 16: 431-6,; Zeschnigk et al. (1997) Hum Mol Genet. 6: 387-95; Feil et al. (1994) Nucleic Acids Res. 22: 695; Martin et al. (1995) Gene 157: 261-4; International Publication No. WO 97 / 46705; International Publication No. WO 95 / 15373).
[0155] A variety of methylation analysis methods are known in the art and can be used in combination with bisulfite treatment according to the present technology. These analyses can determine the methylation state of one or more CpG dinucleotides (e.g., CpG islands) within a nucleic acid sequence. Such analyses include, among other techniques, sequencing of bisulfite-treated nucleic acids, PCR (for sequence-specific amplification), Southern blot analysis, and the use of methylation-sensitive restriction enzymes.
[0156] For example, gene sequencing is simplified for the analysis of methylation patterns, and 5-methylcytosine is distinguished by using bisulfite treatment (Frommer et al. (1992) Proc. Natl. Acad. Sci. USA 89: 1827-1831). Further, restriction enzyme digestion of PCR products amplified from bisulfite-converted DNA is found to be useful in assessing the methylation state, as described, for example, by Sadri & Hornsby (1997) Nucl. Acids Res. 24: 5058-5059, or as carried out in a method known as COBRA (Combined Bisulfite Restriction Analysis) (Xiong & Laird (1997) Nucleic Acids Res. 25: 2532-2534).
[0157] COBRA™ analysis is a quantitative methylation analysis useful for determining DNA methylation levels at specific loci in small amounts of genomic DNA (Xiong & Laird, Nucleic Acids Res. 25:2532-2534, 1997). Briefly, restriction enzyme digestion is used to reveal methylation-dependent sequence differences in the PCR products of sodium bisulfite-treated DNA. Methylation-dependent sequence differences are first introduced into genomic DNA by standard bisulfite treatment by the method described by Frommer et al. (Proc. Natl. Acad. Sci. USA 89:1827-1831, 1992). PCR amplification of bisulfite-converted DNA is then carried out using specific primers for the CpG island of interest, followed by restriction endonuclease digestion, gel electrophoresis, and detection using a specific labeled hybridization probe. The methylation level in the original DNA sample is represented by the relative amounts of digested and undigested PCR products in a linear quantification method across a broad spectrum of DNA methylation levels. Furthermore, this technique can be reliably applied to DNA obtained from paraffin-embedded tissue samples that have been microdissected.
[0158] Typical reagents for COBRA (trademark) analysis (such as those that may be found in a typical COBRA (trademark)-based kit) may include, but are not limited to, PCR primers for specific loci (such as specific genes, markers, DMRs, gene regions, marker regions, bisulfite-treated DNA sequences, CpG islands, etc.); restriction enzymes and appropriate buffers, gene hybridization-forming oligonucleotides; control hybridization-forming oligonucleotides; a kinase labeling kit for oligonucleotide probes; and labeled nucleotides. Further, bisulfite conversion reagents may include a DNA denaturation buffer, a sulfonation buffer, a DNA repair reagent, or a kit (such as precipitation, ultrafiltration, affinity column); a desulfonation buffer, and a DNA repair composition.
[0159] Preferably, assays such as "MethyLight" (trademark) (fluorescence-based real-time PCR technology) (Eads et al., Cancer Res. 59:2302-2306, 1999), Ms-SNuPE (trademark) (methylation-sensitive single nucleotide primer extension) reaction (Gonzalgo & Jones, Nucleic Acids Res. 25:2529-2531, 1997), methylation-specific PCR ("MSP"; Herman et al., Proc. Natl. Acad. Sci. USA 93:9821-9826, 1996; U.S. Pat. No. 5,786,146), and methylated CpG island amplification ("MCA"; Toyota et al., Cancer Res. 59:2307-12, 1999) are used alone or in combination with one or more of these methods.
[0160] The "HeavyMethyl (trademark)" analysis and technology is a quantitative method for assessing methylation differences based on methylation-specific amplification of bisulfite-treated DNA. Methylation-specific selective amplification of a nucleic acid sample is enabled by CpG positions covered between amplification primers or by methylation-specific blocking probes (the "blockers") covered by such primers.
[0161] The term "HeavyMethyl (trademark) MethyLight (trademark)" assay refers to the HeavyMethyl (trademark) MethyLight (trademark) assay, which is a variant of the MethyLight (trademark) assay that combines methylation-specific blocking probes that cover CpG positions between amplification primers. The HeavyMethyl (trademark) assay may also be used in combination with methylation-specific amplification primers.
[0162] Typical reagents for HeavyMethyl (trademark) analysis (such as may be found in a typical MethyLight (trademark)-based kit) may include, but are not limited to, PCR primers for a particular locus (such as a particular gene, marker, DMR, region of a gene, region of a marker, bisulfite-treated DNA sequence, CpG island, or bisulfite-treated DNA sequence or CpG island, etc.); oligonucleotide blocking; optimized PCR buffer and deoxynucleotides; and Taq polymerase.
[0163] Without being affected by the use of methylation-sensitive restriction enzymes, the methylation status of virtually any base of CpG sites within CpG islands can be assayed by MSP (methylation-specific PCR) (Herman et al. Proc. Natl. Acad. Sci. USA 93:9821-9826, 1996; U.S. Pat. No. 5,786,146). Briefly, DNA is modified by sodium bisulfite to convert unmethylated cytosine, but not methylated cytosine, to uracil, and the product is successively amplified with specific primers for methylated DNA versus unmethylated DNA. MSP requires only small amounts of DNA, has a sensitivity of 0.1% for methylated alleles at a given CpG island locus, and can be performed on DNA extracted from paraffin-embedded samples. Typical reagents for MSP analysis (such as may be found in a typical MSP-based kit) include, but are not limited to, methylation and unmethylated PCR primers for a specific locus (such as a specific gene, marker, DMR, region of a gene, region of a marker, bisulfite-treated DNA sequence, CpG island, etc.); optimized PCR buffer and deoxynucleotides, and may also include specific probes.
[0164] The MethyLight™ assay is a high-throughput quantitative methylation assay that utilizes fluorescence-based real-time PCR (e.g., TaqMan®) that does not require additional manipulation after the PCR step (Eads et al., Cancer Res. 59:2302-2306, 1999). Briefly, the MethyLight™ process begins with a pooled genomic DNA sample that is converted, by standard methods, in a sodium bisulfite reaction, into a pooled sample with methylation-dependent sequence differences (the bisulfite process converts unmethylated cytosine residues to uracil). Fluorescence-based PCR is then performed in a "biased" reaction, for example, with PCR primers that partially overlap a known CpG dinucleotide. Sequence discrimination occurs at both the level of the amplification step and the level of the fluorescence detection step.
[0165] The MethyLight™ assay is used as a quantitative test of methylation patterns in nucleic acids (e.g., genomic DNA samples), and sequence discrimination occurs at the level of probe hybridization. In the quantitative version, the PCR reaction gives methylation-specific amplification in the presence of a fluorescent probe that partially overlaps a specific putative methylation site. A bias-free control of the DNA input is provided by a reaction where neither the primers nor the probe overlap any CpG dinucleotide. Also, the quantitative test of genomic methylation is accomplished by exploring a biased PCR pool with either control nucleotides that do not cover known methylation sites (e.g., the fluorescence-based versions of the HeavyMethyl™ and MSP technologies) or oligonucleotides that cover possible methylation sites.
[0166] The MethyLight™ procedure can be used with any preferred probe (e.g., “TaqMan®” probe, Lightcycler® probe, etc.). For example, in some applications, double-stranded genomic DNA is treated with sodium bisulfite and is the target of one of two sets of probes for a PCR reaction using TaqMan® in combination with, for example, MSP primers and / or HeavyMethyl blocker oligonucleotides and TaqMan® probes. TaqMan® probes are doubly labeled with fluorescent “reporter” and “quencher” molecules and are designed to be specific for relatively high GC content regions and thus melt at a temperature approximately 10° C. higher during the PCR cycle than forward or reverse primers. This allows the TaqMan® probe to remain fully hybridized during the PCR annealing / extension step. Taq polymerase can enzymatically synthesize new strands during PCR and eventually reach the finally annealed TaqMan® probe. The Taq polymerase 5′-3′ endonuclease activity may then replace the TaqMan® probe by digesting it, releasing a fluorescent reporter molecule for quantitative detection of a signal that does not immediately quench using a real-time fluorescence detection system.
[0167] Typical reagents for MethyLight™ analysis (e.g., as may be found in a typical MethyLight™-based kit) may include, but are not limited to, PCR primers for a particular locus (e.g., a particular gene, marker, DMR, region of a gene, region of a marker, bisulfite-treated DNA sequence, CpG island, etc.); TaqMan® or Lightcycer® probes; optimized PCR buffer and deoxynucleotides; and Taq polymerase.
[0168] The QM (trademark) (Quantitative Methylation) assay is an alternative quantitative test of methylation patterns in genomic DNA samples, where sequence discrimination occurs at the level of probe hybridization. In this quantitative version, the PCR reaction gives unbiased amplification in the presence of a fluorescent probe that partially overlaps a specific putative methylation site. An unbiased control for DNA input is provided by a reaction where neither the primers nor the probe overlap any CpG dinucleotide. Also, the quantitative test of genomic methylation is accomplished by exploring a biased PCR pool either with a control oligonucleotide that does not cover known methylation sites (the fluorescent-based version of HeavyMethyl (trademark) and MSP technology) or an oligonucleotide that covers possible methylation sites.
[0169] In the amplification step of the QM (trademark) process, any suitable probe (e.g., "TaqMan (registered trademark)" probe, Lightcycer (registered trademark) probe) can be used. For example, double-stranded genomic DNA is treated with sodium bisulfite and exposed to unbiased primers and TaqMan (registered trademark) probes. The TaqMan (registered trademark) probe can remain fully hybridized during the PCR annealing / extension step. During PCR, as Taq polymerase enzymatically synthesizes new strands, it can eventually approach the finally annealed TaqMan (registered trademark) probe. The Taq polymerase 5'-3' endonuclease activity may then replace it by digesting it, releasing a fluorescent reporter molecule for quantitatively detecting a signal that does not immediately quench using a real-time fluorescence detection system. Typical reagents for QM (trademark) analysis (e.g., as may be found in a typical QM (trademark)-based kit) may include, but are not limited to, PCR primers for specific loci (e.g., specific genes, markers, DMRs, regions of genes, regions of markers, bisulfite-treated DNA sequences, CpG islands, etc.); TaqMan (registered trademark) or Lightcycer (registered trademark) probes; optimized PCR buffers and deoxynucleotides; and Taq polymerase.
[0170] The Ms-SNuPE™ technology is a quantitative method for assaying methylation differences at specific CpG sites based on bisulfite treatment of DNA followed by single-stranded nucleotide primer extension (Gonzalgo & Jones, Nucleic Acids Res. 25:2529-2531, 1997). Briefly, genomic DNA reacts with sodium bisulfite to convert unmethylated cytosine to uracil while leaving 5-methylcytosine unchanged. Amplification of the desired target sequence is then performed using PCR specific for the bisulfite-converted DNA, and the resulting product is isolated and used as a template for methylation analysis at the CpG site of interest. It can be analyzed with small amounts of DNA (e.g., microdissected pathological sections), avoiding the use of restriction enzymes to determine the methylation status at CpG sites.
[0171] Typical reagents for Ms-SNuPE™ analysis (e.g., as found in typical Ms-SNuPE™-based kits) may include, but are not limited to, PCR primers for specific loci (e.g., specific genes, markers, DMRs, regions of genes, regions of markers, bisulfite-treated DNA sequences, CpG islands, etc.); optimized PCR buffers and deoxynucleotides; gel extraction kits; positive control primers; Ms-SNuPE™ primers for specific loci; reaction buffer (Ms-SNuPE reaction); and labeled nucleotides. Additionally, bisulfite conversion reagents may include DNA denaturation buffer; sulfonation buffer; DNA repair reagents or kits (e.g., precipitation, ultrafiltration, affinity column); desulfonation buffer, and DNA repair compositions.
[0172] Reduced representation bisulfite sequencing (RRBS) starts with bisulfite treatment of nucleic acids to convert all unmethylated cytosines to uracil, followed by restriction enzyme digestion (e.g., by an enzyme that recognizes sites containing the CG sequence such as MspI), and after binding to adapter ligands, sequencing of the fragments is completed. The choice of restriction enzyme reduces the number of repetitive sequences that can be mapped to the positions of multiple genes during analysis and enriches for fragments of CpG-dense regions. Thus, RRBS reduces the complexity of the nucleic acid sample by selecting a portion of the restriction fragments for sequencing (e.g., by size selection using preparative gel electrophoresis). In contrast to whole-genome bisulfite sequencing, all fragments produced by restriction enzyme digestion contain DNA methylation information for at least one CpG dinucleotide. Thus, RRBS provides an assay for assessing the methylation state of one or more genomic loci by enriching for samples of promoters, CpG islands, and other genomic features with a high frequency of restriction enzyme cleavage sites in these regions.
[0173] A typical RRMS protocol includes steps of digesting a nucleic acid sample with a restriction enzyme such as MspI, steps to fill in overhangs and A-tailing, steps to ligate adapters, a bisulfite conversion step, and a PCR step. See, for example, (2005) “Genome-scale DNA methylation mapping of clinical samples at single-nucleotide resolution” Nat Methods 7: 133-6; Meissner et al. (2005) “Reduced representation bisulfite sequencing for comparative high-resolution DNA methylation analysis” Nucleic Acids Res. 33: 5868-77, etc.
[0174] In some embodiments, the Quantitative Allele-Specific Real-Time Target and Signal amplification (QuARTS) assay is used to evaluate methylation state. In each QuARTS assay, three reactions occur sequentially, including amplification (Reaction 1) and target probe cleavage (Reaction 2) in the first reaction; FRET cleavage and fluorescence signal generation (Reaction 3) in the second reaction. When the target nucleic acid is amplified with specific primers, a specific detection probe with a flap sequence binds loosely to the amplification product. The presence of a specific invading oligonucleotide at the target binding site causes cleavage to release the flap sequence by cleaving between the detection probe and the flap sequence. The flap sequence is complementary to the non-hairpin portion corresponding to the FRET cassette. Thus, the flap sequence functions as an invading oligonucleotide on the FRET cassette, resulting in cleavage between the FRET cassette fluorescent dye molecule and the quencher and generating a fluorescence signal. The cleavage reaction can cleave multiple probes for each target, and thus release multiple fluorescent dye molecules for each flap, providing a rapidly increasing signal amplification. QuARTS can detect multiple targets in a single reaction by using FRET cassettes with well-differentiated dyes. See, for example, Zou et al. (2010) “Sensitive quantification of methylated markers with a novel methylation specific technology” Clin Chem 56: A199; U.S. Patent Applications Nos. 12 / 946,737, 12 / 946,745, 12 / 946,752, and 61 / 548,639.
[0175] The term "bisulfite reagent" refers to a reagent containing bisulfite, disulfite, hydrogen sulfite, or a combination thereof, and as disclosed herein, is useful for distinguishing between methylated and non-methylated CpG dinucleotide sequences. Methods for such treatment are known in the art (e.g., PCT / EP2004 / 011715, incorporated herein by reference). The bisulfite treatment is preferably carried out in the presence of a denaturing solvent such as, but not limited to, an n-alkylene glycol or diethylene glycol dimethyl ether (DME), or in the presence of dioxane or a dioxane derivative. In some embodiments, the denaturing solvent is used at a concentration of 1% to 35% (v / v). In some embodiments, the bisulfite reaction is carried out in the presence of a scavenger such as, but not limited to, a chroman derivative (e.g., 6-hydroxy-2,5,7,8,-tetramethylchroman-2-carboxylate or trihydroxybenzoic acid and their derivatives, e.g., gallic acid) (see: PCT / EP2004 / 011715, incorporated herein by reference). The bisulfite conversion is preferably carried out at a temperature of 30°C to 70°C, and during the reaction, the temperature is raised above 80°C for a short time (see: PCT / EP2004 / 011715, incorporated herein by reference). The bisulfite-treated DNA is preferably purified before quantification. This may be done by any means known in the art such as, but not limited to, ultrafiltration (e.g., by a Microcon™ column (manufactured by Millipore™)). Purification has been carried out by a modified manufacturing protocol (see, e.g., PCT / EP2004 / 011715, incorporated herein by reference).
[0176] In some embodiments, the processed DNA fragments are amplified using the primer oligonucleotides according to the invention (see, for example, Tables 3 and / or 5) and a set of amplification enzymes. Amplification of several DNA segments can be carried out continuously in one and the same reaction vessel. Typically, amplification is carried out using the polymerase chain reaction (PCR). The amplification products typically have a length of 100 to 2000 base pairs.
[0177] In other embodiments of the method, the methylation status of CpG positions within or near a marker containing a DMR (e.g., DMR1-274 as given in Tables 1, 2, 4, and 7) can be detected by use of methylation-specific primer oligonucleotides. This technique (MSP) is described in Herman, U.S. Patent No. 6,265,171. Use of methylation-status-specific primers for amplification of bisulfite-treated DNA allows discrimination between methylated and unmethylated nucleic acids. The MSP primer pair includes at least one primer that hybridizes to a bisulfite-treated CpG dinucleotide. Therefore, the sequence of the primer includes at least one CpG dinucleotide. MSP primers specific for unmethylated DNA contain a "T" at the position of the C in CpG.
[0178] The fragments obtained by amplification can be carried to a directly or indirectly detectable label. In some embodiments, the label is a fluorescent label, a radionuclide, or a detectable molecular fragment having a typical mass detectable by a mass spectrometer. The label is a mass label, and some embodiments provide that the labeled amplification product has a single positive or negative net charge, which can improve detectability in a mass spectrometer. Detection can be carried out, for example, by means using matrix-assisted laser desorption ionization (MALDI) or electrospray ionization (ESI) and may be visualized.
[0179] Methods for isolating DNA suitable for these assay techniques are known in the art. In particular, some embodiments include the isolation of nucleic acids as described in U.S. Patent Application No. 13 / 470,251, "Isolation of Nucleic Acids" (incorporated herein by reference).
[0180] 〔Method〕 In some embodiments of the present technology, a method is provided that includes the following steps: 1) Contacting a nucleic acid obtained from a subject (e.g., genomic DNA, e.g., isolated from a body fluid such as a fecal sample, gastric tissue, plasma sample, etc.) with at least one reagent or a series of reagents that distinguish between methylation and non-methylation in at least one marker containing a DMR (e.g., DMR1-274, e.g., as shown in Tables 1, 2, 4, and 7). 2) Detecting a neoplastic or proliferative disease (e.g., having a sensitivity greater than or equal to 80% and a specificity greater than or equal to 80%).
[0181] In some embodiments of the present technology, a method is provided that includes the following steps: 1) Contacting a nucleic acid obtained from a subject (e.g., genomic DNA, e.g., isolated from a body fluid such as a fecal sample or gastric tissue) with at least one reagent or a series of reagents that distinguish between methylated CpG dinucleotides and non-methylated CpG dinucleotides in at least one marker selected from the group consisting of ARHGEF4, ELMO1, ABCB1, CLEC11A, ST8SIA1, SFMBT2, CD1D, CYP26C1, ZNF569, and C13ORF18. 2) Detecting gastric cancer (e.g., having a sensitivity greater than or equal to 80% and a specificity greater than or equal to 80%).
[0182] In some embodiments of the present technology, a method is provided that includes the following steps: 1) Contacting nucleic acid obtained from a subject (e.g., genomic DNA, e.g., isolated from a plasma sample) with at least one reagent or a series of reagents that distinguish methylated CpG dinucleotides from unmethylated CpG dinucleotides in at least one marker selected from chromosomal regions having an annotation selected from the group consisting of ELMO1, ARHGEF4, EMX1, SP9, CLEC11A, ST8SIA1, BMP3, KCNA3, DMRTA2, KCNK12, CD1D, PRKCB, CYP26C1, ZNF568, ABCB1, ELOVL2, PKIA, SFMBT2(893), PCBP3, MATK, GRN2D, NDRG4, DLX4, PPP2R5C, FGF14, ZNF132, CHST2(7890), FLI1, c13 or f18, or ZNF569; 2) Detecting gastric cancer (e.g., having a sensitivity greater than or equal to 80% and a specificity greater than or equal to 80%).
[0183] In some embodiments of the present technology, a method is provided that includes the following steps: 1) Contacting nucleic acid obtained from a subject (e.g., genomic DNA, e.g., isolated from a body fluid such as a stool sample or gastric tissue) with at least one reagent or a series of reagents that distinguish methylated CpG dinucleotides from unmethylated CpG dinucleotides in at least one marker selected from chromosomal regions having an annotation selected from the group consisting of ELMO1, ZNF569, C13orf18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4, and ST8SIA1; 2) Detecting gastric cancer (e.g., having a sensitivity greater than or equal to 80% and a specificity greater than or equal to 80%).
[0184] Preferably, the sensitivity is about 70% to about 100%, or about 80% to about 90%, or about 80% to about 85%. Preferably, the specificity is 70% to about 100%, or about 80% to about 90%, or about 80% to about 85%.
[0185] Genomic DNA can be isolated by any method, including the use of commercially available kits. Briefly, the DNA of interest is encapsulated by the cell membrane, and the biological sample must be disrupted enzymatically, chemically, or mechanically and then lysed. The DNA solution is then treated to remove proteins and other contaminants, for example, by digestion with Proteinase K. Subsequently, the genomic DNA is eluted from the solution. This can be done by various methods, including salting out, organic solvent extraction, and binding of the DNA to a solid-phase support. The choice of method will be influenced by multiple factors, including time, cost, and the required quality of the DNA. All types of clinical samples, including those containing tumorigenic or pre-tumorigenic events, are suitable for use in this method (e.g., cell lines, histological slides, biopsies, paraffin-embedded tissues, body fluids, feces, gastric tissue, colonic fluid, urine, plasma, serum, whole blood, isolated blood cells, cells isolated from blood, and combinations thereof).
[0186] The above techniques are not limited to methods used for sample preparation and nucleic acid preparation for experimentation. For example, in some specific cases, direct gene capture, such as using the detailed US Patent Application 61 / 485386 and related methods, DNA is isolated from stool samples, blood, and plasma samples.
[0187] The genomic DNA sample is treated with at least one reagent or a series of reagents that distinguish between methylated CpG dinucleotides and unmethylated CpG dinucleotides among at least one marker, including DMRs (e.g., DMR1-274, e.g., based on Tables 1, 2, 4, and 7).
[0188] In some embodiments, the reagent converts the demethylated cytosine base at the 5' position to uracil, thymine, or another base not similar to cytosine by a hybridization reaction. However, in some embodiments, the reagent can be a methylation-sensitive restriction enzyme.
[0189] In some embodiments, the genomic DNA sample is treated by a method that converts the demethylated cytosine base by reaction to a base other than uracil, thymine, or cytosine, which is the demethylated cytosine base at the 5' position. In some embodiments, this treatment is performed according to bisulfite (hydrogen, sulfuric acid, bisulfate) alkaline hydrolysis.
[0190] The treated nucleic acid analyzes the methylation status of a target gene sequence (nucleotides from at least one gene, genomic sequence, or marker for comparing DMRs, for example, at least one DMR selected from DMR1-274, for example, based on Tables 1, 2, 4, or 7). The above methods of analysis are selected, for example, from those known among these techniques, which are included in the lists of QuARTS and MSP in the present invention.
[0191] Abnormal methylation, and more particularly, hypermethylation of markers for comparing DMRs (at least one DMR selected from DMR1-274, for example, based on Tables 1, 2, 4, or 7) is associated with gastric cancer and more particularly with the predicted tumor site.
[0192] The present invention is related to the analysis of several samples related to gastric cancer. For example, in some embodiments, the sample includes tissue or biological fluid derived from a patient. In some embodiments, the sample includes secretions. In some embodiments, the sample includes blood, serum, plasma, gastric secretions, pancreatic juice, gastrointestinal biopsy samples, cells microdissected from gastrointestinal biopsies, gastrointestinal cells discarded into the gastrointestinal lumen, or gastrointestinal cells recovered from feces. In some embodiments, the subject is human. These samples can be derived from the upper gastrointestinal tract, the lower gastrointestinal tract, or include cells, tissues, and / or secretions from both the upper and lower gastrointestinal tracts. The sample includes cells, secretions, or tissues from the liver, bile duct, pancreas, stomach, colon, rectum, esophagus, small intestine, appendix, duodenum, polyp, gallbladder, anus, and / or peritoneum. In some embodiments, the sample includes cytosol, ascites, urine, excreta, pancreatic juice, fluid obtained during endoscopy, blood, mucus, or saliva. In some embodiments, the sample is feces.
[0193] Such samples can be obtained by many any methods known in the art (e.g., obvious to those skilled in the art). For example, urine samples and fecal samples are readily available, and blood samples, ascites samples, plasma samples, or pancreatic juice samples can be obtained parenterally, for example, by using needles and syringes. Cell-free or substantially cell-free samples can be obtained by subjecting the sample to various techniques known in the art (including but not limited to centrifugation and filtration). It is generally preferred that non-invasive techniques be used to obtain the sample, but it is also preferred to obtain samples such as tissue homogenates, tissue pieces, and biopsy specimens.
[0194] In some embodiments, the technology relates to a method for treating a patient (e.g., a patient with gastric cancer (e.g., stomach cancer), a patient with early-stage gastric cancer, or a patient who may progress to gastric cancer), the method comprising determining the methylation status of a DMR as shown herein and administering a treatment to the patient based on the result of the determined methylation status. The treatment may be the administration of a pharmaceutical compound, a vaccine, the performance of a surgical procedure, the imaging diagnosis of a patient, the performance of other tests. Preferably, the use is in a clinical screening method, a prognostic evaluation method, a method of observing the result of treatment, a method for identifying patients who are likely to respond to a particular treatment, a method for imaging diagnosing a patient or subject, a method for screening or developing a drug.
[0195] In some embodiments of the present technology, a method for diagnosing gastric cancer (e.g., stomach cancer) is provided. The terms "diagnosing" and "diagnosis" as used herein refer to a method by which one of ordinary skill in the art can estimate and further determine whether a subject is in a given disease or condition or is likely to progress to a given disease or condition in the future. One of ordinary skill in the art often makes a diagnosis based on one or more indicators of diagnosis (e.g., biomarkers (e.g., DMRs as disclosed herein), the presence of a condition, the methylation status indicating severity or absence).
[0196] Along with diagnosis, clinical cancer prognosis, clinical cancer prediction relates to determining the pathogenicity and recurrence rate of cancer and planning the most effective treatment. When a more accurate diagnosis can be made or the risk of cancer progression can be evaluated, an appropriate treatment, and in some cases a less risky treatment, can be selected. Findings of cancer biomarkers (e.g., measurement of methylation status) are useful for separating subjects with a good prognosis and / or a low risk of cancer progression (who require no treatment or little treatment) from subjects who are likely to progress or recur with cancer (who benefit from more intensive treatment).
[0197] Such "making a diagnosis" or "diagnosis" in the present invention further includes determining the risk of cancer progression or determining the prognosis. It is for selecting an appropriate treatment (or whether it is an effective treatment), or monitoring the progress of the current treatment and possibly changing the treatment, based on the measurement of the diagnosis of the biomarker (e.g., DMR) reported in the present invention, to predict the clinical outcome (whether to perform medical treatment or not). Further, in some embodiments of the present disclosure, a long-term composite determination of biomarkers can facilitate diagnosis and / or prognostic diagnosis. Temporary changes in biomarkers could predict clinical outcomes and monitor the progression of gastric cancer or directly monitor the effect of an appropriate treatment on cancer. In some such embodiments, for example, a biological sample observes and predicts changes in the methylation state of one or more biomarkers (e.g., DMR) (monitored or perhaps one or more biomarkers) disclosed in the present invention during the progress of an effective treatment over a long period.
[0198] Furthermore, to more specifically disclose the patient's condition immediately, there is a method for determining whether to initiate or continue cancer prevention or treatment in a patient. Specifically, in some cases, the method includes a system of biological samples until the period from the patient ends; in each of the biological samples with respect to each other, at least one biomarker disclosed in the present invention has a measurement of changes in the methylation state. Some of the changes in the methylation state of the biomarker have been used to predict the risk of cancer progression, predict clinical outcomes, and determine whether to initiate or continue cancer prevention or treatment and whether the current treatment is an effective treatment for cancer. For example, it is first selected for the start of treatment and then selected again some time after the start of treatment. The methylation state can be quantified at different times or measured for samples with respect to each other obtained from quantitatively different characteristics. Changes in the methylation state of biomarker levels from different samples can show a correlation with the risk of gastric cancer, prognosis prediction, determination of effective treatment, or prediction of cancer in a patient.
[0199] Rather specifically, it is preferred that the methods and properties of the present invention are for the treatment and diagnosis of diseases at an early stage, for example, at the signs before the disease appears. Specifically, in some cases, the methods and properties of the present invention are for the treatment and diagnosis of diseases at the clinical stage.
[0200] More specifically, many determinations of the diagnosis and prognosis of one or more biomarkers have caused changes over time in the markers used to make the diagnosis and prognosis determinations. For example, diagnostic markers have made determinations at an initial stage and again at a second stage. Specifically, in some cases, an increase in the marker from the initial stage to the second stage can diagnose a specific type or severity of cancer, or a prognosis. Similarly, a decrease in the marker from the initial stage to the second stage can indicate a specific type or severity of cancer, or a prognosis. Furthermore, the degree of change in one or more markers is related to the severity of cancer and future adverse events. Those skilled in the above techniques will understand these, and reliable specific measurements can be compared to create similar biomarkers at multiple stages, and it is possible to measure a biomarker at one stage and a second biomarker at a second stage, and the comparison of these markers provides information useful for diagnosis.
[0201] As used in the present invention, the term "diagnostic determination" is related to a method by which a person with skilled technique can predict the course and outcome of a patient's condition. The "prediction of course" period is not related to the ability to predict the course and outcome of a condition with 100% accuracy. Or, it can be predicted that the course and outcome are caused more or less based on the methylation state of a biomarker (e.g., DMR). Instead, a person with the above-mentioned skilled technique will be able to understand that the "prediction of course" period is related to an increase in the probability that a certain course and outcome will occur; and that the course and outcome are more likely to occur in patients showing the condition as compared to these individuals not showing the condition. For example, in individuals not showing a condition (e.g., having a standard methylation state of one or more DMRs), the likelihood of an outcome (e.g., suffering from gastric cancer) may be very low.
[0202] Specifically, statistical analysis is related to an indicator of prediction of course in which the result has a tendency to progress. For example, specifically, a methylation state different from that of a normal reference sample derived from a patient without cancer indicates that patients with a methylation state at a level similar to that in the reference sample have a statistically significant level of determination compared to patients suffering from cancer. Additionally, a change in methylation state from a reference level (e.g., "normal") can be reflected in the prediction of the patient's course, and the degree of change in the methylation state is related to the severity of the opposite result. Statistical significance is usually determined by comparing two or more groups of people and further by determination of the difference in reliability or p-value. For example, Dowdy and Wearden, Statistics for Research, John Wiley & Sons, New York, 1983 is incorporated by reference in its entirety. Typical differences in reliability for the current patient's condition are 90%, 95%, 97.5%, 98%, 99%, 99.5%, 99.9%, 99.99%, and all typical p-values are 0.1, 0.05, 0.025, 0.02, 0.01, 0.005, 0.001, 0.0001.
[0203] Specifically, the degree of change in the methylation state for predicting the course or diagnosing a biomarker (e.g., DMR) disclosed in the present invention can be estimated, and the degree of methylation of the biomarker in a biological sample can be easily compared with the degree of the methylation state threshold. The preferred changes in the methylation state for the biomarkers shown in the present invention are about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 50%, about 75%, about 100%, about 150%. However, in other embodiments, a "nomogram" can be provided. This is because the methylation state serving as an indicator (biomarker or combination of biomarkers) for follow-up observation or diagnosis is directly related to the property of leading to a certain result. A person skilled in the art is proficient in and understands that such a nomogram used in relation to two numerical values is as uncertain as the marker concentration measurement, because the measurement of an individual sample, rather than the population average, is referred to.
[0204] Specifically, in some cases, a reference sample is analyzed simultaneously with the biological sample. Then, the result obtained from the biological sample is compared with the result obtained from the reference sample. Additionally, it is expected to form a standard distribution curve, and the results of the assay for the biological sample are compared. Such a standard distribution curve indicates the methylation state of the biomarker as a function of the assay unit. For example, when a fluorescent label is used, it is the intensity of the fluorescent signal. Samples provided and used from a number of donors define the standard methylation state of one or more biomarkers in normal tissue. It is similar to the "at-risk" level of one or more biomarkers in tissue provided from donors with dysplasia or gastric cancer donors. In one embodiment of the above method, a patient is determined to have dysplasia showing an abnormal methylation state of one or more DMRs shown in the present invention in a biological sample taken from the patient. Further, in another embodiment of the above method, as a result of detecting an abnormal methylation state of one or more such biomarkers in a biological sample derived from the patient, the patient is considered to have cancer.
[0205] The above analysis of markers is performed separately or simultaneously with additional markers within one test sample. For example, each marker performs one test simultaneously for efficient operation of the sample or for making a more accurate diagnosis or prognosis. Additionally, a certain skill in this technology recognizes the usefulness of testing multiple samples from the same patient (e.g., when time is continuous). Such testing of a series of samples is recognized as being identical to changes in methylation markers over time. Changes in the methylation state are useful information about the state of the disease including it, but not limited to this, and confirm approximately the same time from the start of the above state. It is the confirmation of the patient's outcome including its presence or the amount of recovered tissue, the suitability of drug treatment, the effectiveness of various treatments, and the risk of future states.
[0206] The above analysis of biomarkers is performed in various practical systems. For example, the use of microplates and full automation have facilitated the operation of numerous test samples. Alternatively, the type of one sample has evolved rapidly for immediate treatment and diagnosis. For example, it is temporary transportation and the installation of an emergency treatment room.
[0207] In some embodiments, when comparing the methylation status of the reference, the above patient is diagnosed as if having gastric cancer. This is because there is a certain difference in the methylation status of at least one biomarker in the above sample. Conversely, when there is no change in the methylation status, it is identified among biological samples. It is confirmed whether the patient has gastric cancer, has no cancer risk, or has a low cancer risk. In this regard, the patients with the above cancer or its risk can be generally distinguished from patients who do not have cancer or have a low risk. These patients at risk of gastric cancer progression are placed in a more intensive or normal screening program including endoscopic observation. On the other hand, patients with generally no or low risk can avoid endoscopic examination until the time of future screening, for example, screening consistent with current technology, indicating that the risk of gastric cancer appears in those patients.
[0208] As already mentioned, the specific development of the current technology method is that detecting changes in the methylation status of one or more biomarkers enables qualitative or quantitative measurement. Since the patient is at the stage of diagnosis of the risk of progression, gastric cancer (e.g., gastric cancer), the reliable threshold measurement indicates so. For example, it is the methylation status of one or more biomarkers in various biological samples from the pre-determined reference methylation status. In some embodiments of the method, the reference methylation status is several detectable methylation status biomarkers. In other embodiments of the method, in the case of a reference sample experimented with at the same time as the biological sample, the pre-determined reference methylation status is the methylation status in the reference sample. In other embodiments of the method, the pre-determined methylation status may be based later or is determined by a standard distribution curve. In other embodiments of the method, the pre-determined methylation is a clear state or a range of states. The above pre-determined methylation is selected based on a part of the embodiment of the method, such as the desired specificity that is more proficient within the clearly acceptable limitations for those skilled in the art.
[0209] To the extent related to diagnostic methods, the preferred patients are vertebrate patients. The preferred vertebrates are warm-blooded; the preferred warm-blooded vertebrates are mammals. The preferred mammals are mostly human. As used in the present invention, "patient" includes both human and animal patients. Thus, veterinary treatments are indicated by the present invention. Such state-of-the-art is for the diagnosis of mammals such as humans, but also for valuable mammals at risk of extinction, such as the Siberian tiger; animals on farms for human consumption that are of economic importance; or animals of social importance to humans, such as pets or animals kept in zoos. Examples of such animals include, but are not limited to: carnivorous mammals such as cats and dogs; domestic pigs, piglets, male pigs for meat, and also wild male pigs; ruminants or ungulates such as domestic cows, bulls, sheep, giraffes, deer, goats, bison, camels, and also horses. Thus, the above-mentioned diagnosis and treatment of domestic animals are also indicated, but are not limited to, domesticated pigs, ruminants, ungulates, horses (including racehorses), and animals comparable thereto. The immediately indicated patient condition further includes a system for the diagnosis of gastric cancer (e.g., stomach cancer) in a patient. The above system is defined, for example, as a mass-produced kit that has been used to screen for the risk of gastric cancer or the diagnosis of gastric cancer in a patient from whom a biological sample has been collected. The exemplary system is consistent with the present technique for measuring DMRs in a methylated state based on Tables 1, 2, 4, or 7.
Examples
[0210] (Example 1 - Identification of Markers for Detecting Gastric Cancer) In the process of developing embodiments for this technology, experiments were conducted that identified 123 DNA methylation markers for gastric cancer (corresponding to 248 DMRs (e.g., DMR numbers 1 - 248; Table 1)). Such DNA methylation markers were identified from data generated by enrichment of CpG islands with parallel large-scale sequencing of patient-control tissue sample sets. Controls included DNA obtained from normal gastric tissue, normal colonic epithelium, and normal white blood cells. The method utilized reduced representation bisulfite sequencing (RRBS). The third step involved analyzing data for coverage cutoffs, excluding all non-beneficial sites, comparing methylation percentages between diagnostic subgroups using logistic regression, creating computer-generated CpG "islands" based on clear grouping of adjacent methylation sites, and receiver operating characteristic analysis. The fourth step involved filtering methylation percentage data (both individual CpGs and clusters within the computer) to maximize the signal-to-noise ratio, minimize background methylation, reveal tumor heterogeneity, and prioritize ROC performance. Further, this analysis method ensured the identification of CpG hotspots across the entire determined DNA length for easy and optimal design of downstream marker assays (methylation-specific PCR, small fragment deep sequencing).
[0211] Each sample yielded approximately 2 to 3 million good-quality CpGs, which, after analysis and filtering, became less than 1000 individually distinct sites that were highly discriminative. These were concentrated in 123 localized regions of specific methylation (some spanning 30 - 40 bases and some kilobases) (see Table 1). All DMRs had an AUC of 0.8 or more (some showed complete discrimination of patients from controls (AUC of 1.0)). In addition to the 0.8 AUC threshold, the identified gastric cancer markers needed to show a methylation density in tumors that was more than 20 times that of normal gastric mucosa or colonic mucosa and less than 1.0% methylation in non-neoplastic gastric mucosa or colon.
[0212]
Table 1
[0213]
Table 2
[0214]
Table 3
[0215]
Table 4
[0216]
Table 5
[0217]
Table 6
[0218] Of the 123 markers disclosed, 22 were selected based on AUC and fold change rate (cancer vs. normal stomach + normal colon) and incorporated into a larger tissue validation study along with 73 exploratory markers from other GI cancer sites (colon, pancreas, bile duct, esophagus). Samples in these different cohorts were between 15 - 50 samples in cancer, normal, and, where applicable, adenoma. The validation platform was quantitative methylation-specific PCR using primers optimized by a population of methylation controls. All assays were performed on Roche 480 LightCyclers. Results were analyzed for ROC characteristics, fold change, and complementarity. The top 30, as evaluated by sensitivity (for cancer and adenoma), are shown in Table 2. Table 3 shows primer information for the markers shown in Table 2. Of these, 10 were selected for further tissue validation (see Example 2) in a joint study with Mayo Korean.
[0219]
Table 7
[0220]
Table 8
[0221]
Table 9
[0222] (Example 2 - Gastric Cancer Detection by Novel Methylated DNA Markers: Tissue Validation in Patient Cohorts from the United States and Korea) Experiments conducted during the development of embodiments for this technology identified novel methylated DNA by whole methylome sequencing and then validated optimal candidates in tissues from patient cohorts in the United States and Korea (see Tables 4, 5, and 6).
[0223] After attempts to find the overall methylome using reduced representation bisulfite sequencing, 17 DNA methylation marker candidates were selected for blind testing in gastric tissue collections from patients in the United States and South Korea. DNA from microdissected tissue was quantified by methylation-specific polymerase chain reaction, and marker levels were normalized by beta-actin (total human DNA). The GC cases consisted of 35 (United States) and 50 (South Korea) patients pathologically confirmed as untreated adenocarcinoma. Controls included pathologically normal epithelium obtained from sites distant from primary GC from 65 (United States) and 50 (South Korea) healthy patients demographically matched to the cases. For each marker, the area under the ROC curve (AUC) was calculated from nominal logistic regression.
[0224] Overall, the top 10 markers included ARHGEF4, ELMO1, ABCB1, CLEC11A, ST8SIA1, SFMBT2, CD1D, CYP26C1, ZNF569, and C13ORF18 (see Table 4). At 95% specificity, this panel of markers detected GC in 100% of the US cohort and 94% of the South Korean cohort. The individual AUCs were 0.92, 0.91, 0.89, 0.88, 0.85, 0.82, 0.82, 0.81, 0.75, and 0.66 in the US cohort and 0.73, 0.90, 0.78, 0.75, 0.80, 0.94, 0.87, 0.79, 0.75, and 0.79 in the South Korean cohort (see Table 6). Some markers, such as ELMO1, showed similarly high discriminatory ability in each patient cohort, while other markers, such as ARHGEF4, showed more variable discriminatory ability due to high background from normal controls (NL) across the patient cohorts (see Figure 1). Supplement 5 shows the forward and reverse primer information for the 10 markers shown in Supplement 4.
[0225]
Table 10
[0226]
Table 11
[0227]
Table 12
[0228] (Detection of Gastric Cancer by Evaluation of Novel Methylated DNA Markers in Plasma - Example 3) This example investigated the use of selected methylated DNA marker (MDM) candidates from Examples I and II for plasma - based evaluation as an approach to detect gastric adenocarcinoma (GC). Indeed, this example demonstrated that a panel of novel MDMs evaluated in plasma can accurately discriminate GC cases from controls. It was shown that marker levels were negligible in controls, elevated in cases, and increased gradually according to the GC stage.
[0229] Plasma samples from the repository that met the case inclusion criteria were included in the study from a single comprehensive facility. The cases included 37 patients with pathologically confirmed adenocarcinoma, covering all stages, histological types, and gastric sites. Plasma samples of the cases were collected before resection, chemotherapy, or radiation. As controls, plasma samples from the repository were selected from 38 healthy volunteers matched for age and gender. DNA was extracted from 2 ml of plasma and bisulfite treatment. Then, the selected MDMs (ELMO1, ZNF569, C13orf18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4, and ST8SIA1) (Table 7 shows DMR information for LRCC4) were quantified using a quantitative allele-specific real-time target and signal amplification method without optimization. These markers have demonstrated high performance in tissues, as evaluated by clinical sensitivity and specificity, fold change in methylation signal when comparing gastric cancer to normal gastric tissue, and very low background methylation in DNA derived from blood cells.
[0230] To facilitate the highest analytical performance in plasma, these markers (previously examined in tissues using methylation-specific PCR assays) were redesigned using the original DMR sequences while multiplexing with the QuARTs (quantitative allele-specific real-time target and signal amplification) assay. The QuARTs assay contains PCR primers (see Table 8), detection probes (see Table 8), and an invasive oligo (Integrated DNA Technologies), GoTaq DNA polymerase (Promega), Cleavase II (Hologic), as well as a fluorescence resonance energy transfer reporter cassette (FRET) (Biosearch Technologies) containing FAM, HEX, and Quasar 670 dyes.
[0231] Figure 2 shows the oligonucleotide sequences for the FRET cassette used for the detection of the characteristics of methylated DNA by the QuARTs (quantitative allele-specific real-time target and signal amplification) assay. Each FRET sequence contains fluorophores and quenchers that can both be multiplexed in three separate assays.
[0232] The standard was increased from a plasmid control and derived from the digested sequences. The absolute copy number of the serially diluted standard was obtained from Poisson modeling.
[0233] DNA was purified from 2 mL of archived plasma samples and bisulfite converted using in-house methods and automated instrumentation. A processing control gene was added upfront to all samples to correct for variability in target recovery during the test period.
[0234] The samples were pre-amplified using 12 primer sets, diluted, and subjected to QuARTs. The platform for the subsequent stage was LightCycler 480 (Roche). The strand counts of each marker were normalized against the strands from β-actin included in the multiplex reaction.
[0235] The results were analyzed logistically to evaluate complementarity and plotted in matrix format.
[0236] Figure 3 shows the receiver operating characteristic curves highlighting the performance of three marker panels (ELMO1, ZNF569, and c13orf18) in plasma compared to individual marker curves. At 100% specificity, the panel detected gastric cancer at 86% sensitivity.
[0237] Figure 4A shows the stage-dependent log-scale absolute strand counts of methylated ELMO1 in plasma, which explains the progression from normal mucosa to stage 4 gastric cancer. The quantitative marker levels increased with the GC stage.
[0238] Figure 4B shows a bar graph (100% specificity) demonstrating the sensitivity of gastric cancer in plasma according to the stage of three marker panels (ELMO1, ZNF569, and c13orf18).
[0239] Indeed, ELMO1 is the most discriminatory marker, with an AUC of 94% (95% It had a CI of 89 - 99%. At 100% specificity, a panel of three MDMs (ELMO1, ZNF569, and C13orf18) detected GC with a sensitivity of 86% (71 - 95%). According to the GC stage (Figure 4A), the sensitivity by the panel at 100% specificity was 50%, 92%, 100%, and 100% for stage 1 (n = 8), 2 (n = 13), 3 (n = 4), and 4 (n = 11), respectively (p = 0.01 for trend). The quantitative MDM levels increased directly according to the GC stage. As shown by the distribution of methylated ELOMO1 (Figure 4B), the median / ml of plasma in the plasma strand / control was 0.4, and the strand counts in stage 1, 2, 3, and 4 of CG cases were 16, 111, 101, and 213, respectively (p = 0.01 for trend). Neither age nor gender affected the marker levels.
[0240] Figures 5A, B, and C show the performance of 12 gastric cancer markers (ELMO1, ZNF569, C13orf18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4, and ST8SIA1; ELMO1 was performed twice, and the second time was in a multiplex format). Matrix format at 90% specificity (A), 95% specificity (B), and 100% specificity (C) in 74 plasma samples. The markers were arranged vertically and the samples horizontally. The samples were aligned with normal on the left and cancer on the right. Positive hits are bright gray and misses are dark gray. Here, the top - performing marker (ELMO1) was first described at 90% specificity, and the remaining panel was at 100% specificity. This plot enables the markers to be evaluated in a combinatorial fashion.
[0241] Figure 6 shows the performance of three gastric cancer markers (first panel) in 74 plasma samples with 100% specificity. The markers are arranged vertically and the samples horizontally. The cancers are aligned according to stage. Positive hits are in light gray and misses are in dark gray (ELMO1, ZNF569, and c13orf18).
[0242] Figures 7A - M show box plots (linear scale) of 12 gastric cancer markers (ELMO1, ZNF569, C13orf18, CD1D, ARHGEF4, SFMBT2, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4, and ST8SIA1; ELMO1 was performed twice, the second time in a biplex format) in plasma. Control samples (N = 38) are shown on the left and gastric cancer cases (N = 36) on the right. The vertical axis is the % methylation normalized to the β - actin strand.
[0243]
Table 13
[0244]
Table 14
[0245]
Table 15
[0246] The publications and patents described in the above specification are hereby incorporated by reference in their entirety for all purposes. Various modifications and variations of the described compositions, methods, and uses of the technology will be apparent to those skilled in the art without departing from the scope and spirit of the technology as described. Although the technology has been described in connection with exemplary specific embodiments, it is understood that the invention as claimed should not be unduly limited to such specific embodiments. Indeed, various modifications of the described modes for carrying out the invention that are obvious to those skilled in the pharmacological, biochemical, medical, or related fields are intended to be within the scope of the following claims.
[0247] [Others] [Aspect 1] A method for characterizing a biological sample, comprising: treating genomic DNA of the biological sample with a reagent that modifies DNA in a methylation-specific manner; amplifying the bisulfite-treated genomic DNA using primers specific for at least one methylation variable region (DMR) derived from at least one gene selected from SFMBT2, NDRG4, and FER1L4; and measuring the methylation profile of the at least one DMR using methylation-specific PCR, quantitative methylation-specific PCR, methylation-sensitive DNA restriction enzyme analysis, quantitative bisulfite pyrosequencing, and / or bisulfite genomic sequencing PCR. [Aspect 2] The method according to aspect 1, wherein the biological sample is derived from a human subject having or suspected of having a gastric neoplasm. [Aspect 3] The method according to aspect 1, wherein the biological sample is a tissue sample, fecal sample, urine sample, blood sample, plasma sample, or serum sample. [Aspect 4] The method according to aspect 3, wherein the tissue sample is a gastric tissue sample. [Aspect 5] The method according to embodiment 1, further comprising the step of extracting genomic DNA from the biological sample. [Embodiment 6] The at least one DMR is associated with the area under the ROC curve (AUC) that is 0.5 or more, The ROC curve is the method according to embodiment 1 for discriminating a human subject having or suspected of having a gastric neoplasm from a control DNA sample. [Embodiment 7] The method according to embodiment 1, wherein the at least one DMR has a higher methylation rate compared to a control DNA sample. [Embodiment 8] The method according to embodiment 1, wherein the at least one DMR has a higher frequency of hypermethylation compared to a control DNA sample. [Embodiment 9] The method according to any one of embodiments 6 to 8, wherein the control DNA sample is derived from a subject without a gastric neoplasm. [Embodiment 10] The method according to embodiment 1, comprising the step of amplifying a DMR derived from at least two genes selected from SFMBT2, NDRG4 and FER1L4. [Embodiment 11] The method according to embodiment 1, comprising the step of amplifying a DMR derived from each of SFMBT2, NDRG4 and FER1L4. [Embodiment 12] The method according to embodiment 1, comprising the step of amplifying a DMR derived from at least one additional gene selected from the group consisting of ELMO1, ZNF569, C13orf18, CD1D, ARHGEF4, PPP25RC, CYP26C1, PKIA, CLEC11A, LRRC4 and ST8SIA1. [Embodiment 13] The method according to embodiment 1, wherein the reagent for modifying DNA in the methylation-specific method comprises one or more of a methylation-sensitive restriction enzyme, a methylation-dependent restriction enzyme and a bisulfite reagent. [Embodiment 14] The method according to embodiment 1, wherein the DMR is present in a coding region or a control region. [Embodiment 15] (i) The primers specific to SFMBT2 include SEQ ID NO: 5 and 6, or SEQ ID NO: 83 and 84, and (ii) The method according to embodiment 1, wherein the primers specific to NDRG4 include SEQ ID NO: 57 and 58.
Claims
1. A method for assessing the methylation level of at least one gene in a sample from a subject having or suspected of having a gastric neoplasm, comprising: treating the DNA of said sample with a reagent that modifies DNA in a methylation-specific manner; amplifying the treated DNA using primers specific for a methylation variable region (DMR) from the at least one gene; and measuring the methylation level of a DMR from said at least one gene using methylation-specific PCR, quantitative methylation-specific PCR, methylation-sensitive DNA restriction enzyme analysis, quantitative bisulfite pyrosequencing and / or bisulfite genomic sequencing PCR; the at least one gene is selected from FAIM2, ZNF569, NDRG4, CLEC11A, ST8SIA1, CD1D, FER1L4, EMX1, OSR2, GRIN2D, and / or ARHGEF4; method.
2. The method of claim 1, wherein the sample is a tissue sample, a fecal sample, a urine sample, a blood sample, a plasma sample, and / or a serum sample.
3. The method described in claim 2, wherein the tissue sample is a gastric tissue sample.
4. The method described in claim 1, further comprising the step of extracting genomic DNA from the sample.
5. The method described in claim 1, further comprising a step of comparing the methylation level with the methylation level of a corresponding gene from a control sample not having a gastric neoplasm.
6. The method described in claim 5, wherein the methylation level is increased compared to the control sample.
7. The method described in claim 1, wherein the step of measuring the methylation level includes measuring a methylation score and / or measuring a methylation frequency.
8. The method described in claim 1, wherein the DMR is present in a coding region or a control region of the at least one gene.
9. The method described in claim 1, wherein the reagent that modifies DNA in the methylation-specific manner includes a methylation-sensitive restriction enzyme, a methylation-dependent restriction enzyme and / or a bisulfite reagent.
10. The methylation level of the DMR is associated with an area under the receiver operating characteristic curve (AUC) of 0.5 or greater; 2. The method of claim 1, wherein the ROC curve distinguishes between DNA from a sample derived from a human subject having or suspected of having a gastric neoplasia and a control DNA sample, the control DNA sample being derived from a human subject not having a gastric neoplasia.
11. The primers specific for FAIM2 comprise SEQ ID NOs: 5 and 6; Primers specific to ZNF569 include SEQ ID NOs: 11 and 12; Primers specific for NDRG4 include SEQ ID NOs: 57 and 58; Primers specific for CLEC11A include SEQ ID NOs: 29 and 30; Primers specific to ST8SIA1 include SEQ ID NOs: 13 and 14; Primers specific for CD1D include SEQ ID NOs: 39 and 40; Primers specific for EMX1 include SEQ ID NOs: 25 and 26; Primers specific for GRIN2D include SEQ ID NOs: 55 and 56; and Primers specific for ARHGEF4 include SEQ ID NOs: 23 and 24; The method of claim 1.