Compositions and methods for detecting esophageal cancer
By detecting novel differentially methylated regions and copy number variations in esophageal samples, the accuracy problem of diagnosing esophageal precancerous lesions in existing technologies has been solved, and accurate distinction between non-dysplasia and high-grade dysplasia has been achieved, supporting early cancer detection and treatment.
Patent Information
- Application Number
- CN202380087261.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-20
- Filing Date
- 2023-12-20
- Publication Date
- 2025-09-19
AI Technical Summary
Existing diagnostic methods for esophageal cancer and precancerous lesions are limited by esophageal sampling errors and subtle histological changes, leading to delayed treatment of dysplasia and cancer. In particular, it is difficult to accurately distinguish non-dysplasia from high-grade dysplasia and adenocarcinoma in Barrett's esophagus.
By detecting novel differentially methylated regions (DMRs) in esophageal samples, using methylation-specific PCR, quantitative methylation-specific PCR and other methods, combined with copy number variation and aneuploidy scores, we can distinguish non-dysplastic Barrett's esophagus from high-grade dysplasia or esophageal adenocarcinoma.
It improves the diagnostic sensitivity of esophageal cancer and precancerous lesions, provides an area under the ROC curve higher than 0.5 to 0.95, ensures accurate classification of esophageal samples, and supports early detection and treatment.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 433,784, filed December 20, 2022, which is incorporated herein by reference in its entirety for all purposes.
[0003] Electronically Submitted Materials Incorporated by Reference
[0004] Incorporated herein by reference in its entirety is the computer-readable nucleotide / amino acid sequence listing filed concurrently herewith and identified as follows: a 10,938-byte file named "41473-601_SEQUENCE_LISTING" created on December 14, 2023. Technical Field
[0005] The present disclosure provides compositions and methods for distinguishing non-cancerous, precancerous, and cancerous conditions in the esophagus. Specifically, the present disclosure provides compositions and methods for distinguishing non-dysplastic Barrett's esophagus (NDBE) and / or normal esophageal samples from precancerous (e.g., low-grade or high-grade dysplasia) and / or cancerous (e.g., esophageal adenocarcinoma) samples based on methylation status and / or DNA copy number aberrations. Background Art
[0006] Barrett's esophagus (BE) is a condition in which tissue similar to the intestinal lining changes or replaces the lining of the esophagus. Although BE can potentially progress to low-grade dysplasia (LGD) or high-grade dysplasia (HGD), precancerous conditions that can transform into esophageal adenocarcinoma (EAC), in most cases, dysplasia may not be present or identified on biopsy specimens. This condition is referred to as non-dysplastic BE (NDBE). Endoscopic surveillance is generally recommended to identify and differentiate between dysplasia, precancerous lesions, and cancerous conditions. However, this diagnostic paradigm is limited by esophageal sampling errors and subtle histologic changes, leading to delayed treatment of dysplasia and EAC. Summary of the Invention
[0007] Embodiments of the present disclosure provide methods, compositions and systems for screening various types of esophageal cancer in biological samples. According to these embodiments, the present disclosure includes but is not limited to methods and compositions for detecting the presence of esophageal cancer or precancerous lesions from biological samples. In some embodiments, the biological sample is a tissue sample, a blood sample, a plasma sample, a serum sample, a whole blood sample, a buffy coat sample, a secretion sample, an organ secretion sample, a cerebrospinal fluid (CSF) sample, a saliva sample, a urine sample and / or a fecal sample. In some embodiments, the tissue sample is an esophageal tissue sample. In some embodiments, the esophageal sample is obtained from an esophageal biopsy, or obtained by wiping, brushing or using a sponge capsule device.
[0008] As further described herein, embodiments of the present disclosure include novel differentially methylated regions (DMRs), each DMR individually capable of distinguishing esophageal cancer or precancerous lesions from controls or benign tissue. In some embodiments, the novel DMRs are capable of distinguishing high-grade dysplastic Barrett's esophagus (HGD-BE) or esophageal adenocarcinoma (EAC) from non-dysplastic Barrett's esophagus (NDBE) or normal esophageal controls. According to these embodiments, the novel DMR is from a gene selected from the group consisting of ACVRL1, ADAMTS8, ADAP2, ADRBK2, AKR1B1, ANK1, ANKRD13B, ANXA6, ARNT2, B4GALNT2, BACH2, BCL11A, BSCL2, C12orf53, C14orf82, C17orf107, C18orf1, C1orf95, C5orf42, CACNA1C, CAMK1D, CAMTA1, CBX6, CCDC85A, CCKBR, CD38, CD KN2A, CH25H, CHST1, CHST15, CNTLN, CRHR1, CRTC1, CXCR4, CYP1B1, DCTN2, DIDO1, DMKN, DSE, DYNC1I1, EML6, ENOX1, EPHA4, ESRRG, FAM176A, FAM78B, FBXO10, FERMT2, FHOD3, FLJ45079, FMNL1, FOXP2, FRMD4B, FZD8, GALNTL1, GALNTL4, GAS1, GBGT1, GLIPR2, GNAI1 , GNAL, GPR37, GRASP, GRID1, GRM8, GSC, HAR1A, HCN2, HEY2, HIST1H2BE, HS3ST3B1, HTR7, IGFBP2, INHA, INSRR, IRX3, ISM2, KCNG3, KCNK4, KCNS2, KCTD15, KIAA1522, KIAA1614, KIF26A, KL, KLF15, KLHL10, KRT77, KSR2, LBH, LMX1A, LMX1B, LOC100526820, LONRF2, LRFN2, LRRN1, MAF, MAFB, MARK1, ADAMTSL4-AS1, PGBD5, HSPA12A, SFTPD, LOC107984507, LOC100128253, CISTR, SLC16A7, CTXND1, MAX.chr15.4912, ZNF423, RBFOX1, LOC105376772, GSE1, MAX.chr17.8070, ZNF709, MAX.chr19.3439, GREB1, CYP27C1, AC068134.6. LOC388780, STOX2, PPARGC1A, MAX.chr4.4552, PDZD2, MAX.chr6.6743, LPAL2, CDK14, MAX.chr8.3003, MCOLN2, MEGF11, MFSD11, MRC2, MSX1, NAT8L , NAV1, NBEA, NCRNA00092, NECAB2, NEURL, NGEF, NHLH2, NKD1, NLGN1, NOG, NPPC, NR3C1, NRXN2, NTN1, OBSCN, OCA2, OSBPL1A, OXR1, P2RY1, PDE2A, PDE9A , PDGFRA, PID1, PLCG2, PLCL1, PNPLA3, POU3F1, POU3F2, PRKACB, PRKAR2B, PRKG1, PRR18, PRR5L, PTHLH, PYGL, RARG, RHBDL3, RIMS2, ROR2, RPRML, SDK2, SOX9, SYNGR1, TAC4, TNFRSF19, TRANK1, TSPAN33, TSPAN4, TSPAN5, UBE2E2, UCHL1, UNC5A, VASH2, VIM, WNT6, ZBTB10, ZNF680, ZNF738, ZNF808, and ZNF845 (Table 1), including any combination thereof. In some embodiments, the novel DMR is from any gene selected from Table 1, including any combination thereof. Each novel DMR alone is capable of distinguishing HGD-BE and / or EAC from NDBE and / or control samples, and combining two or more novel DMRs may provide improved sensitivity. Therefore, combinations of two or more novel DMRs selected from Table 1 are provided.
[0009] Embodiments of the present disclosure also include novel differentially methylated regions (DMRs), each DMR individually capable of distinguishing esophageal cancer or precancerous lesions from controls or benign tissue. In some embodiments, the novel DMRs are capable of distinguishing high-grade dysplastic Barrett's esophagus (HGD-BE) or esophageal adenocarcinoma (EAC) from non-dysplastic Barrett's esophagus (NDBE) or normal esophageal controls. According to these embodiments, the novel DMRs are from genes selected from the group consisting of ACVRL1, ADAMTS8, ADAP2, ADRBK2, AKR1B1, ANK1, ANKRD13B, ANXA6, ARNT2, B4GALNT2, BACH2, BCL11A, BSCL2, C14orf82, C18orf1, C1orf95, C5orf42, CAMK1D, CAMTA1, CCDC85A, CD38, C DKN2A, CHST1, CHST15, CRHR1, CYP1B1, DIDO1, DSE, DYNC1I1, EML6, ENOX1, EPHA4, FAM176A, FBXO10, FERMT 2. FHOD3, FLJ45079, FMNL1, FRMD4B, FZD8, GALNTL1, GALNTL4, GAS1, GBGT1, GLIPR2, GNAI1, GNAL, GRASP, GR ID1, GRM8, GSC, HAR1A, HCN2, HEY2, HIST1H2BE, HS3ST3B1, HTR7, IGFBP2, INHA, IRX3, ISM2, KCNG3, KCNK4, KCNS2, KCTD15, KIAA1614, KL, KLHL10, KSR2, LBH, LMX1A, LMX1B, LOC100526820, LONRF2, LRFN2, MAFB, MAR K1, PGBD5, HSPA12A, SFTPD, LOC107984507, LOC100128253, CISTR, SLC16A7, CTXND1, ZNF423, LOC1053767 72. GSE1, MAX.chr19.3439, GREB1, CYP27C1, AC068134.6, LOC388780, STOX2, PPARGC1A, PDZD2, MAX.chr6.6743, LPAL2, CDK14, MCOLN2, MEGF11, MRC2, MSX1, NAT8L, NAV1, NBEA, NCRNA00092, NEURL, NGEF, NHLH2, NKD1, NL GN1, NOG, NPPC, NR3C1, NRXN2, OBSCN, OCA2, OSBPL1A, OXR1, P2RY1, PDE2A, PDGFRA, PID1, PLCG2, PLCL1, PNPLA3, POU3F1, POU3F2, PRKACB, PRKAR2B, PRKG1, PRR18, PRR5L, PTHLH, PYGL, RHBDL3, RIMS2, ROR2, RPRML, SDK2, SOX9, SYNGR1, TAC4, TRANK1, TSPAN4, UBE2E2, UCHL1, UNC5A, VASH2, ZBTB10, ZNF680, ZNF709, ZNF738, ZNF808, and ZNF845 (Table 2), including any combination thereof. In some embodiments, the novel DMR is from any gene selected from Table 2, including any combination thereof. Each novel DMR alone is capable of distinguishing HGD-BE and / or EAC from NDBE and / or control samples, and combining two or more novel DMRs can provide improved sensitivity. Thus, combinations of two or more novel DMRs selected from Table 2 are provided.
[0010] Embodiments of the present disclosure also include novel differentially methylated regions (DMRs), each DMR individually capable of distinguishing esophageal cancer or precancerous lesions from controls or benign tissue. In some embodiments, the novel DMRs are capable of distinguishing high-grade dysplastic Barrett's esophagus (HGD-BE) or esophageal adenocarcinoma (EAC) from non-dysplastic Barrett's esophagus (NDBE) or normal esophageal controls. According to these embodiments, the novel DMRs are from genes selected from the group consisting of BACH2, C5orf42, FHOD3, HIST1H2BE, IRX3, KIAA1614, LONRF2, MAFB, PDGFRA, PID1, POU3F1, PRR5L, RHBDL3, and SDK2 (Table 3), including any combination thereof. In some embodiments, the novel DMRs are from any gene selected from Table 3, including any combination thereof. Each novel DMR individually is capable of distinguishing HGD-BE and / or EAC from NDBE and / or control samples, and combining two or more novel DMRs can provide improved sensitivity. Therefore, a combination of two or more novel DMRs selected from Table 3 is provided.
[0011] Embodiments of the present disclosure also include novel differentially methylated regions (DMRs), each DMR individually capable of distinguishing esophageal cancer or precancerous lesions from controls or benign tissue. In some embodiments, the novel DMRs are capable of distinguishing high-grade dysplastic Barrett's esophagus (HGD-BE) or esophageal adenocarcinoma (EAC) from non-dysplastic Barrett's esophagus (NDBE) or normal esophageal controls. According to these embodiments, the novel DMRs are from genes selected from the group consisting of KL, PGBD5, ROR2, and LMX1B (Example 4), including any combination thereof. In some embodiments, the novel DMRs are from any gene selected from Example 4, including any combination thereof. Each novel DMR individually is capable of distinguishing HGD-BE and / or EAC from NDBE and / or control samples, and combining two or more novel DMRs can provide increased sensitivity. Thus, a combination of two or more novel DMRs selected from Example 4 is provided.
[0012] In some embodiments, at least one of methylation-specific PCR, quantitative methylation-specific PCR, methylation-specific DNA restriction enzyme analysis, quantitative bisulfite pyrosequencing, flap endonuclease assay, PCR-flap assay, and bisulfite genomic sequencing PCR is used to validate a novel DMR capable of distinguishing esophageal cancer or precancerous lesions from controls or benign tissues based on at least one of the area under the ROC curve (AUC), methylation fold change, methylation percentage, and / or hypermethylation ratio between the test sample and the control sample.
[0013] According to the above, the control sample includes a sample from a subject who does not suffer from cancer, a sample from a subject who does not suffer from esophageal cancer, a sample from a subject who does not suffer from esophageal precancerous lesions, or a sample from a subject with a cancer type of non-esophageal cancer or non-precancerous lesions. In some embodiments, the control sample includes a sample from a subject with non-dysplastic Barrett's esophagus (NDBE). In some embodiments, the control sample is from a tissue sample, a blood sample, a plasma sample, a serum sample, a whole blood sample, a buffy coat sample, a secretion sample, an organ secretion sample, a cerebrospinal fluid (CSF) sample, a saliva sample, a urine sample, and a stool sample. In some embodiments, the tissue sample is an esophageal tissue sample. In some embodiments, the esophageal sample is obtained from an esophageal biopsy, or obtained by wiping, brushing, or using a sponge capsule device.
[0014] In some embodiments, a novel DMR that can distinguish between esophageal cancer or a precancer lesion and a control sample is associated with an area under the ROC curve (AUC) greater than or equal to 0.5, wherein the ROC curve distinguishes subjects having or suspected of having esophageal cancer or a precancer lesion from a control DNA sample. In some embodiments, a novel DMR that can distinguish between esophageal cancer or a precancer lesion and a control sample is associated with an area under the ROC curve (AUC) greater than or equal to 0.6, wherein the ROC curve distinguishes subjects having or suspected of having esophageal cancer or a precancer lesion from a control DNA sample. In some embodiments, a novel DMR that can distinguish between esophageal cancer or a precancer lesion and a control sample is associated with an area under the ROC curve (AUC) greater than or equal to 0.7, wherein the ROC curve distinguishes subjects having or suspected of having esophageal cancer or a precancer lesion from a control DNA sample. In some embodiments, a novel DMR that can distinguish between esophageal cancer or a precancer lesion and a control sample is associated with an area under the ROC curve (AUC) greater than or equal to 0.75, wherein the ROC curve distinguishes subjects having or suspected of having esophageal cancer or a precancer lesion from a control DNA sample. In some embodiments, a novel DMR that can distinguish between esophageal cancer or a precancer lesion and a control sample is associated with an area under the ROC curve (AUC) greater than or equal to 0.8, wherein the ROC curve distinguishes subjects having or suspected of having esophageal cancer or a precancer lesion from a control DNA sample. In some embodiments, a novel DMR that can distinguish between esophageal cancer or a precancer lesion and a control sample is associated with an area under the ROC curve (AUC) greater than or equal to 0.85, wherein the ROC curve distinguishes subjects having or suspected of having esophageal cancer or a precancer lesion from a control DNA sample. In some embodiments, a novel DMR capable of distinguishing esophageal cancer or a precancerous lesion from a control sample is associated with an area under the ROC curve (AUC) greater than or equal to 0.9, wherein the ROC curve distinguishes subjects having or suspected of having esophageal cancer or a precancerous lesion from a control DNA sample. In some embodiments, a novel DMR capable of distinguishing esophageal cancer or a precancerous lesion from a control sample is associated with an area under the ROC curve (AUC) greater than or equal to 0.95, wherein the ROC curve distinguishes subjects having or suspected of having esophageal cancer or a precancerous lesion from a control DNA sample.
[0015] In some embodiments, the novel DMR capable of distinguishing esophageal cancer or precancerous lesions from control samples comprises an increased percentage of methylation compared to a control DNA sample. In some embodiments, the novel DMR capable of distinguishing esophageal cancer or precancerous lesions from control samples comprises an increased rate of hypermethylation compared to a control DNA sample.
[0016] Embodiments of the present disclosure also include methods and compositions for characterizing a biological sample and determining a methylation profile in at least one differentially methylated region (DMR) of a DNA sample obtained from a subject having or suspected of having esophageal cancer or a precancerous lesion, by treating the sample with an agent that modifies DNA in a methylation-specific manner. In some embodiments, the method includes detecting the presence of esophageal cancer or a precancerous lesion from a biological sample. In some embodiments, at least one DMR is capable of distinguishing high-grade dysplastic Barrett's esophagus (HGD-BE) or esophageal adenocarcinoma (EAC) from non-dysplastic Barrett's esophagus (NDBE) or a normal esophageal control.
[0017] According to these embodiments, the method includes determining copy number variations (CNVs) in a DNA sample from a subject. In some embodiments, the CNVs can distinguish subjects having or suspected of having high-grade dysplastic Barrett's esophagus or esophageal adenocarcinoma (EAC) from control DNA samples. In some embodiments, at least one DMR comprises an increased CNV compared to the control DNA sample.
[0018] In some embodiments, the method includes determining an aneuploidy score (AS) for a DNA sample from a subject. In some embodiments, the AS can distinguish subjects with or suspected of having high-grade dysplasia, Barrett's esophagus, or esophageal adenocarcinoma (EAC) from control DNA samples. In some embodiments, at least one DMR comprises an increased AS compared to a control DNA sample.
[0019] As further described herein, the present disclosure provides materials and methods for copy number aberration (CNA) determination using sequencing reads from genomes modified in a methylation-specific manner (e.g., genomes converted to cytosine or 5-methylcytosine), including polyploidy and aneuploidy (e.g., determining an aneuploidy score). Most currently available technologies use direct NGS for wild-type, unconverted DNA. However, embodiments of the present disclosure include the ability to perform methylation and CNV analysis from the same chemistry / dataset. According to these embodiments, the same DNA sample obtained from a subject can be used to determine the methylation profile in at least one DMR and CNV and / or AS. In some embodiments, a single DNA sample obtained from a subject can be used to determine the methylation profile in at least one DMR and CNV and / or AS. In some embodiments, the sample (e.g., the same sample or a single sample) has been treated with a reagent that modifies DNA in a methylation-specific manner.
[0020] In some embodiments, the control DNA sample used in the methods and compositions of the present disclosure is from a subject who does not suffer from esophageal cancer or precancerous lesions. In some embodiments, the control DNA sample is from a subject who has non-dysplastic Barrett's esophagus (NDBE). In some embodiments, the control DNA sample is selected from a tissue sample, a blood sample, a plasma sample, a serum sample, a whole blood sample, a buffy coat sample, a secretion sample, an organ secretion sample, a cerebrospinal fluid (CSF) sample, a saliva sample, a urine sample, and a stool sample. In some embodiments, the tissue sample is an esophageal tissue sample or a gastric cardia sample.
[0021] In some embodiments, the biological sample is obtained from a human subject. In some embodiments, the biological sample is obtained from a human subject, and the method includes extracting a DNA sample from the biological sample. In some embodiments, the biological sample is collected using a collection device. For example, the biological sample can be an esophageal sample obtained from an esophageal biopsy, or an esophageal sample obtained by wiping, brushing, or using a sponge capsule device.
[0022] In some embodiments, the methods of the present disclosure include the use of a reagent that modifies DNA in a methylation-specific manner. In some embodiments, the reagent is a borane reducing agent. In some embodiments, the reagent that modifies DNA in a methylation-specific manner comprises one or more of a methylation-sensitive restriction enzyme, a methylation-dependent restriction enzyme, and a bisulfite reagent.
[0023] In some embodiments, determining the methylation profile of at least one DMR comprises amplifying at least a portion of the DMR using a set of primers. In some embodiments, determining the methylation profile of at least one DMR comprises performing at least one of methylation-specific PCR, quantitative methylation-specific PCR, methylation-specific DNA restriction enzyme analysis, quantitative bisulfite pyrosequencing, flap endonuclease assay, PCR-flap assay, and bisulfite genomic sequencing PCR. In some embodiments, determining the methylation profile of at least one DMR comprises determining the presence or absence of methylation at CpG sites. In some embodiments, one or more CpG sites are present in the coding region, non-coding region, and / or regulatory region of a gene (e.g., any one of the genes disclosed herein).
[0024] Embodiments of the present disclosure also include a method for identifying esophageal cancer or precancerous lesions. According to these embodiments, the method includes determining the methylation profile in at least one differentially methylated region (DMR) of the sample obtained from a subject suffering from or suspected of suffering from esophageal cancer or precancerous lesions by treating a reagent that modifies DNA in a methylation-specific manner. In some embodiments, the methylation profile indicates that the individual suffers from esophageal cancer (e.g., esophageal adenocarcinoma (EAC)) or esophageal precancerous lesions (e.g., high-grade dysplasia Barrett's esophagus (HGD-BE)). In some embodiments, the method includes determining the copy number variation (CNV) of the DNA sample from the subject. In some embodiments, the method includes determining the aneuploidy score (AS) of the DNA sample from the subject. In some embodiments, the method also includes treating the subject with anticancer therapy. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 : Representative bar graph of total CNV results for chromosome arm gain and loss levels. (The higher the signal, the greater the degree of chromosome instability and aneuploidy).
[0026] Figure 2 : Representative violin plots of total CNV results for chromosome arm gain and loss levels. (The higher the signal, the greater the degree of chromosome instability and aneuploidy).
[0027] Figure 3 : Representative chromosome scatter plots illustrating normal euploidy (two chromosomes / gene copies) of the normal esophagus (NE) cohort generated from whole-genome sequencing data. Each dot represents a binned genomic segment.
[0028] Figure 4 : Representative chromosome scatter plots illustrating normal euploidy (two chromosomes / gene copies) for the non-dysplastic Barrett's esophagus (NDBE) cohort generated from whole-genome sequencing data. Each dot represents a binned genomic segment.
[0029] Figure 5 Representative chromosome scatter plots illustrating aneuploidy in the high-grade dysplasia (HGD) cohort. These were generated from whole-genome NGS data. Each dot represents a binned genomic segment.
[0030] Figure 6 Representative chromosome scatter plots illustrating aneuploidy in the esophageal adenocarcinoma (EAC) cohort. These were generated from whole-genome NGS data. Each dot represents a binned genomic segment.
[0031] Figure 7: Representative heatmap of DMRs for the MAFB gene (a top methylation marker candidate identified in this disclosure) for high-grade dysplasia (HGD) and esophageal adenocarcinoma (EAC) cohorts. Columns represent samples, and rows represent individual CpGs in genomic order that constitute DMRs. Dark green indicates no methylation; increasing red hues indicate increased methylation intensity.
[0032] Figure 8 : Representative heatmap of DMRs for the MAFB gene (a top methylation marker candidate identified in this disclosure) for normal esophagus (NE) and non-dysplastic Barrett's esophagus (NDBE) cohorts. Columns represent samples, and rows represent individual CpGs in genomic order that constitute DMRs. Dark green indicates no methylation; increasing red hues indicate increased methylation intensity.
[0033] Figure 9 : Representative heatmaps illustrating the complementarity between CNV (aneuploidy) and methylation analysis in esophageal adenocarcinoma (EAC), high-grade dysplasia (HGD), and non-dysplastic Barrett's esophagus (NDBE) cohorts. DETAILED DESCRIPTION
[0034] Barrett's esophagus (BE) is the greatest risk factor and the only known precursor to esophageal adenocarcinoma (EAC), a lethal malignancy with a poor survival rate (less than 20% at 5 years) if detected after symptom onset. The incidence of esophageal adenocarcinoma has increased nearly 600% in the population over the past three decades. BE progresses to EAC through a stepwise pathway from no dysplasia (also known as non-dysplastic Barrett's esophagus or NDBE) to low-grade dysplasia (LGD) to high-grade dysplasia (HGD) to cancer. This progression from metaplasia to dysplasia to carcinoma has led several national gastroenterological associations to recommend BE screening for high-risk individuals with multiple risk factors, followed by endoscopic surveillance (depending on the extent of dysplasia) to detect the development of dysplasia or carcinoma at an early stage. Endoscopic treatments for LGD, HGD, and early-stage cancer have been developed and have proven effective in reducing cancer incidence and improving survival in patients with BE.
[0035] Currently, screening for BE is performed using conventional sedated endoscopy (sEGD), which shows that the normal squamous lining of the esophagus is replaced by metaplastic columnar epithelium in patients with BE. However, sedated endoscopy is expensive in both direct and indirect costs and is not suitable for widespread use. It is also associated with potential complications. Other technologies, such as unsedated transnasal endoscopy (uTNE), have comparable accuracy to sEGD at a lower cost but are still not considered a widely available tool by healthcare providers. Despite the availability of uTNE devices, its utilization by referring physicians remains limited. The lack of accurate risk stratification tools to determine BE risk and target screening efforts is an additional limitation to the widespread use of BE screening.
[0036] Currently, in addition to careful examination of the BE segment using high-resolution white-light imaging and advanced imaging techniques, endoscopic detection of dysplasia is performed using random four-quadrant biopsies of the BE segment every 1–2 cm. Despite recommendations from gastroenterological societies, adherence to these recommendations among practicing gastroenterologists remains low. In fact, adherence decreases with increasing BE segment length, leading to an increased incidence of missed dysplasia. Other challenges in detecting dysplasia in BE include the speckled distribution of dysplasia in BE, which can lead to sampling error, poor interobserver agreement among pathologists in grading dysplasia, and the relatively poor sensitivity of current surveillance strategies in detecting prevalent dysplasia or cancer. The utility of advanced imaging techniques in the community remains uncertain, with only one-third of practicing gastroenterologists reporting routine use in BE surveillance. Recently, a sponge-on-a-string device has been studied for BE screening. This device consists of a polyurethane foam sponge enclosed in a gelatin capsule attached to a string. The capsule is swallowed by the patient. The gelatin shell of the capsule dissolves in gastric fluid, releasing a spherical foam device that is then pulled out using an attached string, providing a brushing / cytology sample of the proximal stomach and esophagus. These samples can then be studied for biomarkers to detect BE.
[0037] As further described herein, BE is a metaplastic change in the epithelial lining of the distal esophagus characterized by replacement of normal squamous epithelium (NE) with specialized intestinal metaplasia. The presence of Barrett's esophagus increases the risk of developing esophageal adenocarcinoma several-fold. The present disclosure relates to the detection of high-grade dysplasia (HGD) Barrett's esophagus, including adenocarcinoma (EAC), which is distinct from non-dysplastic Barrett's esophagus (NDBE). Understanding the dysplasia status of Barrett's patients is critical for subsequent clinical monitoring and treatment strategies. As further described herein, RRBS analysis was performed on tissue biopsies and whole genome sequencing was performed on endoscopic brushing samples to generate methylation and copy number variation profiles of BE patients. Using a proprietary analysis algorithm, 199 methylated DNA markers (MDMs) were identified, 156 of which were confirmed in two NGS datasets. Additionally, aneuploidy scores were developed using WGS data, and these scores demonstrate how this marker class complements MDM analysis. For aneuploidy / CNV analysis, calls were made using genomes converted from methylation analysis, a method not currently available. Therefore, this technology could be implemented in a clinical test format for endoscopic brushings and non-endoscopic, non-invasive esophageal sponge samples.
[0038] The section headings as used in this section and the entire disclosure herein are for organizational purposes only and are not intended to be limiting.
[0039] 1. Definition
[0040] Throughout the specification and claims, unless the context clearly dictates otherwise, the following terms have the meanings clearly associated herein. As used herein, the phrase "in one embodiment" does not necessarily refer to the same embodiment, although it may refer to the same embodiment. In addition, as used herein, the phrase "in another embodiment" does not necessarily refer to a different embodiment, although it may refer to a different embodiment. Therefore, as described below, various embodiments of the present invention can be easily combined without departing from the scope or spirit of the present invention.
[0041] In addition, as used herein, the term "or" is an inclusive "or" operator and is equivalent to the term "and / or" unless the context clearly dictates otherwise. The term "based on" is not exclusive and allows for being based on other factors not described unless the context clearly dictates otherwise. In addition, throughout the specification, the meanings of "a," "an," and "the" include plural meanings. The meaning of "in..." includes "in..." and "on..."
[0042] The transition phrase "consisting essentially of" as used in the claims of this application limits the scope of the claim to the specified materials or steps "and those that do not materially affect" the "basic and novel characteristics" of the claimed invention, as discussed in In re Herz, 537 F.2d 549, 551-52, 190 USPQ 461, 463 (CCPA 1976). For example, a composition that "consists essentially of the recited elements" may contain unrecited contaminants at levels such that the contaminants, while present, do not alter the function of the recited composition as compared to the pure composition (i.e., a composition "consisting of the recited components").
[0043] As used herein, the term "one or more" refers to a number greater than one. For example, the term "one or more" encompasses any of two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, twelve or more, thirteen or more, fourteen or more, fifteen or more, twenty or more, fifty or more, one hundred or more, or even more.
[0044] The terms "one or more but less than a higher number," "two or more but less than a higher number," "three or more but less than a higher number," "four or more but less than a higher number," "five or more but less than a higher number," "six or more but less than a higher number," "seven or more but less than a higher number," "eight or more but less than a higher number," "nine or more but less than a higher number," "ten or more but less than a higher number," "eleven or more but less than a higher number," "twelve or more but less than a higher number," "thirteen or more but less than a higher number," "fourteen or more but less than a higher number," or "fifteen or more but less than a higher number" are not limited to higher numbers. For example, the higher number may be 10,000, 1,000, 100, 50, etc. For example, the higher number can be about 50 (e.g., 50, 49, 48, 47, 46, 45, 44, 43, 42, 41, 40, 39, 38, 37, 36, 35, 34, 33, 32, 31, 32, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or 2).
[0045] The terms "one or more methylation marks" or "one or more DMRs" or "one or more genes" or "one or more markers" or "multiple methylation marks" or "multiple markers" or "multiple genes" or "multiple DMRs" are likewise not limited to a specific numerical combination. In fact, any numerical combination of methylation marks is contemplated (e.g., 1-2 methylation marks, 1-3, 1-4, 1-5, 1-6, 1-7, 1-8, 1-9, 1-10, 1-11, 1-12, 1-13, 1-14, 1-15, 1-16, 1-17, 1-18, 1-19, 1-20, 1-21, 1-22, 1-23, 1-24, 1-25, 1-26, 1-27, 1-28, 1-29, 1-30, 1-31, 1-32, 1-33, 1-34, 1-35, 1-36, 1-37, 1-38) (e.g., 2-3, 2-4, 2-5, 2-6, 2-7, 2-8, 2-9, 2-10, 2-11, 2-12, 2-13, 2-14, 2-15, 2-16, 2-17, 2-18, 2-19, 2-20, 2-21, 2-22, 2-23, 2-24, 2-25, 2-26, 2-27, 2-28, 2-29, 2-30, 2-31, 2-32, 2-33, 2-34, 2-35, 2-36, 2-37, 2-38) (For example, 3-4, 3-5, 3-6, 3-7, 3-8, 3-9, 3-10, 3-11, 3-12, 3-13, 3-14, 3-15, 3-16, 3-17, 3-18, 3-19, 3-20, 3-21, 3-22, 3-23, 3-24, 3-25, 3-26, 3-27, 3-28, 3-29, 3-30, 3-31, 3-32, 3-33, 3-34, 3-35, 3-36, 3-37, 3-38) (For example, 4-5, 4-6, 4-7, 4-8, 4-9, 4-10, 4-11, 4-12, 4-13, 4-14, 4-15, 4-16, 4-17, 4-18, 4-19, 4-20, 4-21, 4-22, 4-23, 4-24, 4-25, 4-26, 4-27, 4-28, 4-29, 4-30, 4-31, 4-32, 4-33, 4-34, 4-35, 4-36, 4-37, 4-38) (e.g., 5-6, 5-7, 5-8, 5-9, 5-10, 5-11, 5-12, 5-13, 5-14, 5-15, 5-16, 5-17, 5-18, 5-19, 5-20, 5-21, 5-22, 5-23, 5-24, 5-25, 5-26, 5-27, 5-28, 5-29, 5-30, 5-31, 5-32, 5-33, 5-34, 5-35, 5-36, 5-37, 5-38) (e.g.,6-7, 6-8, 6-9, 6-10, 6-11, 6-12, 6-13, 6-14, 6-15, 6-16, 6-17, 6-18, 6-19, 6-20, 6-21, 6-22, 6-23, 6-24, 6-25, 6-26, 6-27, 6-28, 6-29, 6-30, 6-31, 6-32, 6-33, 6-34, 6-35, 6-36, 6-37, 6-38) (e.g., 7-8, 7-9, 7-10, 7-11, 7-12, 7-13, 7-14, 7-15, 7-16, 7-17, 7-18, 7-19, 7-20, 7-21, 7-22, 7-23, 7-24, 7-25, 7-26, 7-27, 7-28, 7-29, 7-30, 7-31, 7-32, 7-33, 7-34, 7-35, 7-36, 7-37, 7-38) (e.g., 8-9, 8-10, 8-11, 8-12, 8-13, 8-14, 8-15, 8-16, 8-17, 8-18, 8-19, 8-20, 8-21, 8-22, 8-23, 8-24, 8-25, 8-26, 8-27, 8-28, 8-29, 8-30, 8-31, 8-32, 8-33, 8-34, 8-35, 8-36, 8-37, 8-38) (For example, 9-10, 9-11, 9-12, 9-13, 9-14, 9-15, 9-16, 9-17, 9-18, 9-19, 9-20, 9-21, 9-22, 9-23, 9-24, 9-25, 9-26, 9-27, 9-28, 9-29, 9-30, 9-31, 9-32, 9-33, 9-34, 9-35, 9-36, 9-37, 9-38) (e.g., 10-11, 10-12, 10-13, 10-14, 10-15, 10-16, 10-17, 10-18, 10-19, 10-20, 10-21, 10-22, 10-23, 10-24, 10-25, 10-26, 10-27, 10-28, 10-29, 10-30, 10-31, 10-32, 10-33, 10-34, 10-35, 10-36, 10-37, 10-38) (For example, 11-12, 11-13, 11-14, 11-15, 11-16, 11-17, 11-18, 11-19, 11-20, 11-21, 11-22, 11-23, 11-24, 11-25, 11-26, 11-27, 11-28, 11-29, 11-30, 11-31, 11-32, 11-33, 11-34, 11-35, 11-36, 11-37, 11-38) (For example,12-13, 12-14, 12-15, 12-16, 12-17, 12-18, 12-19, 12-20, 12-21, 12-22, 12-23, 12-24, 12-25, 12-26, 12-27, 12-28, 12-29, 12-30, 12-31, 12-32, 12-33, 12-34, 12-35, 12-36, 12-37, 12-38) (For example, 13-14, 13-15, 13-16, 13-17, 13-18, 13-19, 13-20, 13-21, 13-22, 13-23, 13-24, 13-25, 13-26, 13-27, 13-28, 13-29, 13-30, 13-31, 13-32, 13-33, 13-34, 13-35, 13-36, 13-37, 13-38) (For example, 14-15, 14-16, 14-17, 14-18, 14-19, 14-20, 14-21, 14-22, 14-23, 14-24, 14-25, 14-26, 14-27, 14-28, 14-29, 14-30, 14-31, 14-32, 14-33, 14-34, 14-35, 14-36, 14-37, 14-38) (e.g., 15-16, 15-17, 15-18, 15-19, 15-20, 15-21, 15-22, 15-23, 15-24, 15-25, 15-26, 15-27, 15-28, 15-29, 15-30, 15-31, 15-32, 15-33, 15-34, 15-35, 15-36, 15-37, 15-38) (For example, 16-17, 16-18, 16-19, 16-20, 16-21, 16-22, 16-23, 16-24, 16-25, 16-26, 16-27, 16-28, 16-29, 16-30, 16-31, 16-32, 16-33, 16-34, 16-35, 16-36, 16-37, 16-38) (For example, 17-18, 17-19, 17-20, 17-21, 17-22, 17-23, 17-24, 17-25, 17-26, 17-27, 17-28, 17-29, 17-30, 17-31, 17-32, 17-33, 17-34, 17-35, 17-36, 17-37, 17-38) (e.g., 18-19, 18-20, 18-21, 18-22, 18-23, 18-24, 18-25, 18-26, 18-27, 18-28, 18-29, 18-30, 18-31, 18-32, 18-33, 18-34, 18-35, 18-36, 18-37, 18-38) (e.g.,19-20, 19-21, 19-22, 19-23, 19-24, 19-25, 19-26, 19-27, 19-28, 19-29, 19-30, 19-31, 19-32, 19-33, 19-34, 19-35, 19-36, 19-37, 19-38) (e.g., 20-21, 20-22, 20-23, 20-24, 20-25, 20-26, 20-27, 20-28, 20-29, 20-30, 20-31, 20-32, 20-33, 20-34, 20-35, 20-36, 20-37, 20-38) (For example, 21-22, 21-23, 21-24, 21-25, 21-26, 21-27, 21-28, 21-29, 21-30, 21-31, 21-32, 21-33, 21-34, 21-35, 21-36, 21-37, 21-38) (For example, 22-23, 22-24, 22-25, 22-26, 22-27, 22-28, 22-29, 22-30, 22-31, 22-32, 22-33, 22-34, 22-35, 22-36, 22-37, 22-38) (For example, 23-24, 23-25, 23-26, 23-27, 23-28, 23-29, 23-30, 23-31, 23-32, 23-33, 23-34, 23-35, 23-36, 23-37, 23-38) (For example, 24-25, 24-26, 24-27, 24-28, 24-29, 24-30, 24-31, 24-32, 24-33, 24-34, 24-35, 24-36, 24-37, 24-38) (For example, 25-26, 25-27, 25-28, 25-29, 25-30, 25-31, 25-32, 25-33, 25-34, 25-35, 25-36, 25-37, 25-38) (For example, 26-27, 26-28, 26-29, 26-30, 26-31, 26-32, 26-33, 26-34, 26-35, 26-36, 26-37, 26-38) (For example, 27-28, 27-29, 27-30, 27-31, 27-32, 27-33, 27-34, 27-35, 27-36, 27-37, 27-38) (e.g., 28-29, 28-30, 28-31, 28-32, 28-33, 28-34, 28-35, 28-36, 28-37, 28-38) (e.g., 29-30, 29-31, 29-32, 29-33, 29-34, 29-35, 29-36, 29-37, 29-38) (e.g.,30-31, 30-32, 30-33, 30-34, 30-35, 30-36, 30-37, 30-38) (For example, 31-32, 31-33, 31-34, 31-35, 31-36, 31-37, 31-38) (For example, 32-33, 32-34, 32-35, 32-36, 32-37, 32-38) (For example, 33-34, 33-35, 33-36, 33-37, 33-38) (For example, 34-35, 34-36, 34-37, 34-38) (For example, 35-36, 35-37, 35-38) (For example, 36-37, 36-38) (For example, 37-38) (e.g., 38 or less; 37 or less; 36 or less; 35 or less; 34 or less; 33 or less; 32 or less; 31 or less; 30 or less; 29 or less; 28 or less; 27 or less; 26 or less; 25 or less; 24 or less; 23 or less; 22 or less; 21 or less; 20 or less; 19 or less; 18 or less; 17 or less; 16 or less; 15 or less; 14 or less; 13 or less; 12 or less; 11 or less; 10 or less; 9 or less; 8 or less; 7 or less; 6 or less; 5 or less; 4 or less; 3 or less; 2 or 1).
[0046] As used herein, "nucleic acid" or "nucleic acid molecule" generally refers to any ribonucleic acid or deoxyribonucleic acid, which can be unmodified or modified DNA or RNA. "Nucleic acid" includes but is not limited to single-stranded and double-stranded nucleic acids. As used herein, the term "nucleic acid" also includes DNA as described above containing one or more modified bases. Therefore, DNA whose backbone is modified for stability or other reasons is a "nucleic acid". As used herein, the term "nucleic acid" encompasses such chemically, enzymatically or metabolically modified forms of nucleic acids, as well as chemical forms of DNA characteristic of viruses and cells (including, for example, simple and complex cells).
[0047] The term "oligonucleotide" or "polynucleotide" or "nucleotide" or "nucleic acid" refers to a molecule having two or more, preferably more than three and usually more than ten deoxyribonucleotides or ribonucleotides. The exact size will depend on many factors, which in turn depend on the ultimate function or use of the oligonucleotide. Oligonucleotides can be produced in any manner, including chemical synthesis, DNA replication, reverse transcription, or a combination thereof. The typical deoxyribonucleotides of DNA are thymine, adenine, cytosine, and guanine. The typical ribonucleotides of RNA are uracil, adenine, cytosine, and guanine.
[0048] As used herein, the term "locus" or "region" of a nucleic acid refers to a subregion of a nucleic acid, such as a gene on a chromosome, a single nucleotide, a CpG island, and the like.
[0049] The terms "complementary" and "complementarity" refer to nucleotides (e.g., one nucleotide) or polynucleotides (e.g., a sequence of nucleotides) that are related by the base pairing rules. For example, the sequence 5'-AGT-3' is complementary to the sequence 3'-TCA-5'. Complementarity can be "partial," where only some of the nucleic acid bases match according to the base pairing rules. Alternatively, there can be "complete" or "total" complementarity between nucleic acids. The degree of complementarity between nucleic acid chains affects the efficiency and intensity of hybridization between nucleic acid chains. This is particularly important in amplification reactions and detection methods that rely on binding between nucleic acids.
[0050] The term "gene" refers to a nucleic acid (e.g., DNA or RNA) sequence comprising the coding sequence necessary to produce RNA or a polypeptide or its precursor. A functional polypeptide can be encoded by the full-length coding sequence or by any portion of the coding sequence, as long as the desired activity or functional properties of the polypeptide (e.g., enzymatic activity, ligand binding, signal transduction, etc.) are retained. When used to refer to a gene, the term "portion" refers to a fragment of the gene. The size of a fragment can vary from a few nucleotides to the entire gene sequence minus one nucleotide. Therefore, "nucleotides comprising at least a portion of a gene" can include gene fragments or the entire gene.
[0051] The term "gene" encompasses the coding region of a structural gene and includes sequences adjacent to the coding region at both the 5' and 3' ends, such that the gene corresponds to the length of the full-length mRNA (e.g., including coding, regulatory, structural, and other sequences). Sequences located 5' to the coding region and present on the mRNA are referred to as 5' non-translated or untranslated sequences. Sequences located 3' or downstream of the coding region and present on the mRNA are referred to as 3' non-translated or 3' untranslated sequences. The term "gene" encompasses both cDNA and genomic forms of a gene. In some organisms (e.g., eukaryotes), genomic forms or clones of a gene contain coding regions interrupted by non-coding sequences known as "introns," "insertion regions," or "insertion sequences." Introns are segments of a gene that are transcribed into nuclear RNA (hnRNA); introns may contain regulatory elements, such as enhancers. Introns are removed or "spliced out" from the nuclear or primary transcript; therefore, they are not present in the messenger RNA (mRNA) transcript. mRNA functions during translation to specify the sequence or order of amino acids in a nascent polypeptide. Based on the present disclosure, it will be understood by those skilled in the art that one or more CpG sites in a DMR can be located in the coding region of a gene, the non-coding regulatory region of a gene, or a non-coding region unknown to be associated with a specific gene, such as a region comprising a long non-coding RNA (lncRNA). In some embodiments, sequences corresponding to these regions can be obtained using accession numbers (see, e.g., Table 1) corresponding to genomic databases (e.g., GenBank, NCBI / Ensembl, UniProt, etc.). In some embodiments, one or more CpG sites in a DMR can be located in an unannotated genomic region. As further provided herein, SEQ ID NOs (see, e.g., Table 1; SEQ ID NOs: 1-6) can be used to describe an unannotated genomic region comprising one or more CpG sites in a DMR.
[0052] Based on the present disclosure, one of ordinary skill in the art will recognize that a variety of techniques can be used to determine the location of one or more CpG sites (e.g., CpG islands) within a gene or region and their association with a disease or disorder, including but not limited to those disclosed in Chen et al., “Methods for identifying differentially methylated regions for sequence- and array-based data,” Briefings in Functional Genomics, Vol. 15, No. 6, November 2016, pp. 485-490, which is incorporated herein by reference in its entirety for all purposes.
[0053] The term "wild-type" when referring to a gene refers to a gene that has the characteristics of a gene isolated from a naturally occurring source. The term "wild-type" when referring to a gene product refers to a gene product that has the characteristics of a gene product isolated from a naturally occurring source. The term "wild-type" when referring to a protein refers to a protein that has the characteristics of a naturally occurring protein. The term "naturally occurring" as applied to an object refers to the fact that the object can be found in nature. For example, a polypeptide or polynucleotide sequence present in an organism (including a virus) that can be isolated from a natural source and has not been intentionally modified by a laboratory worker is naturally occurring. A wild-type gene is generally the gene or allele most commonly observed in a population and is therefore arbitrarily designated as the "normal" or "wild-type" form of a gene. In contrast, the terms "modified" or "mutated" when referring to a gene or gene product refer to a gene or gene product that exhibits modifications in sequence and / or functional properties (e.g., altered characteristics) compared to the wild-type gene or gene product, respectively. Note that naturally occurring mutants can be isolated; these are identified by the fact that they have altered characteristics compared to the wild-type gene or gene product.
[0054] The term "allele" refers to a variation of a gene; said variation includes, but is not limited to, variants and mutants, polymorphic loci and single nucleotide polymorphic loci, frameshift and splice mutations. An allele may occur naturally in a population or may occur during the lifetime of any particular individual in a population.
[0055] Thus, the terms "variant" and "mutant" when used in reference to nucleotide sequences refer to a nucleic acid sequence that differs from another, generally related, nucleotide sequence by one or more nucleotides. A "variation" is a difference between two different nucleotide sequences; typically, one sequence is a reference sequence.
[0056] The term "primer" refers to an oligonucleotide, whether naturally occurring (e.g., a nucleic acid fragment from a restriction digest) or synthetically produced, which can serve as a starting point for synthesis when placed under conditions that induce the synthesis of primer extension products complementary to a nucleic acid template strand (e.g., in the presence of nucleotides and an inducing agent such as a DNA polymerase, and at a suitable temperature and pH). In order to achieve maximum amplification efficiency, the primer is preferably single-stranded, but may also be double-stranded. If double-stranded, the primer is first treated to separate its chain before it can be used to prepare an extension product. Preferably, the primer is an oligodeoxyribonucleotide. The primer must be long enough to initiate the synthesis of an extension product in the presence of an inducing agent. The exact length of the primer will depend on many factors, including the use of temperature, primer source, and method. In some embodiments, the primer is specific to the differentially methylated region (e.g., the DMR in Tables 1, 2, and 3) and specifically binds to at least a portion of the genetic region comprising the DMR.
[0057] The term "probe" refers to an oligonucleotide (e.g., a nucleotide sequence), whether naturally occurring (e.g., in a purified restriction digest) or synthesized, recombinant, or produced by PCR amplification, which is capable of hybridizing to another oligonucleotide of interest. The probe can be single-stranded or double-stranded. The probe can be used to detect, identify, and isolate a specific gene sequence (e.g., a "capture probe"). It is envisioned that in some embodiments, any probe used in the embodiments of the present disclosure may be labeled with any "reporter molecule" so that it can be detected in any detection system, including but not limited to enzymes (e.g., ELISA, and enzyme-based histochemical assays), fluorescence, radioactivity, and luminescence systems. The various embodiments of the present disclosure are not limited to any particular detection system or label.
[0058] As used herein, the term "target" refers to a nucleic acid that is sought to be separated from other nucleic acids, such as by probe binding, amplification, separation, capture, etc. For example, when used in reference to a polymerase chain reaction, "target" refers to the region of nucleic acid to which the primers used in the polymerase chain reaction bind, while when used in an assay that does not amplify target DNA, such as in some embodiments of an invasive cleavage assay, the target includes the site where the probe and invasive oligonucleotide (e.g., INVADER oligonucleotide) bind to form an invasive cleavage structure, thereby detecting the presence of the target nucleic acid. A "segment" is defined as a region of nucleic acid within a target sequence.
[0059] Thus, as used herein, "non-target," for example, when used to describe nucleic acids (e.g., DNA), refers to nucleic acids that may be present in a reaction but are not the subject of detection or characterization by the reaction. In some embodiments, non-target nucleic acids can refer to nucleic acids present in a sample that do not contain, for example, a target sequence, while in some embodiments, non-target can refer to exogenous nucleic acids, i.e., nucleic acids that are not derived from a sample containing or suspected of containing a target nucleic acid, and that are added to a reaction, for example, to normalize the activity of an enzyme (e.g., a polymerase) to reduce variation in enzyme performance in the reaction.
[0060] As used herein, "methylation" refers to methylation of cytosine at positions C5 or N4 of cytosine, N6 of adenine, or other types of nucleic acid methylation. In vitro amplified DNA is typically non-methylated because typical in vitro DNA amplification methods do not retain the methylation pattern of the amplified template. However, "unmethylated DNA" or "methylated DNA" may also refer to amplified DNA that is unmethylated or methylated, respectively, from the original template.
[0061] As used herein, the term "amplification reagents" refers to reagents required for amplification (deoxyribonucleoside triphosphates, buffer, etc.) excluding primers, nucleic acid templates, and amplification enzymes. Typically, amplification reagents are placed together with other reaction components and contained in a reaction vessel.
[0062] As used herein, the term "control" when used to refer to nucleic acid detection or analysis refers to a nucleic acid with known characteristics (e.g., a known sequence, a known number of copies per cell) for comparison with an experimental target (e.g., a nucleic acid of unknown concentration). The control can be an endogenous, preferably unchanging gene, for which the test nucleic acid or target nucleic acid in the assay can be standardized. This standardization controls for inter-sample variations that may occur, for example, in sample processing, assay efficiency, etc., and allows accurate inter-sample data comparison. Genes that can be used to standardize nucleic acid detection assays for human samples include, for example, b-actin, ZDHHC1, and B3GALT6 (see, for example, U.S. patent application serial numbers 14 / 966,617 and 62 / 364,082, each of which is incorporated herein by reference). As used herein, "ZDHHC1" refers to a gene encoding a protein characterized by a zinc finger, containing DHHC type 1, which is located on Chr 16 (16q22.1) in human DNA and belongs to the DHHC palmitoyltransferase family.
[0063] The control can also be external. For example, in quantitative assays (such as qPCR, QuARTS, etc.), "calibrator" or "calibration control" is a nucleic acid having a known sequence, for example, a sequence identical to a portion of an experimental target nucleic acid, and having a known concentration or a range of concentrations (for example, a serial dilution control target for generating a calibration curve in quantitative PCR). Typically, the calibration control is analyzed using the same reagents and reaction conditions as the experimental DNA. In certain embodiments, the measurement of the calibrator is carried out simultaneously with the experimental assay, for example, in the same thermal cycler. In a preferred embodiment, a plurality of calibrators can be included in a single plasmid so that different calibrator sequences are easily provided in equimolar amounts. In a particularly preferred embodiment, the plasmid calibrator is digested, for example, with one or more restriction enzymes to release the calibrator portion from the plasmid vector. See, for example, WO 2015 / 066695, which is incorporated herein by reference.
[0064] As used herein, "methylated nucleotide" or "methylated nucleotide base" refers to the presence of a methyl moiety on a nucleotide base, wherein the methyl moiety is not present in recognized typical nucleotide bases. For example, cytosine does not contain a methyl moiety on its pyrimidine ring, but 5-methylcytosine contains a methyl moiety at position 5 of its pyrimidine ring. Therefore, cytosine is not a methylated nucleotide, and 5-methylcytosine is a methylated nucleotide. In another example, thymine contains a methyl moiety at position 5 of its pyrimidine ring; however, for the purposes of this article, when thymine is present in DNA, it is not considered a methylated nucleotide because thymine is a typical nucleotide base of DNA.
[0065] As used herein, a "methylated nucleic acid molecule" refers to a nucleic acid molecule containing one or more methylated nucleotides.
[0066] As used herein, the "methylation state," "methylation profile," and "methylation status" of a nucleic acid molecule refers to the presence or absence of one or more methylated nucleotide bases in a nucleic acid molecule. For example, a nucleic acid molecule containing methylated cytosine is considered methylated (e.g., the methylation state of the nucleic acid molecule is methylated). A nucleic acid molecule that does not contain any methylated nucleotides is considered unmethylated.
[0067] As used herein, the term "methylation level" as applied to a methylation marker refers to the amount of methylation within a particular methylation marker. Methylation level can also refer to the amount of methylation within a particular methylation marker compared to a defined standard or control. Methylation level can also refer to whether one or more cytosine residues present in a CpG environment have or do not have a methylated group. Methylation level can also refer to the proportion of cells in a sample that have or do not have a methylated group on these cytosines. Methylation level can also alternatively describe whether a single CpG dinucleotide is methylated.
[0068] The methylation state of a particular nucleic acid sequence (e.g., a gene marker or DNA region as described herein) can indicate the methylation state of each base in the sequence, or can indicate the methylation state of a subset of bases within the sequence (e.g., one or more cytosines), or can indicate information about the methylation density of a region within the sequence, with or without providing precise information about the location within the sequence where methylation occurs.
[0069] The methylation state of a nucleotide locus in a nucleic acid molecule refers to the presence or absence of a methylated nucleotide at a specific locus in the nucleic acid molecule. For example, when the nucleotide present at the 7th nucleotide in a nucleic acid molecule is 5-methylcytosine, the methylation state of the cytosine at the 7th nucleotide in the nucleic acid molecule is methylated. Similarly, when the nucleotide present at the 7th nucleotide in a nucleic acid molecule is cytosine (rather than 5-methylcytosine), the methylation state of the cytosine at the 7th nucleotide in the nucleic acid molecule is unmethylated.
[0070] The methylation status can optionally be expressed or indicated using a "methylation value" (e.g., representing a methylation frequency, score, ratio, percentage, etc.). A methylation value can be generated, for example, by quantifying the amount of intact nucleic acid present after restriction digestion with a methylation-dependent restriction enzyme, or by comparing amplification profiles after a bisulfite reaction, or by comparing the sequences of bisulfite-treated and untreated nucleic acids, or by comparing TET-treated and untreated nucleic acids. Thus, a value, such as a methylation value, represents the methylation status and can therefore be used as a quantitative indicator of the methylation status across multiple copies of a locus. This is particularly useful when it is necessary to compare the methylation status of a sequence in a sample to a threshold or reference value.
[0071] As used herein, "methylation frequency" or "percent (%) methylation" refers to the number of instances in which a molecule or locus is methylated relative to the number of instances in which the molecule or locus is unmethylated.
[0072] As used herein, the term "methylation score" is a score indicating the methylation events detected in a marker or marker panel compared to the median methylation events of a marker or marker panel from a random population of mammals that do not have the specific neoplasm of interest (e.g., a random population of 10, 20, 30, 40, 50, 100, or 500 mammals). The methylation score that is elevated in a marker or marker panel can be any score as long as the score is greater than the corresponding reference score. For example, the methylation score that is elevated in a marker or marker panel can be 0.5, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more times higher than the reference methylation score.
[0073] Therefore, methylation state describes the methylation state of nucleic acid (such as genomic sequence). In addition, methylation state refers to the characteristics of the nucleic acid segment associated with methylation on a specific genomic locus. Such characteristics include, but are not limited to, whether any cytosine (C) residue in this DNA sequence is methylated, the position of the methylated C residue, the frequency or percentage of methylated C in any specific region of the nucleic acid, and the allelic differences of methylation due to, for example, differences in allelic origin. The terms "methylation state", "methylation overview" and "methylation status" also refer to the relative concentration, absolute concentration or pattern of methylated C or unmethylated C in any specific region of the nucleic acid in a biological sample. For example, if the cytosine (C) residue in the nucleic acid sequence is methylated, it can be referred to as "hypermethylation" or with "increased methylation", and if the cytosine (C) residue in the DNA sequence is not methylated, it can be referred to as "hypomethylation" or with "reduced methylation". In some embodiments, the present invention provides the methylation pattern of the present invention.For example, the methylation pattern of the present invention is the methylation pattern of the present invention.For example, the methylation pattern of the present invention is the methylation pattern of the present invention. For ... The term "differential methylation" refers to the difference in the level or pattern of nucleic acid methylation in a cancer-positive sample compared to the level or pattern of nucleic acid methylation in a cancer-negative sample. It may also refer to the difference in level or pattern between patients whose cancer relapsed after surgery and those who did not. Differential methylation and specific levels or patterns of DNA methylation are prognostic and predictive biomarkers, for example, once the correct cutoff or predictive features are defined. In some embodiments, one or more CpG sites in a DMR can be located in a non-coding region, such as a region corresponding to a long non-coding RNA (lncRNA).
[0074] Methylation state frequency can be used for describing individual colonies or samples from a single individual.For example, the methylation state frequency is that 50% nucleotide locus is methylated in 50% of the cases, and is unmethylated in 50% of the cases.Such frequency can be used for example to describe the degree of methylation of nucleotide locus or nucleic acid region in an individual colony or nucleic acid set.Therefore, when the methylation in the first colony or the pond of nucleic acid molecule is different from the methylation in the second colony or the pond of nucleic acid molecule, the methylation state frequency of the first colony or the pond will be different from the methylation state frequency of the second colony or the pond.Such frequency can also be used for example to describe the degree of methylation of nucleotide locus or nucleic acid region in a single individual.For example, such frequency can be used for describing the degree of methylation of nucleotide locus or nucleic acid region from a group of cells of a tissue sample or unmethylated at nucleotide locus or nucleic acid region.
[0075] Typically, methylation of human DNA occurs on dinucleotide sequences comprising adjacent guanine and cytosine, with cytosine located 5' to guanine (also known as CpG dinucleotide sequences). In the human genome, most cytosines within CpG dinucleotides are methylated, but in specific CpG dinucleotide-rich genomic regions (called CpG islands), some cytosines remain unmethylated (e.g., see Antequera et al. (1990) Cell 62: 503–514).
[0076] As used herein, "CpG island" or "cytosine-phosphate-guanine island" refers to a G:C-rich region of genomic DNA that contains an increased amount of CpG dinucleotides relative to the total genomic DNA. A CpG island can be at least 100, 200, or more base pairs in length, wherein the G:C content of the region is at least 50% and the ratio of the observed CpG frequency to the expected frequency is 0.6; in some cases, a CpG island can be at least 500 base pairs in length, wherein the G:C content of the region is at least 55% and the ratio of the observed CpG frequency to the expected frequency is 0.65. The ratio of the observed CpG frequency to the expected frequency can be calculated according to the method provided in Gardiner-Garden et al. (1987) J. Mol. Biol. 196: 261–281. For example, the ratio of the observed CpG frequency to the expected frequency can be calculated according to the formula R = (A × B) / (C × D), where R is the ratio of the observed CpG frequency to the expected frequency, A is the number of CpG dinucleotides in the analyzed sequence, B is the total number of nucleotides in the analyzed sequence, C is the total number of C nucleotides in the analyzed sequence, and D is the total number of G nucleotides in the analyzed sequence. Methylation status is typically determined in CpG islands, such as in promoter regions. However, it should be recognized that other sequences in the human genome are also susceptible to DNA methylation, such as CpA and CpT (see Ramsahoye (2000) Proc. Natl. Acad. Sci. USA 97: 5237–5242; Salmon and Kaye (1970) Biochim. Biophys. Acta. 204: 340-351; Grafstrom (1985) Nucleic Acids Res. 13: 2827-2842; Nyce (1986) Nucleic Acids Res. 14: 4353-4367; Woodcock (1987) Biochem. Biophys. Res. Commun. 145: 888-894).
[0077] As used herein, "methylation-specific reagent" refers to a reagent that modifies the nucleotides of a nucleic acid molecule according to the methylation state of the nucleic acid molecule, or a methylation-specific reagent refers to a compound or composition or other agent that can change the nucleotide sequence of a nucleic acid molecule in a manner that reflects the methylation state of the nucleic acid molecule. The method for treating a nucleic acid molecule with such a reagent may comprise contacting the nucleic acid molecule with the reagent, and adding additional steps, if necessary, to achieve the desired nucleotide sequence change. Such methods can be applied in a manner that unmethylated nucleotides (e.g., each unmethylated cytosine) are modified into different nucleotides. For example, in some embodiments, such reagents can deaminize unmethylated cytosine nucleotides to produce deoxyuracil residues. Examples of such reagents include, but are not limited to, methylation-sensitive restriction enzymes, methylation-dependent restriction enzymes, bisulfite reagents, TET enzymes, and borane reducing agents.
[0078] Alteration of the nucleic acid nucleotide sequence by a methylation-specific agent can also result in each methylated nucleotide in the nucleic acid molecule being modified to a different nucleotide.
[0079] The term "methylation assay" refers to any assay used to determine the methylation status of one or more CpG dinucleotide sequences within a nucleic acid sequence.
[0080] The term "MS AP-PCR" (methylation-sensitive arbitrarily primed polymerase chain reaction) refers to an art-recognized technique that allows global scanning of the genome using CG-rich primers to focus on regions most likely to contain CpG dinucleotides, as described in Gonzalgo et al. (1997) Cancer Research 57: 594–599.
[0081] The term "MethyLight™" refers to the art-recognized fluorescence-based real-time PCR technology described by Eads et al. (1999) Cancer Res. 59: 2302–2306.
[0082] The term "HeavyMethyl™" refers to an assay in which a methylation-specific blocking probe (also referred to herein as a blocker) covering a CpG position between or covered by amplification primers enables methylation-specific selective amplification of a nucleic acid sample.
[0083] The term "HeavyMethyl™ MethyLight™" assay refers to a HeavyMethyl™ MethyLight™ assay, which is a variation of the MethyLight™ assay in which the MethyLight™ assay is combined with a methylation-specific blocking probe covering the CpG positions between the amplification primers.
[0084] The term "Ms-SNuPE" (Methylation-Sensitive Single Nucleotide Primer Extension) refers to the art-recognized assay described by Gonzalgo and Jones (1997) Nucleic Acids Res. 25: 2529–2531.
[0085] The term "MSP" (methylation-specific PCR) refers to the art-recognized methylation assay described by Herman et al. (1996) Proc. Natl. Acad. Sci. USA 93: 9821-9826 and US Patent No. 5,786,146.
[0086] The term "COBRA" (combined bisulfite restriction analysis) refers to the art-recognized methylation assay described by Xiong and Laird (1997) Nucleic Acids Res. 25: 2532–2534.
[0087] The term "MCA" (methylated CpG island amplification) refers to the methylation assay described in Toyota et al. (1999) Cancer Res. 59: 2307-12 and WO 00 / 26401A1.
[0088] As used herein, "selected nucleotide" refers to one of the four typically occurring nucleotides in a nucleic acid molecule (C, G, T, and A for DNA and C, G, U, and A for RNA), and may include methylated derivatives of typically occurring nucleotides (e.g., when C is the selected nucleotide, both methylated and unmethylated C are included within the meaning of the selected nucleotide), while a methylated selected nucleotide specifically refers to a methylated typically occurring nucleotide and an unmethylated selected nucleotide specifically refers to an unmethylated typically occurring nucleotide.
[0089] The term "methylation-specific restriction enzyme" refers to a restriction enzyme that selectively digests nucleic acids according to the methylation state of the nucleic acid recognition site. In the case of a restriction enzyme that specifically cuts when the recognition site is not methylated or hemimethylated (methylation-sensitive enzyme), if the recognition site is methylated on one or both chains, cutting will not occur (or the efficiency is significantly reduced). In the case of a restriction enzyme that specifically cuts when only the recognition site is methylated (methylation-dependent enzyme), if the recognition site is not methylated, cutting will not occur (or it will occur, but the efficiency is significantly reduced). Preferably, the methylation-specific restriction enzyme has a recognition sequence containing a CG dinucleotide (for example, a recognition sequence, such as CGCG or CCCGGG). Further preferred for some embodiments is a restriction enzyme that will not cut if the cytosine in this dinucleotide is methylated at the carbon atom C5 place.
[0090] As used herein, the terms "copy number variation," "CNV," "copy number aberration," and "CNA" generally refer to a change in the number of copies of a nucleic acid sequence present in a test sample compared to the number of copies of a nucleic acid sequence present in a reference sample. In some cases, the nucleic acid sequence is an entire chromosome or a significant portion thereof. "Copy number variant" refers to a nucleic acid sequence that is found to have a copy number difference by comparing the expected level of the nucleic acid sequence of interest with the nucleic acid sequence of interest in a test sample. For example, the level of the nucleic acid sequence of interest in a test sample is compared with the level of the nucleic acid sequence present in a qualified sample. Copy number variants / variations include deletions (including microdeletions), insertions (including microinsertions), duplications, proliferations, and translocations. CNVs include chromosomal aneuploidy, partial aneuploidy, polyploidy, and partial polyploidy. In some cases, analyzing a CNV of a nucleic acid sample refers to characterizing the state of a chromosome or fragment aneuploidy by one of three types of calls (e.g., normal or unaffected, affected, and no calls). Normal and affected thresholds are typically set. Parameters associated with aneuploidy or other copy number variations can be measured in a sample, and the measured values are compared with the thresholds. For example, for repetitive aneuploidy, if the chromosome or segment dose (or other measured value sequence content) is higher than the definition threshold of the affected sample, the call is affected. For such aneuploidy, if the chromosome or segment dose is lower than the threshold value set for the normal sample, the call is normal. In contrast, for deletion aneuploidy, if the chromosome or segment dose is lower than the definition threshold of the affected sample, the call is affected, and if the chromosome or segment dose is higher than the threshold value set for the normal sample, the call is normal. For example, in the presence of trisomy, "normal" is called to be determined by a parameter value, for example, a test chromosome dose lower than a user-defined reliability threshold, and "affected" is called to be determined by a parameter value, for example, a test chromosome dose higher than a user-defined reliability threshold. The "no call" result is determined by a parameter, for example, a test chromosome dose between the thresholds for making a "normal" or "affected" call. The term "no call" can be used interchangeably with "unclassified".
[0091] The term "aneuploidy" herein generally refers to the imbalance of genetic material, such as caused by the loss or gain of a whole chromosome or a part of a chromosome. The terms "partial aneuploidy" and "partial chromosome aneuploidy" herein refer to the imbalance of genetic material caused by the loss or gain of a part of a chromosome, such as partial monosomy and partial trisomy, and include the imbalance caused by translocation, deletion and insertion. The terms "chromosome aneuploidy" and "complete chromosome aneuploidy" herein refer to the imbalance of genetic material caused by the loss or gain of a whole chromosome, and include germline aneuploidy and mosaic aneuploidy. For example, aneuploidy may be due to a set of extra chromosomes (such as meiotic errors) causing, which may lead to congenital diseases. This aneuploidy may be due to chromosomes failing to separate correctly during meiosis, or sperm fertilizing an egg with more than one set of chromosomes. In other cases, aneuploidy is caused by chromosomal instability (CIN) due to failure of the mitotic checkpoint, resulting in chromosome missegregation (e.g., mitotic errors), leading to gain of oncogenes or loss of tumor suppressor genes in cancer-associated disease states.
[0092] As described herein, the term "aneuploidy score" or "AS" generally refers to the total number of chromosome arms altered in a sample, ranging from 0 (no arms) to 39 (all arms - the long and short arms of each non-acromial chromosome, and the long arms of chromosomes 13, 14, 15, 21, and 22). As further described herein, the aneuploidy score (AS) calculated for each sample is the total number of chromosome arm-level gains and losses adjusted for ploidy.
[0093] As used herein, the "sensitivity" of a given marker (or a set of markers used together) refers to the percentage of samples reporting DNA methylation values above a threshold value that distinguishes neoplastic samples from non-neoplastic samples. In some embodiments, a positive is defined as a histologically confirmed neoplasm formation with a reported DNA methylation value above a threshold value (e.g., a range associated with the disease), and a false negative is defined as a histologically confirmed neoplasm formation with a reported DNA methylation value below a threshold value (e.g., a range associated with no disease). Thus, the value of sensitivity reflects the probability that the DNA methylation measurement value of a given marker obtained from a known disease sample is within the disease-related measurement range. As defined herein, the clinical relevance of a calculated sensitivity value represents an estimate of the probability of detecting the presence of a given marker when applied to a subject with a clinical disease.
[0094] As used herein, the "specificity" of a given marker (or a set of markers used together) refers to the percentage of non-neoplastic samples that report DNA methylation values below a threshold value that distinguishes neoplastic samples from non-neoplastic samples. In some embodiments, a negative is defined as a histologically confirmed non-neoplastic sample that reports a DNA methylation value below a threshold value (e.g., a range associated with the absence of disease), and a false positive is defined as a histologically confirmed non-neoplastic sample that reports a DNA methylation value above a threshold value (e.g., a range associated with disease). Thus, the value of specificity reflects the probability that a DNA methylation measurement value for a given marker obtained from a known non-neoplastic sample is within a non-disease-related measurement range. As defined herein, the clinical relevance of a calculated specificity value represents an estimate of the probability of detecting the absence of a given disease when the given marker is applied to a patient who does not suffer from the clinical disease.
[0095] As used herein, the term "AUC" is an abbreviation for "area under the curve." It refers in particular to the area under the receiver operating characteristic (ROC) curve. The ROC curve is a graph of the true positive rate and the false positive rate for different possible cut-off points of a diagnostic test. It shows the trade-off between sensitivity and specificity (any increase in sensitivity will be accompanied by a decrease in specificity) based on the selected cut-off point. The area under the ROC curve (AUC) is a measure of the accuracy of a diagnostic test (the larger the area, the better; the optimal value is 1; the ROC curve for a random test is on the diagonal with an area of 0.5; Reference: JP Egan. (1975) Signal Detection Theory and ROC Analysis, Academic Press, New York).
[0096] As used herein, the term "neoplasm" refers to any new abnormal growth of tissue. Thus, a neoplasm can be a precancerous neoplasm or a malignant neoplasm.
[0097] The term "vegetation-specific marker" as used herein refers to any biological material or element that can be used to indicate the presence of a vegetation. Examples of biological materials include, but are not limited to, nucleic acids, polypeptides, carbohydrates, fatty acids, cell components (e.g., cell membranes and mitochondria) and whole cells. In some cases, a marker is a specific nucleic acid region (e.g., a gene, an intragenic region, a specific locus, etc.). A nucleic acid region as a marker can be referred to as, for example, a "marker gene," "marker region," "marker sequence," "marker locus," etc.
[0098] As used herein, the term "adenoma" refers to a benign tumor of glandular origin. Although these growths are benign, over time they may develop into a malignant tumor (e.g., esophageal adenocarcinoma or EAC).
[0099] The term "precancerous" or "preneoplastic" and equivalents thereof refer to any cell proliferative disorder that is undergoing malignant transformation. For example, as further described herein, low-grade dysplasia (LGD) BE and high-grade dysplasia (HGD) BE are considered precancerous lesions.
[0100] As used herein, the term "esophageal disorder" refers to a type of disorder associated with the esophagus and / or esophageal tissue. Examples of esophageal disorders include, but are not limited to, Barrett's esophagus (BE), non-dysplastic Barrett's esophagus (NDBE), Barrett's esophagus dysplasia (BED), Barrett's esophagus low-grade dysplasia (BE-LGD), Barrett's esophagus high-grade dysplasia (BE-HGD), and esophageal adenocarcinoma (EAC).
[0101] The "site" of a neoplasm, adenoma, cancer, etc. is the tissue, organ, cell type, anatomical region, body part, etc., within a subject where the neoplasm, adenoma, cancer, etc. is located.
[0102] As used herein, "diagnostic" test applications include detecting or identifying a disease state or condition in a subject, determining the likelihood that a subject is infected with a given disease or condition, determining the likelihood that a subject with a disease or condition will respond to therapy, determining the prognosis of a subject with a disease or condition (or its likely progression or regression), and determining the effect of a treatment on a subject with a disease or condition. For example, a diagnostic test can be used to detect the presence or likelihood of a subject being infected with a neoplasm, or the likelihood that such a subject will respond favorably to a compound (e.g., a drug, e.g., pharmaceutical) or other treatment.
[0103] When the term "isolated" is applied to a nucleic acid (e.g., an "isolated oligonucleotide"), it refers to a nucleic acid sequence that has been identified and separated from at least one contaminating nucleic acid with which it is normally associated in its natural source. An isolated nucleic acid is present in a form or setting different from that in which it is found in nature. In contrast, unisolated nucleic acids, such as DNA and RNA, are found in the state in which they exist in nature. Examples of unisolated nucleic acids include a given DNA sequence (e.g., a gene) found adjacent to a gene on a host cell chromosome; an RNA sequence, such as a specific mRNA sequence encoding a specific protein, is found in a cell as a mixture with many other mRNAs encoding a variety of proteins. However, an isolated nucleic acid encoding a specific protein includes, for example, such a nucleic acid in a cell that normally expresses the protein, where the nucleic acid is in a chromosomal location different from that of the natural cell, or is otherwise flanked by nucleic acids different from those found in nature. An isolated nucleic acid or oligonucleotide can exist in single-stranded or double-stranded form. When an isolated nucleic acid or oligonucleotide is used to express a protein, the oligonucleotide will contain at least the sense strand or coding strand (i.e., the oligonucleotide can be single-stranded), but may contain both the sense strand and the antisense strand (i.e., the oligonucleotide can be double-stranded). An isolated nucleic acid can be combined with other nucleic acids or molecules after being separated from its natural or typical environment. For example, an isolated nucleic acid can be present in a host cell into which it is placed, for example, for heterologous expression.
[0104] The term "purified" refers to a molecule, whether a nucleic acid or an amino acid sequence, that is removed, separated or isolated from its natural environment. Thus, an "isolated nucleic acid sequence" can be a purified nucleic acid sequence. A "substantially purified" molecule is at least 60% free, preferably at least 75% free, and more preferably at least 90% free from other components with which it is naturally associated. As used herein, the term "purified" or "purifying" also refers to the removal of contaminants from a sample. Removal of contaminating proteins increases the percentage of the polypeptide or nucleic acid of interest in the sample. In another example, the recombinant polypeptide is expressed in a plant, bacterial, yeast or mammalian host cell, and the polypeptide is purified by removing host cell proteins; thereby increasing the percentage of the recombinant polypeptide in the sample.
[0105] The term "composition comprising" a given polynucleotide sequence or polypeptide broadly refers to any composition containing the given polynucleotide sequence or polypeptide. The composition may comprise an aqueous solution containing a salt (e.g., NaCl), a detergent (e.g., SDS), and other ingredients (e.g., Denhardt's solution, powdered milk, salmon sperm DNA, etc.).
[0106] The term "sample" is used in its broadest sense. In one sense, it can refer to animal cells or tissues. In another sense, it refers to specimens or cultures obtained from any source, as well as biological samples and environmental samples. Biological samples can be obtained from plants or animals (including humans) and encompass fluids, solids, tissues, and gases. Environmental samples include environmental materials such as surface materials, soil, water, and industrial samples. These examples should not be construed as limiting the sample types applicable to the various embodiments of the present disclosure.
[0107] As used herein, "remote sample" as used in some instances relates to a sample that is collected indirectly from a site that is not the source of the cells, tissues, or organs of the sample. For example, when a sample material originating from the pancreas is assessed in a stool sample, the sample is a remote sample.
[0108] As used herein, the term "patient" or "subject" refers to an organism to be subjected to the various tests described herein. The term "subject" includes animals, preferably mammals, including humans. In a preferred embodiment, the subject is a primate. In an even more preferred embodiment, the subject is a human. Further with respect to the diagnostic methods, the preferred subject is a vertebrate subject. The preferred vertebrate is warm-blooded; the preferred warm-blooded vertebrate is a mammal. The preferred mammal is most preferably a human. As used herein, the term "subject" includes both human and animal subjects. Thus, veterinary therapeutic uses are provided herein. Thus, the present disclosure provides diagnostics for mammals, such as humans, as well as those mammals that are important because they are endangered, such as Siberian tigers; mammals of economic importance, such as animals raised on farms for human consumption; and / or animals of social importance to humans, such as animals kept as pets or in zoos. Examples of such animals include, but are not limited to, carnivorous plants such as cats and dogs; swine, including pigs, hogs, and wild boars; ruminants and / or ungulates, such as cattle, bulls, sheep, giraffes, deer, goats, bison, and camels; pinnipeds; and horses. Thus, diagnostics and treatments for livestock are also provided, including, but not limited to, domestic pigs, ruminants, ungulates, horses (including racehorses), and the like.
[0109] As used herein, the term "test kit" refers to any delivery system for delivering materials. In the case of a reaction assay, such a delivery system includes a system that allows storage of reaction reagents (e.g., oligonucleotides, enzymes, etc. in appropriate containers) and / or support materials (e.g., buffers, written instructions for performing an assay, etc.), transporting them from one location or delivering them to another. For example, a test kit includes one or more housings (e.g., boxes) containing relevant reaction reagents and / or support materials. As used herein, the term "dispersed test kit" refers to a delivery system comprising two or more separate containers, each containing a sub-portion of all test kit components. These containers can be delivered to a predetermined recipient together or individually. For example, a first container may contain an enzyme for assaying, while a second container contains oligonucleotides. The term "dispersed test kit" is intended to encompass test kits containing analyte-specific reagents (ASRs) regulated by Section 520 (e) of the Federal Food, Drug, and Cosmetic Act, but is not limited thereto. In fact, any delivery system comprising two or more separate containers each containing a sub-portion of all test kit components is included in the term "dispersed test kit." In contrast, a "combination kit" refers to a delivery system that contains all components of a reaction assay in a single container (eg, a single box containing each required component). The term "kit" includes both discrete kits and combination kits.
[0110] As used herein, the term "information" refers to a collection of any facts or data. When referring to information stored or processed using a computer system (including but not limited to the Internet), the term refers to any data stored in any format (such as analog, digital, optical, etc.). The term "information related to a subject" as used herein refers to facts or data belonging to a subject (such as humans, plants, or animals). The term "genomic information" refers to information related to a genome, including but not limited to nucleic acid sequences, genes, methylation percentages, allele frequencies, RNA expression levels, protein expression, phenotypes related to genotypes, etc. "Allele frequency information" refers to facts or data related to allele frequencies, including but not limited to statistical correlations between allele identity, the presence of alleles and a characteristic of a subject (such as a human subject), the presence or absence of alleles in an individual or population, the probability percentage of the presence of alleles in an individual with one or more specific characteristics, etc.
[0111] 2. Methylated DNA Markers and Biomarker Panels
[0112] A genome-wide methylation approach is used to simultaneously query copy number aberrations (CNA) and DNA methylation to distinguish non-dysplastic Barrett's esophagus (NDBE) from high-grade dysplastic Barrett's esophagus (HGD-BE) or esophageal adenocarcinoma (EAC), and potentially as an adjunct to endoscopic histological monitoring. As further described herein, the present disclosure provides materials and methods for measuring copy number aberrations (CNA) using sequencing reads from genomes modified in a methylation-specific manner (e.g., genomes converted to cytosine or 5-methylcytosine), including polyploidy and aneuploidy (e.g., determining aneuploidy scores). Most currently available technologies use direct NGS for wild-type, unconverted DNA. However, embodiments of the present disclosure include the ability to perform methylation and CNV analysis from the same chemistry / dataset without the need to split the sample to perform the corresponding analysis separately. That is, methylation analysis and CNV analysis can be performed simultaneously on the same converted DNA sample. According to these embodiments, the same DNA sample obtained from the subject can be used to determine the methylation profile in at least one DMR and CNV and / or AS. In some embodiments, a single DNA sample obtained from a subject can be used to determine the methylation profile in at least one DMR and CNV and / or AS. In some embodiments, the sample (e.g., the same sample or a single sample) has been treated with an agent that modifies DNA in a methylation-specific manner.
[0113] Therefore, as further described herein, embodiments of the present disclosure provide methods, compositions and systems for screening various types of esophageal disorders in biological samples. According to these embodiments, the present disclosure includes but is not limited to methods and compositions for detecting the presence of esophageal cancer or precancerous lesions from biological samples. In some embodiments, the biological sample is a tissue sample, a blood sample, a plasma sample, a serum sample, a whole blood sample, a buffy coat sample, a secretion sample, an organ secretion sample, a cerebrospinal fluid (CSF) sample, a saliva sample, a urine sample and / or a fecal sample. In some embodiments, the tissue sample is an esophageal tissue sample. In some embodiments, the esophageal sample is obtained from an esophageal biopsy, or obtained by wiping, brushing or using a sponge capsule device. In some embodiments, the subject is a human being.
[0114] As further described herein, embodiments of the present disclosure include novel differentially methylated regions (DMRs), each DMR individually capable of distinguishing esophageal cancer or precancerous lesions from controls or benign tissue. In some embodiments, the novel DMRs are capable of distinguishing high-grade dysplastic Barrett's esophagus (HGD-BE) or esophageal adenocarcinoma (EAC) from non-dysplastic Barrett's esophagus (NDBE) or normal esophageal controls. According to these embodiments, the novel DMR is from a gene selected from the group consisting of: ACVRL1, ADAMTS8, ADAP2, ADRBK2, AKR1B1, ANK1, ANKRD13B, ANXA6, ARNT2, B4GALNT2, BACH2, BCL11A, BSCL2, C12orf53, C14orf82, C17orf107, C18orf1, C1orf95, C5orf42, CACNA1C, CAMK1D, CAMTA1, CBX6, CCDC85A, CCKBR, CD38, CDKN 2A, CH25H, CHST1, CHST15, CNTLN, CRHR1, CRTC1, CXCR4, CYP1B1, DCTN2, DIDO1, DMKN, DSE, DYNC1I1, EML6, ENOX1, EPHA4, ESRRG, FA M176A, FAM78B, FBXO10, FERMT2, FHOD3, FLJ45079, FMNL1, FOXP2, FRMD4B, FZD8, GALNTL1, GALNTL4, GAS1, GBGT1, GLIPR2, GNAI1, GN AL, GPR37, GRASP, GRID1, GRM8, GSC, HAR1A, HCN2, HEY2, HIST1H2BE, HS3ST3B1, HTR7, IGFBP2, INHA, INSRR, IRX3, ISM2, KCNG3, KCN K4, KCNS2, KCTD15, KIAA1522, KIAA1614, KIF26A, KL, KLF15, KLHL10, KRT77, KSR2, LBH, LMX1A, LMX1B, LOC100526820, LONRF2, LRFN 2. LRRN1, MAF, MAFB, MARK1, ADAMTSL4-AS1, PGBD5, HSPA12A, SFTPD, LOC107984507, LOC100128253, CISTR, SLC16A7, CTXND1, MAX. chr15.4912, ZNF423, RBFOX1, LOC105376772, GSE1, MAX.chr17.8070, ZNF709_6125, MAX.chr19.3439, GREB1, CYP27C1, AC068134.6. LOC388780, STOX2_8441, PPARGC1A, MAX.chr4.4552, PDZD2, MAX.chr6.6743, LPAL2, CDK14, MAX.chr8.3003, MCOLN2, MEGF11, MFSD11, MRC2, MSX1, NAT8L, N AV1, NBEA, NCRNA00092, NECAB2, NEURL, NGEF, NHLH2, NKD1, NLGN1, NOG, NPPC, NR3C1, NRXN2, NTN1, OBSCN, OCA2, OSBPL1A, OXR1, P2RY1, PDE2A, PDE9A, PDGFRA, P 808, and ZNF845 (Table 1), including any combination thereof. In some embodiments, the novel DMR is from any gene selected from Table 1, including any combination thereof. Each novel DMR alone is capable of distinguishing HGD-BE and / or EAC from NDBE and / or control samples, and combining two or more novel DMRs may provide improved sensitivity. Therefore, combinations of two or more novel DMRs selected from Table 1 are provided.
[0115] Embodiments of the present disclosure also include novel differentially methylated regions (DMRs), each DMR individually capable of distinguishing esophageal cancer or precancerous lesions from controls or benign tissue. In some embodiments, the novel DMRs are capable of distinguishing high-grade dysplastic Barrett's esophagus (HGD-BE) or esophageal adenocarcinoma (EAC) from non-dysplastic Barrett's esophagus (NDBE) or normal esophageal controls. According to these embodiments, the novel DMRs are from genes selected from the group consisting of ACVRL1, ADAMTS8, ADAP2, ADRBK2, AKR1B1, ANK1, ANKRD13B, ANXA6, ARNT2, B4GALNT2, BACH2, BCL11A, BSCL2, C14orf82, C18orf1, C1orf95, C5orf42, CAMK1D, CAMTA1, CCDC85A, CD38, CD KN2A, CHST1, CHST15, CRHR1, CYP1B1, DIDO1, DSE, DYNC1I1, EML6, ENOX1, EPHA4, FAM176A, FBXO10, FERMT2, FHOD3, FLJ45079, FMNL1, FRMD4B, FZD8, GALNTL1, GALNTL4, GAS1, GBGT1, GLIPR2, GNAI1, GNAL, GRASP, GRID 1. GRM8, GSC, HAR1A, HCN2, HEY2, HIST1H2BE, HS3ST3B1, HTR7, IGFBP2, INHA, IRX3, ISM2, KCNG3, KCNK4, KCN S2, KCTD15, KIAA1614, KL, KLHL10, KSR2, LBH, LMX1A, LMX1B, LOC100526820, LONRF2, LRFN2, MAFB, MARK1, P GBD5, HSPA12A, SFTPD, LOC107984507, LOC100128253, CISTR, SLC16A7, CTXND1, ZNF423, LOC105376772, GS E1, MAX.chr19.3439, GREB1, CYP27C1, AC068134.6, LOC388780, STOX2_8441, PPARGC1A, PDZD2, MAX.chr6.6743, LPAL2, CDK14, MCOLN2, MEGF11, MRC2, MSX1, NAT8L, NAV1, NBEA, NCRNA00092, NEURL, NGEF, NHLH2, NKD1, NLGN1, NOG, NPPC, NR3C1, NRXN2, OBSCN, OCA2, OSBPL1A, OXR1, P2RY1, PDE2A, PDGFRA, PID1, PLCG2, PLCL1, PNPLA3, POU3F1, P OU3F2, PRKACB, PRKAR2B, PRKG1, PRR18, PRR5L, PTHLH, PYGL, RHBDL3, RIMS2, ROR2, RPRML, SDK2, SOX9, STOX2_6730, SYNGR1, TAC4, TRANK1, TSPAN4, UBE2E2, UCHL1, UNC5A, VASH2, ZBTB10, ZNF680, ZNF709_4918, ZNF738, ZNF808, and ZNF845 (Table 2), including any combination thereof. In some embodiments, the novel DMR is from any gene selected from Table 2, including any combination thereof. Each novel DMR alone is capable of distinguishing HGD-BE and / or EAC from NDBE and / or control samples, and combining two or more novel DMRs can provide improved sensitivity. Thus, combinations of two or more novel DMRs selected from Table 2 are provided.
[0116] Embodiments of the present disclosure also include novel differentially methylated regions (DMRs), each DMR individually capable of distinguishing esophageal cancer or precancerous lesions from controls or benign tissue. In some embodiments, the novel DMRs are capable of distinguishing high-grade dysplastic Barrett's esophagus (HGD-BE) or esophageal adenocarcinoma (EAC) from non-dysplastic Barrett's esophagus (NDBE) or normal esophageal controls. According to these embodiments, the novel DMRs are from genes selected from the group consisting of BACH2, C5orf42, FHOD3, HIST1H2BE, IRX3, KIAA1614, LONRF2, MAFB, PDGFRA, PID1, POU3F1, PRR5L, RHBDL3, and SDK2 (Table 3), including any combination thereof. In some embodiments, the novel DMRs are from any gene selected from Table 3, including any combination thereof. Each novel DMR individually is capable of distinguishing HGD-BE and / or EAC from NDBE and / or control samples, and combining two or more novel DMRs can provide improved sensitivity. Therefore, a combination of two or more novel DMRs selected from Table 3 is provided.
[0117] Embodiments of the present disclosure also include novel differentially methylated regions (DMRs), each DMR individually capable of distinguishing esophageal cancer or precancerous lesions from controls or benign tissue. In some embodiments, the novel DMRs are capable of distinguishing high-grade dysplastic Barrett's esophagus (HGD-BE) or esophageal adenocarcinoma (EAC) from non-dysplastic Barrett's esophagus (NDBE) or normal esophageal controls. According to these embodiments, the novel DMRs are from genes selected from the group consisting of KL, PGBD5, ROR2, and LMX1B (Example 4), including any combination thereof. In some embodiments, the novel DMRs are from any gene selected from Example 4, including any combination thereof. Each novel DMR individually is capable of distinguishing HGD-BE and / or EAC from NDBE and / or control samples, and combining two or more novel DMRs can provide increased sensitivity. Thus, a combination of two or more novel DMRs selected from Example 4 is provided.
[0118] According to the above, the control sample includes a sample from a subject who does not suffer from cancer, a sample from a subject who does not suffer from esophageal cancer, a sample from a subject who does not suffer from esophageal precancerous lesions, or a sample from a subject with a cancer type of non-esophageal cancer or non-precancerous lesions. In some embodiments, the control sample includes a sample from a subject with non-dysplastic Barrett's esophagus (NDBE). In some embodiments, the control sample is from a tissue sample, a blood sample, a plasma sample, a serum sample, a whole blood sample, a buffy coat sample, a secretion sample, an organ secretion sample, a cerebrospinal fluid (CSF) sample, a saliva sample, a urine sample, and a stool sample. In some embodiments, the tissue sample is an esophageal tissue sample. In some embodiments, the esophageal sample is obtained from an esophageal biopsy, or obtained by wiping, brushing, or using a sponge capsule device.
[0119] In some embodiments, the present disclosure provides compositions and methods for identifying, determining and / or classifying esophageal cancer or precancerous lesions from biological samples (e.g., tissue samples, blood samples, plasma samples, serum samples, whole blood samples, buffy coat samples, secretion samples, organ secretion samples, cerebrospinal fluid (CSF) samples, saliva samples, urine samples and / or stool samples). The method generally includes determining the methylation profile of at least one methylation marker in a biological sample isolated from a subject. In some embodiments, a change in the methylation state or profile of the marker indicates the presence, category or location of esophageal cancer or precancerous lesions. Typically, such methods can be used to detect the presence or absence of esophageal cancer (e.g., esophageal adenocarcinoma (EAC)) or precancerous lesions (e.g., high-grade dysplastic Barrett's esophagus (HGD-BE)).
[0120] In some embodiments, a method is provided comprising contacting nucleic acid (e.g., genomic DNA) in a biological sample obtained from a subject with at least one reagent or a series of reagents that distinguishes between methylated and unmethylated nucleotides (e.g., CpG dinucleotides) within at least one methylation marker; and detecting the presence or absence of esophageal cancer or precancerous lesions (e.g., with a sensitivity greater than or equal to 80% and a specificity greater than or equal to 80%).
[0121] In some embodiments, methods are provided that include measuring one or both of the methylation levels of one or more genes or methylated DNA markers in a biological sample from a human individual by treating genomic DNA in the biological sample with an agent that modifies DNA in a methylation-specific manner; and determining the methylation levels of the one or more genes or methylation markers.
[0122] In some embodiments, a method is provided that includes: measuring the amount of one or more methylated DNA markers or genes in DNA from a biological sample; measuring the amount of at least one reference marker in the DNA; and calculating a percentage value of the amount of the at least one methylated marker gene measured in the DNA to the amount of the reference marker gene measured in the DNA, wherein the value represents the amount of the at least one methylated marker DNA measured in the biological sample.
[0123] In some embodiments, a method is provided comprising: measuring the methylation level of CpG sites of one or more genes in a biological sample from a human individual by treating genomic DNA in the biological sample with a bisulfite reagent that is capable of modifying DNA in a methylation-specific manner; amplifying the modified genomic DNA using a set of primers for the selected one or more genes; and determining the methylation level of the CpG sites of the selected one or more genes.
[0124] In some embodiments, the present disclosure provides a method for characterizing a biological sample, the method comprising measuring one or both of the methylation levels of CpG sites of one or more genes in a biological sample of a human individual by treating genomic DNA in the biological sample with bisulfite; amplifying the bisulfite-treated genomic DNA using a set of primers for the selected one or more genes; and determining the methylation levels of the CpG sites. In some embodiments, the method comprises comparing one or both of the methylation levels of the methylation markers with the methylation levels of a set of corresponding genes in a control sample that does not have a specific type of cancer; and / or determining that the subject has esophageal cancer or a precancerous lesion when one or both of the methylation levels measured in the one or more genes are higher than the methylation levels measured in the corresponding control samples.
[0125] In some embodiments, the present disclosure provides methods comprising: measuring one or both of the methylation levels of one or more genes or markers in a biological sample by treating genomic DNA in the biological sample with bisulfite; amplifying the bisulfite-treated genomic DNA using a set of primers for the selected gene or genes; and determining the methylation levels of the one or more genes or markers.
[0126] In some embodiments, the present disclosure provides methods for screening for esophageal cancer or precancerous lesions in a sample obtained from a subject. According to these embodiments, the method comprises one or both of the following: determining the methylation state or profile of one or more methylated DNA markers; and identifying the subject as having esophageal cancer or precancerous lesions when the methylation state or profile of the markers is different from the methylation state or profile of the markers determined in a subject who does not have esophageal cancer or precancerous lesions.
[0127] In some embodiments, the present disclosure provides methods comprising: measuring the methylation level of one or more genes or markers in a biological sample of a human individual by treating genomic DNA in the biological sample with a reagent that modifies DNA in a methylation-specific manner; amplifying the treated genomic DNA using a set of primers for the selected one or more genes or markers; and determining the methylation level of the one or more genes or markers.
[0128] In some embodiments, the present disclosure provides a method for characterizing a biological sample, the method comprising measuring the amount of at least one methylated DNA marker in DNA extracted from the biological sample; treating the genomic DNA in the biological sample with bisulfite; amplifying the bisulfite-treated genomic DNA using primers specific for the CpG site of each marker, wherein the primers specific for each marker are capable of binding to the amplicon bound by the primer sequence; and determining the methylation level of the CpG sites of one or more genes.
[0129] In some embodiments, the present disclosure provides a method comprising: extracting genomic DNA from a biological sample of a human individual suspected of having or having esophageal cancer or a precancerous lesion, measuring the methylation level of one or more methylated DNA markers in the DNA extracted from the biological sample; treating the extracted genomic DNA with bisulfite, amplifying the bisulfite-treated genomic DNA using primers specific for the one or more markers, wherein the primers specific for the one or more markers are capable of binding to at least a portion of the bisulfite-treated genomic DNA for a chromosomal region of a marker listed in Table 1 or 2; and measuring the methylation level of the one or more methylated markers.
[0130] In some embodiments, the present disclosure provides a method comprising: extracting genomic DNA from a biological sample of a human individual suspected of having or having esophageal cancer or a precancerous lesion, measuring the methylation level of one or more methylated DNA markers in the DNA extracted from the biological sample; treating the extracted genomic DNA with bisulfite, amplifying the bisulfite-treated genomic DNA using primers specific for the one or more markers, wherein the primers specific for the one or more markers are capable of binding to at least a portion of the bisulfite-treated genomic DNA for a chromosomal region of the markers listed in Table 1; and measuring the methylation level of the one or more methylation markers.
[0131] In some embodiments, the present disclosure provides a method comprising: extracting genomic DNA from a biological sample of a human individual suspected of having or having esophageal cancer or a precancerous lesion, measuring the methylation level of one or more methylated DNA markers in the DNA extracted from the biological sample; treating the extracted genomic DNA with bisulfite, amplifying the bisulfite-treated genomic DNA using primers specific for the one or more markers, wherein the primers specific for the one or more markers are capable of binding to at least a portion of the bisulfite-treated genomic DNA for a chromosomal region of the markers listed in Table 2; and measuring the methylation level of the one or more methylation markers.
[0132] In some embodiments, the present disclosure provides a method comprising: extracting genomic DNA from a biological sample of a human individual suspected of having or having esophageal cancer or a precancerous lesion, measuring the methylation level of one or more methylated DNA markers in the DNA extracted from the biological sample; treating the extracted genomic DNA with bisulfite, amplifying the bisulfite-treated genomic DNA using primers specific for the one or more markers, wherein the primers specific for the one or more markers are capable of binding to at least a portion of the bisulfite-treated genomic DNA for a chromosomal region of the markers listed in Table 3; and measuring the methylation level of the one or more methylation markers.
[0133] In some embodiments, the present disclosure provides methods comprising extracting genomic DNA from a biological sample of a human individual suspected of having or having cancer, treating the extracted genomic DNA with bisulfite, amplifying the bisulfite-treated genomic DNA using separate primers specific for CpG sites of one or more methylated DNA markers, and measuring the methylation level of the CpG sites for each of the one or more markers.
[0134] As further described herein, embodiments of the present disclosure include methods and compositions for characterizing a biological sample and determining a methylation profile in at least one differentially methylated region (DMR) of a DNA sample obtained from a subject having or suspected of having esophageal cancer or a precancerous lesion, by treating the sample with an agent that modifies DNA in a methylation-specific manner. In some embodiments, the method includes detecting the presence of esophageal cancer or a precancerous lesion from a biological sample. In some embodiments, at least one DMR is capable of distinguishing high-grade dysplastic Barrett's esophagus (HGD-BE) or esophageal adenocarcinoma (EAC) from non-dysplastic Barrett's esophagus (NDBE) or a normal esophageal control.
[0135] According to these embodiments, the method further comprises assessing a DNA sample from the subject for copy number variation (CNV) or copy number aberration (CNA). In some embodiments, the CNV can distinguish subjects having or suspected of having high-grade dysplasia Barrett's esophagus or esophageal adenocarcinoma (EAC) from a control DNA sample. In some embodiments, at least one DMR comprises an increased CNV compared to the control DNA sample.
[0136] In some embodiments, the method includes determining an aneuploidy score (AS) for a DNA sample from a subject. In some embodiments, AS can distinguish subjects suffering from or suspected of having high-grade dysplasia Barrett's esophagus or esophageal adenocarcinoma (EAC) from control DNA samples. In some embodiments, at least one DMR comprises an increased aneuploidy score compared to a control DNA sample. As described herein, an aneuploidy score refers to the total number of chromosome arms that have changed in a sample, and determining an aneuploidy score can provide an additional and / or alternative method for evaluating esophageal cancer or precancerous lesions in a DNA sample compared to a control.
[0137] In some embodiments, assessing CNV and / or determining aneuploidy scores can supplement methylation profile determination of DNA samples obtained from subjects suffering from or suspected of having esophageal cancer or precancerous lesions. In some embodiments, assessing CNV or determining aneuploidy scores involves obtaining sequencing reads from the same cytosine-converted genome used to assess the methylation profile of the DNA sample. While most current methods use direct NGS for unconverted DNA. However, in the methods of the present disclosure, methylation and CNV / aneuploidy reads are obtained simultaneously from the same chemistry / dataset.
[0138] Based on the present disclosure, it will be understood by those skilled in the art that the various methods described herein are not limited to the use of any one specific methylated DNA marker, methylated marker gene, methylated gene and / or DMR. That is, one or more of the methylated DNA markers, methylated marker genes, methylated genes and / or DMRs disclosed herein can be used to distinguish and / or identify esophageal cancer or precancerous lesions from controls, including any combination thereof. In addition, the methylated DNA markers, methylated marker genes, methylated genes and / or DMRs disclosed herein may include regions or subregions (e.g., genes on chromosomes, single nucleotides, CpG islands, etc.) of any marker listed in Tables 1, 2, and 3.
[0139] In some embodiments, the DMR is from a gene selected from the group consisting of: ACVRL1, ADAMTS8, ADAP2, ADRBK2, AKR1B1, ANK1, ANKRD13B, ANXA6, ARNT2, B4GALNT2, BACH2, BCL11A, BSCL2, C12orf53, C14orf82, C17orf107, C18orf1, C1orf95, C5orf42, CACNA1C, CAMK1D, CAMTA1, CBX6, CCDC85A, CCKBR, CD38, CDKN2A, CH25H, CHST1, CHST15, C NTLN, CRHR1, CRTC1, CXCR4, CYP1B1, DCTN2, DIDO1, DMKN, DSE, DYNC1I1, EML6, ENOX1, EPHA4, ESRRG, FAM176A, FAM78B, FBXO10, FERMT2, FHOD3, F LJ45079, FMNL1, FOXP2, FRMD4B, FZD8, GALNTL1, GALNTL4, GAS1, GBGT1, GLIPR2, GNAI1, GNAL, GPR37, GRASP, GRID1, GRM8, GSC, HAR1A, HCN2, HEY 2. HIST1H2BE, HS3ST3B1, HTR7, IGFBP2, INHA, INSRR, IRX3, ISM2, KCNG3, KCNK4, KCNS2, KCTD15, KIAA1522, KIAA1614, KIF26A, KL, KLF15, KLHL1 0. KRT77, KSR2, LBH, LMX1A, LMX1B, LOC100526820, LONRF2, LRFN2, LRRN1, MAF, MAFB, MARK1, ADAMTSL4-AS1, PGBD5, HSPA12A, SFTPD, LOC107984 507, LOC100128253, CISTR, SLC16A7, CTXND1, MAX.chr15.4912, ZNF423, RBFOX1, LOC105376772, GSE1, MAX.chr17.8070, ZNF709_6125, MAX.ch r19.3439, GREB1, CYP27C1, AC068134.6, LOC388780, STOX2_8441, PPARGC1A, MAX.chr4.4552, PDZD2, MAX.chr6.6743, LPAL2, CDK14, MAX.chr8.3003, MCOLN2, MEGF11, MFSD11, MRC2, MSX1, NAT8L, NAV1, NBEA, NCRNA00092, NECAB2, NEURL, NGEF, NHLH2, NKD1, NLGN1, NOG, NPPC, N R3C1, NRXN2, NTN1, OBSCN, OCA2, OSBPL1A, OXR1, P2RY1, PDE2A, PDE9A, PDGFRA, PID1, PLCG2, PLCL1, PNPLA3, POU3F1, POU3F2, PRKACB , PRKAR2B, PRKG1, PRR18, PRR5L, PTHLH, PYGL, RARG, RHBDL3, RIMS2, ROR2, RPRML, SDK2, SOX9, STOX2_6730, SYNGR1, TAC4, TNFRSF19, TRANK1, TSPAN33, TSPAN4, TSPAN5, UBE2E2, UCHL1, UNC5A, VASH2, VIM, WNT6, ZBTB10, ZNF680, ZNF709_4918, ZNF738, ZNF808 and ZNF845 (Table 1); and the subject has or is suspected of having esophageal cancer (e.g., esophageal adenocarcinoma (EAC)) or esophageal precursor lesions (e.g., high-grade dysplastic Barrett's esophagus (HGD-BE)). In some embodiments, determining the methylation profile of a DMR comprises comparing the methylation profile to a corresponding region of a control DNA sample (e.g., non-dysplastic Barrett's esophagus (NDBE) or a normal esophagus control).
[0140] In some embodiments, the DMR is from a gene selected from the group consisting of: ACVRL1, ADAMTS8, ADAP2, ADRBK2, AKR1B1, ANK1, ANKRD13B, ANXA6, ARNT2, B4GALNT2, BACH2, BCL11A, BSCL2, C14orf82, C18orf1, C1orf95, C5orf42, CAMK1D, CAMTA1, CCDC85A, CD38, CDKN 2A, CHST1, CHST15, CRHR1, CYP1B1, DIDO1, DSE, DYNC1I1, EML6, ENOX1, EPHA4, FAM176A, FBXO10, FERMT2, F HOD3, FLJ45079, FMNL1, FRMD4B, FZD8, GALNTL1, GALNTL4, GAS1, GBGT1, GLIPR2, GNAI1, GNAL, GRASP, GRID1 , GRM8, GSC, HAR1A, HCN2, HEY2, HIST1H2BE, HS3ST3B1, HTR7, IGFBP2, INHA, IRX3, ISM2, KCNG3, KCNK4, KCN S2, KCTD15, KIAA1614, KL, KLHL10, KSR2, LBH, LMX1A, LMX1B, LOC100526820, LONRF2, LRFN2, MAFB, MARK1, P GBD5, HSPA12A, SFTPD, LOC107984507, LOC100128253, CISTR, SLC16A7, CTXND1, ZNF423, LOC105376772, GS E1, MAX.chr19.3439, GREB1, CYP27C1, AC068134.6, LOC388780, STOX2_8441, PPARGC1A, PDZD2, MAX.chr6.6743, LPAL2, CDK14, MCOLN2, MEGF11, MRC2, MSX1, NAT8L, NAV1, NBEA, NCRNA00092, NEURL, NGEF, NHLH2, NKD1, NLGN1, NOG, NPPC, NR3C1, NRXN2, OBSCN, OCA2, OSBPL1A, OXR1, P2RY1, PDE2A, PDGFRA, PID1, PLCG2, PLCL1, PNPLA3, POU3F1, P OU3F2, PRKACB, PRKAR2B, PRKG1, PRR18, PRR5L, PTHLH, PYGL, RHBDL3, RIMS2, ROR2, RPRML, SDK2, SOX9, STOX2_6730, SYNGR1, TAC4, TRANK1, TSPAN4, UBE2E2, UCHL1, UNC5A, VASH2, ZBTB10, ZNF680, ZNF709_4918, ZNF738, ZNF808 and ZNF845 (Table 2); and the subject has or is suspected of having esophageal cancer (e.g., esophageal adenocarcinoma (EAC)) or esophageal precursor lesions (e.g., high-grade dysplastic Barrett's esophagus (HGD-BE)). In some embodiments, determining the methylation profile of a DMR comprises comparing the methylation profile to a corresponding region of a control DNA sample (e.g., non-dysplastic Barrett's esophagus (NDBE) or a normal esophagus control).
[0141] In some embodiments, the DMR is from a gene selected from the group consisting of BACH2, C5orf42, FHOD3, HIST1H2BE, IRX3, KIAA1614, LONRF2, MAFB, PDGFRA, PID1, POU3F1, PRR5L, RHBDL3, and SDK2 (Table 3); and the subject has or is suspected of having esophageal cancer (e.g., esophageal adenocarcinoma (EAC)) or esophageal precancerous lesions (e.g., high-grade dysplastic Barrett's esophagus (HGD-BE)). In some embodiments, determining the methylation profile of the DMR comprises comparing the methylation profile to a corresponding region of a control DNA sample (e.g., non-dysplastic Barrett's esophagus (NDBE) or a normal esophageal control).
[0142] In some embodiments, a novel DMR that can distinguish between esophageal cancer or a precancer lesion and a control sample is associated with an area under the ROC curve (AUC) greater than or equal to 0.5, wherein the ROC curve distinguishes subjects having or suspected of having esophageal cancer or a precancer lesion from a control DNA sample. In some embodiments, a novel DMR that can distinguish between esophageal cancer or a precancer lesion and a control sample is associated with an area under the ROC curve (AUC) greater than or equal to 0.6, wherein the ROC curve distinguishes subjects having or suspected of having esophageal cancer or a precancer lesion from a control DNA sample. In some embodiments, a novel DMR that can distinguish between esophageal cancer or a precancer lesion and a control sample is associated with an area under the ROC curve (AUC) greater than or equal to 0.7, wherein the ROC curve distinguishes subjects having or suspected of having esophageal cancer or a precancer lesion from a control DNA sample. In some embodiments, a novel DMR that can distinguish between esophageal cancer or a precancer lesion and a control sample is associated with an area under the ROC curve (AUC) greater than or equal to 0.75, wherein the ROC curve distinguishes subjects having or suspected of having esophageal cancer or a precancer lesion from a control DNA sample. In some embodiments, a novel DMR that can distinguish between esophageal cancer or a precancer lesion and a control sample is associated with an area under the ROC curve (AUC) greater than or equal to 0.8, wherein the ROC curve distinguishes subjects having or suspected of having esophageal cancer or a precancer lesion from a control DNA sample. In some embodiments, a novel DMR that can distinguish between esophageal cancer or a precancer lesion and a control sample is associated with an area under the ROC curve (AUC) greater than or equal to 0.85, wherein the ROC curve distinguishes subjects having or suspected of having esophageal cancer or a precancer lesion from a control DNA sample. In some embodiments, a novel DMR capable of distinguishing esophageal cancer or a precancerous lesion from a control sample is associated with an area under the ROC curve (AUC) greater than or equal to 0.9, wherein the ROC curve distinguishes subjects having or suspected of having esophageal cancer or a precancerous lesion from a control DNA sample. In some embodiments, a novel DMR capable of distinguishing esophageal cancer or a precancerous lesion from a control sample is associated with an area under the ROC curve (AUC) greater than or equal to 0.95, wherein the ROC curve distinguishes subjects having or suspected of having esophageal cancer or a precancerous lesion from a control DNA sample.
[0143] In some embodiments, the novel DMR capable of distinguishing esophageal cancer or precancerous lesions from control samples comprises an increased percentage of methylation compared to a control DNA sample. In some embodiments, the novel DMR capable of distinguishing esophageal cancer or precancerous lesions from control samples comprises an increased rate of hypermethylation compared to a control DNA sample.
[0144] In some embodiments, determining the methylation profile of at least one DMR includes amplifying at least a portion of the DMR using a set of primers. In some embodiments, determining the methylation profile of at least one DMR includes performing at least one of methylation-specific PCR, quantitative methylation-specific PCR, methylation-specific DNA restriction enzyme analysis, quantitative bisulfite pyrosequencing, flap endonuclease assay, PCR-flap assay, and bisulfite genomic sequencing PCR. In some embodiments, determining the methylation profile of at least one DMR includes determining the presence or absence of methylation at CpG sites. In some embodiments, one or more CpG sites are present in the coding region, non-coding region, and / or regulatory region of a gene (e.g., any one of the genes disclosed herein). In some embodiments, at least one of methylation-specific PCR, quantitative methylation-specific PCR, methylation-specific DNA restriction enzyme analysis, quantitative bisulfite pyrosequencing, flap endonuclease assay, PCR-flap assay, and bisulfite genomic sequencing PCR can be used to verify the DMR that can distinguish esophageal cancer or precancerous lesions from control samples. In some embodiments, a DMR capable of distinguishing esophageal cancer or precancerous lesions from control samples can be evaluated based on at least one of the area under the ROC curve (AUC), methylation fold change, methylation percentage, and / or hypermethylation ratio between the test sample and the control sample.
[0145] Those skilled in the art will appreciate based on this disclosure that esophageal cancer or precancerous lesions can be predicted by various marker combinations (e.g., as determined by statistical techniques related to the specificity and sensitivity of the prediction). Embodiments of the present disclosure provide methods for identifying predictive combinations for esophageal cancer or precancerous lesions and validating such predictive combinations.
[0146] Such methods are not limited to a particular manner or technique for determining, characterizing, measuring or assaying the methylation of one or more methylation markers, methylation marker genes, genes, DMRs and / or DNA methylation markers. In some embodiments, such techniques are based on analysis of the methylation status (e.g., CpG methylation status) of at least one marker comprising a DMR, a marked region, or a marked base.
[0147] In some embodiments, measuring the methylation state or profile of a methylated DNA marker in a sample comprises determining the methylation state of one nucleotide base. In some embodiments, measuring the methylation state of a methylated DNA marker in a sample comprises determining the degree of methylation at a plurality of nucleotide bases. Additionally, in some embodiments, the methylation state or profile of a methylated DNA marker comprises an increase in methylation of the marker relative to a normal methylation state or profile of the marker. In some embodiments, the methylation state or profile of the marker comprises a decrease in methylation of the marker relative to a normal methylation state of the marker. In some embodiments, the methylation state or profile of the marker comprises a different pattern of methylation of the marker relative to a normal methylation state or profile of the marker.
[0148] In addition, in some embodiments, the marker is a region of 100 or less nucleotide bases. In some embodiments, the marker is a region of 500 or less nucleotide bases. In some embodiments, the marker is a region of 1000 or less nucleotide bases. In some embodiments, the marker is a region of 5000 or less nucleotide bases. In some embodiments, the marker is a single nucleotide base. In some embodiments, the marker is in a high CpG density promoter region.
[0149] In certain embodiments, methods for analyzing the presence of 5-methylcytosine in nucleic acids involve treating the DNA with a reagent that modifies the DNA in a methylation-specific manner. Examples of such reagents include, but are not limited to, methylation-sensitive restriction enzymes, methylation-dependent restriction enzymes, bisulfite reagents, TET enzymes, and borane reducing agents.
[0150] A common method for analyzing the presence of 5-methylcytosine in nucleic acids is based on the bisulfite method described by Frommer et al. (Frommer et al. (1992) Proc. Natl. Acad. Sci. USA 89: 1827–31, incorporated herein by reference in its entirety for all purposes) or variants thereof. The bisulfite method for mapping 5-methylcytosine is based on the observation that cytosine (rather than 5-methylcytosine) reacts with bisulfite ions (also known as bisulfite). The reaction is generally carried out according to the following steps: first, cytosine reacts with bisulfite to form sulfonated cytosine. Next, spontaneous deamination of the sulfonated reaction intermediate produces sulfonated uracil. Finally, the sulfonated uracil is desulfonated under alkaline conditions to form uracil. Detection is possible because uracil pairs with adenine bases (thus behaving like thymine), while 5-methylcytosine pairs with guanine bases (thus behaving like cytosine). This makes it possible to distinguish methylated from unmethylated cytosines by, for example, bisulfite genomic sequencing (Grigg G and Clark S, Bioessays (1994) 16: 431–36; Grigg G, DNA Seq. (1996) 6: 189–98), methylation-specific PCR (MSP) (as disclosed in U.S. Pat. No. 5,786,146), or using assays involving sequence-specific probe cleavage, such as the QuARTS flap endonuclease assay (see, e.g., Zou et al. (2010) “Sensitive quantification of methylated markers with a novel methylationspecific technology” Clin Chem 56: A199; and U.S. Pat. Nos. 8,361,720; 8,715,937; 8,916,344; and 9,212,392).
[0151] In some embodiments, conventional techniques include methods that include encapsulating the DNA to be analyzed in an agarose matrix to prevent DNA diffusion and renaturation (bisulfite reacts only with single-stranded DNA) and replacing precipitation and purification steps with rapid dialysis (Olek A et al., (1996) "A modified and improved method forbisulfite based cytosine methylation analysis" Nucleic Acids Res. 24: 5064-6). This makes it possible to analyze the methylation status of individual cells, demonstrating the practicality and sensitivity of the method. Rein, T. et al., (1998) Nucleic Acids Res. 26: 2255 provides an overview of conventional methods for detecting 5-methylcytosine.
[0152] Bisulfite techniques typically involve amplifying short, specific fragments of a known nucleic acid after bisulfite treatment, followed by sequencing (Olek and Walter (1997) Nat. Genet. 17: 275–6) or assaying the products using primer extension reactions (Gonzalgo and Jones (1997) Nucleic Acids Res. 25: 2529–31; WO 95 / 00669; U.S. Patent No. 6,251,594) to analyze individual cytosine positions. Some methods use enzymatic digestion (Xiong and Laird (1997) Nucleic Acids Res. 25: 2532–4). Detection by hybridization has also been described in the art (Olek et al., WO 99 / 28498). In addition, the use of the bisulfite technique to detect methylation of individual genes has been described (Grigg and Clark (1994) Bioessays 16: 431-6; Zeschnigk et al. (1997) Hum Mol Genet. 6: 387-95; Feil et al. (1994) Nucleic Acids Res. 22: 695; Martin et al. (1995) Gene 157: 261-4; WO 9746705; WO 9515373).
[0153] According to embodiments of the present disclosure, various methylation assays can be used in conjunction with bisulfite treatment. These assays allow determination of the methylation status of one or more CpG dinucleotides (e.g., CpG islands) in a nucleic acid sequence. Such assays involve sequencing of bisulfite-treated nucleic acids, PCR (for sequence-specific amplification), Southern blot analysis, and the use of methylation-specific restriction enzymes (e.g., methylation-sensitive or methylation-dependent enzymes).
[0154] For example, genomic sequencing has been simplified by the use of bisulfite treatment to analyze methylation patterns and 5-methylcytosine distribution (Frommer et al. (1992) Proc. Natl. Acad. Sci. USA 89: 1827–1831). Additionally, restriction enzyme digestion of PCR products amplified from bisulfite-converted DNA can be used to assess methylation status, for example, as described in Sadri and Hornsby (1997) Nucl. Acids Res. 24: 5058–5059, or as embodied in a method known as COBRA (Combined Bisulfite Restriction Analysis) (Xiong and Laird (1997) Nucleic Acids Res. 25: 2532–2534).
[0155] COBRA™ analysis is a quantitative methylation assay that can be used to determine the DNA methylation level at a specific locus in a small amount of genomic DNA (Xiong and Laird, Nucleic Acids Res. 25:2532-2534, 1997). In short, restriction enzyme digestion is used to reveal methylation-dependent sequence differences in PCR products of sodium bisulfite-treated DNA. First, according to the procedure described by Frommer et al. (Proc. Natl. Acad. Sci. USA 89:1827-1831, 1992), methylation-dependent sequence differences are introduced into genomic DNA by standard bisulfite treatment. The bisulfite-converted DNA is then PCR amplified using primers specific for the CpG island of interest, followed by restriction enzyme digestion, gel electrophoresis, and detection using specific, labeled hybridization probes. The methylation level in the original DNA sample is represented by the relative amounts of digested and undigested PCR products in a linear quantitative manner over a wide range of DNA methylation levels. In addition, this technique can be reliably applied to DNA obtained from microdissected paraffin-embedded tissue samples.
[0156] Typical reagents for COBRA™ analysis (e.g., those found in typical COBRA™-based kits) may include, but are not limited to: PCR primers for a specific locus (e.g., a specific gene, marker, DMR, gene region, marker region, bisulfite-treated DNA sequence, CpG island, etc.); restriction enzymes and appropriate buffers; gene hybridization oligonucleotides; control hybridization oligonucleotides; kinase labeling kit for oligonucleotide probes; and labeled nucleotides. In addition, bisulfite conversion reagents may include DNA denaturation buffers; sulfonation buffers; DNA recovery reagents or kits (e.g., precipitation, ultrafiltration, affinity columns); desulfonation buffers; and DNA recovery components.
[0157] Assays such as "MethyLight™" (a fluorescence-based real-time PCR technology) (Eads et al., Cancer Res. 59:2302-2306, 1999), the Ms-SNuPE™ (methylation-sensitive single nucleotide primer extension) reaction (Gonzalgo and Jones, Nucleic Acids Res. 25:2529-2531, 1997), methylation-specific PCR ("MSP"; Herman et al., Proc. Natl. Acad. Sci. USA 93:9821-9826, 1996; U.S. Pat. No. 5,786,146), and methylated CpG island amplification ("MCA"; Toyota et al., Cancer Res. 59:2307-12, 1999) can be used alone or in combination with one or more of these methods.
[0158] The "HeavyMethyl™" assay is a quantitative method for assessing methylation differences based on methylation-specific amplification of bisulfite-treated DNA. Methylation-specific blocking probes ("blockers") that overlap CpG sites between, or are covered by, amplification primers enable methylation-specific, selective amplification of nucleic acid samples.
[0159] The term "HeavyMethyl™ MethyLight™" assay refers to a HeavyMethyl™ MethyLight™ assay, which is a variant of the MethyLight™ assay in which the MethyLight™ assay is combined with a methylation-specific blocking probe covering the CpG position between the amplification primers. The HeavyMethyl™ assay can also be used in combination with methylation-specific amplification primers.
[0160] Typical reagents for a HeavyMethyl™ assay (such as those found in a typical MethyLight™-based kit) may include, but are not limited to: PCR primers for a specific locus (e.g., a specific gene, marker, gene region, marker region, bisulfite-treated DNA sequence, CpG island or bisulfite-treated DNA sequence or CpG island, etc.); blocking oligonucleotides; optimized PCR buffers and deoxynucleotides; and Taq polymerase.
[0161] MSP (methylation-specific PCR) can assess the methylation status of almost any group of CpG sites within a CpG island without relying on the use of methylation-sensitive restriction enzymes (Herman et al. Proc. Natl. Acad. Sci. USA 93:9821-9826, 1996; U.S. Patent No. 5,786,146). Briefly, DNA is modified with sodium bisulfite to convert unmethylated cytosine (but not methylated cytosine) to uracil, and the product is then amplified using primers that are specific for methylated DNA compared to unmethylated DNA. MSP requires only a small amount of DNA, is sensitive to 0.1% methylated alleles of a given CpG island locus, and can be performed on DNA extracted from paraffin-embedded samples. Typical reagents for MSP analysis (such as those found in typical MSP-based kits) may include, but are not limited to, methylated and unmethylated PCR primers for specific loci (e.g., specific genes, markers, gene regions, marker regions, bisulfite-treated DNA sequences, CpG islands, etc.); optimized PCR buffers and deoxynucleotides, and specific probes.
[0162] The MethyLight™ assay is a high-throughput quantitative methylation assay that utilizes fluorescence-based real-time PCR (e.g., TaqMan®) with no further manipulation required after the PCR step (Eads et al., Cancer Res. 59:2302-2306, 1999). Briefly, the MethyLight™ process begins with a mixed sample of genomic DNA that is converted into a mixed pool of methylation-dependent sequence differences in a sodium bisulfite reaction according to standard procedures (the bisulfite process converts unmethylated cytosine residues into uracil). Fluorescence-based PCR is then performed with a "biased" reaction, for example using PCR primers that overlap with known CpG dinucleotides. Sequence discrimination occurs at the level of the amplification process and the level of the fluorescence detection process.
[0163] MethyLight ™ measures the methylation pattern in the quantitative test nucleic acid (such as genomic DNA sample), wherein sequence differentiation occurs at the probe hybridization level. In the quantitative version, the PCR reaction provides methylation-specific amplification in the presence of a fluorescent probe overlapping with a specific putative methylation site. Unbiased control of the input DNA amount is provided by the reaction that does not cover any CpG dinucleotide by primers and probes. Alternatively, qualitative testing of genomic methylation is achieved by detecting a biased PCR pool with a control oligonucleotide that does not cover known methylation sites (e.g., HeavyMethyl™ and MSP technology versions based on fluorescence) or an oligonucleotide that covers potential methylation sites.
[0164] The MethyLight™ process can be used with any suitable probe (e.g., TaqMan® probes, Lightcycler® probes, etc.). For example, in some applications, double-stranded genomic DNA is treated with sodium bisulfite, and one of two PCR reactions is performed using a TaqMan® probe, e.g., with MSP primers and / or Heavy Methyl blocking oligonucleotides and a TaqMan® probe. TaqMan® probes are dual-labeled with a fluorescent "reporter molecule" and a "quencher" molecule and are specifically designed for regions with relatively high GC content, resulting in a melting temperature approximately 10°C higher than that of the forward or reverse primers during PCR cycles. This allows the TaqMan® probe to remain fully hybridized during the PCR melt / extension step. As Taq polymerase enzymatically synthesizes a new strand during PCR, it eventually reaches the melted TaqMan® probe. The 5' to 3' endonuclease activity of the Taq polymerase then displaces the TaqMan® probe by digesting it, releasing the fluorescent reporter molecule for quantitative detection of its now unquenched signal using a real-time fluorescence detection system.
[0165] Typical reagents for MethyLight™ analysis (such as those found in typical MethyLight™-based kits) may include, but are not limited to: PCR primers for a specific locus (e.g., a specific gene, marker, gene region, marker region, bisulfite-treated DNA sequence, CpG island, etc.); TaqMan® or Lightcycler® probes; optimized PCR buffers and deoxynucleotides; and Taq polymerase.
[0166] QM ™ (quantitative methylation) assay is an alternative quantitative test for methylation patterns in genomic DNA samples, in which sequence differentiation occurs at the probe hybridization level. In this quantitative version, the PCR reaction provides unbiased amplification in the presence of a fluorescent probe that overlaps with a specific putative methylation site. Unbiased control of the amount of input DNA is provided by reactions in which neither primers nor probes cover any CpG dinucleotides. Alternatively, a qualitative test for genomic methylation is achieved by probing a biased PCR pool with a control oligonucleotide that does not cover known methylation sites (based on fluorescence HeavyMethyl ™ and MSP technology versions) or an oligonucleotide that covers potential methylation sites.
[0167] During amplification, the QM™ process can be used with any suitable probe, such as a "TaqMan®" probe or a Lightcycler® probe. For example, double-stranded genomic DNA is treated with sodium bisulfite and primed with unbiased primers and a TaqMan® probe. TaqMan® probes are dual-labeled with a fluorescent "reporter molecule" and a "quencher" molecule and are specifically designed for regions with relatively high GC content, resulting in a melting temperature approximately 10°C higher than that of the forward or reverse primer during PCR cycles. This allows the TaqMan® probe to remain fully hybridized during the PCR melt / extension step. As the Taq polymerase enzymatically synthesizes a new strand during PCR, it eventually reaches the melted TaqMan® probe. The Taq polymerase's 5' to 3' endonuclease activity then displaces the TaqMan® probe by digesting it, releasing the fluorescent reporter molecule for quantitative detection of its now unquenched signal using a real-time fluorescence detection system. Typical reagents for QM™ analysis (such as those found in a typical QM™-based kit) may include, but are not limited to: PCR primers for a specific locus (e.g., a specific gene, marker, gene region, marker region, bisulfite-treated DNA sequence, CpG island, etc.); TaqMan® or Lightcycler® probes; optimized PCR buffers and deoxynucleotides; and Taq polymerase.
[0168] Ms-SNuPE™ technology is a quantitative method for assessing differential methylation at specific CpG sites based on bisulfite treatment of DNA followed by single-nucleotide primer extension (Gonzalgo and Jones, Nucleic Acids Res. 25:2529-2531, 1997). Briefly, genomic DNA is reacted with sodium bisulfite to convert unmethylated cytosine to uracil while leaving 5-methylcytosine unchanged. PCR primers specific for bisulfite-converted DNA are then used to amplify the desired target sequence, and the resulting product is isolated and used as a template for methylation analysis at the CpG site of interest. Small amounts of DNA (e.g., microdissected pathology slides) can be analyzed, and the use of restriction enzymes to determine the methylation status of CpG sites is avoided.
[0169] Typical reagents for Ms-SNuPE™ analysis (e.g., those found in typical Ms-SNuPE™-based kits) may include, but are not limited to: PCR primers for a specific locus (e.g., a specific gene, marker, gene region, marker region, bisulfite-treated DNA sequence, CpG island, etc.); optimized PCR buffer and deoxynucleotides; gel extraction kit; positive control primers; Ms-SNuPE™ primers for a specific locus; reaction buffer (for the Ms-SNuPE reaction); and labeled nucleotides. In addition, bisulfite conversion reagents may include DNA denaturation buffer; sulfonation buffer; DNA recovery reagents or kits (e.g., precipitation, ultrafiltration, affinity columns); desulfonation buffer; and DNA recovery components.
[0170] Reduced representation bisulfite sequencing (RRBS) starts with bisulfite treatment of the nucleic acid, converts all unmethylated cytosines into uracil, and then performs restriction enzyme digestion (e.g., by an enzyme that recognizes sites including CG sequences, such as MspI), and completes sequencing of the fragments after coupling with an adapter ligand. The choice of restriction enzymes enriches fragments in CpG-dense regions, reducing the number of redundant sequences that may be mapped to multiple gene positions during the analysis process. Therefore, RRBS reduces the complexity of the nucleic acid sample by selecting a subset of restriction fragments (e.g., by size selection using preparative gel electrophoresis) for sequencing. In contrast to whole-genome bisulfite sequencing, each fragment produced by restriction enzyme digestion contains DNA methylation information for at least one CpG dinucleotide. Therefore, RRBS enriches the promoters, CpG islands, and other genomic features of the sample through high-frequency restriction enzyme sites in these regions, and therefore provides a measure for evaluating the methylation status of one or more genomic loci.
[0171] A typical protocol for RRBS includes digestion of nucleic acid samples with a restriction enzyme (such as Mspl), filling in overhangs and A-tails, ligating adapters, bisulfite conversion, and PCR. See, for example, Meissner et al. (2005) “Genome-scale DNA methylation mapping of clinical samples at single-nucleotide resolution” NatMethods 7: 133–6; Meissner et al. (2005) “Reduced representation bisulfite sequencing for comparative high-resolution DNA methylation analysis” NucleicAcids Res. 33: 5868–77.
[0172] In some embodiments, a quantitative allele-specific real-time target and signal amplification (QuARTS) assay is used to assess methylation status. In each QuARTS assay, three reactions occur sequentially, including amplification (reaction 1) and target probe cleavage (reaction 2) in the primary reaction; and FRET cleavage and fluorescence signal generation (reaction 3) in the secondary reaction. When the target nucleic acid is amplified with specific primers, a specific detection probe with a flap sequence loosely binds to the amplicon. The presence of a specific invasive oligonucleotide at the target binding site allows a 5' nuclease (e.g., FEN-1 endonuclease) to release the flap sequence by cutting between the detection probe and the flap sequence. The flap sequence is complementary to the non-hairpin portion of the corresponding FRET box. Therefore, the flap sequence acts as an invasive oligonucleotide on the FRET box and cleaves the FRET box fluorophore and quencher, thereby generating a fluorescent signal. The cleavage reaction can cut multiple probes for each target, and therefore each flap releases multiple fluorophores, providing exponential signal amplification. QuARTS can detect multiple targets in a single reaction well by using FRET boxes with different dyes. See, e.g., Zou et al. (2010) “Sensitive quantification of methylated markers with a novel methylation specific technology” Clin Chem 56: A199) and U.S. Pat. Nos. 8,361,720, 8,715,937, 8,916,344, and 9,212,392, each of which is incorporated herein by reference for all purposes.
[0173] The term "bisulfite reagent" refers to a reagent comprising bisulfite, disulfite, hydrogen sulfite, or a combination thereof, as disclosed herein, that can be used to distinguish methylated from unmethylated CpG dinucleotide sequences. The method of the treatment is known in the art (e.g., PCT / EP2004 / 011715 and WO 2013 / 116375, each of which is incorporated by reference in its entirety). In some embodiments, the bisulfite treatment is carried out in the presence of a denaturing solvent (e.g., but not limited to, n-alkyl glycol or diethylene glycol dimethyl ether (DME)), or in the presence of dioxane or a dioxane derivative. In some embodiments, the denaturing solvent is used at a concentration between 1% and 35% (v / v). In some embodiments, the bisulfite reaction is carried out in the presence of a scavenger, such as, but not limited to, a chroman derivative, for example, 6-hydroxy-2,5,7,8,-tetramethylchroman 2-carboxylic acid or trihydroxybenzoic acid and its derivatives, for example, gallic acid (see: PCT / EP2004 / 011715, which is incorporated by reference in its entirety). In certain preferred embodiments, the bisulfite reaction comprises treatment with ammonium bisulfite, for example, as described in WO 2013 / 116375.
[0174] In some embodiments, according to the methods and compositions described herein, primer oligonucleotide sets and amplification enzymes are used to amplify the fragments of the DNA processed. The amplification of several DNA segments can be carried out simultaneously in the same reaction vessel. Typically, polymerase chain reaction (PCR) is used to amplify. The length of the amplicon is generally 100 to 2000 base pairs.
[0175] In some embodiments of the method, the methylation status or profile of CpG positions within or near the differentially methylated regions (e.g., Tables 1, 2, and 3) can be detected by using methylation-specific primer oligonucleotides. This technology (MSP) has been described in U.S. Pat. No. 6,265,171 to Herman. Amplification of bisulfite-treated DNA using methylation-specific primers can distinguish between methylated and unmethylated nucleic acids. An MSP primer pair contains at least one primer that hybridizes to a bisulfite-treated CpG dinucleotide. Thus, the sequence of the primer contains at least one CpG dinucleotide. An MSP primer that is specific for unmethylated DNA contains a "T" at the C position of the CpG.
[0176] Such methods are not limited to specific types or classes of primers or primer pairs associated with one or more methylation markers, methylation marker genes, genes, DMRs, and / or methylated DNA markers. In some embodiments, the primers or primer pairs specific for each methylation marker gene are capable of binding to an amplicon bound by the primer sequence, wherein the amplicon bound by the primer sequence for the marker gene is at least a portion of a gene region for a methylation marker gene listed in Table 1, 2, or 3. In some embodiments, the primers or primer pairs for the methylation marker are a set of primers that specifically bind to at least a portion of a gene region comprising a methylation marker listed in Table 1, 2, or 3.
[0177] In another embodiment, the present disclosure provides a method for converting oxidized 5-methylcytosine residues in cell-free DNA into dihydrouracil residues (see Liu et al., 2019, Nat Biotechnol. 37, pp. 424-429; U.S. Patent Application Publication No. 202000370114). The method involves reacting an oxidized 5mC residue selected from 5-formylcytosine (5fC), 5-carboxymethylcytosine (5caC), and combinations thereof with a borane reducing agent. The oxidized 5mC residue can be naturally occurring or, more typically, is the result of prior oxidation of the 5mC or 5hmC residue, such as oxidation of 5mC or 5hmC with a TET family enzyme (e.g., TET1, TET2, or TET3), or chemical oxidation of 5mC or 5hmC, such as with potassium perruthenate (KRuO4) or inorganic peroxy compounds or compositions such as peroxytungstate (see, e.g., Okamoto et al. (2011) Chem. Commun. 47:11231-33) and copper(II) perchlorate / 2,2,6,6-tetramethylpiperidin-1-oxyl (TEMPO) combination (see Matsushita et al. (2017) Chem. Commun. 53:5756-59).
[0178] Borane reducing agents may be characterized by a complex of borane with a nitrogen-containing compound selected from nitrogen heterocycles and tertiary amines. The nitrogen heterocycle may be monocyclic, bicyclic or polycyclic, but is typically monocyclic in the form of a 5-membered or 6-membered ring containing a nitrogen heteroatom and optionally one or more additional heteroatoms selected from N, O and S. The nitrogen heterocycle may be aromatic or alicyclic. Preferred nitrogen heterocycles herein include 2-pyrroline, 2H-pyrrole, 1H-pyrrole, pyrazolidine, imidazolidine, 2-pyrazoline, 2-imidazoline, pyrazole, imidazole, 1,2,4-triazole, 1,2,4-triazole, pyridazine, pyrimidine, pyrazine, 1,2,4-triazine and 1,3,5-triazine, any of which may be unsubstituted or substituted with one or more non-hydrogen substituents. Typical non-hydrogen substituents are alkyl groups, particularly lower alkyl groups, such as methyl, ethyl, n-propyl, isopropyl, n-butyl, isobutyl, tert-butyl, etc. Exemplary compounds include pyridine borane, 2-methylpyridine borane (also known as 2-picoline borane), and 5-ethyl-2-pyridine.
[0179] The reaction of oxidized 5mC residues in cell-free DNA with borane reducing agents is advantageous because nontoxic reagents and mild reaction conditions can be employed; no bisulfate is required, nor are any other potentially DNA-degrading agents. Furthermore, the conversion of oxidized 5mC residues to dihydrouracil using borane reducing agents can be performed in a "one-pot" or "one-tube" reaction without the need to isolate any intermediates. This is significant because the conversion involves multiple steps, namely (1) reduction of the olefin bond connecting C-4 and C-5 in oxidized 5mC, (2) deamination, and (3) decarboxylation if the oxidized 5mC is 5caC, or deformylation if the oxidized 5mC is 5fC.
[0180] In addition to providing a method for converting oxidized 5-methylcytosine residues in cell-free DNA into dihydrouracil residues, the present disclosure also provides a reaction mixture related to the aforementioned method. The reaction mixture comprises a cell-free DNA sample containing at least one oxidized 5-methylcytosine residue selected from 5caC, 5fC, and a combination thereof, and a borane reducing agent that can effectively reduce, deaminize, decarboxylate, or deformylate the at least one oxidized 5-methylcytosine residue. The borane reducing agent is a complex of borane and a nitrogen-containing compound selected from nitrogen heterocycles and tertiary amines, as described above. In a preferred embodiment, the reaction mixture is substantially free of bisulfite, meaning substantially free of bisulfite ions and bisulfite. Ideally, the reaction mixture is free of bisulfite.
[0181] In a related aspect of the present disclosure, a kit for converting 5mC residues in cell-free DNA to dihydrouracil residues is provided, wherein the kit includes a reagent for blocking 5hmC residues, a reagent for oxidizing 5mC residues beyond hydroxymethylation to provide oxidized 5mC residues, and a borane reducing agent effective for reducing, deaminating, and decarboxylating or deformylating the oxidized 5mC residues. The kit may also include instructions for using the components to perform the above method.
[0182] In another embodiment, a method utilizing the above-described oxidation reaction is provided. The method is capable of detecting the presence and location of 5-methylcytosine residues in cell-free DNA and comprises the following steps: (a) modifying 5hmC residues in fragmented, adaptor-ligated cell-free DNA to provide an affinity tag thereon, wherein the affinity tag is capable of removing DNA containing modified 5hmC from the cell-free DNA; (b) removing DNA containing modified 5hmC from the cell-free DNA, leaving DNA containing unmodified 5mC residues; (c) oxidizing the unmodified 5mC residues to provide DNA containing oxidized 5mC residues selected from 5caC, 5fC, and combinations thereof; (d) contacting the DNA containing the oxidized 5mC residues with a borane reducing agent that is effective to reduce, deaminize, decarboxylate, or deformylate the oxidized 5mC residues, thereby providing DNA containing dihydrouracil residues in place of the oxidized 5mC residues; (e) amplifying and sequencing the DNA containing the dihydrouracil residues; and (f) determining the 5-methylation pattern based on the sequencing results in (e).
[0183] In some embodiments, the present disclosure provides a method for identifying 5-methylcytosine (5mC) or 5-hydroxymethylcytosine (5hmC) in a target nucleic acid. In some embodiments, the method includes providing a biological sample comprising a target nucleic acid, modifying the target nucleic acid by converting 5mC and 5hmC in the nucleic acid sample into 5-carboxylcytosine (5caC) and / or 5-formylcytosine (5fC) by contacting the nucleic acid sample with a TET enzyme, thereby generating one or more 5caC or 5fC residues, and converting 5caC and / or 5fC into dihydrouracil (DHU) by treating the target nucleic acid with a borane reducing agent to provide a modified nucleic acid sample comprising a modified target nucleic acid, and detecting the sequence of the modified target nucleic acid; wherein the conversion of cytosine (C) to thymine (T) or the conversion of cytosine (C) to DHU in the modified target nucleic acid sequence compared to the target nucleic acid provides the position of 5mC or 5hmC in the target nucleic acid. In some embodiments, the borane reducing agent is 2-methylpyridine borane.
[0184] In some embodiments, detecting the sequence of the modified target nucleic acid comprises one or more of chain termination sequencing, microarray, high throughput sequencing, and restriction enzyme analysis. In some embodiments, the TET enzyme is selected from the group consisting of human TET1, TET2, and TET3; murine TET1, TET2, and TET3; Naegleria TET (NgTET); and Coprinopsis cinerea (CcTET). In some embodiments, the method further comprises a step of blocking one or more modified cytosines. In some embodiments, the blocking step comprises adding a sugar to 5hmC. In some embodiments, the method further comprises a step of amplifying the number of copies of one or more nucleic acid sequences. In some embodiments, the oxidant is potassium perruthenate or Cu(II) / TEMPO (2,2,6,6-tetramethylpiperidine-1-oxyl).
[0185] Cell-free DNA is typically extracted from a biological sample from a subject, wherein the sample can be whole blood, buffy coat, plasma, urine, saliva, mucosal secretions, organ secretions, sputum, feces or tears. In some embodiments, the cell-free DNA is derived from a tumor (e.g., an esophageal tumor). In other embodiments, the cell-free DNA is from a patient suffering from a disease or other pathogenic condition. The cell-free DNA may or may not be derived from a tumor. In some embodiments, the cell-free DNA to be modified with 5hmC residues is in a purified, fragmented form and is adapter-ligated. DNA purification in this context can be performed using any suitable method known to those of ordinary skill in the art and / or described in the relevant literature, and although the cell-free DNA itself can be highly fragmented, further fragmentation may sometimes be required, such as described in U.S. Patent Publication No. 2017 / 0253924. The size of the cell-free DNA fragments is generally in the range of about 20 nucleotides to about 500 nucleotides, more typically in the range of about 20 nucleotides to about 250 nucleotides. The purified cell-free DNA fragments modified in step (a) have been end-repaired using conventional methods (e.g., restriction enzymes) so that the fragments have blunt ends at each 3' and 5' end. In a preferred method, as described in WO 2017 / 176630, a polymerase (such as Taq polymerase) is also used to provide the blunted fragments with 3' overhangs containing a single adenine residue. This facilitates the subsequent connection of selected universal adapters, i.e., adapters that are connected to both ends of the cell-free DNA fragments and contain at least one molecular barcode, such as Y-type adapters or hairpin adapters. The use of adapters can also achieve selective PCR enrichment of adapter-connected DNA fragments.
[0186] In some embodiments, "purified, fragmented cell-free DNA" includes adaptor-ligated DNA fragments. 5hmC residues in these cell-free DNA fragments are modified with an affinity tag so that the DNA containing modified 5hmC can be subsequently removed from the cell-free DNA. In one embodiment, the affinity tag comprises a biotin moiety, such as biotin, desthiobiotin, oxybiotin, 2-iminobiotin, diaminobiotin, biotin sulfoxide, biocytin, etc. The use of a biotin moiety as an affinity tag allows for easy removal using streptavidin (e.g., streptavidin beads, magnetic streptavidin beads, etc.).
[0187] Labeling of 5hmC residues with a biotin moiety or other affinity tag is accomplished by covalently attaching a chemoselective group to the 5hmC residues in the DNA fragment, wherein the chemoselective group is capable of reacting with the functionalized affinity tag, thereby attaching the affinity tag to the 5hmC residue. In one embodiment, the chemoselective group is UDP glucose-6-azide, which undergoes a spontaneous 1,3-cycloaddition reaction with an alkyne-functionalized biotin moiety, as described in Robertson et al. (2011) Biochem. Biophys. Res. Comm. 411(1):40-3, U.S. Patent No. 8,741,567, and WO2017 / 176630. Thus, addition of the alkyne-functionalized biotin moiety results in a covalent attachment of the biotin moiety to each 5hmC residue.
[0188] In one embodiment, the affinity-tagged DNA fragments can then be pulled down using streptavidin in the form of streptavidin beads, magnetic streptavidin beads, etc., and, if desired, retained for later analysis. The supernatant remaining after removal of the affinity-tagged fragments contains DNA with unmodified 5mC residues and no 5hmC residues.
[0189] In some embodiments, the unmodified 5mC residue is oxidized using any suitable method to provide a 5caC residue and / or a 5fC residue. An oxidizing agent is selected to oxidize the 5mC residue beyond hydroxymethylation, i.e., to provide a 5caC and / or 5fC residue. Oxidation can be performed enzymatically using a catalytically active TET family enzyme. The term "TET family enzyme" or "TET enzyme" as used herein refers to the catalytically active "TET family protein" or "TET catalytically active fragment" defined in U.S. Patent No. 9,115,386, the disclosure of which is incorporated herein by reference. In this case, the preferred TET enzyme is TET2; see Ito et al. (2011) Science 333(6047):1300-1303. Oxidation can also be performed chemically using a chemical oxidant, as described in the previous section. Examples of suitable oxidizing agents include, but are not limited to, perruthenate anions in the form of inorganic or organic perruthenates, including metal perruthenates such as potassium perruthenate (KRuO4), tetraalkylammonium perruthenates such as tetrapropylammonium perruthenate (TPAP) and tetrabutylammonium perruthenate (TBAP), and polymer-supported perruthenates (PSP); and inorganic peroxy compounds and compositions such as peroxytungstate or copper (II) perchlorate / TEMPO combinations. There is no need to separate the 5fC-containing fragment from the 5caC-containing fragment at this point, as both the 5fC residue and the 5caC residue are converted to dihydrouracil (DHU) in the next step of the process.
[0190] In some embodiments, 5-hydroxymethylcytosine residues are blocked with β-glucosyltransferase (β3GT), while 5-methylcytosine residues are oxidized with a TET enzyme that effectively provides a mixture of 5-formylcytosine and 5-carboxymethylcytosine. The mixture containing these two oxidizing species can be reacted with 2-pyridine borane or another borane reducing agent to obtain dihydrouracil. In a variant of this embodiment, fragments containing 5hmC are not removed. Instead, "TET-assisted methylpyridine borane sequencing (TAPS)" enzymatically oxidizes fragments containing 5mC and fragments containing 5hmC together to provide fragments containing 5fC and 5caC. The reaction with 2-methylpyridine borane produces DHU residues where the 5mC and 5hmC residues were originally present. "Chemically assisted methylpyridine borane sequencing (CAPS)" involves selectively oxidizing fragments containing 5hmC with potassium perruthenate, leaving the 5mC residues unchanged.
[0191] As disclosed in International PCT Application PCT / US2019 / 012627, which is incorporated herein by reference in its entirety, TAPS includes the use of mild enzymatic and chemical reactions to directly and quantitatively detect 5mC and 5hmC at base resolution without affecting unmodified cytosine. In a related embodiment, the above method also includes identifying the hydroxymethylation pattern in DNA containing 5hmC removed from cell-free DNA. This can be performed using the technology described in detail in WO 2017 / 176630. The process can be performed in a single tube method without removing or separating intermediates. For example, first, cell-free DNA fragments (preferably adapter-connected DNA fragments) are functionalized with βGT-catalyzed uridine diphosphate glucose 6-azide and then biotinylated by chemically selective azide groups. This procedure produces covalently linked biotin at each 5hmC site. In the next step, the biotinylated chain and the chain containing unmodified (natural) 5mC are pulled down simultaneously for further processing. As is known in the art, the native 5mC-containing chain is pulled down using an anti-5mC antibody or a methyl-CpG binding domain (MBD) protein. Then, in the presence of blocking 5hmC residues, unmodified 5mC residues are selectively oxidized using any suitable technique to convert 5mC to 5fC and / or 5caC, as described elsewhere herein.
[0192] The fragment obtained by amplification can be with direct or indirect detectable label.In some embodiments, label is fluorescent label, radionuclide or has the separable molecular fragment of the typical mass that can be detected in mass spectrometer.When described label is mass label, the amplicon of some embodiments regulation mark has single positive or negative net charge, thereby allows to have better detectability in mass spectrometer.Can detect and visualize by for example matrix-assisted laser desorption / ionization mass spectrometry (MALDI) or use electrospray mass spectrometry (ESI).
[0193] Methods for isolating DNA suitable for these assay techniques are known in the art. Specifically, some embodiments include nucleic acid isolation as described in U.S. patent application Ser. No. 13 / 470,251 ("nucleic acid isolation"), which is incorporated herein by reference in its entirety.
[0194] In some embodiments, mark as described herein can be used for the QUARTS determination that stool sample is carried out.In some embodiments, provide for the method for producing DNA sample, particularly for the method for producing the DNA sample of the highly purified low abundance nucleic acid comprising small volume (for example less than 100 microlitres, less than 60 microlitres) and substantially and / or effectively do not contain the material of the mensuration (for example PCR, INVADER, QuARTS determination etc.) that suppresses for testing DNA sample.Such DNA sample can be used for diagnostic determination, and its qualitative detection is taken from the gene, gene variant (for example allele) or genetic modification (for example methylation) present in the sample of patient, or quantitatively measures its activity, expression or amount.For example, some cancers are relevant to the existence of specific mutant allele or specific methylation state, and therefore detection and / or quantification of such mutant allele or methylation state has predictive value in the diagnosis and treatment of cancer.
[0195] Many valuable genetic markers are present at very low levels in samples, and many events that produce such markers are rare. Therefore, even sensitive detection methods such as PCR require large amounts of DNA to provide enough low-abundance targets to meet or replace the detection threshold of the assay. In addition, even the presence of small amounts of inhibitory substances can compromise the accuracy and precision of these assays for detecting such low-amount targets. Therefore, this article provides methods for providing the necessary management of volume and concentration to produce such DNA samples.
[0196] In some embodiments, the biological sample is a tissue sample, a blood sample, a plasma sample, a serum sample, a whole blood sample, a buffy coat sample, a secretion sample, an organ secretion sample, a cerebrospinal fluid (CSF) sample, a saliva sample, a urine sample, and / or a stool sample. In some embodiments, the tissue sample is an esophageal tissue sample. In some embodiments, the tissue sample is an endoscopic esophageal brushing sample. In some embodiments, the sample includes esophageal tissue, and / or the sample is obtained by endoscopic brushing or non-endoscopic whole esophageal brushing or swabbing using a tethered device (e.g., a capsule sponge, balloon, or other device).
[0197] In some embodiments, the sample comprises tissue and / or biological fluid obtained from a human patient. In some embodiments, the sample comprises esophageal tissue. In some embodiments, the sample comprises esophageal tissue obtained by swabbing or brushing the entire esophagus. In some embodiments, the sample comprises secretions. In some embodiments, the sample comprises blood, serum, plasma, gastric secretions, pancreatic juice, gastrointestinal biopsy samples, microdissected cells from esophageal biopsy, esophageal cells shed into the gastrointestinal lumen, and / or esophageal cells recovered from feces. These samples can be derived from the upper digestive tract, the lower digestive tract, or comprise cells, tissues, and / or secretions from the upper and lower digestive tracts. In some embodiments, the sample comprises cell fluid, ascites, urine, feces, pancreatic juice, fluid obtained during endoscopy, blood, mucus, or saliva. In some embodiments, the sample is a fecal sample. Such samples can be obtained by a variety of methods known in the art, such as methods obvious to those skilled in the art. For example, urine and fecal samples are easily obtained, while blood, ascites, serum, or pancreatic juice samples can be obtained parenterally using, for example, a needle and syringe. Cell-free or substantially cell-free samples can be obtained by subjecting the sample to various techniques known to those skilled in the art, including but not limited to centrifugation and filtration. Although it is generally preferred not to use invasive techniques to obtain the sample, it may still be preferred to obtain samples such as tissue homogenates, tissue sections, and biopsy specimens. In some embodiments, the sample is obtained by esophageal swabbing or brushing or using a sponge capsule device.
[0198] Such sample can be obtained by multiple methods known in the art, such as methods apparent to those skilled in the art.Can be by making sample stand various techniques known to those skilled in the art, including but not limited to centrifugation and filtration to obtain acellular or substantially acellular sample.Although it is usually preferred not to use invasive techniques to obtain sample, it is still possible to be preferred to obtain sample such as tissue homogenate, tissue section and biopsy specimen.Described technology is not limited by the method for preparing sample and providing the nucleic acid for test.For example, in some embodiments, DNA is separated from sample (such as tissue sample, blood sample, plasma sample, serum sample, whole blood sample, buffy coat sample, secretion sample, organ secretion sample, cerebrospinal fluid (CSF) sample, saliva sample, urine sample and / or fecal sample) using direct gene capture, for example, as described in detail in U.S. Patent No. 8,808,990 and No. 9,169,511 and WO 2012 / 155072, or by related methods.
[0199] The analysis of mark can be carried out separately, or carry out simultaneously with other marks in a test sample.For example, several marks can be combined into a kind of test, to effectively process multiple samples, and may provide higher diagnosis and / or prognosis accuracy.In addition, those skilled in the art will recognize the value of testing multiple samples (for example, at continuous time points) from the same experimenter.Serial samples are carried out this type of test and can identify the change of mark methylation state over time.The change of methylation state and the no change of methylation state can provide useful information about disease state, include but not limited to the approximate time that identification event starts, the existence and amount of salvageable tissue, the suitability of drug therapy, the effectiveness of various therapies, and the result of identifying experimenter, including the risk of future events.
[0200] Analysis of biomarkers can be performed in a variety of physical formats. For example, the use of microtiter plates or automation can be used to facilitate the processing of large numbers of test samples. Alternatively, a single sample format can be developed to facilitate immediate treatment and diagnosis in a timely manner, such as in ambulatory transport or emergency room settings.
[0201] Genomic DNA can be separated by any means, including the use of commercially available test kits. In short, when the DNA of interest is wrapped by the cell membrane, the biological sample must be destroyed and dissolved by enzymatic, chemical or mechanical means. Proteins and other contaminants can then be removed from the DNA solution, for example, by digestion with Proteinase K. Genomic DNA is then recovered from the solution. This can be achieved by a variety of methods, including salting out, organic extraction or combining the DNA with a solid support. The choice of method will be affected by several factors, including time, cost and the required amount of DNA. All clinical sample types containing neoplastic material or pre-neoplastic material are suitable for the method of the present invention, such as cell lines, tissue sections, biopsies, endoscopic brushing samples, paraffin-embedded tissues, body fluids, feces, tissues, colon effluent, urine, plasma, serum, whole blood, buffy coat, separated blood cells, cells separated from blood, and combinations thereof.
[0202] The technology is not limited by the method used to prepare the sample and provide the nucleic acid for testing. For example, in some embodiments, DNA is isolated from a stool sample or from a blood sample or from a plasma sample using direct gene capture, e.g., as described in detail in U.S. patent application Ser. No. 61 / 485,386, or by related methods.
[0203] The genomic DNA sample is then treated with at least one reagent or series of reagents that distinguishes between methylated and unmethylated CpG dinucleotides within at least one marker comprising a DMR (eg, a DMR in Table 1, 2, or 3).
[0204] In some embodiments, the reagent converts an unmethylated cytosine base at the 5'-position to uracil, thymine, or another base that differs from cytosine in hybridization behavior. However, in some embodiments, the reagent may be a methylation-sensitive restriction enzyme.
[0205] In some embodiments, the genomic DNA sample is treated in a manner such that unmethylated cytosine bases at the 5'-position are converted to uracil, thymine, or another base that differs from cytosine in hybridization behavior. In some embodiments, this treatment is performed with a bisulfite (e.g., hydrogen sulfite, disulfite) followed by alkaline hydrolysis.
[0206] The processed nucleic acid is then analyzed to determine the methylation status of the target gene sequence (at least one gene, genomic sequence, or nucleotide from a marker comprising a DMR, such as at least one DMR selected from Tables 1, 2, or 3). The analysis method can be selected from those known in the art, including those listed herein, such as QuARTS and MSP described herein.
[0207] Such samples can be obtained by a variety of methods known in the art, such as methods that will be apparent to those skilled in the art. For example, urine and fecal samples are readily available, while blood, ascites, serum, or pancreatic juice samples can be obtained parenterally using, for example, a needle and syringe. Cell-free or substantially cell-free samples can be obtained by subjecting the sample to various techniques known to those skilled in the art, including, but not limited to, centrifugation and filtration. Although it is generally preferred not to use invasive techniques to obtain samples, it may still be preferred to obtain samples such as tissue homogenates, tissue sections, and biopsy specimens.
[0208] 3. Treatment
[0209] In some embodiments, the present disclosure provides a method for treating a subject (e.g., a patient suffering from or suspected of having esophageal cancer or precancerous lesions). According to these embodiments, the method includes determining the methylation state or profile of one or more methylated DNA markers provided herein, and administering treatment to the patient based on the result of determining the methylation state. In some embodiments, the method includes assessing copy number variation (CNV) or copy number aberration (CNA) of a DNA sample from a subject, and treating the patient according to the assessment results. In some embodiments, the method includes assessing an aneuploidy score (AS) in a DNA sample from a subject, and treating the patient according to the assessment results.
[0210] Treatment can be administering a pharmaceutical compound, a vaccine, performing surgery, imaging the patient, performing another test. In some embodiments, treating a subject includes methods for clinical screening, methods for prognostic assessment, methods for monitoring the outcome of therapy, methods for identifying patients most likely to respond to a particular therapeutic treatment, methods for imaging a patient or subject, and methods for drug screening and development.
[0211] In some embodiments, a method for diagnosing a particular type of cancer in a subject is provided. As used herein, the terms "diagnosis (diagnosing)" and "diagnosis (diagnosis)" refer to methods by which a skilled person can estimate and even determine whether a subject suffers from a given disease or illness or whether a given disease or illness may develop in the future. Skilled persons typically diagnose based on one or more diagnostic indices, such as one or more biomarkers (e.g., one or more methylation markers, methylation marker genes, genes, DMRs and / or DNA methylation markers as disclosed herein), whose methylation states indicate the presence, severity or absence of an illness. In some embodiments, diagnosis can be performed according to one or more diagnostic indices (e.g., CNV or aneuploidy), which can indicate the presence, severity or absence of an illness (e.g., esophageal cancer or precancerous lesions).
[0212] Along with diagnosis, clinical cancer prognosis involves determining the aggressiveness of a cancer and the likelihood of tumor recurrence in order to plan the most effective therapy. If a more accurate prognosis could be made, or even the potential risk of developing cancer could be assessed, appropriate therapy could be selected for the patient, and in some cases, a less severe therapy could be selected. Assessment of cancer biomarkers (e.g., determining methylation profiles and / or aneuploidy scores) can be used to distinguish subjects with a good prognosis and / or a low risk of developing cancer who will not require treatment or will require limited treatment from those subjects who are more likely to develop cancer or suffer a recurrence of cancer who may benefit from more intensive treatment.
[0213] Thus, "making a diagnosis" or "diagnosing" as used herein also includes determining the risk of developing a cancer or determining a prognosis based on measurements of the diagnostic biomarkers (e.g., DMRs) and / or aneuploidy scores disclosed herein, which can provide a prediction of clinical outcome (with or without medical treatment), selection of an appropriate treatment (or whether a treatment is effective), or monitoring a current treatment and potentially changing treatment. In addition, in some embodiments of the presently disclosed subject matter, multiple measurements of the biomarkers may be made over time to facilitate diagnosis and / or prognosis. Temporal changes in methylation profiles and / or aneuploidy scores can be used to predict clinical outcome, monitor the progression of a cancer or a subclass of cancer, and / or monitor the efficacy of appropriate therapies for the cancer. For example, in such embodiments, it may be desirable to see changes in the methylation state of one or more biomarkers (e.g., DMRs) and / or aneuploidy scores disclosed herein (and potentially one or more additional biomarkers, if monitored).
[0214] In some embodiments, the presently disclosed subject matter also provides a method for determining whether to initiate or continue prevention or treatment of cancer in a subject. In some embodiments, the method comprises providing a series of biological samples from a subject over a period of time; analyzing the series of biological samples to determine the methylation profile and / or aneuploidy score as disclosed herein in each biological sample; and comparing any measurable changes in the methylation profile and / or aneuploidy score in each biological sample. Any changes over a period of time can be used to predict the risk of developing cancer, predict clinical outcomes, determine whether to initiate or continue prevention or treatment of cancer, and whether current therapies are effectively treating cancer. For example, a first time point can be selected before the start of treatment, and a second time point can be selected at some time after the start of treatment. The methylation profile and / or aneuploidy score can be measured in each sample taken at different time points, and qualitative and / or quantitative differences can be recorded. Changes in the methylation status and / or aneuploidy score of biomarker levels from different samples can be correlated with a particular cancer risk, prognosis, determination of treatment efficacy, and / or cancer progression in the subject. In some embodiments, the methods and compositions of the present disclosure are used to treat or diagnose a disease at an early stage, such as before disease symptoms appear. In some embodiments, the methods and compositions of the present disclosure are used to treat or diagnose a disease at a clinical stage.
[0215] In some embodiments, multiple measurements of one or more diagnostic or prognostic biomarkers may be performed, and changes in the markers over time may be used to determine a diagnosis or prognosis. For example, a diagnostic marker may be determined at an initial time and again at a second time. In such embodiments, an increase in the marker from the initial time to the second time may be diagnostic of a particular type or severity of cancer, or a given prognosis. Similarly, a decrease in the marker from the initial time to the second time may be indicative of a particular type or severity of cancer, or a given prognosis. Furthermore, the degree of change in one or more markers may correlate with the severity of the cancer and future adverse events. Those skilled in the art will appreciate that while in certain embodiments, comparative measurements of the same biomarker may be performed at multiple time points, it is also possible to measure a given biomarker at one time point and a second biomarker at a second time point, and to compare these markers to provide diagnostic information.
[0216] As used herein, the phrase "determining a prognosis" refers to a method by which one skilled in the art can predict the course or outcome of a condition in a subject. The term "prognosis" does not mean that the course or outcome of a condition can be predicted with 100% accuracy, nor does it even mean that the methylation status and / or aneuploidy score of a biomarker (e.g., DMR) can predict whether a given condition or outcome is more or less likely to occur. Instead, one skilled in the art will understand that the term "prognosis" refers to an increased likelihood of a process or outcome occurring; that is, a process or outcome is more likely to occur in a subject who exhibits the condition when compared to those individuals who do not exhibit the given condition. For example, in an individual who does not exhibit the condition (e.g., has a normal methylation status and / or a normal aneuploidy score for one or more DMRs), the chance of a given outcome (e.g., developing a particular type of cancer) may be very low.
[0217] In some embodiments, statistical analysis associates prognostic indicators with the tendency of unfavorable outcomes. For example, in some embodiments, the methylation status different from the methylation profile and / or aneuploidy score in the normal control sample obtained from a patient who has never suffered from cancer can show that, compared with the subject with the level of methylation profile and / or aneuploidy score more similar to the control sample, the subject is more likely to suffer from cancer, as determined by statistical significance level. In addition, the change of methylation profile and / or aneuploidy score relative to baseline (for example, "normal") level can reflect the prognosis of the subject, and the degree of change of methylation profile and / or aneuploidy score can be related to the severity of adverse events. Statistical significance is generally determined by comparing two or more colonies and determining confidence interval and / or p value. See, for example, Dowdy and Wearden, Statistics for Research, John Wiley & Sons, New York, 1983, which is incorporated herein by reference in its entirety. Exemplary confidence intervals for the present subject matter are 90%, 95%, 97.5%, 98%, 99%, 99.5%, 99.9%, and 99.99%, while exemplary p-values are 0.1, 0.05, 0.025, 0.02, 0.01, 0.005, 0.001, and 0.0001.
[0218] In other embodiments, a threshold degree of change in the methylation state and / or aneuploidy score of a prognostic or diagnostic biomarker (e.g., a DMR) disclosed herein can be established, and the degree of change in the methylation profile and / or aneuploidy score in the biological sample can be simply compared to the threshold degree of change in the methylation profile and / or aneuploidy score. For example, preferred threshold changes in the methylation state of the biomarkers provided herein are about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 50%, about 75%, about 100%, and about 150%. In other embodiments, a "nomogram" can be established by which the methylation state of a prognostic or diagnostic indicator (biomarker or combination of biomarkers) is directly correlated with the associated tendency of a given outcome. Those skilled in the art are familiar with the use of such nomograms to relate two values and understand that the uncertainty in this measurement is the same as the uncertainty in the marker concentration because reference is made to a single sample measurement, not a population average.
[0219] In some embodiments, the control sample is analyzed simultaneously with the biological sample so that the results obtained from the biological sample can be compared with the results obtained from the control sample. In addition, it is contemplated that a standard curve can be provided, which can be compared with the measurement results of the biological sample. If a fluorescent label is used, such a standard curve presents the methylation state and / or aneuploidy score of the biomarker as a function of the measurement unit, such as the fluorescence signal intensity. Using samples collected from multiple donors, a standard curve can be provided to control the methylation state and / or aneuploidy score in normal tissue. In certain embodiments of the method, after identifying the abnormal methylation state and / or abnormal aneuploidy score of one or more DMRs provided herein in a biological sample obtained from the subject, the subject is identified as having cancer. In other embodiments of the method, detecting the abnormal methylation state and / or abnormal aneuploidy score of one or more such biomarkers in a biological sample obtained from the subject results in the subject being identified as having cancer.
[0220] The analysis of mark can be carried out separately, or carry out simultaneously with other marks in a test sample.For example, several marks can be combined into a kind of test, to effectively process multiple samples, and may provide higher diagnosis and / or prognosis accuracy.In addition, those skilled in the art will recognize the value of testing multiple samples (for example, at continuous time points) from the same experimenter.Serial samples are carried out this type of test and can identify the change of mark methylation state over time.The change of methylation state and the no change of methylation state can provide useful information about disease state, include but not limited to the approximate time that identification event starts, the existence and amount of salvageable tissue, the suitability of drug therapy, the effectiveness of various therapies, and the result of identifying experimenter, including the risk of future events.
[0221] Analysis of biomarkers can be performed in a variety of physical formats. For example, the use of microtiter plates or automation can be used to facilitate the processing of large numbers of test samples. Alternatively, a single sample format can be developed to facilitate immediate treatment and diagnosis in a timely manner, such as in ambulatory transport or emergency room settings.
[0222] As described above, according to embodiments of the methods disclosed herein, detecting changes in methylation profiles and / or aneuploidy scores can be qualitative or quantitative. Thus, the step of diagnosing a subject as having a particular type of cancer or being at risk for a particular type of cancer indicates that certain threshold measurements have been performed (e.g., the methylation state of one or more biomarkers in a biological sample is different from a predetermined control methylation state). In some embodiments of the method, the control methylation profile and / or aneuploidy score is any detectable methylation profile and / or aneuploidy score. In other embodiments of the method, wherein the control sample is tested simultaneously with the biological sample, the predetermined methylation profile and / or aneuploidy score is the methylation profile and / or aneuploidy score in the control sample. In other embodiments of the method, the predetermined methylation profile and / or aneuploidy score is based on a standard curve and / or is identified by a standard curve. In other embodiments of the method, the predetermined methylation profile and / or aneuploidy score is a specific profile or profile range. Thus, the predetermined methylation profile and / or aneuploidy score can be selected, within acceptable limits apparent to those skilled in the art, based in part on, among other things, the embodiment and desired specificity of the method being practiced.
[0223] Further with respect to the diagnostic methods, the preferred subject is a vertebrate subject. The preferred vertebrate is warm-blooded; the preferred warm-blooded vertebrate is a mammal. The preferred mammal is most preferably a human. As used herein, the term "subject" includes both human and animal subjects. Thus, veterinary therapeutic uses are provided herein. Thus, embodiments of the present disclosure provide for the diagnosis of mammals, such as humans, as well as those mammals that are important due to endangerment, such as Siberian tigers; mammals of economic importance, such as animals raised on farms for human consumption; and / or animals of social importance to humans, such as animals kept as pets or in zoos. Examples of such animals include, but are not limited to: carnivorous plants, such as cats and dogs; swine, including pigs, hogs, and wild boars; ruminants and / or ungulates, such as cattle, bulls, sheep, giraffes, deer, goats, bison, and camels; and horses. Thus, diagnosis and treatment of livestock are also provided, including, but not limited to, domestic pigs, ruminants, ungulates, horses (including racehorses), and the like.
[0224] 4. Samples, Kits, and Controls
[0225] Embodiments of the present disclosure provide technologies for screening various types of esophageal cancer and precancerous lesions from biological samples. According to these embodiments, the present disclosure includes but is not limited to methods and compositions for detecting the presence of various types and / or subclasses of esophageal cancer or precancerous lesions from biological samples. In some embodiments, the biological sample is a tissue sample, a blood sample, a plasma sample, a serum sample, a whole blood sample, a buffy coat sample, a secretion sample, an organ secretion sample, a cerebrospinal fluid (CSF) sample, a saliva sample, a urine sample and / or a fecal sample. In some embodiments, the tissue sample is an esophageal tissue sample. In some embodiments, the esophageal sample is obtained from an esophageal biopsy, or obtained by wiping, brushing or using a sponge capsule device. In some embodiments, the subject is a human being.
[0226] In other embodiments, "sample", "test sample" and "biological sample" refer to a fluid sample containing or suspected of containing the methylated DNA markers of the present disclosure. The sample may be derived from any suitable source. In some cases, the sample may comprise a liquid, a flowing particulate solid or a fluid suspension of solid particles. In some cases, the sample may be processed before the analysis described herein. For example, the sample may be separated or purified from its source before analysis. In a specific example, the source is a mammalian (e.g., human) body material (e.g., body fluids, blood such as whole blood, buffy coat, serum, plasma, urine, saliva, sweat, sputum, semen, mucus, tears, lymph, amniotic fluid, interstitial fluid, cerebrospinal fluid, feces, tissue, organ, one or more dried blood spots, etc.). Tissues may include, but are not limited to, esophageal tissue, gastric tissue, pancreatic tissue, bile duct / liver tissue and colorectal tissue. The sample may be a liquid sample, or a liquid extract of a solid sample. In some embodiments, the source of the sample can be an organ or tissue, such as a biopsy sample and / or an endoscopic brushing sample (eg, an endoscopic esophageal brushing sample), which can be lysed by tissue disintegration / cell lysis.
[0227] Various volumes of fluid samples can be analyzed. In some exemplary embodiments, the sample volume can be about 0.5 nL, about 1 nL, about 3 nL, about 0.01 μL, about 0.1 μL, about 1 μL, about 5 μL, about 10 μL, about 100 μL, about 1 mL, about 5 mL, about 10 mL, etc. In some cases, the volume of the fluid sample is between about 0.01 μL and about 10 mL, between about 0.01 μL and about 1 mL, between about 0.01 μL and about 100 μL, or between about 0.1 μL and about 10 μL.
[0228] In some cases, fluid sample can be diluted before being used to measure.For example, in the embodiment that the source containing methylated DNA mark is body fluid (for example blood, serum, secretion), available suitable solvent (for example buffer, such as PBS buffer) dilution body fluid.Fluid sample can be diluted about 1 times, about 2 times, about 3 times, about 4 times, about 5 times, about 6 times, about 10 times, about 100 times or more times before use.In other cases, fluid sample is not diluted before being used to measure.
[0229] In some cases, sample may be processed before analysis. Processing before analysis can provide additional functions, such as non-specific protein removal and / or the mixing function of effective but cheap realization. The general method of processing before analysis can comprise and use electric capture, AC electric, surface acoustic wave, isotachophoresis, dielectrophoresis, electrophoresis or other pre-concentration technology known in the art. In some cases, fluid sample can be concentrated before being used to measure. For example, in the embodiment that the source that contains methylated DNA mark is human body fluid (for example blood, serum, secretion), can concentrate fluid by precipitation, evaporation, filtration, centrifugation or its combination. Fluid sample can be concentrated about 1 times, about 2 times, about 3 times, about 4 times, about 5 times, about 6 times, about 10 times, about 100 times or more times before use.
[0230] It may be necessary to include a control. The control can be analyzed simultaneously with the sample from the subject, as described above. The results obtained from the subject sample can be compared with the results obtained from the control sample. A standard curve can be provided, which can be compared with the assay results of the sample. Such a standard curve shows the relationship between the assay unit and the level of one or more methylated DNA markers. Using samples taken from multiple donors, a standard curve can be provided for reference levels of methylated DNA markers in normal healthy tissue and for "risk" levels of methylated DNA markers in tissue taken from donors who may have one or more characteristics of esophageal cancer or precancerous lesions.
[0231] The embodiments of the present disclosure also include test kits for performing the methods described herein. Test kits include embodiments of compositions, devices, equipment, etc. as described herein, and instructions for use of test kits. Such instructions describe appropriate methods for preparing analytes from samples, such as collecting samples and preparing nucleic acids from samples. The individual components of the test kit are packaged in appropriate containers and packaging (e.g., vials, boxes, blister packs, ampoules, cans, bottles, tubes, etc.), and the components are packaged together in appropriate containers (e.g., one or more boxes) to facilitate storage, transportation, and / or user use of the test kit. It should be understood that liquid components (e.g., buffer) can be provided in lyophilized form for user recovery. The test kit may include controls or references for evaluating, verifying, and / or ensuring test kit performance. For example, a test kit for determining the amount of nucleic acid present in a sample may include controls comprising the same or another nucleic acid of known concentration for comparison, and in some embodiments, further includes a detection reagent (e.g., primer) having specificity for the control nucleic acid. The test kit is suitable for use in clinical settings, and in some embodiments, is suitable for use at home. In some embodiments, the components of the test kit provide the function of a system for preparing nucleic acid solutions from a sample. In some embodiments, certain components of the system are provided by the user.
[0232] In some embodiments, the present disclosure provides compositions (e.g., reaction mixtures). In some embodiments, the present disclosure provides compositions comprising nucleic acids comprising DMRs and reagents capable of modifying DNA in a methylation-specific manner (e.g., methylation-sensitive restriction enzymes, methylation-dependent restriction enzymes, and bisulfite reagents) (e.g., methylation-sensitive restriction enzymes, methylation-dependent restriction enzymes, 10-11 translocation (TET) enzymes (e.g., human TET1, human TET2, human TET3, murine TET1, murine TET2, murine TET3, Nasella TET (NgTET), Coprinus cinereus (CcTET)) or variants thereof), borane reducing agents). Some embodiments provide compositions comprising nucleic acids comprising DMRs and oligonucleotides, as described herein. Some embodiments provide compositions comprising nucleic acids comprising DMRs and methylation-sensitive restriction enzymes. Some embodiments provide compositions comprising nucleic acids comprising DMRs and polymerases.
[0233] In some embodiments, the technology described herein is associated with a programmable machine that is designed to perform a series of arithmetic or logical operations provided by the methods described herein. For example, some embodiments of this technology are associated with computer software and / or computer hardware (for example, executed therein). In one aspect, the technology relates to a computer comprising a form of memory, an element for performing arithmetic and logical operations, and a processing element (for example, a microprocessor) for executing a series of instructions (for example, as provided herein) to read, operate, and store data. In some embodiments, the microprocessor is part of a system for the following: determining a methylation profile (for example, a methylation profile of one or more DMRs in Tables 1, 2, or 3) and / or aneuploidy score; comparing methylation states; generating a standard curve; determining a Ct value; calculating a score, frequency, or percentage of methylation; identifying a CpG island; determining the specificity and / or sensitivity of an assay or label; calculating an ROC curve and associated AUC; sequence analysis; all as described herein or known in the art. In some embodiments, the microprocessor is part of a system for: determining a methylation profile (e.g., a methylation profile of one or more DMRs in Tables 1, 2, or 3) and / or an aneuploidy score; comparing methylation profiles and / or aneuploidy scores; generating a standard curve; determining Ct values; calculating the fraction, frequency, or percentage of methylation; identifying CpG islands; determining the specificity and / or sensitivity of an assay or marker; calculating ROC curves and associated AUCs; and sequence analysis; all as described herein or known in the art.
[0234] In some embodiments, a software or hardware component receives the results of multiple assays and determines a single value result based on the results of the multiple assays to report to the user, indicating a risk of cancer (e.g., determining the methylation status and / or aneuploidy score of one or more DMRs in Tables 1, 2, or 3). Related embodiments calculate a risk factor (e.g., determining the methylation status and / or aneuploidy score of one or more DMRs in Tables 1, 2, or 3) based on a mathematical combination (e.g., a weighted combination, a linear combination) of the results from multiple assays. In some embodiments, the methylation state of the DMR defines a dimension and may have a value in a multidimensional space and the coordinates defined by the methylation state of the multiple DMRs are results (e.g., reported to the user, or associated with a risk of cancer).
[0235] Various embodiments of the present disclosure are associated with a plurality of programmable devices that operate in concert to perform the methods described herein. For example, in some embodiments, a plurality of computers (e.g., connected via a network) can work in parallel to collect and process data, for example, in the implementation of cluster computing or grid computing or some other distributed computer architecture that relies on a complete computer (with onboard CPU, storage, power supply, network interface, etc.) connected to a network (private, public, or the Internet) via conventional network interfaces (e.g., Ethernet, fiber optic) or wireless networking technology.
[0236] For example, some embodiments provide a computer including a computer-readable medium. The embodiment includes a random access memory (RAM) coupled to a processor. The processor executes computer-executable program instructions stored in the memory. Such processors may include a microprocessor, an ASIC, a state machine, or other processors, and may be any of a variety of computer processors, such as processors from Intel Corporation (Intel Corporation) in Santa Clara, California and Motorola Corporation (Motorola Corporation) in Schaumburg, Illinois. Such processors include a medium (e.g., a computer-readable medium) or may communicate therewith, the medium storing instructions, when executed by the processor, causing the processor to perform steps described herein.
[0237] In some embodiments, the computer is connected to a network. The computer may also include many external or internal devices, such as a mouse, CD-ROM, DVD, keyboard, display or other input or output device. Examples of computers are personal computers, digital assistants, personal digital assistants, cellular phones, mobile phones, smart phones, pagers, digital tablets, notebook computers, internet appliances and other processor-based devices. Generally speaking, the computer related to the aspects of the technology provided herein can be any type of processor-based platform, which operates on any operating system that can support one or more programs including the technology provided herein, such as Microsoft Windows, Linux, UNIX, Mac OS X, etc. Some embodiments include a personal computer that executes other application programs (such as application programs). Application programs may be included in a memory and may include, for example, word processing applications, electronic spreadsheet applications, email applications, instant messaging applications, presentation applications, internet browser applications, calendar / organizer applications and any other applications that can be executed by a client device. All such components, computers and systems associated with the technology described herein may be logical or virtual.
[0238] In some embodiments, present disclosure provides a system for screening esophageal cancer or precancerous lesions in a sample obtained from a subject. An exemplary embodiment of the system includes, for example, a system for screening esophageal cancer or precancerous lesions in a sample (such as a tissue sample, a blood sample, a plasma sample, a serum sample, a whole blood sample, a buffy coat sample, a secretion sample, an organ secretion sample, a cerebrospinal fluid (CSF) sample, a saliva sample, a urine sample and / or a stool sample) obtained from a subject. In some embodiments, the system includes a software component configured to determine the methylation state and / or aneuploidy score of one or more methylation markers in the sample, a software component configured to compare the methylation state and / or aneuploidy score of the one or more methylation markers in the sample with a control sample or a reference sample recorded in a database, and an alarm component configured to warn a user of a cancer-related state.
[0239] In some embodiments, an alert is determined by a software component that receives results from multiple assays (e.g., determines the methylation status and / or aneuploidy score of one or more methylation markers) and calculates a value or result to be reported based on the multiple results.
[0240] Some embodiments provide a database of weighted parameters associated with each methylation marker provided herein, for calculating values or results and / or alerts to report to users (e.g., as physicians, nurses, clinicians, etc.). In some embodiments, all results from multiple assays are reported. In some embodiments, one or more results are used to provide a score, value (e.g., an aneuploidy score) or result, which is based on a combination of one or more results from multiple assays, indicating a cancer risk for the subject. Such methods are not limited to specific methylation markers. In such methods and systems, one or more methylation markers include bases in the DMR selected from the DMRs in Tables 1, 2, and 3.
[0241] In this detailed description of the various embodiments, for purposes of explanation, numerous specific details are set forth to provide a thorough understanding of the disclosed embodiments. However, those skilled in the art will appreciate that the various embodiments may be practiced with or without these specific details. In other cases, structures and devices are shown in block diagram form. Furthermore, those skilled in the art will readily appreciate that the particular order in which the methods are presented and performed is exemplary, and it is contemplated that these orders may be varied and still remain within the spirit and scope of the various embodiments disclosed herein.
[0242] The various components of the test kit can be placed in suitable containers as needed. The test kit may also include containers for holding or storing samples (e.g., containers or boxes for urine, whole blood, buffy coat, plasma, serum samples, tissue or body secretion samples). Where appropriate, the test kit may also optionally contain reaction vessels, mixing containers and other components that help prepare reagents or test samples. The test kit may also include one or more instruments for assisting in obtaining test samples, such as syringes, pipettes, tweezers, measuring spoons, etc. In some embodiments, the instrument is a collection device. In some embodiments, the test kit includes a collection device for endoscopic scrubbing or non-endoscopic whole esophageal scrubbing or wiping using a tethering device (e.g., a capsule sponge, balloon or other device). In some embodiments, the biological sample is obtained from a subject, and the method further includes extracting a DNA sample from the biological sample. In some embodiments, the biological sample is collected using a collection device having an absorbent member capable of collecting the biological sample upon contact. In some embodiments, the absorbent member is a sponge configured to be inserted into an orifice.
[0243] 5. Examples
[0244] It will be apparent to those skilled in the art that other suitable modifications and adaptations of the disclosed methods described herein are readily applicable and understandable, and can be made using suitable equivalents without departing from the scope of the disclosure or the aspects and embodiments disclosed herein. Having now described the disclosure in detail, it will be more clearly understood by reference to the following examples, which are intended only to illustrate some aspects and embodiments of the disclosure and should not be construed as limiting the scope of the disclosure. The disclosures of all journal references, U.S. patents, and publications cited herein are hereby incorporated by reference in their entirety.
[0245] The present disclosure has several aspects, illustrated by the following non-limiting examples.
[0246] Example 1
[0247] Experiments were performed to identify a panel of genetic markers that could distinguish non-dysplastic Barrett's esophagus (NDBE) from precancerous conditions, such as low-grade dysplasia (LGD) or high-grade dysplasia (HGD), and / or cancerous conditions, such as esophageal adenocarcinoma (EAC), based on differentially methylated regions (DMRs) and / or DNA copy number variations (CNVs) within the candidate genetic markers.
[0248] Reduced representation bisulfite sequencing (RRBS) studies were performed on tissue biopsy samples, and whole genome sequencing (WGS) studies were performed on endoscopic brushing samples. Briefly, for the RRBS tissue biopsy studies, an average of approximately 50 million mapped reads were generated per sample, and 4 million CpGs were generated per sample at 10x or higher coverage. After applying a proprietary marker discovery pipeline algorithm, logical analysis, and initial filtering (as described in the Materials and Methods below), 4501 DMRs were identified, with an average length of 156 bp. These mapped to 1549 annotated genes and 758 unannotated regions of the genome. Further filtering to remove control sample CpG methylation and applying a case / control FC cutoff brought the number of annotated and unannotated regions to 1196. For the endoscopic brushing WGS studies, an average of approximately 500 million mapped deduplicated reads were generated per sample, and 24 million CpGs were generated per sample at 10x or higher coverage. A total of 7345 DMRs were identified with an average length of 72 bp. These mapped to 2351 genes and 1374 unannotated regions. After applying subsequent filters and cutoffs, a total of 3534 regions remained. Given that RRBS is an enzyme-mediated genome enrichment method that enriches the genome from 3.2 billion bases to less than 5 million bases, the number of WGS is expected to exceed the number of RRBS. Due to the size selection step and the CCGG specificity of the MspI enzyme, RRBS misses potential CpG islands and promoter regions that do not tightly accommodate these tetramers.
[0249] The genomic regions of 1196 RRBS DMRs were merged with 3534 WGS DMRs to identify common hypermethylated sites. Regions had to overlap or have endpoints within 500 bases of each other. A total of 227 DMRs met these criteria. This was then narrowed down to 156 DMRs. Most of the exclusions were for regions that did not show any substantial consistent CpG methylation, which is a sign of true functional events in methylation-mediated inactivation of tumor suppressor genes or activation of oncogenes and enhancer elements. Other regions were regions with strong consistent methylation in the NDBE cohort. Some regions were unusually short (e.g., <60 bases), which limited their information content and imposed limitations on capture and methylation-specific amplification. In addition to the 156 merged regions, 43 additional candidate WGS DMRs were identified that were not in the RRBS data (probably due to the enrichment exclusion step in the protocol). Table 1 lists the complete set of 199 DMRs with gene annotations or SEQ ID NOs.
[0250] Table 1: Differentially methylated regions (DMRs) that distinguish Barrett's esophagus (BE) and / or normal esophageal samples from premalignant high-grade dysplasia (HGD) and / or esophageal adenocarcinoma (EAC). These DMRs were identified by whole-genome sequencing (WGS) studies using esophageal endoscopic brushing samples.
[0251]
[0252]
[0253]
[0254]
[0255]
[0256]
[0257]
[0258] The performance characteristics of DMRs in the whole-genome sequencing (WGS) study are provided in Table 2. In addition, 156 DMRs were also identified in a separate reduced representation bisulfite sequencing (RRBS) study using independent tissue biopsies, thus validating the results of the WGS study.
[0259] Table 2: Performance characteristics (AUC and fold change) of DMRs capable of distinguishing Barrett's esophagus (BE) and / or normal esophageal samples from premalignant high-grade dysplasia (HGD) and / or esophageal adenocarcinoma.
[0260]
[0261]
[0262]
[0263]
[0264]
[0265] Table 3 highlights the top DMRs with AUC > 0.75 and FC > 5 in the independent WGS and RRBS datasets. The DMRs in Table 3 represent the subset of DMRs that showed the highest performance in both studies.
[0266] Table 3: Performance characteristics (AUC and fold change) of the highest performing DMRs capable of distinguishing Barrett's esophagus (BE) and / or normal esophageal samples from premalignant high-grade dysplasia (HGD) and / or esophageal adenocarcinoma.
[0267]
[0268] Taken together, the combined data validate 156 DMRs. The two sequencing studies employed independent, yet distinct, sample sets. RRBS data were derived from clinical tissue biopsies, while WGS data were derived from endoscopic esophageal brushings. However, shared genes were identified in both studies, not only within widely divergent regions of these genes but also, in many cases, within overlapping sequence segments. Both studies also shared similar performance characteristics.
[0269] Example 2
[0270] Experiments were also performed to assess copy number variation (CNV), including ploidy and aneuploidy determinations, to differentiate Barrett's esophagus (BE) and / or normal esophageal samples from precancerous high-grade dysplasia (HGD) and / or esophageal adenocarcinoma. Specifically, another important aspect of the present disclosure is the ability to utilize sequencing reads from cytosine-converted genomes to perform copy number aberration (CNA) determinations, including ploidy and aneuploidy. Most technologies use direct NGS on wild-type, unconverted DNA. Here, methylation and CNV reads are obtained simultaneously from the same chemistry / dataset. For WGS scrubbing studies, aneuploidy calls were made using CNVkit and AneuploidyScore software and numerical scores were generated with the NE cohort as reference. These normal esophageal (NE) samples were defined as euploid (2 copies of the 22 somatic chromosomes) and confirmed by chromosome self-reference scatter plots (see, e.g. Figure 3 ). Mapped deduplicated reads were used for analysis, and the reads were a priori corrected for repetitive sequences, GC content, and PCR duplications. The aneuploidy score (AS) calculated for each sample is the total number of chromosome arm-level gains and losses adjusted for ploidy. See, e.g. Figure 1 To understand the arm-level and segmental aneuploidy scores (y-axis, signal) for EAC, HGD, and NDBE samples.
[0271] Figure 2 yes Figure 1 Violin plots of the data from [ ] show increasing aneuploidy levels and frequencies from nondysplastic to HGD to EAC. Chromosome arm gains or losses were observed in 14 (78% [52-94%]) patients with EAC, 7 (39% [17-64%]) patients with HGD, and 4 (22% [6-48%]) patients with NDBE (specificity 78%). The calls were confirmed by visual analysis of the CNVpytor chromosome scatter plots for each sample. See [ ] Figure 3 (AS 0), Figure 4 (AS 0), Figure 5(AS 4) and Figure 6 Representative images from (AS 9).
[0272] Example 3
[0273] Experiments were also performed to determine whether copy number aberrations (CNAs) and DRMs could complement each other to detect BE-associated HGD and EAC in NDBE (and NE) samples. For this example, one DMR, MAFB (v-maf musculoaponeurotic fibrosarcoma oncogene homolog B), was evaluated, which was the marker with the highest AUC in the above data and was detected in both RRBS and WGS studies. The discriminating region was approximately 1400 bp in length and was highly informative ( Figure 7 and 8 MABF hypermethylation was observed in 16 (89% [65-99%]) patients with EAC brushing, 14 (78% [52-94%]) patients with HGD brushing, and 3 (17% [4-42%]) patients with NDBE brushing (specificity 83%). Incorporating methylation and aneuploidy, 17 (94% [73-100%]) EAC and 16 (89% [65-99%]) HGD in NDBE samples were correctly classified with a fixed specificity of 85% [58-96%]. One EAC and two HGD were missed by both marker categories (Tables 4 and Figure 9 ).
[0274] Table 4. MAFB methylation-complemented CNV results. Negative samples (i.e., "neg" in the CNV call column) are samples where the CNV was missed but was identified by evaluating the methylation status (i.e., "pos" in the Methylation call column).
[0275]
[0276]
[0277] If an additional DMR (e.g., LONRF2) had been added to the panel with 100% specificity, one of the missed HGD patients would have been correctly called without a false-positive result. Furthermore, a subsequent estimated cell fraction analysis (using TGCA epigenetic profiles of different cell types) was performed on all WGS samples (data not shown) and revealed no clear columnar epithelial signal for the two remaining miscalled cases, whereas squamous signal was strong in both. These two brushes may have been improperly sampled or misclassified. The other cases and controls in the WGS study had appropriate epigenetic profiles.
[0278] Taken together, CNV and DMR analysis complement the classification of BE-HGD / EAC versus NDBE. Both CNV and DMR analysis can be performed on endoscopic brushing specimens and could enhance endoscopic surveillance. For example, aneuploidy occurs late in tumorigenesis; DNA methylation is a very early event, which explains why even NDBE samples are often (but not always) heavily hypermethylated and why the number of dysplasia-specific DMRs is lower than that typically observed in cancer / normal cohorts. Combining these genomic and epigenetic alterations is a powerful strategy for detecting dysplasia.
[0279] Example 4
[0280] The markers identified in Example 1 were validated by applying the 199 identified DMRs in Table 1 to an independent set of 169 endoscopic esophageal brushing samples. These samples were also subjected to full methylome sequencing (shallow layer) to confirm the previous aneuploidy results. Based on cross-validated DMR selection, a four-DMR model including KL, PGBD5, ROR2 and LMX1B achieved a cross-validated AUC of 0.87 (0.81-0.94, 95% CI) for the identification of HGD / EAC. Combined with the tMAD score of CNA, a cross-validated AUC of 0.92 (0.87-0.98) was achieved for the identification of HGD / EAC. At a specificity threshold of 80%, the 4-DMR model identified 35 patients (85% [71-94%]) with EAC, 23 patients (88% [70-98%]) with HGD, 14 patients (63% [41-83%]) with LGD, and 8 patients (19% [9-34%]) with NDBE (observed specificity, 81%). The 4-DMR + tMAD model identified 38 patients (93% [80-98%]) with EAC, 23 patients (88% [70-98%]) with HGD, 11 patients (50% [28-72%]) with LGD, and 8 patients (19% [9-34%]) with NDBE (observed specificity, 81%).
[0281] In this validation experiment, CNA and methylated DNA markers (alone and in combination) showed promising ability to distinguish BE-HGD / EAC from NDBE. Because both data types can be analyzed from endoscopic brushing specimens, they can be used for molecular augmentation of histological analysis in endoscopic surveillance.
[0282] 6. Materials and Methods
[0283] Marker Discovery. The following materials and methods were used to identify various DNA methylation markers that could distinguish NDBE, HGD, and EAC in subject biospecimens. Briefly, DNA was extracted using prospectively collected esophageal brushings from patients with NDBE (18), HGD (18), and EAC (18), as well as 17 non-BE squamous epithelial samples (NE). For NE patients, high-volume endoscopic brushings were obtained from the squamous epithelium and the distal 5 cm of the cardia, and for BE patients, from visible BE mucosa. Before the experiment, clinical, endoscopic, and histological data were extracted, and histology was verified by a dedicated gastrointestinal pathologist. Enzymatic methylation sequencing (New England Biolabs) libraries were prepared from the extracted DNA and sequenced on an Illumina NovaSeq 6000 system. Differentially methylated regions (DMRs) were identified from CpGs with 5× or greater read coverage. The positivity rate of DMRs in the NDBE group was set at 85% specificity. CNVkit used the NE group as a reference to call CNAs and determined aneuploidy events using the AneuploidyScore using default settings. The combination of methylation and aneuploidy status was assessed using logistic regression.
[0284] Patient samples. Tissue samples were obtained from the Mayo Clinic Biorepository under institutional IRB oversight. Endoscopic esophageal brushings were collected under Mayo IRB protocol 15-004540. Both sample types were unmatched to the patients and were from completely independent individuals. Samples included esophageal biopsies (FFPE) from patients with and without Barrett's esophagus / EAC, esophageal brushings from patients with and without Barrett's esophagus / EAC, and normal gastric cardia biopsies (FFPE).
[0285] Sample selection strictly adhered to the subject research authorization and inclusion / exclusion criteria. Exclusion criteria included: patient age <18 years; patients who had received chemotherapy for primary esophageal cancer before tissue collection; patients who had received therapeutic radiation or ablation (photodynamic therapy, radiofrequency, or cryotherapy) for primary esophageal cancer, LGD, HGD, or squamous dysplasia / carcinoma before sample collection (note: EMR or ESD alone were not excluded); patients with a history of disease grade (i.e., low-grade / high-grade / adenocarcinoma) higher than the target pathology; and patients with a history of pancreatic cancer, gastric cancer, liver cancer (HCC or cholangiocarcinoma), ampullary cancer, or duodenal cancer within the past 5 years before sample collection.
[0286] Exclusion criteria for normal squamous cell carcinoma and normal cardiac carcinoma included: patients with a history of esophageal squamous cell dysplasia or carcinoma or BE before sample collection; patients with a history of eosinophilic esophagitis before sample collection; patients with erosive esophagitis at the time of sample collection; patients with a history of gastrointestinal metaplasia or dysplasia before sample collection; patients who had received ablative treatment (photodynamic therapy, radiofrequency, or cryotherapy) for esophageal or gastric diseases.
[0287] Exclusion criteria for non-dysplastic BE (NDBE) included: patients with documented dysplasia sampled from Barrett's esophagus (either before or after the date of the sample used in this project); patients who had undergone surgery / biopsy / ablation (photodynamic therapy, radiofrequency, or cryotherapy) to treat (not only sample but also eradicate) neoplastic lesions in the esophagus (note: EMR or ESD were not excluded).
[0288] Inclusion criteria for normal squamous cell and normal cardiac carcinomas included patients with normal histology in their squamous cell and / or cardiac tissue (biopsy or other specimen types).
[0289] Inclusion criteria for nondysplastic BE (NDBE) included patients with histology of Barrett's esophagus without dysplasia and one visit before and one visit after the target visit with no evidence of dysplasia.
[0290] Inclusion criteria for LEG, HGD, and EAC included patients with Barrett's esophagus with dysplasia or esophageal adenocarcinoma histology.
[0291] Brushing samples included 18 cases of esophageal adenocarcinoma (EAC), 18 cases of Barrett's esophagus with high-grade dysplasia (HGD), 18 cases of non-dysplastic Barrett's esophagus (NDBE), and 17 cases of normal esophageal squamous epithelium (NE). The latter were sampled from the distal 5 cm of the esophagus, including some of the cardia. Clinical, endoscopic, and histological data were extracted and verified by an independent pathologist.
[0292] Table 5. Patient characteristics.
[0293]
[0294] Tissues consisted of the following: 3 EAC, 3 gastroesophageal junction carcinoma (GEJC), 18 HGD, 13 Barrett's esophagus with low-grade dysplasia (LGD), 16 NDBE, 10 NE, and 17 cystic gastric cardia (GC) specimens. Tissues were macrodissected, and histology was reviewed by a dedicated gastrointestinal pathologist. Samples were age-matched, randomized, and blinded. DNA was purified from tissue using the Qiagen QIAmp FFPE Tissue Kit (Qiagen, Germantown, MD). DNA was repurified using AMPure XP beads (Beckman-Coulter, Brea, CA) and quantified using PicoGreen (Thermo-Fisher, Waltham, MA). DNA integrity was assessed using qPCR. Brush washes were collected and stored in 1 mL of Qiagen Cell Lysis Buffer at −80°C. DNA was subsequently extracted using the Qiagen Puregene kit modified for a 1 mL sample volume.
[0295] Sequencing. Reduced representation bisulfite sequencing (RRBS) sequencing libraries were prepared from tissue genomic DNA using a modified NuGEN Ovation RRBS Methyl-Seq kit (Tecan Genomics, Redwood City CA). Samples were pooled in 4-plex format and sequenced on an Illumina HiSeq4000 instrument (Illumina, San Diego CA) by the Mayo Genomics Facility. Reads were processed by the Illumina pipeline module for image analysis and base calling. Secondary analysis was performed using the Mayo-developed bioinformatics suite SAAP-RRBS. Briefly, reads were cleaned using Trim-Galore and aligned to the GRCh37 / hg19 reference genome constructed using BSMAP. For CpGs with coverage ≥10X and base quality scores ≥20, methylation ratios were determined by calculating C / (C+T) or, for reads mapped to the reverse strand, G / (G+A).
[0296] For endoscope brush samples, 120 ng of DNA was sheared to approximately 300 bp using a Covaris LE220 ultrasonic generator (Covaris, Woburn MA). The samples were concentrated to 50 uL, and whole-genome sequencing (WGS) libraries were prepared using the NEBNext Enzymatic Methyl-seq kit (New England Biolabs, Ipswich MA). Briefly, the sheared samples were end-repaired, A-tailed, and index adapter-ligated. The libraries were then subjected to an enzymatic conversion step, amplified, and differentiated between methylated and unmethylated cytosines. The converted libraries were amplified, quantified, and pooled as a standardized input in a 24-plex format. A total of three pools were prepared and sent to the Mayo Genome Analysis Core for sequencing on an Illumina NovaSeq system (Illumina, San Diego CA) using an S4 flow cell (PE 150 cycles). Primary and secondary analyses were performed as described above.
[0297] Biomarker Selection. RRBS and WGS data were analyzed separately, but using similar metrics. Included CpGs had a coverage depth of 5 reads (or more) and variance >0 between subgroups. CpGs were then ranked based on their high methylation rate, defined as the ratio of the number of methylated cytosines at a given site to the total cytosine count at that site. For NE and GC controls, the ratio was required to be ≤ 0.01 (1%) and ≤ 0.05 (5%) (for NDBE samples). CpGs that did not meet these criteria were discarded from the results. Candidate CpGs were then genomically mapped into DMRs (differentially methylated regions), ranging from approximately 40 to 2200 bp, with a minimum cutoff of 5 CpGs / 200 bp region. These regions had to have an AUC > 0.60 and a methylation fold change (FC) > 3 between cases and NDBE controls (> 5 for RRBS data; tissues have higher cellular purity). For each candidate region, a 2D matrix was created to compare individual CpGs in cases and controls on a sample-by-sample basis. Final selection required that cases demonstrate consistent and continuous stretches of hypermethylation across a single CpG across the DMR sequence at the per-sample level. In contrast, for control samples, the methylation CpG pattern must be highly random. DMRs were discarded if the mean methylation value of one or more CpGs in the NDBE samples was greater than the mean methylation value of all EACs and HGDs for that DMR.
[0298] After regression, DMRs were ranked by area under the receiver operating characteristic curve (AUC) and fold change difference between cases and NDBE controls. No adjustment for false discovery was performed at this stage because independent validation was planned a priori.
[0299] Biomarker merging / validation. A subset of DMRs was selected for further development. These were identified by merging filtered DMRs from two independent patient and sample type discovery datasets, which were generated and analyzed separately. This was done not only at the gene annotation level, but also at specific genomic locations within genes. Therefore, eligible regions were those where the two discovered genomic coordinates overlapped or where the endpoints were no more than 500 bp apart.
[0300] Copy number variation. Copy number variation events in WGS next-generation sequencing (NGS) scrubbed data were called by CNVkit (github.com / etal / cnvkit) using the NE group as a reference, and aneuploidy events were determined by AneuploidyScore (github.com / quevedor2 / aneuploidy_score) using default settings. Scatter plots were generated using CNVpytor software (github.com / abyzovlab / CNVpytor), with each sample as its own reference and a genomic bin size of 10K.
[0301] Statistics. RRBS and WGS results were logistically analyzed for individual DMR / MDM performance, CNVs based on the log2 score of chromosomal ARM gain / loss, and combinations of methylation and aneuploidy status.
[0302] Validation. For validation, DNA was extracted from prospectively collected esophageal brushings from 169 patients with NDBE (42 patients), LGD (23 patients), HGD (26 patients), and EAC (41 patients), as well as from 37 non-BE squamous epithelial (NE) samples. For NE patients, endoscopic brushings were obtained from the squamous epithelium and distal 5 cm of the cardia, and for BE patients, from the visible BE mucosa, using one brush every 5 cm. Clinical and endoscopic data were extracted, and histology was confirmed by a gastrointestinal pathologist.
[0303] The same sequencing library preparation was used for both methylation and aneuploidy assays, but for the methylation assay, a target capture step was performed before sequencing. For the target capture step, IDT prepared custom DNA capture probes for the 199 DMRs listed in Table 1. Targeted (DMR) and shallow whole-genome (CNA) sequencing was performed on the enzyme-converted DNA libraries using the Illumina NextSeq platform. The number of DMRs was reduced by removing DMRs with low coverage and high variability in the control group (NDBE). Further variable selection was performed using the VSURF package in R, and the final model was generated using a random forest comparing NDBE with HGD / EAC. CNA burden was estimated using the trimmed median absolute deviation (tMAD) score in the ichorCNA package.
[0304] Table 6. Patient characteristics of the validation study. One LGD sample was excluded due to poor sequencing. NE: normal esophagus, NDBE: non-dysplastic BE, LGD: low-grade dysplasia, HGD: high-grade dysplasia, EAC: esophageal adenocarcinoma
[0305]
[0306] Sequences. Various nucleotide sequences cited in this disclosure are provided below (see, e.g., Human 2009 February (GRCh37 / hg19) Assembly).
[0307] MAX.chr15.4912 (DMR 110):
[0308] CGTGGTTGCTGCGTGTGTCCCGCAGCCCCCTGCAGCCAGCGTCTCCTCGGCCCCGGCCCGACCCTGCCGTCCCTGCTCTGCTCCCTCAGCGACCGCAGACCCCTCACGCACATGCCCAGCCCTGCAGTCCTACTCTGCCACCCCAAGGGGCTTGGGCTTCAGTGGGGGCAGCGTGCGGGGCGTGGAGAGGGAGACGTGTAAAGCCCTGAGAGGCTCTGCAGCCACAGTCCCGCAGTCAGCGTTCCCTCTCGGCCAAGCCTCCTGGCTGCACATCGTTTCCCGGGCATCTCAGTGAGACCGGCCTCGGCTCACAGCGGCACGTTGTTTTTACACTTCGGGGACTCTGGTGGGACCGCGGGAAGCCAGGAGGCGGGCGCGCCCCGGGAGCCGATAGGAAAGTGCAACAGCGCCATCTAGTGGCCGCGCGGGGAGGCTGCGGAGCGCGCGCCGCGACCCGGGATCCACCGTCCGCAGGAAAACGTGTTCCCTCCGATGCGTGGTCACAGCGCGGCAGCGCGGCCTTGGGAGCTTCTTAAACAACAGGTTAGCTATCGCTGCCGCCTCAGTAAGAGAAGTCAAGGCCAAAACGCTTAAAGGATTTTGACCTATTCGGCTTGGTGGAGAGCGCTAAAAGGGGACGCTATAGCGGGATTCTCAGCTCTCCTGGGCTAGTGGGACCCCCGGTGCGCCGCCGCCGATGGGGTCTTAGGGCCTGTCTGTCTGGGCTAGAGCCGGCCCGGGAGCCTGTTTGCGGGGAGTGCG (SEQ ID NO: 1)。
[0309] MAX.chr17.8070 (DMR 115):
[0310] ATCTCCGATACTCCTCTCCTCAGCTTCCCGGGGGCAGTATGTCCCTTTGGGTCATGGCCTGTGGAGTGGCCATGACCCAAAGTTCAATTAAGGCTGGCGGGTCCCCAGAGGGGCCTGCTTCCATTCCCCCTTCCCCGCAGCCCTGCAGGCGTCAGCATGGGGCGAGGTTACCAGTTCCTGACCCAGCAGCTCGTCCTCAGGTGGTCCTGCTGCTCGGGAGGTCAGCTGCGTGGGGAGCCTGTCCACCTGGCTGATGGGGGTGATGGCCGCGTGGCTGGGATTCCTCCTCCAGCCCTGCCCTTCACCGCCATACAGTCTGGGCAGCCCCTTCCTGTCCCGTCCTCCTGCCTCGCCCTCTGCTGGGGTCTACGGGGCCCAAGTCTTGCTGGAACACAGGCAGGACAAGGGTTCTGGAGGTTCAGGCCGGAAGTGAAATAAAAGTCGAGAGGGGCCGGGCCCGGCACCCACCGCATCCTGCCTCCGGGCCTCCTCGGGCCCCTCACACGGTGGGGCAGCCCCTCCTCCAGTGGAGCGGAACGGTGGGCCCCGCTGCCCCTCCACCCCTGCGTGTGACACGGACGTGCTTGGAAGCTGGGTTCGTCCTTTCAGTTTTGTTTCCTGGTTGGAAACCCCTGAGACTCCTCCCCCTCCCCCCGCCCCTGCCCACCTGTGCTTTCCCAGGCTGGAAGGGGCAGCCCCTCTCAGCACGTGGGTGGTTAGGCTGTGACGTCCCCAGGCCTCCTGGGCCTGCACCGGGGTCCACCCAGGTGTGGGCCATGCCCAGGTGCGGG (SEQ ID NO: 2)。
[0311] MAX.chr19.3439 (DMR 117):
[0312] CGGTTTCGCTCCCTATCGGGGCCAGGGACGCCTCAGGCTGTCTCGGTCAGGACCTACAGCTCCGGTTGTTGTCCCAGGCTCTTCCGCGAGGTGCTCTCCTGTCTCCTGACCACCCCCATTCCTCCCCACTCCAGCTCCTCAGCGAGCGGCTGCGAAGGACGCGCACAACAACACTCCGCGCAGCGGGAAGGTACCGAACTCGCATCGCAGCCTGAAACTGCTCAACTAAGCTCCCGCCCCGCTGCTCTCTGGTCAATCTAAAGCGAAGACGAGCCTTAGGGCCAATCAGAAGCGACAGCGGTGGAGTCATGCCCGCCTGCTTGAGGCGCCCTGGCGTCTCATTGGCTATGCTTGAGAACGAATCCCAGGCTAAGCCACTTAGAAAGGAGCGGAGCCAGCCAATCAGCGGTGCAACCGCCGCGGGGGCCGGGCCAGAAGCCCCGCAGACAAGCACCTCGGGAGACTGGCGAGGGGCGAGCTCGCAGCTTGTTAGCCCCGAAGCCCTGCCCAGGGGGAAACCCTGCTGGAGGGAGTCTAACCCCCGGGCCAGTTAATGTTGGGGCAGCGCAATCGGCCGTTCCCTCTGCGGTATGGTTGGAGGTGGGAGTGCTGACACGTCCGCGCG (SEQ ID NO: 3)。
[0313] MAX.chr4.4552 (DMR 124):
[0314] GCCTGGGGCCACTGCTCCTGGGTCCTCAGGACTGCCTGGGGGAAGGTAGTGCATTGTGCAGCGCGCGGTCCAGAAGTGAAAAGGGAGGCGCGGAGATAAGCTGCCGGCGGAAGTTCCCTCTCCTGCCTGGGCCGACCCGGCGCTTTACTGCTTCTCACGAAGGTGCGCCGGCTGCTCCAGAAATCGCAGACTGCCTCCAGGAAGAACTTGCTGGAGTCACAGCAGCTTCTCAGCGACTTGACAGCAGTGATTCAGACTTCAACTTGGGCGGAGGGGCGGGGGAGGAGAAAGAGATTTCCAGAGAAAACGACTGAGCGGTAGGGAGGGGAAGAGAGACCGAGCCACGCGCCTCGAGAAGCAGTGCAGAGAGCGGAAGAGACAGAGGCTGCGAGACCTACCCACAGAGACCAAGAGAGACGCTCGGAGGGGAGACCGCCTTAGGCGCAGAGATTCAGAGGCAGACAGACAGGCAGACAGAAGGATACAGGAAAGGAATGTCGCCGAAAGGCAGGGACAAACCTGAGTCCCAGAAAAATAAGAGACAACTCCCACACACCAGGCTGTCCGCGGGCCGCCTTGTCACAGAAAGGCAGCTCCCCAGCCCCGCAGAGTCCCGACAGCTGCCCCCGCGAAGGTGGGGCGAGGGGCGGCTTTTCCG (SEQ ID NO: 4)。
[0315] MAX.chr6.6743 (DMR 126):
[0316]
[0317] MAX.chr8.3003 (DMR 129):
[0318] CGCGCCTCCCGGAACCACGCGTCTCTGTGCACAGACATTCCTGGGGGCAGGCTCCTGTCCTTTAACACAGTCTCAAAGGAGTCTTCAAAAACAAAAAGTTTACAAGCACTGATAGGGAAAACAGAAGGATGATAACACCAAGAGGGACTTTCCCGGGAGAGGACGGCTTCCAGGGACTGAGAAAGGATGGGCAAGTGGGCGGGGCCCGGCGCCGTGCGGGCAGGGCTGGCGCCGGGAGTCCCCAGACTCCCCCGCAGTGGGAAGCACCTCTCCCATTCACGCCGGGCAGGACACCTGGCCGGGCGGGGGAGGCAGCGCAAGGGCCGGCCGGGGAGTACGGGACTCGAGCCGGGGACCTGAGGCAGGAGCCAAGCATCGTCGCAGGGCAACCAGCAGAACGGAGAGGGAGGCGCGGGGGCGAAGGCTGGCGGGAGCCGCGCTGAGGGCAGGAGCCCGGAGCCCCCTAGGGCAGCGCCGATCCGCCCGCCCCGTCCCGCCGAGCTGGGCCTCCGTCTGTGGCCTGCGCAGCCAGGGTCGCCAAGCCG (SEQ ID NO: 6)。
Claims
1. A method for characterizing a biological sample, the method comprising: The methylation profile in at least one differentially methylated region (DMR) of a DNA sample obtained from a subject having or suspected of having esophageal cancer or a precancerous lesion is determined by treating the sample with an agent that modifies DNA in a methylation-specific manner.
2. The method of claim 1, wherein the methylation profile in the at least one DMR indicates that the subject has or is suspected of having high-grade dysplastic Barrett's esophagus or esophageal adenocarcinoma (EAC).
3. The method of claim 1 or claim 2, wherein the at least one DMR is from a gene selected from the group consisting of: ACVRL1, ADAMTS8, ADAP2, ADRBK2, AKR1B1, ANK1, ANKRD13B, ANXA6, ARNT2, B4GALNT2, BACH2, BCL11A, BSCL2, C12orf53, C14orf82, C17orf107, C18orf1, C1orf95, C5orf42, CACNA1C, CAMK1D, CAMTA1, CBX6, CCDC85A, CCKBR, CD38, CDKN2A, CH2 5H, CHST1, CHST15, CNTLN, CRHR1, CRTC1, CXCR4, CYP1B1, DCTN2, DIDO1, DMKN, DSE, DYNC1I1, EML6, ENOX1, EPHA4, ESRRG, FAM176A, FAM78B, FBXO1 0. FERMT2, FHOD3, FLJ45079, FMNL1, FOXP2, FRMD4B, FZD8, GALNTL1, GALNTL4, GAS1, GBGT1, GLIPR2, GNAI1, GNAL, GPR37, GRASP, GRID1, GRM8, GSC, HAR1A, HCN2, HEY2, HIST1H2BE, HS3ST3B1, HTR7, IGFBP2, INHA, INSRR, IRX3, ISM2, KCNG3, KCNK4, KCNS2, KCTD15, KIAA1522, KIAA1614, KIF26A, K L, KLF15, KLHL10, KRT77, KSR2, LBH, LMX1A, LMX1B, LOC100526820, LONRF2, LRFN2, LRRN1, MAF, MAFB, MARK1, ADAMTSL4-AS1, PGBD5, HSPA12A, SFT PD, LOC107984507, LOC100128253, CISTR, SLC16A7, CTXND1, MAX.chr15.4912, ZNF423, RBFOX1, LOC105376772, GSE1, MAX.chr17.8070, ZNF709, MAX.chr19.3439, GREB1, CYP27C1, AC068134.6, LOC388780, STOX2, PPARGC1A, MAX.chr4.4552, PDZD2, MAX.chr6.6743, LPAL2, CDK14, MAX.chr8.<h2 style=";text-align:left;direction:ltr">3003、MCOLN2、MEGF11、MFSD11、MRC2、MSX1、NAT8L、NAV1、NBEA、NCRNA00092、NECAB2、NEURL、NGEF、NHLH2、NKD1、NLGN1、NOG、N PPC、NR3C1、NRXN2、NTN1、OBSCN、OCA2、OSBPL1A、OXR1、P2RY1、PDE2A、PDE9A、PDGFRA、PID1、PLCG2、PLCL1、PNPLA3、POU3F1、POU 3F2、PRKACB、PRKAR2B、PRKG1、PRR18、PRR5L、PTHLH、PYGL、RARG、RHBDL3、RIMS2、ROR2、RPRML、SDK2、SOX9、SYNGR1、TAC4、TNFR SF19、TRANK1、TSPAN33、TSPAN4、TSPAN5、UBE2E2、UCHL1、UNC5A、VASH2、VIM、WNT6、ZBTB10、ZNF680、ZNF738、ZNF808、ZNF845。.
4. The method of claim 1 or claim 2, wherein the at least one DMR is from a gene selected from the group consisting of ACVRL1, ADAMTS8, ADAP2, ADRBK2, AKR1B1, ANK1, ANKRD13B, ANXA6, ARNT2, B4GALNT2, BACH2, BCL11A, BSCL2, C14orf82, C18orf1, C1orf95, C5orf42, CAMK1D, CAMTA1, C CDC85A, CD38, CDKN2A, CHST1, CHST15, CRHR1, CYP1B1, DIDO1, DSE, DYNC1I1, EML6, ENOX1, EPHA4, FAM176A, F BXO10, FERMT2, FHOD3, FLJ45079, FMNL1, FRMD4B, FZD8, GALNTL1, GALNTL4, GAS1, GBGT1, GLIPR2, GNAI1, GNAL , GRASP, GRID1, GRM8, GSC, HAR1A, HCN2, HEY2, HIST1H2BE, HS3ST3B1, HTR7, IGFBP2, INHA, IRX3, ISM2, KCNG3 , KCNK4, KCNS2, KCTD15, KIAA1614, KL, KLHL10, KSR2, LBH, LMX1A, LMX1B, LOC100526820, LONRF2, LRFN2, MAFB , MARK1, PGBD5, HSPA12A, SFTPD, LOC107984507, LOC100128253, CISTR, SLC16A7, CTXND1, ZNF423, LOC10537 6772, GSE1, MAX.chr19.3439, GREB1, CYP27C1, AC068134.6, LOC388780, STOX2, PPARGC1A, PDZD2, MAX.chr6.6743, LPAL2, CDK14, MCOLN2, MEGF11, MRC2, MSX1, NAT8L, NAV1, NBEA, ncRNA00092, NEURL, NGEF, NHLH2, NKD1, NLGN1, NOG, NPPC, NR3C1, NRXN2, OBSCN, OCA2, OSBPL1A, OXR1, P2RY1, PDE2A, PDGFRA, PID1, PLCG2, PLCL1, PNPLA3, POU3F1, POU3F2, PRKACB, PRKAR2B, PRKG1, PRR18, PRR5L, PTHLH, PYGL, RHBDL3, RIMS2, ROR2, RPRML, SDK2, SOX9, SYNGR1, TAC4, TRANK1, TSPAN4, UBE2E2, UCHL1, UNC5A, VASH2, ZBTB10, ZNF680, ZNF709, ZNF738, ZNF808, and ZNF845.
5. The method of claim 1 or claim 2, wherein the at least one DMR is from a gene selected from the group consisting of BACH2, C5orf42, FHOD3, HIST1H2BE, IRX3, KIAA1614, LONRF2, MAFB, PDGFRA, PID1, POU3F1, PRR5L, RHBDL3, and SDK2.
6. The method of claim 1 or claim 2, wherein the at least one DMR is from a gene selected from the group consisting of: KL, PGBD5, ROR2, and LMX1B.
7. The method of any one of claims 1 to 6, wherein the at least one DMR is associated with an area under the ROC curve (AUC) greater than or equal to 0.5, and wherein the ROC curve distinguishes subjects having or suspected of having high-grade dysplastic Barrett's esophagus or esophageal adenocarcinoma (EAC) from control DNA samples.
8. The method of any one of claims 1 to 6, wherein the at least one DMR is associated with an area under the ROC curve (AUC) greater than or equal to 0.75, and wherein the ROC curve distinguishes subjects having or suspected of having high-grade dysplastic Barrett's esophagus or esophageal adenocarcinoma (EAC) from control DNA samples.
9. The method of any one of claims 1 to 6, wherein the at least one DMR is associated with a methylation fold change (FC) ratio greater than or equal to 3.0, and wherein the FC ratio distinguishes subjects having or suspected of having high-grade dysplastic Barrett's esophagus or esophageal adenocarcinoma (EAC) from control DNA samples.
10. The method of any one of claims 1 to 6, wherein the at least one DMR is associated with a methylation fold change (FC) ratio greater than or equal to 5.0, and wherein the FC ratio distinguishes subjects having or suspected of having high-grade dysplastic Barrett's esophagus or esophageal adenocarcinoma (EAC) from control DNA samples.
11. The method of any one of claims 1 to 10, wherein the method further comprises determining a copy number variation (CNV) of the DNA sample of the subject; and wherein the CNV distinguishes subjects having or suspected of having high-grade dysplastic Barrett's esophagus or esophageal adenocarcinoma (EAC) from a control DNA sample.
12. The method of any one of claims 1 to 10, wherein the method further comprises determining an aneuploidy score (AS) of the DNA sample of the subject; and wherein the AS distinguishes subjects having or suspected of having high-grade dysplastic Barrett's esophagus or esophageal adenocarcinoma (EAC) from control DNA samples.
13. The method of any one of claims 1 to 12, wherein the at least one DMR comprises an increased percentage of methylation compared to a control DNA sample.
14. The method of any one of claims 1 to 12, wherein the at least one DMR comprises an increased hypermethylation ratio compared to a control DNA sample.
15. The method of any one of claims 1 to 12, wherein the at least one DMR comprises increased AS compared to the control DNA sample.
16. The method of any one of claims 7 to 15, wherein the control DNA sample is from a subject who does not have esophageal cancer or precancerous lesions.
17. The method of any one of claims 7 to 15, wherein the control DNA sample is from a subject with non-dysplastic Barrett's esophagus (NDBE).
18. The method of any one of claims 7 to 17, wherein the control DNA sample is selected from a tissue sample, a blood sample, a plasma sample, a serum sample, a whole blood sample, a buffy coat sample, a secretion sample, an organ secretion sample, a cerebrospinal fluid (CSF) sample, a saliva sample, a urine sample, and a stool sample. The method of claim 18 , wherein the tissue sample is an esophageal tissue sample or a gastric cardia sample.
20. The method of any one of claims 1 to 19, wherein the biological sample is selected from the group consisting of a tissue sample, a blood sample, a plasma sample, a serum sample, a whole blood sample, a buffy coat sample, a secretion sample, an organ secretion sample, a cerebrospinal fluid (CSF) sample, a saliva sample, a urine sample, and a stool sample.
21. The method of claim 20, wherein the tissue sample is an esophageal tissue sample.
22. The method of claim 20, wherein the tissue sample is an endoscopic esophageal brushing sample.
23. The method of any one of claims 1 to 13, wherein the subject is a human.
24. The method of any one of claims 1 to 13, wherein the biological sample is obtained from the subject, and wherein the method further comprises extracting the DNA sample from the biological sample.
25. The method of any one of claims 1 to 24, wherein the biological sample is collected using a collection device.
26. The method of any one of claims 1 to 25, wherein the agent that modifies DNA in a methylation-specific manner is a borane reducing agent.
27. The method of any one of claims 1 to 25, wherein the reagent that modifies DNA in a methylation-specific manner comprises one or more of a methylation-sensitive restriction enzyme, a methylation-dependent restriction enzyme, and a bisulfite reagent.
28. The method of any one of claims 1 to 27, wherein determining the methylation profile of at least one DMR comprises amplifying at least a portion of the DMR using a set of primers.
29. The method of any one of claims 1 to 28, wherein determining the methylation profile of at least one DMR comprises performing at least one of methylation-specific PCR, quantitative methylation-specific PCR, methylation-specific DNA restriction enzyme analysis, quantitative bisulfite pyrosequencing, a flap endonuclease assay, a PCR-flap assay, and bisulfite genomic sequencing PCR.
30. The method of any one of claims 1 to 29, wherein determining the methylation profile of at least one DMR comprises determining the presence or absence of methylation at one or more CpG sites.
31. The method of claim 30, wherein the one or more CpG sites are present in the coding region, non-coding region and / or regulatory region of a gene.
32. The method of any one of claims 1 to 31, wherein determining the methylation profile of at least one DMR comprises determining a methylation frequency.
33. The method of any one of claims 1 to 31, wherein determining the methylation profile of at least one DMR comprises determining a methylation pattern.
34. The method of any one of claims 1 to 33, wherein the methylation profile in the at least one DMR and the CNV and / or the AS is determined using the same DNA sample obtained from the subject.
35. The method of any one of claims 1 to 33, wherein the methylation profile in the at least one DMR and the CNV and / or the AS is determined using a single DNA sample obtained from the subject.
36. The method of claim 34 or 35, wherein the sample has been treated with an agent that modifies DNA in a methylation-specific manner.
Citation Information
Patent Citations
Isolation of nucleic acids
US20120288868A1
Compositions and methods for detecting epithelial cell DNA
US20160194721A1
Method for genomic profiling of DNA 5-methylcytosine and 5-hydroxymethylcytosine
US20170253924A1
Bisulfite-free, base-resolution identification of cytosine modifications
US20200370114A1
Method of detection of methylated nucleic acid using agents which modify unmethylated cytosine and distinguishing modified methylated and non-methylated nucleic acids
US5786146A