Compositions and methods for detecting esophageal cancer

Novel DMRs and genetic markers enhance the detection of esophageal cancer and precancerous conditions by improving diagnostic accuracy and sensitivity, addressing the limitations of current methods.

JP2026502875APending Publication Date: 2026-01-27MAYO FOUNDATION FOR MEDICAL EDUCATION & RESEARCH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025536701
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-12-20
Filing Date
2023-12-20
Publication Date
2026-01-27

Smart Images

  • Figure 2026502875000023
    Figure 2026502875000023
  • Figure 2026502875000024
    Figure 2026502875000024
  • Figure 2026502875000025
    Figure 2026502875000025
Patent Text Reader

Abstract

The present disclosure provides compositions and methods for distinguishing between non-cancerous, pre-cancerous, and cancerous conditions in the esophagus. In particular, the present disclosure provides compositions and methods for distinguishing non-dysplastic Barrett's esophagus (NDBE) samples from pre-cancerous (e.g., low-grade or high-grade dysplasia) and / or cancerous (e.g., esophageal adenocarcinoma) samples based on methylation status and / or DNA copy number abnormalities.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 433,784, filed December 20, 2022, the entire contents of which are incorporated herein by reference for all purposes.

[0002] Incorporation by Reference of Electronically Submitted Materials Incorporated herein by reference in its entirety is the computer readable nucleotide / amino acid sequence listing, filed concurrently herewith, identified as follows: One 10,938 byte file entitled "41473-601_SEQUENCE_LISTING", created on December 14, 2023.

[0003] The present disclosure provides compositions and methods for distinguishing between non-cancerous, pre-cancerous, and cancerous conditions in the esophagus. In particular, the present disclosure provides compositions and methods for distinguishing non-dysplastic Barrett's esophagus (NDBE) and / or normal esophageal samples from pre-cancerous (e.g., low-grade or high-grade dysplasia) and / or cancerous (e.g., esophageal adenocarcinoma) samples based on methylation status and / or DNA copy number abnormalities. [Background technology]

[0004] Barrett's esophagus (BE) is a condition in which the esophageal lining is altered or replaced by tissue similar to that of the intestine. BE can progress to low-grade dysplasia (LGD) or high-grade dysplasia (HGD), precancerous conditions that can transform into esophageal adenocarcinoma (EAC). However, in most cases, dysplasia is not evident or identifiable in biopsy specimens. This condition is referred to as non-dysplastic BE (NDBE). Endoscopic surveillance is generally recommended to identify and differentiate between dysplastic, precancerous, and cancerous conditions. However, this diagnostic paradigm is limited by esophageal sampling errors and subtle histologic changes, leading to delayed treatment for dysplasia and EAC. Summary of the Invention

[0005] Embodiments of the present disclosure provide methods, compositions, and systems for screening for various types of esophageal cancer in biological samples. According to these embodiments, the present disclosure includes, but is not limited to, methods and compositions for detecting the presence of esophageal cancer or precancer from a biological sample. In some embodiments, the biological sample is a tissue sample, a blood sample, a plasma sample, a serum sample, a whole blood sample, a buffy coat sample, a secretion sample, an organ secretion sample, a cerebrospinal fluid (CSF) sample, a saliva sample, a urine sample, and / or a stool sample. In some embodiments, the tissue sample is an esophageal tissue sample. In some embodiments, the esophageal sample is obtained from an esophageal biopsy or by swabbing, brushing, or using a sponge capsule device.

[0006] As further described herein, embodiments of the present disclosure include novel differentially methylated regions (DMRs) that can distinguish esophageal cancer or precancer from control or benign tissue. In some embodiments, the novel DMRs can distinguish high-grade dysplastic Barrett's esophagus (HGD-BE) or esophageal adenocarcinoma (EAC) from non-dysplastic Barrett's esophagus (NDBE) or normal esophageal controls. According to these embodiments, the novel DMR(s) may be ACVRL1, ADAMTS8, ADAP2, ADRBK2, AKR1B1, ANK1, ANKRD13B, ANXA6, ARNT2, B4GALNT2, BACH2, BCL11A, BSCL2, C12orf53, C14orf82, C17orf107, C18orf1, C1orf95, C5orf42, CACNA1C, CAMK1D, CAMTA1, CBX6, CCDC85A, CCK BR, CD38, CDKN2A, CH25H, CHST1, CHST15, CNTLN, CRHR1, CRTC1, CXCR4, CYP1B1, DCTN2, DIDO1, DMKN, DSE, DYNC1I1, EML6, ENOX 1, EPHA4, ESRRG, FAM176A, FAM78B, FBXO10, FERMT2, FHOD3, FLJ45079, FMNL1, FOXP2, FRMD4B, FZD8, GALNTL1, GALNTL4, GAS1, GBGT1, GLIPR2, GNAI1, GNAL, GPR37, GRASP, GRID1, GRM8, GSC, HAR1A, HCN2, HEY2, HIST1H2BE, HS3ST3B1, HTR7, IGFBP2, INHA, INSRR, IRX3, ISM2, KCNG3, KCNK4, KCNS2, KCTD15, KIAA1522, KIAA1614, KIF26A, KL, KLF15, KLHL10, KRT77, KSR2, LBH, LMX1A, LMX1B, LOC100526820, LONRF2, LRFN2, LRRN1, MAF, MAFB, MARK1, ADAMTSL4-AS1, PGBD5, HSPA12A, SFTPD, LOC107984507, LOC1 00128253, CISTR, SLC16A7, CTXND1, MAX.chr15.4912, ZNF423, RBFOX1, LOC105376772, GSE1, MAX.chr17.8070, ZNF709, MAX.chr19.3439, GREB1, CYP27C1, AC068134.6, LOC388780, STOX2, PPARGC1A, MAX.chr4.4552, PDZD2, MAX.chr6.6743, LPAL2, CDK14, MAX.chr8.3003, MCOLN2, MEGF11, MFSD11 , MRC2, MSX1, NAT8L, NAV1, NBEA, NCRNA00092, NECAB2, NEURL, NGEF, NHLH2, NKD1, NLGN1, NOG, NPPC, NR3C1, NRXN2, NTN1, OBSCN, OCA2, OSBPL1A, OXR1, P2RY1, PDE2A, PDE9A, In some embodiments, the novel DMR(s) are derived from a gene selected from PDGFRA, PID1, PLCG2, PLCL1, PNPLA3, POU3F1, POU3F2, PRKACB, PRKAR2B, PRKG1, PRR18, PRR5L, PTHLH, PYGL, RARG, RHBDL3, RIMS2, ROR2, RPRML, SDK2, SOX9, SYNGR1, TAC4, TNFRSF19, TRANK1, TSPAN33, TSPAN4, TSPAN5, UBE2E2, UCHL1, UNC5A, VASH2, VIM, WNT6, ZBTB10, ZNF680, ZNF738, ZNF808, and ZNF845 (Table 1), including any combination thereof. In some embodiments, the novel DMR(s) are derived from a gene selected from Table 1, including any combination thereof. While each novel DMR alone can distinguish HGD-BE and / or EAC from NDBE and / or control samples, combining two or more of the novel DMRs can improve sensitivity. Thus, combinations of two or more novel DMRs selected from Table 1 are provided.

[0007] Embodiments of the present disclosure also include novel differentially methylated regions (DMRs), each of which can distinguish esophageal cancer or precancer from control or benign tissue. In some embodiments, the novel DMRs can distinguish high-grade dysplastic Barrett's esophagus (HGD-BE) or esophageal adenocarcinoma (EAC) from non-dysplastic Barrett's esophagus (NDBE) or normal esophageal controls. According to these embodiments, the novel DMR(s) can be ACVRL1, ADAMTS8, ADAP2, ADRBK2, AKR1B1, ANK1, ANKRD13B, ANXA6, ARNT2, B4GALNT2, BACH2, BCL11A, BSCL2, C14orf82, C18orf1, C1orf95, C5orf42, CAMK1D, CAMTA1, CCDC85A, CD38, CDKN2A, CHST1, CHST15, CRHR1, CYP1B1, DIDO1, DSE, DYNC1I1, EML6, ENOX1, EPHA4, FAM176A, FBXO10, FERM T2, FHOD3, FLJ45079, FMNL1, FRMD4B, FZD8, GALNTL1, GALNTL4, GAS1, GBGT1, GLIPR2, GNAI1, GNAL, GRASP, G RID1, GRM8, GSC, HAR1A, HCN2, HEY2, HIST1H2BE, HS3ST3B1, HTR7, IGFBP2, INHA, IRX3, ISM2, KCNG3, KCNK4 , KCNS2, KCTD15, KIAA1614, KL, KLHL10, KSR2, LBH, LMX1A, LMX1B, LOC100526820, LONRF2, LRFN2, MAFB, MAR K1, PGBD5, HSPA12A, SFTPD, LOC107984507, LOC100128253, CISTR, SLC16A7, CTXND1, ZNF423, LOC1053767 72, GSE1, MAX.chr19.3439, GREB1, CYP27C1, AC068134.6, LOC388780, STOX2, PPARGC1A, PDZD2, MAX.chr6.6743, LPAL2, CDK14, MCOLN2, MEGF11, MRC2, MSX1, NAT8L, NAV1, NBEA, NCRNA00092, NEURL, NGEF, NHLH2, NKD1, NLGN1, NOG, N PPC, NR3C1, NRXN2, OBSCN, OCA2, OSBPL1A, OXR1, P2RY1, PDE2A, PDGFRA, PID1, PLCG2, PLCL1, PNPLA3, POU3F1, POU3F2, PRKA The novel DMR(s) are derived from a gene selected from CB, PRKAR2B, PRKG1, PRR18, PRR5L, PTHLH, PYGL, RHBDL3, RIMS2, ROR2, RPRML, SDK2, SOX9, SYNGR1, TAC4, TRANK1, TSPAN4, UBE2E2, UCHL1, UNC5A, VASH2, ZBTB10, ZNF680, ZNF709, ZNF738, ZNF808, and ZNF845 (Table 2), including any combination thereof. In some embodiments, the novel DMR(s) are derived from any gene selected from Table 2, including any combination thereof. While each novel DMR alone can distinguish HGD-BE and / or EAC from NDBE and / or control samples, combining two or more of the novel DMRs can improve sensitivity. Thus, combinations of two or more novel DMRs selected from Table 2 are provided.

[0008] Embodiments of the present disclosure also include novel variably methylated regions (DMRs), each capable of distinguishing esophageal cancer or precancer from control or benign tissue. In some embodiments, the novel DMRs can distinguish high-grade dysplastic Barrett's esophagus (HGD-BE) or esophageal adenocarcinoma (EAC) from non-dysplastic Barrett's esophagus (NDBE) or normal esophageal controls. According to these embodiments, the novel DMR(s) are derived from genes selected from BACH2, C5orf42, FHOD3, HIST1H2BE, IRX3, KIAA1614, LONRF2, MAFB, PDGFRA, PID1, POU3F1, PRR5L, RHBDL3, and SDK2 (Table 3), including any combination thereof. In some embodiments, the novel DMR(s) are derived from any gene selected from Table 3, including any combination thereof. While each novel DMR alone can distinguish HGD-BE and / or EAC from NDBE and / or control samples, combining two or more of the novel DMRs can improve sensitivity. Thus, combinations of two or more novel DMRs selected from Table 3 are provided.

[0009] Embodiments of the present disclosure also include novel variably methylated regions (DMRs), each capable of distinguishing esophageal cancer or precancer from control or benign tissue. In some embodiments, the novel DMRs can distinguish high-grade dysplastic Barrett's esophagus (HGD-BE) or esophageal adenocarcinoma (EAC) from non-dysplastic Barrett's esophagus (NDBE) or normal esophageal controls. According to these embodiments, the novel DMR(s) are derived from genes selected from KL, PGBD5, ROR2, and LMX1B (Example 4), including any combination thereof. In some embodiments, the novel DMR(s) are derived from any gene selected from Example 4, including any combination thereof. While each novel DMR alone can distinguish HGD-BE and / or EAC from NDBE and / or control samples, combining two or more of the novel DMRs can improve sensitivity. Accordingly, combinations of two or more novel DMRs selected from Example 4 are provided.

[0010] In some embodiments, the novel DMR(s) capable of distinguishing esophageal cancer or precancer from control or benign tissue were validated based on at least one of the area under the receiver operating characteristic curve (AUC), fold change in methylation, percentage of methylation, and / or percentage of hypermethylation between test and control samples using at least one of methylation-specific PCR, quantitative methylation-specific PCR, methylation-specific DNA restriction enzyme analysis, quantitative bisulfite pyrosequencing, flap endonuclease assay, PCR flap assay, and bisulfite genomic sequencing PCR.

[0011] According to the above, the control sample includes a sample from a subject without cancer, a sample from a subject without esophageal cancer, a sample from a subject without esophageal precancer, or a sample from a subject with a type of cancer that is neither esophageal cancer nor precancer. In some embodiments, the control sample includes a sample from a subject with non-dysplastic Barrett's esophagus (NDBE). In some embodiments, the control sample is derived from a tissue sample, a blood sample, a plasma sample, a serum sample, a whole blood sample, a buffy coat sample, a secretion sample, an organ secretion sample, a cerebrospinal fluid (CSF) sample, a saliva sample, a urine sample, and a stool sample. In some embodiments, the tissue sample is an esophageal tissue sample. In some embodiments, the esophageal sample is obtained from an esophageal biopsy or by swabbing, brushing, or using a sponge capsule device.

[0012] In some embodiments, the novel DMR(s) that can distinguish esophageal cancer or precancer from control samples are associated with an area under the ROC curve (AUC) of 0.5 or greater, where the ROC curve distinguishes between subjects having or suspected of having esophageal cancer or precancer and control DNA samples. In some embodiments, the novel DMR(s) that can distinguish esophageal cancer or precancer from control samples are associated with an area under the ROC curve (AUC) of 0.6 or greater, where the ROC curve distinguishes between subjects having or suspected of having esophageal cancer or precancer and control DNA samples. In some embodiments, the novel DMR(s) that can distinguish esophageal cancer or precancer from control samples are associated with an area under the ROC curve (AUC) of 0.7 or greater, where the ROC curve distinguishes between subjects having or suspected of having esophageal cancer or precancer and control DNA samples. In some embodiments, the novel DMR(s) that can distinguish esophageal cancer or precancer from control samples are associated with an area under the ROC curve (AUC) of 0.75 or greater, where the ROC curve distinguishes between subjects having or suspected of having esophageal cancer or precancer and control DNA samples. In some embodiments, the novel DMR(s) that can distinguish esophageal cancer or precancer from control samples are associated with an area under the ROC curve (AUC) of 0.8 or greater, where the ROC curve distinguishes between subjects having or suspected of having esophageal cancer or precancer and control DNA samples. In some embodiments, the novel DMR(s) that can distinguish esophageal cancer or precancer from control samples are associated with an area under the ROC curve (AUC) of 0.85 or greater, where the ROC curve distinguishes between subjects having or suspected of having esophageal cancer or precancer and control DNA samples. In some embodiments, the novel DMR(s) that can distinguish esophageal cancer or precancer from control samples are associated with an area under the ROC curve (AUC) of 0.9 or greater, where the ROC curve distinguishes between subjects having or suspected of having esophageal cancer or precancer and control DNA samples.In some embodiments, the novel DMR(s) that can distinguish esophageal cancer or precancer from control samples are associated with an area under the ROC curve (AUC) of 0.95 or greater, where the ROC curve distinguishes between subjects having or suspected of having esophageal cancer or precancer and control DNA samples.

[0013] In some embodiments, the novel DMR(s) that can distinguish esophageal cancer or precancer from a control sample comprise an increased rate of hypermethylation compared to a control DNA sample. In some embodiments, the novel DMR(s) that can distinguish esophageal cancer or precancer from a control sample comprise an increased rate of hypermethylation compared to a control DNA sample.

[0014] Embodiments of the present disclosure also include methods and compositions for characterizing a biological sample and determining the methylation profile of at least one variably methylated region (DMR) in a DNA sample obtained from a subject having or suspected of having esophageal cancer or precancer by treating the sample with a reagent that modifies DNA in a methylation-specific manner. In some embodiments, the method includes detecting the presence of esophageal cancer or precancer from the biological sample. In some embodiments, the at least one DMR can distinguish high-grade dysplastic Barrett's esophagus (HGD-BE) or esophageal adenocarcinoma (EAC) from non-dysplastic Barrett's esophagus (NDBE) or normal esophageal controls.

[0015] According to these embodiments, the method comprises determining copy number variation (CNV) for DNA samples from subjects.In some embodiments, the CNV distinguishes between subjects who have or are suspected of having high-grade dysplastic Barrett's esophagus or esophageal adenocarcinoma (EAC) and control DNA samples.In some embodiments, at least one DMR comprises increased CNV compared to control DNA samples.

[0016] In some embodiments, the method comprises determining the aneuploidy score (AS) for DNA samples from subjects.In some embodiments, the AS distinguishes between subjects who have or are suspected of having high-grade dysplastic Barrett's esophagus or esophageal adenocarcinoma (EAC) and control DNA samples.In some embodiments, at least one DMR comprises increased AS compared with control DNA samples.

[0017] As further described herein, the present disclosure provides materials and methods for determining copy number abnormalities (CNAs), including ploidy and aneuploidy (e.g., determining aneuploidy scores), using sequencing reads from genomes modified in a methylation-specific manner, e.g., cytosine- or 5-methylcytosine-converted genomes. Most currently available technologies use NGS directly on wild-type, unconverted DNA. However, embodiments of the present disclosure include the ability to perform methylation and CNV analysis from the same chemistry / dataset. According to these embodiments, the methylation profile at at least one DMR and CNV and / or AS can be determined using the same DNA sample obtained from a subject. In some embodiments, the methylation profile at at least one DMR and CNV and / or AS can be determined using a single DNA sample obtained from a subject. In some embodiments, the samples (e.g., the same sample or a single sample) have been treated with reagents that modify DNA in a methylation-specific manner.

[0018] In some embodiments, the control DNA sample used in the methods and compositions of the present disclosure is derived from a subject without esophageal cancer or precancer. In some embodiments, the control DNA sample is derived from a subject with non-dysplastic Barrett's esophagus (NDBE). In some embodiments, the control DNA sample is selected from a tissue sample, a blood sample, a plasma sample, a serum sample, a whole blood sample, a buffy coat sample, a secretion sample, an organ secretion sample, a cerebrospinal fluid (CSF) sample, a saliva sample, a urine sample, and a stool sample. In some embodiments, the tissue sample is an esophageal tissue sample or a gastric cardia sample.

[0019] In some embodiments, the biological sample is obtained from a human subject. In some embodiments, the biological sample is obtained from a human subject, and the method comprises extracting a DNA sample from the biological sample. In some embodiments, the biological sample is collected with a collection device. For example, the biological sample can be an esophageal sample obtained from an esophageal biopsy or by using a swabbing, brushing, or sponge capsule device.

[0020] In some embodiments, the methods of the present disclosure include using a reagent that modifies DNA in a methylation-specific manner. In some embodiments, the reagent is a borane reducing agent. In some embodiments, the reagent that modifies DNA in a methylation-specific manner includes one or more of a methylation-sensitive restriction enzyme, a methylation-dependent restriction enzyme, and a bisulfite reagent.

[0021] In some embodiments, determining the methylation profile of at least one DMR comprises amplifying at least a portion of the DMR using a set of primers. In some embodiments, determining the methylation profile of at least one DMR comprises performing at least one of methylation-specific PCR, quantitative methylation-specific PCR, methylation-specific DNA restriction enzyme analysis, quantitative bisulfite pyrosequencing, flap endonuclease assay, PCR-flap assay, and bisulfite genomic sequencing PCR. In some embodiments, determining the methylation profile of at least one DMR comprises determining the presence or absence of methylation at CpG sites. In some embodiments, the one or more CpG sites are present in a coding region, a non-coding region, and / or a regulatory region of a gene (e.g., any one of the genes disclosed herein).

[0022] Embodiments of the present disclosure also include methods for identifying esophageal cancer or precancer. According to these embodiments, the method includes determining a methylation profile in at least one variably methylated region (DMR) of a DNA sample obtained from a subject having or suspected of having esophageal cancer or precancer by treating the DNA with a reagent that modifies DNA in a methylation-specific manner. In some embodiments, the methylation profile indicates that the subject has esophageal cancer (e.g., esophageal adenocarcinoma (EAC)) or esophageal precancer (e.g., high-grade dysplastic Barrett's esophagus (HGD-BE)). In some embodiments, the method includes determining copy number variation (CNV) for the DNA sample from the subject. In some embodiments, the method includes determining an aneuploidy score (AS) for the DNA sample from the subject. In some embodiments, the method further includes treating the subject with an anti-cancer therapy. [Brief explanation of the drawings]

[0023] [Figure 1] Representative bar graph of total CNV results at the level of chromosome arm gains and losses. (Higher signals indicate greater degrees of chromosomal instability and aneuploidy.) [Figure 2] Representative violin plot of total CNV results at the level of chromosome arm gains and losses. (Higher signals indicate greater degrees of chromosomal instability and aneuploidy.) [Figure 3] Representative chromosome scatter plot illustrating normal euploidy (two copies of a chromosome / gene) for the normal esophagus (NE) cohort generated from whole-genome sequencing data. Each point represents a binned genomic segment. [Figure 4] Representative chromosome scatter plot showing normal euploidy (two copies of a chromosome / gene) for a non-dysplastic Barrett's esophagus (NDBE) cohort generated from whole-genome sequencing data. Each point represents a binned genomic segment. [Figure 5]Representative chromosome scatter plots showing aneuploidy for the high-grade dysplasia (HGD) cohort. These were generated from whole-genome NGS data. Each point represents a binned genomic segment. [Figure 6] Representative chromosome scatter plots showing aneuploidy for the esophageal adenocarcinoma (EAC) cohort. These were generated from whole-genome NGS data. Each point represents a binned genomic segment. [Figure 7] Representative heatmap of DMRs from the MAFB gene, a top methylation marker candidate identified in this disclosure, for high-grade dysplasia (HGD) and esophageal adenocarcinoma (EAC) cohorts. Columns represent samples, and rows represent individual CpGs in genomic order that make up the DMR. Dark green indicates no methylation, and increasing shades of red indicate increasing methylation intensity. [Figure 8] Representative heatmap of DMRs from the MAFB gene, a top candidate methylation marker identified in this disclosure, for normal esophagus (NE) and non-dysplastic Barrett's esophagus (NDBE) cohorts. Columns represent samples, and rows represent individual CpGs, in genomic order, that make up the DMR. Dark green indicates no methylation, and increasing shades of red indicate increasing methylation intensity. [Figure 9] Representative heatmap showing complementarity between CNV (aneuploidy) and methylation analysis among esophageal adenocarcinoma (EAC), high-grade dysplasia (HGD), and non-dysplastic Barrett's esophagus (NDBE) cohorts. DETAILED DESCRIPTION OF THE INVENTION

[0024] Barrett's esophagus (BE) is the strongest risk factor and the only known precursor of esophageal adenocarcinoma (EAC), a lethal malignancy with poor survival rates (<20% at 5 years) if detected after the onset of symptoms. The incidence of esophageal adenocarcinoma has increased by nearly 600% over the past 30 years in the general population. BE progresses to EAC via a stepwise pathway from no dysplasia (also known as non-dysplastic Barrett's esophagus or NDBE) to low-grade dysplasia (LGD), high-grade dysplasia (HGD), and then carcinoma. This metaplasia-to-dysplasia-to-carcinoma sequence has prompted several national gastroenterological societies to recommend BE screening in high-risk subjects with multiple risk factors, followed by endoscopic surveillance (depending on the grade of dysplasia) to detect early development of dysplasia or carcinoma. Endoscopic treatments for LGD, HGD, and early carcinoma have been developed and have been shown to be effective in reducing the incidence of carcinoma and improving survival in subjects with BE.

[0025] Screening for BE is currently performed using conventional sedated endoscopy (sEGD), which reveals that the normal squamous epithelial lining of the esophagus is replaced by columnar metaplasia in subjects with BE. However, sedated endoscopy is expensive in both direct and indirect costs, making it unsuitable for widespread application. It is also associated with potential complications. Other techniques, such as unsedated transnasal endoscopy (uTNE), are less costly and have similar accuracy to sEGD, but provider acceptance as a widely applicable tool remains low. Despite adequate access to uTNE devices, utilization of uTNE by referring physicians remains limited. The lack of accurate risk stratification tools to determine BE risk and targeted screening efforts is a further limitation to widely applicable BE screening.

[0026] Endoscopic detection of dysplasia is currently performed using four quadrant random biopsies every 1–2 cm of the BE segment, in addition to careful examination of the BE segment with high-resolution white-light imaging and advanced imaging techniques. Although this is recommended by the GI Society, compliance with these recommendations among practicing gastroenterologists remains low. In fact, compliance decreases as the length of the BE segment increases, leading to an increased rate of missed dysplasia. Other challenges in detecting dysplasia in BE include the punctate distribution of dysplasia in BE, which leads to sampling error, poor interobserver agreement between pathologists while classifying dysplasia, and the relatively low sensitivity of current surveillance strategies in detecting prevalent dysplasia or carcinoma. The utility of advanced imaging technologies in the community remains unclear, with only one-third of practicing gastroenterologists reporting their routine use in BE surveillance. Recently, a sponge-on-a-string device has been investigated for BE screening. This device consists of a polyurethane foam sponge compressed within a gelatin capsule and attached to a string. The capsule is swallowed by the patient. The capsule's gelatin shell dissolves in gastric juice, releasing the foam device as a sphere, which is then withdrawn with an attached string to provide brushing / cytology samples of the proximal stomach and esophagus. Biomarker studies can then be performed on these samples to detect BE.

[0027] As further described herein, BE is a metaplastic change in the distal esophageal epithelial lining, characterized by the replacement of normal squamous epithelium (NE) with specific intestinal metaplasia. The presence of Barrett's esophagus increases the risk of esophageal adenocarcinoma several-fold. This disclosure addresses the detection of high-grade dysplasia (HGD) Barrett's esophagus, including adenocarcinoma (EAC), which is distinct from non-dysplastic Barrett's esophagus (NDBE). Knowing the dysplasia status of Barrett's patients is important for follow-up clinical surveillance and treatment strategies. As further described herein, RRBS analysis was performed on tissue biopsies and whole-genome sequencing was performed on endoscopic brushing samples to generate methylation and copy number mutation profiles of BE patients. Using a proprietary analytical algorithm, 199 methylated DNA markers (MDMs) were identified, 156 of which were confirmed in both NGS datasets. Furthermore, aneuploidy scores were developed using WGS data, and these scores demonstrated how this marker class complements MDM analysis. For aneuploidy / CNV analysis, calls were made using the converted genome from the methylation analysis, an approach not currently used. Therefore, this technology can be implemented in a clinical trial format for use with endoscopic brushing samples and non-endoscopic, non-invasive esophageal sponge samples.

[0028] The section headings used in this section and throughout this disclosure are for organizational purposes only and are not intended to be limiting.

[0029] 1.Definition Throughout the specification and claims, the following terms have the meanings expressly associated therewith, unless the context clearly dictates otherwise. As used herein, the phrase "in one embodiment" may refer to the same embodiment, but does not necessarily refer to the same embodiment. Furthermore, as used herein, the phrase "in another embodiment" may refer to a different embodiment, but does not necessarily refer to a different embodiment. Thus, as described below, various embodiments of the invention can be readily combined without departing from the scope or spirit of the invention.

[0030] Additionally, as used herein, the term "or" is an inclusive "or" operator and is synonymous with the term "and / or" unless the context clearly dictates otherwise. The term "based on" is not exclusive and acknowledges that a relationship may be based on additional unlisted factors unless the context clearly dictates otherwise. Additionally, throughout this specification, the meanings of "a," "an," and "the" include plural referents. The meaning of "in" includes "in" and "on."

[0031] The transitional phrase "consisting essentially of," when used in the claims of this application, limits the scope of the claim to the particular materials or steps of the claimed invention and "which do not materially affect the basic and novel characteristic(s)," as discussed in In re Herz, 537 F.2d 549,551-52,190 USPQ 461,463 (CCPA 1976). For example, a composition "consisting essentially of" the recited elements may contain unrecited contaminants at levels such that, although the contaminants are present, the contaminants do not alter the function of the recited composition compared to the pure composition (i.e., a composition "consisting of" the recited components).

[0032] The term "one or more," as used herein, refers to a number greater than 1. For example, the term "one or more" includes any of the following: 2 or more, 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, 10 or more, 12 or more, 13 or more, 14 or more, 15 or more, 20 or more, 50 or more, 100 or more, or even greater numbers.

[0033] The terms "one or more, but less than a larger number," "two or more, but less than a larger number," "three or more, but less than a larger number," "four or more, but less than a larger number," "five or more, but less than a larger number," "six or more, but less than a larger number," "seven or more, but less than a larger number," "eight or more, but less than a larger number," "nine or more, but less than a larger number," "ten or more, but less than a larger number," "eleven or more, but less than a larger number," "twelve or more, but less than a larger number," "thirteen or more, but less than a larger number," "fourteen or more, but less than a larger number," or "fifteen or more, but less than a larger number" are not limited to the larger number. For example, a larger number can be 10,000, 1,000, 100, 50, etc. For example, some higher number can be about 50 (e.g., 50, 49, 48, 47, 46, 45, 44, 43, 42, 41, 40, 39, 38, 37, 36, 35, 34, 33, 32, 31, 32, 30, 29, 28, 27, 26, 25, 24, 23, 22, 21, 20, 19, 18, 17, 16, 15, 14, 13, 12, 11, 10, 9, 8, 7, 6, 5, 4, 3, or 2).

[0034] The terms "one or more methylation markers" or "one or more DMRs" or "one or more genes" or "one or more markers" or "multiple methylation markers" or "multiple markers" or "multiple genes" or "multiple DMRs" are likewise not limited to a specific numerical combination. In fact, any numerical combination of methylation markers is envisaged (e.g., 1–2 methylation markers, 1–3, 1–4, 1–5, 1–6, 1–7, 1–8, 1–9, 1–10, 1–11, 1–12, 1–13, 1–14, 1–15, 1–16, 1–17, 1–18, 1–19, 1–20, 1–21, 1–22, 1–23, 1–24, 1–25, 1–26, 1–27, 1–28, 1–29, 1–30, 1–31, 1–32, 1–33, 1–34, 1–35, 1–36, 1–37, 1–38, 1–39, 1–40, 1–41, 1–42, 1–43, 1–44, 1–45, 1–46, 1–47, 1–48, 1–49, 1–50, 1–51, 1–52, 1–53, 1–54, 1–55, 1–56, 1–57, 1–58, 1–59, 1–60, 1–61, 1–62, 1–63, 1–64, 1–65, 1–66, 1–67, 1–68, 1–69, 1–70, 1–71, 1–72, 1–73, 1–74, 1–75, 1–76, 1–77, 1–78, 1–79, 1–80, 1–81, 1–82, 1–8 7, 1-38) (e.g., 2-3, 2-4, 2-5, 2-6, 2-7, 2-8, 2-9, 2-10, 2-11, 2-12, 2-13, 2-14, 2-15, 2-16, 2-17, 2-18, 2-19, 2-20, 2-21, 2-22, 2-23, 2-24, 2-25, 2-26, 2-27, 2-28, 2-29, 2-30, 2-31, 2-32, 2-33, 2-34, 2-35, 2-36, 2-37, 2-38) (e.g., 3-4, 3-5, 3-6, 3-7, 3-8, 3-9, 3-1 0, 3-11, 3-12, 3-13, 3-14, 3-15, 3-16, 3-17, 3-18, 3-19, 3-20, 3-21, 3-22, 3-23, 3-24, 3-25, 3-26, 3-27, 3-28, 3-29, 3-30, 3-31, 3-32, 3-33, 3-34, 3-35, 3-36, 3-37, 3-38) (e.g., 4-5, 4-6, 4-7, 4-8, 4-9, 4-10, 4-11, 4-12, 4-13, 4-14, 4-15, 4-16, 4-17, 4-18, 4-19) , 4-20, 4-21, 4-22, 4-23, 4-24, 4-25, 4-26, 4-27, 4-28, 4-29, 4-30, 4-31, 4-32, 4-33, 4-34, 4-35, 4-36, 4-37, 4-38) (e.g., 5-6, 5-7, 5-8, 5-9, 5-10, 5-11, 5-12, 5-13, 5-14, 5-15, 5-16, 5-17, 5-18, 5-19, 5-20, 5-21, 5-22, 5-23, 5-24, 5-25, 5-26, 5-27, 5-28, 5-29,5-30, 5-31, 5-32, 5-33, 5-34, 5-35, 5-36, 5-37, 5-38) (e.g., 6-7, 6-8, 6-9, 6-10, 6-11, 6-12, 6-13, 6-14, 6-15, 6-16, 6-17, 6-18, 6-19, 6-20, 6-21, 6- 22, 6-23, 6-24, 6-25, 6-26, 6-27, 6-28, 6-29, 6-30, 6-31, 6-32, 6-33, 6-34, 6-35, 6-36, 6-37, 6-38) (e.g., 7-8, 7-9, 7-10, 7-11, 7-12, 7-13, 7-14, 7-15 , 7-16, 7-17, 7-18, 7-19, 7-20, 7-21, 7-22, 7-23, 7-24, 7-25, 7-26, 7-27, 7-28, 7-29, 7-30, 7-31, 7-32, 7-33, 7-34, 7-35, 7-36, 7-37, 7-38) (e.g., 8-9 , 8-10, 8-11, 8-12, 8-13, 8-14, 8-15, 8-16, 8-17, 8-18, 8-19, 8-20, 8-21, 8-22, 8-23, 8-24, 8-25, 8-26, 8-27, 8-28, 8-29, 8-30, 8-31, 8-32, 8-33, 8-34 , 8-35, 8-36, 8-37, 8-38) (e.g., 9-10, 9-11, 9-12, 9-13, 9-14, 9-15, 9-16, 9-17, 9-18, 9-19, 9-20, 9-21, 9-22, 9-23, 9-24, 9-25, 9-26, 9-27, 9-28, 9-29, 9, 9-30, 9-31, 9-32, 9-33, 9-34, 9-35, 9-36, 9-37, 9-38) (e.g., 10-11, 10-12, 10-13, 10-14, 10-15, 10-16, 10-17, 10-18, 10-19, 10-20, 10-21, 10-22, 1 0-23, 10-24, 10-25, 10-26, 10-27, 10-28, 10-29, 10-30, 10-31, 10-32, 10-33, 10-34, 10-35, 10-36, 10-37, 10-38) (e.g., 11-12, 11-13, 11-14, 11-15, 1 1-16, 11-17, 11-18, 11-19, 11-20, 11-21, 11-22, 11-23, 11-24, 11-25, 11-26, 11-27, 11-28, 11-29, 11-30, 11-31, 11-32, 11-33, 11-34, 11-35, 11-36,11-37, 11-38) (e.g., 12-13, 12-14, 12-15, 12-16, 12-17, 12-18, 12-19, 12-20, 12-21, 12-22, 12-23, 12-24, 12-25, 12-26, 12-27, 12-28, 12-29, 12-30, 12-31, 12-32, 12-33, 12-34, 12-35, 12-36, 12-37, 12-38) (e.g., 13-14, 13-15, 13-16, 13-17, 13-18, 13-19, 13-20, 13-21, 13-22, 13-23, 13-24, 13-25, 13-26, 13-27, 13-28, 13-29, 13-30, 13-31, 13-32, 13-33, 13-34, 13-35, 13-36, 13-37, 13-38) (e.g., 14-15, 14-16, 14-17, 14-18, 14-19, 14-20, 14-21, 14-22, 14-23, 14-24, 14-25, 14-26, 14-27, 14-28, 14-29, 14-30, 14-31, 14-32, 14-33, 14-34, 14-35, 14-36, 14-37, 14-38) (e.g., 15-16, 15-17, 15-18, 15-19, 15-20, 15-21, 15-22, 15-23, 15-24, 15-25, 15-26, 15-27, 15-28, 15-29, 15-30, 15-31, 15-32, 15-33, 15-34, 15-35, 15-36, 15-37, 15-38) (e.g., 16-17, 16-18, 16-19, 16-20, 16-21, 16-22, 16-23, 16-24, 16-25, 16-26, 16-27, 16-28, 16-29, 16-30, 16-31, 16-32, 16-33, 16-34, 16-35, 16-36, 16-37 , 16-38) (e.g., 17-18, 17-19, 17-20, 17-21, 17-22, 17-23, 17-24, 17-25, 17-26, 17-27, 17-28, 17-29, 17-30, 17-31, 17-32, 17-33, 17-34, 17-35, 17-36 , 17-37, 17-38) (e.g., 18-19, 18-20, 18-21, 18-22, 18-23, 18-24, 18-25, 18-26, 18-27, 18-28, 18-29, 18-30, 18-31, 18-32, 18-33, 18-34, 18-35, 18-36,18-37, 18-38) (e.g., 19-20, 19-21, 19-22, 19-23, 19-24, 19-25, 19-26, 19-27, 19-28, 19-29, 19-30, 19-31, 19-32, 19-33, 19-34, 19-35, 19-36, 19-37, 19-38) (e.g., 20-21, 20-22, 20-23, 20-24, 20-25, 20-26, 20-27, 20-28, 20-29, 20-30, 20-31, 20-32, 20-33, 20-34, 20-35, 20-36, 20-37, 20-38) (e.g., 21-22, 21-23, 21-24, 21-25, 21-26, 21-27, 21-28, 21-29, 21-30, 21-31, 21-32, 21-33, 21-34, 21-35, 21-36, 21-37, 21-38) (e.g., 22-23, 22-24, 22-25 , 22-26, 22-27, 22-28, 22-29, 22-30, 22-31, 22-32, 22-33, 22-34, 22-35, 22-36, 22-37, 22-38) (e.g., 23-24, 23-25, 23-26, 23-27, 23-28, 23-29, 23-30 , 23-31, 23-32, 23-33, 23-34, 23-35, 23-36, 23-37, 23-38) (e.g., 24-25, 24-26, 24-27, 24-28, 24-29, 24-30, 24-31, 24-32, 24-33, 24-34, 24-35, 24-36, 24-37, 24-38) (e.g., 25-26, 25-27, 25-28, 25-29, 25-30, 25-31, 25-32, 25-33, 25-34, 25-35, 25-36, 25-37, 25-38) (e.g., 26-27, 26-28, 26-29, 26-30 , 26-31, 26-32, 26-33, 26-34, 26-35, 26-36, 26-37, 26-38) (e.g., 27-28, 27-29, 27-30, 27-31, 27-32, 27-33, 27-34, 27-35, 27-36, 27-37, 27-38) (e.g., 28-29, 28-30, 28-31, 28-32, 28-33, 28-34, 28-35, 28-36, 28-37, 28-38) (e.g., 29-30, 29-31, 29-32, 29-33, 29-34, 29-35, 29-36, 29-37, 29-38) (e.g.,30-31, 30-32, 30-33, 30-34, 30-35, 30-36, 30-37, 30-38) (e.g., 31-32, 31-33, 31-34, 31-35, 31-36, 31-37, 31-38) (e.g., 32-33, 32-34, 32-35, 32-36, 32-37, 32-38) (e.g., 33-34, 33-35, 33-36, 33-37, 33-38) (e.g., 34-35, 34-36, 34-37, 34-38) (e.g., 35-36, 35-37, 3 5-38) (e.g., 36-37, 36-38) (e.g., 37-38) (e.g., 38 or less; 37 or less; 36 or less; 35 or less; 34 or less; 33 or less; 32 or less; 31 or less; 30 or less; 29 or less; 28 or less; 27 or less; 26 or less; 25 or less; 24 or less; 23 or less; 22 or less; 21 or less; 20 or less; 19 or less; 18 or less; 17 or less; 16 or less; 15 or less; 14 or less; 13 or less; 12 or less; 11 or less; 10 or less; 9 or less; 8 or less; 7 or less; 6 or less; 5 or less; 4 or less; 3 or less; 2, or 1).

[0035] As used herein, "nucleic acid" or "nucleic acid molecule" generally refers to any ribonucleic acid or deoxyribonucleic acid, which may be unmodified or modified DNA or RNA. "Nucleic acid" includes, but is not limited to, single-stranded and double-stranded nucleic acids. As used herein, the term "nucleic acid" also includes DNA, as described above, containing one or more modified bases. Thus, DNA with backbone modifications for stability or other reasons is a "nucleic acid." As used herein, the term "nucleic acid" encompasses chemically, enzymatically, or metabolically modified forms of nucleic acid, as well as chemical forms of DNA characteristic of viruses and cells, including, for example, simple and complex cells.

[0036] The terms "oligonucleotide," "polynucleotide," "nucleotide," or "nucleic acid" refer to a molecule containing two or more deoxyribonucleotides or ribonucleotides, preferably three or more, and usually ten or more. The exact size will depend on many factors and is dependent on the ultimate function or use of the oligonucleotide. Oligonucleotides can be produced in any manner, including chemical synthesis, DNA replication, reverse transcription, or a combination thereof. Typical deoxyribonucleotides of DNA are thymine, adenine, cytosine, and guanine. Typical ribonucleotides of RNA are uracil, adenine, cytosine, and guanine.

[0037] As used herein, the term "locus" or "region" of a nucleic acid refers to a small region of nucleic acid, e.g., a gene on a chromosome, a single nucleotide, a CpG island, and the like.

[0038] The terms "complementary" and "complementarity" refer to nucleotides (e.g., single nucleotides) or polynucleotides (e.g., nucleotide sequences) related by the base-pairing rules. For example, the sequence 5'-AGT-3' is complementary to the sequence 3'-TCA-5'. Complementarity can be "partial," in which only a portion of the nucleic acid bases match according to the base-pairing rules. Alternatively, there can be "complete" or "total" complementarity between nucleic acids. The degree of complementarity between nucleic acid strands determines the efficiency and strength of hybridization between nucleic acid strands. This is particularly important in amplification reactions and detection methods that rely on binding between nucleic acids.

[0039] The term "gene" refers to a nucleic acid (e.g., DNA or RNA) sequence that comprises coding sequences necessary for the production of an RNA or polypeptide or its precursor. A functional polypeptide can be encoded by a full-length coding sequence or by any portion of the coding sequence, so long as the desired activity or functional property of the polypeptide (e.g., enzymatic activity, ligand binding, signal transduction, etc.) is retained. When used in reference to a gene, the term "portion" refers to fragments of that gene. These fragments can range in size from a few nucleotides to the entire gene sequence minus one nucleotide. Thus, "nucleotides comprising at least a portion of a gene" can include fragments of a gene or the entire gene.

[0040] The term "gene" includes the coding region of a structural gene and includes sequences located adjacent to the coding region at both the 5' and 3' ends (e.g., including coding, regulatory, structural, and other sequences) such that the gene corresponds to the length of the full-length mRNA. Sequences located 5' of the coding region and present on the mRNA are referred to as 5' untranslated or non-translated sequences. Sequences located 3' or downstream of the coding region and present on the mRNA are referred to as 3' untranslated or 3' non-translated sequences. The term "gene" encompasses both cDNA and genomic forms of a gene. In some organisms (e.g., eukaryotes), genomic forms or clones of a gene contain coding regions interrupted by non-coding sequences termed "introns" or "intervening regions" or "intervening sequences." Introns are segments of a gene that are transcribed into nuclear RNA (hnRNA) and may contain regulatory elements such as enhancers. Introns are removed or "spliced ​​out" from the nuclear or primary transcript; therefore, introns are absent in the messenger RNA (mRNA) transcript. mRNA functions during translation to specify the sequence or order of amino acids in a nascent polypeptide. As will be understood by those skilled in the art based on the present disclosure, one or more CpG sites in a DMR may be located in a coding region not known to be associated with a particular gene, such as a coding region of a gene, a non-coding regulatory region of a gene, or a region containing long non-coding RNA (lncRNA). In some embodiments, sequences corresponding to these regions can be obtained using the corresponding accession numbers (see, e.g., Table 1) in genome databases (e.g., GenBank, NCBI / Ensembl, UniProt, etc.). In some embodiments, one or more CpG sites in a DMR may be located within an unannotated genomic region. As further provided herein, an unannotated genomic region containing one or more CpG sites within a DMR can be described using a sequence number (see, e.g., Table 1, SEQ ID NOS: 1-6).

[0041] As will be understood by one of skill in the art based on the present disclosure, the location of one or more CpG sites within a gene or region (e.g., a CpG island) and its association with a disease or condition can be determined using a variety of techniques, including, but not limited to, those described in Chen et al., "Methods for identifying differentially methylated regions for sequence- and array-based data," Briefings in Functional Genomics, Volume 15, Issue 6, November 2016, Pages 485-490, which is incorporated herein by reference in its entirety for all purposes.

[0042] The term "wild-type," when used in reference to a gene, refers to a gene having the characteristics of a gene isolated from a naturally occurring source. The term "wild-type," when used in reference to a gene product, refers to a gene product having the characteristics of a gene product isolated from a naturally occurring source. The term "wild-type," when used in reference to a protein, refers to a protein having the characteristics of a naturally occurring protein. The term "naturally occurring," when applied to an object, refers to the fact that the object can be found in nature. For example, a polypeptide or polynucleotide sequence present in an organism (including a virus) that can be isolated from a natural source and has not been intentionally modified by human hands in the laboratory is naturally occurring. A wild-type gene is often that gene or allele that is most frequently observed in a population and is therefore arbitrarily referred to as the "normal" or "wild-type" form of the gene. In contrast, when referring to a gene or gene product, the terms "modified" or "mutant" refer to a gene or gene product, respectively, that exhibits modifications in sequence and / or functional properties (i.e., altered characteristics) when compared to the wild-type gene or gene product. Note that naturally occurring variants can be isolated. These are identified by the fact that they have altered properties when compared to the wild-type gene or gene product.

[0043] The term "allele" refers to a genetic variation, including, without limitation, variants and mutations, polymorphic and single nucleotide polymorphic loci, frameshifts, and splice variants. Alleles may occur naturally within a population or may arise during the lifetime of any particular individual in a population.

[0044] Thus, when used in reference to a nucleotide sequence, the terms "variant" and "mutant" refer to a nucleic acid sequence that differs by one or more nucleotides from another, usually related, nucleotide sequence. A "mutation" is a difference between two different nucleotide sequences, typically one sequence being a reference sequence.

[0045] The term "primer" refers to an oligonucleotide, whether naturally occurring as a nucleic acid fragment from a purified restriction digest or synthesized, that can act as a point of initiation of synthesis when placed under conditions that induce synthesis of a primer extension product complementary to a template strand of nucleic acid (e.g., in the presence of nucleotides and an inducing agent such as DNA polymerase, at a suitable temperature and pH). Primers are preferably single-stranded to maximize amplification efficiency, but may alternatively be double-stranded. If double-stranded, the primer is first treated to separate its strands before being used to prepare extension products. Preferably, the primer is an oligodeoxyribonucleotide. The primer must be sufficiently long to prime the synthesis of an extension product in the presence of the inducing agent. The exact length of the primer will depend on many factors, including temperature, primer source, and the use of the method. In some embodiments, the primer pair is specific for a variably methylated region (e.g., a DMR in Tables 1, 2, and 3) and specifically binds to at least a portion of a gene region that includes the DMR.

[0046] The term "probe" refers to an oligonucleotide (e.g., a series of nucleotides) capable of hybridizing to another oligonucleotide of interest, whether naturally occurring, as in a purified restriction digest, or produced synthetically, recombinantly, or by PCR amplification. Probes can be single-stranded or double-stranded. Probes are useful for the detection, identification, and isolation of specific gene sequences (e.g., "capture probes"). It is contemplated that any probe used in embodiments of the present disclosure can, in some embodiments, be labeled with any "reporter molecule" and thus be detectable in any detection system, including, but not limited to, enzymatic (e.g., ELISA and enzyme-based histochemical assays), fluorescent, radioactive, and luminescent systems. It is not intended that the various embodiments of the present disclosure be limited to any particular detection system or label.

[0047] The term "target," as used herein, refers to a nucleic acid that is sought to be sorted out from other nucleic acids, e.g., by probing, amplifying, isolating, capturing, etc. For example, when used in reference to the polymerase chain reaction, "target" refers to the region of nucleic acid bound by the primers used in the polymerase chain reaction; however, in some embodiments of assays in which the target DNA is not amplified, e.g., invasive cleavage assays, the target includes the site where the probe and invasive oligonucleotide (e.g., INVADER oligonucleotide) bind to form an invasive cleavage structure such that the presence of the target nucleic acid can be detected. A "segment" is defined as a region of nucleic acid within the target sequence.

[0048] Thus, as used herein, "non-target," when used to describe a nucleic acid such as, for example, DNA, refers to a nucleic acid that may be present in a reaction but is not the subject of detection or characterization by the reaction. In some embodiments, non-target nucleic acid can refer to a nucleic acid present in a sample that does not contain, for example, a target sequence, although in some embodiments, non-target can also refer to an exogenous nucleic acid, i.e., a nucleic acid that is not derived from a sample that contains or is suspected of containing a target nucleic acid, and that is added to a reaction, for example, to reduce variability in the performance of an enzyme (e.g., a polymerase) in the reaction, to normalize the activity of the enzyme.

[0049] As used herein, "methylation" refers to cytosine methylation at the C5 or N4 position of cytosine, the N6 position of adenine, or other types of nucleic acid methylation. In vitro amplified DNA is typically unmethylated, since typical in vitro DNA amplification methods do not preserve the methylation pattern of the amplified template. However, "unmethylated DNA" or "methylated DNA" can also refer to amplified DNA in which the original template was unmethylated or methylated, respectively.

[0050] As used herein, the term "amplification reagents" refers to the reagents needed for amplification (deoxyribonucleoside triphosphates, buffers, etc.) excluding primers, nucleic acid template, and amplification enzymes. Typically, amplification reagents are placed and contained within a reaction vessel along with other reaction components.

[0051] As used herein, the term "control," when used in reference to nucleic acid detection or analysis, refers to a nucleic acid with known characteristics (e.g., known sequence, known copy number per cell) used for comparison with an experimental target (e.g., a nucleic acid of unknown concentration, etc.). A control may be an endogenous, preferably invariant, gene to which a test or target nucleic acid in an assay can be normalized. Such normalization controls for sample-to-sample variations that may arise, for example, from sample processing, assay efficiency, etc., allowing for accurate data comparison between samples. Genes useful for normalizing nucleic acid detection assays for human samples include, for example, b-actin, ZDHHC1, and B3GALT6 (see, e.g., U.S. Patent Application Nos. 14 / 966,617 and 62 / 364,082, each of which is incorporated herein by reference). As used herein, "ZDHHC1" refers to a gene located on chromosome 16 (16q22.1) of human DNA that encodes a protein belonging to the DHHC palmitoyltransferase family, characterized as zinc finger DHHC-type containing 1.

[0052] A control can also be an external control. For example, in quantitative assays such as qPCR and QuARTS, a "calibrator" or "calibration control" is a nucleic acid of known sequence, e.g., a portion of an experimental target nucleic acid, with a known concentration or series of concentrations (e.g., a control target serially diluted for generating a calibration curve in quantitative PCR). Typically, the calibration control is analyzed using the same reagents and reaction conditions as those used for the experimental DNA. In certain embodiments, the measurement of the calibrator is performed simultaneously with the experimental assay, e.g., in the same thermal cycler. In preferred embodiments, multiple calibrators can be included in a single plasmid, so that different calibrator sequences can be easily provided in equimolar amounts. In particularly preferred embodiments, the plasmid calibrator is digested, e.g., with one or more restriction enzymes, to release the calibrator portion from the plasmid vector. See, e.g., WO2015 / 066695, incorporated herein by reference.

[0053] As used herein, "methylated nucleotide" or "methylated nucleotide base" refers to the presence of a methyl moiety on a nucleotide base, which is not present in recognized typical nucleotide bases.For example, cytosine does not contain a methyl moiety on its pyrimidine ring, but 5-methylcytosine contains a methyl moiety at the 5th position of its pyrimidine ring.Therefore, cytosine is not a methylated nucleotide, but 5-methylcytosine is a methylated nucleotide.In another example, thymine contains a methyl moiety at the 5th position of its pyrimidine ring, but because thymine is a typical nucleotide base of DNA, for the purposes of this specification, thymine is not considered a methylated nucleotide when present in DNA.

[0054] As used herein, a "methylated nucleic acid molecule" refers to a nucleic acid molecule that contains one or more methylated nucleotides.

[0055] As used herein, the "methylation state," "methylation profile," and "methylation status" of a nucleic acid molecule refer to the presence or absence of one or more methylated nucleotide bases in a nucleic acid molecule. For example, a nucleic acid molecule containing a methylated cytosine is considered to be methylated (e.g., the methylation state of the nucleic acid molecule is methylated). A nucleic acid molecule that does not contain any methylated nucleotides is considered to be unmethylated.

[0056] As used herein, the term "methylation level" applied to a methylation marker refers to the amount of methylation within a particular methylation marker. Methylation level may also refer to the amount of methylation within a particular methylation marker compared to an established standard or control. Methylation level may also refer to whether one or more cytosine residues present in a CpG context have a methylation group. Methylation level may also refer to the proportion of cells in a sample that have or do not have a methylation group at such cytosine. Methylation level may also represent whether a single CpG dinucleotide is methylated.

[0057] The methylation state of a particular nucleic acid sequence (e.g., a genetic marker, or a DNA region, as described herein) can indicate the methylation state of all bases in the sequence, or it can indicate the methylation state of a subset of these bases (e.g., one or more cytosines) within the sequence, or it can indicate information about the methylation density of a region within the sequence, with or without providing precise information about the position within the sequence where methylation occurs.

[0058] The methylation state of a nucleotide locus in a nucleic acid molecule refers to the presence or absence of a methylated nucleotide at a particular locus in the nucleic acid molecule. For example, the methylation state of the cytosine at the seventh nucleotide in a nucleic acid molecule is methylated if the nucleotide present at the seventh nucleotide in the nucleic acid molecule is 5-methylcytosine. Similarly, the methylation state of the cytosine at the seventh nucleotide in a nucleic acid molecule is unmethylated if the nucleotide present at the seventh nucleotide in the nucleic acid molecule is cytosine (and not 5-methylcytosine).

[0059] Methylation status can optionally be expressed or indicated by a "methylation value" (e.g., representing a methylation frequency, fraction, proportion, percent, etc.). Methylation values ​​can be generated, for example, by quantifying the amount of intact nucleic acid present after restriction digestion with a methylation-dependent restriction enzyme, or by comparing amplification profiles after a bisulfite reaction, or by comparing sequences of bisulfite-treated nucleic acid with sequences of untreated nucleic acid, or by comparing TET-treated nucleic acid with untreated nucleic acid. Thus, a value, e.g., a methylation value, represents methylation status and can thereby be used as a quantitative indicator of methylation status across multiple copies of a locus. This is of particular use when it is desirable to compare the methylation status of sequences in a sample to a threshold or reference value.

[0060] As used herein, "methylation frequency" or "percent (%) methylation" refers to the number of instances where a molecule or locus is methylated relative to the number of instances where the molecule or locus is unmethylated.

[0061] As used herein, the term "methylation score" refers to a score indicating the number of methylation events detected in a marker or panel of markers compared to the median number of methylation events for that marker or panel of markers from a randomized population of mammals (e.g., a random population of 10, 20, 30, 40, 50, 100, or 500 mammals) that do not have a particular tumor of interest. A high methylation score for a marker or panel of markers can be any score, as long as the score is greater than the corresponding reference score. For example, a high methylation score for a marker or panel of markers can be 0.5, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more times greater than the reference methylation score.

[0062] Thus, methylation status refers to the methylation state of a nucleic acid (e.g., a genomic sequence). Furthermore, methylation status refers to characteristics of a nucleic acid segment at a particular genomic locus related to methylation. Such characteristics include, but are not limited to, whether any of the cytosine (C) residues within the DNA sequence are methylated, the location of the methylated C residue(s), the frequency or percentage of methylated C throughout any particular region of the nucleic acid, and allelic differences in methylation (e.g., due to differences in allelic origin). Furthermore, the terms "methylation status," "methylation profile," and "methylation status" refer to the relative concentration, absolute concentration, or pattern of methylated or unmethylated C throughout any particular region of a nucleic acid in a biological sample. For example, if a cytosine (C) residue(s) within a nucleic acid sequence are methylated, it can be referred to as "hypermethylated" or "increased methylation," whereas if a cytosine (C) residue(s) within a DNA sequence are unmethylated, it can be referred to as "hypomethylated" or "decreased methylation." Similarly, if a cytosine (C) residue(s) in a nucleic acid sequence are methylated compared to another nucleic acid sequence (e.g., from a different region or from a different individual), the sequence is considered to be hypermethylated, or have high methylation, compared to the other nucleic acid sequence. Alternatively, if a cytosine (C) residue(s) in a DNA sequence are unmethylated compared to another nucleic acid sequence (e.g., from a different region or from a different individual), the sequence is considered to be hypomethylated, or have low methylation, compared to the other nucleic acid sequence. Additionally, as used herein, the term "methylation pattern" refers to the collection of methylated and unmethylated nucleotides across a nucleic acid region. Two nucleic acids can have the same or similar methylation frequency or methylation pattern, but have different methylation patterns when the numbers of methylated and unmethylated nucleotides are the same or similar throughout the region but the positions of the methylated and unmethylated nucleotides are different.Sequences are referred to as having "variable methylation," "differences in methylation," or "differential methylation states" when they differ in the degree (e.g., one has increased or decreased methylation compared to the other), frequency, or pattern of methylation. The term "differential methylation" refers to the difference in the level or pattern of nucleic acid methylation in a cancer-positive sample compared to the level or pattern of nucleic acid methylation in a cancer-negative sample. The term can also refer to the difference in the level or pattern between patients who experience cancer recurrence after surgery and those who do not. Variable methylation and specific levels or patterns of DNA methylation can be prognostic and predictive biomarkers, for example, if precise cutoffs or predictive characteristics are defined. In some embodiments, one or more CpG sites in a DMR can be located within a non-coding region, such as a region corresponding to a long non-coding RNA (lncRNA).

[0063] Methylation state frequencies can be used to describe a population of individuals or a sample from a single individual. For example, a nucleotide locus with a methylation state frequency of 50% is methylated in 50% of cases and unmethylated in 50% of cases. Such frequencies can be used, for example, to represent the degree to which a nucleotide locus or nucleic acid region is methylated in a population of individuals or a collection of nucleic acids. Thus, if the methylation in a first population or pool of nucleic acid molecules is different from the methylation in a second population or pool of nucleic acid molecules, the frequency of the methylation state in the first population or pool will be different from the frequency of the methylation state in the second population or pool. Such frequencies can also be used, for example, to describe the degree to which a nucleotide locus or nucleic acid region is methylated in a single individual. For example, such frequencies can be used to describe the degree to which a group of cells from a tissue sample is methylated or unmethylated at a nucleotide locus or nucleic acid region.

[0064] Typically, methylation of human DNA occurs at dinucleotide sequences containing adjacent guanines and cytosines, where the cytosine is located 5' of the guanine (also called CpG dinucleotide sequences). In the human genome, most cytosines within CpG dinucleotides are methylated, although some remain unmethylated in certain CpG dinucleotide-rich genomic regions known as CpG islands (see, e.g., Antequera, et al. (1990) Cell 62:503-514).

[0065] As used herein, "CpG island" or "cytosine-phosphate-guanine-island" refers to a G:C-rich region of genomic DNA that contains more CpG dinucleotides than the total genomic DNA. A CpG island can be at least 100, 200 base pairs long, or longer, where the G:C content of the region is at least 50% and the ratio of observed CpG frequency to expected frequency is 0.6; in some cases, a CpG island can be at least 500 base pairs long, where the G:C content of the region is at least 55% and the ratio of observed CpG frequency to expected frequency is 0.65. The observed CpG frequency to expected frequency can be calculated according to the method provided in Gardiner-Garden et al. (1987) J.Mol.Biol.196:261-281. For example, the observed CpG frequency relative to the expected frequency can be calculated according to the formula R = (A x B) / (C x D), where R is the ratio of the observed CpG frequency to the expected frequency, A is the number of CpG dinucleotides in the analyzed sequence, B is the total number of nucleotides in the analyzed sequence, C is the total number of C nucleotides in the analyzed sequence, and D is the total number of G nucleotides in the analyzed sequence. Methylation status is usually determined in CpG islands, for example, in promoter regions. However, it will be appreciated that other sequences in the human genome, such as CpA and CpT, are also subject to DNA methylation (see Ramsahoye (2000) Proc. Natl. Acad. Sci. USA 97:5237-5242; Salmon and Kaye (1970) Biochim. Biophys. Acta. 204:340-351; Grafstrom (1985) Nucleic Acids Res. 13:2827-2842; Nyce (1986) Nucleic Acids Res. 14:4353-4367; Woodcock (1987) Biochem. Biophys. Res. Commun. 145:888-894).

[0066] As used herein, a "methylation-specific reagent" refers to a reagent that modifies the nucleotides of a nucleic acid molecule as a function of the methylation state of the nucleic acid molecule, or a methylation-specific reagent refers to a compound or composition, or other agent, that is capable of altering the nucleotide sequence of a nucleic acid molecule in a manner that reflects the methylation state of the nucleic acid molecule. Methods of treating nucleic acid molecules with such reagents can include contacting the nucleic acid molecule with the reagent, optionally in combination with additional steps, to achieve a desired change in nucleotide sequence. Such methods can be applied in a manner that results in the modification of unmethylated nucleotides (e.g., each unmethylated cytosine) to a different nucleotide. For example, in some embodiments, such reagents can deaminate unmethylated cytosine nucleotides to generate deoxyuracil residues. Examples of such reagents include, but are not limited to, methylation-sensitive restriction enzymes, methylation-dependent restriction enzymes, bisulfite reagents, TET enzymes, and borane reducing agents.

[0067] Alteration of a nucleic acid nucleotide sequence with a methylation-specific reagent can also result in a nucleic acid molecule in which each methylated nucleotide is modified to a different nucleotide.

[0068] The term "methylation assay" refers to any assay for determining the methylation status of one or more CpG dinucleotide sequences within a nucleic acid sequence.

[0069] The term "MS AP-PCR" (methylation-sensitive arbitrarily primed polymerase chain reaction) refers to an art-recognized technique that uses CG-rich primers to scan the entire genome and focus on regions most likely to contain CpG dinucleotides, and is described by Gonzalgo et al. (1997) Cancer Research 57:594-599.

[0070] The term "MethyLight™" refers to the art-recognized fluorescence-based real-time PCR technology described by Eads et al. (1999) Cancer Res. 59:2302-2306.

[0071] The term "HeavyMethyl™" refers to an assay in which methylation-specific inhibitory probes (also referred to herein as inhibitors) that cover the CpG positions between or covered by the amplification primers enable methylation-specific selective amplification of a nucleic acid sample.

[0072] The term "HeavyMethyl™ MethyLight™" assay refers to the HeavyMethyl™ MethyLight™ assay, which is a variation of the MethyLight™ assay in which the MethyLight™ assay is combined with a methylation-specific blocking probe that covers the CpG positions between the amplification primers.

[0073] The term "Ms-SNuPE" (methylation-sensitive single-nucleotide primer extension) refers to the art-recognized assay described in Gonzalgo & Jones (1997) Nucleic Acids Res. 25:2529-2531.

[0074] The term "MSP" (methylation-specific PCR) refers to the art-recognized methylation assay described in Herman et al. (1996) Proc. Natl. Acad. Sci. USA 93:9821-9826 and U.S. Pat. No. 5,786,146.

[0075] The term "COBRA" (Combined Bisulfite Restriction Analysis) refers to an art-recognized methylation assay described in Xiong & Laird (1997) Nucleic Acids Res. 25:2532-2534.

[0076] The term "MCA" (methylated CpG island amplification) refers to the methylation assay described in Toyota et al. (1999) Cancer Res. 59:2307-12 and WO00 / 26401A1.

[0077] As used herein, a "selected nucleotide" refers to one of the four nucleotides typically occurring in a nucleic acid molecule (C, G, T, and A for DNA and C, G, U, and A for RNA), and can include methylated derivatives of a typically occurring nucleotide (e.g., when C is a selected nucleotide, both methylated and unmethylated C are included in the meaning of the selected nucleotide), but a methylated selected nucleotide specifically refers to a methylated typically occurring nucleotide, and an unmethylated selected nucleotide specifically refers to an unmethylated typically occurring nucleotide.

[0078] The term "methylation-specific restriction enzyme" refers to a restriction enzyme that selectively digests nucleic acids depending on the methylation state of its recognition site. For restriction enzymes that specifically cleave when the recognition site is unmethylated or hemimethylated (methylation-sensitive enzymes), cleavage does not occur (or occurs much less efficiently) when the recognition site is methylated on one or both strands. For restriction enzymes that specifically cleave only when the recognition site is methylated (methylation-dependent enzymes), cleavage does not occur (or occurs much less efficiently) when the recognition site is unmethylated. Methylation-specific restriction enzymes are preferred, and their recognition sequences contain a CG dinucleotide (e.g., a recognition sequence such as CGCG or CCCGGG). More preferred in some embodiments are restriction enzymes that do not cleave when the cytosine in this dinucleotide is methylated at the C5 carbon atom.

[0079] As used herein, the terms "copy number variation," "CNV," "copy number abnormality," and "CNA" generally refer to a variation in the copy number of a nucleic acid sequence present in a test sample compared to the copy number of the nucleic acid sequence present in a reference sample. In some cases, the nucleic acid sequence is an entire chromosome or a significant portion thereof. A "copy number variant" refers to a nucleic acid sequence in which a copy number difference is found by comparing the nucleic acid sequence of interest in a test sample with the expected level of the nucleic acid sequence of interest. For example, the level of the nucleic acid sequence of interest in a test sample is compared with the level of the nucleic acid sequence of interest present in a qualified sample. Copy number variants / mutations include deletions, including microdeletions, insertions, including microinsertions, duplications, multiplications, and translocations. CNVs include chromosomal aneuploidy, partial aneuploidy, polyploidy, and partial polyploidy. In some cases, analyzing a nucleic acid sample for CNVs refers to characterizing the status of a chromosome or segmental aneuploidy by one of three types of calls (e.g., normal or unaffected, affected, and no call). Usually, a threshold is set for determining whether a sample is normal or abnormal. Parameters related to aneuploidy or other copy number variations can be measured in the sample, and the measured value can be compared with the threshold. For example, in the case of duplication aneuploidy, if the dosage of a chromosome or segment (or other measured sequence content) exceeds the defined threshold set for abnormal samples, a call for abnormality is made. In the case of such aneuploidy, if the dosage of a chromosome or segment is below the threshold set for normal samples, a call for normality is made. In contrast, in the case of deletion aneuploidy, if the dosage of a chromosome or segment is below the defined threshold for abnormal samples, a call for abnormality is made, and if the dosage of a chromosome or segment is above the threshold set for normal samples, a call for normality is made. For example, in the presence of trisomy, a call for "normal" is determined by a parameter, for example, a test chromosome dosage value that is below a user-defined confidence threshold, and a call for "abnormality" is determined by a parameter, for example, a test chromosome dosage value that is above a user-defined confidence threshold.A "no call" result is determined by a parameter, e.g., test chromosome dosage, that falls between the thresholds for making a "normal" or "abnormal" call. The term "no call" is used interchangeably with "non-classification."

[0080] The term "aneuploidy" as used herein generally refers to an imbalance of genetic material caused, for example, by the loss or gain of an entire chromosome or a portion of a chromosome. The terms "partial aneuploidy" and "partial chromosomal aneuploidy" as used herein refer to an imbalance of genetic material caused by the loss or gain of a portion of a chromosome, such as partial monosomy and partial trisomy, and include imbalances resulting from translocations, deletions, and insertions. The terms "chromosomal aneuploidy" and "complete chromosomal aneuploidy" as used herein refer to an imbalance of genetic material caused by the loss or gain of an entire chromosome, and include germline aneuploidy and mosaic aneuploidy. For example, aneuploidy can result from an extra set of chromosomes (e.g., meiotic errors) that can cause congenital disorders. This type of aneuploidy can be caused by chromosomes not properly separating during meiosis or by sperm fertilizing an egg with more than one set of chromosomes. In other cases, aneuploidy results from chromosomal instability (CIN) due to failure of mitotic checkpoints, resulting in chromosomal missegregation (e.g., mitotic errors), which leads to the gain of oncogenes or loss of tumor suppressors in cancer-related disease states.

[0081] As described herein, the term "aneuploidy score" or "AS" generally refers to the total number of altered chromosome arms in a sample, ranging from 0 (no arms) to 39 (all arms - the long and short arms for each non-acrocentric chromosome, and only the long arms for chromosomes 13, 14, 15, 21, and 22). As further described herein, the aneuploidy score (AS) calculated for each sample is the sum of gains and losses at the chromosome arm level, adjusted for ploidy.

[0082] As used herein, the "sensitivity" of a particular marker (or set of markers used in combination) refers to the percentage of samples reporting DNA methylation values ​​above a threshold that distinguishes between tumor and non-tumor samples. In some embodiments, a positive result is defined as a histologically confirmed tumor reporting a DNA methylation value above a threshold (e.g., a range associated with disease), and a false negative result is defined as a histologically confirmed tumor reporting a DNA methylation value below a threshold (e.g., a range associated with non-disease). Thus, a sensitivity value reflects the probability that a DNA methylation measurement value for a given marker from a known diseased sample will fall within the range of disease-associated measurements. As defined herein, the clinical significance of a calculated sensitivity value represents an estimate of the probability that a given marker will detect the presence of a clinical condition when applied to subjects with that condition.

[0083] As used herein, the "specificity" of a given marker (or a set of markers used together) refers to the proportion of non-neoplastic samples reporting DNA methylation values ​​below a threshold that distinguishes between neoplastic and non-neoplastic samples. In some embodiments, a negative is defined as a histologically confirmed non-neoplastic sample reporting a DNA methylation value below the threshold (e.g., a range not associated with any disease), and a false positive is defined as a histologically confirmed non-neoplastic sample reporting a DNA methylation value above the threshold (e.g., a range associated with a disease). Thus, the specificity value reflects the probability that a DNA methylation measurement value of a given marker from a known non-neoplastic sample will fall within the range of non-disease-associated measurements. As defined herein, the clinical relevance of a calculated specificity value represents an estimate of the probability that a given marker will detect the absence of a clinical condition when applied to patients without that condition.

[0084] The term "AUC" as used herein is an abbreviation for "area under the curve." In particular, AUC refers to the area under the receiver operating characteristic (ROC) curve. An ROC curve is a plot of the true positive rate against the false positive rate for different possible cut points of a diagnostic test. The ROC curve shows the trade-off between sensitivity and specificity depending on the cut point selected (any increase in sensitivity will be accompanied by a decrease in specificity). The area under the ROC curve (AUC) is a measure of the accuracy of a diagnostic test (the larger the area, the better, with 1 being optimal, and a random test has a ROC curve located on the diagonal, with an area of ​​0.5. See: J.P. Egan. (1975) Signal Detection Theory and ROC Analysis, Academic Press, New York).

[0085] As used herein, the term "tumor" refers to any new, abnormal growth of tissue. Thus, a tumor can be a pre-malignant tumor or a malignant tumor.

[0086] The term "tumor-specific marker," as used herein, refers to any biological material or element that can be used to indicate the presence of a tumor. Examples of biological materials include, but are not limited to, nucleic acids, polypeptides, carbohydrates, fatty acids, cellular components (e.g., cell membranes and mitochondria), and whole cells. In some cases, a marker is a specific nucleic acid region (e.g., a gene, a region within a gene, a specific locus, etc.). A region of a nucleic acid that is a marker may be referred to, for example, as a "marker gene," a "marker region," a "marker sequence," a "marker locus," etc.

[0087] As used herein, the term "adenoma" refers to a benign tumor of glandular origin. These growths are benign, although over time they can progress to become malignant (e.g., esophageal adenocarcinoma or EAC).

[0088] The terms "precancerous" or "preneoplastic" and their equivalents refer to any cell proliferative disorder undergoing malignant transformation. For example, as further described herein, low-grade dysplasia (LGD) BE and high-grade dysplasia (HGD) BE are considered precancerous conditions.

[0089] As used herein, the term "esophageal disorder" refers to a type of disorder associated with the esophagus and / or esophageal tissue. Examples of esophageal disorders include, but are not limited to, Barrett's esophagus (BE), non-dysplastic Barrett's esophagus (NDBE), Barrett's esophageal dysplasia (BED), Barrett's low-grade esophageal dysplasia (BE-LGD), Barrett's high-grade esophageal dysplasia (BE-HGD), and esophageal adenocarcinoma (EAC).

[0090] The "site" of a tumor, adenoma, cancer, etc. is the tissue, organ, cell type, anatomical region, body part, etc. within a subject's body in which the tumor, adenoma, cancer, etc. is located.

[0091] As used herein, the application of a "diagnostic" test includes detecting or identifying a disease state or condition in a subject, determining the likelihood that a subject will suffer from a given disease or condition, determining the likelihood that a subject with a disease or condition will respond to a treatment, determining the prognosis (or likelihood of progression or regression) of a subject with a disease or condition, and determining the effectiveness of a treatment for a subject with a disease or condition. For example, a diagnostic test can be used to detect the presence or likelihood of a subject suffering from a tumor, or the likelihood that such a subject will respond favorably to a compound (e.g., a pharmaceutical, e.g., a drug) or other treatment.

[0092] The term "isolated," when used with reference to a nucleic acid, such as "isolated oligonucleotide," refers to a nucleic acid sequence that is identified and separated from at least one contaminant nucleic acid normally associated with its natural source. An isolated nucleic acid exists in a form or setting that is different from that in which it is found in nature. In contrast, non-isolated nucleic acids, such as DNA and RNA, are found in the state in which they exist in nature. Examples of non-isolated nucleic acids include a given DNA sequence (e.g., a gene) found adjacent to neighboring genes on a host cell chromosome; an RNA sequence, such as a particular mRNA sequence encoding a particular protein, that is found in a cell as a mixture with many other mRNAs encoding many proteins. However, an isolated nucleic acid encoding a particular protein includes, for example, a nucleic acid in a cell that normally expresses that protein, where the nucleic acid is in a location different from its chromosomal location in natural cells or is otherwise flanked by different nucleic acid sequences than that in which it is found in nature. An isolated nucleic acid or oligonucleotide can exist in single-stranded or double-stranded form. When an isolated nucleic acid or oligonucleotide is used to express a protein, the oligonucleotide will minimally contain a sense or coding strand (i.e., the oligonucleotide may be single-stranded), but may also contain both a sense and an antisense strand (i.e., the oligonucleotide may be double-stranded). An isolated nucleic acid may be combined with other nucleic acids or molecules after isolation from its natural or typical environment. For example, an isolated nucleic acid may be present in a host cell, for example, for heterologous expression.

[0093] The term "purified" refers to a molecule, either a nucleic acid or an amino acid sequence, that has been removed, isolated, or separated from its natural environment. Thus, an "isolated nucleic acid sequence" can be a purified nucleic acid sequence. "Substantially purified" molecules are at least 60% free, preferably at least 75% free, and more preferably at least 90% free from other components with which they are naturally associated. As used herein, the terms "purified" or "to purify" also refer to the removal of contaminants from a sample. Removal of contaminating proteins results in an increase in the percentage of the polypeptide or nucleic acid of interest in a sample. In another example, recombinant polypeptides are expressed in plant, bacterial, yeast, or mammalian host cells, and these polypeptides are purified by removal of host cell proteins, thereby increasing the percentage of recombinant polypeptide in a sample.

[0094] The term "composition comprising" a given polynucleotide sequence or polypeptide refers broadly to any composition that includes the given polynucleotide sequence or polypeptide. Compositions can include aqueous solutions containing salts (e.g., NaCl), detergents (e.g., SDS), and other components (e.g., Denhardt's solution, milk powder, salmon sperm DNA, etc.).

[0095] The term "sample" is used in its broadest sense. In one sense, it can refer to animal cells or tissues. In another sense, it refers to specimens or cultures obtained from any source, as well as biological and environmental samples. Biological samples can be obtained from plants or animals (including humans) and can include fluids, solids, tissues, and gases. Environmental samples include environmental materials such as surface material, soil, water, and industrial samples. These examples should not be construed as limiting the sample types applicable to various embodiments of the present disclosure.

[0096] As used herein, a "remote sample," as used in some contexts, refers to a sample that is indirectly collected from a site that is not the sample source of the cell, tissue, or organ. For example, if sample material derived from the pancreas is evaluated in a stool sample, the sample is a remote sample.

[0097] As used herein, the term "patient" or "subject" refers to an organism that is the subject of the various tests described herein. The term "subject" includes animals, preferably mammals, including humans. In preferred embodiments, the subject is a primate. In even more preferred embodiments, the subject is a human. Furthermore, with respect to diagnostic methods, preferred subjects are vertebrate subjects. Preferred vertebrates are warm-blooded animals, and preferred warm-blooded vertebrates are mammals. Preferred mammals are most preferably humans. As used herein, the term "subject" includes both human and animal subjects. Accordingly, veterinary uses are provided herein. Thus, the present disclosure provides for the diagnosis of mammals, such as humans, as well as mammals of endangered importance, such as the Amur tiger, mammals of economic importance, such as animals raised on farms for human consumption, and / or animals of social importance to humans, such as animals kept as pets or in zoos. Examples of such animals include, but are not limited to, carnivores such as cats and dogs; swine such as pigs, hogs, and wild boars; ruminants and / or ungulates such as cows, oxen, sheep, giraffes, deer, goats, bison, and camels; pinnipeds; and horses. Thus, diagnostics and treatments for livestock, including but not limited to domestic pigs, ruminants, ungulates, horses (including racehorses), and the like, are also provided.

[0098] As used herein, the term "kit" refers to any delivery system for delivering materials. In the context of a reaction assay, such delivery systems include systems that allow for the storage, transport, or delivery of reaction reagents (e.g., oligonucleotides, enzymes, etc. in appropriate containers) and / or supporting materials (e.g., buffers, written instructions for conducting the assay, etc.) from one location to another. For example, a kit may include one or more enclosed members (e.g., boxes) containing the relevant reaction reagents and / or supporting materials. As used herein, the term "fragmented kit" refers to a delivery system that includes two or more separate containers, each containing a subportion of the overall kit components. The containers can be delivered to the intended recipient together or separately. For example, a first container may contain an enzyme for use in an assay, while a second container contains an oligonucleotide. The term "separated kit" is intended to encompass, but is not limited to, a kit containing analyte-specific reagents (ASRs) regulated under Section 520(e) of the Federal Food, Drug, and Cosmetic Act. Indeed, any delivery system comprising two or more separate containers, each housing a portion of the overall kit's components, is encompassed by the term "fragmented kit." In contrast, a "combined kit" refers to a delivery system that contains all components of a reaction assay in a single container (e.g., in a single box housing each of the desired components). The term "kit" encompasses both fragmented and combined kits.

[0099] As used herein, the term "information" refers to any collection of facts or data. With respect to information stored or processed using computer system(s), including but not limited to the Internet, the term refers to any data stored in any format (e.g., analog, digital, optical, etc.). As used herein, the term "information about a subject" refers to facts or data about a subject (e.g., a human, plant, or animal). The term "genomic information" refers to information related to a genome, including, but not limited to, nucleic acid sequences, genes, methylation rates, allele frequencies, RNA expression levels, protein expression, phenotypes correlated with genotypes, etc. "Allele frequency information" refers to facts or data about allele frequencies, including, but not limited to, allele identities, statistical correlations between the presence of alleles and characteristics of a subject (e.g., a human subject), the presence or absence of alleles in an individual or population, the likelihood that an allele is present in an individual with one or more particular characteristics, etc.

[0100] 2. Methylation DNA markers and biomarker panels To distinguish non-dysplastic Barrett's esophagus (NDBE) from high-grade dysplastic Barrett's esophagus (HGD-BE) or esophageal adenocarcinoma (EAC), a whole-genome methylation approach was used to simultaneously query copy number abnormalities (CNAs) and DNA methylation, potentially serving as an adjunct to endoscopic histological surveillance. As further described herein, the present disclosure provides materials and methods for determining copy number abnormalities (CNAs), including ploidy and aneuploidy (e.g., determining an aneuploidy score), using sequencing reads from genomes modified in a methylation-specific manner, e.g., cytosine or 5-methylcytosine-converted genomes. Most currently available technologies use NGS directly on wild-type, unconverted DNA. However, embodiments of the present disclosure include the ability to perform methylation and CNV analysis from the same chemistry / dataset without the need to split the sample to perform each respective analysis separately. That is, methylation analysis and CNV analysis can be performed simultaneously on the same converted DNA sample. According to these embodiments, the methylation profile at at least one DMR and CNV and / or AS can be determined using the same DNA sample obtained from the subject. In some embodiments, the methylation profile at at least one DMR and CNV and / or AS can be determined using a single DNA sample obtained from the subject. In some embodiments, the samples (e.g., the same sample or a single sample) have been treated with a reagent that modifies DNA in a methylation-specific manner.

[0101] Thus, as further described herein, embodiments of the present disclosure provide methods, compositions, and systems for screening for various types of esophageal cancer in biological samples. According to these embodiments, the present disclosure includes, but is not limited to, methods and compositions for detecting the presence of esophageal cancer or precancer from a biological sample. In some embodiments, the biological sample is a tissue sample, a blood sample, a plasma sample, a serum sample, a whole blood sample, a buffy coat sample, a secretion sample, an organ secretion sample, a cerebrospinal fluid (CSF) sample, a saliva sample, a urine sample, and / or a stool sample. In some embodiments, the tissue sample is an esophageal tissue sample. In some embodiments, the esophageal sample is obtained from an esophageal biopsy or by swabbing, brushing, or using a sponge capsule device. In some embodiments, the subject is a human.

[0102] As further described herein, embodiments of the present disclosure include novel differentially methylated regions (DMRs) that can distinguish esophageal cancer or precancer from control or benign tissue. In some embodiments, the novel DMRs can distinguish high-grade dysplastic Barrett's esophagus (HGD-BE) or esophageal adenocarcinoma (EAC) from non-dysplastic Barrett's esophagus (NDBE) or normal esophageal controls. According to these embodiments, the novel DMR(s) may be ACVRL1, ADAMTS8, ADAP2, ADRBK2, AKR1B1, ANK1, ANKRD13B, ANXA6, ARNT2, B4GALNT2, BACH2, BCL11A, BSCL2, C12orf53, C14orf82, C17orf107, C18orf1, C1orf95, C5orf42, CACNA1C, CAMK1D, CAMTA1, CBX6, CCDC85A, CCKBR , CD38, CDKN2A, CH25H, CHST1, CHST15, CNTLN, CRHR1, CRTC1, CXCR4, CYP1B1, DCTN2, DIDO1, DMKN, DSE, DYNC1I1, EML6, ENOX1, EPHA4, ESRRG, FAM176A, FAM78B, FBXO10, FERMT2, FHOD3, FLJ45079, FMNL1, FOXP2, FRMD4B, FZD8, GALNTL1, GALNTL4, GAS1, GBG T1, GLIPR2, GNAI1, GNAL, GPR37, GRASP, GRID1, GRM8, GSC, HAR1A, HCN2, HEY2, HIST1H2BE, HS3ST3B1, HTR7, IGFBP2, INHA, INS RR, IRX3, ISM2, KCNG3, KCNK4, KCNS2, KCTD15, KIAA1522, KIAA1614, KIF26A, KL, KLF15, KLHL10, KRT77, KSR2, LBH, LMX1A, LMX1 B, LOC100526820, LONRF2, LRFN2, LRRN1, MAF, MAFB, MARK1, ADAMTSL4-AS1, PGBD5, HSPA12A, SFTPD, LOC107984507, LOC10012 8253, CISTR, SLC16A7, CTXND1, MAX.chr15.4912, ZNF423, RBFOX1, LOC105376772, GSE1, MAX.chr17.8070, ZNF709_6125, MAX.chr19.3439, GREB1, CYP27C1, AC068134.6, LOC388780, STOX2_8441, PPARGC1A, MAX.chr4 .4552, PDZD2, MAX.chr6.6743, LPAL2, CDK14, MAX.chr8.3003, MCOLN2, MEGF11, MFSD11, M RC2, MSX1, NAT8L, NAV1, NBEA, NCRNA00092, NECAB2, NEURL, NGEF, NHLH2, NKD1, NLGN1, NOG, NPPC, NR3C1, NRXN2, NTN1, OBSCN, OCA2, OSBPL1A, OXR1, P2RY1, PDE2A, PDE9A, PDGFRA, PI In some embodiments, the novel DMR(s) are derived from a gene selected from: D1, PLCG2, PLCL1, PNPLA3, POU3F1, POU3F2, PRKACB, PRKAR2B, PRKG1, PRR18, PRR5L, PTHLH, PYGL, RARG, RHBDL3, RIMS2, ROR2, RPRML, SDK2, SOX9, STOX2_6730, SYNGR1, TAC4, TNFRSF19, TRANK1, TSPAN33, TSPAN4, TSPAN5, UBE2E2, UCHL1, UNC5A, VASH2, VIM, WNT6, ZBTB10, ZNF680, ZNF709_4918, ZNF738, ZNF808, and ZNF845 (Table 1), including any combination thereof. In some embodiments, the novel DMR(s) are derived from a gene selected from Table 1, including any combination thereof. While each novel DMR alone can distinguish HGD-BE and / or EAC from NDBE and / or control samples, combining two or more of the novel DMRs can improve sensitivity. Thus, combinations of two or more novel DMRs selected from Table 1 are provided.

[0103] Embodiments of the present disclosure also include novel variably methylated regions (DMRs), each capable of distinguishing esophageal cancer or precancer from control or benign tissue. In some embodiments, the novel DMRs can distinguish high-grade dysplastic Barrett's esophagus (HGD-BE) or esophageal adenocarcinoma (EAC) from non-dysplastic Barrett's esophagus (NDBE) or normal esophageal controls. According to these embodiments, the novel DMR(s) include ACVRL1, ADAMTS8, ADAP2, ADRBK2, AKR1B1, ANK1, ANKRD13B, ANXA6, ARNT2, B4GALNT2, BACH2, BCL11A, BSCL2, C14orf82, C18orf1, C1orf95, C5orf42, CAMK1D, CAMTA1, CCDC85A, CD38, C DKN2A, CHST1, CHST15, CRHR1, CYP1B1, DIDO1, DSE, DYNC1I1, EML6, ENOX1, EPHA4, FAM176A, FBXO10, FERMT2 , FHOD3, FLJ45079, FMNL1, FRMD4B, FZD8, GALNTL1, GALNTL4, GAS1, GBGT1, GLIPR2, GNAI1, GNAL, GRASP, GRI D1, GRM8, GSC, HAR1A, HCN2, HEY2, HIST1H2BE, HS3ST3B1, HTR7, IGFBP2, INHA, IRX3, ISM2, KCNG3, KCNK4, KC NS2, KCTD15, KIAA1614, KL, KLHL10, KSR2, LBH, LMX1A, LMX1B, LOC100526820, LONRF2, LRFN2, MAFB, MARK1, PGBD5, HSPA12A, SFTPD, LOC107984507, LOC100128253, CISTR, SLC16A7, CTXND1, ZNF423, LOC105376772, G SE1, MAX.chr19.3439, GREB1, CYP27C1, AC068134.6, LOC388780, STOX2_8441, PPARGC1A, PDZD2, MAX.chr6.6743, LPAL2, CDK14, MCOLN2, MEGF11, MRC2, MSX1, NAT8L, NAV1, NBEA, NCRNA00092, NEURL, NGEF, NHLH2, NKD1, NLGN1, NOG, NPPC, NR3C1, NRXN2, OBSCN, OCA2, OSBPL1A, OXR1, P2RY1, PDE2A, PDGFRA, PID1, PLCG2, PLCL1, PNPLA3, POU3F1, POU3F2, PRKACB, PRKAR The novel DMR(s) are derived from a gene selected from: 2B, PRKG1, PRR18, PRR5L, PTHLH, PYGL, RHBDL3, RIMS2, ROR2, RPRML, SDK2, SOX9, STOX2_6730, SYNGR1, TAC4, TRANK1, TSPAN4, UBE2E2, UCHL1, UNC5A, VASH2, ZBTB10, ZNF680, ZNF709_4918, ZNF738, ZNF808, and ZNF845 (Table 2), including any combination thereof. In some embodiments, the novel DMR(s) are derived from any gene selected from Table 2, including any combination thereof. While each novel DMR alone can distinguish HGD-BE and / or EAC from NDBE and / or control samples, combining two or more of the novel DMRs can improve sensitivity. Thus, combinations of two or more novel DMRs selected from Table 2 are provided.

[0104] Embodiments of the present disclosure also include novel variably methylated regions (DMRs), each capable of distinguishing esophageal cancer or precancer from control or benign tissue. In some embodiments, the novel DMRs can distinguish high-grade dysplastic Barrett's esophagus (HGD-BE) or esophageal adenocarcinoma (EAC) from non-dysplastic Barrett's esophagus (NDBE) or normal esophageal controls. According to these embodiments, the novel DMR(s) are derived from genes selected from BACH2, C5orf42, FHOD3, HIST1H2BE, IRX3, KIAA1614, LONRF2, MAFB, PDGFRA, PID1, POU3F1, PRR5L, RHBDL3, and SDK2 (Table 3), including any combination thereof. In some embodiments, the novel DMR(s) are derived from any gene selected from Table 3, including any combination thereof. While each novel DMR alone can distinguish HGD-BE and / or EAC from NDBE and / or control samples, combining two or more of the novel DMRs can improve sensitivity. Thus, combinations of two or more novel DMRs selected from Table 3 are provided.

[0105] Embodiments of the present disclosure also include novel variably methylated regions (DMRs), each capable of distinguishing esophageal cancer or precancer from control or benign tissue. In some embodiments, the novel DMRs can distinguish high-grade dysplastic Barrett's esophagus (HGD-BE) or esophageal adenocarcinoma (EAC) from non-dysplastic Barrett's esophagus (NDBE) or normal esophageal controls. According to these embodiments, the novel DMR(s) are derived from genes selected from KL, PGBD5, ROR2, and LMX1B (Example 4), including any combination thereof. In some embodiments, the novel DMR(s) are derived from any gene selected from Example 4, including any combination thereof. While each novel DMR alone can distinguish HGD-BE and / or EAC from NDBE and / or control samples, combining two or more of the novel DMRs can improve sensitivity. Accordingly, combinations of two or more novel DMRs selected from Example 4 are provided.

[0106] According to the above, the control sample includes a sample from a subject without cancer, a sample from a subject without esophageal cancer, a sample from a subject without esophageal precancer, or a sample from a subject with a type of cancer that is neither esophageal cancer nor precancer. In some embodiments, the control sample includes a sample from a subject with non-dysplastic Barrett's esophagus (NDBE). In some embodiments, the control sample is derived from a tissue sample, a blood sample, a plasma sample, a serum sample, a whole blood sample, a buffy coat sample, a secretion sample, an organ secretion sample, a cerebrospinal fluid (CSF) sample, a saliva sample, a urine sample, and a stool sample. In some embodiments, the tissue sample is an esophageal tissue sample. In some embodiments, the esophageal sample is obtained from an esophageal biopsy or by swabbing, brushing, or using a sponge capsule device.

[0107] In some embodiments, the present disclosure provides compositions and methods for identifying, determining, and / or classifying esophageal cancer or precancer from a biological sample (e.g., a tissue sample, a blood sample, a plasma sample, a serum sample, a whole blood sample, a buffy coat sample, a secretion sample, an organ secretion sample, a cerebrospinal fluid (CSF) sample, a saliva sample, a urine sample, and / or a stool sample). The methods generally involve determining a methylation profile of at least one methylation marker in a biological sample isolated from a subject. In some embodiments, a change in the methylation status or profile of the marker indicates the presence, class, or site of esophageal cancer or precancer. Generally, such methods are useful for detecting the presence or absence of esophageal cancer (e.g., esophageal adenocarcinoma (EAC)) or precancer (e.g., high-grade dysplastic Barrett's esophagus (HGD-BE)).

[0108] In some embodiments, methods are provided that include contacting nucleic acid (e.g., genomic DNA) in a biological sample obtained from a subject with at least one reagent or set of reagents that distinguish between methylated and unmethylated nucleotides (e.g., CpG dinucleotides) within at least one methylation marker, and detecting the presence or absence of esophageal cancer or precancer (e.g., providing a sensitivity of 80% or greater and a specificity of 80% or greater).

[0109] In some embodiments, methods are provided that include measuring one or both of the methylation levels of one or more genes or methylated DNA markers in a biological sample from a human individual by treating genomic DNA in the biological sample with a reagent that modifies the DNA in a methylation-specific manner, and determining the methylation levels of the one or more genes or methylation markers.

[0110] In some embodiments, a method is provided that includes measuring the amount of one or more methylated DNA markers or genes in DNA from a biological sample, measuring the amount of at least one reference marker in the DNA, and calculating the amount of the at least one methylation marker gene measured in the DNA as a percentage of the amount of the reference marker gene measured in the DNA, wherein the value indicates the amount of the at least one methylation marker DNA measured in the biological sample.

[0111] In some embodiments, methods are provided that include measuring the methylation level of CpG sites for one or more genes in a biological sample of a human individual by treating genomic DNA in the biological sample with bisulfite, a reagent capable of modifying DNA in a methylation-specific manner; amplifying the modified genomic DNA using a set of primers for the selected one or more genes; and determining the methylation level of CpG sites for the selected one or more genes.

[0112] In some embodiments, the present disclosure provides a method for characterizing a biological sample, comprising measuring one or both methylation levels of CpG sites of one or more genes in a biological sample from a human individual by treating genomic DNA in the biological sample with bisulfite, amplifying the bisulfite-treated genomic DNA using a set of primers for one or more selected genes, and determining the methylation levels of the CpG sites. In some embodiments, the method comprises comparing the methylation levels of one or both of the methylation markers with the methylation levels of a corresponding set of genes in a control sample without the particular type of cancer, and / or determining that the subject has esophageal cancer or precancer if one or both of the methylation levels measured in the one or more genes are higher than the methylation levels measured in the respective control samples.

[0113] In some embodiments, the present disclosure provides methods that include one or both of measuring the methylation level of one or more genes or markers in a biological sample by treating genomic DNA in the biological sample with bisulfite, amplifying the bisulfite-treated genomic DNA using a set of primers for one or more selected genes, and determining the methylation level of the one or more genes or markers.

[0114] In some embodiments, the present disclosure provides a method for screening for esophageal cancer or precancer in a sample obtained from a subject. According to these embodiments, the method includes one or both of assaying the methylation status or profile of one or more methylated DNA markers, and identifying the subject as having esophageal cancer or precancer if the methylation status or profile of the markers differs from the methylation status or profile of the markers assayed in a subject without esophageal cancer or precancer.

[0115] In some embodiments, the disclosure provides methods that include measuring the methylation level of one or more genes or markers in a biological sample from a human individual by treating genomic DNA in the biological sample with a reagent that modifies the DNA in a methylation-specific manner, amplifying the treated genomic DNA using a set of primers for selected one or more genes or markers, and determining the methylation level of the one or more genes or markers.

[0116] In some embodiments, the present disclosure provides methods for characterizing a biological sample, comprising measuring the amount of at least one methylated DNA marker in DNA extracted from the biological sample; treating genomic DNA in the biological sample with bisulfite; amplifying the bisulfite-treated genomic DNA using primers specific for CpG sites for each marker, wherein the primers specific for each marker are capable of binding to amplicons bounded by the primer sequences; and determining the methylation level of the CpG sites for one or more genes.

[0117] In some embodiments, the present disclosure provides a method comprising: extracting genomic DNA from a biological sample of a human individual having or suspected of having esophageal cancer or precancer, thereby measuring the methylation level of one or more methylated DNA markers in the DNA extracted from the biological sample; treating the extracted genomic DNA with bisulfite; and amplifying the bisulfite-treated genomic DNA with primers specific for one or more markers, wherein the primers specific for the one or more markers are capable of binding to at least a portion of the bisulfite-treated genomic DNA for chromosomal regions of the markers listed in Table 1 or 2; and measuring the methylation level of the one or more methylation markers.

[0118] In some embodiments, the present disclosure provides a method comprising: extracting genomic DNA from a biological sample of a human individual having or suspected of having esophageal cancer or precancer, thereby measuring the methylation level of one or more methylated DNA markers in the DNA extracted from the biological sample; treating the extracted genomic DNA with bisulfite; and amplifying the bisulfite-treated genomic DNA with primers specific for one or more markers, wherein the primers specific for the one or more markers are capable of binding to at least a portion of the bisulfite-treated genomic DNA for chromosomal regions of the markers listed in Table 1; and measuring the methylation level of the one or more methylation markers.

[0119] In some embodiments, the present disclosure provides a method comprising: extracting genomic DNA from a biological sample of a human individual having or suspected of having esophageal cancer or precancer, thereby measuring the methylation level of two or more methylated DNA markers in the DNA extracted from the biological sample; treating the extracted genomic DNA with bisulfite; and amplifying the bisulfite-treated genomic DNA with primers specific for one or more markers, wherein the primers specific for the one or more markers are capable of binding to at least a portion of the bisulfite-treated genomic DNA for chromosomal regions of the markers listed in Table 1; and measuring the methylation level of the one or more methylation markers.

[0120] In some embodiments, the present disclosure provides a method comprising: extracting genomic DNA from a biological sample of a human individual having or suspected of having esophageal cancer or precancer, thereby measuring the methylation level of three or more methylated DNA markers in the DNA extracted from the biological sample; treating the extracted genomic DNA with bisulfite; and amplifying the bisulfite-treated genomic DNA with primers specific for one or more markers, wherein the primers specific for the one or more markers are capable of binding to at least a portion of the bisulfite-treated genomic DNA for chromosomal regions of the markers listed in Table 1; and measuring the methylation level of the one or more methylation markers.

[0121] In some embodiments, the present disclosure provides a method comprising extracting genomic DNA from a biological sample of a human individual suspected of having or having cancer, treating the extracted genomic DNA with bisulfite, amplifying the bisulfite-treated genomic DNA using separate primers specific for one or more CpG sites of the methylated DNA markers, and measuring the methylation level of each CpG site of the one or more markers.

[0122] As further described herein, embodiments of the present disclosure include methods and compositions for characterizing a biological sample and determining the methylation profile of at least one variably methylated region (DMR) in a DNA sample obtained from a subject having or suspected of having esophageal cancer or precancer by treating the sample with a reagent that modifies DNA in a methylation-specific manner. In some embodiments, the method includes detecting the presence of esophageal cancer or precancer from the biological sample. In some embodiments, the at least one DMR can distinguish high-grade dysplastic Barrett's esophagus (HGD-BE) or esophageal adenocarcinoma (EAC) from non-dysplastic Barrett's esophagus (NDBE) or normal esophageal controls.

[0123] According to these embodiments, the method further comprises evaluating the DNA sample from the subject for copy number variation (CNV) or copy number abnormality (CNA).In some embodiments, the CNV distinguishes the subject who has or is suspected of having high-grade dysplastic Barrett's esophagus or esophageal adenocarcinoma (EAC) from the control DNA sample.In some embodiments, at least one DMR comprises increased CNV compared to the control DNA sample.

[0124] In some embodiments, the method includes determining an aneuploidy score (AS) for a DNA sample from a subject. In some embodiments, the AS distinguishes between a subject with or suspected of having high-grade dysplastic Barrett's esophagus or esophageal adenocarcinoma (EAC) and a control DNA sample. In some embodiments, at least one DMR comprises an increased aneuploidy score compared to a control DNA sample. As described herein, the aneuploidy score refers to the total number of altered chromosomal arms in a sample, and determining the aneuploidy score can provide an additional and / or alternative method for evaluating a DNA sample for esophageal cancer or precancer compared to a control.

[0125] In some embodiments, assessing CNV and / or determining an aneuploidy score can complement determining a methylation profile for a DNA sample obtained from a subject with or suspected of having esophageal cancer or precancer. In some embodiments, assessing CNV or determining an aneuploidy score involves obtaining sequencing reads from the same cytosine-converted genome used to assess the methylation profile of the DNA sample. While most current methods use direct NGS on unconverted DNA, in the disclosed method, methylation and CNV / aneuploidy reads are obtained simultaneously from the same chemistry / dataset.

[0126] As will be understood by those skilled in the art based on the present disclosure, the various methods described herein are not limited to the use of any one particular methylated DNA marker, methylation marker gene, methylation gene, and / or DMR. That is, one or more of the methylated DNA markers, methylation marker genes, methylation genes, and / or DMRs disclosed herein (including any combination thereof) can be used to distinguish and / or identify esophageal cancer or precancer from controls. In addition, the methylated DNA markers, methylation marker genes, methylation genes, and / or DMRs disclosed herein can include regions or subregions (e.g., genes on chromosomes, single nucleotides, CpG islands, etc.) of any of the markers listed in Tables 1, 2, and 3.

[0127] These include DMR, ACVRL1, ADAMTS8, ADAP2, ADRBK2, AKR1B1, ANK1, and A NKRD13B, ANXA6, ARNT2, B4GALNT2, BACH2, BCL11A, BSCL2, C12orf53, C1 4orf82, C17orf107, C18orf1, C1orf95, C5orf42, CACNA1C, CAMK1D, CAM TA1, CBX6, CCDC85A, CCKBR, CD38, CDKN2A, CH25H, CHST1, CHST15, CNTLN CRHR1, CRTC1, CXCR4, CYP1B1, DCTN2, DIDO1, DMKN, DSE, DYNC1I1, EML 6. ENOX1, EPHA4, ESRRG, FAM176A, FAM78B, FBXO10, FERMT2, FHOD3, FLJ4 5079 FMNL1 FOXP2 FRMD4B FZD8 GALNTL1 GALNTL4 GAS1 GBGT1 GLI PR2、GNAI1、GNAL、GPR37、GRASP、GRID1、GRM8、GSC、HAR1A、HCN2、HEY2、H IST1H2BE, HS3ST3B1, HTR7, IGFBP2, INHA, INSRR, IRX3, ISM2, KCNG3, K CNK4, KCNS2, KCTD15, KIAA1522, KIAA1614, KIF26A, KL, KLF15, KLHL10 KRT77, KSR2, LBH, LMX1A, LMX1B, LOC100526820, LONRF2, LRFN2, LRRN1 MAF, MAFB, MARK1, ADAMTSL4-AS1, PGBD5, HSPA12A, SFTPD, LOC10798450 7. LOC100128253, CISTR, SLC16A7, CTXND1, MAX.chr15.4912, ZNF423. RBFOX1, LOC105376772, GSE1, MAX.chr17.8070, ZNF709_6125, MAX.chr 19.3439, GREB1, CYP27C1, AC068134.6, LOC388780, STOX2_8441, PPARG C1A, MAX.chr4.4552, PDZD2, MAX.chr6.6743, LPAL2, CDK14, MAX.chr8.3003, MCOLN2, MEGF11, MFSD11, MRC2, MSX1, NAT8L, NAV1, NBEA, NCRNA00092, NECAB2, NEURL, NGEF, NHLH2, NKD1, NLGN1, NOG, NPPC, NR3C1, NRXN2, NTN1, OBSCN, OC A2, OSBPL1A, OXR1, P2RY1, PDE2A, PDE9A, PDGFRA, PID1, PLCG2, PLCL1, PNPLA3, POU3F1, POU3F2, PRKACB, PRKAR2B, PRKG1, PRR18, PRR5L, PTHLH, PYGL, RARG, RHBDL The DMRs are derived from genes selected from the group consisting of 3, RIMS2, ROR2, RPRML, SDK2, SOX9, STOX2_6730, SYNGR1, TAC4, TNFRSF19, TRANK1, TSPAN33, TSPAN4, TSPAN5, UBE2E2, UCHL1, UNC5A, VASH2, VIM, WNT6, ZBTB10, ZNF680, ZNF709_4918, ZNF738, ZNF808, and ZNF845 (Table 1), and the subject has or is suspected of having esophageal cancer (e.g., esophageal adenocarcinoma (EAC)) or esophageal precancer (e.g., high-grade dysplastic Barrett's esophagus (HGD-BE)). In some embodiments, determining the methylation profile of the DMR comprises comparing the methylation profile to a corresponding region from a control DNA sample (e.g., non-dysplastic Barrett's esophagus (NDBE) or normal esophageal control).

[0128] In some embodiments, the DMRs are ACVRL1, ADAMTS8, ADAP2, ADRBK2, AKR1B1, ANK1, ANKRD13B, ANXA6, ARNT2, B4GALNT2, BACH2, BCL11A, BSCL2, C14orf82, C18orf1, C1orf95, C5orf42, CAMK1D, CAMTA1, CCDC85A, CD38, CDKN2A, C HST1, CHST15, CRHR1, CYP1B1, DIDO1, DSE, DYNC1I1, EML6, ENOX1, EPHA4, FAM176A, FBXO10, FERMT2, FHOD3 , FLJ45079, FMNL1, FRMD4B, FZD8, GALNTL1, GALNTL4, GAS1, GBGT1, GLIPR2, GNAI1, GNAL, GRASP, GRID1, GR M8, GSC, HAR1A, HCN2, HEY2, HIST1H2BE, HS3ST3B1, HTR7, IGFBP2, INHA, IRX3, ISM2, KCNG3, KCNK4, KCNS2, KCTD15, KIAA1614, KL, KLHL10, KSR2, LBH, LMX1A, LMX1B, LOC100526820, LONRF2, LRFN2, MAFB, MARK1, PGB D5, HSPA12A, SFTPD, LOC107984507, LOC100128253, CISTR, SLC16A7, CTXND1, ZNF423, LOC105376772, GSE 1, MAX.chr19.3439, GREB1, CYP27C1, AC068134.6, LOC388780, STOX2_8441, PPARGC1A, PDZD2, MAX.chr6.6743, LPAL2, CDK14, MCOLN2, MEGF11, MRC2, MSX1, NAT8L, NAV1, NBEA, NCRNA00092, NEURL, NGEF, NHLH2, NKD1, NLGN1, NOG, NPPC, NR3C1, NRXN2, OBSC N, OCA2, OSBPL1A, OXR1, P2RY1, PDE2A, PDGFRA, PID1, PLCG2, PLCL1, PNPLA3, POU3F1, POU3F2, PRKACB, PRKAR2B, PRKG1, PRR18, PRR5L, PTHLH, PYGL, The DMRs are derived from genes selected from RHBDL3, RIMS2, ROR2, RPRML, SDK2, SOX9, STOX2_6730, SYNGR1, TAC4, TRANK1, TSPAN4, UBE2E2, UCHL1, UNC5A, VASH2, ZBTB10, ZNF680, ZNF709_4918, ZNF738, ZNF808, and ZNF845 (Table 2), and the subject has or is suspected of having esophageal cancer (e.g., esophageal adenocarcinoma (EAC)) or esophageal precancer (e.g., high-grade dysplastic Barrett's esophagus (HGD-BE)). In some embodiments, determining the methylation profile of the DMR comprises comparing the methylation profile to a corresponding region from a control DNA sample (e.g., non-dysplastic Barrett's esophagus (NDBE) or normal esophageal control).

[0129] In some embodiments, the DMR is derived from a gene selected from BACH2, C5orf42, FHOD3, HIST1H2BE, IRX3, KIAA1614, LONRF2, MAFB, PDGFRA, PID1, POU3F1, PRR5L, RHBDL3, and SDK2 (Table 3), and the subject has or is suspected of having esophageal cancer (e.g., esophageal adenocarcinoma (EAC)) or esophageal precancer (e.g., high-grade dysplastic Barrett's esophagus (HGD-BE)). In some embodiments, determining the methylation profile of the DMR comprises comparing the methylation profile to a corresponding region from a control DNA sample (e.g., non-dysplastic Barrett's esophagus (NDBE) or a normal esophagus control).

[0130] In some embodiments, the novel DMR(s) that can distinguish esophageal cancer or precancer from control samples are associated with an area under the ROC curve (AUC) of 0.5 or greater, where the ROC curve distinguishes between subjects having or suspected of having esophageal cancer or precancer and control DNA samples. In some embodiments, the novel DMR(s) that can distinguish esophageal cancer or precancer from control samples are associated with an area under the ROC curve (AUC) of 0.6 or greater, where the ROC curve distinguishes between subjects having or suspected of having esophageal cancer or precancer and control DNA samples. In some embodiments, the novel DMR(s) that can distinguish esophageal cancer or precancer from control samples are associated with an area under the ROC curve (AUC) of 0.7 or greater, where the ROC curve distinguishes between subjects having or suspected of having esophageal cancer or precancer and control DNA samples. In some embodiments, the novel DMR(s) that can distinguish esophageal cancer or precancer from control samples are associated with an area under the ROC curve (AUC) of 0.75 or greater, where the ROC curve distinguishes between subjects having or suspected of having esophageal cancer or precancer and control DNA samples. In some embodiments, the novel DMR(s) that can distinguish esophageal cancer or precancer from control samples are associated with an area under the ROC curve (AUC) of 0.8 or greater, where the ROC curve distinguishes between subjects having or suspected of having esophageal cancer or precancer and control DNA samples. In some embodiments, the novel DMR(s) that can distinguish esophageal cancer or precancer from control samples are associated with an area under the ROC curve (AUC) of 0.85 or greater, where the ROC curve distinguishes between subjects having or suspected of having esophageal cancer or precancer and control DNA samples. In some embodiments, the novel DMR(s) that can distinguish esophageal cancer or precancer from control samples are associated with an area under the ROC curve (AUC) of 0.9 or greater, where the ROC curve distinguishes between subjects having or suspected of having esophageal cancer or precancer and control DNA samples.In some embodiments, the novel DMR(s) that can distinguish esophageal cancer or precancer from control samples are associated with an area under the ROC curve (AUC) of 0.95 or greater, where the ROC curve distinguishes between subjects having or suspected of having esophageal cancer or precancer and control DNA samples.

[0131] In some embodiments, the novel DMR(s) that can distinguish esophageal cancer or precancer from control samples comprise an increased rate of hypermethylation compared to control DNA samples. In some embodiments, the novel DMR(s) that can distinguish esophageal cancer or precancer from control samples comprise an increased rate of hypermethylation compared to control DNA samples.

[0132] In some embodiments, determining the methylation profile of at least one DMR comprises amplifying at least a portion of the DMR using a set of primers. In some embodiments, determining the methylation profile of at least one DMR comprises performing at least one of methylation-specific PCR, quantitative methylation-specific PCR, methylation-specific DNA restriction enzyme analysis, quantitative bisulfite pyrosequencing, flap endonuclease assay, PCR-flap assay, and bisulfite genomic sequencing PCR. In some embodiments, determining the methylation profile of at least one DMR comprises determining the presence or absence of methylation at CpG sites. In some embodiments, the one or more CpG sites are present in a coding region, a non-coding region, and / or a regulatory region of a gene (e.g., any one of the genes disclosed herein). In some embodiments, the DMR(s) that can distinguish esophageal cancer or precancer from control samples can be validated using at least one of methylation-specific PCR, quantitative methylation-specific PCR, methylation-specific DNA restriction enzyme analysis, bisulfite pyrosequencing, flap endonuclease assay, PCR flap assay, and bisulfite genomic sequencing PCR. In some embodiments, the DMR(s) that can distinguish esophageal cancer or precancer from control samples can be assessed based on at least one of the area under the receiver operating characteristic curve (AUC), methylation fold change, methylation percentage, and / or hypermethylation percentage between the test sample and the control sample.

[0133] As one of skill in the art will appreciate based on the present disclosure, esophageal cancer or precancer can be predicted by various combinations of markers (e.g., identified by statistical techniques related to the specificity and sensitivity of the prediction). Embodiments of the present disclosure provide methods for identifying predictive combinations and validated predictive combinations for esophageal cancer or precancer.

[0134] Such methods are not limited to a particular modality or technique for determining the methylation characterization, measurement, or assay for one or more methylation markers, methylation marker genes, genes, DMRs, and / or DNA methylation markers. In some embodiments, such techniques are based on analysis of the methylation status of at least one marker, marker region, or marker base (e.g., CpG methylation status) that comprises a DMR.

[0135] In some embodiments, measuring the methylation state or profile of a methylated DNA marker in a sample comprises determining the methylation state of a single nucleotide base. In some embodiments, measuring the methylation state of a methylated DNA marker in a sample comprises determining the degree of methylation at multiple nucleotide bases. Further, in some embodiments, the methylation state or profile of a methylated DNA marker comprises an increase in methylation of the marker relative to the marker's normal methylation state or profile. In some embodiments, the methylation state or profile of a marker comprises a decrease in methylation of the marker relative to the marker's normal methylation state. In some embodiments, the methylation state or profile of a marker comprises a different pattern of methylation of the marker compared to the marker's normal methylation state or profile.

[0136] Further, in some embodiments, the marker is a region of 100 or fewer nucleotide bases. In some embodiments, the marker is a region of 500 or fewer nucleotide bases. In some embodiments, the marker is a region of 1000 or fewer nucleotide bases. In some embodiments, the marker is a region of 5000 or fewer nucleotide bases. In some embodiments, the marker is one nucleotide base. In some embodiments, the marker is within a high CpG density promoter region.

[0137] In certain embodiments, methods for analyzing nucleic acids for the presence of 5-methylcytosine include treating DNA with reagents that modify DNA in a methylation-specific manner, including, but not limited to, methylation-sensitive restriction enzymes, methylation-dependent restriction enzymes, bisulfite reagents, TET enzymes, and borane reducing agents.

[0138] A frequently used method for analyzing nucleic acids for the presence of 5-methylcytosine is based on the bisulfite method described by Frommer et al. for detecting 5-methylcytosine or its mutations in DNA (Frommer et al. (1992) Proc. Natl. Acad. Sci. USA 89:1827-31, expressly incorporated herein by reference in its entirety for all purposes). The bisulfite method for mapping 5-methylcytosine is based on the finding that cytosine, but not 5-methylcytosine, reacts with hydrogen sulfite ions (also known as bisulfite). The reaction is usually carried out according to the following steps: first, cytosine reacts with bisulfite to form sulfonated cytosine; second, spontaneous deamination of the sulfonated reaction intermediate results in sulfonated uracil; and finally, the sulfonated uracil is desulfonated under alkaline conditions to form uracil. Detection is possible because uracil base pairs with adenine (and therefore behaves like thymine), whereas 5-methylcytosine base pairs with guanine (and therefore behaves like cytosine). This allows methylated cytosine to be distinguished from unmethylated cytosine, for example, by using bisulfite genomic sequencing (Grigg G, & Clark S, Bioessays (1994) 16:431-36; Grigg G, DNA Seq. (1996) 6:189-98), methylation-specific PCR (MSP) (e.g., as disclosed in U.S. Pat. No. 5,786,146), or assays involving sequence-specific probe cleavage (e.g., the QuARTS flap endonuclease assay (see, e.g., Zou et al. (2010) "Sensitive quantification of methylated markers with a novel methylation-specific technology" Clin Chem 56:A199), and U.S. Pat. Nos. 8,361,720, 8,715,937, 8,916,344, and 9,212,392).

[0139] In some embodiments, conventional techniques include methods that involve encapsulating the DNA to be analyzed in an agarose matrix, thereby preventing DNA diffusion and renaturation (bisulfite only reacts with single-stranded DNA), and replacing precipitation and purification steps with high-speed dialysis (Olek A, et al. (1996) "A modified and improved method for bisulfite-based cytosine methylation analysis" Nucleic Acids Res. 24:5064-6). Thus, it is possible to analyze individual cells for methylation status, demonstrating the utility and sensitivity of the method. A summary of conventional methods for detecting 5-methylcytosine is provided in Rein, T., et al. (1998) Nucleic Acids Res. 26:2255.

[0140] Bisulfite methods generally involve bisulfite treatment followed by amplification of short, specific fragments of known nucleic acids, followed by assay of the products by sequencing (Olek & Walter (1997) Nat Genet. 17:275-6) or primer extension reactions (Gonzalgo & Jones (1997) Nucleic Acids Res. 25:2529-31, WO 95 / 00669, U.S. Pat. No. 6,251,594) to analyze individual cytosine positions. Some methods use enzymatic digestion (Xiong & Laird (1997) Nucleic Acids Res. 25:2532-4). Hybridization detection has also been described in the art (Olek et al., WO 99 / 28498). Furthermore, the use of bisulfite technology for methylation detection of individual genes has been described (Grigg & Clark (1994) Bioessays 16:431-6; Zeschnigk et al. (1997) Hum Mol Genet. 6:387-95; Feil et al. (1994) Nucleic Acids Res. 22:695; Martin et al. (1995) Gene 157:261-4; WO9746705; WO9515373).

[0141] Various methylation assay techniques can be used in conjunction with bisulfite treatment in accordance with the techniques of the present invention. These assays allow for the determination of the methylation status of one or more CpG dinucleotides (e.g., CpG islands) within a nucleic acid sequence. Such assays involve, among other techniques, sequencing of the bisulfite-treated nucleic acid, PCR (for sequence-specific amplification), Southern blot analysis, and the use of methylation-specific enzymes, e.g., methylation-sensitive or methylation-dependent enzymes.

[0142] For example, the use of bisulfite treatment has simplified genome sequencing for the analysis of methylation patterns and 5-methylcytosine distribution (Frommer et al. (1992) Proc. Natl. Acad. Sci. USA 89:1827-1831). In addition, restriction enzyme digestion of PCR products amplified from bisulfite-converted DNA has been used to assess methylation status, for example, as described in Sadri & Hornsby (1997) Nucl. Acids Res. 24:5058-5059 or embodied in a method known as COBRA (Combined Bisulfite Restriction Analysis) (Xiong & Laird (1997) Nucleic Acids Res. 25:2532-2534).

[0143] The COBRA™ analysis is a quantitative methylation assay useful for determining DNA methylation levels at specific loci in small amounts of genomic DNA (Xiong & Laird, Nucleic Acids Res. 25:2532-2534, 1997). Briefly, restriction enzyme digestion is used to reveal methylation-dependent sequence differences in PCR products of sodium bisulfite-treated DNA. Methylation-dependent sequence differences are first introduced into genomic DNA by standard bisulfite treatment, following the procedure described by Frommer et al. (Proc. Natl. Acad. Sci. USA 89:1827-1831, 1992). PCR amplification of the bisulfite-converted DNA is then performed using primers specific for the CpG island of interest, followed by restriction endonuclease digestion, gel electrophoresis, and a specifically labeled hybridization probe. Methylation levels in the original DNA sample are represented by the relative amounts of digested and undigested PCR products, providing a linear method for quantitating DNA methylation levels over a wide range. Additionally, this method can be reliably applied to DNA obtained from microdissected, paraffin-embedded tissue samples.

[0144] Typical reagents for COBRA™ analysis (e.g., as found in a typical COBRA™-based kit) may include, but are not limited to, PCR primers for specific loci (e.g., specific genes, markers, DMRs, regions of genes, regions of markers, bisulfite-treated DNA sequences, CpG islands, etc.), restriction enzymes and appropriate buffers, gene hybridization oligonucleotides, control hybridization oligonucleotides, kinase-labeling kits of oligonucleotide probes, and labeled nucleotides. Additionally, bisulfite conversion reagents may include DNA denaturing buffers; sulfonation buffers; DNA recovery reagents or kits (e.g., precipitation, ultrafiltration, affinity columns); desulfonation buffers; and DNA recovery components.

[0145] Assays such as "MethyLight™" (fluorescence-based real-time PCR technology) (Eads et al., Cancer Res. 59:2302-2306, 1999), Ms-SNuPE™ (methylation-sensitive single nucleotide primer extension) reactions (Gonzalgo & Jones, Nucleic Acids Res. 25:2529-2531, 1997), methylation-specific PCR ("MSP"; Herman et al., Proc. Natl. Acad. Sci. USA 93:9821-9826, 1996, U.S. Pat. No. 5,786,146), and methylated CpG island amplification ("MCA"; Toyota et al., Cancer Res. 59:2307-12, 1999) are used alone or in combination with one or more of these methods.

[0146] The "HeavyMethyl™" assay technique is a quantitative method for assessing methylation differences based on methylation-specific amplification of bisulfite-treated DNA. Methylation-specific inhibitor probes ("inhibitors") covering CpG positions between or covered by the amplification primers allow for methylation-specific selective amplification of the sample.

[0147] The term "HeavyMethyl™ MethyLight™" assay refers to the HeavyMethyl™ MethyLight™ assay, which is a variation of the MethyLight™ assay in which the MethyLight™ assay is combined with a methylation-specific blocking probe that covers the CpG positions between the amplification primers. The HeavyMethyl™ assay can also be used in combination with methylation-specific amplification primers.

[0148] Typical reagents for HeavyMethyl™ analysis (e.g., as may be found in a typical MethyLight™-based kit) may include, but are not limited to, the following: PCR primers for a specific locus (e.g., a specific gene, marker, region of a gene, region of a marker, bisulfite-treated DNA sequence, CpG island, or bisulfite-treated DNA sequence or CpG island, etc.), blocking oligonucleotides, optimized PCR buffer and deoxynucleotides, and Taq polymerase.

[0149] MSP (methylation-specific PCR) allows the assessment of the methylation status of virtually any group of CpG sites within a CpG island, independent of the use of methylation-sensitive restriction enzymes (Herman et al. Proc. Natl. Acad. Sci. USA 93:9821-9826, 1996; U.S. Patent No. 5,786,146). Briefly, DNA is modified with sodium bisulfite, which converts unmethylated cytosines, but not methylated cytosines, to uracil, and these products are then amplified with primers specific for methylated versus unmethylated DNA. MSP requires only small amounts of DNA, is sensitive to 0.1% methylated alleles of a given CpG island locus, and can be performed on DNA extracted from paraffin-embedded samples. Typical reagents for MSP analysis (e.g., those that may be found in a typical MSP-based kit) may include, but are not limited to, methylated and unmethylated PCR primers, optimized PCR buffers and deoxynucleotides, and specific probes for specific loci (e.g., specific genes, markers, regions of genes, regions of markers, bisulfite-treated DNA sequences, CpG islands, etc.).

[0150] The MethyLight™ assay is a high-throughput quantitative methylation assay that utilizes fluorescence-based real-time PCR (e.g., TaqMan®) without further manipulation after the PCR step (Eads et al., Cancer Res. 59:2302-2306, 1999). Briefly, the MethyLight™ process begins with a mixed sample of genomic DNA that is converted into a mixed pool of methylation-dependent sequence differences in a sodium bisulfite reaction according to standard procedures (the bisulfite process converts unmethylated cytosine residues to uracil). Fluorescence-based PCR is then performed in a "biased" reaction, e.g., with PCR primers that overlap known CpG dinucleotides. Sequence discrimination occurs both at the level of the amplification process and at the level of the fluorescence detection process.

[0151] The MethyLight™ assay is used as a quantitative test for methylation patterns in nucleic acid, e.g., genomic DNA samples, in which sequence discrimination occurs at the level of probe hybridization. In the quantitative version, the PCR reaction results in methylation-specific amplification in the presence of fluorescent probes that overlap specific putative methylation sites. An unbiased control for input DNA amount is provided by a reaction in which neither the primers nor the probe overlap any CpG dinucleotides. Alternatively, a quantitative test for genomic methylation is achieved by probing a biased PCR pool with either control oligonucleotides that do not cover known methylation sites (e.g., fluorescent-based versions of the HeavyMethyl™ and MSP techniques) or oligonucleotides that cover potential methylation sites.

[0152] The MethyLight™ process can be used with any suitable probe (e.g., TaqMan® probe, Lightcycler® probe, etc.). For example, in some applications, double-stranded genomic DNA is treated with sodium bisulfite and subjected to one of two PCR reactions using a TaqMan® probe, e.g., with MSP primers and / or HeavyMethyl blocker oligonucleotides and a TaqMan® probe. The TaqMan® probe is dual-labeled with fluorescent "reporter" and "quencher" molecules and is designed to be specific for relatively GC-rich regions, so it melts during PCR cycles at a higher temperature, approximately 10°C, than the forward or reverse primers. This allows the TaqMan® probe to remain fully hybridized during the PCR annealing / extension step. Taq polymerase enzymatically synthesizes new strands during PCR, ultimately reaching the annealed TaqMan® probe. Taq polymerase 5' to 3' endonuclease activity then displaces the TaqMan® probe by digesting it, releasing a fluorescent reporter molecule for quantitative detection of its unquenched signal using a real-time fluorescence detection system.

[0153] Typical reagents for MethyLight™ analysis (e.g., as may be found in a typical MethyLight™-based kit) may include, but are not limited to: PCR primers for specific loci (e.g., specific genes, markers, regions of genes, regions of markers, bisulfite-treated DNA sequences, CpG islands, etc.), TaqMan® or Lightcycler® probes, optimized PCR buffer and deoxynucleotides, and Taq polymerase.

[0154] The QM™ (Quantitative Methylation) Assay is an alternative quantitative test for methylation patterns in genomic DNA samples, in which sequence discrimination occurs at the level of probe hybridization. In this quantitative version, PCR reactions result in unbiased amplification in the presence of fluorescent probes that overlap specific putative methylation sites. An unbiased control for input DNA amount is provided by reactions in which neither the primers nor the probe overlap any CpG dinucleotides. Alternatively, quantitative testing for genomic methylation is achieved by probing biased PCR pools with either control oligonucleotides that do not cover known methylation sites (fluorescence-based versions of HeavyMethyl™ and MSP techniques) or oligonucleotides that cover potential methylation sites.

[0155] The QM™ process can be used with any suitable probe, such as a TaqMan® probe, a Lightcycler® probe, or the like, in the amplification process. For example, double-stranded genomic DNA is treated with sodium bisulfite and subjected to unbiased primers and a TaqMan® probe. The TaqMan® probe is dual-labeled with fluorescent reporter and quencher molecules and is designed to be specific to relatively GC-rich regions, so it melts during PCR cycles at a temperature approximately 10°C higher than the forward or reverse primers. This allows the TaqMan® probe to remain fully hybridized during the PCR annealing / extension step. Taq polymerase enzymatically synthesizes new strands during PCR, eventually reaching the annealed TaqMan® probe. The 5' to 3' endonuclease activity of Taq polymerase then displaces the TaqMan® probe by digesting it, releasing a fluorescent reporter molecule for quantitative detection of its unquenched signal using a real-time fluorescence detection system. Typical reagents for QM™ analysis (e.g., as found in a typical QM™-based kit) can include, but are not limited to, PCR primers for specific loci (e.g., specific genes, markers, regions of genes, regions of markers, bisulfite-treated DNA sequences, CpG islands, etc.), TaqMan® or Lightcycler® probes, optimized PCR buffer and deoxynucleotides, and Taq polymerase.

[0156] The Ms-SNuPE™ technology is a quantitative method for assessing methylation differences at specific CpG sites based on bisulfite treatment of DNA followed by single-nucleotide primer extension (Gonzalgo & Jones, Nucleic Acids Res. 25:2529-2531, 1997). Briefly, genomic DNA is reacted with sodium bisulfite to convert unmethylated cytosines to uracil, while leaving 5-methylcytosines unchanged. Next, PCR primers specific to the bisulfite-converted DNA are used to amplify the desired target sequence, and the resulting product is isolated and used as a template for methylation analysis at the CpG sites of interest. This allows for the analysis of small amounts of DNA (e.g., microdissected pathology slices), avoiding the use of restriction enzymes to determine the methylation status at CpG sites.

[0157] Typical reagents for Ms-SNuPE™ analysis (e.g., as might be found in a typical Ms-SNuPE™-based kit) may include, but are not limited to, PCR primers for specific loci (e.g., specific genes, markers, regions of genes, regions of markers, bisulfite-treated DNA sequences, CpG islands, etc.), optimized PCR buffers and deoxynucleotides, gel extraction kits, positive control primers, Ms-SNuPE™ primers for specific loci, reaction buffers (for the Ms-SNuPE reaction), and labeled nucleotides. Additionally, bisulfite conversion reagents may include DNA denaturing buffers; sulfonation buffers; DNA recovery reagents or kits (e.g., precipitation, ultrafiltration, affinity columns); desulfonation buffers; and DNA recovery components.

[0158] Reduced representation bisulfite sequencing (RRBS) begins with bisulfite treatment of nucleic acids to convert all unmethylated cytosines to uracils, followed by restriction enzyme digestion (e.g., with an enzyme that recognizes sites containing CG sequences, such as MspI) and completion of fragment sequencing after coupling to an adaptor ligand. The choice of restriction enzyme enriches for fragments in CpG-dense regions, reducing the number of redundant sequences that may map to multiple gene locations during analysis. As such, RRBS reduces the complexity of a nucleic acid sample by selecting a subset of restriction fragments for sequencing (e.g., by size selection using preparative gel electrophoresis). In contrast to whole-genome bisulfite sequencing, all fragments generated by restriction enzyme digestion contain DNA methylation information for at least one CpG dinucleotide. As such, RRBS enriches samples for promoters, CpG islands, and other genomic features with high frequency of restriction enzyme cleavage sites in these regions, providing an assay to assess the methylation status of one or more genomic loci.

[0159] A typical protocol for RRBS includes the steps of digesting a nucleic acid sample with a restriction enzyme such as MspI, filling in overhangs and A-tailing, adapter ligation, bisulfite conversion, and PCR (see, e.g., Meissner et al. (2005) "Genome-scale DNA methylation mapping of clinical samples at single-nucleotide resolution" Nat Methods 7:133-6; Meissner et al. (2005) "Reduced representation bisulfite sequencing for comparative high-resolution DNA methylation analysis" Nucleic Acids Res. 33:5868-77).

[0160] In some embodiments, quantitative allele-specific real-time target and signal amplification (QuARTS) assays are used to assess methylation status. Each QuARTS assay involves three sequential reactions: a primary reaction involving amplification (reaction 1) and target probe cleavage (reaction 2), and a secondary reaction involving FRET cleavage and fluorescent signal generation (reaction 3). When a target nucleic acid is amplified with specific primers, a specific detection probe with a flap sequence loosely binds to the amplicon. The presence of a specific invasive oligonucleotide at the target binding site allows a 5' nuclease, such as FEN-1 endonuclease, to cleave the gap between the detection probe and the flap sequence, thereby releasing the flap sequence. The flap sequence is complementary to the non-hairpin portion of the corresponding FRET cassette. Thus, the flap sequence functions as an invasive oligonucleotide on the FRET cassette, resulting in cleavage between the FRET cassette fluorophore and quencher, generating a fluorescent signal. The cleavage reaction can cleave multiple probes per target, thus liberating multiple fluorophores per flap, providing exponential signal amplification. QuARTS can detect multiple targets in a single reaction well by using FRET cassettes with different dyes (see, e.g., Zou et al. (2010) "Sensitive quantification of methylated markers with a novel methylation-specific technology" Clin Chem 56:A199), and U.S. Patent Nos. 8,361,720, 8,715,937, 8,916,344, and 9,212,392, each of which is incorporated by reference herein for all purposes).

[0161] The term "bisulfite reagent" refers to a reagent containing bisulfite, disulfite, hydrogen sulfite, or a combination thereof, useful for distinguishing between methylated and unmethylated CpG dinucleotide sequences, as disclosed herein. Methods for such treatment are known in the art (e.g., PCT / EP2004 / 011715 and WO2013 / 116375, each of which is incorporated by reference in its entirety). In some embodiments, the bisulfite treatment is carried out in the presence of a denaturing solvent, such as, but not limited to, n-alkylene glycol or diethylene glycol dimethyl ether (DME), or in the presence of dioxane or a dioxane derivative. In some embodiments, the denaturing solvent is used at a concentration of 1% to 35% (v / v). In some embodiments, the bisulfite reaction is carried out in the presence of a scavenger, such as, but not limited to, a chroman derivative, such as 6-hydroxy-2,5,7,8-tetramethylchroman 2-carboxylic acid or trihydroxybenzone, and derivatives thereof, such as gallic acid (see PCT / EP2004 / 011715, incorporated by reference in its entirety). In certain preferred embodiments, the bisulfite reaction involves treatment with ammonium bisulfite, for example, as described in WO2013 / 116375.

[0162] In some embodiments, fragments of the treated DNA are amplified using a set of primer oligonucleotides and an amplification enzyme according to the methods and compositions described herein. Amplification of several DNA segments can be performed simultaneously in one reaction vessel, and in the same reaction vessel. Typically, amplification is performed using the polymerase chain reaction (PCR). Amplicons are typically 100-2000 base pairs in length.

[0163] In some embodiments of the method, the methylation status or profile of CpG positions within or near variably methylated regions (e.g., Tables 1, 2, and 3) can be detected by using methylation-specific primer oligonucleotides. This technique (MSP) is described in U.S. Patent No. 6,265,171 to Herman. The use of methylation status-specific primers for amplification of bisulfite-treated DNA allows differentiation between methylated and unmethylated nucleic acids. An MSP primer pair contains at least one primer that hybridizes to a bisulfite-treated CpG dinucleotide. Thus, the primer sequence contains at least one CpG dinucleotide. MSP primers specific for unmethylated DNA contain a "T" at the C position in the CpG.

[0164] Such methods are not limited to a particular type or kind of primer or primer pair associated with one or more methylation markers, methylation marker genes, genes, DMRs, and / or methylated DNA markers. In some embodiments, a primer or primer pair specific to each methylation marker gene can bind to an amplicon bound by the primer sequence, wherein the amplicon bound by the primer sequence of a marker gene is at least a portion of a gene region of a methylation marker gene listed in Table 1, 2, or 3. In some embodiments, the methylation marker primer or primer pair is a set of primers that specifically bind to at least a portion of a gene region comprising a methylation marker listed in Table 1, 2, or 3.

[0165] In another embodiment, the present disclosure provides a method for converting oxidized 5-methylcytosine residues in cell-free DNA to dihydrouracil residues (see Liu et al., 2019, Nat Biotechnol. 37, pp. 424-429; U.S. Patent Application Publication No. 202000370114). The method includes reacting an oxidized 5mC residue selected from 5-formylcytosine (5fC), 5-carboxymethylcytosine (5caC), and combinations thereof with a borane reducing agent. Oxidized 5mC residues may be naturally occurring or, more typically, may result from prior oxidation of a 5mC or 5hmC residue, e.g., oxidation of 5mC or 5hmC by a TET family enzyme (e.g., TET1, TET2, or TET3), or chemical oxidation of 5mC or 5hmC using, for example, inorganic peroxo compounds or compositions such as potassium perruthenate (KRuO4) or peroxotungstate (see, e.g., Okamoto et al. (2011) Chem. Commun. 47:11231-33) and copper(II) perchlorate / 2,2,6,6-tetramethylpiperidine-1-oxyl (TEMPO) combinations (see, e.g., Matsushita et al. (2017) Chem. Commun. 53:5756-59).

[0166] The borane reducing agent can be characterized as a complex of borane and a nitrogen-containing compound selected from a nitrogen heterocycle and a tertiary amine. The nitrogen heterocycle can be monocyclic, bicyclic, or polycyclic, but is typically a monocyclic ring in the form of a five- or six-membered ring containing a nitrogen heteroatom and, optionally, one or more additional heteroatoms selected from N, O, and S. The nitrogen heterocycle can be aromatic or alicyclic. Preferred nitrogen heterocycles herein include 2-pyrroline, 2H-pyrrole, 1H-pyrrole, pyrazolidine, imidazolidine, 2-pyrazoline, 2-imidazoline, pyrazole, imidazole, 1,2,4-triazole, 1,2,4-triazole, pyridazine, pyrimidine, pyrazine, 1,2,4-triazine, and 1,3,5-triazine, any of which may be unsubstituted or substituted with one or more non-hydrogen substituents. Typical non-hydrogen substituents are alkyl groups, particularly lower alkyl groups such as methyl, ethyl, n-propyl, isopropyl, n-butyl, isobutyl, t-butyl, etc. Exemplary compounds include pyridine borane, 2-methylpyridine borane (also called 2-picoline borane), and 5-ethyl-2-pyridine.

[0167] The reaction of oxidized 5mC residues in cell-free DNA with borane reducing agents is advantageous insofar as it utilizes nontoxic reagents and mild reaction conditions, eliminating the need for hydrogen sulfate or other potentially DNA-degrading reagents. Furthermore, the conversion of oxidized 5mC residues to dihydrolauracils with borane reducing agents can be carried out in a "one-pot" or "one-tube" reaction without the need for isolation of intermediates. This is crucial because this conversion involves multiple steps: (1) reduction of the alkene bond connecting C-4 and C-5 of oxidized 5mC, (2) deamination, and (3) either decarboxylation if the oxidized 5mC is 5caC or deformylation if the oxidized 5mC is 5fC.

[0168] In addition to a method for converting oxidized 5-methylcytosine residues in cell-free DNA to dihydrouracil residues, the present disclosure also provides a reaction mixture related to the aforementioned method. The reaction mixture includes a sample of cell-free DNA containing at least one oxidized 5-methylcytosine residue selected from 5caC, 5fC, and combinations thereof, and a borane reducing agent effective in reducing, deaminating, and decarboxylating or deformylating the at least one oxidized 5-methylcytosine residue. As explained above, the borane reducing agent is a complex of borane and a nitrogen-containing compound selected from nitrogen heterocycles and tertiary amines. In a preferred embodiment, the reaction mixture is substantially bisulfite-free, meaning that it is substantially free of bisulfite ions and bisulfite salts. Ideally, the reaction mixture is free of bisulfite salts.

[0169] In a related aspect of the present disclosure, a kit for converting 5mC residues in cell-free DNA to dihydrouracil residues is provided, the kit including a reagent for blocking 5mC residues, a reagent for oxidizing 5mC residues beyond hydroxymethylation to provide oxidized 5mC residues, and a borane reducing agent effective for reducing, deaminating, and decarboxylating or deformylating the oxidized 5mC residues. The kit may also include instructions for using the components to practice the method.

[0170] In another embodiment, a method utilizing the above-described oxidation reaction is provided, which allows for the detection of the presence and location of 5-methylcytosine residues in cell-free DNA and includes the following steps: (a) modifying 5hmC residues in fragmented, adaptor-ligated cell-free DNA to provide affinity tags thereon, which enable the modified 5hmC-containing DNA to be removed from the cell-free DNA; (b) removing the modified 5hmC-containing DNA from the cell-free DNA, leaving DNA containing unmodified 5mC residues; and (c) oxidizing the unmodified 5mC residues to produce 5caC. , 5fC, and combinations thereof; (d) contacting the DNA containing the oxidized 5mC residues with a borane reducing agent effective to reduce, deaminate, and decarboxylate or deformylate the oxidized 5mC residues, thereby providing DNA containing dihydrosyl residues in place of the oxidized 5mC residues; (e) amplifying and sequencing the DNA containing the dihydrosyl residues; and (f) determining the 5-methylation pattern from the sequencing results of (e).

[0171] In some embodiments, the present disclosure provides a method for identifying 5-methylcytosine (5mC) or 5-hydroxymethylcytosine (5hmC) in a target nucleic acid. In some embodiments, the method includes providing a biological sample containing the target nucleic acid; modifying the target nucleic acid by converting 5mC and 5hmC in the nucleic acid sample to 5-carboxylcytosine (5caC) and / or 5-formylcytosine (5fC) to generate one or more 5caC or 5fC residues by contacting the nucleic acid sample with a TET enzyme; treating the target nucleic acid with a borane reducing agent to convert the 5caC and / or 5fC to dihydrouracil (DHU) to provide a modified nucleic acid sample containing the modified target nucleic acid; and detecting the sequence of the modified target nucleic acid, wherein a cytosine (C) to thymine (T) transition or a cytosine (C) to DHU transition in the sequence of the modified target nucleic acid indicates the location of either a 5mC or a 5hmC in the target nucleic acid, relative to the target nucleic acid. In some embodiments, the borane reducing agent is 2-picoline borane.

[0172] In some embodiments, detecting the sequence of the modified target nucleic acid comprises one or more of chain termination sequencing, microarray, high-throughput sequencing, and restriction enzyme analysis. In some embodiments, the TET enzyme is selected from the group consisting of human TET1, TET2, and TET3, mouse TET1, TET2, and TET3, Naegleria TET (NgTET), and Coprinopsis cinerea (CcTET). In some embodiments, the method further comprises blocking one or more modified cytosines. In some embodiments, the blocking comprises adding a sugar to 5hmC. In some embodiments, the method further comprises amplifying the copy number of one or more nucleic acid sequences. In some embodiments, the oxidizing agent is potassium perruthenate or Cu(II) / TEMPO (2,2,6,6-tetramethylpiperidine-1-oxyl).

[0173] Cell-free DNA is typically extracted from a biological sample from a subject, which may be whole blood, buffy coat, plasma, urine, saliva, mucosal excretions, organ secretions, sputum, stool, or tears. In some embodiments, the cell-free DNA is derived from a tumor (e.g., an esophageal tumor). In other embodiments, the cell-free DNA is derived from a patient with a disease or other pathological condition. The cell-free DNA may or may not be derived from a tumor. In some embodiments, the cell-free DNA modified at 5hmC residues is purified, in fragmented form, and adapter-ligated. DNA purification in this context can be performed using any suitable method known to those of skill in the art and / or described in the relevant literature. While the cell-free DNA itself may be highly fragmented, further fragmentation may sometimes be desirable, as described, for example, in U.S. Patent Publication No. 2017 / 0253924. Cell-free DNA fragments generally range in size from about 20 nucleotides to about 500 nucleotides, more typically from about 20 nucleotides to about 250 nucleotides. The purified cell-free DNA fragments modified in step (a) are end-repaired using conventional means (e.g., restriction enzymes) to ensure that the fragments have blunt ends at each 3' and 5' end. In a preferred method, as described in WO 2017 / 176630, the blunted fragments are also provided with 3' overhangs containing a single adenine residue using a polymerase such as Taq polymerase. This facilitates subsequent ligation of a selected universal adapter, i.e., a Y adapter or hairpin adapter, which ligates to both ends of the cell-free DNA fragments and contains at least one molecular barcode. The use of adapters also allows for selective PCR enrichment of adapter-ligated DNA fragments.

[0174] In some embodiments, "purified fragmented cell-free DNA" comprises adaptor-ligated DNA fragments. The 5hmC residues of these cell-free DNA fragments are modified with an affinity tag to enable subsequent removal of the modified 5hmC-containing DNA from the cell-free DNA. In one embodiment, the affinity tag comprises a biotin moiety, such as biotin, desthiobiotin, oxybiotin, 2-iminobiotin, diaminobiotin, biotin sulfoxide, or biocytin. The use of a biotin moiety as an affinity tag allows for easy removal with streptavidin (e.g., streptavidin beads, magnetic streptavidin beads, etc.).

[0175] Tagging of 5hmC residues with biotin moieties or other affinity tags is achieved by covalently attaching a chemoselective group to the 5hmC residues in the DNA fragment, which can react with a functionalized affinity tag to attach the affinity tag to the 5hmC residue. In one embodiment, the chemoselective group is UDP-glucose-6-azide, which undergoes spontaneous 1,3-cycloaddition with an alkyne-functionalized biotin moiety, as described in Robertson et al. (2011) Biochem. Biophys. Res. Comm. 411(1):40-3, U.S. Patent No. 8,741,567, and WO2017 / 176630. Thus, the addition of the alkyne-functionalized biotin moiety results in the covalent attachment of the biotin moiety to each 5hmC residue.

[0176] The affinity-tagged DNA fragments can then be removed, in one embodiment, using streptavidin in the form of streptavidin beads, magnetic streptavidin beads, etc., and saved for later analysis if desired. The supernatant remaining after removal of the affinity-tagged fragments contains DNA with unmodified 5mC residues but no 5hmC residues.

[0177] In some embodiments, unmodified 5mC residues are oxidized to yield 5caC and / or 5fC residues using any suitable means. The oxidizing agent is selected to oxidize the 5mC residue beyond hydroxymethylation, i.e., to yield 5caC and / or 5fC residues. Oxidation can be performed enzymatically using a catalytically active TET family enzyme. As these terms are used herein, "TET family enzyme" or "TET enzyme" refers to a catalytically active "TET family protein" or "TET catalytically active fragment," as defined in U.S. Patent No. 9,115,386, the disclosure of which is incorporated herein by reference. A preferred TET enzyme in this context is TET2 (see Ito et al. (2011) Science 333(6047):1300-1303). Oxidation can also be performed chemically using a chemical oxidizing agent, as described in the previous section. Examples of suitable oxidizing agents include, but are not limited to, perruthenate anions in the form of inorganic or organic perruthenates, including metal perruthenates such as potassium perruthenate (KRuO), tetraalkylammonium perruthenates such as tetrapropylammonium perruthenate (TPAP) and tetrabutylammonium perruthenate (TBAP), and polymer-supported perruthenate (PSP); and inorganic peroxo compounds and compositions, such as peroxotungstate or a combination of copper(II) perchlorate / TEMPO. It is not necessary to separate the 5fC-containing fragments from the 5caC-containing fragments at this point, as long as both the 5fC and 5caC residues are converted to dihydrouracil (DHU) in the next step of the process.

[0178] In some embodiments, 5-hydroxymethylcytosine residues are blocked with β-glucosyltransferase (β3GT), while 5-methylcytosine residues are oxidized with a TET enzyme, which is effective in generating a mixture of 5-formylcytosine and 5-carboxymethylcytosine. A mixture containing both of these oxidized species can be reacted with 2-picoline borane or another borane reducing agent to yield dihydrouracil. In a variation of this embodiment, 5hmC-containing fragments are not removed. Instead, in "TET-assisted picoline borane sequencing (TAPS)," 5mC- and 5hmC-containing fragments are enzymatically oxidized together to yield 5fC- and 5caC-containing fragments. Reaction with 2-picoline borane generates DHU residues where the 5mC and 5hmC residues originally resided. In "chemically assisted picoline borane sequencing (CAPS)," 5hmC-containing fragments are selectively oxidized with potassium perruthenate, leaving the 5mC residues unchanged.

[0179] As disclosed in International PCT Application PCT / US2019 / 012627 (incorporated herein by reference in its entirety), TAPS involves the use of mild enzymatic and chemical reactions to directly and quantitatively detect 5mC and 5hmC at base resolution without affecting unmodified cytosines. In a related embodiment, the above method further includes identifying hydroxymethylation patterns in 5hmC-containing DNA removed from cell-free DNA. This can be done using the techniques described in detail in WO 2017 / 176630. This process can be performed in a one-tube manner without intermediate removal or isolation. For example, cell-free DNA fragments, preferably adapter-ligated DNA fragments, are first functionalized with βGT-catalyzed uridine diphosphoglucose 6-azide and then biotinylated with a chemoselective azide group. This procedure covalently attaches biotin to each 5hmC site. In a next step, the biotinylated strand and the strand containing unmodified (native) 5mC are simultaneously removed for further processing. Native 5mC-containing chains are pulled down using anti-5mC antibodies or methyl-CpG binding domain (MBD) proteins, as known in the art. Then, with 5hmC residues blocked, unmodified 5mC residues are selectively oxidized using any suitable technique to convert 5mC to 5fC and / or 5caC, as described elsewhere herein.

[0180] These fragments obtained by amplification can have directly or indirectly detectable labels.In some embodiments, these labels are fluorescent labels, radionuclides, or detachable molecular fragments, which have a typical mass that can be detected in mass spectrometer.When these labels are mass labels, some embodiments provide that the labeled amplicons have a single positive or negative effective charge, which allows for better detectability in mass spectrometer.For example, this detection can be performed and visualized by matrix-assisted laser desorption / ionization mass spectrometry (MALDI) or by electron spray mass spectrometry (ESI).

[0181] Methods for isolating DNA suitable for these assay techniques are known in the art. In particular, some embodiments involve the isolation of nucleic acids as described in U.S. Patent Application No. 13 / 470,251 ("Isolation of Nucleic Acids"), which is incorporated herein by reference in its entirety.

[0182] In some embodiments, the markers described herein are used in a QUARTS assay performed on a stool sample. In some embodiments, methods are provided for generating DNA samples, particularly DNA samples containing highly purified, low-abundance nucleic acids in small volumes (e.g., less than 100 microliters, less than 60 microliters) and free of substances that substantially and / or effectively inhibit assays used to test the DNA sample (e.g., PCR, INVADER, QuARTS assays, etc.). Such DNA samples are utilized in diagnostic tests that qualitatively detect the presence or quantitatively measure the activity, expression, or amount of genes, genetic variants (e.g., alleles), or genetic modifications (e.g., methylation) present in a sample collected from a patient. For example, some cancers are correlated with the presence of specific mutant alleles or specific methylation states; therefore, detection and / or quantification of such mutant alleles or methylation states has predictive value in cancer diagnosis and treatment.

[0183] Many useful genetic markers are present in very small amounts in samples, and many of the events that generate such markers are rare. Therefore, even highly sensitive detection methods such as PCR require a large amount of DNA to provide targets with low enough abundance to meet or exceed the detection threshold of the assay. Furthermore, the presence of even a small amount of inhibitors can impair the accuracy and precision of these assays aimed at detecting such low-abundance targets. Therefore, the present specification provides a method for generating such DNA samples, providing the necessary volume and concentration control.

[0184] In some embodiments, the biological sample is a tissue sample, a blood sample, a plasma sample, a serum sample, a whole blood sample, a buffy coat sample, a secretion sample, an organ secretion sample, a cerebrospinal fluid (CSF) sample, a saliva sample, a urine sample, and / or a stool sample. In some embodiments, the tissue sample is an esophageal tissue sample. In some embodiments, the tissue sample is an endoscopic esophageal brushing sample. In some embodiments, the sample comprises esophageal tissue and / or the sample is obtained by endoscopic brushing or non-endoscopic whole esophageal brushing or swabbing using a tethered device (e.g., a capsule sponge, balloon, or other device).

[0185] In some embodiments, the sample comprises tissue and / or biological fluid obtained from a human patient. In some embodiments, the sample comprises esophageal tissue. In some embodiments, the sample comprises esophageal tissue obtained by whole esophageal swabbing or brushing. In some embodiments, the sample comprises secretions. In some embodiments, the sample comprises blood, serum, plasma, gastric secretions, pancreatic juice, gastrointestinal biopsy samples, microscopically dissected cells from esophageal biopsies, esophageal cells shed into the gastrointestinal lumen, and / or esophageal cells recovered from stool. These samples can be derived from the upper gastrointestinal tract, the lower gastrointestinal tract, or include cells, tissue, and / or secretions from both the upper and lower gastrointestinal tract. In some embodiments, the sample comprises cellular fluid, peritoneal fluid, urine, feces, pancreatic juice, fluid obtained during endoscopy, blood, mucus, or saliva. In some embodiments, the sample is a stool sample. Such samples can be obtained by any number of means known in the art, as will be apparent to those skilled in the art. For example, urine and fecal samples are readily available, while blood, ascites, serum, or pancreatic juice samples can be obtained parenterally, for example, by using a needle and syringe. Cell-free or substantially cell-free samples can be obtained by subjecting the sample to a variety of techniques, including, but not limited to, centrifugation and filtration. While obtaining samples without the use of invasive techniques is generally preferred, it may still be preferable to obtain samples such as tissue homogenates, tissue sections, and biopsy specimens. In some embodiments, samples are obtained by esophageal swabbing or brushing, or by use of a sponge capsule device.

[0186] Such samples can be obtained by any number of means known in the art, as will be apparent to those skilled in the art. Cell-free or substantially cell-free samples can be obtained by subjecting the sample to a variety of techniques, including, but not limited to, centrifugation and filtration. While obtaining samples without invasive techniques is generally preferred, it may still be preferable to obtain samples such as tissue homogenates, tissue sections, and biopsy specimens. The present technology is not limited by the method used to prepare the sample and provide nucleic acids for testing. For example, in some embodiments, DNA is isolated from a sample (e.g., a tissue sample, a blood sample, a plasma sample, a serum sample, a whole blood sample, a buffy coat sample, a secretion sample, an organ secretion sample, a cerebrospinal fluid (CSF) sample, a saliva sample, a urine sample, and / or a stool sample) using direct gene capture or related methods, as described, for example, in U.S. Pat. Nos. 8,808,990 and 9,169,511 and WO 2012 / 155072.

[0187] Marker analysis can be performed separately or simultaneously with additional markers within a single test sample. For example, it is possible to combine several markers in one test to efficiently process multiple samples and potentially provide higher diagnostic and / or prognostic accuracy. Furthermore, those skilled in the art will recognize the value of testing multiple samples from the same subject (e.g., at successive time points). Testing such serial samples allows for the identification of changes in the methylation status of markers over time. Changes in methylation status, and the absence of changes in methylation status, can provide useful information about disease states, including, but not limited to, identifying the subject's outcome, including the approximate time since the occurrence of this event, the presence and amount of recoverable tissue, the suitability of drug therapy, the effectiveness of various therapies, and the risk of future events.

[0188] Biomarker analysis can be performed in a variety of physical formats. For example, microtiter plates or automated applications can be used to facilitate the processing of large numbers of test samples. Alternatively, single sample formats can be developed to facilitate immediate treatment and diagnosis in a timely manner, for example, in an outpatient or emergency room setting.

[0189] Genomic DNA can be isolated by any means, including the use of commercially available kits. Briefly, if the DNA of interest is encapsulated in a cell membrane, the biological sample must be disrupted and dissolved by enzymatic, chemical, or mechanical means. Proteins and other contaminants can then be removed from the DNA solution, for example, by digestion with proteinase K. The genomic DNA is then recovered from the solution. This can be done by a variety of methods, such as salting out, organic extraction, or binding of DNA to a solid support. The choice of method is influenced by several factors, including time, cost, and the amount of DNA required. All clinical sample types, including tumorous or pre-neoplastic material, are suitable for use with this method, including cell lines, histological slides, biopsies, esophageal brushings, paraffin-embedded tissues, body fluids, stool, tissues, colonic effluent, urine, plasma, serum, whole blood, buffy coat, isolated blood cells, cells isolated from blood, and combinations thereof.

[0190] The present technology is not limited to the method used to prepare the sample and provide the nucleic acid for testing. For example, in some embodiments, DNA is isolated from a stool sample, or from a blood sample, or from a plasma sample, using direct gene capture or related methods, for example, as detailed in U.S. Patent Application No. 61 / 485,386.

[0191] The genomic DNA sample is then treated with at least one reagent, or a series of reagents, that distinguishes between methylated and unmethylated CpG dinucleotides within at least one marker that comprises a DMR (e.g., a DMR in Tables 1, 2, or 3).

[0192] In some embodiments, the reagent converts unmethylated cytosine bases at the 5' position to uracil, thymine, or another base that differs from cytosine in terms of hybridization behavior, although in some embodiments the reagent may be a methylation-sensitive restriction enzyme.

[0193] In some embodiments, the genomic DNA sample is treated to convert 5'-unmethylated cytosine bases to uracil, thymine, or another base that differs from cytosine in terms of hybridization behavior. In some embodiments, this treatment is carried out with bisulfite (e.g., bisulfite, disulfite) followed by alkaline hydrolysis.

[0194] The processed nucleic acid is then analyzed to determine the methylation status of the target gene sequence (at least one gene, genomic sequence, or nucleotide from a marker comprising at least one DMR, e.g., a DMR selected from the DMRs in Tables 1, 2, or 3). The analytical method can be selected from those known in the art, including those listed herein, e.g., QuARTS and MSP as described herein.

[0195] Such samples can be obtained by any number of means known in the art, as will be apparent to those skilled in the art. For example, urine and fecal samples are readily available, while blood, ascites, serum, or pancreatic juice samples can be obtained parenterally, for example, by using a needle and syringe. Cell-free or substantially cell-free samples can be obtained by subjecting the sample to a variety of techniques, including, but not limited to, centrifugation and filtration. While it is generally preferred to obtain samples without the use of invasive techniques, it may still be preferable to obtain samples such as tissue homogenates, tissue sections, and biopsy specimens.

[0196] 3. Treatment method In some embodiments, the present disclosure provides methods for treating a subject (e.g., a patient having or suspected of having esophageal cancer or precancer). According to these embodiments, the method comprises determining the methylation status or profile of one or more methylated DNA markers provided herein and administering a treatment to the patient based on the results of determining the methylation status. In some embodiments, the method comprises evaluating a DNA sample from the subject for copy number variations (CNVs) or copy number abnormalities (CNAs) and administering a treatment to the patient based on the results of the evaluation. In some embodiments, the method comprises evaluating an aneuploidy score (AS) in a DNA sample from the subject and administering a treatment to the patient based on the results of the evaluation.

[0197] The treatment can be administering a pharmaceutical compound, administering a vaccine, performing surgery, imaging the patient, performing another test, etc. In some embodiments, treating a subject includes methods of clinical screening, methods of prognostic evaluation, methods of monitoring treatment outcome, methods of identifying patients most likely to respond to a particular therapeutic treatment, methods of imaging a patient or subject, and methods for drug screening and development.

[0198] In some embodiments, a method for diagnosing a specific type of cancer in a subject is provided. As used herein, the terms "diagnosing" and "diagnosis" refer to a method by which a skilled artisan can estimate, and even determine, whether a subject suffers from a given disease or condition, or whether a subject is likely to develop a given disease or condition in the future. Those skilled in the art often make a diagnosis based on one or more diagnostic indicators, such as one or more biomarkers (e.g., one or more methylation markers, methylation marker genes, genes, DMRs, and / or DNA methylation markers, such as those disclosed herein), where the methylation status of the biomarkers indicates the presence, severity, or absence of the condition. In some embodiments, a diagnosis can be made based on one or more diagnostic indicators, such as CNV or aneuploidy, which can indicate the presence, severity, or absence of a condition (e.g., esophageal cancer or precancer).

[0199] Along with diagnosis, clinical prognosis of cancer is related to determining the aggressiveness of cancer and the possibility of tumor recurrence in order to plan the most effective treatment.If a more accurate prognosis can be made or even the potential risk of developing cancer can be evaluated, appropriate therapy, and in some cases, a less harsh therapy for the patient, can be selected.Evaluation of cancer biomarkers (for example, determining methylation profile and / or aneuploidy score) is useful for distinguishing subjects with good prognosis and / or low risk of developing cancer, who do not need therapy or require limited therapy, from subjects who are more likely to develop cancer or suffer from cancer recurrence, who can benefit from more intensive treatment.

[0200] As such, "making a diagnosis" or "diagnosing," as used herein, further includes determining the risk of developing cancer or determining a prognosis, which may be provided to predict a clinical outcome (with or without medical treatment), select an appropriate treatment (or whether treatment is effective), or monitor current treatment to potentially modify treatment, based on the measurement of a diagnostic biomarker (e.g., DMR) and / or aneuploidy score disclosed herein. Furthermore, in some embodiments of the subject matter disclosed herein, multiple determinations of biomarkers over time can be made to facilitate diagnosis and / or prognosis. Changes in the methylation profile and / or aneuploidy score over time can be used to predict clinical outcomes, monitor the progression of the cancer or cancer subtype, and / or monitor the effectiveness of appropriate therapies for the cancer. In such embodiments, for example, one may expect to see changes in the methylation status of one or more biomarkers (e.g., DMR) and / or aneuploidy score (and potentially one or more additional biomarkers, if monitored) disclosed herein.

[0201] The presently disclosed subject matter further provides, in some embodiments, a method for determining whether to initiate or continue cancer prevention or treatment in a subject. In some embodiments, the method includes providing a series of biological samples from a subject over a period of time, analyzing the series of biological samples to determine a methylation profile and / or an aneuploidy score in each of the biological samples, and comparing any measurable changes in the methylation profile and / or the aneuploidy score in each of the biological samples. Any changes over a period of time can be used to predict the risk of developing cancer, predict clinical outcomes, determine whether to initiate or continue cancer prevention or treatment, and determine whether a current treatment is effectively treating the cancer. For example, a first time point can be selected before the start of treatment, and a second time point can be selected sometime after the start of treatment. The methylation profile and / or the aneuploidy score can be measured in each of the samples collected at different time points, and qualitative and / or quantitative differences can be recorded. Changes in the methylation status and / or the aneuploidy score of biomarker levels from different samples can be correlated with a particular cancer risk, prognosis, determining treatment efficacy, and / or cancer progression in the subject. In some embodiments, the disclosed methods and compositions are for the treatment or diagnosis of disease at an early stage, e.g., before disease symptoms appear, hi some embodiments, the disclosed methods and compositions are for the treatment or diagnosis of disease at a clinical stage.

[0202] In some embodiments, multiple determinations of one or more diagnostic or prognostic biomarkers can be performed, and changes in the markers over time can be used to determine a diagnosis or prognosis. For example, a diagnostic marker can be determined a first time and then again a second time. In such embodiments, an increase in a marker from the first time to the second time can be diagnostic of a particular type or severity of cancer, or a given prognosis. Similarly, a decrease in a marker from the first time to the second time can indicate a particular type or severity of cancer, or a given prognosis. Furthermore, the degree of change in one or more markers can be related to the severity of cancer and future adverse events. Those skilled in the art will understand that, in certain embodiments, comparative measurements of the same biomarker can be performed at multiple time points, but a given biomarker can also be measured at one time point and a second biomarker at a second time point, and the comparison of these markers can provide diagnostic information.

[0203] As used herein, the phrase "determining a prognosis" refers to a method by which a person skilled in the art can predict the course or outcome of a condition in a subject. The term "prognosis" does not refer to the ability to predict the course or outcome of a condition with 100% accuracy, or that a given course or outcome is more or less likely to occur in a predictable manner based on the methylation status and / or aneuploidy score of a biomarker (e.g., DMR). Instead, those skilled in the art will understand that the term "prognosis" refers to a high probability of a certain course or outcome occurring, i.e., a high probability of a certain course or outcome occurring in a subject exhibiting a given condition compared to individuals who do not exhibit the condition. For example, the likelihood of a given outcome (e.g., suffering from a particular type of cancer) may be very low in individuals who do not exhibit the condition (e.g., who have a normal methylation status and / or a normal aneuploidy score of one or more DMRs).

[0204] In some embodiments, statistical analysis correlates the prognostic indicator with a predisposition to adverse outcomes. For example, in some embodiments, a methylation profile and / or aneuploidy score that differs from that in a normal control sample obtained from a patient without cancer, as determined by the level of statistical significance, may indicate that the subject is more likely to suffer from cancer than a subject with a methylation profile and / or aneuploidy score more similar to that in the control sample. In addition, changes in the methylation profile and / or aneuploidy score from baseline (e.g., "normal") levels may reflect the subject's prognosis, and the degree of change in the methylation profile and / or aneuploidy score may be related to the severity of adverse events. Statistical significance is often determined by comparing two or more populations and determining a confidence interval and / or p-value. See, for example, Dowdy and Wearden, Statistics for Research, John Wiley & Sons, New York, 1983, incorporated herein by reference in its entirety. Exemplary confidence intervals of the present subject matter are 90%, 95%, 97.5%, 98%, 99%, 99.5%, 99.9%, and 99.99%, while exemplary p-values ​​are 0.1, 0.05, 0.025, 0.02, 0.01, 0.005, 0.001, and 0.0001.

[0205] In other embodiments, a threshold degree of change in the methylation status and / or aneuploidy score of a prognostic or diagnostic biomarker (e.g., DMR) disclosed herein can be established, and the degree of change in the methylation profile and / or aneuploidy score in a biological sample is simply compared to the threshold degree of change in the methylation profile and / or aneuploidy score. For example, preferred threshold changes in the methylation status of the biomarkers provided herein are about 5%, about 10%, about 15%, about 20%, about 25%, about 30%, about 50%, about 75%, about 100%, and about 150%. In yet another embodiment, a "nomogram" can be established, whereby the methylation status of a prognostic or diagnostic indicator (biomarker or combination of biomarkers) is directly correlated with the associated predisposition for a given outcome. Those skilled in the art are familiar with the use of such nomograms to correlate two values, with the understanding that because measurements are referenced to individual samples rather than population averages, the uncertainty in the measurements is the same as the uncertainty in the marker concentration.

[0206] In some embodiments, a control sample is analyzed simultaneously with the biological sample to allow results obtained from the biological sample to be compared to those obtained from the control sample. It is further contemplated that a standard curve can be provided and assay results of the biological sample can be compared to the standard curve. Such a standard curve displays the methylation status and / or aneuploidy score of the biomarker according to assay units, e.g., fluorescent signal intensity if a fluorescent label is used. A standard curve can be provided for control methylation status and / or aneuploidy score in normal tissue using samples collected from multiple donors. In certain embodiments of the method, a subject is identified as having cancer by identifying an abnormal methylation status and / or aneuploidy score of one or more DMRs provided herein in a biological sample obtained from the subject. In other embodiments of the method, detecting an abnormal methylation status and / or an abnormal aneuploidy score of one or more such biomarkers in a biological sample obtained from the subject results in the subject being identified as having cancer.

[0207] Marker analysis can be performed separately or simultaneously with additional markers within a single test sample. For example, it is possible to combine several markers in one test to efficiently process multiple samples and potentially provide higher diagnostic and / or prognostic accuracy. Furthermore, those skilled in the art will recognize the value of testing multiple samples from the same subject (e.g., at successive time points). Testing such serial samples allows for the identification of changes in the methylation status of markers over time. Changes in methylation status, and the absence of changes in methylation status, can provide useful information about disease states, including, but not limited to, identifying the subject's outcome, including the approximate time since the occurrence of this event, the presence and amount of recoverable tissue, the suitability of drug therapy, the effectiveness of various therapies, and the risk of future events.

[0208] Biomarker analysis can be performed in a variety of physical formats. For example, microtiter plates or automated applications can be used to facilitate the processing of large numbers of test samples. Alternatively, single sample formats can be developed to facilitate immediate treatment and diagnosis in a timely manner, for example, in an outpatient or emergency room setting.

[0209] As described above, depending on the embodiment of the disclosed method, detecting a change in the methylation profile and / or aneuploidy score can be a qualitative or quantitative determination. Thus, diagnosing a subject as having or at risk of developing a particular type of cancer indicates that a certain threshold measurement is made (e.g., the methylation state of one or more biomarkers in a biological sample changes from a predetermined control methylation state). In some embodiments of the method, the control methylation profile and / or aneuploidy score is any detectable methylation profile and / or aneuploidy score. In other embodiments of the method in which a control sample is tested simultaneously with the biological sample, the predetermined methylation profile and / or aneuploidy score is the methylation profile and / or aneuploidy score in the control sample. In other embodiments of the method, the predetermined methylation profile and / or aneuploidy score is based on and / or determined by a standard curve. In other embodiments of the method, the predetermined methylation profile and / or aneuploidy score is a specific profile or range of profiles. Thus, a predetermined methylation profile and / or aneuploidy score may be selected based in part on the embodiment of the method being performed, the desired specificity, etc., within acceptable limits that will be apparent to one of skill in the art.

[0210] Furthermore, with respect to diagnostic methods, preferred subjects are vertebrate subjects. Preferred vertebrates are warm-blooded animals, and preferred warm-blooded vertebrates are mammals. Preferred mammals are most preferably humans. As used herein, the term "subject" includes both human and animal subjects. Accordingly, veterinary uses are provided herein. Accordingly, embodiments of the present disclosure provide for the diagnosis of mammals, such as humans, as well as mammals of endangered importance, such as the Amur tiger, mammals of economic importance, such as animals raised on farms for human consumption, and / or animals of social importance to humans, such as animals kept as pets or in zoos. Examples of such animals include, but are not limited to, carnivores, such as cats and dogs; swine, including pigs, hogs, and wild boars; ruminants and / or ungulates, such as cows, oxen, sheep, giraffes, deer, goats, bison, and camels; and horses. Thus, diagnostics and treatments for livestock, including but not limited to domestic pigs, ruminants, ungulates, horses (including racehorses), and the like, are also provided.

[0211] 4. Samples, Kits, and Controls Embodiments of the present disclosure provide techniques for screening for multiple types of esophageal cancer or precancer from a biological sample. According to these embodiments, the present disclosure includes, but is not limited to, methods and compositions for detecting the presence of multiple types and / or subtypes of esophageal cancer or precancer from a biological sample. In some embodiments, the biological sample is a tissue sample, a blood sample, a plasma sample, a serum sample, a whole blood sample, a buffy coat sample, a secretion sample, an organ secretion sample, a cerebrospinal fluid (CSF) sample, a saliva sample, a urine sample, and / or a stool sample. In some embodiments, the tissue sample is an esophageal tissue sample. In some embodiments, the esophageal sample is obtained from an esophageal biopsy or by swabbing, brushing, or using a sponge capsule device. In some embodiments, the subject is a human.

[0212] In other embodiments, "sample," "test sample," and "biological sample" refer to a fluid sample containing or suspected of containing a methylated DNA marker of the present disclosure. A sample may be derived from any suitable source. In some cases, a sample may include a liquid, a fluid particulate solid, or a solid particle suspension. In some cases, a sample may be processed prior to analysis as described herein. For example, a sample may be separated or purified from its source prior to analysis. In certain examples, the source is a mammalian (e.g., human) bodily substance (e.g., bodily fluid, blood (whole blood, buffy coat, serum, plasma, etc.), urine, saliva, sweat, sputum, semen, mucus, tears, lymph, amniotic fluid, interstitial fluid, cerebrospinal fluid, feces, tissue, organ, one or more dried blood spots, etc.). Tissues include, but are not limited to, esophageal tissue, stomach tissue, pancreatic tissue, bile duct / liver tissue, and colon tissue. A sample may be a liquid sample or a liquid extract of a solid sample. In some embodiments, the source of the sample may be an organ or tissue, such as a biopsy and / or esophageal brushing (e.g., an endoscopic esophageal brushing), which may be solubilized by tissue disruption / cell lysis.

[0213] A wide range of fluid sample volumes can be analyzed. In some exemplary embodiments, the sample volume can be about 0.5 nL, about 1 nL, about 3 nL, about 0.01 μL, about 0.1 μL, about 1 μL, about 5 μL, about 10 μL, about 100 μL, about 1 mL, about 5 mL, about 10 mL, etc. In some cases, the volume of the fluid sample is about 0.01 μL to about 10 mL, about 0.01 μL to about 1 mL, about 0.01 μL to about 100 μL, or about 0.1 μL to about 10 μL.

[0214] In some cases, the fluid sample may be diluted before use in the assay. For example, in embodiments where the source containing the methylated DNA marker is a human bodily fluid (e.g., blood, serum, secretions), the bodily fluid may be diluted with an appropriate solvent (e.g., a buffer such as PBS buffer). The fluid sample may be diluted about 1-fold, about 2-fold, about 3-fold, about 4-fold, about 5-fold, about 6-fold, about 10-fold, about 100-fold, or more before use. In other cases, the fluid sample is not diluted before use in the assay.

[0215] In some cases, the sample may undergo pre-analysis treatment. Pre-analysis treatment may provide additional functions, such as removal of nonspecific proteins and / or effective yet inexpensively implemented mixing functions. Common methods of pre-analysis treatment may include the use of electrokinetic trapping, AC electrokinetics, surface acoustic waves, isotachophoresis, dielectrophoresis, electrophoresis, or other pre-concentration techniques known in the art. In some cases, the fluid sample may be concentrated before use in the assay. For example, in embodiments where the source containing the methylated DNA marker is a human bodily fluid (e.g., blood, serum, secretions), the bodily fluid may be concentrated by precipitation, evaporation, filtration, centrifugation, or a combination thereof. The fluid sample may be concentrated about 1-fold, about 2-fold, about 3-fold, about 4-fold, about 5-fold, about 6-fold, about 10-fold, about 100-fold, or more, prior to use.

[0216] It may be desirable to include a control. The control may be analyzed simultaneously with the sample from the subject, as described above. Results obtained from the subject sample may be compared to results obtained from the control sample. A standard curve may be provided to which the assay results of the sample may be compared. Such a standard curve shows the level of one or more methylated DNA markers as a function of assay units. Samples from multiple donors can be used to provide standard curves for reference levels of methylated DNA markers in normal, healthy tissue, as well as "risk" levels of methylated DNA markers in tissue from donors who may have one or more characteristics of esophageal cancer or precancer.

[0217] Embodiments of the present disclosure also include kits for carrying out the methods described herein. The kits include embodiments of the compositions, devices, apparatus, etc. described herein, as well as instructions for using the kit. Such instructions describe appropriate methods for preparing an analyte from a sample, e.g., methods for collecting a sample and preparing nucleic acid from the sample. Individual components of the kit are packaged in suitable containers and packaging (e.g., vials, boxes, blister packs, ampoules, jars, bottles, tubes, etc.), and the components are packaged together in suitable containers (e.g., box(es)) for convenient storage, transportation, and / or use by the user of the kit. It is understood that liquid components (e.g., buffers) may be provided in lyophilized form to be reconstituted by the user. The kit may also include a control or reference for assessing, validating, and / or ensuring the performance of the kit. For example, a kit for assaying the amount of nucleic acid present in a sample may include a control containing a known concentration of the same or another nucleic acid for comparison, and in some embodiments, a detection reagent (e.g., primers) specific for the control nucleic acid. The kit is suitable for use in a clinical setting and, in some embodiments, for use in the user's home. The components of the kit, in some embodiments, provide the functionality of a system for preparing a nucleic acid solution from a sample. In some embodiments, certain components of the system are provided by the user.

[0218] In some embodiments, the disclosure provides compositions (e.g., reaction mixtures). In some embodiments, the disclosure provides compositions comprising a nucleic acid containing a DMR and a reagent capable of modifying DNA in a methylation-specific manner (e.g., a methylation-sensitive restriction enzyme, a methylation-dependent restriction enzyme, and a bisulfite reagent) (e.g., a methylation-sensitive restriction enzyme, a methylation-dependent restriction enzyme, a ten-eleven translocation (TET) enzyme (e.g., human TET1, human TET2, human TET3, mouse TET1, mouse TET2, mouse TET3, Naegleria TET (NgTET), Coprinopsis cinerea (CcTET)), or a variant thereof), a borane reducing agent). Some embodiments provide compositions comprising a nucleic acid containing a DMR and an oligonucleotide described herein. Some embodiments provide compositions comprising a nucleic acid containing a DMR and a methylation-sensitive restriction enzyme. Some embodiments provide compositions comprising a nucleic acid containing a DMR and a polymerase.

[0219] In some embodiments, the technology described herein is associated with a programmable machine designed to perform an array of arithmetic or logical operations, such as those provided by the methods described herein. For example, some embodiments of the technology are associated with (e.g., implemented in) computer software and / or computer hardware. In one aspect, the technology relates to a computer that includes a form of memory, elements for performing arithmetic and logical operations, and a processing element (e.g., a microprocessor) for executing a set of instructions for reading, manipulating, and storing data (e.g., the methods provided herein). In some embodiments, the microprocessor is part of a system for determining methylation profiles and / or aneuploidy scores (e.g., of one or more DMRs in Tables 1, 2, or 3), comparing methylation profiles and / or aneuploidy scores, generating standard curves, determining Ct values, calculating methylation fractions, frequencies, or percentages, identifying CpG islands, determining assay or marker specificity and / or sensitivity, calculating ROC curves and associated AUCs, analyzing sequences, or any of the foregoing. In some embodiments, the microprocessor is part of a system for determining methylation profiles and / or aneuploidy scores (e.g., of one or more DMRs in Tables 1, 2, or 3), comparing methylation profiles and / or aneuploidy scores, generating standard curves, determining Ct values, calculating methylation fractions, frequencies, or percentages, identifying CpG islands, determining assay or marker specificity and / or sensitivity, calculating ROC curves and associated AUCs, analyzing sequences, or any of the foregoing described herein or known in the art.

[0220] In some embodiments, the software or hardware component receives the results of the multiple assays and determines to report a single value result indicative of cancer risk to a user based on the results of the multiple assays (e.g., determining the methylation status and / or aneuploidy score of one or more DMRs in Table 1, 2, or 3). Related embodiments calculate a risk factor (e.g., determining the methylation status and / or aneuploidy score of one or more DMRs in Table 1, 2, or 3) based on a mathematical combination (e.g., weighted combination, linear combination) of the results from the multiple assays. In some embodiments, the methylation status of the DMRs defines a dimension and can have values ​​in a multidimensional space, and the coordinate defined by the methylation status of the multiple DMRs is a result (e.g., for reporting to a user or related to cancer risk).

[0221] Various embodiments of the present disclosure relate to multiple programmable devices operating in concert to perform the methods described herein. For example, in some embodiments, multiple computers (e.g., connected by a network) can operate in parallel to collect and process data, for example, in an implementation of cluster computing or grid computing or some other distributed computing architecture that relies on complete computers (on-board CPU, storage, power, network interfaces, etc.) connected to a network (private, public, or the Internet) by traditional network interfaces such as Ethernet, fiber optics, etc., or by wireless networking technology.

[0222] For example, some embodiments provide a computer including a computer-readable medium. The embodiment includes a random access memory (RAM) coupled to a processor. The processor executes computer-executable program instructions stored in the memory. Processors such as these may include microprocessors, ASICs, state machines, or other processors, and may be any of a number of computer processors, such as processors from Intel Corporation of Santa Clara, California, or Motorola Corporation of Schaumburg, Illinois. Processors such as these may include or be in communication with a medium, such as a computer-readable medium, that stores instructions that, when executed by the processor, cause the processor to perform the steps described herein.

[0223] In some embodiments, the computer is connected to a network. The computer may also include multiple external or internal devices, such as a mouse, CD-ROM, DVD, keyboard, display, or other input or output devices. Examples of computers include personal computers, digital assistants, personal digital assistants, cellular phones, mobile phones, smartphones, pagers, digital tablets, laptop computers, Internet appliances, and other processor-based devices. Generally, these computers related to aspects of the technology provided herein can be any type of processor-based platform, running any operating system, such as Microsoft Windows, Linux, UNIX, or Mac OS X, and capable of supporting one or more programs, including the technology provided herein. Some embodiments include personal computers running other application programs (e.g., applications). Applications can be stored in memory and can include, for example, word processing applications, spreadsheet applications, email applications, instant messenger applications, presentation applications, Internet browser applications, calendar / organizer applications, and any other application executable by a client device. All such components, computers, and systems described herein as related to the technology may be logical or virtual.

[0224] In some embodiments, the present disclosure provides a system for screening for esophageal cancer or precancer in a sample obtained from a subject. Exemplary embodiments of the system include, for example, a system for screening for esophageal cancer or precancer in a sample obtained from a subject (e.g., a tissue sample, a blood sample, a plasma sample, a serum sample, a whole blood sample, a buffy coat sample, a secretion sample, an organ secretion sample, a cerebrospinal fluid (CSF) sample, a saliva sample, a urine sample, and / or a stool sample). In some embodiments, the system includes an analysis component configured to determine a methylation status and / or an aneuploidy score of one or more methylation markers in the sample, a software component configured to compare the methylation status and / or the aneuploidy score of the one or more methylation markers in the sample with a control or reference sample recorded in a database, and an alert component configured to alert a user of a cancer-related condition.

[0225] In some embodiments, the alert is determined by a software component that receives results from multiple assays (e.g., determines the methylation status and / or aneuploidy score of one or more methylation markers), calculates a value or result, and reports based on the multiple results.

[0226] Some embodiments provide a database of weighting parameters associated with each methylation marker provided herein for use in calculating a value or result and / or reporting an alert to a user (e.g., a doctor, nurse, clinician, etc.). In some embodiments, all results from multiple assays are reported. In some embodiments, one or more results are used to provide a score, value (e.g., an aneuploidy score), or result based on a combination of one or more results from multiple assays that indicates risk of cancer in a subject. Such methods are not limited to particular methylation markers. In such methods and systems, the one or more methylation markers comprise bases in a DMR selected from the DMRs of Tables 1, 2, and 3.

[0227] In this detailed description of various embodiments, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the disclosed embodiments. However, those skilled in the art will understand that these various embodiments may be practiced with or without these specific details. In other instances, structures and mechanisms are shown in block diagram form. Furthermore, those skilled in the art will readily appreciate that the specific order in which the methods are presented and performed is illustrative, and that the order can be changed and still be within the spirit and scope of the various embodiments disclosed herein.

[0228] The various components of the kit may optionally be provided in suitable containers. The kit may further include a container for holding or storing a sample (e.g., a container or cartridge for urine, whole blood, buffy coat, plasma, serum sample, tissue, or bodily secretion sample). If desired, the kit may also optionally contain reaction vessels, mixing vessels, and other components to facilitate preparation of reagents or test samples. The kit may also include one or more instruments to aid in obtaining a test sample, such as a syringe, pipette, forceps, measuring spoon, etc. In some embodiments, the instrument is a collection device. In some embodiments, the kit includes a collection device for endoscopic brushing or non-endoscopic whole esophageal brushing or swabbing using a tethered device (e.g., a capsule sponge, balloon, or other device). In some embodiments, a biological sample is obtained from a subject, and the method further includes extracting a DNA sample from the biological sample. In some embodiments, the biological sample is collected with a collection device having an absorbent member capable of collecting the biological sample upon contact. In some embodiments, the absorbent member is a sponge configured for insertion into an orifice. [Example]

[0229] 5. Working Example It will be appreciated by those skilled in the art that other suitable modifications and adaptations of the disclosed methods described herein are readily applicable and discernible and may be made using suitable equivalents without departing from the scope of the disclosure or the aspects and embodiments disclosed herein. Having thus described the disclosure in detail, it will be more clearly understood by reference to the following examples. These examples are intended merely to illustrate some aspects and embodiments of the disclosure and should not be construed as limiting the scope of the disclosure. The disclosures of all journal references, U.S. patents, and publications referenced herein are incorporated herein by reference in their entirety.

[0230] The present disclosure has multiple aspects, illustrated by the following non-limiting examples.

[0231] Example 1 Experiments were conducted to identify a panel of genetic markers that can distinguish non-dysplastic Barrett's esophagus (NDBE) from precancerous conditions such as low-grade dysplasia (LGD) or high-grade dysplasia (HGD), and / or cancerous conditions such as esophageal adenocarcinoma (EAC), based on variably methylated regions (DMRs) and / or DNA copy number variations (CNVs) in candidate genetic markers.

[0232] Reduced representation bisulfite sequencing (RRBS) studies were performed on tissue biopsy samples, along with whole-genome sequencing (WGS) studies on endoscopic brushing samples. Briefly, for RRBS tissue biopsy studies, an average of approximately 50 million mapped reads were generated per sample, with 4 million CpGs per sample at over 10x coverage. As described in Materials and Methods below, after employing a proprietary marker discovery pipeline algorithm, logistic analysis, and initial filtering, 4501 DMRs with an average length of 156 bp were identified. These were mapped to 1549 annotated genes and 758 unannotated regions of the genome. Further filtering to remove control sample CpG methylation and imposing a case / control FC cutoff resulted in a total of 1196 annotated and unannotated regions. The endoscopic brushing WGS study generated approximately 500 million mapped, de-duplicated reads per sample on average, yielding 24 million CpGs per sample at over 10x coverage. A total of 7,345 DMRs were identified with an average length of 72 bp. These were mapped to 2,351 genes and 1,374 unannotated regions. After applying subsequent filters and cutoffs, a total of 3,534 regions remained. Given that RRBS is an enzyme-mediated enrichment of a genome of 3.2 billion bases to less than 5 million, the WGS count was expected to exceed the RRBS count. Due to the size selection step and the CCGG specificity of the MspI enzyme, RRBS lacks potential CpG islands and promoter regions that do not harbor these tetramers in close proximity.

[0233] Genomic regions from 1,196 RRBS DMRs were merged with 3,534 WGS DMRs to determine common sites of hypermethylation. Regions were required to overlap or have endpoints within 500 bases of each other. A total of 227 DMRs met these criteria. These were then narrowed down to 156 DMRs. The majority of exclusions were regions that did not show substantial concordant CpG methylation, characteristic of bona fide functional events in methylation-mediated tumor suppressor gene inactivation or oncogene enhancer element activation. Others were regions with strong concordant methylation in the NDBE cohort. Some regions were unusually short (e.g., less than 60 bases), limiting their informativeness and imposing constraints on capture and methylation-specific amplification. In addition to the 156 merged regions, 43 additional candidate WGS DMRs were identified that were not present in the RRBS data (presumably due to the enrichment exclusion step in the protocol). The complete set of 199 DMRs with gene annotations or SEQ ID NOs are listed in Table 1. [Table 1-1] [Table 1-2] [Table 1-3] [Table 1-4] [Table 1-5] [Table 1-6] [Table 1-7] [Table 1-8] [Table 1-9] [Table 1-10]

[0234] DMR performance characteristics from the whole genome sequencing (WGS) study are provided in Table 2. Additionally, 156 DMRs were also identified in a separate reduced bisulfite sequencing (RRBS) study using independent tissue biopsies, thus validating the results of the WGS study. [Table 2-1] [Table 2-2] [Table 2-3] [Table 2-4] [Table 2-5] [Table 2-6]

[0235] Table 3 highlights the top DMRs with AUC > 0.75 and FC > 5 in both the independent WGS and RRBS datasets. The DMRs in Table 3 represent the subset of DMRs that showed the highest performance in both studies. [Table 3]

[0236] Taken together, the merged data validate 156 DMRs. The two sequencing studies used different, independent sample sets: the RRBS data were from clinical tissue biopsies, while the WGS data were from endoscopic esophageal brushings. However, common genes were found in both of these sequences, often in overlapping sequence fragments rather than simply in widely divergent regions, and with similar performance characteristics.

[0237] Example 2 To distinguish Barrett's esophagus (BE) and / or normal esophageal samples from precancerous high-grade dysplasia (HGD) and / or esophageal adenocarcinoma, we also performed experiments to assess copy number variations (CNV), including ploidy and aneuploidy determinations. Notably, another important aspect of the present disclosure is the ability to utilize sequencing reads from cytosine-converted genomes to determine copy number abnormalities (CNAs), including ploidy and aneuploidy. Most techniques use direct NGS on wild-type, unconverted DNA. Here, methylation and CNV readouts were simultaneously acquired from the same chemistry / dataset. In WGS brushing studies, aneuploidy calls were performed using CNVkit and AneuploidyScore software, and a numerical score was generated using the NE cohort as a reference. These normal esophageal (NE) samples were defined as integer multiples (two copies of the 22 somatic chromosomes) and confirmed with chromosome self-referencing scatter plots (see, for example, Figure 3). The mapped, de-duplicated reads were used for analysis and were a priori corrected for repetitive sequences, GC content, and PCR duplication. The aneuploidy score (AS) calculated for each sample is the total number of gains and losses at the chromosome arm level, adjusted for ploidy. For example, see Figure 1 for EAC, HGD, and NDBE sample arm levels and segmented aneuploidy scores (y-axis, signal).

[0238] Figure 2 is a violin plot of the data in Figure 1, showing the increasing level and frequency of aneuploidy from non-dysplastic to HGD to EAC. The presence of chromosome arm gains or losses was observed in 14 (78% [52-94%]) EAC, 7 (39% [17-64%]) HGD, and 4 (22% [6-48%]) NDBE patients (specificity 78%). These calls were confirmed by visual analysis of CNVpytor chromosome scatter plots for all samples. See representative plots in Figure 3 (AS0), Figure 4 (AS0), Figure 5 (AS4), and Figure 6 (AS9).

[0239] Example 3 We also conducted experiments to determine whether copy number abnormalities (CNAs) and DRMs could complement each other for the detection of BE-associated HGD and EAC from NDBE (and NE) samples. In this example, one DMR, MAFB (v-maf myomedullary fibrosarcoma oncogene homolog B), was evaluated, with the highest AUC in the above data and detected in both RRBS and WGS studies. The discriminatory region is approximately 1400 bp in length and is highly informative (Figures 7 and 8). Hypermethylation of MABF was observed in 16 (89% [65-99%]) EAC, 14 (78% [52-94%]) HGD, and 3 (17% [4-42%]) NDBE brushing patients (specificity 83%). Combining methylation and aneuploidy, 17 (94% [73-100%]) EAC and 16 (89% [65-99%]) HGD were correctly classified with a fixed specificity of 85% [58-96%] in the NDBE samples. One EAC and two HGD were missed by both marker classes (Table 4 and Figure 9). [Table 4-1] [Table 4-2] [Table 4-3]

[0240] When an additional DMR (e.g., LONRF2) was added to the panel with 100% specificity, one of the missed HGD patients was correctly called without any false-positive hits. Furthermore, subsequent inferential cell fraction analysis (using TGCA epigenetic profiles of different cell types) was performed on all WGS samples (data not shown) and found that for the remaining two incorrectly called cases, no columnar epithelial signal was evident, but both squamous epithelial signals were very strong. These two brushes may have been improperly sampled or intermixing may have occurred during classification. The other cases and controls in the WGS study had appropriate epigenetic profiles.

[0241] Taken together, CNV and DMR analysis complement the classification of BE-HGD / EAC from NDBE. Both CNV and DMR analysis can be assayed from endoscopic brush specimens, enhancing endoscopic surveillance. For example, aneuploidy develops late in tumor development, while DNA methylation is a very early event, which explains why even NDBE samples are generally (but not always) heavily hypermethylated and why the number of dysplasia-specific DMRs is lower than the number of DMRs typically observed in cancer / normal cohorts. Combining these genomic and epigenetic changes is an advantageous strategy for dysplasia detection.

[0242] Example 4 Validation of the markers identified in Example 1 was performed by applying the 199 DMRs identified from Table 1 to an independent set of 169 endoscopic esophageal brushing samples. Shallow whole-methylome sequencing was also performed on these samples to confirm the preliminary aneuploidy results. Based on cross-validated DMR selection, a model of four DMRs, including KL, PGBD5, ROR2, and LMX1B, achieved a cross-validated AUC of 0.87 (0.81-0.94, 95% CI) for identifying HGD / EAC. Incorporating the tMAD score of CNAs achieved a cross-validated AUC of 0.92 (0.87-0.98) for identifying HGD / EAC. At an 80% specificity threshold, the 4-DMR model identified 35 (85% [71-94%]) EAC, 23 (88% [70-98%]) HGD, 14 (63% [41-83%]) LGD, and 8 (19% [9-34%]) NDBE patients (observed specificity 81%). The 4-DMR+tMAD model identified 38 (93% [80-98%]) EAC, 23 (88% [70-98%]) HGD, 11 (50% [28-72%]) LGD, and 8 (19% [9-34%]) NDBE patients (observed specificity 81%).

[0243] In this validation study, CNA and methylated DNA markers (individually and in combination) showed promising discriminatory properties for distinguishing BE-HGD / EAC from NDBE. Both types of markers can be assayed from endoscopic brush specimens and therefore can be utilized for molecular augmentation of histological analysis in endoscopic surveillance.

[0244] 6. Materials and Methods Marker Discovery. The following materials and methods were used to identify various DNA methylation markers capable of distinguishing NDBE, HGD, and EAC from controls in biological samples. Briefly, DNA was extracted from prospectively collected esophageal brushing samples from patients with NDBE (18), HGD (18), and EAC (18), along with 17 non-BE squamous epithelial samples (NE). High-volume endoscopic brushings were obtained from the squamous epithelium and distal 5 cm of the heart of NE patients and from visible BE mucosa of BE patients. Clinical, endoscopic, and histological data were abstracted, and histology was verified by an expert gastrointestinal pathologist prior to experimental execution. Enzyme-converted Methyl-seq (New England Biolabs) libraries were prepared from the extracted DNA and sequenced on an Illumina NovaSeq 6000 system. Variable methylation regions (DMRs) were identified from CpGs with ≥5x read coverage. The positivity rate of DMRs was set at 85% specificity in the NDBE group. CNAs were called by CNV Kit using the NE group as the reference, and aneuploidy events were determined by AneuploidyScore using their default settings. Logistic regression was used to evaluate combinations of methylation and aneuploidy status.

[0245] Patient Samples. Tissue samples were obtained from the Mayo Clinic Biospecimen Repository under institutional IRB oversight. Endoscopic esophageal brushing samples were collected in accordance with Mayo IRB protocol 15-004540. The two sample types were not patient-matched but were derived from completely independent individuals. Samples included FFPE esophageal biopsies from patients with and without Barrett's esophagus / EAC, esophageal brushings from patients with and without Barrett's esophagus / EAC, and normal gastric cardia biopsies (FFPE).

[0246] Samples were selected in strict compliance with the target study approval and inclusion / exclusion criteria. Exclusion criteria included: patients under 18 years of age; patients receiving chemotherapy-class drugs to treat primary esophageal cancer before tissue collection; patients receiving therapeutic radiation therapy or ablation therapy (photodynamic therapy, radiofrequency, or cryotherapy) for primary esophageal cancer or LGD or HGD or squamous cell dysplasia / carcinoma before sample collection (note: EMR or ESD alone does not exclude); patients having a history of disease of higher grade than the target pathology (i.e., low-grade / high-grade / adenocarcinoma); and patients having a history of pancreatic cancer, gastric cancer, liver cancer (HCC or cholangiocarcinoma), pustular carcinoma, or duodenal cancer within the last 5 years prior to sample collection.

[0247] Exclusion criteria for normal squamous epithelium and normal heart included the following: patients had a history of esophageal squamous dysplasia or carcinoma or a history of BE before sample collection; patients had a history of eosinophilic esophagitis before sample collection; patients had erosive esophagitis at the time of sample collection; patients had a history of gastrointestinal metaplasia or dysplasia before sample collection; patients had undergone ablation therapy (photodynamic therapy, radiofrequency therapy, or cryotherapy) for esophageal or gastric disease.

[0248] Exclusion criteria for non-dysplastic BE (NDBE) include: the patient has documented dysplasia sampled from Barrett's esophagus (either before or after the date of the specimen used for this project); the patient has undergone surgical treatment / biopsy / ablation (photodynamic therapy, radiofrequency, or cryotherapy) to treat (not just sample) a tumorous lesion in the esophagus (note: EMR or ESD is not excluded).

[0249] Inclusion criteria for normal squamous epithelium and normal heart include patients having normal histology in squamous epithelium and / or cardiac tissue (biopsy or other sample type).

[0250] Inclusion criteria for non-dysplastic BE (NDBE) include: patients have Barrett's histology without dysplasia; patients have one visit before the target visit and one visit after the target visit with no evidence of dysplasia.

[0251] Inclusion criteria for LEG, HGD, and EAC included patients having Barrett's histology with dysplasia or esophageal adenocarcinoma.

[0252] Brushings consisted of 18 esophageal adenocarcinomas (EAC), 18 Barrett's with high-grade dysplasia (HGD), 18 non-dysplastic Barrett's esophagus (NDBE), and 17 normal esophageal squamous epithelium (NE). The latter were collected from the distal 5 cm of the esophagus and included several cardias. Clinical, endoscopic, and histological data were abstracted, and the latter were reviewed by an independent pathologist. [Table 5]

[0253] Tissues consisted of three EACs, three gastroesophageal junction cancers (GEJCs), 18 HGDs, 13 Barrett's with low-grade dysplasia (LGDs), 16 NDBEs, 10 NEs, and 17 cystic gastric cardia (GC) specimens. Tissues were grossly dissected and histologically reviewed by an expert gastrointestinal pathologist. Samples were age-matched, randomized, and blinded. DNA from the tissues was purified using the Qiagen QIAmp FFPE Tissue Kit (Qiagen, Germantown, MD). DNA was repurified with AMPure XP beads (Beckman-Coulter, Brea, CA) and quantified with PicoGreen (Thermo-Fisher, Waltham, MA). DNA integrity was assessed using qPCR. Brushing samples were collected and stored in 1 mL of Qiagen cell lysis buffer, stored at -80°C. Subsequently, DNA was extracted using a Qiagen Puregene kit, modified to a sample volume of 1 ml.

[0254] Sequencing. Reduced bisulfite sequencing (RRBS) sequencing libraries from tissue genomic DNA were prepared using the NuGEN Ovation RRBS Methil-Seq kit with modifications (Tecan Genomics, Redwood City, CA). Samples were aligned in a 4-plex format and sequenced by the Mayo Genomics Facility on an Illumina HiSeq 4000 instrument (Illumina, San Diego, CA). Reads were processed using Illumina pipeline modules for image analysis and base calling. Secondary analysis was performed using SAAP-RRBS, a bioinformatics suite developed by Mayo. Briefly, reads were cleaned using Trim-Galore and aligned to the GRCh37 / hg19 reference genome constructed with BSMAP. For CpGs with coverage ≥10x and base quality score ≥20, methylation rates were determined by calculating C / (C+T) or, conversely, G / (G+A) in the case of reads mapping to the opposite strand.

[0255] For endoscopic brushing samples, 120 ng of DNA was sheared to approximately 300 bp using a Covaris LE220 sonicator (Covaris, Woburn, MA). The sample was concentrated into 50 μL and whole-genome sequencing (WGS) libraries were prepared using the NEBNext Enzymatic Methyl-seq kit (New England Biolabs, Ipswich, MA). Briefly, the sheared samples underwent end repair, A-tailing, and indexed adapter ligation. The libraries were then subjected to an enzymatic conversion step, amplification, and differentiation between methylated and unmethylated cytosines. The converted libraries were amplified, quantified, and pooled to normalized input in a 24plex format. A total of three pools were prepared and sent to the Mayo Genomics Core for sequencing on an Illumina NovaSeq system (Illumina, San Diego, CA) using an S4 flow cell (PE 150 cycles). Primary and secondary analyses were performed as described above.

[0256] Biomarker Selection. RRBS and WGS data were analyzed separately but with similar metrics. CpGs for inclusion had a coverage depth of 5 reads (or more) and a variance across subgroups of >0. Specifically, CpGs were ranked by their hypermethylation rate, i.e., the number of methylated cytosines at a given locus relative to the total number of cytosines at that site. For NE and GC controls, the ratios for NDBE samples had to be ≤0.01 (1%) and ≤0.05 (5%). CpGs that did not meet these criteria were discarded from the results. Candidate CpGs were then linked to differentially methylated regions (DMRs) ranging from approximately 40 to 2200 bp, depending on genomic location, with a minimum cutoff of 5 CpGs / 200 bp region. The area under the curve (AUC) of these regions had to be greater than 0.60, and the methylation fold change (FC) between cases and NDBE controls had to be greater than 3 (compared to greater than 5 for RRBS data, where tissues have higher cellular purity). For each candidate region, a 2D matrix was created comparing individual CpGs on a sample-by-sample basis for both cases and controls. Final selection required cases showing, at the sample-by-sample level, stretches of coordinated and continuous hypermethylation of individual CpGs across the entire DMR sequence. Conversely, for control samples, the methylation CpG pattern had to be highly probabilistic. DMRs in which one or more NDBE samples had a CpG mean methylation value greater than the mean methylation of all EACs and HGDs combined for that DMR were discarded.

[0257] Following regression, DMRs were ranked by the area under the receiver operating characteristic curve (AUC) and the difference in fold change between cases and controls. Because independent validation was planned in advance, no adjustment for false positives was performed at this stage.

[0258] Biomarker fusion / validation. A subset of DMRs was selected for further expansion. These were identified by merging filtered DMRs from two independent patient and sample type discovery datasets that were generated and analyzed separately. This was not done solely at the gene annotation level, but at specific genomic locations within genes. Thus, eligible regions were those in which the genomic coordinates from the two discoveries overlapped or whose endpoints were 500 bp or less apart.

[0259] Copy number variation. Copy number variation events from WGS next-generation sequencing (NGS) brushing data were called by CNVkit (github.com / etal / cnvkit) using the NE group as a reference, and aneuploidy events were determined by AneuploidyScore using default settings (github.com / quevedor2 / aneuploidy_score). Scatter plots were generated using CNVpytor software (github.com / abyzovlab / CNVpytor) with each sample as its own reference, using a 10K genome bin size.

[0260] Statistics. RRBS and WGS results were logistically analyzed for the performance of individual DMR / MDM, CNVs based on chromosome arm gain / loss log2 scores, and combinations of methylation and aneuploidy status.

[0261] Validation. For the validation study, DNA was extracted from prospectively collected esophageal brushing samples from 169 patients with NDBE (42), LGD (23), HGD (26), and EAC (41), along with 37 non-BE squamous epithelial samples (NE). Endoscopic brushings were obtained from the squamous epithelium and distal 5 cm of the esophagus from patients with NE and from the visible BE mucosa from patients with BE, using one brush for each 5 cm. Clinical and endoscopic data were abstracted, and histological data were confirmed by a gastrointestinal pathologist.

[0262] The same sequencing library preparation was used for both the methylation and aneuploidy assays, except that a target capture step was performed before sequencing for the methylation assay. For the target capture step, custom DNA capture probes were prepared by IDT for the 199 DMRs listed in Table 1. Targeted (DMR) and shallow whole-genome (CNA) sequencing on the enzyme-converted DNA library was performed using the Illumina NextSeq platform. The number of DMRs was reduced by removing those with low coverage and high variability in the control group (NDBE). Further variable selection was performed using the VSURF package in R, and the final model was generated using a random forest comparing NDBE versus HGD / EAC. CNA burden was estimated using trimmed median absolute deviation (tMAD) scores from the ichorCNA package. [Table 6]

[0263] Sequencing. Various nucleotide sequences referred to in this disclosure are provided below (see, e.g., Human Feb. 2009 (GRCh37 / hg19) Assembly).

[0264] MAX.chr15.4912(DMR110): CGTGGTTGCTGCGTGTGTCCCGCAGCCCCCTGCAGCCAGCGTCTCCTCGGCCCCGGCCCGACCCTGCCGTCCCTGCTCTGCTCCCTCAGCGACCGCAGACCCCTCACGCACATGCCCAGCCCTGCAGTCCTACTCTGCCACCCCAAGGGGCTTGGGCTTCAGTGGGGGCAGCGTGCGGGGCGTGGAGAGGGAGACGTGTAAAGCCCTGAGAGGCTCTGCAGCCACAGTCCCGCAGTCAGCGTTCCCTCTCGGCCAAGCCTCCTGGCTGCACATCGTTTCCCGGGCATCTCAGTGAGACCGGCCTCGGCTCACAGCGGCACGTTGTTTTTACACTTCGGGGACTCTGGTGGGACCGCGGGAAGCCAGGAGGCGGGCGCGCCCCGGGAGCCGATAGGAAAGTGCAACAGCGCCATCTAGTGGCCGCGCGGGGAGGCTGCGGAGCGCGCGCCGCGACCCGGGATCCACCGTCCGCAGGAAAACGTGTTCCCTCCGATGCGTGGTCACAGCGCGGCAGCGCGGCCTTGGGAGCTTCTTAAACAACAGGTTAGCTATCGCTGCCGCCTCAGTAAGAGAAGTCAAGGCCAAAACGCTTAAAGGATTTTGACCTATTCGGCTTGGTGGAGAGCGCTAAAAGGGGACGCTATAGCGGGATTCTCAGCTCTCCTGGGCTAGTGGGACCCCCGGTGCGCCGCCGCCGATGGGGTCTTAGGGCCTGTCTGTCTGGGCTAGAGCCGGCCCGGGAGCCTGTTTGCGGGGAGTGCG(SEQ ID NO: 1).

[0265] MAX.chr17.8070(DMR115): ATCTCCGATACTCCTCTCCTCAGCTTCCCGGGGGCAGTATGTCCCTTTGGGTCATGGCCTGTGGAGTGGCCATGACCCAAAGTTCAATTAAGGCTGGCGGGTCCCCAGAGGGGCCTGCTTCCATTCCCCCTTCCCCGCAGCCCTGCAGGCGTCAGCATGGGGCGAGGTTACCAGTTCCTGACCCAGCAGCTCGTCCTCAGGTGGTCCTGCTGCTCGGGAGGTCAGCTGCGTGGGGAGCCTGTCCACCTGGCTGATGGGGGTGATGGCCGCGTGGCTGGGATTCCTCCTCCAGCCCTGCCCTTCACCGCCATACAGTCTGGGCAGCCCCTTCCTGTCCCGTCCTCCTGCCTCGCCCTCTGCTGGGGTCTACGGGGCCCAAGTCTTGCTGGAACACAGGCAGGACAAGGGTTCTGGAGGTTCAGGCCGGAAGTGAAATAAAAGTCGAGAGGGGCCGGGCCCGGCACCCACCGCATCCTGCCTCCGGGCCTCCTCGGGCCCCTCACACGGTGGGGCAGCCCCTCCTCCAGTGGAGCGGAACGGTGGGCCCCGCTGCCCCTCCACCCCTGCGTGTGACACGGACGTGCTTGGAAGCTGGGTTCGTCCTTTCAGTTTTGTTTCCTGGTTGGAAACCCCTGAGACTCCTCCCCCTCCCCCCGCCCCTGCCCACCTGTGCTTTCCCAGGCTGGAAGGGGCAGCCCCTCTCAGCACGTGGGTGGTTAGGCTGTGACGTCCCCAGGCCTCCTGGGCCTGCACCGGGGTCCACCCAGGTGTGGGCCATGCCCAGGTGCGGG(SEQ ID NO: 2).

[0266] MAX.chr19.3439(DMR117): CGGTTTCGCTCCCTATCGGGGCCAGGGACGCCTCAGGCTGTCTCGGTCAGGACCTACAGCTCCGGTTGTTGTCCCAGGCTCTTCCGCGAGGTGCTCTCCTGTCTCCTGACCACCCCCATTCCTCCCCACTCCAGCTCCTCAGCGAGCGGCTGCGAAGGACGCGCACAACAACACTCCGCGCAGCGGGAAGGTACCGAACTCGCATCGCAGCCTGAAACTGCTCAACTAAGCTCCCGCCCCGCTGCTCTCTGGTCAATCTAAAGCGAAGACGAGCCTTAGGGCCAATCAGAAGCGACAGCGGTGGAGTCATGCCCGCCTGCTTGAGGCGCCCTGGCGTCTCATTGGCTATGCTTGAGAACGAATCCCAGGCTAAGCCACTTAGAAAGGAGCGGAGCCAGCCAATCAGCGGTGCAACCGCCGCGGGGGCCGGGCCAGAAGCCCCGCAGACAAGCACCTCGGGAGACTGGCGAGGGGCGAGCTCGCAGCTTGTTAGCCCCGAAGCCCTGCCCAGGGGGAAACCCTGCTGGAGGGAGTCTAACCCCCGGGCCAGTTAATGTTGGGGCAGCGCAATCGGCCGTTCCCTCTGCGGTATGGTTGGAGGTGGGAGTGCTGACACGTCCGCGCG(SEQ ID NO: 3).

[0267] MAX.chr4.4552(DMR124): GCCTGGGGCCACTGCTCCTGGGTCCTCAGGACTGCCTGGGGGAAGGTAGTGCATTGTGCAGCGCGCGGTCCAGAAGTGAAAAGGGAGGCGCGGAGATAAGCTGCCGGCGGAAGTTCCCTCTCCTGCCTGGGCCGACCCGGCGCTTTACTGCTTCTCACGAAGGTGCGCCGGCTGCTCCAGAAATCGCAGACTGCCTCCAGGAAGAACTTGCTGGAGTCACAGCAGCTTCTCAGCGACTTGACAGCAGTGATTCAGACTTCAACTTGGGCGGAGGGGCGGGGGAGGAGAAAGAGATTTCCAGAGAAAACGACTGAGCGGTAGGGAGGGGAAGAGAGACCGAGCCACGCGCCTCGAGAAGCAGTGCAGAGAGCGGAAGAGACAGAGGCTGCGAGACCTACCCACAGAGACCAAGAGAGACGCTCGGAGGGGAGACCGCCTTAGGCGCAGAGATTCAGAGGCAGACAGACAGGCAGACAGAAGGATACAGGAAAGGAATGTCGCCGAAAGGCAGGGACAAACCTGAGTCCCAGAAAAATAAGAGACAACTCCCACACACCAGGCTGTCCGCGGGCCGCCTTGTCACAGAAAGGCAGCTCCCCAGCCCCGCAGAGTCCCGACAGCTGCCCCCGCGAAGGTGGGGCGAGGGGCGGCTTTTCCG (SEQ ID NO: 4).

[0268] MAX.chr6.6743(DMR126):

[0269] MAX.chr8.3003(DMR129): CGCGCCTCCCGGAACCACGCGTCTCTGTGCACAGACATTCCTGGGGGCAGGCTCCTGTCCTTTAACACAGTCTCAAAGGAGTCTTCAAAAACAAAAAGTTTACAAGCACTGATAGGGAAAACAGAAGGATGATAACACCAAGAGGGACTTTCCCGGGAGAGGACGGCTTCCAGGGACTGAGAAAGGATGGGCAAGTGGGCGGGGCCCGGCGCCGTGCGGGCAGGGCTGGCGCCGGGAGTCCCCAGACTCCCCCGCAGTGGGAAGCACCTCTCCCATTCACGCCGGGCAGGACACCTGGCCGGGCGGGGGAGGCAGCGCAAGGGCCGGCCGGGGAGTACGGGACTCGAGCCGGGGACCTGAGGCAGGAGCCAAGCATCGTCGCAGGGCAACCAGCAGAACGGAGAGGGAGGCGCGGGGGCGAAGGCTGGCGGGAGCCGCGCTGAGGGCAGGAGCCCGGAGCCCCCTAGGGCAGCGCCGATCCGCCCGCCCCGTCCCGCCGAGCTGGGCCTCCGTCTGTGGCCTGCGCAGCCAGGGTCGCCAAGCCG (SEQ ID NO: 6).

Claims

1. 1. A method for characterizing a biological sample, said method comprising: The method comprises determining a methylation profile in at least one variably methylated region (DMR) of a DNA sample obtained from a subject having or suspected of having esophageal cancer or precancer by treating the sample with a reagent that modifies DNA in a methylation-specific manner.

2. 10. The method of claim 1, wherein the methylation profile in the at least one DMR indicates that the subject has or is suspected of having high-grade dysplastic Barrett's esophagus or esophageal adenocarcinoma (EAC).

3. The at least one DMR is ACVR1L, ADAMTS8, ADAP2, ADRBK2, AKR1B1, ANK1, ANKRD13B, ANXA6, ARNT2, B4GALNT2, BACH2, BCL11A, BSCL2, C12orf53, C14orf82, C17orf107, C18orf1, C1orf95, C5orf42, CACNA1C, CAMK1D, CAMTA1, CBX6, CCDC85A, CCKBR, CD38, CDKN2A, CH25H, CHST1, CHST15, CNTLN, CRHR1, CRTC1, CXCR4, CYP1B1, DCTN2, DDO1, DMKN, DSE, DYNC1I1, EML6, ENOX1, EPHA4, ESRRG, FAM176A, FAM78B, FBXO10, FERMT2, FHOD3, FLJ45079, FMNL1, FOXP2, FRMD4B, FZD8, GALNTL1, GALNTL4, GAS1, GBGT1, GLIPR2, GNAI1, GNAL, GPR37, GRASP, GRID1, GRM8, GSC, HAR1A, HCN2, HEY2, HIST1H2BE, HS3ST3B1, HTR7, IGFBP2, INHA, INSRR, IRX3, ISM2, KCNG3, KCNK4, KCNS2, KCTD15, KIAA1522, KIAA1614, KIF26A, KL, KLF15, KLHL10, KRT77, KSR2, LBH, LMX1A, LMX1B, LOC100526820, LONRF2, LRFN2, LRRN1, MAF, MAFB, MARK1, ADAMTSL4-AS1, PGBBD5, HSP A12A, SFTPD, LOC107984507, LOC100128253, CISTR, SLC16A7, CTXN D1, MAX.chr15.4912, ZNF423, RBFOX1, LOC105376772, GSE1, MAX.chr17.8070, ZNF709, MAX.chr19.3439, GREB1, CYP27C1, AC068134.6, LOC388780, STOX2, PPARG C1A, MAX.chr4.4552, PDZD2, MAX.chr6.6743, LPAL A2, CDK14, MAX.chr8.3003, MCOLN2, MEGF11, MFSD1, MRC2, MSX1, NAT8L, NAV1, NBEA, ncRNA00092, NECAB2,NEURL, NGEF, NHLH2, NKD1, NLGN1, NOG, NPPC, NR3C1, NRXN2, NTN1, OBSCN, OCA2, OSBPL1A, OXR1, P2RY1, PDE2A, PDE 9A, PDGFRA, PID1, PLCG2, PLCL1, PNPLA3, POU3F1, POU3F2, PRKACB, PRKAR2B, PRKG1, PRR18, PRR5L, PTHLH, PYGL, RA RG, RHBDL3, RIMS2, ROR2, RPRML, SDK2, SOX9, SYNGR1, TAC4, TNFRSF19, TRANK1, TSPAN33, TSPAN4, TSPAN5, UBE2E2, UCHL1, UNC5A, VASH2, VIM, WNT6, ZBTB10, ZNF680, ZNF738, ZNF808, and ZNF845.

4. The at least one DMR is ACVR1L, ADAMTS8, ADAP2, ADRBK2, AKR1B1, ANK1, ANKRD13B, ANXA6, ARNT2, B4GALNT2, BACH2, BCL11A, BSCL2, C14orf82, C18orf1, C1orf95, C5orf42, CAMK1D, CAMTA1, CCDC85A, CD38, CDKN2A, CHST1, CHST15, CRHR1, CYP1B1, DIO1, DSE, DYNC1I1, EML6, ENOX1, EPHA4, FAM176A, FBXO10, FERMT2, FHOD3, FLJ45079, FMNL1, FRMD4B, FZD8, GALNTL1, GALNTL4, GAS1, GBGT1, GLIPR2, GNAI1, GNAL, GRASP, GRID1, GRM8, GSC, HAR1A, HCN2, HEY2, HIST1H2BE, HS3ST3B1, HTR7, IGFBP2, INHA, IRX3, ISM2, KCNG3, KCNK4, KCNS2, KCTD15, KIAA1614, KL, KLHL10, KSR2, LBH, LMX1A, LMX1B, LOC100526820, LONRF2, LRFN2, MAFB, MARK1, PGBBD5, HSP12A, SFTPD, LOC107984507, LOC100128253, CISTR, SLC16A7, CTXN1, ZNF423, LOC105376772, GSE1, MAX.chr19.3439, GREB1, CYP27C1, AC068134.6, LOC388780, STOX2, PPARG1A, PDZD2, MAX.chr6.6743, LPAL2, CDK14, MCOLN2, MEGF11, MRC2, MSX1, NAT8L, NAV1, NBEA, ncRNA00092, NEURL, NGFF, NHLH2, NKD1, NLGN1, NOG, Nppc, NR3C1, NRXN2, OBSCN, OCA2, OSBPL1A, OXR1, P2RY1, PDE2A, PDGFRA, PID1, PLCG2, PLCL1, PNPLA3, POU3F1, POU3F2, PRKACB, PRKAR2B, PRKG1, PRR18, PRR5L, PTHLH, PYGL, RHBDL3, RIMS2, ROR2, RPRML, SDK2, SOX9, SYNGR1, TAC4, TRAN1, TSPAN4, UBE2E2,The method according to claim 1 or 2, wherein the gene is derived from a gene selected from UCHL1, UNC5A, VASH2, ZBTB10, ZNF680, ZNF709, ZNF738, ZNF808, and ZNF845.

5. 3. The method of claim 1 or 2, wherein the at least one DMR is derived from a gene selected from BACH2, C5orf42, FHOD3, HIST1H2BE, IRX3, KIAA1614, LONRF2, MAFB, PDGFRA, PID1, POU3F1, PRR5L, RHBDL3, and SDK2.

6. 3. The method of claim 1 or 2, wherein the at least one DMR is derived from a gene selected from KL, PGBD5, ROR2, and LMX1B.

7. 7. The method of any one of claims 1-6, wherein the at least one DMR is associated with an area under the ROC curve (AUC) of 0.5 or greater, and the ROC curve distinguishes between subjects having or suspected of having high-grade dysplastic Barrett's esophagus or esophageal adenocarcinoma (EAC) and control DNA samples.

8. 7. The method of any one of claims 1-6, wherein the at least one DMR is associated with an area under the ROC curve (AUC) of 0.75 or greater, and the ROC curve distinguishes between subjects with or suspected of having high-grade dysplastic Barrett's esophagus or esophageal adenocarcinoma (EAC) and control DNA samples.

9. 7. The method of any one of claims 1 to 6, wherein the at least one DMR is associated with a methylation fold change (FC) ratio of 3.0 or greater, and the FC ratio distinguishes between subjects with or suspected of having high-grade dysplastic Barrett's esophagus or esophageal adenocarcinoma (EAC) and control DNA samples.

10. 7. The method of any one of claims 1 to 6, wherein the at least one DMR is associated with a methylation fold change (FC) ratio of 5.0 or greater, and the FC ratio distinguishes between subjects with or suspected of having high-grade dysplastic Barrett's esophagus or esophageal adenocarcinoma (EAC) and control DNA samples.

11. 11. The method of any one of claims 1 to 10, wherein the method further comprises determining copy number variations (CNV) for the DNA sample from the subject, wherein the CNV distinguishes between subjects having or suspected of having high-grade dysplastic Barrett's esophagus or esophageal adenocarcinoma (EAC) and control DNA samples.

12. 11. The method of any one of claims 1 to 10, wherein the method further comprises determining an aneuploidy score (AS) for the DNA sample from the subject, wherein the AS distinguishes between subjects having or suspected of having high-grade dysplastic Barrett's esophagus or esophageal adenocarcinoma (EAC) and control DNA samples.

13. 13. The method of any one of claims 1 to 12, wherein the at least one DMR comprises an increased percentage of methylation compared to a control DNA sample.

14. 13. The method of any one of claims 1 to 12, wherein the at least one DMR comprises an increased hypermethylation rate compared to a control DNA sample.

15. 13. The method of any one of claims 1 to 12, wherein the at least one DMR comprises an increased AS compared to the control DNA sample.

16. The method of any one of claims 7 to 15, wherein the control DNA sample is derived from a subject who does not have esophageal cancer or precancer.

17. The method of any one of claims 7 to 15, wherein the control DNA sample is derived from a subject with non-dysplastic Barrett's esophagus (NDBE).

18. 18. The method of any one of claims 7 to 17, wherein the control DNA sample is selected from a tissue sample, a blood sample, a plasma sample, a serum sample, a whole blood sample, a buffy coat sample, a secretion sample, an organ secretion sample, a cerebrospinal fluid (CSF) sample, a saliva sample, a urine sample, and a stool sample.

19. 19. The method of claim 18, wherein the tissue sample is an esophageal tissue sample or a gastric cardia sample.

20. 20. The method of any one of claims 1 to 19, wherein the biological sample is selected from a tissue sample, a blood sample, a plasma sample, a serum sample, a whole blood sample, a buffy coat sample, a secretion sample, an organ secretion sample, a cerebrospinal fluid (CSF) sample, a saliva sample, a urine sample, and a stool sample.

21. 21. The method of claim 20, wherein the tissue sample is an esophageal tissue sample.

22. 21. The method of claim 20, wherein the tissue sample is an endoscopic esophageal brushing sample.

23. The method of any one of claims 1 to 13, wherein the subject is a human.

24. The method of any one of claims 1 to 13, wherein the biological sample is obtained from the subject, and the method further comprises extracting the DNA sample from the biological sample.

25. The method of any one of claims 1 to 24, wherein the biological sample is collected with a collection device.

26. 26. The method of any one of claims 1 to 25, wherein the reagent that modifies DNA in a methylation-specific manner is a borane reducing agent.

27. 26. The method of any one of claims 1 to 25, wherein the reagent that modifies DNA in a methylation-specific manner comprises one or more of a methylation-sensitive restriction enzyme, a methylation-dependent restriction enzyme, and a bisulfite reagent.

28. 28. The method of any one of claims 1 to 27, wherein determining the methylation profile of at least one DMR comprises amplifying at least a portion of the DMR using a set of primers.

29. 29. The method of any one of claims 1 to 28, wherein determining the methylation profile of at least one DMR comprises performing at least one of methylation-specific PCR, quantitative methylation-specific PCR, methylation-specific DNA restriction enzyme analysis, quantitative bisulfite pyrosequencing, flap endonuclease assay, PCR flap assay, and bisulfite genomic sequencing PCR.

30. 30. The method of any one of claims 1 to 29, wherein determining the methylation profile of at least one DMR comprises determining the presence or absence of methylation at one or more CpG sites.

31. 31. The method of claim 30, wherein the one or more CpG sites are present in a coding region, a non-coding region, and / or a regulatory region of a gene.

32. 32. The method of any one of claims 1 to 31, wherein determining the methylation profile of at least one DMR comprises determining a methylation frequency.

33. 32. The method of any one of claims 1 to 31, wherein determining the methylation profile of at least one DMR comprises determining a methylation pattern.

34. 34. The method of any one of claims 1 to 33, wherein the methylation profile in the at least one DMR and the CNV and / or AS are determined using the same DNA sample obtained from the subject.

35. 34. The method of any one of claims 1 to 33, wherein the methylation profile in the at least one DMR and the CNV and / or AS are determined using a single DNA sample obtained from the subject.

36. 36. The method of claim 34 or 35, wherein the sample has been treated with a reagent that modifies DNA in a methylation-specific manner.