Improved methods for detecting esophageal neoplasias and / or metaplasias in the esophagus
The method of enzymatic or chemical conversion and PCR amplification of vimentin and CCNA1 sequences addresses the invasiveness and low compliance of current BE and EAC detection methods, improving sensitivity and specificity for early detection.
Patent Information
- Application Number
- PCT/US2025/032668
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-07
- Filing Date
- 2025-06-06
- Publication Date
- 2025-12-11
AI Technical Summary
Current methods for detecting esophageal neoplasias such as Barrett's Esophagus and Esophageal Adenocarcinoma are invasive and have low compliance due to their nature, necessitating improved non-endoscopic screening methods with higher sensitivity and specificity for early detection.
A method involving enzymatic or chemical conversion of DNA sequences in tissue samples to produce modified samples, followed by polymerase chain reaction amplification and analysis using machine learning models to determine methylation status, specifically targeting vimentin and CCNA1 nucleic acid sequences for accurate detection of BE and EAC.
Enhances the sensitivity and specificity of detecting BE and EAC, enabling robust large-scale screening with high success rates, reducing the invasiveness of current endoscopic procedures.
Smart Images

Figure US2025032668_11122025_PF_FP_ABST
Abstract
Description
IMPROVED METHODS FOR DETECTING ESOPHAGEAL NEOPLASIAS AND / OR METAPLASIAS IN THE ESOPHAGUSCROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of and priority to U.S. Provisional Application No. 63 / 657,635, filed June 7, 2024, and U.S. Provisional Application No. 63 / 657,662, filed June 7, 2024, the entire contents of each of which are incorporated herein by reference.SEQUENCE LISTING
[0002] This application contains a Sequence Listing which has been submitted electronically in XML format and is hereby incorporated by reference in its entirety. The XML copy, created on June 6, 2025, is titled 182338-011902_PCT.xml and is 5,031 bytes in size.TECHNICAL FIELD
[0003] Example embodiments relate generally to identifying disease within a sample of tissue. More particularly, the example embodiments relate to analyzing samples gathered from patients to detect the presence of diseases such as Barrett’s Esophagus (BE) or Esophageal Adenocarcinoma (EAC).BACKGROUND
[0004] Esophageal Adenocarcinoma (EAC) is the second deadliest cancer in the United States, with a five-year survival rate of less than 20%. Barrett's Esophagus (BE) is the precursor to EAC, with well-defined stages of disease progression. BE advances from Benign metaplasia, i.e., non-dysplastic BE (NDBE) to low-grade dysplasia (LGD), to high-grade dysplasia (HGD), and finally to cancer. EAC has poor prognosis (<20% 5 year survival rate) even when detected at early stages such as stage I and II. In contrast, when detected at precancerous stage (BE) (precancer) can be successfully treated (80-90% success rate) using endoscopy modalities such as ablation reliably halting malignant progression. Therefore, screening for BE has been recommended by many societal guidelines. However, compliance for BE screening using upper endoscopy (UE) has been poor (<10%) due to its invasive nature and logistical challenges. In recent update to American College of Gastroenterology (ACG) and American Gastroenterological Association (AGA) guidelines (2022), non-endoscopic screening methods for Barrett’s esophagus have been recommended. As with other cancers, the rate of deaths caused by EAC may be decreased by improved methods for diagnosis. EsoCheck is an FDAapproved, 510k cleared, balloon device that can collect targeted, circumferential sample from distal esophagus and protect sample in attached capsule during sample retrieval. The sample collection using EsoCheck can be performed in non-sedated manner in office settings in less than 5 minutes, making it highly suitable device for BE screening at large scale in high-risk population. A molecular biomarker epigenetic test using VIM and CCNA1 genes has shown high sensitivity and specificity for detecting BE in non-endoscopically collected samples, for example as discussed in the paper titled “Identifying DNA Methylation Biomarkers for Non- Endoscopic Detecting of Barrett’s Esophagus” in the National Library of Medicine and authored by Moinova, et al, and the paper titled “Clinical Utility Study of EsoGuard on Samples Collected with EsoCheck as a Triage Test for Endoscopy to Identify Barrett’s Esophagus - Interim Data of the First 15 Subjects” in medRxiv and authored by Lister, et al. However, there remains a need for improved methods for more robust biomarker testing that can be performed at large scale for detection of BE and EAC.SUMMARY
[0005] Aspects of the present disclosure relate to a method for detecting methylation status of at least one sequence in a tissue sample of a subject. The method includes providing the tissue sample obtained from the subject, treating the tissue sample to convert DNA sequences within the tissue sample with at least one of enzymatic conversion and chemical conversion to produce a modified sample having converted DNA sequences, and replicating the converted DNA to produce at least three copies of each sequence of the converted DNA sequences. The method further includes amplifying the converted DNA sequences with a polymerase chain reaction such that a plurality of converted amplicons is produced and analyzing the plurality of converted amplicons to determine the methylation status of the sample.
[0006] In some embodiments, treating the sample includes treating the sample with enzymatic bisulfite conversion. In other embodiments, treating the sample includes treating the sample with chemical bisulfite conversion through applying sodium bisulfite to the sample. In further embodiments, analyzing the plurality of converted amplicons includes using a machine learning model to determine whether the sample is methylated. In further embodiments, analyzing the plurality of converted amplicons includes, for each replicate, measuring the number of methylated cytosines within each of the plurality of converted amplicons, wherein a read is classified as a methylated read when at least 70% of the cytosines of each of the plurality of converted amplicon are methylated and calculating percent of total reads are methylated reads from the amplifying and measuring steps, wherein the replicate is considereda positive sample when at least 0.5% of the total reads are methylated reads. In some embodiments, the plurality of converted amplicons are vimentin nucleic acid amplicons, a read is classified as a methylated read when at least 80% of the cytosines of the individual amplicon are methylated, and the sample obtained from the subject is classified as a methylated sample when at least 1.0% of the total reads are methylated reads. In some embodiments, the plurality of converted amplicons is a plurality of CCNA nucleic acid amplicons. In further embodiments, if a majority of the plurality of converted amplicons are identified as methylated reads, the sample is identified as a methylated sample.
[0007] Aspects of the present disclosure is directed to a method for detecting methylation status of at least one sequence in a human subject to determine whether the human subject is healthy or diseased. The method includes providing a tissue sample obtained from the human subject, treating the tissue sample to convert DNA within the tissue sample with at least one of enzymatic conversion and chemical conversion to produce a modified sample containing a plurality of converted vimentin nucleic acid sequences and a plurality of converted CCNA1 nucleic acid sequences, replicating the plurality of converted vimentin nucleic acid sequences and the plurality of converted CCNA1 nucleic acid sequences, such that each of the plurality of converted CCNA1 nucleic acid sequences has at least three replicates and such that each of the plurality of converted vimentin nucleic acid sequences has at least three replicates, amplifying each of the at least three replicates of each of the plurality of converted vimentin nucleic acid sequences such that a plurality of converted vimentin amplicons is produced, and amplifying each of the at least three replicates of each of the plurality of converted CCNA1 nucleic acid sequences such that a plurality of converted CCNA1 amplicons is produced. The method further includes analyzing the plurality of converted vimentin amplicons and the plurality of converted CCNA1 amplicons to determine the methylation status of the sample.
[0008] In some embodiments, amplifying each of the at least three copies of the plurality of converted vimentin nucleic acid sequences and the plurality of converted CCNA1 amplicons includes undergoing polymerase chain reaction amplification. In further embodiments, treating the sample includes treating the sample with enzymatic bisulfite conversion. In other embodiments, treating the sample includes treating the sample with chemical bisulfite conversion through applying bisulfite to the sample. In further embodiments, prior to the amplifying step, the method further includes applying a solution for PCR preparation including a magnesium concentration of between approximately 2 mM and approximately 4 mM.
[0009] Aspects of the present disclosure is directed to a method for detecting methylationstatus of CpG dinucleotide sites within a sequence in a human subject to determine whether the human subject is healthy or diseased. The method includes providing a tissue sample obtained from the human subject, treating the tissue sample to convert DNA within the tissue sample with at least one of enzymatic conversion and chemical conversion to produce a modified sample containing a plurality of converted vimentin nucleic acid sequences and a plurality of converted CCNA1 nucleic acid sequences, replicating the plurality of converted vimentin nucleic acid sequences and the plurality of converted CCNA1 nucleic acid sequences, such that each of the plurality of converted CCNA1 nucleic acid sequences has at least three copies and such that each of the plurality of converted vimentin nucleic acid sequences has at least three copies, and amplifying each of the at least three copies of each of the plurality of converted vimentin nucleic acid sequences and each of the at least three copies of each of the plurality of converted CCNA1 nucleic acid sequences. The method further includes identifying whether each of the at least three copies of each of the plurality of converted CCNA 1 amplicons is considered a methylated read, identifying whether each of the at least three copies of each of the plurality of converted vimentin amplicons is considered a methylated read, and identifying whether each of the at least three copies of each of the plurality of converted vimentin amplicons is considered a methylated read.
[0010] In some embodiments, the method further includes, prior to the amplifying steps, applying a solution to the sample for PCR preparation including a magnesium concentration of between approximately 2 mM and approximately 4 mM. In further embodiments, treating the sample includes treating the sample with enzymatic bisulfite conversion. In other embodiments, treating the sample includes treating the sample with chemical bisulfite conversion.BRIEF DESCRIPTION OF THE DRAWINGS
[0011] FIG. 1 is a flowchart illustrating a method for detecting methylation status of at least one sequence in a human subject, in accordance with the present disclosure.
[0012] FIG. 2A is a table providing area under the curve, sensitivity and specific values observed for VIM and CCNA as biomarkers for BE and EAC using a singleplex assay, in accordance with Example 1. FIG. 2B is a table illustrating the clinical specificity and sensitivity for detecting BE and EAC using cytology brushing samples as well as non-endoscopically collected EsoCheck balloon samples, in accordance with Example 1.
[0013] FIG. 3 A is a table providing the analytical sensitivity, specificity, and accuracy values of the singleplex assay v 1.0.
[0014] FIG. 3B is a table comparing the reference data set with the results obtained from the assay, in accordance with Example 1.
[0015] FIGs. 4A-4B are tables providing the stability of epigenetic signal in VIM and CCNA1 genes measured over the course of 14 days, in accordance with Example 1.
[0016] FIG. 5 is a table illustrating the analytical sensitivity, analytical specificity, accuracy, intra run precision and inter run precision of the improved version of singleplex assay (vl.2), in accordance with Example 1.
[0017] FIG. 6A is a table illustrating limit of detection of the singleplex assay (v 1.2) showing sample number, in dilution series of contrived specimens consisting of mixture of positive cell line with negative cell line, in accordance with Example 1.
[0018] FIG. 6B is an assay linearity graph illustrating the R2of the singleplex assay, in accordance with Example 1.
[0019] FIG. 7A is a table illustrating the performance of a single-plex assay versus the multiplex assay for VIM and CCNA1 genes, in accordance with Example 1 .
[0020] FIG. 7B is an assay linearity graph of values for multiplex testing of VIM and CCNA1 genes, in accordance with Example 1.
[0021] FIG. 8 illustrates the assay quantity not sufficient (QNS) rate, in accordance with Example 1.
[0022] FIG. 9 is a graph showing the linearity of methylation percentage obtained for vimentin and CCNA1 genes, in accordance with Example 2.
[0023] FIG. 10 are graphs showing the sample stability in transport media at room temperature. Panel A shows the methylation percentage for vimentin (VIM). Panel B shows the methylation percentage for CCNA1.
[0024] FIG. 11 are graphs showing the EsoGuard® assay performance in the presence of interfering substances. Panel A shows the methylation percentage for vimentin (VIM). Panel B shows the methylation percentage for CCNA1. Con, H, B,UB and L represent control, heme, Conjugated bilirubin, unconjugated bilirubin and triglyceride rich lipoproteins.
[0025] FIG. 12 is Sanger sequencing data showing all 3 ICpG sites are methylated in SK-GT- 4 cell line and not methylated in Hl 975 cell line.
[0026] FIG. 13 are graphs showing EsoGuard® bioinformatic pipeline (V2.0) in comparisonto previous version of the pipeline. Panel A shows the methylation for vimentin (VIM). Panel B shows the methylation for CCNA1.
[0027] FIG. 14 are graphs showing accuracy of NextSeq 1000 platform against MISeq platform. Panel A shows the methylation for vimentin (VIM). Panel B shows the methylation for CCNA1.DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] Various exemplary embodiments will be described more fully hereinafter with reference to the accompanying drawings, in which some example embodiments are shown. The present inventive concept may, however, be embodied in many different forms and should not be construed as limited to the example embodiments set forth herein. Rather, these example embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the present inventive concept to those skilled in the art.
[0029] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this inventive concept belongs. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and will not be interpreted in an idealized or overly formal sense unless expressly so defined herein.
[0030] Throughout this specification, the word “comprise” or variations such as “comprises” or “comprising” will be understood to imply the inclusion of a stated integer or groups of integers but not the exclusion of any other integer or group of integers.
[0031] The term “including” is used herein to mean, and is used interchangeably with, the phrase “including but not limited to.”
[0032] The articles “a” and “an” are used herein to refer to one or to more than one (i.e., to at least one) of the grammatical object of the article. By way of example, “an element” means one element or more than one element.
[0033] The terms “adenoma” is used herein to describe any precancerous neoplasia or benign tumor of epithelial tissue, for example, a precancerous neoplasia of the gastrointestinal tract, pancreas, and / or the bladder.
[0034] The term “esophagus” is intended to encompass the upper portion of the digestive system spanning from the back of the oral cavity, passing downwards through the rear part ofthe mediastinum, through the diaphragm and into the stomach. Chronic Gastroesophageal reflux disease (GERD) represents reflux condition in which stomach’s acidic content move up to esophagus, which increases the risk for developing esophageal cancer.
[0035] The term “esophageal cancer” is used herein to refer to any cancerous neoplasia of the esophagus. “Barrett's esophagus” as used herein refers to an abnormal change (metaplasia) in the cells of the lower portion of the esophagus. Barrett's is characterized the finding of intestinal metaplasia in the esophagus. “Dysplasia” means presence of cells of abnormal type in a tissue which may or may not develop cancer. “Non-dysplastic Barrett’s Esophagus” (NDBE) refers to Barrett’s esophagus where no dysplastic stages are found. “High Grade Dysplasia” (HGD) refers to more advanced stages of Esophagus dysplasia for developing cancerous stage than “LGD” or Low-Grade Dysplasia. The term “esophageal cancer” or “EAC” is used herein to refer to any cancerous neoplasia of the esophagus.
[0036] The term “sample collection” or “brushing” of the esophagus, as referred to herein, may be obtained using any of the means known in the art. In some embodiments, a brushing is obtained by contacting the esophagus with a brush, a cytology brush, a sponge, a balloon, or with any other device or substance that contacts the esophagus and obtains an esophageal sample.
[0037] “Cells,” “host cells” or “recombinant host cells” are terms used interchangeably herein. It is understood that such terms refer not only to the particular subject cell but to the progeny or potential progeny of such a cell. Because certain modifications may occur in succeeding generations due to either mutation or environmental influences, such progeny may not, in fact, be identical to the parent cell, but are still included within the scope of the term as used herein.
[0038] The terms “compound”, “test compound,” “agent”, and “molecule” are used herein interchangeably and are meant to include, but are not limited to, peptides, nucleic acids, carbohydrates, small organic molecules, natural product extract libraries, and any other molecules (including, but not limited to, chemicals, metals, and organometallic compounds).
[0039] The term “converted DNA” herein refers to DNA that has been treated such that unmethylated C bases in DNA are converted to a different nucleotide base and methylated C bases in DNA remain as C bases. For example, converted DNA may refer to DNA that has been treated with enzymatic conversion.
[0040] The term “compound-converted DNA” herein refers to DNA that has been treated or reacted with an enzymatic / chemical compound that converts unmethylated C bases in DNA toa different nucleotide base. For example, one such compound is sodium bisulfite, which converts unmethylated C to U. If DNA that contains conversion-sensitive cytosine is treated with sodium bisulfite, the compound-converted DNA will contain U in place of C. If the DNA which is treated with sodium bisulfite contains only methylcytosine, the compound-converted DNA will not contain uracil in place of the methylcytosine and will instead remain as methylcytosine. Further, in another example, another such compound is Tet methylcytosine dioxygenase 2 (TET2). TET2 causes enzymatic oxidation of 5-methylcystosine (5mC) and 5- hydroxymethylcytosine (5hmC) int 5 -carboxy cytosine (5caC) to provide compound-converted DNA.
[0041] The term “modified sample” herein refers to genetic material that has been treated or reacted with a chemical compound, enzyme, or by any form of mechanical / enzymatic shearing or enriching by any forms of hybridization or pull-down steps that converts sample material to allow analysis of any one of the biomarkers. For example, one such compound is sodium bisulfite, which converts unmethylated C to U. If DNA that contains conversion-sensitive cytosine is treated with sodium bisulfite, the compound-converted DNA will contain U in place of C. If the DNA which is treated with sodium bisulfite contains only methylcytosine, the compound-converted DNA will not contain uracil in place of the methylcytosine. Further, in another example for enzymatic conversion, another such compound is Tet methylcytosine dioxygenase 2 (TET2). TET2 causes enzymatic oxidation of 5-methylcystosine (5mC) and 5- hydroxymethylcytosine (5hmC) int 5 -carboxy cytosine (5caC) to produce a modified sample.
[0042] The term “de-methylating agent” as used herein refers agents that restore activity and / or gene expression of target genes silenced by methylation upon treatment with the agent. Examples of such agents include without limitation de-sulphonating agents for chemical modification or 5-azacytidine and 5-aza-2' -deoxy cytidine.
[0043] The term “Biomarker” is meant to include, but is not limited to, peptides, nucleic acids, carbohydrates, small organic molecules, natural product extract libraries, and any other molecules (including, but not limited to DNA, RNA, Protein, noncoding RNAs like miRNA). The presence of a biomarker beyond a limit is indicative of onset of a disease or advanced disease stage.
[0044] The term “detection” is used herein to refer to any process of observing a marker, or a change in a marker (such as for example the change in the methylation state of the marker), in a biological sample, whether or not the marker or the change in the marker is actuallydetected. In other words, the act of probing a sample for a marker or a change in the marker, is a “detection” even if the marker is determined to be not present or below the level of sensitivity. Detection may be a quantitative, semi -quantitative or non-quantitative observation.
[0045] The term “neoplasia” as used herein refers to an abnormal growth of tissue. As used herein, the term “neoplasia” may be used to refer to cancerous and non-cancerous tumors, as well as to Barrett's esophagus (which may also be referred to herein as a metaplasia) and Barrett's esophagus with dysplasia. In some embodiments, the Barrett's esophagus with dysplasia is Barrett's esophagus with high grade dysplasia. In some embodiments, the Barrett's esophagus with dysplasia is Barrett's esophagus with low grade dysplasia. In some embodiments, the neoplasia is a cancer (e.g., esophageal adenocarcinoma).
[0046] “Gastrointestinal neoplasia” refers to neoplasia of the upper and lower gastrointestinal tract. As commonly understood in the art, the upper gastrointestinal tract includes the esophagus, stomach, and duodenum; the lower gastrointestinal tract includes the remainder of the small intestine and all of the large intestine.
[0047] The term “methylation-specific PCR” (“MSP”) herein refers to a polymerase chain reaction in which amplification of the modified sample is performed. Two sets of primers are designed for use in MSP. Each set of primers comprises a forward primer and a reverse primer. One set of primers, called methylation-specific primers (see below), will amplify the compound-converted template sequence if C bases in CpG dinucleotides within the DNA are methylated. Another set of primers, called unmethylation-specific primers or primers for unmethylated sequences and the like (see below), will amplify the compound-converted template sequences if C bases in CpG dinucleotides within the DNA are not methylated.
[0048] “SqBE 18” used herein refers to a nucleotide sequence and may be used with the term “CCNA1” interchangeably throughout the present disclosure.
[0049] “Operably linked” when describing the relationship between two DNA regions simply means that they are functionally related to each other. For example, a promoter or other transcriptional regulatory sequence is operably linked to a coding sequence if it controls the transcription of the coding sequence.
[0050] The terms “healthy”, “normal,” and “non-neoplastic” are used interchangeably herein to refer to a subject or particular cell or tissue that is devoid (at least to the limit of detection) of a disease condition, such as a neoplasia.
[0051] As used herein, the term “nucleic acid” refers to polynucleotides such asdeoxyribonucleic acid (DNA), and, where appropriate, ribonucleic acid (RNA). The term should also be understood to include, as equivalents, analogs of either RNA or DNA made from nucleotide analogs, and, as applicable to the embodiment being described, single -stranded (such as sense or antisense) and double -stranded polynucleotides.
[0052] The term, “about” a number, as used herein, refers to range including the number and ranging from 10% below that number to 10% above that number. “About” a range refers to 10% below the lower limit of the range, spanning to 10% above the upper limit of the range.
[0053] As applied to polypeptides, the term “substantial sequence identity” means that two peptide sequences, when optimally aligned such as by the programs GAP or BESTFIT using default gap, share at least 90 percent sequence identity, in some embodiments, at least 95 percent sequence identity, or at least 99 percent sequence identity or more. In some embodiments, residue positions which are not identical differ by conservative amino acid substitutions. For example, the substitution of amino acids having similar chemical properties such as charge or polarity is not likely to affect the properties of a protein. Examples include glutamine for asparagine or glutamic acid for aspartic acid.
[0054] An “informative loci” as used herein, refers to any of the nucleic acid sequences disclosed herein that may have altered (e.g., increased) methylation in a sample (e.g., an esophageal tissue sample) from a subject having Barrett's esophagus and / or an esophageal neoplasia as compared to the methylation patterns of the corresponding nucleic acid sequence in a sample from a healthy control subject. An example of an informative loci is vimentin.
[0055] The term “or” is used herein to mean, and is used interchangeably with, the term “and / or”, unless context clearly indicates otherwise.
[0056] A “sample” includes any material that is obtained or prepared for detection of a molecular marker or a change in a molecular marker such as for example the methylation state, or any material that is contacted with a detection reagent or detection device for the purpose of detecting a molecular marker or a change in the molecular marker.
[0057] As used herein, “obtaining a sample” includes directly retrieving a sample from a subject to be assayed, or directly retrieving a sample from a subject to be stored and assayed at a later time. Alternatively, a sample may be obtained via a second party. That is, a sample may be obtained via, e.g., shipment, from another individual who has retrieved the sample, or otherwise obtained the sample.
[0058] A “subject” is any organism of interest, generally a mammalian subject, such as a mouse, and in particular embodiments, a human subject.
[0059] Aspects of the disclosure is based at least in part on the recognition that differential methylation of particular genomic loci (e.g., vimentin and / or CCNA1) may be indicative of a neoplasia or metaplasia of the upper gastrointestinal tract, e.g., esophagus. The present findings demonstrate that methylation at these genomic loci may be a useful biomarker of neoplasia in the upper gastrointestinal tract. In some embodiments, the vimentin methylation is detected in a manner consistent with that described in Li et al. (Li M, et al. (2009) Sensitive digital quantification of DNA methylation in clinical samples. Nat Biotechnol 27(9): 858-863).
[0060] In other embodiments, the disclosure also provides the methylated forms of the nucleotide sequences, wherein the cytosine bases of the CpG islands present in the sequences are methylated. In other words, the nucleotide sequences or fragments or complements thereof may be either in the methylated status (e.g., as seen in neoplasias) or in the unmethylated status (e.g., as seen in normal cells). In further embodiments, the nucleotide sequences of the disclosure can be isolated, recombinant, and / or fused with a heterologous nucleotide sequence, or in a DNA library.
[0061] In certain embodiments, the present disclosure provides bisulfite-converted nucleotide sequences, for example, bisulfite-converted sequences.
[0062] A fragment / portion of any of the nucleotide sequences disclosed herein may be of any length, so long as the methylation status of that nucleotide sequence may be determined. In some embodiments, the nucleotide sequence is at least 10, 15, 25, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 250, 300, 350, 400, 450, 500, 600, 700, 800, 900, 1000, 1200, 1400, 1500, 1700, or 2000 nucleotides in length, in some embodiments, the nucleotide sequence is at least 10-2000, 10-1000, 10-500, 10-200, 10-150, 10-100, 50-2000, 50-1000, 50-500, 50-200, 50-150, 50-100, 80-2000, 80-1000, 80-500, 80-150, 80-100, 100- 2000, 100-1000, 100-500, 100-200, or 100-150 nucleotides in length.
[0063] In general, neoplasias may develop through one of at least three different pathways, termed chromosomal instability, microsatellite instability, and the CpG island methylator phenotype (CIMP). Although there is some overlap, these pathways tend to present somewhat different biological behavior.
[0064] Aspects of the disclosure is based, at least in part, on the recognition that certain target genes may be silenced or inactivated by the methylation of CpG islands in the 5 ' flanking orpromoter regions of the target gene. CpG islands are clusters of cytosine-guanosine residues in a DNA sequence, which are prominently represented in the 5 -flanking region or promoter region of about half the genes in our genome. In particular, this application is based at least in part on the recognition that differential methylation of particular genomic loci may be indicative of neoplasia of the upper gastrointestinal tract including, but not limited to, esophageal neoplasia.
[0065] Aspects of the present disclosure relates at least in part to the identification of genomic loci whose altered DNA methylation is indicative of the presence of esophageal neoplasias and / or metaplasias that include Barrett's esophagus (BE) and / or esophageal adenocarcinoma (EAC). In some embodiments, the Barrett's esophagus is associated with dysplasia. In some embodiments, the dysplasia is high-grade dysplasia. In some embodiments, the dysplasia is low-grade dysplasia. In some embodiments, the methylation patterns of the informative loci as disclosed herein are determined in a sample taken from a subject as described herein and may be used to distinguish between subjects having Barrett's esophagus and subjects having high grade dysplasia and / or low-grade dysplasia and / or esophageal adenocarcinoma. Examples of the informative loci are provided herein.
[0066] The present disclosure relates at least in part to identifying methylation of vimentin (VIM) and cyclin Al (CCNA1) genes in esophageal cells as a biomarker for detection of BE and EAC. More particularly, the present disclosure teaches a Next Generation Sequencing (NGS) assay that is used to analyze a sample collected from a patient. The sample may be a tissue sample taken from the esophageal tissue of the patient. In some embodiments, this NGS assay may be combined with a custom bioinformatics algorithm, or machine learning model, in order to determine methylation of VIM and / or CCNA1 genes in esophageal cells as biomarkers for the detection of BE and / or EAC. In these embodiments, the present disclosure may provide for a method of detecting biomarkers that maintain high sensitivity while maintaining specificity.
[0067] For example, FIG. 1 is a schematic illustrating an embodiment of a method 100 that may be used for identifying methylated cytosines in CpG dinucleotides in target nucleic acid sequences or informative loci. This may include vimentin nucleic acid sequences or portions thereof, or a CCNA1 nucleic acid sequence, or portion thereof, as a biomarker for detection of a disease. In this embodiment, the disease may be at least one of BE and EAC, however, the method may be used for detection of other diseases.
[0068] With reference to FIG. 1, the method 100 includes a step 102 of providing a tissue sample from a subject. In some embodiments, the tissue sample refers to a collection of cells that may be acquired from the tissue of a patient through various collection means. The sample may be collected through the use of cytology brushing, through a balloon collection device such as the use of EsoCheck®, through a sponge collection device such as the use of Esophacap® or Cytosponge® or through a biopsy procedure. In further embodiments, other established methods for collecting a sample may be used. The sample of cells may be taken from the tissue lining the esophagus or the upper stomach of the patient. However, in further embodiments, the sample may be taken from any taiget location within the patient in order to evaluate the tissue for biomarkers of a disease. Step 102 may further include isolating the DNA within the sample through at least one of a variety of methods, including but not limited to a manual, semi-automated, or automated high-throughput method. These methods may be well- known by those of ordinary skill in the art and may be used to isolate and purify the DNA.
[0069] At step 104, the method 100 further includes treating the sample to convert DNA within the sample. In these embodiments, the DNA may contain CpG islands of a VIM nucleic acid sequence and / or a CCNA1 nucleic acid sequence. Step 104 of treating the sample may include one of chemically converting and enzymatically converting the DNA within the sample . Converting the sample, either chemically or enzymatically, includes converting the sample to differentiate methylated DNA bases from unmethylated DNA bases, and more particularly differentiating methylated cytosine from unmethylated cytosine. For example, chemically converting the sample may include treating the sample with a bisulfite solution such that unmethylated Cytosine bases within CpG islands are converted to other bases while methylated Cytosine bases remain. The process of chemically converting the sample results in modified sample having a plurality of bisulfite converted sequences. As previously disclosed herein, the step 104 of treating the sample may include treating the sample with enzymatic bisulfite conversion. For example, in these embodiments, enzymatic bisulfite conversion includes treating the sample with Tet methylcytosine dioxygenase 2 (TET2). TET2 causes enzymatic oxidation of 5-methylcystosine (5mC) and 5 -hydroxymethylcytosine (5hmC) int 5- carboxycytosine (5caC). In some embodiments, NEBNext® Enzymatic Methyl-seq Conversion Module or equivalent module may be used to complete the enzymatic conversion. Enzymatic bisulfite conversion also results in a modified sample having a plurality of bisulfite converted sequences, some of which may be differentiated as methylated sequences, which can be detected in later steps of the method 100. In some embodiments, enzymatic conversion maybe preferred in comparison to chemical conversion as it provides the same result of a modified sample but may be less harsh on the DNA sequence which may protect the signal. In some embodiments, additional methods of methylation detection may be incorporated into the method 100. For example, targeted capture of methylated DNA using hybridization probes may be used for detecting methylation within the sample.
[0070] With continued reference to the flow chart of FIG. 1, method 100 further includes step 106 of replicating the converted DNA sequences. In embodiments, this includes replicating the converted DNA sequences into at least three copies to form at least triplicates of the converted DNA sequences. This may provide the benefit of optimizing the VIM and CCNA1 genes within the DNA sample for a multiplexing PCR process that will subsequently be applied. More specifically, this allows for PCR amplification to be applied to the same converted DNA sequences within the sample, at least three independent times. As will be described further herein, this may allow for consensus calling to be utilized when identifying the disease state of the sample.
[0071] Further, in some embodiments, step 106 may additionally include reducing an amplicon size of the nucleic acid sequence of interest (e.g. CCNA1 nucleic acid sequence, VIM nucleic acid sequence) within the sample. The amplicon size may be maintained at a value that is less than or equal to approximately 150 base pairs. This process mitigates the occurrence of DNA degradation during the testing process which may otherwise occur more frequently. This DNA degradation may reduce amplification efficiency when the length of the nucleic acid sequence of interest (e.g. CCNA1 amplicon, VIM amplicon) is over approximately 180 base pairs. Thus, in some embodiments, step 106 may include reducing the size of the nucleic acid of interest (e.g. CCNA1 nucleic acid, VIM nucleic acid) such that the length of the nucleic acid sequence of interest (e.g. CCNA1 nucleic acid sequence) is between approximately 120 base pairs and approximately 150 base pairs. In some embodiments, the nucleic acid sequence of interest (e.g. CCNA1 nucleic acid sequence) may have a length of between approximately 120 base pairs and approximately 145 base pairs. In some embodiments, the nucleic acid sequence of interest (e.g. CCNA1 nucleic acid sequence) may have a length of between approximately 120 base pairs and approximately 140 base pairs. In some embodiments, the nucleic acid sequence of interest (e.g. CCNA1 nucleic acid sequence) may have a length of between approximately 120 base pairs and approximately 135 base pairs. In some embodiments, the nucleic acid sequence of interest (e.g. CCNA1 nucleic acid sequence) may have a length of between approximately 120 base pairs and approximately 130 base pairs. In some embodiments,the nucleic acid sequence of interest (e.g. CCNA1 nucleic acid sequence) may have a length of between approximately 120 base pairs and approximately 125 base pairs. In some embodiments, the nucleic acid sequence of interest (e.g. CCNA1 nucleic acid sequence) may have a length of between approximately 125 base pairs and approximately 145 base pairs. In some embodiments, the nucleic acid sequence of interest (e.g. CCNA1 nucleic acid sequence) may have a length of between approximately 125 base pairs and approximately 140 base pairs. In some embodiments, the nucleic acid sequence of interest (e.g. CCNA1 nucleic acid sequence) may have a length of between approximately 125 base pairs and approximately 135 base pairs. In some embodiments, the nucleic acid sequence of interest (e.g. CCNA1 nucleic acid sequence) may have a length of between approximately 125 base pairs and approximately 130 base pairs. In some embodiments, the nucleic acid sequence of interest (e.g. CCNA1 nucleic acid sequence) may have a length of between approximately 130 base pairs and approximately 145 base pairs. In some embodiments, the nucleic acid sequence of interest (e.g. CCNA1 nucleic acid sequence) may have a length of between approximately 130 base pairs and approximately 140 base pairs. In some embodiments, the nucleic acid sequence of interest (e.g. CCNA1 nucleic acid sequence) may have a length of between approximately 130 base pairs and approximately 135 base pairs. In some embodiments, the nucleic acid sequence of interest (e.g. CCNA1 nucleic acid sequence) may have a length of between approximately 135 base pairs and approximately 145 base pairs. In some embodiments, the nucleic acid sequence of interest (e.g. CCNA1 nucleic acid sequence) may have a length of between approximately 135 base pairs and approximately 140 base pairs. In some embodiments, the nucleic acid sequence of interest (e.g. CCNA1 nucleic acid sequence) may have a length of between approximately 140 base pairs and approximately 145 base pairs. In some embodiments, the nucleic acid sequence of interest (e.g. CCNA1 nucleic acid sequence) may have a length of between approximately 145 base pairs and approximately 150 base pairs. In some embodiments, the nucleic acid sequence of interest (e.g. CCNA1 nucleic acid sequence) may have a length of between approximately 140 base pairs and approximately 150 base pairs. In some embodiments, the nucleic acid sequence of interest (e.g. CCNA 1 nucleic acid sequence) may have a length of between approximately 122 base pairs and approximately 135 base pairs. In further embodiments, the CCNA1 nucleic acid sequence may have a length of between approximately 122 base pairs and approximately 135 base pairs.
[0072] In some embodiments, method 100 then includes applying conditions for PCR amplification. During this step, magnesium concentrations may be adjusted to have values of 2 mM and approximately 4 mM and / or primer concentrations may be adjusted to have valuesof between approximately 80 nM and approximately 300 nM and may then be applied to the converted DNA sequences prior to amplification. Optimization of magnesium concentration is important for efficient amplification by both the vimentin primer and CCNA1 primer, while keeping off target amplification levels low. The above-described value ranges for the magnesium concentration and the primer concentrations may improve the amplification coverage of vimentin and CCNA1 amplicons. Applying the above-described conditions may optimize the performance of method 100, particularly for step 108 of amplification of the converted DNA sequences.
[0073] With continued reference to FIG. 1, method 100 may then include step 108 of amplifying the replicates of the converted DNA sequences to produce converted amplicons. More particularly, multiplex PCR amplification is conducted on the for example triplicates formed from steps 104, 106. In other words, for example PCR amplification is conducted three times on each of the VIM and CCNA1 genes since it is conducted on each of the triplicates of the DNA. With reference still to step 108, amplifying the triplicates allows for the signal to be detected from both the VIM gene and the CCNA1 gene using the same DNA template. In some embodiments, additional biomarker genes may be added to the assay to further improve the accuracy without requiring additional DNA input. For example, the DNA input can be about 50 ng. In some embodiments, the DNA input is from about 40 ng to about 50 ng, about 50 ng to about 60 ng, about 60 ng to about 70 ng, about 70 ng to about 80 ng, about 80 ng to about 90 ng, about 90 ng to about 100 ng ot higher, for example, about 40 ng, about 50 ng, about 60 ng, about 70 ng, about 80 ng, about 90 ng, about 100 ng or higher, While described herein as using multiplex PCR, the detecting of DNA methylation may be completed by any one or more of DNA sequencing, next generation sequencing, methylation specific PCR, methylation specific PCR combined with a fluorogenic hybridization probe, real time methylation specific PCR, hybridization to an array, or any combination of the above.
[0074] With continued reference to FIG. 1, step 110 of method 100 further includes analyzing the converted amplicons formed from the amplifying step 108 to determine a methylation status of the sample. More specifically, each individual converted amplicon that corresponds to a triplicate is analyzed in order to determine if it contains methylated CpG islands such that a disease state (positive (i.e., diseased) or negative (not diseased)) is identified for each triplicate. More particularly, this may include measuring the amount of methylated cytosines in CpG islands of each replicate of the vimentin nucleic acid sequence, or portion thereof, and / or of the CCNA1 nucleic acid sequence, or portion thereof.
[0075] For example, the method 100 includes analyzing each triplicate to determine the amount of methylated cytosines within the CpG islands of the triplicate. If the amount of methylated cytosines within these CpG islands is equal to or higher than the predetermined cut off value, the triplicate may be identified as diseased, while if the amount of methylated CpG islands within the triplicate is lower than the predetermined cut off value, the triplicate may be identified as healthy. More particularly, if at least 65%, 70%, 75%, or 80% of the cytosines in CpG dinucleotides in the each of the triplicate amplicons, or portion thereof, are methylated, then the triplicate amplicon or portion thereof, is considered a methylated read. Similarly, if at least 60%, 70%, or at least 75% of the cytosines in CpG dinucleotides in each CCNA1 nucleic acid sequence triplicate amplicon are methylated, then the CCNA1 nucleic acid sequence triplicate amplicon, or portion thereof, is considered a methylated read. Further, if at least 80% or at least 75% of the cytosines in CpG dinucleotides are methylated, then the CCNA1 nucleic acid sequence triplicate amplicon, or portion thereof, is considered a methylated read.
[0076] Furthermore, once the number of triplicate amplicons within each triplicate are identified as methylated reads, the number of methylated reads is determined and compared to the cut off in order to identify whether the triplicate should be labeled as a methylation positive (diseased) or methylation negative (not diseased). If at least 0.5%, 1.0%, 1.5%, 2.0%, 2.5%, or 3.0% of the CCNA1 nucleic acid sequence triplicate amplicons, or portions thereof, are designated as methylated reads, then the triplicate is identified as being methylation positive. Similarly, if at least 0.5%, 1.0%, 1.5%, 2.0%, 2.5%, or 3.0% of the vimentin nucleic acid sequence triplicate amplicons, or portions thereof, are identified as methylated reads, then the triplicate is identified as being methylation positive. The above-described process of analyzing methylation values is conducted for each triplicate.
[0077] As previously described, this measuring process of analyzing whether each sequence is a methylated read is conducted for each converted copy of the triplicate that is produced from steps 106, 108. If a majority of the triplicates of an amplicon are determined as methylated reads, then the amplicon is designated as the methylated read. However, if a majority of the triplicates are determined to not be methylated reads, then the amplicon is not designated as a methylated read.
[0078] The method 100 further includes step 110 which includes identifying if the sample is one of BE and / or EAC. More particularly, this is completed by looking at the status of the plurality of triplicates for each of the CCNA1 nucleic acid sequence and the vimentin nucleic acid sequence. For example, the three triplicates of the CCNA1 nucleic acid sequence arereviewed. If at least two of the three triplicates were identified as methylation positive, the CCNA1 nucleic acid sequence, and thus the sample, is identified as diseased. Similarly, the three triplicates of the vimentin nucleic acid are reviewed and if at least two of the three triplicates were identified as methylation positive, the vimentin nucleic acid, and thus the sample, is identified as diseased.
[0079] If at least 0.5% or at least 3% of the vimentin nucleic acid sequences, portions thereof, in the sample are designated as methylated reads, then the subject is determined to have an esophageal neoplasia or metaplasia. Similarly, if at least 0.5%, 1%, 1.5%, 2.0%, 2.5%, or 3% of the CCNA 1 nucleic acid sequence, or portions thereof, in the sample are methylated reads, then the subject is determined to have an esophageal neoplasia or metaplasia. In some embodiments, if at least 3% of the CCNA1 nucleic acid sequence, or portions thereof, in the sample are methylated reads, then the subject is determined to have an esophageal neoplasia or metaplasia.
[0080] In further embodiments, step 110 may include running an Artificial Intelligence / Machine Learning (AI / ML) model with at least the results from step 108 to aid in identifying whether or not the sample is diseased or healthy based on risk factors, which may include, but are not limited to, age, gender, BMI or obesity, or ethnicity as described below). In some embodiments, the AI / ML model may help predict a disease stage of the sample. In some embodiments, the AI / ML model may incorporate additional characteristics for each patient to better predict whether the sample is healthy or diseased. For example, the patient’s age, gender, and various other risk factors for BE / EAC as mentioned in established guidelines, may be inputted into the AI / ML to better predict the disease status of a patient. However, various other features or factors may be incorporated to aid in the robustness and accuracy of the model. Vimentin and CCNA1 methylation percentages with or without other risk factors may be used to train the data for the machine learning model against available endoscopy data such that the model with the highest or most accurate predictability may be used to test separate validation data sets for robustness of the tests.EXAMPLES
[0081] The following examples should not be construed as limiting the scope of this disclosure.Example 1Method
[0082] The validation of the assay incorporating embodiments of method 100, involved the assessment of its performance using both discarded clinical specimens and contrived cell line samples. This evaluation aimed to measure accuracy, sensitivity, specificity, reproducibility, and the limit of detection (LOD). For the collection of clinical samples, EsoCheck®, a non- invasive, FDA 510(k)-cleared, balloon cell collection device, was employed. After collection, the balloon was placed in a proprietary collection medium and sent to the Central Lab. The balloon containing the sample was stored in proprietary preservative media at room temperature until DNA extraction. The cells were harvested by centrifugation from the preservative media at 4000 rpm for 5 min after removing the balloon by disposable forceps. Cells were lysed at 56 ± 2°C in shaking heat block (1200 rpm for 30 min) in presence of proteinase K. The lysed cells were subjected to automated bead -based (NuCleoMag, Macherey-Nagel) purification in KingFisher Apex instrument (Thermo Fisher). Eluted DNA was quantified by Qubit™ dsDNA HS Assay reagent (Thermo Fisher) following standard protocol (Thermo Scientific).
[0083] In the case of contrived samples, a 100% methylation-specific cell line (SKTG-4, ATCC) DNA was spiked into 0% methylation cell line (NCI-H1975, ATCC) DNA at variable concentrations. For stability assessment both cell lines were mixed to create a 1% methylated sample, incubated with preservative up to 14 days, and then subjected to extraction.
[0084] The extracted DNA underwent bisulfite modification, followed by bead-based cleanup. Each marker in the assay was PCR amplified using indexed primers. Pooled libraries were ligated with adapters, and a final amplification step prepared them for sequencing. A validated in-house pipeline processed the reads, and samples were assigned a positive or a negative status based on pre-established cutoff.Results
[0085] The historical assay performance against endoscopy is shown in FIGs. 2A-2B using a singleplex amplification method. More particularly, FIG. 2A illustrates the area under the curve (AUC), sensitivity, and specificity values observed for VIM and CCNA1 as biomarkers for BE and EAC, with the performance of VIM and CCNA1 being analyzed separately. Clinical specificity and sensitivity for detecting BE and / or cancer using EsoCheck® or cytology brushing samples is shown in FIG. 2B as discussed in the paper titled “Identifying DNA Methylation Biomarkers for Non-Endoscopic Detecting of Barrett’s Esophagus” in the National Library of Medicine and authored by Moinova, et al.
[0086] Analytical validation of the assay (vl.O), which uses singleplex amplification, showed 98.5% sensitivity (n=66), 86.9% specificity (n=23) and 95.5% accuracy (n=89) compared to the reference dataset which included 49 DNA samples collected using EsoCheck® device and 40 contrived DNA samples, as is shown in FIGs. 3A and 3B.
[0087] FIGS. 4A and 4B illustrate the stability of methylated signal in both VIM and CCNA1 genes measured over the course of 14 days. As shown, the methylation signal was stable for up to 14 days in preservative at room temperature.
[0088] As shown in FIG. 5, validation of the assay which included high-throughput bisulfite conversion along with automated bioinformatics pipeline provided accuracy for 98%, with sensitivity and specificity at 100% and 95.5%, respectively in comparison to v 1.0 assay. Further, as shown in the table of FIG. 6A, the assay displayed limit of detection (LOD) as low as 1 methylated cell in the background of 400 unmethylated cells (0.25 p).
[0089] Additionally, the assay was further optimized and validated for multiplexed testing of VIM and CCNA1 genes at accuracy of 97% (n=173) in comparison to the single-plex assay, as shown in FIG. 7A. Further, the multiplexed testing of VIM and CCNA1 genes showed improved operational efficiency.
[0090] FIG. 6B is an assay linearity graph illustrating the R2of the singleplex assay. FIG. 7B is an assay linearity graph of values for multiplex testing of VIM and CCNA 1 genes.
[0091] The assay quantity not sufficient (QNS) rate was calculated at 5.5%, among 4609 samples tested from Jan 2023 - July 2023 for 100 ng minimal input, as shown in FIG. 8.
[0092] The assay has been successfully transferred from academic lab to a central CAP accredited, CLIA certified lab (LucidDx Labs, Lake Forest, CA). The assay has been optimized for samples collected and shipped at RT making it applicable in real-world. The methylation signal is stable until fourteen days at RT in preservative. Limit of detection of the tests is one methylated cell in the background of four hundred unmethylated cells, making it a highly sensitive test that has ability to detect early molecular changes indicative of BE And / or EAC. The assay has been further optimized for high-throughput, robust testing and has continued to show high performance (greater than 95% accuracy) in updated version (vl .2 and v2.0) with a QNS rate of 5.5%. The unevaluable rate, which is a measure of the robustness of the EsoGuard® assay reduced to 4.6% (n=2100) to 1.4% (n=1772). Overall, EsoGuard® is a highly sensitive and accurate biomarker test which has the potential to improve BE screening when used in combination with non-invasive sample collection using EsoCheck®.
[0093] In order to improve the assay robustness at LOD, each sample could be tested in triplicate as technical replicates by splitting the bisulfite converted consensus DNA and employing consensus calling for determining the disease state of the sample.Example 2: Analytical Validation of a DNA Methylation Biomarker Test for Diagnosis of BE and EAC from Samples Collected Using EsoCheck®, a Non-endoscopic Esophageal Cell Collection Device
[0094] This example presents analytical validation studies including analytical accuracy, analytical sensitivity, analytical specificity, linearity, lower limit of detection (LLOD), limit of blank (LOB), inter and intra-assay precision, reference interval and reportable range for EsoGuard® assay. Additionally, the example presents data for sample stability in preservative media during sample transportation as well as interfering substance study.Materials & MethodsCell culture:
[0095] NCI-H1975 cell line (ATCC, CRL-5908) and SK-TG-4 cell line (Millipore Sigma, 11012007) were utilized to create contrived specimens. The cells were cultured in RPMI-1640 medium (ATCC, 30-2001) supplemented with 10% fetal bovine serum (ATCC, 30-2020) in presence of lx antibiotics mixture (Penicillin and Streptomycin). The cells were passaged either in 1 :2 or 1 :3 ratio following trypsinization and total cell count was determined by trypan blue staining with counting in Countess 3 FL-Automated Cell counter (Themo Scientific, CA). Each of the cell lines (SK-TG-4 (100% methylated) and NCI-H1975 (0% methylated)) were mixed in 1: 100 ratio for creating 1% spike-in.Sample Collection:
[0096] Esophageal samples were collected using EsoCheck® (K230339) as per its Instructions for Use (IFU). The balloon containing the sample was stored in proprietary preservative media at room temperature until DNA extraction.DNA extraction from Balloon samples:
[0097] The cells were harvested by centrifugation from the preservative media at 4000 rpm for 5 min after removing the balloon by disposable forceps. Cells were lysed at 56 ± 2°C in shaking heat block (1200 rpm for 30 min) in presence of proteinase K. The lysed cells were subjected to automated bead-based (NuCleoMag, Macherey-Nagel) purification in KingFisher Apex instrument (Thermo Fisher). Eluted DNA was quantified by Qubit™ dsDNA HS Assayreagent (Thermo Fisher) following standard protocol (Thermo Scientific).EsoGuard® Assay:
[0098] The extracted DNA was bisulfite converted by using Lightning bisulfite conversion kit (Zymo, D5046) and cleaned up using KingFisher Apex instrument. Purified bisulfite converted DNA was either tested in singleplex or multiplex for VIM and CCNA1 genes using primers and polymerase chain reaction (PCR) conditions as previously described [Moinova, H.R., et al., 2018], In multiplex testing, each bisulfite converted DNA sample was divided into three technical replicates prior to PCR. The indexed amplicon library was pooled from multiple samples and subjected to AMPureXP bead clean up followed by end-repair, A-tailing step (NEBNext Ultra II kit (E7645L)) and dual index adapter ligation (NEXTflex dual index adapter (Perkin Elmer)). The final library was quantified by Qubit™ dsDNAHS Assay and sequenced on MiSeq orNextSeq sequencing platform (Illumina).
[0099] Each EsoGuard® run included one low positive control (1% contrived), one negative control (0% contrived), one PCR blank and one bisulfite blank. Each run was evaluated for controls prior to analyzing results of samples.Bioinformatics Analysis:
[0100] DNA sequencing reads were processed in data analysis pipeline as described previously [Moinova, H.R., et al., 2018], The methylation status of each gene was determined based on cut-off as established from clinical validations [Moinova, H.R., et al., 2018, Moinova, H.R., et al., Official journal of the American College of Gastroenterology] . Final EsoGuard® result was considered positive if either VIM or CCNA 1 gene were positive and negative if both genes were negative. In case of triplicate results, the methylation status of each sample was determined based on majority call from all three technical replicates of a sample. Each sample was evaluated for mapping efficiency (>80%), bisulfite conversion efficiency (>95%), % reads covering full length CpG (>90%) and depth of coverage of each gene (>10,000X coverage) as QC parameters. If the sample failed any of these QC criteria, the sample was reported by pipeline as non-diagnostic.Sanger Sequencing:
[0101] The extracted DNA from SK-TG-4 and NCI-H1975 were bisulfite converted as above and amplified with specific primers for VIM and CCNA1 separately. The amplified PCR products were cleaned up by NucleoSpin Gel and PCR cleanup kit (Macherey -Nagel) following manufacturer’s recommendations and the cleaned-up PCR product was sent tooutside sequencing laboratory (Genewiz Inc.) for sequencing of both forward and reverse strands. The full-length DNA sequence was aligned by blastn (NCBI).Analytical Performances:
[0102] The analytical performances of the sample cohort were determined by comparing with the orthogonal tests for TP (true positive), true negative (TN), false positive (FP), or false negative (FN) samples. The assay sensitivity and specificity were calculated by TP / (TP+FN) and FN / (TN+FP) calls. For LOD, the DNA from SK-TG-4 was serially diluted from 1%, 0.5%, 0.25% and 0.13% in NCI-H1975. For stability in preservative media, 0.5 million cells mix (1%), and 0%controls were stored in preservative media for specific time and condition and subsequently DNA was extracted from the cell mixture for EsoGuard® assay. For the interfering substances Heme (7.0 mg / dL), Bilirubin (conjugated and unconjugated) (0.7 mg / dL) were used to incubate the cell line for 48 hours (Sun Diagnostics- Assurance, INT-01) at room temperature in the preservative media until extraction following standard protocol as mentioned above.Statistical Methods:
[0103] The sample mean was calculated by averaging the replicates and plotted with standard deviation. Assay linearity was assessed using R2 value from scatterplot.ResultsAccuracy of Methylation Calling:
[0104] The accuracy of the methylation calls made by EsoGuard® NGS assay at all 31 CpG sites (10 VIM CpG sites and 21 CCNA1 CpG sites) was confirmed against sanger sequencing as a gold standard. A total of 62 unique data points were evaluated (2 cell lines at 31 CpG sites) in triplicates. All thirty-one CpG sites showed methylation level as expected based on sanger sequencing (FIG. 12). FIG. 12 shows the following Homo sapiens genomic DNA sequences in order of appearance:SEQ ID NO: 1:CacagttgtttaggttgtaggtgcngggtggatgtagttatgtagttttggttggagtttgnttgntttgtggtgtttgggttgttgaatatttSEQ ID NO: 2:Gcataggttgtttaggttgtaggtgcngggtggacgtagttacgtagtttcggttggagttcggtcggttcgcggtgttcgggtcgtcga atattttSEQ ID NO: 3:Cgattgtatttggggtagttttgttgtgttttagttgttttttggtaggaagtgtaggtgtgtgagttgatttggagtgagttgtgtttttgggtta gtgtgggtagggtgttgtagtttgtgtagttttgaggattttgtgttgtttttttgagttagggtttttaggagcgggtgtgtataSEQ ID NO: 4: tggcatnngngattgtatttggggtagtttcgtcgcgttttagtcgtttttcggtaggaagcgtaggtgtgtgagtcgattcgnagcgagtc gcgttttcgggttagcgtgggtagggcgtcgtagtttgcgtagtttcgaggatttcgcgtcgttttttcgagttagggtttttagga
[0105] The methylation percentages at each CpGs site as calculated by EsoGuard® assay pipeline for both VIM and CCNA1 genes were within the expected range of variance (Table 1A, Table IB). The difference between average methylation % computed by EsoGuard® assay and sanger sequencing is expected given differences in sensitivity of two methods (Table 1C).Table 1A: Methylation percentage for all 10 CpG sites measured on Vimentin gene usingEsoGuard® assay.Table IB: Methylation percentage for all 21 CpG sites measured on CCNA1 gene usingEsoGuard® assay.* average %CV performed on average methylation % of all CpG sites.Table 1C: Accuracy of EsoGuard® assay in comparison to Sanger Sequencing across all31 CpGs sites.Analytical Sensitivity, Specificity and Accuracy:
[0106] Analytical accuracy of the EsoGuard® assay was first measured by sample exchange with Case Western Reserve University (CWRU), where the method was developed. The accuracy data set consisted of blinded patient DNA specimens collected using EsoCheck® balloon (n=49) and contrived cell line mixtures (n=40). The balloon samples showed analytical sensitivity, specificity, and accuracy of 97%, 81% and 92%, respectively, while the contrived cell lines showed 100% sensitivity, specificity, and accuracy. Overall assay sensitivity, specificity and accuracy was 99%, 87% and 96% when compared to reference lab (Table 2 A,B and C).Table 2A: Analytical accuracy of EsoGuard® assay in comparison to reference lab in balloon samplesTable 2B: Analytical accuracy of EsoGuard® assay in comparison to reference lab in contrived samples.Table 2C: Analytical performance of EsoGuard® assay (singleplex) against reference lab
[0107] Recently, EsoGuard® assay was upgraded to allow multiplex testing of VIM and CCNA1 genes and sequencing on MiSeq or NextSeq platform based on sample volume. Additionally, samples are now tested in triplicates instead of singlicate. The analytical sensitivity, specificity, and accuracy of the multiplex EsoGuard® assay was performed against the singleplex assay. The analytical accuracy study consisted of a total of 77 EsoCheck® collected samples (27 positive and 50 negative). The assay displayed 89% analytical sensitivity, 100% analytical specificity, and 96 % analytical accuracy (Table 3). Three false negative samples were discordant for VIM gene only and methylation percentage were close to the cutoff of >1.0% for VIM (1.1%, 1.1% and 1.2%) (Table 8).Table 3: Analytical accuracy of multiplex EsoGuard® assay in comparison to singleplex assay.Table 8 Analytical accuracy of multiplex EsoGuard® (EG) assay in comparison to singleplex.*na=non-diagnostic result due to QC failureBioinformatic Pipeline Accuracy:
[0108] The accuracy for bioinformatic pipeline was evaluated against reference laboratory (CWRU) bioinformatics pipeline by testing same data set (n=32) between the sites. Theaccuracy was found at 100% for the pipeline with R2 of 1 suggesting identical data obtained using both pipelines (data not shown). Recent updates to the pipeline which included updates to the bioinformatics tools and automated initiation of data analysis pipeline were all validated against previous version of the pipeline and showed accuracy of 100% (n= 207) (FIG. 13).Accuracy for Sequencing Platform:
[0109] Accuracy of NextSeqlOOO sequencing platform against MiSeq platform was performed. The same library was sequenced both in MiSeq andNextSeq platforms. The sample cohort consisted of 62 EsoCheck® collected samples (33 positive and 29 negative) and 30 contrived cell line samples. The accuracy for the data set between the platforms was 100% with R2of 1 (FIG. 14).Intra-Assay and Inter-Assay Precision:EsoGuard® Assay Precision:
[0110] The intra-assay and inter-assay precision was performed by testing a sample cohort of 10 samples (four contrived and six EsoCheck® collected samples). The contrived sample set consisted of a mixture of positive and negative cell line mimicking medium and low positive samples and a negative sample ( 10%, 1%, 0.5%, and 0%), while EsoCheck® samples consisted of three positive and three negative clinical discard specimens. The intra-assay precision was performed by testing the same samples three to four times on same day, on same run, by a single operator, while inter-assay precision was performed by testing same samples three times, on three different days, by different operators. The inter and intra-assay precision for the EG assay was 100% for all samples, and the %CV ranged from 1 %-34% for Vim and 6%-67% for CCNA1 in positive samples and it was 0% for all negative samples (Table 4 and 5). As expected, the CV% was higher in the low positive samples, however, the EsoGuard® assay binary results were consistent across all replicates, showing the higher CV% does not impact EsoGuard® results as the results are based on both Vimentin and CCNA1 genes and triplicate data points.Table 4: Intra-assay precision.*CV% which could not be calculated due to division by 0 have been converted to 0% manuallyTable 5: Inter-assay precision.*CV% which could not be calculated due to division by 0 have been converted to 0% manuallyAssay Linearity and Limit of Detection (LOD):
[0112] The assay linearity was measured by testing serially diluted contrived cell line mixture (1% ,0.5% ,0.25% and 0.13% spike-in). Each dilution was tested in at least 20 replicates for each dilution. A positive correlation was observed for both VIM and CCNA1 genes with R2 at 0.99 and 0.92 respectively suggesting that the assay is linear (FIG. 9). The LOD of the assay was determined to be 0.25% spike-in or one methylated cell in the background of 400 unmethylated cells as 0.25% was the lowest dilution at which >95% data points were positive (Table 6).Table 6: Limit of Detection (LOD).Assay Input Range:
[0113] The DNA input range for the EsoGuard® assay was tested using EsoCheck® collected samples (three positive and three negative) along with contrived specimens mimicking low positive samples (0.5% and 1%, spike-in). All samples were tested at 50, 80,100, and 300 ng DNA input. EsoGuard® assay concordance was 100% for the range (data not shown) . However, for the EsoGuard® assay minimum DNA requirement was kept at 100 ng as per clinical validations which were performed at 100 ng DNA input.Reference Range:
[0114] The assay reference range was determined by testing 82 samples that were confirmed negative for BE by endoscopy. In the sample cohort 69 samples were detected as true negative, whereas 13 samples were false positive suggesting a specificity at 84.1% in normal healthy patients (Table 7).Table 7: Reference Range.Reportable Range:
[0115] The EG assay report provides a qualitative, binary (“positive” or “negative”) result. Apositive test result suggests the presence of BE (metaplasia or dysplasia) or EAC and confirmatory EGD is recommended. A negative result does not warrant any further testing unless medically necessary. Infrequently, cell samples may have DNA “Quantity Not Sufficient” for EG analysis or the samples may have quality issues prohibiting analysis. These are reported as “QNS” or “Unevaluable” results, respectively. If this occurs, patients and providers are given the option to repeat the EC cell collection to provide an analyzable sample.Limit of Blank (LOB):
[0116] Negative cell-line control with known methylation status (H-1975) was tested in multiple replicates across multiple runs (n=60) by different operators. VIM and CCNA1 methylation percentages for all the data points were reported at 0% both for VIM and CCNA1 (data not shown).Sample Stability in Preservative Media:
[0117] Sample stability in preservative media was tested for up to 21 days. Low positive cell line spike-in (1% contrived) were used along with negative cell line (0%). Each cell line sample was incubated for 0, 2, 7, 14, and 21 days (n=6 for each) at room temperature (25 ± 3 °C) in preservative media. The EsoGuard® assay data showed 100% concordance for both positive and negative samples till 21 days (Table 9). The dot plot shows the distribution of methylation reads at different days for both genes (FIG. 10, panels A and B). Further the effect of any extreme temperature during shipping was evaluated by incubating the low positive and negative cell line controls at -20°C, 4°C, Room Temperature (25 ± 3 °C), 37°C and 50 °C for 2 days followed by incubation at room temperature for 12 days. The assay displayed 100% concordance at all temperatures (Table 10).Table 9: Sample Stability in Preservative media. 1% contrived specimen (1 P) and 0% contrived specimen (OP) were tested at Day 0 (DO), Day 2(D2), Day 14 (D14) and Day 21 (D21)Table 10: Sample Stability in Preservative media at different temperature on day 14.Interference Testing:
[0118] Some of the substances that may be present in samples can interfere with test results. Interference on EsoGuard® assay by substances like hemolysate (heme), bile (conjugated or unconjugated bilirubin) and triglyceride rich lipoprotein was evaluated as these substances are generally present in blood and if high amount of blood contaminates EsoCheck® sample during collection then they may interfere with test results. Known negative (0%) (n=3) and low positive contrived cell line samples (1%) (n=3) were incubated in the preservative media for 2 days with the interfering substances and tested in the EsoGuard® assay after extraction. The methylation data from the EsoGuard® assay confirmed 100% concordance indicating nointerference was observed due to any of the substances tested (FIG. 11, panels A, and B) for both negative and positive samples.Discussion
[0119] EsoGuard® is first of its kind DNA methylation biomarker assay which shows unprecedented clinical sensitivity (81% - 92%) and specificity (72% -92%) for detecting esophageal pre-cancer (BE), in non-endoscopically collected esophageal cell samples collected using EsoCheck® balloon device [Moinova, H.R., et al., Identifying DNA methylation biomarkers for non-endoscopic detection of Barrett's esophagus. Sci Transl Med, 2018. 10(424); Moinova, H.R., et al., Multicenter, Prospective Trial Of Non-Endoscopic Biomarker-Driven Detection Of Barrett’s Esophagus And Esophageal Adenocarcinoma. Official journal of the American College of Gastroenterology | ACG, 9900: p. 10.14309 / ajg.0000000000002850; Greer, K., et al., Acceptability Of Non-Endoscopic Screening For Barrett's Esophagus (Be) Among Veterans Eligible For Be Screening. Gastrointestinal Endoscopy, 2023. 97(6): p. AB 1057], EsoGuard® assay investigates methylation at 31 CpG sites in the regulatory regions of Vimentin and CCNA1 genes, which has been correlated with metaplastic to neoplastic changes in distal esophageal cells. It utilizes a proprietary algorithm that calculates methylation per read for both genes and provides percentage of sequencing reads that are methylated. If either of the VIM or CCNA1 gene has methylated reads above pre-established cutoff based on prior clinical validation [Moinova, H.R., et al., Official journal of the American College of Gastroenterology] then EsoGuard® assay is reported positive.
[0120] In this study we have shown that EsoGuard® DNA methylation test has high analytical sensitivity (89%), specificity (100%), and accuracy (96%). Additionally, the assay is linear (VIM, R2=0.998 and CCNA R2=0.918) and has a limit of detection as low as one methylated cell in the background of four hundred unmethylated cells as measured using contrived cell line mixtures. The assay displayed 100% inter and intra-assay precision in all samples (contrived and clinical samples) tested in this study. The EsoCheck® samples collected and preserved in EsoGuard® preservative media displayed expected result for as long as 21 days at room temperature. The assay has shown concordant results in DNA input as low as 50 ng. However, as the clinical validation studies were performed using minimal 100 ng DNA input, the minimal DNA input requirement for the assay is kept at 100 ng now which may be further reduced in future based on additional clinical data.
[0121] Overall EsoGuard® analytical validation studies were performed as per CLSIguidelines [Evaluation of Qualitative, Binary Output Examination Performance. Clinical and Laboratory Standards Institute, 2023. 3rd ed] and it met all requirement as per the standards of College of American Pathology (CAP), Clinical Laboratory Improvement Amendment (CLIA), and New York state to be approved as a Laboratory Developed Test (LDT).
[0122] EsoGuard® assay being a qualitative, binary assay could have a limitation that results near the cutoff may have lower precision. To quantify the percentage of samples in the real world that may fall in this grey zone area or zone of uncertainty, we first determined grey zone area for each gene using clinical specimens near cutoff and tested in multiple replicates (333 data points). The grey zone for VIM was determined to be 0.58%- 1.95 %(95 %CI) and CCNA1 was 0.24-1.57%(95% CI). 2617 samples that were previously tested in clinical lab were retrospectively evaluated if the methylation % underlying EG result falls in the grey zone for either VIM, CCNA1 or both. It was determined that only 1.7% of the samples had both Vim and CCNA1 genes in grey zone of methylation and could be impacted by imprecision near the cutoff. Therefore, EsoGuard® binary results provide sufficient confidence in the data and grey zone does not impact significant number of samples. Additionally, recent screening study performed in intended use population using EsoGuard® assay run at LucidDx labs has shown high sensitivity of the assay (92.9%) at specificity of 72.2% with PPV of 32.5% and NPV of 98.6% [Greer, K., et al., 2022] indicating that EsoGuard® binary results are optimized for high sensitivity and high NPV of the assay provides confidence that patients with negative results do not need any further testing for BE screening. As all positive samples do get tested by EGD, binary results are appropriate and sufficient for EsoGuard®.
[0123] Another limitation of the EsoGuard® assay is that ~3 - 6% of EsoCheck® collected samples does not yield enough DNA for EsoGuard® testing as shown by recent clinical utility studies [Lister, D., et al., Clinical Utility of EsoGuard® on Samples Collected with EsoCheck® as a Triage to Endoscopy for Identification of Barrett’s Esophagus - Interim Data from the CLUE Study. Archives of clinical and biomedical research, 2023. 7(7): p. 626-634; Richard Englehardt et al., Real World Experience and Clinical Utility of Esoguard® - Interim Data from the Lucid Registry. Journal of Gastroenterology & Digestive Systems, 2023. 7(2): p. 43-53.]. This is a significant improvement from previous clinical validation study in which QNS rate was reported at 14%, where column based DNA extraction method was utilized as opposed to recent studies where magnetic bead based DNA extraction method was employed [Moinova, H.R., et al., Official journal of the American College of Gastroenterology]. However, there is further room for improvement either in cell collection method and / or in DNA extractionmethods. EsoCheck® cell collection device is very gentle where no abrasion are observed in 98.5% of patients who went through EGD right after EsoCheck® (internal non-published data). In future versions of EsoCheck® device, the cell collection device could be made more abrasive to collect more sample. Additionally, any cell loss during transfer of cell pellet from preservative media tube (50 ml conical tube) to sample processing tube (2 ml Eppendorf) can be prevented by transporting samples in smaller tube. Additionally, EsoGuard® assay has shown consistent performance in DNA input as low as 50 ng. Therefore, in future with further clinical studies DNA input requirement could be further brought down. All these improvements in combination can further bring the QNS rate down in future.
[0124] In summary, EsoGuard® assay is an accurate, robust, and highly sensitive test which in combination with EsoCheck® cell collection device could provide as an effective triage tool for screening of BE by EGD and improve overall compliance with BE screening guidelines.
[0125] While the present disclosure has been described with reference to certain embodiments thereof, it should be understood by those skilled in the art that various changes may be made and equivalents may be substituted without departing from the true spirit and scope of the disclosure. In addition, many modifications may be made to adapt to a particular situation, indication, material and composition of matter, process step or steps, without departing from the spirit and scope of the present disclosure. All such modifications are intended to be within the scope of the claims appended hereto.INCORPORATION BY REFERENCE
[0126] All publications, patents and sequence database entries mentioned herein are hereby incorporated by reference in their entirety as if each individual publication or patent was specifically and individually indicated to be incorporated by reference.
Claims
CLAIMS1. A method for detecting methylation status of at least one sequence in a tissue sample of a subject, comprising: providing the tissue sample obtained from the subject; treating the tissue sample to convert DNA sequences within the tissue sample with at least one of enzymatic conversion and chemical conversion to produce a modified sample having converted DNA sequences; replicating the converted DNA to produce at least three copies of each sequence of the converted DNA sequences; amplifying the converted DNA sequences with a polymerase chain reaction such that a plurality of converted amplicons is produced; and analyzing the plurality of converted amplicons to determine the methylation status of the sample.
2. The method of claim 1, wherein treating the sample includes treating the sample with enzymatic bisulfite conversion.
3. The method of claim 1, wherein treating the sample includes treating the sample with chemical bisulfite conversion through applying sodium bisulfite to the sample.
4. The method of claim 1, wherein analyzing the plurality of converted amplicons includes using a machine learning model to determine whether the sample is methylated.
5. The method of any one of claims 1-4, wherein analyzing the plurality of converted amplicons comprises, for each replicate: measuring a number of methylated cytosines within each of the plurality of converted amplicons, wherein a read is classified as a methylated read when at least 70% of the cytosines of each of the plurality of converted amplicon are methylated; and calculating percent of total reads are methylated reads from the amplifying and measuring steps, wherein the replicate is considered a positive sample when at least 0.5% of the total reads are methylated reads.
6. The method of claim 5, wherein the plurality of converted amplicons are vimentin nucleic acid amplicons, wherein a read is classified as a methylated read when at least 80% of the cytosines of the individual amplicon are methylated, and wherein the sample obtained from the subject is classified as a methylated sample when at least 1.0% of the total reads are methylated reads.
7. The method of claim 5 or claim 6. wherein the plurality of converted amplicons is a plurality of CCNA1 nucleic acid amplicons.
8. The method of claim 5 or claim 6, wherein if a majority of the plurality of converted amplicons are identified as methylated reads, the sample is identified as a methylated sample.
9. A method for detecting methylation status of at least one sequence in a human subject to determine whether the human subject is healthy or diseased, comprising: providing a tissue sample obtained from the human subject; treating the tissue sample to convert DNA within the tissue sample with at least one of enzymatic conversion and chemical conversion to produce a modified sample containing a plurality of converted vimentin nucleic acid sequences and a plurality of converted CCNA1 nucleic acid sequences; replicating the plurality of converted vimentin nucleic acid sequences and the plurality of converted CCNA1 nucleic acid sequences, such that each of the plurality of converted CCNA1 nucleic acid sequences has at least three replicates and such that each of the plurality of converted vimentin nucleic acid sequences has at least three replicates; amplifying each of the at least three replicates of each of the plurality of converted vimentin nucleic acid sequences such that a plurality of converted vimentin amplicons is produced; amplifying each of the at least three replicates of each of the plurality of converted CCNA1 nucleic acid sequences such that a plurality of converted CCNA1 amplicons is produced; and analyzing the plurality of converted vimentin amplicons and the plurality of converted CCNA1 amplicons to determine the methylation status of the sample.
10. The method of claim 9, wherein amplifying each of the at least three replicates of the plurality of converted vimentin amplicons and the plurality of converted CCNA1 amplicons includes undergoing polymerase chain reaction amplification.
11. The method of claim 9, wherein treating the sample includes treating the sample with enzymatic bisulfite conversion.
12. The method of claim 9, wherein treating the sample includes treating the sample with chemical bisulfite conversion through applying bisulfite to the sample.
13. The method of claim 9, wherein prior to the amplifying step, the method further comprises applying a solution for PCR preparation including a magnesium concentration of between approximately 2 mM and approximately 4 mM.
14. A method for detecting methylation status of CpG dinucleotide sites within a sequence in a human subject to determine whether the human subject is healthy or diseased, comprising: providing a tissue sample obtained from the human subject; treating the tissue sample to convert DNA within the tissue sample with at least one of enzymatic conversion and chemical conversion to produce a modified sample containing a plurality of converted vimentin nucleic acid sequences and a plurality of converted CCNA1 nucleic acid sequences; replicating the plurality of converted vimentin nucleic acid sequences and the plurality of converted CCNA1 nucleic acid sequences, such that each of the plurality of converted CCNA1 nucleic acid sequences has at least three copies and such that each of the plurality of converted vimentin nucleic acid sequences has at least three copies; amplifying each of the at least three copies of each of the plurality of converted vimentin nucleic acid sequences and each of the at least three copies of each of the plurality of converted CCNA1 nucleic acid sequences; identifying whether each of the at least three copies of each of the plurality of converted CCNA1 amplicons is considered a methylated read; and identifying whether each of the at least three copies of each of the plurality of converted vimentin amplicons is considered a methylated read; and identifying the sample as methylated if a majority of the at least three copies of the plurality of converted CCNA1 amplicons are identified as methylated reads or if a majority of the at least three copies of the plurality of converted vimentin amplicons are identified as methylated reads.
15. The method of claim 14, wherein prior to the amplifying step, the method further comprises applying a solution to the sample for PCR preparation including a magnesium concentration of between approximately 2 mM and approximately 4 mM.
16. The method of claim 14, wherein treating the sample includes treating the sample with enzymatic bisulfite conversion.
17. The method of claim 14, wherein treating the sample includes treating the sample with chemical bisulfite conversion.
Citation Information
Patent Citations
Cancer detection methods
US20180216195A1
Compositions and methods for preserving DNA methylation
US20220243281A1
Random epigenomic sampling
WO2023028270A1