Detection of methylation changes in DNA samples using restriction enzymes and high-throughput sequencing.
Patent Information
- Application Number
- JP2023530742
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-11-19
- Filing Date
- 2021-11-18
- Publication Date
- 2026-10-01
- Estimated Expiration
- 2041-11-18
Smart Images

Figure 0007927313000006 
Figure 0007927313000007 
Figure 0007927313000008
Abstract
Description
Technical Field
[0001] The present invention relates to methods and systems for profiling genetic and epigenetic characteristics of DNA samples, and particularly cell-free DNA samples obtained from body fluids (e.g., plasma and urine). The methods and systems of the present invention comprise digestion of DNA with methylation-sensitive or methylation-dependent restriction enzymes, preparation of sequencing libraries, high-throughput sequencing (e.g., next-generation sequencing), and analysis of sequence reads. Advantageously, the methods and systems of the present invention are sensitive yet accurate, allow working with very small amounts of DNA, and yield a vast amount of information including methylation data, mutation data, and the like based on sequencing data from a single run. The methods and systems of the present invention are useful, for example, for both the discovery of novel methylation markers and diagnostic applications in clinical settings.
Background Art
[0002] Genetic and epigenetic changes including mutations, DNA methylation alterations (e.g., hypomethylation of isolated CpGs and hypermethylation occurring primarily at CpG islands), copy number variations, and the like are known to occur in many types of cancer. For example, hypermethylation of CpG islands in the promoter region of tumor suppressor genes, which leads to gene silencing, has been extensively studied and demonstrated in various types of cancer.
[0003] Tumors release DNA fragments or "cell-free DNA" into bodily fluids, allowing for the detection of genetic and epigenetic changes in tumor-derived DNA molecules in "liquid biopsies" obtained from bodily fluids such as plasma and urine. In contrast to conventional biopsies, liquid biopsies are non-invasive and may better represent the complete genetic spectrum of tumor subclones. Consequently, the detection of cancer-related genetic and epigenetic changes in liquid biopsies holds great promise for early detection, prognosis, and therapeutic monitoring. However, detecting tumor-derived DNA in liquid biopsies requires highly sensitive biochemical techniques because the concentration of cell-free DNA in bodily fluids may be low, and tumor DNA may be present in very small amounts due to a large background of normal DNA.
[0004] Several techniques have been developed for detecting methylated DNA molecules in liquid biopsies, based on sodium bisulfite treatment of DNA to convert unmethylated cytosine to uracil, followed by quantitative PCR or sequencing of the converted DNA to detect methylation changes. The converted base is identified as thymine in the sequencing data (post-PCR), and the percentage (%) of methylated cytosine can be determined using the read count. Bisulfite conversion sequencing can be performed using targeted methods or whole-genome bisulfite sequencing. Advances in high-throughput sequencing, such as next-generation sequencing (NGS), enable both genome-wide and targeted approaches for identifying and analyzing methylation patterns at the single-nucleotide level.
[0005] Despite its popularity, DNA conversion with sodium bisulfite is a cumbersome assay, with drawbacks including degradation of template DNA, nonspecific or incomplete conversion that introduces noise into the assay, and a reduction in genomic complexity from a 4-nucleotide genome to approximately a 3-nucleotide genome, leading to decreased PCR specificity, increased bias in DNA amplification, and increased noise levels in DNA sequencing. Furthermore, bisulfite treatment alters the DNA sequence, hindering mutation analysis because transition events become obscured in one of the DNA strands due to the sequence change.
[0006] Ball et al. (2009) Nat Biotechnol., 27(4):361-368 reported two techniques for cytosine methylation profiling using next-generation sequencing technology: bisulfite padlock probe (BSPP) and methyl sensitive cut counting (MSCC).
[0007] Brunner et al. (2009) Genome Res., 19(6):1044-1056 reported Methyl-seq, a method for assaying DNA methylation in over 90,000 regions across the entire genome. Methyl-seq combines DNA digestion by methyl-sensitive enzymes with next-generation DNA sequencing technology.
[0008] Jelinek et al. (2012) Epigenetics, 7:12, 1368-1378 report a method titled Digital Restriction Enzyme Analysis of Methylation (DREAM), which is based on next-generation sequencing analysis of methylation-specific signatures created by sequential digestion of genomic DNA with a pair of neoschizomeric restriction enzymes that recognize the same sequence, Sma (methylation-sensitive enzyme) and XmaI (methylation-insensitive enzyme).
[0009] Marsh and Pasqualone (2014) Front Physiol, 5:173 report the characterization of methyl-cytosine composition patterns in the marine polychaete Spiophanes tcherniai from McMurdo Sound, Antarctica. The methylation patterns were characterized using DNA digestion with methylation-sensitive restriction endonucleases and subsequent next-generation sequencing.
[0010] Marsh et al., 2016, Front Genet., 7:191, reported a quantitative methodology for computer-based reconstruction of site-specific CpG methylation status from next-generation sequencing (NGS) data using methyl-sensitive restriction endonucleases (MSREs).
[0011] Viswanathan et al. (2019) Nucleic Acids Research, 47(19):e122 reports a single-tube enzymatic method, DNA Analysis by Restriction Enzymes (DARE), which enables quantitative analysis of both unmethylated and methylated DNA in the same sample. Information on both methylation states is captured by tagging of differential adapters of DNA fragments that are sequentially digested by pairs of methylation-sensitive and non-methylation-sensitive restriction enzymes.
[0012] Pereira et al. (2020) PLoS ONE, 15(6):e0233800 report a technique called methyl-sensitive DArT-seq (MS-DArT-seq) based on a combination of genome dual digestion followed by special adapter ligation and next-generation sequencing. Two libraries are constructed in parallel using restriction enzymes that target the CCGG site and exhibit contrasting methylation sensitivity (MspI, methylation-insensitive, and HpaII, which does not cleave when internal cytosine is 5'-methylated).
[0013] Tanaka et al. (2020) Analytical Biochemistry, 609:113977 reports an approach combining methylation-sensitive restriction enzymes (MSREs) and next-generation sequencing (NGS) to identify differentially methylated regions between chorionic villi (CVs) and maternal blood cells (MBCs).
[0014] U.S. Patent No. 10,392,666 discloses the analysis of a biological sample (e.g., plasma) containing a mixture of DNA from different genomes (e.g., fetal and maternal, or tumor and normal cells) for determining the methylation pattern (methylome) of DNA, more specifically, the methylation pattern (methylome) of a small number of genomes.
[0015] International Publication No. 2016 / 061624 discloses a method for identifying gene or genomic sites and regions suitable for methylation analysis. This method enables efficient genome-wide identification of target restriction sites and fragments that provide targets for subsequent analysis.
[0016] International Publication No. 2018 / 195211 discloses compositions, kits, and methods for constructing libraries for the simultaneous detection of genomic variants and DNA methylation status in limited DNA inputs, such as circulating polynucleotide fragments in the body of a subject, including circulating tumor DNA.
[0017] International Publications 2011 / 070441, 2017 / 006317, 2019 / 142193, and 2020 / 188561, assigned to the applicant of the present invention, disclose a method for detecting methylation changes in a DNA sample.
[0018] Having methods and systems that can generate various types of genetic and epigenetic data from cell-free DNA obtained from a single sample from a subject, and from sequencing data from a single run, would be extremely beneficial. [Overview of the Initiative]
[0019] The present invention relates to a method and system for profiling the genetic and epigenetic properties of DNA samples, and in particular cell-free DNA samples obtained from bodily fluids (e.g., plasma and urine). The method and system of the present invention includes digestion with at least one methylation-sensitive restriction enzyme, preferably multiple methylation-sensitive restriction enzymes applied simultaneously, preparation of a sequencing library using a library preparation method that preserves sequence information at the ends of DNA molecules in the sample, and analysis of high-throughput sequencing and sequence reading.
[0020] The present invention provides a simpler and more accurate assay compared to previously described methods, yielding higher quality sequencing data compared to bisulfite sequencing and enabling highly sensitive detection of cancer-related changes. A vast amount of information, including methylation data, mutation data, etc., can be obtained from a single run based on the same sequencing data, thus avoiding the need for parallel assays to obtain comprehensive genetic and epigenetic information. Importantly, it has been found that high-quality sequencing data can be obtained even from very small amounts of DNA without the need for amplification before library preparation. As illustrated below, the method disclosed herein can detect early-stage cancer genetic and epigenetic changes based on the amount of cell-free DNA obtained from a single standard blood test tube, even when the amount of tumor-derived DNA in plasma is very low.
[0021] The methods and systems of the present invention do not involve or require bisulfite conversion. The methods and systems of the present invention do not require alteration of the DNA sequence and enable simultaneous analysis of, for example, methylation, mutation, copy number, and nucleosome positioning based on the same sequencing data.
[0022] As illustrated below, a comparison of sequencing data obtained from cell-free DNA samples subjected to high-throughput sequencing following methylation-sensitive enzyme digestion with sequencing data obtained after bisulfite conversion and high-throughput sequencing showed that enzyme-treated DNA exhibited significantly better sequencing metrics (read count, mapping rate, etc.), copy number completeness, and nucleosome positioning completeness compared to bisulfite-treated DNA. Analysis of pooled cell-free DNA samples compared to individual samples showed that information loss occurred when small amounts of DNA were used. However, the quality of sequencing data from enzyme-treated DNA samples remained high, enabling reliable analysis, while bisulfite sequencing showed significantly reduced read count and mapping rate, effectively losing all copy number and nucleosome positioning information. Furthermore, sequencing noise was high in bisulfite-treated DNA samples, and as a result, mutations could not be distinguished from sequencing noise. Regarding methylation detection, methylation analysis of enzymatically treated DNA samples detected significantly more methylation changes in plasma compared to bisulfite-treated DNA samples.
[0023] As further illustrated below, bisulfite-treated DNA samples yielded a broader CG range at the lower end of sequencing depth compared to enzymatically treated samples, but a continuous and sharp decrease was observed in the number of CGs covered in bisulfite-treated DNA samples as depth increased. In contrast, enzymatically treated samples showed substantially constant coverage even at depths exceeding 250–300. At high depths, methylation-sensitive digestion provided significantly better CG coverage compared to bisulfite. Methylation-sensitive digestion provided coverage of millions of CGs at very high depths, thus enabling the detection of rare methylation signals, e.g., tumor-derived methylated DNA molecules in plasma in the early stages of tumors, which may be present in very low amounts in plasma, i.e., 1% or less of total cell-free DNA. The data showed that bisulfite did not provide sufficient coverage at the depths required to identify rare signals, and such rare signals are likely to be missed when using bisulfite sequencing for small amounts of DNA.
[0024] The present invention further discloses an improved method for determining the methylation level of a target genomic locus. The methylation analysis according to the present invention is performed on restriction loci, i.e., the restriction sites of restriction enzymes used in the assay. The methylation analysis disclosed herein is based on analyzing an alignment covering a given genomic region of at least 50 bp, preferably at least 100 bp, that contains the target restriction locus, and determining the read count of sequence reads covering the given genomic region. Such an alignment represents a DNA molecule of at least 50 bp (preferably at least 100 bp), where the analyzed restriction locus, as well as any additional restriction loci within the DNA molecule, are all methylated in the DNA sample, and therefore this DNA molecule remains intact after digestion with the enzyme used in this assay. Analyzing an alignment containing multiple restriction loci of at least 50 or at least 100 bp in length and all methylated in the DNA sample increases the specificity of cancer-associated hypermethylation signals, enabling improved and more accurate detection of differences between normal and cancerous samples. In addition, since the copy number of such alignments reflects the nucleosome boundary, analysis of such alignments is advantageous for evaluating nucleosome arrangement in cell-free DNA, in addition to methylation, where high copy numbers are typical in the center of the nucleosome and low copy numbers are typical at the boundaries between nucleosomes.
[0025] The present invention further discloses a method for directly calculating both methylated and unmethylated levels of DNA based on sequencing data generated after methylation-sensitive / independent restriction of a DNA sample. Advantageously, the method and system of the present invention enable the independent determination of methylated and unmethylated levels of DNA based on the same sequencing data in a single assay, thus providing improved identification of methylation changes. More specifically, the method and system disclosed herein, according to some embodiments, includes digestion of a DNA sample with at least one methylation-sensitive restriction endonuclease, followed by high-throughput sequencing to generate a plurality of sequence reads. The sequence reads can be aligned to a reference genome, selecting and analyzing restriction loci, i.e., restriction sites within the genome. The level of methylated DNA at the selected restriction loci is determined based on the read count for each restriction locus, where the read count represents the number of DNA molecules in the sample that were methylated and therefore remained intact at the restriction locus. The level of unmethylated DNA at the selected restriction loci is determined by a specific analysis of the ends of the sequence reads by determining the number of reads that begin or end with nucleotides within each restriction locus. This reading count represents the number of DNA molecules in the sample whose restriction loci were not methylated and therefore cleaved by restriction endonucleases. Direct analysis of such unmethylated DNA molecules is advantageous over indirect assessment based on methylated DNA levels, as performed by existing methods. Direct determination of unmethylation in addition to methylation using the same sequencing data provides complementary methylation information for genomic regions, thus offering improved methylation profiling, more accurate and effective assessment of potential DNA methylation markers, and better detection of methylation differences between samples. It also provides increased sensitivity for methylation analysis, which is particularly beneficial for genomic regions with very high or very low methylation levels.
[0026] According to one aspect, the present invention provides a method for profiling genetic and epigenetic characteristics of a cell-free DNA (cfDNA) sample derived from a subject, the method comprising: (a) subjecting a cell-free DNA sample to digestion with at least one methylation-sensitive restriction endonuclease to obtain restriction endonuclease-treated DNA in which methylated restriction sites remain intact and unmethylated restriction sites are cleaved; (b) preparing a sequencing library from the restriction endonuclease-treated DNA while preserving sequence information at the ends of DNA molecules, wherein preparing the sequencing library comprises ligating sequencing adapters to DNA molecules in the restriction endonuclease-treated DNA, and each adapter can be ligated to both digested DNA molecules and undigested DNA molecules; (c) sequencing the sequencing library by high-throughput sequencing to provide sequencing data; (d) determining, from the sequencing data, a methylation value for at least one restriction locus, and optionally, at least one additional genetic or epigenetic characteristic of the cell-free DNA sample selected from the group consisting of DNA mutations, copy number variations and nucleosome positioning; an amount of cell-free DNA comprising 3000 haploid equivalents is sufficient for the method, the cell-free DNA sample is not subjected to amplification prior to library preparation, and determining the methylation value and the at least one additional genetic or epigenetic characteristic of the cell-free DNA sample is performed based on the same sequencing data.
[0027] According to another aspect, the present invention provides a method for processing a cell-free DNA sample to obtain sequencing data for genetic analysis and epigenetic analysis, the method comprising: (a) subjecting a cell-free DNA sample to digestion with at least one methylation-sensitive restriction endonuclease to obtain restriction endonuclease-treated DNA in which methylated restriction sites remain intact and unmethylated restriction sites are cleaved; (b) preparing a sequencing library from restriction endonuclease-treated DNA while preserving sequence information at the ends of DNA molecules, wherein preparing the sequencing library comprises ligating sequencing adapters to DNA molecules in the restriction endonuclease-treated DNA, each adapter being capable of ligating to both digested DNA molecules and undigested DNA molecules; (c) sequencing the sequencing library by high-throughput sequencing to obtain sequencing data, wherein the method comprises: the amount of cell-free DNA comprising 3000 haploid equivalents is sufficient to achieve at least one of: a unique mapping rate of at least 85%, copy number completeness characterized by a Pearson correlation of at least 0.65 compared to an undigested sample, and nucleosome positioning completeness characterized by a Pearson correlation of at least 0.55 compared to an undigested sample, genetic and epigenetic analysis is performed based on the same sequencing data.
[0028] In some embodiments, the amount of cell-free DNA comprising 6,000 haploid equivalents is sufficient for the methods disclosed herein.
[0029] In some embodiments, the cell-free DNA is plasma cell-free DNA, and the amount of cell-free DNA is the amount obtained from 9 to 10 mL of blood.
[0030] In some embodiments, the amount of cell-free DNA is 10 to 200 ng. In additional embodiments, the amount of cell-free DNA is 20 to 100 ng.
[0031] In some embodiments, the at least one methylation-sensitive restriction endonuclease generates non-blunt ends, and the method further comprises subjecting the restriction endonuclease-treated DNA to end repair prior to ligation for sequencing to obtain DNA molecules having blunt ends.
[0032] In some embodiments, high-throughput sequencing is whole-genome high-throughput sequencing.
[0033] In some embodiments, high-throughput sequencing is high-throughput sequencing that targets only the target.
[0034] In some embodiments, determining the methylation level of at least one restriction locus is possible. (i) Select at least one restriction locus and determine the number of sequence reads that cover a given genomic region of at least 50 bp in length that contains the restriction locus, (ii) Comparing the methylation value of at least one restriction locus based on the read count and reference read count determined in step (i),
[0035] In some embodiments, step (i) includes determining the number of sequence reads that cover a given genomic region of at least 100 bp in length that contains a restriction locus.
[0036] In some embodiments, at least one restriction locus is multiple restriction loci.
[0037] In some embodiments, at least one methylation-sensitive limiting endonuclease is multiple methylation-sensitive limiting endonucleases, and digestion by multiple methylation-sensitive limiting endonucleases is simultaneous digestion.
[0038] In some embodiments, the multiple methylation-sensitive limiting endonucleases include HinP1I. In additional embodiments, the multiple methylation-sensitive limiting endonucleases include AciI. In additional embodiments, digestion is carried out using HinP1I and AciI. In some embodiments, digestion is carried out using HinP1I and AcII in a ratio of 1:1 to 5:1 (enzyme units) (HinP:AciI).
[0039] In some embodiments, subjecting a cell-free DNA sample to digestion with at least one methylation-sensitive restriction endonuclease further includes determining the digestive efficacy and proceeding to the preparation of a sequencing library if the digestive efficacy exceeds a predetermined threshold.
[0040] In another aspect, the present invention provides a method for detecting cancer-related genetic and epigenetic alterations in a cell-free DNA sample (cfDNA) derived from a subject, the method comprising: obtaining a genetic and epigenetic profile of the cfDNA sample by profiling the methylation of the cfDNA sample and optionally at least one additional genetic and epigenetic characteristic as disclosed herein; and detecting cancer-related genetic and epigenetic alterations in the cfDNA sample by comparing the genetic and epigenetic profile of the cfDNA sample with one or more reference genetic and epigenetic profiles selected from cancer profiles and non-cancer profiles.
[0041] In some embodiments, the cell-free DNA sample is from a subject suspected of having cancer or at risk of having cancer, and the method further includes subjecting the subject to active cancer surveillance and follow-up examinations if cancer-related changes are detected, the cancer surveillance and follow-up examinations include one or more of the following: blood tests, urine tests, cytology, imaging, endoscopy, and biopsy.
[0042] In an additional aspect, the present invention provides a method for evaluating the presence or absence of cancer in a subject, the method being: (a) The target cell-free DNA (cfDNA) sample is subjected to digestion with at least one methylation-sensitive restriction endonuclease to obtain restriction endonuclease-treated DNA in which methylation restriction sites are intact and unmethylation restriction sites are cleaved. (b) Restriction endonuclease-treated DNA is sequenced by a high-throughput sequencing method, (c) Select at least one multi-omics genomic region that includes tumor hypermethylation restriction loci and tumor mutation loci that are 150 bp apart from each other, (d) Determining the likelihood that a subject has cancer based on an analysis of a sequence reading covering at least one multi-omics region.
[0043] In some embodiments, at least one multi-omics region includes a tumor hypermethylation restriction locus and a tumor mutation locus that are within 100 bp of each other.
[0044] In some embodiments, the analysis of a sequence read covering at least one multi-omics region is performed. -For each multi-omics domain, (i) The number of methylation mutation sequence reads that cover a multi-omics region, including all nucleotides of the restriction locus and presenting a mutant genotype at the mutant locus, (ii) The number of methylated wild-type sequence readouts covering a multi-omics region that includes all nucleotides of the restriction locus and presents the wild-type genotype at the mutant locus, (iii) the number of unmethylated mutant sequence reads covering a multi-omics region that begin or end with a nucleotide within a restriction locus and present a mutant genotype at a mutant locus, and (iv) Determine at least one of the number of unmethylated wild-type sequence reads covering a multi-omics region that begin or end with a nucleotide in a restriction locus and present the wild-type genotype at the mutant locus, - This includes comparing the number of readings determined in (i) to (iv) with the reference values for cancer patients and / or healthy individuals in order to assess the likelihood that the subject has cancer.
[0045] In an additional aspect, the present invention provides a method for characterizing cell-free DNA (cfDNA) samples from subjects suspected of having cancer or at risk of having cancer, the method being: (a) To obtain restriction endonuclease-treated DNA in which the methylation restriction sites are intact and the unmethylated sites are cleaved, by subjecting a cell-free DNA sample to digestion with at least one methylation-sensitive restriction endonuclease, (b) Restriction endonuclease-treated DNA is sequenced by a high-throughput sequencing method, (c) Select at least one multi-omics genomic region that includes tumor hypermethylation restriction loci and tumor mutation loci that are 150 bp apart from each other, (d) For each multi-omics domain, (i) The number of methylation mutation sequence reads that cover a multi-omics region, including all nucleotides of the restriction locus and presenting a mutant genotype at the mutant locus, (ii) The number of methylated wild-type sequence readouts covering a multi-omics region that includes all nucleotides of the restriction locus and presents the wild-type genotype at the mutant locus, (iii) the number of unmethylated mutant sequence reads covering a multi-omics region that begin or end with a nucleotide within a restriction locus and present a mutant genotype at a mutant locus, and (iv) Determining at least one of the number of unmethylated wild-type sequence reads covering a multi-omics region that begin or end with a nucleotide in a restriction locus and present the wild-type genotype at the mutant locus, This allows for the characterization of cell-free DNA samples.
[0046] In another aspect, the present invention provides a method for profiling the methylation of a DNA sample from a subject, the method being: (a) To obtain restriction endonuclease-treated DNA in which the methylation restriction sites are intact and the unmethylated sites are cleaved, by subjecting the DNA sample to digestion with at least one methylation-sensitive restriction endonuclease. (b) Preparing a sequencing library from restriction endonuclease-treated DNA, wherein the preparation of the sequencing library includes ligating sequencing adapters to restriction endonuclease-treated DNA fragments, each adapter capable of ligating to both digested and undigested DNA molecules. (c) To obtain sequence readings by sequencing a sequencing library using a high-throughput sequencing method, (d) Selecting at least one restriction locus and determining the number of sequence reads to cover a given genomic region of at least 50 bp in length that contains the restriction locus, (e) calculating the methylation value of at least one restriction locus based on the read count and reference read count determined in step (d), This allows for profiling of methylation in cell-free DNA samples.
[0047] In some embodiments, a predetermined region covering a restriction locus begins at least 25 bp upstream of the cleavage site within the restriction locus and ends at least 25 bp downstream of the cleavage site within the restriction locus.
[0048] In some embodiments, step (d) includes determining a number of sequence reads that cover a given genomic region of at least 100 bp in length containing a restriction locus. In some embodiments, the given region covering the restriction locus begins at least 50 bp upstream of the cleavage site within the restriction locus and ends at least 50 bp downstream of the cleavage site within the restriction locus.
[0049] In some embodiments, at least one restriction locus is located within the CG island.
[0050] In some embodiments, the reference read count is a read count determined for a predetermined genomic region of at least 50 bp in length containing restriction loci in an undigested control DNA sample, and is optionally corrected for sequencing depth differences.
[0051] In some embodiments, the reference read count is a read count determined using a reference region of at least 50 bp in length that contains a reference locus that is not cleaved by a restriction endonuclease.
[0052] In some embodiments, the reference read count is an average read count determined using multiple reference regions of at least 50 bp in length that contain reference loci that are not cleaved by restriction endonucleases.
[0053] In some embodiments, calculating the methylation value includes normalizing the read count determined in step (d) with respect to the median read count of the DNA sample to obtain a normalized read count, and calculating the ratio of the normalized read count to the normalized reference read count.
[0054] In another aspect, the present invention provides a method for genetic and epigenetic profiling of a DNA sample, comprising determining the methylation value of at least one restriction locus disclosed herein, and further determining at least one additional genetic or epigenetic characteristic of the DNA sample selected from sequencing data from DNA mutations, copy number variations, and nucleosome positioning.
[0055] In some embodiments, the DNA is cell-free DNA extracted from a biological fluid sample. In additional embodiments, the DNA is DNA extracted from a tumor sample.
[0056] In an additional aspect, the present invention provides a method for identifying a genomic region differentially methylated between a first source and a second source of DNA, the method being: To obtain a first DNA methylation profile by profiling the methylation of at least one DNA sample from a first source disclosed herein, To obtain a second DNA methylation profile by profiling the methylation of at least one DNA sample from a second source disclosed herein, This includes comparing the first and second DNA methylation profiles to identify genomic regions that are differentially methylated between the first and second DNA sources.
[0057] In some embodiments, the first DNA source is cancer DNA, and the second DNA source is non-cancerous DNA. In some embodiments, the first DNA source is cell-free plasma DNA from a cancer patient, and the second DNA source is cell-free plasma DNA from one or more healthy individuals. In additional embodiments, the first and second DNA sources are from different stages of cancer.
[0058] In a further embodiment, the present invention provides a method for profiling the methylation of a DNA sample from a subject, the method being: (a) To provide a DNA sample derived from the target, (b) Digesting the DNA sample with at least one methylation-sensitive restriction endonuclease to obtain restriction endonuclease-treated DNA containing DNA fragments generated by the restriction endonuclease, (c) Perform high-throughput sequencing of endonuclease-treated DNA to obtain sequence readings, (d) Determining the read count of at least one restriction locus from the sequence reading, wherein the read count represents the number of DNA molecules in the DNA sample in which at least one restriction locus was methylated and therefore remained intact. (e) Determining from the sequence reads the read count of sequence reads ending in a nucleotide within at least one restriction locus, wherein the read count represents the number of DNA molecules in the DNA sample where at least one restriction locus is not methylated and therefore cleaved by restriction endonucleases. (f) Calculate the level of methylated DNA at at least one restriction locus based on the read count of at least one restriction locus determined in step (d), and calculate the level of unmethylated DNA at at least one restriction locus based on the read count of sequence reads that begin or end with a nucleotide in at least one restriction locus determined in step (e), This includes profiling the methylation of DNA samples.
[0059] In some embodiments, steps (c) to (e) are: Using a sequencing adapter ligated to multiple restriction endonuclease-generating DNA fragments, a sequencing library is prepared from restriction endonuclease-treated DNA, and the sequencing library is subjected to high-throughput sequencing to obtain sequence readings. This involves mapping multiple sequence reads to a reference genome to generate mapped sequence reads, and selecting at least one restriction locus within the reference genome. Determining the read count of at least one restriction locus from mapped sequence reads, wherein the read count represents the number of DNA molecules in the DNA sample in which at least one restriction locus was methylated and therefore remained intact. The method comprises determining, from mapped sequence reads, read counts for sequence reads that begin or end with a nucleotide in at least one restriction locus, wherein the read counts represent the number of DNA molecules in the DNA sample where at least one restriction locus is not methylated and therefore cleaved by a restriction endonuclease.
[0060] In some embodiments, high-throughput sequencing is whole-genome high-throughput sequencing. In other embodiments, high-throughput sequencing is high-throughput sequencing targeting only a specific target.
[0061] In some embodiments, the reference genome is the complete human genome.
[0062] In some embodiments, the DNA is cell-free DNA extracted from a biological fluid sample. In some embodiments, the biological fluid sample is plasma, serum, or urine. Each possible biological sample is a separate embodiment of the present invention.
[0063] In some embodiments, the DNA is DNA extracted from a tumor sample.
[0064] In some embodiments, calculating the level of methylated DNA at at least one restriction locus involves calculating the ratio of the read count of at least one restriction locus determined in step (d) to the expected read count of at least one restriction locus.
[0065] In some embodiments, calculating the level of unmethylated DNA at at least one restriction locus involves calculating the difference between the read count of sequence reads beginning or ending with a nucleotide in at least one restriction locus determined in step (e) and the expected read count of sequences beginning or ending with a nucleotide in at least one restriction locus, and then dividing the difference by the expected read count of at least one restriction locus.
[0066] In some embodiments, calculating the level of methylated DNA at at least one restriction locus is possible. The total number of fragments is determined by summing the read counts of at least one restriction locus determined in step (d) and the read counts of sequence reads that begin or end with a nucleotide in at least one restriction locus determined in step (e), and then subtracting the expected read counts of sequences that begin or end with a nucleotide in a restriction locus from the sum. This includes dividing the read count of at least one restriction locus determined in step (d) by the total number of fragments.
[0067] In some embodiments, calculating the level of unmethylated DNA at at least one restriction locus is possible. To determine the total number of fragments described herein, Calculate the difference between the read count of sequence reads that begin or end with a nucleotide in at least one restriction locus determined in step (e) and the expected read count of sequences that begin or end with a nucleotide in the restriction locus, This includes dividing the difference by the total number of fragments.
[0068] In some embodiments, the expected read count is a read count determined using a reference locus of the same length as at least one restriction locus that is not cleaved by a restriction endonuclease.
[0069] In some embodiments, the expected read count is an average read count determined using multiple reference loci of the same length as at least one restriction locus that are not cleaved by restriction endonucleases.
[0070] In some embodiments, the expected read count is the read count determined for at least one restriction locus in the undigested control DNA sample, and is optionally corrected for differences in sequencing depth.
[0071] In some embodiments, at least one restriction locus is multiple restriction loci.
[0072] In some embodiments, at least one methylation-sensitive limiting endonuclease is multiple methylation-sensitive limiting endonucleases.
[0073] In some embodiments, a method for profiling methylation further includes the step of identifying the presence or absence of disease in a subject based on the methylation profile of a DNA sample by comparing the methylation profile of the DNA sample with one or more reference methylation profiles.
[0074] In some embodiments, the method further includes generating a report in paper or electronic format based on the methylation profile and communicating the report to the subject and / or the subject's healthcare provider.
[0075] In another aspect, the present invention provides a method for detecting methylation changes in a DNA sample, the method comprising profiling the methylation of the DNA sample as disclosed herein to obtain a methylation profile of the DNA sample, and detecting methylation changes in the DNA sample by comparing the methylation profile of the DNA sample with one or more reference methylation profiles.
[0076] In some embodiments, one or more reference methylation profiles include healthy DNA methylation profiles. In additional embodiments, one or more reference methylation profiles include disease DNA methylation profiles. In some embodiments, the DNA sample originates from a subject suspected of having a disease and / or a subject at risk of developing a disease, and the step of detecting methylation changes includes determining whether the DNA sample is a healthy DNA sample or a disease DNA sample. In some embodiments, the disease is cancer.
[0077] In an additional aspect, the present invention provides a method for identifying a genomic region differentially methylated between a first source and a second source of DNA, the method being: To obtain a first DNA methylation profile by profiling the methylation of at least one DNA sample from a first source according to the method disclosed herein, Obtaining a second DNA methylation profile by profiling the methylation of at least one DNA sample from a second source according to the method disclosed herein, This includes comparing the first and second DNA methylation profiles to identify genomic regions that are differentially methylated between the first DNA source and the second DNA source.
[0078] In some embodiments, the first DNA source is disease DNA, and the second DNA source is non-disease DNA. In additional embodiments, the first and second DNA sources are from different stages of the disease. In some embodiments, the disease is cancer.
[0079] In an additional aspect, the present invention provides a method for profiling genetic and epigenetic characteristics in a DNA sample, the method being: To profile the methylation of DNA samples disclosed herein, Determining at least one additional genetic or epigenetic characteristic of a DNA sample, wherein the at least one additional genetic or epigenetic characteristic is selected from DNA mutations, copy number variations, and nucleosome positioning, Methylation profiling and determination of at least one additional genetic or epigenetic trait were performed using the same sequencing data. This includes profiling and determining the genetic and epigenetic characteristics of a DNA sample. These and additional aspects and features of the present invention will become apparent.
[0080] The following detailed description, examples, and claims are derived from the above. [Brief explanation of the drawing]
[0081] [Figure 1A] Copy number data for pooled plasma cell-free DNA samples that were subjected to methylation-sensitive digestion, bisulfite conversion, or not before sequencing. Data are presented as hit counts per genomic location. [Figure 1B] Correlation of hits between test (treated) pooled plasma cell-free DNA samples and control (untreated) pooled plasma cell-free DNA samples. [Figure 2A] Nucleosome positioning of pooled plasma cell-free DNA samples subjected to methylation-sensitive digestion, bisulfite conversion, or no treatment prior to sequencing. Data are presented as "hit span 100" (= number of reads starting >50 bp upstream and ending >50 bp downstream of the analyzed genomic location). [Figure 2B] Correlation of "hit span 100" between test (treated) pooled plasma cell-free DNA samples and control (untreated) pooled plasma cell-free DNA samples. [Figure 3] Copy number integrity of plasma cell-free DNA from patient BMD LNG165(3A) and patient BMD LNG166(3B) subjected to methylation-sensitive digestion or bisulfite conversion before sequencing. [Figure 4] Nucleosome positioning integrity of plasma cell-free DNA from patient BMD LNG165(4A) and patient BMD LNG166(4B) subjected to methylation-sensitive digestion or bisulfite conversion prior to sequencing. Data are presented as "hit span 100" (= number of reads starting >50 bp upstream and ending >50 bp downstream of the analyzed genomic location). [Figure 5] CG depth of cell-free plasma DNA from patient BMD LNG165(5A) and patient BMD LNG166(5B) subjected to methylation-sensitive digestion or bisulfite conversion before sequencing. [Figure 6]Detection of hypermethylated marker gene loci in patient BMD LNG165 and patient BMD LNG166 cell-free plasma DNA using DNA methylation-sensitive digestion or bisulfite conversion. [Figure 7] Detection of tumor mutations in plasma cell-free DNA of patient BMD LNG165 compared to control (7A) and in plasma cell-free DNA of patient BMD LNG166 compared to control (7B) using DNA methylation-sensitive digestion or bisulfite conversion. [Figure 8] Sample preparation for genetic and epigenetic profiling. (8A) Lung cancer sample; (8B) Control sample. [Figure 9] Clinical and methylation data for patient BMD LNG165(9A) and BMD LNG166(9B). [Figure 10] Methylation loci with methylation signals in patient BMD LNG165(10A) and BMD LNG166(10B). [Figure 11] Patient BMD LNG165(11A) and BMD LNG166(11B) mutation data. [Figure 12] Multi-omics region within patient BMD LNG165. [Figure 13] Types of multi-omics alignment. [Figure 14] Diagram illustrating the pre- and post-digestion and terminal repair of methylation-sensitive HinP1I sites. [Figure 15] A diagram showing DNA fragments obtained after digestion and end repair of DNA molecules spanning HinP1I restriction sites, where methylation or demethylation occurs at the cleavage site. [Figure 16] Analysis of sequence readings according to embodiments of the present invention. (16A) Restriction locus reading count, (16B) Sequence reading count of sequences beginning with a nucleotide within a restriction locus, (16C) Sequence reading count of sequences ending with a nucleotide within a restriction locus. [Figure 17]Analysis of sequence readings according to embodiments of the present invention for exemplary loci CG#1 (Figure 17A), exemplary loci CG#4 (Figure 17B), and exemplary loci CG#5 (Figure 17C). [Figure 18] A flowchart illustrating an exemplary method for profiling methylation of a DNA sample in a lung cancer-related genomic region according to embodiments of the present invention. [Figure 19] A flowchart illustrating additional exemplary methods for profiling methylation of DNA samples in lung cancer-related genomic regions according to embodiments of the present invention. [Figure 20] A flowchart illustrating an exemplary method for determining whether a DNA sample is positive or negative for lung cancer, according to embodiments of the present invention. [Figure 21] A flowchart illustrating an additional exemplary method for determining whether a DNA sample is positive or negative for lung cancer, according to embodiments of the present invention. [Modes for carrying out the invention]
[0082] This invention relates to a method and system for profiling the genetic and epigenetic properties of DNA samples, particularly cell-free DNA samples, using methylation-sensitive / methylation-dependent restriction enzyme digestion of DNA, followed by high-throughput sequencing and sequence reading analysis. Advantageously, the method and system of the present invention are highly sensitive yet accurate, allow working with very small amounts of DNA, and provide a vast amount of information, including methylation data, mutation data, etc., based on sequencing data from a single run.
[0083] Notably, even when very small amounts of DNA may be used, the quality of the sequencing data, and therefore the genetic and epigenetic information that can be derived from it, is very high, enabling highly sensitive and comprehensive identification of cancer-related changes.
[0084] The methods disclosed herein, according to some embodiments, require preserving sequence information at the 5' and / or 3' ends of a DNA molecule, including natural ends (e.g., for evaluating nucleosome positioning of cell-free DNA) and ends generated after digestion with restriction enzymes disclosed herein (e.g., for analyzing DNA molecules that were not methylated in a DNA sample). According to the present invention, preserving sequence information at the ends of a DNA molecule, or "end preservation," encompasses avoiding PCR for enriching a genomic region of interest and / or introducing a sequencing adapter. In some specific embodiments, end preservation according to the present invention involves preserving sequence information at the ends of a DNA molecule relating to the methylation status of the DNA molecule.
[0085] In some embodiments, the library preparation according to the present invention is carried out in a terminal preservation manner, which means that the library preparation process does not involve PCR for enriching the genomic region of interest and / or introducing sequencing adapters. According to these embodiments, the library preparation includes adding sequencing adapters via ligation (e.g., enzymatic ligation). If enrichment of a specific genomic region is desired, the library preparation according to these embodiments includes enriching the genomic region of interest using a capture agent.
[0086] The method of the present invention does not require the use of restriction enzyme isoschisomers (where one enzyme recognizes both the methylated and unmethylated forms of the restriction site, while the other recognizes only the unmethylated form), or requires the use of a combination of a methylation-sensitive restriction enzyme and a methylation-insensitive restriction enzyme.
[0087] Furthermore, the method of the present invention does not require or use size selection of DNA fragments within a specific size range after digestion, or filtering of read counts having a specific size range after sequencing.
[0088] In some embodiments, the present invention provides an improved method for determining the methylation value of a target genomic locus. The improved method is based on determining the read count of sequence reads covering a given genomic region of at least 50 bp in length, preferably at least 100 bp in length, which includes the target restriction locus.
[0089] In some embodiments, the present invention relates to systems and methods for high-resolution DNA methylation profiling. In some embodiments, the present invention provides the use of methylation-sensitive / methylation-dependent restriction enzymes and high-throughput sequencing in the analysis of DNA methylation. In some specific embodiments, the present invention provides the use of methylation-sensitive / methylation-dependent restriction enzymes and high-throughput sequencing for directly calculating methylated and unmethylated DNA levels.
[0090] Methylation in the human genome occurs in the form of 5-methylcytosine and is limited to cytosine residues that are part of CG sequences, also known as CpG dinucleotides (cytosine residues that are part of other sequences are not methylated). Some CG dinucleotides in the human genome are methylated, while others are not. Furthermore, because methylation is cell and tissue specific, a particular CG dinucleotide may be methylated in certain cells but not in others, or methylated in certain tissues but not in others. DNA methylation is an important regulator of gene transcription.
[0091] The methylation patterns of cancer DNA differ from those of normal DNA, with some loci being hypermethylated and others hypomethylated. In some embodiments, the present invention provides methods and systems for highly sensitive detection of differentially methylated (e.g., hypermethylated) genomic loci associated with cancer.
[0092] In some embodiments, the present invention provides a method for profiling the genetic and epigenetic properties of cell-free DNA (cfDNA) samples derived from a subject, the method being: (a) To obtain restriction endonuclease-treated DNA in which the methylation restriction sites are intact and the unmethylated sites are cleaved, by subjecting a cell-free DNA sample to digestion with at least one methylation-sensitive restriction endonuclease, (b) Preparing a sequencing library from restriction endonuclease-treated DNA while preserving the sequence information of the ends of the DNA molecules, wherein the preparation of the sequencing library includes ligating sequencing adapters to DNA molecules in the restriction endonuclease-treated DNA, and each adapter can ligate both digested and undigested DNA molecules. (c) To sequence a sequencing library using a high-throughput sequencing method and provide sequencing data, (d) Determining the methylation value of at least one restriction locus from sequencing data, and optionally, determining at least one additional genetic or epigenetic characteristic of a cell-free DNA sample selected from DNA mutations, copy number variations, and nucleosome positioning, The amount of cell-free DNA containing 3000 haploid equivalents is sufficient for the method, the cell-free DNA sample is not subjected to amplification before library preparation, and the methylation level and at least one additional genetic or epigenetic characteristic of the cell-free DNA sample are determined based on the same sequencing data.
[0093] In some embodiments, a method for processing cell-free DNA samples to obtain sequencing data for genetic and epigenetic analysis, (a) To obtain restriction endonuclease-treated DNA in which the methylation restriction sites are intact and the unmethylated sites are cleaved, by subjecting a cell-free DNA sample to digestion with at least one methylation-sensitive restriction endonuclease, (b) Preparing a sequencing library from restriction endonuclease-treated DNA while preserving the sequence information of the ends of the DNA molecules, wherein the preparation of the sequencing library includes ligating sequencing adapters to DNA molecules in the restriction endonuclease-treated DNA, and each adapter can ligate both digested and undigested DNA molecules. (c) The process includes sequencing a sequencing library using a high-throughput sequencing method to obtain sequencing data, The amount of cell-free DNA containing 3000 haploid equivalents is sufficient to achieve at least one of the following: a unique mapping rate of at least 85%, copy number completeness characterized by at least 0.65 Pearson correlations compared to the undigested sample, and nucleosome positioning completeness characterized by at least 0.55 Pearson correlations compared to the undigested sample. Genetic and epigenetic analyses are performed based on the same sequencing data.
[0094] As used herein, 3.3 pg of DNA corresponds to one haploid equivalent.
[0095] In some embodiments, 10 ng of DNA is sufficient for the methods disclosed herein. In additional embodiments, 20 ng of DNA is sufficient for the methods disclosed herein. In additional embodiments, the methods disclosed herein are carried out using initial amounts of DNA ranging from 10 to 200 ng, for example, 20 to 200 ng, 20 to 100 ng (including each value within the range). Each possibility represents a separate embodiment.
[0096] In some embodiments, 3,000 haploid equivalents are sufficient for the methods disclosed herein. In additional embodiments, 6,000 haploid equivalents are sufficient for the methods disclosed herein. In additional embodiments, the methods disclosed herein are carried out using an initial amount of DNA containing 3,000 to 60,000 haploid equivalents, for example, haploid equivalents between 6,000 and 60,000, and haploid equivalents between 6,000 and 30,000 (including each value within the range). Each possibility represents a separate embodiment.
[0097] In some embodiments, the amount of cell-free DNA disclosed herein is sufficient to achieve intrinsic mapping rates of at least 85%, at least 86%, at least 87%, at least 88%, and at least 89%. Each possibility represents a separate embodiment.
[0098] In some embodiments, the amount of cell-free DNA disclosed herein is sufficient to achieve copy number completeness characterized by Pearson correlations of at least 0.6, for example, at least 0.65, at least 0.66, at least 0.67, at least 0.68, and at least 0.69 compared to the undigested sample. Each possibility represents a separate embodiment.
[0099] In some embodiments, the amount of cell-free DNA disclosed herein is sufficient to achieve nucleosome positioning integrity characterized by Pearson correlations of at least 0.55, for example, at least 0.56, at least 0.57, at least 0.58, and at least 0.59 compared to the undigested sample. Each possibility represents a separate embodiment.
[0100] In some embodiments, the present invention provides a method for profiling the methylation of a DNA sample derived from a subject, the method being: (a) To obtain restriction endonuclease-treated DNA in which the methylation restriction sites are intact and the unmethylated sites are cleaved, by subjecting the DNA sample to digestion with at least one methylation-sensitive restriction endonuclease. (b) Preparing a sequencing library from restriction endonuclease-treated DNA, (c) To obtain sequence readings by sequencing a sequencing library using a high-throughput sequencing method, (d) Selecting at least one restriction locus and determining the number of sequence reads to cover a given genomic region of at least 50 bp in length that contains the restriction locus, (e) Determining the methylation value of at least one restriction locus based on the read count and reference read count determined in step (d), This allows for profiling of methylation in cell-free DNA samples.
[0101] In some embodiments, profiling the methylation of a DNA sample involves determining the number of sequence reads covering a given genomic region of at least 60 bp in length containing a restriction locus, for example, a given genomic region of at least 70 bp, at least 80 bp, at least 90 bp, at least 100 bp, 50–150 bp, 50–120 bp, or 50–100 bp containing a restriction locus. Each possibility represents a separate embodiment.
[0102] In some embodiments, at least one restriction locus is located within a CG island. A "CG island" (or CpG island) is a region of DNA with a high G / C content and high frequency of CG dinucleotides relative to the entire genome of the organism of interest. CG islands are typically 200–3,000 bp in length and are typically characterized by a GC content of over 50% and observed to be a predicted CG ratio greater than 0.6. Genomic regions with lower CG density are called "CG oceans" and comprise the majority of the genome.
[0103] In some embodiments, a method is provided for profiling the methylation of a DNA sample derived from a subject, the method comprising (i) digesting the DNA sample derived from the subject with at least one methylation-sensitive / dependent restriction endonuclease, and (ii) sequencing the digested DNA by a high-throughput sequencing method, the method enabling independent determination of the methylation and demethylation levels of the DNA based on the same sequencing data in a single assay.
[0104] In some embodiments, a method is provided for profiling the methylation of a DNA sample derived from a target, and the method is (a) To provide a DNA sample derived from the target, (b) Digesting the DNA sample with at least one methylation-sensitive restriction endonuclease to obtain restriction endonuclease-treated DNA containing DNA fragments generated by the restriction endonuclease, (c) Perform high-throughput sequencing of endonuclease-treated DNA to obtain multiple sequence readings, (d) Determining the read count of at least one restriction locus from the sequence reading, wherein the read count represents the number of DNA molecules in the DNA sample in which at least one restriction locus was methylated and therefore remained intact. (e) Determining from sequence reads the read count of sequence reads that begin or end with a nucleotide in at least one restriction locus, wherein the read count represents the number of DNA molecules in the DNA sample that are not methylated at least one restriction locus and are therefore cleaved by restriction endonucleases. (f) profiling the methylation of the DNA sample by calculating the level of methylated DNA at at least one restriction locus based on the read count of at least one restriction locus determined in step (d), and by calculating the level of unmethylated DNA at at least one restriction locus based on the read count of sequence reads that begin or end with a nucleotide in at least one restriction locus determined in step (e).
[0105] In additional embodiments, a method is provided for profiling the methylation of a DNA sample derived from a subject, the method being: (A) To provide DNA samples derived from the target, (B) Digesting the DNA sample with at least one methylation-sensitive restriction endonuclease to obtain restriction endonuclease-treated DNA containing multiple restriction endonuclease-generating DNA fragments, (C) Preparing a sequencing library from restriction endonuclease-treated DNA, and obtaining sequence readings by subjecting the sequencing library to high-throughput sequencing, (D) Mapping multiple sequence reads to a reference genome to generate mapped sequence reads, and selecting at least one restriction locus in the reference genome, (E) Determining the read count of at least one restriction locus from the mapped sequence reads, wherein the read count represents the number of DNA molecules in the DNA sample in which at least one restriction locus was methylated and therefore remained intact. (F) Determining, from the mapped sequence reads, the read count of sequence reads that begin or end with a nucleotide in at least one restriction locus, wherein the read count represents the number of DNA molecules in the DNA sample where at least one restriction locus is not methylated and therefore cleaved by restriction endonucleases. (G) Comparing the level of methylated DNA at at least one restriction locus based on the read count of at least one restriction locus determined in step (E), and compiling the level of unmethylated DNA at at least one restriction locus based on the read count of sequence reads that begin or end with a nucleotide in at least one restriction locus determined in step (F), thereby profiling the methylation of the DNA sample.
[0106] In some embodiments, a method for profiling methylation of a DNA sample derived from a subject, comprising: subjecting the DNA sample to digestion with at least one methylation-sensitive restriction endonuclease; performing high-throughput sequencing of the digested sample; determining the read count of at least one restriction locus; and calculating the level of methylated DNA at at least one restriction locus based on the read count, wherein the improvement is: Determining the read count of sequence reads that begin or end at a nucleotide within at least one restriction locus, wherein the read count represents the number of DNA molecules in the DNA sample that are not methylated at least one restriction locus and are therefore cleaved by restriction endonucleases. Based on the read count of sequence reads that begin or end with a nucleotide in at least one restriction locus, the level of unmethylated DNA in at least one restriction locus is calculated, This includes profiling the methylation of a DNA sample using the levels of methylated and unmethylated DNA at at least one restriction locus.
[0107] As used herein, the term “multiple” means “at least two” or “two or more.”
[0108] In some embodiments, a method is provided for identifying the presence or absence of a disease in a subject, the method comprising profiling the methylation of a DNA sample derived from the subject disclosed herein, comparing the methylation profile of the DNA sample with one or more reference methylation profiles, and determining the presence or absence of a disease in the subject based on the comparison.
[0109] In some embodiments, a method is provided for identifying DNA methylation markers indicating the source of a DNA sample, comprising profiling the methylation as disclosed herein. In additional embodiments, a method is provided herein for evaluating the quality of DNA methylation markers, comprising profiling the methylation as disclosed herein. In some embodiments, the DNA methylation marker is a marker indicating the presence or absence of a disease, such as a type of cancer. In additional embodiments, the DNA methylation marker is a marker indicating the stage of a disease, such as a stage of cancer. In additional embodiments, the DNA methylation marker is a marker indicating the type of tissue (e.g., lung tissue, breast tissue, colon tissue, etc.).
[0110] In some embodiments, the use of (i) at least one methylation-sensitive restriction enzyme and / or at least one methylation-dependent restriction enzyme, and (ii) high-throughput sequencing is provided for directly determining the methylated and unmethylated DNA levels of at least one restriction locus in a DNA sample.
[0111] In some embodiments, a use is provided for profiling the methylation of a DNA sample by directly determining the methylated and unmethylated DNA levels of at least one restriction locus in the DNA sample, wherein the determination of methylated and unmethylated DNA levels is based on the same sequencing data.
[0112] Some embodiments provide a use for profiling the methylation of a DNA sample by directly determining the methylated and unmethylated DNA levels of at least one restriction locus in the DNA sample, the digestion of the DNA sample with at least one methylation-sensitive restriction enzyme and / or at least one methylation-dependent restriction enzyme, and sequence readings generated after high-throughput sequencing, wherein the determination of methylated and unmethylated DNA levels is based on the same sequencing data.
[0113] Generally, embodiments that can be carried out using methylation-sensitive restriction enzymes can instead be carried out using methylation-dependent restriction enzymes, with downstream steps adjusted accordingly. For example, in some embodiments, following high-throughput sequencing and sequence reading generation, a method for profiling methylation according to the present invention includes selecting at least one restriction locus, determining the number of sequence readings covering a given genomic region of at least 50 bp in length containing the restriction locus, and calculating a methylation value based on the read count and reference read count of the given genomic region, the calculated methylation value reflecting the number of molecules in the DNA sample that were not methylated and therefore remained intact after digestion by methylation-dependent restriction enzymes.
[0114] As another example, in some embodiments, to calculate the level of methylated DNA at a restriction locus, following high-throughput sequencing and generation of sequence reads, the method includes determining from the sequence reads a read count of sequence reads that begin or end with a nucleotide in the restriction locus, wherein the read count represents the number of DNA molecules in the DNA sample that were methylated, i.e., cleaved by a restriction endonuclease at the restriction locus, and calculating the level of methylated DNA at the restriction locus based on the determined read count of sequence reads that begin or end with a nucleotide in the restriction locus. To calculate the level of unmethylated DNA at a restriction locus, in some embodiments, the method includes determining a read count of the restriction locus from the sequence reads, wherein the read count represents the number of DNA molecules in the DNA sample that were unmethylated and therefore intact at the restriction locus, and calculating the level of unmethylated DNA at the restriction locus based on the determined read count of the restriction locus.
[0115] DNA sample DNA samples for use according to the present invention can be obtained from any biological sample from which nucleic acids can be obtained, including biological fluid samples such as blood, plasma, serum, urine, cerebrospinal fluid, semen, feces, sputum, and amniotic fluid. Each possibility represents an individual embodiment of the present invention. Biological samples also include tissue and organ samples.
[0116] The subjects according to the present invention are typically human subjects. The subjects may be suspected of having a particular disease. In some embodiments, the subjects are diagnosed with the disease of interest. In other embodiments, the subjects are healthy subjects who do not have the disease of interest. The subjects may also be at risk of developing the disease based on, for example, a history of the disease, a genetic predisposition, and / or a family history, and / or subjects that exhibit suspicious clinical signs of the disease and / or subjects suspected of having the disease based on other previous assays, for example, testing of other biomarkers. In some embodiments, the subjects are at risk of disease recurrence. In some embodiments, the subjects exhibit at least one symptom or feature of the disease. In other embodiments, the subjects are asymptomatic.
[0117] In some embodiments, the DNA sample is cell-free DNA extracted from a biological fluid sample. The term “cell-free DNA” (abbreviated as “cfDNA”) refers to DNA molecules that circulate freely in body fluids and are not contained within intact cells. The origin of cfDNA is not fully understood, but it is thought to be related to apoptosis, necrosis, and active release from cells. cfDNA is released by both normal and tumor cells. cfDNA is highly fragmented, with fragments typically ranging in length from 120 to 220 bp, and most commonly from 150 to 180 bp. As used herein, the term “cell-free DNA” should be understood to refer to DNA that is already cell-free within the body of the subject. With respect to cell-free DNA samples, “restriction endonuclease-treated DNA” should be understood to include fragments produced as a result of digestion, and naturally occurring cell-free DNA fragments, e.g., cell-free DNA fragments that do not contain the recognition sequences of the enzymes used in the assay, and cell-free DNA fragments that are fully methylated and therefore not cleaved by the enzymes, and that contain one or more recognition sequences of the enzymes.
[0118] Alternatively, the DNA sample may be DNA extracted from cells, such as DNA extracted from tissue or organ samples or blood cells. Typically, cell lysis is required to extract the DNA. The DNA may be obtained from tumor samples or healthy tissue. As used herein, “tumor sample” includes the whole or a portion of a tumor surgically removed. “Tumor sample” also includes samples taken from a tumor by biopsy, and samples taken from lesions or tissues suspected to be cancerous. Tumor samples for use according to the present invention include fresh tumor samples and frozen / preserved tumor samples.
[0119] For DNA extracted from cells, fragmentation of the DNA into fragments suitable for high-throughput sequencing can be performed before, after, or during digestion with at least one methylation-sensitive or methylation-dependent restriction endonuclease according to the present invention to simplify downstream processing and preparation of sequencing libraries. Such fragmentation can be performed, for example, using sonication, or using a restriction endonuclease that is insensitive to methylation, i.e., a restriction endonuclease that cleaves its recognition sequence regardless of its methylation status. It can also be performed using a restriction endonuclease having a recognition sequence that does not contain CG dinucleotides.
[0120] The present invention encompasses whole-genome sequencing and target-only sequencing (e.g., sequencing of CpG islands, exons, or specific loci of interest). For target-only sequencing, the genomic region of interest is enriched using a capture agent, such as a sequence-specific probe attached to beads. Typically, enrichment of the genomic region of interest is performed after methylation-sensitive / unmethylation-dependent digestion according to the present invention and after sequencing library preparation, as described in more detail below. In some embodiments, enrichment may be performed before digestion and library preparation.
[0121] Therefore, in some embodiments, the DNA sample subjected to methylation-sensitive or methylation-dependent digestion according to the present invention is an untreated DNA sample, i.e., a DNA sample extracted from a biological sample. In other embodiments, the DNA sample is a treated DNA sample, which, for example, has been concentrated and / or fragmented and reduced in size for a specific region of interest before digestion by at least one methylation-sensitive or methylation-dependent restriction endonuclease according to the present invention.
[0122] Preferably, the DNA sample on which methylation analysis is performed is substantially free of single-stranded DNA (ssDNA). As used herein, “substantially free of ssDNA” or “substantially lacking ssDNA” refers to a DNA sample in which less than 7% of the DNA is ssDNA, preferably less than 5% of the DNA is ssDNA, and more preferably less than 1% of the DNA is ssDNA (i.e., at least 99% of the DNA is double-stranded). In some embodiments, the DNA sample contains less than 0.1% ssDNA. In some embodiments, the DNA sample contains less than 0.01% ssDNA. In some embodiments, the DNA sample contains no ssDNA at all (ssDNA-free). DNA extraction to obtain a DNA sample substantially free of ssDNA is described, for example, in International Publication No. 2020 / 188561, assigned to the applicant of the present invention. An exemplary kit for extracting cell-free DNA suitable for use in the method of the present invention is the QIAamp® Circulating Nucleic Acid Kit (QIAGEN, Hilden, Germany). An exemplary kit for extracting DNA from cells is the QIAamp® Blood Mini Kit.
[0123] DNA digestion According to the present invention, following extraction (and optionally fragmentation to concentrate and / or reduce the size of the region of interest), the DNA is subjected to digestion with at least one methylation-sensitive restriction endonuclease and / or at least one methylation-dependent restriction endonuclease, preferably multiple methylation-sensitive restriction endonucleases (or multiple methylation-dependent restriction endonucleases) applied simultaneously. As used herein, “simultaneously applied restriction endonucleases” or “simultaneous digestion” means that the enzymes are present together in the reaction mixture in their active form without inactivating one enzyme before the application of another.
[0124] For example, one, two, three, four, or five methylation-sensitive or methylation-dependent restriction endonucleases can be used. Each number of endonucleases used in the assay represents a distinct embodiment of the present invention.
[0125] In some embodiments, the entire extracted DNA is used in the digestion step. In some embodiments, the DNA is not quantified before being digested. In other embodiments, the DNA is quantified before its digestion. In some embodiments, the DNA is divided into a first aliquot to be digested and a second aliquot to be maintained as an undigested control.
[0126] In this specification, “restriction enzyme” and “restriction endonuclease” as used interchangeably refer to enzymes that cleave DNA at or near a specific recognition sequence, also known as a restriction site. Restriction sites are typically 4 to 8 nucleotides long and are usually palindromic (i.e., the DNA sequence is the same in both directions).
[0127] A "methylation-sensitive" restriction endonuclease is a restriction endonuclease that cleaves its recognition sequence only if it is not methylated (methylated sites remain intact). Therefore, the degree of digestion of a DNA sample by a methylation-sensitive restriction endonuclease depends on the methylation level; the higher the methylation level, the more protected it is from cleavage, and therefore the less digested it is. A DNA sample treated with a methylation-sensitive restriction endonuclease is characterized by intact methylated sites and cleaved unmethylated sites. It should be understood that 100% digestion efficiency is not required, and therefore some unmethylated sites may remain intact. In some embodiments, the method of the present invention includes determining the digestion efficiency and proceeding to the preparation of a sequencing library if the digestion efficiency exceeds a predetermined threshold / level.
[0128] A "methylation-dependent" restriction endonuclease is one that cleaves its recognition sequence only if it is methylated (unmethylated sites remain intact). Therefore, the degree of digestion of a DNA sample by a methylation-dependent restriction endonuclease depends on the methylation level; higher methylation levels result in more extensive digestion.
[0129] The methylation-sensitive limiting endonucleases for use according to the present invention are AatII, Acc65I, AccI, Acil, ACII, Afel, Agel, Apal, ApaLI, AscI, AsiSI, Aval, AvaII, BaeI, BanI, BbeI, BceAI, BcgI, BfuCI, BglI, BmgBI, BsaAI, BsaBI, BsaHI, BsaI, BseYI, B siEI, BsiWI, BslI, BsmAI, BsmBI, BsmFI, BspDI, BsrBI, BsrFI, BssHII, BssKI, BstAPI, BstBI, BstUI, BstZ l7I, Cac8I, ClaI, DpnI, DrdI, EaeI, EagI, Eagl-HF, EciI, EcoRI, EcoRI-HF, FauI, Fnu4HI, FseI, FspI, Hae II, HgaI, HhaI, HincII, HincII, Hinfl, HinPlI, HpaI, HpaII, Hpyl66ii, Hpyl88iii, Hpy99I, HpyCH4IV, Ka sI, MluI, MmeI, MspAlI, MwoI, NaeI, NacI, NgoNIV, Nhe-HFI, NheI, NlaIV, NotI, NotI-HF, NruI, Nt.BbvCI, The group can be selected from Nt.BsmAI, Nt.CviPII, PaeR7I, PleI, PmeI, PmlI, PshAI, PspOMI, PvuI, RsaI, RsrII, SacII, Sall, SalI-HF, Sau3AI, Sau96I, ScrFI, SfiI, SfoI, SgrAI, SmaI, SnaBI, TfiI, TscI, TseI, TspMI, and ZraI. Each possibility represents an individual embodiment of the present invention. In some specific embodiments, at least one methylation-sensitive limiting endonuclease comprises HinP1I. In additional specific embodiments, at least one methylation-sensitive limiting endonuclease comprises HhaI. In yet another specific embodiment, at least one methylation-sensitive limiting endonuclease comprises AciI.
[0130] The methylation-dependent restriction endonuclease can be selected from the group consisting of McrBC, McrA, and MrrA. Each possibility represents an individual embodiment of the present invention.
[0131] In some embodiments, the DNA sample of the present invention is digested by a single methylation-sensitive restriction endonuclease. In some specific embodiments, the methylation-sensitive restriction endonuclease is HinP1I. In additional specific embodiments, the methylation-sensitive restriction endonuclease is HhaI. In additional embodiments, the DNA sample is digested by two methylation-sensitive restriction endonucleases.
[0132] In some specific embodiments, methylation-sensitive limiting endonucleases HinP1I and AciI are used.
[0133] In some embodiments, a method is provided for profiling the methylation of a DNA sample, comprising: subjecting the DNA sample to digestion with the methylation-sensitive restriction endonucleases HinP1I and AciI; and analyzing the methylation of at least one restriction locus of HinP1I and / or at least one restriction locus of AciI to profile the methylation of the DNA sample. In some embodiments, the method comprises: subjecting the DNA sample to digestion with the methylation-sensitive restriction endonucleases HinP1I and AciI; and determining the level of methylated DNA and optionally unmethylated DNA at at least one restriction locus of HinP1I and / or at least one restriction locus of AciI to profile the methylation of the DNA sample. In some embodiments, the DNA sample is cell-free DNA extracted from a biological fluid.
[0134] In some embodiments, HinP1I and AciI are used with the methods and systems of the present invention in ratios of 1:1 to 5:1 (enzyme units) (HInP:AciI), for example, 2:1, 2.5:1, 3:1, 3.5:1, 4:1, and 4.5:1 (enzyme units) (HinP:AciI). Each possibility represents a separate embodiment of the present invention. In some embodiments, HinP1I and AciI are used with the methods of the present invention in ratios of 2:1 to 4.5:1 (enzyme units) (HInP:AciI).
[0135] In some embodiments, a method is provided for detecting methylation changes in a DNA sample, comprising profiling the methylation of the DNA sample using HinP1I and AciI digestion, and comparing the methylation profile with one or more reference methylation profiles. In some embodiments, the DNA sample is cell-free DNA extracted from a biological fluid.
[0136] In some embodiments, a method is provided for profiling the methylation of a DNA sample, the method comprising: subjecting the DNA sample to digestion with methylation-sensitive restriction endonucleases HinP1I and AciI to obtain restriction endonuclease-treated DNA containing DNA fragments produced by the restriction endonucleases; performing high-throughput sequencing of the endonuclease-treated DNA to obtain multiple sequence reads; and determining from the sequence reads the level of methylated DNA and optionally unmethylated DNA at at least one restriction locus of HinP1I and / or at least one restriction locus of AciI, thereby profiling the methylation of the DNA sample. In some embodiments, the DNA sample is cell-free DNA extracted from a biological fluid.
[0137] In some embodiments, a reaction mixture is provided comprising human cell-free DNA extracted from a biological fluid and the methylation-sensitive restriction endonucleases HinP1I and AciI. The reaction mixture further comprises a buffer suitable for the activity of HinP1I and AciI. In some embodiments, HinP1I and AciI are present in the reaction mixture in ratios of 1:1 to 5:1 (enzyme units) (Hinp:AciI), for example, 2:1, 2.5:1, 3:1, 3.5:1, 4:1, and 4.5:1 (enzyme units) (Hinp:AciI). Each possibility represents a separate embodiment of the present invention. In some embodiments, the reaction mixture contains HinP1I and AciI in ratios of 2:1 to 4.5:1 (enzyme units) (Hinp:AciI).
[0138] In some embodiments, methods are provided for processing cell-free DNA samples for genetic and epigenetic analysis, the methods comprising: providing a reaction mixture disclosed herein; incubating the reaction mixture to obtain restriction endonuclease-treated cell-free DNA in which methylation restriction sites are intact and unmethylation restriction sites are cleaved; and subjecting the restriction endonuclease-treated cell-free DNA to high-throughput sequencing.
[0139] Digestion efficiency can be evaluated internally or externally for the test sample. Internal evaluation can be performed by measuring intact cleavage sites at genomic locations known to be ubiquitously unmethylated. Examples of such loci may be any site on mitochondrial DNA. External evaluation of digestion efficiency can be performed by including an unmethylated sample in the digestion step, digesting both samples in parallel, and then verifying that the unmethylated sample was actually digested (by measuring the number of intact cleavage sites). Such an unmethylated sample may be, for example, a PCR amplicon, plasmid DNA, a commercially available unmethylated DNA species, or cell line DNA known to be unmethylated at a specific genomic location. Alternatively, external evaluation of digestion efficiency can be performed in a single step by adding the unmethylated sample to the investigation sample and measuring the digestion of the unmethylated DNA sample in the same step as the investigation sample. For this purpose, all types of unmethylated DNA species described above can be used. In some embodiments, the use of small targets (e.g., PCR amplicons or plasmid DNA) is preferred.
[0140] In some embodiments, DNA digestion may be carried out until complete digestion is achieved. In some embodiments, the methylation-sensitive restriction endonuclease may be HinP1I and / or AciI, and complete digestion may be achieved after 1-2 hours of incubation with the enzyme at 37°C.
[0141] Library preparation and sequencing High-throughput sequencing (also known as next-generation sequencing) involves sequencing methods that use techniques to determine a large number (typically thousands to billions) of nucleic acid sequences in parallel. High-throughput sequencing generally involves three basic steps: library preparation, sequencing, and data analysis. Examples of high-throughput sequencing techniques include synthetic and ligation sequencing (e.g., employed by Illumina Inc., Life Technologies Inc., and Roche), nanopore sequencing, and electron detection-based methods such as Ion Torrent® technology (Life Technologies Inc.).
[0142] Library preparation for major high-throughput sequencing platforms requires ligation of specific adapter oligonucleotides to the DNA fragments to be sequenced. As disclosed herein, restriction digestion is preferably performed prior to adapter ligation to avoid possible enzymatic digestion of the adapters. Digestion of DNA by methylation-sensitive / dependent restriction endonucleases disclosed herein typically does not result in uniform blunt-end fragments. Therefore, end repair is required to ensure that each DNA molecule is free of overhangs and contains a 5' phosphate group and a 3' hydroxyl group. Typical blunting enzyme mixtures include polymerases and polynucleotide kinases, e.g., T4 DNA polymerase and T4 polynucleotide kinase (PNK). T4 DNA polymerase (in the presence of dNTPs) can fill the 5' overhang and trim the 3' overhang to the dsDNA interface to produce blunt ends. T4 PNK can then phosphorylate the 5' terminal nucleotide. For Illumina libraries, the incorporation of non-templated deoxyadenosine 5'-monophosphate (dAMP) into the 3' end of the blunted DNA fragment (a process known as dA tailing) is also required for library preparation. The dA tail prevents concatemer formation during the downstream ligation step and allows the DNA fragment to be ligated to an adapter oligonucleotide with a complementary dT overhang.
[0143] As disclosed herein, adapter oligonucleotides, also called “sequencing adapters,” are ligated to DNA fragments using end-preservation methods such as enzymatic ligation, in which a ligase enzyme covalently binds the sequencing adapter to the DNA fragment to create a complete library molecule. The sequencing adapter is ligated to the 5' and 3' ends of each DNA fragment in the sequencing library. The sequencing adapter typically contains platform-specific sequences for fragment recognition by a particular sequencer, such as sequences that enable the library fragment to be bound to a flow cell on an Illumina platform. Each sequencing instrument provider typically uses a specific set of sequences for this purpose.
[0144] Sequence adapters may also include sample indices. A "sample index," also known as a "sample barcode," is a sequence that allows multiple samples to be sequenced together (i.e., multiplexed) on the same instrument flow cell or chip. Each sample index (typically 6-10 base pairs) is unique to a particular sample library and is used during data analysis to demultiplex individual sequence reads to assign them to the correct sample. Sequence adapters may include single or dual sample indices, depending on the number of libraries being combined and the desired level of precision.
[0145] Sequencing adapters may include unique molecular identifiers (UMIs). UMIs are a type of molecular barcode that provides molecular tracking, error correction, and improved accuracy during sequencing. UMIs are typically short sequences of 5–20 nucleotides in length and are used to uniquely tag each molecule in a sample library. Because each nucleic acid in the starting material is tagged with a unique molecular barcode, bioinformatics software can filter out duplicate reads and PCR errors with a high level of accuracy, report unique reads, and eliminate errors identified before final data analysis.
[0146] In some embodiments, both the sample barcode sequence and UMI are incorporated into the nucleic acid target molecule.
[0147] The method disclosed herein does not require differential adapter tagging of digested DNA molecules versus undigested DNA molecules (i.e., differential adapter tagging of methylated DNA molecules versus unmethylated DNA molecules), and the same adapter population is used throughout the sample so that any adapter in the mixture can ligate both digested and undigested DNA.
[0148] The high-throughput sequencing according to the present invention can be performed using a variety of high-throughput sequencing instruments and platforms, including but not limited to Novaseq®, Nextseq®, and MiSeq® (Illumina), 454 Sequencing (Roche), Ion Chef® (ThermoFisher), SOLiD® (ThermoFisher), and Sequel II® (Pacific Biosciences). A sequencing adapter with an appropriate platform design is used to prepare the sequencing library.
[0149] In some embodiments, whole-genome sequencing is performed on a library prepared from endonuclease-treated DNA. The library is prepared using a sequencing adapter suitable for the sequencing platform used.
[0150] In other embodiments, a region of interest in endonuclease-treated DNA may be captured using, for example, a solution-phase or solid-phase hybridization-based process and subsequently high-throughput sequencing. The enrichment of the region of interest and subsequent high-throughput sequencing are referred to herein as “targeted high-throughput sequencing.” Targeted high-throughput sequencing includes, for example, CpG island sequencing and exome sequencing. Targeted high-throughput sequencing also includes sequencing of specific beneficial genomic regions, e.g., regions known to be differentially methylated between cancerous and non-cancerous tissues. Capture of genomic regions for targeted sequencing is typically performed after library preparation. In some embodiments, the methods disclosed herein include enriching the genomic region of interest. To preserve the ends of DNA fragments in a DNA sample (e.g., to enable analysis of sequences that begin or end at nucleotides within restriction loci), enrichment according to the present invention is typically not performed using PCR amplification of the genomic region of interest.
[0151] In some embodiments, the method for genetic and epigenetic profiling of DNA samples according to the present invention is Extracting DNA from biological samples, The extracted DNA is subjected to digestion with at least one methylation-sensitive restriction endonuclease, thereby obtaining restriction endonuclease-treated DNA. The process involves preparing a sequencing library from restriction endonuclease-treated DNA using a sequencing adapter ligated to a DNA fragment in restriction endonuclease-treated DNA, and The method involves enriching at least one (preferably more) target genomic regions from a sequencing library using a capture agent to obtain a sequencing library enriched with at least one (preferably more) target genomic regions, The sequencing library enriched with at least one (preferably more) target genomic regions is subjected to high-throughput sequencing, This includes determining the methylation value of at least one restriction locus from sequencing data, and optionally, determining at least one additional genetic or epigenetic characteristic of a cell-free DNA sample selected from DNA mutations, copy number variations, and nucleosome positioning disclosed herein.
[0152] In some embodiments, the method for profiling methylation according to the present invention is Extracting DNA from biological samples, The extracted DNA is subjected to digestion with at least one methylation-sensitive restriction endonuclease, thereby obtaining restriction endonuclease-treated DNA. The process involves preparing a sequencing library from restriction endonuclease-treated DNA using a sequencing adapter ligated to DNA fragments in DNA generated by restriction endonucleases, and The method involves enriching at least one (preferably more) target genomic regions from a sequencing library using a capture agent to obtain a sequencing library enriched with at least one (preferably more) target genomic regions, The method involves subjecting a sequencing library enriched with at least one (preferably more) target genomic regions to high-throughput sequencing to obtain sequence readings, This includes determining the levels of methylated and unmethylated DNA at at least one restriction locus within a genomic region of interest, as disclosed herein.
[0153] Sequence reading analysis In some embodiments, a “sequence read” (or simply “read”), i.e., a nucleotide sequence generated by a sequencing process, is mapped to a reference genome. As used herein, a “reference genome” refers to a previously identified genome sequence, whether partial or complete, assembled as a representative example of a species or subject. A reference genome is typically haploid and does not typically represent the genome of a single individual of a species, but rather a mosaic of genomes from several individuals. The reference genome for the methods of the present invention is typically a human reference genome. In some embodiments, the reference genome is a complete human genome, such as a human genome assembly available on the National Center for Biotechnology Information (NCBI) or the University of California, Santa Cruz (UCSC) Genome Browser website. An example of a suitable reference genome for human studies is the “hg18” genome assembly. As an alternative, a more recent GRCh38 major assembly can be used (up to patch p13).
[0154] Read mapping is the process of aligning a read on a reference genome to identify its location within the reference genome. A sequence read to be aligned is designated as “mapped.” The alignment process aims to maximize the possibility of obtaining regions of sequence identity across various sequences during alignment, and to allow for mismatches, indels, and / or clipping of several short fragments on the two ends of a read. The number of reads mapped to a specific genomic locus of interest is referred herein to as the “read count” or “copy number” of that locus. Computer software can be used to analyze sequence reads, map them to a reference genome, and quantify the number of reads.
[0155] As used herein, the terms “genomic locus” and “locus” are interchangeable and refer to a DNA sequence at a specific location on a chromosome. A “locus” may include a single location (a single nucleotide at a defined location in the genome) or a stretch of nucleotides that begin and end at defined locations in the genome. A specific location can be identified by the molecular location, i.e., by the number of start and end base pairs on the chromosome. Variants of a DNA sequence at a given genomic location are called alleles. Alleles of a locus are located at the same site on homologous chromosomes. Genomic loci include gene sequences as well as other genetic elements (such as intergeneric sequences).
[0156] In this specification, the term "restriction locus" is used to describe a genomic locus that is a restriction site for a methylation-sensitive / independent restriction endonuclease applied in the digestion step according to the present invention. Restriction loci according to the present invention may be differentially methylated between normal DNA and disease DNA, meaning that for a given disease in which the analysis is performed, e.g., a particular type of cancer, the restriction locus may have different levels of methylation between normal DNA and DNA derived from cancer cells. For example, DNA from cancer cells may have increased levels of methylation at the restriction locus compared to normal non-cancerous DNA. More specifically, the restriction locus contains a CG dinucleotide that is more methylated in cancer DNA compared to normal non-cancerous DNA. According to the present invention, the differentially methylated CG dinucleotide is located within the recognition site of at least one restriction enzyme applied in the digestion step.
[0157] In some embodiments, the restriction loci according to the present invention contain CG dinucleotides that are more methylated in the cell-free DNA of a subject with a particular type of cancer, such as plasma DNA, than in the cell-free DNA of a healthy subject. In some embodiments, plasma samples from cancer patients contain a higher proportion of methylated DNA molecules at the restriction locus compared to plasma samples from a healthy subject.
[0158] In additional embodiments, restriction loci according to the present invention include CG dinucleotides that are more methylated in DNA derived from cancerous tissue (e.g., tumor samples) than in DNA derived from non-cancerous tissue, meaning that a larger proportion of DNA molecules are methylated at this position in cancerous tissue compared to non-cancerous tissue.
[0159] Methylation-sensitive restriction enzymes cleave their recognition sequences only if they are not methylated. Methylation-dependent restriction enzymes cleave their recognition sequences only if they are methylated. Therefore, differences in methylation levels between samples result in differences in the degree of digestion, and subsequently, differences in the amount of sequence read in subsequent sequencing and quantification steps. Such differences allow for the distinction between DNA from different samples, for example, between DNA samples from subjects with cancer and DNA samples from healthy subjects.
[0160] The terms “level of methylated DNA” of a restriction locus, “methylation level” or “methylation value” are numerical values representing the number of DNA molecules in a sample that are methylated at that restriction locus (i.e., methylated with CG dinucleotides within the restriction locus) out of the total number of DNA molecules containing the restriction locus in the sample. In some embodiments, the level of methylated DNA at a restriction locus is calculated, as herein by reference, from the read count of the restriction locus after digestion with at least one methylation-sensitive restriction endonuclease. In additional embodiments, the level of methylated DNA at a restriction locus is calculated, as herein by reference, from the read count of a given genomic region of at least 50 bp containing the restriction locus. Since methylation-sensitive restriction endonucleases cleave recognition sequences only if their recognition sequences are not methylated, the read count of a restriction locus represents the number of DNA molecules in the DNA sample where the restriction locus was methylated and therefore remained intact.
[0161] In some embodiments, the methylation level of a restriction locus is calculated by dividing the read count of the restriction locus, or the read count of a given genomic region of at least 50 bp containing the restriction locus, by the expected read count of the restriction locus or the given genomic region of at least 50 bp containing the restriction locus. The expected read count of the restriction locus / given genomic region may be determined, for example, using (i) the read count of a reference locus / genomic region of the same length as the restriction locus / given region that is not cleaved by restriction endonucleases, (ii) the average read count of multiple reference loci / genomic regions of the same length as the restriction locus / given region that are not cleaved by restriction endonucleases, or (iii) the read count of the restriction locus / given genomic region in an undigested control DNA sample, optionally corrected for differences in sequencing depth. Exemplary calculations are described in the following Examples section of this specification.
[0162] In additional embodiments, the methylation level is calculated by determining the total number of fragments, which is determined from the restriction locus read count and the read count of sequence reads that begin or end with a nucleotide within the restriction locus. Exemplary calculations are described in the following Examples section of this specification.
[0163] In some embodiments, the methylation level is expressed as a percentage of methylation (%), representing the percentage of DNA molecules that are methylated at the restriction locus out of the total number of DNA molecules containing the restriction locus in the sample.
[0164] The term “level of unmethylated DNA” or “level of unmethylation” for a restriction locus is a numerical value representing the number of DNA molecules that are not methylated at that restriction locus (i.e., not methylated at the CG dinucleotide within the restriction locus) out of the total number of DNA molecules containing the restriction locus in the sample. As disclosed herein, the level of unmethylated DNA for a restriction locus is calculated from the number of reads that begin or end at a nucleotide within the restriction locus after digestion with at least one methylation-sensitive restriction endonuclease and any subsequent end repair. The exact nucleotide within the restriction locus at which a sequence read begins or ends depends on the type of restriction endonuclease used in the digestion step and the length of its recognition sequence. For example, for a restriction endonuclease that produces a non-blunt end with a 5' overhang, digestion and end repair will produce fragments that begin at the second nucleotide of the recognition sequence and fragments that end at the second-to-last nucleotide of the recognition sequence. For example, for a four-base cutter that generates a non-blunt end with a 5' overhang, digestion and end repair produce a fragment that begins at the second nucleotide of the recognition sequence and a fragment that ends at the third nucleotide of the recognition sequence (Figure 15). Therefore, for restriction endonucleases that generate a non-blunt end with a 5' overhang, the "start" analysis of the restriction locus is performed on sequence reads that begin at the second nucleotide of the restriction locus (the second nucleotide of the recognition sequence), and the "end" analysis is performed on sequence reads that end at the second-to-last nucleotide of the restriction locus (the second-to-last nucleotide of the recognition sequence).
[0165] Methylation-sensitive restriction endonucleases cleave the recognition sequence only if that recognition sequence is not methylated. Therefore, the number of reads that begin or end at a nucleotide within the restriction locus represents the number of DNA molecules in the DNA sample where the restriction locus was not methylated and thus cleaved by the restriction endonuclease.
[0166] Each DNA molecule cleaved by a restriction endonuclease as disclosed herein yields two fragments (one beginning with a nucleotide in the restriction locus and the other ending with a nucleotide in the restriction locus). Therefore, it may be possible to obtain two different sequence reads for a given DNA molecule. For accurate analysis of the number of unmethylated DNA molecules present in a sample, the level of unmethylation can be calculated based on the number of sequence reads beginning with a restriction locus, the number of sequence reads ending with a restriction locus, or by the average between the two values, but not by the sum of the values. Calculating the level of unmethylated DNA at a restriction locus based on the read count of sequence reads beginning or ending with a nucleotide in the restriction locus, as disclosed herein, involves calculating the level of unmethylated DNA using the average between the two values.
[0167] It is further noted that some library preparation methods may result in depletion of small fragments that are subsequently not sequenced. Such depletion can lead to underestimation of non-methylation levels and overestimation of methylation levels. In addition, the number of sequence reads beginning at restriction loci may differ from the number of sequence reads ending at restriction loci. The present invention favorably addresses such library preparation bias. To reduce this bias and achieve more accurate results, it is preferable to determine both the number of reads beginning and ending at restriction loci and then select an orientation that provides more reads for additional analysis and calculation, or to calculate an average between the two values and use the average for additional analysis and calculation.
[0168] Accordingly, in some embodiments, the method of the present invention includes determining the number of sequence readings that begin with a nucleotide in a restriction locus, determining the number of sequence readings that end with a nucleotide in a restriction locus, and calculating the level of unmethylated DNA in the restriction locus using an orientation that provides a larger number of sequence readings. In additional embodiments, the method of the present invention includes determining the number of sequence readings that begin with a nucleotide in a restriction locus, determining the number of sequence readings that end with a nucleotide in a restriction locus, calculating the average between the two values, and using the average to calculate the level of unmethylated DNA in the restriction locus.
[0169] The number of sequence reads that begin or end with nucleotides within a restriction locus can be normalized by subtracting the expected number of sequence reads that begin or end with nucleotides within a restriction locus. The expected number of sequence reads that begin or end with nucleotides within a restriction locus can be determined, for example, using: (i) the number of sequence reads that begin or end with reference loci of the same size as the restriction locus and are not cleaved by restriction endonucleases; (ii) the average number of sequence reads that begin or end with multiple reference loci of the same size as the restriction locus and are not cleaved by the enzyme; or (iii) the number of reads in an undigested control DNA sample that begin or end with a restriction locus, optionally corrected for differences in sequencing depth. Exemplary calculations are described in the following Examples section of this specification. Using the normalized values, the level of unmethylated DNA can be calculated by creating a ratio between the normalized number of sequence reads that begin or end with nucleotides within a restriction locus and the expected read count of the restriction locus.
[0170] In some embodiments, the level of unmethylated DNA is obtained by calculating the difference between the number of reads that begin or end with a nucleotide in the restriction locus and the expected number of reads that begin or end with a nucleotide in the restriction locus, and then dividing the difference by the expected read count of the restriction locus.
[0171] In additional embodiments, the level of unmethylated DNA is calculated by determining the total number of fragments, which is determined from the restriction locus read counts and the read counts of sequence reads that begin or end with nucleotides within the restriction locus. Exemplary calculations are described in the following Examples section of this specification.
[0172] In some embodiments, the level of unmethylated DNA is expressed as a percentage (%) of unmethylation, representing the percentage of DNA molecules that are unmethylated at the restriction locus out of the total number of DNA molecules containing restriction loci in the sample.
[0173] Methylation levels (or levels of unmethylated DNA) can also be calculated for regions in the genome that span multiple restriction loci (i.e., genomic regions containing multiple restriction sites). Genomic regions spanning multiple restriction loci can include genes, intergeneric regions, promoter regions, parts of chromosomes (e.g., chromosomal arms), or entire chromosomes. Each possibility represents a separate embodiment of the present invention.
[0174] Detection of methylation changes As used herein, “detection of methylation changes” means detecting whether a tested DNA sample contains methylation changes compared to one or more reference DNA samples, detecting whether a DNA sample is characterized by different methylation profiles at selected genomic loci, and / or determining whether the methylation profile of a DNA sample is normal or contains methylation changes that indicate the presence of disease. Each possibility represents a separate embodiment of the invention. Detection of methylation changes also includes comparing methylation data obtained as disclosed herein between samples to identify differentially methylated genomic regions between samples, which can be used as DNA methylation markers. For example, methylation data obtained as disclosed herein can be analyzed to identify differentially methylated genomic regions between different types of tissues, between cancer DNA and non-cancerous DNA, between different types of cancer, or between different stages of a particular type of cancer. In some embodiments, the methods disclosed herein provide genome-wide methylation analysis. In other embodiments, the methods disclosed herein provide targeted methylation analysis. Computer software can be used in sequencing and analysis of methylation data.
[0175] The methods of the present invention may be applied to identify and analyze DNA methylation marker regions that can be used as pan-cancer diagnostic markers, i.e., DNA methylation markers indicating a group of cancer types. For example, in some embodiments, the pan-cancer markers according to the present invention indicate multiple cancer types selected from lung cancer, colorectal cancer, liver cancer, breast cancer, pancreatic cancer, uterine cancer, ovarian cancer, head and neck cancer, gastric cancer, esophageal cancer, hematological cancers (e.g., lymphoma), and sarcoma. The methods may also be applied to identify differential methylation between different types of cancer, for example, to determine methylation profiles characteristic of different types of cancer that can distinguish between different types of cancer. The methods disclosed herein are applicable to any type of cancer, including but not limited to lung cancer, bladder cancer, breast cancer, colorectal cancer, prostate cancer, gastric cancer, skin cancer (e.g., melanoma), cancers affecting the nervous system, bone cancer, ovarian cancer, liver cancer (e.g., hepatocellular carcinoma), hematological malignancies, pancreatic cancer, kidney cancer, and cervical cancer. Each cancer type represents a separate embodiment of the present invention. The method of the present invention can also be applied to identify tissue-specific methylation markers, for example, to identify methylation markers specific to lung, bladder, breast, colorectal, prostate, stomach, ovary, pancreas, kidney, and cervical tissues. Each tissue type represents a separate embodiment of the present invention. Such markers can be used, for example, to identify tissue sources of circulating cell-free DNA.
[0176] The methods of the present invention may also be applied to identify diseases (e.g., cancer) in subjects. As used herein, the term "identify" encompasses one or more of the following actions in a subject: screening for a disease, detecting the presence or absence of a disease, detecting disease recurrence, detecting susceptibility to a disease, detecting a response to treatment, determining the effectiveness of treatment, determining the stage (severity) of a disease, and determining the prognosis and early diagnosis of a disease. Each possibility represents an individual embodiment of the present invention.
[0177] As used herein, “assess cancer,” “assess the presence of cancer,” or “assess the presence or absence of cancer” refers to determining the likelihood that a subject has cancer. These terms encompass determining whether a subject should undergo confirmatory cancer tests, such as confirmatory blood tests, urine tests, cytology, imaging, endoscopy, and / or biopsy, to confirm (or rule out) the presence of cancer. These terms further encompass assisting in the diagnosis of cancer in a subject. This term further encompasses quantifying cancer-related changes in cell-free DNA samples indicating the presence of cancer. The assessment of cancer presence according to the present invention includes one or more of the following in a subject: screening for cancer, assessment of cancer recurrence, assessment of susceptibility or risk to cancer, assessment and / or monitoring of response to treatment, assessment of treatment effectiveness, assessment of cancer severity (stage), and assessment of cancer prognosis. Each possibility represents an individual embodiment of the present invention. It should be understood that a negative result in the assays disclosed herein is still considered an assessment of cancer presence according to the present invention.
[0178] The method of the present invention may further include the step of determining the tumor fraction or the fraction concentration of tumor DNA. The tumor fraction is the proportion of tumor molecules in the cfDNA sample.
[0179] Determining a “methylation profile” (or “DNA methylation profile” or “methylation profile of a DNA sample”) as disclosed herein means determining the methylation levels at one or more restriction loci, preferably multiple restriction loci. In some embodiments, determining a methylation profile includes determining the levels of methylated and unmethylated DNA at one or more restriction loci, preferably multiple restriction loci.
[0180] The “reference methylation profile” disclosed herein refers to a methylation profile determined in DNA from a known source. A “reference DNA sample” is a DNA sample from a known source. In some embodiments, the reference methylation profile is a profile determined in multiple reference DNA samples. In addition, the methods of the present invention can be used to analyze (e.g., measure) methylation changes between DNA samples taken from a single subject at different time points, for example, at different stages of a disease, or before and after treatment of a disease. The methylation profile of a DNA sample taken at a first time point can be used as a reference for the methylation profile of a DNA sample taken at a second (later) time point.
[0181] The "reference methylation level" for a specific restriction locus or a specific genomic region spanning multiple restriction loci is the level of methylation measured for that specific restriction locus / genomic region in DNA from a known source. The "reference methylation value" for a specific restriction locus or a specific genomic region spanning multiple restriction loci is a numerical value representing the level of methylation for that specific restriction locus / genomic region in DNA from a known source.
[0182] The "reference level of unmethylated DNA" for a specific restriction locus or a particular genomic region spanning multiple restriction loci is the level of unmethylated DNA measured for that specific restriction locus / genomic region in DNA from a known source.
[0183] The reference methylation / demethylation level / value may be a distribution of methylation / demethylation levels / values determined for a specific restriction locus or specific genomic region in a large set of DNA samples from a known source. In some embodiments, the methylation / demethylation level / value may be a baseline scale.
[0184] A reference scale for a specific restriction locus / genomic region may include methylation / demethylation levels / values measured for that restriction locus in multiple DNA samples from the same reference source. For example, a reference scale for a reference cancer patient or a reference scale for a reference healthy individual. Alternatively, a reference scale for a given restriction locus may include a single scale combining methylation / demethylation levels / values from both healthy and diseased individuals, i.e., reference methylation values from both sources. Generally, when a single scale is used, the values are distributed such that values from healthy individuals are at one end of the scale, for example, below a cutoff value, and values from patients are at the other end of the scale, for example, above a cutoff value. In some embodiments, methylation / demethylation levels / values calculated for a test DNA sample from an unknown source can be compared to a reference scale of health and / or disease reference values, and the score can be assigned to the methylation / demethylation levels / values calculated based on their relative position within the scale.
[0185] The terms “disease-reference methylation” (e.g., “cancer-reference methylation”), “disease-reference demethylation”, or “reference methylation (or demethylation) in disease DNA” (e.g., “reference methylation in cancer DNA”) are interchangeable and refer to methylation and / or demethylation values measured for specific restriction loci or specific genomic regions in DNA samples from subjects with the disease being analyzed, e.g., subjects with certain types of cancer. Disease-reference methylation and / or demethylation represent methylation / demethylation values in disease DNA, i.e., DNA from samples from subjects with the disease. Disease-reference methylation / demethylation levels / values can be a single value or multiple values (e.g., a distribution), as detailed above.
[0186] The term “disease DNA methylation profile” (e.g., “cancer DNA methylation profile”) refers to methylation and / or demethylation levels at multiple restriction loci determined from a sample (e.g., a plasma sample) of an individual with the disease being analyzed, for example, an individual with the specific type of cancer being analyzed.
[0187] The terms “healthy reference methylation,” “normal reference methylation,” or “reference methylation in healthy / normal DNA” are interchangeable and refer to methylation values measured for specific restriction loci / genomic regions in DNA samples derived from healthy individuals. Similarly, “healthy reference demethylation,” “normal reference demethylation,” or “reference demethylation in healthy / normal DNA” are interchangeable and refer to demethylation values measured for specific restriction loci / genomic regions in DNA samples derived from healthy individuals. “Normal” or “healthy” is defined with respect to the specific disease on which the analysis is performed. A “healthy” or “normal” individual is defined herein as an individual without detectable symptoms and / or pathological findings of disease, as determined by conventional diagnostic methods. Healthy reference methylation / demethylation levels / values may be a single ratio, a statistical value, or multiple ratios (e.g., distributions), as detailed above.
[0188] The term "healthy DNA methylation profile" or "normal DNA methylation profile" refers to the methylation and / or demethylation values at multiple restriction loci determined from a DNA sample of a normal individual, as defined above.
[0189] In some embodiments, the diagnostic methods disclosed herein include pre-determining baseline methylation / demethylation levels / values from diseased DNA. In some embodiments, the diagnostic methods of the present invention include pre-determining baseline methylation / demethylation levels / values from normal DNA, as disclosed herein.
[0190] Tissue-specific methylation profiles can also be characterized using the methods disclosed herein to establish a normal non-cancerous DNA methylation profile of the tissue. Alternatively, or further, tissue-specific methylation profiles can be characterized to identify tissue sources of circulating cell-free DNA.
[0191] In some embodiments, the detection of methylation changes according to the present invention includes identifying the presence or absence of a specific disease in a subject based on the methylation profile of a DNA sample derived from that subject.
[0192] In some embodiments, methods are provided for identifying the cellular or tissue source of a DNA sample (e.g., identifying the type of tissue from which the DNA originates, and / or identifying whether the DNA originates from normal or diseased cells / tissues).
[0193] Those skilled in the art will understand that a comparison of DNA methylation / demethylation levels / values calculated for a corresponding reference value with one or more corresponding reference values can be performed in several ways using various statistical means.
[0194] In some embodiments, comparing calculated test methylation / demethylation values for a specific restriction locus / genomic region to a reference value includes comparing the test value to a single reference value. The single reference value may correspond to the mean value obtained for reference methylation / demethylation levels / values from a large population of healthy individuals or subjects with the disease being analyzed. In other embodiments, comparing the test value to a reference value includes comparing the test value to a distribution or scale of multiple reference values. Known statistical methods can be used to determine whether the calculated value for a test sample corresponds to a disease reference value or a normal reference value.
[0195] In some embodiments, the diagnosis according to the present invention is based on an analysis of whether the methylation / demethylation level / value of a test DNA sample is a disease value, i.e., whether it indicates the disease in question. In some embodiments, the method includes the step of comparing the calculated value to its corresponding healthy reference value to obtain a score that reflects the likelihood that the calculated value is a disease value. In some embodiments, the method disclosed herein includes comparing the calculated value to its corresponding disease reference value to obtain a score that reflects the likelihood that the calculated value is a disease value. In some embodiments, a higher score indicates a higher likelihood that the calculated value is a disease value. In some embodiments, the score is based on the relative position of the calculated value within a distribution of disease values.
[0196] In some embodiments, the method involves comparing multiple values calculated for multiple restriction loci to their corresponding healthy and / or disease reference values. In some embodiments, patterns of values may be analyzed using statistical means and computerized algorithms to determine whether they represent a disease pattern or a normal, healthy pattern. Exemplary algorithms include, but are not limited to, machine learning and pattern recognition algorithms.
[0197] In some exemplary embodiments, the values calculated for a tested sample may be compared against a scale of reference values generated from a large set of cancer samples, non-cancer samples, or both. The scale may represent a threshold, hereafter referred to as the “cutoff” or “predefined threshold,” above which is the reference value corresponding to cancer, below which is the reference value corresponding to a healthy individual, or vice versa. In some embodiments, lower values at the bottom of the scale and / or below the cutoff may be from samples of normal individuals (healthy, i.e., not suffering from the cancer in question), and higher values at the top of the scale and / or above a given cutoff may be from cancer patients. In diagnoses based on the analysis of multiple restriction loci, the values calculated for each locus may be assigned a score based on its relative position on the scale, and the individual scores (for each locus) may be combined to give a single score. In some embodiments, the individual scores may be summed to give a single score. In other embodiments, the individual scores may be averaged to give a single score. In some embodiments, the single score may be used to determine whether a subject has the cancer in question, with a score above a given threshold indicating cancer.
[0198] In additional exemplary embodiments, for diagnosis based on the analysis of multiple restriction loci, for each calculated value, the probability that it represents cancer DNA may be determined based on a comparison with the corresponding cancer reference value and / or normal reference value. A score can be assigned to each locus, and then the individual scores calculated for each locus are combined (e.g., summed or averaged) to obtain a combined score. The combined score can be used to determine whether a subject is positive or negative for cancer, and a combined score above a predetermined threshold indicates cancer. Thus, in some embodiments, a threshold or cutoff score is determined, and subjects above (or below) it are identified as positive for the disease in question, e.g., the cancer in question. The threshold score distinguishes between a population of healthy individuals and a population of unhealthy individuals.
[0199] In some embodiments, the diagnostic method of the present invention includes providing a threshold score.
[0200] In many cases, statistical significance is determined by comparing two or more populations and determining confidence intervals (CIs) and / or p-values. In some embodiments, statistically significant values refer to confidence intervals (CIs) of about 90%, 95%, 97.5%, 98%, 99%, 99.5%, 99.9%, and 99.99%, and preferred p-values are about 0.1, 0.05, 0.025, 0.02, 0.01, 0.005, less than 0.001, or less than 0.0001. Each possibility represents a separate embodiment of the invention. According to some embodiments, the p-value for the threshold score is at most 0.05.
[0201] In some embodiments, the diagnostic sensitivity of the diagnostic method disclosed herein is at least 75%. In some embodiments, the diagnostic sensitivity is at least 80%. In some embodiments, the diagnostic sensitivity is at least 85%. In some embodiments, the diagnostic sensitivity of the method is at least about 90%.
[0202] In some embodiments, the “diagnostic sensitivity” of a diagnostic assay used herein refers to the percentage of patients diagnosed as positive (the percentage of “true positives”). Therefore, patients not detected by the assay are “false negatives.” Subjects who are not affected and test negative in the assay are called “true negatives.” The “specificity” of a diagnostic assay is 1 minus the false positive rate, where the “false positive” rate is defined as the proportion of individuals who do not have the disease but test positive. A particular diagnostic method may not provide a definitive diagnosis of a condition, but it may be sufficient if it provides a positive indicator useful for diagnosis.
[0203] In some embodiments, the diagnostic specificity of the diagnostic methods disclosed herein may be at least about 65%. In some embodiments, the diagnostic specificity of the methods may be at least about 70%. In some embodiments, the diagnostic specificity of the methods may be at least about 75%. In some embodiments, the diagnostic specificity of the methods may be at least about 80%.
[0204] In some embodiments, the diagnostic method according to the present invention includes generating a report (on paper or electronically) based on the methylation profile. The report may be communicated to the subject and / or the subject's healthcare provider.
[0205] In some embodiments, the diagnostic method according to the present invention includes referring the subject to follow-up examinations and screenings.
[0206] Additional genetic and epigenetic characterization In addition to DNA methylation / demethylation values, information regarding DNA mutations, copy number variations, and nucleosome positioning of cell-free DNA can be obtained from the same sequencing data disclosed herein. Generally, cell-free DNA circulates in fragments ranging between 120 and 220 bp. This pattern corresponds to the length of DNA wrapped around a single nucleosome, in addition to a short stretch of approximately 20 bp (linker DNA) bound to histones. Since nucleosome positioning varies between different tissues and in malignant cells, fragmentation patterns have been shown to be useful in determining the dominant cell type of origin contributing to the cfDNA pool.
[0207] Advantageously, the determination of DNA methylation profiles and at least one additional genetic or epigenetic trait disclosed herein can be performed based on the same sequencing data.
[0208] In some embodiments, the sequence-based assays disclosed herein combine the detection of methylation changes with mutation detection and analysis of additional epigenetic properties, all in a single assay. This assay advantageously enables the combined analysis of small amounts of DNA in a single assay.
[0209] Combination analysis of methylation and additional genetic and epigenetic traits is useful for enhancing the detection of cancer (or any other condition / tissue source).
[0210] In some exemplary embodiments, a method for detecting the presence or absence of cancer in a subject is: (A) Profile the methylation of DNA samples disclosed herein to detect the presence or absence of hypermethylation in one or more cancer-related genomic regions, and (B) One or more of the following: To determine the presence or absence of one or more cancer-related mutations (e.g., cancer-related mutations in oncogenes / tumor suppressors), To determine the presence or absence of cancer-related copy number variations, and This includes determining the presence or absence of cancer-related nucleosome positioning, (A) and (B) are performed using the same sequencing data. The presence of hypermethylation in one or more cancer-related genomic regions, as well as the determination of at least one of one or more cancer-related mutations, cancer-related copy number variations, and cancer-related nucleosome positioning, indicates the presence of cancer in the subject.
[0211] Non-methylated cancer-associated changes can be combined with methylation information in a dependent or independent manner, depending on whether the cancer-associated changes are found on the same DNA fragment, and changes found on the same fragment provide a stronger indicator of cancer presence.
[0212] In some embodiments, methods are provided for profiling the genetic and epigenetic properties of a DNA sample, the methods comprising profiling the methylation of the DNA sample disclosed herein and determining at least one additional genetic or epigenetic property of the DNA sample, the at least one additional genetic or epigenetic property being selected from DNA mutation, copy number variation and nucleosome positioning, and the profiling of methylation and the determination of the at least one additional genetic or epigenetic property being performed using the same sequencing data, thereby profiling the genetic and epigenetic properties of the DNA sample.
[0213] In some embodiments, a method is provided for detecting the presence or absence of a disease in a subject, the method comprising profiling the methylation of a DNA sample disclosed herein and determining at least one additional genetic or epigenetic characteristic of the DNA sample, the at least one additional genetic or epigenetic characteristic being selected from DNA mutation, copy number variation and nucleosome positioning, the profiling of methylation and the determination of the at least one additional genetic or epigenetic characteristic being performed using the same sequencing data to obtain the genetic and epigenetic characteristics of the DNA sample, comparing the genetic and epigenetic characteristics of the DNA sample to one or more reference genetic and epigenetic characteristics, and determining the presence or absence of a disease based on the comparison. In some embodiments, the disease is cancer.
[0214] Systems and kits In some embodiments, systems for detecting methylation changes in DNA samples are provided herein. In some embodiments, systems and methods for detecting genetic and epigenetic changes in DNA samples are provided herein. In additional embodiments, kits for detecting methylation changes in DNA samples are provided herein. In additional embodiments, kits for detecting genetic and epigenetic changes in DNA samples are provided herein.
[0215] The system according to the present invention includes a computer processor for performing calculations, for example, for performing an assay and / or processing the results. In some embodiments, computer implementation methods are provided herein.
[0216] In some embodiments, the system and kit are for profiling the methylation of a DNA sample according to the methods disclosed herein. In some embodiments, the system and kit are for profiling the genetic and epigenetic properties of a DNA sample according to the methods disclosed herein. In additional embodiments, the system and kit are for detecting methylation changes in a DNA sample according to the methods disclosed herein. In additional embodiments, the system and kit are for detecting genetic and epigenetic changes in a DNA sample according to the methods disclosed herein.
[0217] In some embodiments, the system according to the present invention DNA sample and A methylation-sensitive restriction endonuclease and / or a methylation-dependent restriction endonuclease for digesting a DNA sample, Components for preparing a sequencing library containing DNA fragments generated by multiple restriction endonucleases, A high-throughput sequencer for determining sequences in a sequence determination library and generating sequence reads, The invention includes computer software stored on a non-temporary computer-readable medium, which instructs a computer processor to profile the genetic and epigenetic properties of a DNA sample based on a plurality of sequence readings in accordance with a method disclosed herein. In some embodiments, the computer software instructs the computer processor to profile the methylation of the DNA sample based on a plurality of sequence readings in accordance with a method disclosed herein.
[0218] In some embodiments, computer software stored on a non-temporary computer-readable medium instructs a computer processor to determine genetic and epigenetic changes in a DNA sample based on multiple sequence readings, according to a method disclosed herein. In some embodiments, computer software stored on a non-temporary computer-readable medium instructs a computer processor to determine methylation changes in a DNA sample based on multiple sequence readings, according to a method disclosed herein.
[0219] As used herein, “components” for preparing a sequencing library includes biochemical components (e.g., enzymes, nucleotides), chemical components (e.g., buffers), and technical components (e.g., equipment such as tubes, vials, and pipettes).
[0220] In some embodiments, the kit or system according to the present invention includes, in addition to restriction enzymes, one or more buffers and other components necessary for DNA digestion.
[0221] In some embodiments, a system is provided for profiling the genetic and epigenetic properties of a cell-free DNA sample, the system comprising a cell-free DNA sample and computer software stored on a non-temporary computer-readable medium, which, when executed, sends the following steps to a computer processor: (i) Preparing a sequencing library comprising receiving sequencing data of a library of DNA molecules obtained after digestion of a cell-free DNA sample with at least one methylation-sensitive restriction endonuclease, and ligating sequencing adapters to DNA molecules in the restriction endonuclease-treated DNA, wherein each adapter is capable of ligating to both digested and undigested DNA molecules, (ii) Determining the methylation value of at least one restriction locus from sequencing data, and at least one additional genetic or epigenetic characteristic of a cell-free DNA sample selected at will from DNA mutations, copy number variations, and nucleosome positioning, Computer software including instructions configured or directed to perform, The amount of cell-free DNA containing 3000 haploid equivalents is sufficient for the method, the cell-free DNA sample has not been subjected to amplification before library preparation, and the methylation level and at least one additional genetic or epigenetic characteristic of the cell-free DNA sample are determined based on the same sequencing data.
[0222] In some embodiments, a system for profiling the methylation of a DNA sample is provided herein, the system comprising computer software stored on a non-temporary computer-readable medium, which, when executed, sends the following steps to a computer processor: (i) to receive a sequence reading of a library of DNA molecules obtained after digestion of the DNA sample with at least one methylation-sensitive restriction endonuclease, (ii) Selecting at least one restriction locus and determining the number of sequence reads to cover a given genomic region of at least 50 bp in length that contains the restriction locus, (iii) Complying with the calculation of the methylation value of at least one restriction locus based on the read count and reference read count determined in step (ii).
[0223] In some embodiments, a system for profiling the methylation of a DNA sample is provided herein, the system comprising computer software stored on a non-temporary computer-readable medium, which, when executed, sends the following steps to a computer processor: (i) to receive a sequence reading of a library of DNA fragments obtained after digestion of the DNA sample with at least one methylation-sensitive restriction endonuclease, (ii) Mapping multiple sequence reads to a reference genome to generate mapped sequence reads, and selecting at least one restriction locus in the reference genome, (iii) Determining the read count of at least one restriction locus from the mapped sequence reads, wherein the read count represents the number of DNA molecules in the DNA sample in which at least one restriction locus was methylated and therefore remained intact. (iv) Determining from mapped sequence reads the read count of sequence reads that begin or end with a nucleotide in at least one restriction locus, wherein the read count represents the number of DNA molecules in the DNA sample that are not methylated at least one restriction locus and are therefore cleaved by restriction endonucleases. (v) Instructions configured or directed to perform the following: calculate the level of methylated DNA at at least one restriction locus based on the read count of at least one restriction locus determined in step (iii); and calculate the level of unmethylated DNA at at least one restriction locus based on the read count of sequence reads that begin or end with a nucleotide in at least one restriction locus determined in step (iv).
[0224] In some embodiments, the computer software further instructs the computer processor to compare the genetic and epigenetic profiles of the tested DNA sample with one or more reference genetic and epigenetic profiles, and based on the comparison, to output whether the DNA sample is a normal DNA sample or a diseased DNA sample.
[0225] In some embodiments, the computer software further instructs the computer processor to compare the methylation profile of the tested DNA sample with one or more reference methylation profiles and, based on the comparison, output whether the DNA sample is a normal DNA sample or a diseased DNA sample.
[0226] In some embodiments, computer software according to the present invention receives raw input data for a high-throughput sequencing run. In some embodiments, the computer software instructs a computer processor to analyze the sequencing data to determine genetic and epigenetic profiles as disclosed herein. In some embodiments, the computer software instructs a computer processor to analyze the sequencing data to determine DNA methylation and / or DNA demethylation values, as disclosed herein.
[0227] Computer software includes processor-executable instructions stored on a non-temporary computer-readable medium. Computer software may also include stored data. Computer-readable mediums are tangible computer-readable media such as compact discs (CDs), magnetic storage devices, optical storage devices, random access memory (RAM), read-only memory (ROM), or any other tangible medium of representation.
[0228] It is understood that the computer-related methods, steps, and processes described herein are performed using software that, when executed, is configured to perform instructions or is stored in non-volatile or non-temporary computer-readable instructions that instruct a computer processor or computer to do so.
[0229] Each of the systems, servers, computer devices, and computers described herein can be implemented in one or more computer systems and configured to communicate over a network. They can also all be implemented in a single computer system. In one embodiment, the computer system includes a bus or other communication mechanism for communicating information and a hardware processor coupled to the bus for processing information.
[0230] A computer system also includes main memory, such as random access memory (RAM) or other dynamic storage devices, coupled to a bus for storing information and instructions executed by the processor. Main memory can also be used to store temporary variables or other intermediate information during the execution of instructions by the processor. Once such instructions are stored in a non-temporary storage medium accessible to the processor, they render the computer system into dedicated devices customized to perform the operations specified by the instructions.
[0231] The computer system further includes read-only memory (ROM) or other static storage devices coupled to the bus for storing static information and instructions to the processor. Storage devices such as magnetic disks or optical disks are provided and coupled to the bus for storing information and instructions.
[0232] Computer systems can be connected to displays via buses to display information to computer users.
[0233] Input devices, including alphanumeric and other keys, are coupled to a bus to communicate information and command selections to the processor. Another type of user input device is cursor control, such as a mouse, trackball, or cursor directional keys, which transmits directional information and command selections to the processor and controls the movement of a cursor on the display.
[0234] According to one embodiment, the technology described herein is executed by a computer system in response to a processor executing one or more sequences of one or more instructions contained in a main memory. Such instructions can be read into the main memory from another storage medium such as a storage device. Execution of the sequence of instructions contained in the main memory causes the processor to execute the process steps described herein. In alternative embodiments, hardwired circuitry may be used in place of, or in combination with, software instructions.
[0235] As used herein, the term storage medium refers to any non-transitory medium that stores data and / or instructions that cause a machine to operate in a specific manner. Common forms of storage media include, for example, a floppy disk, a flexible disk, a hard disk, a solid state drive, magnetic tape, or any other magnetic data storage medium, a CD-ROM, any other optical data storage medium, any physical medium with patterns of holes, RAM, PROM, and EPROM, FLASH-EPROM, NVRAM, any other memory chip or cartridge.
[0236] A storage medium is different from a transmission medium, but can be used in combination with a transmission medium. Transmission media participate in transferring information between storage media. For example, transmission media include coaxial cables, copper wire, and optical fibers, including the wires that make up a bus.
[0237] The following examples are presented to more fully illustrate certain embodiments of the present invention. However, they should in no way be construed as limiting the broad scope of the present invention. Those skilled in the art can readily devise many variations and modifications of the principles disclosed herein without departing from the scope of the present invention. EXAMPLES
[0238] Example 1 - Methylation-sensitive DNA digestion followed by next-generation sequencing (NGS) versus bisulfite treatment + NGS In the following examples, sequencing data obtained from cell-free DNA samples subjected to methylation-sensitive enzyme digestion followed by NGS were compared with sequencing data obtained after bisulfite conversion and NGS. First, the sequencing data of pooled plasma DNA samples were examined. Next, the sequencing data of individual plasma DNA samples, each containing approximately 10–200 ng of DNA (corresponding to approximately 3,000–60,000 1ploid equivalents of DNA), were examined.
[0239] A. Pooled plasma samples from healthy individuals DNA was extracted and pooled from plasma samples of 56-60 healthy control subjects using the QIAamp® Circulating Nucleic Acid Kit (QIAGEN, Hilden, Germany). 500 ng of aliquots were retained as untreated control DNA, 1700 ng of aliquots were subjected to bisulfite conversion using the EZ DNA Methylation-Gold® Kit (Zymo Research), and the remaining DNA (770 ng) was subjected to digestion with methylation-sensitive restriction enzymes HinP1I and AciI. Methylation-sensitive digestion was performed by incubating the sample with 10 units of HinP1I and 5 units of AciI at 37°C for 2 hours, followed by inactivation at 65°C for 20 minutes.
[0240] Next, sequencing libraries were prepared from each sample (enzyme-treated, bisulfite-treated, and untreated control sample) using the NEBNext Ultra DNA Library Prep Kit for the enzyme-treated and untreated control samples, and the ACCEL-NGS® METHYL-SEQ DNA LIBRARY kit (swift) for the bisulfite-treated sample. The sequencing libraries were prepared while preserving information at the ends of the DNA molecules by adding an Illumina platform sequencing adapter using enzymatic ligation. The libraries were subjected to whole-genome next-generation sequencing using an Illumina NovaSeq 6000 sequencing platform equipped with an S4 flow cell. Sequence readings from each sample were mapped to the complete human genome (hg18 genome construction).
[0241] Sequence determination metrics Table 1 and Figures 1A-1B and 2A-2B summarize the sequencing metrics, copy number completeness data, and nucleosome positioning completeness data obtained for pooled plasma DNA samples. For the analysis of copy number completeness, the number of hits at each genomic position located >100 bp from restriction loci in methylation-sensitive digested aliquots was compared to the corresponding number of hits obtained from untreated aliquots. The same analysis was performed for bisulfite-treated aliquots (compared to untreated aliquots). Pearson correlations were calculated for all data points in each experimental setup (methylation-sensitive digested DNA and bisulfite treatment). Pearson correlations yielded numbers between -1 and 1, with numbers closer to 1 indicating better correlation.
[0242] The procedure was the same as for copy number analysis, except that for nucleosome positioning completeness analysis, we used "hit span 100" (= the number of reads that start >50 bp upstream and end >50 bp downstream of the analyzed genomic location) instead of hits.
[0243] [Table 1]
[0244] As the data shows, methylation-sensitive digestion yielded substantially the same mapping and unique mapping rates as those obtained for the untreated control sample, reaching a unique mapping rate of over 92%. In contrast, the bisulfite-treated sample showed a significant loss of information, with a unique mapping rate of only about 80%.
[0245] The loss of information in bisulfite-treated samples was further demonstrated by copy number and nucleosome positioning integrity data: as seen in Figures 1A and 2A, similar patterns of copy number and nucleosome positioning were observed in methylation-sensitive digested samples and untreated samples. These patterns were not maintained in bisulfite-treated samples. Pearson correlation analysis showed a correlation of 0.9 for copy number and 0.88 for nucleosome positioning between methylation-sensitive digested samples and untreated control samples. In contrast, correlations of 0.67 (copy number) and 0.58 (nucleosome positioning) were obtained between bisulfite-treated samples and untreated control samples (Figures 1B and 2B).
[0246] B. Plasma samples from lung cancer patients Section A above provides results using pooled plasma samples containing relatively large amounts of DNA. For analysis, it would be interesting to check the differences between methylation-sensitive digestion and bisulfite conversion using individual plasma samples containing much smaller amounts of DNA.
[0247] To this end, DNA was extracted from plasma samples of treatment-naïve patients with non-small cell lung cancer (NSCLC). DNA was extracted as described in Section A above, and subjected to bisulfite conversion or digestion with methylation-sensitive restriction enzymes HinP1I and AciI. The amount of DNA extracted from plasma samples ranged from approximately 10 to 200 ng of DNA per sample (corresponding to approximately 3,000 to 60,000 haploid equivalents of DNA). Following enzymatic digestion or bisulfite conversion, samples were subjected to library preparation and sequencing as described above.
[0248] Exemplary results from two patients identified as BMD LNG165 (26 ng of cell-free DNA) and BMD LNG166 (94 ng of cell-free DNA) are provided below.
[0249] Sequencing metrics Tables 2 to 3 and Figures 3A to 3B and 4A to 4B summarize the obtained sequencing metrics, copy number completeness data and nucleosome positioning completeness data for each plasma DNA sample.
[0250] [Table 2]
[0251] [Table 3]
[0252] The results show that for small amounts of DNA, the number of reads obtained is significantly reduced when bisulfite treatment versus methylation-sensitive digestion is used: that is, the number of reads obtained for bisulfite-treated DNA, and importantly, the number of uniquely mapped reads, was less than half the amount obtained for methylation-sensitive digested DNA. Furthermore, while methylation-sensitive digested DNA exhibited a unique mapping rate of approximately 90%, the unique mapping rate for bisulfite-treated DNA was only 77 to 78%.
[0253] Significant loss of information in bisulfite-treated samples was further demonstrated by copy number and nucleosome positioning completeness data: Pearson correlation analysis showed correlations of 0.735 and 0.693 for copy number and 0.647 and 0.595 for nucleosome positioning between methylation-sensitive digested DNA and untreated control samples. In contrast, correlations of 0.196 and 0.161 (copy number) and 0.19 and 0.176 (nucleosome positioning) were obtained between bisulfite-treated DNA and untreated control samples (Figures 3A-3B, 4A-4B). Bisulfite sequencing resulted in the loss of virtually all copy number and nucleosome positioning information when small amounts of DNA were used in the assay.
[0254] CG coverage Figures 5A and 5B show the distribution of CG depths in bisulfite-treated DNA and methylation-sensitive digested DNA. More specifically, the graphs show the number of CG sites in the genome covered at each depth by each method. Figure 5A shows data obtained from a sample from patient BMD LNG165. Figure 5B shows data obtained from a sample from patient BMD LNG166.
[0255] Genome-wide methylation analysis using methylation-sensitive enzyme digestion is limited to CGs located within the recognition site of the enzyme used in the assay, whereas bisulfite sequencing, in principle, covers all CG sites in the genome. The ability to investigate only fractions of CG sites in the genome has been considered one of the main limitations of restriction enzyme-based methylation analysis. However, the data presented herein show that bisulfite provides a broader CG range at the lower end of depth compared to methylation-sensitive digestion, while there is a continuous and sharp decrease in the number of CGs covered in bisulfite-treated DNA as depth increases. In contrast, methylation-sensitive digestion shows substantially constant coverage even at depths beyond 250–300. At high depths, methylation-sensitive digestion provides significantly better CG coverage compared to bisulfite.
[0256] For example, in DNA samples from patient BMD LNG165 (Figure 5A), methylation-sensitive digestion covered more genomic CGs than bisulfite at depths greater than 165. At a depth of 300, the methylation-sensitive digestion sample covered 4.16 M of CG, compared to only 44 K of CG in the bisulfite-treated sample. In DNA samples from patient BMD LNG166 (Figure 5B), methylation-sensitive digestion covered more genomic CGs than bisulfite at depths greater than 255. At a depth of 400, the methylation-sensitive digestion sample covered 4.24 M of CG, compared to only 65 K of CG in the bisulfite-treated sample.
[0257] Therefore, methylation-sensitive digestion provides coverage of millions of CGs at very high depths, enabling the detection of rare methylation signals, such as tumor-derived methylated DNA molecules in plasma in the early stages of tumors, which may be present in very low amounts—less than 1%. The data showed that bisulfites do not provide sufficient coverage at the depth required to identify rare signals, and such rare signals are likely to be missed when using bisulfite sequencing for small amounts of DNA.
[0258] Detection of methylation changes A set of low-background hypermethylation marker loci was compiled, which exhibit hypermethylation in tumor versus normal tissue and are characterized by low background methylation in the plasma of healthy individuals. Methylation levels were determined as described in Example 2 below. This set of marker loci was compiled based on samples from two lung cancer patients (BMD LNG165 and patient BMD LNG166) and pooled plasma samples from healthy individuals, and included low-background hypermethylation loci observed using both detection methods, i.e., methylation-sensitive digestion + NGS and bisulfite conversion + NGS. In addition, a set of isomethylation marker loci, i.e., loci that do not exhibit different methylation levels between tumor and normal tissue, was compiled.
[0259] Next, plasma DNA from each patient was analyzed using methylation-sensitive digestion + NGS or bisulfite conversion + NGS to determine the methylation levels of low-background hypermethylated marker loci in the patient's plasma. A threshold methylation level was set, and if it was exceeded, the marker locus was considered "detected." To obtain a detection specificity of 95%, the threshold was determined based on the set of isomethylated marker loci. The number of marker loci that exceeded the threshold (i.e., were detected) was compared in methylation-sensitive digestion DNA and bisulfite-converted DNA from each patient. The results are summarized in Figure 6. Methylation analysis using methylation-sensitive digestion + NGS detected significantly more methylation changes in plasma in both samples compared to bisulfite + NGS.
[0260] Mutation detection Tumor mutations were defined as genotypes found in tumor DNA that differ from the most dominant genotype in the corresponding normal tissue from the same patient. The proportion of reads with the mutated genotype in tumor DNA represented the tumor mutation level, and the proportion of reads with the same mutated genotype in the patient's plasma represented the plasma mutation level. For each sample, the mean tumor and plasma mutation levels were calculated across all mutations to calculate the tumor mutation load (i.e., mean plasma mutation level / mean tumor mutation level). The tumor mutation load represents the proportion of tumor DNA in the patient's plasma. To control for sequencing noise, the tumor mutation load of patient A was compared to the control tumor mutation load calculated from patient B's tumor mutations (i.e., mean mutation level of patient B's tumor mutations in patient A's plasma / patient A's tumor mutation level).
[0261] The results are summarized in Figures 7A and 7B. Tumor mutations were detected in plasma at levels clearly exceeding sequencing noise by methylation-sensitive digestion + NGS, but bisulfite + NGS mutations could not be distinguished from high sequencing noise.
[0262] Example 2 - Genetic and epigenetic profiling of tumor and plasma DNA in lung cancer patients Methylation and mutational analysis were performed on samples from two lung cancer patients identified as BMD LNG165 and BMD LNG166. Clinical data from each patient are detailed in Figures 9A and 9B.
[0263] Sample preparation for analysis is shown in Figure 8A. Normal lung tissue samples, tumor lung tissue samples, and blood samples were provided to each patient. Blood samples were separated into buffy coat and plasma samples. DNA was extracted from each sample as shown in the figure. Normal tissue DNA, tumor tissue DNA, and buffy coat DNA were fragmented by sonication. Next, the DNA was digested and purified using methylation-sensitive restriction enzymes HinP1I and AciI, as in Example 1. Aliquots of normal tissue DNA from each patient were left undigested and stored as controls. The purified DNA samples were subjected to library preparation and sequencing as described above.
[0264] Figure 8B shows the sample preparation of control samples taken from 100 healthy control subjects. The control samples included buffy-coated samples and plasma samples from each control subject. As shown in the figure, DNA was extracted from each sample. Buffy-coated DNA was fragmented by sonication, digested with HinP1I and AciI, and then purified. Aliquots of buffy-coated DNA from each control subject were left undigested and stored as controls. Plasma DNA was digested with HinP1I and AciI and purified. Aliquots of plasma DNA were taken for quality control (e.g., evaluation of the quality of plasma separation) and to create an undigested control pool of plasma DNA. The purified DNA samples were subjected to library preparation and sequencing as described above.
[0265] Sequence readings from each sample were mapped to the complete human genome (hg18 genome construction). Alignments with CIGAR&MAPQ>0&abs(TLEN)≦500bp were selected for additional methylation and mutation analysis to identify methylation changes and mutations in tumors and their presentation in plasma.
[0266] For each genome location, we determined the "hit span 100," which is the number of reads that start >50 bp upstream and end >50 bp downstream of the genome location.
[0267] "Hitspan 100" is an alignment of at least 100 bp, representing DNA molecules of at least 100 bp in length in the DNA sample remaining after methylation-sensitive digestion and library preparation. In addition, since the copy number of such alignments reflects the nucleosome boundary, analysis of such alignments is advantageous for evaluating nucleosome arrangement in cell-free DNA in addition to methylation, where high copy numbers are typical in the center of the nucleosome and low copy numbers are typical at the boundaries between nucleosomes.
[0268] Furthermore, many cancer-associated methylation changes occur within CG islands, i.e., within the genomic region that reaches the CG site receiving methylation. The “hitspan 100” region surrounding the analyzed CG site located within the restriction locus of the enzyme used in the assay typically includes additional restriction loci of the enzyme containing additional CG sites. Thus, a “hitspan 100” alignment represents a DNA molecule at least 100 bp in length, where the analyzed restriction locus, as well as any additional restriction loci within the DNA molecule, were all methylated in the DNA sample and remained intact after digestion with the enzyme used in the assay. Analyzing alignments that are at least 100 bp in length and contain multiple restriction loci that are all methylated in the DNA sample increases the specificity of cancer-associated hypermethylation signals, enabling improved and more accurate detection of differences between normal and cancerous samples. Such methylation analysis is particularly advantageous for CG sites located within CG islands.
[0269] The "hitspan 100" value was normalized relative to the median "hitspan 100" value in the same sample. For example: the normalized "hits within range 100" at a specific locus = the number of "hits within range 100" at that locus / the median number of "hits within range 100" in the sample. Normalization was performed relative to the median across chromosomes 1 to 22.
[0270] Whole-genome methylation analysis Methylated loci were defined as restriction loci having a number of normalized "hit span 100" that exceeds a predetermined threshold in the undigested normal tissue pool.
[0271] The background methylation levels of methylated loci were determined as follows: Normalized "hitspan 100" in a pool of digested normal plasma / Normalized "hitspan 100" in a pool of undigested normal plasma.
[0272] If the normalized "hitspan 100" in the undigested normal plasma pool was 0, the value 1 / median "hitspan 100" was used instead.
[0273] The tumor methylation levels of methylation loci were determined as follows: Normalized "hitspan 100" in tumors / Normalized "hitspan 100" in the undigested normal tissue pool.
[0274] The normal methylation levels of methylation loci were determined as follows: Normalization of normal tissue with a "hit span of 100" / Normalization of undigested normal tissue pool with a "hit span of 100".
[0275] The plasma methylation levels of methylation loci were determined as follows: Normalized "hitspan 100" in plasma / Normalized "hitspan 100" in undigested normal plasma pool.
[0276] For each patient, a set of low-background methylation loci was compiled by selecting methylation loci whose background methylation levels were below a predetermined threshold.
[0277] In addition, a set of hypermethylation loci exhibiting hypermethylation in tumor versus normal tissue was compiled. The set of hypermethylation loci was compiled by determining the tumor-normal differential methylation level (=tumor methylation level - normal methylation level) of hypermethylation loci and selecting methylation loci with tumor-normal differential methylation levels exceeding a predetermined threshold.
[0278] A set of hypomethylated loci exhibiting low methylation in tumor versus normal tissue was edited. A set of hypermethylated loci was edited by determining the tumor-normal differential methylation level (=tumor methylation level - normal methylation level) of hypomethylated loci and selecting methylated loci with tumor-normal differential methylation levels below a predetermined threshold.
[0279] We edited a set of isomethylation loci that do not exhibit different methylation levels between tumor and normal tissue. This set of isomethylation loci was edited by determining the tumor-normal differential methylation level (= tumor methylation level - normal methylation level) of the methylation loci and selecting methylation loci that are neither tumor-normal hypermethylated nor tumor-normal hypomethylated.
[0280] The analysis results for each patient are shown in Figures 9A and 9B. Millions of hypermethylation and hypomethylation events were detected in the tumors of each patient. Furthermore, thousands of low-background hypermethylation events were detected in the plasma of each patient. The detected events represent putative methylation markers.
[0281] Figures 10A and 10B show that we were able to identify thousands of methylation loci that exhibit particularly strong hypermethylation signals in plasma.
[0282] Whole-genome mutation analysis Tumor mutations were defined as genotypes found in tumor DNA that differ from the most dominant genotype in the corresponding normal tissue from the same patient.
[0283] The proportion of reads with the mutated genotype in tumor DNA represented the tumor mutation level, and the proportion of reads with the same mutated genotype in the patient's plasma represented the plasma mutation level. For each sample, the mean tumor and plasma mutation levels were calculated across all mutations to calculate the tumor mutation load (i.e., mean plasma mutation level / mean tumor mutation level). The tumor mutation load represents the proportion of tumor DNA in the patient's plasma.
[0284] For each patient, a set of low-background mutations was compiled by selecting mutations with a plasma mutation background (= proportion of mutations in the normal plasma pool) below a predetermined threshold. Furthermore, the mean mutation rate in plasma was determined for each patient. The analysis results for each patient are shown in Figures 11A and 11B.
[0285] Multi-omics domain A multi-omics region is defined herein as a genomic region having tumor hypermethylation sites (highly methylated in tumors compared to normal tissue) and tumor mutation sites within a given distance. The method of the present invention aims to detect cancer-related genetic and epigenetic changes in cell-free DNA samples. Therefore, multi-omics regions up to 150 bp are preferred for identifying DNA containing both tumor hypermethylation sites and tumor mutation sites on the same sequence readout. In this example, multi-omics regions in which tumor hypermethylation loci and tumor mutations are within 100 bp of each other were searched for in tumor samples from patient BMD LNG165 and patient BMD LNG166.
[0286] Analysis identified 6,060 multi-omics regions in patient BMD LNG165 and 9,471 multi-omics regions in patient BMD LNG166. An example of a multi-omics region in BMD LNG165 is shown in Figure 12 (chr.7 pos.150220856-150220921).
[0287] Multi-omics alignment was defined as an alignment across a multi-omics region where CIGAR&MAPQ>0&TLEN>0&TLEN≦500bp. Examples of multi-omics alignment types are shown in Figure 13 and include: Matched methylation alignments in which the cancer phenotype is observed at both methylation sites (where the site is methylated) and mutation sites (where the mutant is present): The alignments include all hypermethylation restriction sites (all letters of the recognition sequence of the restriction enzyme used in the assay are present in the alignment, e.g., GCGC for HinP1I) and contain the mutant genotype during reading.
[0288] Mismatched methylation alignments where the cancer phenotype is observed at the methylation site (the site is methylated) and the normal phenotype is observed at the mutation site (the WT variant is present): The alignment spans all hypermethylation restriction sites (e.g., all letters of GCGC are present in the alignment) and contains the WT (reference) genotype during reading.
[0289] Matched unmethylated alignments in which the normal phenotype is found at both the methylation site (where this site is unmethylated) and the mutation site (where the WT variant exists): The alignments begin or end at the exact cleavage site (starting at position n of the restriction site or ending at position n+1) and include the WT (reference) genotype during reading.
[0290] Mismatched nonmethylated alignments where the normal phenotype is seen at the methylation site (where it is unmethylated) and the cancer phenotype is seen at the mutation site (where a mutant variant is present): The alignments start or end at the exact cleavage site (starting at position n of the restriction site or ending at position n+1) and contain the mutant genotype during reading.
[0291] The above demonstrates that the method disclosed herein, employing methylation-sensitive digestion followed by next-generation sequencing, works with small amounts of DNA and is sufficiently sensitive yet accurate to receive vast amounts of information, including methylation data, mutation data, and more. This method is advantageous for both, for example, the discovery of new methylation markers and for clinical diagnostic applications. This method enables the detection of signals that cannot be detected by bisulfites.
[0292] Example 3 - Direct calculation of methylated and unmethylated DNA levels In the following examples, methylation / demethylation calculations were performed by digesting cell-free DNA derived from plasma samples with methylation-sensitive restriction enzymes HinP1I and AciI, followed by library preparation, next-generation sequencing, and sequence reading analysis.
[0293] Figure 14 shows the methylation-sensitive HinP1I site before and after digestion and end repair. Cell-free DNA molecules that are not methylated at the HinP1I site undergo digestion, producing double-stranded DNA molecules with blunt (sticky) ends corresponding to the HinP1I cleavage site. Specifically, since HinP1I has a 4-base cleavage site, digestion produces a pair of double-stranded DNA molecules, one with a 2-base 5' overhang and the other with a complementary 5' overhang. The blunt ends are subjected to end repair (e.g., using the NEBNext Ultra DNA Library Prep Kit) to produce blunt-ended DNA molecules. After end repair, two types of DNA fragments are obtained: a fragment ending at the third nucleotide (G nucleotide) of the HinP1I recognition sequence (3' end), and a fragment beginning at the second nucleotide (C nucleotide) of the HinP1I recognition sequence (5' end).
[0294] Figure 15 shows the differences in DNA fragments obtained after digestion and end repair of cell-free DNA molecules spanning HinP1I restriction sites that are methylated or unmethylated at the cleavage site. Black circles represent methylation. DNA molecules methylated at the cleavage site remain intact after digestion, yielding DNA fragments that span the cleavage site. DNA molecules that are not methylated at the cleavage site are digested by the enzyme. After end repair, the result is DNA fragments that start or end at the recognition sequence (specifically, fragments ending at the third nucleotide G of the recognition sequence and fragments starting at the second nucleotide C).
[0295] Experimental Procedure Plasma samples were collected from 56 healthy control subjects. DNA was extracted from the plasma samples using the QIAamp® Circulating Nucleic Acid Kit (QIAGEN, Hilden, Germany) and pooled. 450 ng of aliquots were retained as undigested control DNA, and the remaining DNA (800 ng) was subjected to digestion: the samples were incubated with 10 units of HinP1I and 5 units of AciI at 37°C for 2 hours, followed by inactivation at 65°C for 20 minutes. After digestion, sequencing libraries were prepared using the NEBNext Ultra DNA Library Prep Kit. The libraries were subjected to next-generation sequencing using the Illumina NovaSeq 6000 sequencing platform with an S4 flow cell. Sequence readings from digested and undigested DNA samples were mapped to the complete human genome (hg18 genome construction).
[0296] Calculation of methylated DNA levels To calculate the level of methylated DNA, sequence reads were plotted as read counts per 4bp locus. Loci corresponding to restriction enzyme (HinP1I) cleavage sites were analyzed, and the number of reads extending to the completely intact site was recorded. Figure 16A shows an exemplary analysis of a 4bp locus corresponding to a HinP1I site in a digested DNA sample (marked by a rectangle). As can be seen from the figure, a decrease in read count is observed at this HinP1I site. The decrease indicates enzymatic digestion, and the read count at this location corresponds to the number of DNA fragments where the locus remains intact. As mentioned above, HinP1I is methylation-sensitive and therefore does not cleave methylated DNA. Thus, the read count at this restriction locus corresponds to the number of DNA molecules in the DNA sample where the restriction locus is methylated.
[0297] The level of methylated DNA at restriction loci is calculated as follows:
[0298] (Math 1) Methylated DNA level = Actual read count of restriction loci Predictive reading count of restriction loci Here, the expected read count for restriction loci can be calculated from the following: (i) Read count of 4bp reference loci that are not cleaved by enzymes, (ii) The average reading count of multiple 4bp reference loci that are not cleaved by the enzyme (ideally consisting of loci with the same copy number as the restriction loci in the undigested sample), or (iii) Reading count of restriction loci in undigested control samples (correctable for depth differences).
[0299] The reference locus may be a 4bp stretch located immediately upstream or downstream of the restriction locus, or a 4bp locus located at a more distant location within the genome.
[0300] Depth difference correction can be performed as follows: Predicted reading count for restriction loci = Restriction locus reading count in undigested sample / (Average sequencing depth in undigested sample / Average sequencing depth in digested sample) Average sequence determination depth = Total number of reads * Average read length / genome size
[0301] By multiplying the obtained level of methylated DNA by 100, the percentage (%) of methylated DNA at the tested HinP1I locus in the original DNA sample can be obtained.
[0302] Calculation of the level of unmethylated DNA To calculate the level of unmethylated DNA, sequence reads were plotted as read counts ending at each base across the entire genome. Alternatively or additionally, sequence reads could be plotted as read counts starting at each base across the genome. Genomic loci corresponding to restriction enzyme (HinP1I) cleavage sites were analyzed.
[0303] Figure 16B shows the "start" analysis for the HinP1I site and adjacent regions analyzed above. As can be seen from the figure, a peak is observed at the second nucleotide (C nucleotide) of the cleavage site. The peak height, i.e., the number of sequence reads starting at the second nucleotide of the cleavage site, corresponds to the number of DNA fragments cleaved by the enzyme. As mentioned above, HinP1I is methylation-sensitive and therefore cleaves unmethylated DNA. Thus, the peak height corresponds to the number of DNA molecules in the DNA sample where the restriction locus is not methylated.
[0304] Figure 16C shows the "termination" analysis of the HinP1I site analyzed in the above and adjacent regions. As can be seen from the figure, a peak is observed at the third nucleotide (G nucleotide) of the cleavage site. The peak height, i.e., the number of sequence reads ending at the third nucleotide of the cleavage site, corresponds to the number of DNA fragments cleaved by the enzyme at this cleavage site. As mentioned above, HinP1I is methylation-sensitive and therefore cleaves unmethylated DNA. Thus, the peak height corresponds to the number of DNA molecules in the DNA sample where the restriction locus is not methylated.
[0305] The level of methylated DNA at restriction loci is calculated as follows: Level of unmethylated DNA = (Actual number of reads starting or ending at the restriction locus - Expected number of reads starting or ending at the restriction locus) / Expected number of reads at the restriction locus
[0306] The expected number of reads that begin or end at a restriction locus can be calculated as follows: (i) The number of reads that start or end at a 4bp reference locus that is not cleaved by enzymes, (ii) The average number of reads that start or end in a family of 4bp reference loci that are not cleaved by enzymes, or (iii) The number of readouts that start or end at restriction loci in the undigested control sample (correctable for depth differences).
[0307] The predicted read count for restriction loci can be calculated from the following: (i) Read count of 4bp reference loci that are not cleaved by enzymes, (ii) The average reading count of multiple 4bp reference loci that are not cleaved by the enzyme (ideally consisting of loci with the same copy number as the restriction loci in the undigested sample), or (iii) Reading count of restriction loci in undigested control samples (correctable for depth differences).
[0308] The reference locus may be a 4bp stretch located immediately upstream or downstream of the restriction locus, or a 4bp locus located at a more distant location within the genome.
[0309] Each DNA molecule cleaved by a restriction endonuclease as disclosed herein yields two fragments, one beginning with a nucleotide in a restriction locus and the other ending with a nucleotide in a restriction locus. Therefore, it may be possible to obtain two different sequence reads for a given DNA molecule. For accurate analysis of the number of unmethylated DNA molecules present in a sample, the unmethylated DNA level may be calculated based on the number of sequence reads beginning with a restriction locus, the number of sequence reads ending with a restriction locus, or the average of the two values, but not necessarily based on the sum of the values. When using a library preparation method that depletes small fragments, it is preferable to calculate the unmethylated DNA level using an orientation with a larger number of sequence reads to avoid bias due to the depletion of small fragments.
[0310] By multiplying the obtained level of unmethylated DNA by 100, the percentage (%) of methylated DNA at the tested HinP1I locus in the original DNA sample can be obtained.
[0311] Such methylation / demethylation analysis is particularly advantageous for CG sites located in genomic regions with low CG content.
[0312] Simultaneous calculation of methylated and unmethylated DNA levels To simultaneously calculate the levels of methylated and unmethylated DNA, first calculate the "total number of fragments" as follows: Total number of fragments = Restriction locus read count + Number of reads starting or ending at restriction loci - Estimated number of reads starting or ending at restriction loci
[0313] The expected number of reads that begin or end at restriction loci is calculated as described above.
[0314] The levels of methylated and unmethylated DNA are calculated using the total number of fragments as follows:
[0315] (Math 2) Methylated DNA level = Restriction locus read count Total number of fragments Level of unmethylated DNA = (Number of reads that start or end at restriction loci - Estimated number of reads that start or end at restriction loci) / Total number of fragments
[0316] By multiplying the obtained levels of methylated and unmethylated DNA by 100, the percentage (%) of methylated and unmethylated DNA at the tested HinP1I locus in the original DNA sample can be obtained.
[0317] Example 4 - Analysis of restriction loci using methylated and unmethylated DNA levels The levels of methylated and unmethylated DNA were calculated for eight CG dinucleotides located within the HinP1I restriction locus identified as CG#1-8 in the pooled plasma DNA sample described in Example 3 (Table 4). As detailed in Example 3, the pooled DNA sample was digested with methylation-sensitive restriction enzymes HinP1I and AciI, followed by library preparation, next-generation sequencing, and alignment of the sequence readings against the complete human genome.
[0318] Exemplary raw data for CG#1 (highly methylated), CG#4 (not highly methylated), and CG#5 are shown in Figures 17A–17C. The upper panel of each figure shows the number of reads per 4bp locus to determine the number of reads for restriction loci. Restriction loci are indicated by rectangles. The lower panel of each figure shows the read counts for each base in the reference genome to determine the read counts for sequence reads that start or end at restriction loci. The presentation of "end" or "start" follows the direction that provided more reads.
[0319] The level of methylated DNA at each restriction locus was calculated by dividing the read count of the restriction locus by the expected read count (read count of the control locus) and multiplying by 100 to obtain the percentage of methylated DNA at the restriction locus.
[0320] The level of unmethylated DNA at each restriction locus was calculated by subtracting the expected number of reads that begin or end at the restriction locus, then dividing by the expected number of reads at the restriction locus, and multiplying by 100 to obtain the percentage of unmethylated DNA at the restriction locus. For each restriction locus, the number of reads that begin and end at the restriction locus was determined, and further calculations were performed based on the larger number of reads.
[0321] The discrepancy level (%) was calculated for each restriction locus by determining the difference between the sum of the methylation and demethylation percentages calculated in this embodiment and the expected sum of 100%. Mismatch % = (Methylation % + Unmethylation %) - 100
[0322] The results are summarized in Table 4. Restriction loci are listed in Table 4 in ascending order according to the level of discrepancy. The level of discrepancy can be used when evaluating and selecting potential DNA methylation markers, and loci with lower levels of discrepancy may be preferred. The level of discrepancy can also be used as an indicator of appropriate sample preparation and analysis for already identified DNA methylation markers, where lower levels of discrepancy indicate appropriate sample preparation and analysis.
[0323] [Table 4]
[0324] The results demonstrate that direct determination of demethylation in addition to methylation using the same sequencing data provides complementary methylation information for genomic regions, enabling improved methylation profiling, more accurate and effective evaluation of potential DNA methylation markers, and more precise subsequent analysis of methylation differences between samples.
[0325] Example 5 - Methylation profiling and diagnosis of lung cancer using lung cancer DNA methylation markers The methylation profiles of DNA samples extracted from plasma samples are determined in six genomic regions containing HinP1I restriction loci that are differentially methylated between lung cancer DNA and normal non-lung cancer DNA. The genomic regions previously disclosed in International Publication No. 2019 / 142193, assigned to the applicant of the present invention, are identified as Sequence IDs 1-6 and are detailed in Table 5.
[0326] [Table 5] * Starting position. The description refers to the position on the hg18 genome build.
[0327] Figure 18 is a flowchart illustrating an exemplary method for profiling methylation of a DNA sample according to an embodiment of the present invention. The exemplary method includes the following steps: The DNA sample is digested with the 1801-methylation-sensitive restriction endonuclease HinP1I. 1802 - Prepare a sequencing library from digested DNA using an adapter ligated to multiple DNA fragments. 1803 - High-throughput sequencing of a sequencing library for obtaining sequence readings. 1804 - Mapping sequence reading for the complete human genome, Select genomic regions 1-6 containing restriction loci differentially methylated between 1805-lung cancer DNA and normal non-lung cancer DNA. For each restriction locus in genome regions 1-6 of 1806, determine the read count of the restriction locus. For each restriction locus in genome regions 1-6, determine the read count for sequence reads starting at the second nucleotide of the restriction locus and the read count for sequence reads ending at the second-to-last nucleotide of the restriction locus, and select the orientation with the larger read count. For each restriction locus in genome regions 1-6, the level of methylated DNA is calculated based on the read count of the restriction locus. For each restriction locus in genome regions 1-6, the level of unmethylated DNA is calculated based on the read count of sequence reads that begin or end with a nucleotide within the restriction locus. This allows us to obtain methylation profiles of DNA samples in genomic regions 1-6 (step 1910).
[0328] Figure 19 is a flowchart illustrating an additional exemplary method for profiling methylation of a DNA sample according to embodiments of the present invention. The exemplary method includes the following steps: The DNA sample is digested with the 1901-methylation-sensitive restriction endonuclease HinP1I. 1902 - Prepare a sequencing library from digested DNA using an adapter ligated to multiple DNA fragments. 1903 enriches genomic regions 1-6 containing restriction loci differentially methylated between lung cancer DNA and normal non-lung cancer DNA. 1904 - High-throughput sequencing of enriched sequencing libraries for obtaining sequence reads. 1905 - Assign sequence reading to one of genome regions 1-6. Follow steps 1906-1910 above to obtain methylation profiles of DNA samples in genomic regions 1-6.
[0329] Figure 20 is a flowchart illustrating an exemplary method for determining whether a DNA sample is positive or negative for lung cancer according to an embodiment of the present invention. The exemplary method includes the following steps: 2001- As described above, by combining the levels of methylated and unmethylated DNA, a methylation profile of DNA samples in genomic regions 1-6 is obtained. Compare the methylation profile of the 2002-DNA sample with at least one reference DNA methylation profile in genomic regions 1-6 (e.g., lung cancer reference profile and / or healthy non-lung cancer methylation profile), and Based on 2003-comparison, DNA samples are identified as positive or negative for lung cancer.
[0330] Figure 21 is a flowchart illustrating an additional exemplary method for determining whether a DNA sample is positive or negative for lung cancer, according to embodiments of the present invention. The exemplary method includes the following steps: 2101-As described above, the methylation profiles of DNA samples in genomic regions 1-6 are obtained by combining the levels of methylated and unmethylated DNA. The score is calculated based on the methylation profile in 2102-genome regions 1-6. 2103 - Compare the score to the cutoff value, and 2104-Based on comparison, DNA samples are identified as positive or negative for lung cancer.
[0331] The foregoing description of specific embodiments is intended to fully illustrate the general nature of the invention, and others can be readily modified and / or adapted to various uses of such specific embodiments by applying current knowledge without excessive experimentation and without departing from the general concept. Therefore, such adaptations and modifications should be understood, and are intended, within the meaning and scope of the equivalents of the disclosed embodiments. It should be understood that any expressions or terms used herein are for illustrative purposes only and not for limitation. Means, materials, and steps for carrying out the various chemical structures and functions disclosed can take various alternative forms without departing from the invention.
Claims
1. A method for profiling the genetic and epigenetic characteristics of cell-free DNA (cfDNA) samples derived from a target, (a) To obtain restriction endonuclease-treated DNA in which the cell-free DNA sample is digested with at least one methylation-sensitive restriction endonuclease, in which the methylation sites are intact and the unmethylation sites are cleaved. (b) Preparing a sequencing library from restriction endonuclease-treated DNA while preserving the sequence information of the ends of the DNA molecules, wherein the preparation of the sequencing library includes ligating sequencing adapters to DNA molecules in the restriction endonuclease-treated DNA, and each adapter can ligate both digested and undigested DNA molecules. (c) To sequence the sequence determination library using a high-throughput sequence determination method and provide sequence determination data, (d) Determining the methylation value of at least one restriction locus from the sequencing data, The amount of cell-free DNA containing 3000 haploid equivalents is sufficient for the method, and the cell-free DNA sample is not subjected to amplification before library preparation, Determining the methylation level of at least one restriction locus is necessary. (i) Selecting at least one restriction locus and determining the number of sequence readouts to cover a predetermined genomic region of at least 50 bp in length that contains the restriction locus, (ii) calculating the methylation value of the at least one restriction locus based on the read count and reference read count determined in step (i), The aforementioned reference read count is A) A read count determined for the predetermined genomic region containing the restriction locus in the undigested control DNA sample, which is at least 50 bp in length. B) A read count determined for the predetermined genomic region containing the restriction locus in the undigested control DNA sample, which is at least 50 bp in length and corrected for the difference in sequencing depth. C) A read count determined using a reference region of at least 50 bp in length containing a reference locus that is not cleaved by the restriction endonuclease, or D) A method comprising an average read count determined using a plurality of reference regions, each at least 50 bp in length, containing a reference locus that is not cleaved by the restriction endonuclease.
2. The method according to claim 1, wherein the predetermined genomic region begins at least 25 bp upstream of the cleavage site within the restriction locus and ends at least 25 bp downstream of the cleavage site within the restriction locus.
3. The method according to claim 1, wherein step (i) is to determine a number of sequence reads that cover a predetermined genomic region of at least 100 bp in length containing the restriction locus.
4. The method according to claim 3, wherein the predetermined genomic region begins at least 50 bp upstream of the cleavage site within the restriction locus and ends at least 50 bp downstream of the cleavage site within the restriction locus.
5. The method according to any one of claims 1 to 4, wherein the amount of cell-free DNA containing 6,000 haploid equivalents is sufficient for the method.
6. The method according to any one of claims 1 to 4, wherein the cell-free DNA is plasma cell-free DNA, and the amount of the cell-free DNA is the amount obtained from 9 to 10 mL of blood.
7. The method according to any one of claims 1 to 4, wherein the amount of cell-free DNA is 10 to 200 ng.
8. The method according to any one of claims 1 to 4, wherein the amount of cell-free DNA is 20 to 100 ng.
9. The method according to any one of claims 1 to 4, further comprising the method of providing the DNA treated with the restriction endonucleases to end repair prior to ligation of the sequencing adapter to obtain a DNA molecule having blunt ends, wherein the at least one methylation-sensitive restriction endonuclease generates a non-blunt end, and the method further comprises the method of providing the DNA treated with the restriction endonuclease to end repair prior to ligation of the sequencing adapter to obtain a DNA molecule having blunt ends.
10. The method according to any one of claims 1 to 4, wherein the high-throughput sequencing is whole-genome high-throughput sequencing.
11. The method according to any one of claims 1 to 4, wherein the high-throughput sequencing is high-throughput sequencing targeting only the target.
12. The method according to any one of claims 1 to 4, wherein the at least one restriction locus is a plurality of restriction loci.
13. The method according to any one of claims 1 to 4, wherein the at least one methylation-sensitive limiting endonuclease is a plurality of methylation-sensitive limiting endonucleases, and the digestion by the plurality of methylation-sensitive limiting endonucleases is simultaneous digestion.
14. The method according to claim 13, wherein the plurality of methylation-sensitive limiting endonucleases include HinP1I.
15. The method according to claim 13, wherein the plurality of methylation-sensitive limiting endonucleases include AciI.
16. The method according to claim 13, wherein the digestion is carried out using HinP1I and AciI.
17. The method according to any one of claims 1 to 4, wherein the step of subjecting the cell-free DNA sample to digestion with at least one methylation-sensitive limiting endonuclease further comprises determining the digestive efficacy and proceeding to the preparation of a sequencing library if the digestive efficacy exceeds a predetermined threshold.
18. A method for detecting cancer-related genetic and epigenetic changes in a cell-free DNA sample (cfDNA) derived from a subject, comprising: performing the method according to any one of claims 1 to 4 to obtain a genetic and epigenetic profile of the cfDNA sample; and detecting cancer-related genetic and epigenetic changes in the cfDNA sample by comparing the genetic and epigenetic profile of the cfDNA sample with one or more reference genetic and epigenetic profiles selected from cancer profiles and non-cancer profiles.
19. The method according to any one of claims 1 to 4, wherein the at least one restriction locus is located within a CG island.
20. The method according to any one of claims 1 to 4, wherein calculating the methylation value includes normalizing the read count determined in step (d) with respect to the median read count of the cfDNA sample to obtain a normalized read count, and calculating the ratio of the normalized read count to a normalized reference read count.
21. A method for genetic and epigenetic profiling of a cfDNA sample, comprising: determining the methylation value of at least one restriction locus as described in any one of claims 1 to 4; and further determining at least one additional genetic or epigenetic characteristic of the cfDNA sample selected from DNA mutations, copy number variations, and nucleosome positioning from sequencing data.
22. A method for identifying genomic regions differentially methylated between first and second cfDNA sources, A first cfDNA methylation profile is obtained by profiling the methylation of at least one cfDNA sample from the first source according to the method of any one of claims 1 to 4, A second cfDNA methylation profile is obtained by profiling the methylation of at least one cfDNA sample from the second source according to the method of any one of claims 1 to 4, A method comprising comparing the first and second cfDNA methylation profiles to identify genomic regions that are differentially methylated between the first and second cfDNA sources.
23. The method according to claim 22, wherein the first source of cfDNA is plasma cfDNA from a cancer patient, and the second source of cfDNA is plasma cfDNA from one or more healthy individuals.
24. The method according to claim 22, wherein the first and second cfDNA sources are from different stages of cancer.
Citation Information
Patent Citations
Non-invasive determination of fetal or tumor methylome using plasma.
JP2015536639A
Processes and compositions for methylation-based enrichment of fetal nucleic acid from maternal sample useful for non-invasive prenatal diagnoses
JP2018064594A
Epigenetic discrimination of DNA
WO2018035125A1
Methods and systems for detecting methylation changes in DNA samples
WO2020188561A1