Useful combinations of restriction enzymes
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- NUCLEIX LTD
- Filing Date
- 2023-05-22
- Publication Date
- 2026-05-26
AI Technical Summary
In the prior art, when analyzing cell free DNA (cfDNA), chemical methods damage the integrity and deviation of DNA, enzymatic methods require multiple steps and are complex, making it difficult to efficiently and accurately detect the CpG methylation state.
The two restriction enzymes, HinP1I and AciI, were digested by digesting cfDNA, combined with heat treatment, completely inactivated enzyme activity, avoid interference from downstream steps, and analyze the CpG methylation state through high-throughput sequencing and real-time PCR.
Efficient digestion and accurate methylation detection of cfDNA are achieved, reducing chemical damage, simplifying the process, and improving the accuracy and reliability of the detection.
Abstract
Description
Technical Field
[0001] (Incorporation by Reference) All documents and online information cited in this specification are incorporated by reference in their entirety.
[0002] (Field of the Invention) The present invention belongs to the field of analyzing the methylation of cytosine residues in DNA, and is particularly useful for analyzing human cell-free DNA found, for example, in plasma.
Background Art
[0003] Various techniques for analyzing the methylation of cytosine residues in DNA are known. One common method involves bisulfite conversion, which uses bisulfite to convert unmethylated cytosine to uracil. The converted DNA is then analyzed, and by comparing the bisulfite-treated DNA with the non-bisulfite-treated DNA, it is revealed which cytosine residues were not converted to uracil (and thus were methylated). One major drawback of this technique is that bisulfite conversion is chemically harsh and results in a high level of degradation of the source material, which is a problem when using small amounts of source DNA. Also, the chemical conversion is biased and inherently noisy.
[0004] Another technique uses methylation-sensitive restriction enzymes (MSREs), where the activity of the enzyme is blocked when cytosine within the CpG site in the recognition sequence is methylated. Various MSRE-based techniques are available using either a single enzyme or a combination. For example, the HELP assay uses a combination of HpaII and MspI. The recognition sequence for both of these enzymes is CCGG, but HpaII is methylation-sensitive. Thus, comparison of the digestion products for the two enzymes can reveal which CCGG sites are CpG-methylated. Other MSRE-based assays using multiple enzymes are known, including methods using three or four (or more) different enzymes.
[0005] It is also possible to use methylation-dependent restriction enzymes (MDREs), which digest their recognition sequence only when cytosine is methylated, i.e., the reverse of an MSRE-based assay.
[0006] These enzyme-based techniques are also used to analyze the methylation of cell-free DNA (cfDNA), as in the EpiCheck platform commercially available from Nucleix. The use of multiple enzymes for the digestion of cfDNA has also been reported. For example, a mixture of HhaI, HpaII, and exonuclease I has been used to digest cfDNA, and a mixture of two or three of BstUI, HhaI, and / or HpaII has been used to analyze fetal cfDNA in maternal blood. Methods for digesting cfDNA using BstUI alone or in combination with HhaI, HpaII, or HpaII + HinP1I are known, such as methods using Bsh1236I + HhaI, Bsh1236I + HpaII + HinP1I, HinP1I + HhaI, HhaI + HpaII, or HhaI + AccII, and methods using HpaII + AciI + HpyCH4IV and HinP1I. Also known are methods using a mixture of two or three of BstUI + HpaII, HpaII + CfoI, or AccII, HpaII, and HpyCH4IV, and similarly, methods using BstUI + MluI, BstUI + HpaII, or NauI + MbuBI. Many of these methods involve a downstream PCR step, and thus it is necessary to inactivate the MSRE before PCR so that the amplicon (which is not methylated) is not digested.
[0007] There is still a need for additional combinations of MSREs and / or MDREs that are useful for digesting cfDNA and can provide various advantages over the combinations already in use, including techniques involving downstream PCR. There is also a need for additional techniques for analyzing CpG methylation using combinations of restriction enzymes that can provide advantages over known methods. SUMMARY OF THE INVENTION
[0008] The present invention provides a method for digesting cfDNA using a combination of restriction enzymes including HinP1I and AciI, the method comprising: (i) digesting the cfDNA with the restriction enzymes; and (ii) inactivating the restriction enzymes by heating for longer than 15 minutes. It is preferred to inactivate the restriction enzymes by heating for longer than 20 minutes. This inactivation period can achieve complete inactivation and, unlike the 15-minute heat inactivation used in some known methods, ensures that residual enzyme activity does not persist in downstream steps.
[0009] The present invention also provides a method for digesting cfDNA using a combination of restriction enzymes including HinP1I and AciI, the method comprising: (i) digesting the cfDNA with the restriction enzymes for no more than 11 hours; and (ii) completely inactivating the restriction enzymes by heating. A 11-hour digestion is more than sufficient for digesting all cfDNA in a typical sample, and the 14-hour or 16-hour digestion times in some known methods are unnecessarily long.
[0010] Digestion in less than 11 hours is sufficient for the complete digestion of cfDNA in a typical sample obtained from blood collected in a typical collection tube containing an anticoagulant such as a K2EDTA collection tube, although digestion has been found to be inhibited in some types of collection tubes. In particular, this inhibition can be seen in collection tubes containing an anticoagulant and an agent that inhibits the release of genomic DNA from white blood cells, such as BCT from Streck (see below), into the plasma component of the blood sample. Thus, when blood is stored in such tubes, it may be useful to increase the digestion time compared to blood stored in a K2EDTA collection tube. Accordingly, the present invention also provides a method for digesting cfDNA using a combination of restriction enzymes including HinP1I and AciI, the method comprising: (i) providing a blood sample contained within a collection tube containing an anticoagulant and an agent that inhibits the release of genomic DNA from white blood cells in the sample into the plasma component of the blood sample; (ii) preparing plasma from the blood sample; and (iii) digesting the cfDNA with the restriction enzyme for at least 2 hours. The method may further comprise (iv) inactivating the restriction enzyme by heating, as disclosed elsewhere herein. The use of this type of collection tube advantageously allows the blood sample to be stored and / or transported (e.g., at room temperature) after collection while inhibiting contamination of the cfDNA by genomic DNA that may otherwise be released from white blood cells into the plasma during storage. It has been found that the 1-hour digestion time used in some known methods after blood collection from pregnant women in such tubes may be too short to achieve consistent and reliable results with cfDNA, and step (iii) may involve a digestion for a period longer than 2 hours (e.g., at least 4, 6, 8, 10, 12, 14 hours, or more, e.g., about 16 hours), ideally of a length sufficient to provide substantially complete digestion of the cfDNA.In contrast to previous reports that certain cfDNA blood collection tubes are not suitable for downstream analysis of methylated sequences in cfDNA, it has now been found that blood samples stored in these tubes can in fact be subjected to the cfDNA analysis disclosed herein, particularly when digestion is carried out for two hours or more.
[0011] The present invention likewise provides a method for digesting cfDNA using a combination of restriction enzymes including HinP1I and AciI, the method comprising the step of digesting the cfDNA with the restriction enzymes for at least two hours, wherein the cfDNA is derived from plasma prepared from a blood sample contained in a collection tube comprising an anticoagulant and an agent that inhibits the release of genomic DNA from leukocytes in the sample into the plasma component of the blood sample. The method may further comprise the step of inactivating the restriction enzymes by heating, as disclosed elsewhere herein.
[0012] The present invention likewise provides a method for digesting cfDNA using a combination of restriction enzymes including HinP1I and AciI, the method comprising: (i) preparing plasma from a blood sample contained in a collection tube comprising an anticoagulant and an agent that inhibits the release of genomic DNA from leukocytes in the sample into the plasma component of the blood sample; and (ii) digesting the cfDNA with the restriction enzymes for at least two hours. The method may comprise either or both of the step of receiving the collection tube prior to (i) and the step of inactivating the restriction enzymes by heating after step (ii), as disclosed elsewhere herein.
[0013] The present invention also provides a method for analyzing cfDNA, the method comprising: (i) digesting the cfDNA using a combination of restriction enzymes including HinP1I and AciI; and (ii) sequencing the digested cfDNA. In particular, step (ii) may involve next-generation sequencing in which the digested cfDNA is converted into a sequencing library and then a sequencing reaction is performed on the library.
[0014] The present invention also provides a method for analyzing cfDNA, the method comprising sequencing a digested cfDNA sample, the sample having been pre-digested using a combination of restriction enzymes comprising HinP1I and AciI.
[0015] The present invention also provides a method for analyzing cfDNA, the method comprising: (i) digesting cfDNA for 11 hours or less using a combination of restriction enzymes comprising HinP1I and AciI; (ii) completely inactivating the restriction enzymes; and (iii) performing real-time PCR on the digested cfDNA. As described above, a 11-hour digestion is sufficient for digestion of a typical amount of cfDNA, particularly when followed by downstream real-time PCR analysis.
[0016] The present invention also provides a method for analyzing cfDNA, the method comprising: (i) digesting cfDNA for 11 hours or less using a combination of restriction enzymes comprising HinP1I and AciI; (ii) completely inactivating the restriction enzymes; and (iii) performing real-time PCR on the digested cfDNA. As described above, a 11-hour digestion is sufficient for digestion of a typical amount of cfDNA, particularly when followed by downstream real-time PCR analysis.
[0017] The present invention also provides a method for analyzing data derived from digested cfDNA, the data comprising real-time PCR quantification cycle data or sequence read data from high-throughput sequencing, the cfDNA having been digested by a method for digesting cfDNA as disclosed herein. The method can be used to provide the methylation status of one or more CpG sites of interest in cfDNA.
[0018] The present invention also provides a composition comprising a plurality of restriction enzymes, wherein the plurality of restriction enzymes consists of MSRE and / or MDRE, and (i) at least two different restriction enzymes in the plurality of restriction enzymes have different recognition sequences, and (ii) the restriction enzymes can be completely inactivated by heating at 65°C. By recognizing different sequences, the number of genomic CpG sites that can be analyzed is increased compared to using a combination of enzymes having the same recognition sequence (e.g., HhaI and HinP1I as used by Zhao et al., both of which recognize GCGC but have different cleavage sites therein). Inactivation at 65°C is milder (e.g., with respect to the generation of unwanted ssDNA) and easier (e.g., lower energy intensity) than inactivation of mixtures containing enzymes such as HpaII, AvaI, HaeII, or MluI (which require heating at 80°C for inactivation according to the suppliers of such enzymes), and provides a distinct advantage over mixtures containing enzymes that cannot be easily heat-inactivated (e.g., as reported for BstUI, PvuI, and HhaI).
[0019] This composition can be based on MSRE without the need for MDRE. Thus, the present invention also provides a composition comprising a plurality of MSRE, wherein (i) at least two different MSRE in the plurality of MSRE have different recognition sequences, and (ii) the plurality of MSRE can be completely inactivated by heating at 65°C. This composition may not contain MDRE.
[0020] The present invention also provides a composition comprising HinP1I and AciI as the only two restriction enzymes in the composition. The pairing of these enzymes covers over 99% of the CpG islands in the human genome and is easier to prepare with higher accuracy than more complex mixtures that have been used occasionally. A composition containing only two restriction enzymes is advantageous in that, without impairing the enzyme activity, it is easier to adjust the digestion conditions (e.g., incubation temperature, reaction buffer, etc.) to a combination of two enzymes as compared to a combination of three or more enzymes. For example, Bsh1236I, HpaII, and HinP1I as used by Ellinger et al. require different optimal buffers for 100% activity. ThermoFisher recommends "Tango buffer" as optimal for its HpaII and HinP1I, but recommends "Buffer R" for Bsh1236I. As reported by Ellinger et al., preparing a reaction mixture containing this combination of enzymes using only Tango buffer impairs Bsh1236I activity.
[0021] Both HinP1I and AciI exhibit 100% activity at 37°C and exhibit 100% activity in the same reaction buffer, namely rCutSmart™ (NEB). Both of these enzymes also use the same diluent (Diluent A; NEB) and can be completely inactivated by heating to 65°C.
[0022] Compositions containing only two restriction enzymes are also advantageous in downstream library preparation methods involving depletion of small DNA fragments that are not subsequently sequenced. Prior to sequencing, small DNA fragments are typically depleted to remove free (i.e., unligated) sequencing adapters and / or adapter dimers that can otherwise interfere with the efficiency and / or quality of DNA sequencing. When starting cfDNA molecules are cleaved at two or more sites, a greater number of DNA fragments are provided, some of which are very small and are removed (and thus not sequenced) during this process. By increasing the number of restriction enzymes used for digestion, the likelihood of generating these small DNA fragments increases, resulting in bias in library preparation and underestimation of unmethylated CpG sites. Compositions containing only two restriction enzymes limit this bias.
[0023] The present invention also provides a composition comprising HinP1I and AciI, wherein the ratio of HinP1I to AciI is at least 1.2:1 (measured in enzyme units). The use of an excess amount of HinP1I has been found to give better results than the previously used ratios of 0.5:1 or 1:1. Without wishing to be bound by theory, it is thought that improvement may occur because AciI can cleave the human genome more frequently than HinP1I and a single cleavage is sufficient to impair PCR amplification, so less AciI activity is required to achieve the same damage.
[0024] In these various methods and compositions, (a) the ratio of HinP1I to AciI is at least 2:1, (b) a source of Mg ++ ions is provided, (c) the restriction enzymes are used at a pH above 7, for example, in the range of 7.5 to 8.5, (d) the cfDNA is human cfDNA, for example, human plasma cfDNA, and / or (e) the amount of cfDNA subjected to digestion is preferably from 10 to 400 ng, for example, from 10 to 250 ng or from 10 to 200 ng.
[0025] The present invention also provides further methods and compositions that include or use these compositions and / or methods, as detailed below.
DETAILED DESCRIPTION OF THE INVENTION
[0026] Methylation The methods and compositions disclosed herein are useful for the analysis of DNA methylation, particularly for analyzing the presence or absence of 5-methyl modification of cytosine with respect to CG dinucleotide sequences (commonly referred to as "CpG" dinucleotides or "CpG sites") in eukaryotic DNA. CpG sites are not randomly distributed throughout the eukaryotic genome, but are frequently found in clusters known as "CpG islands". These islands are formally defined as regions that are at least 200 bp in length, have a GC content of greater than 50%, and have an observed CpG to predicted CpG ratio of greater than 60% (i.e., the number of CpG sites multiplied by the length of the sequence, divided by the product of the number of C's and the number of G's, is greater than 0.6) (Gardiner-Garden & Frommer (1987) J Mol Biol 196:261-82). CpG islands are often found near the start of genes in mammalian genomes, and approximately 70% of the promoters near transcription start sites in the human genome contain CpG islands. Methylation of multiple CpG sites within the CpG island of a promoter is generally associated with stable silencing of gene expression from that promoter.
[0027] The human genome sequence contains approximately 28,000,000 CpG sites (per haploid genome) and has approximately 30,000 CpG islands. In any given nucleated cell, some CpG sites are methylated and others are not. The methylation pattern can vary between different cells and tissues within a subject, and thus a particular CpG may be methylated in one cell or tissue but not in a different cell or tissue within the same subject.
[0028] Tumors are known to exhibit different methylation patterns compared to non-tumor cells (or compared to other types of tumors). Some sites can be hypermethylated in tumors, while other sites can be hypomethylated, and differences in these patterns are used to assist in tumor diagnosis.
[0029] Cell-free DNA The methods and compositions disclosed herein are particularly useful for analyzing cell-free DNA (cfDNA) found in vivo in body fluids rather than in intact cells in animals, i.e., fragmented genomic DNA. The origin of cfDNA is not fully understood but is generally thought to be released from cells in processes such as apoptosis and necrosis. CfDNA is highly fragmented compared to intact genomic DNA (see, e.g., Alcaide et al. (2020) Scientific Reports 10, article 12564) and generally circulates as fragments 120 - 220 bp in length, with a peak at about 168 bp (in humans).
[0030] CfDNA is present in many body fluids including, but not limited to, blood and urine, and the methods and compositions disclosed herein can use any suitable source of cfDNA, e.g., a blood sample (such as venous blood) or a urine sample. Ideally, cfDNA is isolated from blood, and the blood can be treated to obtain plasma (i.e., the liquid remaining after a whole blood sample has been subjected to a separation process to remove blood cells, typically involving centrifugation) or serum (i.e., plasma without clotting factors such as fibrinogen). Thus, the methods and compositions disclosed herein can be used as part of a so-called liquid biopsy test and can be performed using plasma or serum cfDNA. Accordingly, the methods disclosed herein can include the step of purifying cfDNA from a blood, plasma, or serum sample to provide cfDNA for digestion and analysis. The method can also include the step of obtaining a blood sample and preparing plasma or serum therefrom, thus providing a source for downstream purification of cfDNA.
[0031] Blood can be collected in a tube containing an anticoagulant and an agent that inhibits the release of genomic DNA from white blood cells in the sample into the plasma component of the blood sample. Such tubes are commercially available from Streck (La Vista, NE) as glass cfDNA “blood collection tubes” or “BCTs”, as considered by Diaz et al. (2016) PLoS One 11(11):e0166354, and they can stabilize cfDNA in blood for up to 14 days at 6 - 37°C (thus providing an advantage compared to a typical K2EDTA collection tube). Useful anticoagulants include, but are not limited to, EDTA, heparin, or citrate. Agents useful for inhibiting the release of genomic DNA from white blood cells include, but are not limited to, diazolidinyl urea, imidazolidinyl urea, dimethylol - 5,5 - dimethylhydantoin, dimethylol urea, 2 - bromo - 2 - nitropropane - 1,3 - diol, oxazolidine, sodium hydroxymethylglycinate, 5 - hydroxy - methoxymethyl - 1 - 1 - aza - 3,7 - dioxabicyclo[3.3.0]octane, 5 - hydroxymethyl - 1 - 1 - aza - 3,7 - dioxabicyclo[3.3.0]octane, 5 - hydroxypoly[methyleneoxy]methyl - 1 - 1 - aza - 3,7 - dioxabicyclo[3.3.0] - octane, quaternary adamantines, and mixtures thereof. Other useful components can include quenchers (e.g., lysine, ethylenediamine, arginine, urea, adenine, guanine, cytosine, thymine, spermidine, or any combination thereof) that can weaken the reaction of free aldehydes with DNA in the sample, aurintricarboxylic acid, metabolic inhibitors (e.g., glyceraldehyde and / or sodium fluoride), and / or nuclease inhibitors. For example, the tube can contain imidazolidinyl urea (or diazolidinyl urea), EDTA, and glycine. Further information on suitable collection tubes can be found in International Publication No. WO 2013 / 123030 and U.S. Patent Application Publication No. US 2010 / 0184069.
[0032] Various plastic tubes: the "cell-free DNA collection tube" from Roche, made of PET; the "LBgard blood tube" from Biomatrica, made of plastic and suitable for up to 8.5 mL of blood; and other useful collection tubes including, but not limited to, the "PAXgene Blood DNA Tube" from PreAnalytiX or Qiagen are available. These various tubes are discussed in more detail in Kerachian et al. (2021) Clinical Epigenetics 13, 193, Schmidt et al. (2017) Clinica Chimica Acta 269:94-8, and Grolz et al. (2018) Current Pathobiology Reports 6:275-86.
[0033] These various tubes can store up to 8.5 mL of blood, or in some cases up to 10 mL of blood. Thus, a blood sample collected from a subject can typically have a volume of 5-10 mL.
[0034] A 10 mL blood sample typically yields 10-500 ng of cfDNA, but in some cases, particularly in certain cancer patients, substantially higher amounts, such as up to about 10 μg, can be obtained. The methods disclosed herein can be performed on the amount of cfDNA contained in a 10 mL blood sample. The methods and compositions disclosed herein can typically use 10-400 ng of cfDNA, such as 10-250 ng or 10-200 ng of cfDNA.
[0035] Analysis of cfDNA derived from plasma is preferred. Kits for purifying cfDNA from plasma (and other body fluids) are readily available, such as the MagMAX cfDNA Isolation Kit from ThermoFisher, the Maxwell RSC ccfDNA Plasma Kit from Promega, the Apostle MiniMax High Efficiency Isolation Kit from Beckman Coulter, or the QIAamp or EZ1 products from Qiagen.
[0036] Accordingly, the methods and compositions disclosed herein can utilize cfDNA extracted from a biological fluid sample of a subject, typically a plasma or serum sample. The methods can start with cfDNA that has already been prepared or can include upstream steps for preparing cfDNA. Similarly, the methods can include upstream steps for obtaining a plasma sample prior to the step of preparing cfDNA from the plasma sample.
[0037] Preferably, the cfDNA utilized in the methods and compositions disclosed herein substantially does not contain single-stranded DNA (ssDNA), i.e., less than 7% of the cfDNA molecules (count units) are single-stranded, preferably less than 5% or less than 1% (i.e., such that at least 99% of the cfDNA molecules are double-stranded). In some embodiments, the cfDNA may contain less than 0.1% ssDNA, less than 0.01% ssDNA, or even may not contain ssDNA (i.e., is ssDNA-free). Extraction of cfDNA to obtain a cfDNA sample substantially free of ssDNA is described, for example, in International Publication No. WO 2020 / 188561. By ensuring low levels of ssDNA, potential inhibition of restriction digestion is avoided, and it is useful after digestion as ssDNA can interfere with downstream steps such as ligation and amplification. Commercially available kits for quantifying single-stranded in a sample, such as the Promega QuantiFluor™ kit, are available.
[0038] In some embodiments, all of the extracted cfDNA is used in the methods disclosed herein. In other embodiments, the cfDNA is divided into multiple fractions, and one or more of the fractions are not used in the methods disclosed herein but may instead be used in other analytical methods or retained for use in control experiments or for other purposes.
[0039] In some embodiments, the cfDNA is quantified prior to digestion (e.g., in units of weight, concentration, etc.). In other embodiments, the cfDNA is not quantified prior to digestion.
[0040] The cfDNA used with the methods and compositions disclosed herein can be obtained from any eukaryotic subject, such as a mammal, and is preferably obtained from a human subject. In some embodiments, the human subject may be known or suspected to have a disease (e.g., cancer). In other embodiments, the human subject may be known to be healthy. In some embodiments, the subject is not pregnant.
[0041] Restriction Enzymes and Digestion The methods and compositions disclosed herein use restriction enzymes that recognize specific sequences in double-stranded DNA and introduce double-stranded breaks into the DNA. The enzymes have recognition sites that contain CpG sequences. Type II restriction enzymes are particularly useful, i.e., enzymes in which the double-stranded break is introduced within the recognition site. By using multiple restriction enzymes, parallel and simultaneous digestion within a sample becomes possible.
[0042] More specifically, the methods and compositions disclosed herein use methylation-sensitive restriction enzymes and / or methylation-dependent restriction enzymes. MSREs cleave target DNA only when the CpG within their recognition site is not methylated, and methylation inhibits cleavage. Conversely, MDREs cleave target DNA only when the CpG within their recognition site is methylated. MSREs and MDREs are readily available from well-known commercial suppliers such as ThermoFisher, New England Biolabs, and Promega.
[0043] Examples of MSREs include, but are not limited to, AatII, AccII, AciI, AclI, AfeI, AgeI, Aor13HI, Aor51HI, AscI, AsiSI, AvaI, BceAI, BmgBI, BsaAI, BsaHI, BsiEI, BsiWI, BsmBI, BspDI, BspT104I, BssHII, BstBI, BstUI, Cfr10I, ClaI, CpoI, DpnII, EagI, Eco52I, FauI, FseI, FspI, HaeII, HapII, HgaI, HhaI, HinP1I, HpaII, Hpy99I, HpyCH4IV, KasI, MluI, NaeI, NarI, NgoMIV, NotI, NruI, NsbI, PaeR7I, PluTI, PmaCI, PmlI, Psp1406I, PvuI, RsrII, SacII, SalI, ScrFI, SfoI, SgrAI, SmaI, SnaBI, SrfI, TspMI, ZraI.
[0044] Examples of MDREs include, but are not limited to, BspEI, BtgZI, FspEI, GlaI, LpnPI, McrBC, MspJI, XhoI, XmaI.
[0045] The methods and compositions disclosed herein may comprise a plurality of restriction enzymes, which may consist of MSREs and / or MDREs. Thus, the plurality of restriction enzymes may include only MSREs, only MDREs, or a mixture of both (e.g., one or more MSREs + one or more MDREs). However, in general, it is preferred to act with MSREs without the need for MDREs, and thus the plurality of restriction enzymes includes two or more MSREs. By using MSREs, cfDNA is provided in which methylated CpG sites are intact but unmethylated CpG sites are digested. Thus, for any particular CpG-containing restriction site in a cfDNA sample, a higher percentage of methylation at this site results in a lower degree of digestion compared to a cfDNA sample containing a higher percentage of methylation at this site.
[0046] Preferred multiple MSREs include both HinP1I and AciI. In some embodiments, it is possible to use one or more MSREs in addition to HinP1I and AciI, but it is more preferred to use HinP1I and AciI as the only two restriction enzymes for the digestion of cfDNA. The pairing of these enzymes covers more than 99% of the CpG islands in the human genome. In this MSRE pairing, it is preferred to include HinP1I in an excess amount (measured in enzyme units) relative to AciI, and ideally at least 1.2:1 (i.e., at least 1.2 units of HinP1I for 1 unit of AciI), for example, at least 1.5:1, at least 1.75:1, at least 2:1, at least 3:1, at least 4:1, or at least 5:1. A ratio of 2:1 to 5:1 is particularly useful for human cfDNA, and an excess of about 4.5, for example, 4.4 to 4.6 is preferred. The digestion can be carried out at about 37 °C until completion. Incubation at 37 °C for 2 hours is typically sufficient for the complete digestion of cfDNA samples using HinP1I and AciI as described herein, but longer digestion may be used, for example, when digesting cfDNA obtained from blood stored in a blood collection tube containing an anticoagulant and an agent that inhibits the release of genomic DNA from white blood cells into the plasma component of the blood sample (see other parts of this specification).
[0047] In some embodiments, HhaI, AspLEI, or CfoI is used as one of the restriction enzymes for digestion of cfDNA. Each of HhaI, AspLeI, and CfoI is an MSRE, and each recognizes the same recognition sequence as HinP1I, but does not necessarily cleave at the same cleavage site within the recognition sequence. In some embodiments, SsiI is used as one of the restriction enzymes for digestion of cfDNA. SsiI is an MSRE and recognizes the same recognition sequence as AciI, but does not necessarily cleave at the same cleavage site within the recognition sequence. In some embodiments, HinP1I and SsiI are used as the only restriction enzymes for digestion of cfDNA. In some embodiments, AciI and any one of HhaI, AspLEI, and CfoI are used as the only restriction enzyme for digestion of cfDNA. Advantageously, AciI and HinP1I can be completely inactivated by heating to 65° C., while HhaI and AspLEI are insensitive to heat inactivation. In addition, both AciI and HinP1I use the same optimal buffer for 100% activity, but the optimal reaction buffers recommended for each of SsiI, AspLEI, and CfoI are different, making it more difficult to adjust the digestion conditions for these alternative enzyme combinations without compromising enzyme activity.
[0048] In some embodiments, any combination of MSREs that recognize the same recognition sequence as Hinp1I or AciI, but do not necessarily cleave at the same cleavage site within the recognition sequence, can be used for digestion of cfDNA.
[0049] The concentration of the restriction enzyme can be selected according to a specific experiment in progress. Typically, HinP1I can be used at 10 - 450 units per μg of cfDNA, AciI can be used at 2.5 - 100 units per μg of cfDNA, and for example, has a ratio of 4.5 units of HinP1I per unit of AciI. In other embodiments, HinP1I can be used at 500 - 2500 units per μg of cfDNA, AciI can be used at 100 - 500 units per μg of cfDNA, and for example, has a ratio of 4.4 - 4.6 units (such as 4.5 units) of HinP1I per unit of AciI. Regarding the solution concentration, HinP1I can be used at 35 - 45 units / ml, AciI can be used at 5 - 15 units / mL, and for example, has a ratio of 4.5 units of HinP1I per unit of AciI.
[0050] HinP1I recognizes the sequence GCGC and cuts after the first G, leaving a 2 - nucleotide 5’ overhang (5’-G / CGC). HinP1I cuts well at 37°C and can be heat - inactivated by heating at 65°C for 20 minutes. For HinP1I, NEB recommends the use of its rCutSmart™ buffer (50 mM potassium acetate, 20 mM Tris - acetate, 10 mM magnesium acetate, 100 μg / mL recombinant albumin, pH 7.9). One unit of HinP1I is defined as the amount of enzyme required to digest 1 μg of λDNA in a total reaction volume of 50 μl at 37°C for 1 hour.
[0051] Several commercial suppliers offer the Hin6I enzyme instead of HinP1I. These two enzymes have essentially the same properties, i.e., they have the same recognition sequence, the same optimal digestion temperature, and they can both be inactivated at 65°C in 20 minutes. Also, one unit of Hin6I is defined similarly to one unit of HinP1I. Thus, the terms HinP1I and Hin6I are used interchangeably herein, and any combination of enzymes disclosed as using HinP1I should be understood as disclosing the same combination using Hin6I instead. For example, the present invention provides a combination of AciI and Hin6I, as disclosed herein for AciI and HinP1I.
[0052] AciI recognizes the sequence CCGC and cuts after the first C, leaving a 2-nucleotide 5’ overhang (5’-C / CGC). AciI cuts well at 37°C and can be heat inactivated by heating at 65°C for 20 minutes. For AciI, NEB recommends the use of its rCutSmart™ buffer (50 mM potassium acetate, 20 mM Tris-acetate, 10 mM magnesium acetate, 100 μg / mL recombinant albumin, pH 7.9). One unit of AciI is defined as the amount of enzyme required to digest 1 μg of λDNA in 1 hour at 37°C in a total reaction volume of 50 μl. Its recognition site is non-palindromic.
[0053] λDNA is a commonly used DNA substrate extracted from bacteriophage lambda (cI857ind 1 Sam 7) and is 48502 bp in length. It is usually stored in 10 mM Tris-HCl (pH 8.0), 1 mM EDTA and is widely available from commercial suppliers, e.g., NEB, catalog number N3011S.
[0054] Since HinP1I and AciI share essentially the same conditions for digestion and inactivation, they create pairs useful for digesting DNA. In contrast, enzymes such as HpaII, AvaI, HaeII, and MluI require heating to 80 °C for inactivation. BstUI and PvuI, as well as HhaI, are less affected by heat inactivation. BstUI cuts optimally at 60 °C and shows only 10% of its full activity at 37 °C. PvuI shows only 10% of its full activity in NEB's rCutSmart™ buffer.
[0055] After digestion, it is preferable to inactivate the restriction enzyme, especially when using downstream amplification steps such as PCR. Heat inactivation is particularly preferred, and both HinP1I and AciI can be inactivated by heating the composition at 65 °C for at least 20 minutes, for example, 20 - 60 minutes. Further details on inactivation are shown below.
[0056] Other useful enzyme combinations include, or consist of, (i) HinP1I + AciI + McrBC, (ii) HinP1I + AciI + MspJI, (iii) HinP1I + AciI + HpaII + HpyCH4IV + BstUI, (iv) HinP1I + AciI + HpaII + HpyCH4IV + AvaI, (v) MspJI + FspEI, (vi) MspJI + HinP1I + AciI, (vii) MspJI + FspEI + HinP1I + AciI, or (viii) MspJI + FspEI + HinP1I + AciI + HpyCH4IV.
[0057] Other useful combinations of enzymes include, or consist of, MSREs that recognize the same recognition sequences as HinP1I and / or AciI, but do not necessarily cleave at the same cleavage sites within the recognition sequences, and include (i) AciI+HhaI, (ii) AciI+AspLEI, (iii) AciI+CfoI, (iv) SsiI+HinP1, (v) SsiI+HhaI, (vi) SsiI+AspLEI, (vii) SsiI+CfoI, (viii) SsiI+HinP1I. MspJI shares essentially the same conditions as HinP1I and AciI with respect to digestion and inactivation (e.g., it is active at 37° C. in rCutSmart™ and can be inactivated at 65° C.). These three enzymes are particularly useful because they can provide 85% CpG coverage and 100% CpG island coverage.
[0058] Two further useful combinations include, or consist of, (i) HinP1I+AciI+HpaII, or (ii) HinP1I+AciI+HpaII+HpyCH4IV. For these two combinations, the methods and compositions of the invention should use at least one of the following additional features, as discussed elsewhere herein: (a) HinP1I is used in excess relative to AciI in terms of enzyme units, (b) digestion is carried out for no more than 11 hours, (c) the digested cfDNA is subjected to sequencing.
[0059] Other useful combinations of enzymes include (i) AciI and one of AfeI, Aor51HI, AscI, BssHII, PauI, HaeII, Eco47III, EheI, FspAI, GIaI, KasI, MteI, NarI, NsbI, PluTI, SfoI, or SgsI; (ii) BsrBI and one of AfeI, Aor51HI, AscI, BssHII, PauI, HaeII, Eco47III, EheI, FspAI, GIaI, KasI, MteI, NarI, NsbI, PluTI, SfoI, or SgsI; (iii) MbiI and one of AfeI, Aor51HI, AscI, BssHII, PauI, HaeII, Eco47III, EheI, FspAI, GIaI, KasI, MteI, NarI, NsbI, PluTI, SfoI, or SgsI; (iv) NotI and one of AfeI, Aor51HI, AscI, BssHII, PauI, HaeII, Eco47III, EheI, FspAI, GIaI, KasI, MteI, NarI, NsbI, PluTI, SfoI, or SgsI; (v) SacII and one of AfeI, Aor51HI, AscI, BssHII, PauI, HaeII, Eco47III, EheI, FspAI, GIaI, KasI, MteI, NarI, NsbI, PluTI, SfoI, or SgsI; (vi) Cfr42I and one of AfeI, Aor51HI, AscI, BssHII, PauI, HaeII, Eco47III, EheI, FspAI, GIaI, KasI, MteI, NarI, NsbI, PluTI, SfoI, or SgsI; (vii) SgrBI and one of AfeI, Aor51HI, AscI, BssHII, PauI, HaeII, Eco47III, EheI, FspAI, GIaI, KasI, MteI, NarI, NsbI, PluTI, SfoI, or SgsI, and include, or consist of, MSREs that recognize recognition sequences containing recognition sequences of HinP1I and / or AciI.
[0060] Other useful combinations of enzymes are: (i) HinP1I and one of BsrBI, MbiI, NotI, SacII, Cfr42I, or SgrBI; (ii) AfeI and one of BsrBI, MbiI, NotI, SacII, Cfr42I, or SgrBI; (iii) Aor51HI and one of BsrBI, MbiI, NotI, SacII, Cfr42I, or SgrBI; (iv) AscI and one of BsrBI, MbiI, NotI, SacII, Cfr42I, or SgrBI; (v) BssHII and one of BsrBI, MbiI, NotI, SacII, Cfr42I, or SgrBI; (vi) PauI and one of BsrBI, MbiI, NotI, SacII, Cfr42I, or SgrBI; (vii) HaeII and one of BsrBI, MbiI, NotI, SacII, Cfr42I, or SgrBI; (viii) Eco47III and one of BsrBI, MbiI, NotI, SacII, Cfr42I, or SgrBI; (ix) EheI and one of BsrBI, MbiI, NotI, SacII, Cfr42I, or SgrBI; (x) FspAI and one of BsrBI, MbiI, NotI, SacII, Cfr42I, or SgrBI; (xi) GIaI and one of BsrBI, MbiI, NotI, SacII, Cfr42I, or SgrBI; (xii) KasI and one of BsrBI, MbiI, NotI, SacII, Cfr42I, or SgrBI; (xiii) Mtel and one of BsrBI, MbiI, NotI, SacII, Cfr42I, or SgrBI; (xv) Narl and one of BsrBI, MbiI, NotI, SacII, Cfr42I, or SgrBI; (xvi) NsbI and one of BsrBI, MbiI, NotI, SacII, Cfr42I, or SgrBI; (xvii) PIutI and one of BsrBI, MbiI, NotI, SacII, Cfr42I, or SgrBI; (xviii) SfoI and one of BsrBI, MbiI, NotI, SacII, Cfr42I, or SgrBI; (ixx) SgsI and one of BsrBI, MbiI, NotI, SacII, Cfr42I,comprises or consists of an MSRE that recognizes a recognition sequence comprising a recognition sequence of HinP1I and / or AciI and comprising one of SgrBI or the like.
[0061] Where a method is described herein as involving "digestion", this term (and also "digesting", etc.) refers to the mixing of a restriction enzyme and DNA under conditions where digestion can occur. If the recognition site of the restriction enzyme is absent (e.g., if it is an MSRE and all of the recognition sequences are fully methylated), the "digestion" step is still carried out even if DNA cleavage does not occur.
[0062] method Disclosed herein are various methods for digesting cfDNA using combinations of restriction enzymes (e.g., combinations of MSREs).
[0063] The enzyme and cfDNA are typically incubated for a period of time long enough to effect substantially complete digestion, i.e., further incubation does not result in any measurable increase in cfDNA cleavage. For a typical sample, this can be achieved by incubation at 37°C for 2 hours, but if desired, longer digestions, e.g., 3 hours, 4 hours, or longer times (e.g., overnight) can be performed. In some embodiments, the digestion is carried out for 11 hours or less. Thus, in some embodiments, the digestion can be carried out for 2 - 11 hours, e.g., 2 - 10 hours, 2 - 9 hours, 2 - 8 hours, or 2 - 4 hours. In other embodiments (e.g., when using collection tubes as discussed herein), the digestion can be carried out for a longer period, e.g., 12 hours or more.
[0064] After digestion, especially when using downstream amplification steps, it is preferred to inactivate the restriction enzymes. Both HinP1I and AciI can be inactivated by heating them to 65°C, for example, by immersing the reaction mixture in a 65°C water bath. The digestion reaction mixture with cfDNA tends to have a low volume so that the temperature of the entire reaction mixture reaches 65°C very rapidly, resulting in enzyme inactivation. In some embodiments, the heating at this temperature is carried out for longer than 15 minutes, ideally for at least 20 minutes, for example, 20 - 60 minutes. The temperature can exceed 65°C if desired, but this is not essential. This heating step is sufficient for complete inactivation of the restriction enzyme, i.e., sufficient such that the digestion activity of the enzyme against the target cfDNA molecules that were cleavable under the digestion conditions used prior to heating can no longer be measurably detected.
[0065] The present invention also provides a method for analyzing cfDNA, comprising digestion of cfDNA as discussed above, followed by downstream analysis steps such as amplification steps (PCR, especially real-time PCR, etc.), ligation steps (ligation of sequencing adapters, etc.), DNA sequencing steps, etc. See also below.
[0066] The present invention also provides a method for evaluating the methylation status of one or more CpG sites in cfDNA, comprising digestion of cfDNA as discussed above, followed by a downstream analysis step of quantifying the extent of digestion at one or more CpG sites. The extent of digestion can be determined individually for each site or can be determined in total.
[0067] The present invention also provides a method for diagnosing the presence or absence of cancer in a subject, the method comprising evaluating the methylation status of one or more CpG sites in cfDNA as discussed above, wherein hypermethylation and / or hypomethylation of the one or more CpG sites is associated with cancer. In some embodiments, the method includes creating a report in paper or electronic form based on the evaluation of the presence or absence of cancer, and optionally communicating the report to the subject and / or the subject's healthcare provider.
[0068] The present invention also provides a method for treating or managing cancer in a subject, the method comprising diagnosing the presence of cancer as described above and administering a suitable anti-cancer treatment to the subject. The treatment may include one or more of surgical resection, chemotherapy, radiation therapy, immunotherapy, and / or targeted therapy.
[0069] Preferred methods do not include a bisulfite conversion step. Other preferred methods do not include a step of chemically modifying the nucleobases in the DNA, for example, bisulfite conversion, TAPS conversion, etc. are not performed. TAPS conversion refers to TET-assisted pyridine borane sequencing.
[0070] Preferred methods do not use restriction enzyme isoschizomers in which one of the enzymes recognizes both the methylated and non-methylated forms of a restriction site, while the other enzyme recognizes only one of these forms.
[0071] Preferred methods do not use a mixture of restriction enzymes in which at least one enzyme is a restriction enzyme that contains CpG but is neither an MSRE nor an MDRE, i.e., an enzyme that digests regardless of the CpG methylation status.
[0072] Some methods do not include a step of heating the sample containing purified cfDNA prior to digestion. Other preferred methods do not include such a pre-digestion heating step that includes heating the sample above 40°C, above 50°C, above 60°C, above 70°C, or above 80°C. Other preferred methods do not include a pre-digestion heating step that includes heating the sample at 80°C or higher for 20 minutes or more.
[0073] Composition A variety of compositions containing multiple restriction enzymes (e.g., multiple MSREs) are disclosed herein. They are typically aqueous compositions containing the enzyme in a soluble active form, along with other components such as salts, buffers, cofactors, etc.
[0074] These compositions may contain salts and / or buffers in an aqueous solution. For example, the composition may contain 50 mM potassium acetate, 20 mM Tris-acetate, 10 mM magnesium acetate, 100 μg / mL recombinant albumin, pH 7.9 (i.e., the composition of commercially available rCutSmart™ buffer). Alternatively, the composition may contain 50 mM Tris-HCl, 10 mM MgCl₂, 100 mM NaCl, 100 μg / mL recombinant albumin, pH 7.9 (i.e., the composition of commercially available NEBuffer™ r3.1 product). The pH is measured at 25 °C.
[0075] The composition may contain cfDNA, especially when used for digestion. As discussed above, in some compositions, HinP1I is present at 10 - 450 units per μg of cfDNA, AciI is present at 2.5 - 100 units per μg of cfDNA, and has a ratio of, for example, 4.5 units of HinP1I per unit of AciI. In other embodiments, HinP1I can be used at 500 - 2500 units per μg of cfDNA, AciI can be used at 100 - 500 units per μg of cfDNA, and has a ratio of, for example, 4.4 - 4.6 units (such as 4.5 units) of HinP1I per unit of AciI. With respect to the solution concentration, HinP1I can be present at 35 - 45 units / ml, AciI can be present at 5 - 15 units / mL of cfDNA, and has a ratio of, for example, 4.5 units of HinP1I per unit of AciI.
[0076] Accordingly, one useful composition of the present invention comprises HinP1I and AciI (e.g., having an excess amount of HinP1I as described herein), potassium acetate, Tris-acetate, magnesium acetate, albumin, pH 7.8 - 8.0 (and optionally, cfDNA to be digested). For example, the composition can comprise 4 - 5 units of HinP1I, 0.5 - 1.5 units of AciI, 50 mM potassium acetate, 20 mM Tris-acetate, 10 mM magnesium acetate, 100 μg / mL albumin, pH 7.9, and cfDNA.
[0077] The restriction enzymes in the composition are preferably present in an enzymatically active form, which enables digestion of cfDNA by the use of the restriction enzymes. However, after digestion, the composition can be heated (e.g., to 65°C) to inactivate the enzymes, and thus, in some embodiments, the restriction enzymes are present in a heat-inactivated form.
[0078] In some embodiments, the composition can also comprise PCR reagents, such as suitable buffer / salt components (if necessary in addition to the buffer / salt that persists after digestion), DNA polymerase (such as Taq polymerase), dNTPs, primers, probes, etc.
[0079] In some embodiments, the composition can also comprise sequencing reagents, such as one or more of sequencing adapters, DNA ligase (such as T4 ligase), Klenow fragment of DNA polymerase I, A-tailing enzyme (such as Taq polymerase), blunt-end polymerase (such as T4 DNA polymerase), kinase (such as T4 polynucleotide kinase), etc.
[0080] In some embodiments, the composition can also comprise control DNA as discussed below.
[0081] As described above, when the composition contains HinP1I and AciI, HinP1I is preferably present in an excess (measured in enzyme units) relative to AciI, and is preferably present in an excess of at least 1.2:1, for example at least 1.5:1, at least 1.75:1, at least 2:1, at least 3:1, at least 4:1, or at least 5:1. For example, when analyzing human cfDNA, a ratio of at least 2:1 is often useful, and when digesting human cfDNA from plasma, a ratio of about 4.5:1 has been found to be useful.
[0082] Preferred compositions do not contain restriction enzyme isoschizomers where one enzyme recognizes both methylated and non-methylated forms of the restriction site and another enzyme recognizes only one of these forms.
[0083] Preferred compositions do not contain a mixture of restriction enzymes where at least one enzyme is a restriction enzyme that contains CpG but is neither an MSRE nor an MDRE, i.e., an enzyme that digests regardless of the CpG methylation state.
[0084] Downstream amplification After digestion, the methods disclosed herein may include a step of amplification (e.g., PCR) performed on the digested cfDNA. Typically, this amplification is targeted to one or (preferably) multiple loci of interest, e.g., loci containing CpG sites whose methylation status is known or predicted to be associated with a particular biological state (e.g., a cancer of interest). Thus, upstream and downstream primers flanking the CpG site of interest are used, and intervening CpG-containing sequences are amplified if they have not been digested by the restriction enzyme. The resulting amplicons can then be detected, for example, using a labeled probe complementary to a subsequence within the amplicon of interest.
[0085] Accordingly, the method can include the step of adding, after digestion, PCR reagents, such as a suitable buffer / salt component (if necessary in addition to the buffer / salt remaining from digestion), a DNA polymerase (such as Taq polymerase), dNTPs, primers, and optionally a probe. Alternatively, one or more of these components may be present during digestion, for example, it is possible to use a hot start PCR protocol, and thus, the PCR reagents are already present during the digestion step, but they do not become active until the reaction mixture is heated (for example, during heat inactivation of the restriction enzyme).
[0086] Restriction digestion is typically carried out in the presence of a high level of Mg ++ PCR usually depends on Mg ++ Therefore, a standard PCR buffer contains Mg ++ However, in this situation, the addition of a standard PCR buffer can result in an excess amount of Mg ++ which can inhibit the efficiency of amplification. Thus, the PCR reagents added may contain a lower level of Mg ++ than normal.
[0087] If PCR primers and probes are present during MSRE digestion, they should be designed such that their sequences do not contain the recognition site of the MSRE being used.
[0088] Amplification and detection of amplicons can be performed by conventional PCR using fluorescently labeled primers followed by capillary electrophoresis of the amplification products. In some embodiments, after amplification, the amplification products are separated by capillary electrophoresis and the fluorescence signal is quantified. An electropherogram can be generated that plots the change in the fluorescence signal as a function of size (bp) or time from injection, and each peak in the electropherogram corresponds to the amplification product of a single locus. The height of the peak (provided, for example, using "relative fluorescent unit" (rFU)) can represent the intensity of the signal from the amplified locus. Computer software can be used to detect the peaks, calculate the fluorescence intensity (peak height) of a series of loci whose amplification products were run on a capillary electrophoresis machine, and then calculate the ratio between the signal intensities.
[0089] A preferred PCR technique is real-time PCR (also known as qPCR) in which co-amplification and detection of the amplification products is performed. Real-time PCR can be used with non-specific detection or sequence-specific detection. Non-specific detection (using, for example, a dsDNA-binding dye such as SYBR Green) can be used within the methods disclosed herein, but is not ideal if it is desired to distinguish multiple different amplicons in the same reaction. Thus, it is more typical to use sequence-specific detection, and the methods and compositions can use labeled oligonucleotide probes that are complementary to a specific sequence within the nucleic acid amplicon of interest (usually having a fluorophore and a fluorescence quencher on the same probe, as in the TaqMan system). Different probes for amplicons derived from different target CpGs can be labeled with different fluorophores so that multiple different amplicons can be distinguished.
[0090] Thus, real-time PCR can be achieved by using hydrolysis probes based on a combination of reporter molecules and quencher molecules. In such an assay, the oligonucleotide probe has a fluorescent moiety (fluorophore) attached to its 5' end and a quencher attached to its 3' end. During PCR amplification, the polynucleotide probe selectively hybridizes to its target sequence on the template, and when the polymerase replicates the template, the 5'-nuclease activity of the polymerase also cleaves the polynucleotide probe. When the polynucleotide probe is intact, the very close proximity between the quencher and the fluorescent moiety usually results in a low level of background fluorescence. When the polynucleotide probe is cleaved, the quencher is separated from the fluorescent moiety and the fluorescence intensity increases. The fluorescence signal correlates with the amount of amplification product, i.e., the signal increases as the amplification product accumulates.
[0091] Suitable fluorophores include, but are not limited to, fluorescein, FAM, lysamine, phycoerythrin, rhodamine, Cy2, Cy3, Cy3.5, Cy5, Cy5.5, Cy7, FluorX, JOE, HEX, NED, VIC, and ROX. Suitable fluorophore / quencher pairs, including but not limited to FAM-TAMRA, FAM-BHQ1, Yakima Yellow-BHQ1, ATTO550-BHQ2, and ROX-BHQ2, are known in the art.
[0092] Fluorescence can be monitored during each PCR cycle, providing an amplification plot showing the change in fluorescence signal from the probe as a function of the number of cycles. The following terms are used with respect to real-time PCR.
[0093] The "quantification cycle" ("Cq") refers to the number of cycles in which fluorescence increases automatically by software or is manually set by the user beyond a threshold. In some embodiments, the threshold can be constant for each CpG locus of interest and can be preset before performing amplification and detection. In other embodiments, the threshold can be defined separately for each CpG locus after execution based on the maximum fluorescence level detected for this locus during the amplification cycle.
[0094] The "threshold" refers to the value of fluorescence used for Cq determination. In some embodiments, the threshold can be a value exceeding the baseline fluorescence and / or a value exceeding the background noise, and can be a value within the exponential growth phase of the amplification plot.
[0095] The "baseline" refers to the initial cycles of PCR where there is little or no change in fluorescence.
[0096] Computer software for analyzing the amplification plot and determining the baseline, threshold, and Cq is readily available.
[0097] If the CpG site is not digested and thus amplified in subsequent PCR, a detectable amplification product accumulates after a relatively small number of amplification cycles, resulting in a relatively low Cq value. Conversely, if the amplicon is present at a lower level (e.g., because some CpG loci of interest are digested), fewer amplicons are seen and the Cq value is higher.
[0098] Therefore, these results can indicate, for any given CpG site, the proportion of cfDNA molecules in the methylated / unmethylated sample at that CpG site. These numbers can be expressed as percentages, fractions, normalized values, etc.
[0099] Primers can vary in length depending on the specific assay format and specific requirements. In some embodiments, the primer can be at least 15 nucleotides in length, for example, 15-25 nucleotides in length or 18-25 nucleotides in length. The primer can be adapted to suit the selected amplification system.
[0100] Primers can be designed to generate amplicons that are 60-150 bp in length (when the relevant CpG site is intact), for example, amplicons that are 70-140 bp in length.
[0101] Oligonucleotide probes can vary in length. In some embodiments, the probe can comprise 15-30 nucleotides, 20-30 nucleotides, or 25-30 nucleotides.
[0102] Oligonucleotide probes can be designed to bind to either strand of the double-stranded amplicon. Additional considerations include the melting temperature of the probe, which should preferably be compatible with the melting temperature of the primer.
[0103] When multiple CpG sites are analyzed in parallel by co-amplification of two or more targets in the same reaction mixture using different primer pairs for each CpG site of interest, these different primers can be designed to act at the same annealing temperature during amplification. Thus, for example, primers with similar melting temperatures (Tm) can be designed within 3°C to 5°C of each other. Similar considerations apply when multiple probes are used. + 3°C to 5°C of each other, primers with similar melting temperatures (Tm) can be designed. Similar considerations apply when multiple probes are used.
[0104] Computer software for the routine design of primers and probes that meet the various requirements of any particular experiment is readily available.
[0105] Downstream sequencing After digestion, the methods disclosed herein can include steps of DNA sequencing, such as steps using next-generation sequencing (“NGS”) technology (also known as high-throughput sequencing). NGS generally involves three basic steps: library preparation, sequencing, and data processing. Examples of NGS technologies include sequencing-by-synthesis and sequencing-by-ligation (e.g., employed by Illumina Inc., Life Technologies Inc., PacBio, and Roche), nanopore sequencing methods, and electronic detection-based methods such as Ion Torrent™ technology (Life Technologies Inc.). NGS can be performed using a variety of high-throughput sequencing instruments and platforms, including but not limited to Novaseq™, Nextseq™, and MiSeq™ (Illumina), 454 Sequencing (Roche), Ion Chef™ (ThermoFisher), SOLiD® (ThermoFisher), and Sequel II™ (Pacific Biosciences). Appropriate platform-designed sequencing adapters are used to prepare sequencing libraries and are readily available from the manufacturers of the platforms.
[0106] Library preparation for the major high-throughput sequencing platforms involves ligation of specific adapter oligonucleotides (also referred to as “sequencing adapters”) to the DNA fragments to be sequenced. Sequencing adapters typically include platform-specific sequences for fragment recognition by a particular sequencer, e.g., sequences that enable ligated molecules to bind to a flow cell of an Illumina platform (e.g., P5 and P7 sequences). The suppliers of each sequencing instrument typically sell a specific set of sequences for this purpose. Further details of library preparation are considered below.
[0107] The sequencing adapter may include a site for binding to a universal set of PCR primers. This allows multiple adapter-ligated DNA molecules to be amplified in parallel by PCR using a single set of primers.
[0108] The sequencing adapter may include a sample index, which is a sequence that allows multiple samples to be combined and then sequenced together (i.e., multiplexed) on the same flow cell or chip of the instrument. Each sample index, typically 6 - 10 nucleotides, is specific to a given sample and is used to demultiplex and assign individual sequence reads to the correct sample during downstream data analysis. The sequencing adapter may contain single or dual sample indexes depending on the number of libraries to be combined and the desired level of accuracy.
[0109] The sequencing adapter may include a unique molecular identifier (UMI) to provide molecular tracking, error correction, and improved accuracy during sequencing. The UMI is typically a short sequence 5 - 20 bases in length and is used to uniquely identify the original molecules in the sample library. Since each nucleic acid in the starting material is tagged to provide a unique molecular barcode, bioinformatics software can filter out duplicate reads and PCR errors with a high level of accuracy, report unique reads, and remove identified errors prior to final data analysis.
[0110] In some embodiments, the sequencing adapter includes both a sample barcode sequence and a UMI.
[0111] In some embodiments, the sequencing adapter enables paired-end sequencing.
[0112] In some embodiments, the compositions and methods disclosed herein use a Y-shaped sequencing adapter, i.e., an adapter consisting of two single-stranded oligonucleotides that anneal to provide a double-stranded stem and two single-stranded "arms". In other embodiments, the compositions and methods disclosed herein use a hairpin sequencing adapter, i.e., a single-stranded oligonucleotide in which the 5' and 3' ends anneal to provide a double-stranded stem. For both the Y-shaped adapter and the hairpin adapter, the double-stranded stem can include a short single-stranded overhang, e.g., a single A or T nucleotide. For both the Y-shaped adapter and the hairpin adapter, the double-stranded stem can be ligated to cfDNA fragments to prepare a sequencing library.
[0113] Thus, sequencing adapters suitable for use in the compositions and methods disclosed herein can be TruSeq™ or AmpliSeq™ or TruSight™ adapters (for use on the Illumina platform) or SMRTbell™ adapters (for use on the PacBio platform).
[0114] When the sequencing adapter is added by ligation, this is typically done at both ends of the DNA to be sequenced.
[0115] Restriction digestion can leave blunt ends but typically generates single-stranded overhangs. The library preparation step can preserve this overhang (i.e., add complementary nucleotides) or remove it. Since the sequence of the post-digestion terminal single-stranded overhang can contain useful information, it is preferred to add the sequencing adapter so as to preserve the overhang, e.g., using enzymatic ligation where a ligase enzyme covalently links the sequencing adapter to the DNA fragment (where the terminal sequence of the adapter is complementary to the terminal sequence obtained using a restriction enzyme), or by adding complementary nucleotides using a polymerase to generate blunt-ended fragments.
[0116] In addition to the removal or filling of single-stranded overhangs, the end repair method performed prior to adapter ligation can ensure that the DNA molecule contains a 5' phosphate group and a 3' hydroxyl group.
[0117] For some libraries, the incorporation of non-template deoxyadenosine 5'-monophosphate (dAMP) into the 3' ends of blunt-ended DNA fragments is used in library preparation (a process known as dA tailing). The dA tail prevents concatemer formation during downstream ligation steps and allows the DNA fragments to ligate to adapter oligonucleotides having complementary dT overhangs.
[0118] As described above, restriction digestion is typically performed in the presence of a high level of Mg ++ . Since sequencing library preparation may also be Mg ++ -dependent, standard library preparation buffers contain Mg ++ . However, in this situation, the addition of standard library preparation buffer can result in an excess amount of Mg ++ , which can inhibit the efficiency of downstream steps. Thus, the reagents added may contain a lower level of Mg ++ than is normal for library preparation.
[0119] As an alternative approach to using a lower level of Mg ++ , it is possible to add a chelating agent after digestion, which can eliminate the need for removal or dilution of excess Mg ++ for downstream amplification steps. It has been found that the addition of the chelating agent at the concentrations disclosed herein does not impair such amplification steps or subsequent sequencing. A chelating agent can be added to provide an amplification reaction mix containing the chelating agent and the divalent cation in a molar ratio of 1:20 to 2:1. For example, the reaction mix can contain 8-20 mM Mg ++, for example, may contain about 10 mM of magnesium. For example, the amplification can be carried out in a reaction mix containing 3-4 mM of a chelating agent and 4 mM of Mg ++ and can be carried out in a reaction mix containing one or both of EDTA and EGTA. The chelating agent may contain one or both of EDTA and EGTA.
[0120] After library preparation, the prepared DNA molecules can be sequenced to provide a plurality of "sequence reads". These sequence reads can then be subjected to data processing, for example, to remove sequences that do not meet the desired quality criteria, remove duplicates, correct sequencing errors, map the sequences onto a reference genome, count the number of sequence reads, and so on. Computer software for performing these steps is readily available.
[0121] Any particular CpG site can be characterized by a plurality of sequence reads, which can be sequence reads derived from the same original cfDNA molecule and / or from different cfDNA molecules spanning the same CpG site. Sequencing is preferably performed such that the CpG site of interest is seen in at least 100 sequence reads, for example, at least 200, 300, 400, 500, 600, 700, or more sequence reads. The number of sequence reads spanning a particular genomic locus (e.g., a particular CpG site) is referred to as the sequencing "depth" or "coverage". The average number of sequence reads spanning a particular genomic locus (e.g., a particular CpG site) is referred to as the average sequencing "depth" or "coverage". The term "average" refers to the arithmetic mean. Thus, the average depth can be calculated by dividing the total number of nucleotides in the sequence reads mapped to the human genome by the length of the (diploid) human genome.
[0122] Higher sequencing depth improves accuracy by enabling better discrimination of signal from noise. This is because high-throughput DNA sequencing is error-prone (about 0.1 - 10% of all called bases are inaccurate), and thus sequence reads mapped to a particular locus often contain (apparent) mutations compared to the reference sequence at that site and / or other reads mapped to that site.
[0123] High depth is useful for determining whether differences in sequence reads relative to the reference sequence reflect the underlying sequence (signal) of the sample DNA or are due to errors (noise) during sequencing. Thus, high depth helps to draw meaningful conclusions from sequencing data, particularly with respect to the presence of rare signals such as those arising from tumor DNA. Methylated cfDNA molecules from tumors can be present in plasma in amounts as low as less than 1% of total cfDNA. Bisulfite sequencing does not provide sufficient depth across a sufficient number of genomic sites to reliably detect these rare signals. In contrast, the method of the present invention examines CpG sites in the entire human genome at a very high average depth and thus enables detection of these rare signals.
[0124] Array reads can be mapped to a previously identified genomic sequence, whether partial or complete, that is assembled as a representative example of a species or subject, the reference genome. The reference genome is typically diploid and typically does not represent the genome of a single individual of the species, but rather is a mosaic of the genomes of several individuals. The reference genome for the methods of the present invention is typically a human reference genome such as the human genome assembly available on the website of the National Center for Biotechnology Information or the University of California, Santa Cruz, Genome Browser, for example, the complete human genome. An example of a reference genome suitable for human studies is the "hg18" genome assembly. Alternatively, the more recent GRCh38 primary assembly can be used (up to patch p13).
[0125] Mapping aligns the array reads to the reference genome to identify the position of the reads within the reference genome. Array reads that are aligned are designated as "mapped". The alignment process aims to maximize the likelihood of obtaining regions of sequence identity across the various sequences being aligned and to allow for mismatches, indels, and / or clipping of some short fragments at the two ends of the read. The number of array reads mapped to a particular genomic locus is referred to as the "read count" or "copy number" of this genomic locus. It is not necessary to map all of the array reads obtained. In fact, it is not uncommon for a portion of the array reads obtained in any given experiment to be unmappable.
[0126] The term "genomic locus" refers to a specific position within the genome and can include a single position (a single nucleotide at a defined position in the genome) or a stretch of nucleotides that begins and ends at a defined position in the genome. The specific position can be identified by the position of the molecule, i.e., by the chromosome and the number of start and end base pairs on the chromosome. The genomic locus for the purposes herein contains at least one CpG site.
[0127] When restriction digestion uses an MSRE, sequence reads spanning a particular CpG site are from molecules that were not digested, i.e., molecules that were methylated at that CpG site (by complete digestion). The methylation level of this CpG site can be calculated by dividing its read count by the expected read count for this site (e.g., the read count expected if it were fully methylated and thus not digested). The expected read count can be determined, for example, by (i) the read count of a control locus not cleaved by the restriction endonuclease, (ii) the average read count of a plurality of such control loci, or (iii) the read count of the same CpG site in an undigested control sample, optionally corrected for sequencing depth differences.
[0128] Alternatively, the expected read count for a CpG site can be determined as the sum of the read counts at this CpG site (indicating methylation) + the sum of the read counts whose ends map to this CpG site (indicating non-methylation), taking into account any end repair that occurred during library preparation if necessary.
[0129] To avoid double counting, non-methylated CpG sites can be considered as sequencing reads with their 5' ends mapping to the site, as sequencing reads with their 3' ends mapping to the site, or as half the total of sequencing reads with either their 5' or 3' ends mapping to the site. Some library preparation methods can result in depletion of small fragments, which then will not be sequenced (e.g., in CpG islands where the starting cfDNA molecule is cleaved by MSRE at two or more non-methylated sites, thus providing three or more restriction fragments, some of which are very small), and the observed number of non-methylated CpG sites can be lower than the true value in the original sample. This distortion can be addressed to some extent by using the larger of the number of reads with their 3' ends mapping to the site and the number of reads with their 5' ends mapping to the site (or by using the average).
[0130] Thus, these calculations can provide, for any given CpG site, the proportion of cfDNA molecules in the sample that are methylated at that CpG site. Conversely, similar calculations can provide the proportion of a particular CpG site that is not methylated. These numbers can be expressed as percentages, fractions, normalized values, etc.
[0131] One way to represent the coverage of a particular CpG site is referred to as "Hitspan100", which refers to the number of sequence reads that span a particular CpG position and have at least 50 nucleotides both upstream and downstream. For example, 90 Hitspan100 at a particular CpG site means that there are 90 sequence reads that span this site and have at least 50 nucleotides both upstream and downstream.
[0132] The methods disclosed herein do not require differential adapter tagging of methylated versus non-methylated DNA molecules. The same population of adapters can be used for all molecules.
[0133] Control The methods disclosed herein may utilize positive and negative controls. In some embodiments, parallel analysis can be performed on one or more of the following: · A DNA control that does not contain the recognition sequence of the restriction enzyme used for digestion. If this DNA is digested, this indicates that the method was not performed accurately. · A DNA control that contains the fully methylated recognition sequence of the restriction enzyme used for digestion. If this DNA is digested when the method uses only MSRE, this indicates that the method was not performed accurately (the opposite is true for MDRE). · A DNA control that contains the non-fully methylated recognition sequence of the restriction enzyme used for digestion. If this DNA is not digested when the method uses only MSRE, this indicates that the method was not performed accurately (the opposite is true for MDRE).
[0134] These DNA controls can also be used as reference points for the analysis, such as to check the completeness of digestion. As described above, for example, when fragments are obtained using MSRE digestion, knowing the predicted read count can be useful in downstream NGS experiments, and one way to obtain this value is to examine the read count of DNA that does not contain the recognition sequence of MSRE, or of DNA that contains the recognition sequence but is fully methylated.
[0135] For these purposes, the DNA control is preferably similar in size and composition to the cfDNA molecule containing the CpG site of interest. Thus, it is possible to use synthetic DNA or PCR amplicons or bacterial plasmid DNA as non-methylated controls, but these are more useful when they have a size similar to cfDNA (e.g., long synthetic DNA, or appropriately sized restriction fragments prepared from plasmids).
[0136] Control experiments can be performed internally or externally in the sample. For internal controls, the control DNA can already be present in the sample (e.g., cfDNA containing CpG sites that are known to be ubiquitously (un)methylated, or cfDNA that does not contain the recognition sequence of the restriction enzyme used), and / or can be added (e.g., synthetic DNA added to the cfDNA). Thus, the control DNA can be processed in combination with the cfDNA and experience the same conditions as the cfDNA, and the method can involve co-amplification of the restriction locus and the control locus. For external controls, the control DNA is subjected to the same treatment as the cfDNA, but not as part of the same reaction mixture.
[0137] Thus, control DNA such as cfDNA can be digested with a restriction enzyme and then subjected to downstream analysis steps such as amplification, DNA sequencing, etc. Real-time PCR of suitable control loci can yield results that can be used as reference points. For example, signals obtained from cfDNA at the CpG site of interest and from control DNA (especially control DNA not digested by the restriction enzyme used) can be compared, and since the signal ratio reflects the methylation ratio, the signal ratio can be used to determine the degree of methylation at the CpG site of interest. Thus, the method disclosed herein can be carried out by calculating the signal ratio between the analyzed genomic locus and the control, without the need to evaluate the absolute methylation level at the genomic locus. This is in contrast to some conventional methods of methylation analysis for distinguishing tumor-derived DNA from normal DNA, which require determining the actual methylation level at a specific genomic locus. Thus, the method disclosed herein can eliminate the need for standard curves and / or additional cumbersome steps that involved determining absolute methylation levels, thereby providing a simple and cost-effective procedure. An additional advantage of using an internal control is that the signal ratio is obtained for loci amplified in the same reaction mixture under the same reaction conditions, which can help to eliminate potential sources of error (e.g., differences in template concentration, enzymes, etc. between reaction mixtures).
[0138] Accordingly, a method using qPCR may involve calculating the signal intensity ratio between co-amplified CpG sites after digestion of DNA as disclosed herein, thereby providing the methylation status for the CpG sites. This methylation status can then be compared to a reference value (e.g., a value obtained from a healthy subject or a subject with a known disease), and a diagnostic result can be derived based on the comparison. Accordingly, the method may involve co-amplifying CpG sites and a control locus from DNA digested with a restriction endonuclease, thereby generating co-amplified products, determining the signal intensity of each generated co-amplified product, and calculating the ratio between the signal intensities of the co-amplified products of the CpG site and the control locus.
[0139] The ratio between the signal intensities of the co-amplified products can be calculated by determining the quantification cycle (Cq) for each locus and calculating 2 (Cq対照遺伝子座-Cq CpG部位) Thereby. In other words, determine the reduction in Cq relative to the control locus and use this value as the exponent of 2 to calculate the ratio.
[0140] Accordingly, it is possible to derive a numerical value representing the degree of methylation of a CpG site in a cfDNA sample based on the degree of digestion at any particular CpG site using qPCR or sequencing. This value can be represented in various ways, for example, as the ratio or percentage of cfDNA molecules methylated at the CpG site, or as the intensity of the signal obtained from a particular CpG site, or as the ratio between the CpG site and a control locus, etc.
[0141] Systems and Kits The invention also provides various systems and kits.
[0142] The system may include a computer processor for performing and / or controlling the methods disclosed herein and / or for processing the results, e.g., for performing calculations based on the results. At least partially computer-implemented methods are provided.
[0143] A system or kit can include a blood, plasma, or serum sample from a human subject, components for performing the methods disclosed herein on at least one CpG site, and computer software stored on a non-transitory computer-readable medium that can instruct a computer processor to determine methylation values for at least one CpG locus based on a methylation assay. The software can also link the methylation values to a diagnostic result or prediction, for example, by comparing one or more methylation values to one or more reference values, to evaluate the presence of a disease in the subject. The computer software can receive data from qPCR and / or NGS experiments.
[0144] Components for performing the methods disclosed herein include biochemical components (e.g., enzymes, primers, probes, NTPs, etc.), chemical components (e.g., buffers, reagents), and technical components (e.g., PCR systems such as real-time PCR systems, and instruments such as tubes, vials, plates, pipettes, etc.).
[0145] Based on the methylation values, the system can generate and / or communicate a report to the subject and / or the subject's healthcare provider.
[0146] The computer software includes processor-executable instructions stored on a non-transitory computer-readable medium. The computer software can also include stored data. The computer-readable medium is a tangible computer-readable medium such as a compact disc (CD), magnetic storage device, optical storage device, random access memory (RAM), read only memory (ROM), or any other tangible expression medium.
[0147] The computer-related methods and processes described herein, when executed, are implemented using software stored in non-volatile or non-transitory computer-readable instructions that configure or direct a computer processor or computer to execute the instructions.
[0148] Each of the systems, servers, computer devices, and computers described herein may be implemented on one or more computer systems and may be configured to communicate via a network. They may also all be implemented on a single computer system. In one embodiment, a computer system includes a bus or other communication mechanism for communicating information, and a hardware processor coupled to the bus for processing information.
[0149] The computer system also includes main memory, such as random access memory (RAM) or other dynamic storage device, coupled to the bus for storing information and instructions to be executed by the processor. The main memory may also be used to store temporary variables or other intermediate information during execution of instructions by the processor. When such instructions are stored on a non-transitory storage medium accessible to the processor, the computer system is rendered a special-purpose device customized to perform the operations specified in the instructions.
[0150] The computer system may include read-only memory (ROM) or other static storage device coupled to the bus for storing static information and instructions for the processor. A storage device such as a magnetic disk or optical disk is provided and coupled to the bus for storing information and instructions.
[0151] The computer system may be coupled to a display via the bus for displaying information to a user of the computer.
[0152] An input device that includes alphanumeric and other keys can be coupled to a bus to communicate information and command selections to a processor. Another type of user input device is cursor control such as a mouse, trackball, or cursor direction keys that communicate direction information and command selections to a processor and control the movement of a cursor on a display.
[0153] The methods disclosed herein can be performed by a computer system in response to a processor executing one or more sequences of one or more instructions included in a main memory. Such instructions can be read into the main memory from another storage medium, such as a storage device. Execution of the sequences of instructions included in the main memory causes the processor to perform the process steps described herein. In alternative embodiments, hardwired circuitry may be used in place of, or in combination with, software instructions.
[0154] Suitable storage media include any non-transitory media that store data and / or instructions for operating a machine in a particular manner. Common forms of storage media include, for example, floppy disks, flexible disks, hard disks, solid state drives, magnetic tape, or any other magnetic data storage medium, CD-ROM, any other optical data storage medium, any physical medium with hole patterns, RAM, PROM, and EPROM, FLASH-EPROM, NVRAM, any other memory chip, or cartridge.
[0155] A storage medium is different from, but can be used in combination with, a transmission medium. A transmission medium participates in the transfer of information between storage media. For example, transmission media include coaxial cables, copper wires, and fiber optics, including the wires that make up a bus.
[0156] The present invention also provides a kit comprising: (i) a composition comprising a plurality of restriction enzymes as contemplated above; and (ii) components for analyzing cfDNA digested with the composition. These components can be, for example, components for performing PCR or for preparing a sequencing library from the digested cfDNA. For example, the kit can comprise: (a) a buffer having, for example, 50 mM potassium acetate, 20 mM Tris-acetate, 10 mM magnesium acetate, 100 μg / mL recombinant albumin, pH 7.9, or 50 mM Tris-HCl, 10 mM MgCl₂, 100 mM NaCl, 100 μg / mL recombinant albumin, pH 7.9; (b) DNA polymerase, dNTPs, primers, and optionally one or more probes; (c) sequencing adapters; (d) an enzyme solution comprising DNA ligase and / or DNA polymerase; and / or (e) one or more of control DNAs. Further details of these components (a)-(e) are discussed elsewhere in this specification.
[0157] The kit can include instructions for carrying out the methods as disclosed herein.
[0158] The kit can include a non-transitory computer-readable medium storing computer software that, when executed, configures or instructs a computer processor to perform the method steps as disclosed herein.
[0159] Disclaimer In some cases, the disclosure of International Publication No. 2022 / 107145 (PCT / IL2021 / 051382) is excluded.
[0160] In some embodiments, the compositions and methods disclosed herein do not use a mixture of cfDNA from 50 to 60 individuals.
[0161] In some embodiments, the compositions and methods disclosed herein do not use 760 to 810 ng of cfDNA, particularly 760 to 810 ng of cfDNA from healthy patients.
[0162] In some embodiments, the compositions and methods disclosed herein do not use 26 ng or 94 ng of cfDNA from untreated non-small cell lung cancer patients.
[0163] In some embodiments, the compositions and methods disclosed herein do not use 25 - 95 ng of cfDNA from untreated non-small cell lung cancer patients.
[0164] In some embodiments, the compositions and methods disclosed herein do not use a panel consisting of eight CpG sites located at positions chr1-11397653, chr17-17362652, chr17-71690026, chr3-121760779, chr12-49705230, chr1-8120128, chr2-39309230, and chr12-84283776 (disclosed in Table 4 of PCT / IL2021 / 051382) in the hg18 human genome assembly.
[0165] In some embodiments, the compositions and methods disclosed herein do not use a mixture of 10 units of HinP1I and 5 units of AciI.
[0166] In some embodiments, the compositions and methods disclosed herein do not use an activity ratio of HinP1I:AciI of 2:1.
[0167] Generally The practice of the present invention uses conventional methods of chemistry, biochemistry, and molecular biology within the skill of the art, unless otherwise indicated. Such techniques are well described in the literature. See, for example, Methods In Enzymology (Academic Press, Inc.), Green & Sambrook (2012) Molecular Cloning: A Laboratory Manual, 4th edition (Cold Spring Harbor Press), Ausubel et al. (eds) Short protocols in molecular biology, 5th edition (Current Protocols), Molecular Biology Techniques: An Intensive Laboratory Course, (Ream & Field, eds., 1998, Academic Press), Wilson and Walker’s Principles and Techniques of Biochemistry and Molecular Biology (Hodmann & Clokie, 2018), Basic Molecular Biology & Techniques - Recent Advances: Molecular Biology & Its Technique (Singh et al., 2021), etc.
[0168] The term "comprising" encompasses "including" and "consisting of"; for example, a composition "comprising" X can consist of only X or can contain some additional, for example X+Y.
[0169] The term "about" in relation to a numerical value x is optional and means, for example, x ± 10%.
[0170] The term "substantially" does not exclude that a composition "substantially free of" Y may not be completely free of Y; i.e., it may contain some Y. If necessary, the term "substantially" may be omitted from the definitions of the present invention.
[0171] The term "between" with respect to two values includes those two values. For example, the range "between" 10 mg and 20 mg includes, inter alia, 10, 15, and 20 mg.
[0172] Unless otherwise specified, a method that includes the step of mixing two or more components does not require any particular mixing order. Thus, the components can be mixed in any order. If there are three components, two of the components can be combined with each other and then that combination can be combined with the third component, etc.
[0173] The steps of various methods can be performed by the same or different people or entities at the same or different times at the same or different geographical locations, e.g., countries.
Examples
[0174] The human genome sequence was analyzed for the presence of various MSRE and MDRE recognition sequences. The percentage of all CpG sites (about 28,000,000) in the genome accessible to various different MSRE and MDRE is shown in Table 1 and is as follows.
[0175]
Table 1
[0176] CpG coverage increases when combinations of enzymes are used, and these combinations can provide recognition sites in more than 99% of the CpG islands in the human genome, as shown in Table 2.
[0177]
Table 2
[0178] * ( *) The combination of enzymes shown in represents a combination of enzymes for which the CpG coverage was calculated based on the total number of recognition sequences present in the human genome (hg18 genome build) for each enzyme. When the recognition sequence of one enzyme overlaps with the recognition sequence of another enzyme and the same CpG site is encompassed by each recognition sequence, an overestimation of the CpG coverage is calculated. The CpG coverage calculated for such a combination of enzymes represents the maximum CpG coverage (i.e., when there is no overlap between recognition sequences across the entire human genome). The maximum CpG coverage is indicated by "≦" in Table 2 above and is, for example, "≦13%". * ) For combinations of enzymes not shown in , overlap in the recognition sequences is taken into account when calculating the CpG coverage / CpG island coverage, and thus the relevant values in Table 2 above represent the absolute CpG coverage / CpG island coverage values.
[0179] The methylation status of multiple sites within a single CpG island tends to be the same (referred to as co-methylation). Thus, a single target site within a CpG island may be sufficient to obtain an image of the methylation status of the entire island. The pair of just two enzymes, HinP1I + AciI, provides coverage of over 99% of CpG islands with minimal complexity and rapid digestion in the reaction mixture. When using other MSREs with the same recognition sequence (e.g., HhaI, AspLEI, or CfoI instead of HinP1I, and / or SsiI instead of AciI), the same high level of coverage is provided, but the preferred combination of AciI and HinP1I offers the advantage that both enzymes can be completely inactivated by heating to 65°C, while HhaI and AspLEI are insensitive to heat inactivation. In addition, both AciI and HinP1I use the same optimal buffer for 100% activity, but the optimal reaction buffers recommended by the suppliers for each of SsiI, AspLEI, and CfoI are different, making it more difficult to adjust the digestion conditions for these alternative enzyme combinations without compromising enzyme activity.
[0180] HinP1I and AciI can be completely inactivated by heating at 65 °C for 20 minutes, and the improved coverage resulting from adding HpaII or HpaII + HpyCH4IV or HhaI + BstUI + HpaII (as in certain known methods) does not justify the negative aspect resulting from the requirement of 80 °C to inactivate HpaII. The same applies to the addition of BstUI, which cannot even be heat-inactivated.
[0181] HinP1I and AciI were used individually for human cfDNA digestion, followed by qPCR for various loci. To obtain equivalent ΔCq values for each enzyme, it was necessary to use more units of HinP1I. This may be due to the fact that AciI cleaves the human genome more frequently than HinP1I and is sufficient to prevent PCR amplification in one cleavage. In the mixture of HinP1I + AciI, better results were found when using an excess amount (enzyme units) of HinP1I, a 4 - to 6-fold excess was the most useful, and a 4.5-fold excess provided the best results (regarding ΔCq for various different loci versus control loci).
[0182] The HinP1I + AciI pair with an excess of HinP1I has been used to digest cfDNA derived from pooled human plasma prepared from approximately 60 subjects. The purified cfDNA was mixed with the enzyme, incubated at 37 °C for 2 hours, and then inactivated at 65 °C for 20 minutes. Two hours was long enough to achieve complete digestion (except when blood was collected in BCT from Streck, in which case typically a longer period was required to achieve complete digestion).
[0183] A useful digestion mix is prepared by mixing 11 μL of rCutSmart™ buffer (10x strength), 4.5 μL of HinP1I (10,000 units / mL), 1 μL of AciI (10,000 units / mL), and 93.5 μL of cfDNA solution (containing total cfDNA extracted from a single blood collection tube). Good results are provided when digested at 37 °C for 2 hours followed by heating at 65 °C for 20 minutes.
[0184] Using the digested cfDNA, sequencing libraries were prepared using the NEBNext Ultra DNA Library Prep Kit. The sequencing libraries were prepared by adding Illumina platform sequencing adapters using enzymatic ligation while preserving the information at the ends of the DNA molecules. The libraries were subjected to whole-genome NGS using the Illumina NovaSeq 6000 sequencing platform with an S4 flow cell. Sequence reads from each sample were mapped against the complete human genome (hg18 genome build). From more than 18×10 9 sequence reads, 98.4% were mapped to the reference genome.
[0185] Any enzyme having a recognition sequence longer than but encompassing the recognition sequence of any of the enzymes listed in Table 1 or Table 2 will inherently have a lower genome-wide CpG coverage / CpG island coverage than the relevant enzyme. Recognition sequences longer than 4 nucleobases are statistically less likely to occur in the genome than recognition sequences that are 4 nucleobases in length, so the longer the recognition sequence, the lower the likelihood of occurrence in the genome. Thus, combinations of enzymes having recognition sequences longer than but encompassing the recognition sequence of AciI or HinP1I can have high CpG coverage / CpG island coverage, but this will not be as high as that achieved using the preferred combination of AciI and HinP1I. Lists of MSRE / MDREs having recognition sequences encompassing the recognition sequence of AciI (CCGC) or HinP1I (GCGC) are provided in Table 3A and Table 3B below, respectively.
[0186]
Table 3
[0187]
Table 4
[0188] The CpG coverage and CpG island coverage of each of the enzymes listed in Table 3A and Table 3B are lower than those of AciI and HinP1I, respectively.
[0189] It is understood that the research of the present inventors is described above only by way of example and that modifications may be made while remaining within the scope and spirit of the present invention.
[0190] References WO 2005 / 090607 WO 2011 / 109529 WO 2014 / 078913 WO 2015 / 169947 WO 2020 / 188561 WO 2022 / 073012 WO 2022 / 107145 U.S. Patent No. 10,801,060 U.S. Patent Application Publication No. 2020 / 0283840 Aleman et al. (2008) Br J Cancer; 98(2), 466 - 473 Alhonen - Hongisto et al. (1987) Biochem J; 242, 205 - 210 Bait et al. (2020) Genome; 64(5): 533 - 546 Bartlett et al. (1991) Somatic Cell & Molecular Genetics; 17(1): 35 - 47 Beikircher et al. (2018) Chapter 21 in DNA Methylation Protocols (ed. Jorg Tost), Methods in Molecular Biology, vol. 1708: 407 - 424 Ellinger et al. (2009) J Urol 182: 324 - 29 Ghosh & Sen - Mandi (2018) Tropical Plant Research; 5(1): 1 - 7 Gong et al. (2002) Cell; 11(6); 803 - 814 Khulan et al. (2006) Genome Res 16: 1046 - 55 Larsson et al. (2012) Tree Genetics & Genomes; 9: 601 - 612 List et al. (1994) J. Biological Chemistry; 269(16): 11902 - 11911 Nalabothula et al. (2015) PLoS ONE 10(8): e0135410 Schmidt et al. (2017) Clinica Chimica Acta 469: 94 - 8 Van Roon et al. (2013) Clinical Epigenetics; 5(2): 1 - 10 van Zogchel et al. (2021) JCO Precision Oncology 5: 1738 - 1748 Wielscher et al. (2015) EbioMedicine 2: 929 - 36 Yan et al. (2009) Chapter 8 in DNA Methylation: Methods and Protocols (ed. Jorg Tost), Methods Molecular Biology; 507: 89 - 106 Zhao et al. (2010) Prenat Diagn 30: 778 - 82
Claims
1. A method for digesting cfDNA using a combination of restriction enzymes comprising HinP1I and AciI, wherein the method comprises (i) digesting the cfDNA with the restriction enzymes and (ii) inactivating the restriction enzymes by heating them for a longer period than 15 minutes, wherein the restriction enzymes are completely inactivated by the heating.
2. A method for digesting cfDNA using a combination of restriction enzymes comprising HinP1I and AciI, the method comprising: (i) providing a blood sample contained in a collection tube comprising an anticoagulant and an agent that inhibits the release of genomic DNA from leukocytes in the sample into the plasma component of the blood sample; (ii) preparing plasma from the blood sample; and (iii) digesting the plasma cfDNA with the restriction enzymes for at least two hours.
3. A method for analyzing cfDNA, the method comprising: (i) digesting the cfDNA using a combination of restriction enzymes consisting of HinP1I and AciI to provide digested cfDNA; and (ii) sequencing the digested cfDNA.
4. A method for digesting cfDNA using a combination of restriction enzymes comprising HinP1I and AciI, wherein the method comprises (i) digesting the cfDNA with the restriction enzymes for 11 hours or less, and (ii) inactivating the restriction enzymes by heating, wherein the restriction enzymes are completely inactivated by the heating.
5. A method for analyzing cfDNA, the method comprising: (i) digesting the cfDNA for 11 hours or less using a combination of restriction enzymes including HinP1I and AciI to provide digested cfDNA; (ii) completely inactivating the restriction enzymes; and (iii) performing real-time PCR on the digested cfDNA.
6. The method according to claim 4 or 5, wherein the combination of restriction enzymes consists of HinP1I and AciI.
7. A composition comprising HinP1I and AciI as only two restriction enzymes.
8. A composition comprising HinP1I and AciI, wherein the ratio of the activity of HinP1I to AciI is at least 1.2:
1.
9. A composition comprising a plurality of restriction enzymes, wherein the plurality of restriction enzymes consist of MSRE and / or MDRE, (i) at least two different restriction enzymes in the plurality of restriction enzymes have different recognition sequences, and (ii) the restriction enzymes can be completely inactivated by heating to 65°C.
10. The composition according to claim 9, comprising a plurality of MSREs, wherein (i) at least two different MSREs in the plurality of MSREs have different recognition sequences, and (ii) the plurality of MSREs can be completely inactivated by heating to 65°C.
11. (i) at least one salt and / or at least one buffer; (ii) 50 mM potassium acetate, 20 mM tris-acetic acid, 10 mM magnesium acetate, 100 μg / mL recombinant albumin, pH 7.9; or (iii) 50 mM Tris-HCl, 10 mM MgCl 2 100 mM NaCl, 100 μg / mL recombinant albumin, pH 7.9; A composition according to any one of claims 7 to 9, comprising:
12. (i) comprising cfDNA; (ii) comprising PCR reagents and / or sequencing reagents; and / or (iii) (a) the ratio of HinP1I to AciI is at least 2:1, (b) the composition contains Mg++ ions, and / or (c) the pH of the composition is greater than 7; The composition according to any one of claims 7 to 9.
13. A method for analyzing cfDNA, comprising digesting the cfDNA by the method of any one of claims 1, 2, or 4, and subsequently (a) performing real-time PCR on the digested cfDNA, or (b) sequencing the digested cfDNA.
14. A method for evaluating the methylation status of one or more CpG sites in cfDNA, comprising digesting the cfDNA by the method of any one of claims 1, 2, or 4, and subsequently quantifying the degree of digestion of one or more of the one or more CpG sites.
15. The method according to any one of claims 1 to 5, wherein the restriction enzyme is inactivated by heating the composition at 65°C for at least 20 minutes after digestion of cfDNA, and the restriction enzyme is completely inactivated by the heating.
16. (a) The ratio of HinP1I to AciI activity is at least 2:1, and (b) the restriction enzyme contains Mg during digestion. ++ The method according to any one of claims 1 to 5, wherein an ion source is provided, (c) digestion is carried out at a pH greater than 7, (d) the cfDNA is human cfDNA, and / or (e) the amount of cfDNA to be digested is 10 to 400 ng.
17. (i) cfDNA is human plasma cfDNA; (ii) HinP1I is present in an excess amount of 2:1 to 5:1 relative to AciI (measured with respect to enzyme units); and / or (iii) HinP1I is replaced by an MSRE that recognizes the same recognition sequence as HinP1I, and / or AciI is replaced by an MSRE that recognizes the same recognition sequence as AciI; The composition according to any one of claims 7 to 9, or the method according to any one of claims 1 to 5.
18. A composition comprising, as two restriction enzymes, (i) HinP1I or an MSRE that recognizes the same recognition sequence as HinP1I, and (ii) AciI or an MSRE that recognizes the same recognition sequence as AciI.
19. The composition according to claim 18, wherein (i) the MSRE that recognizes the same recognition sequence as HinP1I is HhaI, AspLEI, or CfoI, and / or (ii) the MSRE that recognizes the same recognition sequence as AciI is SsiI.
20. A composition comprising (i) one of HhaI, AspLEI, and CfoI as two restriction enzymes, and (ii) SsiI.