A method for library preparation to enrich informative DNA fragments using enzymatic digestion.
The method of enzymatic digestion with restriction enzymes and adapters enriches informative DNA fragments, particularly those with CpG-rich regions, addressing the low enrichment challenge in cfDNA, enhancing methylation profiling and clinical applications.
Patent Information
- Application Number
- JP2022509564
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-04-28
- Filing Date
- 2020-04-28
- Publication Date
- 2025-12-18
- Estimated Expiration
- 2040-04-28
AI Technical Summary
Existing methods for enriching informative regions of cell-free DNA (cfDNA) face challenges due to fragmented species exhibiting a characteristic peak around 166 base pairs, leading to low levels of enrichment when using size-selection processes, and typical enrichment approaches select all or nearly all of the cfDNA population.
A method involving enzymatic digestion with restriction enzymes like MspI and adapters, followed by adapter ligation, amplification, and reduction of adapter dimers, to enrich informative DNA fragments, particularly those with CpG-rich regions, using enzymes such as BspDI, ClaI, NarI, Xhol, SmlI, HpyF30I, PaeR7I, and Sfr274I to digest adapter dimers during and/or after adapter ligation.
Enhances the enrichment of informative DNA fragments, facilitating efficient methylation profiling and clinical applications like cancer diagnosis and screening by reducing adapter dimers and improving the specificity of sequencing library preparation.
Smart Images

Figure 0007788372000001 
Figure 0007788372000002 
Figure 0007788372000003
Abstract
Description
[Technical Field]
[0001] cross reference This application claims the benefit of U.S. Provisional Patent Application No. 62 / 839,719, filed April 28, 2019, which is incorporated herein by reference in its entirety.
[0002] Technical Field Aspects of the present disclosure include at least the fields of nucleic acid preparation and analysis, sequencing, molecular biology, cell biology, and medicine. [Background technology]
[0003] background With the rapid development of next-generation sequencing (NGS) technology, analysis of genomic modifications in deoxyribonucleic acid (DNA) has become routine to provide diagnostic information about disease (e.g., cancer) or other health conditions (e.g., fetal genetic material in maternal blood). A typical sequencing library preparation technique may include one or more steps, such as DNA fragmentation, fragment end repair, dA tailing, adapter ligation, and polymerase chain reaction (PCR) enrichment, as well as one or more purification steps.
[0004] Certain health conditions, such as cancer or infectious diseases, can cause DNA to be released into the bloodstream or lymphatic system, where tumor or microbiome DNA can become part of cell-free DNA (cfDNA) circulating in bodily fluids such as plasma or urine. Such cfDNA can be subjected to genomic or epigenomic profiling for clinical applications such as cancer screening, microbial detection, or prenatal testing. For example, whole-genome bisulfite sequencing (WGBS) can provide a comprehensive view of the DNA methylome, but deep sequencing the entire genome can be expensive. Methods for enriching informative regions of cell-free DNA could advantageously enable genomic or epigenomic profiling for clinical diagnostic applications. Using restriction enzymes, clustered regularly interspaced short palindromic repeats (CRISPR), transposases, or other techniques, intact DNA can be fragmented so that informative fragments can be enriched by size-based selection. For example, MspI enzyme digestion can enrich for CpG-rich regions by generating smaller fragments that can be used for methylation profiling.
[0005] Fragmented species of cell-free DNA, which can exhibit a characteristic peak around 166 base pairs (bp), pose a challenge to typical enrichment approaches based on restriction enzyme digestion. Any size-selection process to select informative DNA may select all or nearly all of the cfDNA population, thus resulting in low levels of enrichment.
[0006] The present disclosure provides improvements to methods and compositions for nucleic acid library preparation. Summary of the Invention
[0007] overview The present disclosure provides methods for preparing nucleic acid libraries using restriction enzymes and adapters, which represent an improvement in the art. In some embodiments, these methods include library preparation with reduced amounts of adapter dimers. Once prepared, the nucleic acid library can be used for any purpose, including, for example, next-generation sequencing. In some embodiments, the present disclosure relates to methods for preparing libraries from informative deoxyribonucleic acid (DNA) fragments, whose sequences, modification states, and / or levels indicate a medical condition or the risk or susceptibility thereof. As used herein, "informative fragments" refers to fragments generated by cleavage with a restriction enzyme (e.g., multiple CpG sites after MspI restriction enzyme digestion). While the nucleic acid can be of any type, in some embodiments, the nucleic acid comprises DNA, including cell-free DNA (cfDNA). In some embodiments, the library is used for methylation profiling of cfDNA.
[0008] In one aspect, the present disclosure provides a method for preparing a library of nucleic acids (e.g., derived from a plurality of deoxyribonucleic acid (DNA) molecules of a subject) (e.g., for analysis, including by sequencing), the method comprising: subjecting the plurality of DNA molecules to enzymatic digestion, which fragments at least a subset of the DNA molecules, to generate DNA fragments having overhangs on one or both ends; ligating adapters having overhangs that complement the overhangs of the DNA fragments to generate a plurality of tagged DNA molecules; concentrating the plurality of tagged DNA molecules (which may be referred to herein as adapter-ligated DNA molecules or fragments) before or after reducing the number of adapter dimers; and optionally, subjecting the plurality of tagged DNA molecules, or derivatives thereof, to nucleic acid sequencing to obtain a plurality of sequence reads.
[0009] In some embodiments, restriction enzymes such as BspDI, ClaI, AclI, NarI, Xhol, SmlI, HpyF30I, PaeR71, Sfr274I, or combinations thereof are used to digest the adapter dimer during and / or after adapter ligation. In some embodiments, the subjecting and ligating steps are performed in the same reaction using (1) one or more restriction enzymes, e.g., MspI and / or HpaII and / or Taqα1, and (2) a ligase, e.g., T7 ligase and / or T4 ligase. In some embodiments, the subjecting, ligating, and reducing steps are performed in the same reaction using (1) one or more restriction enzymes, such as MspI, and / or HpaII, and / or TaqαI, and / or BspD1, and / or ClaI, and / or AclI, and / or NarI, and / or XhoI, and / or SmlI, and / or HpyF30I, and / or PaeR7I, and / or Sfr274I, and (2) a ligase, such as T7 ligase and / or T4 ligase. In some embodiments, the enriching step for the plurality of tagged DNAs includes an amplification step using, for example, polymerase chain reaction (PCR). In some embodiments, the enriching step for the plurality of tagged DNAs includes targeted capture. In some embodiments, the plurality of tagged DNAs undergo bisulfite conversion. In some embodiments, primers for PCR are designed to recognize the junction between the adapter and the target DNA, but not the junction between the adapters.
[0010] In one aspect, the present disclosure provides a method for enriching a plurality of DNA fragments from a plurality of cfDNA molecules of a subject, the method comprising: subjecting a plurality of cfDNA molecules to enzymatic digestion, wherein the plurality of cfDNA molecules is fragmented to generate fragments comprising one or more regions of interest; ligating adapters having overhangs complementary to overhangs of the plurality of fragmented cfDNA molecules to provide a plurality of tagged DNA molecules; reducing the number of adapter dimers; optionally, subjecting the plurality of tagged DNA molecules or derivatives thereof to nucleic acid sequencing to obtain a plurality of sequence reads; and processing the plurality of sequence reads to achieve one or more clinical applications.
[0011] In another aspect, the disclosure provides a method for preparing a library of nucleic acids, the method comprising: (a) providing a plurality of DNA molecules; (b) subjecting the molecules to digestion with a first one or more restriction enzymes, which generates DNA fragments; (c) subjecting the DNA fragments to (e.g., ligating) adapters, which generates a mixture of adapter-ligated DNA fragments and adapter dimers; and (d) amplifying the adapter-ligated DNA fragments to generate amplified adapter-ligated DNA fragments, optionally, the method further comprises, in some embodiments, performing (b) and (c) in the same operation, the method further comprising reducing the generated adapter dimers, the reducing step being performed during or after (c) and / or after (d), the reducing step comprising distinguishing between a junction between an adapter and a DNA fragment and a junction between an adapter and another adapter.
[0012] In another aspect, the disclosure provides a method for preparing a library of nucleic acids, the method comprising: (a) providing a plurality of DNA molecules; (b) subjecting the molecules to digestion with a first one or more restriction enzymes, which produces DNA fragments; (c) subjecting the DNA fragments to (e.g., ligating) adapters, which produces a mixture of adapter-ligated DNA fragments and adapter-dimers; and (d) amplifying the adapter-ligated DNA fragments to produce amplified adapter-ligated DNA fragments, with one or more of the following proviso: (1) performing (d) with primers that bind to the junction between the ends of the DNA fragments and the adapters, but that do not bind to the junction between the end of one adapter and the end of another adapter; (2) producing the mixture of adapter-ligated DNA fragments and adapter-dimers. (3) performing (b) in the same reaction as (c) in the presence of a second restriction enzyme or enzymes that digest the junction between the end of one adapter and the end of another adapter but not the junction between the end of the DNA fragment and the adapter; (4) the adapters are by design adapter dimers, and a third restriction enzyme or enzymes digest the junction between the end of one adapter and the end of another adapter but not the junction between the end of the DNA fragment and the adapter; and / or (5) the amplifying step also produces amplified adapter dimers that are digested by a fourth restriction enzyme or enzymes that digest the junction between the end of one adapter and the end of another adapter.
[0013] In some embodiments, subjecting the plurality of cfDNA molecules to enzymatic digestion comprises performing digestion with one or more restriction enzymes on the plurality of cell-free DNA molecules. In some embodiments, the method comprises: digesting the plurality of cfDNA molecules with one or more restriction enzymes, such as AcII, HindIII, MluCI, PciI, AgeI, BspMI, BfuAI, SexAI, MluI, BceAI, HpyCH4IV, HpyCH4III, BaeI, BsaXI, AflIII, SpeI, BsrI, BmrI, BglII, BspDI, PI-SceI, NsiI, AseI, CspCI, MfeI, BssS α I, DraIII, EcoP15I, AlwNI, BtsIMutI, NdeI, CviAII, FatI, NlaIII, FspEI, XcmI, BstXI, PflMI, BccI, NcoI, BseYI, FauI, TspMI, XmaI, LpnPI, AclI, ClaI , SacII, HpaII, MspI, ScrFI, StyD4I, BsaJI, BslI, BtgI, NciI, AvrII, MnlI, BbvCI, SbfI, Bpu10I, Bsu36I, EcoNI, HpyAV, BstNI, PspGI, StyI, BcgI, PvuI , EagI, RsrII, BsiEI, BsiWI, BsmBI, Hpy99I, AbaSI, MspJI, SgrAI, BfaI, BspCNI, XhoI, PaeR7I, EarI, AcuI, PstI, BpmI, DdeI, SfcI, AflII, BpuEI, SmlI, AvaI, BsoBI, MboII, BbsI, BsmI, EcoRI, HgaI, AatII, PflFI, Tth111I, AhdI, DrdI, SacI, BseRI, PleI, HinfI, Sau3AI, MboI, DpnII, TfiI, BsrDI, BbvI, Bts αI, BstAPI, SfaNI, SphI, NmeAIII, NgoMIV, BglI, AsiSI, BtgZI, HhaI, HinP1I, BssHII, NotI, Fnu4HI, MwoI , BmtI, NheI, BspQI, BlpI, TseI, ApeKI, Bsp1286I, AlwI, BamHI, BtsCI, FokI, FseI, SfiI, NarI, PluTI, Kas I, AscI, EciI, BsmFI, ApaI, PspOMI, Sau96I, KpnI, Acc65I, BsaI, HphI, BstEII, AvaII, BanI, BaeGI, BsaH I, BanII, CviQI, BciVI, SalI, BcoDI, BsmAI, ApaLI, BsgI, AccI, Tsp45I, BsiHKAI, TspRI, ApoI, NspI, BsrF α I, BstYI, HaeII, EcoO109I, PpuMI, I-CeuI, I-SceI, BspHI, BspEI, MmeI, Taq α 1, Hpy188I, Hpy188III, XbaI, BclI, PI-PspI, BsrGI, MseI, PacI, BstBI, PspXI, BsaWI, EaeI, HpyF30I, Sfr274I, and combinations thereof are utilized. In certain embodiments, it is contemplated that 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more of these, or any range derivable therein, may be excluded.
[0014] In some embodiments, subjecting a plurality of cfDNA molecules to enzymatic digestion comprises using a CRISPR (clustered regularly interspaced short palindromic repeats)-Cas9 system or its functional derivative to cleave the cell-free DNA molecules. In some embodiments, subjecting a plurality of cfDNA molecules to enzymatic digestion comprises using one or more transposases or their functional derivatives to cleave the cfDNA molecules.
[0015] In some embodiments, the method further comprises subjecting a plurality of tagged DNA fragments or derivatives thereof to conditions sufficient to allow differentiation between methylated and unmethylated nucleobases in the tagged DNA fragments.In some embodiments, subjecting a plurality of tagged DNA fragments or derivatives thereof to conditions for distinguishing between methylated and unmethylated bases comprises performing bisulfite conversion on the plurality of tagged DNA fragments.In some embodiments, subjecting a plurality of tagged DNA fragments or derivatives thereof to conditions for distinguishing between methylated and unmethylated bases comprises an enzymatic and / or chemical reaction for oxidizing methylated cytosine nucleobases and / or hydroxymethylated cytosine nucleobases, followed by reduction and / or deamination of the oxidation reaction product.
[0016] In some embodiments, restriction enzymes such as BspDI, ClaI, AclI, NarI, Xhol, SmlI, HpyF30I, PaeR7I, and / or Sfr274I are utilized to digest adapter dimers during and / or after adapter ligation, and / or after the PCR amplification step, and / or after both bisulfite conversion and PCR amplification.
[0017] In some embodiments, the enzyme digestion of cfDNA and the ligation of adapters are carried out in the same reaction.Furthermore, in some embodiments, the enzyme used is MspI and / or BspDI, and the ligase can be any ligase, including T7 DNA ligase and / or T4 DNA ligase.
[0018] In some embodiments, the step of enriching a plurality of tagged DNAs includes amplification, such as PCR. In some embodiments, primers for PCR are designed to recognize (e.g., bind to) the junction between the adapter and the target DNA, but not the junction between two adapters bound to each other. In some embodiments, primers for PCR are designed to recognize the junction between the adapter and the target DNA after bisulfite conversion, but not the junction between the adapter and the adapter after bisulfite conversion. In some embodiments, primers for PCR are designed to recognize the junction between the adapter and the target DNA after an enzymatic and / or chemical reaction to oxidize methylated cytosine nucleobases and / or hydroxymethylated cytosine nucleobases, followed by reduction and / or deamination of the oxidation reaction product, but not the junction between the adapter and the adapter after the enzymatic and / or chemical reaction.
[0019] The libraries generated may contain one or more regions of interest, which may be of any type, and in some embodiments, the libraries also contain one or more CpG sites.
[0020] The adaptors for ligating to the DNA fragments may themselves be designed as adaptor-adaptor dimers, for example to obtain long-term stability.
[0021] In some embodiments, the present disclosure provides a sequencing library preparation method, which facilitates DNA fragmentation and adapter ligation, and reduces adapter dimer.The example of the application of the method of the present disclosure is that after the library preparation method of the present disclosure, cfDNA methylome profiling is carried out for cancer diagnosis and screening.
[0022] In another aspect, the disclosure provides a method for preparing a library of nucleic acids, the method comprising: (a) providing a plurality of DNA molecules; (b) subjecting the molecules to digestion with a first one or more restriction enzymes, which produces DNA fragments; (c) subjecting the DNA fragments to (e.g., ligating) adapters and a ligase, which produces a mixture of adapter-ligated DNA fragments and adapter-dimers; and (d) amplifying the adapter-ligated DNA fragments to produce amplified adapter-ligated DNA fragments, the method further comprising reducing the amount of adapter-dimers produced, the method further comprising performing the reducing step during or after (c) and / or (d), the reducing step comprising distinguishing between a junction between an adapter and a DNA fragment and a junction between an adapter and another adapter.
[0023] The first one or more restriction enzymes are AcII, HindIII, MluCI, PciI, AgeI, BspMI, BfuAI, SexAI, MluI, BceAI, HpyCH4IV, HpyCH4III, BaeI, BsaXI, AflIII, SpeI, BsrI, BmrI, BglII, BspDI, PI-SceI, NsiI, AseI, CspCI, MfeI, BssS αI、DraIII、EcoP15I、AlwNI、BtsIMutI、NdeI、CviAII、FatI、NlaIII、FspEI、X cmI, BstXI, PflMI, BccI, NcoI, BseYI, FauI, TspMI, XmaI, LpnPI, AclI, ClaI 、SacII、HpaII、MspI、ScrFI、StyD4I、BsaJI、BslI、BtgI、NciI、AvrII、MnlI、 BbvCI, SbfI, Bpu10I, Bsu36I, EcoNI, HpyAV, BstNI, PspGI, StyI, BcgI, PvuI 、EagI、RsrII、BsiEI、BsiWI、BsmBI、Hpy99I、AbaSI、MspJI、SgrAI、BfaI、Bsp CNI、XhoI、PaeR7I、EarI、AcuI、PstI、BpmI、DdeI、SfcI、AflII、BpuEI、SmlI、 AvaI, BsoBI, MboII, BbsI, BsmI, EcoRI, HgaI, AatII, PflFI, Tth111I, AhdI DrdI, SacI, BseRI, PleI, HinfI, Sau3AI, MboI, DpnII, TfiI, BsrDI, BbvI, Bts α I、BstAPI、SfaNI、SphI、NmeAIII、NgoMIV、BglI、AsiSI、BtgZI、HhaI、HinP1I、BssHII、NotI、Fnu4HI、MwoI BmtI, NheI, BspQI, BlpI, TseI, ApeKI, Bsp1286I, AlwI, BamHI, BtsCI, FokI, FseI, SfiI, NarI, PluTI I、AscI、EciI、BsmFI、ApaI、PspOMI、Sau96I、KpnI、Acc65I、BsaI、HphI、BstEII、AvaII、BanI、BaeHGI and BsaHGI I、BanII、CviQI、BciVI、SalI、BcoDI、BsmAI、ApaLI、BsgI、AccI、Tsp45I、BsiHKAI、TspRI、ApoI、NspIF、BsrIF α I、BstYI、HaeII、EcoO109I、PpuMI、I-CeuI、I-SceI、BspHI、BspEI、MmeI、Taq αThe amino acids may include I, Hpy188I, Hpy188III, XbaI, BclI, PI-PspI, BsrGI, MseI, PacI, BstBI, PspXI, BsaWI, EaeI, HpyF30I, Sfr274I, or a combination thereof. It is contemplated that in certain embodiments, one, two, three, four, five, six, seven, eight, nine, ten, or more of these may be excluded. In some embodiments, (b) and (c) are carried out in the same reaction mixture. In some embodiments, (b) is carried out at a different temperature than (c) or the same temperature as (c).
[0024] In some embodiments, distinguishing between a junction between an adaptor and a DNA fragment and a junction between an adaptor and another adaptor further comprises using an adaptor that is designed to be digested by a second one or more restriction enzymes when in a dimerized configuration but is incapable of being digested by the second one or more restriction enzymes when the adaptor is ligated to the end of the DNA fragment. Distinguishing between a junction between an adaptor and a DNA fragment and a junction between an adaptor and another adaptor may further comprise using an adaptor that is designed such that a primer for amplification can initiate polymerization at the junction between the adaptor and the DNA fragment but cannot initiate polymerization at the junction between the adaptor and another adaptor.
[0025] In another aspect, the disclosure provides a method for preparing a library of nucleic acids, the method comprising: (a) providing a plurality of DNA molecules; (b) subjecting the molecules to digestion with a first one or more restriction enzymes, which generates DNA fragments; generating the DNA fragments; (c) subjecting the DNA fragments to (e.g., ligating) adapters (e.g., adapters can comprise known sequences, unique sequences, or random sequences), which generates a mixture of adapter-ligated DNA fragments and adapter dimers; and (d) amplifying the adapter-ligated DNA fragments to generate amplified adapter-ligated DNA fragments, with one or more of the following proviso: (1) performing (d) with primers that bind to the junction between the ends of the DNA fragments and the adapters but do not bind to the junction between the end of one adapter and the end of another adapter; (2) amplifying the adapter-ligated DNA fragments to generate amplified adapter-ligated DNA fragments. (3) subjecting the mixture of ligated DNA fragments and adapter dimers to the influence of a second restriction enzyme or enzymes that digest the junction between the end of one adapter and the end of another adapter but not the junction between the end of the DNA fragment and the adapter; (4) performing (b) and (c) in the same reaction mixture, and the second restriction enzyme or enzymes digest the junction between the end of one adapter and the end of another adapter but not the junction between the end of the DNA fragment and the adapter; (5) the adapters are designed adapter dimers, and the second restriction enzyme or enzymes digest the junction between the end of one adapter and the end of another adapter but not the junction between the end of the DNA fragment and the adapter; and / or (6) amplifying, which also produces amplified adapter dimers that are digested by a third restriction enzyme or enzymes that digest the junction between the end of one adapter and the end of another adapter.
[0026] In some embodiments, the method further comprises subjecting the adaptor-ligated fragments to conditions sufficient to distinguish methylated nucleobases from unmethylated nucleobases. The conditions sufficient to distinguish methylated nucleobases from unmethylated nucleobases may comprise subjecting the adaptor-ligated fragments to bisulfite conversion. The conditions sufficient to distinguish methylated nucleobases from unmethylated nucleobases may comprise subjecting the adaptor-ligated fragments to, for example, one or more enzymatic and / or chemical reactions to oxidize methylated cytosine nucleobases and / or hydroxymethylated cytosine nucleobases, followed by reduction and / or deamination of the oxidation reaction products. Deamination of the oxidation reaction products may be performed using an apolipoprotein B mRNA editing enzyme catalytic polypeptide-like (APOBEC) that deaminates cytosine nucleobases. Reduction and / or deamination of the oxidation reaction products may be performed using pyridine borane. In some embodiments, the method further comprises performing β-glucosyltransferase treatment prior to the one or more enzymatic and / or chemical reactions.
[0027] In some embodiments, some or all of the amplified, adapter-ligated DNA fragments are further analyzed, modified, or both. The analysis may include sequencing, such as next-generation sequencing. In some embodiments, targeted capture is performed before next-generation sequencing to further enrich the adapter-ligated fragments. In some embodiments, size selection is performed before next-generation sequencing to further enrich the adapter-ligated fragments. The analysis may include analyzing the methylation patterns of the amplified, adapter-ligated DNA fragments. The adapters may include GC (3' to 5') overhangs. The first one or more restriction enzymes may include MspI, HpaII, TaqαI, or functional analogs thereof, or mixtures thereof. The second one or more restriction enzymes may include one or more of BspD1, ClaI, AclI, NarI, XhoI, SmlI, HpyF30I, PaeR7I, Sfr274I, or functional analogs thereof, or mixtures thereof. In some embodiments, the ligase is T7 DNA ligase, T4 DNA ligase, T3 DNA ligase, Taq DNA ligase, or a functional analog thereof, or a mixture thereof.
[0028] In some embodiments, the plurality of DNA molecules comprises cell-free DNA.In some embodiments, the method further comprises obtaining cfDNA from the sample obtained or derived from the sample of subject or individual.The cfDNA can be enriched for molecules with one or more CpG sites.The sample can be any type, for example, can be derived from plasma, serum, bone marrow, cerebrospinal fluid, pleural fluid, saliva, stool or urine.In some embodiments, the method further comprises obtaining sample from subject or individual.
[0029] In another aspect, the present disclosure provides a method for preparing a library of nucleic acids, comprising the steps of: (a) providing a plurality of DNA molecules; (b) subjecting the molecules to digestion with a first one or more restriction enzymes, which produces DNA fragments; (c) subjecting the DNA fragments to (e.g., ligating) adapters, which produces a mixture of adapter-ligated DNA fragments and adapter dimers; and (d) amplifying the adapter-ligated DNA fragments to produce amplified adapter-ligated DNA fragments, which uses a set of primers that bind to the junction between the ends of the DNA fragments and the adapters but do not bind to the junction between the end of one adapter and the end of another adapter. In some embodiments, the first one or more restriction enzymes comprise one or more of MspI, HpaII, TaqαI, or functional analogs thereof, or a mixture thereof.
[0030] In some embodiments, the method further comprises carrying out (b) and (c) in the same reaction mixture. In some embodiments, the method further comprises subjecting the adaptor-ligated fragments to conditions sufficient to allow methylated nucleobases to be distinguished from unmethylated nucleobases. Conditions sufficient to allow methylated nucleobases to be distinguished from unmethylated nucleobases may comprise subjecting the adaptor-ligated fragments to bisulfite conversion, or subjecting the adaptor-ligated fragments to one or more enzymatic and / or chemical reactions (e.g., those to oxidize methylated and / or hydroxymethylated cytosine nucleobases, followed by reduction and / or deamination of the oxidation reaction product).
[0031] In some embodiments, the oxidation is carried out using a ten-eleven translocation (TET) enzyme. In some embodiments, the oxidation is carried out using potassium perruthenate. In some embodiments, deamination of the oxidation reaction product is carried out using an APOBEC that deaminates cytosine nucleobases, or deamination of the oxidation reaction product may be carried out using pyridine borane. In some embodiments, the method further comprises performing β-glucosyltransferase treatment prior to one or more enzymatic or chemical reactions. In some embodiments, the adapter comprises a GC overhang.
[0032] In another aspect, the present disclosure provides a method for preparing a library of nucleic acids, comprising the steps of: (a) providing a plurality of DNA molecules; (b) subjecting the molecules to digestion with a first one or more restriction enzymes, which produces DNA fragments; (c) subjecting (e.g., ligating) the DNA fragments to adapters (which may, for example, contain GC overhangs), which produces a mixture of adapter-ligated DNA fragments and adapter dimers; (d) subjecting the mixture of adapter-ligated DNA fragments and adapter dimers to a second one or more restriction enzymes that digest the junction between the end of one adapter and the end of another adapter but do not digest the junction between the end of the DNA fragment and the adapter; and (e) amplifying the adapter-ligated DNA fragments to produce amplified adapter-ligated DNA fragments. In some embodiments, the method further comprises performing (b), (c), and (d) in the same reaction mixture.
[0033] In some embodiments, the first restriction enzyme(s) comprise one or more of MspI, HpaII, TaqαI, or functional analogs thereof, or mixtures thereof. The second restriction enzyme(s) may comprise one or more of BspDI, ClaI, AclI, NarI, XhoI, SmlI, HpyF30I, PaeR7I, Sfr274I, or functional analogs thereof, or mixtures thereof.
[0034] In some embodiments, the method further comprises subjecting the adaptor-ligated fragments to conditions sufficient to distinguish methylated nucleobases from unmethylated nucleobases, e.g., subjecting the adaptor-ligated fragments to bisulfite conversion, or subjecting the adaptor-ligated fragments to one or more enzymatic and / or chemical reactions, e.g., to oxidize methylated and / or hydroxymethylated cytosine nucleobases, followed by reduction and / or deamination of the oxidation reaction products (e.g., performed using APOBEC or pyridine borane). In some embodiments, the oxidation is performed using a ten-eleven translocation (TET) enzyme. In some embodiments, the oxidation is performed using potassium perruthenate. In some embodiments, the method further comprises performing β-glucosyltransferase treatment prior to the one or more enzymatic or chemical reactions. In some embodiments, the adaptor comprises a GC overhang.
[0035] In another aspect, the present disclosure provides a method for preparing a library of nucleic acids, comprising the steps of: (a) providing a plurality of DNA molecules; (b) subjecting the molecules to digestion with a first one or more restriction enzymes, which produces DNA fragments; (c) subjecting the DNA fragments to (e.g., ligating) adapters that are by-design adapter dimers, comprising subjecting the adapter dimers to the influence of a second one or more restriction enzymes to produce adapters, which further produces a mixture of adapter-ligated DNA fragments and adapter dimers, wherein the second one or more restriction enzymes digest the junction between the end of one adapter and the end of another adapter, but do not digest the junction between the end of the DNA fragment and the adapter; and (d) amplifying the adapter-ligated DNA fragments to produce amplified adapter-ligated DNA fragments. In some embodiments, the method further comprises performing (b) and (c) in the same reaction mixture.
[0036] In some embodiments, the first one or more restriction enzymes comprise one or more of MspI, HpaII, TaqαI, or functional analogs thereof, or mixtures thereof. The second one or more restriction enzymes may comprise one or more of BspDI, ClaI, AclI, NarI, XhoI, SmlI, HpyF30I, PaeR7I, Sfr274I, or functional analogs thereof, or mixtures thereof. In some embodiments, the method further comprises performing (b) and (c) in the same reaction mixture.
[0037] In some embodiments, the method further comprises subjecting the adaptor-ligated fragments to conditions sufficient to distinguish methylated nucleobases from unmethylated nucleobases. The conditions may include subjecting the adaptor-ligated fragments to bisulfite conversion, or subjecting the adaptor-ligated fragments to one or more enzymatic and / or chemical reactions, for example, to oxidize methylated and / or hydroxymethylated cytosine nucleobases, followed by reduction and / or deamination of the oxidation reaction products (e.g., using APOBEC or pyridine borane). In some embodiments, the oxidation is carried out using a ten-eleven translocation (TET) enzyme. In some embodiments, the oxidation is carried out using potassium perruthenate. In some embodiments, the method further comprises performing β-glucosyltransferase treatment before the one or more enzymatic or chemical reactions. In some embodiments, GC overhangs are created by digestion of the adaptor dimer of the second adaptor with a second restriction enzyme or enzymes.
[0038] In another aspect, the present disclosure provides a method for preparing a library of nucleic acids, comprising the steps of: (a) providing a plurality of DNA molecules; (b) subjecting the molecules to digestion with a first one or more restriction enzymes, which produces DNA fragments; (c) subjecting the DNA fragments to (e.g., ligating) adapters, which produces a mixture of adapter-ligated DNA fragments and adapter-dimers; and (d) amplifying the adapter-ligated DNA fragments to produce amplified adapter-ligated DNA fragments and amplified adapter-dimers, which are digested with a second one or more restriction enzymes that digest the junction between the end of one adapter and the end of another adapter. In some embodiments, the method further comprises performing (b) and (c) in the same reaction mixture.
[0039] In some embodiments, the first restriction enzyme(s) comprise one or more of MspI, HpaII, TaqαI, or functional analogs thereof, or mixtures thereof. The second restriction enzyme(s) may comprise one or more of BspDI, ClaI, AclI, NarI, XhoI, SmlI, HpyF30I, PaeR7I, Sfr274I, or functional analogs thereof, or mixtures thereof.
[0040] In some embodiments, the method further comprises subjecting the adaptor-ligated fragments to conditions sufficient to distinguish methylated nucleobases from unmethylated nucleobases, for example, by subjecting the adaptor-ligated fragments to bisulfite conversion or subjecting the adaptor-ligated fragments to one or more enzymatic and / or chemical reactions. The one or more enzymatic reactions may be for oxidizing methylated cytosine nucleobases and / or hydroxymethylated cytosine nucleobases, followed by reduction and / or deamination of the oxidation reaction product, for example, using APOBEC or pyridine borane. In some embodiments, the oxidation is carried out using a ten-eleven translocation (TET) enzyme. In some embodiments, the oxidation is carried out using potassium perruthenate. In some embodiments, the method further comprises performing β-glucosyltransferase treatment prior to the one or more enzymatic or chemical reactions. In some embodiments, the adaptor comprises a GC overhang.
[0041] In another aspect, the present disclosure provides a method for preparing a library of nucleic acids, comprising: (a) providing a plurality of DNA molecules; (b) subjecting the molecules to digestion with a first restriction enzyme in the presence of a ligase, wherein the digestion produces DNA fragments; (c) subjecting the DNA fragments to the influence of (e.g., ligating) adapters, which produces a mixture of adaptor-ligated DNA fragments and adaptor dimers, wherein a second restriction enzyme or enzymes digest the junction between the end of one adaptor and the end of another adaptor, but not the junction between the end of the DNA fragment and the adaptor; and (d) amplifying the adaptor-ligated DNA fragments to produce amplified adaptor-ligated DNA fragments.
[0042] In some embodiments, the method further comprises performing (b) and (c) in the same reaction mixture. In some embodiments, the method further comprises subjecting the adapter-ligated fragments to bisulfite conversion. The adapters may comprise GC overhangs.
[0043] In some embodiments of the methods provided herein, the restriction enzyme that digests the junction between the end of one adapter and the end of another adapter is replaced with a CRISPR-associated endonuclease and a specifically designed guide RNA.
[0044] [The present invention 1001] 1. A method for preparing a library of nucleic acids, comprising the steps of: (a) digesting a plurality of DNA molecules with a first one or more restriction enzymes to generate DNA fragments; (b) ligating adapters to the DNA fragments by incubation with a ligase to produce a mixture of adapter-ligated DNA fragments and adapter dimers; (c) amplifying the adaptor-ligated DNA fragments to produce amplified adaptor-ligated DNA fragments; and (d) reducing the amount of adapter dimers either after or during (b) and / or after (c), wherein the reducing step comprises distinguishing between a junction between an adapter and a DNA fragment and a junction between an adapter and another adapter. [The present invention 1002] The first one or more restriction enzymes are AcII, HindIII, MluCI, PciI, AgeI, BspMI, BfuAI, SexAI, MluI, BceAI, HpyCH4IV, HpyCH4III, BaeI, BsaXI, AflIII, SpeI, BsrI, BmrI, BglII, BspDI, PI-SceI, NsiI, AseI, CspCI, MfeI, BssS α I、DraIII、EcoP15I、AlwNI、BtsIMutI、NdeI、CviAII、FatI、NlaIII、FspEI、X cmI, BstXI, PflMI, BccI, NcoI, BseYI, FauI, TspMI, XmaI, LpnPI, AclI, ClaI 、SacII、HpaII、MspI、ScrFI、StyD4I、BsaJI、BslI、BtgI、NciI、AvrII、MnlI、 BbvCI, SbfI, Bpu10I, Bsu36I, EcoNI, HpyAV, BstNI, PspGI, StyI, BcgI, PvuI 、EagI、RsrII、BsiEI、BsiWI、BsmBI、Hpy99I、AbaSI、MspJI、SgrAI、BfaI、Bsp CNI、XhoI、PaeR7I、EarI、AcuI、PstI、BpmI、DdeI、SfcI、AflII、BpuEI、SmlI、 AvaI, BsoBI, MboII, BbsI, BsmI, EcoRI, HgaI, AatII, PflFI, Tth111I, AhdI DrdI, SacI, BseRI, PleI, HinfI, Sau3AI, MboI, DpnII, TfiI, BsrDI, BbvI, Bts α I、BstAPI、SfaNI、SphI、NmeAIII、NgoMIV、BglI、AsiSI、BtgZI、HhaI、HinP1I、BssHII、NotI、Fnu4HI、MwoI BmtI, NheI, BspQI, BlpI, TseI, ApeKI, Bsp1286I, AlwI, BamHI, BtsCI, FokI, FseI, SfiI, NarI, PluTI I、AscI、EciI、BsmFI、ApaI、PspOMI、Sau96I、KpnI、Acc65I、BsaI、HphI、BstEII、AvaII、BanI、BaeHGI and BsaHGI I、BanII、CviQI、BciVI、SalI、BcoDI、BsmAI、ApaLI、BsgI、AccI、Tsp45I、BsiHKAI、TspRI、ApoI、NspIF、BsrIF α I、BstYI、HaeII、EcoO109I、PpuMI、I-CeuI、I-SceI、BspHI、BspEI、MmeI、Taq α1, Hpy188I, Hpy188III, XbaI, BclI, PI-PspI, BsrGI, MseI, PacI, BstBI, PspXI, BsaWI, EaeI, HpyF30I, Sfr274I, or a combination thereof. [The present invention 1003] The process of any one of claims 1001 and 1002, further comprising carrying out (a) and (b) in the same reaction mixture. [The present invention 1004] 1004. The method of claim 1003, wherein (a) is carried out at a temperature different from (b). [The present invention 1005] The method of claim 1003, wherein (a) is carried out at the same temperature as (b). [The present invention 1006] Any of the methods of inventions 1001 to 1005, wherein distinguishing between a junction between an adaptor and a DNA fragment and a junction between an adaptor and another adaptor further comprises using an adaptor that is designed to be digested by a second restriction enzyme or enzymes when in a dimerized configuration, but that is incapable of being digested by the second restriction enzyme or enzymes when the adaptor is ligated to an end of the DNA fragment. [The present invention 1007] Any of the methods of the present inventions 1001 to 1006, wherein (d) comprises using a primer in the amplification step that can initiate polymerization at the junction between the adapter and the DNA fragment but cannot initiate polymerization at the junction between the adapter and another adapter. [The present invention 1008] 1. A method for preparing a library of nucleic acids, comprising: (a) digesting a plurality of DNA molecules with a first one or more restriction enzymes to generate DNA fragments; (b) ligating adapters to the DNA fragments by incubation with a ligase to produce a mixture of adapter-ligated DNA fragments and adapter dimers; and (c) amplifying the adaptor-ligated DNA fragments to produce amplified adaptor-ligated DNA fragments. Including, A method, subject to one or more of the following: (1) performing (c) with one or more primers that bind to the junction between the ends of the DNA fragments and the adapters, but not to the junction between the end of one adapter and the end of another adapter; (2) digesting the mixture of adaptor-ligated DNA fragments and adaptor dimers with a second restriction enzyme or enzymes that digest the junction between the end of one adaptor and the end of another adaptor but do not digest the junction between the end of the DNA fragment and the adaptor; (3) performing (a) and (b) in the same reaction mixture and further comprising the step of digesting the mixture with a second restriction enzyme or enzymes that digest the junction between the end of one adapter and the end of another adapter but do not digest the junction between the end of the DNA fragment and the adapter; (4) the adapters are adapter dimers by design, and the method further comprises the step of digesting the mixture of adapter-ligated DNA fragments and adapter dimers with a second restriction enzyme or enzymes that digest the junction between the end of one adapter and the end of another adapter but do not digest the junction between the end of the DNA fragment and the adapter; and / or (5) Amplified adapter dimers are produced by (c) that are digested with a third restriction enzyme or enzymes that digest the junction between the end of one adapter and the end of another adapter. [The present invention 1009] 1008. The method of claim 10, further comprising the step of distinguishing between methylated and unmethylated nucleobases in the adaptor-ligated fragments. [The present invention 1010] 1009. The method of claim 10, further comprising the step of subjecting the adaptor-ligated fragments to bisulfite conversion. [The present invention 1011] The method of any one of claims 1009 to 1010, further comprising the step of subjecting the adaptor-ligated fragments to one or more enzymatic and / or chemical reactions. [The present invention 1012] The method of claim 1011, further comprising the steps of oxidizing the methylated cytosine nucleobase and / or the hydroxymethylated cytosine nucleobase to produce an oxidation reaction product, followed by the steps of reducing and / or deaminating said oxidation reaction product. [The present invention 1013] 1012. The method of claim 1012, wherein the oxidizing step is carried out using a ten-eleven translocation (TET) enzyme. [The present invention 1014] The process of claim 1012, wherein the oxidizing step is carried out using potassium perruthenate. [The present invention 1015] The method of claim 1012, wherein the step of deaminating the oxidation reaction product is carried out using an apolipoprotein B mRNA editing enzyme catalytic polypeptide-like (APOBEC). [The present invention 1016] The method of claim 1012, wherein the step of reducing and / or deaminating the oxidation reaction product is carried out using pyridine borane. [The present invention 1017] The method of any of claims 1011 to 1016, further comprising the step of performing β-glucosyltransferase treatment prior to one or more enzymatic and / or chemical reactions. [The present invention 1018] 10. The method of any of claims 1008 to 1017, wherein some or all of the amplified, adaptor-ligated DNA fragments are analyzed, modified, or both. [The present invention 1019] The method of claim 1018, wherein the analyzing comprises sequencing. [The present invention 1020] The method of the present invention 1019, wherein the sequencing is next generation sequencing. [The present invention 1021] The method of claim 1020, further comprising performing targeted capture prior to next generation sequencing to further enrich for adaptor-ligated fragments. [The present invention 1022] 1022. The method of claim 1020 or 1021, further comprising the step of performing size selection prior to next generation sequencing to further enrich for adaptor-ligated fragments. [The present invention 1023] 1023. The method of any of claims 1018 to 1022, further comprising analyzing the amplified, adaptor-ligated DNA fragments to generate a methylation profile. [The present invention 1024] 10. The method of any one of claims 1008 to 1023, wherein in (1), (2), (3), or (5), the adapter comprises a GC (3' to 5') overhang. [The present invention 1025] 1025. The method of any of claims 1008 to 1024, wherein the first one or more restriction enzymes comprise MspI, HpaII, TaqαI, or functional analogs thereof, or mixtures thereof. [The present invention 1026] 1026. The method of any of claims 1008 to 1025, wherein the second one or more restriction enzymes comprise one or more of BspD1, XhoI, SmlI, HpyF30I, PaeR7I, Sfr274I, or functional analogs thereof, or mixtures thereof. [The present invention 1027] 1027. The method of any one of claims 1008 to 1026, wherein the ligase is T7 DNA ligase, T4 DNA ligase, T3 DNA ligase, Taq DNA ligase, or a functional analogue thereof, or a mixture thereof. [The present invention 1028] 1028. The method of any one of claims 1008 to 1027, wherein the plurality of DNA molecules comprises cell-free DNA. [The present invention 1029] The method of claim 1028, further comprising the step of obtaining cfDNA. [The present invention 1030] The method of claim 1029, wherein the cfDNA is obtained from or derived from a sample derived from a subject or individual. [The present invention 1031] The method of claim 1030, wherein the sample is obtained from or derived from plasma, serum, bone marrow, cerebrospinal fluid, pleural fluid, saliva, stool, or urine. [The present invention 1032] The method of any one of claims 1030 to 1031, further comprising the step of obtaining a sample from the subject or individual. [The present invention 1033] The method of any of claims 1008 to 1032, wherein the adaptor comprises a known sequence. [The present invention 1034] The method of any one of claims 1008 to 1032, wherein the adapter comprises a unique sequence. [This invention 1035] The method of any of claims 1008 to 1034, wherein the nucleic acid is enriched for molecules having one or more CpG sites. [The present invention 1036] 1. A method for preparing a library of nucleic acids, comprising the steps of: (a) digesting a plurality of DNA molecules with a first one or more restriction enzymes to generate DNA fragments; (b) ligating adapters to the DNA fragments by incubation with a ligase to produce a mixture of adapter-ligated DNA fragments and adapter dimers; and (c) amplifying the adaptor-ligated DNA fragments by utilizing one or more primers that bind to the junction between the ends of the DNA fragments and the adaptors but do not bind to the junction between the end of one adaptor and the end of another adaptor to produce amplified adaptor-ligated DNA fragments. [This invention 1037] 1036. The method of claim 1036, wherein the first one or more restriction enzymes comprise one or more of MspI, HpaII, TaqαI, or functional analogs thereof, or mixtures thereof. [The present invention 1038] The process of any one of claims 1036 to 1037, further comprising the step of carrying out (a) and (b) in the same reaction mixture. [This invention 1039] The method of any of claims 1036 to 1038, further comprising the step of distinguishing between methylated and unmethylated nucleobases in the adaptor-ligated fragment. [The present invention 1040] The method of claim 1039, further comprising the step of subjecting the adaptor-ligated fragments to bisulfite conversion. [The present invention 1041] The method of any one of claims 1039 to 1040, further comprising the step of subjecting the adaptor-ligated fragments to one or more enzymatic and / or chemical reactions. [The present invention 1042] The method of claim 1041, further comprising the steps of oxidizing the methylated cytosine nucleobase and / or the hydroxymethylated cytosine nucleobase to produce an oxidation reaction product, followed by the steps of reducing and / or deaminating the oxidation reaction product. [This invention 1043] The method of claim 1042, wherein the oxidizing step is carried out using a ten-eleven translocation (TET) enzyme. [This invention 1044] The process of claim 1042, wherein the oxidizing step is carried out using potassium perruthenate. [This invention 1045] The method of claim 1042, wherein the step of reducing and / or deaminating the oxidation reaction product is carried out using an APOBEC. [The present invention 1046] The process of claim 1042, wherein the step of reducing and / or deaminating the oxidation reaction product is carried out using pyridine borane. [This invention 1047] The method of any of claims 1041 to 1046, further comprising the step of performing β-glucosyltransferase treatment prior to one or more enzymatic or chemical reactions. [This invention 1048] 1048. The method of any one of claims 1036 to 1047, wherein the adapter comprises a GC overhang. [This invention 1049] 1. A method for preparing a library of nucleic acids, comprising the steps of: (a) digesting a plurality of DNA molecules with a first one or more restriction enzymes to generate DNA fragments; (b) ligating adapters to the DNA fragments by incubation with a ligase to produce a mixture of adapter-ligated DNA fragments and adapter dimers; (c) digesting the mixture of adaptor-ligated DNA fragments and adaptor dimers with a second restriction enzyme or enzymes that digest the junction between the end of one adaptor and the end of another adaptor but do not digest the junction between the end of the DNA fragment and the adaptor; and (d) amplifying the adaptor-ligated DNA fragments to produce amplified adaptor-ligated DNA fragments. [The present invention 1050] 1049. The method of claim 1049, wherein the first one or more restriction enzymes comprise one or more of MspI, HpaII, TaqαI, or functional analogs thereof, or mixtures thereof. [This invention 1051] 1050. The method of claim 1049 or 1050, wherein the second one or more restriction enzymes are one or more of BspDI, ClaI, AclI, NarI, XhoI, SmlI, HpyF30I, PaeR7I, Sfr274I, or functional analogs thereof, or mixtures thereof. [This invention 1052] The method of any of claims 1049 to 1051, further comprising the steps of carrying out (a), (b), and (c) in the same reaction mixture. [This invention 1053] The method of any of claims 1049 to 1052, further comprising the step of distinguishing between methylated and unmethylated nucleobases in the adaptor-ligated fragment. [This invention 1054] The method of claim 1053, further comprising the step of subjecting the adaptor-ligated fragments to bisulfite conversion. [This invention 1055] The method of claim 1053, further comprising the step of subjecting the adaptor-ligated fragments to one or more enzymatic and / or chemical reactions. [This invention 1056] The method of claim 1055, further comprising the step of oxidizing the methylated cytosine nucleobase and / or the hydroxymethylated cytosine nucleobase to produce an oxidation reaction product, followed by the step of reducing and / or deaminating the oxidation reaction product. [This invention 1057] The method of claim 1056, wherein the oxidizing step is carried out using a ten-eleven translocation (TET) enzyme. [This invention 1058] The process of claim 1056, wherein the oxidizing step is carried out using potassium perruthenate. [This invention 1059] The method of claim 1056, wherein the step of reducing and / or deaminating the oxidation reaction product is carried out using an APOBEC. [The present invention 1060] The process of claim 1056, wherein the step of reducing and / or deaminating the oxidation reaction product is carried out using pyridine borane. [The present invention 1061] The method of any of claims 1055 to 1060, further comprising the step of performing β-glucosyltransferase treatment prior to one or more enzymatic and / or chemical reactions. [This invention 1062] 1062. The method of any one of claims 1049 to 1061, wherein the adapter comprises a GC overhang. [This invention 1063] 1. A method for preparing a library of nucleic acids, comprising the steps of: (a) digesting a plurality of DNA molecules with a first one or more restriction enzymes to generate DNA fragments; (b) ligating the DNA fragments and the first adapters that are designed adapter dimers by incubating with a ligase to create second adapters and also to generate a mixture of DNA fragments ligated to the second adapters and adapter dimers of the second adapters, and subjecting the designed adapter dimers to the influence of a second restriction enzyme or enzymes that digest the junction between the end of one second adapter and the end of another second adapter but do not digest the junction between the end of the DNA fragment and the second adapter; and (c) amplifying the second adaptor-ligated DNA fragment to produce an amplified adaptor-ligated DNA fragment. [This invention 1064] 1064. The method of claim 1063, wherein the first one or more restriction enzymes comprise one or more of MspI, HpaII, TaqαI, or functional analogs thereof, or mixtures thereof. [This invention 1065] 1065. The method of claim 1063 or 1064, wherein the second one or more restriction enzymes comprise one or more of BspDI, ClaI, AclI, NarI, XhoI, SmlI, HpyF30I, PaeR7I, Sfr274I, or functional analogs thereof, or mixtures thereof. [The present invention 1066] The method of any one of claims 1063 to 1065, further comprising the step of carrying out (a) and (b) in the same reaction mixture. [This invention 1067] The method of any of claims 1063 to 1066, further comprising the step of distinguishing between methylated and unmethylated nucleobases in the DNA fragment ligated to the second adaptor. [The present invention 1068] The method of claim 1067, further comprising the step of subjecting the DNA fragment ligated to the second adaptor to bisulfite conversion. [This invention 1069] The method of claim 1067, further comprising the step of subjecting the DNA fragment ligated to the second adaptor to one or more enzymatic and / or chemical reactions. [The present invention 1070] The method of claim 1069, further comprising the steps of oxidizing the methylated cytosine nucleobase and / or the hydroxymethylated cytosine nucleobase to produce an oxidation reaction product, followed by the steps of reducing and / or deaminating the oxidation reaction product. [This invention 1071] 1070. The method of claim 1070, wherein the oxidizing step is carried out using a ten-eleven translocation (TET) enzyme. [This invention 1072] The process of claim 1070, wherein the oxidizing step is carried out using potassium perruthenate. [This invention 1073] The method of claim 1070, wherein the step of reducing and / or deaminating the oxidation reaction product is carried out using an APOBEC. [This invention 1074] The process of claim 1070, wherein the step of reducing and / or deaminating the oxidation reaction product is carried out using pyridine borane. [This invention 1075] The method of any of claims 1069 to 1074, further comprising the step of performing β-glucosyltransferase treatment prior to one or more enzymatic or chemical reactions. [This invention 1076] 1076. The method of any of claims 1063 to 1075, wherein the GC overhang is created by digestion of the adapter dimer of the second adapter with a second restriction enzyme or enzymes. [This invention 1077] 1. A method for preparing a library of nucleic acids, comprising the steps of: (a) digesting a plurality of DNA molecules with a first one or more restriction enzymes to generate DNA fragments; (b) ligating adapters to the DNA fragments to produce a mixture of adapter-ligated DNA fragments and adapter dimers; (c) amplifying the adaptor-ligated DNA fragments to produce a mixture of amplified adaptor-ligated DNA fragments and amplified adaptor-dimers; and (d) digesting the mixture of amplified, adaptor-ligated DNA fragments and amplified adaptor-dimers with a second restriction enzyme or enzymes that digest the junction between the end of one adaptor and the end of another adaptor but do not digest the junction between the end of the DNA fragment and the adaptor. [This invention 1078] 1078. The method of claim 1077, wherein the first one or more restriction enzymes comprise one or more of MspI, HpaII, TaqαI, or functional analogs thereof, or mixtures thereof. [This invention 1079] 1079. The method of claim 1077 or 1078, wherein the second one or more restriction enzymes comprise one or more of BspDI, ClaI, AclI, NarI, XhoI, SmlI, HpyF30I, PaeR7I, Sfr274I, or functional analogs thereof, or mixtures thereof. [The present invention 1080] The method of any of claims 1077 to 1079, further comprising the step of carrying out (a) and (b) in the same reaction mixture. [This invention 1081] The method of any of claims 1077 to 1080, further comprising the step of distinguishing between methylated and unmethylated nucleic acid bases in the adaptor-ligated DNA fragment. [This invention 1082] 1081. The method of claim 1081, further comprising the step of subjecting the adaptor-ligated fragments to bisulfite conversion. [This invention 1083] The method of claim 1081, further comprising the step of subjecting the adaptor-ligated fragments to one or more enzymatic and / or chemical reactions. [This invention 1084] The method of claim 1083, further comprising the steps of oxidizing the methylated cytosine nucleobase and / or the hydroxymethylated cytosine nucleobase to produce an oxidation reaction product, followed by the steps of reducing and / or deaminating the oxidation reaction product. [This invention 1085] The method of claim 1084, wherein the oxidizing step is carried out using a ten-eleven translocation (TET) enzyme. [This invention 1086] The process of claim 1084, wherein the oxidizing step is carried out using potassium perruthenate. [This invention 1087] The method of claim 1084, wherein the step of reducing and / or deaminating the oxidation reaction product is carried out using an APOBEC. [This invention 1088] The process of claim 1084, wherein the step of reducing and / or deaminating the oxidation reaction product is carried out using pyridine borane. [This invention 1089] The method of any of claims 1083 to 1088, further comprising the step of performing β-glucosyltransferase treatment prior to one or more enzymatic and / or chemical reactions. [The present invention 1090] 1089. The method of any one of claims 1077 to 1089, wherein the adapter comprises a GC overhang. It is specifically contemplated that any limitations discussed with respect to one embodiment of the invention may also apply to any other embodiment of the invention. Furthermore, any composition of the invention may be used in any method of the invention, and any method of the invention may be used to produce or utilize any composition of the invention. Aspects of one embodiment described in an example are also embodiments that may be practiced in conjunction with embodiments discussed elsewhere in another example or elsewhere in this application, e.g., in the Summary of the Invention, Detailed Description of the Embodiments, Claims, and Figure Legends.
[0045] The foregoing has outlined rather broadly the features and technical advantages of the present disclosure in order that the detailed description that follows may be better understood. Additional features and advantages of the present invention that form the subject of the claims of the present invention will be described hereinafter. It should be understood by those skilled in the art that the conception and specific embodiments disclosed may readily be utilized as a basis for modifying or designing other structures for carrying out the same purposes of the inventive design. It should also be understood by those skilled in the art that such equivalent constructions do not depart from the spirit and scope as set forth in the appended claims. The novel features believed characteristic of the presently disclosed design, both as to its organization and manner of operation, together with further objects and advantages, will be better understood by considering the following description in conjunction with the accompanying drawings. It is to be expressly understood, however, that each of the drawings is provided for the purpose of illustration and description only and is not intended as a definition of the limits of the disclosure. Additional objects, features, aspects, and advantages of the present invention will be set forth in part in the description that follows, and in part will be obvious from the description or may be learned by the practice of the invention. Various aspects of the disclosure have been described in sufficient detail to enable those skilled in the art to practice the invention, and it should be understood that other aspects may be utilized and changes may be made without departing from the scope of the invention. Therefore, the following detailed description is not to be taken in a limiting sense, and the invention of the present disclosure is best defined by the appended claims. [Brief explanation of the drawings]
[0046] The novel features of the present invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments in which the principles of the invention are utilized, as well as the accompanying drawings (also referred to herein as "Figure" and "FIG."), which are as follows:
[0047] [Figure 1]FIG. 1 shows an example of library generation from nucleic acids, such as cell-free DNA (cfDNA) and / or genomic DNA (gDNA), using restriction enzyme digestion and adapter ligation, followed by amplification of the generated molecules, such as by polymerase chain reaction (PCR). [Figure 2] 2A-2R provide examples of adapters that can be used in the methods of the present disclosure. [Figure 3] 3A-3E provide examples of adaptor-adaptor dimers that can be digested by restriction enzymes. [Figure 4] FIG. 4 illustrates one example of a method of the present disclosure in which cfDNA is enzymatically digested, followed by adapter ligation, bisulfite conversion, and PCR amplification, where PCR primers target the junctions between the adapters and the DNA fragments, but not the junctions between the adapters. [Figure 5] FIG. 5 illustrates one example of a method of the present disclosure in which cfDNA is enzymatically digested, followed by adapter ligation, bisulfite conversion, and PCR amplification, whereby the adapter dimers are selectively digested with an appropriate restriction enzyme. [Figure 6] FIG. 6 illustrates one example of a method of the present disclosure, in which cfDNA is enzymatically digested in the presence of a ligase and an adapter, followed by bisulfite conversion and PCR amplification, where one restriction enzyme digests the junction between the adapters, and the enzyme that originally digested the starting DNA also digests the ligated target DNA fragment without the adapter; the junction of the adapter and target DNA ligation product does not have a recognition site that can be digested by either enzyme. [Figure 7]FIG. 7 illustrates one example of a method of the present disclosure, in which cfDNA is enzymatically digested in the presence of a ligase and an adapter, followed by bisulfite conversion and PCR amplification, where one restriction enzyme (e.g., BspDI) digests the junction between the adapters, and the enzyme (e.g., MspI) that originally digested the starting DNA also digests the ligated target DNA fragment without the adapter; the junction of the adapter and target DNA ligation product does not have a recognition site that can be digested by either restriction enzyme (e.g., BspDI or MspI). [Figure 8] FIG. 8 illustrates one example of a method of the present disclosure in which cfDNA is enzymatically digested, followed by adapter ligation, bisulfite conversion, and PCR amplification, wherein the adapter dimers are digested after PCR amplification. [Figure 9A] Figures 9A-9C illustrate results from an example of the disclosed method performed on 10 nanograms (ng) of cfDNA from three plasma samples. The restriction enzyme MspI was used in the reaction. Both TBE-urea-polyacrylamide gel analysis of the library size (Figure 9A) and fragment size analysis based on the sequencing data (Figure 9B) show a typical reduced-representation bisulfite sequencing (RRBS) library with three characteristic peaks associated with Alu repeats at approximately 68 bp, 135 bp, and 202 bp. Figure 9C shows a summary of the sequencing results, including the total number of sequencing reads; the percentage of sequencing reads remaining untrimmed in the QC pipeline; the percentage of duplicates; R1 sequencing reads beginning with a CGG sequence; R2 sequencing reads beginning with a CGG sequence; and the percentage of sequencing reads mapping to the characteristic RRBS region. [Figure 9B] See the legend to Figure 9A. [Figure 9C] See the legend to Figure 9A.
[0048] While various aspects of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that such aspects are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the aspects of the present disclosure described herein may be employed. DETAILED DESCRIPTION OF THE INVENTION
[0049] Detailed Description I Definition Example In keeping with long-standing patent law practice, the words "a" and "an," when used in conjunction with the word "comprising" herein, including the claims, mean "one or more." Some embodiments of the present disclosure may consist of, or consist essentially of, one or more elements, method steps, and / or methods of the present disclosure. It is contemplated that any method or composition described herein can be implemented with respect to any other method or composition described herein, and that various embodiments may be combined.
[0050] As used herein, the terms "or" and "and / or" are used to describe multiple components in combination or mutually exclusive. For example, "x, y, and / or z" can mean "x" only, "y" only, "z" only, "x, y, and z," "(x and y) or z," "x or (y and z)," or "x or y or z." It is specifically contemplated that x, y, or z may be specifically excluded from an embodiment.
[0051] Throughout this application, the term "about" is used in accordance with its clear and ordinary meaning within the field of cell and molecular biology to indicate that a value includes the standard deviation of error for the device or method being employed to determine the value.
[0052] The term "comprising" is synonymous with "including," "containing," or "characterized by" and is inclusive or open-ended and does not exclude additional, unrecited elements or method steps. The phrase "consisting of" excludes any element, step, or ingredient not specified. The phrase "consisting essentially of" limits the scope of the described subject matter to the specified materials or steps and those that do not materially affect its basic and novel characteristics. It is contemplated that embodiments described in the context of the term "comprising" may also be implemented in the context of the terms "consisting of" or "consisting essentially of."
[0053] Reference throughout this specification to "one embodiment," "an embodiment," "a particular embodiment," "a related embodiment," "a certain embodiment," "an additional embodiment," or "a further embodiment," or combinations thereof, means that a particular feature, structure, or attribute described in connection with that embodiment is included in at least one embodiment of the invention. Thus, the appearances of such phrases in various places throughout this specification are not necessarily all referring to the same embodiment. Furthermore, particular features, structures, or attributes may be combined in any suitable manner in one or more embodiments.
[0054] Various aspects of the present invention can be described in range format. It should be understood that the description in range format is merely for convenience and brevity and should not be construed as an inflexible limitation on the scope of the present disclosure. Thus, the description of a range should be considered to have specifically disclosed all possible subranges and individual numerical values within that range, as if they were explicitly written out in full. For example, a description of a range such as 1 to 6 should be considered to have specifically disclosed subranges such as 1 to 3, 1 to 4, 1 to 5, 2 to 4, 2 to 6, 3 to 6, etc., as well as individual numbers within that range, for example, 1, 2, 3, 4, 5, and 6. This is true regardless of the breadth of the range. When ranges are present, they may include both ends of the range.
[0055] The term "adapter dimer," as used herein, refers to the molecule that results when a first adapter molecule is ligated to a second adapter molecule.
[0056] As used herein, the term "subject" generally refers to an individual having a biological sample to be processed or analyzed. The subject can be an animal or a plant. The subject can be a mammal, such as a human, dog, cat, horse, pig, or rodent. The subject can be, for example, a patient who has, is suspected of having, or is at risk of having a disease or disorder. These diseases or disorders can be, for example, one or more cancers (e.g., brain cancer, breast cancer, cervical cancer, colorectal cancer, endometrial cancer, esophageal cancer, gastric cancer, hepatobiliary cancer, leukemia, liver cancer, lung cancer, lymphoma, ovarian cancer, pancreatic cancer, skin cancer, urinary tract cancer, testicular cancer, kidney cancer, sarcoma, bile duct cancer, thyroid cancer, gallbladder cancer, spleen cancer, or prostate cancer, and the cancer may or may not include a solid tumor), one or more infectious diseases, one or more genetic disorders, or one or more tumors, or any combination thereof. For a subject that has or is suspected to have one or more tumors, the tumors can be of one or more types.The subject can have disease or disorder, or be suspected to have said disease or disorder.The subject can not have disease or disorder, or not be suspected to have said disease or disorder.The subject can be a healthy control.The subject can be asymptomatic for a specific disease or disorder.
[0057] As used herein, the term "sample" generally refers to a biological sample. A sample may be taken from tissues and / or cells or from the environment of tissues and / or cells. In some examples, a sample may include or be derived from a tissue biopsy, a cell biopsy, blood (e.g., whole blood), plasma, serum, bone marrow, cerebrospinal fluid, pleural fluid, saliva, stool, urine, extracellular fluid, dried blood spot, cultured cells, culture medium, discarded tissue, plant matter, synthetic protein, bacterial and / or viral sample, fungal tissue, archaea, or protozoa. A sample may be isolated from a source before collection. A sample may include forensic evidence. Non-limiting examples include fingerprints, saliva, urine, blood, stool, semen, or other bodily fluids isolated from a primary source before collection. In some examples, a sample is isolated from its primary source (cells, tissues, bodily fluids such as blood, environmental samples, etc.) during sample preparation. The sample may be from an extinct species, including, but not limited to, samples from fossils. The sample may or may not be purified or otherwise enriched from its primary source. In some embodiments, the primary source is homogenized before further processing. The sample may be filtered or centrifuged to remove buffy coat, lipids, or particulate matter. The sample may also be purified or enriched for nucleic acids, or treated with RNase or DNase. The sample may include intact, fragmented, or partially degraded tissues and / or cells.
[0058] The sample may be obtained from a subject with a disease or disorder, and the subject may or may not have been diagnosed with the disease or disorder. The subject may need a second opinion. The disease or disorder may be an infectious disease, an immune disorder or disease, cancer, a genetic disease, a degenerative disease, a lifestyle-related disease, or an injury. The infectious disease may be caused by bacteria, viruses, fungi, and / or parasites. Non-limiting examples of cancer include pancreatic cancer, liver cancer, lung cancer, colorectal cancer, leukemia, bladder cancer, bone cancer, brain cancer, breast cancer, cervical cancer, endometrial cancer, esophageal cancer, gastric cancer, head and neck cancer, melanoma, ovarian cancer, testicular cancer, kidney cancer, thyroid cancer, gallbladder cancer, spleen cancer, and prostate cancer. Some examples of genetic diseases or disorders include, but are not limited to, cystic fibrosis, Charcot-Marie-Tooth disease, Huntington's disease, Peutz-Jeghers syndrome, Down syndrome, rheumatoid arthritis, and Tay-Sachs disease. Non-limiting examples of lifestyle-related diseases include obesity, diabetes, arteriosclerosis, heart disease, stroke, high blood pressure, cirrhosis of the liver, nephritis, cancer, chronic obstructive pulmonary disease (COPD), hearing problems, and chronic back pain. Some examples of injuries include, but are not limited to, abrasions, brain injuries, contusions, burns, concussions, congestive heart failure, construction injuries, dislocations, flail chest, fractures, hemothorax, herniated discs, iliac crest contusions, hypothermia, lacerations, nerve compression, pneumothorax, rib fractures, sciatica, spinal cord injuries, tendon, ligament, and fascial injuries, traumatic brain injury, and whiplash. The sample may be collected before and / or after treatment of a subject with a disease or disorder. The sample may be collected before and / or after treatment of the subject's disease or disorder. The sample may be collected during the course of treatment or treatment regimen. Multiple samples may be collected from a subject to monitor the effectiveness of treatment over a period of time, including beginning before treatment begins. The sample may be collected from a subject known or suspected to be suffering from an infectious disease for which a diagnostic antibody may or may not be available. The sample may be collected from a subject to monitor aberrant tissue-specific cell death or organ transplantation.
[0059] The sample may be collected from a subject suspected of having a disease or disorder. The sample may be collected from a subject experiencing unexplained symptoms, such as fatigue, nausea, weight loss, pain, soreness, weakness, or memory loss. The sample may be collected from a subject with explained symptoms. The sample may be collected from a subject at risk of developing a disease or disorder due to one or more factors, such as family history and / or personal history, age, environmental exposure, lifestyle risk factors, the presence of other known risk factors, or a combination thereof.
[0060] Samples may be collected from healthy subjects or individuals. In some embodiments, samples may be collected from the same subject or individual over a long period of time. In some embodiments, samples obtained over a long period of time may be analyzed to monitor an individual's health status and for early detection of health problems (e.g., early diagnosis of cancer). In some embodiments, samples may be collected at home or at a care site and then transported by mail, courier, or other transportation method before analysis. For example, a home user may collect a blood spot sample by pricking their fingertip, which may be dried and then transported by mail before analysis. In some embodiments, samples obtained over a long period of time may be used to monitor response to stimuli expected to affect health status, athletic performance, or cognitive performance. Non-limiting examples include response to drug therapy, diet, and / or exercise therapy. In some embodiments, individual samples are versatile and allow for methylation profiling to obtain clinically relevant information, but are also used for information about an individual's personal or family ancestry. In some embodiments, samples may be collected from a pregnant woman and / or her fetus.
[0061] In some embodiments, biological sample is a nucleic acid sample containing one or more nucleic acid molecules.Nucleic acid molecule can be cell-free or substantially cell-free nucleic acid molecule, for example, cell-free DNA (cfDNA) or cell-free RNA (cfRNA), or a mixture thereof.Nucleic acid molecule can be derived from various sources, including human, mammal, non-human mammal, ape, monkey, chimpanzee, reptile, amphibian, or bird.In addition, sample can be extracted from various animal body fluids containing cell-free sequences, including but not limited to blood, serum, plasma, bone marrow, vitreous humor, sputum, feces, urine, tears, sweat, saliva, semen, mucous excretion, mucus, cerebrospinal fluid, pleural fluid, amniotic fluid, and lymphatic fluid.Sample can be collected from embryo, fetus, or pregnant woman.In some examples, sample can be isolated from maternal plasma. In some examples, the sample may contain cell-free nucleic acids (e.g., cfDNA) derived from a fetus (via a bodily sample obtained from the pregnant subject) or from the subject's own tissues or cells.
[0062] Components of a sample (including nucleic acids) may be tagged, e.g., with identifiable tags, e.g., to allow for sample multiplexing. Some non-limiting examples of identifiable tags include fluorophores, magnetic nanoparticles, and nucleic acid barcodes. Fluorophores may include fluorescent proteins such as GFP, YFP, RFP, eGFP, mCherry, tdtomato; FITC, Alexa Fluor 350, Alexa Fluor 405, Alexa Fluor 488, Alexa Fluor 532, Alexa Fluor 546, Alexa Fluor 555, Alexa Fluor 568, Alexa Fluor 594, Alexa Fluor 647, Alexa Fluor 680, Alexa Fluor 750, Pacific Blue, coumarin, BODIPY FL, Pacific Green, Oregon Green, Cy3, Cy5, Pacific Orange, TRITC, Texas Red, phycoerythrin, allophcocyanin, or other fluorophores. One or more barcode tags may be attached (e.g., by coupling or ligation) to cell-free nucleic acids (e.g., cfDNA) in the sample prior to sequencing. The barcode may uniquely tag the cfDNA molecules in the sample. Alternatively, the barcode may non-uniquely tag the cfDNA molecules in the sample. The barcode may non-uniquely tag the cfDNA molecules in the sample, so that additional information obtained from the cfDNA molecules (for example, at least one part of the endogenous sequence of the cfDNA molecules) can be combined with the non-unique tag to function as a unique identifier for the cfDNA molecules in the sample (for example, to uniquely identify them compared to other molecules). For example, the cfDNA sequence reads with unique identity (for example, derived from a given template molecule) can be detected based at least in part on the sequence information comprising one or more consecutive base regions at one or both ends of the sequence reads, the length of the sequence reads, and / or the sequence of the barcodes attached to one or both ends of the sequence reads.DNA molecules can be uniquely identified without tagging by partitioning a DNA (e.g., cfDNA) sample into a large number (e.g., at least about 50, at least about 100, at least about 500, at least about 1000, at least about 5000, at least about 10000, at least about 5000, at least about 100000, at least about 50000, or at least about 100000) of different individual subunits (e.g., compartments, wells, or droplets) prior to amplification such that the amplified DNA molecules can be uniquely distinguished and identified as derived from each individual input molecule of DNA.
[0063] Any number of samples can be multiplexed. For example, multiplex analysis can include at least about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 25, about 30, about 35, about 40, about 45, about 50, about 55, about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, about 100 or more samples, or any range derivable therein. Identifiable tags can provide a way to determine the origin of each sample, or can direct various samples to separate into different regions or solid supports.
[0064] Any number of samples may be mixed prior to analysis without tagging or multiplexing. For example, a multiplex analysis may include at least about 2, about 3, about 4, about 5, about 6, about 7, about 8, about 9, about 10, about 11, about 12, about 13, about 14, about 15, about 16, about 17, about 18, about 19, about 20, about 25, about 30, about 35, about 40, about 45, about 50, about 55, about 60, about 65, about 70, about 75, about 80, about 85, about 90, about 95, about 100, or more samples, or any range derivable therein. Samples may be multiplexed without tagging using a combinatorial pooling design, where samples are mixed into pools in a manner that allows for the signal from individual samples to be separated from the pool being analyzed using computational multiplexing.
[0065] The sample may be enriched before sequencing. For example, cfDNA molecules may be selectively enriched or non-selectively enriched in one or more regions from the subject's genome or transcriptome. For example, cfDNA molecules may be selectively enriched in one or more regions from the subject's genome or transcriptome by targeted sequence capture (e.g., using a panel), selective amplification, or targeted amplification. As another example, cfDNA molecules may be non-selectively enriched in one or more regions from the subject's genome or transcriptome by global amplification. In some embodiments, amplification includes global amplification, whole genome amplification, or non-selective amplification. cfDNA molecules may be size-selected to obtain fragments having a predetermined range of lengths. For example, size selection can be performed on DNA fragments prior to adapter ligation to obtain lengths ranging from 40 base pairs (bp) to 250 bp, or any range derivable therein. As another example, size selection can be performed on DNA fragments after adapter ligation to obtain lengths ranging from 160 bp to 400 bp, or any range derivable therein.
[0066] As used herein, the term "nucleic acid" or "polynucleotide" generally refers to a molecule comprising one or more nucleic acid subunits or nucleotides. A nucleic acid may comprise one or more nucleotides selected from adenosine (A), cytosine (C), guanine (G), thymine (T), and uracil (U), or variants thereof. Typically, a nucleotide comprises a nucleoside and at least one, two, three, four, five, six, seven, eight, nine, ten, or more phosphate (PO) groups, or any range derivable therein. A nucleotide may comprise a nucleobase, a five-carbon sugar (either ribose or deoxyribose), and one or more phosphate groups, either individually or in combination.
[0067] As used herein, the terms "nucleic acid molecule," "nucleic acid sequence," "nucleic acid fragment," "oligonucleotide," and "polynucleotide" generally refer to polynucleotides, such as deoxyribonucleotides (DNA) or ribonucleotides (RNA), or analogs and / or combinations thereof (e.g., mixtures of DNA and RNA). Nucleic acid molecules can be of various lengths. Nucleic acid molecules can be at least 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 200, 300, 400, 500, 1 kilobase (kb), 2 kb, 3 kb, 4 kb, 5 kb, 10 kb, or 50 kb in length, or any range derivable therein, or the nucleic acid molecule can have any number of bases between any two of the aforementioned values. Oligonucleotides typically consist of specific sequences of the four nucleotide bases: adenine (A); cytosine (C); guanine (G); and thymine (T) (if the polynucleotide is RNA, uracil (U) is substituted for thymine (T)). Thus, the terms "nucleic acid molecule," "nucleic acid sequence," "nucleic acid fragment," "oligonucleotide," and "polynucleotide" are intended, at least in part, to be alphabetical representations of polynucleotide molecules. Alternatively, these terms may be applied to the polynucleotide molecule itself. This alphabetical representation can be input into a database in a computer having a central processing unit and / or used for bioinformatics applications such as functional genomics and homology searching. An oligonucleotide may contain one or more non-standard nucleotides, nucleotide analogs, and / or modified nucleotides.
[0068] As used herein, term " cell-free DNA " or " cfDNA " refers to the DNA that normally circulates in free form in bodily fluids, such as bloodstream or the plasma derived from it.In some embodiments of the method used herein, cfDNA encompasses certain types of cfDNA, such as circulating tumor DNA (ctDNA), which is the tumor-derived fragmented DNA in bloodstream that is not associated with cells.cfDNA can be double-stranded or single-stranded, or have both characteristics.
[0069] As used herein, the term "CpG site" generally refers to a position on a nucleic acid molecule that contains adjacent cytosine (C) and guanine (G) in the 5' to 3' direction. A nucleic acid molecule may contain at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 100, 500, 1000, 10000, or more CpG sites, or any range of numbers derivable therein. Such CpG sites in the 3' to 5' direction of a nucleic acid molecule may be referred to as "GpC sites."
[0070] As used herein, the term "CpG island" generally refers to a contiguous region of genomic DNA that meets the following criteria: (1) an "observed-to-expected ratio" of CpG dinucleotides greater than about 0.6; (2) a "GC content" greater than about 0.5; and (3) a length of at least about 0.2 kilobases (kb). However, there may be exceptions where repetitive regions that meet these criteria are excluded or ignored. Criteria for identifying CpG islands are described, for example, by Gardiner-Garden et al. (J. Mol. Biol., 196:262-282, 1987), the entire contents of which are incorporated herein by reference.
[0071] As used herein, the term "CpG-rich" generally refers to a genomic region with a high CpG content, where most of the DNA methylation may occur.A region with a high CpG content may have a CpG content of at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, or any range derivable therein, or higher.In some embodiments, such a CpG content is higher than 1%.In some embodiments, a CpG-rich region may include a CpG island and a promoter region.A CpG-rich region may include any length (for example, but not limited to, at least 0.2 kb).
[0072] As used herein, the term "bisulfite conversion" generally refers to a biochemical process for converting an unmethylated base (e.g., a cytosine base) to a uracil base, while preserving the methylation information (e.g., a methylated cytosine). Examples of reagents for bisulfite conversion include sodium bisulfite, magnesium bisulfite, and trialkylammonium bisulfite.
[0073] II. Method Examples The present disclosure provides a method for preparing a nucleic acid library, which has improvements to enrich informative fragments and reduce adapter dimers, which can reduce the efficiency of library preparation. In some embodiments, the DNA source from which the library is made includes any type of DNA, particularly cell-free DNA (cfDNA). In some embodiments, the library is made after manual DNA fragmentation. However, in some embodiments, the starting nucleic acid material itself may contain fragments (such as fragmentation of natural sources (including cancer cell DNA, cell apoptosis, or necrosis)).
[0074] The present disclosure provides methods that use a series of steps to generate a desired library. In some embodiments, the method includes digesting DNA, ligating adapters to the ends of the digested DNA, amplifying the adapter-ligated DNA, and sequencing the amplified adapter-ligated DNA, whereby adapter-dimers are generated as by-products at some point during the method steps, and the adapter-dimers are also reduced in number (e.g., by digesting or otherwise destroying them) to increase the efficiency of the method. The method may also include bisulfite conversion when it is desired to measure the methylation state of nucleic acids, including when it is desired to generate a methylome, for example.
[0075] In some embodiments, for example, because the read length of current next-generation sequencing (NGS) sequencer is limited, DNA fragmentation is carried out for preparing sequencing library.In the previous method for preparing NGS library, fragmented DNA can be end-repaired and / or added with dA tail before being linked to adapter, and typical adapter can be blunt-ended or have overhang.
[0076] As shown in Figure 1, previous library methods utilize enzymatic approaches, such as the use of one or more restriction enzymes, to fragment DNA with overhangs at both ends. Specifically designed adapters are ligated to the fragments. The adapter-tagged (also called adapter-ligated) fragments can then be amplified by PCR, with or without adapter removal, to sufficient quantities for use in NGS methods.
[0077] As the first or initial step in these methods, DNA digestion with restriction enzymes can be used not only to fragment DNA but also as a general approach to enrich for DNA of interest. For example, the use of MspI enzyme in reduced-representation bisulfite sequencing (RRBS) enriches DNA fragments containing CpG sites for methylation profiling. While most fragments generated from genomic DNA (gDNA) may have been cut twice by restriction enzymes after size selection, this may not be the case for cfDNA fragments. Typically, cell-free DNA contains DNA molecules with a size distribution centered around 166 base pairs. As shown in Figure 1, some fragments contain no restriction enzyme recognition sites, while others contain only one restriction enzyme recognition site. Therefore, such fragments are uninformative or less informative than target fragments with multiple restriction enzyme recognition sites. Some restriction enzymes create 3' or 5' overhangs after digestion. FIG. 1 shows two representative overhangs ("PQ" and "NM") generated by cleavage with two restriction enzymes and their complementary sequences ("pq" and "nm").
[0078] In some embodiments, the present disclosure provides a specifically designed adapter for library preparation that is ligated to enzyme-digested DNA fragments, thereby enriching fragments with multiple restriction enzyme recognition sites from cfDNA, particularly for genome and epigenomic analysis.As illustrated in Figure 1, these adapters are designed with overhangs (e.g., "MN", "pq") that are complementary to the overhangs of the restriction enzyme-digested fragments.After the adapter is ligated to the target fragment, subsequent library preparation steps, such as PCR, selectively amplify only the fragments that are ligated at both ends to the specifically designed adapter, and thus can enrich informative fragments from cfDNA for subsequent analysis, including sequencing.
[0079] One problem encountered with previous library preparation methods, such as the example shown in FIG. 1 , and other library preparation methods is the unwanted formation of numerous adapter dimers. The method of the present disclosure overcomes at least this problem. Unlike conventional adapters, whose overhangs are not complementary to each other and therefore cannot easily form dimers, the adapters used to ligate to fragments digested by restriction enzymes can ligate to each other to form adapter dimers. The adapter dimers in the final library may be sequenced, which may adversely affect the yield of sequenced reads for the target fragments and may result in spurious sequenced reads resulting from the adapter dimers. The adapter dimers in the generated library may be present in large quantities and cannot be easily removed by a simple purification step, such as Ampure bead purification.
[0080] The present disclosure provides a method for avoiding or reducing the amount of adapter-dimer in nucleic acid library, including the library that is prepared for sequencing.In some embodiments, the method of the present disclosure uses PCR primers that are specifically designed to selectively amplify the DNA fragments that have adapters at both ends, and / or uses one or more restriction enzymes to cut adapter-dimer during or after adapter ligation.These methods are exemplified in the following specific applications when the method of the present disclosure is used to prepare nucleic acid library.After preparation, the library can be used for any purpose, for example, for profiling cfDNA, including cfDNA methylome.
[0081] Figure 2 illustrates examples of engineered adapters that can be used to enrich informative fragments from restriction enzyme-digested cfDNA. A variety of restriction enzymes can be used to generate informative cfDNA fragments, including but not limited to AcII, HindIII, MluCI, PciI, AgeI, BspMI, BfuAI, SexAI, MluI, BceAI, HpyCH4IV, HpyCH4III, BaeI, BsaXI, AflIII, SpeI, BsrI, BmrI, BglII, BspDI, PI-SceI, NsiI, AseI, CspCI, MfeI, BssSαI, DraIII, EcoP15I, AlwNI, Bts IMutI, NdeI, CviAII, FatI, NlaIII, FspEI, XcmI, BstXI, PflMI, BccI, NcoI, BseYI, FauI, TspMI, XmaI, LpnPI, AclI, ClaI, SacII, HpaII, MspI , ScrFI, StyD4I, BsaJI, BslI, BtgI, NciI, AvrII, MnlI, BbvCI, SbfI, Bpu10I, Bsu36I, EcoNI, HpyAV, BstNI, PspGI, StyI, BcgI, PvuI, EagI, Rsr II, BsiEI, BsiWI, BsmBI, Hpy99I, AbaSI, MspJI, SgrAI, BfaI, BspCNI, XhoI, PaeR7I, EarI, AcuI, PstI, BpmI, DdeI, SfcI, AflII, BpuEI, SmlI, AvaI, BsoBI, MboII, BbsI, BsmI, EcoRI, HgaI, AatII, PflFI, Tth111I, AhdI, DrdI, SacI, BseRI, PleI, HinfI, Sau3AI, MboI, DpnII, TfiI, BsrD I, BbvI, BtsαI, BstAPI, SfaNI, SphI, NmeAIII, NgoMIV, BglI, AsiSI, BtgZI, HhaI, HinP1I, BssHII, NotI, Fnu4HI, MwoI, BmtI, NheI, BspQI, Bl pI, TseI, ApeKI, Bsp1286I, AlwI, BamHI, BtsCI, FokI, FseI, SfiI, NarI, PluTI, KasI, AscI, EciI, BsmFI, ApaI, PspOMI, Sau96I, KpnI, Acc65I,BsaI, HphI, BstEII, AvaII, BanI, BaeGI, BsaHI, BanII, CviQI, BciVI, SalI, BcoDI, BsmAI, ApaLI, BsgI, AccI, Tsp45I, BsiHKAI, TspRI, ApoI, NspI, BsrFαI, BstYI, HaeII, Eco Includes O109I, PpuMI, I-CeuI, I-SceI, BspHI, BspEI, MmeI, TaqαI, Hpy188I, Hpy188III, XbaI, BclI, PI-PspI, BsrGI, MseI, PacI, BstBI, PspXI, BsaWI, EaeI, HpyF30I, Sfr274I. It is contemplated that in certain embodiments, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more of these, or any range derivable therein, may be excluded.
[0082] The adapters used for adapter ligation in the disclosed methods can be of any type, but in some embodiments, the adapters are specifically designed to correspond to the ends of the fragments to which they can be ligated. In some embodiments, the adapters are configured to ligate to nucleic acid molecules digested by enzymes. For example, once the starting nucleic acid material (e.g., cfDNA from a sample) is obtained, the nucleic acid is digested with one or more specific enzymes. The enzyme(s) may be selected to enrich for a specific type of nucleic acid molecule (e.g., CpG-rich), to generate fragments substantially within a specific size range, or a combination thereof. As an example, the initial cfDNA may be digested with MspI. In such cases, the adapters correspond to the MspI-digested DNA ends of the DNA, which contain CG (5' to 3') overhangs at the digested ends. In such embodiments, the adapters have GC (3' to 5') overhangs at the ends so that they are complementary to and can be ligated to the MspI-digested DNA ends.
[0083] The adapters in Figure 2 are merely examples of adapters that can be used in the methods of the present disclosure. Adapters for the methods of the present disclosure may include standard adapters having overhangs with sequences complementary to the overhangs of restriction enzyme-digested fragments (e.g., Figures 2A, 2B, 2C, 2D, and 2E), or standard adapters having sequences complementary to the overhangs of restriction enzyme-digested fragments plus random and / or fixed sequences (e.g., Figures 2F, 2G, 2H, 2I, 2J, 2K, 2L, 2M, 2N, and 2O). In some embodiments, the adapters can be in the form of adapter-adapter dimers, including, for example, those adapters described above (e.g., Figures 2P, 2Q, and 2R); this type of adapter is specifically designed to be an adapter dimer, as opposed to one generated as a by-product of a method such as the methods of the present disclosure.
[0084] In some embodiments, adapter is designed so that when adapter dimer with two molecules of adapter is generated, it can be digested by restriction enzyme.In addition to this feature, adapter can be designed so that when adapter is ligated to the end of cfDNA fragment, the restriction enzyme that digests adapter dimer itself cannot digest the junction between adapter and the end of digested cfDNA fragment.Figure 3A-3E show some examples of adapter dimer that can be digested by restriction enzyme.
[0085] Figure 4 illustrates the use of an exemplary library preparation method of the present disclosure for methylation profiling of cfDNA for applications such as cancer diagnosis. For example, the restriction enzyme MspI is used to digest cfDNA at the recognition site CCGG. Fragments cleaved at both ends by MspI contain CpG-rich sites that are useful for methylation profiling. Designed adapters with GC (3' to 5') overhangs can be ligated to the ends cleaved by MspI. The digested fragments are then subjected to bisulfite treatment to distinguish methylated nucleobases from unmethylated nucleobases. After bisulfite conversion, PCR can be performed to enrich fragments ligated at both ends to specifically designed adapters for sequencing and methylation profiling.
[0086] In the example of Figure 4, an adapter with a GC overhang can form an adapter dimer with a 5'-CTCGAG-3' sequence at the junction of the adapter dimer. However, as a product of restriction enzyme digestion (in this example, MspI), the junction between the adapter and the end of the target DNA fragment contains a different sequence, such as 5'-CTCGG-3' (or 5'-CTTGG-3' after bisulfite conversion if the central C is not methylated). In some embodiments, PCR primers are designed to recognize the junction between the adapter and the target DNA but not the junction between the adapter molecules in the adapter dimer. With this design, only ligation products with a target DNA fragment inserted between the adapters can be amplified and enriched for, for example, sequencing.
[0087] Figure 5 illustrates an example of a method for enzymatic library preparation, for example, for the purpose of cfDNA methylome profiling. In this example, after MspI digestion, target DNA fragments with GC overhangs at their 5' ends are directly ligated to adapters with GC overhangs at their 5' ends. Other restriction enzymes (e.g., Xhol, SmlI, and TaqαI) that recognize the junction sequence between the adapters but not between the adapters and the target DNA are used to cleave the adapter dimers after the ligation reaction. The ligation products digested by the restriction enzymes can then be subjected to bisulfite conversion and PCR enrichment.
[0088] Figure 6 illustrates an example method for enzymatic library preparation, including, for example, for cfDNA methylome profiling. In this example, enzymatic digestion and adapter ligation are performed in the same reaction. In this reaction, the restriction enzyme MspI, for example, generates target DNA fragments with GC (3' to 5') overhangs at their 5' ends. Adapters with GC overhangs at their 5' ends are ligated to the target DNA fragments in the presence of DNA ligase in the same reaction. Adapters and adapters, or fragments digested with MspI, may also be ligated to each other. To avoid adapter-adapter ligation, which would result in adapter dimer formation, the restriction enzyme BspDI, which can recognize and cleave the junction between the adapters, is added to the reaction. Meanwhile, MspI in the mixture can digest the ligated target DNA fragments. However, the junction between the adapter and the target DNA ligation product does not have a recognition site that can be digested by a restriction enzyme such as BspDI or MspI. Restriction enzyme digestion and adapter ligation in this reaction can be performed at the same or different temperatures.
[0089] Figure 7 illustrates an example of a method for enzymatic library preparation for cfDNA methylome profiling. Similar to the method illustrated in Figure 6, enzymatic digestion and adapter ligation are performed in the same reaction. The adapters used in this reaction are synthesized in the form of adapter-adapter dimers, for example, as shown in Figures 2P, 2Q, and 2R, for long-term stability. In the disclosed method, the adapters, which are adapter dimers used to perform the method steps, are specifically designed manually (rather than being created as part of the method steps). The BspDI restriction enzyme can digest this type of adapter to form an adapter product that can be used to ligate to a target DNA fragment.
[0090] Enzymatic digestion of the adapter dimers can be performed after bisulfite conversion or after PCR amplification. As shown in Figure 8, the restriction enzyme SmlI (as an example) can be used to cleave the adapter dimers after the PCR enrichment step.
[0091] 9A-9C illustrate an example of the preparation of a typical reduced representation bisulfite sequencing (RRBS) library from cfDNA based on the methods of the present disclosure. In this example, enzymatic library preparation was performed to generate three RRBS libraries from cfDNA extracted from patient plasma. The complete protocol includes: 1. Enzyme Reaction: The enzyme reaction solution may contain 10 nanograms (ng) of cfDNA, HO, 10x CutSmart, ATP, DTT, PEG, adapter, MspI, BspDI, and ligase. After mixing, the solution is placed in a thermal cycler and run through the following program: 17 cycles of 30 minutes at 37°C, 30 minutes at 25°C, 90 minutes at 37°C, and infinity at 4°C. The enzyme reaction product is purified using Ampure XP beads (Beckman Coulter). 2 Bisulfite conversion: Bisulfite conversion can be performed using the EpiTect Bisulfite Kit (Qiagen) according to the manufacturer's protocol. 3. PCR Amplification: The bisulfite-converted products are amplified to enrich for adapter-containing fragments in the final library. The PCR reaction solution may contain the bisulfite-converted products, NEB index primer, NEB universal primer, KAPA HiFI uracil ready mix, and HO. After mixing, the solution is placed in a thermal cycler and run the following program: 98°C for 45 minutes, 15 cycles of (98°C for 15 minutes, 60°C for 30 minutes, 72°C for 30 minutes), 72°C for 60 minutes, and 4°C for infinity. The PCR reaction products are purified using Ampure XP beads (Beckman Coulter), and the purified library is ready for sequencing on platforms such as the Illumina HiSeq 2000.
[0092] In this example, the restriction enzyme MspI is used in an enzymatic reaction to digest cfDNA fragments. After bisulfite conversion and PCR amplification, this example method produces a sequencing library equivalent to a conventional RRBS library prepared from intact DNA. As shown in Figures 9A-9C, both TBE-urea-polyacrylamide gel analysis of the library size (Figure 9A) and fragment size analysis based on the sequencing data (Figure 9B) indicate a typical RRBS library with three characteristic peaks associated with Alu repeats at around 68 bp, 135 bp, and 202 bp. Figure 9C shows a summary of the sequencing results, including the total number of sequencing reads; the percentage of sequencing reads remaining untrimmed in the QC pipeline; the percentage of duplicates; R1 sequencing reads beginning with a CGG sequence; R2 sequencing reads beginning with a CGG sequence; and the percentage of sequencing reads mapping to the characteristic RRBS region.
[0093] Aspects of the present disclosure include at least methods for preparing a nucleic acid library, methods for making a plurality of polynucleotides, methods for making double-stranded DNA, the use of nucleotides to generate libraries, methods for preparing sequencing libraries, methods for applying sequencing libraries, and the like.
[0094] Embodiments include methods that include one, two, three, four, five, six, seven, eight, nine, ten, or more of any of the following steps, or any range derivable therein: providing a plurality of DNA molecules, isolating the DNA molecules, ligating DNA fragments to adapters, digesting a plurality of DNA molecules, amplifying the DNA molecules, ligating adapters to DNA fragments, analyzing any type of DNA molecule, using a ligase, generating a mixture of DNA molecules, generating a mixture of ligated molecules, generating a mixture of adapter-ligated molecules or fragments, enriching a population of particular DNA molecules (including, by way of example, molecules that are not adapter dimers), performing one or more specific steps, identifying particular DNA molecules, distinguishing between methylated and unmethylated bases, subjecting particular DNA molecules to bisulfite conversion, enzymatic reactions, and / or chemical reactions, etc.
[0095] III. Nucleic acid molecules for sequencing library preparation In some embodiments, the nucleic acid molecule from which the sequencing library is prepared is DNA, and in some embodiments, the DNA is cell-free DNA (cfDNA). cfDNA can be obtained from a subject or individual, including a mammal. cfDNA can be derived from a subject or individual who needs to analyze cfDNA to determine a health condition, such as detecting a disease state or the risk or susceptibility thereof. cfDNA can be obtained from or derived from one or more samples from an individual. Samples can be obtained from or derived from plasma, serum, bone marrow, cerebrospinal fluid, pleural fluid, saliva, stool, or urine. The cfDNA from which the library is prepared can be double-stranded, single-stranded (preceding steps in the method can include polymerization of the second strand), or a mixture thereof.
[0096] In some embodiments, the nucleic acid molecule that is desired to prepare library can be modified before being used in the method of the present disclosure.For example, nucleic acid molecule can be enriched for specific type of nucleic acid molecule, specific size of nucleic acid molecule, or combination thereof.In some embodiments, nucleic acid molecule is, for example, the cfDNA that is enriched for the molecule of specific size and / or for the molecule that has one or more specific characteristics, for example, that comprises one or more methylation sites.
[0097] IV. Applications of Sequencing Libraries The present disclosure provides methods, systems, and compositions related to the preparation of molecules for any type of analysis, including for sequencing, for measuring the quality or quantity of methylation, etc. In some embodiments, the molecules comprise cfDNA, and in some embodiments, the cfDNA is obtained from or derived from an individual (for example, blood or plasma or urine (or a combination thereof) sample from a subject or individual). In some embodiments, after library preparation, the present disclosure provides methods and systems for evaluating DNA methylation in cfDNA molecules, for example, in the CpG-rich regions of cfDNA molecules.
[0098] The present disclosure relates to various aspects of methods for providing methylation information about cfDNA. Some embodiments include methods for assessing DNA methylation in CpG-rich regions of cfDNA.
[0099] In the embodiment of the present disclosure that is related to diseases such as cancer, detecting and characterizing cfDNA in suitable samples can be an effective method for obtaining information.For example, after preparing library, the sequencing library that is prepared can be used to determine whether an individual has, or is at risk of, or is susceptible to certain diseases or medical conditions.In one example, an individual has, or is suspected to have, or is at risk of having cancer, and analyzing the library of prepared cfDNA molecules can help determine whether the individual has, or is suspected to have, or is at risk of having cancer.
[0100] In some embodiments, post-library preparation methods include non-invasive cancer screening, including identifying the tissue of origin of the tumor. Liquid biopsies (also called liquid biopsies or liquid-phase biopsies), such as blood draws, are useful for identifying a variety of different malignancies, unlike traditional tissue biopsies, and can be utilized in the methods of the present disclosure.
[0101] In some embodiments, at least a subset of the plurality of DNA fragments has methylated nucleobases. In some embodiments, the starting cfDNA molecule may have 0, 1, or multiple CpG sites, and the method includes identifying cell-free DNA molecules with 2, 3, 4, or more CpG sites. In some embodiments, the method further includes placing the cfDNA molecule or its derivative (adapter-linked DNA fragment, or amplified adapter-linked DNA fragment) under sufficient conditions to distinguish methylated nucleobases from unmethylated nucleobases in the molecule. In some embodiments, placing the plurality of cfDNA molecules, DNA fragments, or derivatives thereof under sufficient conditions includes performing bisulfite conversion on the plurality of DNA fragments. In some embodiments, placing the plurality of cfDNA molecules, DNA fragments, or derivatives thereof under sufficient conditions includes performing an enzymatic and / or chemical reaction to oxidize methylated cytosine nucleobases and / or hydroxymethylated cytosine nucleobases, followed by reduction and / or deamination of the oxidation reaction product.
[0102] In some embodiments, the method further comprises measuring the methylation status of at least one portion of a plurality of DNA fragments or at least one portion of a plurality of adaptor-ligated DNA fragments to provide a methylation profile of at least one portion of a plurality of DNA fragments. In some embodiments, the method further comprises measuring the methylation status of at least some of the adaptor-ligated DNA fragments or amplified adaptor-ligated DNA fragments to provide a methylation profile of the cfDNA. In some embodiments, the method further comprises processing the methylation profile against one or more references. The methylation profile can include information about any number of CpG sites, CpG-rich sequences, and / or CpG islands (including the presence and / or absence of specific methylation sites). In some embodiments, the reference comprises a reference methylation profile of cfDNA molecules from one or more additional subjects. The subjects from which the reference methylation profile of cfDNA is obtained can be, for example, healthy, cancer-free, have cancer, or be at increased risk of having cancer.
[0103] In some embodiments, the plurality of cfDNA molecules are obtained from the subject's body sample.In some embodiments, the body sample is selected from the group consisting of plasma, serum, bone marrow, cerebrospinal fluid, pleural fluid, saliva, stool, sputum, nipple aspirate, biopsy, oral scraping, urine, and combinations thereof.In some embodiments, the method further comprises processing molecules having one or more CpG sites to create a methylation profile of the plurality of cfDNA molecules.In some embodiments, the method further comprises analyzing and processing the methylation profile to determine the possibility that the subject has disease or disorder, or the possibility that the subject is suspected to have disease or disorder.When the methylation profile obtained from an individual's sample is compared with one or more references, the source of the one or more reference samples may or may not be the same source as the individual's sample.
[0104] In some embodiments, the disease or disorder for which information is desired is selected from the group consisting of cancer, multiple sclerosis, traumatic or ischemic brain injury, diabetes, pancreatitis, Alzheimer's disease, and fetal abnormalities. In some embodiments, the disease or disorder is a cancer selected from the group consisting of pancreatic cancer, liver cancer, lung cancer, colorectal cancer, leukemia, bladder cancer, bone cancer, brain cancer, breast cancer, cervical cancer, endometrial cancer, esophageal cancer, gastric cancer, head and neck cancer, melanoma, ovarian cancer, testicular cancer, kidney cancer, sarcoma, bile duct cancer, thyroid cancer, gallbladder cancer, spleen cancer, and prostate cancer.
[0105] In some embodiments, the methylation pattern of cfDNA molecules obtained from the subject's body sample can be used to monitor aberrant tissue-specific cell death or organ transplantation.
[0106] In some embodiments, the library made by using the method or system provided herein for enriching CpG-rich regions or CpG islands in cfDNA is used for some purposes.In some embodiments, the library is examined for one or more characteristics.The library can be examined to determine the amount and / or location of methylation sites in some or all of the molecules of the library.In some embodiments, the methylation pattern is determined for at least some of some or all of the molecules of the library, including one or more specific sites.Methylation profiling can be carried out for at least some of some or all of the molecules of the library.
[0107] In some embodiments, the one or more methylation sites or markers may include plasma methylation biomarkers for various specific diseases or disorders, including cancer. Differentially methylated biomarkers can be identified by comparing methylation profile data obtained from patients with certain disease or disorder characteristics (cancer type, stage, prognosis, treatment response, etc.) with methylation profile data obtained from healthy controls. By identifying various methylation profiles specific to various cancers or tissue types, embodiments disclosed herein can detect many types of cancer and provide tumor location information for further detailed clinical investigation based on simple non-invasive liquid biopsy. Methylation profiles can be used to detect any disease or disorder, for example, based on non-invasive liquid biopsy.
[0108] In some embodiments, cfDNA methylation profile can be used to diagnose subject or patient, at least in part, based on whether the subject has a cfDNA methylation profile that suggests a disease or disorder.In some aspects, the present disclosure provides a method for diagnosing subject based on cfDNA methylation profile, comprising determining whether the patient has cancer, and making a cfDNA methylation profile that suggests cancer.In some embodiments, cfDNA methylation profile is made by processing the biological sample from patient that comprises cell-free DNA using the methods, compositions and systems included herein.
[0109] In some embodiments, cfDNA methylation profiles can be used to diagnose patients with cancer symptoms, patients without cancer symptoms, patients with a family history or personal history of cancer, patients at risk of cancer, or patients diagnosed with cancer. Patients can be mammalian, but in most embodiments, patients are human. Cancer can be malignant, benign, metastatic, or precancerous. In yet another embodiment, cancer is melanoma, non-small cell lung cancer, small cell lung cancer, lung cancer, hepatocellular carcinoma, retinoblastoma, astrocytoma, glioblastoma, gum cancer, tongue cancer, leukemia, neuroblastoma, head cancer, neck cancer, breast cancer, pancreatic cancer, prostate cancer, kidney cancer, bone cancer, testicular cancer, ovarian cancer, liver cancer, mesothelioma, cervical cancer, gastrointestinal cancer, lymphoma, brain cancer, colon cancer, sarcoma, gallbladder thyroid cancer, spleen cancer, or bladder cancer. Cancer can include tumors composed of tumor cells.
[0110] In some aspects, the present disclosure provides a method for treating cancer in cancer patients after determining the need for treatment, based on the method and system herein for concentrating DNA containing CpG islands or CpG-rich DNA for cancer diagnosis.Such a treatment method can include, after determining that a patient has cancer based on the method disclosed herein, administering to the patient an effective amount of chemotherapy, radiotherapy, hormone therapy, targeted therapy, or immunotherapy (or a combination thereof).Can identify the location of cancer occurrence, and in this case, treatment is tailored to the cancer that occurs there.In some embodiments, tumor resection can be performed as a treatment, or can be part of a treatment that is combined with one of other treatments.Examples of chemotherapeutic agents include alkylating agents such as bifunctional alkylating agents (e.g., cyclophosphamide, mechlorethamine, chlorambucil, melphalan) or monofunctional alkylating agents (e.g., dacarbazine (DTIC), nitrosoureas, temozolomide (oral dacarbazine)); anthracyclines (e.g., daunorubicin, doxorubicin, epirubicin, idarubicin, mitoxantrone, and valrubicin); cytoskeleton-interfering taxanes (e.g., paclitaxel, docetaxel, abraxane, taxotere); epothilones; histone dehydrogenase inhibitors (H2O3, H2O4, H2O5, H3O6, H4O7, H5O8, H5O9, H5O10, H5O11, H5O12, H5O13, H5O14, H5O15, H5O26, H5O16, H5O17, H5O18, H5O19, H5O21, H5O22, H5O23, H5O24, H5O25, H5O36, H5O47, H5O16, H5O18, H5O19, H5O25, H5O16, H5O18, H5O19, H5O15, H5O16 ... topoisomerase I inhibitors (e.g., irinotecan, topotecan); topoisomerase II inhibitors (e.g., etoposide, teniposide, tafluposide); kinase inhibitors (e.g., bortezomib, erlotinib, gefitinib, imatinib, vemurafenib, and vismodegib); nucleotide analogs and nucleotide precursor analogs (e.g., azacitidine, azathioprine, capecitabine, cytarabine, doxifluridine, fluorouracil, gemcitabine, hydrochloride, peptide antibiotics (e.g., bleomycin, actinomycin); platinum-based antineoplastic agents (e.g., carboplatin, cisplatin, oxaliplatin); retinoids (e.g., retinoin, alitretinoin, bexarotene); and vinca alkaloids (e.g., vinblastine, vincristine, vindesine, and vinorelbine). Examples of immunotherapeutics include, but are not limited to, cell therapies such as dendritic cell therapies (e.g., using chimeric antigen receptors); antibody therapies (e.g., alemtuzumab, atezolizumab, ipilimumab, nivolumab, ofatumumab, pembrolizumab, rituximab, or other antibodies with the same target as one of these antibodies, e.g., CTLA-4, PD-1, PD-L1, or other checkpoint inhibitors); and cytokine therapies (e.g., interferons or interleukins).
[0111] In some embodiments, the method of using cfDNA methylation profiling to diagnose an object can further comprise, before or after measuring the methylation profile of a patient, performing a biopsy, performing a CAT scan, performing a mammogram, performing an ultrasound diagnosis, or otherwise evaluating the suspected cancerous tissue.In some embodiments, the cancer that is found is classified into cancer classification or stage classification (for example, stage I, stage II, stage III, or stage IV).
[0112] In some embodiments, the cfDNA methylation profile obtained by the method and system for enriching CpG islands in cfDNA is used to monitor therapy and / or monitor tumor progression, including during and / or after treatment.For example, blood can be taken at various times during one or more treatment plans to monitor tumor progression, and the cfDNA derived therefrom can be analyzed.
[0113] In some embodiments, the cfDNA methylation profiles obtained by the disclosed methods and systems can be utilized for disease staging or as prognostic biomarkers, for example, when tissue biopsy is not possible or archived tumor samples are not available for genetic analysis.
[0114] In some embodiments, the cfDNA methylation profile obtained by the method and system for enriching CpG-rich regions in cfDNA provided herein can be used for screening and early detection of cancer.For example, blood sampling can be carried out regularly in individuals who do not have any symptoms of cancer, in order to detect cancer early or to identify the predisposition to cancer.
[0115] In some embodiments, the cfDNA methylation profiles obtained by the methods and systems for enriching CpG-rich regions in cfDNA provided herein can be used for prenatal testing of fetal DNA obtained from maternal plasma or serum to identify Down syndrome and other chromosomal abnormalities in the fetus.
[0116] In some embodiments, the cfDNA methylation profiles obtained by the methods and systems for enriching CpG-rich regions in cfDNA provided herein can be used for organ transplant monitoring.
[0117] In some embodiments, the cfDNA methylation profile obtained by the methods and systems for enriching CpG-rich regions in cfDNA provided herein can be used for the diagnosis of other types of diseases, such as multiple sclerosis, traumatic / ischemic brain injury, diabetes, pancreatitis, or Alzheimer's disease, or infectious diseases.
[0118] It is contemplated that any aspect discussed herein can be implemented with respect to any method, system, kit, computer-readable medium, or device of the invention, and vice versa. Additionally, devices used in this disclosure can be used to practice the methods of the disclosure.
[0119] In some embodiments, the one or more CpG sites comprise two or more, three or more, or four or more CpG sites. In some embodiments, the method further comprises generating a report, for example, electronically outputting a report showing the methylation profile. In some embodiments, the method further comprises analyzing and processing the methylation profile to determine the likelihood or risk that the subject has at least one disease or disorder, or the likelihood or risk that the subject is suspected of having at least one disease or disorder. In some embodiments, the disease or disorder is selected from the group consisting of cancer, multiple sclerosis, traumatic or ischemic brain injury, diabetes, pancreatitis, Alzheimer's disease, and fetal abnormalities. In some embodiments, the disease or disorder is a cancer selected from the group consisting of pancreatic cancer, liver cancer, lung cancer, colorectal cancer, leukemia, bladder cancer, bone cancer, brain cancer, breast cancer, cervical cancer, endometrial cancer, esophageal cancer, gastric cancer, head and neck cancer, melanoma, ovarian cancer, testicular cancer, kidney cancer, sarcoma, bile duct cancer, thyroid cancer, spleen cancer, gallbladder cancer, and prostate cancer.
[0120] In some embodiments, the one or more CpG sites comprise two or more CpG sites. In some embodiments, one or more computer processors are individually or collectively programmed to electronically output a report showing the methylation profile. In some embodiments, one or more computer processors are individually or collectively programmed to analyze and process the methylation profile to determine the likelihood or risk that the subject has one or more diseases or disorders, or the likelihood or risk that the subject is suspected of having one or more diseases or disorders. In some embodiments, the disease or disorder is selected from the group consisting of cancer, multiple sclerosis, traumatic or ischemic brain injury, diabetes, pancreatitis, Alzheimer's disease, and fetal abnormality. In some embodiments, the disease or disorder is a cancer selected from the group consisting of pancreatic cancer, liver cancer, lung cancer, colorectal cancer, leukemia, bladder cancer, bone cancer, brain cancer, breast cancer, cervical cancer, endometrial cancer, esophageal cancer, gastric cancer, head and neck cancer, melanoma, ovarian cancer, testicular cancer, kidney cancer, sarcoma, bile duct cancer, thyroid cancer, spleen cancer, gallbladder cancer, and prostate cancer.
[0121] In another aspect, the present disclosure provides a non-transitory computer-readable medium comprising machine-executable code, the machine-executable code, when executed by one or more computer processors, performing a method for analyzing or analysing a plurality of cfDNA molecules that have been subjected to the library preparation method of the present disclosure, the method comprising: (a) collecting a plurality of sequence reads produced by a sequencer, wherein at least a subset of the plurality of sequence reads comprises individual sequence reads that include (i) sequences derived from the plurality of cfDNA molecules and (ii) adapter sequences not derived from the plurality of cell-free DNA molecules, located at both ends of each of the individual sequence reads; (b) analyzing the plurality of sequence reads to (i) identify one or more sequence reads derived from the plurality of sequence reads that have the adapter sequences at both ends and (ii) confirm that the one or more sequence reads are associated with one or more CpG sites of the plurality of cell-free DNA molecules; and (c) using the one or more CpG sites identified in (b) to generate a methylation profile of the plurality of cell-free DNA molecules.
[0122] The library prepared from the cfDNA sample from subject can be subjected to any kind of analysis, including methylation profiling, for example, for screening, diagnosis, prognosis, treatment selection or treatment monitoring for tumor or non-solid cancer.For example, analysis can suggest that the patient with a certain methylation profile can best respond to surgery, chemotherapy, radiotherapy, targeted therapy, hormone therapy, immunotherapy or their combination.By accurately profiling the methylation of cfDNA sample, it can prevent the treatment that may be ineffective from being prescribed and administered to the patient.
[0123] Methylation profiling of V sequencing libraries After preparing a molecular library using the methods encompassed herein, methylation profiling may be performed on the enriched DNA molecules. For example, sequencing reads may be generated from the enriched DNA molecules using any suitable sequencing method. The sequencing method may be a first-generation sequencing method such as Maxam-Gilbert sequencing or Sanger sequencing, or a high-throughput sequencing method (e.g., next-generation sequencing or NGS). A high-throughput sequencing method may simultaneously (or substantially simultaneously) sequence at least 10,000, 100,000, 1 million, 10 million, 100 million, 1 billion, or more polynucleotide molecules. Sequencing methods may include, but are not limited to, pyrosequencing, sequencing by synthesis, single molecule sequencing, nanopore sequencing, semiconductor sequencing, sequencing by ligation, sequencing by hybridization, digital gene expression (Helicos), massively parallel sequencing, such as sequencing using Helicos, clonal single molecule arrays (Solexa / Illumina), PacBio, SOLiD, Ion Torrent, or Nanopore platforms, BGISEQ, or combinations thereof.
[0124] In some embodiments, sequencing comprises whole genome sequencing (WGS). In some embodiments, sequencing comprises whole genome bisulfite sequencing (WGBS), for example, for reference DNA samples. In some embodiments, sequencing comprises reduced representation bisulfite sequencing (RRBS), for example, for reference DNA samples. In some embodiments, sequencing comprises targeted sequencing using a panel that comprises multiple loci. Sequencing can be performed to a sufficient depth to perform methylation profiling in subjects with desired performance (for example, accuracy, sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), or area under the curve (AUC) of receiver operating characteristic (ROC)). In some embodiments, sequencing is performed to a depth of at least about 5x, at least about 10x, at least about 20x, at least about 50x, at least about 75x, at least about 100x, at least about 125x, at least about 150x, at least about 175x, or at least about 200x, or any range derivable therein.
[0125] In some embodiments, the multiple loci may correspond to coding and / or non-coding genomic regions of a genome, such as CpG islands, hypermethylated and / or hypomethylated regions, and / or regions adjacent to such hypermethylated and / or hypomethylated regions. These genomic regions may correspond to cancer-related (or tumor-related) coding and / or non-coding genomic regions of a genome, such as cancer driver mutations or genetic variants. Genetic variants may include, for example, single nucleotide variants (SNVs), copy number variants (CNVs), insertions or deletions (indels), fusion genes, hypermethylation, and hypomethylation.
[0126] In some embodiments, performing methylation profiling of a subject may include aligning cfDNA sequencing reads to a reference genome. The reference genome may include at least one portion of a genome (e.g., a human genome). The reference genome may include the entire genome (e.g., the entire human genome). In some embodiments, the reference genome may include multiple loci corresponding to coding and / or non-coding genomic regions of a genome, such as CpG-rich regions, CpG islands, hypermethylated and / or hypomethylated regions, and / or regions adjacent to such hypermethylated and / or hypomethylated regions. These multiple genomic regions may correspond to cancer-related (or tumor-related) coding and / or non-coding genomic regions of a genome, such as cancer driver mutations or gene variants. Genetic variants may include, for example, single nucleotide variants (SNVs), copy number variants (CNVs), insertions or deletions (indels), fusion genes, hypermethylation, and hypomethylation. Alignment can be performed using, for example, the Burrows-Wheeler algorithm or other alignment algorithms (e.g., suitable for bisulfite-converted reads).
[0127] In some embodiments, carrying out methylation profiling in subject can comprise making the quantitative index of cfDNA sequencing reads for each of a plurality of loci.The quantitative index of cfDNA sequencing reads can be made, for example, the total number of DNA sequencing reads that are aligned with given loci (for example, CpG-rich region, CpG island, hypermethylated region, hypomethylated region, the region adjacent to hypermethylated region, or the region adjacent to hypomethylated region).For example, the cfDNA sequencing reads that have some or all of the sequencing reads that are aligned with given CpG-rich region or CpG island can be calculated for the quantitative index of this CpG-rich region or CpG island.
[0128] The combination of specific CpG-rich region and / or CpG island and non-specific CpG-rich region and / or CpG island patterns can form the methylation profile of a subject.The time-dependent changes in these patterns of CpG-rich region and / or CpG island can indicate the change in the methylation profile of a subject.This change can include the absence of methylation at one or more specific CpG sites, the increase in the methylation level at specific CpG-rich site or CpG island, the decrease in the methylation level at specific CpG-rich site or CpG island, etc.
[0129] In some embodiments, binding measurement can be carried out for methylation profiling, and binding measurement can include using a probe that is selective for multiple CpG-rich regions and / or CpG islands in multiple enriched cfDNA fragments to assay enriched cfDNA fragments.In some embodiments, the probe is the nucleic acid molecule that has sequence complementarity with the nucleic acid sequence of CpG-rich regions and / or CpG islands.In some embodiments, the nucleic acid molecule is primer or enrichment sequence.In some embodiments, said assaying comprises using array hybridization or polymerase chain reaction (PCR), or nucleic acid sequencing.
[0130] In some embodiments, the library can be enriched for at least a portion of a plurality of loci.In some embodiments, the enrichment can comprise amplifying a plurality of library molecules.For example, a plurality of cfDNA molecules can be amplified by selective amplification (for example, by using a set of primers or probes that comprise nucleic acid molecules that have sequence complementarity with the nucleic acid sequence of CpG islands).Alternatively, or in combination, a plurality of cfDNA molecules can be amplified by universal amplification (for example, by using universal primers).In some embodiments, the enrichment comprises selectively isolating at least a portion of a plurality of cfDNA molecules.
[0131] In some embodiments, performing methylation profiling on a subject comprises analyzing and processing sequence reads from a library to obtain a quantitative indicator of deviation.In some embodiments, the quantitative indicator of deviation is a z-score based on one or more reference cfDNA samples.The reference cfDNA sample can be obtained from a subject with a specific methylation profile and / or from a subject without a specific methylation profile.The reference cfDNA sample can be obtained from a subject with or without a certain type of cancer (for example, pancreatic cancer, liver cancer, lung cancer, colorectal cancer, leukemia, bladder cancer, bone cancer, brain cancer, breast cancer, cervical cancer, endometrial cancer, esophageal cancer, gastric cancer, head and neck cancer, melanoma, ovarian cancer, testicular cancer, kidney cancer, sarcoma, bile duct cancer, thyroid cancer, gallbladder cancer, spleen cancer, and prostate cancer). The reference cfDNA sample may be obtained from a subject suffering from a specific stage of cancer or a subject not suffering from a specific stage of cancer (including stage I, stage II, stage III, or stage IV).The reference cfDNA sample may be obtained from a subject experiencing abnormal tissue-specific cell death.
[0132] In some embodiments, performing methylation profiling on subject comprises measuring the deviation of the cfDNA methylation profile of subject when the quantitative indicator of deviation meets a predetermined criterion.In some embodiments, the predetermined criterion is that the z-score (or the quantitative indicator calculated from multiple z-scores) of subject's methylation profile is greater than or less than a predetermined value.The predetermined value can be about 0.1, about 0.2, about 0.5, about 1, about 1.5, about 2, about 2.5, about 3, about 3.5, about 4, about 4.5, about 5 or a value greater than about 5.
[0133] In some embodiments, the prepared sequencing library is analyzed for one or more specific loci.In some embodiments, the plurality of loci comprises CpG-rich regions, CpG islands, hypermethylated regions and / or hypomethylated regions, and / or regions adjacent to such hypermethylated regions and / or hypomethylated regions.The plurality of loci comprises at least about 10 different loci, at least about 20 different loci, at least about 30 different loci, at least about 40 different loci, at least about 50 different loci, at least about 75 different loci, at least about 100 different loci, at least about 500 different loci, at least about 1000 different loci, at least about 5000 different loci, at least about 100,000 different loci, at least about 5 ... The number of different loci may include at least about 1 million different loci, at least about 1 million different loci, at least about 2 million different loci, at least about 3 million different loci, at least about 4 million different loci, at least about 5 million different loci, at least about 10 million different loci, at least about 25 million different loci, at least about 50 million different loci, at least about 75 million different loci, at least about 100 million different loci, or more than 100 million different loci, or any range derivable therein. The positions of the different loci may or may not be in the same gene, on the same chromosome, or on different chromosomes.
[0134] In some embodiments, measuring the subject's deviated cfDNA methylation profile is performed with a sensitivity of at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%, or any range derivable therein.
[0135] In some embodiments, measuring the subject's deviated cfDNA methylation profile is performed with a specificity of at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%, or any range derivable therein.
[0136] In some embodiments, measuring the subject's deviated cfDNA methylation profile is performed with a positive predictive value (PPV) of at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%, or any range derivable therein.
[0137] In some embodiments, measuring the subject's deviated cfDNA methylation profile is performed with a negative predictive value (NPV) of at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%, or any range derivable therein.
[0138] In some embodiments, measuring the subject's deviated cfDNA methylation profile is performed with an area under the curve (AUC) of the receiver operating characteristic (ROC) of at least about 0.5, at least about 0.6, at least about 0.7, at least about 0.75, at least about 0.8, at least about 0.85, at least about 0.9, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99, or any range derivable therein.
[0139] In some embodiments, performing methylation profiling on a subject comprises measuring the normal cfDNA methylation profile of the subject when the quantitative indicator of deviation meets a predetermined criterion.In some embodiments, the predetermined criterion is that the z-score (or the quantitative indicator calculated from multiple z-scores) of the subject's methylation profile is greater than or less than a predetermined value.The predetermined value can be about 0.1, about 0.2, about 0.5, about 1, about 1.5, about 2, about 2.5, about 3, about 3.5, about 4, about 4.5, about 5, or greater than about 5, or any range that can be derived therein.
[0140] In some embodiments, measuring the subject's normal cfDNA methylation profile is performed with a sensitivity of at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%, or any range derivable therein.
[0141] In some embodiments, measuring the subject's normal cfDNA methylation profile is performed with a specificity of at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%, or any range derivable therein.
[0142] In some embodiments, measuring the subject's normal cfDNA methylation profile is performed with a positive predictive value (PPV) of at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%, or any range derivable therein.
[0143] In some embodiments, measuring the subject's normal cfDNA methylation profile is performed with a negative predictive value (NPV) of at least about 10%, at least about 20%, at least about 30%, at least about 40%, at least about 50%, at least about 60%, at least about 70%, at least about 80%, at least about 90%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, or at least about 99%, or any range derivable therein.
[0144] In some embodiments, measuring the subject's normal cfDNA methylation profile is performed with an area under the curve (AUC) of the receiver operating characteristic (ROC) of at least about 0.5, at least about 0.6, at least about 0.7, at least about 0.75, at least about 0.8, at least about 0.85, at least about 0.9, at least about 0.95, at least about 0.96, at least about 0.97, at least about 0.98, or at least about 0.99, or any range derivable therein.
[0145] In some embodiments, the subject has been diagnosed with cancer, or is suspected to be affected by cancer, or is at risk of being affected by cancer.For example, the cancer can be one or more types, including brain cancer, breast cancer, cervical cancer, colorectal cancer, endometrial cancer, esophageal cancer, gastric cancer, hepatobiliary cancer, leukemia, liver cancer, lung cancer, lymphoma, ovarian cancer, pancreatic cancer, skin cancer, testicular cancer, kidney cancer, sarcoma, bile duct cancer, prostate cancer, thyroid cancer, gallbladder cancer, spleen cancer or urinary tract cancer.
[0146] In some embodiments, based on the cfDNA methylation profile obtained from the subject (e.g., by measuring a deviated cfDNA methylation profile or a normal cfDNA methylation profile), the method of the present disclosure comprises administering a therapeutically effective amount of one or more therapies to treat the subject's disease or disorder (e.g., cancer). In some embodiments, the treatment comprises chemotherapy, radiation therapy, targeted therapy, immunotherapy, or a combination thereof. Based on the obtained methylation profile of the subject, the existing treatment for the subject may be discontinued, and another treatment may be administered to the subject. Alternatively, based on the obtained methylation profile of the subject, the existing treatment for the subject may be continued and / or another treatment may be administered to the subject. Some individuals may be deemed unresponsive to one or more therapies based on the results of the methylation profile, and as a result, the treatment may never be administered, or may be administered but discontinued, or discontinued after a certain number of times and / or a certain period of time, based on the results of subsequent methylation profiles of the same individual.
[0147] The obtained cfDNA methylation profile of the subject can be evaluated to determine the diagnosis of cancer, the prognosis of cancer, or the sign of tumor progression or regression.In addition, based on cfDNA methylation profile evaluation or monitoring (for example, the difference between the cfDNA methylation profile at two or more time points), one or more clinical outcomes can be assigned.Such clinical outcomes can include one or more of the following: diagnosing the subject suffering from cancer, comprising one or more types of tumors; diagnosing the subject suffering from cancer, comprising one or more types and / or stages of tumors; determining the prognosis of the subject suffering from cancer (for example, indicating, instructing, or administering a clinical treatment course (for example, surgery, chemotherapy, radiotherapy, hormone therapy, targeted therapy, immunotherapy, or other treatment) to the subject; indicating, instructing, or administering a different clinical course of action to the subject (for example, not treating, for example, continuing to monitor at a specified time interval, stopping current treatment, switching to another treatment); or indicating the expected survival time to the subject.
[0148] In some embodiments, measuring the cfDNA methylation profile of a subject comprises measuring one or more predetermined thresholds for one or more loci (for example, a plurality of CpG-rich regions and / or CpG islands).The predetermined threshold (for example, for each of a plurality of CpG-rich regions and / or CpG islands) can be determined by performing cfDNA methylation profiling on one or more samples from one or more control subjects (for example, patients known to have or not have a specific disease or disorder, patients known to have or not have a specific tumor type, patients known to have or not have a specific tumor type at a specific stage, or healthy subjects who have not been diagnosed with a disease or disorder or do not show any clinical symptoms of a disease or disorder), and determining an appropriate predetermined threshold based on the cfDNA methylation profiling of the control samples.
[0149] The predetermined threshold may be adjusted based on the desired sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), or accuracy of measuring a subject's deviated cfDNA methylation profile or measuring a normal cfDNA methylation profile. For example, if high sensitivity is desired for measuring a subject's deviated cfDNA methylation profile status, the predetermined threshold may be adjusted to a lower value. Alternatively, if high specificity is desired for measuring a subject's deviated cfDNA methylation profile, the predetermined threshold may be adjusted to a higher value. The predetermined threshold may be adjusted to maximize the area under the curve (AUC) of the receiver operating characteristic (ROC) of the control sample obtained from the control subject. The predetermined threshold may be adjusted to achieve a desired balance between false positives (FP) and false negatives (FN) when measuring a subject's deviated cfDNA methylation profile.
[0150] In some embodiments, measuring the cfDNA methylation profile of the subject further comprises repeating cfDNA methylation profiling at a later second time point.The second time point can be selected to appropriately compare the cfDNA methylation profile with that of the first time point.Examples of the second time point can correspond to the time after surgical resection, the time during or after administering treatment to treat disease or disorder (for example, cancer) in the subject to monitor the efficacy of treatment, or the time after disease or disorder (for example, cancer) becomes undetectable in the subject after treatment, for example, to monitor residual disease or cancer recurrence in the subject.
[0151] In some embodiments, measuring the cfDNA methylation profile of the subject further comprises determining the difference between the first cfDNA methylation profile and the second cfDNA methylation profile, which indicates the progression or regression of the subject's tumor.Alternatively or in combination, the method can further comprise using a computer processor to generate a plot of the first cfDNA methylation profile and the second cfDNA methylation profile as a function of the first time point and the second time point.This plot can indicate the progression or regression of the subject's tumor.For example, the computer processor can generate a plot in which the time corresponding to the collection time of the data corresponding to two or more cfDNA methylation profiles is plotted on the x-axis, and two or more cfDNA methylation profiles are plotted on the y-axis.
[0152] The plot illustrating the difference between the first cfDNA methylation profile and the second cfDNA methylation profile can show the progression or regression of the tumor of the subject.For example, if the deviation of the second cfDNA methylation profile is greater than the deviation of the first cfDNA methylation profile, this difference can indicate, for example, tumor progression, the ineffectiveness of a certain treatment for the tumor of the subject, the resistance of the tumor to ongoing treatment, the metastasis of the tumor to other parts of the subject, or the residual disease or cancer recurrence in the subject.Alternatively, if the deviation of the second cfDNA methylation profile is smaller than the deviation of the first cfDNA methylation profile, this difference can indicate, for example, tumor regression, the surgical resection of the tumor of the subject is effective, the effectiveness of a certain treatment for the disease or disorder (for example, cancer) of the subject, or the absence of residual disease or cancer recurrence in the subject.
[0153] After assessing and / or monitoring the cfDNA methylation profile, one or more clinical outcomes can be assigned based on the cfDNA methylation profile assessment or monitoring (e.g., the difference between the cfDNA methylation profiles at two or more time points). Such clinical outcomes can include one or more of the following: diagnosing a subject suffering from cancer, including one or more types of tumors; diagnosing a subject suffering from cancer, including one or more types and / or stages of tumors; determining the prognosis of a subject suffering from cancer (e.g., suggesting, prescribing, or administering a clinical treatment course (e.g., surgery, chemotherapy, radiation therapy, targeted therapy, immunotherapy, or other treatment) to the subject; suggesting, prescribing, or administering an alternative clinical course of action to the subject (e.g., no treatment, e.g., continued monitoring at specified time intervals, discontinuing the current treatment, switching to another treatment); or suggesting an expected survival time to the subject.
[0154] VI Kit Any of the compositions described herein can be included in kit.In a non-limiting example, kit can include cfDNA; one or more devices for collecting cfDNA; enzyme; adapter; primer; dNTP; buffer solution; and other chemicals, including ATP, DTT, sodium bisulfite, etc.
[0155] The components of the kit may be packaged either dissolved in aqueous media or in lyophilized form. The container means of the kit may include at least one vial, test tube, flask, bottle, or other container means into which the components may be placed, preferably suitably aliquoted. Where multiple components are present in the kit, the kit may also include second, third, or other additional containers into which the additional components may be placed separately. However, various combinations of components may be contained in a single vial. The kits of the present disclosure may also include means for containing the components in close confinement for commercial sale. Such containers may include blow-molded plastic containers into which the desired vials are retained.
[0156] The kit of the present disclosure can comprise the instruction manual for carrying out the method provided herein, for example, the method for digesting and concentrating cfDNA and the method for subjecting the concentrated cfDNA to further analysis (for example, PCR, nucleic acid array, next-generation sequencing).This instruction manual can be in physical form (for example, printed instruction manual) or electronic form.
[0157] The kit of the present disclosure can comprise a software package or a web link to a server or cloud computing platform, which is used to analyze the sequencing data generated from the sequencing library prepared by the kit.By analyzing, the information about the quality control of the kit can be obtained, such as digestion efficiency, bisulfite conversion efficiency, etc., and the methylation profile of enriched cfDNA can be obtained.
[0158] The kits of the present disclosure may include a report generated by a software package provided with the kit or by a server or cloud computing platform. The report may provide information for (1) diagnosis and / or prevention of a medical condition; (2) therapy for a medical condition; (3) therapy monitoring, etc. For example, the report may provide information about the presence or risk of cancer, including specific types of cancer.
[0159] Although the present disclosure and its advantages have been described in detail, it should be understood that various changes, substitutions, and alterations can be made therein without departing from the spirit and scope of the design as defined by the appended claims. Moreover, the scope of the present application is not limited to the particular embodiments of the processes, machines, manufacture, compositions of matter, means, methods, and steps described in the specification. As those skilled in the art will readily understand from this disclosure, any existing or later-developed processes, machines, manufacture, compositions of matter, means, methods, or steps that perform substantially the same function or achieve substantially the same result as the corresponding embodiments described herein can be utilized in accordance with the present disclosure. Accordingly, the appended claims are intended to include within their scope such processes, machines, manufacture, compositions of matter, means, methods, or steps.
Claims
1. 1. A method for preparing a sequencing library of nucleic acids, comprising the steps of: (a) digesting a plurality of DNA molecules with a first one or more restriction enzymes to generate DNA fragments; (b) ligating adapters to the DNA fragments by incubation with a ligase to produce a mixture comprising the adapter-ligated DNA fragments and adapter dimers; (c) digesting the mixture in (b) with a second restriction enzyme or enzymes that digest the junction between the end of one adapter and the end of another adapter but do not digest the junction between the end of the DNA fragment and the adapter; (d) subjecting the adaptor-ligated DNA fragments to conditions sufficient to allow for differentiation of methylated and unmethylated nucleobases in the adaptor-ligated DNA fragments by sequencing; and (e) amplifying the adaptor-ligated DNA fragments to produce amplified adaptor-ligated DNA fragments.
2. AcII, HindIII, MluCI, PciI, A geI, BspMI, BfuAI, SexAI, MluI, BceAI, HpyCH4IV, H pyCH4III, BaeI, BsaXI, AflIII, SpeI, BsrI, BmrI, BglII, BspDI, PI-SceI, NsiI, AseI, CspCI, MfeI, BssS α I、DraIII、EcoP15I、AlwNI、BtsIMutI、NdeI、CviAII、FatI、NlaIII、FspEI、X cmI, BstXI, PflMI, BccI, NcoI, BseYI, FauI, TspMI, XmaI, LpnPI, AclI, ClaI 、SacII、HpaII、MspI、ScrFI、StyD4I、BsaJI、BslI、BtgI、NciI、AvrII、MnlI、 BbvCI, SbfI, Bpu10I, Bsu36I, EcoNI, HpyAV, BstNI, PspGI, StyI, BcgI, PvuI 、EagI、RsrII、BsiEI、BsiWI、BsmBI、Hpy99I、AbaSI、MspJI、SgrAI、BfaI、Bsp CNI、XhoI、PaeR7I、EarI、AcuI、PstI、BpmI、DdeI、SfcI、AflII、BpuEI、SmlI、 AvaI, BsoBI, MboII, BbsI, BsmI, EcoRI, HgaI, AatII, PflFI, Tth111I, AhdI DrdI, SacI, BseRI, PleI, HinfI, Sau3AI, MboI, DpnII, TfiI, BsrDI, BbvI, Bts α I、BstAPI、SfaNI、SphI、NmeAIII、NgoMIV、BglI、AsiSI、BtgZI、HhaI、HinP1I、BssHII、NotI、Fnu4HI、MwoI BmtI, NheI, BspQI, BlpI, TseI, ApeKI, Bsp1286I, AlwI, BamHI, BtsCI, FokI, FseI, SfiI, NarI, PluTI I、AscI、EciI、BsmFI、ApaI、PspOMI、Sau96I、KpnI、Acc65I、BsaI、HphI、BstEII、AvaII、BanI、BaeHGI and BsaHGI I、BanII、CviQI、BciVI、SalI、BcoDI、BsmAI、ApaLI、BsgI、AccI、Tsp45I、BsiHKAI、TspRI、ApoI、NspIF、BsrIF α I、BstYI、HaeII、EcoO109I、PpuMI、I-CeuI、I-SceI、BspHI、BspEI、MmeI、Taq α I、Hpy188I、Hpy188III、XbaI、BclI、PI-PspI、BsrGI、MseI、PacI、BstB I、PspXI、BsaWI、EaeI、HpyF30I、Sfr274I、またはそれら
3. 3. The method of claim 1 or 2, further comprising carrying out (a) and (b) in the same reaction mixture.
4. 4. The method of claim 3, wherein (a) is carried out at a different temperature than (b).
5. 4. The method of claim 3, wherein (a) is carried out at the same temperature as (b).
6. 6. The method of any one of claims 1 to 5, wherein (e) comprises utilizing a primer during the amplifying step that can initiate polymerization at a junction between an adaptor and a DNA fragment but not at a junction between an adaptor and another adaptor.
7. 1. A method for preparing a library of nucleic acids, comprising: (a) digesting a plurality of DNA molecules with a first one or more restriction enzymes to generate DNA fragments; (b) ligating adapters to the DNA fragments by incubation with a ligase to produce a mixture of adapter-ligated DNA fragments and adapter dimers; (c) digesting the mixture of adaptor-ligated DNA fragments and adaptor dimers with a second restriction enzyme or enzymes that digest the junction between the end of one adaptor and the end of another adaptor but do not digest the junction between the end of the DNA fragment and the adaptor; (d) subjecting the adaptor-ligated DNA fragments to conditions sufficient to allow for differentiation of methylated and unmethylated nucleobases in the adaptor-ligated DNA fragments by sequencing; and (e) amplifying the adaptor-ligated DNA fragments to produce amplified adaptor-ligated DNA fragments. The method comprising:
8. 8. The method of claim 7, wherein said subjecting step comprises bisulfite conversion.
9. 8. The method of claim 7, wherein said subjecting step comprises one or more enzymatic and / or chemical reactions.
10. 10. The method of claim 9, further comprising oxidizing the methylated cytosine nucleobase and / or the hydroxymethylated cytosine nucleobase to produce an oxidation reaction product, followed by reducing and / or deaminating the oxidation reaction product.
11. 11. The method of claim 10, wherein the oxidizing step is carried out using a ten-eleven translocation (TET) enzyme.
12. 11. The method of claim 10, wherein the oxidizing step is carried out using potassium perruthenate.
13. 11. The method of claim 10, wherein the step of deaminating the oxidation reaction product is carried out using an apolipoprotein B mRNA editing enzyme catalytic polypeptide-like (APOBEC).
14. 11. The method of claim 10, wherein the step of reducing and / or deaminating the oxidation reaction product is carried out using pyridine borane.
15. 15. The method of any one of claims 9 to 14, further comprising a step of performing β-glucosyltransferase treatment prior to one or more enzymatic and / or chemical reactions.
16. 8. The method of claim 7, wherein the sequencing is next generation sequencing.
17. 17. The method of claim 16, further comprising performing targeted capture prior to next generation sequencing to further enrich for adaptor-ligated fragments.
18. 18. The method of claim 16 or 17, further comprising performing size selection prior to next generation sequencing to further enrich for adaptor-ligated fragments.
19. 19. The method of any one of claims 7 to 18, further comprising analyzing the amplified, adaptor-ligated DNA fragments to generate a methylation profile.
20. A method described in any one of claims 7 to 19, wherein the adapter comprises a GC (3' to 5') overhang.
21. 21. The method of any one of claims 7 to 20, wherein the first one or more restriction enzymes comprise MspI, HpaII, TaqαI, or functional analogs thereof, or mixtures thereof.
22. 22. The method of any one of claims 7 to 21, wherein the second one or more restriction enzymes comprise one or more of BspD1, XhoI, SmlI, HpyF30I, PaeR7I, Sfr274I, or functional analogs thereof, or mixtures thereof.
23. 23. The method of any one of claims 7 to 22, wherein the ligase is T7 DNA ligase, T4 DNA ligase, T3 DNA ligase, Taq DNA ligase, or a functional analogue thereof, or a mixture thereof.
24. 24. The method of any one of claims 7-23, wherein the plurality of DNA molecules comprises cell-free DNA (cfDNA).
25. 25. The method of claim 24, wherein the cfDNA is obtained from or derived from a sample derived from a subject or individual.
26. 26. The method of claim 25, wherein the sample is obtained from or derived from plasma, serum, bone marrow, cerebrospinal fluid, pleural fluid, saliva, stool, or urine.
27. 27. The method of any one of claims 7 to 26, wherein the adaptor comprises a known sequence.
28. 28. The method of any one of claims 7 to 27, wherein the adapter comprises a unique sequence.
29. 29. The method of any one of claims 7 to 28, wherein the nucleic acids are enriched for molecules having one or more CpG sites.
30. 1. A method for preparing a library of nucleic acids, comprising the steps of: (a) digesting a plurality of DNA molecules with a first one or more restriction enzymes to generate DNA fragments; (b) ligating adapters to the DNA fragments by incubation with a ligase to produce a mixture of adapter-ligated DNA fragments and adapter dimers; (c) digesting the mixture of adaptor-ligated DNA fragments and adaptor dimers with a second restriction enzyme or enzymes that digest the junction between the end of one adaptor and the end of another adaptor but do not digest the junction between the end of the DNA fragment and the adaptor; (d) subjecting the adaptor-ligated DNA fragments to conditions sufficient to allow for differentiation of methylated and unmethylated nucleobases in the adaptor-ligated DNA fragments by sequencing; and (e) amplifying the adaptor-ligated DNA fragments to produce amplified adaptor-ligated DNA fragments.
31. 31. The method of claim 30, wherein the first one or more restriction enzymes comprise one or more of MspI, HpaII, TaqαI, or functional analogs thereof, or mixtures thereof.
32. 32. The method of claim 30 or 31, wherein the second one or more restriction enzymes are one or more of BspDI, ClaI, AclI, NarI, XhoI, SmlI, HpyF30I, PaeR7I, Sfr274I, or functional analogs thereof, or mixtures thereof.
33. 33. The method of any one of claims 30-32, further comprising carrying out (a), (b), and (c) in the same reaction mixture.
34. 34. The method of any one of claims 30 to 33, wherein said subjecting step comprises subjecting the adaptor-ligated fragments to bisulfite conversion.
35. 34. The method of any one of claims 30 to 33, further comprising subjecting the adaptor-ligated fragments to one or more enzymatic and / or chemical reactions.
36. 36. The method of claim 35, further comprising oxidizing the methylated cytosine nucleobase and / or the hydroxymethylated cytosine nucleobase to produce an oxidation reaction product, followed by reducing and / or deaminating the oxidation reaction product.
37. 37. The method of claim 36, wherein the oxidizing step is carried out using a ten-eleven translocation (TET) enzyme.
38. 37. The method of claim 36, wherein the oxidizing step is carried out using potassium perruthenate.
39. 37. The method of claim 36, wherein the step of reducing and / or deaminating the oxidation reaction product is carried out using an APOBEC.
40. 37. The method of claim 36, wherein the step of reducing and / or deaminating the oxidation reaction product is carried out using pyridine borane.
41. 41. The method of any one of claims 35 to 40, further comprising a step of performing β-glucosyltransferase treatment prior to one or more enzymatic and / or chemical reactions.
42. 42. The method of any one of claims 30 to 41, wherein the adapter comprises a GC overhang.
43. 1. A method for preparing a library of nucleic acids, comprising the steps of: (a) digesting a plurality of DNA molecules with a first one or more restriction enzymes to generate DNA fragments; (b) ligating the DNA fragments and the first adapters, which are designed adapter dimers, by incubating with a ligase to create second adapters and also to generate a mixture of DNA fragments ligated to the second adapters and adapter dimers of the second adapters, and subjecting the designed adapter dimers to the influence of a second restriction enzyme or enzymes, wherein the second restriction enzyme or enzymes digest the junction between the end of one second adapter and the end of another second adapter, but do not digest the junction between the end of the DNA fragment and the second adapter; (c) subjecting the adaptor-ligated DNA fragments to conditions sufficient to allow for differentiation of methylated and unmethylated nucleobases in the adaptor-ligated DNA fragments by sequencing; and (d) amplifying the second adaptor-ligated DNA fragment to produce an amplified adaptor-ligated DNA fragment.
44. 44. The method of claim 43, wherein the first one or more restriction enzymes comprise one or more of MspI, HpaII, TaqαI, or functional analogs thereof, or mixtures thereof.
45. 45. The method of claim 43 or 44, wherein the second one or more restriction enzymes comprise one or more of BspDI, ClaI, AclI, NarI, XhoI, SmlI, HpyF30I, PaeR7I, Sfr274I, or functional analogs thereof, or mixtures thereof.
46. 46. The method of any one of claims 43-45, further comprising carrying out (a) and (b) in the same reaction mixture.
47. 47. The method of any one of claims 43 to 46, further comprising distinguishing between methylated and unmethylated nucleobases in the DNA fragments ligated to the second adaptor.
48. 48. The method of claim 47, further comprising subjecting the DNA fragment ligated to the second adaptor to bisulfite conversion.
49. 48. The method of claim 47, further comprising subjecting the DNA fragments ligated to the second adaptors to one or more enzymatic and / or chemical reactions.
50. 50. The method of claim 49, further comprising oxidizing the methylated cytosine nucleobase and / or the hydroxymethylated cytosine nucleobase to produce an oxidation reaction product, followed by reducing and / or deaminating the oxidation reaction product.
51. 51. The method of claim 50, wherein the oxidizing step is carried out using a ten-eleven translocation (TET) enzyme.
52. 51. The method of claim 50, wherein the oxidizing step is carried out using potassium perruthenate.
53. 51. The method of claim 50, wherein the step of reducing and / or deaminating the oxidation reaction product is carried out using an APOBEC.
54. 51. The method of claim 50, wherein the step of reducing and / or deaminating the oxidation reaction product is carried out using pyridine borane.
55. 55. The method of any one of claims 49 to 54, further comprising the step of performing β-glucosyltransferase treatment prior to one or more enzymatic or chemical reactions.
56. 56. The method of any one of claims 43-55, wherein the GC overhang is created by digestion of the adapter dimer of the second adapter with a second restriction enzyme or enzymes.
57. 1. A method for preparing a library of nucleic acids, comprising the steps of: (a) digesting a plurality of DNA molecules with a first one or more restriction enzymes to generate DNA fragments; (b) ligating adapters to the DNA fragments to produce a mixture of adapter-ligated DNA fragments and adapter dimers; (c) amplifying the adaptor-ligated DNA fragments to produce a mixture of amplified adaptor-ligated DNA fragments and amplified adaptor-dimers; (d) digesting the mixture of amplified, adaptor-ligated DNA fragments and amplified adaptor-dimers with a second restriction enzyme or enzymes that digest the junction between the end of one adaptor and the end of another adaptor but do not digest the junction between the end of the DNA fragment and the adaptor; and (e) subjecting the adaptor-ligated DNA fragments to conditions sufficient to allow differentiation of methylated and unmethylated nucleobases in the adaptor-ligated DNA fragments by sequencing.
58. 58. The method of claim 57, wherein the first one or more restriction enzymes comprise one or more of MspI, HpaII, TaqαI, or functional analogs thereof, or mixtures thereof.
59. 59. The method of claim 57 or 58, wherein the second one or more restriction enzymes comprise one or more of BspDI, ClaI, AclI, NarI, XhoI, SmlI, HpyF30I, PaeR7I, Sfr274I, or functional analogs thereof, or mixtures thereof.
60. 60. The method of any one of claims 57-59, further comprising carrying out (a) and (b) in the same reaction mixture.
61. 61. The method of any one of claims 57 to 60, further comprising distinguishing between methylated and unmethylated nucleobases in the adaptor-ligated DNA fragments.
62. 62. The method of claim 61, further comprising subjecting the adaptor-ligated fragments to bisulfite conversion.
63. 58. The method of claim 57, wherein said subjecting step comprises oxidizing the methylated cytosine nucleobases and / or hydroxymethylated cytosine nucleobases to produce oxidation reaction products, followed by reducing and / or deaminating the oxidation reaction products.
64. 64. The method of claim 63, wherein the oxidizing step is carried out using a ten-eleven translocation (TET) enzyme.
65. 64. The method of claim 63, wherein the oxidizing step is carried out using potassium perruthenate.
66. 64. The method of claim 63, wherein the step of reducing and / or deaminating the oxidation reaction product is carried out using an APOBEC.
67. 64. The method of claim 63, wherein the step of reducing and / or deaminating the oxidation reaction product is carried out using pyridine borane.
68. 68. The method of any one of claims 57 to 67, further comprising the step of performing β-glucosyltransferase treatment prior to one or more enzymatic and / or chemical reactions.
69. 69. The method of any one of claims 57-68, wherein the adapter comprises a GC overhang.
Citation Information
Patent Citations
Method of library preparation avoiding the formation of adaptor dimers
US20100167954A1
Compositions and methods for improving sample identification in indexed nucleic acid libraries
US20180305753A1
Enrichment of short nucleic acid fragments in sequencing library preparation
US20190093102A1
Methods for removal of adaptor dimers from nucleic acid sequencing preparations
WO2018024671A1