Methods for preparing hybrid ssDNA- and dsDNA-NGS libraries
A hybrid library preparation method combining dsDNA and ssDNA techniques improves molecular recovery and accuracy in DNA methylation detection, addressing low recovery rates in bisulfite sequencing and enhancing cancer detection.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-04
- Publication Date
- 2026-03-27
AI Technical Summary
Current DNA methylation detection methods, such as bisulfite sequencing, suffer from low molecular recovery rates due to the harsh chemical treatment that degrades DNA, limiting analytical and clinical sensitivity, especially in the context of cancer detection and minimal residual disease monitoring.
A hybrid library preparation method combining aspects of both dsDNA and ssDNA library preparations, leveraging molecular barcodes and adapters to improve molecular recovery and provide information on DNA molecular topology, enabling multi-omics workflows.
Enhances molecular recovery and accuracy in detecting cancer-specific biomarkers, reducing unnecessary treatments and improving patient outcomes by providing deeper insights into cancer-related DNA and protein changes.
Smart Images

Figure 2026510027000001_ABST
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 63 / 488,898, filed Mar. 7, 2023, and U.S. Provisional Patent Application No. 63 / 502,826, filed May 17, 2023.
[0002] Field of the Invention Methods and compositions related to the detection of nucleic acids and the preparation for the preservation, enrichment, storage, and sequencing of samples are described herein.
Background Art
[0003] Background DNA methylation detection by massively parallel sequencing is a promising methodology for highly sensitive detection of the presence of cancer in liquid biopsies and is applicable for screening and early detection. The gold - standard single - site methylation sequencing, such as bisulfite sequencing, provides high - resolution methylation status and changes in cell - free DNA (cfDNA) molecules, but many molecules are lost, thereby limiting analytical and clinical sensitivity. This is due to the fact that bisulfite is a harsh chemical treatment that non - specifically degrades DNA in the cytosine deamination reaction used to resolve methylated and non - methylated cytosine bases. To boost the molecular recovery rate in bisulfite sequencing, an improved library preparation methodology is needed.
[0004] As a system of next-generation sequencing (NGS) library preparation methods, there are double-stranded DNA library preparation (dsDNA-LP) and single-stranded DNA library preparation (ssDNA-LP), which operate using either double-stranded DNA (dsDNA) molecules or single-stranded DNA (ssDNA) molecules as substrates, respectively. These have different advantages and suitability, so one or the other is used depending on the given situation. There is a need for a “hybrid” library preparation workflow that combines aspects of both workflows, improving molecular recovery, providing information on DNA molecular topology, and enabling novel multi-omics workflows. [Overview of the project] [Means for solving the problem]
[0005] Summary of the Invention This disclosure provides methods and systems for boosting molecular recovery, for example, using bisulfite sequencing, by leveraging the features of both dsDNA and ssDNA library preparations. Such methods, which combine aspects of both workflows in “hybrid” library preparation, can improve molecular recovery, provide information on DNA molecular topology, and enable novel multi-omics workflows. High molecular recovery is critical in early cancer detection and minimal residual disease (MRD) monitoring. Integrating genomic, epigenomic, and molecular topological information can further amplify disease-specific signals to monitor treatment response and resistance, and / or identify markers of disease onset, progression, and metastasis in cfDNA or tissue. Furthermore, these methods can lead to a deeper understanding of cancer-causing DNA and protein changes, potentially enabling the identification of tumor-specific biomarkers and the design of treatments targeting those proteins. The methods can also improve the recovery of tumor-specific biomarkers that can inform treatment choices. Such treatments may include small molecule drugs or monoclonal antibodies. This method can also improve biomarker testing in individuals affected by disease, helping to determine whether an individual is a candidate for a particular drug or drug combination based on the presence or absence of a biomarker. Furthermore, this method can improve the identification of mutations that contribute to the development of resistance to targeted therapy. Consequently, this analytical technique may reduce unnecessary or untimely treatment interventions, patient suffering, and patient mortality.
[0006] Therapies can work by assisting the immune system in destroying cancer cells. For example, certain targeted therapies can mark cancer cells so that the immune system can destroy them. Other targeted therapies can support the immune system to make its action against cancer more effective. Still other therapies can stop the growth of cancer cells by interfering with cancer cell surface markers to prevent cancer cell division. Furthermore, therapies can inhibit signals that promote angiogenesis. Such angiogenesis inhibitors disrupt the blood supply to the tumor, thereby preventing tumor growth. Other targeted therapies can deliver toxic substances to the tumor. Examples include toxins, chemotherapy, or monoclonal antibodies combined with radiation. Some targeted therapies induce apoptosis in cancer or deplete hormones in cancer cells.
[0007] In some embodiments, the treatment is a PARP inhibitor, e.g., olaparib (Lynparza), lucaparib (Rubraca), niraparib (Zejula), and talazoparib (Talzenna). In some embodiments, the treatment includes immunotherapy and / or immune checkpoint inhibitors (ICIS), e.g., anti-PD-1 / PD-L1 treatments including pembrolizumab (Keytruda), nivolumab (Opdivo), and semiprimab (Libtayo), atezolizumab (Tecentriq), durvalumab (Imfinzi), and avelumab (Bavencio). In some embodiments, the treatment targets variant forms of the EGFR protein. Examples of such treatments include osimertinib (Tagrisso), erlotinib (Tarceva), and gefitinib (Iressa).
[0008] In some embodiments, the treatment may include one or more of the following targeted therapies: abemaciclib (Verzenio), abiraterone acetate (Zytiga), acalabrutinib (Calquence), adaglacib (Krazati), and ado-trastuzumab emtansine (ado-TrastuzumabEmtansine (Kadcyla), afatinib dimaleate (Gilotrif), alectinib (Alecensa), alemtuzumab (Campath), alitretinoin (Panretin), alpelisib (Piqray), amivantamab-vmjw (Rybrevant), anastrozole (Arimidex), apalutamide (Erleada), aciminib hydrochloride (Scemblix), atezolizumab (Tecentriq), atezoli Zumab (Tecentriq), avapritinib (Ayvakit), avelumab (Bavencio), axicatebutagensilolucel (Yescarta), axitinib (Inlyta), belinostat (Beleodaq), verzutifan (Welireg), bevacizumab (Avastin), bexarotene (Targretin), binimetinib (Mektovi), blinatumomab (Blincyto), bortezomib (Velcade), bosutinib ( Bosulif), brentuximab vedotin (Adcetris), brexcabutadiene oatlucel (Tecartus), brigatinib (Alunbrig), cabazitaxel (Jevtana), cabozantinib-s-malate (Cabometyx), cabozantinib-s-malate (Cometriq), capmatinib hydrochloride (Tabrecta), carfilzomib (Kyprolis), semiprimab-rwlc (Libtayo), ceritinib (Z (ykadia), cetuximab (Erbitux), siltacabutage autolucer (Carvykti), cobimetinib fumarate (Cotellic), copanlisib hydrochloride (Aliqopa), crizotinib (Xalkori), dabrafenib (Tafinlar), dabrafenib mesylate (Tafinlar), dacomitinib (Vizimpro), daratumumab (Darzalex), daratumumab and hyaluronidase-fihj (Darzalex)Faspro), darolutamide (Nubeqa), dasatinib (Sprycel), deniroikin difutitox (Ontak), denosumab (Xgeva), dinutuximab (Unituxin), dostallumab-gxly (Jemperli), durvalumab (Imfinzi), duvelisib (Copiktra), elacestrant dihydrochloride fam-trastuzumab deruxtecan-nxki (Enhertu), Latinib hydrochloride (Inrebic), fulvestrant (Faslodex), futivatinib (Lytgobi), gefitinib (Iressa), gemtuzumab ozogamicin (Mylotarg), gilteritinib fumarate (Xospata), gladegib maleate (Daurismo), ibritumomab tiuxetan (Zevalin), ibrutinib (Imbruvica), idekabutagenbiculucel (Abecma), idelalisib (Zydelig), imatinib mesylate (Gleevec), infiglatinibrinate (Truseltiq), inotuzumab ozogamicin (Besponsa), iobenguan I131 (Azedra), Ipilimumab (Yervoy), Isatuximab-irfc (Sarclisa), Ivosidenib (Tibsovo), Ixazomib citrate (Ninlaro), Lanreotide acetate (SomatulineDepot), Lapatinib ditosylate (Tykerb), Lalotrectinib sulfate (Vitrakvi), Lenvatinib mesylate (Lenvima), Letrozole (Femara), Lysokabutagenmaleucel (Breyanzi), Loncastuximab tesirin-lpyl (Zynlonta), Lorlatinib (Lorbrena), Lutetium Lu 177 Bipivotide tetraxetan (Pluvicto), Lutetium Lu 177-Doteate (Lutathera), Margetuximab-cmkb (Margenza), Midostaurin (Rydapt), Milbetuximab sorabtansine-gynx (Elahere), Mobocertinib succinate (Exkivity), Mogamulizumab-kpkc (Poteligeo), Mosnetuzumab-axgb (Lunsumio), Moxetumomab Pasudotox-tdfk (Lumoxiti), Naxitamab-gqgk (Danyelza), Necitumumab (Portrazza), Neratinib maleate (Nerlynx), Nilotinib (Tasigna), Niraparib tosylate monohydrate (niraparib tosylate)(monohydrate) (Zejula), nivolumab (Opdivo), nivolumab and relatrimab-rmbw (Opdualag), obinutuzumab (Gazyva), ofatumumab (Arzerra), olaparib (Lynparza), ortasidenib (Rezlidhia), osimertinib mesylate (Tagrisso), pacritinib citrate (Vonjo), palbociclib (Ibrance), panitumumab (Vectibix), pazopanib hydrochloride (Votrient), pembrolizumab (Keytruda), pemigatinib (Pemazyre), pertuzumab (Perjeta), pertuzumab, trastuzumab Mab, and hyaluronidase-zzxf (Phesgo), pexidartinib hydrochloride (Turalio), pirtobrutinib (Jaypirca), polatuzumab vedotin-piiq (Polivy), ponatinib hydrochloride (Iclusig), pralatrexate (Folotyn), prarlsetinib (Gavreto), radium-223 dichloride (Xofigo), ramucirumab (Cyramza), regorafenib (Stivarga), retifanrimab-dlwr (Zynyz), ribociclib (Kisqali), ripretinib (Qinlock), rituximab (Rituxan), rituximab and hyaluronidase human (Rituxan)Hycela), Romidepsin (Istodax), Rucaparibucansilate (Rubraca), Ruxolitinibulinate (Jakafi), Sacituzumab Govitecan-hziy (Trodelvy), Selinexor (Xpovio), Serpercatinib (Retevmo), Selumetinib sulfate (Koselugo), Siltuximab (Sylvant), Sirolimus protein-binding particles (Fyarro), Sonidegib (Odo mzo), sorafenib tosylate (Nexavar), sotrasib (Lumakras), sunitinib malate (Sutent), tafacitamab-cxix (Monjuvi), tagraxofusp-erzs (Elzonris), talazoparib tosylate (Talzenna), tamoxifen citrate (Soltamox), tazemetostat hydrobromide (Tazverik), teventafusp-tebn (Kimm Trak, Teclistamagb-CQYV (Tecvayli), Temsirolimus (Torisel), Tepotinib hydrochloride (Tepmetko), Tisagenlecleucel (Kymriah), Tisotumab vedotin-TFTV (Tivdak), Tivozanib hydrochloride (Fotivda), Toremifene (Fareston), Trametinib (Mekinist), Trametinib dimethyl sulfoxide (Mekinist), Trastuz Mab (Herceptin), Tremelimumab-actl (Imjudo), Tretinoin (Vesanoid), Tucatinib (Tukysa), Vandetanib (Caprelsa), Vemurafenib (Zelboraf), Venetoclax (Venclexta), Bismodegib (Erivedge), Vorinostat (Zolinza), Zanubrutinib (Brukinsa), Ziv-Aflibercept (Zaltrap).
[0009] In one embodiment, the present disclosure is a method for preparing a sequencing library from DNA molecules in a sample, comprising: (a) providing a first population of DNA molecules derived from the sample, wherein the first population of DNA molecules comprises double-stranded DNA and single-stranded DNA; (b) ligating a first set of adapters, each comprising a molecular barcode configured to attach to a plurality of double-stranded DNA molecules, to produce a second population comprising a plurality of double-stranded DNA molecules to which the adapters are ligated to both ends or one end of the double-stranded DNA molecules, and a plurality of unligated DNA molecules; and (c) subjecting the second population to a process that denatures and fragments the plurality of adapter-ligated DNA molecules and unligated DNA molecules, to produce a third population of DNA molecules comprising single-stranded DNA molecules having adapters at both ends, having an adapter at one end, and / or having no adapter at either end, as well as fragmented single-stranded DNA molecules, wherein the fragmented single-stranded DNA molecules have one at one end. The method provides a sequencing library derived from a population of DNA molecules in a sample, comprising the steps of: (d) ligating a second set of adapters to a subset of molecules in a third population, where the adapters are ligated to the fragment or the adapters are not ligated to the fragment, thereby generating a tagged DNA molecule comprising at least two of the following: (i) adapter-ligated single-stranded DNA, each containing adapters from the first set of adapters ligated to both ends of the molecule; (ii) adapter-ligated single-stranded DNA, each containing one adapter from the first set of adapters ligated to one end of the molecule and one adapter from the second set of adapters ligated to the other end of the molecule; and (iii) adapter-ligated single-stranded DNA, each containing adapters from the second set of adapters ligated to both ends of the molecule.In various embodiments, the adapter is a hairpin.
[0010] In some embodiments, the first population of DNA molecules includes double-stranded and single-stranded cell-free DNA (cfDNA).
[0011] In some embodiments, the first set of adapters is a Y-shaped adapter. In some embodiments, the first set of adapters is protected from processing in (c). In some embodiments, the first set of adapters further includes single-stranded ends protected from ligation using modifications containing a 5'OH and / or 3'P. In some embodiments, the first set of adapters includes single-stranded ends protected from ligation without modifications containing a 5'OH and / or 3'P when T4 PNK is used in (d). In some embodiments, the first set of adapters further includes single-stranded ends protected from ligation using modifications containing a 5'C3 spacer, a 5' inverted dideoxy base, another 5' spacer, a 3'C3 spacer, a 3' inverted dT, a 3' dideoxy base, and another 3' spacer when T4 PNK is used in (d). In some embodiments, the first set of adapters includes a universal amplification sequence. In some embodiments, the molecular barcode distinguishes the molecule ligated in (b) from the molecule ligated in (d).
[0012] In some embodiments, the population of DNA molecules is phosphorylated using T4 PNK before ligation in (b). In some embodiments, the population of DNA molecules is phosphorylated using T4 PNK before ligation in (d).
[0013] In some embodiments, the process of denaturing and fragmenting a second population of DNA molecules includes at least one of bisulfite conversion, Tet-assisted bisulfite conversion, and Tet-assisted conversion using a substituted borane reducing agent, the substituted borane reducing agent being optionally 2-picoline borane, borampyridine, tert-butylamine borane, or ammonia borane. In some embodiments, the process of denaturing and fragmenting a second population of DNA molecules includes chemical-assisted conversion using a substituted borane reducing agent, the substituted borane reducing agent being optionally 2-picoline borane, borampyridine, tert-butylamine borane, or ammonia borane.
[0014] In some embodiments, the first set of adapters includes methylated cytosine to protect them from processing. In some embodiments, the second set of adapters is a “sprint” adapter containing a 5' or 3' overhang. In some embodiments, the second set of adapters includes a universal amplification sequence. In some embodiments, the second set of adapters includes (i) an adapter having a double-stranded portion and a single-stranded overhang containing a randommer sequence on the 5' side of the reverse strand, and (ii) an adapter having a double-stranded portion and a single-stranded overhang containing a randommer sequence on the 3' side of the reverse strand. In some embodiments, the second set of adapters selectively tags only the ends of DNA molecules that lack the first adapter sequence.
[0015] In some embodiments, the method further includes the step of amplifying the molecules in (d)(i) to (iii) to generate replicated DNA molecules. In some embodiments, the method further includes the step of sequencing the amplified DNA molecules to generate sequencing reads. In some embodiments, the amplified molecules are captured before sequencing to enrich one or more target regions. In some embodiments, the amplified molecules include an index before sequencing.
[0016] In some embodiments, unique molecules are degraded from PCR replicas by analyzing sequencing reads to identify molecules containing the same terminal coordinates and / or molecular barcodes that map to a reference sequence.
[0017] In some embodiments, a DNA molecule having a first adapter at one end and a second adapter at the other end has greater end coordinate diversity than the molecule in (d)(i), and this diversity of fragment ends improves the accuracy of degrading specific molecules from PCR replicas. In some embodiments, a DNA molecule having second adapters at both ends of (d)(iii) has greater end coordinate diversity than the molecules in (d)(i) and (ii), and this diversity improves the accuracy of degrading specific molecules from PCR replicas. In some embodiments, end coordinate diversity improves the accuracy of degrading specific molecules from PCR replicas. In some embodiments, end coordinates and molecular barcodes improve the accuracy of degrading specific molecules from PCR replicas.
[0018] In some embodiments, the symmetry between the forward and reverse strands improves the accuracy of detecting methylation status. In some embodiments, the symmetry between the forward and reverse strands improves the accuracy of detecting somatic mutations. In some embodiments, for DNA molecules having a first adapter at both ends of (d)(i), a consensus is generated that explains the symmetry between the forward and reverse strands to determine methylation status and / or somatic mutations.
[0019] In some embodiments, prior to (b), DNA molecules within the first population undergo end repair and / or A-tail addition.
[0020] In another embodiment, the present disclosure is a method for analyzing a population of DNA molecules in a sample, comprising: (a) ligating a first set of adapters containing molecular barcodes to at least a subset of DNA molecules in the population of DNA molecules, wherein the adapters are ligated to both ends of the DNA molecules to produce ligated DNA molecules; (b) subjecting the population of DNA molecules to a biochemical treatment that denatures and randomly fragments a plurality of DNA molecules, thereby producing fragmented DNA molecules having two or fewer first adapters at their ends; and (c) ligating a second set of adapters to their ends The present invention provides a method comprising the steps of: ligating DNA molecules in a population having fewer than two first adapter sequences to generate a population of tagged DNA molecules including at least two of the following: (i) DNA molecules having a first adapter at both ends, (ii) DNA molecules having a first adapter at one end and a second adapter at the other end, and (iii) DNA molecules having a second adapter at both ends; (d) sequencing the population of tagged DNA molecules from (c) to generate sequencing reads; and (e) analyzing the sequencing reads to detect parallel signals associated with disease. In various embodiments, the adapter is a hairpin.
[0021] In some embodiments, the disease-associated signals further include one or more of the following: the ratio of short cfDNA fragments to long cfDNA fragments, differences in chromatin architecture, nucleosome structure, epigenetics, and / or genetic information.
[0022] In yet another aspect, the Disclosure provides a method for preparing a sequencing library from a population of DNA molecules in a sample, comprising the steps of (a) ligating a subset of DNA molecules in a population a molecular barcode pair comprising (i) a first set of molecular barcodes including a ligable end and a 3' single-strand overhang at the opposite end and (ii) a second set of molecular barcodes including a ligable end and a 5' single-strand overhang at the opposite end to generate a plurality of tagged DNA molecules having molecular barcodes at each end; and (b) ligating a set of adapters to the plurality of tagged DNA molecules and a plurality of untagged DNA molecules that were not ligated in (a) to generate (i) adapter-tagged molecules containing molecular barcodes and (ii) adapter-tagged molecules that do not contain barcodes, thereby providing a sequencing library derived from a population of DNA molecules in a sample.
[0023] In some embodiments, 5' and 3' overhangs prevent pairs of barcodes from ligating each other or self-ligating. In some embodiments, ligable ends include a T-tail or an A-tail. In some embodiments, ligable ends include a blunt end.
[0024] In some embodiments, the barcode further includes an adapter sequence. In some embodiments, the adapter attaches to the 3' and 5' single-stranded overhangs of the tagged DNA molecule.
[0025] In some embodiments, the set of adapters includes (i) an adapter having a double-stranded portion and a single-stranded overhang including a randomer sequence on the 5' side of the reverse strand, and (ii) an adapter having a double-stranded portion and a single-stranded overhang including a randomer sequence on the 3' side of the reverse strand. In some embodiments, the adapter is a "sprint" adapter containing a 5' or 3' overhang. In some embodiments, the adapter includes a universal amplification sequence. In some embodiments, the method further includes amplifying a sequencing library to generate amplified (i) adapter-tagged molecules containing a molecular barcode and (ii) adapter-tagged molecules not containing a barcode.
[0026] In some embodiments, the method further includes selectively enriching the library to isolate a subset of the molecules of (i) and (ii) above. In some embodiments, the method further includes sequencing a library of the molecules of (i) and (ii) above to generate sequencing reads. In some embodiments, the selective enrichment is performed by hybridization or amplification techniques. In some embodiments, prior to sequencing, the amplified molecules include an index.
[0027] In some embodiments, the enriched subset of the molecules of (i) and (ii) above is associated with a disease. In some embodiments, the disease is cancer, Alzheimer's disease, hypertriglyceridemia, coronary artery disease.
[0028] In yet another aspect, the present disclosure provides a method for preparing a sequencing library from a population of DNA molecules in a sample, the method comprising: (a) ligating a first set of adapters comprising a capture label to at least a subset of the DNA molecules to generate a first subset of adapter-ligated DNA molecules comprising the capture label, wherein the adapters are ligated to both ends of the DNA molecules; (b) separating the adapter-ligated DNA molecules comprising the capture label by contacting the population of DNA molecules with a capture molecule, thereby generating (i) a first subset of adapter-ligated DNA molecules comprising the capture label, wherein the label is bound to the capture molecule, and (ii) non-adapted DNA molecules; and (c) ligating a second set of adapters to at least a subset of the non-adapted DNA molecules to generate a second subset of adapter-ligated DNA molecules, wherein the adapters are ligated to both ends of the DNA molecules, thereby providing a sequencing library derived from the population of DNA molecules in the sample.
[0029] In some embodiments, the first set of adapters further comprises a Y-shaped adapter. In some embodiments, the first set of adapters further comprises a molecular barcode.
[0030] In some embodiments, the capture label comprises an affinity ligand. In some embodiments, the affinity ligand is biotin or photocleavable biotin. In some embodiments, the photocleavable biotin is biotin-UTP. In some embodiments, the separating in (b) further comprises affinity purification using at least one capture molecule. In some embodiments, the capture molecule is streptavidin or magnetic beads coated with streptavidin.
[0031] In some embodiments, the adapter is subjected to a process that digests a first set of ligated molecules into unmethylated DNA. In some embodiments, the captured molecules are treated with SMRE before, during, or after streptavidin treatment or application of streptavidin-coated magnetic beads.
[0032] In some embodiments, the second set of adapters is a “sprint” adapter containing a 5' or 3' overhang. In some embodiments, the second set of adapters includes a universal amplification sequence. In some embodiments, the second set of adapters includes (i) an adapter having a double-stranded portion and a single-stranded overhang containing a randommer sequence on the 5' side of the reversed strand, and (ii) an adapter having a double-stranded portion and a single-stranded overhang containing a randommer sequence on the 3' side of the reversed strand.
[0033] In some embodiments, the method further includes the step of sequencing the library to generate sequencing reads. In some embodiments, the method further includes the step of amplifying the library before sequencing. In some embodiments, the method further includes the step of selectively enriching the library before sequencing.
[0034] In some embodiments, the method further includes the step of analyzing the sequencing reads to determine the methylation status at one or more gene loci. In some embodiments, the method further includes the step of analyzing the sequencing reads to determine the molecular topology (e.g., molecules originating from double strands and molecules originating from single strands) of a second subset of DNA molecules to which the adapter in (c) is ligated.
[0035] In some embodiments, the method further includes the step of analyzing sequencing reads to detect parallel signals associated with a disease. In some embodiments, the disease includes cancer, Alzheimer's disease, hypertriglyceridemia, or coronary artery disease.
[0036] Additional aspects and advantages of the present disclosure will be readily apparent to those skilled in the art from the following detailed description, which is provided only for illustrative purposes. As will be understood, other different embodiments of the present disclosure are possible, and some of its details may be modified in various obvious ways, all of which will not deviate from the present disclosure. Accordingly, the drawings and description should be construed as illustrative in nature and not restrictively.
[0037] Novel features of this disclosure are described in detail in the appended claims. A better understanding of the features and advantages of this disclosure can be obtained by referring to the following detailed description, which includes exemplary embodiments in which the principles of this disclosure are utilized, and to the accompanying figures (also referred to herein as "Figure" and "FIG."). [Brief explanation of the drawing]
[0038] [Figure 1-1] Hybrid dsDNA- and ssDNA-LPs for increasing bisulfite-seq recovery. As shown in Figure 1A, a hybrid library preparation workflow for increasing bisulfite-seq molecular recovery is illustrated. In Figure 1B, ssDNA-LPs are added after denaturation and bisulfite treatment. Figure 1C, PCR amplification and depiction of amplicon libraries from selected molecules. [Figure 1-2] Same as above. [Figure 1-3] Same as above.
[0039] [Figure 2-1]Molecular topology. A hybrid library preparation workflow is shown to improve single-stranded DNA recovery and obtain molecular topology information. Figure 2A shows a method for adding cell-free single-stranded DNA (cf-ssDNA) to a liquid biopsy workflow using double-stranded DNA barcode ligation. Figure 2B shows a method for adding cf-ssDNA to a liquid biopsy workflow using hairpin dsDNA ligation. [Figure 2-2] Same as above.
[0040] [Figure 3-1] The biotinylated NGS adapter separates the dsDNA MSRE and ssDNA LP processing. As shown in Figure 3A, a hybrid library preparation workflow is shown to enable a multi-omics workflow that requires / is convenient for separate dsDNA-LP processing and ssDNA-LP processing. As shown in Figure 3B, the addition of the biotinylated NGS adapter is shown. [Figure 3-2] Same as above. [Modes for carrying out the invention]
[0041] Embedding by reference All publications, patents, and patent applications referenced herein are incorporated by reference to the same extent as any individual publication, patent, or patent application is specifically and individually incorporated by reference.
[0042] Detailed explanation In bisulfite sequencing (bisulfite-Seq) assays, such as methyl-Seq (Urich, M., Nery, J., Lister, R. et al. MethylC-seq library preparation for base-resolution whole-genome bisulfite sequencing. Nat Protoc 10, 475-483 (2015), incorporated herein by reference), dsDNA methylation adapter ligation is performed before the bisulfite conversion process. While this significantly increases the achievable adapter ligation efficiency (60-70%), the ligated DNA molecules become unamplified and unsequenable during bisulfite processing due to cleavage. Molecular recovery rates using these methods range from less than 1% to approximately 6% according to the literature. A newer bisulfite-Seq method exists that uses adapter tagging after bisulfite processing to enable sequencing of many molecules cleaved during bisulfite processing. IDT / Swift's 1D Accel Methyl-NGS and Claret Bio's SRSLY are examples of commercially available ssDNA-LPs.
[0043] Standard dsDNA-LP conversion (ligation) efficiency / recovery rate exceeds 60%. For methylation detection, conventional bisulfite treatment is required after dsDNA-LP / ligation, resulting in a very low recovery rate (less than 5%). On the other hand, while bisulfite treatment creates nicks in DNA molecules / cleaves DNA molecules (making the library molecules unamplified), DNA is not lost during the process (BS input quantification (quant) and BS output quantification are almost equal). ssDNA-LP can be applied after BS, but the recovery rate is limited (15-20%).
[0044] Comparing the potential performance of BS treatment and LP efficiency, dsDNA-LP shows better recovery when the number of nicks in the molecule is 0; ssDNA-LP shows better (non-zero) recovery when the number of nicks in the molecule is 1; and ssDNA-LP shows better (non-zero) recovery when the number of nicks in the molecule is 2 or more.
[0045] The "combined" library preparation (dsDNA-ligation and backup ssDNA-LP step) yields the highest overall recovery rate, and BS molecular results include: if there are 0 nicks in BS, the molecule is prepared by dsDNA-LP; if there is 1 nick in BS, one adapter is added with high efficiency by dsDNA-LP and a second adapter is added with low efficiency by ssDNA-LP, which is higher than ssDNA-LP alone; if there are 2+ nicks in BS, the "inner" fragment is prepared by ssDNA-LP (the only option), and the "outer" fragment is processed as described above ("if there is 1 nick in BS").
[0046] Therefore, while the recovery rate of ssDNA-LP is higher than that of dsDNA-LP, a significantly higher recovery rate may be achieved through composite library preparation, depending on the number of molecules with zero and one nicks. Furthermore, since non-random UMIs (multiple UMIs) can be obtained by dsDNA ligation for molecules with zero and one nicks, it is highly likely that complete degradation of molecules by UMI is not necessary in ssDNA-LP. In addition, there are start and stop coordinates, and in molecules with one nick, diversity should be added by the "nick" coordinate, and in molecules with 2+ nicks, there should be significant diversity in the terminal coordinates from the nick. Moreover, since molecules with zero and one nicks can be completely defined, fewer molecules need to be degraded.
[0047] Figure 1 shows a hybrid library preparation workflow to increase bisulfite-seq molecule recovery. These methods leverage the characteristics of each LP type to boost molecule recovery in bisulfite-seq. Before bisulfite, dsDNA (methylated) adapters are efficiently ligated to dsDNA molecules. Then, bisulfite is performed, which cleaves the dsDNA adapter-attached molecules at zero, one, or more than one site. After bisulfite, ssDNA-LP is performed to selectively rescue molecular fragments that have fewer than two dsDNA adapter sequences at their ends (adapter tagging). If a molecule tagged with a dsDNA adapter is not cleaved by bisulfite, it is unaffected by ssDNA LP; however, if the molecule contains one cleavage, ssDNA preparation tags the cleaved end with an adapter, preparing the molecule for sequencing. The resulting library molecules will have an adapter generated by dsDNA-LP at one end and an adapter generated by ssDNA-LP at the other end. If a DNA molecule is cleaved at more than one site by bisulfite, a molecule with bisulfite-generated cleavage at both ends will have ssDNA-LP-generated adapters at both ends. Non-random molecular barcodes can be used with the dsDNA-LP adapters, enabling molecular resolution power for DNA fragments with zero cleavage and DNA fragments with one cleavage in sequencing. Furthermore, by attaching this adapter type to each end of a DNA molecule, it is possible to determine whether the molecule originates from a fragment end or from bisulfite-generated cleavage. This information is identifiable in sequencing (by checking for the presence or absence of molecular barcodes).
[0048] Aspects of this disclosure provide a method for preparing a sequencing library from DNA molecules in a sample using a hybrid library preparation method. Prior to bisulfite treatment, end-prepared double-stranded DNA molecules are tagged by ligation (T4 DNA ligase) using a Y-shaped NGS adapter having a molecular barcode. Cytosine bases in the Y-shaped adapter are protected from bisulfite conversion, and single-stranded ends are protected from ligation. Bisulfite conversion is performed, thereby fragmenting a portion of the NGS adapter-attached DNA molecule. The bisulfite-converted DNA is subjected to ssDNA-LP using a “sprint” adapter containing a 5' or 3' overhang, and a suitable NGS universal amplification sequence also contained in the Y-adapter. The sprint adapter selectively tags only the DNA ends lacking the Y-adapter sequence. This occurs because the first ligation fails or the DNA molecule is fragmented by bisulfite. The latter is a significant group that the present invention aims to rescue. After sprint adapter ssDNA ligation, all molecules with adapters (of either type) at the 5' and 3' ends can be amplified using universal primers and processed downstream for NGS, i.e., enriched by hybrid capture before sequencing. After sequencing, molecules can be identified and degraded by analysis using DNA end coordinates and molecular barcodes, if present in the molecule / read. Bisulfite fragmentation introduces greater diversity at the fragment ends, and therefore, molecules fragmented by bisulfite (without molecular barcodes) can be degraded without molecular barcodes. Molecules that are not completely fragmented will have minimal diversity with respect to end coordinates and will have two molecular barcode tags used in conjunction with the fragment ends to degrade the molecule. For molecules where one end is fragmented while the low-diversity end contains a molecular barcode, the molecule can be identified by a combination of the fragment end and a single molecular barcode.Thus, the "hybrid" library preparation protocol also improves the accuracy of molecular degradation in the bisulfite sequencing workflow.
[0049] Modification of the ssDNA terminal of the Y adapter.
[0050] Commercial ssDNA-LP "sprint adapter ligation" kits include a step that occurs concurrently with ligation, which phosphorylates the 5' end and removes the 3' phosphate group if necessary, thereby preparing the dsDNA for ligation. This step can be omitted, as it is performed during the upstream "end repair" that prepares the dsDNA for Y adapter ligation, in the case of the hybrid workflow described herein. In this case, the simplest and most economical modifications to use are the 5' hydroxyl group and the 3' phosphate group. However, if the ssDNA-LP T4 PNK step is left in the protocol, these two modifications cannot be used because they become substrates for T4 PNK activity and are converted into ligable ends. In protocols involving the ssDNA-LP PNK step, similar suitable modifications to the ssDNA ends of the Y adapter are a 5' C3 spacer, a 5' inverted dideoxy base, another 5' spacer, a 3' C3 spacer, a 3' inverted dT, a 3' dideoxy base, and another 3' spacer. This is because these interfere with / inhibit T4 PNK activity.
[0051] Loss of ssDNA molecule and DNA molecule topology information (whether of ssDNA origin or dsDNA origin) in gene sequencing workflows. Cell-free DNA (cfDNA) primarily exists as dsDNA and is effectively sequenced by the dsDNA-LP method. However, trace amounts of potentially biologically important cf-ssDNA, as well as cf-dsDNA molecules that are denatured to ssDNA before LP (the extraction process slightly denatures them), are not sequenced by the dsDNA-LP process. Using the ssDNA-LP process alone to recover these ssDNA and dsDNA is not ideal because the molecular recovery rate is relatively low and information about the molecular topology of origin (whether it is ssDNA or dsDNA) is not preserved.
[0052] Figure 2, including Figures 2A and 2B, illustrates a hybrid library preparation workflow for improving single-stranded DNA recovery and obtaining molecular topology information. Figure 2A shows an exemplary method for recovering cf-ssDNA from a liquid biopsy workflow using double-stranded DNA barcode ligation. In this embodiment, double-stranded DNA molecules are ligated with a barcode tag that identifies dsDNA, and then the sequencing reads are analyzed bioinformatics to extract the original topology (single-stranded or double-stranded) of the DNA molecule. After ligation of the dsDNA barcode tag, the sample is then subjected to ssDNA-LP. This hybrid LP method amplifies and sequences both ssDNA and dsDNA-derived DNA molecules, and the original topology (ss or ds) of the molecule is identified by a specific dsDNA barcode tag.
[0053] The barcode tag identifying the dsDNA provides information that the molecule originates from dsDNA. Specifically, multiple dsDNA barcode tags can be used, and they can also function as molecular barcodes. dsDNA ligation may simply involve attaching the "dsDNA barcode tag," or it may involve attaching the barcode tag as part of a Y adapter. After dsDNA barcode tag ligation, the sample is then subjected to ssDNA-LP. If only the "tag" is attached in the "dsDNA barcode" ligation step, the ssDNA-LP attaches amplification and sequencing adapters to these molecules, as well as molecules originating from ssDNA. Molecules originating from ssDNA do not contain the dsDNA barcode tag and are therefore degradable during sequencing. If a Y adapter (with a dsDNA barcode) is ligated to a dsDNA molecule in the first step, the same Y adapter modifications described above are applied here to prevent undesirable ligation with respect to the ssDNA-LP. Figure 2B illustrates an alternative method for recovering cf-ssDNA from a liquid biopsy workflow using hairpin adapter ligation. In this embodiment, a dsDNA molecule is ligated to a T-tailed (optional) hairpin NGS adapter (such as those used in the NEBNext kit) modified to have a molecular barcode, and then subjected to the ssDNA LP method using sprint adapter ligation. An exemplary workflow is shown in Example 2.
[0054] Incompatibility with certain multi-omics workflows using standard LP methods. Methylation-sensitive restriction enzyme (MSRE) sequencing workflows offer the advantage of high molecular recovery in methylation analysis. To achieve the highest accuracy of methylation calls with the lowest sequencing load, dsDNA-LP is used, and MSRE processing is performed after Y adapter ligation but before amplification. Thus, only molecules with fully methylated MSRE recognition sites are amplified and sequenced. Such workflows require dsDNA-LP, which hinders / blocks the simultaneous detection of other omic signals that require / are favorable to ssDNA-LP workflows (e.g., fragment mix investigating the ratio of possible short cfDNA fragments to long cfDNA fragments, as well as other fragment mix analyses investigating subnucleosomal cfDNA molecules and features defining subnucleosomal diseases in cell-free DNA, and nucleosome footprints that provide information about the tissue of origin).
[0055] Figure 3 shows a hybrid library preparation workflow to enable a multi-omics workflow that requires / is convenient for separate dsDNA-LP and ssDNA-LP processing. To prepare ssDNA molecules for sequencing while allowing the removal of MSRE processing from a sequencing library of unmethylated (dsDNA) molecules, enrichable reagents can be incorporated into the dsDNA ligation and / or MSRE cleavage repair step to separate their substrates from the ssDNA molecules, which are then prepared for sequencing by an ssDNA-LP step. If this enrichment / separation step is not performed, MSRE cleavage products may be prepared for sequencing by the undesirable ssDNA-LP step.
[0056] Use of biotinylated Y adapters for isolating ssDNA material. Double-stranded DNA molecules are ligated with a biotinylated Y adapter after end repair / tail addition. Streptovidin-magnetic beads and a magnet are applied to immobilize / pull down the ligated dsDNA, and the ssDNA molecules in the supernatant are removed and placed in a separate reaction vessel. In a parallel reaction, MSRE treatment is applied to the adapter-ligated dsDNA, and the resulting library, lacking molecules unmethylated at the MSRE site, is amplified for downstream processing. The reaction vessel containing the ssDNA is then subjected to ssDNA-LP. After the ligation step in the ssDNA-LP step, it is safe to combine the ssDNA library with the MSRE-post-dsdDNA library for co-library amplification and / or hybrid capture and / or sequencing.
[0057] Alternatively, in some embodiments, biotin dATP can be added to the A-tail addition step to biotin-tagged the dsDNA library molecules. This method of biotinylating dsDNA molecules may negatively affect amplification efficiency and, consequently, process sensitivity. The use of photocleavable or standard biotin in the Y-adapter and the addition of a cleavable base (i.e., uracil) is an optional method that would allow the dsDNA library molecules to be released from immobilization after the ssDNA molecules have been removed from the sample. This may be beneficial if immobilization causes steric effects that inhibit the MSRE and / or PCR amplification efficiency of the dsDNA library molecules.
[0058] Specific biotinylation of MSRE products to remove them from ssDNA-LP. In this method, the Y adapter may contain the previously considered ssDNA end modification to prevent the end from being modified in downstream enzymatic steps. The dsDNA molecule is repaired / A-tailed, ligated with the modified Y adapter, and then subjected to MSRE treatment. After MSRE treatment, the DNA is repaired and A-tailed again, but this time containing biotinylated nucleotides. The use of biotin-dATP for A-tail addition is necessary to tag the MSRE reaction product (e.g., HpaII) with a remaining 5' overhang. The Y adapter ssDNA ends, as well as the ssDNA molecules in the sample, should be inert to this repair / A-tail addition. The MSRE product is immobilized / pulled down by applying streptavidin-magnetic beads and a magnet, and the undigested dsDNA library molecules and ssDNA molecules are removed as supernatant and placed in a new reaction vessel. ssDNA-LP is performed on the new reaction vessel to selectively prepare the ssDNA molecules. In the same tube, co-amplification is performed on both the dsDNA library molecules and the ssDNA library molecules after MSRE, and the products are taken for downstream processing such as hybrid capture and / or sequencing.
[0059] This disclosure provides a method for developing next-generation sequencing libraries useful for cancer screening tests (cancer-specific tests or tests for multiple cancers) and for incorporating key genomic regions associated with disease into targeted sequencing workflows (e.g., hybrid capture panels). This includes monitoring genetic and epigenetic changes related to treatment response and resistance, and / or identifying markers of disease onset, progression, and metastasis in cfDNA or tissue.
[0060] 1. Sample The sample may be any biological sample isolated from the subject. Examples of samples include body tissues, such as known or suspected solid tumors, whole blood, platelets, serum, plasma, stool, red blood cells, white blood cells or leucocytes, endothelial cells, tissue biopsy material, cerebrospinal fluid, synovial fluid, lymph, ascites, interstitial fluid or extracellular fluid, intercellular fluids including gingival crevicular exudate, bone marrow, pleural fluid, cerebrospinal fluid, saliva, mucus, sputum, semen, sweat, and urine. The sample is preferably a body fluid, particularly blood and its fractions, as well as urine. The sample may be in the form originally isolated from the subject and may have been subjected to further processing to remove or add components such as cells, or to enrich one component by comparing it with another. Therefore, the preferred body fluid for analysis is plasma or serum containing cell-free nucleic acids.
[0061] The volume of plasma may depend on the desired read depth of the region to be sequenced. Exemplary volumes are 0.4–40 mL, 5–20 mL, and 10–20 mL. For example, the volume could be 0.5 mL, 1 mL, 5 mL, 10 mL, 20 mL, 30 mL, or 40 mL. The volume of the sampled plasma could be, for example, 5–20 mL.
[0062] The sample may contain varying amounts of nucleic acids containing genome equivalents. For example, a sample of approximately 30 ng of DNA may contain approximately 10,000 haploid human genome equivalents, or approximately 200 billion individual nucleic acid molecules in the case of cell-free DNA. Similarly, a sample of approximately 100 ng of DNA may contain approximately 30,000 haploid human genome equivalents, or approximately 600 billion individual molecules in the case of cell-free DNA. Some samples contain 1-500, 2-100, or 5-150 ng of cell-free DNA, for example, 5-30 ng or 10-150 ng of cell-free DNA.
[0063] The sample may contain nucleic acids from different sources. For example, the sample may contain germline DNA or somatic DNA. The sample may contain nucleic acids with mutations. For example, the sample may contain DNA with germline mutations and / or somatic mutations. The sample may also contain DNA with cancer-related mutations (e.g., cancer-related somatic mutations).
[0064] Exemplary amounts of cell-free nucleic acids in the sample before amplification can range from about 1 fg to about 1 ug, for example, 1 pg to 200 ng, 1 ng to 100 ng, or 10 ng to 1000 ng. For example, the amount could be up to about 600 ng, up to about 500 ng, up to about 400 ng, up to about 300 ng, up to about 200 ng, up to about 100 ng, up to about 50 ng, or up to about 20 ng of cell-free nucleic acid molecules. The amount could be at least 1 fg, at least 10 fg, at least 100 fg, at least 1 pg, at least 10 pg, at least 100 pg, at least 1 ng, at least 10 ng, at least 100 ng, at least 150 ng, or at least 200 ng of cell-free nucleic acid molecules. The quantity may be up to 1 femtogram (fg), 10 fg, 100 fg, 1 picogram (pg), 10 pg, 100 pg, 1 ng, 10 ng, 100 ng, 150 ng, or 200 ng of cell-free nucleic acid molecules. The method may include obtaining 1 femtogram (fg) to 200 ng.
[0065] An example sample is 5–10 ml of whole blood, plasma, or serum containing approximately 30 ng of DNA, or approximately 10,000 haploid genome equivalents.
[0066] Cell-free nucleic acids are nucleic acids that are not contained within cells and are not bound to cells in any other form, or in other words, nucleic acids remaining in a sample after intact cells have been removed. Examples of cell-free nucleic acids include genomic DNA, mitochondrial DNA, siRNA, miRNA, circulating RNA (cRNA), tRNA, rRNA, small nucleolar RNA (snoRNA), Piwi-interacting RNA (piRNA), long non-coding RNA (long ncRNA), or DNA, RNA, and hybrids thereof, including fragments of any of these. Cell-free nucleic acids may be double-stranded, single-stranded, or hybrids thereof. For any method disclosed herein, a double-stranded DNA molecule having at least a single-stranded overhang is a preferred form of cell-free DNA. Cell-free nucleic acids may be released into body fluids by secretion or cell death processes, such as cell necrosis and apoptosis. Some cell-free nucleic acids, such as circulating tumor DNA (ctDNA), are released into body fluids from cancer cells. Other cell-free nucleic acids are released from healthy cells.
[0067] Cell-free nucleic acids may have one or more epigenetic modifications. For example, cell-free nucleic acids may be acetylated, methylated, ubiquitinated, phosphorylated, SUMOylated, ribosylated, and / or citrullinated.
[0068] The size distribution of cell-free nucleic acids is approximately 100–500 nucleotides, particularly 110–230 nucleotides, with a mode of approximately 168 nucleotides, and a second minor peak in the range of 240–440 nucleotides.
[0069] Cell-free nucleic acids can be isolated from body fluids by a partitioning step, which separates the cell-free nucleic acids found in the solution from intact cells and other insoluble components of the body fluid. Partitioning may involve techniques such as centrifugation or filtration. Alternatively, cells in the body fluid can be lysed, and the cell-free and cellular nucleic acids can be processed together. Generally, after buffer addition and washing steps, the nucleic acids can be precipitated with alcohol. Further clarification steps, such as silica-based columns for removing impurities or salts, can be used. For example, non-specific bulk carrier nucleic acids can be added throughout the reaction to optimize certain aspects of the procedure, such as yield.
[0070] After such processing, the sample may contain nucleic acids in various forms, including double-stranded DNA, single-stranded DNA, and single-stranded RNA. If necessary, single-stranded DNA and RNA can be converted to double-stranded forms, and therefore, the subsequent processing and analysis steps will include the double-stranded forms.
[0071] 2. Linking the sample nucleic acid molecule to the adapter. As described above, nucleic acids present in pre-processed or unprocessed samples typically contain substantial portions of the molecule in the form of partially double-stranded molecules with single-stranded overhangs. As shown at the top of Figure 1, such molecules can be converted to blunt-ended double-stranded molecules by processing them in the presence of one or more enzymes that impart 5'-3' polymerase and 3'-5' exonuclease (or proofreading function) in the presence of all four standard nucleotide types. Such a combination of activities can elongate a chain with a concave 3' end, resulting in the end being tightly aligned with the 5' end of the opposite chain (in other words, producing a blunt end), or digest a chain with a 3' overhang, resulting in the end being tightly aligned with the 5' end of the opposite chain. If necessary, both activities can be conferred by a single polymerase. It is preferable that the polymerase is heat-sensitive and therefore its activity can be terminated when the temperature is increased. Examples of suitable polymerases are Klenow large fragment and T4 polymerase.
[0072] It is preferable that one or more enzymes conferring 5'-3' polymerase activity and 3'-5' exonuclease activity be denatured by increasing the temperature or by other means. For example, denaturation can be brought about by raising the temperature to, for example, 75°C to 80°C. The sample is then treated with a polymerase lacking proofreading function (Figure 1, center). It is preferable that this polymerase be thermally stable, for example, so that its activity is maintained when the temperature is increased. Examples of such polymerases are Taq, Bst large fragment, and Tth polymerase. The second polymerase brings about the untemplated addition of a single nucleotide to the 3' end of the blunt-ended nucleic acid. The reaction mixture typically contains equal molar concentrations of each of the four standard nucleotide types from the previous step, but these four nucleotide types are not added to the 3' end in equal proportions. If anything, A is added most frequently, followed by G, then C and T.
[0073] In certain embodiments, after tail-addition to the sample molecule, and after subsequent purification of the tail-added sample molecule, or without purification, the tail-added sample molecule is brought into contact with an adapter to which complementary T and C nucleotides are tail-added to one end of the adapter (Figure 1, bottom). The adapters are typically formed by the separate synthesis and annealing of their respective chains. Therefore, additional T and C tails can be added as additional nucleotides in the synthesis of one of the chains. Typically, adapters with G and A tail-added are not included because these adapters can anneal to sample molecules with C and T tail-added, respectively, but also to other adapters. Adapter molecules and sample molecules having complementary nucleotides at the 3' end (i.e., TA and CG) can anneal and ligate with each other. The percentage of C-tailed adapters relative to T-tailed adapters ranges from approximately 5% to 40% molarly, for example, 10-35%, 15-25%, 20-35%, 25-35%, or approximately 30%. Since the non-template directed addition of a single nucleotide to the 3' end of the sample molecule does not proceed to completion, the sample also contains some blunt-ended sample molecules without tail additions. These molecules can also be recovered by supplying the sample with adapters having one, preferably only one, blunt end. Blunt-ended adapters are typically supplied in molar ratios of 0.2-20%, or 0.5-15%, or 1-10% relative to adapters having T-tailed and C-tailed adapters. Blunt-ended adapters can be supplied simultaneously with, before, or after, the T-tailed and C-tailed adapters. A blunt-ended sample molecule is ligated to a blunt-ended adapter, which then brings back the sample molecule sandwiched between the adapters on both sides. These molecules lack the AT or CG nucleotide pairs between the sample and the adapter that are present when a tail-attached sample molecule is ligated to a tail-attached adapter.
[0074] The adapters used in these reactions may have a tail-added T or C at only one end, or one end may be blunt, and thus may ligate with the sample molecule in only one orientation. The adapter may be, for example, a Y-shaped adapter with one end tail-added or blunt and the other end having two single strands. An exemplary Y-shaped adapter has the following sequence and has a tag (6 bases). The upper oligonucleotide contains a single-base T tail.
[0075] Customized combinations of such oligonucleotides, including oligonucleotides having both a T-tail and a C-tail, can be synthesized for use in this method.
[0076] Shortened versions of these adapter sequences are described in Rohland et al., Genome Res. 2012 May; 22(5): 939-946.
[0077] The adapter may be bell-shaped, with only one end tail-added or blunt. The adapter may include a primer binding site for amplification, a binding site for sequencing primers, and / or a nucleic acid tag for identification purposes. The same or different adapters may be used in a single reaction.
[0078] If the adapter contains an identification tag and the adapter is attached to each end of the nucleic acid in the sample, the number of potential identifier combinations increases exponentially with the number of unique tags supplied (i.e., nn combinations, where n is the number of unique identification tags). In some methods, the number of unique tag combinations is statistically sufficient to ensure that all or substantially all (e.g., at least 90%) of the different double-stranded DNA molecules in the sample receive a different tag combination. In some methods, the number of unique identifier tag combinations is less than the number of unique double-stranded DNA molecules in the sample (e.g., 5 to 10,000 different tag combinations).
[0079] The kit that provides enzymes suitable for carrying out the above method is the NEBNext® Ultra® II DNA Library Prep Kit for Illumina®. This kit provides the following reagents:
[0080] NEBNext Ultra II End Prep Enzyme Mix, NEBNext Ultra II End Prep Reaction Buffer, NEBNext Ligation Enhancer, NEBNext Ultra II Ligation Master Mix -20, NEBNext® Ultra II Q5® Master Mix.
[0081] Blurtening and tail addition of sample nucleic acids can be performed in a single tube. It is not necessary to separate the blunt-ended nucleic acid from the enzyme(s) that performed the blunt-ending before the tail addition reaction. If necessary, all enzymes, nucleotides, and other reagents can be supplied together before the blunt-ending reaction. Supplying together means that everything is introduced into the sample within a sufficiently short time, and therefore everything is present when the sample is incubated for blunt-ending. If necessary, nothing should be removed from the sample after supplying the enzymes, nucleotides, and other reagents until at least both the blunt-ending incubation and the tail addition incubation are complete. In many cases, the tail addition reaction is performed at a higher temperature than the blunt-ending reaction. For example, the blunt-end reaction can be carried out at ambient temperature where the 5'-3' polymerase and 3'-5' exonuclease are active and the thermostable polymerase is inactive or minimally active, while the end-tail addition reaction can be carried out at elevated temperature, e.g., above 60°C, where the 5'-3' polymerase and 3'-5' exonuclease are inactive and the thermostable polymerase is active.
[0082] 3. Amplification Sample nucleic acids sandwiched between adapters can be amplified by PCR and other amplification methods, typically by priming them with primers that bind to the primer binding sites of the adapters surrounding the nucleic acids to be amplified. The amplification method may involve thermocycling, extension, denaturation, and annealing cycles, or it may be isothermal, as in transcription-mediated amplification. Other amplification methods include ligase chain reaction, strand displacement amplification, nucleic acid sequence-based amplification, and autologous persistent sequence-based replication.
[0083] In this method, it is preferable that at least 75, 80, 85, 90%, or 95% of the double-stranded nucleic acids in the sample are linked to the adapter. By using T-tail addition and C-tail addition, it is preferable that the percentage of double-stranded nucleic acids in the sample linked to the adapter increases by at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10% compared to a control method performed using a T-tailed adapter alone (an increase in yield from 75% to 80% is considered a 5% increase). By using T-tail addition and C-tail addition in combination with a blunt-ended adapter, it is preferable that the percentage of double-stranded nucleic acids linked to the adapter increases by at least 5, 10, 15, 20, or 25%. The percentage of nucleic acids linked to the adapter can be determined by gel electrophoresis comparing the original sample with the processed sample after linkage with the adapter is complete.
[0084] In this method, it is preferable that sequencing is achieved for at least 75, 80, 85, 90, or 95% of the available double-stranded molecules in the sample. By using T-tail addition and C-tail addition, it is preferable that the percentage of double-stranded nucleic acids to be sequenced in the sample increases by at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10% compared to a control method performed using a T-tailed adapter alone. By using T-tail addition and C-tail addition in combination with a blunt-end adapter, it is preferable that the percentage of double-stranded nucleic acids to be sequenced in the sample increases by at least 5, 10, 15, 20, or 25% compared to a control method performed using a T-tailed adapter alone. The percentage of nucleic acids to be sequenced can be determined by comparing the number of molecules actually sequenced with the number that could have been sequenced based on the input nucleic acid and the genomic region targeted for sequencing.
[0085] 4. Tags Tags providing molecular identifiers or barcodes can be incorporated into adapters by ligation, overlap extension PCR, or other methods, among others. Generally, the assignment of unique or non-unique identifiers or molecular barcodes in a reaction follows the methods and systems described in U.S. Patent Applications Nos. 20010053519, 20030152490, 20110160078, and U.S. Patents Nos. 6,582,908 and 7,537,898.
[0086] The tags can be attached to the sample nucleic acids randomly or non-randomly. In some cases, the tags are introduced in an expected ratio of unique identifiers to microwells. For example, unique identifiers can be loaded so that each genomic sample is loaded with more than approximately 1, more than approximately 2, more than approximately 3, more than approximately 4, more than approximately 5, more than approximately 6, more than approximately 7, more than approximately 8, more than approximately 9, more than approximately 10, more than approximately 20, more than approximately 50, more than approximately 100, more than approximately 500, more than approximately 1000, more than approximately 5000, more than approximately 50000, more than approximately 50000, more than approximately 50000, more than approximately 1,000,000, more than approximately 10,000,000, more than approximately 50,000,000, or more than approximately 1,000,000,000 unique identifiers. In some cases, unique identifiers can be loaded so that each genome sample is loaded with fewer than approximately 2, fewer than approximately 3, fewer than approximately 4, fewer than approximately 5, fewer than approximately 6, fewer than approximately 7, fewer than approximately 8, fewer than approximately 9, fewer than approximately 10, fewer than approximately 20, fewer than approximately 50, fewer than approximately 100, fewer than approximately 500, fewer than approximately 1000, fewer than approximately 5000, fewer than approximately 50000, fewer than approximately 50000, fewer than approximately 500000, fewer than approximately 1,000000, fewer than approximately 10,000000, fewer than approximately 50,000000, or fewer than approximately 1,000,00000.In some cases, the average number of unique identifiers loaded per sample genome is less than approximately 1, less than approximately 2, less than approximately 3, less than approximately 4, less than approximately 5, less than approximately 6, less than approximately 7, less than approximately 8, less than approximately 9, less than approximately 10, less than approximately 20, less than approximately 50, less than approximately 100, less than approximately 500, less than approximately 1000, less than approximately 5000, less than approximately 50000, less than approximately 50000, less than approximately 1,00000, less than approximately 10,000,000, and less than approximately 50,000,000. It is a unique identifier that is either none or less than approximately 1,000,000,000, or more than approximately 1, more than approximately 2, more than approximately 3, more than approximately 4, more than approximately 5, more than approximately 6, more than approximately 7, more than approximately 8, more than approximately 9, more than approximately 10, more than approximately 20, more than approximately 50, more than approximately 100, more than approximately 500, more than approximately 1,000, more than approximately 5,000, more than approximately 10,000, more than approximately 500,000, more than approximately 500,000, more than approximately 1,000,000, more than approximately 10,000,000, more than approximately 50,000,000, or more than approximately 1,000,000,000.
[0087] In some cases, the unique identifier may be an oligonucleotide of a predetermined, random, or semi-random sequence. In other cases, multiple barcodes may be used, and therefore the barcodes are not necessarily unique to one another. In this example, barcodes can be ligated to individual molecules, and thus the combination of a barcode and a sequence that can be ligated to it creates a unique sequence that can be tracked individually. As described herein, by detecting a non-unique barcode in combination with sequence data of the beginning and ending portions of a sequence read, it may be possible to assign a unique identity to a particular molecule. It may also be possible to assign a unique identity to such a molecule using the length or number of base pairs of individual sequence reads. As described herein, a single-stranded nucleic acid fragment assigned to a unique identity may thereby enable the identification of the fragment from the subsequent parent strand.
[0088] Polynucleotides in a sample can be tagged with a sufficient number of different tags, and therefore, it is highly probable that all polynucleotides mapped to a particular genomic region will have different identification tags (the molecules within the region are substantially uniquely tagged) (e.g., at least 90%, at least 95%, at least 98%, at least 99%, at least 99.9%, or at least 99.99%). The genomic region to which polynucleotides are mapped may be, for example, (1) an entire panel of genes to be sequenced, (2) some part of that panel, e.g., mapping to within a single gene, within an exon, or within an intron, (3) a single nucleotide coordinate (e.g., at least one nucleotide in a polynucleotide is mapped to a coordinate, e.g., a start position, a stop position, a midpoint, or any position in between), or (4) a specific pair of start / stop (beginning / end) nucleotide coordinates. The number of different identifiers required to substantially uniquely tag a polynucleotide (tag count) is a function of how many original polynucleotide molecules in the sample are mapped to that region. This, in turn, is a function of several factors. One factor is the total number of haploid genomic equivalents included in the assay. Another factor is the average size of polynucleotide molecules. Another factor is the distribution of molecules across regions, which is now a function of the cleavage pattern. Cleavage occurs primarily between nucleosomes, and therefore it can be expected that more polynucleotides may be mapped across nucleosome locations than between nucleosomes. Another factor is the distribution of barcodes in the pool and the ligation efficiency of individual barcodes, which potentially cause differences in effective concentration between one barcode and another. Another factor is the size of the regions (e.g., the same start / stop or the same exon) to which uniquely tagged molecules are localized.
[0089] An identifier can be a single barcode attached to one end of a molecule, or two barcodes attached to different ends of the molecule. By independently attaching barcodes to both ends of the molecule, the number of possible identifiers increases exponentially. In this case, the number of different barcodes is chosen such that the combination of barcodes at each end of a particular polynucleotide is highly likely to be unique to other polynucleotides mapped to the same selected genomic region.
[0090] In a particular embodiment, the number of different identifier or barcode combinations used (tag count) may be at least 64, 100, 400, 900, 1400, 2500, 5625, 10,000, 14,400, 22,500, or 40,000, and any of 90,000 or less, 40,000 or less, 22,500 or less, 14,400 or less, or 10,000 or less. For example, the number of identifier or barcode combinations may be between 64 and 90,000, between 400 and 22,500, between 400 and 14,400, or between 900 and 14,400.
[0091] In a sample containing fragmented genomic DNA derived from multiple genomes, such as cell-free DNA (cfDNA), there is some possibility that more polynucleotides will share the same start and stop positions than one from a different genome ("replicas" or "cognates"). The number of possible replicas starting at any given position is a function of the number of haploid genomic equivalents and the distribution of fragment sizes in the sample. For example, cfDNA has a fragment peak of about 160 nucleotides, with the majority of fragments within this peak ranging from about 140 to 180 nucleotides. Thus, cfDNA derived from a genome of about 3 billion bases (e.g., the human genome) can consist of nearly 20 million (2 × 10⁷) polynucleotide fragments. A sample of about 30 ng of DNA may contain about 10,000 haploid human genomic equivalents (similarly, a sample of about 100 ng of DNA may contain about 30,000 haploid human genomic equivalents). A sample containing approximately 10,000 (10⁴) haploid genome equivalents of such DNA may have approximately 200 billion (2 × 10¹¹) individual polynucleotide molecules. It has been empirically determined that in a sample of approximately 10,000 haploid genome equivalents of human DNA, there are approximately 3 replicated polynucleotides beginning at any given position. Therefore, such a collection may contain a diversity of approximately 6 × 10¹¹ to 8 × 10¹¹ (approximately 60 billion to 80 billion, e.g., approximately 70 billion (7 × 10¹¹)) differently sequenced polynucleotide molecules.
[0092] The probability of correctly identifying a molecule depends on the initial number of genome equivalents, the distribution of sequenced molecule lengths, sequence uniformity, and the number of tags. This number can be calculated using a Poisson distribution. A tag count equal to 1 is equivalent to having no unique tags or being untagged. Table 1 below lists the probabilities of correctly identifying molecules as unique, assuming a typical cell-free size distribution as described above. Table 1. Probability of correctly identifying molecules [Table 1]
[0093] In this case, it may not be possible to determine which sequence read originates from which parent molecule during genomic DNA sequencing. This problem can be mitigated by tagging parent molecules with a sufficient number of unique identifiers (e.g., tag counts) so that two replicated molecules, i.e., molecules with the same start and stop positions, may have different unique identifiers, and thus sequence reads can be traced back to specific parent molecules. One approach to this problem is to uniquely tag all or nearly all different parent molecules in the sample. However, depending on the number of haploid gene equivalents and the distribution of fragment sizes in the sample, billions of different unique identifiers may be required.
[0094] This method can be cumbersome and costly. In some embodiments, methods and compositions are provided herein for tagging a population of polynucleotides in a sample of fragmented genomic DNA with n distinct unique identifiers, where n is at least 2 and no more than or equal to 100,000*z, and z is a measure (e.g., mean, median, mode) of the central tendency of the expected number of replicated molecules having the same start and stop positions. In certain embodiments, n is at least 2*z, 3*z, 4*z, 5*z, 6*z, 7*z, 8*z, 9*z, 10*z, 11*z, 12*z, 13*z, 14*z, 15*z, 16*z, 17*z, 18*z, 19*z, 20*z, or 100*z (e.g., lower limit). In other embodiments, n is less than or equal to 100,000*z, less than or equal to 10,000*z, less than or equal to 2,000*z, less than or equal to 1,000*z, less than or equal to 500*z, or less than or equal to 100*z (e.g., upper limit). Thus, n can span any combination between these lower and upper limits. In certain embodiments, n is between 100*z and 1,000*z, between 5*z and 15*z, between 8*z and 12*z, or about 10*z. For example, a haploid human genome equivalent has about 3 picograms of DNA. A sample of about 1 microgram of DNA contains about 300,000 haploid human genome equivalents. The number n can be between 15 and 45, between 24 and 36, between 64 and 2,500, between 625 and 31,000, or about 900 and 4,000. As long as at least a portion of the replicated or homogeneous polynucleotides have unique identifiers, i.e., different tags, sequencing improvements can be achieved. However, in certain embodiments, the number of tags used is selected such that there is at least a 95% probability that all replicated molecules starting from any one position will have unique identifiers. For example, a sample containing approximately 10,000 haploid human genome equivalents of cfDNA can be tagged with approximately 36 unique identifiers. Each unique identifier may include six unique DNA barcodes. Attaching these to both ends of a polynucleotide results in 36 possible unique identifiers.Samples tagged in this manner may contain fragmented polynucleotides, such as genomic DNA, such as cfDNA, in amounts ranging from approximately 10 ng to approximately 100 ng, approximately 1 μg, or approximately 10 μg.
[0095] Accordingly, the Disclosure also provides compositions of tagged polynucleotides. Polynucleotides may include fragmented DNA, such as cfDNA. A set of polynucleotides in a composition can be non-specifically tagged to map to mappable base positions in the genome; that is, the number of different identifiers can be at least two and less than the number of polynucleotides mapped to mappable base positions. A composition between about 10 ng and about 10 μg (e.g., any of about 10 ng to 1 μg, about 10 ng to 100 ng, about 100 ng to 10 μg, about 100 ng to 1 μg, or about 1 μg to 10 μg) may have different identifiers between 2, 5, 10, 50 or 100, and any of 100, 1000, 10,000 or 100,000. For example, polynucleotides in such a composition can be tagged using different identifiers between 5 and 100 or between 100 and 4000.
[0096] The event in which different molecules are mapped to the same coordinates (in this case, having the same start / stop position) and have the same tag rather than different tags is called a "molecular collision." In certain cases, the actual number of molecular collisions may be greater than the theoretical number of collisions calculated, for example, as described above. This may be a function of the uneven distribution of molecules across coordinates, differences in ligation efficiency between barcodes, and other factors. In this case, an empirical method can be used to determine the number of barcodes required to approach the theoretical number of collisions. In one embodiment, a method is provided herein for determining the number of barcodes required to reduce barcode collisions for a given haploid genome equivalent, based on the distribution of lengths and sequence uniformity of the sequenced molecules. The method includes the steps of: creating multiple pools of nucleic acid molecules; tagging each pool with an increasing number of barcodes; and determining the optimal number of barcodes to reduce the number of barcode collisions to a theoretical level, which may be due to differences in effective barcode concentration resulting from differences in pools and ligation efficiency, for example.
[0097] In one embodiment, the number of identifiers required to substantially uniquely tag polynucleotides mapped to a given region can be empirically determined. For example, a selected number of different identifiers can be attached to molecules in the sample, and the number of different identifiers for molecules mapped to that region can be counted. If an insufficient number of identifiers are used, several polynucleotides mapped to that region will have the same identifier. In that case, the number of identifiers counted will be less than the number of original molecules in the sample. The number of different identifiers used can be iteratively increased for a given sample type until no further identifiers representing new original molecules are detected. For example, in the first iteration, five different identifiers representing at least five different original molecules may be counted. In the second iteration, more barcodes are used, and seven different identifiers representing at least seven different original molecules are counted. In the third iteration, more barcodes are used, and ten different identifiers representing at least ten different original molecules are counted. In the fourth iteration, more barcodes are used, and ten different identifiers are counted again. At this point, adding more barcodes is unlikely to increase the number of original molecules detected.
[0098] 5. Sequence determination Sample nucleic acids, sandwiched between adapters with or without prior amplification, can be subjected to sequencing. Sequencing methods include, for example, Sanger sequencing, high-throughput sequencing, pyrosequencing, synthesis sequencing, single-molecule sequencing, nanopore sequencing, semiconductor sequencing, ligation sequencing, hybridization sequencing, RNA-Seq (Illumina), Digital Gene Expression (Helicos), next-generation sequencing, Single Molecule Sequencing by Synthesis (SMSS) (Helicos), large-scale parallel sequencing, Clonal Single Molecule Array (Solexa), shotgun sequencing, Ion Torrent, Oxford Nanopore, Roche Genia, Maxim-Gilbert sequencing, primer walking, PacBio, SOLiD, Ion Torrent, or sequencing using the Nanopore platform. The sequencing reaction can be carried out in various sample processing units, which may be multiple lanes, multiple channels, multiple wells, or other means for processing multiple sample sets substantially simultaneously. The sample processing unit may also include multiple sample chambers to enable simultaneous processing of multiple trials.
[0099] Sequencing reactions can be performed on one or more fragment types known to contain markers for cancer or other diseases. Sequencing reactions can also be performed on any nucleic acid fragment present in the sample. Sequencing reactions may result in genomic sequence coverage of at least 5%, 10%, 15%, 20%, 25%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 99%, 99.9%, or 100%. In other cases, genomic sequence coverage may be less than 5%, less than 10%, less than 15%, less than 20%, less than 25%, less than 30%, less than 40%, less than 50%, less than 60%, less than 70%, less than 80%, less than 90%, less than 95%, less than 99%, less than 99.9%, or less than 100%.
[0100] Multiplex sequencing can be used to perform simultaneous sequencing reactions. In some cases, cell-free nucleic acids can be sequenced in at least 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 50,000, and 100,000 sequencing reactions. In other cases, cell-free polynucleotides can be sequenced in fewer than 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 50,000, and 100,000 sequencing reactions. Sequencing reactions may be performed sequentially or simultaneously. Subsequent data analysis can be performed on all or part of the sequencing reactions. In some cases, data analysis can be performed on at least 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 50,000, and 100,000 sequencing reactions. In other cases, data analysis can be performed on fewer than 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 50,000, and 100,000 sequencing reactions.
[0101] The sequencing method may be large-scale parallel sequencing, that is, sequencing at least 100, 1,000, 10,000, 100,000, 1 million, 10 million, 100 million, or 1 billion nucleic acid molecules simultaneously (or in succession).
[0102] 6.Analysis Genetic data can be used to characterize specific forms of cancer. Cancers are often heterogeneous in terms of both composition and staging. Genetic profiling data may enable the characterization of specific subtypes of cancer, which may be important in the diagnosis or treatment of those specific subtypes. This information also provides clues to the prognosis of a particular type of cancer for the patient or practitioner, allowing them to adapt treatment options as the disease progresses. Some cancers progress, become more invasive, and become genetically unstable. Other cancers may remain benign, inactive, or quiescent. The systems and methods of this disclosure may be useful in determining disease progression.
[0103] This method is useful in determining the effectiveness of specific treatment options. It can also be used to detect genetic variations in non-cancerous conditions. Immune cells, such as B cells, can undergo rapid clonal expansion in the presence of certain diseases. Clonal expansion can be monitored using copy number variation detection, and certain immune states can also be monitored. In this example, copy number variation analysis can be performed over time to create a profile of how a particular disease may progress.
[0104] Furthermore, the method of this disclosure can be used to characterize heterogeneity of abnormal conditions in a subject. The method includes the step of generating a gene profile of extracellular polynucleotides in a subject, wherein the gene profile includes multiple data obtained from copy number variation and rare mutation analysis. Diseases, including cancer, can be heterogeneous, in some cases but not limited to these. Disease cells may not be identical. In the case of cancer, it has been found that some tumors contain different types of tumor cells, and some cells are at different stages of cancer. In other cases, heterogeneity may include multiple lesions of the disease. Again, in the case of cancer, multiple tumor lesions may be present, in which case one or more lesions may be the result of metastasis that has spread from the primary site.
[0105] This method can be used to generate or profile a fingerprint or set of data, which is the sum of genetic information derived from different cells in heterogeneous diseases. This set of data may include copy number variation and rare mutation analysis, either alone or in combination.
[0106] This method can be used to diagnose, determine the prognosis of, monitor, or observe cancer or other diseases.
[0107] 7. Kit This disclosure also provides a kit for carrying out any of the methods described above. An exemplary kit includes oligonucleotide probes for targeting and capturing the sequencing panel discussed herein. The kit may also include packaging, leaflets, CDs, etc., providing instructions for carrying out the claimed methods.
[0108] All patent applications, websites, other publications, accession numbers, etc., cited above or below are incorporated by reference in whole for any purpose to the same extent as each individual item is specifically and individually indicated to be incorporated by reference. If different versions of an array are associated with an accession number at different times, the version associated with the accession number on the effective filing date of this application is meant. The effective filing date means the earlier of the actual filing date or, where applicable, the filing date of the priority application referencing the accession number. Similarly, if different versions of a publication, website, etc., are published at different times, unless otherwise specified, the most recently published version on the effective filing date of this application is meant. Unless otherwise indicated, any feature, step, element, embodiment, or aspect of the present invention may be used in combination with any other. Although the present invention has been described in some detail by illustration and example for clarity and understanding, it will become apparent that certain changes and modifications within the scope of the appended claims can be implemented. [Examples]
[0109] Workflow example (Example 1) Improvement of bisulfite sequencing recovery 1) Obtain a cfDNA sample.
[0110] 2) Prepare the ends for Y adapter ligation (T4 DNA ligase): “End repair” using a dNTP mix containing dmCTP to mark the bases synthesized during the process (improving the accuracy of methylation calls), and A tail addition. Commercial ssDNA-LP “sprint adapter ligation” kits, such as Claret Biosciences’ SRSLY, have a step that occurs concurrently with ligation, which phosphorylates the 5' end and removes the 3' phosphate group if necessary, thereby preparing the ends for ligation. This step can be omitted for the purposes of the hybrid workflow described herein, as it is performed during the upstream “end repair” that prepares the dsDNA for Y adapter ligation. In this case, the 5' hydroxyl group and 3' phosphate group are the simplest and most economical modifications to use. However, if the ssDNA-LP T4 PNK step is left in the protocol, these two modifications cannot be used because they become substrates for T4 PNK activity and are converted into ligable ends. In protocols involving the ssDNA-LP PNK step, suitable modifications to the ssDNA end of the Y adapter include a 5'C3 spacer, a 5' inverted dideoxy base, another 5' spacer, a 3'C3 spacer, a 3' inverted dT, a 3' dideoxy base, and another 3' spacer. These are suitable because they interfere with / inhibit T4 PNK activity.
[0111] 3) Ligate the ds-cfDNA molecule with a Y adapter containing methylated cytosine to protect it from downstream bisulfite conversion. The ssDNA ends of Y contain specific modifications (e.g., 5'OH, 3'P) to prevent ligation through these ends. The Y adapter contains a DNA barcode that functions as a molecular barcode and / or identifies the molecule as originating from Y adapter ligation.
[0112] 4) Post-ligation cleanup. This step is optional and can be omitted, potentially improving recovery.
[0113] 5) Denaturate the DNA and perform bisulfite treatment.
[0114] 6) ssDNA library preparation is performed using two types of sprint adapters: (1) a double-stranded DNA hybrid with an SP5 / read 1-primer sequence and a randommer sequence overhang at the 5' end relative to the SP5 reverse complement, and (2) a double-stranded hybrid with an SP7 / read 2-primer sequence and a randommer sequence overhang at the 3' end relative to the SP7 / read 2-primer sequence. Both adapters have modifications (e.g., 5'OH, 3'P) at the double-stranded blunt ends and therefore cannot be ligated at these ends.
[0115] 7) Denaturate the DNA to ssDNA and apply single-strand binding protein (SSB).
[0116] 8) Ligate ssDNA that does not have a Y-adapter sequence at its terminus with a sprint adapter. These are DNA fragment ends that were lost in the first ligation or created by bisulfite-induced fragmentation. Note that this ssDNA LP method using a sprint adapter deviates from published protocols / kits that have a phosphorylation step before / concurrent with ligation. This modification interferes with recovery in standard ssDNA LP, but is essential, and recovery is improved in this method by blocking the Y-adapter with a 5'OH modification. Alternatively, as described above, alternative Y-adapter modifications can be used to block ligation of these ends that are less susceptible to downstream phosphorylation.
[0117] 9) Post-ligation cleanup. This step is optional and can be omitted, potentially improving recovery.
[0118] 10) Amplify all DNA library molecules that contain either a Y adapter sequence or a sprint adapter sequence at both ends. Enrich the target genomic region for sequencing by oligo-hybrid capture as needed.
[0119] 11) Sequence all target DNA library molecules using an NGS platform.
[0120] 12) Analyze sequencing reads to identify the molecular topology of origin and degrade unique molecules from PCR replicas: (1) "Full-length" DNA molecules that originate from dsDNA and have Y-adapter barcodes at both molecular ends (the barcodes may be molecular barcodes or simply Y-adapter ligations that identify the barcodes) and have not been fragmented by bisulfite in the process. These types of molecules are degraded by a combination of barcodes on the molecular ends and genomic coordinates at the ends. (2) Partially bisulfite-fragmented DNA molecules that originate from dsDNA and have Y-adapter barcodes at only one end and have been fragmented by bisulfite. These types of molecules are degraded by a combination of barcodes on one end and genomic coordinates at the ends. Note that bisulfite fragmentation resulting in one end should increase the diversity of fragment ends, which increases the resolution when using genomic coordinates. (3) A DNA molecule that originates from dsDNA and is fragmented by bisulfite at both ends, or from ssDNA (in trace amounts) that is identified by the absence of molecular barcodes at both ends, and is completely fragmented by bisulfite. The molecule is degraded by the genomic coordinates at the ends, which will have greater diversity due to further fragmentation.
[0121] From multiple NGS reads of the same molecule, a consensus sequence is generated that increases the accuracy of the base, including the methylation state.
[0122] For "full-length" molecules where both the top and bottom strands have NGS reads, further consensus can be generated to explain dsDNA strand symmetry (in methylation events and somatic mutations).
[0123] (Example 2) Ligation using only double-stranded DNA barcodes 1) Perform dsDNA end repair and A tail addition (if necessary).
[0124] 2) Ligate the dsDNA molecule with a T-tailed (optional) dsDNA barcode tag containing a series of non-complementary 5' and 3' overhang sequences (similar to the "Y" in the Y adapter) to prevent the barcodes from ligating each other / self-ligating. Note that, unlike other methods, the tag ends should be ligable in downstream steps. Thus, the barcode tag may contain 5'P or 3'OH at all ends, not just the desired ligation junction side (the T-tail side of the barcode). Clean up the first ligation as needed.
[0125] 3) The ssDNA is subjected to the ssDNA LP method using the sprint adapter ligation method described above, but with an important difference: a phosphorylation step can be included before / simultaneously with the (second) ligation step, thereby potentially improving recovery.
[0126] 4) Amplify all DNA library molecules containing sprint adapter sequences at both ends. If necessary, enrich the target genomic region for sequencing by oligohybrid capture.
[0127] 5) Sequence all target DNA library molecules using an NGS platform.
[0128] 6) Analyze sequencing reads to identify the molecular topology of origin and degrade unique molecules from PCR replicas: (2) Molecules originating from dsDNA - contain barcodes / tags at both ends of the molecule. The molecule can be identified using the combination of DNA barcodes and the genomic coordinates of the ends. (2) Molecules originating from ssDNA - do not contain barcodes / tags at either end of the molecule. The molecule can be identified using the genomic coordinates of the ends. Note: If a molecule has a barcode / tag at only one end, the dsDNA molecule is likely to undergo partial ligation due to inefficient end repair, A tail addition, and first ligation.
[0129] (Example 3) Hairpin dsDNA-Ligation 1) Perform dsDNA end repair and A tail addition (if necessary).
[0130] 2) Ligate the dsDNA molecule with a T-tailed (optional) hairpin NGS adapter (such as those used in the NEBNext kit) that has been modified to have a molecular barcode. Clean up the first ligation as needed.
[0131] 3) The ssDNA is subjected to the ssDNA LP method using the sprint adapter ligation method described above, but with an important difference: a phosphorylation step can be included before / simultaneously with the (second) ligation step, thereby potentially improving recovery.
[0132] 4) The sample is subjected to "User treatment" (NEB; uracil DNA glycosylase, followed by endonuclease III) to cleave the hairpin-adapter-attached molecules into linear monomer library molecules that are easily amplified.
[0133] 5) Amplify all DNA library molecules containing sprint adapter sequences at both ends. If necessary, enrich the target genomic region for sequencing by oligo-hybrid capture.
[0134] 6) Sequence all target DNA library molecules using an NGS platform.
[0135] 7) Identify the molecular topology of origin by sequencing analysis and degrade unique molecules from PCR replicas: (1) Molecules originating from dsDNA - contain barcodes / tags at both ends of the molecule. The molecule can be identified using the combination of DNA barcodes and the genomic coordinates of the ends. (2) Molecules originating from ssDNA - do not contain barcodes / tags at either end of the molecule. The molecule can be identified using the genomic coordinates of the ends. Note: If a molecule has a barcode / tag at only one end, the dsDNA molecule is likely to undergo partial ligation due to inefficient end repair, A tail addition, and first ligation.
[0136] (Example 4) Use of biotinylated Y adapters for isolating ssDNA material 1) Perform dsDNA end repair and A tail addition (if necessary).
[0137] 2) Ligate the dsDNA molecule with a T-tailed (if necessary) Y-adapter having a biotin-modified (e.g., oligo-modified base) molecular barcode. Clean up the first ligation as needed.
[0138] 3) The sample is subjected to separation based on streptavidin magnetic beads, and the supernatant containing ssDNA molecules is taken out and placed in a second reaction vessel. Modifications to the workflow as needed include: (1) MSRE treatment applied before streptavidin-based separation. Any double-digested DNA fragments are retained in the ssDNA reaction vessel and are potentially converted into a library. However, these do not contain molecular barcodes (and are treated as fully methylated molecules) and are potentially identifiable by fragment end coordinates that match the MSRE sites, so they do not contribute to noise for methylation measurements. Furthermore, if separate enrichment is performed on the ssDNA and post-MSRE dsDNA libraries, the ssDNA enrichment panel can be designed so that MSRE fragment information is avoided if it has no informational value. (2) The dsDNA library is released from immobilization before MSRE treatment. This can be done by using a photocleavable biotin or desbiotin (tracked by biotin) or other cleavable types of biotin bases (i.e., biotin-UTP, and applying "user processing") in the adapter.
[0139] 4) The first reaction vessel containing the dsDNA library molecules immobilized on beads is subjected to MSRE (methylation-sensitive restriction enzyme) treatment to digest the library molecules containing unmethylated MSRE sites.
[0140] 5) The second reaction vessel is subjected to ssDNA-LP to prepare the contained ssDNA molecules for NGS. The first and second reaction vessels can be amplified separately or jointly (using streptavidin beads). If necessary, enrich the target genome region for sequencing by oligohybrid capture. Note that separate enrichment reactions and regions for enrichment can be selected for each library if the ssDNA library and the post-MSRE dsDNA library are kept separately in the previous step.
[0141] 6) Sequence all target DNA library molecules using an NGS platform.
[0142] 7) Analyze the sequencing reads to break down two types of molecules: (1) molecules originating from "fully methylated" double-stranded DNA, due to the presence of molecular barcodes at both ends of the molecule, and (2) molecules originating from single-stranded DNA (ssDNA), due to the absence of molecular barcodes at both ends of the molecule.
[0143] Throughout this specification, many variations of the method, particularly with respect to the sprint adapter method, are indicated with an asterisk. Different methods exist for protecting the first ligation (Y adapter) from re-ligation in the sprint adapter step. Furthermore, in the subsection “Loss of ssDNA Molecules and DNA Molecular Topological Information (whether of ssDNA or dsDNA Origin) in Gene Sequence Workflows,” it is noted that in the first “dsDNA” ligation step, the Y adapter can be used to generate a fully formed library molecule at this stage, or simply to ligate a dsDNA identification tag / barcode (without any end protection), and an NGS adapter can be further added to these molecules during ssDNA LP. The former may be advantageous in terms of higher molecular recovery (of dsDNA molecules), while the latter may be preferable due to the use of simpler / less expensive reagents. In the latter case, the barcode identifying the dsDNA may further be a molecular barcode. Furthermore, the first ligation can be performed using a hairpin adapter (with molecular / identification barcode) that does not require adapter end protection, which simplifies several aspects of the workflow.
[0144] It is also possible to use a different ssDNA LP method instead of the "sprint adapter" method which is the focus. Swift / IDT has an ssDNA LP method in which, in the first adapter ligation step, the 3' end of the ssDNA molecule is extended with terminal transferase polymerase and then / simultaneously ligated with a sprint adapter having an overhang complementary to the extended template. Next, a primer matching the 3' adapter end is extended with polymerase to generate a blunt-ended or A-tailed dsDNA molecule again. Then, a second ligation is performed using a blunt-ended or T-tailed adapter. This method is interchangeable with the "sprint adapter" method in "Improving ssDNA recovery using a 'hybrid' LP workflow to obtain molecular topology information (ssDNA and dsDNA)". This can also be used to improve bisulfite-Seq molecule recovery in the method of the present disclosure, but it is only possible to recover half of the molecule that is "partially fragmented with bisulfite," i.e., only the molecule that has a 3' end exposed to bisulfite.
[0145] The method disclosed herein can be applied to single-molecule sequencing, thereby obtaining single-site methylation information with added molecular topology information.
Claims
1. A method for preparing a sequencing library from DNA molecules in a sample, (a) A step of providing a first population of DNA molecules derived from the sample, wherein the first population of DNA molecules includes double-stranded DNA and single-stranded DNA; (b) Ligating a first set of adapters, each having a molecular barcode configured to attach to a plurality of the double-stranded DNA molecules, to generate a second population comprising a plurality of double-stranded DNA molecules to which the adapters are ligated, and a plurality of DNA molecules that are not ligated, wherein the adapters are ligated to both or one end of the double-stranded DNA molecules; (c) A step of subjecting the second population to a process that denatures and fragments a plurality of DNA molecules with adapters ligated and a plurality of DNA molecules without adapters to generate a third population of DNA molecules, including single-stranded DNA molecules having adapters at both ends, having an adapter at one end, and / or having no adapters at either end, and fragmented single-stranded DNA molecules, wherein the fragmented single-stranded DNA molecules include fragments having one adapter ligated at one end and / or not having an adapter ligated; (d) Ligate a second set of adapters to a subset of molecules in the third population that are either ligated with adapters to the fragment or not ligated with adapters, thereby (i) A single-stranded DNA to which adapters have been ligated, including adapters from a first set of adapters, which have been ligated to both ends of the molecule. (ii) A single-stranded DNA to which adapters are ligated, comprising one adapter from a first set of adapters ligated to one end of the molecule and one adapter from a second set of adapters ligated to the other end of the molecule, and (iii) Adapter-ligated single-stranded DNA, including adapters from a second set of adapters ligated to both ends of the molecule. A step of generating a tagged DNA molecule containing at least two of the following: Includes, A method for providing a sequencing library derived from a population of DNA molecules in the aforementioned sample.
2. The method according to claim 1, wherein the first group of DNA molecules includes double-stranded and single-stranded cell-free DNA (cfDNA).
3. The method according to claims 1 to 2, wherein the first set of the adapters is a Y-shaped adapter.
4. The method according to claims 1 to 3, wherein the first set of adapters is protected from the processing in (c).
5. The method according to any one of claims 1 to 4, wherein the first set of adapters further comprises single-stranded ends protected from ligation using modifications including 5'OH and / or 3'P.
6. The method according to claims 1 to 4, wherein the first set of the adapter includes single-stranded ends that are protected from ligation without using modifications including 5'OH and / or 3'P when T4 PNK is used in (d).
7. The method according to claims 1 to 4, wherein the first set of the adapter further comprises single-stranded ends protected from ligation using modifications including a 5'C3 spacer, a 5' inverted dideoxy base, another 5' spacer, a 3'C3 spacer, a 3' inverted dT, a 3' dideoxy base, and another 3' spacer, when T4 PNK is used in (d).
8. The method according to any one of claims 1 to 7, wherein the first set of the adapters includes a universal amplification array.
9. The method according to any one of claims 1 to 8, wherein the molecular barcode distinguishes the molecule ligated in (b) from the molecule ligated in (d).
10. The method according to any one of claims 1 to 4 or 6 to 9, wherein the population of DNA molecules is phosphorylated using T4 PNK prior to ligation in (b).
11. The method according to any one of claims 1 to 4 or 6 to 10, wherein the population of DNA molecules is phosphorylated using T4 PNK prior to ligation in (d).
12. The method according to any one of claims 1 to 11, wherein the process of denaturing and fragmenting the second group of DNA molecules comprises at least one of bisulfite conversion, Tet-assisted bisulfite conversion, and Tet-assisted conversion using a substituted borane reducing agent, wherein the substituted borane reducing agent is optionally 2-picoline borane, borampyridine, tert-butylamine borane, or ammonia borane.
13. The method according to any one of claims 1 to 12, wherein the process of denaturing and fragmenting the second group of DNA molecules includes a chemically assisted transformation using a substituted borane reducing agent, wherein the substituted borane reducing agent is optionally 2-picoline borane, borampyridine, tert-butylamine borane, or ammonia borane.
14. The method according to any one of claims 1 to 13, wherein the first set of adapters comprises methylated cytosine for protecting them from the processing.
15. The method according to any one of claims 1 to 14, wherein the second set of adapters is a “sprint” adapter containing a 5' or 3' overhang.
16. The method according to any one of claims 1 to 15, wherein the second set of the adapters includes a universal amplification array.
17. The method according to any one of claims 1 to 16, wherein the second set of adapters includes (i) an adapter having a double-stranded portion and a single-stranded overhang having a randommer sequence on the 5' side of the reversed strand, and (ii) an adapter having a double-stranded portion and a single-stranded overhang having a randommer sequence on the 3' side of the reversed strand.
18. The method according to any one of claims 1 to 17, wherein the second set of adapters selectively tags only DNA molecular ends that lack the first adapter sequence.
19. The method according to any one of the claims, further comprising the step of amplifying the molecules in (i) to (iii) to generate a replicated DNA molecule.
20. The method according to claim 19, further comprising the step of sequencing the amplified DNA molecule to generate sequencing reads.
21. The method according to claim 20, wherein the amplified molecules are captured and enriched with one or more target regions before sequencing.
22. The method according to claim 21, wherein a specific molecule is degraded from the PCR replica by analyzing the sequencing reads to identify molecules containing the same terminal coordinates and / or molecular barcode.
23. The method according to any one of claims 1 to 22, wherein a DNA molecule having the first adapter at one end and the second adapter at the other end has greater end coordinate diversity than the molecule of (d)(i), and the diversity of the fragment ends improves the accuracy of degrading specific molecules from PCR replicas.
24. The method according to any one of claims 1 to 23, wherein the DNA molecule having the second adapter at both ends of (d) and (iii) has greater end coordinate diversity than the molecules of (d)(i) and (ii), and the diversity improves the accuracy of degrading specific molecules from PCR replicas.
25. The method according to any one of claims 1 to 24, wherein the diversity of terminal coordinates improves the accuracy of degrading specific molecules from PCR replicas.
26. The method according to any one of claims 1 to 25, wherein the terminal coordinates and molecular barcode improve the accuracy of degrading specific molecules from PCR replicas.
27. The method according to any one of claims 1 to 26, wherein the symmetry between the forward and reverse chains improves the accuracy of detecting the methylation state.
28. The method according to any one of claims 1 to 27, wherein the symmetry between the forward strand and the reverse strand improves the accuracy of detecting somatic mutations.
29. The method according to any one of the claims, for a DNA molecule having the first adapter at both ends of (d) and (i), generating a consensus that describes the symmetry of the forward and reverse strands in order to determine the methylation state and / or somatic mutation.
30. The method according to any one of the claims, wherein, prior to (b), the DNA molecules in the first population are end-repaired and / or A-tailed.
31. A method for analyzing a collection of DNA molecules in a sample, (a) A step of ligating a first set of adapters containing molecular barcodes to at least a subset of DNA molecules within the population of DNA molecules, The steps include: ligating the adapter to both ends of the DNA molecule to generate a DNA molecule with the adapter ligated; (b) The step of subjecting the group of DNA molecules to a biochemical treatment that denatures and randomly fragments a plurality of the DNA molecules, thereby producing fragmented DNA molecules having two or fewer first adapters at their ends; (c) Ligating a second set of adapters to the DNA molecules in the population having fewer than two first adapter sequences at their ends, (i) A DNA molecule having the first adapter at both ends, (ii) A DNA molecule having the first adapter at one end and the second adapter at the other end, (iii) A DNA molecule having the second adapter at both ends. The steps include: generating a collection of tagged DNA molecules that include at least two of the following; (d) The step of sequencing the group of tagged DNA molecules from (c) to generate sequencing reads; (e) The step of analyzing the sequence determination reads to detect parallel signals associated with the disease. Methods that include...
32. The method according to claim 31, wherein the signal associated with the disease further comprises one or more of the ratio of short cfDNA fragments to long cfDNA fragments, differences in chromatin architecture, nucleosome structure, epigenetics, and / or genetic information.
33. A method for preparing a sequencing library from a collection of DNA molecules in a sample, (a) a step of ligating a pair of molecular barcodes, each having a molecular barcode at each end, to a subset of DNA molecules in the population, the pair further comprising (i) a first set of molecular barcodes including a ligable end and a 3' single-stranded overhang at the opposite end, and (ii) a second set of molecular barcodes including a ligable end and a 5' single-stranded overhang at the opposite end; (b) Ligating a set of adapters to a plurality of the tagged DNA molecules and a plurality of untagged DNA molecules that were not ligated in (a) to produce (i) adapter-tagged molecules containing a molecular barcode and (ii) adapter-tagged molecules that do not contain the barcode. Includes, A method for providing a sequencing library derived from a population of DNA molecules in the sample.
34. The method according to claim 33, wherein the 5' and 3' overhang arrangements prevent the pairs of barcodes from ligating to each other or self-ligating.
35. The method according to claim 33 or 34, wherein the ligable end includes a T-tail or an A-tail.
36. The method according to any one of claims 33 to 35, wherein the ligable end includes a blunt end.
37. The method according to any one of claims 33 to 36, wherein the barcode further comprises an adapter array.
38. The method according to any one of claims 33 to 37, wherein the adapter attaches to the 3' and 5' single-stranded overhangs of the tagged DNA molecule.
39. The method according to any one of claims 33 to 38, wherein the set of adapters includes (i) an adapter having a double-stranded portion and a single-stranded overhang having a randommer sequence on the 5' side of the reversed strand, and (ii) an adapter having a double-stranded portion and a single-stranded overhang having a randommer sequence on the 3' side of the reversed strand.
40. The method according to any one of claims 33 to 39, wherein the adapter is a “sprint” adapter containing a 5' or 3' overhang.
41. The method according to any one of claims 33 to 40, wherein the adapter includes a universal amplification array.
42. The method according to any one of claims 33 to 41, further comprising the step of amplifying the sequencing library to generate amplified molecules, (i) adapter-tagged molecules including a molecular barcode and (ii) adapter-tagged molecules not including the barcode.
43. The method according to any one of claim 33 or 42, further comprising the step of selectively enriching the library to isolate a subset of the molecules of b(i)(ii).
44. The method according to any one of claims 33 to 43, further comprising the step of sequencing the library of molecules b(i)(ii) to generate sequencing reads.
45. The method according to claim 44, wherein the selective enrichment is carried out by hybridization or amplification techniques.
46. The method according to any one of claims 43 to 45, wherein the enriched subset of the molecules of b(i)(ii) is associated with a disease.
47. The method according to claim 46, wherein the disease is cancer, Alzheimer's disease, hypertriglyceridemia, or coronary artery disease.
48. A method for preparing a sequencing library from a collection of DNA molecules in a sample, (a) Ligating a first set of adapters including capture labels to at least a subset of DNA molecules to generate a first subset of DNA molecules to which the adapters are ligated, the adapters including capture labels, wherein the adapters are ligated to both ends of the DNA molecules; (b) Separating the DNA molecules ligated with the adapter, including the capture label, by contacting the collection of DNA molecules with the capture molecule, to produce (i) a first subset of the DNA molecules ligated with the adapter, including the capture label, wherein the label is bound to the capture molecule; and (ii) DNA molecules that have not been adapter-attached; (c) A step of ligating a second set of adapters with at least a subset of DNA molecules that have not been adapter-attached to generate a second subset of DNA molecules to which the adapters have been ligated, wherein the adapters are ligated to both ends of the DNA molecules. Includes, A method for providing a sequencing library derived from a population of DNA molecules in the sample.
49. The method according to claim 48, wherein the first set of adapters further includes a y-shaped adapter.
50. The method according to any one of claims 48 to 49, wherein the first set of adapters further includes a molecular barcode.
51. The method according to any one of claims 48 to 50, wherein the capture label comprises an affinity ligand.
52. The method according to any one of claims 51, wherein the affinity ligand is biotin or photocleavable biotin.
53. The method according to claim 52, wherein the photocleavable biotin is biotin-UTP.
54. The method according to any one of claims 48 to 53, wherein the separation in (b) further comprises affinity purification using at least one capture molecule.
55. The method according to any one of claims 48 to 54, wherein the capture molecule is streptavidin or a magnetic bead coated with streptavidin.
56. The method according to any one of claims 48 to 55, wherein the adapter is subjected to a process to digest a first set of ligated molecules into unmethylated DNA.
57. The method according to any one of claims 48 to 56, wherein the second set of adapters is a “sprint” adapter containing a 5' or 3' overhang.
58. The method according to claims 48 to 57, wherein the second set of the adapters includes a universal amplification array.
59. The method according to claims 48 to 58, wherein the second set of adapters includes (i) an adapter having a double-stranded portion and a single-stranded overhang having a randommer sequence on the 5' side of the reversed strand, and (ii) an adapter having a double-stranded portion and a single-stranded overhang having a randommer sequence on the 3' side of the reversed strand.
60. The method according to any one of claims 48 to 59, further comprising the step of sequencing the library to generate sequencing reads.
61. The method according to claim 60, further comprising the step of amplifying the library before sequencing.
62. The method according to claim 61, further comprising the step of selectively enriching the library before sequencing.
63. The method according to any one of claims 60 to 62, further comprising the step of analyzing the sequencing reads to determine the methylation status at one or more gene loci.
64. The method according to any one of claims 60 to 62, further comprising the step of analyzing the sequencing reads to determine the molecular topology of a second subset of DNA molecules ligated by the adapter in (c) (e.g., molecules originating from double strands and molecules originating from single strands).
65. The method according to any one of claims 60 to 64, further comprising the step of analyzing the sequencing reads to detect parallel signals associated with a disease.
66. The method according to any one of claims 61 to 65, wherein the disease includes cancer, Alzheimer's disease, hypertriglyceridemia, or coronary artery disease.