Single-tube bead-based DNA co-barcoding for accurate and economical efficient sequencing, haplotyping and assembly
By inserting hybridization sequences on beads and using transposons to integrate unique barcodes into long genomic DNA molecule subfragments, combined with 3' branch connection and PCR amplification, the problem of lack of continuous block information transmission in whole genome sequences in existing technologies is solved, and efficient and economical whole genome sequence analysis is achieved.
Patent Information
- Application Number
- CN202510210888.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2018-06-19
- Filing Date
- 2019-05-07
- Publication Date
- 2025-09-12
AI Technical Summary
Existing technologies lack sequential information about single-base to multi-base variations transmitted in continuous blocks on homologous chromosomes in achieving whole-genome sequences. In addition, most methods are costly, have low data quality, and are difficult to provide unique common barcodes, which limits their use.
The single-tube long fragment read (stLFR) technology is used to insert hybridization sequences on beads and use transposons to integrate unique barcode sequences into long genomic DNA molecule subfragments, combined with 3' branch ligation and PCR amplification to achieve efficient co-barcoded sequencing.
It enables efficient and economical generation of whole genome sequence information, reduces costs, improves data quality and variant calling accuracy, simplifies experimental processes, and is suitable for standard equipment in almost all molecular biology laboratories.
Smart Images

Figure CN120624607A_ABST
Abstract
Description
This application is a divisional application of the invention patent application with application number 201980046212.4, application date May 7, 2019, applicants are Shenzhen BGI Intelligent Manufacturing Technology Co., Ltd. and Shenzhen BGI Life Sciences Institute, and the invention name is "Single-tube bead-based DNA co-barcoding for accurate and cost-effective sequencing, haplotype typing and assembly". Citations of previous applications
[0001] This application claims the benefit of priority to U.S. Provisional Patent Application No. 62 / 668,757, filed May 8, 2018; U.S. Provisional Patent Application No. 62 / 672,501, filed May 16, 2018; and U.S. Provisional Patent Application No. 62 / 687,159, filed June 19, 2018. The above priority applications are incorporated herein by reference in their entirety for all purposes. Background Art
[0002] To date, the vast majority of individual whole genome sequences lack information about the order of single-base to multi-base variants that are transmitted as contiguous blocks on homologous chromosomes. A number of technologies have recently been developed to achieve this. Most of these technologies are based on co-barcoding methods (13), that is, the same barcode is added to subfragments of a single long genomic DNA molecule. After sequencing, the barcode information can be used to determine which reads are derived from the original long DNA molecule. This method was originally described by Drmanac (14) and implemented as a 384-well plate assay by Peters et al. (6). However, these methods are technically challenging to implement, expensive, have low data quality, do not provide unique co-barcodes, or some combination of these four. In practice, most of these methods require the generation of a separate whole genome sequence by standard methods to improve variant calling. This has limited the use of these methods, as cost and ease of use are the main factors affecting the choice of technology to use for WGS. Summary of the Invention Figures and Tables
[0003] Figures 1(A) to 1(D). Overview of stLFR. Figure 1(A): The first step of stLFR involves inserting a hybridization sequence approximately every 200-1000 base pairs into a long genomic DNA molecule. This is achieved using a transposon. The transposon-integrated DNA is then mixed with beads, each containing 400,000 copies of an adapter sequence containing a unique barcode shared by all adapters on the bead, a common PCR primer site, and a common capture sequence complementary to the sequence on the integrated transposon. After genomic DNA is captured to the beads, the transposon is ligated to the barcoded adapters. There are several additional library processing steps, and then the co-barcoded subfragments are sequenced on a BGISEQ-500 or equivalent sequencer. Figure 1(B): Mapping read data by barcode results in read clustering analysis within a 10 to 350 kb region of the genome. Total and barcode coverage are shown for the four barcodes in a small region on Chr11 for 1 ng of the stLFR-1 library. Most barcodes are associated with only one read cluster in the genome. Figure 1(C): The number of original long DNA fragments per barcode is plotted for 1 ng of libraries stLFR-1 and stLFR-2 (orange) and 10 ng of stLFR libraries stLFR-3 and stLFR-4. More than 80% of the fragments in 1 ng of stLFR library were co-barcoded by a single unique barcode. Figure 1(D): The fraction of non-overlapping sequence reads covering each original long DNA fragment and captured subfragments (orange) are plotted for 1 ng of stLFR-1 library. See also Figure 14 .
[0004] Figures 2(A) to 2(D). SV detection. The previously reported deletion in NA12878 was also discovered using stLFR data. A heat map of barcode sharing for each deletion can be seen in Figure 10. Figure 2(A) plots the barcode sharing on chromosome 8 using the previously described Jaccard index (12). Heatmap of shared barcodes within a 2 kb window of a 150 kb heterozygous deletion region. Regions of high overlap are shown in dark red. Those without overlap are shown in beige. Arrows indicate how regions that are spatially distant from each other on chromosome 8 increase overlap, marking the location of the deletion. Figure 2(B) Co-barcoded reads are separated by haplotype and plotted by unique barcode on the y-axis and by chromosome 8 position on the x-axis. Heterozygous deletions are found in a single haplotype. Figure 2(C) also plots heatmaps for overlapping barcodes between chromosomes 5 and 12 for a patient cell line with a known translocation (26); and Figure 2(D) plots heatmaps for GM20759, a cell line with a known transversion on chromosome 2 (27).
[0005] Figures 3(A) and 3(B): Coverage distribution plots. Coverage is plotted for the stLFR-2 (A) and standard (B) libraries sequenced on the BGISEQ500. Coverage for both samples was reduced to 30x. The Poisson distribution for the 30x genome is plotted in blue.
[0006] Figures 4(A) and 4(B): FP overlap between libraries. (A) The FPs for each stLFR library, the BGISEQ-500 standard library, and the PCR-free library sequenced by Illumina (library "HiSeq2500-TruSeq_PCR-Free_DNA_2x251_NA12878" downloaded from basespace) were plotted in a Venn diagram. 2078 FPs were shared between the four stLFR libraries. (B) Overlap of the stLFR library FPs and the Chromium library FPs shows that 1194 FPs were shared between the two different technologies, both of which used DNA isolated from GM12878 rather than the GIAB reference material for NA12878. 884 FPs were unique to the stLFR libraries.
[0007] Figures 5(A) to 5(D): stLFR-1 variant metrics. Read depth and barcode depth were analyzed for the reference allele and variant allele for all true positive variants, false positive variants, and shared false positive variants (green). Read depth for the reference allele (A) and alternative allele (B), as well as barcode counts for the reference allele (C) and alternative allele (D) are plotted. In general, shared false positives look more like true positives, indicating that there is some filtering criteria that can distinguish these variants from unshared false positives.
[0008] Figures 6(A) to 6(D): stLFR-3 variant metrics. Read depth and barcode depth were analyzed for reference and variant alleles for all true positive variants, false positive variants, and shared false positive variants (green). Read depth for the reference allele (A) and alternative allele (B), as well as barcode counts for the reference allele (C) and alternative allele (D), are plotted. In general, shared false positives look more like true positives, suggesting that some filtering criteria exist to distinguish these variants from unshared false positives.
[0009] Figures 7(A) and 7(B): Distribution of shared false-positive variants. The genomic distances separating the 2078 shared FP variants were summed in consecutive bins of 100 bp (dark blue), 1000 bp (orange), 10,000 bp, 100,000 bp, and 1,000,000 bp. Five sets of 2078 randomly selected variants from the stLFR-1 library are also plotted. For each sample, the total number of positions or total number of variants is plotted. Only variants within a bin or bins where 2 or more variants were found were summed. (A) Before filtering, 219 shared FPs appeared to be tightly clustered and were most likely the result of mapping errors. The remaining 1,859 variants appeared to have a distribution similar to that of a random set of variants. (B) After filtering, 1,738 shared FPs remained, but only 72 were tightly clustered.
[0010] Figures 8(A) to 8(T): NA12878 deletion detection using barcode sharing heatmaps. Deletions in the stLFR-1 library were detected using the following data at the following positions: chr3:65189000-65213999 using 230 Gb (A) or 100 Gb (B) of reads, chr4:116167000-116176999 using 230 Gb (E) or 100 Gb (F) of reads, chr4:187094000-187097999 using 230 Gb (G) or 100 Gb (H) of reads, chr7:110182000-110187999 using 230 Gb (I) or 100 Gb (J) of reads. At chr16:62545000-62549999, 230Gb(K) or 100Gb(L) of read segment data is used at chr1:189704509-189783359, 230Gb(M) or 100Gb(N) of read segment data is used at chr3:162512134-162569235, 230Gb(O) or 100Gb(P) of read segment data is used at chr5:104432113-104467893, at chr6:78967194-79001807, and 230Gb(S) or 100Gb(T) of read segment data is used at chr8:39232074-39309652.
[0011] Figures 9(A) to 9(L): Translocation and inversion detection using stLFR. Patient cell lines and cell line GM20759, carrying a translocation between chromosomes 5 and chromosome 12 and an inversion on chromosome 2, respectively, were analyzed with stLFR. For each library, the total sequence coverage was downsampled to investigate the detection capability at lower coverage. The translocation between chromosomes 12 and chromosome 5 was readily detected at total sequence coverage of 40 Gb (A), 20 Gb (B), 10 Gb (C), and even 5 Gb (D). The inversion of GM20759 was also readily detected at total sequence coverage of 46 Gb (E), 20 Gb (F), 10 Gb (G), and 5 Gb (H). In addition, we investigated these regions in the GM12878 cell line, which is not known to contain any of these SVs. No translocations between chromosomes 5 and 12 were evident from either 1 ng of the stLFR library with 230 Gb coverage (I) or 10 ng of the stLFR library with 126 Gb coverage (J). No transversions were found in the stLFR-1 library (K) or the stLFR-4 library (L).
[0012] Figures 10(A) to 10(C): Alignment dot plots of NA12878 scaffolds. SALSA scaffolds from the stLFR-1 (A) and stLFR-4 (B) libraries were mapped against the hg37 reference human genome. 734 million HiC reads obtained from Dixon et al. (29) were also used to generate scaffolds and mapped against hg37 (C). In all cases, only scaffolds covering 5% or more of the chromosome were mapped.
[0013] Figure 11 : LongHap phasing. A complete description of the phasing algorithm applied with LongHap can be found in the "Methods and Materials" section.
[0014] Figure 12 : Barcode sequence assembly. Three ligations were required to generate approximately 3.6 billion different barcodes. The expected sequences for each step of barcode assembly are shown as SEQ ID NOs: 1-13.
[0015] Figure 13 : Barcode sequence assembly flow chart.
[0016] Figure 14 : An exemplary flow chart of a barcoding experiment plan.
[0017] Figure 15 : Animation of the hybridization steps.
[0018] Figure 16 : Animation of the ligation and degradation steps. The final steps showing denaturation and C-tailing are optional and are not described here.
[0019] Figure 17 : Production and use of double-chain DNBs.
[0020] Figure 18 : Amplification of long molecules.
[0021] Figure 19 : Random nickase method
[0022] Figure 20 : Hairpin adapter method.
[0023] Figures 21(A) and 21(B): Figure 21(A) is a schematic diagram of a ligation assay on different DNA substrates. The blunt-ended DNA donor is a synthetic partial dsDNA molecule with a dideoxy 3'-end (solid circle) to prevent the adapter from self-ligating. The long arm of the adapter has been 5'-phosphorylated. The DNA acceptor is assembled using 2 or 3 oligonucleotides (black, red and orange lines) to form a nick (no phosphate), a gap (1 or 8 nt) or a 36 nt 5'-overhang. All strands of the substrate are unphosphorylated and the scaffold strand is 3' dideoxy protected. Figure 21(B) shows the size migration of the ligation products analyzed using a 6% denaturing polyacrylamide gel. Negative no-ligase controls (lanes 1, 3, 4, 6, 7, 9, 10, 12 and 13) were loaded at 1 / 2 the volume of the corresponding experimental test. If ligation occurs, the substrate size will shift up by 22 nt. The red arrow corresponds to the substrate, and the blue arrow corresponds to the adaptor-ligated substrate. M2 = Thermo Fisher's 25 bp DNA ladder (C) Table of expected ligation product sizes and estimation of ligation efficiency using ImageJ. Ligation efficiency was estimated by dividing the intensity of the ligated product by the total intensity of the ligated and unligated products.
[0024] Figures 22(A) to 22(D): Gel analysis of size migration of ligation products using 6% TBE polyacrylamide gels. Red arrows correspond to substrates, and blue arrows correspond to adaptor-ligated substrates: nick (A, left), 5'-overhang (A, right), 1 nt gap (B), 2 nt gap (C), and 3 nt gap (D). M2 = Thermo Fisher 25 bp DNA ladder. The sequences of the two adaptors (Ad1 and Ad2) were compared, and the different bases (A or G) at the 5' end of the ligation junction of Ad2 were also examined. **(e) Ligation efficiency table calculated from band intensities using ImageJ.
[0025] Figures 23(A) and 23(B): Figure 23(A) shows a schematic diagram of 3' branch ligation on a DNA / RNA hybrid with a 20 bp complementary region. We tested whether blunt-ended adapters would ligate to the 3'-end of DNA at the 5'-RNA overhang and / or to the 3'-end of RNA at the 5'-DNA overhang. (B) Gel analysis of size migration of ligation products using a 6% denaturing polyacrylamide gel. Red arrows correspond to RNA substrates (29 nt) and green arrows correspond to DNA substrates (80 nt). Blue arrows correspond to adapter-ligated RNA substrates. If ligation occurs, the substrate size will shift up by 20 nt. Reactions 1 and 2 were performed in duplicate. M2 = 25 bp DNA ladder from Thermo Fisher.
[0026] Figures 24(A) to 24(C): Figure 24(A) is a schematic diagram of transposon insertion followed by 3'-3' branch ligation and PCR amplification using Pr-A (blue arrow) and Pr-B (green arrow). (B) Amplification products after transposon insertion using TnA and / or TnB and / or 3' branch ligation using AdB primers pr-A, pr-B, or both. Products were electrophoresed on a 6% polyacrylamide gel. M1 = ThermoFisher MassRuler low range DNA ladder. (C) Graph of amplification signals using pr-A and pr-B after various transposon insertion and 3' branch ligation conditions.
[0027] Figure 25 : Mid-length mark.
[0028] Figures 26(A) and 26(B): 3'-branched ligation of unconventional DNA ends formed at nicks, gaps, and overhangs using T4 DNA ligase. (A) Schematic diagram of ligation assays on different DNA acceptor types. The blunt-ended DNA donor is a synthetic, partially dsDNA molecule with a dideoxy 3'-end (solid circles) to prevent self-ligation of the DNA donor. The long arm of the donor is 5' phosphorylated. DNA acceptors are assembled using two or three oligonucleotides to form a nick (no phosphate), a gap (1 or 8 nt), or a 36-nt 3'-recessive end. All strands of the substrate are unphosphorylated, and the scaffold strand is protected by a 3' dideoxy bond. (B) Size migration analysis of ligation products from substrates 1, 2, 3, and 4, respectively, using a 6% denaturing polyacrylamide gel. Negative no-ligase controls (lanes 1, 3, 4, 6, 7, 9, 10, 12, and 13) were loaded at one-fold or one-half the volume of the sample tested in the corresponding experiment. If ligation occurs, the substrate size will shift up by 22 nt. The red arrow corresponds to the substrate, and the purple arrow corresponds to the donor-ligated substrate. A 25 bp DNA ladder from Thermo Fisher was used. Donor and substrate sequences are in Table S1. Table 8 shows the expected sizes of substrates and ligation products in each experimental group, as well as the approximate ligation efficiency. The intensity of each band was estimated using ImageJ and normalized by its expected size. The ligation efficiency was estimated by dividing the normalized intensity of the ligated product by the normalized total intensity of the ligated and unligated products.
[0029] Figures 27(A) to 27(E): Gel analysis of size migration of ligation products using 6% TBE polyacrylamide gel. Red arrows correspond to substrates and purple arrows correspond to substrates for donor ligation: substrate 5 (nicked) (A), substrate 6 (1-nt gap) (B), substrate 7 (2-nt gap) (C), substrate 8 (3-nt gap) (D), and substrate 9 (3'-recessed end) (E). A 25 bp DNA ladder from Thermo Fisher was used. Three DNA donors with different bases at the 5' end (T, A, or GA) of the ligation junction were examined. Table 9 shows the ligation efficiency calculated based on normalized band intensities using ImageJ.
[0030] Figures 28(A) to 28(D). 3' branched ligation of RNA 3' ends in DNA / RNA hybrids. Schematic diagram of 3' branched ligation on DNA / RNA hybrids with a 20 bp complementary region. We tested whether a blunt-ended DNA donor would ligate to the 3' recessed end of DNA and / or the 3' recessed end of RNA. DNA (ON-21) hybridized to the RNA strand (A), while DNA (ON-23) failed to hybridize to the RNA strand (B). Figures 28(C) and (D) show gel analysis of the size migration of ligation products using a 6% denaturing polyacrylamide gel. The red arrow corresponds to the RNA substrate (29 nt) and the green arrow corresponds to the DNA substrate (80 nt). The purple arrow corresponds to the donor-ligated RNA substrate. If ligation occurs, the substrate size shifts up by 20 nt. (c) Lanes 1 and 2, replicate experiments; lanes 7-10, no ligase control; 10% PEG with T4 DNA ligase was added. (d) Lane 1, no ligase control; lanes 2, 3, and 8, T4 DNA ligase with 10% PEG; lanes 4, 5, and 9, T4 RNA ligase 1 with 20% DMSO; lanes 6, 7, and 10, T4 RNA ligase 2 with 20% DMSO. A 25 bp DNA ladder from Thermo Fisher was used. This corresponds to Figure 23 but is not entirely accurate in several respects.
[0031] Figures 29(A) to 29(C) show schematic diagrams of three transposon tagging methods followed by PCR amplification using Pr-A (blue arrows) and Pr-B (green arrows). Two transposon methods (A); a single Y transposon tag with 3' gap filling (B); and a single transposon method with adapter ligation at the 3' gap (C). Figure 29(D) shows a graph of the amplification signal after purification using either Pr-A or Pr-A and Pr-B after various tagging and gap ligation conditions. This corresponds to Figure 23, but is not entirely accurate in several respects.
[0032] Figures 30(A) to 30(C). Base distribution deviations for Tn5 gap junctions (A), two transposons (B), and conventional TA junctions (C). Only the first 20 bases at each end of the junction are shown; adenine, blue; cytosine, orange; guanine, gray; thymine, yellow; averages and standard deviations are given for five independent libraries. Not present at all.
[0033] Figures 31(A) and 31(B). DNA 3' branch ligation under different addition conditions. (A) Ligation was performed on 5'-overhanging DNA with titrated ATP concentrations. The experiment was repeated with 0.01 mM ATP (lanes 4 and 5) and 0.005 mM ATP (lanes 6 and 7). Lane 9 is a control without donor. (B) DNA 3' branch ligation at a nick, a 1-nt gap, an 8-nt gap, a 5'-overhang, and a blunt end with or without SSB and ligase. Red arrows correspond to substrate, purple arrows correspond to donor-ligated substrate. Not present at all.
[0034] Table 1: Phasing and variant calling statistics. Unless otherwise stated, reads were mapped to Hg37 with bait sequences, and variants were called using GATK with default settings for all libraries. SNPs from the GIAB high-confidence variant calling VCF were used as input for phasing.
[0035] Table 2: Scaffold statistics.
[0036] Table 3: Filtering reduces false positive calls. Final FP calls were calculated by subtracting 1666 from the filtered FPs, with the exception of the STD library, which by definition does not share any of these FPs with the stLFR library, as the stLFR library was made from the GIAB reference.
[0037] Table 4: LongHap SNP and Indel phasing.
[0038] Table 5: Filtering criteria. Various filtering criteria described in the Materials and Methods section were used to remove FPs.
[0039] Table 6: Exemplary sequences. DETAILED DESCRIPTION
[0040] 1.stLFR library process
[0041] 1.1 Introduction
[0042] Here, we describe the implementation of single-tube long fragment read (stLFR) technology (15), an efficient DNA co-barcoding technology that can achieve millions of barcodes in a single tube. See WO 2014 / 145820A2 (2014), incorporated herein by reference for all purposes. This is achieved by using the surface of microbeads instead of compartments (e.g., the wells of a 384-well plate). Each bead carries many copies of a unique barcode sequence, which is transferred to subfragments of each long DNA molecule. These co-barcoded subfragments are then analyzed on a common short-read sequencing instrument (e.g., BGISEQ-500 or equivalent). In this implementation of this method, we used a ligation-based combinatorial barcode generation strategy to create more than 1.8 billion different barcodes through three ligation steps. For a single sample, we used beads with approximately 10-50 million barcodes to capture approximately 10-100 million long DNA molecules in a single tube. Few two beads will share the same barcode because we sampled 10-50 million beads from such a large total barcode library. Furthermore, using 50 million beads and 10 million long genomic DNA fragments, the vast majority of subfragments of each long DNA fragment are co-barcoded by a unique barcode. This is similar to single-molecule sequencing of long reads and may provide a powerful informatics approach for de novo assembly. Importantly, stLFR is easy to perform and can be achieved with relatively small oligonucleotide inputs to generate barcoded beads. Furthermore, stLFR uses standard equipment found in almost all molecular biology laboratories and can be analyzed with almost any sequencing strategy. Finally, stLFR replaces standard NGS library preparation methods, requires only 1 ng of DNA, and does not significantly increase the cost of whole genome or whole exome analysis, with a total cost of less than $30 per sample.
[0043] As used herein, "single tube" refers to the analysis of a large number of individual DNA fragments without the need to separate the fragments into separate tubes, containers, aliquots, wells or droplets during the labeling step. Instead, the surface of the beads can replace the compartments.
[0044] The first step in stLFR is to insert hybridization sequences, preferably at regular intervals, along the genomic DNA fragment. The appropriate spacing may vary depending on the application and the desired outcome, but is typically between 100-1500 bp, typically 200-1000 bp. This is accomplished by transposition of the bound DNA sequence. In one embodiment, the transposase is Tn3, Tn5, Tn7, or Mu. Typically, Tn5 transposase is used (see Picelli et al., 2014, incorporated herein by reference for all purposes). The transposed DNA or insert sequence contains a single-stranded region for hybridization (the "hybridization sequence") and a double-stranded mosaic sequence (the "mosaic sequence") that is recognized by the enzyme and enables the transposition reaction. Figure 1A This transposition step is performed in solution (as opposed to attaching the insert sequence directly to the beads). This allows for very efficient incorporation of the hybridizing sequence along the genomic DNA molecule. As previously observed (10), transposases have the property of remaining bound to genomic DNA after the transposition event, effectively leaving the long genomic DNA molecule intact after transposon integration.
[0045] After the DNA is treated with, for example, Tn5, it is diluted in hybridization buffer and combined with the cloned barcoded beads. In one approach (Example below), 50 million 2.8um cloned barcoded beads. Each bead contains approximately 400,000 capture adapters (also called capture oligonucleotides or capture oligonucleotides), each containing the same barcode sequence. A portion of the capture adapters contains uracil nucleotides to enable destruction of unused adapters in subsequent steps. For example, the capture adapters can be 5-50% uracil, more often 5-50%, and more often 5-20%. The mixture is incubated under optimized temperature and buffer conditions, during which time the transposon-inserted DNA is captured onto the beads via hybridization sequences.
[0046] It has been proposed that genomic DNA in solution forms a ball with two tails sticking out (16). This can enable the capture of long DNA fragments towards one end of the molecule, followed by a rolling motion that wraps the genomic DNA molecule around the bead. Each bead has a capture oligonucleotide approximately every 7.8 nm on its surface. This allows for very uniform and high-rate subfragment capture. A 100 kb genomic fragment will wrap around a 2.8 μm bead approximately three times. In our data, the longest fragment captured was 300 kb in size, suggesting that larger beads may be required to capture longer DNA molecules.
[0047] In alternative embodiments, parameters such as bead size, oligonucleotide spacing, or the number of different oligonucleotides in each mixture can be varied. For example, the diameter of the beads used can be 1-20 μm, or 2-8 μm, 3-6 μm, or 1-3 μm. For example, the spacing of the barcoded oligonucleotides on the beads can be at least 1 nm, at least 2 nm, at least 3 nm, at least 4 nm, at least 5 nm, at least 6 nm, or at least 7 nm. In some embodiments, the spacing is less than 10 nm (e.g., 5-10 nm), less than 15 nm, less than 20 nm, less than 30 nm, less than 40 nm, or less than 50 nm. In some embodiments, the number of different barcodes used in each mixture can be >1M, >10M, >30M, >100M, >300M, or >1B. As described below, for example, using the methods described herein, a large number of barcodes can be generated for use in the present invention. In some embodiments, the number of different barcodes used per mixture can be >1M, >10M, >30M, >100M, >300M, or >1B, and they are sampled from a pool with at least 10-fold greater diversity (e.g., >10M, >0.1B, 0.3B, >0.5B, >1B, >3B, >10B of different barcodes from beads).
[0048] Individual barcode sequences are transferred at regular intervals by ligating the 3' end of the capture adaptor to the 5' end of the hybridization sequence of the transposon insert, mediated by a bridge or splint oligonucleotide (the terms are used interchangeably) having a first region complementary to the capture adaptor and a second region complementary to the hybridization sequence ( Figure 1A and Figure 15 The beads are collected and the DNA / transposase complex is disrupted, generating subfragments less than 1 kb in size.
[0049] If necessary, sample barcodes can be obtained in this step. Transposons with unique barcodes between the chimeric sequence and the hybridization sequence are used. These can be synthesized in 96, 384 or 1536 plate formats, with each well containing many copies of transposons carrying the same barcode, and each barcode is different between the wells. Using these barcode transposons, different DNA samples can be subjected to transposon insertion in 96, 384 or 1536 plate formats. Samples labeled with sample barcodes can be multiplexed in any manner.
[0050] Due to the large number of beads and the high capture oligonucleotide density per bead, the amount of excess adapter is four orders of magnitude greater than the amount of product. This large amount of unused adapter can overwhelm the following steps. To circumvent this, we designed beads with 5' end-ligated capture oligonucleotides. This enabled the development of an exonuclease strategy that specifically degrades excess, unused capture oligonucleotides. See Figure 14 and 16 Uracil DNA glycosylase (UDG) can also be used to degrade excess adaptors.
[0051] In one aspect, the method comprises combining (i) a first fragment of a target nucleic acid and (ii) a population of beads into a single mixture, wherein each bead comprises an oligonucleotide immobilized thereon, the oligonucleotide comprising a tag-containing sequence (or a barcode adapter), wherein each tag-containing sequence comprises a tag sequence, wherein the oligonucleotides immobilized on the same single bead comprise the same tag-containing sequence, and a majority of the beads have different tag sequences. In some embodiments, the DNA fragment is a concatamer of at least 2, at least 10, at least 30, or at least 100 copies of a DNA or cDNA molecule. The length of the nucleic acid monomer can be 0.5 kb to 10 kb, or >1 kb, or >10 kb. In some methods, the sequence of >50% or >70% >90%, 95%, >99%, 100% of the bases of the DNA or cDNA molecules in the mixture is determined.
[0052] 1.1.1 Two transposon methods
[0053] In one approach to achieving stLFR, two different transposons are used in the initial insertion step, allowing PCR to be performed after exonuclease treatment. However, this approach results in coverage of approximately 50% or less per long DNA molecule because it requires the insertion of two different transposons adjacent to each other to generate a suitable PCR product.
[0054] 1.1.2 Single transposon method using 3′ branch ligation
[0055] To achieve the highest coverage per genomic DNA fragment, we used a single transposon in the initial insertion step and added additional adapters by ligation. This noncanonical ligation, called 3' branched ligation, involves covalently attaching the 5' phosphate from the blunt-ended adapter to the recessed 3' hydroxyl group of the genomic DNA ( Figure 1A). Branch connection is described in Example 3 below. See also U.S. Patent Publication No. US2018 / 0044668 and International Application No. WO2016 / 037418, which are incorporated herein by reference for all purposes. See also U.S. Patent Publication No. 2018 / 0044667, which is incorporated by reference for all purposes. Using this method, it is theoretically possible to amplify and sequence all subfragments of a captured genomic molecule.
[0056] Additionally, this ligation step enables sample barcodes to be placed adjacent to genomic sequences for multiplexed sampling. The benefit of using these adapters for sample barcoding is that the barcodes can be placed adjacent to the genomic DNA, allowing the barcodes and genomic DNA to be sequenced using the same primers, without the need for additional sequencing primers to read the barcodes. Sample barcoding allows preparations from multiple samples to be pooled and differentiated by barcode prior to sequencing. 3' branched ligation adapters can be synthesized in 96, 384, or 1536 plate formats, with each well containing many copies of the adapter carrying the same barcode, and each barcode being different between wells. After capture on the beads, these adapters can be used for ligation in 96, 384, or 1536 plate formats.
[0057] After this ligation step, PCR is performed and the library is ready to enter any standard next generation sequencing (NGS) workflow. It will be appreciated that a first primer that hybridizes to a site on the capture oligonucleotide or its complement (see Figure 1A ) and a second primer that hybridizes to a site of the 3' branched ligation adapter or its complement. In the case of the BGISEQ-500, the library is circularized as previously described (17). DNA nanoballs are prepared from the single-stranded circles and loaded onto patterned nanoarrays (17). These nanoarrays are then sequenced on the BGISEQ-500 using combined probe-anchor synthesis (cPAS) (18-20). After sequencing, the barcode sequences are extracted. Mapping the read data by unique barcodes shows that most reads with the same barcode are clustered in the genomic region that corresponds to the DNA length used during library preparation ( Figure 1B ). This method and the experimental protocol for preparing beads are described in detail in Examples 1 and 2.
[0058] In some embodiments, >50%, >70%, >80%, >90%, or >95% of the barcoded DNA fragments are barcoded with unique barcodes. In some embodiments, >50%, >70%, >80%, >90% of the subfragments in the fragments are linked to barcode oligonucleotides. In some embodiments, an average of >10%, >20%, >40%, >50%, >60% of the subfragments in the long fragments are sequenced.
[0059] 1.2stLFR read coverage and variant calling
[0060] To demonstrate stLFR phasing and variant calling, we generated four libraries using 1 ng (stLFR-1 and stLFR-2) and 10 ngs (stLFR-3 and stLFR-4) of DNA from NA12878. The number of beads used varied, namely 10 million (stLFR-3), 30 million (stLFR-4), and 50 million (stLFR-1 and stLFR-2). Finally, the 3' branch joining method (stLFR-1, stLFR-2, and stLFR-3) and two transposon methods (stLFR-4) were tested. The sequencing depth of stLFR-1 and stLFR-2 reached 336 Gb and 660 Gb of total base coverage, respectively. We also analyzed the coverage of downsampling. stLFR-3 and stLFR-4 sequenced to a moderate level of 117 Gb and 126 Gb, respectively. The co-barcoded reads were mapped to build 37 of the human reference genome using BWA-MEM (21). Since stLFR does not require any pre-amplification step, the distribution of read coverage across the genome was close to a Poisson distribution (Figure 3). The non-replicated coverage was 34-58X, and the number of long DNA molecules per barcode was 1.2-6.8 (Tables 1 and Figure 1C As expected, the stLFR library made from 50 million beads and 1 ng of genomic DNA had the highest single unique barcode co-barcoding rate of over 80% ( Figure 1C These libraries also observed the highest average non-overlapping read coverage of 10.7–12.1% per long DNA molecule and the highest average non-overlapping base coverage of 17.9–18.4% per captured subfragment per long DNA molecule (Fig. 1d). This coverage is approximately 10-fold higher than previously demonstrated using 3 ng of DNA and transposons attached to beads (12).
[0061] For each library, variants were called using GATK (22) with default settings. SNP and indel calls were compared to genome in a bottle (GIAB) (23) to determine false positive (FP) and false negative (FN) rates (Table 1). In addition, we performed variant calls using the same settings in GATK on a standard non-stLFR library made from approximately 1000 times more genomic DNA and also sequenced on a BGISEQ-500 (STD) and on a Chromium library from 10X genome (11). We also compared the accuracy and sensitivity reported by Zhang et al. in a bead haplotype library study (12), which is incorporated herein by reference for all purposes. Our stLFR method and the method described by Zhang et al. showed lower SNP and indel FP rates compared to the Chromium library. The FP and FN rates for stLFR were 2-fold higher than those for the STD library, and the FN rates were higher or lower than those for the Chromium library depending on the specific stLFR library and filtering criteria. The higher FN rate in the stLFR library compared to the standard library is mainly due to the shorter average insert size ( 200 bp vs. 300 bp in the standard library). That is, the FN rates for SNPs and indels were much lower in stLFR than those of Zhang et al., and for indels, the FN rate was much lower than that of the Chromium library (Table 1). Overall, our stLFR library performed better than the results published by Zhang et al. or the Chromium library for most metrics of variant calling, especially when non-optimized mapping and variant calling methods were used (Table 1, "No filter").
[0062] A potential issue with using GIAB data to measure FP rates is that we were unable to use the GIAB reference material (NIST RM 8398) due to the small size of the isolated DNA fragments. Therefore, we used the GM12878 cell line and isolated DNA using a dialysis-based method that produces very high-molecular-weight DNA (see Methods). However, our GM12878 cell line isolates may harbor many unique somatic mutations compared to the GIAB reference material, leading to an increased number of FPs in our stLFR libraries. To further examine this, we compared the overlap of single-nucleotide FP variants between the four stLFR libraries and the two non-LFR libraries (Figure 4a). Overall, 544 FP variants were shared among the six libraries, while 2078 FP variants were unique to the four stLFR libraries. We also compared the stLFR FPs with the Chromium libraries and found that more than half of these shared FPs (1194) were also present in the Chromium libraries (Figure 4b). Inspection of the read and barcode coverage of these shared variants revealed that they were more similar to TP variants (Figures 5-6). We also examined the distribution of these shared FP variants with 2078 randomly selected variants across the genome (Figure 7a). This analysis showed that 219 variants were found in clusters where two or more of these FPs were within 100bp of each other. However, the distribution of most (90%) variants did not appear to be different from that of randomly selected variants. In addition, of those FPs shared between the stLFR and Chromium libraries, only 41 were found to be clustered (Figure 7a). Ultimately, GIAB called 96 of these variants, but with different zygosity than those called in the stLFR library.
[0063] If we accept the evidence that these shared FP variants are largely real and not present in the GIAB reference material, the FP rate of stLFR may be up to 1859 fewer variants than reported for SNP detection in Table 1. This is still thousands of single nucleotide variants more than the standard BGISEQ-500 library. To further improve the FP rate in the stLFR library, we tested a number of different filtering strategies to eliminate errors. Ultimately, by applying a number of filtering criteria based on reference and variant allele ratios and barcode counts (see Examples), we were able to remove 3647-13840 FP variants depending on the library and coverage. Importantly, this was achieved while only increasing the FN rate by 0.10-0.29% in the stLFR library. After this filtering step, we examined the shared FPs between the four stLFR libraries. Filtering only removed 340 shared FP variants, of which 147 were clustered within 100 base pairs of each other and were likely not real (Figure 7b). This further suggests that the majority of these shared FPs are real variants. Taking these variations into account and the reduction in the number of FP variants after filtering, the FP rate was similar and the FN rate was 2-3 times higher compared to the filtered STD library using SNP calling (Table 3). This increased FN rate was mainly due to the increase in non-unique mapping of partner pairs with short insert lengths in the stLFR library.
[0064] 1.3st LFR phasing performance
[0065] To assess the phasing performance of the variants, high-confidence variants from GIAB were phased using the publicly available software package HapCut2 (24). Depending on the library type and the amount of sequence data, more than 99% of all heterozygous SNPs were placed in contigs with N50s ranging from 0.6 to 15.1 Mb (Table 1). The stLFR-1 library, with 336 Gb of total read coverage (44× unique genome coverage), achieved the highest phasing performance with an N50 of 15.1 Mb. The N50 length appears to be primarily influenced by the length and coverage of long genomic fragments. This can be seen from the reduced N50 for stLFR-2, as the DNA used for this sample was slightly older and more fragmented than the material used for stLFR-1 (Table 1, average fragment length 52.5 kb vs. 62.2 kb), while the N50 for the 10 ng libraries (stLFR-3 and 4) was 10 times shorter. Comparison with GIAB data showed low short and long switch error rates, comparable to previous studies (11, 12, 25). stLFR performance was very similar to Chromium libraries. Because no read data were available for the bead haplotype method by Zhang et al., we could only compare our results with those of a phasing algorithm written and optimized specifically for their data. This showed that the stLFR-1 and stLFR-2 libraries had longer N50s, similar short switch error rates, but higher long switch error rates. stLFR-3 and stLFR-4, which use more DNA, had N50s similar to those of Zhang et al. However, direct comparison is difficult due to differences in DNA input and coverage.
[0066] It should be noted that this phasing result was obtained using a program that was not written for stLFR data. To see if this result could be improved, we developed a phasing program, LongHap, and optimized it specifically for stLFR data. Using GIAB variants, LongHap was able to phase 99% of the SNPs into contigs with an N50 of 18.1 Mb (Table 1). Importantly, this increase in the length of these contigs was achieved while reducing short and long switching errors (Table 1). LongHap can also phase indels. Applying LongHap to stLFR-1 using GIAB SNPs and indels resulted in an N50 of 23.4 Mb, but also resulted in an increased switching error rate (Table 4).
[0067] 1.4 Structural variation detection
[0068] Previous studies have shown that long-range information can improve the detection of structural variants (SVs) and have described large deletions (4–155 kb) in NA12878 (11, 12). To demonstrate the ability of stLFR to detect SVs, we examined barcode overlap data from stLFR-1 and stLFR-4 libraries in these regions as previously described (12). In each case, deletions were observed in the stLFR-1 data, even at lower coverage (Figures 2a and 8). Closer inspection of coverage of chromosome 8 The co-barcoded sequence reads for the 150 kb deletion indicated that the deletion was heterozygous and present in a single haplotype (Figure 2b-c). The 10 ng stLFR-4 library also detected most deletions, but the smallest three were difficult to identify due to the lower per-fragment coverage (and therefore less barcode overlap) of this library.
[0069] To evaluate the performance of stLFR for detecting other types of SVs, we prepared libraries for cell lines from patients with known translocations between chromosomes 5 and 12 (26) and for the cell line GM20759, which has a known inversion in chromosome 2 (27). The stLFR libraries were able to identify both inversions and translocations in the corresponding cell lines (Fig. 2d). Downsampling the read amount for each library showed that even with only 5 Gb of read data, strong signals for translocations were detected ( Finally, examination of the two SVs in the stLFR-1 library did not result in a clear pattern (Figure 9i-l), indicating that the false positive rate for detecting these types of SVs is low.
[0070] 1.5 Scaffold contigs using stLFR
[0071] stLFR is a powerful method, in part because it uses a very large number (e.g., approximately 1.8 billion) of unique barcodes and is able to achieve co-barcoding that is unique to each individual long genomic DNA molecule. Such data should be beneficial for de novo genome assembly and improved scaffolding. To demonstrate how stLFR can be used to improve genome assembly, we used reads from the stLFR-1 and stLFR-4 libraries and SALSA (28), a program designed for chromatin conformation capture (Hi-C) data, to support single-molecule real-time (SMRT) read assembly of NA12878 (29). SALSA was not designed for stLFR data, so it was necessary to modify the stLFR data to resemble the Hi-C structure. This was achieved by selecting read pairs that share the same barcode and are located at the ends of the captured long DNA molecules. These were then labeled as read pairs for the SALSA program. Replacing the Hi-C data with stLFR data produced excellent scaffolding. Using only 60 million stLFR reads, 1411 contigs were connected to 597 scaffolds with an N50 of 44.7 Mb. These scaffolds covered 2.84 Gb of the genome. These metrics compared very favorably with those generated in the SALSA manuscript using the same contigs and 10x (734 million) Hi-C reads generated from human embryonic stem cells (Table 2). The quality of the stLFR scaffolds was further analyzed by aligning the stLFR scaffolds to build 37 of the human reference genome and comparing them with the dnadiff program (31). In general, the stLFR scaffolds matched the reference genome very well, and the number of breakpoints, translocations, relocations, and inversions was similar to that of scaffolds generated using Hi-C reads (Table 2). The alignment dot plots further demonstrated the high degree of continuity between the stLFR scaffolds and the reference genome (Figure 10).
[0072] 1.6 Discussion
[0073] Here, we describe an efficient whole-genome sequencing library preparation technology, stLFR, that enables co-barcoding of subfragments of long genomic DNA molecules with individual, unique clonal barcodes in a single-tube approach. Using microbeads as miniaturized compartments, a virtually unlimited number of clonal barcodes can be used per sample at negligible cost. Our optimized hybridization-based capture of transposon-inserted DNA on beads, combined with 3' branch ligation and exonuclease degradation of an extreme excess of capture adapters, successfully barcoded up to approximately 20% of subfragments from DNA molecules up to 300 kb in length. Importantly, this is achieved without DNA amplification of the original long DNA fragments and the attendant representation bias. Thus, stLFR addresses the cost and limited co-barcoding capabilities of emulsion-based methods.
[0074] The quality of variant calling using stLFR is very high and, with further optimization, could potentially approach that of standard WGS methods, but with the added benefit of shared barcoding enabling advanced informatics applications. We demonstrate high-quality, near-complete genome phasing into long contigs with very low error rates, detection of SVs, and scaffolding of contigs enabling de novo assembly applications. All of this is achieved with a single library that requires neither specialized equipment nor the added cost of library preparation.
[0075] Due to effective barcoding, we successfully prepared stLFR libraries using as little as 1 ng of human DNA (600X genome coverage) and obtained high-quality WGS, with most subfragments uniquely co-barcoded. Less DNA can be used, but stLFR does not employ DNA amplification during the co-barcoding process, and therefore does not generate overlapping subfragments from each individual long DNA molecule. Therefore, whole-genome coverage suffers as the amount of DNA decreases. Furthermore, because stLFR currently retains 10–20% of each original long DNA molecule for subsequent PCR amplification, a sampling issue arises. This results in relatively high read duplication rates and increases sequencing costs, but improvements are possible. One obvious solution is to remove the PCR step. This would eliminate sampling but would also significantly reduce false-positive and false-negative error rates. Furthermore, improvements such as optimizing the insertion distance between transposons and increasing the length of paired-end sequencing reads to 200 bases should be readily implementable and would increase coverage and overall quality. For some applications, such as structural variation detection, using less DNA and reduced coverage may be desirable. As we demonstrate here, only 5 Gb of sequence coverage can faithfully detect inter- and intrachromosomal translocations, where the duplication rate is negligible. Indeed, stLFR may represent a simple and cost-effective alternative to long mate pair libraries in the clinical setting.
[0076] Additionally, we believe that this type of data could enable complete de novo assembly of diploids from a single stLFR library without the need for long physical reads, such as those produced by SMRT or nanopore technologies. An interesting feature of transposon insertion is that it produces a 9-base overlap of sequence between adjacent subfragments. Typically, these adjacent subfragments are captured and sequenced, doubling the synthetic length of the read (e.g., for a 200-base read, two adjacent captured subfragments create two 200-base reads, overlapping by 9 bases, or 391 bases). stLFR does not require special equipment like droplet-based microfluidics methods, and the cost per sample is minimal. In this paper, we demonstrate that 50 million beads can be used, but many more beads can be used. This will enable many types of cost-effective analyses, in which hundreds of millions of barcodes would be useful. We envision that such inexpensive large-scale barcoding could be used for RNA analysis, such as full-length mRNA sequencing of 1000s cells in combination with single-cell technologies or deep population sequencing of 16S RNA in microbial samples. Using stLFR, phase-angle chromatin mapping or methylation studies can also be performed using assays for transposase-accessible chromatin (ATAC-seq) (32).
[0077] 1.7 Target Nucleic Acid
[0078] As used herein, the term "target nucleic acid" (or polynucleotide) or "nucleic acid of interest" refers to any nucleic acid (or polynucleotide) suitable for processing and sequencing by the methods described herein. Nucleic acids can be single-stranded or double-stranded and can include DNA, RNA or other known nucleic acids. Target nucleic acids can be nucleic acids from any organism, including but not limited to viruses, bacteria, yeast, plants, fish, reptiles, amphibians, birds and mammals (including but not limited to mice, rats, dogs, cats, goats, sheep, cattle, horses, pigs, rabbits, monkeys and other non-human primates and humans). Target nucleic acids can be obtained from an individual or multiple individuals (i.e., a population). The sample from which nucleic acid is obtained can contain nucleic acids from cells or even a mixture of organisms, for example: a human saliva sample comprising human cells and bacterial cells; a mouse xenograft comprising mouse cells and transplanted human tumor cells, etc. The target nucleic acid may not be amplified, or it may be amplified by any suitable nucleic acid amplification method known in the art. Target nucleic acids can be purified according to methods known in the art to remove cellular and subcellular contaminants (lipids, proteins, carbohydrates, nucleic acids other than the nucleic acid to be sequenced, etc.), or they can be unpurified, i.e., include at least some cellular and subcellular contaminants, including but not limited to intact cells, which are disrupted to release their nucleic acids for processing and sequencing. Target nucleic acids can be obtained from any suitable sample using methods known in the art. These samples include but are not limited to: tissues, isolated cells or cell cultures, body fluids (including but not limited to blood, urine, serum, lymph, saliva, anal and vaginal secretions, sweat and semen); air samples, agricultural samples, water samples and soil samples, etc. Non-limiting examples of target nucleic acids include "circulating nucleic acids" (CNA), which are nucleic acids circulating in human blood or other body fluids (including but not limited to lymph, cerebrospinal fluid (liquor), ascites, milk, urine, feces, and bronchial lavage) and can be distinguished as cell-free (CF) nucleic acids or cell-associated nucleic acids (reviewed in Pinzani et al., Methods 50:302-307, 2010).
[0079] Target nucleic acid can be genomic DNA (for example, from a single individual), cDNA, and / or can be composite nucleic acid, including nucleic acid from multiple individuals or genomes.The example of composite nucleic acid includes microbiome, circulating fetal cells in pregnant women's blood (see, for example, Kavanagh et al., J.Chromatol.B 878:1905-1911,2010), from tumor cells (CTC) circulating in cancer patients' blood (see, for example, Allard et al., Clin Cancer Res.10:6897-6904,2004). Another example is the genomic DNA from a single cell or a small amount of cells, such as from biopsy tissue (for example, fetal cells from blastocyst trophectoderm biopsy; Cancer cells aspirated from a needle of a solid tumor, etc.). Another example is pathogens in tissue, blood or other body fluids, such as bacterial cells, viruses or other pathogens. As used herein, the term "composite nucleic acid" refers to a large number of different nucleic acids or polynucleotides. In certain embodiments, the target nucleic acid is genomic DNA; exome DNA (a subset of complete genomic DNA enriched for transcribed sequences, comprising a set of exons in a genome); a transcriptome (i.e., the collection of all mRNA transcripts produced in a cell or cell population, or cDNA produced from such mRNA); a methylome (i.e., the population of methylation sites and the pattern of methylation in a genome); an exome (i.e., the protein-coding regions of a genome selected by exon capture or enrichment methods); a microbiome; a mixture of genomes from different organisms; a mixture of genomes from different cell types of an organism; and other complex nucleic acid mixtures comprising a large number of different nucleic acid molecules (examples include, but are not limited to, microbiomes, xenografts, solid tumor biopsies comprising normal and tumor cells, etc.), including subsets of the aforementioned types of complex nucleic acids. In one embodiment, such complex nucleic acids have a complete sequence comprising at least one gigabase (Gb) (a diploid human genome comprises approximately 6 Gb of sequence).
[0080] In some cases, the target nucleic acid or the first fragment is a genomic fragment. In some embodiments, the genomic fragment is longer than 10kb, for example, 10-100kb, 10-500kb, 20-300kb or longer than 100kb. The quantity of the DNA (such as human genomic DNA) used in a single mixture can be <10ng, <3ng, <1ng, <0.3nm or <0.1ng of DNA. In some cases, the length of the target nucleic acid or the first fragment is 5000 to 100000KB.
[0081] 1.8 Other methods
[0082] Although the working examples described herein use polymerase chain reaction, other nucleic acid amplification methods may be used. It is within the ability of those skilled in the art to make modifications suitable for appropriate amplification techniques.
[0083] Figure 17 Other methods are shown. Figure 17 The production of double-stranded DNA molecules (DNBs) is shown, which can be inserted by transposons and captured by stLFR beads. Up to several thousand copies (e.g., 10-10,000 copies, such as 10-1,000 copies or 100-1,000 copies) can be replicated on the same DNA strand. This allows for high coverage of the original molecule by stLFR sequencing. Figure 18 It is demonstrated that when limited amounts of template DNA are available, a limited pre-amplification step can be used before stLFR. Figure 19 A method is described in which a low concentration of random nicking enzyme, a moderate concentration of Klenow fragment, and a high concentration of ligase are used. The concentrations of beads and DNA are suitable for stLFR. When the nick is formed by Klenow and opens to the gap, ligation occurs immediately, locking the long fragment to the bead. The nicking is allowed to proceed, and more gaps are opened to allow more adapters to be connected in the gaps. Primer extension produces fragments of approximately 500 base pairs. A second adapter is ligated to the blunt end, and the library is sequenced. Figure 20 The ligation of hairpin adapters on long DNA is shown, and the use of primers and Ph29 or similar polymerases in the loop to generate concatemerized dsDNA before barcoding. In addition to improving the read coverage of each molecule, an interesting result of this method is that the total "length" (number of bases) of each concatemer is similar at the end of the 0.5-3h polymerase reaction, regardless of the starting fragment length. This provides the option of using barcode beads with a binding capacity corresponding to the size of the concatemer, thereby preventing each bead from binding multiple concatemers. This will reduce the number of beads required for each reaction, further reducing costs.
[0084] Figure 25A method for intermediate length labeling is shown. In one method, 96 or more different barcode transposons are connected in groups of 10 or less by a linker portion (e.g., DNA, a long inert molecule such as dextrin or polyethylene glycol (PEG), or a long protein such as keratin or collagen). Hybridization and connection can be used to connect the transposon to the linker DNA. Other methods can be by chemical bonds or by connecting avidin to these molecules, and then biotin is connected to the transposon. This achieves two things, it controls the insertion distance between transposons and provides proximity information (10kb or less) of the intermediate reads. This is useful for analyzing repetitive sequences (tandem repeats, trinucleotide mapping, etc.). As with other stLFR methods described herein and elsewhere, DNA containing the inserted sequence can be captured on beads. See Joseph C. Mellor et al., “Phased NGS Library Generation via Tethered Synaptic Complexes,” seqWell (2017), available on the World Wide Web at (http: / / )seqwell.com / wp-content / uploads / 2017 / 02 / seqWell_LongBow_poster_AGBT2017.pdf (last accessed May 16, 2018).
[0085] 1.9 References for Part 1
[0086] 1.K.Zhang et al., Long-range polony haplotyping of individual humanchromosome molecules. Nat Genet 38, 382-387 (2006).
[0087] 2. L.Ma et al., Direct determination of molecular haplotypes bychromosome microdissection. Nat Methods 7, 299-301 (2010).
[0088] 3.JOKitzman et al., Haplotype-resolved genome sequencing ofaGujarati Indian individual.Nat Biotechnol 29,59-63(2011).
[0089] 4.E.K.Suk et al.,A comprehensively molecular haplotype-resolvedgenome of a European individual.Genome Res 21,1672-1685(2011).
[0090] 5.H.C.Fan,J.Wang,A.Potanina,S.R.Quake,Whole-genome molecularhaplotyping of single cells.Nat Biotechnol 29,51-57(2011).
[0091] 6.B.A.Peters et al.,Accurate whole-genome sequencing and haplotypingfrom 10to 20human cells.Nature 487,190-195(2012).
[0092] 7.J.Duitama et al.,Fosmid-based whole genome haplotyping of aHapMaptrio child:evaluation of Single Individual Haplotyping techniques.NucleicAcids Res 40,2041-2053(2012).
[0093] 8.S.Selvaraj,R.D.J,V.Bansal,B.Ren,Whole-genome haplotypereconstruction using proximity-ligation and shotgun sequencing.Nat Biotechnol31,1111-1118(2013).
[0094] 9.V.Kuleshov et al.,Whole-genome haplotyping using long reads andstatistical methods.Nat Biotechnol 32,261-266(2014).
[0095] 10.S.Amini et al.,Haplotype-resolved whole-genome sequencingbycontiguity-preserving transposition and combinatorial indexing.Nat Genet46,1343-1349(2014).
[0096] 11.G.X.Zheng et al.,Haplotyping germline and cancer genomes withhigh-throughput linked-read sequencing.Nat Biotechnol,(2016).
[0097] 12.F.Zhang et al.,Haplotype phasing of whole human genomes usingbead-based barcode partitioning in a single tube.Nat Biotechnol 35,852-857(2017).
[0098] 13.B.A.Peters,J.Liu,R.Drmanac,Co-barcoded sequence reads fromlong DNAfragments:a cost-effective solution for"perfect genome"sequencing.Frontiersin genetics 5,466(2014).
[0099] 14.R.Drmanac.Nucleic Acid Analysis by Random Mixtures of Non-Overlapping Fragments.WO 2006 / 138284 A2(2006).
[0100] 15.R.Drmanac,Peters,B.A.,Alexeev,A.Multiple tagging of longDNAfragments.WO 2014 / 145820 A2(2014).
[0101] 16.K.Jo,Y.L.Chen,J.J.de Pablo,D.C.Schwartz,Elongation andmigration ofsingle DNA molecules in microchannels using oscillatory shear flows.Lab Chip9,2348-2355(2009).
[0102] 17.R.Drmanac et al.,Human genome sequencing using unchained basereadson self-assembling DNA nanoarrays.Science 327,78-81(2010).
[0103] 18.T.Fehlmann et al.,cPAS-based sequencing on the BGISEQ-500toexplore small non-coding RNAs.Clin Epigenetics 8,123(2016).
[0104] 19.J.Huang et al.,A reference human genome dataset of the BGISEQ-500sequencer.Gigascience 6,1-9(2017).
[0105] 20.S.S.T.Mak et al.,Comparative performance of the BGISEQ-500vsIllumina HiSeq2500 sequencing platforms for palaeogenomicsequencing.Gigascience 6,1-13(2017).
[0106] 21.H.Li,R.Durbin,Fast and accurate short read alignment withBurrows-Wheeler transform.Bioinformatics 25,1754-1760(2009).
[0107] 22.A.McKenna et al.,The Genome Analysis Toolkit:a MapReduceframeworkfor analyzing next-generation DNA sequencing data.Genome Res 20,1297-1303(2010).
[0108] 23.J.M.Zook et al.,Integrating human sequence data sets providesaresource of benchmark SNP and indel genotype calls.Nat Biotechnol 32,246-251(2014).
[0109] 24.P.Edge,V.Bafna,V.Bansal,HapCUT2:robust and accuratehaplotypeassembly for diverse sequencing technologies.Genome Res 27,801-812(2017).
[0110] 25.Q.Mao et al.,The whole genome sequences and experimentallyphasedhaplotypes of over 100 personal genomes.Gigascience 5,1-9(2016).
[0111] 26.Z.Dong et al.,Low-pass whole-genome sequencing inclinicalcytogenetics:a validated approach.Genet Med 18,940-948(2016).
[0112] 27.Z.Dong et al.,Identification of balanced chromosomalrearrangementspreviously unknown among participants in the 1000 GenomesProject:implicationsfor interpretation of structural variation in genomes andthe future of clinicalcytogenetics.Genet Med,(2017).
[0113] 28.J.Ghurye,M.Pop,S.Koren,D.Bickhart,C.S.Chin,Scaffolding oflong readassemblies using long range contact information.BMC Genomics 18,527(2017).
[0114] 29.M.Pendleton et al.,Assembly and diploid architecture of anindividualhuman genome via single-molecule technologies.Nat Methods 12,780-786(2015).
[0115] 30.J.R.Dixon et al.,Topological domains in mammaliangenomesidentified by analysis of chromatin interactions.Nature 485,376-380(2012).
[0116] 31.A.M.Phillippy,M.C.Schatz,M.Pop,Genome assembly forensics:findingthe elusive mis-assembly.Genome biology 9,R55(2008).
[0117] 32.JDBuenrostro,PGGiresi,LCZaba,HYChang,WJGreenleaf,Transposition of native chromatin for fast and sensitive epigenomic profiling of open chromatin,DNA-binding proteins and nucleosome position.NatMethods 10,1213-1218(2013). Example
[0118] 2. Example 1: Methods and Materials
[0119] 2.1. Isolation of High Molecular Weight DNA
[0120] Follow RecoverEase TM A modified version of the DNA Isolation Kit (Agilent Technologies, La Jolla, CA) protocol was used to isolate long genomic DNA from cell lines (1).
[0121] Briefly, approximately one million cells were pelleted and lysed with 500 μl of lysis buffer. After incubation at 4°C for 10 minutes, 20 μl of RNase-IT ribonuclease cocktail in 4 mL of digestion buffer was added directly to the lysed cells and incubated on a heat block at 50°C. After 5 minutes, 4.5 mL of proteinase K solution ( 1.1 mg / mL proteinase K, 0.56% SDS and 0.89X TE), and the mixture was incubated for another 2 hours at 50° C. The genomic DNA was then transferred to dialysis tubing with a molecular weight cutoff of 1000 kD (Spectrum Laboratories, Inc., Rancho Dominguez, CA) and dialyzed against 0.5X TE buffer at room temperature overnight.
[0122] 2.2 Barcoded Bead Construction
[0123] Barcoded beads were constructed using a split-and-pool-based strategy using three sets of double-stranded barcoded DNA molecules. Figure 12 and 13 Attach a common adapter sequence containing PCR primer annealing sites to Dynabeads with a 5' double biotin linker TMM-280 streptavidin (ThermoFisher, Waltham, MA) magnetic beads. Three sets of 1536 barcode oligonucleotides containing overlapping sequence regions were constructed by Integrated DNA Technologies (Coralville, IA). Ligation was performed in a 384-well plate in a 15 μL reaction containing 50 mM Tris-HCl (pH 7.5), 10 mM MgCl2, 1 mM ATP, 2.5% PEG-8000, 571 units of T4 ligase, 580 pmol of barcode oligonucleotide, and 65 million M-280 beads. The ligation reaction was incubated at room temperature on a rotator for 1 hour. Between ligations, beads were pooled into a single container by centrifugation, collected to the side of the container using a magnet, and washed once with high-salt wash buffer (50 mM Tris-HCl (pH 7.5), 500 mM NaCl, 0.1 mM EDTA, and 0.05% Tween 20) and twice with low-salt wash buffer (50 mM Tris-HCl (pH 7.5), 150 mM NaCl, and 0.05% Tween 20). The beads were resuspended in 1X ligation buffer and distributed in a 384-well plate, and the ligation step was repeated.
[0124] Some of the "barcodes" referred to herein are "tripartate barcodes." Tripartate refers to their structure and / or their composition. Figure 12 As shown, a triple barcode can be synthesized by continuously connecting shorter (e.g., 4-20 nucleotide) sequences. In one embodiment, the shorter barcode is 10 bases in length. As shown, an exemplary structure includes CS1-BC1-CS2-BC2-CS3-BC3-CS4, where CS is a constant sequence present on all capture adapters, and the BC sequence is a diversified 10-base barcode as described herein. As shown, a partially double-stranded oligonucleotide having the structure CSa-BC-CSb can be annealed to a shorter oligonucleotide (i.e., the complementary sequence of BC (i.e., BC')) to construct a triple barcode.
[0125] On the one hand, the present invention provides a composition comprising beads having capture oligonucleotides comprising attached clonal barcodes, wherein the composition comprises more than 3 billion different barcodes, and wherein the barcode is a triple barcode having the structure 5'-CS1-BC1-CS2-BC2-CS3-BC3-CS4. In some embodiments, CS1 and CS4 are longer than CS2 and CS3. In some embodiments, CS2 and CS3 are 4-20 bases, CS1 and CS4 are 5 or 10 to 40 bases, for example 20-30 bases, and the length of the BC sequence is 4-20 bases (for example 10 bases). In some embodiments, CS4 is complementary to the splint oligonucleotide. In some embodiments, the composition comprises a bridge oligonucleotide. In some embodiments, the composition comprises a bridge oligonucleotide, beads comprising a triple barcode as described above, and genomic DNA comprising a hybridization sequence having a region complementary to the bridge oligonucleotide.
[0126] 2.3 stLFR using two transposons
[0127] 2 pmol of Tn5-coupled transposon was inserted into 40 ng of genomic DNA in a 60 μL reaction of 10 mM TAPS-NaOH (pH 8.5), 5 mM MgCl2, and 10% DMF at 55°C for 10 minutes. 1.5 μL of transposon-inserted DNA was transferred to 248.5 μL of hybridization buffer consisting of 50 mM Tris-HCl (pH 7.5), 100 mM MgCl2, and 0.05% 20 composition. 10-50 million barcoded beads are resuspended in the same hybridization buffer. The diluted DNA is added to the barcoded beads, and the mixture is heated to 60°C for 10 minutes with occasional gentle mixing. The DNA-bead mixture is transferred to a tube revolver in a laboratory oven and incubated at 45°C for 50 minutes. 500 μL of a ligation mixture comprising 50mM Tris-HCl (pH 7.8), 10mM DTT, 1mM ATP, 2.5% PEG-8000 and 4000 units of T4 ligase is added directly to the DNA-bead mixture. The ligation reaction is incubated on a rotator at room temperature for 1 hour. 110 μL of 1% SDS is added, and the mixture is incubated at room temperature for 10 minutes to remove the Tn5 enzyme. The beads are collected on the side of the tube via a magnet and washed once with low salt wash buffer and once with NEB2 buffer (New England Biolabs, Ipswich, MA). Excess barcode oligonucleotides were removed using 10 units of UDG (MA, Ipswich, New England Biolabs, MA) and 30 units of APE1 (New England Biolabs, Ipswich, MA) and 40 units of Exonuclease 1 (New England Biolabs, Ipswich, MA) in 100 uL 1X NEB2 buffer. The reaction was incubated at 37°C for 30 minutes. The beads were collected on the side of the tube and washed once with low salt wash buffer, then washed once with 1X PCR buffer (1X PfuCx buffer (Agilent Technologies, La Jolla, CA), 5% DMSO, 1M Betaine, 6mM MgSO4 and 600 μM dNTPs). The PCR mixture containing 1X PCR buffer, 400 pmol of each primer and 6 μL PfuCx enzyme (Agilent Technologies, La Jolla, CA) was heated to 95°C for 3 minutes and then cooled to room temperature. This mixture was used to resuspend the beads, and the combined mixture was incubated at 72°C for 10 minutes, followed by 12 cycles of 95°C for 10 seconds, 58°C for 30 seconds, and then 72°C for 2 minutes.
[0128] 2.4 stLFR with 3' branched ligation adapter
[0129] The method starts with identical hybridization insertion conditions, but only uses one transposon instead of two transposons. As mentioned above, after capture and barcode connection step, beads are collected to the side of pipe and washed with low salt wash buffer. The adapter digestion mixture of 90 units of exonuclease I (New England Biolabs, Ipswich, MA) and 100 units of exonuclease III (New England Biolabs, Ipswich, MA) in 100 μ L 1X TA buffer (Teknova, Hollister, CA) is added to beads and incubated at 37 DEG C for 10 minutes. The reaction is terminated, and the Tn5 enzyme is removed by adding 1% SDS of 11 μ L. Beads are collected to the side of pipe, and washed once with low salt wash buffer, then washed once with 1X NEB2 buffer (New England Biolabs, Ipswich, MA). Excess capture oligonucleotides were removed by adding 10 units of UDG (New England Biolabs, Ipswich, MA) and 30 units of APE1 (New England Biolabs, Ipswich, MA) in 100 uL of 1X NEB2 buffer (New England Biolabs, Ipswich, MA) and incubating at 37°C for 30 minutes. The beads were collected on the side of the tube and washed once with high salt wash buffer and once with low salt wash buffer. 300 pmol of the second adaptor was ligated to the beads of the bound subfragment with 4000 units of T4 ligase in 100 uL of 50 mM Tris-HCl (pH 7.8), 10 mM MgCl , 0.5 mM DTT, 1 mM ATP, and 10% PEG-8000 at room temperature on a rotator for 2 hours. The beads were collected on the side of the tube and washed once in high salt wash buffer and once in 1X PCR buffer. The PCR mixture and conditions were the same as for the two transposon method described above.
[0130] Exemplary 3' branched ligation adapters include 3' branched ligation adapter-F ( / 5Phos / CTGATGGCGCGAGGGAGGC) and 3' branched ligation adapter-R (TCGCGCCATCA / 3'dd / G) oligonucleotides as shown in Table 6. For example, the adapter F sequence includes a PCR primer annealing sequence. Optionally, a barcode (e.g., a sample barcode) can be included between the 5' phosphate and the sequence shown. In this example, the adapter R sequence is shorter than the primer annealing sequence so that it will melt under conditions of PCR primer annealing.
[0131] 2.5 Sequence Mapping and Variant Calling
[0132] First, the raw read data were demultiplexed by their associated barcode sequences using the barcode segmentation tool (available from GitHub https: / / github.com / stLFR / stLFR_read_demux). The barcode-assigned and sheared reads were mapped to the hs37d5 reference genome using BWA-MEM (2). The resulting BAM files were then sorted by chromosome coordinates using SAMtools (3), and duplicates were marked using the Picard MarkDuplicate function (http: / / broadinstitute.github.io / picard). Short variant (SNP and indel) calling was performed using HaplotypeCaller in GATK4.0.3.0 (4). The vcf files generated in the above steps were then benchmarked against the Genome in a Bottle (GIAB) high-confidence variant list (ftp: / / ftp-trace.ncbi.nlm.nih.gov / giab / ftp / release / NA12878_HG001 / latest / GRCh37 / ) (5) using the rtgtools vcfeval function (6). After benchmarking, the stLFR libraries were analyzed using GATK VariantRecalibrator and Gaussian mixture models were trained using the GIAB truth set. The VCFs were then filtered using GATK ApplyVQSR. In almost all cases, 99.9 tranches were applied to the original vcfs, with the exception of the 100Gb stLFR-1 library and the STD library, where 100 tranches were applied. We then established and applied further hard filtering criteria based on the GQ score, the ratio of reference depth to alternative depth, and barcode support listed in Table 5:
[0133] 2.6 Variant phasing using Hapcut2
[0134] SNPs were phased using Hapcut2 (https: / / github.com / vibansal / HapCUT2)(7) using its 10XGenomics data pipeline. The BAM files were first converted to a format with barcode information similar to the 10X Genomics barcode BAM. Specifically, a "BX" field was added to each line to reflect the barcode information of that read. GIAB variants or variants called by GATK for each library were used as input for phasing, and the calculate_haplotype_statistics.py tool of Hapcut2 was used to summarize the phasing results and compare them with the GIAB phased vcf files (5).
[0135] 2.7.LongHap
[0136] The seed extension strategy is used during the phasing process of LongHap. It initially starts with a pair of seeds consisting of the most upstream heterozygous variants in the chromosome. The seeds are extended by connecting other downstream candidate variants until no more variants can be added to the extended seeds ( Figure 11 ). In this expansion process, candidate variants located at different loci will not be treated equally (i.e., upstream variants have higher priority than downstream variants across the chromosome). Every two heterozygous loci have two possible combinations along two different alleles. Take the variants T2 / G2 and G3 / C3 as an example ( Figure 11 ), one combination pattern is T2-G3 and G2-C3, and the other is T2-C3 and G2-G3. The score of each combination is calculated by the number of long DNA fragments spanning the two loci, which is equal to the number of unique barcodes with reads mapped to the two loci. Figure 11 As shown, the final score of the former combination is 3, which is three times that of the latter. The variant T2 / G2 is added to the extended seed and the process is repeated. It is worth noting that if any barcode supports two alleles at a specific locus, it will be ignored when calculating the connection score. This helps to reduce the switching error rate. When conflicts occur when connecting downstream candidate variants, such as Figure 11 As shown in variants A4 / C4, a simple decision will be made by comparing the number of linked loci to allow further expansion of candidate variants. In this case, there are two linked loci in the left scenario, while there is only one in the right scenario. LongHap will select the left combined pattern as the final phasing result.
[0137] 2.8SV Detection
[0138] As described above, structural variation is detected by counting shared barcodes between genomic regions. Duplicate reads are first removed. Mapped co-barcoded reads are scanned using a sliding window along the genome (the default value is 2kb). For each window, the number of barcodes found in this 2kb window is recorded, and the Jaccard index is calculated for the shared barcode ratio between window pairs. Structural variation events are identified by the Jaccard index sharing metric between window pairs.
[0139] For each window pair (X, Y) on the genome, the Jaccard index is calculated as follows: X=(x1,x2,...x n ); Y=(y1,y2,...y n ) Jaccard Index
[0140] 2.9 Contig scaffolding using SALSA
[0141] Sequencing reads from the stLFR library were used to support the NA12878 assembly containing 18,903 contigs with an NG50 of 26.83 Mb (9) (contigs downloaded from the NCBI genome website using the scaffolding program SALSA (10)). To mimic the HiC sequence structure suitable for SALSA, stLFR sequence reads were selected from fragments of size >= 5 kb. From each fragment of length >= 5 kb, the "first" and "last" reads were selected to form a read pair. These artificial read pairs were then selected by shifting the fragments inward at intervals of 2 kb. These read pairs were then mapped to the NA12878 contigs and scaffolded using SALSA. The resulting scaffolds were then aligned and compared to the hg19 reference genome using nucmer and dnadiff from the MUMmer 4 program (11).
[0142] 2.10 References for Example 1
[0143] 1.I.Agent Technologies,RecoverEase DNA Isolation Kit.Revision C.0,(2015).
[0144] 2.H.Li,R.Durbin,Fast and accurate short read alignment with Burrows-Wheeler transform.Bioinformatics 25,1754-1760(2009).
[0145] 3.H.Li et al.,The Sequence Alignment / Map format andSAMtools.Bioinformatics 25,2078-2079(2009).
[0146] 4.A.McKenna et al.,The Genome Analysis Toolkit:a MapReduce frameworkfor analyzing next-generation DNA sequencing data.Genome Res 20,1297-1303(2010).
[0147] 5.J.M.Zook et al.,Integrating human sequence data sets providesaresource of benchmark SNP and indel genotype calls.Nat Biotechnol 32,246-251(2014).
[0148] 6.J.G.Cleary et al.,Comparing Variant Call Files for PerformanceBenchmarking of Next-Generation Sequencing Variant Calling Pipelines.bioRxiv,(2015).
[0149] 7.P.Edge,V.Bafna,V.Bansal,HapCUT2:robust and accurate haplotypeassembly for diverse sequencing technologies.Genome Res 27,801-812(2017).
[0150] 8.F.Zhang et al.,Haplotype phasing of whole human genomes using bead-based barcode partitioning in a single tube.Nat Biotechnol 35,852-857(2017).
[0151] 9. M. Pendleton et al., Assembly and diploid architecture of anindividual human genome via single-molecule technologies. Nat Methods 12, 780-786 (2015).
[0152] 10.J.Ghurye,M.Pop,S.Koren,D.Bickhart,CSChin,Scaffolding of longread assemblies using long range contact information.BMC Genomics 18,527(2017).
[0153] 11.S.Kurtz et al.,Versatile and open software for comparing largegenomes.Genome biology 5,R12(2004).
[0154] Example 2: Detailed experimental protocol
[0155] 3.1 Materials
[0156] 1Kb Plus DNA Ladder (ThermoFisher, Cat. No. 10787018) 100Kd MWCO Biotech CE dialysis tubing (Spectrum Labs, Cat. No. 131486) 384-well Armadillo PCR plate (ThermoFisher, Cat. No. AB2384) AMPure XP beads (Beckman Coulter, Cat. No. A63882) APE 1 (10,000 units / mL) (New England Biolabs, catalog number M0282L) ATP (100 mM) (Teknova, Catalog No. A1210) Barcoded bead construction oligonucleotides (IDT) (see Notes) Betaine (5M) (Sigma-Aldrich, Catalog No. B0300-5VL) BSA (20 mg / mL) (New England Biolabs, Cat. No. B9000S) Universal Adapter Oligonucleotide (IDT) DMF (~100%) (Sigma-Aldrich, Product No. D4551-250ML) DMSO (100%) (Sigma-Aldrich, Catalog No. D9170-5VL) dNTPs (25 mM) (ThermoFisher, Cat. No. R1121) Dialysis tubing (1000 kD MWCO) (Spectrum Laboratories, Inc., Cat. No. 131486) DTT (Sigma-Aldrich, Catalog No. 11583786001) Dynabeads TM M-280 Streptavidin (ThermoFisher, Cat. No. 60210) EDTA (0.5 M, pH 8.0) (Sigma-Aldrich, Cat. No. 03690-100ML) Exonuclease I (20,000 units / mL) (New England Biolabs, cat. no. M0293L) Exonuclease III (100,000 units / mL) (New England Biolabs, cat. no. M0206L) Formamide (100%, 250 mL) (Sigma-Aldrich, Product No. 47671-250ML-F) Glycerol (100%) (Sigma-Aldrich, Product No. G5516-100ML) KCl (Sigma-Aldrich, Product No. P9333-1KG) KH2PO4 (Sigma-Aldrich, Catalog No. 795488-1KG) KOH (Sigma-Aldrich, Product No. P5958-1KG) MgCl2 (1M) (Sigma-Aldrich, Cat. No. 63069-500ML) MgSO4 (1M) (Sigma-Aldrich, Product No. M3409-100ML) MicroAmp clear adhesive plate sealing film (ThermoFisher, Catalog No. 4306311) NaCl (5M) (ThermoFisher, Catalog No. AM9760G) Na2HPO4 (Sigma-Aldrich, Product No. S7907-1KG) NaOH (10M) (Sigma-Aldrich, Catalog No. 72068-100ML) NEB2 buffer (10X) (New England Biolabs, cat. no. B7002S) PEG-8000 (50%) (Rigaku, cat. no. 1008063) Pfu Turbo Cx Hotstart DNA Polymerase (Agilent, Cat. No. 600414) Proteinase K, recombinant, PCR-grade solution (14-22 mg / mL) (Roche, Catalog No. 03115844001); RiboRuler Low Range RNA Ladder (Thermofisher, Catalog No. SM1831); RNase-IT Ribonuclease Mix (Agilent, Catalog No. 400720) SDS (10%) (ThermoFisher, cat. no. 15553027) Sucrose (Sigma-Aldrich, Catalog No. S7903-1KG) T4 DNA ligase (2x10 6 Units / mL) (New England Biolabs, Catalog No. M0202M) TA buffer (10X) (Teknova, Catalog No. T0379) TAPS-NaOH (1 M, pH 8.5) (Boston BioProducts, Catalog No. BB-2375) TBE (10X) (ThermoFisher, Catalog No. 15581028) TE buffer (10X) (Fisher Scientific, Cat. No. BP24771) Tn5 enzyme Induced Transposon Oligonucleotides (IDTs) Tris-HCl (1 M, pH 7.5) (ThermoFisher, Cat. No. 15567027) Tris-HCl (2M, pH 7.8) (Amresco, Cat. No. J837-500ML) Triton TM X-100 (10%) (Sigma-Aldrich, Cat. No. 93443-100ML) 20 (10%) (Roche, Cat. No. 11332465001) UDG (5,000 units / mL) (New England Biolabs, Cat. No. M0280L)
[0157] 3.2 Equipment
[0158] 2.4 L tall polystyrene container (Click Clack, Cat. No. 659030) or equivalent DynaMag TM -2 magnet (ThermoFisher, Cat. No. 12321D) Easy 50EasySep TM Magnet (Stem Cell Technologies, Cat. No. 18002) or equivalent Laboratory oven that can accommodate a tube rotator / rotator Magnetic plate stirrer Medium magnetic stirring bar Standard laboratory vortex mixer Tetrad PCR Thermal Cycler (Bio-Rad, Cat. No. PTC0240) or equivalent, with a reaction volume of up to 100 μL per well Tube rotator / rotator (Thermo Fisher, Cat. No. 88881001) or equivalent
[0159] 3.3 Reagent Setup
[0160] Annealing buffer (3X) 3 mL of 1 M Tris-HCl, pH 7.5 6 mL of 5 M NaCl 91 mL of sterile dH2O Store at room temperature for 1 year.
[0161] 3.4 Buffer D (10X) 224 mg of KOH 50 μL of 0.5 M EDTA 2.45 mL of sterile dH2O Aliquot and store at -20°C for 1 month.
[0162] 3.5 Coupling buffer (1X) 5 mL of 1X TE 5 mL of 100% glycerol Store at -20°C for 1 year.
[0163] 3.6 Digestion buffer (1X, pH 8.0) 1.75g Na2HPO4 0.2 g of KCl 0.2g KH2PO4 27.4 mL 5 M NaCl 20 mL of 0.5 M EDTA (pH 8.0) 800 mL of sterile dH2O Adjust the pH to 8.0 with 1 M NaOH. Add sterile dH2O to a final volume of 1 L. Filter sterilize. Store at room temperature for 1 year.
[0164] 3.7 3' branch ligation buffer (3X) 6 mL of 50% PEG-8000 0.75 mL 2 M Tris-HCl (pH 7.8) 0.3 mL 1 M MgCl2 0.3 mL 0.1 M ATP 15 μL 1M DTT 75 μL 20 mg / mL BSA 2.560 mL of sterile dH2O Store at -20°C for 1 year.
[0165] 3.8 High Salt Bead Binding Buffer (1X) 5 mL of 1 M Tris-HCl (pH 7.5) 6 mL of 5 M NaCl 20 μL 0.5 M EDTA 88.98 mL of sterile dH2O Store at room temperature for 1 year.
[0166] 3.9 High Salt Wash Buffer (1X) 5 mL of 1 M Tris-HCl, pH 7.5 10 mL of 5 M NaCl 20 μL 0.5 M EDTA 10% of 0.5mL 20 84.48 mL of sterile dH2O Store at room temperature for 1 year.
[0167] 3.10 Hybridization buffer (1X) 50 mL of 1 M Tris-HCl, pH 7.5 100 mL of 1 M MgCl2 10% of 5mL 20 845mL water Store at room temperature for 1 year.
[0168] 3.11 Ligation buffer (10X) 25 mL of 50% PEG-8000 12.5 mL of 2 M Tris-HCl (pH 7.8) 5 mL of 100 mM ATP 5 mL of 1 M MgCl2 2.5 mL of sterile dH2O Store at -20°C for 1 year.
[0169] 3.12 Ligation buffer, MgCl2-free (10X) 25 mL of 50% PEG-8000 12.5 mL of 2 M Tris-HCl (pH 7.8) 5 mL of 100 mM ATP 5 mL of 1M DTT 2.5 mL of sterile dH2O Store at -20°C for 1 year.
[0170] 3.13 Low Salt Wash Buffer (1X) 5 mL of 1 M Tris-HCl, pH 7.5 3 mL of 5 M NaCl 10% of 0.5mL 20 91.5 mL of sterile dH2O Store at room temperature for 1 year.
[0171] 3.14 Lysis buffer (1X, pH 8.3) 0.22 g of KCl 120g sucrose 13 mL of 1 M Tris-HCl (pH 7.5) 2 mL of 0.5 M EDTA (pH 8.0) 28 mL of 5 M NaCl 10mL X-100 800 mL of sterile dH2O Adjust pH to 8.3 Add sterile dH2O to a final volume of 1 L. Filter sterilize. Store at 4°C for 1 year.
[0172] 3.15 Transposase Buffer (5X) 0.5 mL of 1 M TAPS-NaOH (pH 8.5) 0.25 mL of 1 M MgCl2 5 mL of 100% DMF 4.25 mL of sterile dH2O Store at -20°C for 1 year.
[0173] 3.16PfuCx mixture (2X) 2 mL of 10X PfuCx buffer (contains enzyme) 0.5 mL of 100% DMSO 2 mL of 5 M betaine 60 μL of 1M MgSO4 240 μL of 25 mM dNTPs 5.2 mL of sterile dH2O
[0174] 3.16 Barcoded Bead Construction Oligonucleotides
[0175] All barcoded oligonucleotides were synthesized in a 384-well format at a 100 nmol scale by standard desalting methods and delivered by Integrated DNA Technologies (Coralville, IA) at a concentration of 200 μM in 1X TE (pH 8.0). There are a total of 1536 unique barcode oligonucleotides per barcode set, and there are 3 barcode sets. This enables up to 3.6 billion different barcode combinations. This may not be necessary for some applications, and fewer barcode combinations can be achieved by ordering fewer oligonucleotide plates. This particular design does require the use of at least one barcode oligonucleotide from each set to produce an appropriate final sequence, however, slight modifications to the 6-base overlap between barcode sets can be made to remove entire barcode sets.
[0176] 3.2 Procedure
[0177] Isolation of high molecular weight DNA from cells
[0178] This method is based on RecoverEase TM DNA Isolation Kit protocol 26, but using much larger volumes to reduce the viscosity of the resulting solution.
[0179] 1. Precipitate up to 1 × 10 in a 15 or 50 mL conical tube. 7Disperse the nucleated cells (500 x g for 5 minutes). Remove the supernatant. Add 500 μL of lysis buffer to the cell pellet and briefly vortex the sample at medium speed for 3-5 seconds, then place the conical tube in the refrigerator for approximately 10 minutes with occasional swirling.
[0180] 2. Prepare Proteinase K solution by combining 250 μL of 10% SDS, 250 μL of Proteinase K, and 4 mL of 1X TE. Place on a heating block at 50°C and heat briefly ( 5 minutes).
[0181] 3. Prepare digestion solution by mixing 20 μL of RNase-It Ribonuclease Mix with 4 mL of Digestion Buffer.
[0182] 4. Add approximately 4 mL of the prepared digestion solution to the lysed cells and buffer from step 1 and gently shake the conical tube.
[0183] After 5.5 minutes, place the conical tube in a 50°C heat block and add 4.5 mL of heated Proteinase K solution to the free-floating pellet. Gently swirl the conical tube to mix.
[0184] 6. Recap the tube and incubate in a heat block at 50 °C for 2 h, gently swirling the tube every 30 min.
[0185] 7. Cut approximately 13 cm of dialysis tubing (its capacity is approximately 1 mL / cm). Allow to equilibrate in 0.5X TE for 30 minutes. Seal one end with a dialysis clip.
[0186] 8. Pour at least 1 L of 0.5X TE buffer into the dialysis cell.
[0187] 9. Carefully pour the viscous genomic DNA from the conical tube into the open end of the dialysis tubing. Seal the open end of the dialysis tubing with a dialysis clamp. Secure a float to one of the clamps. Place the dialysis tubing with the float in the dialysis tank.
[0188] 10. Dialyze the genomic DNA at room temperature for 24 to 48 hours while gently stirring the buffer with a magnetic stir bar. Change the TE buffer once during the dialysis to maximize the purity of the recovered DNA.
[0189] 11. After dialysis is complete, remove the dialysis tubing from the TE buffer, remove the float and clamp from the top of the dialysis tubing, and gently pour into a 15 mL conical tube. The DNA is ready for immediate use without shearing.
[0190] 3.3 Barcoded Beads
[0191] The barcoded beads were constructed using a split-and-merge strategy of three sets of double-stranded barcoded DNA molecules. Full-length adapters were constructed by successive ligation ( Figure 12 and Figure 13 Barcode oligonucleotides are provided in 384-well plates (see "Reagent Instructions"). Common adapter oligonucleotides are provided in tubes. Depending on the sequencing technology used, it may be necessary to vary the PCR primer sequences within the common adapter oligonucleotides.
[0192] 12. Combine 10 μL of complementary oligonucleotide from each well of the source 384-well plate with 10 μL of 3X Annealing Buffer in a 384-well PCR plate. Combine 30 μL of common adapter oligonucleotide in one well of an 8-well PCR strip.
[0193] 13. Incubate at 70°C for 3 minutes in a PCR thermocycler, then slowly increase the temperature to 20°C at a rate of 0.1°C / s. The final concentration of the hybridized barcode oligonucleotide is 66 μM.
[0194] 14. Mix 4.725 mL (157.5 μmol) of hybridized bead linker containing 5' bi-biotin with 3.225 mL of ligation buffer (10X), 460.8 μL (921,600 units) of T4 DNA ligase, and 9.67 mL of dH2O to a total volume of 18.081 mL.
[0195] 15. Dispense 11.2 μL of the ligation mix into each well of four new 384-well PCR plates. Then, add 8.8 μL (580 pmol) from each well of the first hybridized barcode plate to each well containing the bead-adapter ligation mix. Seal with MicroAmp clear adhesive sealer, vortex, centrifuge, and incubate at room temperature for 1 hour.
[0196] 16. Collect 100 billion (143 mL) of M-280 Streptavidin-coated magnetic beads by transferring 50 mL of beads into an empty 50 mL centrifuge tube. Place the 50 mL tube with beads in the Easy 50 EasySep TM Place on the magnet for 5 minutes to collect the beads to the side of the tube. Carefully remove the supernatant with a pipette. Transfer the second 50 mL of beads to the tube on the magnet. Let it sit on the magnet for 5 minutes, then carefully remove the supernatant. Transfer the final 43 mL of beads to the 50 mL tube. Let it sit on the magnet for 5 minutes, then carefully remove the supernatant. Wash the beads twice with low-salt wash buffer, then thoroughly resuspend in 8 mL of high-salt bead binding buffer.
[0197] 17. Dispense 5 μL of beads in high salt bead binding buffer into each well of the plate containing the ligated product. During the dispensing process, occasionally vortex the bead source tube to keep the beads well suspended.
[0198] 18. Seal the plate with MicroAmp clear adhesive sealer, vortex, and incubate on a tube rotator in "shake" mode at room temperature for 1 hour.
[0199] 19. Centrifuge the plate at 300 x g for 5 seconds to remove the beads from the seal, but do not allow pelleting. Remove the seal and add 2.8 μL of 0.1% SDS to each well. Seal the plate again with MicroAmp clear adhesive sealer, vortex briefly, and incubate at room temperature for 10 minutes.
[0200] 20. Vortex, then centrifuge the plate at 300 x g for 5 seconds to remove the beads from the plate seal. Remove the seal from each plate and invert the plate onto a collection plate. Centrifuge at 500 x g for 2 minutes. Using a 10 ml serological pipette, collect the beads into a new 50 ml tube.
[0201] 21. Collect the beads into Easy 50EasySep TM Magnetize the tube sideways for 5 minutes. Discard the supernatant. Wash once with 10 mL of high-salt wash buffer, then wash twice with low-salt wash buffer. Resuspend the beads in 8 mL of 1X ligation buffer.
[0202] 22. Dispense 5 μl of beads into each well of four new 384-well PCR plates. During the dispensing process, occasionally vortex the bead source tube to keep the beads well suspended.
[0203] 23. To ligate the second set of barcodes, bring the total volume to 10.02 mL with a mixture containing 3.225 mL of 10X ligation buffer, 460.8 μL (921,600 units) of T4 DNA ligase, and 6.33 mL of dHO. Dispense 6.2 μL of the second ligation mixture into each well of four 384-well PCR plates containing beads. Next, add 8.8 μL (580 pmol) from each well of the hybridized second barcode plate to the corresponding well of the 384-well PCR plate containing beads and ligation mixture.
[0204] 24. Repeat steps 18-22.
[0205] 25. To ligate the third set of barcodes, bring the total volume of the ligation mix to 10.02 mL, containing 3.225 mL of ligation buffer (10X), 460.8 μL (921,600 units) of T4 DNA ligase, and 6.33 mL of dHO. Dispense 6.2 μL of the third ligation mix into each well of four 384-well PCR plates containing beads. Next, add 8.8 μL (580 pmol) from each well of the hybridized third barcode plate to the corresponding well of the 384-well PCR plate containing beads and ligation mix.
[0206] 26. Repeat steps 18-22. The beads can now be stored at 4°C for up to one year. In their current form, the beads are almost completely double-stranded and not yet in the correct form for stLFR.
[0207] 27. Count the beads using a hemocytometer and remove 5 million beads for the QC step. Place the tube with beads in the DynaMag TM -2 magnet for 5 minutes. Discard the supernatant. Add 5 μL 100% formamide, 4 μL dH2O, and 1 μL 10X loading buffer. Incubate at 95°C for 3 minutes on a PCR thermocycler. Immediately place on ice for 2 minutes. Place the tube with the beads on a DynaMag TM -2 magnet for 5 minutes. Collect the supernatant, load onto a 15% TBU gel, and run at 200V for 40 minutes to check the length and quantity of the oligonucleotides. Alternatively, the beads can be examined using flow cytometry by hybridizing fluorescently labeled oligonucleotides to the 3' end of the bead adapter sequence. We typically see approximately 25% of the total streptavidin binding sites harbor the full-length constructed adapter sequence.
[0208] 3.20 stLFR Bead Preparation
[0209] To prepare beads for stLFR, they must first be denatured into single-stranded DNA and then rehybridized with the bridging oligonucleotide.
[0210] 28. Pipette 500 million constructed barcoded beads from step 26 of the previous section into a standard 1.5 mL microcentrifuge tube.
[0211] 29. Place on DynaMag TM Place on a -2 magnet for 2 minutes to collect the beads on the side of the tube. Remove the supernatant.
[0212] 30. Add 1 mL of the 1X dilution of Buffer D. Vortex briefly and incubate at room temperature for 2 minutes.
[0213] 31. Place on DynaMag TMPlace on a -2 magnet for 2 minutes to collect the beads on the side of the tube. Remove the supernatant.
[0214] 32. Repeat steps 30 and 31 once more.
[0215] 33. Wash once in 1X annealing buffer. Place on DynaMag TM Place on a -2 magnet for 2 minutes to collect the beads on the side of the tube. Remove the supernatant.
[0216] 34. Combine 36 μL of 100 μM Bridge Oligo, 333.33 μL of Annealing Buffer (3X), and 630.67 μL of dH2O to a final volume of 1 mL. Add the mixture to the beads and vortex briefly.
[0217] 35. Incubate at 60°C for 5 minutes and at room temperature for 50 minutes.
[0218] 36. Place on DynaMag TM Place on a -2 magnet for 2 minutes to collect the beads to the side of the tube. Remove the supernatant and resuspend in 500 μL of low-salt wash buffer. The beads are now ready for stLFR and can be stored at 4°C for 3 months.
[0219] 3.21 Two transposon stLFR experimental protocols
[0220] This protocol utilizes two transposons to generate hybridization sequences and PCR primer sites along the length of a genomic DNA molecule. This is the simplest and fastest stLFR method, but coverage of each long DNA fragment may be reduced by 50% compared to the 3' branch ligation protocol. For compatibility with sequencing technologies other than BGISEQ-500, it may be necessary to alter some transposon sequences after the chimeric region. Before ordering these oligonucleotides, check the sequencing primers used. Information on all oligonucleotide sequences is available in the supplementary material.
[0221] 37. Hybridize the captured transposon oligonucleotides by combining 10 μL of transposon 1T (100 μM), 10 μL of transposon B (100 μM), and 10 μL of annealing buffer (3X) into the first well of an 8-well PCR strip. Hybridize the non-captured transposon oligonucleotides by combining 10 μL of transposon 1T (100 μM), 10 μL of transposon B (100 μM), and 10 μL of annealing buffer (3X) into the second well of the same PCR strip.
[0222] 38. Incubate at 70°C for 3 minutes on a PCR thermocycler, then slowly increase the temperature to 20°C at 0.1°C / s. Combine the two transposons into the third well of the PCR strip.
[0223] 39. Couple the Tn5 enzyme to the transposon mix by combining 9.6 μL of the mixed transposon with 23.53 μL of Tn5 (13.6 pmol / μL) and 46.87 μL of coupling buffer (1X).
[0224] 40. Incubate at 30°C for 1 hour. Use immediately or store at -20°C for up to 1 month. For optimal performance and consistency between experiments, we recommend aliquoting before storage.
[0225] 41. Incorporate the transposon into the long genomic DNA by combining 12 μL of transposase buffer (5X), 0.5 μL of the coupled transposon from step 40, and 40 ng of DNA in one well of an 8-well tube strip for a total volume of 60 μL. Note: The amount of DNA and the amount of coupled transposon can be adjusted at this step. Because there may be variability between batches, it may be necessary to titrate the amount of Tn5 enzyme used. Again, it is possible to start with less DNA, but for titration purposes, it is useful to use 40 ng so that some material can be run on an agarose gel to determine the efficiency of transposon incorporation (see later steps).
[0226] 42. Incubate at 55°C for 10 minutes.
[0227] 43. Transfer 40 μL of the transposon-incorporated material to one well of a new 8-well tube strip, add 4 μL of 1% SDS, and incubate at room temperature for 10 minutes.
[0228] 44. Load the material from step 43 onto a 0.5X TBE 1% agarose gel and run at 150V for 40 minutes. The transposed DNA should appear between 200 and 1500 bp on the gel. We typically expect to see the brightest portion of the DNA smear around 600 bp, although this may vary depending on the sequencing technology chosen. We typically load a control that has undergone the same steps but lacks the transposon, Tn5 enzyme, or genomic DNA. If the transposon integration product appears to be the correct size, proceed to step 45. Otherwise, repeat the above steps, adjusting the concentration of the coupled product, until the smear reaches the desired size.
[0229] 45. Dilute 1.5 μL of the remaining product from step 42 with 248.5 μL of 1x hybridization buffer.
[0230] 46. Transfer 50 μL of beads (50 million) from step 36 to a 1.5 mL microcentrifuge tube. Place in DynaMag TM -2 magnet for 2 minutes to collect the beads to the side of the tube. Remove the supernatant and resuspend in 250 μL hybridization buffer (1X).
[0231] 47. Heat the DNA and beads separately at 60°C for 30 seconds.
[0232] 48. Add 250 μL of diluted DNA to 250 μL of beads, gently mix by tapping the bottom of the tube with your finger, and continue incubating at 60°C for 10 minutes. Gently mix the tube with your finger every few minutes.
[0233] 49. Place on a tube rotator and incubate in a 45°C oven on "shake" mode for 50 minutes.
[0234] 50. Prepare the ligation buffer by combining 100 μL of ligation buffer, 10X MgCl2-free, 2 μL of T4 DNA ligase (2x10 6 Prepare the ligation mix by adding 1 mL of 5% dHO (100 μL / mL) and 398 μL of dH 2 O. Remove the tube from the rotator and add the ligation mix to a total volume of 1 mL.
[0235] 51. Incubate at room temperature on a tube rotator in "shake" mode for 1 hour.
[0236] 52. Add 110 μL of 1% SDS to the tube and incubate at room temperature for 10 minutes.
[0237] 53. Place on DynaMag TM -2 magnet for 2 minutes to collect the beads to the side of the tube. Remove the supernatant and wash once with 500 μL low salt wash buffer and once with 500 μL NEB2 buffer (1X).
[0238] 54. Prepare the capture oligonucleotide digestion mix by combining 10 μL NEB2 buffer (10X), 2 μL UDG (5,000 U / mL), 3 μL APE1 (10,000 U / mL), 2 μL Exonuclease 1 (20,000 units / mL), and 83 μL dH 2 O. Remove the wash buffer and add the digestion mix to the beads.
[0239] 55. Vortex gently to resuspend the beads and incubate at 37°C for 30 minutes.
[0240] 56. Place on DynaMag TM -2 magnet for 2 minutes to collect the beads to the side of the tube. Remove the supernatant and wash once with 500 μL low salt wash buffer and once with 500 μL PfuCx buffer (1X).
[0241] 57. Prepare a PCR premix by adding 150 μL PCR mix (2X), 4 μL PCR primer 1 (100 μM), 4 μL PCR primer 2 (100 μM), 6 μL PfuCx enzyme, and 136 μL dH2O. Preheat the PCR premix at 95°C for 3 minutes. Place in a DynaMag TM Place on a -2 magnet for 2 minutes to collect the beads to the side of the tube. Remove the wash buffer and add the PCR master mix to the beads.
[0242] 58. Gently vortex to resuspend the beads and cycle the PCR reaction under the following conditions:
[0243] 59. PCR should yield approximately 500 ng of DNA. Run 20 ng of the product on a 0.5X TBE 1% agarose gel for 40 minutes at 150 V. The material should be a smear with a peak around 500 bp.
[0244] 60. Purify the PCR product using 300 μL of Agencourt XP beads according to the manufacturer's protocol. This purified product is now ready for sequencing.
[0245] Single transposon 3' branch ligation stLFR protocol
[0246] This protocol, based on single transposon insertion in DNA nicks and a novel adapter ligation method, can result in higher coverage per fragment, which can be important for certain sequencing strategies, such as de novo assembly. This strategy is slightly more expensive due to the addition of additional reagents and takes longer than 2.5 hours.
[0247] 61. Hybridize the captured transposon oligonucleotides by combining 10 μL of Transposon 1T (100 μM), 10 μL of Transposon B (100 μM), and 10 μL of Annealing Buffer (3X) into the first well of an 8-well PCR strip. Hybridize the gap ligation adapters by combining 10 μL of BranchT (100 μM), 10 μL of BranchB (100 μM), and 10 μL of Annealing Buffer (3X) into the second well of the same PCR strip.
[0248] 62. Incubate at 70°C for 3 minutes on a PCR thermocycler, then slowly increase the temperature to 20°C at a rate of 0.1°C / s.
[0249] 63. Couple the Tn5 enzyme to the transposon by combining 9.6 μL of the hybridization-captured transposon in step 61 with 23.53 μL of Tn5 (13.6 pmol / μL) and 46.87 μL of coupling buffer (1X).
[0250] 64. Incubate at 30°C for 1 hour. Use immediately or store at -20°C for up to 1 month.
[0251] 65. Execute steps 41-51.
[0252] 66. Place on DynaMag TM Place on a -2 magnet for 2 minutes to collect the beads to the side of the tube. Remove the supernatant and wash once with 500 μL of low salt wash buffer.
[0253] 67. Prepare the adapter oligonucleotide digestion mix by combining 10 μL TA buffer (10X), 4.5 μL Exonuclease I (20,000 U / mL), 1 μL Exonuclease III (100,000 U / mL), and 74.5 μL dH 2 O. Remove the wash buffer and add the digestion mix to the beads.
[0254] 68. Gently vortex to resuspend the beads and incubate on a tube rotator at 37°C for 10 minutes on the "shake" mode.
[0255] 69. Add 11 μL of 1% SDS and incubate at room temperature for 10 minutes.
[0256] 70. Place on DynaMag TM -2 magnet for 2 minutes to collect the beads to the side of the tube. Remove the supernatant and wash once with 500 μL low salt wash buffer and once with 500 μL NEB2 buffer (1X).
[0257] 71. Prepare capture oligonucleotide digestion mix by combining 10 μL NEB2 buffer (10X), 2 μL UDG (5000 U / mL), 3 μL APE1 (10000 U / mL), and 85 μL dH 2 O. Remove wash buffer and add digestion mix to beads.
[0258] 72. Gently vortex to resuspend the beads and incubate at 37°C for 30 minutes.
[0259] 73. Place on DynaMag TM -2 magnet for 2 minutes to collect the beads to the side of the tube. Remove the supernatant and wash once with 500 μL high salt wash buffer and once with 500 μL low salt wash buffer (1X).
[0260] 74. Prepare the 3' branch ligation buffer (3X) by combining 33.4 μL of 3' branch ligation adapter (16.7 μM) prepared in step 61, 2 μL of T4 DNA ligase (2x106 The 3' branch ligation mix was prepared by adding 46.6 μL of dH 2 O. The wash buffer was removed and the ligation mix was added to the beads.
[0261] 75. Gently vortex to resuspend the beads and incubate on a tube rotator at 25°C for 2 hours on the "shake" mode.
[0262] 76. Place on DynaMag TM -2 magnet for 2 minutes to collect the beads to the side of the tube. Remove the supernatant and wash once with 500 μL high salt wash buffer and once with 500 μL PCR buffer (1X).
[0263] 77. Prepare a PCR master mix by adding 150 μL of 2X PCR buffer, 4 μL of PCR primer 1 (100 μM), 4 μL of PCR primer 2 (100 μM), 6 μL of PCR enzyme, and 136 μL of dH 2 O. Remove the wash buffer and add the PCR master mix to the beads.
[0264] 78. Gently vortex to resuspend the beads and cycle the PCR reaction under the following conditions:
[0265] 79. Perform steps 59-60 above.
[0266] 3.4 Analysis of stLFR Data
[0267] 7 The starting point for this process is a FASTQ file. This is the standard format for read data generated by most sequencing technologies. The software we use to deconvolute the barcode information takes a FASTQ file and expects the 42 bases of the barcode and the universal adapter sequence to be appended to the end of the first read. It matches the barcoded read data to the expected 1536 sequences at each barcode position. The barcoding strategy used by stLFR can error-correct barcodes with single-base mismatches. The final output of our software is a FASTQ file with the barcode information appended to the end of the read ID in the format #Barcode1ID_Barcode2ID_Barcode3ID, where BarcodeID is a number ranging from 0 to 1536. A barcode ID of zero indicates that it does not match any expected barcode sequence. We recommend using BWA-mem27 for mapping, GATK28 for variant calling, and HapCUT229 for phasing. We also recommend using a decoy sequence to map to Hg19.
[0268] 3.5 References for Example 2
[0269] 1Zhang,K.et al.Long-range polony haplotyping of individual humanchromosome molecules.Nat Genet 38,382-387(2006).
[0270] 2Ma,L.et al.Direct determination of molecular haplotypes bychromosome microdissection.Nat Methods 7,299-301(2010).
[0271] 3Kitzman,J.O.et al.Haplotype-resolved genome sequencing of aGujaratiIndian individual.Nat Biotechnol 29,59-63(2011).
[0272] 4Suk,E.K.et al.A comprehensively molecular haplotype-resolved genomeof a European individual.Genome Res 21,1672-1685(2011).
[0273] 5Fan,H.C.,Wang,J.,Potanina,A.&Quake,S.R.Whole-genome molecularhaplotyping of single cells.Nat Biotechnol 29,51-57(2011).
[0274] 6Peters,B.A.et al.Accurate whole-genome sequencing and haplotypingfrom 10to 20human cells.Nature 487,190-195(2012).
[0275] 7Duitama,J.et al.Fosmid-based whole genome haplotyping of aHapMaptrio child:evaluation of Single Individual Haplotyping techniques.NucleicAcids Res 40,2041-2053(2012).
[0276] 8Selvaraj,S.,J,R.D.,Bansal,V.&Ren,B.Whole-genome haplotypereconstruction using proximity-ligation and shotgun sequencing.Nat Biotechnol31,1111-1118(2013).
[0277] 9Kuleshov,V.et al.Whole-genome haplotyping using long reads andstatistical methods.Nat Biotechnol 32,261-266(2014).
[0278] 10Amini,S.et al.Haplotype-resolved whole-genome sequencing bycontiguity-preserving transposition and combinatorial indexing.Nat Genet 46,1343-1349(2014).
[0279] 11Zheng,G.X.et al.Haplotyping germline and cancer genomes with high-throughput linked-read sequencing.Nat Biotechnol(2016).
[0280] 12Zhang,F.et al.Haplotype phasing of whole human genomes using bead-based barcode partitioning in a single tube.Nat Biotechnol 35,852-857(2017).
[0281] 13 Peters,B.A.,Liu,J.&Drmanac,R.Co-barcoded sequence reads fromlongDNA fragments:a cost-effective solution for"perfect genome"sequencing.Frontiers in genetics 5,466(2014).
[0282] 14 Drmanac,R.Nucleic Acid Analysis by Random Mixtures of Non-Overlapping Fragments.WO 2006 / 138284 A2(2006).
[0283] 15 McElwain,M.A.,Zhang,R.Y.,Drmanac,R.&Peters,B.A.LongFragment Read(LFR)Technology:Cost-Effective,High-Quality Genome-WideMolecularHaplotyping.Methods Mol Biol 1551,191-205(2017).
[0284] 16 Schaaf,C.P.et al.Truncating mutations of MAGEL2 cause Prader-Williphenotypes and autism.Nat Genet 45,1405-1408(2013).
[0285] 17 Peters,B.A.et al.Detection and phasing of single base denovomutations in biopsies from human in vitro fertilized embryos by advancedwhole-genome sequencing.Genome Res 25,426-434(2015).
[0286] 18 Ciotlos,S.et al.Whole genome sequence analysis of BT-474usingcomplete Genomics'standard and long fragment readtechnologies.Gigascience 5,8(2016).
[0287] 19 Hellner,K.et al.Premalignant SOX2 overexpression in thefallopiantubes of ovarian cancer patients:Discovery and validationstudies.EBioMedicine10,137-149(2016).
[0288] 20 Mao,Q.et al.The whole genome sequences and experimentallyphasedhaplotypes of over 100 personal genomes.Gigascience 5,1-9(2016).
[0289] 21 Gulbahce,N.et al.Quantitative Whole Genome SequencingofCirculating Tumor Cells Enables Personalized Combination Therapy ofMetastaticCancer.Cancer Res 77,4530-4541(2017).
[0290] 22 Walker,R.F.et al.Clinical and genetic analysis of a raresyndromeassociated with neoteny.Genetics In Medicine(2017).
[0291] 23 Mao,Q.et al.Advanced Whole-Genome Sequencing and Analysis ofFetalGenomes from Amniotic Fluid.Clinical chemistry(2018).
[0292] 24 Drmanac,R.,Peters,B.A.,Alexeev,A.Multiple tagging of individuallong DNA fragments.WO 2014 / 145820 A2(2013).
[0293] 25Picelli,S.et al.Tn5 transposase and tagmentation procedures formassively scaled sequencing projects.Genome Res 24,2033-2040(2014).
[0294] 26Agent Technologies,I.RecoverEase DNA Isolation Kit.Revision C.0(2015).
[0295] 27Li,H.&Durbin,R.Fast and accurate short read alignment with Burrows-Wheeler transform.Bioinformatics 25,1754-1760(2009).
[0296] 28McKenna,A.et al.The Genome Analysis Toolkit:a MapReduce frameworkfor analyzing next-generation DNA sequencing data.Genome Res 20,1297-1303(2010).
[0297] 29Edge,P.,Bafna,V.&Bansal,V.HapCUT2:robust and accurate haplotypeassembly for diverse sequencing technologies.Genome Res 27,801-812(2017).
[0298] The BAM file "NA12878_WGS_v2_phased_possorted_bam.bam" from the recent Chromium dataset was downloaded from the 10X Genomics website and processed in the same manner as the stLFR library. For filtering, we used the VCF file "NA12878_WGS_v2_phased_variants.vcf.gz" from the same Chromium library. This VCF contains data processed through the 10X Genomics optimized pipeline. Fragment sizes for the Chromium library were copied from the 10X Genomics website. 10Genomics calculates fragment sizes using a length-weighted average, which may be larger than the average fragment size. 2 No read data are available yet, as reported by Zhang et al. (12). 3 Data from a standard library processed on BGISEQ-500.
[0299] Table 6 shows exemplary sequences that can be used in the stLFR methods described herein.
[0300] Example 3: 3' branch ligation, a novel method for ligating DNA to the 3'OH terminus of DNA or RNA and its applications
[0301] 4.1 Introduction
[0302] This example generally describes 3' branch ligation. In the stLFR embodiments described herein, 3' branch ligation is used to add additional adaptors (3' branch ligation adaptors). See, eg, §1.1.2.
[0303] Ligases connect breaks in nucleic acids, which is essential for cell viability and vitality. DNA ligases catalyze the formation of phosphodiester bonds between DNA ends and play a key role in DNA repair, recombination, and replication in vivo. RNA ligases join 5'-phosphodiester (5'PO4) and 3'-hydroxyl (3'OH) RNA ends via phosphodiester bonds and participate in RNA repair, splicing, and editing. Organisms from all three kingdoms (bacteria, archaea, and eukaryotes) can be used as important molecular tools in vitro for applications such as cloning, ligase-based amplification or detection, and synthetic biology.
[0304] One of the most widely used ligases in vitro is bacteriophage T4 DNA ligase, a single 55-kDa polypeptide that requires ATP as an energy source. T4 DNA ligase typically ligates adjacent 5'PO4 and 3'OH termini of double-stranded DNA. In addition to sealing nicks or ligating sticky ends, T4 DNA ligase can also efficiently catalyze blunt-end ligation, a process not seen in all other DNA ligases. Several unusual catalytic properties of this ligase have been previously reported, including sealing single-stranded gaps in double-stranded DNA, sealing nicks near abasic sites in double-stranded DNA (dsDNA), promoting intramolecular loop formation in partially double-stranded DNA, and ligating DNA strands containing 3' branch extensions. (Nilsson and Magnusson, Nucleic Acids Res 10:1425-1437, 1982; Goffin et al., Nucleic Acids Res 15:8755-8771, 1987; Mendel-Hartvig et al., Nucleic Acids Res. 32:e2, 2004; Western and Rose, Nucleic Acids Res.,19:809-813,1991). Researchers have also observed template-independent ligations mediated by T4 ligases, such as sealing of mispaired nicks in dsDNA (Alexander, 2003, Nucleic Acids Res. 2003 Jun 15; 31(12): 3208-16) and even single-stranded DNA (ssDNA) ligation, although with very low efficiency (H. Kuhn, 2005, FEBS J. 2005 Dec; 272(23): 5991-6000). These results suggest that perfect complementary base pairing at or near the ligation junction is not critical for some unconventional T4 DNA ligase activities. T4 RNA ligases 1 and 2 are products of T4 bacteriophage genes 63 and 24, respectively. Both require adjacent 5'PO4 and 3'OH termini to successfully hydrolyze ATP to AMP and PPi. The substrates of T4 RNA ligase 1 include single-stranded RNA and DNA, while T4 RNA ligase 2 preferentially seals nicks on dsRNA rather than ligating the ends of ssRNA.
[0305] Here, we demonstrate an unconventional end-joining event mediated by T4 DNA ligase, which we term 3'-branched ligation (3'BL). It can join DNA or DNA / RNA fragments at nicks, single-stranded gaps, or 5'-overhang regions to form branched structures. This report extensively investigates various ligation cofactors and activators and optimizes the ligation conditions for this novel ligation. Using our 3'BL protocol, base pairing is not required, and ligation is achieved with over 90% completeness, even in the presence of 1nt gaps. One application is the ligation of adapters to DNA or RNA in next-generation sequencing (NGS) library preparation. Several genomic structures previously considered unligatable are now substrates for 3'BL, resulting in high conversion rates from input DNA to adapter-ligated molecules while avoiding chimeras. We demonstrate that 3'BL can be combined with transposon insertion. Our proposed targeted transposon insertion strategy can theoretically generate 100% sequence-usable templates. Applications include microRNAs. Our studies demonstrate the value of this novel technology for NGS library preparation and its potential to facilitate numerous other molecular applications, such as radiolabeling the 3' ends of RNA.
[0306] 4.2 3' branch ligation, a novel method for joining DNA ends
[0307] Typically, DNA ligation involves joining the 5'PO4 and 3'OH DNA ends of sticky or blunt-ended fragments. Sticky-end ligation is generally faster than blunt-end ligation and is less dependent on enzyme concentration. Bacteriophage T4 DNA ligase can catalyze both processes, but bacteriophage T4 DNA ligase uses ATP as a cofactor for energy production and requires Mg. 2+ . T4 DNA ligase has also been reported to ligate specific or degenerate single-stranded oligonucleotides to partially single-stranded substrates by hybridization. Here, we demonstrate unprecedented T4 DNA ligase-mediated ligation that does not require complementary base pairing and can ligate a blunt-ended DNA donor to the 3'OH terminus of a double-stranded DNA acceptor at a nick, a gap, or a 5'-overhang to form a branched structure (Figure 21a). Therefore, we use the term 3'-branched ligation (3'BL) to describe these ligations. The synthetic donor DNA we used contains a blunt-ended duplex end and an ssDNA end. The acceptor substrate contains one of the following structures: a dephosphorylated nick, a 1- or 8-nucleotide (nt) gap, or a 36nt 5'-overhang. T4 ligase helps to connect the 5'PO4 of the adapter chain to the only ligatable 3'OH of the substrate chain, thereby forming a forked ligation product.
[0308] To optimize ligation efficiency, we extensively tested many factors that influence general ligation efficiency, including adapter::DNA substrate ratio, T4 ligase amount, final ATP concentration, Mg 2+ The concentration, pH, incubation time and presence of different additives such as polyethylene glycol-8000 (PEG-8000) and single-stranded binding protein (SSB) were analyzed (Supplementary Figures 1 and 2). We found that adding PEG-8000 to a final concentration of 10% significantly increased the ligation efficiency from less than 10% to greater than 90% (Figure 21). 3' branch ligation also worked over a wide range of ATP concentrations (from 1 uM to 1 mM) and Mg 2+ The ligase concentration (from 3mM to 10mM) of the 3'BL was effective. The amount of ligase required for 3'BL was comparable to that for blunt-end ligation. Under our optimized conditions, we used an adapter::substrate DNA molar ratio of 10 to 100 and reacted with 1mM ATP, 10mM MgCl2 and 10% PEG-8000 at pH 7.8 at 37°C for one hour. The same adapter was ligated to a blunt-ended substrate and no ligase reaction was used as a control. To determine the yield of the ligation product, the reaction was run on a denaturing polyacrylamide gel (Figure 21b) or a TBE gel (Figures 22a-d). The ratio of product to substrate intensity was used to quantify the ligation efficiency by ImageJ (Figure 21b). The 5'-overhang ligation (lane 11 in Figure 21b) appeared to be more than 90% complete, even higher than the blunt-end ligation control (lane 14, 76.9%), indicating that the ligation efficiency of the DNA 5'-overhang was very high. The 1 or 8 nt gap substrates (lanes 5 and 8) showed good ligation efficiency of approximately 60%. However, the nick ligation (lane 2) had the lowest efficiency, approximately 20%. However, if we incubate the nick ligation for 12 hours, the ligation yield can be improved, indicating that the kinetics of the nick ligation reaction are slower (Figure 22).
[0309] We also extended our studies to different adapter and substrate sequences (Figure 22). The 5'PO4 termini of three different adapters (Ad-T, Ad-A, or Ad-GA) contained either a T, an A, or the dinucleotide GA before the consensus CTGCTGA sequence. These were ligated to the 3'OH terminus of the acceptor template, with a T at the ligation junction. Overall, Ad-T and Ad-A achieved higher ligation efficiencies (70-90%) than Ad-GA in all cases except for nick ligation (Figure 22), suggesting certain nucleotide preferences of T4 DNA ligase at the ligation junction. Regardless of the adapter and substrate sequences, ligations with 5' overhangs or 3' branches consistently showed higher efficiencies (60-90%), whereas nick ligations had significantly lower efficiencies after 1 hour of incubation. We hypothesize that these differences in ligation efficiencies are due to DNA bending, i.e., bending at the beginning of the nick / gap / overhang, exposing the 3'OH group for ligation. Longer ssDNA regions may make the 3' terminus more accessible for ligation and lead to higher ligation efficiencies. We also tested whether similar end-joining events might occur with 5' branched ligations. In contrast, no significant ligation of blunt-ended adapters to 5'PO4 ends was observed at gaps or 3' overhangs, suggesting that T4 DNA ligase may have more stringent tertiary structure requirements at the 5' end of the donor than at the 3' end.
[0310] 4.3.3′ Branch Ligation: A Novel Method for Connecting DNA to RNA
[0311] We further investigated the 3'BL on DNA / RNA hybrids (ON21 / 22), which form one DNA and one RNA 5'-overhang (Figure 23a). Negative ligation controls included DNA / RNA hybrids, ssDNA or ssRNA oligonucleotides, incubated alone or with adapters (lanes 3, 4 and 5 in Figure 23b). Interestingly, when the DNA / RNA hybrids were incubated with adapters, we saw that the size of the RNA oligonucleotide changed from the original 29nt to 49nt with an efficiency of >90%, indicating that T4 DNA ligase can effectively connect the adapter to the RNA. However, the DNA substrate remained unchanged (lanes 1 and 2 in Figure 23b). This indicates that the blunt-ended DNA adapter ligated to the 3' end of the RNA at the 5'-DNA overhang, but the 5'-RNA overhang ligated to the 3' end of the DNA. To confirm the required 5'-overhang structure for 3'BL, we performed the same ligation reaction, replacing the original DNA oligonucleotide (ON21) with another long DNA template (ON23) that is not complementary to the ON22 RNA. Not surprisingly, no ligation was observed using the ON23 DNA template, indicating that 3'BL can only occur with a 5'-overhang. Our findings suggest that T4 DNA ligase has certain substrate preferences, likely due to differences in protein-substrate binding affinity.
[0312] Previous studies have shown that T4 DNA ligase and T4 RNA ligase 2, but not T4 RNA ligase 1, can ligate 5'PO4 DNA ends to juxtaposed 3'OH DNA or RNA ends on RNA / DNA duplexes, but not to RNA 3'OH (Bullard 2006, Biochem J 398: 135-144). We performed the same ligation test using T4 RNA ligases 1 and 2 ( Figure 24C It appears that R4 RNA ligases 1 and 2 can ligate blunt-ended adaptors to RNA, but the ligation efficiency is very low (<10%).
[0313] 4.4 Construction of directional transposon insertion library
[0314] Since 3' branch ligation has been shown to be useful for efficiently ligating adapters to several genomic constructs, we explored its application in NGS workflows. Transposon-based library construction methods are time-saving and consume less input DNA than traditional NGS library preparation. However, using commercial transposon-based library preparation systems, only half of the labeled molecules are flanked by two different adapter sequences, while the labeled DNA is flanked by self-complementary regions that may form stable hairpin structures, which may compromise sequencing quality (Gorbacheva, 2015, Biotechniques Apr;58(4):200-202). In addition, PCR-mediated incorporation of adapter sequences is not suitable for whole-genome bisulfite sequencing or PCR-free NGS library construction.
[0315] To overcome these limitations, we developed a new experimental protocol for transposon-based NGS library construction combined with 3'BL. Both Tn5 and MuA transposons work through a "cut and paste" mechanism, in which the transposon adapter sequence end is ligated to the 5' end of the target DNA, creating a 9 bp or 5 bp gap at the 3' end of the genomic DNA, respectively (Figure 24). 3'BL is then used to add another adapter sequence to the 3' end of the genomic DNA at the gap to complete the ligation of the directional adapter. We compared the efficiency of the 3'BL method with that of a dual transposon insertion method, which uses two different Tn5-based adapters, TnA and TnB. Human genomic DNA was incubated with either a single TnA transposome complex or equimolar amounts of TnA and TnB transposome complexes. The TnA transposome fragmentation products were further used in 3'BL with the blunt-end adapter AdB, which shares a common adapter sequence with TnB. PCR amplification using two primers, Pr-A and Pr-B, designed for TnA and AdB / TnB adapters, respectively, showed similar PCR yields ( FIG. 24 b , lanes 9 and 10, and FIG. 24 c ), indicating that the two methods have the same efficiency. No significant amplification was observed when only one primer specific for either the TnA or AdB / TnB adapter was used ( FIG. 24 b and FIG. 24 c ). As expected, due to PCR inhibition, both the 3′-ligation method and the dual transposon insertion method showed significantly higher PCR efficiency compared to transposon insertion reactions using only the TnA or TnB transposome complex alone ( FIG. 24 b , lanes 3 and 8, and FIG. 24 c ).
[0316] 4.5 Materials and Methods
[0317] 3' branch ligation of double-stranded DNA
[0318] Substrates for 3'BL were prepared by mixing 2 pmol of ON1 or ON9 with 4 pmol of one or two additional oligonucleotides in Tris-EDTA (TE) buffer, pH 8 (Life Technologies). Substrates 1 and 5 (nicks): ON1 / 2 / 3 and ON9 / 10 / 11; substrates 2 and 6 (1 bp gap): ON1 / 2 / 4 and ON9 / 10 / 12; substrate 3 (8 bp gap): ON1 / 4 / 5; substrates 4 and 9 (5' overhang): ON1 / 2 and ON9 / 10; substrate 7 (2 bp gap): ON9 / 10 / 13; substrate 8 (3 bp gap): ON9 / 10 / 14; blunt-end control: ON1 / 6 (Figure 1, Supplementary Table 1). The template was ligated to 180 pmol of adaptor (Ad-C: ON7 / 8, Ad-T: ON15 / 16, Ad-A: ON17 / 8, or Ad-GA: ON19 / 20) using 2400 units of T4 ligase (Enzymatics Inc) in 3'BL buffer [0.05 mg / ml BSA (New England Biolabs), 50 mM Tris-Cl pH 7.8 (Amresco), 10 mM MgCl2 (EMD Millipore), 0.5 mM DTT (VWR Scientific), 10% PEG-8000 (Sigma Aldrich), and 1 mM ATP (Sigma Aldrich)]. The ATP concentration was varied from 1 μM to 1 mM, and the MgCl2 concentration was adjusted to 1 μM. 2+ Optimization tests were performed by varying the concentration, pH, temperature from 12°C to 42°C, and additives such as PEG-8000 from 2.5% to 10% and SSB from 2.5 to 20 ng / ul. The ligation mixture was prepared on ice and incubated at 37°C for 1 to 12 hours and then heat-inactivated at 65°C for 15 minutes. The samples were purified using Axygen beads (Corning) and eluted into 40 μL TE buffer. All ligation reactions were performed on 6% TBE or denaturing polyacrylamide gels (Life Technologies) and visualized on an Alpha Imager (Alpha Innotech). The amount of input control loaded was equal to or half the amount of template used for ligation. Ligation efficiency was estimated using ImageJ Software (NIH) by dividing the intensity of the ligated product by the total intensity of the ligated and unligated products.
[0319] 3' branch ligation of DNA / RNA hybrids
[0320] The substrate for 3'BL was a mixture of 10 pmol ON22 RNA oligonucleotide and 2 pmol ON21 or ON23 DNA oligonucleotide. For T4 DNA ligase-mediated 3'BL, the substrate was incubated with Ad-T (ON15 / 16) in 3'BL buffer as described above and incubated at 37°C for 1 hour. 3'BL was performed using T4 RNA ligase 1 or 2 in their own 1x RNA ligase buffer (NEB) with 20% DMSO. All ligation products were analyzed on 6% denaturing polyacrylamide gels.
[0321] Construction of a directional transposon insertion library
[0322] Transposon oligonucleotides used in this experiment were synthesized by Sangon Biotech. For two-transposon experiments using TnA and TnB, TnA, TnB, and MErev oligonucleotides were annealed at a ratio of 1:1:2. For single-transposon experiments using tn1, tn1 and MErev were annealed at a ratio of 1:1.
[0323] Transposome assembly was performed by mixing 15 pmol of pre-annealed adapters, 7 ul Tn5 transposase (Vazemy) and 5.5 ul glycerol to obtain a 20 ul reaction, which was incubated at 30°C for 1 hour. Transposon insertion of genomic DNA was performed in a 20 ul reaction comprising 100 ng gDNA, TAG buffer (Vazyme) and 2 ul of assembled transposomes (Coriell 19240). The reaction was incubated at 55°C for 10 minutes, followed by addition of 100 ul of PB buffer (Qiagen) to remove the transposome complex from the tagged DNA, and purified using Agencourt AMPure XP beads (Beckman Coulter). Incubated at 25°C for 1 hour in a reaction containing 100 pmol of adapters, 600 U of T4 DNA ligase (Enzymatics Inc.) and 3' BL buffer, AdB (ONB1, ONB2) was 3' branched to the tag DNA. The reaction was purified using AMPure XP beads. PCR amplification of labeled and gap-ligated DNA was performed in 50 μl reactions containing 2 μl of labeled or gap-ligated DNA, TAB buffer, 1 μl of TruePrep Amplification Enzyme (Vazyme), 200 mM dNTPs (Enzymatics Inc.), and 400 mM each of primers Pr-A and Pr-B. The labeling reaction was performed as follows: 3 minutes at 72°C; 30 seconds at 98°C; 8 cycles of 10 seconds at 98°C, 30 seconds at 58°C, and 2 minutes at 72°C; and a 10-minute extension at 72°C. The gap-ligation reaction was performed using the same procedure, but without the initial 3-minute extension at 72°C. The PCR reactions were purified using AMPure XP beads with either single-step size selection or by dual fractionation. The purified products were quantified using the Qubit High Sensitivity DNA Kit (Invitrogen).
[0324] Example 4: 3' Branch Ligation: A Novel Method for Ligating Non-Complementary DNA to Recessed or Internal 3'OH Ends of DNA or RNA
[0325] Nucleic acid ligases are key enzymes that repair breaks in DNA or RNA during synthesis, repair, and recombination. A variety of molecular tools have been developed leveraging the diverse activities of DNA / RNA ligases. However, additional ligase activities remain to be discovered. Here, we demonstrate the unconventional ability of T4 DNA ligase to ligate 5'-phosphorylated blunt-ended double-stranded DNA to DNA breaks at 3' recessed ends, gaps, or nicks, forming 3'-branched structures. This base-pairing-independent ligation is therefore termed 3'-branched ligation (3'BL). In an extensive study of optimal ligation conditions similar to blunt-end ligation, the presence of 10% PEG-8000 in the ligation buffer significantly improved ligation efficiency. Using different synthetic DNAs, some nucleotide preferences were observed at the ligation site, suggesting a ligation bias in the 3'BL. Furthermore, we found that T4 DNA ligase efficiently ligated DNA to the 3' end of RNA in DNA / RNA hybrids, whereas RNA ligases are less efficient in this reaction. These novel properties of T4 DNA ligase can be exploited in a wide range of molecular techniques for many important applications. We performed a proof-of-concept study of a new directional tagging protocol for next-generation sequencing (NGS) library construction that eliminates reverse adapters and allows sample barcodes to be inserted adjacent to genomic DNA. While single-transposon tagging followed by 3'BL theoretically yields 100% usable template, our empirical data demonstrate that this new approach produces higher yields compared to traditional dual-transposon or Y-transposon tagging. We further explored the potential use of the 3'BL for preparing targeted RNA NGS libraries with reduced structure-based bias and adapter dimerization.
[0326] 5.1 Introduction
[0327] Ligases repair breaks in nucleic acids, and this activity is essential for cell viability and vitality. DNA ligases catalyze the formation of phosphodiester bonds between DNA ends and play a key role in DNA repair, recombination, and replication in vivo. 1-3 RNA ligases join 5'-phosphodiester (5'PO4) and 3'-hydroxyl (3'OH) RNA ends via phosphodiester bonds and are involved in RNA repair, splicing, and editing. 4 Ligases from all three kingdoms (bacteria, archaea, and eukaryotes) can be used as important molecular tools in vitro for applications such as cloning, ligase-based amplification or detection, and synthetic biology. 5-7
[0328] One of the most widely used ligases in vitro is bacteriophage T4 DNA ligase, a single 55-kDa polypeptide that requires ATP as an energy source. 8 T4 DNA ligase typically ligates adjacent 5'PO4 and 3'OH termini of duplex DNA. In addition to sealing nicks and ligating sticky ends, T4 DNA ligase can also efficiently catalyze blunt-end ligation, a property not observed with any other DNA ligase. 9,10 Several unusual catalytic properties of this ligase have been previously reported, such as sealing single-stranded gaps in duplex DNA, sealing nicks near abasic sites in duplex DNA (dsDNA), promoting intramolecular loop formation in partially duplex DNA, and ligating DNA strands containing 3' branched extensions. 11-13 Template-independent ligations mediated by T4 ligase have also been observed, such as sealing mispaired nicks in dsDNA or even single-stranded DNA (ssDNA), albeit with very low efficiencies. 15 These results suggest that perfect complementary base pairing at or near the ligation junction is not required for some of the unconventional T4 DNA ligase activities. T4 RNA ligases 1 and 2 are products of T4 bacteriophage genes 63 and 24, respectively. Both require adjacent 5'PO4 and 3'OH termini for successful ligation, with ATP hydrolyzed to AMP and PPi. Substrates for T4 RNA ligase 1 include single-stranded RNA and DNA, while T4 RNA ligase 2 preferentially seals nicks in dsRNA rather than ligating the ends of ssRNA. 16,17
[0329] Here, we demonstrate an unconventional end-joining event mediated by T4 DNA ligase, which we term 3'-branched ligation (3'BL). This method can ligate DNA or DNA / RNA fragments at nicks, single-stranded gaps, or 3' recessed ends to form branched structures. This report includes an extensive investigation of various ligation cofactors and activators, as well as optimization of ligation conditions for this novel ligation. Using our 3'BL protocol, base pairing is not required, and in most cases, including 1nt gaps, ligation can achieve 70-90% completion rates. One application of this method is the ligation of adapters to DNA or RNA during next-generation sequencing (NGS) library preparation. Several genomic structures previously considered unligatable can now be used as substrates for 3'BL, resulting in high conversion rates from input DNA to adapter-ligated molecules while avoiding chimeras. We demonstrate that 3'BL can be combined with transposon tagmentation to increase library yields. Our proposed targeted tagging strategy theoretically yields 100% sequence-usable templates. Our studies demonstrate the value of this novel technology for NGS library preparation and its potential to advance numerous other molecular applications.
[0330] 5.2 Results: 3′ branch ligation, a novel method for joining DNA ends
[0331] Typically, DNA ligation involves joining the 5'PO4 and 3'OH DNA ends of sticky or blunt-ended fragments. Sticky-end ligation is generally faster and less dependent on enzyme concentration than blunt-end ligation. Both processes can be catalyzed by bacteriophage T4 DNA ligase, which uses ATP as a cofactor for energy production and requires Mg. 2+ 8. T4 DNA ligase has also been reported to ligate specific or degenerate single-stranded oligonucleotides to partially single-stranded substrates by hybridization. 18,19 Here, we demonstrate an unconventional T4 DNA ligase-mediated ligation that does not require complementary base pairing and can ligate a blunt-ended DNA donor to the 3'OH terminus of a double-stranded DNA acceptor at a 3' recessed strand, a gap, or a nick (Figure 26a). Therefore, we use the term 3'-branched ligation (3'BL) to describe these ligations. The synthetic donor DNA we used contained a 5' blunt double-stranded end and a 3' ssDNA end. The acceptor substrate contained one of the following structures: a dephosphorylated nick, a 1- or 8-nucleotide (nt) gap, or a 3' 36-nt recessed end (Supplementary Table 1). T4 ligase facilitates the ligation of the 5'PO4 of the donor strand to the only ligatable 3'OH of the acceptor strand, forming a branched ligation product.
[0332] To optimize ligation efficiency, we extensively tested many factors that affect general ligation efficiency, including adapter:DNA substrate ratio, T4 ligase amount, final ATP concentration, Mg 2+ The concentration, pH, incubation time and different additives such as polyethylene glycol-8000 (PEG-8000) and single-stranded binding protein (SSB) were used. Adding PEG-8000 to a final concentration of 10% significantly increased the ligation efficiency from less than 10% to more than 80% (Figures 26 and 27). A wide range of ATP concentrations (from 1 uM to 1 mM) and Mg 2+ The concentration (3mM to 10mM) is compatible with 3'BL. The amount of ligase required for 3'BL is comparable to that for blunt-end ligation. Under our optimized conditions, we used a 30:100 donor:substrate DNA molar ratio and performed the reaction at 37°C, pH 7.8, with 1mM ATP, 10mM MgCl2, and 10% PEG-8000 for one hour. Ligation of the same donor with a blunt-end substrate and a reaction without ligase served as positive and negative controls, respectively.
[0333] The ligation donor (Ad-G) is double-stranded at one end (5' phosphorylated and 3' dideoxy protected) and single-stranded at the other end (3' dideoxy protected) (Figure 26). The ligation substrate consists of the same bottom chain (ON1) and different top chains to form the nick, vacancy and protrusion structures. To quantify the yield of the ligation product, the reaction products were separated on a 6% denaturing polyacrylamide gel (Figure 26b). The ligation efficiency was calculated as the ratio of product to substrate intensity using ImageJ (Figure 26b). The 3'-recessed ligation (lane 11 in Figure 26b) appeared to be about 90% complete, even higher than the blunt-end ligation control (lane 14, 72.74%), and indicated that the ligation efficiency with the 3'-recessed DNA end was very high. The 1-nt or 8-nt vacancy substrates (lanes 5 and 8) showed good ligation efficiency of about 45%. The nicked ligation (lane 2) had the lowest efficiency, about 13%. However, the ligation yield improved when the nick ligation reaction was incubated for a longer time, indicating that the kinetics of the nick ligation reaction were slower.
[0334] We also extended our studies to different adapter and substrate sequences (Figure 27). The 5'PO4 termini of three different adapters (Ad-T, Ad-A, or Ad-GA in Supplementary Table 1) contained a single T, A, or the dinucleotide GA at the ligation junction preceding the consensus CTGCTGA sequence. These 5'PO4 termini were ligated to the 3'OH terminus of the acceptor template, with a T added at the ligation junction. Overall, high ligation efficiencies (70–90%) were observed in most cases, with the exception of nick ligation or 3'BL with Ad-GA (Figure 27), suggesting that T4 DNA ligase has some nucleotide preference at the ligation junction. Independent of adapter and substrate sequence, ligation with 3' recessed ends or gaps consistently showed higher efficiencies (60–90%) over a 1-hour incubation, while nick ligation efficiency was very low. We hypothesize that these differences in ligation efficiency are due to DNA bending at the beginning of the nick / gap / overhang, exposing the 3'OH group for ligation. Longer ssDNA regions may make the 3' end more accessible for ligation and lead to higher ligation efficiency. We also tested whether similar end-joining events might occur with 5' branched ligations. In contrast to 3'BL, no significant ligation of blunt-ended adapters to 5'PO4 ends was observed at gaps or 5' recessed ends. This result suggests that steric hindrance to T4 DNA ligase is greater at the 5' end of the donor than at the 3' end.
[0335] 5.3: 3' branch ligation to connect DNA to RNA
[0336] We further investigated the 3'BL on DNA / RNA hybrids (ON-21 / ON-23 in Table 3), which form one DNA and one RNA 5'-overhang (Figure 28a). Ligation on the DNA / DNA hybrid was used as a positive control, while negative ligation controls included DNA / RNA hybrids, ssDNA, or ssRNA oligonucleotides incubated alone or with adapters (lanes 3, 4, and 5 in Figure 28c). Interestingly, when the DNA / RNA hybrid was incubated with a blunt-ended dsDNA donor, we observed that the size of the RNA oligonucleotide changed from the original 29nt to 49nt after ligation. However, the DNA substrate remained unchanged (lanes 1 and 2 in Figure 28c). This result indicates that the blunt-ended dsDNA donor ligates to the 3'-end of the RNA at the 3'-recessed DNA end, but not to the 3'-end of the DNA at the 3'-recessed RNA end. As a positive control, DNA / DNA hybrids with 3' recessed ends on each side showed a band shift efficiency of nearly 100% toward the larger species on both chains. To confirm that 3'BL requires a 3' recessed structure, we performed the same ligation reaction while replacing the original DNA oligonucleotide (ON-21) with another long DNA template (ON-23) that is not complementary to ON-22 RNA (Figure 28b). Not surprisingly, no ligation was observed using the ON-23 DNA template (lanes 10-13 in Figure 28c). Our findings indicate that T4 DNA ligase can promote 3'BL on DNA / RNA hybrids and that this activity has a certain spatial substrate preference, which may be affected by the difference in T4 DNA ligase binding affinity to the substrate.
[0337] Previous studies have reported that when the complementary strand is RNA rather than DNA, T4 DNA ligase and T4 RNA ligase 2, but not T4 RNA ligase 1, can efficiently ligate 5'PO4 DNA ends to juxtaposed 3'OH DNA or RNA ends to seal the nick in the DNA / RNA hybrid. 17 Therefore, we performed the same ligation assay using T4 RNA ligases 1 and 2 in 20% DMSO (Figure 28d) or 10% PEG. In both assays, T4 RNA ligase 1 and T4 RNA ligase 2 slightly ligated blunt-ended adapters to the 3' end of RNA in the DNA / RNA hybrid. Notably, in the RNA-only control, T4 RNA ligase 2 could ligate blunt-ended dsDNA adapters to ssRNA. In summary, T4 DNA ligase, but not T4 RNA ligase, is able to efficiently ligate blunt-ended dsDNA to the 3' end of RNA via 3'BL.
[0338] 5.4 Construction of Directed Marker Library
[0339] Because 3'BL can be used to efficiently ligate adapters to multiple genomic constructs, we explored its application in NGS workflows. Compared to traditional NGS library preparation, transposon-based library construction is rapid and consumes less input DNA. However, using commercial transposon-based library preparation systems, only half of the labeled molecules are flanked by two different adapter sequences (Figure 29a), and labeled DNA is flanked by self-complementary regions that can form stable hairpin structures and compromise sequencing quality. 20 Furthermore, PCR-mediated incorporation of adapter sequences is not yet suitable for whole-genome bisulfite sequencing or PCR-free NGS library construction.
[0340] To overcome these limitations, we developed a new experimental protocol for transposon-based NGS libraries by incorporating 3'BL. Both Tn5 and MuA transposons work by a "cut and paste" mechanism in which the transposon adapter sequence end is ligated to the 5' end of the target DNA, creating a 9-bp or 5-bp gap at the 3'-end of the genomic DNA, respectively (Figure 29a). Subsequently, 3'BL can be used to add another adapter sequence to the 3' end of the genomic DNA at the gap to complete the ligation of the directional adapter (Figure 29c). In this manuscript, we used the Tn5 transposon to compare the efficiency of the single-tag + 3'BL method (Figure 29c) with the efficiency of the dual-tag method (Figure 29a), which uses two different Tn5-based adapters TnA and TnB, as well as with the efficiency of another directional single-tag strategy using a Y adapter containing two different adapter sequences (Figure 29b). Human genomic DNA was incubated with equimolar amounts of TnA and TnB transposon complexes, or with TnA transposome complexes alone, or with TnY (TnA / B) transposome complexes.
[0341] Only the product of TnA transposome fragmentation was further used as a 3'BL template for the blunt-end adapter AdB, which shares a common adapter sequence with TnB. PCR amplification was performed using two primers, Pr-A and Pr-B, designed to recognize TnA and AdB / TnB adapters, respectively. Quantitative data showed that TnA&AdB had the highest efficiency compared to TnA&TnB and TnY (TnA / B) (Figure 29d). No significant amplification was observed when only one primer specific for the TnA adapter was used (Figure 29d). As expected, due to PCR inhibition, the TnA-3'BL method, the dual tag method, and the TnY method all showed significantly higher PCR efficiency than the tag reaction using only TnA or TnB transposome complexes alone (Figure 29d).
[0342] We also sequenced these libraries using BGISEQ-500 and compared the base position deviations between the transposon interference end, the 3'BL end, and the conventional TA junction end (Figure 30). Clearly, the position deviation of the 3'BL end is less than that of the Tn5 end (Figures 30a-b), because the 3'BL end is affected by the transposon interruption and the 3'BL. Because only the first 6nt (positions 1-6) of the 3'BL end showed base deviation, and the deviation was similar but not identical to that of the hybrid Tn5 end (positions 30-35 after the 9-nt overhang), we believe that the position deviation observed at the 3'BL end is mainly caused by the Tn5 transposon. Therefore, the 3'BL causes the least deviation and is similar to the conventional TA junction (Figure 30c).
[0343] 5.5 Discussion
[0344] A key property of T4 DNA ligase is its efficient ligation of blunt-ended dsDNA,21,22 which has not been observed with other DNA ligases. This ligase has also been reported to mediate several unconventional catalytic events, such as ligating single-stranded gaps or base mismatches in duplex DNA,11,12 forming stem-loop molecules from partially duplex DNA,13 or inefficiently ligating ssDNA in a template-independent manner,20.
[0345] Here, we demonstrate that T4 DNA ligase catalyzes the ligation of blunt-ended dsDNA to the 3'OH termini of nicked dsDNA, as well as the ligation of partially single-stranded double-stranded DNA with nicks or 5'-overhangs. In contrast, ligation to 5'PO4 ends within 5' recessed ends or gaps was not observed, suggesting that after binding to the 5'PO4 ends of dsDNA adapters, T4 DNA ligase can access the 3' recessed ends upon DNA bending. Using our 3'BL method, which does not require base pairing, optimized conditions can achieve completion rates exceeding 70% even with a 1 nt gap. However, varying ligation efficiencies were observed for ligating 5'T, A, or GA to 3'T (Figure 2), suggesting some sequence preference at the ligation junction. Despite the well-recognized ligation bias,23 T4 DNA ligase is commonly used in the adapter addition step during NGS library preparation. The ability of T4 ligase to perform 3'BL has enabled adapter ligation into several genomic constructs previously thought to be unligatable, thereby improving template utilization. 3'BL can also be combined with transposon tags. With the traditional dual-transposon strategy, only 50% of the labeled molecules are suitable for the subsequent amplification step. However, when DNA tagging is performed using a single transposon followed by 3'BL, an increased yield of molecules with different adapters at each insert end can be obtained (Figure 4). Furthermore, the tagged 3'BL product can be directly loaded into an Illumina flow cell as a PCR-free WGS library, which is difficult to achieve using the dual-transposon strategy.
[0346] Other targeted transposon protocols have been proposed that use Y transposons composed of two different adapter sequences or replace the unligated strand of a single transposon with a second adapter oligonucleotide, followed by gap filling and ligation. 24 However, these approaches continue to retain the inverted adapter sequence and are unable to insert sample barcodes near the genomic DNA as in the tagged-3′BL protocol. Based on NGS data, genomic ends ligated with 3′BLs also showed fewer positions with positional base composition bias, and the first 6-nt bias was mild, mainly caused by transposon interruption, indicating that 3′BLs have minimal positional bias. Using this new library construction approach, Wang et al. successfully achieved highly accurate and complete variant calling in WGS and nearly perfectly phased variants into long contigs with an N50 size of up to 23.4 Mb, which can be used for long fragment reads (BioRxiv, https: / / doi.org / 10.1101 / 324392).
[0347] In this study, we also investigated 3'BL using templates that form chimeric DNA / RNA duplexes with 5' DNA and 5' RNA overhangs (Figure 3). Unexpectedly, blunt-ended dsDNA was efficiently ligated to the 3' end of RNA rather than DNA, indicating a preference for ternary complex formation by T4 ligase. Ligation efficiency was significantly reduced if T4 RNA ligase I or II was used to ligate the ends. Another preliminary but important application of 3'BL with T4 DNA ligase is the enrichment of mRNA or the construction of targeted RNA libraries, particularly for miRNAs, small regulatory RNAs whose unregulated expression contributes to numerous diseases. 25, 26 Therefore, our 3'BL technology could be readily applied to miRNA-based cancer and Alzheimer's disease assays. Hybridization with DNA probes targeting poly(A) tails or specific miRNA sequences can be used to create DNA-RNA hybrids with DNA 5' overhangs, which can then be ligated to adapter sequences barcoded with sample and / or UIDs via 3'BL. These universal sequences can then be reverse transcribed to generate cDNA targeting the RNA sequence. Compared with current miRNA capture technologies, the use of 3'BL mediated by T4 DNA ligase may provide some advantages for the construction of NGS RNA libraries. First, hybridization with the DNA chain will prevent the RNA chain from forming secondary structures, thereby alleviating the bias introduced by other experimental protocols. Second, T4 DNA ligase can achieve efficient adapter addition through 3'BL, thereby avoiding the intramolecular RNA interactions that RNA ligase can promote. Third, adapter dimers can be effectively eliminated, which may make unnecessary gel purification unnecessary. This new method can improve unbiased microRNA expression profiling through a simple and scalable workflow, so large-scale studies will become more economical and affordable.
[0348] Our findings add to the growing understanding of T4 DNA ligase activity. We foresee that 3' branch ligation will become a versatile tool in molecular biology, driving the development of new DNA engineering methods beyond previously described NGS applications.
[0349] 5.6 Materials and Methods
[0350] 3' branch ligation of double-stranded DNA
[0351] The substrates for 3'BL were composed of 2 pmol of ON1 or ON9 mixed with 4 pmol of one or two additional oligonucleotides in pH 8 Tris-EDTA (TE) buffer (Life Technologies) as follows: substrates 1 and 5 (nicks), ON-1 / 2 / 3 and ON-9 / 10 / 11; substrates 2 and 6 (1-nt gap), ON1 / 2 / 4 and ON9 / 10 / 12; substrate 3 (8-nt gap), ON1 / 4 / 5; substrates 4 and 9 (5' overhang), ON1 / 2 and ON9 / 10; substrate 7 (2-nt gap), ON9 / 10 / 13; substrate 8 (3-nt gap), ON9 / 10 / 14; blunt-end control, ON1 and ON6 (Figure 26, Supplementary Table 1). The template was ligated to 180 pmol of adapter (Ad-G: ON7 / 8, Ad-T: ON15 / 16, Ad-A: ON17 / 8, or Ad-GA: ON19 / 20) using 2400 units of T4 ligase (Enzymatics Inc.) in 3'BL buffer [0.05 mg / ml BSA (New England Biolabs), 50 mM Tris-Cl pH 7.8 (Amresco), 10 mM MgCl2 (EMD Millipore), 0.5 mM DTT (VWR Scientific), 10% PEG-8000 (Sigma Aldrich), and 1 mM ATP (Sigma Aldrich)]. The ATP concentration was varied from 1 μM to 1 mM. 2+ Optimization tests were performed by varying the concentration from 3 mM to 10 mM, the pH from 3 to 9, the temperature from 12°C to 42°C, and adjusting the additives (e.g., PEG-8000 from 2.5% to 10%, SSB from 2.5 to 20 ng / μL). The ligation mixture was prepared on ice and incubated at 37°C for 1 to 12 hours, followed by heat inactivation at 65°C for 15 minutes. The samples were purified using Axygen beads (Corning) and eluted into 40 μL of TE buffer. All ligation reactions were performed on 6% TBE or denaturing polyacrylamide gels (Life Technologies) and visualized on an Alpha Imager (Alpha Innotech). The amount of input control loaded was equal to or half the amount of template used for ligation. Ligation efficiency was estimated using ImageJ Software (NIH) by dividing the intensity of the ligated product by the total intensity of the ligated and unligated products.
[0352] 3' branch ligation of DNA / RNA hybrids
[0353] The substrate for 3'BL was a mixture of 10 pmol ON-21 RNA oligonucleotide and 2 pmol ON-21 or ON-23 DNA oligonucleotide. For T4 DNA ligase-mediated 3'BL, the substrate was incubated with Ad-T (ON15 / 16) in 3'BL buffer as described above and incubated at 37°C for 1 hour. 3'BL was performed using T4 RNA ligase 1 or 2 in 1x RNA ligase buffer (NEB) with 20% DMSO or 25% PEG. All ligation products were analyzed on 6% denaturing polyacrylamide gels.
[0354] Directed marker library construction
[0355] Transposon oligonucleotides used in this experiment were synthesized by Sangon Biotech. For dual transposon experiments using TnA / TnB, oligonucleotides for TnA (ON24), TnB (ON25), and MErev (ON26) were annealed at a ratio of 1:1:2. For single transposon experiments using TnA, ON24 and ON26 were annealed at a ratio of 1:1. For Y (TnA & TnB) transposon experiments, ON24 and ON27 were annealed at a ratio of 1:1.
[0356] Transposon assembly was performed by mixing 100 pmol of pre-annealed adapters, 7 μL of Tn5 transposase, and enough glycerol to obtain a total of 20 μL of reaction and incubating it at 30°C for 1 hour. Labeling of genomic DNA (Coriell12878) was performed in a 20 μL reaction containing 100 ng of gDNA, TAG buffer (homemade), and 1 μL of assembled transposon. The reaction was incubated at 55°C for 10 minutes; 40 μL of 6 M guanidine hydrochloride (Sigma) was then added to remove the transposon complex from the labeled DNA, and the DNA was purified using Agencourt AMPure XP beads (Beckman Coulter). In a reaction containing 100 pmol of adapters, 600 U of T4 DNA ligase (Enzymatics Inc.), and 3'BL buffer, AdB (ON28 and ON29) was gap-ligated to the labeled DNA at 25°C for 1 hour. The reaction was purified using AMPure XP beads. PCR amplification of labeled and gap-ligated DNA was performed in 50 μL reactions containing 2 μL of labeled or gap-ligated DNA, TAB buffer, 1 μL of TruePrep Amplification Enzyme (Vazyme), 200 mM dNTPs (Enzymatics Inc.), and 400 mM each of primers Pr-A and Pr-B. The labeling reaction was incubated as follows: 72°C for 3 minutes; 98°C for 30 seconds; 8 cycles of: 98°C for 10 seconds, 58°C for 30 seconds, and 72°C for 2 minutes; and 72°C for 10 minutes extension. The gap-ligated reaction was performed using the same procedure, but without the initial extension at 72°C for 3 minutes. PCR reactions using prA (ON30) or prA and prB (ON31) were purified using AMPure XP beads. The purified products were quantified using the Qubit High Sensitivity DNA Kit (Invitrogen).
[0357] 5.6 References for Example 4
[0358] 1.Lehnman,IRDNA ligase:structure,mechanism,and function.Science(80-.).186,790-797(1974).
[0359] 2.Tomkinson,AE&Mackey,ZBStructure and function of mammalian DNAligases.Mutat.Res.Repair 407,1-9(1998).
[0360] 3.Timson,D.J.,Singleton,M.R.&Wigley,D.B.DNA ligases in the repair andreplication of DNA.Mutat.Res.Repair 460,301-318(2000).
[0361] 4.Ho,C.K.,Wang,L.K.,Lima,C.D.&Shuman,S.Structure and mechanism of RNAligase.Structure 12,327-339(2004).
[0362] 5.Tomkinson,A.E.,Vijayakumar,S.,Pascal,J.M.&Ellenberger,T.DNAligases:structure,reaction mechanism,and function.Chem.Rev.106,687-699(2006).
[0363] 6.Pascal,J.M.DNA and RNA ligases:structural variations and shared mechanisms.Curr.Opin.Struct.Biol.18,96-105(2008).
[0364] 7.Shuman,S.DNA ligases:progress and prospects.J.Biol.Chem.284,17365-17369(2009).
[0365] 8.Dickson,K.S.,Burns,C.M.&Richardson,J.P.Determination of the free-energy change for repair of a DNA phosphodiester bond.J.Biol.Chem.275,15828-15831(2000).
[0366] 9.Cai,L.,Hu,C.,Shen,S.,Wang,W.&Huang,W.Characterization ofbacteriophage T3 DNA ligase.J.Biochem.135,397-403(2004).
[0367] 10. Thermostable DNA Ligase.Available at:http: / / www.epibio.com / enzymes / ligases-kinases-phosphatases / dna-ligases / ampligase-thermostable-dna-ligase?details.
[0368] 11.Nilsson,S.V&Magnusson,G.Sealing of gaps in duplex DNA by T4DNAligase.Nucleic Acids Res.10,1425-1437(1982).
[0369] 12.Goffin,C.,Bailly,V.&Verly,W.G.Nicks 3′or 5′to AP sites or tomispaired bases,and one-nucleotide gaps can be sealed by T4 DNAligase.NucleicAcids Res.15,8755-8771(1987).
[0370] 13.Mendel-Hartvig,M.,Kumar,A.&Landegren,U.Ligase-mediatedconstructionof branched DNA strands:a novel DNA joining activity catalyzed byT4 DNAligase.Nucleic Acids Res.32,e2-e2(2004).
[0371] 14.Alexander,R.C.,Johnson,A.K.,Thorpe,J.A.,Gevedon,T.&Testa,S.M.Canonical nucleosides can be utilized by T4 DNA ligase asuniversaltemplate bases at ligation junctions.Nucleic Acids Res.31,3208-3216(2003).
[0372] 15.Kuhn,H.&Frank-Kamenetskii,M.D.Template-independentligation ofsingle-stranded DNA by T4 DNA ligase.FEBS J.272,5991-6000(2005).
[0373] 16.Ho,C.K.&Shuman,S.Bacteriophage T4 RNA ligase 2(gp24.1)exemplifiesa family of RNA ligases found in all phylogeneticdomains.Proc.Natl.Acad.Sci.99,12709-12714(2002).
[0374] 17.Bullard,D.R.&Bowater,R.P.Direct comparison of nick-joiningactivityof the nucleic acid ligases from bacteriophage T4.Biochem.J.398,135-144(2006).
[0375] 18.Broude,N.E.,Sano,T.,Smith,C.L.&Cantor,C.R.Enhanced DNAsequencingby hybridization.Proc.Natl.Acad.Sci.91,3072-3076(1994).
[0376] 19.Gunderson,K.L.et al.Mutation detection by ligation to complete n-mer DNA arrays.Genome Res.8,1142-1153(1998).
[0377] 20.Gorbacheva,T.,Quispe-Tintaya,W.,Popov,V.N.,Vijg,J.&Maslov,A.Y.Improved transposon-based library preparation for the Ion Torrentplatform.Biotechniques 58,200(2015).
[0378] 21.Sgaramella,V.&Khorana,H.G.CXII.Total synthesis of thestructuralgene for an alanine transfer RNA from yeast.Enzymic joining of thechemicallysynthesized polydeoxynucleotides to form the DNA duplexrepresentingnucleotide sequence 1 to 20.J.Mol.Biol.72,427-444(1972).
[0379] 22.SGARAMELLA,V.&EHRLICH,S.D.Use of the T4 PolynucleotideLigase inThe Joining of Flush-Ended DNA Segments Generated by RestrictionEndonucleases.FEBS J.86,531-537(1978).
[0380] 23.Seguin-Orlando,A.et al.Ligation bias in illumina next-generationDNA libraries:implications for sequencing ancient genomes.PLoS One 8,e78575(2013).
[0381] 24.Goryshin,I.,Baas,B.,Vaidyanathan,R.&Maffitt,M.Oligonucleotidereplacement for di-tagged and directional libraries.(2016).
[0382] 25.Bushati,N.&Cohen,S.M.microRNA functions.Annu.Rev.Cell Dev.Biol.23,175-205(2007).
[0383] 26. Mallory, AC & Vaucheret, H. Functions of microRNAs and related smallRNAs in plants. Nat. Genet. 38, S31 (2006).
[0384] While the present invention has been disclosed with reference to certain aspects and embodiments, it will be apparent that other embodiments and variations of the invention may be devised by others skilled in the art without departing from the true spirit and scope of the invention.
[0385] Each publication and patent document cited in this disclosure is incorporated herein by reference for all purposes in the United States of America to the same extent as if each such publication or document was specifically and individually indicated to be incorporated herein by reference. Citation of publications and patent documents is not intended to indicate that any such document is relevant prior art and does not constitute an admission as to its contents or date. Table 1B Table 2. Scaffolding statistics 1 HiC read pairs from human embryonic stem cells (hESCs) (30) were downloaded and used to scaffold SMRT reads using SALSA (28) and the same process as for the stLFR library. 2 The results reported by Ghurye et al. used the same HiC read pairs to scaffold SMRT reads using SALSA. Table 3 Table 4 Table 5
Claims
1. A method for preparing a sequencing library for sequencing a target nucleic acid without using nanodroplets, comprising: (a) synthesizing (i) a first fragment of a target nucleic acid comprising a nick and / or a gap and (ii) a population of beads into a single mixture, Each bead comprises a plurality of 3' branched ligation adapters, each of which comprises an adapter oligonucleotide whose 3' end is fixed to the bead, wherein each adaptor oligonucleotide comprises a barcode sequence, wherein the 3' branched ligation adaptors immobilized on the same bead comprise the same barcode sequence, and wherein a majority of the beads have different barcode sequences, and (b) ligating the adaptor oligonucleotide of the branched ligation adaptor of the single bead to the first fragment of the target nucleic acid at the nick and / or gap by 3' branched ligation.
2. The method of claim 1, wherein the 3' branched ligation adapter oligonucleotide is a blunt-ended adapter, and wherein the 3' branched ligation comprises covalently linking the 5' phosphate of the blunt-ended adapter to the recessed 3' hydroxyl group at the nick of the first fragment.
3. The method according to claim 1 or 2, wherein the method further comprises performing primer extension to form an extension product having a blunt end, and ligating a second adaptor to the blunt end of the extension product.
4. The method of any one of claims 1 to 3, wherein the nicks are produced by a random nicking enzyme, and wherein ligation occurs upon nick formation.
5. The method according to any one of claims 1 to 4, wherein step (b) occurs in the presence of polyethylene glycol-8000.
6. A method according to any one of claims 1 to 4, comprising wrapping the target nucleic acid around the beads.
7. The method of claim 1, wherein the mixture comprises polyethylene glycol-8000 (PEG-8000), and the concentration of PEG-8000 is 10%.
8. The method according to any one of claims 1 to 4, comprising amplifying the extension products to generate amplicons, and sequencing the amplicons to generate sequence reads, wherein sequence reads having the same barcode sequence are derived from the same first fragment.
9. The method according to claim 8, comprising the additional step of: (e) assigning the majority of sequence reads to the corresponding first fragment; and (f) assembling the sequence reads to generate an assembled sequence of the target.
10. The method according to any one of claims 1 to 4, wherein the target nucleic acid is genomic DNA, optionally from a eukaryote, optionally from a mammal, optionally from human genomic DNA, optionally the target nucleic acid is from a prokaryote or a heterogeneous population of prokaryotes, such as from the microbiome of an individual.
11. The method according to any one of claims 1 to 4, wherein the target nucleic acid is a tandem dsDNA.
12. The method according to any one of claims 1 to 4, wherein the plurality of first fragments are derived from a single cell.
13. The method according to any one of claims 1 to 4, wherein the single mixture contains 5-100 human DNA genome equivalents.
14. The method of any one of the preceding claims, wherein the average spacing of the adaptor oligonucleotides on the beads is <10 nm, <15 nm, <20 nm, <30 nm, <40 nm, or 50 nm.
15. A composition comprising: (1) Random nicking enzyme, (2) ligase; (3) a first fragment of a plurality of nucleic acid fragments comprising a nick, (4) a branched ligation adaptor comprising a barcode oligonucleotide, wherein the barcode oligonucleotide is attached to a bead and comprises a barcode, The 5' end of the double-stranded blunt end is ligated to the 3' end of at least some of the first fragments.
Citation Information
Patent Citations
Method for constructing nucleic acid single-stranded cyclic library and reagents thereof
US20180044667A1
Mate pair library construction
US20180044668A1
Nucleic acid analysis by random mixtures of non-overlapping fragments
WO2006138284A2
Fuel cell system, and vehicle
WO2010137149A1
Multiple tagging of long DNA fragments
WO2014145820A2