An in-solution positional covercoding method for sequencing long DNA molecules
The method generates nested sets of nucleic acid constructs with unique barcodes and varying lengths to address sequencing inefficiencies in long DNA molecules, achieving accurate and cost-effective sequencing of entire genomic DNA.
Patent Information
- Application Number
- JP2025504308
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-25
- Filing Date
- 2023-07-20
- Publication Date
- 2025-08-13
AI Technical Summary
Sequencing long genomic DNA is challenging due to inefficiencies in existing platforms like DNBseq and Illumina sequencers, which struggle with read lengths and repetitive sequences, leading to difficulties in accurately deciphering sequence information.
A method for generating nested sets of single-stranded or double-stranded nucleic acid constructs with unique barcodes and varying lengths, using adapter sequences and enzymatic processes to create constructs that can be sequenced without nanodrops, allowing for accurate and efficient sequencing of long DNA molecules.
Enables sequencing of entire long genomic DNA fragments by retaining location information, overcoming repetitive sequence challenges and providing complete sequence coverage at a lower cost compared to bead-based methods.
Smart Images

Figure 2025526393000001_ABST
Abstract
Description
[Technical Field]
[0001] Related Applications
[0002] This application claims priority to and the benefit of U.S. Provisional Application No. 63 / 369,346, filed July 25, 2022, the entire contents of which are incorporated herein by reference for all purposes. [Background technology]
[0003] Sequencing long genomic DNA can be challenging. Many sequencing platforms, such as DNBseq and Illumina sequencers, are not designed for sequencing long DNA molecules. For example, generating DNBs from long DNA molecules with enough template copies for high-quality sequencing can be challenging. Illumina sequencers typically require bridge amplification, and bridge amplification for long DNA molecules tends to be inefficient. Furthermore, the read lengths possible with these systems are typically less than 500 bases, making it impossible to sequence the central portions of these molecules. Thus, while these MPS sequencing platforms can be cost-effective and efficient, the sequence reads obtained from these platforms are limited in length. Sequencing DNA molecules with highly repetitive base sequences presents additional challenges. For example, it is often difficult to decipher whether two identical sequence reads relate to different locations in the genome or simply represent duplicated sequence reads at the same location in the genome. Therefore, there remains a need to prepare sequencing libraries that allow for accurate and efficient collection of sequence information of long DNA molecules from short sequence reads. Summary of the Invention
[0004] In one aspect, disclosed herein is a method for generating single-stranded, adapted constructs for sequencing, optionally without the use of nanodrops, comprising: Preparing multiple nested sets of single-stranded nucleic acid constructs, optionally in a single mixture, each single-stranded nucleic acid construct in each nested set comprises a target sequence portion flanked by a first adaptor sequence at the 5' end and a second adaptor sequence at the 3' end; the first adapter sequence comprises, from 5' to 3', a primer binding sequence, a barcode sequence, and a first hybridization sequence, and the second adapter sequence comprises a second hybridization sequence; the first and second hybridization sequences are complementary to each other; each target sequence portion having a first end and a second end; the distance between the first end and the barcode sequence is shorter than the distance between the second end and the barcode sequence; the single-stranded nucleic acid constructs in each nested set share the same barcode sequence, and the single-stranded nucleic acid constructs in different nested sets have different barcode sequences; preparing, for each nested set of single-stranded nucleic acid constructs, target sequence portions of the nested set that have identical nucleotide sequences near a first end and differ from one another by truncation near a second end, such that each nested set of single-stranded nucleic acid constructs contains multiple target sequence portions having different lengths.
[0005] In another aspect, disclosed herein is a method for generating single-stranded DNA circles containing single-stranded, adapted constructs for sequencing, optionally without the use of nanodrops, comprising: Preparing multiple nested sets of single-stranded nucleic acid constructs in a single mixture, comprising: each single-stranded nucleic acid construct of each nested set comprises a target sequence portion flanked by a first adaptor sequence and a second adaptor sequence; the first adapter sequence comprises a barcode sequence and a primer binding sequence; each target sequence portion having a first end and a second end; the distance between the first end and the barcode sequence is shorter than the distance between the second end and the barcode sequence; the single-stranded nucleic acid constructs of each nested set share the same barcode sequence, and the single-stranded nucleic acid constructs of different nested sets have different barcode sequences; For each nested set of single-stranded nucleic acid constructs: (a) the target sequence portions of a nested set of single-stranded nucleic acid constructs have identical nucleotide sequences near a first end and differ from one another by truncations near a second end, such that each nested set includes multiple target sequence portions having different lengths; (b) circularizing the single-stranded nucleic acid constructs of each nested set to generate single-stranded DNA circles to which the first adaptor sequence and the second adaptor sequence are joined; preparing a method of
[0006] In another aspect, disclosed herein is a method for generating single-stranded DNA circles containing single-stranded, adapted constructs for sequencing, optionally without the use of nanodrops, comprising: Preparing multiple nested sets of single-stranded nucleic acid constructs in a single mixture, comprising: each single-stranded nucleic acid construct of each nested set comprises a target sequence portion flanked by a first adaptor sequence and a second adaptor sequence; the first adapter sequence comprises a barcode sequence and a primer binding sequence; each target sequence portion having a first end and a second end; the distance between the first end and the barcode sequence is shorter than the distance between the second end and the barcode sequence; the single-stranded nucleic acid constructs of each nested set share the same barcode sequence, and the single-stranded nucleic acid constructs of different nested sets have different barcode sequences; For each nested set of single-stranded nucleic acid constructs: (a) the target sequence portions of a nested set of single-stranded nucleic acid constructs have identical nucleotide sequences near a first end and differ from one another by truncations near a second end, such that each nested set includes multiple target sequence portions having different lengths; (b) circularizing the single-stranded nucleic acid constructs of each nested set to generate single-stranded DNA circles to which the first adaptor sequence and the second adaptor sequence are joined; preparing a method of
[0007] In another aspect, disclosed herein is a method for generating double-stranded, adapted constructs for sequencing, optionally without the use of nanodrops, said method comprising: (i) amplifying a plurality of genomic fragments, each genomic fragment comprising a target sequence, to generate a plurality of sets of amplified nucleic acid fragments in the mixture, wherein the amplified nucleic acid fragments of each set share the same target sequence, and optionally amplification is performed using target-specific primers for each set; The method further comprises: (ii) contacting the amplified nucleic acid fragments with an enzyme to introduce cleavage in the amplified nucleic acid fragments; (iii) distributing the mixture of fragments into a plurality of aliquots; (iv) performing nick translation on aliquots of the fragments and synthesizing DNA strands under conditions such that the DNA strands synthesized in different aliquots have different lengths; each of the DNA strands comprises a target sequence portion at a first end and a second end; and wherein the DNA strands of the different aliquots share the same sequence near a first end and have different sequences near a second end; (v) for each aliquot, ligating a second adapter to the second end of the DNA strand synthesized in step (iv) via branched ligation, each second adaptor is a partially double-stranded adaptor comprising a first adaptor oligonucleotide and a second adaptor oligonucleotide; the first adaptor oligonucleotide and the second adaptor oligonucleotide are both complementary and hybridize to each other; each of the second adaptors comprises a positional barcode sequence; each ligation comprises joining the 5-prime end of a first adapter oligonucleotide of a second adapter to the second end of a synthesized DNA strand; wherein the first adaptor oligonucleotides ligated to the second ends of the synthesized DNA strands of different aliquots comprise different position barcode sequences, and the first adaptor oligonucleotides ligated to the second ends of the synthesized DNA strands of the same aliquot share the same position barcode sequence; (vi) combining the synthesized DNA strands ligated with the second adapter from different aliquots from step (v) in a single mixture; (vii) extending a primer hybridized to the first adaptor oligonucleotide ligated to the synthesized DNA strand to generate a double-stranded fragment with blunt ends; and (viii) optionally selecting double-stranded fragments of step (vii) in the size range of 200 bp to 1.5 kb, e.g., 300 bp to 1.2 kb, 300 bp to 1 kb, or 500 to 1000 bp, from a single mixture; and (ix) ligating a third adaptor to the blunt end of the double-stranded fragment, thereby generating a double-stranded adapted construct.
[0008] In another aspect, disclosed herein is a method for preparing multiple nested sets of adapted fragments, optionally without the use of Nanodrops, comprising: each adapted fragment is a single-stranded nucleic acid comprising a target sequence fragment having a first end and a second end, a 5-prime adaptor sequence, a 3-prime adaptor sequence, a primer sequence, and a barcode sequence; in each nested set of adapted fragments, the target sequence fragments have identical nucleotide sequences at a first end and differ from one another by truncation at a second end such that each nested set of adapted fragments comprises a plurality of target sequence fragments having different lengths; the first end is closer to the barcode sequence than the second end; The method comprises: (a) providing a population of single-stranded DNA concatemers in a single mixture reaction, each concatemer comprises a plurality of identical monomers, each monomer comprising a complement of a target sequence, a complement of a barcode sequence that identifies the concatemer, and a primer binding sequence shared by the population of single-stranded concatemers; the primer binding sequence comprises a sequence that is complementary to the primer sequence; the complement of the primer binding sequence and the barcode sequence are both 3-prime to the complement of the target sequence; (b) annealing a primer comprising a primer sequence to the primer binding sequences of the plurality of monomers of each of the plurality of concatemers; (c) extending at least a portion of the primers hybridized to the primer binding sequences with a DNA polymerase having 5' to 3' exonuclease activity and no strand displacement activity, said extension producing a plurality of extended primers, each extended primer comprising a target sequence fragment having a barcode sequence and a primer sequence; The extended primer is hybridized to the concatemer, the extended primers are separated by an interval; and (d) contacting the plurality of extended primers with 5-prime adapters comprising 5-prime adapter sequences, 3-prime adapters comprising 3-prime adapter sequences, a DNA ligase, and an exonuclease having single-stranded DNA exonuclease activity under conditions where the exonuclease degrades a portion of the target sequence fragment of the extended primer to generate shortened extended primers, 5-prime adapters linked to the 5' ends of the shortened extended primers, and 3-prime adapters linked to the 3' ends of the shortened extended primers; thereby generating a group of multiple nested sets of adapted fragments.
[0009] In another aspect, disclosed herein is a method for preparing multiple nested sets of adapted fragments, optionally without the use of Nanodrops, comprising: each adapted fragment is a single-stranded nucleic acid comprising a target sequence fragment having a first end and a second end, a 5-prime adaptor sequence, a 3-prime adaptor sequence, a primer binding sequence, and a complement of a barcode sequence; in each nested set of adapted fragments, the target sequence fragments have identical nucleotide sequences at a first end and differ from each other at a second end such that each nested set of adapted fragments comprises a plurality of target sequence fragments having different lengths; the first end is closer to the barcode sequence than the second end; The method comprises: (a) providing barcoded fragments comprising a barcode sequence, a target sequence, and a primer binding sequence, wherein the barcoded fragments are immobilized at one end on beads; (b) annealing a primer comprising a 5-prime adapter sequence to the primer binding sequence of the barcoded fragment, the 5-prime adapter sequence comprises i) the complement of the barcode sequence, and ii) a primer sequence that is complementary to the primer binding sequence of the barcoded fragment; (c) extending the primer to generate an extended primer comprising the complement of the target sequence fragment and the barcode sequence; (d) contacting the extended primer with a branched adapter comprising a 3-prime adapter sequence to generate an adapted fragment; (e) separating the adapted fragments from the barcoded fragments that remain immobilized on the beads; (f) repeating steps (b)-(e) for one or more cycles under controlled extension conditions to generate one or more adapted fragments; the adapted fragments generated from step (e) and the adapted fragments generated from step (f) constitute a nested set of adapted fragments; wherein the adapted fragments of each nested set comprise target sequence fragments having different lengths.
[0010] In another aspect, disclosed herein is a method for preparing multiple sets of adapted fragments, optionally without the use of Nanodrops, comprising: each adapted fragment is a single-stranded nucleic acid comprising a target sequence fragment having a first end and a second end, a 5-prime adaptor sequence, a 3-prime adaptor sequence, a primer binding sequence, and a complement of a barcode sequence; The method comprises: (a) providing barcoded fragments comprising a barcode sequence, a target sequence, and a primer binding sequence, wherein the barcoded fragments are immobilized at one end on beads; (b) annealing a primer comprising a 5-prime adapter sequence to the primer binding sequence of the barcoded fragment, the 5-prime adapter sequence comprises i) the complement of the barcode sequence, and ii) a primer sequence that is complementary to the primer binding sequence of the barcoded fragment; (c) extending the primer to generate an extended primer comprising the complement of the target sequence fragment and the barcode sequence; (d) contacting the extended primer with a first branched adapter comprising a 3-prime portion that includes a degenerate sequence region, thereby forming a first extension product that includes a degenerate sequence region at the 3-prime portion; hybridizing the 3-prime portion to the barcoded fragment via the degenerate sequence region; (e) extending the 3-prime portion of the first extension product to generate a second extension product; and (f) contacting the second extension product with a second branched adapter to generate an adapted fragment.
[0011] In another aspect, disclosed herein is a DNA complex comprising: (a) a barcoded fragment immobilized on a solid support, the barcode fragment comprising a barcode sequence and a target sequence; (b) a polynucleotide hybridized to the barcoded fragment, the polynucleotide comprises a 5-prime portion comprising the complement of the barcode sequence, a 3-prime portion comprising a target sequence fragment, A DNA complex comprising a polynucleotide in which a 5-prime portion and a 3-prime portion anneal to the barcoded fragment, leaving a central portion that does not anneal to the barcoded fragment, thereby forming a bubble.
[0012] In another aspect, disclosed herein is a composition comprising a nested set of adapted fragments, each comprising a barcode sequence and a target sequence fragment having a first end and a second end, a 5-prime adapter sequence, and a 3-prime adapter sequence; the target sequence fragments have identical nucleotide sequences at a first end and differ from one another by truncation at a second end such that the nested set of adapted fragments comprises target sequence fragments having a plurality of different lengths; A composition in which the adapted fragments share the same barcode sequence. [Brief explanation of the drawings]
[0013] The drawings and their description depict exemplary embodiments of the present disclosure. The methods and compositions provided in this disclosure are not limited to the embodiments shown in these drawings.
[0014] [Figure 1] Figure 1 illustrates one embodiment of the disclosed method. The top panel depicts an adapted, double-stranded genomic fragment containing a target sequence having a first end and a second end. The target sequence is flanked by a 3-prime adapter 1 and a 5-prime adapter 3. The first adapter contains a primer binding site and a barcode sequence, with the primer binding site located 3-prime relative to the barcode sequence (not shown).
[0015] [Figure 2] 2 illustrates one exemplary embodiment of a method in the present disclosure. Various method steps are shown, including adding barcodes to genomic fragments and amplifying the barcoded genomic fragments.
[0016] [Figure 3A] FIG. 3A shows one embodiment of a DNA circle-based scheme for generating single-stranded DNA circles from the amplified genomic fragments of FIG.
[0017] [Figure 3B] FIG. 3B shows another embodiment of a DNA circle-based scheme for generating single-stranded DNA circles from the amplified genomic fragments of FIG.
[0018] [Figure 3C] FIG. 3C shows another embodiment of a DNA circle-based scheme for generating single-stranded DNA circles from the amplified genomic fragments of FIG.
[0019] [Figure 4A] FIG. 4A shows one embodiment of a DNA circle-based scheme for generating double-stranded, adapted constructs from single-stranded DNA circles formed as shown in FIG. 3A or 3B.
[0020] [Figure 4B] FIG. 4B shows another embodiment of a DNA circle-based scheme for generating double-stranded, adapted constructs from single-stranded DNA circles formed as shown in FIG. 3A or FIG. 3B.
[0021] [Figure 5] FIG. 5 shows one embodiment of a linear DNA-based scheme for generating double-stranded, adapted constructs for sequencing.
[0022] [Figure 6A]Figures 6A and 6B illustrate the concatemer-based method of the present invention. Figure 1A shows that a double-stranded DNA molecule containing a barcode 110 and a target sequence 120 is denatured into single-stranded nucleic acid. The single-stranded nucleic acid is circularized and amplified by rolling circle replication to form concatemers containing multiple monomers, each containing a complement of the target nucleic acid sequence, a complement of the barcode sequence 121, and a primer binding sequence 131. In each monomer, a primer 130 is annealed to the primer binding sequence 131 that is 3-prime to the complement of the barcode sequence 111 and extended using a polymerase that does not have strand displacement activity but does have 5-prime to 3-prime exonuclease activity. The extended primers 150 are separated by a gap 160. Optionally, the gap 160 is widened by a gapping enzyme, resulting in a gap 170. If the primer 130 is an RNA primer, the gapping enzyme can be RNase H. An L-adapter is then ligated to the 5-prime of the extended primer, followed by a branched adapter ligation to the 3-prime of the extended primer in the presence of a 3-prime to 5-prime exonuclease. Note that while extension, gapping, and ligation are shown as separate steps, in some embodiments, reagents used for one or more of these reactions can be added simultaneously to a single reaction mixture. After denaturation, a nested set of single-stranded adapter fragments 191 is generated, each with an L-adapter sequence at the 5-prime and a branched adapter sequence at the 3-prime. Barcodes are located on the 5-prime portion of the adapted fragments. [Figure 6B] Same as above
[0023] [Figure 7A]Figures 7A-7B illustrate another concatemer-based method of the present invention. Similar to Figure 6, a polymerase is used to extend primer 230 annealed to primer binding sequence 231. The DNA polymerase does not have strand displacement activity, but does have 5-prime to 3-prime exonuclease activity. 210 is a barcode, and 220 is a target sequence. Unlike Figure 6, primer binding sequence 231 is 5-prime to the complement of barcode sequence 211 (primer binding sequence 131 is 3-prime to the complement of barcode sequence 111 in Figure 6). Also unlike Figure 6, ligation of L-adapter 280 and branched adapter 290 is performed in the presence of a 5-prime to 3-prime exonuclease (instead of a 3-prime to 5-prime exonuclease as in Figure 6). After denaturation, a nested set of adapted fragments 291 is formed, each with an L-adapter sequence at the 5-prime and a branched adapter sequence at the 3-prime. The adapted fragments include target sequence fragments 222, 223, 224, and 225. These target sequence fragments are generated in different lengths. Barcodes 210 are located on the 3-prime portions of the adapted fragments. [Figure 7B] Same as above
[0024] [Figure 8] Figure 8 illustrates one embodiment of the combinatorial scheme-based method of the present invention. Genomic DNA is first fragmented to generate staggered fragments with single-strand breaks, as disclosed in Figure 2 of U.S. Provisional Application No. 63 / 224,731. StLFR is performed to generate co-barcoded fragments containing a branched adapter sequence at one end and an L-adapter sequence at the other end. These co-barcoded fragments are then released from the beads and circularized and processed according to the procedures described in Figures 6A-6B or 7A-7B.
[0025] [Figure 9]Figure 9 shows another exemplary embodiment of the combinatorial scheme-based method of the present invention. A barcoded fragment comprising a barcode sequence 410, a primer binding sequence 433, and a target sequence 420 is immobilized on a bead, and the 3-prime end of the barcoded fragment is also immobilized on the bead. A tailed primer comprises an optional tail 431 that does not hybridize to the barcode fragment. The tailed primer comprises a primer sequence 432 that hybridizes to the primer binding sequence 433 and the complement of the barcode sequence 411 that hybridizes to the barcode sequence of the barcoded fragment. This tailed primer is extended to generate an extended tailed primer 435 that comprises a target sequence fragment 436 and the complement of the barcode sequence 411. A branched adapter 440 is ligated to the 3-prime end of the target sequence fragment 436 to generate an adapted fragment 450. The adapted fragment 450 is then separated from the barcoded fragment 460, which remains immobilized on the bead. The barcoded fragments are used as templates for subsequent extension cycles of tailed primer 430, generating multiple extended tailed primers containing target sequence fragments 451-453. The extension was controlled so that the target sequence fragments 451-453 had different lengths.
[0026] [Figure 10A]Figures 10A-10C show another exemplary embodiment of a combinatorial scheme-based method. Barcoded fragments and tailed primers are provided as described in Figure 9. Figure 10A shows that after an initial period of extension with regular deoxynucleotides, uracil is added to the reaction mixture to generate an extended tailed primer 550 containing uracil, followed by the addition of regular deoxynucleotides (e.g., deoxynucleotides without uracil) and a reversible terminator. The terminator can be added at different concentrations in different cycles to generate extended tailed primers containing target sequence fragments with different lengths. Figure 10B shows that the extended tailed primer 550 is then ligated to a branched adapter 540 to generate an adapted fragment 551. If a reversible terminator is used, it must be reversed before ligation. The adapted fragment 551 is then digested with USER to remove the uracil, leaving a gap 560 flanked by an exposed 3-prime end 570 and an exposed 5-prime end 580. An internal branched adaptor 551 is then ligated to the exposed 3-prime end 570, and an L-adaptor 552 is ligated to the gap-exposed 5-prime end 580. A splint oligo 590 is then hybridized to the 3-prime portion of the internal branched adaptor 551 and the 5-prime portion of the L-adaptor 552, allowing ligation between the two. This ligation forms a shortened adapted fragment 600 and a loop 591 of the barcoded fragment (which remains immobilized). The shortened adapted fragment 600 can be separated from the barcoded fragment upon denaturation. [Figure 10B] Same as above
[0027] [Figure 10C] Figure 10C shows fragments generated from multiple cycles of the process depicted in Figure 10B, where each shortened, adapted fragment contains a shortened target sequence fragment (e.g., 610 or 620, or 630) generated from a different cycle.
[0028] [Figure 11] Figure 11 shows that truncated target sequence fragments, e.g., 610, 620, and 630, have sequences that correspond to different regions of the target sequence generated from the process depicted in Figures 10A-10C. Because the truncated, adapted fragments all contain the same complement of barcode sequence 511, sequencing reads from these adapted fragments can be assembled based on the shared barcode sequence to achieve complete coverage across the entire target sequence.
[0029] [Figure 12] Figure 12 shows another exemplary embodiment of a combinatorial scheme-based method. A primer annealed to the primer binding sequence of a barcoded fragment is extended in a first extension step under extension-controlled conditions. The extended primer is ligated to a branched adapter 720 having a degenerate sequence region 730 in its 3-prime portion. The first branched adapter can hybridize to a random position of the barcoded fragment via the degenerate sequence region, forming a loop 740 and skipping replication of some random portion of the barcoded fragment. Next, a second extension step is performed under extension-controlled conditions by extending the 3-prime end of the first branched adapter 750 to form a second extension product. A second branched adapter is ligated to the 3-prime end of the second extension product to generate an adapted fragment. 710 is a barcode. 711 is the complement of the barcode.
[0030] [Figure 13A] 13A and 13B show another embodiment of the invention for preparing a nested set of target sequence fragments for a loop-mediated complete stLFR. [Figure 13B] Same as above
[0031] [Figure 14] FIG. 14 shows one embodiment of a loop-mediated complete stLFR.
[0032] [Figure 15] FIG. 15 shows an embodiment of preparing molecules generated from the loop-mediated complete stLFR shown in FIG. 14 for sequencing. DETAILED DESCRIPTION OF THE INVENTION
[0033] 1. Overview
[0034] The methods disclosed herein involve the preparation of libraries using massively parallel short-read sequencing to sequence entire long molecules. These long DNA molecules typically range in length from 1 to 20 kb, e.g., greater than 1,000 bp, or greater than 1,500 bp, or greater than 2,000 bp, or greater than 3,000 bp. These strategies disclosed herein do not require clonally barcoded beads and can be performed entirely in solution, i.e., genomic fragments and adapters are present in solution throughout the entire library preparation. Therefore, they can be conveniently used to barcode large numbers of molecules (e.g., 1 million to 10 million to 100 million to 1 billion molecules) in a single library at a lower cost than strategies requiring barcoded beads.
[0035] The methods disclosed herein generate a nested set of nucleic acid constructs for each genome fragment, and generate multiple nested sets for multiple genome fragments. The nucleic acid constructs can be single-stranded or double-stranded. Each nucleic acid construct in each nested set includes a barcode and a target sequence portion, and the nucleic acid constructs within each nested set have different lengths. The nucleic acid constructs in each nested set share a unique barcode sequence. The target sequence portion has a first end and a second end. The nucleic acid constructs in each nested set share an identical nucleotide sequence near the first end but differ in nucleotide sequence near the second end. This method allows sequencing of both the first and second ends of all nucleic acid constructs in a nested set, and the sequence reads can be assembled to generate sequence information for the entire long genomic DNA fragment. Various approaches to achieving this goal are described below.
[0036] In some approaches, this method provides a way to retain information that can be used to identify the location of each nucleic acid sequence (corresponding to each sequence read) within the original long DNA genomic molecule. This location information is useful for deciphering the sequence information of long DNA molecules with repetitive sequences.
[0037] 2. Definition
[0038] Components or reactions in a "single reaction mixture" means that the reaction occurs in a single mixture without compartmentalization into separate tubes, containers, aliquots, wells, chambers, or droplets during the tagging step. Components can be added simultaneously or in any order to create a single reaction mixture.
[0039] As used herein, the terms "first end" and "second end" are used to define the two ends of each nucleic acid molecule in a nested set of nucleic acid molecules. The target sequences near the first end of the nucleic acid molecule share the same nucleotide sequence, but the nucleotide sequences near the second end are different. In a double-stranded DNA molecule, the first end can be either a 5-prime end or a 3-prime end. Similarly, in a double-stranded DNA molecule, the second end can be either a 5-prime end or a 3-prime end. Relative to the second end of the same molecule, the first end is closer to the barcode sequence.
[0040] As used herein, a "unique molecular identifier" (UMI) refers to a sequence of nucleotides present in a DNA molecule that can be used to distinguish individual DNA molecules from one another. See, e.g., Kivioja, Nature Methods 9, 72-74 (2012). UMIs can be sequenced along with the DNA sequence to which they are associated to identify sequence reads that originate from the same source nucleic acid. The term "UMI" is used herein to refer to both the nucleotide sequence and the physical nucleotides of a UMI, as is clear from the context. A UMI can be random, pseudorandom, or partially random, or a non-random nucleotide sequence inserted into an adapter or otherwise incorporated into the source nucleic acid (molecule) to be sequenced. In some embodiments, each UMI is expected to uniquely identify any given source DNA molecule present in a sample. For purposes of this disclosure, the term "UMI" is used interchangeably with the term "barcode."
[0041] As used herein, the term "single-tube LFR" or "stLFR" refers to the process described, for example, in U.S. Patent Publication No. 2014 / 0323316 and Wang et al., Genome Research, 29: 798-808 (2019), the entire contents of each of which are incorporated herein by reference. In stLFR, multiple copies of the same unique barcode sequence (or "tag") are associated with each long nucleic acid fragment. In one embodiment of single-tube LFR, long nucleic acid fragments are labeled with barcodes at regular intervals. In one embodiment, barcodes are introduced into long nucleic acid molecules using one or more enzymes, such as transposase, nickase, and ligase. The barcode sequence between nucleic acid fragments can be conveniently performed, for example, within a single container, without compartmentalization. This process allows for the analysis of multiple individual DNA fragments without the need to separate the fragments into separate tubes, containers, aliquots, wells, or droplets during the tagging step.
[0042] As used herein, a "unique" barcode refers to a nucleotide sequence used to identify and distinguish a group of individual polynucleotides from other polynucleotides among a group of mixtures. For example, a unique barcode for a nested set of nucleic acid constructs means that the barcode sequence associated with one nested set is different from the barcode sequences associated with at least 90% of the entire nested set, more frequently at least 99% of the entire nested set, even more frequently at least 99.5% of the entire nested set, and most frequently at least 99.9% of the entire nested set. In some embodiments, unique barcodes are used to identify the location of a group of nucleic acid fragments relative to the genomic DNA from which they are derived. This type of barcode is also referred to as a location barcode in the present disclosure. In some cases, different groups of nucleic acid fragments, each with a unique location barcode, are present in a single mixture. See, for example,
[0316] in Figure 3C. In some cases, different groups of nucleic acid fragments, each carrying a unique location barcode, are contained in separate aliquots, which can then be combined into a single mixture. For example, see
[0504] and
[0505] in Figure 5.
[0043] The term "in solution" when used with respect to an adaptor (or any other nucleic acid construct or polynucleotide complex) used in the methods or compositions disclosed herein refers to the adaptor (or any other polynucleotide or polynucleotide complex) not being immobilized on a substrate and being free to move in solution. When used to describe a reaction, as in "a reaction performed in solution," it refers to the reaction occurring between nucleic acids that are entirely in solution.
[0044] The terms "adapted nucleic acid fragment" and "adapted fragment" are used interchangeably and refer to a polynucleotide comprising a target nucleic acid fragment and one or more adapter sequences.
[0045] The term "adapter sequence," as the context may dictate, refers to the sequence on either strand of the adapter. That is, "adapter sequence" can refer to both the sequence of the adapter on one strand and the complementary sequence on the second strand. Similarly, the term "barcode sequence" refers to the sequence of a barcode on one strand or its complementary strand.
[0046] The terms "reversible terminator nucleotide" and "reversible terminator" are used interchangeably and refer to a nucleotide having a 3-prime reversible blocking group. A "reversible blocking group" refers to a group that can be cleaved to provide a hydroxyl group at the 3' position of a nucleotide that can be linked to the 5-prime phosphate group of another nucleotide. The reversible blocking group can be cleaved by enzymes, chemical reactions, heat, and / or light. Exemplary nucleotides having 3-prime reversible blocking groups are known in the art and are disclosed in U.S. Pat. No. 10,988,501, the entire disclosure of which is incorporated herein by reference.
[0047] The term "target sequence" refers to the sequence information of a DNA molecule, such as a genomic DNA fragment. The methods and compositions provided herein can be used to determine the target sequence.
[0048] The term "target sequence portion" refers to the entire target sequence or a portion of the complement of a target sequence. Multiple nucleic acid fragments can contain sequences corresponding to different portions of the same target sequence.
[0049] The term "extended primer" refers to a DNA strand produced by extending a primer annealed to a template.
[0050] The term "copy" refers to generating a complementary nucleotide strand of the template by primer extension.
[0051] The term "corresponding" means that a DNA sequence has an identical or complementary sequence to another DNA sequence.
[0052] The term "near" when referring to a sequence near a reference point (e.g., a nucleotide sequence near a first end) refers to a nucleotide sequence within a specified length from the reference point. The specified length is typically less than 200 bases, less than 100 bases, less than 50 bases, less than 20 bases, or less than 10 bases. In some embodiments, the specified length is in the range of 1 to 50 bases, e.g., 1 to 30 bases, or 1 to 20 bases.
[0053] The term "exposed 5-prime" refers to the 5-prime end of a DNA fragment that is formed after the bond between two nucleotides of another contiguous DNA strand is broken. Similarly, the term "exposed 3-prime" refers to the 3-prime end of a DNA fragment that is formed only after the bond between two nucleotides of another contiguous DNA strand is broken.
[0054] As used herein, the term "length suitable for sequencing" refers to a DNA strand having a length equal to the length of the sequence read generated by MPS sequencing. This length is determined by the sequencing method, but generally, the length of a single DNA strand suitable for sequencing is within the range of 200 to 1.5 base pairs, e.g., 300 to 1000 base pairs, 300 to 500 base pairs, 400 to 600 base pairs, or 500 to 1000 base pairs. The length of a DNA duplex suitable for sequencing is within the range of 200 to 1.5 base pairs, e.g., 300 to 1000 base pairs, 300 to 500 base pairs, 400 to 600 base pairs, or 500 to 1000 base pairs.
[0055] The term "conjugated" as used in connection with a polynucleotide and a substrate (e.g., a bead) refers to the direct contact or covalent linkage of a polynucleotide (or one end of a polynucleotide) with the substrate. For example, a surface may have reactive functional groups that react with functionality on a polynucleotide molecule to form a covalent linkage. As an illustrative example, barcoded fragments are conjugated to the beads shown in Figure 9. The term "conjugated" can also be used to refer to the linking of one polynucleotide to another to form a single contiguous polypeptide, e.g., 551 and 552 in Figure 10B are conjugated to form a single contiguous, adapted fragment.
[0056] As used in this context, a "fragment" is single-stranded, but as discussed above and elsewhere herein, a fragment may hybridize with a complementary strand, e.g., to form a nucleic acid complex. The term "fragment" is generally used interchangeably with the term "polynucleotide."
[0057] As used in this context, "barcode region" refers to the region within a DNA molecule where the barcode or the complement of the barcode is located.
[0058] The term "barcoded fragment" refers to a fragment that includes a barcode sequence or the complement of a barcode sequence.
[0059] The term "branched adapter" refers to a partially double-stranded adapter that includes (i) a double-stranded blunt end comprising the 5'-end of one strand and the 3'-end of the complementary strand, and (ii) a single-stranded region comprising a barcode sequence. The 5'-end of the double-stranded region of the branched adapter can be ligated to the 3'-end of a nucleic acid fragment via a branched ligation, as further described below.
[0060] The term "nested set" refers to multiple nucleic acid fragments that (i) have different lengths, (ii) share the same nucleotide sequence at one end, and (iii) have a different nucleotide sequence at the other end due to truncation. An example of a nested set is shown in Figure 6B as 191.
[0061] The term "5-prime portion" of a polynucleotide refers to a contiguous nucleotide sequence region of a polynucleotide that includes a 5-prime end. The 5-prime portion may comprise at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, or at least 70% of the total length of the polynucleotide. The "5-prime portion" of a polynucleotide does not include the 3-prime end of the polynucleotide.
[0062] The term "3-prime portion" of a polynucleotide refers to the contiguous nucleotide sequence region of the polynucleotide that includes the 3-prime end. The 5-prime portion may comprise at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, or at least 70% of the total length of the polynucleotide. The "3-prime portion" of a polynucleotide does not include the 5-prime end of the polynucleotide.
[0063] The term "intermediate portion" of a polynucleotide refers to the portion between the 3-prime portion and the 5-prime portion.
[0064] The term "bubble" refers to a DNA structural configuration consisting of two DNA strands, containing a non-hybridizing region flanked by two double-stranded regions. The non-hybridizing regions contain two single-stranded loops that do not anneal to each other due to insufficient complementarity. Figure 10C shows a diagram of the bubble region.
[0065] The term "interval" refers to the space separating two single-stranded nucleic acid fragments.
[0066] The term "gap" refers to a widened interval (used interchangeably with the term "stretched"). An interval widens to form a gap. However, this does not necessarily mean that an interval is always of a smaller length than a gap. For example, a particular interval may be longer in length than a gap formed by widening another interval.
[0067] The process of preparing a library of long fragments for sequencing can be carried out according to various schemes.The following is an exemplary embodiment of the method.Those skilled in the art of molecular biology and sequencing, guided by this disclosure, will recognize many variations of the individual steps and reagents that can be incorporated into the following scheme.
[0068] 3. Method
[0069] 3.1 Adding adapters to the ends of nucleic acid molecules
[0070] A variety of approaches can be used to add adapter sequences to one or both ends of a nucleic acid molecule, e.g., a genomic fragment, for example, via adapter ligation, PCR amplification, and other methods known in the art.
[0071] In some approaches, each of the multiple genome fragments is linked to an adapter that contains a barcode sequence unique to each genome fragment. This unique barcode sequence can then be used to identify all reads originating from a particular genome fragment. Methods for labeling each genome fragment with a unique barcode are well known, see the section entitled "Barcodes," further described below.
[0072] In various approaches, each genomic fragment is ligated at one end to a first adaptor and at the other end to a third adaptor, and is amplified by extending a primer hybridized to the two adaptors. The terms "first," "second," and "third" are arbitrary and are used to refer to different adaptors. Unless specifically defined in the context of this disclosure, they do not imply a specific physical relationship between the positions appearing on the genomic fragment, nor do they imply a specific order of the adaptors used in this method.
[0073] In some approaches, the first adapter comprises a barcode sequence and a primer binding sequence, as described above, configured such that when a primer that binds to the primer binding sequence is extended, the extension product comprises the primer sequence, the barcode sequence, and the target sequence, in 5-prime to 3-prime order.
[0074] 3.2 Cyclization
[0075] Methods for generating DNA circles are known. In one exemplary embodiment, splint oligonucleotides, for example, 8-40 bases long, are annealed to both ends of a single-stranded molecule. These annealed oligos allow for a 1-10 base overlap between the two ends of the product. Ligation can then be performed using T4 DNA ligase to create a single-stranded circle with a small region of double-stranded DNA at the ligation site.
[0076] Circularization of single-stranded DNA molecules can be performed using methods well known in the art. In some approaches, splint oligos are then added to each of several single-stranded nucleic acid molecules, hybridizing to adapter sequences added to both ends of the target nucleic acid fragment. The single-stranded nucleic acids are then circularized in the presence of a ligase (e.g., T4 or Taq ligase). The DNA polymerase used for PCR can be any DNA polymerase with strand displacement activity, such as Phi29, Bst DNA polymerase, the Klenow fragment of DNA polymerase I, and Deep-Vent® DNA polymerase (NEB#MO258). These DNA polymerases are known to have different strengths of strand displacement activity. It is within the ability of one skilled in the art to select one or more DNA polymerases suitable for the methods and compositions disclosed herein.
[0077] 3.3 Aliquots
[0078] Some approaches disclosed herein involve aliquoting the reaction mixture. The term "aliquot" is used interchangeably with "pool" to refer to the distribution of the entire mixture. Different aliquots of the entire mixture are similar in volume and composition at the time the aliquots are formed. As used herein, different aliquots may be subjected to different processing steps, resulting in different compositions. For example, in some approaches disclosed herein, adapters with different positional barcodes are added to different aliquots, resulting in aliquots with different compositions. Preferably, the DNA fragments in each aliquot are of similar length. For aliquots containing long and short DNA, specific methods can be used to minimize overcoverage of short fragments. The products are amplified by PCR and divided into 10-20 pools, followed by controlled extension, ExoIII digestion, or controlled nick translation, with different pools undergoing different time periods. Short DNA fragments are fully extended to form blunt ends, and these blunt-ended fragments can be blocked from branch ligation using methods known in the art, such as DNA tailing or 3' blocking with terminal transferase. Exemplary methods for blocking short fragments from ligation are disclosed in WO2023001262, e.g., in Section 7.2 entitled "Remove excel adapters," the entire disclosure of which is incorporated herein by reference.
[0079] 3.4 Controlled Elongation
[0080] In some approaches, the method involves extending a primer hybridized to a DNA fragment under conditions that allow control of the extent of the extension reaction. These extension control conditions include, but are not limited to, selecting a polymerase(s) with appropriate polymerization rates or other properties, and using various reaction parameters, including (but not limited to) reaction temperature, reaction duration, primer composition, DNA polymerase, primer and nucleotide concentrations, additives, and buffer composition. In some cases, extension can be controlled by a mixture of reversible terminators and normal nucleotides for extension. The ratio of the amount of reversible terminator nucleotides to the amount of normal nucleotides can be adjusted to achieve the desired degree of extension. Generally, a higher ratio of the amount of reversible terminator nucleotides to the amount of normal nucleotides results in less than complete extension.
[0081] In some approaches, the amplified genome fragments are distributed into multiple aliquots, and each aliquot of the amplified genome fragments is subjected to different extension control conditions so that the extension products of different aliquots have different lengths. Each aliquot can be placed in a different container or well. Each aliquot can also be in a different distribution (e.g., droplet) within the same container.
[0082] The number of aliquots required depends on the length of the target sequence and the length of the sequence reads generated from the sequencing platform. Typically, the larger the amplicon size, the more aliquots are required. In one illustrative example, for a 5 kb amplicon with 500 bases per read (250 base paired-end read length or 500 base single-end read length), 10-20 aliquots are typically used in this method. In some approaches, there are at least three, at least four, at least five, at least six, at least seven, at least eight, or at least nine aliquots. In some approaches, the number of aliquots can range from 3-100, e.g., 5-50, 6-40, or 10-20.
[0083] To generate extension products in each aliquot, a primer is annealed to the primer binding sequence of the adapted genomic fragment, and the primer is extended to copy the barcode sequence and beyond, i.e., extend into the target sequence portion of the adapted genomic fragment. The extension reaction in each aliquot is controlled as described above, resulting in extension products with different sequences near their ends. In some approaches, extension of different aliquots is terminated at different times. In some approaches, each aliquot is extended for increasing amounts of time; for example, the first aliquot is extended for 2 minutes, the second aliquot for 4 minutes, and so on (Figure 3B). The extension time for each aliquot can range from 10 seconds to 20 minutes. Extension can also be controlled by limiting the nucleotide concentration, such that extension stops at 100-1000 bases as a result of exhausting the nucleotide supply. In some approaches, extension is performed using a polymerase with 5' to 3' exo-activity, which may be suitable for nick translation. A non-limiting example of a DNA polymerase that can be used in this method includes E. coli DNA polymerase 1.
[0084] 3.5 Controlled Digestion
[0085] In some approaches, the amplified genomic fragments are distributed into multiple aliquots and then digested with a nuclease. In some embodiments, the nuclease is a double-stranded DNA nuclease with 3' to 5' nuclease activity, such as ExoIII or Klenow. In some approaches, the digestion is controlled so that the lengths of the polynucleotides remaining after digestion are different in different aliquots. The extent of digestion can be controlled by parameters such as reaction temperature, digestion time, and nuclease concentration. In one exemplary approach, the digestion time in each aliquot is varied so that the polynucleotides remaining after digestion have different lengths. In one approach, the digestion of each aliquot is carried out over increasingly longer time intervals so that the fragments after digestion in different aliquots are progressively shorter in length, for example, so that the fragments in different aliquots differ in length by 500 bases. Then, after digestion via branched ligation in each aliquot, a second adaptor is ligated to the newly formed ends, thereby generating single-stranded nucleic acid constructs, each containing a portion of the target sequence flanked by a first adaptor sequence and a second adaptor sequence.
[0086] 3.6 Size Selection
[0087] In some approaches, the adaptor fragment or the amplified adaptor fragment of length that is suitable for sequencing is selected.The method of selecting DNA fragment with desired length is well known.One exemplary method is to use AMPure XP beads, for example, available from Pacific Biosciences (Menlo Park, CA), product number 100-265-900, to select the fragment with desired length.
[0088] 3.7 Amplification
[0089] Various methods involve amplification, such as the amplification of genomic fragments or adapted DNA fragments. Such amplification methods include, but are not limited to, multiple displacement amplification (MDA), polymerase chain reaction (PCR), ligation chain reaction (also known as oligonucleotide ligase amplification (OLA)), cycling probe technology (CPT), strand displacement assay (SDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), rolling circle amplification (RCR) (for circularized fragments), and invasive cleavage technology. Amplification can be performed after fragmentation, or before or after any of the steps outlined herein.
[0090] In some approaches, amplification is performed on adapted genomic fragments by extending primers annealed to the adapter sequences. In some approaches, genomic fragments with different target sequences are ligated to adapters at both ends, and the adapters share a common sequence. The genomic fragments are then amplified using primers hybridized to the adapters at both ends. In some approaches, at least one of the adapters contains a barcode.
[0091] In some approaches, amplification is performed using target-specific primers, i.e., primers that hybridize to target sequences in genomic DNA. In some approaches, the target-specific primers contain a common adapter tag with a random barcode to amplify a specific region.
[0092] In some approaches, amplification can be performed using multiplex PCR, i.e., multiple primer pairs targeting different target sequences in genomic DNA. In some approaches, amplification is multiplex PCR, in which 2-1000 different target regions are amplified in a single reaction using target-specific primers, such that the reaction mixture contains amplified genomic fragments with different target sequences.
[0093] In some approaches, adapted or genomic fragments can be amplified using rolling circle amplification (RCR). The genomic fragment is first denatured into a single-stranded nucleic acid molecule. A splint oligo is added and hybridizes to the adapter sequence adjacent to the target sequence. The single-stranded nucleic acid is then circularized in the presence of a ligase (e.g., T4 or Taq ligase). The DNA polymerase used for RCR can be any DNA polymerase with strand displacement activity; exemplary DNA polymerases include Phi29, Bst DNA polymerase, Klenow fragment of DNA polymerase I, and Deep-Vent® DNA polymerase (NEB#MO258). These DNA polymerases are known to have different strengths of strand displacement activity. Selecting one or more appropriate DNA polymerases for use in the present invention is within the capabilities of one of ordinary skill in the art.
[0094] 3.8 Nicking
[0095] In various embodiments of the present disclosure, genomic fragments or amplified genomic fragments (including fragments incorporating one or more adapter sequences) are combined with one or more nicking agents to create nicks in the genomic fragments. In some approaches, the nicking agent is an enzyme (commonly referred to as a "nickase"), such as an endonuclease that cleaves a phosphodiester bond within a polynucleotide strand or removes one or more adjacent nucleotides from a polynucleotide strand. In some cases, the nickase is a non-sequence-specific endonuclease that nicks the DNA strand at random locations. Non-limiting examples of nicking agents include Vibrio vulnificus nuclease (Vvn), shrimp dsDNA-specific endonuclease, and DNAse I. In some approaches, the nicking agent is a site-specific nuclease, such as a restriction endonuclease, or a sequence-specific nuclease that nicks DNA at its recognition sequence. Non-limiting examples of site-specific nickases include Nt.CviPII(CCD), Nt.BspQI, and Nt.BbvCI, described in Shuang-yong Xu, BioMol Concepts 2015; 6(4): 253-267, the entire disclosure of which is incorporated herein by reference.
[0096] In some approaches, the nicking agents disclosed herein are chemical nicking agents, non-limiting examples of which include the dipeptide seryl-histidine (Ser-His), Fe / H2O2, or Cu(II) complex / H2O2.
[0097] In some approaches, the methods use two or more nicking agents. In some approaches, the methods use two or more nicking agents from the same category of nicking agents, such as non-specific nickases, site-specific nickases, or chemical nickases. In some approaches, the methods use nicking agents from different categories.
[0098] The length of the genomic fragments separated by nicks after treatment can vary. Typically, the higher the concentration of the nicking agent, the more nicks the nicking agent will generate, resulting in shorter fragments. Longer treatment times will similarly generate more nicks and result in shorter fragments. By adjusting one or more of these parameters, the length of the fragments can be controlled within a desired range. In some approaches, the average length of the nucleic acid fragments resulting from nicking is 200 to 10,000 nucleotides, e.g., 200 to 500 nucleotides, 400 to 1,000 nucleotides, or 1,000 to 10,000 nucleotides. One exemplary embodiment of using a nicking agent to generate nicks in genomic fragments is shown in Figure 3A.
[0099] In some embodiments, the nick created by the nickase is extended (widened) by an exonuclease to form a gap. This process can be referred to as "gapping," and the exonuclease used in the process can be referred to as a "gapping enzyme." Examples of enzymes with 3' exonuclease activity include DNA polymerase I, Klenow fragment (when no nucleotides are present), exonuclease III, and others known in the art. Examples of enzymes with 5' exonuclease activity include Bst DNA polymerase, T7 exonuclease, truncated exonuclease VIII, lambda exonuclease, T5 exonuclease, and other exonucleases known in the art. Low-processivity exonucleases (i.e., exonucleases that remove nucleotides from the ends of polynucleotides at a relatively low rate) are preferred because they open short gaps (e.g., 2-7 bases, 3-10 bases, or 3-20 bases) that dissociate from the DNA and allow adapter ligation. When using exonucleases, protecting the DNA adapter from exonuclease digestion, if necessary, can be achieved by introducing phosphorothioate bonds between bases (or modified bases) at the 5' and 3' ends of the adapter.
[0100] Nicking and gapping of the amplified double-stranded genomic fragments generates DNA fragments of different lengths with exposed 3-prime ends, and a second adapter containing a second adapter sequence can be ligated to the 3-prime ends via branched ligation. This process produces at least some ligation products flanked by the first adapter sequence containing the barcode sequence and the second adapter sequence. The ligation products are separated from their hybridized complementary strands by denaturation, forming a nested set of single-stranded nucleic acid constructs containing portions of the target sequence.
[0101] 3.9 Nick Translation
[0102] Nick translation is performed at a nick in a DNA strand by a DNA polymerase (e.g., E. coli DNA polymerase I). DNA polymerases suitable for nick translation typically have three activities: (1) a 5' to 3' polymerase activity, which requires a single-stranded template and a primer with a 3' hydroxyl group and synthesizes a new nucleotide strand complementary to the template; (2) a 5' to 3' exonuclease activity, which degrades double-stranded DNA from the free 5' end; and (3) a 3' to 5' exonuclease activity, which degrades double- or single-stranded DNA from the free 3' hydroxyl end. This latter activity is a proofreading or editing function. On double-stranded DNA, the 3' to 5' exonuclease activity is blocked by the 5' to 3' polymerase activity. During nick translation, the 5' to 3' polymerase activity of a DNA polymerase adds a nucleotide to the 3'-OH created by the nicking, while simultaneously the 5' to 3' exonuclease activity removes a nucleotide from the 5' side of the nick. These concerted activities result in the removal of a nucleotide from the 5' side of the nick while adding a nucleotide to the 3' side of the nick. This results in the nick moving—or translating—along the DNA. See Susan J. Karcher, Molecular Biology, A Project Approach, 1995, pages 135-192, relevant portions of which are incorporated herein by reference. Nick translation can be used, for example, in various embodiments of the method of FIG. 13A.
[0103] 3.10 Branching and Concatenation
[0104] Branched ligation, also known as "3-prime ligation" or "3-prime branched ligation," utilizes the properties of T4 ligase to ligate double-stranded DNA adapters to the 3-prime ends of DNA at intervals or gaps. See Wang et al., DNA Research, 2019 Feb 1 16(1):45-53, the entire disclosure of which is incorporated herein by reference. Branched ligation is efficient for adapter ligation because degenerate single-stranded bases at the adapter ends do not need to hybridize to the gaps.
[0105] Adapters suitable for use in branched ligation typically include (i) a double-stranded blunt end containing a 5-prime end on one strand and a 3-prime end on the complementary strand, and (ii) a single-stranded region containing a barcode sequence. The double-stranded blunt end provides a 5-prime phosphate that can be ligated to the 3-prime of a target nucleic acid fragment via 3-prime branched ligation. In some embodiments, the double-stranded blunt end provides a 3-prime that is blocked from ligation by a dideoxynucleotide, a 3' phosphate group, a 3' overhang, or the like. 3-prime branched ligation covalently joins the 5-prime phosphate of the blunt-ended adapter (donor DNA) to the 3-prime hydroxyl end of the double-stranded DNA acceptor at the 3-prime recessed strand, gap, or interval. In contrast to conventional DNA ligation, 3-prime branched ligation does not require complementary base pairing. 3-prime branched ligation is described in Wang et al., DNA Res. 26(1):45-53, doi:10.1093 / dnares / dsy037; PCT Publication WO2019 / 217452; U.S. Patent Publication US2018 / 0044668, and International Application WO2016 / 037418, U.S. Patent Publication 2018 / 0044667, all of which are incorporated by reference for all purposes.
[0106] In various embodiments, branched ligation is used to join adapters to genome fragments. In some approaches, a nick is introduced into the amplified genome fragment to generate exposed 3-prime and 5-prime ends, and then a second adapter is ligated to the nick via branched ligation to form an adapted fragment. In some approaches, controlled extension or digestion is performed using the genome fragment as a template, and a second adapter is ligated to the newly formed 3-prime end of the extension product. Ligation thus generates an adapter fragment with a barcode sequence at one end and a second adapter sequence at the second end.
[0107] The adapters used in branched ligation may contain additional information useful for assembling sequence reads. In some approaches, the second adapter contains a positional barcode specific to each aliquot. Fragments incorporating the second adapter from different aliquots contain different positional barcode sequences, while fragments incorporating the second adapter from the same aliquot share the same positional barcode sequence. Aliquots of DNA fragments incorporating positional barcodes can be combined and sequenced. The presence of positional barcodes can be used to identify long genomic fragments containing highly repetitive sequences. For example, the same sequence read from two aliquots can be assigned as a duplicate of two different genomic locations rather than being mistakenly treated as a single sequence read from a single genomic location. The methods and compositions disclosed herein can accurately determine sequence information for highly repetitive sequences, making them useful for sequencing target sequences located at genomic loci where highly repetitive sequences are found, such as DNA fragments near telomeres. Methods and compositions using these positional barcodes may also be useful for sequencing duplicates of target genes correlated with disease states.
[0108] In some approaches, genomic fragments are ligated with a first adapter containing a barcode, then ligated with a second adapter via branched ligation. In some approaches, the second adapter contains a positional barcode, and branched ligation with the second adapter results in genomic fragments flanked by a first adapter sequence containing a barcode sequence unique to each genomic fragment and a second adapter sequence containing a positional barcode unique to each aliquot. Dual barcoding allows all aliquots from all nested sets to be sequenced together in a single reaction, significantly improving sequencing efficiency.
[0109] 3.11 Sequencing
[0110] The library of adapted fragments can be sequenced using sequencing methods well known in the art, including, but not limited to, polymerase-based sequencing by synthesis (e.g., HiSeq 2500 system, Illumina, San Diego, CA), ligation-based sequencing (e.g., SOLiD 5500, Life Technologies Corporation, Carlsbad, CA), ion semiconductor sequencing (e.g., Ion PGM or Ion Proton Sequencer, Life Technologies Corporation, Carlsbad, CA), zero-mode waveguide (e.g., PacBio RS sequencer, Pacific Biosciences, Menlo Park, CA), nanopore sequencing (e.g., Oxford Nanopore Technologies Ltd., Oxford, United Kingdom), pyrosequencing (e.g., 454 Life Sciences, Branford, CT), or other sequencing technologies. Some of these sequencing technologies are short read technologies, while others generate long reads, such as GS FLX+ (454 Life Sciences; up to 1000 bp), PacBio RS (Pacific Biosciences; approximately 1000 bp), and nanopore sequencing (Oxford Nanopore Technologies Ltd.; 100 kb). For haplotype phasing, longer reads are advantageous and require fewer computations, but there may be a higher error rate in longer reads and errors may need to be identified and corrected according to the methods described herein before haplotype phasing.
[0111] In some approaches, sequencing is performed using combinatorial probe-anchor ligation (cPAL), as described, for example, in U.S. Patent No. 20140051588 and U.S. Patent No. 20130124100, both of which are incorporated by reference in their entireties for all purposes.
[0112] In some approaches, sequencing is performed using a DNBseq sequencer. The adapted fragments or their amplification products are denatured to generate single-stranded molecules. These loops are then used to fabricate DNA nanoballs (DNBs) for the DNB sequencer.
[0113] In some approaches, the adapted fragments or their amplification products are sequenced on Illumina or other systems that do not require circularization.
[0114] In some approaches, the sequencing is paired-end sequencing, which involves sequencing from either end of the same DNA fragment. In some approaches, a first read is generated by extending a sequencing primer annealed to an adapter sequence closer to the first end of the target sequence fragment than the second end ("first read sequencing"), and a second sequencing read is generated by extending a sequencing primer annealed to an adapter sequence closer to the second end of the target sequence fragment than the first end ("second read sequencing"). The first read sequencing will generate a barcode sequence. The second read sequencing will generate overlapping reads that substantially or completely cover molecules up to 500 bp, 700 bp, or 1000 bp in length. These overlapping sequencing reads will be clustered based on the barcode sequences determined by the first read sequencing in a de novo assembly.
[0115] In some approaches, the sequencing is single-end sequencing, where the sequence information of a genomic fragment is determined based solely on the first read sequencing.
[0116] 3.12 Assembly of sequence information
[0117] Sequence reads from the same nested set of nucleic acid constructs (derived from a single genomic fragment being sequenced and linked to adapters with unique barcode sequences) can be aligned based on the presence of the same barcode sequence. The sequence reads, each containing sequence information near the first end (which is the same end) and the second end (which is variable), are assembled to provide the full-length sequence of the long genomic DNA fragment.
[0118] 4. Composition
[0119] 4.1 Sample
[0120] The sample containing target nucleic acid can be obtained from any suitable source. For example, the sample can be obtained or provided from any organism of interest. Such organisms include, for example, plants; animals (e.g., mammals, including humans and non-human primates); or pathogens, such as bacteria and viruses. In some cases, the sample can be or can be obtained from the cells, tissues, or polynucleotides of the population of such organisms of interest. As another example, the sample can be a microbiome or a microbiota. Optionally, the sample is an environmental sample, such as a water, air, or soil sample.
[0121] Samples from a target organism or a group of such target organisms include, but are not limited to, bodily fluid samples (including but not limited to blood, urine, serum, lymph, saliva, anal and vaginal secretions, sweat and semen); cells; tissues; biopsies, research samples (for example, products of nucleic acid amplification reactions such as PCR amplification reactions); purified samples, for example, purified genomic DNA; RNA preparations; and raw samples (bacteria, viruses, genomic DNA, etc.).Methods for obtaining target polynucleotides (for example, genomic DNA) from organisms are well known in the art.
[0122] 4.2. Target nucleic acid
[0123] As used herein, the term "target nucleic acid" (or polynucleotide) or "nucleic acid of interest" refers to any nucleic acid (or polynucleotide) suitable for processing and sequencing by the methods described herein. In some approaches, the target nucleic acid is a genomic fragment generated by fragmenting genomic DNA extracted from a sample. While genomic fragments are used to describe the methods and compositions disclosed herein, it should be noted that sequencing libraries can also be prepared using these methods and compositions to sequence any target nucleic acid or fragment thereof, including those containing nucleotide modifications, such as nucleotide analogs.
[0124] Nucleic acids can be single-stranded or double-stranded and can include DNA, RNA, or other known nucleic acids. Target nucleic acids can be from any organism, including, but not limited to, viruses, bacteria, yeast, plants, fish, reptiles, amphibians, birds, and mammals (including, but not limited to, mice, rats, dogs, cats, goats, sheep, cows, horses, pigs, rabbits, monkeys and other non-human primates, and humans). Target nucleic acids can be obtained from an individual or from multiple individuals (i.e., a population). Samples from which nucleic acids are obtained can contain nucleic acids from a mixture of cells or even organisms, such as a human saliva sample containing human cells and bacterial cells; a mouse xenograft containing mouse cells and cells from a transplanted human tumor. Target nucleic acids can be unamplified or amplified by any suitable nucleic acid amplification method known in the art. The target nucleic acid can be purified according to methods known in the art to remove cellular and intracellular contaminants (such as lipids, proteins, carbohydrates, and nucleic acids other than the nucleic acid to be sequenced), or can be unpurified, i.e., contain at least some cellular and intracellular contaminants, including, but not limited to, intact cells that are disrupted to release their nucleic acids for processing and sequencing.The target nucleic acid can be obtained from any suitable sample using methods known in the art.Such samples include, but are not limited to, biological samples, such as tissues, isolated cells or cell cultures, bodily fluids (including, but not limited to, blood, urine, serum, lymph, saliva, anal and vaginal secretions, sweat, and semen); and environmental samples, such as air, agricultural, water, and soil samples.
[0125] The target nucleic acid can be genomic DNA (e.g., from a single individual), cDNA, and / or a complex nucleic acid containing nucleic acids from multiple individuals or genomes. Examples of complex nucleic acids include the microbiota in the bloodstream of a pregnant mother, circulating fetal cells (see, for example, Kavanagh et al., J. Chromatol. B 878: 1905-1911, 2010), and circulating tumor cells (CTCs) from the bloodstream of a cancer patient. In one embodiment, such complex nucleic acids have a complete sequence containing at least 1 gigabase (Gb) (the diploid human genome contains approximately 6 Gb of sequence).
[0126] In some cases, the target nucleic acid is a genomic fragment. In some approaches, the genomic fragment is longer than 10 kb, for example, 10-100 kb, 10-500 kb, 20-300 kb, 50-200 kb, 100-400 kb, or longer than 500 kb. In some cases, the target nucleic acid is 5,000-100,000 kb. In some approaches, the target nucleic acid is 500 bases to 50,000 bases in length, for example, 1,000 bases to 20,000 bases, or 5,000 bases to 10,000 bases. The amount of DNA (e.g., human genomic DNA) used in a single mixture can be <10 ng, <3 ng, <1 ng, <0.3 ng, or <0.1 ng of DNA. In some approaches, the amount of DNA used in a single mixture can be less than 3,000x, e.g., less than 900x, less than 300x, less than 100x, or less than 30x the amount of haploid DNA. In some approaches, the amount of DNA used in a single mixture can be at least 1x the amount of haploid DNA, e.g., at least 2x, or at least 10x the amount of haploid DNA.
[0127] The target nucleic acid can be isolated using conventional techniques, such as those disclosed in Sambrook and Russell, Molecular Cloning: A Laboratory Manual, supra. In some cases, particularly when small amounts of nucleic acid are employed in a particular step, it is advantageous to provide carrier DNA, e.g., unrelated circular synthetic double-stranded DNA, for use in admixture with the sample nucleic acid whenever only small amounts of sample nucleic acid are available and there is a risk of loss through nonspecific binding, e.g., to container walls, etc.
[0128] According to some embodiments of the present invention, genomic DNA or other complex target nucleic acids are obtained from individual cells or small numbers of cells, with or without purification, by any known method.
[0129] As described above, the disclosed methods are useful for sequencing long nucleic acid fragments. Long fragments of genomic DNA can be isolated from cells by any known method. Protocols for isolating long genomic DNA fragments from human cells are described, for example, in Peters et al., Nature 487:190-195 (2012). In one embodiment, cells are lysed and intact nuclei are pelleted by a gentle centrifugation step. The genomic DNA is then released via digestion with proteinase K and RNase over several hours. The material can be treated to reduce the concentration of remaining cellular waste products, for example, by dialysis and / or dilution for a period of time (i.e., 2-16 hours). Because such methods do not require the use of many disruptive processes (e.g., ethanol precipitation, centrifugation, and vortexing), the genomic nucleic acid remains largely intact, resulting in a majority of fragments greater than 150 kilobases in length. In some approaches, the fragments are approximately 5 to approximately 750 kilobases in length. In further embodiments, fragments are about 150 to about 600, about 200 to about 500, about 250 to about 400, and about 300 to about 350 kilobases in length. The smallest fragments that can be used for haplotyping are about 2 to 5 kb; although there is no maximum theoretical size, fragment length can be limited by shearing due to manipulation of the starting nucleic acid preparation.
[0130] In other embodiments, long DNA fragments are isolated and manipulated in a way that minimizes shearing or absorption of the DNA into the container, including, for example, isolating cells in agarose gel plugs or agarose in oil, or using specially coated tubes and plates.
[0131] According to another embodiment, to obtain uniform genome coverage in the case of samples containing a small number of cells (e.g., 1, 2, 3, 4, 5, 10, 10, 15, 20, 30, 40, 50, or 100 cells from a microbial biopsy or circulating tumor or fetal cells), all long fragments obtained from the cells are barcoded using the methods disclosed herein.
[0132] Barcode
[0133] According to one embodiment, a barcode-containing sequence having two, three, or more segments is used, of which, for example, one is a barcode sequence. For example, the introduced sequence may contain one or more regions of a known sequence and one or more regions of a degenerate sequence acting as a barcode(s) or tag(s). The known sequence (B) may include, for example, a PCR primer binding site, a transposon end, a restriction endonuclease recognition sequence (e.g., a site for a rare cutter, such as Not I, Sac II, Mlu I, BssH II, etc.), or other sequences. The degenerate sequence (N) acting as a tag is long enough to provide a population of different sequence tags equal to or preferably greater than the number of fragments of the target nucleic acid being analyzed. The higher the N value, the less likely two molecules will share the same barcode.
[0134] According to one embodiment, the barcode-containing sequence comprises one region of a known sequence of any selected length. According to another embodiment, the barcode-containing sequence comprises two regions of a known sequence of a selected length, namely, B and B, flanking a region of a degenerate sequence of a selected length. n N n B n where N can be of a length sufficient to tag long fragments of target nucleic acid, for example, but not limited to, N=10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20, and B can be of any length that accommodates desired sequences such as transposon ends, primer binding sites, etc. For example, such an embodiment may include B 20 N 15 B 20 It could be.
[0135] In one embodiment, a two- or three-segment design is utilized for the barcodes used to tag the long fragments. This design allows for a wider range of possible barcodes by allowing combined barcode segments to be generated by linking different barcode segments together to form a complete barcode segment or by using segments as reagents in oligonucleotide synthesis. This combinatorial design provides a larger repertoire of possible barcodes while reducing the number of full-size barcodes that need to be generated. In a further embodiment, unique identification of each long fragment is achieved with an 8-12 base pair (or longer) barcode.
[0136] In one embodiment, two different barcode segments are used. The A and B segments are easily modified to each contain a different half-barcode sequence, generating thousands of combinations. In a further embodiment, the barcode sequences are incorporated onto the same adapter. This can be achieved by splitting the B adapter into two parts, each with a half-barcode sequence separated by a common overlapping sequence used for ligation. The two tag components each contain 4-6 bases. An 8-base (2 x 4 base) tag set can uniquely tag 65,000 sequences. Both the 2 x 5 base and 2 x 6 base tags can include the use of degenerate bases (i.e., "wildcards") to achieve optimal decoding efficiency.
[0137] In a further embodiment, unique identification of each sequence is achieved using an 8-12 base pair error-correcting barcode, which can have a length of, by way of example and not limitation, 5-20 informative bases, typically 8-16 informative bases.
[0138] 4.4.UMI
[0139] In various embodiments, unique molecular identifiers (UMIs) are used to distinguish individual DNA molecules from one another. A collection of adapters is generated, each with a UMI. These adapters are attached to the fragments to be sequenced or other source DNA molecules, and each sequenced molecule has a UMI that helps distinguish it from all other fragments. In such implementations, a large number of different UMIs (e.g., thousands to millions) can be used to uniquely identify DNA fragments in a sample. An exemplary embodiment of a method using UMIs is described in Example 2.
[0140] The UMI is long enough to ensure the uniqueness of each and every source DNA molecule. In some approaches, the unique molecular identifier is about 3-12 nucleotides in length, or 3-5 nucleotides in length. In some cases, each unique molecular identifier is about 3-12 nucleotides in length, or 3-5 nucleotides in length. Thus, a unique molecular identifier can be 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, or more nucleotides in length.
[0141] The process of preparing a long fragment library for sequencing can be carried out according to various schemes. These schemes can be used to generate nested sets of nucleic acid constructs for each genome fragment, allowing sequencing near both ends of each nucleic acid construct in each nested set. The nucleic acid constructs in each nested set can be either single-stranded or double-stranded. These approaches allow for efficient generation of sequence information for long genome fragments. Some of these approaches involve the fabrication of DNA circles. Other approaches use linear DNA molecules. The following are exemplary embodiments of the method. Those skilled in the art of molecular biology and sequencing technology, guided by the present disclosure, will recognize numerous variations of the individual steps and reagents that can be incorporated into the following scheme.
[0142] 4.5. Barcode beads
[0143] The beads are barcoded with adapter barcode oligonucleotides immobilized thereon. Each bead contains multiple adapters, and therefore multiple barcode oligonucleotides. Each barcode oligonucleotide contains at least one barcode. Barcode oligonucleotides on the same bead share the same barcode sequence, while barcode oligonucleotides on different beads have different barcode sequences. As such, each bead carries multiple copies of a unique barcode sequence that can be transferred to target nucleic acid fragments using the methods described above.
[0144] The beads used can have diameters ranging from 1 to 20 μm, alternatively 2 to 8 μm, 3 to 6 μm, or 1 to 3 μm (e.g., about 2.8 μm). For example, the spacing between barcoded oligonucleotides on the beads can be at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7 nm. In some embodiments, the spacing is less than 10 nm (e.g., 5 to 10 nm), less than 15 nm, less than 20 nm, less than 30 nm, less than 40 nm, or less than 50 nm. In some embodiments, the number of different barcodes used per mixture can be >1 M, >10 M, >30 M, >100 M, >300 M, or >1 B. As discussed below, very large numbers of barcodes can be generated for use in the present invention, for example, using the methods described herein. In some embodiments, the number of different barcodes used per mixture can be >1M, >10M, >30M, >100M, >300M, or >1B, sampled from a pool of at least 10-fold greater diversity (e.g., >10M, >0.1B, >0.3B, >0.5B, >1B, >3B, >10B different barcodes on beads). In some embodiments, the number of barcodes per bead is between 100k and 10M (e.g., between 200k and 1M, 300k and 800k, or about 400k).
[0145] In some embodiments, the barcode region is about 3-15 nucleotides in length, e.g., 5-12, 8-12, or 10 nucleotides in length. In some cases, each barcode in the barcode region is about 3-12 nucleotides in length, or 3-5 nucleotides in length. Thus, a barcode, whether a sample barcode, cell barcode, or other barcode, can be 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more nucleotides in length. In one example, each barcode region includes three barcodes, each consisting of 10 bases, and the three barcodes are separated by 6 bases of common sequence.
[0146] The barcode beads are transferred to the target nucleic acid sequence. In some embodiments, the transfer occurs at regular intervals via ligation of the 3' end of an adapter oligonucleotide to the nucleic acid fragment created by the disclosed nicking and gapping.
[0147] In some embodiments, barcode beads are constructed using a split-and-pool ligation strategy using three sets of double-stranded barcode DNA molecules. In some embodiments, each set of double-stranded barcode DNA molecules consists of 10 base pairs, and the three sets have different nucleic acid sequences. An exemplary split-and-pool ligation method for generating barcode beads is described in PCT Publication No. 2019 / 217452, the disclosure of which is incorporated herein by reference in its entirety. Figures 12 and 13 of WO 2019 / 217452 also illustrate the split-and-pool methodology. In one approach, a common adapter sequence containing a PCR primer annealing site was attached to Dynabeads™ M-280 streptavidin (ThermoFisher, Waltham, MA) magnetic beads with a 5' double biotin linker. Three sets of 1,536 barcode oligos containing regions of overlapping sequence were constructed by Integrated DNA Technologies (Coralville, IA). Ligations were performed in 384-well plates in 15 μL reactions containing 50 mM Tris-HCl (pH 7.5), 10 mM MgCl2, 1 mM ATP, 2.5% PEG-8000, 571 units of T4 ligase, 580 pmol of barcode oligo, and 65 million M-280 beads. Ligation reactions were incubated on a rotator for 1 hour at room temperature. During ligation, beads were pooled into a single container by centrifugation, collected on the side of the container using a magnet, and washed once with high-salt wash buffer (50 mM Tris-HCl (pH 7.5), 500 mM NaCl, 0.1 mM EDTA, and 0.05% Tween 20) and twice with low-salt wash buffer (50 mM Tris-HCl (pH 7.5), 150 mM NaCl, and 0.05% Tween 20). The beads were resuspended in 1x ligation buffer and distributed into a 384-well plate, and the ligation step was repeated.
[0148] In one aspect, the present invention provides a composition comprising beads having adapter oligonucleotides comprising clonal barcodes attached thereto, the composition comprising over 3 billion distinct barcodes, wherein the barcodes are tripartite barcodes having the structure 5'-CS1-BC1-CS2-BC2-CS3-BC3-CS4. In some embodiments, CS1 and CS4 are longer than CS2 and CS3. In some embodiments, CS2 and CS3 are 4-20 bases long, CS1 and CS4 are 5 or 10-40 bases long (e.g., 20-30), and the BC sequence is 4-20 bases long (e.g., 10 bases). In some embodiments, CS4 is complementary to a splint oligonucleotide. In some embodiments, the composition comprises a bridging oligonucleotide. In some embodiments, the composition comprises a bridging oligonucleotide, beads comprising the tripartite barcodes described above, and genomic DNA comprising a hybridization sequence having a region complementary to the bridging oligonucleotide.
[0149] Other sources of clonal barcodes, such as beads or other supports associated with multiple copies of tags, can be prepared by emulsion PCR or CPG (controlled pore glass) or chemical synthesis, and other particles can be prepared with copies of the corresponding barcodes prepared. A population of tag-containing DNA sequences can be PCR-amplified on beads in a water-in-oil (w / o) emulsion by known methods. For example, see Tawfik and Griffiths Nature Biotechnology 16: 652-656 (1998); Dressman et al., Proc. Natl. Acad. Sci. USA 100:8817-8820, 2003; and Shendure et al., Science 309:1728-1732 (2005). This results in multiple copies of each single tag-containing sequence on each bead.
[0150] Another method for generating a source of clonal barcodes is the "mix and split" combinatorial process by oligonucleotide synthesis on microbeads or CPGs. This process can be used to create a set of beads, each carrying a population of barcode copies. For example, on average, 100 beads each contain approximately 1 billion copies of all B16-B ... 20 N 15 B 20 To create the , we start with about 100 billion beads and then mix them all together to create the B 20 Synthesize the common sequences (adapters), split them into 1024 synthesis columns to make a different 5-mer on each, then mix them together, then split them again into 1024 columns to make additional 5-mers, then do it again to complete N15, then mix them together, and add the final B 20 can be synthesized as the second adapter. Thus, in 3050 synthesis, one large emulation PCR reaction would generate approximately 1,000 billion beads (1 12 A "clonal" set of barcodes can be generated that is identical to the set of beads (each bead) so that only one in ten beads has the starting template (the other nine do not), preventing having two templates with different barcodes per bead.
[0151] An exemplary process for barcode sequence assembly is described in PCT Application Publication No. 2019 / 217452, the disclosure of which is incorporated herein by reference.
[0152] 4.6 Reaction mixture
[0153] Provided herein is a reaction mixture useful for preparing a library of polynucleotides. The reaction mixture includes: 1) a polymerase lacking 5'-3' exonuclease activity and no strand displacement activity; and 2) a DNA complex hybridizing to one or more monomers of a DNA concatemer and containing multiple fragments separated by nicks or gaps. In some embodiments, some or all of the fragments are generated by extending an RNA primer, thus incorporating an RNA sequence at the 5-prime end. In some embodiments, the reaction mixture further includes one or more gapping enzymes as described herein. In some embodiments, the gapping enzyme has 5' to 3' exonuclease activity. In some embodiments, the gapping enzyme has 3' to 5' exonuclease activity.
[0154] In some embodiments, each fragment is ligated to an L-adapter at its 5-prime end and to a branching adapter at its 3-prime end.
[0155] Also disclosed herein are DNA complexes comprising a barcoded fragment immobilized on a solid support (e.g., a bead) and a fragment hybridized to the barcoded fragment. In some embodiments, the fragment comprises multiple uracils. In some embodiments, the fragment comprises a 5-prime portion, a 3-prime portion, and an intermediate portion located therebetween, wherein the intermediate portion of the fragment does not hybridize to the barcoded fragment. In some embodiments, the 5-prime portion of the fragment is an adapter sequence, and the 3-prime portion of the fragment comprises a branched adapter sequence. One exemplary embodiment is shown in Figure 10A.
[0156] Also provided herein are compositions comprising a set of DNA fragments having overlapping target sequences. In some embodiments, the fragments vary in target sequence length but share the same barcode sequence. In certain embodiments, the fragments share a common adapter sequence at the 5-prime end and a common adapter sequence at the 3-prime end. In some embodiments, a set of fragments share a common sequence in the 5-prime portion that includes the barcode sequence. An exemplary embodiment is shown in Figure 6B. In some embodiments, a set of fragments share a common sequence in the 3-prime portion that includes the barcode sequence. An exemplary embodiment is shown in Figure 7B.
[0157] Also provided herein is a composition comprising multiple nested sets of single-stranded DNA loops, each loop comprising a target sequence portion flanked by a first adapter sequence and a second adapter sequence. The first adapter sequence comprises, from 5' to 3', a primer binding sequence, a barcode sequence, and a first hybridization sequence, and the second adapter sequence comprises a second hybridization sequence. The first hybridization sequence and the second hybridization sequence hybridize to each other, thereby forming a loop. Each target sequence portion has a first end and a second end, and the distance between the first end and the barcode sequence is shorter than the distance between the second end and the barcode sequence. The single-stranded nucleic acid constructs of each nested set share the same barcode sequence, and the single-stranded nucleic acid constructs of different nested sets have different barcode sequences. For each nested set of single-stranded nucleic acid constructs, the target sequence portions of that nested set have identical nucleotide sequences near a first end and differ from each other by truncations near a second end, such that the nested set of single-stranded nucleic acid DNA loops includes multiple target sequence portions having different lengths.
[0158] 5. DNA ring-based scheme
[0159] In some approaches, the method uses a DNA circle-based approach. This method involves circularizing each nested set of nucleic acid constructs so that the two ends of each nucleic acid construct are joined together. See Figure 1. Each nucleic acid construct contains a target sequence flanked by a first adaptor ("Adapter 1") sequence and a third adaptor ("Adapter 3") sequence. By circularizing the nucleic acid constructs, sequences near both ends are included in a single sequence read. Target sequence portions from different nucleic acid constructs of the same nested set can be assembled to generate sequence information corresponding to the entire target sequence of a genomic fragment. This scheme is illustrated in more detail below.
[0160] 5.1 Ligation of adapters to both ends of genomic fragments and amplified genomic fragments
[0161] An example of adding adapters to both ends of genomic fragment
[0201] is shown in Figure 2. Step (i) shows the ligation of a double-stranded genomic fragment to double-stranded adapters at both ends to generate an adapted genomic fragment
[0202] . Step (ii) shows the amplification of the adapted double-stranded genomic fragment. Amplification can be performed using primers that hybridize to the first and third adapter sequences (not shown). As a result of this step, the amplified genomic fragment
[0203] has blunt ends.
[0162] 5.2 Generation and Circularization of Nested Sets of Single-Stranded Nucleic Acid Constructs Containing Target Sequences
[0163] In the circularization approach, the amplified genomic fragments are processed to generate nested sets of single-stranded nucleic acid constructs (e.g.,
[0303] in Figure 3A), and the single-stranded nucleic acid constructs are circularized. Each single-stranded DNA construct contains a target sequence portion having a first end and a second end. The target sequence portion is adjacent to a first adapter sequence abutting the first end. A second adapter sequence (adapter 2) is inserted adjacent to the second end. In some approaches, the first adapter sequence is located at the 5-prime position of the single-stranded DNA construct, and the second adapter sequence is located at the 3-prime position. The first adapter sequence contains a barcode sequence that is unique to each nested set; i.e., single-stranded DNA constructs within the same nested set share the same barcode sequence, while single-stranded DNA constructs from different nested sets have different barcode sequences. In some cases, the first end of the single-stranded DNA construct is closer to the barcode sequence than the second end.
[0164] Various approaches can be used to generate single-stranded nucleic acid constructs from amplified genomic fragments. In some approaches, amplified genomic fragments are contacted with a nicking agent to introduce nicks into the target sequence, generating a nested set of single-stranded DNA constructs. A second adaptor is then ligated at the nicks via branched ligation. An example of such an approach is shown in Figure 3A. Step (iii) in Figure 3A shows contacting the amplified genomic fragments (generated from step (ii) in Figure 2) with a nicking agent to generate nicks at random positions in the target sequence of the amplified genomic fragment. Each amplified genomic fragment can be nicked one or more times, and the resulting nicks can be extended using one or more exonucleases to form gaps. This process produces one or more fragments, but only one of the fragments contains the first adaptor sequence (including the barcode sequence
[0311] ). Many parameters can affect the length of the nucleic acid fragments separated by the nicks and / or gaps. Typically, the higher the concentration of the nicking agent and the longer the treatment time with the nicking agent, the shorter the fragments. By adjusting one or more of these parameters, the length of the fragments can be controlled within a desired range. In some embodiments, the average length of the nucleic acid fragments generated by nicking is between 200 and 10,000 nucleotides, e.g., between 200 and 500 nucleotides, or between 400 and 1,000 nucleotides, or between 1,000 and 10,000 nucleotides. Step (iv) depicts ligating a second adaptor comprising a second adaptor sequence to the nick via branched ligation to form ligation products
[0302] , each of a portion of the ligation products comprising the first adaptor sequence and the second adaptor sequence. Step (v) depicts denaturing the ligation products to form single-stranded nucleic acid constructs
[0303] , each of a portion of the ligation products comprising the first adaptor sequence and the second adaptor sequence. The population of single-stranded nucleic acid constructs represents a nested set of constructs comprising the target sequence.
[0165] In some other approaches, generating a nested set of single-stranded DNA constructs involves annealing a primer to the primer binding sequence of a first adapter ligated to the genomic fragment and extending the primer to generate primer extension products. An example of such an approach is shown in Figure 3B. Step (iii) of Figure 3B illustrates distributing the amplified, double-stranded, adapted genomic fragment
[0203] into multiple aliquots, as shown in Figure 2. Step (iv) involves denaturing the amplification product of each aliquot to form a single-stranded molecule
[0304] and then hybridizing a primer to the primer binding sequence. The primer is extended in the presence of a polymerase and dNTPs, and the extension is controlled so that the extension products of different aliquots have different lengths, thereby forming a nested set of extension products. Each extension product has a first end (where the primer begins) and a second end (where the extension ends), and the extension products share the same sequence near the first end and different sequences near the second end. Step (v) depicts the addition of a second adapter to the second end by branched ligation. In some embodiments, these second adapters contain positional barcode sequences
[0312] that are unique to each aliquot. As a result, single-stranded nucleic acid constructs
[0305] formed as a result of branched ligation in different aliquots contain different positional barcode sequences. Single-stranded nucleic acid constructs in the same aliquot share the same positional barcode sequence. The aliquots are then combined into a single mixture
[0306] . The dashed oval represents a single mixture. All subsequent steps are performed within the single mixture. Step (vi) depicts the denaturation of the product in the single mixture to form adapted fragments
[0307] .
[0166] In some other approaches, generating a nested set of single-stranded DNA constructs involves ligating adapters via branched ligation with positional barcode sequences after each specific period. One exemplary approach is shown in Figure 3C. Unlike the approach in Figure 3B, the reaction in Figure 3C is performed in a single tube throughout the entire procedure, and no aliquots are required. Step iii depicts the primer being extended for a first time, followed by ligation of a second adapter to the extended primer, resulting in a reaction mixture of fragments ligated with the second adapter
[0314] ; step iv depicts the primer being extended for an additional time, followed by ligation of an additional adapter to the extended primer (in the same reaction mixture), resulting in a mixture of fragments ligated with the second adapter
[0314] and the first additional adapter
[0315] ; and step v depicts the primer being extended for yet another time, followed by addition of a second adapter, resulting in a mixture of fragments ligated with the second adapter
[0314] , the first additional adapter
[0315] , the second additional adapter
[0316] , etc. The process of adding adapters with unique positional barcodes can be repeated 3 to 50 times, for example, 10 to 40 times, or 10 to 20 times. The second adapter, the first additional adapter, the second additional adapter, and the further additional adapter each contain a unique position barcode sequence, and ligation of each of these adapters to the extended primer is via branched ligation.
[0167] The molar amount of each adapter used in branch ligation (as shown in steps (iii)-(v) of Figure 3C) is a small percentage of the total molar amount of the amplified genomic fragment, such that only a portion of the extended primers available for branch ligation in each round are ligated with adapters. In some embodiments, the molar amount of adapter used is 1-20%, e.g., 1-10%, or 2-15%, of the total molar amount of the amplified genomic fragment. In some embodiments, the amount of adapter (e.g., second adapter, first additional adapter, second additional adapter) used in different rounds is the same. In some embodiments, the amount of adapter used in different rounds is different. Step (vi) of Figure 3C shows denaturing the reaction mixture
[0316] to generate single-stranded, adapted fragments
[0317] . The single-stranded fragments are circularized as described below.
[0168] 5.3 Cyclization
[0169] The single-stranded nucleic acid constructs are then circularized to form single-stranded circles. Methods for circularizing single-stranded nucleic acids are well known and are described in Section 3.2. At least some of these single-stranded DNA fragments contain a barcode sequence and a target sequence portion. In each nested set, the target sequence portions share the same nucleotide sequence near their first end but have different nucleotide sequences near their second end. Upon circularization, the first and second adapter sequences of each single-stranded nucleic acid construct are joined, bringing the first and second ends of the target sequence portions into close proximity with each other, allowing sequence information near both ends to be identified with a single sequence read. An exemplary approach is shown in step (vi) of Figure 3A and step (vii) of Figure 3B.
[0170] 5.4 Generation of linear double-stranded adapted constructs from single-stranded circles
[0171] A variety of approaches can be used to generate linear double-stranded, adapted constructs (e.g., those generated using the methods of Figure 3A, Figure 3B, or Figure 3C) from single-stranded DNA circles. These approaches include randomly fragmenting the single-stranded circle to generate fragments and then generating double-stranded, adapted constructs from these fragments. The DNA circle can be fragmented using methods known in the art, such as sonication, to generate multiple single-stranded DNA fragments. Each single-stranded DNA fragment contains the barcode sequence that was present in the DNA circle. Optionally, size selection is performed to select fragments with lengths suitable for sequencing.
[0172] An example of such an approach is shown in Figure 4A. A single-stranded DNA fragment is used as a template to synthesize a complementary strand, forming multiple double-stranded fragments
[0402] . Step (viii) shows ligating adapters to both ends of the double-stranded fragments to generate adapted double-stranded constructs
[0402] . Step (ix) shows size selection to generate nucleic acid constructs
[0403] with lengths suitable for sequencing.
[0173] In some other approaches, a primer hybridized to the circle is extended under controlled extension conditions to generate a linear, adaptor-modified, double-stranded construct, generating an extended primer of a length suitable for sequencing. An illustrative example is shown in steps (vii) and (viii) of Figure 4B. Typically, controlled extension does not copy the entire template sequence; rather, the extended primer remains hybridized to the circle
[0404] and has an exposed 3-prime end (i.e., a 3-prime recessed end), ready for branch ligation. In some cases, the extended primer has a length in the range of 300-1000 bases, e.g., 300-500 bases or 400-600 bases, to achieve more efficient sequencing. Short artifact products can be removed by exonuclease treatment or purification. Therefore, size selection is not required with this approach, and all extension products can be used to generate sequence reads.
[0174] Next, a second adapter is ligated to the recessed 3-prime end of the extended primer via branched ligation, forming adapted extended primers, each having a second adapter sequence at one end and a primer binding sequence and barcode sequence at the other end ( Figure 4B , step (ix)
[0405] ). The adapted extended primers
[0406] are recovered ( Figure 4B , step (x)), and primer extension is performed using the adapted extended primers as templates to generate complementary strands, forming double-stranded DNA fragments
[0407] ( Figure 4B , step (xi)). The double-stranded DNA fragments can be amplified and sequenced.
[0175] A nested set of linear double-stranded fragments for each genomic fragment to be sequenced can be generated using the DNA circle-based scheme described above. Each double-stranded DNA fragment in the nested set contains a different target sequence portion of the genomic fragment, and by assembling these different target sequence portions together, the sequence of the original long DNA molecule can be deciphered. See the "Assembly of Sequence Information" section above.
[0176] 6. Linear DNA-based scheme
[0177] In some approaches, sequence libraries containing double-stranded, adapted constructs containing target sequences are generated using a linear DNA-based approach, meaning that no DNA circles are generated during the process.
[0178] 6.1 Addition of adapters to both ends of genomic fragments and amplified genomic fragments
[0179] First, adapters are added to both ends of a genomic fragment, as shown in Figure 2 and also described in Section 5.1 above. In some approaches, the genomic fragment is amplified similarly to Section 5.1 above, except that amplification is performed with a polymerase in a reaction mixture containing uracil or with primers containing uracil, thereby generating amplified nucleic acid fragments with uracil incorporated into the reaction mixture. An example is shown in step (i) of Figure 5, where uracil is part of the amplification primer (not shown) and is incorporated into the amplified genomic fragment during amplification.
[0180] 6.2 Nick creation
[0181] Next, nicks are introduced into the amplified genomic fragments. In some approaches, amplification is performed in the presence of uracil, as described above, and nicks can be introduced into the amplified genomic fragments containing uracil by contacting them with uracil-DNA glycosylase. Uracil glycosylase can remove uracil and form abasic sites. An enzyme (e.g., APE1 or EndoIV) is also added to the reaction to remove the sugar group from the abasic sites. This treatment of the uracil-containing genomic fragments with the enzymes described above results in nicks in the extension products of the regions containing uracil bases, with each nick flanked by a 5-prime and a 3-prime exposed end.
[0182] Preferably, uracil is spiked into the amplification reaction after the extension of the amplification primer passes through the barcode region but before it reaches approximately the desired read length (also called the length suitable for sequencing). The length suitable for sequencing depends on the read length determined by the sequencing method, but can range from 25 to 1000 bases. In some approaches, this is achieved by spiking uracil into the reaction mixture after extension has already begun, i.e., after all other components required for amplification have already been added to the reaction mixture. In some approaches, uracil is spiked into the reaction mixture approximately 10 seconds to 10 minutes after extension has begun.
[0183] In other approaches, the primers used to amplify the genomic fragments contain uracil, which is incorporated into the amplified genomic fragments. In some embodiments, the forward primers contain one or more uracils. In some embodiments, each forward primer contains a single uracil, such that one nick is generated in each of the double-stranded nucleic acid fragments (after enzymatic treatment to remove the uracil, as described above).
[0184] 6.3 Aliquots
[0185] The reaction mixture is then divided into multiple aliquots, see step (iii) of Figure 5.
[0186] 6.4 Nick Translation to Generate Nested Sets of Nucleic Acid Constructs
[0187] Next, nick translation is performed on the aliquots using a DNA polymerase with 5' to 3' exonuclease activity to synthesize a DNA strand with a newly formed end (second end). Non-limiting examples of DNA polymerases include DNA Pol1, Taq, full-length Bst, and Pfu DNA polymerase. The end opposite the second end is the first end. Extension is controlled so that the DNA strands synthesized in different aliquots have different lengths. Each synthesized DNA strand contains a first end and a second end, and the DNA strands of different aliquots share the same sequence near the first end and have different sequences near the second end.
[0503] Each synthesized DNA strand contains a target sequence portion with a first end and a second end, and the second end is the end formed by nick translation, and the first end is the end opposite the second end. The DNA strands of different aliquots share the same sequence near the first end and have different sequences near the second end. An illustrative example is shown in step (iv) of FIG.
[0188] 6.5 Branch ligation in individual aliquots
[0189] Adapters (second adapters) are added to the aliquot after the completion of the nick translation reaction. These second adapters are ligated to the second end of the newly synthesized DNA strand. Each second adapter is partially double-stranded and contains a first adapter oligonucleotide and a second adapter oligonucleotide. The first and second adapter oligonucleotides are complementary and hybridize to each other. During branch ligation, the 5-prime end of the first adapter oligonucleotide is joined to the 3-prime end of the DNA strand synthesized via nick translation as described above (e.g.,
[0504] in Figure 5).
[0190] In some approaches, the second adapter contains a positional barcode that is unique to the aliquot. In some approaches, aliquots containing unique positional barcodes are combined into a single reaction mixture (e.g.,
[0505] in Figure 5).
[0191] In some approaches, the second adaptor further comprises an anchor component for separating fragments ligated to the second adaptor from fragments not ligated to the second adaptor. In some approaches, the anchor component allows the adapted fragments to be captured by a solid support, allowing the captured adapted fragments to be separated from other reagents in solution. In some approaches, the anchor component can be biotin, and the solid support is coated with streptavidin. In some approaches, the anchor component is an oligonucleotide of the second adaptor, and the solid support is a magnetic bead with the oligonucleotide immobilized thereon.
[0192] Next, the synthesized DNA strands ligated with the second adaptor from the different aliquots from (v) are combined to form a single mixture. (For example,
[0505] . Figure 5, step (vi), the dashed oval represents a single mixture.) All subsequent steps are performed on the single mixture.
[0193] 6.6 Extending the second adapter oligonucleotide to generate a double-stranded fragment
[0194] By branched ligation, a first adaptor oligonucleotide is attached to the nucleic acid construct, and a second adaptor oligonucleotide is not attached but remains hybridized to the attached first adaptor oligonucleotide. A primer is then hybridized to the first adaptor oligonucleotide, and the hybridized primer is extended to generate double-stranded fragments. In some approaches, the double-stranded fragments thus generated have blunt ends. In some approaches, the double-stranded fragments thus generated contain location barcodes unique to each aliquot. An illustrative example is shown in step (vii) of Figure 5.
[0195] 6.7 Generation of linear double-stranded adapted constructs
[0196] In some embodiments, double-stranded DNA molecules having a length suitable for sequencing are selected. In some approaches, double-stranded fragments having a length in the range of 200 bp to 1.5 kb, e.g., 500 to 1000 bp, are selected. In some approaches, the selected double-stranded fragments are ligated to an adapter ("third adapter"), e.g., by blunt-end ligation, thereby generating a double-stranded adapted construct. See step (viii) of Figure 5. The double-stranded adapted construct can then be sequenced as disclosed herein.
[0197] The sequences of the double-stranded fragments of each aliquot near their positional barcodes can be determined by sequencing, and the sequence reads corresponding to different target sequence portions of each nucleic acid construct are assembled to generate sequence information for the entire target sequence.
[0198] 7. Loop-mediated complete stLFR
[0199] Also provided is a loop-mediated complete stLFR method for preparing libraries for sequencing long DNA molecules. In one embodiment of this method, loop-mediated complete stLFR involves preparing multiple nested sets of single-stranded nucleic acid adaptor constructs using any of the methods disclosed herein. Each single-stranded nucleic acid construct in each nested set comprises a target sequence portion of a long DNA molecule flanked by a first adaptor sequence at the 5' end and a second adaptor sequence at the 3' end (see, e.g., Figure 13B). The first adaptor sequence comprises, from 5' to 3', a primer binding sequence (e.g., 1311 in Figure 13A), a barcode sequence (e.g., 1319 in Figure 13A), and a first hybridization sequence (e.g., 1432 in Figure 13A). The second adaptor sequence comprises a second hybridization sequence. The first and second hybridization sequences are complementary to each other. Each target sequence portion has a first end and a second end, and the distance between the first end and the barcode sequence is shorter than the distance between the second end and the barcode sequence. The single-stranded nucleic acid constructs of each nested set share the same barcode sequence, and the single-stranded nucleic acid constructs of different nested sets have different barcode sequences. For each nested set of single-stranded nucleic acid constructs, the target sequence portions of that nested set have identical nucleotide sequences near the first end and differ from each other by truncation near the second end, and each nested set of single-stranded nucleic acid constructs includes multiple target sequence portions of different lengths. As further described below, various schemes can be used to generate nested sets of target sequence fragments.
[0200] In some embodiments, the method further includes subjecting multiple nested sets of single-stranded nucleic acid constructs to hybridization conditions in a reaction, whereby a first adapter sequence hybridizes to a second adapter sequence, thereby forming a loop (e.g., 1431 in Figure 14). The method further includes extending the second adapter sequence using a DNA polymerase to copy the barcode sequence and primer binding sequence of the first adapter sequence. The method further includes denaturing the reaction, whereby the loop opens and linear single-stranded DNA constructs are formed. Each linear single-stranded DNA construct comprises a barcode sequence and a primer binding sequence at its 3' end. In some embodiments, the method further includes annealing a primer to the primer binding sequence 3' of the linear single-stranded DNA construct and extending the primer to generate an extension product having a length suitable for sequencing. Details of the complete loop-mediated stLFR method are described further below.
[0201] An exemplary embodiment of loop-mediated complete stLFR involves ligating two partially double-stranded blunt-end adapters (each comprising a first adapter sequence and a third adapter sequence) to an end-repaired DNA fragment having a 5'-phosphate group to prepare an adapted double-stranded genomic fragment, each comprising a target sequence flanked by the first adapter sequence and the third adapter sequence. In some embodiments, the third adapter sequence is added to the DNA fragment by nick translation, as described in Figure 13A and Example 4.
[0202] In some embodiments, double-adapted double-stranded genomic fragments are amplified. Random nicking and gapping are performed on the double-adapted double-stranded genomic fragments. Random nicking generated a nested set of fragments with target sequence portions of different lengths and sharing a common barcode sequence at 5-prime. One such fragment is shown in Figure 13B as 1341.
[0203] When gaps are opened using exonucleases as described above, protection of the DNA adapter can be achieved, if necessary, via phosphorothioate bonds between bases at the 5' and 3' ends of the adapter and / or between modified bases.
[0204] A second adaptor (e.g., AD153UMI_5R shown in Figure 13B) is then ligated to the 3' side of the nick or ssDNA gap of the adaptor-ligated DNA fragment via bifurcated ligation. This results in a set of fragments with different lengths of target sequence portions, each flanked by two adaptor sequences (e.g., AD153UMI_5 and AD153UMI_5R in Figure 13B). One such fragment is shown as 1342 in Figure 13B. The second adaptor is a partially double-stranded DNA adaptor molecule comprising a long strand (1320.1 in Figure 13B) and a short strand (1320.2 in Figure 13B). The long strand has a 5'-phosphate. The long strand further comprises a first adaptor sequence comprising a first hybridization sequence (e.g., 1432 in Figure 14) located 3' to the barcode sequence (e.g., 1319 in Figure 14). The longer strand of the second adapter comprises a second adapter sequence that comprises a second hybridization sequence (e.g., 1433 in Figure 14). The first and second hybridization sequences are complementary and can hybridize to each other under conditions suitable for hybridization, as further disclosed below.
[0205] Ligation of the first and / or second adaptors to DNA fragments via branched ligation can be performed in solution or on beads. When performed on beads, the adapted DNA fragments are preloaded onto the beads with a high concentration of PEG (5%-15%) before adding other reaction components. In some embodiments, the branched ligation reaction is performed in the presence of additives (e.g., polyethylene glycol or betaine) to enhance the activity of the ligation and / or nicking enzymes. The reaction can be incubated at pH 5.0-9.0, at room temperature, 37°C, or cycled between various temperatures, such as 5-15°C and 37°C. Incubation can last from 5 minutes to several hours. The incubation time and nickase activity depend on the desired number of nicks per DNA fragment. The reaction can be terminated using DNA purification methods (e.g., Ampure XP beads) if performed in solution, or by a simple wash step with Tris-NaCl buffer containing PEG (5%-15%) if performed on beads.
[0206] In some embodiments, after branch ligation, the DNA fragments are denatured to generate single-stranded DNA molecules containing a target sequence flanked by adapter sequences, each of which has a single-stranded hybridization sequence (e.g., the molecule shown in the lower left of Figure 13B). In some embodiments, after branch ligation, the branched-ligated DNA fragments can be heat denatured (90°C to 95°C). Alternatively, the branched-ligated DNA fragments can be denatured with an alkaline agent (e.g., 0.05M to 0.2M NaOH or KOH) and further neutralized with a neutralizing agent (e.g., HCl, Tris-HCl, MOPS). A single-stranded DNA molecule (e.g., 1343 in Figure 13B) containing a target sequence flanked by adapter sequences is formed.
[0207] Alternatively, in some embodiments, instead of denaturing, the branched-ligated DNA fragments are digested with one or more dsDNA-specific exonucleases with 3'-5' exonuclease activity (e.g., exonuclease III) to expose the 5' single-stranded first hybridization sequence of the first adapter (e.g., 1432 in Figure 13 or Figure 14) available for hybridization with the second hybridization sequence of the second adapter at the 3' end of the DNA fragment.
[0208] Hybridization between the first and second hybridization sequences can be carried out in a hybridization buffer containing a buffer (e.g., Tris-HCl, MOPS, sodium phosphate) and / or salt. In some embodiments, the hybridization buffer also contains cofactors such as MgCl2 and dNTPs for the subsequent enzymatic reaction.
[0209] Following the DNA hybridization step, the barcode and primer binding sequence on the first adapter (e.g., AD153 UMI_5 shown in Figure 14) are copied by extending the hybridized 3' end of the branched adapter (e.g., AD153UMI_5R shown in Figure 14). In some embodiments, to increase specificity, extension is performed using one or more DNA polymerases lacking 3'-exonuclease activity. Exemplary DNA polymerases include, but are not limited to, Taq DNA polymerase, Klenow fragment (3' to 5' exonuclease), and Bst DNA polymerase, large fragment. In some embodiments, extension is performed at a temperature suitable for the polymerase to perform the polymerization reaction. In some embodiments, the temperature ranges from 30°C to 75°C.
[0210] The linear extension product (e.g., 1431 in Figure 14) is in the form of a looped double-stranded or partially double-stranded DNA molecule. The double-stranded or partially double-stranded DNA molecule contains a double-stranded adapter containing a barcode sequence, which is attached to the target sequence portion of the long DNA molecule. The double-stranded or partially double-stranded DNA molecule is then denatured to open the loop and form a single-stranded fragment (e.g., 1441 in Figure 14) containing the target sequence portion flanked by the first and second adapter sequences. By copying the barcode, the ends of target sequence portions of different lengths, such as 1411, 1421, and 1431, are brought near the barcode sequence 1319 on the same DNA strand, allowing these ends (with different target sequences) to be sequenced and assembled based on the common barcode sequence 1319. The adapter sequence at the 3-prime of the target sequence portion contains a copy of the barcode sequence and a copy of the primer binding sequence. In some embodiments, the primer binding site is recognized by a universal amplification primer. A primer can be annealed to the primer binding sequence and extended to generate an extension product of a length suitable for sequencing. The extension product is ligated to a fourth adapter (e.g., Ad153_3 in Figure 15) via branched ligation, and the fragment thus generated (e.g., 1510 in Figure 15) is amplified by PCR and circularized for DNB sequencing.
[0211] Other exemplary schemes for generating nested sets of target sequence fragments are disclosed in Section 8, "Concatemer-Based Methods," and Section 9, "Combinatorial Schemes." The first scheme starts with a circle of ssDNA or dsDNA, while the second scheme starts with linear ssDNA or dsDNA. Both schemes involve the addition of adapter sequences to both ends of the molecule of interest, which can be done via adapter ligation, amplicon PCR, or many other strategies. One of these adapters carries a barcode sequence that can later be used to identify all reads originating from this particular DNA fragment / molecule.
[0212] 8. Concatamer-based methods
[0213] 8.1.1 Concatemer
[0214] In some embodiments, provided herein are concatemer-based methods that generate nested sets of adapted fragments with target sequence fragments of different lengths. In some embodiments, concatemers are generated by rolling circle replication of single-stranded circular templates. Single-stranded circular templates can be generated by circularizing single-stranded DNA molecules using methods well known in the art. For example, circularization can be achieved by using splint oligos with sequences complementary to adapter sequences at both ends of the molecule, and the 5' and 3' ends can be ligated together. DNA concatemers can then be generated by extending primers annealed to the circular template sequences using a DNA polymerase with strand displacement activity.
[0215] The circular DNA template disclosed herein contains a barcode, a primer sequence, and a target sequence. Once the circle is created, the next step is to form DNA concatemers, e.g., DNA nanoballs (DNBs). The incubation time for creating concatemers can range from 20 minutes to several hours. Longer concatemer creation times can result in very long concatemers (over 100 kb) that can be cleaved into separate concatemers. Because each circle contains a unique barcode, this cleavage is not a problem, as all reads from these separate concatemers can still be properly identified using the barcode information. Each concatemer contains multiple identical monomers, each of which contains a complement of the target sequence, a complement of the barcode sequence that identifies the DNB, and a primer binding sequence. The primer binding sequence contains a sequence complementary to the primer sequence. In some embodiments, the primer binding sequence is shared by a population of single-stranded concatemers.
[0216] 8.1.2. Generation of Multiple Extended Primers Separated by Spacing
[0217] After concatemers are formed, they are converted to dsDNA by extending primers annealed to the primer binding sequences of the concatemers. In some embodiments, primers are used at a concentration sufficient to ensure that almost all primer binding sites on the DNA fragments are occupied by extended primers. Extension can be performed using a polymerase that lacks 5-3' exonuclease activity and does not have strand displacement activity. This results in the formation of a DNA complex containing multiple extended primers complementary to one or more monomers of the DNA concatemers. These extended primers in the DNA complex are hybridized to the DNA concatemers and separated by intervals. See, for example, Figure 6A. Each extended primer contains a target sequence fragment. In some embodiments, the primers are DNA primers. In some embodiments, the primers are RNA primers. In some embodiments, the primers are a mixture of RNA and DNA primers.
[0218] DNA polymerases lacking 5-3' exonuclease activity and lacking strand displacement activity are known, and non-limiting examples include Klenow Exo, Q5, Hemocline Taq, T7 polymerase, and T4 polymerase. Readily available commercial products are available, for example, from New England BioLabs (Ipswich, Massachusetts).
[0219] 8.1.3. Generating Gaps for Adapter Addition (in Concatamer-Based Methods)
[0220] In some embodiments, the spacing between fragments of such a DNA complex is extended (widened) by an exonuclease to form a gap. See, for example, Figure 6B. This process is sometimes referred to as "gapping," and the exonuclease used in this process is sometimes referred to as a "gapping enzyme." Examples of enzymes with 3' exonuclease activity include DNA polymerase I, Klenow fragment (in the absence of nucleotides), exonuclease III, and others known in the art. Examples of enzymes with 5' exonuclease activity include Bst DNA polymerase, T7 exonuclease, exonuclease VIII truncate, lambda exonuclease, T5 exonuclease, and other exonucleases known in the art. Low processivity exonucleases (i.e., exonucleases that remove nucleotides from the ends of polynucleotides at a relatively slow rate) are preferred for opening short gaps (e.g., 2-7 bases, 3-10 bases, 3-20 bases) to allow for adapter ligation. Exemplary exonucleases that can be used are listed in Table 1.
[0221] [Table 1]
[0222] In the scenario where fragments are generated by extending an RNA primer as described above, RNase H can be added to degrade the RNA primer, extending the gap and forming a gap. The gap typically has the length of the RNA primer (e.g., 8-40 bases, 10-35 bases, or 10-25 bases). The 5' and 3' ends of the gap can be ligated with an L adaptor and a 3' branched ligation adaptor, respectively.
[0223] Figures 6A and 6B illustrate the process of using one or more gapping enzymes (e.g., exonucleases with 3' to 5' exonuclease activity) to widen spacing (160), resulting in gap (170). Figures 7A and 7B illustrate the process of using one or more gapping enzymes (e.g., exonucleases with 5' to 3' exonuclease activity) to widen spacing (260), resulting in gap (270).
[0224] To maintain the integrity of the barcode sequence, the exonuclease used should only digest the end far from the barcode sequence. For example, if the barcode sequence is close to the 5-prime end of the target sequence fragment (as in Figure 6B), an exonuclease with 3' to 5' exonuclease activity should be used to initiate digestion of the fragment from the 3-prime end. On the other hand, if the barcode sequence is close to the 3-prime end of the target sequence fragment (as in Figure 7B), an exonuclease with 5' to 3' exonuclease activity should be used to initiate digestion of the fragment from the 5-prime end.
[0225] If the primer binding sequence is located 3-prime to the complement of each monomer's barcode sequence (i.e., if the extension primer is located 5-prime to the barcode sequence of the extended primer), adding a 3'-5' exonuclease during the ligation step can create target sequence fragments with different sequences cleaved at the 3-prime end but identical sequences at the 5-prime end (Figure 6B). Importantly, this results in adaptored fragments that all contain the barcode and have full-length coverage of the original molecule, including the target sequence. Furthermore, L-oligo adaptors can be designed to recognize portions of the adaptor sequence on the concatemer, improving ligation efficiency.
[0226] If the primer is positioned 5-prime to the complement of the barcode sequence of each monomer (i.e., the configuration of the extension primer is positioned 3-prime to the barcode sequence of the extended primer), a 5'-3' exo can be used instead, which generates target sequence fragments with different sequences by cleavage at the 5-prime end and an identical sequence at the 3-prime end (Figure 7B).
[0227] If the primers used are DNA extension primers, a low concentration of exonuclease can be added before adding the ligase, adapter, and exonuclease. This allows most of the gap to open before the ligase has a chance to reseal the gap. If the primer binding sequence is located 3-prime to the complement of each monomer's barcode sequence (as in Figure 6A), an exonuclease with 3' to 5' exonuclease activity is used (Figure 6B). If the primers are located 5-prime to the complement of each monomer's barcode sequence (as in Figure 7A), an exonuclease with 5' to 3' exonuclease activity is used (Figure 7B).
[0228] 8.1.4 Simultaneous Ligation and Exonuclease Treatment
[0229] In some embodiments, exonucleases are used to generate target sequence fragments of different sizes. Due to the stochastic nature of exonucleases, exonuclease treatment results in a distribution of cleaved and extended primers of different sizes containing target sequence fragments. These target sequence fragments have identical nucleotide sequences at their first ends and differ from each other by cleavage at their second ends, resulting in sets of adapted fragments containing target sequence fragments of different lengths. In some embodiments, the first ends are 5-prime ends of the target sequence fragments, as shown in (191) of Figure 6B. In some embodiments, the first ends are 3-prime ends of the target sequence fragments, as illustrated in (291) of Figure 7B.
[0230] In some embodiments, exonuclease treatment of the extended primer and ligation of the extended primer to the adapter occur in the same reaction mixture. In some embodiments, ligation comprises ligating at least a branched adapter to the nucleic acid fragment. In some embodiments, ligation comprises ligating both a branched adapter and an L-adapter to the extended primer in solution. The following are exemplary conditions under which exonuclease treatment and ligation can occur in the same reaction:
[0231] temperature
[0232] The reaction can be maintained at a temperature ranging from 5 to 65°C, e.g., from 5 to 42°C, from 10 to 37°C, or from 5 to 15°C. In some embodiments, the reaction is maintained at room temperature, 37°C. In some embodiments, when using a thermostable ligase and exonuclease, the reaction can be maintained at a temperature greater than 37°C.
[0233] pH
[0234] In some embodiments, the pH of the reaction mixture is maintained within the range of 5.0 to 9.0, e.g., 7.0 to 9.0, to accommodate all enzymatic functions required for library preparation. The duration of the exonuclease treatment and ligation reaction can vary depending on the desired size of the nucleic acid fragments and other conditions, such as the concentration of enzymes (including polymerase, exonuclease, or both), time, temperature, and amount of input DNA.
[0235] time
[0236] Typically, the duration of the nicking and exonuclease treatment reaction can be 5 minutes to 5 hours, e.g., 15 to 90 minutes, or 30 to 120 minutes. The reaction can be terminated using methods well known in the art. In some embodiments, exonuclease treatment and ligation are performed in solution, and the reaction can be terminated via DNA purification methods (e.g., Ampure XP beads from Beckman Coulter). In some embodiments, exonuclease treatment and ligation are performed on beads, and the reaction can be terminated by washing the beads with a buffer (e.g., Tris-NaCl buffer) to remove enzymes and components necessary for the nicking and ligation reactions.
[0237] 8.1.5. Chaining Adapters
[0238] As described above and shown in FIG. 6B, extended primers (130) annealed to the adapter sequences of the DNA concatemers generate a plurality of extended primers (150) having respective 5' and 3' ends. In some embodiments, at least some of the extended primers are each ligated with two adapters, one at either end. In some embodiments, an L-adapter is ligated to the 5' end and a branched adapter is ligated to the 3' end of the extended primer. This results in a plurality of adapted fragments with two different adapter sequences, and all adapted fragments generated in the reaction have the same defined configuration (e.g., an L-adapter at the 5' end and a branched adapter at the 3' end).
[0239] 9. Combination Schemes
[0240] In some embodiments, the methods disclosed herein can be combined with the stLFR method to sequence long genomic DNA, e.g., genomic fragments with lengths of 20 kb to 200 kb. In this case, stLFR is performed, e.g., as disclosed in U.S. Patent No. 9,328,382 B2 and PCT Publication No. WO 2023001262, followed by the above-described method to adjust the size of the barcoded fragments to approximately 300 bp to 1,000 bp, 500 bp to 1,500 bp, or 1 kb to 3 kb. The advantage of using long inserts is their resistance to enzyme bias (e.g., transposases and DNA nicking enzymes), which allows for co-barcoding of stLFR. It is also easy to remove stLFR adapter-adapter artifacts by size selection (i.e., removing DNA less than 300 bp in length). The barcodes provided by the beads will be the barcodes used for each circle. After performing the above process, each 2 kb fragment will have greater than 1x read coverage and can be combined with other fragments sharing the same stLFR barcode to create longer fragments (possibly up to several megabytes) with close to 1x or greater coverage across the entire fragment (average read coverage 5-15x). An exemplary method according to this embodiment is shown in Figure 8.
[0241] Overall, with over 10x coverage in long fragments per haplotype and approximately 5-10x read coverage in each long fragment, a total read coverage of 50-100x per haplotype would be required. Such complete and accurate WGS will become more affordable as the cost of MPS (NGS) continues to decrease.
[0242] In certain embodiments, this method begins with any linear DNA molecule that has an adapter at at least one end. In some embodiments, this method begins with a PCR amplicon to provide sufficient copies of each barcode molecule. In some embodiments, this process can be performed in solution. In some embodiments, this method can be performed on beads onto which one end (5-prime or 3-prime) of the adapter of the linear DNA molecule is immobilized (Figure 9), and the adapter contains a unique barcode. In some embodiments, the barcoded fragments comprise a barcode sequence, a target sequence, and a primer binding sequence, and the 3-prime end of the barcoded fragment is immobilized on the bead. In some embodiments, the barcoded fragments comprise a barcode sequence, a target sequence, and a primer binding sequence, and the 5-prime end of the barcoded fragment is immobilized on the bead. In some embodiments, the primer binding sequence is 3-prime relative to the barcode sequence.
[0243] Polynucleotides (e.g., barcoded fragments) can be immobilized on beads by various methods, including covalent and non-covalent attachment. In some embodiments, the 3' or 5' end of the polynucleotide's adapter is attached to biotin, and the barcoded fragments are captured on streptavidin-coated beads. In some embodiments, the polynucleotide is conjugated to a substrate (e.g., a bead), i.e., one end of the polynucleotide is directly contacted or linked to the substrate. For example, the surface may have reactive functional groups that react with complementary functional groups on the polynucleotide molecule to form covalent bonds. Long DNA molecules, e.g., those several nucleotides or larger, can also be efficiently attached to hydrophobic surfaces, e.g., clean glass surfaces with low concentrations of various reactive functional groups, such as -OH groups. In yet another embodiment, polynucleotide molecules can be adsorbed to the surface through non-specific interactions with the surface or through non-covalent interactions, such as hydrogen bonding or van der Waals forces.
[0244] In some embodiments, polynucleotides (e.g., barcoded fragments) are immobilized to a surface by hybridizing to a capture oligonucleotide on the surface and forming a complex, e.g., a double-stranded duplex or a partially double-stranded duplex, with a component of the capture oligonucleotide.
[0245] In some embodiments, the method uses a primer comprising a primer sequence complementary to a primer sequence of the barcode fragment. In some embodiments, the primer is a tailed primer and comprises a tail that is not complementary to the barcoded fragment. In some embodiments, the tail comprises a common adapter sequence.
[0246] In some embodiments, extension is controlled so that the polymerase extends the primer past the barcode region on the barcoded fragment. In some embodiments, the polymerase extends the primer past the barcode region to a length approximately equal to the length suitable for sequencing (also known as the sequencing read length), e.g., in the range of 25-1000 bases. Figure 9. The extension product, i.e., the extended primer, can be separated from the barcoded fragment, leaving the barcoded fragment as a template for a subsequent extension reaction cycle. Various methods of controlling the extension reaction can be used, and are further described in Section 3.4 ("Controlled Extension") below. In some embodiments, primer extension (one or more cycles) can be controlled so that the next extension cycle generates extension products that are longer or shorter than the products of the previous extension cycle. In some embodiments, controlled extension is performed in the presence of a reversible terminator. In some embodiments, controlled extension is performed by using a different polymerase capable of longer or shorter extensions. One advantage of using terminators is that the length of the additional polymerization can be controlled by the concentration of the terminator, which is easier to control than the time. At the end of the polymerization, the terminators are replaced and the beads are washed. In some embodiments, the beads can be washed with a buffered NaCl solution. This is followed by 3' branch ligation of a branched adapter (e.g., 440 in Figure 9). After denaturing the DNA under either heat or alkaline conditions, the supernatant can be collected and purified.
[0247] For example, this process, involving primer extension and denaturation, can be repeated many times to generate nested sets of adapted fragments containing target sequence fragments of various lengths, such that the entire original DNA molecule is covered. These adapted fragments also share the same barcode sequence. This method yields variable target sequence fragments with sizes ranging from approximately 100 bp to 5000 bp, 100 bp to 3000 bp, 100 bp to 1000 bp, 100 bp to 750 bp, or 100 bp to 500 bp.
[0248] The target sequence fragments generated above can be circularized for DNA nanoball (DNB) preparation and sequencing. As described above, these target sequence fragments have identical nucleotide sequences at their first ends (the ends closer to the barcode sequence) and differ from each other by truncation at their second ends. In some embodiments, the sequencing is paired-end sequencing, which involves sequencing from either end of the same DNA fragment. In some embodiments, a first read is generated by extending a sequencing primer annealed to an adapter sequence closer to the first end of the target sequence fragment than to the second end ("first read sequencing"), and a second sequencing read is generated by extending a sequencing primer annealed to an adapter sequence closer to the second end of the target sequence fragment than to the first end ("second read sequencing"). The first read sequencing generates a barcode sequence. The second read sequencing generates overlapping reads that substantially or completely cover molecules up to 500 bp, 700 bp, or 1000 bp in length. These overlapping sequencing reads are clustered in a de novo assembly based on the barcode sequences determined by the first read sequencing.
[0249] In some embodiments, uracil is incorporated midway through the extension step to more efficiently cover the entire molecule but avoid generating fragments with excessive length (e.g., greater than 700 bp, or greater than 1000 bp, or greater than 1500 bp, or greater than 2000 bp, or greater than 3000 bp). For example, uracil can be added to the reaction after extension has passed the barcode region but before reaching approximately the desired read length (e.g., 25-1000 bases, depending on the read length determined by the sequencing method). This can be achieved by extending the primers through a first extension period, then spiking the reaction mixture with uracil to allow the primers to continue extending through a second extension period, washing the beads, and adding a new uracil-free deoxynucleotide mix (normal deoxynucleotides) to the reaction to allow the primers to further extend through a third extension period (Figures 10A-10B). In some embodiments, this third extension reaction (a uracil-free extension) is carried out in the presence of a reversible terminator, for example, using a mixture of regular nucleotides and reversible terminators.
[0250] After the extension reaction is complete, terminators, if used, are replaced, the beads are washed, and the extension products are ligated with 3' branched ligation adapters. Next, uracil glycosylase, which removes uracil to form abasic sites, can be added to the reaction. An enzyme capable of removing sugar groups from abasic sites can be added to the reaction. This results in fragmentation of the extension products at the region containing the uracil base. Non-limiting examples of enzymes capable of removing sugar groups from abasic sites include APE1 and EndoIV. Removal of these fragmented products leaves gaps flanked by exposed 5-prime and 3-prime ends. L-adapters and internal branched adapters are ligated to the exposed 5-prime and 3-prime ends. In some embodiments, the L-adapters, internal branched adapters, and adapter sequences at the 5' and 3' ends of the extended fragments are all distinguishable from one another. The 5' and 3' ends of these two adapters can then be joined via a splint oligonucleotide and ligated by T4 ligase to rejoin the 5' and 3' ends of the extended fragment (the single-stranded portion of the template molecule folds the two adapters into close proximity and hybridizes with the splint oligonucleotide, Figure 5B).
[0251] The products are then denatured, separated from the beads, and recovered. In some cases, the beads can be reused for one or more cycles (see Figure 10C). The denatured products from all rounds are collected and sequenced. This procedure reduces the overall length of the target sequence fragment for each adapted fragment by removing the middle portion of the extension product (see Figure 10C). Figure 11 shows truncated target sequence fragments with sequences corresponding to different regions of the target sequence of the original molecule. This allows molecules up to 1500 bp, 2000 bp, 3000 bp, 4000 bp, or 5000 bp to be read (covered) by MPS. To achieve 10x coverage at 300 bases per read, a 3 kb molecule requires 100 reads. Due to losses during library preparation and DNA recombination (DNB) or clustering, it is preferable to start the library process with at least 300, 1000, or 10,000 copies of each DNA molecule. Reusing beads reduces the copy number by 2-10x or more. If longer sequencing reads (e.g., 600 bases) are available and shorter template molecules (e.g., 1500 bp) are reused 25-100 times, it is possible to perform this process without first amplifying the molecules by linear or exponential PCR or other methods. However, the amplified molecules allow for multiple (2 or more, 3 or more, 4 or more, 2-6, or 4-8) reactions in parallel with longer uracil regions to optimally cover the entire region of a DNA molecule with a length ranging from 1 kb to 10 kb, e.g., 1 kb to 5 kb, or 1 kb to 3 kb.
[0252] Another exemplary solution to the problem (i.e., the original molecule containing the target sequence is too long for sequencing) is to use a first branched adapter with a degenerate sequence region in its 3-prime portion. This first branched adapter is ligated to the extended tailed primer formed after the first extension, as described above. The first extension is controlled to extend the primer past the barcode region. In some embodiments, the degenerate sequence region contains 3 to 10, e.g., 3 to 8, 5 to 10, or 6 to 10 degenerate nucleotides. The first branched adapter can hybridize to a random position in the barcoded fragment through the degenerate sequence region, thereby skipping replication of a random portion of the barcoded fragment. Next, a second controlled extension is performed by extending the 3-prime end of the first branched adapter. The second extension can be performed to add 100 to 300 bases to the 3-prime end to form a second extension product. A second branched adapter can then be ligated to the 3-prime end of the second extension product to generate an adapted fragment. See Figure 12.
[0253] In some embodiments, the adapted fragments are denatured and released from the beads. The barcoded fragments can be used as extension templates for additional extension cycles to generate more adapted fragments. The first and second extensions in each cycle are controlled so that adapted fragments are generated from cycles with overlapping target sequence fragments. These adapted fragments can be sequenced, and the sequencing reads of the overlapping target sequence fragments can be assembled to generate sequence information for the entire target sequence.
[0254] In some embodiments, the barcoded fragments are amplified such that multiple copies of the barcoded fragments are used as templates for extension (e.g., to extend primers annealed to the barcoded fragments). In some embodiments, these multiple copies are immobilized on the same bead. In some embodiments, these multiple copies are immobilized on one or more beads. These copies are distinguishable by the same barcode they share. In this embodiment, one cycle (comprising a first extension, ligation with a first branched adapter, a second extension, and ligation with a second branched adapter) is often sufficient to generate overlapping target sequence fragments. However, if necessary, the extension products can be denatured and released from the beads, and the barcoded fragments can be reused for additional cycles to generate additional adapted fragments as described above.
[0255] 10. Exemplary Embodiments of the Present Disclosure
[0256] Embodiment 1 is a method of generating a single-stranded, adapted construct for sequencing, comprising: Preparing a plurality of nested sets of single-stranded nucleic acid constructs, comprising: each single-stranded nucleic acid construct in each nested set comprises a target sequence portion flanked by a first adaptor sequence at the 5' end and a second adaptor sequence at the 3' end; the first adapter sequence comprises, from 5' to 3', a primer binding sequence, a barcode sequence, and a first hybridization sequence, and the second adapter sequence comprises a second hybridization sequence; the first and second hybridization sequences are complementary to each other; each target sequence portion having a first end and a second end; the distance between the first end and the barcode sequence is shorter than the distance between the second end and the barcode sequence; the single-stranded nucleic acid constructs in each nested set share the same barcode sequence, and the single-stranded nucleic acid constructs in different nested sets have different barcode sequences; preparing, for each nested set of single-stranded nucleic acid constructs, target sequence portions of the nested set that have identical nucleotide sequences near a first end and differ from one another by truncation near a second end, such that each nested set of single-stranded nucleic acid constructs contains multiple target sequence portions having different lengths.
[0257] Embodiment 2 is the method of Embodiment 1, further comprising subjecting the plurality of nested sets of single-stranded nucleic acid constructs to hybridization conditions, whereby a first adaptor sequence hybridizes to a second adaptor sequence, thereby forming a loop.
[0258] Embodiment 3 is the method of embodiment 2, further comprising extending the second adapter sequence with a DNA polymerase to copy the barcode sequence and primer binding sequence of the first adapter sequence to form an extension product.
[0259] Embodiment 4 is the method of embodiment 3, further comprising denaturing the extension products to open the loops, thereby forming linear single-stranded DNA constructs, each linear single-stranded DNA construct comprising a barcode sequence and a primer binding sequence, the primer binding sequence being located 3' to the barcode sequence. In some embodiments, the primer binding sequence at the 3' end of the linear single-stranded DNA construct.
[0260] Embodiment 5 is the method of embodiment 4, further comprising annealing a primer to a primer binding sequence 3' of the linear single-stranded DNA construct and extending the primer to generate an extension product having a length appropriate for sequencing.
[0261] Embodiment 6 is a method for generating single-stranded DNA circles comprising single-stranded, adapted constructs for sequencing, comprising: Preparing a plurality of nested sets of single-stranded nucleic acid constructs, comprising: each single-stranded nucleic acid construct of each nested set comprises a target sequence portion flanked by a first adaptor sequence and a second adaptor sequence; the first adapter sequence comprises a barcode sequence and a primer binding sequence; each target sequence portion having a first end and a second end; the distance between the first end and the barcode sequence is shorter than the distance between the second end and the barcode sequence; the single-stranded nucleic acid constructs of each nested set share the same barcode sequence, and the single-stranded nucleic acid constructs of different nested sets have different barcode sequences; For each nested set of single-stranded nucleic acid constructs: (a) the target sequence portions of a nested set of single-stranded nucleic acid constructs have identical nucleotide sequences near a first end and differ from one another by truncations near a second end, such that each nested set includes multiple target sequence portions having different lengths; (b) circularizing the single-stranded nucleic acid constructs of each nested set to generate single-stranded DNA circles to which the first adaptor sequence and the second adaptor sequence are joined; preparing a method of
[0262] Embodiment 7 is the method of any of Embodiments 1 to 6, wherein each nested set of single-stranded nucleic acid constructs comprises: (i) preparing an adapted double-stranded genomic fragment comprising a target sequence flanked by a first adaptor sequence and a third adaptor sequence; (ii) amplifying the adapted double-stranded genomic fragment using primers hybridized to the first and third adapter sequences to generate an amplified genomic fragment; (iii) contacting the amplified genomic fragment from step (ii) with a nicking agent to generate a nick at the target sequence on one strand of the amplified genomic fragment; (iv) ligating a second adapter comprising a second adapter sequence at the nick of step (iii) via branch ligation to form a ligated product; and (v) denaturing the ligated products from step (iv) to form single-stranded nucleic acid constructs each comprising the first adapter sequence and the second adapter sequence.
[0263] Embodiment 8 is the method of any of Embodiments 1 to 6, wherein each nested set of single-stranded nucleic acid constructs comprises: (i) preparing adapted double-stranded genomic fragments, each comprising a target sequence flanked by a first adaptor sequence and a third adaptor sequence; (ii) amplifying the adapted double-stranded genomic fragments to produce amplified genomic fragments; (iii) distributing the amplified genomic fragments into multiple aliquots; (iv) denaturing the amplified genomic fragments of step (iii) to prepare single-stranded genomic fragments, at least a portion of each of the single-stranded genomic fragments comprising a primer binding sequence; (iv) extending the primer hybridized to the primer binding sequence under extension control conditions such that the extension products from the different aliquots differ in length, thereby generating extension products with newly formed termini, the extension products having different sequences near the newly formed termini in the different aliquots; each extension product comprising a target sequence portion; (v) ligating a second adapter comprising a second adapter sequence at the newly formed end via branched ligation in each aliquot, thereby generating single-stranded nucleic acid constructs each comprising the first adapter sequence and the second adapter sequence; This is a method for preparing the compound.
[0264] Embodiment 9 is the method of any of Embodiments 1 to 6, wherein each nested set of single-stranded nucleic acid constructs comprises: (i) preparing adapted double-stranded genomic fragments, each comprising a target sequence flanked by a first adaptor sequence and a third adaptor sequence; (ii) amplifying the adapted double-stranded genomic fragments to produce amplified genomic fragments; (iii) distributing the amplified genomic fragments into multiple aliquots; (iv) adding a double-stranded DNA nuclease having 3' to 5' nuclease activity, varying the aliquots under controlled conditions such that the length of the product is maintained after double-stranded DNA nuclease digestion in the different aliquots, thereby generating digestion products with newly formed ends having different sequences in the different aliquots; each digestion product containing a portion of the target sequence; and (v) ligating a second adapter comprising a second adapter sequence at the newly formed end via branched ligation in each aliquot, thereby generating single-stranded nucleic acid constructs each comprising the first adapter sequence and the second adapter sequence; This is a method for preparing the compound.
[0265] 7. The method of any one of embodiments 1 to 6, wherein each nested set of single-stranded nucleic acid constructs comprises: (i) preparing adapted double-stranded genomic fragments, each comprising a target sequence flanked by a first adaptor sequence and a third adaptor sequence; (ii) amplifying the adapted double-stranded genomic fragments to produce amplified genomic fragments; (iii) denaturing the amplified genomic fragments to prepare single-stranded genomic fragments, at least a portion of each of the single-stranded genomic fragments comprising a primer binding sequence; (iv) for each single-stranded genomic fragment: extending the primer hybridized to the primer binding sequence for an initial period of time to produce an extended primer; the extension is incomplete such that the length of the extended primer is a fraction of the length of the single-stranded genomic fragment; the extended primer comprises a portion of the target sequence, and ligating a second adaptor to the end of the extended primer formed by the extension via branched ligation; thereby generating single-stranded nucleic acid constructs in a single reaction mixture, each comprising a first adapter sequence and a second adapter sequence; (v) repeating step (iv) multiple times, each time further extending the primer for an additional period of time and ligating additional adapters having unique positional barcodes to the further extended primers; Additional adapters are used in molar amounts that are fractions of the total molar amount of amplified genomic fragments; thereby generating a mixture of a nested set of single-stranded nucleic acid constructs; The method for preparing the compound according to the present invention is as follows.
[0266] Embodiment 11 is a method for detecting a target sequence comprising the steps of: the second adapter comprises a positional barcode sequence that is unique for each aliquot; 9. The method of embodiment 8, wherein the single-stranded nucleic acid constructs formed in step (v) in different aliquots comprise different positional barcode sequences, and the single-stranded nucleic acid constructs in the same aliquot share the same positional barcode sequence.
[0267] Embodiment 12 is the method of embodiment 6, wherein the primer binding sequence is 3-prime to the barcode sequence.
[0268] Embodiment 13 further comprises (vi) fragmenting the single-stranded DNA circle to generate a plurality of single-stranded DNA fragments, at least some of which comprise a barcode sequence; (vii) generating double-stranded DNA fragments from the single-stranded DNA fragments from step (vi); 7. The method of embodiment 6, further comprising the step of: (vii) ligating a second adaptor to each of the double-stranded DNA fragments from step (vii), thereby generating double-stranded, adapted fragments.
[0269] Embodiment 14 further comprises (viii) amplifying the double-stranded, adapted fragments, optionally 14. The method of embodiment 13, further comprising the step of (ix) selecting the amplified double-stranded adapted fragments having a length in the range of 300 to 1000 bases.
[0270] Embodiment 15 further comprises (vi) hybridizing a primer to a primer binding sequence in each of the single-stranded DNA circles; (vii) extending the primer under extension-controlled conditions using each of the single-stranded DNA circles as a template; the extension generates an extended primer hybridized to the single-stranded DNA circle, thereby generating a plurality of extended primers having different lengths; extending, wherein each of the extended primers comprises a barcode sequence and a primer binding sequence; (viii) ligating a second adaptor to the plurality of extended primers via branched ligation to generate adapted extended primers. 7. The method of embodiment 6, further comprising:
[0271] Embodiment 16 is the method of any of embodiments 6 to 15, further comprising the steps of amplifying the adapted extended primer to generate amplified double-stranded fragments; selecting amplified double-stranded fragments having a length in the range of 300 bases to 1000 bases; and and sequencing the selected amplified double-stranded, adapted fragments.
[0272] Embodiment 17 is the method of embodiments 1-16, wherein the single-stranded DNA circle is prepared in solution without a solid support.
[0273] Embodiment 18 is the method of embodiment 6, wherein the first end or the second end is attached to a solid support.
[0274] Embodiment 19 is a method for generating a double-stranded adapted construct for sequencing, said method comprising: (i) amplifying a plurality of genomic fragments, each genomic fragment comprising a target sequence, to generate a plurality of sets of amplified nucleic acid fragments in the mixture, wherein the amplified nucleic acid fragments of each set share the same target sequence, and optionally amplification is performed using target-specific primers for each set; The method further comprises: (ii) contacting the amplified nucleic acid fragments with an enzyme to introduce cleavage in the amplified nucleic acid fragments; (iii) distributing the mixture of fragments into a plurality of aliquots; (iv) performing nick translation on aliquots of the fragments and synthesizing DNA strands under conditions such that the DNA strands synthesized in different aliquots have different lengths; each of the DNA strands comprises a target sequence portion at a first end and a second end; and wherein the DNA strands of the different aliquots share the same sequence near a first end and have different sequences near a second end; (v) for each aliquot, ligating a second adapter to the second end of the DNA strand synthesized in step (iv) via branched ligation, each second adaptor is a partially double-stranded adaptor comprising a first adaptor oligonucleotide and a second adaptor oligonucleotide; the first adaptor oligonucleotide and the second adaptor oligonucleotide are both complementary and hybridize to each other; each of the second adaptors comprises a positional barcode sequence; each ligation comprises joining the 5-prime end of a first adapter oligonucleotide of a second adapter to the second end of a synthesized DNA strand; wherein the first adaptor oligonucleotides ligated to the second ends of the synthesized DNA strands of different aliquots comprise different position barcode sequences, and the first adaptor oligonucleotides ligated to the second ends of the synthesized DNA strands of the same aliquot share the same position barcode sequence; (vi) combining the synthesized DNA strands ligated with the second adapter from different aliquots from step (v) in a single mixture; (vii) extending a primer hybridized to the first adaptor oligonucleotide ligated to the synthesized DNA strand to generate a double-stranded fragment with blunt ends; and (viii) optionally selecting double-stranded fragments of step (vii) in the size range of 200 bp to 1.5 kb (the present disclosure provides an optimal range of around 500 to 1000 bp) from a single mixture; and (ix) ligating a third adaptor to the blunt end of the double-stranded fragment, thereby generating a double-stranded adapted construct.
[0275] Embodiment 20 is a method for preparing a nucleic acid fragment encoding a uracil-containing nucleic acid fragment, wherein step (i) comprises amplifying a plurality of genomic fragments in a mixture comprising uracil, thereby generating amplified nucleic acid fragments having incorporated uracil; and 20. The method of embodiment 19, wherein step (ii) comprises contacting the amplified nucleic acid fragments with uracil-DNA glycosylase, wherein the uracil-DNA glycosylase removes uracil from the amplified genomic fragments.
[0276] Embodiment 21 is the method of embodiment 19, wherein amplifying the plurality of genomic fragments in step (i) is performed using primers that contain uracil, thereby generating a plurality of sets of amplified nucleic acid fragments that contain uracil.
[0277] Embodiment 22 is the method of embodiment 21, wherein each of the plurality of genomic fragments is amplified using a forward primer and a reverse primer, each forward primer comprising one or more uracils.
[0278] Embodiment 23 is the method of embodiment 22, wherein each of the plurality of genomic fragments is amplified using a forward primer and a reverse primer, and each reverse primer contains a single uracil.
[0279] Embodiment 24 is the method of embodiment 19, wherein step (ii) comprises contacting the amplified genomic fragments with an endonuclease, wherein the endonuclease randomly cleaves the amplified genomic fragments.
[0280] Embodiment 25 is the method of embodiment 24, wherein the endonuclease is EndoIV or APE1.
[0281] Embodiment 26 is a reaction mixture containing the single-stranded DNA circle produced in paragraph 6.
[0282] Embodiment 27 is a reaction mixture comprising the combined synthesized DNA strands from step (vi) of paragraph 18.
[0283] Embodiment 28 is a method for preparing multiple nested sets of adapted fragments, comprising: each adapted fragment is a single-stranded nucleic acid comprising a target sequence fragment having a first end and a second end, a 5-prime adaptor sequence, a 3-prime adaptor sequence, a primer sequence, and a barcode sequence; in each nested set of adapted fragments, the target sequence fragments have identical nucleotide sequences at a first end and differ from one another by truncation at a second end such that each nested set of adapted fragments comprises a plurality of target sequence fragments having different lengths; the first end is closer to the barcode sequence than the second end; The method comprises: (a) providing a population of single-stranded DNA concatemers in a reaction, each concatemer comprises a plurality of identical monomers, each monomer comprising a complement of a target sequence, a complement of a barcode sequence that identifies the concatemer, and a primer binding sequence shared by the population of single-stranded concatemers; the primer binding sequence comprises a sequence that is complementary to the primer sequence; the complement of the primer binding sequence and the barcode sequence are both 3-prime to the complement of the target sequence; (b) annealing a primer comprising a primer sequence to the primer binding sequences of the plurality of monomers of each of the plurality of concatemers; (c) extending at least a portion of the primers hybridized to the primer binding sequences with a DNA polymerase having 5' to 3' exonuclease activity and no strand displacement activity, said extension producing a plurality of extended primers, each extended primer comprising a target sequence fragment having a barcode sequence and a primer sequence; The extended primer is hybridized to the concatemer, the extended primers are separated by an interval; and (d) contacting the plurality of extended primers with 5-prime adapters comprising 5-prime adapter sequences, 3-prime adapters comprising 3-prime adapter sequences, a DNA ligase, and an exonuclease having single-stranded DNA exonuclease activity under conditions where the exonuclease degrades a portion of the target sequence fragment of the extended primer to generate shortened extended primers, 5-prime adapters linked to the 5' ends of the shortened extended primers, and 3-prime adapters linked to the 3' ends of the shortened extended primers; thereby generating a group of multiple nested sets of adapted fragments.
[0284] Embodiment 29 is the method of embodiment 28, wherein the population of single-stranded DNA concatemers is generated by rolling circle replication of circular templates, each of the circular templates comprising a target sequence, a barcode sequence, and a primer sequence.
[0285] Embodiment 30 is the method of embodiment 28, wherein the 5-prime adapter is an L-adapter and the 3-prime adapter is a branched adapter.
[0286] Embodiment 31 is the method of embodiment 28, further comprising adding a nuclease to extend the gap formed in step (c), wherein the nuclease has single-stranded exonuclease activity.
[0287] Embodiment 32 is the method of embodiment 31, wherein at least some of the primers are RNA primers, the nuclease is RNAse H, and the RNAse H digests the RNA primers, thereby extending the spacing.
[0288] Embodiment 33 is the method of embodiment 28, wherein the primer binding sequence is located 3-prime to the complementary strand of the barcode sequence of step (a); the exonuclease has 3' to 5' exonuclease activity, and A method in which the barcode sequence of each of the set of adapted fragments is located 5-prime to the target sequence fragment.
[0289] Embodiment 34 is the method of embodiment 28, wherein the primer binding sequence is located 5-prime to the complementary strand of the barcode sequence of step (a); the exonuclease has 5' to 3' exonuclease activity, and A method in which the barcode sequence is 3-prime to the target sequence fragment of each of the adapted fragments.
[0290] Embodiment 35 is the method of any preceding paragraph, wherein both the 5-prime adapter and the 3-prime adapter are present in solution.
[0291] Embodiment 36 is the method of embodiment 35, wherein the reaction does not involve a solid support.
[0292] Embodiment 37 is the method of any preceding clause, wherein the target sequence has a length between 500 bases and 50 kilobases.
[0293] Embodiment 38 is a method for preparing a blunt-ended branched adaptor comprising: a 5'-end of one strand and a 3'-end of a complementary strand; 31. The method of embodiment 30, wherein the 5' end of the strand at the blunt end of the double strand is ligated to the 3' end of at least one of the extended primers via a branch ligation.
[0294] Embodiment 39 is the method of embodiment 30, wherein the L-adapter is 3-prime and comprises 1 to 10 degenerate bases.
[0295] Embodiment 40 is a method for preparing multiple nested sets of adapted fragments, comprising: each adapted fragment is a single-stranded nucleic acid comprising a target sequence fragment having a first end and a second end, a 5-prime adaptor sequence, a 3-prime adaptor sequence, a primer binding sequence, and a complement of a barcode sequence; in each nested set of adapted fragments, the target sequence fragments have identical nucleotide sequences at a first end and differ from each other at a second end such that each nested set of adapted fragments comprises a plurality of target sequence fragments having different lengths; the first end is closer to the barcode sequence than the second end; The method comprises: (a) providing barcoded fragments comprising a barcode sequence, a target sequence, and a primer binding sequence, wherein the barcoded fragments are immobilized at one end on beads; (b) annealing a primer comprising a 5-prime adapter sequence to the primer binding sequence of the barcoded fragment, the 5-prime adapter sequence comprises i) the complement of the barcode sequence, and ii) a primer sequence that is complementary to the primer binding sequence of the barcoded fragment; (c) extending the primer to generate an extended primer comprising the complement of the target sequence fragment and the barcode sequence; (d) contacting the extended primer with a branched adapter comprising a 3-prime adapter sequence to generate an adapted fragment; (e) separating the adapted fragments from the barcoded fragments that remain immobilized on the beads; (f) repeating steps (b) through (e) for one or more cycles under controlled extension conditions to generate one or more adapted fragments; the adapted fragments generated from step (e) and the adapted fragments generated from step (f) constitute a nested set of adapted fragments; wherein the adapted fragments of each nested set comprise target sequence fragments having different lengths.
[0296] Embodiment 41 is the method of embodiment 40, wherein the primer is extended with uracil in one or more extension cycles under extension control conditions to generate an extended primer, thereby generating an adapted fragment incorporating uracil at the 5-prime portion of the target sequence fragment; (g) contacting the adapted fragment with an enzyme that removes incorporated uracil, thereby creating at least one gap flanked by an exposed 3-prime end and an exposed 5-prime end of the adapted fragment; (h) ligating an internal branched adaptor to the exposed 3-prime end in at least one interval and ligating an L-adaptor to the exposed 5-prime end in said interval; (i) joining the internal branched adaptor ligated to the exposed 3-prime end in step (h) with the L-adaptor ligated to the exposed 5-prime end, thereby creating a shortened adapted fragment; This method generates a set of shortened, adapted fragments comprising shortened target sequence fragments having sequences that correspond to different regions of the target sequence, the different regions overlapping.
[0297] Embodiment 42 is the method of embodiment 41, wherein the step of ligating the internal branched adapter and the L-adapter comprises contacting the internal branched adapter and the L-adapter with a splint oligonucleotide; the splint oligonucleotide comprises a 5-prime portion that is complementary to the sequence of the internal branch adapter and a 3-prime portion that is complementary to the L-adapter; This method involves hybridizing the splint oligonucleotide to the internal branch adapter via the 5-prime portion and the splint oligonucleotide to the L-adapter via the 3-prime portion, thereby linking the internal branch adapter and the L-adapter.
[0298] Embodiment 43 is a method for preparing a plurality of sets of adapted fragments, comprising: each adapted fragment is a single-stranded nucleic acid comprising a target sequence fragment having a first end and a second end, a 5-prime adaptor sequence, a 3-prime adaptor sequence, a primer binding sequence, and a complement of a barcode sequence; The method comprises: (a) providing barcoded fragments comprising a barcode sequence, a target sequence, and a primer binding sequence, wherein the barcoded fragments are immobilized at one end on beads; (b) annealing a primer comprising a 5-prime adapter sequence to the primer binding sequence of the barcoded fragment, the 5-prime adapter sequence comprises i) the complement of the barcode sequence, and ii) a primer sequence that is complementary to the primer binding sequence of the barcoded fragment; (c) extending the primer to generate an extended primer comprising the complement of the target sequence fragment and the barcode sequence; (d) contacting the extended primer with a first branched adapter comprising a 3-prime portion that includes a degenerate sequence region, thereby forming a first extension product that includes a degenerate sequence region at the 3-prime portion; hybridizing the 3-prime portion to the barcoded fragment via the degenerate sequence region; (e) extending the 3-prime portion of the first extension product to generate a second extension product; and (f) contacting the second extension product with a second branched adapter to generate an adapted fragment.
[0299] Embodiment 44 is the method of embodiment 43, further comprising the step of (g) denaturing and separating the adapted fragments from the barcoded fragments.
[0300] Embodiment 45 is the method of embodiment 44, further comprising repeating steps (b) through (g) for one or more cycles under controlled extension conditions to generate one or more adapted fragments.
[0301] Embodiment 46 is a DNA complex comprising a plurality of fragments hybridized to one or more monomers of DNA concatemers, The fragments are separated by intervals, each of the plurality of fragments comprises a barcode sequence and a target sequence fragment having a first end and a second end; A DNA complex in which the target sequence fragments of the plurality of fragments have the same nucleotide sequence at a first end and differ from one another by cleavage at a second end such that the target sequence fragments of the plurality of fragments have different lengths.
[0302] Embodiment 47 is the DNA complex of embodiment 46, wherein each of the plurality of fragments is ligated at its 5-prime end to an L-adaptor and at its 3-prime end to a branched adaptor.
[0303] Embodiment 48 is a DNA complex comprising: (a) a barcoded fragment immobilized on a solid support, the barcode fragment comprising a barcode sequence and a target sequence; (b) a polynucleotide hybridized to the barcoded fragment, the polynucleotide comprises a 5-prime portion comprising the complement of the barcode sequence, a 3-prime portion comprising a target sequence fragment, A DNA complex comprising a polynucleotide in which a 5-prime portion and a 3-prime portion anneal to the barcoded fragment, leaving a central portion that does not anneal to the barcoded fragment, thereby forming a bubble.
[0304] Embodiment 49 is the plurality of DNA complexes of any of embodiments 46 to 48, wherein the DNA complexes share the same barcode sequence.
[0305] Embodiment 50 is a composition comprising a nested set of adapted fragments, each comprising a barcode sequence and a target sequence fragment having a first end and a second end, a 5-prime adapter sequence, and a 3-prime adapter sequence; the target sequence fragments have identical nucleotide sequences at a first end and differ from one another by truncation at a second end such that the nested set of adapted fragments comprises target sequence fragments having a plurality of different lengths; A composition in which the adapted fragments share the same barcode sequence.
[0306] Embodiment 2.1. A method for generating single-stranded DNA circles containing single-stranded, adapted constructs for sequencing, comprising:
[0307] Preparing a plurality of nested sets of single-stranded nucleic acid constructs, comprising: each single-stranded nucleic acid construct of each nested set comprises a target sequence portion flanked by a first adaptor sequence and a second adaptor sequence; the first adapter sequence comprises a barcode sequence and a primer binding sequence;
[0308] each target sequence portion having a first end and a second end;
[0309] the distance between the first end and the barcode sequence is shorter than the distance between the second end and the barcode sequence;
[0310] the single-stranded nucleic acid constructs of each nested set share the same barcode sequence, and the single-stranded nucleic acid constructs of different nested sets have different barcode sequences;
[0311] For each nested set of single-stranded nucleic acid constructs:
[0312] (a) the target sequence portions of a nested set of single-stranded nucleic acid constructs have identical nucleotide sequences near a first end and differ from one another by truncations near a second end, such that each nested set includes multiple target sequence portions having different lengths;
[0313] (b) circularizing the single-stranded nucleic acid constructs of each nested set to generate single-stranded DNA circles to which the first adaptor sequence and the second adaptor sequence are joined; preparing a method comprising the steps of:
[0314] Embodiment 2.2. The method of embodiment 2.1, wherein each nested set of single-stranded nucleic acid constructs comprises:
[0315] (i) preparing an adapted double-stranded genomic fragment comprising a target sequence flanked by a first adaptor sequence and a third adaptor sequence;
[0316] (ii) amplifying the adapted double-stranded genomic fragment using primers hybridized to the first and third adapter sequences to generate an amplified genomic fragment;
[0317] (iii) contacting the amplified genomic fragment from step (ii) with a nicking agent to generate a nick at the target sequence on one strand of the amplified genomic fragment;
[0318] (iv) ligating a second adapter comprising a second adapter sequence at the nick of step (iii) via branch ligation to form a ligated product; and
[0319] (v) denaturing the ligated products from step (iv) to form single-stranded nucleic acid constructs each comprising a first adaptor sequence and a second adaptor sequence; The method of claim 1,
[0320] Embodiment 2.3. The method of embodiment 2.1, wherein each nested set of single-stranded nucleic acid constructs comprises:
[0321] (i) preparing adapted double-stranded genomic fragments, each comprising a target sequence flanked by a first adaptor sequence and a third adaptor sequence;
[0322] (ii) amplifying the adapted double-stranded genomic fragments to produce amplified genomic fragments;
[0323] (iii) distributing the amplified genomic fragments into multiple aliquots;
[0324] (iv) denaturing the amplified genomic fragments of step (iii) to prepare single-stranded genomic fragments, at least a portion of each of the single-stranded genomic fragments comprising a primer binding sequence;
[0325] (iv) extending the primer hybridized to the primer binding sequence under extension control conditions such that the extension products from the different aliquots differ in length, thereby generating extension products with newly formed termini, the extension products having different sequences near the newly formed termini in the different aliquots;
[0326] each extension product comprising a target sequence portion;
[0327] (v) ligating a second adapter comprising a second adapter sequence at the newly formed end via branched ligation in each aliquot, thereby generating single-stranded nucleic acid constructs each comprising the first adapter sequence and the second adapter sequence; The method for preparing the compound according to the present invention is as follows.
[0328] Embodiment 2.4. The method of embodiment 2.1, wherein each nested set of single-stranded nucleic acid constructs comprises:
[0329] (i) preparing adapted double-stranded genomic fragments, each comprising a target sequence flanked by a first adaptor sequence and a third adaptor sequence;
[0330] (ii) amplifying the adapted double-stranded genomic fragments to produce amplified genomic fragments;
[0331] (iii) distributing the amplified genomic fragments into multiple aliquots;
[0332] (iv) adding a double-stranded DNA nuclease having 3' to 5' nuclease activity, varying the multiple aliquots under controlled conditions such that the length of the product is maintained after digestion in the different aliquots, thereby generating digestion products with newly formed ends having different sequences in the different aliquots;
[0333] each digestion product containing a portion of the target sequence; and
[0334] (v) ligating a second adapter comprising a second adapter sequence at the newly formed end via branched ligation in each aliquot, thereby generating single-stranded nucleic acid constructs each comprising the first adapter sequence and the second adapter sequence; The method for preparing the compound according to the present invention is as follows.
[0335] Embodiment 2.5. The method of embodiment 2.1, wherein each nested set of single-stranded nucleic acid constructs comprises:
[0336] (i) preparing adapted double-stranded genomic fragments, each comprising a target sequence flanked by a first adaptor sequence and a third adaptor sequence;
[0337] (ii) amplifying the adapted double-stranded genomic fragments to produce amplified genomic fragments;
[0338] (iii) denaturing the amplified genomic fragments to prepare single-stranded genomic fragments, at least a portion of each of the single-stranded genomic fragments comprising a primer binding sequence;
[0339] (iv) for each single-stranded genomic fragment:
[0340] extending the primer hybridized to the primer binding sequence for an initial period of time to produce an extended primer;
[0341] the extension is incomplete such that the length of the extended primer is a fraction of the length of the single-stranded genomic fragment;
[0342] the extended primer comprises a portion of the target sequence, and
[0343] ligating a second adaptor via branched ligation at the newly formed end of the extended primer formed by the extension; thereby generating single-stranded nucleic acid constructs in a single reaction mixture, each comprising a first adapter sequence and a second adapter sequence;
[0344] (v) repeating step (iv) multiple times, each time further extending the primer for an additional period of time and ligating additional adapters having unique positional barcodes to the further extended primers;
[0345] Additional adapters are used in molar amounts that are fractions of the total molar amount of amplified genomic fragments;
[0346] thereby generating a mixture of a nested set of single-stranded nucleic acid constructs; The method for preparing the compound according to the present invention is as follows.
[0347] Embodiment 2.6. The method of embodiment 2.3, wherein the target sequence comprises a repetitive sequence; the second adapter comprises a positional barcode sequence that is unique for each aliquot;
[0348] The method, wherein the single-stranded nucleic acid constructs formed in step (v) in different aliquots comprise different positional barcode sequences, and the single-stranded nucleic acid constructs in the same aliquot share the same positional barcode sequence.
[0349] Embodiment 2.7. The method of embodiment 2.1, wherein the primer binding sequence is 3-prime to the barcode sequence.
[0350] Embodiment 2.8. The method of embodiment 2.1, comprising:
[0351] (vi) fragmenting the single-stranded DNA circle to generate a plurality of single-stranded DNA fragments, at least some of which comprise the barcode sequence;
[0352] (vii) generating double-stranded DNA fragments from the single-stranded DNA fragments from step (vi);
[0353] (vii) ligating a second adaptor to each of the double-stranded DNA fragments from step (vii), thereby generating double-stranded adapted fragments.
[0354] Embodiment 2.9. The method of embodiment 2.8, comprising: (viii) amplifying the double-stranded adapted fragments;
[0355] Optionally, the method further comprises (ix) selecting amplified double-stranded adapted fragments having a length in the range of 300 to 1000 bases.
[0356] Embodiment 2.10. The method of embodiment 2.1, comprising:
[0357] (vi) hybridizing a primer to the primer binding sequence in each of the single-stranded DNA circles;
[0358] (vii) extending the primer under extension-controlled conditions using each of the single-stranded DNA circles as a template;
[0359] the extension generates an extended primer hybridized to the single-stranded DNA circle, thereby generating a plurality of extended primers having different lengths;
[0360] extending, wherein each of the extended primers comprises a barcode sequence and a primer binding sequence;
[0361] (viii) ligating a second adaptor to the plurality of extended primers via branched ligation to generate adapted extended primers. The method further comprises:
[0362] Embodiment 2.11. The method of embodiment 2.10, comprising:
[0363] amplifying the adapted extended primer to generate amplified double-stranded fragments;
[0364] Selecting amplified double-stranded fragments having a length in the range of 300 bases to 1000 bases (the present disclosure discloses an optimal length nesting range of approximately 600 bases); and
[0365] and sequencing the selected amplified double-stranded, adapted fragments.
[0366] Embodiment 2.12. The method of any one of embodiments 2.1 to 2.11, wherein the single-stranded DNA circle is prepared in solution without a solid support.
[0367] Embodiment 2.13. The method of embodiment 2.1, wherein the first end or the second end is attached to a solid support.
[0368] Embodiment 2.14. A method of generating a double-stranded adapted construct for sequencing, said method comprising:
[0369] (i) amplifying a plurality of genomic fragments, each genomic fragment comprising a target sequence, to generate a plurality of sets of amplified nucleic acid fragments in a mixture, wherein the amplified nucleic acid fragments of each nested set share the same target sequence, and optionally the amplification is performed using target-specific primers;
[0370] For each nested set, the method further comprises:
[0371] (ii) contacting the amplified nucleic acid fragments with an enzyme to introduce cleavage in the amplified nucleic acid fragments;
[0372] (iii) distributing the mixture of fragments into a plurality of aliquots;
[0373] (iv) performing nick translation on aliquots of the fragments and synthesizing DNA strands under conditions such that the DNA strands synthesized in different aliquots have different lengths; each of the DNA strands comprises a target sequence portion at a first end and a second end; and wherein the DNA strands of the different aliquots share the same sequence near a first end and have different sequences near a second end;
[0374] (v) for each aliquot, ligating a second adapter to the second end of the DNA strand synthesized in step (iv) via branched ligation, each second adaptor is a partially double-stranded adaptor comprising a first adaptor oligonucleotide and a second adaptor oligonucleotide;
[0375] the first adaptor oligonucleotide and the second adaptor oligonucleotide are both complementary and hybridize to each other;
[0376] each of the second adaptors comprises a positional barcode sequence;
[0377] each ligation comprises joining the 5-prime end of a first adapter oligonucleotide of a second adapter to the second end of a synthesized DNA strand;
[0378] wherein the first adaptor oligonucleotides ligated to the second ends of the synthesized DNA strands of different aliquots comprise different position barcode sequences, and the first adaptor oligonucleotides ligated to the second ends of the synthesized DNA strands of the same aliquot share the same position barcode sequence;
[0379] (vi) combining the synthesized DNA strands ligated with the second adapter from different aliquots from step (v) in a single mixture;
[0380] (vii) extending a primer hybridized to the first adaptor oligonucleotide ligated to the synthesized DNA strand to generate a double-stranded fragment with blunt ends; and
[0381] (viii) optionally selecting double-stranded fragments of step (vii) in the size range of 200 bp to 1.5 kb (the present disclosure provides an optimal range of around 500 to 1000 bp) from a single mixture; and
[0382] (ix) ligating a third adaptor to the blunt end of the double-stranded fragment, thereby generating a double-stranded adapted construct.
[0383] Embodiment 2.15. The method of embodiment 2.14, comprising: step (i) comprises amplifying a plurality of genomic fragments in a mixture comprising uracil, thereby producing amplified nucleic acid fragments having uracil incorporated therein; and
[0384] The method, wherein step (ii) comprises contacting the amplified nucleic acid fragments with uracil-DNA glycosylase, wherein the uracil-DNA glycosylase removes uracil from the amplified genomic fragments.
[0385] Embodiment 2.16. The method of embodiment 2.14, comprising: A method wherein the step of amplifying the plurality of genomic fragments in step (i) is carried out using primers that contain uracil, thereby generating a plurality of sets of amplified nucleic acid fragments that contain uracil.
[0386] Embodiment 2.17. The method of embodiment 2.16, comprising: A method in which each of the plurality of genomic fragments is amplified using a forward primer and a reverse primer, each forward primer containing one or more uracils.
[0387] Embodiment 2.18. The method of embodiment 2.17, comprising: A method in which each of the plurality of genomic fragments is amplified using a forward primer and a reverse primer, each reverse primer containing a single uracil.
[0388] Embodiment 2.19. The method of embodiment 2.14, comprising: The method, wherein step (ii) comprises contacting the amplified genomic fragments with an endonuclease, wherein the endonuclease randomly cleaves the amplified genomic fragments.
[0389] Embodiment 2.20. The method of embodiment 2.19, wherein the endonuclease is EndoIV or APE1.
[0390] Embodiment 2.21. A reaction mixture comprising the single-stranded DNA circle produced in embodiment 2.7.
[0391] Embodiment 2.22. A reaction mixture comprising the combined synthesized DNA strands from step (vi) of embodiment 1.19. [Example]
[0392] Example
[0393] 1. Example 1. Cover Coding in Solution Using Unscheduled Random Nicking
[0394] This example describes full-coverage generation of DNA molecules between 1 and 20 kb. This approach can be useful for most sequence assembly, especially using sequencing platforms such as SE400-SE1000 and PE300+MPS reads. Positional co-barcoding, as described herein, is only necessary if the target nucleic acid contains highly repetitive sequences. This method begins by ligating a barcoded adapter to one end of the molecule and a non-barcoded adapter to the opposite end, achieved by Y-adapter ligation or other commonly used methods. This method can also be used for target sequences by adding a common adapter tag to each PCR primer and including a barcode on one of the adapter tags in each PCR primer pair. After PCR, the molecules are treated with a nonspecific nicking enzyme at low concentration, low temperature, and / or short time to introduce a nick within each template. If necessary, this nick can be widened to a gap of several bases in the sequence using 5' or 3' exonuclease or a dNTP-free polymerase. Branch ligation then occurs, and another adapter is added. The molecule is then circularized using a splint oligonucleotide between the branched ligation adapter and the adapter containing the barcode (Figure 3A). The circle is then fragmented into 500-1000 base pairs, followed by primer extension from the barcode adapter so that the barcode is copied. After another round of adapter ligation, the molecule can be directly sequenced or PCR-amplified and sequenced.
[0395] Another embodiment of this process uses controlled extension of approximately 600 bases after circularization, followed by 3' branch ligation and PCR (Figure 4B). This has the advantage of producing products within a relatively narrow size range, as opposed to the broad size distribution that random fragmentation produces. Short artifacts are removed via exonuclease treatment or purification.
[0396] The end result of this process is a series of overlapping sequence reads from each original DNA molecule, all of which share the same barcode sequence. Random nicking provides similar coverage of short and long DNA molecules from the same pool, allowing for the complete reassembly of each original DNA molecule.
[0397] 2. Example 2. Positional Covercoding of 1-20 kb Fragments in Solution
[0398] This example is similar to Example 1, except that it incorporates a barcode (co-barcode) that can be shared among all subfragments of the original molecule. The process begins by using target PCR primers containing a common adapter tag and a random barcode to amplify specific regions, or by ligating adapters with a common sequence and a random barcode to dsDNA fragments of 1-20 kb in length. Preferably, pools contain DNA fragments of similar lengths. For pools containing both long and short DNA fragments, specific methods can be used to minimize overcoverage of the short fragments. The products are PCR amplified and then divided into 10-20 pools, each of which undergoes different amounts of controlled extension or ExoIII digestion (as described above), or controlled nick translation. Short DNA fragments are fully extended to form blunt ends, and these blunt-ended fragments can be blocked from branch ligation using methods known in the art, such as DNA tailing or 3' blocking with terminal transferase. This step is followed by 3' branch ligation, where adapters containing a common sequence and a barcode sequence specific to each pool (position barcode) are added. Unlike the above, the product is circularized, linking the DNA molecule barcode and the location barcode with a common adapter sequence between them. The circle is then fragmented into 500-1000 base pairs, and the fragments are primer-extended with primers 5' to both the molecule and the location barcode. After extension, a third adapter can be ligated and sequenced, or PCR and sequencing can be performed. Alternatively, a blunt-ended third adapter with an unphosphorylated 5' end can be ligated and PCR and sequencing can be performed. (Figure 3B)
[0399] Another approach that can be adopted based on this process is to perform the above reactions in a single tube rather than in separate pools. This can be achieved by adding limited amounts of 3' branched ligation adapters with different sequences after each time interval. For example, the first 3' branched ligation adapter is added 10 minutes after the start of elongation, and the second 3' branched ligation adapter is added 20 minutes after the start of elongation, with the first and second 3' branched ligation adapters having different sequences. The ideal amount of 3' branched ligation adapter is enough to ligate 1-10% of the total number of molecules (depending on the length of the molecules used). This process of adding limited amounts of adapters can be repeated 10-20 times or more (corresponding to the total number of pools used in other approaches). This has the advantage of being performed in a single tube, but requires multiple pipetting of adapters into the same tube.
[0400] 3. Example 3. Target-enriched 2-20 kb pool for unrelated sequences
[0401] The method described in this example involves performing long-range PCR on target regions of the genome. In some cases, multiplex PCR can be performed, with 100–1000 different target regions amplified in one or more reactions. After amplification, the products are separated into different pools. The number of pools varies depending on the size of the amplicon and the length of the sequencing reads to be used. However, for a 5 kb product with approximately 500 bases of reads (250 paired-end or 500 single-end), approximately 10–20 pools are sufficient. Each pool is subjected to timed digestion with ExoIII or controlled extension with a polymerase with 5'-3' exoactivity (e.g., E. coli DNA polymerase 1). Importantly, the duration of ExoIII treatment is varied in steps (e.g., 1, 2, 3, 4 minutes) for each pool, ensuring that the amount of digestion between each pool is approximately 500 bases in this example. Similarly, when performing controlled extension or nick translation, similar results can be obtained by varying the duration or ratio of dNTPs. It is important to note that the amount of extension and digestion in each pool varies, resulting in a range of products rather than a specific length. This results in overlapping fragments between different pools, which makes in silico assembly of each original molecule much easier after sequencing. After this step, 3' branch ligation is performed, and adapters with a common sequence and barcode sequence unique to each pool are added. These adapters can contain biotin at the 3' end to aid in the purification step. The pools are then combined, and the products are fragmented to approximately 500-1000 base pairs, followed by primer extension from the adapter sequence and ligation of a second common adapter. The products can be PCR amplified or directly circularized for DNA fragment formation and sequencing. Figure 5.
[0402] 4. Example 4. Loop-extended complete stLFR
[0403] The complete loop-mediated stLFR involves ligating two functionalized partially double-stranded blunt-end adapters to an end-repaired DNA fragment bearing a 5′-phosphate group.
[0404] Ligation of blunt-end adapters to end-repaired DNA fragments bearing 5'-phosphate groups
[0405] The first partially double-stranded blunt-end adapter (e.g., AD153UMI_5, Figure 13A) has a long strand (1313) and a short strand (1314) that anneal to form a blunt end and an unpaired end. The blunt ends are ligatable. The long strand (1313) includes a single-stranded 5'-overhang, which includes one or more barcodes (UIDs) (1319), such as unique molecular identifiers (UMIs) or multiplexed sample barcodes (1319). The single-stranded 5'-overhang may include a sequence complementary to a universal amplification primer. The barcode sequence may exist as a single sequence or as several separate sequences. Optionally, the single-stranded 5'-overhang includes a single T-overhang at its 3' end for ligation to an A-tailed DNA fragment. The shorter strand (1314) annealed to the longer strand (1313) contains a 5'-end that does not contain a 5'-phosphate group and therefore cannot ligate to the DNA fragment. The 3'-end of the shorter strand (1314) is also modified to prevent ligation. 3'-modifications to prevent ligation include, but are not limited to, inverted nucleosides, dideoxynucleosides, 3'-amino groups, and 3'-phosphate groups. When the first partially double-stranded blunt-end adaptor contacts the genomic DNA (1310) under ligation conditions, only the longer strand (1313) ligates to the gDNA, while the shorter strand (1314), which has neither a ligatable 5' nor a 3' end, cannot ligate to the genomic DNA fragment, leaving a nick (1317) between the shorter strand and the gDNA strand. See Figure 13A.
[0406] The second partially double-stranded blunt-end adapter (eg, Ad183 shown in Figure 13) is designed similarly to the first, with the exception that it does not contain a barcode sequence (UID).
[0407] The genomic fragments are then ligated with the adapters described above and purified using SPRI bead purification (Beckman Coulter Life Sciences, Indianapolis, IN). The adapter-ligated DNA molecules are then enzymatically extended from the 3'-end of the genomic DNA fragment to the unligatable 5'-end of the short strand of the partially double-stranded adapter (1314) using a DNA polymerase with strand displacement activity (e.g., Bst DNA polymerase, large fragment; Phi29 DNA polymerase; Bsu DNA polymerase, large fragment; Bsm DNA polymerase, large fragment) or a DNA polymerase with 5'-3' exonuclease activity (e.g., rTaq DNA polymerase, E. coli DNA polymerase I). This results in the formation of a double-stranded adapter-ligated DNA molecule (1318) with the double-stranded adapter attached to the DNA fragment sequence. The 3' end of the short strand of the branched adapter (1320.2) contains a 15-20 base sequence complementary to the 5' end of the long strand of the branched adapter. The long strand of the branched adapter contains a barcode sequence and has a melting temperature (Tm) of 50°C to 70°C. The 3' end of the short strand (1320.2) is blocked with a 3'-terminal modification (e.g., dideoxynucleoside, 3'-amino group, 3'-phosphate group) that prevents ligation. (See Figure 13B.) The adapter sequence derived from the first partially double-stranded adapter contains a site for a universal amplification primer; therefore, the double-stranded adapter-ligated DNA molecule (1318) can be amplified by PCR or other amplification methods that rely on two priming sequences. (See Figure 13A.)
[0408] Random nicking and gapping occurs. This can be achieved by using nonspecific nicking nucleases (e.g., Vvn and mutants, shrimp dsDNA-specific endonuclease, DNAse I) that cleave only one strand of the DNA backbone per catalytic reaction. Alternatively, a mixture of multiple nicking enzymes, such as several site-specific nickases (e.g., CCD), can be used. In some cases, additional enzymes with 3' exoactivity (e.g., DNA polymerase I, nucleotide-free Klenow fragment, exonuclease III, or similar) or 5' exoactivity (nucleotide-free Bst DNA polymerase full-length, T7 exonuclease, exonuclease VIII truncated, lambda exonuclease, T5 exonuclease, or similar) can be added to increase nick opening and allow for branched adapter ligation. Exonucleases with low processivity are preferred because they open short gaps (e.g., 2–7 bases) and dissociate from the DNA to allow adapter ligation. Figures 13A and 13B. Random nicking generates a set of fragments with target sequence fragments of different lengths, sharing a common barcode sequence 5'. One such fragment is shown as 1341 in Figure 13B.
[0409] Next, a branched adapter (e.g., AD153UMI_5R shown in Figure 13B) is ligated to the 3' side of the nick or ssDNA gap of the adapter-ligated DNA fragment in the presence of T4 DNA ligase. The branched adapter (1320) is a partially double-stranded DNA adapter molecule with a 5'-phosphate on the longer strand (1320.1). This generates a set of fragments with target sequence fragments of different lengths, each flanked by the adapter sequences AD153UMI_5 and AD153UMI_5R, one of which is shown as 1342 in Figure 13B.
[0410] The longer strand of the first partially double-stranded blunt-end adapter and the longer strand of the branched adapter each comprise a first hybridization sequence (1432) and a second hybridization sequence (1433). The first hybridization sequence (1432) is located 3' to the barcode sequence (1319).
[0411] When gaps are opened using exonucleases as described above, protection of the DNA adapter can be achieved, if necessary, via phosphorothioate bonds between bases at the 5' and 3' ends of the adapter and / or between modified bases.
[0412] Ligation of branched adapters to adapted DNA fragments (e.g., adapted genomic DNA fragments) can be performed in solution or on beads. If the process is performed on beads, the adapted DNA fragments are preloaded onto the beads with high-concentration PEG (5%–15%) before adding other reaction components.
[0413] In some embodiments, the branched ligation reaction is performed in the presence of additives (e.g., polyethylene glycol or betaine) to enhance the activity of the ligation and / or nicking enzymes. The reaction can be incubated at room temperature, 37°C, or cycled between various temperatures, such as 5-15°C and 37°C, at pH 5.0-9.0. The incubation time and nickase activity vary depending on the desired number of nicks per DNA fragment. The reaction can be terminated using a DNA purification method (e.g., Ampure XP beads) if performed in solution, or by a simple wash step with Tris-NaCl buffer containing PEG (5%-15%) if performed on beads.
[0414] In one embodiment, after branch ligation, the DNA fragments are denatured and the reaction mixture is heated to 90°C to 95°C (inclusive). Alternatively, the branched and ligated DNA fragments can be denatured with an alkaline agent (e.g., 0.05M to 0.2M NaOH or KOH) and further neutralized with a neutralizing agent (e.g., HCl, Tris-HCl, MOPS). A single-stranded DNA molecule (1343) containing genomic DNA and adapter sequences at both ends is formed. Figure 13B.
[0415] Alternatively, in some embodiments, instead of denaturation, the 5' single-stranded tail of the adapter-ligated DNA fragment required for hybridization of the 3' end of the branched adapter (e.g., 1432 in Figure 13 or Figure 14) can be generated by digestion with one or more dsDNA-specific exonucleases with 3'-5' exonuclease activity (e.g., exonuclease III).
[0416] DNA loop formation and enzymatic extension of the 3' end of the hybridized branched adapter
[0417] The long strand of the branched adapter and the long strand of the first partially double-stranded blunt-end adapter contain complementary sequences (1432 and 1433) and can hybridize to each other. Hybridization is performed in a hybridization buffer containing a buffer (e.g., Tris-HCl, MOPS, sodium phosphate), salt, and cofactors required for the subsequent enzymatic reaction, such as MgCl2 and dNTPs. Figure 14.
[0418] Following the DNA hybridization step, a linear extension step is performed on the 3' end of the hybridized branched adapter (e.g., AD153UMI_5R shown in Figure 14) to copy the barcode onto the first adapter, AD153UMI_5. To enhance specificity, linear extension is performed using a DNA polymerase lacking 3'-exonuclease activity, such as Taq DNA polymerase, Klenow fragment (3'→5' exonuclease), or Bst DNA polymerase, large fragment. Depending on the DNA polymerase used, extension can be performed at different temperatures ranging from 30°C to 75°C.
[0419] The products of linear extension (1431) represent partially double-stranded DNA molecules with double-stranded adapters containing UID sequences attached to the DNA fragments. The products are then denatured to form single-stranded sequences with adapter sequences at both ends (1441), bringing the ends of target sequence fragments 1411, 1421, 1431, etc., closer to the barcode sequence 1319. The adapter sequences contain sites for the barcode sequence and universal amplification primers and can therefore be used in the next step, controlled primer extension, to generate fragments of suitable length for sequencing.
[0420] ***
[0421] While the present invention has been disclosed with reference to specific aspects and embodiments, it will be apparent that other embodiments and variations of the present invention may be devised by those skilled in the art without departing from the true spirit and scope of the invention.
[0422] Each and every publication and patent document cited in this disclosure is herein incorporated by reference to the same extent as if each such publication or document was specifically and individually indicated to be incorporated by reference. Citation of publications and patent documents is not intended as an indication that any such document is pertinent prior art, nor does it constitute an admission as to the contents or date thereof.
Claims
1. 1. A method of generating a single-stranded, adapted construct for sequencing, comprising: Preparing a plurality of nested sets of single-stranded nucleic acid constructs, comprising: each single-stranded nucleic acid construct in each nested set comprises a target sequence portion flanked by a first adaptor sequence at the 5' end and a second adaptor sequence at the 3' end; the first adapter sequence comprises, from 5' to 3', a primer binding sequence, a barcode sequence, and a first hybridization sequence, and the second adapter sequence comprises a second hybridization sequence; the first and second hybridization sequences are complementary to each other; each target sequence portion having a first end and a second end; the distance between the first end and the barcode sequence is shorter than the distance between the second end and the barcode sequence; the single-stranded nucleic acid constructs in each nested set share the same barcode sequence, and the single-stranded nucleic acid constructs in different nested sets have different barcode sequences; preparing, for each nested set of single-stranded nucleic acid constructs, target sequence portions of the nested set having identical nucleotide sequences near a first end and differing from one another by truncations near a second end, such that each nested set of single-stranded nucleic acid constructs comprises a plurality of target sequence portions having different lengths; A method comprising:
2. 2. The method of claim 1, further comprising subjecting the plurality of nested sets of single-stranded nucleic acid constructs to hybridization conditions, whereby a first adaptor sequence hybridizes to a second adaptor sequence, thereby forming a loop.
3. 3. The method of claim 2, further comprising extending the second adapter sequence with a DNA polymerase to copy the barcode sequence and primer binding sequence of the first adapter sequence and form an extension product.
4. 4. The method of claim 3, further comprising the step of denaturing the extension products to open the loops, thereby forming linear single-stranded DNA constructs, each linear single-stranded DNA construct comprising a barcode sequence and a primer binding sequence, the primer binding sequence being located 3' to the barcode sequence.
5. 5. The method of claim 4, further comprising annealing a primer to a primer binding sequence 3' of the linear single-stranded DNA construct and extending the primer to generate an extension product having a length appropriate for sequencing.
6. 1. A method for generating single-stranded DNA circles containing single-stranded, adapted constructs for sequencing, comprising: Preparing a plurality of nested sets of single-stranded nucleic acid constructs, comprising: each single-stranded nucleic acid construct of each nested set comprises a target sequence portion flanked by a first adaptor sequence and a second adaptor sequence; the first adapter sequence comprises a barcode sequence and a primer binding sequence; each target sequence portion having a first end and a second end; the distance between the first end and the barcode sequence is shorter than the distance between the second end and the barcode sequence; the single-stranded nucleic acid constructs of each nested set share the same barcode sequence, and the single-stranded nucleic acid constructs of different nested sets have different barcode sequences; For each nested set of single-stranded nucleic acid constructs: (a) the target sequence portions of a nested set of single-stranded nucleic acid constructs have identical nucleotide sequences near a first end and differ from one another by truncations near a second end, such that each nested set includes a plurality of target sequence portions having different lengths; (b) circularizing the single-stranded nucleic acid constructs of each nested set to generate single-stranded DNA circles to which the first adaptor sequence and the second adaptor sequence are joined; preparing a method comprising the steps of:
7. 7. The method of claim 1, wherein each nested set of single-stranded nucleic acid constructs comprises: (i) preparing an adapted double-stranded genomic fragment comprising a target sequence flanked by a first adaptor sequence and a third adaptor sequence; (ii) amplifying the adapted double-stranded genomic fragment using primers hybridized to the first and third adapter sequences to generate an amplified genomic fragment; (iii) contacting the amplified genomic fragment from step (ii) with a nicking agent to generate a nick at the target sequence on one strand of the amplified genomic fragment; (iv) ligating a second adapter comprising a second adapter sequence at the nick of step (iii) via branch ligation to form a ligated product; and (v) denaturing the ligated products from step (iv) to form single-stranded nucleic acid constructs each comprising a first adaptor sequence and a second adaptor sequence; The method of claim 1,
8. 7. The method of claim 1, wherein each nested set of single-stranded nucleic acid constructs comprises: (i) preparing adapted double-stranded genomic fragments, each comprising a target sequence flanked by a first adaptor sequence and a third adaptor sequence; (ii) amplifying the adapted double-stranded genomic fragments to produce amplified genomic fragments; (iii) distributing the amplified genomic fragments into multiple aliquots; (iv) denaturing the amplified genomic fragments of step (iii) to prepare single-stranded genomic fragments, at least a portion of each of the single-stranded genomic fragments comprising a primer binding sequence; (iv) extending the primer hybridized to the primer binding sequence under extension control conditions such that the extension products from the different aliquots differ in length, thereby generating extension products with newly formed termini, the extension products having different sequences near the newly formed termini in the different aliquots; each extension product comprising a target sequence portion; (v) ligating a second adapter comprising a second adapter sequence at the newly formed end via branched ligation in each aliquot, thereby generating single-stranded nucleic acid constructs each comprising the first adapter sequence and the second adapter sequence; The method for preparing the compound according to the present invention is as follows.
9. 7. The method of claim 1, wherein each nested set of single-stranded nucleic acid constructs comprises: (i) preparing adapted double-stranded genomic fragments, each comprising a target sequence flanked by a first adaptor sequence and a third adaptor sequence; (ii) amplifying the adapted double-stranded genomic fragments to produce amplified genomic fragments; (iii) distributing the amplified genomic fragments into multiple aliquots; (iv) adding a double-stranded DNA nuclease having 3' to 5' nuclease activity, varying the aliquots under controlled conditions such that the length of the product is maintained after double-stranded DNA nuclease digestion in the different aliquots, thereby generating digestion products with newly formed ends having different sequences in the different aliquots; each digestion product containing a portion of the target sequence; and (v) ligating a second adapter comprising a second adapter sequence at the newly formed end via branched ligation in each aliquot, thereby generating single-stranded nucleic acid constructs each comprising the first adapter sequence and the second adapter sequence; The method for preparing the compound according to the present invention is as follows.
10. 7. The method of claim 1, wherein each nested set of single-stranded nucleic acid constructs comprises: (i) preparing adapted double-stranded genomic fragments, each comprising a target sequence flanked by a first adaptor sequence and a third adaptor sequence; (ii) amplifying the adapted double-stranded genomic fragments to produce amplified genomic fragments; (iii) denaturing the amplified genomic fragments to prepare single-stranded genomic fragments, at least a portion of each of the single-stranded genomic fragments comprising a primer binding sequence; (iv) for each single-stranded genomic fragment: extending the primer hybridized to the primer binding sequence for an initial period of time to produce an extended primer; the extension is incomplete such that the length of the extended primer is a fraction of the length of the single-stranded genomic fragment; the extended primer comprises a portion of the target sequence, and ligating a second adaptor to the end of the extended primer formed by the extension via a branched ligation; thereby generating single-stranded nucleic acid constructs in one reaction mixture, each comprising a first adapter sequence and a second adapter sequence; (v) repeating step (iv) multiple times, each time further extending the primer for an additional period of time and ligating additional adapters with unique positional barcodes to the further extended primers; Additional adapters are used in molar amounts that are fractions of the total molar amount of amplified genomic fragments; thereby generating a mixture of a nested set of single-stranded nucleic acid constructs; The method for preparing the compound according to the present invention is as follows.
11. 9. The method of claim 8, the target sequence comprises a repetitive sequence, the second adapter comprises a positional barcode sequence that is unique for each aliquot; The method, wherein the single-stranded nucleic acid constructs formed in step (v) in different aliquots comprise different positional barcode sequences, and the single-stranded nucleic acid constructs in the same aliquot share the same positional barcode sequence.
12. 7. The method of claim 6, wherein the primer binding sequence is 3-prime to the barcode sequence.
13. 7. The method of claim 6, (vi) fragmenting the single-stranded DNA circle to generate a plurality of single-stranded DNA fragments, at least some of which comprise the barcode sequence; (vii) generating double-stranded DNA fragments from the single-stranded DNA fragments from step (vi); (vii) ligating a second adaptor to each of the double-stranded DNA fragments from step (vii), thereby generating double-stranded adapted fragments.
14. 14. The method of claim 13, (viii) amplifying the double-stranded, adapted fragments; optionally, (ix) selecting amplified double-stranded, adapted fragments having a length in the range of 300 to 1000 bases.
15. 7. The method of claim 6, (vi) hybridizing a primer to the primer binding sequence in each of the single-stranded DNA circles; (vii) extending the primer under extension-controlled conditions using each of the single-stranded DNA circles as a template; the extension generates an extended primer hybridized to the single-stranded DNA circle, thereby generating a plurality of extended primers having different lengths; extending, wherein each of the extended primers comprises a barcode sequence and a primer binding sequence; (viii) ligating a second adaptor to the plurality of extended primers via branched ligation to generate adapted extended primers. The method further comprises:
16. A method according to any one of claims 6 to 15, amplifying the adapted extended primer to generate amplified double-stranded fragments; Selecting amplified double-stranded fragments having a length in the range of 300 bases to 1000 bases (an optimal length nesting range of approximately 600 bases is disclosed herein); and and sequencing the selected amplified double-stranded, adapted fragments.
17. The method according to claims 1 to 16, wherein the single-stranded DNA circle is prepared in solution without a solid support.
18. The method of claim 6 , wherein the first end or the second end is attached to a solid support.
19. 1. A method for generating a double-stranded, adapted construct for sequencing, said method comprising: (i) amplifying a plurality of genomic fragments, each genomic fragment comprising a target sequence, to generate a plurality of sets of amplified nucleic acid fragments in the mixture, wherein each set of amplified nucleic acid fragments shares the same target sequence, and optionally the amplification is performed using target-specific primers; For each set, the method further comprises: (ii) contacting the amplified nucleic acid fragments with an enzyme to introduce cleavage in the amplified nucleic acid fragments; (iii) distributing the mixture of fragments into a plurality of aliquots; (iv) performing nick translation on aliquots of the fragments and synthesizing DNA strands under conditions such that the DNA strands synthesized in different aliquots have different lengths; each of the DNA strands comprises a target sequence portion at a first end and a second end; and wherein the DNA strands of the different aliquots share the same sequence near a first end and have different sequences near a second end; (v) for each aliquot, ligating a second adaptor to the second end of the DNA strand synthesized in step (iv) via branched ligation, each second adaptor is a partially double-stranded adaptor comprising a first adaptor oligonucleotide and a second adaptor oligonucleotide; the first adaptor oligonucleotide and the second adaptor oligonucleotide are both complementary and hybridize to each other; each of the second adapters comprises a positional barcode sequence; each ligation comprises joining the 5-prime end of a first adapter oligonucleotide of a second adapter to the second end of a synthesized DNA strand; wherein first adapter oligonucleotides ligated to the second ends of the synthesized DNA strands of different aliquots contain different position barcode sequences, and first adapter oligonucleotides ligated to the second ends of the synthesized DNA strands of the same aliquot share the same position barcode sequence; (vi) combining in a single mixture the synthesized DNA strands ligated with second adapters from different aliquots from step (v); (vii) extending a primer hybridized to the first adaptor oligonucleotide ligated to the synthesized DNA strand to generate a double-stranded fragment with blunt ends; and (viii) optionally selecting double-stranded fragments of step (vii) in the size range of 200 bp to 1.5 kb from a single mixture; and (ix) ligating a third adaptor to the blunt end of the double-stranded fragment, thereby generating a double-stranded adapted construct.
20. 20. The method of claim 19, step (i) comprising amplifying a plurality of genomic fragments in a mixture comprising uracil, thereby producing amplified nucleic acid fragments having uracil incorporated therein; and The method, wherein step (ii) comprises contacting the amplified nucleic acid fragments with uracil-DNA glycosylase, wherein the uracil-DNA glycosylase removes uracil from the amplified genomic fragments.
21. 20. The method of claim 19, wherein amplifying the plurality of genomic fragments in step (i) is performed using primers that contain uracil, thereby generating a plurality of sets of amplified nucleic acid fragments that contain uracil.
22. 22. The method of claim 21, wherein each of the plurality of genomic fragments is amplified using a forward primer and a reverse primer, each forward primer comprising one or more uracils.
23. 23. The method of claim 22, wherein each of the plurality of genomic fragments is amplified using a forward primer and a reverse primer, each reverse primer containing a single uracil.
24. 20. The method of claim 19, wherein step (ii) comprises contacting the amplified genomic fragments with an endonuclease, wherein the endonuclease randomly cleaves the amplified genomic fragments.
25. 25. The method of claim 24, wherein the endonuclease is EndoIV or APE1.
26. A reaction mixture comprising the single-stranded DNA circle produced according to claim 6.
27. 20. A reaction mixture comprising the combined synthesized DNA strands from step (vi) of claim 19.
28. 1. A method for preparing multiple nested sets of adapted fragments, comprising: each adapted fragment is a single-stranded nucleic acid comprising a target sequence fragment having a first end and a second end, a 5-prime adapter sequence, a 3-prime adapter sequence, a primer sequence, and a barcode sequence; in each nested set of adapted fragments, the target sequence fragments have identical nucleotide sequences at a first end and differ from one another by truncation at a second end such that each nested set of adapted fragments comprises a plurality of target sequence fragments having different lengths; the first end is closer to the barcode sequence than the second end; The method comprises: (a) providing a population of single-stranded DNA concatemers in a reaction, each concatemer comprises a plurality of identical monomers, each monomer comprising a complement of a target sequence, a complement of a barcode sequence that identifies the concatemer, and a primer binding sequence shared by the population of single-stranded concatemers; the primer binding sequence comprises a sequence that is complementary to the primer sequence; the complement of the primer binding sequence and the barcode sequence are both 3-prime to the complement of the target sequence; (b) annealing a primer comprising a primer sequence to the primer binding sequences of the plurality of monomers of each of the plurality of concatemers; (c) extending at least a portion of the primers hybridized to the primer binding sequences with a DNA polymerase having 5' to 3' exonuclease activity and no strand displacement activity, said extension producing a plurality of extended primers, each extended primer comprising a target sequence fragment having a barcode sequence and a primer sequence; The extended primer is hybridized to the concatemer, the extended primers are separated by an interval; and (d) contacting the plurality of extended primers with 5-prime adapters comprising 5-prime adapter sequences, 3-prime adapters comprising 3-prime adapter sequences, a DNA ligase, and an exonuclease having single-stranded DNA exonuclease activity under conditions where the exonuclease degrades a portion of the target sequence fragment of the extended primers to generate shortened extended primers, 5-prime adapters linked to the 5' ends of the shortened extended primers, and 3-prime adapters linked to the 3' ends of the shortened extended primers; thereby generating a group of multiple nested sets of adapted fragments.
29. 29. The method of claim 28, wherein the population of single-stranded DNA concatemers is generated by rolling circle replication of circular templates, each of the circular templates comprising a target sequence, a barcode sequence, and a primer sequence.
30. 29. The method of claim 28, wherein the 5-prime adapter is an L-adapter and the 3-prime adapter is a branched adapter.
31. 29. The method of claim 28, further comprising adding a nuclease to extend the gap formed in step (c), wherein the nuclease has single-stranded exonuclease activity.
32. 32. The method of claim 31, wherein at least some of the primers are RNA primers, the nuclease is RNAse H, and the RNAse H digests the RNA primers, thereby extending the spacing.
33. 29. The method of claim 28, the primer binding sequence is located 3-prime to the complementary strand of the barcode sequence of step (a); the exonuclease has 3' to 5' exonuclease activity, and A method wherein the barcode sequence of each of the set of adapted fragments is located 5-prime relative to the target sequence fragment.
34. 29. The method of claim 28, the primer binding sequence is located 5-prime to the complementary strand of the barcode sequence of step (a); the exonuclease has 5' to 3' exonuclease activity, and The method, wherein the barcode sequence is 3-prime to the target sequence fragment of each of the adapted fragments.
35. 10. The method of any preceding claim, wherein both the 5-prime adapter and the 3-prime adapter are present in solution.
36. 36. The method of claim 35, wherein the reaction does not involve a solid support.
37. 10. The method of any preceding claim, wherein the target sequence has a length between 500 bases and 50 kilobases.
38. the branched adapter comprises a double-stranded blunt end comprising a 5'-end of one strand and a 3'-end of a complementary strand; 31. The method of claim 30, wherein the 5' end of the strand at the blunt end of the double strand is linked to the 3' end of at least one of the extended primers via a branched ligation.
39. 31. The method of claim 30, wherein the L-adapter is 3-prime and comprises 1 to 10 degenerate bases.
40. 1. A method for preparing multiple nested sets of adapted fragments, comprising: each adapted fragment is a single-stranded nucleic acid comprising a target sequence fragment having a first end and a second end, a 5-prime adaptor sequence, a 3-prime adaptor sequence, a primer binding sequence, and a complement of a barcode sequence; in each nested set of adapted fragments, the target sequence fragments have identical nucleotide sequences at a first end and differ from each other at a second end such that each nested set of adapted fragments comprises a plurality of target sequence fragments having different lengths; the first end is closer to the barcode sequence than the second end; The method comprises: (a) providing barcoded fragments comprising a barcode sequence, a target sequence, and a primer binding sequence, wherein the barcoded fragments are immobilized at one end on beads; (b) annealing a primer comprising a 5-prime adapter sequence to the primer binding sequence of the barcoded fragment, a 5-prime adapter sequence comprising i) the complement of the barcode sequence, and ii) a primer sequence that is complementary to the primer binding sequence of the barcoded fragment; (c) extending the primer to generate an extended primer comprising the complement of the target sequence fragment and the barcode sequence; (d) contacting the extended primer with a branched adapter comprising a 3-prime adapter sequence to generate an adapted fragment; (e) separating the adapted fragments from the barcoded fragments that remain immobilized on the beads; (f) repeating steps (b) through (e) for one or more cycles under controlled extension conditions to generate one or more adapted fragments; the adapted fragments generated from step (e) and the adapted fragments generated from step (f) constitute a nested set of adapted fragments; wherein the adapted fragments of each nested set comprise target sequence fragments having different lengths.
41. 41. The method of claim 40, extending the primer with uracil in one or more extension cycles under extension control conditions to generate an extended primer, thereby generating an adapted fragment incorporating uracil at the 5-prime portion of the target sequence fragment; (g) contacting the adapted fragment with an enzyme that removes incorporated uracil, thereby creating at least one space flanked by an exposed 3-prime end and an exposed 5-prime end of the adapted fragment; (h) ligating an internal branched adaptor to the exposed 3-prime end in at least one interval and ligating an L-adaptor to the exposed 5-prime end in said interval; (i) joining the internal branched adaptor ligated to the exposed 3-prime end in step (h) with the L-adaptor ligated to the exposed 5-prime end, thereby creating a shortened adapted fragment; The method thereby generates a set of shortened, adapted fragments comprising shortened target sequence fragments having sequences corresponding to different regions of the target sequence, the different regions overlapping.
42. 42. The method of claim 41 , the step of ligating the internal branched adapter and the L-adapter comprises contacting the internal branched adapter and the L-adapter with a splint oligonucleotide; the splint oligonucleotide comprises a 5-prime portion that is complementary to the sequence of the internal branch adapter and a 3-prime portion that is complementary to the L-adapter; A method whereby the splint oligonucleotide hybridizes to the internal branch adapter via the 5-prime portion and the splint oligonucleotide hybridizes to the L-adapter via the 3-prime portion, thereby ligating the internal branch adapter and the L-adapter.
43. 1. A method for preparing multiple sets of adapted fragments, comprising: each adapted fragment is a single-stranded nucleic acid comprising a target sequence fragment having a first end and a second end, a 5-prime adapter sequence, a 3-prime adapter sequence, a primer binding sequence, and a complement of a barcode sequence; The method comprises: (a) providing barcoded fragments comprising a barcode sequence, a target sequence, and a primer binding sequence, wherein the barcoded fragments are immobilized at one end on beads; (b) annealing a primer comprising a 5-prime adapter sequence to the primer binding sequence of the barcoded fragment, a 5-prime adapter sequence comprising i) the complement of the barcode sequence, and ii) a primer sequence that is complementary to the primer binding sequence of the barcoded fragment; (c) extending the primer to generate an extended primer comprising the complement of the target sequence fragment and the barcode sequence; (d) contacting the extended primer with a first branched adapter comprising a 3-prime portion that includes a degenerate sequence region, thereby forming a first extension product that includes a degenerate sequence region at the 3-prime portion; 3-prime portion hybridizing to the barcoded fragment via the degenerate sequence region; (e) extending the 3-prime portion of the first extension product to produce a second extension product; and (f) contacting the second extension product with a second branched adapter to generate an adapted fragment.
44. 44. The method of claim 43, further comprising the step of (g) denaturing to separate the adapted fragments from the barcoded fragments.
45. 45. The method of claim 44, further comprising repeating steps (b) through (g) for one or more cycles under controlled extension conditions to generate one or more adapted fragments.
46. A DNA complex comprising a plurality of fragments hybridized to one or more monomers of DNA concatemers, The fragments are separated by intervals, each of the plurality of fragments comprises a barcode sequence and a target sequence fragment having a first end and a second end; A DNA complex in which the target sequence fragments of the plurality of fragments have the same nucleotide sequence at a first end and differ from one another by cleavage at a second end such that the target sequence fragments of the plurality of fragments have different lengths.
47. 47. The DNA complex of claim 46, wherein each of the plurality of fragments is ligated at its 5-prime end to an L-adaptor and at its 3-prime end to a branched adaptor.
48. A DNA complex comprising: (a) a barcoded fragment immobilized on a solid support, the barcode fragment comprising a barcode sequence and a target sequence; (b) a polynucleotide hybridized to the barcoded fragment, the polynucleotide comprises a 5-prime portion comprising the complement of the barcode sequence, a 3-prime portion comprising a target sequence fragment; A DNA complex comprising a polynucleotide in which a 5-prime portion and a 3-prime portion anneal to a barcoded fragment, leaving a central portion that does not anneal to the barcoded fragment, thereby forming a bubble.
49. 49. The plurality of DNA complexes according to any one of claims 46 to 48, wherein the DNA complexes share the same barcode sequence.
50. 1. A composition comprising a nested set of adapted fragments, each comprising a barcode sequence and a target sequence fragment having a first end and a second end, a 5-prime adapter sequence, and a 3-prime adapter sequence; the target sequence fragments have identical nucleotide sequences at a first end and differ from one another by truncation at a second end such that the nested set of adapted fragments comprises target sequence fragments having a plurality of different lengths; A composition wherein the adapted fragments share the same barcode sequence.