Method for creating asymmetric adapter at terminal of polynucleotide template including hairpin loop, and sequencing from the adapter

Asymmetric adapters with hairpin loops at the ends of polynucleotide templates address the issue of Poisson artifacts in sequencing, improving yield and accuracy by ensuring consistent primer and polymerase binding, and enabling in-platform library preparation.

JP2025170232APending Publication Date: 2025-11-18ILLUMINA CAMBRIDGE LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025111971
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-10-25
Filing Date
2025-07-02
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Current sequencing methods using closed-end double-stranded templates require symmetric adapters at both ends, leading to Poisson artifacts due to inconsistent primer and polymerase binding, which affects sequencing accuracy and yield.

Method used

Generate asymmetric adapters with hairpin loops at the ends of polynucleotide templates using methods that include attaching a nucleic acid-based hairpin or dumbbell adapter, extending complementary sequences, and closing the free end to form an asymmetric closed-end template, which can be sequenced using a processive polymerase.

Benefits of technology

This approach eliminates Poisson artifacts, improves sequencing library conversion yield, and allows library preparation steps to be performed directly on the sequencing platform, enhancing sequencing accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025170232000001
    Figure 2025170232000001
  • Figure 2025170232000002
    Figure 2025170232000002
  • Figure 2025170232000003
    Figure 2025170232000003
Patent Text Reader

Abstract

To provide a method for creating two different adapter sequences to the end of a sequence determination template, which is required for a sequencing platform using a closed ended double-stranded template.SOLUTION: A method for creating an asymmetric closed ended double-stranded nucleic acid template from a double-stranded nucleic acid template having free 5' and 3' terminals by using a hairpin or a dumbbell adapter, and sequencing from the template.SELECTED DRAWING: None
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] (CROSS-REFERENCE TO RELATED APPLICATIONS) This application claims priority under 35 U.S.C. §119(e) from U.S. Provisional Application No. 62 / 926,360, filed October 25, 2019, the disclosure of which is incorporated herein by reference for all purposes.

[0002] (Summary of the Invention) The present disclosure generally relates to methods for generating asymmetric adapters at the ends of polynucleotides that contain hairpin loops and for sequencing therefrom. Including an array list by reference

[0003] This application is accompanied by a Sequence Listing entitled "Sequence-Listing_ST25.txt," created on October 19, 2020, and containing 784 bytes of data machine-formatted for an IBM-PC, MS-Windows operating system. The Sequence Listing is hereby incorporated by reference in its entirety for all purposes. [Background technology]

[0004] Sequencing platforms that use closed-end double-stranded templates ideally require two distinct adapter sequences at the end of the sequencing template. This asymmetry of the ends allows a single sequencing primer and polymerase to bind to only one end of the closed-end template, generating only one sequence read per template molecule. Currently, this is not possible with standard library preparation, where the same sequence is added to both ends of the template, thus introducing Poisson artifacts into sequencing: some templates require no primer / polymerase binding, some require one primer / polymerase binding, and some require one primer / polymerase binding (one at each end). Summary of the Invention

[0005] The present disclosure provides methods for generating asymmetric adapters at the ends of polynucleotide templates containing hairpin loops. Enabling asymmetric adapters can avoid Poisson artifacts inherent in standard methods, thereby improving the conversion yield of sequencing libraries. Furthermore, the methods disclosed herein simplify the library preparation workflow, allowing some steps of library preparation to be performed on the sequencing platform.

[0006] In one embodiment, the disclosure provides a method for generating an asymmetric closed-end double-stranded nucleic acid template from a double-stranded nucleic acid template having free 5' and 3' ends, the method comprising: (A) attaching a first nucleic acid-based hairpin or dumbbell adapter to the 3' end of the double-stranded nucleic acid template comprising the free 5' and 3' ends; and (B) extending a sequence complementary to the double-stranded nucleic acid template from each 3' end of the nucleic acid hairpin using a processive polymerase to generate two long hairpin double-stranded templates, wherein one end of the double-stranded template is attached to a 3' end of the double-stranded nucleic acid template. (C) closing the free end of each long hairpin double-stranded template by (i) ligating a second nucleic acid-based hairpin or dumbbell adapter to the free end of the double-stranded template, or (ii) using TelN protelomerase to close the free end of the double-stranded template designed to contain a TelN recognition sequence, to form an asymmetric closed-end double-stranded nucleic acid template. In further embodiments of the embodiments provided herein, the double-stranded nucleic acid template is a double-stranded DNA template. In other embodiments of the embodiments provided herein, the 5' and 3' ends of the double-stranded nucleic acid template are dephosphorylated and end-repaired. In yet other embodiments of the embodiments provided herein, the double-stranded nucleic acid template has blunt 5' and 3' ends. In other embodiments of the embodiments provided herein, the double-stranded nucleic acid template has an A-tail at its 3' end. In further embodiments of the embodiments provided herein, a first nucleic acid-based hairpin or dumbbell adapter is ligated to the 3' end of the double-stranded nucleic acid template using a ligase. In yet other embodiments of the embodiments provided herein, the ligase is T4 DNA ligase or T3 DNA ligase. In other embodiments of the embodiments provided herein, the first nucleic acid-based hairpin or dumbbell adapter comprises a blunt end or a T-tail end. In yet another embodiment of the embodiments provided herein, the first nucleic acid-based hairpin adapter is a Y-shaped adapter.In certain embodiments provided herein, dimers formed from two first nucleic acid-based hairpin or dumbbell adapters linked to each other are removed by using size selection or size exclusion techniques. In other embodiments provided herein, the processive polymerase is Phi29 polymerase. In yet other embodiments provided herein, the second nucleic acid-based hairpin or dumbbell adapter comprises a blunt end or a T-tail end. In further embodiments provided herein, prior to step (C), the long hairpin duplex template is digested with a restriction enzyme that generates a 5' overhang. In yet other embodiments provided herein, prior to step (C), the long hairpin duplex template is digested with a restriction enzyme that generates a 3' overhang. In other embodiments provided herein, the second nucleic acid-based hairpin or dumbbell adapter comprises an overhang sequence complementary to the overhang sequence of the digested long hairpin duplex template. In other embodiments provided herein, a second nucleic acid-based hairpin or dumbbell adapter is ligated to the free end of the double-stranded template using polynucleotide kinase and ligase. In yet another embodiment of the embodiments provided herein, dimers formed from two linked second nucleic acid-based hairpin or dumbbell adapters are removed by using size selection or size exclusion techniques. In yet another embodiment of the embodiments provided herein, the TelN protelomerase is derived from phage N15, and the TelN protelomerase cleaves the long hairpin double-stranded template at the TelN recognition sequence, leaving a covalently closed end at the cleavage site. In yet another embodiment of the embodiments provided herein, the method disclosed herein further includes (C') using rolling circle replication to generate nanoball complexes containing the polycistronic amplified asymmetric closed-end double-stranded nucleic acid template. In other embodiments of the embodiments presented herein, the methods disclosed herein further comprise the step of (D) sequencing the asymmetric closed-ended double-stranded nucleic acid template or nanoball complex using a sequencing primer and a polymerase.In yet another embodiment of the embodiments provided herein, step (B), step (C)(ii), and step (D) can be combined together as a one-pot reaction. In a further embodiment of the embodiments provided herein, step (B), step (C)(ii), and step (D) are performed in the wells of an automated sequencing platform.

[0007] In certain embodiments, the present disclosure also provides a method for generating an asymmetric double-stranded nucleic acid template from tagged DNAs comprising complementary hairpin loops, the method comprising: (I) generating tagged DNAs comprising complementary hairpin loops at the 5' end of each strand, wherein the hairpin loops comprise a base-paired transposase recognition sequence, with a single-stranded sequence gap between the 5' and 3' ends of the tagged DNA; (II) filling the gap between the 5' and 3' ends of the tagged DNAs using a gap-fill ligation reaction to form closed-end tagged DNAs; and (III) generating a nick in the top strand of each hairpin region of the closed-end tagged DNAs. (IV) using a processive polymerase to extend a sequence complementary to the double-stranded nucleic acid template from each nick to generate two long hairpin double-stranded templates, wherein one end of the double-stranded template comprises a closed hairpin ("hairpin end") and the other end of the duplex comprises a 3'-strand end and a 5'-strand end ("free end"); and (V) closing the free end of each long hairpin double-stranded template by (a) ligating a nucleic acid-based hairpin or dumbbell adapter to the free end of the double-stranded template or (b) using TelN protelomerase to close the free end of the double-stranded template designed to contain a TelN recognition sequence, thereby forming an asymmetric closed-end double-stranded nucleic acid template. In further embodiments of the embodiments presented herein, the transposase recognition sequence is a 19-bp mosaic end sequence. In yet other embodiments of the embodiments presented herein, the gap is 9 base pairs in length. In other embodiments of the embodiments provided herein, the gap-fill ligation reaction comprises Klenow fragment. In yet other embodiments of the embodiments provided herein, the gap-fill ligation reaction comprises T4 DNA polymerase and ampligase. In certain embodiments of the embodiments provided herein, the nick is generated by using a site-specific endonuclease. In further embodiments of the embodiments provided herein, the processive polymerase is Phi29 polymerase.In yet other embodiments of the embodiments provided herein, the nucleic acid-based hairpin or dumbbell adapter comprises a blunt end or a T-tail end. In other embodiments of the embodiments provided herein, prior to step (V), the long hairpin double-stranded template is digested with a restriction enzyme that generates a 5' overhang. In still other embodiments of the embodiments provided herein, prior to step (V), the long hairpin double-stranded template is digested with a restriction enzyme that generates a 3' overhang. In still other embodiments of the embodiments provided herein, the nucleic acid-based hairpin or dumbbell adapter comprises an overhang sequence complementary to the overhang sequence of the digested long hairpin double-stranded template. In still other embodiments of the embodiments provided herein, the nucleic acid-based hairpin or dumbbell adapter is ligated to the free end of the double-stranded template using polynucleotide kinase and ligase. In other embodiments of the embodiments provided herein, dimers formed from two nucleic acid-based hairpin or dumbbell adapters that are bound to each other are removed using size selection or size exclusion techniques. In yet other embodiments of the embodiments presented herein, the TelN protelomerase is derived from phage N15, and the TelN protelomerase cleaves the long hairpin double-stranded template at the TelN recognition sequence, leaving a covalently closed end at the cleavage site. In further embodiments of the embodiments presented herein, the method further comprises (V') using rolling circle replication to generate nanoball complexes containing the polycistronic amplified asymmetric closed-end double-stranded nucleic acid template. In yet other embodiments of the embodiments presented herein, the method further comprises (VI) using a sequencing primer and a polymerase to sequence the asymmetric closed-end double-stranded nucleic acid template or nanoball complex. In other embodiments of the embodiments presented herein, step (V)(ii) and step (VI) can be combined together into a single step. In yet other embodiments of the embodiments presented herein, step (V)(ii) and step (VII) are performed within the wells of an automated sequencing platform.

[0008] In certain embodiments, the present disclosure also provides a method for sequencing tagged DNA comprising complementary hairpin loops, the method comprising: (I) generating tagged DNA comprising complementary hairpin loops at the 5' end of each strand, wherein the hairpin loops comprise a base-paired transposase recognition sequence, with a single-stranded sequence gap between the 5' and 3' ends of the tagged DNA; (II) using a polymerase to form two stretches of tagged DNA comprising a transposase recognition sequence and a complementary transposase recognition sequence at the 5' and 3' ends of each stretch of tagged DNA; (III) separating the stretches of tagged DNA and rehybridizing the transposase recognition sequence and the complementary transposase recognition sequence at each end of the stretch of tagged DNA to form complementary hairpin loops; and (IV) using a sequencing polymerase to sequence the stretches of tagged DNA comprising complementary hairpin loops. In other embodiments of the embodiments provided herein, step (IV) is performed in the wells of an automated sequencing platform. In yet other embodiments of the embodiments provided herein, prior to the sequencing of step (IV), (III') nanoball complexes comprising polycistronic amplified extension strands of tagged DNA are generated. [Brief explanation of the drawings]

[0009] [Figure 1]

[0023] Figure 1 is a schematic diagram showing a method for generating template end asymmetry by ligating the 5' end of a hairpin to the 3' end of a double-stranded polynucleotide template and using a processive polymerase to extend the 3' end of the hairpin, thereby generating a hairpin end and a free end. An additional hairpin or dumbbell can be ligated to the free end of the asymmetric template to provide a closed-end template containing an asymmetric region. Small adapter dimers can be removed using size selection or size exclusion techniques.

[0010] [Figure 2]FIG. 1 is a schematic diagram showing a method for generating template end asymmetry by ligating the 5′ of a Y-shaped adaptor to the 3′ end of a double-stranded polynucleotide template and using a processive polymerase to extend the 3′ end of the Y-shaped adaptor, thereby generating a hairpin end and a free end.

[0011] [Figure 3] FIG. 1 is a schematic diagram showing blunt-end ligation of dumbbell or simple hairpin adapters to the "free ends" of an asymmetric template.

[0012] [Figure 4] FIG. 1 is a schematic diagram showing blunt-end ligation of dumbbell or simple hairpin adapters with complementary overhang sequences to a restriction-digested asymmetric template with a 5′ or 3′ overhang.

[0013] [Figure 5] FIG. 1 is a schematic diagram showing the inclusion of a TelN protelomerase site into an asymmetric template sequence and subsequent circularization (SEQ ID NOs: 1 and 2).

[0014] [Figure 6] FIG. 1 is a schematic diagram showing how a closed-end template containing an asymmetric region can be sequenced using an automated sequencing platform, and thus at least some steps can be performed in situ in the cell of the sequencing platform, as shown.

[0015] [Figure 7] Schematic diagram showing how to generate asymmetric ends from tagged DNA by creating nicks in the top strand using a gap-filling reaction. Nicks in the top strand can be generated using several approaches, including the use of site-specific endonucleases, user digestion, RNA bases and ribonucleases, diols, etc.

[0016] [Figure 8]FIG. 1 is a schematic showing the use of a polymerase Flash-based method to generate 3′-terminal hairpin loops from tagged DNA.

[0017] [Figure 9] FIG. 1 is a schematic diagram showing the method for generating and sequencing instrument templates from tagged DNA. DETAILED DESCRIPTION OF THE INVENTION

[0018] As used herein, the terms "includes," "including," "includes," "including," "contains," "containing," "have," "having," and any variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, product defined by a process, or composition of matter that includes or contains an element or list of elements may include not only those elements, but also other elements not expressly listed in or inherent in such process, method, product defined by a process, or composition of matter. Similarly, "comprise," "comprises," "comprising," "include," "includes," and "including" are interchangeable and are not intended to be limiting.

[0019] Where the descriptions of various embodiments use the term "comprising," those skilled in the art will further understand that in some specific instances, the embodiments may alternatively be described using the phrases "consisting essentially of" or "consisting of."

[0020] As used herein, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Thus, for example, reference to a "protein" includes a mixture of two or more proteins, and so on.

[0021] Also, unless stated otherwise, the use of "or" means "and / or."

[0022] Other than in the operating examples, or unless otherwise specified, all numerical values ​​expressing quantities of ingredients or reaction conditions used herein should be understood to be modified in all instances by the term "about." When used to describe embodiments of the present disclosure, the term "about" refers to percentages and means ±1%, ±2%, ±3%, ±4%, or ±5%. As used herein, the term "about" can mean within an acceptable error range of a particular value as determined by one of ordinary skill in the art, which may depend in part on how the value is measured or determined, such as limitations of the measurement system. Alternatively, "about" can mean a range of plus or minus 20%, plus or minus 10%, plus or minus 5%, or plus or minus 1% of a given value. Alternatively, particularly with respect to biological systems or processes, the term can mean within 10-fold, within 5-fold, or within 2-fold. When specific values ​​are described in the application and claims, unless otherwise specified, the term "about" can be assumed to mean within an acceptable error range of the particular value. Also, when ranges and / or sub-ranges of values ​​are provided, the ranges and / or sub-ranges may include the endpoints of the ranges and / or sub-ranges. In some examples, variations may include amounts or concentrations of 20%, 10%, 5%, 1%, 0.5%, or even 0.1% of the specified amount.

[0023] For the recitation of numerical ranges herein, each intervening number is expressly contemplated to the same degree of precision. For example, in the range 6 to 9, the numbers 7 and 8 are contemplated in addition to 6 and 9, and in the range 6.0 to 7.0, the numbers 6.0, 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.8, 6.9, and 7.0 are expressly contemplated.

[0024] All publications mentioned herein are incorporated by reference in their entirety for the purpose of describing and disclosing methodologies that might be used in connection with the description herein. Furthermore, with respect to any terms presented in one or more publications that are similar or identical to terms explicitly defined in this disclosure, the definition of the term expressly provided in this disclosure shall control in all respects.

[0025] It is to be understood that this disclosure is not limited to the particular methodology, protocols, and reagents, etc., described herein and as such may vary. The terminology used herein is for the purpose of describing particular embodiments or aspects only and is not intended to limit the scope of the present disclosure.

[0026] As used herein, the term "adaptor" refers to a single- or double-stranded nucleic acid molecule that can be ligated to the end of another nucleic acid. For purposes of this disclosure, "adaptor" includes hairpins or dumbbell loops, unless otherwise specified. In certain embodiments, adapters of the present disclosure are double-stranded nucleic acids (e.g., oligonucleotides) that include single-stranded nucleotide overhangs at the 5' and / or 3' ends. In further embodiments, the single-stranded overhangs are 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or 20 nucleotides.

[0027] As used herein, the term "complementary" when used with respect to a polynucleotide is intended to mean a polynucleotide comprising a nucleotide sequence that can selectively anneal to an identified region of a target polynucleotide under specific conditions. As used herein, the term "substantially complementary" and grammatical equivalents are intended to mean a polynucleotide comprising a nucleotide sequence that can specifically anneal to an identified region of a target polynucleotide under specific conditions. Annealing refers to the nucleotide base-pairing interaction between one nucleic acid and another, resulting in the formation of a duplex, triplex, or other higher-order structure. Primary interactions are typically nucleotide base-specific, e.g., A:T, A:U, and G:C, via Watson-Crick and Hoogsteen hydrogen bonding. In certain embodiments, base stacking and hydrophobic interactions may also contribute to the stability of the duplex. Conditions under which a polynucleotide will anneal to a complementary or substantially complementary region of a target nucleic acid are well known in the art, as described, for example, in Nucleic Acid Hybridization, A Practical Approach, Hames and Higgins, eds., IRL Press, Washington, DC (1985) and Wetmur and Davidson, Mol. Biol. 31:349 (1968). Annealing conditions will depend on the particular application and can be routinely determined by one of ordinary skill in the art without undue experimentation.

[0028] As used herein, the term "dNTP" refers to deoxynucleoside triphosphate. NTP refers to ribonucleotide triphosphate. Purine bases (Pu) include adenine (A), guanine (G), and their derivatives and analogs. Pyrimidine bases (Py) include cytosine (C), thymine (T), uracil (U), and their derivatives and analogs. Examples of such derivatives or analogs include, but are not limited to, those modified with reporter groups, biotinylated, amine-modified, radiolabeled, alkylated, and the like, including phosphorothioates, phosphites, and derivatives modified at ring atoms. Reporter groups can be fluorescent groups such as fluorescein, chemiluminescent groups such as luminol, terbium chelators such as N-(hydroxyethyl)ethylenediaminetriacetic acid, which allow detection by delayed fluorescence, and the like.

[0029] As used herein, the term "hybridization" refers to the process by which two single-stranded polynucleotides non-covalently bind to form a stable double-stranded polynucleotide. The resulting double-stranded polynucleotide is a "hybrid" or "double-stranded." Hybridization conditions typically include a salt concentration of less than about 1 M, more usually less than about 500 mM, and can be less than about 200 mM. The hybridization buffer includes a buffered salt solution such as 5% SSPE or other such buffers known in the art. Hybridization temperatures can be as low as 5°C, but are typically greater than 22°C, more typically greater than about 30°C, and typically greater than 37°C. Hybridization is usually performed under stringent conditions, i.e., conditions under which a probe hybridizes to its target subsequence but not to other non-complementary sequences. Stringent conditions are sequence-dependent, vary in different circumstances, and can be routinely determined by one of skill in the art.

[0030] As used herein, the term "label" refers to a process by which a component, e.g., an adapter, is modified, e.g., attached to another molecule, to facilitate separation of the component and its associated elements.

[0031] As used herein, the terms "ligation," "ligating," and their grammatical equivalents are intended to mean forming a covalent bond or linkage between the ends of two or more nucleic acids, e.g., oligonucleotides and / or polynucleotides, typically in a template-driven reaction. The nature of the bond or linkage can vary widely, and ligation can be performed enzymatically or chemically. As used herein, ligation is typically performed enzymatically, forming a phosphodiester bond between the 5' carbon terminal nucleotide of one oligonucleotide and the 3' carbon of another nucleotide. Template-driven ligation reactions are described in references such as U.S. Pat. No. 4,883,750, U.S. Pat. No. 5,476,930, U.S. Pat. No. 5,593,826, and U.S. Pat. No. 5,871,921, which are incorporated herein by reference in their entireties. The term "ligation" also encompasses the non-enzymatic formation of phosphodiester bonds, as well as the formation of non-phosphodiester covalent bonds between the ends of oligonucleotides, such as phosphorothioate bonds and disulfide bonds.

[0032] As used herein, the term "nucleic acid" refers to 2'-deoxyribonucleotides (DNA) and ribonucleotides (RNA) linked by internucleotide phosphodiester bonds or internucleotide analogs, and associated counterions, e.g., H + , N.H. 4+ , trialkylammonium, tetraalkylammonium, Mg 2+ , Na +Nucleic acids refer to single- and double-stranded polymers of nucleotide monomers, including, but not limited to, ribonucleotides, ribonucleotides, and chimeric mixtures thereof. Nucleic acids can be polynucleotides or oligonucleotides. Nucleic acids can be composed entirely of deoxyribonucleotides, entirely of ribonucleotides, or chimeric mixtures thereof. Nucleic acid monomer units can include any of the nucleotides described herein, including, but not limited to, naturally occurring nucleotides and nucleotide analogs. Nucleic acids typically range in size from a few monomeric units, e.g., 5-40, to several thousand monomeric nucleotide units. Nucleic acids include, but are not limited to, genomic DNA, eDNA, hnRNA, mRNA, rRNA, tRNA, fragmented nucleic acids, nucleic acids obtained from intracellular organelles such as mitochondria and chloroplasts, and nucleic acids obtained from microorganisms or DNA or RNA viruses that may be present on or in a biological sample.

[0033] As used herein, the term "nucleotide analog" refers to synthetic analogs having modified nucleotide base moieties, modified pentose moieties, and / or modified phosphate moieties, and, in the case of polynucleotides, modified internucleotide linkages as generally described elsewhere (e.g., Scheit, Nucleotide Analogs, John Wiley, New York, 1980; Englisch, Angew. Chem. Int. Ed. Engl. 30:613-29, 1991; Agarwal, Protocols for Polynucleotides and Analogs, Humana Press, 1994; and S. Verma and F. Eckstein, Ann. Rev. Biochem. 67:99-134, 1998). Exemplary phosphate analogs include phosphorothioates, phosphorodithioates, phosphoroselenoates, phosphorodiselenoates, phosphoroanilothioates, phosphoranilidates, phosphoramidates, boranophosphates, and the associated counterions, e.g., H, when such counterions are present. + , NH4 + , Na +Exemplary modified nucleotide base moieties include, but are not limited to, 5-methylcytosine (5mC); C-5-propynyl analogs, including, but not limited to, C-5 propynyl-C and C-5 propynyl-U; 2,6-diaminopurine (also known as 2-aminoadenine or 2-amino-dA); hypoxanthine, pseudouridine, 2-thiopyrimidine, isocytosine (isoC), 5-methylisoC, and isoguanine (isoG; see, e.g., U.S. Pat. No. 5,432,272). Exemplary modified pentose moieties include, but are not limited to, locked nucleic acid (LNA) analogs, including Bz-A-LNA, 5-Me-Bz-C-LNA, dmf-G-LNA, and T-LNA (see, e.g., The Glen Report, 16(2):5, 2003; Koshkin et al., Tetrahedron 54:3607-30, 1998), and 2'- or 3'-modifications in which the 2'- or 3'-position is hydrogen, hydroxy, alkoxy (e.g., methoxy, ethoxy, allyloxy, isopropoxy, butoxy, isobutoxy, and phenoxy), azido, amino, alkylamino, fluoro, chloro, or bromo. Modified internucleotide linkages include phosphate analogs, analogs with achiral and uncharged intersubunit linkages (e.g., Sterchak, EP et al., Organic Chem., 52:4202, 1987), and uncharged morpholino-based polymers with achiral intersubunit linkages (see, e.g., U.S. Pat. No. 5,034,506). Some internucleotide linkage analogs include morpholidate, acetal, and polyamide-linked heterocycles.

[0034] As used herein, the terms "variant" and "derivative" in the context of "polynucleotides" refer to a polynucleotide comprising a nucleotide sequence of the polynucleotide or a fragment of the polynucleotide that has been altered by the introduction of nucleotide substitutions, deletions, or additions. A variant or derivative of a polynucleotide can be a fusion polynucleotide that contains a portion of the nucleotide sequence of the polynucleotide. As used herein, the terms "variant" or "derivative" also refer to a polynucleotide or fragment thereof that has been chemically modified, for example, by the covalent attachment of any type of molecule to the polynucleotide. For example, but not limited to, a polynucleotide or fragment thereof can be chemically modified by, for example, acetylation, phosphorylation, methylation, etc. A variant or derivative is modified in a manner that differs from the naturally occurring or starting nucleotide or polynucleotide, either in the type or location of the attached molecule. A variant or derivative further includes deletion of one or more chemical groups naturally occurring on the nucleotide or polynucleotide. A variant or derivative of a polynucleotide or a fragment of a polynucleotide may be chemically modified by chemical modification using techniques known to those of skill in the art, including, but not limited to, specific chemical cleavage, acetylation, formulation, etc. Additionally, a variant or derivative of a polynucleotide or a fragment of a polynucleotide may contain one or more dNTPs or nucleotide analogs. A polynucleotide variant or derivative may have a similar or identical function as the polynucleotides or fragments of polynucleotides described herein. A polynucleotide variant or derivative may have additional or different functions compared to the polynucleotides or fragments of polynucleotides described herein.

[0035] As used herein, the terms "tagmentation," "tagment," or "tagmenting" refer to the conversion of nucleic acids, e.g., DNA, into adapter-modified templates in solution ready for clustering and sequencing using transposase-mediated fragmentation and tagging. This process often involves modification of the nucleic acid with a transposon complex containing a transposase enzyme complexed with adapters containing transposon end sequences. Tagging simultaneously fragments the nucleic acid and ligates adapters to the 5' ends of both strands of the double-stranded fragments. Following a purification step to remove the transposase enzyme, additional sequences are added to the ends of the adapted fragments by PCR.

[0036] "Transposase" refers to an enzyme that can form a functional complex with a transposon end-containing composition (e.g., a transposon, a transposon end, a transposon end composition) and catalyze the insertion or transposition of the transposon end-containing composition into a double-stranded target nucleic acid with which it is incubated, e.g., in an in vitro transposition reaction. Transposases provided herein can also include integrases from retrotransposons and retroviruses. Transposases, transposomes, and transposome complexes are generally known to those of skill in the art, as exemplified by the disclosure of U.S. Patent Application Publication No. 2010 / 0120098, the entire contents of which are incorporated herein by reference. While many embodiments described herein refer to Tn5 transposase and / or hyperactive Tn5 transposase, it is understood that any transposition system capable of inserting transposon ends with sufficient efficiency to 5' tag and fragment target nucleic acids for the intended purpose can be used in the present invention. In certain embodiments, a preferred transposition system can insert transposon ends in a random or near-random manner to 5' tag and fragment target nucleic acids.

[0037] As used herein, the term "transposition reaction" refers to a reaction in which one or more transposons are inserted into a target nucleic acid, for example, at random or near-random sites. The essential components of a transposition reaction are a transposase and a DNA oligonucleotide representing the nucleotide sequence of the transposon, including the transferred transposon sequence and its complement (the non-transferred transposon end sequence), as well as other components necessary to form a functional transposition or transposon complex. The DNA oligonucleotide may further comprise additional sequences (e.g., adapter or primer sequences) as needed or desired. In some embodiments, the methods provided herein are exemplified using transposition complexes formed by hyperactive Tn5 transposase and Tn5-type transposon ends (Goryshin and Reznikoff, 1998, J. Biol. Chem., 273:7367) or by MuA transposase and Mu transposon ends containing R1 and R2 end sequences (Mizuuchi, 1983, Cell, 35:785; Savilahti et al., 1995, EMBO J., 14:4893). However, any transposition system capable of inserting transposon ends in a random or near-random manner with sufficient efficiency to 5'-tag and fragment target DNA for the intended purpose can be used in the present invention.Examples of transposition systems known in the art that can be used in the methods of the present invention include Staphylococcus aureus Tn552 (Colegio et al., 2001, J. Bacterid., 183:2384-8; Kirby et al., 2002, Mol. Microbiol., 43:173-86), TyI (Devine and Boeke, 1994, Nucleic Acids Res., 22:3765-72 and International Patent Application No. 95 / 23875), transposon Tn7 (Craig, 1996, Science. 271:1512; Craig, 1996, Review in: Curr. Top Microbiol. Immunol., 204:27-48), TnIO and ISLO (Kleckner et al., 1996, Curr. Top Microbiol. Immunol., 204:27-48). Immunol, 204:49-82), mariner transposase (Lampe et al., 1996, EMBO J., 15:5470-9), Tci (Plasterk, 1996, Curr Top Microbiol Immunol, 204:125-43), P elements (Gloor, 2004, Methods Mol Biol, 260:97-114), TnJ (Ichikawa and Ohtsubo, 1990, J Biol Chem. 265:18829-32), bacterial insertion sequences (Ohtsubo and Sekine, 1996, Curr. Top. Microbiol. Immunol. 204:1-26), retroviruses (Brown et al., 1989, Proc Natl Acad Sci USA, 86:2525-9), and yeast retrotransposons (Boeke and Corces, 1989, Annu Rev Microbiol. 43:403-34. Methods for inserting transposon ends into target sequences can be performed in vitro using any suitable transposon system for which a suitable in vitro transposition system is available or which can be developed based on knowledge in the art.In general, a suitable in vitro transposition system for use in the methods provided herein requires, at a minimum, a transposase enzyme of sufficient purity, sufficient concentration, and sufficient in vitro transposition activity, and transposase ends that form a functional complex with the respective transposase that can catalyze a transposition reaction. Suitable transposase transposon end sequences that can be used in the present invention include, but are not limited to, wild-type, derivative, or mutant transposon end sequences that form a complex with a transposase selected from a wild-type, derivative, or mutant of the transposase.

[0038] As used herein, the term "transposome complex" refers to a transposase enzyme that non-covalently binds to double-stranded nucleic acid. For example, the complex can be a transposase enzyme pre-incubated with double-stranded transposon DNA under conditions that support non-covalent complex formation. The double-stranded transposon DNA can include, but is not limited to, Tn5 DNA, a portion of Tn5 DNA, a transposon end composition, a mixture of transposon end compositions, or other double-stranded DNA that can interact with a transposase, such as a hyperactive Tn5 transposase.

[0039] The term "transposon end" (TE) refers to a double-stranded nucleic acid, e.g., a double-stranded DNA that exhibits only the nucleotide sequences ("transposon end sequences") necessary to form a complex with a transposase or integrase enzyme that functions in an in vitro transposition reaction. In some embodiments, the transposon end is capable of forming a functional complex with a transposase in a transposition reaction. As non-limiting examples, as described in the disclosure of U.S. Patent Application Publication No. 2010 / 0120098, which is incorporated herein by reference in its entirety, the transposon end can include a 19-bp outer end ("OE") transposon end, an inner end ("IE") transposon end, or a "mosaic end" ("ME") transposon end recognized by wild-type or mutant Tn5 transposase, or an R1 and R2 transposon end. The transposon end can comprise any nucleic acid or nucleic acid analog suitable for forming a functional complex with a transposase or integrase enzyme in an in vitro transposition reaction. For example, transposon ends may comprise DNA, RNA, modified bases, unnatural bases, modified backbones, and may contain nicks in the single or double strands. Although the term "DNA" is sometimes used in this disclosure in reference to transposon end compositions, it should be understood that any suitable nucleic acid or nucleic acid analog may be utilized in the transposon ends.

[0040] Sequencing platforms that use closed-end double-stranded sequencing templates typically require two different adapter sequences at the ends of the sequencing template. This asymmetry allows a single sequencing polymer and polymerase to bind to only one end of the closed-end sequencing template, generating only one sequence read per template molecule. However, such asymmetry is not currently used by such sequencing platforms for library preparation. Instead, these platforms have symmetric ends, allowing sequencing primers to bind to either end of the sequencing template, introducing Poisson artifacts into the sequencing. In other words, some sequencing templates have no primer / polymerase attached, some have one primer attached, and some have two primers attached to both ends of the sequencing template. Using closed-end sequencing templates with asymmetric ends can avoid these Poisson artifacts. Thus, the disclosed methods offer a significant improvement over the state of the art by providing for the generation of closed-ended double-stranded templates with asymmetric ends, thereby significantly increasing the conversion yield of sequencing libraries. Furthermore, the disclosed methods provide steps that can be performed in situ within the wells of a sequencing platform, allowing library preparation to be performed on the sequencing platform.

[0041] As shown in Figure 1, a nucleic acid-based hairpin or dumbbell adapter is attached to a double-stranded nucleic acid template with free 5' and 3' ends. The double-stranded nucleic acid template can be dsDNA, dsRNA, or a chimeric mixture of DNA and RNA that forms a double-stranded molecule. The nucleotides of the nucleic acid can be composed of naturally occurring nucleotides, such as A, G, C, T, and U, linked together via phosphodiester bonds. Alternatively, one or more nucleotides of the nucleic acid can be modified in some way (e.g., with a nucleotide analog). Examples of nucleotide analogs include nucleotides or ribonucleotides modified at the 2' position of the ribose or deoxyribose to have a -methoxy-ethyl group, an -O-methyl group, a fluoro group, 2-aminopurine, 5-bromoduline, deoxyuridine, 2,6-diaminopurine, deoxyinosine, hydroxymethyl dC, 5-methyl dC, 5-nitroindole, 5-hydroxybutyne-2'-deoxyuridine, and 8-aza-7-deazaguanosine. Furthermore, nucleotides can be linked together by phosphorothioate bonds in addition to phosphodiester bonds. The free 5' and 3' ends of the double-stranded nucleic acid template can be blunt or can have one or more mismatched base overhangs (e.g., 3' overhangs or 5' overhangs). In certain embodiments, the free 5' and / or 3' ends of the double-stranded nucleic acid template comprise one or more 3' overhangs of adenine bases. In further embodiments, the 5' and / or 3' ends of the double-stranded nucleic acid template are dephosphorylated and end-repaired.

[0042] Nucleic acid-based hairpin or dumbbell adapters contain a sequence that loops back on itself so that the ends can bind to a double-stranded nucleic acid template. Nucleic acid-based hairpin or dumbbell adapters can have any sequence or can be designed to contain a target sequence. Examples of target sequences include, but are not limited to, a consensus primer sequence, a universal sequencing primer sequence, a barcode sequence, a restriction enzyme sequence, a TelN recognition sequence, or any combination of the foregoing. Nucleic acid-based hairpin or dumbbell adapters can be attached to a double-stranded nucleic acid template using a ligase. Examples of ligases include, but are not limited to, DNA ligases such as T4 DNA ligase, E. coli DNA ligase, Ampligase DNA ligase, T3 DNA ligase, T7 DNA ligase, and Taq DNA ligase, and RNA ligases such as T4 RNA ligase 1, T4 RNA ligase 2, RtcB ligase, and M. thermoautotrophicum ligase. A phosphorylation or dephosphorylation step may be performed, if necessary, before attaching the nucleic acid-based hairpin or dumbbell adapter to the double-stranded nucleic acid template or tagged DNA. For purposes of this disclosure, any of the aforementioned enzymes may be further modified by genetic engineering techniques to enhance one or more functionalities of the enzyme, such as processivity, thermostability, fidelity, etc. The nucleic acid-based hairpin or dumbbell adapter may further comprise a label to enable detection and / or purification of the adapter-containing sequence, such as the adapter attached to the nucleic acid template. The nucleic acid-based hairpin or dumbbell adapter comprises a sequence that loops back on itself so that its ends can bind to the double-stranded template sequence or tagged DNA. For purposes of this disclosure, the nucleic acid-based hairpin or dumbbell adapter may be a Y-shaped adapter, comprising a portion of its sequence that forms a hairpin loop. As shown in Figure 2, a Y-shaped adapter can be used to generate an asymmetric, closed-end double-stranded nucleic acid template from a double-stranded nucleic acid template with free 5' and 3' ends. Additionally, dimers formed by ligating two nucleic acid-based hairpin or dumbbell adapters together can be removed using size selection or size exclusion techniques if such dimers are indicative of downstream reactions.

[0043] The present disclosure further provides extending a sequence complementary to a double-stranded nucleic acid template or tagged DNA from each 3' end of the nucleic acid hairpin using a processive polymerase and free nucleotides (e.g., dNTPs or NTPs) to generate two long hairpin double-stranded templates, where one end of the double-stranded template comprises a closed hairpin ("hairpin end") and the other end of the double-stranded template comprises a free 3' strand end and a free 5' strand end ("free end"). Examples of processive polymerases include, but are not limited to, Phi29 polymerase, SP6 RNA polymerase, and T7 RNA polymerase.

[0044] As shown in Figures 3 and 4, the present disclosure also provides for closing the "free end" of a long hairpin double-stranded template by attaching a second nucleic acid-based hairpin or dumbbell adapter to the double-stranded template to form an asymmetric closed-end double-stranded nucleic acid template. The second nucleic acid-based hairpin or dumbbell adapter may contain a sequence that loops back on itself, thereby attaching the end to the double-stranded template. As shown in Figure 3, the second nucleic acid-based hairpin or dumbbell adapter to the double-stranded template may have a blunt end and may be attached to a double-stranded template with a blunt end. As shown in Figure 4, the second nucleic acid-based hairpin or dumbbell adapter may have a 5' overhang or a 3' overhang end and may be attached to a double-stranded template with a complementary overhang end. The overhang end of the double-stranded template may be generated by digesting with a restriction enzyme and designing the second nucleic acid-based hairpin or dumbbell adapter to have a complementary overhang end. The second nucleic acid-based hairpin or dumbbell adapter can have any sequence or can be designed to include a target sequence. Examples of target sequences include, but are not limited to, a consensus primer sequence, a universal sequencing primer sequence, a barcode sequence, a restriction enzyme sequence, a TelN recognition sequence, or any combination of the foregoing. The second nucleic acid-based hairpin or dumbbell adapter attached to the template can include a sequence that is less than 50%, more than 50%, more than 70%, more than 80%, more than 90%, more than 95%, more than 98%, or 100% identical to the nucleic acid-based hairpin or dumbbell adapter attached to the double-stranded nucleic acid template or tagged DNA. The second nucleic acid-based hairpin or dumbbell adapter can be attached to the double-stranded template using a ligase. Examples of ligases include, but are not limited to, DNA ligases such as T4 DNA ligase, E. coli DNA ligase, Ampligase DNA ligase, T3 DNA ligase, T7 DNA ligase, and Taq DNA ligase; and RNA ligases such as T4 RNA ligase 1, T4 RNA ligase 2, RtcB ligase, and M. thermoautotrophicum ligase. A phosphorylation or dephosphorylation step may be performed, if necessary, before ligating the nucleic acid-based hairpin or dumbbell adapter to the double-stranded template.For purposes of this disclosure, any of the aforementioned enzymes can be further modified by genetic engineering techniques to enhance one or more functionalities of the enzyme, such as processivity, thermostability, fidelity, etc. The second nucleic acid-based hairpin or dumbbell adapter can further comprise a label to enable detection and / or purification of the adapter-containing sequence, such as the adapter attached to the double-stranded template. Dimers formed by ligating two second nucleic acid-based hairpin or dumbbell adapters together can be removed using size selection or size exclusion techniques if such dimers are indicative of downstream reactions. The resulting asymmetric, closed-end double-stranded nucleic acid template can be amplified by rolling circle amplification to form nanoballs. The nanoballs are loaded into wells of an automated sequencing platform, from which sequencing occurs. The nanoball structure can be of sufficient size to exclude all but one template nanoball from entering the sequencing well, thus avoiding multiple interfering sequencing reactions from individual wells. Furthermore, the nanoball workflow steps can be prepared in solution or in tubes.

[0045] An alternative embodiment for closing a double-stranded template is shown in Figure 5. The double-stranded template comprises a TelN recognition sequence, and the double-stranded template can be "closed" by using a TelN protelomerase (e.g., phage N15 protelomerase). In certain embodiments, the TelN recognition sequence is 5'-TATCAGCACACAATTGCCCATTATACGCGCGTATAATGGACTATTGTGTGCTGATA-3' (SEQ ID NO: 1) It contains a 56-bp sequence of 5'-ATAGTCGTGTGTTAACGGGTAATATGCGCGCATATTACCTGATAACACACGACTAT-3' (SEQ ID NO: 2).

[0046] As shown in Figure 6, many of the steps of the methods disclosed herein can be performed in wells of a sequencing platform, such as zero-mode waveguide (ZMW) wells. For example, a ZMW well can be loaded with a double-stranded nucleic acid template containing an attached hairpin or dumbbell adapter, a polymerase, TelN protelomerase, one or more sequencing primers, and a reaction buffer containing dNTPS or NTPS. In a first read, the 3' end of the nucleic acid hairpin or dumbbell sequence is extended to generate a double-stranded template containing a closed hairpin ("hairpin end"), and the other end of the duplex contains a free 3'-strand end and a free 5'-strand end ("free end"). This generates a TelN recognition sequence upon which the TelN enzyme acts, thereby forming a closed-end double-stranded nucleic acid template to which a sequencing primer binds, allowing sequencing to be performed on the automated platform in a subsequent read.

[0047] As shown in Figures 7-9, nucleic acid-based hairpin or dumbbell adapters are attached to double-stranded nucleic acid templates via a transposase-mediated tagging or transposition reaction. Examples of such reactions are described in U.S. Publication No. 2010 / 0120098, which is incorporated by reference in its entirety. Transoposomes have free DNA ends that are randomly inserted into DNA in a "cut and paste" reaction. Because the DNA ends are free, they effectively fragment the DNA while adding nucleic acid-based hairpin or dumbbell adapters. Exemplary transposition complexes suitable for use in the methods provided herein include, but are not limited to, those formed by a hyperactive Tn5 transposase and Tn5-type transposon ends, or a MuA transposase and Mu transposon ends, including R1 and R2 end sequences, transposase Tn3, and Sleeping Beauty transposase (see, e.g., Goryshin and Reznikoff, J. Biol. Chem. 273:7367, 1998 and Mizuuchi, Cell 35:785, 1983; Savilahti et al., EMBO J. 14:4893, 1995, which are incorporated herein by reference in their entireties). However, the transposition system can insert transposon ends with sufficient efficiency to join nucleic acid-based hairpin or dumbbell adapters to the 5' end of a double-stranded nucleic acid template.Other examples of known transposition systems that can be used in the provided methods include Staphylococcus aureus Tn552, Tyl, transposons Tn7, Tn / O and IS10, mariner transposase, Tel, P elements, Tn3, bacterial insertion sequences, retroviruses, and yeast retrotransposons (e.g., Colegio et al., 2001, J. Bacteriol. 183:2384-8; Kirby et al., 2002, Mol. Microbiol. 43:173-86; Devine and Boeke, 1994, Nucleic Acids Res., 22:3765-72; International Patent Application No. 95 / 23875; Craig, 1996, Science 271:1512; Craig, 1996, Review in: Curr. Top Microbiol. Immunol. 204:27-48; Kleckner et al. al.,1996,Curr Top Microbiol Immunol.204:49-82, Lampe et al.,1996,EMBO J.15:5470-9,Plasterk,1996,Curr Top Microbiol Immunol 204:125-43,Gloor,2004,Methods Mol.Biol.260:97-114, Ichikawa and Ohtsubo,1990,J Biol.Chem.265:18829-32, Ohtsubo and Sekine,1996,Curr.Top.Microbiol.Immunol.204:1-26,Brown et al.,1989,Proc Natl Acad Sci USA 86:2525-9, Boeke and Corces, 1989, Annu Rev. Microbiol. 43:403-34, which are incorporated herein by reference in their entireties. In certain embodiments, the tagged template comprises a mosaic end (ME) sequence and the transposase is a Tn5 transposase.

[0048] After the transposase reaction, the resulting tagged DNA, which contains complementary hairpin loops at the 5' end of each strand, contains single-stranded gaps (e.g., 9 bp gaps) that can be filled using a gap-filling reaction using Klenow fragment, T4 DNA polymerase, and / or Ampligase (see Figure 7). The resulting closed-end tagged DNA is then nicked in the top strand using any number of techniques, including the use of site-specific endonucleases, user digestion, incorporation of RNA bases followed by ribonuclease, diols, etc. In certain embodiments, a site-specific endonuclease is used to generate a nick in the top strand. The nicked tagged template can be extended from each nick, a sequence complementary to the double-stranded nucleic acid template, by using a processive polymerase to generate two long hairpin double-stranded templates, where one end of the double-stranded template contains a closed hairpin ("hairpin end") and the other end of the duplex contains a 3'-strand end and a 5'-strand end ("free end"). If the double-stranded template is designed to contain a TelN recognition sequence, the double-stranded template can be "closed" using the same methods described above, including the use of TelN protelomerase.

[0049] Alternatively, a polymerase flash can be used after the transposase reaction (see Figure 8). Tagged DNA containing complementary hairpin loops at the 5' end of each strand is first generated. Then, two extended strands of tagged DNA containing transposase recognition sequences complementary to the transposase recognition sequences at the 5' and 3' ends of each extended strand of tagged DNA are synthesized using a polymerase. The extended strands are then separated, and the transposase recognition sequences complementary to the transposase recognition sequences at the ends of the tagged DNA are reformed into hairpin loops. The extended strands containing the hairpin loops are then loaded into the wells of an automated sequencing platform and sequenced using a sequencing polymerase (see Figure 9).

[0050] In some embodiments, sequencing the asymmetric closed-ended double-stranded nucleic acid template comprises using one or more of sequencing-by-synthesis, bridge PCR, chain termination sequencing, sequencing-by-hybridization, nanopore sequencing, and sequencing-by-ligation.

[0051] In some embodiments, the sequencing methodology used in the methods provided herein is sequencing-by-synthesis (SBS). SBS involves monitoring the extension of a nucleic acid primer along a nucleic acid template (e.g., a target nucleic acid or an amplicon thereof) to determine the sequence of nucleotides in the template. The underlying chemical process can be polymerization (e.g., catalyzed by a polymerase enzyme). In certain polymer-based SBS embodiments, fluorescently labeled nucleotides are added to the primer (thereby extending the primer) in a template-dependent manner, such that detection of the order and type of nucleotides added to the primer can be used to determine the sequence of the template.

[0052] Other sequencing procedures using cyclic reactions can be used, such as pyrosequencing. Pyrosequencing detects the release of inorganic pyrophosphate (PPi) when a specific nucleotide is incorporated into a nascent nucleic acid chain (Ronaghi, et al., Analytical Biochemistry 242(1), 84-9 (1996); Ronaghi, Genome Res. 11(1), 3-11 (2001); Ronaghi et al. Science 281(5375), 363 (1998); U.S. Patent No. 6,210,891, U.S. Patent No. 6,258,568, and U.S. Patent No. 6,274,320, each of which is incorporated herein by reference). In pyrosequencing, the released PPi can be detected by being immediately converted into adenosine triphosphate (ATP) by ATP sulfurylase, and the level of generated ATP can be detected through luciferase-generated photons. Therefore, the sequencing reaction can be monitored via a luminescence detection system. The excitation radiation source used in fluorescence-based detection systems is not required for the pyrosequencing procedure. Useful fluid systems, detectors, and procedures that can be adapted to the application of pyrosequencing to amplicons generated by the present disclosure are described, for example, in International Application No. PCT / US11 / 57111, U.S. Patent Application Publication No. 2005 / 0191698A1, U.S. Patent No. 7,595,883, and U.S. Patent No. 7,244,559, each of which is incorporated herein by reference.

[0053] Some embodiments can utilize methods involving real-time monitoring of DNA polymerase activity. For example, nucleotide incorporation can be detected via fluorescence resonance energy transfer (FRET) interactions between fluorophore-bearing polymerases and γ-phosphate-labeled nucleotides or using zero-mode waveguides (ZMW). Techniques and reagents for FRET-based sequencing are described, for example, in Levene et al. Science 299, 682-686 (2003), Lundquist et al. Opt. Lett. 33, 1026-1028 (2008), and Korlach et al. Proc. Natl. Acad. Sci. USA 105, 1176-1181 (2008), the disclosures of which are incorporated herein by reference.

[0054] Some SBS embodiments involve the detection of protons released upon incorporation of a nucleotide into an extension product. For example, sequencing based on detection of released protons can use commercially available electrical detectors and related technology from IonTorrent (Guilford, CT, a Life Technologies subsidiary), or the sequencing methods and systems described in U.S. Patent Application Publication Nos. 2009 / 0026082A1, 2009 / 0127589A1, 2010 / 0137143A1, or 2010 / 0282617A1, each of which is incorporated herein by reference. The methods described herein for amplifying target nucleic acids using equilibrium exclusion can be readily adapted to substrates used to detect protons. More specifically, the methods described herein can be used to generate clonal populations of amplicons used to detect protons.

[0055] Another useful sequencing technique is nanopore sequencing (see, e.g., Deamer et al. Trends Biotechnol. 18, 147-151 (2000); Deamer et al. Acc. Chem. Res. 35:817-825 (2002); Li et al. Nat. Mater. 2:611-615 (2003), the disclosures of which are incorporated herein by reference). In some nanopore embodiments, the target nucleic acid or individual nucleotides removed from the target nucleic acid pass through the nanopore. As the nucleic acid or nucleotide passes through the nanopore, each nucleotide species can be identified by measuring the fluctuations in the electrical conductance of the pore. (U.S. Pat. No. 7,001,792; Soni et al. Clin. Chem. 53, 1996-2001 (2007); Healy, Nanomed. 2, 459-481 (2007); Cockroft et al. J. Am. Chem. Soc. 130, 818-820 (2008), the disclosures of which are incorporated herein by reference).

[0056] From the foregoing description, it will be apparent that the invention described herein may be adapted for various uses and conditions with variations and modifications, and such embodiments also fall within the scope of the following claims.

[0057] The recitation of a list of elements in any definition of a variable herein includes definitions of that variable as any single element or combination (or subcombination) of the listed elements. The recitation of an embodiment herein includes that embodiment as any single embodiment or in combination with any other embodiment or portion thereof.

[0058] All patents and publications mentioned in this specification are herein incorporated by reference to the same extent as if each individual patent and publication was specifically and individually indicated to be incorporated by reference.

Claims

1. 1. A method for generating an asymmetric closed-end double-stranded nucleic acid template from a double-stranded nucleic acid template having free 5′ and 3′ ends, comprising: (A) attaching a first nucleic acid-based hairpin or dumbbell adaptor to the 3' end of a double-stranded nucleic acid template comprising free 5' and 3' ends; (B) extending a sequence complementary to the double-stranded nucleic acid template from each 3' end of the nucleic acid-based hairpin or dumbbell adaptor using a processive polymerase to generate two long hairpin double-stranded templates, one end of the double-stranded template comprising a closed hairpin ("hairpin end") and the other end of the double-stranded template comprising a free 3' strand end and a free 5' strand end ("free end"); (C) and ligating a second nucleic acid-based hairpin or dumbbell adaptor to the free ends of the double-stranded templates, wherein the second nucleic acid-based hairpin or dumbbell adaptor is different from the first nucleic acid-based hairpin or dumbbell adaptor, and wherein a single sequencing oligonucleotide binds to only one terminal adaptor of the closed-ended double-stranded nucleic acid template, thereby closing the free end of each long hairpin double-stranded template and forming an asymmetric closed-ended double-stranded nucleic acid template.

2. 2. The method of claim 1, wherein the double-stranded nucleic acid template is a double-stranded DNA template.

3. 3. The method of claim 1, wherein the 5' and 3' ends of the double-stranded nucleic acid template are dephosphorylated and end-repaired.

4. The method of any one of claims 1 to 3, wherein the double-stranded nucleic acid template has blunt 5' and 3' ends.

5. The method of any one of claims 1 to 3, wherein the double-stranded nucleic acid template has an A-tailed 3' end.

6. 6. The method of any one of claims 1 to 5, wherein the first nucleic acid-based hairpin or dumbbell adaptor is ligated to the 3' end of the double-stranded nucleic acid template using a ligase.

7. The method of claim 6, wherein the ligase is T4 DNA ligase or T3 DNA ligase.

8. The method of any one of claims 1 to 7, wherein the first nucleic acid-based hairpin or dumbbell adaptor comprises a blunt end or a T-tail end.

9. The method of any one of claims 1 to 7, wherein the first nucleic acid-based hairpin adaptor is a Y-shaped adaptor.

10. 10. The method of any one of claims 1 to 9, wherein dimers formed from two first nucleic acid-based hairpin or dumbbell adaptors bound to each other are removed by using size selection or size exclusion techniques.

11. The method of any one of claims 1 to 10, wherein the processive polymerase is Phi29 polymerase.

12. The method of any one of claims 1 to 11, wherein the second nucleic acid-based hairpin or dumbbell adaptor comprises a blunt end or a T-tail end.

13. 12. The method of any one of claims 1 to 11, wherein prior to step (C), the long hairpin double-stranded template is digested with a restriction enzyme that generates a 5' overhang.

14. 12. The method of any one of claims 1 to 11, wherein prior to step (C), the long hairpin double-stranded template is digested with a restriction enzyme that generates a 3' overhang.

15. 15. The method of claim 13 or 14, wherein the second nucleic acid-based hairpin or dumbbell adaptor comprises an overhang sequence complementary to the overhang sequence of the digested long hairpin double-stranded template.

16. 16. The method of any one of claims 1 to 15, wherein the second nucleic acid-based hairpin or dumbbell adaptor is ligated to the free end of the double-stranded template using polynucleotide kinase and ligase.

17. 17. The method of any one of claims 1 to 16, wherein dimers formed from two second nucleic acid-based hairpin or dumbbell adaptors bound to each other are removed by using size selection or size exclusion techniques.

18. 18. The method of any one of claims 1 to 17, further comprising the step of (C') using rolling circle replication to generate nanoball complexes comprising polycistronic amplification asymmetric closed-ended double-stranded nucleic acid templates.

19. 19. The method of any one of claims 1 to 18, further comprising the step of (D) sequencing the asymmetric closed-end double-stranded nucleic acid template or nanoball complex using a sequencing primer and a polymerase.

20. 1. A method for generating an asymmetric double-stranded nucleic acid template from tagged DNA containing complementary hairpin loops, comprising: (I) generating tagged DNA comprising a complementary hairpin loop at the 5' end of each strand, the hairpin loop comprising a base-paired transposase recognition sequence, and a single-stranded sequence gap between the 5' end and the 3' end of the tagged DNA; (II) filling the gap between the 5′ end and the 3′ end of the tagged DNA using a gap-fill ligation reaction to form a closed-end tagged DNA; (III) generating a nick in the top strand of each hairpin region of the closed-end tagged DNA; (IV) using a processive polymerase to extend a sequence complementary to the double-stranded nucleic acid template from each nick to generate two long hairpin double-stranded templates, one end of the double-stranded template comprising a closed hairpin (a "hairpin end") and the other end of the double-stranded template comprising a 3' strand end and a 5' strand end (a "free end"); (V) and ligating nucleic acid-based hairpin or dumbbell adaptors to the free ends of the double-stranded templates, wherein the nucleic acid-based hairpin or dumbbell adaptors are distinct from hairpin loops, and wherein a single sequencing oligonucleotide binds to only one hairpin or dumbbell of the asymmetric double-stranded nucleic acid template, thereby closing the free ends of each long hairpin double-stranded template and forming an asymmetric closed-end double-stranded nucleic acid template.

21. 21. The method of claim 20, wherein the transposase recognition sequence is a 19-bp mosaic terminal sequence.

22. 22. The method of claim 21, wherein the gap is 9 base pairs in length.

23. The method of any one of claims 20 to 22, wherein the gap-fill ligation reaction comprises Klenow fragment.

24. The method of any one of claims 20 to 22, wherein the gap-fill ligation reaction comprises T4 DNA polymerase and ampligase.

25. The method of any one of claims 20 to 24, wherein the nick is generated using a site-specific endonuclease.

26. 26. The method of any one of claims 20 to 25, wherein the processive polymerase is Phi29 polymerase.

27. 27. The method of any one of claims 20 to 26, wherein the nucleic acid-based hairpin or dumbbell adaptor comprises a blunt end or a T-tail end.

28. 27. The method of any one of claims 20 to 26, wherein prior to step (V), the long hairpin double-stranded template is digested with a restriction enzyme that generates a 5' overhang.

29. 27. The method of any one of claims 20 to 26, wherein prior to step (V), the long hairpin double-stranded template is digested with a restriction enzyme that generates a 3' overhang.

30. 30. The method of claim 28 or 29, wherein the nucleic acid-based hairpin or dumbbell adaptor comprises an overhang sequence complementary to the overhang sequence of the digested long hairpin double-stranded template.

31. 31. The method of any one of claims 20 to 30, wherein the nucleic acid-based hairpin or dumbbell adaptor is ligated to the free end of the double-stranded template using polynucleotide kinase and ligase.

32. 32. The method of any one of claims 20 to 31, wherein dimers formed from two nucleic acid-based hairpin or dumbbell adaptors that are linked to each other are removed by using size selection or size exclusion techniques.

33. (V') The method of any one of claims 20 to 23, further comprising the step of using rolling circle replication to generate nanoball complexes comprising polycistronic amplification asymmetric closed-ended double-stranded nucleic acid templates.

34. (VI) sequencing the asymmetric closed-end double-stranded nucleic acid template or nanoball complex using a sequencing primer and a polymerase.

35. 1. A method for sequencing tagged DNA containing complementary hairpin loops, comprising: (I) generating tagged DNA comprising a complementary hairpin loop at the 5' end of each strand, the hairpin loop comprising a base-paired transposase recognition sequence, and a single-stranded sequence gap between the 5' end and the 3' end of the tagged DNA; (II) using a polymerase to form two stretches of tagged DNA, each stretch of tagged DNA comprising a transposase recognition sequence and a complementary transposase recognition sequence at the 5' and 3' ends; (III) separating the stretches of tagged DNA and rehybridizing the transposase recognition sequence and the complementary transposase recognition sequence at each end of the stretches of tagged DNA to form complementary hairpin loops; (IV) using a sequencing polymerase to sequence the extended strand of tagged DNA that includes the complementary hairpin loop.

36. 36. The method of claim 35, wherein step (IV) is performed in a well of an automated sequencing platform.

37. Prior to the sequencing step (IV), (III') The method of claim 35, wherein nanoball complexes comprising polycistronic amplified extension strands of tagged DNA are generated.

Citation Information

Patent Citations

  • Systems and methods for clonal replication and amplification of nucleic acid molecules for genomic and therapeutic applications

    JP2017512071A

  • Methods and compositions for generating asymmetrically-tagged nucleic acid fragments

    US20170362639A1

  • Methods and systems for evaluating DNA methylation in cell-free DNA

    WO2019006269A1

  • Target enrichment by unidirectional dual probe primer extension

    WO2019121842A1