Methods and kits for preparing engineered nucleic acids

The method of tagmentation, denaturation, and ligation with blocking molecules and splint oligonucleotides addresses the issue of redundant tagging in conventional library construction, enhancing sequencing yield and sensitivity by producing high-quality nucleic acid libraries with reduced redundancy and bias.

WO2026107435A1PCT designated stage Publication Date: 2026-05-21SEQWELL INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SEQWELL INC
Filing Date
2025-11-17
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

Conventional transposase-based library construction methods result in 50% of DNA molecules being non-functional due to redundant tagging with the same adapter sequence, requiring excessive DNA amounts and empirical data correction, decreasing unique molecules, and increasing amplified duplicates, thus limiting sequencing yield and sensitivity for low-frequency variants.

Method used

A method involving tagmentation with a transposome and a first nucleic acid adapter, followed by denaturation, blocking molecule complex formation, exposure to splint oligonucleotides with a degenerate sequence, and ligation of a second adapter to the 3' end, producing an engineered nucleic acid sample with reduced redundancy.

Benefits of technology

This approach enhances sequencing yield and sensitivity by reducing redundant tagging, improving library quality, and minimizing amplified duplicates, thereby increasing the number of unique molecules and reducing sequencing bias.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000028_0001_TABLE
    Figure IMGF000028_0001_TABLE
  • Figure IMGF000029_0001_TABLE
    Figure IMGF000029_0001_TABLE
  • Figure IMGF000030_0001_TABLE
    Figure IMGF000030_0001_TABLE
Patent Text Reader

Abstract

The invention provides kits comprising transposomes, artificial nucleic acids that include nucleic acid adapters, blocking molecules, and as well as methods of using the same, for example, for preparation of engineered nucleic acid samples, including nucleic acid libraries for sequencing.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] PATENT

[0002] Attorney Docket No.: 51178-019WO2

[0003] METHODS AND KITS FOR PREPARING ENGINEERED NUCLEIC ACIDS

[0004] SEQUENCE LISITING

[0005] The instant application contains a Sequence Listing which has been submitted electronically in XML format and is hereby incorporated by reference in its entirety. Said XML copy, created on November 13, 2025, is named “51178-019WO2_Sequence_Listing_11_13_25.xml” and is 16,144 bytes in size.

[0006] FIELD OF THE INVENTION

[0007] The present invention relates generally to methods of preparing nucleic acid samples, such as a nucleic acid library and kits that include artificial nucleic acids, proteins including transposases, and complexes including the same.

[0008] BACKGROUND

[0009] Conventional transposase-based library construction employs hyperactive Tn5 loaded with canonical mosaic end sequence and two different adapters encoding sequencing primers. Both adapter sequences must be present to produce a functional molecule for sequencing. However, current strategies rely on simultaneous transposition of two adapters from transposomes, which renders 50% of resultant DNA molecules non-functional due to “redundant” tagging with the same adapter sequence on both ends of a molecule. These strategies, therefore, present numerous challenges for production of high-quality sequencing libraries, including requiring greater amounts of DNA in a sample for library construction and requiring empirical data correction for library quantification since redundantly tagged molecules will contribute more noise in a data set.

[0010] Additionally, redundant tagging decreases the number of unique molecules in each library and increases the fraction of amplified duplicates observed during sequencing. Thus, strategies that require simultaneous transposition of two adapters to generate a fragmented DNA library decrease overall sequencing yield and have limited sensitivity for low frequency variants in a mixture of a DNA sample. Therefore, there is a need for new methods and reagents for library preparation.

[0011] SUMMARY OF THE INVENTION

[0012] The present invention provides compositions and methods of use thereof, e.g., for nucleic acid library preparation and sequencing.

[0013] In one aspect, the invention features a method of preparing an engineered nucleic acid sample, the method including:

[0014] (a) exposing a sample double stranded nucleic acid to a transposome including a transposase and a first nucleic acid adapter under conditions and for a time sufficient for the transposome to carry out a tagmentation reaction, wherein the tagmentation reaction produces a nucleic acid fragment including the first nucleic acid adapter transposed to a 5’ end of a fragment of the sample double-stranded nucleic acid, and wherein the nucleic acid fragment is double-stranded in part; PATENT

[0015] Attorney Docket No.: 51178-019WO2

[0016] (b) denaturing the nucleic acid fragment to generate a single-stranded nucleic acid fragment, and exposing the single-stranded nucleic acid fragment to a plurality of blocking molecules, wherein the blocking molecules form a complex with the single-stranded nucleic acid fragment;

[0017] (c) exposing the complex including the single-stranded nucleic acid fragment and the blocking molecule to a plurality of splint oligonucleotides including a second nucleic acid adapter, wherein the splint oligonucleotides include a degenerate nucleic acid sequence at the 3’ end, wherein one of the plurality of splint oligonucleotides hybridizes to the 3’ end of the single-stranded nucleic acid fragment via the degenerate nucleic acid sequence; and

[0018] (d) ligating the second nucleic acid adapter to the 3’ end of the single-stranded nucleic acid fragment using a ligase,

[0019] thereby producing an engineered nucleic acid sample.

[0020] In some embodiments, in step (b), the nucleic acid fragment is denatured via heat or chemical denaturation.

[0021] In some embodiments, the degenerate nucleic acid sequence is between 5 to 15 nucleotides in length (e.g., 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, or 15 nucleotides in length).

[0022] In some embodiments, the plurality of blocking molecules includes a protein.

[0023] In some embodiments, the protein is a single-stranded nucleic acid binding protein.

[0024] In some embodiments, the plurality of blocking molecules includes single-stranded oligonucleotides having different sequences.

[0025] In some embodiments, the single-stranded oligonucleotides include a 3’ end modification to prevent extension by a polymerase. In some embodiments, the 3’ end modification is selected from the group consisting of an amino group modification, a phosphate group, a three-carbon chain (C3), a dideoxynucleotide, an inverted nucleotide, or a locked nucleic acid.

[0026] In some embodiments, the single-stranded oligonucleotides are between 5 and 15 nucleotides in length (e.g., 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, or 15 nucleotides in length).

[0027] In some embodiments, the single-stranded oligonucleotides are between 5 and 10 nucleotides in length (e.g., 5, 6, 7, 8, 9, or 10 nucleotides in length).

[0028] In some embodiments, the sample double-stranded nucleic acid is DNA or a DNA / RNA hybrid (e.g., an R-loop). In some embodiments, the DNA is genomic DNA, mitochondrial DNA, complementary DNA (cDNA), or plasmid DNA.

[0029] In some embodiments, the transposase is selected from the group consisting of a Tn5 transposase, a hyperactive Tn5 transposase, and a TnX transposase.

[0030] In some embodiments, the first nucleic acid adapter and the second nucleic acid adapter each include a priming region. In some embodiments, the priming region of the first nucleic acid adapter has a different nucleotide sequence and is not the reverse complement of the priming region of the second nucleic acid adapter.

[0031] In any one of the preceding embodiments, the method further includes the step of amplifying the engineered nucleic acid sample. In some embodiments, the amplification is performed via polymerase chain reaction (PCR), multiple annealing and looping-based amplification cycles (MALBAC), or multiple displacement amplification (MDA). PATENT

[0032] Attorney Docket No.: 51178-019WO2

[0033] In another aspect, the invention features a kit including:

[0034] (i) a transposome including a transposase and a first nucleic acid adapter;

[0035] (ii) a plurality of blocking molecules; and

[0036] (iii) a second nucleic acid adapter including a splint oligonucleotide including a degenerate nucleic acid sequence at the 3’ end.

[0037] In some embodiments, the kit further includes a ligase. In some embodiments, the kit further includes a suitable buffer.

[0038] In some embodiments, the degenerate nucleic acid sequence is between 5 to 15 nucleotides in length (e.g., 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, or 15 nucleotides in length).

[0039] In some embodiments, the plurality of blocking molecules includes a protein. In some embodiments, the protein is a single-stranded nucleic acid binding protein.

[0040] In some embodiments, the plurality of blocking molecule includes single-stranded oligonucleotides having different sequences.

[0041] In some embodiments, the single-stranded oligonucleotides include a 3’ end modification to prevent extension by a polymerase. In some embodiments, the 3’ end modification is selected from the group consisting of an amino group modification, a phosphate group, a C3, a dideoxynucleotide, an inverted nucleotide, or a locked nucleic acid.

[0042] In some embodiments, the single-stranded oligonucleotides are between 5 and 15 nucleotides in length (e.g., 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, or 15 nucleotides in length).

[0043] In some embodiments, the single-stranded oligonucleotides are between 5 and 10 nucleotides in length (e.g., 5, 6, 7, 8, 9, or 10 nucleotides in length).

[0044] In some embodiments, the transposase is selected from the group consisting of a Tn5 transposase, a hyperactive Tn5 transposase, and a TnX transposase.

[0045] In some embodiments, the first nucleic acid adapter and the second nucleic acid adapter each include a priming region. In some embodiments, the priming region of the first nucleic acid adapter has a different nucleic acid sequence and is not the reverse complement of the priming region of the second nucleic acid adapter.

[0046] Definitions

[0047] The term “about”, as applied to a numeric value, includes ± 10% of the recited value.

[0048] Unless otherwise specified, the terms “a” or “an” mean “one or more” throughout this application.

[0049] As used herein, the term “adapter” or “adapter oligonucleotide” refers to any nucleic acid used to modify a target nucleic acid to make it suitable for amplification or nucleic acid sequencing. In some instances, an adapter may include a nucleic acid sequence for binding transposase, e.g., the transposase mosaic end (ME) sequence. In some instances, an adapter may include a nucleic acid sequence that is homologous or complementary to a nucleic acid sequence used for sequencing. In some instances, an adapter may include a barcode sequence. In some instances, an adapter may include a nucleic acid sequence for amplification. In some instances, an adapter may be bound to a solid surface. In some instances, an adapter may be bound to a soluble molecular scaffold. PATENT

[0050] Attorney Docket No.: 51178-019WO2

[0051] As used herein, the term “amplify” or “amplification” refers to the act or method of generating copies (i.e., amplicons) of a nucleic acid molecule. Methods of nucleic acid amplification are known in the art and include polymerase chain reaction (PCR), ligase chain reaction (LCR), looping-based amplification cycles (MALBAC), and multiple displacement amplification (MDA. In some instances, PCR may be performed using one or more pairs of sequencing oligonucleotides and / or one or more pairs of barcoding oligonucleotides as primers.

[0052] As used herein, the term “blocking molecule” refers to a molecule that binds or forms a complex with a single-stranded portion of a nucleic acid, e.g., to minimize or prevent the formation of secondary structure in the nucleic acid. A suitable blocking molecule of the invention includes a protein or an oligonucleotide (e.g., a single-stranded oligonucleotide). A blocking molecule may be a protein that is highly stable and is resistant to unfolding or denaturing in the presence of heat (e.g., temperatures greater than or equal to 85°C) or in the presence of chemical agents commonly used for nucleic acid denaturation (e.g., urea, sodium hydroxide, guanidine hydrochloride, dimethyl sulfoxide, propylene glycol, or formamide). A blocking molecule may be an oligonucleotide such as a singlestranded oligonucleotide that is between about 5 to about 500 nucleotides in length (e.g., 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41 , 42, 43, 44, 45, 46, 47, 48, 49, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490 or 500 nucleotides in length) and can hybridize to a single-stranded portion of a nucleic acid, such as a single-stranded nucleic acid produced from one or more methods described herein.

[0053] As used herein, the term “barcode” refers to a unique oligonucleotide sequence that may allow the corresponding sample to be identified. In some embodiments, the barcode may be located at a specific position in a longer nucleic acid sequence.

[0054] As used herein, the terms “complement,” “complementary,” or “complementarity” in reference to nucleic acid sequences means that a sequence of a first nucleic acid in relation to that of a second nucleic acid will form hydrogen bonds with that of an opposing (i.e., antiparallel) nucleic acid strand. The 5’ end of the first nucleic acid when aligned to the 3’ end of the second nucleic acid, and vice versa to each other will have complementary structures following a lock-and-key principle (i.e., A will be paired with U or T and G will be paired with C).

[0055] As used herein, the term “degenerate” in reference to a nucleic acid or portion of a nucleic acid, refers to a nucleic acid sequence or a portion thereof that contains random nucleotides, such that a population of degenerate nucleic acids includes a degenerate sequence region in each degenerate nucleic acid that differs in one or more nucleotides. In some embodiments, degenerate nucleic acids of a given length include nucleic acids having combinatorial sequences of the given length (e.g., comprising a large number of or all of the possible canonical nucleotide sequences of the given length (e.g., >50%, >60%, >70%, >80%, >90%, or 100% of the possible sequences).

[0056] As used herein, the term “flank” refers to the relative positions of three nucleic acid regions. A first and second nucleic acid region is said to flank a third nucleic acid region if the first and second PATENT

[0057] Attorney Docket No.: 51178-019WO2

[0058] regions lie upstream and downstream of the third nucleic acid region. The first and second nucleic acid regions may be directly adjacent to the third nucleic acid region.

[0059] As used herein, the term “homologous” refers to having substantially the same sequence. Homologous sequences may differ by up to one third of nucleotide bases. For example, two sequences that are nine bases in length may differ at most by 3, at most by 2, at most by 1 , or at most by 0 nucleotide bases, and remain homologous to one another.

[0060] As used herein, the term “hybridization” refers to a process in which two single-stranded nucleic acids bind non-covalently by base pairing to form a stable double-stranded nucleic acid. Hybridization may occur for the entire lengths of the two nucleic acids, or only for a portion or subregion of one or both of the nucleic acids. The resulting double-stranded nucleic acid molecule or region is a “duplex.”

[0061] As used herein, the terms “increased enzymatic activity” or “hyperactive enzyme” refers to an improved property of an enzyme (e.g. a transposase), which can be represented by an increase in specific activity (e.g., product yield and / or kinetics) or an increase in percent conversion of the substrate to the product (e.g., percent conversion of starting amount of substrate to product in a specified time period using a specified amount of transposase) as compared to the reference transposase enzyme. Any property relating to enzyme activity may be affected, including the classical enzyme properties of Km, Vmax or kcat, changes of which can lead to increased enzymatic activity. Improvements in enzyme activity can be from about 1.2 times the enzymatic activity of the corresponding wild-type enzyme, to as much as 2 times, 5 times, 10 times, 20 times, 25 times, 50 times or more enzymatic activity than the naturally occurring or another engineered transposase from which the transposase polypeptides were derived. Transposase activity can be measured by any one of standard assays, such as by monitoring changes in properties of substrates, cofactors, or products. In some embodiments, the amount of products generated can be measured by Liquid Chromatography-Mass Spectrometry (LC-MS), HPLC, capillary electrophoresis or other methods, as known in the art. Comparisons of enzyme activities are made using a defined preparation of enzyme, a defined assay under a set condition, and one or more defined substrates, as further described in detail herein. Generally, when lysates are compared, the numbers of cells and the amount of protein assayed are determined as well as use of identical expression systems and identical host cells to minimize variations in amount of enzyme produced by the host cells and present in the lysates.

[0062] As used herein, the terms “library” or “fragment library” refers to a collection of nucleic acids (e.g., engineered nucleic acids) derived from one or more nucleic acid samples, in which fragments of nucleic acid have been modified, generally by incorporating terminal adapter sequences comprising one or more domains to which one or more primers can bind and / or identifiable sequence tags (e.g., barcodes).

[0063] As used herein, the term “nucleic acid” refers to a polymeric molecule of at least two linked nucleotides. The terms include, for example, deoxyribonucleic acid (DNA) and ribonucleic acid (RNA), as well as hybrids and mixtures thereof. A nucleic acid may be double-stranded, single-stranded, or contain a mix of regions or portions of both single-stranded or double-stranded sequences. The nucleotides in a nucleic acid are usually linked by phosphodiester bonds, though “nucleic acid” may PATENT

[0064] Attorney Docket No.: 51178-019WO2

[0065] also refer to other molecular analogs having other types of chemical bonds or backbones, including, but not limited to, phosphoramide, phosphorothioate, phosphorodithioate, O-methyl phosphoramidate, morpholino, locked nucleic acid (LNA), glycerol nucleic acid (GNA), threose nucleic acid (TNA), and peptide nucleic acid (PNA) linkages or backbones. Nucleic acids may contain any combination of deoxyribonucleotides, ribonucleotides, or non-natural analogs thereof. Examples of nucleic acids include, but are not limited to, a gene, a gene fragment, a genomic gap, an exon, an intron, intergenic DNA (including, without limitation, heterochromatic DNA), messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozymes, small interfering RNA (siRNA), miRNA, small nucleolar RNA (snoRNA), cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, isolated DNA of a sequence, isolated RNA of a sequence, nucleic acid probes, and primers.

[0066] As used herein, a “nucleic acid sample” or a “target nucleic acid” refers to any nucleic acid (e.g., DNA) of interest that is selected for a method of the invention (e.g., tagmentation and / or splint ligation) using one or more compositions described herein. The present methods can be carried out using nucleic acid samples (e.g., DNA samples) pooled from more than one source. It is to be understood that a nucleic acid sample may be DNA or RNA, for example. In some instances, RNA may be converted to cDNA prior to being selected for a method of the invention (e.g., tagmentation and / or splint ligation). A nucleic acid sample may be used to carry out a method of the invention (e.g., tagmentation and / or splint ligation) to generate an engineered nucleic acid sample, which includes nucleic acid fragments appended with nucleic acid adapters on the 5’ and 3’ ends. Such engineered nucleic acid samples may be amplified by any nucleic acid amplification method known in the art or described herein and / or may constitute a nucleic acid library for sequencing methods known in the art or described herein.

[0067] As used herein, the term “nucleotide” or “nt” refers to any deoxyribonucleotide, ribonucleotide, non-standard nucleotide, modified nucleotide, or nucleotide analog. Nucleotides include adenine, thymine, cytosine, guanine, and uracil. Examples of modified nucleotides include, but are not limited to, diaminopurine, 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 5-(carboxyhydroxylmethyl) uracil, 5-carboxymethylaminomethyl-2-thiouridine, 5-carboxymethylaminomethyluracil, dihydrouracil, beta-D-galactosylqueosine, inosine, N6-isopentenyladenine, 1-methylguanine, 1 -methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N6-adenine, 7-methylguanine, 5-methylaminomethyluracil, 5-methoxyaminomethyl-2-thiouracil, beta-D-mannosylqueosine, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, 2-methylthio-N6-isopentenyladenine, uracil-5-oxyacetic acid (v), wybutoxosine, pseudouracil, queosine, 2-thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, uracil-5-oxyacetic acid methylester, 5-methyl-2-thiouracil, and 3-(3-amino-3-N-2-carboxypropyl) uracil.

[0068] As used herein, the term “oligonucleotide” refers to a nucleic acid up to 500 nucleotides in length. Oligonucleotides may be synthetic. Oligonucleotides may contain one or more chemical modifications, whether on the 5’ end, the 3’ end, or internally. Examples of chemical modifications include, but are not limited to, addition of functional groups (e.g., biotins, amino modifiers, alkynes, thiol modifiers, phosphates, or azides), fluorophores (e.g., quantum dots or organic dyes), spacers PATENT

[0069] Attorney Docket No.: 51178-019WO2

[0070] (e.g., C3 spacer (e.g., 3-hydroxy propyl), dSpacer, photo-cleavable spacers), modified bases, or modified backbones.

[0071] As used herein the terms "a portion," “a part,” and / or grammatical equivalents thereof can refer to any fraction of a whole amount. For example, "a portion" can refer to at least about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 99%, 99.9% or 100% of a whole amount.

[0072] As used herein, the term “primer” refers to an oligonucleotide, either natural or synthetic, that is capable of forming a duplex with a nucleic acid template and then acting as a point of initiation of nucleic acid synthesis for extension from its 3’ end along the template nucleic acid so that an extended duplex is formed, e.g., for nucleic acid amplification or sequencing by synthesis. The sequence of nucleotides added during the extension process is determined by the sequence of the template polynucleotide. Typically, primers are extended by a DNA polymerase. Primers usually have a length in the range of between 5 to 36 nucleotides (e.g., 5-12, 12-18, 18-24, 24-30, or 30-36 nucleotides). In certain aspects, primers are universal primers or non-universal primers. Pairs of primers can flank a sequence of interest or a set of sequences of interest. Primers can be degenerate in sequence to increase the number of regions that are hybridized in a nucleic acid sample or template, e.g., for non-sequence-specific nucleic acid amplification or sequencing. Primers may be universal primers, or primers that have nucleotide sequences that are well-known and commonly used in the art for sequencing or amplification methods. A nucleic acid (e.g., from a nucleic acid sample such as an engineered nucleic acid sample described herein), includes a priming region at its 5’ and 3’ ends, in which the priming region of the 5’ end corresponds to a primer sequence and the priming region of the 3’ end corresponds to the reverse complement of the primer.

[0073] As used herein, the term “priming region” refers to a sequence to which a primer can bind or to a sequence that corresponds to a primer sequence.

[0074] As used herein, the terms “recombinant,” “engineered,” or “non-naturally occurring,” when used with reference to a cell, nucleic acid, or polypeptide, refers to a material, or a material corresponding to the natural or native form of the material, that has been produced by human intervention. In some instances, the cell, nucleic acid, or polypeptide may be modified in a manner that would not otherwise exist in nature. In some instances, the cell, nucleic acid, or polypeptide is identical to a naturally occurring cell, nucleic acid, or polypeptide, but is produced or derived from synthetic materials and / or by manipulation using recombinant techniques. Non-limiting examples include synthesized oligonucleotides such as a nucleic acid adapter described herein or a nucleic acid fragment library generated from a nucleic acid sample. Other examples include recombinant cells expressing genes that are not found within the native (non-recombinant) form of the cell or expressed native genes that are otherwise expressed at a different level.

[0075] As used herein, the term “redundant” in reference to a nucleic acid, a fragment thereof (e.g., a fragment generated for nucleic acid library preparation), or a nucleic acid adapter refers to multiple copies of a particular nucleotide sequence in a given sample (e.g., a generated sample for nucleic acid sequencing). Redundancy may arise due to identical nucleic acid adapters being appended to a particular nucleic acid fragment. In contrast, a “non-redundant” nucleic acid, fragment thereof, or a PATENT

[0076] Attorney Docket No.: 51178-019WO2

[0077] nucleic acid adapter refers to a nucleic acid with a unique identity (i.e., does not have an identical nucleic acid sequence) compared to other nucleic acids, fragments thereof, or nucleic acid adapters that are represented in a given sample (e.g., a generated sample for nucleic acid sequencing).

[0078] As used herein, the term “stability” in reference to a protein refers to the propensity of a protein to remain in its folded three-dimensional conformation in an environment. Particularly, stability refers to the equilibrium of a protein in its folded and its unfolded state in a solution and depends on a protein’s primary structure (e.g., amino acid sequence), disulfide bond content, or other modifications (e.g., post-translational modifications; e.g., methylation, glycosylation, or oxidation). Determination of a protein’s stability is routine in the art, and may be performed by measuring the melting temperature (Tm) of a purified or substantially purified protein solution via a thermal shift assay or differential scanning calorimetry.

[0079] As used herein, the terms “synaptic complex,” “transposome,” and “transposase complex” refers to a protein-nucleic acid complex including one or more transposases and one or more oligonucleotides. In some instances, the one or more oligonucleotides of the synaptic complex are inserted into a nucleic acid sequence of a nucleic acid sample by transposase activity. In some instances, the synaptic complex may include a dimer of transposase bound to two or more oligonucleotides (e.g., two transposases and two or more nucleic acid adapters that are the same or two transposases and two or more nucleic acid adapters that are mixed or different). The addition of oligonucleotides into the nucleic acid sequence of the nucleic acid sample is preceded by fragmentation of the nucleic acid at the site of insertion by transposase. In some instances, the transposase may be Tn5 transposase or an engineered transposase variant. In some instances, the oligonucleotides may be adapter sequences. In some instances, the synaptic complex is preassembled. In some instances, the synaptic complex may be bound to a solid surface. In some instances, the synaptic complex may be bound to a soluble molecular scaffold.

[0080] “Transferred” or “transposed” nucleic acid is any nucleic acid that is chemically bound to a nucleic acid sample (e.g., a double-stranded nucleic acid sample, e.g., DNA) in a transposition event (e.g., a tagmentation event).

[0081] The term “transposase” refers to a protein that can catalyze movement of a transposase binding site (TBS) as well as associated transposable nucleic acid sequence to a different nucleic acid (e.g., DNA) molecule. In nature, transposases bind to TBSs at the ends of a transposon (also known as a transposable element) prior to catalyzing movement of the transposon to a different location of the host genome. Transposases typically effect transposition of nucleic acid (e.g., DNA) sequences using a cut and paste mechanism or a replicative transposition mechanism. Transposases typically catalyze nucleic acid transposition as oligomers. For example, Tn5 transposases catalyze transposition as a dimer, with a monomer binding each TBS. Other transposases, such as Mu (also referred to as MuA), catalyze transposition as a tetramer (dimer of dimers), with a dimer binding each TBS. The term “transposase,” as used herein, refers to the minimal unit that binds to a TBS, and may include, for example, one transposase protein (e.g., a monomer) or more than one transposase protein (e.g., a dimer). Transposases are members of the RnaseH superfamily of proteins, which is characterized by an active site that includes DDE residues that chelate two Mg++ions, which are PATENT

[0082] Attorney Docket No.: 51178-019WO2

[0083] critical for catalysis, and the overall architecture and active site DDE are considered to be nearly identical to that of retroviral integrases, RuvC, and RnaseH (see, e.g., Reznikoff, Mol. Microbiol. 47(5):1199-1206, 2003). Given that transposases and retroviral integrases share common active site architecture (including the DDE active site) as well as catalytic mechanisms (e.g., transposon-donor backbone DNA nicking and strand transfer), it is expressly contemplated that retroviral integrases (e.g., human immunodeficiency virus (HIV)-1, HIV-2, simian immunodeficiency virus (SIV), and Rous sarcoma virus integrases) and other related integrases (e.g., integrases of retrotransposons, for example, yeast Ty integrases (e.g., Ty1 , Ty2, Ty3, Ty4, and Ty5 integrase)) may also be used in the context of the invention as falling within the scope of “transposase.”

[0084] A “transposase binding site” (TBS) is a nucleic acid (e.g., DNA) sequence that can be selectively bound by a transposase. In particular embodiments, the sequence is a DNA sequence. In some instances, under at least a condition specified herein and / or in the context of a sequencing method of the invention, transposase binding sites attached to the target nucleic acid (e.g., DNA) by transposase activity may remain selectively bound by transposases within the synaptic complex (e.g., transposome).

[0085] As used herein, the terms “tagmentation” or “tagment” refer to a reaction for modifying a nucleic acid (e.g., a double-stranded nucleic acid; e.g., DNA) by a synaptic complex that includes a transposase and a nucleic acid adapter. Tagmentation results in the simultaneous fragmentation of the nucleic acid sample (e.g., a double-stranded nucleic acid, e.g., DNA) and covalent attachment of the nucleic acid adapter to the 5’ end of fragmented nucleic acid sample. A suitable nucleic acid sample includes genomic DNA, mitochondrial DNA, or a mixture thereof.

[0086] BRIEF DESCRIPTION OF THE DRAWINGS FIG. 1A is a schematic diagram showing an exemplary double-stranded nucleic acid sample that has been fragmented and a first nucleic adapter that has been transposed to its 5’ end following a tagmentation reaction.

[0087] FIG. 1B is a schematic diagram showing an exemplary single-stranded nucleic acid fragment to which a plurality of blocking molecules, exemplified in this diagram as blocking oligonucleotides, can hybridize during denaturation.

[0088] FIG. 1C is a schematic diagram showing an exemplary splinted ligation reaction, in which a second nucleic acid adapter is added to the 3’ end of the single-stranded nucleic acid fragment.

[0089] FIG. 1D is a schematic diagram illustrating an exemplary engineered nucleic acid sample produced from the steps shown in FIGS. 1A-1C.

[0090] FIG.2A is a graph showing normalized sequence coverage by guanine-cytosine (GO) content in a DNA sample that included blocking proteins. Mean base quality is shown in the solid line. The histogram shows genomic bins with varied GC-content, and the mean normalized coverage across the GC-bins is shown as circles.

[0091] FIG.2B is a graph showing normalized sequence coverage by GO content in a DNA sample that included blocking oligonucleotides. Mean base quality is shown in the solid line. The histogram PATENT

[0092] Attorney Docket No.: 51178-019WO2

[0093] shows genomic bins with varied GO content, and the mean normalized coverage across the GC-bins is shown as circles.

[0094] FIG.2C is a graph showing normalized sequence coverage by GO content in a DNA sample that did not contain any blocking molecules. Mean base quality is shown in the solid line. The histogram shows genomic bins with varied GC-content, and the mean normalized coverage across the GC-bins is shown as circles.

[0095] FIG.3 is a bar graph comparing next-generation sequence total library yield in samples containing blocking proteins (cross-hatched bars) or blocking oligonucleotides (solid bars) across two replicates.

[0096] DETAILED DESCRIPTION

[0097] Described herein are methods of preparing an engineered nucleic acid sample, e.g., that is suitable for massively parallel or next generation sequencing methods. Such methods may be used for preparing a nucleic acid sample that includes fragmented nucleic acids appended with adapter molecules using tagmentation and ligation in the presence of blocking molecules, e.g., wherein the tagmentation and ligation are used to add two different adapters to the fragmented nucleic acids. In one embodiment, tagmentation adds a sequence corresponding to a primer, and ligation adds a primer binding sequence. This arrangement allows for the use of primer pairs to ensure that only sequences that have undergone both tagmentation and ligation are amplified. Additionally, provided herein are kits that include compositions for use in one or more methods described herein.

[0098] Specifically, the compositions include one or more nucleic acids or proteins or combinations thereof, such as a nucleic acid adapter, a plurality of blocking molecules (e.g., proteins and / or nucleic acids that form a complex with single-stranded nucleic acids), a transposome, a transposase, a ligase, or one or more additional components (e.g., a reaction vessel, a buffered solution, or a cofactor) useful for performing the methods described herein.

[0099] I. Methods

[0100] Conventional transposase-based library preparation strategies have presented challenges for producing high quality library samples with minimal amounts of nucleic acid starting material. Some methods have employed single-stranded ligation using T4 DNA Ligase. However, ligation of singlestranded DNA is problematic due to potential for: intramolecular folding into secondary structures, non-specific hybridization between templates blocking available 3’ ends for ligation, chimera formation, and anti-GC bias related to reassociation of GC-rich templates. Provided herein are methods and compositions for preparing nucleic acid libraries with improved coverage and sequencing quality as well as reduced bias (e.g., reduced sequence bias or GO bias).

[0101] Provided herein are methods of preparing an engineered nucleic acid sample of nucleic acid fragments that includes appended nucleic acid adapters on the 5’ and the 3’ ends of the fragment. The methods include (a) performing a tagmentation reaction for fragmenting a double-stranded nucleic acid sample (e.g., DNA, e.g., genomic DNA, mitochondrial DNA, or cDNA) into doublestranded nucleic acid fragments and transposing a first nucleic acid adapter to the 5’ end of each PATENT

[0102] Attorney Docket No.: 51178-019WO2

[0103] fragment; (b) generating single-stranded nucleic acid fragments from the double-stranded nucleic acid fragments via denaturation; and (c) performing a ligation reaction with the single-stranded nucleic acid fragments, a splint oligonucleotide that includes a second nucleic acid adapter, and a ligase for ligating the second nucleic acid adapter to the 3’ end of the single-stranded nucleic acid fragment. These methods are useful for preparing a nucleic acid library with reduced guanine-cytosine (GO) bias for further sequencing (e.g., next-generation sequencing) modalities.

[0104] A. Methods of Preparing an Engineered Nucleic Acid Sample

[0105] / '. Methods of Tagmenting a Nucleic Acid Sample

[0106] Methods of the invention described herein include a tagmentation reaction, wherein a nucleic sample is contacted with a synaptic complex (e.g., a transposome) that includes a complex of transposases each appended with a nucleic acid adapter. The tagmentation reaction generates a collection of nucleic acid fragments, in which each fragment includes the nucleic acid adapter transposed to its 5’ end.

[0107] In some embodiments, the tagmentation reaction occurs at a temperature between 25-65°C (e.g., between 35-65°C, between 40-65 °C, between 45-65°C, between 50-65°C, between 55-65°C, between 60-65°C, between 25-60°C, between 25-55 °C, between 25-50°C, between 25-45°C, between 25-40°C, between 25-35 °C, between 25-30°C, between 40-50 °C, or between 53-57°C; e.g., at about 25°C, at about 30 °C, about 35 °C, about 40°C, about 45°C, about 50°C, about 53°C, about 54°C, about 55°C, about 56°C, about 57°C, about 60°C, or about 65°C). In some instances, a tagmentation reaction temperature is selected following consideration of the rate of reaction and transposase activity at a given temperature or range of temperatures. In some embodiments, a specific temperature is maintained over the course of the enzymatic reaction to maximize enzymatic activity or reaction efficiency. In other embodiments, the temperature increases or decreases over a range of temperatures to maximize enzymatic activity or reaction efficiency.

[0108] In some instances, the tagmentation reaction occurs for a first reaction duration between 1 and 30 minutes (e.g., between 1 and 25 minutes, between 1 and 20 minutes, between 1 and 15 minutes, between 1 and 10 minutes, between 1 and 5 minutes, between 5 and 10 minutes, between 5 and 20 minutes, between 5 and 30 minutes, between 10 and 30 minutes, between 15 and 30 minutes, between 20 and 30 minutes, between 25 and 30 minutes, between 10 and 20 minutes, between 15 and 25 minutes; e.g., about 1 minute, about 5 minutes, about 10 minutes, about 13 minutes, about 14 minutes, about 15 minutes, about 16 minutes, about 17 minutes, about 20 minutes, about 25 minutes, or about 30 minutes). In some instances, the duration of the tagmentation reaction is determined by the temperature or range of temperatures at which the transposition reaction occurs. In some embodiments, the duration of the tagmentation reaction is determined by the expected efficiency or stability of a transposase.

[0109] In some embodiments, a tagmentation reaction occurs in a total volume between about 10 pL and 100 pL. In some embodiments, the tagmentation reaction occurs in a total volume of 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, PATENT

[0110] Attorney Docket No.: 51178-019WO2

[0111] 40, 41 , 42, 43, 44, 45, 46, 47, 48, 49, 50, 51 , 52, 53, 54, 55, 56, 57, 58, 59, 60, 61 , 62, 63, 64, 65, 66, 67, 68, 69, 70, 71 , 72, 73, 74, 75, 76, 77, 78, 79, 80, 81 , 82, 83, 84, 85, 86, 87, 88, 89, 90, 91 , 92, 93, 94, 95, 96, 97, 98, 99, or 100 |_iL.

[0112] Nucleic acid samples can be obtained from any source. For example, a nucleic acid sample may be prepared from nucleic acid molecules obtained from a single organism or from populations of nucleic acid molecules obtained from one or more organisms. A nucleic acid sample may include nucleic acid molecules from one or more organs, tissues, cells, or subcellular compartments. In preferred embodiments, a nucleic acid sample is double-stranded. In some embodiments, a nucleic acid sample is DNA. In some embodiments, a nucleic acid sample includes genomic DNA. In some embodiments, a nucleic acid sample includes mitochondrial DNA. In some embodiments, a nucleic acid sample is a DNA / RNA hybrid (e.g., an R-loop). In some embodiments, a nucleic sample includes complementary DNA (cDNA) obtained from the reverse transcription of RNA (e.g., total RNA or subtypes of RNA that have been selectively isolated (e.g., messenger RNA (mRNA), ribosomal RNA (rRNA), small nuclear RNA (snRNA), transfer RNA (tRNA), among others)). The nucleic acid sample may represent at least a portion of an organism’s genome (e.g., at least about 1%, 5%, 10%, 20%, 25%, 30%, 40%, 50%, 75%, 80%, 90%, 95%, 99%, or 100% of the organism’s genome). The nucleic acid sample may be a chromosome or chromatin. The target nucleic acid may include genomic DNA or cDNAs from a single cell. The nucleic acid sample may include nucleic acids from a plurality of haplotypes. In some embodiments, the target nucleic acid may be contained within a cell or a subcellular compartment.

[0113] A nucleic acid sample may be provided as a purified nucleic acid sample or as a partially purified nucleic acid sample (e.g., partially purified from a cell or a tissue (e.g., a lysate)). In some embodiments, a nucleic acid sample is about 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% pure. Nucleic acid purity may be evaluated, for example, by UV spectrum absorbance and the absorbance ratio at a wavelength of 260 nm to a wavelength of 280 nm. For instance, an absorbance ratio of about 1.8 is a generally accepted value for pure DNA, while an absorbance ratio of about 2.0 is a generally accepted value for pure RNA.

[0114] In some embodiments, a nucleic acid sample is provided in an amount of about 10 ng, about 15 ng, about 20 ng, about 25 ng, about 30 ng, about 35 ng, about 40 ng, about 45 ng, about 50 ng, about 55 ng, about 60 ng, about 65 ng, about 70 ng, about 75 ng, about 80 ng, about 85 ng, about 90 ng, about 95 ng, about 100 ng, about 105 ng, about 110 ng, about 115 ng, about 120 ng, about 125 ng, about 150 ng, about 175 ng, about 200 ng, about 225 ng, about 250 ng, about 275 ng, about 300 ng, about 325 ng, about 350 ng, about 375 ng, about 400 ng, about 425 ng, about 450 ng, about 475 ng, about 500 ng, about 525 ng, about 550 ng, about 575 ng, about 600 ng, about 625 ng, about 650 ng, about 675 ng, about 700 ng, about 725 ng, about 750 ng, about 775 ng, about 800 ng, about 825 ng, about 850 ng, about 875 ng, about 900 ng, about 925 ng, about 950 ng, about 975 ng, about 1 ,000 ng, or more.

[0115] In some instances, a nucleic acid sample is present in the tagmentation reaction at variable amounts. The sample may be present at low concentrations at about 1 to 10 ng / pL (e.g., about 10 ng / pL, about 9 ng / pL, about 8 ng / pL, about 7 ng / pL, about 6 ng / pL, about 5 ng / pL, about 4 ng / pL, PATENT

[0116] Attorney Docket No.: 51178-019WO2

[0117] about 3 ng / pL, about 2 ng / pL, about 1 ng / pL, or less than 1 ng / pL). The sample may be present at medium concentrations at about 10 ng / pL to 100 ng / pL (e.g., about 100 ng / pL, about 90 ng / pL, about 80 ng / pL, about 70 ng / pL, about 60 ng / pL, about 50 ng / pL, about 40 ng / pL, about 30 ng / pL, about 20 ng / pL, or about 10 ng / pL). Further, the sample may be present at high concentrations at about 100 ng / pL or higher (e.g., about 100 ng / pL, about 150 ng / pL, about 200 ng / pL, about 250 ng / pL, about 300 ng / pL, or higher).

[0118] In some embodiments, normalization of varied genomic DNA inputs is effectively achieved by precipitating DNA onto magnetic beads before tagmentation. This limits the amount of template accessible to the transposase and ensures limited variation in insert size irrespective of the ratio of transposase to genomic DNA. Such normalization strategy negates the need for quantification and normalization of input DNA before library construction and reduces variation in PCR yields allowing post-PCR pooling without quantification and normalization.

[0119] A synaptic complex (e.g., a transposome) is a complex that includes one or more transposases (e.g., 2, 3, or 4 transposases). Any of the transposases described herein may be used in the methods of the invention, including those described in further detail herein. For example, the transposase may be TnX, Tn3, Tn5, Tn9, Tn10, gamma-delta, Mu, piggyBac, Minos, Tc1 , or Sleeping Beauty transposase or a variant thereof that has increased enzymatic activity (e.g., a hyperactive variant). Other transposases are known in the art and may also be used in the invention.

[0120] A synaptic complex (e.g., a transposome) further includes one or more nucleic acid adapters in complex with each transposase, in which a nucleic acid adapter (i.e., a first nucleic acid adapter) is transposed to the 5’ end of a fragmented double-stranded nucleic acid sample in the tagmentation reaction. A nucleic acid adapter may be between 10 and 100 nucleotides, between about 10 and about 20 nucleotides, between about 15 and about 30 nucleotides, between about 25 and about 50 nucleotides, between about 40 and about 60 nucleotides, between about 50 and about 70 nucleotides, between about 60 and about 80 nucleotides, between about 70 and about 90 nucleotides, or between about 80 and about 100 nucleotides in length. In some embodiments, the nucleic acid adapter is 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, 25, 26, 27, 28, 29, 30, 31 , 32, 33, 34, 35, 36, 37, 38, 39, 40, 41 , 42, 43, 44, 45, 46, 47, 48, 49, 50, 51 , 52, 53, 54, 55, 56, 57, 58, 59, 60, 61 , 62, 63, 64, 65, 66, 67, 68, 69, 70, 71 , 72, 73, 74, 75, 76, 77, 78, 79, 80, 81 , 82, 83, 84, 85, 86, 87, 88, 89, 90, 91 , 92, 93, 94, 95, 96, 97, 98, 99, or 100 nucleotides in length.

[0121] In some embodiments, the sense strand of the nucleic acid adapter includes a free 5’-hydroxyl group and is not phosphorylated. In some embodiments, the sense strand of the nucleic acid adapter includes a free 3’-hydroxyl group. In some embodiments, the sense strand of the nucleic acid adapter includes a modification that is amenable to click chemistry such as an azide or an alkyne (e.g., a strain-promoted azido-alkyne or a metal-catalyzed azido-alkyne) for immobilization or for appending one or more additional molecules (e.g., a nucleic acid, a lipid, a glycan or carbohydrate, a peptide, a fluorophore or a fluorophore quencher, biotin, or a radiolabel). In some embodiments, the antisense strand of the nucleic acid adapter includes a phosphate group on the 5’ end. In some embodiments, the antisense strand of the nucleic acid adapter includes a modification at the 3’ end. In some embodiments, the modification at the 3’ end is an amino group modification, a phosphate group, PATENT

[0122] Attorney Docket No.: 51178-019WO2

[0123] a three-carbon chain (03), e.g., 3-hydroxy propyl, a dideoxynucleotide (e.g., a dideoxycytosine (ddC), a dideoxyguanosine (ddG), a dideoxythymidine (ddT), a dideoxyadenosine (ddA), or a dideoxyuridine (ddU)), an inverted nucleotide, or a locked nucleic acid. In some embodiments, 3’ end of the antisense strand includes an inverted nucleotide, e.g., T.

[0124] In some embodiments, the nucleic acid adapter includes a priming region (e.g., a nucleic acid sequence corresponding to a primer or the reverse complement thereto (e.g., a domain for primer binding)). The nucleic acid adapter added to the nucleic acid during the tagmentation reaction (e.g., the first nucleic acid adapter) is initially bound to an oligonucleotide to produce a double-stranded portion that includes a transposase binding site (TBS; e.g., a mosaic end sequence) that binds the nucleic acid adapter to the transposase of the synaptic complex (e.g., the transposome), e.g., for loading of the nucleic acid adapter into the synaptic complex. In some embodiments, the double stranded portion does not extend for the entire length of the adapter, resulting in a single-stranded portion on the 5’ end (e.g., an overhang at the 5’ end). In some embodiments, the nucleic acid adapter, e.g., in the single-stranded portion, includes a domain for primer binding or the reverse complement thereto. In some embodiments, a nucleic acid adapter includes a barcode, an affinity tag, and / or a reporter moiety. In some embodiments, the adapter does not include a barcode, an affinity tag, and / or a reporter moiety. In some embodiments, it can be advantageous for each

[0125] nucleic acid adapter to include at least one universal primer sequence or a reverse complement thereto (e.g., a domain for universal primer binding). Universal primers have nucleic acid sequences that are well-established in the art and can have various applications, such as use in amplifying, sequencing, and / or identifying one or more nucleic acid samples. In some embodiments, the first nucleic acid adapter includes, from 5’ to 3’, a priming region (e.g., an amplification or sequencing primer; e.g., a Read 1 or Read 2 primer) and a TBS, e.g., mosaic end sequence. In some embodiments, the nucleic acid adapter includes 5-methylated cytosines. In some embodiments, the nucleic acid adapter includes unmethylated cytosines. In some embodiments, the nucleic acid adapter does not include a priming region and is suitable for library sample preparation that does not require amplification (e.g., a PCR-free preparation).

[0126] A tagmentation reaction of the invention generates a collection of nucleic acid fragments (e.g., double-stranded nucleic acid fragments or partially double-stranded nucleic acid fragments) from the nucleic acid sample, in which a nucleic acid fragment includes a nucleic acid adapter (e.g., a first nucleic acid adapter) transposed to a 5’ end. In some embodiments, the nucleic acid fragment includes a portion that is double-stranded (e.g., on the 3’ end) and a portion that is single-stranded (e.g., on the 5’ end) that further includes a primer sequence. In some embodiments, the nucleic acid fragment includes a portion that is double-stranded and a portion that is single-stranded on the 5’ end of each strand (e.g., a sticky end). An exemplary nucleic acid fragment following a tagmentation reaction is illustrated in FIG. 1A.

[0127] In some embodiments, a nucleic acid fragment is between about 5 to about 20,000 nucleotides in length. In some embodiments, a nucleic acid fragment is about 5 to about 10 nucleotides in length. In some embodiments, a nucleic acid fragment is about 10 to about 20 nucleotides in length. In some embodiments, a nucleic acid fragment is about 20 to about 40 PATENT

[0128] Attorney Docket No.: 51178-019WO2

[0129] nucleotides in length. In some embodiments, a nucleic acid fragment is about 30 to about 50 nucleotides in length. In some embodiments, a nucleic acid fragment is about 50 to about 70 nucleotides in length. In some embodiments, a nucleic acid fragment is about 70 to about 90 nucleotides in length. In some embodiments, a nucleic acid fragment is about 80 to about 100 nucleotides in length. In some embodiments, a nucleic acid fragment is about 100 to about 150 nucleotides in length. In some embodiments, a nucleic acid fragment is about 150 to about 200 nucleotides in length. In some embodiments, a nucleic acid fragment is about 200 to about 250 nucleotides in length. In some embodiments, a nucleic acid fragment is about 250 to about 300 nucleotides in length. In some embodiments, a nucleic acid fragment is about 300 to about 350 nucleotides in length. In some embodiments, a nucleic acid fragment is about 350 to about 400 nucleotides in length. In some embodiments, a nucleic acid fragment is about 400 to about 450 nucleotides in length. In some embodiments, a nucleic acid fragment is about 450 to about 500 nucleotides in length. In some embodiments, a nucleic acid fragment is about 500 to about 750 nucleotides in length. In some embodiments, a nucleic acid fragment is about 500 nucleotides to about 1 ,000 nucleotides in length. In some embodiments, a nucleic acid fragment is about 1 ,000 to about 2,000 nucleotides in length. In some embodiments, a nucleic acid fragment is about 2,000 to about 5,000 nucleotides in length. In some embodiments, a nucleic acid fragment is about 5,000 to about 10,000 nucleotides in length. In some embodiments, a nucleic acid fragment is about 10,000 to about 15,000 nucleotides in length. In some embodiments, a nucleic acid fragment is about 15,000 to about 20,000 nucleotides in length.

[0130] / ' / . Methods of Generating a Single-Stranded Nucleic Acid Fragment from a Double-Stranded Nucleic Acid Fragment

[0131] Methods of the invention further include denaturing the nucleic acid fragment (e.g., the nucleic acid fragment produced via tagmentation) to generate a single-stranded nucleic acid fragment (e.g., a single-stranded DNA fragment). Methods of nucleic acid denaturation are routine in the art and may include an application of heat (e.g., temperatures of about 70°C, about 75°C, about 80°C, about 85°C, about 90°C, about 95°C, about 100°C, or greater) and / or exposure to one or more chemical reagents (e.g., urea, sodium hydroxide, guanidine hydrochloride, dimethyl sulfoxide, propylene glycol, or formamide) in a solution such as a basic solution (e.g., a solution that has a pH>7.4) in an amount that can effectively denature a nucleic acid.

[0132] Methods of the invention include providing a plurality of blocking molecules (e.g., blocking molecules described herein) prior to or during the denaturation step, e.g., prior to tagmentation or after tagmentation. A plurality of blocking molecules is provided prior to denaturing the nucleic acid fragment such that the plurality of blocking molecules binds to the generated single-stranded nucleic acid fragment, e.g., to prevent or reduce the likelihood of secondary structure formation (e.g., internal loops, hairpins, stems, and / or bulges) in the generated single-stranded nucleic acid fragment, as compared to a single-stranded nucleic acid fragment not exposed to a plurality of blocking molecules.

[0133] In some embodiments, a plurality of blocking molecules is provided to a solution containing a nucleic acid fragment at a concentration between about 1 pM to about 1 M, such that the final PATENT

[0134] Attorney Docket No.: 51178-019WO2

[0135] concentration of blocking molecules in a reaction is between about 100 nM to about 100 pM. In some embodiments, a plurality of blocking molecules is provided to a solution containing the nucleic acid fragment (e.g., a double-stranded nucleic acid fragment prior to denaturation or a single-stranded nucleic acid fragment after denaturation) at a concentration of about 1 pM to about 5 pM, about 5 pM to about 10 pM, about 10 pM to about 50 pM, about 50 pM to about 100 pM, about 100 pM to about 200 pM, about 200 pM to about 300 pM, about 300 pM to about 400 pM, about 400 pM to about 500 pM, about 500 pM to about 600 pM, about 600 pM to about 700 pM, about 700 pM to about 800 pM, about 800 pM to about 900 pM, about 900 pM to about 1 M. In some embodiments, a plurality of blocking molecules is provided to a solution containing the nucleic acid fragment at a concentration of about 1 pM, about 5 pM, about 10 pM, about 20 pM, about 30 pM, about 40 pM, about 50 pM, about 60 pM, about 70 pM, about 80 pM, about 90 pM, about 100 pM, about 110 pM, about 120 pM, about 130 pM, about 140 pM, about 150 pM, about 160 pM, about 170 pM, about 180 pM, about 190 pM, about 200 pM, about 250 pM, about 300 pM, about 350 pM, about 400 pM, about 450 pM, about 500 pM, about 550 pM, about 600 pM, about 650 pM, about 700 pM, about 750 pM, about 800 pM, about 850 pM, about 900 pM, about 950 pM, or about 1 M.

[0136] In some embodiments, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, 25, 26, 27, 28, 29, 30, 31 , 32, 33, 34, 35, 36, 37, 38, 39, 40, 41 , 42, 43, 44, 45, 46, 47, 48, 49, 50, or more blocking molecules (e.g., blocking proteins or blocking oligonucleotides) simultaneously contact (e.g., bind or hybridize to) a single-stranded nucleic acid fragment, e.g., to reduce the likelihood of secondary structure formation across the length of the single-stranded nucleic acid fragment. An example of a plurality of blocking molecules is illustrated in FIG. 1B.

[0137] Hi. Methods of Adding a Second Nucleic Acid Adapter to a Single-Stranded Nucleic Acid Fragment via Ligation

[0138] Methods of the invention also include appending a second nucleic acid adapter to the 3’ end of a nucleic acid fragment (e.g., a single-stranded nucleic acid fragment). In some embodiments, the second nucleic adapter is appended to the 3’ end of the nucleic acid fragment (e.g., the singlestranded nucleic acid fragment) in a splint ligation reaction. In some embodiments, a second nucleic acid adapter is appended to the 3’ end of the nucleic acid fragment (e.g., the single-stranded nucleic acid) fragment by exposing the nucleic acid fragment (e.g., the single-stranded nucleic acid fragment in complex with a plurality of blocking molecules) to a splint oligonucleotide that includes the second nucleic acid adapter. The splint oligonucleotide contacts the nucleic acid fragment (e.g., singlestranded nucleic acid fragment) by hybridizing to the 3’ end of the fragment via a degenerate nucleic acid sequence at the 3’ end of the splint oligonucleotide and places the second nucleic acid adapter in close proximity to (e.g., adjacent to) the 3’ end of the nucleic acid fragment (e.g., the single-stranded nucleic acid fragment). The second nucleic acid adapter is then covalently bound to the 3’ end of the nucleic acid fragment (e.g., the single-stranded nucleic acid fragment) via a ligase. An exemplary splint ligation reaction is illustrated in FIG. 1C.

[0139] A splint oligonucleotide is an oligonucleotide that is double-stranded in part and includes a nucleic acid adapter (e.g., a second nucleic acid adapter) on one strand and, on the other, strand PATENT

[0140] Attorney Docket No.: 51178-019WO2

[0141] includes the reverse complement of at least a portion of the nucleic acid adapter (e.g., the second nucleic acid adapter) and a degenerate nucleic acid sequence on the 3’ end. In some embodiments the degenerate nucleic acid sequence is between about 5 to about 25 nucleotides in length. In some embodiments, the degenerate nucleic acid sequence is about 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, or 25 nucleotides in length. In some embodiments, a population of splint oligonucleotides having degenerate nucleic acid sequences of different lengths is suitable for methods of adding a second nucleic acid adapter to a nucleic acid fragment (e.g., a single-stranded nucleic acid fragment). In some embodiments, the splint oligonucleotide or nucleic acid adapter therein includes 5-methylated cytosines. In some embodiments, the splint oligonucleotide or nucleic acid adapter therein includes unmethylated cytosines.

[0142] In general, a degenerate nucleic acid sequence of a splint oligonucleotide can hybridize to the 3’ end of a nucleic acid fragment (e.g., a single-stranded nucleic acid fragment) to bring the nucleic acid adapter in proximity to the fragment (e.g., adjacent to a single-stranded nucleic acid fragment). In some embodiments, the degenerate nucleic acid sequence has complementarity to a portion of equal length of the 3’ end of a nucleic acid fragment (e.g., a single-stranded nucleic acid fragment). In some embodiments, the degenerate nucleic acid sequence has about 50% complementarity, about 55% complementarity, about 60% complementarity, about 65% complementarity, about 70% complementarity, about 75% complementarity, about 80% complementarity, about 85% complementarity, about 90% complementarity, about 95% complementarity, or 100% complementarity to a portion of equal length of the 3’ end of a nucleic acid fragment (e.g., a singlestranded nucleic acid fragment). In some embodiments, inosine, 5-nitroindole, 2,6-diaminopurine, 2'-deoxy-pseudoisocytidine, 2'-deoxy-pseudoisoguanosine or other modified bases are incorporated into a splint oligonucleotide, e.g., at the 3’ end, to facilitate ligation, especially in situations where mismatches need to be tolerated at the ligation junction. Inosine, in particular, is often used due to its ability to pair with A, C, T, or G, which can help bridge mismatches and increase the efficiency of ligation.

[0143] Following contact of a splint oligonucleotide having a degenerate sequence complementary to the single-stranded nucleic acid fragment, the second nucleic acid adapter is appended to the 3’ end of the single-stranded nucleic acid fragment. In some embodiments, the second nucleic acid adapter is enzymatically linked via a covalent bond to the 3’ end of the single stranded nucleic acid fragment by a ligase (e.g., a DNA or an RNA ligase). A ligase is used to link (e.g., covalently bind) a 5'-phosphorylated nucleic acid molecule to the 3'-hydroxyl group of a nucleic acid (e.g., a linear nucleic acid) forming a new phosphodiester linkage. In some embodiments, the ligase is a splint ligase, such as, e.g., SplintR® ligase, or the ligase is an RNA ligase II, a T4 RNA ligase, or a T4 DNA ligase. In other embodiments, ligation is carried out with Taq ligase or other template-dependent ligases, singlestranded ligase, or non-template-dependent ligases other than T4 DNA ligase. Similarly, attachment of adapters may be performed using a gap-fill ligation reaction where denatured DNA is extended with a polymerase and ligated in a template-dependent manner.

[0144] The ligation reaction that includes the single-stranded nucleic acid fragments, a splint oligonucleotide, a plurality of blocking molecules, a ligase, or a combination thereof in a suitable PATENT

[0145] Attorney Docket No.: 51178-019WO2

[0146] reaction buffer may be performed at a temperature for a duration of time suitable for the ligation reaction to occur. In some embodiments, the ligation reaction is performed at ambient room temperature (e.g., about 22-28°C). In some embodiments, the ligation reaction is performed at about 15°C, about 20°C, at about 25°C, at about 30°C, or at about 35°C. In some embodiments, the ligation reaction is performed for 10-20 minutes, for 15-25 minutes, for 20-30 minutes, for 30-45 minutes, or for 45-60 minutes. In some embodiments, the ligation reaction is performed for 1 , 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or 16 hours.

[0147] In some instances, the ligation reaction is carried out in the presence of a ligation enhancer, which is a reagent that brings the reaction components (e.g., the single-stranded nucleic acid fragments, the splint nucleotide, the plurality of blocking molecules, and the ligase) in close proximity to enhance the efficiency of the ligation reaction. A ligation enhancer may be a plurality of small molecules or a plurality of macromolecules (e.g., polymers) that induce molecular crowding. As an example, polyethylene glycol (PEG) may be included in the ligation reaction at a concentration of about 5% to about 15% v / v. Exemplary ligation enhancers are described elsewhere such as U.S. Patent No. 8,697,408, which is hereby incorporated by reference.

[0148] In some embodiments, the ligation reaction occurs in a total volume between about 20 pL and 100 pL. In some embodiments, the ligation reaction occurs in a total volume of 20, 21 , 22, 23, 24, 25, 26, 27, 28, 29, 30, 31 , 32, 33, 34, 35, 36, 37, 38, 39, 40, 41 , 42, 43, 44, 45, 46, 47, 48, 49, 50, 51 , 52, 53, 54, 55, 56, 57, 58, 59, 60, 61 , 62, 63, 64, 65, 66, 67, 68, 69, 70, 71 , 72, 73, 74, 75, 76, 77, 78, 79, 80, 81 , 82, 83, 84, 85, 86, 87, 88, 89, 90, 91 , 92, 93, 94, 95, 96, 97, 98, 99, or 100 pL.

[0149] B. Methods of Amplifying an Engineered Nucleic Acid Sample

[0150] In some embodiments, the methods of the invention do not require amplification. In some embodiments, an amplification-free (e.g., PCR-free) library may be constructed by transposing and ligating full length adapter sequences. Libraries generated without PCR could also be generated without purification by employing a heat-kill after ligation. Similarly, libraries generated without PCR could be generated without purification through chelating agents such as EDTA, detergents such as sodium dodecyl sulfate, chaotropes such as guanidine, or denaturants such as sodium hydroxide.

[0151] In some embodiments, the methods of the invention further include amplification of the engineered nucleic acid sample (e.g., nucleic acid fragments appended with a nucleic acid adapter (e.g., two nucleic acid adapters, e.g., a first nucleic acid adapter at the 5’ end and a second nucleic acid adapter at the 3’ end)). An exemplary engineered nucleic acid fragment appended with two nucleic acid adapters is illustrated in FIG. 1D. The engineered nucleic acid sample may be amplified by any nucleic acid amplification method known in the art. In some embodiments, the engineered nucleic acid sample is amplified by polymerase chain reaction (PCR), multiple annealing and loopingbased amplification cycles (MALBAC), multiple displacement amplification (MDA), ligase chain reaction (LCR), loop mediated isothermal amplification (LAMP), rolling circle amplification (RCA), or strand displacement amplification (SDA).

[0152] For example, a PCR reaction generates a plurality of amplicons (e.g., amplified products) of the nucleic acid fragments using amplifier oligonucleotides as primers. In some instances, the PATENT

[0153] Attorney Docket No.: 51178-019WO2

[0154] amplicons are generated in a PCR reaction with a first primer and a second primer, wherein the first primer and the second primer bind to respective priming regions added via the respective nucleic acid adapters. In some embodiments, the tagmentation reaction adds a primer sequence at the 5’ end of the fragment, and the ligation reaction adds a primer binding sequence to the 3’ end. When a primer binds to the primer binding sequence and is extended, the amplicon produced now includes a primer binding region at its 3’ end that is the reverse complement of the primer sequence added during tagmentation. A primer can then bind to this primer binding region and be extended. The primers used for extending the original strand and its reverse complement can be the same or different. Use of different primers can allow for only fragments that have undergone tagmentation and ligation to be amplified. In some embodiments, the first primer and / or the second primer selected for amplification are universal primers (e.g., primers that have nucleic acid sequences that are well-known and commonly utilized in the art), and the first and / or second nucleic acid adapter include priming regions corresponding to universal primers (e.g., nucleic acid sequences corresponding to universal primer nucleic acid sequences or reverse complements thereof).

[0155] In some embodiments, the first primer and / or second primer of an amplification reaction includes a barcode, an affinity tag, and / or a reporter moiety such that the amplicons of the nucleic acid fragments include a barcode, an affinity tag, and / or a reporter moiety following amplification. In some embodiments, a barcode, an affinity tag, and / or a reporter moiety permits attachment of a nucleic acid to a surface (e.g., a bead, a resin, a surface of a reaction vessel (e.g., a microplate), or a flow cell). Exemplary nucleotide sequences that permit attachment of nucleic acids (e.g., amplicons) to a surface include P5 and P7 sequences.

[0156] In some embodiments, the method includes surveying the methylation state of one or more cytosine nucleotides in the nucleic acid sample (e.g., DNA). Detecting methylated versus unmethylated cytosines in a nucleic acid sample may include the step of contacting the nucleic acid with sodium bisulfite prior to amplification (e.g., prior to tagmentation, prior to splint ligation, or prior to amplification), which deaminates unmethylated cytosines to uracil and leaves methylated cytosines (e.g., 5-mC or 5-hmC) intact. Upon conversion to uracil, unmethylated cytosines are detected as thymine nucleotides following sequencing analysis after amplification (e.g., PCR).

[0157] In some instances, the amplicons include (a) a nucleic acid sequence including a first priming region, a first amplifier barcode region (e.g., a pool tag, e.g., an I5 or an I7 tag), a homologous sequence of a first nucleic acid fragment, a complement sequence of the second amplifier barcode region (e.g., a pool tag, e.g., an I5 or an I7 tag), and a second priming region (e.g., the reverse complement of a second primer); and (b) the reverse complement sequence thereof. In some instances, the PCR reaction includes 1-35 cycles (e.g., 1 , 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, or 35 cycles). In some instances, the PCR reaction includes more than 35 cycles. In some instances, the PCR reaction may be monitored by tracking the cycle threshold (Ct) via fluorescence detection using well-established methods in the art. In some instances, the resulting sequence library following transposition, tagmentation, ligation, and amplification may be evaluated for quality (e.g., size distribution, purity, concentration, among other quality considerations). In some instances, the sequence library may be PATENT

[0158] Attorney Docket No.: 51178-019WO2

[0159] assessed for the size distribution of the resulting fragments by gel electrophoresis compared to a reference sample (e.g., a DNA ladder).

[0160] C. Methods of Sequencing an Engineered Nucleic Acid Sample

[0161] Methods described herein can be used for preparing engineered nucleic acid samples for determining the nucleic acid sequences or amplicons thereof via nucleic acid sequencing (e.g., nextgeneration sequencing (NGS)) or other methods known in the art. In some instances, sequencing can be performed by various systems that are currently available, e.g., a sequencing system by Pacific Biosciences (PACBIO®), ILLUMINA®, Oxford NANOPORE®, Element Biosciences, or ThermoFisher (ION TORRENT®). Alternatively or in addition, sequencing may be performed using nucleic acid amplification, sequencing by ligation (e.g., SOLiD or polony-based sequencing), sequencing by synthesis (e.g., Illumina dye sequencing, single-molecule real-time (SMRT) sequencing, or pyrosequencing), nanopore sequencing, polymerase chain reaction (PCR) (e.g., digital PCR, quantitative PCR, or real time PCR), or isothermal amplification. In some instances, the sequencing oligonucleotides and amplicons described herein can be uniquely identified based on the nucleic acid sequences of the nucleic acid fragment and the nucleic acid sequences of the nucleic acid adapters or amplifier barcode regions (e.g., pool tags, e.g., I5 or I7 tags) of the sequencing oligonucleotides or amplicons, respectively.

[0162] D. Reaction Considerations

[0163] For the reactions of the methods described herein, various ranges of suitable reaction conditions can be used in the processes, include but are not limited to, substrate loading, co-substrate loading, pH, temperature, buffer, solvent system, a cofactor, polypeptide loading, and reaction time. Further suitable reaction conditions for carrying out the process for generation of engineered nucleic acid samples from a nucleic acid sample (e.g., a DNA sample; e.g., genomic DNA, mitochondrial DNA, or cDNA) can be readily optimized in view of the guidance provided herein by routine experimentation that includes, but is not limited to, contacting a protein or protein complex (e.g., a transposase, a ligase, or a polymerase) and one or more substrate compounds (e.g., nucleic acid samples) under experimental reaction conditions of concentration, pH, temperature, and solvent conditions, and detecting the product compound.

[0164] In some instances, one or more enzymes (e.g., a transposase, a ligase, or a polymerase) are inactivated by the application of heat or an inhibitor following a chemical reaction (e.g., tagmentation, ligation, or amplification). In some instances, one or more enzymes are inactivated by changes in buffer or solvent condition (e.g., changes in pH or changes in salt concentration). In some instances, the nucleic acid sample is purified or partially purified following a chemical reaction. In some instances, the nucleic acid sample is concentrated or diluted following a chemical reaction.

[0165] In some instances, one or more enzymes (e.g., a transposase, a ligase, or a polymerase) are immobilized to improve reaction efficiency, enzyme activity, or sample recovery. In some instances, an enzyme is immobilized on a solid support, a surface, a resin, a membrane, a bead, a nanoparticle, a flow cell, a biological cofactor (e.g., a protein, an antibody or antibody fragment, a peptide, an PATENT

[0166] Attorney Docket No.: 51178-019WO2

[0167] oligonucleotide, an aptamer), or a combination thereof. In some instances, an enzyme is immobilized through an irreversible or a reversible reaction. In some instances, an enzyme is immobilized by cross-linking, adsorption, entrapment, encapsulation, click chemistry reactions, ionic bonding, covalent bonding, hydrophobic interactions, affinity interaction, among other non-limiting methods of immobilization.

[0168] In some instances, further processing steps are used for library preparation, including gap filling, normalization, pooling of tagged fragment samples, and / or ligation of additional adapters to create dual-tagged libraries. A variety of parameters may be manipulated to create sequencing libraries with various characteristics. For example, varying the concentration of transposase and a first nucleic acid adapter can change the fragment length. Additionally, concentration of a second nucleic acid adapter and ligase can alter reaction efficiency. Copy number and sample coverage may also be influenced or controlled using methods known to those of skill in the art.

[0169] Sample quality (e.g., quality of an engineered nucleic acid sample and / or amplicons thereof) generated by one or more methods described herein may be evaluated by methods of nucleic acid sequencing (e.g., massively parallel sequencing, e.g., next-generation sequencing) or gel electrophoresis.

[0170] E. Exemplary Embodiments of the Methods

[0171] A 24 pL tagmentation reaction may be prepared with 100 ng of human genomic DNA diluted to 12.5 ng / pL in 1X TE. Next, 8 pL of diluted genomic DNA may be mixed with the following: 2.4 pL 10X T4 DNA Ligase Buffer, 2 pL of 10 pM 3’-ddC-blocked random nonamers, 8 pL 125 nM synaptic complex containing Illumina Readl sequencing primer and 5.6 pL nuclease-free water. The reaction may be incubated for 15 minutes at 55 °C to allow tagmentation to occur. Next, the reaction may be heat inactivated and denatured at 95 °C for 5 minutes, followed by a 2-minute incubation at 4°C. A ligation mixture containing the following components may added directly to heat-denatured DNA: 26 pL Blunt / TA Ligation Mix (New England Biolabs) and 2 pL of 50 pM Illumina Read2 sequencing adapter (IDT) containing an 8-base 3’-ddC blocked degenerate overhang and 5’ phosphorylation (IDT). Ligation may be performed at 25 °C for 60 minutes, followed by purification with 1.25X MAGwise beads (seqWell, Inc.) and elution into 22.5 pl 10 mM Tris-HCI, pH 8.0. Purified ligation products may then be amplified using 25 pL 2X Kapa HiFi Readymix (Roche), 2.5 pL Indexed P5 / P7 primer mix at 10 pM each (IDT) under the following conditions: 95 °C for 3 minutes, 10 Cycles: 98 °C for 30s, 60 °C for 30s, and 72 °C for 30s, final extension for 3 minutes at 72°C, followed by hold at 4°C. Amplification produces PCR products that have P5 and P7 sequences on the 5’ end and the 3’ end, respectively, which may be used to attach the PCR products to a flowcell for next-generation sequencing. PCR products may be purified using 1.25X MAGwise beads (seqWell, Inc.) and eluted into 22 pl 10 mM Tris-HCI, pH 8.0 prior to sequencing (e.g., next-generation sequencing). PATENT

[0172] Attorney Docket No.: 51178-019WO2

[0173] II. Compositions

[0174] A. Transposases

[0175] In some embodiments, the tagmentation reaction for fragmenting and appending a first nucleic acid adapter to a nucleic acid is performed with a synaptic complex (e.g., a transposome) that includes a Tn5 transposase. Tn5 is a well-studied transposition system derived from E. co / / 'which can be used in the context of the invention (see, e.g., Reznikoff, Mol. Microbiol. 47(5):1199-206, 2003). NCBI Accession No. U00004 provides the nucleic acid sequence of the E. coli n5 transposon. Tn5 encodes the transposase TnpA (UniProt Accession No. Q46731), which is also referred to herein as Tn5 transposase. The amino acid sequence of wild-type Tn5 transposase is shown below:

[0176] MITSALHRAADWAKSVFSSAALGDPRRTARLVNVAAQLAKYSGKSITISSEGSEAMQEGA YRFIRNPNVSAEAIRKAGAMQTVKLAQEFPELLAIEDTTSLSYRHQVAEELGKLGSIQDK SRGWWVHSVLLLEATTFRTVGLLHQEWWMRPDDPADADEKESGKWLAAAATSRLRMGSMM SNVIAVCDREADIHAYLQDKLAHNERFVVRSKHPRKDVESGLYLYDHLKNQPELGGYQIS IPQKGVVDKRGKRKNRPARKASLSLRSGRITLKQGNITLNAVLAEEINPPKGETPLKWLL LTSEPVESLAQALRVIDIYTHRWRIEEFHKAWKTGAGAERQRMEEPDNLERMVSILSFVA VRLLQLRESFTLPQALRAQGLLKEAEHVESQSAETVLTPDECQLLGYLDKGKRKRKEKAG SLQWAYMAIARLGGFMDSKRTGIASWGALWEGWEALQSKLDGFLAAKDLMAQGIKI

[0177] (SEQ ID NO: 1)

[0178] Biologically active variants of Tn5 transposase, including variants with amino acid substitutions, insertions, and / or deletions, may be used in the compositions of the invention.

[0179] Biologically active Tn5 transposase variants with amino acid substitutions are known in the art. In some instances, a biologically active variant has an enhanced transposition rate relative to wild-type Tn5, and is thus considered hyperactive (see, e.g., U.S. Pat. Nos. 5,965,443; 5,925,545; and 6,159,736). For example, substitution of a lysine residue at amino acid 54 in place of the glutamic acid found in wild-type Tn5 transposase (E54K) has been shown to improve the avidity of the modified transposase for OE termini and to increase the transposition rate approximately 10-fold. Other mutations that have been associated with Tn5 transposase hyperactivity include a substitution of amino acid 372 (leucine) with proline (L372P) and a substitution of amino acid 56 (methionine) with alanine (M56A). The substitution mutations may be relative to the exemplary wild-type sequence of Tn5 transposase shown in SEQ ID NO: 1. A biologically active variant may include any combination of the preceding substitution mutations. For example, in some instances, the Tn5 transposase includes the substitution mutations E54K, M56A, and L372P. In other instances, the Tn5 transposase includes the substitution mutations E54K and L372P. Hyperactive Tn5 transposase proteins are commercially available, for example, Ez-Tn5™ transposase and Ez-Tn5™ Custom Transposome Construction Kits (Epicentre).

[0180] It is generally understood that to carry out tagmentation, Tn5 transposases bind nucleic acid adapters via transposase binding sites (TBSs), which are a pair of inverted repeat nucleotide sequences. The inverted repeat sequences of the Tn5 transposase binding sites are referred to as the outside end (OE) (5’-CTGACTCTTATACACAAGT-3’ (SEQ ID NO: 2)) and inside end (IE) (5’-CTGTCTCTTGATCAGATCT-3’ (SEQ ID NO: 3)) (see, e.g., U.S. Pat. No. 5,965,443). Biologically active variants of a Tn5 TBS may be used, including end sequence variants that are associated with higher rates of transposition, for example, the hyperactive hybrid of the outside and inside ends (also PATENT

[0181] Attorney Docket No.: 51178-019WO2

[0182] referred to as “mosaic end” (ME)) 5’-CTGTCTCTTATACACATCT-3’ (SEQ ID NO: 4), which differs from the wild-type OE sequence at positions 4, 17, and 18, as well as 5’-CTGTCTCTTATACAGATCT-3’ (SEQ ID NO: 5), which differs from the wild-type OE sequence at positions 4, 15, 17, and 18 (see, e.g., U.S. Pat. No. 5,925,545). In some instances, a nucleic acid of the invention may include one or more Tn5 TBSs having a nucleic acid sequence selected from SEQ ID NOs: 2-5 and / or a biologically active variant thereof.

[0183] In some embodiments, the tagmentation reaction for fragmenting and appending a first nucleic acid adapter to a nucleic acid is performed with a synaptic complex (e.g., a transposome) that includes a TnX transposase, which is a commercially available high-performance transposase developed by seqWell Inc. (Beverly, MA, USA) and Codexis (Redwood City, CA, USA).

[0184] In some embodiments, the tagmentation reaction for fragmenting and appending a first nucleic acid adapter to a nucleic acid is performed with a synaptic complex (e.g., a transposome) that includes a transposase selected from Tn3, Tn9, Tn10, gamma-delta, Mu, piggyBac, Minos, Tc1 , and Sleeping Beauty transposase or a variant thereof that has increased enzymatic activity (e.g., a hyperactive variant). Other transposases are known in the art and may also be used in the invention.

[0185] B. Blocking Molecules

[0186] In some embodiments, a plurality of blocking molecules includes a protein. In some embodiments, the protein is thermostable and resists unfolding at temperatures of about 60°C, about 65°C, about 70°C, or about 75°C. In some embodiments, the protein is highly thermostable and resists unfolding at temperatures greater than 75°C or at temperatures of about 80°C, about 85°C, about 90°C, about 95°C, or about 100°C or more.

[0187] A blocking protein may be a prokaryotic protein (e.g., a bacterial or archaeal protein) or eukaryotic protein that binds a single-stranded portion of a nucleic acid (e.g., a single-stranded nucleic acid binding protein (SSB)). Examples of blocking proteins include E. coli SSB, E. coli RecA, Extreme Thermostable Single-Stranded DNA Binding Protein (ET SSB), Thermus thermophilus (Tth) RecA, T4 Gene 32 Protein, replication protein A (RPA — a eukaryotic SSB), among others. ET SSB, Tth RecA, E. co / / ' RecA, T4 Gene 32 Protein, as well buffers and detailed protocols for preparing SSB-bound ssNA using such SSBs are commercially available (e.g., New England Biolabs).

[0188] In some embodiments, a plurality of blocking molecules includes single-stranded oligonucleotides that have different sequences (e.g., random nucleic acid sequences) to maximize the number of binding sites (e.g., regions of hybridization) to a single-stranded nucleic acid fragment. A single-stranded oligonucleotide has complementarity to and can hybridize a portion of the singlestranded nucleic acid fragment. In some embodiments, a single-stranded oligonucleotide has about 30% complementarity to a portion of equal length of the single-stranded nucleic acid fragment. In some embodiments, a single-stranded oligonucleotide has about 40% complementarity to a portion of equal length of the single-stranded nucleic acid fragment. In some embodiments, a single-stranded oligonucleotide has about 50% complementarity to a portion of equal length of the single-stranded nucleic acid fragment. In some embodiments, a single-stranded oligonucleotide has about 60% complementarity to a portion of equal length of the single-stranded nucleic acid fragment. In some PATENT

[0189] Attorney Docket No.: 51178-019WO2

[0190] embodiments, a single-stranded oligonucleotide has about 70% complementarity to a portion of equal length of the single-stranded nucleic acid fragment. In some embodiments, a single-stranded oligonucleotide has about 80% complementarity to a portion of equal length of the single-stranded nucleic acid fragment. In some embodiments, a single-stranded oligonucleotide has about 90% complementarity to a portion of equal length of the single-stranded nucleic acid fragment. In some embodiments, a single-stranded oligonucleotide has 100% complementarity to a portion of equal length of the single-stranded nucleic acid fragment.

[0191] In some embodiments, the plurality of blocking molecules includes single-stranded oligonucleotides of the same length. In some embodiments, the plurality of blocking molecules includes single-stranded oligonucleotides of different or mixed lengths.

[0192] In some embodiments, a single-stranded oligonucleotide that is used as a blocking molecule is about 5 to about 500 nucleotides in length (e.g., about 5 to about 10, about 5 to about 25 nucleotides, about 25 to about 50 nucleotides, about 50 to about 75 nucleotides, about 75 to about 100 nucleotides, about 100 to about 150 nucleotides, about 150 to about 200 nucleotides, about 200 to about 250 nucleotides, about 250 to about 300 nucleotides, about 300 to about 350 nucleotides, about 350 to about 400 nucleotides, about 400 to about 450 nucleotides, or about 450 to about 500 nucleotides in length). In some embodiments, a single-stranded oligonucleotide is 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, 25, 26, 27, 28, 29, 30, 31 , 32, 33, 34, 35, 36, 37, 38, 39, 40, 41 , 42, 43, 44, 45, 46, 47, 48, 49, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 105, 110, 115, 120, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490 or 500 nucleotides in length.

[0193] In some embodiments, a single-stranded oligonucleotide that is used as a blocking molecule has GC-content of about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, or about 90%. In some embodiments, the plurality of blocking molecules includes single-stranded oligonucleotides that have the same percentage of GC-content or different percentages of GC-content. In other embodiments, the plurality of blocking molecules includes singlestranded oligonucleotides that have similar percentages of GC-content such that the single-stranded oligonucleotides have relative GC-content that fits in a desired range (e.g., from about 10% to about 20%, from about 20% to about 30%, from about 30% to about 40%, from about 40% to about 50%, from about 50% to about 60%, from about 60%, to about 70%, from about 70% to about 80%, or from about 80% to about 90%).

[0194] In some embodiments, a single-stranded oligonucleotide that is used as a blocking molecule includes a 3’ end modification to prevent or reduce the likelihood of nucleic acid extension by a polymerase or covalent biding by a ligase, as compared to an unmodified single-stranded oligonucleotide having the same nucleotide sequence. In some embodiments, the 3’ end modification is a 3’-dideoxynucleotide (e.g., a dideoxycytosine (ddC), a dideoxyguanosine (ddG), a dideoxythymidine (ddT), a dideoxyadenosine (ddA), or a dideoxyuridine (ddll)). In some embodiments, the 3’ end modification is an amino group. In some embodiments, the 3’ end PATENT

[0195] Attorney Docket No.: 51178-019WO2

[0196] modification is a phosphate group. In some embodiments, the 3’ end modification is an inverted nucleotide. In some embodiments, the 3’ end modification is a three-carbon chain (03), e.g., 3-hydroxy propyl. In some embodiments, the 3’ end modification is a locked nucleic acid.

[0197] In one example, random oligonucleotides comprising 3’ end modifications described above (e.g., a population of oligonucleotides with different sequences comprising 3’ end modifications described above) that are 9 nucleotides or fewer in length (e.g., 5, 6, 7, 8, or 9 nucleotides) may be used as blocking molecules to reduce the likelihood of secondary structure formation in a generated single-stranded nucleic acid fragment. In other instances, non-degenerate or semi-degenerate base compositions of variable, fixed, or mixed lengths may be used as blocking molecules.

[0198] In some embodiments, more than one (e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 , 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, or more) blocking molecules (e.g., blocking proteins or blocking oligonucleotides) can bind to a single-stranded nucleic acid or single-stranded portion of a nucleic acid.

[0199] In some embodiments, about 10% to about 90% of the nucleotides of a single-stranded nucleic acid bind to a plurality of blocking molecules. In some embodiments, about 10% to about 50%, about 25% to about 75%, or about 50% to 90%, e.g., about 10%, about 15%, about 20%, about 25%, about 30%, about 35%, about 40%, about 45%, about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, or about 90%, of the nucleotides of a single-stranded nucleic acid bind to a plurality of blocking molecules.

[0200] C. Nucleic Acid Adapters

[0201] In some embodiments a nucleic acid adapter (e.g., the first or second nucleic acid adapter) may be between 10 and 100 nucleotides in length. The nucleic acid adapter may be between 10 and 100 nucleotides, between about 10 and about 20 nucleotides, between about 15 and about 30 nucleotides, between about 25 and about 50 nucleotides, between about 40 and about 60 nucleotides, between about 50 and about 70 nucleotides, between about 60 and about 80 nucleotides, between about 70 and about 90 nucleotides, or between about 80 and about 100 nucleotides in length. In some embodiments, the nucleic acid adapter is 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21 , 22, 23, 24, 25, 26, 27, 28, 29, 30, 31 , 32, 33, 34, 35, 36, 37, 38, 39, 40, 41 , 42, 43, 44, 45, 46, 47, 48, 49, 50, 51 , 52, 53, 54, 55, 56, 57, 58, 59, 60, 61 , 62, 63, 64, 65, 66, 67, 68, 69, 70, 71 , 72, 73, 74, 75, 76, 77, 78, 79, 80, 81 , 82, 83, 84, 85, 86, 87, 88, 89, 90, 91 , 92, 93, 94, 95, 96, 97, 98, 99, or 100 nucleotides in length. In some embodiments, a first nucleic acid adapter (e.g., a nucleic acid adapter added upon tagmentation) and a second nucleic acid adapter (e.g., a nucleic acid adapter added during a splint-ligation reaction) have the same length or have different lengths.

[0202] When used in tagmentation, an adapter will include a TBS, e.g., the mosaic end sequence. In such embodiments, the TBS will be bound to an oligonucleotide to produce a double stranded region for loading into a synaptic complex. The oligonucleotide in the TBS may be modified at the 3’ end to prevent extension. A modification at the 3’ end to prevent extension may be an amino group modification, a phosphate group, a three-carbon chain (C3), e.g., 3-hydroxy propyl, a dideoxynucleotide (e.g., a dideoxycytosine (ddC), a dideoxyguanosine (ddG), a dideoxythymidine PATENT

[0203] Attorney Docket No.: 51178-019WO2

[0204] (ddT), a dideoxyadenosine (ddA), or a dideoxyuridine (ddU)), an inverted nucleotide, or a locked nucleic acid. In some embodiments, the first adapter will include a free hydroxyl group at the 5’ end by lacking a phosphate group. It will also be understood that the first adapter sequence binds only at one end to a transposase, so that the transposition of the adapter results in tagmentation. In some embodiments, the first adapter consists of a TBS and a priming region. In some embodiments, the nucleic acid adapter includes a priming region (e.g., a nucleic acid sequence of a primer or the reverse complement thereto (e.g., a domain for primer binding)). In some embodiments, the nucleic acid adapter does not include a priming region and is suitable for amplification-free reactions (e.g., PCR-free). In some embodiments, a nucleic acid adapter includes a barcode, an affinity tag, and / or a reporter moiety. In some embodiments, a nucleic acid adapter does not include a barcode, an affinity tag, and / or a reporter moiety. In some embodiments, an adapter includes a sequence that permits attachment of a nucleic acid to a surface (e.g., a bead, a resin, a surface of a reaction vessel (e.g., a microplate), or a flow cell). Exemplary nucleotide sequences that permit attachment of nucleic acids to a surface include P5 and P7 sequences. In some embodiments, it can be advantageous for each nucleic acid adapter to include the nucleic acid sequence of at least one universal primer or the reverse complement thereto (e.g., the domain for universal primer binding). Universal primers can have various applications, such as use in amplifying, sequencing, and / or identifying one or more nucleic acid samples

[0205] In some embodiments, the sense strand of the second adapter includes a phosphate group at the 5’ end. In some embodiments, the sense strand and / or the antisense of the second adapter include a modification at the 3’ end to prevent extension by a polymerase selected from an amino group modification, a phosphate group, a three-carbon chain (C3), e.g., 3-hydroxy propyl, a dideoxynucleotide (e.g., a dideoxycytosine (ddC), a dideoxyguanosine (ddG), a dideoxythymidine (ddT), a dideoxyadenosine (ddA), or a dideoxyuridine (ddU)), an inverted nucleotide, and a locked nucleic acid. In some embodiments, the modification at the 3’ end is an amino group modification. In some embodiments, the antisense strand includes a free 5’-hydroxyl group and is not phosphorylated.

[0206] In some embodiments, a nucleic acid adapter includes 5-methylated cytosines. In some embodiments, the splint oligonucleotide or nucleic acid adapter therein includes unmethylated cytosines.

[0207] In preferred embodiments, the nucleic acid sequence of the first and second nucleic acid adapters are different from each other and are not reverse complements. In some embodiments, the nucleic acid sequence that includes a priming region (e.g., a nucleic acid sequence of a primer or the reverse complement thereto (e.g., a domain for primer binding) is different in the first nucleic acid adapter and the second nucleic acid adapter such that two different primers can be used to amplify the engineered nucleic acid sample (e.g., the nucleic acid fragments appended with two nucleic acid adapters). In some embodiments, a primer having the following nucleic acid sequences selected from group consisting of

[0208] 5’-AATGATACGGCGACCACCGAGATCTACACGTAACACAGATCGTCGGCAGCGTCAGATGTG-3’ (SEQ ID NO: 9); PATENT

[0209] Attorney Docket No.: 51178-019WO2

[0210] 5’-AATGATACGGCGACCACCGAGATCTACACTAAGTTGTGGTCGTCGGCAGCGTCAGATGTG-3’ (SEQ ID NO: 10);

[0211] 5’-AATGATACGGCGACCACCGAGATCTACACCACCTACCTCTCGTCGGCAGCGTCAGATGTG-3’ (SEQ ID NO: 11);

[0212] 5’-AATGATACGGCGACCACCGAGATCTACACATATCAGTCCTCGTCGGCAGCGTCAGATGTG-3’ (SEQ ID NO: 12);

[0213] 5’-CAAGCAGAAGACGGCATACGAGATTGGACTTGACGTCTCGTGGGCTCGGAGATGTG-3’ (SEQ ID NO: 13);

[0214] 5’-CAAGCAGAAGACGGCATACGAGATGAGTTAGTTGGTCTCGTGGGCTCGGAGATGTG-3’ (SEQ ID NO: 14);

[0215] 5’-CAAGCAGAAGACGGCATACGAGATGTCAGGTTATGTCTCGTGGGCTCGGAGATGTG-3’ (SEQ ID NO: 15); and

[0216] 5’-CAAGCAGAAGACGGCATACGAGATGAAGTACCTGGTCTCGTGGGCTCGGAGATGTG-3’ (SEQ ID NO: 16) can bind to one of the nucleic acid adapters.

[0217] In some embodiments, a first nucleic acid adapter has the nucleic acid sequence 5’-TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG-3’ (SEQ ID NO: 6) and contains a 5’ amino modifier C12 modification. In some embodiments, a second nucleic acid adapter has the nucleic acid sequence 5’-CTGTCTCTTATACACATCTCCGAGCCCACGAGAC-3’ (SEQ ID NO: 7) and contains a 5’ phosphorylation and a 3’ amino modifier.

[0218] III. Kits

[0219] The invention provides kits that include one or more compositions (e.g., synaptic complexes, blocking molecules, and / or nucleic acid adapters) to perform any of the methods described herein. The kits may include one or more additional reagents that are useful, for example, for carrying out the methods of the invention. The kit may include one or more containers (e.g., one or more reaction vessels) for holding the components of the kit (e.g., tubes (e.g., microcentrifuge tubes), plates (e.g., microtiter plates or PCR plates), trays, packaging materials, and the like). In some embodiments, the kit includes a reaction vessel that is packaged with one or more reagent already added to the reaction vessel. For example, the kit may include a PCR plate that includes a volume (e.g., 1 pL, 2 pL, 3 pL, 4 pL, 5 pL, 6 pL, 7 pL, 8 pL, 9 pL, or 10 pL) of amplification primers (e.g., indexing primers) added to one or more wells of the PCR plate. The kit may also include instructions (e.g., printed instructions for using the kit).

[0220] A kit may include any of the nucleic acids described herein. For example, a kit may include a first nucleic acid adapter, a plurality of blocking molecules that are single-stranded oligonucleotides, and / or a splint oligonucleotide that includes a second nucleic acid adapter (e.g., a single-stranded nucleic acid adapter that is added to the 3’ end of a nucleic acid fragment that is at least singlestranded at the 3’ end). A kit may include a splint oligonucleotide as single-stranded elements that hybridize in situ and can contact a single-stranded nucleic acid fragment. A kit may include any of the proteins described herein. For example, a kit may include a transposase, a ligase, a plurality of blocking molecules that are proteins, and / or a polymerase suitable for amplification. PATENT

[0221] Attorney Docket No.: 51178-019WO2

[0222] A kit may include one or more additional reagents. For example, the one or more additional reagents may include a cofactor, a buffered solution (e.g., a ligation buffer, a Tris buffer, or a phosphate buffer), a denaturing solution, a sodium bisulfite solution, nuclease free water, ethanol, and / or a reference nucleic acid. The cofactor may be a divalent metal cation (e.g., a magnesium cation or a calcium cation). Any of the kits may also include a reagent for nucleic acid sequencing, which may include, for example, oligonucleotide primer(s), a substrate, an enzyme (e.g., a DNA polymerase), a mixture of nucleotides, and / or a reference nucleic acid.

[0223] A kit may be packaged in one or more boxes or containers that includes a 96-well PCR plate that includes primers (e.g., index primers) at a volume between 1 and 10 pL per well (e.g., 1 pL, 2 pL, 3 pL, 4 pL, 5 pL, 6 pL, 7 pL, 8 pL, 9 pL, or 10 pL) and five 0.5 mL tubes (e.g., microcentrifuge tubes or screw-cap tubes) that contain between 20 pL and 300 pL (e.g., 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 125, 130, 135, 140, 145, 150, 155, 160, 165, 170, 175, 180, 185, 190, 195, 200, 205, 210, 215, 220, 225, 230, 235, 240, 245, 250, 255, 260, 265, 270, 275, 280, 285, 290, 295, or 300 pL) that contain a first nucleic acid adapter, a ligation enhancer, a reaction buffer (e.g., a 5X reaction buffer), a splint oligonucleotide that includes a second nucleic acid adapter therein, or a ligase in each tube. The kit may further include a 2 mL tube (e.g., a microcentrifuge tube or a screw-cap tube) that includes between 500 pL and 1.5 mL (e.g., 500 pL, 525 pL, 550 pL, 575 pL, 600 pL, 625 pL, 650 pL, 675 pL, 700 pL, 725 pL, 750 pL, 775 pL, 800 pL, 825 pL, 850 pL, 875 pL, 900 pL, 925 pL, 950 pL, 975 pL, 1 mL, 1.1 mL, 1.15 mL, 1.2 mL, 1.25 mL, 1.3 mL, 1.35 mL, 1.4 mL, 1.45 mL, or 1.5 mL) of a PCR reaction buffer (e.g., a 2X Amplification ReadyMix) and a bottle (e.g., a 3 mL to 20 mL bottle) that contains a solution of magnetic beads (e.g., MAGwise™, seqWell Inc. magnetic beads) at a volume between 1 mL and 10 mL (e.g., 1 , 2, 3, 4, 5, 6, 7, 8, 9, or 10 mL). One or more of the kit components may be lyophilized.

[0224] An exemplary kit and components thereof are described in Table 1 below.

[0225] Table 1 : TnX DNA Prep Kit Components (24 Reactions)

[0226]

[0227] PATENT

[0228] Attorney Docket No.: 51178-019WO2

[0229]

[0230] EXAMPLES

[0231] The invention is described by the following non-limiting examples.

[0232] Example 1 : Tagmentation and splinted ligation with single-stranded nucleic acid binding blocking molecules for DNA library preparation

[0233] This Example describes preparing an engineered nucleic acid sample that includes DNA fragments with appended nucleic acid adapters added by performing tagmentation and splinted ligation reactions on a DNA sample in the presence of various blocking molecules to inhibit or reduce secondary structure formation of the single-stranded DNA fragments and a negative control sample that did not include blocking molecules. An overview of the steps of preparing an engineered nucleic acid sample is outlined in FIGS. 1A-1D. The blocking molecules utilized included a population of extreme thermostable single-stranded binding proteins (ET SSB) and a population of 3’-ddC-blocked random nonamers (e.g., 9 base pair oligonucleotides having different sequences).

[0234] In a microcentrifuge tube, a tagmentation reaction with a human DNA sample and synaptic complexes was prepared with the reagents shown in Table 2 for the blocked samples and the negative control. Each synaptic complex included two TnX transposases and a partially doublestranded, 5’-end phosphorylated DNA adapter containing the Illumina Read 1 sequencing primer and having the nucleic acid sequence 5’-tcgtcggcagcgtcAGATGTGTATAAGAGACAG-3’ (SEQ ID NO: 6), in which the capitalized nucleotides in the sequence denotes the double-stranded portion of the adapter that includes the mosaic end that binds the TnX transposases, and lowercase nucleotides in the sequence denotes the single-stranded portion of the adapter. The reaction was incubated at 55°C for 5 minutes to fragment the human DNA and add the nucleic acid adapter to the 5’ end of each generated double-stranded fragment. PATENT

[0235] Attorney Docket No.: 51178-019WO2

[0236] Table 2: Tagmentation Reaction

[0237]

[0238] The sense strand containing the single-stranded portion of the DNA adapter was transposed to the 5’ end of each strand of the generated double-stranded DNA fragments. The reverse complement of the double-stranded portion of the adapter was adjacent to the 3’ end of each doublestranded DNA fragment, in which a gap was present between the 3’ end of the adapter and the 3’ end of the fragment.

[0239] Following this reaction, the synaptic complexes were heat inactivated, and the doublestranded DNA fragments were denatured into single-stranded fragments by incubating the microcentrifuge tube at a temperature of 95 °C for 5 minutes. Upon denaturing into single-stranded DNA fragments, the reverse complement of the double-stranded portion of the DNA adapter released from the sense strand of the adapter. Additionally, multiple ET SSB molecules or 3’-ddC nonamers bound to each fragment to reduce the formation of DNA secondary structure and maintain a linear conformation. After heat inactivation, the reaction was incubated at 4 °C for 2 minutes.

[0240] Next, ligation reactions were performed to add a second DNA adapter to the 3’ end of each single-stranded DNA fragment in the blocked samples and the negative control. The second DNA adapter was provided in a plurality of splint oligonucleotides, each of which was a partially doublestranded DNA molecule including a single-stranded portion of 8 nucleotides having a degenerate sequence that could hybridize to the 3’ ends of the single-stranded DNA fragments that were generated during the DNA denaturation step. The sense strand of the splint oligonucleotide contained the second DNA adapter and had the nucleic acid sequence PATENT

[0241] Attorney Docket No.: 51178-019WO2

[0242] 5’-CTGTCTCTTATACACATCTCCGAGCCCACGAGAC-3’ (SEQ ID NO: 7), while the antisense strand of the splint oligonucleotide had the nucleic acid sequence

[0243] 5’-GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAGNNNNNNNN-3’ (SEQ ID NO: 8), in which N denotes any nucleotide selected from A, T, G, and C. The antisense strand of the splint oligonucleotide also had phosphorylated 5’ and 3’ ends to prevent extension by a polymerase.

[0244] The ligation reaction was performed in the same microcentrifuge tubes as the tagmentation reaction without any reaction clean up between the reactions. To each of the microcentrifuge tubes, 2 pL of 50 pM splint oligonucleotide and 26 pL of a 2X Blunt / TA Mastermix (New England BioLabs) was added. The microcentrifuge tube was then incubated at 25 °C for 60 minutes so the ligase could catalytically form a phosphodiester bond between the 3’ end of the single-stranded DNA fragments and the 5’ end of the second DNA adapter.

[0245] The DNA fragments in each sample were purified via incubation with about a one volume equivalent of magnetic beads (MAGwise™ Paramagnetic Beads, seqWell Inc.) for about 5 minutes and subsequent elution in 22.5 pL 10 mM Tris-HCI, pH 8.0 buffer.

[0246] Following purification of DNA fragments, amplification was performed via PCR with 100 ng DNA fragments, 10 pM UDI primers, 25 pL Kapa Hi Fi 2X Readymix reaction buffer, and less than 1 pL Kapa H i Fi DNA polymerase in each sample. The PCR primers included were as follows:

[0247] 5’-AATGATACGGCGACCACCGAGATCTACACGTAACACAGATCGTCGGCAGCGTCAGATGTG-3’ (SEQ ID NO: 9);

[0248] 5’-AATGATACGGCGACCACCGAGATCTACACTAAGTTGTGGTCGTCGGCAGCGTCAGATGTG-3’ (SEQ ID NO: 10);

[0249] 5’-AATGATACGGCGACCACCGAGATCTACACCACCTACCTCTCGTCGGCAGCGTCAGATGTG-3’ (SEQ ID NO: 11);

[0250] 5’-AATGATACGGCGACCACCGAGATCTACACATATCAGTCCTCGTCGGCAGCGTCAGATGTG-3’ (SEQ ID NO: 12);

[0251] 5’-CAAGCAGAAGACGGCATACGAGATTGGACTTGACGTCTCGTGGGCTCGGAGATGTG-3’ (SEQ ID NO: 13);

[0252] 5’-CAAGCAGAAGACGGCATACGAGATGAGTTAGTTGGTCTCGTGGGCTCGGAGATGTG-3’ (SEQ ID NO: 14);

[0253] 5’-CAAGCAGAAGACGGCATACGAGATGTCAGGTTATGTCTCGTGGGCTCGGAGATGTG-3’ (SEQ ID NO: 15); and

[0254] 5’-CAAGCAGAAGACGGCATACGAGATGAAGTACCTGGTCTCGTGGGCTCGGAGATGTG-3’ (SEQ ID NO: 16).

[0255] The PCR was performed in a thermocycler under the conditions outlined in Table 3 below. The fragment amplicons were then purified via magnetic beads as described above and eluted in 20 pL 10 mM Tris-HCI, pH 8.0 buffer. PATENT

[0256] Attorney Docket No.: 51178-019WO2

[0257] Table 3: PCR Conditions

[0258]

[0259] After amplification, the fragments were sequenced via next-generation sequencing in which the DNA quality of the samples that included blocking molecules and the negative control sample were compared. Libraries were captured using the Twist Exome Enrichment Panel and sequenced on the NEXTSEQ™ 2000. Data from each captured library was subsampled to 115 million reads and run through the Picard Hybrid Selection Analysis pipeline (Broad Institute).

[0260] When comparing sequence coverage between blocked and unblocked DNA samples, nearly twice as many GC-binds had less than 50% of mean coverage in unblocked samples as compared to samples that included blocking molecules. Table 4 below shows the counts of GC-bins above 1.5X or less than 0.5X normalized means between blocked and unblocked DNA samples.

[0261] Table 4: GC-Bin Counts

[0262] > <

[0263]

[0264] When comparing GO bias, samples that included blocking molecules in the reaction showed reduced GO bias and improved sample quality as compared to negative control samples that did not contain blocking molecules (FIG.2A-2C; Table 4). Additionally, library yields from samples containing blocking proteins and blocking oligonucleotides were compared in FIG. 3.

[0265] Example 2: Tagmentation with or without ligation for DNA library preparation

[0266] This Example describes preparing an engineered nucleic acid sample using bead-linked tagmentation or tagmentation and ligation in the presence of blocking molecules.

[0267] Human genomic DNA (100 ng) from cell line NA12878 (Coriell) was used to generate libraries with bead-linked synaptic complexes (Illumina DNA Prep S) (i.e., tagmentation without blocking molecules) or with combined tagmentation and ligation in the presence of blocking oligonucleotides (n=4 per library preparation method). The blocking oligonucleotides were 3’-ddC-blocked random nonamers (e.g., 9 base pair oligonucleotides having different sequences). PATENT

[0268] Attorney Docket No.: 51178-019WO2

[0269] Libraries were captured using the Twist Exome Enrichment Panel and sequenced on the NEXTSEQ™ 2000. Data from each captured library was subsampled to 115M reads and run through the Picard Hybrid Selection Analysis pipeline (Broad Institute). The library coverage of each method is shown below in Table 5. Comparable mean target coverage was observed for both library preparation methods; however, the tagmentation and ligation approach showed improved target coverage >1X and >10X.

[0270] Table 5: Exome Sequencing Metrics

[0271] > >

[0272]

[0273] Other Embodiments

[0274] All publications, patents, and patent applications mentioned in this specification are incorporated herein by reference to the same extent as if each independent publication or patent application was specifically and individually indicated to be incorporated by reference.

[0275] While the invention has been described in connection with specific embodiments thereof, it will be understood that it is capable of further modifications and this application is intended to cover any variations, uses, or adaptations following, in general, the principles and including such departures from the invention that come within known or customary practice within the art to which the invention pertains and may be applied to the essential features hereinbefore set forth, and follows in the scope of the claims.

[0276] Other embodiments are within the claims.

Claims

1. PATENT2.Attorney Docket No.: 51178-019WO23.CLAIMS1. A method of preparing an engineered nucleic acid sample comprising:5.(a) exposing a sample double stranded nucleic acid to a transposome comprising a transposase and a first nucleic acid adapter under conditions and for a time sufficient for the transposome to carry out a tagmentation reaction, wherein the tagmentation reaction produces a nucleic acid fragment comprising the first nucleic acid adapter transposed to a 5’ end of a fragment of the sample double-stranded nucleic acid, and wherein the nucleic acid fragment is double-stranded in part;6.(b) denaturing the nucleic acid fragment to generate a single-stranded nucleic acid fragment, and exposing the single-stranded nucleic acid fragment to a plurality of blocking molecules, wherein the blocking molecules form a complex with the single-stranded nucleic acid fragment;7.(c) exposing the complex comprising the single-stranded nucleic acid fragment and the blocking molecule to a plurality of splint oligonucleotides comprising a second nucleic acid adapter, wherein the splint oligonucleotides comprise a degenerate nucleic acid sequence at the 3’ end, wherein one of the plurality of splint oligonucleotides hybridizes to the 3’ end of the single-stranded nucleic acid fragment via the degenerate nucleic acid sequence; and8.(d) ligating the second nucleic acid adapter to the 3’ end of the single-stranded nucleic acid fragment using a ligase,9.thereby producing an engineered nucleic acid sample.

2. The method of claim 1 , wherein, in step (b), the nucleic acid fragment is denatured via heat or chemical denaturation.

3. The method of claim 1 or 2, wherein the degenerate nucleic acid sequence is between 5 to 15 nucleotides in length.

4. The method of any one of claims 1-3, wherein the plurality of blocking molecules comprises a protein.

5. The method of claim 4, wherein the protein is a single-stranded nucleic acid binding protein.

6. The method of any one of claims 1-3, wherein the plurality of blocking molecules comprises singlestranded oligonucleotides having different sequences.

7. The method of claim 6, wherein the single-stranded oligonucleotides comprise a 3’ end modification to prevent extension by a polymerase.PATENT16.Attorney Docket No.: 51178-019WO28. The method of claim 7, wherein the 3’ end modification is selected from the group consisting of an amino group modification, a phosphate group, a three-carbon chain (C3), a dideoxynucleotide, an inverted nucleotide, and a locked nucleic acid.

9. The method of any one of claims 6-8, wherein the single-stranded oligonucleotides are between 5 and 15 nucleotides in length.

10. The method of claim 9, wherein the single-stranded oligonucleotides are between 5 and 10 nucleotides in length.

11. The method of any one of claims 1 -10, wherein the sample double-stranded nucleic acid is DNA or a DNA / RNA hybrid.

12. The method of claim 11 , wherein the DNA is genomic DNA, mitochondrial DNA, cDNA, or plasmid DNA.

13. The method of any one of claims 1-12, wherein the transposase is selected from the group consisting of a Tn5 transposase, a hyperactive Tn5 transposase, and a TnX transposase.

14. The method of any one of claims 1-13, wherein the first nucleic acid adapter and the second nucleic acid adapter each comprise a priming region.

15. The method of claim 14, wherein the priming region of the first nucleic acid adapter has a different nucleotide sequence and is not the reverse complement of the priming region of the second nucleic acid adapter.

16. The method of any one of claims 1-15, further comprising the step of amplifying the engineered nucleic acid sample.

17. The method of claim 16, wherein the amplification is performed via polymerase chain reaction (PCR), multiple annealing and looping-based amplification cycles (MALBAC), or multiple displacement amplification (MDA).

18. A kit comprising:28.(i) a transposome comprising a transposase and a first nucleic acid adapter;29.(ii) a plurality of blocking molecules; and30.(Hi) a second nucleic acid adapter comprising a splint oligonucleotide comprising a degenerate nucleic acid sequence at the 3’ end.

19. The kit of claim 18, further comprising a ligase.PATENT32.Attorney Docket No.: 51178-019WO220. The kit of claim 18 or 19, further comprising a suitable buffer.

21. The kit of any one of claims 18-20, wherein the degenerate nucleic acid sequence is between 5 to 15 nucleotides in length.

22. The kit of any one of claims 18-20, wherein the plurality of blocking molecules comprises a protein.

23. The kit of claim 22, wherein the protein is a single-stranded nucleic acid binding protein.

24. The kit of any one of claims 18-23, wherein the plurality of blocking molecule comprises singlestranded oligonucleotides having different sequences.

25. The kit of claim 24, wherein the single-stranded oligonucleotides comprise a 3’ end modification to prevent extension by a polymerase.

26. The kit of claim 25, wherein the 3’ end modification is selected from the group consisting of an amino group modification, a phosphate group, a C3, a dideoxynucleotide, an inverted nucleotide, and a locked nucleic acid.

27. The kit of any one of claims 24-26, wherein the single-stranded oligonucleotides are between 5 and 15 nucleotides in length.

28. The kit of claim 27, wherein the single-stranded oligonucleotides are between 5 and 10 nucleotides in length.

29. The kit of any one of claims 18-28, wherein the transposase is selected from the group consisting of a Tn5 transposase, a hyperactive Tn5 transposase, and a TnX transposase.

30. The kit of any one of claims 18-29, wherein the first nucleic acid adapter and the second nucleic acid adapter each comprise a priming region.

31. The kit of claim 30, wherein the priming region of the first nucleic acid adapter has a different nucleic acid sequence and is not the reverse complement of the priming region of the second nucleic acid adapter.