Methods and compositions for nucleic acid sequencing
The integration of barcode and capture sequences with transpososomes in nucleic acid molecules using active transposases addresses the inefficiencies in current sequencing methods, enabling precise and high-throughput nucleic acid sequencing.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- PLACE GENOMICS CORP
- Filing Date
- 2025-10-14
- Publication Date
- 2026-04-23
AI Technical Summary
Current nucleic acid sequencing methods lack efficient and precise methods for integrating barcode sequences and capture sequences with transpososomes to facilitate high-throughput sequencing and accurate identification of nucleic acid sequences.
A composition comprising nucleic acid molecules with barcode and capture sequences, coupled with transpososomes containing active transposases and polynucleotides with transposon end recognition sequences, allows for the integration of barcode sequences into target nucleic acids, enabling precise sequencing and identification.
Enables high-throughput nucleic acid sequencing by generating barcoded nucleic acid molecules with integrated barcode and target sequences, allowing for accurate sequencing and identification of nucleic acid sequences.
Smart Images

Figure US2025050961_23042026_PF_FP_ABST
Abstract
Description
Atty Dkt No.: 68244-703601METHODS AND COMPOSITIONS FOR NUCLEIC ACID SEQUENCINGCROSS REFERENCE
[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 707,681, filed October 15, 2024, which is incorporated by reference herein in its entirety for all purposes.SEQUENCE LISTING
[0002] The instant application contains a Sequence Listing which has been submitted electronically in XML format and is hereby incorporated by reference in its entirety. Said XML copy, created on October 9, 2025, is named 68244-703_601_SL.xml and is 6,292 bytes in size.BACKGROUND
[0003] Nucleic acids may be sequenced for various purposes. Nucleic acid sequencing can have genomics applications and provide insight into genetic markers that underpin health and disease. Barcode sequences can provide identifying information for nucleic acid processing and analysis.SUMMARY
[0004] In one aspect, disclosed herein is a composition, comprising: (i) a nucleic acid molecule, comprising a plurality of barcode sequences and a plurality of capture sequences; and (ii) a plurality of transpososomes coupled to the nucleic acid molecule, wherein each transpososome comprises an active transposase and a plurality of polynucleotides, wherein the plurality of polynucleotides comprises: (A) a pair of transposon end recognition sequences that are bound to the active transposase and (B) a polynucleotide comprising a binding sequence that is hybridized to a capture sequence of the plurality of capture sequences of the nucleic acid molecule.
[0005] In some embodiments, a transposon end recognition sequence of the pair of transposon end recognition sequences comprises a mosaic end (ME) sequence or a reverse complement thereof. In some embodiments, a transposon end recognition sequence of the pair of transposon end recognition sequences comprises an inside end (IE) sequence, an outside end (OE) sequence, or a reverse complement thereof. In some embodiments, a polynucleotide of the plurality of polynucleotides further comprises a unique molecular index. In some embodiments, a polynucleotide of the plurality of polynucleotides further comprises a primer binding sequence or reverse complement thereof. In some embodiments, a polynucleotide of the plurality of polynucleotides further comprises an adapter sequence. In some embodiments, the binding sequence of the polynucleotide is on a 3’ end of the polynucleotide. In some embodiments, aAtty Dkt No.: 68244-703601 polynucleotide of the plurality of polynucleotides comprises a transposon end recognition sequence of the pair of transposon end recognition sequences and the binding sequence. In some embodiments, a polynucleotide of the plurality of polynucleotides comprises a transposon end recognition sequence of the pair of transposon end recognition sequences, the binding sequence, and a unique molecular index. In some embodiments, the active transposase comprises a protein dimer.
[0006] In some embodiments, the nucleic acid molecule comprises at least 10 barcode sequences. In some embodiments, the nucleic acid molecule comprises at least 100 barcode sequences. In some embodiments, the plurality of barcode sequences comprises a same barcode sequence. In some embodiments, the plurality of barcode sequences comprises different barcode sequences. In some embodiments, the plurality of capture sequences comprises at least 10 capture sequences. In some embodiments, the plurality of capture sequences comprises at least 100 capture sequences. In some embodiments, the nucleic acid molecule is at least 500 nucleotides in length. In some embodiments, a capture sequence of the plurality of capture sequences is on a 3’ side of a barcode sequence of the plurality of barcode sequences of the nucleic acid molecule. In some embodiments, a capture sequence of the plurality of capture sequences is adjacent to a barcode sequence of the plurality of barcode sequences of the nucleic acid molecule. In some embodiments, a barcode sequence of the plurality of barcode sequences is located between two capture sequences of the plurality of capture sequences of the nucleic acid molecule. In some embodiments, each transpososome of the plurality of transpososomes comprises two polynucleotides, each comprising a binding sequence that is hybridized to a separate capture sequence of the plurality of capture sequences of the nucleic acid molecule. In some embodiments, each separate capture sequence that is hybridized to a polynucleotide of the two polynucleotides is on a 3’ side of a barcode sequence of the plurality of barcode sequences. In some embodiments, each separate capture sequence that is hybridized to a polynucleotide of the two polynucleotides is adjacent to a barcode sequence of the plurality of barcode sequences.
[0007] In some embodiments, the plurality of polynucleotides comprises four nucleic acid strands comprising: a first nucleic acid strand comprising a first transposon end recognition sequence; a second nucleic acid strand comprising a reverse complement of the first transposon end recognition sequence and a first binding sequence that is hybridized to a first capture sequence of the plurality of capture sequences of the nucleic acid molecule; a third nucleic acid strand comprising a second transposon end recognition sequence; and a fourth nucleic acid strand comprising a reverse complement of the second transposon end recognition sequence, and a second binding sequence that is hybridized to a second capture sequence of the plurality of capture sequences of the nucleic acid molecule. In some embodiments, the first nucleic acid strand or theAtty Dkt No.: 68244-703601 third nucleic acid strand further comprises an adapter sequence. In some embodiments, the second nucleic acid strand or the fourth nucleic acid strand further comprises a unique molecular index. In some embodiments, the second nucleic acid strand and the fourth nucleic acid strand each comprises a unique molecular index.
[0008] In some embodiments, the plurality of polynucleotides comprises three nucleic acid strands comprising: a first nucleic acid strand comprising (A) a first transposon end recognition sequence, (B) a reverse complement of a second transposon end recognition sequence; a second nucleic acid strand comprising a reverse complement of the first transposon end recognition sequence, wherein a portion of the second nucleic acid strand is hybridized to a first segment of the first nucleic acid strand; and a third nucleic acid strand comprising the second transposon end recognition sequence, wherein a portion of the third nucleic acid strand is hybridized to a second segment of the first nucleic acid strand. In some embodiments, the second nucleic acid strand or the third nucleic acid strand further comprises a binding sequence that is hybridized to a capture sequence of the plurality of capture sequences of the nucleic acid molecule. In some embodiments, the first nucleic acid strand further comprises a unique molecular index located between the first transposon recognition sequence and the second transposon recognition sequence. In some embodiments, the first nucleic acid strand further comprises two discrete unique molecular indexes located between the first transposon recognition sequence and the reverse complement of the second transposon recognition sequence.
[0009] In some embodiments, the plurality of polynucleotides comprises four nucleic acid strands comprising: a first nucleic acid strand comprising a first transposon end recognition sequence, wherein the first nucleic acid strand comprises a first region and a second region; a second nucleic acid strand comprising second transposon end recognition sequence, wherein the fourth nucleic acid strand comprises a third region and a fourth region; wherein the first region of the first nucleic acid strand is hybridized to the third region of the second nucleic acid strand; a third nucleic acid strand comprising a reverse complement of the first transposon end recognition sequence, wherein a portion of the third nucleic acid strand is hybridized to the second region of the first nucleic acid strand; and a fourth nucleic acid strand comprising a reverse complement of the second transposon end recognition sequence, wherein a portion of the fourth nucleic acid strand is hybridized to the fourth region of the second nucleic acid strand. In some embodiments, the third nucleic acid strand or the fourth nucleic acid strand comprises a binding sequence that is hybridized to a capture sequence of the plurality of capture sequences of the nucleic acid molecule. In some embodiments, the first nucleic acid strand or the second nucleic acid strand further comprises a unique molecular index.
[0010] In some embodiments, the nucleic acid molecule is single-stranded.Atty Dkt No.: 68244-703601
[0011] In some embodiments, barcode sequences of the plurality of barcode sequences are separated by a non-canonical nucleotide. In some embodiments, the non-canonical nucleotide is non-amplifiable. In some embodiments, barcode sequences of the plurality of barcode sequences are separated by a recognition sequence bound to a DNA-binding protein. In some embodiments, the DNA-binding protein comprises a CRISPR-Cas domain, and wherein the recognition sequence is hybridized to a guide RNA complexed with the CRISPR-Cas domain. In some embodiments, the CRISPR-Cas domain comprises an inactive nuclease domain. In some embodiments, the CRISPR-Cas domain comprises a Cas9 domain.
[0012] In another aspect, disclosed herein is a method, comprising: (a) providing: (i) a first nucleic acid, comprising a plurality of barcode sequences and a plurality of capture sequences; (ii) a plurality of transpososomes coupled to the first nucleic acid via the plurality of capture sequences, wherein each transpososome comprises an active transposase and a plurality of polynucleotides, wherein the plurality of polynucleotides comprises a pair of transposon end recognition sequences that are bound to the active transposase; and (iii) a second nucleic acid; and (b) generating a barcoded nucleic acid molecule using the first nucleic acid, the second nucleic acid, the active transposase, and a polynucleotide of the plurality of polynucleotides; wherein the barcoded nucleic acid molecule comprises (I) a barcode sequence of the plurality of barcode sequences of the first nucleic acid or a reverse complement thereof, and (II) a nucleic acid sequence of the second nucleic acid or a reverse complement thereof.
[0013] In some embodiments, (b) comprises (i) performing a nucleic acid extension reaction on the polynucleotide or a derivative thereof using a segment of the first nucleic acid as a template, and (ii) performing a transposase-mediated insertion reaction using the active transposase to integrate the polynucleotide or a derivative thereof into the second nucleic acid. In some embodiments, (b)(i) and (b)(ii) occur simultaneously. In some embodiments, (b)(i) and (b)(ii) occur sequentially. In some embodiments, (b)(i) occurs prior to (b)(ii). In some embodiments, (b)(i) occurs after (b)(ii). In some embodiments, in (b), the active transposase integrates a barcoded derivative of the polynucleotide into the second nucleic acid in a transposase-mediated insertion reaction, wherein the barcoded derivative comprises the barcode sequence or a reverse complement thereof.
[0014] In some embodiments, the plurality of polynucleotides comprises a polynucleotide comprising a binding sequence that is hybridized to a capture sequence of the plurality of capture sequences. In some embodiments, the binding sequence of the polynucleotide is on a 3’ end of the polynucleotide, and (b) comprises performing a nucleic acid extension reaction using the polynucleotide of the plurality of polynucleotides as a primer and a segment of the first nucleic acid as a template to generate the barcoded derivative of the polynucleotide. In some embodiments,Atty Dkt No.: 68244-703601 the active transposase integrates the barcoded derivative of the polynucleotide into the second nucleic acid into a random location in the second nucleic acid. In some embodiments, (b) further comprises covalently linking an end of the polynucleotide or a barcoded derivative of the polynucleotide and the nucleic acid sequence of the second nucleic acid or a reverse complement thereof in a ligation reaction or a nucleic acid gap fill-in reaction. In some embodiments, (b) further comprises incorporating a dideoxynucleotide (ddNTP) at a 3’ end of an extension product of polynucleotide or a derivative thereof. In some embodiments, the barcoded nucleic acid molecule comprises a length of from 100 to 600 bp.
[0015] In some embodiments, the method further comprises (c) sequencing the barcoded nucleic acid molecule or a derivative thereof. In some embodiments, the barcode sequence identifies the second nucleic acid, and wherein the method further comprises (d) identifying the barcode sequence or the reverse complement thereof in the barcoded nucleic acid molecule or the derivative thereof, thereby identifying its associated nucleic acid sequence as being derived from the second nucleic acid.
[0016] In some embodiments, (b) further comprises generating a plurality of barcoded nucleic acid molecules using the first nucleic acid, the second nucleic acid, and the plurality of transpososomes; wherein the plurality of barcoded nucleic acid molecules each comprises (I) a barcode sequence of the plurality of barcode sequences of the first nucleic acid or a reverse complement thereof, and (II) a nucleic acid sequence of the second nucleic acid or a reverse complement thereof. In some embodiments, the plurality of barcoded nucleic acid molecules comprises different nucleic acid sequences of the second nucleic acid. In some embodiments, the method further comprises (c) sequencing the plurality of barcoded nucleic acid molecules. In some embodiments, (d) further comprises linking nucleic acid sequences of the plurality of barcoded nucleic acid molecules in an assembled sequence. In some embodiments, the linking in (d) is based at least in part on associated barcode sequences of the plurality of barcoded nucleic acid molecules. In some embodiments, in (a), the plurality of barcode sequences comprises identical barcode sequences.
[0017] In some embodiments, in (a), the plurality of polynucleotides of a given transpososome of the plurality of transpososomes comprises a unique molecular index; wherein, in (b), the plurality of barcoded nucleic acid molecules comprises (I) a first barcoded nucleic acid molecule comprising the unique molecular index and a first sequence of the second nucleic acid and (II) a second barcoded nucleic acid molecule comprising the unique molecular index and a second sequence of the second nucleic acid; and wherein (d) comprises identifying the first sequence and the second sequence as being adjacent to each other on the second nucleic acid based on the unique molecular index.Atty Dkt No.: 68244-703601
[0018] In some embodiments, in (a), the plurality of polynucleotides of a given transpososome of the plurality of transpososomes comprises a first unique molecular index and a second unique molecular index; wherein, in (b), the plurality of barcoded nucleic acid molecules comprises (I) a first barcoded nucleic acid molecule comprising the first unique molecular index and a first sequence of the second nucleic acid and (II) a second barcoded nucleic acid molecule comprising the second unique molecular index and a second sequence of the second nucleic acid; and wherein (d) comprises identifying the first sequence and the second sequence as being adjacent to each other on the second nucleic acid based on the first unique molecular index and the second unique molecular index. In some embodiments, the first unique molecular index and the second unique molecular index are on a same polynucleotide of the plurality of polynucleotides.
[0019] In some embodiments, the method further comprises, prior to (a), generating a polynucleotide of the plurality of polynucleotides. In some embodiments, the method further comprises, prior to (a), generating the first nucleic acid. In some embodiments, the generating the first nucleic acid comprises performing a nucleic acid amplification reaction on a template nucleic acid molecule comprising a barcode sequence of the plurality of barcode sequences and a capture sequence of the plurality of capture sequences. In some embodiments, the nucleic acid amplification reaction comprises rolling circle amplification. In some embodiments, the second nucleic acid is double-stranded. In some embodiments, the second nucleic acid comprises genomic DNA. In some embodiments, the second nucleic acid is in a cell.
[0020] In another aspect, disclosed herein is a method, comprising: (a) providing: (i) a nucleic acid molecule, comprising a plurality of barcode sequences and a plurality of capture sequences; (ii) a plurality of transpososomes coupled to the nucleic acid molecule, wherein each transpososome comprises an active transposase and a plurality of polynucleotides, wherein the plurality of polynucleotides comprises (A) a pair of transposon end recognition sequences that are bound to the active transposase, and (B) a polynucleotide comprising a binding sequence that is hybridized to a capture sequence of the plurality of capture sequences of the nucleic acid molecule; and (b) performing a nucleic acid extension reaction to extend a polynucleotide of a transpososome of the plurality of transpososomes that is hybridized to a capture sequence of the nucleic acid molecule to yield an extended polynucleotide comprising (I) the binding sequence and (II) a reverse complement of a barcode sequence of the plurality of barcode sequences.
[0021] In another aspect, disclosed herein is a method, comprising: (a) providing: (i) a nucleic acid molecule, comprising a plurality of barcode sequences and a plurality of capture sequences; (ii) a plurality of transpososomes coupled to the nucleic acid molecule, wherein each transpososome comprises an active transposase and a plurality of polynucleotides, wherein the plurality of polynucleotides comprises (A) a pair of transposon end recognition sequences that areAtty Dkt No.: 68244-703601 bound to the active transposase, and (B) two polynucleotides each comprising a binding sequence that is hybridized to a capture sequence of the plurality of capture sequences of the nucleic acid molecule; and (b) attaching the two polynucleotides together.
[0022] In another aspect, disclosed herein is a nucleic acid comprising: a first nucleic acid strand comprising (A) a first transposon end recognition sequence, (B) a reverse complement of a second transposon end recognition sequence, and (C) two discrete barcode sequences located between the first transposon end recognition sequence and the second transposon end recognition sequence; a second nucleic acid strand comprising a reverse complement of the first transposon end recognition sequence, wherein a portion of the second nucleic acid strand is hybridized to a first segment of the first nucleic acid strand; and a third nucleic acid strand comprising the second transposon end recognition sequence, wherein a portion of the third nucleic acid strand is hybridized to a second segment of the first nucleic acid strand; wherein the second nucleic acid strand and the third nucleic acid strand are separate nucleic acid strands.
[0023] In some embodiments, the two discrete barcode sequences are identical sequences. In some embodiments, the two discrete barcode sequences comprise different sequences. In some embodiments, the first nucleic acid strand further comprises a functional sequence located between the first transposon end recognition sequence and a discrete barcode sequence of the two discrete barcode sequences. In some embodiments, the functional sequence of the first nucleic acid strand distinguishes the first nucleic acid strand from a reverse complement of the first nucleic acid strand. In some embodiments, the functional sequence of the first nucleic acid strand is at most 3 nucleotides in length. In some embodiments, the first nucleic acid strand further comprises (i) a first functional sequence located between the first transposon end recognition sequence and a first discrete barcode sequence of the two discrete barcode sequences and (ii) a second functional sequence located between the second transposon end recognition sequence and a second discrete barcode of the two discrete barcode sequences. In some embodiments, the first nucleic acid strand further comprises a primer binding sequence or reverse complement thereof located between the two discrete barcode sequences. In some embodiments, the first nucleic acid strand further comprises two distinct primer binding sequences or reverse complements thereof located between the two discrete barcode sequences.
[0024] In some embodiments, the first nucleic acid strand further comprises a hairpin region located between the two discrete barcode sequences, wherein the hairpin region comprises a first portion of the first nucleic acid strand that is hybridized to a second portion of the first nucleic acid strand.
[0025] In some embodiments, the second nucleic acid strand further comprises a reverse complement of a discrete barcode sequence of the two discrete barcode sequences. In someAtty Dkt No.: 68244-703601 embodiments, the second nucleic acid strand does not comprise a reverse complement of a discrete barcode sequence of the two discrete barcode sequences. In some embodiments, the second nucleic acid strand and the third nucleic acid strand do not comprise a reverse complement of a discrete barcode sequence of the two discrete barcode sequences.
[0026] In some embodiments, the first nucleic acid strand comprises a non-hybridizing region that is not hybridized to the second nucleic acid strand and that is not hybridized to the third nucleic acid strand. In some embodiments, the non-hybridizing region of the first nucleic acid strand comprises a primer binding sequence or reverse complement thereof. In some embodiments, the non-hybridizing region of the first nucleic acid strand comprises two distinct primer binding sequences or reverse complements thereof. In some embodiments, the non-hybridizing region of the first nucleic acid strand comprises a discrete barcode sequence of the two discrete barcode sequences. In some embodiments, the non-hybridizing region of the first nucleic acid strand is located between the first segment and the second segment of the first nucleic acid strand. In some embodiments, the non-hybridizing region of the first nucleic acid strand is located between two hybridizing regions of the first nucleic acid strand, wherein the two hybridizing regions are each hybridized to the second nucleic acid strand.
[0027] In some embodiments, the first nucleic acid strand comprises two distinct nonhybridizing regions that are not hybridized to the second nucleic acid strand and that are not hybridized to the third nucleic acid strand. In some embodiments, the two distinct non-hybridizing regions of the first nucleic acid strand each comprises a primer binding sequence or reverse complement thereof. In some embodiments, the two distinct non-hybridizing regions of the first nucleic acid strand comprise a primer binding sequence or reverse complement thereof and a discrete barcode sequence of the two discrete barcode sequences. In some embodiments, a first non-hybridizing region of the two non-hybridizing regions of the first nucleic acid strand is located between the first segment and the second segment of the first nucleic acid strand; and a second non-hybridizing region of the two non-hybridizing regions of the first nucleic acid strand is located between two hybridizing regions of the first nucleic acid strand, wherein the two hybridizing regions are each hybridized to the second nucleic acid strand.
[0028] In some embodiments, the second nucleic acid strand comprises a non-complementary region that is not hybridized to the first nucleic acid strand. In some embodiments, the non- complementary region of the second nucleic acid strand comprises a primer binding sequence or reverse complement thereof. In some embodiments, the non-complementary region of the second nucleic acid strand comprises a binding sequence that is hybridized to a splint oligonucleotide. In some embodiments, the non-complementary region of the second nucleic acid strand comprises a binding sequence that is hybridized to a nucleic acid molecule comprising a third discrete barcodeAtty Dkt No.: 68244-703601 sequence. In some embodiments, the nucleic acid molecule comprises a plurality of barcode sequences. In some embodiments, the non-complementary region of the second nucleic acid strand is on a 3’ end of the second nucleic acid strand.
[0029] In some embodiments, the second nucleic acid strand comprises two distinct non- complementary regions that are not hybridized to the first nucleic acid strand. In some embodiments, the two distinct non-complementary regions comprise a primer binding sequence or reverse complement thereof. In some embodiments, a first non-complementary region of the two non-complementary regions of the second nucleic acid strand is located on a 3’ end of the second nucleic acid strand; and a second non-complementary region of the two non- complementary regions of the second nucleic acid strand is located in an internal segment of the second nucleic acid strand.
[0030] In some embodiments, the third nucleic acid strand further comprises a non- complementary region that is not hybridized to the first nucleic acid strand. In some embodiments, the non-complementary region of the third nucleic acid strand comprises a primer binding sequence or reverse complement thereof. In some embodiments, the non-complementary region of the third nucleic acid strand is on a 3’ end of the third nucleic acid strand.
[0031] In some embodiments, disclosed herein is a method, comprising (a) providing the nucleic acid, an active transposase, and a target nucleic acid molecule comprising a target sequence; and (b) generating a barcoded nucleic acid molecule using the nucleic acid, the active transposase, and the target nucleic acid molecule, wherein the barcoded nucleic acid molecule comprises (i) a discrete barcode sequence of the two discrete barcode sequences, and (ii) the target sequence. In some embodiments, the active transposase integrates the nucleic acid into the target nucleic acid molecule in a transposase-mediated insertion reaction. In some embodiments, the method further comprising (c) sequencing the barcoded nucleic acid molecule or a derivative thereof. In some embodiments, (b) further comprises generating an additional barcoded nucleic acid molecule using the nucleic acid, wherein the additional barcoded nucleic acid molecule comprises (I) a discrete barcode sequence of the two discrete barcode sequences, and (II) an additional target sequence of the target nucleic acid molecule.
[0032] In some embodiments, the method further comprises (c) sequencing the barcoded nucleic acid molecule or a derivative thereof and sequencing the additional barcoded nucleic acid molecule or a derivative thereof. In some embodiments, the method further comprises (d) linking a nucleic acid sequence of the barcoded nucleic acid molecule and a nucleic acid sequence of the additional barcoded nucleic acid molecule in an assembled sequence. In some embodiments, the discrete barcode sequence of the barcoded nucleic acid molecule and the discrete barcode sequence of the additional barcoded nucleic acid molecule identify the barcoded nucleic acid molecule andAtty Dkt No.: 68244-703601 the additional barcoded nucleic acid molecule as being generated by a common transpososome comprising the nucleic acid and the active transposase. In some embodiments, (d) comprises identifying the target sequence and the additional target sequence as being adjacent to each other on the target nucleic acid molecule based on the discrete barcode sequence of the barcoded nucleic acid molecule and the discrete barcode sequence of the additional barcoded nucleic acid molecule.
[0033] In some embodiments, the nucleic acid is integrated into a target nucleic acid molecule.
[0034] In another aspect, disclosed herein is a nucleic acid comprising: a first nucleic acid strand comprising a first transposon end recognition sequence, wherein the first nucleic acid strand comprises a first region and a second region; a second nucleic acid strand comprising second transposon end recognition sequence, wherein the fourth nucleic acid strand comprises a third region and a fourth region; wherein the first region of the first nucleic acid strand is hybridized to the third region of the second nucleic acid strand; a third nucleic acid strand comprising a reverse complement of the first transposon end recognition sequence, wherein a portion of the third nucleic acid strand is hybridized to the second region of the first nucleic acid strand; a fourth nucleic acid strand comprising a reverse complement of the second transposon end recognition sequence, wherein a portion of the fourth nucleic acid strand is hybridized to the fourth region of the second nucleic acid strand.
[0035] In some embodiments, the first nucleic acid strand or the second nucleic acid strand further comprises a barcode sequence. In some embodiments, the first nucleic acid strand further comprises a hairpin region.
[0036] In some embodiments, disclosed herein is a method, comprising: (a) providing the nucleic acid, an active transposase, and a target nucleic acid molecule comprising a target sequence; and (b) generating a barcoded nucleic acid molecule using the nucleic acid, the active transposase, and the target nucleic acid molecule, wherein the barcoded nucleic acid molecule comprises (i) a discrete barcode sequence of the two discrete barcode sequences, and (ii) the target sequence. In some embodiments, the active transposase integrates the nucleic acid into the target nucleic acid molecule in a transposase-mediated insertion reaction. In some embodiments, the method further comprises (c) sequencing the barcoded nucleic acid molecule or a derivative thereof. In some embodiments, (b) further comprises generating an additional barcoded nucleic acid molecule using the nucleic acid, wherein the additional barcoded nucleic acid molecule comprises (I) a discrete barcode sequence of the two discrete barcode sequences, and (II) an additional target sequence of the target nucleic acid molecule.
[0037] In some embodiments, the method further comprises (c) sequencing the barcoded nucleic acid molecule or a derivative thereof and sequencing the additional barcoded nucleic acid molecule or a derivative thereof. In some embodiments, the method further comprises (d) linkingAtty Dkt No.: 68244-703601 a nucleic acid sequence of the barcoded nucleic acid molecule and a nucleic acid sequence of the additional barcoded nucleic acid molecule in an assembled sequence.
[0038] In some embodiments, the discrete barcode sequence of the barcoded nucleic acid molecule and the discrete barcode sequence of the additional barcoded nucleic acid molecule identify the barcoded nucleic acid molecule and the additional barcoded nucleic acid molecule as being generated by a common transpososome comprising the nucleic acid and the active transposase. In some embodiments, (d) comprises identifying the target sequence and the additional target sequence as being adj acent to each other on the target nucleic acid molecule based on the discrete barcode sequence of the barcoded nucleic acid molecule and the discrete barcode sequence of the additional barcoded nucleic acid molecule.
[0039] In some embodiments, the nucleic acid is integrated into a target nucleic acid molecule.
[0040] In one aspect, the present disclosure provides a composition comprising: (i) a nucleic acid strand comprising two identical copies of a transposon end recognition sequence; (ii) a first oligonucleotide comprising a reverse complement of the transposon end recognition sequence; and (iii) a second oligonucleotide comprising a reverse complement of the transposon end recognition sequence. In some embodiments, the transposon end recognition sequence comprises any one of SEQ ID NOs: 1-6. In some embodiments, the nucleic acid strand comprises (I) a first segment comprising a first barcode sequence and a first copy of the two identical copies of the transposon end recognition sequence, and (II) a second segment comprising a second barcode sequence and a second copy of the two identical copies of the transposon end recognition sequence. In some embodiments, the first barcode sequence and the second barcode sequence are identical sequences. In some embodiments, the first barcode sequence and the second barcode sequence are different sequences. In some embodiments, the nucleic acid strand comprises, in order from 5’ to 3’ : the first barcode sequence, the first copy of the two identical copies of the transposon end recognition sequence, the second barcode sequence, and the second copy of the two identical copies of the transposon end recognition sequence.
[0041] In some embodiments, the nucleic acid strand further comprises a restriction enzyme recognition sequence. In some embodiments, the first oligonucleotide further comprises a reverse complement of the restriction enzyme recognition sequence. In some embodiments, the nucleic acid strand further comprises a restriction enzyme cut site located adjacent and 3’ to the first copy of the two identical copies of the transposon end recognition sequence. In some embodiments, the restriction enzyme recognition sequence is a PvuII restriction enzyme recognition sequence. In some embodiments, the nucleic acid strand further comprises an additional restriction enzyme recognition sequence, and wherein the second oligonucleotide further comprises a reverse complement of the additional restriction enzyme recognition sequence. In some embodiments, theAtty Dkt No.: 68244-703601 nucleic acid strand further comprises an additional restriction enzyme cut site located adjacent and 3’ to the second copy of the two identical copies of the transposon end recognition sequence. In some embodiments, the additional restriction enzyme recognition sequence is a PvuII restriction enzyme recognition sequence. In some embodiments, the nucleic acid strand does not comprise a restriction enzyme recognition sequence.
[0042] In some embodiments, the composition further comprises one or more linker polynucleotides coupling the first segment of the nucleic acid strand to the second segment of the nucleic acid strand, wherein the one or more linker polynucleotides is distinct from the nucleic acid strand. In some embodiments, the one or more linker polynucleotides comprises a unimolecular linker polynucleotide comprising (i) a first capture sequence hybridized to a first binding site in the first segment of the nucleic acid strand; and (ii) a second capture sequence hybridized to a second binding site in the second segment of the nucleic acid strand. In some embodiments, the composition comprises a plurality of linker polynucleotide molecules coupling the first segment of the nucleic acid strand to the second segment of the nucleic acid strand, wherein the plurality of linker polynucleotides comprises (i) a first linker polynucleotide molecule hybridized to a first binding site in the first segment of the nucleic acid strand; and (ii) a second linker polynucleotide molecule hybridized to a second binding site in the second segment of the nucleic acid strand. In some embodiments, the one or more linker polynucleotides comprises a linker polynucleotide that comprises the first oligonucleotide.
[0043] In some embodiments, the nucleic acid strand further comprises one or more primer binding sequences. In some embodiments, the nucleic acid strand further comprises primer binding sequences at its 5’ end and its 3’ end. In some embodiments, the nucleic acid strand is coupled to a biotin. In some embodiments, the nucleic acid strand is further coupled to a support. In some embodiments, the support is a bead. In some embodiments, the nucleic acid strand is releasably coupled to the support via an enzymatically labile or photo-labile linker.
[0044] In some embodiments, the composition further comprises a transposase. In some embodiments, the transposase is a transposase dimer. In some embodiments, the transposase is a Tn5 transposase. In some embodiments, the composition further comprises a restriction enzyme. In some embodiments, the restriction enzyme is PvuII.
[0045] In some embodiments, the present disclosure provides a method, comprising generating, from the composition, (I) a first adapter construct comprising (A) a first doublestranded transposon end comprising the first copy of the transposon end recognition sequence, and (B) the first barcode sequence; and (II) a second adapter construct comprising (A) a second doublestranded transposon end comprising the second copy of the transposon end recognition sequence, and (B) the second barcode sequence. In some embodiments, the method further comprisesAtty Dkt No.: 68244-703601 generating a transpososome comprising (I) the first adapter construct, (II) the second adapter construct, and (III) a transposase dimer. In some embodiments, generating the first adapter construct and the second adapter construct comprises cleaving the nucleic acid strand at a first site adjacent to and 3’ of the first copy of the transposon end recognition sequence. In some embodiments, the method further comprises performing the cleaving at the first site using the transposase dimer. In some embodiments, the method further comprises performing the cleaving at the first site using a restriction enzyme prior to generating the transpososome. In some embodiments, the restriction enzyme is PvuII.
[0046] In some embodiments, generating the first adapter construct and the second adapter construct further comprises cleaving the nucleic acid strand at a second site adjacent to and 3’ of the second copy of the transposon end recognition sequence. In some embodiments, the method further comprises performing the cleaving at the second site using the transposase dimer. In some embodiments, the method comprises performing the cleaving at the second site using a restriction enzyme prior to generating the transpososome. In some embodiments, the restriction enzyme is PvuII.
[0047] In some embodiments, the present disclosure provides a method comprising, (a) generating, from the composition, (I) a first adapter construct comprising (A) a first doublestranded transposon end comprising the first copy of the transposon end recognition sequence, and (B) the first barcode sequence; and (II) a second adapter construct comprising (A) a second doublestranded transposon end comprising the second copy of the transposon end recognition sequence, and (B) the second barcode sequence; and (b) removing the one or more linker polynucleotides from the first adapter construct and the second adapter construct. In some embodiments, the method further comprises performing (b) after generating a transpososome comprising (I) the first adapter construct, (II) the second adapter construct, and (III) a transposase dimer.
[0048] In some embodiments, the nucleic acid strand of the composition is releasably coupled to a support, wherein the method further comprises: (i) hybridizing the first oligonucleotide and the second oligonucleotide to the nucleic acid strand to yield a hybridized complex; (ii) releasing the hybridized complex from the support; and (iii) after (ii) generating the first adapter construct and the second adapter construct from the hybridized complex.
[0049] In some embodiments, the method further comprises, prior to (i), coupling the nucleic acid strand to the support. In some embodiments, the method further comprises, prior to generating the first adapter construct and the second adapter construct, sequencing the nucleic acid strand or a copy or derivative thereof to obtain pairing information identifying the first barcode sequence as paired with the second barcode sequence.Atty Dkt No.: 68244-703601
[0050] In some embodiments, the method further comprises performing a transposition reaction using the first adapter construct, the second adapter construct, and a target nucleic acid molecule, thereby generating a plurality of transposition products comprising (I) a first transposition product comprising the first barcode sequence and a first target sequence of the target nucleic acid molecule or reverse complement thereof, and (II) a second transposition product comprising the second barcode sequence and a second target sequence of the target nucleic acid molecule or reverse complement thereof. In some embodiments, the method further comprises sequencing the first transposition product or derivative thereof and the second transposition product or derivative thereof. In some embodiments, the method further comprises identifying the first target sequence and the second target sequence as being adjacent to each other on the target nucleic acid molecule based on the first barcode sequence and the second barcode sequence. In some embodiments, the method further comprises generating an assembled sequence comprising the first target sequence and the second target sequence using the first barcode sequence and the second barcode sequence.
[0051] In some embodiments, the method further comprises: generating a plurality of unique pairs of adapter constructs having unique pairs of barcode sequences; performing transposition reactions using the plurality of unique pairs of adapter constructs and the target nucleic acid molecule or derivative thereof to generate transposition products comprising barcode sequences of the unique pairs of barcode sequences; and generating an assembled sequence from the transposition products using the barcode sequences in the transposition products based on pairing information of the unique pairs of barcode sequences.
[0052] In another aspect, the present disclosure provides a composition comprising an adapter pair comprising: (i) a first adapter construct comprising a first double-stranded transposon end comprising a first transposon end recognition sequence; (ii) a second adapter construct comprising a second double-stranded transposon end comprising a second transposon recognition sequence; wherein the first adapter construct and the second adapter construct are coupled to each other via one or more linker polynucleotides.
[0053] In some embodiments, the one or more linker polynucleotides comprises a unimolecular linker polynucleotide comprising (i) a first capture sequence hybridized to a first binding site in the first adapter construct; and (ii) a second capture sequence hybridized to a second binding site in the second adapter construct. In some embodiments, the composition further comprises a plurality of linker polynucleotide molecules coupling the first segment of the nucleic acid strand to the second segment of the nucleic acid strand, wherein the plurality of linker polynucleotides comprises (i) a first linker polynucleotide molecule hybridized to a first bindingAtty Dkt No.: 68244-703601 site in the first adapter construct; and (ii) a second linker polynucleotide molecule hybridized to a second binding site in the second adapter construct.
[0054] In some embodiments, the first transposon end recognition sequence and the second transposon end recognition sequence are identical sequences. In some embodiments, the first transposon end recognition sequence and the second transposon end recognition sequence are different sequences. In some embodiments, the first transposon end recognition sequence and the second transposon end recognition sequence each comprise any one of SEQ ID NOs: 1-6.
[0055] In some embodiments, the first adapter construct further comprises a first barcode sequence and wherein the second adapter construct further comprises a second barcode sequence. In some embodiments, the first barcode sequence is located 5’ of the first transposon end recognition sequence, and wherein the second barcode sequence is located 5’ of the second transposon end recognition sequence. In some embodiments, the first barcode sequence and the second barcode sequence are identical sequences. In some embodiments, the first barcode sequence and the second barcode sequence are different sequences. In some embodiments, the composition is among one or more different adapter constructs that are not coupled to the first adapter construct or the second adapter construct. In some embodiments, the one or more different adapter constructs comprise one or more barcode sequences different from the first barcode sequence and the second barcode sequence. In some embodiments, the first barcode sequence and the second barcode sequence together identify the first adapter construct and the second adapter construct as associated with each other.
[0056] In some embodiments, the first adapter construct and the second adapter construct each further comprises one or more primer binding sequences. In some embodiments, the composition further comprises a transposase. In some embodiments, the adapter pair is bound to a transposase dimer. In some embodiments, the transposase is a Tn5 transposase.
[0057] In some embodiments, the present disclosure provides a system comprising the composition, and further comprising an additional adapter pair that is different from the adapter pair, wherein the additional adapter pair comprises a third adapter construct and a fourth adapter construct, wherein the third adapter construct and the fourth adapter construct are coupled to each other via one or more additional linker polynucleotides. In some embodiments, the third adapter construct comprises a third barcode sequence, wherein the fourth adapter construct comprises a fourth barcode sequence, wherein the third barcode sequence and the fourth barcode sequence together identify the third adapter construct and the fourth adapter construct as associated with each other. In some embodiments, the system further comprises at least 10 unique pairs of adapter constructs having unique pairs of barcode sequences.Atty Dkt No.: 68244-703601
[0058] In some embodiments, the present disclosure provides a method, comprising: (a) providing a transpososome comprising the composition and a transposase dimer; and (b) removing the one or more linker polynucleotides from the first adapter construct and the second adapter construct in the transpososome. In some embodiments, the method further comprises generating the transpososome. In some embodiments, the method further comprises, after (b), performing a transposition reaction using the first adapter construct, the second adapter construct, and a target nucleic acid molecule, thereby generating a plurality of transposition products comprising (I) a first transposition product comprising the first barcode sequence and a first target sequence of the target nucleic acid molecule or reverse complement thereof, and (II) a second transposition product comprising the second barcode sequence and a second target sequence of the target nucleic acid molecule or reverse complement thereof. In some embodiments, the method further comprises sequencing the first transposition product or derivative thereof and the second transposition product or derivative thereof. In some embodiments, the method further comprises identifying the first target sequence and the second target sequence as being adjacent to each other on the target nucleic acid molecule based on the first barcode sequence and the second barcode sequence. In some embodiments, the method further comprises generating an assembled sequence comprising the first target sequence and the second target sequence using the first barcode sequence and the second barcode sequence.
[0059] In some embodiments, the method further comprises: (c) providing an additional transpososome comprising an additional adapter pair and a transposase dimer, wherein the additional adapter pair comprises a third adapter construct and a fourth adapter construct, wherein the third adapter construct comprises a third barcode sequence and the fourth adapter construct comprises a fourth barcode sequence, and wherein the third adapter construct and the fourth adapter construct are coupled to each other via one or more additional linker polynucleotides; and (d) removing the one or more additional linker polynucleotides from the third adapter construct and the fourth adapter construct in the additional transpososome. In some embodiments, the method further comprises, after (d), performing an additional transposition reaction using the third adapter construct, the fourth adapter construct, and the target nucleic acid molecule or a derivative thereof, thereby generating an additional plurality of transposition products comprising (I) a third transposition product comprising the third barcode sequence and a third target sequence of the target nucleic acid molecule or reverse complement thereof, and (II) a fourth transposition product comprising the fourth barcode sequence and a fourth target sequence of the target nucleic acid molecule or reverse complement thereof. In some embodiments, the method further comprises sequencing the third transposition product or derivative thereof and the fourth transposition product or derivative thereof. In some embodiments, the method further comprises identifying theAtty Dkt No.: 68244-703601 third target sequence and the fourth target sequence as being adjacent to each other on the target nucleic acid molecule based on the third barcode sequence and the fourth barcode sequence. In some embodiments, the method further comprises generating an assembled sequence comprising the first target sequence, the second target sequence, the third target sequence, and the fourth target sequence using the first barcode sequence, the second barcode sequence, the third barcode sequence, and the fourth barcode sequence.
[0060] In some embodiments, the method further comprises: providing at least 100 to 100 billion unique pairs of adapter constructs having unique pairs of barcode sequences; performing transposition reactions using the at least 100 to 1000 unique pairs of adapter constructs and the target nucleic acid molecule or derivative thereof to generate a plurality of transposition products comprising barcode sequences of the unique pairs of barcode sequences; and generating an assembled sequence from the plurality of transposition products using the barcode sequences in the plurality of transposition products based on pairing information of the unique pairs of barcode sequences. In some embodiments, the assembled sequence has a length of 500 to 10,000 nucleotides. In some embodiments, generating the assembled sequence comprises sequencing the plurality of transposition products or derivatives thereof to generate sequence reads, and assembling the sequence reads in silico to generate the assembled sequence. In some embodiments, the method further comprises generating the at least 100 to 1000 unique pairs of adapter constructs using a plurality of nucleic acid strands each comprising a unique pair of the unique pairs of barcode sequences; and obtaining the pairing information of the unique pairs of barcode sequences by sequencing the plurality of nucleic acid strands or copies thereof.
[0061] In another aspect, the present disclosure provides a method comprising: (a) providing a composition comprising: (i) a nucleic acid strand comprising a first transposon end recognition sequence and a second transposon end recognition sequence; (ii) a first oligonucleotide that is hybridized to the first transposon end recognition sequence; and (iii) a second oligonucleotide that is hybridized to the second transposon end recognition sequence; wherein the first oligonucleotide and the second oligonucleotide are separate oligonucleotides; and (b) from the composition, generating: (I) a first adapter construct comprising a first double-stranded transposon end comprising the first transposon end recognition sequence; and (II) a second adapter construct comprising a second double-stranded transposon end comprising the second transposon end recognition sequence. In some embodiments, the nucleic acid strand comprises (I) a first segment comprising a first barcode sequence and the first transposon end recognition sequence, and (II) a second segment comprising a second barcode sequence and the second transposon end recognition sequence. In some embodiments, the nucleic acid strand comprises, in order from 5’ to 3’ : the first barcode sequence, the transposon end recognition sequence, the second barcode sequence, and theAtty Dkt No.: 68244-703601 second copy of the two identical copies of the transposon end recognition sequence. In some embodiments, the method further comprises generating a transpososome comprising (I) the first adapter construct, (II) the second adapter construct, and (III) a transposase dimer. In some embodiments, generating the first adapter construct and the second adapter construct comprises cleaving the nucleic acid strand at a first site adjacent to and 3’ of the first transposon end recognition sequence. In some embodiments, the method further comprises performing the cleaving at the first site using the transposase dimer. In some embodiments, the method further comprises performing the cleaving at the first site using a restriction enzyme prior to generating the transpososome. In some embodiments, the restriction enzyme is PvuII.
[0062] In some embodiments, generating the first adapter construct and the second adapter construct further comprises cleaving the nucleic acid strand at a second site adjacent to and 3’ of the second transposon end recognition sequence. In some embodiments, the method further comprises performing the cleaving at the second site using the transposase dimer. In some embodiments, the method further comprises performing the cleaving at the second site using a restriction enzyme prior to generating the transpososome. In some embodiments, the restriction enzyme is PvuII.
[0063] In some embodiments, the method further comprises performing a transposition reaction using the first adapter construct, the second adapter construct, and a target nucleic acid molecule, thereby generating a plurality of transposition products comprising (I) a first transposition product comprising the first barcode sequence and a first target sequence of the target nucleic acid molecule or reverse complement thereof, and (II) a second transposition product comprising the second barcode sequence and a second target sequence of the target nucleic acid molecule or reverse complement thereof. In some embodiments, the method further comprises sequencing the first transposition product or derivative thereof and the second transposition product or derivative thereof. In some embodiments, the method further comprises identifying the first target sequence and the second target sequence as being adjacent to each other on the target nucleic acid molecule based on the first barcode sequence and the second barcode sequence. In some embodiments, the method further comprises generating an assembled sequence comprising the first target sequence and the second target sequence using the first barcode sequence and the second barcode sequence.
[0064] In some embodiments, the method further comprises: generating a plurality of unique pairs of adapter constructs having unique pairs of barcode sequences; performing transposition reactions using the plurality of unique pairs of adapter constructs and the target nucleic acid molecule or derivative thereof to generate a plurality of transposition products comprising barcode sequences of the unique pairs of barcode sequences; and generating an assembled sequence fromAtty Dkt No.: 68244-703601 the plurality of transposition products using the barcode sequences in the plurality of transposition products based on pairing information of the unique pairs of barcode sequences.
[0065] In some embodiments, the assembled sequence has a length of 500 to 10,000 nucleotides. In some embodiments, generating the assembled sequence comprises sequencing the plurality of transposition products or derivatives thereof to generate sequence reads, and assembling the sequence reads in silico to generate the assembled sequence. In some embodiments, generating the assembled sequence does not comprise further barcoding transposition products or derivatives thereof of the plurality of transposition products using additional barcode sequences that are different from the barcode sequences of the unique pairs of barcode sequences. In some embodiments, generating the assembled sequence does not comprise sequence assembly based on an overlap sequence of the target nucleic acid molecule that is common to transposition products of the plurality of transposition products. In some embodiments, generating the assembled sequence does not comprise reference-based assembly. In some embodiments, generating the assembled sequence does not comprise sequence alignment of transposition products or derivatives thereof of the plurality of transposition products.
[0066] In some embodiments, the method further comprises generating the plurality of unique pairs of adapter constructs using a plurality of nucleic acid strands each comprising a unique pair of the unique pairs of barcode sequences; and obtaining the pairing information of the unique pairs of barcode sequences by sequencing the plurality of nucleic acid strands or copies thereof. In some embodiments, in (b), the first adapter construct and the second adapter construct are coupled to each other via one or more linker polynucleotides. In some embodiments, in (a), the first segment of the nucleic acid strand and the second segment of the nucleic acid strand are coupled to each other via the one or more linker polynucleotides. In some embodiments, method, further comprises (c) removing the one or more linker polynucleotides from the first adapter construct and the second adapter construct. In some embodiments, (c) occurs after generating a transpososome comprising (I) the first adapter construct, (II) the second adapter construct, and (III) a transposase dimer. In some embodiments, (c) occurs prior to performing a transposition reaction using the first adapter construct, the second adapter construct, and the target nucleic acid molecule. In some embodiments, prior to or during (a), the nucleic acid strand of the composition is releasably coupled to a support, wherein the method further comprises releasing the composition from the support.
[0067] In some embodiments, the method further comprises performing the sequencing of the nucleic acid strand or a copy or derivative thereof using a primer comprising a locked nucleic acid (LNA) nucleotide. In some embodiments, the method performs the sequencing of the first transposition product or derivative thereof and the second transposition product or derivativeAtty Dkt No.: 68244-703601 thereof using a primer comprising a locked nucleic acid (LNA) nucleotide. In some embodiments, the method comprises performing the sequencing of the transposition products or derivatives thereof using a primer comprising a locked nucleic acid (LNA) nucleotide. In some embodiments, the method further comprises performing the sequencing of the plurality of nucleic acid strands or copies thereof using a primer comprising a locked nucleic acid (LNA) nucleotide.
[0068] In another aspect, the present disclosure provides a library of nucleic acid strands, wherein each nucleic acid strand comprises two identical copies of a transposon end recognition sequence and a unique pair of barcode sequences. In some embodiments, each nucleic acid strand comprises, in order from 5’ to 3 ’ : a first barcode sequence of the unique pair of barcode sequences, a first copy of the two identical copies of the transposon end recognition sequence, a second barcode sequence of the unique pair of barcode sequences, and a second copy of the two identical copies of the transposon end recognition sequence.
[0069] In some embodiments, the library comprises a nucleic acid strand comprising a restriction enzyme recognition sequence. In some embodiments, the nucleic acid strand comprises a restriction enzyme cut site located adjacent and 3’ to the first copy of the two identical copies of the transposon end recognition sequence. In some embodiments, the restriction enzyme recognition sequence is a PvuII restriction enzyme recognition sequence. In some embodiments, the nucleic acid strand further comprises an additional restriction enzyme recognition sequence. In some embodiments, the nucleic acid strand further comprises an additional restriction enzyme cut site located adjacent and 3’ to the second copy of the two identical copies of the transposon end recognition sequence. In some embodiments, the additional restriction enzyme recognition sequence is a PvuII restriction enzyme recognition sequence.
[0070] In another aspect, the present disclosure provides a method, comprising: (a) generating a plurality of transposition products via transposase-mediated fragmentation of a target nucleic acid molecule, wherein the plurality of transposition products comprises a plurality of unique pairs of barcode sequences, wherein a unique pair of barcode sequences of the plurality of unique pairs of barcode sequences identifies a first transposition product and a second transposition product of the plurality of transposition products as derived from sequences adjacent to each other within the target nucleic acid molecule; (b) generating a plurality of sequence reads from the plurality of transposition products; and (c) generating an assembled sequence from the plurality of sequence reads using the plurality of unique pairs of barcode sequences. As described herein, “transposase- mediated fragmentation” refers to fragmenting a nucleic acid molecule into separate nucleic acid fragments by one or more transposases.
[0071] In some embodiments, the method does not comprise further fragmentation of the plurality of transposition products by a non-transposase enzyme. In some embodiments, theAtty Dkt No.: 68244-703601 method does not comprise further ligating together transposition products of the plurality of transposition products or derivatives thereof. In some embodiments, the transposase-mediated fragmentation comprises performing a transposition reaction between the target nucleic acid molecule and a transposon composition, wherein the transposon composition comprises two double-stranded transposon ends but is not fully double-stranded.
[0072] In some embodiments, the transposon composition comprises (I) a first adapter construct comprising a first double-stranded transposon end comprising a first transposon end recognition sequence and a first barcode sequence of a unique pair of barcode sequences of the plurality of unique barcode sequences; and (II) a second adapter construct comprising a second double-stranded transposon end comprising a second transposon recognition sequence and a second barcode sequence of the unique pair of barcode sequences.
[0073] In some embodiments, the first adapter construct and the second adapter construct are assembled together with a same transposase dimer. In some embodiments, the first adapter construct and the second adapter construct are coupled to each other via one or more linker polynucleotides. In some embodiments, the method comprises generating the transposon composition from a precursor composition comprising: (i) a nucleic acid strand comprising two identical copies of a transposon end recognition sequence; (ii) a first oligonucleotide comprising a reverse complement of the transposon end recognition sequence; and (iii) a second oligonucleotide comprising a reverse complement of the transposon end recognition sequence. In some embodiments, each unique pair of barcodes sequences of the plurality of unique pairs of barcode sequences does not comprise complementary sequences.
[0074] In another aspect, the present disclosure provides a method, comprising generating an assembled sequence from a plurality of sequence reads using a plurality of unique pairs of barcode sequences, wherein each unique pair of barcode sequences of the plurality of unique pairs of barcode sequences identifies two sequence reads of the plurality of sequence reads as being paired, wherein each unique pair of barcodes sequences does not comprise identical or complementary barcode sequences.
[0075] In some embodiments, the plurality of unique pairs of barcode sequences comprises at least 1 million unique pairs of barcode sequences. In some embodiments, the method does not comprise sequence assembly based on an overlap sequence of the target nucleic acid molecule that is common to transposition products of the plurality of transposition products. In some embodiments, the method does not comprise reference-based assembly. In some embodiments, the method does not comprise sequence alignment of sequence reads of the plurality of sequence reads.Atty Dkt No.: 68244-703601INCORPORATION BY REFERENCE
[0076] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference.BRIEF DESCRIPTION OF THE DRAWINGS
[0077] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings of which:
[0078] FIG. 1 A shows an example flowchart for barcoding target nucleic acid using an adapter construct and target nucleic acid sequence reconstruction. FIG. IB shows a schematic of an example adapter construct and transposase-mediated insertion of the adapter construct. FIG. 1C shows an example adapter construct.
[0079] FIG. 2 shows a schematic of barcoding an adapter construct using a barcode construct.
[0080] FIG. 3 shows an example adapter construct that results in a barcode construct connection.
[0081] FIG. 4 shows an example adapter construct that results in a barcode construct connection.
[0082] FIG. 5 shows an example adapter construct that results in a combined barcode construct and cis transposase monomer connection.
[0083] FIG. 6 shows an example adapter construct that results in a cis transposase monomer connection and a combined barcode construct and cis transposase monomer connection.
[0084] FIG. 7 shows an example adapter construct that results in a cis transposase monomer connection.
[0085] FIG. 8 shows an example adapter construct that results in a trans transposase monomer connection, and combined barcode construct and cis transposase monomer connection, and a combined cis and trans transposase monomer connection.
[0086] FIG. 9 shows an example adapter construct that results in a combined cis transposase monomer connection and trans transposase monomer connection.
[0087] FIG. 10 shows an example adapter construct that results in a combined barcode construct and trans transposase monomer connection.
[0088] FIG. 11 shows an example adapter construct that results in a trans transposase monomer connection, a combined barcode construct connection and cis transposase monomerAtty Dkt No.: 68244-703601 connection, and a combined cis transposase monomer connection and trans transposase monomer connection.
[0089] FIG. 12 shows an example adapter construct that results in a combined cis transposase monomer connection and trans transposase monomer connection.
[0090] FIG. 13A shows an example adapter construct that results in a combined cis transposase monomer connection and trans transposase monomer connection. FIG. 13B shows an example adapter construct that results in a trans transposase monomer connection.
[0091] FIG. 14 shows a schematic of information flow in a barcode construct connection.
[0092] FIGs. 15A and 15B show an example of transposase-mediated insertion of an adapter construct resulting in a barcode construct connection. FIG. 15A shows an adapter construct coupled to a barcode construct. FIG. 15B shows a plurality of barcoded adapter constructs inserted in target nucleic acid after transposase-mediated insertion.
[0093] FIG. 16 shows a schematic of information flow in a cis transposase monomer connection.
[0094] FIGs. 17A and 17B show an example of transposase-mediated insertion of an adapter construct resulting in a combined barcode construct connection and a cis transposase monomer connection. FIG. 17A shows generation of an adapter construct and coupling to a barcode construct. FIG. 17B shows a plurality of barcoded adapter constructs inserted in target nucleic acid after transposase-mediated insertion.
[0095] FIG. 18 shows a schematic of information flow in a trans transposase monomer connection.
[0096] FIGs. 19A and 19B show an example of transposase-mediated insertion of an adapter construct resulting in a trans transposase monomer connection. FIG. 19A shows generation of a barcoded adapter construct. FIG. 19B shows insertion of a barcoded adapter construct in target nucleic acid.
[0097] FIGs. 20A-20C show example of transposase-mediated insertion of an adapter construct resulting in a barcode construct connection and a trans transposase monomer connection. FIG. 20A shows generation of a barcoded adapter construct. FIG. 20B shows coupling to a barcode construct and insertion of a plurality of barcoded adapter constructs in target nucleic acid. FIG. 20C shows generation of a barcoded fragment for sequencing.
[0098] FIG. 21 A shows an example of edge detection using a barcoded adapter construct using an end-sequence. FIG. 21B shows an example of edge detection using a barcoded adapter construct using a UMI sequence.
[0099] FIGs. 22A-22H show an example workflow of adapter construct generation, transposase-mediated insertion, and sequence library preparation. FIG. 22A shows an exampleAtty Dkt No.: 68244-703601 precursor of an adapter construct. FIG. 22B shows an example adapter construct coupled to a barcode construct. FIG. 22C shows a second example adapter construct coupled to a barcode construct. FIG. 22D shows insertion products of transposase-mediated insertion of adapter constructs shown in FIG. 22B and FIG. 22C. FIG. 22E shows processed insertion products after extension and ligation. FIG. 22F shows amplification of top strands of processed insertion products. FIG. 22G shows amplification of bottom strands of processed insertion products. FIG. 22H shows example library sequence structures for sequencing.
[0100] FIGs. 23A-23G show an example workflow of adapter construct generation, transposase-mediated insertion, and sequence library preparation. FIG. 23A shows an example precursor of an adapter construct. FIG. 23B shows coupling of an example adapter construct to a barcode construct. FIG. 23C shows insertion products of transposase-mediated insertion of adapter constructs shown in FIG. 23B. FIG. 23D shows processed insertion products after extension and ligation. FIG. 23E shows amplification of top strands of processed insertion products. FIG. 23F shows amplification of bottom strands of processed insertion products. FIG. 23G shows example library sequence structures for sequencing.
[0101] FIGs. 24A-24E show an example workflow of transposase-mediated insertion of an adapter construct and sequence library preparation. FIG. 24A shows an example of an adapter construct. FIG. 24B shows insertion products of transposase-mediated insertion of an example adapter construct. FIG. 24C shows example processed insertion products of an adapter construct. FIGs. 24D and FIG. 24E show example library sequence structures for sequencing.
[0102] FIG. 25 provides examples of barcode constructs.
[0103] FIG. 26 provides an example schematic of generating a barcode construct.
[0104] FIG. 27 provides an example of generating adapter constructs from an example composition comprising two identical copies of a transposase recognition sequence.
[0105] FIG. 28 provides an example of generating adapter constructs from an example composition comprising a linker polynucleotide.
[0106] FIG. 29 provides an example of generating adapter constructs from an example composition comprising a linker polynucleotide. In this figure, the linker polynucleotide 2903 comprises oligonucleotide 2901.
[0107] FIG. 30 provides an example of generating adapter constructs from an example composition comprising two linker polynucleotides.
[0108] FIG. 31 provides an example of generating adapter constructs from an example composition comprising two linker polynucleotides. In this figure, the linker polynucleotide 3104 comprises oligonucleotide 3101.Atty Dkt No.: 68244-703601
[0109] FIGs. 32A-E provide an example workflow of transposase-mediated insertion of adapter constructs and sequence library preparation. FIG. 32A shows insertion products of transposase-mediated insertion. FIG. 32B show example processed insertion products following DNA polymerase strand displacement extension. FIG. 32C shows PCR amplification of the top strands. FIG. 32D shows PCR amplification of the bottom strands. FIG. 32E shows example library sequence structures for sequencing.
[0110] FIGs. 33A and 33B provide an example workflow of generating an example composition comprising two identical copies of a transposase recognition sequence coupled to a bead and characterization. FIG. 33A shows PCR amplification and splitting a PCR amplified library for sequencing characterization and transposon preparation. FIG. 33B show an example of transposon preparation via coupling to a bead.
[0111] FIG. 34 provides an example of generating adapter constructs from an example composition comprising two identical copies of a transposase recognition sequence. In this figure, a transposase cleaves at the cut sites indicated by *.
[0112] FIGs. 35A and 35B provide an example of generating adapter constructs from an example composition comprising two identical copies of a transposase recognition sequence. In this figure, a restriction enzyme PvuII cleaves at the cut sites indicated by *.DETAILED DESCRIPTIONOverview
[0113] Recognized herein is a need for improved methods for nucleic acid sequencing and for genomics applications to improve accuracy, efficiency, and throughput. Many sequencing methods are limited by read length and face complications in the reconstruction short reads of amplified fragments into the original sequence of the long nucleic acid molecule. Assembling short reads into long reads can benefit from information about the relative positions of the short reads. Recognized herein is a need to preserve information about the origin (e.g., identifying information about the biological sample from which the reads are derived). Current de novo assembly techniques rely on extensive overlapping regions, which introduces redundancy and inefficiency in sequencing library preparation. Using overlapping regions between sequencing reads to assemble short reads into long reads can be extremely computationally intensive, scaling to N square with N being the number of reads) and can be error-prone even after expending a lot of resources due to repetitive sequences. Recognized herein is a need to provide higher accuracy and contiguity, improved efficiency, and simpler workflows.
[0114] In some aspects, the present disclosure provides long-read sequencing methods and compositions used for long-read sequencing methods that do not require overlap-based assemblyAtty Dkt No.: 68244-703601 or even any sequence alignment. The long-read sequencing methods described herein can reconstruct long (e.g., over 10,000 bp) molecules by concatenating reads with barcode pairs. Computationally, the assembly methods described herein can scale linearly with the number of reads, which is far more efficient than the quadratic read assembly methods of many other sequencing technologies. In some aspects, by utilizing barcode pairs to link together reads, the methods described herein do not require overlaps between reads or sequence alignment to assemble the reads. Compared to traditional overlap-based methods, the methods described herein can also achieve far higher accuracy and eliminate error-correction heuristics.
[0115] In some aspects, long-read sequencing methods described herein are able to reconstruct long molecules using a high number (e.g., 10-10,000) of reads with high accuracy by utilizing multiple unique pairs of barcodes during tagmentation and controlling the tagmentation reactions such that a given unique pair of barcodes is assembled together with a same transposase dimer for tagmentation to ensure that adjacent reads are tagged by a given unique pair of barcodes. In some cases, each barcode of a pair of barcodes is associated with and linked to a separate 3’ transposon end during tagmentation, such that the two barcodes of the pair end up in separate transferred strands after tagmentation. Identifying the barcodes in the transposition products allow adjacent reads to be linked together computationally without performing additional gap-fill and ligation reactions on the transposition products. The methods described herein also allow adjacent reads to be linked together computationally without performing additional fragmentation reactions on the transposition products. In some aspects, the long-read sequencing methods described herein utilize at least 1 million unique pairs of barcodes. In some aspects, the present disclosure provides unique methods and compositions to control the assembly of a given pair of barcodes with a same transposase dimer for tagmentation such that the pairing information is maintained when tagging adjacent reads. For example, in some aspects, the method can comprise assembling a transposase dimer together with a pair of adapter constructs (each comprising a barcode of the given pair of barcodes) that are coupled to each other via one or more linker polynucleotides. The pair of adapter constructs can stay assembled with the transposase dimer after the one or more linker polynucleotides are removed. As another example, in some aspects, the method can comprise assembling a transposase dimer together with an adapter precursor of the pair of adapter constructs each comprising a barcode of the given pair of barcodes. This can be followed by processing the adapter precursor (e.g., via cleaving) to yield the pair of adapter constructs that are assembled with the transposase dimer. In some aspects, the present disclosure provides a unique composition comprising a pair of adapter constructs that are coupled together via one or more linker polynucleotides. In some aspects, the present disclosure provides a unique adapter precursorAtty Dkt No.: 68244-703601 composition that is configured to assemble with a transposase dimer and is configured to be cleaved to yield a pair of adapter constructs that can be used for tagmentation.
[0116] Barcoding at tagmentation (by providing barcode pairs in the adapter constructs for augmentation) can have advantages compared to barcoding after tagmentation (e.g., via beads carrying many copies of a single barcode that are used to barcode multiple fragmented transposition products of a single long nucleic acid molecule). For example, the methods described herein impart unique pairing information of adjacent reads to achieve continuous long read assembly in contrast to discontiguous short reads that originate from the same parent molecule. This unique pairing information may not be achieved by using beads carrying copies of a single barcode and transferring the monoclonal barcode information to multiple already fragmented transposition products. Furthermore, the methods described herein can achieve high synthetic read lengths (N50) (e.g., at least 800 nucleotides). The methods described herein can also result in barcoding efficiency of almost 100%, which can be double the efficiency of bead workflows for barcoding after tagmentation. Additionally, the methods described herein may not incur the efficiency losses that result in barcoding with beads after tagmentation, and can reduce reagent and time costs by at least 30% compared to bead-based barcoding after tagmentation.
[0117] In some aspects, the present disclosure provides compositions and methods for transposase-mediated insertion (e.g., tagmentation) of nucleic acid molecules to barcode target DNA sequences for sequencing nucleic acid molecules. The compositions and methods described herein can utilize unique barcoded adapter constructs to enable accurate and efficient reconstruction of short reads into long assembled sequences. FIG. 1A shows an example workflow of a method of barcoding target nucleic acid using an adapter construct and sequence assembly. In the example provided in FIG. 1A, an adapter construct comprising a barcode sequence is provided. The method can comprise inserting the adapter construct into a target nucleic acid using transposase-mediated insertion to generate barcoded insertion products comprising the barcode sequence or reverse complement thereof. The method can comprise sequencing the barcoded insertion product to generate short reads. The method can further comprise linking the short reads using the barcode sequence to generate an assembled sequence.
[0118] In some aspects, the unique structure of the adapter construct and the barcodes provided in the adapter constructs preserve information about the relative positions of different short reads and enable the linking of adjacent short reads during reconstruction of the original sequence. In some aspects, the combination of adapter constructs and additional barcode constructs (e.g., clonal webs or DNA origami constructs) are utilized to capture further identifying information (e.g., origin information) about the short read sequence. The methods and compositions described herein can permit analysis of large structural variation, repeat expansions,Atty Dkt No.: 68244-703601 and haplotypic phasing in human genomes. Additionally, the methods and compositions described herein can permit analysis of nucleic acid molecules from multiple biological samples and preserve sample-identifying information in the assembled sequence.
[0119] In some aspects, the present disclosure provides compositions (e.g., adapter constructs) for transposase-mediated insertion of nucleic acid molecules. In some aspects, the present disclosure provides nucleic acid molecules comprising a transposon end recognition sequence (e.g., a mosaic end sequence, an inside end sequence, an outside end sequence, a reverse complement thereof, etc.). A transposon end recognition sequence can bind to a transposase domain. In some cases, a pair of DNA transposon ends, each comprising a transposon end recognition sequence, can assemble with an active transposase to form a transpososome complex for transposase-mediated insertion into a target nucleic acid molecule. An active transposase can be a dimer of two transposase monomers. In some cases, each transposase monomer binds to a DNA transposon end of the pair of DNA transposon ends in the transpososome complex. A DNA transposon end can be single-stranded. Alternatively, a DNA transposon end can be doublestranded.
[0120] In some aspects, the present disclosure provides a nucleic acid construct or a plurality of polynucleotides comprising a pair of transposon end recognition sequences. In some cases, the nucleic acid construct or the plurality of polynucleotides comprises a single-stranded DNA transposon end. In some cases, the nucleic acid construct or the plurality of polynucleotides comprises a pair of double-stranded DNA transposon ends. In some aspects, the two DNA transposon ends are directly coupled. The two DNA transposon ends can be coupled covalently or non-covalently.
[0121] In some aspects, a nucleic acid construct or a plurality of polynucleotides described herein further comprises one or more barcode sequences. The one or more barcode sequences can be used for barcoding one or more target sequences of a target nucleic acid molecule upon transposase-mediated insertion. FIG. IB shows a schematic of transposase-mediated insertion of an adapter construct. In FIG. IB, adapter construct 110 comprising transposon end recognition sequences 113 and 114 assembles with active transposases 101 to form a transpososome complex. The adapter construct 110 further comprises barcode sequences 111 and 112. The active transposases 101 inserts adapter construct 110 into a target nucleic acid 123 to generate an insertion product comprising the barcode sequences 111 and 112.
[0122] In some aspects, a nucleic acid construct or plurality of polynucleotides described herein enable copying of barcode sequences from a barcode construct (e.g., a clonal web or a DNA origami construct) comprising one or more barcode sequences and barcoding one or more targetAtty Dkt No.: 68244-703601 sequences of a target nucleic acid molecule with the copied barcode sequences upon transposase- mediated insertion.
[0123] FIG. 2 shows an example of an adapter construct 230 comprising two transposon ends comprising transposon end recognition sequences 233 and 234. Adapter construct 230 further comprises barcode sequences 231 and 232. Adapter construct 2310 further comprises binding sequences 235 and 236. Binding sequence 235 binds to capture sequence 241 and binding sequence 236 binds to capture sequence 242 of a barcode construct 240. Adapter construct 230 can enable copying barcode sequence 243 of the barcode construct by complementary DNA (cDNA) synthesis to generate the reverse complement of barcode sequence 243.
[0124] In some aspects, the nucleic acid construct or a plurality of polynucleotides described herein enables a barcode construct connection (e.g., type A connection) or the sharing / copying of barcode information between the nucleic acid construct or the plurality of polynucleotides and a separate barcode construct comprising one or more barcodes, as depicted in FIG. 14. A barcode construct connection (e.g., type A connection) can group reads having a same barcode sequence originating from the barcode construct from which the barcode sequence was copied. A type A linkage can enable construction of long reads by assembling short reads that share a barcode sequence copied from the barcode construct.
[0125] In some cases, barcodes within the nucleic acid construct and the architecture of the nucleic acid construct enable a trans transposase monomer connection (e.g., type C linkage) or the sharing of barcode information across two transposase monomers of a same transposase dimer for sequence assembly of barcoded short reads, as depicted in FIG. 18. A trans transposase monomer connection (e.g., type C connection) pairs two adjacent reads or fragments generated by two different monomers of the same transposase dimer. A trans transposase monomer connection (e.g., type C connection) can enable the linking of adjacent reads made by monomers of the same transposase dimer and can enable assembly of adjacent reads without overlap.
[0126] In some aspects, barcodes within the nucleic acid construct and the architecture of the nucleic acid construct enable a cis transposase monomer connection (e.g., type B connection) or the copying / sharing of barcode information between strands that bind to the same transposase monomer, as depicted in FIG. 16. A cis transposase monomer connection (e.g., type B connection) pairs reads made from the sense and antisense strands of the same double stranded DNA fragment. A cis transposase monomer connection (e.g., type B connection) can enable the linking of reads from different barcode constructs that tag the same target nucleic acid molecule. This can enable use of multiple barcode constructs to tag a target nucleic acid molecule. A cis transposase monomer connection (e.g, type B connection) can also enable linking between reads that have type C linkages to enable longer sequence continuity in assembly. A cis transposase monomerAtty Dkt No.: 68244-703601 connection (e.g., type B connection) and a trans transposase monomer connection (e.g., type C connection) can enable long read assembly without a barcode construct connection (e.g., type A connection).COMPOSITIONS AND SYSTEMSA. Adapter constructs and precursors thereof
[0127] In some aspects, the present disclosure provides compositions (e.g., adapter constructs) for transposase-mediated insertion of nucleic acid molecules. In some aspects, the present disclosure provides nucleic acid molecules comprising a transposon end recognition sequence. A transposon end recognition sequence can be an inside end (IE) sequence or an outside end (OE) sequence, or a reverse complement thereof. In some cases, the transposon end recognition sequence comprises any one of SEQ ID NOs: 1, 2, 4, or 5. In some cases, the transposon end recognition sequence is a mosaic end sequence (ME) or a reverse complement thereof. The mosaic sequence can be a hyperactive related sequence of the IE or OE sequence. In some cases, the mosaic end sequence comprises SEQ ID NO: 3 or 6. A transposon end recognition sequence can bind to a transposase domain. In some cases, a pair of DNA transposon ends, each comprising a transposon end recognition sequence, can assemble with an active transposase to form a transpososome complex for transposase-mediated insertion into a target nucleic acid molecule. An active transposase can be a dimer of two transposase monomers. In some cases, each transposase monomer binds to a DNA transposon end of the pair of DNA transposon ends in the transpososome complex. A DNA transposon end can be single-stranded. Alternatively, a DNA transposon end can be double-stranded. An adapter construct used for transpose-mediated insertion can comprise a transposon end recognition sequence.Table 1. Transposon end recognition sequences
[0128] In some aspects, the present disclosure provides an adapter construct comprising plurality of polynucleotides comprising a pair of transposon end recognition sequences. In some cases, the adapter construct comprises a single-stranded DNA transposon end. In some cases, theAtty Dkt No.: 68244-703601 adapter construct comprises a double-stranded DNA transposon end. In some cases, the adapter construct comprises a pair of transposon ends configured to assemble with an active transposase to form a transpososome complex. An example is shown in FIG. IB. The adapter construct in FIG. IB comprises a first double-stranded DNA transposon end 110A comprising transposon end recognition sequence 113 and a second double-stranded DNA transposon end HOB comprising transposon end recognition sequence 114. In FIG. IB, the two DNA transposon ends bind to active transposase 101.
[0129] In some aspects, the adapter construct further comprises one or more functional sequences. The one or more functional sequences can be configured to enable further processing of nucleic acid products generated after transposase-mediated insertion of the adapter construct into a target nucleic acid. For example, a functional sequence can be a primer binding sequence or reverse complement thereof, an adapter sequence, barcode sequence, or a distinguishing sequence that can distinguish a nucleic acid strand from a reverse complement. The one or more functional sequences provided by the adapter construct can be important for sequencing library preparation.
[0130] In some aspects, the adapter construct further comprises one or more barcode sequences (e.g., a unique molecular index or a sample barcode sequence). In some aspects, the adapter construct comprises a binding sequence configured to bind to a barcode construct described elsewhere herein to facilitate barcoding of the adapter construct or a derivative thereof (e.g., via nucleic acid extension to generate a barcoded product comprising a reverse complement of a barcode sequence of the barcode construct). In some cases, an adapter construct comprises a barcoded derivative of another adapter construct described herein. For example, a second adapter construct can be generated from a first adapter construct and a barcode construct described elsewhere herein following hybridization of a binding sequence of the first adapter construct to a capture sequence of the barcode construct and nucleic acid extension and complementary DNA synthesis using barcode construct as a template. The second adapter construct can comprise a barcode sequence or reverse complement thereof of the barcode construct.
[0131] In some aspects, the adapter construct further comprises a binding sequence that is configured to hybridize to an additional nucleic acid molecule. The additional nucleic acid molecule can be a barcode construct described elsewhere herein. In some cases, the binding sequence of the adapter construct hybridizes to a capture sequence of the barcode construct. In some embodiments, the binding sequence is on a 3’ end of a polynucleotide of the adapter construct. A same polynucleotide can comprise both the binding sequence and a transposon end recognition sequence. A same polynucleotide can comprise both the binding sequence and a reverse complement of a transposon end recognition sequence. In some cases, a sameAtty Dkt No.: 68244-703601 polynucleotide comprises a transposon end recognition sequence, a binding sequence, and a barcode sequence (e.g., a unique molecular index).
[0132] In some aspects, the two DNA transposon ends are directly coupled. The two DNA transposon ends can be coupled covalently or non-covalently. In some aspects, the two DNA transposon ends are covalently coupled. For example, FIG. IB provides a nucleic acid construct 110 comprising three nucleic acid strands HOi, HOii, and HOiii. The first double-stranded transposon end 110A is formed by the ends of nucleic acid strands 1 lOi and 1 lOii, and the second double-stranded transposon end 11 OB is formed by the ends of nucleic acid strands 1 lOi and 1 lOiii. The two transposon ends in FIG. IB are covalently coupled via the nucleic acid strand HOi. In other aspects, the two DNA transposon ends are non-covalently coupled. For example, FIG. 1C provides a nucleic acid construct comprising four nucleic acid strands 140i, 140ii 140iii, and 140iv. The first transposon end 140A is formed by the ends of nucleic acid strands 140i and 140iii, and the second transposon end 140B is formed by the ends of nucleic acid strands 140ii and 140iv. The two transposon ends in FIG. 1C are coupled via hybridization between nucleic acid strands 140i and 140ii.
[0133] In one aspect, the present disclosure provides a nucleic acid construct (e.g., an adapter construct) comprising: a first nucleic acid strand comprising (A) a first transposon end recognition sequence, and (B) a reverse complement of a second transposon end recognition sequence. The first nucleic acid strand can comprise (C) two discrete barcode sequences located between the first transposon end recognition sequence and the second transposon end recognition sequence. The nucleic acid construct can comprise a second nucleic acid strand comprising a reverse complement of the first transposon end recognition sequence, wherein a portion of the second nucleic acid strand is hybridized to a first segment of the first nucleic acid strand. The nucleic acid construct can comprise a third nucleic acid strand comprising the second transposon end recognition sequence, wherein a portion of the third nucleic acid strand is hybridized to a second segment of the first nucleic acid strand. In some cases, the second nucleic acid strand and the third nucleic acid strand are separate nucleic acid strands.
[0134] In some embodiments, the second nucleic acid strand or the third nucleic acid strand further comprises a reverse complement of a discrete barcode sequence of the two discrete barcode sequences. In some cases, the second nucleic acid strand and the third nucleic acid strand each comprises a reverse complement of a discrete barcode sequence of the two discrete barcode sequences. In other embodiments, the second nucleic acid strand does not comprise a reverse complement of a discrete barcode sequence of the two discrete barcode sequences. In other embodiments, the second nucleic acid strand and the third nucleic acid strand do not comprise a reverse complement of a discrete barcode sequence of the two discrete barcode sequences.Atty Dkt No.: 68244-703601
[0135] An example of a nucleic acid construct described herein is shown in FIG. 22B. The nucleic acid construct 2298 comprises a first nucleic acid strand 2201 comprising a first transposon end recognition sequence 2202 and a reverse complement 2203RC of a second transposon end recognition sequence 2203. The nucleic acid strand 2201 further comprises two discrete barcode sequences 2204 and 2205 located between the first transposon end recognition sequence 2202 and the reverse complement 2203RC of the second transposon end recognition sequence. The nucleic acid construct in FIG. 22B further comprises a second nucleic acid strand 2214 comprising a reverse complement 2202RC of the first transposon end recognition sequence, wherein a portion of the second nucleic acid strand is hybridized to a first segment of the first nucleic acid strand. The nucleic acid construct in FIG. 22B further comprises a third nucleic acid strand 2211 comprising the second transposon end recognition sequence 2203, wherein a portion of the third nucleic acid strand is hybridized to a second segment of the first nucleic acid strand.
[0136] In some embodiments, the first nucleic acid strand comprises a self-complementary, a self-hybridizing region, or a hairpin region. In some embodiments, the first nucleic acid strand comprises a self-complementary region. In some cases, the first segment of the first nucleic acid comprises a complementary region to the second segment of the first nucleic acid strand. An isolated first nucleic acid strand can comprise a hairpin that is opened up upon hybridization to the second nucleic acid strand and / or the third nucleic acid strand.
[0137] In an example construct shown in FIG. 9, the first nucleic acid strand 900 comprises segments (F and Frc) that are complementary to each other. When strand 900 is an isolated nucleic acid strand, it forms a hairpin structure, wherein segments F and Frc are hybridized to each other. The hairpin is opened up upon hybridization of the nucleic acid strand 900 to the second nucleic acid strand 906 and the third nucleic acid strand 904.
[0138] In some embodiments, the self-complementary region, the self-hybridizing region, or the hairpin region is located between the first segment and the second segment of the first nucleic acid strand. In some cases, the hairpin region remains a hairpin structure after hybridization of the first nucleic acid strand to the second nucleic acid strand and the third nucleic acid strands.
[0139] In an example construct shown in FIG. 12, the hairpin region of the first nucleic acid strand 1203 is located between the first segment 1201 and the second segment 1207.
[0140] In some embodiments, the first nucleic acid strand comprises a non-hybridizing region that is not hybridized to the second nucleic acid strand and that is not hybridized to the third nucleic acid strand. In some embodiments, the non-hybridizing region of the first nucleic acid strand can comprise a discrete barcode sequence of the two discrete barcode sequences. In other embodiments, the non-hybridizing region does not comprise a barcode sequence. In some embodiments, the non-hybridizing region comprises a functional sequence. The functionalAtty Dkt No.: 68244-703601 sequence can be a primer binding sequence or reverse complement thereof. In some cases, the nonhybridizing region comprises at least two, at least three, at least four, or at least five distinct primer binding sequence or reverse complement thereof.
[0141] In some embodiments, the non-hybridizing region or the self-hybridizing region is located between the first segment and the second segment of the first nucleic acid strand. In the example shown in FIG. 6, the first nucleic acid strand 600 comprises a non-hybridizing region 601 that comprises two functional sequences B and C. The non-hybridizing region 601 is located between the first segment 602 and the second segment 603 of the first nucleic acid strand.
[0142] In some embodiments, the non-hybridizing region is located between two hybridizing regions of the first nucleic acid strand, wherein the two hybridizing regions are each hybridized to the second nucleic acid strand.
[0143] In some embodiments, the first nucleic acid strand comprises two distinct nonhybridizing regions that are not hybridized to the second nucleic acid strand and that are not hybridized to the third nucleic acid strand. In some cases, a first non-hybridizing region of the two non-hybridizing regions of the first nucleic acid strand is located between the first segment and the second segment of the first nucleic acid strand; and a second non -hybridizing region of the two non-hybridizing regions of the first nucleic acid strand is located between two hybridizing regions of the first nucleic acid strand, wherein the two hybridizing regions are each hybridized to the second nucleic acid strand.
[0144] In an example shown in FIG. 7, the nucleic acid construct comprises a first nucleic acid strand 700 comprising two distinct non-hybridizing regions 701 and 703. The first nonhybridizing region 701 is located between the first segment 712 and the second segment 705 of the first nucleic acid strand. The second non-hybridizing region 703 is located between two hybridizing regions 704 and 702 of the first nucleic acid strand, wherein the two hybridizing regions are each hybridized to a second nucleic acid strand 707.
[0145] In some embodiments, the two distinct non-hybridizing regions of the first nucleic acid strand comprise a primer binding sequence or reverse complement thereof and a discrete barcode sequence of the two discrete barcode sequences. For example, non-hybridizing region 701 in FIG. 7 comprises a primer binding sequence or reverse complement thereof while non-hybridizing region 703 comprises a barcode sequence.
[0146] In some embodiments, the two distinct non-hybridizing regions of the first nucleic acid strand each comprises a primer binding sequence or reverse complement thereof. In an example shown in FIG. 8, non-hybridizing regions 801 and 803 each comprise a primer binding sequence or reverse complement thereof.Atty Dkt No.: 68244-703601
[0147] In some embodiments, the second nucleic acid strand or the third nucleic acid strand of the nucleic acid construct comprises a non-complementary region that is not hybridized to the first nucleic acid strand. In some embodiments, the second nucleic acid strand and the third nucleic acid strand of the nucleic acid construct each comprises a non-complementary region that is not hybridized to the first nucleic acid strand. For example, in FIG. 6, the second nucleic acid strand 607 comprises non-complementary region 606 and the third nucleic acid strand 605 comprises non-complementary region 604.
[0148] The non-complementary region of the second nucleic acid strand or the third nucleic acid strand can comprise a functional sequence. In some cases, the non-complementary region comprises an additional barcode sequence. In some cases, the non-complementary region comprises a binding sequence that is configured to hybridize or is hybridized to another nucleic acid molecule or a barcode construct (e.g., clonal web or DNA origami construct) described elsewhere herein. The binding sequence of the second nucleic acid strand or the third nucleic acid strand can be hybridized to a nucleic acid molecule comprising a discrete barcode sequence. In some cases, the binding sequence is hybridized to a nucleic acid molecule comprising a plurality of barcode sequences (e.g., a clonal web or DNA origami construct), as described elsewhere herein.
[0149] In some cases, the non-complementary region comprises a binding sequence that is hybridized to a splint oligonucleotide that couples the nucleic acid construct to another nucleic acid molecule or a barcode construct (e.g., clonal web or DNA origami construct) described elsewhere herein.
[0150] The non-complementary region of the second nucleic acid strand or the third nucleic acid strand can be on a 3’ end of its respective nucleic acid strand. This can enable 5’ to 3’ nucleic acid extension from the 3’ end, for example, using a template sequence provided by a barcode construct (e.g., clonal web or DNA origami construct) described elsewhere herein.
[0151] In the example shown in FIG. 6, the non-complementary region 606 of the second nucleic acid strand 607 comprises a binding sequence that hybridizes to a capture sequence 610 of a barcode construct 609. The barcode construct 609 can comprise one or more barcode sequences. In FIG. 6, the non-complementary region 604 of the third nucleic acid strand 605 comprises a primer binding sequence or reverse complement thereof.
[0152] In some embodiments, the second nucleic acid strand comprises two distinct non- complementary regions that are not hybridized to the first nucleic acid strand. In some cases, a first non-complementary region of the two non-complementary regions of the second nucleic acid strand is located on a 3’ end of the second nucleic acid strand, and a second non-complementary region of the two non-complementary regions of the second nucleic acid strand is located in anAtty Dkt No.: 68244-703601 internal segment of the second nucleic acid strand. For example, in FIG. 8, the second nucleic acid strand 810 comprises two distinct non-complementary regions 809 and 808. Non-complementary region 808 is located on a 3’ end of the second nucleic acid strand 810, while non-complementary region 809 is located in an internal segment of the second nucleic acid strand. In FIG. 8, the non- complementary region 809 comprise primer binding sequence or reverse complement thereof.
[0153] In some embodiments, the third nucleic acid strand further comprises a non- complementary region that is not hybridized to the first nucleic acid strand. The non- complementary region of the third nucleic acid strand can comprise a functional sequence (e.g., a binding sequence or a primer binding sequence or reverse complement thereof). The non- complementary region of the third nucleic acid strand can be on a 5’ end of the third nucleic acid strand, as shown, for example, in FIG. 8.
[0154] The two discrete barcode sequences of the first nucleic acid strand can be identical sequences. Alternatively, the two discrete barcode sequences of the first nucleic acid strand can comprise different sequences. In some embodiments, the first nucleic acid strand comprises at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7 discrete barcode sequences.
[0155] A functional sequence can be located between the two discrete barcode sequences in the first nucleic acid strand. A functional sequence can be located between the first transposon end recognition sequence and a discrete barcode sequence of the two discrete barcode sequences in the first nucleic acid strand. In some cases, the first nucleic acid strand comprises (i) a first functional sequence located between the first transposon end recognition sequence and a first discrete barcode sequence of the two discrete barcode sequences and (ii) a second functional sequence located between the second transposon end recognition sequence and a second discrete barcode of the two discrete barcode sequences. In some cases, the first nucleic acid strand comprises two distinct functional sequences located between the two discrete barcode sequences. The first nucleic acid strand can comprise two distinct primer binding sequence or reverse complement thereof located between the two discrete barcode sequences
[0156] In another aspect, the present disclosure provides a nucleic acid construct (e.g., an adapter construct) comprising: a first nucleic acid strand comprising a first transposon end recognition sequence, wherein the first nucleic acid strand comprises a first region and a second region. The nucleic acid construct can further comprise a second nucleic acid strand comprising second transposon end recognition sequence, wherein the second nucleic acid strand comprises a third region and a fourth region, wherein the first region of the first nucleic acid strand is hybridized to the third region of the second nucleic acid strand. The nucleic acid construct can further comprise a third nucleic acid strand comprising a reverse complement of the first transposon end recognition sequence, wherein a portion of the third nucleic acid strand isAtty Dkt No.: 68244-703601 hybridized to the second region of the first nucleic acid strand. The nucleic acid construct can further comprise a fourth nucleic acid strand comprising a reverse complement of the second transposon end recognition sequence, wherein a portion of the fourth nucleic acid strand is hybridized to the fourth region of the second nucleic acid strand.
[0157] In another aspect, the present disclosure provides a nucleic acid construct (e.g., an adapter construct) comprising: a first nucleic acid strand comprising a reverse complement of a first transposon end recognition sequence, wherein the first nucleic acid strand comprises a first region and a second region. The nucleic acid construct can further comprise a second nucleic acid strand comprising a reverse complement of a second transposon end recognition sequence, wherein the second nucleic acid strand comprises a third region and a fourth region, wherein the first region of the first nucleic acid strand is hybridized to the third region of the second nucleic acid strand. The nucleic acid construct can further comprise a third nucleic acid strand comprising the first transposon end recognition sequence, wherein a portion of the third nucleic acid strand is hybridized to the second region of the first nucleic acid strand. The nucleic acid construct can further comprise a fourth nucleic acid strand comprising the second transposon end recognition sequence, wherein a portion of the fourth nucleic acid strand is hybridized to the fourth region of the second nucleic acid strand.
[0158] In some aspects, the first nucleic acid strand or the second nucleic acid strand of the nucleic acid construct further comprises a barcode sequence. In some embodiments, the barcode sequence is located in the first region of the first nucleic acid strand. The third region of the second nucleic acid strand can comprise a reverse complement of the barcode sequence. In some embodiments, the first nucleic acid strand or the second nucleic acid strand further comprises a functional sequence (e.g., a primer binding sequence or reverse complement thereof). In some embodiments, the first nucleic acid strand or the second nucleic acid strand further comprises a single-stranded non-hybridizing region that is not hybridized to another nucleic acid strand. The single-stranded non-hybridizing region can be at a 5’ end or a 3’ end of the nucleic acid strand. In some cases, the single-stranded non-hybridizing region is located in an internal region of the nucleic acid strand. In some embodiments, the first nucleic acid strand further comprises a hairpin region.
[0159] In some aspects, the third nucleic acid strand further comprises a non-complementary region that is not hybridized to the first nucleic acid strand and is not hybridized to the second nucleic acid strand. In some embodiments, the fourth nucleic acid strand further comprises a non- complementary region that is not hybridized to the first nucleic acid strand and is not hybridized to the second nucleic acid strand. The non-complementary region of the third nucleic acid strand or the fourth nucleic acid strand can be at a 5’ end or a 3’ end of the respective nucleic acid strand.Atty Dkt No.: 68244-703601
[0160] FIG. 11 shows an example of a nucleic acid construct 11 described herein. The nucleic acid construct comprises a first nucleic acid strand 1100, comprising a reverse complement 113 IRC of a first transposon end recognition sequence 1131 and a first region 1101 and a second region 1102. The nucleic acid construct 11 further comprises a second nucleic acid strand 1120 comprising a reverse complement 1141RC of a second transposon end recognition sequence 1141, wherein the second nucleic acid strand comprises a third region 1101RC and a fourth region 1141RC, wherein the first region 1101 of the first nucleic acid strand is hybridized to the third region 110 IRC of the second nucleic acid strand. The nucleic acid construct 11 further comprises a third nucleic acid strand 1130 comprising the first transposon end recognition sequence 1131, wherein a portion 1102RC of the third nucleic acid strand is hybridized to the second region 1102 of the first nucleic acid strand. The third nucleic acid strand 1130 further comprises a non- complementary region that is not hybridized to the first nucleic acid strand or the second nucleic acid strand. The nucleic acid construct 11 further comprises a fourth nucleic acid strand 1140 comprising the second transposon end recognition sequence 1141, wherein a portion 1141 of the fourth nucleic acid strand is hybridized to the fourth region 1141RC of the second nucleic acid strand 1120. The fourth nucleic acid strand further comprises a non-complementary region that is not hybridized to the first nucleic acid strand or the second nucleic acid strand.
[0161] In some aspects, the third nucleic acid strand comprises a portion that is hybridized to the first nucleic acid strand and further comprises a portion that is hybridized to the second nucleic acid strand. In some embodiments, the fourth nucleic acid strand comprises a portion that is hybridized to the second nucleic acid strand and further comprises a portion that is hybridized to the first nucleic acid strand.
[0162] FIG. 10 shows an example of a nucleic acid construct described herein. The nucleic acid construct 10 comprises a first nucleic acid strand 1000, comprising a first transposon end recognition sequence 1002 and a first region 1001 and a second region 1002. The nucleic acid construct 10 further comprises a second nucleic acid strand 1010 comprising the second transposon end recognition sequence 1011, wherein the second nucleic acid strand comprises a third region 100 IRC and a fourth region 1011, wherein the first region 1001 of the first nucleic acid strand 1000 is hybridized to the third region 100 IRC of the second nucleic acid strand 1010. The nucleic acid construct 10 further comprises a third nucleic acid strand 1020 comprising a reverse complement 1002RC of the first transposon end recognition sequence 1002, wherein a portion 1002RC of the third nucleic acid strand is hybridized to the second region 1002 of the first nucleic acid strand 1000. The nucleic acid construct 10 further comprises a fourth nucleic acid strand 1030 comprising a reverse complement 101 IRC of the second transposon end recognition sequenceAtty Dkt No.: 68244-7036011011, wherein a portion 101 IRC of the fourth nucleic acid strand is hybridized to the fourth region 1011 of the second nucleic acid strand 1010.
[0163] In some aspects, the third nucleic acid strand or the fourth nucleic acid strand comprises a barcode sequence. In some embodiments, the third nucleic acid strand or the fourth nucleic acid strand comprises a functional sequence (e.g., a binding sequence or a primer binding sequence or reverse complement thereof). The third nucleic acid strand or the fourth nucleic acid strand can comprise a binding sequence that is configured to hybridize or is hybridized to another nucleic acid molecule or a barcode construct (e.g., clonal web or DNA origami construct) described elsewhere herein. The binding sequence of the third nucleic acid strand or the fourth nucleic acid strand can be hybridized to a nucleic acid molecule comprising a discrete barcode sequence. In some cases, the binding sequence is hybridized to a nucleic acid molecule comprising a plurality of barcode sequences (e.g., a clonal web or DNA origami construct), as described elsewhere herein.
[0164] In some aspects, the binding sequence is hybridized to a splint oligonucleotide that couples the nucleic acid construct to another nucleic acid molecule or a barcode construct (e.g., clonal web or DNA origami construct) described elsewhere herein.
[0165] The binding sequence of the third nucleic acid strand or the fourth nucleic acid strand can be on a 3’ end of its respective nucleic acid strand. This can enable 5’ to 3’ nucleic acid extension from the 3 ’ end, for example, using a template sequence provided by a barcode construct (e.g., clonal web or DNA origami construct) described elsewhere herein.
[0166] In some aspects, the third nucleic acid strand and the fourth nucleic acid strand both comprise binding sequences that are hybridized to a same nucleic acid molecule or a barcode construct (e.g., clonal web or DNA origami construct) described elsewhere herein. In other embodiments, the third nucleic acid strand and the fourth nucleic acid strand both comprise binding sequences that are hybridized different nucleic acid molecules or a barcode constructs (e.g., clonal web or DNA origami construct) described elsewhere herein.
[0167] In the example shown in FIG. 10, the third nucleic acid strand 1020 and the fourth nucleic acid strand comprise binding sequences that are hybridized to different regions 1051 and 1052 of a same barcode construct. In the example shown in FIG. 11, the second nucleic acid strand 1120 comprises a binding sequence that is hybridized to a barcode construct 1150. The third nucleic acid strand 1130 comprises a primer binding sequence or reverse complement thereof.
[0168] In some aspects, a nucleic acid strand of a nucleic acid construct described herein can comprise a functional sequence. A functional sequence can be a primer binding sequence or reverse complement thereof, an adapter sequence, a barcode sequence, or a distinguishing sequence. For example, a functional sequence in the first nucleic acid strand can distinguish theAtty Dkt No.: 68244-703601 first nucleic acid strand from a reverse complement of the first nucleic acid strand. A functional sequence can be at least 1, at least 2, at least 3, at least 4, at least 5, at most 2, at most 3, at most 4, at most 5, at most 10, or at most 15 nucleotides in length. In some cases, a nucleic acid strand of the nucleic acid construct comprises at least 2, at least 3, at least 4, at least 5, at least 6, or at least 7 functional sequences. In some cases, more than one nucleic acid strand of the nucleic acid construct comprises a functional sequence.
[0169] In some aspects, the structure of an adapter construct described herein can establish a barcode construct connection (e.g., type A connection), a cis transposase monomer connection (e.g., type B connection), or a trans transposase monomer connection (e.g., type C linkage). The adapter constructs shown in FIGs. 3-13A are examples of adapter constructs that can establish a barcode construct connection via hybridization to a barcode construct and are configured to copy a barcode sequence from the barcode construct. The adapter constructs shown in FIGs. 5-9, and FIG. 11-13A are examples of adapter constructs that can establish a cis transposase monomer connection. In these examples, a transposase monomer is bound to a double-stranded transposon end formed by two nucleic strands, wherein one nucleic acid strand comprises a barcode sequence and the other nucleic acid strand comprises the reverse complement of the barcode sequence. In these examples, barcode information is shared between the two strands that bind to the same transposase monomer. For example, in FIG. 5, strands 507 and 508 form a transposon end that is bound to a transposase monomer. Strand 507 comprises barcode sequence 506, and strand 508 comprises the reverse complement 506RC, and barcode information is shared between strand 507 and 508. One of the two strands that bind to a same transposase monomer can be generated by nucleic acid extension using the other of the two strands as a template for cDNA synthesis. For example, strand 507 can be generated by nucleic acid extension using a primer that binds to E,rc of strand 508 and nucleic acid extension using strand 508 as a template. This can enable barcode information to be copied between strands that bind to the same transposase monomer.
[0170] Examples of adapter constructs that can establish a trans transposase monomer connection include adapter constructs represented in FIGs. 6-13B. In these examples, the structure of adapter construct enables barcode information across two transposase monomers of a same transposase dimer is linked for sequence assembly of barcoded short reads. For example, in FIG. 10, strands 1000 and 1020 form a first transposon end that binds a first transposase monomer of a transposase dimer. Strands 1030 and 1010 form a second transposon end that binds to a second transposase monomer of the transposase dimer. Strand 1000 comprises a barcode sequence and strand 1010 comprises reverse complement of the barcode sequence. In this example, barcode information is shared across two transposase monomers of a same transposase dimer.Atty Dkt No.: 68244-703601
[0171] In some cases, for example, in FIG. 5, the structure of an adapter construct described herein can establish a combined barcode construct connection (e.g., type A connection) and cis transposase monomer connection (e.g., type B connection). In some cases, for example, in FIG. 10, the structure of an adapter construct described herein can establish a combined barcode construct connection (e.g., type A connection) and trans transposase monomer connection (e.g., type C connection). In some cases, for example, in for example, in FIGs. 9 and 12, the structure of an adapter construct described herein can establish a combined cis transposase monomer connection (e.g., type B connection) and trans transposase monomer connection (e.g., type C connection). In some cases, the structure of an adapter construct described herein can establish a barcode construct connection, cis transposase monomer connection, and trans transposase monomer connection. For example, in FIGs. 8 and 11, the structure of an adapter construct can establish a type C connection, a combined type B connection and barcode construct connection, and a combined type B and type C connection.
[0172] In some aspects, an adapter construct described herein is integrated into a target nucleic acid molecule. A derivative of an adapter construct can be an integrated adapter construct. An integrated adapter construct can be an insertion product of a transposase-mediated insertion reaction.
[0173] In some aspects, the present disclosure provides a composition that is a precursor to one or more adapter constructs described herein. In some aspects, the present disclosure provides a composition that is configured to be cleaved (e.g., via a transposase or via a restriction enzyme) to yield a pair of adapter constructs described herein. In one aspect, the present disclosure provides a composition comprising (i) a nucleic acid strand comprising two identical copies of a transposon end recognition sequence; (ii) a first oligonucleotide comprising a reverse complement of the transposon end recognition sequence; and (iii) a second oligonucleotide comprising a reverse complement of the transposon end recognition sequence.
[0174] Non-limiting examples of a transposon end recognition sequence include SEQ ID NOs: 1-6. The first transposon end recognition sequence and the second transposon end recognition sequence can each comprise any one of SEQ ID NOs: 1-6. In some cases, the first transposon end recognition sequence and the second transposon end recognition sequence each comprises SEQ ID NO: 1. In some cases, the first transposon end recognition sequence and the second transposon end recognition sequence each comprises SEQ ID NO: 2. In some cases, the first transposon end recognition sequence and the second transposon end recognition sequence each comprises SEQ ID NO: 3. In some cases, the first transposon end recognition sequence and the second transposon end recognition sequence each comprises SEQ ID NO: 4. In some cases, the first transposon end recognition sequence and the second transposon end recognition sequence each comprises SEQAtty Dkt No.: 68244-703601ID NO: 5. In some cases, the first transposon end recognition sequence and the second transposon end recognition sequence each comprises SEQ ID NO: 6.
[0175] In some cases, the nucleic acid strand comprises (I) a first segment comprising a first barcode sequence and a first copy of the two identical copies of the transposon end recognition sequence, and (II) a second segment comprising a second barcode sequence and a second copy of the two identical copies of the transposon end recognition sequence. The first barcode sequence and the second barcode sequence can be identical sequences. Alternatively, the first barcode sequence and the second barcode sequence can be different sequences. In some cases, the first barcode sequence and the second barcode sequence are not identical sequences and are also not reverse complements of each other. In some cases, the nucleic acid strand comprises, in order from 5’ to 3’ : the first barcode sequence, the first copy of the two identical copies of the transposon end recognition sequence, the second barcode sequence, and the second copy of the two identical copies of the transposon end recognition sequence.
[0176] In some cases, the nucleic acid strand comprises a transposase cut site (e.g., a Tn5 transposase cut site) located adjacent and 3’ to the first copy of the two identical copies of the transposon end recognition sequence. In some cases, the nucleic acid strand comprises a transposase cut site (e.g., a Tn5 transposase cut site) located adjacent and 3’ to the second copy of the two identical copies of the transposon end recognition sequence. In some cases, the nucleic acid strand does not comprise a restriction enzyme recognition sequence. In some cases, the nucleic acid strand does not comprise a PvuII restriction enzyme recognition sequence.
[0177] In some cases, the nucleic acid strand comprises a restriction enzyme recognition sequence. The restriction enzyme recognition sequence can be, for example, a PvuII restriction enzyme recognition sequence, e.g., CAGCTG. In some cases, the first oligonucleotide comprises a reverse complement of the restriction enzyme recognition sequence. In some cases, the nucleic acid strand comprises a restriction enzyme cut site, e.g., at the cut site indicated by * in CAG*CTG. In some cases, the nucleic acid strand comprises a restriction enzyme cut site located adjacent and 3’ to the first copy of the two identical copies of the transposon end recognition sequence.
[0178] The nucleic acid strand can further comprise an additional restriction enzyme recognition sequence (e.g., an additional PvuII restriction enzyme recognition sequence). In some cases, the second oligonucleotide comprises a reverse complement of the additional restriction enzyme recognition sequence. In some cases, the nucleic acid strand further comprises an additional restriction enzyme cut site located adjacent and 3’ to the second copy of the two identical copies of the transposon end recognition sequence.Atty Dkt No.: 68244-703601
[0179] An example of a composition described herein is shown in FIG. 27, which depicts a nucleic acid strand 2700 comprising two identical copies of a transposon end recognition sequence ME. The composition further comprises a first oligonucleotide 2701 comprising a reverse complement of the transposon end recognition sequence ME, rc and a second oligonucleotide 2702 comprising a reverse complement of the transposon end recognition sequence ME, rc. The nucleic acid strand 2700 comprises a first segment 2711 comprising a first barcode sequence UMI1-SI and a first copy of the two identical copies of the transposon end recognition sequence ME and a second segment 2712 comprising a second barcode sequence UMI2-SI and a second copy of the two identical copies of the transposon end recognition sequence ME. The nucleic acid strand 2700 comprises, in order from 5’ to 3’ : the first barcode sequence UMI1-SI, the first copy of the two identical copies of the transposon end recognition sequence ME, the second barcode sequence UMI2-SI, and the second copy of the two identical copies of the transposon end recognition sequence ME.
[0180] In FIG. 27, the nucleic acid strand 2700 comprises an enzymatic cut site (*) located adjacent and 3’ to the first copy of the two identical copies of the transposon end recognition sequence ME. The enzymatic cut site can be a transposase cut site and / or a restriction enzyme cut site. Additionally, the nucleic acid strand 2700 further comprises an enzymatic cut site (*) located adjacent and 3’ to the second copy of the two identical copies of the transposon end recognition sequence ME. The nucleic acid strand 2700 further comprises a first primer binding sequence Pl at its 5’ end and a second primer binding sequence P2 at its 3’ end.
[0181] In some cases, the composition further comprises a linker coupling the first segment of the nucleic acid strand to the second segment of the nucleic acid strand. In some cases, the composition further comprises one or more linker polynucleotides coupling the first segment of the nucleic acid strand to the second segment of the nucleic acid strand, wherein the one or more linker polynucleotides is distinct from the nucleic acid strand. The coupling can comprise covalent linkage. Alternatively, the coupling can be a noncovalent coupling, e.g., via hybridization.
[0182] In some cases, the composition comprises a unimolecular linker polynucleotide comprising (i) a first capture sequence hybridized to a first binding site in the first segment of the nucleic acid strand; and (ii) a second capture sequence hybridized to a second binding site in the second segment of the nucleic acid strand. An example of a composition comprising a unimolecular linker polynucleotide is 2850 depicted in FIG. 28. The composition 2850 comprises a nucleic acid strand 2800 comprising two identical copies of a transposon end recognition sequence ME. The composition 2850 further comprises a first oligonucleotide 2801 comprising a reverse complement of the transposon end recognition sequence ME, rc and a second oligonucleotide 2802 comprising a reverse complement of the transposon end recognitionAtty Dkt No.: 68244-703601 sequence ME, rc. The nucleic acid strand 2800 comprises a first segment 2811 comprising a first barcode sequence UMI1-SI and a first copy of the two identical copies of the transposon end recognition sequence ME and a second segment 2812 comprising a second barcode sequence UMI2-SI and a second copy of the two identical copies of the transposon end recognition sequence ME. The nucleic acid strand 2800 comprises, in order from 5’ to 3’ : the first barcode sequence UMI1-SI, the first copy of the two identical copies of the transposon end recognition sequence ME, the second barcode sequence UMI2-SI, and the second copy of the two identical copies of the transposon end recognition sequence ME. The nucleic acid strand 2800 further comprises a first primer binding sequence Pl at its 5’ end and a second primer binding sequence P2 at its 3’ end
[0183] The nucleic acid strand 2800 further comprises enzymatic cut site (*) located adjacent and 3’ to the two identical copies of the transposon end recognition sequence ME. The enzymatic cut site can be a transposase cut site and / or a restriction enzyme cut site.
[0184] In this example, the composition 2850 further comprises a unimolecular linker polynucleotide 2803 coupling the first segment 2811 of the nucleic acid strand 2800 to the second segment 2812 of the nucleic acid strand 2800. The linker polynucleotide 2803 is distinct from the nucleic acid strand 2800 and comprises a first capture sequence (X, rc and Pl, rc) hybridized to a first binding site (X and Pl) in the first segment 2811 of the nucleic acid strand 2800 and a second capture sequence tY, rc hybridized to a second binding site in Y in the second segment 2812 of the nucleic acid strand 2800.
[0185] In other cases, the composition comprises a plurality of linker polynucleotide molecules coupling the first segment of the nucleic acid strand to the second segment of the nucleic acid strand, wherein the plurality of linker polynucleotides comprises (i) a first linker polynucleotide molecule hybridized to a first binding site in the first segment of the nucleic acid strand; and (ii) a second linker polynucleotide molecule hybridized to a second binding site in the second segment of the nucleic acid strand. The plurality of linker polynucleotides can comprise at least 2, at least 3, at least 4, at least 5, or at least 6 linker polynucleotides.
[0186] An example of a composition comprising a plurality of linker polynucleotide molecules is 3050 shown in FIG. 30. The composition 3050 comprises a nucleic acid strand 3000 comprising two identical copies of a transposon end recognition sequence ME. The composition further comprises a first oligonucleotide 3001 comprising a reverse complement of the transposon end recognition sequence ME, rc and a second oligonucleotide 3002 comprising a reverse complement of the transposon end recognition sequence ME, rc. The nucleic acid strand 3000 comprises a first segment 3011 comprising a first barcode sequence UMI1-SI and a first copy of the two identical copies of the transposon end recognition sequence ME and a second segment 3012 comprising aAtty Dkt No.: 68244-703601 second barcode sequence UMI2-SI and a second copy of the two identical copies of the transposon end recognition sequence ME. The nucleic acid strand 3000 further comprises an enzymatic cut site (*) located adjacent and 3’ to the two identical copies of the transposon end recognition sequence ME. The enzymatic cut site can be a transposase cut site and / or a restriction enzyme cut site. In this example, the composition 3050 further comprises a plurality of linker polynucleotide molecules 3003 and 3004 coupling the first segment 3011 of the nucleic acid strand 3000 to the second segment 3012 of the nucleic acid strand 3000. The linker polynucleotide molecules 3003 and 3004 are distinct from the nucleic acid strand 3000 and comprise a first linker polynucleotide molecule 3003 hybridized to a first binding site X in the first segment 3011 of the nucleic acid strand 3000 and a second linker polynucleotide molecule 3004 hybridized to a second binding site in Y in the second segment 3012 of the nucleic acid strand 3000.
[0187] In some cases, the one or more linker polynucleotides comprises a linker polynucleotide that comprises the first oligonucleotide comprising the reverse complement of the transposon end recognition sequence. For example, the first oligonucleotide can be a segment of the linker polynucleotide. In some cases, the first oligonucleotide is covalently or noncovalently linked to a linker polynucleotide.
[0188] FIG. 29 depicts a unimolecular linker polynucleotide 2903 that comprises the first oligonucleotide 2901. FIG. 29 depicts a composition 2950 comprising a nucleic acid strand 2900 comprising two identical copies of a transposon end recognition sequence ME. The composition 2950 further comprises a first oligonucleotide 2901 comprising a reverse complement of the transposon end recognition sequence ME, rc and a second oligonucleotide 2902 comprising a reverse complement of the transposon end recognition sequence ME, rc. The nucleic acid strand 2900 comprises a first segment 2911 comprising a first barcode sequence UMI1-SI and a first copy of the two identical copies of the transposon end recognition sequence ME and a second segment 2912 comprising a second barcode sequence UMI2-SI and a second copy of the two identical copies of the transposon end recognition sequence ME. The nucleic acid strand 2900 further comprises an enzymatic cut site (*) located adjacent and 3’ to the two identical copies of the transposon end recognition sequence ME. The enzymatic cut site can be a transposase cut site and / or a restriction enzyme cut site. In this example, the composition 2950 further comprises a unimolecular linker polynucleotide 2903 coupling the first segment 2911 of the nucleic acid strand 2900 to the second segment 2912 of the nucleic acid strand 2900. The linker polynucleotide 2903 is distinct from the nucleic acid strand 2900 and comprises a first capture sequence (X, rc and Pl, rc) hybridized to a first binding site (X and Pl) in the first segment 2911 of the nucleic acid strand 2900 and a second capture sequence Y, rc hybridized to a second binding site Y in the secondAtty Dkt No.: 68244-703601 segment 2912 of the nucleic acid strand 2900. In this example, the linker polynucleotide 2903 comprises the first oligonucleotide 2901.
[0189] Another example of a linker polynucleotide that comprises the first oligonucleotide is shown in FIG. 31, which depicts a linker polynucleotide 3104 comprising the first oligonucleotide 3101. In this example, the composition 3150 comprises nucleic acid strand 3100 comprising two identical copies of a transposon end recognition sequence ME. The composition 3150 further comprises a first oligonucleotide 3101 comprising a reverse complement of the transposon end recognition sequence ME, rc and a second oligonucleotide 3102 comprising a reverse complement of the transposon end recognition sequence ME, rc. The nucleic acid strand 3100 comprises a first segment 3111 comprising a first barcode sequence UMI1-SI and a first copy of the two identical copies of the transposon end recognition sequence ME and a second segment 3112 comprising a second barcode sequence UMI2-SI and a second copy of the two identical copies of the transposon end recognition sequence ME. The nucleic acid strand 3100 further comprises an enzymatic cut site (*) located adjacent and 3’ to the two identical copies of the transposon end recognition sequence ME. The enzymatic cut site can be a transposase cut site and / or a restriction enzyme cut site. In this example, the composition further comprises a plurality of linker polynucleotide molecules 3103 and 3104 coupling the first segment 3111 of the nucleic acid strand 3100 to the second segment 3112 of the nucleic acid strand 3100. The linker polynucleotide molecules 3103 and 3104 are distinct from the nucleic acid strand 3100 and comprise a first linker polynucleotide molecule 3103 hybridized to a first binding site X in the first segment 3111 of the nucleic acid strand 3100 and a second linker polynucleotide molecule 3104 hybridized to a second binding site Y in the second segment 3112 of the nucleic acid strand 3100. In this example, the second linker polynucleotide 3104 comprises the first oligonucleotide 3101.
[0190] The nucleic acid strand in the composition can further comprise other structural and / or functional features. For example, the nucleic acid strand can further comprise one or more primer binding sequences. In some cases, the nucleic acid strand further comprises primer binding sequences at its 5’ end and its 3’ end (e.g., Pl and P2 in FIGs. 27-31). In some cases, the nucleic acid strand further comprises an internal primer binding sequence (e.g., A in FIGs. 27-31).
[0191] In some cases, the nucleic acid strand is further coupled to a support. The support can be, for example, a bead. The nucleic acid strand can be releasably coupled to the support, for example, via an enzymatically labile or photo-labile linker. In some cases, the nucleic acid strand is coupled to a biotin that can couple to an avidin on a support, for example, as shown in FIG.33B
[0192] In another aspect, the present disclosure provides composition comprising an adapter pair comprising: (i) a first adapter construct comprising a first double-stranded transposon endAtty Dkt No.: 68244-703601 comprising a first transposon end recognition sequence; (ii) a second adapter construct comprising a second double-stranded transposon end comprising a second transposon recognition sequence; wherein the first adapter construct and the second adapter construct are coupled to each other via one or more linker polynucleotides.
[0193] In some cases, the first transposon end recognition sequence and the second transposon end recognition sequence are identical sequences. In other cases, the first transposon end recognition sequence and the second transposon end recognition sequence are different sequences. Non-limiting examples of a transposon end recognition sequence include SEQ ID NOs: 1-6. The first transposon end recognition sequence and the second transposon end recognition sequence can each comprise any one of SEQ ID NOs: 1-6. In some cases, the first transposon end recognition sequence and the second transposon end recognition sequence each comprises SEQ ID NO: 1. In some cases, the first transposon end recognition sequence and the second transposon end recognition sequence each comprises SEQ ID NO: 2. In some cases, the first transposon end recognition sequence and the second transposon end recognition sequence each comprises SEQ ID NO: 3. In some cases, the first transposon end recognition sequence and the second transposon end recognition sequence each comprises SEQ ID NO: 4. In some cases, the first transposon end recognition sequence and the second transposon end recognition sequence each comprises SEQ ID NO: 5. In some cases, the first transposon end recognition sequence and the second transposon end recognition sequence each comprises SEQ ID NO: 6.
[0194] In some cases, the first adapter construct further comprises a first barcode sequence and wherein the second adapter construct further comprises a second barcode sequence. The first barcode sequence can be located 5’ of the first transposon end recognition sequence. The second barcode sequence can be located 5’ of the second transposon end recognition sequence. The first barcode sequence and the second barcode sequence can be identical sequences. Alternatively, the first barcode sequence and the second barcode sequence can be different sequences. In some cases, the first barcode sequence and the second barcode sequence are not identical sequences and are also not reverse complements of each other.
[0195] The first barcode sequence and the second barcode sequence can be associated with each other as a linked pair. In some cases, the composition is among one or more different adapter constructs that are not coupled to the first adapter construct or the second adapter construct. The one or more different adapter constructs can comprise one or more barcode sequences different from the first barcode sequence and the second barcode sequence. In some cases, the first barcode sequence and the second barcode sequence together identify the first adapter construct and the second adapter construct as associated with each other.Atty Dkt No.: 68244-703601
[0196] In some cases, the first adapter construct and the second adapter construct each further comprises one or more primer binding sequences. The primer binding sequence can be at a 5’ or a 3’ end. Alternatively, the primer binding sequence can be an internal primer binding sequence.
[0197] In some cases, the composition comprises a unimolecular linker polynucleotide comprising (i) a first capture sequence hybridized to a first binding site in the first adapter construct; and (ii) a second capture sequence hybridized to a second binding site in the second adapter construct.
[0198] Examples are shown in FIGs. 28 and 29. FIG. 28 depicts an adapter pair comprising a first adapter construct 2821 and a second adapter construct 2822. The first adapter construct 2821 comprises a first double-stranded transposon end comprising a first transposon end recognition sequence ME. The second adapter construct 2822 comprises a second double-stranded transposon end comprising a second transposon recognition sequence ME. The first adapter construct 2821 and the second adapter construct 2822 are coupled to each other via a unimolecular linker polynucleotide 2803. In this example, the unimolecular linker polynucleotide 2803 comprises a first capture sequence (X, rc and Pl, rc) hybridized to a first binding site (X and Pl) in the first adapter construct 2821 and a second capture sequence tY, rc hybridized to a second binding site in Y in the second adapter construct 2822. The first adapter construct 2821 further comprises a first barcode sequence UM1-SI, and the second adapter construct 2822 further comprises a second barcode sequence UM2-SI. The first barcode sequence UM1-SI is located 5’ of the first transposon end recognition sequence ME, and the second barcode sequence UM2-SI is located 5’ of the second transposon end recognition sequence ME. The first adapter construct 2821 further comprises primer binding sequences Pl and A. The second adapter construct 2822 further comprises a primer binding sequence A.
[0199] FIG. 29 depicts an adapter pair comprising a first adapter construct 2921 and a second adapter construct 2922. The first adapter construct 2921 comprises a first double-stranded transposon end comprising a first transposon end recognition sequence ME. The second adapter construct 2922 comprises a second double-stranded transposon end comprising a second transposon recognition sequence ME. The first adapter construct 2921 and the second adapter construct 2922 are coupled to each other via a unimolecular linker polynucleotide 2923.
[0200] In other cases, the composition comprises a plurality of linker polynucleotide molecules coupling the first segment of the nucleic acid strand to the second segment of the nucleic acid strand, wherein the plurality of linker polynucleotides comprises (i) a first linker polynucleotide molecule hybridized to a first binding site in the first adapter construct; and (ii) a second linker polynucleotide molecule hybridized to a second binding site in the second adapterAtty Dkt No.: 68244-703601 construct. The plurality of linker polynucleotides can comprise at least 2, at least 3, at least 4, at least 5, or at least 6 linker polynucleotides.
[0201] Examples are shown in FIGs. 30 and 31. FIG. 30 depicts an adapter pair comprising a first adapter construct 3021 and a second adapter construct 3022. The first adapter construct 3021 comprises a first double-stranded transposon end comprising a first transposon end recognition sequence ME. The second adapter construct 3022 comprises a second double-stranded transposon end comprising a second transposon recognition sequence ME. The first adapter construct 3021 and the second adapter construct 3022 are coupled to each other via a plurality of linker polynucleotide molecules 3003 and 3004. In this example, the two linker polynucleotide molecules 3003 and 3004 couple the first segment 3011 of the nucleic acid strand to the second segment 3012 of the nucleic acid strand. The first linker polynucleotide molecule 3003 is hybridized to a first binding site X in the first adapter construct 3021. The second linker polynucleotide molecule 3004 is hybridized to a second binding site in Y in the second adapter construct 3022. The first adapter construct 3021 further comprises a first barcode sequence UM1-SI, and the second adapter construct 3022 further comprises a second barcode sequence UM2-SI. The first barcode sequence UM1-SI is located 5’ of the first transposon end recognition sequence ME, and the second barcode sequence UM2-SI is located 5’ of the second transposon end recognition sequence ME. The first adapter construct 3021 further comprises primer binding sequences Pl and A. The second adapter construct 3022 further comprises a primer binding sequence A.
[0202] FIG. 31 depicts an adapter pair comprising a first adapter construct 3121 and a second adapter construct 3122. The first adapter construct 3121 comprises a first double-stranded transposon end comprising a first transposon end recognition sequence ME. The second adapter construct 3122 comprises a second double-stranded transposon end comprising a second transposon recognition sequence ME. The first adapter construct 3121 and the second adapter construct 3122 are coupled to each other via a plurality of linker polynucleotide molecules 3103 and 3123.
[0203] In some aspects, a composition described herein further comprises a transposase. The transposase can be a Tn5 transposase. In some cases, the composition further comprises a restriction enzyme. The restriction enzyme can be, for example, PvuII.
[0204] In some cases, a composition described herein or a derivative thereof is bound to a transposase. For example, the composition can comprise a complex comprising the nucleic acid strand, the first oligonucleotide, and the second oligonucleotide described herein, and a transposase. As another example, the composition can comprise a complex comprising the first adapter construct and the second adapter construct described herein, and a transposase. The transposase can be a transposase dimer. In some cases, the complex further comprises one or moreAtty Dkt No.: 68244-703601 linker polynucleotides described herein. In other cases, the complex does not comprise one or more linker polynucleotides coupling the first adapter construct and the second adapter construct.B. Systems and kits
[0205] In some aspects, the present disclosure provides a system or kit comprising multiple nucleic acid strands described elsewhere herein. The system or kit can comprise a library of nucleic acid strands, wherein each nucleic acid strand comprises two identical copies of a transposon end recognition sequence and a unique pair of barcode sequences. In some cases, each nucleic acid strand comprises, in order from 5’ to 3’ : a first barcode sequence of the unique pair of barcode sequences, a first copy of the two identical copies of the transposon end recognition sequence, a second barcode sequence of the unique pair of barcode sequences, and a second copy of the two identical copies of the transposon end recognition sequence. In some cases, the library comprises one or more nucleic acid strands comprising a restriction enzyme recognition sequence. A nucleic acid strands can comprise a restriction enzyme cut site located adjacent and 3’ to the first copy of the two identical copies of the transposon end recognition sequence. The restriction enzyme recognition sequence can be a PvuII restriction enzyme recognition sequence. In some cases, the nucleic acid strand further comprises an additional restriction enzyme recognition sequence. The nucleic acid strand can further comprise an additional restriction enzyme cut site located adjacent and 3’ to the second copy of the two identical copies of the transposon end recognition sequence. The additional restriction enzyme recognition sequence can be a PvuII restriction enzyme recognition sequence.
[0206] The system or kit can comprise at least 2, at least 3, at least 4, at least 5, at least 6, at least 8, at least 10, at least 15, at least 20, at least 25, at least 30, at least 40, at least 50, at least 100, at least 200, at least 500, at least 1000, at least 2000, at least 5000, at least 10,000, at least 50,000, at least 100,000, at least 500,000, at least 1 million, at least 5 million, at least 10 million, at least 25 million, at least 50 million, at least 100 million, at least 500 million, at least 1 billion, at least 10 billion, or at least 100 billion unique nucleic acid strands each comprising two identical copies of a transposon end recognition sequence and a unique pair of barcode sequences.
[0207] The system or kit can further comprise multiple oligonucleotides comprising a reverse complement of the transposon end recognition sequence. The system or kit can further comprise one or more linker polynucleotides described elsewhere herein, e.g., one or more linker polynucleotides coupling the first segment of the nucleic acid strand to the second segment of the nucleic acid strand, wherein the one or more linker polynucleotides is distinct from the nucleic acid strand.
[0208] In some aspects, the present disclosure provides a system comprising multiple adapter pairs. The system can comprise one or more adapter pairs described elsewhere herein. In someAtty Dkt No.: 68244-703601 cases, the system comprises a first adapter pair and a second adapter pair, wherein the second adapter pair is different from the first adapter pair. The first adapter pair can comprise a first adapter construct and a second adapter construct, wherein the first adapter construct and the second adapter construct are coupled to each other via one or more linker polynucleotides. The second adapter pair can comprises a third adapter construct and a fourth adapter construct, wherein the third adapter construct and the fourth adapter construct are coupled to each other via one or more additional linker polynucleotides. In some cases, the first adapter construct comprises a first barcode sequence and the second adapter construct comprises a second barcode sequence, wherein the first barcode sequence and the second barcode sequence together identify the first adapter construct and the second adapter construct as associated with each other. In some cases, the third adapter construct comprises a third barcode sequence, wherein the fourth adapter construct comprises a fourth barcode sequence, wherein the third barcode sequence and the fourth barcode sequence together identify the third adapter construct and the fourth adapter construct as associated with each other.
[0209] The system can comprise at least 2, at least 3, at least 4, at least 5, at least 6, at least 8, at least 10, at least 15, at least 20, at least 25, at least 30, at least 40, or at least 50, at least 100, at least 200, at least 500, at least 1000, at least 2000, at least 5000, at least 10,000, at least 50,000, at least 100,000, at least 500,000, at least 1 million, at least 5 million, at least 10 million, at least 25 million, at least 50 million, at least 100 million, at least 500 million, at least 1 billion, at least 10 billion, or at least 100 billion unique pairs of adapter constructs having unique pairs of barcode sequences.
[0210] In some aspects, a system or kit described herein further comprises one or more enzymes. The system or kit can further comprise a transposase (e.g., Tn5). The system or kit can further comprise a restriction enzyme (e.g., PvuII).C. Barcode constructs
[0211] In some aspects, the present disclosure provides a composition comprising a barcode construct. The barcode construct can be a nucleic acid barcode construct. The barcode construct can comprise DNA or RNA. The barcode construct can comprise single-stranded nucleic acid or double-stranded nucleic acid. The barcode construct can be a linear nucleic acid construct or a circular nucleic acid construct. The barcode construct can comprise one or more barcode sequences (e.g., a UMI). In some embodiments, the barcode construct comprises a plurality of barcode sequences. In some cases, the plurality of barcode sequences comprises a same barcode sequence. In some cases, all barcode sequences of the plurality of barcode sequences are identical. In some cases, the plurality of barcode sequences comprise different barcode sequences. The barcodeAtty Dkt No.: 68244-703601 construct can comprise at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, or at least 250 barcode sequences.
[0212] The barcode construct can comprise one or more capture sequences configured to hybridize to an additional nucleic acid molecule. The one or more capture sequences in the barcode construct can be configured to enable hybridization and barcoding of an additional nucleic acid molecule via a nucleic acid extension reaction. In some cases, the additional nucleic acid molecule is a polynucleotide of an adapter construct or a derivative thereof described elsewhere herein. In some cases, the additional nucleic acid molecule is a product of transposase-mediated insertion of an adapter construct described elsewhere herein into a target nucleic acid. In some cases, a capture sequence is on a 3 ’ side of a barcode sequence in the barcode construct. This can enable generation of an extension product that comprises the reverse complement of the barcode sequence. In some embodiments, a barcode construct is configured to barcode a target nucleic acid via hybridization to an adapter construct described elsewhere herein used for transposase mediated insertion into the target nucleic acid. The barcoding can occur via a barcode construct connection (e.g., type A connection), for example as shown in FIGs. 15A and 15B. An adapter construct or a derivative thereof described elsewhere herein can comprise a binding sequence that hybridizes to a capture sequence of a barcode construct. The segment of the adapter construct that is hybridized to the capture sequence can be extended in a nucleic acid extension reaction to generate an extension product comprising a reverse complement of a barcode sequence in the barcode construct.
[0213] In some embodiments, the barcode construct comprises a plurality of capture sequences. In some cases, the plurality of capture sequences comprises a same capture sequence. In some cases, all capture sequences of the plurality of capture sequences are identical. In some cases, the plurality of capture sequences comprise different sequences. The barcode construct can comprise at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 150, at least 200, or at least 250 capture sequences. A capture sequence can be adjacent to a barcode sequence in the barcode construct. In some cases, a capture sequence is located between two barcode sequences in the barcode construct. In some cases, a barcode sequence is located between two capture sequences in the barcode construct. In some embodiments, the barcode construct is configured to enable hybridization of multiple additional nucleic acid molecules to the barcode construct. The barcode construct can enable barcoding of the multiple additional nucleic acid molecules via nucleic acid extension reactions. In some cases, the barcoding of the multiple additional nucleic acid molecules occur simultaneously. In other cases, the barcoding of the multiple additional nucleic acid molecules occur sequentially.Atty Dkt No.: 68244-703601
[0214] In some embodiments, the barcode construct is configured to stop a nucleic acid extension reaction at a given site. In some embodiments, the barcode construct is configured to stop a nucleic acid extension reaction at a given site via a non-canonical nucleotide that is non- amplifiable by a polymerase. In other embodiments, the barcode construct is configured to stop a nucleic acid extension reaction at a given site via a recognition sequence in the barcode construct that is configured to bind to a DNA-binding protein and prevent nucleic acid extension beyond the recognition sequence. In some cases, the recognition sequence is bound to the DNA-binding protein, which prevents nucleic acid extension beyond the recognition sequence. As an example, the recognition sequence can be a complementary sequence to a guide RNA sequence of a guide RNA complexed with a CRISPR-Cas domain. In some cases, the recognition sequence is bound to a guide RNA complexed with a CRISPR-Cas domain. The CRISPR-Cas domain comprise an inactive nuclease domain. In some cases, the CRISPR-Cas domain comprises a Cas9 domain.
[0215] In some cases, the barcode construct comprises a plurality of barcode sequences, wherein barcode sequences of the plurality of barcode sequences are separated by a non-canonical nucleotide. In other cases, the barcode construct comprises a plurality of barcode sequences, wherein barcode sequences of the plurality of barcode sequences are separated by a recognition sequence configured to prevent nucleic acid extension beyond the recognition sequence. In some cases, the barcode construct comprises a plurality of capture sequences, wherein capture sequences of the plurality of capture sequences are separated by a non-canonical nucleotide. In other cases, the barcode construct comprises a plurality of capture sequences, wherein capture sequences of the plurality of capture sequences are separated by a recognition sequence configured to prevent nucleic acid extension beyond the recognition sequence.
[0216] The barcode construct can be at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 700, at least 800, at least 900, at least 1000, at least 1500, at least 2000, at least 2500 nucleotides, at least 3000 nucleotides, at least 3000 nucleotides, at least 4000 nucleotides, at least 5000 nucleotides, at least 6000 nucleotides, at least 7000 nucleotides, at least 8000 nucleotides, at least 9000 nucleotides, or at least 10,000 nucleotides in length.
[0217] In some embodiments, the barcode construct comprises a concatemer or a continuous nucleic acid molecule that contains multiple copies of a same sequence linked in series. The concatemer can comprise at least 5, at least 8, at least 10, at least 15, at least 20, at least 30, at least 40, at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 120, at least 140, or at least 160 repeat sequences. The concatemer can comprise multiple copies of a same plurality of barcode sequences (e.g., a clonal web or DNA origami construct). For example, a barcode construct can be a nucleic acid amplification product, such as a rolling circle amplification (RCA)Atty Dkt No.: 68244-703601 product. FIG. 25 provides an example of a circular barcode construct 2501 comprising five different barcode sequences and a linear barcode construct 2502 that is a rolling circle amplification product of 2501. FIG. 26 provides another example of a linear barcode construct 2502. The linear barcode construct can be circularized to generate a circular barcode construct. An example of a circular barcode construct is 2603 in FIG. 26. Barcode construct 2604 is another example of a barcode construct that is a rolling circle amplification product of circular construct 2603.
[0218] In some embodiments, the barcode construct comprises a scaffolded structure. The barcode construct can comprise a DNA origami structure. In some cases, the barcode construct comprises a plurality of short DNA oligonucleotides to fold a single-stranded scaffold DNA strand.
[0219] In some embodiments, a barcode sequence in the barcode construct identifies the barcode construct or provides information about the barcode construct from which the one or more barcode sequences were derived. In some embodiments, a barcode construct is used to generate one or more barcoded nucleic acid molecules (e.g., via hybridization and nucleic acid extension). The one or more barcoded nucleic acid molecules can comprise one or more barcode sequences originating from the barcode construct from which the one or more barcode sequences were copied. A barcode sequence of a barcoded nucleic acid molecule can identify the barcoded nucleic acid molecule as having been generated from the barcode construct or resulting from a barcode construct connection (e.g., a type A connection).
[0220] In some embodiments, a barcode sequence in the barcode construct comprises a sample barcode sequence identifying a sample from which a target nucleic acid is derived. A target nucleic acid from a given sample can be barcoded by the barcode construct comprising the sample barcode to generate barcoded nucleic acid molecule comprising a target nucleic acid sequence from the target nucleic acid and the sample barcode sequence. In some embodiments, the barcode construct comprises both a UMI and a sample barcode sequence.D. Adapter construct coupled to a barcode construct
[0221] In some aspects, the present disclosure provides a composition comprising one or more adapter constructs or derivatives thereof described elsewhere herein coupled to a barcode construct described elsewhere herein. The coupling of the adapter construct or derivative thereof and barcode construct can enable generation of a barcoded nucleic acid molecule comprising a barcode sequence of the barcode construct or a reverse complement thereof (e.g., via nucleic acid extension). The composition can comprise an adapter construct coupled to a barcode construct prior to transposase-mediated insertion. The composition can enable barcoding of the adapter construct to generate a barcoded adapter construct comprising a reverse complement of the barcode construct. Alternatively, the composition can comprise a transposase-mediated insertion productAtty Dkt No.: 68244-703601 of the adapter construct and a target nucleic acid, wherein the transposase-mediated insertion product is coupled to barcode construct. The composition can enable barcoding of the transposase- mediated insertion product to generate a barcoded insertion product comprising (i) a target nucleic acid sequence of the target nucleic acid or a reverse complement thereof, and (ii) a barcode sequence of the barcode construct or a reverse complement thereof.
[0222] The barcode construct can comprise single stranded or double-stranded nucleic. The barcode construct can be a circular or linear nucleic acid construct. In some embodiments, the barcode construct comprises a barcode sequence as described elsewhere herein. The barcode construct can comprise a plurality of barcode sequences. In some cases, the barcode construct comprises a capture sequence as described elsewhere herein. The barcode construct can comprise a plurality of capture sequences. An adapter construct can be coupled to a capture sequence of the barcode construct. In some embodiments, a plurality of adapter constructs are coupled to a single barcode construct. In some embodiments, a single adapter construct is coupled to a plurality of barcode constructs.
[0223] An adapter construct coupled to a barcode construct can comprise a plurality of polynucleotides comprising one or more transposon end recognition sequences. The plurality of polynucleotides can comprise a pair of transposon end recognition sequences. As described elsewhere herein, the plurality of polynucleotides of an adapter construct can comprise a polynucleotide comprising a binding sequence that is hybridized to a capture sequence of the barcode construct. The capture sequence can be on a 3’ side of a barcode sequence in the barcode construct, enabling generation of a barcoded nucleic acid molecule comprising a reverse complement of the barcode sequence via nucleic acid extension. In some cases, the capture sequence is adjacent to a barcode sequence.
[0224] In some embodiments, the adapter construct comprises a plurality of polynucleotides each comprising a binding sequence that is hybridized to a portion of the barcode construct. For example, the adapter construct can comprise two polynucleotides that each comprise a binding sequence that is hybridized to a separate capture sequence of the barcode construct. An example is shown in FIG. 10, where the non-complementary regions of the nucleic acid strands 1020 and 1030 comprise binding sequences that are hybridized to capture sequences 1041 and 1052 of a barcode construct 1050 comprising one or more barcode sequences. Each separate capture sequence that is hybridized to a polynucleotide of the adapter construct can be on a 3’ side of a barcode sequence of the barcode construct. In some cases, each separate capture sequence is adjacent to a barcode sequence.
[0225] In some embodiments, the adapter construct coupled to the barcode construct comprises four nucleic acid strands. The adapter construct can comprise a first nucleic acid strandAtty Dkt No.: 68244-703601 comprising a first transposon end recognition sequence. The adapter construct can further comprise a second nucleic acid strand comprising a reverse complement of the first transposon end recognition sequence. The second nucleic acid strand can further comprise a first binding sequence that is hybridized to a first capture sequence of the plurality of capture sequences of the barcode construct. The adapter construct can further comprise a third nucleic acid strand comprising a second transposon end recognition sequence. The adapter construct can further comprise a fourth nucleic acid strand comprising a reverse complement of the second transposon end recognition sequence. The fourth nucleic acid strand can further comprise a second binding sequence that is hybridized to a second capture sequence of the plurality of capture sequences of the nucleic acid molecule. In some cases, the first nucleic acid strand or the third nucleic acid strand further comprises an adapter sequence. In some cases, the second nucleic acid strand or the fourth nucleic acid strand further comprises a unique molecular index. In some cases, the second nucleic acid strand and the fourth nucleic acid strand each comprises a unique molecular index.
[0226] In some embodiments, the adapter construct coupled to the barcode construct comprises three nucleic acid strands. The adapter construct can comprise a first nucleic acid strand comprising (A) a first transposon end recognition sequence, (B) a reverse complement of a second transposon end recognition sequence. The adapter construct can further comprise a second nucleic acid strand comprising a reverse complement of the first transposon end recognition sequence, wherein a portion of the second nucleic acid strand is hybridized to a first segment of the first nucleic acid strand. The adapter construct can further comprise a third nucleic acid strand comprising the second transposon end recognition sequence, wherein a portion of the third nucleic acid strand is hybridized to a second segment of the first nucleic acid strand. An example is shown in FIG. 6. The second nucleic acid strand or the third nucleic acid strand can further comprise a binding sequence that is hybridized to a capture sequence of the plurality of capture sequences of the barcode construct. In some cases, the first nucleic acid strand further comprises a unique molecular index located between the first transposon recognition sequence and the second transposon recognition sequence. In some cases, the first nucleic acid strand further comprises two discrete unique molecular indexes located between the first transposon recognition sequence and the reverse complement of the second transposon recognition sequence.
[0227] In some embodiments, the adapter construct coupled to the barcode construct comprises four nucleic acid strands comprising a first nucleic acid strand comprising a first transposon end recognition sequence, wherein the first nucleic acid strand comprises a first region and a second region. The adapter construct can further comprise a second nucleic acid strand comprising second transposon end recognition sequence, wherein the fourth nucleic acid strand comprises a third region and a fourth region. In some cases, the first region of the first nucleic acidAtty Dkt No.: 68244-703601 strand is hybridized to the third region of the second nucleic acid strand. The adapter construct can further comprise a third nucleic acid strand comprising a reverse complement of the first transposon end recognition sequence, wherein a portion of the third nucleic acid strand is hybridized to the second region of the first nucleic acid strand. The adapter construct can further comprise a fourth nucleic acid strand comprising a reverse complement of the second transposon end recognition sequence, wherein a portion of the fourth nucleic acid strand is hybridized to the fourth region of the second nucleic acid strand. The third nucleic acid strand or the fourth nucleic acid strand can further comprise a binding sequence that is hybridized to a capture sequence of the plurality of capture sequences of the barcode construct. In some cases, the first nucleic acid strand or the second nucleic acid strand further comprises a unique molecular index.E. Transpososomes and transpososomes coupled to a barcode construct
[0228] In some aspects, the present disclosure provides a transpososome comprising a transposase and an adapter construct or a barcoded derivative of an adapter construct described elsewhere herein. The transposase can be an active transposase. In some cases, the transposase is a dimer comprising two transposase monomers. The transposase can be, for example, a Tn5 transposase. The Tn5 transposase can be a modified or engineered Tn5 transposase. In some cases, the transposase is a hyperactive transposase. An example of a hyperactive mutant is a Tn5 protein comprising three missense mutations E54K, M56A, L372P.
[0229] In some cases, the transpososome comprises a transposase bound to one or more transposon end recognition sequences. A transposon end recognition sequence can be an outside end sequence, an inside end sequence, or a mosaic end sequence. In some cases, the transpososome comprises a transposase bound to a plurality of polynucleotides comprising a pair of transposon end recognition sequences. The transpososome can comprise a transposase bound to a pair of transposon ends, each comprising a transposon end recognition sequence. The transpososome can comprise an active transposase bound to an adapter construct or a barcoded derivative of an adapter construct described elsewhere herein and can be capable of carrying out transposase-mediated insertion of the adapter construct or the barcoded derivative thereof into a target nucleic acid molecule.
[0230] In some aspects, the transpososome comprises a transposase bound to an adapter construct described herein. In some aspects, the transpososome comprises a transposase bound to a barcoded derivative of an adapter construct described herein. For example, a barcoded derivative of an adapter construct can be generated via hybridization of the adapter construct to a barcode construct described elsewhere herein and nucleic acid extension and complementary DNA (cDNA) synthesis using the barcode construct as a template. The barcoded derivative of the adapter construct can comprise a barcode sequence of the barcode construct or a reverse complementAtty Dkt No.: 68244-703601 thereof. In some aspects, the transpososome comprises a transposase bound to an adapter construct that is integrated into a target nucleic acid. The transpososome can comprise a transposase bound to an insertion product of transposase-mediated insertion of an adapter construct described herein into a target nucleic acid.
[0231] In some aspects, the present disclosure provides a composition comprising an adapter construct or derivative thereof coupled to barcode construct, as described elsewhere herein, wherein the composition further comprises a transposase. The transposase can be bound to the adapter construct or derivative thereof in a transpososome complex. In some cases, the transpososome comprises a transposase bound to an adapter construct described elsewhere herein, wherein the adapter construct is coupled to the barcode construct. In other cases, the transpososome comprises a transposase bound to a transposase-mediated insertion product of an adapter construct described elsewhere herein, wherein the transposase-mediated insertion product is coupled to the barcode construct. The adapter construct or the transposase-mediated insertion product of the adapter construct can comprise a binding sequence that is hybridized to a capture sequence of the barcode construct.
[0232] The coupling of the transposase, the adapter construct, and the barcode construct can enable generation of a barcoded transposase-mediated insertion product comprising (i) a barcode sequence of the barcode construct or a reverse complement thereof (e.g., via nucleic acid extension and cDNA synthesis) and (ii) a target sequence of the target nucleic acid.
[0233] In one aspect, the present disclosure provides a composition comprising a nucleic acid molecule (e.g., a barcode construct), comprising a plurality of barcode sequences and a plurality of capture sequences. The composition can further comprise a plurality of transpososomes coupled to a nucleic acid molecule, wherein each transpososome comprises an active transposase and a plurality of polynucleotides, wherein the plurality of polynucleotides comprises: (A) a pair of transposon end recognition sequences that are bound to the active transposase. The plurality of polynucleotides can further comprise (B) a polynucleotide comprising a binding sequence that is hybridized to a capture sequence of the plurality of capture sequences of the nucleic acid molecule.
[0234] In some embodiments, each transpososome of the plurality of transpososomes comprises two polynucleotides, each comprising a binding sequence that is hybridized to a separate capture sequence of the plurality of capture sequences of the nucleic acid molecule. Each separate capture sequence that is hybridized to a polynucleotide of the two polynucleotides can be on a 3’ side of a barcode sequence of the plurality of barcode sequences. Each separate capture sequence that is hybridized to a polynucleotide of the two polynucleotides can be adjacent to a barcode sequence of the plurality of barcode sequences.Atty Dkt No.: 68244-703601METHODS
[0235] In some aspects, the present disclosure provides a method comprising (a) generating a plurality of transposition products, wherein the plurality of transposition products comprises a plurality of unique pairs of barcode sequences, wherein a unique pair of barcode sequences of the plurality of unique pairs of barcode sequences identifies a first transposition product and a second transposition product of the plurality of transposition products as derived from sequences adjacent to each other within the target nucleic acid molecule; (b) generating a plurality of sequence reads from the plurality of transposition products; and (c) generating an assembled sequence from the plurality of sequence reads using the plurality of unique pairs of barcode sequences.
[0236] In an aspect, the present disclosure provides a method comprising (a) generating a plurality of transposition products via transposase-mediated fragmentation of a target nucleic acid molecule, wherein the plurality of transposition products comprises a plurality of unique pairs of barcode sequences, wherein a unique pair of barcode sequences of the plurality of unique pairs of barcode sequences identifies a first transposition product and a second transposition product of the plurality of transposition products as derived from sequences adjacent to each other within the target nucleic acid molecule; (b) generating a plurality of sequence reads from the plurality of transposition products; and (c) generating an assembled sequence from the plurality of sequence reads using the plurality of unique pairs of barcode sequences. As described herein, “transposase- mediated fragmentation” refers to fragmenting a nucleic acid molecule into separate nucleic acid fragments by one or more transposases. For example, the separate nucleic acid fragments may be separate molecules in the post-transposition product mixture.
[0237] In some aspects, the method does not comprise further fragmentation of the plurality of transposition products by a non-transposase enzyme (e.g., a restriction enzyme). In some aspects, the method does not comprise further ligating together transposition products of the plurality of transposition products or derivatives thereof.
[0238] The transposase-mediated fragmentation can comprise performing a transposition reaction between the target nucleic acid molecule and a transposon composition or derivative thereof, wherein the transposon composition comprises two double-stranded transposon ends but is not fully double-stranded. In some cases, between the segments containing the transposon recognition sequences, the transposon composition comprises a middle segment that is singlestranded. In some cases, the transposon composition does not self-tagment and / or does not get transposed into another copy of the transposon composition. In some cases, the transposon composition does not get transposed into another copy of the transposon composition, because it is not fully double-stranded. In some cases, the transposon composition does not get inserted into a target sequence via a transposition reaction if the target sequence is not double-stranded. ForAtty Dkt No.: 68244-703601 example, the transposon composition may not get inserted into a single-stranded region of another transposon composition.
[0239] In some cases, the transposon composition is a precursor of one or more adapter compositions described elsewhere herein. In some cases, the transposon composition is a precursor of a pair of adapter compositions described elsewhere herein. In some cases, the transposon composition comprises (I) a first adapter construct comprising a first double-stranded transposon end comprising a first transposon end recognition sequence and a first barcode sequence of a unique pair of barcode sequences of the plurality of unique barcode sequences; and (II) a second adapter construct comprising a second double-stranded transposon end comprising a second transposon recognition sequence and a second barcode sequence of the unique pair of barcode sequences. In some cases, the first adapter construct and the second adapter construct are assembled together with a same transposase dimer. In some cases, the first adapter construct and the second adapter construct are coupled to each other via one or more linker polynucleotides.
[0240] In some cases, the method comprises generating the transposon composition from a precursor composition comprising (i) a nucleic acid strand comprising two identical copies of a transposon end recognition sequence; (ii) a first oligonucleotide comprising a reverse complement of the transposon end recognition sequence; and(iii) a second oligonucleotide comprising a reverse complement of the transposon end recognition sequence. In some cases, each unique pair of barcodes sequences of the plurality of unique pairs of barcode sequences does not comprise complementary sequences
[0241] In another aspect, the present disclosure provides a method, comprising generating an assembled sequence from a plurality of sequence reads using a plurality of unique pairs of barcode sequences, wherein each unique pair of barcode sequences of the plurality of unique pairs of barcode sequences identifies two sequence reads of the plurality of sequence reads as being paired, wherein each unique pair of barcodes sequences does not comprise identical or complementary barcode sequences. The plurality of unique pairs of barcode sequences can comprises at least 1 million unique pairs of barcode sequences.
[0242] In some aspects, the method does not comprise sequence assembly based on an overlap sequence of the target nucleic acid molecule that is common to transposition products of the plurality of transposition products. In some aspects, the method does not comprise reference-based assembly. In some aspects, the method does not comprise sequence alignment of sequence reads of the plurality of sequence reads.A. Generating an adapter construct
[0243] In some aspects, the present disclosure provides a method of generating an adapter construct described elsewhere herein or generating a polynucleotide of an adapter constructAtty Dkt No.: 68244-703601 described elsewhere herein. The method can comprise a nucleic acid extension reaction. In some cases, a nucleic acid extension reaction is conducted using a primer hybridized to a template nucleic acid strand comprising a transposon end recognition sequence or a reverse complement thereof to generate a double-stranded transposon end. The primer can comprise a 3’ end that hybridizes to template nucleic acid strand and a non-hybridizing region that does not hybridize to the template nucleic acid strand. The non-hybridizing region of the primer can comprise a binding sequence for an additional primer or nucleic acid molecule. The template nucleic acid strand can further comprise a functional sequence. In some cases, the template nucleic acid strand further comprises one or more discrete barcode sequences.
[0244] The method can comprise a ligation reaction. In some cases, a ligation reaction ligates two nucleic acid oligonucleotides to generate a polynucleotide of the adapter construct. In some cases, the method comprises both a nucleic acid extension reaction and a ligation reaction to generate a polynucleotide of the adapter construct. The method can comprise providing a primer and an oligonucleotide hybridized to a nucleic acid strand comprising a transposon end recognition sequence or reverse complement thereof. The nucleic acid strand can further comprise a functional sequence. In some cases, the nucleic acid strand further comprises one or more discrete barcode sequences. In some cases, the primer hybridizes to the transposon end recognition sequence or reverse complement thereof. In some cases, the oligonucleotide is hybridized to another region of the nucleic acid strand. The oligonucleotide can comprise a non-hybridizing region that is not hybridized to the nucleic acid strand. The non-hybridizing region can comprise a binding sequence for an additional primer or nucleic acid molecule (e.g., a barcode construct described elsewhere herein).
[0245] FIGs. 22A and 22B show an example of generating an adapter construct 2298 using nucleic acid strands 2201 and oligonucleotides 2208, 2209, and 2210. As shown in FIG. 22A, nucleic acid strand 2201 comprises a first transposon end recognition sequence 2202, a reverse complement 2203RC of a second transposon end recognition sequence. In FIG. 22A, oligonucleotide 2208 comprises a reverse complement 2202RC of the transposon end recognition sequence 2202 and hybridizes to the nucleic acid strand 2201. Oligonucleotides 2210 and 2209 hybridize to two other regions 2206 and 2207 that are located between the two discrete barcode sequences 2204 and 2205 of nucleic acid strand 2201. Extension / ligation reactions are carried out to generate the adapter construct 2298. Oligonucleotide 2209 is extended to generate an extended strand 2211. Oligonucleotide 2208 is extended and linked to oligonucleotide 2210 in an extensionligation reaction to generate strand 2214.Atty Dkt No.: 68244-703601
[0246] FIGs. 23A and 24A show other examples of generating adapter constructs described herein. In FIG. 23 A, strand 2307 is extended and ligated to strand 2308. In FIG. 24A, strand 2404 is extended and ligated to strand 2403.
[0247] In an aspect, the present disclosure provides a method comprising (a) providing a composition comprising: (i) a nucleic acid strand comprising a first transposon end recognition sequence and a second transposon end recognition sequence; (ii) a first oligonucleotide that is hybridized to the first transposon end recognition sequence; and (iii) a second oligonucleotide that is hybridized to the second transposon end recognition sequence, wherein the first oligonucleotide and the second oligonucleotide are separate oligonucleotides; and (b) from the composition, generating: (I) a first adapter construct comprising a first double-stranded transposon end comprising the first transposon end recognition sequence; and (II) a second adapter construct comprising a second double-stranded transposon end comprising the second transposon end recognition sequence.
[0248] In some aspects, the nucleic acid strand comprises (I) a first segment comprising a first barcode sequence and the first transposon end recognition sequence, and (II) a second segment comprising a second barcode sequence and the second transposon end recognition sequence. The nucleic acid strand can comprise, in order from 5 ’ to 3 ’ : the first barcode sequence, the transposon end recognition sequence, the second barcode sequence, and the second copy of the two identical copies of the transposon end recognition sequence.
[0249] The composition can be a composition (e.g., an adapter precursor composition) described elsewhere herein. In some cases, the composition comprises a nucleic acid strand comprising a first barcode sequence and a second barcode sequence, and method generates (I) a first adapter construct comprising (A) a first double-stranded transposon end comprising the first copy of the transposon end recognition sequence, and (B) the first barcode sequence; and (II) a second adapter construct comprising (A) a second double-stranded transposon end comprising the second copy of the transposon end recognition sequence, and (B) the second barcode sequence.
[0250] In some aspects, the method comprises generating a transpososome comprising (I) the first adapter construct, (II) the second adapter construct, and (III) a transposase dimer. In some aspects, generating the first adapter construct and the second adapter construct comprises cleaving the nucleic acid strand at a first site adjacent to and 3’ of the first transposon end recognition sequence. Generating the first adapter construct and the second adapter construct can further comprise cleaving the nucleic acid strand at a second site adjacent to and 3’ of the second transposon end recognition sequence. The cleaving at the first site and / or the second site can be performed by the transposase dimer (e.g., after transposase synapse formation).Atty Dkt No.: 68244-703601
[0251] An example method of generating a first adapter construct and a second adapter construct via cleaving by a transposase dimer upon transposase synapse formation is shown in FIG. 34. The method shown in FIG. 34 provides the advantage of keeping originally paired barcodes together in the same transpososome complex even after cutting the original nucleic acid strand comprising the paired barcodes and separating the barcodes into two separate adapter constructs. This preserves the pairing of the barcodes even after the nucleic acid strand comprising the paired barcodes is cut. FIG. 34 provides a composition comprises a nucleic acid strand 3400 comprising, in order from 5’ to 3’ : the first barcode sequence 24N, the transposon end recognition sequence ME, the second barcode sequence 24N, and the second copy of the two identical copies of the transposon end recognition sequence ME. The composition in FIG. 34 further comprises a first oligonucleotide 3401 that is hybridized to the first transposon end recognition sequence ME and a second oligonucleotide 3402 that is hybridized to the second transposon end recognition sequence ME. The first oligonucleotide 3401 and the second oligonucleotide 3402 are separate oligonucleotides. FIG. 34 shows the generation of a first adapter construct 3431 and a second adapter construct 3432 from the composition via cutting at the cut sites (*) by a transposase. The composition assembles with a Tn5 transposase dimer, and upon transposase synapse formation, the transposase cleaves the nucleic acid strand 3400 at a first site adjacent to and 3’ of the first transposon end recognition sequence ME and at a second site adjacent to and 3’ of the second transposon end recognition sequence ME. The cutting yields the first adapter construct 3431 and the second adapter construct 3432, which are assembled in a transpososome with the transposase dimer. As shown in FIG. 34, the first adapter construct 3431 comprises a first double- stranded transposon end comprising the first transposon end recognition sequence ME and the first barcode sequence. The second adapter construct 3432 comprises a second double-stranded transposon end comprising the second transposon end recognition sequence ME and the second barcode sequence. The first barcode sequence and the second barcode sequence are in separate adapters in the same transpososome complex, thus preserving the pairing of the barcodes and allowing any resulting sequencing reads of transposition products tagged by the adapters to be linked together via the barcode sequences.
[0252] Alternatively, the cleaving at the first site and / or the second site can be performed by a restriction enzyme prior to generating the transpososome or prior to transposase synapse formation. The restriction enzyme can be, for example, PvuII.
[0253] An example method of generating a first adapter construct and a second adapter construct via cleaving by a restriction enzyme is shown in FIGs. 35A and 35B. The method shown in FIG. 35 also provides the advantage of keeping originally paired barcodes together in the same transpososome complex even after cutting the original nucleic acid strand comprising the pairedAtty Dkt No.: 68244-703601 barcodes and separating the barcodes into two separate adapter constructs. This preserves the pairing of the barcodes even after the nucleic acid strand comprising the paired barcodes is cut. FIG. 35A provides a composition comprises a nucleic acid strand 3500 comprising, in order from 5’ to 3’ : the first barcode sequence 24N, the transposon end recognition sequence ME, the second barcode sequence 24N, and the second copy of the two identical copies of the transposon end recognition sequence ME. The composition in FIGs. 35A and 35B further comprises a first oligonucleotide 3501 that is hybridized to the first transposon end recognition sequence ME and a second oligonucleotide 3502 that is hybridized to the second transposon end recognition sequence ME. The first oligonucleotide 3501 and the second oligonucleotide 3502 are separate oligonucleotides. In FIG. 35A, the first oligonucleotide 3501 is in a linker polynucleotide 3503 coupling a first segment 3511 and a second segment 3512 of the nucleic acid strand. FIGs. 35A and 35B show the generation of a first adapter construct 3531 and a second adapter construct 3532 from the composition via cutting at the cut sites (*) by a restriction enzyme. In this example, the restriction enzyme PvuII performs the cleaving at a first site adjacent to and 3’ of the first transposon end recognition sequence ME and at a second site adjacent to and 3’ of the second transposon end recognition sequence ME prior to generating the transpososome. The cutting yields the first adapter construct 3531 and the second adapter construct 3532, which are coupled to each other via 3503c, a derivative of the linker polynucleotide 3503. As shown in FIG. 35, the first adapter construct 3531 comprises a first double-stranded transposon end comprising the first transposon end recognition sequence ME and the first barcode sequence. The second adapter construct 3532 comprises a second double-stranded transposon end comprising the second transposon end recognition sequence ME and the second barcode sequence.
[0254] The first adapter construct 3531 and the second adapter construct 3532 then assemble with a Tn5 transposase dimer to form a transpososome, followed by removal of the 3503c linker polynucleotide. The first barcode sequence and the second barcode sequence are in separate adapters in the same transpososome complex, thus preserving the pairing of the barcodes and allowing any resulting sequencing reads of transposition products tagged by the adapters to be linked together via the barcode sequences.
[0255] In some aspects, the first adapter construct and the second adapter construct that are generated via cleaving at the first site and / or the second site are coupled to each other via one or more linker polynucleotides. In some aspects, in the composition pre-cleaving (e.g., the adapter precursor composition), the first segment of the nucleic acid strand and the second segment of the nucleic acid strand are coupled to each other via the one or more linker polynucleotides. In some cases, the one or more linker polynucleotides or one or more derivatives thereof remain bound toAtty Dkt No.: 68244-703601 the first segment and the second segment post-cleaving, thereby coupling the first adapter construct and the second adapter construct that are generated after cleaving.
[0256] In some aspects, the method further comprises (c) removing the one or more linker polynucleotides or derivatives thereof from the first adapter construct and the second adapter construct, for example, as shown in FIG. 35B. The removing can occur after generating a transpososome comprising (I) the first adapter construct, (II) the second adapter construct, and (III) a transposase dimer. In some cases, the removing occurs prior to performing a transposition reaction using the first adapter construct, the second adapter construct, and the target nucleic acid molecule, as described elsewhere herein.
[0257] In some aspects, prior to or during (a), the method further comprises preparing the composition provided in (a). For example, the method can comprise generating the nucleic acid strand in the composition via nucleic acid amplification. The method can further comprise hybridizing the first oligonucleotide and the second oligonucleotide to the nucleic acid strand to yield a hybridized complex. In some cases, prior to hybridizing to the first oligonucleotide and the second oligonucleotide, the nucleic acid strand is coupled to a support. The nucleic acid strand can be releasably coupled to the support, for example, via a labile linker or via a noncovalent interaction (e.g., biotin-streptavidin binding). In some cases, one or more other nucleic acid strands are washed away from the nucleic acid strand coupled to the support prior to hybridizing the nucleic acid strand to the first oligonucleotide and the second oligonucleotide to yield the hybridized complex. In some cases, the hybridized complex is released from the support. The hybridized complex can used to generate the first adapter construct and the second adapter construct in (b). An example is shown in FIG. 33 A, which depicts PCR amplification of a nucleic acid strand comprising, in order of 5’ to 3’ : the first barcode sequence 24N, a first copy of the transposon end recognition sequence ME, the second barcode sequence 24N, and a second copy of the transposon end recognition sequence ME using a biotinylated primer. In FIG. 33 A, the PCR amplified library is split for into NGS characterization and transposon preparation (e.g., the preparation of the adapter precursor composition). The double-stranded biotinylated nucleic acid molecule shown in FIG. 33B is coupled to a magnetic bead, followed by washing away the bottom strand. The single-stranded nucleic strand coupled to the magnetic bead in FIG. 33B can then hybridized to oligonucleotides comprising the reverse complement of the transposon recognition sequence, for example, as shown in FIG. 35A, or to yield the starting composition shown in FIG. 34
[0258] In some aspects, the method further comprises characterizing one or more portions of the composition provided in (a). The method can comprise identifying a pair of barcodes in the composition provided in (a). For example, the method can comprise sequencing the nucleic acidAtty Dkt No.: 68244-703601 strand or a copy or derivative thereof (e.g., via NGS characterization as shown in FIG. 33 A) to obtain pairing information identifying the first barcode sequence as paired with the second barcode sequence. The sequencing can be performed using a primer comprising a locked nucleic acid (LNA) nucleotide.
[0259] In some aspects, the method can comprise generating at least 2, at least 3, at least 4, at least 5, at least 6, at least 8, at least 10, at least 15, at least 20, at least 25, at least 30, at least 40, at least 50, at least 100, at least 200, at least 500, at least 1000, at least 2000, at least 5000, at least 10,000, at least 50,000, at least 100,000, at least 500,000, at least 1 million, at least 5 million, at least 10 million, at least 25 million, at least 50 million, at least 100 million, at least 500 million, at least 1 billion, at least 10 billion, or at least 100 billion unique nucleic acid strands each comprising two identical copies of a transposon end recognition sequence and a unique pair of barcode sequences. The barcode sequences can be different random sequences and can be randomly generated. The random barcode sequences can be determined by sequencing the nucleic acid strands.B. Generating a barcode construct
[0260] In some aspects, the method further comprises generating a barcode construct described elsewhere herein. The barcode construct can be a linear nucleic acid construct or a circular nucleic acid construct that comprises one or more barcode sequences (e.g., a UMI). In some cases, the barcode construct is single-stranded. In other cases, the barcode construct is double-stranded. Generating a circular barcode construct can comprise splint-ligation, for example, as shown in FIG. 26. In FIG. 26, barcode construct 2602 is hybridized to a splint molecule 2601 and circularized via ligation, for example by a ligase, to generate circular barcode construct 2603.
[0261] The barcode construct can further comprise one or more capture sequences configured to hybridize to a binding sequence of another nucleic acid (e.g., an adapter construct). In some cases, the barcode construct comprises a plurality of barcode sequences and a plurality of capture sequences. Generating the barcode construct can comprise performing a nucleic acid amplification reaction on a template nucleic acid molecule comprising a barcode sequence of the plurality of barcode sequences and a capture sequence of the plurality of capture sequences. The nucleic acid amplification reaction can comprise rolling circle amplification. FIG. 26 provides an example of generating a barcode construct 2604 via rolling circle amplification of circular construct 2603.C. Barcoding an adapter construct using a barcode construct
[0262] In some aspects, the present disclosure provides a method of barcoding an adapter construct or a derivative thereof as described elsewhere herein using a barcode construct. The method can comprise using a barcode construct to barcode an adapter construct to generate aAtty Dkt No.: 68244-703601 barcoded adapter construct. Alternatively, the method can comprise using a barcode construct to barcode a transposase-mediated insertion product of the adapter construct.
[0263] The method can comprise providing a composition comprising an adapter construct or derivative thereof coupled to a barcode construct, as described elsewhere herein. The composition can be generated by hybridizing a binding sequence of any one of the adapter constructs described elsewhere herein or a derivative thereof to a capture sequence of a barcode construct.
[0264] The method can comprise providing an adapter construct coupled to a barcode construct prior to transposase-mediated insertion of the adapter construct into a target nucleic acid. A binding sequence of a strand of the adapter construct can be hybridized to a capture sequence of the barcode construct. The binding sequence can be on a 3’ end of the strand. The capture sequence can be on a 3’ side of a barcode sequence in the barcode construct. The method can further comprise performing a nucleic acid extension reaction and complementary DNA synthesis using the barcode construct as a template and the strand of the adapter construct comprising the binding sequence as a primer. Extension of the strand of the adapter construct can generate a barcoded adapter construct comprising a reverse complement of the barcode sequence of the barcode construct.
[0265] The method can comprise providing a plurality of adapter constructs that couple to a barcode construct and generating a plurality of barcoded nucleic adapter constructs. The plurality of barcoded nucleic acid adapter constructs can be generated sequentially. Alternatively, the plurality of barcoded nucleic acid adapter constructs can be generated simultaneously. The barcode construct can comprise a plurality of capture sequences and a plurality of barcode sequences. In some cases, the plurality of adapter constructs simultaneously hybridize to a plurality of capture sequences in the barcode construct and initiate nucleic acid extension reactions to copy barcode sequences of the barcode construct and generate the plurality of barcoded nucleic acid adapter constructs.
[0266] Alternatively, the method can comprises providing an insertion product of the adapter construct that is coupled to a barcode construct after transposase-mediated insertion of the adapter construct into a target nucleic acid. The insertion product can comprise a target sequence of the target nucleic acid. A binding sequence of a strand of the insertion product can be hybridized to a capture sequence of the barcode construct. The binding sequence can be on a 3’ end of the strand. The capture sequence can be on a 3’ side of a barcode sequence in the barcode construct. The method can further comprise performing a nucleic acid extension reaction and complementary DNA (cDNA) synthesis using the barcode construct as a template and the strand of the insertion product comprising the binding sequence as a primer. Nucleic acid extension and cDNA synthesis can generate a barcoded product comprising a reverse complement of the barcode sequence of theAtty Dkt No.: 68244-703601 barcode construct and the target sequence of the target nucleic acid or a reverse complement thereof.
[0267] The method can comprise generating a plurality of barcoded products from the plurality of insertion products. The plurality of barcoded products can be generated sequentially. Alternatively, the plurality of barcoded products can be generated simultaneously. The barcode construct can comprise a plurality of capture sequences and a plurality of barcode sequences. In some cases, the plurality of insertion products simultaneously hybridize to a plurality of capture sequences in the barcode construct and initiate nucleic acid extension reactions to copy barcode sequences of the barcode construct and generate the plurality of barcoded products.
[0268] The nucleic acid extension reaction that generates the barcoded adapter construct or the barcoded product can be halted a given site by a non-canonical nucleotide or a DNA-binding protein bound to a specific recognition sequence. In some embodiments, the barcode construct comprises a non-canonical nucleotide that is non-amplifiable by a polymerase, and the non- canonical nucleotide stops the nucleic acid extension reaction. In some embodiments, the method comprises providing a DNA-binding protein (e.g., a CRISPR-Cas domain) that binds to a recognition sequence in the barcode construct and prevents nucleic acid extension beyond the recognition sequence.
[0269] The barcode construct can comprise a plurality of barcode sequences, wherein barcode sequences of the plurality of barcode sequences are separated by a non-canonical nucleotide. In other cases, the barcode construct comprises a plurality of barcode sequences, wherein barcode sequences of the plurality of barcode sequences are separated by a recognition sequence configured to prevent nucleic acid extension beyond the recognition sequence. In some cases, the barcode construct comprises a plurality of capture sequences, wherein capture sequences of the plurality of capture sequences are separated by a non-canonical nucleotide. In other cases, the barcode construct comprises a plurality of capture sequences, wherein capture sequences of the plurality of capture sequences are separated by a recognition sequence configured to prevent nucleic acid extension beyond the recognition sequence.
[0270] In some aspects, during barcoding of the adapter construct or derivative thereof, the adapter construct or derivative thereof is further bound to a transposase in a transpososome complex.
[0271] In one aspect, the present disclosure provides a method comprising: (a) providing: (i) a nucleic acid molecule, comprising a plurality of barcode sequences and a plurality of capture sequences; and (ii) a plurality of transpososomes coupled to the nucleic acid molecule, wherein each transpososome comprises an active transposase and a plurality of polynucleotides. The nucleic acid molecule can be a barcode construct described elsewhere herein. In some cases, theAtty Dkt No.: 68244-703601 plurality of polynucleotides is an adapter construct described elsewhere herein. In some cases, the plurality of polynucleotides a derivative of an adapter construct described elsewhere herein. For example, the plurality of polynucleotides can be part of a transposase-mediated insertion product of an adapter construct described elsewhere herein. The plurality of polynucleotides can comprise(A) a pair of transposon end recognition sequences that are bound to the active transposase, and(B) a polynucleotide comprising a binding sequence that is hybridized to a capture sequence of the plurality of capture sequences of the nucleic acid molecule (e.g., barcode construct).
[0272] The method can further comprise (b) performing a nucleic acid extension reaction to extend a polynucleotide of a transpososome of the plurality of transpososomes that is hybridized to a capture sequence of the nucleic acid molecule to yield an extended polynucleotide comprising (I) the binding sequence and (II) a reverse complement of a barcode sequence of the plurality of barcode sequences. In some cases, (b) occurs prior to transposase-mediated insertion of an adapter construct comprising the polynucleotide, and (b) results in the generation of a barcoded adapter construct. In some cases, (b) occurs after transposase-mediated insertion of an adapter construct and the polynucleotide is a part of an insertion product that is bound to the transposase, and (b) results in the generation of a barcoded insertion product.D. Transposase-mediated insertion of adapter construct and generation of barcoded insertion products
[0273] In some aspects, the present disclosure provides a method of inserting an adapter construct or barcoded adapter construct described elsewhere herein into a target nucleic acid using a transposase. The target nucleic acid can be double- stranded DNA. In some cases, the target nucleic acid is circular nucleic acid. In some cases, the target nucleic acid is linear nucleic acid. The target nucleic acid can comprise genomic DNA. The target nucleic acid can be at least 1000, at least 1500, at least 2000, at least 2500, at least 3000, at least 3500, at least 4000, at least 5000, at least 6000, at least 7000, at least 8000, or at least 9000 bp in length. The target nucleic acid can be at least 10 kb, 15 kb, 20 kb, 25 kb, 30 kb, 35 kb, 40 kb, 50 kb, 60 kb, 70 kb, 80 kb, 90 kb, 100 kb, 1 Mb, 1.5 Mb, or 2 Mb in length.
[0274] The transposase can be a dimer comprising two transposase monomers. In some cases, the transposase is a Tn5 transposase or a modified or engineered variant thereof. In some cases, each transposase monomer of a transposase dimer binds to a transposon end comprising a transposon end recognition sequence. The transposon end can comprise double-stranded DNA. The transposase can assemble with an adapter construct described elsewhere herein to form a transpososome.
[0275] The transposase can insert the adapter construct into the target nucleic acid. Insertion of the adapter construct can comprise cleaving the target nucleic acid and ligating a strand of theAtty Dkt No.: 68244-703601 adapter construct to a fragment of the target nucleic acid, thereby generating an insertion product comprising the adapter construct and a target sequence of the target nucleic acid.
[0276] In some aspects, the method comprises (a) providing an adapter construct described elsewhere herein, an active transposase, and a target nucleic acid molecule comprising a target sequence; and (b) integrating the nucleic acid into a target nucleic acid molecule using an active transposase. The active transposase can integrate the adapter construct into the target nucleic acid molecule in a transposase-mediated insertion reaction.
[0277] In some aspects, the present disclosure provides a method of generating a barcoded nucleic acid molecule using transposase-mediated insertion of an adapter construct.
[0278] In some aspects, the method comprises (a) providing an adapter construct described elsewhere herein, an active transposase, and a target nucleic acid molecule comprising a target sequence; and (b) generating a barcoded nucleic acid molecule using the nucleic acid, the active transposase, and the target nucleic acid molecule. In some embodiments, the adapter construct provided in (a) comprises two discrete barcode sequences, and the barcoded nucleic acid molecule comprises (i) a discrete barcode sequence of the two discrete barcode sequences, and (ii) the target sequence. In some embodiments, (a) further comprises providing a barcode construct described elsewhere herein and (b) comprises using the barcode construct to generate the barcoded nucleic acid molecule. In some cases, (b) comprise hybridizing a binding sequence of the adapter construct to a capture sequence of the barcode construct and performing a nucleic acid extension reaction to copy a reverse complement of the barcode sequence. In some cases, (b) comprise hybridizing a binding sequence of an insertion product of the adapter construct to a capture sequence of the barcode construct and performing a nucleic acid extension reaction to copy a reverse complement of the barcode sequence.
[0279] In one aspect, the method comprises (a) providing: (i) a first nucleic acid, comprising a plurality of barcode sequences and a plurality of capture sequences; (ii) a plurality of transpososomes coupled to the first nucleic acid via the plurality of capture sequences, wherein each transpososome comprises an active transposase and a plurality of polynucleotides, wherein the plurality of polynucleotides comprises a pair of transposon end recognition sequences that are bound to the active transposase; and (iii) a second nucleic acid. The first nucleic acid can be a barcode construct described elsewhere herein. The second nucleic acid can be a target nucleic acid. The second nucleic acid can be double-stranded. In some cases, the second nucleic acid comprises genomic DNA. In some cases, the second nucleic acid is in a cell. The plurality of polynucleotides can be an adapter construct described elsewhere herein.
[0280] The method can further comprise (b) generating a barcoded nucleic acid molecule using the first nucleic acid, the second nucleic acid, the active transposase, and a polynucleotide of theAtty Dkt No.: 68244-703601 plurality of polynucleotides; wherein the barcoded nucleic acid molecule comprises (I) a barcode sequence of the plurality of barcode sequences of the first nucleic acid or a reverse complement thereof, and (II) a nucleic acid sequence (e.g., a target sequence) of the second nucleic acid or a reverse complement thereof.
[0281] In some embodiments, (b) comprises (i) performing a nucleic acid extension reaction on the polynucleotide or a derivative thereof using a segment of the first nucleic acid as a template. In some embodiments, (b) further comprises (ii) performing a transposase-mediated insertion reaction using the active transposase to integrate the polynucleotide or a derivative thereof into the second nucleic acid.
[0282] In some embodiments, the nucleic acid extension reaction in (b)(i) and the transposase- mediated insertion reaction in (b)(ii) occur simultaneously. The nucleic acid extension reaction in (b)(i) and the transposase-mediated insertion reaction in (b)(ii) can occur in a one-pot reaction. In other embodiments, the nucleic acid extension reaction in (b)(i) and the transposase-mediated insertion reaction in (b)(ii) occur sequentially. For example, the nucleic acid extension reaction in (b)(i) can occur prior to the transposase-mediated insertion reaction in (b)(ii). Alternatively, the nucleic acid extension reaction in (b)(i) can occur after the transposase-mediated insertion reaction in (b)(ii).
[0283] In some embodiments, in (b), the active transposase integrates a barcoded derivative of the polynucleotide into the second nucleic acid in a transposase-mediated insertion reaction to generate the barcoded nucleic acid molecule, wherein the barcoded derivative comprises the barcode sequence or a reverse complement thereof. The plurality of polynucleotides can comprise a polynucleotide comprising a binding sequence that is hybridized to a capture sequence of the plurality of capture sequences. In some cases, the binding sequence of the polynucleotide is on a 3’ end of the polynucleotide, and (b) comprises performing a nucleic acid extension reaction using the polynucleotide of the plurality of polynucleotides as a primer and a segment of the first nucleic acid as a template to generate the barcoded derivative of the polynucleotide. The active transposase can integrate the barcoded derivative of the polynucleotide into the second nucleic acid into a random location in the second nucleic acid. In some embodiments, (b) further comprises covalently linking an end of the polynucleotide or a barcoded derivative of the polynucleotide and the nucleic acid sequence of the second nucleic acid or a reverse complement thereof in a ligation reaction or a nucleic acid gap fill-in reaction to generate single barcoded nucleic acid strand comprising the barcode sequence or a reverse complement thereof and the nucleic acid sequence (e.g., target sequence) of the second nucleic acid or a reverse complement thereof.
[0284] In some embodiments, in (b), the active transposase integrates the polynucleotide into the second nucleic acid in a transposase-mediated insertion reaction to generate an insertionAtty Dkt No.: 68244-703601 product that is further barcoded to generate the barcoded nucleic acid molecule. The insertion product can comprise a polynucleotide comprising a binding sequence that is hybridized to a capture sequence of the plurality of capture sequences. In some cases, the binding sequence of the polynucleotide is on a 3’ end of the polynucleotide, and (b) comprises performing a nucleic acid extension reaction using the polynucleotide as a primer and a segment of the first nucleic acid as a template to generate the barcoded nucleic acid molecule. In some embodiments, (b) further comprises covalently linking an end of the polynucleotide comprising the barcode sequence or reverse complement thereof and the nucleic acid sequence of the second nucleic acid or a reverse complement thereof in a ligation reaction or a nucleic acid gap fill-in reaction to generate single barcoded nucleic acid strand comprising the barcode sequence or a reverse complement thereof and the nucleic acid sequence (e.g., target sequence) of the second nucleic acid or a reverse complement thereof.
[0285] In some embodiments, (b) further comprises incorporating a dideoxynucleotide (ddNTP) at a 3’ end of an extension product of polynucleotide or a derivative thereof.
[0286] In some embodiments, the barcoded nucleic acid molecule that is generated after (b) comprises a length of from 100 to 600 bp.
[0287] In some aspects, the method comprises performing a transposition reaction using a first adapter construct, a second adapter construct, and a target nucleic acid molecule. The first adapter construct and the second adapter constructs can be a pair of adapter constructs described elsewhere herein. As described elsewhere herein, the first adapter construct and the second adapter construct can be generated from a composition comprising: (i) a nucleic acid strand comprising a first transposon end recognition sequence, a first barcode sequence, a second transposon end recognition sequence, and a second barcode sequence; (ii) a first oligonucleotide that is hybridized to the first transposon end recognition sequence; and (iii) a second oligonucleotide that is hybridized to the second transposon end recognition sequence, wherein the first oligonucleotide and the second oligonucleotide are separate oligonucleotides. The first adapter construct can comprise a first double-stranded transposon end comprising the first transposon end recognition sequence and the first barcode sequence; and the second adapter construct can comprise a second double-stranded transposon end comprising the second transposon end recognition sequence and the second barcode sequence.
[0288] As described elsewhere herein, the manner in which the first adapter construct and the second adapter construct are generated from the composition can keep the first adapter construct and the second adapter construct in a same transpososome complex, allowing transposition products tagged by the first adapter construct and the second adapter construct to be linked together. For example, the first adapter construct and the second adapter construct can be generatedAtty Dkt No.: 68244-703601 from a single composition via transposase synapse formation and cleavage, thereby keeping the first adapter construct and the second adapter construct in a same transpososome complex. Alternatively, the first adapter construct and the second adapter construct can be linked together via one or more linker polynucleotides during assembly with a transposase, thereby keeping the first adapter construct and the second adapter construct in a same transpososome complex.
[0289] Performing a transposition reaction with a transpososome complex comprising the first adapter construct and the second adapter construct and the target nucleic acid molecule can thereby generate a plurality of transposition products comprising (I) a first transposition product comprising the first barcode sequence and a first target sequence of the target nucleic acid molecule or reverse complement thereof, and (II) a second transposition product comprising the second barcode sequence and a second target sequence of the target nucleic acid molecule or reverse complement thereof, wherein the first target sequence and the second target sequence were adjacent to each other on the target nucleic acid molecule.
[0290] The method can further comprise sequencing the first transposition product or derivative thereof and the second transposition product or derivative thereof. The method can further comprise identifying the first target sequence and the second target sequence as being adjacent to each other on the target nucleic acid molecule based on the first barcode sequence and the second barcode sequence. The method can further comprise generating an assembled sequence comprising the first target sequence and the second target sequence using the first barcode sequence and the second barcode sequence.
[0291] The method can further comprise providing an additional transpososome comprising an additional adapter pair and a transposase dimer, wherein the additional adapter pair comprises a third adapter construct and a fourth adapter construct, wherein the third adapter construct comprises a third barcode sequence and the fourth adapter construct comprises a fourth barcode sequence. The manner in which the third adapter construct and the fourth adapter construct are generated can keep the third adapter construct and the fourth adapter construct in a same transpososome complex, allowing transposition products tagged by the third adapter construct and the fourth adapter construct to be linked together. For example, the third adapter construct and the fourth adapter construct can be generated from a single composition via transposase synapse formation and cleavage, thereby keeping the third adapter construct and the fourth adapter construct in a same transpososome complex. Alternatively, the third adapter construct and the fourth adapter construct can be linked together via one or more linker polynucleotides during assembly with a transposase, thereby keeping the third adapter construct and the fourth adapter construct in a same transpososome complex.Atty Dkt No.: 68244-703601
[0292] In some cases, the third adapter construct and the fourth adapter construct are coupled to each other via one or more additional linker polynucleotides, which are removed from the third adapter construct and the fourth adapter construct in the additional transpososome prior to performing an additional transposition reaction using the additional transpososome.
[0293] The additional transposition reaction between the third adapter construct, the fourth adapter construct, and the target nucleic acid molecule or a derivative thereof, can thereby generate an additional plurality of transposition products comprising (I) a third transposition product comprising the third barcode sequence and a third target sequence of the target nucleic acid molecule or reverse complement thereof, and (II) a fourth transposition product comprising the fourth barcode sequence and a fourth target sequence of the target nucleic acid molecule or reverse complement thereof, wherein the third target sequence and the fourth target sequence were adjacent to each other on the target nucleic acid molecule.
[0294] The method can further comprise sequencing the third transposition product or derivative thereof and the fourth transposition product or derivative thereof. The method can further comprise identifying the third target sequence and the fourth target sequence as being adjacent to each other on the target nucleic acid molecule based on the third barcode sequence and the fourth barcode sequence. The method can further comprise generating an assembled sequence comprising the first target sequence, the second target sequence, the third target sequence, and the fourth target sequence using the first barcode sequence, the second barcode sequence, the third barcode sequence, and the fourth barcode sequence.
[0295] An example method of performing transposition reactions using a first transpososome comprising a first adapter construct and a second adapter construct and a second transpososome comprising a third adapter construct and a fourth adapter construct and the resulting sequencing library are shown in FIGs. 32A-32E.
[0296] In some aspects, the method comprises generating a plurality of unique pairs of adapter constructs having unique pairs of barcode sequences; performing transposition reactions using the plurality of unique pairs of adapter constructs and the target nucleic acid molecule or derivative thereof to generate a plurality of transposition products comprising barcode sequences of the unique pairs of barcode sequences; and generating an assembled sequence from the plurality of transposition products using the barcode sequences in the plurality of transposition products based on pairing information of the unique pairs of barcode sequences. In some cases, at least 5, at least 10, at least 25, at least 100, at least 1000, at least 10,000, at least 100,000, at least 1 million, at least 10 million, at least 100 million, or at least 1 billion unique pairs of barcodes are provided.
[0297] Generating the plurality of unique pairs of adapter constructs can comprise using a plurality of nucleic acid strands each comprising a unique pair of the unique pairs of barcodeAtty Dkt No.: 68244-703601 sequences. In some cases the method further comprises obtaining the pairing information of the unique pairs of barcode sequences by sequencing the plurality of nucleic acid strands or copies thereof.
[0298] The resulting assembled sequence can have a length of 500 to 10,000 nucleotides, 1000 to 10,000 nucleotides, 2000 to 10,000 nucleotides, 3000 to 10,000 nucleotides, 4000 to 10,000 nucleotides, 5000 to 10,000 nucleotides, 6000 to 10,000 nucleotides, 7000 to 10,000 nucleotides, 8000 to 10,000 nucleotides, 500 to 30,000 nucleotides, 1000 to 30,000 nucleotides, 2000 to 30,000 nucleotides, 3000 to 30,000 nucleotides, 4000 to 30,000 nucleotides, 5000 to 30,000 nucleotides, 6000 to 30,000 nucleotides, 7000 to 30,000 nucleotides, 8000 to 30,000 nucleotides, 10,000 to 30,000 nucleotides, or 15,000 to 30,000 nucleotides. Generating the assembled sequence can comprise sequencing the plurality of transposition products or derivatives thereof to generate sequence reads, and assembling the sequence reads in silico to generate the assembled sequence.
[0299] In some aspects, generating the assembled sequence does not comprise further barcoding transposition products or derivatives thereof of the plurality of transposition products using additional barcode sequences that are different from the barcode sequences of the unique pairs of barcode sequences. In some aspects, generating the assembled sequence does not comprise sequence assembly based on an overlap sequence of the target nucleic acid molecule that is common to transposition products of the plurality of transposition products. In some aspects, generating the assembled sequence does not comprise reference-based assembly. In some aspects, generating the assembled sequence does not comprise sequence alignment of transposition products or derivatives thereof of the plurality of transposition products. In some cases, generating the assembled sequence does not comprise performing additional ligation reactions on the transposition products. In some cases, generating the assembled sequence does not comprise additional fragmentation reactions on the transposition products. In some cases, generating the assembled sequence comprises using at least 5, at least 10, at least 25, at least 100, at least 1000, at least 10,000, at least 100,000, at least 1 million, at least 10 million, at least 100 million unique pairs of barcodes. In some aspects, the method further comprises generating a plurality of assembled sequences using at least 5, at least 10, at least 25, at least 100, at least 1000, at least 10,000, at least 100,000, at least 1 million, at least 10 million, at least 100 million unique pairs of barcodes. In some aspects, the method further comprises sequencing an entire genome using at least 5, at least 10, at least 25, at least 100, at least 1000, at least 10,000, at least 100,000, at least 1 million, at least 10 million, at least 100 million unique pairs of barcodes.E. Processing barcoded products, sequencing, and linking sequence reads
[0300] In some aspects, the method further comprises processing a barcoded nucleic acid molecule resulting from a transposase-mediated insertion reaction, as described elsewhere herein.Atty Dkt No.: 68244-703601The barcoded nucleic acid molecule can be a barcoded insertion product of an adapter construct or barcoded adapter construct described elsewhere herein. In some cases, the barcoded nucleic acid molecule comprises a barcode sequence of the adapter construct or a reverse complement thereof. In some cases, the barcoded nucleic acid molecule comprises a barcode sequence of a barcode construct, as described elsewhere herein, or a reverse complement thereof.
[0301] The processing can comprise one or more nucleic acid extension reactions, nucleic acid amplification reactions, and / or ligation reactions to prepare derivatives of the barcoded nucleic acid molecule for sequencing. For example, the processing can comprise a polymerase extension and ligation reaction for gap sealing, shown in Example 5 and FIGs. 22D and 22E. The processing can link two strands of a barcoded nucleic acid molecule together, for example a strand comprising a barcode sequence or a reverse complement thereof and a strand comprising a target sequence or reverse complement thereof to generate a nucleic acid strand comprising (i) a barcode sequence or a reverse complement thereof and (ii) a target sequence or reverse complement thereof. The processing can further comprise a PCR amplification reaction to amplify one or more strands.
[0302] In some aspects, the method further comprises sequencing the barcoded nucleic molecule or a derivative thereof. In some aspects, the method comprises generating two barcoded nucleic acid molecules, sequencing the two barcoded nucleic acid molecules or derivatives thereof to generate sequence reads, and linking the sequence reads based in part on barcode sequences in the sequence reads in an assembled sequence (e.g., a contig).
[0303] In some aspects, the method comprises (a) providing a nucleic acid (e.g., an adapter construct or integrated adapter construct), an active transposase, and a target nucleic acid molecule comprising a target sequence; (b) generating a barcoded nucleic acid molecule using the nucleic acid, the active transposase, and the target nucleic acid molecule, wherein the barcoded nucleic acid molecule comprises (i) a barcode sequence of the nucleic acid, and (ii) the target sequence; and (c) sequencing the barcoded nucleic acid molecule or derivative thereof. The nucleic acid can be integrated into the target nucleic acid molecule by the active transposase in a transposase- mediated insertion reaction during or prior to (b). The method can further comprise, in (b), generating an additional barcoded nucleic acid molecule using the nucleic acid, wherein the additional barcoded nucleic acid molecule comprises (I) a barcode sequence of the nucleic acid, and (II) an additional target sequence of the target nucleic acid molecule. The method can further comprise sequencing the additional barcoded nucleic acid molecule or a derivative thereof. The method can further comprise linking a first nucleic acid sequence of the barcoded nucleic acid molecule and a second nucleic acid sequence of the additional nucleic acid molecule in an assembled sequence (e.g., in a contig). The linking can be based at least in part on the barcodeAtty Dkt No.: 68244-703601 sequence of the barcoded nucleic acid and the barcode sequence of the additional barcoded nucleic acid molecule.
[0304] The barcode sequence of the barcoded nucleic acid molecule and the barcode sequence of the additional barcoded nucleic acid molecule can identify the barcoded nucleic acid molecule and the additional barcoded nucleic acid molecule as being generated by a common transpososome comprising the nucleic acid and the active transposase. The barcode sequence of the barcoded nucleic acid molecule and the barcode sequence of the additional barcoded nucleic acid molecule can identify the barcoded nucleic acid molecule and the additional barcoded nucleic acid molecule as being generated by a common nucleic acid (e.g., adapter construct). The method can further comprise, in (d), identifying the target sequence and the additional target sequence as being adjacent to each other on the target nucleic acid molecule based on the discrete barcode sequence of the barcoded nucleic acid molecule and the discrete barcode sequence of the additional barcoded nucleic acid molecule. For example, in a trans transposase monomer connection (e.g., type C connection), barcode information can be shared across two transposase monomers of a same transposase dimer for sequence assembly of barcoded short reads. A trans transposase monomer connection (e.g., type C connection) can enable the linking of adjacent reads made by monomers of the same transposase dimer and can enable assembly of adjacent reads without overlap.
[0305] In some cases, the barcode sequence of the nucleic acid is copied from a barcode construct, as described elsewhere herein. In other cases, the barcode sequence of the nucleic acid is not copied from a barcode construct.
[0306] The barcoded nucleic acid molecule and / or the additional barcoded nucleic acid molecule can further comprise one or more additional barcode sequences that are copied or derived from a barcode construct. A barcode sequence derived from a barcode construct can identify the barcode construct from which it was derived. In some cases, the barcoded nucleic acid molecule and the additional barcoded nucleic acid molecule both comprise barcode sequences that identify the two barcoded nucleic acid molecules as having been barcoded by a common barcode construct. In some cases, the barcoded nucleic acid molecule and the additional barcoded nucleic acid molecule both comprise a common barcode sequence that identifies target sequences of the two barcoded nucleic acid molecules as being derived from a common target nucleic acid. In some cases, the barcoded nucleic acid molecule and the additional barcoded nucleic acid molecule both comprise a sample barcode sequence. The barcoded nucleic acid molecule and the additional barcoded nucleic acid can be generated by barcoding target nucleic acid fragments with a sample barcode sequence that identifies the sample from which the target nucleic acid fragments are derived.Atty Dkt No.: 68244-703601
[0307] In one aspect, the method comprises (a) providing: (i) a first nucleic acid (e.g., a barcode construct), comprising a plurality of barcode sequences; (ii) a plurality of transpososomes coupled to the first nucleic acid via the plurality of capture sequences, wherein each transpososome comprises an active transposase and a plurality of polynucleotides, wherein the plurality of polynucleotides comprises a pair of transposon end recognition sequences that are bound to the active transposase; and (iii) a second nucleic acid (e.g., target nucleic acid); (b) generating a barcoded nucleic acid molecule using the first nucleic acid, the second nucleic acid, the active transposase, and a polynucleotide of the plurality of polynucleotides; wherein the barcoded nucleic acid molecule comprises (I) a barcode sequence of the plurality of barcode sequences of the first nucleic acid or a reverse complement thereof, and (II) a nucleic acid sequence of the second nucleic acid or a reverse complement thereof; and (c) sequencing the barcoded nucleic acid molecule or a derivative thereof. In some cases, the barcode sequence identifies the first nucleic acid (e.g., barcode construct). In some cases, the barcode sequence of the barcoded nucleic acid molecule identifies the second nucleic acid (e.g., target nucleic acid), and the method further comprises identifying the barcode sequence or the reverse complement thereof in the barcoded nucleic acid molecule or derivative thereof, thereby identifying its associated nucleic acid sequence as being derived from the second nucleic acid. In some cases, the barcode sequence of the barcoded nucleic acid molecule is a sample barcode sequence, and the method further comprises identifying the barcode sequence or the reverse complement thereof in the barcoded nucleic acid molecule or derivative thereof, thereby identifying its associated nucleic acid sequence as being derived from the sample.
[0308] The method can further comprise, in (b), generating a plurality of barcoded nucleic acid molecules using the first nucleic acid, the second nucleic acid, and the plurality of transpososomes; wherein the plurality of barcoded nucleic acid molecules each comprises (I) a barcode sequence of the plurality of barcode sequences of the first nucleic acid or a reverse complement thereof, and (II) a nucleic acid sequence of the second nucleic acid or a reverse complement thereof. The plurality of barcoded nucleic acid molecules can comprise different nucleic acid sequences of the second nucleic acid. The method can comprise sequencing the plurality of barcoded nucleic acid molecules. The method can further comprise linking nucleic acid sequences of the plurality of barcoded nucleic acid molecules in an assembled sequence, for example, in a contig. The linking can be based at least in part on associated barcode sequences of the plurality of barcoded nucleic acid molecules. In some cases, the plurality of barcode sequences provided in the first nucleic acid in (a) comprises identical barcode sequences. In other cases, the plurality of barcode sequences provided in the first nucleic acid in (a) comprises different barcode sequences.Atty Dkt No.: 68244-703601
[0309] The method can further comprise identifying barcode sequences in the plurality of barcoded nucleic acid molecules or derivatives thereof. In some cases, the method identifies the plurality of barcoded nucleic acid molecules as having been barcoded by a common barcode construct. In some cases, the method identifies the nucleic acid sequences in the plurality of barcoded nucleic acid molecules as being derived from a common nucleic acid molecule. In some cases, the method identifies the nucleic acid sequences in the plurality of barcoded nucleic acid molecules as being derived from the second nucleic acid. In some cases, the plurality of barcoded nucleic acid molecules each comprise a sample barcode sequence, and the method identifies the nucleic acid sequences in the plurality of barcoded nucleic acid molecules as being derived from the sample.
[0310] In some cases, the plurality of polynucleotides of a given transpososome comprises a unique molecular index. The plurality of barcoded nucleic acid molecules generated in (b) can comprise (I) a first barcoded nucleic acid molecule comprising the unique molecular index or reverse complement thereof and a first sequence of the second nucleic acid or reverse complement thereof and (II) a second barcoded nucleic acid molecule comprising the unique molecular index or reverse complement thereof and a second sequence of the second nucleic acid or reverse complement thereof. In some cases, the method further comprises identifying the first sequence and the second sequence as being adjacent to each other on the second nucleic acid based on the unique molecular index.
[0311] In some cases, the plurality of polynucleotides of a given transpososome of the plurality of transpososomes comprises a first unique molecular index and a second unique molecular index. The plurality of barcoded nucleic acid molecules generated in (b) can comprise (I) a first barcoded nucleic acid molecule comprising the first unique molecular index or reverse complement thereof and a first sequence of the second nucleic acid or reverse complement thereof and (II) a second barcoded nucleic acid molecule comprising the second unique molecular index or reverse complement thereof and a second sequence of the second nucleic acid or reverse complement thereof. In some cases, the method further comprises identifying the first sequence and the second sequence as being adjacent to each other on the second nucleic acid based on the first unique molecular index and the second unique molecular index. The first unique molecular index and the second unique molecular index can comprise a same sequence. The first unique molecular index and the second unique molecular index can comprise different sequences. In some cases, the first unique molecular index and the second unique molecular index are on a same polynucleotide of the plurality of polynucleotides. In other cases, the first unique molecular index and the second unique molecular index are on a different polynucleotides of the plurality of polynucleotides.Atty Dkt No.: 68244-703601
[0312] As an example, in a trans transposase monomer connection (e.g., type C connection), barcode information can be shared across two transposase monomers of a same transposase dimer for sequence assembly of barcoded short reads. A trans transposase monomer connection (e.g., type C connection) can enable the linking of adjacent reads made by monomers of the same transposase dimer and can enable assembly of adjacent reads without overlap.
[0313] In some aspects, information across two transposase monomers is shared by physically attaching two polynucleotides of an adapter construct described elsewhere herein. The two polynucleotides of the adapter construct can both comprise a binding sequence that is hybridized to a capture sequence of a barcode construct. The method can comprise (a) providing: (i) a nucleic acid molecule, comprising a plurality of barcode sequences and a plurality of capture sequences; (ii) a plurality of transpososomes coupled to the nucleic acid molecule, wherein each transpososome comprises an active transposase and a plurality of polynucleotides, wherein the plurality of polynucleotides comprises (A) a pair of transposon end recognition sequences that are bound to the active transposase, and (B) two polynucleotides each comprising a binding sequence that is hybridized to a capture sequence of the plurality of capture sequences of the nucleic acid molecule; and (b) attaching the two polynucleotides together. The two polynucleotides can be attached by a nucleic acid extension reaction, a ligation reaction, or a combination thereof.
[0314] In some aspects, the structure of an adapter construct described herein can establish a barcode construct connection (e.g., type A connection) or the sharing / copying of barcode information between the nucleic acid construct or the plurality of polynucleotides and a separate barcode construct comprising one or more barcodes, as depicted in FIG. 14. This sharing of barcode information can be used to link reads for sequence assembly. Examples of adapter constructs that can establish a barcode construct connection include adapter constructs represented in FIGs. 3-13A. Nucleic acid construct 3 and nucleic acid construct 4 in FIG. 3 and FIG. 4 can establish 100% type A barcode construct connections.
[0315] In some aspects, the structure of an adapter construct described herein enables a cis transposase monomer connection (e.g., type B connection) for transposase-mediated insertion or the copying / sharing of barcode information between strands that bind to the same transposase monomer, as depicted in FIG. 16. This sharing of barcode information can be used to link reads for sequence assembly. Examples of adapter constructs that can establish a cis transposase monomer connection include adapter constructs represented in FIGs. 5-9, and FIG. 11-13A. An adapter construct can establish a cis transposase monomer connection (e.g., a type B connection) and a barcode construct connection. In some cases, adapter construct can establish a combined cis transposase monomer connection (e.g., a type B connection) and a barcode construct connection. For example, nucleic acid construct 5 in FIG. 5 can establish 100% combined Type B (cisAtty Dkt No.: 68244-703601 transposase monomer connection) and barcode construct connection. As another example, nucleic acid construct 6 in FIG. 6 can establish 50% combined type B cis transposase monomer connection and barcode construct connections and 50% type B cis transposase monomer connection only. Nucleic acid construct 7 in FIG. 7 can establish 50% combined UMI and barcode construct connection and 50% type B cis transposase monomer connection only.
[0316] In some aspects, the structure of an adapter construct described herein enables a trans transposase monomer connection (e.g., type C linkage) for transposase-mediated insertion or the sharing of barcode information across two transposase monomers of a same transposase dimer for sequence assembly of barcoded short reads, as depicted in FIG. 18. This sharing of barcode information can be used to link reads for sequence assembly. Examples of adapter constructs that can establish a trans transposase monomer connection include adapter constructs represented in FIGs. 6-13B. In some cases, barcode information is linked with pre-sequencing of the adapter construct. For example, adapter constructs shown in FIG. 6 and FIG. 7 can result in a type C connection with pre-sequencing of the adapter construct. In other cases, barcode information is shared without pre-sequencing of the adapter construct. For example, adapter constructs shown in FIGs. 8-13B can result in a type C connection without pre-sequencing.
[0317] An adapter construct can establish a trans transposase monomer connection (e.g., a type C connection) and a barcode construct connection. In some cases, adapter construct can establish a combined trans transposase monomer connection (e.g., a type C connection) and a barcode construct connection. For example, nucleic acid construct 10 in FIG. 10 can establish 100% combined type C trans transposase monomer connection and barcode construct connection. As another example, with pre-sequencing of the barcode pairing in nucleic acid construct 6 or nucleic acid construct 7, 100% type C trans transposase monomer connection can be established by providing information sharing across transposase monomers of a same transposase dimer.
[0318] The structure of an adapter construct described herein can establish both a cis transposase monomer connection (e.g., a type B connection) and a trans transposase monomer connection (e.g., a type C connection) that can be used in sequence assembly. In some cases, the adapter construct can establish a combined cis transposase monomer connection and trans transposase monomer connection. For example, nucleic acid constructs 9, 12, and 13 in FIGs. 9, 12, and 13A can establish 100% combined type B and type C connection. As another example, nucleic acid constructs 8 and 11 in FIGs. 8 and 11 can establish 25% type C trans transposase monomer connection, 25% combined type B cis transposase monomer connection and barcode construct connection, 25% combined UMI and barcode construct connection, and 25% combined type B and type C connection.Atty Dkt No.: 68244-703601
[0319] In some aspects, the method further comprises edge detection for sequence assembly by determining if a read is a terminal read. The method can comprise using an end-sequence that is provided in an adapter construct for edge detection. An example is shown in FIG. 21A for an adapter construct 2100 that establishes a barcode construct connection (e.g., a Type A connection) and that comprises end-sequence 2101. The presence of end-sequence 2101 in sequence reads resulting from insertion of adapter construct 2101 that comprise end-sequence 2101 can indicate that the read is a terminal read. Alternatively, the method can comprise using a UMI sequence in the adapter construct for edge detection. This can be used for adapter constructs that establish a cis or trans transposase monomer connection. An example is shown in FIG. 21B for an adapter construct 2110 that comprises UMI sequence 2111. The presence of the same UMI sequence 2111 on both ends of a read resulting from transposase-mediated insertion of 2110 can indicate that the read is a terminal read.ENUMERATED EMBODIMENTS1. A composition, comprising:(i) a nucleic acid molecule, comprising a plurality of barcode sequences and a plurality of capture sequences; and(ii) a plurality of transpososomes coupled to the nucleic acid molecule, wherein each transpososome comprises an active transposase and a plurality of polynucleotides, wherein the plurality of polynucleotides comprises: (A) a pair of transposon end recognition sequences that are bound to the active transposase and (B) a polynucleotide comprising a binding sequence that is hybridized to a capture sequence of the plurality of capture sequences of the nucleic acid molecule.2. The composition of embodiment 1, wherein a transposon end recognition sequence of the pair of transposon end recognition sequences comprises a mosaic end (ME) sequence or a reverse complement thereof.3. The composition of embodiment 1, wherein a transposon end recognition sequence of the pair of transposon end recognition sequences comprises an inside end (IE) sequence, an outside end (OE) sequence, or a reverse complement thereof.4. The composition of any one of embodiments 1-3, wherein a polynucleotide of the plurality of polynucleotides further comprises a unique molecular index.5. The composition of any one of embodiments 1-4, wherein a polynucleotide of the plurality of polynucleotides further comprises a primer binding sequence or reverse complement thereof.Atty Dkt No.: 68244-7036016. The composition of any one of embodiments 1-5, wherein a polynucleotide of the plurality of polynucleotides further comprises an adapter sequence.7. The composition of any one of embodiments 1-6, wherein the binding sequence of the polynucleotide is on a 3’ end of the polynucleotide.8. The composition of any one of embodiments 1-7, wherein a polynucleotide of the plurality of polynucleotides comprises a transposon end recognition sequence of the pair of transposon end recognition sequences and the binding sequence.9. The composition of any one of embodiments 1-8, wherein a polynucleotide of the plurality of polynucleotides comprises a transposon end recognition sequence of the pair of transposon end recognition sequences, the binding sequence, and a unique molecular index.10. The composition of any one of embodiments 1-9, wherein the active transposase comprises a protein dimer.11. The composition of any one of embodiments 1-10, wherein the nucleic acid molecule comprises at least 10 barcode sequences.12. The composition of any one of embodiments 1-11, wherein the nucleic acid molecule comprises at least 100 barcode sequences.13. The composition of any one of embodiments 1-12, wherein the plurality of barcode sequences comprises a same barcode sequence.14. The composition of any one of embodiments 1-12, wherein the plurality of barcode sequences comprises different barcode sequences.15. The composition of any one of embodiments 1-14, wherein the plurality of capture sequences comprises at least 10 capture sequences.16. The composition of any one of embodiments 1-15, wherein the plurality of capture sequences comprises at least 100 capture sequences.17. The composition of any one of embodiments 1-16, wherein the nucleic acid molecule is at least 500 nucleotides in length.18. The composition of any one of embodiments 1-17, wherein a capture sequence of the plurality of capture sequences is on a 3’ side of a barcode sequence of the plurality of barcode sequences of the nucleic acid molecule.Atty Dkt No.: 68244-70360119. The composition of any one of embodiments 1-18, wherein a capture sequence of the plurality of capture sequences is adjacent to a barcode sequence of the plurality of barcode sequences of the nucleic acid molecule.20. The composition of any one of embodiments 1-19, wherein a barcode sequence of the plurality of barcode sequences is located between two capture sequences of the plurality of capture sequences of the nucleic acid molecule.21. The composition of any one of embodiments 1-20, wherein each transpososome of the plurality of transpososomes comprises two polynucleotides, each comprising a binding sequence that is hybridized to a separate capture sequence of the plurality of capture sequences of the nucleic acid molecule.22. The composition of embodiment 21, wherein each separate capture sequence that is hybridized to a polynucleotide of the two polynucleotides is on a 3’ side of a barcode sequence of the plurality of barcode sequences.23. The composition of embodiment 21 or 22, wherein each separate capture sequence that is hybridized to a polynucleotide of the two polynucleotides is adjacent to a barcode sequence of the plurality of barcode sequences.24. The composition of any one of embodiments 1-23, wherein the plurality of polynucleotides comprises four nucleic acid strands comprising: a first nucleic acid strand comprising a first transposon end recognition sequence; a second nucleic acid strand comprising a reverse complement of the first transposon end recognition sequence and a first binding sequence that is hybridized to a first capture sequence of the plurality of capture sequences of the nucleic acid molecule; a third nucleic acid strand comprising a second transposon end recognition sequence; and a fourth nucleic acid strand comprising a reverse complement of the second transposon end recognition sequence, and a second binding sequence that is hybridized to a second capture sequence of the plurality of capture sequences of the nucleic acid molecule.25. The composition of embodiment 24, wherein the first nucleic acid strand or the third nucleic acid strand further comprises an adapter sequence.26. The composition of embodiment 24 or 25, wherein the second nucleic acid strand or the fourth nucleic acid strand further comprises a unique molecular index.27. The composition of any one of embodiments 24-26, wherein the second nucleic acid strand and the fourth nucleic acid strand each comprises a unique molecular index.Atty Dkt No.: 68244-70360128. The composition of any one of embodiments 1-23, wherein the plurality of polynucleotides comprises three nucleic acid strands comprising: a first nucleic acid strand comprising (A) a first transposon end recognition sequence, (B) a reverse complement of a second transposon end recognition sequence; a second nucleic acid strand comprising a reverse complement of the first transposon end recognition sequence, wherein a portion of the second nucleic acid strand is hybridized to a first segment of the first nucleic acid strand; and a third nucleic acid strand comprising the second transposon end recognition sequence, wherein a portion of the third nucleic acid strand is hybridized to a second segment of the first nucleic acid strand.29. The composition of embodiment 28, wherein the second nucleic acid strand or the third nucleic acid strand further comprises a binding sequence that is hybridized to a capture sequence of the plurality of capture sequences of the nucleic acid molecule.30. The composition of embodiment 28 or 29, wherein the first nucleic acid strand further comprises a unique molecular index located between the first transposon recognition sequence and the second transposon recognition sequence.31. The composition of any one of embodiment 28-30, wherein the first nucleic acid strand further comprises two discrete unique molecular indexes located between the first transposon recognition sequence and the reverse complement of the second transposon recognition sequence.32. The composition of any one of embodiments 1-23, wherein the plurality of polynucleotides comprises four nucleic acid strands comprising: a first nucleic acid strand comprising a first transposon end recognition sequence, wherein the first nucleic acid strand comprises a first region and a second region; a second nucleic acid strand comprising second transposon end recognition sequence, wherein the fourth nucleic acid strand comprises a third region and a fourth region; wherein the first region of the first nucleic acid strand is hybridized to the third region of the second nucleic acid strand; a third nucleic acid strand comprising a reverse complement of the first transposon end recognition sequence, wherein a portion of the third nucleic acid strand is hybridized to the second region of the first nucleic acid strand; and a fourth nucleic acid strand comprising a reverse complement of the second transposon end recognition sequence, wherein a portion of the fourth nucleic acid strand is hybridized to the fourth region of the second nucleic acid strand.Atty Dkt No.: 68244-70360133. The composition of embodiment 32, wherein the third nucleic acid strand or the fourth nucleic acid strand comprises a binding sequence that is hybridized to a capture sequence of the plurality of capture sequences of the nucleic acid molecule.34. The composition of embodiment 32 or 33, wherein the first nucleic acid strand or the second nucleic acid strand further comprises a unique molecular index.35. The composition of any one of embodiments 1-34, wherein the nucleic acid molecule is single-stranded.36. The composition of any one of embodiments 1-35, wherein barcode sequences of the plurality of barcode sequences are separated by a non-canonical nucleotide.37. The composition of embodiment 36, wherein the non-canonical nucleotide is non- amplifiable.38. The composition of any one of embodiments 1-37, wherein barcode sequences of the plurality of barcode sequences are separated by a recognition sequence bound to a DNA-binding protein.39. The composition of embodiment 38, wherein the DNA-binding protein comprises a CRISPR-Cas domain, and wherein the recognition sequence is hybridized to a guide RNA complexed with the CRISPR-Cas domain.40. The composition of embodiment 39, wherein the CRISPR-Cas domain comprises an inactive nuclease domain.41. The composition of embodiment 39 or 40, wherein the CRISPR-Cas domain comprises a Cas9 domain.42. A method, comprising:(a) providing:(i) a first nucleic acid, comprising a plurality of barcode sequences and a plurality of capture sequences;(ii) a plurality of transpososomes coupled to the first nucleic acid via the plurality of capture sequences, wherein each transpososome comprises an active transposase and a plurality of polynucleotides, wherein the plurality of polynucleotides comprises a pair of transposon end recognition sequences that are bound to the active transposase; and(iii) a second nucleic acid; andAtty Dkt No.: 68244-703601(b) generating a barcoded nucleic acid molecule using the first nucleic acid, the second nucleic acid, the active transposase, and a polynucleotide of the plurality of polynucleotides; wherein the barcoded nucleic acid molecule comprises (I) a barcode sequence of the plurality of barcode sequences of the first nucleic acid or a reverse complement thereof, and (II) a nucleic acid sequence of the second nucleic acid or a reverse complement thereof.43. The method of embodiment 42, wherein (b) comprises (i) performing a nucleic acid extension reaction on the polynucleotide or a derivative thereof using a segment of the first nucleic acid as a template, and (ii) performing a transposase-mediated insertion reaction using the active transposase to integrate the polynucleotide or a derivative thereof into the second nucleic acid.44. The method of embodiment 43, wherein (b)(i) and (b)(ii) occur simultaneously.45. The method of embodiment 43, wherein (b)(i) and (b)(ii) occur sequentially.46. The method of embodiment 45, wherein (b)(i) occurs prior to (b)(ii).47. The method of embodiment 45, wherein (b)(i) occurs after (b)(ii).48. The method of any one of embodiments 42-47, wherein, in (b), the active transposase integrates a barcoded derivative of the polynucleotide into the second nucleic acid in a transposase- mediated insertion reaction, wherein the barcoded derivative comprises the barcode sequence or a reverse complement thereof.49. The method of any one of embodiments 48, wherein the plurality of polynucleotides comprises a polynucleotide comprising a binding sequence that is hybridized to a capture sequence of the plurality of capture sequences.50. The method of embodiment 49, wherein the binding sequence of the polynucleotide is on a 3’ end of the polynucleotide, and (b) comprises performing a nucleic acid extension reaction using the polynucleotide of the plurality of polynucleotides as a primer and a segment of the first nucleic acid as a template to generate the barcoded derivative of the polynucleotide.51. The method of any one of embodiments 48-50, wherein the active transposase integrates the barcoded derivative of the polynucleotide into the second nucleic acid into a random location in the second nucleic acid.52. The method of any one of embodiment 42-51, wherein (b) further comprises covalently linking an end of the polynucleotide or a barcoded derivative of the polynucleotide and the nucleicAtty Dkt No.: 68244-703601 acid sequence of the second nucleic acid or a reverse complement thereof in a ligation reaction or a nucleic acid gap fill-in reaction.53. The method of any one of embodiments 42-52, wherein (b) further comprises incorporating a dideoxynucleotide (ddNTP) at a 3’ end of an extension product of polynucleotide or a derivative thereof.54. The method of any one of embodiments 42-53, wherein the barcoded nucleic acid molecule comprises a length of from 100 to 600 bp.55. The method of any one of embodiments 42-54, further comprising (c) sequencing the barcoded nucleic acid molecule or a derivative thereof.56. The method of embodiment 55, wherein the barcode sequence identifies the second nucleic acid, and wherein the method further comprises (d) identifying the barcode sequence or the reverse complement thereof in the barcoded nucleic acid molecule or the derivative thereof, thereby identifying its associated nucleic acid sequence as being derived from the second nucleic acid.57. The method of any one of embodiments 42-56, wherein (b) further comprises generating a plurality of barcoded nucleic acid molecules using the first nucleic acid, the second nucleic acid, and the plurality of transpososomes; wherein the plurality of barcoded nucleic acid molecules each comprises (I) a barcode sequence of the plurality of barcode sequences of the first nucleic acid or a reverse complement thereof, and (II) a nucleic acid sequence of the second nucleic acid or a reverse complement thereof.58. The method of embodiment 57, wherein the plurality of barcoded nucleic acid molecules comprise different nucleic acid sequences of the second nucleic acid.59. The method of embodiment 57 or 58, further comprising (c) sequencing the plurality of barcoded nucleic acid molecules.60. The method of embodiment 59, further comprising (d) linking nucleic acid sequences of the plurality of barcoded nucleic acid molecules in an assembled sequence.61. The method of embodiment 60, wherein the linking in (d) is based at least in part on associated barcode sequences of the plurality of barcoded nucleic acid molecules.62. The method of any one of embodiments 57-61, wherein, in (a), the plurality of barcode sequences comprises identical barcode sequences.63. The method of any one of embodiments 60-62, wherein, in (a), the plurality of polynucleotides of a given transpososome of the plurality of transpososomes comprises a unique molecular index;Atty Dkt No.: 68244-703601 wherein, in (b), the plurality of barcoded nucleic acid molecules comprises (I) a first barcoded nucleic acid molecule comprising the unique molecular index and a first sequence of the second nucleic acid and (II) a second barcoded nucleic acid molecule comprising the unique molecular index and a second sequence of the second nucleic acid; and wherein (d) comprises identifying the first sequence and the second sequence as being adjacent to each other on the second nucleic acid based on the unique molecular index.64. The method of any one of embodiments 60-62, wherein, in (a), the plurality of polynucleotides of a given transpososome of the plurality of transpososomes comprises a first unique molecular index and a second unique molecular index; wherein, in (b), the plurality of barcoded nucleic acid molecules comprises (I) a first barcoded nucleic acid molecule comprising the first unique molecular index and a first sequence of the second nucleic acid and (II) a second barcoded nucleic acid molecule comprising the second unique molecular index and a second sequence of the second nucleic acid; and wherein (d) comprises identifying the first sequence and the second sequence as being adjacent to each other on the second nucleic acid based on the first unique molecular index and the second unique molecular index.65. The method of embodiment 64, wherein the first unique molecular index and the second unique molecular index are on a same polynucleotide of the plurality of polynucleotides.66. The method of any one of embodiments 42-65, further comprising, prior to (a), generating a polynucleotide of the plurality of polynucleotides.67. The method of any one of embodiments 42-66, further comprising, prior to (a), generating the first nucleic acid.68. The method of embodiment 67, wherein the generating the first nucleic acid comprises performing a nucleic acid amplification reaction on a template nucleic acid molecule comprising a barcode sequence of the plurality of barcode sequences and a capture sequence of the plurality of capture sequences.69. The method of embodiment 68, wherein the nucleic acid amplification reaction comprises rolling circle amplification.70. The method of any one of embodiments 42-69, wherein the second nucleic acid is doublestranded.71. The method of any one of embodiments 42-70, wherein the second nucleic acid comprises genomic DNA.72. The method of any one of embodiments 42-71, wherein the second nucleic acid is in a cell.Atty Dkt No.: 68244-70360173. A method, comprising:(a) providing:(i) a nucleic acid molecule, comprising a plurality of barcode sequences and a plurality of capture sequences;(ii) a plurality of transpososomes coupled to the nucleic acid molecule, wherein each transpososome comprises an active transposase and a plurality of polynucleotides, wherein the plurality of polynucleotides comprises (A) a pair of transposon end recognition sequences that are bound to the active transposase, and (B) a polynucleotide comprising a binding sequence that is hybridized to a capture sequence of the plurality of capture sequences of the nucleic acid molecule; and(b) performing a nucleic acid extension reaction to extend a polynucleotide of a transpososome of the plurality of transpososomes that is hybridized to a capture sequence of the nucleic acid molecule to yield an extended polynucleotide comprising (I) the binding sequence and (II) a reverse complement of a barcode sequence of the plurality of barcode sequences.74. A method, comprising:(a) providing:(i) a nucleic acid molecule, comprising a plurality of barcode sequences and a plurality of capture sequences;(ii) a plurality of transpososomes coupled to the nucleic acid molecule, wherein each transpososome comprises an active transposase and a plurality of polynucleotides, wherein the plurality of polynucleotides comprises (A) a pair of transposon end recognition sequences that are bound to the active transposase, and (B) two polynucleotides each comprising a binding sequence that is hybridized to a capture sequence of the plurality of capture sequences of the nucleic acid molecule; and(b) attaching the two polynucleotides together.75. A nucleic acid comprising:Atty Dkt No.: 68244-703601 a first nucleic acid strand comprising (A) a first transposon end recognition sequence, (B) a reverse complement of a second transposon end recognition sequence, and (C) two discrete barcode sequences located between the first transposon end recognition sequence and the second transposon end recognition sequence; a second nucleic acid strand comprising a reverse complement of the first transposon end recognition sequence, wherein a portion of the second nucleic acid strand is hybridized to a first segment of the first nucleic acid strand; and a third nucleic acid strand comprising the second transposon end recognition sequence, wherein a portion of the third nucleic acid strand is hybridized to a second segment of the first nucleic acid strand; wherein the second nucleic acid strand and the third nucleic acid strand are separate nucleic acid strands.76. The nucleic acid of embodiment 75, wherein the two discrete barcode sequences are identical sequences.77. The nucleic acid of embodiment 75, wherein the two discrete barcode sequences comprise different sequences.78. The nucleic acid of any one of embodiments 75-77, wherein the first nucleic acid strand further comprises a functional sequence located between the first transposon end recognition sequence and a discrete barcode sequence of the two discrete barcode sequences.79. The nucleic acid of embodiment 78, wherein the functional sequence of the first nucleic acid strand distinguishes the first nucleic acid strand from a reverse complement of the first nucleic acid strand.80. The nucleic acid of embodiment 78 or 79, wherein the functional sequence of the first nucleic acid strand is at most 3 nucleotides in length.81. The nucleic acid of any one of embodiments 75-77, wherein the first nucleic acid strand further comprises (i) a first functional sequence located between the first transposon end recognition sequence and a first discrete barcode sequence of the two discrete barcode sequences and (ii) a second functional sequence located between the second transposon end recognition sequence and a second discrete barcode of the two discrete barcode sequences.82. The nucleic acid of any one of embodiments 75-81, wherein the first nucleic acid strand further comprises a primer binding sequence or reverse complement thereof located between the two discrete barcode sequences.Atty Dkt No.: 68244-70360183. The nucleic acid of any one of embodiments 75-82, wherein the first nucleic acid strand further comprises two distinct primer binding sequences or reverse complements thereof located between the two discrete barcode sequences.84. The nucleic acid of any one of embodiments 75-83, wherein the first nucleic acid strand further comprises a hairpin region located between the two discrete barcode sequences, wherein the hairpin region comprises a first portion of the first nucleic acid strand that is hybridized to a second portion of the first nucleic acid strand.85. The nucleic acid of any one of embodiments 75-84, wherein the second nucleic acid strand further comprises a reverse complement of a discrete barcode sequence of the two discrete barcode sequences.86. The nucleic acid of any one of embodiments 75-84, wherein the second nucleic acid strand does not comprise a reverse complement of a discrete barcode sequence of the two discrete barcode sequences.87. The nucleic acid of any one of embodiments 75-86, wherein the second nucleic acid strand and the third nucleic acid strand do not comprise a reverse complement of a discrete barcode sequence of the two discrete barcode sequences.88. The nucleic acid of any one of embodiments 75-87, wherein the first nucleic acid strand comprises a non-hybridizing region that is not hybridized to the second nucleic acid strand and that is not hybridized to the third nucleic acid strand.89. The nucleic acid of embodiment 88, wherein the non-hybridizing region of the first nucleic acid strand comprises a primer binding sequence or reverse complement thereof.90. The nucleic acid of embodiment 88 or 89, wherein the non-hybridizing region of the first nucleic acid strand comprises two distinct primer binding sequences or reverse complements thereof.91. The nucleic acid of any one of embodiments 88-90, wherein the non-hybridizing region of the first nucleic acid strand comprises a discrete barcode sequence of the two discrete barcode sequences.92. The nucleic acid of any one of embodiments 88-91, wherein the non-hybridizing region of the first nucleic acid strand is located between the first segment and the second segment of the first nucleic acid strand.Atty Dkt No.: 68244-70360193. The nucleic acid of any one of embodiments 88-91, wherein the non-hybridizing region of the first nucleic acid strand is located between two hybridizing regions of the first nucleic acid strand, wherein the two hybridizing regions are each hybridized to the second nucleic acid strand.94. The nucleic acid of any one of embodiments 75-93, wherein the first nucleic acid strand comprises two distinct non-hybridizing regions that are not hybridized to the second nucleic acid strand and that are not hybridized to the third nucleic acid strand.95. The nucleic acid of embodiment 94, wherein the two distinct non -hybridizing regions of the first nucleic acid strand each comprises a primer binding sequence or reverse complement thereof.96. The nucleic acid of embodiment 94 or 95, wherein the two distinct non-hybridizing regions of the first nucleic acid strand comprise a primer binding sequence or reverse complement thereof and a discrete barcode sequence of the two discrete barcode sequences.97. The nucleic acid of any one of embodiments 94-96, wherein: a first non-hybridizing region of the two non-hybridizing regions of the first nucleic acid strand is located between the first segment and the second segment of the first nucleic acid strand; and a second non-hybridizing region of the two non -hybridizing regions of the first nucleic acid strand is located between two hybridizing regions of the first nucleic acid strand, wherein the two hybridizing regions are each hybridized to the second nucleic acid strand.98. The nucleic acid of any one of embodiments 75-97, wherein the second nucleic acid strand comprises a non-complementary region that is not hybridized to the first nucleic acid strand.99. The nucleic acid of embodiment 98, wherein the non-complementary region of the second nucleic acid strand comprises a primer binding sequence or reverse complement thereof.100. The nucleic acid of embodiment 98 or 99, wherein the non-complementary region of the second nucleic acid strand comprises a binding sequence that is hybridized to a splint oligonucleotide.101. The nucleic acid of embodiment 98 or 99, wherein the non-complementary region of the second nucleic acid strand comprises a binding sequence that is hybridized to a nucleic acid molecule comprising a third discrete barcode sequence.102. The nucleic acid of embodiment 101, wherein the nucleic acid molecule comprises a plurality of barcode sequences.Atty Dkt No.: 68244-703601103. The nucleic acid of any one of embodiments 98-102, wherein the non-complementary region of the second nucleic acid strand is on a 3’ end of the second nucleic acid strand.104. The nucleic acid of any one of embodiments 75-103, wherein the second nucleic acid strand comprises two distinct non-complementary regions that are not hybridized to the first nucleic acid strand.105. The nucleic acid of embodiment 104, wherein the two distinct non-complementary regions comprise a primer binding sequence or reverse complement thereof.106. The nucleic acid of embodiment 104 or 105, wherein: a first non-complementary region of the two non-complementary regions of the second nucleic acid strand is located on a 3’ end of the second nucleic acid strand; and a second non-complementary region of the two non-complementary regions of the second nucleic acid strand is located in an internal segment of the second nucleic acid strand.107. The nucleic acid of any one of embodiments 75-106, wherein the third nucleic acid strand further comprises a non-complementary region that is not hybridized to the first nucleic acid strand.108. The nucleic acid of embodiment 107, wherein the non-complementary region of the third nucleic acid strand comprises a primer binding sequence or reverse complement thereof.109. The nucleic acid of embodiment 107 or 108, wherein the non-complementary region of the third nucleic acid strand is on a 3’ end of the third nucleic acid strand.110. A method, comprising:(a) providing the nucleic acid of any one of embodiments 75-109, an active transposase, and a target nucleic acid molecule comprising a target sequence; and(b) generating a barcoded nucleic acid molecule using the nucleic acid, the active transposase, and the target nucleic acid molecule, wherein the barcoded nucleic acid molecule comprises (i) a discrete barcode sequence of the two discrete barcode sequences, and (ii) the target sequence.111. The method of embodiment 110, wherein the active transposase integrates the nucleic acid into the target nucleic acid molecule in a transposase-mediated insertion reaction.112. The method of embodiment 111, further comprising (c) sequencing the barcoded nucleic acid molecule or a derivative thereof.Atty Dkt No.: 68244-703601113. The method of any one of embodiments 110 or 111, wherein (b) further comprises generating an additional barcoded nucleic acid molecule using the nucleic acid, wherein the additional barcoded nucleic acid molecule comprises (I) a discrete barcode sequence of the two discrete barcode sequences, and (II) an additional target sequence of the target nucleic acid molecule.114. The method of embodiment 113, further comprising (c) sequencing the barcoded nucleic acid molecule or a derivative thereof and sequencing the additional barcoded nucleic acid molecule or a derivative thereof.115. The method of embodiment 114, further comprising (d) linking a nucleic acid sequence of the barcoded nucleic acid molecule and a nucleic acid sequence of the additional barcoded nucleic acid molecule in an assembled sequence.116. The method of any one of embodiments 113-115, wherein the discrete barcode sequence of the barcoded nucleic acid molecule and the discrete barcode sequence of the additional barcoded nucleic acid molecule identify the barcoded nucleic acid molecule and the additional barcoded nucleic acid molecule as being generated by a common transpososome comprising the nucleic acid and the active transposase.117. The method of any one of embodiments 113-116, wherein (d) comprises identifying the target sequence and the additional target sequence as being adjacent to each other on the target nucleic acid molecule based on the discrete barcode sequence of the barcoded nucleic acid molecule and the discrete barcode sequence of the additional barcoded nucleic acid molecule.118. The nucleic acid of any one of embodiments 75-109, wherein the nucleic acid is integrated into a target nucleic acid molecule.119. A nucleic acid comprising: a first nucleic acid strand comprising a first transposon end recognition sequence, wherein the first nucleic acid strand comprises a first region and a second region; a second nucleic acid strand comprising second transposon end recognition sequence, wherein the fourth nucleic acid strand comprises a third region and a fourth region; wherein the first region of the first nucleic acid strand is hybridized to the third region of the second nucleic acid strand; a third nucleic acid strand comprising a reverse complement of the first transposon end recognition sequence, wherein a portion of the third nucleic acid strand is hybridized to the second region of the first nucleic acid strand; andAtty Dkt No.: 68244-703601 a fourth nucleic acid strand comprising a reverse complement of the second transposon end recognition sequence, wherein a portion of the fourth nucleic acid strand is hybridized to the fourth region of the second nucleic acid strand.120. The nucleic acid of embodiment 119, wherein the first nucleic acid strand or the second nucleic acid strand further comprises a barcode sequence.121. The nucleic acid of embodiment 119 or 120, wherein the first nucleic acid strand further comprises a hairpin region.122. A method, comprising:(a) providing the nucleic acid of any one of embodiments 119-121, an active transposase, and a target nucleic acid molecule comprising a target sequence; and(b) generating a barcoded nucleic acid molecule using the nucleic acid, the active transposase, and the target nucleic acid molecule, wherein the barcoded nucleic acid molecule comprises (i) a discrete barcode sequence of the two discrete barcode sequences, and (ii) the target sequence.123. The method of embodiment 122, wherein the active transposase integrates the nucleic acid into the target nucleic acid molecule in a transposase-mediated insertion reaction.124. The method of embodiment 123, further comprising (c) sequencing the barcoded nucleic acid molecule or a derivative thereof.125. The method of any one of embodiments 122 or 123, wherein (b) further comprises generating an additional barcoded nucleic acid molecule using the nucleic acid, wherein the additional barcoded nucleic acid molecule comprises (I) a discrete barcode sequence of the two discrete barcode sequences, and (II) an additional target sequence of the target nucleic acid molecule.126. The method of embodiment 125, further comprising (c) sequencing the barcoded nucleic acid molecule or a derivative thereof and sequencing the additional barcoded nucleic acid molecule or a derivative thereof.127. The method of embodiment 126, further comprising (d) linking a nucleic acid sequence of the barcoded nucleic acid molecule and a nucleic acid sequence of the additional barcoded nucleic acid molecule in an assembled sequence.128. The method of any one of embodiments 124-127, wherein the discrete barcode sequence of the barcoded nucleic acid molecule and the discrete barcode sequence of the additional barcoded nucleic acid molecule identify the barcoded nucleic acid molecule and the additional barcodedAtty Dkt No.: 68244-703601 nucleic acid molecule as being generated by a common transpososome comprising the nucleic acid and the active transposase.129. The method of any one of embodiments 124-128, wherein (d) comprises identifying the target sequence and the additional target sequence as being adjacent to each other on the target nucleic acid molecule based on the discrete barcode sequence of the barcoded nucleic acid molecule and the discrete barcode sequence of the additional barcoded nucleic acid molecule.130. The nucleic acid of any one of embodiments 119-121, wherein the nucleic acid is integrated into a target nucleic acid molecule.EXAMPLESExample 1. Barcode Construct Connection (E.G., Type A Connection) Enables Sharing / Copying Of Barcode Information Between Target Nucleic Acid And A Barcode Construct
[0320] This example provides a method of generating a barcode construct connection (e.g., Type A connection), as described herein, for the sharing / copying of barcode information between target nucleic acid and a separate barcode construct comprising one or more barcodes. A schematic of a barcode construct connection is provided in FIG. 14, which shows information flow from a barcode construct to target nucleic acid enabled by transposase-mediated insertion of an adapter construct described herein. This example uses a combination of barcoding of the adapter construct and transposase-mediated insertion of the adapter construct into the target nucleic acid.
[0321] In this example, an adapter construct 1510 is barcoded using barcode construct 1520, shown in FIG. 15A. The adapter construct 1510 comprises a transposon end comprising transposon end recognition sequence 1511 and a binding sequence 1512 that binds to a capture sequence 1521 of a subregion 1520A of the barcode construct 1520. Adapter construct 1510 enables copying of barcode sequence 1522 and additional functional sequences 1523 and 1524 of the barcode construct by complementary DNA (cDNA) synthesis.
[0322] As shown in FIG. 15B, barcode construct 1520 comprises a plurality of subregions 1520A, each comprising a capture sequence 1521, a barcode sequence 1522, and functional sequences 1523 and 1524. A plurality of adapter constructs 1510 hybridize to a plurality of capture sequences on the barcode construct and are extended to generate barcoded adapter constructs comprising cDNA segment 1530 using the barcode construct as a template. The cDNA segment 1530 comprises the reverse complement of the barcode sequence 1522 and the reverse complements of functional sequences 1523 and 1524. The barcoded adapter constructs are integrated into target nucleic acid 1540 using transposase-mediated insertion by transposases 1550.Atty Dkt No.: 68244-703601As an alternative, the adapter constructs 1510 can also be integrated into target nucleic acid prior to barcoding by the barcode construct, and the integrated adapter constructs can be barcoded after transposase-mediated insertion.
[0323] The adapter construct 1510 in this example generates a barcode construct connection (e.g., Type A connection) that enables sharing / copying of barcode information between target nucleic acid 1540 and the barcode construct 1520.Example 2. Barcode Construct Connection (E.G., Type A) And Cis Transposase Monomer Connection (E.G., Type B Connection) Enables Sharing / Copying Of Barcode Information Between Target Nucleic Acid And A Barcode Construct And Sharing / Copying Of Barcode Information Between Strands That Bind To The Same Transposase Monomer
[0324] This example provides a method of generating a barcode construct connection (e.g., Type A connection) and a cis transposase monomer connection (e.g., Type B connection), as described herein. This method shares barcode information between target nucleic acid and a separate barcode construct and additionally shares barcode information between nucleic acid strands that bind to the same transposase monomer for transposase-mediated insertion. A schematic of a cis transposase monomer connection is provided in FIG. 16, showing information flow between strands that bind to a same transposase monomer.
[0325] In this example, an adapter construct 1700 comprising strand 1710 and oligonucleotide 1720 is provided, shown in FIG. 17A. Strand 1710 comprises a barcode sequence 1711 and a sequence 1712RC that is a reverse complement of a transposon end recognition sequence. Strand 1720 further comprises a binding sequence 1713 configured to bind to a capture sequence 1751 of a subregion 1750A of a barcode construct 1750. Oligonucleotide 1720 is hybridized to a portion of strand 1710. In this example, oligonucleotide 1720 is extended in a nucleic acid extension reaction to generate an extended strand 1730 comprising the transposon end recognition sequence and a reverse complement of barcode sequence 1711. The binding sequence 1713 hybridizes to a capture sequence 1751 of barcode construct 1750. Adapter construct 1700 enables copying of barcode sequence 1711 in strand 1710 in the adapter construct and copying of barcode sequence 1752 and functional sequences 1753 and 1754 of the barcode construct by complementary DNA (cDNA) synthesis. The copying of barcode sequence 1711 and the copying of barcode sequence 1752 can occur simultaneously or sequentially in separate steps.
[0326] As shown in FIG. 17B, the barcode construct 1750 comprises a subregions 1750A, each comprising a capture sequence 1751, a sample barcode sequence 1752, and functional sequences 1753 and 1754. Binding sequences 1713 hybridize to a plurality of capture sequences on the barcode construct and are extended to generate barcoded adapter constructs comprising theAtty Dkt No.: 68244-703601 reverse complements of barcode sequence 1752 and functional sequences 1753 and 1754 using the barcode construct as a template. The barcoded adapter constructs are integrated into target nucleic acid 1760 using transposase-mediated insertion by transposases 1770. The copying of barcode sequences 1711 and 1752 can both occur prior to transposase-mediated insertion into a target nucleic acid. Alternatively, copying of barcode sequence 1711 can occur prior to transposase-mediated insertion and copying of sample barcode sequence 1752 can occur after transposase-mediated insertion.
[0327] The insertion products are then processed and sequenced. The short reads resulting from the fragments are assembled together in an assembled sequence based in part on barcode sequence 1711 and sample barcode sequence 1752. The sample barcode sequence identifies each short read as being derived from a common sample.
[0328] In this example, the barcode construct connection (e.g. Type A connection) shares barcode information (e.g., sample barcode sequence 1752) between target nucleic acid 1760 and the barcode construct 1750. The cis transposase monomer connection (e.g. Type B connection) shares barcode information (e.g., barcode sequence 1711) between strands that bind to the same transposase monomer.Example 3. Trans Transposase Monomer Connection (E.G., Type C Connection) Enables Sharing / Copying Of Barcode Information Across Two Transposase Monomers Of A Same Transposase Dimer And Sequence Assembly Of Short Reads Without Overlap
[0329] This example provides a method of generating a trans transposase monomer connection (e.g., Type C connection), as described herein. This method shares of barcode information across two transposase monomers of a same transposase dimer in a transpososome complex for transposase-mediated insertion. A schematic of a trans transposase monomer connection is provided in FIG. 18, showing information flow across two transposase monomers of a same transposase dimer in a transpososome complex.
[0330] In this example, an adapter construct 1900 comprising three nucleic acid strands 1900i, 1910ii, and 1900iii is provided. Nucleic acid strand 1900i comprises (i) a sequence 1914RC that is a reverse complement of transposon end recognition sequence 1914 and (ii) a barcode sequence 1911. Nucleic acid strand 1910ii comprises (i) a sequence 1913RC that is a reverse complement of a transposon end recognition sequence 1913 and (ii) a sequence 1912RC that is a reverse complement of barcode sequence 1912RC. Nucleic acid strand 1900iii and 1910ii hybridize to different portions of nucleic acid strand 1900i. Adapter construct 1900 is used to generate adapter construct 1910 via nucleic acid extension of strands 1900i and 1900iii. Nucleic acid strand 1900i is extended to generate nucleic acid strand 1910i that comprises (i) barcode sequence 1911, (ii) barcode sequence 1912 and (iii) a transposon end sequence 1913. Nucleic acid strand 1900iii isAtty Dkt No.: 68244-703601 extended to generate nucleic acid strand 1910iii that comprises (i) a reverse complement of barcode sequence 1911, and (ii) transposon end sequence 1914.
[0331] In this example, active transposases 1901 inserts a plurality of different adapter constructs 1910, 1920, and 1930 into a target nucleic acid 1940. Adapter construct 1920 comprises barcode sequences 1921 and 1922, and adapter construct 1930 comprises barcode sequences 1931 and 1932. The insertion products generated by adapter construct 1910 comprise a strand comprising barcode sequences 1911 and 1912 and a strand comprising the reverse complement of 1911. The insertion products generated by adapter construct 1920 comprise a strand comprising barcode sequences 1921 and 1922 and a strand comprising the reverse complement of 1921. The insertion products generated by adapter construct 1930 comprise a strand comprising barcode sequences 1931 and 1932 and a strand comprising the reverse complement of 1931.
[0332] The insertion products are then processed and sequenced. Insertion products are processed by nucleic acid extension and ligation and further amplification for preparation of the sequencing library. A processing reaction links a strand comprising the reverse complement of 1912 with a strand comprising the reverse complement of 1921. The short reads resulting from the fragments are assembled together in an assembled sequence based in part on the barcode sequences.
[0333] The trans transposase monomer connections generated by each of these adapter constructs pairs two adjacent reads or fragments generated by two different monomers of the same transposase dimer. For instance, a read comprising the barcode sequences 1911 and 1912 is paired with a read comprising the reverse complement of 1911. Adjacent reads made by monomers of the same transposase dimer are linked together using the barcode information, allowing for assembly of adjacent reads without overlap.Example 4. Barcode Construct Connection (E.G., Type A Connection) And Trans Transposase Monomer Connection (E.G., Type C Connection) Enables Sharing / Copying Of Barcode Information Between Target Nucleic Acid And A Barcode Construct And Sharing / Copying Of Barcode Information Across Two Transposase Monomers Of A Same Transposase Dimer
[0334] This example provides a method of generating barcode construct connection (e.g., TypeA connection) a trans transposase monomer connection (e.g., Type C connection), as described herein. This method shares barcode information between target nucleic acid and a separate barcode construct and additionally shares of barcode information across two transposase monomers of a same transposase dimer in a transpososome complex for transposase-mediated insertion.Atty Dkt No.: 68244-703601
[0335] In this example, an adapter construct 2000 comprising three nucleic acid strands 2000i, 201 Oii, and 2000iii is provided, as shown in FIG. 20A. Nucleic acid strand 2000i comprises (i) a sequence 2014RC that is a reverse complement of transposon end recognition sequence 2014, and (ii) a barcode sequence 2011. Nucleic acid strand 201 Oii comprises (i) a sequence 2013RC that is a reverse complement of a transposon end recognition sequence 2013, (ii) a sequence 2012RC that is a reverse complement of barcode sequence 2012RC, and (iii) a binding sequence 2015. Nucleic acid strand 2010iii comprises binding sequence 2016. Nucleic acid strand 2000iii and 1910ii hybridize to different portions of nucleic acid strand 2000i. Adapter construct 2000 is used to generate adapter construct 2010 via nucleic acid extension of strands 2000i and 2000iii. Nucleic acid strand 2000i is extended to generate nucleic acid strand 1910i that comprises (i) barcode sequence 2011, (ii) barcode sequence 2012 and (iii) a transposon end sequence 2013. Nucleic acid strand 2000iii is extended to generate nucleic acid strand 2010iii that comprises (i) a reverse complement of barcode sequence 2011, and (ii) transposon end sequence 2014.
[0336] The binding sequences 2015 and 2016 hybridize to capture sequences 2051 of barcode construct 2050. Adapter construct 2000 enables copying of barcode sequences 2011 and 2012 in the adapter construct and copying of sample barcode sequence 2052 and additional functional sequences of the barcode construct by complementary DNA (cDNA) synthesis. The copying of barcode sequences 2011 and 2012 and the copying of barcode sequence 2052 can occur simultaneously or sequentially in separate steps.
[0337] As shown in FIG. 20B, binding sequences 2015 and 2016 in a plurality of adaptor constructs hybridize to a plurality of capture sequences on the barcode construct and are extended to generate barcoded adapter constructs comprising the reverse complements of barcode sequence 2052 and additional functional sequences using th...
Claims
Atty Dkt No.: 68244-703601CLAIMSWHAT IS CLAIMED IS:
1. A composition comprising:(i) a nucleic acid strand comprising two identical copies of a transposon end recognition sequence;(ii) a first oligonucleotide comprising a reverse complement of the transposon end recognition sequence; and(iii) a second oligonucleotide comprising a reverse complement of the transposon end recognition sequence.
2. The composition of claim 1, wherein the transposon end recognition sequence comprises any one of SEQ ID NOs: 1-6.
3. The composition of claim 1 or 2, wherein the nucleic acid strand comprises (I) a first segment comprising a first barcode sequence and a first copy of the two identical copies of the transposon end recognition sequence, and (II) a second segment comprising a second barcode sequence and a second copy of the two identical copies of the transposon end recognition sequence.
4. The composition of claim 3, wherein the first barcode sequence and the second barcode sequence are identical sequences.
5. The composition of claim 3, wherein the first barcode sequence and the second barcode sequence are different sequences.
6. The composition of any one of claims 3-5, wherein the nucleic acid strand comprises, in order from 5’ to 3’ : the first barcode sequence, the first copy of the two identical copies of the transposon end recognition sequence, the second barcode sequence, and the second copy of the two identical copies of the transposon end recognition sequence.
7. The composition of claim 6, wherein the nucleic acid strand further comprises a restriction enzyme recognition sequence.
8. The composition of claim 7, wherein the first oligonucleotide further comprises a reverse complement of the restriction enzyme recognition sequence.
9. The composition of claim 7 or 8, wherein the nucleic acid strand further comprises a restriction enzyme cut site located adjacent and 3’ to the first copy of the two identical copies of the transposon end recognition sequence.
10. The composition of any one of claims 7-9, wherein the restriction enzyme recognition sequence is a PvuII restriction enzyme recognition sequence.Atty Dkt No.: 68244-70360111. The composition of any one of claims 7-10, wherein the nucleic acid strand further comprises an additional restriction enzyme recognition sequence, and wherein the second oligonucleotide further comprises a reverse complement of the additional restriction enzyme recognition sequence.
12. The composition of claim 11, wherein the nucleic acid strand further comprises an additional restriction enzyme cut site located adjacent and 3’ to the second copy of the two identical copies of the transposon end recognition sequence.
13. The composition of claim 12, wherein the additional restriction enzyme recognition sequence is a PvuII restriction enzyme recognition sequence.
14. The composition of claim 6, wherein the nucleic acid strand does not comprise a restriction enzyme recognition sequence.
15. The composition of any one of claims 3-14, further comprising one or more linker polynucleotides coupling the first segment of the nucleic acid strand to the second segment of the nucleic acid strand, wherein the one or more linker polynucleotides is distinct from the nucleic acid strand.
16. The composition of claim 15, wherein the one or more linker polynucleotides comprises a unimolecular linker polynucleotide comprising (i) a first capture sequence hybridized to a first binding site in the first segment of the nucleic acid strand; and (ii) a second capture sequence hybridized to a second binding site in the second segment of the nucleic acid strand.
17. The composition of claim 15, further comprising a plurality of linker polynucleotide molecules coupling the first segment of the nucleic acid strand to the second segment of the nucleic acid strand, wherein the plurality of linker polynucleotides comprises (i) a first linker polynucleotide molecule hybridized to a first binding site in the first segment of the nucleic acid strand; and (ii) a second linker polynucleotide molecule hybridized to a second binding site in the second segment of the nucleic acid strand.
18. The composition of any one of claims 15-17, wherein the one or more linker polynucleotides comprises a linker polynucleotide that comprises the first oligonucleotide.
19. The composition of any one of claims 1-18, wherein the nucleic acid strand further comprises one or more primer binding sequences.
20. The composition of any one of claims 1-19, wherein the nucleic acid strand further comprises primer binding sequences at its 5’ end and its 3’ end.
21. The composition of any one of claims 1-20, wherein the nucleic acid strand is coupled to a biotin.Atty Dkt No.: 68244-70360122. The composition of any one of claims 1-21, wherein the nucleic acid strand is further coupled to a support.
23. The composition of claim 22, wherein the support is a bead.
24. The composition of claim 22 or 23, wherein the nucleic acid strand is releasably coupled to the support via an enzymatically labile or photo-labile linker.
25. The composition of any one of claims 1-24, further comprising a transposase.
26. The composition of claim 25, wherein the transposase is a transposase dimer.
27. The composition of claim 25 or 26, wherein the transposase is a Tn5 transposase.
28. The composition of any one of claims 1-27, further comprising a restriction enzyme.
29. The composition of claim 28, wherein the restriction enzyme is PvuII.
30. A method, comprising generating, from the composition of any one of claims 3-29,(I) a first adapter construct comprising (A) a first double-stranded transposon end comprising the first copy of the transposon end recognition sequence, and (B) the first barcode sequence; and(II) a second adapter construct comprising (A) a second double-stranded transposon end comprising the second copy of the transposon end recognition sequence, and (B) the second barcode sequence.
31. The method of claim 30, further comprising generating a transpososome comprising (I) the first adapter construct, (II) the second adapter construct, and (III) a transposase dimer.
32. The method of claim 31, wherein generating the first adapter construct and the second adapter construct comprises cleaving the nucleic acid strand at a first site adjacent to and 3’ of the first copy of the transposon end recognition sequence.
33. The method of claim 32, further comprising performing the cleaving at the first site using the transposase dimer.
34. The method of claims 32, further comprising performing the cleaving at the first site using a restriction enzyme prior to generating the transpososome.
35. The method of claim 34, wherein the restriction enzyme is PvuII.
36. The method of any one of claims 32-35, wherein generating the first adapter construct and the second adapter construct further comprises cleaving the nucleic acid strand at a second site adjacent to and 3’ of the second copy of the transposon end recognition sequence.
37. The method of claim 36, further comprising performing the cleaving at the second site using the transposase dimer.
38. The method of claim 36, further comprising performing the cleaving at the second site using a restriction enzyme prior to generating the transpososome.Atty Dkt No.: 68244-70360139. The method of claim 38, wherein the restriction enzyme is PvuII.
40. A method comprising,(a) generating, from the composition of any one of claims 15-29,(I) a first adapter construct comprising (A) a first double-stranded transposon end comprising the first copy of the transposon end recognition sequence, and (B) the first barcode sequence; and(II) a second adapter construct comprising (A) a second double-stranded transposon end comprising the second copy of the transposon end recognition sequence, and (B) the second barcode sequence; and(b) removing the one or more linker polynucleotides from the first adapter construct and the second adapter construct.
41. The method of claim 40, further comprising performing (b) after generating a transpososome comprising (I) the first adapter construct, (II) the second adapter construct, and (III) a transposase dimer.
42. The method of any one of claims 30-41, wherein the nucleic acid strand of the composition is releasably coupled to a support, wherein the method further comprises:(i) hybridizing the first oligonucleotide and the second oligonucleotide to the nucleic acid strand to yield a hybridized complex;(ii) releasing the hybridized complex from the support; and(iii) after (ii) generating the first adapter construct and the second adapter construct from the hybridized complex.
43. The method of claim 42, further comprising, prior to (i), coupling the nucleic acid strand to the support.
44. The method of any one of claims 30-43, further comprising, prior to generating the first adapter construct and the second adapter construct, sequencing the nucleic acid strand or a copy or derivative thereof to obtain pairing information identifying the first barcode sequence as paired with the second barcode sequence.
45. The method of any one of claims 30-44, further comprising performing a transposition reaction using the first adapter construct, the second adapter construct, and a target nucleic acid molecule, thereby generating a plurality of transposition products comprising (I) a first transposition product comprising the first barcode sequence and a first target sequence of the target nucleic acid molecule or reverse complement thereof, and (II) a second transposition product comprising the second barcode sequence and a second target sequence of the target nucleic acid molecule or reverse complement thereof.Atty Dkt No.: 68244-70360146. The method of claim 45, further comprising sequencing the first transposition product or derivative thereof and the second transposition product or derivative thereof.
47. The method of claim 46, further comprising identifying the first target sequence and the second target sequence as being adjacent to each other on the target nucleic acid molecule based on the first barcode sequence and the second barcode sequence.
48. The method of claim 45 or claim 46, further comprising generating an assembled sequence comprising the first target sequence and the second target sequence using the first barcode sequence and the second barcode sequence.
49. The method of claim 48, further comprising: generating a plurality of unique pairs of adapter constructs having unique pairs of barcode sequences; performing transposition reactions using the plurality of unique pairs of adapter constructs and the target nucleic acid molecule or derivative thereof to generate transposition products comprising barcode sequences of the unique pairs of barcode sequences; and generating an assembled sequence from the transposition products using the barcode sequences in the transposition products based on pairing information of the unique pairs of barcode sequences.
50. A composition comprising an adapter pair comprising:(i) a first adapter construct comprising a first double-stranded transposon end comprising a first transposon end recognition sequence;(ii) a second adapter construct comprising a second double-stranded transposon end comprising a second transposon recognition sequence; wherein the first adapter construct and the second adapter construct are coupled to each other via one or more linker polynucleotides.
51. The composition of claim 50, wherein the one or more linker polynucleotides comprises a unimolecular linker polynucleotide comprising (i) a first capture sequence hybridized to a first binding site in the first adapter construct; and (ii) a second capture sequence hybridized to a second binding site in the second adapter construct.
52. The composition of claim 50, further comprising a plurality of linker polynucleotide molecules coupling the first segment of the nucleic acid strand to the second segment of the nucleic acid strand, wherein the plurality of linker polynucleotides comprises (i) a first linker polynucleotide molecule hybridized to a first binding site in the first adapter construct; and (ii) a second linker polynucleotide molecule hybridized to a second binding site in the second adapter construct.Atty Dkt No.: 68244-70360153. The composition of any one of claims 50-52, wherein the first transposon end recognition sequence and the second transposon end recognition sequence are identical sequences.
54. The composition of any one of claims 50-52, wherein the first transposon end recognition sequence and the second transposon end recognition sequence are different sequences.
55. The composition of any one of claims 50-54, wherein the first transposon end recognition sequence and the second transposon end recognition sequence each comprise any one of SEQ ID NOs: 1-6.
56. The composition of any one of claims 50-55, wherein the first adapter construct further comprises a first barcode sequence and wherein the second adapter construct further comprises a second barcode sequence.
57. The composition of claim 56, wherein the first barcode sequence is located 5’ of the first transposon end recognition sequence, and wherein the second barcode sequence is located 5’ of the second transposon end recognition sequence.
58. The composition of claim 56 or 57, wherein the first barcode sequence and the second barcode sequence are identical sequences.
59. The composition of claim 56 or 57, wherein the first barcode sequence and the second barcode sequence are different sequences.
60. The composition of any one of claims 56-59, wherein the composition is among one or more different adapter constructs that are not coupled to the first adapter construct or the second adapter construct.
61. The composition of 60, wherein the one or more different adapter constructs comprise one or more barcode sequences different from the first barcode sequence and the second barcode sequence.
62. The composition of any one of claims 56-61, wherein the first barcode sequence and the second barcode sequence together identify the first adapter construct and the second adapter construct as associated with each other.
63. The composition of any one of claims 50-62, wherein the first adapter construct and the second adapter construct each further comprises one or more primer binding sequences.
64. The composition of any one of claims 50-63, further comprising a transposase.
65. The composition of claim 64, wherein the adapter pair is bound to a transposase dimer.
66. The composition of claim 64 or 65, wherein the transposase is a Tn5 transposase.
67. A system comprising the composition of any one of claims 56-66, further comprising an additional adapter pair that is different from the adapter pair, wherein the additional adapter pair comprises a third adapter construct and a fourth adapter construct, whereinAtty Dkt No.: 68244-703601 the third adapter construct and the fourth adapter construct are coupled to each other via one or more additional linker polynucleotides.
68. The system of claim 67, wherein the third adapter construct comprises a third barcode sequence, wherein the fourth adapter construct comprises a fourth barcode sequence, wherein the third barcode sequence and the fourth barcode sequence together identify the third adapter construct and the fourth adapter construct as associated with each other.
69. The system of claim 68, further comprising at least 10 unique pairs of adapter constructs having unique pairs of barcode sequences.
70. A method, comprising:(a) providing a transpososome comprising the composition of any one of claims 57- 63 and a transposase dimer; and(b) removing the one or more linker polynucleotides from the first adapter construct and the second adapter construct in the transpososome.
71. The method of claim 70, further comprising generating the transpososome.
72. The method of claim 70 or 71, further comprising, after (b), performing a transposition reaction using the first adapter construct, the second adapter construct, and a target nucleic acid molecule, thereby generating a plurality of transposition products comprising (I) a first transposition product comprising the first barcode sequence and a first target sequence of the target nucleic acid molecule or reverse complement thereof, and (II) a second transposition product comprising the second barcode sequence and a second target sequence of the target nucleic acid molecule or reverse complement thereof.
73. The method of claim 72, further comprising sequencing the first transposition product or derivative thereof and the second transposition product or derivative thereof.
74. The method of claim 73, further comprising identifying the first target sequence and the second target sequence as being adjacent to each other on the target nucleic acid molecule based on the first barcode sequence and the second barcode sequence.
75. The method of claim 72 or claim 73, further comprising generating an assembled sequence comprising the first target sequence and the second target sequence using the first barcode sequence and the second barcode sequence.
76. The method of any one of claims 70-75, further comprising:(c) providing an additional transpososome comprising an additional adapter pair and a transposase dimer, wherein the additional adapter pair comprises a third adapter construct and a fourth adapter construct,Atty Dkt No.: 68244-703601 wherein the third adapter construct comprises a third barcode sequence and the fourth adapter construct comprises a fourth barcode sequence, and wherein the third adapter construct and the fourth adapter construct are coupled to each other via one or more additional linker polynucleotides; and(d) removing the one or more additional linker polynucleotides from the third adapter construct and the fourth adapter construct in the additional transpososome.
77. The method of claim 76, further comprising, after (d), performing an additional transposition reaction using the third adapter construct, the fourth adapter construct, and the target nucleic acid molecule or a derivative thereof, thereby generating an additional plurality of transposition products comprising (I) a third transposition product comprising the third barcode sequence and a third target sequence of the target nucleic acid molecule or reverse complement thereof, and (II) a fourth transposition product comprising the fourth barcode sequence and a fourth target sequence of the target nucleic acid molecule or reverse complement thereof.
78. The method of claim 77, further comprising sequencing the third transposition product or derivative thereof and the fourth transposition product or derivative thereof.
79. The method of claim 78, further comprising identifying the third target sequence and the fourth target sequence as being adjacent to each other on the target nucleic acid molecule based on the third barcode sequence and the fourth barcode sequence.
80. The method of claim 78 or claim 79, further comprising generating an assembled sequence comprising the first target sequence, the second target sequence, the third target sequence, and the fourth target sequence using the first barcode sequence, the second barcode sequence, the third barcode sequence, and the fourth barcode sequence.
81. The method of claim 80, further comprising: providing at least 100 to 100 billion unique pairs of adapter constructs having unique pairs of barcode sequences; performing transposition reactions using the at least 100 to 1000 unique pairs of adapter constructs and the target nucleic acid molecule or derivative thereof to generate a plurality of transposition products comprising barcode sequences of the unique pairs of barcode sequences; and generating an assembled sequence from the plurality of transposition products using the barcode sequences in the plurality of transposition products based on pairing information of the unique pairs of barcode sequences.
82. The method of claim 81, wherein the assembled sequence has a length of 500 to 10,000 nucleotides.Atty Dkt No.: 68244-70360183. The method of claim 81 or 82, wherein generating the assembled sequence comprises sequencing the plurality of transposition products or derivatives thereof to generate sequence reads, and assembling the sequence reads in silico to generate the assembled sequence.
84. The method of any one of claims 81-83, further comprising generating the at least 100 to 1000 unique pairs of adapter constructs using a plurality of nucleic acid strands each comprising a unique pair of the unique pairs of barcode sequences; and obtaining the pairing information of the unique pairs of barcode sequences by sequencing the plurality of nucleic acid strands or copies thereof.
85. A method comprising:(a) providing a composition comprising:(i) a nucleic acid strand comprising a first transposon end recognition sequence and a second transposon end recognition sequence;(ii) a first oligonucleotide that is hybridized to the first transposon end recognition sequence; and(iii) a second oligonucleotide that is hybridized to the second transposon end recognition sequence; wherein the first oligonucleotide and the second oligonucleotide are separate oligonucleotides; and(b) from the composition, generating:(I) a first adapter construct comprising a first double-stranded transposon end comprising the first transposon end recognition sequence; and(II) a second adapter construct comprising a second double-stranded transposon end comprising the second transposon end recognition sequence.
86. The method of claim 85, wherein the nucleic acid strand comprises (I) a first segment comprising a first barcode sequence and the first transposon end recognition sequence, and (II) a second segment comprising a second barcode sequence and the second transposon end recognition sequence.
87. The method of any one of claims 85 or 86, wherein the nucleic acid strand comprises, in order from 5’ to 3’ : the first barcode sequence, the transposon end recognition sequence, the second barcode sequence, and the second copy of the two identical copies of the transposon end recognition sequence.Atty Dkt No.: 68244-70360188. The method of any one of claims 85-87, further comprising generating a transpososome comprising (I) the first adapter construct, (II) the second adapter construct, and (III) a transposase dimer.
89. The method of any one of claims 85-88, wherein generating the first adapter construct and the second adapter construct comprises cleaving the nucleic acid strand at a first site adjacent to and 3’ of the first transposon end recognition sequence.
90. The method of claim 89, further comprising performing the cleaving at the first site using the transposase dimer.
91. The method of claims 89, further comprising performing the cleaving at the first site using a restriction enzyme prior to generating the transpososome.
92. The method of claim 91, wherein the restriction enzyme is PvuII.
93. The method of any one of claims 89-92, wherein generating the first adapter construct and the second adapter construct further comprises cleaving the nucleic acid strand at a second site adjacent to and 3’ of the second transposon end recognition sequence.
94. The method of claim 93, further comprising performing the cleaving at the second site using the transposase dimer.
95. The method of claim 93, further comprising performing the cleaving at the second site using a restriction enzyme prior to generating the transpososome.
96. The method of claim 95, wherein the restriction enzyme is PvuII.
97. The method of any one of claims 86-96, further comprising performing a transposition reaction using the first adapter construct, the second adapter construct, and a target nucleic acid molecule, thereby generating a plurality of transposition products comprising (I) a first transposition product comprising the first barcode sequence and a first target sequence of the target nucleic acid molecule or reverse complement thereof, and (II) a second transposition product comprising the second barcode sequence and a second target sequence of the target nucleic acid molecule or reverse complement thereof.
98. The method of claim 97, further comprising sequencing the first transposition product or derivative thereof and the second transposition product or derivative thereof.
99. The method of claim 98, further comprising identifying the first target sequence and the second target sequence as being adjacent to each other on the target nucleic acid molecule based on the first barcode sequence and the second barcode sequence.
100. The method of claim 98 or 99, further comprising generating an assembled sequence comprising the first target sequence and the second target sequence using the first barcode sequence and the second barcode sequence.
101. The method of any one of claims 97-100, further comprising:Atty Dkt No.: 68244-703601 generating a plurality of unique pairs of adapter constructs having unique pairs of barcode sequences; performing transposition reactions using the plurality of unique pairs of adapter constructs and the target nucleic acid molecule or derivative thereof to generate a plurality of transposition products comprising barcode sequences of the unique pairs of barcode sequences; and generating an assembled sequence from the plurality of transposition products using the barcode sequences in the plurality of transposition products based on pairing information of the unique pairs of barcode sequences.
102. The method of claim 101, wherein the assembled sequence has a length of 500 to 10,000 nucleotides.
103. The method of claim 101 or 102, wherein generating the assembled sequence comprises sequencing the plurality of transposition products or derivatives thereof to generate sequence reads, and assembling the sequence reads in silico to generate the assembled sequence.
104. The method of any one of claims 101-103, wherein generating the assembled sequence does not comprise further barcoding transposition products or derivatives thereof of the plurality of transposition products using additional barcode sequences that are different from the barcode sequences of the unique pairs of barcode sequences.
105. The method of any one of claims 101-104, wherein generating the assembled sequence does not comprise sequence assembly based on an overlap sequence of the target nucleic acid molecule that is common to transposition products of the plurality of transposition products.
106. The method of any one of claims 101-105, wherein generating the assembled sequence does not comprise reference-based assembly.
107. The method of any one of claims 101-106, wherein generating the assembled sequence does not comprise sequence alignment of transposition products or derivatives thereof of the plurality of transposition products.
108. The method of any one of claims 101-107, further comprising generating the plurality of unique pairs of adapter constructs using a plurality of nucleic acid strands each comprising a unique pair of the unique pairs of barcode sequences; and obtaining the pairing information of the unique pairs of barcode sequences by sequencing the plurality of nucleic acid strands or copies thereof.Atty Dkt No.: 68244-703601109. The method of any one of claims 86-108, wherein, in (b), the first adapter construct and the second adapter construct are coupled to each other via one or more linker polynucleotides.
110. The method of claim 109, wherein, in (a), the first segment of the nucleic acid strand and the second segment of the nucleic acid strand are coupled to each other via the one or more linker polynucleotides.
111. The method of claim 109 or 110, further comprising (c) removing the one or more linker polynucleotides from the first adapter construct and the second adapter construct.
112. The method of claim 111, wherein (c) occurs after generating a transpososome comprising (I) the first adapter construct, (II) the second adapter construct, and (III) a transposase dimer.
113. The method of claim 111 or 112, wherein (c) occurs prior to performing a transposition reaction using the first adapter construct, the second adapter construct, and the target nucleic acid molecule.
114. The method of any one of claims 85-113, wherein, prior to or during (a), the nucleic acid strand of the composition is releasably coupled to a support, wherein the method further comprises releasing the composition from the support.
115. The method of claim 44, further comprising performing the sequencing of the nucleic acid strand or a copy or derivative thereof using a primer comprising a locked nucleic acid (LNA) nucleotide.
116. The method of any one of claims 46, 73, or 98, further comprising performing the sequencing of the first transposition product or derivative thereof and the second transposition product or derivative thereof using a primer comprising a locked nucleic acid (LNA) nucleotide.
117. The method of claim 83 or 103, further comprising performing the sequencing of the transposition products or derivatives thereof using a primer comprising a locked nucleic acid (LNA) nucleotide.
118. The method of claim 84 or 108, further comprising performing the sequencing of the plurality of nucleic acid strands or copies thereof using a primer comprising a locked nucleic acid (LNA) nucleotide.
119. A library of nucleic acid strands, wherein each nucleic acid strand comprises two identical copies of a transposon end recognition sequence and a unique pair of barcode sequences.
120. The library of claim 119, wherein each nucleic acid strand comprises, in order from 5’ to 3’: a first barcode sequence of the unique pair of barcode sequences, a firstAtty Dkt No.: 68244-703601 copy of the two identical copies of the transposon end recognition sequence, a second barcode sequence of the unique pair of barcode sequences, and a second copy of the two identical copies of the transposon end recognition sequence.
121. The library of claim 119 or 120, wherein the library comprises a nucleic acid strand comprising a restriction enzyme recognition sequence.
122. The library of claim 121, wherein the nucleic acid strand comprises a restriction enzyme cut site located adjacent and 3’ to the first copy of the two identical copies of the transposon end recognition sequence.
123. The library of claim 122, wherein the restriction enzyme recognition sequence is a PvuII restriction enzyme recognition sequence.
124. The library of any one of claims 121-123, wherein the nucleic acid strand further comprises an additional restriction enzyme recognition sequence.
125. The library of claim 124, wherein the nucleic acid strand further comprises an additional restriction enzyme cut site located adjacent and 3’ to the second copy of the two identical copies of the transposon end recognition sequence.
126. The library of claim 125, wherein the additional restriction enzyme recognition sequence is a PvuII restriction enzyme recognition sequence.
127. A method, comprising:(a) generating a plurality of transposition products via transposase-mediated fragmentation of a target nucleic acid molecule, wherein the plurality of transposition products comprises a plurality of unique pairs of barcode sequences, wherein a unique pair of barcode sequences of the plurality of unique pairs of barcode sequences identifies a first transposition product and a second transposition product of the plurality of transposition products as derived from sequences adjacent to each other within the target nucleic acid molecule;(b) generating a plurality of sequence reads from the plurality of transposition products; and(c) generating an assembled sequence from the plurality of sequence reads using the plurality of unique pairs of barcode sequences.
128. The method of claim 127, wherein the method does not comprise further fragmentation of the plurality of transposition products by a non-transposase enzyme.
129. The method of any one of claims 127 or 128, wherein the method does not comprise further ligating together transposition products of the plurality of transposition products or derivatives thereof.Atty Dkt No.: 68244-703601130. The method of any one of claims 127-129, wherein the transposase-mediated fragmentation comprises performing a transposition reaction between the target nucleic acid molecule and a transposon composition, wherein the transposon composition comprises two double-stranded transposon ends but is not fully double-stranded.
131. The method of claim 130, wherein the transposon composition comprises (I) a first adapter construct comprising a first double-stranded transposon end comprising a first transposon end recognition sequence and a first barcode sequence of a unique pair of barcode sequences of the plurality of unique barcode sequences; and (II) a second adapter construct comprising a second double-stranded transposon end comprising a second transposon recognition sequence and a second barcode sequence of the unique pair of barcode sequences.
132. The method of claim 131, wherein the first adapter construct and the second adapter construct are assembled together with a same transposase dimer.
133. The method of claim 131 or 132, wherein the first adapter construct and the second adapter construct are coupled to each other via one or more linker polynucleotides.
134. The method of any one of claims 130-133, wherein the method comprises generating the transposon composition from a precursor composition comprising:(i) a nucleic acid strand comprising two identical copies of a transposon end recognition sequence;(ii) a first oligonucleotide comprising a reverse complement of the transposon end recognition sequence; and(iii) a second oligonucleotide comprising a reverse complement of the transposon end recognition sequence.
135. The method any one of claims 127-134, wherein each unique pair of barcodes sequences of the plurality of unique pairs of barcode sequences does not comprise complementary sequences.
136. A method, comprising generating an assembled sequence from a plurality of sequence reads using a plurality of unique pairs of barcode sequences, wherein each unique pair of barcode sequences of the plurality of unique pairs of barcode sequences identifies two sequence reads of the plurality of sequence reads as being paired, wherein each unique pair of barcodes sequences does not comprise identical or complementary barcode sequences.Atty Dkt No.: 68244-703601137. The method of any one of claims 127-136, wherein the plurality of unique pairs of barcode sequences comprises at least 1 million unique pairs of barcode sequences.
138. The method of any one of claims 127-137, wherein the method does not comprise sequence assembly based on an overlap sequence of the target nucleic acid molecule that is common to transposition products of the plurality of transposition products.
139. The method of any one of claims 127-138, wherein the method does not comprise reference-based assembly.
140. The method of any one of claims 127-139, wherein the method does not comprise sequence alignment of sequence reads of the plurality of sequence reads.
Citation Information
Patent Citations
Transposon end compositions and methods for modifying nucleic acids
WO2010048605A1
Linking sequence reads using paired code tags
WO2012061832A1
Methods and compositions for nucleic acid sequencing
WO2014142850A1
Methods and compositions for analyzing cellular components
WO2016130704A2
Modified transposons, compositions and uses thereof
WO2023225519A1