Dual-tagmentation single-cell dnaseq
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- BIO RAD LABORATORIES INC
- Filing Date
- 2024-07-25
- Publication Date
- 2026-06-03
AI Technical Summary
Current DNA sequencing technologies face limitations in sequencing long DNA fragments, as many platforms can only process short reads, requiring compilation of shorter sequences into contigs for longer sequence determination.
A dual-tagmentation method is employed, involving a first transposase with homoadaptor oligonucleotides to introduce breaks in genomic DNA and insert barcoded adapters, followed by a second transposase with heteroadaptor oligonucleotides to generate shorter fragments suitable for short-read sequencing while preserving long-read linkage information.
This approach allows for the generation of short-read sequencer-compatible DNA fragments while maintaining the long-range connectivity of the genome, facilitating accurate whole-genome assembly and variant analysis.
Smart Images

Figure US2024039617_30012025_PF_FP_ABST
Abstract
Description
DUAL-TAGMENTATION SINGLE-CELL DNASEQCROSS-REFERENCE TO RELATED PATENT APPLICATIONS
[0001] The present application claims benefit of priority to U.S. Provisional Patent Application No. 63 / 529,093, filed July 26. 2023. which is incorporated by reference for all purposes.BACKGROUND OF THE INVENTION
[0002] Next-generation sequencing allows for rapid and efficient nucleotide sequencing of DNA. However, for many sequencing platforms, including sequencing-by-synthesis sequencing methods, there is a limit to the length of DNA fragments that can be sequenced. If long sequences are to be determined, one must compile shorter sequences into contigs.
[0003] Chromosomal DNA is bound to a variety of proteins, including but not limited to histones. The presence or absence of chromosomal DNA-binding proteins can affect cell phenoty pes. In some cases, one can use Assay for Transposase Accessible Chromatin with high-throughput sequencing (ATAC-seq) to use transposases to selectively cleave in regions of chromosomal DNA not protected by chromosomal proteins, resulting in a variety of lengths of chromosomal fragments, many of which are too long for short-read sequencing. While longer read sequencing methods can be employed, those can be much more burdensome than short-read sequencing.BRIEF SUMMARY OF THE INVENTION
[0004] In some embodiments, methods of generating DNA sequencing reads are provided. In some embodiments, the method comprises, providing a solution comprising permeabilized cells or permeabilized isolated nuclei comprising genomic DNA; diffusing into the permeabilized cells or permeabilized isolated nuclei a first transposase. carrying homoadaptor oligonucleotides, that introduces breaks in the genomic DNA to form first DNA fragments and inserts homoadaptor oligonucleotides at the breaks, wherein thehomoadaptor oligonucleotides comprise a first strand comprising, 5’ to 3 ' an optional spacer sequence and a mosaic end (ME) sequence, and a second strand comprising an antisense ME sequence, wherein 3’ ends of the first strand of the adaptor oligonucleotides are covalently linked to 5’ ends of each strand of double- stranded DNA fragments to form double-stranded fragments having 5’ overhangs; linking cell-specific barcode sequences to 5 ' ends of the homoadaptor oligonucleotides on the double-stranded fragments to generate barcoded double-stranded fragments and forming a bulk solution of barcoded double-stranded fragments; in the bulk solution lysing the cells or nuclei and digesting protein in the solution; in the bulk solution extending 3’ ends of the barcoded double-stranded fragments using the 5’ overhangs as a template, thereby forming end-tagged first and second barcoded DNA fragment strands comprising. 5’ to 3’: the barcode sequence, the ME sequence and a DNA fragment sequence, wherein the DNA fragment sequences from the first and second strands are reverse complements; independently mutating the end-tagged first and second barcoded DNA fragment strands to generate (i) randomly mutated end-tagged first barcoded DNA fragments having a first pattern of introduced mutations and (ii) randomly mutated end-tagged second barcoded DNA fragments having a second pattern of introduced mutations; optionally further amplifying the randomly mutated end-tagged first barcoded DNA fragments and the randomly mutated end-tagged second barcoded DNA fragments to produce copies of the randomly mutated end-tagged first barcoded DNA fragments and the randomly mutated end-tagged second barcoded DNA fragments; contacting the randomly mutated end-tagged first barcoded DNA fragments and the randomly mutated end-tagged second barcoded DNA fragments with a second transposase. carrying heteroadaptor oligonucleotides, that introduces breaks in the DNA and inserts heteroadaptor oligonucleotides at the breaks to form heteroadaptor-linked DNA fragments, wherein the heteroadaptor oligonucleotides comprise a first strand comprising 5’ to 3‘ a heteroadaptor sequence and a ME sequence, and a second strand comprising an antisense ME sequence, wherein 3’ ends of the first strand of the adaptor oligonucleotides are covalently linked to 5’ ends of each strand of double-stranded DNA fragments to form new double-stranded fragments; andnucleotide sequencing the new double-stranded fragments to generate a plurality of sequencing reads.
[0005] In some embodiments, the providing comprises providing a solution comprising fixed and permeabilized cells. In some embodiments, the linking comprises synthesizing cellspecific barcodes on the 5’ ends of the homoadaptor oligonucleotides on the double-stranded fragments using split-pooling.
[0006] In some embodiments, the providing comprises providing a solution comprising permeabilized cells.
[0007] In some embodiments, the providing comprises providing a solution comprising permeabilized isolated nuclei. In some embodiments, the permeabilized cells or permeabilized nuclei are encapsulated in a hydrogel bead. In some embodiments, the linking comprises: introducing or forming partitions comprising (i) single permeabilized cells or single nuclei and (ii) a bead linked to 5’ ends of a plurality of clonal barcoding oligonucleotides, the barcoding oligonucleotides comprising a 5’ PCR handle sequence, a 3’ capture sequence and a barcode sequence unique to the bead to which the barcode oligonucleotide is linked; optionally releasing the clonal barcoding oligonucleotides; linking the clonal barcoding oligonucleotides to the 5’ ends of the homoadaptor oligonucleotides on the double-stranded fragments.
[0008] In some embodiments, the 5’ ends of the homoadaptor oligonucleotides are phosphorylated and the linking of the clonal barcoding oligonucleotides to the 5’ ends of the homoadaptor oligonucleotides on the double-stranded fragments comprises ligating the 5’ ends of the homoadaptor oligonucleotides to 3’ ends of the clonal barcoding oligonucleotides.
[0009] In some embodiments, the 5’ ends of the homoadaptor oligonucleotides comprises an alkyne moiety and 3’ ends of the clonal barcoding oligonucleotides comprise an azide moiety and the linking of the clonal barcoding oligonucleotides to the 5’ ends of the homoadaptor oligonucleotides on the double-stranded fragments comprises reacting the azide moiety with the alkyne moiety via a click chemistry reaction.
[0010] In some embodiments, the method further comprises assembling sequences for the first DNA fragments by assembling the sequencing reads based upon (i) the pattern of introduced mutations, (ii) sequence of the fragment, (iii) a location of transposase fragmentation and (iv) the sequence of the cell-specific barcode.
[0011] In some embodiments, the introducing mutations comprises no more than 2, 3, 4, 5, 6. 7, 8, 9, or 10 cycles of thermocycling.
[0012] In some embodiments, the introducing mutations while selectively amplifying comprises amplifying in the presence of a universal nucleotide.
[0013] In some embodiments, the introducing mutations while selectively amplifying comprises performing error-prone PCR.
[0014] In some embodiments, the contacting of the DNA with the first transposase generates fragments of greater than 500, 750 or 1000 bp on average.
[0015] In some embodiments, the partitions are droplets or hydrogel beads or microwells.
[0016] In some embodiments, the independently mutating comprises primer extension from a single primer that anneals to both of the end-tagged first and second barcoded DNA fragment strands to generate (i) the randomly mutated end-tagged first barcoded DNA fragments having the first pattern of introduced mutations and (ii) the randomly mutated end- tagged second barcoded DNA fragments having the second pattern of introduced mutations.BRIEF DESCRIPTION OF THE DRAWINGS
[0017] FIG. 1 depicts a bulk solution (meaning not within partitions) tagmentation of chromosomal DNA that is associated with chromosomal proteins in various locations. See step 1. Step 2 depicts tagmentation, which occurs more frequently between chromosomal binding proteins. Different sized DNA fragments are generated, each having the same adaptor-modified ends comprising the mosaic end (ME) sequence. Step 3 shows a blown-up cartoon of one fragment and shows the ME sequence and an optional spacer sequence linked to the 5’ ends of the fragments generated by the tagmentase. ”S" refers to ‘"sense strand" and ”AS refers to “antisense strand.” As explained herein, tagmentation can be achieved by introducing the transposase (tagmentase) carrying oligonucleotides into permeabilized cells or permeabilized isolated nuclei such that the tagmentase fragments the chromosomal DNA therein and introduces the oligonucleotides to the 5’ ends of the resulting fragments.
[0018] FIG. 2 depicts additional steps after those shown in FIG. 1. Step 4 shows linkage of a cell-specific barcode (CBC) to the resulting fragments such that fragments are barcoded with a barcode specific for the cell from which it was derived. In embodiments in the which the starting material is permeabilized cells, the cells can be separated into different partitionssuch partitions contain single (only one) cells and at least one bead linked to clonal copies of oligonucleotides having a bead-specific barcode, which is subsequently linked to the DNA fragments (not shown) by any one of several possible mechanisms (e.g., ligation, click chemistry, amplification). In embodiments in the which the starting material is fixed and permeabilized cells, the fixed cells can themselves act as partitions and split-pooling can be used to synthesize cell-specific barcodes on the DNA fragments (not shown). Step 5 shows an optional embodiment in which the cell-specific barcodes are linked via a chemical modification or bases on the template strand to block polymerization, which can be beneficial, though not necessary, later in the method. In step 6, Barcode-linked single-cell fragments (long frag) are combined into a common pool solution and optionally any protein associated with the DNA is degraded and / or removed. As shown in step 7, gap-filling (filling 5’ single stranded regions by extension of 3’ ends) can be performed with a polymerase so that each strand is flanked by a sense and antisense ME sequence. Optionally (as shown) a nucleotide or linkage occurs that prevents extension of the strands beyond the ME sequence (and optional spacer).
[0019] FIG. 3 depicts additional steps after those shown in FIG. 2. The products from the gap-filling form the molecules depicted in step 8, in which fragments, including longer fragments, generated from the initial tagmentation step have common adaptor / single-cell barcodes linked to both ends as shown. In step 9, barcoded sense and antisense strand are subjected to separate mutagenic PCR amplified to generate separate "signature" (or "landmark”) mutations in the different strands. Vertical lines in the amplified products are intended to indicate the location of mutations introduced by the mutagenic PCR step. In step 10, the products of the mutagenic PCR can be further amplified under non-mutagenic conditions to create multiple copies of the separate mutagenized strands.
[0020] FIG. 4 depicts additional steps after those shown in FIG. 3. In step 11 (2nd tagmentation), the amplified signature-marked and single-cell barcoded fragments are fragments with a tagmentase. The tagmentase in this step can have, for example two separate adaptor oligonucleotides (heteroadaptors). The second tagmentation generates shorter fragments that can be sequenced in short read sequencing methods. In step 12, sequence reads are generated and aligned based on based upon (i) the pattern of introduced mutations (signature), (ii) sequence of the fragment, (iii) a location of transposase fragmentation and (iv) the sequence of the cell-specific barcode. In step 13, from some or all of the criteria above, one can reconstruct the longer initial fragment from the initial tagmentation step.Because in some embodiments both strands are reconstructed, one can improve base calling accuracy.
[0021] FIG. 5 depicts exemplary methods of linking barcode oligonucleotides to the fragmented DNA using a ligase. In the embodiments depicted, a splint oligonucleotide is included in the reaction mixture that anneals to the barcoding oligonucleotide and the 5’ overhang of the fragmented DNA allowing the 3’ of the barcoding oligonucleotide and the 5’ overhang to be ligated.
[0022] FIG. 6 depicts exemplary methods of linking barcode oligonucleotides to the fragmented DNA using click chemistry.DEFINITIONS
[0023] Unless defined otherwise, all technical and scientific terms used herein generally have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Generally, the nomenclature used herein and the laboratory procedures in cell culture, molecular genetics, organic chemistry, and nucleic acid chemistry and hybridization described below are those well-known and commonly employed in the art. Standard techniques are used for nucleic acid and peptide synthesis. The techniques and procedures are generally performed according to conventional methods in the art and various general references (see generally, Sambrook et al. MOLECULAR CLONING: A LABORATORY MANUAL, 2d ed. (1989) Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y., which is incorporated herein by reference), which are provided throughout this document. The nomenclature used herein and the laboratory procedures in analytical chemistry, and organic synthetic described below are those well-known and commonly employed in the art.
[0024] The term "amplification reaction" refers to any in vitro means for multiplying the copies of a target sequence of nucleic acid in a linear or exponential manner. Such methods include but are not limited to polymerase chain reaction (PCR); DNA ligase chain reaction (see U.S. Pat. Nos. 4,683,195 and 4,683,202; PCR Protocols: A Guide to Methods and Applications (Innis et al., eds, 1990)) (LCR); QBeta RNA replicase and RNA transcriptionbased amplification reactions (e.g., amplification that involves T7, T3, or SP6 primed RNA polymerization), such as the transcription amplification system (TAS), nucleic acid sequence based amplification (NASBA). and self-sustained sequence replication (3SR); isothermalamplification reactions (e.g., single-primer isothermal amplification (SPIA)); as well as others known to those of skill in the art.
[0025] "Amplifying" refers to a step of submitting a solution to conditions sufficient to allow for amplification of a polynucleotide if all of the components of the reaction are intact. Components of an amplification reaction include, e.g, primers, a polynucleotide template, polymerase, nucleotides, and the like. The term "amplifying" typically refers to an "exponential" increase in target nucleic acid. However, "amplifying" as used herein can also refer to linear increases in the numbers of a select target sequence of nucleic acid, such as is obtained with cycle sequencing or linear amplification. In an exemplary embodiment, amplifying refers to PCR amplification using a first and a second amplification primer.
[0026] The term "amplification reaction mixture" refers to an aqueous solution comprising the various reagents used to amplify a target nucleic acid. These include enz mes, aqueous buffers, salts, amplification primers, target nucleic acid, and nucleoside triphosphates. Amplification reaction mixtures may also further include stabilizers and other additives to optimize efficiency and specificity. Depending upon the context, the mixture can be either a complete or incomplete amplification reaction mixture.
[0027] "Polymerase chain reaction" or "PCR" refers to a method whereby a specific segment or subsequence of a target double-stranded DNA, is amplified in a geometric progression. PCR is well known to those of skill in the art; see, e.g., U.S. Pat. Nos. 4,683,195 and 4,683,202; and PCR Protocols: A Guide to Methods and Applications, Innis et al., eds, 1990. Exemplary PCR reaction conditions typically comprise either two or three step cycles. Two step cycles have a denaturation step followed by a hybridization / elongation step. Three step cycles comprise a denaturation step followed by a hybridization step followed by a separate elongation step.
[0028] A "primer" refers to a polynucleotide sequence that hybridizes to a sequence on a target nucleic acid and serves as a point of initiation of nucleic acid synthesis. Primers can be of a variety of lengths and are often less than 50 nucleotides in length, for example 12-30 nucleotides, in length. The length and sequences of primers for use in PCR can be designed based on principles known to those of skill in the art, see, e.g., Innis et al., supra. Primers can be DNA, RNA, or a chimera of DNA and RNA portions. In some cases, primers can include one or more modified or non-natural nucleotide bases. In some cases, primers are labeled.
[0029] ‘‘Primer extension” refers to any method in which a primer is extended in a template-specific manner. Examples of primer extension include, for example, methods in which a primer hybridizes to a template nucleic acid and a polymerase extends the primer in a template-specific manner. In some embodiments, the template is DNA and the polymerase is a DNA polymerase. In some embodiments, the template is RNA and the polymerase is a reverse-transcriptase. Primer extension can also include, for example, template switching (see, e.g.. Zhu YY. Machleder EM, et al. (2001) Biotechniques. 30(4): 892-897; Ramskold D, Luo S, et al. (2012) Nat Biotechnol , 30(8):777-78, and nick polymerization (also referred to as nick translation), the latter involving nicking one strand of a nucleic acid duplex and using the nicked strand as a primer that is extended using the other strand as a template (see, e.g.. Leonard G. Davis Ph.D., et al, in Basic Methods in Molecular Biology, 1986).
[0030] A nucleic acid, or a portion thereof, ‘‘hybridizes” or “anneals” to another nucleic acid under conditions such that non-specific hybridization is minimal at a defined temperature in a physiological buffer (e.g., pH 6-9, 25-150 mM chloride salt or in a PCR reaction mixture). In some cases, a nucleic acid, or portion thereof, hybridizes to a conserved sequence shared among a group of target nucleic acids. In some cases, a primer, or portion thereof, can hybridize to a primer binding site if there are at least about 6, 8, 10, 12, 14, 16, or 18 contiguous complementary nucleotides, including “universal” nucleotides that are complementary to more than one nucleotide partner. Alternatively, a primer, or portion thereof, can hybridize to a primer binding site if there are fewer than 1 or 2 complementarity mismatches over at least about 12. 14, 16, or 18 contiguous complementary nucleotides. In some embodiments, the defined temperature at which specific hybridization occurs is room temperature. In some embodiments, the defined temperature at which specific hybridization occurs is higher than room temperature. In some embodiments, the defined temperature at which specific hybridization occurs is at least about 37, 40, 42, 45, 50, 55, 60, 65, 70, 75. or 80 °C. In some embodiments, the defined temperature at which specific hybridization occurs is 37, 40, 42, 45, 50, 55, 60, 65, 70, 75, or 80 °C.
[0031] A "template" refers to a polynucleotide sequence that comprises the polynucleotide to be amplified, flanked by or a pair of primer hybridization sites. Thus, a "target template" comprises the target polynucleotide sequence adjacent to at least one hybridization site for a primer. In some cases, a "target template" comprises the target polynucleotide sequence flanked by a hybridization site for a “forward” primer and a “reverse” primer.
[0032] As used herein, "nucleic acid" means DNA, RNA, single-stranded, double-stranded, or more highly aggregated hybridization motifs, and any chemical modifications thereof. Modifications include, but are not limited to, those providing chemical groups that incorporate additional charge, polarizability, hydrogen bonding, electrostatic interaction, points of attachment and functionality7to the nucleic acid ligand bases or to the nucleic acid ligand as a whole. Such modifications include, but are not limited to, peptide nucleic acids (PNAs), phosphodiester group modifications (e.g.. phosphorothioates, methylphosph onates). 2'-position sugar modifications, 5-position pyrimidine modifications, 8-position purine modifications, modifications at exocyclic amines, substitution of 4-thiouridine, substitution of 5-bromo or 5-iodo-uracil; backbone modifications, methylations, unusual base-pairing combinations such as the isobases, isocytidine and isoguanidine and the like. Nucleic acids can also include non-natural bases, such as, for example, nitroindole. Modifications can also include 3' and 5' modifications including but not limited to capping with a fluoroph ore (e.g., quantum dot) or another moiety7.
[0033] A "polymerase" refers to an enzyme that performs template-directed synthesis of polynucleotides, e.g.. DNA and / or RNA. The term encompasses both the full length polypeptide and a domain that has polymerase activity. DNA polymerases are well-known to those skilled in the art, including but not limited to DNA polymerases isolated or derived from Pyrococcus furiosus, Thermococcus litoralis, and Thermotoga maritime, or modified versions thereof. Additional examples of commercially available polymerase enzymes include, but are not limited to: Klenow fragment (New England Biolabs® Inc.). Taq DNA polymerase (QIAGEN), 9 °N™ DNA polymerase (New England Biolabs® Inc ), Deep Vent™ DNA polymerase (New England Biolabs® Inc.), Manta DNA polymerase (Enzymatics®), Bst DNA polymerase (New England Biolabs® Inc.), and phi29 DNA polymerase (New England Biolabs® Inc.).
[0034] Polymerases include both DNA-dependent polymerases and RNA-dependent polymerases such as reverse transcriptase. At least five families of DNA-dependent DNA polymerases are known, although most fall into families A, B and C. Other ty pes of DNA polymerases include phage polymerases. Similarly, RNA polymerases typically include eukaryotic RNA polymerases I, II. and III. and bacterial RNA polymerases as well as phage and viral polymerases. RNA polymerases can be DNA-dependent and RNA-dependent.
[0035] As used herein, the term "partitioning" or "partitioned" refers to separating a sample into a plurality of portions, or "partitions." Partitions are generally physical, such that a sample in one partition does not, or does not substantially, mix with a sample in an adjacent partition. Partitions can be solid or fluid. In some embodiments, a partition is a solid partition, e.g., a microchannel. In some embodiments, a partition is a fluid partition, e.g, a droplet. In some embodiments, a fluid partition (e.g., a droplet) is a mixture of immiscible fluids (e.g., water and oil). In some embodiments, a fluid partition (e.g., a droplet) is an aqueous droplet that is surrounded by an immiscible carrier fluid (e.g.. oil).
[0036] As used herein a “barcode” is a short nucleotide sequence (e.g., at least about 4, 6, 8, 10, 12, 14, 16, 18, 20 or more nucleotides long) that identifies a molecule to which it is conjugated. Barcodes can be used, e.g., to identify molecules in a partition. Such a partitionspecific barcode should be unique for that partition as compared to barcodes present in other partitions. For example, partitions containing target RNA from single-cells can subject to reverse transcription conditions using primers that contain a different partition-specific barcode sequence in each partition, thus incorporating a copy of a unique “cellular barcode” into the reverse transcribed nucleic acids of each partition. Thus, nucleic acid from each cell can be distinguished from nucleic acid of other cells due to the unique “cellular barcode.” In some cases, the cellular barcode is provided by a “bead barcode” that is present on oligonucleotides conjugated to a bead, wherein the bead barcode is shared by (e.g., identical or substantially identical amongst) all, or substantially all, of the oligonucleotides conjugated to that bead but is different from most or substantially all oligonucleotides conjugated to other beads. Thus, cellular and bead barcodes can be present in a partition, attached to a bead, or bound to cellular nucleic acid as multiple copies of the same barcode sequence. Cellular or bead barcodes of the same sequence can be identified as deriving from the same cell, partition, or bead. Such partition-specific, cellular, or bead barcodes can be generated using a variety of methods, which methods result in the barcode conjugated to or incorporated into a solid or hydrogel support (e.g., a solid bead or particle or hydrogel bead or particle). In some cases, the partition-specific, cellular, or bead barcode is generated using a split and mix (also referred to as split and pool) synthetic scheme as described herein. A partition-specific barcode can be a cellular barcode and / or a bead barcode (for example when associated with a cell or partition or both). Similarly, a cellular barcode can be a partition specific barcode (when provided in a partition) and / or a bead barcode (when delivered by abead). Additionally, a bead barcode can be a cellular barcode and / or a partition-specific barcode.
[0037] In other cases, barcodes uniquely identify the molecule to which it is conjugated and are referred to as a unique molecular identifier (UMI). The number of nucleotides of the UMI, which can be continuous, or discontinuous, will depend on the number of UMI sequences required. In some embodiments, the number of UMIs available are many times (e.g., 2X, 10X, 100X, etc.) higher than possible conjugation partners, thereby reducing the chance of rare duplicates being linked to different molecules. In some embodiments, pools of different UMIs are present in a partition and the composition of the pool acts as an identifiers for the partition, with some UMIs being in common with some other partitions but the total pool of UMIs being unique or substantially unique between partitions. UMI sequences can be generated for example as random sequences of a set length, and in some embodiments is identified by a flanking known sequence.
[0038] The length of the barcode sequence determines how many unique samples can be differentiated. For example, a 1 nucleotide barcode can differentiate 4, or fewer, different samples or molecules; a 4-nucleotide barcode can differentiate 44or 256 samples or less; a 6 nucleotide barcode can differentiate 4096 different samples or less; and an 8 nucleotide barcode can index 65,536 different samples or less. Additionally, barcodes can be attached to both strands either through barcoded primers for both first and second strand synthesis, through ligation, or in a tagmentation reaction.
[0039] Barcodes are typically synthesized and / or polymerized (e.g.. amplified) using processes that are inherently inexact. Thus, barcodes that are meant to be uniform e.g., a cellular, particle, or partition-specific barcode shared amongst all barcoded nucleic acid of a single partition, cell, or bead) can contain various N-l deletions or other mutations from the canonical barcode sequence. Thus, barcodes that are referred to as ‘‘identical” or “substantially identical” copies refer to barcodes that differ due to one or more errors in, e.g.. synthesis, polymerization, or purification errors, and thus contain various N-l deletions or other mutations from the canonical barcode sequence. Moreover, the random conjugation of barcode nucleotides during synthesis using e.g., a split and pool approach and / or an equal mixture of nucleotide precursor molecules as described herein, can lead to low probability events in which a barcode is not absolutely unique (e.g.. different from all other barcodes of a population or different from barcodes of a different partition, cell, or bead). However, suchminor variations from theoretically ideal barcodes do not interfere with the high-throughput sequencing analysis methods, compositions, and kits described herein. Therefore, as used herein, the term “unique’’ in the context of a particle, cellular, partition-specific, or molecular barcode encompasses various inadvertent N-l deletions and mutations from the ideal barcode sequence. In some cases, issues due to the inexact nature of barcode synthesis, polymerization, and / or amplification, are overcome by oversampling of possible barcode sequences as compared to the number of barcode sequences to be distinguished (e.g.. at least about 2-, 5-, 10-fold or more possible barcode sequences). For example, 10,000 cells can be analyzed using a cellular barcode having 9 barcode nucleotides, representing 262,144 possible barcode sequences. The use of barcode technology is well known in the art, see for example Katsuyuki Shiroguchi, et al. Proc Natl Acad Sci U S A., 2012 Jan 24;109(4): 1347- 52; and Smith, AM et al.. Nucleic Acids Research Can 11, (2010). Further methods and compositions for using barcode technology include those described in U.S. 2016 / 0060621.
[0040] A “transposase” or “tagmentase” means an enzyme that is capable of forming a functional complex with a transposon end-containing composition and catalyzing insertion or transposition of the transposon end-containing composition into the double-stranded target DNA with which it is incubated in an in vitro transposition reaction.
[0041] The term “transposon end” means a double-stranded DNA that exhibits the nucleotide sequences (the “transposon end sequences”) that are necessary to form the complex with the transposase that is functional in an in vitro transposition reaction. A transposon end forms a “complex” or a “synaptic complex” or a “transposome complex” or a “transposome composition” with a transposase or integrase that recognizes and binds to the transposon end, and which complex is capable of inserting or transposing the transposon end into target DNA with which it is incubated in an in vitro transposition reaction. A transposon end exhibits two complementary sequences consisting of a “transferred transposon end sequence” or “transferred strand” and a “non-transferred transposon end sequence,” or “non transferred strand” For example, one transposon end that forms a complex with a hyperactive Tn5 transposase (e.g., EZ-Tn5™ Transposase, EPICENTRE Biotechnologies, Madison, Wis., USA) that is active in an in vitro transposition reaction comprises a transferred strand that exhibits a “transferred transposon end sequence” as follows:5' AGATGTGTATAAGAGACAG 3' (SEQ ID NO: 1),and a non-transferred strand that exhibits a “non-transferred transposon end sequence” as follows:5' CTGTCTCTTATACACATCT 3' (SEQ ID NO: 2).
[0042] The 3'-end of a transferred strand is joined or transferred to target DNA in an in vitro transposition reaction. The non-transferred strand, which exhibits a transposon end sequence that is complementary to the transferred transposon end sequence, is not joined or transferred to the target DNA in an in vitro transposition reaction.
[0043] The term “solid support” refers to the surface of a bead, microtiter well or other surface that is useful for attaching a nucleic acid, such as an oligonucleotide or polynucleotide. The surface of the solid support can be treated to facilitate attachment of a nucleic acid, such as a single stranded nucleic acid.
[0044] The term “bead” refers to any solid support that can be in a partition, e.g., a small particle or other solid support. In some embodiments, the beads comprise an alginate matrix, i.e.. calcium alginate. In some embodiments, the beads comprise polyacrylamide. For example, in some embodiments, the beads incorporate barcode oligonucleotides into the gel matrix through an acrydite chemical modification attached to each oligonucleotide.Exemplary beads can include hydrogel beads. In some cases, the hydrogel is in sol form. In some cases, the hydrogel is in gel form. An exemplary hydrogel is an agarose hydrogel. Other hydrogels include, but are not limited to, those described in, e.g., U.S. Patent Nos. 4,438,258; 6,534,083; 8,008,476; 8,329,763; U.S. Patent Appl. Nos. 2002 / 0,009,591; 2013 / 0,022,569; 2013 / 0,034,592; and International Patent Publication Nos.WO / 1997 / 030092; and WO / 2001 / 049240.
[0045] It will be understood that any range of numerical values disclosed herein can include the endpoints of the range, and any values or subranges in between the endpoints. For example, the range 1 to 10 includes the endpoints 1 and 10, and any value between 1 and 10. The values typically include one significant digit.
[0046] The term “sample” refers to a biological composition, such as a cell, comprising a target nucleic acid.
[0047] The term “about” refers to the usual error range for the respective value that is known by a person of ordinary skill in the art for this technical field, for example, a range of± 10%, ± 5%, or ± 1% can encompass the recited value, even if the recited value is not modified by the term "about."
[0048] All ranges described herein can include the end point values of the range, and any sub-range of values included between the endpoints of the range, where the values include the first significant digit. For example, a range of 1 to 10 includes a range from 2 to 9, 3 to 8, 4 to 7, 5 to 6, 1 to 5, 2 to 5, 2 to 10, 3 to 10, and so on.DETAILED DESCRIPTION OF THE INVENTION
[0049] The inventors have discovered a methodology of using a dual -tagmentation process to generate short-read sequencer-compatible single-cell gDNA, useful for example for library generation, wherein the gDNA has preserved long-read linkage information for whole genome assembly and sequence / variant analysis. An advantage of the methodology described herein is that chromatin can remain intact at the initial stages, allowing for initial steps in situ in the cells without removal or with only partial removal of proteins that bind genomic DNA until barcoding. Moreover, the initial tagmentation step described herein is a single-adaptor (homo adaptor) tagmentation, thereby preserving the entire genome for analysis. The methods described herein can allow for. if desired, for example, include variant calling of (1) single nucleotide polymorphisms (SNPs), (2) insertions or deletions (indels), (3) copy number variation (CNV), (4) translocations, (5) inversions, and (6) homozygous / heterozygous determination. Nevertheless, the final nucleic acid product of the methodology is compatible with short-read sequencing.
[0007] The methods provided herein allow one to begin with fully or partially intact chromatin, for example, contained in permeabilized cells or nuclei, and diffusing in a transposase, carry ing homoadaptor oligonucleotides, that can fragment the genomic DNA, being more active in areas of the chromatin having fewer interfering chromatin proteins, and being less active in areas of heterochromatin. The resulting gDNA fragments can be longer than can be processed by short read sequencing, for example at least 500, 750 or 1000 base pairs in length on average.
[0007] The methods described herein can comprise permeabilized cells or permeabilized nuclei, which for example can be fixed or encapsulated in a hydrogel bead, for example such that the chromosomal DNA of the cells or nuclei are compartmentalized from each other in abulk mixture (e.g., by the structure of the fixed or encapsulated cells or nuclei) and the cells or nuclei can subsequently be introduced into separate partitions.
[0050] Thus, in some embodiments, permeabilized cells are provided. The cells can be permeabilized to allow for entry of reagents while the cells themselves remain substantially intact. Permeabilization can remove cellular membrane lipids to allow large molecules such as enzymes to enter the cell. In some embodiments, a detergent is used for permeabilization. Exemplary detergents can include, for example, Triton X-100 and NP-40 are used for permeabilization (for example, at 0. 1-0.5% (v / v, in PBS). In some embodiments, a steroidal saponin (or saraponin) is used to solubilize lipid, resulting in permeabilization. An exemplary saraponin is Digitonin. The appropriate permeabilization reagent can be selected to be compatible with the integrity of partition, if used.
[0051] In some embodiments, the permeabilized cells are fixed cells. For example, in some embodiments, the cells are formalin-fixed, paraffin-embedded (FFPE) samples. In embodiments where the cells are permeabilized and fixed, the cells themselves can act as partitions. In other embodiments, the permeabilized cells are not fixed. In these embodiments, the permeabilized cells can be provided encapsulated in a hydrogel bead, allowing for containment of the cell and its contents while allowing for diffusion of reagents into the cell. Examples of methods for hydrogel bead-encapsulation of cells are described in, for example, Utrech et al., Advanced Healthcare Materials, Volume 4, Issue 11, August 5, 2015, pages 1628-1633.
[0052] Any type of cells can be used according to the methods and compositions described herein. In some embodiments, the cells are mammalian, for example human cells. In some embodiments, the cells are from a biological sample. Biological samples can be obtained from any biological organism, e.g., an animal, plant, fungus, pathogen (e.g., bacteria or virus), or any other organism. In some embodiments, the biological sample is from an animal, e.g., a mammal (e.g., a human or anon-human primate, a cow, horse, pig, sheep, cat, dog, mouse, or rat), a bird (e.g., chicken), or a fish. A biological sample can be any tissue or bodily fluid obtained from the biological organism, e.g., blood, a blood fraction, or a blood product (e.g., serum, plasma, platelets, red blood cells, and the like), sputum or saliva, tissue (e.g., kidney, lung, liver, heart, brain, nervous tissue, thyroid, eye. skeletal muscle, cartilage, or bone tissue); cultured cells, e.g.. primary cultures, explants, and transformed cells, stem cells, or cells found in stool, urine, etc.
[0053] In some embodiments, isolated nuclei from cells are provided. Methods of forming isolated nuclei are known and can be used as desired. Exemplary methods of generating isolated nuclei include those described in U.S. Patent No. 8546134; Gaublomme, et al., Nature Communications volume 10, Article number: 2907 (2019).
[0054] A transposase, carrying oligonucleotides, that introduced breaks in DNA and introduces oligonucleotides into the break sites is introduced into the bulk solution comprising the above-described permeabilized cells or isolated nuclei such that the transpose accesses genomic DNA in the cells or nuclei. The action of some transposases is sometimes referred to as “tagmentation” and can involve introduction of different adaptor sequences on different sides of a DNA breakage point or the adaptor sequences added can be identical. Homoadaptor-loaded tagmentases are tagmentases that contain adaptors of only one sequence, which adaptor is added to both ends of a tagmentase-induced breakpoint in the genomic DNA. Heteroadaptor-loaded tagmentases are tagmentases that contain two different adaptors, such that a different adaptor sequence is added to the two DNA ends created by a tagmentase-induced breakpoint in the DNA. Adaptor loaded tagmentases are further described, e.g., in U.S. Patent Publication Nos: 2010 / 0120098; 2012 / 0301925; and 2015 / 0291942 and U.S. Patent Nos: 5,965,443; U.S. 6,437,109; 7,083,980; 9,005,935; and 9,238,671, the contents of each of which are hereby incorporated by reference in the entirety for all purposes.
[0055] A tagmentase is an enzy me that is capable of forming a functional complex with a transposon end-containing composition and catalyzing insertion or transposition of the transposon end-containing composition into the double-stranded target DNA with which it is incubated in an in vitro transposition reaction. Exemplary' transposases include but are not limited to modified Tn5 transposases that are hyperactive compared to wildtype Tn5, for example can have one or more mutations selected from E54K, M56A, or L372P. Wild-type Tn5 transposon is a composite transposon in which two near-identical insertion sequences (IS50L and IS50R) are flanking three antibiotic resistance genes (Reznikoff WS. Annu Rev Genet 42: 269-286 (2008)). Each IS50 contains two inverted 19-bp end sequences (ESs), an outside end (OE) and an inside end (IE). However, wild-ty pe ESs have a relatively low activity and were replaced in vitro by hyperactive mosaic end (ME) sequences. A complex of the transposase with the 19-bp ME is thus all that is necessary for transposition to occur, provided that the intervening DNA is long enough to bring two of these sequences close together to form an active Tn5 transposase homodimer (Reznikoff WS.. Mol Microbiol 47:1199-1206 (2003)). Transposition is a very infrequent event in vivo, and hyperactive mutants were historically derived by introducing three missense mutations in the 476 residues of the Tn5 protein (E54K, M56A, L372P), which is encoded by IS50R (Goryshin IY, Reznikoff WS. 1998. J Biol Chem 273: 7367-7374 (1998)). Transposition works through a ’‘cut-and- paste” mechanism, where the Tn5 excises itself from the donor DNA and inserts into a target sequence, creating a 9-bp duplication of the target (Schaller H. Cold Spring Harb Symp Quant Biol 43: 401-408 (1979); Rezmkoff WS.. A / w / 7tev Ge T 42: 269-286 (2008)). In current commercial solutions (Nextera™ DNA kits, Illumina), free synthetic ME adaptors are end-joined to the 5'-end of the target DNA by the transposase (tagmentase).
[0056] In this initial tagmentation, the transposase (e.g., tagmentase) inserts homoadaptor oligonucleotides at the breaks, wherein the homoadaptor oligonucleotides comprise a first strand comprising, 5’ to 3’ an optional spacer sequence (for example, which might be or include a priming adaptor sequence) and a mosaic end (ME) sequence (e.g., AGATGTGTATAAGAGACAG), and a second strand comprising an antisense (i.e., reverse complementary to the) ME sequence, wherein 3’ ends of the first strand of the adaptor oligonucleotides are covalently linked to 5?ends of each strand of double-stranded DNA fragments to form double-stranded fragments having 5’ overhangs. “Homoadaptor” in this context means that the transposase is loaded with two copies of the oligonucleotides such that an identical first strand of the adaptor oligonucleotides is linked to the 5’ ends of each strand of double-stranded DNA fragments. See for example FIG. 1. As noted above, the first strand of the homoadaptor oligonucleotides can optionally include a spacer sequence. If included, the spacer sequence can include, in some non-limiting embodiments, a PCR handle sequence allowing for later annealing to the PCR handle sequence or its reverse complement, for example allowing use of universal priming across the fragments at a later step.
[0057] As shown in FIG. 1, the initial tagmentation is performed under conditions such that tagmentation occurs preferentially in some locations of the chromatin compared to others. In some embodiments, the amount of tagmentase enzyme used is less in this initial tagmentation step compared to the second tagmentation step described below7. In some embodiments, the amount of tagmentase enzy me used results on average one in cleavage of the chromatin every75-10 kb, 10-20 kb, 20-40 kb. 40-60 kb, 10-50 kb, or 20-40 kb. The amount of tagmentase to achieve these results can be determined empirically. Typically this is interpreted as being more active in euchromatin sites and less active in heterochromatin sites where more chromosomal DNA-binding proteins make the DNA less accessible to the transposases. Thisis also observed for example in ATACseq methods. The result of tagmentation will be double-stranded chromosomal DNA fragments having 5' overhangs from the introduced first strand of the adaptor oligonucleotides. See FIG. 1. A diversity of DNA fragments will be generated. In some embodiments, the DNA fragments are greater than 500, 750 or 1000 bp on average.
[0058] Following generation of the DNA fragments in the permeabilized cells or isolated nuclei, cell-specific barcode sequences are linked to the 5'' ends of the homoadaptor oligonucleotides on the double-stranded fragments to generate barcoded double-stranded fragments. This barcoding reaction can occur in partitions or in a bulk solution (not partitions) depending on the method of barcoding used, and following barcoding a bulk solution of barcoded double-stranded fragments is formed. If the barcoding occurs in bulk (for example, using a split-pooling approach) then the product is in bulk. If the barcoding occurs in partitions, the contents of the partitions can subsequently be combined to form a bulk solution. Cell-specific barcode sequences can be linked to the 5’ overhangs as desired. In some embodiments, the cell-specific barcode sequence can be added by ligation or by primer extension. In some embodiments, cell-specific barcode sequences are ligated to the 5?overhang using a splint oligonucleotide that anneals both to a 3’ sequence of an oligonucleotide having the cell-specific barcode sequence and at least a portion of the 5’ overhang. An example of this is depicted in FIG. 5. As shown in FIG. 5, the splint oligonucleotide anneals to both the barcoding oligonucleotide 3’ end and the 5‘ end of the 5‘ overhang, allowing for a ligase to ligate the 3’ end and the 5’ end. Exemplary ligases can include, but are not limited to, T4 ligase. The aspects depicted in FIG. 5 correspond to item 4 in FIG. 2.
[0059] In some embodiments, the permeabilized cells or nuclei, which may or may not be fixed, can be partitioned such that individual cells or nuclei are the only cell or nuclei within a particular partition. Exemplary partitions can include but are not limited to droplets within an emulsion or microwells. In embodiments in which partitions are employed, a plurality of copies of barcoding oligonucleotides linked to a bead (e.g., a hydrogel bead) can be delivered to the partitions, wherein different partitions receive different beads and accompanying barcoding oligonucleotides. The bead can be attached to multiple copies of the same oligonucleotide, for example, at least about 10, 50, 100, 500, 1000, 5000, 10,000, 50,000, 100,000, 500,000, 1,000,000, 5,000,000, 10,000,000, 108, 109, 1010or more copies of the same or substantially identical oligonucleotide can be attached to one (e.g., the same) bead.The barcoding oligonucleotides will comprise at least a bead-specific barcode sequence and (i) a 3' end comprising an azide moiety allowing for click chemistry’ reactions as described herein or (ii) a capture sequence for annealing to a target sequence on the adaptor oligonucleotides. In some embodiments, the barcoding oligonucleotides further comprise a 5" PCR handle sequence, allowing for a common 5’ sequence between sequences comprising different cell-specific barcodes, allowing them all to be amplified with a universal primer that anneals to the PCR handle sequence. Optionally, once the bead is present in the partition, the oligonucleotides can be cleaved from the bead prior to linking the oligonucleotides to the DNA fragments.
[0060] Each oligonucleotide can be linked at its 5’ end or elsewhere on the oligonucleotide to the bead and in some embodiments can include a cleavable moiety to remove the oligonucleotides from the bead, e.g.. before the oligonucleotides are linked to the DNA fragments comprising the adaptor sequence. In some embodiments, the cleavable linker comprises a uridine incorporated site in a portion of a nucleotide sequence. A uridine incorporated site can be cleaved, for example, using a uracil glycosylase enzyme (e.g., a uracil N-glycosylase enzyme or uracil DNA glycosylase (UDG) enzyme). In some embodiments, the cleavable linker comprises a photocleavable nucleotide. Photocleavable nucleotides include, for example, photocleavable fluorescent nucleotides and photocleavable biotinylated nucleotides. See, e.g., Li et al., PNAS, 2003, 100:414-419; Luo et s .. Methods Enzymol, 2014. 549: 115-131. In some cases, the oligonucleotides are attached to bead through a disulfide linkage (e.g.. through a disulfide bond between a sulfide of the solid support and a sulfide covalently attached to the 5’ or 3’ end, or an intervening nucleic acid, of the oligonucleotide). In such cases, the oligonucleotide can be cleaved from the solid support by contacting the solid support with a reducing agent such as a thiol or phosphine reagent, including but not limited to a beta mercaptoethanol, dithiothreitol (DTT), or tris(2- carboxyethyl)phosphine (TCEP).
[0061] The oligonucleotides from the bead will include, for example, a bead-specific barcode such that the bead-specific barcode sequence on a first oligonucleotide can be used to distinguish it from a bead-specific barcode from a second oligonucleotide from a different bead. The 3’ end of the oligonucleotides can comprise the reverse complement of a (e.g., ■‘universal”) sequence from the adaptor sequences added to the fragments such that the oligonucleotides from the beads can be used as primers in a primer extension (e.g., PCR) reaction using one strand of the fragments having adaptor sequences at their ends as atemplate, or the 3’ end will have an appropriate click chemistry moiety (e.g., an azide moiety).
[0062] In some embodiments, the cell-specific barcode sequences can be linked to the 5’ overhangs with click chemistry. Click chemistry involves an azide alkyne Huisgen cycloaddition reaction, wherein in the embodiments described herein, a first oligonucleotide having an end alkyne moiety is reacted with a second oligonucleotide having an end azide moiety and promoting or catalyzing a reaction between the moieties to covalently-link the two oligonucleotides. In some embodiments, the azide-alkyne Huisgen cycloaddition is a 1,3-dipolar cycloaddition between an azide and a terminal or internal alky ne to give a 1,2,3- triazole. The reaction can in some embodiments be catalyzed by copper (Copper(I)-catalyzed azide-alkyne cycloaddition (CuAAC)) or promoted by a strained difluorooctyne (DIFO) (Strain-promoted azide-alkyne cycloaddition (SPAAC)) for example. See, e.g.. Chemical Reviews: Click Chemistry, June 23, 2021, Volume 121, Issue 12, Pages 6697-7248, including for example, Fantoni et al., “A Hitchhiker’s Guide to Click-Chemistry with Nucleic Acids” Chem. Rev. 2021, 121, 7122-7154. Thus in some embodiments, click chemistry ligation chemically attaches a 5' alkyne moiety of the 5’ overhang with a 3' azido moiety at the terminus of a cell-specific oligonucleotide. The click ligation can be proceeded in some embodiments, by an incubation at 45 degrees C, for example, for 2 hours in the presence of Vitamin C, Cu(II)-TBTA, MgSO4, THPTA, and DMSO. FIG. 6 depicts an example of linkage of the cell-specific barcode sequences can be linked to the 5’ overhangs with click chemi stry.
[0063] In these embodiments, the 5’ end of the oligonucleotide transferred strand (also referred to herein as the “first strand”) by the transposase to the DNA (which is fragmented in the process) comprises one of the reactive moieties for click chemistry' (i.e.. an azide or an alkyne moiety) and a "3 end of the barcoding oligonucleotide has the corresponding click chemistry’ moiety (i.e., an alkyne or an azide, respectively). In some embodiments, the 5’ end of the oligonucleotide transferred strand comprises an alkyne moiety and the ‘3 end of the barcoding oligonucleotide comprises an azide moiety. In some embodiments, an alkyne adaptor (e.g., subsequently loaded onto a transposase) is chemically synthesized to add 5' Hexynyl to the 5' phosphate of oligonucleotide. In some embodiments, an azide is added to a’ssDNA oligonucleotide by terminal transferase TdT through an azide-ddNTP, which as N3 (azide) placed at the original 3' OH position.
[0064] In yet other embodiments, the cell-specific barcode sequence can be added to the 5‘ ends of the DNA fragments using a split-pool method to synthesize cell-specific barcodes on the 5’ ends of the homoadaptor oligonucleotides on the double-stranded fragments. In embodiments in which the permeabilized and fix cells or nuclei are involved and thus the DNA fragments do not substantially diffuse from the cells or nuclei, mixtures of cells can be added to different aliquots, for example in a multi-well plate, and different oligonucleotide sequences can be added to the 5’ ends of the DNA fragments within the cells or nuclei. The aliquots can then be combined together in a mixture and aliquoted again to add a second oligonucleotide to the 5’ end. This split-pool approach can be repeated as desired such that a cell-specific barcode is synthesized on the 5‘ end each of the DNA fragments. Exemplary split-pool methods for adding cell-specific barcode sequences that can be used in the present methods have been described, for example in O’Huallochain et al., Communications Biology volume 3, Article number: 213 (2020).
[0065] In some embodiments, the adaptor oligonucleotides include one or more (e.g., 2, 3, 4, or more) uracils or other non-natural nucleotides or a carbon spacer between the ME sequence and the 5?spacer sequence. See. e.g., FIG. 2, step 5. Once the barcoding oligonucleotides are linked (including but not limited to synthesized) to the end of the DNA fragments, a uracil (or non-natural nucleotide or a carbon spacer)-sensitive polymerase can be used to amplify the barcoded DNA fragments. In some embodiments, a carbon spacer is between the ME and spacer sequences. Exemplary carbon spacers include but are not limited to 3C, 9C, 18C. where the number indicates the number of carbons. See, e.g., Wang et a!., Bioorganic & Medicinal Chemistry Letters, Volume 18, Issue 12, 15 June 2008, Pages 3597- 3602.
[0066] Exemplary7polymerases that stop or stall at uracils or non-standard linkers such as spacer carbon chains include but are not limited to archaeal DNA polymerase from Pyrococcus furiosus (Horvath et al.. Nucleic Acids Res. 2010 Nov; 38(21): e!96.) and Vent (Greagg et al., PNAS USA 96 (16) 9045-9050 (1999)). By gap-filling with such polymerases, extension stops at the uracil or non-natural nucleotide or carbon spacer. See, e.g., FIG. 2, step 7. In addition, the displacing activity of the polymerase will displace the antisense ME sequence, which is not covalently linked to other sequences in the reaction.
[0067] The resulting product will be a double-stranded fragment comprising the homodaptor sequence and a cell-specific barcode. An exemplary product is shown in FIG. 3,step 9, though it will be appreciated that the uracil or non-natural nucleotide is optional. In various embodiments, the product has blunt ends or a 5’ overhang as depicted. Because the initial tagmentation uses a homoadaptor, both strands have the same ends, i.e., 5’-3’: cellspecific barcode, mosaic end (ME) sequence, DNA strand sequence (sense or antisense depending on strand) and an antisense ME sequence. This will allow at a later step for copying of both strands using a single primer that comprises the ME sequence and that will anneal to the antisense ME sequence.
[0068] Once the DNA fragments have been linked to a cell-specific barcode, the contents of the cells can be combined into a bulk solution comprising a mixture of barcoded DNA fragments from different cells. See, e.g., FIG 2, step 6. In some embodiments, the barcoded DNA fragments are separated from cellular components such as protein or other cellular debris. In some embodiments, cell lysis is induced to release the contents of the cells, protein is stripped from the DNA and DNA can be purified for example using phase-, bead- and / or column-based DNA purification methods. For example, Solid-phase reversible immobilization (SPRI) can be used to purify the DNA from other cellular components.
[0069] The different strands of the tagmented DNA fragments can be independently mutated to generate “landmark7’ mutations on different strands. Subsequent sequencing reads generated from different strands will allow for better resolution of sequences because complementary strands will be sequenced for many DNA regions and where some strands are lost in the method, the complementary strand will nevertheless be available to provide sequence information, for example single nucleotide polymorphisms. Independent mutations of the different strands can be achieved as desired. In some embodiments, common sequences on both strands can be used to initiate primer extension from a single primer. In some embodiments, for example, one can take advantage of there being the same antisense ME sequence at the 3’ end of both strands (see, e.g.. FIG. 3. step 9) to initiate primer extension from a primer comprising the ME sequence, annealing the primer to the antisense ME sequence on both strands, and then initiating primer extension that results in mutations in the strand. Because extension of each strand is independent, mutations in one strand will not necessarily occur for the other strand, resulting in different “landmark” mutations on each strand. In some embodiments, mutations are introduced using error-prone PCR. Error prone PCR is a random mutagenesis technique for generating amino acid substitutions in proteins by introducing mutations into a template during PCR. In some embodiments, error prone PCR has been carried out with Taq polymerase, either by using Mn2+and unbalanced dNTPlevels, or mutants of Taq. In some embodiments, a mutant polymerase adapted to misincorporate a greater number of incorrect bases than compared to a wild-type archaeal DNA polymerase is employed. See, e.g., W02006 / 030174. In other embodiments, the amplification is performed in the presence of a universal nucleotide that at same rate is introduced into the amplicons. Universal nucleotides inclusion will result in subsequent replacement during amplification of any nucleotide, thereby introducing mutations. Exemplary universal nucleotides include those described in, e.g.. Lee et al., Biochim Biophys Acta. 2010 May; 1804(5): 1064-1080. The result of the mutagenesis will be a first set of mutations generated on a first strand and a second different set of mutations on the second strand. The first and second sets of mutations can be used to identify the products from each strand.
[0070] Depending on how much product comprising the sets of mutations have been generated by the mutagenesis, additional rounds of normal (non-error prone) amplification can be used to amplify the mutated products. Thus the amplification produces copies of the randomly mutated end-tagged first barcoded DNA fragments and the randomly mutated end- tagged second barcoded DNA fragments, as referenced elsewhere herein. See, e.g., FIG. 3, step 10 and FIG. 4, step 11.
[0071] Subsequently, a second transposase (tagmentation) step is performed on the two barcoded and mutated fragments. In this step the DNA fragments are no longer bound by chromatin-associated proteins and thus transposase cleavage of the DNA fragments is relatively uniformly-distributed in the fragments compared to the initial tagmentation. In this second transposase reaction, the transposase carries heteroadaptor sequences, meaning that the transposase carries two oligonucleotides having different sequences. The second transposase introduces breaks in the DNA and inserts heteroadaptor oligonucleotides at the breaks to form heteroadaptor-linked DNA fragments. See, e.g., FIG 4, step 11. In some embodiments, the heteroadaptor oligonucleotides comprise a first strand comprising 5’ to 3’ a heteroadaptor sequence and a ME sequence, and a second strand comprising an antisense ME sequence. The 3 ' ends of the first strand of the adaptor oligonucleotides are covalently linked by the transposase to 5’ ends of each strand of double-stranded DNA fragments to form new double-stranded fragments that in some embodiments have different ends due to introduction in the second tagmentation of different heteroadaptor sequences at the ends of the fragments. See, e.g., FIG. 4, step 12. The first strands of the heteroadaptors can include sequences used for next generation sequencing (NGS). For example, when Illumina-based sequencing isused the oligonucleotide primer can first strands of the heteroadaptors can include a 5’ P5 or P7 sequence, respectively. Thus the resulting products can be directly applied to nucleotide sequencing instruments. The average length of fragments generated by this second tagmentation step will be useful for NGS, for example in some embodiments, averaging between 100-1000 nucleotides in length.
[0072] Sequencing platforms can be selected as desired to generate sequencing reads. In some embodiments, Illumina™-supported sequencing methods are employed. See. e.g., U.S. Patent Nos 11,029,513; US 11,150,179; 11,308,640; and 11,473,067 and citations therein. Exemplary DNA sequencing techniques include fluorescence-based sequencing methodologies (See, e.g., Birren et al., Genome Analysis: Analyzing DNA, 1, Cold Spring Harbor, N.Y.; herein incorporated by reference in its entirety). In some embodiments, automated sequencing techniques understood in that art are utilized. In some embodiments, the present technology provides parallel sequencing of partitioned amplicons (PCT Publication No. WO 2006 / 0841,32, herein incorporated by reference in its entirety)- In some embodiments, DNA sequencing is achieved by parallel oligonucleotide extension (See, e.g., U.S. Pat. Nos. 5,750.341; and 6,306.597, both of which are herein incorporated by reference in their entireties). Additional examples of sequencing techniques include the Church polony technology (Mitra et al., 2003, Analytical Biochemistry 320, 55-65; Shendure et al., 2005 Science 309, 1728-1732; and U.S. Pat. Nos. 6,432,360; 6,485.944; 6,511,803; herein incorporated by reference in their entireties), the 454 picotiter pyrosequencing technology (Margulies et al.. 2005 Nature 437, 376-380; U.S. Publication No. 2005 / 0130173; herein incorporated by reference in their entireties), the Solexa single base addition technology (Bennett et al., 2005, Pharmacogenomics, 6, 373-382; U.S. Pat. Nos. 6,787,308; and 6,833,246; herein incorporated by reference in their entireties), the Lynx massively parallel signature sequencing technology (Brenner et al. (2000). Nat. Biotechnol. 18:630-634; U.S. Pat. Nos. 5,695,934; 5,714,330; herein incorporated by reference in their entireties), and the Adessi PCR colony technology7(Adessi et al. (2000). Nucleic Acid Res. 28, E87; WO 2000 / 018957; herein incorporated by reference in its entirety).
[0073] Once generated, sequencing reads can be grouped with reference to the assaygenerated mutation landmarks and the location on genome. In some embodiments, the sequences can be aligned to assemble a "contig" consensus sequence based upon the aligned sequences. This is depicted, for example, as items 12-13 of FIG. 4. Assembling sequences for the first DNA fragments from the first tagmentation can be achieved, for example, byassembling the sequencing reads based upon (i) the pattern of introduced mutations as discussed above, (ii) the sequence of the fragment, and (iii) the location of transposase fragmentation and (iv) the sequence of the cell-specific barcode. The resulting consensus contig sequence represents the sequencing of the original fragments generated in the first tagmentation reactions, allowing one to generate the equivalent of a long sequencing read based on short read sequencing techniques.
[0074] Reaction mixtures comprising the reagents of any of the steps of the methods described herein are also provided in this disclosure. For example, a reaction mixture is provided comprising a solution comprising (i) permeabilized cells or permeabilized isolated nuclei comprising genomic DNA as described herein and (ii) a first transposase, carrying homoadaptor oligonucleotides, that introduces breaks in the genomic DNA to form first DNA fragments and inserts homoadaptor oligonucleotides at the breaks, wherein the homoadaptor oligonucleotides comprise a first strand comprising, 5' to 3' an optional spacer sequence and a mosaic end (ME) sequence, and a second strand comprising an antisense ME sequence, wherein 3’ ends of the first strand of the adaptor oligonucleotides are covalently linked to 5’ ends of each strand of double-stranded DNA fragments to form double-stranded fragments having 5’ overhangs.
[0075] Also provided, for example, a reaction mixture (from the bulk solution as described herein) comprising (i) randomly mutated end-tagged first barcoded DNA fragments having a first pattern of introduced mutations; and (ii) randomly mutated end-tagged second barcoded DNA fragments, having a second pattern of introduced mutations as described herein.
[0076] Kits providing one or more reagent for performing the described methods are also provided. For example, in some embodiments, the kit comprises the first and second tagmentase as described herein. For example, the first tagmentase comprises homoadaptor oligonucleotides that comprise a first strand comprising. 5’ to 3’ an optional spacer sequence and a mosaic end (ME) sequence, and a second strand comprising an antisense ME sequence and the second tagmentase comprises heteroadaptor oligonucleotides comprising a first strand comprising 5’ to 3’ a heteroadaptor sequence and a ME sequence, and a second strand comprising an antisense ME sequence.
[0077] Although the foregoing disclosure has been described in some detail by way of illustration and example for purposes of clarity of understanding, one of skill in the art will appreciate that certain changes and modifications may be practiced within the scope of theappended claims. In addition, each reference provided herein, including patents, patent applications, non-patent literature, and Genbank accession numbers, is incorporated by reference in its entirety to the same extent as if each reference was individually incorporated by reference. Where a conflict exists between the instant application and a reference provided herein, the instant application shall dominate.
Claims
WHAT IS CLAIMED IS:
1. A method of generating DNA sequencing reads, the method comprising, providing a solution comprising permeabilized cells or permeabilized isolated nuclei comprising genomic DNA; diffusing into the permeabilized cells or permeabilized isolated nuclei a first transposase, carrying homoadaptor oligonucleotides, that introduces breaks in the genomic DNA to form first DNA fragments and inserts homoadaptor oligonucleotides at the breaks, wherein the homoadaptor oligonucleotides comprise a first strand comprising, 5’ to 3’ an optional spacer sequence and a mosaic end (ME) sequence, and a second strand comprising an antisense ME sequence, wherein 3’ ends of the first strand of the adaptor oligonucleotides are covalently linked to 5’ ends of each strand of double-stranded DNA fragments to form double-stranded fragments having 5‘ overhangs; linking cell-specific barcode sequences to 5’ ends of the homoadaptor oligonucleotides on the double-stranded fragments to generate barcoded double-stranded fragments and forming a bulk solution of barcoded double-stranded fragments; in the bulk solution lysing the cells or nuclei and digesting protein in the solution; in the bulk solution extending 3’ ends of the barcoded double-stranded fragments using the 5’ overhangs as a template, thereby forming end-tagged first and second barcoded DNA fragment strands comprising, 5’ to 3’: the barcode sequence, the ME sequence, a DNA fragment sequence and an antisense ME sequence, wherein the DNA fragment sequences from the first and second strands are reverse complements; independently mutating the end-tagged first and second barcoded DNA fragment strands to generate (i) randomly mutated end-tagged first barcoded DNA fragments having a first pattern of introduced mutations and (ii) randomly mutated end-tagged second barcoded DNA fragments having a second pattern of introduced mutations; optionally further amplifying the randomly mutated end-tagged first barcoded DNA fragments and the randomly mutated end-tagged second barcoded DNA fragments to produce copies of the randomly mutated end-tagged first barcoded DNA fragments and the randomly mutated end-tagged second barcoded DNA fragments; contacting the randomly mutated end-tagged first barcoded DNA fragments and the randomly mutated end-tagged second barcoded DNA fragments with a secondtransposase, carry ing heteroadaptor oligonucleotides, that introduces breaks in the DNA and inserts heteroadaptor oligonucleotides at the breaks to form heteroadaptor-linked DNA fragments, wherein the heteroadaptor oligonucleotides comprise a first strand comprising 5’ to 3’ a heteroadaptor sequence and a ME sequence, and a second strand comprising an antisense ME sequence, wherein 3’ ends of the first strand of the adaptor oligonucleotides are covalently linked to 5’ ends of each strand of double-stranded DNA fragments to form new double-stranded fragments; and nucleotide sequencing the new double-stranded fragments to generate a plurality' of sequencing reads.
2. The method of claim 1, wherein the providing comprises providing a solution comprising fixed and permeabilized cells.
3. The method of claim 2, wherein the linking comprises synthesizing cell-specific barcodes on the 5’ ends of the homoadaptor oligonucleotides on the doublestranded fragments using split-pooling.
4. The method of claim 1, wherein the providing comprises providing a solution comprising permeabilized cells.
5. The method of claim 1, wherein the providing comprises providing a solution comprising permeabilized isolated nuclei.
6. The method of claim 1 , wherein the permeabilized cells or permeabilized nuclei are encapsulated in a hydrogel bead.
7. The method of claim 4 or 5. wherein the linking comprises: introducing or forming partitions comprising (i) single permeabilized cells or single nuclei and (ii) a bead linked to 5’ ends of a plurality' of clonal barcoding oligonucleotides, the barcoding oligonucleotides comprising a 5’ PCR handle sequence, a 3’ capture sequence and a barcode sequence unique to the bead to which the barcode oligonucleotide is linked; optionally releasing the clonal barcoding oligonucleotides; linking the clonal barcoding oligonucleotides to the 5’ ends of the homoadaptor oligonucleotides on the double-stranded fragments.
8. The method of claim 7, wherein the 5?ends of the homoadaptor oligonucleotides are phosphorylated and the linking of the clonal barcoding oligonucleotides to the 5 ’ ends of the homoadaptor oligonucleotides on the double-stranded fragments comprises ligating the 5’ ends of the homoadaptor oligonucleotides to 3’ ends of the clonal barcoding oligonucleotides.
9. The method of claim 7, wherein the 5’ ends of the homoadaptor oligonucleotides comprises an alkyne moiety and 3’ ends of the clonal barcoding oligonucleotides comprise an azide moiety and the linking of the clonal barcoding oligonucleotides to the 5’ ends of the homoadaptor oligonucleotides on the double-stranded fragments comprises reacting the azide moiety with the alkyne moiety via a click chemistry reaction.
10. The method of claim 1. further comprising assembling sequences for the first DNA fragments by assembling the sequencing reads based upon (i) the pattern of introduced mutations, (ii) sequence of the fragment, (iii) a location of transposase fragmentation and (iv) the sequence of the cell-specific barcode.1 1 . The method of claim 1 , wherein the introducing mutations comprises no more than 2, 3, 4, 5, 6, 7, 8, 9, or 10 cycles of thermocycling.
12. The method of claim 1, wherein the introducing mutations while selectively amplifying comprises amplifying in the presence of a universal nucleotide.
13. The method of claim 1, wherein the introducing mutations while selectively amplifying comprises performing error-prone PCR.
14. The method of claim 1, wherein the contacting of the DNA with the first transposase generates fragments of greater than 500, 750 or 1000 bp on average.
15. The method of claim 7, wherein the partitions are droplets or hydrogel beads or microwells.
16. The method of any one of claims 1-15, wherein the independently mutating comprises primer extension from a single primer that anneals to both of the end- tagged first and second barcoded DNA fragment strands to generate (i) the randomly mutatedend-tagged first barcoded DNA fragments having the first pattern of introduced mutations and (ii) the randomly mutated end-tagged second barcoded DNA fragments having the second pattern of introduced mutations.