Single-cell sequencing using multiple transposase adapters
Patent Information
- Application Number
- PCT/US2025/021282
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-27
- Filing Date
- 2025-03-25
- Publication Date
- 2025-10-30
AI Technical Summary
Next-generation sequencing (NGS) technologies struggle to accurately resolve complex transcript structures due to short-read limitations, leading to inaccurate isoform identification and difficulty in distinguishing between closely related isoforms, particularly in single-cell isoform RNA-seq, which is essential for understanding gene regulation and disease mechanisms.
A method involving transposases with different adaptor oligonucleotides introduces breaks in RNA/first strand cDNA hybrids, forming double-stranded double-barcoded cDNA fragments, which are then amplified and sequenced, allowing for accurate alignment and assembly of long-range transcript information.
This approach enhances the ability to resolve complex transcript structures and accurately identify isoforms, providing insights into gene regulation and disease mechanisms by improving sequencing accuracy and resolution.
Smart Images

Figure US2025021282_30102025_PF_FP_ABST
Abstract
Description
SINGLE-CELL SEQUENCING USING MULTIPLE TRANSPOSASE ADAPTERSCROSS-REFERENCE TO RELATED PATENT APPLICATIONS
[0001] The present application claims benefit of priority to U.S. Provisional Patent Application No. 63 / 570,618, filed March 27, 2025, which is incorporated by reference for all purposes.BACKGROUND OF THE INVENTION
[0002] Next-generation sequencing (NGS) allows for rapid and efficient nucleotide sequencing of DNA, including for example cDNAs. However, for many sequencing platforms, including sequencing-by-synthesis sequencing methods, there is a limit to the length of DNA fragments that can be sequenced. If long sequences are to be determined, one must compile shorter sequences into contigs.
[0003] However, single-cell isoform RNA-seq cannot be fulfilled by NGS (which is shortread sequencing) technology, due to the nature of isoform analysis, which requires long-range information of a transcript in order to differentiate various transcript variants. Nearly all human multi-exon genes are alternatively spliced allowing a single gene to generate multiple RNA isoforms that give rise to different protein isoforms which consequently drive phenotypic complexity in eukaryotes. Detecting these isoforms helps unravel the intricate mechanisms behind gene regulation, offering insights into how cells finely tune gene expression to respond to diverse physiological conditions.
[0004] Moreover, aberrant splicing or expression of specific transcript isoforms is also known to be associated with numerous diseases, including cancer, neurodegenerative disorders, and genetic syndromes. Furthermore, transcript isoforms can create phenotypic variations within a population by modulating gene function.
[0005] As a result, identifying and characterizing these isoforms aids in understanding disease mechanisms, and the genetic basis of phenotypic diversity. However, using nextgeneration sequencing (NGS) to determine isoforms presents several challenges.
[0006] Short-read NGS platforms typically generate short sequencing reads, which can be insufficient for accurately resolving complex transcript structures including genes with paralogous sequence, long introns and extensive alternative splicing events spanning multiple exons. The ambiguity' of mapping short sequencing reads to multiple locations within the genome leading to inaccurate isoform identification during transcript assembly, which makes it difficult to confidently distinguish between closely related isoforms of the full-length transcripts.
[0007] Tagmentation is often used in sequencing to generate fragments of target DNA short enough for short reads in NGS. Some workflows of NGS use tagmentases loaded with either a first or a second oligonucleotide adaptor and allowing for subsequent amplification using a first and second primer. However, inherent in these methods is that a high number of fragments generated have the first oligonucleotide adaptor on both ends or the second oligonucleotide adaptor on both ends, and thus are lost in the methods which require different adaptors on different ends of the tagmentation products.BRIEF SUMMARY OF THE INVENTION
[0008] In some embodiments, methods of generating double-stranded double-barcoded cDNA fragments are provided.
[0009] In some embodiments, the method comprises, providing fixed and permeabilized cells; diffusing reverse transcriptase and nucleotides and a primer into the fixed and permeabilized cells and performing reverse transcription of RNA in the cells to form RNA / first strand cDNA hybrids; diffusing into the fixed and permeabilized cells a set of at least three (e.g., 3-20, at least 3, 4, 5, 6, 7, 8, 9, 10, etc.) transposases, each of the at least three transposases carrying different adaptor oligonucleotides, that introduce breaks in the RNA / first strand cDNA hybrids to form RNA / first strand cDNA hybrid fragments and inserts homoadaptor oligonucleotides at the breaks, wherein the adaptor oligonucleotides comprise a first adaptor strand comprising, 5' to 3’:an Rn sequence, wherein Rn sequences differ for the adaptor oligonucleotides of each of the at least three transposases, optionally, one or more barcode sequences and a mosaic end (ME) sequence, and a second adaptor strand comprising an antisense ME sequence, wherein one of the transposases covalently links 3’ ends of the first adaptor strand of the adaptor oligonucleotides to 5’ ends of each strand of the RNA / first strand cDNA hybrid fragments to form RNA / first strand cDNA hybrid fragments having 5’ overhangs.
[0010] In some embodiments, the method further comprises displacing the second adaptor strand and (i) extending 3’ ends of the first strand cDNAs with a polymerase using the RNA and first adaptor strand linked thereto as a template to form first strand cDNA fragments linked to a 5’ first Rn sequence (Rl) and a 3’ antisense second Rn sequence (R2) or (ii) extending 3’ ends of the RNAs with a polymerase that initiates from RNA using the first strand cDNAs and first adaptor strand linked thereto as a template to form RNA fragments linked to a 5’ first Rn sequence (Rl) and a 3’ antisense second Rn sequence (R2) and then performing a second reverse transcription to form the first strand cDNA fragments linked to a 5‘ first Rn sequence (Rl) and a 3’ antisense second Rn sequence (R2);
[0011] In some embodiments, the method further comprises forming partitions comprising (i) single permeabihzed cells comprising the first strand cDNA fragments and (ii) a bead linked to a plurality of clonal barcoding oligonucleotides having free 3’ ends, the barcoding oligonucleotides comprising a first 5’ PCR handle sequence, a 3’ capture sequence and a barcode sequence unique to the bead to which the barcode oligonucleotide is linked, wherein the plurality of clonal barcoding oligonucleotides have different 3?capture sequences such that the plurality includes all of the different Rn sequences on the transposases; optionally releasing the clonal barcoding oligonucleotides in the partitions; in the partitions, linking a first of the clonal barcoding oligonucleotide 3’ capture sequences to the second Rn sequence (or complement thereol) of the first strand cDNA fragments to form an adaptor-linked intermediate second strand cDNAs that comprise 5^-3 the first 5’ PCR handle sequence, the barcode sequence unique to the bead, the second Rn sequence, optionally the one or more optional barcode sequences, the ME sequence, a second strand cDNA fragment, and an antisense of the ME sequence, an antisense of the first 5’ PCRhandle sequence, optionally an antisense of the one or more optional barcode sequences, an antisense of the first Rn sequence.
[0012] In some embodiments, the method further comprises, in the partitions, annealing a second of the clonal barcoding oligonucleotide 3’ capture sequences to the antisense of the first Rn sequence on the adaptor-linked intermediate second strand cDNAs and extending the second of the clonal barcoding oligonucleotides using the adaptor-linked intermediate second strand cDNAs as a template with the DNA-dependent DNA polymerase to form doublebarcoded first strand cDNA fragments that comprise 5 ’-3’: the first 5’ PCR handle sequence, the barcode sequence unique to the bead, the first Rn sequence, optionally the one or more optional barcode sequences, the ME sequence, the first strand cDNA fragment, and an antisense of the ME sequence, an antisense of the first 5’ PCR handle sequence, optionally an antisense of the one or more optional barcode sequences, and the antisense of the second Rn sequence;
[0013] In some embodiments, the method further comprises extending the adaptor-linked intermediate second strand cDNAs using the double-barcoded first strand cDNAs as a template to form double-stranded double-barcoded cDNA fragments.
[0014] In some embodiments, the method comprises, providing fixed and permeabilized cells; diffusing reverse transcriptase and nucleotides and a pnmer into the fixed and permeabilized cells and performing reverse transcription of RNA in the cells to form RNA / first strand cDNA hybrids; diffusing into the fixed and permeabilized cells a set of at least three (e.g., 3-20, at least 3, 4, 5, 6, 7, 8, 9, 10. etc.) transposases, each of the at least three transposases carrying different adaptor oligonucleotides, that introduce breaks in the RNA / first strand cDNA hybrids to form RNA / first strand cDNA hybrid fragments and inserts homoadaptor oligonucleotides at the breaks, wherein the adaptor oligonucleotides comprise a first adaptor strand comprising, 5’ to 3': an Rn sequence, wherein Rn sequences differ for the adaptor oligonucleotides of each of the at least three transposases, optionally, one or more barcode sequences anda mosaic end (ME) sequence, and a second adaptor strand comprising an antisense ME sequence, wherein one of the transposases covalently links 3?ends of the first adaptor strand of the adaptor oligonucleotides to 5’ ends of each strand of the RNA / first strand cDNA hybrid fragments to form RNA / first strand cDNA hybrid fragments having 5’ overhangs; displacing the second adaptor strand and (i) extending 3 ’ ends of the first strand cDNAs with a polymerase using the RNA and first adaptor strand linked thereto as a template to form first strand cDNA fragments linked to a 5?first Rn sequence (Rl) and a 3’ antisense second Rn sequence (R2) or (ii) extending 3’ ends of the RNAs with a polymerase that initiates from RNA using the first strand cDNAs and first adaptor strand linked thereto as a template to form RNA fragments linked to a 5’ first Rn sequence (Rl) and a 3’ antisense second Rn sequence (R2) and then performing a second reverse transcription to form the first strand cDNA fragments linked to a 5’ first Rn sequence (Rl) and a 3’ antisense second Rn sequence (R2) ; forming partitions comprising (i) single permeabilized cells comprising the first strand cDNA fragments and (ii) a bead linked to a plurality' of clonal barcoding oligonucleotides having free 3?ends, the barcoding oligonucleotides comprising a first 5' PCR handle sequence, a 3’ capture sequence and a barcode sequence unique to the bead to which the barcode oligonucleotide is linked, wherein the plurality of clonal barcoding oligonucleotides have different 3’ capture sequences such that the plurality includes all of the different Rn sequences on the transposases; optionally releasing the clonal barcoding oligonucleotides in the partitions; in the partitions, annealing a first of the clonal barcoding oligonucleotide 3’ capture sequences to the antisense second Rn sequence of the first strand cDNA fragments and extending the first of the clonal barcoding oligonucleotides using the first strand cDNA fragments as a template with a RNA / DNA-dependent DNA polymerase to form an adaptor- linked intermediate second strand cDNAs that comprise 5 ’-3’: the first 5’ PCR handle sequence, the barcode sequence unique to the bead, the second Rn sequence, optionally the one or more optional barcode sequences, the ME sequence, a second strand cDNA fragment, and an antisense of the ME sequence, an antisense of the first 5 ' PCR handle sequence, optionally an antisense of the one or more optional barcode sequences, an antisense of the first Rn sequence; andin the partitions, annealing a second of the clonal barcoding oligonucleotide 3 ’ capture sequences to the antisense of the first Rn sequence on the adaptor-linked intermediate second strand cDNAs and extending the second of the clonal barcoding oligonucleotides using the adaptor-linked intermediate second strand cDNAs as a template with the DNA-dependent DNA polymerase to form double-barcoded first strand cDNA fragments that comprise 5 ’-3’: the first 5’ PCR handle sequence, the barcode sequence unique to the bead, the first Rn sequence, optionally the one or more optional barcode sequences, the ME sequence, the first strand cDNA fragment, and an antisense of the ME sequence, an antisense of the first 5’ PCR handle sequence, optionally an antisense of the one or more optional barcode sequences, and the antisense of the second Rn sequence; and extending the adaptor-linked intermediate second strand cDNAs using the doublebarcoded first strand cDNAs as a template to form double-stranded double-barcoded cDNA fragments.
[0015] In some embodiments, the method comprises, providing fixed and permeabilized cells; diffusing reverse transcriptase and nucleotides and a primer into the fixed and permeabilized cells and performing reverse transcription of RNA in the cells to form RNA / first strand cDNA hybrids; diffusing into the fixed and permeabilized cells a set of at least three (e.g., 3-20, at least 3, 4. 5, 6, 7, 8, 9, 10, etc.) transposases, each of the at least three transposases carrying different adaptor oligonucleotides, that introduce breaks in the RNA / first strand cDNA hybrids to form RNA / first strand cDNA hybrid fragments and inserts homoadaptor oligonucleotides at the breaks, wherein the adaptor oligonucleotides comprise a first adaptor strand comprising, 5’ to 3’: an Rn sequence, wherein Rn sequences differ for the adaptor oligonucleotides of each of the at least three transposases, optionally, one or more barcode sequences and a mosaic end (ME) sequence, and a second adaptor strand comprising an antisense ME sequence, wherein one of the transposases covalently links 3?ends of the first adaptor strand of the adaptoroligonucleotides to 5’ ends of each strand of the RNA / first strand cDNA hybrid fragments to form RNA / first strand cDNA hybrid fragments having 5' overhangs; displacing the second adaptor strand and (i) extending 3?ends of the first strand cDNAs with a polymerase using the RNA and first adaptor strand linked thereto as a template to form first strand cDNA fragments linked to a 5’ first Rn sequence (Rl) and a 3’ antisense second Rn sequence (R2) or (ii) extending 3’ ends of the RNAs with a polymerase that initiates from RNA using the first strand cDNAs and first adaptor strand linked thereto as a template to form RNA fragments linked to a 5’ first Rn sequence (Rl) and a 3’ antisense second Rn sequence (R2) and then performing a second reverse transcription to form the first strand cDNA fragments linked to a 5’ first Rn sequence (Rl) and a 3’ antisense second Rn sequence (R2); forming partitions comprising (i) single permeabilized cells comprising the first strand cDNA fragments and (ii) a bead linked to a plurality of clonal barcoding oligonucleotides having free 3’ ends, the barcoding oligonucleotides comprising a first 5’ PCR handle sequence, a 3’ capture sequence and a barcode sequence unique to the bead to which the barcode oligonucleotide is linked, wherein the partitions further comprise a plurality of different bridging oligonucleotides having a 5?end sequence and a 3’ end sequence, wherein the plurality of different bridging oligonucleotides include end sequences that anneal to or comprise all of the different Rn sequences on the transposases; optionally releasing the clonal barcoding oligonucleotides in the partitions; and in the partitions, linking barcoding oligonucleotides to both ends of the first strand cDNA fragments linked to a 5’ first Rn sequence (Rl) and a 3’ antisense second Rn sequence (R2) via annealing and polymerase extension using the bridging oligonucleotides to form doublestranded double-barcoded cDNA fragments that comprise 5 ’-3’: the first 5’ PCR handle sequence, the barcode sequence unique to the bead, the first Rn sequence, optionally the one or more optional barcode sequences, the ME sequence, the first strand cDNA fragment, and an antisense of the ME sequence, an antisense of the first 5’ PCR handle sequence, optionally an antisense of the one or more optional barcode sequences, and the antisense of the second Rn sequence.
[0016] In some embodiments, the method further comprises generating barcoded cDNA fragments from the double-stranded double-barcoded cDNA fragments, the generating comprising,in the partitions, optionally amplifying the double-stranded double-barcoded cDNA fragments with primers that anneal to the first 5’ PCR handle sequences; combining contents of the partitions to form a bulk solution; in the bulk solution, annealing primers comprising 5’-3’ a second PCR handle sequence and the ME sequence to the antisense of the ME sequences on strands of the double-stranded double-barcoded cDNAs and extending the primers using a strand of the double-stranded cDNA fragments as a template with a DNA polymerase to form strands comprising 5’-3‘ the second PCR handle sequence and the ME sequence, a first strand cDNA or second strand cDNA fragment, the antisense of the ME sequence, the antisense of the first 5’ PCR handle sequence, optionally the antisense of the one or more optional barcode sequences, the antisense of the first or second Rn sequence, and the antisense of the barcode sequence unique to the bead; treating the bulk solution with an exonuclease to digest single stranded 3’ overhangs, thereby forming singly barcoded cDNA fragments comprising a first and second PCR handle sequence.
[0017] In some embodiments, the method further comprises amplifying the barcoded cDNA fragments with primers that add sequencing adapter sequences on one or both ends. In some embodiments, the method further comprises sequencing cDNA fragments and linked barcode sequences to form sequencing reads. In some embodiments, the method further comprises aligning sequencing reads to a reference genome sequence and compiling larger cDNA sequence from sequencing reads based on location of sequence reads on the reference genome, common Rn end sequences, common overlapping end sequences, common barcode sequence and / or cDNA end sequences mapping less than 10 nucleotides apart on the reference genome.
[0018] In some embodiments, the polymerase is Bst3.0, Bst2.0, Superscript II, or Superscript III.
[0019] In some embodiments, the displacing and extending is performed with a blend of polymerases, wherein the blend has RNA- and DNA-dependent DNA polymerase activity.
[0020] In some embodiments, the adaptor oligonucleotides comprise one or more barcodes. In some embodiments, the one or more barcodes comprises a unique molecular identifier(UMI) barcode. In some embodiments, the one or more barcodes comprises a sample barcode.
[0021] In some embodiments, the partitions are droplets, impermeable capsules or semi- permeable capsules or microwells.
[0022] Also provided is a plurality of partitions as described herein. In some embodiments, the partitions comprise,(i) a plurality of clonal barcoding oligonucleotides having free 3’ ends, the barcoding oligonucleotides linked to or cleaved from a bead in the partitions, the barcoding oligonucleotides comprising a first 5‘ PCR handle sequence, a 3’ capture sequence and a barcode sequence unique to the bead to which the barcode oligonucleotide is linked, wherein the plurality of clonal barcoding oligonucleotides have different 3’ capture sequences such that the plurality includes different Rn sequences; or a plurality of clonal barcoding oligonucleotides having free 3' ends, the barcoding oligonucleotides linked to or cleaved from a bead in the partitions, the barcoding oligonucleotides comprising a first 5’ PCR handle sequence, a 3’ capture sequence and a barcode sequence unique to the bead to which the barcode oligonucleotide is linked, wherein the partitions further comprise a plurality of different bridging oligonucleotides having a 5‘ end sequence and a 3?end sequence, wherein the 3’ end sequence anneals to the 3’ capture sequence of the barcoding oligonucleotides and the 5’ end sequences anneal to different Rn sequences; and(ii) double-barcoded cDNA fragments that comprise 5’-3’: the first 5’ PCR handle sequence, the barcode sequence unique to the bead, the first Rn sequence, optionally the one or more optional barcode sequences, the ME sequence, the first strand cDNA fragment, and an antisense of the ME sequence, an antisense of the first 5’ PCR handle sequence, optionally an antisense of the one or more optional barcode sequences, and the antisense of the second Rn sequence.
[0023] Also provided is a method of sequencing cDNA. In some embodiments, the method comprises, providing fixed and permeabilized cells; performing reverse transcription in the cells to form RNA / first strand cDNA hybrids;diffusing into the fixed and permeabilized cells a set of at least three different transposases each of the at least three transposases earning adaptor oligonucleotides that comprise a mosaic end (ME) sequence and different Rn sequences to form RNA / first strand cDNA hybrid fragments having 5’ overhangs; extending 3’ ends of the first strand cDNAs with a polymerase using the RNA and first adaptor strand linked thereto as a template to form first strand cDNA fragments linked to a 5 ’ first Rn sequence and a 3?antisense second Rn sequence; forming partitions comprising (i) single permeabilized cells comprising the first strand cDNA fragments and (ii) a bead linked to a plurality of clonal barcoding oligonucleotides having free 3’ ends, the barcoding oligonucleotides comprising a first 5’ PCR handle sequence, a 3’ capture sequence and a barcode sequence unique to the bead to which the barcode oligonucleotide is linked, wherein the plurality of clonal barcoding oligonucleotides have different 3’ capture sequences such that the plurality includes all of the different Rn sequences on the transposases; optionally releasing the clonal barcoding oligonucleotides in the partitions; in the partitions forming double-stranded double-barcoded cDNA fragments that comprise 5'- 3‘: the first 5‘ PCR handle sequence, the barcode sequence unique to the bead, a first Rn sequence, the ME sequence, a cDNA fragment sequence, and an antisense of the ME sequence, an antisense of the first 5’ PCR handle sequence, and the antisense of the second Rn sequence, wherein the forming comprises annealing and extending with a polymerase clonal barcoding oligonucleotide 3’ capture sequences to the antisense of the first Rn sequence; combining contents of the partitions to form a bulk solution; in the bulk solution, annealing primers comprising 5’-3’ a second PCR handle sequence and the ME sequence to the antisense of the ME sequences on strands of the double-stranded double-barcoded cDNAs and extending the primers using a strand of the double-stranded cDNA fragments as a template with a DNA polymerase to form strands comprising 5’-3‘ the second PCR handle sequence and the ME sequence, a first strand cDNA or second strand cDNA fragment, the antisense of the ME sequence, the antisense of the first 5’ PCR handle sequence, the antisense of the first or second Rn sequence, and the antisense of the barcode sequence unique to the bead;treating the bulk solution with an exonuclease to digest single stranded 3’ overhangs; optionally amplifying the barcoded cDNA fragments with primers that add sequencing adapter sequences on one or both end; and sequencing cDNA fragments and linked barcode sequences to form sequencing reads.
[0024] In some embodiments, the polymerase is Bst3.0, Bst2.0, Superscript II, or Superscript III.
[0025] In some embodiments, the displacing and extending is performed with a blend of polymerases, wherein the blend has RNA- and DNA-dependent DNA polymerase activity.
[0026] In some embodiments, the extending with the DNA-dependent DNA polymerase occurs at a higher temperature than the displacing and annealing.
[0027] In some embodiments, the adaptor oligonucleotides comprise one or more barcodes. In some embodiments, the one or more barcodes comprises a unique molecular identifier (UMI) barcode. In some embodiments, the one or more barcodes comprises a sample barcode.
[0028] In some embodiments, the partitions are droplets, impermeable capsules or semi- permeable capsules or microwells.
[0029] Also provided are methods of generating cDNA fragments for sequencing. In some embodiments, the method comprises, providing fixed and permeabilized cells; performing reverse transcription in the cells to form RNA / first strand cDNA hybrids; diffusing into the fixed and permeabilized cells a set of at least three different transposases each of the at least three (e.g., 3-20, at least 3, 4. 5, 6, 7, 8, 9, 10, etc.) transposases carrying adaptor oligonucleotides that comprise a mosaic end (ME) sequence and different Rn sequences to form RNA / first strand cDNA hybrid fragments having 5’ overhangs comprising Rn sequences.
[0030] Also provided is a plurality of double-stranded double-barcoded cDNA fragments as described herein. In some embodiments, the fragments comprise: a first 5’ PCR handle sequence, a barcode sequence (e.g., unique to a bead), a first Rn sequence, optionally one or more optional barcode sequences, a mosaic end (ME) sequence, a first strand cDNAfragment, and an antisense of the ME sequence, an antisense of the first 5’ PCR handle sequence, optionally an antisense of the one or more optional barcode sequences, and an antisense sequence of a second Rn sequence.
[0031] In some embodiments, at least 4, 5,6, 7, 8, 9, or 10 different double-stranded double-barcoded cDNA fragments are present in the plurality, wherein each of the at least 4, 5,6, 7, 8, 9, or 10 double-stranded double-barcoded cDNA fragments have at least one different Rn sequence from the others of the at least 4. 5,6, 7, 8, 9, or 10. In some embodiments, the one or more optional barcodes are unique molecular identifier (UMI) barcodes, sample barcodes, or both.BRIEF DESCRIPTION OF THE DRAWINGS
[0032] FIG. 1 depicts options in which only two different tagmentases (i.e. , loaded with different oligonucleotide adaptors) are used on RNA / first strand cDNA hybrids, and the resulting fragments that are available for downstream processing (left) and demonstrates that with an increased number of tagmentases (i.e., each '‘different” tagmentase has different oligonucleotide adaptors), more RNA / first strand cDNA fragments generated by tagmentation can be amplified, leaving only a small fraction of fragments not available for downstream processing. Rn (e.g., Rl, R2. R3. etc.) refer to oligonucleotide adaptors of different sequence, for example having different 5’ sequences that can be used subsequently to amplify fragments with different Rn ends.
[0033] FIG. 2 depicts an example in which nine different tagmentase oligonucleotide adaptors (R1-R9) are used, each delivered by a different homo-adaptor-loaded tagmentase. The figure depicts an exemplary fragment bound by different “Rn” adaptors. Each fragment will occur in the mixture but for convenience only one is blown up in the depiction. In section A, a product of tagmentation is shown in w hich different adaptors have been added to the fragment (each end resulting from a tagmentase having different oligonucleotides adaptors on it (i.e., having 5’ ends Rl and R2). In embodiments described herein, tagmentation occurs in permeabilized and fixed cells following reverse transcription. Cells are distributed into partitions (i.e., so that the majority (e.g., 90, 95, 99% or more cells) occur as single cells in a partition. This can be achieved for example by allowing for more partitions than cells, such that most cells occur in separate partitions, with some partitions being empty. Tagmentation can occur in partitions. In section B, the results of a gap-filling reaction is depicted, displacing antisense mosaic end (ME) sequences (which are delivered bythe tagmentase) and extending the 3‘ ends of the RNA or first strand cDNA fragments. Gapfilling can be performed, for example with both RNA- and DNA-dependent DNA polymerase activity. Extension of the cDNA will require using at least a few nucleotides of the RNA as a template due to the nature of tagmentation cleavage. Sections C and D depict the two separate strands generated. An apostrophe (“ ' ’") is used to depict antisense (reverse complement) sequences.
[0034] FIG. 3 depicts reactions in partitions that start with cDNA fragments comprising different Rn sequences on different sides of the cDNA fragment and result in doublebarcoded polynucleotides.
[0035] FIG. 4 depicts reactions in bulk that are used to generate singly-bead-barcoded cDNA fragments flanked by different Rn sequences.
[0036] FIG. 5 depicts methods of adding PCT handle sequences (P5 and P7 as depicted) on opposite ends of the polynucleotides.
[0037] FIG. 6 depicts an example of two products that can subsequently be sequenced. While two products are shown, depending on the number of oligonucleotides used in the tagmentase (i.e., number N of Rn sequences, many products can be generated.
[0038] FIG. 7 depicts exemplary options for read stitching, i.e., forming a cDNA sequence based on the sequences of various fragments and their linked Rn sequences.
[0039] FIG. 8 depicts an embodiment in which instead of the option depicted at the top of FIG. 3, the bridging oligonucleotides anneals to the tagmented polynucleotides and the barcoding oligonucleotides, allowing for linkage of the barcoding oligonucleotide to the tagmented polynucleotide.
[0040] FIG. 9 depicts an alternative embodiment to FIG. 8 in which instead of the option depicted at the top of FIG. 3. the bridging oligonucleotides anneals to the tagmented polynucleotides and the barcoding oligonucleotides, allowing for linkage of the barcoding oligonucleotide to the tagmented polynucleotide.
[0041] FIG. 10 depicts results of the experiments described in the Example, for example size of fragments generated in various steps, showing that the size of fragments shifted as predicted.DEFINITIONS
[0042] Unless defined otherwise, all technical and scientific terms used herein generally have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Generally, the nomenclature used herein and the laboratory procedures in cell culture, molecular genetics, organic chemistry', and nucleic acid chemistry and hybridization described below are those well-known and commonly employed in the art. Standard techniques are used for nucleic acid and peptide synthesis. The techniques and procedures are generally performed according to conventional methods in the art and various general references (see generally, Sambrook et al. MOLECULAR CLONING: A LABORATORY MANUAL, 2d ed. (1989) Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y.. which is incorporated herein by reference), which are provided throughout this document. The nomenclature used herein and the laboratory procedures in analytical chemistry, and organic synthetic described below are those well-known and commonly employed in the art.
[0043] The term "amplification reaction" refers to any in vitro means for multiplying the copies of a target sequence of nucleic acid in a linear or exponential manner. Such methods include but are not limited to polymerase chain reaction (PCR); DNA ligase chain reaction (see U.S. Pat. Nos. 4,683,195 and 4,683,202; PCR Protocols: A Guide to Methods and Applications (Innis et al., eds, 1990)) (LCR); QBeta RNA replicase and RNA transcriptionbased amplification reactions (e.g. amplification that involves T7, T3, or SP6 primed RNA polymerization), such as the transcnption amplification system (TAS), nucleic acid sequence based amplification (NASBA), and self-sustained sequence replication (3SR); isothermal amplification reactions (e.g., single-primer isothermal amplification (SPIA)); as well as others known to those of skill in the art.
[0044] "Amplifying" refers to a step of submitting a solution to conditions sufficient to allow for amplification of a polynucleotide if all of the components of the reaction are intact. Components of an amplification reaction include, e.g., primers, a polynucleotide template, polymerase, nucleotides, and the like. The term "amplifying" ty pically refers to an "exponential" increase in target nucleic acid. However, "amplifying" as used herein can also refer to linear increases in the numbers of a select target sequence of nucleic acid, such as is obtained with cycle sequencing or linear amplification. In an exemplary embodiment, amplifying refers to PCR amplification using a first and a second amplification primer.
[0045] The term "amplification reaction mixture" refers to an aqueous solution comprising the various reagents used to amplify a target nucleic acid. These include enzymes, aqueous buffers, salts, amplification primers, target nucleic acid, and nucleoside triphosphates. Amplification reaction mixtures may also further include stabilizers and other additives to optimize efficiency and specificity. Depending upon the context, the mixture can be either a complete or incomplete amplification reaction mixture.
[0046] "Polymerase chain reaction" or "PCR" refers to a method whereby a specific segment or subsequence of a target double-stranded DNA, is amplified in a geometric progression. PCR is well known to those of skill in the art; see, e.g., U.S. Pat. Nos. 4,683,195 and 4,683,202; and PCR Protocols: A Guide to Methods and Applications, Innis et al., eds, 1990. Exemplary PCR reaction conditions typically comprise either two or three step cycles. Two step cycles have a denaturation step followed by a hybridization / elongation step. Three step cycles comprise a denaturation step followed by a hybridization step followed by a separate elongation step.
[0047] A "primer" refers to a polynucleotide sequence that hybridizes to a sequence on a target nucleic acid and serves as a point of initiation of nucleic acid synthesis. Primers can be of a variety of lengths and are often less than 50 nucleotides in length, for example 12-30 nucleotides, in length. The length and sequences of primers for use in PCR can be designed based on principles known to those of skill in the art, see, e.g., Innis et al., supra. Primers can be DNA, RNA, or a chimera of DNA and RNA portions. In some cases, primers can include one or more modified or non-natural nucleotide bases. In some cases, primers are labeled.
[0048] ‘ ‘Primer extension” refers to any method in which a primer is extended in a template-specific manner. Examples of primer extension include, for example, methods in which a primer hybridizes to a template nucleic acid and a polymerase extends the primer in a template-specific manner. In some embodiments, the template is DNA and the polymerase is a DNA polymerase. In some embodiments, the template is RNA and the polymerase is a reverse-transcriptase. Primer extension can also include, for example, template switching (see, e.g., Zhu YY, Machleder EM, et al. (2001) Biotechniques , 30(4): 892-897; Ramskold D, Luo S, et al. (2012) Nat Biotechnol, 30(8):777-78, and nick polymerization (also referred to as nick translation), the latter involving nicking one strand of a nucleic acid duplex and usingthe nicked strand as a primer that is extended using the other strand as a template (see, e.g., Leonard G. Davis Ph.D., et al, in Basic Methods in Molecular Biology, 1986).
[0049] A nucleic acid, or a portion thereof, “hybridizes’7or “anneals” to another nucleic acid under conditions such that non-specific hybridization is minimal at a defined temperature in a physiological buffer (e.g., pH 6-9, 25-150 mM chloride salt or in a PCR reaction mixture). In some cases, a nucleic acid, or portion thereof, hybridizes to a conserved sequence shared among a group of target nucleic acids. In some cases, a primer, or portion thereof, can hybridize to a primer binding site if there are at least about 6, 8, 10, 12, 14, 16, or 18 contiguous complementary nucleotides, including “universal” nucleotides that are complementary to more than one nucleotide partner. Alternatively, a primer, or portion thereof, can hybridize to a primer binding site if there are fewer than 1 or 2 complementarity mismatches over at least about 12, 14, 16, or 18 contiguous complementary nucleotides. In some embodiments, the defined temperature at which specific hybridization occurs is room temperature. In some embodiments, the defined temperature at which specific hybridization occurs is higher than room temperature. In some embodiments, the defined temperature at which specific hybridization occurs is at least about 37, 40, 42, 45, 50, 55, 60, 65, 70, 75. or 80 °C. In some embodiments, the defined temperature at which specific hybridization occurs is 37, 40, 42, 45, 50, 55, 60, 65, 70, 75, or 80 °C.
[0050] A "template" refers to a polynucleotide sequence that comprises the polynucleotide to be amplified, flanked by or a pair of primer hybridization sites. Thus, a "target template" comprises the target polynucleotide sequence adjacent to at least one hybridization site for a primer. In some cases, a "target template" comprises the target polynucleotide sequence flanked by a hybridization site for a “forward” primer and a “reverse” primer.
[0051] As used herein, "nucleic acid" means DNA, RNA, single-stranded, double-stranded, or more highly aggregated hybridization motifs, and any chemical modifications thereof. Modifications include, but are not limited to, those providing chemical groups that incorporate additional charge, polarizability, hydrogen bonding, electrostatic interaction, points of attachment and functionality to the nucleic acid ligand bases or to the nucleic acid ligand as a whole. Such modifications include, but are not limited to, peptide nucleic acids (PNAs), phosphodiester group modifications (e.g.. phosphorothioates, methylphosphonates), 2'-position sugar modifications, 5-position pyrimidine modifications, 8-position purine modifications, modifications at exocyclic amines, substitution of 4-thiouridine, substitution of5-bromo or 5-iodo-uracil; backbone modifications, methylations, unusual base-pairing combinations such as the isobases, isocytidine and isoguanidine and the like. Nucleic acids can also include non-natural bases, such as, for example, nitroindole. Modifications can also include 3' and 5' modifications including but not limited to capping with a fluorophore (e.g., quantum dot) or another moiety.
[0052] A "polymerase" refers to an enzyme that performs template-directed synthesis of polynucleotides, e.g.. DNA and / or RNA. The term encompasses both the full length polypeptide and a domain that has polymerase activity. DNA polymerases are well-known to those skilled in the art, including but not limited to DNA polymerases isolated or derived from Pyrococcus fur iosus. Thermococcus litoralis, and Thermotoga maritime, or modified versions thereof. Additional examples of commercially available polymerase enzymes include, but are not limited to: Klenow fragment (New England Biolabs® Inc.). Taq DNA polymerase (QIAGEN), 9 °N™ DNA polymerase (New England Biolabs® Inc ), Deep Vent™ DNA polymerase (New England Biolabs® Inc.), Manta DNA polymerase (Enzymatics®), Bst DNA polymerase (New England Biolabs® Inc.), and phi29 DNA polymerase (New England Biolabs® Inc.).
[0053] Polymerases include both DNA-dependent polymerases and RNA-dependent polymerases such as reverse transcriptase. At least five families of DNA-dependent DNA polymerases are known, although most fall into families A, B and C. Other t pes of DNA polymerases include phage polymerases. Similarly, RNA polymerases typically include eukaryotic RNA polymerases I, II, and III. and bacterial RNA polymerases as well as phage and viral polymerases. RNA polymerases can be DNA-dependent and RNA-dependent.
[0054] As used herein, the term "partitioning" or "partitioned" refers to separating a sample into a plurality of portions, or "partitions." Partitions are generally physical, such that a sample in one partition does not, or does not substantially, mix with a sample in an adjacent partition. Partitions can be solid or fluid. In some embodiments, a partition is a solid partition, e.g., a microchannel. In some embodiments, a partition is a fluid partition, e.g., a droplet. In some embodiments, a fluid partition (e.g., a droplet) is a mixture of immiscible fluids (e.g., water and oil). In some embodiments, a fluid partition (e.g., a droplet) is an aqueous droplet that is surrounded by an immiscible carrier fluid (e.g., oil). Other partitions can include, but are not limited to wells (e.g., microwells) and capsules, including capsules that can later be degraded or semi-permeable capsules, that retain larger molecules such aspolynucleotides but allow for diffusion of reagents such as enzymes. Exemplary array of wells and well descriptions can be found for example in U.S. Patent No. 9,103,754 and 10,391,493. The array of wells (set of nanowells, microwells, wells) can function to capture the solid supports, optionally in addressable, known locations. As such, the array of wells can be configured to facilitate bead capture in at least one of a single-solid support format or optionally in small groups of solid supports. Exemplary microwell arrays and methods of delivery of beads to the microwells and analysis thereof is described in. e.g., PCT / US2021 / 034152.
[0055] As used herein a “barcode” is a short nucleotide sequence (e.g. , at least about 4, 6, 8, 10, 12, 14, 16, 18, 20 or more nucleotides long) that identifies a molecule to which it is conjugated. Barcodes can be used, e.g., to identify molecules in a partition. Such a partitionspecific barcode should be unique for that partition as compared to barcodes present in other partitions. For example, partitions containing target RNA from single-cells can subject to reverse transcription conditions using primers that contain a different partition-specific barcode sequence in each partition, thus incorporating a copy of a unique “cellular barcode” into the reverse transcribed nucleic acids of each partition. Thus, nucleic acid from each cell can be distinguished from nucleic acid of other cells due to the unique “cellular barcode.” In some cases, the cellular barcode is provided by a “bead barcode” that is present on oligonucleotides conjugated to a head, wherein the bead barcode is shared by (e.g., identical or substantially identical amongst) all, or substantially all, of the oligonucleotides conjugated to that bead but is different from most or substantially all oligonucleotides conjugated to other beads. Thus, cellular and bead barcodes can be present in a partition, attached to a bead, or bound to cellular nucleic acid as multiple copies of the same barcode sequence. Cellular or bead barcodes of the same sequence can be identified as deriving from the same cell, partition, or bead. Such partition-specific, cellular, or bead barcodes can be generated using a variety of methods, which methods result in the barcode conjugated to or incorporated into a solid or hydrogel support (e.g., a solid bead or particle or hydrogel bead or particle). In some cases, the partition-specific, cellular, or bead barcode is generated using a split and mix (also referred to as split and pool) synthetic scheme as described herein. A partition-specific barcode can be a cellular barcode and / or a bead barcode (for example when associated with a cell or partition or both). Similarly, a cellular barcode can be a partition specific barcode (when provided in a partition) and / or a bead barcode (when delivered by abead). Additionally, a bead barcode can be a cellular barcode and / or a partition-specific barcode.
[0056] In other cases, barcodes uniquely identify the molecule to which it is conjugated and are referred to as a unique molecular identifier (UMI). The number of nucleotides of the UMI, which can be continuous, or discontinuous, will depend on the number of UMI sequences required. In some embodiments, the number of UMIs available are many times (e.g., 2X, 10X, 100X, etc.) higher than possible conjugation partners, thereby reducing the chance of rare duplicates being linked to different molecules. In some embodiments, pools of different UMIs are present in a partition and the composition of the pool acts as an identifiers for the partition, with some UMIs being in common with some other partitions but the total pool of UMIs being unique or substantially unique between partitions. UMI sequences can be generated for example as random sequences of a set length, and in some embodiments is identified by a flanking known sequence.
[0057] The length of the barcode sequence determines how many unique samples can be differentiated. For example, a 1 nucleotide barcode can differentiate 4, or fewer, different samples or molecules; a 4-nucleotide barcode can differentiate 44or 256 samples or less; a 6 nucleotide barcode can differentiate 4096 different samples or less; and an 8 nucleotide barcode can index 65,536 different samples or less. Additionally, barcodes can be attached to both strands either through barcoded primers for both first and second strand synthesis, through ligation, or in a tagmentation reaction.
[0058] Barcodes are typically synthesized and / or polymerized (e.g.. amplified) using processes that are inherently inexact. Thus, barcodes that are meant to be uniform (e.g., a cellular, particle, or partition-specific barcode shared amongst all barcoded nucleic acid of a single partition, cell, or bead) can contain various N-l deletions or other mutations from the canonical barcode sequence. Thus, barcodes that are referred to as “identical” or “substantially identical” copies refer to barcodes that differ due to one or more errors in, e.g.. synthesis, polymerization, or purification errors, and thus contain various N-l deletions or other mutations from the canonical barcode sequence. Moreover, the random conjugation of barcode nucleotides during synthesis using e.g., a split and pool approach and / or an equal mixture of nucleotide precursor molecules as described herein, can lead to low probability events in which a barcode is not absolutely unique (e.g.. different from all other barcodes of a population or different from barcodes of a different partition, cell, or bead). However, suchminor variations from theoretically ideal barcodes do not interfere with the high-throughput sequencing analysis methods, compositions, and kits described herein. Therefore, as used herein, the term “unique’’ in the context of a particle, cellular, partition-specific, or molecular barcode encompasses various inadvertent N-l deletions and mutations from the ideal barcode sequence. In some cases, issues due to the inexact nature of barcode synthesis, polymerization, and / or amplification, are overcome by oversampling of possible barcode sequences as compared to the number of barcode sequences to be distinguished (e.g.. at least about 2-, 5-, 10-fold or more possible barcode sequences). For example, 10,000 cells can be analyzed using a cellular barcode having 9 barcode nucleotides, representing 262,144 possible barcode sequences. The use of barcode technology is well known in the art, see for example Katsuyuki Shiroguchi, et al. Proc Natl Acad Sci U S A., 2012 Jan 24;109(4): 1347- 52; and Smith, AM et al., Nucleic Acids Research Can 11, (2010). Further methods and compositions for using barcode technology include those described in U.S. 2016 / 0060621.
[0059] A “transposase” or “tagmentase” means an enzyme that is capable of forming a functional complex with a transposon end-containing composition and catalyzing insertion or transposition of the transposon end-containing composition into the double-stranded target DNA with which it is incubated in an in vitro transposition reaction.
[0060] The term “transposon end” means a double-stranded DNA that exhibits the nucleotide sequences (the “transposon end sequences”) that are necessary to form the complex with the transposase that is functional in an in vitro transposition reaction. A transposon end forms a “complex” or a “synaptic complex” or a “transposome complex” or a “transposome composition” with a transposase or integrase that recognizes and binds to the transposon end, and which complex is capable of inserting or transposing the transposon end into target DNA with which it is incubated in an in vitro transposition reaction. A transposon end exhibits two complementary sequences consisting of a “transferred transposon end sequence” or “transferred strand” and a “non-transferred transposon end sequence,” or “non transferred strand” For example, one transposon end that forms a complex with a hyperactive Tn5 transposase (e.g., EZ-Tn5™ Transposase, EPICENTRE Biotechnologies, Madison, Wis., USA) that is active in an in vitro transposition reaction comprises a transferred strand that exhibits a “transferred transposon end sequence” as follows:5' AGATGTGTATAAGAGACAG 3' (SEQ ID NO: 1),and a non-transferred strand that exhibits a “non-transferred transposon end sequence” as follows:5' CTGTCTCTTATACACATCT 3' (SEQ ID NO: 2).
[0061] The 3'-end of a transferred strand is joined or transferred to target DNA in an in vitro transposition reaction. The non-transferred strand, which exhibits a transposon end sequence that is complementary to the transferred transposon end sequence, is not joined or transferred to the target DNA in an in vitro transposition reaction.
[0062] The term "‘solid support” refers to the surface of a bead, microtiter well or other surface that is useful for attaching a nucleic acid, such as an oligonucleotide or polynucleotide. The surface of the solid support can be treated to facilitate attachment of a nucleic acid, such as a single stranded nucleic acid.
[0063] The term “bead” refers to any solid support that can be in a partition, e.g., a small particle or other solid support. In some embodiments, the beads comprise an alginate matrix, i.e.. calcium alginate. In some embodiments, the beads comprise polyacrylamide. For example, in some embodiments, the beads incorporate barcode oligonucleotides into the gel matrix through an acrydite chemical modification attached to each oligonucleotide.Exemplary beads can include hydrogel beads. In some cases, the hydrogel is in sol form. In some cases, the hydrogel is in gel form. An exemplary hydrogel is an agarose hydrogel. Other hydrogels include, but are not limited to, those described in, e.g., U.S. Patent Nos. 4,438,258; 6,534,083; 8,008,476; 8,329,763; U.S. Patent Appl. Nos. 2002 / 0,009,591;2013 / 0,022,569; 2013 / 0,034,592; and International Patent Publication Nos. WO / 1997 / 030092; and WO / 2001 / 049240.
[0064] It will be understood that any range of numerical values disclosed herein can include the endpoints of the range, and any values or subranges in between the endpoints. For example, the range 1 to 10 includes the endpoints 1 and 10, and any value between 1 and 10. The values typically include one significant digit.
[0065] The term “sample” refers to a biological composition, such as a cell, comprising a target nucleic acid.
[0066] The term “about” refers to the usual error range for the respective value that is known by a person of ordinary skill in the art for this technical field, for example, a range of± 10%, ± 5%, or ± 1% can encompass the recited value, even if the recited value is not modified by the term "about."
[0067] All ranges described herein can include the end point values of the range, and any sub-range of values included between the endpoints of the range, where the values include the first significant digit. For example, a range of 1 to 10 includes a range from 2 to 9, 3 to 8, 4 to 7, 5 to 6, 1 to 5, 2 to 5, 2 to 10, 3 to 10, and so on.DETAILED DESCRIPTION OF THE INVENTION
[0068] The disclosure provides single cell RNA sequencing (scRNA-seq) methods comprising multiple (for example, N>2) transposomes rather than two transposomes for tagmentation. While in general the transposases will be the same, the transposomes will vary based on different oligonucleotides having different sequences carried by the different transposases. Whereas two different transposomes have been used in the past to generate amplicons in later steps of an NGS workflow, the inventors have determined that use of three or more transposomes (where the three or more each deliver a different oligonucleotide sequence, explained more herein) can reduce read loss from 50% to 1 / N, where N is the number of different transposomes used. These transposomes can also introduce indexes, thereby labelling tagmentation cut sites. In some embodiments, after tagmentation, cells together with barcode beads are encapsulated into partitions, in which tagged cells are lysed and transposome-indexed RNA:cDNA molecules are amplified and barcoded during by PCR or primer extension in partitions. Barcoded cDNA fragments from the partitions are pooled together and further amplified to obtain a NGS library. One can compile larger cDNA sequences from resulting sequencing reads of the cDNA fragments based on one or more of: location of sequence reads on the reference genome, identification of common end sequences (referred to herein as Rn sequences) introduced from oligonucleotides attached to fragment ends by a transposase, common overlapping end sequences, common barcode sequence and / or cDNA end sequences mapping less than 10 nucleotides apart on the reference genome, each of which can assist in identifying when cDNA fragment reads came from the same larger cDNA sequence.
[0069] The methods and compositions described herein provide for a number of different useful aspects. Methods are provided for generating double-stranded cDNA fragments having a partition-specific barcode sequences at both ends of the molecule. Moreover, asexplained herein, the methods involve using three or more transposases (e.g. tagmentases) each of which are loaded with different oligonucleotides having different 5’ Rn sequences, which allows for detection of a larger number of cDNA fragments compared to methods involving fewer transposases with fewer oligonucleotide sequences. Further, the methods allow for splicing longer cDNA sequences from sequencing reads from shorter cDNA fragments in view of use of multiple criteria to align fragments from the same cDNA in the same cell with each other. In one aspect, the methods comprise performing reverse transcription in fixed and permeabilized cells and then introducing N number (N>2, e.g., 3, 4, 5, 6, 7, 8, 9, 10, e.g., 3-10, 3-20, etc.) different transposomes, i.e., transposases loaded with different oligonucleotides, wherein the oligonucleotides of the N different transposomes have different 5’ end sequences (referred to below as Rn sequences) such that RNA / first strand cDNA hybrid fragments are generated that have a random assortment of Rn different ends.
[0007] The methods described herein can comprise permeabilized cells (e.g., which can comprise the N different transposomes described herein) which for example can be fixed or encapsulated in a hydrogel bead, for example such that the RNA of the cells are compartmentalized from each other in a bulk mixture (e.g., by the structure of the fixed or encapsulated cells ) and the cells can subsequently be introduced into separate partitions.
[0070] Thus, in some embodiments, permeabilized cells are provided. The cells can be permeabilized to allow for entry of reagents while the cells themselves remain substantially intact. Permeabilization can remove cellular membrane lipids to allow molecules such as enzymes to enter the cell while substantially retaining RNA (e.g., >100 or 500 nucleotides long). In some embodiments, a detergent is used for permeabilization. Exemplary detergents can include, for example, Tween-20, Triton X-100, and NP-40 are used for permeabilization (for example, at 0. 1-0.5% (v / v, in PBS). In some embodiments, a steroidal saponin (or saraponin) is used to solubilize lipid, resulting in permeabilization. An exemplary saraponin is Digitonin. The appropriate permeabilization reagent can be selected to be compatible with the integrity of a dow stream partition, if used.
[0071] In some embodiments, the permeabilized cells are fixed cells. For example, in some embodiments, the cells are formalin-fixed, paraffin-embedded (FFPE) samples. In embodiments where the cells are permeabilized and fixed, the cells themselves can act as partitions. In other embodiments, the permeabilized cells are not fixed. In these embodiments, the permeabilized cells can be provided encapsulated in a hydrogel bead,allowing for containment of the cell and its contents while allowing for diffusion of reagents into the cell. Examples of methods for hydrogel bead-encapsulation of cells are described in, for example, Utrech et al.. Advanced Healthcare Materials, Volume 4, Issue 11, August 5, 2015, pages 1628-1633. Alternatively, in some embodiments, the cells, permeabilized and / or fixed or not, can be encompassed in a semi-permeable capsule (SPC).
[0072] Any type of cells can be used according to the methods and compositions described herein. In some embodiments, the cells are mammalian, for example human cells. In some embodiments, the cells are from a biological sample. Biological samples can be obtained from any biological organism, e.g., an animal, plant, fungus, pathogen (e.g., bacteria or virus), or any other organism. In some embodiments, the biological sample is from an animal, e.g., a mammal (e.g., a human or a non-human primate, a cow, horse, pig. sheep, cat, dog, mouse, or rat), a bird (e.g.. chicken), or a fish. A biological sample can be any tissue or bodily fluid obtained from the biological organism, e.g., blood, a blood fraction, or a blood product (e.g., serum, plasma, platelets, red blood cells, and the like), sputum or saliva, tissue (e.g., kidney, lung, liver, heart, brain, nervous tissue, thyroid, eye, skeletal muscle, cartilage, or bone tissue); cultured cells, e.g.. primary cultures, explants, and transformed cells, stem cells, or cells found in stool, urine, etc.
[0073] As noted above, reagents can be added to the bulk solution comprising the permeabilized cells allowing diffusion of reagents such as enzymes, nucleotides and smaller reagents into the cells for molecular activities. Thus, an appropriate concentration of reagents are added to the bulk solution and allowed to diffuse into the cells, optionally with stirring or other types of mixing.
[0074] Reverse transcription in the permeabilized will result in generation of first strand cDNAs. The RNA from the lysed cells can be reverse transcribed in the cells to form cDNAs. Reverse transcription (RT) is an amplification method that copies RNA into DNA. The disclosure provides for reverse transcribing one or more RNA in the permeabilized cells under conditions to allow7for reverse transcription and generation of a first strand cDNA. The RT reaction can be primed with primers to prime an RT reaction from at least one target RNA molecule. The primers can have a 3’ end sequence that anneals to target RNA. For example the 3’ end sequence of the primers can be one that is random, an oligo dT (also referred to herein as a “poly I ”) sequence, or an RNA-specific sequence. Oligo dT sequences are single stranded sequences of deoxythymine (dT). The length of the oligo dTsequence can vary, for example, from 6 bases to 30 bases, and may be a mixture of oligo dT sequences with different lengths. Components and conditions for RT reactions are generally known. The components for the RT reaction, such as for example the RTase and nucleotides, and buffers can be applied to a solution comprising the permeabilized cells and then passively diffused into the cells such that the RNA in the cells are reverse transcribed. Nucleotides used can be deoxyribonucleotides (dNTPs) or can be ribonucleotides or modified non-natural nucleotides or mixtures thereof. Depending on the nucleotides types used, specific RT enzymes may need to be selected that can incorporate the nucleotides into a cDNA or otherwise complementary' polynucleotide.
[0075] Suitable reverse transcriptases can include but are not limited to Maxima RNAse+ (Thermo), Maxima RNAse- (Thermo), murine leukemia virus (MLV) reverse transcriptase (Gerard and Grandgenett, Journal of Virology 15:785-797, 1975; Verma, Journal of Virology’ 15:843-854, 1975) or feline leukemia virus (FLV) reverse transcriptase (Rho and Gallo, Cancer Lett., 10:207-221, 1980 or SEQ ID NO: 1, bovine leukemia virus (BLV) (Demirhan et al., Anticancer Res., 16:2501-5, 1996; Drescher et al., Arch Geschwulstforsch., 49:569-79, 1979), Avian Myeloblastosis Virus (AMV) reverse transcriptase, Respiratory Syncytial Virus (RSV) reverse transcriptase, Equine Infectious Anemia Virus (EIAV) reverse transcriptase, Rous-associated Virus-2 (RAV2) reverse transcriptase, SUPERSCRIPT II reverse transcriptase, SUPERSCRIPT III reverse transcriptase (US8541219, US7056716, US7078208), THERMOSCRIPT reverse transcriptase and MMLV RNase H- reverse transcriptase and Sensiscnpt (Qiagen).
[0076] Follo ving formation of the first strand cDNA / RNA hybrid, tagmentation of the hybrid can take place in the permeabilized cells. Tagmentation results in fragmentation of polynucleotides and addition of oligonucleotides on the ends of the resulting fragments. A transposase, carrying two oligonucleotides, that introduces breaks in cDNA / RNA hybrids and introduces oligonucleotides into the break sites is referred to as a “tagmentase” and the action of the tagmentase is referred to as “tagmentation” and can involve introduction of different adaptor sequences on different sides of a DNA breakage point or the adaptor sequences added by a transposases can be identical. A “transposome’' refers to a tagmentase loaded with oligonucleotide adaptors. The transposomes described herein are homoadaptor-loaded tagmentases, i.e., tagmentases that contain t vo adaptors of the same sequence (except optionally carrying different UMI barcodes), yvhich adaptor is added to both ends of a tagmentase-induced breakpoint in the genomic DNA. Adaptor loaded tagmentases are furtherdescribed, e.g., in U.S. Patent Publication Nos: 2010 / 0120098; 2012 / 0301925; and 2015 / 0291942 and U.S. Patent Nos: 5,965.443; U.S. 6,437,109; 7,083,980; 9.005,935; and 9,238,671, the contents of each of which are hereby incorporated by reference in the entirety for all purposes. Tagmentation of RNA / DNA hybrids is described in, e.g., Bo LuLiting et al., eLife 9:e54919 (2020).
[0077] A tagmentase is an enzyme that is capable of forming a functional complex with a transposon end-containing composition and catalyzing insertion or transposition of the transposon end-containing composition into the double-stranded target DNA with which it is incubated in an in vitro transposition reaction. Exemplary' transposases include but are not limited to modified Tn5 transposases that are hyperactive compared to wildtype Tn5, for example can have one or more mutations selected from E54K, M56A, or L372P. Wild-type Tn5 transposon is a composite transposon in which two near-identical insertion sequences (IS50L and IS50R) are flanking three antibiotic resistance genes (Reznikoff WS. Annu Rev Genet 42: 269-286 (2008)). Each IS50 contains two inverted 19-bp end sequences (ESs), an outside end (OE) and an inside end (IE). However, wild-ty pe ESs have a relatively low activity and were replaced in vitro by’ hyperactive mosaic end (ME) sequences. A complex of the transposase with the 19-bp ME is thus all that is necessary for transposition to occur, provided that the intervening DNA is long enough to bring two of these sequences close together to form an active Tn5 transposase homodimer (Reznikoff WS., Mol Microbiol 47 : 1199-1206 (2003)). Transposition is a very infrequent event in vivo, and hyperactive mutants were historically derived by introducing three missense mutations in the 476 residues of the Tn5 protein (E54K, M56A, L372P), which is encoded by IS50R (Goryshin IY, Reznikoff WS. 1998. J Biol Chem 273: 7367-7374 (1998)). Transposition works through a “cut-and- paste"’ mechanism, where the Tn5 excises itself from the donor DNA and inserts into a target sequence, creating a 9-bp duplication of the target (Schaller H. Cold Spring Harb Symp Quant Biol 43: 401-408 (1979); Reznikoff S.,Annu Rev Genet 42: 269-286 (2008)). In current commercial solutions (Nextera™ DNA kits, Illumina), free synthetic ME adaptors are end-joined to the 5 '-end of the target DNA by the transposase (tagmentase).
[0078] As described herein, at least three (e.g., 3, 4, 5, 6, 7, 8, 9, 10, e.g., 3-10, 3-20, etc.) different tagmentases are used to cleave the cDNA / RNA hybrids in the permeabilized cells. See, e.g., FIG. 1. “Different tagmentases” means for the purposes of this disclosure that the tagmentases (i. e. , copies of the same enzyme) carry different oligonucleotides differing at least by their 5" end Rn sequences. This is shown schematically in the top right of FIG. 1,showing the result of random fragmentation with different tagmentases carrying 9 different adaptor oligonucleotides. As shown, a large majority of fragments formed result, due to random cleavage, from cleavage by two different tagmentases that deliver different Rn sequences. As only fragments having different Rn sequences will be amplified downstream, the use of more than three tagmentases and their different adaptor sequences allows for more fragments being available for sequencing. This can be contrasted with, for example, the depiction top left of FIG. 1 in which only two different 5’ adaptor oligonucleotide sequences are used, resulting in -50% of fragment not being amplified in subsequent steps.
[0079] Adaptor oligonucleotides delivered by the tagmentases can have the following sequence 5'-3’: an Rn sequence, wherein Rn sequences differ for the adaptor oligonucleotides of each of the at least transposases, optionally a spacer sequence, optionally one or more barcode sequences and a mosaic end (ME) sequence (e.g., 5'- CTGTCTCTTATACACATCT-3'). The Rn sequence will be used as an adaptor sequence for annealing in dow nstream step(s) after the Rn sequence is added as part of the adaptor oligonucleotide to DNA or RNA fragments by the tagmentase. The length and complexity of the Rn sequence can be selected based on various criteria, e.g., such as the number of Rn sequences used. The Rn sequence will in some embodiments be at least 14 nucleotides long, e.g., 14-30 nucleotides long.
[0080] The adaptor oligonucleotides loaded on the tagmentase will including a doublestranded portion comprising the ME sequence, and thus a short antisense ME sequence is present and is delivered to the fragment but is not covalently linked to the fragment, and instead remains solely by base pairing with the ME sequence, which has been covalently linked to the fragment. See, e.g., FIG. 2. As part of fragmentation with the tagmentase a 9 base pair sequence at the site of cleavage is replaced such that both fragments formed by the tagmentase comprise the same 9 bp sequence. As discussed later, this can be used in sequencing to splice sequencing reads into longer sequences.
[0081] In some embodiments, the adaptor oligonucleotides delivered by the tagmentases will comprise one or more barcode, i.e., between the Rn and ME sequences. See, e.g., FIG.2. In some embodiments, the adaptor oligonucleotides cany' a unique molecular identifier (UMI) barcode sequence allowing for a unique sequence for different oligonucleotides, which as used herein can be used to identify specific molecules to which the UMI are linked. The UMI is labeled as “Ui” in the figures. In some embodiments, alternatively, or in addition tothe UMIs, the adaptor oligonucleotides carry a sample barcode sequence allowing for identification of which sample was reverse transcribed and allowing for cDNA fragments from different samples to be combined downstream while allowing for their association with a particular sample or patient. The sample barcode is labeled as “Ti” in the figures.
[0082] The number of different transposomes (i.e., the number of transposases carrying different adaptor oligonucleotides) over two can be selected as preferred. As depicted in FIG.2, even inclusion of three different adaptor oligonucleotides allows for increased number of resulting fragments that can be accesses in the workflow. Whereas use of only two adaptor oligonucleotides leaves half of the fragments not amplifiable and accessible to the downstream workflow, use of four adaptor oligonucleotides leaves only a quarter of the fragments inaccessible. Thus in some embodiments, at least 3, 4, 5. 6, 7, 8, 9, 10 or more different adaptor oligonucleotides, e.g., 3-20, 3-15, 3-10, 4-10, 8-20, etc., are used, where each is carried by a different transposase. To some extent the number of different adaptor oligonucleotides will be limited by synthesis of different sequences or generation of barcoding oligonucleotide or bridge oligonucleotides to capture the fragments in downstream products. The predicted percentage of fragments that are detected in the method will increase as more adaptor oligonucleotides are used. Whereas a large improvement of the percentage of fragments will occur for example between use, for example, of two different adaptors (predicted recovery 50%) and four adaptors (75%), beyond inclusion of 20 (predicted 95% coverage) the improvement in coverage only increases by smaller margins.
[0083] In view of the above, in some embodiments, the methods comprise generating an RNA / cDNA hybrid in permeabilized cells and then contacting the RNA / cDNA hybrids with different transposases earn ing three or more different adaptor oligonucleotides, allowing for the fragments to be accessible for any variety' of downstream workflows for ultimate sequencing of the resulting fragments. Exemplary doyvnstream workflows are described below.
[0084] Gap-filling can occur following tagmentation to produce blunt ends, allowing for example, for primer annealing sites (i.e., antisense Rn sequences), and can occur before or after partitioning as discussed below. See, e.g., FIG. 2. Because of the nature of the tagmentation of a cDNA / RNA hybrid, the gap-filling is performed with RNA- and DNA- dependent polymerase activity, which can be provided in a single enzyme or a cocktail of different enzy mes. Exemplary polymerases having RNA- and DNA-dependent DNApolymerase activity can include, for example, Bst 2.0, Bst 3.0, Superscript II, Superscript III and polymerases described in for example US Patent Publication No. 20230287364.
[0085] The gap-filling enzymes’ activity will displace the short antisense ME oligonucleotide sequence. As depicted in FIG. 2, extension of the 3’ end of the cDNA fragment can be performed using the 5’ end of the complementary' RNA fragment and the attached adaptor oligonucleotide as a template. Because the 5’ end of the RNA is initially used as a template, RNA-dependent DNA polymerase activity is used. Gap-filling activity on the cDNA results in first strand cDNA fragments linked to a 5’ first R sequence (Rl) and any barcodes and a 3’ antisense second R sequence (R2) and antisense barcodes if present, e.g., as depicted at the top of FIG. 3.
[0086] Gap-filling activity can also be used to form a polynucleotide strand at the 3’ end of the RNA fragment using the cDNA fragment and its linked adaptor oligonucleotide as a template (bottom strand in FIG. 2) and if desired the resulting polynucleotide can be used as a template to form a complementary cDNA fragment with the respective end sequences. In these aspects, a second reverse transcription reaction, optionally in the permeabilized cells, can be used to generate a DNA copy, e.g., a first cDNA fragment having the end sequences, e.g., as depicted at the top of FIG. 3. For example, RT can be used to extend 3’ ends of the RNAs with a reverse transcriptase using the first strand cDNAs and first adaptor strand linked thereto as a template to form RNA fragments linked to a 5’ first R sequence (Rl) and a 3’ antisense second R sequence (R2).
[0087] Before or after gap-filling (typically after), the tagmented products and the permeabilized cells that contain them can be partitioned to form partitions containing single cells and one or more bead comprising a plurality of clonal barcoding oligonucleotides having free 3’ ends and used subsequently to add a barcode sequence specific for the bead, allowing for partition-specific barcoding.
[0088] In some embodiments, the permeabilized cells can be partitioned such that individual cells are the only cell within a particular partition. Exemplary partitions can include but are not limited to droplets within an emulsion, microwells, capsules, including but not limited to semi-permeable capsules (SPCs). Methods and compositions for partitioning are described, for example, in published patent applications WO 2010 / 036352, US 2010 / 0173394, US 2011 / 0092373, and US 2011 / 0092376.
[0089] In some embodiments, one or more reagents are added during droplet formation or to the droplets after the droplets are formed. Methods and compositions for delivering reagents to one or more partitions include microfluidic methods as known in the art; droplet or microcapsule combining, coalescing, fusing, bursting, or degrading (e.g., as described in U.S. 2015 / 0027,892; US 2014 / 0227,684; WO 2012 / 149,042; and WO 2014 / 028,537); droplet injection methods (e.g., as described in WO 2010 / 151,776); and combinations thereof. In some embodiments, the droplets described herein are relatively stable and have minimal coalescence between two or more droplets. In some embodiments, less than 0.0001%, 0.0005%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10% of droplets generated from a sample coalesce with other droplets. The emulsions can also have limited flocculation, a process by which the dispersed phase comes out of suspension in flakes. In some cases, such stability or minimal coalescence is maintained for up to 4, 6, 8, 10, 12, 24, or 48 hours or more (e.g., at room temperature, or at about 0, 2, 4, 6, 8, 10, or 12 °C). In some embodiments, the droplet is formed by flowing an oil phase through an aqueous sample or reagents.
[0090] The oil phase of an emulsion can comprise a fluorinated base oil which can additionally be stabilized by combination with a fluorinated surfactant such as a perfluorinated polyether. Exemplar}' oil phase compositions along these lines are described in, e.g., PCT WO 2020 / 247950 and US Patent Publication No. 2017 / 0175179.
[0091] In some embodiments, the methods described herein take advantage of partition aspects available from semi-permeable capsules (SPCs) while allowing for molecular manipulations such as can be available in bulk processing. For example, SPCs can function to retain macromolecules in the SPCs, thus functioning like partitions while allowing for manipulation of a solution containing many SPCs as if it were a bulk solution because reagents added to such a solution can readily diffuse into the SPCs. Accordingly, methods and compositions are provided that allow for single cells (for example but not limited to. mammalian cells) and one or more bead that delivers barcoding oligonucleotides in SPCs. An SPC is a capsule having a semi-permeable shell that allows for small molecules to pass through the shell while substantially retaining larger molecules, such as DNA (e.g., having at least 100 or at least 500 nucleotides), mRNA and optionally some proteins. For example WO-2023117364 describes SPCs with a semipermeable shells comprising a gel formed from a polyampholyte and / or a polyelectrolyte, wherein the polyampholyte and / or the polyelectrolyte in the gel is covalently cross-linked. In some embodiments, as described inWO-2023117364, the SPCs comprise an inner core in a liquid form, or in a hydrogel form and optionally enriched in polyhydroxy compounds belonging to the class of polysaccharides, oligosaccharides, carbohydrates, or sugars. For example, the core can comprise a polyhydroxy compound and / or an antichaotropic agent. In other embodiments, WO- 2023099610 describes processes for manufacturing core-shell microcapsules and methods for using core-shell microcapsules to compartmentalize and optionally process biological entities and molecules. Other SPC formulations can comprise poly(ethylene glycol) diacrylate (PEGDA), for example as described in Michielin and Maerkl, Set Rep. 2022; 12: 21391. Bomi et al. , Langmuir 2015, 31, 6027-6034 describes yet other SPCs based on water-in-oil- in-water droplets that have a middle layer composed of photocurable resin and inert oil.
[0092] In some embodiments, the sample is partitioned into, or into at least, 500 partitions. 1000 partitions. 2000 partitions, 3000 partitions. 4000 partitions, 5000 partitions. 6000 partitions, 7000 partitions, 8000 partitions, 10,000 partitions, 15,000 partitions, 20,000 partitions, 30,000 partitions, 40,000 partitions, 50,000 partitions, 60,000 partitions, 70,000 partitions, 80.000 partitions, 90,000 partitions, 100,000 partitions, 200,000 partitions, 300,000 partitions, 400,000 partitions, 500,000 partitions, 600,000 partitions, 700,000 partitions, 800,000 partitions, 900,000 partitions, 1,000,000 partitions, 2,000,000 partitions, 3,000,000 partitions, 4,000,000 partitions, 5,000,000 partitions, 10,000,000 partitions, 20,000,000 partitions, 30,000,000 partitions, 40,000,000 partitions, 50,000,000 partitions, 60,000,000 partitions, 70,000,000 partitions. 80,000,000 partitions, 90,000,000 partitions, 100,000,000 partitions. 150,000.000 partitions, or 200,000,000 partitions.
[0093] As discussed further herein, plurality of copies of barcoding oligonucleotides linked to a head (e.g., a hydrogel bead), and optionally bridging oligonucleotides if used, can be delivered to the partitions (including for example forming the partitions with the beads and the cells), wherein different partitions receive different beads and accompanying barcoding oligonucleotides. The bead can be attached to multiple copies of the same oligonucleotide, for example, at least about 10, 50, 100, 500, 1000, 5000, 10,000, 50,000, 100,000, 500,000, 1,000,000, 5,000,000, 10,000,000, 108, 109, 1010or more copies of the same or substantially identical oligonucleotide can be attached to one (e g., the same) bead. The barcoding oligonucleotides will comprise at least a bead-specific barcode sequence and a 3’ capture sequence for annealing to a target sequence (e.g., an Rn sequence) on the adaptor oligonucleotides or alternatively to a bridging oligonucleotide as described below. In some embodiments, the barcoding oligonucleotides further comprise a 5’ PCR handle sequence,allowing for a common 5’ sequence between sequences comprising different cell-specific barcodes, allowing them all to be amplified with a universal primer that anneals to the PCR handle sequence. Optionally, once the bead is present in the partition, and before linkage to target fragments, the oligonucleotides can be cleaved from the bead prior to linking the oligonucleotides to the DNA fragments.
[0094] In some embodiments, the 3’ capture sequence on the barcoding oligonucleotides anneal to a reverse complement (i.e.. antisense) of a Rn sequence as it occurs linked to a first strand cDNA fragment. See, e.g., FIG. 3. In other embodiments, the 3’ capture sequence on the barcoding oligonucleotides anneals to a sense Rn sequence. In general, in these aspects, the barcoding oligonucleotides on the bead will comprise a mixture of oligonucleotides having the same barcode sequence, but varying in that different 3‘ capture sequence sequences, that can anneal to different antisense Rn sequences, are present. Thus, for example, if there are four Rn sequences used (i.e., four tagmentation oligonucleotides) then there will be at least four types of barcoding oligonucleotides on the bead, one type for each Rn sequence, (i.e., in this example, type 1 will have antisense R1 capture sequences, type 2 will have antisense R2 capture sequences, type 3 will have antisense R3 capture sequences, and ty pe 4 will have antisense R4 capture sequences). This option is depicted in the figures.
[0095] In an alternative option, the barcoding oligonucleotides have a common 3’ capture sequence that can anneal to a 3’ end of the bridging oligonucleotides that have been provided (diffused in, added upon formation, etc.) in the partition. The 5’ ends of the bridging oligonucleotide vary and in total will represent alternative sequences that can anneal to any of the Rn sequences at the end of the tagmented polynucleotides. FIGs. 8 and 9 depict embodiments in which instead of the option depicted at the top of FIG. 3, the bridging oligonucleotides are used to link the barcoding oligonucleotide to the tagmented polynucleotide.
[0096] Each barcoding oligonucleotide can be linked at its 5’ end or elsewhere on the oligonucleotide to the bead and in some embodiments can include a cleavable moiety to remove the oligonucleotides from the bead, e.g., before the oligonucleotides are linked to the DNA fragments comprising the adaptor sequence. In some embodiments, the cleavable linker comprises a uridine incorporated site in a portion of a nucleotide sequence. A uridine incorporated site can be cleaved, for example, using a uracil glycosylase enzyme (e.g., a uracil N-glycosylase enzyme or uracil DNA glycosylase (UDG) enzyme). In someembodiments, the cleavable linker comprises a photocleavable nucleotide. Photocleavable nucleotides include, for example, photocleavable fluorescent nucleotides and photocleavable biotinylated nucleotides. See, e.g., Li et al., PNAS, 2003, 100:414-419; Luo et al.. Methods Enzymol, 2014, 549: 115-131. In some cases, the oligonucleotides are attached to bead through a disulfide linkage (e.g., through a disulfide bond between a sulfide of the solid support and a sulfide covalently attached to the 5’ or 3’ end. or an intervening nucleic acid, of the oligonucleotide). In such cases, the oligonucleotide can be cleaved from the solid support by contacting the solid support with a reducing agent such as a thiol or phosphine reagent, including but not limited to a beta mercaptoethanol, dithiothreitol (DTT), or tris(2- carboxyethyl)phosphine (TCEP).
[0097] The barcoding oligonucleotides from the bead will include, for example, a beadspecific barcode such that the bead-specific barcode sequence on a first oligonucleotide can be used to distinguish it from a bead-specific barcode from a second oligonucleotide from a different bead. The length of the barcode can depend on the number of different barcodes desired, and in some embodiments, can be between 4-20 nucleotides, which can be contiguous or non-contiguous. The 3’ end of the barcoding oligonucleotides can comprise the Rn sequence from the adaptor oligonucleotides sequences added to the fragments by the tagmentases such that the oligonucleotides from the beads can be used as primers in a primer extension (e.g., PCR) reaction using one strand of the fragments having adaptor sequences at their ends as a template (see, e.g., FIG. 3), or the barcoding oligonucleotide 3’ capture sequence can anneal to a 3’ universal sequence on a bridging oligonucleotide (see, g., FIG. 8 or 9).
[0098] The barcoding oligonucleotides can be covalently or non-covalently linked to the solid support(s). Oligonucleotides can be linked to beads as desired. Methods of linking oligonucleotides to beads are described in, e.g., WO 2015 / 200541. In some embodiments, the oligonucleotide configured to link a hydrogel bead to the barcode is covalently linked to the hydrogel. Numerous methods for covalently linking an oligonucleotide to one or more hydrogel matrices are known in the art. As but one example, aldehyde derivatized agarose can be covalently linked to a 5’ -amine group of a synthetic oligonucleotide.
[0099] Any bead of useful size and composition for deliver)’ to partitions can be used. The particle or bead can be any particle or bead having a solid support surface. Solid supports suitable for particles include controlled pore glass (CPG)(available from Glen Research,Sterling, Va.), oxalyl-controlled pore glass See, e.g., Alul, et al., Nucleic Acids Research 1991, 19, 1527), TentaGel Support-an aminopolyethyleneglycol derivatized support {See, e.g., Wright, et al, Tetrahedron Letters 1993, 34, 3373), polystyrene, Poros (a copolymer of polystyrene / divinylbenzene), or reversibly cross-linked acrylamide. Many other solid supports are commercially available and amenable to the present methods. In some embodiments, the bead material is a polystyrene resin or poly(methyl methacrylate) (PMMA). The bead material can be metal. In some embodiments, the particle or bead comprises hydrogel or another similar composition. In some cases, the hydrogel is in sol form. In some cases, the hydrogel is in gel form. An exemplary hydrogel is an agarose hydrogel. Other hydrogels include, but are not limited to, those described in, e.g., U.S. Patent Nos. 4,438.258; 6,534,083; 8,008,476; 8,329,763; U.S. Patent Appl. Nos. 20020009591; 20130022569; 20130034592; and International Patent Publication Nos. W01997030092; and WO2001049240. Additional compositions and methods for making and using hydrogels, such as barcoded hydrogels, include those described in, e.g., Klein et al., Cell, 2015 May 21; 161(5): 1187-201.
[0100] As noted herein, in some embodiments, bridging oligonucleotides are provided into the partitions to bridge linkage between the tagmented and gap-filled cDNA fragments and the barcoding oligonucleotides. An advantage of use of a bridging oligonucleotide is that the barcoding oligonucleotides can be synthesized with a common 3’ end, potentially allowing for simpler manufacture of the beads and their oligonucleotides. The bridging oligonucleotides can be linked to the tagmented and gap-filled cDNA fragments and the barcoding oligonucleotides via primer extension (e.g., PCR) and / or ligation. In some embodiments, the bridging oligonucleotides are provided to have (i) a common (universal) 3’ end sequence that will anneal to a common 3’ sequence on the barcoding oligonucleotides and (ii) different 5’ ends that can anneal to the tagmentase-fragmented cDNA fragments via the Rn sequence. Depending on the strand, the 3 ’ ends of different bridging oligonucleotides can have different Rn sequences or antisense sequences thereof, allowing for the population of bridging oligonucleotides to anneal to any fragment produced by the tagmentases. See, e.g., FIG. 8-9. Optionally the bridging oligonucleotides can be blocked on their 3' ends such that the bridging oligonucleotides themselves are not extended. See, e.g., FIGs, 8-9. In yet other embodiments, one or both ends of the bridging oligonucleotides, or reverse complements thereof formed by primer extension, are ligated to the tagmented and gap-filledcDNA fragments and the barcoding oligonucleotides, optionally before or after a primer extension step.
[0101] In embodiments in which the bridging oligonucleotides are used, the bridging oligonucleotides anneal to both the barcoding oligonucleotides and the tagmented cDNA fragments allowing for linkage of the barcoding oligonucleotide to the tagmented cDNA fragments, e.g., by using ligation and / or one or more DNA polymerization reaction to add the barcoding oligonucleotides to the end of both strands of the tagmented cDNA fragments.
[0102] Whether the bridging oligonucleotides are used to link the barcoding oligonucleotides to the tagmented cDNA fragments or not, a product can be generated in which a barcoding oligonucleotide has been linked to both ends of the tagmented cDNA fragment. In some embodiments, a further amplification can occur for example to add a first PCR handle sequence. See e.g., bottom of FIG. 3. An example of this is shown at the bottom of FIG. 3. Thus the disclosure provides, for example, double-barcoded first strand cDNA fragments that comprise 5’-3’: a first 5’ PCR handle sequence, the barcode sequence unique to the bead, a first Rn sequence, optionally one or more optional barcode sequences e.g., a UMI and / or sample barcode), the ME sequence, the first strand cDNA fragment, and an antisense of the ME sequence, an antisense of the first 5?PCR handle sequence, optionally an antisense of the one or more optional barcode sequences, and an antisense sequence of a second Rn sequence (which is different from the first Rn sequence). See, e.g., bottom of FIG. 3. Once this double-stranded double-barcoded cDNA fragment has been generated, the contents of partitions can be combined (e.g.. different microwell contents can be mixed, droplets or SPCs can be disrupted, etc.) such that the double-stranded double-barcoded cDNA fragments from different partitions are together in a bulk solution.
[0103] Reactions can then be performed in the bulk solution as desired to generate sequences that can be used in sequencing reactions, e.g.. next generation sequencing reactions, which can include for example sequencing by synthesis of other methods. In some embodiments, primer extension or PCR reactions are performed to add one or more PCR handle sequences to one or both ends of the double-barcoded cDNA fragments. In some embodiments, in the bulk solution one anneals primers comprising 5 ’-3’ a second PCR handle sequence and the ME sequence to the antisense of the ME sequences on strands of the double-stranded double-barcoded cDNAs and extending the pnmers using a strand of the double-stranded cDNA fragments as a template with a DNA polymerase. See, e.g., FIG. 4.This reaction can form strands comprising 5 ’-3’ the second PCR handle sequence and the ME sequence, a first strand cDNA or second strand cDNA fragment, the antisense of the ME sequence, the antisense of the first 5’ PCR handle sequence, optionally the antisense of the one or more optional barcode sequences, optionally the antisense of the spacer sequence, the antisense of the first or second Rn sequence, and the antisense of the barcode sequence unique to the bead. This effectively forms a cDNA fragment linked to only one bead-specific barcode sequence, albeit at the 3’ end of the molecules formed. See, e.g.. FIG. 4. In some embodiments, the primers comprising 5’-3’ a second PCR handle sequence and the ME sequence are used in a higher concentration than that of the underlying template, for example to facilitate annealing of the primers compared to potential competition from self- annealing of the template. Subsequent introduction of a 3’-5?exonuclease that digests 3’ singlestranded overhangs but not double-stranded DNA, can be used to remove sequences 3’ of the antisense ME sequence on the double-stranded cDNAs. Exemplary 3’-5’ endonucleases can include, for example, Macosko et al. , RESOURCE\ Vol. 161, Issue 5, P1202-1214, MAY 21, 2015. The result is a singly -barcoded cDNA fragment. Optionally different PCR handle sequences (e.g., Illumina P5 and P7 sequences) can be added to different ends of the fragments using primers that are used to amplify the singly-barcoded fragment. Resulting products will be double-stranded DNA molecules comprising the cDNA fragments, a beadspecific barcode and different Rn sequences at either end of the cDNA fragment sequence. Because the original tagmentase reaction comprises N different Rn sequences, the products will be mixtures of double stranded DNA molecules that include different combinations of Rn sequences flanking different cDNA fragment sequences. Two such exemplary sequences are depicted in FIG. 4, but is should be appreciated that many more combinations of Rn sequences can occur as more Rn sequences are used in the tagmentation.
[0104] Sequencing of the double-stranded DNA molecules can be performed as desired. Methods for high throughput sequencing and genotyping are known in the art. For example, such sequencing technologies include, but are not limited to, pyrosequencing, sequencing-by- ligation, single molecule sequencing, sequence-by -synthesis (SBS), massive parallel clonal, massive parallel single molecule SBS, massive parallel single molecule real-time, massive parallel single molecule real-time nanopore technology, etc. Morozova and Marra provide a review of some such technologies in Genomics, 92: 255 (2008), herein incorporated by reference in its entirety.
[0105] Exemplary' DNA sequencing techniques include fluorescence-based sequencing methodologies (See, e.g.. Birren et al., Genome Analysis: Analyzing DNA, 1, Cold Spring Harbor, N.Y.; herein incorporated by reference in its entirety). In some embodiments, automated sequencing techniques understood in that art are utilized. In some embodiments, the present technology7provides parallel sequencing of partitioned amplicons (PCT Publication No. WO 2006 / 0841,32, herein incorporated by reference in its entirety). In some embodiments, DNA sequencing is achieved by parallel oligonucleotide extension (See. e.g., U.S. Pat. Nos. 5,750,341 ; and 6,306,597, both of which are herein incorporated by reference in their entireties). Additional examples of sequencing techniques include the Church polony technology (Mitra et al., 2003, Analytical Biochemistry 320, 55-65; Shendure et al., 2005 Science 309, 1728-1732; and U.S. Pat. Nos. 6.432,360; 6,485.944; 6,511.803; herein incorporated by reference in their entireties), the 454 picotiter pyrosequencing technology (Margulies et al., 2005 Nature 437, 376-380; U.S. Publication No. 2005 / 0130173; herein incorporated by reference in their entireties), the Solexa single base addition technology7(Bennett et al., 2005, Pharmacogenomics, 6, 373-382; U.S. Pat. Nos. 6,787.308; and 6.833,246; herein incorporated by reference in their entireties), the Lynx massively parallel signature sequencing technology (Brenner et al. (2000). Nat. Biotechnol. 18:630-634; U.S. Pat. Nos. 5,695,934; 5,714,330; herein incorporated by reference in their entireties), and the Adessi PCR colony technology (Adessi et al. (2000). Nucleic Acid Res. 28. E87; WO 2000 / 018957; herein incorporated by reference in its entirety).
[0106] Typically, high throughput sequencing methods share the common feature of massively parallel, high-throughput strategies, with the goal of lower costs in comparison to older sequencing methods (See, e.g., Voelkerding et al., Clinical Chem, 55: 641-658, 2009; MacLean et al.. Nature Rev. Microbiol., 7:287-296; each herein incorporated by reference in their entirety). Such methods can be broadly divided into those that typically use template amplification and those that do not. Amplification-requiring methods include pyrosequencing commercialized by Roche as the 454 technology7platforms (e.g., GS 20 and GS FLX), the Solexa platform commercialized by Illumina, and the Supported Oligonucleotide Ligation and Detection (SOLiD) platform commercialized by Applied Biosystems. Non-amphfication approaches, also known as single-molecule sequencing, are exemplified by the HeliScope platform commercialized by Heli cos BioSciences, and platforms commercialized by VisiGen, Oxford Nanopore Technologies Ltd., Life Technologies / Ion Torrent, and Pacific Biosciences, respectively.
[0107] Once sequencing reads are generated that comprise the Rn sequences, any barcodes, and the cDNA fragment, the different sequences can be aligned as desired to a reference genome or cDNA sequence. This will allow for example one to identify a likely identity of which cDNA the fragment originated from. Fragments from the cDNA should align with the same reference sequence. In addition, fragments from the precise original cDNA, i.e., the original RNA molecule, can be identified and subsequently stitched together. An example of stitching is shown e.g.. in FIG. 7. Various criteria can be used including one or some or all of:The location of sequence reads on the reference genome. The ends should map to the same place in the reference genome if they are from the same cDNA molecule.Common Rn end sequences. If a tagmentase carrying homoadaptors (both oligonucleotides on a transposases have identical Rn sequences) then adjacent fragments will have the same Rn sequence on properly aligned ends of the cDNA fragments on the reference genome.Common bead barcode sequence. cDNA fragments linked to the same bead barcode sequence can be expected to have originated from the same cell. cDNA fragments linked to different bead barcodes may be from different cells or may come from the same cell that w s in a partition with more than one bead. cDNA end sequences map less than 10 nucleotides apart on the reference genome. Tagmentation results in insert of a 9 bp repeat on either site of the cleavage site of the tagmentase. resulting in cDNA fragments with the same 9 bp end sequence (one on the 5 ’ end and one on the 3’ end of the two fragments).
[0108] Accordingly, the above criteria can be employed to compile cDNA sequences from the resulting sequencing reads, thereby generating sequences for cDNAs.Kits
[0109] Also provided are kits that can be used for performing part or all of the methods described herein. In some embodiments, the kit will comprise one or more container for holding various reagents, and optionally instructions. Exemplary components of a kit, which can be provided separately or in a mixture, as the methods described herein allow, can include, for example, one or more of cell fixatives, cell permeation reagents, adaptor-loaded transposases (optionally more than one), one or more polymerase, which can include forexample a reverse transcriptase and a DNA polymerase, one or more exonuclease, primers for primer extension and / or amplification as described herein, reagents for performing droplet-based amplification (dPCR) and optionally a reaction mixture for performing a second PCR amplification. Other reagents as described herein or as useful for performing the steps described herein can also be included in the kit.EXAMPLES
[0110] In brief. mRNA was reverse-transcribed to cDNA. forming cDNA / RNA hybrid structure. The cDNA / RNA hybrid structure was further tagmented by transposases carrying 19 different adaptor oligonucleotides (FIG. 2). The tagmented product was gap-filled by Bst 3.0 (FIG. 2), followed by the 1st PCR using these 19 different adaptors as primers (as shown in FIG. 3 but in a bulk reaction and without CBC). After purification, the 1st PCR product was denatured, annealed with primer R2, and extended with Bst 3.0 (steps shown in FIG. 4). Leftover primers and single stranded 3’ overhangs were digested by exonuclease (FIG. 4). Full-length sequencing adapters were added during the 2nd PCR (steps show n in FIG. 5).[OHl] Sizes of DNA fragments from 1st PCR product, pre-2nd PCR product, and 2nd PCR product were measured by Bioanalyzer. Lower size marker (~35bp) and Upper size (~10380bp) marker were clearly labeled respectively at the left and the right of each traces as shown in FIG. 10. The broad peaks in the middle of each traces of FIG. 10 were the DNA fragments. As show n in the upper figure of FIG. 10, the presence of amplified DNA fragments after 1 st PCR indicates the feasibility of the tagmentation reaction with multiadapter-assembled transposes (depicted in FIG. 2) and the amplification reaction of these tagged fragments (depicted in FIG. 3). The measured peak shift between the upper figure and the middle figure of FIG. 10 matches the theoretical number, showing the pre-2nd PCR and the exonuclease reactions (see e.g., FIG. 4) are working. Similarly, the measured peak shift between the middle figure and the lower figure of FIG. 10 matches the theoretical number, indicating the 2nd PCR (depicted in FIG. 5) is working.
[0112] Although the foregoing disclosure has been described in some detail by way of illustration and example for purposes of clarity of understanding, one of skill in the art will appreciate that certain changes and modifications may be practiced within the scope of the appended claims. In addition, each reference provided herein, including patents, patent applications, non-patent literature, and Genbank accession numbers, is incorporated by reference in its entirety to the same extent as if each reference w as individually incorporatedby reference. Where a conflict exists between the instant application and a reference provided herein, the instant application shall dominate.
Claims
WHAT IS CLAIMED IS:
1. A method of generating double-stranded double-barcoded cDNA fragments, the method comprising, providing fixed and permeabilized cells; diffusing reverse transcriptase and nucleotides and a primer into the fixed and permeabilized cells and performing reverse transcription of RNA in the cells to form RNA / first strand cDNA hybrids; diffusing into the fixed and permeabilized cells a set of at least three (e.g., 3- 20, at least 3, 4, 5, 6, 7, 8, 9, 10, etc.) transposases, each of the at least three transposases carrying different adaptor oligonucleotides, that introduce breaks in the RNA / first strand cDNA hybrids to form RNA / first strand cDNA hybrid fragments and inserts homoadaptor oligonucleotides at the breaks, wherein the adaptor oligonucleotides comprise a first adaptor strand comprising, 5 ' to 3 ’ : an Rn sequence, wherein Rn sequences differ for the adaptor oligonucleotides of each of the at least three transposases, optionally, one or more barcode sequences and a mosaic end (ME) sequence, and a second adaptor strand comprising an antisense ME sequence, wherein one of the transposases covalently links 3?ends of the first adaptor strand of the adaptor oligonucleotides to 5’ ends of each strand of the RNA / first strand cDNA hybrid fragments to form RNA / first strand cDNA hybrid fragments having 5’ overhangs; displacing the second adaptor strand and (i) extending 3 ' ends of the first strand cDNAs with a polymerase using the RNA and first adaptor strand linked thereto as a template to form first strand cDNA fragments linked to a 5’ first Rn sequence (Ri) and a 3’ antisense second Rn sequence (R2) or (ii) extending 3’ ends of the RNAs with a polymerase that initiates from RNA using the first strand cDNAs and first adaptor strand linked thereto as a template to form RNA fragments linked to a 5’ first Rn sequence (Ri) and a 3' antisense second Rn sequence (R2) and then performing a second reverse transcription to form the first strand cDNA fragments linked to a 5’ first Rn sequence (Ri) and a 3’ antisense second Rn sequence (R2) ; forming partitions comprising (i) single permeabilized cells comprising the first strand cDNA fragments and (ii) a bead linked to a plurality of clonal barcoding oligonucleotides having free 3’ ends, the barcoding oligonucleotides comprising a first 5’PCR handle sequence, a 3’ capture sequence and a barcode sequence unique to the bead to which the barcode oligonucleotide is linked, wherein the plurality of clonal barcoding oligonucleotides have different 3’ capture sequences such that the plurality includes all of the different Rn sequences on the transposases; optionally releasing the clonal barcoding oligonucleotides in the partitions; in the partitions, annealing a first of the clonal barcoding oligonucleotide 3‘ capture sequences to the antisense second Rn sequence of the first strand cDNA fragments and extending the first of the clonal barcoding oligonucleotides using the first strand cDNA fragments as a template with a RNA / DNA-dependent DNA polymerase to form an adaptor- linked intermediate second strand cDNAs that comprise 5 ’-3’: the first 5’ PCR handle sequence, the barcode sequence unique to the bead, the second Rn sequence, optionally the one or more optional barcode sequences, the ME sequence, a second strand cDNA fragment, and an antisense of the ME sequence, an antisense of the first 5’ PCR handle sequence, optionally an antisense of the one or more optional barcode sequences, an antisense of the first Rn sequence; and in the partitions, annealing a second of the clonal barcoding oligonucleotide 3‘ capture sequences to the antisense of the first Rn sequence on the adaptor-linked intermediate second strand cDNAs and extending the second of the clonal barcoding oligonucleotides using the adaptor-linked intermediate second strand cDNAs as a template with the DNA- dependent DNA polymerase to form double-barcoded first strand cDNA fragments that comprise 5’-3’: the first 5’ PCR handle sequence, the barcode sequence unique to the bead, the first Rn sequence, optionally the one or more optional barcode sequences, the ME sequence, the first strand cDNA fragment, and an antisense of the ME sequence, an antisense of the first 5’ PCR handle sequence, optionally an antisense of the one or more optional barcode sequences, and the antisense of the second Rn sequence; and extending the adaptor-linked intermediate second strand cDNAs using the double-barcoded first strand cDNAs as a template to form double-stranded double-barcoded cDNA fragments.
2. A method of generating double-stranded double-barcoded cDNA fragments, the method comprising, providing fixed and permeabilized cells;diffusing reverse transcriptase and nucleotides and a primer into the fixed and permeabilized cells and performing reverse transcription of RNA in the cells to form RNA / first strand cDNA hybrids; diffusing into the fixed and permeabilized cells a set of at least three (e.g., 3- 20, at least 3, 4, 5, 6, 7, 8, 9, 10, etc.) transposases, each of the at least three transposases carrying different adaptor oligonucleotides, that introduce breaks in the RNA / first strand cDNA hybrids to form RNA / first strand cDNA hybrid fragments and inserts homoadaptor oligonucleotides at the breaks, wherein the adaptor oligonucleotides comprise a first adaptor strand comprising, 5 ’ to 3 ’ : an Rn sequence, wherein Rn sequences differ for the adaptor oligonucleotides of each of the at least three transposases. optionally, one or more barcode sequences and a mosaic end (ME) sequence, and a second adaptor strand comprising an antisense ME sequence, wherein one of the transposases covalently links 3’ ends of the first adaptor strand of the adaptor oligonucleotides to 5‘ ends of each strand of the RNA / first strand cDNA hybrid fragments to form RNA / first strand cDNA hybrid fragments having 5’ overhangs; displacing the second adaptor strand and (i) extending 3 ’ ends of the first strand cDNAs with a polymerase using the RNA and first adaptor strand linked thereto as a template to form first strand cDNA fragments linked to a 5’ first Rn sequence (Ri) and a 3’ antisense second Rn sequence (R2) or (ii) extending 3’ ends of the RNAs with a polymerase that initiates from RNA using the first strand cDNAs and first adaptor strand linked thereto as a template to form RNA fragments linked to a 5’ first Rn sequence (Ri) and a 3’ antisense second Rn sequence (R2) and then performing a second reverse transcription to form the first strand cDNA fragments linked to a 5’ first Rn sequence (Ri) and a 3’ antisense second Rn sequence (R2); forming partitions comprising (i) single permeabilized cells comprising the first strand cDNA fragments and (ii) a bead linked to a plurality of clonal barcoding oligonucleotides having free 3’ ends, the barcoding oligonucleotides comprising a first 5’ PCR handle sequence, a 3’ capture sequence and a barcode sequence unique to the bead to which the barcode oligonucleotide is linked, wherein the partitions further comprise a plurality' of different bridging oligonucleotides having a 5’ end sequence and a 3’ end sequence, wherein the plurality of different bridging oligonucleotides include end sequences that anneal to or comprise all of the different Rn sequences on the transposases;optionally releasing the clonal barcoding oligonucleotides in the partitions; and in the partitions, linking barcoding oligonucleotides to both ends of the first strand cDNA fragments linked to a 5’ first Rn sequence (Ri) and a 3’ antisense second Rn sequence (R2) via annealing and polymerase extension using the bridging oligonucleotides to form double-stranded double-barcoded cDNA fragments that comprise 5 ‘-3’: the first 5’ PCR handle sequence, the barcode sequence unique to the bead, the first Rn sequence, optionally the one or more optional barcode sequences, the ME sequence, the first strand cDNA fragment, and an antisense of the ME sequence, an antisense of the first 5’ PCR handle sequence, optionally an antisense of the one or more optional barcode sequences, and the antisense of the second Rn sequence.
3. The method of claim 1 or 2, further comprising generating barcoded cDNA fragments from the double-stranded double-barcoded cDNA fragments, the generating comprising, in the partitions, optionally amplifying the double-stranded double-barcoded cDNA fragments with primers that anneal to the first 5’ PCR handle sequences; combining contents of the partitions to form a bulk solution; in the bulk solution, annealing primers comprising 5 '-3’ a second PCR handle sequence and the ME sequence to the antisense of the ME sequences on strands of the double-stranded double-barcoded cDNAs and extending the primers using a strand of the double-stranded cDNA fragments as a template with a DNA polymerase to form strands comprising 5’ -3’ the second PCR handle sequence and the ME sequence, a first strand cDNA or second strand cDNA fragment, the antisense of the ME sequence, the antisense of the first 5’ PCR handle sequence, optionally the antisense of the one or more optional barcode sequences, the antisense of the first or second Rn sequence, and the antisense of the barcode sequence unique to the bead; treating the bulk solution with an exonuclease to digest single stranded 3' overhangs, thereby forming singly barcoded cDNA fragments comprising a first and second PCR handle sequence.
4. The method of claim 3, further comprising amplifying the barcoded cDNA fragments with primers that add sequencing adapter sequences on one or both ends.
5. The method of claim 4, further comprising sequencing cDNA fragments and linked barcode sequences to form sequencing reads.
6. The method of claim 5, further comprising aligning sequencing reads to a reference genome sequence; compiling larger cDNA sequence from sequencing reads based on location of sequence reads on the reference genome, common Rn end sequences, common overlapping end sequences, common barcode sequence and / or cDNA end sequences mapping less than 10 nucleotides apart on the reference genome.
7. The method of claim 1 or 2, wherein the polymerase is Bst3.0, Bst2.0, Superscript II, or Superscript III.
8. The method of claim 1 or 2, wherein the displacing and extending is performed with a blend of polymerases, wherein the blend has RNA- and DNA-dependent DNA polymerase activity.
9. The method of claim 1 or 2. wherein the adaptor oligonucleotides comprise one or more barcodes.
10. The method of claim 9, wherein the one or more barcodes comprises a unique molecular identifier (UMI) barcode.
11. The method of claim 9 or 10, wherein the one or more barcodes comprises a sample barcode.
12. The method of any one of claims 1-11, wherein the partitions are droplets, impermeable capsules or semi-permeable capsules or microwells.
13. A plurality’ of partitions, the partitions comprising,(i) a plurality of clonal barcoding oligonucleotides having free 3 ’ ends, the barcoding oligonucleotides linked to or cleaved from a bead in the partitions, the barcoding oligonucleotides comprising a first 5’ PCR handle sequence, a 3’ capture sequence and a barcode sequence unique to the bead to which the barcode oligonucleotide is linked, wherein theplurality of clonal barcoding oligonucleotides have different 3‘ capture sequences such that the plurality includes different Rn sequences; or a plurality of clonal barcoding oligonucleotides having free 3’ ends, the barcoding oligonucleotides linked to or cleaved from a bead in the partitions, the barcoding oligonucleotides comprising a first 5' PCR handle sequence, a 3’ capture sequence and a barcode sequence unique to the bead to which the barcode oligonucleotide is linked, wherein the partitions further comprise a plurality of different bridging oligonucleotides having a 5’ end sequence and a 3’ end sequence, wherein the 3‘ end sequence anneals to the 3‘ capture sequence of the barcoding oligonucleotides and the 5’ end sequences anneal to different Rn sequences; and(ii) double-barcoded cDNA fragments that comprise 5 ’-3’: the first 5’ PCR handle sequence, the barcode sequence unique to the bead, the first Rn sequence, optionally the one or more optional barcode sequences, the ME sequence, the first strand cDNA fragment, and an antisense of the ME sequence, an antisense of the first 5’ PCR handle sequence, optionally an antisense of the one or more optional barcode sequences, and the antisense of the second Rn sequence.
14. A method of sequencing cDNA, the method comprising, providing fixed and permeabilized cells; performing reverse transcription in the cells to form RNA / first strand cDNA hybrids; diffusing into the fixed and permeabilized cells a set of at least three different transposases each of the at least three transposases carrying adaptor oligonucleotides that comprise a mosaic end (ME) sequence and different Rn sequences to form RNA / first strand cDNA hybrid fragments having 5’ overhangs; extending 3' ends of the first strand cDNAs with a polymerase using the RNA and first adaptor strand linked thereto as a template to form first strand cDNA fragments linked to a 5’ first Rn sequence and a 3’ antisense second Rn sequence; forming partitions comprising (i) single permeabilized cells comprising the first strand cDNA fragments and (ii) a bead linked to a plurality of clonal barcoding oligonucleotides having free 3’ ends, the barcoding oligonucleotides comprising a first 5’PCR handle sequence, a 3’ capture sequence and a barcode sequence unique to the bead to which the barcode oligonucleotide is linked, wherein the plurality of clonal barcoding oligonucleotides have different 3’ capture sequences such that the plurality includes all of the different Rn sequences on the transposases; optionally releasing the clonal barcoding oligonucleotides in the partitions; in the partitions forming double-stranded double-barcoded cDNA fragments that comprise 5?-3’: the first 5’ PCR handle sequence, the barcode sequence unique to the bead, a first Rn sequence, the ME sequence, a cDNA fragment sequence, and an antisense of the ME sequence, an antisense of the first 5’ PCR handle sequence, and the antisense of the second Rn sequence, wherein the forming comprises annealing and extending with a polymerase clonal barcoding oligonucleotide 3’ capture sequences to the antisense of the first Rn sequence; combining contents of the partitions to form a bulk solution; in the bulk solution, annealing primers comprising 5’ -3’ a second PCR handle sequence and the ME sequence to the antisense of the ME sequences on strands of the double-stranded double-barcoded cDNAs and extending the pnmers using a strand of the double-stranded cDNA fragments as a template with a DNA polymerase to form strands comprising 5?-3’ the second PCR handle sequence and the ME sequence, a first strand cDNA or second strand cDNA fragment, the antisense of the ME sequence, the antisense of the first 5’ PCR handle sequence, the antisense of the first or second Rn sequence, and the antisense of the barcode sequence unique to the bead; treating the bulk solution with an exonuclease to digest single stranded 3’ overhangs; optionally amplifying the barcoded cDNA fragments with primers that add sequencing adapter sequences on one or both end; and sequencing cDNA fragments and linked barcode sequences to form sequencing reads.
15. The method of claim 14, wherein the polymerase is Bst3.0, Bst2.0, Superscript II, or Superscript III.
16. The method of claim 14, wherein the displacing and extending is performed with a blend of polymerases, wherein the blend has RNA- and DNA-dependent DNA polymerase activity.
17. The method of claim 14, wherein the extending with the DNA- d ependent DNA polymerase occurs at a higher temperature than the displacing and annealing.
18. The method of claim 14, wherein the adaptor oligonucleotides comprise one or more barcodes.
19. The method of claim 18, wherein the one or more barcodes comprises a unique molecular identifier (UMI) barcode.
20. The method of claim 18 or 19, wherein the one or more barcodes comprises a sample barcode.
21. The method of any one of claims 1-11, wherein the partitions are droplets, impermeable capsules or semi-permeable capsules or microwells.
22. A method of generating cDNA fragments for sequencing, the method comprising, providing fixed and permeabilized cells; performing reverse transcription in the cells to form RNA / first strand cDNA hybrids; diffusing into the fixed and permeabilized cells a set of at least three different transposases each of the at least three (e.g., 3-20, at least 3, 4, 5, 6, 7, 8, 9, 10, etc.) transposases carrying adaptor oligonucleotides that comprise a mosaic end (ME) sequence and different Rn sequences to form RNA / first strand cDNA hybrid fragments having 5‘ overhangs comprising Rn sequences.
23. A plurality of double-stranded double-barcoded cDNA fragments comprising: a first 5’ PCR handle sequence, a barcode sequence (e.g., unique to a bead), a first Rn sequence, optionally one or more optional barcode sequences, a mosaic end (ME) sequence, a first strand cDNA fragment, and an antisense of the ME sequence, an antisense of the first 5?PCR handle sequence, optionally an antisense of the one or more optional barcode sequences, and an antisense sequence of a second Rn sequence.
24. The plurality’ of claim 23, wherein at least 4, 5,6, 7, 8, 9, or 10 different double-stranded double-barcoded cDNA fragments are present in the plurality, wherein each of the at least 4, 5,6, 7, 8, 9, or 10 double-stranded double-barcoded cDNA fragments have at least one different Rn sequence from the others of the at least 4, 5,6, 7, 8, 9, or 10 .
25. The plurality of claim 23 or claim 24, wherein the one or more optional barcodes are unique molecular identifier (UMI) barcodes, sample barcodes, or both.
Citation Information
Patent Citations
Single cell transcriptional library building and sequencing method
CN116949133A
Single tube bead-based DNA co-barcoding for accurate and cost-effective sequencing, haplotyping, and assembly
US20210115595A1
Transposome enabled DNA / RNA-sequencing (ted RNA-SEQ)
US20210222163A1
B(ead-based) a(tacseq) p(rocessing)
US20230235391A1
Dual-tagmentation single-cell dnaseq
WO2025024703A1