Method for sequencing RNA oligonucleotide

Scifi-RNA-seq addresses throughput limitations in single-cell RNA sequencing by pre-indexing cells and using microfluidics to load high cell concentrations, achieving efficient and clean sequencing data with existing platforms.

JP2025181919APending Publication Date: 2025-12-11チェムフォルシュングスツェントルンフュルモレクラーレメディツィンゲゼルシャフトミットベシュレンクテルハフツング
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025156105
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-12-16
Filing Date
2025-09-19
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

Current single-cell RNA sequencing technologies face limitations in throughput due to the need for limiting dilution to avoid cell doublets, leading to inefficiencies and high costs, and combinatorial indexing methods suffer from material loss, synthesis errors, and labor-intensive processes.

Method used

A method called scifi-RNA-seq, which involves pre-indexing the transcriptome with a first barcode before microfluidic sequencing, allowing multiple cells in a droplet to be deconvoluted using a combination of barcodes, and utilizing a microfluidic system to load cells at higher concentrations without clogging.

Benefits of technology

Scifi-RNA-seq achieves a 15- to 25-fold increase in throughput by enabling efficient sequencing of large numbers of cells without cell doublets, providing cleaner data and reduced technical noise, and is compatible with existing platforms like the 10xGenomics Chromium.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025181919000001_ABST
    Figure 2025181919000001_ABST
Patent Text Reader

Abstract

To provide a method for sequencing an oligonucleotide including RNA.SOLUTION: In a method, two index sequences are introduced into an RNA oligonucleotide. The present invention further relates to use of the method and to an apparatus used for the method. There is also provided a kit including one or more components used in the method of the present invention.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a method for sequencing an oligonucleotide comprising RNA, the method comprising the steps of: (a) providing permeabilized cells and / or nuclei comprising a first oligonucleotide comprising RNA; (b) combining the cells and / or nuclei of (a) with a second oligonucleotide comprising DNA, wherein the second oligonucleotide comprises at least a first sequence at least partially complementary to the sequence of the first oligonucleotide, a second sequence comprising an index sequence, and a third sequence comprising a primer binding site, in a first reaction compartment under conditions that allow the first sequence of the second oligonucleotide to anneal to the first oligonucleotide; (c) reverse transcribing the first oligonucleotide in the cells and / or nuclei to obtain an extended second oligonucleotide; and (d) microtransferring the cells and / or nuclei obtained in step (c) in a second reaction compartment. (e) amplifying the DNA oligonucleotides obtained in step (d); and (f) sequencing the amplified DNA oligonucleotides. The present invention further relates to the use of such methods and to devices for use therewith. Also provided are kits comprising one or more components for use in the methods of the present invention. [Background technology]

[0002] Cell atlas projects (e.g., the Human Cell Atlas (Rozenblatt-Rosen et al. (2017) Nature 550, 451-3) and single-cell CRISPR screens (e.g., using CROP-seq (Datlinger et al. (2017) Nat Methods 14, 297-301)) are reaching the limits of current technology because they require profiling millions of single cells. Most single-cell RNA-seq studies that have scaled beyond what is feasible using standard microtiter (96- or 384-well) plates are currently based either on subnanoliter well plates or on microfluidic droplet generators. Both technologies are built on a microstructuring method called soft lithography.

[0003] In subnanoliter well-based scRNA-seq (Cyto-Seq (Chen et al. (2015) Science 348, aaa6090), Seq-Well (Gierahn et al. (2017) Nat Methods 14, 395-8), Microwell-Seq (Han et al. (2018) Cell 172, 1091-1107), and sci-RNA-seq (Cao et al. (2017) Science 357, 661-7)), plates with miniaturized reaction compartments in the subnanoliter range are fabricated from materials such as PDMS and agarose. Beads and cells are loaded by gravity. Beads are typically loaded near saturation, while cells are loaded at limiting dilution (i.e., very low concentration) to avoid cells from entering the same reaction compartment. If two cells do indeed land in the same well on the plate, they will have the exact same cell barcode and will be indistinguishable in downstream analyses. Once on the plate, the cells are lysed and their transcriptomes anneal to complementary oligonucleotides on microbeads. Typically, the beads are then harvested and reverse transcription is performed in bulk. Currently, due to the lack of well-validated, readily available protocols and commercially available solutions, most laboratories prefer microfluidic droplet generators (described next).

[0004] Soft lithography is not limited to open designs such as subnanoliter well plates. When using PDMS as a material, the open side can be sealed by bonding it to a glass slide, enabling complex channel designs. This has enabled the fabrication of microfluidic droplet generators for scRNA-seq (Drop-seq (Macosko et al. (2015) Cell 161, 1202-14), inDrop (Klein et al. (2015) Cell 161, 1187-1201), and 10xGenomics Chromium (Zheng et al. (2017) Nat. Commun. 8, 14049)). A typical microfluidic device for scRNA-seq has four inputs (for cells, barcoded microbeads, reverse transcription reagents, and carrier oil) and one output (for the droplet emulsion). The reverse transcription reaction is typically performed within the droplets. While deformable beads can be loaded near saturation, cells are delivered at limiting dilution to reduce the chance of two cells landing in the same droplet. If two cells do land in the same droplet, they will receive the exact same cell barcode and thus will be indistinguishable in downstream analysis. As a result, while most droplets contain both reagents and beads and are therefore fully functional, they ultimately go unused because they contain no cells.

[0005] The throughput of sub-nanoliter well plates and microfluidic droplet generators is limited by the requirement to load cells at limiting dilution to avoid cell doublets. These platforms typically reach a throughput of approximately 10,000 cells per experiment (e.g., per sub-nanoliter well plate or per channel on a 10xGenomics Chromium chip), which can be scaled up through parallelization (multiple plates, multiple channels on a microfluidic device). However, this is often costly and labor-intensive.

[0006] With combinatorial indexing, the number of profiled cells can exponentially increase with the number of barcoding rounds. Two rounds of barcoding (using 384 × 384 barcodes) allow profiling of roughly 10,000 cells (thus requiring significant manual effort), but offer no advantage over subnanoliter well plates or droplet generators. Only when a third round of indexing is introduced does processing of more than one million cells become possible. The current largest dataset generated with sci-RNA-seq v3 contains 2 million single-cell transcriptomes from developing mouse embryos (Cao et al. (2019) Nature 566, 496–502). However, this comes with several drawbacks: (1) most NGS library preparation protocols are not readily compatible with three-round combinatorial indexing (e.g., assays such as ATAC-seq, DNA methylation profiling, and Hi-C). (2) In each barcoding step, nuclei or cells must remain intact despite the strong reaction buffers and high-temperature incubation. Three barcoding rounds typically result in material loss exceeding 90%. (3) Designing an elegant library read structure for cost-effectively sequencing three barcode combinations is challenging (this is particularly problematic when concatenated overhangs must be sequenced along with the barcodes, for example, in SPLIT-seq or sci-RNA-seq v3). (4) Synthesis and sequencing errors accumulate in barcodes, preventing reliable assignment to a larger proportion of reads. (5) Performing reactions on intact cells or nuclei is only partially efficient. The more reactions must be performed in this manner, the lower the overall efficiency of library preparation and the quality of the resulting single-cell transcriptome. (6) To achieve large cell numbers, a large number of indexes must be used in each barcoding round. As an example, 384 x 384 x 768 barcode combinations were used to generate a 2 million cell dataset.This is labor intensive and wasteful of the volumes of reagents required. Given these drawbacks, it is difficult to imagine that the published methods for combinatorial indexing scRNA-seq will be widely adopted in laboratories or become commercially successful.

[0007] In a typical experiment, a cell suspension is loaded onto the surface of a microfluidic chip along with a population of microbeads bearing unique DNA barcodes, reverse transcription reagents, and carrier oil (Figure 1a). When the aqueous and oil phases are combined at a controlled flow rate, emulsion droplets simultaneously encapsulate individual cells and individual microbeads. Due to the buffer composition, the cells lyse, releasing cellular macromolecules into the droplets. Cell transcripts anneal to complementary bead-tethered primers bearing unique cellular barcodes. For full transcriptome coverage, these primers contain oligo-dT stretches complementary to the poly-A tails in messenger RNA. However, in principle, any capture sequence can be used to selectively enrich specific transcripts or RNAs. In some embodiments, the microbeads are melted under reducing conditions or UV light for more efficient transcript capture. In most protocols, emulsion droplets are used as reaction compartments for the reverse transcription reaction, which incorporates barcodes into the cellular transcriptome.

[0008] Importantly, if two cells were to enter the same droplet or the same well on a subnanoliter well plate, their transcriptomes would be labeled with the exact same cell barcode, resulting in a cell doublet that would confound the analysis. To circumvent this issue, prior art droplet generators deliver cell suspensions at limiting dilutions, with most droplets containing zero or one cell. This makes microfluidic scRNA-seq highly inefficient. While most emulsion droplets are fully functional (containing both barcoded microbeads and reverse transcription reagents), they do not accept cells and therefore do not constitute productive library preparation events.

[0009] Thus, there is a need for improved methods for analyzing RNA oligonucleotides, particularly methods that allow for high-throughput analysis.

[0010] The technical problem is solved by the embodiments presented herein, particularly as presented in the claims. Summary of the Invention

[0011] The present invention particularly relates to the following items:

[0012] 1. A method for sequencing an oligonucleotide comprising RNA, comprising: (a) providing permeabilized cells and / or nuclei containing a first oligonucleotide comprising RNA; (b) combining the cells and / or nuclei of (a) in a first reaction compartment with a second oligonucleotide comprising DNA, wherein the second oligonucleotide comprises at least a first sequence at least partially complementary to a sequence of the first oligonucleotide, a second sequence comprising an index sequence, and a third sequence comprising a primer binding site, under conditions that allow the first sequence of the second oligonucleotide to anneal to the first oligonucleotide; (c) reverse transcribing the first oligonucleotide in the cells and / or nuclei to obtain an extended second oligonucleotide; (d) reacting the cells and / or nuclei obtained in step (c) with a third oligonucleotide bound to microbeads in a second reaction compartment, wherein the third oligonucleotide is (i) comprises a first sequence corresponding to a fourth sequence contained in the second oligonucleotide used in step (b); (ii) a fourth oligonucleotide comprising a first sequence complementary to a first sequence of the fourth oligonucleotide, the fourth oligonucleotide further comprising a second sequence at least partially complementary to the third sequence of the second oligonucleotide; For (i), the method further comprises a step of second strand DNA synthesis after step (c) and before step (d), and for (ii), the method further comprises a step of DNA ligation; the third oligonucleotide further comprises a second sequence comprising an index sequence and a third sequence comprising a primer binding site; (e) amplifying the DNA oligonucleotides obtained in step (d); and (f) sequencing the amplified DNA oligonucleotides.

[0013] 2. The method of item 1, wherein in step (c), a non-templated nucleotide is added to the 3' end of the second oligonucleotide.

[0014] 3. The method of paragraph 2, wherein second strand DNA synthesis comprises the use of a primer that contains a sequence complementary to the added non-templated nucleotide.

[0015] 4. The method of item 2, wherein a primer containing an RNA nucleotide complementary to the added non-templated nucleotide is added for extension.

[0016] 5. Second strand DNA synthesis occurs (a) introducing a nick into the first oligonucleotide; (b) extending the nicked oligonucleotide; (c) ligating the extended oligonucleotides.

[0017] 6. The method of item 1 or 5, further comprising the step of introducing a non-template nucleotide into the 5' end of the synthesized second strand DNA after or simultaneously with second strand DNA synthesis.

[0018] 7. The method of paragraph 6, wherein a transposase enzyme, particularly Tn5 transposase, is used to introduce non-templated nucleotides.

[0019] 8. The method of item 1, further comprising a step of linear extension after DNA ligation, wherein the linear extension comprises adding a primer comprising RNA nucleotides and adding a reverse transcriptase.

[0020] 9. The method of item 1, further comprising a linear extension step comprising adding a primer containing random nucleotides.

[0021] 10. The method according to any one of items 1 to 9, wherein the sequence of the first oligonucleotide to which the first sequence of the second oligonucleotide binds is located at the 3' end of this first oligonucleotide.

[0022] 11. The method of any one of paragraphs 1 to 10, wherein the first sequence of the second oligonucleotide is complementary to the 3' poly-A-tail of the first oligonucleotide.

[0023] 12. The method of any one of paragraphs 1 to 11, wherein the first reaction compartment comprises permeabilized intact cells and / or nuclei.

[0024] 13. The method of any one of paragraphs 1 to 12, wherein the first reaction compartment contains 5,000 to 10,000 cells.

[0025] 14. The method of any one of paragraphs 1 to 13, wherein the second reaction compartment contains lysed cells and / or nuclei.

[0026] 15. The method of any one of paragraphs 1 to 14, wherein the second reaction compartment comprises two or more cells and / or nuclei per microbead, preferably 10 cells and / or nuclei per microbead.

[0027] 16. The method of any one of paragraphs 1 to 15, wherein the second reaction compartment is a microfluidic droplet or well on a microtiter plate, particularly a sub-nanoliter well plate.

[0028] 17. The method of paragraph 16, wherein the second reaction compartment is a microfluidic droplet and the third oligonucleotide is released from the microbead upon formation of the droplet.

[0029] 18. The method of any one of paragraphs 1 to 17, wherein the second oligonucleotide further comprises a unique molecular identifier (UMI).

[0030] 19. The method of any one of paragraphs 1 to 18, wherein the cells / nuclei are obtained from an in vitro culture or from a fresh or frozen sample.

[0031] 20. The cell / nucleus is (a) derived from organoids or xenografts obtained from existing cell lines, primary cells, blood cells, or somatic cells; (b) is a CAR-T cell, a CAR-NK cell, an engineered T cell, a B cell, an NK cell, or an immune cell, or is isolated from a patient treated with such a product; or (c) The method of any one of paragraphs 1 to 19, wherein the stem cells are pluripotent stem cells (iPS) or embryonic stem cells that have undergone natural differentiation or artificially induced reprogramming or transdifferentiation.

[0032] 21. The method of any one of items 1 to 20, wherein the DNA ligation utilizes a thermostable DNA ligase.

[0033] 22. The method of any one of paragraphs 1 to 21, in particular using a microfluidic system for generating microfluidic droplets or for delivering materials to a microfluidic well-based device.

[0034] 23. Use of paragraph 22, wherein the microfluidic system is a droplet generator.

[0035] 24. The use of paragraph 22, wherein the microfluidic system comprises a sub-nanoliter well plate.

[0036] 25. A kit comprising a second oligonucleotide as defined in paragraph 1, preferably together with instructions for using the method of any one of paragraphs 1 to 21.

[0037] 26. The kit of paragraph 25, further comprising a transposase enzyme.

[0038] 27. The kit of paragraph 25, further comprising a second strand synthesis reagent and / or a thermostable ligase.

[0039] 28. The kit of any one of paragraphs 25 to 27, further comprising a fourth oligonucleotide.

[0040] The present invention relates to a method for sequencing an oligonucleotide comprising RNA, the method comprising the steps of: (a) providing permeabilized cells and / or nuclei comprising a first oligonucleotide comprising RNA; (b) combining the cells and / or nuclei of (a) with a second oligonucleotide comprising DNA, wherein the second oligonucleotide comprises at least a first sequence at least partially complementary to the sequence of the first oligonucleotide, a second sequence comprising an index sequence, and a third sequence comprising a primer binding site, in a first reaction compartment under conditions that allow the first sequence of the second oligonucleotide to anneal to the first oligonucleotide; (c) reverse transcribing the first oligonucleotide in the cells and / or nuclei to obtain an extended second oligonucleotide; and (d) microtransferring the cells and / or nuclei obtained in step (c) in a second reaction compartment. and (f) sequencing the amplified DNA oligonucleotides. The method of the present invention may further comprise the steps of: combining the bead-bound third oligonucleotide (wherein the third oligonucleotide (i) comprises a first sequence corresponding to the fourth sequence contained in the second oligonucleotide used in step (b); or (ii) comprises a first sequence complementary to the first sequence of the fourth oligonucleotide, the fourth oligonucleotide further comprising a second sequence at least partially complementary to the third sequence of the second oligonucleotide; for (i), the method further comprises a step of second-strand DNA synthesis after step (c) and before step (d); and for (ii), the method further comprises a step of DNA ligation; the third oligonucleotide further comprises a second sequence comprising an index sequence and a third sequence comprising a primer binding site); (e) amplifying the DNA oligonucleotides obtained in step (d); and (f) sequencing the amplified DNA oligonucleotides. The method of the present invention provided herein may also comprise the additional step of fixing the permeabilized cells and / or nuclei containing the first oligonucleotides comprising RNA. Corresponding embodiments are also provided below.

[0041] We surprisingly found that microfluidic scRNA-seq can be used to its full potential when the entire transcriptome is pre-indexed with a first barcode before the microfluidic run (Figure 1b). Even if multiple cells end up in the same droplet and receive the same second microfluidic barcode, their transcriptomes can still be deconvoluted using the first barcode. Importantly, this concept is completely different from cell hashing, which uses DNA-tagged antibodies (Stoeckius et al. (2018) Genome Biol. 19, 224) or lipids (McGinnis et al. (2019) Nature Methods 16, 619-626). In cell hashing, the cell transcriptome is never barcoded. Therefore, cell doublets can only be detected, not separated, and must be discarded in the analysis.

[0042] The method presented herein for ultra-high-throughput single-cell RNA sequencing is named scifi-RNA-seq (for single-cell combinatorial indexing with fluidic indexing RNA sequencing). The method of the present invention is an extension of prior art droplet-based scRNA-seq with single-round combinatorial pre-indexing, thereby increasing throughput by at least 15-fold, at least 20-fold, at least 25-fold, or more. This is achieved primarily because a large number of cells can be loaded into a single droplet without generating indistinguishable labeled readouts.

[0043] In scifi-RNA-seq (Figure 1b), cells or nuclei are permeabilized, and their transcriptomes are pre-indexed by reverse transcription in partitioned pools (i.e., in many physically separated bulk aliquots on a plate of microwells containing, for example, 384 pre-indexed (round 1) barcodes). The cells or nuclei containing the pre-indexed cDNA are then pooled, randomly mixed, and packaged using a microfluidic droplet generator, filling most droplets so that multiple cells or nuclei occupy the same droplet. Within the droplets, transcripts are labeled with microfluidic (round 2) barcodes. Importantly, neither of the two barcodes is dedicated to a single cell; instead, they are shared among all cells in each reaction compartment (plate well in round 1, droplet in round 2). Yet, because cells or nuclei are randomly mixed between barcoding rounds, the combination of the two barcodes uniquely identifies a single cell.

[0044] The means and methods presented herein can be used in particular with the Chromium platform ("Chromium™"), currently the most popular scRNA-seq platform, commercially available from 10xGenomics. However, the methods of the present invention can be employed to increase the throughput of any microfluidic or plate-based platform, particularly nano- and / or subnanoliter microplate-based platforms, and / or any protocol involving barcoding (e.g., combinatorial indexing protocols). For example, the methods of the present invention can be used to improve results using the Becton Dickinson Rhapsody system (see, e.g., Shum et al. (2019) Adv Exp Med Biol, 1129:63-79 / "BD Rhapsody™"). Such improvements are particularly evident with substantially larger cell / nuclei inputs and / or potential multiplexing of hundreds or thousands of samples, since the methods of the present invention do not require individual channels for evaluation. The present invention also provides cleaner data (e.g., greater single-cell purity). Furthermore, the inventors have shown that the methods of the present invention overcome various shortcomings of standard methods used in conventional systems (such as the Chromium™ platform from 10xGenomics mentioned above). These surprising improvements over prior art such as Chromium™ include, for example, reduced "background" (often resulting from artifacts in free-floating RNA or cell preparations) and / or improved (single) cell purity (as shown in particular in Figure 39, e.g., Figures 39a and / or b).

[0045] Therefore, the scifi-RNA-seq method presented herein and its variations, i.e., the method of the present invention, can be particularly useful in single-cell sequencing projects at the organ and / or organism scale (e.g., human cell atlases) and / or developmental studies at the organ and / or organism level. The method of the present invention can also be used to identify extremely rare and / or transient cell types, developmental stages, and / or cell phenotypes. Such applications can include identifying extremely rare reprogramming and / or transdifferentiation events that have previously been difficult to capture using selectable marker proteins. A further application of the method of the present invention envisions combining CRISPR single-cell sequencing (e.g., by CROP-seq, Perturb-seq, CRISP-seq, or Mosaic-seq) with whole-transcriptome and / or CRISPR gRNA readout. As a further example, the methods of the present invention can be used to perform CRISPR single-cell sequencing (e.g., by CROP-seq, Perturb-seq, CRISP-seq, or Mosaic-seq) in combination with a single transcript and CRISPR gRNA readout, or with a panel of transcripts and CRISPR gRNA readout. Furthermore, it is conceivable to combine scifi-RNA-seq with CRISPR single-cell sequencing with CRISPR activation to profile the response of the entire transcriptome or a subset of the transcriptome to perturbation. The scifi-RNA-seq method presented herein and its variations, i.e., the methods of the present invention, can also be used for drug screening and / or compound testing (e.g., testing the ability of a compound to uncover opportunities in cellular expression profiles, etc.). Thus, screening methods are also provided by the present invention. The means and methods presented herein are also useful in biological / biochemical research approaches, particularly in elucidating ligand-receptor relationships and / or signal cascades and their (cellular) consequences.

[0046] Our method, scifi-RNA-seq, can serve as a readout for CRISPR single-cell sequencing where there are many perturbations per cell, requiring ultra-high throughput to capture all possible combinations.

[0047] The methods of the present invention can be combined with single-cell ATAC-seq for an integrated transcriptome / epigenome readout. The methods of the present invention can also be combined with lineage tracing methods for an integrated readout of lineage information and / or transcriptome.

[0048] Further provided is the use of the method of the present invention, scifi-RNA-seq, for ultra-high throughput sequencing of the immune repertoire by specific enrichment of transcripts encoding B cell receptors, T cell receptors, or other related proteins (Figure 17).

[0049] Also provided is the use of the method of the invention, scifi-RNA-seq, for integrated sequencing of the transcriptome and immune repertoire.

[0050] Further provided is the use of the method of the invention, scifi-RNA-seq, and variations thereof, to identify antigen-specific reactive T cells, B cells, and / or other immune cells, e.g., by their activation signatures. Also provided is the use to detect barcoded antibodies or other biomolecules that interact with extracellular and / or intracellular partners (e.g., targets and / or antigens).

[0051] Also provided is the combination of the methods of the invention with enrichment of transcripts of interest (e.g., single transcripts, panels of transcripts, CRISPR gRNAs, feature barcodes, particularly obtained from barcoded antibodies or other biomolecules), for example, by specific PCR or transcript capture, including diagnostic applications.

[0052] The means and methods of the present invention are also useful for assessing cell-cell interactions and / or profiling cell-cell interactions. According to this embodiment of the present invention, cells are not separated but can physically interact. The cell-cell interaction allows the cells to pass through the same first reaction compartment. The interaction between the cells can be stabilized by fixation methods.

[0053] Specifically, in the first experiment, the loading capacity of the microfluidic system was tested by replacing the lysis reagent with standard EB buffer. Therefore, the number of nuclei contained in the microfluidic droplets could be counted under an optical microscope. As shown in Figure 7, 15,300, 191,250, 382,500, 765,000, and 1,530,000 cell nuclei were loaded per microfluidic channel. Surprisingly, even when loading up to 1,530,000 nuclei per channel (100 times the maximum recommended amount), all test conditions resulted in significant overloading of the device, resulting in stable droplet emulsions without clogging the microfluidic channels. When loading 1,530,000 nuclei per channel, an average of 9.6 nuclei per droplet was observed. Thus, the 10xGenomics Chromium platform was demonstrated to be able to tolerate loading concentrations 100 times greater than typically used without clogging the microfluidic channels, thereby achieving stable droplet emulsions with the desired random loading distribution.

[0054] In the second experiment, we introduced the first barcode index using the specialized library preparation method shown in Figure 2. Alternative method designs are shown in Figures 3-6. Our protocol works with permeabilized cells and / or nuclei distributed in plates, e.g., 96-well, 384-well, or 1536-well. In this representative setup, each well contained a DNA primer containing (1) an oligo-dT stretch for transcript capture, (2) a unique well-specific round 1 index, (3) an optional unique molecular identifier for PCR replication removal, (4) a primer binding site for an NGS sequencing primer, and (5) a primer binding site for linear barcoding (pR1N) in the microfluidic device. After reverse transcription, RNase H was used to introduce nicks into the template mRNA, DNA polymerase extended the nicks, and DNA ligase sealed the nicks to form double-stranded cDNA.

[0055] The next step in this representative protocol of the method of the present invention was to introduce a second defined end for the subsequent enrichment PCR reaction. This was achieved using a custom-made Tn5 transposase loaded with Illumina-compatible i7-specific adapters. Alternatives to the method of the present invention for achieving the same result include template cross-linking with reverse transcriptase, particularly when appropriate oligonucleotides are provided; random priming with Klenow exoenzyme or similar enzymes; and single-strand ligation with or without RNA base tailing.

[0056] Importantly, and advantageously over prior art methods, nuclei and / or cells remain intact throughout the entire process and are loaded onto the surface of the microfluidic device at unusually high concentrations, facilitating the loading of a large number of cells per droplet. In our method, a single microbead is encapsulated with a large number of barcoded cells / nuclei. Due to the buffer composition, the nuclei are lysed, allowing the transcriptome to anneal to the oligos tethered to the microbeads. The microfluidic droplets are then subjected to many rounds of linear extension to introduce a second (microfluidic) barcode into the transcriptome. After this reaction, the droplet emulsion is broken, and the sequencing library is enriched by PCR, thereby allowing the introduction of additional channel-specific barcodes. While both the first and second barcodes may be shared by multiple cells, the combination of these two barcodes is unique to each individual cell. During bioinformatics analysis, cells were identified by their cellular barcodes, which contain both the plate-based first barcode and the microfluidic second barcode. The combination of both led to the surprising results presented here. In particular, the results of a typical library preparation experiment are shown in Figures 13a and 13b. Sequencing metrics for the Illumina NextSeq 500 and NovaSeq 6000 platforms are shown in Figures 13c and 13d.

[0057] For several reasons, the field believed that combinatorial indexing RNA-seq could not be combined with droplet microfluidics. Most importantly, it was believed that performing reverse transcription, second strand synthesis, and tagmentation (tagging and fragmentation) on cells or nuclei would inevitably cause damage. Therefore, it was surprising and unexpected that the method of the present invention represents a significant improvement over prior art methods.

[0058] The accompanying examples demonstrate that the 10xGenomics Chromium assay can be overloaded with nuclei 100-fold higher than the maximum recommended amount. Surprisingly, stable droplet emulsions were achieved without clogging of the microfluidic channels, even at the highest loading concentrations. Detailed metrics are provided for nuclei loading across a range of high loading concentrations, demonstrating that it can be tightly controlled even at unusually high loading concentrations. For example, a stable average loading rate of 9.6 cells per droplet was achieved when loading 1.53 million nuclei per channel (100 times the maximum recommended amount). It also demonstrates that there are no physical constraints on filling droplets with nuclei. For example, loading 1.53 million nuclei per channel resulted in a 95.5% loading rate.

[0059] Furthermore, the accompanying examples demonstrate that nuclei undergoing combinatorial pre-indexing rounds are sufficiently stable to withstand the pressure and shear stress within a microfluidic device. This was unexpected, since nuclei undergo three enzymatic reactions in some cases of the present invention: reverse transcription, second-strand synthesis, and tagmentation. These steps involve high-temperature incubation and strong buffers that would be expected to compromise the integrity of the nuclei. Therefore, combining the pre-indexing step with microfluidics was not obvious. Surprisingly, the optimized workflow for scif-RNA-seq presented herein recovers pre-indexed cells / nuclei at rates comparable to standard microfluidic scRNA-seq.

[0060] The methods of the present invention constitute the first use of linear barcoding for single-cell transcriptome sequencing. In some cases, the present invention also provides the first use of a thermostable ligase for next-generation sequencing library preparation. Linear barcoding refers to the introduction of cellular barcodes by annealing to bead-tethered oligonucleotides followed by linear extension using an appropriate DNA polymerase. Linear barcoding has recently been described for single-cell ATAC-seq, but has not been suggested for scRNA-seq. Prior to the present invention, no other scRNA-seq methods have utilized linear barcoding. Through the inventions described herein, linear barcoding has been demonstrated to be effective for preparing single-cell transcriptome libraries. The resulting data is of high quality and complexity, with minimal technical noise or sequencing artifacts. Similarly, prior to the present invention, no other scRNA-seq methods have utilized a thermostable ligase. For the related methods presented herein, the use of a thermostable ligase has been demonstrated to be effective for preparing single-cell transcriptome libraries. The data obtained are of high quality and complexity, with minimal technical noise or sequencing artifacts.

[0061] By using droplet microfluidics for the second indexing, approximately 750,000 sequences can be used in the second combinatorial barcoding round in the methods of the present invention. This results in roughly 288 million (384 x 750,000) possible barcodes when using a 384-well plate for the first indexing round. Two rounds of state-of-the-art combinatorial indexing in a 384-well plate yield only 147,456 combinations. The combination of combinatorial indexing with a microfluidic droplet generator also enables scaling of NGS protocols that, by design, are not readily compatible with three rounds of indexing.

[0062] In summary, our method utilizes a pre-indexing step to barcode the entire transcriptome of a single cell prior to a microfluidic run. Our method does not suffer from the above limitations because cells are distinguishable even when contained within the same droplet. Therefore, microfluidic droplet generators (as well as sub-nanoliter well plates) can be loaded with much larger numbers of cells than in existing protocols.

[0063] Thus, the methods of the invention can be used, inter alia, as a high-content readout for saturation mutagenesis, e.g., for the experimental annotation of genetic variants in cells. The methods of the invention can also be used as a high-content readout for synthetic biology, e.g., when large numbers of synthetic DNA modules are introduced into cells (both natural and artificial).

[0064] Thus, in a first embodiment, the present invention relates to a method for sequencing an RNA-containing oligonucleotide, the method comprising the steps of: (a) providing permeabilized cells and / or nuclei containing a first RNA-containing oligonucleotide; (b) combining the cells and / or nuclei of (a) in a first reaction compartment with a second DNA-containing oligonucleotide, wherein the second oligonucleotide comprises at least a first sequence at least partially complementary to the sequence of the first oligonucleotide, a second sequence comprising an index sequence, and a third sequence comprising a primer binding site, under conditions allowing the first sequence of the second oligonucleotide to anneal to the first oligonucleotide; (c) reverse transcribing the first oligonucleotide in the cells and / or nuclei to obtain an extended second oligonucleotide; and (d) dissolving the cells and / or nuclei obtained in step (c) in a second reaction compartment. (e) amplifying the DNA oligonucleotides obtained in step (d); and (f) sequencing the amplified DNA oligonucleotides. As discussed herein, the permeabilized cells and / or nuclei containing the RNA-containing first oligonucleotide can also be fixed to cellular or nuclear structures, for example, via chemical cross-linking of the RNA to be analyzed. Details of this embodiment of the additional fixation step are also provided herein below.The fixation step may be of particular interest when fresh samples, such as unprotected cells / nuclei (e.g. material not previously formalin fixed), are analyzed according to the means and methods of the present invention.

[0065] Thus, in general, the present invention relates to methods for sequencing oligonucleotides, including RNA. The term "sequence" refers to the sequence information associated with an oligonucleotide, or any portion of an oligonucleotide that is two or more units (nucleotides) in length. The term can be used to refer to the oligonucleotide itself or a relevant portion thereof.

[0066] The information of oligonucleotide sequence, as in the method of the present invention, relates to the sequence of nucleotide bases in oligonucleotide, particularly RNA, particularly the RNA of the first oligonucleotide.For example, if oligonucleotide contains base adenine, guanine, cytosine and / or uracil, or its chemical analogue, the oligonucleotide sequence can be represented by the corresponding sequence of letters A, G, C or U.Such oligonucleotide can be sequenced using the method of the present invention.

[0067] Thus, in the first step, the method of the present invention includes providing permeabilized cells and / or nuclei containing a first oligonucleotide comprising RNA. The first oligonucleotide comprises RNA. However, the method of the present invention is not limited by the type of first oligonucleotide or RNA contained in the cells / nuclei used in the method of the present invention. Therefore, the RNA can be of any type known to those skilled in the art. The RNA may preferably be messenger RNA. It may preferably represent part or the entire transcriptome, preferably the entire transcriptome, contained in the cells / nuclei used in the method of the present invention. Therefore, the RNA contained in the first oligonucleotide is preferably in the form of messenger RNA (mRNA). As will be appreciated by those skilled in the art, mRNA generally contains a polyadenylation tail at its 3' end. Therefore, the first sequence of the second oligonucleotide is preferably at least partially complementary to the 3' end, i.e., the poly-A tail, of the first oligonucleotide. However, the method of the present invention is not limited to binding to the 3' end. Rather, the first sequence of the second oligonucleotide can be at least partially complementary to the sequence of the first oligonucleotide, and said sequence is located in the 5' direction from the 3' end of the first oligonucleotide. This is particularly useful when the target sequence is known or at least partially known.

[0068] The cells / nuclei can be present in a variety of states and can be obtained from samples of a variety of states or origins.

[0069] For example, in one embodiment, cells and / or nuclei are obtained from in vitro cultures or from fresh or frozen samples. Cells / nuclei could be obtained from preserved tissue samples, such as formalin-fixed, paraffin-embedded (FFPE) material.

[0070] Within the scope of the present invention, cells / nuclei can be of any origin, as long as they contain oligonucleotides, including RNA. For example, cells can be derived from cell lines, primary cells, blood cells, somatic cells, organoids, or xenografts. Furthermore, cells can be obtained from cell preparations used in immuno-oncology, such as CAR-T cells, CAR-NK cells, modified T cells, B cells, NK cells, or other immune cells, or isolated from patients treated with such products. Furthermore, cells can be naturally differentiated or artificially induced pluripotent stem cells (iPS) or embryonic stem cells that have undergone artificially induced reprogramming or transdifferentiation. Thus, nuclei can be derived from any of the above cells, including, for example, blood cells, somatic cells, induced pluripotent stem cells (iPS), or embryonic stem cells. As such, the methods of the present invention may be used, inter alia, in immuno-oncology (CAR-T cells, CAR-NK cells, bispecific engagers, BiTEs, immune checkpoint blockade, cancer vaccines delivered as mRNA), molecularly targeted cancer therapy, dissecting mechanisms of drug resistance and toxicity, and / or target discovery and / or validation.

[0071] In a further embodiment, the cells and / or nuclei can be obtained from biological materials used in forensic medicine, reproductive medicine, regenerative medicine, or immuno-oncology. Thus, the cells and / or nuclei can be cells / nuclei derived from tumors, blood, bone marrow aspirates, lymph nodes, and / or cells / nuclei obtained from microdissected tissues, embryonic blastomeres or blastocysts, sperm cells, amniotic fluid, or buccal swabs. The tumor cells / nuclei are preferably disseminated tumor cells / nuclei, circulating tumor cells / nuclei, or cells / nuclei from tumor biopsies. Furthermore, the blood cells / nuclei are preferably peripheral blood cells / nuclei or cells / nuclei obtained from umbilical cord blood. It is particularly preferred that the RNA oligonucleotides contained in the cells / nuclei represent the transcriptome of the cells / nuclei.

[0072] Within the scope of the method of the present invention, cells / nuclei are provided in a permeabilized state. Those skilled in the art are familiar with suitable methods for providing cells / nuclei in said state. For example, methanol permeabilization can be used for whole cells, while incomplete lysis using detergents (such as Igepal CA-630, digitonin, or Tween-20) can be used. Thus, the first reaction compartment can contain permeabilized intact cells and / or nuclei.

[0073] There is no particular limit to the number of cells in the first reaction compartment. However, the total number of cells will depend on the lengths selected for the first and second index sequences and the number of unique first and second indexes to ensure proper sample attribution. Typically, in the methods of the present invention, the first reaction compartment contains 5,000 to 10,000 cells.

[0074] In a second step of the method of the present invention, cells and / or nuclei containing a first oligonucleotide comprising RNA are combined in a first reaction compartment with a second oligonucleotide comprising DNA, wherein the second oligonucleotide comprises at least a first sequence at least partially complementary to the sequence of the first oligonucleotide, a second sequence comprising an index sequence, and a third sequence comprising a primer binding site, under conditions that allow the first sequence of the second oligonucleotide to anneal to the first oligonucleotide.

[0075] In a preferred embodiment of the invention, cells and / or nuclei comprising a first oligonucleotide comprising RNA are combined in a first reaction compartment with a second oligonucleotide comprising DNA, wherein said second oligonucleotide comprises at least a first sequence at least partially complementary to the 3' end of said first oligonucleotide, a second sequence comprising an index sequence, and a third sequence comprising a primer binding site, under conditions that allow said first sequence of said second oligonucleotide to anneal to said 3' end of said first oligonucleotide.

[0076] As described in more detail above, the method of the present invention surprisingly allows for high throughput of cells / nuclei to be analyzed / sequenced. This is due, at least in part, to the incorporation of at least two index sequences into the RNA-containing oligonucleotide to be analyzed / sequenced. The first of the at least two index sequences is introduced by combining the cells / nuclei containing the first oligonucleotide with a second oligonucleotide containing DNA, in a first reaction compartment, wherein the second oligonucleotide comprises at least a first sequence at least partially complementary to the sequence of the first oligonucleotide, a second sequence comprising an index sequence, and a third sequence comprising a primer binding site, under conditions that allow the first sequence of the second oligonucleotide to anneal to the first oligonucleotide. In one particular embodiment, the first of the at least two index sequences is introduced by combining, in a first reaction compartment, a cell / nucleus comprising said first oligonucleotide comprising RNA with a second oligonucleotide comprising DNA, wherein said second oligonucleotide comprises at least a first sequence at least partially complementary to the 3' end of said first oligonucleotide, a second sequence comprising an index sequence, and a third sequence comprising a primer binding site, under conditions that allow said first sequence of said second oligonucleotide to anneal to said 3' end of said first oligonucleotide.

[0077] Thus, a second oligonucleotide is used in the method of the present invention. The second oligonucleotide comprises DNA and at least three functional sequences / portions. The first sequence of the second oligonucleotide is at least partially complementary to the sequence of the first oligonucleotide, preferably the 3' end of the first oligonucleotide. As mentioned above, within the scope of the present invention, it is preferred that the first oligonucleotide comprising RNA, for example, comprises a polyadenylated 3' end, which is generally found in mRNA. Therefore, the first sequence of the second oligonucleotide used in the method of the present invention comprises a sequence at least partially complementary to the 3' end of the first oligonucleotide (particularly a sequence containing or consisting mainly of thymine residues). Therefore, the first sequence of the second oligonucleotide can anneal partially or completely to the 3' end of the first oligonucleotide. Therefore, a method is provided in which the first sequence of the second oligonucleotide is complementary to the 3' poly-A tail of the first oligonucleotide. However, as also presented herein, the method of the present invention is not limited to the first sequence of the second oligonucleotide being at least partially complementary to the poly-A tail of the first oligonucleotide. The first sequence of the second oligonucleotide can be at least partially complementary to the sequence extending from the 5' end to the 3' end of the first oligonucleotide.

[0078] The second sequence / portion of the second oligonucleotide comprises or consists of an index sequence. Although the term "index sequence" is known to those skilled in the art, it is surprising that an index sequence is used as a portion of the second oligonucleotide used in the methods of the present invention.

[0079] An "index sequence" according to the present invention is understood to be a sequence of nucleotides, which may be known or unknown, in which each position in the sequence has an independent and equal probability of occurrence of every nucleotide. In a preferred embodiment of the method of the present invention, the first index sequence is known, and the second index sequence can be known or unknown. The nucleotides of the index sequence can be any nucleotide (e.g., G, A, C, T, U, or chemical analogs thereof) and can be in any order. It is understood that G represents a guanyl nucleotide, A represents an adenyl nucleotide, T represents a thymidyl nucleotide, C represents a cytidyl nucleotide, and U represents a uracil nucleotide. Those skilled in the art will recognize that known oligonucleotide synthesis methods inherently result in unequal occurrence of the nucleotides G, A, C, T, or U. For example, synthesis can result in an over-representation of certain nucleotides, such as G, in a randomized DNA sequence. This can reduce the number of unique sequences predicted based on equal nucleotide occurrence. However, those skilled in the art are well aware that the total number of unique sequences contained in the second oligonucleotide used in the method of the present invention is generally sufficient to clearly identify each target RNA containing the oligonucleotide. This is because those skilled in the art are also aware that the length of the index sequence may vary depending on the expected number of first oligonucleotides. The expected number of first oligonucleotides can be derived from the number of genes expected to be expressed and / or the number of cells / nuclei expected to be analyzed / sequenced. Therefore, the possibility that nucleotides do not appear equally in the index sequence of the second oligonucleotide used in the method of the present invention is due to the unequal coupling efficiency of nucleotides in known standard oligonucleotide synthesis methods, which can be easily considered by those skilled in the art based on their general knowledge. In particular, those skilled in the art are well aware that the length of the index sequence may be increased to increase the number of unique sequences.

[0080] The third sequence of the second oligonucleotide used in the method of the present invention comprises a primer binding site. Those skilled in the art will be familiar with suitable sequences. Therefore, any sequence can be used as long as the primer used in the method of the present invention can bind to the third sequence of the second oligonucleotide used in the method of the present invention.

[0081] Within the scope of the method of the present invention, the first sequence of the second oligonucleotide is permitted to anneal to the sequence contained in the first oligonucleotide, preferably to the 3' end of the first oligonucleotide. Those skilled in the art are familiar with the conditions that allow these sequences to anneal to each other. Within the scope of the present invention, the composition of the first sequence of the second oligonucleotide promotes this annealing. That is, the first sequence of the second oligonucleotide primarily contains nucleotides complementary to nucleotides contained in the target sequence of the first oligonucleotide, and preferably constitutes the 3' end of the first oligonucleotide. In a preferred embodiment, the 3' end of the first oligonucleotide contains an adenine nucleotide, which will anneal to a thymine nucleotide contained in the first sequence of the second oligonucleotide.

[0082] In some embodiments of the invention, the second oligonucleotide further comprises a unique molecular identifier (UMI).

[0083] After annealing the first sequence of the second oligonucleotide to the first oligonucleotide, preferably the 3' end of the first oligonucleotide, the method of the present invention comprises a step of reverse transcribing the first oligonucleotide in the cell / nucleus to obtain an extended second oligonucleotide. Those skilled in the art are familiar with the means and methods available for reverse transcribing the first oligonucleotide in the method of the present invention. More specifically, the reaction will generally involve the use of a reverse transcriptase. In some embodiments of the present invention, a reverse transcriptase capable of adding non-templated nucleotides may be preferred.

[0084] Reverse transcriptases are enzymes composed of separate domains that exhibit different biochemical activities. RNA-dependent DNA polymerase activity and RNase H activity are the primary functions of reverse transcriptases, although there are variations in function (e.g., DNA-dependent DNA polymerase activity) depending on the organism from which they are derived. The reverse transcription process typically involves multiple steps:

[0085] In the presence of an annealed primer, reverse transcriptase binds to the RNA template and initiates the reaction. RNA-dependent DNA polymerase activity synthesizes a complementary DNA (cDNA) strand and incorporates dNTPs. Optional RNase H activity degrades the RNA template in the DNA:RNA complex. DNA-dependent DNA polymerase activity (if present) recognizes the single-stranded cDNA as a template and, using the RNA fragment as a primer, synthesizes second-strand cDNA to form double-stranded cDNA. Various types of reverse transcriptases can be used in the methods of the present invention, particularly enzymes with only RNA-dependent DNA polymerase activity or with RNA-dependent DNA polymerase activity combined with RNase H activity. Enzymes with all three of the above activities can also be used.

[0086] For example, the methods of the present invention can be carried out by incubating a first reaction compartment (e.g., a multi-well plate) at elevated temperatures for a given time (e.g., at about 55°C for 5 minutes or more) to disrupt RNA secondary structures. After the secondary structures are disrupted, the first reaction compartment can be placed on ice to prevent their reformation. A reaction mixture containing buffer, dNTPs, and reverse transcriptase can then be added to initiate the reverse transcription reaction. Additives such as RNase inhibitors or DTT can be added to the reaction. The reaction is preferably carried out at increasing temperatures, starting at about 4°C and gradually increasing to a temperature of about 55°C.

[0087] Some reverse transcriptases can also exhibit terminal nucleotidyl transferase (TdT) activity, resulting in the non-template addition of nucleotides to the 3' end of the synthesized DNA. TdT activity occurs only when the reverse transcriptase reaches the 5' end of the RNA template and adds excess nucleotides to the cDNA end, demonstrating specificity for double-stranded nucleic acid substrates (e.g., DNA:RNA during first-strand cDNA synthesis and DNA:DNA during second-strand cDNA synthesis). One representative reverse transcriptase with such activity is Maxima H Minus RT. Although this activity is often undesirable because the added nucleotides do not correspond to the template, the methods of the invention can include the use of such enzymes. Therefore, in a particular embodiment, the methods of the invention include a step (c) in which non-templated nucleotides are added to the 3' end of the second oligonucleotide. In a more particular embodiment of the invention, second-strand DNA synthesis can then include the use of a primer containing a sequence complementary to the added non-templated nucleotide.

[0088] Thus, after reverse transcription, the method of the invention may include a step of second strand DNA synthesis to obtain double-stranded cDNA.

[0089] After reverse transcription and / or second-strand DNA synthesis, the method of the present invention involves transferring the permeabilized cells / nuclei to a second reaction compartment. At this stage, the cells / nuclei are permeabilized but preferably remain intact and not lysed. Therefore, the method of the present invention allows for the use of permeabilized intact cells / nuclei during the first indexing reaction, whereas prior art methods include a lysis step prior to the first indexing reaction.

[0090] The second reaction compartment can be a microfluidic droplet or a microtiter plate. The microtiter plate can be a miniaturized microtiter plate. In another embodiment of the present invention, both the first and second reaction compartments can be generated by a microfluidic droplet generator or can be miniaturized plates. Within the scope of the present invention, both reaction compartments can be standard microwell plates. Exemplary plates include Seq-Well (Gierahn et al. (2017) Nature Methods 14, 395-8) or Microwell-seq (Han et al. (2018) Cell 172(5), 1091-1107).

[0091] In a second reaction compartment, the cells and / or nuclei obtained in step (c) are reacted with a third oligonucleotide bound to microbeads, wherein said third oligonucleotide is (i) comprises a first sequence corresponding to a fourth sequence contained in the second oligonucleotide used in step (b); (ii) a fourth oligonucleotide comprising a first sequence complementary to a first sequence of the fourth oligonucleotide, the fourth oligonucleotide further comprising a second sequence at least partially complementary to the third sequence of the second oligonucleotide; For (i), the method further comprises a step of second strand DNA synthesis after step (c) and before step (d), and for (ii), the method further comprises a step of DNA ligation; The third oligonucleotide is combined with a second sequence that includes an index sequence and further includes a third sequence that includes a primer binding site.

[0092] The cells / nuclei can be lysed after being transferred to the second reaction compartment, so that the second reaction compartment contains the lysed cells / nuclei.

[0093] The third oligonucleotide used in the method of the present invention contains at least three functional moieties / sequences and is first bound to a microbead. In the second reaction compartment, the microbead can be dissolved to release the third oligonucleotide. The first sequence contained in the third oligonucleotide is used to directly or indirectly target the cDNA contained in the cells / nuclei obtained in the previous step of the method to the third oligonucleotide bound to the microbead.

[0094] Whether the first sequence of the third oligonucleotide binds directly or indirectly to the cDNA depends on the presence of a second-strand DNA synthesis step prior to combining the cDNA with the microbead-bound third oligonucleotide. In one embodiment, the first sequence of the third oligonucleotide can correspond to the fourth sequence portion of the second oligonucleotide. As will be appreciated by those skilled in the art, the sequence corresponding to a portion of the second oligonucleotide will be complementary to the synthesized second-strand DNA. Therefore, this embodiment of the invention includes a second-strand DNA synthesis step after step (c) and before step (d).

[0095] In a preferred embodiment of the present invention, second strand DNA synthesis comprises introducing a nick into a first oligonucleotide; extending the nicked oligonucleotide; and ligating the extended oligonucleotide. Nicks can be introduced by adding an additional enzyme (e.g., RNase H). As detailed above, reverse transcriptase can have RNase H activity and can therefore also be used to introduce nicks into the first oligonucleotide. The nicked oligonucleotides are then extended by reverse transcriptase and / or additional enzymes (such as DNA polymerase) and then ligated to form cDNA oligonucleotides for further processing.

[0096] The method of the present invention can further comprise the step of introducing a non-template nucleotide into the 5'-end of the synthesized second strand DNA after or simultaneously with second strand DNA synthesis. The non-template nucleotide is preferably introduced using a transposase enzyme (particularly Tn5 transposase).

[0097] Transposases are enzymes that bind to the ends of transposons and catalyze their movement to another part of the genome by a cut-and-paste or replication-transcription mechanism. Transposases are classified under the EC 2.7.7. Genes encoding transposases are widespread in the genomes of most organisms and are among the most abundant known genes. One preferred transposase in the context of the present invention is the transposase (Tnp) Tn5, specifically a customized transposase. Tn5 is a member of the RNase superfamily of proteins, which includes retroviral integrases. Tn5 can be found in the bacteria Shewanella and Escherichia coli. Transposons encode antibiotic resistance to kanamycin and other aminoglycoside antibiotics. Tn5 and other transposases are significantly inactive. Because DNA transcription events are inherently mutagenic, low transposase activity is required to reduce the risk of introducing lethal mutations into the host, thereby eliminating transcribable elements. One reason Tn5 is so unresponsive is that its N- and C-termini are located relatively close to each other, tending to inhibit each other. This was revealed by the characterization of several mutations that result in hyperactive forms of the transposase. One such mutation, L372P, is at amino acid 372 in Tn5 transposase. This amino acid is typically a leucine residue in the center of the alpha helix. Replacing this leucine with a proline residue breaks the alpha helix, altering the conformation of the C-terminal domain and separating it sufficiently from the N-terminal domain to render the protein more active. Therefore, it is preferable to use such altered transposases, which have greater activity than native Tn5 transposase. Additionally, it is particularly preferred that the transposase used in the methods of the present invention be loaded with an oligonucleotide, preferably a non-templated nucleotide, to be inserted into the target double-stranded oligonucleotide.

[0098] Therefore, it is preferable to use a hyperactive Tn5 transposase and a Tn5-type transposase recognition site (Goryshin and Reznikoff, J. Biol. Chem., 273:7367 (1998)), or a MuA transposase and a Mu transposase recognition site containing R1 and R2 end sequences (Mizuuchi, K., Cell, 35:785, 1983; Savilahti, H, et al., EMBO J., 14:4893, 1995).Further examples of transcription systems that can be used in the methods of the invention include Staphylococcus aureus Tn552 (Colegio et al, J. Bacteriol, 183: 2384-8, 2001; Kirby C et al, Mol. Microbiol, 43: 173-86, 2002), Tyl (Devine & Boeke, Nucleic Acids Res., 22: 3765-72, 1994 and International Publication No. WO 95 / 23875), transposon Tn7 (Craig, NL, Science. 271: 1512, 1996; Craig, NL, Review in: Curr Top Microbiol Immunol, 204: 27-48, 1996), Tn / O and IS 10 (Kleckner N, et al, Curr Top Microbiol Immunol, 204:49-82, 1996), Mariner transposase (Lampe DJ, et al, EMBO J., 15: 5470-9, 1996), Tel (Plasterk RH, Curr. Topics Microbiol. Immunol, 204: 125-43, 1996), P element (Gloor, GB, Methods Mol. Biol, 260: 97-1 14, 2004), Tn3 (Ichikawa & Ohtsubo, J Biol. Chem. 265: 18829-32, 1990), bacterial insertion sequence (Ohtsubo & Sekine, Curr. Top. Microbiol. Immunol. 204: 1-26, 1996), retrovirus (Brown, et al, Proc Natl Acad Sci USA, 86:2525-9, 1989), and yeast retrotransposons (Boeke & Corces, Annu Rev Microbiol. 43:403-34, 1989).Further examples include IS5, TnlO, Tn903, IS91 1, and engineered versions of enzymes from the transposase family (Zhang et al, (2009) PLoS Genet. 5:el000689. Epub 2009 Oct 16; Wilson C. et al (2007) J. Microbiol. Methods 71 :332-5), and those described in U.S. Patent Nos. 5,925,545; 5,965,443; 6,437,109; 6,159,736; 6,406,896; 7,083,980; 7,316,903; 7,608,434; 6,294,385; 7,067,644; 7,527,966; and International Patent Publication No. WO2012103545, all of which are specifically incorporated by reference in their entireties.

[0099] While any buffer suitable for the transposase used can be used in the methods of the present invention, it is preferable to use a buffer that is particularly suitable for efficient enzymatic reaction of the transposase used. In this regard, a buffer containing dimethylformamide is particularly preferred for use in the methods of the present invention, especially during the transposase reaction. In addition, buffers containing alternative buffer systems, including TAPS, Tris-acetate, or similar systems, can be used. Furthermore, crowding agents such as polyethylene glycol (PEG) are particularly useful for increasing the tagmentation efficiency of very small amounts of DNA. Particularly useful conditions for the tagmentation reaction are described by Picelli et al. (2014) Genome Res. 24:2033-2040.

[0100] Transposase enzymes catalyze the insertion of nucleic acids (especially DNA) into target nucleic acids (especially target DNA). The transposases used in the methods of the present invention are loaded with oligonucleotides, which are inserted into the target nucleic acid (especially target DNA). The complex of transposase and oligonucleotide is also called a transposome. Preferably, the transposome is a heterodimer containing two different oligonucleotides for integration. In this regard, the oligonucleotide loaded onto the surface of the transposase contains multiple sequences. In particular, the oligonucleotide contains at least a first sequence and a second sequence. The first sequence is necessary for loading the oligonucleotide onto the surface of the transposase. Exemplary sequences for loading oligonucleotides onto the surface of the transposase are provided in US2010 / 0120098. The second sequence contains a linker sequence necessary for primer binding during amplification (especially during PCR amplification) and optionally further contains non-template nucleotides. Thus, the oligonucleotide containing the first and second sequences is inserted into the target nucleic acid (especially target DNA) by the transposase enzyme. The oligonucleotide may further comprise a sequence containing a barcode sequence. The barcode sequence may be a random sequence or a defined sequence. In this regard, the term "random sequence" according to the present invention is understood to mean a sequence of nucleotides in which any nucleotide occurs at each position with independent and equal probability. Random nucleotides may be any nucleotide (e.g., G, A, C, T, U) or their chemical analogs, in any order (wherein G represents a guanyl nucleotide, A represents an adenyl nucleotide, T represents a thymidyl nucleotide, C represents a cytidyl nucleotide, and U represents a uracil nucleotide). Those skilled in the art will recognize that known oligonucleotide synthesis methods may inherently lead to unequal occurrence of the nucleotides G, A, C, T, or U. For example, synthesis may result in an over-representation of nucleotides such as G in a randomized DNA sequence. This may result in a lower number of unique random sequences than would be expected based on equal occurrence of nucleotides.The oligonucleotide for insertion into a target nucleic acid (particularly DNA) can further comprise a sequencing adaptor.

[0101] Those skilled in the art are well aware that the time required for the transposase used to efficiently integrate into nucleic acids (especially DNA in a target nucleic acid (especially target DNA)) may vary depending on various parameters (buffer components, temperature, etc.). Therefore, those skilled in the art are well aware that various incubation times may be tested / applied before finding the optimal incubation time. Another possible factor is the ratio of transposomes to tagmented DNA. In this regard, optimal refers to the optimal time taking into account integration efficiency and / or the time required to perform the method of the present invention.

[0102] Alternatively, the first sequence of the third oligonucleotide can be complementary to the first sequence of a fourth oligonucleotide present in the second reaction compartment. Thus, the third oligonucleotide can comprise a first sequence complementary to the first sequence of the fourth oligonucleotide, where the fourth oligonucleotide further comprises a second sequence at least partially complementary to the third sequence of the second oligonucleotide. The presence of the fourth oligonucleotide directs the second oligonucleotide to the third oligonucleotide. In this embodiment, the second oligonucleotide is then ligated to the third oligonucleotide. As will be appreciated by those skilled in the art, in this embodiment, the second oligonucleotide comprises a 5' phosphorylation for ligation. In this embodiment, the fourth oligonucleotide is preferably blocked at its 3' end to prevent extension by DNA polymerase. Thus, in this embodiment, the method further comprises a step of DNA ligation to obtain an oligonucleotide comprising the second and third oligonucleotides. In a preferred embodiment of the present invention, the ligase is thermostable. Non-limiting examples of representative thermostable ligases include Ampligase (Lucigen) or Taq HiFi DNA Ligase (New England Biolabs). This allows for the annealing of second, third, and fourth oligonucleotides using thermal denaturation and cooling, i.e., temperature cycling, without impairing ligase activity. Specifically, emulsion droplets containing the oligonucleotides and ligase enzyme can be subjected to multiple rounds of thermal cycling between thermal denaturation and annealing, thereby enabling efficient annealing and ligation.

[0103] In the methods of the present invention, the third oligonucleotide further comprises a second sequence containing an index sequence and a third sequence containing a primer binding site. Therefore, a second index sequence is introduced in the methods of the present invention. The combined use of the first and second index sequences allows the methods of the present invention to achieve high cell / nuclei throughput. Because of the presence of two independent index sequences, the second reaction compartment in the methods of the present invention contains two or more cells / nuclei per microbead, preferably 10 cells / nuclei per microbead. Prior art methods allow for much lower throughput because the number of cells / nuclei is theoretically limited to one cell / nuclei per microbead to ensure that the RNA in the cells / nuclei accepts a unique index sequence. In practice, conventional methods are further limited for practical reasons to 0.1-0.2 cells / nuclei per microbead.

[0104] The method of the present invention further comprises the step of amplifying the DNA oligonucleotide obtained by combining the second and third oligonucleotides, optionally together with a fourth oligonucleotide, which step comprises linear extension for incorporation of a second index sequence contained in the third oligonucleotide, and amplification for sequencing.

[0105] The method of the present invention then includes the step of sequencing the amplified DNA oligonucleotides.

[0106] Those skilled in the art are familiar with the method suitable for sequencing DNA oligonucleotides.Representative non-limiting methods used to determine the sequence of oligonucleotides include, for example, nucleic acid sequencing methods (for example, Sanger dideoxy sequencing), massively parallel sequencing methods (such as pyrosequencing, reverse dye terminator, proton detection, phospho-linked fluorescent nucleotide or nanopore sequencing).

[0107] In particular, the resulting amplified oligonucleotides can be subjected to either conventional Sanger-based dideoxynucleotide sequencing methods or novel massively parallel sequencing methods ("next generation sequencing"), such as those marketed by Roche (454 technology), Illumina (e.g., Solexa technology, sequencing-by-synthesis technology), ABI (solid-state technology), Oxford Nanopore (e.g., nanopore sequencing), or Pacific Biosciences (SMRT technology). Sequencing is preferably performed using the Illumina NextSeq 500 / 550 platform, the Illumina NovaSeq 6000 platform, or the NextSeq 1000 / 2000 platform.

[0108] Various steps of the methods of the invention involve the generation and / or amplification of oligonucleotides. In addition to such reactions, sequencing reactions can involve the use of primer sequences.

[0109] Thus, the present invention relates to oligonucleotides capable of specifically amplifying the oligonucleotides of the present invention. Thus, oligonucleotides within the meaning of the present invention can function as starting points for amplification, i.e., as primers. Such oligonucleotides can contain oligoribonucleotides or deoxyribonucleotides complementary to a region of one strand of the oligonucleotide. According to the present invention, those skilled in the art will readily understand that the term "primer" can also refer to a pair of primers oriented in opposite directions relative to the complementary region of the oligonucleotide, for example, to enable amplification by polymerase chain reaction (PCR). It is generally considered that primers are purified before use in the methods of the present invention. Such purification steps can include HPLC (high-performance liquid chromatography) or PAGE (polyacrylamide gel electrophoresis) and are known to those skilled in the art.

[0110] The term "specifically" when used in the context of a primer means that the desired oligonucleotide described herein is preferably or exclusively amplified. Thus, the primer of the present invention is preferably a primer that binds to a region of the oligonucleotide that is exclusively for this molecule. With respect to a pair of primers, according to the present invention, one of the pair of primers can be specific in the above sense, or both of the pair of primers can be specific.

[0111] The 3'-OH end of the primer is utilized by the polymerase and extended by sequential incorporation of nucleotides. The primer or pair of primers of the present invention is preferably utilized for an amplification reaction on a template oligonucleotide. The term "template" refers to an oligonucleotide or fragment thereof of any source or composition that contains the target oligonucleotide sequence. It is known that the length of a primer is determined by various parameters (Gillam, Gene 8 (1979), 81-97; Innis, PCR protocol: A guide to methods and applications, Academic Press, San Diego, USA (1990)). It is preferred that a primer hybridizes or binds only to a specific region of the target oligonucleotide. The length of a primer that statistically hybridizes to only one region of the target nucleotide sequence can be calculated by the following formula: (1 / 4) x (where x is the length of the primer). However, it is known that a primer that exactly matches the complementary template strand must be at least 9 base pairs in length, otherwise a stable double strand cannot be formed (Goulian, Biochemistry 12 (1973), 2893-2901). Computer-based algorithms can be used to design primers that can amplify DNA. It is also possible to label a primer or a pair of primers. Labels can be, for example, radioactive labels ( 32 P, 33 P, or 35In a preferred embodiment of the invention, the label is a non-radioactive label (e.g., digoxigenin, biotin, and fluorescent dyes).

[0112] The present invention further relates to the use of microfluidic systems (especially microfluidic droplet generators) in the methods of the invention. Microfluidic systems can be used, in particular, to generate (microfluidic) droplets or deliver materials to well- or chamber-based devices (such as microfluidic well-based devices). Such devices are known in the art and are based, in particular, on integrated fluidic circuit technology. An example of a provider of such devices is Fluidigm Corporation / USA. Thus, the generation of (microfluidic) droplets or the delivery of materials to well- or chamber-based devices can also be part of the methods of the invention. A representative example of a droplet generator is the Chromium™ controller provided by 10xGenomics (Pleasanton, CA). Further examples include the Drop-seq platform and the inDrop platform. Furthermore, the present invention can be used to increase the throughput of sub-nanoliter well-based platforms such as CytoSeq (Fan et al., 2015), Seq-Well (Gierahn et al., 2017), Microwell-Seq (Han et al., 2018), or microfluidic systems with built-in reaction chambers. One compatible commercially available version is the BD Rhapsody™ system mentioned above, and it is in this system that the method of the present invention can be shown to provide surprising results.

[0113] The methods of the present invention can further include an additional layer of multiplexing through cell hashing.

[0114] As presented herein, the methods of the present invention can be used in synthetic biology. For example, the methods of the present invention can be used with gene panel readouts (e.g., tens to hundreds of specifically interrogated genes instead of whole transcriptome readouts). Therefore, devices are provided that utilize the methods of the present invention, single-cell RNA-seq, instead of flow cytometry, as a key diagnostic assay for cancer, immune disorders, and many other diseases (especially when combined with barcoded antibodies and / or TCR / BCR immune repertoire profiling). In a further contemplated embodiment, the methods of the present invention are combined with guide-RNA enrichment for large-scale CRISPR single-cell sequencing (e.g., CROP-seq, Perturb-seq, etc.) with hypothesis-driven gene set / pathway readouts using CRISPR knockout, CRISPR activation, CRISPR knockdown, CRISPR knockin of natural or synthetic sequences, CRISPR epigenome editing, saturation mutagenesis, or similar assays for perturbation steps.

[0115] Further provided is the method of the invention combined with ChIPmentation, as described in WO2017 / 025594 as a separate assay based on the same technology (e.g., for single-cell epigenomic profiling) or combined with guide-RNA enrichment (e.g., for epigenomic-based CROP-seq screens).

[0116] The methods of the present invention also find use in drug discovery, drug screening, compound testing, and / or target validation. Therefore, the methods of the present invention can particularly derive relevant screening signatures directly from the transcriptome of control cells, without requiring prior knowledge of the mechanism of action of the drug and / or test compound. Furthermore, the single-cell resolution of the methods of the present invention makes it possible to evaluate the effects of the drug / test compound being screened on different cell types in a complex mixture (e.g., but not limited to, PBMCs) or on a mixture of cells from different donors.

[0117] Accordingly, provided herein is a method for identifying and / or screening a test compound capable of altering the transcriptome of a cell, the method comprising: (a) contacting cells and / or nuclei containing a first oligonucleotide comprising RNA with one or more test compounds to be identified and / or screened; (b) permeabilizing said cells and / or nuclei containing said first oligonucleotide comprising RNA; (c) combining the cells and / or nuclei of (b) in a first reaction compartment with a second oligonucleotide comprising DNA, wherein the second oligonucleotide comprises at least a first sequence at least partially complementary to a sequence of the first oligonucleotide, a second sequence comprising an index sequence, and a third sequence comprising a primer binding site, under conditions that allow the first sequence of the second oligonucleotide to anneal to the first oligonucleotide; (d) reverse transcribing the first oligonucleotide in the cells and / or nuclei to obtain an extended second oligonucleotide; (e) reacting the cells and / or nuclei obtained in step (d) with a third oligonucleotide bound to microbeads in a second reaction compartment, wherein the third oligonucleotide is (i) comprises a first sequence corresponding to a fourth sequence contained in the second oligonucleotide used in step (c); (ii) a fourth oligonucleotide comprising a first sequence complementary to a first sequence of the fourth oligonucleotide, the fourth oligonucleotide further comprising a second sequence at least partially complementary to the third sequence of the second oligonucleotide; For (i), the method further comprises a step of second strand DNA synthesis after step (d) and before step (e), and for (ii), the method further comprises a step of DNA ligation; the third oligonucleotide further comprises a second sequence comprising an index sequence and a third sequence comprising a primer binding site; (f) amplifying the DNA oligonucleotide obtained in step (e); (g) sequencing the amplified DNA oligonucleotides; and (h) identifying the test compound as a compound capable of altering the transcriptome of a cell if the sequenced DNA oligonucleotide is different from the sequenced DNA oligonucleotide obtained by the method without step (a). Includes:

[0118] In the above method, the "first oligonucleotide comprising RNA" contained in the cell and / or nucleus can be natural RNA, but can also be synthetic, chimeric, and / or artificial RNA constructs (such as guide RNA and / or shRNA used in CRISPR technology), viruses, or virus-derived nucleic acids, particularly those used for gene transfer. Non-limiting examples of such "first oligonucleotide comprising RNA" include the natural transcriptome of a cell, other natural or artificial small RNAs (tRNA, snRNA, snoRNA, microRNA, rRNA, etc.), synthetic biology tools (such as riboswitches and RNA aptamers), combinations of RNAs used in CRISPR technology (such as combinations of guide RNAs or shRNAs in the same cell) (e.g., co-essentiality, combined action), libraries of synthetic genes and mutagenized synthetic genes, RNA barcodes (e.g., original sample, spatial RNA barcodes can be used to identify target locations, treatments, transgenes, RNA barcodes from lineage tracing experiments, RNA barcodes attached to antibodies expressed in a given cell, RNA barcodes indicating location on a tissue section, RNA barcodes indicating cell-cell interactions, RNA barcodes that label (e.g., via antibodies) (cell surface) proteins, intracellular proteins, or modified amino acid residues, RNA barcodes used as synthetic readers of biological processes, e.g., viral RNA to assess the infection status of cells, immune receptors (such as chimeric antigen receptors or T cell receptors), (synthetic) transcription factors, (synthetic) homing receptors, etc.

[0119] As with all means and methods provided by the present invention, the method for identifying and / or screening test compounds capable of altering the transcriptome of a cell as presented herein, which includes the step (in step (b) above of the present specification) of "permeabilizing cells and / or nuclei containing a first oligonucleotide comprising RNA," can also include an additional optional step of fixing the cells / nuclei. Fixation of cells / nuclei is known in the art and particularly, but preferably, includes chemical cross-linking (e.g., using formaldehyde or an alcohol (e.g., methanol)). This fixation step, in the context of the methods presented herein, can include fixing RNA to be analyzed in its cellular context to the surface of a structural element, such as a cell / nucleus. Such an optional fixation step also has the advantage of being able to protect / preserve the cells / nuclei and / or allow these fixed cells / nuclei to be used / analyzed at a later time. Such protection / preservation can include freezing the permeabilized and fixed cells / nuclei.

[0120] The one or more test compounds to be screened / verified / identified and / or used in the above-described methods can be selected from the group of small molecules, large molecules, RNA, DNA, and other synthetic substances (including chemical compounds and / or pharmaceuticals). However, biological materials and / or pathogens can also be the "test compounds" to be screened / identified and / or used in the methods of the present invention. Such biological materials and / or pathogens can include bacteria, viruses, fungi, and / or other biological materials (multicellular pathogens such as nematodes and jellyfish). The term "biological materials and / or pathogens" also includes parts of the materials / pathogens (e.g., proteins, peptides, nucleic acids, etc.), mixtures, extracts, etc. of such materials / pathogens. The test compound can also be a compound or a group of compounds that cause genetic perturbations (e.g., CRISPR modification and / or editing in the genome of a cell and / or nucleus).

[0121] Further non-limiting examples of "test compounds" used in the methods of the present invention include compounds that lead to alterations and / or changes in the state within a given cell (such as a change in the differentiation state or transition to apoptosis). "Test compounds" can also be mRNAs, plasmids, viral vectors, etc., introduced into cells / nuclei. Such compounds can also be used, in particular, for gene transfer. Such "encoding" nucleic acids and / or gene transfer shuttles can encode transcription factors, epigenetic regulators, kinases, homing receptors for controlling cell localization within organs or tissues, immune co-stimulatory domains (such as 41BB, CD27, CD28, OX40, CD2, or CD40L), or immune co-inhibitory domains (such as BTLA, CTLA4, LAG3, LAIR1, PD-1, TIGIT, or TIM3). Components of receptor / ligand systems (or isolated portions thereof, such as extracellular domains and / or soluble portions)) can also be used as "test compounds." Non-limiting examples of such receptor / ligand systems include, in particular, molecules of signal transduction pathways and / or immunoregulatory pathways (such as the PD-1 / PD-L1 / PD-L2 system, or the CD40 / CD40L system, B7-1, B7-2, etc.).

[0122] As is clear from the present specification and the context of the present invention, the examples of "test compounds" provided herein above are not limited to the above-mentioned "method for identifying and / or screening test compounds capable of altering the transcriptome of a cell." These "test compounds" can also be used in the general method for sequencing oligonucleotides provided herein, i.e., the scifi-RNA-seq method of the present invention and its variations.

[0123] The methods of the present invention can also combine various steps, as shown herein and in the accompanying examples. Particularly preferred are versions of the present invention, such as EXT-TN5 (Example 3), LIG-TS (Example 4), EXT-RP (Example 5), LIG-RP (Example 6), and / or EXT-TS (Example 7). Each of these versions of the means and methods of the present invention is particularly useful for increasing the number of uniquely labeled cells, and therefore throughput, compared to existing methods.

[0124] Thus, in one particular embodiment, the present invention relates to a method for sequencing an oligonucleotide comprising RNA (EXT-TN5), comprising: (a) providing permeabilized cells and / or nuclei containing a first oligonucleotide comprising RNA; (b) combining the cells and / or nuclei of (a) in a first reaction compartment with a second oligonucleotide comprising DNA, wherein the second oligonucleotide comprises at least a first sequence at least partially complementary to a sequence of the first oligonucleotide, a second sequence comprising an index sequence, and a third sequence comprising a primer binding site, under conditions that allow the first sequence of the second oligonucleotide to anneal to the first oligonucleotide; (c) reverse transcribing the first oligonucleotide in the cells and / or nuclei to obtain an extended second oligonucleotide; (d) synthesizing a second strand of DNA and introducing a non-templated nucleotide at the 5' end of the synthesized second strand DNA using a transposase enzyme, particularly Tn5 transposase; (e) combining the cells and / or nuclei obtained in step (d) with a third oligonucleotide bound to a microbead in a second reaction compartment, wherein the third oligonucleotide comprises a first sequence corresponding to a fourth sequence contained in the second oligonucleotide used in step (b), and the third oligonucleotide further comprises a second sequence comprising an index sequence and a third sequence comprising a primer binding site; (f) amplifying the DNA oligonucleotides obtained in step (e); and (g) sequencing the amplified DNA oligonucleotides Includes:

[0125] In a particular embodiment, the present invention relates to a method for sequencing RNA-containing oligonucleotides (LIG-TS), comprising: (a) providing permeabilized cells and / or nuclei containing a first oligonucleotide comprising RNA; (b) combining the cells and / or nuclei of (a) in a first reaction compartment with a second oligonucleotide comprising DNA, wherein the second oligonucleotide comprises at least a first sequence at least partially complementary to a sequence of the first oligonucleotide, a second sequence comprising an index sequence, and a third sequence comprising a primer binding site, under conditions that allow the first sequence of the second oligonucleotide to anneal to the first oligonucleotide; (c) reverse transcribing the first oligonucleotide in the cell and / or nucleus to obtain an extended second oligonucleotide, wherein a non-templated nucleotide is added to the 3' end of the second oligonucleotide; (d) combining the cells and / or nuclei obtained in step (c) with third oligonucleotides bound to microbeads in a second reaction compartment, wherein the third oligonucleotides comprise a first sequence complementary to the first sequence of the fourth oligonucleotide, the fourth oligonucleotides further comprise a second sequence at least partially complementary to the third sequence of the second oligonucleotides, and the third oligonucleotides further comprise a second sequence comprising an index sequence and a third sequence comprising a primer binding site; (e) ligating the second and third oligonucleotides using a DNA ligase, preferably a thermostable DNA ligase; (f) adding a primer containing RNA nucleotides and extending the ligated oligonucleotides by adding a reverse transcriptase; (g) amplifying the DNA oligonucleotides obtained in step (f); and (h) sequencing the amplified DNA oligonucleotides Includes:

[0126] In a particular embodiment, the present invention relates to a method for sequencing oligonucleotides comprising RNA (EXT-RP), the method comprising: (a) providing permeabilized cells and / or nuclei containing a first oligonucleotide comprising RNA; (b) combining the cells and / or nuclei of (a) in a first reaction compartment with a second oligonucleotide comprising DNA, wherein the second oligonucleotide comprises at least a first sequence at least partially complementary to a sequence of the first oligonucleotide, a second sequence comprising an index sequence, and a third sequence comprising a primer binding site, under conditions that allow the first sequence of the second oligonucleotide to anneal to the first oligonucleotide; (c) reverse transcribing the first oligonucleotide in the cells and / or nuclei to obtain an extended second oligonucleotide; (d) synthesizing a second DNA strand; (e) combining the cells and / or nuclei obtained in step (d) with a third oligonucleotide bound to a microbead in a second reaction compartment, wherein the third oligonucleotide comprises a first sequence corresponding to a fourth sequence contained in the second oligonucleotide used in step (b), and the third oligonucleotide further comprises a second sequence comprising an index sequence and a third sequence comprising a primer binding site; (f) adding a primer containing random nucleotides for linear extension; (g) amplifying the DNA oligonucleotides obtained in step (f); and (h) sequencing the amplified DNA oligonucleotides Includes:

[0127] In a particular embodiment, the present invention relates to a method for sequencing oligonucleotides containing RNA (LIG-RP), the method comprising: (a) providing permeabilized cells and / or nuclei containing a first oligonucleotide comprising RNA; (b) combining the cells and / or nuclei of (a) in a first reaction compartment with a second oligonucleotide comprising DNA, wherein the second oligonucleotide comprises at least a first sequence at least partially complementary to a sequence of the first oligonucleotide, a second sequence comprising an index sequence, and a third sequence comprising a primer binding site, under conditions that allow the first sequence of the second oligonucleotide to anneal to the first oligonucleotide; (c) reverse transcribing the first oligonucleotide in the cells and / or nuclei to obtain an extended second oligonucleotide; (e) combining the cells and / or nuclei obtained in step (d) with third oligonucleotides bound to microbeads in a second reaction compartment, wherein the third oligonucleotides comprise a first sequence complementary to the first sequence of the fourth oligonucleotide, the fourth oligonucleotides further comprise a second sequence at least partially complementary to the third sequence of the second oligonucleotides, and the third oligonucleotides further comprise a second sequence comprising an index sequence and a third sequence comprising a primer binding site; (f) ligating the second and third oligonucleotides using a DNA ligase, preferably a thermostable DNA ligase; (g) adding a primer containing random nucleotides for linear extension; (h) amplifying the DNA oligonucleotides obtained in step (g); and (i) sequencing the amplified DNA oligonucleotides Includes:

[0128] In a particular embodiment, the present invention relates to a method for sequencing an oligonucleotide comprising RNA (EXT-TS), the method comprising: (a) providing permeabilized cells and / or nuclei containing a first oligonucleotide comprising RNA; (b) combining the cells and / or nuclei of (a) in a first reaction compartment with a second oligonucleotide comprising DNA, wherein the second oligonucleotide comprises at least a first sequence at least partially complementary to a sequence of the first oligonucleotide, a second sequence comprising an index sequence, and a third sequence comprising a primer binding site, under conditions that allow the first sequence of the second oligonucleotide to anneal to the first oligonucleotide; (c) reverse transcribing the first oligonucleotide in the cell and / or nucleus to obtain an extended second oligonucleotide, wherein a non-template nucleotide is added to the 3' end of the second oligonucleotide, and a primer containing an RNA nucleotide complementary to the added non-template nucleotide is added for extension; (d) combining the cells and / or nuclei obtained in step (d) with a third oligonucleotide bound to a microbead in a second reaction compartment, wherein the third oligonucleotide comprises a first sequence corresponding to a fourth sequence contained in the second oligonucleotide used in step (b), and the third oligonucleotide further comprises a second sequence comprising an index sequence and a third sequence comprising a primer binding site; (e) amplifying the DNA oligonucleotides obtained in step (d); and (f) sequencing the amplified DNA oligonucleotides Includes:

[0129] The above-mentioned versions of the invention, such as EXT-TN5 (also shown in the accompanying Example 3), LIG-TS (also shown in the accompanying Example 4), EXT-RP (also shown in the accompanying Example 5), LIG-RP (also shown in the accompanying Example 6), and EXT-TS (also shown in the accompanying Example 7), can optionally also include the additional step of fixing the permeabilized cells and / or nuclei containing the first oligonucleotide comprising RNA before carrying out subsequent steps. Thus, if desired, an optional fixation step can be carried out after step (a) mentioned in relation to the scifi-RNA-seq method and variants thereof presented herein above.

[0130] The present invention also relates to kits, particularly research kits. The kits of the present invention include the second oligonucleotide of the present invention, preferably together with instructions for use in the method of the present invention. The kits of the present invention may further include a hyperreactive transposase, preferably loaded with an oligonucleotide, and / or reagents for second strand synthesis. The kits of the present invention may also include a ready-to-use form of the transposase enzyme. Further included may be one or more of the other oligonucleotides used in the present invention (e.g., a fourth oligonucleotide) and / or a thermostable ligase. The kits of the present invention may be particularly useful for research and for sequencing RNA molecules, etc.

[0131] In a particularly preferred embodiment of the present invention, the kits of the present invention (provided in the context of) or the methods and uses of the present invention may further comprise or be equipped with instruction manuals. For example, said instruction manuals may instruct a person skilled in the art how to use the kits of the present invention in accordance with the present invention in the diagnostic applications presented herein. In particular, said instruction manuals may comprise guides for using or applying the methods or uses presented herein.

[0132] The kits of the present invention (provided in the context) may further comprise substances / compounds and / or equipment suitable / necessary for carrying out the methods and uses of the present invention, such as solvents, diluents, and / or buffers for stabilizing and / or storing and / or enabling or terminating enzymatic reactions, compounds necessary for the uses presented herein (e.g., stabilizing and / or storing the chemicals and / or transposases included in the kits of the present invention).

[0133] Further embodiments are illustrated in the scientific part. The accompanying drawings provide an illustration of the invention. However, the experimental data shown in the examples and in the accompanying drawings should not be considered limiting. The technical information contained herein forms part of the present invention.

[0134] The present invention therefore also covers all additional features individually shown in the drawings, but which may not be described in the foregoing or following description, and single alternatives of the embodiments and features thereof described in the drawings and specification may be excluded from the subject matter of other aspects of the invention. [Brief explanation of the drawings]

[0135] [Figure 1] Single-cell combinatorial indexing with fluidic indexing (scifi) combines whole-transcriptome pre-indexing with droplet-based single-cell RNA-seq. a) Standard droplet-based scRNA-seq using microfluidic droplet generators is highly inefficient when using droplets. Most droplets contain both barcoded microbeads and reverse transcription reagents (and therefore function well), but never accept cells; furthermore, the reagents within a droplet are sufficient to barcode more than one cell. b) scifi-RNA-seq unleashes the full capabilities of microfluidic droplet generators. Prior to the microfluidic run, the whole transcriptome is pre-indexed by reverse transcription inside permeabilized cells or nuclei (round 1 barcodes are indicated by letters A–F). Pools of differently barcoded cells / nuclei are loaded, for example, at a loading rate of approximately 10 per droplet. Cells within the same emulsion droplet are labeled with the same microfluidic (round 2) barcode but can still be distinguished through their transcriptome (round 1) index.

[0136] [Figure 2]scifi-RNA-seq (version EXT-TN5) based on linear extension and custom-built Tn5 transposomes. mRNA is reverse transcribed within intact cells or nuclei. Second-strand synthesis is performed by introducing a nick into the RNA template, extending it with polymerase, and sealing the nick with ligase. The double-stranded cDNA undergoes tagmentation using custom-built i7-specific Tn5 transposomes. In the second reaction compartment, a round 2 index is introduced by linear extension with polymerase. The final library is enriched by PCR and sequenced.

[0137] [Figure 3] In scifi-RNA-seq (EXT-TS) based on linear extension and template transfer, mRNA is reverse transcribed within intact cells or nuclei. Second-strand synthesis is performed by introducing a nick into the RNA template, extending it with polymerase, and sealing the nick with ligase. In the second reaction compartment, a round 2 index is introduced by linear extension with polymerase. P7 sequencing adapters are introduced by random priming. The final library is enriched by PCR and sequenced.

[0138] [Figure 4] In linear extension and template crossover-based scifi-RNA-seq (EXT-TS), mRNA within intact cells or nuclei is reverse transcribed under conditions that allow for the addition of non-templated C bases. Template crossover extends the cDNA molecules at the 3' end. In the second reaction compartment, double-stranded cDNA is generated by extension of the TSO enrichment primer, and round 2 barcodes are introduced by extension using a polymerase. The cDNA library is then enriched by PCR and can be further processed by established methods (such as commercial or custom transposomes, adapter ligation after fragmentation, or random priming). The final library is enriched by PCR and sequenced.

[0139] [Figure 5] In scifi-RNA-seq (version LIG-TS), mRNA is reverse transcribed from intact cells or nuclei using a 5'-phosphorylated reverse transcription primer under conditions that allow for the addition of non-templated C bases. In the second reaction compartment, round 2 barcodes are introduced by ligating indexed oligonucleotides with a ligase (preferably a thermostable ligase). This ligation requires a matching bridge oligo, preferably blocked at the 3' end. Template transfer then extends the cDNA molecules at the 3' end. The cDNA library is then enriched by PCR and can be further processed by established methods (such as commercial or custom transposomes, adapter ligation after fragmentation, or random priming). The final library is enriched by PCR and sequenced.

[0140] [Figure 6] In scifi-RNA-seq (version LIG-RP), mRNA is reverse transcribed using a 5'-phosphorylated reverse transcription primer within intact cells or nuclei. In the second reaction compartment, round 2 barcodes are introduced by ligating indexed oligonucleotides with a ligase (preferably a thermostable ligase). This ligation requires a matching bridge oligo, preferably blocked at the 3' end. Random priming then introduces a P7 sequencing adapter at the 3' end. The cDNA library is then enriched by PCR and can be further processed by established methods (such as commercial or custom transposomes, adapter ligation after fragmentation, or random priming). The final library is enriched by PCR and sequenced.

[0141] [Figure 7]a) By omitting the lysis reagent, intact nuclei can be imaged within emulsion droplets, confirming the feasibility of the overloading microfluidic droplet generator. Representative droplets containing 1-10 nuclei are shown. b) Overloading increases the percentage of droplets filled with nuclei from 16.4% (10x Genomics maximum) to 95.5% (100x overloading with 1.53 million nuclei per channel). c) Overloading increases the average number of nuclei per droplet in a controlled manner while maintaining the desired random loading distribution.

[0142] [Figure 8] a) Expected doublet rate as a function of cell / nuclei loading concentration per channel for a defined set of round 1 barcodes. The cell / nuclei loading rate was modeled as a zero-excess Poisson distribution. b) Due to the large number of microfluidic round 2 barcodes, two-level scifi exceeds the barcode combinations for three-level combinatorial indexing.

[0143] [Figure 9] a) Cells / nuclei pre-processed with the scifi-RNA-seq protocol are stable during a microfluidic run. Plotting barcode rank versus sequenced reads on a logarithmic scale identifies a characteristic inflection point that separates cells / nuclei from noise. Results demonstrate that scifi-RNA-seq can efficiently recover input cells / nuclei. b) Round 1 transcriptome indexing can deconvolute a large number of cells / nuclei per droplet into their respective single-cell transcriptomes. 125,000 nuclei / cells and nuclei from a 1:1 mixture of human (Jurkat) and mouse (3T3) cells were processed and demultiplexed based on microfluidic round 2 barcodes alone (left) or a combination of round 1 and round 2 barcodes (right).

[0144] [Figure 10]a) Performance plot showing unique molecular identifiers (UMIs) per cell / nucleus as a function of sequencing coverage, with the proportion of unique reads shown as the slope. b) UMIs per cell / nucleus are plotted against the number of cells / nuclei contained in each droplet, demonstrating that a large number of cells / nuclei per droplet does not decrease library complexity.

[0145] [Figure 11-1] a) Optimization of fixation and permeabilization conditions for processing human primary T cells. A single freeze-thaw cycle had no negative impact on data quality; therefore, sampling and library preparation can be performed on different days or in different laboratories, adding to the ease and flexibility of the assay. [Figure 11-2] b) Primary human T cell nuclei were visualized in a Fuchs-Rosenthal counting chamber after reverse transcription and second-strand synthesis. Nuclei were stabilized using an optimized protocol: fixation with 4% formaldehyde, freezing at -80°C, and permeabilization with digitonin and Tween-20. c) Detected cell barcodes (x-axis) are ranked according to the number of sequenced reads per barcode (y-axis). A distinctive inflection point indicates that the dataset contains roughly 250,000 cells. At low sequencing coverage, 32,745 cells had more than 100 UMIs, and 124,474 cells had more than 50 UMIs. d) Our human primary T cell dataset contains a complex transcriptome signature. 10,000 sequence reads correspond to 1,332 UMIs and 616 genes. Neither graph is saturated, suggesting that deeper sequencing would recover more UMIs per cell.

[0146] [Figure 12-1] a) By replacing the nuclei suspension with 1x nuclei buffer and omitting reducing agent B, intact gel beads could be visualized within the emulsion droplets. The bead packing fraction based on 1,265 droplet images evaluated is shown. [Figure 12-2] b) By omitting the lysis reagent, intact nuclei could be imaged using a standard microscope. For droplets in the correct focal plane, this allows for accurate counting of nuclei per droplet. The results for loading concentrations of 15,300, 191,000, 383,000, 765,000, and 1,530,000 cells / nuclei per channel are summarized in a histogram. [Figure 12-3] b) By omitting the lysis reagent, intact nuclei could be imaged using a standard microscope. For droplets in the correct focal plane, this allows for accurate counting of nuclei per droplet. The results for loading concentrations of 15,300, 191,000, 383,000, 765,000, and 1,530,000 cells / nuclei per channel are summarized in a histogram. [Figure 12-4] b) By omitting the lysis reagent, intact nuclei could be imaged using a standard microscope. For droplets in the correct focal plane, this allows for accurate counting of nuclei per droplet. The results for loading concentrations of 15,300, 191,000, 383,000, 765,000, and 1,530,000 cells / nuclei per channel are summarized in a histogram. [Figure 12-5] c) Stable droplet emulsions were obtained under all conditions investigated despite substantial overloading of the microfluidic device. d) Computer modeling of cell / nuclei loading as a zero-excess Poisson function. [Figure 12-6] e) Nuclear loading exhibits super-Poisson properties. f) Independent estimation of cell doublet fraction through Monte Carlo simulation of scifi processes.

[0147] [Figure 13-1]a) Enrichment of a primary human T-cell library containing 250,000 cells in seven qPCR reactions. Amplification was monitored based on the SYBR green signal, and reactions were removed from the thermocycler as soon as saturation was reached (cycle 14). b) Typical size distribution of the final scifi-RNA-seq library. A library made from 250,000 primary human T-cells is shown. [Figure 13-2] c-d) Key metrics from next-generation sequencing runs on the Illumina NextSeq 500 and NovaSeq 6000 platforms. [Figure 13-3] c-d) Key metrics from next-generation sequencing runs on the Illumina NextSeq 500 and NovaSeq 6000 platforms. [Figure 13-4] e) The relationship between the percentage of occupied cluster positions and the percentage or number of reads passing the filter on the Illumina NovaSeq 6000 platform. The type of patterned flow cell (SP, S2) is color-coded. This information is intended to help users find optimal loading for scifi-RNA-seq libraries. [Figure 13-5] f) NGS performance statistics for key scifi-RNA-seq experiments.

[0148] [Figure 14-1] a) Percentage of all reads that perfectly matched plate-based round 1 barcodes or microfluidic round 2 barcodes. Calculated separately for all detected barcodes (including background) or for barcodes corresponding to real cells (top 125,000 or 250,000, depending on the experiment). b) Matching barcodes show the expected random base distribution for bases 1-11, with a fixed V (not T) base detected at position 12. Sequences that do not match the reference barcode are biased towards A. [Figure 14-2]c) The abundance of well-specific round 1 barcodes is similarly distributed across the seven scifi-RNA-seq experiments. d) Percentage of reads that aligned exclusively to the human or mouse transcriptome across a total of six scifi-RNA-seq runs. [Figure 14-3] e) In one scifi-RNA-seq experiment containing a 1:1 mixture of human (Jurkat) and mouse (3T3) cells and nuclei, nuclei perform slightly better than whole cells. f) Percentage of cell doublets in species-mixed experiments versus transcriptome purity threshold.

[0149] [Figure 15-1]a) Reverse transcription reactions were performed on 200,000 nuclei isolated from human Jurkat cells (Superscript IV without template transfer, Maxima H Minus without template transfer, and Maxima H Minus with template transfer). The number of intact nuclei was then quantified by flow cytometry using fluorescent counting beads and visualized in a bar graph. The condition Beads_Only was a negative control reaction containing only counting beads. In similar experiments, nuclei were instead resuspended in 1x Ampligase buffer (Lucigen), 1x Taq HiFi buffer (NEB), or 1x Nuclei buffer (10x Genomics) and kept at 4°C for 1 hour. Nuclei were surprisingly stable under these conditions but dissolved during thermal cycling of the thermal ligation reaction (which is expected to release cellular macromolecules into the emulsion droplets). b) In vitro transcribed polyadenylated BFP mRNA was reverse transcribed using the 5' phosphorylated scifi-RNA-seq LIG reverse transcription primer and thermally ligated using HiFi Taq ligase. Two amplicons were amplified in qPCR reactions: "Positive" is a positive control for the RT reaction, in which both primers bind to BFP, and "Test" uses the BFP-FWD primer and a partial P5 primer, allowing amplification of only successfully ligated products. Reactions were performed without a bridging oligo, with a mismatched bridging oligo, or with the correct bridging oligo. No ligation product was formed when no bridging oligo or a mismatched bridging oligo was used, demonstrating the high specificity of the thermal ligation reaction. Importantly, when the correct bridging oligo was used, the expected ligation product (indicated by the arrow) was formed. This was also true when single-cell ATAC gel beads plus reducing agent B (both from 10x Genomics) were used instead of the soluble oligonucleotide substrate. Interestingly, when the reverse transcription primer did not have a phosphate group or no ligase was used, some residual tagged product was present, likely due to annealing in the qPCR reaction, but this product was much less abundant (13.38 and 16.74 amplification cycles instead of 5.93 for the complete reaction). [Figure 15-2] c) Same experiment as in b) with a broader range of primer binding sites (indrop, dropseq, truseq), a thermostable ligase (Taq HiFi, Ampligase), and with or without reducing agent B. Top: Experiment performed with polyadenylated BFP mRNA. Bottom: Experiment performed with polyadenylated MS2-p65-HSF1 mRNA. In all cases, the desired ligation product is formed (indicated by the arrow).

[0150] [Figure 16]BFP experiments: Polyadenylated BFP mRNA was reverse transcribed with the 5'-phosphorylated scifi-RNA-seq LIG reverse transcription primer using Maxima H Minus reverse transcriptase, which adds a non-templated cytosine base when it reaches the end of the transcript. Thermal ligation was performed on the cDNA using the thermostable Taq HiFi ligase to generate tagged oligonucleotides and matching bridge oligos. Subsequently, tags were added to the 3' end of the cDNA by template crossover. Three amplicons were enriched by PCR: test_RT, a positive control for reverse transcription, uses forward and reverse primers specific for BFP; test_LIG uses a partial P5 primer and a BFP-FWD primer and can only be formed if thermal ligation is successful; and test_TS uses a partial P5 primer and a TSO enrichment primer and can only be formed if thermal ligation and template crossover thermal ligation are successful. Taken together, the experiments with BFP mRNA demonstrate the success of both tagging reactions. Total RNA experiments: The same experiments were performed with total RNA isolated from human Jurkat-Cas9-TCR cells. A cDNA library was obtained as a result of PCR amplification using partial P5 and TSO enrichment primers. This demonstrates that both tagging reactions function efficiently when using total RNA as the starting material. Single-cell experiments: Similar experiments were performed with a 1:1 mixture of nuclei isolated from human Jurkat-Cas9-TCR cells and mouse 3T3 cells. Reverse transcription reactions were performed on 10,000 intact nuclei per well in a 10 μl reaction volume. Nuclei were then pooled, enriched, and resuspended in a thermal ligation master mix using Taq HiFi or Ampligase enzymes with their corresponding reaction buffers to obtain matching crosslinked oligonucleotides. Nuclei from the reaction mixture were then encapsulated in microfluidic droplets on a 10xGenomics Chromium-controlled chip E with single-cell ATAC gel beads and Partitioning Oil (10xGenomics). After the emulsion droplets were incubated, the emulsion was broken.The cleaned sample was subjected to template transfer and then cleaned. cDNA was enriched using partial P5 primers and TSO enrichment primers. This experiment demonstrates that intact nuclei can be used as starting material and that thermal ligation can be performed within emulsion droplets.

[0151] [Figure 17-1] a) Design for enriching specific transcripts from scifi-RNA-seq libraries. While CRISPR gRNA enrichment is shown as an example, the same strategy can be used to enrich for specific transcripts (e.g., T cell and B cell immune repertoires), entire gene panels, or feature barcodes. Briefly, reverse transcription and thermal ligation steps are performed as previously described. 3'-end tagging via template transfer is not required. Instead, the P7 end of the library is introduced by PCR enrichment using transcript-specific primers with 5' extensions for next-generation sequencing. [Figure 17-2] b) Testing of four different primers specific for the hU6 promoter on CRISPR gRNA transcripts (e.g., obtained by CROP-seq (Datlinger et al., 2017)). The four primers differ in the length of the P7 extension. This experiment demonstrates that it is possible to introduce the full P7 sequencing adapter in a single-step PCR (primer hU6 full Nextera). [Figure 17-3] c) Enrichment of CRISPR gRNAs using partial P5 primers and hU6 full Nextera primers starting from cDNA obtained in a single-cell scifi-RNA-seq experiment (1:1 mixture of Jurkat-Cas9-TCR and 3T3 cells).

[0152] [Figure 18-1]Sequencing results for scifi-RNA-seq based on thermal ligation and template crossover. a) Percentage of exact matches for barcodes in round 1 and round 2. b) Experimental performance of a typical scifi-RNA-seq experiment based on thermal ligation and template crossover. Left: Reads per cell plotted against unique UMIs per cell reveals the high complexity of single-cell transcriptomes. Right: The percentage of unique reads per cell averages approximately 90% across the broad range of reads sequenced. [Figure 18-2] c) Ranked barcodes plotted against reads reveal characteristic inflection points that separate cells from background noise. In this particular experiment, 15,300 nuclei were loaded into the microfluidic device. d) Species mixture plot for a 1:1 mixture of human (Jurkat-Cas9-TCR) and mouse (3T3) nuclei.

[0153] [Figure 19] a) A 1:1 mixture of human and mouse nuclei (Jurkat and 3T3, respectively) was processed with scifi-RNA-seq, and 15,300, 383,000, and 765,000 nuclei were loaded into a single microfluidic channel of a Chromium instrument. Total detected barcodes, ranked by frequency, are plotted against the number of unique molecular identifiers (UMIs) per barcode, identifying a characteristic inflection point that separates nuclei from background noise. b) Distribution of the number of nuclei (round 1 index) per droplet (round 2 barcode) at increasing nuclei loading concentrations. The average number of nuclei per droplet and nuclei loading concentration per channel are shown.

[0154] [Figure 20]Round 1 transcriptome indexing can deconvolute a large number of nuclei per droplet into their respective single-cell transcriptomes. 765,000 pre-indexed nuclei from a mixture of human (Jurkat) and mouse (3T3) cells were processed in a single microfluidic channel and demultiplexed based on microfluidic Round 2 barcodes alone (left panel) or a combination of Round 1 and Round 2 barcodes (right panel). The percentage of interspecies collisions detected is shown by a pie chart.

[0155] [Figure 21] The ratio of UMIs per cell and unique reads per cell was plotted against the number of nuclei contained in each droplet, demonstrating that the complexity of single-cell transcriptomes does not decrease when many cells simultaneously occupy the same droplet. This analysis is based on the largest human / mouse mixed experiment, with 765,000 nuclei per microfluidic channel.

[0156] [Figure 22] a) Four human cell lines (HEK293T, Jurkat, K562, and NALM6) were processed with scifi-RNA-seq using a defined set of round 1 barcodes for each cell line. Considering only round 1 barcodes, the dataset yields the averaged pseudo-bulk RNA-seq profile of the cell lines, which is plotted here. b) 151,788 single-cell transcriptomes derived from a human cell line mixture are displayed in 2D projections using the UMAP algorithm and colored based on round 1 barcodes corresponding to cell line (left), UMI per cell (top right), or marker gene expression (bottom right).

[0157] [Figure 23]a) Heatmap showing single-cell expression levels of the top 100 specific genes for each cell line. Equal numbers of single-cell transcriptomes per cell line were randomly sampled without filtering for transcriptome quality. b) Gene set enrichment analysis of differentially expressed genes clearly identifies cell lines.

[0158] [Figure 24] a) Single-cell transcriptomes of human primary T cells processed using scifi-RNA-seq with or without T cell receptor stimulation are displayed as UMAP projections (color-coded based on stimulation status). b) Expression levels of four genes induced by TCR stimulation are overlaid on the UMAP projections.

[0159] [Figure 25] a) UMAP projection with single cells colored according to the clusters assigned by graph-based clustering using the Leiden algorithm. b) Gene set enrichment analysis of differentially expressed genes within each cluster according to panel k.

[0160] [Figure 26] a) Typical size distribution of enriched cDNA obtained using scifi-RNA-seq. b) Typical size distribution of the final scifi-RNA-seq library for next-generation sequencing.

[0161] [Figure 27] a) Distribution of DNA bases along scifi-RNA-seq sequencing reads showing characteristic sequence patterns of UMIs, round 1 barcodes, round 2 barcodes, sample barcodes, and transcripts. b) Heatmap showing sequencing quality (Qscore) for each sequencing cycle.

[0162] [Figure 28]This table summarizes all NovaSeq 6000 sequencing runs performed as part of this study. scifi-RNA-seq was thoroughly probed with NovaSeq SP, S1, and S2 reagents. The table also summarizes the percentage of reads that matched the sample (i7) barcode, pre-indexed (Round 1) barcode, and microfluidic (Round 2) barcode exactly, as well as the percentage of reads with the correct combination of all three barcodes.

[0163] [Figure 29] Nuclear recovery after pre-indexing of the whole transcriptome by reverse transcription. scifi-RNA-seq achieves high recovery rates for both cell lines and raw materials.

[0164] [Figure 30] Nuclei with pre-indexed transcriptomes were visualized under a microscope in the counting chamber before loading into the microfluidic device. Selected images show nuclei from human primary T cells.

[0165] [Figure 31]A mixture of human (Jurkat) and mouse (3T3) cells was prepared, and scifi-RNA-seq was performed on whole cells permeabilized with methanol, freshly isolated nuclei, and nuclei cryopreserved, rehydrated, and permeabilized after fixation with 1% or 4% formaldehyde. During reverse transcription in 96-well plates, each sample was assigned a specific set of round 1 barcodes. All wells were then pooled, and 15,300 cells / nuclei were loaded onto a single channel of the Chromium instrument. The following performance plots are presented: (i) ranked barcodes plotted against reads, unique molecular identifiers (UMIs), or detected genes (distinguishing single-cell transcriptomes from background noise) aligned to the mouse genome (x-axis) and human genome (y-axis); (ii) reads plotted against UMIs; (iii) reads plotted against the number of detected genes; (iv) reads plotted against the proportion of unique reads; and (v) a species mixture plot showing the number of UMIs per cell. To facilitate comparisons between different types of input material, the axes of the performance plots use the same scale across conditions.

[0166] [Figure 32] 15,300 pre-indexed nuclei from a mixture of human (Jurkat) and mouse (3T3) cells are processed within a single microfluidic channel and demultiplexed based on microfluidic round 2 barcodes alone (left) or a combination of round 1 and round 2 barcodes (right). At the standard loading concentration of the Chromium instrument (15,300 nuclei per channel), the microfluidic (round 2) index provides sufficient complexity to resolve single cells, while the combination of round 1 and round 2 barcodes again reduces background noise.

[0167] [Figure 33]Coverage along human and mouse transcripts from 200 bp upstream of the transcription start site (TSS) to 200 bp downstream of the transcription termination site (TES) is shown for methanol-permeabilized whole cells, freshly isolated nuclei, and nuclei cryopreserved, rehydrated, and permeabilized after fixation with 1% or 4% formaldehyde. Freshly isolated nuclei show the strongest 3' enrichment.

[0168] [Figure 34] Boxplots summarizing sequence alignment metrics for different types of input material: total reads sequenced, percentage of uniquely mapped reads, percentage of multiple mappings, percentage of alignments to exons + introns, percentage of alignments to exons, and percentage of spliced ​​reads. Freshly isolated nuclei showed the best performance for these alignment metrics.

[0169] [Figure 35] Principal component analysis of a scifi-RNA-seq experiment on a 1:1:1:1 mixture of four uniquely characterized human cell lines. a) Variance explained by the top 30 principal components. b) Principal component analysis (PCA) projection of 151,788 single cells, color-coded using the number of UMIs per cell (top row) and round 1 barcode representing the cell line.

[0170] [Figure 36-1] The expression values ​​of 72 additional cell lineage-specific genes were mapped onto the UMAP projection as shown in FIG. [Figure 36-2] The expression values ​​of 72 additional cell lineage-specific genes were mapped onto the UMAP projection as shown in FIG. [Figure 36-3] The expression values ​​of 72 additional cell lineage-specific genes were mapped onto the UMAP projection as shown in FIG. [Figure 36-4]The expression values ​​of 72 additional cell lineage-specific genes were mapped onto the UMAP projection as shown in FIG.

[0171] [Figure 37] Principal component analysis of scifi-RNA-seq experiments on primary human T cells with or without T cell receptor stimulation. a) Variance explained by the top 30 principal components. b) PCA projections of 62,558 single cells. From top to bottom, the following variables were mapped onto these projections: logarithm of UMI per cell, cluster ID, donor ID, and T cell receptor (TCR) stimulation status.

[0172] [Figure 38] UMAP projections of 62,558 single cells (as shown in Figure 24), with additional variables mapped onto these projections: donor ID, logarithm of UMIs per cell, logarithm of detected genes per cell, proportion of unique reads per cell, mitochondrial expression rate, and ribosomal expression rate.

[0173] [Figure 39]a) Equal mixtures of four human cell lines (HEK293T, Jurkat, K562, and NALM6) were processed in parallel by scifi-RNA-seq and 10xGenomics v3 profiling using either intact cells / nuclei or methanol-fixed cells as input. To allow direct comparison between platforms, a standardized concentration of 7,500 cells / nuclei was loaded per microfluidic channel. To assess cell / nuclei recovery, all detected barcodes ranked by frequency were plotted against the number of unique molecular identifiers (UMIs) per barcode. b) Dimensionality reduction (UMAP) and clustering using the Leiden algorithm readily identified the four cell lines across all samples. For the Chromium line, an additional false cluster (gray) representing a mixture of cell lines was detected, which is completely absent in the scifi-RNA-seq data. c) Cell lines are recovered at the same rate despite their significantly different transcript contents. d) Clustering of gene expression profiles based on Pearson correlation allowed samples to be grouped by cell line regardless of the technology or cell preparation method used.

[0174] [Figure 40-1]a) Cas9-expressing human Jurkat cells were transduced in an array format with lentiviral constructs encoding 48 different gRNAs. After efficient genome editing, samples were split and stimulated with anti-CD3 / CD28 beads to activate the T cell receptor (TCR) or left untreated. Plates were processed with scifi-RNA-seq, labeled with CRISPR perturbations, and processed with specific round 1 reverse transcription barcodes. This proof-of-concept screen demonstrates the potential of scifi-Round 1 multiplexing for genetic perturbation and drug screens across hundreds to thousands of conditions, making it useful for drug development. b) Principal component analysis of 96 bulk transcriptomes, colored by treatment and labeled with genetic perturbations. Key activators of the TCR pathway are highlighted by circles. c) The top 300 genes differentially expressed between stimulated and unstimulated control cells were used as a screening signature. A heatmap was generated for this gene set (data not shown). Genetic perturbations were assigned a TCR activation score based on the expression of these genes. Samples were sorted by TCR activation score. Some gene knockouts result in a lower TCR activation score, similar to unstimulated samples. d) Transcriptome-based TCR activation scores were plotted against amplification scores derived from cell counts. [Figure 40-2] e) Single-cell transcriptomes derived from a CRISPR screen are displayed in 2D projections using the UMAP algorithm and colored based on TCR treatment. f) Cells assigned to control gRNA or gRNAs targeting ZAP70, LCK, or LAT are highlighted in black. g) Enrichment of gRNAs in the Leiden cluster is identified as the stimulated / unstimulated cluster. gRNAs targeting ZAP70, LAT, and LCK are highlighted with circles.

[0175] [Figure 41-1]a) The droplet overloading experiment was repeated on the Chromium NextGEM platform. By omitting the lysis reagent, the nuclei remained intact, allowing for counting of nuclei per droplet by imaging using a standard microscope. Histograms show results for loading concentrations of 15,300, 191,000, 383,000, 765,000, and 1,530,000 nuclei per channel. For each loading concentration, the number of droplet images evaluated, the droplet fill ratio, and the average number of nuclei per droplet are shown. Furthermore, by replacing the nuclei suspension with 1x nuclei buffer and omitting reducing agent B, intact gel beads were observed within the emulsion droplets. The bead fill ratio based on 1,610 droplet images evaluated is shown. [Figure 41-2] b) Despite substantial droplet overloading, stable droplet emulsions were obtained under all conditions investigated. c) Droplet diameters were compared between scATAC 1.0 and scATAC 1.1 (NextGEM platform) for increasing loading concentrations. 100 droplets were measured per condition. d) Droplet diameters shown as a histogram. Data at different loading concentrations were pooled for a total of 500 droplets per platform. e) Nuclei loading exhibits Poisson-like distribution characteristics. Mean values ​​are plotted on the x-axis and variance on the y-axis. [Figure 41-3] f) Computer modeling of nuclei loading as a zero-excess Poisson function. g) Posterior probability distributions of lambda and psi sampled with Markov Chain Monte Carlo (MCMC). h) Droplet overloading increases the fraction of droplets filled with nuclei for the NextGEM platform. i) Droplet overloading increases the average number of nuclei per droplet in a controlled manner while maintaining the desired Poisson-like loading distribution. [Figure 41-4]j) Expected collision rate as a function of cell / nuclei loading concentration per channel for standard Chromium profiling and a defined set of round 1 barcodes. The cell / nuclei loading rate was modeled as a zero-excess Poisson distribution.

[0176] [Figure 42-1] a) Frequency-ranked cell barcodes versus UMIs per cell. b) Reads per cell are plotted against UMIs per cell to assess the level of sequencing saturation. [Figure 42-2] c) Reads per cell are plotted against the ratio of unique reads per cell to assess PCR replication and library complexity. d) Alignment to the human genome versus the mouse genome. [Figure 42-3] e) Alignment metrics are compared between scATAC 1.0 and 1.1 (NextGEM) platforms. [Figure 42-4] f) Frequency-ranked cell barcodes versus UMIs per cell. g) Reads per cell are plotted against UMIs per cell to assess the level of sequencing saturation. [Figure 42-5] h) To assess PCR replication and library complexity, reads per cell are plotted against the rate of unique reads per cell. i) Alignment metrics for scifi-RNA-seq using Maxima H Minus are compared with Superscript IV reverse transcriptase for the reverse transcription step. Template switching was performed in both cases using Maxima H Minus reverse transcriptase.

[0177] [Figure 43-1]Equal mixtures of four human cell lines (HEK293T, Jurkat, K562, and NALM-6) were processed in parallel with scifi-RNA-seq and the Chromium v3 single-cell gene expression kit. a) Single-cell transcriptomes are displayed in 2D projections using the UMAP algorithm, with the number of UMIs per cell mapped on top. [Figure 43-2] b) Single cell clustering using the Leiden algorithm, with cluster IDs mapped onto the UMAP projection. [Figure 43-3] c) Enrichment of cell lineage signatures obtained from the ARCHS4 database for identified Leiden clusters. These results can be used to label clusters with their respective cell lineages and to identify pseudo-clusters of doublet cells. [Figure 43-4] d) Overlap rate of the top 100 differentially expressed genes between samples.

[0178] [Figure 44-1] Technology comparison between scifi-RNA-seq and existing multi-round combinatorial indexing approaches or the 10xGenomics Chromium platform. For this comparison, publicly available combinatorial indexing data was obtained, including data from Cao et al., 2017. The Cao et al., 2017 dataset is highlighted in this figure. A species mixture of human Jurkat cells and mouse 3T3 cells was processed in parallel with the method of the present invention and the 10xGenomics Chromium workflow. a) Detected cell barcodes, ranked by frequency, are plotted against the number of unique molecular identifiers (UMIs) per barcode. b) UMI counts are summarized as a bar graph. c) Reads per cell are plotted against UMIs per cell to assess sequencing saturation. d) UMI / read ratio as a metric for PCR replication. e) Reads per cell are plotted against the rate of unique reads per cell. f) The rate of unique reads is summarized as a bar graph. [Figure 44-2] g) Comparison of alignments to the human genome and the mouse genome. [Figure 44-3] h) Barcoding combinations in the largest experiment actually performed versus the total number of sequencing cycles used in that experiment. The gray line represents the 138 sequencing cycles included in the NovaSeq 100-cycle kit. i) Sequencing cycles used to read the composite cell barcode (excluding UMIs). Non-informative sequencing cycles from ligated overhangs, primer binding sites, and transposase mosaic ends are shown in gray, providing the percentage of non-informative sequencing cycles. In summary, we consistently demonstrate that our method achieves superior data quality compared to the Cao et al. (2017) method and all other published combinatorial indexing methods. scifi-RNA-seq also provides at least a 15-fold increase in cell throughput compared to 10xGenomics Chromium.

[0179] [Figure 45-1] a) Diffusion map of 96 bulk transcriptomes (48 CRISPR knockouts, 2 treatments), colored by treatment and labeled by genetic perturbation. Key regulators of the T cell receptor (TCR) pathway are highlighted with circles. Knockout of ZAP70, LAT, and LCK brings cells closer to the unstimulated sample. b) TCR activation signature, as defined in Figure 3c, mapped onto a schematic of TCR pathway activation. [Figure 45-2] c) Enrichment of cells with the indicated gRNA in the stimulated group over the unstimulated group, which is a measure of proliferation, as opposed to TCR activation, which we defined based on the transcriptome. DETAILED DESCRIPTION OF THE INVENTION

[0180] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Although methods and materials similar or equivalent to those described herein can be used to practice or test the present invention, suitable methods and materials are described below. In case of conflict, the present specification, including definitions, will control. Additionally, the materials, methods, and examples are illustrative only and not intended to be limiting.

[0181] The methods and techniques of the present invention will generally be carried out according to conventional methods well known in the art and described in various general and more specific references cited and discussed throughout the specification, unless otherwise indicated. See, e.g., Sambrook et al., Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (1989); Ausubel et al., Current Protocol in Molecular Biology, Greene Publishing Associates (1992); and Harlow and Lane Antibodies: A Laboratory Manual, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, NY (1990).

[0182] While the present invention has been shown and described in detail in the drawings and foregoing description, such illustration and description are to be considered illustrative or representative, and not restrictive. Those skilled in the art will recognize that changes and modifications are possible within the scope and spirit of the following claims. In particular, the present invention covers further embodiments having any combination of features from the various embodiments described above and below.

[0183] The present invention also covers all additional features shown individually in the drawings, but which may not be described above or in the following description, and single alternatives of the embodiments and features thereof described in the drawings and description may be excluded from the subject matter of other aspects of the invention.

[0184] Furthermore, the term "comprising" in the claims does not exclude other elements or steps, and the indefinite article "a" does not exclude a plurality. A single unit may fulfill the functions of several features recited in a claim. Terms such as "substantially," "about," and "roughly," particularly relating to an attribute or value, may also specify exactly that attribute or exactly that value, respectively. Any reference markings in the claims should not be construed as limiting the scope.

[0185] Example 1 – Cell / Nuclei Preparation

[0186] 1.1 Preparation of permeabilized whole cells from human and mouse cell lines

[0187] Five million cells were washed with 10 ml of ice-cold 1x PBS (Gibco catalog no. 14190-094, centrifugation: 300 rcf, 5 min, 4°C) and fixed in 5 ml of ice-cold methanol (Fisher Scientific catalog no. M / 4000 / 17) for 10 min at -20°C. After two additional washes (centrifugation: 300 rcf, 5 min, 4°C) with 5 ml of ice-cold PBS-BSA-SUPERase (1x PBS supplemented with 1% w / v BSA (Sigma catalog no. A8806-5) and 1% v / v SUPERase-In RNase inhibitor (Thermo Fisher Scientific catalog no. AM2696)), the permeabilized cells were resuspended in 200 μl of ice-cold PBS-BSA-SUPERase and filtered through a cell strainer (40 μM or 70 μM depending on cell size). A 10 μl sample was counted using a CASY instrument (Scharfe System) and diluted to 5,000 cells per μl with ice-cold PBS-BSA-SUPERase, which was then immediately subjected to the reverse transcription step.

[0188] 1.2 Preparation of fresh nuclei from human and mouse cell lines

[0189] Five million cells were washed with 10 ml of ice-cold 1×PBS (Gibco Cat. No. 14190-094, 300 rcf, 5 min, 4° C.). Nuclei were prepared by resuspending cells in 500 μl of ice-cold nuclei preparation buffer (10 mM Tris-HCl pH 7.5 (Sigma catalog no. T2944-100ML), 10 mM NaCl (Sigma catalog no. S5150-1L), 3 mM MgCl2 (Ambion catalog no. AM9530G), 1% w / v BSA (Sigma catalog no. A8806-5), 1% v / v SUPERase-In RNase inhibitor (Thermo Fisher Scientific catalog no. AM2696), 0.1% v / v Tween-20 (Sigma catalog no. P7949-500ML), 0.1% v / v IGEPAL CA-630 (Sigma catalog no. I8896-50ML), 0.01% v / v digitonin (Promega catalog no. G944A)) and incubating on ice for 5 minutes. Plasma membrane lysis was stopped by adding 5 ml of ice-cold nuclei wash buffer (10 mM Tris-HCl pH 7.5, 10 mM NaCl, 3 mM MgCl2, 1% w / v BSA, 1% v / v SUPERase-In RNase inhibitor, 0.1% v / v Tween-20). Nuclei were collected by centrifugation (500 rcf, 5 min, 4°C), resuspended in 200 μl of ice-cold PBS-BSA-SUPERase (1x PBS supplemented with 1% w / v BSA and 1% v / v SUPERase-In RNase inhibitor (20 U / μl, catalog number)), and filtered through a cell strainer (40 μM or 70 μM depending on cell size). A 10 μl sample was counted using a CASY instrument (Scharfe System) and diluted to 5,000 cells per μl with ice-cold PBS-BSA-SUPERase, which was then immediately subjected to the reverse transcription step.

[0190] 1.3 Preparation of nuclei from primary cells using formaldehyde fixation and permeabilization

[0191] Five million primary cells were washed with 10 ml of ice-cold 1x PBS (Gibco Cat. No. 14190-094, centrifugation: 300 rcf, 5 min, 4°C). Nuclei were prepared by resuspending cells in 500 μl of ice-cold nuclei preparation buffer (10 mM Tris-HCl pH 7.5 (Sigma Catalog No. T2944-100ML), 10 mM NaCl (Sigma Catalog No. S5150-1L), 3 mM MgCl2 (Ambion Catalog No. AM9530G), 1% w / v BSA (Sigma Catalog No. A8806-5), 1% v / v SUPERase-In RNase Inhibitor (Thermo Fisher Scientific Catalog No. AM2696), 0.1% v / v IGEPAL CA-630 (Sigma Catalog No. I8896-50ML)) without digitonin and without Tween-20, followed by incubation on ice for 5 minutes. Plasma membrane lysis was stopped by adding 5 ml of Nuclei Wash Buffer (10 mM Tris-HCl pH 7.5, 10 mM NaCl, 3 mM MgCl2, 1% w / v BSA, 1% v / v SUPERase-In RNase inhibitor) without Tween-20. Nuclei were collected by centrifugation (500 rcf, 5 min, 4°C) and fixed for 15 min in 5 ml of ice-cold 1x PBS containing 4% formaldehyde (Thermo Fisher Scientific catalog number 28908) on ice. Fixed nuclei were collected (500 rcf, 5 min, 4°C), and the pellet was resuspended in 1.5 ml of ice-cold Nuclei Wash Buffer without Tween-20 and transferred to a 1.5 ml tube. After another wash with 1.5 ml of ice-cold nuclei wash buffer without Tween-20 (500 rcf, 5 min, 4°C), the fixed nuclei were resuspended in 200 μl of nuclei wash buffer without Tween-20, flash-frozen in liquid nitrogen, and stored at −80°C.

[0192] For processing with scifi-RNA-seq, frozen samples were thawed for exactly 1 minute in a 37°C water bath and immediately placed on ice. After centrifugation (500 rcf, 5 minutes, 4°C), fixed nuclei were resuspended in 250 μl of ice-cold permeabilization buffer (10 mM Tris-HCl, 10 mM NaCl, 3 mM MgCl2, 1% w / v BSA, 1% v / v SUPERase-In RNase inhibitor, 0.01% v / v digitonin (Promega catalog no. G944A), 0.1% v / v Tween-20 (Sigma catalog no. P7949-500ML)). After a 5-minute incubation on ice, 250 μl of Nuclei Wash Buffer without Tween-20 was added per sample, and nuclei were collected (500 rcf, 5 minutes, 4°C). After another wash with 250 μl of Tween-20-free nuclei wash buffer, nuclei were taken up in 100 μl of 1× PBS containing 1% w / v BSA and 1% v / v SUPERase-In RNase inhibitor. A 5 μl sample was used for cell counting using a CASY instrument (Scharfe Systems) and diluted to 5,000 cells per μl with PBS-BSA-SUPERase. It was then immediately subjected to the reverse transcription step.

[0193] Example 2 - Device Testing

[0194] 2.1 Testing the kernel loading capabilities of the Chromium controller

[0195] Human Jurkat cells (clone E6-1) were cultured in RPMI medium (Gibco catalog no. 21875-034) supplemented with 10% FCS (Sigma) and penicillin-streptomycin (Gibco catalog no. 15140122). Fresh nuclei were isolated as described above. Samples of 15.3k, 191k, 383k, 765k, and 1.53M nuclei were then prepared and 1.5 μl of Reducing Agent B (10x Genomics catalog no. 2000087) and 1x Nuclei Buffer (10x Genomics catalog no. 2000153) were added to a total volume of 80 μl. Because this buffer does not contain detergents, nuclei remained intact throughout the microfluidic run and could be visualized within the emulsion droplets using a standard light microscope. Simultaneously, reducing agent B dissolved the gel beads, which would otherwise have obscured the field of view. A microfluidic chip (Single Cell E-Chip, 10xGenomics 2000121) was loaded as follows: 75 μl of nuclear sample at the indicated loading concentration was loaded into inlet 1, 40 μl of single-cell ATAC gel beads (10xGenomics catalog no. 2000132) into inlet 2, and 240 μl of Partitioning Oil (10xGenomics catalog no. 220088) into inlet 3. To image the resulting droplets, 15 μl of Partitioning Oil was pipetted onto the surface of a glass slide, followed by a 5 μl emulsion droplet. Images were captured at 10x magnification. An average of 653 droplets were counted per condition.

[0196] 2.2 Measuring the bead filling rate of the Chromium control device

[0197] To measure bead packing, a single-cell E-chip (10xGenomics 2000121) was loaded with 80 μl of 1× Nuclei Buffer (10xGenomics catalog no. 2000153) through inlet 1, 40 μl of single-cell ATAC gel beads (10xGenomics catalog no. 2000132) through inlet 2, and 240 μl of Partitioning Oil (10xGenomics catalog no. 220088) through inlet 3. Removal of reducing agent B ensured that the gel beads remained intact throughout the microfluidic run, allowing them to be visualized within the emulsion droplets using a standard light microscope. Packing calculations are based on a total of 1,265 droplets.

[0198] Example 3 – scifi-RNA-seq based on linear extension and custom-made Tn5 transposomes (version EXT-TN5)

[0199] Reverse transcription: Sets of 96 and 384 indexed reverse transcription primers were synthesized by Sigma-Aldrich and shipped at 100 μM in EB buffer in 96-well plates. The primers had the sequence (5'-TCGTCGGCAGCGTCGGATGCTGAGTGATTGCTTGTGACGCCTTCNNNNNNNN N XXXXXXXXXXX VThe primers had the following structure: TTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTVN-3' (where N represents a random base, the underlined bases are known for a given primer, and X represents a primer-specific index sequence 11 bases in length). A 96-well plate with barcoded oligo-dT primers was prepared prior to the experiment and stored at -20°C (1 μl of 25 μM per well). 10,000 permeabilized cells or nuclei (2 μl of a 5,000 cell / μl suspension) were added to the pre-dispensed primers, and the well assignments were recorded. The plate was incubated at 55°C for 5 minutes (to disrupt RNA secondary structures) and then immediately placed on ice (to prevent their reformation). A mixture of 3 μl of nuclease-free water, 2 μl of 5× Superscript IV buffer, 0.5 μl of 100 mM DTT, 0.5 μl of 10 mM dNTPs (Invitrogen catalog no. 18427-088), 0.5 μl of RNaseOUT RNase inhibitor (40 U / ml, Invitrogen catalog no. 10777019), and 0.5 μl of Superscript IV reverse transcriptase (200 U / ml, Thermo Fisher Scientific catalog no. 18090200) was added per well. The reverse transcripts were incubated (heated lid set to 60°C) for 2 minutes at 4°C, 2 minutes at 10°C, 2 minutes at 20°C, 2 minutes at 30°C, 2 minutes at 40°C, 2 minutes at 50°C, 2 minutes at 55°C for 15 minutes, and then stored at 4°C.

[0200] Second-Strand Synthesis and Cell / Nuclei Harvesting: For second-strand synthesis, a mixture of 1.33 μl of Second-Strand Synthesis Reaction Buffer and 0.67 μl of Second-Strand Synthesis Enzyme Mix (NEB Catalog No. E6111L) was added per well and then incubated at 16°C for 2 hours. Treated nuclei were harvested from the plates and pooled into one 15 ml tube per plate. The wells were washed with 1x PBS-1% BSA, and the PBS-1% BSA was transferred to the same tube to maximize harvest. The volume was increased to 10 ml with 1x PBS-1% BSA, and nuclei were harvested (500 rcf, 5 minutes, 4°C). Two additional wash steps with 1x PBS-1% BSA were used to remove cellular debris. The resulting pellet was resuspended in 1.5 ml of 1x nuclear buffer (10xGenomics catalog number 2000153), transferred to a 1.5 ml tube, and centrifuged (500 rcf, 5 minutes, 4°C). The supernatant was completely removed, and the tube was briefly centrifuged (500 rcf, 30 seconds, 4°C) to collect the liquid remaining at the bottom of the tube. This typically resulted in less than 10 μl of highly concentrated suspension, which was diluted 1:50 and counted in a Fuchs Rosenthal counting chamber (Incyto catalog number DHC-F01).

[0201] Tagmentation: For tagmentation, processed nuclei were combined with 1x nuclear buffer to a total volume of 5 μl and mixed with 7 μl of ATAC buffer (10x Genomics catalog number 2000122) and 6 μl of custom-made i7-specific transposomes (prepared as described below). Double-stranded cDNA in processed nuclei was tagmented for 1 hour at 37°C and then stored at 4°C.

[0202] Linear barcoding: An unused channel of a Chromium Chip E (10xGenomics catalog number 2000121) was filled with 75 μl (inlet 1), 40 μl (inlet 2), or 240 μl (inlet 3) of 50% glycerol solution (Sigma catalog number G5516-100ML). Immediately before loading onto the chip, a mixture of 61.5 μl of barcoding reagent, 1.5 μl of reducing agent B, and 2.0 μl of barcoding enzyme (all from 10xGenomics catalog number 1000110) was added per tagmentation reaction. A microfluidic chip was loaded with 75 μl of barcoding mix containing tagmented nuclei (inlet 1), 40 μl of single-cell ATAC gel beads (inlet 2, 10xGenomics catalog no. 2000132), and 240 μl of Partitioning Oil (inlet 3, 10xGenomics catalog no. 220088), and run on a 10xGenomics Chromium controller. The linear barcoding reaction was incubated as follows: (heated lid set to 105°C, volume set to 125 μl), 72°C for 5 minutes, 98°C for 30 seconds, 12x (98°C for 10 seconds, 59°C for 30 seconds, 72°C for 1 minute), and stored at 15°C. The emulsion was broken by adding 125 μl of recovery agent (10xGenomics catalog no. 220016), and 125 μl of the pink oil phase was removed by pipetting. The remaining sample was mixed with 200 μl of Dynabead cleanup master mix (per reaction: 182 μl cleanup buffer (10xGenomics catalog no. 2000088), 8 μl Dynabeads MyOne silane (Thermo Fisher Scientific catalog no. 37002D), 5 μl reducing agent B (10xGenomics catalog no. 2000087), 5 μl nuclease-free water).After a 10-minute incubation at room temperature, the sample was washed twice with 200 μl of freshly prepared 80% ethanol (Merck Catalog No. 603-002-00-5) and eluted in 40.5 μl of EB buffer (Qiagen Catalog No. 19086) containing 0.1% Tween (Sigma Catalog No. P7949-500ML) and 1% v / v reducing agent B. Clumps of beads were sheared with a 10 μl pipette or needle. 40 μl of sample was transferred to a new tube strip, processed with SPRIselect beads (Beckman Coulter Catalog No. B23318) at 1.2× cleanup, and eluted in 40.5 μl of EB buffer.

[0203] Enrichment PCR: Each sample was enriched in eight separate PCR reactions containing 50 μl of NEBNext High Fidelity 2x Master Mix (NEB catalog no. M0541S), 5 μl of primer 06-11_Partial-P5 (10 μM, 5'-AATGATACGGCGACCACCGAGA-3'), 1 μl of 100x SYBR Green in DMSO (Life Technologies catalog no. S7563), 34 μl of water, 5 μl of indexed 06-11_P7-Read2N-00X primer (10 μM, 5'-CAAGCAGAAGACGGCATACGAGAT[indexi7] GTCTCGTGGGCTCGG-3'), and 5 μl of sample from the previous step. Reactions were incubated in the qPCR machine: 98°C for 45 seconds, 40x (plate read after 98°C for 20 seconds, 67°C for 30 seconds, and 72°C for 30 seconds). Fluorescence signal was monitored during the run, and samples were removed from the thermocycler when saturation was reached. To complete any unfinished PCR products, samples were incubated in another thermocycler at 72°C for 2 minutes.

[0204] Size selection and quality control: PCR reactions were cleaned with 0.7x standard SPRI cleanup followed by a dual-pronged 0.5x / 0.7x SPRI cleanup. Library size distribution was examined using a Bioanalyzer HS chip (Agilent catalog numbers 5067-4626 and 5067-4627), and dsDNA concentration was measured using the Qubit dsDNA HS assay (Thermo Fisher Scientific catalog number Q32854).

[0205] Example 4 - scifi-RNA-seq (version LIG-TS) based on thermal cycling ligation and template crossing

[0206] Reverse transcription: Sets of 96 and 384 indexed reverse transcription primers were synthesized by Sigma-Aldrich and shipped at 100 μM in EB buffer in a 96-well plate. The primers had the sequence (5'-[phos]ACACTCTTTCCCTACACGACGCTCTTCCGATCTNNNNNNNN N XXXXXXXXXXX VThe oligos had the following structure: TTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTVN-3' (where N represents a random base, the underlined base is known for a given primer, and X represents a primer-specific index sequence 11 bases long, with the 5' phosphate group enabling ligation of this oligo). A 96-well plate with barcoded oligo-dT primers was prepared prior to the experiment and stored at -20°C (1 μl of 25 μM per well). 10,000 permeabilized cells or nuclei (2 μl of a 5,000 / μl suspension) were added to the pre-dispersed primers, and the well assignments were recorded. The plate was incubated at 55°C for 5 minutes (to disrupt RNA secondary structures) and then immediately placed on ice (to prevent their reformation). A mixture of 3 μl of nuclease-free water, 2 μl of 5× reverse transcription buffer, 0.5 μl of 100 mM DTT, 0.5 μl of 10 mM dNTPs (Invitrogen catalog no. 18427-088), 0.5 μl of RNaseOUT RNase inhibitor (40 U / ml, Invitrogen catalog no. 10777019), and 0.5 μl of Maxima H Minus reverse transcriptase (200 U / ml, Thermo Fisher Scientific catalog no. EP0753) was added per well. The reverse transcripts were incubated as follows: (heated lid set to 60°C), 50°C for 10 minutes, 3 cycles of 8°C for 12 seconds, 15°C for 45 seconds, 20°C for 45 seconds, 30°C for 30 seconds, 42°C for 2 minutes, 50°C for 3 minutes, 50°C for 5 minutes, and stored at 4°C.

[0207] Cell / Nuclei Harvesting and Pooling: Treated cells / nuclei were harvested from plates and pooled into one 15 ml tube per plate. Wells were washed with 1x PBS-1% BSA, which was then transferred to the same tube to maximize recovery. The volume was increased to 15 ml with 1x PBS-1% BSA, and nuclei were harvested (500 rcf, 5 min, 4°C). The resulting pellet was resuspended in 1.0 ml of 1x HiFi Taq DNA Ligase Buffer (NEB #M0647S) or 1x Ampligase Reaction Buffer (Lucigen #A0102K) and filtered through a cell strainer (40 μm or 70 μm, depending on cell / nuclei size) into a 1.5 ml tube and centrifuged (500 rcf, 5 min, 4°C). The supernatant was completely removed, and the tube was briefly centrifuged (500 rcf, 30 seconds, 4°C) to collect any liquid remaining at the bottom of the tube. This typically resulted in less than 10 μl of highly concentrated suspension, which was diluted 1:50 and counted in a Fuchs Rosenthal counting chamber (Incyto catalog number DHC-F01). The desired number of cells / nuclei was brought to a volume of 15 μl with 1× HiFi Taq DNA ligase buffer (NEB #M0647S) or 1× Ampligase reaction buffer (Lucigen #A0102K).

[0208] Microfluidic thermal linkage barcoding: Unused channels of a Chromium Chip E (10xGenomics catalog number 2000121) were filled with 75 μl (inlet 1), 40 μl (inlet 2), or 240 μl (inlet 3) of 50% glycerol solution (Sigma catalog number G5516-100ML). Immediately prior to loading onto the chip, a mixture of 47.4 μl of nuclease-free water, 11.5 μl of either HiFi Taq DNA ligase buffer (10×, NEB #M0647S) or Ampligase reaction buffer (10×, Lucigen #A0102K), 2.3 μl of either HiFi Taq DNA ligase (NEB #M0647S) or Ampligase (Lucigen #A0102K), 1.5 μl of reducing agent B (10× Genomics catalog number 2000087), and 2.3 μl of bridge oligo (100 μM, 5′-CGTCGTGTAGGGAAAGAGTGTGACGCTGCCGACGA[ddC]-3′) was added per sample. The microfluidic chip was loaded with 75 μl of cell / nuclei-containing thermal ligation mix (inlet 1), 40 μl of single-cell ATAC gel beads (inlet 2, 10xGenomics catalog no. 2000132), and 240 μl of Partitioning Oil (inlet 3, 10xGenomics catalog no. 220088), and run on a 10xGenomics Chromium controller. The thermal ligation barcoding reaction was incubated as follows: (heated lid set to 105°C, volume set to 100 μl), 12× (98°C for 30 seconds, 59°C for 2 minutes), and stored at 15°C. The emulsion was broken by adding 125 μl of recovery agent (10xGenomics catalog no. 220016), and 125 μl of the pink oil phase was removed by pipetting.The remaining sample was mixed with 200 μl of Dynabead cleanup master mix (per reaction: 182 μl cleanup buffer (10xGenomics catalog no. 2000088), 8 μl of Dynabeads MyOne silane (Thermo Fisher Scientific catalog no. 37002D), 5 μl of reducing agent B (10xGenomics catalog no. 2000087), and 5 μl of nuclease-free water). After a 10-minute incubation at room temperature, the sample was washed twice with 200 μl of freshly prepared 80% ethanol (Merck catalog no. 603-002-00-5) and eluted in 40.5 μl of EB buffer (Qiagen catalog no. 19086) containing 0.1% Tween (Sigma catalog no. P7949-500ML) and 1% v / v reducing agent B. The clumps of beads were sheared with a 10 μl pipette or needle. 40 μl of sample was transferred to a new tube strip and treated with 1.0× cleanup with SPRIselect beads (Beckman Coulter catalog number B23318) and eluted in 22 μl of EB buffer.

[0209] Template transfer: 20 μl of the sample from the previous step was mixed with 10 μl of 5x reverse transcription buffer, 10 μl of Ficoll PM-400 (20%, Sigma #F5415-50ML), 5 μl of 10 mM dNTPs (Invitrogen catalog no. 18427-088), 1.25 μl of recombinant ribonuclease inhibitor (Takara #2313A), 1.25 μl of template transfer oligo (100 μM, 5'-AAGCAGTGGTATCAACGCAGAGTGAATrGrGrG-3', where r indicates the RNA base), and 2.5 μl of Maxima H Minus reverse transcriptase (200 U / ml, Thermo Fisher Scientific catalog no. EP0753). The template transfer reaction was incubated at 25°C for 30 minutes, 42°C for 90 minutes, stored at 4°C, cleaned up with 1.0x SPRI cleanup, and eluted in 17 μl of EB buffer.

[0210] cDNA enrichment: 15 μl of the above sample was mixed with 33 μl of nuclease-free water, 50 μl of NEBNext High Fidelity 2× Master Mix (NEB #M0541S), 0.5 μl of partial P5 primer (100 μM, 5′-AATGATACGGCGACCACCGAGA-3′), 0.5 μl of TSO enrichment primer (100 μM, 5′-AAGCAGTGGTATCAACGCAGAGT-3′), and 1 μl of SYBR Green (100× in DMSO). The cDNA was amplified in a thermocycler at 98°C for 30 seconds, repeated cycles (98°C for 20 seconds, 65°C for 30 seconds, 72°C for 3 minutes) until the fluorescent signal exceeded 2000 RFU, then stored at 4°C for 5 minutes in another thermocycler. The cDNA was cleaned by one 0.8× SPRI cleanup followed by a 0.6× SPRI cleanup, quantified with the Qubit HS assay (ThermoFisher Scientific # Q32854), and 1.5 ng was probed on a Bioanalyzer high sensitivity DNA chip (Agilent #5067-4626 and #5067-4627).

[0211] Library Preparation: cDNA can be converted into NGS-ready libraries by a variety of established methods: (i) tagmentation of double-stranded cDNA using commercially available (e.g., Illumina Nextera) or custom-made Tn5 transposases (instructions for preparing transposomes are included below), followed by PCR enrichment; (ii) fragmentation of double-stranded cDNA by physical (e.g., sonication) or enzymatic (e.g., NEB dsDNA fragmentase) means, followed by end-repair, A-tailing, adapter ligation, and PCR enrichment; (iii) linear extension by random priming using a processivity polymerase (e.g., Klenow fragment), followed by PCR enrichment.

[0212] Example 5 – Linear extension and random priming based scifi-RNA-seq (EXT-RP)

[0213] Random priming (RP) provides an alternative method for introducing defined sequences into the ends of library fragments distal to sequences captured during reverse transcription (e.g., poly-A tails). It is compatible with version TN5 (where it replaces the tagmentation step) and version LIG (where it replaces the template transfer step). Reverse transcription, second-strand synthesis, and cell / nuclei recovery and counting were performed as described above for version EXT-TN5 (Example 3). However, tagmentation was no longer required. Instead, a total volume of 11 μl of 1× nuclear buffer containing treated cells / nuclei was mixed with 7 μl of ATAC buffer (10xGenomics catalog no. 2000122), 61.5 μl of barcoding reagent, 1.5 μl of reducing agent B, and 2.0 μl of barcoding enzyme (all from 10xGenomics catalog no. 1000110), loaded into the microfluidic chip, and run as described above. The sample was cleaned by silane and SPRI bead cleanup as described above for version EXT-TN5 and eluted in a volume of 43 μl of nuclease-free water. 41.75 μl of the cleaned sample was diluted with 5 μl of Blue Buffer (10x, Enzymatics #P7010-HC-L), 1.25 μl of 10 mM dNTPs (Invitrogen catalog number 18427-088), and 1 μl of random primer (100 μM, 5'-[Btn]GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG NNNNThe DNA was mixed with a 100-kDa DNA fragment containing 100 kDa DNA (where the underlined portion corresponds to a random stretch of bases, ideally 4-8 bases in length, optimally biotin-modified). The sample was then denatured at 95°C for 5 minutes and immediately chilled on ice to prevent reformation of secondary structures and allow random primer annealing. One microliter of Klenow exo-polymerase (50 U / μl, Enzymatics #P7010-HC-L) was then added, and the reaction was mixed by pipetting and incubated in a thermocycler: 4°C for 15 minutes, then ramped to 37°C at 1°C / minute, 37°C for 1 hour, then 70°C for 10 minutes (enzyme inactivation), and stored at 4°C. Excess random primers were removed by adding 2.5 μl of Exonuclease I (20 U / μl, NEB #M0293S) and 1.25 μl of rSAP (1 U / μl, NEB #M0371S), followed by incubation at 37°C for 1 hour, heat inactivation at 80°C for 20 minutes, and storage at 4°C. After 0.8x SPRI cleanup or streptavidin-bead cleanup, the library was enriched by PCR as described above for version EXT-TN5.

[0214] Example 6: scifi-RNA-seq based on thermal cycling ligation and random priming (version LIG-RP):

[0215] Reverse transcription, cell / nuclei recovery and counting, microfluidic device thermal ligation barcoding, and silane cleanup were performed as described above for version LIG (Example 4). After SPRI cleanup, the sample was eluted in 43 μl of nuclease-free water. Instead of the template transfer step, random priming was performed as follows: 41.75 μl of the cleaned sample was diluted with 5 μl of Blue Buffer (10x, Enzymatics #P7010-HC-L), 1.25 μl of 10 mM dNTPs (Invitrogen catalog number 18427-088), and 1 μl of random primer (100 μM, 5'-[Btn]GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG NNNN (where the underlined portion corresponds to a random stretch of bases, ideally 4-8 bases in length, optimally biotin-modified). The sample was then denatured at 95°C for 5 minutes, immediately chilled on ice to prevent reformation of secondary structures and allow random primer annealing. Next, 1 μl of Klenow exo-polymerase (50 U / μl, Enzymatics #P7010-HC-L) was added, and the reaction was mixed by pipetting and incubated in a thermocycler: 4°C for 15 minutes, ramped to 37°C at 1°C / min, 37°C for 1 hour, 70°C for 10 minutes (enzyme inactivation), and stored at 4°C. Excess random primers were removed by adding 2.5 μl of Exonuclease I (20 U / μl, NEB #M0293S) and 1.25 μl of rSAP (1 U / μl, NEB #M0371S), followed by incubation at 37°C for 1 hour, heat inactivation at 80°C for 20 minutes, and storage at 4°C. After 0.8x SPRI cleanup or streptavidin-bead cleanup, the library was enriched by PCR as described above for version EXT-TN5.

[0216] Example 7 – Linear extension and template crossover based scifi-RNA-seq (EXT-TS)

[0217] Template transfer (TS) provides an alternative means for introducing defined sequences into the ends of library fragments distal to the captured sequence (e.g., poly-A-tail) during reverse transcription. TS is already used in version LIG-TS, but is also compatible with version EXT-TN5 as described below. Reverse transcription is performed with a reverse transcriptase or an alternative reverse transcriptase that adds a non-templated C base to the cDNA when the end of the transcript is reached. The reverse transcription primer has the sequence (5'-TCGTCGGCAGCGTCGGATGCTGAGTGATTGCTTGTGACGCCTTCNNNNNNNN N XXXXXXXXXXX V The primers have the following structure: TTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTTVN-3' (where N represents a random base, the underlined bases are known for a given primer, and X represents a primer-specific index sequence 11 bases in length). A 96-well plate with barcoded oligo-dT primers was prepared prior to the experiment and stored at -20°C (1 μl of 25 μM per well). 10,000 permeabilized cells or nuclei (2 μl of a 5,000 / μl suspension) were added to the pre-dispersed primers, and well assignments were recorded. The plate was incubated at 55°C for 5 minutes (to disrupt RNA secondary structures) and then immediately placed on ice (to prevent their reformation).

[0218] Per well, add a mixture of 1 μl of 5× reverse transcription buffer, 1 μl of Ficoll PM-400 (20%, Sigma #F5415-50ML), 0.5 μl of 10 mM dNTPs (Invitrogen catalog number 18427-088), 0.125 μl of recombinant ribonuclease inhibitor (Takara #2313A), 0.125 μl of template transfer oligo (100 μM, 5′-AAGCAGTGGTATCAACGCAGAGTGAATrGrGrG-3′, where r represents the RNA base), and 0.25 μl of Maxima H Minus reverse transcriptase (200 U / ml, Thermo Fisher Scientific catalog number EP0753). The combined reverse transcription and template transfer reaction is incubated as follows: (heated lid set to 60°C), 25°C for 30 minutes, 42°C for 90 minutes, and stored at 4°C. Cell / nuclei harvesting and counting are performed as described above for version EXT-TN5, except that tagmentation is no longer necessary. Instead, a total volume of 9.7 μl of 1× nuclear buffer containing the treated cells / nuclei is mixed with 7 μl of ATAC buffer (10xGenomics catalog no. 2000122), 61.5 μl of barcoding reagent, 1.5 μl of reducing agent B, 2.0 μl of barcoding enzyme (all from 10xGenomics catalog no. 1000110), and 1.3 μl of TSO enrichment primer (100 μM, 5'-AAGCAGTGGTATCAACGCAGAGT-3'). The microfluidic chip is loaded, run, and the droplet emulsion is incubated as previously described. The sample is cleaned with silane and SPRI bead cleanup as described above for version EXT-TN5. cDNA is amplified and libraries are prepared as described above for version LIG-TS.

[0219] Example 8 - Assembly of custom i7-specific transposomes

[0220] Oligonucleotides Tn5-top_ME (5'-[Phos]CTGTCTCTTATACACATCT-3') and Tn5-bottom_Read2N (5'-GTTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG-3') were synthesized by Sigma-Aldrich and reconstituted at 100 μM in EB buffer (Qiagen catalog no. 19086). 22.5 μl of each oligonucleotide was mixed with 5 μl of 10× oligonucleotide annealing buffer (10 mM Tris-HCl (Sigma catalog no. T2944-100ML), 50 mM NaCl (Sigma catalog no. S5150-1L), 1 mM EDTA (Invitrogen catalog no. AM9260G)) and annealed in a thermocycler: 95°C for 3 min, 70°C for 3 min, ramped down to 25°C at 2°C / min. The annealing reaction was then diluted by adding 180 μl of water. At this point, the diluted oligonucleotide cassettes could be aliquoted and frozen for future transposome assembly. To load the Tn5 transposase, 20 μl of the diluted oligonucleotide cassette from the previous step was mixed with 20 μl of 100% glycerol (Sigma catalog no. G5516-100ML) and 10 μl of EZ-Tn5 transposase (Lucigen catalog no. TNP92110) and incubated at 25°C for 30 minutes in a thermocycler. The resulting 50 μl of assembled transposomes is sufficient for eight scifi-RNA-seq reactions (6 μl per reaction) using the EXT-TN5 protocol, or for over 200 library preps for implementing scifi-RNA-seq with cDNA enrichment. Transposomes can be stored at -20°C for at least one month.

[0221] Example 9 - Activity check of custom-made i7-specific transposomes by qPCR

[0222] Tagmented DNA flanked by two Illumina i7 adapters is suppressed in PCR reactions due to competition between intramolecular annealing and primer binding. Therefore, custom-made i7-specific transposomes were tested in a previously described negative qPCR assay (Rykalina et al., 2017). Briefly, one tagmentation reaction and one no-enzyme control reaction are performed with a defined PCR product. Both samples are then reamplified using the same primers in a qPCR reaction. Because tagmentation fragments the PCR product, the corresponding reaction should yield a larger Ct value. Tagmentation efficiency can then be calculated from the shift in Ct value: Tagmentation efficiency [%] = 100 / [2^(mean Ct tagmentation - mean Ct no-enzyme control)].

[0223] PCR product generation: Oligonucleotides pUC19-FWD (5'-AAGTGCCACCTGACGTCTAAG-3') and pUC19-REV (5'-CAACAATTAATAGACTGGATGGAGGCGG-3') were synthesized by Sigma-Aldrich and reconstituted at 100 μM in EB buffer (Qiagen catalog no. 19086). A 1,961 bp PCR product was then generated by mixing 1.65 μl each of primers pUC19-FWD and pUC19-REV (100 μM) in combination with 128.7 μl of water, 33 μl of 50 pg / μl pUC19 plasmid (NEB catalog no. N3041S), and 165 μl of 2x Q5 HotStart High-Fidelity Master Mix (NEB catalog no. M0494L). The resulting 6.6x master mix was dispensed into tube strips (six 50 μl reactions) and amplified in a thermocycler: 98°C for 30 seconds; 31x (98°C for 10 seconds, 68°C for 30 seconds, 72°C for 1 minute), 72°C for 2 minutes, and stored at 12°C. To each 50 μl PCR reaction, 6.25 μl of 10x CutSmart buffer and 6.25 μl of DpnI (NEB catalog no. R0176L) were added and incubated at 37°C for 1 hour to digest the PCR template plasmid. The six PCR reactions were pooled and purified using a QiaQuick PCR Purification Kit (Qiagen catalog no. 28106) using two columns and eluted with 30 μl of EB buffer per column. The eluates were pooled, and the purity of the PCR fragments was checked on a 1% agarose gel containing ethidium bromide. The concentration of dsDNA was then measured using a Qubit HS assay (Thermo Fisher Scientific catalog number Q32854), and the PCR product was diluted to 25 ng / μl in EB buffer.

[0224] Tagmentation: The tagmentation reaction was prepared by mixing 2 μl of the 25 ng / μl pUC19 PCR product from the previous step, 7 μl of ATAC buffer (10x Genomics catalog no. 2000122), and either 6 μl of custom-made i7-specific transposomes (tagmentation reaction) or 6 μl of water (no enzyme control reaction). After incubation at 37°C for 60 minutes, the Tn5 enzyme was stripped from the DNA by adding 1.75 μl of 1% SDS solution (Sigma catalog no. 71736-100ML), followed by incubation at 70°C for 10 minutes. Two reactions were diluted 1 / 100 in EB buffer, and qPCR reactions (2 μl of the 1 / 100 diluted reaction, 10 μl of 2× GoTaq qPCR Master Mix (Promega catalog no. A600A), 0.1 μl each of 100 μM pUC19-FWD and pUC19-REV primers, and 7.8 μl of water) were prepared in triplicate. qPCR reactions were incubated as follows: 95°C for 2 minutes, 40× (95°C for 30 seconds, 68°C for 30 seconds, 72°C for 2 minutes, and then the plate was read).

[0225] Example 10 – Next Generation Sequencing

[0226] The resulting scifi-RNA-seq libraries were sequenced on an Illumina NextSeq 500 platform using High Output v2.5 reagents (75 cycles, Illumina catalog number 20024906). Custom sequencing primers were used: 18-12_scifi_SEQ_inDrop_read1 (5'-GGATGCTGAGTGATTGCTTGTGACGCC*T*T*C (where * represents a phosphorothioate linkage)) for read 1 and 18-12_scifi_SEQ_inDrop_index2 (5'-GCATCCGACGCTGCCGA*C*G*A-3') for index 2. The machine was configured to read lengths of 21 bases (read 1), 47 bases (read 2), 8 bases (index 1, i7), and 16 bases (index 12, i5).

[0227] Large single-cell libraries were sequenced on the Illumina NovaSeq 6000 platform using NovaSeq 6000 SP Reagents (100 cycles, Illumina catalog number 20027464) or S2 Reagents (100 cycles, Illumina catalog number 20012862). A custom sequencing primer, 18-12_scifi_SEQ_inDrop_read1 (5'-GGATGCTGAGTGATTGCTTGTGACGCC*T*T*C, where * represents a phosphorothioate linkage)), was applied to read 1. Due to the different sequencing chemistry, index 2 can be read with standard NovaSeq primers. The sequencer was configured to read structures of 21 bases (read 1), 55 bases (read 2), 8 bases (index 1, i7), and 16 bases (index 2, i5).

[0228] In some implementations of scifi-RNA-seq, custom primers were no longer required as primer binding sites compatible with standard Illumina sequencing primers were used.

[0229] Example 11 - scifi-RNA-seq on a 1:1 mixture of human and mouse cells

[0230] Cell Culture: Human Jurkat-Cas9-TCRlib cells were cultured in RPMI medium (Gibco #21875-034) containing 10% FCS (Sigma) and penicillin-streptomycin and sequentially selected with 25 μg / ml blasticidin (Invivogen #ant-bl-5) and 2 μg / ml puromycin (Fisher Scientific #A1113803). Mouse 3T3 cells were cultured in DMEM medium (Gibco #10569010) containing 10% FCS (Sigma) and penicillin-streptomycin.

[0231] Single-cell RNA-seq: Nuclei suspensions from human Jurkat-Cas9-TCRlib cells and mouse 3T3 cells were freshly prepared as described in Example 1.2 (above). To evaluate the performance of scifi-RNA-seq as a function of droplet overloading, 15,300, 383,000, or 765,000 pre-indexed nuclei were loaded into a single channel of the Chromium system. Both the number of single-cell transcriptomes and the average number of nuclei in each droplet scaled linearly with the amount loaded (Figure 19). Additionally, this dataset, based on a 1:1 mixture of human and mouse cell lines, allowed validation of our pre-indexing strategy for correctly assigning transcripts to single cells. To that end, we compared the number of human-mouse cell doublets based solely on microfluidic (Round 2) barcodes with the number of such doublets based on a combination of pre-indexed (Round 1) and microfluidic (Round 2) barcodes (Figure 20). As expected at a loading rate of 765,000 nuclei per channel, almost all droplets contained both human and mouse cells (Figure 20, left), but the majority of these doublets could be distinguished when both round 1 and round 2 barcodes were considered (Figure 20, right). As expected, the surprising effect of pre-indexing was only seen when the droplet generator was overloaded, whereas only the microfluidic round 2 barcodes were sufficient to minimize cell doublets at a standard loading rate of 15,300 nuclei per channel (Figure 33).

[0232] Finally, this dataset allowed us to definitively address a third feasibility concern for scifi-RNA-seq: whether the reagents in each droplet are sufficient for effective barcoding of transcriptomes from a large number of nuclei. When plotting UMI counts and the ratio of unique reads per cell against the number of nuclei per droplet (Figure 21), we observed no trend toward lower transcriptome complexity in droplets containing up to 15 distinct nuclei. This strongly suggests that the reagents for droplet-based indexing are not a limiting factor in scifi-RNA-seq.

[0233] Example 12 – scifi-RNA-seq based on a mixture of four human cell lines

[0234] Cell Culture: Jurkat-Cas9-TCRlib, K562, and NALM-6 cell lines were cultured in RPMI medium (Gibco #21875-034) containing 10% FCS (Sigma) and penicillin-streptomycin. Jurkat-Cas9-TCRlib cells were sequentially selected with 25 μg / ml blasticidin (Invivogen #ant-bl-5) and 2 μg / ml puromycin (Fisher Scientific #A1113803). HEK293T cells were cultured in DMEM medium (Gibco #10569010) containing 10% FCS (Sigma) and penicillin-streptomycin.

[0235] Single-cell RNA-seq: Nuclei suspensions from four uniquely characterized human cell lines (Jurkat, K562, NALM-6, and HEK293T) were freshly prepared as described in Example 1.2 above. These nuclei were then subjected to scifi-RNA-seq, as described in Example 4 above, following a protocol based on temperature cycling ligation and template transfer (LIG-TS). Each cell line was assigned a specific set of pre-indexed (round 1) barcodes during the reverse transcription step in 384-well plates. After pre-indexing, samples were pooled, and 383,000 nuclei were loaded into a single microfluidic channel of the Chromium system. 151,788 single-cell transcriptomes passed quality control (Figures 22-23, 35-36). This represents a 15-fold increase compared to the output of the standard Chromium protocol. This experiment also demonstrated that the method of the present invention inherently supports multiplexing up to 384 different samples in a single experiment.

[0236] Example 13 – scifi-RNA-seq on primary human T cells

[0237] Isolation of primary human T cells: Peripheral blood from healthy donors was obtained from blood packs containing buffered sodium citrate as an anticoagulant. For each donor, T cells were prepared from 3 x 15 ml of peripheral blood according to the following protocol: 15 ml of peripheral blood was mixed with 750 μl of RosetteSep Human T Cell Enrichment Cocktail (Stemcell #15061). After 10 minutes of incubation at room temperature, the sample was diluted by adding 15 ml of 1x PBS (Gibco #14190-094) containing 2% v / v FCS (Sigma). A SepMate tube (Stemcell #86450) was loaded with 15 ml of Lymphoprep density gradient medium (Stemcell #07851), and the blood sample was poured onto the top. After centrifugation (1,200 rcf, 10 min, room temperature, brake set at 9), the supernatant was transferred to a new 50 ml tube, brought up to 50 ml with 1x PBS containing 2% FCS, and centrifuged (1,200 rcf, 10 min, room temperature, brake set at 3). After one additional wash with 50 ml of 1x PBS containing 2% FCS (1,200 rcf, 10 min, room temperature, brake set at 3), T cells were resuspended in 10 ml of 1x PBS containing 2% FCS, filtered through a 40 μM cell strainer, and counted using a CASY instrument (Scharfe Systems). For accurate cell counts, it was important to exclude contaminating red blood cells, as they would be lysed during subsequent nuclei preparation.

[0238] Anti-CD3 / CD28 stimulation of human T cells: Freshly isolated primary human T cells were resuspended at a density of 1 million cells per ml in human T cell medium (OpTmizer medium (Thermo Fisher #A1048501) containing 1 / 38.5 volume of OpTmizer supplement, 1x GlutaMax (Thermo Fisher #35050061), 1x penicillin / streptomycin (Thermo Fisher #15140122), 2% heat-inactivated human AB serum (Fisher Scientific #MT35060CI), and 10 ng / ml recombinant human IL-2 (PeproTech #200-02)) at a density of 1 million cells per ml. The culture was split into two flasks, and one was treated with human T-activator CD3 / CD28 Dynabeads (25 μl beads per million cells, Thermo Fisher #11131D). After 16 hours, formaldehyde-fixed nuclei were prepared as described herein and the nuclei suspension was flash frozen.

[0239] Flow cytometry analysis of T cell populations: A total of 1 million primary human T cells were washed twice with 1× PBS containing 0.1% BSA and 5 mM EDTA (PBS-BSA-EDTA). Single cell suspensions were incubated with anti-CD16 / CD32 (clone 93, 1:200, Biolegend #101301) to block nonspecific binding, and then incubated with CD4 (PE-TxRed, clone OKT4, 1:200, Biolegend #317448), CD8 (APC-Cy7, clone SK1, 1:150, Biolegend #344746), CD25 (PE-Cy7, clone BC96, 1:100, Biolegend #302612), CD45RA (PerCp-Cy5.5, clone HI100, 1:100, Biolegend #304122), CD45RO (AF700, clone UCHL1, 1:100, Biolegend #304218), CD69 (AF488, clone FN50, 1:100, Biolegend #304218), and CD69 (AF488, clone FN50, 1:100, Biolegend #304218). The cells were stained for 30 minutes at 4°C with a combination of antibodies against CD4+ T cells (CD45RA+ CCR7+), CD127 (APC, clone A019D5, 1:100, Biolegend #351342), CD197 (CCR7, PE, clone G043H7, 1:100, Biolegend #353204), and DAPI viability dye (Biolegend #422801). After washing twice with PBS-BSA-EDTA, cells were acquired using an LSRFortessa cell analyzer (BD). CD4+ and CD8+ T cells were divided into naive T cells (CD45RA+ CCR7+), effector memory T cells (CD45RA- CCR7-), central memory T cells (CD45RA- CCR7+), and TEMRA cells (CD45RA+ CCR7-). T cell receptor-mediated activation of CD4+ and CD8+ T cells was assessed based on the expression of CD25 and CD69.

[0240] Single-cell RNA-seq: scifi-RNA-seq was performed as described in Example 4, following a protocol based on temperature cycling ligation and template transfer (LIG-TS). During the reverse transcription step on a 384-well plate, donor identity and TCR stimulation status were barcoded with a set of unique round 1 pre-indexes. After pre-indexing, samples were pooled, and 765,000 nuclei were loaded into a single microfluidic channel on a Chromium system. Results are shown in Figures 24-25 and 37-38.

[0241] Example 14 – Comparison with existing combinatorial indexing protocols

[0242] In this experiment, we compared the performance of our method with existing multi-round combinatorial indexing technologies. Publicly available data were obtained for sci-RNA-seq v1 (Cao, Packer et al., 2017), SPLiT-seq (Rosenberg, Roco et al., 2018), sci-RNA-seq v3 (Cao, Spielmann et al., 2019), and sci-Plex (Srivatsan, McFaline-Figueroa, Ramani et al., 2020). Using mouse 3T3 cells as a common reference point, we demonstrated that library quality for sci-RNA-seq was consistently superior to that of sci-RNA-seq v1, sci-RNA-seq v3, and sci-Plex (Figures 44a-f), with a significantly reduced percentage of doublet cells observed in human / mouse mixed-species experiments (Figure 44g). The data quality of scifi-RNA-seq was more reproducible than SPLIT-seq, which produced highly variable results across two replicate samples (Figure 4a-f).

[0243] Additionally, we compared library design and sequencing read structure between the methods to assess their cost-effectiveness. Because scifi-RNA-seq does not read uninformative ligated overhangs, all sequencing cycles utilized for cell barcodes are informative, in contrast to sci-RNA-seq v1 (58% informative), sci-RNA-seq v3 and sci-Plex (87% informative), and SPLiT-seq (33% informative). As a result, scifi-RNA-seq significantly reduces the sequencing cost barrier for ultra-high-throughput single-cell RNA-seq (Figure 4h-i). In summary, we found that scifi-RNA-seq produces superior data quality and reproducibility, significantly reduces experimental workload, and can be performed faster than existing methods.

[0244] Example 15 – Comparison with 10xGenomics Chromium Platform

[0245] To compare with microfluidic single-cell RNA-seq, we benchmarked scifi-RNA-seq against the widely used 10xGenomics technology, utilizing the latest v3 chemistry. In a series of novel wet-lab experiments, test samples were split and processed in parallel with both assays, loading the same number of nuclei / cells (7,500) per microfluidic channel. Results were compared between permeabilized nuclei, methanol-fixed cells, and intact cells. Equal mixtures of four human cell lines (K562, HEK293T, Jurkat, and NALM-6) differing in transcript content were used, as well as a cross-species mixture of human (Jurkat) and mouse (3T3) cells. This setup allowed us to separate the effects of permeabilization method, technology platform, cell type, species, and transcript content.

[0246] In summary, these experiments demonstrated the following: (i) pre-indexed cells / nuclei were recovered in scifi-RNA-seq at approximately the same rate as the original cells / nuclei in the 10xGenomics system. Because minimal sequencing coverage was utilized relative to background, this could be offset by increasing the loading concentration (Figure 39a). (ii) The washing and filtration steps in scifi-RNA-seq efficiently removed permeabilization artifacts (e.g., free-floating RNA and cell fragments common in 10xGenomics data for methanol-fixed nuclei and cells), further demonstrating the advantages of our protocol (Figure 39a). (iii) False clusters of doublet cells were frequently detected in the 10xGenomics data but completely absent in the scifi-RNA-seq data, demonstrating the enormous barcoding capabilities of our method (Figure 39b and Figures 43a-c). (iv) Four human cell lines were recovered at the same rate. This indicates little cell-type-specific sampling bias or bias due to transcript content (Figure 39c). (v) Gene expression profiles correlate across cell lines regardless of technology (scifi-RNA-seq vs. 10xGenomics) and sample preparation method (nuclei, methanol-fixed cells, whole cells). (v) By design, none of the combined indexing methods are expected to reach the library complexity of direct single-cell RNA-seq using the latest 10xGenomics v3 chemistry, yet they offer greatly increased cell throughput (at least 15x more cells per run) without compromising scifi-RNA-seq data quality.

[0247] Example 16 - Compatibility with Chromium single cell ATAC v.1.1 (NextGEM) design

[0248] Droplet overloading by our method was shown to be compatible with the Chromium Single Cell ATAC v.1.1 (NextGEM) kit (Figure 41a). All loading concentrations tested resulted in stable, monodispersed droplet emulsions (Figure 41b), and the droplet filling rate and number of nuclei per droplet increased with loading concentration in a controlled manner, ranging from 15,000 to 1.53 million nuclei per channel. Design-specific differences of the NextGEM compared to the original chip design were identified, particularly the higher number of nuclei per droplet, better bead loading rates, and significantly reduced number of empty droplets. It was also demonstrated that droplet diameters were similar between platforms and did not change when droplets were overloaded with nuclei (Figures 41c-d). Based on NextGEM-specific data, we computationally modeled the loading of nuclei into droplets (Figure 41e-g), visualized droplet filling rates and nuclei loading distributions (Figure 41h-i), and determined the expected percentage of cell doublets for different numbers of round 1 pre-indexed barcodes (Figure 41j). Finally, we applied our method in parallel using scATAC v1.0 and scATAC v1.1 reagents (NextGEM), demonstrating comparable data quality and single-cell purity (Figure 42a-e). In conclusion, these experiments demonstrated that our method is fully compatible with the NextGEM chip design with respect to both droplet overloading and enzymatic reactions.

[0249] Example 17 - scifi multiplexing enables large-scale perturbation screens at the single-cell level

[0250] The advantages of the whole-transcriptome pre-indexing step in scifi-RNA-seq are twofold. First, barcoded cells / nuclei can be loaded into the second compartment at a large number of cells / nuclei per compartment, enabling ultra-high-throughput processing of samples. Second, round 1 pre-indexing can label hundreds to thousands of experimental conditions, thereby enabling large-scale perturbation studies (e.g., drug screens or genetic perturbation screens) at the single-cell level.

[0251] To demonstrate the multiplexing capabilities of the present invention and the benefits of profiling large numbers of single cells for drug development and target discovery, the following experiment was performed. A human Jurkat cell line was transduced with a lentiviral vector expressing the Cas9 nuclease. These cells were further modified with a second lentiviral vector expressing 48 different CRISPR guide RNAs (gRNAs), each targeting 20 genes with two gRNAs plus eight non-targeting control gRNAs. Efficient genome editing under antibiotic selection was allowed for 10 days. The 48 single knockout cell lines were then split into two portions and stimulated with anti-CD3 / CD28 beads to activate the T cell receptor (TCR) or left untreated. For the resulting 96 samples, methanol-fixed cells were prepared and scifi-RNA-seq was performed according to the method described in the present invention (Figure 40a). A signature of 300 genes differentially expressed under stimulated and unstimulated conditions was used to define a T cell receptor activation score for each gene knockout (Figure 40c).Transcriptome data from this screen were used to identify key regulators of the T cell receptor pathway, including the kinases ZAP70 and LCK, the adaptor protein LAT, and the phosphatase PTPN11, at both the bulk transcriptome (Figure 40b-d) and single-cell levels (Figure 40e-g).

[0252] The above highlights the potential of the methods of the present invention for drug discovery and target validation. Because the methods of the present invention derive relevant screening signatures directly from the transcriptomes of control cells, prior knowledge of the drug's mechanism of action is not required, thereby saving valuable time in prioritizing lead candidates and bringing drug products to market. Furthermore, the single-cell resolution of the methods of the present invention makes it possible to evaluate the effect of drug treatments on different cell types within a complex mixture (e.g., PBMCs) or on mixtures of cells from different donors.

Claims

1. 1. A method for sequencing an oligonucleotide comprising RNA, comprising: (a) providing permeabilized cells and / or nuclei containing a first oligonucleotide comprising RNA; (b) combining the cells and / or nuclei of (a) in a first reaction compartment with a second oligonucleotide comprising DNA, wherein the second oligonucleotide comprises at least a first sequence at least partially complementary to a sequence of the first oligonucleotide, a second sequence comprising an index sequence, and a third sequence comprising a primer binding site, under conditions that allow the first sequence of the second oligonucleotide to anneal to the first oligonucleotide; (c) reverse transcribing the first oligonucleotide in the cells and / or nuclei to obtain an extended second oligonucleotide; (d) reacting the cells and / or nuclei obtained in step (c) with a third oligonucleotide bound to microbeads in a second reaction compartment, wherein the third oligonucleotide is (i) comprises a first sequence that corresponds to a fourth sequence contained in the second oligonucleotide used in step (b); (ii) a first sequence complementary to a first sequence of a fourth oligonucleotide, the fourth oligonucleotide further comprising a second sequence at least partially complementary to the third sequence of the second oligonucleotide; For (i), the method further comprises a step of second strand DNA synthesis after step (c) and before step (d), and for (ii), the method further comprises a step of DNA ligation; the third oligonucleotide further comprises a second sequence comprising an index sequence and a third sequence comprising a primer binding site; (e) amplifying the DNA oligonucleotides obtained in step (d); and (f) sequencing the amplified DNA oligonucleotides.

2. 2. The method of claim 1, wherein in step (c) a non-templated nucleotide is added to the 3' end of the second oligonucleotide.

3. 3. The method of claim 2, wherein second strand DNA synthesis comprises the use of a primer that contains a sequence complementary to the added non-templated nucleotide.

4. 3. The method of claim 2, wherein a primer containing an RNA nucleotide complementary to the added non-templated nucleotide is added for extension.

5. Second strand DNA synthesis occurs (a) introducing a nick into the first oligonucleotide; (b) extending the nicked oligonucleotide; 10. The method of claim 1, further comprising (c) ligating the extended oligonucleotides.

6. The method of claim 1 or 5, further comprising the step of introducing a non-template nucleotide into the 5' end of the synthesized second strand DNA after or simultaneously with second strand DNA synthesis.

7. 7. The method of claim 6, wherein a transposase enzyme, particularly a Tn5 transposase, is used to introduce non-templated nucleotides.

8. 2. The method of claim 1, further comprising a step of linear extension after DNA ligation, wherein the linear extension comprises adding a primer comprising RNA nucleotides and adding a reverse transcriptase.

9. 10. The method of claim 1, further comprising a linear extension step comprising adding a primer containing random nucleotides.

10. 10. The method of any one of claims 1 to 9, wherein the sequence of the first oligonucleotide to which the first sequence of the second oligonucleotide binds is located at the 3' end of the first oligonucleotide.

11. 11. The method of any one of claims 1 to 10, wherein the first sequence of the second oligonucleotide is complementary to the 3' poly-A-tail of the first oligonucleotide.

12. 12. The method of any one of claims 1 to 11, wherein the first reaction compartment comprises permeabilized intact cells and / or nuclei.

13. 13. The method of any one of claims 1 to 12, wherein said first reaction compartment comprises between 5,000 and 10,000 cells.

14. The method of any one of claims 1 to 13, wherein the second reaction compartment contains lysed cells and / or nuclei.

15. 15. The method of any one of claims 1 to 14, wherein the second reaction compartment comprises more than one cell and / or nucleus per microbead, preferably 10 cells and / or nuclei per microbead.

16. The method of any one of claims 1 to 15, wherein said second reaction compartment is a microfluidic droplet or well on a microtiter plate, in particular a sub-nanoliter well plate.

17. 17. The method of claim 16, wherein the second reaction compartment is a microfluidic droplet and the third oligonucleotide is released from the microbead upon formation of the droplet.

18. 18. The method of any one of claims 1 to 17, wherein the second oligonucleotide further comprises a unique molecular identifier (UMI).

19. 19. The method of any one of claims 1 to 18, wherein said cells / nuclei are obtained from an in vitro culture or from a fresh or frozen sample.

20. The cell / nucleus is (a) derived from organoids or xenografts derived from existing cell lines, primary cells, blood cells, or somatic cells; (b) is a CAR-T cell, a CAR-NK cell, an engineered T cell, a B cell, an NK cell, or an immune cell, or is isolated from a patient treated with such a product; or (c) The method of any one of claims 1 to 19, wherein the stem cells are pluripotent stem cells (iPS) or embryonic stem cells that have undergone natural differentiation or artificially induced reprogramming or transdifferentiation.

21. The method of any one of claims 1 to 20, wherein the DNA ligation utilizes a thermostable DNA ligase.

22. The method of any one of claims 1 to 21, using a microfluidic system, in particular for generating microfluidic droplets or for delivering materials to a microfluidic well-based device.

23. 23. The use of claim 22, wherein the microfluidic system is a droplet generator.

24. 23. The use of claim 22, wherein the microfluidic system comprises a sub-nanoliter well plate.

25. A kit comprising a second oligonucleotide as defined in claim 1, preferably together with instructions for use in the method of any one of claims 1 to 21.

26. 26. The kit of claim 25, further comprising a transposase enzyme.

27. 26. The kit of claim 25, further comprising second strand synthesis reagents and / or a thermostable ligase.

28. 28. The kit of any one of claims 25 to 27, further comprising a fourth oligonucleotide.