High-throughput multiomic readout of RNA and genomic DNA within single cells
Patent Information
- Application Number
- PCT/US2024/029950
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-05-19
- Filing Date
- 2024-05-17
- Publication Date
- 2025-09-04
AI Technical Summary
Current methods for simultaneously reading out RNA and genomic DNA (gDNA) within single cells are laborious, low-throughput, and limited in scalability, particularly when trying to link genetic information to transcriptomic signatures, especially in non-coding regions, and existing high-throughput methods suffer from low sensitivity and high sequencing costs.
A method involving a single cell suspension of fixed and permeabilized cells, followed by in-situ reverse transcription to generate cDNA, and subsequent multiplexed PCR in droplets using primers with specific overhang sequences to amplify both cDNA and gDNA simultaneously, allowing for targeted and high-sensitivity detection of multiple RNA transcripts and gDNA loci.
Enables high-throughput, targeted, and sensitive detection of RNA and gDNA within single cells, improving the linkage of genomic information to transcriptomic signatures with reduced sequencing costs and increased coverage, suitable for various applications including genome-wide association studies and precision editing.
Smart Images

Figure US2024029950_04092025_PF_FP_ABST
Abstract
Description
PATENT Attorney Docket No.079445-012310PC-1444403 Client Ref. No. S23-117 HIGH-THROUGHPUT MULTIOMIC READOUT OF RNA AND GENOMIC DNA WITHIN SINGLE CELLS CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Application No.63 / 503,366, filed May 19, 2023, the disclosure of which is herein incorporated by reference in its entirety for all purposes. BACKGROUND
[0002] Recent advances in droplet-based or split-pooling single cell technologies have led to a revolution in the ease of assessing cellular heterogeneity in an unprecedented way (1, 4, 14). DNA, RNA and protein can be currently read-out individually or in a combinatorial fashion. Many of the multiomic combinatorial readouts exist already, i.e., RNA and chromatin accessibility, RNA and DNA methylation status, DNA and DNA methylation status, RNA and protein and DNA and protein. Combined protein readout is mostly achieved by using oligo- linked antibodies that recognize cell membrane proteins, which in turn limits the type and number of proteins that can be assessed in addition to either RNA or DNA (15, 16). In addition, methods exist to simultaneously assess mutational status within transcribed RNA together with transcriptomic signatures (17). However, a high-throughput method that enables scientists to assess genetic information also in non-coding regions and link this information to a transcriptomic signature with high sensitivity is currently lacking.
[0003] Current methods to achieve simultaneous readout of both RNA and genomic DNA (gDNA) are laborious and low-throughput (1, 2, 18). Current approaches are microtiter plate well-based and RNA and DNA is either split before downstream analysis or processed and tagged without separation. In all cases the number of cells retrieved after these methods is very low, which confounds interpretation of the data significantly, makes is unusable for highthroughput screening purposes or other large-scale experimental approaches. More scalable methods like the recently published method of Olsen et al., 2023, uses nucleosome depletion and Tn5-based tagmentation to enable simultaneous gDNA and RNA read-out in single cells (3). These Tn5-based approaches have the limitation that not all the gDNA withina single cell gets tagmented at each locus, limiting the coverage per given target site. In addition, these methods need to read-out the entire tagmented gDNA, which means very high sequencing costs.
[0004] This present disclosure describes a method that enables read out of a multitude of RNA transcripts and gDNA loci simultaneously within single cells in a targeted fashion with high coverage in all cells. It has numerous potential applications that involve linking genomic information to transcriptomic signatures. BRIEF SUMMARY
[0005] The present disclosure provides methods and compositions that are useful for simultaneously detecting RNA and genomic DNA (gDNA) within the same cell. In one aspect, the disclosure provides a method for simultaneously detecting RNA and genomic DNA (gDNA) within the same cell, comprising: (i) providing a single cell suspension comprising fixed and permeabilized cells; (ii) performing an in-situ reverse transcription (RT) step to generate cDNA molecules by contacting the cells with reverse transcriptase and a RT primer; (iii) lysing the cells in a first droplet comprising a first reverse primer that hybridizes to the cDNA molecules and a second reverse primer that hybridizes to the gDNA molecules, wherein the first or second reverse primer comprises a R2N or R2 overhang sequence; (iv) fusing the first droplet to a second droplet comprising PCR reagents and a forward primer that hybridizes to both the cDNA and gDNA molecules and comprises a capture sequence (CS) overhang sequence; and (v) simultaneously amplifying the cDNA and gDNA molecules, thereby detecting both RNA and gDNA within the same cell.
[0006] In some embodiments, the single cell suspension is fixed with paraformaldehyde (PFA) or glyoxal.
[0007] In some embodiments, the R2N overhang sequence comprises the nucleic acid sequence GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG (SEQ ID NO:2), and the R2 overhang sequence comprises the nucleic acid sequence GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT (SEQ ID NO:3).
[0008] In some embodiments, the RT primer comprises a capture sequence (CS), a sample barcode (SBC) sequence, a unique molecular identifier (UMI) sequence, and a sequence that binds to mRNA.
[0009] In some embodiments, the CS sequence hybridizes to a complementary sequence attached to a solid support, the SBC sequence comprises a known sequence of variable length, and the UMI sequence comprises a random sequence of variable length.
[0010] In some embodiments, the CS comprises the nucleic acid sequence GTACTCGCAGTAGTC (SEQ ID NO:1).
[0011] In some embodiments, the SBC sequence is from 4 to 50 nucleotides in length. In some embodiments, the SBC sequence comprises a nucleic acid sequence selected from the group consisting of TCGCCTTA (SEQ ID NO:4), CTAGTACG (SEQ ID NO:5), TTCTGCCT (SEQ ID NO:6), GCTCAGGA (SEQ ID NO:7), AGGAGTCC (SEQ ID NO:8), CATGCCTA (SEQ ID NO:9), GTAGAGAG (SEQ ID NO:10), and CCTCTCTG (SEQ ID NO:11).
[0012] In some embodiments, the UMI sequence is from 4 to 50 nucleotides in length. In some embodiments, the UMI sequence comprises the nucleic acid sequence NNNNNNNN (SEQ ID NO:12).
[0013] In some embodiments, the sequence that binds to mRNA comprises oligo(dT) or a sequence that hybridizes to any region of the mRNA.
[0014] In some embodiments, step (v) of the method comprises PCR conditions that favor binding of forward and reverse primers to the cDNA and gDNA molecules, thereby producing a first set of PCR products each comprising the CS.
[0015] In some embodiments, the second droplet further comprises cell barcode oligonucleotides each comprising a cell barcode (CBC) and a sequence complementary to the CS, thereby the first set of PCR products hybridize to the cell barcode oligonucleotides, producing a second set of PCR products each comprising the CS and the CBC.
[0016] In some embodiments, the cell barcode oligonucleotides are attached to a solid support in the second droplet prior step (iv). In some embodiments, the cell barcode oligonucleotides are released from the solid support after the second droplet is fused with the first droplet.
[0017] In some embodiments, simultaneous amplification in step (v) produces a cDNA library and a gDNA library.
[0018] In some embodiments, the method further comprises sequencing the cDNA and gDNA libraries.
[0019] In some embodiments, the cDNA library is sequenced by a first sequencing modality, and the gDNA library is sequenced by a second sequencing modality. In some embodiments, the first and second sequencing modalities are the same or different. In some embodiments, the first sequencing modality comprises scRNA-seq and the second sequencing modality comprises scDNA-seq. In some embodiments, the first sequencing modality comprises Illumina® next generation sequencing (NGS), and the second sequencing modality comprises Nextera® NGS, or the first sequencing modality comprises Nextera® NGS, and the second sequencing modality comprises Illumina® NGS.
[0020] In some embodiments, step (iii) of the method comprises contacting the cells with proteinase K to lyse the cells.
[0021] In some embodiments, the cell is a prokaryotic or eukaryotic cell. In some embodiments, the cell is an induced pluripotent stem cell (iPSC). In some embodiments, wherein the cell is a genetically modified cell. In some embodiments, the cell is modified using a CRISPR / Cas gene editing system. In some embodiments, the CRISPR / Cas gene editing system is CRISPR interference (CRISPRi).
[0022] In some embodiments, RNA expressed by the genetically modified cell is sequenced by scRNA-seq and genomic DNA is sequenced by scDNA-seq.
[0023] In another aspect, the disclosure provides a reaction mixture comprising RNA and gDNA from a single cell, reverse transcriptase, and a RT primer comprising a capture sequence (CS), a sample barcode (SBC) sequence, a unique molecular identifier (UMI) sequence, and a sequence that binds to mRNA.
[0024] In some embodiments, the CS hybridizes to a complementary sequence attached to a solid support, the SBC sequence comprises a known sequence of variable length, and the UMI sequence comprises a random sequence of variable length, and the sequence that binds to mRNA comprises oligo(dT) or a sequence that hybridizes to any region of the mRNA.
[0025] In some embodiments, the CS comprises the nucleic acid sequence GTACTCGCAGTAGTC (SEQ ID NO:1). In some embodiments, the SBC sequence is from 4 to 50 nucleotides in length. In some embodiments, the SBC sequence comprises a nucleic acid sequence selected from the group consisting of TCGCCTTA (SEQ ID NO:4), CTAGTACG (SEQ ID NO:5), TTCTGCCT (SEQ ID NO:6), GCTCAGGA (SEQ ID NO:7), AGGAGTCC (SEQ ID NO:8), CATGCCTA (SEQ ID NO:9), GTAGAGAG (SEQ ID NO:10), and CCTCTCTG (SEQ ID NO:11). In some embodiments, the UMI sequence is from 4 to 50 nucleotides in length. In some embodiments, the UMI sequence comprises the nucleic acid sequence NNNNNNNN (SEQ ID NO:12).
[0026] In some embodiments, the reaction mixture comprises cDNA molecules produced by reverse transcription of the RNA and a reverse primer that hybridizes to both the cDNA and the gDNA molecules and comprises a R2N or R2 overhang sequence.
[0027] In some embodiments, the reverse primer comprising the R2N overhang sequence comprises the nucleic acid sequence GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG (SEQ ID NO:2) and the reverse primer comprising the R2 overhang sequence comprises the nucleic acid sequence GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT (SEQ ID NO:3).
[0028] In some embodiments, the reaction mixture comprises a, one or more, or a plurality of polynucleotides selected from a sequence in any one of SEQ ID Nos: 1-83.
[0029] In another aspect, the disclosure provides a polynucleotide comprising a capture sequence (CS).
[0030] In some embodiments, the polynucleotide further comprises a sample barcode (SBC) sequence, a unique molecular identifier (UMI) sequence, and / or a sequence that binds to mRNA.
[0031] In some embodiments, the SBC sequence of the polynucleotide comprises a known sequence of variable length, and the UMI sequence of the polynucleotide comprises a random sequence of variable length.
[0032] In some embodiments, the polynucleotide comprises a sequence in any one of SEQ ID Nos: 1-83.
[0033] In any of the embodiments of the disclosure, the CS can comprise the nucleic acid sequence GTACTCGCAGTAGTC (SEQ ID NO:1).
[0034] In some embodiments, the SBC sequence is from 4 to 50 nucleotides in length. In any of the embodiments of the disclosure, the SBC sequence can comprise a nucleic acid sequence selected from the group consisting of TCGCCTTA (SEQ ID NO:4), CTAGTACG (SEQ ID NO:5), TTCTGCCT (SEQ ID NO:6), GCTCAGGA (SEQ ID NO:7), AGGAGTCC (SEQ ID NO:8), CATGCCTA (SEQ ID NO:9), GTAGAGAG (SEQ ID NO:10), and CCTCTCTG (SEQ ID NO:11).
[0035] In some embodiments, the UMI sequence is from 4 to 50 nucleotides in length. In any of the embodiments of the disclosure, the UMI comprises the nucleic acid sequence NNNNNNNN (SEQ ID NO:12).
[0036] In any of the embodiments of the disclosure, the sequence that binds to mRNA can comprise oligo(dT) or a sequence that hybridizes to any region of the mRNA. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Fig. 1. Overview of targeted scDNA-scRNA-seq method with high sensitivity read- out. Main steps involve fixation of a single-cell suspension and in-situ reverse transcription (RT) followed by a multiplexed PCR within individual droplets. Both cDNA and gDNA targets are amplified at the same time. UMI, unique molecular identifier; SBC, sample barcode; CS, capture sequence; CBC, cell barcode.
[0038] Figs. 2A-2I. scDNA-scRNA-seq proof-of principle experiment. (A) Outline of experiment. Fixation conditions tested are PFA and Glyoxal. 28 gDNA and 30 cDNA targets were amplified. (B) Ranked cell barcodes. Automatic threshold is indicated. (C) Number of cells found per fixation condition. Number of gDNA (D) and cDNA (G) targets found per cell. Number of gDNA (E) and cDNA (H) target reads or UMIs / cell. gDNA (F) or cDNA (I) target coverage per cell.
[0039] Figs. 3A-3H. Expression / Coverage of cDNA and gDNA targets and comparison to bulk RNA-seq data. Average expression of gDNA (A) and cDNA (B) targets. Size indicates fraction of cells target is expressed / found in; Color indicates expression (log10). (C) Comparison of expressed cDNA targets to bulk RNA-seq data. Z-score, data is scaled by column. (D) Clustered based on expressed cDNA targets. Example UMAPs of gDNA (E, F) and cDNA (G, H) targets.
[0040] Figs. 4A-4E. CRSIPRi perturbation screen using scDNA-scRNA-seq. (A) Modification of the technology to enable separate sequencing of both cDNA and gDNA libraries. R2 overhangs for cDNA and gDNA are distinct (Illumina and Nextera, respectively) enabling separate NGS library generation and sequencing. (B) Screen target selection. Number of non-targeting control (NTC), CRISPRi control and eQTL gRNAs is shown. cDNA and gDNA targets amplified are indicated. (C) Experimental outline. Lentiviral CROP-seq gRNA library is used to infect transgenic iPSCs that express CRISPRi from the AAVS1 locus. (D) gRNA coverage. Color indicated type of gRNA. (E) Screening results. NTC, CRISPRi control and eQTL gRNAs are shown. Color indicates that this gRNA showed a significant downregulation of its target gene.
[0041] Figs. 5A-5G. Editing screen using scDNA-scRNA-seq. (A) Screen target selection. Number of coding controls (introducing STOP codons) and eQTL pegRNAs is shown. cDNA and gDNA targets amplified are indicated. (B) Experimental outline. pegRNA library is used to transfect two transgenic iPSCs that express PEmax or PEmax and MLH1dn from the AAVS1 locus. (C) Editing efficiencies for both PEmax cell lines shown. (D) Dimplot indicating clusters found. UMAPs showing SOX11 (E) and ATF4 (F) edited cells highlighting homozygous (E+ / E+), heterozygous (E+ / E-) and non-edited (E- / E-) cells. Color scale indicates fraction of gDNA reads with edit per cell. (G) ATF4 gene expression correlated with fraction of gDNA reads with edit per cell. P-value < 0.05; Linear regression model.
[0042] Fig. 6 shows an exemplary barcode bead of the disclosure. The oligonucleotide is attached to the barcoding bead by a photocleavable linker (PhotoC-E), and comprises an R1N sequence, followed by a first cell barcode (CBC1, – 9 bp), an off-staggered constant space linker (CSL,– 14 to 17 bp), a second cell barcode (CBC2, – 9 bp) and a capture sequence (CS).
[0043] Figs.7A-7N. SDR-seq to simultaneously gDNA targets and gene expression in single cells. (A), Overview of targeted scDNA-scRNA-seq (SDR-seq) method. Main steps involve fixation of a single-cell suspension and in-situ reverse transcription (RT) followed by a multiplexed PCR within individual droplets. Both RNA and gDNA targets are amplified at the same time. UMI, unique molecular identifier; BC, barcode. (B), Outline of proof-of-concept (POP) experiment. Fixation conditions and number or gDNA / RNA targets are indicated. (C), Knee plot of ranked lineages by sequencing depth. (D), Number of cells found per fixation condition. (E), Correct sample BC detection per cell. Reads for maximum sample BC found was divided by the total amount of reads for all RNA targets found per cell. (F), Reads / cell forgDNA targets vs UMIs / cell for genes for per fixation condition. Color indicates % of maximum sample BC of total reads. (G and H), Number of gDNA targets / cell (G) and reads / cell for gDNA targets. i, Individual gDNA targets are shown per fixation condition. Size indicates fraction of cells detected in. Color indicates read coverage. (J and K), Number of genes per cell (J) and UMIs / cell for genes measured. (L), Individual gDNA targets are shown per fixation condition. Size indicates fraction of cells detected in. Color indicates UMI coverage. (M), Comparison of expressed genes to bulk RNA-seq data. Z-score, data is scaled by row. (N), Pearson correlation of expressed genes to bulk RNA-seq data.
[0044] Figs. 8A-8H. Quality controls and comparison of PFA and glyoxal fixation conditions. a-d, Quality control plots before (A, B) and after (C, D) filtering for low quality cells. Color indicates either fraction of detected gDNA or RNA targets. e, f, gDNA (E) or RNA (F) coverage and detection. g, h, Comparison of coverage and detection between PFA and Glyoxal for each gDNA (G) and RNA (H) target.
[0045] Figs. 9A-9D. UMAPs and clustering of POP data. (A), Clustering of POP SDR-seq data. Clusters are indicated. (B), Color coding of clusters in UMAP by fixation condition. (C), UMAP plots of ubiquitously expressed genes GAPDH, POU5F1 and SOX2. (D), UMAP plots of cluster-specifically expressed genes KLF5, SALL4 and ESRRB.
[0046] Figs. 10A-10J. SDR-seq is amendable for hundreds of targets simultaneously. (A), Outline of panel size testing experiments. gDNA and RNA targets are equal within panels; Shared targets are indicated. (B, C), Pearson correlation of detection (B) and coverage (C) of shared gDNA targets between panels. (D, E), Pearson correlation of detection (B) and coverage (C) of shared genes between panels. (F), Outline of chromatin sites tested. OEG, overlapping expressed gene. NOEG, non-overlapping expressed gene. PLS, promotor-like sequence. pELS, proximal enhancer-like sequence. dELS, distal enhancer-like sequence. (G), Detection of different chromatin sites between panels. Size indicates fraction of cells detected in. Color indicates read coverage. (H), Outline of expression levels tested. (I), Detection of genes with different expression levels between panels. Size indicates fraction of cells detected in. Color indicates read coverage. (J), Heatmap of expression of all shared genes between panels. Z- score, data is scaled by row.
[0047] Figs. 11A-11F. SDR-seq with separate library generation for RNA and gDNA targets. (A), Overview of SDR-seq with separate library generation for RNA and gDNA. Modification includes a distinct R2 or R2N overhang for each library. Sequencing readylibraries can be generated using specific library primers binding to R2 or R2N, respectively. (B), Specificity of gDNA or RNA NGS libraries. Data from gDNA or RNA libraries was mapped to either gDNA or RNA references. (C, D), Subsampled reads / cell for all gDNA targets (C) or shared gDNA targets (D). (E, F), Subsampled reads / cell for all RNA targets (E) or shared RNA targets (F).
[0048] Figs. 12A-12J. Metrics for target detection and coverage across differently sized target panels. (A), Quality metrics of panel size experiment. Color indicates fraction of gDNA targets / cell recovered. (B), gDNA coverage and detection for all targets across panels tested. (C), gDNA target detection per cell for all targets across panels tested. (D), gDNA coverage and detection for shared targets across panels tested. (E), gDNA target detection per cell for shared targets across panels tested. (F), Quality metrics of panel size experiment. Color indicates fraction of RNA targets / cell recovered. (G), RNA coverage and detection for all targets across panels tested. (H), RNA target detection per cell for all targets across panels tested. (I), RNA coverage and detection for shared targets across panels tested. (J), RNA target detection per cell for shared targets across panels tested.
[0049] Figs.13A-13B. Comparison of target detection and coverage across differently sized target panels. (A, B), Comparison of coverage and detection for shared targets across panels tested for each gDNA (A) and RNA (B) target.
[0050] Figs.14A-14K. SDR-seq is sensitive to detect gene expression changes and link them to variants. (A), Outline of the CRISPRi screen. NTC, non-targeting control. (B), Experimental outline of the CRISPRi screen. (C), Volcano plot for CRISPRi screen with different gRNA classes indicating foldchange and P-value. Significant hits (P-value < 0.05) are colored. For NTCs all genes measured are shown. For other gRNA classes only the intended target for each gRNA is shown. (D), Outline of prime editing (PE) screen. (E), Experimental outline of the PE screen. (F), Volcano plot for PE screen with different gRNA classes indicating foldchange and P-value. Significant hits (P-value < 0.05) are colored. Comparison between the different alleles is shown as shapes. (G), Alleles and gene expression for SOX11, ATF4 and MYH10 STOP controls are shown. ***P < 10í4by MAST. (H), Outline of base editing (BE) screen. (I), Experimental outline of the BE screen. (J), Volcano plot for different gRNA classes indicating foldchange and P-value. Significant hits (P-value < 0.05) are colored. Comparison between the different alleles is shown as shapes. (K), Variants in POU5F1 locus shown and their impact on gene expression shown. Cell numbers are indicated on the left. Impact of variant is color coded.REF, HET and ALT alleles are shown for each genotype. Foldchange is indicated in color (green), P-value as size.
[0051] Figs. 15A-15E. Quality metrics for CRISPRi screen and functional testing of PE iPSCs. (A), Overview of gRNA assignment for CRISPRi screen. (B), Coverage for each gRNA in CRISPRi screen. (C), Relative distance of gRNA binding site to transcription start site (TSS). Positive values are after TSS (within transcript), negative values before transcript. Size indicates P-value. Significant hits (P-value < 0.05) are colored. (D), Outline of testing for PE iPSCs. Editing can be measured by repairing a non-functional EGFP that was integrated via a lentivirus. (E), Flow cytometry indicating editing in PE iPSCs.
[0052] Figs. 16A-16E. Editing efficiency and genotyping in PE screen. (A), Editing efficiency in PE screen. CRISPRi is indicated as a control. (B), Editing efficiency for each locus that was assessed. Loci that were either HET or ALT in the PE iPCSs are not shown. Color indicates either PE cell line or CRISPRi. (C-E), Variant allele frequency (VAF) and genotype quality shown for SOX11 (C), ATF4 (D) and MYH10 (E). Called genotypes shown in Figure 3 are indicated as color.
[0053] Figs.17A-17B. POU5F1 locus in BE screen. (A), Intended edit is shown at POU5F1 locus and its impact on gene expression. (B), All measured variants shown for POU5F1 with their impact on gene expression. DETAILED DESCRIPTION Definitions
[0054] It is to be understood that this disclosure is not strictly limited to particular embodiments described, as such may of course vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present invention will be limited only by the claims.
[0055] It must be noted that as used herein and in the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise. It should further be understood that as used herein, the term “a” entity or “an” entity refers to one or more of that entity. For example, a nucleic acid molecule refers to one or more nucleic acid molecules. As such, the terms “a”, “an”, “one or more” and “at least one” can be used interchangeably. Similarly, the terms “comprising”, “including” and “having” can be used interchangeably.
[0056] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present invention, the preferred methods and materials are now described. All publications mentioned herein are incorporated herein by reference to disclose and describe the methods and / or materials in connection with which the publications are cited. The publications discussed herein are provided solely for their disclosure prior to the filing date of the present application. Nothing herein is to be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention. Further, the dates of publication provided may be different from the actual publication dates, which may need to be independently confirmed.
[0057] It is appreciated that certain features of the disclosure, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the invention, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable sub- combination. All combinations of the embodiments are specifically embraced by the present invention and are disclosed herein just as if each and every combination was individually and explicitly disclosed. In addition, all sub-combinations are also specifically embraced by the present invention and are disclosed herein just as if each and every such sub-combination was individually and explicitly disclosed herein.
[0058] It is further noted that the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as “solely,” “only” and the like in connection with the recitation of claim elements or use of a “negative” limitation.
[0059] As used herein, the term “about” means a range of values including the specified value, which a person of ordinary skill in the art would consider reasonably similar to the specified value. In embodiments, about means within a standard deviation using measurements generally acceptable in the art. In embodiments, about means a range extending to + / - 10% of the specified value (e.g., + / - 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10% of the specified value). In embodiments, about means the specified value.
[0060] As used herein, a “barcode” refers to a nucleotide sequence that is used to identify an individual group of polynucleotides and distinguish it from other groups of polynucleotidesamong a mixture of groups. A “sample barcode” or “SBC” refers to a unique barcode for a sample (e.g., a biological sample, a tissue, a biopsy, a cell) that is different from the barcode sequence associated with other samples (e.g., a biological sample, a tissue, a biopsy, a cell). A “cell barcode” or “CBC” refers to a unique barcode for a cell that is different from the barcode sequence associated with other cells.
[0061] As used herein, “unique molecular identifier” (UMI) refers to sequences of nucleotides present in DNA molecules that may be used to distinguish individual DNA molecules from one another. See, e.g., Kivioja, Nature Methods 9, 72-74 (2012). UMIs may be sequenced along with the DNA sequences with which they are associated to identify sequencing reads that are from the same source nucleic acid. The term “UMI” is used herein to refer to both the nucleotide sequence of the UMI and the physical nucleotides, as will be apparent from context. UMIs may be random, pseudo-random, or partially random, or nonrandom nucleotide sequences that are inserted into adapters or otherwise incorporated in source nucleic acid (e.g., DNA or RNA) molecules to be sequenced. In some embodiments, each UMI is expected to uniquely identify any given source molecule present in a sample. In some embodiments, a probe can further include a cleavage domain and / or a functional domain (e.g., a primer-binding site, such as for next generation sequencing (NGS).
[0062] The term “capture sequence” or “CS” refers to a nucleic acid sequence that capable of capturing, directly or indirectly, and / or labelling a target RNA or DNA in a sample. A capture sequence consists of a pre-defined and known nucleic acid sequence but may vary in length. In some embodiments, the capture sequence has a length of 10 to 100 nucleotides, 10 to 80 nucleotides, 10 to 60 nucleotides, 10 to 40 nucleotides, 20 to 40 nucleotides, or 15 to 35 nucleotides. In some cases, the capture sequence comprises random nucleic acid sequences for capturing total RNA or genomic DNA sequence. In some cases, the capture sequence comprises specific target nucleic acid sequences for capturing target RNA or DNA sequences.
[0063] The terms “overhang sequence” refers to a nucleic acid sequence that is located at the 5’ end of an oligonucleotide or primer that, at least initially, does not hybridize to the target sequence that is amplified during the first round of primer extension.
[0064] The term “genome editing” refers to a type of genetic engineering in which DNA is inserted, replaced, or removed from a target DNA (e.g., the genome of a cell) using one or more nucleases and / or nickases. The nucleases create specific double-strand breaks (DSBs) at desired locations in the genome and harness the cell's endogenous mechanisms to repair theinduced break by homology-directed repair (HDR) (e.g., homologous recombination) or by nonhomologous end joining (NHEJ). The nickases create specific single-strand breaks at desired locations in the genome. In one non-limiting example, two nickases can be used to create two single-strand breaks on opposite strands of a target DNA, thereby generating a blunt or a sticky end. Any suitable DNA nuclease can be introduced into a cell to induce genome editing of a target DNA sequence.
[0065] As used herein, the term “reverse transcriptase” refers to its plain and ordinary meaning as an enzyme used to generate complementary DNA (cDNA) from an RNA template, a process termed reverse transcription.
[0066] The terms "polypeptide" and "protein" refer to a polymer of amino acid residues and are not limited to a minimum length. Thus, peptides, oligopeptides, dimers, multimers, and the like, are included within the definition. Both full length proteins and fragments thereof are encompassed by the definition. The terms also include post expression modifications of the polypeptide, for example, glycosylation, acetylation, phosphorylation, hydroxylation, and the like. Furthermore, for purposes of the present disclosure, a "polypeptide" refers to a protein which includes modifications, such as deletions, additions and substitutions to the native sequence, so long as the protein maintains the desired activity. These modifications may be deliberate, as through site directed mutagenesis, or may be accidental, such as through mutations of hosts which produce the proteins or errors due to PCR amplification.
[0067] As used herein, “sequence specific endonuclease” refers to an enzyme that cleaves at a specific sequence within a polynucleotide sequence. In some aspects, the nuclease activity can be partially or completed inhibited, so that only one of the two strands or neither strand is cleaved,. Non-limiting examples of sequence specific endonucleases include CRISPR associated (Cas) nuclease, a Zinc-finger nuclease, a Transcription activator-like effector nuclease (TALEN), or a meganuclease.
[0068] The term "Cas9" as used herein encompasses type II clustered regularly interspaced short palindromic repeats (CRISPR) system of Cas9 endonucleases from any species, and also includes biologically active fragments, variants, analogs, and derivatives thereof that retain Cas9 endonuclease activity (i.e., catalyze site-directed cleavage of DNA to generate double- strand breaks). A Cas9 endonuclease binds to and cleaves DNA at a site comprising a sequence complementary to its bound guide RNA (gRNA).
[0069] A Cas9 polynucleotide, nucleic acid, oligonucleotide, protein, polypeptide, or peptide refers to a molecule derived from any source. The molecule need not be physically derived from an organism but may be synthetically or recombinantly produced. Cas9 sequences from a number of bacterial species are well known in the art and listed in the National Center for Biotechnology Information (NCBI) database. See, for example, NCBI entries for Cas9 from: Streptococcus pyogenes (WP_002989955, WP_038434062, WP_011528583); Campylobacter jejuni (WP_022552435, YP_002344900), Campylobacter coli (WP_060786116); Campylobacter fetus (WP_059434633); Corynebacterium ulcerans (NC_015683, NC_017317); Corynebacterium diphtheria (NC_016782, NC_016786); Enterococcus faecalis (WP_033919308); Spiroplasma syrphidicola (NC_021284); Prevotella intermedia (NC_017861); Spiroplasma taiwanense (NC_021846); Streptococcus iniae (NC_021314); Belliella baltica (NC_018010); Psychroflexus torquisI (NC_018721); Streptococcus thermophilus (YP_820832), Streptococcus mutans (WP_061046374, WP_024786433); Listeria innocua (NP_472073); Listeria monocytogenes (WP_061665472); Legionella pneumophila (WP_062726656); Staphylococcus aureus (WP_001573634); Francisella tularensis (WP_032729892, WP_014548420), Enterococcus faecalis (WP_033919308); Lactobacillus rhamnosus (WP_048482595, WP_032965177); and Neisseria meningitidis (WP_061704949, YP_002342100); all of which sequences (as entered by the date of filing of this application) are herein incorporated by reference. Any of these sequences or a variant thereof comprising a sequence having at least about 70-100% sequence identity thereto, including any percent identity within this range, such as 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, or 99% sequence identity thereto, can be used for genome editing, as described herein, wherein the variant retains biological activity, such as Cas9 site-directed endonuclease activity. See also Fonfara et al. (2014) Nucleic Acids Res. 42(4):2577-90; Kapitonov et al. (2015) J. Bacteriol. 198(5):797- 807, Shmakov et al. (2015) Mol. Cell.60(3):385-397, and Chylinski et al. (2014) Nucleic Acids Res. 42(10):6091-6105); for sequence comparisons and a discussion of genetic diversity and phylogenetic analysis of Cas9.
[0070] The terms "polynucleotide," "oligonucleotide," "nucleic acid" and "nucleic acid molecule" are used herein to include a polymeric form of nucleotides of any length, either ribonucleotides or deoxyribonucleotides. This term refers only to the primary structure of the molecule. Thus, the term includes triple-, double- and single-stranded DNA, as well as triple- , double- and single-stranded RNA. It also includes modifications, such as by methylationand / or by capping, and unmodified forms of the polynucleotide. More particularly, the terms "polynucleotide," "oligonucleotide," "nucleic acid" and "nucleic acid molecule" include polydeoxyribonucleotides (containing 2-deoxy-D-ribose), polyribonucleotides (containing D- ribose), any other type of polynucleotide which is an N- or C-glycoside of a purine or pyrimidine base, and other polymers containing nonnucleotidic backbones, for example, polyamide (e.g., peptide nucleic acids (PNAs)) and polymorpholino (commercially available from the Anti-Virals, Inc., Corvallis, Oreg., as Neugene) polymers, and other synthetic sequence-specific nucleic acid polymers providing that the polymers contain nucleobases in a configuration which allows for base pairing and base stacking, such as is found in DNA and RNA. There is no intended distinction in length between the terms "polynucleotide," "oligonucleotide," "nucleic acid" and "nucleic acid molecule," and these terms will be used interchangeably. Thus, these terms include, for example, 3ƍ-deoxy-2',5ƍ-DNA, oligodeoxyribonucleotide N3ƍ P5ƍ phosphoramidates, 2'-O-alkyl-substituted RNA, double- and single-stranded DNA, as well as double- and single-stranded RNA, microRNA, DNA:RNA hybrids, and hybrids between PNAs and DNA or RNA, and also include known types of modifications, for example, labels which are known in the art, methylation, "caps," substitution of one or more of the naturally occurring nucleotides with an analog (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolo-pyrimidine, 3-methyl adenosine, C5-propynylcytidine, C5- propynyluridine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-methylcytidine, 7- deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine, and 2-thiocytidine), internucleotide modifications such as, for example, those with uncharged linkages (e.g., methyl phosphonates, phosphotriesters, phosphoramidates, carbamates, etc.), with negatively charged linkages (e.g., phosphorothioates, phosphorodithioates, etc.), and with positively charged linkages (e.g., aminoalklyphosphoramidates, aminoalkylphosphotriesters), those containing pendant moieties, such as, for example, proteins (including nucleases, toxins, antibodies, signal peptides, poly-L-lysine, etc.), those with intercalators (e.g., acridine, psoralen, etc.), those containing chelators (e.g., metals, radioactive metals, boron, oxidative metals, etc.), those containing alkylators, those with modified linkages (e.g., alpha anomeric nucleic acids, etc.), as well as unmodified forms of the polynucleotide or oligonucleotide. The term also includes locked nucleic acids (e.g., comprising a ribonucleotide that has a methylene bridge between the 2'-oxygen atom and the 4'-carbon atom). See, for example, Kurreck et al. (2002) Nucleic Acids Res.30: 1911-1918; Elayadi et al. (2001) Curr. Opinion Invest. Drugs 2: 558-561; Orum et al. (2001) Curr. Opinion Mol. Ther. 3: 239-243; Koshkin et al. (1998)Tetrahedron 54: 3607-3630; Obika et al. (1998) Tetrahedron Lett.39: 5401-5404. All nucleic acid sequences are provided from 5’ to 3’ unless otherwise indicated.
[0071] The terms "hybridize" and "hybridization" refer to the formation of complexes between nucleotide sequences which are sufficiently complementary to form duplexes via Watson-Crick base pairing.
[0072] In general, "identity" refers to an exact nucleotide to nucleotide or amino acid to amino acid correspondence of two polynucleotides or polypeptide sequences, respectively. Percent identity can be determined by a direct comparison of the sequence information between two molecules by aligning the sequences, counting the exact number of matches between the two aligned sequences, dividing by the length of the shorter sequence, and multiplying the result by 100. Readily available computer programs can be used to aid in the analysis, such as ALIGN, Dayhoff, M.O. in Atlas of Protein Sequence and Structure M.O. Dayhoff ed., 5 Suppl. 3:353358, National Biomedical Research Foundation, Washington, DC, which adapts the local homology algorithm of Smith and Waterman Advances in Appl. Math. 2:482489, 1981 for peptide analysis. Programs for determining nucleotide sequence identity are available in the Wisconsin Sequence Analysis Package, Version 8 (available from Genetics Computer Group, Madison, WI) for example, the BESTFIT, FASTA and GAP programs, which also rely on the Smith and Waterman algorithm. These programs are readily utilized with the default parameters recommended by the manufacturer and described in the Wisconsin Sequence Analysis Package referred to above. For example, percent identity of a particular nucleotide sequence to a reference sequence can be determined using the homology algorithm of Smith and Waterman with a default scoring table and a gap penalty of six nucleotide positions. Unless otherwise indicated, all nucleic acid and / or amino acid sequences include sequences that are at least 90% identical to a nucleic acid sequence disclosed herein, e.g., at least 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 99.9%, identical to a nucleic acid and / or amino acid sequence disclosed herein as determined by the Smith and Waterman algorithm.
[0073] As used herein, the terms "complementary" or "complementarity" refers to polynucleotides that are able to form base pairs with one another. Base pairs are typically formed by hydrogen bonds between nucleotide units in an anti-parallel orientation between polynucleotide strands. Complementary polynucleotide strands can base pair in a Watson- Crick manner (e.g., A to T, A to U, C to G), or in any other manner that allows for the formation of duplexes. As persons skilled in the art are aware, when using RNA as opposed to DNA,uracil (U) rather than thymine (T) is the base that is considered to be complementary to adenosine. However, when a uracil is denoted in the context of the present disclosure, the ability to substitute a thymine is implied, unless otherwise stated. "Complementarity" may exist between two RNA strands, two DNA strands, or between a RNA strand and a DNA strand. It is generally understood that two or more polynucleotides may be "complementary" and able to form a duplex despite having less than perfect or less than 100% complementarity. Two sequences are "perfectly complementary" or "100% complementary" if at least a contiguous portion of each polynucleotide sequence, comprising a region of complementarity, perfectly base pairs with the other polynucleotide without any mismatches or interruptions within such region. Two or more sequences are considered "perfectly complementary" or "100% complementary" even if either or both polynucleotides contain additional non-complementary sequences as long as the contiguous region of complementarity within each polynucleotide is able to perfectly hybridize with the other. "Less than perfect" complementarity refers to situations where less than all of the contiguous nucleotides within such region of complementarity are able to base pair with each other. Determining the percentage of complementarity between two polynucleotide sequences is a matter of ordinary skill in the art. For purposes of Cas9 targeting, a gRNA may comprise a sequence "complementary" to a target sequence (e.g., major or minor allele), capable of sufficient base-pairing to form a duplex (i.e., the gRNA hybridizes with the target sequence). Additionally, the gRNA may comprise a sequence complementary to a sequence adjacent to a PAM sequence, wherein the gRNA also hybridizes with the sequence adjacent to a PAM sequence in a target DNA.
[0074] A "target site" or "target sequence" is the nucleic acid sequence recognized (i.e., sufficiently complementary for hybridization) by a guide RNA (gRNA) or a homology arm of a donor polynucleotide. The target site may be allele-specific (e.g., a major or minor allele).
[0075] The terms “target edit site” or “target edit locus” or “edit locus” refer to a target site in the host cell genome comprising a nucleic acid sequence recognized by a guide RNA (gRNA) or a homology arm of a donor polynucleotide that is or was edited by the methods of the disclosure.
[0076] "Recombinant" as used herein to describe a nucleic acid molecule means a polynucleotide of genomic, cDNA, viral, semisynthetic, or synthetic origin which, by virtue of its origin or manipulation, is not associated with all or a portion of the polynucleotide with which it is associated in nature. The term "recombinant" as used with respect to a protein orpolypeptide means a polypeptide produced by expression of a recombinant polynucleotide. In general, the gene of interest is cloned and then expressed in transformed organisms, as described further below. The host organism expresses the foreign gene to produce the protein under expression conditions.
[0077] The term "transformation" refers to the insertion of an exogenous polynucleotide into a host cell, irrespective of the method used for the insertion. For example, direct uptake, transduction or f-mating are included. The exogenous polynucleotide may be maintained as a non-integrated vector, for example, a plasmid, or alternatively, may be integrated into the host genome.
[0078] "Recombinant host cells," "host cells," "cells." "cell lines," "cell cultures," and other such terms denoting microorganisms or higher eukaryotic cell lines cultured as unicellular entities refer to cells which can be, or have been, used as recipients for recombinant vector or other transferred DNA, and include the original progeny of the original cell which has been transfected.
[0079] A "coding sequence" or a sequence which "encodes" a selected polypeptide, is a nucleic acid molecule which is transcribed (in the case of DNA) and translated (in the case of mRNA) into a polypeptide in vivo when placed under the control of appropriate regulatory sequences (or "control elements"). The boundaries of the coding sequence can be determined by a start codon at the 5ƍ (amino) terminus and a translation stop codon at the 3ƍ (carboxy) terminus. A coding sequence can include, but is not limited to, cDNA from viral, prokaryotic or eukaryotic mRNA, genomic DNA sequences from viral or prokaryotic DNA, and even synthetic DNA sequences. A transcription termination sequence may be located 3ƍ to the coding sequence. The coding sequence may be interrupted by introns which can be self- splicing group I or group II introns or those which are spliced out by the host cell splicing machinery,
[0080] Typical "control elements," include, but are not limited to, transcription promoters, transcription enhancer elements, introns (located anywhere in the transcript), transcription termination signals, polyadenylation sequences (located 3ƍ to the translation stop codon), sequences for optimization of initiation of translation (located 5ƍ to the coding sequence), and translation termination sequences.
[0081] "Operably linked" refers to an arrangement of elements wherein the components so described are configured so as to perform their usual function. Thus, a given promoter operablylinked to a coding sequence is capable of effecting the expression of the coding sequence when the proper enzymes are present. The promoter need not be contiguous with the coding sequence, so long as it functions to direct the expression thereof. Thus, for example, intervening untranslated yet transcribed sequences can be present between the promoter sequence and the coding sequence and the promoter sequence can still be considered "operably linked" to the coding sequence.
[0082] "Expression cassette" or "expression construct" refers to an assembly which is capable of directing the expression of the sequence(s) or gene(s) of interest. An expression cassette generally includes control elements, as described above, such as a promoter which is operably linked to (so as to direct transcription of) the sequence(s) or gene(s) of interest, and often includes a polyadenylation sequence as well. Within certain embodiments of the present disclosure, the expression cassette described herein may be contained within a plasmid or viral vector construct (e.g., a vector for genome modification comprising a genome editing cassette comprising a promoter operably linked to a polynucleotide encoding a guide RNA and a donor polynucleotide). In addition to the components of the expression cassette, the construct may also include, one or more selectable markers, a signal which allows the construct to exist as single stranded DNA (e.g., a M13 origin of replication), at least one multiple cloning site, and a "mammalian" origin of replication (e.g., a SV40 or adenovirus origin of replication) or “yeast” origin of replication (e.g. a 2-micron vector or centromeric vector with an autonomously replicating sequence (ARS)).
[0083] The term "transfection" is used to refer to the uptake of foreign DNA by a cell. A cell has been "transfected" when exogenous nucleic acids have been introduced inside the cell membrane. A number of transfection techniques are generally known in the art. See, e.g., Graham et al. (1973) Virology, 52:456, Sambrook et al. (2001) Molecular Cloning, a laboratory manual, 3rd edition, Cold Spring Harbor Laboratories, New York, Davis et al. (1995) Basic Methods in Molecular Biology, 2nd edition, McGraw-Hill, and Chu et al. (1981) Gene 13:197. Such techniques can be used to introduce one or more exogenous nucleic acids moieties into suitable host cells. The term refers to both stable and transient uptake of the genetic material and includes uptake of peptide- or antibody-linked nucleic acids.
[0084] A "vector" is capable of transferring nucleic acid sequences to target cells (e.g., viral vectors, non-viral vectors, particulate carriers, and liposomes). Typically, "vector construct," "expression vector," and "gene transfer vector," mean any nucleic acid construct capable ofdirecting the expression of a nucleic acid of interest and which can transfer nucleic acid sequences to target cells. Thus, the term includes cloning and expression vehicles, as well as plasmid and viral vectors.
[0085] The terms "variant", "analog" and "mutein" refer to modifications of a nucleic acid and / or amino acid sequence disclosed herein. For example, a nucleic acid sequence of the disclosure can include a variant having one, two, three, four, five or more nucleotide substitutions of the reference sequence. In some embodiments, a variant of a nucleic acid sequence of the disclosure can hybridize to a sequence that is complementary to the original reference sequence. The terms "variant", "analog" and "mutein" also refer biologically active derivatives of the reference molecule that retain desired activity, such as site-directed Cas9 endonuclease activity. For example, the terms "variant" and "analog" refer to compounds having a native polypeptide sequence and structure with one or more amino acid additions, substitutions (generally conservative in nature) and / or deletions, relative to the native molecule, so long as the modifications do not destroy biological activity and which are "substantially identical" to the reference molecule as defined below. In general, the amino acid sequences of such variants and analogs will have a high degree of sequence identity to the reference sequence, e.g., an amino acid sequence having greater than or equal to 50%, 60%, 70%, 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% sequence identity, when the two sequences are aligned. Often, the variants and analogs will include the same number of amino acids but will include substitutions, as explained herein. The term "mutein" further includes polypeptides having one or more amino acid-like molecules including but not limited to compounds comprising only amino and / or imino molecules, polypeptides containing one or more analogs of an amino acid (including, for example, unnatural amino acids, etc.), polypeptides with substituted linkages, as well as other modifications known in the art, both naturally occurring and non-naturally occurring (e.g., synthetic), cyclized, branched molecules and the like. The term also includes molecules comprising one or more N-substituted glycine residues (a "peptoid") and other synthetic amino acids or peptides. (See, e.g., U.S. Patent Nos.5,831,005; 5,877,278; and 5,977,301; Nguyen et al., Chem. Biol. (2000) 7:463-473; and Simon et al., Proc. Natl. Acad. Sci. USA (1992) 89:9367–9371 for descriptions of peptoids). Methods for making polypeptide analogs and muteins are known in the art.
[0086] As explained above, analogs generally include substitutions that are conservative in nature, i.e., those substitutions that take place within a family of amino acids that are related in their side chains. Specifically, amino acids are generally divided into four families: (1) acidic-- aspartate and glutamate; (2) basic -- lysine, arginine, histidine; (3) non-polar -- alanine, valine, leucine, isoleucine, proline, phenylalanine, methionine, tryptophan; and (4) uncharged polar -- glycine, asparagine, glutamine, cysteine, serine threonine, and tyrosine. Phenylalanine, tryptophan, and tyrosine are sometimes classified as aromatic amino acids. For example, it is reasonably predictable that an isolated replacement of leucine with isoleucine or valine, an aspartate with a glutamate, a threonine with a serine, or a similar conservative replacement of an amino acid with a structurally related amino acid, will not have a major effect on the biological activity. For example, the polypeptide of interest may include up to about 5-10 conservative or non-conservative amino acid substitutions, or even up to about 15-25 conservative or non-conservative amino acid substitutions, or any integer between 5-25, so long as the desired function of the molecule remains intact. One of skill in the art may readily determine regions of the molecule of interest that can tolerate change by reference to Hopp / Woods and Kyte-Doolittle plots, well known in the art.
[0087] "Gene transfer" or "gene delivery" refers to methods or systems for reliably inserting DNA or RNA of interest into a host cell. Such methods can result in transient expression of non-integrated transferred DNA, extrachromosomal replication and expression of transferred replicons (e.g., episomes), or integration of transferred genetic material into the genomic DNA of host cells. Gene delivery expression vectors include, but are not limited to, vectors derived from bacterial plasmid vectors, viral vectors, non-viral vectors, adenoviruses, retroviruses, alphaviruses, pox viruses, and vaccinia viruses.
[0088] As used herein, the term “heterologous” refers to biological material that is introduced, inserted, or incorporated into a recipient (e.g., host) organism that originates from another organism. Typically, the heterologous material that is introduced into the recipient organism (e.g., a host cell) is not normally found in that organism. Heterologous material can include, but is not limited to, nucleic acids, amino acids, peptides, proteins, and structural elements such as genes, promoters, and cassettes. A host cell can be, but is not limited to, a bacterium, a yeast cell, a mammalian cell, or a plant cell. The introduction of heterologous material into a host cell or organism can result, in some instances, in the expression of additional heterologous material in or by the host cell or organism. As a non-limiting example, the transformation of a yeast host cell with an expression vector that contains DNA sequences encoding a bacterial protein may result in the expression of the bacterial protein by the yeast cell. The incorporation of heterologous material may be permanent or transient. Also, the expression of heterologous material may be permanent or transient.Methods for simultaneously detecting RNA and genomic DNA (gDNA) within the same cell
[0089] Described herein are compositions and methods for simultaneously detecting RNA and genomic DNA (gDNA) within the same cell. The compositions and methods can be used for multiplexed polymerase chain reaction (PCR) and cell barcoding of PCR fragments within droplets. The instant disclosure provides the following advantages. First, the inventors have demonstrated that targeted read-out of both transcriptomic and gDNA is feasible using the methods of the disclosure. Second, separate library construction enables sequencing using different sequencing modalities and coverage. Third, the methods of the disclosure are sensitive enough to allow detection of gene expression changes and is therefore amendable for screening purposes. Fourth, the scale of the methods described herein (10,000 cells in one experiment) is a significant improvement over existing microtiter plate well-based methods (2), while coverage of targets is better than in splitseq based approaches (3). Thus, the instant methods have the advantage of directly reading out the loci of interest in a targeted fashion having high sensitivity and cost-effective sequencing.
[0090] In one aspect, the method for simultaneously detecting RNA and genomic DNA (gDNA) within the same cell comprises providing a single cell suspension as a starting material for obtaining genomic DNA and RNA molecules. The single cell suspension can comprise fixed and permeabilized cells, which allows enzymes and oligonucleotides to enter the cells which maintain cell membrane integrity. Fixed and permeabilized cells can then be contacted with a reverse transcriptase, dNTPs and a primer (e.g., an RT primer) in a container to generate cDNA molecules by in-situ reverse transcription (RT) of the RNA molecules in the cell. In some embodiments, the RT primer comprises a sequence that binds to mRNA. In some embodiments, the sequence that binds to mRNA comprises oligo(dT) or a sequence that hybridizes to any region of the mRNA.
[0091] In some embodiments, the RT primer further comprises a sequence selected from the group consisting of a unique molecular identifier (UMI) sequence, an extension sequence for primer binding, a sample barcode (SBC) sequence, a capture sequence (CS), and combinations thereof. In some embodiments, the CS provides an overhang sequence that is incorporated into a double stranded PCR product (see below) and is complementary to a sequence attached to a solid support, such as a barcoded bead. In some embodiments, the RT primer comprises a capture sequence (CS), an extension sequence for primer binding, a sample barcode sequence, a unique molecular identifier (UMI) sequence, and a poly dT sequence.
[0092] In some embodiments, the RT primer comprises, from 5’ to 3’, a CS with an optional extension sequence for an optimal primer binding site, a SBC sequence, a UMI sequence, and a sequence that binds to mRNA (e.g., oligo(dT) or a sequence that hybridizes to any region of the mRNA). In some embodiments, the RT primer comprises a nucleic acid sequence selected from Table 1 (SEQ ID NOs:14-21).
[0093] In some embodiments, the CS comprises the nucleic acid sequence GTACTCGCAGTAGTC (SEQ ID NO:1). In some embodiments, the CS hybridizes to a complementary sequence attached to a solid support described herein. In some embodiments, the complementary sequence attached to the solid support further comprises a cell barcoding sequence.
[0094] In some embodiments, the SBC sequence comprises a known sequence of variable length. In some embodiments, the SBC sequence is from 4 to 50 nucleotides in length, e.g., 4 to 50, 4 to 40, 4 to 30, 4 to 20, 4 to 15, 4 to 14, 4 to 13, 4 to 12 nucleotides in length, any sub- range between 4 and 50 nucleotides in length, or 4, 5, 6, 7, 8, 9, 10, 11, or 12 nucleotides in length. In some embodiments, the SBC comprises a sequence selected from the group consisting of TCGCCTTA (SEQ ID NO:4), CTAGTACG (SEQ ID NO:5), TTCTGCCT (SEQ ID NO:6), GCTCAGGA (SEQ ID NO:7), AGGAGTCC (SEQ ID NO:8), CATGCCTA (SEQ ID NO:9), GTAGAGAG (SEQ ID NO:10), and CCTCTCTG (SEQ ID NO:11).
[0095] In some embodiments, the UMI sequence comprises a random sequence of variable length. In some embodiments, the UMI sequence is from 4 to 50 nucleotides in length, e.g., 4 to 50, 4 to 40, 4 to 30, 4 to 20, 4 to 15, 4 to 14, 4 to 13, 4 to 12 nucleotides in length, any sub- range between 4 and 50 nucleotides in length, or 4, 5, 6, 7, 8, 9, 10, 11, or 12 nucleotides in length. In some embodiments, the UMI sequence comprises the nucleic acid sequence NNNNNNNN (SEQ ID NO:12), where N is any nucleotide.
[0096] In some embodiments, the extension sequence for primer binding comprises the sequence GACACGTC (SEQ ID NO:29).
[0097] The container for the RT reaction can be any suitable space for holding a liquid volume that is isolated from other containers, including but not limited to a well of a microtiter plate or a microfluidic droplet. Following the RT step, the fixed and permeabilized cells are lysed in a reaction mixture comprising a second primer (e.g., a reverse primer) that hybridizes to both the cDNA and gDNA molecules and contains either a R2N or R25’ overhang sequence.The reaction mixture can be present in a separate (second) container, such as a microfluidic droplet (e.g., a first droplet).
[0098] The reaction mixture comprising the cell lysate and reverse primer can then be combined with PCR reagents and two different forward primers, where one forward primer hybridizes to the cDNA molecules and the second forward primer hybridizes to the gDNA molecules. In some embodiments, the forward primer comprises a capture sequence (CS) 5’ overhang sequence. In some embodiments, the forward primer that hybridizes to the cDNA molecules comprises a 5’ CS overhang and a sample barcode (SBC) sequence. In some embodiments, the SBC sequences comprises a known sequence of variable length. In some embodiments, the SBC sequence is from 4 to 50 base-pairs in length, e.g., 4 to 50, 4 to 40, 4 to 30, 4 to 20, 4 to 15, 4 to 14, 4 to 13, or 4 to 12 base-pairs in length, or any sub-range between 4 and 50 base-pairs in length. In some embodiments, the SBC comprises a sequence selected from the group consisting of TCGCCTTA (SEQ ID NO:4), CTAGTACG (SEQ ID NO:5), TTCTGCCT (SEQ ID NO:6), and GCTCAGGA (SEQ ID NO:7). In some embodiments, the forward primer that hybridizes to the gDNA molecules comprises a 5’ CS overhang and a gDNA specific sequence.
[0099] Reverse transcription of mRNA in the cell produces first strand cDNA molecules that can then be amplified with appropriate forward and reverse primers. The gDNA present in the same cell can also be amplified with appropriate forward and reverse primers. As described above, the amplification reaction can take place in a separate or different container or droplet than the RT reaction. In some embodiments, the forward and / or reverse primers used to amplify the cDNA and gDNA molecules comprise the same 5’ sequence region and a different 3’ sequence region that hybridizes to complementary sequences in the cDNA or gDNA.
[0100] Thus, in some embodiments, the forward primer used to amplify cDNA comprises a 5’ CS and a 3’ SBC sequence, and the reverse primer used to amplify cDNA comprises a 5’ R2N sequence and a 3’ sequence that hybridizes to complementary sequences in the cDNA. In some embodiments, the forward primer used to amplify cDNA comprises a sequence selected from the group consisting of SEQ ID NOs: 22-25. In some embodiments, the reverse primer used to amplify cDNA comprises GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG (SEQ ID NO:2).
[0101] In some embodiments, the forward primer used to amplify cDNA comprises a 5’ CS and a 3’ extension sequence for primer binding, and the reverse primer used to amplify cDNAcomprises a 5’ R2 sequence and a 3’ sequence that hybridizes to complementary sequences in the cDNA. In some embodiments, the forward primer used to amplify cDNA comprises GTACTCGCAGTAGTCGACACGTC (SEQ ID NO:30), and the reverse primer used to amplify cDNA comprises GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT (SEQ ID NO:3) and a 3’ sequence that hybridizes to complementary sequences in the cDNA.
[0102] In some embodiments, the forward primer used to amplify gDNA comprises a 5’ CS and a 3’ sequence that hybridizes to complementary sequences in the gDNA, and the reverse primer used to amplify gDNA comprises a 5’ R2N and a 3’ sequence that hybridizes to complementary sequences in the gDNA. In some embodiments, the forward primer used to amplify gDNA comprises SEQ ID NO:1 and a 3’ sequence that hybridizes to complementary sequences in the gDNA. In some embodiments, the reverse primer used to amplify gDNA comprises SEQ ID NO:2 and a 3’ sequence that hybridizes to complementary sequences in the gDNA.
[0103] In some embodiments, the second droplet further comprises one or more cell barcode oligonucleotides. Each cell barcode oligonucleotide comprises a cell barcode (CBC) sequence and a sequence complementary to the CS. Therefore, when the first set of PCR products are generated within the second droplet, the PCR products bind to the cell barcode oligonucleotides in the hybridization of the CS, thereby introducing the CBC to the PCR products, resulting in a second set of PCR products. As disclosed herein, each second set of PCR product comprises the CS and the CBC. In some embodiments, the cell barcode oligonucleotides are attached to a solid support.
[0104] In some embodiments, the PCR reagents and forward primer are present in a second microfluidic droplet that is fused or merged with the first droplet comprising the reaction mixture comprising the cell lysate and reverse primer. Exemplary embodiments are illustrated in Figs.1 and 7a.
[0105] Exemplary sequences useful in the practice of the methods are provided in Table 1 below. Table 1. Exemplary sequences for amplifying cDNA and gDNA from the same cell.Solid Support
[0106] In some embodiments, the container comprising the PCR reagents (e.g., in the second droplet) comprises a solid support attached to oligonucleotides comprising sequences that are complementary to and hybridize to the reverse-complement of the CS sequence present in the double-stranded PCR product. In some embodiments, the oligos attached to the solid support comprise an R1N sequence (that is useful for Illumina® based sequencing), a cell barcode sequence, and a UMI sequence. In some embodiments, the oligos attached to the solid support comprise an R1N sequence (that is useful for Illumina® based sequencing), a first cell barcode sequence (CBC1), an off-staggered constant space linker (CSL), a second cell barcode sequence (CBC2), and a capture sequence (CS). In some embodiments, the off-staggered CSL can be a variable-length nucleic acid sequence. In some embodiments, the CSL has a length of 5 to 50 bp, 10 to 40 bp, 10 to 30 bp, or 10 to 20 bp, 20 to 40 bp, or 15 to 35 bp. In some embodiments, the CSL has a length of 14 to 17 bp. The off-staggered CSL separates the two CBCs for better sequencing quality. The CS facilitates capture target RNAs or gDNAs for cell barcoding.
[0107] Exemplary sequences attached to the solid support are provided in Table 2 below. Table 2. Exemplary sequences attached to solid support.
[0108] In some embodiments, the solid support is a bead, such as a gel bead. In some embodiments, the oligonucleotides are removed from the bead prior to the PCR reaction. Exemplary beads include those provided by 10X Genomics (San Francisco, CA) and Mission Bio (South San Francisco, CA). An exemplary barcoded bead is shown in Fig.6.
[0109] In some embodiments of the method, the first step comprises contacting fixed and permeabilized cells with reverse transcriptase, dNTPs, and a reverse transcription primer to generate cDNA (“reverse transcription step”). Two microfluidic droplets are then generated consecutively and targeted multiplexed PCR is performed in each droplet for both the cDNA and gDNA targets. In some embodiments, the first droplet contains Proteinase K and one or more reverse primers. Thus, in some embodiments, in the first droplet cells from the RT step are lysed and different reverse primers comprising a R2N 5’ overhangs are used to amplify the cDNA targets or the gDNA targets. The reverse primers used to amplify the cDNA targets comprise a R2N 5’ overhang sequence and a cDNA specific sequence. The reverse primers used to amplify the gDNA targets comprise a R2N 5’ overhang sequence and a gDNA specific sequence. The second drop is produced by fusing or combining the first droplet with a droplet comprising PCR reagents, a forward primer comprising a CS 5’ overhang for amplifying the target cDNA or gDNA sequence, and a solid support attached to oligonucleotides comprising sequences that are complementary to and hybridize to the reverse-complement of the CS sequence present in the double-stranded PCR product. PCR reagents typically include athermostable polymerase and dNTPs. In some embodiments, the oligonucleotides attached to the solid support comprises a cell barcode sequence. In some embodiments, the solid support is referred to as a cell barcoding bead.
[0110] In another aspect, the method is modified to enable separate sequencing of both cDNA and gDNA libraries from the same single cell. Thus, in some embodiments, the first droplet comprises a first plurality of reverse primers comprising a R25’ overhang and a cDNA specific sequence that is used to amplify each cDNA target, and a second plurality of reverse primers comprising a R2N 5’ overhang and a gDNA specific sequence that is used to amplify each gDNA target. As above, the first droplet also contains reagents for lysing cells, such as proteinase-K, following the RT step. As above, the first droplet is fused or combined with a second droplet comprising PCR reagents, a forward primer comprising a CS overhang for amplifying the cDNA and gDNA sequences, and a solid support attached to oligonucleotides comprising sequences that are complementary to and hybridize to the reverse-complement of the CS sequence present in the double-stranded PCR product. In some embodiments, the oligonucleotides attached to the solid support comprise a cell barcode sequence.
[0111] In some embodiments, the solid support comprises a bead, and the oligonucleotides attached to the bead comprise a photocleavable linker (see exemplary oligonucleotides in Fig. 6). During the second droplet step, the oligonucleotides attached to the bead can be contacted with UV light, thereby releasing the oligonucleotides from the bead surface into solution. Without being bound by theory, releasing the oligonucleotides from the bead into solution may increase the PCR efficiency. Examples of suitable beads are provided by Mission Bio, Inc. (South San Francisco, CA). In some embodiments, each droplet contains only 1 bead on average in order to provide single cell resolution.
[0112] In some embodiments, the solid support comprises a bead that is capable of being dissolved by heat, and the oligonucleotides attached to the bead can be released from the bead surface into solution by heating the droplet during the second droplet step. Examples of suitable beads are provided by 10X Genomics (Pleasanton, CA). Screening of Genetically Modified Cells
[0113] In another aspect, the methods of the disclosure further comprise screening cDNA and gDNA from genetically modified (e.g., genetically edited) cells. In some embodiments, the cells are genetically modified using CRISPR / Cas technology. In some embodiments, the cells are genetically modified using CRISP interference (“CRISPRi”). The CRISPRi systemuses a catalytically inactive Cas9 (dCas9) protein (e.g., comprising two point mutations D10A and H840A) that lacks endonuclease activity to regulate genes in an RNA-guided manner, and can be used to block transcription of target genes without cutting the target DNA. The dCas9- single guide RNA (sgRNA) complex binds to DNA elements complementary to the sgRNA and causes a steric block that halts transcript elongation by RNA polymerase, resulting in repression of the target gene. In some embodiments, the CRISPRi system comprises a dCas9 fused to repressor proteins (e.g., SALL1 and SDS3), and a guide RNA specifically designed to target the region immediately downstream of a gene’s transcriptional start site (TSS). The guide RNA associates with dCas9-SALL1-SDS3 and directs the repressor complex to the DNA target site. In some embodiments, the guide RNAs are designed to target the transcription start site of genes. In some embodiments, the guide RNAs are designed to target sequences that are transcribed into mRNA (referred to as “cDNA targets”).
[0114] Following genetic modification of the cells, the cells are fixed and permeabilized as described above, and the cells are contacted with reverse transcriptase, dNTPS, and a RT primer to produce cDNA. In some embodiments, the RT primer comprises one or more of a CS sequence, an extension sequence for primer binding, a sample barcode sequence, and a poly dT sequence. After first strand cDNA is produced, the cells are added to a first droplet comprising a first reverse primer comprising a CS sequence and R2 overhang that is used to amplify each cDNA target and a second reverse primer comprising a CS sequence and R2N overhang that is used to amplify each gDNA target. In some embodiments, the first droplet also contains reagents for lysing cells, such as proteinase-K. The first droplet is the fused or combined with a second droplet comprising PCR reagents, a forward primer comprising a CS overhang for amplifying the cDNA and gDNA target sequences, and a solid support attached to oligonucleotides and a solid support attached to oligonucleotides comprising sequences that are complementary to and hybridize to the reverse-complement of the CS sequence present in the double-stranded PCR product. In some embodiments, the oligonucleotides attached to the solid support comprises a cell barcode sequence.
[0115] The amplified cDNA and gDNA libraries can then be sequenced to determine RNA expression levels of the targeted genes. In some embodiments, the cDNA and gDNA libraries are sequenced using next generation sequencing (Illumina®).
[0116] The RNA-guided nuclease (e.g., dCas9) can be provided in the form of a protein, such as the nuclease complexed with a gRNA, or provided by a nucleic acid encoding the RNA-guided nuclease, such as an RNA (e.g., messenger RNA) or DNA (expression vector). Codon usage may be optimized to improve production of an RNA-guided nuclease in a particular cell or organism. For example, a nucleic acid encoding an RNA-guided nuclease can be modified to substitute codons having a higher frequency of usage in a yeast cell, a bacterial cell, a human cell, a non-human cell, a mammalian cell, a rodent cell, a mouse cell, a rat cell, or any other host cell of interest, as compared to the naturally occurring polynucleotide sequence. When a nucleic acid encoding the RNA-guided nuclease is introduced into cells, the protein can be transiently, conditionally, or constitutively expressed in the cell.
[0117] Donor polynucleotides and gRNAs are readily synthesized by standard techniques, e.g., solid phase synthesis via phosphoramidite chemistry, as disclosed in U.S. Patent Nos. 4,458,066 and 4,415,732, incorporated herein by reference; Beaucage et al., Tetrahedron (1992) 48:2223-2311; and Applied Biosystems User Bulletin No. 13 (1 April 1987). Other chemical synthesis methods include, for example, the phosphotriester method described by Narang et al., Meth. Enzymol. (1979) 68:90 and the phosphodiester method disclosed by Brown et al., Meth. Enzymol. (1979) 68:109. In view of the short lengths of gRNAs (typically about 20 nucleotides in length) and donor polynucleotides (typically about 100-150 nucleotides), gRNA-donor polynucleotide cassettes can be produced by standard oligonucleotide synthesis techniques and subsequently ligated into vectors. Moreover, libraries of gRNA-donor polynucleotide cassettes directed against thousands of genomic targets can be readily created using highly parallel array-based oligonucleotide library synthesis methods (see, e.g., Cleary et al. (2004) Nature Methods 1:241-248, Svensen et al. (2011) PLoS One 6(9):e24906).
[0118] In addition, adapter sequences can be added to oligonucleotides to facilitate high- throughput amplification or sequencing. For example, a pair of adapter sequences can be added at the 5ƍ and 3ƍ ends of an oligonucleotide to allow amplification or sequencing of multiple oligonucleotides simultaneously by the same set of primers. Additionally, restriction sites can be incorporated into oligonucleotides to facilitate cloning of oligonucleotides into vectors. For example, oligonucleotides comprising gRNA-donor polynucleotide cassettes can be designed with a common 5ƍ restriction site and a common 3ƍ restriction site to facilitate ligation into the genome modification vectors. A restriction digest that selectively cleaves each oligonucleotide at the common 5ƍ restriction site and the common 3ƍ restriction site is performed to produce restriction fragments that can be cloned into vectors (e.g., plasmids or viral vectors), followed by transformation of cells with the vectors comprising the gRNA-donor polynucleotide cassettes. A restriction site can also be added in between the gRNA and donor polynucleotidesequences to enable a second cloning step for the introduction of a guide RNA scaffold sequence or other constructs into the vector.
[0119] Amplification of polynucleotides encoding gRNA-donor polynucleotide cassettes may be performed, for example, before ligation into genome modification vectors or before sequencing and after barcoding. Any method for amplifying oligonucleotides may be used, including, but not limited to polymerase chain reaction (PCR), isothermal amplification, nucleic acid sequence-based amplification (NASBA), transcription mediated amplification (TMA), strand displacement amplification (SDA), and ligase chain reaction (LCR). In one embodiment, the genome editing cassettes comprise common 5ƍ and 3ƍ priming sites to allow amplification of the gRNA-donor polynucleotide sequences in parallel with a set of universal primers. In another embodiment, a set of selective primers is used to selectively amplify a subset of the gRNA-donor polynucleotides from a pooled mixture.
[0120] Genome editing may be performed on a single cell or a population of cells of interest and can be performed on any type of cell, including any cell from a prokaryotic, eukaryotic, or archaeon organism, including bacteria, archaea, fungi, protists, plants, and animals. Cells from tissues, organs, and biopsies, as well as recombinant cells, genetically modified cells, cells from cell lines cultured in vitro, and artificial cells (e.g., nanoparticles, liposomes, polymersomes, or microcapsules encapsulating nucleic acids) may all be used in the practice of the present disclosure. The methods of the disclosure are also applicable to editing of nucleic acids in cellular fragments, cell components, or organelles comprising nucleic acids (e.g., mitochondria in animal and plant cells, plastids (e.g., chloroplasts) in plant cells and algae). Cells may be cultured or expanded prior to or after performing genome editing as described herein. In one embodiment, the cells are induced pluripotent stem cells (iPSCs).
[0121] The Cas9 nuclease and gRNAs can be delivered to cells in vitro or in vivo using any method known in the art. In some embodiments, a DNA expression vector comprising a nucleic acid sequence encoding the Cas9 protein operably linked to transcriptional and / or translation control elements such as promoters and terminators is transfected into cells. In some embodiments, an mRNA molecule encoding the Cas9 protein is transfected into cells. In some embodiments, purified Cas9 protein is delivered to the cell cytoplasm or nucleus, for example by encapsulating the protein in a vesicle. Suitable methods include using viral and nonviral vectors and physical methods. Viral vectors include lentiviral, adenovirus (AV), adeno-associated virus (AAV) and retroviral vectors. Physical methods include microinjectionand electroporation of DNA, RNA or protein, cell penetrating peptides, lipofection, lipid-based nanoparticles comprising DNA, RNA or protein, and gold based nanoparticles. Other suitable methods include virus-like particles, extracellular vesicles, transposon systems, chemically- induced transformation, typically using divalent cations (e.g., CaCl2), microinjection. See, e.g., Yip BH. Recent Advances in CRISPR / Cas9 Delivery Strategies. Biomolecules. 2020 May 30;10(6):839. doi: 10.3390 / biom10060839; Sambrook et al. (2001) Molecular Cloning, a laboratory manual, 3rdedition, Cold Spring Harbor Laboratories, New York, Davis et al. (1995) Basic Methods in Molecular Biology, 2ndedition, McGraw-Hill, and Chu et al. (1981) Gene 13:197; which are herein incorporated by reference in their entireties.
[0122] In some embodiments, the gRNAs or libraries thereof are delivered into cells using a viral vector, such as a lentiviral vector. In some embodiments, the gRNAs or libraries thereof are delivered into cells using lipofection. In some embodiments, the gRNAs are prime editing guide RNAs (pegRNAs) that are extended on the 3’ end to install edits using a prime editing system. The prime editing system comprises a modified version of the Cas9 enzyme fused with reverse transcriptase. Compositions Reaction Mixtures
[0123] In another aspect, the disclosure provides reaction mixtures that are useful for amplifying cDNA and gDNA from the same cell. Reaction mixtures typically include a nucleic acid template (e.g., RNA or DNA), an RNA or DNA polymerase, one or more oligo nucleotide primers, and dNTPs. In some embodiments, the reaction mixture comprises RNA and gDNA from a single cell, reverse transcriptase, and a RT primer comprising a sequence that binds to mRNA. In some embodiments, the sequence that binds to mRNA comprises oligo(dT) or a sequence that hybridizes to any region of the mRNA. In some embodiments, the RT primer further comprises a sequence selected from the group consisting of a unique molecular identifier (UMI) sequence, a sample barcode (SBC) sequence, a capture sequence (CS), and combinations thereof. In some embodiments, the RT primer comprises, from 5’ to 3’, a CS, a SBC sequence, a UMI sequence, and a sequence that binds to mRNA (e.g., oligo(dT) or a sequence that hybridizes to any region of the mRNA). In some embodiments, the RT primer comprises, from 5’ to 3’, a CS, an extension sequence for primer binding, a SBC sequence, a UMI sequence, and a sequence that binds to mRNA (e.g., oligo(dT) or a sequence that hybridizes to any region of the mRNA).
[0124] In some embodiments, the CS comprises the nucleic acid sequence GTACTCGCAGTAGTC (SEQ ID NO:1).
[0125] In some embodiments, the SBC sequence comprises a known sequence of variable length. In some embodiments, the SBC sequence is from 4 to 50 nucleotides in length, e.g., 4 to 50, 4 to 40, 4 to 30, 4 to 20, 4 to 15, 4 to 14, 4 to 13, 4 to 12 nucleotides in length, any sub- range between 4 and 50 nucleotides in length, or 4, 5, 6, 7, 8, 9, 10, 11, or 12 nucleotides in length. In some embodiments, the SBC sequences comprises a known sequence of variable length. In some embodiments, the SBC comprises a sequence selected from the group consisting of TCGCCTTA (SEQ ID NO:4), CTAGTACG (SEQ ID NO:5), TTCTGCCT (SEQ ID NO:6), GCTCAGGA (SEQ ID NO:7), AGGAGTCC (SEQ ID NO:8), CATGCCTA (SEQ ID NO:9), GTAGAGAG (SEQ ID NO:10), and CCTCTCTG (SEQ ID NO:11).
[0126] In some embodiments, the UMI sequence comprises a random sequence of variable length. In some embodiments, the UMI sequence is from 4 to 50 nucleotides in length, e.g., 4 to 50, 4 to 40, 4 to 30, 4 to 20, 4 to 15, 4 to 14, 4 to 13, 4 to 12 nucleotides in length, any sub- range between 4 and 50 nucleotides in length, or 4, 5, 6, 7, 8, 9, 10, 11, or 12 nucleotides in length. In some embodiments, the UMI sequence comprises the nucleic acid sequence NNNNNNNN (SEQ ID NO:12), where N is any nucleotide.
[0127] In some embodiments, the extension sequence for primer binding comprises the sequence GACACGTC (SEQ ID NO:29).
[0128] In some embodiments, the forward and / or reverse primers used to amplify the cDNA and gDNA molecules comprise the same 5’ sequence region and a different 3’ sequence region that hybridizes to complementary sequences in the cDNA or gDNA.
[0129] Thus, in some embodiments, the forward primer used to amplify cDNA comprises a 5’ CS and a 3’ SBC sequence, and the reverse primer used to amplify cDNA comprises a 5’ R2N sequence and a 3’ sequence that hybridizes to complementary sequences in the cDNA. In some embodiments, the forward primer used to amplify cDNA comprises a sequence selected from the group consisting of SEQ ID NOs:22-25. In some embodiments, the reverse primer used to amplify cDNA comprises SEQ ID NO:2 and a 3’ sequence that hybridizes to complementary sequences in the cDNA.
[0130] In some embodiments, the forward primer used to amplify cDNA comprises a 5’ CS and a 3’ extension sequence for primer binding, and the reverse primer used to amplify cDNAcomprises a 5’ R2 sequence and a 3’ sequence that hybridizes to complementary sequences in the cDNA. In some embodiments, the forward primer used to amplify cDNA comprises GTACTCGCAGTAGTCGACACGTC (SEQ ID NO:30), and the reverse primer used to amplify cDNA comprises SEQ ID NO:3 and a 3’ sequence that hybridizes to complementary sequences in the cDNA.
[0131] In some embodiments, the forward primer used to amplify gDNA comprises a 5’ CS and a 3’ sequence that hybridizes to complementary sequences in the gDNA, and the reverse primer used to amplify gDNA comprises a 5’ R2N and a 3’ sequence that hybridizes to complementary sequences in the gDNA. In some embodiments, the forward primer used to amplify gDNA comprises SEQ ID NO:1 and a 3’ sequence that hybridizes to complementary sequences in the gDNA. In some embodiments, the reverse primer used to amplify gDNA comprises SEQ ID NO:2 and a 3’ sequence that hybridizes to complementary sequences in the gDNA.
[0132] Thus, in some embodiments, the reaction mixture comprises cDNA molecules produced by reverse transcription of the RNA and a reverse primer that hybridizes to both the cDNA and the gDNA molecules and comprises a R2N or R2 overhang sequence. In some embodiments, the reverse primer comprising the R2N overhang sequence comprises the nucleic acid sequence GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG (SEQ ID NO:2) and the reverse primer comprising the R2 overhang sequence comprises the nucleic acid sequence GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT SEQ ID NO:3).
[0133] In some embodiments, the reaction mixture comprises a polynucleotide of the disclosure. In some embodiments, the reaction mixture comprises a (one or more) polynucleotide(s) of any one of SEQ ID Nos: 1-83. Polynucleotides
[0134] In another aspect, the disclosure provides polynucleotides that are useful in the methods of the disclosure. In some embodiments, the polynucleotide can be used to prime reverse transcription of an mRNA molecule.
[0135] In some embodiments, the polynucleotide comprises a capture sequence (CS). In some embodiments, the CS comprises the nucleic acid sequence GTACTCGCAGTAGTC (SEQ ID NO:1).
[0136] In some embodiments, the polynucleotide further comprises a sample barcode (SBC) sequence, a unique molecular identifier (UMI) sequence, and / or a sequence that binds to mRNA. In some embodiments, the SBC sequence comprises a known sequence of variable length, and the UMI sequence comprises a random sequence of variable length. In some embodiments, the SBC sequence is from 4 to 50 nucleotides in length. In some embodiments, the UMI sequence is from 4 to 50 nucleotides in length, e.g., 4 to 50, 4 to 40, 4 to 30, 4 to 20, 4 to 15, 4 to 14, 4 to 13, 4 to 12 nucleotides in length, any sub-range between 4 and 50 nucleotides in length, or 4, 5, 6, 7, 8, 9, 10, 11, or 12 nucleotides in length. In some embodiments, the SBC sequence comprises a nucleic acid sequence selected from the group consisting of TCGCCTTA (SEQ ID NO:4), CTAGTACG (SEQ ID NO:5), TTCTGCCT (SEQ ID NO:6), GCTCAGGA (SEQ ID NO:7), AGGAGTCC (SEQ ID NO:8), CATGCCTA (SEQ ID NO:9), GTAGAGAG (SEQ ID NO:10), and CCTCTCTG (SEQ ID NO:11). In some embodiments, the UMI sequence is from 4 to 50 nucleotides in length. In some embodiments, the UMI sequence is from 4 to 50 nucleotides in length, e.g., 4 to 50, 4 to 40, 4 to 30, 4 to 20, 4 to 15, 4 to 14, 4 to 13, 4 to 12 nucleotides in length, any sub-range between 4 and 50 nucleotides in length, or 4, 5, 6, 7, 8, 9, 10, 11, or 12 nucleotides in length. In some embodiments, the UMI comprises the nucleic acid sequence NNNNNNNN (SEQ ID NO:12). In some embodiments, the sequence that binds to mRNA comprises oligo(dT) or a sequence that hybridizes to any region of the mRNA. In some embodiments, the polynucleotide comprises a sequence in any one of SEQ ID Nos: 1-83.
[0137] In some embodiments, the polynucleotide comprises a capture sequence (CS), a sample barcode (SBC) sequence, a unique molecular identifier (UMI) sequence, and nucleic acid sequence that binds to mRNA. In some embodiments, the polynucleotide comprises SEQ ID NO:1. In some embodiments, the polynucleotide comprises a nucleic acid sequence selected from the group consisting of SEQ ID NO:1, SEQ ID NOs: 5-11, SEQ ID NO:12, and SEQ ID NO:13.
[0138] In some embodiments, the polynucleotide comprises, from 5’ to 3’, a CS and a sample barcode (SBC) sequence. In some embodiments, the polynucleotide comprises SEQ ID NOs:1 and 4. In some embodiments, the polynucleotide comprises SEQ ID NOs:1 and 5. In some embodiments, the polynucleotide comprises SEQ ID NOs:1 and 6. In some embodiments, the polynucleotide comprises SEQ ID NOs:1 and 7. In some embodiments, the polynucleotide comprises SEQ ID NOs:1 and 8. In some embodiments, the polynucleotide comprises SEQ ID NOs:1 and 9. In some embodiments, the polynucleotide comprises SEQ ID NOs:1 and 10. Insome embodiments, the polynucleotide comprises SEQ ID NOs:1 and 11. In some embodiments, the polynucleotide comprises a nucleic acid sequence selected from a sequence in Table 3. Table 3.
[0139] In some embodiments, the polynucleotide comprises, from 5’ to 3’, a CS, an SBC sequence, and a UMI. In some embodiments, the polynucleotide comprises SEQ ID NOs:1, 4 and 12. In some embodiments, the polynucleotide comprises SEQ ID NOs:1, 5 and 12. In some embodiments, the polynucleotide comprises SEQ ID NOs:1, 6 and 12. In some embodiments, the polynucleotide comprises SEQ ID NOs:1, 7 and 12. In some embodiments, the polynucleotide comprises SEQ ID NOs:1, 8 and 12. In some embodiments, the polynucleotide comprises SEQ ID NOs:1, 9 and 12. In some embodiments, the polynucleotide comprises SEQ ID NOs:1, 10 and 12. In some embodiments, the polynucleotide comprises SEQ ID NOs:1, 11 and 12. In some embodiments, the polynucleotide comprises a nucleic acid sequence selected from a sequence in Table 4.Table 4.
[0140] In some embodiments, the polynucleotide comprises an extension sequence for primer binding. In some embodiments, polynucleotide comprises the sequence GACACGTC (SEQ ID NO:29). In some embodiments, the polynucleotide comprises, from 5’ to 3’, a CS and an extension sequence for primer binding. In some embodiments, the polynucleotide comprises SEQ ID NOs:1 and 29 (GTACTCGCAGTAGTCGACACGTC (SEQ ID NO:30)).
[0141] In some embodiments, the polynucleotide comprises, from 5’ to 3’, a CS, an extension sequence for primer binding, and an SBC sequence. In some embodiments, the polynucleotide comprises a nucleic acid sequence selected from a sequence in Table 5. Table 5.
[0142] In some embodiments, the polynucleotide comprises, from 5’ to 3’, a CS, an extension sequence for primer binding, an SBC sequence, and a UMI sequence. In some embodiments, the polynucleotide comprises a nucleic acid sequence selected from a sequence in Table 6. Table 6.
[0143] In some embodiments, the polynucleotide comprises, from 5’ to 3’, a CS, an extension sequence for primer binding, a SBC sequence, a UMI sequence, and a sequence that binds to an mRNA. In some embodiments, the polynucleotide comprises a nucleic acid sequence selected from a sequence in Table 7.Table 7.EXAMPLES Example 1.
[0144] This example describes an exemplary materials and methods of the disclosure. Fixation
[0145] WTC-11 induced pluripotent stem cells were dissociated into single cells using Accutase (StemCell Technologies - #07922), filtered through a 40 μm cell strainer, counted. 1.5 x10^6 cells were transferred to a new 15 ml conical tube and spun at 500 g for 3 min.
[0146] For the glyoxal fixation condition the supernatant was removed, cells were resuspended in 200 μl of Glyoxal solution fixation solution (3 % Glyoxal, 20 % EtOH, 0.75 % Acetic Acid – Glacial, pH = 4.0) and incubated for 7 min at room temperature. 1 ml ice-cold wash buffer 1 (1x PBS with 2 % BSA, 1 mM DTT and 0.5 U / μl RNasin® Plus Ribonuclease Inhibitor - Promega #N2615) was added and cells were spun at 500 g for 3 min at 4 C. Supernatant was carefully removed and wash step was repeated with wash buffer 1 for a total of 2 washes. Cells were resuspended in 175 μl ice-cold permeabilization buffer (10 mM TRIS- HCl pH 7.5, 10 mM NaCl, 3 mM MgCl2, 0.1 % Tween-20, 0.2 U / μl RNasin® Plus Ribonuclease Inhibitor, 1mM DTT, 2 % BSA, 0.1 % IGEPAL CA-630 and 0.01 % Digitonin) and incubated for 4 min on ice. 1 ml of ice-cold wash buffer 2 (10 mM TRIS pH 7.5, 10 mM NaCl, 3 mM MgCl2, 0.1 % Tween-20, 0.2 U / μl RNasin® Plus Ribonuclease Inhibitor, 1mMDTT and 2 % BSA) was added and tube was gently inverted 4-6 times. Cells were spun at 500 g for 5 min at 4 C and resuspended in ice-cold resuspension buffer (1x PBS, 2 % BSA, 1 mM DTT and 0.2 U / μl RNasin® Plus Ribonuclease Inhibitor), filtered through a 40 μm strainer, counted and diluted to 1.4 x106cells / ml.
[0147] PFA fixation conditions were followed as described elsewhere with adaptions (Rosenberg et al., 2018). In short, the supernatant was removed, cells were resuspended in 1 ml 1x PBS with 0.2 U / μl RNasin® Plus Ribonuclease Inhibitor and 3 ml of 1.3 % PFA solution (in 1x PBS) were added. Cells were fixed for 10 min on ice.160 μl of permeabilization buffer (5 % Trition-X 100 with 0.2 U / μl RNasin® Plus Ribonuclease Inhibitor) was added and cells were incubated for 3 min on ice. Cells were spun at 500 g for 3 min at 4 C, supernatant was carefully removed and cells were resuspended in 500 μl of 1x PBS with 0.2 U / μl RNasin® Plus Ribonuclease Inhibitor. 500 μl of ice-cold 100 mM TRIS-HCl at pH 8.0 was added and mixed by inverting the tube.20 μl of permeabilization buffer was added and mixed by inverting the tube. Cells were spun at 500 g for 3 min at 4 C, supernatant removed, resuspended in 300 μl of 0.5x PBS with 0.2 U / μl RNasin® Plus Ribonuclease Inhibitor, filtered through a 40 μm strainer, counted and diluted to 1.4 x106cells / ml. In-situ reverse transcription
[0148] Reverse transcription master mix consisting of a final concentration of 1x RT Buffer, 0.25 U / μl Enzymatics RNAse Inhibitor (Biozym – 180520), 0.2 U / μl RNasin® Plus Ribonuclease Inhibitor, 500 mM dNTPs and 20 U / μl Maxima H Minus Reverse Transcriptase (ThermoFisher - EP0752) was prepared on ice for in 8 μl for a total reaction volume of 20 μl. 4μl of reverse transcription oligo (12.5 μM) were combined in each 96-well plate with 8 μl reverse transcription master mix. 8 μl fixed and permeabilized cells (10000 cells total) were added to each well, yielding a total reaction volume of 20 μl. Reverse transcription was performed in a thermocycler using the program in the Table 8. All reverse transcription reactions were pooled into a 15 ml conical tube, 1 ml of ice-cold 1x PBS with 1 % BSA was added and cells were spun at 500 g for 5 min and supernatant was removed. Table 8. In-situ reverse transcription protocolReverse transcription oligos
[0149] Common sequence (CS) GTACTCGCAGTAGTC (underlined)
[0150] Extension for primer binding ACACGTC (italicized)
[0151] Sample barcode (SBC) (SSSSSSSS)
[0152] Unique molecular identifier (UMI)
[0153] PolydTVN (TTTTTTTTTTTTTTTTVN)
[0154] RT_primer _V1: GTACTCGCAGTAGTCSSSSSSSSNNNNNNNNTTTTTTTTTTTTTTTTVN
[0155] RT_primer_V2: GTACTCGCAGTAGTCGACACGTCSSSSSSSSNNNNNNNNTTTTTTTTTTTTTTTTVN Droplet-based multiplexed PCR
[0156] Samples were processed using the Tapestri microfluidic device from Mission Bio (MB51-0007, MB51-0010, MB51-0009) according to manufactures protocol withmodifications. In-situ RT processed cell pellet from previous step was resuspended in the cell buffer of Mission Bio, cells were counted and diluted to the appropriate concentration of 3000- 4000 cells / μl. Custom primers were used in the multiplexed droplet PCR amplification step. cDNA primers were designed using the TAPseq primer prediction tool with at targeted Tm of 60 °C (Schraivogel et al., 2020). gDNA primers were designed using the Tapestri Designer (see the internet at designer.missionbio.com). Version 1 gDNA and cDNA both had CS and R2N overhangs. Version 2 gDNA primers had CS and R2N, whereas cDNA primers had CS and R2 overhangs. Forward and reverse stock primers had a concentration of 16 μM and 90 μM for both versions, respectively. Lysis Mix was prepared by mixing the reverse primer mix with lysis buffer and the cartridge was prepared according to the Mission Bio user guide. After the first droplet generation the emulsion was incubated at 50 °C for 60 min and 80 °C for 10 min. Barcoding master mix was prepared and the cartridge was prepared according to the Mission Bio user guide. Second droplets were generated and emulsion was subjected to the targeted PCR protocol (Table 9). Table 9: Droplet-based multiplexed PCR protocolLibrary preparation was followed using the Mission Bio user guide. For Version 1 final sequencing libraries were generated according to the Mission Bio user guide, using the primers listed in Table 10. For Version 2 cDNA and gDNA sequencing libraries were generated separately using the corresponding library amplification primers (Table 11). Droplet amplification oligos Common sequence (CS) GTACTCGCAGTAGTC (underlined). Extension for primer binding ACACGTC (italicized). Sample barcode (SBC) (SSSSSSSS) R2N: GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG (Bold) R2: GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT Table 10: Droplet amplification primers (Version1):Table 11: Droplet amplification primers (Version2):Example 2.
[0157] This example describes an exemplary method of the disclosure.
[0158] We have developed a method to measure RNA and gDNA simultaneously within the same cell in a targeted fashion. This is achieved by combining in-situ reverse transcription (RT) (4) with a multiplexed PCR within droplets (Fig. 1). In a first proof-of-principle we used the commercial droplet generation device of the Tapestri platform from Mission Bio. Cells are dissociated into a single-cell suspension and fixed followed by several washes and a permeabilization step to enable enzymes to enter the cells while maintaining cell membrane integrity. This is followed by an in-situ RT step generating cDNA. Primers used for this RT step contain oligo(dT)s to bind to polyadenylated mRNAs, a unique molecular identifier (UMI), a sample barcode (SBC) that can be used for different experimental conditions and a capture sequence (CS) that is used to introduce a cell barcode (CBC) in subsequent droplet steps. After this the cells contain gDNA and cDNA with an UMI, a SBC and a CS.
[0159] These cells are used as an input for the Tapestri platform from Mission Bio. Within two droplets that are generated consecutively a targeted multiplexed PCR is performed in each droplet for both cDNA and gDNA targets. During the first droplet generation cells are treated with proteinase K yielding a cell lysate and reverse primer is added for each cDNA or gDNA target containing a R2N overhang. This is followed by a second droplet generation where the first droplet is fused with a cell barcoding bead, PCR reagents and a forward primer for each target that contains a CS overhang. Tm conditions of the multiplexed PCR favor binding of the primers to its cDNA and gDNA targets for the first 10 PCR cycles. Tm conditions for the second 10 PCR cycles favor binding of those PCR products to the CS barcoding overhangs introducing a unique CBC within each separate droplet. Note that only cDNA targets receive an UMI and SBC as this is introduced in the in-situ RT step (Fig.1).
[0160] In a first proof-of-principle experiment we tested two different fixation conditions – PFA and glyoxal (Fig.2A). PFA is commonly used in in-situ RT reactions (4), but it is known to have detrimental effects on RNA quality. On the contrary, glyoxal does not crosslink nucleic acids and has been described to yield better RNA quality (5). We aimed to amplify a small number of both gDNA (28) and cDNA (30) targets using induced pluripotent stem cells (iPSCs)in this experiment. The different fixing conditions could be discriminated by using two distinct SBCs for each condition in the in-situ RT reaction (Fig. 1). Analysis was performed using UMI-tools (6).
[0161] An initial ranking of cells per reads yielded 24,468 cells that lie above an automatically determined threshold (Fig.2B). After filtering out low quality cells (less than 10 cDNA / gDNA targets and less than 80% of reads associated with a given SBC) we obtained 9,364 cells that are evenly distributed over both fixation conditions (Fig. 2C). The type of fixation did not have an impact for gDNA target amplification (Fig.2D-F). However, glyoxal fixation yielded higher cDNA targets / cell and UMIs / cell than PFA (Fig. 2G-I) indicating that sensitivity is enhanced using glyoxal as a fixative.
[0162] Expected coverage for gDNA targets would be uniform as each cell has the same input of gDNA. Coverage for cDNA targets is expected to be varying, as not all genes are expressed in all cells and to the same degree and targets were chosen based on a range of expression values to check the sensitivity of the method. 23 out of 28 gDNA targets (82%) showed high coverage and were found in almost all the cells sequenced (Fig. 3A, E and F). 4 gDNA targets (14%) showed only limited coverage, while 1 gDNA target was not covered at all. This contrasts existing tagmentation-based single cell methods where individual gDNA targets are only covered in a fraction of the cells (3). As expected, individual cDNA targets showed varying levels of expression while some were expressed only in a fraction of the cells (Fig. 3B, G and H). Housekeeping genes like GAPDH and TUBB2B were expressed in all cells, in addition to genes like POU5F1, SOX2 and STAT3 that are crucial for iPSC maintenance. Comparison of gene expression levels to bulk RNA-seq data of summed up cDNA UMIs or reads of all cells showed comparable expression levels for the vast majority of targets (Fig. 3C). Unbiased clustering based on the expressed cDNA targets revealed several clusters (Fig. 3D), while clusters were mainly driven by genes that were expressed only in a subpopulation of cells. This is the case for NANOG that is exclusively expressed in cluster 3 (Fig. 3H). This is in line with published scRNA-seq data of iPSCs, where NANOG is also expressed only a fraction of all cells (7).
[0163] Next, we wanted to test if the system is sensitive enough to detect gene expression changes. In this experiment we adapted the primer overhangs of the cDNA targets and gDNA targets to contain two different R2 overhangs (Illumina and Nextera, respectively) to enable separate sequencing of both target libraries (Fig. 4A). This enables sequencing of both cDNAand gDNA libraries with different sequencing modalities regarding length and coverage. For this experiment we designed gRNAs to perturb predicted eQTLs in iPSCs (Fig. 4B) (8). As controls we designed gRNAs using a CRISPRi prediction tool (9, 10) to perturb the transcription start site of all the predicted genes affected by the eQTLs (CRISPRi controls) and non-targeting control (NTC) gRNAs that should not show any effect. We designed primer panels against all the eQTL target sites (28 gDNA targets) and target genes (21 cDNA targets). In addition, we designed cDNA target primers against genes in proximity to the predicted eQTLs, housekeeping genes and the CROP-seq gRNA transcript to correctly assign cells to an expressed gRNA yielding 48 cDNA target primers in total.
[0164] The lentiviral CROP-seq gRNA library was used to infect cells that constitutively express a CRISPRi transgene from the AAVS1 locus in iPSCs (Fig. 4C). Infected cells were selected for by flow cytometry and the scDNA-scRNA-seq method was performed as described above. gRNAs were confidently assigned to cells with a mean coverage of 46 cells / gRNA (Fig. 4D) and tested if the cause any expression changes within the cells. NTC gRNAs did not show any effect on any of the expressed genes, while the 20 out of 21 CRISPRi controls (95 %) showed a significant down-regulation of the intended target gene.7 out of 30 predicted eQTLs (23 %) showed an effect on their predicted target gene. This highlights the power of the method to detect expression changes in a quantitative manner.
[0165] Next, we aimed to directly install eQTLs and read out associated gene expression changes (Fig.5A). For this we screened the 29 eQTLs from the previous CRISPRi perturbation screen while including 5 coding controls that install STOP codons in the beginning of the targeted transcript. We expected to observe non-sense mediated decay in these controls. Same primer panels as in the previous experiment were used for both cDNA and gDNA targets. Transgenic iPSCs that expressed an improved version of the prime editor (PEmax) or PEmax together with MLH1dn (11) were lipofected with a construct expressing pegRNAs that aim to introduce the specific edits (Fig.5B). Lipofected cells were selected for by flow cytometry and scDNA-scRNA-seq was performed. Cells were assigned to be either homozygous- (E+ / E+), heterozygous- (E+ / E-) and non-edited (E- / E-) depending on the fraction of edited reads found for a given target. Overall, we observed limited editing efficiency with both PEmax cell lines tested (Fig. 5C). SOX11 and ATF4 coding controls showed highest editing rates and edited cells were distributed over all clusters found (Fig. 5D-F). ATF4 coding control edits had a significant effect on ATF4 gene expression in an editing dependent manner (Fig. 5G). Other sites could not be confidently tested for as limited editing efficiency confounded theinterpretation of the data. However, we could confidently assign cells based on edited genotypes and show that edited cells can exert an effect on target gene expression. Commercial Applications
[0166] The described method could be of high interest for any application that aims to link genetic information to transcriptomic signatures. Genome wide association studies (GWAS) produced a wealth of information on associations between distinct genetic loci and human diseases. Over 90 percent of predicted variants for common diseases are located in the non- coding genome mainly affecting cis-regulatory elements (CREs) like enhancers (12). Single nucleotide polymorphisms (SNPs) in enhancers can affect gene expression and contribute to disease mechanisms, but for the majority of SNPs their functional contribution to disease is not understood. Current high-throughput methods to study enhancers use CRISPRi to perturb them entirely (13). This neglects the differential influence of individual SNPs at CREs on disease- relevant gene expression. By combining precision genome editing with the targeted scDNA- scRNAseq readout described in this disclosure it will be possible to faithfully link variable genomic editing outcomes with disease-relevant gene expression. This single-cell multiomic readout is crucial to precisely link genotype with gene expression profiles due to differential editing efficiencies at distinct loci.
[0167] Other applications include to characterize patient samples for mutational status and associated gene expression. This could lead to better predictions for treatment of cancer patients, while it will drive mechanistic insights of disease-relevant eQTL mappings. In addition, it enables to perform lineage tracing analysis using endogenous gDNA or mtDNA loci. Due to high mutational rates observed in mtDNA this could already enable lineage dissection of patient samples without any further modifications of the described method. Example 3.
[0168] This example illustrates the methods of developing targeted droplet-based scDNA- scRNA-seq (SDR-seq) that is amenable to screen genetic variation in high-throughput. Abstract
[0169] Genomic variation ranging from single nucleotide polymorphisms to structural variants in both coding and non-coding regions can have an impact on gene function and gene expression thereby contributing to disease mechanism. Current precision editing tools to study these variants systematically have limited efficiency and variable editing outcome inmammalian cells, whereas primary tumor samples have heterogenous mutational profiles that are difficult to assess with current single cell technologies in high-throughput. We developed a droplet-based multiomic targeted scDNA-scRNAseq (SDR-seq) method to precisely link genotype with gene expression profiles in high-throughput. We show that we can assess up to 480 RNA and gDNA targets simultaneously with high coverage and sensitivity in all cells enabling to detect subtle gene expression changes. Using this technology, we can associate variants in both coding and non-coding regions with distinct gene expression profiles in iPSCs and primary B-cell leukemia patient samples. This method has the potential to provide functional insights into the regulatory mechanisms encoded by gDNA variants that govern gene expression at diverse loci. Introduction
[0170] Genomic variation in both coding and non-coding regions of the genome is influencing human population differences and driving diseases in a mono-genetic or complex manner1–3. Genetic loss-of-function screening of coding genes and CRISPRi / CRISPRa screens in non-coding regions have contributed valuable insights into disease mechanisms but lacks taking into account information about precise genomic variation4–6. This potentially masks more complex cellular disease phenotypes that are caused by individual variants in a gene or in non-coding regions7. Existing precision genome editing tools have limited efficiency and variable editing outcome in mammalian cells hindering a systematic study of genetic variation and its impact on disease-relevant gene expression8–10. To confidently link a precise genotype to gene expression a combined single-cell RNA and gDNA is required. Current technologies for simultaneous high-sensitivity readout of both RNA and gDNA are well-based and laborious with limited throughput11–14. High-throughput droplet-based or split-pooling approaches enabling to measure thousands of cells simultaneously are currently possible for a multitude of readouts but are lacking for a combined high-sensitivity and tagmentation-independent readout of RNA and gDNA15–17. Here we developed targeted droplet-based scDNA-scRNA-seq (SDR- seq) that is amenable to screen genetic variation in high-throughput and link them to gene expression and distinct cellular states. Results Droplet-based scDNA-scRNA-seq (SDR-seq)
[0171] We developed SDR-seq to simultaneously measure RNA and gDNA in the same cell with high coverage in all cells. It combines an in-situ reverse transcription (RT) step of fixedcells with a multiplexed PCR in droplets based on the Tapestri technology from MissionBio (Fig.7a). Cells are dissociated into a single-cell suspension, fixed, permeabilized and subjected to an in-situ RT using polydT primers. This step adds a unique molecular identifier (UMI), a sample barcode (sample BC) and a capture sequence (CS) to cDNA molecules. In a first droplet generation cells are lysed, treated with proteinase K and mixed with reverse primer for each intended gDNA or cDNA target. This is followed by a second droplet generation where the first droplet is fused with a cell barcoding bead, PCR reagents and a forward primer for each target with a CS overhang. In a multiplexed PCR both gDNA and cDNA targets are amplified and cell barcoding occurs by complementary CS overhangs of the PCR amplicons and oligos from the cell barcoding bead. Emulsions are broken after the PCR and sequencing ready libraries generated.
[0172] In a first proof-of-principle experiment we compared two different fixation conditions – PFA and glyoxal (Fig. 7B). PFA is commonly used in in-situ RT reactions but can have detrimental effects on gDNA and RNA quality as it crosslinks nucleic acids18. Glyoxal does not crosslink nucleic acids and is therefore expected to yield a more sensitive readout19,20. In this first experiment we aimed to amplify a small number of both gDNA (28) and RNA (30) targets using induced pluripotent stem cells (iPSCs). After filtering for high-quality cells, we obtained around 9,500 cells from a single run that were equally distributed over the two fixation conditions and could be discriminated by using distinct sample BCs for each fixation condition (Figs. 7C, D, Figs. 8A-D). On average over 95 % of reads / cell mapped to the correct sample BC (Figs.7E, F) and contaminating reads were removed per cell. Expected coverage for gDNA targets would be uniform as each cell has the same input of gDNA. Coverage for RNA targets is expected to be varying, as not all genes are expressed in all cells and to the same degree and targets were chosen based on a range of expression values to check the sensitivity of the method.23 out of 28 gDNA targets (82 %) showed high coverage and were found in almost all the cells sequenced (Figs. 7G-I). Little differences were observed in gDNA target detection and coverage using either PFA or glyoxal (Figs.11E, G). As expected, individual RNA targets showed varying levels of expression while some were expressed only in a fraction of the cells (Figs.7J-L). Both RNA targets detected and UMI coverage were increased when using glyoxal compared to the PFA for fixation (Figs. 8F, H). A highly sensitive gene expression readout is crucial due to the potential limited impact of individual variants on gene expression. Ubiquitously expressed housekeeping or iPSC-maintenance genes were detected in all cells, while other genes showed specific expression only in a subset of cells (Fig.9). Comparison ofgene expression levels to bulk RNA-seq data of summed up UMIs of all cells showed comparable expression levels for the vast majority of targets with high correlation (Figs. 7M, N). This shows that SDR-seq is capable of a highly sensitive readout of a multitude of DNA and RNA targets in thousands of single cells in a single experiment with the potential to link those modalities in a high-throughput fashion. SDR-seq is scalable to detect hundreds of targets with high confidence
[0173] Different sequencing modalities in terms of sequencing length and depth would be preferred for either gDNA or RNA targets. For gDNA targets the entire amplicon should be sequenced to ensure that each possible variant is covered. On the contrary, for RNA targets only the barcode structure (cell BC, sample BC and UMI) and the transcript information is required while adjusting the sequencing depth would allow for higher sensitivity to detect potential gene expression changes. To enable this, we modified the overhangs of the reverse primers to contain either R2N (gDNA) or R2 (RNA) overhangs allowing to generate separate NGS libraries (Fig. 11A). We confirmed that the vast majority of reads aligned to the correct references in both the gDNA and RNA libraries (Fig.11B).
[0174] Next, we wanted to test if SDR-seq is scalable to detect a range of hundreds of gDNA or RNA targets simultaneously. We designed an experiment with panel sizes of 120, 240 and 480 targets consisting of equal parts of gDNA and RNA targets in iPSCs (Fig.10A). To allow for cross comparison between panels 60 gDNA and 30 RNA targets were shared between panels. To keep the results comparable between panels, reads were subsampled for gDNA and RNA according to panel size to achieve on average an equal read coverage / cell for the shared targets between panels (Figs.11C-F). A gDNA target was counted as detected per cell if it was covered with more than 5 reads per cell, while an RNA target was counted if it was detected by 1 UMI per cell. Overall, over 80% of all gDNA targets were detected with high confidence in more than 80% of cells across all panels with only a minor decrease in detection for bigger panel sizes (Fig.12A-C). Detection and coverage of shared gDNA targets between panels was well correlated indicating reproducible detection of gDNA targets independent of the panel size (Figs. 10B, C). The minor decrease in detection rate for the bigger panel sizes predominantly affected targets with lower coverage across panels (Figs. 12D, E, Fig. 13A). Checking all RNA targets we observed that there was as only a minor decrease in targets detected in larger panels compared to the 120 panel (Extended Data Fig.5F-H). Detection and gene expression of shared RNA targets were highly correlated between all panels (Fig. 10D,E, Figs. 12I, J) showing a robust and sensitive detection of gene expression readouts. Similar to gDNA targets, variability was predominantly observed for lowly expressed genes (Fig.13B).
[0175] To check whether chromosomal context has an influence on detection we included targets sites that were either overlapping (OEG) or not overlapping expressed genes (NOEG), and were tested for different chromatin marks indicating different regulatory elements depending on their proximity to the TSS (Fig 10E)21. We could not see a strong impact on detection and coverage across panels depending on OEG or NOEG location (Fig. 10F). Importantly there was no regulatory element that showed a systematic bias to be detected, while also sites with low-DNase signal were confidently recovered.
[0176] Genes were chosen based on a range of expression dividing them into high, medium and lowly expressed genes (Fig. 10G). While highly and medium expressed genes were detected in all almost all cells, lowly expressed genes showed a decreased detection rate (Fig. 10H). This is in line with published data, where some genes are not found to be expressed in all cells in iPSCs22. Overall expression levels of shared genes were highly similar across different the different panel sizes tested (Fig.10I).
[0177] We show that SDR-seq is scalable to assay hundreds of gDNA and RNA targets simultaneously with high reproducibility and sensitivity across different panel sizes. This makes it a versatile tool to both assay variants in a multitude of loci in single cells, while linked gene expression changes and cell identity markers can be measured. SDR-seq is sensitive to confidently detect gene expression changes
[0178] Genomic variants might increase or decrease gene expression levels of distinct genes, while effect sizes are expected to be small. Therefore, it is critical to have a highly sensitive gene expression readout that can confidently detect changes. To test if SDR-seq is sensitive to measure gene expression changes, we designed a CRISPRi experiment consisting of non- targeting control gRNAs (NTC), gRNAs targeting high-confidence eQTLs, gRNAs targeting the transcription start site (TSS) of genes predicted to be affected by those eQTLs (CRISPRi controls) and gRNAs that target the gene body of a transcript where a possible STOP codon could be introduced via prime editing (STOP controls) (Fig. 14A). The primer panel was reading out the gDNA sites of the eQTLs together with associated transcripts, the CROP-seq transcript and multiple housekeeping genes to enable better normalization of the gene expression data. iPSCs expressing the CRISPRi transgene from the safe-harbor locus AAVS1 were infected with the lentiviral CROP-seq gRNA library, selected via FACS and SDR-seqwas performed (Fig. 14B). Cells were assigned to gRNAs (75%) with an average coverage of 30 cells / gRNA (Figs. 15A, B). NTC gRNAs did not show a significant effect on any of the genes that were measured, while the vast majority of CRISPRi control gRNAs targeting the TSS of a gene showed a strong effect on the expression levels of the associated target gene (Fig. 14C). 7 eQTL gRNAs and 3 STOP control gRNAs showed a significant effect on target gene expression. Significantly scoring eQTL and STOP control gRNAs were located within a two kb window of the TSS indicating a direct inhibitory effect of CRISPRi comparable to the CRISPRi control gRNAs (Fig. 15C). This indicates the importance of directly assessing the variants to check their impact on gene expression when approximating variant effects that are close to the TSS with CRISPRi experiments.
[0179] Next, we aimed to directly install the eQTL variant and measure its effect sizes on gene expression. For this we generated two iPSCs cell lines that express a transgene to enable prime editing (PE), with or without the simultaneous expression of a dominant negative regulator of the mismatch repair pathway that was shown to increase editing efficiency (PEmax or PEmax-MLH1dn)10. We tested these PE iPSCs with a fluorescent lentiviral reporter system that enables to measure editing efficiencies via the reconstitution of a non-functional EGFP (Fig. 15D). We observed around 50 % editing efficiency after lipofection of a prime editing gRNA (pegRNA) that repairs the EGFP fluorescent protein (Fig. 15E) indicating the system has the potential to enable editing in iPSCs. Next, we lipofected these validated PE iPSCs with a pegRNA library that aimed to introduce the same eQTLs as in the CRISPRi screen or STOP codons to assay nonsense mediated decay. We enriched cells that were lipofected with flow cytometry and performed SDR-seq (Fig.14E). Overall, we observed limited editing efficiency with both PE cell lines for the variants we aimed to introduce (Figs. 16A, B). We could only observe significant gene expression changes for the STOP control (Fig. 14F), but limited editing efficiency for many of the eQTLs confounds the interpretation of those variants. Genotypes could be confidently called discriminating between reference (REF), heterozygous (HET) and alternative (ALT) alleles (Figs. 16C-E). Depending on the position of the STOP codon within the transcript, effects of nonsense mediated decay on transcript levels can vary23. For SOX11 we observed no changes on transcript levels, while STOP codons introduced in ATF4 and MYH10 had a significant impact on gene expression (Fig.14G).
[0180] As the PE iPSCs exhibited only limited editing efficiency we aimed to install eQTLs via base editing in a follow up experiment. We selected 56 high confidence eQTLs to be installed with either ABE8e or CBE base editors (Fig.14H). After introducing gRNA librariesinto iPSCs, cells were selected and SDR-seq performed (Fig.14I). In this experiment we found several eQTL variants having a significant effect on target gene expression (Fig. 14J). In particular a synonymous variant in POU5F1 showed a significant effect on gene expression (Fig. 17A). However, after assessing variants along the entire amplicon of POU5F1 we found a diverse set of variants that had an impact on gene expression (Fig. 14K, Fig. 17B). In particular variants in the 3’ UTR were associated with differential transcript levels. This highlights the importance of directly assessing variants at the locus of interest to get a better resolution on the impact on gene expression.
[0181] We can confidently call variants on a single cell level and associate them gene expression and are sensitive to detect even subtle changes. Overall editing efficiency was limited in our experiments confounding the interpretation of many tested eQTLs. Discussion
[0182] Here were developed SDR-seq to directly measure gDNA and RNA in single cells in high-throughput and sensitivity. The method is scalable to assess hundreds of gDNA targets and genes in each single cell, allowing for accurate variant calling, gene expression measurement and cell type identification. Materials and Methods Cell Culture
[0183] WTC-11 iPSCs (Coriell Institute for Medical Research – GM25256) were verified to display a normal karyotype, contamination free and regularly tested for mycoplasma. They were cultured in Essential 8™ Medium (E8) (Thermo Fisher Scientific, Cat. #A1517001) on Vitronectin XF™ (Stem Cell Technologies, Cat. #07180) coated tissue culture plates. Cells were maintained at 37° C at 5% CO2. iPSCs were split using Accutase (StemCell Technologies - #07922) and E8 supplemented with 10 μM Y-27632 dihydrochloride (RI) (Tocris, Cat. No 1254). After dissociation into single cells, cells 1 volume of E8+RI was added, cells spun at 200 g for 5 min and resuspended and plated in E8+RI. Cloning, molecular biology and generating transgenic iPSCs
[0184] For the constitutive CRISPRi, PEmax and PEmax-MLHdn1 cell lines the corresponding transgene was inserted into the AAVS1 locus in WTC-11 iPSCs as previously described using specific TALENs24. The AAVS1 targeting vector containing the homologyarms and the CAGG promotor was a kind gift from J. A. Knoblich (Institute of Molecular Biotechnology of the Austrian Academy of Science -IMBA, Vienna BioCenter, Vienna, Austria). SDR-seq
[0185] WTC-11 iPSCs were dissociated into single cells using Accutase, filtered through a 40 μm cell strainer and counted. 1.5 x10^6 cells were transferred to a new 15 ml conical tube and spun at 500 g for 3 min.
[0186] For the glyoxal fixation condition the supernatant was removed, cells were resuspended in 200 μl of glyoxal solution fixation solution (3% glyoxal, 20% EtOH, 0.75% Acetic Acid – Glacial, pH = 4.0) and incubated for 7 min at room temperature. 1 ml ice-cold wash buffer 1 (1x PBS with 2% BSA, 1 mM DTT and 0.5 U / μl RNasin® Plus Ribonuclease Inhibitor - Promega #N2615) was added and cells were spun at 500 g for 3 min at 4 C. Supernatant was carefully removed and wash step was repeated with wash buffer 1 for a total of 2 washes. Cells were resuspended in 175 μl ice-cold permeabilization buffer (10 mM TRIS- HCl pH 7.5, 10 mM NaCl, 3 mM MgCl2, 0.1% Tween-20, 0.2 U / μl RNasin® Plus Ribonuclease Inhibitor, 1mM DTT, 2 % BSA, 0.1 % IGEPAL CA-630 and 0.01 % Digitonin) and incubated for 4 min on ice. 1 ml of ice-cold wash buffer 2 (10 mM TRIS pH 7.5, 10 mM NaCl, 3 mM MgCl2, 0.1% Tween-20, 0.2 U / μl RNasin® Plus Ribonuclease Inhibitor, 1mM DTT and 2 % BSA) was added and tube was gently inverted 4-6 times. Cells were spun at 500 g for 5 min at 4 °C and resuspended in ice-cold resuspension buffer (1x PBS, 2% BSA, 1 mM DTT and 0.2 U / μl RNasin® Plus Ribonuclease Inhibitor), filtered through a 40 μm strainer, counted and diluted to 1.4 x10^6 cells / ml.
[0187] PFA fixation conditions were followed as described elsewhere with adaptions (Rosenberg et al., 2018). In short, the supernatant was removed, cells were resuspended in 1 ml 1x PBS with 0.2 U / μl RNasin® Plus Ribonuclease Inhibitor and 3 ml of 1.3% PFA solution (in 1x PBS) were added. Cells were fixed for 10 min on ice.160 μl of permeabilization buffer (5% Trition-X 100 with 0.2 U / μl RNasin® Plus Ribonuclease Inhibitor) was added, the tube gently inverted for 4-6 times and cells were incubated for 3 min on ice. Cells were spun at 500 g for 3 min at 4 C, supernatant was carefully removed and cells were resuspended in 500 μl of 1x PBS with 0.2 U / μl RNasin® Plus Ribonuclease Inhibitor.500 μl of ice-cold 100 mM TRIS- HCl at pH 8.0 was added and mixed by inverting the tube.20 μl of permeabilization buffer was added and mixed by inverting the tube 4-6 times. Cells were spun at 500 g for 3 min at 4 °C,supernatant removed, resuspended in 300 μl of 0.5x PBS with 0.2 U / μl RNasin® Plus Ribonuclease Inhibitor, filtered through a 40 μm strainer, counted and diluted to 1.4 x10^6 cells / ml.
[0188] RT master mix consisting of a final concentration of 1x RT Buffer, 0.25 U / μl Enzymatics RNAse Inhibitor (Biozym – 180520), 0.2 U / μl RNasin® Plus Ribonuclease Inhibitor, 500 mM dNTPs and 20 U / μl Maxima H Minus Reverse Transcriptase (ThermoFisher - EP0752) was prepared on ice for in 8 μl for a total reaction volume of 20 μl. 4μl of reverse transcription oligo (12.5 μM) were combined in each 96-well plate with 8 μl reverse transcription master mix. Sequences for RT 8 μl fixed and permeabilized cells (10000 cells total) were added to each well, yielding a total reaction volume of 20 μl. Reverse transcription was performed in a thermocycler using the protocol shown in Table 12. All RT reactions were pooled into a 15 ml conical tube, 1 volume of ice-cold 1x PBS with 1% BSA was added and cells were spun at 500 g for 5 min and supernatant was removed. Table 12. Reverse transcription (RT) protocol
[0189] Samples were processed using the Tapestri microfluidic device from Mission Bio (MB51-0007, MB51-0010, MB51-0009) according to manufactures protocol withmodifications. In-situ RT processed cell pellet from previous step was resuspended in the cell buffer of Mission Bio, cells were counted and diluted to the appropriate concentration of 3000- 4000 cells / μl. Custom primers were used in the multiplexed droplet PCR amplification step. cDNA primers were designed using the TAP-seq primer prediction tool with at targeted Tm of 60°C4. gDNA primers were designed using the Tapestri Designer (https: / / designer.missionbio.com). Version 1 gDNA and cDNA primers both had CS and R2N overhangs (only used in proof-of-concept experiment). Version 2 gDNA primers had CS and R2N, whereas cDNA primers had CS and R2 overhangs. Forward and reverse stock primers had a concentration of 16 μM and 90 μM for both versions, respectively. For Version 1 final sequencing libraries were generated according to the Mission Bio user guide. For Version 2 cDNA and gDNA sequencing libraries were generated separately using the corresponding library amplification primers (gDNA: R1N – R2N, cDNA: R1N – R2). References for Example 3 1. Landrum, M. J. et al. ClinVar: public archive of relationships among sequence variation and human phenotype. Nucleic Acids Res.42, D980–D985 (2014). 2. Buniello, A. et al. The NHGRI-EBI GWAS Catalog of published genome-wide association studies, targeted arrays and summary statistics 2019. Nucleic Acids Res. 47, D1005–D1012 (2019). 3. Eichler, E. E. Genetic Variation, Comparative Genomics, and the Diagnosis of Disease. N. Engl. J. Med.381, 64–74 (2019). 4. Schraivogel, D. et al. Targeted Perturb-seq enables genome-scale genetic screens in single cells. Nat. Methods 17, 629–635 (2020). 5. Esk, C. et al. A human tissue screen identifies a regulator of ER secretion as a brain- size determinant. Science 370, 935–941 (2020). 6. Morris, J. A. et al. Discovery of target genes and pathways at GWAS loci by pooled single-cell CRISPR screens. Science 380, eadh7699 (2023). 7. Kornienko, J. et al. Mislocalization of pathogenic RBM20 variants in dilated cardiomyopathy is caused by loss-of-interaction with Transportin-3. Nat. Commun. 14, 1–20 (2023).8. Anzalone, A. V. et al. Search-and-replace genome editing without double-strand breaks or donor DNA. Nature 576, 149–157 (2019). 9. Porto, E. M., Komor, A. C., Slaymaker, I. M. & Yeo, G. W. Base editing: advances and therapeutic opportunities. Nat. Rev. Drug Discov.19, 839–859 (2020). 10. Chen, P. J. et al. Enhanced prime editing systems by manipulating cellular determinants of editing outcomes. Cell 184, 5635-5652.e29 (2021). 11. Yu, L. et al. scONE-seq: A single-cell multi-omics method enables simultaneous dissection of phenotype and genotype heterogeneity from frozen tumors. Sci. Adv.9, eabp8901 (2023). 12. Rodriguez-Meira, A. et al. Unravelling Intratumoral Heterogeneity through High- Sensitivity Single-Cell Mutational Analysis and Parallel RNA Sequencing. Mol. Cell 73, 1292- 1305.e8 (2019). 13. Kong, S. L. et al. Concurrent Single-Cell RNA and Targeted DNA Sequencing on an Automated Platform for Comeasurement of Genomic and Transcriptomic Signatures. Clin. Chem.65, 272–281 (2019). 14. Zachariadis, V., Cheng, H., Andrews, N. & Enge, M. A Highly Scalable Method for Joint Whole-Genome Sequencing and Gene-Expression Profiling of Single Cells. Mol. Cell 80, 541-553.e5 (2020). 15. Lee, J., Hyeon, D. Y. & Hwang, D. Single-cell multiomics: technologies and data analysis methods. Exp. Mol. Med.52, 1428–1442 (2020). 16. Yin, Y. et al. High-Throughput Single-Cell Sequencing with Linear Amplification. Mol. Cell 76, 676-690.e10 (2019). 17. Olsen, T. R. et al. Scalable co-sequencing of RNA and DNA from individual nuclei. 2023.02.09.527940 Preprint at https: / / doi.org / 10.1101 / 2023.02.09.527940 (2023). 18. Phan, H. V. et al. High-throughput RNA sequencing of paraformaldehyde-fixed single cells - Nature Communications. Nat. Commun.12, 5636 (2021). 19. Richter, K. N. et al. Glyoxal as an alternative fixative to formaldehyde in immunostaining and super^resolution microscopy. EMBO J.37, 139–159 (2018).20. Channathodiyil, P. & Houseley, J. Glyoxal fixation facilitates transcriptome analysis after antigen staining and cell sorting by flow cytometry. PLOS ONE 16, e0240769 (2021). 21. Expanded encyclopaedias of DNA elements in the human and mouse genomes. Nature https: / / www.nature.com / articles / s41586-020-2493-4 (2020). 22. Nguyen, Q. H. et al. Single-cell RNA-seq of human induced pluripotent stem cells reveals cellular heterogeneity and cell state transitions between subpopulations. Genome Res. 28, 1053–1066 (2018). 23. Lindeboom, R. G. H., Supek, F. & Lehner, B. The rules and impact of nonsense- mediated mRNA decay in human cancers. Nat. Genet.48, 1112–1118 (2016). 24. Hockemeyer, D. et al. Genetic engineering of human pluripotent cells using TALE nucleases. Nat. Biotechnol.29, 731–734 (2011). REFERENCES 1. L. Chappell, A. J. C. Russell, T. Voet, Single-Cell (Multi)omics Technologies. Annu. Rev. Genomics Hum. Genet.19, 15–41 (2018). 2. L. Yu, X. Wang, Q. Mu, S. S. T. Tam, D. S. C. Loi, A. K. Y. Chan, W. S. Poon, H.-K. Ng, D. T. M. Chan, J. Wang, A. R. Wu, scONE-seq: A single-cell multi-omics method enables simultaneous dissection of phenotype and genotype heterogeneity from frozen tumors. Sci. Adv.9, eabp8901 (2023). 3. T. R. Olsen, P. Talla, J. Furnari, J. N. Bruce, P. Canoll, S. Zha, P. A. Sims, Scalable co- sequencing of RNA and DNA from individual nuclei (2023), p. 2023.02.09.527940, doi:10.1101 / 2023.02.09.527940. 4. A. B. Rosenberg, C. M. Roco, R. A. Muscat, A. Kuchina, P. Sample, Z. Yao, L. T. Graybuck, D. J. Peeler, S. Mukherjee, W. Chen, S. H. Pun, D. L. Sellers, B. Tasic, G. Seelig, Single-cell profiling of the developing mouse brain and spinal cord with split-pool barcoding. Science. 360, 176–182 (2018). 5. P. Channathodiyil, J. Houseley, Glyoxal fixation facilitates transcriptome analysis after antigen staining and cell sorting by flow cytometry. PLOS ONE.16, e0240769 (2021). 6. T. S. Smith, A. Heger, I. Sudbery, Genome Res., in press, doi:10.1101 / gr.209601.116.7. Q. H. Nguyen, S. W. Lukowski, H. S. Chiu, A. Senabouth, T. J. C. Bruxner, A. N. Christ, N. J. Palpant, J. E. Powell, Single-cell RNA-seq of human induced pluripotent stem cells reveals cellular heterogeneity and cell state transitions between subpopulations. Genome Res. 28, 1053–1066 (2018). 8. C. DeBoever, H. Li, D. Jakubosky, P. Benaglio, J. Reyna, K. M. Olson, H. Huang, W. Biggs, E. Sandoval, M. D’Antonio, K. Jepsen, H. Matsui, A. Arias, B. Ren, N. Nariai, E. N. Smith, A. D’Antonio-Chronowska, E. K. Farley, K. A. Frazer, Large-Scale Profiling Reveals the Influence of Genetic Variation on Gene Expression in Human Induced Pluripotent Stem Cells. Cell Stem Cell.20, 533-546.e7 (2017). 9. K. R. Sanson, R. E. Hanna, M. Hegde, K. F. Donovan, C. Strand, M. E. Sullender, E. W. Vaimberg, A. Goodale, D. E. Root, F. Piccioni, J. G. Doench, Optimized libraries for CRISPR- Cas9 genetic screens with multiple modalities. Nat. Commun.9, 1–15 (2018). 10. J. G. Doench, N. Fusi, M. Sullender, M. Hegde, E. W. Vaimberg, K. F. Donovan, I. Smith, Z. Tothova, C. Wilen, R. Orchard, H. W. Virgin, J. Listgarten, D. E. Root, Optimized sgRNA design to maximize activity and minimize off-target effects of CRISPR-Cas9. Nat. Biotechnol. 34, 184–191 (2016). 11. P. J. Chen, J. A. Hussmann, J. Yan, F. Knipping, P. Ravisankar, P.-F. Chen, C. Chen, J. W. Nelson, G. A. Newby, M. Sahin, M. J. Osborn, J. S. Weissman, B. Adamson, D. R. Liu, Enhanced prime editing systems by manipulating cellular determinants of editing outcomes. Cell.184, 5635-5652.e29 (2021). 12. A. Buniello, J. A. L. MacArthur, M. Cerezo, L. W. Harris, J. Hayhurst, C. Malangone, A. McMahon, J. Morales, E. Mountjoy, E. Sollis, D. Suveges, O. Vrousgou, P. L. Whetzel, R. Amode, J. A. Guillen, H. S. Riat, S. J. Trevanion, P. Hall, H. Junkins, P. Flicek, T. Burdett, L. A. Hindorff, F. Cunningham, H. Parkinson, The NHGRI-EBI GWAS Catalog of published genome-wide association studies, targeted arrays and summary statistics 2019. Nucleic Acids Res.47, D1005–D1012 (2019). 13. C. P. Fulco, M. Munschauer, R. Anyoha, G. Munson, S. R. Grossman, E. M. Perez, M. Kane, B. Cleary, E. S. Lander, J. M. Engreitz, Systematic mapping of functional enhancer- promoter connections with CRISPR interference. Science.354, 769–773 (2016). 14. E. Z. Macosko, A. Basu, R. Satija, J. Nemesh, K. Shekhar, M. Goldman, I. Tirosh, A. R. Bialas, N. Kamitaki, E. M. Martersteck, J. J. Trombetta, D. A. Weitz, J. R. Sanes, A. K. Shalek,A. Regev, S. A. McCarroll, Highly Parallel Genomewide Expression Profiling of Individual Cells Using Nanoliter Droplets. Cell.161, 1202–1214 (2015). 15. M. Stoeckius, C. Hafemeister, W. Stephenson, B. Houck-Loomis, P. K. Chattopadhyay, H. Swerdlow, R. Satija, P. Smibert, Simultaneous epitope and transcriptome measurement in single cells. Nat. Methods.14, 865–868 (2017). 16. B. Demaree, C. L. Delley, H. N. Vasudevan, C. A. C. Peretz, D. Ruff, C. C. Smith, A. R. Abate, Joint profiling of DNA and proteins in single cells to dissect genotype-phenotype associations in leukemia. Nat. Commun.12, 1583 (2021). 17. A. S. Nam, N. Dusaj, F. Izzo, R. Murali, R. M. Myers, T. H. Mouhieddine, J. Sotelo, S. Benbarche, M. Waarts, F. Gaiti, S. Tahri, R. Levine, O. Abdel-Wahab, L. A. Godley, R. Chaligne, I. Ghobrial, D. A. Landau, Single-cell multi-omics of human clonal hematopoiesis reveals that DNMT3A R882 mutations perturb early progenitor states through selective hypomethylation. Nat. Genet., 1–13 (2022). 18. S. L. Kong, H. Li, J. A. Tai, E. T. Courtois, H. M. Poh, D. P. Lau, Y. X. Haw, N. G. Iyer, D. S. W. Tan, S. Prabhakar, D. Ruff, A. M. Hillmer, Concurrent Single-Cell RNA and Targeted DNA Sequencing on an Automated Platform for Comeasurement of Genomic and Transcriptomic Signatures. Clin. Chem.65, 272–281 (2019). EXEMPLARY EMBODIMENTS
[0190] Exemplary embodiments provided in accordance with the presently disclosed subject matter include, but are not limited to, the claims and the following embodiments:
[0191] Embodiment 1. A method for simultaneously detecting RNA and genomic DNA (gDNA) within the same cell, comprising: (i) providing a single cell suspension comprising fixed and permeabilized cells; (ii) performing an in-situ reverse transcription (RT) step to generate cDNA molecules by contacting the cells with reverse transcriptase and a RT primer; (iii) lysing the cells in a first droplet comprising a first reverse primer that hybridizes to the cDNA molecules and a second reverse primer that hybridizes to the gDNA molecules, wherein the first or second reverse primer comprises a R2N or R2 overhang sequence;(iv) fusing the first droplet to a second droplet comprising PCR reagents and a forward primer that hybridizes to both the cDNA and gDNA molecules and comprises a capture sequence (CS) overhang sequence; and (v) simultaneously amplifying the cDNA and gDNA molecules, thereby detecting both RNA and gDNA within the same cell.
[0192] Embodiment 2. The method of embodiment 1, wherein the single cell suspension is fixed with paraformaldehyde (PFA) or glyoxal.
[0193] Embodiment 3. The method of embodiment 1 or 2, wherein the R2N overhang sequence comprises the nucleic acid sequence GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG (SEQ ID NO:2), and the R2 overhang sequence comprises the nucleic acid sequence GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT (SEQ ID NO:3).
[0194] Embodiment 4. The method of any one of embodiments 1 to 3, wherein the RT primer comprises a capture sequence (CS), a sample barcode (SBC) sequence, a unique molecular identifier (UMI) sequence, and a sequence that binds to mRNA.
[0195] Embodiment 5. The method of embodiment 4, wherein the CS sequence hybridizes to a complementary sequence attached to a solid support, the SBC sequence comprises a known sequence of variable length, and the UMI sequence comprises a random sequence of variable length.
[0196] Embodiment 6. The method of any one of embodiments 1 to 5, wherein the CS comprises the nucleic acid sequence GTACTCGCAGTAGTC (SEQ ID NO:1).
[0197] Embodiment 7. The method of any one of embodiments 4 to 6, wherein the SBC sequence is from 4 to 50 nucleotides in length.
[0198] Embodiment 8. The method of embodiment 7, wherein the SBC sequence comprises a nucleic acid sequence selected from the group consisting of TCGCCTTA (SEQ ID NO:4), CTAGTACG (SEQ ID NO:5), TTCTGCCT (SEQ ID NO:6), GCTCAGGA (SEQ ID NO:7), AGGAGTCC (SEQ ID NO:8), CATGCCTA (SEQ ID NO:9), GTAGAGAG (SEQ ID NO:10), and CCTCTCTG (SEQ ID NO:11).
[0199] Embodiment 9. The method of any one of embodiments 4 to 8, wherein the UMI sequence is from 4 to 50 nucleotides in length.
[0200] Embodiment 10. The method of embodiment 9, wherein the UMI sequence comprises the nucleic acid sequence NNNNNNNN (SEQ ID NO:12).
[0201] Embodiment 11. The method of embodiment 4, wherein the sequence that binds to mRNA comprises oligo(dT) or a sequence that hybridizes to any region of the mRNA.
[0202] Embodiment 12. The method of any one of embodiments 1 to 11, wherein step (v) comprises PCR conditions that favor binding of forward and reverse primers to the cDNA and gDNA molecules thereby producing a first set of PCR products each comprising the CS.
[0203] Embodiment 13. The method of embodiment 12, wherein the second droplet further comprises cell barcode oligonucleotides each comprising a cell barcode (CBC) and a sequence complementary to the CS, thereby the first set of PCR products hybridize to the cell barcode oligonucleotides, producing a second set of PCR products each comprising the CS and the CBC.
[0204] Embodiment 14. The method of embodiment 13, wherein the cell barcode oligonucleotides are attached to a solid support in the second droplet prior step (iv) and released from the solid support after the second droplet is fused with the first droplet.
[0205] Embodiment 15. The method of any one of embodiments 1 to 14, wherein simultaneous amplification in step (v) produces a cDNA library and a gDNA library.
[0206] Embodiment 16. The method of embodiment 15, further comprising sequencing the cDNA and gDNA libraries.
[0207] Embodiment 17. The method of embodiment 16, wherein the cDNA library is sequenced by a first sequencing modality, and the gDNA library is sequenced by a second sequencing modality.
[0208] Embodiment 18. The method of embodiment 17, wherein the first and second sequencing modalities are the same or different.
[0209] Embodiment 19. The method of embodiment 17 or 18, wherein the first sequencing modality comprises scRNA-seq and the second sequencing modality comprises scDNA-seq.
[0210] Embodiment 20. The method of embodiment 17 or 18, whereinthe first sequencing modality comprises Illumina® next generation sequencing (NGS), and the second sequencing modality comprises Nextera® NGS, or the first sequencing modality comprises Nextera® NGS, and the second sequencing modality comprises Illumina® NGS.
[0211] Embodiment 21. The method of any one of embodiments 1 to 20, wherein step (iii) comprises contacting the cells with proteinase K to lyse the cells.
[0212] Embodiment 22. The method of any one of embodiments 1 to 21, wherein the cell is a prokaryotic or eukaryotic cell.
[0213] Embodiment 23. The method of any one of embodiments 1 to 22, wherein the cell is an induced pluripotent stem cell (iPSC).
[0214] Embodiment 24. The method of any one of embodiments 1 to 23, wherein the cell is a genetically modified cell.
[0215] Embodiment 25. The method of embodiment 24, wherein the cell is modified using a CRISPR / Cas gene editing system.
[0216] Embodiment 26. The method of embodiment 25, wherein the CRISPR / Cas gene editing system is CRISPR interference (CRISPRi).
[0217] Embodiment 27. The method of any one of embodiments 24 to 26, wherein RNA expressed by the genetically modified cell is sequenced by scRNA-seq and genomic DNA is sequenced by scDNA-seq.
[0218] Embodiment 28. A reaction mixture comprising RNA and gDNA from a single cell, a reverse transcriptase, and a RT primer comprising a capture sequence (CS), a sample barcode (SBC) sequence, a unique molecular identifier (UMI) sequence, and a sequence that binds to mRNA.
[0219] Embodiment 29. The reaction mixture of embodiment 28, wherein the CS hybridizes to a complementary sequence attached to a solid support, SBC sequence comprises a known sequence of variable length, and the UMI sequence comprises a random sequence of variable length, and the sequence that binds to mRNA comprises oligo(dT) or a sequence that hybridizes to any region of the mRNA.
[0220] Embodiment 30. The reaction mixture of embodiment 28 or 29, wherein the CS comprises the nucleic acid sequence GTACTCGCAGTAGTC (SEQ ID NO:1).
[0221] Embodiment 31. The reaction mixture of any one of embodiments 28 to 30, wherein the SBC sequence is from 4 to 50 nucleotides in length.
[0222] Embodiment 32. The reaction mixture of any one of embodiments 28 to 31, wherein the SBC sequence comprises a nucleic acid sequence selected from the group consisting of TCGCCTTA (SEQ ID NO:4), CTAGTACG (SEQ ID NO:5), TTCTGCCT (SEQ ID NO:6), GCTCAGGA (SEQ ID NO:7), AGGAGTCC (SEQ ID NO:8), CATGCCTA (SEQ ID NO:9), GTAGAGAG (SEQ ID NO:10), and CCTCTCTG (SEQ ID NO:11).
[0223] Embodiment 33. The reaction mixture of any one of embodiments 28 to 32, wherein the UMI sequence is from 4 to 50 nucleotides in length.
[0224] Embodiment 34. The reaction mixture of any one of embodiments 28 to 33, wherein the UMI sequence comprises the nucleic acid sequence NNNNNNNN (SEQ ID NO:12).
[0225] Embodiment 35. The reaction mixture of any one of embodiments 28 to 34, further comprising cDNA molecules produced by reverse transcription of the RNA and a reverse primer that hybridizes to both the cDNA and the gDNA molecules and comprises a R2N or R2 overhang sequence.
[0226] Embodiment 36. The reaction mixture of embodiment 35, wherein the reverse primer comprising the R2N overhang sequence comprises the nucleic acid sequence GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG (SEQ ID NO:2) and the reverse primer comprising the R2 overhang sequence comprises the nucleic acid sequence GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT (SEQ ID NO:3).
[0227] Embodiment 37. The reaction mixture of any one of embodiments 28 to 36, comprising a, one or more, or a plurality of polynucleotides selected from a sequence in any one of SEQ ID Nos: 1-83.
[0228] Embodiment 38. A polynucleotide comprising a capture sequence (CS).
[0229] Embodiment 39. The polynucleotide of embodiment 38, further comprising a sample barcode (SBC) sequence, a unique molecular identifier (UMI) sequence, and / or a sequence that binds to mRNA.
[0230] Embodiment 40. The polynucleotide of embodiment 39, wherein the SBC sequence comprises a known sequence of variable length, and the UMI sequence comprises a random sequence of variable length.
[0231] Embodiment 41. The polynucleotide of any one of embodiments 38 to 40, wherein the CS comprises the nucleic acid sequence GTACTCGCAGTAGTC (SEQ ID NO:1).
[0232] Embodiment 42. The polynucleotide of any one of embodiments 39 to 41, wherein the SBC sequence is from 4 to 50 nucleotides in length.
[0233] Embodiment 43. The polynucleotide of any one of embodiments 39 to 42, wherein the SBC sequence comprises a nucleic acid sequence selected from the group consisting of TCGCCTTA (SEQ ID NO:4), CTAGTACG (SEQ ID NO:5), TTCTGCCT (SEQ ID NO:6), GCTCAGGA (SEQ ID NO:7), AGGAGTCC (SEQ ID NO:8), CATGCCTA (SEQ ID NO:9), GTAGAGAG (SEQ ID NO:10), and CCTCTCTG (SEQ ID NO:11).
[0234] Embodiment 44. The polynucleotide of any one of embodiments 39 to 43, wherein the UMI sequence is from 4 to 50 nucleotides in length.
[0235] Embodiment 45. The polynucleotide of any one of embodiments 39 to 44, wherein the UMI comprises the nucleic acid sequence NNNNNNNN (SEQ ID NO:12).
[0236] Embodiment 46. The polynucleotide of any one of embodiments 39 to 45, wherein the sequence that binds to mRNA comprises oligo(dT) or a sequence that hybridizes to any region of the mRNA.
[0237] Embodiment 47. The polynucleotide of any one of embodiments 38 to 46, wherein the polynucleotide comprises a sequence in any one of SEQ ID Nos: 1-83.
[0238] All publications and patent applications mentioned in this disclosure are herein incorporated by reference to the same extent as if each individual publication or patent application was specifically and individually indicated to be incorporated by reference.
[0239] No admission is made that any reference cited herein constitutes prior art. The discussion of the references states what their authors assert, and the Applicant reserves the right to challenge the accuracy and pertinence of the cited documents. It will be clearly understood that, although a number of information sources, including scientific journal articles, patent documents, and textbooks, may be referred to herein; this reference does not constitute an admission that any of these documents forms part of the common general knowledge in the art.
[0240] The discussion of the general methods given herein is intended for illustrative purposes only. Other alternative methods and alternatives will be apparent to those of skill in the art upon review of this disclosure and are to be included within the scope of this application.
[0241] While particular alternatives of the present disclosure have been disclosed, it is to be understood that various modifications and combinations are possible and are contemplated within the scope of the appended claims. There is no intention, therefore, of limitations to the exact abstract and disclosure herein presented.
Claims
WHAT IS CLAIMED IS:
1. A method for simultaneously detecting RNA and genomic DNA (gDNA) within the same cell, comprising: (i) providing a single cell suspension comprising fixed and permeabilized cells; (ii) performing an in-situ reverse transcription (RT) step to generate cDNA molecules by contacting the cells with reverse transcriptase and a RT primer; (iii) lysing the cells in a first droplet comprising a first reverse primer that hybridizes to the cDNA molecules and a second reverse primer that hybridizes to the gDNA molecules, wherein the first or second reverse primer comprises a R2N or R2 overhang sequence; (iv) fusing the first droplet to a second droplet comprising PCR reagents and a forward primer that hybridizes to both the cDNA and gDNA molecules and comprises a capture sequence (CS) overhang sequence; and (v) simultaneously amplifying the cDNA and gDNA molecules, thereby detecting both RNA and gDNA within the same cell.
2. The method of claim 1, wherein the single cell suspension is fixed with paraformaldehyde (PFA) or glyoxal.
3. The method of claim 1 or 2, wherein the R2N overhang sequence comprises the nucleic acid sequence GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG (SEQ ID NO:2), and the R2 overhang sequence comprises the nucleic acid sequence GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT (SEQ ID NO:3).
4. The method of claim 1, wherein the RT primer comprises a capture sequence (CS), a sample barcode (SBC) sequence, a unique molecular identifier (UMI) sequence, and a sequence that binds to mRNA.The method of claim 4, wherein the CS sequence hybridizes to a complementary sequence attached to a solid support, the SBC sequence comprises a known sequence of variable length, and the UMI sequence comprises a random sequence of variable length.
6. The method of claim 1, wherein the CS comprises the nucleic acid sequence GTACTCGCAGTAGTC (SEQ ID NO:1).
7. The method of claim 4, wherein the SBC sequence is from 4 to 50 nucleotides in length.
8. The method of claim 7, wherein the SBC sequence comprises a nucleic acid sequence selected from the group consisting of TCGCCTTA (SEQ ID NO:4), CTAGTACG (SEQ ID NO:5), TTCTGCCT (SEQ ID NO:6), GCTCAGGA (SEQ ID NO:7), AGGAGTCC (SEQ ID NO:8), CATGCCTA (SEQ ID NO:9), GTAGAGAG (SEQ ID NO:10), and CCTCTCTG (SEQ ID NO:11).
9. The method of claim 4, wherein the UMI sequence is from 4 to 50 nucleotides in length.
10. The method of claim 9, wherein the UMI sequence comprises the nucleic acid sequence NNNNNNNN (SEQ ID NO:12).
11. The method of claim 4, wherein the sequence that binds to mRNA comprises oligo(dT) or a sequence that hybridizes to any region of the mRNA.
12. The method of claim 1, wherein step (v) comprises PCR conditions that favor binding of forward and reverse primers to the cDNA and gDNA molecules thereby producing a first set of PCR products each comprising the CS.
13. The method of claim 12, wherein the second droplet further comprises cell barcode oligonucleotides each comprising a cell barcode (CBC) and a sequence complementary to the CS, thereby the first set of PCR products hybridize to the cell barcode oligonucleotides, producing a second set of PCR products each comprising the CS and the CBC.
14. The method of claim 13, wherein the cell barcode oligonucleotides are attached to a solid support in the second droplet prior step (iv) and released from the solid support after the second droplet is fused with the first droplet.
15. The method of claim 1, wherein simultaneous amplification in step (v) produces a cDNA library and a gDNA library.
16. The method of claim 15, further comprising sequencing the cDNA and gDNA libraries.
17. The method of claim 16, wherein the cDNA library is sequenced by a first sequencing modality, and the gDNA library is sequenced by a second sequencing modality.
18. The method of claim 17, wherein the first and second sequencing modalities are the same or different.
19. The method of claim 17 or 18, wherein the first sequencing modality comprises scRNA-seq and the second sequencing modality comprises scDNA-seq.
20. The method of claim 17 or 18, wherein the first sequencing modality comprises Illumina® next generation sequencing (NGS), and the second sequencing modality comprises Nextera® NGS, or the first sequencing modality comprises Nextera® NGS, and the second sequencing modality comprises Illumina® NGS.
21. The method of claim 1, wherein step (iii) comprises contacting the cells with proteinase K to lyse the cells.
22. The method of claim 1, wherein the cell is a prokaryotic or eukaryotic cell.
23. The method of claim 1, wherein the cell is an induced pluripotent stem cell (iPSC).
24. The method of claim 1, wherein the cell is a genetically modified cell.
25. The method of claim 24, wherein the cell is modified using a CRISPR / Cas gene editing system.
26. The method of claim 25, wherein the CRISPR / Cas gene editing system is CRISPR interference (CRISPRi).
27. The method of claim 24, wherein RNA expressed by the genetically modified cell is sequenced by scRNA-seq and genomic DNA is sequenced by scDNA-seq.
28. A reaction mixture comprising RNA and gDNA from a single cell, a reverse transcriptase, and a RT primer comprising a capture sequence (CS), a sample barcode (SBC) sequence, a unique molecular identifier (UMI) sequence, and a sequence that binds to mRNA.
29. The reaction mixture of claim 28, wherein the CS hybridizes to a complementary sequence attached to a solid support, SBC sequence comprises a known sequence of variable length, and the UMI sequence comprises a random sequence of variable length, and the sequence that binds to mRNA comprises oligo(dT) or a sequence that hybridizes to any region of the mRNA.
30. The reaction mixture of claim 28 or 29, wherein the CS comprises the nucleic acid sequence GTACTCGCAGTAGTC (SEQ ID NO:1).
31. The reaction mixture of claim 28, wherein the SBC sequence is from 4 to 50 nucleotides in length.
32. The reaction mixture of claim 28, wherein the SBC sequence comprises a nucleic acid sequence selected from the group consisting of TCGCCTTA (SEQ ID NO:4), CTAGTACG (SEQ ID NO:5), TTCTGCCT (SEQ ID NO:6), GCTCAGGA (SEQ ID NO:7), AGGAGTCC (SEQ ID NO:8), CATGCCTA (SEQ ID NO:9), GTAGAGAG (SEQ ID NO:10), and CCTCTCTG (SEQ ID NO:11).
33. The reaction mixture of claim 28, wherein the UMI sequence is from 4 to 50 nucleotides in length.
34. The reaction mixture of claim 28, wherein the UMI sequence comprises the nucleic acid sequence NNNNNNNN (SEQ ID NO:12).
35. The reaction mixture of claim 28, further comprising cDNA molecules produced by reverse transcription of the RNA and a reverse primer that hybridizes to both the cDNA and the gDNA molecules and comprises a R2N or R2 overhang sequence.
36. The reaction mixture of claim 35, wherein the reverse primer comprising the R2N overhang sequence comprises the nucleic acid sequence GTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG (SEQ ID NO:2) and the reverseprimer comprising the R2 overhang sequence comprises the nucleic acid sequence GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT (SEQ ID NO:3).
37. The reaction mixture of claim 28, comprising a, one or more, or a plurality of polynucleotides selected from a sequence in any one of SEQ ID Nos: 1-83.
38. A polynucleotide comprising a capture sequence (CS).
39. The polynucleotide of claim 38, further comprising a sample barcode (SBC) sequence, a unique molecular identifier (UMI) sequence, and / or a sequence that binds to mRNA.
40. The polynucleotide of claim 39, wherein the SBC sequence comprises a known sequence of variable length, and the UMI sequence comprises a random sequence of variable length.
41. The polynucleotide of claim 38, wherein the CS comprises the nucleic acid sequence GTACTCGCAGTAGTC (SEQ ID NO:1).
42. The polynucleotide of claim 39, wherein the SBC sequence is from 4 to 50 nucleotides in length.
43. The polynucleotide of claim 39, wherein the SBC sequence comprises a nucleic acid sequence selected from the group consisting of TCGCCTTA (SEQ ID NO:4), CTAGTACG (SEQ ID NO:5), TTCTGCCT (SEQ ID NO:6), GCTCAGGA (SEQ ID NO:7), AGGAGTCC (SEQ ID NO:8), CATGCCTA (SEQ ID NO:9), GTAGAGAG (SEQ ID NO:10), and CCTCTCTG (SEQ ID NO:11).
44. The polynucleotide of claim 39, wherein the UMI sequence is from 4 to 50 nucleotides in length.
45. The polynucleotide of claim 39, wherein the UMI comprises the nucleic acid sequence NNNNNNNN (SEQ ID NO:12).
46. The polynucleotide of claim 39, wherein the sequence that binds to mRNA comprises oligo(dT) or a sequence that hybridizes to any region of the mRNA.
47. The polynucleotide of claim 38, wherein the polynucleotide comprises a sequence in any one of SEQ ID Nos: 1-83.