Single-cell RNA profiling
The method addresses the limitations of existing RNA sequencing by using single-cell lysing and amplification techniques with sample barcodes, enabling efficient profiling of diverse RNAs for disease and drug analysis.
Patent Information
- Application Number
- PCT/IB2025/051788
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-20
- Filing Date
- 2025-02-19
- Publication Date
- 2025-08-28
AI Technical Summary
Existing RNA sequencing methods, such as microarrays and bulk RNA sequencing, are limited in their ability to profile non-coding RNAs and require complex sample preparation, making them inefficient for single-cell transcription profiling.
A method for sequencing RNA molecules from single cells involving lysing cells with a non-ionic surfactant and temperature variation, adding a crowding reagent, and using priming oligonucleotides and template switching oligonucleotides to synthesize cDNA, followed by amplification and sequencing, with sample barcodes for deconvolution.
Enables efficient sequencing of RNA molecules from single cells, allowing for accurate profiling of both coding and non-coding RNAs without target-sequence specific primers, suitable for disease diagnosis and drug characterization.
Smart Images

Figure IMGF000032_0001 
Figure IMGF000032_0002 
Figure IMGF000033_0001
Abstract
Description
SINGLE-CELL RNA PROFILINGCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application is a non-provisional of and claims the benefit of 63 / 555,796, filed February 20, 2024, which is incorporated by reference in its entirety for all purposes.SEQUENCE LISTING
[0002] This application includes sequences in an XML filed named 624982SEQLST of 33.9KB created February 13, 2025, which is incorporated by reference.BACKGROUND
[0003] Some RNAs are translated into proteins, others serve a structural function, for example, rRNAs in the assembly of ribosomes, others are transporters, e.g., tRNAs, and others serve regulatory functions, for example, short interfering RNA or long non-coding RNAs. Both coding and non-coding RNAs play roles in human diseases such as cancer, cardiovascular, and neurological disorders. Transcription profiling of multiple RNAs within a cell is used in diverse areas of biomedical research, including disease diagnosis, biomarker discovery, and risk assessment of new drugs or environmental chemicals.
[0004] The early approach to transcription profiling used microarrays, a set of defined sequences arranged on a solid substrate. Microarrays almost exclusively represented mRNAs, that is, genes that are translated into proteins.
[0005] The microarray approach has largely been supplanted by high-throughput RNA sequencing, RNA-Seq, which can also detect non-coding RNAs. In this methodology, bulk RNA is extracted from a sample and copied into double-stranded cDNA, which is then sequenced. The sequences obtained can be aligned to reference genome sequences, available in data banks, to identify which genes are transcribed and levels of transcription.SUMMARY OF THE CLAIMED INVENTION
[0006] The invention provides a method of sequencing RNA molecules from single cells, comprising:(a) providing a plurality of isolated single cells;(b) separately lysing the single cells with a non-ionic surfactant and temperature variation to provide separate samples from the single cells comprising RNA populations;(c) adding a crowding reagent to the samples to decrease volume occupied by the RNA populations in the samples;(d) adding at least 5 consecutive identical nucleotides to 3'-termini of RNA molecules of the populations or fragments thereof in the samples;(e) contacting the samples with a priming oligonucleotide comprising a sequence complementary to the sequence of the at least 5 consecutive identical nucleotides and a sample barcode, wherein different samples receive priming oligonucleotides with different sample barcodes, wherein the priming oligonucleotide hybridizes to the sequence of at least 5 consecutive nucleotides added to the RNA molecules or their fragments, and optionally wherein a 5' end of the priming oligonucleotide is linked to a first member of a binding pair;(f) contacting the samples with a reverse transcriptase to synthesize cDNA sequences primed from the priming oligonucleotide hybridized to the RNA molecules or their fragments, which serve as templates for synthesis to obtain double-stranded nucleic acid molecules, followed by addition of untemplated nucleotides forming an overhang at 3' ends of the cDNA sequences;(g) hybridizing a template switching oligonucleotide (TSO) comprising a template switching motif sequence for hybridization to the overhang of untemplated nucleotides linked to a TSO molecular bar code made of at least 6 bases varying among molecules of the TSO, optionally wherein the TSO is linked at its 5' end to the first member of the binding pair provided at least one of the priming oligonucleotide and TSO is linked to the first member of the binding pair;(h) extending the 3' ends of the cDNA sequences with the TSO serving as a template to synthesize extended double-stranded nucleic acids;(i) contacting the extended-double stranded nucleic acids with a second member of the binding pair linked to a support with which to immobilize the extended double-stranded nucleic acids;(j) performing a first amplification of the extended double-stranded nucleic acids from primers having sequences of the priming oligonucleotide and the TSO respectively, the primers having 5' tails;(k) performing a second amplification of amplification products of the first amplification from primers having sequences of the 5' tails;(l) pooling samples from the single cells, wherein the pooling is performed after step (e) and before step (m); preferably between steps (i) and (j) and / or between steps (k) and (m);(m) sequencing amplification products of the second amplification to provide sequencing reads; and(n) deconvoluting sample barcodes from the sequencing reads and thereby assigning the sequencing reads to single cells from which they originated, and thereby determining sequences of RNA molecules from the single cells.
[0007] Optionally, the steps with a possible exception of step (I) are performed sequentially. Optionally, at least one of the primers used in step (k) includes a primer molecular barcode. Optionally, both of the primers used in step (k) include primer molecular barcodes. Optionally, the primer molecular barcodes are i5 and i7, respectively. Optionally, the pooling is performed after step (i).
[0008] Optionally, the deconvoluting step further comprises deconvoluting sequences of the primer molecular bar codes (e.g., i5 and i7 added in step (k)) from the sequencing reads, wherein sequencing reads are assigned to the single cells from which they originated by a combination of sequences of the primer molecule bar codes and a sequence of the sample bar code (added in step (e)).
[0009] Optionally, the method further comprises counting instances of an RNA molecule in a sample from a single cell from a number of different sequences of the TSO molecular barcode associated with sequencing reads of the RNA molecule. Optionally, the method further comprises grouping the sequencing reads from step (m) so that members of a group have the same combination of sequences of primer molecular barcodes, and determining a consensus sequence from sequencing reads in the same group.
[0010] Optionally, the lysis is performed with a lysis solution comprising 1% TERGITOL™ 15- S-9, lOOmM Tris-HCI, 1 U / pl of RNase inhibitor and the temperature variation comprises 2 cycles freeze / thaw at -80°C and heat denaturation at 72°C for 5 minutes. Optionally, thecrowding reagent is a polyol, preferably PEG, optionally PEG8000. Optionally, the PEG is added to 2-25% concentration by weight / 100 mL.
[0011] Optionally, the pooling pools samples from 2-50, optionally 32, single cells.
[0012] Optionally, the first member of the binding pair is biotin and the second member is streptavidin linked to beads. Optionally, 0.9 pg of streptavidin linked to beads is used for each single cell being pooled.
[0013] Optionally, the priming oligonucleotide and TSO comprise a P7 sequence of SEQ ID NO:8 or 12 and a P5 sequence of SEQ ID NO:6 or 11 respectively. Optionally, the TSO molecular barcode has 8-12 nucleotides. Optionally, the TSO molecular barcode has a more or less equal distribution of the four standard nucleotide types.
[0014] Optionally, the TSO molecular barcode has the sequence NNNNNNNNNNNN (SEQ ID NO:17) or NNNNBBNNNBBB (SEQ ID NO:18), wherein N and B are EUPAC-IUB ambiguity codes. Optionally, the TSO has a length of 29-48 nucleotides. Optionally, the TSO is linked by its 5' end to a blocker made of a chemical group. Optionally, the chemical group is a 5'-end abasic site, a 5'-end spacer, and a 5'-end monophosphate or 5'-end biotin.
[0015] Optionally, the first amplification is performed with primers having sequences of SEQ ID NOS:16 andl4respectively. Optionally, the TSO comprises a sequence of SEQ ID NO:1 or SEQ ID NO:2. Optionally, the TSO has a sequence consisting of sequence SEQ ID NO:3.
[0016] Optionally, the at least 5 consecutive identical nucleotides are ribonucleotides, deoxy-ribonucleotides or dideoxy-ribonucleotides of A, T, C, G or U.
[0017] Optionally, the sequencing is performed by adding dNTPs incorporated by a polymerase, each dNTP being conjugated to a label and containing an extension terminator, wherein unincorporated dNTPs are washed, wherein image is captured, wherein dye and terminator are cleaved and wherein these steps are repeated until sequencing is complete. Optionally, the label is a fluorophore and the fluorophore is excited with a laser that emits light of a specific wavelength. Optionally, fluorescence emission from the fluorophore is captured on high resolution CCD camera.
[0018] Optionally, the RNA populations comprise mRNAs, Inc (long non-coding) RNAs, miRNAs, small RNAs, piRNAs, or bisulfite-converted RNAs, or any mixture thereof.
[0019] Optionally, the method further comprising one or more of the additional steps of: treating with DNase; denaturing double strand nucleic acids; fragmenting the RNA molecules; and end-repairing double stranded nucleic acids.
[0020] BRIEF DESCRIPTION OF THE FIGURES
[0021] Fig. 1 shows an exemplary workflow for RNA sequencing.
[0022] Fig. 2 shows first and stage amplification and resulting amplification product.AAAAAAAAAAA is SEQ ID NO:19 and TTTTTTTTTTTTT is SEQ ID NO:20.
[0023] Figs. 3A-D show exemplary sequences of primers and their sequence alignments.Fig. 3A shows a construct after cDNA synthesis. The lower strand includes a priming oligonucleotide at its 5' end and the upper strand includes a TSO at its 5' end. Upper strand in Fig. 3A is 5' to 3' (SEQ ID NO:1)RNA sequence(SEQ ID NO:21)[cell ID]SEQ ID NO:22). Lower strand in Fig. 3A is 5' to 3' (SEQ ID NO:23)[cell ID](SEQ ID NO:24)lstcDNA strand(SEQ ID NO:25). Fig. 3B shows alignment of first-stage amplification primers with the construct of Fig. 3A. Upper strand in Fig. 3B is 5' to 3' (SEQ ID NO:l)2ndcDNA strand(SEQ ID NO:26)[cell ID](SEQ ID NO:27). R-pre-amp primer in Fig. 3B is 5' to 3' (SEQ ID NO:14). F-pre-amp primer in Fig. 3B is 5' to 3' (SEQ ID NO:16). Lower strand in Fig. 3B is 5' to 3' (SEQ ID NO:23)[cell ID](SEQ ID NO:24)lstcDNA strand(SEQ ID NO:25). Fig. 3C shows the alignment of second-stage amplification primers with the amplification product generated in Fig. 3B. R(U)DI amplification primer in Fig. 3C is 5' to 3' (SEQ ID NO:7)[i7](SEQ ID NO:8). Upper strand in Fig. 3C is 5' to 3' (SEQ ID NO:28)2ndcDNA strand(SEQ ID NO:29) [cell I D] (SEQ ID NQ:30). F(U)DI amplification primer in Fig. 3C is 5' to 3' (SEQ ID NO:9)[i5](SEQ ID NQ:10). Lower strand in Fig. 3C is 5' to 3' (SEQ ID NO:8)[cell ID](SEQ ID NO:24)lstcDNA strand(SEQ ID NO:25). Fig. 3D shows the amplification product resulting from second-stage amplification. Upper strand in Fig. 3D is 5' to 3' (SEQ ID NO:9) [i5] (SEQ ID NO:28)2ndcDNA strand(SEQ ID NO:29)[cell ID](SEQ ID NQ:30)[i7](SEQ ID NO:31). Lower strand in Fig. 3D is 5' to 3' (SEQ ID NO:7)[i7](SEQ ID NO:8) [cell ID](SEQ ID NO:32)lstcDNA strand(SEQ ID NO:33)[i5](SEQ ID NO:34).DEFINITIONS
[0024] The term "support" refers to any solid or semi-solid article on which reagents such as nucleic acid molecules can be immobilized, optionally through a binding pair one member of which is linked to the nucleic acid and the other to the support. A support can be in the form of beads (e.g., magnetic beads), spheres, particles, granules, a gel, or a porous matrix. A support can be a component of a flow cell and / or can be included within or adapted to be received by a sequencing instrument. A support can include a polymer, a glass, or a metallic material, organic polymers such as polystyrene, polyethylene, polypropylene, polyfluoroethylene, polyethyleneoxy, and polyacrylamide (e.g., polyacrylamide gel), as well as co-polymers and grafts thereof, latex, dextran, silica, gold, controlled-pore-glass (CPG), or reverse-phase silica. A support can be porous or non-porous and can have swelling or non-swelling characteristics. A support can be shaped to comprise one or more wells, depressions, or other containers, vessels, features, or locations. A plurality of supports can be configured in an array at various locations.
[0025] A target nucleic acid refers to a nucleic acid whose sequence is to be at least partially determined.
[0026] The term "about" when referring to a measurable value such as an amount of a compound, dose, time and the like is meant to encompass 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.5% or 0.1% of the specified amount or value.
[0027] The term "nucleic acid" comprises polymeric or oligomeric macromolecules, including DNA (deoxyribonucleic acid) and RNA (ribonucleic acid), formed of nucleotide units. The four standard nucleotides have bases Adenine (A), Cytosine (C), Guanine (G) and Thymine (T), in DNA and A, C, G and Uracil (U) in RNA. A nucleic acid refers to a multimeric compound comprising nucleotides or analogs that have nitrogenous heterocyclic bases or base analogs linked together to form a polymer, including conventional RNA, DNA, mixed RNA-DNA, and analogs thereof.
[0028] The nitrogenous heterocyclic bases can be referred to as nucleobase units.Nucleobase units can be conventional DNA or RNA bases (A, G, C, T, U), base analogs, e.g.,inosine, 5-nitroindazole and others (The Biochemistry of the Nucleic Acids 5-36, Adams et al., ed., 11th ed., 1992; van Aerschott et al., 1995, Nucl. Acids Res. 23(21): 4363-70), imidazole-4- carboxamide (Nair et al., 2001, Nucleosides Nucleotides Nucl. Acids, 20(4-7):735-8), pyrimidine or purine derivatives, e.g., modified pyrimidine base 6H,8H-3,4-dihydropyrimido[4,5- c][l,2]oxazin-7-one (sometimes designated "P" base that binds A or G) and modified purine base N6-methoxy-2,6-diaminopurine (sometimes designated "K" base that binds C or T), hypoxanthine (Hill et al., 1998, Proc. Natl. Acad. Sci. USA 95(8):4258-63, Lin and Brown, 1992, Nucl. Acids Res. 20(19):5149-52), 2-amino-7-deaza-adenine (which pairs with C and T; Okamoto et al., 2002, Bioorg. Med. Chem. Lett. 12(l):97-9), N-4-methyl deoxygaunosine, 4-ethyl-2'- deoxycytidine (Nguyen et al., 1998, Nucl. Acids Res. 26(18):4249-58), 4,6-difluorobenzimidazole and 2,4-difluorobenzene nucleoside analogues (Kiopffer & Engels, 2005, Nucleosides Nucleotides Nucl. Acids, 24(5-7) 651-4), pyrene-functionalized LNA nucleoside analogues (Babu & Wengel, 2001, Chem. Commun. (Camb.) 20: 2114-5; Hrdlicka et al., 2005, J. Am. Chem. Soc. 127(38): 13293-9), deaza- or aza-modified purines and pyrimidines, pyrimidines with substituents at the 5 or 6 position and purines with substituents at the 2, 6 or 8 positions, 2- aminoadenine (nA), 2-thiouracil (sU), 2-amino-6-methylaminopurine, O-6-methylguanine, 4- thio-pyrimidines, 4-amino-pyrimidines, 4-dimethylhydrazine-pyrimidines, and O-4-alkyl- pyrimidines (U.S. Pat. No. 5,378,825; WO 93 / 13121; Gamper et al., 2004, Biochem. 43(31): 10224-36), and hydrophobic nucleobase units that form duplex DNA without hydrogen bonding (Berger et al., 2000, Nucl. Acids Res. 28(15): 2911-4). Many derivatized and modified nucleobase units or analogues are commercially available (e.g., Glen Research, Sterling, Va.).
[0029] A nucleobase unit attached to a sugar can be referred to as a nucleobase unit, or monomer. Sugar moieties of a nucleic acid can be ribose, deoxyribose, or similar compounds, e.g., with 2' methoxy or 2' halide substitutions. Nucleotides and nucleosides are examples of nucleobase units.
[0030] The nucleobase units can be joined by a variety of linkages or conformations, including phosphodiester, phosphorothioate or methylphosphonate linkages, peptide-nucleic acid linkages (PNA; Nielsen et al., 1994, Bioconj. Chem. 5(1): 3-7; PCT No. WO 95 / 32305), and alocked nucleic acid (LNA) conformation in which nucleotide monomers with a bicyclic furanose unit are locked in an RNA mimicking sugar conformation (Vester et al., 2004, Biochemistry 43(42):13233-41; Hakansson & Wengel, 2001, Bioorg. Med. Chem. Lett. 11 (7):935-8), or combinations of such linkages in a nucleic acid strand. Nucleic acids may include one or more "abasic" residues, e.g., the backbone includes no nitrogenous base for one or more positions (U.S. Pat. No. 5,585,481).
[0031] A nucleic acid may include only conventional RNA or DNA sugars, bases and linkages, or may include both conventional components and substitutions (e.g., conventional RNA bases with 2'-O-methyl linkages, or a mixture of conventional bases and analogs). For example, some otherwise conventional deoxyribonucleotides can include deoxyuridine in place of deoxythymidine. Inclusion of PNA, 2'-methoxy or 2'-fluoro substituted RNA, or structures that affect the overall charge, charge density, or steric associations of a hybridization complex, including oligomers that contain charged linkages (e.g., phosphorothioates) or neutral groups (e.g., methylphosphonates) may affect the stability of duplexes formed by nucleic acids.
[0032] Genetic amplification is a biochemical technology used in molecular biology to amplify by primers a single or few copies of a piece or portion of DNA by replication and copy across several orders of magnitude, generating thousands to millions or more of copies of particular DNA sequence. The most widely use used genetic amplification technology is the polymerase chain reaction, as described in US patent 4, 683, 195-B2 and US 4, 683, 202-B2, using two primers sequences and the heat stable DNA polymerase, such as the Taq polymerase obtained from bacterium Thermus aquatica allowing thermal cycling.
[0033] Transcription mediated amplification (TMA) is an isothermal nucleic-acid-based method that can amplify RNA or DNA targets a billion-fold in less than one hour's time. TMA technology uses two primers and two enzymes: RNA polymerase and reverse transcriptase. One primer contains a promoter sequence for RNA polymerase. In the first step of amplification, this primer hybridizes to the target RNA at a defined site. Reverse transcriptase creates a DNA copy of the target RNA by extension from the 3' end of the promoter primer. The RNA in the resulting RNA:DNA duplex is degraded by the RNase activity of the reversetranscriptase. Next, a second primer binds to the DNA copy. A new strand of DNA is synthesized from the end of this primer by reverse transcriptase, creating a double-stranded DNA molecule. RNA polymerase recognizes the promoter sequence in the DNA template and initiates transcription. Each of the newly synthesized RNA amplicons reenters the TMA process and serves as a template for a new round of replication.
[0034] Reverse transcriptase PCR (RT-PCR) includes three major steps. The first step is reverse transcription (RT), in which RNA is reverse transcribed to cDNA using reverse transcriptase. The RT step can be performed in the same tube with PCR (using a temperature between 40°C and 50°C, depending on the properties of the reverse transcriptase used. The next step involves the denaturation of the dsDNA at temperature at or about 95°C, so that the two strands separate, and the primers can bind again at lower temperatures and begin a new chain reaction. Then, the temperature is decreased until it reaches the annealing temperature which can vary depending on the set of primers used, their concentration, the probe and its concentration (if used), and the cations concentration. An annealing temperature about 5 °C below the lowest Tm of the pair of primers is usually used (e.g., at or around 60 °C). RT-PCR utilizes a pair of primers, which are respectively complementary to sequence on each of the two strands of the cDNA. The final step of PCR amplification is DNA extension from the primers with a DNA polymerase, preferably a thermostable taq polymerase, usually at or around 72°C, the temperature at which the enzyme works optimally. The length of the incubation at each temperature, the temperature alterations, and the number of cycles is controlled by a programmable thermal cycler.
[0035] Complementarity of nucleic acids means that a nucleotide sequence in one strand of nucleic acid, due to orientation of its nucleobase groups, hydrogen bonds to another sequence on an opposing nucleic acid strand. The complementary bases typically are, in DNA, A with T and C with G, and, in RNA, C with G, and U with A. Complementarity can be perfect or substantial / sufficient. Perfect complementarity between two nucleic acids means that the two nucleic acids can form a duplex in which every base in the duplex is bonded to a complementary base by Watson-Crick pairing. "Substantial" or "sufficient" complementarymeans that a sequence in one strand is not completely and / or perfectly complementary to a sequence in an opposing strand, but that sufficient bonding occurs between bases on the two strands to form a stable hybrid complex in set of hybridization conditions (e.g., salt concentration and temperature). Such conditions can be predicted by using the sequences and standard mathematical calculations to predict the Tm of hybridized strands, or by empirical determination of Tm by using routine methods. Tm refers to the temperature at which a population of hybridization complexes formed between two nucleic acid strands are 50% denatured. At a temperature below the Tm, formation of a hybridization complex is favored, whereas at a temperature above the Tm, melting or separation of the strands in the hybridization complex is favored. Tm may be estimated for a nucleic acid having a known G+C content in an aqueous 1 M NaCI solution by using, e.g., Tm=81.5+0.41(% G+C), although other known Tm computations take into account nucleic acid structural characteristics.
[0036] "Hybridization condition" refers to the cumulative environment in which one nucleic acid strand bonds to a second nucleic acid strand by complementary strand interactions and hydrogen bonding to produce a hybridization complex. Such conditions include the chemical components and their concentrations (e.g., salts, chelating agents, formamide) of an aqueous or organic solution containing the nucleic acids, and the temperature of the mixture. Other factors, such as the length of incubation time or reaction chamber dimensions may contribute to the environment (e.g., Sambrook et al., Molecular Cloning, A Laboratory Manual, 2. sup. nd ed., pp. 1.90-1.91, 9.47-9.51, 11.47-11.57 (Cold Spring Harbor Laboratory Press, Cold Spring Harbor, N.Y., 1989)).
[0037] Specific binding of nucleic acids refers to binding between exactly or substantially complementary segment of the nucleic acids to form a stable duplex. Such binding is detectably stronger (higher signal or melting temperature) than binding to other nucleic acids in the sample lacking a segment exactly or substantially complementary to the subject nucleic acids. Lack of binding between nucleic acids can be manifested by binding indistinguishable from nonspecific binding occurring between a randomly selected pair of nucleic acids lackingsubstantial complementarity but of the same lengths as the nucleic acids showing specific binding.
[0038] A primer refers to an oligonucleotide comprising or consisting of a segment, usually of about 12-25 nucleotides, hybridizing specifically to a sequence of interest and which functions as a substrate onto which nucleotides can be polymerized by a polymerase.
[0039] A Template Switching Oligonucleotide (TSO) is an oligonucleotide that hybridizes to untemplated C nucleotides added by a reverse transcriptase during reverse transcription and provides a template for further extension of a cDNA synthesized from an RNA template.
[0040] A "template switching motif sequence" corresponds to the 3' end of a TSO designed to match the overhang nucleotides (that binds to the added bases) by reverse transcription during the template switch (by the reverse transcriptase at the 3' end of the cDNA after first strand synthesis) as described by M. Matz et al (Nucleic Acids Research, vol 27, N° 6 p 1558 - 1560) (1999)).
[0041] A subject refers to an animal, such as a mammalian species (preferably human) or avian (e.g., bird) species, or other organism, such as a plant. More specifically, a subject can be a vertebrate, e.g., a mammal such as a mouse, a primate, a simian or a human. Animals include farm animals, sport animals, and pets. A subject can be a healthy individual, an individual that has symptoms or signs or is suspected of having a disease or a predisposition to the disease, or an individual that needs therapy or is suspected of needing therapy.
[0042] A barcode is a short nucleic acid (e.g., less than 20, 15, 10 or 5 nucleotides long), used to label nucleic acid molecules to distinguish nucleic acids from different samples (a sample barcode), or different nucleic acid molecules in the same sample (a molecular barcode). Barcodes are typically provided as set of molecules having different sequences. In general, sample barcodes are used to distinguish molecules originating from one sample from molecules originating from a different sample. For example, molecules of a first sample typically receive multiple molecules all of the same sample barcode sequence and molecules of second sample receive multiple molecules all of a different sample barcode sequence than received by the first sample. Molecular barcodes are configured to distinguish molecules within the same sample.That is, different molecules within the same sample receive barcode molecules of different sequence. If different and non-overlapping sets of molecular barcode molecules are used for different samples, molecular barcodes can also be used to track sample of origin as well as molecules within the samples. Molecular barcodes can be unique with respect to RNA molecules in a sample in the sense that each RNA molecule, or at least a high percentage of the RNA molecules (e.g., at least 90%, 95% or 99%) receive different molecular barcode sequences from each other. Molecular barcodes can also be non-unique with respect to RNA samples, in which case the number of molecular barcodes is typically sufficient that each instance of the same RNA molecule in the sample receives a molecular bar code sequence different from other instances of that RNA molecule. Barcodes can be single-stranded, double-stranded or have both single and double-stranded components. Barcodes can have the same or different lengths within a set. Barcodes can be random, non-random or semi-random sequences in which at least one position is randomly selected and at least one is not. Barcodes can be synthesized together with pooling of nucleotides at random positions, or individually. Some sets of barcodes having sequences selected such that there is a Hamming distance of at least 2, 3, 4 or 5 nucleotides between each barcode in a set. Barcodes can also be selected to avoid sequences of self- complementary, to avoid hybridizing to other barcode sequences or their complements, or other molecules within a reaction, to avoid sequences subject to sequencing errors, or sequences subject to confusion with sequences of other barcodes and their complements.
[0043] Reference to the sequence of a barcode should be understood as encompassing the sequence of a single-stranded barcode and its exact complement because either or both can be read depending on which strand of a duplex is read and either can be readily translated into the other. For example, if nucleic acids from the same sample linked to the same sample barcode are read some sequencing reads may contain a barcode sequence that is the exact complement of other barcode sequences due to the respective sequences being read from opposing orientations. In such case, the sequencing reads associated with complementary sample barcode sequences are associated with the same cell of origin.
[0044] Exemplary primer sequences use for second-stage amplification are:Single indexing:- P5: 5' AATGATACGGCGACCACCGAGATCTACACGTTCAGAGTTCTACAGTCCGA 3' (whole sequence SEQ ID NO:5, underlined sequence SEQ ID NO:6), and- P7: 5' CAAGCAGAAGACGGCATACGAGATfSEQ. ID NO:7)[index]GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT 3' (underlined sequence SEQ ID NO:8).Dual indexing:- P5: 5' AATGATACGGCGACCACCGAGATCTACACfSEQID NO:9)[i51ACACTCTTTCCCTACACGACGCTCTTCCGATCT(SEQ ID NO:1Q) 3' underlined portion of SEQ ID NQ:10 is SEQ ID NO:11), and-P7: 5'CAAGCAGAAGACGGCATACGAGAT(SEQ ID NO:7)[i71GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT(SEQ ID NO:8) 3' (underlined portion of SEQ ID NO:8 is SEQ ID NO:12). An underlined segment of P5 sequence (SEQ ID NO:6 or 11) can also be included in the TSO and an underlined segments of P7 sequences (SEQ ID NO:8 or 12) can also be included in the priming oligonucleotide as further described below. Exemplary primer sequences used for first-stage amplification when the second amplification with the dual indexing primers shown above are F pre-amp PCR primerACACTCTTTCCCTACACGACGCTCTTCCGATC-31(whole sequence SEQ ID NO:16, underlined SEQ ID NO:13) R pre-amp PCR primer: 5'-GTGACTGGAGTTCAGACGTGTGCTCTTCCGATC-3' (whole sequence SEQ ID NO:14, underlined SEQ ID NO:15). The underlined segment in the forward primer (SEQ ID NO:13) shows sequence in common with SEQ ID NO:11 of the TSO and P5 primer used in second-stage amplification. The underlined segment in the reverse primer (SEQ ID NO:15) shows sequence in common with SEQ ID NO:12 of the priming oligonucleotide and P7 primer used in second-stage amplification. DeoxyUridine (ideoxyU) can be substituted for thymidine in the above sequences.
[0045] Reference to a range of values also includes integers within the range and subranges defined by integers in the range.DETAILED DESCRIPTIONI. General
[0046] Provided are methods for sequencing RNA molecules from single cells. The methods are configured such that initial steps are performed on separate populations of RNA from individual cells, but sequencing and sometimes other steps, such as amplification, can be performed after pooling of the populations from individual cells. Sequences can be deconvoluted to their original cells from barcodes added at one or more steps during the process. The methods do not require target-sequence specific primers or ligation. Single-cell RNA profiles determined by such methods can be useful for diagnosis and characterization of disease, and for characterization of drugs, toxins or environmental materials, forensic analysis or in further research to identify specific RNAs associated with disease or a developmental stage.
[0047] In the description that follows steps can be performed sequentially as described or in other orders including combining two or more steps, so they are performed together. Pooling can be performed at various stages, as further described below.II. Cell types and separation into single cells
[0048] The methods can be practiced on any type of cell including eukaryotic, prokaryotic and eubacterial. For example, cells can mammalian, rodent, insect, avian, amphibian, fish, or fungal. Mammalian cells can be for example, human, non-human primate, bovine, porcine, sheep, goat, rat, mouse or rabbit. Cells can be obtained from a body fluid, tissue or cell culture among other sources. Cells can be primary cells or a cell line. Cells can be a heterogeneous mixture of different cell types or can be enriched for or consist exclusively or essentially of a single cell type. Cells can be from a normal (undiseased) subject or from a subject having a pathological condition, such as a cancer, infection or immune disorder, in which case cells can be obtained from a body fluid or tissue sample, e.g., biopsy expected to include cells characterized by the disorder. If some cells are from a subject having a pathological condition, other cells can be an undiseased subject as a control.
[0049] If cells are provided in the form of tissue, such as a tumor biopsy, the tissue is separated into isolated individual cells. Such can be performed by mechanical mincing of tissue followed by enzymatic digestion (e.g., with dispase) followed by filtration through a stainless steel or nylon mesh or the like to separate dispersed cells from tissue fragments. Cells resulting from this process, or cells obtained directly in the form of a body fluid, can then optionally be subject to enrichment for certain cell types, using techniques, such as immunomagnetic cell separation or FACS™ separation. The resulting cells can then be subject to dilution and the diluted solution pipetted into reaction vessels (e.g., wells of a multiwell plate), such that at least most wells received their own individual cell as shown in Fig. 1(1).III. Lysis
[0050] Lysis can be performed with a lysis solution including a non-ionic surfactant, such as TERGITOL® 15-S-9 (a nonionic secondary alcohol ethoxylate surfactant, CAS number 84133-50- 6, Ci2 i4H25-29O[CH2CH2O]xH, mean molecular weight about 595 Da) or TRITON-X-100®, a buffer, such as Tris and optionally an RNase inhibitor. An exemplary lysis reagent is 1% TERGITOL® 15- S-9, 100 mM Tris-HCI, 1 U / pl RNase inhibitor. Lysis is also facilitated by temperature variation between a temperature below freezing, e.g., below -25, -50 or -75 °C and a temperature above 40, 50, 60, or 70 °C. An exemplary regime of temperature variation comprises 2 cycles freeze / thaw at -80°C and heat denaturation at 72°C for 5 minutes. RNA released by the lysis can be subject to various purification or enrichment procedures, including removal of cell debris, proteins and DNA, digestions with protease or DNase, fractionation of RNA, end-repair of double-stranded RNA structures, denaturation of double-stranded RNA structures, bisulfite conversion, and / or fractionation of RNA by size, class, and / or secondary structure. Classes of RNA that can be analyzed include mRNA, miRNA, small RNA (less than 200 nucleotides), piRNA, small interfering RNA, long non-coding RNA of >200 nucleotides, tRNA rRNA, and bisulfite- converted RNA and mixtures thereof. Fig. 1(2) shows a population of RNA molecules released following cell lysis.IV. Crowding reagent
[0051] Released RNA can be treated with a crowding reagent to decrease the effective volume occupied by RNA thereby increasing the effective concentration of the RNA and thereby increase the rate of subsequent reactions. Crowding reagents include polyols, preferably, polyethylene glycol (PEG), e.g., PEG8000. For example, the crowding reagent can be added at a concentration range of 2-25% by weight per 100 mL of sample.V. Nucleotide tailing to 3' end of RNA
[0052] Released RNAs or their fragments are tailed at their 3' termini with at least 5, at least 10, at least 15, at least 20, at least 25, at least 30, or at least 35, but preferably less than 100, less than 75, or less than 50 consecutive nucleotides as described in WO2015 / 173402-A1 and W02018 / 035170 and illustrated by Fig. 1(3). Tailing is preferably with a homopolymer sequences of identical nucleotides, e.g., polyA. The added nucleotides are preferably selected from ribonucleotides, deoxy-ribonucleotides or dideoxy-ribonucleotides of A, T, C, G or U, and these nucleotides are preferably added by an enzyme, this enzyme being any of a poly(A)— polymerase, poly(U)-polymerase, poly(G)-polymerase, terminal transferase, DNA ligase, RNA ligase and dinucleotides and trinucleotides RNA ligases.VI. Priming oligonucleotide
[0053] Samples are contact with a priming oligonucleotide to hybridize to the nucleotide tail added in the previous step and prime synthesis of a cDNA sequence directed by an RNA template to which the nucleotide tail was added (see e.g., WO WO2015 / 173402-A1 and W02018 / 035170 and Fig. 1(4) . The priming oligonucleotide has several components, which include at the 3' end a binding segment complementary to the nucleotide tail and a sample bar code 5' to the binding segment. The binding segment is preferably of the same or about the same length (e.g., + / -10%) of the nucleotide tail to which it is to hybridize. The sample barcode is a segment of nucleotides constituting a code that can be deconvoluted. Molecules of priming oligonucleotide used for an RNA population from a single cell typically have the same sample bar code sequence. However, the priming oligonucleotide used for an RNA population from one single cell can have a different barcode to the RNA population from a different cell so thatthe different RNA populations can be deconvoluted after pooling. The sample bar codes sequences typically have at least 2, 3, 4, 5 or 6 nucleotides and no more than 10 or 15 nucleotides. The sequences can be randomly generated sequences or can be designed sequences for the assay. Desirable design features include lack of complementarity to other nucleic acids used in the assay and their complements, lack of self-complementary and a different of at least 2 and preferably 3 nucleotides between different sample bar code and their complements to sequences to avoid sequencing errors causing misassignment of RNA populations to the wrong initial cell. The binding segment of the priming oligonucleotide is configured 3' to the sample barcode.
[0054] The priming oligonucleotide includes a primer sequence at its 5' end, which corresponds to the complement of a primer binding site. A primer with the primer sequence is used in subsequent amplification. Optionally, the primer sequence is an Illumina P7 primer sequence, such as SEQ ID NO:8 or 12. The primer sequence in the priming oligonucleotide is preferably positioned 5' to the sample barcode. The primer sequence in the priming oligonucleotide can be a separate from or overlapping with the sample barcode and binding segment. The priming oligonucleotide can also include a first member of a binding pair at its 5' end for subsequent immobilization of the priming oligonucleotide and cDNA primed from it. Examples of binding pairs include biotin and related molecules (e.g., NHS-desthiobiotin) binding to avidin and related molecules (e.g., Streptavidin or NeutrAvidin™) (Liu et al., Chem Soc Rev. 2017 May 9; 46(9): 2391-2403) , fos-jun and coiled-coil peptide pairs, maltose and maltose binding protein, chitin and chitin binding protein, glutathione and glutathione S-transferase, and antibody / epitope combinations. Optionally 0.9 pg of Streptavidin linked to beads is used for single cell being pooled.VII. cDNA synthesis
[0055] RNA molecules hybridized to a priming oligonucleotide are subject to a cDNA synthesis reaction in the presence of a reverse transcriptase and nucleoside triphosphates (usually dATP, dCTP, dGTP and dTTP) contacted with the sample to form a suitable reaction mixture as shown in Fig. 1(5). Examples of reverse transcriptases used in molecular biologyinclude AMV / MAV™, MuLV™, Tth™, MonsterScript™ RT, and Klenow fragment of C. therm. Extension is primed from the 3' end of the priming oligonucleotide with the RNA molecule linked to the priming oligonucleotide serving as a template. cDNA synthesis can proceed in parallel in this reaction for any or all RNA molecules from a single-cell sample. This step generates double-stranded nucleic acid molecules in which one strand comprises an RNA molecule and the other a cDNA. After cDNA synthesis, the reverse transcriptase continues by adding a plurality of untemplated nucleotides, usually C's, sufficient to support sequencespecific hybridization to a TSO. Preferably at least three untemplated C nucleotides are added at the 3' end of the cDNA.VIII. Template switching oligonucleotide
[0056] Samples are contact with a template switching oligonucleotide TSO (e.g., as shown in Fig. 1(5) to hybridize to the non-templated nucleotides added in the previous step and provide a template for further extension of the cDNA to include other useful sequences for subsequent sequence analysis including a primer binding site or complement thereof and optionally an index sequence. The TSO can also include a member of a binding pair as disclosed above for the priming oligonucleotide. Typically, the priming oligonucleotide or the TSO or both includes a member of a binding pair. The TSO is configured to include a segment complementary to the non-templated nucleotides at its 3' end (e.g., polyG), followed by a molecular barcode, followed by a primer sequence (e.g., complement of primer binding site) at the 5' end of the TSO. In other words, the molecular barcode is flanked by a primer sequence at the 5' end and the segment complementary to the non-templated nucleotides at its 3' end. The molecular barcode can serve several purposes. As described in W02018 / 035170, inclusion of a molecular barcode, which can include any of the four common nucleotides at any of its positions, is advantageous in the Illumina platform for initiating a sequencing run, which would otherwise start with a homopolymeric segment. The molecular barcode can also be used in molecular counting of instances of the same RNA molecule in the same cell as further described below. The molecular barcode can also be used in combination with the sample barcode of the priming oligonucleotide to permit tracing RNA molecules to a single cell of origin. Themolecular barcode of the TSO can also be used in grouping amplicons of the same RNA molecule, which can be useful for distinguishing sequencing errors from genuine molecular variation.
[0057] The length of the molecular barcode of the TSO can be the same or similar to the length of the sample barcode, as discussed above, e.g., from 6-12 or 8-12 nucleotides. The molecules of the molecular barcode of the TSO can be random or can be selected for the assay. In either event, preferably each of the four standard nucleotides (A, C, G and T / U) is represented at each position in molecules of a molecular barcode. Preferably the representation is more or less equal (e.g., all four standard nucleotides present at each position with a representation of between 15-35%, preferably 20-30%).
[0058] Preferably the primer sequence or complement of primer binding site present in the TSO is an Illumina P5 sequence (e.g., SEQ ID NO:6 or 11).
[0059] The TSO can have a total length of e.g., 20-50 or 29-48 nucleotides. An exemplary TSO comprises from the 5' end toward the 3' end, a primer sequence, having a length preferably of 10-40 or 18-32 nucleotides, this primer sequence preferably being an ILLUMINA ® P5 a sequence, such as SEQ ID NO:6 or 8 for genetic material amplification, preferably for PCR amplification, flow cell clustering and sequencing primer for read one hybridization, and a template switch motif sequence, having a length preferably of 1-6 or 3-5 nucleotides, this template switch motif sequence being present in this TSO construct, and designed to match the overhang nucleotides added by the reverse transcriptase at the 3' end of a cDNA after first strand synthesis, wherein this template switch motif sequence is linked to this primer sequence by a molecular barcode preferably of about 6-12 nucleotides, which are sequence in the first 8 cycles to 12 cycles of the sequencing.
[0060] Optionally, the 5' end of the TSO is linked to a member of a binding pair, and / or a chemical blocker comprising a 5'-end abasic site ( / dSp / ), a 5'-end spacer (C3, C6, C9) or a 5'-end monophosphate or biotin ( / BioSg / ). BioSg means biotin with a C6 spacer in oligo configuration and dSp means a stable abasic site.
[0061] Optionally, a TSO comprises the sequence SEQ ID NO:1, the sequence SEQ ID NO:2, or the sequence SEQ.ID.NO:3, comprising a molecular barcode of 12 nucleotides having the sequence: NNNNNNNNNNNN (SEQ ID NO:17) but also less or more nucleotides, wherein N is selected from Adenosine (A), Thymidine (T), Guanine (G) or Cytosine (C).
[0062] Optionally, the molecular barcode includes more or less equal representation of each of the four standard nucleotides (e.g., 15-35% or 20-30%) with the possible exception for certain positions where A or C can be omitted or present with less than 15% representation. For example, A can be omitted or present at less than 15% representation at any or all of the 5th, 6th, 7th, 8th, 10th, 11thand 12thpositions. Alternatively or additionally C can be present at less than 15% representation at any or all of the 7th, 8th, 9thor 10thpositions.
[0063] An exemplary TSO comprises or preferably consist of the following sequence: SEQ ID NO:1: 5' / 5BioSg / CACGACGCTC / ideoxyU / / ideoxyU / CCGA / ideoxyU / C / ideoxyU / NNNNBBNNNBBBrGrGrG 3', or 5'- / 5BioSg / TAC ACG TTC AGA GTT CTA CAG TCC GAC GAT CNN NNB BNN NBB BrG rGrG - 3' (SEQ ID NO:2), wherein the primer sequence is present in the known Illumina® adaptor sequence with dual indexing (CATS-ILMN (TruSeq® HT) and without dual indexing (CATS-IMN) (TruSeq®smRNA)) and, S' / SBiotin / GAACGACATGGCTACGATCCGACTT NN NNB BNN NBB BrG rGrG 3' (SEQ ID NO:3) is the construct according to the disclosure, wherein the primer sequence is present in the known MGI ® adaptor sequence (CATS-MGI).
[0064] The sequence rGrGrG is the consensus template switch oligo sequence, rG is a specific nucleoside: a ribonucleotide of the base guanine (G).
[0065] SEQ ID NO:1, SEQ ID NO:2, and / or SEQ ID NO:3 can also or alternatively be linked at the 5' end of the primer sequence to a specific chemical label group, such as a 5'-end abasic site.
[0066] The TSO can be formed entirely of deoxyribonucleotides or can include one or more ribonucleotides. Optionally, the TSO includes at one or more positions an ideoxyU base excisable after reverse transcription by a cocktail of enzymes including any or all of Uracil DNAglycosylase (UDG) and the DNA glycosylase-lyase Endonuclease VIII or Antarctic Thermolabile Uracil DNA glycosylase (UDG) and Endonuclease III.
[0067] After hybridization of the TSO to the non-templated nucleotides, a cDNA is further extended by the reverse transcriptase or other polymerase using the TSO as a template. The further extended nucleic acid now includes from 5' to 3', the priming oligonucleotide, a sequence complementary to the RNA molecule linked to the priming oligonucleotide and a sequence complementary to the TSO. The further extended nucleic acid is duplexed to the RNA, which serves as a template for its synthesis. Optionally, the further extended nucleic acid can be denatured from this RNA.
[0068] The type of structure resulting after templated directed synthesis, extended doublestranded nucleic acid, is shown in Fig. 1 (6). The lower strand from left to right includes the priming oligonucleotide, cDNA templated by an RNA molecule, nontemplated C nucleotides and extended of the cDNA templated by the TSO. The upper strand contains from left to right, identical nucleotides added to RNA, RNA (or second strand cDNA) and the TSO.IX. Immobilization
[0069] After synthesis of extended nucleic acids, the nucleic acids can be immobilized via the binding pair member linked to the 5' end of the priming oligonucleotide or the TSO, if it remains hybridized to the extended nucleic acids. One member of the binding pair (whichever is not linked to the priming oligonucleotide or TSO) is attached to a solid support, such as a magnetic bead and contacted with the samples. Association of the binding pairs immobilizes the other member of the binding pair and an extended nucleic acid linked or otherwise associated with it to the support. Extended nucleic acids linked to the support can be washed to remove reagents used in their synthesis or purification so far in the procedure.
[0070] After immobilization and washing, samples deriving from different individual cells can be pooled as shown in Fig. 1(6) and (7). Different colors are used to represent different sample barcodes.X. Amplification
[0071] Extended nucleic acids are subjected to a two-stage amplification (Fig. 1(8)) and Fig.2. Amplification can be initiated with or without the extended nucleic acids being released from the supports to which they are immobilized in a prior step. First-stage amplification is primed by two primers having sequences common to the priming oligonucleotide and TSO respectively (Fig. 2 (1)). In other words, one primer hybridizes with the complement of the priming oligonucleotide and the other primer hybridizes with the complement of the TSO. Each of the primers has a 5' tail, the complement of which serves as a primer binding site in second stage amplification. Exemplary primers for the first amplification are F pre-amp PCR primer: 5’- ACACTCTTTCCCTACACGACGCTCTTCCGATC-3' (SEQ ID NO:16) and R pre-amp PCR primer: 5'- GTGACTGGAGTTCAGACGTGTGCTCTTCCGATC-3' (SEQ ID NO:14). Second-stage amplification is primed by two primers, which are complementary and hybridize to complements of the 5' tails of the primers (Fig. 2(2)). In other words, the primers used in second-stage amplification have sequences common to the 5' tails on the primers used in the first-stage amplification. First- and second-stage amplification can proceed under typical conditions for PCR including buffer, polymerase and temperature cycling. Optionally, first amplification is conducted for 15 to 25 cycles and second stage amplification is conducted for 2 to 10 cycles. Either or both of the primers used in the first or second stage amplification can include molecular barcodes (primer molecular barcodes). Preferably, the primers used in the second stage amplification include molecular barcodes, preferably Illumina i5 and i7 barcodes. The molecular barcodes can be used as an additional means of tracking sample of origin of an RNA population or for grouping amplicons of the same molecule, as further described below. After first or second stage amplification or both, amplification products can be purified by immobilization and washing e.g., on AMPure XP beads.
[0072] An exemplary result of amplification is shown in Fig. 1(9) and Fig. 2(3). The lower strand includes from right to left, a P7 sequence including a primer molecular barcode, a sample barcode, a homopolymeric T region contributing by the priming oligonucleotide, a first cDNA strand complementary to an RNA in a sample, non-templated nucleotides, a TSOmolecular index and a P5 sequence including a primer molecular barcode. The upper strand, which is the complement of the lower strand includes from right to left, a P7 complement sequence, a sample barcode, a polyA segment added to an RNA, a 2ndcDNA strand, a TSO sequence and a P5 complement sequence.
[0073] Figs. 3A-D show exemplary sequences of primers and their sequence alignments. Fig. 3A shows a construct after cDNA synthesis. The lower strand includes a priming oligonucleotide at its 5' end and the upper strand includes a TSO at its 5' end. Fig. 3B shows alignment of first-stage amplification primers with the construct of Fig. 3A. Fig. 3C shows the alignment of second-stage amplification primers with the amplification product generated in Fig. 3B. Fig. 3D shows the amplification product resulting from second-stage amplification.XI. Pooling
[0074] As previously mentioned, the method includes incorporation of barcodes at various stages. Barcodes are incorporated in the priming oligonucleotide and the TSO and can also be incorporated in any of the primers used in amplification. Any or all of these barcodes can be used to track populations of RNA to their original cell of origin after pooling of such populations from multiple cells. Pooling of samples increases efficiency because multiple RNA populations can be sequenced or otherwise processed in the same reaction as distinct from conducting a separate reaction for each population from each cell. Pooling is generally performed at one or more stages after hybridizing priming oligonucleotides to RNA populations and before sequencing. For example, pooling can be done immediately after hybridizing, after cDNA synthesis, after completion of further extended nucleic acids, before or after immobilization, before first stage amplification, before second stage amplification or between second stage amplification and sequencing. Pooling can also be performed at multiple stages. The number of populations pooled depends in part on the numbers of different barcodes or index sequences available. In some methods, 2, 4, 8, 16, 32, 48, 64 or 92 populations are pooled from a corresponding number of different cells. In some methods, at least 2, 5, 8, 16, 32, 64 or 128 samples are pooled. In some methods, 2-100, 10-50, or 8-64 samples are pooled. In somemethods, pooling is done in two stages. A first pooling step generates packs of RNA populations from individual cells. A second pooling step combines the packs. Sample barcodes from the priming oligonucleotide are used to deconvolute RNA populations to cells from the first pooling and barcodes from primers used in second-stage amplification are used to deconvolute the different packs. Optionally, the first pooling step is performed after immobilization but before amplification and the second pooling step is performed after second-stage amplification and before sequencing. Optionally, the first pooling step combines RNA populations to form multiple packs, each pack representing a combination of RNA populations from 32 or more cells, and the second pooling step combines the packs, optionally wherein ten or more packs are combined.
[0075] XII. Sequencing
[0076] The term sequencing refers to a process for generating a sequence of a target nucleic acid, here an RNA molecule. Sequencing can be performed on individual molecules (single molecule sequencing) or clonal populations originating from the same molecule (clonal sequencing). Sequencing can be part of a template-directed primer extension (sequencing-by- synthesis) in which successive nucleotides incorporated into a nascent chain are detected. Sequencing can also be independent of extension, such as in sequencing by hybridization or ligation. Sequence generated from a single template or clonal population of the same templates can be referred to as a sequencing read. A sequencing read can include sequence of part or all of an RNA target molecule as well as linked sequences including primer sequences or their complement and barcodes within the primer sequences. For example, a sequencing read can include at least 5, 10, 15, 20, 50, 100, 250, or 500 nucleotides of sequence from a target RNA The nucleotides in a sequence read can correspond to contiguous nucleotides in an RNA target or can have contiguous stretches separated by gap(s) of one or more positions where the identity of a nucleotide is not determined. A sequencing read can be formed of the nucleotide types present in a single-stranded target nucleic acids or alternatively formed of their complementary nucleotide types. A forward and reverse sequencing reads refers to sequencegenerated in opposing directions with respect to a single-stranded target nucleic acid. Sequencing reads can be corresponded to specific RNAs by alignment with database sequences.
[0077] 454 pyrosequencing is one example of a sequencing method (Siqueira et al., J OralMicrobiol. 2012;4:10.3402 / jom.v4i0.10743. doi:10.3402 / jom.v4i0.10743). Individual nucleic acids to be sequenced are ligated to adapters and tethered to individual beads, preferably one original fragment per bead. Nucleic acids are then amplified on the beads such that a clonal population forms on the beads. Amplification is followed by chemical detection of DNA synthesis reactions primed by the amplicons in a miniature chamber where pyrophosphate release is measured. By consecutively flooding such a chamber with sequencing reagents containing one of the 4 standard nucleotides, when the correct nucleotide is incorporated in the synthesized strand, pyrophosphate release is measured utilizing a light-generating reaction. The intensity of light also provides information concerning homopolymer runs of nucleotides in the sequence. Pyrosequencing was developed by Pyrosequencing AB, and subsequently acquired by Qiagen who licensed it to 454 Life Sciences, before it was ultimately acquired by Roche.
[0078] Ion Torrent is another example of a sequencing method (see Hu et al., Human Immunology 82, 801-811 (2021)). Nucleic acids to be sequenced are attached to adapters. The adapted DNA fragments are then attached to beads, preferably one bead for one fragment. Fragments are then amplified on the beads by emulsion amplification to generate large populations of beads each having a clonal population of the same original molecule. The beads are then flowed across the chip containing the wells configured such that only one bead can enter an individual well. When the sequencing reagents are then flowed across the wells, when the appropriate nucleotide is incorporated, a hydrogen ion is given off and the signal recorded. Nucleotide incorporation is directly converted to voltage which is recorded directly. The Ion Torrent system is sold by Thermo-Fisher and several versions of the platform are available, including Ion Personal Genome Machine™ (PGM™) System, Ion Proton™ System, Ion S5 system and ION S5 XL system, each with different throughput characteristics (see world wide web thermofisher.com / us / en / home / life-science / sequencing / next-generation-sequencing.html). Anautomated library and template preparation system is also available (Ion Chef™). A large number of applications are supported, including targeted and de novo DNA and RNA sequencing, transcriptome sequencing, microbial sequencing, copy number variation detection, small RNA and miRNA sequencing and CHIP-seq (chromatin immunoprecipitation sequencing).
[0079] Illumina sequencing is based on a technique known as bridge amplification in which nucleic acids with appropriate adapters ligated on each end are used as substrates for repeated amplification synthesis reactions on a solid support (glass slide) that contains primers complementary to a ligated adapter (see, e g., Slatko et al., Curr. Protoc. Mol. Biol. 122(1), e59 (2018)). The primers on the slide are spaced such that the nucleic acids, which is then subjected to repeated rounds of amplification, creates clonal "clusters" consisting of about 1000 copies. Each glass slide can support millions of parallel cluster reactions. During the synthesis reactions, proprietary modified nucleotides, corresponding to each of the four bases, each with a different fluorescent label, are incorporated and then detected. The nucleotides also act as terminators of synthesis for each reaction, which are unblocked after detection for the next round of synthesis.
[0080] Another sequencing method developed by the BGI group is called combinatorial probe-anchor synthesis (cPAS). This consists of rolling circle replication with the Phi 29 DNA polymerase, which synthesizes a long, single-stranded DNA that self assembles into a nanoball (around 300 nm across). Fluorescent probes are incorporated, and the nanoballs are attached to a silicon wafer flow cell where they selectively bind to the positively charged material in a highly ordered pattern. The emission of fluorescence is then imaged and measured to record the base position, the rolling cycle amplification is obtained by addition of a sufficient amount of the Phi 29 DNA polymerase, this polymerase enzyme allows the production of clusters, concatemers or DNA nanoballs (DNBs) into a long single stranded DNA sequence, this sequence comprising several head-to-tail copies of the circular template and wherein the resulting nanoparticle self-assembles into a tight ball of DNA. In this embodiment, the polymerase replicates the looped DNA and when it finishes one circle, it does not stop it, but it continues the replication by peeling off its -previously copied DNA. This copying process continues overand over, thereby forming the DNA cluster or DNA nanoball, as a large mass of repeating DNA to be sequenced all connected together.
[0081] Preferably, the flow cell is a silicon wafer coated with silicon dioxide, titanium, hexamethyldisilazane (HDMS) and a photoresist material and each DNA nanoball selectively binds to the positively charged amino silane according to the pattern.
[0082] In some methods, including Illumina and BGI, sequencing is obtained by adding dNTP incorporated by polymerase, wherein each dNTP is conjugated to a particular label, preferably a label being a fluorophore or dye and containing an extension terminator, wherein unincorporated dNTPs are washed, wherein image is captured, wherein dye and terminator are cleaved and wherein these steps are repeated until sequencing is complete. The added fluorophore can then be excited with a laser that sends light of a specific wavelength so that fluorescence emission from each DNA cluster or DNA nanoball is captured on high resolution CCD camera and wherein the color of each DNA cluster or each DNA nanoball corresponds to a base at an interrogative position, so that a computer can record a base corresponding to that position in generating a sequence.
[0083] Other sequencing methods include, for example, Sanger sequencing, Maxim-Gilbert sequencing, single molecule real time sequencing (Pac-Bio), ONT-sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing direct sequencing, random shotgun sequencing, whole genome sequencing, capillary electrophoreses, gel electrophoresis, duplex sequencing, cycle sequencing, co-amplification at lower denaturation temperature-PCT (COLD-PCR), sequencing by reversible dye terminator, paired- end sequencing, near-term sequencing, exonuclease sequencing, sequencing by ligation, shortread sequencing, nanopore sequencing, semiconductor sequencing, sequencing-by-ligation, sequencing-by-hybridization, Digital Gene Expression (Helicos), next generation sequencing, single molecule sequencing by synthesis (SMSS) (Helicos), massively-parallel sequencing, shotgun sequencing, Oxford Nanopore, Roche Genia, primer walking, MS-PET sequencing or Nanopore platforms, and combinations thereof.XIII. Deconvolution
[0084] Sequencing generates sequencing reads which include sequence of an RNA molecule and at least one barcode. A barcode sequence may have a 1:1 association with a sample of origin. In other words, every sequencing read including a particular barcode sequence derives from the same cell type and every sequencing read with a different barcode sequence derives from a different cell type. Such is typically the case for sample barcodes added as part of a priming oligonucleotide or molecular barcodes in amplification primers. Alternatively molecular barcodes have multiple sequences associated with the same sample. Such is the case for the molecular barcode component of the TSO. In either case, the relationship of a particular barcode sequence to a cell of origin is known allowing deconvolution of sequencing reads so they can be assigned to a cell of origin.XIV. Counting RNA molecules
[0085] Multiple instances of the same nucleic acid type in a sample can be counted by linking each instance to a different barcode sequence before any amplification is performed. The barcodes then provide a means of distinguishing between instances of the same RNA molecule in the original sample (which receive different barcode sequences) and amplification copies, which receive the same barcode sequence as the original RNA from which they were generated. See, e.g., US 8,835,358
[0086] Molecular barcodes incorporated into a TSO provide a means of counting the number of instances of an RNA (same start and stop points) in a population from a single cell. The molecular barcode can be configured such that essentially each cDNA molecule synthesized, or at least each instance of the same type of cDNA molecule synthesized, (same start and stop points) receives a different molecular barcode sequence. Accordingly, counting the number of different molecular barcode sequences linked to cDNA molecule sequences of the same start and end points provides a measure of the number of copies of the corresponding RNA molecule in the original sample. / . Grouping sequencing reads in families
[0087] Molecular barcodes added as a component of the TSO also provide a means of grouping sequencing reads of amplicons of the same original RNA molecule into groups sometimes referred to as families. Members of the same family have the same sequence or combination of sequences of molecular barcodes. Variation among sequencing reads within the same family is likely to arise as a result of amplification errors or sequencing errors. Conversely, a consensus sequence of individual sequences within the same family is likely to represent the actual sequence of an original RNA molecule within a sample. Variation between consensus sequences derived from different families aligning with the same region of RNA is likely to represent variation in sequence among original RNA molecules mapping to the same location of the genome.XVI. Applications
[0088] The present methods indicate the sequences of members of populations of RNA populations present in a single cell type and can also provide a measure of their copy number. The number of RNA molecules included in such a profile can be at least 2, 5, 10, 25, 50, 100, 1000, 10,000 or 100,000. By comparing the populations between cell types of different disease states (e.g., subject with disease and undiseased control), one can identify RNAs that are over- or under-expressed in the disease state relative to normal subject. Disease states that can be analyzed in this manner include cancers, bacterial, fungal or viral infections and immune disorders. Likewise by comparing cells from different stages of development, one can determine RNAs that are over- or under-expressed in one stage relative to another. RNA populations in a patient sample can also be compared with populations from various controls including different diseases or stages of disease and normal subjects as a means of characterizing the patient sample as most similar to one of the controls. Expression profiles determined by the present methods can also be used to characterized drugs, toxins, environmental factors and other chemicals. By comparing the RNA profile of a test molecule with various controls one can determine whether the test molecule has a similar mechanism of action and / or toxicity to the control molecules.
[0089] All publications, patents and patent applications, accession numbers, CAS Registry numbers, websites and the like mentioned in this specification are incorporated by reference to the same extent as if each individual publication, patent or patent application was so individually denoted. To the extent different content is associated with an accession number, CAS Registry number, or other reference at different times, the content in effect as of the effective filing date of this application is meant. The effective filing date is the date of the earliest priority application disclosing the accession number in question. Unless otherwise apparent from the context any element, embodiment, step, feature or aspect of the disclosure can be performed in combination with any other.
[0090] Examples
[0091] Materials & methods+
[0092] D-Plex smRNA-seq libraries were prepared from lOOpg of total RNA extracted and purified from a K562 cell culture using the default miRNeasy protocol according to the manufacturer's instructions (Qiagen, 217084). The total RNA was quantified using a Qubit HS RNA assay (ThermoFisher scientific, Q32852) and qualified using the RNA 6000 Pico kit for Bioanalyzer (Agilent, 5067-1513). The total RNA used for the experiment had a RIN >8. One hundred picograms of the K562 total RNA was introduced for library preparation in the D-Plex smRNA-seq protocol (Diagenode, C05030001) with modifications to produce ultra-low input libraries with and without crowding buffer (CB). The libraries were purified using a l.Ox sample: beads ratio following the AMPure XP beads protocol (Beckman Coulter, A63881). The final elution of the libraries was done in ultrapure water.
[0093] The libraries generated were assayed using the HS DNA analysis kit for Bioanalyzer (Agilent, 5067-4626) and quantified using the Qubit HS dsDNA kit (ThermoFisher Scientific, Q32851).
[0094] The libraries generated using the protocol were assayed using the HS DNA analysis kit for Bioanalyzer (Agilent, 5067-4626) and quantified using the Qubit HS dsDNA kit (ThermoFisher Scientific, Q32851).
[0095] Results
[0096] Crowding buffer
[0097] Tables 1A-D show that the addition of crowding buffer to low input RNA library preparation significantly increases the proportion of positive signal (green square) expected from the library compared to a control without crowding buffer.LYSIS BUFFER AND HEAT DENATURATION
[0098] Material and method
[0099] The total RNA used as a reference point for the RT-qPCR was extracted and purified from a K562 cell culture using the default miRNeasy protocol according to the manufacturer's instructions (Qiagen, 217084). The total RNA was quantified using a Qubit HS RNA assay (ThermoFisher scientific, Q32852) and qualified using the RNA 6000 Pico kit for Bioanalyzer (Agilent, 5067-1513). The total RNA used for the experiment had a RIN >8.
[0100] The miRCURY™ LNA RT kit (Qiagen, Cat No. / ID: 339340) and the miRCURY LNA SYBR® Green PCR Kit (Qiagen, Cat No. / ID: 339346) were used for the RT-qPCR assay alongside the miRCURY™ LNA miRNA PCR Assay (Qiagen, 339306) for the targets: hsa-miR-103a-3p, hsa-miR-486-5p, hsa-let-7a-5p, U6 snRNA and the synthetic control UniSp6 (used as calibrator with and without lysis buffer).
[0101] For the control RNA extracted and purified using the miRNeasy kit ("RNA ext" as referred to in the results section), lOng of total RNA were used in RT and the resulting cDNA solution was diluted 60-fold for the qPCR assay.
[0102] For the test conditions assaying the different lysis conditions, an estimate of 860 K562 cells were used. The resulting lysate was then used for the reverse transcription process. The cDNA solution was also diluted 60-fold for the qPCR measurement of the above-mentioned targets.
[0103] A comparison on hsa-miR103a-3p was done in the same conditions using the NEB cell lysis buffer (NEB, E5530S) to evaluate the efficiency of the lysis procedure relative to a competing solution.
[0104] Appropriate negative controls for the RT reaction and the qPCR reading of the different targets (NT RT ctl and NT qPCR ctl) were used throughout the experiment to evaluate the absence of contaminating DNA or other contaminants possibly influencing the outcome.
[0105] ResultsFormula =LB2 - lysis buffer2 - lx concentration - 1% TERGITOL™ 15S-9, lOOmM Tris-HCI, 1 U / pl of murine RNase inhibitor, nuclease-free waterA - 2 F / T cyclesB - 3 F / T cycles5' - 5 minutes heat denaturation 75°C10' - 10 minutes heat denaturation 75°C
[0106] buffer and heat denaturationTable 2ATable 2B
[0107] Tables 2A-B show that the lysis buffer and physical disruption tested are capable of releasing the cell content and particularly the small RNA content of the cells as indicated by cytoplasmic markers (hsa-miR103a-3p; hsa-miR486-5p; has-let7a) but also a nuclear marker (snU6). The RNA extraction / purification control (RNA ext) for the different markers is showing that the developed formula for the lysis buffer is as efficient as releasing the cell content as a reference extraction method employing acidic phenol, chloroform and ethanol precipitation.
[0108] The hsa-miR103a-3p marker shows that the developed formula is as efficient than a formula marketed by a competitor (NEB single cell lysis module) to lyse cells.
[0109] Cell barcoding and cell multiplexing
[0110] Material and method
[0111] HEK293T cells in culture at Diagenode were sorted in wells of a hard shell 384-well PCR plate (BioRad, HSP3801) so that one event (cell) was directed to a separate well of the plate. The cell sorting was carried out using a SONY MA900 cytometer at the GIGA-Cell imaging platform. Each well was pre-filled with 5 p.1 of cell lysis buffer so that the cell is instantaneously lysed when dropping in a plate well and the cellular RNases are inhibited. The cells were also pre-treated according to the Dead Cell Apoptosis Kit with Annexin V for Flow Cytometry(ThermoFisher scientific, V13242) so that necrotic and apoptotic cells are excluded and not distributed in the recipient plates.
[0112] Upon cell sorting, the methods of the disclosure were applied to generate the libraries starting from different number of cells (20, 24, 32 or 48 cells). The pooling of the different cells' transcriptome was done after reverse transcription and template switching in the protocol and before PCR amplification. Briefly, 2 pl (O.lpg / pl) of washed and conditioned Streptavidin Cl beads (ThermoFisher scientific - Dynabeads™ MyOne™ Streptavidin Cl, 65001) was added to one cDNA reaction (equivalent of one cell) and incubated for 15 minutes at room temperature under mild agitation with an orbital shaker. Once the binding of the cell cDNA was finished, a number of cell-equivalent transcriptomes were pooled together in a 1.5ml microtube. The beads inside the tube were then attracted onto a magnet and the supernatant discarded. The beads pellet was resuspended in 25pl of ultrapure water and transferred in a 0.2ml PCR tube for the amplification of the cell-pool equivalent.
[0113] After amplification, the libraries were purified using a l.Ox sample: beads ratio following the AMPure XP beads protocol (Beckman Coulter, A63881). The final elution of the libraries was done in ultrapure water. Finally, the libraries generated were assayed using the HS DNA analysis kit for Bioanalyzer (Agilent, 5067-4626) and quantified using the Qubit HS dsDNA kit (ThermoFisher scientific, Q32851).Tables 3A-D cell multiplexingTable 4Table 4: % of empty flbrary / artefad according to ceff multiplexing level
[0114] Tables 3A-D and 4 show that the optimal multiplexing level for the cells is 32. This result is taken on the empty library / artefact explaining that the best situation is when the library profile has the least empty library possible.
[0115] Bead capture (indirect: RTP binds to synthetic A tail, fish-out complex with streptavidin beads)
[0116] Material and method
[0117] Upon cell sorting, the D-Plex sc smRNA-seq protocol was applied and the pooling of the different cells' equivalent-transcriptome was done in this experiment either before or after reverse transcription and template switching (capture before RT or capture before PCR).
[0118] Briefly, 2pl (O.lpg / pl) of washed and conditioned Streptavidin Cl beads (ThermoFisher scientific - Dynabeads™ MyOne™ Streptavidin Cl, 65001) was added to onereaction (equivalent of one cell) and incubated for 15 minutes at room temperature under mild agitation with an orbital shaker. Once the binding of the cell cDNA is finished, 32 cell equivalent-transcriptomes were pooled together in a 1.5ml microtube. The beads inside the tube were then attracted onto a magnet and the supernatant was discarded.
[0119] For the experimental condition corresponding to the pooling before RT, the beads pellet was resuspended in 25pl of reverse transcription master mix. For pooling after RT / before PCR, the beads pellet was resuspended in 25 pl of ultrapure water and transferred in a 0.2ml PCR tube for the amplification of the cell-pool equivalent.
[0120] Depending on the experimental condition, the protocol was continued before RT or after RT / before PCR on a pool of 32 cell-transcriptome equivalents.ResultsTables 5A, B and C Bead capture (indirect: RTP binds to synthetic A tail, fish-out complex w streptavidin beads)From AFP21120708 - lOOpg K562 total RNA (23 PCR cycles)
[0121] Tables 5A-C shows that the positive portion of the library profile is significantly higher, and the negative portion of the profile is significantly lower when the bead capture of the cell transcriptome is done before PCR rather than before RT.
[0122] Bead quantity to capture the pooled barcoded transcriptome
[0123] Material and method
[0124] lOOpg of human liver total RNA (TakaraBio, ref. 636531) was used as a proxy to a cell lysate in this experiment. lOOpg of total RNA is used in the D-Plex sc smRNA-seq protocol with bead capture after RT / before PCR but without pooling of the different cells' equivalent- transcriptome as this design rely on a total RNA proxy.
[0125] Briefly, several quantities of washed and conditioned Streptavidin Cl beads (ThermoFisher scientific - Dynabeads™ MyOne™ Streptavidin Cl, 65001) were tested to capture the lOOpg total RNA-equivalent in one binding reaction: 2pg - 0,2pg - 0,02pg. The beads were added to one reaction and incubated for 15 minutes at room temperature under mild agitation with an orbital shaker. Once the binding of the cDNA is finished, the beads were transferred in a 1.5ml microtube. The beads inside the tube were then attracted onto a magnet and the supernatant is discarded. Afterwards, the beads pellet was resuspended in 25pl of ultrapure water and transferred in a 0.2ml PCR tube for the amplification of the cell-pool equivalent.
[0126] The DPLX sc smRNA-seq protocol was continued after RT / before PCR on the lOOpg total RNA-equivalent.
[0127] ResultsTable 6: Bead quantity to capture the pooled barcoded transcriptome
[0128] From AFP21112324 - lOOpg liver total RNA
[0129] Graph 1 highlights that there is an optimum mass of beads to be used for the capture of the cell transcriptome as measured by library output in nanograms. Too many beads create hindrance during the reaction while not enough will not allow to completely capture the transcriptome.
[0130] Bead based library prep <=> in solution library prep
[0131] Material and method
[0132] lOOpg of human liver total RNA (TakaraBio, ref. 636531) is used as a proxy to a cell lysate in this experiment. lOOpg of total RNA is used in the D-Plex sc smRNA-seq protocol either in solution or with a bead capture after RT / before PCR but without pooling of the different cells' equivalent-transcriptome as this design rely on a total RNA proxy.
[0133] Briefly, after RT / before PCR, the cDNA corresponding to lOOpg total RNA is either directly amplified in solution or captured on 0,2pg of washed and conditioned Streptavidin Cl beads (ThermoFisher scientific - Dynabeads™ MyOne™ Streptavidin Cl, 65001). The beads were added to one reaction and incubated for 15 minutes at room temperature under mild agitation with an orbital shaker. Once the binding of the cDNA is finished, the beads were transferred in a 1,5ml microtube. The beads inside the tube are then attracted onto a magnet and the supernatant was discarded. Afterwards, the beads pellet was resuspended in 25pl of ultrapure water and transferred in a 0.2ml PCR tube for the amplification of the cell-pool equivalent.
[0134] The DPLX sc smRNA-seq protocol is continued for the PCR amplification of the cDNA either in solution or captured on beads.ResultsTables 7A-DBead based library prep <=> in solution library prepTables 7A-D indicate that the bead-based library prep is suitable for library preparation compared to reference, in-solution library preparation. Metrics used for this analysis is empty library / artefact presence among the two conditions.
[0135] From AFP21112526 - lOOpg liver total RNA
[0136] Two-step PCR
[0137] Material and method
[0138] After cell sorting, pooling of the 32 different cells' equivalent-transcriptome is done after reverse transcription and template switching (capture before RT or capture before PCR) using 0.2pg of Streptavidin Cl beads (ThermoFisher scientific - Dynabeads™ MyOne™ Streptavidin Cl, 65001). Afterwards, the bead pellet binding the pool of the 32-cell equivalent transcriptome was resuspended in 25 pl of ultrapure water and transferred in a 0.2ml PCR tube for the amplification of the cell pool-equivalent.
[0139] Depending on the experimental conditions, the amplification of the pool of 32-cell transcriptome-equivalent was done either directly in one program of PCR using D-Plex UDI primer pairs (Diagenode, C05030021 / 22 / 23 / 24) (standard PCR) or in a two-step manner.
[0140] In a two-step PCR, the cDNA bound on the beads is amplified in a first round of PCR using a first set of PCR primers (non-indexing). After the first amplification round was completed, the library generated in solution was recovered and cleaned with 1.3x the volume of AMPure XP beads (Beckman Coulter, A63881) following the standard protocol for bead clean-up. The library was then eluted in water and re-amplified in a second round of PCR with a second set of PCR primers (indexing). Upon re-amplification of the cleaned library in the second PCR round, the final library was cleaned with l.Ox the volume of AMPure XP beads (Beckman Coulter, A63881) following the standard protocol for bead clean-up.ResultsTable 8B 2-step PCR> From AFP 220421
[0141] Tables 8A-B indicate that the 2-step PCR is capable of significantly reducing the presence of empty library / artefact compared to the standard PCR amplification. Absence of later peaks shows the library is constructed around co-products, artefacts, oligo-oligo interactions using the standard PCR and not the biological signal coming from the cells. Thepresence of later peaks indicates that the library is constructing around the biological signal coming from the cells and less around artefacts and oligo-oligo interactions.
Claims
What is claimed is:
1. A method of sequencing RNA molecules from single cells, comprising(a) providing a plurality of isolated single cells;(b) separately lysing the single cells with a non-ionic surfactant and temperature variation to provide separate samples from the single cells comprising RNA populations;(c) adding a crowding reagent to the samples to decrease volume occupied by the RNA populations in the samples;(d) adding at least 5 consecutive identical nucleotides to 3'-termini of RNA molecules of the populations or fragments thereof in the samples;(e) contacting the samples with a priming oligonucleotide comprising a sequence complementary to the sequence of the at least 5 consecutive identical nucleotides and a sample barcode, wherein different samples receive priming oligonucleotides with different sample barcodes, wherein the priming oligonucleotide hybridizes to the sequence of at least 5 consecutive nucleotides added to the RNA molecules or their fragments, and optionally wherein a 5' end of the priming oligonucleotide is linked to a first member of a binding pair;(f) contacting the samples with a reverse transcriptase to synthesize cDNA sequences primed from the priming oligonucleotide hybridized to the RNA molecules or their fragments, which serve as templates for synthesis to obtain double-stranded nucleic acid molecules, followed by addition of untemplated nucleotides forming an overhang at 3' ends of the cDNA sequences;(g) hybridizing a template switching oligonucleotide (TSO) comprising a template switching motif sequence for hybridization to the overhang of untemplated nucleotides linked to a TSO molecular bar code made of at least 6 bases varying among molecules of the TSO, optionally wherein the TSO is linked at its 5' end to the first member of the binding pair provided at least one of the priming oligonucleotide and TSO is linked to the first member of the binding pair;(h) extending the 3' ends of the cDNA sequences with the TSO serving as a template to synthesize extended double-stranded nucleic acids;(i) contacting the extended-double stranded nucleic acids with a second member of the binding pair linked to a support with which to immobilize the extended double-stranded nucleic acids;(j) performing a first amplification of the extended double-stranded nucleic acids from primers having sequences of the priming oligonucleotide and the TSO respectively, the primers having 5' tails;(k) performing a second amplification of amplification products of the first amplification from primers having sequences of the 5' tails;(l) pooling samples from the single cells, wherein the pooling is performed after step (e) and before step (m); preferably between steps (i) and (j) and / or between steps (k) and (m);(m) sequencing amplification products of the second amplification to provide sequencing reads; and(n) deconvoluting sample barcodes from the sequencing reads and thereby assigning the sequencing reads to single cells from which they originated, and thereby determining sequences of RNA molecules from the single cells.
2. The method of claim 1, wherein the steps with a possible exception of step (I) are performed sequentially.
3. The method of claim 1, wherein at least one of the primers used in step (k) includes a primer molecular barcode.
4. The method of claim 1, wherein both the primers used in step (k) include primer molecular barcodes.
5. The method of claim 4, where the primer molecular barcodes are i5 and i7. respectively.
6. The method of claim 4, wherein the pooling is performed after step (i).
7. The method of claim 1, wherein the deconvoluting step further comprises deconvoluting sequences of the primer molecular bar codes from the sequencing reads, wherein the sequencing reads are assigned to the single cells from which they originated by acombination of sequences of the primer molecular bar codes and a sequence of the sample bar code.
8. The method of claim 1, further comprising counting instances of an RNA molecule in a sample from a single cell from a number of different sequences of the TSO molecular barcode associated with sequencing reads of the RNA molecule.
9. The method of claim 1, further comprising grouping the sequencing reads from step (m) so that members of a group have the same combination of sequences of primer molecular barcodes, and determining a consensus sequence from sequencing reads in the same group.
10. The method of claim 1, wherein the lysis is performed with a lysis solution comprising 1% TERGITOL™ 15-S-9, lOOmM Tris-HCI, 1 U / pl of RNase inhibitor and the temperature variation comprises 2 cycles freeze / thaw at -80°C and heat denaturation at 72°C for 5 minutes.
11. The method of claim 1, wherein the crowding reagent is a polyol, preferably PEG, optionally PEG8000.
12. The method of claim 11, wherein the PEG is added to 2-25% concentration by weight / 100 mL.
13. The method of claim 1, wherein the pooling pools samples from 2-50, optionally 32, single cells.
14. The method of claim 1, wherein the first member of the binding pair is biotin and the second member is streptavidin linked to beads.
15. The method of claim 14, wherein 0.9 pg of streptavidin linked to beads is used for each single cell being pooled.
16. The method of claim 1, wherein the priming oligonucleotide and TSO comprise a P7 sequence of SEQ ID NO:8 or 12 and a P5 sequence of SEQ ID NO:6 or 11 respectively.
17. The method according to claim 1, wherein the TSO molecular barcode has 8-12 nucleotides.
18. The method according to claim 17, wherein the TSO molecular barcode has a more or less equal distribution of the four standard nucleotide types.
19. The method according to claim 1, wherein the TSO molecular barcode has the sequence NNNNNNNNNNNN (SEQ ID NO:17) or NNNNBBNNNBBB (SEQ ID NO:18), wherein N and B are EUPAC-IUB ambiguity codes.
20. The method according to claim 1, wherein the TSO has a length comprised between 29 nucleotides and 48 nucleotides.
21. The method according to claim 1, wherein the TSO is linked by its 5' end to a blocker made of a chemical group.
22. The method according to claim 21, wherein the chemical group is a 5'-end abasic site, a 5'-end spacer, and a 5'-end monophosphate.
23. The method according to claim 1, wherein the first amplification is performed with primers having sequences of SEQ ID NOS:16 and 14 respectively.
24. The method according to claim 1, wherein the TSO comprises a sequence of SEQ ID NO:1 or SEQ ID NO:2.
25. The method according to claim 1, wherein the TSO has a sequence consisting of SEQ ID NO:3.
26. The method according to claim 1, wherein the at least 5 consecutive identical nucleotides are ribonucleotides, deoxy-ribonucleotides or dideoxy-ribonucleotides of A, T, C, G or U.
27. The method according to claim 1, wherein the sequencing is performed by adding dNTPs incorporated by a polymerase, each dNTP being conjugated to a label and containing an extension terminator, wherein unincorporated dNTPs are washed, wherein image is captured, wherein dye and terminator are cleaved and wherein these steps are repeated until sequencing is complete.
28. The method of claim 27, wherein the label is a fluorophore and the fluorophore is excited with a laser that emits light of a specific wavelength.
29. The method of claim 28, wherein fluorescence emission from the fluorophore is captured on high resolution CCD camera.
30. The method according to claim 1, wherein the RNA populations comprise mRNAs, Inc RNA, miRNAs, small RNAs, piRNAs, or bisulfite-converted RNAs, or any mixture thereof.
31. The method according to claim 1, further comprising one or more of the additional steps of: treating with DNase; denaturing double strand nucleic acids; fragmenting the RNA molecules; and end-repairing double stranded nucleic acids.
Citation Information
Patent Citations
Process for amplifying, detecting, and / or-cloning nucleic acid sequences
US4683195A
Process for amplifying nucleic acid sequences
US4683202A
Backbone modified oligonucleotide analogs
US5378825A
Linking reagents for nucleotide probes
US5585481A
Digital counting of individual molecules by stochastic attachment of diverse labels
US8835358B2