New methods and kits for combinatorial labeling of cellular molecules
The method of generating cDNA in fixed and permeabilized cells using barcoded primers and split-pool labeling addresses the limitations of existing single-cell sequencing, enabling efficient and scalable detection of individual transcriptomes by uniquely labeling molecules within cells.
Patent Information
- Application Number
- PCT/US2025/016912
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-21
- Filing Date
- 2025-02-21
- Publication Date
- 2025-08-28
AI Technical Summary
Existing methods for single-cell sequencing are cumbersome and limited in scalability, often requiring manual separation of cells and specialized equipment, and techniques like microscopy are difficult to implement and limited to sequencing a low number of cells, making it challenging to associate individual transcriptomes with specific cells.
A method involving fixed and permeabilized cells or nuclei, where cDNA is generated using well-specific barcoded primers, followed by split-pool labeling and template switching, allowing for the creation of whole transcriptome sequencing libraries that can be sequenced and analyzed to identify individual cells.
Enables efficient and scalable single-cell sequencing by uniquely labeling molecules, allowing for the specific detection of targets in every cell and associating their expression with the whole transcriptome, improving sequencing efficiency and sensitivity.
Smart Images

Figure US2025016912_28082025_PF_FP_ABST
Abstract
Description
NEW METHODS AND KITS FOR COMBINATORIAL LABELING OF CELLULARMOLECULESFIELD
[0001] The present disclosure relates generally to methods of uniquely labeling or barcoding molecules within or originating from a cell or plurality of cells, a nucleus or plurality of nuclei, or one or more tissues, organs, or organisms. The present disclosure also relates to kits for uniquely labeling molecules within or originating from a cell or plurality of cells, a nucleus or plurality of nuclei, or one or more tissues, organs, or organisms. In particular, the methods and kits may relate to the labeling of RNA, cDNA, genomic DNA, or other molecules within or originating from cells or nuclei for the preparation of sequencing libraries.CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] The present application claims the benefit of and priority to U.S. Provisional Application No. 63 / 556,372, filed February 21, 2024, which is herein incorporated by reference in its entirety.BACKGROUND
[0003] Next Generation Sequencing (NGS) can be used, e.g., to sequence genomic DNA or to identify and / or quantify individual transcripts from a sample of cells. However, such techniques may be complicated to perform on individual cells in large samples. In such methods, RNA transcripts or genomic DNA fragments may be purified from lysed cells (i.e., cells that have been broken apart), and the fragments or cDNA molecules generated from the RNA can then be sequenced using NGS. In such a procedure, all of the cDNA sequences or genomic DNA fragments are mixed together before sequencing, such that RNA expression or genomic DNA sequences are detected for a whole sample and individual sequences cannot be linked back to an individual cell.
[0004] Methods for uniquely labeling or barcoding transcripts or genomic DNA fragments from individual cells can involve the manual separation of individual cells into separate reaction vessels and can require specialized equipment. An alternative approach to sequencing individual transcripts in cells is to use microscopy to identify individual fluorescent bases. However, this technique can be difficult to implement and is limited to sequencing a low number of cells.
[0005] Accordingly, there is a need for new and improved methods of single-cell sequencing that allow the specific detection of particular individual targets in every cell without the need for exhaustive sequencing, together with an ability to associate the presence, expression, and / or identity of the individual targets with the whole transcriptome or a subset thereof. The present disclosure addresses this need and provides other advantages as well.BRIEF DESCRIPTION OF THE DRAWINGS
[0006] FIG. 1 provides an overview of some embodiments of the present method. In the depicted workflow, cells or nuclei are fixed and permeabilized, cDNA is generated within the cells or nuclei by reverse transcription using well-specific barcoded primers, one or more additional barcodes are appended to the cDNA within the cells or nuclei by a split-pool labeling process, the cells or nuclei are lysed, the released barcoded cDNA is isolated, and a second cDNA strand is generated by template switching. The barcoded cDNA is then amplified and used to prepare, in parallel, one or more Whole Transcription (WT) sequencing libraries. Finally, the sequencing libraries are sequenced, and the sequencing reads analyzed.
[0007] FIGS. 2A-2D provides an overview of in situ cell barcoding steps (i.e., combinatorial barcoding, or split-pool labeling) performed in various embodiments of the present methods. FIG. 2A: Round 1 Barcoding. Fixed and permeabilized cells or nuclei are loaded into a multiple wells (e.g., 48 wells) of a Round 1 plate. RNA is reverse transcribed to generate cDNA using reverse transcription primers comprising well-specific barcodes (“BC1”, which may, in some embodiments, be associated with specific samples) and, e.g., a poly(dT) sequence or a random sequence. FIG. 2B: Round 2 Barcoding: The cells / nuclei containing the cDNA are pooled and loaded into a Round 2 Plate. An adapter (e.g., nucleic acid tag as described herein) with a well-specific barcode (“BC2”) is ligated to the cDNA, e.g., to the 5’ end of the cDNA, 5’ to the first barcode. FIG. 2C: Round 3 Barcoding: The cells are pooled and loaded into a Round 3 Plate. A third barcode is ligated to the cDNA (indicated in the darker shade) via another adapter (e.g., nucleic acid tag as described herein), which also contains additional elements such as an Illumina R2 sequence, and biotin. FIG. 2D: Lysis and Sublibrary Generation: Cells are split into multiple sublibraries (or “samples”) (e.g., 8 sublibraries or samples) and lysed.
[0008] FIGS. 3A-3C provide an overview of cDNA capture and amplification steps performed in various embodiments of the present methods. FIG. 3A: cDNA Capture: Following cell lysis, biotinylated cDNA is captured (isolated) in each sublibrary viastreptavidin beads. FIG. 3B: cDNA Template Switch: A template switch (TS) reaction adds an adapter to the 3’ end of the cDNA. FIG. 3C: cDNA Amplification: the whole transcriptome (WT) cDNA is amplified by PCR, e.g., using primers binding to a Template Switch (TS) sequence within the TS adapter (or TS oligo, TSO) and to the Illumina Truseq R2 sequence.
[0009] FIGS. 4A-4C show steps in the preparation of WT sequencing libraries starting from amplified WT cDNA molecules. FIG. 4A: the cDNA molecules are fragmented, and the ends are repaired and then A-tailed. FIG. 4B: Adapter Ligation: An Illumina Truseq R1 Adapter is ligated to the 5’ ends of the DNA. FIG. 4C: Round 4 Barcoding: The sequencing library is amplified, adding P5 / P7 Adapters and a fourth barcode via the UDI - WT Plate.
[0010] FIG. 5 : Whole transcriptome (WT) library structure. The diagram illustrates the composition of the sequencing libraries generated.
[0011] FIGS. 6A-6C illustrate exemplary nucleic acid tags and their use according to certain embodiments of the present disclosure. FIG. 6A depicts a nucleic acid tag, e.g., a Round 2 (R2) nucleic acid tag, with the first and second strands preannealed. The first strand is shown below, and the second strand is shown above. The annealed first and second strands of the nucleic acid tag comprise a central duplex region comprising the annealed first and second strand barcode sequences, and the annealed second strand hybridization sequence and first strand 3’ hybridization sequence. The depicted tag has an overhang at either end of the central duplex, with one overhang comprising the first strand 5’ hybridization region, and the other overhang comprising the second strand overhang sequence. The relative lengths of the sequences in the illustration are not drawn to scale. FIG. 6B shows the nucleic acid tag from FIG. 6A coupled to a barcoded (Rl) reverse transcription (RT) primer. The second strand overhang from the nucleic acid tag of FIG. 6A is annealed to the 5’ overhang sequence of the RT primer. FIG. 6C shows a second nucleic acid tag being coupled to a first nucleic acid tag, e.g., a round 3 tag coupled to a round 2 tag. The second strand overhang sequence from the second nucleic acid tag is annealed to the first strand 5’ hybridization sequence of the first nucleic acid tag.
[0012] FIGS. 7A-7C provide examples of various types of fixed cells following staining with trypan blue. FIG. 7A: high quality fixed samples have single distinct cells with <5% cell aggregation and no debris. FIG. 7B: some aggregation of cells is present. FIG. 7C: cell debris is present; when quantifying fixed cells, it is critical to avoid counting such debris, which could lead to overestimations of the number of cells.
[0013] FIGS. 8A-8B show exemplary cDNA size distributions at two different points in a WT workflow. FIG. 8A. Expected post-amplification sublibrary cDNA size distribution. Example trace of Human cDNA run on a Tapestation. FIG. 8B. Expected Size Distribution before Illumina Sequencing. Example trace of Human DNA from indexed sublibraries run on a TapeStation.
[0014] FIGS. 9A-9B Fixation of cells and nuclei. FIG. 9A: Cells in suspension are fixed and permeabilized before undergoing split-pool combinatorial barcoding steps. FIG. 9B: Nuclei in suspension are fixed and permeabilized before undergoing split-pool combinatorial barcoding steps.
[0015] FIG. 10 A cell / nuclei fixation workflow for 100,000 to 1 million cells, e.g., for whole transcriptome sequencing.
[0016] FIG. 11 A low input cell / nuclei fixation workflow, designed, e.g., for 12 reach ons / samples (in tubes) or 96 reach ons / samples (in plates). In the first step, cells or nuclei present in 20 pL of buffer or media are centrifuged, and the supernatant is removed. In the next step, 60 pL of prefixation buffer is added, and the cells or nuclei are strained. Next, 25 pL of cell / nuclei fixation master mix is added, and the cells or nuclei are incubated on ice for 10 minutes. Next, 8 pL of permeabilization solution is added, and the cells or nuclei are incubated on ice for 3 minutes. Next, 110 pL of fix and perm stop is added. The cells or nuclei can then be frozen, e.g., at -80 °C. At a later time, the tubes or plates are thawed, and 4-12 pL of magnetic cell-binding beads (e.g., ConA beads) are added to each tube or well. The cells / nuclei and beads are mixed 3 times. Next, the tubes or plates are placed on a magnet for 5 minutes, and the supernatant is then removed. The cells / nuclei are then resuspended in cell / nuclei storage buffer, and can then be used in any of a variety of single-cell labeling protocols such as Evercode combinatorial barcoding, immune profiling (e.g., BCR or TCR profiling), or others.
[0017] FIGS. 12A-C. FIGS. 12A-C show a ligation experiment performed to assess the effects of having non-duplexed (FIG. 12A) or duplexed (FIG. 12B) barcode sequences on ligation efficiency after 5 (FIG. 12C) or 30 (FIG. 12D) minutes of ligation. Ligation was quantified by qPCR after ligation reactions of 5 minutes (FIG. 12C) or 30 minutes (FIG. 12D). In FIGS. 12C-12D, ligation efficiencies are reflected in the indicated Cq value (where a lower value indicates a higher level of efficiency). The samples shown are as follows: A, B: neitherBC1 nor BC2 is duplexed; C, D: BC1 duplexed, BC2 non-duplexed; E, F: BC1 non-duplexed, BC2 duplexed; G, H: both BC1 and BC2 duplexed.
[0018] FIGS. 13A-C show a ligation experiment performed to assess the effects of varying lengths of different nucleic acid tag regions on ligation efficiency. The region shown as “A” in FIGS. 13A-13B (corresponding to the 2ndstrand overhang sequence and the 1ststrand 5’ hybridization sequence in FIG. 6C), was varied, as were the regions shown as “B” and “C” in FIGS. 13A-13B (corresponding, together, to the 2ndstrand hybridization and the 1ststrand 3’ hybridization sequence in FIG. 6C). FIG. 13C shows the Cq values obtained for each of the constructs; the lengths of the A and B / C regions for each construct are indicated below the plotted values.SUMMARY
[0019] The present disclosure provides improved methods, compositions, kits, systems, and workflows for labeling nucleic acid molecules such as RNA and genomic DNA, as well as other molecules, within and originating from cells and nuclei.
[0020] In one aspect, the present disclosure provides a nucleic acid tag for barcoding nucleic acid molecules in a cell or nucleus, the tag comprising: a) a first strand comprising: i) a first strand barcode sequence; ii) a first strand 5’ hybridization sequence located 5’ of the first strand barcode sequence; and iii) a first strand 3’ hybridization sequence located 3’ of the first strand barcode sequence; and b) a second strand comprising: i) a second strand barcode sequence, wherein the second strand barcode sequence is complementary to the first strand barcode sequence; ii) a second strand hybridization sequence located 5’ of the second strand barcode sequence, wherein the second strand hybridization sequence is complementary to the first strand 3’ hybridization sequence; and iii) a second strand overhang sequence located 5’ of the second strand hybridization sequence; wherein the first strand and second strand are annealed such that the nucleic acid tag comprises: i) a double-stranded central region comprising the first strand barcode sequence annealed to the second strand barcode sequence and the first strand 3’ hybridization sequence annealed to the second strand hybridization sequence; and ii) a single-stranded overhang located at each end of the nucleic acid, wherein one of the two overhangs comprises the first strand 5’ hybridization sequence, and the other overhang comprises the second strand overhang sequence.
[0021] In some embodiments, the first strand 3’ hybridization sequence and / or the second strand hybridization sequence are from 4-10 nucleotides long. In some embodiments, the firststrand 3’ hybridization sequence is 4, 5, 6, 7, 8, 9, or 10 nucleotides long. In some embodiments, the second strand hybridization sequence is 4, 5, 6, 7, 8, 9, or 10 nucleotides long. In some embodiments, the first strand 3’ hybridization sequence and the second strand hybridization are each 6 nucleotides long. In some embodiments, the first strand 5’ hybridization sequence and / or the second strand overhang sequence are from 5 to 10 nucleotides long. In some embodiments, the first strand 5’ hybridization sequence is 5, 6, 7, 8, 9, or 10 nucleotides long. In some embodiments, the second strand overhang sequence is 5, 6, 7, 8, 9, or 10 nucleotides long. In some embodiments, the first strand 5’ hybridization sequence and the second strand overhang sequence are each 6 nucleotides long.
[0022] In some embodiments, the first strand 3’ hybridization sequence comprises the sequence shown as SEQ ID NO:1 or as SEQ ID NO:2. In some embodiments, the first strand 5’ hybridization sequence comprises the sequence shown as SEQ ID NO: 3 or as SEQ ID NO: 4. In some embodiments, the second strand hybridization sequence comprises the sequence shown as SEQ ID NO:5 or as SEQ ID NO:6. In some embodiments, the second strand overhang sequence comprises the sequence shown as SEQ ID NO:7 or as SEQ ID NO:8. In some embodiments, the first strand barcode sequence and / or the second strand barcode sequence is at least 8 nucleotides long. In some embodiments, the first strand barcode sequence and the second strand barcode sequence are each 8 nucleotides long. In some embodiments, one or more of the first strand or second strand barcode sequences comprise a sequence selected from the group consisting of the barcode sequences present within any of the sequences of Tables 2- 7 and the barcode sequences present within any of the sequences shown as SEQ ID NOS: 11- 590. In some embodiments, the first strand comprises a sequence selected from the group consisting of the barcode sequences present within any of the sequences of Tables 2 and 4 and the barcode sequences present within any of the sequences shown as SEQ ID NOS: 11-105 and SEQ ID NOS: 201-295. In some embodiments, the second strand comprises a sequence selected from the group consisting of the sequences shown in Tables 3 and 5, and the sequences shown as SEQ ID NOS: 106-200 and SEQ ID NOS: 296-400. In some embodiments, the first and / or second strands are DNA molecules. In some embodiments, the first strand further comprises one or more elements located 5’ of the first strand barcode sequence selected from the group consisting of a random nucleotide sequence to prevent counting of PCR duplicates, a capture agent, and a next generation sequencing (NGS) adapter sequence. In some embodiments, the first strand comprises a sequence selected from the group consisting of the sequences shown in Table 4 and the sequences shown as SEQ ID NOS: 201-295.
[0023] In another aspect, the present disclosure provides a set of nucleic acid tags for labeling nucleic acids within a plurality of cells or nuclei, the set comprising a plurality of any of the herein-disclosed nucleic acid tags, wherein all of the nucleic acid tags in the set comprise the same first strand 3’ hybridization sequence, the same first strand 5’ hybridization sequence, the same second strand hybridization sequence, and / or the same second strand overhang sequence; and wherein a plurality of distinct first strand barcode sequences and a plurality of distinct second strand barcode sequences are present among the nucleic acid tags of the set.
[0024] In some embodiments, the nucleic acid tags of the set are distributed into a plurality of aliquots, and wherein the first strand barcode sequences and the second strand barcode sequences present among the nucleic acid tags within each of the aliquots are aliquot-specific. In some embodiments, all of the first strand barcode sequences and all of the second strand barcode sequences present among the nucleic acid tags within any individual aliquot of the plurality of aliquots are the same. In some embodiments, the plurality of aliquots are distributed among the wells of a multi-well plate. In some embodiments, the plurality of aliquots comprises 96 aliquots distributed into the wells of a 96-well plate. In some embodiments, the set of nucleic acid tags comprises at least 96 distinct first strand barcode sequences and 96 distinct second strand barcode sequences. In some embodiments, each of the at least 96 distinct first strand barcode sequences and the at least 96 distinct second strand barcode sequences is present in only one well of the 96-well plate. In some embodiments, the set of nucleic acid tags comprises 96 distinct first strand barcode sequences, such that only one distinct first strand barcode sequence is present in each well of the 96-well plate, and 96 distinct second strand barcode sequences, such that only one distinct second strand barcode sequence is present in each well of the 96-well plate. In some embodiments, the set of nucleic acid tags comprises at least 192 distinct first strand barcode sequences and at least 192 distinct second strand barcode sequences, and wherein two or more distinct first strand barcode sequences and two or more distinct second strand barcode sequences are present in each well of the 96-well plate.
[0025] In another aspect, the present disclosure provides a collection of nucleic acid tags for labeling nucleic acids with cell- or nucleus-specific tags, the collection comprising: two or more sets of any of the herein-disclosed nucleic acid tags, wherein the second strand overhang sequence present in a first set of the two or more sets of nucleic acid tags is complementary to the 5’ hybridization sequence present in a second set of the two or more sets of nucleic acid tags.
[0026] In some embodiments, the nucleic acid tags of each of the two or more sets are distributed into some or all of the wells of a unique multi-well plate; and the first and second strand barcode sequences present among the nucleic acid tags of each set are well-specific.
[0027] In another aspect, the present disclosure provides a collection of nucleic acid tagging agents for labeling nucleic acids with cell- or nucleus-specific tags, the collection comprising: a set of any of the herein-disclosed nucleic acid tags, and a set of reverse transcription (RT) primers each comprising: (i) a poly(T) sequence or a random sequence; (ii) an RT barcode sequence; and (iii) a 5’ overhang comprising a 5’ overhang sequence; wherein the second strand overhang sequence present in the set of nucleic acid tags is complementary to the 5’ overhang sequence present in the set of RT primers.
[0028] In some embodiments, the nucleic acid tags of the set of nucleic acid tags are distributed into some or all of the wells of a first multi-well plate; wherein the nucleic acid tags of the set of RT primers are distributed into some or all of the wells of a second multi -well plate; and
[0029] wherein the first and second strand barcode sequences present among the set of nucleic acid tags and the RT barcode sequences present among the set of RT primers are wellspecific. In some embodiments, one or more of the RT primers present within the set of RT primers comprises a sequence selected from the group consisting of the sequences shown in Tables 6 and 7 and the sequences shown as SEQ ID NOS: 401-590.
[0030] In another aspect, the present disclosure provides a method of labeling nucleic acids with cell- or nucleus-specific tags, the method comprising: (a) providing a plurality of fixed and permeabilized cells or nuclei, each comprising a plurality of RNA molecules; (b) dividing the plurality of cells or nuclei into a first plurality of aliquots, wherein each aliquot comprises more than one cell or nucleus; (c) generating complementary DNA (cDNA) molecules by reverse transcribing RNA molecules within the cells or nuclei of the first plurality of aliquots, wherein the RNA molecules are reverse transcribed using reverse transcription (RT) primers each comprising: (i) a poly(T) sequence or a random sequence; (ii) an RT barcode sequence, wherein the RT barcode sequences present within the RT primers are specific to each aliquot within the first plurality of aliquots; and (iii) a 5’ overhang comprising a 5’ overhang sequence; (d) pooling the cells or nuclei from the first plurality of aliquots; (e) tagging the cDNA molecules within the pooled cells or nuclei from the first plurality of aliquots with one or more nucleic acid tags, thereby generating tagged cDNA molecules, by performing steps (i)through (iii) one or more times: (i) dividing the pooled cells or nuclei into an additional plurality of aliquots; (ii) coupling any of the herein-disclosed nucleic acid tags to the cDNA molecules within the cells or nuclei of the additional plurality of aliquots; and (iii) pooling the cells or nuclei from the additional plurality of aliquots; (f) lysing the cells or nuclei to release the tagged cDNA molecules and produce a lysate comprising the released tagged cDNA molecules; and (g) isolating the released tagged cDNA molecules.
[0031] In some embodiments of the method, step (e)(ii) comprises coupling the cDNA molecules with any of the herein-disclosed sets of nucleic acid tags. In some embodiments, the reverse transcription of step (c) and the tagging of step (e)(ii) are performed using any of the herein-disclosed collections of nucleic acid tagging agents. In some embodiments, steps (e)(i) through (e)(iii) are performed two or more times, thereby generating repeatedly tagged cDNA molecules, and the two or more times are performed using any of the herein-disclosed collections of nucleic acid tags. In some embodiments, the released tagged cDNA molecules are isolated using a binding agent, such that the isolated tagged cDNA molecules are bound to the binding agent. In some embodiments, the isolated tagged cDNA molecules comprise biotin, and the binding agent comprises streptavidin-coated magnetic beads. In some embodiments, the coupling of the nucleic acid tags to the cDNA molecules in step (e)(ii) comprises ligation. In some embodiments, the ligation is performed for a duration of 15 minutes or less. In some embodiments, the method further comprises: (h) generating second strands of the isolated tagged cDNA molecules to produce double-stranded tagged cDNA molecules. In some embodiments, the second strands are generated in step (h) using a template switching oligo (TSO) comprising a TSO adapter sequence. In some embodiments, the method further comprises: amplifying the double-stranded tagged cDNA molecules.
[0032] In some embodiments, the second strands are generated in step (h) using a template switching oligo (TSO) comprising a TSO adapter sequence, wherein the final nucleic acid tags coupled to the tagged cDNA molecules during the one or more times that steps (e)(i) to (e)(iii) are performed comprise an NGS adapter sequence, and wherein one or more rounds of the amplification of the double-stranded tagged cDNA molecules are performed using amplification primers specific to the TSO adapter sequence and the Tag adapter sequence. In some embodiments, prior to lysing the cells or nuclei in step (f), the additional plurality of aliquots is divided into a plurality of sublibraries, and the amplifying of the double-stranded tagged cDNA molecules is performed using amplification primers comprising a sublibraryspecific index sequence. In some embodiments, the method further comprises: preparing asequencing library using the amplified double-stranded tagged cDNA molecules. In some embodiments, preparing the sequencing library comprises fragmenting the amplified doublestranded tagged cDNA molecules and appending a post-fragmentation adapter comprising a post-fragmentation adapter sequence to the fragment ends. In some embodiments, the method further comprises: sequencing the sequencing library.
[0033] In some embodiments, the method further comprises: grouping the sequencing reads according to one or more features selected from the group consisting of RT barcode sequences, first strand barcode sequences, second strand barcode sequences, index sequences, cDNA sequences, and series or combinations of any one or more of RT barcode sequences, first strand barcode sequences, second strand barcode sequences, and index sequences, and cDNA sequences.
[0034] In some embodiments, the method, further comprises: using the grouped sequencing reads to determine the individual cell or nucleus from among the plurality of cells or nuclei from which a given tagged cDNA molecule originated. In some embodiments, the method further comprises: mapping the sequencing reads obtained in the sequencing to a reference genome. In some embodiments, the cells or nuclei comprise mammalian cells or nuclei. In some embodiments, the mammalian cells or nuclei comprise human cells or nuclei or mouse cells or nuclei.
[0035] In another aspect, the present disclosure provides a method of labeling nucleic acids with cell- or nucleus-specific tags, the method comprising: (a) providing a plurality of fixed and permeabilized cells or nuclei, each comprising genomic DNA; (b) fragmenting the genomic DNA within the plurality of cells or nuclei of the first plurality of aliquots to produce a plurality of genomic DNA fragments, and appending a nucleic acid adapter to the ends of the plurality of genomic DNA fragments within the cells or nuclei in a non-target-specific manner, thereby producing a plurality of adapter-coupled genomic DNA fragments; (c) tagging the adapter coupled genomic DNA molecules within the plurality of cells or nuclei with one or more nucleic acid tags, thereby generating tagged genomic DNA fragments, by performing steps (i) through (iii) one or more times: (i) dividing the pooled cells or nuclei into a plurality of aliquots; (ii) coupling any of the herein-disclosed nucleic acid tags to the adapter-coupled genomic DNA fragments within the cells or nuclei of the plurality of aliquots; and (iii) pooling the cells or nuclei from the plurality of aliquots; (d) lysing the cells or nuclei to release thetagged genomic DNA fragments and produce a lysate comprising the released tagged genomic DNA fragments; and (e) isolating the released tagged genomic DNA fragments.
[0036] In some embodiments, step (c)(ii) comprises coupling the adapter-coupled genomic DNA fragments with any of the herein-disclosed sets of nucleic acid tags. In some embodiments, steps (c)(i) through (c)(iii) are performed two or more times, thereby generating repeatedly tagged adapter-coupled genomic DNA molecules, and the two or more times are performed using any of the herein-disclosed collections of nucleic acid tags.
[0037] In another aspect, the present disclosure provides a method of cell- or nucleus- specifically labeling molecules in cells or nuclei, the method comprising: (a) providing a plurality of fixed and permeabilized cells or nuclei; (b) dividing the plurality of cells or nuclei into a first plurality of aliquots, wherein each of the aliquots of the first plurality of aliquots comprises more than one cell or nucleus; (c) coupling a first set of nucleic acid tags as described herein to molecules within the cells or nuclei of the first plurality of aliquots, thereby generating a plurality of tagged molecules, (d) pooling the cells or nuclei from the first plurality of aliquots; (e) dividing the pooled cells from the first plurality of aliquots into a second plurality of aliquots; (f) coupling a second set of nucleic acid tags as described herein to plurality the tagged molecules, thereby generating a plurality of repeatedly tagged molecules, and (g) pooling the cells or nuclei from the second plurality of cells or nuclei.
[0038] In some embodiments, the method further comprises: (h) lysing the cells in the second plurality of aliquots to release the repeatedly tagged molecules; and (i) isolating the released repeatedly tagged molecules. In some embodiments, the molecules are nucleic acids, and the method further comprises: (j) amplifying the isolated repeatedly tagged molecules. In some embodiment, the method further comprises: (k) sequencing the amplified repeatedly tagged molecules. In some embodiments, the method further comprises: (1) grouping the sequencing reads obtained in (k) by first and / or second barcode sequence. In some embodiments, the first barcode sequences are used to identify the individual aliquots among the first plurality of aliquots in which each of the repeatedly tagged molecules was present during step (c); and / or wherein the second barcode sequences are used to identify the individual aliquots among the second plurality of aliquots in which each of the repeatedly tagged molecules was present during step (f). In some embodiments, the first and / or second barcode sequences are used to determine the identities of the individual cells or nuclei from which each of the molecules originated.
[0039] In some embodiments, the molecules are mRNA molecules, and the coupling in step (c) comprises reverse transcribing the nucleic acid molecules using the first set of nucleic acid tags as primers. In some embodiments, the molecules are complementary DNA (cDNA) molecules that have been generated by reverse transcription of RNA within the fixed and permeabilized cells or nuclei prior to step (a). In some embodiments, the molecules comprise genomic DNA, and the coupling in step (c) comprises fragmenting the genomic DNA to generate a plurality of genomic DNA fragments, and appending the first set of nucleic acid tags to the ends of the plurality of genomic DNA fragments. In some embodiments, the molecules comprise adapter-coupled genomic DNA fragments that have been generated by fragmenting genomic DNA and appending adapters to the ends of the genomic DNA fragments within the fixed and permeabilized cells or nuclei prior to step (a). In some embodiments, the molecules comprise polypeptides. In some embodiments, the molecules comprise nucleic acid molecules, and the coupling in step (c) comprises ligating the first set of nucleic acid tags to the nucleic acid molecules. In some embodiments, the coupling in step (f) comprises ligating the second set of nucleic acid tags to the tagged nucleic acid molecules. In some embodiments, the ligating is performed for 15 or fewer minutes. In some embodiments, the ligating is performed for 5, 10, or 15 minutes. In some embodiments, the nucleic acid tags of the first and / or second set of nucleic acid tags are DNA tags.
[0040] In some embodiments of the method, the coupling in step (c) comprises ligating one or more nucleic acid tags as described herein and / or as shown in any of Tables 1-5 or as SEQ ID NOS: 11-400. In some embodiments, 2, 3, 4, 5, 6, 7, 8, 9, or 10 distinct aliquot-specific tag barcode sequences are present in each aliquot among the first plurality of aliquots. In some embodiments, 2, 3, 4, 5, 6, 7, 8, 9, or 10 distinct aliquot-specific amplification barcode sequences are present in each aliquot among the second plurality of aliquots.
[0041] In some embodiments, the nucleic acid molecules comprise genomic DNA, and wherein the coupling in step (c) comprises fragmenting the genomic DNA to generate a plurality of genomic DNA fragments, and appending the first set of nucleic acid tags to the ends of the plurality of genomic DNA fragments. In some embodiments, the fragmenting of the genomic DNA is fragmented using a DNase enzyme or a transposase.
[0042] In some embodiments, the herein-disclosed sets of nucleic acid tags comprises at least 192, 288, 384, 480, 576, 672, 768, 864, or 960 distinct barcode sequences, and / or each aliquot of a plurality of aliquots comprises at least 2, 3, 4, 5, 6, 7, 8, 9 or 10 distinct barcodesequences. In some embodiments, each of the at least 192, 288, 384, 480, 576, 672, 768, 864, or 960 distinct barcode sequences is present in only one well of a 96-well plate. In some embodiments of the herein-disclosed sets of amplification primers, the amplification primers of the set each further comprise a 5’ sequence located 5’ of the barcode sequence that is the same in all of the amplification primers of the set. In some embodiments, the plurality of aliquots are distributed among the wells of a multi-well plate. In some embodiments, the plurality of aliquots comprises 96 aliquots distributed into the wells of a 96-well plate. In some embodiments, the set of amplification primers comprises at least 192, 288, 384, 480, 576, 672, 768, 864, or 960 distinct barcode sequences, and / or wherein each aliquot of the plurality of aliquots comprises at least 2, 3, 4, 5, 6, 7, 8, 9 or 10 distinct barcode sequences. In some embodiments, each of the at least 192, 288, 384, 480, 576, 672, 768, 864, or 960 distinct barcode sequences is present in only one well of the 96-well plate.
[0043] In another aspect, the present disclosure provides a kit for performing any of the herein-described methods, the kit comprising any one or more of the herein-described nucleic acid tags, RT primers, sets of nucleic acid tags or primers, and / or collections of nucleic acid tags or tagging agents.
[0044] In some aspects, implementations of the present methods include hardware, e.g., a device, apparatus, or system configured to perform any of the herein-disclosed methods or processes, and / or computer software on a computer-accessible medium. Such aspects of the present disclosure include computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.DETAILED DESCRIPTION1. Introduction
[0045] The present disclosure relates generally to methods of uniquely labeling or barcoding molecules within a nucleus, a plurality of nuclei, a cell, a plurality of cells, and / or one or more tissues, organs, organisms, or subjects. The present disclosure also relates to kits for uniquely labeling or barcoding molecules within a nucleus, a plurality of nuclei, a cell, a plurality of cells, and / or a tissue, organ or organism. The molecules to be labeled may include, but are not limited to, RNA molecules, cDNA molecules, DNA molecules, proteins, peptides, and / or antigens.
[0046] In various embodiments, the present disclosure provides methods and compositions for creating sequencing libraries, e.g., transcriptome sequencing libraries or genomic DNA libraries. In some embodiments, multiplex libraries are created, e.g., single cell whole transcriptome libraries that are coupled with gene- or vector-enriched libraries, or with libraries related to, e.g., chromatin accessibility or spatial information.
[0047] The present methods and compositions provide improved performance over previous methods of single-cell sequencing, e.g., improved in terms of efficiency, sensitivity, workflow, and other features. The methods and compositions can be used to improve singlecell sequencing of a number of types and used for any of a number of applications. For example, the methods and compositions can be used for RNA sequencing, DNA sequencing (e.g., cDNA or genomic DNA), and in applications including, but not limited to, scRNA-seq, DNase-seq, ATAC-seq, immune profiling (e.g., TCR profiling, BCR profiling), screening (e.g., CRISPR screens such as CROP-seq), validation of biomarkers or drug targets, and others, as well as multiomics applications comprising combinations of any of these types of sequencing and sequencing applications.
[0048] The present methods and compositions, for example, can be used in or applied to any of the methods described in any one or more of U.S. Patent Nos. 10,900,065, 11,634,751, 11,168,355, 11,427,856, 11,555,216, 11,639,519, 12,195,786, 12,043,864, 12,247,247,11,987,838, 12,180,536, 12,227,793, 12,247,248, 11,680,283, 12,234,501, 10,633,648,11,421,221, 12,163,189, PCT Application Nos. PCT / US2019 / 057939, PCT / US24 / 14893, or PCT / US24 / 61803, each of which is herein disclosed by reference in its entirety.
[0049] It will be readily understood that the embodiments, as generally described herein, are exemplary. The following more detailed description of various embodiments is not intended to limit the scope of the present disclosure, but is merely representative of various embodiments. Moreover, the order of the steps or actions of the methods disclosed herein may be changed by those skilled in the art without departing from the scope of the present disclosure. In other words, unless a specific order of steps or actions is required for proper operation of the embodiment, the order or use of specific steps or actions may be modified.2. Definitions
[0050] As will be understood by one of ordinary skill in the art, each embodiment disclosed herein can comprise, consist essentially of, or consist of its particular stated element,step, ingredient, or component. As used herein, the transition term “comprise” or “comprises” means includes, but is not limited to, and allows for the inclusion of unspecified elements, steps, ingredients, or components, even in major amounts. The transitional phrase “consisting of’ excludes any element, step, ingredient or component not specified. The transition phrase “consisting essentially of’ limits the scope of the embodiment to the specified elements, steps, ingredients or components, and to those that do not materially affect the embodiment.
[0051] Unless otherwise indicated, all numbers expressing quantities of ingredients, properties such as molecular weight, reaction conditions, and so forth used in the specification and claims are to be understood as being modified in all instances by the term “about.” Accordingly, unless indicated to the contrary, the numerical parameters set forth in the specification and attached claims are approximations that may vary depending upon the desired properties sought to be obtained by the present disclosure. At the very least, and not as an attempt to limit the application of the doctrine of equivalents to the scope of the claims, each numerical parameter should at least be construed in light of the number of reported significant digits and by applying ordinary rounding techniques. When further clarity is required, the term “about” has the meaning reasonably ascribed to it by a person skilled in the art when used in conjunction with a stated numerical value or range, i.e., denoting somewhat more or somewhat less than the stated value or range, to within a range of, e.g., ±20% of the stated value; ±19% of the stated value; ±18% of the stated value; ±17% of the stated value; ±16% of the stated value; ±15% of the stated value; ±14% of the stated value; ±13% of the stated value; ±12% of the stated value; ±11% of the stated value; ±10% of the stated value; ±9% of the stated value; ±8% of the stated value; ±7% of the stated value; ±6% of the stated value; ±5% of the stated value; ±4% of the stated value; ±3% of the stated value; ±2% of the stated value; or ±1% of the stated value.
[0052] The terms “a,” “an,” “the” and similar referents used in the context of describing the disclosure (especially in the context of the following claims) are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. Recitation of ranges of values herein is merely intended to serve as a shorthand method of referring individually to each separate value falling within the range. Unless otherwise indicated herein, each individual value is incorporated into the specification as if it were individually recited herein. All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise clearly contradicted by context. The use of any and all examples, or exemplary language (e.g., “such as”) provided herein is intendedmerely to better illuminate the disclosure and does not pose a limitation on the scope of the disclosure otherwise claimed. No language in the specification should be construed as indicating any non-claimed element essential to the practice of the disclosure.
[0053] Notwithstanding that the numerical ranges and parameters setting forth the broad scope of the disclosure are approximations, the numerical values set forth in the specific examples are reported as precisely as possible. Any numerical value, however, inherently contains certain errors necessarily resulting from the standard deviation found in their respective testing measurements.
[0054] Groupings of alternative elements or embodiments of the disclosure disclosed herein are not to be construed as limitations. Each group member may be referred to and claimed individually or in any combination with other members of the group or other elements found herein. It is anticipated that one or more members of a group may be included in, or deleted from, a group for reasons of convenience and / or patentability. When any such inclusion or deletion occurs, the specification is deemed to contain the group as modified thus fulfilling the written description of all Markush groups used in the appended claims.
[0055] Definitions and explanations used in the present disclosure are meant and intended to be controlling in any future construction unless clearly and unambiguously modified in the following examples or when application of the meaning renders any construction meaningless or essentially meaningless in cases where the construction of the term would render it meaningless or essentially meaningless, the definition should be taken from Webster's Dictionary, 3rd Edition or a dictionary known to those of ordinary skill in the art, such as the Oxford Dictionary of Biochemistry and Molecular Biology (Ed. Anthony Smith, Oxford University Press, Oxford, 2004).
[0056] The terms “polynucleotide” and “nucleic acid,” used interchangeably herein, refer to a polymeric form of nucleotides of any length, either ribonucleotides or deoxynucleotides. Thus, this term includes, but is not limited to, single-, double-, or multi -stranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, or a polymer comprising purine and pyrimidine bases or other natural, chemically or biochemically modified, non-natural, or derivatized nucleotide bases. The terms “polynucleotide” and “nucleic acid” should be understood to include, as applicable to the embodiment being described, single-stranded (such as sense or antisense) and double-stranded polynucleotides. Unless specifically limited, the term encompasses nucleic acids containing known analogs or derivatives of natural nucleotides, e.g.,molecules that have similar binding properties as the reference nucleic acid. In some embodiments, the nucleic acids can comprise one or more modified nucleotides, e.g., nucleic acids modified at the base moiety, at the sugar moiety, or at the phosphate backbone (e.g., phosphorothioates). In some embodiments, the nucleic acids can comprise one or more moieties to allow or facilitate, e.g., detection, quantification, purification, capture, identification, or selective removal, e.g., biotin, fluorescent labels, etc.
[0057] The term “gene” refers to the segment of DNA involved in producing a polypeptide chain or a non-coding transcript (e.g., mRNA). For coding sequences, it may include regions preceding and following the coding region (leader sequence and / or trailer sequence) as well as intervening sequences (introns) between individual coding segments (exons). A “transgene” refers to a gene that has been introduced into a cell or organism from another source (e.g., from another organism or following synthesis).
[0058] The terms “hybridizable” or “complementary” or “substantially complementary” it is meant that a nucleic acid (e.g. RNA) comprises a sequence of nucleotides that enables it to non-covalently bind, i.e. form Watson-Crick base pairs and / or G / U base pairs, “anneal”, or “hybridize,” to another nucleic acid in a sequence-specific, antiparallel, manner (i.e., a nucleic acid specifically binds to a complementary nucleic acid) under the appropriate in vitro and / or in vivo conditions of temperature and solution ionic strength. As is known in the art, standard Watson-Crick base-pairing includes: adenine (A) pairing with thymidine (T), adenine (A) pairing with uracil (U), and guanine (G) pairing with cytosine (C) [DNA, RNA], In addition, it is also known in the art that for hybridization between two RNA molecules (e.g., dsRNA), guanine (G) base pairs with uracil (U). For example, G / U base-pairing is partially responsible for the degeneracy (i.e., redundancy) of the genetic code in the context of tRNA anti-codon base-pairing with codons in mRNA. In the context of this disclosure, a guanine (G) of a proteinbinding segment (dsRNA duplex) of a subject DNA-targeting RNA molecule is considered complementary to a uracil (U), and vice versa. As such, when a G / U base-pair can be made at a given nucleotide position a protein-binding segment (dsRNA duplex) of a subject DNA- targeting RNA molecule, the position is not considered to be non-complementary, but is instead considered to be complementary. As used herein, the terms “hybridize” or “complementary” refer to a first nucleotide sequence capable of forming non-covalently bind (hydrogen bond) with at least a portion of a specified second nucleotide sequence.
[0059] A "promoter" refers to a set of nucleic acid sequences that direct the transcription of a nucleic acid, e.g., an adjacent coding sequence. Promoters can be constitutive or inducible. As used herein, a promoter includes necessary nucleic acid sequences near the start site of transcription, such as, in the case of a polymerase II type promoter, a TATA element. Promoters as used herein can include bacterial promoters or eukaryotic promoters including RNA polymerase II (e.g., EF-1 alpha) and RNA polymerase III (e.g., U6) promoters. A promoter can also include distal enhancer or repressor elements. The promoter can be a heterologous promoter (i.e., not naturally linked to the coding sequence) or homologous (i.e., the promoter that naturally drives the expression of the transcribed sequence).
[0060] An “expression cassette” is a nucleic acid construct, generated recombinantly or synthetically, with a series of specified nucleic acid elements that permit transcription of a particular polynucleotide sequence in a host cell. An expression cassette may be part of a plasmid, viral genome, or nucleic acid fragment. Typically, an expression cassette includes a polynucleotide to be transcribed (e.g., a protein coding sequence or a non-coding RNA such as a guide RNA), operably linked to a promoter. The promoter can be a heterologous promoter, i.e., a promoter not naturally linked to the transcribed sequence.
[0061] The term “binding” or “coupling” is used broadly throughout this disclosure to refer to any form of attaching or coupling two or more components, entities, or objects. For example, two or more components may be bound to each other via chemical bonds, covalent bonds, non-covalent bonds, ionic bonds, hydrogen bonds, electrostatic forces, Watson-Crick hybridization, nucleic acid sequence complementarity, etc.
[0062] A “barcode” or “index” refers to a nucleotide sequence (the “barcode sequence” or “index sequence”) that is used to label an entity such as a cell, plurality of cells, cell populations, cell compartments, nucleic acids, polypeptides, or other molecules, and that varies among or between cells, cell populations, nucleic acids or other molecules, etc. For example, in some embodiments, a barcode is used to label (or tag) cDNAs generated within a given aliquot of cells, e.g., where all of the cDNA s labeled in the aliquot receive the same barcode, or receive a set of barcodes that is specific to the aliquot, i.e., that the specific set of barcodes used in the aliquot is different from the sets used in the other aliquots. Barcodes can be added to polynucleotides (or other molecules) in any of a number of ways. For polynucleotides, for example, they can be introduced, e.g., in a primer, template, template-switch oligonucleotide (TSO), or other polynucleotide used during a polymerization-based reaction such as reversetranscription, PCR, or other polymerization-based and / or amplification reaction; barcodes can also be added to polynucleotides by hybridization and / or by ligation, e.g., by ligation of an adaptor or other polynucleotide (e.g., a “nucleic acid tag”) comprising a barcode. Such adaptors or other barcode-comprising polynucleotides can be appended to a polynucleotide, e.g., by ligation via blunt-end ligation, ligation to compatible restriction ends, ligation to A-tailed or otherwise tailed ends, using a linker strand, etc. The barcode, or adapter or other polynucleotide comprising the barcode, can be single-stranded, double-stranded, partially double-stranded and partially single-stranded (e.g., comprising one or more overhangs at the 3’ and / or 5’ ends), etc.
[0063] As used herein, when a polynucleotide is said to comprise a “barcode” (or the equivalent term “barcode sequence”) it means that the polynucleotide comprises a sequence of nucleotides that can be used to distinguish the polynucleotide comprising the barcode from one or more other polynucleotides, e.g., from polynucleotides originating from another cell, from polynucleotides labeled in a different aliquot or well, or from all other polynucleotides in a sample. In some embodiments, the barcode alone is sufficient to distinguish the polynucleotide from other polynucleotides, whereas in other embodiments the barcode provides information that can contribute to distinguishing the polynucleotide from other polynucleotides, but is not sufficient on its own (e.g., one or more additional sequence elements, or other markers, are also needed to completely distinguish the polynucleotide).
[0064] It will be appreciated that a “barcode” can refer to a single sequence of contiguous nucleotides, or to a combination of individual sequences of contiguous nucleotides. For example, in certain split-pool labeling methods as described in more detail elsewhere herein, multiple rounds of tagging can be performed, e.g., multiple rounds in each of which cells are divided into aliquots, nucleic acid tags (comprising a barcode) are added to molecules (such as cDNA s) in the cells of each aliquot, and the cells are then recombined (or repooled). As such, the tagged cDNAs in the cells of the aliquots can comprise two nucleic acid tags, each comprising a barcode. In some embodiments, the two barcodes present on the same molecule can be referred to herein as a single “barcode,” even if there are other sequence elements (such as linker sequences, adapter sequences, primer-binding sequences, etc.) intervening between the two barcodes on the molecule.
[0065] Further, it will be appreciated that when a polynucleotide is said to comprise a “barcode,” this can mean that, depending on the context, the specific sequence of the barcode (or “barcode sequence”) can vary between the different polynucleotides comprising thebarcode, or that the specific barcode sequence is the same between the different polynucleotides. For example, in some embodiments, all of the cDNAs in cells of a given aliquot (or sample) are tagged with nucleic acid tags comprising the same barcode sequence, whereas the cDNAs of cells in other aliquots are tagged with nucleic acid tags comprising other barcode sequences. In some embodiments, however, a polynucleotide comprising a barcode is added to molecules in a given aliquot (or sample), wherein the specific barcode sequence differs between the different polynucleotides used in the aliquot or sample; such polynucleotides and barcodes can be used, for example, to distinguish between the different original molecules in the sample (e.g., to be able to detect errors arising during amplification of the original molecules, such as cDNAs derived from original mRNA molecules).
[0066] In some embodiments of the present disclosure, one or more barcodes is included in primers used for reverse transcription (i.e., RT primers) or amplification (e.g., PCR), or in a nucleic acid tag appended to a polynucleotide, e.g., by ligation, where the cells are divided into two or more (e.g., 2, 4, 8, 12, 16, 24, 32, 48, 96, or more) aliquots or wells prior to the reverse transcription, amplification, or ligation, and the barcode sequences used in the cells are aliquot- or well-specific. When barcode or index sequences are said to be “aliquot-specific” or “wellspecific,” this means that there is an association between the sequences used and the presence of the different cells within the two or more aliquots or wells, such that the association can be used to derive information about the location of a given cell within the aliquots or wells based upon the specific sequence. For example, in some embodiments each cell within a given aliquot or well has primers or tags with the same barcode sequence, and the barcode sequences are different between each aliquot or well. However, it will be appreciated that even where a less direct relationship exists between the barcode sequences and the aliquots or wells (e.g., where more than one barcode sequences are used within a given aliquot or well, or where more than one aliquot or well share one or more barcode sequences), the barcodes are still considered aliquot-specific or well-specific, so long that some information can be derived from the barcode sequence about the aliquot or well in which a given cell or nucleus was present.
[0067] “Split-pool labeling” or “split-pool barcoding” or “combinatorial labeling” or “combinatorial barcoding” refers to a cell-specific labeling method involving the use of fixed and permeabilized cells or nuclei as containers, wherein a plurality of the cells or nuclei are first separated into multiple wells (or aliquots), followed by the labeling of RNA or other molecules within each cell or nucleus using a well-specific tag or barcode, followed by the pooling of the cells, and wherein this cycle of separation, tagging, and pooling is repeated oneor more times. In this way, at the end of the process each cell or nucleus within the plurality will comprise a combination of tags or barcodes that will reflect the particular combination of wells or aliquots in which it was present throughout the multiple rounds of tagging. As the number of rounds of tagging and / or the number of wells or aliquots used in each round is increased, the number of potential barcode combinations increases correspondingly. As such, for a given number of cells or nuclei in the plurality a suitable experimental design can be prepared that will generate a high likelihood that tagged molecules, e.g., cDNA s, within each cell or nucleus will have the same combination of barcodes that is unique among the overall population of cells or nuclei. Examples of split-pool labeling methods are disclosed, e.g., in US Patent Nos. 10,900,065, 11,634,751, 11,168,355, 11,427,856, 11,555,216, 11,639,519,12,195,786, 12,043,864, 12,247,247, 11,987,838, 12,180,536, 12,227,793, 12,247,248,11,680,283, 12,234,501, 10,633,648, 11,421,221, 12,163,189, PCT Application Nos.PCT / US2019 / 057939, PCT / US24 / 14893, PCT / US24 / 61803, in Rosenberg et al., Science 360, 176-182 (2018), Rosenberg etal., BioRxiv (2017), “Scaling single cell transcriptomics through split pool barcoding,” doi.org / 10.1101 / 105163, Tran et al. BioRxiv (2022) “High sensitivity single cell RNA sequencing with split pool barcoding,” doi.org / 10.1101 / 2022.08.27.505512, the entire disclosures of all of which are herein incorporated by reference (including all supplemental material).
[0068] As used herein, the term “tagged nucleic acid molecules” refers to nucleic acid molecules (e.g., complementary DNA (cDNA) molecules, genomic DNA fragments) comprising one or more barcodes (e.g., well-specific barcodes), e.g., cDNA molecules or genomic DNA fragments generated within cells or nuclei to which one or more nucleic acid tags have been appended (e.g., by ligation).3. Analysis methods in single cells or nuclei
[0069] One aspect of the present disclosure relates to methods of labeling nucleic acids (or other molecules), e.g., labeling nucleic acids such as RNA, cDNA, and / or genomic DNA in a cell- or nucleus-specific manner, such that the nucleic acids within the cells or nuclei each comprise a cell- or nucleus-specific label.
[0070] It will be appreciated that, as used herein, when a method is said to, e.g., label or tag nucleic acids “in” or “within” a cell or nucleus, this means that at least one of the labeling steps takes place at the interior of the cell or nucleus (e.g., at least a first step or a first set of steps), but does not necessarily mean that all labeling or tagging steps take place at the interiorof the cell or nucleus. For example, in some embodiments of the present methods, one or more tagging steps, such as those involving reverse transcription to generate cDNA molecules and the subsequent coupling of one or more nucleic acid tags to the cDNA molecules, may take place at the interior of the cells or nuclei, and one or more subsequent steps, such as template switching and preamplification / amplification steps, may be performed on tagged cDNA molecules isolated from the cells or nuclei following their lysis.4. Split-pool barcoding
[0071] The present methods involve labeling molecules such as nucleic acids, e.g., RNA, cDNA, or genomic DNA, with barcodes in a process called, alternatively, split-pool labeling, tagging, or barcoding, or combinatorial labeling, tagging, or barcoding. In various embodiments, the split-pool barcoding step of the protocol may be repeated a number of times sufficient to generate a unique combination or series of labeling sequences for the cDNAs, genomic DNA fragments, or other molecules in each sequencing library such that all (or virtually all, or the great majority) of the tagged nucleic acid molecules originating from a given cell (or nucleus) will have the same combination or series of labeling sequences (also referred to as barcode sequences or index sequences), and that the complexity of the combinations or series of labeling sequences is such that each combination or series is unique, or essentially unique, among all of the cells in the plurality of cells. For example, the labeling could be repeated enough times to generate a sufficient number of distinct combinations or series of labeling sequences that each individual combination or series has, e.g., at least a 95%, 96%, 97%, 98%, 99%, or higher probability of being unique among all of the combinations or series in the cells of the plurality.
[0072] Stated another way, the split-pool barcoding may be repeated a number of times such that the tagged nucleic acid molecules in the first cell may have a first unique series of labeling sequences, the tagged nucleic acid molecules in a second cell may have a second unique series of labeling sequences, the tagged nucleic acid molecules in a third cell may have a third unique series of labeling sequences, and so on. The methods of the present disclosure may provide for the labeling of tagged nucleic acid molecules sequences from single cells with unique barcodes, wherein the unique barcodes may identify or aid in identifying the cell from which the tagged nucleic acid molecules originated. In other words, a portion, a majority, or substantially all of the tagged nucleic acid molecules from a single cell may have the same barcode, and that barcode may not be repeated in tagged nucleic acid molecules originatingfrom one or more other cells in a sample (e.g., from a second cell, a third cell, a fourth cell, etc.).
[0073] The barcodes used in the present methods can be added to the tagged nucleic acid molecules at any of a number of steps, e.g., during reverse transcription, during one or more rounds of ligation-based tagging (e.g., appending a nucleic acid tag to a cDNA molecule or genomic DNA fragment), or during amplification (e.g., introduced via a PCR primer).
[0074] The present barcodes offer numerous advantages relative to other barcodes and nucleic acid tags. For example, in some embodiments, the present nucleic acid tags comprise certain features that enhance their utility in ligation-based tagging steps, e.g., improved ligation efficiency. Such improved efficiency can be reflected, e.g., in more rapid ligation reactions (e.g., 5, 10, or 15 minutes) and / or lower temperatures for the ligation reaction (e.g., 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 °C).
[0075] In one aspect, the present barcodes allow a single barcode to be used to determine the cell (or nucleus) origin of molecules such as nucleic acids within or originating from a plurality of cells or nuclei. In particular, the barcodes comprise an additional level of complexity, e.g., as provided by the use of multiple barcodes per aliquot in multiple rounds of a split-pool barcoding process, that allows a determination of the cell (or nuclear) origin of each molecule within the cells (or nuclei), but can also be used to identify duplicates arising from, e.g., PCR errors, i.e., to distinguish such duplicates resulting from technical errors from duplicates representing distinct molecules (sharing the same sequence) that were originally present in a given cell or nucleus. Such barcodes can be generated, e.g., by deliberately designing multiple distinct barcodes for each aliquot, or through a semi-random process, e.g., using an error-prone method of synthesizing or generating barcodes such that a suitable amount of diversity is present among the barcodes used in each aliquot that can be analyzed to distinguish between different original molecules. In such embodiments, the barcodes used in each aliquot would typically be significantly different from the barcodes used in the other aliquots (e.g., barcodes used in different aliquots would differ by at least, e.g., 4, 5, 6, 7, 8 or more nucleotides, such that an error-prone process that introduces, e.g., at most 1-3 nucleotide changes among the barcodes within a given aliquot would still allow the altered barcodes in each aliquot to be distinguished from barcodes in other aliquots).
[0076] In some embodiments, barcoded (or tagged) tagged nucleic acid molecules are mixed together and sequenced (e.g., using NGS), such that data can be gathered regarding, e.g.,RNA expression or chromatin accessibility at the level of a single cell. In some embodiments, nucleic acid molecules are barcoded (or tagged) together, but then sequenced independently. In either case, so long as the tagged nucleic acid molecules originated from the same original cells (or nuclei) and were labeled together such that all of the cDNA (or other nucleic acid) molecules originating from a given original cell or nucleus comprise the same combination of barcodes, the sequencing data can be combined and analyzed together, even if certain tagged nucleic acid molecules from the same individual cells are present in different sequencing libraries and / or sequenced independently. Certain embodiments of the methods of the present disclosure may be useful in assessing, analyzing, or studying the transcriptome (i.e., the different RNA species transcribed from the genome of a given cell or nucleus) of one or more individual cells or nuclei. Some embodiments of the present methods may be useful in assessing, analyzing, or studying the accessibility of genomic loci (e.g., chromatin accessibility) by virtue of, e.g., fragmenting the genomic DNA using a DNase enzyme, MNase enzyme, or transposase.
[0077] As discussed above, for any one or more steps of any the present methods, including, but not limited to, reverse transcription steps, genomic fragmentation and adapter ligation steps, nucleic acid tag coupling steps, cell lysis steps, template switching steps, amplification steps, or enrichment steps, an aliquot or group of cells (or nuclei, or lysates) can be separated into different reaction vessels or containers. Vessels or containers can also be referred to herein as receptacles, samples, and wells, and the terms vessel, container, receptacle, sample, and well may be used interchangeably herein. In some embodiments, cells or nuclei may be separated into a number of different reaction vessels. For example, the number of reaction vessels may include four 1.5 ml microcentrifuge tubes, a plurality of wells of a 96- well plate, or another suitable number and type of reaction vessels. In some embodiments, the reaction vessels or containers include one or more 96-well plates (or, e.g., 6, 12, 24, 48, 384 well plates).
[0078] For combinatorial barcoding (or split-pool tagging), cells or nuclei can be distributed into a plurality of aliquots and polynucleotides within the cells or nuclei labeled with an aliquot-specific barcode (e.g., by reverse transcription of RNA within the cells or nuclei using primers comprising the barcode, or by appending a nucleic acid tag to polynucleotides within the cells or nuclei wherein the tags comprise aliquot-specific barcodes), the aliquots can then be repooled, washed, and separated again into a new plurality of aliquots, and a further set of barcodes can be added to the polynucleotides. In this way, after repeated rounds ofseparating, tagging, and repooling, tagged nucleic acid molecules within each cell or nucleus may be bound to a unique combination or sequence of barcodes, or substantially unique combination or sequence of barcodes. In some embodiments, all (or most, depending, e.g., on the efficiency of the tagging reactions in a given cell or nucleus) of the tagged nucleic acid molecules within any individual cell or nucleus within a plurality of cells or nuclei will comprise the same combination of barcodes (or barcode sequences). In some embodiments, the combination or sequence of barcodes can be used to identify, or help identify, the individual cell from which a given tagged nucleic acid (e.g., cDNA molecule, genomic DNA fragment) originated.
[0079] As described above, in some embodiments, in a given barcoding step cells or nuclei within each well or aliquot are tagged with a different barcode, i.e., all of the barcodes (or barcode sequences) used within the well or aliquot are the same, while the barcode sequences are different in each of the wells or aliquots. However, other barcoding strategies are possible as well, e.g., in which more than one barcode sequence is used within a given well or aliquot, or in which one or more barcode sequences are present in multiple wells or aliquots during the barcoding step. In general, any barcoding protocol can be encompassed by the present disclosure so long that during the protocol the labeled molecules (e.g., RNA molecules, or cDNA molecules produced therefrom, or genomic DNA fragments) within each cell or nucleus acquire a combination of barcodes that reflects the different wells or aliquots in which the cell or nucleus was present.
[0080] The different labeling sequences can be introduced at one or more steps, including during reverse transcription (e.g., wherein each reverse transcription (RT) primer comprises a barcode), during one or more subsequent labeling steps (e.g., ligating, tagmentation, or otherwise coupling a nucleic acid tag comprising a barcode sequence to a cDNA), or during one or more amplification steps (e.g., using one or more primers that include a barcode or index sequence). For example, in some embodiments of the present disclosure, RNA is labeled within cells or nuclei by generating cDNA through reverse transcription (RT) using well-specific barcode-containing primers, and subsequently additional well-specific barcodes are ligated to the cDNA molecules in one or more round of split-pool tagging, and finally yet more barcodes (or indexes) are added to the cDNA molecules during amplification using well- or samplespecific barcoded primers (e.g., unique dual indexes or UDIs).
[0081] For applications involving the labeling or tagging of genomic DNA, the barcodes can be added, e.g., in an initial step in which the genomic DNA is fragmented and an adapter is appended to the fragment ends (e.g., using DNase-seq or ATAC-seq), in one or more subsequent steps in which nucleic acid tags are coupled to the genomic DNA fragments (e.g., via an adapter coupled to the fragment ends), or in one or more amplification steps (e.g., via a PCR primer).
[0082] Accordingly, the number of possible barcode combinations can vary by, e.g., increasing or decreasing the number of wells used for reverse transcription, for appending adapters to genomic DNA fragments, for ligation-based tagging, and / or for indexing during amplification, and / or by changing the number of total barcoding steps, e.g., by varying the number of rounds of split-pool tagging or by omitting barcodes in one or more steps (e.g., by performing RT and / or amplification using non-barcoded primers, or by omitting the ligation tagging steps and / or amplification indexing steps altogether).
[0083] In certain embodiments, steps of the present methods in which a nucleic acid tag is appended or coupled to a cDNA, genomic DNA fragment, or other polynucleotide within a cell or nucleus may be repeated one or more times, e.g., 1, 2, 3, 4, 5 times, or 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 20, 25, 30, 35, 40, 45, 50, 100, or more times. In certain other embodiments, the steps are repeated a sufficient number of times such that the tagged nucleic acid molecules within of each cell or nucleus would be likely to be bound to a unique barcode (e.g., unique among the plurality of cells or nuclei, or among multiple pluralities of cells or nuclei, e.g., in situations where multiple pluralities may be sequenced together). The number of times may be selected to provide a greater than 50% likelihood, greater than 90% likelihood, greater than 95% likelihood, greater than 99% likelihood, or some other probability that the cDNAs in each cell are bound to a unique barcode.
[0084] The number of total possible barcode combinations in the population will be a function of the number of barcode tagging rounds that are performed, and on the number of different barcodes / aliquots included in each round. The total number of possible barcode combinations can be achieved in any of a number of ways. In some embodiments, the number of total possible barcode combinations is greater than the number of different cells in the population, e.g., such that the probability that a given combination of barcodes is unique among all of the cells of the plurality is, e.g., 95%, 96%, 97%, 98%, 99%, or higher. Accordingly, thepresent methods could include, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, or more individual rounds of barcoding.
[0085] The barcodes introduced at any of the steps can be any length, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides or longer. In some embodiments, the barcodes are at least 8 nucleotides long. In some embodiments, the barcodes are 8 nucleotides long. In some embodiments, the barcodes are greater than 8 nucleotides long. The length of any given barcode used in a given embodiment of the method (e.g., a barcode sequence in a first nucleic acid tag) can be independent of the length of a different barcode used in the embodiment (e.g., a barcode in a second nucleic acid tag). For example, in some embodiments, a RT primer barcode, a barcode in an adapter appended to genomic DNA fragment ends, a barcode in a first nucleic acid tag, a barcode in a second nucleic acid tag, and / or a barcode or index sequence in a PCR primer may all be the same length, whereas in some embodiments they may all be different lengths, while in other embodiments some barcodes may be the same length while some are of different lengths.5. Cells and Nuclei
[0086] The present methods can be used to specifically label molecules, e.g., nucleic acids, in any of a wide variety of cell types. Cells suitable for use in the present methods include, e.g., primary cells, cell lines (e.g., HEK293, HEK293T, HEK293F, NIH3T3, Jurkat cells, or others), cells isolated from an organism, organoid, or a tissue, isolated blood cells, and others. In some embodiments, the cells are healthy cells (e.g., wild-type or control cells). In some embodiments, the cells are disease cells (e.g., cancer cells, infected cells). In some embodiments, the cells are stem cells. In various embodiments, the plurality of cells may be eukaryotic cells, vertebrate cells, mammalian cells, human cells, mouse cells, insect cells, plant cells, fungal cells, yeast cells, or bacterial cells.
[0087] In some embodiments, the cells comprise one or more cell types such as blood cells (e.g., peripheral blood mononuclear cells or PBMCs, immune cells such as T cells, B cells, NK cells), brain cells, liver cells, gut cells, bone marrow cells, pancreatic cells, epithelial cells, endothelial cells, neuronal cells, fibroblast cells, bone cells, muscle cells, skin cells, fat cells, lymphocytes, myeloid cells, macrophages, stem cells, and others. In some embodiments, the cells are all from a single source, i.e., from a single individual, organism, or tissue. In some embodiments, the cells are from multiple sources, i.e. from multiple individuals, organisms, or tissues. In some embodiments, the cells are autologous cells. In some embodiments, the cellsare allogeneic cells. In some embodiments, the cells are all or primarily comprise a single cell type (e.g., from a cell line, or a specific cell type isolated from a primary sample). In some embodiments, the cells comprise a mixture of different cell types. In some embodiments, the cells are adherent cells. In some embodiments, the cells are suspension cells.
[0088] In some embodiments, the cells have been previously frozen. In some embodiments, the cells have been previously fixed and frozen, e.g., the methods are performed using multiple samples that have been fixed and / or frozen at different times. In some embodiments, the cells have been previously fixed, permeabilized, and frozen.
[0089] It will be appreciated that in all of the herein-disclosed embodiments, either cells or nuclei obtained from the cells can be used. Nuclei can be prepared, e.g., using standard methods such as by douncing. In some embodiments, nuclei are prepared from frozen cells, tissue samples, or tissue slices or sections. Nuclei can be prepared, e.g., by placing the frozen sample (cells, tissue, minced tissue sample, slice, section, etc.) into a cooled nuclei isolation (NIM) buffer solution (e.g., NIM1 or NIM2 buffer), then transferred to a dounce and homogenized, e.g., using a pestle (e.g., 10 strokes each with a loose and with a tight pestle). The homogenate can then be filtered (e.g., using a 40 um or 70 um filter) and transferred to, e.g. conical tubes. The tube can then be centrifuged, e.g., 200x or 500x g in a pre-cooled swinging bucket centrifuge for, e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, or more minutes at a low temperature, e.g., 4 °C, or at 1 °C, 2 °C, 3 °C, 4 °C, 5 °C, 6 °C, 7 °C, or 8 °C, or at a temperature less than 8 °C, 7 °C, 6 °C, 5 °C, 4 °C, 3 °C, 2 °C, or 1 °C. In some embodiments, the cells are counted before and / or after centrifugation, e.g., to ensure an appropriate number of cells in each aliquot or tube. The pellets can then be resuspended in an appropriate solution or buffer, e.g., a nuclei buffer containing BSA (e.g., 0.75% BSA), and subsequently fixed and stored, e.g., at -80 °C.
[0090] The present methods allow the labeling and multiplex analysis of cells at a broad range of scales. For example, in some embodiments, the methods are used to label up to 10,000 cells (using, e.g., 12 samples in round 1 of the present methods, e.g., as exemplified in Example 1 below). In some embodiments, the methods are used to label up to 100,000 cells or nuclei (using, e.g., 48 samples in round 1 of the present methods, e.g., as exemplified in Example 1 below). In some embodiments, the methods are used to label up to 1,000,000 cells or nuclei (using, e.g., 96 samples in round 1 of the present methods, e.g., as exemplified in Example 1 below). In some embodiments, the methods are used to label up to 1,000,000 cells or nuclei,using up to 384 samples or conditions distributed into four 96-well plates (see, e.g., Example 4, below). In embodiments of the present methods, the numbers of cells or nuclei labeled and analyzed are at least 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,000,000, 1,500,000, 2,000,000, 3,000,000, 4,000,000, 5,000,000, 6,000,000, 7,000,000, 8,000,000, 9,000,000, 10,000,000, 50,000,000, 100,000,000, or more. .6. Fixation and permeabilization
[0091] In some embodiments, the cells (or nuclei) are fixed and / or permeabilized prior to barcoding such that the components are immobilized or held in place. In some embodiments, fixation is performed on single cell (or nuclei) suspensions. In some embodiments, the cells (or nuclei) were previously frozen and are thawed before fixation. In some embodiments, the cells (or nuclei) are used directly, e.g., from culture or after isolation from a biological source, without freezing. In some embodiments, the cells (or nuclei) are assessed for quality before fixation. Cells (or nuclei) can be counted (e.g., using a hemocytometer) and / or otherwise assessed in any of a number of standard ways, e.g., by staining with trypan blue, acridine orange, and / or propidium iodide. In some embodiments, the cell (or nuclei) suspensions show 70% or more viability, the suspensions show 5% and / or less aggregation / debris. In some embodiments, suspensions with at least about 100,000 cells or nuclei are used. In some embodiments, adherent cells are first dissociated using TrypLE Express Enzyme (IX), phenol red (Thermo Fisher Scientific).
[0092] In some embodiments, the present methods comprise providing a plurality of fixed and permeabilized cells. In some embodiments, the present methods comprise fixing and permeabilizing the plurality of cells or nuclei prior to, e.g., generating cDNA in the cells or nuclei by reverse transcribing RNA in the cells or nuclei. In some embodiments, the cells or nuclei may be fixed and permeabilized and frozen, e.g., at -80 °C, prior to, e.g., generating cDNA. In some embodiments, the cells or nuclei are fixed and permeabilized and then directly used in the present methods, i.e., without freezing and storing them.
[0093] The plurality of cells (or nuclei) may be fixed using any of a number of suitable reagents or conditions. For example, in some embodiments, the cells (or nuclei) can be fixed in formaldehyde in phosphate buffered saline (PBS) (e.g., in about 1-4% formaldehyde in PBS). In various embodiments, the plurality of cells (or nuclei) may be fixed using methanol (e.g., 100% methanol) at about -20° C. or at about 25° C. In various other embodiments, theplurality of cells (or nuclei) may be fixed using methanol (e.g., 100% methanol), at between about -20° C. and about 25° C. In yet various other embodiments, the plurality of cells (or nuclei) may be fixed using ethanol (e.g., about 70-100% ethanol) at about -20° C. or at room temperature. In yet various other embodiments, the plurality of cells (or nuclei) may be fixed using ethanol (e.g., about 70-100% ethanol) at between about -20° C. and room temperature. In still various other embodiments, the plurality of cells (or nuclei) may be fixed using acetic acid, for example, at about -20° C. In still various other embodiments, the plurality of cells (or nuclei) may be fixed using acetone, for example, at about -20° C. Other suitable methods of fixing the plurality of cells (or nuclei) are also within the scope of this disclosure.
[0094] In some embodiments, RNases are inactivated or eliminated before, during, and / or after fixation and / or permeabilization using, e.g., RNase decontamination products such as RNaseZap RNAse Decontamination Solution (Thermo Fisher). In some embodiments, BSA (e.g., 5-10%, or 7.5%) is added to the cells or nuclei, e.g., to prevent aggregation.
[0095] In some embodiments, the methods may include fixing and / or permeabilizing the cells or nuclei at a temperature below about 8 °C, below about 7 °C, below about 6 °C, below about 5 °C, below about 4 °C, below about 3 °C, below about 2 °C, below about 1 °C, below about 0 °C, below about -5 °C, below about -10 °C, at about 8 °C, at about 7 °C, at about 6 °C, at about 5 °C, at about 4 °C, at about 3 °C, at about 2 °C, at about 1 °C, or at another suitable temperature. In some embodiments, the cells or nuclei are fixed and / or permeabilized at a temperature of below about 8, 7, 6, 5, 4, 3, 2, 1, 0, -1, -2, -3, or -4 °C, between about -4 to 8, - 4 to 0, 0 to 4, 4 to 8, or 0 to 8 °C, or at about 8, 7, 6, 5, 4, 3, 2, 1, 0, -1, -2, -3, or -4 °C. In some embodiments, the cells or nuclei are fixed and / or permeabilized on ice.
[0096] In some embodiments, the cells are adherent cells (i.e., cells that are adhered to a plate, e.g., adherent mammalian cells). In some such embodiments, adherent cells are fixed, permeabilized, and / or undergo reverse transcription, followed by trypsinization to detach the cells from a surface. Alternatively, the adherent cells may be detached prior to the separation and / or tagging steps. In some other embodiments, the adherent cells may be trypsinized prior to the fixing and / or permeabilizing steps.
[0097] Permeabilization of the cells (or nuclei) can be achieved in any of a number of ways. For example, a detergent or surfactant such as TRITON™ X-100 may be added to the plurality of cells (or nuclei), followed by the optional addition of HC1. In some such embodiments, about 0.2% TRITON™ X-100 is added to the plurality of cells (or nuclei),followed by the addition of about 0.1 N HC1. In some embodiments, the plurality of cells (or nuclei) is permeabilized using ethanol (e.g., about 70% ethanol), methanol (e.g., about 100% methanol), Tween 20 (e.g., about 0.2% Tween 20), and / or NP -40 (e.g., about 0.1% NP-40).
[0098] Exemplary methods of fixing and permeabilizing cells and nuclei in various quantities and concentrations (e.g., 10,000-100,000 cells or nuclei, 100,000-1,000,000 cells or nuclei, greater than 1,000,000 cells or nuclei, in, e.g., 1, 2, 12, 48, 96, 384, or more samples), in plates or in tubes, are described, e.g., in Examples 5 and 6, below.7. Reverse transcription
[0099] In some embodiments, reverse transcription is conducted or performed on the plurality of cells or nuclei. In certain embodiments, reverse transcription may be conducted on a fixed and / or permeabilized plurality of T cells (or nuclei). In some embodiments, variants of M-MuLV reverse transcriptase may be used in the reverse transcription. However, any suitable method of reverse transcription is within the scope of this disclosure. An exemplary illustration of reverse transcription according to the present methods is shown in FIG. 2A.
[0100] In some embodiments, the reverse transcription primers may be configured to reverse transcribe all, or substantially all, RNA in a cell (e.g., a random hexamer with a 5' overhang). In some other embodiments, the reverse transcription primers may be configured to reverse transcribe RNA having a poly(A) tail (e.g., a poly(dT) primer, such as a dT(15) primer or anchored dT(15) primer, with a 5' overhang). In some embodiments, the reverse transcription primers are configured to reverse transcribe both all, or substantially all, RNA in a cell, as well as to reverse transcribe polyadenylated RNA. In some embodiments, reverse transcription primers may be included that are configured to reverse transcribe predetermined RNAs.
[0101] In some embodiments, a portion of a reverse transcription (RT) primer that is configured to bind to RNA and / or initiate reverse transcription may comprise one or more of the following: a random hexamer, random septamer, an octamer, a nonamer, a decamer, or a poly(T) (or polydT) stretch of nucleotides (e.g., comprising 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more consecutive thymine bases). In some embodiments, the poly(T) sequence is anchored, i.e., comprises a base other than T at the 3’ end (e.g., a mixture of anchored primers is used each comprising an A, C, or G at the 3’ end of the poly(T) sequence). In some embodiments, RT primers are used that comprise well-specific barcodes, as described in more detail elsewhere herein.
[0102] In some embodiments, primers comprising a poly(T) sequence are used in the reverse transcription in the absence of primers with random sequences (such as random hexamers). In some embodiments, primers comprising random sequences (e.g., random hexamers) are used in the absence of primers with poly(T) sequences. In some embodiments, the reverse transcription is performed with both primers comprising a random sequence (e.g., random hexamer) and primers comprising a poly(T) sequence (e.g., with 15 consecutive thymidine residues).
[0103] In some embodiments, each of the RT primers comprises a 5’ overhang comprising a 5’ overhang sequence located 5’ of the poly(T) or the random nucleotide sequence, wherein the 5’ overhang sequence is the same in all of the RT primers used in the first plurality of aliquots, and wherein following reverse transcription of the RNA, the 5’ overhang sequence is present at the 5’ end of each of the cDNA molecules. In some embodiments, one or more of the RT primers comprises a sequence as shown in Table 6 or Table 7, as shown in any of SEQ ID NOS: 401-590, or a subsequence thereof (e.g., a subsequence corresponding to an RT barcode or an RT 5’ overhang sequence, as indicated, e.g., in Tables 6 and 7).
[0104] The RT primers can be used at any of a range of suitable concentrations. For example, in some embodiments, the concentration of each RT primer (e.g., a polydT primer, a random hexamer primer, or a target-specific primer) is between about 0.5 pM and about 10 pM. In some embodiments, the concentrations are each between about 1 pM and about 7 pM, between about 1.5 pM and about 4 pM, between about 2 pM and about 3 pM, about 2.5 pM, or another suitable concentration.
[0105] In some embodiments, each of the RT primers comprises a barcode (i.e., a specific barcode sequence) (i.e., an “RT primer barcode”). In some embodiments, the reverse transcription is performed on a population of cells (or nuclei) distributed in a plurality of aliquots or wells, and the RT primer barcodes are aliquot- or well-specific. In some such embodiments, all of the RT primer barcodes used in a given aliquot or well are the same, and a different RT primer barcode is used in each of the aliquots or wells. Stated another way, a first barcode may be added to the cDNA molecules in a first specific container, mixture, reaction, receptacle, sample, well, or vessel (e.g., specific to the given container, mixture, reaction, receptacle, sample, well, or vessel), and a second barcode sequence may be added to the cDNA molecules in a second container, mixture, reaction, receptacle, sample, well, or vessel (and the same for a third, fourth, etc. barcode sequence and third, fourth, etc. container,mixture, reaction, receptacle, sample, well, or vessel). For example, in some embodiments, 48 sets of different well-specific RT primers are used (e.g., in a 48-well plate).
[0106] Accordingly, if there are 48 samples (e.g., cells, tissues, etc.), each sample can get a unique well-specific barcode. However, if there are only four samples, each sample could have 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, or more different sets of well-specific RT primers. A user can know which barcodes correspond to each sample, so the user can recover sample identities. Other numbers of specific barcodes (or well-specific RT primers) are also within the scope of this disclosure. Such a configuration may allow or provide for the multiplexing of the method. In general, any distribution of barcode sequences that can provide some information about the aliquot or well in which a given cell was present can be used. For example, even if more than one barcode sequence is used in one or more aliquots or wells, or if a given barcode sequence is used in more than one aliquot or well, the methods are encompassed by the present disclosure.
[0107] The barcode sequences present within the RT primers can have any of a range of lengths, e.g., 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides or longer. In some embodiments, the RT primer barcode sequences are 8 nucleotides in length. By varying 8 nucleotides, there are 65,536 possible unique sequences. In some embodiments, the RT primer barcodes comprise more than 8 nucleotides. In some other embodiments, the RT primers comprise fewer than 8 nucleotides.
[0108] The plurality of cells can be divided prior to reverse transcription into any of a number of aliquots or wells, and using any of a number of suitable reaction vessels or containers. For example, the plurality of cells or nuclei can be distributed into individual tubes or containers, or into a plurality of wells in a multi-well plate. Any multi-well plate can be used, e.g., 4, 6, 8, 12, 24, 48, 96, 384, or 1536 well plates. In some embodiments, the plurality of cells or nuclei are distributed into one or more 96-well plates. In some embodiments, all 96 wells of a 96-well plate are used (i.e., the plurality of cells or nuclei is divided into 96 aliquots). In some embodiments, a fraction of the wells on the plate are used, e.g., the cells or nuclei are distributed into, e.g., 8, 16, 24, 32, 40, 48, 56, 64, 72, 80, 88, 12, 24, 36, 48, 60, 72, or 84 wells of a 96-well plate. In some embodiments, up to 12 wells are used for up to 10,000 cells or nuclei, or up to 48 wells are used for up to 100,000 cells or nuclei, or up to 96 wells are used for up to 1,000,000 cells or nuclei. In some embodiments, some or all of the different wells (e.g., the 48 wells shown in FIG. 2A) can correspond to different biological samples orexperimental conditions. In other embodiments, the wells contain cells originating from the same samples or experimental conditions.
[0109] Suitable methods of reverse transcribing RNA within cells for use in the herein- described methods are described, e.g., in US Patent Nos. 10,900,065, 11,634,751, 11,168,355, 11,427,856, 11,639,519, 12,043,864, 12,247,247, 12,180,536, 12,227,793, 12,247,248, 11,680,283, 12,234,501, 10,633,648, 11,421,221, 12,163,189, PCT Application Nos. PCT / US2019 / 057939, PCT / US24 / 14893, PCT / US24 / 61803, in Rosenberg et al., Science 360, 176-182 (2018), Rosenberg etal., BioRxiv (2017), “Scaling single cell transcriptomics through split pool barcoding,” doi.org / 10.1101 / 105163, Tran et al. BioRxiv (2022) “High sensitivity single cell RNA sequencing with split pool barcoding,” doi.org / 10.1101 / 2022.08.27.505512, the entire disclosures of all of which are herein incorporated by reference (including all supplemental material).8. Labeling with nucleic acid tags
[0110] As described above, in some embodiments of the present disclosure a plurality of cells or nuclei is subjected to combinatorial barcoding, such that molecules (e.g., cDNAs synthesized in cells by reverse transcription) are labeled with barcodes that when viewed in combination provide a cell- or nucleus-specific label for the labeled molecules. In such labeling steps, a plurality of cells is split into multiple aliquots or wells, labeled with well-specific barcodes (e.g., during reverse transcription using barcoded primers and / or by ligating barcoded nucleic acid tags), and repooled. This cycle of dividing the plurality of cells, well-specific labeling, and repooling can be repeated any number of times, with each round adding more tags to the cDNAs and thereby creating a set of nucleic acid tags that together can act as, e.g., a cell-specific (or nucleus-specific) barcode. As more and more rounds are added, the number of paths that a cell can take increases and consequently the number of possible barcodes that can be created also increases. Given enough rounds and divisions, the number of possible barcodes will be much higher than the number of cells, resulting in a high likelihood that each cell (or nucleus) in the population has a unique barcode. For example, if the division took place in a 96-well plate, after 4 divisions there would be 964=84,934,656 possible barcodes. In particular embodiments of the present methods, enough rounds of labeling is performed such that the number of possible barcodes is 50%, 60%, 70%, 80%, 90%, 100%, 110%, 120%, 130%, 140%, 150%, 200%, 300%, 400%, 500%, 600%, 700%, 800%, 900%, 1000%, or lx, 2x, 3x, 4x, 5x, 6x, 7x, 8x, 9x, lOx, 20x, 30x, 40x, 50x, 60x, 70x, 80x, 90x, lOOx, or more ofthe number of cells (or nuclei) in the plurality of cells (or nuclei). In particular embodiments, enough rounds of labeling are performed such that the likelihood that the tagged molecules within a given cell or nucleus have a unique barcode combination among the plurality of cells is about 95%, 96%, 97%, 98%, 99%, 99.1%, 99.2%, 99.3%, 99.4%, 99.5%, 99.6%, 99.7%, 99.8%, 99.9%, or greater.[OHl] In each round of tagging, the plurality of cells or nuclei can be divided into any of a number of aliquots or wells, and using any of a number of suitable reaction vessels or containers. For example, the plurality of cells or nuclei can be distributed into individual tubes or containers, or into a plurality of wells in a multi-well plate. Any multi-well plate can be used, e.g., 4, 6, 8, 12, 24, 48, 96, 384, or 1536 well plates. In particular embodiments, the plurality of cells or nuclei are distributed into one or more 96-well plates. In some embodiments, all 96 wells of a 96-well plate are used (i.e., the plurality of cells or nuclei is divided into 96 aliquots). In some embodiments, a fraction of the wells on the plate are used, e.g., the cells or nuclei are distributed into, e.g., 8, 16, 24, 32, 40, 48, 56, 64, 72, 80, 88, 12, 24, 36, 48, 60, 72, or 84 wells of a 96-well plate. In some embodiments, 96 distinct barcode sequences are present among the nucleic acid tags used in the plurality of aliquots. In some embodiments, the plurality of aliquots comprises 96 aliquots distributed in a 96-well plate. In some embodiments, each of the 96 distinct barcode sequences are present in only one of the 96 aliquots.
[0112] FIGS. 2A-2D illustrate an exemplary embodiment of labelling or tagging nucleic tags. As shown, during reverse transcription of the mRNA molecules (FIG. 2A), Poly T and random hexamer primers anneal to mRNA within single cells. Each primer contains a barcode and optionally a 5’ overhang comprising a 5’ overhang sequence. Reverse transcriptase extends cDNA to form a cDNA / mRNA hybrid comprising the barcode (BC1), for example, a wellspecific barcode. The cells may be pooled and split and redistributed into individual wells. In subsequent steps, nucleic acid tags can be appended to the cDNA / mRNA hybrid via the 5’ overhang (FIG. 2B). The nucleic acid tags can contain a second barcode (e.g., BC2) and optionally a second DNA linker. The cells can be pooled and split and redistributed into individual wells. A second nucleic acid tag can be ligated to the growing cDNA (see, e.g., FIG. 2C). The adaptor contains a third barcode, an Illumina adaptor sequence (e.g., R2) or a compatible equivalent thereof, and a biotin molecule.
[0113] Suitable methods of combinatorial barcoding are described, e.g., in US Patent Nos. 10,900,065, 11,634,751, 11,168,355, 11,427,856, 11,555,216, 11,639,519, 12,195,786,12,043,864, 12,247,247, 11,987,838, 12,180,536, 12,227,793, 12,247,248, 11,680,283,12,234,501, 10,633,648, 11,421,221, 12,163,189, PCT Application Nos.PCT / US2019 / 057939, PCT / US24 / 14893, PCT / US24 / 61803, in Rosenberg et al., Science 360, 176-182 (2018), Rosenberg etal., BioRxiv (2017), “Scaling single cell transcriptomics through split pool barcoding,” doi.org / 10.1101 / 105163, Tran et al. BioRxiv (2022) “High sensitivity single cell RNA sequencing with split pool barcoding,” doi.org / 10.1101 / 2022.08.27.505512, the entire disclosures of all of which are herein incorporated by reference (including all supplemental material).
[0114] In some embodiments, combinatorial barcoding is accomplished by coupling the cDNA molecules generated in each cell during reverse transcription (or genomic DNA fragments or adapters appended to other molecules as described elsewhere herein) with a nucleic acid tag, wherein each nucleic acid tag comprises a barcode sequence, e.g., a wellspecific barcode sequence or tag barcode. In particular embodiments, coupling the cDNA molecule with the nucleic tag comprises ligating the nucleic tag to the cDNA molecule. In some embodiments, each nucleic acid tag comprises a first strand comprising the barcode sequence, and further comprises a 3’ and / or 5’ region located 3’ and / or 5’ of the barcode. In some embodiments, the first strand comprises a 3' hybridization sequence extending from a 3' end of a labeling (i.e., barcode) sequence and / or a 5' hybridization sequence extending from a 5' end of the labeling (i.e., barcode) sequence. Each nucleic acid tag may also comprise a second (linker) strand including an overhang sequence. The overhang sequence may include (i) a first portion complementary to a 5' hybridization sequence of a different nucleic acid tag (e.g., a nucleic acid tag appended in a previous round of tagging) or to a 5' overhang sequence of an RT primer. The second (linker) strand may also comprise a sequence complementary to the 3' hybridization sequence of the first strand. In some embodiments, the first and second strands are preannealed before being added to the wells or aliquots containing the cells or nuclei.
[0115] In some embodiments, the nucleic acid tags comprise i) a tag barcode sequence, and ii) a 3’ hybridization sequence located 3’ of the barcode sequence and / or a 5’ hybridization sequence located 5’ of the barcode sequence, wherein multiple distinct tag barcode sequences are present among the nucleic acid tags used in the second plurality of aliquots, and wherein the tag barcode sequences present in each individual aliquot of the second plurality of aliquots are specific to the individual aliquot.
[0116] In some embodiments, the 3’ end of the nucleic acid tag is present within the 3’ hybridization sequence, and the 3’ end of the nucleic acid tag is brought into proximity of the 5’ end of the cDNA molecule by being preannealed to a linker nucleic acid strand that is complementary to the 3’ hybridization sequence of the nucleic acid tag and to the 5’ overhang sequence of the RT primer, or to the 3’ hybridization sequence of the nucleic acid tag and to the 5’ hybridization sequence of a previously coupled nucleic acid tag.
[0117] In some embodiments, the nucleic acid tags comprise a) a first strand comprising: i) a first strand barcode sequence; ii) a first strand 5’ hybridization sequence located 5’ of the first strand barcode sequence; and / or iii) a first strand 3’ hybridization sequence located 3’ of the first strand barcode sequence; and b) a second strand comprising: i) a second strand barcode sequence, wherein the second strand barcode sequence is complementary to the first strand barcode sequence; ii) a second strand hybridization sequence located 5’ of the second strand barcode sequence, wherein the second strand hybridization sequence is complementary to the first strand 3’ hybridization sequence; and / or iii) a second strand overhang sequence located 5’ of the second strand hybridization sequence; wherein the first strand and second strand are annealed such that the nucleic acid tag comprises: i) a double-stranded central region comprising the first strand barcode sequence annealed to the second strand barcode sequence and the first strand 3’ hybridization sequence annealed to the second strand hybridization sequence; and ii) a single-stranded overhang located at each end of the nucleic acid, wherein one of the two overhangs comprises the first strand 5’ hybridization sequence, and the other overhang comprises the second strand overhang sequence.
[0118] In some embodiments, the disclosure provides a nucleic acid tag comprising a first strand and a second strand, wherein the first strand comprises: a first strand barcode sequence, and a first strand 3’ hybridization sequence located 3’ of the first strand barcode sequence, and wherein the second strand comprises: a second strand barcode sequence, a second strand hybridization sequence located 5’ of the second strand barcode sequence, and a second strand overhang sequence located 5’ of the second strand hybridization sequence.
[0119] In some embodiments, the nucleic acid tags comprise a first strand comprising a hybridization sequence element as shown in Table 1 or as any of SEQ ID NOS: 1-4, or comprising a sequence as shown in Table 2 or Table 4 or as any of SEQ ID NOS: 11-105 or 201-295, and / or a second strand comprising a hybridization sequence element as shown inTable 1 or as any of SEQ ID NOS: 4-8, or comprising a sequence as shown in Table 3 or Table 5 or as any one of SEQ ID NOS: 106-200 or 296-400.
[0120] Table 2 shows exemplary round 2 nucleic acid tag first strand sequences. In the sequences shown in Table 2, the six 3’ terminal nucleotides (i.e., TCCAAC) represent the First strand 3’ hybridization sequence, the six 5’ terminal nucleotides (i.e., GGTCAG) represent the First strand 5’ hybridization sequence, and the eight nucleotides between the first strand 3’ and 5’ hybridization sequences represent the barcode sequence (e.g., tag barcode sequence, or first strand barcode sequence).
[0121] Table 3 shows exemplary round 2 nucleic acid second strand sequences. In the sequences shown in Table 3, the six 5’ terminal nucleotides (i.e., GAGGTG) represent the second strand overhang sequence, the six nucleotides located immediately 3’ of the second strand overhang sequence (i.e., GTTGGA) represent the second strand hybridization sequence, and the eight 3’ terminal nucleotides represent the barcode sequence (e.g., tag barcode sequence, or second strand barcode sequence).
[0122] Table 4 shows exemplary Round 3 nucleic acid tag first strand sequences. In the sequences shown in Table 4, the six 3’ terminal nucleotides (i.e., ATGAGG) represent the first strand 3’ hybridization sequence, the eight nucleotides immediately 5’ of the first strand 3’ hybridization sequence represents the barcode sequence (e.g., tag barcode sequence, or first strand barcode sequence), and the “N”s correspond to random sequences, used, e.g., to permit detection of PCR duplicates. The primers shown also include a 5’ biotin moiety.
[0123] Table 5 shows exemplary Round 3 nuclei acid tag second strand sequences. In the sequences shown in Table 5, the six 5’ terminal nucleotides (i.e., CTGACC) represent the second strand overhang sequence, the six nucleotides immediately 3’ of the second strandoverhang sequence (i.e., CCTCAT) represents the second strand hybridization sequence, and the eight 3’ terminal nucleotides represent the barcode sequence (e.g., tag barcode sequence, or second strand barcode sequence).
[0124] Table 6 shown exemplary round 1 reverse transcription (RT) primer sequences. In the sequences shown in Table 6, the 15 3’ terminal nucleotides represent the poly(dT) sequence, the eight nucleotides 5’ of the poly(dT) sequence represents the RT barcode sequence, and the six 5’ terminal nucleotides (i.e., CACCTC) represent the 5’ overhang sequence, which may be bound in some embodiments by the second strand overhang of a nucleic acid tag, e.g., R2 nucleic acid tag (see, e.g. FIGS. 6B-6C).
[0125] Table 7 shows exemplary random hexamer RT primer (Round 1) sequences. In the sequences shown in Table 7, the 15 3’ terminal nucleotides represent the poly(dT) sequence, the eight nucleotides 5’ of the poly(dT) sequence represents the RT barcode sequence, and the six 5’ terminal nucleotides (i.e., CACCTC) represent the 5’ overhang sequence, which may be bound in some embodiments by the second strand overhang of a nucleic acid tag, e.g., R2 nucleic acid tag (see, e.g. FIGS. 6B-6C).
[0126] The barcode sequences present within the nucleic acid tags can be any of range of lengths, e.g., 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides or longer. In some embodiments, the barcode sequences are 8 nucleotides in length. By varying 8 nucleotides, there are 65,536 possible unique sequences. In some embodiments, more than 8 nucleotides are used. In some other embodiments, fewer than 8 nucleotides are used. In some embodiments, the first strand of the nucleic acid tag is preannealed to the second (linker) strand. In some embodiments, the linker strand includes sequence complementary to part of the RT primer or to a 5’ region in a previously coupled nucleic acid tag (i.e., appended to the cDNA in a previous round of tagging), thereby allowing it to hybridize and bring the 3' end of the barcodes into close proximity to the 5' end of the reverse transcription primer or the previously added tag. In some embodiments, the phosphate of the reverse transcription primer or previous tag can is ligated to the 3' end of the first-round barcodes by, e.g., T4 DNA ligase. A 5’ region of the nucleic acid tag (e.g., the first strand 5’ hybridization sequence as shown in FIGS. 6A- 6C) can then provide an accessible binding domain for a linker oligo to be used in another round of barcoding (e.g., via the second strand overhang sequence as shown in FIGS. 6A-6C). The nucleic acid tags (barcode oligos) can include a 5' phosphate that can allow ligation to the 3' end of another oligo by T4 DNA ligase.
[0127] In particular embodiments, the nucleic acid tags are ligated to the cDNAs (or adapter molecules, e.g., appended to genomic DNA fragments) during each round of labeling. In some embodiments, the methods of labeling nucleic acids in the first cell may comprise ligating at least two of the nucleic acid tags that are bound to the cDNAs (or genomic DNA fragments). In some embodiments, the nucleic acid tags are hybridized to the cDNAs (or adapters) during each round, and ligation is performed for all of the hybridized tags at a later stage, e.g., following cell lysis, i.e., ligation may be conducted before or after the lysing and / orthe cDNA purification steps. Ligation can comprise covalently linking the 5' phosphate sequences on the nucleic acid tags to the 3' end of an adjacent strand or nucleic acid tag such that individual tags are formed into a continuous, or substantially continuous, barcode sequence that is bound to the 3' end of the cDNA or genomic DNA fragment sequence. In various embodiments, a double-stranded DNA or RNA ligase may be used with an additional linker strand that is configured to hold a nucleic acid tag together with an adjacent nucleic acid in a “nicked” double-stranded conformation. The double-stranded DNA or RNA ligase can then be used to seal the “nick.” In various other embodiments, a single-stranded DNA or RNA ligase may be used without an additional linker.
[0128] In some embodiments, following the ligation of the nucleic acid tags during each round, one or more unbound nucleic acid tags are removed (e.g., by washing the plurality of cells). For example, the methods may comprise removing a portion, a majority, or substantially all of the unbound nucleic acid tags. Unbound nucleic acid tags may be removed such that further rounds of the disclosed methods are not contaminated with one or more unbound nucleic acid tags from a previous round of a given method. In some embodiments, unbound nucleic acid tags may be removed via centrifugation. For example, the plurality of cells can be centrifuged such that a pellet of cells is formed at the bottom of a centrifuge tube. The supernatant (i.e., liquid containing the unbound nucleic acid tags) can be removed from the centrifuged cells. The cells may then be resuspended in a buffer (e.g., a fresh buffer that is free or substantially free of unbound nucleic acid tags). In another example, the plurality of cells may be coupled or linked to magnetic beads that are coated with an antibody that is configured to bind the cell or nuclear membrane. The plurality of cells can then be pelleted using a magnet to draw them to one side of the reaction vessel. In some other embodiments, the plurality of cells may be placed in a cell strainer (e.g., a PLURISTRAINER® cell strainer) and washed with a wash buffer. For example, the plurality of cells may remain in the cell strainer while the wash buffer passes through the cell strainer. Wash buffer may include, e.g., a surfactant, a detergent, and / or about 5-60% formamide.
[0129] In some embodiments, the ligation can be stopped during each round of combinatorial labeling by adding an excess of oligo that is complementary to all or part of the linker (second) strand used during the same round of labeling. To stop each barcode ligation, oligo strands that are fully complementary to the linker oligos can be added. These oligos can bind the linker strands attached to unligated barcodes and displace the unligated barcodes through a strand displacement reaction. The unligated barcodes can then be completely single-stranded. As T4 DNA ligase, for example, is unable to ligate single-stranded DNA to other single-stranded DNA, the ligation reaction will stop progressing. In particular embodiments, to ensure that all linker oligos are bound by the complementary (i.e., stop) oligos, a molar excess of the stop oligos (relative to the linker oligos) is added. In some embodiments, stop ligation strands are diluted into 10X Ligase Buffer and water, e.g., 264 pl stop ligation strand, 300 pl 10X T4 DNA Ligase Buffer, and 636 pl nuclease-free water. In some embodiments, the stop oligo used comprises a sequence as shown as SEQ ID NO: 9 or SEQ ID NO: 10.
[0130] In some embodiments, the nucleic acid tag (e.g., a final nucleic acid tag appended during the last of multiple rounds of tagging) may comprise a capture agent such as, but not limited to, biotin, e.g., a 5' biotin. A cDNA labeled with a 5' biotin-comprising nucleic acid tag may allow or permit the attachment or coupling of the cDNA to a streptavidin-coated magnetic bead, e.g., Cl beads. In some other embodiments, a plurality of beads may be coated with a capture strand (i.e., a nucleic acid sequence) that is configured to hybridize to a final sequence overhang of a barcode. In yet some other embodiments, cDNA may be purified or isolated by use of a commercially available kit (e.g., an RNEASY™ kit).
[0131] In some embodiments, one or more nucleic acid tags may comprise additional elements (in addition to biotin or another capture agent) such as a nucleotide sequence (e.g., comprising random and / or degenerate nucleotides) allowing the detection of PCR duplicates, primer binding sequences, adapter sequences for next-generation sequencing (NGS) (e.g., Illumina adapter sequences), and others. The random nucleotide sequences allow the computational removal of PCR duplicates, since these duplicates will have the same random sequence. In this way, each original transcript will only be counted once, even if multiple PCR duplicates are sequenced. Such sequences can contain any number of random successive nucleotides, e.g., 2, 3, 4, 5, 6, 7, 8, 9, 10 nucleotides or more, and can be used either alone or in conjunction with other sequence elements (such as the other barcode sequences described herein) to allow the identification of the PCR duplicates. The number of different random sequences per cell indicates how many unique RNA molecules can be detected per cell, which is directly related to how efficiently the molecules can be barcoded and processed to enable detection by next generation sequencing. In some embodiments, a plurality of distinct barcodes are used in each aliquot in order to allow the removal of PCR duplicates, as described elsewhere herein.9. Lysis
[0132] In some embodiments, the methods include lysing (i.e., breaking down the cell or nuclear structure) the plurality of cells (or nuclei) to release the tagged cDNA molecules from the plurality of cells or nuclei following combinatorial barcoding, thereby forming a lysate comprising the released tagged cDNA or genomic DNA fragment molecules. In some embodiments, the cells or nuclei comprising the tagged cDNAs (or other molecules) are divided into one or more samples or sublibraries, and the lysis is performed separately on each sample or sublibrary. In some embodiments, as described below, each sample or sublibrary can be tagged with one or more index or barcode sequences during a subsequent step of the herein- described methods, (e.g., using unique dual indices, or UDIs, as described elsewhere herein).
[0133] The total number of sublibraries prepared can depend on various factors, including the number of cells in the plurality of cells or nuclei. For example, in some embodiments, the plurality of cells or nuclei comprises up to 10,000 cells or nuclei, and two sublibraries are prepared. In some embodiments, the plurality of cells or nuclei comprises up to 100,000 cells or nuclei, and 8 sublibraries are prepared. In some embodiments, the plurality of cells or nuclei comprises up to 1,000,000 cells or nuclei, and 16 sublibraries are prepared. Other numbers of cells or nuclei, including less than 10,000 and greater than 1,000,000, or various numbers between 10,000 and 1,000,000, can be used, and a skilled artisan will be able to determine a suitable number of sublibraries.
[0134] Further, different numbers of cells or nuclei can be added to each sublibrary as desired. Sublibraries with small cell or nuclei numbers (e.g., 200-500) will be easier to sequence to saturation, and can serve, e.g., as a good quality control (QC) measure before sequencing additional sublibraries with much larger cell numbers. In some embodiments, the number of cells or nuclei in each library can be, for example, 200, 500, 1,000, 2,000, 3,000, 4,000, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, 12,000, 25,000, 62,000, or more.
[0135] Once prepared, each sublibrary can be immediately processed, e.g., to prepare a sequencing library, or stored, e.g., at -80 °C. Further, while all of the sublibraries can be processed together, each sublibrary can be sequenced separately.
[0136] In some embodiments, the plurality of cells is lysed in a lysis solution (e.g., a solution comprising Tris-HCl, EDTA, NaCl, and SDS, e.g., 10 mM Tris-HCl (pH 7.9), 50 mM EDTA (pH 7.9), 0.2 M NaCl, 2.2% SDS) comprising an RNase inhibitor (e.g., 0.5 mg / mlANTI-RNase, AMBION®) and a protease such as a serine protease, e.g., Proteinase K (e.g., 1000 mg / ml proteinase K (AMBION®)). The skilled artisan will be able to identify suitable conditions for cell (or nuclear) lysis. In some embodiments, lysis is performed at about 55 °C for about 3 hours with shaking (e.g., vigorous shaking). In some other embodiments, the plurality of cells is lysed using ultrasonication and / or by being passed through an 18-25 gauge syringe needle at least once. In yet some other embodiments, the plurality of cells is lysed by being heated to about 70-90 °C. For example, the plurality of cells may be lysed by being heated to about 70-90 °C. for about one or more hours.10. cDNA (or genomic DNA fragment) isolation
[0137] Following cell or nuclear lysis, the cDNA molecules may be isolated from the lysed cells or nuclei. In some embodiments, RNase H (or another RNase) may be added to the cDNA to remove RNA. The methods may further comprise ligating at least two of the nucleic acid tags that are bound to the released cDNA s (i.e., in embodiments in which the tags were not ligated during each round of split-pool tagging). In some such embodiments, the methods may comprise ligating at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more nucleic acid tags that are bound to the cDNA molecules.
[0138] In some embodiments, the tagged cDNA molecules are isolated from the lysis solution with a purification or clean-up step, e.g. an SPRI bead cleanup, before binding the desired nucleic acids (i.e., cDNAs containing 5' biotin) to streptavidin beads. In some embodiments, a protease inhibitor is added to lysates and then streptavidin beads are directly added (i.e., skipping the first SPRI isolation of nucleic acids). The protease inhibitor may include phenylmethanesulfonyl fluoride (PMSF), 4-(2-aminoethyl)benzenesulfonyl fluoride hydrochloride (AEBSF), a combination thereof, and / or another suitable protease inhibitor.
[0139] In some embodiments, the cDNA molecules (or tagged genomic DNA fragments) are isolated using Streptavidin beads, e.g., Cl beads. For example, 20 pl of resuspended DYNABEADS® MYONE™ Streptavidin Cl beads (for each aliquot of cells) can be added to a 1.7 ml microcentrifuge tube (EPPENDORF®). The beads can be washed, e.g., 3 times, with, e.g., l x phosphate buffered saline Tween 20 (PBST) and resuspended in PBST (e.g., 20 pl PBST). In some embodiments, 900 pl PBST is added to the cell aliquot and 20 pl of washed Cl beads are added to the aliquot of lysed cells. The samples can, e.g., be placed on a gentle roller for 15 minutes at room temperature and then washed, e.g., 3 times with 800 pl PBSTusing a magnetic tube rack (EPPENDORF®). The beads can then be resuspended in, e.g., PBS such as 100 pl PBS.
[0140] In some embodiments, a microcentrifuge tube (EPPENDORF®) comprising a sample can be placed against a magnetic tube rack (EPPENDORF®) for, e.g., 2 minutes and then the liquid can be aspirated. In some embodiments, the beads can be resuspended in, e.g., an RNase solution (3 pl RNase Mix (ROCHE™), 1 pl RNase H (NEW ENGLAND BIOLABS®), 5 pl RNase H 10* Buffer (NEW ENGLAND BIOLABS®), and 41 pl nuclease- free water). The sample can be incubated under suitable conditions, e.g., at 37 °C for 1 hour, and then removed from the conditions and placed against a magnetic tube rack (EPPENDORF®) for, e.g., 2 minutes. The sample can be washed with, e.g., 750 pl of nuclease- free water+0.01% Tween 20 (H2O-T), without resuspending the beads and keeping the tube disposed against the magnetic tube rack. The liquid can then be aspirated. The sample can be washed with 750 pl H2O-T without resuspending the beads and while keeping the tube disposed against the magnetic tube rack. Next, the liquid can be aspirated while keeping the tube disposed against the magnetic tube rack. The tube can then be removed from the magnetic tube rack and the sample can be resuspended in 40 pl of nuclease-free water.
[0141] FIGS. 3A-3B illustrate an exemplary embodiment of isolating cDNA molecules. After cell lysis, the biotinylated cDNA / mRNA hybrid binds to a streptavidin binder bead (FIG. 3A). Molecules having biotin are collected and molecules lacking biotin are removed. Next, a template switch reaction is performed (FIG. 3B) using a template switching oligonucleotide (TSO) to append a template switching (TS) adapter comprising, e.g., a primer binding site (e.g., a template switching primer “TS primer”) to the 3’ end of the cDNA molecule for cDNA amplification.11. Second strand synthesis
[0142] In particular embodiments, to facilitate subsequent amplification, a common adapter sequence (or NGS adapter) is added to the 3 '-end of the released cDNA molecules following isolation of the released cDNA (i.e., cDNA / mRNA duplex). In particular embodiments, the common (or NGS) adapter sequence is the same, or substantially the same, for each of the cDNA molecules (i.e., within a given experiment). The addition of the common adapter may be conducted or performed in a solution including up to about 10% w / v of PEG, wherein the molecular weight of the PEG is between about 7,000 g / mol and 9,000 g / mol. In some embodiments, to prevent concatemers of the adapter oligo, dideoxy cytidine (ddC) can beincluded at the 3' end of the adapter oligo. In some embodiments, adapters are used with a phosphate at the 5' end and ddC at the 3' end. Several enzymes are capable of ligating singlestranded oligo to the 3' end of single-stranded DNA, e.g., T4 RNA ligase 1 (NEW ENGLAND BIOLABS®) or thermostable 5' AppDNA / RNA Ligase (NEW ENGLAND BIOLABS®).
[0143] In some embodiments, the adapter sequence is added to the 3 '-end of the released cDNA molecules by template switching (see, e.g., Picelli, S, et al. Nature Methods 10, 1096- 1098 (2013)). For example, template switching can be performed on the cDNA molecules, i.e., the cDNA / RNA duplexes attached to streptavidin beads, using a template switching oligonucleotide (TSO) comprising the adapter sequence. In some embodiments, up to 10% w / v PEG (molecular weight 7000-9000) is used in the template switch reaction. In some embodiments of the present methods, a TS primer sequence present within the adapter sequence introduced by the TSO is used for the amplification of tagged WT cDNA molecules, e.g., as illustrated in FIG. 3C.12. cDNA (or genomic DNA fragment) amplification
[0144] Following the isolation of the tagged cDNA molecules and second strand synthesis, the cDNA molecules are amplified using WT primers (e.g., TSO and R2 primers) configured to amplify all tagged cDNAs in the transcriptome.
[0145] In some embodiments, amplification comprises amplifying cDNA molecules using at least one pair of primers (i.e., whole transcriptome (WT) preamplification primers) configured to broadly amplify tagged cDNA molecules in the mixture. In some embodiments, the at least one pair of WT primers can comprise one primer complementary to an adapter sequence introduced by the TSO, and one primer complementary to an adapter sequence (e.g., R2 sequence) introduced by the last nucleic acid tag appended to the cDNA during split-pool labeling. In embodiments involving genomic DNA fragments, similarly suitable primers are used, e.g., primers specific to adapters appended to the genomic DNA fragment ends and / or introduced via one or more nucleic acid tags.
[0146] In some embodiments, one or more WT preamplification primers used in the preamplification round comprise the sequence of SEQ ID NO:591 or SEQ ID NO:592, or a sequence comprising 80%, 85%, 90%, 95%, 96%, 97%, 98%, or 99% identity to SEQ ID NO:591 or SEQ ID NO:592, or a sequence comprising not more than 1, 2, 3, 4, or 5 mismatches relative to SEQ ID NO:591 or SEQ ID NO:592.
[0147] The primers can be used at a range of concentrations. In some embodiments, the primers are each added at from 1 to 10 |1M, e.g., at 1.2 |1M, 2.4 |1M, 4.8 |1M, 7.2 |1M, 9.6 |1M, or 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more |1M, or from 100 nM to 1 |1M, e.g., 100, 200, 300, 400, 500, 600, 700, 800, 900, or more nM for each target preamplification primer. In some embodiments, the ratio (e.g., molar ratio) of target-specific primers to WT primers used is 1:1000, 1:900, 1:800, 1:700, 1:600, 1:500, 1:400, 1:300, 1:200, 1:100, 1:90, 1:80, 1:70, 1:60, 1:50, 1:40, 1:30, 1:20, 1:10, 1:9, 1:8, 1:7, 1:6, 1:5, 1:4, 1:3, 1:2, 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, 200:1, 300:1, 400:1, 500:1, 600:1, 700:1, 800:1, 900:1, or 1000:1.
[0148] It will be appreciated that primers said to “non-specifically” amplify the tagged cDNAs (e.g., WT preamplification primers) in the mixture means that they are configured broadly amplify all tagged cDNAs generated using the present methods (e.g., all cDNAs comprising a TSO adapter sequence and an R2 adapter sequence). It does not necessarily mean, however, that the primers will have identical affinity for each of the tagged cDNAs in the mixture, or that they will amplify all tagged cDNAs in the mixture at an identical rate. Instead, primers said to “non-specifically” amplify the tagged cDNAs simply means that the primers are not designed to specifically anneal to any particular cDNAs in the mixture, e.g., are not designed to be complementary to a target sequence, such that the primers can generally be expected to amplify the overall collection of tagged cDNA molecules in the mixture.
[0149] In particular embodiments, the preamplification step is performed by placing tubes comprising the (magnetic) beads with bound tagged cDNA molecules against a magnetic rack, removing and discarding the clear supernatant, and resuspending the beads in bind buffer, then removing and discarding the supernatant, then removing the tubes from the magnetic rack and resuspending the beads in amplification reaction solution (e.g., a solution comprising an amplification master buffer, and WT preamplification primers for amplifying the wholetranscriptome. The resuspended beads can then be stored on ice (or a suitable temperature, e.g., at about 0, 1, 2, 3, 4, 5, 6, 7, or 8 °C). The tubes can then be placed in a thermocycler and subjected to suitable conditions for PCR amplification of the cDNAs. Following amplification, the tubes can be removed and stored, e.g., at 4 °C.
[0150] Following preamplification of the cDNA molecules, the PCR products can be cleaned up, e.g., by the addition of solid phase reversible immobilization (SPRI) beads. SPRI beads are used to remove polynucleotides of less than about 200 base pairs, less than about 175 base pairs, or less than about 150 base pairs (see DeAngelis, MM, et al. Nucleic Acids Research (1995) 23(22):4742). In particular embodiments, SPRI beads are used to remove polynucleotides of less than about 200 base pairs. The ratio of SPRI bead solution to amplified cDNA molecule solution may be between about 0.9: 1 and about 0.7: 1, between about 0.875: 1 and about 0.775: 1, between about 0.85: 1 and about 0.75: 1, between about 0.825: 1 and about 0.725:1, about 0.8: 1, or another suitable ratio. In some embodiments, the SPRI clean-up is single-sided. In some embodiments, the SPRI clean-up is double-sided.
[0151] Furthermore, the SPRI bead solution may include between about 1 M and 4 M NaCl, between about 2 M and 3 M NaCl, between about 2.25 M and 2.75 M NaCl, about 2.5 MNaCl, or another suitable amount of NaCl. The SPRI bead solution may also include between about 15% w / v and 25% w / v polyethylene glycol (PEG), wherein the molecular weight of the PEG is between about 7,000 g / mol and 9,000 g / mol (PEG 8000). In various embodiments, the SPRI bead solution may include between about 17% w / v and 23% w / v PEG 8000, between about 18% w / v and 22% w / v PEG 8000, between about 19% w / v and 21% w / v PEG 8000, about 20% w / v PEG 8000, or another suitable % w / v PEG 8000.
[0152] Specifically, 20 pl of the RNase-treated beads can be added to a single PCR tube. 80 pl of ligase mix (5 pl T4 RNA Ligase 1 (NEW ENGLAND BIOLABS®), 10 pl 10X T4 RNA ligase buffer, 5 pl BC_0047 oligo at 50 pM, 50 pl 50% PEG 8000, and 10 pl 10 mM ATP) can be added to the 20 pl of beads in the PCR tube. 50 pl of the ligase mixed with the beads can be transferred into a new PCR tube to prevent too many beads from settling to the bottom of a single tube and the sample can be incubated at 25 °C for 16 hours.13. Preparation of Sequencing libraries
[0153] Following cDNA amplification, the amplification products are used to prepare a sequencing library for each cDNA sublibrary, e.g., 1, 8, 16, or more sublibraries. In someembodiments, during preparation of the WT library or libraries, an additional (e.g., fourth) sublibrary-specific barcode is added to the cDNA molecules. In some embodiments, this additional (e.g., fourth) barcode is an Illumina Unique Dual Index (UDI).Whole transcriptome (WT) Sequencing Libraries
[0154] In some embodiments, to prepare WT sequencing libraries, the amplified WT cDNA molecules are fragmented, an adapter comprising, inter alia, a primer binding sequence is appended to the fragmented ends, and an additional amplification reaction is performed to introduce one or more index sequences (e.g., unique dual indexes, or UDIs) and sequencing primer binding sites for NGS sequencing (see, e.g., FIGS. 4A-4C). The WT cDNAs can be fragmented, e.g., using a fragmentation enzyme and fragmentation buffer. In some embodiments, the preamplified cDNA molecules are fragmented by incubating with the fragmentation enzyme and buffer at 32 °C for 10 minutes and are then held at 65 °C for, e.g., 30 minutes. In some embodiments, following fragmentation of the DNA, the fragment ends are repaired and A-tailed, and the adapter is ligated to the ends. For example, in some embodiments, an Illumina Truseq R1 Adapter is ligated to the 5’ end of the DNA. In some embodiments, the ligation of the adapter to the ends of the fragments of the preamplified cDNA can be preceded and / or followed by an SPRI clean-up step, e.g., using Ampure XP or KAPA Pure Beads.
[0155] In some embodiments, the cleaned-up molecules are then subjected to an additional round of amplification, e.g., adding P5 / P7 adapter sequences. A fourth barcode can also be added. In some embodiments, the additional barcodes correspond to unique dual indexes (UDI), e.g., with different well-specific index primers used for each sublibrary (see, e.g., Example 1 and Table 32). In some embodiments, the indexing round of amplification can be preceded and / or followed by an additional SPRI clean up step (e.g., using Ampure XP or KAPA Pure Beads).
[0156] FIGS. 4A-4C illustrate an exemplary embodiment of preparation of whole transcriptome libraries for sequencing. Sublibrary cDNA is fragmented to a size compatible with a suitable sequencing platform (FIG. 4A), e.g., Illumina sequencing or other compatible sequencing platforms. A second adaptor (e.g., an R1 adapter) is ligated to the fragmented end of the cDNA (FIG. 4B). Lastly, a final PCR (FIG. 4C) amplifies the fragmented cDNA and appends the fourth DNA barcode, UDIs, and P5 and P7 adaptors. The final library structure is shown in FIG. 5.14. Sequencing and analysis
[0157] Sequencing can be performed on any suitable NGS platform, e.g., Illumina. In particular embodiments, libraries are sequenced with paired reads, e.g., using i7 and i5 indexes provided by the UDI barcodes as described elsewhere herein (see, e.g., Table 32).
[0158] In some embodiments, the sequencing reads are subjected to one or more filtration steps, e.g., filtering the sequencing reads to remove transcripts represented by only one read, or by only 1, 2, 3, 4, or 5 reads.
[0159] In some embodiments, sequencing reads are grouped by cell barcodes (e.g., RT primer barcodes, nucleic acid tag barcodes, first barcode, second barcode, UDI barcodes, and combinations thereof). Each barcode combination should correspond to the cDNA from a single cell. In some embodiments, only reads with valid barcodes are retained. For the WT library, the sequencing reads with each barcode combination can be aligned to a reference genome, e.g., to a reference human genome. Multiple reads with the same random identifier sequence are counted as a single read. In some embodiments, reads with random identifier sequences with two or less mismatches are assumed to be generated by sequencing errors and are counted as a single read. In some embodiments, transcripts represented by one read only are filtered out and not included in subsequent analysis of the libraries.15. Kits
[0160] Another aspect of the disclosure relates to kits for preparing and sequencing enriched sequencing libraries, e.g., sequencing libraries for performing whole transcriptome sequencing together or analyzing accessible chromatin regions in genomic DNA as described herein. In some embodiments, the kit may comprise at least one reverse transcription (RT) primer, e.g., an RT primer as disclosed herein comprising an RT barcode, a sequence such as a poly(dT) sequence, a random sequence, or a target sequence, and / or a 5' overhang sequence (e.g., a sequence as disclosed in Table 6-7 or comprising a sequence of any of SEQ ID NOS: 401-590). The kit may also comprise a plurality of nucleic acid tags, e.g., nucleic acid tags with well-specific barcodes (e.g., a nucleic acid sequence comprising any of the sequences shown in any of Tables 1-5 or as SEQ ID NOS: 1-400). Each first nucleic acid tag may comprise a first strand. Each first nucleic acid tag may further comprise a second strand. The kit may also comprise one or more stop oligos according to the present disclosure. In some embodiments, the nucleic acid tags are pre-complexed with their corresponding linker strands in the kit.
[0161] In some embodiments, the kit may comprise one or more sets of nucleic acid tags, e.g., sets of nucleic acid tags configured to be used in a given round of ligation-based tagging. For example, each tag in a given set of nucleic acid tags may comprise the same 5’ and / or 3’ hybridization sequence (i.e., first or second strand hybridization or overhang sequences), with the 5’ and / or 3’ hybridization sequences differing from the 5’ and / or 3’ hybridization sequences used in other sets of tags. Each set of tags may also comprise a plurality of distinct barcode sequences, e.g., 96 (or more) different barcode sequences for distinctively labeling cDNAs present in cells or nuclei within each well of a 96-well plate. The different sets of nucleic acid tags may also differ with respect to the presence or absence of additional elements such as capture agent (e.g., biotin), a random sequence, and / or an adapter sequence such as an NGS adapter sequence.
[0162] In various embodiments, the kit may further comprise at least one of a reverse transcriptase, a fixation agent, a permeabilization agent, a ligation agent, and / or a lysis agent.
[0163] In particular embodiments, the kit comprises primers for amplifying cDNA molecules according to the present disclosure, e.g., one or more primers comprising any of the nucleotide sequence shown as SEQ ID NOS :401-495.
[0164] In some embodiments, the kit comprises one or more reaction vessels or containers for the any one or more of the herein-described compositions or methods. For example, in particular embodiments the kit comprises one or more multi-well plates such as 96-well plates. In some embodiments, the kit comprises one or more multi-well plates pre-loaded with barcoded RT primers, with nucleic acid tags, or with indexed primers (e.g. UDI primers) according to the present disclosure.16. Examples
[0165] The following examples are illustrative of disclosed methods and compositions. In light of this disclosure, those of skill in the art will recognize that variations of these examples and other examples of the disclosed methods and compositions would be possible without undue experimentation.Example 1. Exemplary protocol for cell- or nucleus-specific labeling of the whole transcriptomeWorkflow
[0166] Evercode split-pool combinatorial barcoding brings a simple workflow to large- scale single cell RNA-seq experiments. The Evercode WT v3 kit can profile up to 100,000 cells across up to 48 different biological samples or experimental conditions. Evercode Fixation kits fix and permeabilize cells / nuclei so they act as individual reaction compartments, which eliminates the need for dedicated microfluidics instrumentation. Through four rounds of barcoding, the transcriptome of each fixed cell / nuclei is uniquely labeled. The four rounds of barcoding can yield a very large number of possible barcode combinations that is more than sufficient to uniquely label up to 100,000 cells or nuclei while avoiding doublets. After sequencing, the Parse Biosciences Analysis Pipeline assigns reads that share the same 4 barcode combination to a single cell or nucleus.
[0167] Steps of the process are illustrated, e.g., in FIGS. 1-4, and an overview of the steps is provided in Table 9.General guidelines and tipsSample Input
[0168] This protocol begins with cells or nuclei fixed with, e.g., an Evercode Cell Fixation v3 or Evercode Nuclei Fixation v3 kit (see, e.g., Examples 5-6). It is also compatible with samples fixed with Evercode Cell Fixation v2 or Evercode Nuclei Fixation v2 kits. Even if samples were counted before freezing, we strongly recommend counting cells / nuclei again after thawing to account for any changes during storage and freeze thaw. Typically, a 5-15% decrease after thawing should be expected. These counts are used to determine how cells / nuclei are loaded in the Round 1 Plate, and their accuracy is critical to recover the desired number of cells / nuclei.
[0169] When processing many fixed samples, we recommend aliquoting samples after fixation and counting the aliquots the day before using an Evercode Whole Transcriptome kit. The Evercode Fixation User Manuals outline recommendations for generating aliquots. Because aliquots have undergone a similar storage time and a freeze thaw, cell / nuclei counts from these aliquots will be more representative than using counts from immediately after fixation. Aliquots should be thawed in a water bath set to 37 °C in sets of 2-4 and counted with a hemocytometer or alternative counting device. Counts should be recorded, e.g., in a Sample Loading Table, and any remaining counting aliquots should be discarded. Once fixed samples have been thawed, they should not be refrozen.Cell / Nuclei Counting and Quality Assessment
[0170] We recommend a hemocytometer for counting, but alternative counting devices can also be used. If possible, validate counts from alternative devices to a hemocytometer when using Evercode Whole Transcriptome kits for the first time. When first using Evercode Whole Transcriptome kits, we suggest saving images at each counting step. To assess sample quality, we recommend using viability stains like trypan blue or acridine orange and propidium iodide (AO / PI). High quality fixed samples have single distinct cells / nuclei with <5% cell aggregation and no debris (see, e.g., FIG. 7A). Higher levels of aggregation will lead to elevated doublets after sequencing (see, e.g., FIG. 7B). When quantifying fixed samples, it is critical to avoid counting cell debris to avoid overestimating the number of cells / nuclei (see, e.g., FIG. 7C).Avoiding RNase Contamination
[0171] Take standard precautions to avoid introducing RNases into samples or reagents throughout the workflow. Always wear proper laboratory gloves and use aseptic technique. Although RNases are not inactivated by ethanol or isopropanol, they are inactivated by products such as RNaseZap RNase Decontamination Solution (Thermo Fisher Scientific). These can be sprayed on benchtops and pipettes. Filtered pipette tips should be used to reduce RNase contamination from pipettes.Centrifugation
[0172] A range of centrifugation speeds and durations are provided in this protocol. Optimize centrifugation conditions for each sample type to balance retention and resuspension efficiencies. Use a swinging bucket rotor for all high-speed centrifugation steps in this protocol. A fixed-angle rotor will lead to substantial cell / nuclei loss.Optimizing Cell / Nuclei Recovery
[0173] It is critical to thoroughly resuspend the cells / nuclei after centrifugation throughout the protocol. Resuspend by slowly and repeatedly pipetting up and down until no clumps are visible. Due to cell / nuclei adherence to tubes, carefully pipette up and down along the bottom and sides of tubes to minimize cell / nuclei loss.
[0174] We do not recommend wide bore pipette tips as they make it difficult to resuspend cell / nuclei pellets adequately. Ensure that the 15 mL centrifuge tubes that will be used are polypropylene, as polystyrene tubes will lead to substantial sample loss. The first time usingan Evercode Whole Transcriptome kit, we recommend retaining supernatants after each barcoding step. In the unlikely event of unexpectedly high sample loss, these supernatants can be analyzed to identify points for optimization.Sample Loading Tables
[0175] It may be helpful to complete a spreadsheet such as the Parse Biosciences Evercode WT Sample Loading Table v2 (Excel spreadsheet) before starting the experiment. Such tables help determine how many cells or nuclei to load for each sample, to ensure that an appropriate number of cells or nuclei are barcoded. In the unlikely event that samples are not concentrated enough according to the Sample Loading Table (Excel spreadsheet), choose to either: Decrease the “Max number barcoded cells” until the Sample Loading Table no longer gives an error, which will maintain the desired proportion between samples but barcode fewer cells. Add 14 pL of undiluted sample into each designated well of the Round 1 Plate, which will result in more overall barcoded cells but change the desired proportion between samples.Cell Strainers
[0176] A cell strainer with an appropriately sized mesh should be used throughout the protocol. Although 30-40 pm is appropriate for many cell types, the mesh size should be chosen based on the sample type. To maximize cell retention with cell strainers, press the pipette tip directly against the mesh. Ensure ample pressure is applied to hold contact between the tip and the strainer to force liquid through in ~1 second.Vortexing
[0177] Unless specified in an individual step, we strongly discourage vortexing samples or enzymes throughout the protocol.Plate Sealing
[0178] While sealing or unsealing 96 well plates, do not splash liquid onto the PCR plate seal or between wells. Securing plates in PCR tube racks will minimize this occurrence. PCR plate seals may be difficult to remove. Carefully peel the PCR plate seal while applying downward pressure on the plate to keep it in the PCR tube rack.Magnetic Racks and Bead Cleanups
[0179] The Parse Biosciences Magnetic Rack (Parse Biosciences) uses powerful rare earth magnets for rapid and efficient magnetic bead purifications for 0.2 mL tubes. The rack has high and low magnet positions important for optimal yield at key steps. We generally do not recommend substituting alternative racks, although other racks may be used if determined to be capable of producing acceptable yields. To alternate between the positions, the rack can be flipped upside down so the magnet is closer to the top (high) or bottom (low) of the 0.2 mL tubes.
[0180] To ensure material is not lost during bead purifications, ensure supernatants are completely clear before moving to the next step. The incubation times at each step are recommendations, but visual confirmation of clearing should be used to make the final determination. Discarding any beads in supernatants will result in a reduction of transcripts and genes detected per cell (or nucleus).Sublibrary Loading
[0181] The Evercode WT kit generates 8 sublibraries with distinct Illumina indexing barcodes, which can be processed and sequenced independently or concurrently after Section 1 of the protocol. The number of cells or nuclei per sublibrary is determined when the cells are divided into sublibraries in Section 1.5. When working with a new sample type, it can be beneficial to process a single sublibrary to optimize PCR conditions before running the remaining sublibraries. Sublibraries can be loaded with different numbers of cells, and the maximum number of cells that can be analyzed is the sum of cells / nuclei across all sublibraries. Asymmetric sublibrary loading can enable cost-effective sequencing quality control. One sublibrary can be loaded with a few hundred cells / nuclei and sequenced very deeply. This data can be used to choose an appropriate sequence depth for the remaining sublibraries.Indexing Primers
[0182] The UDI Plate - WT is a 96-well plate containing 48 unique dual indexing (UDI) primers. Each well is a single-use reaction sufficient for one Barcoding Round 4 PCR reaction for a single sublibrary. Thus, the UDI Plate - WT can be used for multiple Evercode Whole Transcriptome kits. The UDI Plate - WT is sealed with a pierceable foil to minimize crosscontamination. The plate seal should be wiped with 70% ethanol and pierced with a new pipette tip immediately prior to use. Avoid splashing or mixing the liquid between individual wells.Once a well has been used, it should not be resealed or reused. We recommend selecting UDIs column-wise from left to right (starting with indices 1-8). UDI sequences can be found in Table 32.Parts ListTable 10 and Table 11 provide a list of components used in the exemplary protocol.Section 1: In situ Cell / Nuclei Bar coding1.1. Set up and Sample Counting
[0183] Prior to barcoding, cells / nuclei are thawed and counted. Appropriate dilutions, loading concentrations, and loading positions are determined (e.g., using a Sample Loading Table). To set up for barcoding: 1. Prepare the Sample Loading Table (or otherwise determine appropriate sample dilutions and plate loading amounts for later steps). 2. Cool a centrifuge with swinging bucket rotors to 4 °C. 3. Set a water bath to 37 °C. 3. 4. Fill a bucket with ice. 5. Prepare a hemocytometer, flow cytometer, or other cell counting device. 6. Gather the following items: Round 1 Plate, Round 2 Plate, Resuspension Buffer, Sample Dilution Buffer, Round 2 Ligation Buffer, Spin Additive.
[0184] 7 Thaw the previously fixed cell / nuclei samples in a water bath set to 37 °C until all ice crystals dissolve. 8. Thoroughly mix each sample by pipetting and store on ice. 8. While minimizing time on ice, count the number of cells / nuclei in the sample with a hemocytometer or alternative cell counting device. Record the cell / nuclei count. Note: When processing many fixed samples, we recommend aliquoting samples after fixation and counting the aliquots theday before using an Evercode Whole Transcriptome kit. See the Guidelines section above for details.
[0185] 9. Record the sample names and cell / nuclei count. 10. Place the Round 1 Plate into a thermocycler and run the program shown in Table 12. 11. Dilute each sample with sample dilution buffer and store on ice. 12. Proceed immediately to Section 1.2.1.2 Barcoding Round 1
[0186] Samples are loaded into Round 1 Plate. An in situ reverse transcription reaction adds well-specific barcodes that also serve as sample barcodes. Cells / nuclei are then pooled, centrifuged, and resuspended. To add round 1 barcodes: 1. Gently remove the Round 1 Plate from the thermocycler, place in a 0.2 mL tube rack, and centrifuge for 1 minute at 100 x g at 4 °C. 2. Remove the Round 1 Plate from the centrifuge, place in a PCR tube rack, remove the plate seal, and store on ice. 3. With the Round 1 Plate on ice, add 14 pL of each diluted sample to the appropriate wells of Round 1 Plate. Mix immediately after dispensing each sample by pipetting 3x. Note: Do not reuse any tips during this step. Different tips must be used when pipetting cells into each well of the plate. Never place a tip that has entered one of the wells into a sample or different well. Note: When pipetting the same sample into many wells, the sample must be mixed by gentle pipetting prior to each transfer to avoid cells or nuclei settling. Do not vortex the samples. 4. While secured in a PCR tube rack on a flat surface, add a new plate seal. 5. Place the Round 1 Plate into a thermocycler and run the program shown in Table 13. Upon completion, proceed immediately to the next step.
[0187] 6. Remove the Round 1 Plate from the thermocycler, place in a PCR tube rack, and store on ice. 7. Place the Round 2 Plate into a thermocycler and run the program shown in Table 14. 8. While secured in a PCR tube rack on a flat surface, remove the plate seal from the Round 1 Plate. 9. With the plate and tube on ice, pool all wells from the Round 1 plate into a 15 mL centrifuge tube, and pipette as follows: i. With a multichannel P200 set to 30 pL, mix the sample in row B of a 96-well plate by pipetting 3x. ii. Transfer 30 pL from row B to row A of the 96-well plate, iii. Repeat i-ii for rows C-D to mix the sample, then transfer to row A. iv. Transfer any residual liquid in rows B-D to row A with a multichannel P20 set to 10 pL. v. Ensure the cells / nuclei in row A are in suspension as described in i. Then, transfer the total volume of each well in row A into the same 15 mL tube with a single channel P200 set to 200 pL.
[0188] 10. Add 9.6 pL of spin additive to the 15 mL tube with pooled cells. Do not discard the spin additive as it will be needed in another step. 11. Mix by gently inverting the tube justonce. 12. Centrifuge the 15 mL tube in a swinging bucket rotor for 5-10 minutes at 200-500 x g at 4°C. Immediately move to the next step after centrifugation. Note: Ideal centrifugation speed and duration should be determined empirically for each sample type to optimize retention and resuspension efficiencies. Overly aggressive centrifugation can damage cell integrity and decrease data quality. If the centrifugation speeds used during fixation gave satisfactory retention, they should be used throughout this protocol. Move quickly and handle the samples gently to avoid dislodging the pellet, which will have a significant impact on data quality.
[0189] 13. Remove the supernatant until about ~40 pL of liquid remains above the pellet.Use a Pl 000 for the first 1 mL and then a P200 for the remaining volume. Note Depending on the number of cells and cell types, a pellet may or may not be visible. 14. Fully but gently resuspend the pellet in 1 mL of resuspension buffer. 15. Add an additional 1 mL of resuspension buffer for a total addition of 2 mL. Store on ice. 16. Proceed immediately to section 1.3. Note: If a low input fixation workflow was performed, execute the following step before proceeding to Section 1.3. 17. 17. Pipette the sample through a cell strainer into a new 15 mL tube with a Pl 000. Before transfer, gently mix the cells by pipetting 2x.1.3 Barcoding Round 2
[0190] The pooled cells are added to the Ligation Master Mix, which is loaded into the Round 2 Plate. An in situ ligation reaction adds a well-specific barcode to the 3’ end of the cDNA. The ligation reaction is quenched with Round 2 stop buffer, and the cells / nuclei are pooled and strained.
[0191] To add round 2 barcodes: 1. Gather the following items: Round 2 Ligation Enzyme, Round 2 Stop Buffer, and Round 3 Plate. 2. On ice, prepare the Round 2 Ligation Master Mix by adding the items shown in Table 15 to the samples in Resuspension Buffer (prepared in Section 1.2). Mix thoroughly by pipetting lOx with a P1000 set to 1000 pL. Store on ice. Note: Prior to use, mix by inverting 3x the Round 2 Ligation Buffer that was gathered and stored on ice in Section 1.1.
[0192] 3.Remove the Round 2 Plate from the thermocycler, place in a PCR tube rack, and centrifuge for 1 minute at 100 x g at 4 °C. 4. While secured in a PCR tube rack on a flat surface, remove the plate seal from the Round 2 Plate and store on ice. 5. Transfer the Round 2 Ligation Master Mix to a basin with a Pl 000. 6. With the Round 2 Plate on ice and the basin on the bench, transfer Round 2 Ligation Master Mix to each well in the Round 2 Plate as follows: i. Mix the sample in the basin by pipetting 2x with a multichannel P200 set to 40 pL. ii. Transfer 40 pL of the mix to row A of the Round 2 Plate and mix by pipetting 2x. iii. Repeat i-ii to mix the sample in the basin then transfer to rows B-H. Note: Do not reuse any tips during this step. Different tips must be used when pipetting cells into each well of the plate. Never place a tip that has entered one of the wells back into the basin or a different well. Note: If the volume is insufficient to transfer the last row with a multichannel, tilt the basin and transfer the remaining volume with a single channel pipette. If the volume is still insufficient to fill every well, a few can be left empty without impacting experimental results.
[0193] 7 While secured in a PCR tube rack on a flat surface, add a new plate seal to theRound 2 Plate. 8. Place the Round 2 Plate into a thermocycler and run the program shown in Table 16. Upon completion, proceed immediately to the next step.
[0194] 9. Briefly vortex the Round 2 stop buffer and ensure there is no precipitate.Transfer the entire volume of this tube to a new basin with a Pl 000. 10. Remove the Round 2 Plate from the thermocycler, place in a PCR tube rack, remove the plate seal, and store on ice. 11. With the Round 2 Plate on ice and the basin on the bench, transfer 10 pL of the Round 2 stop buffer to each well in the Round 2 Plate with a multichannel P20. After each transfer, mixby pipetting exactly 3x. Note: Do not reuse any tips during this step. Different tips must be used when pipetting Round 2 Stop Buffer into each well of the plate. Never place a tip that has entered one of the wells back into the basin or a different well. 12. While secured in a PCR tube rack on a flat surface, add a new plate seal to the Round 2 Plate. 13. Place the Round 2 Plate into a thermocycler and run the program shown in Table 17. Proceed to the next step while the program is still running.
[0195] 14. Place the Round 3 Plate into a second thermocycler and run the program shown in Table 18. Proceed to the next step while the program is still running. Note: If a second thermocycler is not available, the same thermocycler used in step 14 can subsequently be used in step 15. However, the Round 2 Plate should be stored on ice until the Thaw Round 3 Plate program is complete.
[0196] 15. Immediately upon completion of the Round 2 Stop program, transfer the Round2 Plate from the thermocycler to a PCR tube rack, remove the plate seal, and store on ice. 16. With the Round 2 Plate on ice and the basin on the bench, transfer all the liquid in the Round 2 Plate into a new basin as follows: i. With a multichannel P200 set to 50 pL, mix the sample in row A by pipetting 3x. ii. Transfer 50 pL from row A to the basin, iii. Repeat i-ii for rowsB-H to mix the sample then transfer to the basin, iv. Transfer any residual liquid in the Round 2 Plate to the basin with a multichannel P20 set to 10 pL.
[0197] 17. Pipette the sample through a cell strainer into a new basin with a P1000. Before each transfer, gently mix the cells in the basin by pipetting 2x and tilt the basin to recover as much liquid as possible. Note: Do not directly touch the mesh of the cell strainer with gloved hands. To ensure that all of the liquid passes through the strainer, press the tip of the pipette against the mesh to create a tight seal and press the pipette plunger down steadily. All the liquid should pass through the strainer in ~1 second. Note: Bubbles may form while pooling or straining. They will not affect the quality of the experiment. 18. Proceed immediately to Section 1.4.1.4 Barcoding Round 3
[0198] The Round 3 Ligation Enzyme is added to the pooled cells / nuclei, which are then loaded into the Round 3 Plate. A second in situ ligation reaction adds a third well-specific barcode, the Illumina Truseq R2 sequence, and a biotin. The sample is then pooled and strained.
[0199] To add round 3 barcodes: 1. Gather the following items: Round 3 Stop Buffer, PreLysis Wash Buffer, Round 3 Ligation Enzyme, Pre-Lysis Dilution Buffer, Lysis Enzyme, Lysis Buffer. 2. Add 20 pL of Round 3 Ligation Enzyme to the basin containing the strained sample. Mix by gently pipetting 20x with a Pl 000 set to 1000 pL. 3. Remove the Round 3 Plate from the thermocycler, place in a PCR tube rack, and centrifuge for 1 minute at 100 x g at 4°C. 4. While secured in a PCR tube rack on a flat surface, remove the plate seal from the Round 3 Plate.
[0200] 5. With the Round 3 Plate on ice and the basin on the bench, transfer 50 pL from the basin to each well in the Round 3 Plate as follows: i. Mix the sample in the basin by pipetting 2x with a multichannel P200 set to 50 pL. ii. Transfer 50 pL of the mix to row A of the Round 3 Plate and mix by pipetting 2x. iii. Repeat i-ii to mix the sample in the basin then transfer to rows B-H. Note: Do not reuse any tips during this step. Different tips must be used when pipetting cells into each well of the plate. Never place a tip that has entered one of the wells back into the basin or a different well. Note: If the volume is insufficient to transfer the last row with a multichannel, tilt the basin and transfer the remaining volume with a single channel pipette. If the volume is still insufficient to fill every well, a few can be left empty without impacting experimental results. 6. While secured in a PCR tube rack on a flat surface, add anew plate seal to Round 3 Plate. 7. Place the Round 3 Plate into a thermocycler and run the program shown in Table 19.
[0201] 8. Briefly vortex the Round 3 stop buffer. Transfer the entire volume to a new basin with a Pl 000. 9. Remove the Round 3 Plate from the thermocycler, place in a PCR tube rack, remove the plate seal, and store on ice. 10. With the Round 3 Plate on ice and the basin on the bench, transfer 20 pL of the Round 3 stop buffer from the basin to each well in the Round 3 Plate with a multichannel P20. After each transfer, mix by pipetting exactly 3x. Note: Do not reuse any tips during this step. Different tips must be used when pipetting Round 3 Stop Buffer into each well of the plate. Never place a tip that has entered one of the wells back into the basin or a different well. 11. Without incubation, proceed immediately to the next step.
[0202] 12. With the Round 3 Plate on ice and the basin on the bench, transfer all the liquid in the Round 3 Plate into a new basin as follows: i. With a multichannel P200 set to 70 pL, mix the sample in row A by pipetting 3x. ii. Transfer 70 pL from row A to the basin, iii. Repeat i- ii for rows B-H to mix the sample then transfer to the basin, iv. Transfer any residual liquid in the Round 3 Plate to the basin with a multichannel P20 pipette set to 10 pL. 13. Pipette the sample through a cell strainer into a new 15 mL tube with a P1000. Before each transfer, gently mix the cells in the basin by pipetting 2x and tilt the basin to recover as much liquid as possible. Note: Bubbles may form while pooling or straining. They will not affect the quality of the experiment. 14. Proceed immediately to section 1.5.Lysis and Sublibrary Generation
[0203] The cell / nuclei pool is centrifuged, washed, and resuspended in Pre-Lysis Dilution buffer. The cells / nuclei are counted and divided into sublibraries. These sublibraries are lysed and stored at -80°C. To generate and lyse sublibraries: 1. Add 70 pL of spin additive to the 15 mL tube with the sample. Gently invert once to mix. 2. Centrifuge the 15 mL tube in a swingingbucket rotor for 5-10 minutes at 200-500 x g at 4°C. Immediately move to the next step after centrifugation. Note: Move quickly and handle the samples gently to avoid dislodging the pellet, which will impact data quality. 3. Remove the supernatant until about ~40 pL of liquid remains above the pellet. Use a Pl 000 for the first 6 mL and then a P200 for the remaining volume. Note: Depending on the number of cells / nuclei and sample type, the pellet may or may not be visible. 4. Fully but gently resuspend the pellet in 1 mL of Pre-Lyse Wash Buffer. 5. Add an additional 3 mL of Pre-Lyse Wash Buffer for a total addition of 4 mL. 6. Centrifuge the 15 mL tube in a swinging bucket rotor for 5-10 minutes at 200-500 x g at 4°C. Immediately move to the next step after centrifugation. Note: Move quickly and handle the samples gently to avoid dislodging the pellet, which will impact data quality. 7. Remove the supernatant until about ~40 pL of liquid remains above the pellet. Use a Pl 000 for the first 3 mL and then a P200 for the remaining volume. 8. Fully but gently resuspend the pellet in 60 pL of Pre-Lysis Dilution Buffer for a final total volume of 100 pL. Store on ice. Note: Do not discard Pre-Lysis Dilution Buffer as it will be used in another step. 9. While minimizing time on ice, count the number of cells / nuclei in the sample with a hemocytometer or alternative cell counting device. Record the cell / nuclei count. Note: We strongly recommend using a hemocytometer and carefully mixing the sample before collecting an aliquot for counting to ensure accurate sublibrary loading.
[0204] 10. Decide how to divide cells / nuclei across the sublibraries (see, e.g., the section on Sublibrary Loading in the Guidelines above. Note: Do not add more than 12,500 cells / nuclei to a sublibrary. Adding additional cells / nuclei will result in an increased multiplet rate. 11. Ensure the cells / nuclei are in suspension by pipetting 5x with a P200 set to 75 pL prior to each transfer. Add the appropriate volume of sample to 8 different 0.2 mL PCR tubes. 12. Add the appropriate volume of Pre-Lysis Dilution Buffer to the 0.2 mL tubes for a total volume of 25 pL. 13. Prepare the Lysis Master Mix in a new 1.5 mL tube by combining 220 pL of lysis buffer with 44 pl of lysis enzyme, for a total volume of 264 pL. Mix by pipetting 3x with a Pl 000 set to 220 pL. Store at room temperature. Note: Ensure that there is no precipitate before using the Lysis Buffer. Note: Do not place Lysis Master Mix on ice, as a precipitate will form.
[0205] 14. Add 30 pL of Lysis Master Mix to each 0.2 mL tube with diluted cells / nuclei.Store at room temperature. 15. Vortex the 0.2 mL tube(s) for 10 seconds. Briefly centrifuge. 16. Place the tube(s) into a thermocycler and run the program shown in Table 20. If continuing to Section 2 without freezing the sample, proceed to Section 2 while the program is stillrunning. 17. Freeze the lysate(s) at -80°C or proceed to Section 2. Safe stopping point: Sublibrary lysates can be stored at -80°C for up to 6 months.2: cDNA Capture and Amplification2.1. cDNA Capture
[0206] The barcoded cDNA is captured with streptavidin-coated magnetic beads and washed to remove cellular debris. To capture the cDNA: 1. Fill an ice bucket. 2. For each lysate, prepare 400 pL of 85% ethanol per lysate with nuclease-free water. 3. Equilibrate 80 pL of SPRI beads per lysate to room temperature 4. Gather the following equipment: i. Magnetic rack for 1.5 mL tubes; ii. Parse Biosciences magnetic rack for 0.2 mL PCR tubes; iii. Vortex with an adapter for 96 well plates. 5. Gather the following items: Streptavidin beads, Bead wash buffer, Wash buffer 1, Wash buffer 2, Capture enhancer, Binding buffer. 6. Remove the desired tube(s) of lysate from the thermocycler (if continuing directly from Section 1) or storage at - 80 °C. 7. If previously frozen, incubate the tube(s) in water bath or thermocycler at 37°C for 5 minutes. 8. Briefly centrifuge and store at room temperature. 9. Briefly centrifuge Capture Enhancer and gently mix by pipetting 2x with a P20 set to 15 pL. 10. Add 2.5 pL of Capture Enhancer to each tube of lysate and mix by pipetting 5x with a P200 set to 40 pL. Briefly centrifuge. 11. Incubate for 10 minutes at room temperature. Proceed immediately to the next step during the incubation. Note: This incubation can be extended by 5 additional minutes (up to a total of 15 minutes) without negatively impacting performance.
[0207] 12. Vortex Streptavidin Beads until fully mixed. Add an appropriate volume ofStreptavidin beads to a new 1.5 mL tube depending on the number of lysates being processed, e.g., as follows: for 1 lysate being processed, use 44 pL of streptavidin beads; for 8 lysates being processed, use 352 pL of Streptavidin beads. 13. Place the tube on the magnetic rack for 1.5 mL tubes until the solution clears (~2 minutes). 14. Remove and discard the supernatant.15. Remove the tube from the magnetic rack and fully resuspend the bead pellet in an appropriate volume of Bead Wash Buffer, e.g., as follows: for 1 lysate being processed, use 50 pL bead wash buffer; for 8 lysates being processed, use 400 pL bead wash buffer. Note: Ensure no beads are stuck to the sides of the 1.5 mL tube.
[0208] 16. Place the tube on the magnetic rack for 1.5 mL tubes until the solution clears(-2 minutes). 17. Remove and discard the supernatant. 18. Repeat steps 15-17 twice for a total of 3 washes. 19. Remove the tube from the magnetic rack. Fully resuspend the pellet in an appropriate volume of binding buffer, as follows: for 1 lysate being processed, use 55 pL of binding buffer; for 8 lysates being processed, use 440 pL of binding buffer. Store at room temperature.
[0209] 20. Add 50 pL of Streptavidin beads in Binding Buffer to each tube of lysate and fully mix by pipetting 2x with a P200 set to 90 pL. 21. Place the tube(s) into a 96 well PCR tube rack, press to secure, and ensure the caps are secured tightly. Place the lid on the rack. 22. Place the rack onto a vortex mixer with a plate adapter. Push to secure. Vortex on 20% power (-800-1000 RPM) for 30 minutes at room temperature. Note: To ensure the beads are being mixed sufficiently, check that the beads are in solution 10 minutes into the incubation. If settled, increase the vortex mixing speed to keep the beads in solution. 23. Remove the tube(s) from the vortex mixer. 24. Briefly vortex the tube(s) on a standard vortex adapter. Briefly centrifuge without letting beads collect at the bottom of the tube(s). 25. Place the tube(s) on the high magnet position of the Parse Biosciences Magnetic Rack, so the magnet is closer to the top of the 0.2 mL tubes. Incubate until the solution clears (-2 minutes). Note: Ensure the supernatant is completely clear before proceeding. Discarding any beads in the supernatant will result in a reduction of transcripts and genes detected per cell.
[0210] 26. While still on the magnetic rack, remove and discard the supernatant. 27.Remove the tube(s) from the magnetic rack. Fully resuspend each bead pellet with 125 pL Wash Buffer 1. 28. Incubate for 1 minute at room temperature. 29. Return the tube(s) to the high position of the magnetic rack. Incubate until the solution clears (-2 minutes). 30. While still on the magnetic rack, remove and discard the supernatant. 31. Repeat steps 27-30 once for a total of 2 washes with Wash Buffer 1. 32. Remove the tube(s) from the magnetic rack. Fully resuspend each bead pellet with 125 pL Wash Buffer 2. Note: Save Wash Buffer 2 to use for optional storage before cDNA amplification. 33. Incubate for 1 minute at room temperature. 34. Proceed immediately to Section 2.2.2.2 cDNA Template Switch
[0211] After an additional wash, the template switch master mix is added to the captured cDNA. The template switch reaction adds a 5’ adaptor to the cDNA. To perform template switch: 1. Gather the following items: Wash buffer 3, Template switch buffer, Template switch primer, Template switch enzyme. Note: Ensure that there is no precipitate in the Template Switch Buffer before proceeding. 2. Prepare the Template Switch Master Mix in a new 1.5 mL tube as shown in Table 21, depending on the number of sublibraries being processed. 2. Mix by pipetting lOx and store on ice.
[0212] 3. Place each tube of captured cDNA from Section 2.1 on the high position of the magnetic rack. Incubate until the solution clears (~2 minutes). 4. While still on the magnetic rack, remove and discard the supernatant. 5. While still on the magnetic rack, add 125 pL of Wash Buffer 3 to each tube. Note: Do not discard the Wash Buffer 3 as it will be used in another step. 6. Incubate for 1 minute at room temperature. 7. While still on the magnetic rack, remove the Wash Buffer 3. 8. Remove the tube(s) from the magnetic rack. Fully resuspend each bead pellet with 100 pL of the Template Switch Master Mix. Note: Because the Template Switch Master Mix is viscous, it may take time to fully resuspend the beads. 9. Briefly centrifuge without letting beads collect at the bottom of the tube(s). 10. Incubate for 30 minutes at room temperature. 11. Fully resuspend each bead pellet by mixing 5x with a P200 set to 75 pL. 12. Place the tube(s) into a thermocycler and run the program shown in Table 22.
[0213] 13. Proceed immediately to section 2.3. Alternatively, proceed to step 14 to store samples prior to cDNA amplification. 14. Place the tube(s) on the high position of the magnetic rack. Incubate until the solution clears (~2 minutes). Note: Beads may need to be resuspended if they have settled. 15. While still on the magnetic rack, remove and discard the supernatant. 16. Remove the tube(s) from the magnetic rack. Fully resuspend each bead pellet with 125 pL wash buffer 2. Safe stopping point: Template switched cDNA can be stored at 4°C for up to 18 hours. Do not freeze.2.3 cDNA Amplification
[0214] The captured cDNA is washed and amplified with Template Switch Primer and Illumina Truseq Read 2 specific primers. To amplify the cDNA: 1. Gather the following items: cDNA Amp Mix, cDNA Amp Primers. 2. Prepare the Amplification Reaction Solution Master Mix in a new 1.5 mL tube as shown in Table 23. 2. Mix by pipetting lOx and store on ice.
[0215] 3. Place each tube of template switched cDNA from Section 2.2 on the high position of the Parse Biosciences Magnetic Rack. Incubate until the solution clears (~2 minutes). Note: You may need to pipette mix to resuspend settled beads so they separate appropriately. 4. While still on the magnetic rack, remove and discard the supernatant. 5. While still on the magnetic rack, add 125 pL of wash buffer 3 to each tube. 6. Incubate for 1 minute at room temperature. 7. While still on the magnetic rack, remove and discard the wash buffer 3. 8. Remove tube(s) from the magnetic rack. Fully resuspend each bead pellet with 100 pL of the Amplification Master Mix. Store on ice. 9. Determine the number of PCR cycles required for cDNA amplification based on the recommendations in Table 24. Although theserecommendations are appropriate for many cell types, the number of cycles may need to be optimized for your sample type.
[0216] 10. Place the tube(s) into a thermocycler and run the program shown in Table 25.Note: If processing sublibraries with different numbers of cells / nuclei, they should be amplified in separate thermocyclers according to the recommendations in Table 24. Note: Annealing steps 3* and 6* have different time and temperature settings. Ensure these are correct before starting the program. Safe stopping point: Amplified cDNA can be stored at 4°C for up to 18 hours.2.4 Post-Amplification Purification
[0217] Amplified cDNA is purified with a 0.8x SPRI bead cleanup. To purify the cDNA: 1. For each tube of amplified cDNA, gather 400 pL of freshly prepared 85% ethanol. 2. Gather room temperature SPRI beads (80 pL per tube of amplified cDNA). Note: Ensure the SPRI beads have been equilibrated to room temperature for at least 30 minutes. 3. Place each tube of amplified cDNA from section 2.3 on the high position of the magnetic rack. Incubate until the solution clears (~2 minutes). Note: If beads remain in solution after 2-3 minutes, pipette 3x in the bottom of the PCR tube with a P200 set to 40 pL. Then return to the magnet and incubate until the solution clears. 4. While still on the magnetic rack, transfer 90 pL of the supernatant containing the cDNA into a new 0.2 mL tube(s). Store at room temperature. 5. Vortex the SPRI beads until fully mixed. Add 72 pL of SPRI beads to each tube with amplified cDNA. 6. Vortex the tube(s) for 5 seconds. Briefly centrifuge. 7. Incubate for 5 minutes at room temperature. 8. Place the tube(s) on the high position of the magnetic rack. Incubate until the solution clears (~2 minutes). 9. While still on the magnetic rack, remove and discard the supernatant.
[0218] 10. While still on the magnetic rack, add 180 pL of 85% ethanol to each tube. 11.Incubate for 1 minute at room temperature. 12. While still on the magnetic rack, remove and discard the supernatant. 13. Repeat steps 10-12 once for a total of 2 washes. Remove any residual ethanol with a P20. 14. While still on the magnetic rack, air dry the SPRI beads (~2 minutes). Note: Do not over-dry the beads, which can lead to a substantial loss in yield. “Cracking” of the beads is a sign of over-drying. 15. Remove the tube(s) from the magnetic rack. Fully resuspend each bead pellet with 25 pL of nuclease-free water. 16. Incubate for 10 minutes at 37 °C in a thermocycler. 17. Place the tube(s) on the low magnet position of the magnetic rack, so the magnet is closer to the bottom of the 0.2 mL tubes. Incubate until the solution clears (~2 minutes). 18. While still on the magnetic rack, transfer 25 pL of the supernatant containing the purified cDNA into new 0.2 mL tube(s). Store on ice. Safe stopping point: Amplified cDNA can be stored at 4 °C for up to 48 hours or at -20 °C for up to 3 months. Otherwise, proceed immediately to Section 3.2.5 cDNA Quantification
[0219] The concentration and size distribution of the cDNA are measured with fluorescent dyes and capillary electrophoresis. The cDNA is then stored at 4 °C for up to 48 hours or at - 20°C for up to 3 months. To quantify the cDNA: 1. Measure the concentration of each tube of purified cDNA (from section 2.4) with, e.g., the Qubit dsDNA HS (High Sensitivity) AssayKit according to the manufacturer’s instructions. Record the concentration(s), which will be used in section 3. 2. Assess the size distribution of each tube of purified cDNA with a High Sensitivity DNA Kit on the Agilent Bioanalyzer System or High Sensitivity D5000 ScreenTape and Reagents on the Agilent TapeStation System according to the manufacturer’s instructions. See, e.g., FIG. 8A, which shows an example of an expected post-amplification sublibrary cDNA size distribution. Note: Samples may need to be diluted to be within the manufacturer's recommended concentration range. Note: The traces in FIG. 8A are representative of typical TapeStation cDNA traces. The shape and prominence of the trace is dependent on cell type, sublibrary size, and amount of DNA loaded into the TapeStation. Sublibraries with minor deviations can still produce high quality data.3. Sequencing Library Preparation3.1. Fragmentation and End Prep
[0220] Barcoded and amplified cDNA is fragmented, end repaired and A-tailed in a single reaction. To prepare for fragmentation and end prep: 1. For each sublibrary, prepare 1.2 mL of 85% ethanol per sublibrary with nuclease-free water. 2. Equilibrate 180 pL of SPRI beads per sublibrary to room temperature. 3. Fill an ice bucket. 4. Take out the magnetic rack for 0.2 mL PCR tubes. 5. Obtain recorded cDNA concentrations (e.g., from section 2.5). 6. Gather the following items: Fragmentation / end prep buffer, Fragmentation / end prep enzymes. 7. Vortex the tube(s) of cDNA for 5 seconds. Briefly centrifuge. 8. Prepare diluted cDNA in new 0.2 mL tube(s) as follows and store on ice: 10 pL purified cDNA, 25 pL nuclease-free water, for a total volume of 35 pL. Store any remaining sublibrary cDNA at -20°C. 9. Start the program shown in Table 26.
[0221] 10. Vortex the Fragmentation Buffer for 5 seconds. Briefly centrifuge. Note:Confirm the Fragm / End Prep Buffer is fully thawed without precipitation. 11. Prepare the Fragmentation Master Mix in a new 1.5 mL tube as shown in Table 27. Mix by pipetting lOx and store on ice.12. Add 15 pL of Fragmentation Master Mix to each tube of diluted cDNA. Mix by pipetting lOx with a P200 multichannel pipette set to 40 pL. Briefly centrifuge. 13. Place the tube(s) into a cooled thermocycler. Skip the 4°C hold in Step 1 (of the Fragmentation and End Prep Program shown in Table 26 so the thermocycler proceeds to Step 2. 14. As soon as the program reaches step 4 of the thermal cycling program (4 °C), store the tube(s) on ice and proceed immediately to section 3.2.3.2. Post-Fragmentation and End Prep Size Selection
[0222] The fragmented and end prepped DNA is size selected with a double-sided SPRI cleanup. To size select the fragmented and end prepped DNA: 1. Gather freshly prepared 85% ethanol. 2. Gather room temperature SPRI beads (~50 pL per sublibrary). Note: Ensure the SPRI beads have been equilibrated to room temperature for at least 30 minutes. 3. Vortex the SPRI beads until fully mixed. Add 30 pL of SPRI beads to each tube of fragmented and end prepped DNA. 4. Vortex the tube(s) for 5 seconds. Briefly centrifuge. 5. Incubate for 5 minutes at room temperature. 6. Place the tube(s) on the high position of the magnetic rack for 0.2 mL tubes. Incubate until the solution clears (~2 minutes). 7. While still on the magnetic rack, transfer 75 pL of the supernatant containing the fragmented and end prepped DNA into new 0.2 mL tube(s). Discard the tube(s) with bead pellet(s). 8. Add 10 pL of SPRI beads to each tube. 9. Vortex the tube(s) for 5 seconds. Briefly centrifuge. 10. Incubate for 5 minutes at room temperature.
[0223] 11. Place the tube(s) on the high position of the magnetic rack for 0.2 mL tubes.Incubate until the solution clears (~3 minutes). Note: Ensure the solution is completely clearbefore proceeding. 12. While still on the magnetic rack, remove and discard the supernatant. 13. While still on the magnetic rack, add 180 pL of 85% ethanol to each tube. 14. Incubate for 1 minute at room temperature. 15. While still on the magnetic rack, remove and discard the supernatant. 16. Repeat steps 13-15 once for a total of 2 washes. Remove any residual ethanol with a P20. 17. While still on the magnetic rack, air dry the SPRI beads (~30 seconds). Note: Do not over-dry the beads, which can lead to substantial loss in yield. “Cracking” of the beads is a sign of over-drying. 18. Remove the tube(s) from the magnetic rack. Fully resuspend each bead pellet with 50 pL of nuclease-free water. 19. Incubate for 5 minutes at room temperature. 20. Place the tube(s) on the high position of the magnetic rack. Incubate until the solution clears (~2 minutes). 21. While still on the magnetic rack, transfer 50 pL of the supernatant into new 0.2 mL tube(s). Safe stopping point: The size-selected fragmented and end prepped DNA can be stored at 4 °C for up to 18 hours or at -20°C for up to 2 weeks.3.3. Adaptor Ligation
[0224] Adaptors with an Illumina Truseq R2 sequence are ligated to the 5’ end of the fragmented and end prepped DNA. To ligate adaptors: 1. Gather the following items: Ligation adapter, Adapter ligation buffer, Library amp mix, UDI Plate - WT, Adapter ligation enzyme 2. Prepare the Adaptor Ligation Master Mix in a new 1.5 mL tube as shown in Table 28. 2. Mix by pipetting lOx and store on ice.
[0225] 3. Add 50 pL of Adaptor Ligation Master Mix to each tube of purified fragmented and end prepped DNA from section 3.2. Mix by pipetting 1 Ox with a P200 multichannel pipette set to 80 pL. Briefly centrifuge. 4. Place the tube(s) into a thermocycler and run the program shown in Table 29. Note: If the thermocycler's lid can't reach the recommended temperatureof 30 °C, turn the lid heating off. 5. As soon as the program reaches 4 °C, store the adapter ligated DNA on ice and proceed immediately to section 3.4.3.4. Post-Ligation Purification
[0226] Adaptor ligated DNA is purified with a 0.8x SPRI bead cleanup. To purify the ligated DNA: 1. Gather freshly prepared 85% ethanol. 2. Gather room temperature SPRI beads (~90 pL per sublibrary). Note: Ensure the SPRI beads have been equilibrated to room temperature for at least 30 minutes. 3. Vortex the SPRI beads until fully mixed. Add 80 pL of SPRI beads to each tube of adaptor ligated DNA from section 3.3. 4. Vortex the tube(s) for 5 seconds. Briefly centrifuge. 5. Incubate for 5 minutes at room temperature. 6. Place the tube(s) on the high position of the magnetic rack for 0.2 mL tubes. Incubate until the solution clears (~2 minutes). 7. While still on the magnetic rack, remove and discard the supernatant. 8. While still on the magnetic rack, add 180 pL of 85% ethanol to each tube. 9. Incubate for 1 minute at room temperature. 10. While still on the magnetic rack, remove and discard the supernatant.
[0227] 11. Repeat steps 8-10 once for a total of 2 washes. Remove any residual ethanol with a P20. 12. While still on the magnetic rack, air dry the SPRI beads (~2 minutes). 13. Remove the tube(s) from the magnetic rack. Fully resuspend each bead pellet with 23 pL of nuclease-free water. 14. Incubate for 5 minutes at room temperature. 15. Place the tube(s) on the low magnet position of the magnetic rack so the magnet is closer to the bottom of the 0.2 mL tubes. Incubate until the solution clears (~2 minutes). 16. While still on the magnetic rack, transfer exactly 21 pL of the supernatant containing the purified adaptor ligated DNA into new 0.2 mL tube(s). Store on ice. 17. Proceed immediately to section 3.5 (Barcoding Round 4).3.5. Barcoding Round 4
[0228] Purified adaptor ligated DNA is PCR amplified with Illumina Truseq R1 and R2 primers. This indexing PCR generates sequencing libraries and adds i5 / i7 UDIs that act as a fourth cell barcode. To add round 4 barcodes: 1. Centrifuge the UDI Plate - WT at 100 x g for 1 minute. 2. Wipe the surface of the plate with 70% ethanol and allow it to dry. 3. Orient the UDI Plate - WT with the notch on the bottom left. For each sublibrary being processed, choose one unused well of the UDI Plate - WT and record the well position and number for each sublibrary. 4. With a multichannel P20, pierce the seal of the chosen wells in the UDI Plate - WT. 5. With a multichannel P20 and new tips, mix by pipetting 5x then immediately transfer 4 pL from a chosen unused well of the UDI Plate - WT to its corresponding tube of purified adaptor ligated DNA from section 3.4. Note: Only transfer primers from 1 well of the UDI Plate - WT to 1 tube of adapter ligated DNA. 6. If any unused wells remain in the UDI Plate - WT, store the plate at -20°C. Do not reuse wells. 7. Add 25 pL of Library Amp Mix to each tube. Mix by pipetting lOx with a P200 multichannel pipette set to 25 pL. Briefly centrifuge. 8. Determine the number of PCR cycles required for indexing PCR based on the concentration of cDNA added to the fragmentation and end prep reaction as recorded in section 2.5, based on the recommendations shown in Table 30.
[0229] 9. Place the tube(s) into a thermocycler and run the program shown in Table 31.Note: If processing sublibraries with different cDNA concentrations, they should be amplified in separate thermocyclers according to the recommendations shown in Table 30. Safe stopping point: Prior to size selection, sequencing libraries can be stored at 4 °C for up to 18 hours.
[0230] The fourth barcode that tags each sublibrary acts as standard Illumina i7 and i5 indexes. The sequences present in each well of the UDI-WT plate are shown in Table 32. These sequences can be used to demultiplex the WT libraries.3.6. Post-Round 4 Barcoding Size Selection
[0231] The sequencing libraries are size selected with a double sided SPRI cleanup. To size select the sequencing libraries: 1. Gather freshly prepared 85% ethanol. 2. Gather room temperature SPRI beads (~50 pL per sublibrary). Note: Ensure the SPRI beads have been equilibrated to room temperature for at least 30 minutes. 3. Vortex the SPRI beads until fully mixed. Add 30 pL of SPRI beads to each sequencing library tube. 4. Vortex the tube(s) for 5 seconds. Briefly centrifuge. 5. Incubate for 5 minutes at room temperature. 6. Place the tube(s) on the high position of the magnetic rack for 0.2 mL tubes. Incubate until the solution clears (~2 minutes). 7. While still on the magnetic rack, transfer 75 pL of the supernatant containing the DNA into new 0.2 mL tube(s). Discard the tube(s) with bead pellet(s). 8. Add 10 pL of SPRI beads to each tube. 9. Vortex the tube(s) for 5 seconds. Briefly centrifuge. 10. Incubate for 5 minutes at room temperature. 11. Place the tube(s) on the high position of the magnetic rack for 0.2 mL tubes. Incubate until the solution clears (~3 minutes). Note: This may take longer due to the low volume of beads. Ensure the solution is completely clear before proceeding.
[0232] 12. While still on the magnetic rack, remove and discard the supernatant. 13. While still on the magnetic rack, add 180 pL of 85% ethanol to each tube. 14. Incubate for 1 minute at room temperature. 15. While still on the magnetic rack, remove and discard the supernatant.16. Repeat steps 13-15 once for a total of 2 washes. Remove any residual ethanol with a P20.17. While still on the magnetic rack, air dry the SPRI beads (~30 seconds). 18. Remove the tube(s) from the magnetic rack. Fully resuspend each bead pellet with 20 pL of nuclease-free water. 19. Incubate for 5 minutes at room temperature. 20. Place the tube(s) on the low position of the magnetic rack. Incubate until the solution clears (~2 minutes). 21. While still on the magnetic rack, transfer the supernatant into new 0.2 mL tube(s). Store on ice. Safe stopping point: Sequencing libraries can be stored at -20°C for up to 3 months.3. 7. Sequencing Library Quantification
[0233] The concentration and size distribution of the sequencing libraries are measured with fluorescent dyes and capillary electrophoresis. Sequencing libraries can be stored at -20°C for up to 3 months. To quantify the sequencing libraries: 1. Measure the concentration of each purified sequencing library from section 3.6 with, e.g., a Qubit dsDNA HS (High Sensitivity) Assay Kit according to the manufacturer’s instructions. 2. Assess the size distribution of each purified sequencing library with a High Sensitivity DNA Kit on the Agilent Bioanalyzer System or High Sensitivity DI 000 ScreenTape and Reagents on the Agilent TapeStation System according to the manufacturer’s instructions. Note: Samples may need to be diluted to be within the manufacturer's recommended concentration range. Typically, between a 1 :3 to 1 : 10 dilution is appropriate. An exemplary expected size distribution before Illumina sequencing is shown in FIG. 8B. Note: The traces shown in FIG. 8B are representative of typical TapeStation of DNA from indexed sublibraries. There should be a peak between 400- 500 bp. The prominence of the trace is dependent on amount of DNA loaded into the TapeStation. Sublibraries with minor deviations can still produce high quality data. Note: If using a Bioanalyzer, there may be an additional peak present. This typically occurs if products are overamplified, but it should not impact sequencing or data quality (assuming there is still a peak present at 400-500 bp). Do not use this additional peak when estimating amplicon size.Example 2, Sequencing information (e.g., for Illumina sequencing)
[0234] A minimum sequencing depth of 20,000 reads per cell is recommended for whole transcriptome libraries. However, ideal sequencing depth is dependent on the sample type and experimental goals. For example, if trying to capture rare cell types in a heterogenouspopulation, deeper sequencing is required. Alternatively, analyzing a homogenous population, more shallow sequencing may be appropriate. Whole transcriptome sequencing libraries should be diluted and denatured according to the manufacturer’s instructions for the relevant sequencing instrument. If using the Illumina platform, adding 5% PhiX is strongly recommended for optimal sequencing quality.
[0235] Details of the final whole transcriptome sequencing library structure is shown in FIG. 5. Libraries should be sequenced with paired reads using the read structure in Table 33. Read lengths longer than those recommended below are acceptable, but additional bases in read 2 are trimmed by the Parse Analysis Pipeline. The fourth barcode that tags each sublibrary act as standard Illumina i7 and i5 indexes. Refer, e.g., to Table 32 to demultiplex the whole transcriptome libraries.Example 3, Exemplary protocol for cell- or nucleus-specific labeling of the whole transcriptome (up to 10,000 cells)
[0236] For preparing sequencing libraries with fewer cells or nuclei (e.g., up to 10,000 cells or nuclei), the exemplary protocol disclosed in Example 1 can be followed, with certain modifications. The modifications may include, e.g., using 12 wells of the Round 1 plate (instead of, e.g., 48 as in Example 1), and preparing 2 sublibraries (instead of, e.g., 8 as in Example 1), with appropriate changes to the protocol in view of those modifications (e.g., pooling steps after round 1 due to fewer wells and overall volume of pooled cells, potential number of cells / nuclei in each sublibrary, with a maximum of 5,000 cells / nuclei in the present protocol instead of 12,500 in the protocol of Example 1, amount of lysis master mix prepared due to lower volume of pooled cells after Round 3, potential number of PCR cycles in view of lower maximum number of cells / nuclei in sublibraries).Example 4, Exemplary protocols for cell- or nucleus-specific labeling of the whole transcriptome (up to 1,000,000 cells)
[0237] For preparing sequencing libraries with a greater number of cells or nuclei (e.g., up to 1,000,000 cells or nuclei, across up to 96 different biological samples or experimental conditions), the exemplary protocol disclosed in Example 1 can be followed, with certain modifications. The modifications may include, e.g., using 96 wells of the Round 1 plate (instead of, e.g., 48 as in Example 1), and preparing 16 sublibraries (instead of, e.g., 8 as in Example 1), with appropriate changes to the protocol in view of those modifications (e.g., pooling steps after round 1 due to more wells and greater overall volume of pooled cells, potential number of cells / nuclei in each sublibrary, with a maximum of 62,500 cells / nuclei in the present protocol instead of 12,500 in the protocol of Example 1, amount of lysis master mix prepared due to higher volume of pooled cells after Round 3, potential number of PCR cycles in view of higher maximum number of cells / nuclei in sublibraries, e.g., for sublibraries with over 25,000 cells or nuclei, 3 cycles for cells / nuclei with high RNA content such as cell lines, 5 cycles for cells / nuclei with low RNA content such as PBMCs, and 4 cycles for nuclei).
[0238] For preparing sequencing libraries with up to 1,000,000 cells or nuclei, across up to 384 different biological samples or experimental conditions, the exemplary protocol may be further modified, e.g., by using 4 x 96-well Round 1 plates (using all 96 wells of each plate), instead of, e.g., 48 as in Example 1), and preparing 16 sublibraries, with appropriate changes to the protocol in view of those modifications (e.g., pooling steps after round 1 due to more wells and greater overall volume of pooled cells, potential number of cells / nuclei in each sublibrary, with a maximum of 62,500 cells / nuclei in the present protocol instead of 12,500 in the protocol of Example 1, amount of lysis master mix prepared due to higher volume of pooled cells after Round 3, potential number of PCR cycles in view of higher maximum number of cells / nuclei in sublibraries, e.g., for sublibraries with over 25,000 cells or nuclei, 3 cycles for cells / nuclei with high RNA content such as cell lines, 5 cycles for cells / nuclei with low RNA content such as PBMCs, and 4 cycles for nuclei). In addition, the four Round 1 plates can be identical, but can be loaded into the Round 2 plate in a way that allows the samples to be distinguished based on the combination of Round 1 and Round 2 barcodes (e.g., by loading the cells or nuclei from each of the four Round 1 plates into a specific subset of wells within the Round 2 plate).Example 5, Exemplary protocol for cell or nuclei fixation for whole transcriptome sequencing librariesWorkflow
[0239] From a single cell or nuclei suspension, the present protocols generate fixed and permeabilized cells or nuclei that are ready for use in Whole Transcriptome or other labeling proto. 1.5 mL tube based workflows are recommended when processing <12 samples at a time. If processing >12 samples at a time, we recommend mid-throughput plate-based workflows, which streamline fixation when processing more samples. If processing >48 samples, we recommend using high-throughput plate-fixation protocols. Fixation maintains cell / nuclear structure, prevents RNA degradation, and locks the RNA inside the cells or nuclei, which are crucial for downstream processing with Evercode split-pool combinatorial barcoding technology (FIGS. 9A-9B). Because fixed samples are also stable for up to 6 months at -80°C, these protocols provide flexibility by separating sample collection from library preparation. It also enables samples to be stored and batched after fixation so they can be processed through library preparation together, reducing the potential of batch effects.
[0240] FIG. 10 provides an overview of the fixation workflows. Between 100,000 and 1 million cells or nuclei can be fixed in a single reaction. Note that more than 100,000 cells or nuclei may need to be fixed to fully utilize the capacity of the downstream Evercode kits. Details of the total and hands-on time required for the cell or nuclei fixation workflow are shown in Table 34.Sample Input
[0241] These protocol begins with previously prepared single cell or nuclei suspensions. We recommend suspensions with >70% viability (ideally above 90%) and <5% aggregation / debris. If cells or nuclei were previously frozen, ensure the suspension is completely thawed and in suspension before beginning fixation. We recommend minimizing the length of time samples are stored on ice prior to fixation, as it can negatively impact results. Between 100,000 and 1 million cells or nuclei can be fixed in a single reaction. However, we recommend using the highest number available up to 1 million total. Exceeding 1 million cells or nuclei in a single fixation will result in substantially elevated doublet rates. The minimum input into fixation should also be determined based on how the samples will be processed downstream. Table 35 provides guidance on the post-fixation concentrations needed for downstream kits. However, more or less sample input may be required depending on the exact experimental design. Note that retention during fixation varies typically between 40-60%, and some cells or nuclei will be lost when freezing and thawing fixed samples, typically between 5-15%. The final concentration of cells or nuclei post-fixation is also influenced by the resuspension volume used in Step 18 of Section 2. These factors should all be taken into account when determining how much sample input is needed for fixation.Avoiding RNase Contamination
[0242] Standard precautions should be taken to avoid introducing RNases into samples or reagents throughout the workflow. Always wear proper laboratory gloves and use aseptic technique. Although RNases are not inactivated by ethanol or isopropanol, they are inactivated by products such as RNaseZap RNase Decontamination Solution (Thermo Fisher Scientific). These can be sprayed on benchtops and pipettes. Nuclease-free, filtered pipette tips should be used to reduce RNase contamination from pipettes.Cell Detachment
[0243] If using adherent cell line samples, we recommend TrypLE Express Enzyme (IX), phenol red (Thermo Fisher Scientific). Due to high RNase activity, we do not recommend dissociation with standard trypsin, which may reduce gene and transcript detection.Cell Strainers
[0244] To maximize cell / nuclei retention with cell strainers, press the pipette tip directly against the mesh. Ensure ample pressure is applied to hold contact between the tip and the strainer to force liquid through in ~1 second. A cell strainer with an appropriately sized mesh should be used throughout the protocol. Although 30-40 pm is appropriate for many cell types or nuclei, the mesh size should be chosen based on your sample type.Maximizing Cell / Nuclei Recovery
[0245] It is critical to thoroughly resuspend the cells or nuclei after centrifugation throughout the protocol. Resuspend by slowly and repeatedly pipetting until no clumps are visible. Ideally this should be verified with microscopy. To minimize cell / nuclei loss from adherence to tubes, carefully pipette along the bottom and sides of tubes. We do not recommend wide bore pipette tips as they make it difficult to resuspend pellets adequately.Reagent Stability
[0246] Cell / Nuclei Fixation Reagents should not be frozen and thawed more than 3 times. If a kit is going to be used more than 4 times, the reagents should be aliquoted into nuclease- free 1.5 mL tubes and stored at -20°C until use. We do not recommend making single use aliquots to minimize the impact of evaporation during storage. To avoid pipetting < 2 pL of RNase Inhibitor, we do not recommend preparing less than 2 reactions of Storage Master Mix. If required, the master mix can be prepared without DMSO, split into aliquots, and stored at - 20°C until use. DMSO should be added prior to use in the protocol to a final concentration of 5%. With the exception of the partial Storage Master Mix described above, reagent master mixes should be made fresh and used the same day.Storage of Fixed Samples
[0247] Fixed samples can be stored at -80 °C for up to 6 months. Fixed samples should not be refrozen after thawing. When possible, we recommend splitting samples into aliquotsafter fixation in Section 2 Step 21. Aliquots should be at least 50 pL when stored in 1.5 mL tubes.1.1. Block Tubes with BSA
[0248] Although not required, blocking tubes with B SA can increase cell / nuclei retention. When Protein LoBind tubes are not available, we recommend blocking tubes, especially for samples with low cell inputs or cells prone to aggregation. To block tubes: 1. Prepare a fresh 1% BSA as follows, depending on the number of samples being processed: for 1 sample: 2.6 mL nuclease-free water, 400 pL Gibco Bovine Albumin Fraction V (7.5% solution), total volume: 3 mL; for 12 samples: 31.2 mL nuclease-free water, 4.8 mL Gibco Bovine Albumin Fraction V (7.5% solution), total volume: 36 mL. 2. For each sample, fill two 1.5 mL tubes with 1.5 mL of 1% BSA and cap the tubes. 3. Invert once to fully coat the tubes. 4. Incubate the tubes for 30 minutes at room temperature. 5. Remove the 1% BSA with a Pl 000 and discard. 6. Remove any remaining solution from the bottom of the tube with a P200. 7. With the caps removed, air dry the tubes for 30 minutes in a biosafety cabinet at room temperature. 8. Proceed to Section 1.2 or store capped BSA-coated tubes at 4 °C for up to 4 weeks.1.2 Prepare master mixes.
[0249] Master mixes should be prepared just prior to fixation. To prepare master mixes: 1. Fill a bucket with ice. Gather the following items: Prefixation buffer, Storage buffer, Fixative Solution A, Fixative Solution B, Permeabilization Solution, Fix and Perm Stop buffer, DMSO, RNase inhibitor, Prefixation Enhancer. 2. Prepare the Cell or Nuclei Prefixation Master Mix in a new tube as follows: for 1 sample: Prefixation buffer (203.5 pL); RNase Inhibitor (2.75 pL), Prefixation enhancer (13.75 pL), total volume (220 pL); for 12 samples: Prefixation buffer (2.44 mL); RNase Inhibitor (33 pL), Prefixation enhancer (165 pL), total volume (2.64 mL). Mix thoroughly by pipetting, carefully avoiding to create bubbles, and store on ice. 3. Prepare the Cell Fixative Master Mix in a new tube as follows: for 1 sample: Fixative Solution A (36 pL); Fixative Solution B (36 pL), total volume (72 pL); for 12 samples: Fixative Solution A (432 pL); Fixative Solution B (432 pL), total volume (864 pL). Mix thoroughly by pipetting and store on ice. For Nuclei fixation, a Nuclei Fixative solution is used instead of the Cell Fixative Master Mix used with cells.
[0250] 4. Prepare the Cell / Nuclei Storage Master Mix in a new tube as shown in Table36. Mix thoroughly by pipetting and store on ice. Proceed immediately to Section 2.2.1. Cell or Nuclei Fixation
[0251] After the initial centrifugation to remove the buffer / medium from the single cell / nuclei suspension, cells or nuclei are transferred to Cell or Nuclei Prefixation Master Mix. Reagents are added to fix and permeabilize cells, and then to stop these reactions. Cells are resuspended in Cell Storage Master Mix and stored at -80°C or processed immediately with a downstream Evercode kit (e.g., according to the protocols of any of Examples 1-4).
[0252] To fix cells / nuclei: 1. Cool the centrifuge with a swinging bucket rotor to 4°C. 2. Fill a bucket with ice. 3. Prepare a hemocytometer, flow cytometer, or other cell counting device. 4. Place a Mr. Frosty Freezing Container at room temperature. 5. Count the cells in the single cell suspension with a hemocytometer or alternative counting device and record the count. Keep cells on ice during counting and work quickly to minimize time on ice prior to fixation. 6. Transfer 100,000 to 1 million cells from each sample into a Protein LoBind 1.5 mL tube (or a BSA coated tube if prepared in Section 1.1). 7. Centrifuge the tubes in a swinging bucket rotor for 5-10 minutes at 200-500 x g at 4°C. Note: Use of a fixed-angle rotor in this protocol will lead to substantial cell loss. Note: Ideal centrifugation speed and duration should be determined for each sample type to optimize retention and resuspension efficiencies. Move quickly and handle the samples gently to avoid dislodging the pellet, which will impact data quality.
[0253] 8. Slowly aspirate then discard the supernatant. 9. Fully resuspend each pellet in187.5 pL of Cell / Nuclei Prefixation Master Mix. 10. Pipette each sample through a cell strainer into a new 1.5 mL tube and store on ice. Note: Do not directly touch the mesh of cell strainer(s). Note: To ensure that all of the liquid passes through the strainer, press the tip of the pipette against the filter and steadily depress down the pipette plunger. All of the liquid should pass through the strainer in ~1 second. 11. Add 62.5 pL of Cell Fixative Master Mix or NucleiFixative Solution to each tube and mix immediately by pipetting exactly 3x. Note: Do not perform additional mixing at this step. 12. Incubate on ice for 10 minutes. 13. Add 20 pL of Permeabilization Solution to each tube. Immediately mix thoroughly by pipetting 3x with a P200 set to 180 pL. 14. Incubate on ice for 3 minutes. 15. Mix the Fix and Perm Stop Buffer by inverting the tube 5x. Do not vortex. 16. Add 250 pL of Fix and Perm Stop Buffer to each tube and gently pipette mix 3x. 17. Centrifuge in a swinging bucket rotor for 5-10 minutes at 200-500 x g at 4 °C. 18. With a P1000 set to 500 pL, slowly aspirate then discard the supernatant. 19. Fully resuspend each pellet in 50-100 pL Cell or Nuclei Storage Master Mix and store on ice.
[0254] Note: Choose a resuspension volume appropriate for the experimental design and downstream Evercode kits or protocols. 20. Pipette the sample through a cell strainer into a new 1.5 mL tube and store on ice. 21. While minimizing time on ice, count the number of cells or nuclei in the sample with a hemocytometer or alternative counting device and record the cell / nuclei count. Note: Downstream Evercode processing can be streamlined by aliquoting samples at this step. 22. Proceed to the appropriate user guide if immediately processing samples with an Evercode kit. Otherwise, proceed to the next step. 23. Store tubes in a Mr. Frosty Freezing Container (or equivalent device) at -80°C, according to the manufacturer’s instructions. Note: Storing samples directly in the freezer without controlled cooling may lead to cell damage and compromise data quality. Safe stopping point: Samples are stable for up to 6 months at -80°C.Centrifugation Optimization
[0255] When using Evercode Fixation kits for the first time or when testing a new sample type, we recommend optimizing centrifugation conditions. This appendix provides guidelines for optimization, suggestions for common sample types, and an example experiment to optimize centrifugation speed. Note that physical properties of cells may change after the fixation process, which requires centrifugation conditions to be optimized during fixation.
[0256] A range of centrifugation speeds should be tested to identify a speed that maximizes sample retention and permits thorough resuspension into a high quality single cell solution. Cells or nuclei should be examined under a microscope before and after centrifugation to calculate cell / nuclei retention and assess any aggregation or morphological changes. After determining the appropriate centrifugation conditions, we recommend using the same speed and duration throughout this and downstream Evercode User Guides.Typical Sample Retention
[0257] Across a range of samples, cell / nuclei retention post-fixation typically varies between 40-60% of the initial input. Retention is impacted by sample type, sample preparation method, centrifugation conditions, and sample handling.Speed
[0258] Increasing centrifugation speeds can improve retention, but high speeds can complicate the pellet resuspension and damage or even lyse cells / nuclei. The optimal centrifugation speed will generally achieve a greater than 50% retention through the centrifugation step while maintaining membrane integrity. Centrifugation speed depends on cell / nuclei size. Smaller cells / nuclei need faster speeds, and larger cells / nuclei need slower speeds.Duration
[0259] If cells or nuclei are damaged by increased centrifugation speed, centrifugation duration can be adjusted to increase retention without cell / nuclei damage.Temperature
[0260] For most sample types, the centrifugation should be done at 4 °C. However, some sample types may require different temperatures to maximize cell viability or nuclei quality prior to fixation. For example, isolated dendritic cells, myeloid-derived suppressor cells, and macrophages are sensitive to cold temperatures and should be processed at 25 °C until the addition of the Cell Fixative Master Mix. After fixation, the final centrifugation step in this User Guide and all centrifugation steps in the Evercode User Guide should be done at 4 °C to maintain cell / nuclei and RNA integrity.Aggregates After Centrifugation
[0261] If the pellet cannot be resuspended back into a single cell / nuclei suspension and there are aggregates where there were previously not, this is an indication that the sample may have been over centrifuged. Aggregates may also be an indication of insufficient pipette mixing. Gently resuspend the pellet by slowly and repeatedly pipetting until no clumps are visible. This can be visually inspected via microscopy. Aggregates at this stage may also be a result of the sample preparation method used. If none of the above have been successful inremoving the aggregates, a filtering step may help remove aggregates or the sample preparation may require additional optimization.Debris After Centrifugation
[0262] Samples with viability of <70% or with low quality nuclei may result in excessive debris in your fixed sample. Ideally, measures should be taken to optimize sample quality prior to proceeding into fixation. If a sample with minimal debris has significant debris after centrifugation, this may be an indication that the sample has lysed due to over centrifugation and / or overly aggressive resuspension. The centrifugation speed should be reduced and / or pellets should be less aggressively pipetted.Recommendations for Common Sample Types
[0263] These centrifugation conditions can be used as a starting point for common sample types. However, as samples vary, we still recommend following using the optimization protocol below. Sample type: HEK293, 3T3, and other cell lines: 200 x g, 10 min, 4 °C. Sample type: PBMCs: 200-400 x g, 10 min, 4 °C. Sample type: mammalian nuclei: 300-400 x g, 10 min. 4 °C.Centrifugation Optimization Method
[0264] When using Evercode Fixation kits for the first time or when testing a new sample type, we recommend using 1-2 samples to optimize centrifugation conditions prior to processing samples of interest. When this is not possible, centrifugation conditions can be determined while fixing samples of interest.
[0265] Different centrifugation conditions can be tested, for example, by starting centrifugation at a low speed, but retaining the supernatants after each spin. These supernatants are then centrifuged again to recover additional cells. After resuspension, each pellet should be assessed with microscopy to count cells / nuclei, quantify debris, and assess aggregation. Resuspended pellets of high quality (minimal debris, minimal aggregation, and minimal evidence of cell / nuclei damage) are pooled and can be used with downstream Evercode kits or protocols.
[0266] In an example, a sample is first centrifuged at 200 x g for 10 minutes. The pellet is resuspended in Storage Master Mix, and the first supernatant is centrifuged again at 300 x g for 10 minutes. The second pellet is resuspended in Storage Master Mix, and the secondsupernatant is centrifuged again at 400 x g for 10 minutes. This final, third pellet is resuspended in Storage Master Mix and the third supernatant is discarded. The three resuspended pellets are then counted with a hemocytometer. In this example, the cells centrifuged at 400 x g are aggregated with significant debris, so this resuspended pellet should be discarded. Conversely, the cells / nuclei centrifuged at 200 x g and 300 x g were high quality, so they were pooled together. Once pooled, this sample has -50% retention. These results suggest that this sample type should be centrifuged at 300 x g in, e.g., Evercode Cell / Nuclei Fixation and Evercode workflows.
[0267] If desired, 1-4 million cells or nuclei can be fixed in a single reaction. However, this requires the reagent volume to be scaled up 4x, which reduces the total number of samples that can be fixed with a 12 reaction kit to 3 samples.Mid-Throughput and High-Throughput Plate Workflows
[0268] Plate-based versions of the workflows can also be performed, e.g., for processing between 12 and 48 samples (mid-throughput) or more than 48 samples (high throughput). Between 100,000 and 1 million cells or nuclei can be fixed in a single reaction using these protocols.Example 6, Exemplary protocol for low input cell or nuclei fixation
[0269] From single cell or nuclei suspensions, low numbers of fixed and permeabilized cells / nuclei (e.g., less than 100,000) can be generated that are ready for use in downstream assays. This protocol can be used to fix cells or nuclei for various single-cell applications, including single-cell RNA-seq (e.g., Evercode WT), immune profiling applications (e.g., TCR or BCR profiling), or genomic DNA / chromatin accessibility assays (e.g., DNase-seq or ATAC- seq). An overview of the workflow is shown in FIG. 11.
[0270] The workflow is designed to efficiently process between 10,000 and 100,000 cells or nuclei, accommodating up to 12 samples or as many as 96 samples simultaneously. The fixation protocol preserves cell structure, prevents RNA degradation, and locks RNA inside the cells, essential for downstream processing with Evercode's split-pool combinatorial barcoding technology. The workflow allows for high levels of cell / nuclei retention with low numbers of inputted cells / nuclei. Fixed samples are stable for up to 4 months at -80 °C, providing flexibility by decoupling sample collection from library preparation. This allows samples to be stored and batched post-fixation, enabling simultaneous library preparation andminimizing batch effects. The workflow facilitates parallel fixation of multiple samples, streamlining the process when handling up to 96 samples at a time.
[0271] The high retention levels achieved with this protocol using low numbers of inputted cells or nuclei are due in part to the use of magnetic beads capable of binding cells or nuclei. In particular, in the protocol the cells / nuclei are first fixed and permeabilized (in plates or tubes), the permeabilization stopped, and the cells / nuclei directly frozen (i.e., without removing the permeabilization stop solution) for up to 4 months at -80 °C. The frozen cells / nuclei can then later be thawed, bound to the magnetic beads, and localized within the plate / tube by applying a magnet. The supernatant is then removed and the cells / nuclei resuspended in Storage Buffer prior to barcoding. If fixed samples were stored at -80 °C, they will need to be thawed before the capture step. It is recommended to count the captured cells / nuclei using fluorescent-based dyes on an automated cell counter or trypan blue on a hemocytometer prior to input into barcoding. Alternatively, the number may be extrapolated based on the previous count, assuming 65% retention for small cells / nuclei and 80% retention for large cells / nuclei.Example 7, Duplexed barcodes increase ligation efficiency of oligonucleotides as measured by quantitative PCR
[0272] A ligation experiment was performed in which two barcode-containing oligonucleotides were ligated to one another in the presence of a linker strand with complementarity to both of the barcode-containing strands, as well as a additional strand with complementarity to one of the barcode-containing strands only (FIGS. 12A-12B). Different configurations of the complementary and linking strands were assessed, in particular with respect to whether the barcode sequences (BC1 and BC2) were left in an unannealed state (FIG. 12A) or in a duplexed (FIG. 12B) state. Ligation was quantified by qPCR after ligation reactions of 5 minutes (FIG. 12C) or 30 minutes (FIG. 12D). The results (FIG. 12C) showed that when either BC2 alone was duplexed (conditions E and F in FIG. 12C), or when both BC1 and BC2 were duplexed (conditions G and H in FIG. 12C), ligation was substantially more efficient than when neither BC1 nor BC2 were duplexed (conditions A and B in FIG. 12C) or when only barcode BC1 was duplexed (conditions C and D in FIG. 12C).
[0273] The results demonstrate that duplexing barcodes (i.e., using a linker strand that covers both the 3’ region of the barcode-containing strand and the barcode itself) substantially increases ligation efficiency after both 5 and 30 minutes of ligation. The effect at 5 minuteswas as significant as after 30 minutes, indicating that ligation steps in a protocol using nucleic acid tags with duplexed barcodes can be shortened to, e.g., 5, 10, or 15 minutes, while maintaining a high level of ligation efficiency.Example 8, Varying the lengths of different linker constant regions
[0274] Quantitative PCR was used to assess the effects of varying lengths of duplexed linkers and their constant regions on ligation efficiency. In particular, the lengths of the regions marked A and B / C in FIGS. 13A-13B were varied, and their effects on ligation assessed by qPCR. In one experiment, region A was held constant at 4 nucleotides, and region B / C varied between 2, 4, 6, 10, or 15 nucleotides in region B and either 0 or 5 nucleotides in region C (FIG. 13A). In another experiment, region B / C was held at 4 nucleotides (with a B region having 4 nucleotides and a C region with 0 nucleotides) or 15 nucleotides (with a B region of 4 nt and C region of 11 nt), and A was set to 2, 7, or 10 nucleotides.
[0275] The results, shown in FIG. 13C, indicated, first, that when the A region was limited to 4 (or 2) nucleotides, ligation efficiency was significantly lower than when the A region was 7 or 10 nucleotides in length, regardless of the length of the combined BC region. However, when the A region was 7 nucleotides or longer, it was possible to shorten region B (e.g., to 4 nucleotides) while still achieving a high level of ligation efficiency.
[0276] In another experiment, the length of A varied between 4, 6, 7, or 10 nucleotides, and the B / C region was held at 4 nucleotides, and ligation efficiency was assessed by qPCR using two different barcode sequences. The barcode sequences were duplexed in all conditions.
[0277] The results showed that, with both of the barcode sequences, an A region of 6 nucleotides yielded a lower Cq value (and was therefore more efficient) than A regions of 4, 7, or 10 nucleotides (Table 37).
[0278] All references cited herein are incorporated by reference to the same extent as if each individual publication, database entry (e.g., Genbank sequences or GenelD entries), patent application, or patent, was specifically and individually indicated incorporated by reference in its entirety, for all purposes. This statement of incorporation by reference is intended by Applicants, pursuant to 37 C.F.R. § 1.57(b)(1), to relate to each and every individual publication, database entry (e.g., Genbank sequences or GenelD entries), patent application, or patent, each of which is clearly identified in compliance with 37 C.F.R. § 1.57(b)(2), even if such citation is not immediately adj acent to a dedicated statement of incorporation by reference. The inclusion of dedicated statements of incorporation by reference, if any, within the specification does not in any way weaken this general statement of incorporation by reference. Citation of the references herein is not intended as an admission that the reference is pertinent prior art, nor does it constitute any admission as to the contents or date of these publications or documents.
[0279] While the invention has been particularly shown and described with reference to a preferred embodiment and various alternate embodiments, it is understood by persons skilled in the relevant art that various changes in form and details can be made therein without departing from the spirit and scope of the invention
Claims
CLAIMS:
1. A nucleic acid tag for barcoding nucleic acid molecules in a cell or nucleus, the tag comprising: a) a first strand comprising: i) a first strand barcode sequence; ii) a first strand 5’ hybridization sequence located 5’ of the first strand barcode sequence; and iii) a first strand 3’ hybridization sequence located 3’ of the first strand barcode sequence; and b) a second strand comprising: i) a second strand barcode sequence, wherein the second strand barcode sequence is complementary to the first strand barcode sequence; ii) a second strand hybridization sequence located 5’ of the second strand barcode sequence, wherein the second strand hybridization sequence is complementary to the first strand 3’ hybridization sequence; and iii) a second strand overhang sequence located 5’ of the second strand hybridization sequence; wherein the first strand and second strand are annealed such that the nucleic acid tag comprises: i) a double-stranded central region comprising the first strand barcode sequence annealed to the second strand barcode sequence and the first strand 3’ hybridization sequence annealed to the second strand hybridization sequence; and ii) a single-stranded overhang located at each end of the nucleic acid, wherein one of the two overhangs comprises the first strand 5’ hybridization sequence, and the other overhang comprises the second strand overhang sequence.
2. The nucleic acid tag of claim 1, wherein the first strand 3’ hybridization sequence and / or the second strand hybridization sequence are from 4-10 nucleotides long.
3. The nucleic acid tag of claim 2, wherein the first strand 3’ hybridization sequence is 4, 5, 6, 7, 8, 9, or 10 nucleotides long.
4. The nucleic acid tag of claim 2 or 3, wherein the second strand hybridization sequence is 4, 5, 6, 7, 8, 9, or 10 nucleotides long.
5. The nucleic acid tag of any of claims 2-4, wherein the first strand 3’ hybridization sequence and the second strand hybridization are each 6 nucleotides long.
6. The nucleic acid tag of any one of claims 1-5, wherein the first strand 5’ hybridization sequence and / or the second strand overhang sequence are from 5 to 10 nucleotides long.
7. The nucleic acid tag of claim 6, wherein the first strand 5’ hybridization sequence is 5, 6, 7, 8, 9, or 10 nucleotides long.
8. The nucleic acid tag of claim 6 or 7, wherein the second strand overhang sequence is 5, 6, 7, 8, 9, or 10 nucleotides long.
9. The nucleic acid tag of any one of claims 1-8, wherein the first strand 5’ hybridization sequence and the second strand overhang sequence are each 6 nucleotides long.
10. The nucleic acid tag of any one of claims 1-9, wherein the first strand 3’ hybridization sequence comprises the sequence shown as SEQ ID NO: 1 or as SEQ ID NO:2.
11. The nucleic acid tag of any one of claims 1-10, wherein the first strand 5’ hybridization sequence comprises the sequence shown as SEQ ID NO: 3 or as SEQ ID NO: 4.
12. The nucleic acid tag of any one of claims 1-11, wherein the second strand hybridization sequence comprises the sequence shown as SEQ ID NO:5 or as SEQ ID NO:6.
13. The nucleic acid tag of any one of claims 1-12, wherein the second strand overhang sequence comprises the sequence shown as SEQ ID NO:7 or as SEQ ID NO:8.
14. The nucleic acid tag of any one of claims 1-13, wherein the first strand barcode sequence and / or the second strand barcode sequence is at least 8 nucleotides long.
15. The nucleic acid tag of any one of claims 1-14, wherein the first strand barcode sequence and the second strand barcode sequence are each 8 nucleotides long.
16. The nucleic acid tag of any one of claims 1-15, wherein one or more of the first strand or second strand barcode sequences comprise a sequence selected from the group consisting of the barcode sequences present within any of the sequences of Tables 2-7 and the barcode sequences present within any of the sequences shown as SEQ ID NOS: 11-590.
17. The nucleic acid tag of any one of claims 1-16, wherein the first strand comprises a sequence selected from the group consisting of the barcode sequences present within any of the sequences of Tables 2 and 4 and the barcode sequences present within any of the sequences shown as SEQ ID NOS: 11-105 and SEQ ID NOS: 201-295.
18. The nucleic acid tag of any one of claims 1-17, wherein the second strand comprises a sequence selected from the group consisting of the sequences shown in Tables 3 and 5, and the sequences shown as SEQ ID NOS: 106-200 and SEQ ID NOS: 296-400.
19. The nucleic acid tag of any one of claims 1-18, wherein the first and / or second strands are DNA molecules.
20. The nucleic acid tag of claim 19, wherein the first strand further comprises one or more elements located 5’ of the first strand barcode sequence selected from the group consisting of a random nucleotide sequence to prevent counting of PCR duplicates, a capture agent, and a next generation sequencing (NGS) adapter sequence.
21. The nucleic acid tag of claim 20, wherein the first strand comprises a sequence selected from the group consisting of the sequences shown in Table 4 and the sequences shown as SEQ ID NOS: 201-295.
22. A set of nucleic acid tags for labeling nucleic acids within a plurality of cells or nuclei, the set comprising a plurality of the nucleic acid tags of any one of claims 1-21, wherein all of the nucleic acid tags in the set comprise the same first strand 3’ hybridization sequence, the same first strand 5’ hybridization sequence, the same second strand hybridization sequence, and / or the same second strand overhang sequence; and wherein a plurality of distinct first strand barcode sequences and a plurality of distinct second strand barcode sequences are present among the nucleic acid tags of the set.
23. The set of nucleic acid tags of claim 22, wherein the nucleic acid tags of the set are distributed into a plurality of aliquots, and wherein the first strand barcode sequences and the second strand barcode sequences present among the nucleic acid tags within each of the aliquots are aliquot-specific.
24. The set of nucleic acid tags of claim 23, wherein all of the first strand barcode sequences and all of the second strand barcode sequences present among the nucleic acid tags within any individual aliquot of the plurality of aliquots are the same.
25. The set of nucleic acid tags of any one of claims 22-24, wherein the plurality of aliquots are distributed among the wells of a multi-well plate.
26. The set of nucleic acid tags of claim 25, wherein the plurality of aliquots comprises 96 aliquots distributed into the wells of a 96-well plate.
27. The set of nucleic acid tags of claim 26, wherein the set of nucleic acid tags comprises at least 96 distinct first strand barcode sequences and 96 distinct second strand barcode sequences.
28. The set of nucleic acid tags of claim 27, wherein each of the at least 96 distinct first strand barcode sequences and the at least 96 distinct second strand barcode sequences is present in only one well of the 96-well plate.
29. The set of nucleic acid tags of claim 28, wherein the set of nucleic acid tags comprises 96 distinct first strand barcode sequences, such that only one distinct first strand barcode sequence is present in each well of the 96-well plate, and 96 distinct second strand barcode sequences, such that only one distinct second strand barcode sequence is present in each well of the 96-well plate.
30. The set of nucleic acid tags of claim 28 or 29, wherein the set of nucleic acid tags comprises at least 192 distinct first strand barcode sequences and at least 192 distinct second strand barcode sequences, and wherein two or more distinct first strand barcode sequences and two or more distinct second strand barcode sequences are present in each well of the 96-well plate.
31. A collection of nucleic acid tags for labeling nucleic acids with cell- or nucleus-specific tags, the collection comprising: two or more sets of nucleic acid tags according to any one of claims 22-30, wherein the second strand overhang sequence present in a first set of the two or more sets of nucleic acid tags is complementary to the 5’ hybridization sequence present in a second set of the two or more sets of nucleic acid tags.
32. The collection of nucleic acid tags of claim 31, wherein the nucleic acid tags of each of the two or more sets are distributed into some or all of the wells of a unique multi-well plate; and wherein the first and second strand barcode sequences present among the nucleic acid tags of each set are well-specific.
33. A collection of nucleic acid tagging agents for labeling nucleic acids with cell- or nucleus-specific tags, the collection comprising: a set of nucleic acid tags according to any one of claims 22-30, and a set of reverse transcription (RT) primers each comprising: (i) a poly(T) sequence or a random sequence; (ii) an RT barcode sequence; and (iii) a 5’ overhang comprising a 5’ overhang sequence; wherein the second strand overhang sequence present in the set of nucleic acid tags is complementary to the 5’ overhang sequence present in the set of RT primers.
34. The collection of nucleic acid tags of claim 33, wherein the nucleic acid tags of the set of nucleic acid tags are distributed into some or all of the wells of a first multi-well plate; wherein the nucleic acid tags of the set of RT primers are distributed into some or all of the wells of a second multi-well plate; and wherein the first and second strand barcode sequences present among the set of nucleic acid tags and the RT barcode sequences present among the set of RT primers are well-specific.
35. The collection of nucleic tags of claim 33 or claim 34, wherein one or more of the RT primers present within the set of RT primers comprises a sequence selected from the group consisting of the sequences shown in Tables 6 and 7 and the sequences shown as SEQ ID NOS: 401-590.
36. A method of labeling nucleic acids with cell- or nucleus-specific tags, the method comprising:(a) providing a plurality of fixed and permeabilized cells or nuclei, each comprising a plurality of RNA molecules;(b dividing the plurality of cells or nuclei into a first plurality of aliquots, wherein each aliquot comprises more than one cell or nucleus;(c) generating complementary DNA (cDNA) molecules by reverse transcribing RNA molecules within the cells or nuclei of the first plurality of aliquots, wherein the RNA molecules are reverse transcribed using reverse transcription (RT) primers each comprising: (i) a poly(T) sequence or a random sequence; (ii) an RT barcode sequence, wherein the RT barcode sequences present within the RT primers are specific to each aliquot within thefirst plurality of aliquots; and (iii) a 5’ overhang comprising a 5’ overhang sequence;(d) pooling the cells or nuclei from the first plurality of aliquots;(e) tagging the cDNA molecules within the pooled cells or nuclei from the first plurality of aliquots with one or more nucleic acid tags, thereby generating tagged cDNA molecules, by performing steps (i) through (iii) one or more times:(i) dividing the pooled cells or nuclei into an additional plurality of aliquots;(ii) coupling nucleic acid tags of any one of claims 1-21 to the cDNA molecules within the cells or nuclei of the additional plurality of aliquots; and(iii) pooling the cells or nuclei from the additional plurality of aliquots;(f) lysing the cells or nuclei to release the tagged cDNA molecules and produce a lysate comprising the released tagged cDNA molecules; and(g) isolating the released tagged cDNA molecules.
37. The method of claim 36, wherein step (e)(ii) comprises coupling the cDNA molecules with the set of nucleic acid tags of any one of claims 22-30.
38. The method of claim 36 or 37, wherein the reverse transcription of step (c) and the tagging of step (e)(ii) are performed using the collection of nucleic acid tagging agents of any one of claims 33-35.
39. The method of any one of claims 36-38, wherein steps (e)(i) through (e)(iii) are performed two or more times, thereby generating repeatedly tagged cDNA molecules, and wherein the two or more times are performed using the collection of nucleic acid tags of claim 31 or 32.
40. The method of any one of claims 36-39, wherein the released tagged cDNA molecules are isolated using a binding agent, such that the isolated tagged cDNA molecules are bound to the binding agent.
41. The method of claim 40, wherein the isolated tagged cDNA molecules comprise biotin, and wherein the binding agent comprises streptavidin-coated magnetic beads.
42. The method of any one of claims 36-41, wherein the coupling of the nucleic acid tags to the cDNA molecules in step (e)(ii) comprises ligation.
43. The method of claim 42, wherein the ligation is performed for a duration of 15 minutes or less.
44. The method of any one of claims 36-43, further comprising:(h) generating second strands of the isolated tagged cDNA molecules to produce double-stranded tagged cDNA molecules.
45. The method of claim 44, wherein the second strands are generated in step (h) using a template switching oligo (TSO) comprising a TSO adapter sequence.
46. The method of any one of claims 36-45, further comprising: amplifying the double-stranded tagged cDNA molecules.
47. The method of claim 46, wherein the second strands are generated in step (h) using a template switching oligo (TSO) comprising a TSO adapter sequence, wherein the final nucleic acid tags coupled to the tagged cDNA molecules during the one or more times that steps (e)(i) to (e)(iii) are performed comprise an NGS adapter sequence, and wherein one or more rounds of the amplification of the double-stranded tagged cDNA molecules are performed using amplification primers specific to the TSO adapter sequence and the Tag adapter sequence.
48. The method of claim 46 or 47, wherein, prior to lysing the cells or nuclei in step (f), the additional plurality of aliquots is divided into a plurality of sublibraries, and wherein the amplifying of the double-stranded tagged cDNA molecules is performed using amplification primers comprising a sublibrary-specific index sequence.
49. The method of any one of claims 46-48, further comprising: preparing a sequencing library using the amplified double-stranded tagged cDNA molecules.
50. The method of claim 49, wherein preparing the sequencing library comprises fragmenting the amplified double-stranded tagged cDNA molecules and appending a postfragmentation adapter comprising a post-fragmentation adapter sequence to the fragment ends.
51. The method of claim 49 or 50, further comprising: sequencing the sequencing library.
52. The method of claim 51, further comprising: grouping the sequencing reads obtained in the sequencing of claim 51 according to one or more features selected from the group consisting of RT barcode sequences, first strand barcode sequences, second strand barcode sequences, index sequences, cDNA sequences, and series or combinations of any one or more of RT barcode sequences, first strand barcode sequences, second strand barcode sequences, and index sequences, and cDNA sequences.
53. The method of claim 52, further comprising: using the grouped sequencing reads to determine the individual cell or nucleus from among the plurality of cells or nuclei from which a given tagged cDNA molecule originated.
54. The method of any one of claims 51-53, further comprising: mapping the sequencing reads obtained in the sequencing of claim 51 to a reference genome.
55. The method of any one of claims 1-54, wherein the cells or nuclei comprise mammalian cells or nuclei.
56. The method of claim 55, wherein the mammalian cells or nuclei comprise human cells or nuclei or mouse cells or nuclei.
57. A method of labeling nucleic acids with cell- or nucleus-specific tags, the method comprising:(a) providing a plurality of fixed and permeabilized cells or nuclei, each comprising genomic DNA;(b) fragmenting the genomic DNA within the plurality of cells or nuclei of the first plurality of aliquots to produce a plurality of genomic DNA fragments, and appending a nucleic acid adapter to the ends of the plurality of genomic DNA fragments within the cells or nuclei in a non-target-specific manner, thereby producing a plurality of adapter-coupled genomic DNA fragments;(c) tagging the adapter coupled genomic DNA molecules within the plurality of cells or nuclei with one or more nucleic acid tags, therebygenerating tagged genomic DNA fragments, by performing steps (i) through (iii) one or more times:(i) dividing the pooled cells or nuclei into a plurality of aliquots;(ii) coupling nucleic acid tags of any one of claims 1-21 to the adapter-coupled genomic DNA fragments within the cells or nuclei of the plurality of aliquots; and(iii) pooling the cells or nuclei from the plurality of aliquots;(d) lysing the cells or nuclei to release the tagged genomic DNA fragments and produce a lysate comprising the released tagged genomic DNA fragments; and(e) isolating the released tagged genomic DNA fragments.
58. The method of claim 57, wherein step (c)(ii) comprises coupling the adapter- coupled genomic DNA fragments with the set of nucleic acid tags of any one of claims 22- 30.
59. The method of claim 57 or 58, wherein steps (c)(i) through (c)(iii) are performed two or more times, thereby generating repeatedly tagged adapter-coupled genomic DNA molecules, and wherein the two or more times are performed using the collection of nucleic acid tags of any one of claims 33-35.
60. A method of cell- or nucleus-specifically labeling molecules in cells or nuclei, the method comprising:(a) providing a plurality of fixed and permeabilized cells or nuclei;(b) dividing the plurality of cells or nuclei into a first plurality of aliquots, wherein each of the aliquots of the first plurality of aliquots comprises more than one cell or nucleus;(c) coupling a first set of nucleic acid tags of any one of claims 22-30 to molecules within the cells or nuclei of the first plurality of aliquots, thereby generating a plurality of tagged molecules,(d) pooling the cells or nuclei from the first plurality of aliquots;(e) dividing the pooled cells from the first plurality of aliquots into a second plurality of aliquots;(f) coupling a second set of nucleic acid tags of any one of claims 22-30 to plurality the tagged molecules, thereby generating a plurality of repeatedly tagged molecules, and(g) pooling the cells or nuclei from the second plurality of cells or nuclei.
61. The method of claim 60, further comprising:(h) lysing the cells in the second plurality of aliquots to release the repeatedly tagged molecules; and(i) isolating the released repeatedly tagged molecules.
62. The method of claim 61, wherein the molecules are nucleic acids, and wherein the method further comprises:(j) amplifying the isolated repeatedly tagged molecules.
63. The method of claim 62, further comprising:(k) sequencing the amplified repeatedly tagged molecules.
64. The method of claim 63, further comprising:(l) grouping the sequencing reads obtained in (k) by first and / or second barcode sequence.
65. The method of claim 63 or 64, wherein the first barcode sequences are used to identify the individual aliquots among the first plurality of aliquots in which each of the repeatedly tagged molecules was present during step (c); and / or wherein the second barcode sequences are used to identify the individual aliquots among the second plurality of aliquots in which each of the repeatedly tagged molecules was present during step (f).
66. The method of claim 65, wherein the first and / or second barcode sequences are used to determine the identities of the individual cells or nuclei from which each of the molecules originated.
67. The method of any one of claims 60-66, wherein the molecules are mRNA molecules, and wherein the coupling in step (c) comprises reverse transcribing the nucleic acid molecules using the first set of nucleic acid tags as primers.
68. The method of any one of claims 60-67, wherein the molecules are complementary DNA (cDNA) molecules that have been generated by reverse transcription of RNA within the fixed and permeabilized cells or nuclei prior to step (a).
69. The method of any one of claims 60-68, wherein the molecules comprise genomic DNA, and wherein the coupling in step (c) comprises fragmenting the genomic DNA to generate a plurality of genomic DNA fragments, and appending the first set of nucleic acid tags to the ends of the plurality of genomic DNA fragments.
70. The method of any one of claims 60-69, wherein the molecules comprise adapter-coupled genomic DNA fragments that have been generated by fragmenting genomic DNA and appending adapters to the ends of the genomic DNA fragments within the fixed and permeabilized cells or nuclei prior to step (a).
71. The method of any one of claims 60-70, wherein the molecules comprise polypeptides.
72. The method of any one of claims 60-71, wherein the molecules comprise nucleic acid molecules, and wherein the coupling in step (c) comprises ligating the first set of nucleic acid tags to the nucleic acid molecules.
73. The method of any one of claims 60-72, wherein the coupling in step (f) comprises ligating the second set of nucleic acid tags to the tagged nucleic acid molecules.
74. The method of claim 72 or 73, wherein the ligating is performed for 15 or fewer minutes.
75. The method of any one of claims 60-74, wherein the nucleic acid tags of the first and / or second set of nucleic acid tags are DNA tags.
76. A kit for performing the method of any one or more of claims 36-75, the kit comprising the nucleic acid tag of any one of claims 1-21, a stop oligo comprising a sequence as shown in Table 1 and / or comprising the sequence of SEQ ID NO: 9 or SEQ ID NOTO, the set of nucleic acid tags of any one of claims 22-30, and / or the collection of nucleic acid tags of any one of claims 31-35.
Citation Information
Patent Citations
Reagents And Methods For The Analysis of Linked Nucleic Acids
US20220220531A1
Methods and Reagents for Molecular Barcoding
US20230017673A1
Cell barcoding compositions and methods
US20230048356A1