Molecular barcoding at opposite transcript ends
Molecular barcoding at both 5' and 3' ends of nucleic acid targets addresses amplification bias in PCR, enabling accurate counting of nucleic acid molecules by forming stem-loops and using unique molecular labels for precise quantification.
Patent Information
- Application Number
- JP2023165263
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-05-03
- Filing Date
- 2023-09-27
- Publication Date
- 2025-09-03
- Estimated Expiration
- 2039-05-01
AI Technical Summary
Existing methods for determining the absolute number of nucleic acid molecules, such as mRNA, are challenging due to amplification bias and stochastic replication in PCR, leading to inaccurate gene expression measurements.
A method involving molecular barcoding at both 5' and 3' ends of nucleic acid targets using oligonucleotide barcodes, forming stem-loops, and extending these barcodes to generate barcoded nucleic acid molecules, followed by amplification and sequencing to determine the number of targets based on unique molecular label sequences.
This approach corrects for amplification bias and accurately counts nucleic acid targets, providing precise quantification of nucleic acid molecules, especially in low abundance samples.
Smart Images

Figure 0007733704000001 
Figure 0007733704000002 
Figure 0007733704000003
Abstract
Description
[Technical Field]
[0001] Related Applications This application claims the benefit of U.S. Provisional Patent Application No. 62 / 666,506, filed May 3, 2018, which is incorporated herein by reference in its entirety.
[0002] The present disclosure relates generally to the field of molecular biology, and more particularly to multi-omics analysis using molecular barcoding. [Background technology]
[0003] Methods and techniques such as molecular barcoding are useful for single-cell transcriptomics analysis, which involves decoding gene expression profiles to determine the state of a cell, for example, using reverse transcription, polymerase chain reaction (PCR) amplification, and next-generation sequencing (NGS). Molecular barcoding is also useful for single-cell proteomics analysis. Summary of the Invention [Means for solving the problem]
[0004] Disclosed herein is a method for binding oligonucleotide barcodes to targets in a sample. For example, the target can comprise, consist essentially of, or consist of a nucleic acid target. In some embodiments, the method includes the steps of: barcoding copies of the nucleic acid target in the sample using a plurality of oligonucleotide barcodes to generate a plurality of barcoded nucleic acid molecules, each of which comprises the sequence of the nucleic acid target, a molecular label, and a target binding region, wherein at least 10 of the plurality of oligonucleotide barcodes comprise different molecular label sequences; binding oligonucleotides comprising complements of the target binding regions to the plurality of barcoded nucleic acid molecules to generate a plurality of barcoded nucleic acid molecules, each of which comprises a target binding region and a complement of the target binding region; hybridizing the target binding region and the complement of the target binding region in each of the plurality of barcoded nucleic acid molecules to form a stem-loop; and extending the 3' ends of the plurality of barcoded nucleic acid molecules to extend the stem-loop and generate a plurality of extended barcoded nucleic acid molecules, each of which comprises a molecular label and a complement of the molecular label. Optionally, the method further comprises amplifying the plurality of barcoded nucleic acid molecules to generate a plurality of amplified barcoded nucleic acid molecules. Binding an oligonucleotide comprising a complement of the target binding region may comprise binding the oligonucleotide to the plurality of amplified barcoded nucleic acid molecules to generate a plurality of barcoded nucleic acid molecules, each comprising a target binding region and a complement of the target binding region. Optionally, the method further comprises amplifying the plurality of extended barcoded nucleic acid molecules. The number of nucleic acid targets in the sample can be determined after amplifying the plurality of extended barcoded nucleic acid molecules. In some embodiments, the method comprises determining the number of nucleic acid targets in the sample based on the number of molecular labels having unique sequences, their complements, or combinations thereof, associated with the plurality of extended barcoded nucleic acid molecules. In some embodiments, the method comprises determining the number of nucleic acid targets in the sample based on the number of molecular labels having unique sequences associated with the plurality of extended barcoded nucleic acid molecules.
[0005] Disclosed herein are methods for determining the number of targets in a sample. By way of example, the targets can comprise, consist essentially of, or consist of nucleic acids. In some embodiments, the method includes barcoding copies of a nucleic acid target in a sample with a plurality of oligonucleotide barcodes to generate a plurality of barcoded nucleic acid molecules, each comprising the sequence of the nucleic acid target, a molecular label, and a target binding region, wherein at least 10 of the plurality of oligonucleotide barcodes comprise different molecular label sequences; binding oligonucleotides comprising complements of the target binding regions to the plurality of barcoded nucleic acid molecules to generate a plurality of barcoded nucleic acid molecules, each comprising a target binding region and a complement of the target binding region; hybridizing the target binding region and the complement of the target binding region in each of the plurality of barcoded nucleic acid molecules to form a stem-loop; extending the 3' ends of the plurality of barcoded nucleic acid molecules to extend the stem-loop to generate a plurality of extended barcoded nucleic acid molecules, each comprising a molecular label and a complement of the molecular label; and determining the number of nucleic acid targets in the sample based on the number of complements of molecular labels with unique sequences associated with the plurality of extended barcoded nucleic acid molecules. Optionally, the method further comprises amplifying the plurality of barcoded nucleic acid molecules to generate a plurality of amplified barcoded nucleic acid molecules. Binding an oligonucleotide comprising a complement of the target binding region may comprise binding the oligonucleotide to the plurality of amplified barcoded nucleic acid molecules to generate a plurality of barcoded nucleic acid molecules, each comprising a target binding region and a complement of the target binding region. Optionally, the method further comprises amplifying the plurality of extended barcoded nucleic acid molecules. The number of nucleic acid targets in the sample can be determined after amplifying the plurality of extended barcoded nucleic acid molecules. In some embodiments, the method comprises determining the number of nucleic acid targets in the sample based on the number of molecular labels having unique sequences, their complements, or a combination thereof, associated with the plurality of extended barcoded nucleic acid molecules.
[0006] In some embodiments, any of the methods described herein includes barcoding a plurality of copies of a target, the step comprising contacting the copies of the nucleic acid target with a plurality of oligonucleotide barcodes, each of the plurality of oligonucleotide barcodes comprising a target binding region capable of hybridizing to the nucleic acid target; and extending the copies of the nucleic acid target hybridized to the oligonucleotide barcodes to generate a plurality of barcoded nucleic acid molecules. In some embodiments, any of the methods described herein includes barcoding a plurality of copies of a target, the step comprising contacting the copies of the nucleic acid target with a plurality of oligonucleotide barcodes, each of the plurality of oligonucleotide barcodes comprising a target binding region. The target binding region can hybridize to the nucleic acid target. The method may further include extending the copies of the nucleic acid target hybridized to the oligonucleotide barcodes to generate a plurality of barcoded nucleic acid molecules.
[0007] In some embodiments, any of the methods described herein comprises amplifying a plurality of barcoded nucleic acid molecules to generate a plurality of amplified barcoded nucleic acid molecules, wherein binding an oligonucleotide comprising a complement of the target binding region comprises binding an oligonucleotide comprising a complement of the target binding region to the plurality of amplified barcoded nucleic acid molecules to generate a plurality of barcoded nucleic acid molecules each comprising a target binding region and a complement of the target binding region.
[0008] In some embodiments, any of the methods described herein comprises amplifying a plurality of elongated barcoded nucleic acid molecules to generate a plurality of single-labeled nucleic acid molecules, each comprising a complement of a molecular label, wherein determining the number of nucleic acid targets in the sample comprises determining the number of nucleic acid targets in the sample based on the number of complements of molecular labels having unique sequences that are associated with the plurality of single-labeled nucleic acid molecules.
[0009] In some embodiments, the method comprises amplifying a plurality of extended barcoded nucleic acid molecules to generate a plurality of copies of the extended barcoded nucleic acid molecules, wherein determining the number of nucleic acid targets in the sample comprises determining the number of nucleic acid targets in the sample based on the number of complements of molecular labels having unique sequences that are associated with the copies of the plurality of extended barcoded nucleic acid molecules.
[0010] Disclosed herein is a method for determining the number of nucleic acid targets in a sample. In some embodiments, the method includes the steps of contacting copies of the nucleic acid targets in the sample with a plurality of oligonucleotide barcodes, each of the plurality of oligonucleotide barcodes comprising a molecular label and a target binding region capable of hybridizing to the nucleic acid target, wherein at least 10 of the plurality of oligonucleotide barcodes comprise different molecular label sequences; extending the copies of the nucleic acid targets hybridized to the oligonucleotide barcodes to produce a plurality of nucleic acid molecules, each of which comprises a sequence complementary to at least a portion of the nucleic acid target; amplifying the plurality of barcoded nucleic acid molecules to produce a plurality of amplified barcoded nucleic acid molecules; and contacting oligonucleotides comprising complements of the target binding region with the plurality of amplified barcoded nucleic acids. hybridizing the target binding region and the complement of the target binding region in each of the plurality of barcoded nucleic acid molecules to form a stem-loop; extending the 3' ends of the plurality of barcoded nucleic acid molecules to extend the stem-loop and generate a plurality of extended barcoded nucleic acid molecules each comprising a molecular label and a complement of the molecular label; amplifying the plurality of extended barcoded nucleic acid molecules to generate a plurality of single-labeled nucleic acid molecules each comprising a complement of the molecular label; and determining the number of nucleic acid targets in the sample based on the number of complements of molecular labels having unique sequences associated with the plurality of single-labeled nucleic acid molecules.
[0011] In some embodiments, for any of the methods described herein, the molecular label is hybridized to a complement of the molecular label after extending the 3' ends of the plurality of barcoded nucleic acid molecules, and the method includes denaturing the plurality of extended barcoded nucleic acid molecules before amplifying the plurality of extended barcoded nucleic acid molecules to generate a plurality of single-labeled nucleic acid molecules. Contacting copies of the nucleic acid targets in the sample may include contacting the copies of the plurality of nucleic acid targets with a plurality of oligonucleotide barcodes. Extending copies of the nucleic acid targets may include extending copies of the plurality of nucleic acid targets hybridized to the oligonucleotide barcodes to generate a plurality of barcoded nucleic acid molecules, each of which comprises a sequence complementary to at least a portion of one of the plurality of nucleic acid targets. Determining the number of nucleic acid targets may include determining the number of each of the plurality of nucleic acid targets in the sample based on the number of complements of molecular labels having unique sequences associated with the single-labeled nucleic acid molecules of the plurality of single-labeled nucleic acid molecules comprising the respective sequences of the plurality of nucleic acid targets. The sequences of each of the plurality of nucleic acid targets may comprise a respective subsequence of the plurality of nucleic acid targets.
[0012] In some embodiments, for any of the methods described herein, the sequence of the nucleic acid target in the plurality of barcoded nucleic acid molecules comprises a subsequence of the nucleic acid target. The target binding region may comprise a gene-specific sequence. Binding an oligonucleotide comprising a complement of the target binding region may comprise ligating an oligonucleotide comprising a complement of the target binding region to the plurality of barcoded nucleic acid molecules.
[0013] In some embodiments, for any of the methods described herein, the target binding region can comprise a poly(dT) sequence, and wherein binding an oligonucleotide comprising a complement of the target binding region comprises adding a plurality of adenosine monophosphates to the plurality of barcoded nucleic acid molecules using terminal deoxynucleotidyl transferase.
[0014] In some embodiments, for any of the methods described herein, extending a copy of the nucleic acid target hybridized to the oligonucleotide barcode may include reverse transcribing the copy of the nucleic acid target hybridized to the oligonucleotide barcode to generate a plurality of barcoded complementary deoxyribonucleic acid (cDNA) molecules. Extending a copy of the nucleic acid target hybridized to the oligonucleotide barcode may include extending a copy of the nucleic acid target hybridized to the oligonucleotide barcode using a DNA polymerase lacking at least one of 5'-3' exonuclease activity and 3'-5' exonuclease activity. The DNA polymerase may comprise Klenow fragment.
[0015] In some embodiments, any of the methods described herein includes obtaining sequence information for the plurality of extended barcoded nucleic acid molecules. Obtaining the sequence information may include attaching sequencing adapters to the plurality of extended barcoded nucleic acid molecules.
[0016] In some embodiments, for any of the methods described herein, the complement of the target binding region may comprise the reverse complementary sequence of the target binding region. The complement of the target binding region may comprise the complementary sequence of the target binding region. The complement of the molecular label may comprise the reverse complementary sequence of the molecular label. The complement of the molecular label may comprise the complementary sequence of the molecular label.
[0017] In some embodiments, for any of the methods described herein, the plurality of barcoded nucleic acid molecules may comprise barcoded deoxyribonucleic acid (DNA) molecules. The barcoded nucleic acid molecules may comprise barcoded ribonucleic acid (RNA) molecules. The nucleic acid target may comprise a nucleic acid molecule. The nucleic acid molecule may comprise ribonucleic acid (RNA), messenger RNA (mRNA), microRNA, small interfering RNA (siRNA), RNA degradation products, RNA containing a poly(A) tail, or any combination thereof.
[0018] In some embodiments, for any of the methods described herein, the nucleic acid target can include a cellular component binding reagent. The nucleic acid molecule can be associated with the cellular component binding reagent. The method can include dissociating the nucleic acid molecule and the cellular component binding reagent.
[0019] In some embodiments, for any of the methods described herein, each molecular label of the plurality of oligonucleotide barcodes comprises at least 6 nucleotides. The oligonucleotide barcodes may comprise the same sample label. Each sample label of the plurality of oligonucleotide barcodes may comprise at least 6 nucleotides. The oligonucleotide barcodes may comprise the same cell label. Each cell label of the plurality of oligonucleotide barcodes may comprise at least 6 nucleotides.
[0020] In some embodiments, for any of the methods described herein, at least one of the plurality of barcoded nucleic acid molecules becomes associated with the solid support when the target binding region and the complement of the target binding region in each of the plurality of barcoded nucleic acid molecules hybridize to form a stem loop. At least one of the plurality of barcoded nucleic acid molecules can dissociate from the solid support when the target binding region and the complement of the target binding region in each of the plurality of barcoded nucleic acid molecules hybridize to form a stem loop. At least one of the plurality of barcoded nucleic acid molecules can become associated with the solid support when the target binding region and the complement of the target binding region in each of the plurality of barcoded nucleic acid molecules hybridize to form a stem loop.
[0021] In some embodiments, for any of the methods described herein, at least one of the plurality of barcoded nucleic acid molecules is associated with a solid support when the 3' ends of the plurality of barcoded nucleic acid molecules are extended to extend the stem-loop to generate a plurality of extended barcoded nucleic acid molecules each comprising a molecular label and a complement of the molecular label. At least one of the plurality of barcoded nucleic acid molecules may be dissociated from the solid support when the 3' ends of the plurality of barcoded nucleic acid molecules are extended to extend the stem-loop to generate a plurality of extended barcoded nucleic acid molecules each comprising a molecular label and a complement of the molecular label. At least one of the plurality of barcoded nucleic acid molecules may be associated with a solid support when the 3' ends of the plurality of barcoded nucleic acid molecules are extended to extend the stem-loop to generate a plurality of extended barcoded nucleic acid molecules each comprising a molecular label and a complement of the molecular label. The solid support may comprise a synthetic particle. The solid support may comprise a flat surface (e.g., a slide such as a microscope slide and a coverslip). In some embodiments, the solutions described herein can be dispensed into compartments containing one or fewer cells. The compartments can include at least one of a microdroplet, a well (e.g., a microwell) on a substrate, or a chamber of a fluidic device such as a microfluidic device. The microdroplet can include a hydrogel.
[0022] In some embodiments, for any of the methods described herein, at least one of the plurality of barcoded nucleic acid molecules is in solution when the target binding region and the complement of the target binding region in each of the plurality of barcoded nucleic acid molecules hybridize to form a stem-loop. At least one of the plurality of barcoded nucleic acid molecules can be in solution when the 3' ends of the plurality of barcoded nucleic acid molecules are extended to extend the stem-loop and generate a plurality of extended barcoded nucleic acid molecules each comprising a molecular label and a complement of the molecular label.
[0023] In some embodiments, for any of the methods described herein, the sample includes a single cell, and the method includes associating a synthetic particle comprising a plurality of oligonucleotide barcodes with the single cell in the sample. The method may include lysing the single cell after associating the synthetic particle with the single cell. Lysing the single cell may include heating the sample, contacting the sample with a detergent, changing the pH of the sample, or any combination thereof. The synthetic particle and the single cell may be in the same well. The synthetic particle and the single cell may be in the same droplet.
[0024] In some embodiments, for any of the methods described herein, at least one of the plurality of oligonucleotide barcodes may be immobilized on a synthetic particle. At least one of the plurality of oligonucleotide barcodes may be partially immobilized on a synthetic particle. At least one of the plurality of oligonucleotide barcodes may be encapsulated in a synthetic particle. At least one of the plurality of oligonucleotide barcodes may be partially encapsulated in a synthetic particle. The synthetic particle may be disintegrable. The synthetic particle may comprise beads. The beads may comprise sepharose beads, streptavidin beads, agarose beads, magnetic beads, conjugated beads, protein A conjugated beads, protein G conjugated beads, protein A / G conjugated beads, protein L conjugated beads, oligo(dT) conjugated beads, silica beads, silica-like beads, anti-biotin microbeads, anti-fluorescent dye microbeads, or any combination thereof. The synthetic particles may comprise a material selected from the group consisting of polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogel, paramagnetic material, ceramic, plastic, glass, methylstyrene, acrylic polymer, titanium, latex, Sepharose, cellulose, nylon, silicone, and any combination thereof. The synthetic particles may comprise collapsible hydrogel particles. Each of the plurality of oligonucleotide barcodes may comprise a linker functional group. The synthetic particles may comprise a solid support functional group. The support functional group and the linker functional group may be associated with each other. The linker functional group and the support functional group may be independently selected from the group consisting of C6, biotin, streptavidin, primary amine, aldehyde, ketone, and any combination thereof.
[0025] Disclosed herein is a kit for binding oligonucleotide barcodes to targets in a sample, determining the number of targets in a sample, and / or determining the number of nucleic acid targets in a sample. In some embodiments, the kit includes a plurality of oligonucleotide barcodes, each of the plurality of oligonucleotide barcodes comprising a molecular label and a target binding region, wherein at least 10 of the plurality of oligonucleotide barcodes comprise different molecular label sequences; a terminal deoxynucleotidyl transferase or ligase; and a DNA polymerase lacking at least one of 5'-3' exonuclease activity and 3'-5' exonuclease activity. The kit may further include a plurality of oligonucleotides comprising complements of the target binding region. The plurality of oligonucleotides comprising complements of the target binding region are configured for binding to the 3' end of a DNA molecule, such as a cDNA molecule. The DNA polymerase may comprise a Klenow fragment. The kit may include a buffer. The kit may include a cartridge. The kit may include one or more reagents for a reverse transcription reaction. The kit may include one or more reagents for an amplification reaction.
[0026] In some embodiments, for any of the kits or methods described herein, the target binding region comprises a gene-specific sequence, an oligo(dT) sequence, a random multimer, or any combination thereof. The oligonucleotide barcodes may comprise the same sample label and / or the same cell label. Each sample label and / or cell label of the plurality of oligonucleotide barcodes may comprise at least 6 nucleotides. Each molecular label of the plurality of oligonucleotide barcodes may comprise at least 6 nucleotides.
[0027] In some embodiments, for any kit or method described herein, at least one of the plurality of oligonucleotide barcodes is immobilized on a synthetic particle. At least one of the plurality of oligonucleotide barcodes may be partially immobilized on a synthetic particle. At least one of the plurality of oligonucleotide barcodes may be encapsulated in a synthetic particle. At least one of the plurality of oligonucleotide barcodes may be partially encapsulated in a synthetic particle. The synthetic particle may be disintegrable. The synthetic particle may comprise beads. The beads may comprise sepharose beads, streptavidin beads, agarose beads, magnetic beads, conjugated beads, protein A conjugated beads, protein G conjugated beads, protein A / G conjugated beads, protein L conjugated beads, oligo(dT) conjugated beads, silica beads, silica-like beads, anti-biotin microbeads, anti-fluorescent dye microbeads, or any combination thereof. The synthetic particles may comprise a material selected from the group consisting of polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogel, paramagnetic material, ceramic, plastic, glass, methylstyrene, acrylic polymer, titanium, latex, Sepharose, cellulose, nylon, silicone, and any combination thereof. The synthetic particles may comprise collapsible hydrogel particles. Each of the plurality of oligonucleotide barcodes may comprise a linker functional group. The synthetic particles may comprise a solid support functional group. The support functional group and the linker functional group may be associated with each other. The linker functional group and the support functional group may be independently selected from the group consisting of C6, biotin, streptavidin, primary amine, aldehyde, ketone, and any combination thereof. [Brief explanation of the drawings]
[0028] [Figure 1] 1 illustrates a non-limiting exemplary barcode of some embodiments. [Figure 2] 1 illustrates a non-limiting exemplary workflow for barcoding and digital counting of some embodiments. [Figure 3] FIG. 1 is a schematic diagram showing a non-limiting, exemplary process for generating an indexed library of 3′-barcoded targets from multiple targets, according to some embodiments. [Figure 4A] FIG. 1 shows a schematic diagram of a non-limiting exemplary method for gene-specific labeling of nucleic acid targets at the 5′ end according to some embodiments. [Figure 4B] FIG. 1 shows a schematic diagram of a non-limiting exemplary method for gene-specific labeling of nucleic acid targets at the 5′ end according to some embodiments. [Figure 5A] FIG. 1 shows a schematic diagram of a non-limiting exemplary method for labeling nucleic acid targets at the 5′ end for whole transcriptome analysis of some embodiments. [Figure 5B] FIG. 1 shows a schematic diagram of a non-limiting exemplary method for labeling nucleic acid targets at the 5′ end for whole transcriptome analysis of some embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0029] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof. In these drawings, like numerals typically identify like elements, unless the context dictates otherwise. The exemplary embodiments described in the detailed description, drawings, and claims are not intended to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that aspects of the present disclosure, as generally described herein and illustrated in the figures, can be arranged, substituted, combined, separated, and designed in a variety of different configurations, all of which are expressly contemplated herein and form a part of this disclosure.
[0030] All patents, published patent applications, other publications, and sequences from GenBank and other databases mentioned herein are incorporated by reference in their entirety for relevant art.
[0031] Quantifying a small number of nucleic acids, such as messenger ribonucleic acid (mRNA) molecules, is clinically important, for example, to determine the genes expressed in cells at different developmental stages or under different environmental conditions. However, determining the absolute number of nucleic acid molecules (e.g., mRNA molecules) can be very difficult, especially when the number of molecules is very small. One method for determining the absolute number of molecules in a sample is digital polymerase chain reaction (PCR). Ideally, PCR produces identical copies of molecules in each cycle. However, PCR can have drawbacks, such as each molecule replicates with a stochastic probability, and this probability varies depending on the PCR cycle and gene sequence, resulting in amplification bias and inaccurate gene expression measurements. Probabilistic barcodes with unique molecular labels (also called molecular beacons (MIs)) can be used to count molecules and correct for amplification bias. Probabilistic barcoding, such as the Precise™ assay (Cellular Research, Inc., Palo Alto, CA) and the Rhapsody™ assay (Becton, Dickinson and Company, Franklin Lakes, NJ), can correct for biases induced by the PCR and library generation steps by using molecular beacons (MLs) to label mRNA during reverse transcription (RT). Methods and techniques for molecular barcoding of nucleic acid targets at either or both the 5' and 3' ends are needed.
[0032] The Precise™ assay uses a non-depleting pool of stochastic barcodes with a large number (e.g., 6561-65536) of unique molecular tag sequences on poly(T) oligonucleotides to hybridize all poly(A)-mRNAs in a sample during the RT step. The stochastic barcodes may contain universal PCR priming sites. During RT, target gene molecules react randomly with the stochastic barcodes. Each target molecule can hybridize with the stochastic barcode to generate a complementary ribonucleotide acid (cDNA) molecule with a stochastic barcode. 。After labeling, the probabilistically barcoded cDNA molecules from the microwells of the microwell plate can be pooled into a single tube for PCR amplification and sequencing. The raw sequencing data can be analyzed to obtain the number of reads, the number of probabilistic barcodes with unique molecular label sequences, and the number of mRNA molecules.
[0033] Disclosed herein is a method for binding oligonucleotide barcodes to targets in a sample. In some embodiments, the method includes the steps of: barcoding copies of a nucleic acid target in the sample using a plurality of oligonucleotide barcodes to generate a plurality of barcoded nucleic acid molecules, each of which comprises the sequence of the nucleic acid target, a molecular label, and a target binding region, wherein at least 10 of the plurality of oligonucleotide barcodes comprise different molecular label sequences; binding an oligonucleotide comprising a complement of the target binding region to the plurality of barcoded nucleic acid molecules, to generate a plurality of barcoded nucleic acid molecules, each of which comprises the target binding region and the complement of the target binding region; hybridizing the target binding region and the complement of the target binding region in each of the plurality of barcoded nucleic acid molecules to form a stem-loop; and extending the 3' ends of the plurality of barcoded nucleic acid molecules to extend the stem-loop and generate a plurality of extended barcoded nucleic acid molecules, each of which comprises a molecular label and the complement of the molecular label.
[0034] The present disclosure includes a method for determining the number of targets in a sample. In some embodiments, the method includes the steps of: barcoding copies of nucleic acid targets in the sample using a plurality of oligonucleotide barcodes to generate a plurality of barcoded nucleic acid molecules, each of which contains the sequence of the nucleic acid target, a molecular label, and a target binding region, wherein at least 10 of the plurality of oligonucleotide barcodes contain different molecular label sequences; binding oligonucleotides containing complements of the target binding region to the plurality of barcoded nucleic acid molecules to generate a plurality of barcoded nucleic acid molecules, each of which contains a target binding region and a complement of the target binding region; hybridizing the target binding region and the complement of the target binding region in each of the plurality of barcoded nucleic acid molecules to form a stem-loop; extending the 3' ends of the plurality of barcoded nucleic acid molecules to extend the stem-loop to generate a plurality of extended barcoded nucleic acid molecules, each of which contains a molecular label and a complement of the molecular label; and determining the number of nucleic acid targets in the sample based on the number of complements of molecular labels with unique sequences associated with the plurality of extended barcoded nucleic acid molecules.
[0035] Disclosed herein is a method for determining the number of nucleic acid targets in a sample. In some embodiments, the method includes the steps of contacting copies of the nucleic acid targets in the sample with a plurality of oligonucleotide barcodes, each of the plurality of oligonucleotide barcodes comprising a molecular label and a target binding region capable of hybridizing to the nucleic acid target, wherein at least 10 of the plurality of oligonucleotide barcodes comprise different molecular label sequences; extending the copies of the nucleic acid targets hybridized to the oligonucleotide barcodes to produce a plurality of nucleic acid molecules, each of which comprises a sequence complementary to at least a portion of the nucleic acid target; amplifying the plurality of barcoded nucleic acid molecules to produce a plurality of amplified barcoded nucleic acid molecules; and contacting oligonucleotides comprising complements of the target binding region with the plurality of amplified barcoded nucleic acids. hybridizing the target binding region and the complement of the target binding region in each of the plurality of barcoded nucleic acid molecules to form a stem-loop; extending the 3' ends of the plurality of barcoded nucleic acid molecules to extend the stem-loop and generate a plurality of extended barcoded nucleic acid molecules each comprising a molecular label and a complement of the molecular label; amplifying the plurality of extended barcoded nucleic acid molecules to generate a plurality of single-labeled nucleic acid molecules each comprising a complement of the molecular label; and determining the number of nucleic acid targets in the sample based on the number of complements of molecular labels having unique sequences associated with the plurality of single-labeled nucleic acid molecules.
[0036] Disclosed herein is a kit for binding oligonucleotide barcodes to targets in a sample, determining the number of targets in a sample, and / or determining the number of nucleic acid targets in a sample. In some embodiments, the kit includes a plurality of oligonucleotide barcodes, each of the plurality of oligonucleotide barcodes comprising a molecular label and a target binding region, and at least 10 of the plurality of oligonucleotide barcodes comprising different molecular label sequences; a terminal deoxynucleotidyl transferase or ligase; and a DNA polymerase lacking at least one of 5'-3' exonuclease activity and 3'-5' exonuclease activity. The DNA polymerase may comprise a Klenow fragment. The kit may include a buffer. The kit may include a cartridge. The kit may include one or more reagents for a reverse transcription reaction. The kit may include one or more reagents for an amplification reaction.
[0037] Unless otherwise defined, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Examples of references available to those skilled in the art include Singleton et al., Dictionary of Microbiology and Molecular Biology 2nd ed., J. Wiley & Sons (New York, NY 1994); Sambrook et al., Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Press (Cold Spring Harbor, NY 1989).
[0038] As used herein, the term "adapter" has its conventional and ordinary meaning in the art, taking this specification into consideration. It may refer to a sequence for facilitating amplification or sequencing of a related nucleic acid. The related nucleic acid may include a target nucleic acid. The related nucleic acid may include one or more of a spatial label, a target label, a sample label, an index label, or a barcode sequence (e.g., a molecular label). The adapter may be linear. The adapter may be a pre-adenylated adapter. The adapter may be double-stranded or single-stranded. One or more adapters may be located at the 5' or 3' end of a nucleic acid. When an adapter includes a known sequence at the 5' and 3' ends, the known sequence may be the same or different sequences. An adapter located at the 5' and / or 3' end of a polynucleotide may be capable of hybridizing to one or more oligonucleotides immobilized on a surface. In some embodiments, the adapter may include a universal sequence. A universal sequence may be a region of nucleotide sequence common to two or more nucleic acid molecules. The two or more nucleic acid molecules may also have regions of different sequences. Thus, for example, the 5' adapters can contain identical and / or universal nucleic acid sequences, and the 3' adapters can contain identical and / or universal sequences. A universal sequence that can be present in different members of a plurality of nucleic acid molecules can enable replication or amplification of multiple different sequences using a single universal primer complementary to the universal sequence. Similarly, at least one, two (e.g., pairs), or more universal sequences that can be present in different members of a collection of nucleic acid molecules can enable replication or amplification of multiple different sequences using at least one, two (e.g., pairs), or more single universal primers complementary to the universal sequence. Thus, a universal primer comprises a sequence that can hybridize to such a universal sequence. A target nucleic acid sequence-bearing molecule can be modified to attach universal adapters (e.g., non-target nucleic acid sequences) to one or both ends of different target nucleic acid sequences. One or more universal primers attached to the target nucleic acids can provide sites for hybridization of the universal primers.The one or more universal primers bound to the target nucleic acid can be the same or different from each other.
[0039] As used herein, the terms "associated" or "associated with" have their conventional and ordinary meaning in the art, given the present specification. It means that two or more species are identifiable as co-located at some point in time. Association means that two or more species are or were in similar containers. Association can be an informatic association. For example, digital information associated with two or more species can be stored and used to determine that one or more of these species were co-located at some point in time. Association can also be a physical association. In some embodiments, two or more associated species are "tethered," "bound," or "immobilized" to each other or to a common solid or semi-solid surface. Association can refer to covalent or non-covalent means for attaching a label to a solid or semi-solid support, such as a bead. Association can be a covalent bond between a target and a label. Association can include hybridization between two molecules (such as a target molecule and a label).
[0040] As used herein, the term "complementary" has its conventional and usual meaning in the art, taking this specification into consideration. It can refer to the ability for precise pairing between two nucleotides. For example, if a nucleotide at a given position in a nucleic acid is capable of hydrogen bonding with a nucleotide in another nucleic acid, the two nucleic acids are considered to be complementary to each other at that position. Complementarity between two single-stranded nucleic acid molecules can be "partial" when only a portion of the nucleotides bind, or it can be complete when there is overall complementarity between the single-stranded molecules. If a first nucleotide sequence is complementary to a second nucleotide sequence, the first nucleotide sequence can be said to be the "complement" of the second sequence. If a first nucleotide sequence is complementary to the reverse (i.e., the order of the nucleotides is reversed) sequence of the second sequence, the first nucleotide sequence can be said to be the "reverse complement" of the second sequence. As used herein, a "complementary" sequence can refer to the "complement" or "reverse complement" of a sequence. It is understood from this disclosure that when a molecule is capable of hybridizing to another molecule, it may be complementary or partially complementary to the hybridizing molecule.
[0041] As used herein, the term "digital counting" can refer to a method for estimating the number of target molecules in a sample. Digital counting can include determining the number of unique labels associated with targets in a sample. This method can be probabilistic in nature, transforming the problem of counting molecules from one of locating and identifying identical molecules to a series of digital present / absent problems involving the detection of a given set of labels.
[0042] As used herein, the term "label" or "labels" has its conventional and usual meaning in the art in light of this specification. It can refer to a nucleic acid code associated with a target in a sample. A label can be, for example, a nucleic acid label. A label can be a wholly or partially amplifiable label. A label can be a wholly or partially sequenceable label. A label can be a portion of a native nucleic acid that can be uniquely identified. A label can be a known sequence. A label can include a nucleic acid sequence junction, e.g., a junction of a native and a non-native sequence. As used herein, the term "label" can be used synonymously with the terms "index," "tag," or "label tag." A label can convey information. For example, in various embodiments, a label can be used to determine the identity of a sample, the source of the sample, the identity of a cell, and / or a target.
[0043] As used herein, the term "non-depletion reservoir" can refer to a pool of barcodes (e.g., probabilistic barcodes) composed of many different labels. A non-depletion reservoir can contain many different barcodes so that when the non-depletion reservoir is associated with a pool of targets, each target is more likely to be associated with a unique barcode. The uniqueness of each labeled target molecule can be determined by the statistics of random selection and depends on the copy number of the same target molecule in the collection compared to the diversity of the labels. The size of the resulting set of labeled target molecules can be determined by the stochastic nature of the barcoding process, and analysis of the number of detected barcodes then allows for the calculation of the number of target molecules present in the original collection or sample. If the ratio of the copy number of the target molecule present to the number of unique barcodes is low, the labeled target molecule is highly unique (i.e., the probability that two or more target molecules are labeled with a given label is very low).
[0044] As used herein, the term "nucleic acid" has its conventional and usual meaning in the art in light of this specification. It refers to a polynucleotide sequence or a fragment thereof. A nucleic acid can comprise nucleotides. A nucleic acid can be exogenous or endogenous to a cell. A nucleic acid can be present in a cell-free environment. A nucleic acid can be a gene or a fragment thereof. A nucleic acid can be DNA. A nucleic acid can be RNA. A nucleic acid can contain one or more analogs (e.g., modified backbones, sugars, or nucleobases). Some non-limiting examples of analogs include 5-bromouracil, peptide nucleic acids, xenonucleic acids, morpholinos, locked nucleic acids, glycol nucleic acids, threose nucleic acids, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein attached to the sugar), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queosine, and wyosine. "Nucleic acids," "polynucleotides," 」 "Target polynucleotide" and "target nucleic acid" may be used interchangeably.
[0045] Nucleic acids can contain one or more modifications (e.g., base modifications, backbone modifications) to provide nucleic acids with new or improved characteristics (e.g., improved stability). Nucleic acids can contain a nucleic acid affinity tag. A nucleoside can be a base-sugar combination. The base portion of a nucleoside can be a heterocyclic base. The two most common classes of such heterocyclic bases are purines and pyrimidines. A nucleotide can be a nucleoside that further includes a phosphate group covalently linked to the sugar portion of the nucleoside. In nucleosides that include a pentofuranosyl sugar, the phosphate group can be attached to the 2', 3', or 5' hydroxyl moiety of the sugar. In forming nucleic acids, the phosphate group can covalently link adjacent nucleosides to one another to form a linear polymeric compound. Thus, the respective ends of this linear polymeric compound can be further linked to form a circular compound; however, linear compounds are generally preferred. Additionally, linear compounds may have internal nucleotide base complementarity and therefore may fold to produce fully or partially double-stranded compounds. Within nucleic acids, the phosphate groups may generally be referred to as forming the internucleoside backbone of the nucleic acid. The linkage or backbone may be a 3'-5' phosphodiester bond.
[0046] Nucleic acids can contain modified backbones and / or modified internucleoside linkages. Modified backbones can include those that retain a phosphorus atom in the backbone and those that do not have a phosphorus atom in the backbone. Suitable modified nucleic acid backbones containing a phosphorus atom therein can include, for example, phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotriesters, aminoalkylphosphotriesters, methyl and other alkyl phosphonates, such as 3'-alkylene phosphonates, 5'-alkylene phosphonates, chiral phosphonates, phosphinates, phosphoramidates including 3'-aminophosphoramidate and aminoalkylphosphoramidate, phosphorodiamidates, thienophosphoramidates, thienoalkylphosphonates, thienoalkylphosphotriesters, selenophosphates and boranophosphates having normal 3'-5' linkages, 2'-5' linkage analogs, and those having reverse polarity in which one or more internucleotide linkages are 3'-3', 5'-5', or 2'-2' linkages.
[0047] Nucleic acids can contain polynucleotide backbones formed by short chain alkyl or cycloalkyl internucleoside linkages, mixed heteroatom and alkyl or cycloalkyl internucleoside linkages, or one or more short chain heteroatom or heterocyclic internucleoside linkages. These can include morpholino linkages (formed in part from the sugar portion of the nucleoside), siloxane backbones, sulfide, sulfoxide, and sulfone backbones, formacetyl and thioformacetyl backbones, methyleneformacetyl and thioformacetyl backbones, riboacetyl backbones, alkene-containing backbones, sulfamate backbones, methyleneimino and methylenehydrazino backbones, sulfonate and sulfonamide backbones, those with amide backbones, and others with mixed N, O, S, and CH2 moieties.
[0048] Nucleic acids may include nucleic acid mimetics. The term "mimetic" is intended to include polynucleotides in which only the furanose ring or both the furanose ring and the internucleotide linkage are replaced with non-furanose groups; replacement of only the furanose ring may also be referred to as a sugar surrogate. A heterocyclic base moiety or a modified heterocyclic base moiety may be retained for hybridization with an appropriate target nucleic acid. One such nucleic acid may be a peptide nucleic acid (PNA). In a PNA, the sugar backbone of a polynucleotide may be replaced with an amide-containing backbone, particularly an aminoethylglycine backbone. Nucleotides may be retained and directly or indirectly linked to the aza nitrogen atoms of the amide portion of the backbone. The backbone in a PNA compound may contain two or more linked aminoethylglycine units, giving the PNA an amide-containing backbone. A heterocyclic base moiety may be directly or indirectly linked to the aza nitrogen atoms of the amide portion of the backbone.
[0049] The nucleic acid may include a morpholino backbone structure. For example, the nucleic acid may include a six-membered morpholino ring instead of a ribose ring. In some of these embodiments, phosphorodiamidate or other non-phosphodiester internucleoside linkages may replace the phosphodiester linkage.
[0050] Nucleic acids may contain linked morpholino units (e.g., morpholino nucleic acids) having heterocyclic bases attached to the morpholino ring. Linking groups may connect the morpholino monomer units in morpholino nucleic acids. Nonionic morpholino-based oligomeric compounds may have fewer undesirable interactions with cellular proteins. Morpholino-based polynucleotides may be nonionic mimics of nucleic acids. Various compounds within the morpholino class may be linked using different linking groups. A further class of polynucleotide mimetics may be called cyclohexenyl nucleic acids (CeNA). The furanose ring normally present in nucleic acid molecules may be replaced with a cyclohexenyl ring. CeNA DMT-protected phosphoramidite monomers may be prepared and used for oligomeric compound synthesis using phosphoramidite chemistry. Incorporation of CeNA monomers into nucleic acid chains can increase the stability of DNA / RNA hybrids. CeNA oligoadenylates can form complexes with nucleic acid complements with stability similar to that of the native complex. Further modifications may include locked nucleic acids (LNAs), in which a 2'-hydroxyl group is attached to the 4' carbon atom of the sugar ring, forming a 2'-C,4'-C-oxymethylene linkage to form a bicyclic sugar moiety. The linkage may be a methylene (-CH2) group (where n is 1 or 2) bridging the 2' oxygen atom and the 4' carbon atom. LNAs and LNA analogs may exhibit very high duplex thermal stability with complementary nucleic acids (Tm = +3 to +10°C), stability against 3'-exonuclease degradation, and good solubility.
[0051] Nucleic acids may also include nucleobase (often simply referred to as "base") modifications or substitutions. As used herein, "unmodified" or "natural" nucleobases may include purine bases (e.g., adenine (A) and guanine (G)) and pyrimidine bases (e.g., thymine (T), cytosine (C), and uracil (U)). Modified nucleobases include other synthetic and natural nucleobases, such as 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyl (-C=C-CH3) uracil and other alkynyl derivatives of cytosine and pyrimidine bases, 6-azouracil, cytosine ... These may include tosine and thymine, 5-uracil (pseudouracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl and other 8-substituted adenines and guanines, 5-halo, particularly 5-bromo, 5-trifluoromethyl and other 5-substituted uracils and cytosines, 7-methylguanine and 7-methyladenine, 2-F-adenine, 2-aminoadenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine and 3-deazaguanine and 3-deazaadenine.Modified nucleobases include tricyclic pyrimidines, such as phenoxazine cytidine (1H-pyrimido(5,4-b)(1,4)benzoxazin-2(3H)-one), phenothiazine cytidine (1H-pyrimido(5,4-b)(1,4)benzothiazin-2(3H)-one), G-clamps such as substituted phenoxazine cytidines (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzoxazin-2(3H)-one), phenothiazine cytidine (1H-pyrimido(5,4-b)(1,4)benzothiazin-2(3H)-one), G-clamps such as substituted phenoxazine cytidines (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzoxazin-2(3H)-one), carbazole cytidines (2H-pyrimido(4,5-b)indol-2-one), and pyridoindole cytidines (H-pyrido(3',2':4,5)pyrrolo[2,3-d]pyrimidin-2-one).
[0052] As used herein, the term "sample" has its conventional and ordinary meaning in the art in light of this specification. It can refer to a composition containing a target. Samples suitable for analysis by the disclosed methods, devices, and systems include cells, tissues, organs, or organisms. In some embodiments, a sample comprises a single cell. In some embodiments, a sample comprises, consists essentially of, or consists of at least 100,000, 200,000, 300,000, 500,000, 800,000, or 1,000,000 single cells.
[0053] As used herein, the terms "sampling device" or "device" may refer to a device capable of taking a section of a sample and / or depositing a section on a substrate. A sample device may refer to, for example, a fluorescence activated cell sorting (FACS) machine, a cell sorter machine, a biopsy needle, a biopsy device, a tissue sectioning device, a microfluidic device, a blade grid, and / or a microtome.
[0054] As used herein, the term "solid support" has its conventional and ordinary meaning in the art in light of this specification. It can refer to a discrete solid or semi-solid surface to which multiple barcodes (e.g., stochastic barcodes) can be attached. A solid support can include any type of solid, porous, or hollow sphere, ball, bearing, cylinder, or other similar structure composed of plastic, ceramic, metal, or polymeric material (e.g., hydrogel) to which nucleic acids can be immobilized (e.g., covalently or non-covalently). A solid support can include discrete particles that can be spherical (e.g., microspheres) or have non-spherical or irregular shapes, such as cubic, rectangular, pyramidal, cylindrical, conical, ellipsoidal, or disc-shaped. Beads can be non-spherical in shape. A plurality of solid supports spaced apart in an array can also be substrate-free. A solid support can be used synonymously with the term "bead." Wherever a solid support, such as a particle or surface, is described herein (e.g., where a barcode is immobilized on a solid support, particle, bead, etc.), another option would be to dispense a solution into the compartment so as to associate the cell label one-to-one with the compartment (and thus associate the cell label one-to-one with a single cell in the compartment). Examples of compartments may include droplets, such as microdroplets, wells (e.g., microwells), which may be on a substrate such as a multi-well plate, and chambers in a fluidic device (e.g., a microfluidic device). In some embodiments, a solution described herein may be dispensed into a compartment containing no more than one cell. The compartment may comprise at least one microdroplet, microwell, or chamber of a fluidic device, such as a microfluidic device. The microdroplet may comprise a hydrogel. The barcode may be immobilized on the substrate compartment or may be free in solution in the compartment.
[0055] As used herein, the term "probabilistic barcode" may refer to a polynucleotide sequence comprising a label of the present disclosure. A probabilistic barcode may be a polynucleotide sequence that can be used for probabilistic barcoding. A probabilistic barcode may be used to quantify a target in a sample. A probabilistic barcode may be used to control errors that may occur after associating a label with a target. For example, a probabilistic barcode may be used to evaluate amplification or sequencing errors. A probabilistic barcode associated with a target may be referred to as a probabilistic barcode target or a probabilistic barcode tag target.
[0056] As used herein, the term "gene-specific probabilistic barcode" may refer to a polynucleotide sequence that includes a label and a gene-specific target binding region. A probabilistic barcode may be a polynucleotide sequence that can be used for probabilistic barcoding. A probabilistic barcode may be used to quantify a target in a sample. A probabilistic barcode may be used to control errors that may occur after associating a label with a target. For example, a probabilistic barcode may be used to evaluate amplification or sequencing errors. A probabilistic barcode associated with a target may be referred to as a probabilistic barcode target or a probabilistic barcode tag target.
[0057] As used herein, the term "probabilistic barcoding" can refer to random labeling (e.g., barcoding) of nucleic acids. Probabilistic barcoding can use a recursive Poisson strategy to associate a label with a target and quantify the label associated with the target. As used herein, the term "probabilistic barcoding" can be used synonymously with "probabilistic labeling."
[0058] As used herein, the term "target" has its conventional and ordinary meaning in the art in light of this specification. It may refer to a composition that can be associated with a barcode (e.g., a probabilistic barcode). Exemplary targets suitable for analysis by the disclosed methods, devices, and systems include oligonucleotides, DNA, RNA, mRNA, microRNA, tRNA, and the like. A target may be single-stranded or double-stranded. In some embodiments, a target may be a protein, peptide, or polypeptide. In some embodiments, a target is a lipid. As used herein, "target" may be used synonymously with "species."
[0059] As used herein, the term "reverse transcriptase" has its conventional and ordinary meaning in the art in light of this specification. It can refer to a group of enzymes that have reverse transcriptase activity (i.e., that catalyze the synthesis of DNA from an RNA template). Generally, such enzymes include, but are not limited to, retroviral reverse transcriptases, retrotransposon reverse transcriptases, retroplasmid reverse transcriptases, retron reverse transcriptases, bacterial reverse transcriptases, group II intron-derived reverse transcriptases, and mutants, variants, or derivatives thereof. Non-retroviral reverse transcriptases include non-LTR retrotransposon reverse transcriptases, retroplasmid reverse transcriptases, retron reverse transcriptases, and group II intron reverse transcriptases. Examples of group II intron reverse transcriptases include the Lactococcus lactis LI.LtrB intron reverse transcriptase, the Thermosynechococcus elongatus TeI4c intron reverse transcriptase, or the Geobacillus stearothermophilus GsI-IIC intron reverse transcriptase. Other classes of reverse transcriptases can include many classes of non-retroviral reverse transcriptases (i.e., retrons, group II introns, and particularly diversity-generating retroelements).
[0060] The terms "universal adapter primer," "universal primer adapter," or "universal adapter sequence" are used interchangeably to refer to a nucleotide sequence that can be used to hybridize a barcode (e.g., a stochastic barcode) to generate a gene-specific barcode. The universal adapter sequence can be, for example, a known sequence that is universal for all barcodes used in the methods of the present disclosure. For example, when multiple targets are labeled using the methods disclosed herein, each of the target-specific sequences can be attached to the same universal adapter sequence. In some embodiments, two or more universal adapter sequences can be used in the methods disclosed herein. For example, when multiple targets are labeled using the methods disclosed herein, at least two of the target-specific sequences are attached to different universal adapter sequences. The universal adapter primer and its complement can be included in two oligonucleotides, one of which contains the target-specific sequence and the other of which contains the barcode. For example, the universal adapter sequence can be part of an oligonucleotide that contains a target-specific sequence to generate a nucleotide sequence complementary to the target nucleic acid. A second oligonucleotide comprising the complementary sequence of the barcode and universal adapter sequence can hybridize with the nucleotide sequence to generate a target-specific barcode (e.g., a target-specific stochastic barcode). In some embodiments, the universal adapter primer has a different sequence than the universal PCR primer used in the disclosed methods.
[0061] Barcode Barcoding, such as probabilistic barcoding, is described, for example, in U.S. Patent Application Publication No. 2015 / 0299784, WO 2015 / 031691, and Fu et al., Proc Natl Acad Sci USA 2011 May 31;108(22):9026-31 (the contents of each of these publications are incorporated herein by reference in their entirety). In some embodiments, the barcodes disclosed herein can be probabilistic barcodes, which can be polynucleotide sequences that can be used to probabilistically label (e.g., barcode, tag) targets. A barcode may be referred to as a probabilistic barcode if the ratio of the number of distinct barcode sequences in the probabilistic barcode to the number of occurrences of any of the labeled targets can be 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, or a number or range between any two of these values or an approximation of such a value. The targets may be mRNA species that include mRNA molecules with identical or nearly identical sequences. A barcode may be referred to as a probabilistic barcode if the ratio of the number of distinct barcode sequences of the probabilistic barcode to the number of occurrences of any of the labeled targets is at least or at most 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, or 100:1. The barcode sequences of a probabilistic barcode may be referred to as molecular labels.
[0062] A barcode, e.g., a probabilistic barcode, can include one or more labels. Exemplary labels can include a universal label, a cell label, a barcode sequence (e.g., a molecular label), a sample label, a plate label, a spatial label, and / or a pre-spatial label. FIG. 1 shows an exemplary barcode 104 having a spatial label. The barcode 104 can include a 5' amine that can link the barcode to a solid support 105. The barcode can include a universal label, a dimensional label, a spatial label, a cell label, and / or a molecular label. The barcode can include a universal label, a cell label, and a molecular label. The barcode can include a universal label, a spatial label, a cell label, and a molecular label. The barcode can include a universal label, a dimensional label, a cell label, and a molecular label. The order of various labels (including, but not limited to, a universal label, a dimensional label, a spatial label, a cell label, and / or a molecular label) in a barcode can vary. For example, as shown in FIG. 1, the universal label can be the 5'-most label and the molecular label can be the 3'-most label. The spatial label, dimensional label, and cellular label can be in any order. In some embodiments, the universal label, spatial label, dimensional label, cellular label, and molecular label can be in any order. The barcode can include a target binding region. The target binding region can interact with a target in a sample (e.g., a target nucleic acid, RNA, mRNA, DNA). For example, the target binding region can include an oligo(dT) sequence that can interact with the poly(A) tail of an mRNA. In some cases, the labels of the barcode (e.g., the universal label, dimensional label, spatial label, cellular label, and barcode sequence) can be separated by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more nucleotides.
[0063] Labels, e.g., cellular labels, can include a unique set of nucleic acid subsequences of a defined length, e.g., seven nucleotides each (corresponding to the number of bits used in some Hamming error-correcting codes), which can be designed to confer error-correction capabilities. The set of error-correcting subsequences includes seven nucleotide sequences, which can be designed so that any pairwise combination of sequences in the set exhibits a defined "genetic distance" (or number of mismatched bases); for example, a set of error-correcting subsequences can be designed to exhibit a genetic distance of three nucleotides. In this case, review of the error-correcting sequences within a set of sequence data for a labeled target nucleic acid molecule (described in more detail below) allows for the detection or correction of amplification or sequencing errors. In some embodiments, the length of the nucleic acid subsequences used to create the error-correcting code can vary; e.g., they can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 31, 40, 50 nucleotides in length, or a number or range between or approximating any two of these values. In some embodiments, nucleic acid subsequences of other lengths can be used to create error correcting codes.
[0064] The barcode may include a target binding region. The target binding region may interact with a target in a sample. The target may be or include ribonucleic acid (RNA), messenger RNA (mRNA), microRNA, small interfering RNA (siRNA), RNA degradation products, RNA containing a poly(A) tail, or any combination thereof. In some embodiments, the multiple targets may include deoxyribonucleic acid (DNA).
[0065] In some embodiments, the target binding region may include an oligo(dT) sequence that can interact with the poly(A) tail of mRNA. One or more of the labels of the barcode (e.g., universal label, dimensional label, spatial label, cellular label, and barcode sequence (e.g., molecular label)) may be separated from one or two of the remaining labels of the barcode by a spacer. The spacer may be, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more nucleotides. In some embodiments, none of the labels of the barcode are separated by a spacer.
[0066] Universal Signage A barcode may include one or more universal labels. In some embodiments, the one or more universal labels may be the same for all barcodes in a set of barcodes bound to a given solid support. In some embodiments, the one or more universal labels may be the same for all barcodes bound to a plurality of beads. In some embodiments, the universal label may include a nucleic acid sequence capable of hybridizing to a sequencing primer. The sequencing primer may be used to sequence the barcodes comprising the universal label. The sequencing primer (e.g., a universal sequencing primer) may include a sequencing primer associated with a high-throughput sequencing platform. In some embodiments, the universal label may include a nucleic acid sequence capable of hybridizing to a PCR primer. In some embodiments, the universal label may include a nucleic acid sequence capable of hybridizing to a sequencing primer and a PCR primer. The nucleic acid sequence of a universal label capable of hybridizing to a sequencing primer or a PCR primer may be referred to as a primer binding site. The universal label may include a sequence that can be used to initiate transcription of the barcode. The universal label may include a sequence that can be used to extend the barcode or a region within the barcode. A universal label can be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or a number or range between or approximating any two of these values. For example, a universal label can comprise at least about 10 nucleotides. A universal label can be at least or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length. In some embodiments, a cleavable linker or modified nucleotide can be part of the universal label sequence that allows for cleavage and removal of the barcode from the support.
[0067] dimensional indicator A barcode can include one or more dimensional labels. In some embodiments, a dimensional label can include a nucleic acid sequence that provides information about the dimension in which the labeling (e.g., stochastic labeling) occurred. For example, a dimensional label can provide information about the time at which a target was barcoded. A dimensional label can be associated with the time of sample barcoding (e.g., stochastic barcoding). A dimensional label can be activated at the time of labeling. Different dimensional labels can be activated at different time points. A dimensional label provides information about the order in which targets, groups of targets, and / or samples were barcoded. For example, a cell population can be barcoded in the G0 phase of the cell cycle. Cells can be pulsed again with a barcode (e.g., a stochastic barcode) in the G1 phase of the cell cycle. Cells can be pulsed again with a barcode in the S phase of the cell cycle, and so on. The barcode for each pulse (e.g., each phase of the cell cycle) can include a different dimensional label. In this way, the dimensional label provides information about which phase of the cell cycle the target was labeled in. Dimensional labels can probe many different biological time periods. Exemplary biological time periods can include, but are not limited to, cell cycle, transcription (e.g., transcription initiation), and transcript degradation. In another example, a sample (e.g., a cell, a cell population) can be labeled before and / or after drug treatment and / or therapy. Changes in the copy number of unique targets can be an indicator of the sample's response to the drug and / or therapy.
[0068] Dimensional labels may be activatable. Activatable dimensional labels may be activated at specific times. Activatable labels may, for example, be constitutively activated (e.g., do not switch off). Activatable dimensional labels may, for example, be reversibly activated (e.g., they can be switched on and off). Dimensional labels may, for example, be reversibly activatable at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more times. Dimensional labels may, for example, be reversibly activatable at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 or more times. In some embodiments, dimensional labels may be activated by fluorescence, light, a chemical event (e.g., cleavage, ligation of another molecule, addition of a modification (e.g., pegylation, sumoylation, acetylation, methylation, deacetylation, demethylation), a photochemical event (e.g., photocaging), and the introduction of a non-natural nucleotide.
[0069] In some embodiments, the dimension labels may be the same for all barcodes (e.g., stochastic barcodes) attached to a given solid support (e.g., a bead), but may be different for different solid supports (e.g., beads). In some embodiments, at least 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100% of the barcodes on the same solid support may comprise the same dimension label. In some embodiments, at least 60% of the barcodes on the same solid support may comprise the same dimension label. In some embodiments, at least 95% of the barcodes on the same solid support may comprise the same dimension label.
[0070] For multiple solid supports (e.g., beads), 10 6There may be about or more unique dimension label sequences. Dimension labels can be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or any number or range between or approximations of any two of these values. Dimension labels can be at least or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length. Dimension labels can comprise from about 5 to about 200 nucleotides. Dimension labels can comprise from about 10 to about 150 nucleotides. Dimension labels can comprise from about 20 to about 125 nucleotides in length.
[0071] spatial sign A barcode may include one or more spatial labels. In some embodiments, a spatial label may include a nucleic acid sequence that provides information about the spatial orientation of a target molecule associated with the barcode. A spatial label may be associated with a coordinate in a sample. The coordinate may be a fixed coordinate. For example, the coordinate may be fixed relative to a substrate. The spatial label may be referenced to a two-dimensional or three-dimensional grid. The coordinate may be fixed relative to a landmark. A landmark may be identifiable in space. A landmark may be an imageable structure. A landmark may be a biological structure, e.g., an anatomical landmark. A landmark may be a cellular landmark, e.g., an organelle. A landmark may be a non-natural landmark, such as a color code, a barcode, a structure with an identifiable identifier, such as magnetic, fluorescent, radioactive, or a unique size or shape. A spatial label may be associated with a physical compartment (e.g., a well, a container, or a droplet). In some embodiments, multiple spatial labels are used together to code one or more locations in space.
[0072] Spatial labels can be the same for all barcodes attached to a given solid support (e.g., a bead), but can be different for different solid supports (e.g., beads). In some embodiments, the percentage of barcodes on the same solid support that contain the same spatial label can be 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%, or a number or range between any two of these values, or an approximation of such a value. In some embodiments, the percentage of barcodes on the same solid support that contain the same spatial label can be at least or at most 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100%. In some embodiments, at least 60% of the barcodes on the same solid support can contain the same spatial label. In some embodiments, at least 95% of the barcodes on the same solid support can contain the same spatial label.
[0073] For multiple solid supports (e.g., beads), 10 6 There may be about or more unique spatial marker sequences. Spatial markers may be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or any number or range between or approximations of any two of these values. Spatial markers may be at least or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length. Spatial markers may comprise about 5 to about 200 nucleotides. Spatial markers may comprise about 10 to about 150 nucleotides. Spatial markers may comprise about 20 to about 125 nucleotides in length.
[0074] cell labeling A barcode (e.g., a probabilistic barcode) may include one or more cell labels. In some embodiments, the cell label may include a nucleic acid sequence that provides information for determining which target nucleic acid originates from which cell. In some embodiments, the cell label is the same for all barcodes attached to a given solid support (e.g., a bead) but different for different solid supports (e.g., beads). In some embodiments, the percentage of barcodes on the same solid support that contain the same cell label may be 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%, or a number or range between any two of these values, or an approximation of such a value. In some embodiments, the percentage of barcodes on the same solid support that contain the same cell label may be 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100%, or an approximation of such a value. For example, at least 60% of barcodes on the same solid support may contain the same cell label. As another example, at least 95% of the barcodes on the same solid support may contain the same cell label.
[0075] For multiple solid supports (e.g., beads), 10 6 There may be about or more unique cellular marker sequences. A cellular marker can be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or a number or range between any two of these values, or an approximation of such a value. A cellular marker can be at least or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length. For example, a cellular marker can comprise about 5 to about 200 nucleotides. As another example, a cellular marker can comprise about 10 to about 150 nucleotides. As yet another example, a cellular marker can comprise about 20 to about 125 nucleotides in length.
[0076] Barcode sequence The barcode may include one or more barcode sequences. In some embodiments, the barcode sequence may include a nucleic acid sequence that provides information for identifying the specific type of target nucleic acid species hybridized to the barcode. The barcode sequence may include a nucleic acid sequence that provides a counter (e.g., provides an approximation) for the specific presence of the target nucleic acid species hybridized to the barcode (e.g., target binding region).
[0077] In some embodiments, a diverse set of barcode sequences is attached to a given solid support (e.g., a bead). 2 , 10 3 、 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 Or there may be a number or range between any two of these values or an approximation of such values of unique molecular label sequences. For example, the plurality of barcodes may include about 6561 barcode sequences with unique sequences. As another example, the plurality of barcodes may include about 65536 barcode sequences with unique sequences. In some embodiments, there may be at least or at most 10 2 , 10 3 、 10 4 , 10 5 , 10 6 , 10 7 , 10 8 or 10 9 There may be a unique barcode sequence of 1000. The unique molecular tag sequence may be attached to a given solid support (e.g., a bead). In some embodiments, the unique molecular tag sequence is partially or entirely encompassed by a particle (e.g., a hydrogel bead).
[0078] The length of the barcode may vary in different implementations. For example, the barcode may be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or a number or range between any two of these values, or an approximation of such a value. As another example, the barcode may be at least or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length.
[0079] molecular label A barcode (e.g., a probabilistic barcode) can include one or more molecular labels. A molecular label can include a barcode sequence. In some embodiments, a molecular label can include a nucleic acid sequence that provides information for identifying the specific type of target nucleic acid species hybridized to the barcode. A molecular label can include a nucleic acid sequence that provides a counter for the specific presence of a target nucleic acid species hybridized to the barcode (e.g., a target binding region).
[0080] In some embodiments, a diverse set of molecular labels is attached to a given solid support (e.g., a bead). 2 , 10 3 、 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 Or there may be a number or range between any two of these values or an approximation of such values. For example, the plurality of barcodes may include about 6561 molecular labels with unique sequences. As another example, the plurality of barcodes may include about 65536 molecular labels with unique sequences. In some embodiments, at least or at most 10 2 , 10 3 、 10 4 , 10 5 , 10 6 , 10 7 , 10 8 or 109 There can be a unique molecular tag sequence of 100. A barcode having a unique molecular tag sequence can be attached to a given solid support (e.g., a bead).
[0081] In barcoding using multiple probabilistic barcodes (e.g., probabilistic barcoding), the ratio of the number of distinct molecular label sequences to the number of occurrences of any of the targets can be 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, or a number or range between any two of these values or an approximation of such a value. The targets can be mRNA species that include mRNA molecules with identical or nearly identical sequences. In some embodiments, the ratio of the number of different molecular label sequences to the number of occurrences of any of the targets is at least or at most 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, or 100:1.
[0082] A molecular label can be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or a number or range between or approximation of any two of these values. A molecular label can be at least or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length.
[0083] Target binding region The barcode may include one or more target binding regions, such as a capture probe. In some embodiments, the target binding region may hybridize with a target of interest. In some embodiments, the target binding region may include a nucleic acid sequence that specifically hybridizes to a target (e.g., a target nucleic acid, a target molecule, e.g., a cellular nucleic acid to be analyzed), such as a specific gene sequence. In some embodiments, the target binding region may include a nucleic acid sequence that can bind (e.g., hybridize) to a specific position of a specific target nucleic acid. In some embodiments, the target binding region may include a nucleic acid sequence capable of specific hybridization to a restriction enzyme site overhang (e.g., an EcoRI sticky end overhang). The barcode may then be ligated to any nucleic acid molecule that contains a sequence complementary to the restriction site overhang.
[0084] In some embodiments, the target binding region can include a non-specific target nucleic acid sequence. A non-specific target nucleic acid sequence can refer to a sequence that can bind to multiple target nucleic acids regardless of the specific sequence of the target nucleic acid. For example, the target binding region can include a random multimer sequence or an oligo(dT) sequence that hybridizes to the poly(A) tail of an mRNA molecule. The random multimer sequence can be, for example, a random dimer, trimer, quatramer, pentamer, hexamer, septamer, octamer, nonamer, decamer, or higher-order multimer sequence of any length. In some embodiments, the target binding region is the same for all barcodes bound to a given bead. In some embodiments, the target binding regions of multiple barcodes bound to a given bead can include two or more different target binding sequences. The target binding region can be 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or a number or range between or approximating any two of these values. The target binding region can be at most about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 or more nucleotides in length.
[0085] In some embodiments, the target binding region can include an oligo(dT) that can hybridize to an mRNA containing a polyadenylated end. The target binding region can be gene-specific. For example, the target binding region can be configured to hybridize to a specific region of the target. In some embodiments, the target binding region does not include an oligo(dT). The target binding region can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 nucleotides in length, or a number or range between any two of these values, or an approximation of such a value. The target binding region can be at least or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides in length. The target binding region can be about 5-30 nucleotides in length. When a barcode includes a gene-specific target binding region, the barcode may be referred to herein as a gene-specific barcode.
[0086] Orientation A probabilistic barcode (e.g., a probabilistic barcode) can include one or more orientations that can be used to orient (e.g., align) the barcode. The barcode can include a moiety for isoelectric focusing. Different barcodes can include different isoelectric focusing points. When these barcodes are introduced into a sample, the sample can undergo isoelectric focusing to orient the barcodes into a known configuration. In this manner, orientations can be used to create a known map of barcodes in the sample. Exemplary orientations include electrophoretic mobility (e.g., based on the size of the barcode), isoelectric point, spin, conductivity, and / or self-assembly. For example, a barcode with orientations for self-assembly can self-assemble into a specific orientation upon activation (e.g., a nucleic acid nanostructure).
[0087] affinity A barcode (e.g., a probabilistic barcode) may include one or more affinities. For example, a spatial label may include an affinity. The affinity may include a chemical and / or biological moiety that can facilitate binding of the barcode to another entity (e.g., a cellular receptor). For example, the affinity may include an antibody, e.g., an antibody specific to a particular moiety (e.g., a receptor) on a sample. In some embodiments, the antibody may direct the barcode to a particular cell type or molecule. A particular cell type or molecule and / or a target in its vicinity may be labeled (e.g., stochastically labeled). Because the antibody can direct the barcode to a specific location, the affinity, in some embodiments, can provide spatial information in addition to the nucleotide sequence of the spatial label. The antibody may be a therapeutic antibody, e.g., a monoclonal or polyclonal antibody. The antibody may be humanized or chimeric. The antibody may be a naked antibody or a fusion antibody.
[0088] Antibodies can be full-length (i.e., naturally occurring or generated by conventional immunoglobulin gene fragment recombination processes) immunoglobulin molecules (e.g., IgG antibodies) or immunoreactive (i.e., specific binding) portions of immunoglobulin molecules, such as antibody fragments.
[0089] An antibody fragment can be a portion of an antibody, such as, for example, F(ab')2, Fab', Fab, Fv, or sFv. In some embodiments, an antibody fragment can bind to the same antigen recognized by a full-length antibody. An antibody fragment can include an isolated fragment consisting of the variable region of an antibody, such as an "Fv" fragment consisting of the variable regions of the heavy and light chains, and a recombinant single-chain polypeptide molecule in which the variable regions of the light and heavy chains are connected by a peptide linker ("scFv protein"). Exemplary antibodies can include, but are not limited to, antibodies against cancer cells, antibodies against viruses, antibodies that bind to cell surface receptors (CD8, CD34, CD45), and therapeutic antibodies.
[0090] Universal Adapter Primer A barcode can include one or more universal adapter primers. For example, a gene-specific barcode, such as a gene-specific probability barcode, can include a universal adapter primer. A universal adapter primer can refer to a nucleotide sequence that is universal for all barcodes. A universal adapter primer can be used to construct a gene-specific barcode. A universal adapter primer can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 nucleotides in length, or a number or range between any two of these nucleotide lengths, or an approximation of such a value. The universal adapter primer can be at least or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides in length. The universal adapter primer can be 5 to 30 nucleotides in length.
[0091] Linker When a barcode includes two or more types of labels (e.g., two or more cellular labels or two or more barcode sequences, e.g., one molecular label), the labels may incorporate a linker label sequence. The linker label sequence may be at least about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, or more nucleotides in length. The linker label sequence may be at most about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, or more nucleotides in length. In some cases, the linker label sequence is 12 nucleotides in length. The linker label sequence may be used to facilitate synthesis of the barcode. The linker label may include an error-correcting (e.g., Hamming) code.
[0092] solid support In some embodiments, barcodes, such as the stochastic barcodes disclosed herein, can be associated with a solid support. The solid support can be, for example, a synthetic particle. In some embodiments, some or all of the barcode sequences (e.g., first barcode sequences), such as molecular labels of stochastic barcodes, of a plurality of barcodes (e.g., a first plurality of barcodes) on a solid support differ by at least one nucleotide. The cellular labels of barcodes on the same solid support can be the same. The cellular labels of barcodes on different solid supports can differ by at least one nucleotide. For example, a first cellular label of a first plurality of barcodes on a first solid support can have the same sequence, and a second cellular label of a second plurality of barcodes on a second solid support can have the same sequence. The first cellular label of a first plurality of barcodes on a first solid support and the second cellular label of a second plurality of barcodes on a second solid support can differ by at least one nucleotide. The cellular labels can be, for example, about 5-20 nucleotides in length. The barcode sequences can be, for example, about 5-20 nucleotides in length. The synthetic particles can be, for example, beads.
[0093] The beads can be, for example, silica gel beads, controlled pore glass beads, magnetic beads, Dynabeads, Sephadex / Sepharose beads, cellulose beads, polystyrene beads, or any combination thereof. The beads can comprise materials such as polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogels, paramagnetic materials, ceramics, plastics, glass, methylstyrene, acrylic polymers, titanium, latex, Sepharose, cellulose, nylon, silicone, or any combination thereof.
[0094] In some embodiments, the beads may be polymer beads, such as deformable beads or gel beads (e.g., gel beads from 10X Genomics, San Francisco, CA), functionalized with barcodes or stochastic barcodes. In some implementations, the gel beads may comprise a polymer-based gel. Gel beads may be made, for example, by encapsulating one or more polymer precursors in droplets. Gel beads may be made by exposing the polymer precursors to an accelerant (e.g., tetramethylethylenediamine (TEMED)).
[0095] In some embodiments, the particles can be disintegrable (e.g., dissolvable, degradable). For example, polymer beads can dissolve, melt, or decompose under desired conditions. The desired conditions can include environmental conditions. The desired conditions can result in the dissolution, melting, or decomposition of the polymer beads in a controlled manner. Gel beads can dissolve, melt, or decompose due to chemical, physical, biological, thermal, magnetic, electrical, or optical stimuli, or any combination thereof.
[0096] For example, analytes and / or reagents, such as oligonucleotide barcodes, can be coupled / immobilized to the interior surface of gel beads (e.g., the interior accessible by diffusion of the oligonucleotide barcodes and / or the material used to create the oligonucleotide barcodes) and / or to the exterior surface of gel beads or any other microcapsules described herein. Coupling / immobilization can be via any form of chemical bond (e.g., covalent bond, ionic bond) or physical phenomenon (e.g., van der Waals forces, dipole-dipole interactions, etc.). In some embodiments, the coupling / immobilization of reagents to gel beads or any other microcapsules described herein can be reversible, such as, for example, via a labile moiety (e.g., a chemical crosslinker, including those described herein). Upon application of a stimulus, the labile moiety can be cleaved, releasing the immobilized reagent. In some embodiments, the labile moiety is a disulfide bond. For example, if an oligonucleotide barcode is immobilized to a gel bead via a disulfide bond, exposure of the disulfide bond to a reducing agent can cleave the disulfide bond and release the oligonucleotide barcode from the bead. The labile moiety may be included as part of a gel bead or microcapsule, as part of a chemical linker connecting a reagent or analyte to the gel bead or microcapsule, and / or as part of the reagent or analyte. In some embodiments, at least one barcode of the plurality of barcodes may be immobilized on a particle, partially immobilized on a particle, encapsulated in a particle, partially encapsulated in a particle, or any combination thereof.
[0097] In some embodiments, the gel beads may comprise a variety of different polymers, including but not limited to polymers, thermosensitive polymers, photosensitive polymers, magnetic polymers, pH-sensitive polymers, salt-sensitive polymers, chemically sensitive polymers, polyelectrolytes, polysaccharides, peptides, proteins, and / or plastics. The polymer may include, but is not limited to, materials such as poly(N-isopropylacrylamide) (PNIPAAm), poly(styrene sulfonate) (PSS), poly(allylamine) (PAAm), poly(acrylic acid) (PAA), poly(ethyleneimine) (PEI), poly(diallyldimethyl-ammonium chloride) (PDADMAC), poly(pyrrole) (PPy), poly(vinylpyrrolidone) (PVPON), poly(vinylpyridine) (PVP), poly(methacrylic acid) (PMAA), poly(methyl methacrylate) (PMMA), polystyrene (PS), poly(tetrahydrofuran) (PTHF), poly(phthalaldehyde) (PTHF), poly(hexylviologen) (PHV), poly(L-lysine) (PLL), poly(L-arginine) (PARG), poly(lactic-co-glycolic acid) (PLGA).
[0098] Many chemical stimuli can be used to cause bead rupture, dissolution, or degradation. Examples of these chemical changes can include, but are not limited to, pH-mediated changes to the bead wall, bead wall collapse via chemical scission of cross-links, triggering bead wall depolymerization, and bead wall switching reactions. Bulk changes can also be used to trigger bead rupture.
[0099] Bulk or physical changes to microcapsules via various stimuli also offer many advantages in designing capsules to release reagents. Bulk or physical changes occur on a macroscopic scale, where bead rupture is the result of mechanical physical forces induced by the stimulus. These processes can include, but are not limited to, pressure-induced rupture, bead wall melting, or bead wall porosity changes.
[0100] Biological stimuli can also be used to trigger the disruption, dissolution, or degradation of beads. Generally, biological triggers are similar to chemical triggers, but in many instances, biomolecules or molecules commonly present in biological systems, such as enzymes, peptides, sugars, fatty acids, and nucleic acids, are used. For example, beads may contain polymers with peptide crosslinks that are susceptible to cleavage by specific proteases. More specifically, one example may include microcapsules containing GFLGK peptide crosslinks. Addition of a biological trigger, such as the protease cathepsin B, can cause the shell to dissolve. wall The peptide crosslinks are cleaved, releasing the contents of the bead. In other cases, the protease can be heat-activated. In another example, the beads contain a shell wall comprising cellulose. The addition of the hydrolytic enzyme chitosan serves as a biological trigger for cleavage of the cellulose bonds, depolymerization of the shell wall, and release of its internal contents.
[0101] Beads can also be induced to release their contents upon application of a thermal stimulus. A change in temperature can cause various changes in the beads. A change in heat can cause the beads to melt, causing the bead walls to collapse. In other cases, heat can increase the internal pressure of the beads' internal components, causing the beads to break or burst. In still other cases, heat can transform the beads into a shrunken, dehydrated state. Heat can also act on the thermosensitive polymers within the bead walls, causing the beads to break.
[0102] The inclusion of magnetic nanoparticles in the bead walls of microcapsules can trigger bead rupture as well as guide multiple beads. The devices of the present disclosure can include magnetic beads for any purpose. In one example, the incorporation of Fe3O4 nanoparticles into polyelectrolyte-containing beads triggers rupture in the presence of an oscillating magnetic field stimulus.
[0103] Beads can also be disrupted, dissolved, or decomposed as a result of electrical stimulation. Similar to the magnetic particles described in the previous section, electrically sensitive beads can also trigger bead rupture as well as other functions such as alignment under an electric field, conductivity, or redox reactions. In one example, beads containing electrically sensitive materials are aligned under an electric field so that the release of internal reagents can be controlled. In another example, an electric field can induce redox reactions within the bead wall itself, thereby increasing porosity.
[0104] Optical stimulation can also be used to disrupt the beads. Numerous optical triggers are possible, including systems using various molecules such as nanoparticles and chromophores that can absorb photons of specific wavelengths. For example, metal oxide coatings can be used as capsule triggers. UV irradiation of SiO2-coated polyelectrolyte capsules can result in the collapse of the bead wall. In yet another example, photoswitch materials such as azobenzene groups can be incorporated into the bead wall. Upon application of UV or visible light, chemicals such as these undergo reversible cis-trans isomerization upon absorption of a photon. In this embodiment, the incorporation of a photonic switch can cause the bead wall to collapse or become more porous upon application of the optical trigger.
[0105] For example, in a non-limiting example of barcoding (e.g., probabilistic barcoding) shown in FIG. 2, after introducing cells, such as single cells, into multiple microwells of a microwell array in block 208, beads may be introduced into multiple microwells of the microwell array in block 212. Each microwell may contain one bead. The beads may contain multiple barcodes. The barcodes may include 5' amine regions attached to the beads. The barcodes may include a universal label, a barcode sequence (e.g., a molecular label), a target binding region, or any combination thereof.
[0106] The barcodes disclosed herein can be associated with (e.g., attached to) solid supports (e.g., beads). The barcodes attached to the solid supports can include barcode sequences selected from a group including at least 100 or 1000 barcode sequences, each having a unique sequence. In some embodiments, different barcodes attached to the solid supports can include barcodes with different sequences. In some embodiments, a certain percentage of the barcodes attached to the solid supports include the same cell label. For example, the percentage can be 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%, or a number or range between any two of these values, or an approximation of such a value. As another example, the percentage can be at least or at most 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100%. In some embodiments, the barcodes attached to the solid supports can have the same cell label. The barcodes attached to the different solid supports can have different cell markers selected from a group comprising at least 100 or 1000 cell markers with unique sequences.
[0107] The barcodes disclosed herein can be associated with (e.g., attached to) a solid support (e.g., a bead). In some embodiments, barcoding multiple targets in a sample can be performed using a solid support comprising multiple synthetic particles associated with multiple barcodes. In some embodiments, the solid support can comprise multiple synthetic particles associated with multiple barcodes. The spatial labeling of multiple barcodes on different solid supports can differ by at least one nucleotide. The solid support can comprise multiple barcodes, for example, in two or three dimensions. The synthetic particles can be beads. The beads can be silica gel beads, controlled pore glass beads, magnetic beads, Dynabeads, Sephadex / Sepharose beads, cellulose beads, polystyrene beads, or any combination thereof. The solid support can comprise a polymer, matrix, hydrogel, needle array device, antibody, or any combination thereof. In some embodiments, the solid support can be free-floating. In some embodiments, the solid support can be embedded in a semi-solid or solid array. The barcodes need not be attached to the solid support. The barcodes can be individual nucleotides. The barcode may be associated with the substrate.
[0108] As used herein, the terms "tethered," "attached," and "immobilized" are used interchangeably and can refer to covalent or non-covalent means for attaching a barcode to a solid support. Any of a variety of different solid supports can be used as solid supports to attach pre-synthesized barcodes or for in situ solid phase synthesis of barcodes.
[0109] In some embodiments, the solid support is a bead. Beads may include one or more types of solid, porous, or hollow spheres, balls, bearings, cylinders, or other similar structures to which nucleic acids can be immobilized (e.g., covalently or non-covalently). Beads may be composed of, for example, plastic, ceramic, metal, polymeric materials, or any combination thereof. Beads may be or include discrete particles that are spherical (e.g., microspheres) or have non-spherical or irregular shapes, such as cubes, rectangular prisms, pyramidal, cylindrical, conical, ellipsoidal, or discoidal shapes. In some embodiments, beads may be non-spherical in shape.
[0110] The beads may comprise a variety of materials, including, but not limited to, paramagnetic materials (e.g., magnesium, molybdenum, lithium, and tantalum), superparamagnetic materials (e.g., ferrite (Fe3O4; magnetite) nanoparticles), ferromagnetic materials (e.g., iron, nickel, cobalt, some alloys thereof, and some rare earth metal compounds), ceramic, plastic, glass, polystyrene, silica, methylstyrene, acrylic polymers, titanium, latex, sepharose, agarose, hydrogels, polymers, cellulose, nylon, or any combination thereof.
[0111] In some embodiments, the beads (e.g., beads having labels attached thereto) are hydrogel beads. In some embodiments, the beads comprise a hydrogel.
[0112] Some embodiments disclosed herein include one or more particles (e.g., beads). Each of the particles may include a plurality of oligonucleotides (e.g., barcodes). Each of the plurality of oligonucleotides may include a barcode sequence (e.g., a molecular tag sequence), a cell tag, and a target binding region (e.g., an oligo(dT) sequence, a gene-specific sequence, a random multimer, or a combination thereof). The cell tag sequence of each of the plurality of oligonucleotides may be the same. The cell tag sequences of oligonucleotides on different particles may be different so that the oligonucleotides on different particles can be identified. The number of different cell tag sequences may vary in different implementations. In some embodiments, the number of cell labeling sequences is 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 , 10 7 , 10 8 , 10 9 , a number or range between, or exceeding, any two of these values, or an approximation of such a value. In some embodiments, the number of cell labeling sequences is at least or at most 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 , 10 7 , 10 8 or 10 9In some embodiments, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 or more of the plurality of particles comprise oligonucleotides having the same cellular sequence. In some embodiments, at most 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10% or more of the plurality of particles comprise oligonucleotides having the same cellular sequence. In some embodiments, none of the plurality of particles comprise the same cellular targeting sequence.
[0113] The multiple oligonucleotides on each particle can include different barcode sequences (e.g., molecular labels). In some embodiments, the number of barcode sequences is 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 , 10 7 , 10 8 , 10 9 In some embodiments, the number of barcode sequences is at least or at most 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 , 10 7 , 10 8 or 10 9For example, at least 100 of the plurality of oligonucleotides may comprise different barcode sequences. As another example, in a single particle, at least 100, 500, 1000, 5000, 10000, 15000, 20000, 50000, or a number or range between any two of these values, or more, of the plurality of oligonucleotides may comprise different barcode sequences. Some embodiments provide a plurality of particles comprising barcodes. In some embodiments, the ratio of the presence (or copies or number) of labeled targets to distinct barcode sequences can be at least 1:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:11, 1:12, 1:13, 1:14, 1:15, 1:16, 1:17, 1:18, 1:19, 1:20, 1:30, 1:40, 1:50, 1:60, 1:70, 1:80, 1:90, or more. In some embodiments, each of the plurality of oligonucleotides further comprises a sample label, a universal label, or both. The particles can be, for example, nanoparticles or microparticles.
[0114] The size of the beads can vary. For example, the diameter of the beads can range from 0.1 micrometers to 50 micrometers. In some embodiments, the diameter of the beads can be 0.1, 0.5, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50 micrometers, or a number or range between any two of these values, or an approximation of such value.
[0115] The diameter of a bead can be related to the diameter of a well in a substrate. In some embodiments, the diameter of a bead can be 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100% longer or shorter than the diameter of the well, or a number or range between any two of these values, or an approximation of such a value. The diameter of a bead can be related to the diameter of a cell (e.g., a single cell confined in a well in a substrate). In some embodiments, the diameter of a bead can be at least or at most 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% longer or shorter than the diameter of the well. The diameter of a bead can be related to the diameter of a cell (e.g., a single cell confined in a well in a substrate). In some embodiments, the diameter of a bead can be 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 150%, 200%, 250%, 300% longer or shorter than the diameter of a cell, or a number or range between any two of these values or an approximation of such a value. In some embodiments, the diameter of a bead can be at least or at most 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 150%, 200%, 250%, or 300% longer or shorter than the diameter of a cell.
[0116] The beads can be bound and / or embedded in a substrate. The beads can be bound and / or embedded in a gel, hydrogel, polymer, and / or matrix. The spatial location of the beads within the substrate (e.g., gel, matrix, scaffold, or polymer) can be identified using spatial labels present in barcodes on the beads, which can serve as location addresses.
[0117] Examples of beads may include, but are not limited to, streptavidin beads, agarose beads, magnetic beads, Dynabeads®, MACS® microbeads, antibody-conjugated beads (e.g., anti-immunoglobulin microbeads), protein A-conjugated beads, protein G-conjugated beads, protein A / G-conjugated beads, protein L-conjugated beads, oligo(dT)-conjugated beads, silica beads, silica-like beads, anti-biotin microbeads, anti-fluorescent dye microbeads, and BcMag™ carboxyl-terminated magnetic beads.
[0118] The beads may be associated with (e.g., impregnated with) quantum dots or fluorescent dyes to fluoresce in one fluorescent optical channel or multiple optical channels. The beads may be associated with iron oxide or chromium oxide to make them paramagnetic or ferromagnetic. The beads may be identifiable. For example, the beads may be imaged using a camera. The beads may have a detectable code associated with them. For example, the beads may include a barcode. The beads may change size, for example, due to swelling in an organic or inorganic solution. The beads may be hydrophobic. The beads may be hydrophilic. The beads may be biocompatible.
[0119] The solid support (e.g., a bead) can be visualized. The solid support can include a visualization tag (e.g., a fluorescent dye). The solid support (e.g., a bead) can be etched with an identifier (e.g., a number). The identifier can be visualized by imaging the bead.
[0120] A solid support can comprise an insoluble, semi-soluble, or insoluble material. A solid support can be referred to as "functionalized" if it contains a linker, backbone, building block, or other reactive moiety attached thereto, while a solid support can be "non-functionalized" if it does not contain such a reactive moiety attached thereto. A solid support can be used free in solution, such as in a microtiter well format; in a flow-through format, such as in a column; or in a dipstick format.
[0121] The solid support may comprise a membrane, paper, plastic, coated surface, flat surface, glass, slide, chip, or any combination thereof. The solid support may take the form of a resin, gel, microsphere, or other geometric shape. The solid support may comprise a silica chip, microparticle, nanoparticle, plate, array, capillary tube, flat support such as a glass fiber filter, glass surface, metal surface (steel, gold, silver, aluminum, silicon, and copper), glass support, plastic support, silicon support, chip, filter, membrane, microwell plate, slide, plastic material (including multiwell plates or membranes (e.g., formed of polyethylene, polypropylene, polyamide, polyvinylidene difluoride)), and / or a wafer, comb, pin, or needle (e.g., an array of pins suitable for combinatorial synthesis or analysis), or beads in a series of depressions or nanoliter wells on a flat surface such as a wafer (e.g., a silicon wafer), a wafer with depressions with or without a filter bottom.
[0122] The solid support may comprise a polymer matrix (e.g., a gel, a hydrogel). The polymer matrix may be capable of penetrating intracellular spaces (e.g., around organelles). The polymer matrix may also be capable of being transported throughout the circulatory system.
[0123] Substrates and microwell arrays As used herein, a substrate may refer to a type of solid support. A substrate may refer to a solid support that may include a barcode or stochastic barcode of the present disclosure. A substrate may include, for example, a plurality of microwells. For example, a substrate may be a well array including two or more microwells. In some embodiments, a microwell may include a small reaction chamber of a defined volume. In some embodiments, a microwell may confine one or more cells. In some embodiments, a microwell may confine only one cell. In some embodiments, a microwell may confine one or more solid supports. In some embodiments, a microwell may confine only one solid support. In some embodiments, a microwell confines a single cell and a single solid support (e.g., a bead). A microwell may include a barcode reagent of the present disclosure.
[0124] Barcoding methods The present disclosure provides methods for estimating the number of unique targets at unique locations in a bodily sample (e.g., tissue, organ, tumor, cell). The methods may include placing a barcode (e.g., a probabilistic barcode) in proximity to the sample, lysing the sample, associating the unique targets with the barcode, amplifying the targets, and / or digitally counting the targets. The methods may further include analyzing and / or visualizing information obtained from the spatial labeling of the barcode. In some embodiments, the methods include visualizing multiple targets in the sample. Mapping the multiple targets to a map of the sample may include creating a two-dimensional or three-dimensional map of the sample. The two-dimensional and three-dimensional maps may be created before or after barcoding (e.g., probabilistic barcoding) the multiple targets in the sample. Visualizing multiple targets in the sample may include mapping the multiple targets to a map of the sample. Mapping the multiple targets to a map of the sample may include creating a two-dimensional or three-dimensional map of the sample. The two-dimensional and three-dimensional maps can be generated before or after barcoding multiple targets in a sample. In some embodiments, the two-dimensional and three-dimensional maps can be generated before or after lysing the sample. Lysing the sample before or after generating the two-dimensional or three-dimensional map can include heating the sample, contacting the sample with a detergent, changing the pH of the sample, or any combination thereof.
[0125] In some embodiments, barcoding the plurality of targets comprises hybridizing the plurality of barcodes to the plurality of targets to generate barcoded targets (e.g., stochastically barcoded targets). Barcoding the plurality of targets may comprise creating an indexed library of barcoded targets. Creating an indexed library of barcoded targets may be performed using a solid support comprising a plurality of barcodes (e.g., stochastic barcodes).
[0126] Contacting the sample and barcode The present disclosure provides methods for contacting a sample (e.g., cells) with a substrate of the present disclosure. For example, a sample including cells, an organ, or a tissue slice can be contacted with a barcode (e.g., a stochastic barcode). The cells can be contacted, for example, by gravity flow, in which case the cells can settle to form a monolayer. The sample can be a tissue slice. The slice can be disposed on a substrate. The sample can be one-dimensional (e.g., forming a planar surface). The sample (e.g., cells) can be spread across the substrate, for example, by growing / culturing the cells on the substrate.
[0127] When the barcode is in proximity to the target, the target can hybridize to the barcode. The barcodes can be contacted in a non-depleting ratio so that each unique target can bind to a unique barcode of the present disclosure. To ensure efficient binding between the target and the barcode, the target can be cross-linked to the barcode.
[0128] Cell lysis After partitioning the cells and barcodes, the cells can be lysed to release the target molecule. Cell lysis can be achieved by any of a variety of means, such as chemical or biochemical means, osmotic shock, or thermal lysis, mechanical lysis, or optical lysis. Cells can be lysed by adding a cell lysis buffer containing a detergent (e.g., SDS, Li dodecyl sulfate, Triton X-100, Tween-20, or NP-40), an organic solvent (e.g., methanol or acetone), or a digestive enzyme (e.g., proteinase K, pepsin, or trypsin), or any combination thereof. To improve target and barcode association, the diffusion rate of the target molecule can be altered, for example, by lowering the temperature and / or increasing the viscosity of the lysate.
[0129] In some embodiments, the sample can be lysed using filter paper, which can be soaked with a lysis buffer over the filter paper, and the filter paper can be applied to the sample with pressure, which can promote lysis of the sample and hybridization of the target of the sample to the substrate.
[0130] In some embodiments, lysis may be performed by mechanical, thermal, optical, and / or chemical lysis. Chemical lysis may include the use of digestive enzymes such as proteinase K, pepsin, and trypsin. Lysis may be performed by adding a lysis buffer to the substrate. The lysis buffer may include Tris-HCl. The lysis buffer may include at least about 0.01, 0.05, 0.1, 0.5, or 1 M or more Tris-HCl. The lysis buffer may include at most about 0.01, 0.05, 0.1, 0.5, or 1 M or more Tris-HCl. The lysis buffer may include about 0.1 M Tris-HCl. The pH of the lysis buffer may be at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more. The pH of the lysis buffer may be at most about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more. In some embodiments, the pH of the lysis buffer is about 7.5. The lysis buffer may include a salt (e.g., LiCl). The concentration of the salt in the lysis buffer may be at least about 0.1, 0.5, or 1 M or more. The concentration of the salt in the lysis buffer may be at most about 0.1, 0.5, or 1 M or more. In some embodiments, the concentration of the salt in the lysis buffer is about 0.5 M. The lysis buffer may include a detergent (e.g., SDS, Li dodecyl sulfate, triton X, tween, NP-40). The concentration of the detergent in the lysis buffer may be at least about 0.0001%, 0.0005%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, or 7% or more. The concentration of the detergent in the lysis buffer can be at most about 0.0001%, 0.0005%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, or 7% or more. In some embodiments, the concentration of the detergent in the lysis buffer is about 1% Li dodecyl sulfate. The time used in the method for lysis can depend on the amount of detergent used. In some embodiments, the more detergent used, the shorter the time required for lysis. The lysis buffer can include a chelating agent (e.g., EDTA, EGTA).The concentration of the chelating agent in the lysis buffer may be at least about 1, 5, 10, 15, 20, 25, or 30 mM or more. The concentration of the chelating agent in the lysis buffer may be at most about 1, 5, 10, 15, 20, 25, or 30 mM or more. In some embodiments, the concentration of the chelating agent in the lysis buffer is about 10 mM. The lysis buffer may include a reducing reagent (e.g., β-mercaptoethanol, DTT). The concentration of the reducing reagent in the lysis buffer may be at least about 1, 5, 10, 15, or 20 mM or more. The concentration of the reducing reagent in the lysis buffer may be at most about 1, 5, 10, 15, or 20 mM or more. In some embodiments, the concentration of the reducing reagent in the lysis buffer is about 5 mM. In some embodiments, the lysis buffer may comprise about 0.1 M Tris-HCl, about pH 7.5, about 0.5 M LiCl, about 1% lithium dodecyl sulfate, about 10 mM EDTA, and about 5 mM DTT.
[0131] Lysing may be performed at a temperature of about 4, 10, 15, 20, 25, or 30° C. Lysing may be performed for about 1, 5, 10, 15, or 20 minutes or more. Lysed cells may contain at least about 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, or 700,000 or more target nucleic acid molecules. Lysed cells may contain at most about 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, or 700,000 or more target nucleic acid molecules.
[0132] Attaching the barcode to the target nucleic acid molecule After cell lysis and release of nucleic acid molecules therefrom, the nucleic acid molecules can be randomly associated with the barcodes on the co-localized solid support. Association can involve hybridization of the target recognition region of the barcode to a complementary portion of the target nucleic acid molecule (e.g., the oligo(dT) of the barcode can interact with the poly(A) tail of the target). Assay conditions (e.g., buffer pH, ionic strength, temperature, etc.) used for hybridization can be selected to promote the formation of specific, stable hybrids. In some embodiments, nucleic acid molecules released from lysed cells can be associated with (e.g., hybridized to) multiple probes on a substrate. If the probes contain oligo(dT), mRNA molecules can hybridize to the probes and be reverse transcribed. The oligo(dT) portion of the oligonucleotide can act as a primer for first-strand synthesis of cDNA molecules. For example, in the non-limiting example of barcoding shown in Figure 2, block 216, mRNA molecules can hybridize to barcodes on beads. For example, a single-stranded nucleotide fragment can hybridize to the target binding region of the barcode.
[0133] The binding may further include ligating the target recognition region of the barcode and a portion of the target nucleic acid molecule. For example, the target binding region may include a nucleic acid sequence capable of specific hybridization to a restriction site overhang (e.g., an EcoRI sticky end overhang). The assay procedure may further include treating the target nucleic acid with a restriction enzyme (e.g., EcoRI) to generate the restriction site overhang. The barcode may then be ligated to any nucleic acid molecule containing a sequence complementary to the restriction site overhang. A ligase (e.g., T4 DNA ligase) may be used to link the two fragments.
[0134] For example, in a non-limiting example of barcoding shown in Figure 2, block 220, labeled targets (e.g., target-barcode molecules) from multiple cells (or multiple samples) can then be pooled, e.g., in a tube. The labeled targets can be pooled, e.g., by collecting beads to which the barcodes and / or target-barcode molecules are bound.
[0135] Solid support-based collection recovery of bound target-barcode molecules can be achieved through the use of magnetic beads and an externally applied magnetic field. After the target-barcode molecules are pooled, all further processing can proceed in a single reaction vessel. Further processing can include, for example, reverse transcription reactions, amplification reactions, cleavage reactions, dissociation reactions, and / or nucleic acid extension reactions. Further processing reactions can be performed within microwells, i.e., without first pooling the labeled target nucleic acid molecules from multiple cells.
[0136] Reverse transcription The present disclosure provides a method for generating a target-barcode conjugate using reverse transcription (e.g., at block 224 of Figure 2). The target-barcode conjugate can include a barcode and a complementary sequence of all or part of a target nucleic acid (i.e., a barcoded cDNA molecule, such as a stochastically barcoded cDNA molecule). Reverse transcription of the associated RNA molecule can occur by adding a reverse transcription primer along with a reverse transcriptase. The reverse transcription primer can be an oligo(dT) primer, a random hexanucleotide primer, or a target-specific oligonucleotide primer. The oligo(dT) primer can be 12-18 nucleotides in length or approximately such nucleotides and binds to the endogenous poly(A) tail at the 3' end of mammalian mRNA. The random hexanucleotide primer can bind to the mRNA at various complementary sites. The target-specific oligonucleotide primer typically selectively primes the mRNA of interest.
[0137] In some embodiments, reverse transcription of the labeled RNA molecule can occur by adding a reverse transcription primer. In some embodiments, the reverse transcription primer is an oligo(dT) primer, a random hexanucleotide primer, or a target-specific oligonucleotide primer. Generally, oligo(dT) primers are 12-18 nucleotides in length and bind to the endogenous poly(A) tail at the 3' end of mammalian mRNAs. Random hexanucleotide primers can bind to mRNAs at various complementary sites. Target-specific oligonucleotide primers typically selectively prime the mRNA of interest.
[0138] Reverse transcription can be performed repeatedly to generate multiple labeled cDNA molecules. The methods disclosed herein can include performing at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 reverse transcription reactions. The methods can include performing at least about 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 reverse transcription reactions.
[0139] amplification One or more nucleic acid amplification reactions (e.g., at block 228 of FIG. 2 ) can be performed to generate multiple copies of the labeled target nucleic acid molecule. Amplification can be performed in a multiplexed manner, where multiple target nucleic acid sequences are amplified simultaneously. The amplification reaction can be used to add sequencing adapters to the nucleic acid molecule. The amplification reaction can include amplifying at least a portion of the sample label, if present. The amplification reaction can include amplifying at least a portion of the cell label and / or barcode sequence (e.g., molecular label). The amplification reaction can include amplifying at least a portion of the sample tag, cell label, spatial label, barcode sequence (e.g., molecular label), target nucleic acid, or a combination thereof. The amplification reaction may include amplifying 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 100%, or a range or number between any two of these values of the plurality of nucleic acids. The method may further include performing one or more cDNA synthesis reactions to generate one or more cDNA copies of the target-barcode molecule comprising the sample label, cell label, spatial label, and / or barcode sequence (e.g., molecular label).
[0140] In some embodiments, amplification can be performed using polymerase chain reaction (PCR). As used herein, PCR can refer to a reaction for in vitro amplification of specific DNA sequences by simultaneous primer extension of complementary strands of DNA. As used herein, PCR can encompass derivatives of the reaction, including, but not limited to, RT-PCR, real-time PCR, nested PCR, quantitative PCR, multiplex PCR, digital PCR, and assembly PCR.
[0141] Amplification of the labeled nucleic acid may include non-PCR-based methods. Examples of non-PCR-based methods include, but are not limited to, multiple displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), real-time SDA, rolling circle amplification, or circle-circle amplification. Other non-PCR-based amplification methods include DNA-dependent RNA polymerase-driven RNA transcription amplification or multiple cycles of RNA-directed DNA synthesis and transcription to amplify DNA or RNA targets, ligase chain reaction (LCR) and Qβ replicase (Qβ) methods, the use of palindromic probes, strand displacement amplification, oligonucleotide-driven amplification using restriction endonucleases, amplification methods in which a primer is hybridized to a nucleic acid sequence and the resulting duplex is cleaved before extension and amplification, strand displacement amplification using a nucleic acid polymerase lacking 5' exonuclease activity, rolling circle amplification, and branched extension amplification (RAM). In some embodiments, amplification does not produce circularized transcripts.
[0142] In some embodiments, the methods disclosed herein further include performing a polymerase chain reaction on the labeled nucleic acid (e.g., labeled RNA, labeled DNA, labeled cDNA) to generate a labeled amplicon (e.g., a stochastically labeled amplicon). The labeled amplicon can be a double-stranded molecule. The double-stranded molecule can include a double-stranded RNA molecule, a double-stranded DNA molecule, or an RNA molecule hybridized to a DNA molecule. One or both strands of the double-stranded molecule can include a sample label, a spatial label, a cell label, and / or a barcode sequence (e.g., a molecular label). The labeled amplicon can be a single-stranded molecule. The single-stranded molecule can include DNA, RNA, or a combination thereof. The nucleic acids of the present disclosure can include synthetic or modified nucleic acids.
[0143] Amplification may include the use of one or more non-natural nucleotides. Non-natural nucleotides may include photolabile or trigger nucleotides. Examples of non-natural nucleotides may include, but are not limited to, peptide nucleic acids (PNAs), morpholinos, locked nucleic acids (LNAs), glycol nucleic acids (GNAs), and threose nucleic acids (TNAs). Non-natural nucleotides may be added to one or more cycles of the amplification reaction. The addition of non-natural nucleotides may be used to identify products at specific cycles or time points of the amplification reaction.
[0144] The step of performing one or more amplification reactions may include the use of one or more primers. The one or more primers may comprise, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides or more. The one or more primers may comprise at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides or more. The one or more primers may comprise 12 to fewer than 15 nucleotides. The one or more primers may anneal to at least a portion of the multiple labeled targets (e.g., stochastically labeled targets). The one or more primers may anneal to the 3' or 5' ends of the multiple labeled targets. The one or more primers may anneal to an internal region of the multiple labeled targets. The internal region can be at least about 50, 100, 150, 200, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 650, 700, 750, 800, 850, 900, or 1000 nucleotides from the 3' end of the multiple labeled targets. The one or more primers can comprise a fixed panel of primers. The one or more primers can include at least one or more custom primers. The one or more primers may include at least one or more control primers. The one or more primers may include at least one or more gene-specific primers.
[0145] The one or more primers may include a universal primer. The universal primer may anneal to a universal primer binding site. The one or more custom primers may anneal to a first sample label, a second sample label, a spatial label, a cell label, a barcode sequence (e.g., a molecular label), a target, or any combination thereof. The one or more primers may include a universal primer and a custom primer. The custom primers may be designed to amplify one or more targets. The targets may comprise a subset of all nucleic acids in one or more samples. The targets may comprise a subset of all labeled targets in one or more samples. The one or more primers may include at least 96 or more custom primers. The one or more primers may include at least 960 or more custom primers. The one or more primers may include at least 9600 or more custom primers. The one or more custom primers may anneal to two or more different labeled nucleic acids. The two or more different labeled nucleic acids may correspond to one or more genes.
[0146] Any amplification scheme can be used in the disclosed methods. For example, in one scheme, a first round of PCR can amplify molecules bound to beads using a gene-specific primer and a primer for the universal Illumina sequencing primer 1 sequence. A second round of PCR can amplify the first PCR product using a nested gene-specific primer and a primer for the universal Illumina sequencing primer 1 sequence flanked by Illumina sequencing primer 2 sequences. A third round of PCR adds P5, P7, and a sample index to convert the PCR products into an Illumina sequencing library. Sequencing using 150 bp x 2 sequencing can reveal cell label and barcode sequences (e.g., molecular labels) on read 1, genes on read 2, and a sample index on read index 1.
[0147] In some embodiments, nucleic acids can be removed from a substrate using chemical cleavage. For example, chemical groups or modified bases present in the nucleic acid can be used to facilitate its removal from the solid support. For example, enzymes can be used to remove nucleic acids from a substrate. For example, nucleic acids can be removed from a substrate by restriction endonuclease digestion. For example, treatment of nucleic acids containing dUTP or ddUTP with uracil-d-glycosylase (UDG) can be used to remove nucleic acids from a substrate. For example, nucleic acids can be removed from a substrate using enzymes that perform nucleotide excision, such as base excision repair enzymes, such as apurinic / apyrimidinic (AP) endonucleases. In some embodiments, nucleic acids can be removed from a substrate using photocleavable groups and light. In some embodiments, cleavable linkers can be used to remove nucleic acids from a substrate. For example, the cleavable linker can comprise at least one of biotin / avidin, biotin / streptavidin, biotin / neutravidin, Ig-Protein A, a photolabile linker, an acid or base labile linker group, or an aptamer.
[0148] If the probe is gene-specific, the molecule can be hybridized to the probe and reverse transcribed and / or amplified. In some embodiments, after the nucleic acid is synthesized (e.g., reverse transcribed), it can be amplified. Amplification can be performed in a multiplexed manner, where multiple target nucleic acid sequences are amplified simultaneously. Amplification can add sequencing adapters to the nucleic acid.
[0149] In some embodiments, amplification can be performed on the substrate, for example, using bridge amplification. The cDNA can be homopolymer tailed to generate ends compatible with bridge amplification on the substrate using an oligo(dT) probe. In bridge amplification, a primer complementary to the 3' end of the template nucleic acid can be the first primer of each pair covalently attached to a solid particle. When a sample containing the template nucleic acid is contacted with the particle and one thermal cycle is performed, the template molecule can anneal to the first primer, and the first primer can be extended in the forward direction by the addition of nucleotides to form a double-stranded molecule consisting of the template molecule and a newly formed DNA strand complementary to the template. In the heating step of the next cycle, the double-stranded molecule can be denatured, releasing the template molecule from the particle and leaving a complementary DNA strand attached to the particle via the first primer. In the annealing stage of the subsequent annealing and extension step, the complementary strand can hybridize to a second primer complementary to the segment of the complementary strand at the position removed from the first primer. This hybridization allows the complementary strand to form a bridge between the first and second primers, covalently bound to the first primer and hybridized to the second primer. In the extension step, the second primer can be extended in the opposite direction by adding nucleotides to the same reaction mixture, thereby converting the bridge into a double-stranded bridge. The next cycle then begins, and the double-stranded bridge is denatured to produce two single-stranded nucleic acid molecules, each with one end bound to the particle surface via the first and second primers and the other end unbound, respectively. In the annealing and extension step of this second cycle, each strand can hybridize to additional, previously unused, complementary primers on the same particle to form a new single-stranded bridge. The two previously unused hybridized primers can then be extended to convert the two new bridges into double-stranded bridges.
[0150] The amplification reaction can include amplifying at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97% or 100% of the plurality of nucleic acids.
[0151] Amplification of the labeled nucleic acid may include PCR-based or non-PCR-based methods. Amplification of the labeled nucleic acid may include exponential amplification of the labeled nucleic acid. Amplification of the labeled nucleic acid may include linear amplification of the labeled nucleic acid. Amplification may be performed by polymerase chain reaction (PCR). PCR may refer to a reaction for in vitro amplification of specific DNA sequences by simultaneous primer extension of complementary strands of DNA. PCR may encompass derivatives of the reaction, including, but not limited to, RT-PCR, real-time PCR, nested PCR, quantitative PCR, multiplex PCR, digital PCR, suppression PCR, semi-suppressive PCR, and assembly PCR.
[0152] In some embodiments, amplification of the labeled nucleic acid comprises a non-PCR-based method. Examples of non-PCR-based methods include, but are not limited to, multiple displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), real-time SDA, rolling circle amplification, or circle-circle amplification. Other non-PCR-based amplification methods include DNA-dependent RNA polymerase-driven RNA transcription amplification or multiple cycles of RNA-directed DNA synthesis and transcription to amplify DNA or RNA targets, ligase chain reaction (LCR), Qβ replicase (Qβ), the use of palindromic probes, strand displacement amplification, oligonucleotide-driven amplification using restriction endonucleases, amplification methods in which a primer is hybridized to a nucleic acid sequence and the resulting duplex is cleaved before extension and amplification, strand displacement amplification using a nucleic acid polymerase lacking 5' exonuclease activity, rolling circle amplification, and / or branched extension amplification (RAM).
[0153] In some embodiments, the methods disclosed herein further comprise performing a nested polymerase chain reaction on the amplified amplicon (e.g., target). The amplicon may be a double-stranded molecule. The double-stranded molecule may comprise a double-stranded RNA molecule, a double-stranded DNA molecule, or an RNA molecule hybridized to a DNA molecule. One or both strands of the double-stranded molecule may comprise a sample tag or molecular identifier label. Alternatively, the amplicon may be a single-stranded molecule. The single-stranded molecule may comprise DNA, RNA, or a combination thereof. The nucleic acids described herein may include synthetic or modified nucleic acids.
[0154] In some embodiments, the methods include repeatedly amplifying a labeled nucleic acid to generate multiple amplicons. The methods disclosed herein may include performing at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amplification reactions. Alternatively, the methods include performing at least about 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 amplification reactions.
[0155] The amplification may further include adding one or more control nucleic acids to one or more samples containing the plurality of nucleic acids. The amplification may further include adding one or more control nucleic acids to the plurality of nucleic acids. The control nucleic acids may include a control label.
[0156] Amplification may involve the use of one or more non-natural nucleotides. Non-natural nucleotides may include photolabile and / or trigger nucleotides. Examples of non-natural nucleotides include, but are not limited to, peptide nucleic acids (PNAs), morpholinos, locked nucleic acids (LNAs), glycol nucleic acids (GNAs), and threose nucleic acids (TNAs). Non-natural nucleotides may be added in one or more cycles of the amplification reaction. The addition of non-natural nucleotides may be used to identify products at specific cycles or time points of the amplification reaction.
[0157] The step of performing one or more amplification reactions may include the use of one or more primers. The one or more primers may comprise one or more oligonucleotides. The one or more oligonucleotides may comprise at least about 7 to 9 nucleotides. The one or more oligonucleotides may comprise less than 12 to 15 nucleotides. The one or more primers may anneal to at least a portion of the plurality of labeled nucleic acids. The one or more primers may anneal to the 3' and / or 5' ends of the plurality of labeled nucleic acids. The one or more primers may anneal to an internal region of the plurality of labeled nucleic acids. The internal region can be at least about 50, 100, 150, 200, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 650, 700, 750, 800, 850, 900, or 1000 nucleotides from the 3' end of the plurality of labeled nucleic acids. The one or more primers can comprise a fixed panel of primers. The one or more primers can comprise at least one or more custom primers. The one or more primers may include at least one or more control primers. The one or more primers may include at least one or more housekeeping gene primers. The one or more primers may include a universal primer. The universal primer may anneal to a universal primer binding site. The one or more custom primers may anneal to a first sample tag, a second sample tag, a molecular identifier label, a nucleic acid, or a product thereof. The one or more primers may include a universal primer and a custom primer. The custom primer may be designed to amplify one or more target nucleic acids. The target nucleic acids may comprise a subset of the total nucleic acids in one or more samples. In some embodiments, the primers are probes attached to the array of the present disclosure.
[0158] In some embodiments, barcoding (e.g., stochastically barcoding) a plurality of targets in a sample further comprises generating an indexed library of barcoded targets (e.g., stochastically barcoded targets) or barcoded fragments of the targets. The barcode sequences of different barcodes (e.g., molecular labels of different stochastic barcodes) can differ from one another. Generating an indexed library of barcoded targets comprises generating a plurality of indexed polynucleotides from the plurality of targets in the sample. For example, for an indexed library of barcoded targets comprising a first indexed target and a second indexed target, the labeled region of the first indexed polynucleotide can differ from the labeled region of the second indexed polynucleotide by, or by about at least or at most, such a value, or a number or range of nucleotides between any two of these values. In some embodiments, the step of generating an indexed library of barcoded targets includes contacting a plurality of targets, e.g., mRNA molecules, with a plurality of oligonucleotides each comprising a poly(T) region and a label region; and performing first-strand synthesis using a reverse transcriptase to generate single-stranded, labeled cDNA molecules, each comprising a cDNA region and a label region, wherein the plurality of targets comprises at least two mRNA molecules of different sequences and the plurality of oligonucleotides comprises at least two oligonucleotides of different sequences. The step of generating an indexed library of barcoded targets may further include amplifying the single-stranded, labeled cDNA molecules to generate double-stranded, labeled cDNA molecules; and performing nested PCR on the double-stranded, labeled cDNA molecules to generate labeled amplicons. In some embodiments, the method may include generating adapter-labeled amplicons.
[0159] Barcoding (e.g., probabilistic barcoding) can involve using nucleic acid barcodes or tags to label individual nucleic acid (e.g., DNA or RNA) molecules. In some embodiments, it involves adding DNA barcodes or tags to cDNA molecules as they are generated from mRNA. Nested PCR can be performed to minimize PCR amplification bias. Adapters can be added for sequencing, e.g., using next-generation sequencing (NGS). Sequencing results can be used to determine the sequence of cellular labels, molecular labels, and nucleotide fragments of one or more copies of the target, e.g., in block 232 of FIG. 2.
[0160] 3 is a schematic diagram illustrating a non-limiting, exemplary process for generating an indexed library of barcoded targets (e.g., stochastically barcoded targets), such as barcoded mRNAs or fragments thereof. As shown in step 1, the reverse transcription process can encode each mRNA molecule containing a unique molecular tag sequence, a cellular tag sequence, and a universal PCR site. In particular, an RNA molecule 302 can be reverse transcribed by hybridization (e.g., stochastic hybridization) of a set of barcodes (e.g., stochastic barcodes) 310 to a poly(A) tail region 308 of the RNA molecule 302 to generate labeled cDNA molecules 304 containing cDNA regions 306. Each of the barcodes 310 can include a target binding region, e.g., a poly(dT) region 312, a tag region 314 (e.g., a barcode sequence or molecule), and a universal PCR region 316.
[0161] In some embodiments, the cell label sequence can comprise 3 to 20 nucleotides. In some embodiments, the molecular label sequence can comprise 3 to 20 nucleotides. In some embodiments, each of the plurality of stochastic barcodes further comprises one or more of a universal label and a cell label, wherein the universal label is the same for the plurality of stochastic barcodes on the solid support and the cell label is the same for the plurality of stochastic barcodes on the solid support. In some embodiments, the universal label can comprise 3 to 20 nucleotides. In some embodiments, the cell label comprises 3 to 20 nucleotides.
[0162] In some embodiments, label region 314 can include a barcode sequence or molecular label 318 and a cell label 320. In some embodiments, label region 314 can include one or more of a universal label, a dimensional label, and a cell label. Barcode sequence or molecular label 318 can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides in length, or approximately at least or at most such a number of nucleotides in length, or a number or range of nucleotides in length between any of these values. Cell label 320 can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides in length, or approximately at least or at most such a number of nucleotides in length, or a number or range of nucleotides in length between any of these values. The universal label can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides in length, or about at least or at most such values, or a number or range of nucleotides in length between any of these values. The universal label can be the same for multiple probabilistic barcodes on a solid support, and the cell label is the same for multiple probabilistic barcodes on a solid support. The dimensional labels can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides in length, or about at least or at most such values, or a number or range of nucleotides in length between any of these values.
[0163] In some embodiments, label region 314 can include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 different labels, or about such a number of different labels, or at least such a number of different labels, or a number or range of different labels between any of these values, such as barcode sequences or molecular labels 318 and cellular labels 320. Each label can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides in length, or about at least such a number of different labels, or a number or range of different labels between any of these values. The set of barcodes or probabilistic barcodes 310 may be: 10, 20, 40, 50, 70, 80, 90, 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 20 The set of barcodes or probabilistic barcodes 310 may include at least about or at most such number of barcodes or probabilistic barcodes 310, or a number or range of barcodes or probabilistic barcodes 310 between any of these values. The set of barcodes or probabilistic barcodes 310 may, for example, each include a unique labeled region 314. The labeled cDNA molecules 304 may be purified to remove excess barcodes or probabilistic barcodes 310. Purification may include Ampure bead purification.
[0164] As shown in step 2, the products from the reverse transcription process in step 1 can be pooled in one tube and PCR amplified using a first pool of PCR primers and a first universal PCR primer. Pooling is possible due to the uniquely labeled region 314. In particular, the labeled cDNA molecules 304 can be amplified to generate nested PCR-labeled amplicons 322. The amplification can include multiplex PCR amplification. The amplification can include multiplex PCR amplification using 96 multiplex primers in a single reaction volume. In some embodiments, the multiplex PCR amplification can be performed using 10, 20, 40, 50, 70, 80, 90, 10, 25, 30, 45, 50, 60, 75, 80, 90, 100, 150, 250, 300, 450, 500, 600, 750, 800, 900, 1500, 1500, 2500, 3000, 4500, 5000, 6000, 15000, 25000, 30000, 45000, 50000, 50000, 60000, 75000, 8000, 9000, 15000, 15000, 15000, 25000, 30000, 45000, 50 ...0, 250000, 30000, 45000, 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 20 The amplification may include using a first PCR primer pool 324 that includes custom primers 326A-C that target specific genes and a universal primer 328. The custom primer 326 may hybridize to a region within the cDNA portion 306' of the labeled cDNA molecule 304. The universal primer 328 may hybridize to the universal PCR region 316 of the labeled cDNA molecule 304.
[0165] As shown in step 3 of Figure 3, the product from the PCR amplification in step 2 can be amplified using a nested PCR primer pool and a second universal PCR primer. Nested PCR can minimize PCR amplification bias. In particular, nested PCR-labeled amplicons 322 can be further amplified by nested PCR. Nested PCR can include multiplex PCR including a nested PCR primer pool 330 of nested PCR primers 332a-c and a second universal PCR primer 328' in a single reaction volume. The nested PCR primer pool 328 may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 different nested PCR primers 330, or may include approximately such a number of different nested PCR primers 330, or may include at least or at most such a number of different nested PCR primers 330, or a number or range between any of these values. The nested PCR primers 332 may contain an adaptor 334 and hybridize to a region within the cDNA portion 306" of the labeled amplicon 322. The universal primer 328' may contain an adaptor 336 and hybridize to the universal PCR region 316 of the labeled amplicon 322. Thus, step 3 generates adapter-labeled amplicon 338. In some embodiments, nested PCR primer 332 and second universal PCR primer 328' may not contain adapters 334 and 336. Instead, adapters 334 and 336 may be ligated to the product of the nested PCR to generate adapter-labeled amplicon 338.
[0166] As shown in step 4, the PCR products from step 3 can be PCR amplified for sequencing using library amplification primers. In particular, adapters 334 and 336 can be used to perform one or more additional assays on adapter-labeled amplicons 338. Adapters 334 and 336 can be hybridized to primers 340 and 342. One or more of primers 340 and 342 can be PCR amplification primers. One or more of primers 340 and 342 can be sequencing primers. One or more of adapters 334 and 336 can be used for further amplification of adapter-labeled amplicons 338. One or more of adapters 334 and 336 can be used for sequencing of adapter-labeled amplicons 338. Primer 342 can contain plate index 344 so that amplicons or probability barcodes 310 generated using the same set of barcodes can be sequenced in a single sequencing reaction using next-generation sequencing (NGS).
[0167] Barcoding at the 5' end of nucleic acid targets Disclosed herein are systems, methods, compositions, and kits for attaching barcodes (e.g., stochastic barcodes) bearing molecular labels (or molecular beacons) to the 5' ends of nucleic acid targets (e.g., deoxyribonucleic acid molecules and ribonucleic acid molecules) to be barcoded or labeled. The 5'-based transcript counting methods disclosed herein can supplement or complement, for example, 3'-based transcript counting methods (e.g., Rhapsody™ assay (Becton, Dickinson and Company, Franklin Lakes, NJ)), Chromium™ Single Cell 3' Solution (10X Genomics, San Francisco, CA)). Barcoded nucleic acid targets can be used for sequence identification, transcript counting, alternative splicing analysis, mutation screening, and / or full-length sequencing in a high-throughput manner. Transcript counting at the 5' end (5' to the labeled target nucleic acid target) can reveal alternative splice isoforms and variants (including but not limited to splice mutations, single nucleotide polymorphisms (SNPs), insertions, deletions, and substitutions) at or near the 5' end of the nucleic acid molecule. In some embodiments, the method can involve intramolecular hybridization.
[0168] 4A-4B show a schematic diagram of a non-limiting, exemplary method 400 for gene-specific labeling of nucleic acid targets at the 5' end. A barcode 420 (e.g., a stochastic barcode) having a target-binding region (e.g., a poly(dT) tail 422) can bind to a polyadenylated RNA transcript 424 or other nucleic acid target via a poly(dA) tail 426 for labeling or barcoding (e.g., unique labeling). The barcode 420 can include a molecular label (ML) 428 and a sample label (SL) 430 to label the transcript 424 and track the sample origin of the RNA transcript 424, along with one or more additional sequences (e.g., consensus sequences such as adapter sequences 432) flanking the molecular label 428 / sample label 430 region of each barcode 420 for subsequent reactions. The repertoire of molecular label sequences in the barcodes for each sample can be large enough for stochastic labeling of RNA transcripts.
[0169] After cDNA synthesis in block 402 to generate barcoded cDNA molecules 434 comprising RNA transcripts 424 (or portions thereof), gene-specific methods can be used for 5' molecular barcoding. After gene-specific amplification in block 404, which may be optional, terminal transferase and deoxyadenosine triphosphate (dATP) can be added in block 406 to promote 3' poly(dA) tailing, generating amplicons 436 with poly(A) tails 438. A short denaturation step in block 408 allows for separation of the forward strand 436m and reverse strand 436c of amplicon 436 (e.g., barcoded cDNA molecules with poly(dA) tails). In block 410, the reverse strand 436c of amplicon 436 can hybridize intramolecularly via its poly(dA) tail 438 at the 3' end of the strand and the end of poly(dT) region 422 to form a hairpin or stem-loop 440. Next, in block 412, a polymerase (e.g., Klenow fragment) may be used to extend from the poly(dA) tail 438 and replicate the barcode to form an extended barcoded reverse strand 442. Gene-specific amplification in block 414 may (e.g., optionally) be performed to amplify the gene of interest and generate an amplicon 444 having a barcode at the 5' end (relative to the RNA transcript 424) for sequencing in block 416. In some embodiments, method 400 includes one or both of gene-specific amplification of the barcoded cDNA molecule 434 in block 404 and gene-specific amplification of the extended barcoded reverse strand 442 in block 414.
[0170] 5A-5B show a schematic diagram of a non-limiting, exemplary method 500 for labeling nucleic acid targets at the 5' end for whole-transcriptome analysis. A barcode 420 (e.g., a stochastic barcode) having a target binding region (e.g., a poly(dT) tail 422) can bind to a polyadenylated RNA transcript 424 or other nucleic acid target via a poly(dA) tail 426 for labeling or barcoding (e.g., unique barcoding). For example, the barcode 420 having the target binding region can bind to a nucleic acid target for labeling or barcoding. The barcode 420 can include a molecular label (ML) 428 and a sample label (SL) 430. The molecular labels 428 and sample labels 430, along with one or more additional sequences (e.g., consensus sequences such as adapter sequences 432) flanking the molecular label 428 / sample label 430 regions of each barcode 420, can be used to label transcripts 424 or nucleic acid targets (e.g., antibody oligonucleotides, whether associated with or dissociated from antibodies) for subsequent reactions and to track the sample origin of transcripts 424. The repertoire of molecular label 428 sequences in barcodes per sample can be large enough for stochastic labeling of RNA transcripts 424 or nucleic acid targets.
[0171] After cDNA synthesis to generate barcoded cDNA molecules 434 in block 402, a terminal transferase enzyme can be used in block 406 to A-tail the 3' ends of the barcoded cDNA molecules 434 (corresponding to the 5' ends of the labeled RNA transcripts) to generate cDNA molecules 436c, each having a 3' poly(dA) tail 438. Intramolecular hybridization of the cDNA molecule 436c with a 3' poly(dA) tail 438 can be initiated (e.g., by heating and cooling cycles or by diluting the barcoded cDNA molecule 436c with a poly(dA) tail 438), causing the new 3' poly(dA) tail 438 to anneal with the poly(dT) tail 422 of the same labeled cDNA molecule (although for target binding region sequences other than poly(dA), it is contemplated that the corresponding complement of the relevant target binding sequence may anneal to the target binding region), generating a hairpin or stem-loop structure 440 of the barcoded cDNA molecule in block 410. A polymerase (e.g., Klenow enzyme) with dNTPs is added to promote 3' extension beyond the new 3' poly(dA) tail 438, resulting in the barcode (e.g., molecular tag 428) at the 5' end of the labeled cDNA molecule with stem-loop 440 in block 412. )Whole transcriptome amplification (WTA) can be performed in block 414 using primers containing mirrored adapters 432, 432rc or sequences (or subsequences) of adapters 432, 432rc. Methods such as tagging or random priming can be used to generate smaller fragments of amplicon 444 with sequencing adapters (e.g., P5 446 and P7 448 sequences) for sequencing in block 418 (e.g., using an Illumina (San Diego, CA, US) sequencer). In some embodiments, sequencing adapters or sequencers for other sequencing methods (e.g., sequencers from Pacific Biosciences of California, Inc. (Menlo Park, CA, US) or Oxford Nanopore Technologies Limited (Oxford, UK)) can be directly ligated to generate amplicons for sequencing.
[0172] Disclosed herein are methods for determining the number of nucleic acid targets in a sample. In some embodiments, the method includes contacting copies of a nucleic acid target 424 in a sample with a plurality of oligonucleotide barcodes 420, each of the plurality of oligonucleotide barcodes 420 comprising a molecular beacon sequence 428 and a target binding region (e.g., a poly(dT) sequence 422) capable of hybridizing to the nucleic acid target 424, wherein at least ten of the plurality of oligonucleotide barcodes 420 comprise different molecular beacon sequences 428; extending the copies of the nucleic acid target 424 hybridized to the oligonucleotide barcodes 420 at block 402 to generate a plurality of nucleic acid molecules 434, each of which comprises a sequence 450c complementary to at least a portion of the nucleic acid target 424; amplifying the plurality of barcoded nucleic acid molecules 434 at block 404 to generate a plurality of amplified barcoded nucleic acid molecules 436; and coupling oligonucleotides comprising a complement 438 of the target binding region 422 to the plurality of amplified barcoded nucleic acid molecules 436 at block 406. to generate a plurality of barcoded nucleic acid molecules 436c, each of which comprises a target binding region 422 and a complement of the target binding region 438; in block 410, hybridizing the target binding region 422 and the complement of the target binding region 422 in each of the plurality of barcoded nucleic acid molecules 436c to form a stem-loop 440; in block 412, extending the 3' ends of the plurality of barcoded nucleic acid molecules, each of which has a stem-loop 440, to extend the stem-loop 440. In block 414, amplifying the plurality of elongated barcoded nucleic acid molecules 442 to generate a plurality of single-labeled nucleic acid molecules 444c, each of which comprises a molecular label 428 and a molecular label complement 428rc; and determining the number of nucleic acid targets in the sample based on the number of molecular label complements 428rc with unique sequences associated with the plurality of single-labeled nucleic acid molecules.
[0173] In some embodiments, after extending the 3' ends of the plurality of barcoded nucleic acid molecules with stem-loops 440, molecular label 428 is hybridized to complement 428rc of molecular label. The method may include denaturing the plurality of extended barcoded nucleic acid molecules 442 before amplifying the plurality of extended barcoded nucleic acid molecules 442 to generate a plurality of single-labeled nucleic acid molecules 444c (which may be part of amplicons 444c). Contacting copies of nucleic acid targets 424 in the sample may include contacting the copies of the plurality of nucleic acid targets 424 with a plurality of oligonucleotide barcodes 420. Extending copies of nucleic acid targets 424 may include extending copies of the plurality of nucleic acid targets 424 hybridized to oligonucleotide barcodes 420 to generate a plurality of barcoded nucleic acid molecules 436c, each comprising a sequence 450c complementary to at least a portion of one of the plurality of nucleic acid targets 424. Determining the number of nucleic acid targets 424 may include determining the number of each of the plurality of nucleic acid targets 424 in the sample based on the number of complements 428rc of molecular labels having unique sequences associated with single-labeled nucleic acid molecules of the plurality of single-labeled nucleic acid molecules 444c that include respective sequences 452c of the plurality of nucleic acid targets 424. Each sequence 452c of the plurality of nucleic acid targets may include a subsequence (including a complement or reverse complement) of each of the plurality of nucleic acid targets 424.
[0174] Disclosed herein is a method for determining the number of targets in a sample. In some embodiments, the method includes the steps of: barcoding 402 copies of a nucleic acid target 424 in the sample using a plurality of oligonucleotide barcodes 420 to generate a plurality of barcoded nucleic acid molecules 434, each of which comprises a sequence 450c (e.g., a complementary sequence, a reverse complementary sequence, or a combination thereof) of the nucleic acid target 424, a molecular label 428, and a target binding region (e.g., a poly(dT) region 422), wherein at least ten of the plurality of oligonucleotide barcodes 420 comprise different molecular label sequences 428; and attaching 406 an oligonucleotide comprising a complement 438 of the target binding region 422 to the plurality of barcoded nucleic acid molecules 434, to generate a plurality of barcoded nucleic acid molecules 434, each of which comprises a sequence 450c (e.g., a complementary sequence, a reverse complementary sequence, or a combination thereof) of the nucleic acid target 424, a molecular label 428, and a target binding region (e.g., a poly(dT) region 422). generating a plurality of barcoded nucleic acid molecules 436c, each comprising a complement 438; hybridizing 410 the target binding region 422 and the complement 438 of the target binding region in each of the plurality of barcoded nucleic acid molecules 436c to form a stem-loop 440; extending 412 the 3' ends of the plurality of barcoded nucleic acid molecules to extend the stem-loop 440 to generate a plurality of extended barcoded nucleic acid molecules 442, each comprising a molecular label 428 and a complement 428rc of the molecular label; and determining the number of nucleic acid targets 424 in the sample based on the number of complements 428rc of the molecular label having a unique sequence associated with the plurality of extended barcoded nucleic acid molecules 442.
[0175] Disclosed herein is a method for binding oligonucleotide barcodes to targets in a sample. In some embodiments, the method includes the steps of barcoding 402 copies of a nucleic acid target 424 in the sample using a plurality of oligonucleotide barcodes 420 to generate a plurality of barcoded nucleic acid molecules 434, each comprising a sequence 450c of the nucleic acid target 424, a molecular label 428, and a target binding region 422, wherein at least ten of the plurality of oligonucleotide barcodes 420 comprise different molecular label sequences 428; and binding an oligonucleotide comprising a complement 438 of the target binding region 422 to the plurality of barcoded nucleic acid molecules 434 to generate target binding regions 422. The method includes generating a plurality of barcoded nucleic acid molecules 436c, each comprising a region 422 and a complement 438 of the target binding region 422; hybridizing 410 the target binding region 422 and the complement 438 of the target binding region 422 in each of the plurality of barcoded nucleic acid molecules 436c to form a stem-loop 440; and extending 412 the 3' ends of the plurality of barcoded nucleic acid molecules to extend the stem-loop 440 to generate a plurality of extended barcoded nucleic acid molecules 442, each comprising a molecular label 428 and a complement 428rc of the molecular label 428. In some embodiments, the method includes determining the number of nucleic acid targets 424 in the sample based on the number of molecular labels 428 having unique sequences, their complements 428rc, or a combination thereof, associated with the plurality of extended barcoded nucleic acid molecules 442. For example, the number of nucleic acid targets 424 can be determined based on one or both of the molecular beacons 428 having unique sequences, their complements 428rc.
[0176] In some embodiments, the method includes barcoding 402 a plurality of copies of a target 424, which comprises contacting the copies of the nucleic acid target 424 with a plurality of oligonucleotide barcodes 420, each of the plurality of oligonucleotide barcodes 420 comprising a target binding region 422 capable of hybridizing to the nucleic acid target 424; and extending 402 the copies of the nucleic acid target 424 hybridized to the oligonucleotide barcodes 420 to generate a plurality of barcoded nucleic acid molecules 434.
[0177] In some embodiments, the method includes amplifying 404 a plurality of barcoded nucleic acid molecules 434 to generate a plurality of amplified barcoded nucleic acid molecules 436c, and binding an oligonucleotide comprising a complement 438 of the target binding region 422 includes binding an oligonucleotide comprising a complement 438 of the target binding region to the plurality of amplified barcoded nucleic acid molecules to generate a plurality of barcoded nucleic acid molecules 436r, each comprising a target binding region 422 and a complement 438 of the target binding region.
[0178] Gene-specific analysis. In some embodiments, the method (e.g., method 400) includes amplifying 414 a plurality of extended barcoded nucleic acid molecules 442 to generate a plurality of single-labeled nucleic acid molecules 444c, each comprising a complement 428rc of a molecular label 428. The single-labeled nucleic acid molecules 444c may be generated when the amplicons 444 containing them are denatured. Determining the number of nucleic acid targets 424 in the sample may include determining the number of nucleic acid targets 424 in the sample based on the number of complements 428rc of molecular labels 428 having unique sequences that are associated with the plurality of single-labeled nucleic acid molecules 444c.
[0179] Whole Transcriptome Analysis. In some embodiments, the method (e.g., method 500) includes amplifying 414 the plurality of elongated barcoded nucleic acid molecules 442 to generate a plurality of elongated barcoded nucleic acid molecule copies 444c. Determining the number of nucleic acid targets 424 in the sample includes determining the number of nucleic acid targets 424 in the sample based on the number of complements 428rc of molecular labels 428 having unique sequences that are associated with the plurality of elongated barcoded nucleic acid molecule copies 444c. The plurality of elongated barcoded nucleic acid molecule copies 444c can be formed when the amplicons 444 containing them are denatured.
[0180] In some embodiments, the sequence of the nucleic acid target in the plurality of barcoded nucleic acid molecules comprises a subsequence 452c of the nucleic acid target. The target binding region may comprise a gene-specific sequence. Binding 406 an oligonucleotide comprising a complement 438 of the target binding region 422 may comprise ligating an oligonucleotide comprising the complement 438 of the target binding region 422 to the plurality of barcoded nucleic acid molecules 434.
[0181] In some embodiments, the target binding region may comprise a poly(dT) sequence 422 (sometimes referred to herein as an oligo(dT) sequence). Binding an oligonucleotide comprising a complement 438 of the target binding region 422 comprises adding a plurality of adenosine monophosphates to the plurality of barcoded nucleic acid molecules 434 using terminal deoxynucleotidyl transferase. In some embodiments, the target binding region does not comprise a poly(dT) sequence.
[0182] In some embodiments, extending a copy of the nucleic acid target 424 hybridized to the oligonucleotide barcode 420 may include reverse transcribing the copy of the nucleic acid target 424 hybridized to the oligonucleotide barcode 420 to generate a plurality of barcoded complementary deoxyribonucleic acid (cDNA) molecules 434. Extending a copy of the nucleic acid target 424 hybridized to the oligonucleotide barcode 420 may include extending 402 the copy of the nucleic acid target 424 hybridized to the oligonucleotide barcode 420 using a DNA polymerase lacking at least one of 5'-3' exonuclease activity and 3'-5' exonuclease activity. The DNA polymerase may comprise Klenow fragment.
[0183] In some embodiments, the method includes obtaining sequence information of the plurality of extended barcoded nucleic acid molecules 442. Obtaining the sequence information may include attaching sequencing adaptors (e.g., P5 446 and P7 448 adaptors) to the plurality of extended barcoded nucleic acid molecules 442.
[0184] In some embodiments, the complement of the target binding region 438 may comprise the reverse complementary sequence of the target binding region. The complement of the target binding region 438 may comprise the complementary sequence of the target binding region. The complement of the molecular label 428rc may comprise the reverse complementary sequence of the molecular label. The complement of the molecular label may comprise the complementary sequence of the molecular label.
[0185] In some embodiments, the plurality of barcoded nucleic acid molecules 434 may comprise barcoded deoxyribonucleic acid (DNA) molecules. The barcoded nucleic acid molecules 434 may comprise barcoded ribonucleic acid (RNA) molecules. The nucleic acid target 424 may comprise a nucleic acid molecule. The nucleic acid molecule may comprise ribonucleic acid (RNA), messenger RNA (mRNA), microRNA, small interfering RNA (siRNA), RNA degradation products, RNA containing a poly(A) tail, or any combination thereof.
[0186] Antibody oligonucleotides. In some embodiments, the nucleic acid target may include a cellular component binding reagent. Cell-binding reagents associated with nucleic acid targets (e.g., antibody oligonucleotides, such as sample-indexing oligonucleotides) are described in U.S. Patent Application Publication No. 2018 / 0088112; and U.S. Patent Application No. 15 / 937,713, filed March 27, 2018 (the contents of each of these applications are incorporated herein by reference in their entirety). In some embodiments, multi-omics information, such as single-cell genomics, chromatin accessibility, methylomics, transcriptomics, and proteomics, is obtained using the 5' barcoding method of the present disclosure. The nucleic acid molecule may be associated with a cellular component binding reagent. The method may include dissociating the nucleic acid molecule and the cellular component binding reagent.
[0187] In some embodiments, each molecular label 428 of the plurality of oligonucleotide barcodes 420 comprises at least six nucleotides. The oligonucleotide barcodes 420 may comprise the same sample label 430. Each sample label 430 of the plurality of oligonucleotide barcodes 420 may comprise at least six nucleotides. The oligonucleotide barcodes 420 may comprise the same cell label. Each cell label of the plurality of oligonucleotide barcodes 420 may comprise at least six nucleotides.
[0188] In some embodiments, at least one of the plurality of barcoded nucleic acid molecules 436c becomes associated with the solid support when the target binding region and the complement of the target binding region in each of the plurality of barcoded nucleic acid molecules hybridize 410 to form a stem-loop. At least one of the plurality of barcoded nucleic acid molecules 436c can dissociate from the solid support when the target binding region 422 and the complement 438 of the target binding region 422 in each of the plurality of barcoded nucleic acid molecules 436c hybridize 410 to form a stem-loop 440. At least one of the plurality of barcoded nucleic acid molecules 436c can become associated with the solid support when the target binding region 422 and the complement 438 of the target binding region in each of the plurality of barcoded nucleic acid molecules 436c hybridize 410 to form a stem-loop 440.
[0189] In some embodiments, at least one of the plurality of barcoded nucleic acid molecules is associated with the solid support when the 3' ends of the plurality of barcoded nucleic acid molecules are extended 412 to extend the stem-loop 440 to generate a plurality of extended barcoded nucleic acid molecules 442, each comprising a molecular label 428 and a complement of the molecular label 428rc. At least one of the plurality of barcoded nucleic acid molecules may be dissociated from the solid support when the 3' ends of the plurality of barcoded nucleic acid molecules are extended 412 to extend the stem-loop 440 to generate a plurality of extended barcoded nucleic acid molecules 442, each comprising a molecular label 428 and a complement of the molecular label 428rc. At least one of the plurality of barcoded nucleic acid molecules 436c can be associated with a solid support when the 3' end of the plurality of barcoded nucleic acid molecules is extended 412 to extend stem-loop 440 to generate a plurality of extended barcoded nucleic acid molecules 442, each comprising a molecular label 428 and a complement of the molecular label 428rc. The solid support can include a synthetic particle 454. The solid support can include a flat surface or a substantially flat surface (e.g., a slide, such as a microscope slide, or a coverslip).
[0190] In some embodiments, at least one of the plurality of barcoded nucleic acid molecules 436c is in solution when the target binding region 422 and the complement 438 of the target binding region 422 in each of the plurality of barcoded nucleic acid molecules 436c hybridize 410 to form the stem-loop 440. For example, such intramolecular hybridization can occur if the concentration of the plurality of barcoded nucleic acid molecules 436c in solution is sufficiently low. At least one of the plurality of barcoded nucleic acid molecules can be in solution when the 3' ends of the plurality of barcoded nucleic acid molecules are extended 412 to extend the stem-loop 440 and generate a plurality of extended barcoded nucleic acid molecules 442, each comprising a molecular label 428 and a complement 428rc of the molecular label.
[0191] In some embodiments, the sample includes a single cell, and the method includes associating a synthetic particle 454 including a plurality of oligonucleotide barcodes 420 with the single cell in the sample. The method may include lysing the single cell after associating the synthetic particle 454 with the single cell. Lysing the single cell may include heating the sample, contacting the sample with a detergent, changing the pH of the sample, or any combination thereof. The synthetic particle and the single cell may be in the same well. The synthetic particle and the single cell may be in the same droplet.
[0192] In some embodiments, at least one of the plurality of oligonucleotide barcodes 420 may be immobilized on a synthetic particle 454. At least one of the plurality of oligonucleotide barcodes 420 may be partially immobilized on a synthetic particle 454. At least one of the plurality of oligonucleotide barcodes 420 may be encapsulated in a synthetic particle 454. At least one of the plurality of oligonucleotide barcodes 420 may be partially encapsulated in a synthetic particle 454. The synthetic particle 454 may be disintegrable. The synthetic particle 454 may comprise beads. The beads may include sepharose beads, streptavidin beads, agarose beads, magnetic beads, conjugated beads, protein A conjugated beads, protein G conjugated beads, protein A / G conjugated beads, protein L conjugated beads, oligo(dT) conjugated beads, silica beads, silica-like beads, anti-biotin microbeads, anti-fluorescent dye microbeads, or any combination thereof. The synthetic particles 454 may comprise a material selected from the group consisting of polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogel, paramagnetic material, ceramic, plastic, glass, methylstyrene, acrylic polymer, titanium, latex, Sepharose, cellulose, nylon, silicone, and any combination thereof. The synthetic particles 454 may comprise collapsible hydrogel particles. Each of the plurality of oligonucleotide barcodes 420 may comprise a linker functional group. The synthetic particles 454 may comprise a solid support functional group. The support functional group and the linker functional group may be associated with each other. The linker functional group and the support functional group may be independently selected from the group consisting of C6, biotin, streptavidin, primary amine, aldehyde, ketone, and any combination thereof.
[0193] Kit for barcoding the 5' end of nucleic acid targets Disclosed herein is a kit for binding oligonucleotide barcodes 420 to targets 424 in a sample, determining the number of targets 424 in a sample, and / or determining the number of nucleic acid targets 424 in a sample. In some embodiments, the kit includes a plurality of oligonucleotide barcodes 420, each of the plurality of oligonucleotide barcodes 420 comprising a molecular label 428 and a target binding region (e.g., a poly(dT) sequence 422), and at least 10 of the plurality of oligonucleotide barcodes 420 comprising different molecular label sequences 428; a terminal deoxynucleotidyl transferase or ligase; and a DNA polymerase lacking at least one of 5'-3' exonuclease activity and 3'-5' exonuclease activity. The kit may include a plurality of oligonucleotides comprising complements of the target binding regions. The plurality of oligonucleotides comprising complements of the target binding regions may be separated from the plurality of oligonucleotide barcodes. In some embodiments, multiple oligonucleotides comprising a complement of the target binding region are configured for binding to the 3' end of a DNA molecule, such as a cDNA molecule. Multiple oligonucleotides comprising a complement of the target binding region can be used to hybridize to the target binding region so that the DNA molecule forms a hairpin as described herein. The DNA polymerase can include Klenow fragment. The kit can include a buffer. The kit can include a cartridge. The kit can include one or more reagents for a reverse transcription reaction. The kit can include one or more reagents for an amplification reaction.
[0194] In some embodiments, the target binding region comprises a gene-specific sequence, an oligo(dT) sequence, a random multimer, or any combination thereof. The oligonucleotide barcodes may comprise the same sample label and / or the same cell label. Each sample label and / or cell label of the plurality of oligonucleotide barcodes may comprise at least 6 nucleotides. Each molecular label of the plurality of oligonucleotide barcodes may comprise at least 6 nucleotides.
[0195] In some embodiments, at least one of the plurality of oligonucleotide barcodes 420 is immobilized on a synthetic particle 454. At least one of the plurality of oligonucleotide barcodes 420 may be partially immobilized on a synthetic particle 454. At least one of the plurality of oligonucleotide barcodes 420 may be encapsulated in a synthetic particle 454. At least one of the plurality of oligonucleotide barcodes 420 may be partially encapsulated in a synthetic particle 454. The synthetic particle 454 may be disintegrable. The synthetic particle 454 may comprise beads. The beads may include sepharose beads, streptavidin beads, agarose beads, magnetic beads, conjugated beads, protein A conjugated beads, protein G conjugated beads, protein A / G conjugated beads, protein L conjugated beads, oligo(dT) conjugated beads, silica beads, silica-like beads, anti-biotin microbeads, anti-fluorescent dye microbeads, or any combination thereof. The synthetic particles may comprise a material selected from the group consisting of polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogel, paramagnetic material, ceramic, plastic, glass, methylstyrene, acrylic polymer, titanium, latex, Sepharose, cellulose, nylon, silicone, and any combination thereof. The synthetic particles 454 may comprise collapsible hydrogel particles. Each of the plurality of oligonucleotide barcodes may comprise a linker functional group. The synthetic particles 454 may comprise a solid support functional group. The support functional group and the linker functional group may be associated with each other. The linker functional group and the support functional group may be independently selected from the group consisting of C6, biotin, streptavidin, primary amine, aldehyde, ketone, and any combination thereof.
[0196] While various aspects and embodiments are disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for purposes of illustration and not limitation, with the true scope and spirit being indicated by the following claims.
[0197] Those skilled in the art will appreciate that, for this and other processes and methods disclosed herein, the functions performed in these processes and methods may be performed in differing order. Furthermore, the outlined steps and operations are presented by way of example only, and some of the steps and operations may be optional, combined into fewer steps and operations, or expanded into additional steps and operations, without departing from the essence of the disclosed embodiments.
[0198] In connection with the use of substantially all plural and / or singular terms herein, those skilled in the art can convert from plural to singular and / or from singular to plural where appropriate in the context and / or application. Various singular / plural permutations may be expressly set forth herein for clarity.
[0199] In general, it will be understood by those skilled in the art that the terms used in this specification, and particularly in the appended claims (e.g., the body of the appended claims), are generally intended to be "open" terms (e.g., the term "comprising" should be interpreted as "including, but not limited to," the term "having" should be interpreted as "having at least," the term "including" should be interpreted as "including, but not limited to," etc.). Where a specific number of introductory claim recitations is intended, such intention will be explicitly recited in the claim, and it will be further understood by those skilled in the art that, in the absence of such recitation, no such intention exists. For example, as an aid to understanding, the following appended claims may include the use of the introductory phrases "at least one" and "one or more" to introduce claim recitations. However, the use of such phrases indicates that even when the same claim includes the introductory phrases "one or more" or "at least one" and an indefinite article such as "a" or "an" (e.g., "a" and / or "an" should be interpreted to mean "at least one" or "one or more"), the introduction of a claim recitation with the indefinite article "a" or "an" should not be interpreted as meaning to limit any particular claim containing such an introductory claim recitation to embodiments containing only one such recitation; the same applies to the use of a definite article used to introduce a claim recitation. Moreover, even when a specific number of introductory claim recitations is explicitly recited, those skilled in the art will recognize that such recitation should be interpreted to mean at least the recited number (e.g., a minimum recitation of "two recitations," without any other modifiers, means at least two recitations or more than two recitations).Furthermore, when terms similar to "at least one of A, B, and C, etc." are used, such a configuration is generally intended to have the meaning that one of ordinary skill in the art would understand the term (e.g., "a system having at least one of A, B, and C" would include, but is not limited to, a system having A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). When terms similar to "at least one of A, B, or C, etc." are used, such a configuration is generally intended to have the meaning that one of ordinary skill in the art would understand the term (e.g., "a system having at least one of A, B, or C" would include, but is not limited to, a system having A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). It will be further understood by those skilled in the art that virtually all disjunctive words and / or phrases expressing two or more alternative terms, whether in the specification, claims, or drawings, should be understood to contemplate the possibility of including one of the terms, either of the terms, or both terms. For example, the phrase "A or B" will be understood to include the possibilities of "A" or "B" or "A and B."
[0200] Furthermore, when features or aspects of the present disclosure are described in terms of a Markush group, one skilled in the art will recognize that the present disclosure is also described thereby in terms of any individual members or subgroups of the Markush group.
[0201] As will be understood by those of skill in the art, for all purposes, including with respect to the provision of a specification, all ranges disclosed herein encompass all possible subranges and combinations of subranges. It will be readily recognized that any recited range fully describes and allows for the same range to be divided into at least two, three, four, five, ten, etc. As a non-limiting example, each range described herein can be readily divided into a lower third, middle third, and upper third, etc. As will be further understood by those of skill in the art, all terms such as "up to," "at least," etc., are inclusive of the recited number and refer to ranges that can be subsequently divided into the subranges described above. Finally, as will be understood by those of skill in the art, a range includes each individual member. Thus, for example, a group having 1 to 3 cells refers to a group having 1, 2, or 3 cells. Similarly, a group having 1 to 5 cells refers to a group having 1, 2, 3, 4, or 5 cells, etc.
[0202] From the foregoing, it will be understood that various embodiments of the present disclosure have been described herein for purposes of illustration and that various modifications can be made without departing from the scope and spirit of the present disclosure. Accordingly, the various embodiments disclosed herein are not intended to be limiting, with the true scope and spirit being indicated by the following claims. Another aspect of the present invention may be as follows. [1] A method for attaching an oligonucleotide barcode to a nucleic acid target in a sample, comprising: barcoding copies of the nucleic acid target in the sample with a plurality of oligonucleotide barcodes to generate a plurality of barcoded nucleic acid molecules, each comprising the sequence of the nucleic acid target, a molecular label, and a target binding region, wherein at least 10 of the plurality of oligonucleotide barcodes comprise different molecular label sequences; binding an oligonucleotide comprising the complement of the target binding region to the plurality of barcoded nucleic acid molecules to generate a plurality of barcoded nucleic acid molecules each comprising the target binding region and the complement of the target binding region; hybridizing the target binding region and the complement of the target binding region in each of the plurality of barcoded nucleic acid molecules to form a stem-loop; extending the 3' ends of the plurality of barcoded nucleic acid molecules to extend the stem-loop and generate a plurality of extended barcoded nucleic acid molecules each comprising the molecular beacon and a complement of the molecular beacon; A method comprising: [2] The method of [1], further comprising determining the number of nucleic acid targets in the sample based on the number of molecular labels having unique sequences, their complements, or combinations thereof, associated with the plurality of extended barcoded nucleic acid molecules. [3] A method for determining the number of targets in a sample, comprising: barcoding copies of a nucleic acid target in a sample with a plurality of oligonucleotide barcodes to generate a plurality of barcoded nucleic acid molecules, each comprising the sequence of the nucleic acid target, a molecular label, and a target binding region, wherein at least 10 of the plurality of oligonucleotide barcodes comprise different molecular label sequences; binding an oligonucleotide comprising the complement of the target binding region to the plurality of barcoded nucleic acid molecules to generate a plurality of barcoded nucleic acid molecules each comprising the target binding region and the complement of the target binding region; hybridizing the target binding region and the complement of the target binding region in each of the plurality of barcoded nucleic acid molecules to form a stem-loop; extending the 3' ends of the plurality of barcoded nucleic acid molecules to extend the stem-loop to generate a plurality of extended barcoded nucleic acid molecules each comprising the molecular label and the complement of the molecular label; determining the number of nucleic acid targets in the sample based on the number of complements of molecular labels having unique sequences associated with the plurality of elongated barcoded nucleic acid molecules; A method comprising: [4] The step of barcoding the copies of the plurality of targets comprises: contacting copies of the nucleic acid target with the plurality of oligonucleotide barcodes, each of the plurality of oligonucleotide barcodes comprising the target binding region capable of hybridizing to the nucleic acid target; extending the copies of the nucleic acid targets hybridized to the oligonucleotide barcodes to generate the plurality of barcoded nucleic acid molecules; The method according to any one of [1] to [3] above, comprising: [5] The step of barcoding the copies of the plurality of targets comprises: contacting copies of the nucleic acid target with the plurality of oligonucleotide barcodes, each of the plurality of oligonucleotide barcodes comprising the target binding region, the target binding region hybridizing to the nucleic acid target; extending the copies of the nucleic acid targets hybridized to the oligonucleotide barcodes to generate the plurality of barcoded nucleic acid molecules; The method according to any one of [1] to [3] above, comprising: [6] Amplifying the plurality of barcoded nucleic acid molecules to generate a plurality of amplified barcoded nucleic acid molecules. The method according to any one of [1] to [5] above, wherein the step of binding the oligonucleotide comprising the complement of the target binding region comprises the step of binding the oligonucleotide comprising the complement of the target binding region to the plurality of amplified barcoded nucleic acid molecules to generate the plurality of barcoded nucleic acid molecules each comprising the target binding region and the complement of the target binding region. [7] amplifying the plurality of elongated barcoded nucleic acid molecules to generate a plurality of single-labeled nucleic acid molecules, each of which comprises the complement of the molecular label. [6] The method according to any one of [2] to [6] above, wherein determining the number of nucleic acid targets in the sample comprises determining the number of nucleic acid targets in the sample based on the number of complements of molecular labels having unique sequences that are associated with the plurality of single-labeled nucleic acid molecules. [8] amplifying the plurality of extended bar-coded nucleic acid molecules to generate copies of the plurality of extended bar-coded nucleic acid molecules. The method of any one of [2] to [6], wherein determining the number of nucleic acid targets in the sample comprises determining the number of nucleic acid targets in the sample based on the number of complements of molecular labels having unique sequences that are associated with the copies of the plurality of elongated barcoded nucleic acid molecules. [9] A method for determining the number of nucleic acid targets in a sample, comprising: contacting copies of a nucleic acid target in a sample with a plurality of oligonucleotide barcodes, each of the plurality of oligonucleotide barcodes comprising a molecular label and a target binding region capable of hybridizing to the nucleic acid target, wherein at least 10 of the plurality of oligonucleotide barcodes comprise different molecular label sequences; extending the copies of the nucleic acid target hybridized to the oligonucleotide barcodes to generate a plurality of nucleic acid molecules, each of which comprises a sequence complementary to at least a portion of the nucleic acid target; amplifying the plurality of barcoded nucleic acid molecules to generate a plurality of amplified barcoded nucleic acid molecules; ligating an oligonucleotide comprising the complement of the target binding region to the plurality of amplified barcoded nucleic acid molecules to generate a plurality of barcoded nucleic acid molecules each comprising the target binding region and the complement of the target binding region; hybridizing the target binding region and the complement of the target binding region in each of the plurality of barcoded nucleic acid molecules to form a stem-loop; extending the 3' ends of the plurality of barcoded nucleic acid molecules to extend the stem-loop and generate a plurality of extended barcoded nucleic acid molecules each comprising the molecular label and a complement of the molecular label; amplifying the plurality of elongated barcoded nucleic acid molecules to generate a plurality of single-labeled nucleic acid molecules, each of which comprises the complement of the molecular label; determining the number of nucleic acid targets in the sample based on the number of complements of molecular labels having unique sequences associated with the plurality of single-labeled nucleic acid molecules; A method comprising:
[10] The method of any one of [7] to
[10] , wherein the molecular label is hybridized to the complement of the molecular label after extending the 3' ends of the plurality of barcoded nucleic acid molecules, and the method includes a step of denaturing the plurality of extended barcoded nucleic acid molecules before amplifying the plurality of extended barcoded nucleic acid molecules to generate the plurality of single-labeled nucleic acid molecules.
[11] The step of contacting copies of the nucleic acid target in the sample comprises contacting copies of the nucleic acid target with a plurality of oligonucleotide barcodes; extending the copies of the nucleic acid targets comprises extending the copies of the plurality of nucleic acid targets hybridized to the oligonucleotide barcodes to generate a plurality of barcoded nucleic acid molecules, each of the plurality of nucleic acid targets comprising a sequence complementary to at least a portion of one of the plurality of nucleic acid targets; 11. The method of claim 9 or 10, wherein determining the number of nucleic acid targets comprises determining the number of each of the plurality of nucleic acid targets in the sample based on the number of complements of molecular labels having unique sequences that are associated with single-labeled nucleic acid molecules of the plurality of single-labeled nucleic acid molecules that comprise the sequences of each of the plurality of nucleic acid targets.
[12] The method described in
[11] , wherein the sequence of each of the plurality of nucleic acid targets comprises a partial sequence of each of the plurality of nucleic acid targets.
[13] The method described in any one of [1] to
[12] , wherein the sequence of the nucleic acid target in the plurality of barcoded nucleic acid molecules includes a partial sequence of the nucleic acid target.
[14] The method according to any one of [1] to
[13] above, wherein the target binding region comprises a gene-specific sequence.
[15] The method according to any one of [1] to
[14] , wherein the step of binding the oligonucleotide containing the complement of the target binding region comprises the step of ligating the oligonucleotide containing the complement of the target binding region to the plurality of barcoded nucleic acid molecules.
[16] The target binding region comprises a poly(dT) sequence; The method according to any one of [1] to
[12] above, wherein the step of binding the oligonucleotide comprising the complement of the target binding region comprises the step of adding a plurality of adenosine monophosphates to the plurality of barcoded nucleic acid molecules using terminal deoxynucleotidyl transferase.
[17] The method described in any one of [4] to
[16] , wherein the step of extending the copy of the nucleic acid target hybridized to the oligonucleotide barcode includes a step of reverse transcribing the copy of the nucleic acid target hybridized to the oligonucleotide barcode to generate a plurality of barcoded complementary deoxyribonucleic acid (cDNA) molecules.
[18] The method according to any one of [4] to
[16] , wherein the step of extending the copy of the nucleic acid target hybridized to the oligonucleotide barcode comprises a step of extending the copy of the nucleic acid target hybridized to the oligonucleotide barcode using a DNA polymerase lacking at least one of 5'-3' exonuclease activity and 3'-5' exonuclease activity.
[19] The method described in
[18] , wherein the DNA polymerase comprises the Klenow fragment.
[20] The method described in any one of [1] to
[19] above, comprising a step of obtaining sequence information of the plurality of elongated barcoded nucleic acid molecules.
[21] The method described in
[20] , wherein the step of obtaining the sequence information comprises a step of attaching sequencing adapters to the plurality of extended barcoded nucleic acid molecules.
[22] The method according to any one of [1] to
[21] , wherein the complement of the target binding region comprises a reverse complementary sequence of the target binding region.
[23] The method according to any one of [1] to
[21] , wherein the complement of the target binding region comprises a complementary sequence of the target binding region.
[24] The method according to any one of [1] to
[23] , wherein the complement of the molecular label comprises a reverse complementary sequence of the molecular label.
[25] The method according to any one of [1] to
[23] , wherein the complement of the molecular label comprises a complementary sequence of the molecular label.
[26] The method according to any one of [1] to
[25] , wherein the plurality of barcoded nucleic acid molecules include barcoded deoxyribonucleic acid (DNA) molecules.
[27] The method according to any one of [1] to
[25] , wherein the barcoded nucleic acid molecule comprises a barcoded ribonucleic acid (RNA) molecule.
[28] The method according to any one of [1] to
[27] , wherein the nucleic acid target comprises, consists essentially of, or consists of a nucleic acid molecule.
[29] The method described in
[28] , wherein the nucleic acid molecule comprises ribonucleic acid (RNA), messenger RNA (mRNA), microRNA, small interfering RNA (siRNA), RNA degradation products, RNA containing a poly(A) tail, or any combination thereof.
[30] The method according to
[28] or
[29] , wherein the nucleic acid target comprises a cell component binding reagent.
[31] The method of
[30] , wherein the nucleic acid molecule is associated with the cellular component binding reagent.
[32] The method according to
[31] , further comprising a step of dissociating the nucleic acid molecule and the cell component binding reagent.
[33] The method according to any one of [1] to
[32] , wherein each molecular label of the plurality of oligonucleotide barcodes comprises at least 6 nucleotides.
[34] The method according to any one of [1] to
[33] , wherein the oligonucleotide barcodes contain identical sample labels.
[35] The method described in
[34] , wherein each sample label of the plurality of oligonucleotide barcodes comprises at least 6 nucleotides.
[36] The method according to any one of [1] to
[35] , wherein the oligonucleotide barcodes contain identical cell markers.
[37] The method described in
[36] , wherein each cell marker of the plurality of oligonucleotide barcodes comprises at least 6 nucleotides.
[38] The method of any one of [1] to
[37] , wherein at least one of the plurality of barcoded nucleic acid molecules becomes associated with a solid support when the target binding region and the complement of the target binding region in each of the plurality of barcoded nucleic acid molecules hybridize to form the stem-loop.
[39] The method according to any one of [1] to
[37] , wherein at least one of the plurality of barcoded nucleic acid molecules dissociates from the solid support when the target binding region in each of the plurality of barcoded nucleic acid molecules hybridizes with the complement of the target binding region to form the stem loop.
[40] The method of any one of [1] to
[37] , wherein at least one of the plurality of barcoded nucleic acid molecules becomes associated with a solid support when the target binding region and the complement of the target binding region in each of the plurality of barcoded nucleic acid molecules hybridize to form the stem-loop.
[41] The method of any one of [1] to
[37] , wherein at least one of the plurality of barcoded nucleic acid molecules is associated with a solid support when the 3' ends of the plurality of barcoded nucleic acid molecules are extended to extend the stem-loop and generate the plurality of extended barcoded nucleic acid molecules each comprising the molecular label and a complement of the molecular label.
[42] The method according to any one of [1] to
[37] , wherein at least one of the plurality of barcoded nucleic acid molecules dissociates from the solid support when the target binding region in each of the plurality of barcoded nucleic acid molecules hybridizes with the complement of the target binding region to form the stem loop.
[43] The method of any one of [1] to
[37] , wherein at least one of the plurality of barcoded nucleic acid molecules is not associated with a solid support when the 3' ends of the plurality of barcoded nucleic acid molecules are extended to extend the stem-loop and generate the plurality of extended barcoded nucleic acid molecules each comprising the molecular label and a complement of the molecular label.
[44] The method according to any one of
[38] to
[43] above, wherein the solid support comprises synthetic particles.
[45] The method according to any one of
[38] to
[43] above, wherein the solid support comprises a flat surface.
[46] The method of any one of [1] to
[37] , wherein at least one of the plurality of barcoded nucleic acid molecules is in solution when the target binding region and the complement of the target binding region in each of the plurality of barcoded nucleic acid molecules hybridize to form the stem-loop.
[47] The method of any one of [1] to
[37] , wherein at least one of the plurality of barcoded nucleic acid molecules is in solution when the 3' ends of the plurality of barcoded nucleic acid molecules are extended to extend the stem-loop and generate the plurality of extended barcoded nucleic acid molecules each comprising the molecular label and the complement of the molecular label.
[48] The method of
[46] or
[47] , wherein the solution is distributed into compartments containing one or less cells.
[49] The method described in
[48] , wherein the compartment comprises at least one of a microdrop, a microwell, or a chamber of a fluidic device.
[50] The method described in any one of [1] to
[49] , wherein the sample contains a single cell, and the method includes a step of associating synthetic particles containing the plurality of oligonucleotide barcodes with the single cell in the sample.
[51] The method according to
[50] , comprising the step of associating the synthetic particles with the single cell and then lysing the single cell.
[52] The method according to
[51] , wherein the step of lysing the single cell comprises heating the sample, contacting the sample with a detergent, changing the pH of the sample, or any combination thereof.
[53] The method according to any one of
[50] to
[52] , wherein the synthetic particles and the single cells are in the same well.
[54] The method according to any one of
[50] to
[53] , wherein the synthetic particle and the single cell are in the same droplet.
[55] The method according to any one of
[50] to
[54] , wherein at least one of the plurality of oligonucleotide barcodes is immobilized on the synthetic particle.
[56] The method described in any one of
[50] to
[54] , wherein at least one of the plurality of oligonucleotide barcodes is partially immobilized on the synthetic particle.
[57] The method described in any one of
[50] to
[54] , wherein at least one of the plurality of oligonucleotide barcodes is encapsulated in the synthetic particle.
[58] The method described in any one of
[50] to
[54] , wherein at least one of the plurality of oligonucleotide barcodes is partially encapsulated in the synthetic particle.
[59] The method according to any one of
[50] to
[58] , wherein the synthetic particles are disintegrable.
[60] The method according to any one of
[50] to
[59] , wherein the synthetic particles include beads.
[61] The method according to
[60] , wherein the beads include sepharose beads, streptavidin beads, agarose beads, magnetic beads, conjugate beads, protein A conjugate beads, protein G conjugate beads, protein A / G conjugate beads, protein L conjugate beads, oligo(dT) conjugate beads, silica beads, silica-like beads, anti-biotin microbeads, anti-fluorescent dye microbeads, or any combination thereof.
[62] The method of
[60] or
[61] , wherein the synthetic particles comprise a material selected from the group consisting of polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogel, paramagnetic material, ceramic, plastic, glass, methylstyrene, acrylic polymer, titanium, latex, Sepharose, cellulose, nylon, silicone, and any combination thereof.
[63] The method according to any one of
[50] to
[62] , wherein the synthetic particles include disintegrable hydrogel particles.
[64] Each of the plurality of oligonucleotide barcodes comprises a linker functional group; the synthetic particles comprise a solid support functional group; The method according to any one of
[50] to
[63] above, wherein the carrier functional group and the linker functional group are associated with each other.
[65] The method of
[64] , wherein the linker functional group and the carrier functional group are independently selected from the group consisting of C6, biotin, streptavidin, primary amines, aldehydes, ketones, and any combination thereof.
[66] A kit comprising: a plurality of oligonucleotide barcodes, each of the plurality of oligonucleotide barcodes comprising a molecular label and a target binding region, wherein at least 10 of the plurality of oligonucleotide barcodes comprise different molecular label sequences; with terminal deoxynucleotidyl transferase or ligase; A DNA polymerase lacking at least one of 5'-3' exonuclease activity and 3'-5' exonuclease activity. Kit including:
[67] The kit described in
[66] , further comprising a plurality of oligonucleotides comprising complements of the target binding region.
[68] The kit described in
[67] , wherein the plurality of oligonucleotides comprising the complement of the target binding region are configured for binding to the 3' end of a DNA molecule, such as a cDNA molecule.
[69] The kit described in
[66] , wherein the DNA polymerase comprises the Klenow fragment.
[70] The kit according to
[66] or
[67] , which comprises a buffer solution.
[71] The kit according to any one of
[66] to
[68] , which includes a cartridge.
[72] The kit according to any one of
[66] to
[71] above, which comprises one or more reagents for a reverse transcription reaction.
[73] The kit according to any one of
[66] to
[72] above, which comprises one or more reagents for an amplification reaction.
[74] The kit according to any one of
[66] to
[73] , wherein the target binding region comprises a gene-specific sequence, an oligo(dT) sequence, a random multimer, or any combination thereof.
[75] The kit described in any one of
[66] to
[74] , wherein the oligonucleotide barcodes include identical sample labels and / or identical cell labels.
[76] The kit described in
[75] , wherein each sample label and / or cell label of the plurality of oligonucleotide barcodes contains at least 6 nucleotides.
[77] The kit described in any one of
[66] to
[75] , wherein each molecular label of the plurality of oligonucleotide barcodes contains at least 6 nucleotides.
[78] A kit described in any one of
[66] to
[77] , wherein at least one of the plurality of oligonucleotide barcodes is immobilized on a synthetic particle.
[79] A kit described in any one of
[66] to
[78] , wherein at least one of the plurality of oligonucleotide barcodes is partially immobilized on the synthetic particle.
[80] A kit described in any one of
[66] to
[79] , wherein at least one of the plurality of oligonucleotide barcodes is encapsulated in the synthetic particle.
[81] A kit described in any one of
[66] to
[80] , wherein at least one of the plurality of oligonucleotide barcodes is partially encapsulated in the synthetic particle.
[82] The kit described in any one of
[66] to
[81] , wherein the synthetic particles are disintegrating.
[83] The kit according to any one of
[66] to
[82] , wherein the synthetic particles include beads.
[84] The kit described in
[83] , wherein the beads include sepharose beads, streptavidin beads, agarose beads, magnetic beads, conjugate beads, protein A conjugate beads, protein G conjugate beads, protein A / G conjugate beads, protein L conjugate beads, oligo(dT) conjugate beads, silica beads, silica-like beads, anti-biotin microbeads, anti-fluorescent dye microbeads, or any combination thereof.
[85] The kit described in any one of
[66] to
[84] , wherein the synthetic particles comprise a material selected from the group consisting of polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogel, paramagnetic material, ceramic, plastic, glass, methylstyrene, acrylic polymer, titanium, latex, Sepharose, cellulose, nylon, silicone, and any combination thereof.
[86] The kit according to any one of
[66] to
[85] , wherein the synthetic particles include disintegrating hydrogel particles.
[87] Each of the plurality of oligonucleotide barcodes comprises a linker functional group; the synthetic particles comprise a solid support functional group; The kit according to any one of
[66] to
[86] above, wherein the carrier functional group and the linker functional group are associated with each other.
[88] The kit described in
[87] , wherein the linker functional group and the carrier functional group are independently selected from the group consisting of C6, biotin, streptavidin, primary amines, aldehydes, ketones, and any combination thereof.
Claims
1. a plurality of oligonucleotide barcodes, each of the plurality of oligonucleotide barcodes comprising a molecular label and a target binding region, wherein at least 10 of the plurality of oligonucleotide barcodes comprise different molecular label sequences; terminal deoxynucleotidyl transferase or ligase; a DNA polymerase lacking at least one of 5'-3' exonuclease activity and 3'-5' exonuclease activity; A kit comprising: below: barcoding copies of the nucleic acid target in the sample with the plurality of oligonucleotide barcodes to generate a plurality of barcoded nucleic acid molecules, each comprising the sequence of the nucleic acid target, a molecular label, and a target binding region; attaching oligonucleotides comprising a complement of a target binding region to a plurality of bar-coded nucleic acid molecules to generate a plurality of bar-coded nucleic acid molecules, each of which comprises a target binding region and a complement of the target binding region, wherein attaching the oligonucleotides comprising the complement of the target binding region comprises: (a) and / or (b): (a) ligating the oligonucleotides comprising the complements of the target binding regions to the plurality of barcoded nucleic acid molecules using a ligase; and / or (b) adding a plurality of adenosine monophosphates to said plurality of barcoded nucleic acid molecules using terminal deoxynucleotidyl transferase; Includes; hybridizing the target binding region and the complement of the target binding region in each of the plurality of barcoded nucleic acid molecules to form a stem-loop; and extending the 3' ends of the plurality of barcoded nucleic acid molecules with a DNA polymerase to extend the stem-loop and generate a plurality of extended barcoded nucleic acid molecules, each of which comprises the molecular beacon and the complement of the molecular beacon. The kit for use in a method comprising:
2. 10. The kit of claim 1, further comprising a plurality of oligonucleotides comprising the complement of the target binding region.
3. 3. The kit of claim 2, wherein the plurality of oligonucleotides comprising the complement of the target binding region are configured for binding to the 3' end of a DNA molecule, such as a cDNA molecule.
4. The kit according to any one of claims 1 to 3, wherein the DNA polymerase comprises a Klenow fragment.
5. The kit according to any one of claims 1 to 4, comprising (i) a buffer solution, (ii) a cartridge, (iii) one or more reagents for a reverse transcription reaction, and / or (iv) one or more reagents for an amplification reaction.
6. The kit of any one of claims 1 to 5, wherein the target binding region comprises a gene-specific sequence, an oligo(dT) sequence, a random multimer, or any combination thereof.
7. The kit of any one of claims 1 to 6, wherein the oligonucleotide barcodes comprise an identical sample label and / or an identical cell label.
8. 8. The kit of claim 7, wherein each sample label and / or cell label of the plurality of oligonucleotide barcodes comprises at least 6 nucleotides.
9. The kit of any one of claims 1 to 8, wherein each molecular label of the plurality of oligonucleotide barcodes comprises at least 6 nucleotides.
10. 10. The kit of any one of claims 1 to 9, wherein at least one of the plurality of oligonucleotide barcodes is (i) immobilized on a synthetic particle, (ii) partially immobilized on a synthetic particle, (iii) encapsulated in a synthetic particle, and / or partially encapsulated in a synthetic particle.
11. The synthetic particles are disintegrable. the synthetic particles comprise a material selected from the group consisting of polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogels, paramagnetic materials, ceramics, plastics, glass, methylstyrene, acrylic polymers, titanium, latex, sepharose, cellulose, nylon, silicone, and any combination thereof; and / or The synthetic particles include collapsible hydrogel particles. The kit of claim 10.
12. The kit of claim 10 or 11, wherein the synthetic particles comprise beads.
13. 13. The kit of claim 12, wherein the beads comprise sepharose beads, streptavidin beads, agarose beads, magnetic beads, conjugated beads, protein A conjugated beads, protein G conjugated beads, protein A / G conjugated beads, protein L conjugated beads, oligo(dT) conjugated beads, silica beads, silica-like beads, anti-biotin microbeads, anti-fluorescent dye microbeads, or any combination thereof.
14. each of the plurality of oligonucleotide barcodes comprises a linker functional group; the synthetic particles comprise a solid support functional group; and The kit of any one of claims 10 to 13, wherein the carrier functional group and the linker functional group are associated with each other.
15. 15. The kit of claim 14, wherein the linker functional group and the carrier functional group are independently selected from the group consisting of C6, biotin, streptavidin, primary amines, aldehydes, ketones, and any combination thereof.
Citation Information
Patent Citations
Barcoded nucleic acids
JP2015533296A
Error suppression in sequenced DNA fragments using redundant reads with unique molecular indices (UMIS)
WO2016176091A1
Asymmetric templates and asymmetric method of nucleic acid sequencing
WO2018015365A1
Measurement of protein expression using reagents with barcoded oligonucleotide sequences
WO2018058073A2