Aptamer barcoding
Patent Information
- Application Number
- JP2024088926
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-08-17
- Filing Date
- 2024-05-31
- Publication Date
- 2025-06-09
- Estimated Expiration
- 2039-08-14
AI Technical Summary
Current methods struggle to quantitatively analyze protein expression in cells while simultaneously measuring gene expression, particularly in a massively parallel manner, and there is a need for systems and methods that can accurately determine protein and gene expression profiles in individual cells.
The use of aptamer compositions with sample indexing oligonucleotides to barcode individual cells, allowing for the identification of protein targets and subsequent sequencing to determine sample origin and expression levels, combined with barcoding techniques to differentiate between cells from different samples.
Enables the simultaneous and accurate measurement of protein and gene expression in individual cells, improving the ability to analyze and compare samples with reduced amplification bias and increased accuracy.
Abstract
Description
[Technical Field]
[0001] Related Applications This application claims the benefit under 35 U.S.C. §119(e) of U.S. Provisional Patent Application No. 62 / 719,406, filed August 17, 2018, the contents of which are incorporated herein by reference in their entirety for all purposes. Sequence Listing Reference This application is filed with an electronic Sequence Listing, which is provided as a file entitled "SequenceListing," 4 kilobytes in size, created on August 2, 2019. The information in the electronic format of the Sequence Listing is incorporated herein by reference in its entirety.
[0002] The present disclosure relates generally to the field of molecular biology, for example, to using molecular barcoding to distinguish cells from different samples and to determine protein expression profiles in cells. [Background technology]
[0003] Current technology allows for the measurement of gene expression in single cells in a massively parallel manner (e.g., >1000 cells) by attaching cell-specific oligonucleotide barcodes to poly(A) mRNA molecules derived from individual cells while co-localizing each cell in a compartment with a barcoded reagent bead. Gene expression can affect protein expression. Protein-protein interactions can affect gene expression and protein expression. There is a need for systems and methods that can quantitatively analyze protein expression in cells and simultaneously measure protein and gene expression in cells. Summary of the Invention
[0004] The disclosure herein includes embodiments of a method for identifying samples. In some embodiments, the method includes: contacting each of a plurality of samples with a sample indexing composition from a plurality of sample indexing compositions, each of the plurality of samples comprising one or more cells, each comprising one or more protein targets; the sample indexing composition comprises an aptamer composition comprising an aptamer and a sample indexing oligonucleotide, wherein the aptamer is capable of specifically binding to at least one of the one or more protein targets, and the sample indexing oligonucleotide comprises a sample indexing sequence, and the sample indexing sequences of at least two of the plurality of sample indexing compositions comprise different sequences; barcoding the sample indexing oligonucleotides using a plurality of barcodes to generate a plurality of barcoded sample indexing oligonucleotides; obtaining sequencing data of the plurality of barcoded sample indexing oligonucleotides; and identifying the sample origin of at least one cell from the one or more cells based on the sample indexing sequence of at least one barcoded sample indexing oligonucleotide from the plurality of barcoded sample indexing oligonucleotides. In some embodiments, the method includes contacting each of a plurality of samples with a respective sample indexing composition of a plurality of sample indexing compositions, each of the plurality of samples comprising one or more cells, each comprising one or more cellular component targets, the sample indexing composition comprising an aptamer composition comprising an aptamer and a sample indexing oligonucleotide, the aptamer being capable of specifically binding to at least one of the one or more cellular component targets, the sample indexing oligonucleotide comprising a sample indexing sequence, the sample indexing sequences of at least two sample indexing compositions of the plurality of sample indexing compositions comprising different sequences; and identifying the sample origin of at least one cell of the one or more cells based on the sample indexing sequence of at least one sample indexing oligonucleotide of the plurality of sample indexing compositions. In some embodiments, the cellular component target comprises a protein target. In some embodiments, identifying the sample origin of the at least one cell includes barcoding sample indexing oligonucleotides of a plurality of sample indexing compositions using a plurality of barcodes to generate a plurality of barcoded sample indexing oligonucleotides; obtaining sequencing data of the plurality of barcoded sample indexing oligonucleotides; and identifying the sample origin of the cell based on the sample indexing sequence of at least one barcoded sample indexing oligonucleotide of the plurality of barcoded sample indexing oligonucleotides in the sequencing data.
[0005] In some embodiments, identifying the sample origin of the at least one cell comprises identifying the presence or absence of a sample indexing sequence of at least one sample indexing oligonucleotide of the plurality of sample indexing compositions. Identifying the presence or absence of a sample indexing sequence comprises replicating at least one sample indexing oligonucleotide to generate a plurality of replicate sample indexing oligonucleotides; obtaining sequencing data of the plurality of replicate sample indexing oligonucleotides; and identifying the sample origin of the cell based on the sample indexing sequence of the replicate sample indexing oligonucleotide of the plurality of sample indexing oligonucleotides that corresponds to the at least one barcoded sample indexing oligonucleotide in the sequencing data.
[0006] In some embodiments, replicating at least one sample indexing oligonucleotide to generate a plurality of replicate sample indexing oligonucleotides comprises ligating a replication adapter to the at least one barcoded sample indexing oligonucleotide prior to replicating the at least one barcoded sample indexing oligonucleotide, and replicating the at least one barcoded sample indexing oligonucleotide comprises replicating the at least one barcoded sample indexing oligonucleotide using the replication adapter ligated to the at least one barcoded sample indexing oligonucleotide to generate a plurality of replicate sample indexing oligonucleotides. In some embodiments, replicating at least one sample indexing oligonucleotide to generate a plurality of replicate sample indexing oligonucleotides comprises contacting a capture probe with the at least one sample indexing oligonucleotide prior to replicating the at least one barcoded sample indexing oligonucleotide to generate a capture probe hybridized to the sample indexing oligonucleotide; and extending the capture probe hybridized to the sample indexing oligonucleotide to generate a sample indexing oligonucleotide associated with the capture probe. Replicating at least one sample-indexing oligonucleotide may include replicating a sample-indexing oligonucleotide associated with a capture probe to generate a plurality of replicate sample-indexing oligonucleotides.
[0007] In some embodiments, the sample indexing sequences are 6-60 nucleotides in length. The sample indexing oligonucleotides may be 50-500 nucleotides in length. The sample indexing sequences of at least 10, 100, or 1000 sample indexing compositions of the plurality of sample indexing compositions may comprise different sequences.
[0008] In some embodiments, the single polynucleotide comprises a sample indexing oligonucleotide and an aptamer. The aptamer may be located 5' of the sample indexing oligonucleotide in the single polynucleotide. The aptamer may be located 3' of the sample indexing oligonucleotide in the single polynucleotide. The sample indexing oligonucleotide may be associated with the aptamer. The sample indexing oligonucleotide may be attached to the aptamer. The sample indexing oligonucleotide may be covalently attached to the aptamer. The sample indexing oligonucleotide may be conjugated to the aptamer. The sample indexing oligonucleotide may be conjugated to the aptamer via a chemical group selected from the group consisting of a UV photocleavable group, streptavidin, biotin, an amine, and combinations thereof. In some embodiments, the sample indexing oligonucleotide may be non-covalently attached to the aptamer. The sample indexing oligonucleotide may be associated with the aptamer via a linker.
[0009] In some embodiments, the aptamer comprises a nucleotide aptamer. The nucleotide aptamer may comprise deoxyribonucleic acid (DNA), ribonucleic acid (RNA), xenonucleic acid (XNA), or a combination thereof. The nucleotide aptamer may comprise a base analog. The base analog may comprise a fluorescent base analog. The nucleotide aptamer may comprise a fluorophore. The aptamer may comprise a peptide aptamer. In some embodiments, the sample indexing composition may be associated with a first carrier, and different sample indexing compositions of the plurality of sample indexing compositions may be associated with different first carriers.
[0010] In some embodiments, the aptamer composition may be associated with a first carrier. The aptamer composition may include a second aptamer capable of specifically binding to at least one of one or more protein targets or one or more cellular component targets, and a second sample-indexing oligonucleotide including a second sample-indexing sequence. The aptamer and the second aptamer may be associated with a first carrier. The aptamer may be associated with a first carrier, and the second aptamer may be associated with a second carrier. The aptamer and the second aptamer may have at least 60%, 70%, 80%, 90%, or 95% sequence identity. The aptamer and the second aptamer may, for example, be sequence-identical. The aptamer and the second aptamer may be different. The protein targets or cellular component targets of the aptamer and the second aptamer may be identical. The aptamer and the second aptamer may be capable of binding to different regions of the protein or cellular component target. The protein or cellular component targets of the aptamer and the second aptamer may be different.
[0011] In some embodiments, the first carrier comprises a metal nanomaterial. The first carrier may comprise a metal nanostructure, a metal nanoparticle, or a combination thereof. The first carrier may comprise a gold nanomaterial. The first carrier may comprise a gold nanostructure, a gold nanoparticle, or a combination thereof. The first carrier may comprise a lysosome, a micelle, a vesicle, a lipid membrane, a lipid bilayer, a lipid monolayer, or a combination thereof. In some embodiments, the aptamer may be attached to the first carrier. The aptamer may be covalently attached to the first carrier. The aptamer may be conjugated to the first carrier. The aptamer may be conjugated to the first carrier via a chemical group selected from the group consisting of a UV light-cleavable group, streptavidin, biotin, an amine, and combinations thereof. The aptamer may be non-covalently linked to the first carrier. The aptamer may be associated with the first carrier via a linker. The aptamer may be immobilized on the first carrier, partially immobilized on the first carrier, immobilized within the first carrier, partially immobilized within the first carrier, encapsulated within the first carrier, partially encapsulated within the first carrier, embedded within the first carrier, or combinations thereof. In some embodiments, the method includes dissociating the aptamer from the first support. Dissociation may occur after barcoding the sample-indexing oligonucleotide. Dissociation may occur before barcoding the sample-indexing oligonucleotide.
[0012] In some embodiments, contacting each of the plurality of samples with the sample indexing composition comprises contacting cells of one or more cells of the sample with the sample indexing composition. The first carrier may be internalized into the cell. The first carrier may be internalized into the cell by endocytosis, pinocytosis, nanopinocytosis, micropinocytosis, phagocytosis, membrane fusion, or a combination thereof. The first carrier may be internalized into the cell by clathrin-mediated internalization, caveolin-mediated internalization, receptor-dependent internalization, receptor-independent internalization, or a combination thereof. In some embodiments, the method includes removing unbound sample indexing compositions from the plurality of sample indexing compositions. Removing the unbound sample indexing compositions may include washing one or more cells from each of the plurality of samples with a wash buffer. Removing the unbound sample indexing compositions may include using flow cytometry to select cells that bind to at least one aptamer. In some embodiments, the method includes lysing one or more cells from each of the plurality of samples.
[0013] In some embodiments, the sample indexing oligonucleotide is configured to be non-detachable (or may be non-detachable) from the aptamer. The sample indexing oligonucleotide may be configured to be detachable (or may be detachable) from the aptamer. The method may include detaching the sample indexing oligonucleotide from the aptamer. Detaching the sample indexing oligonucleotide may include detaching the sample indexing oligonucleotide from the aptamer by UV light cleavage, chemical treatment, heating, enzymatic treatment, or any combination thereof. In some embodiments, the sample-indexing oligonucleotides are not homologous to any genomic sequence of the one or more cells, are homologous to genomic sequences of the species, or a combination thereof. The species may be a non-mammalian species.
[0014] In some embodiments, a sample of the plurality of samples comprises a plurality of cells, a plurality of single cells, tissues, tumor samples, or any combination thereof. The plurality of samples may comprise mammalian cells, bacterial cells, viral cells, yeast cells, fungal cells, or any combination thereof. In some embodiments, the sample indexing oligonucleotide comprises a sequence complementary to a capture sequence configured to capture (or capable of capturing, hybridizing, or binding to) the sequence of the sample indexing oligonucleotide. The barcode may comprise a target binding region comprising the capture sequence. The target binding region may comprise a poly(dT) region. The sequence of the sample indexing oligonucleotide complementary to the capture sequence may comprise a poly(dA) region.
[0015] In some embodiments, the sample indexing oligonucleotide comprises an alignment sequence flanking a poly(dA) region. The alignment sequence may be one or more nucleotides in length. The alignment sequence may be two or more nucleotides in length. The alignment sequence may comprise guanine, cytosine, thymine, uracil, or a combination thereof. The alignment sequence may comprise a poly(dT) region, a poly(dG) region, a poly(dC) region, a poly(dU) region, or a combination thereof. The sample indexing oligonucleotide may comprise a molecular beacon sequence, a binding site for a universal primer, or both. The molecular beacon sequence may be 2-20 nucleotides in length. The universal primer may be 5-50 nucleotides in length. The universal primer may comprise an amplification primer, a sequencing primer, or a combination thereof.
[0016] In some embodiments, an aptamer is associated with two or more sample indexing oligonucleotides having the same sequence. An aptamer may be associated with two or more sample indexing oligonucleotides having different sample indexing sequences. In some embodiments, a sample indexing composition of the plurality of sample indexing compositions includes a second aptamer that is not associated with a sample indexing oligonucleotide. The aptamer and the second aptamer may be the same. In some embodiments, each of the plurality of sample indexing compositions includes an aptamer.
[0017] In some embodiments, a sample indexing composition of the plurality of sample indexing compositions includes a second aptamer composition comprising a second protein-binding aptamer, wherein the second aptamer is capable of specifically binding to at least one of one or more protein targets or at least one of one or more cellular component targets. The aptamer and the second aptamer may be capable of binding to the same protein target(s) of one or more protein targets or one or more cellular component targets, and the second aptamer is not associated with a sample indexing oligonucleotide. The second aptamer may be associated with a second sample indexing oligonucleotide comprising a second sample indexing sequence. The aptamer and the second aptamer may be at least 60%, 70%, 80%, 90%, or 95% identical in sequence. The aptamer and the second aptamer may, for example, be identical in sequence. The aptamer and the second aptamer may be different. The aptamer and the second aptamer may be capable of binding to different regions of the same protein target or different regions of the same cellular component target. The aptamer and the second aptamer may be capable of binding to different protein targets of one or more protein targets or different cellular component targets of one or more cellular component targets. The sample indexing sequence and the second sample indexing sequence may be the same. The sample indexing sequence and the second sample indexing sequence may be different.
[0018] In some embodiments, the method comprises pooling multiple samples contacted with multiple sample indexing compositions prior to barcoding the sample-indexing oligonucleotides. In some embodiments, a barcode of the plurality of barcodes comprises a target-binding region and a molecular beacon sequence, and the molecular beacon sequences of at least two barcodes of the plurality of barcodes comprise different molecular beacon sequences. The barcode may comprise a cell beacon sequence, a binding site for a universal primer, or any combination thereof. The target-binding region may comprise a poly(dT) region. In some embodiments, the plurality of barcodes are associated with a particle. At least one barcode of the plurality of barcodes may be immobilized on the particle, partially immobilized on the particle, encapsulated within the particle, partially encapsulated within the particle, or a combination thereof. The particle may be destructible. The particle may comprise a bead. The particles may comprise sepharose beads, streptavidin beads, agarose beads, magnetic beads, conjugated beads, protein A conjugated beads, protein G conjugated beads, protein A / G conjugated beads, protein L conjugated beads, oligo(dT) conjugated beads, silica beads, silica-like beads, hydrogel beads, anti-biotin microbeads, anti-fluorescent dye microbeads, or any combination thereof, or the particles comprise a material selected from the group consisting of polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogel, paramagnetic material, ceramic, plastic, glass, methylstyrene, acrylic polymer, titanium, latex, sepharose, cellulose, nylon, silicone, and any combination thereof. The particles may comprise breakable hydrogel beads.
[0019] In some embodiments, the barcode of the particle comprises a molecular label sequence selected from at least 1,000, 10,000, or a combination thereof, different molecular label sequences. The molecular label sequence of the barcode may comprise a random sequence. The particle may comprise at least 10,000 barcodes. In some embodiments, barcoding the sample-indexing oligonucleotides with a plurality of barcodes comprises contacting the plurality of barcodes with the sample-indexing oligonucleotides to generate barcodes hybridized to the sample-indexing oligonucleotides; and extending the barcodes hybridized to the sample-indexing oligonucleotides to generate a plurality of barcoded sample-indexing oligonucleotides.
[0020] In some embodiments, the method includes pooling barcodes hybridized to the sample-indexing oligonucleotides before extending the barcodes hybridized to the sample-indexing oligonucleotides, and extending the barcodes hybridized to the sample-indexing oligonucleotides includes extending the pooled barcodes hybridized to the sample-indexing oligonucleotides to generate a plurality of pooled barcoded sample-indexing oligonucleotides. Extending the barcodes may include extending the barcodes using a DNA polymerase to generate a plurality of barcoded sample-indexing oligonucleotides. Extending the barcodes may include extending the barcodes using a reverse transcriptase to generate a plurality of barcoded sample-indexing oligonucleotides.
[0021] In some embodiments, the method includes amplifying a plurality of barcoded sample-indexing oligonucleotides to produce a plurality of amplicons. Amplifying the plurality of barcoded sample-indexing oligonucleotides may include amplifying at least a portion of the molecular beacon sequences and at least a portion of the sample-indexing oligonucleotides using polymerase chain reaction (PCR). Obtaining sequencing data for the plurality of barcoded sample-indexing oligonucleotides may include obtaining sequencing data for the plurality of amplicons. Obtaining sequencing data may include sequencing at least a portion of the molecular beacon sequences and at least a portion of the sample-indexing oligonucleotides.
[0022] In some embodiments, barcoding the sample-indexing oligonucleotides with a plurality of barcodes to generate a plurality of barcoded sample-indexing oligonucleotides comprises probabilistically barcoding the sample-indexing oligonucleotides with a plurality of stochastic barcodes to generate a plurality of stochastically barcoded sample-indexing oligonucleotides.
[0023] In some embodiments, the method includes barcoding a plurality of cellular targets using a plurality of barcodes to generate a plurality of barcoded targets, each of the plurality of barcodes including a cell labeling sequence, and at least two of the plurality of barcodes including the same cell labeling sequence; and obtaining sequencing data for the barcoded targets. Barcoding the plurality of targets using a plurality of barcodes to generate a plurality of barcoded targets may include contacting copies of the targets with target-binding regions of the barcodes; and reverse transcribing the plurality of targets using the plurality of barcodes to generate a plurality of reverse-transcribed targets. The method may include amplifying the barcoded targets to generate a plurality of amplified barcoded targets before obtaining sequencing data for the plurality of barcoded targets. Amplifying the barcoded targets to generate a plurality of amplified barcoded targets may include amplifying the barcoded targets by polymerase chain reaction (PCR). Barcoding a plurality of targets of a cell using a plurality of barcodes to generate a plurality of barcoded targets may include probabilistically barcoding a plurality of targets of a cell using a plurality of stochastic barcodes to generate a plurality of stochastically barcoded targets.
[0024] Disclosed herein are embodiments of a plurality of sample indexing compositions. In some embodiments, each of the plurality of sample indexing compositions comprises an aptamer composition comprising a first cellular component-binding aptamer and a sample indexing oligonucleotide, wherein the cellular component-binding aptamer is capable of specifically binding to at least one cellular component target, and the sample indexing oligonucleotide comprises a sample indexing sequence for identifying the sample origin of one or more cells of the sample, and the sample indexing sequences of at least two sample indexing compositions of the plurality of sample indexing compositions comprise different sequences.
[0025] In some embodiments, the sample indexing sequences are 6-60 nucleotides in length. The sample indexing oligonucleotides may be 50-500 nucleotides in length. The sample indexing sequences of at least 10, 100, or 1000 sample indexing compositions of the plurality of sample indexing compositions may comprise different sequences. In some embodiments, the single polynucleotide comprises a sample-indexing oligonucleotide and a cellular component-binding aptamer. The aptamer may be located 5' to the sample-indexing oligonucleotide in the single polynucleotide. The aptamer may be located 3' to the sample-indexing oligonucleotide in the single polynucleotide.
[0026] In some embodiments, the sample indexing oligonucleotide is associated with an aptamer. The sample indexing oligonucleotide may be attached to the cellular component-binding aptamer. The sample indexing oligonucleotide may be covalently attached to the cellular component-binding aptamer. The sample indexing oligonucleotide may be conjugated to the cellular component-binding aptamer. The sample indexing oligonucleotide may be conjugated to the cellular component-binding aptamer via a chemical group selected from the group consisting of a UV photocleavable group, streptavidin, biotin, an amine, and combinations thereof. The sample indexing oligonucleotide may be non-covalently attached to the cellular component-binding aptamer. The sample indexing oligonucleotide may be associated with the protein-binding aptamer or the cellular component-binding aptamer via a linker.
[0027] In some embodiments, the protein-binding aptamer or the cellular component-binding aptamer comprises a nucleotide aptamer. The nucleotide aptamer may comprise deoxyribonucleic acid (DNA), ribonucleic acid (RNA), xenonucleic acid (XNA), or a combination thereof. The nucleotide aptamer may comprise a base analog. The base analog may comprise a fluorescent base analog. The nucleotide aptamer may comprise a fluorophore. The cellular component-binding aptamer may comprise a peptide aptamer. In some embodiments, the sample indexing composition is associated with a first carrier. Different sample indexing compositions of the plurality of sample indexing compositions may be associated with different first carriers. In some embodiments, the aptamer composition is associated with a first carrier. In some embodiments, the aptamer composition comprises a second cellular component-binding aptamer capable of specifically binding to at least one of one or more protein targets or one or more cellular component targets, and a second sample-indexing oligonucleotide comprising a second sample-indexing sequence. The cellular component-binding aptamer and the second cellular component-binding aptamer may be associated with a first carrier. The cellular component-binding aptamer may be associated with a first carrier, and the second cellular component-binding aptamer may be associated with a second carrier.
[0028] In some embodiments, the cellular component-binding aptamer and the second cellular component-binding aptamer have at least 60%, 70%, 80%, 90%, or 95% sequence identity. The cellular component-binding aptamer and the second cellular component-binding aptamer may, for example, have identical sequences. The cellular component-binding aptamer and the second cellular component-binding aptamer may also be different. In some embodiments, the protein target or cellular component target of the cellular component-binding aptamer and the second cellular component-binding aptamer are the same. The cellular component-binding aptamer and the second cellular component-binding aptamer may be capable of binding to different regions of the protein target or cellular component target. The protein target or cellular component target of the cellular component-binding aptamer and the second cellular component-binding aptamer may be different.
[0029] In some embodiments, the first carrier comprises a metal nanomaterial. The first carrier may comprise a metal nanostructure, a metal nanoparticle, or a combination thereof. The first carrier may comprise a gold nanomaterial. The first carrier may comprise a gold nanostructure, a gold nanoparticle, or a combination thereof. The first carrier may comprise a lysosome, a micelle, a vesicle, a lipid membrane, a lipid bilayer, a lipid monolayer, or a combination thereof. In some embodiments, the cellular component-binding aptamer is attached to a first carrier. The cellular component-binding aptamer may be covalently attached to the first carrier. The cellular component-binding aptamer may be conjugated to the first carrier. The cellular component-binding aptamer may be conjugated to the first carrier via a chemical group selected from the group consisting of a UV light-cleavable group, streptavidin, biotin, an amine, and combinations thereof. The cellular component-binding aptamer may be non-covalently linked to the first carrier. The cellular component-binding aptamer may be associated with the first carrier via a linker. In some embodiments, the cellular component-binding aptamer is immobilized on the first carrier, partially immobilized on the first carrier, immobilized within the first carrier, partially immobilized within the first carrier, encapsulated within the first carrier, partially encapsulated within the first carrier, embedded within the first carrier, partially embedded within the first carrier, or a combination thereof.
[0030] In some embodiments, the first carrier is configured to be internalized (or be capable of being internalized) into a cell. The first carrier may be internalized into a cell by endocytosis, pinocytosis, nanopinocytosis, micropinocytosis, phagocytosis, membrane fusion, or a combination thereof. The first carrier may be internalized into a cell by clathrin-mediated internalization, caveolin-mediated internalization, receptor-dependent internalization, receptor-independent internalization, or a combination thereof.
[0031] In some embodiments, the sample indexing oligonucleotide is not homologous to any genomic sequence of one or more cells.At least one sample of the plurality of samples may comprise one or more single cells, multiple cells, tissues, tumor samples, or any combination thereof.The sample may comprise a mammalian sample, a bacterial sample, a viral sample, a yeast sample, a fungal sample, or any combination thereof. In some embodiments, the sample indexing oligonucleotide comprises a sequence complementary to a capture sequence configured to capture (or capable of capturing, hybridizing, or binding to) the sequence of the sample indexing oligonucleotide. The target binding region may comprise a poly(dT) region. The sequence of the sample indexing oligonucleotide complementary to the capture sequence may comprise a poly(dA) region.
[0032] In some embodiments, the sample-indexing oligonucleotide comprises an alignment sequence adjacent to the poly(dA) region. The alignment sequence may be 1 or more nucleotides in length. The alignment sequence may be 2 or more nucleotides in length. The alignment sequence may comprise guanine, cytosine, thymine, uracil, or a combination thereof. The alignment sequence may comprise a poly(dT) region, a poly(dG) region, a poly(dC) region, a poly(dU) region, or a combination thereof. In some embodiments, the sample-indexing oligonucleotide comprises a molecular beacon sequence, a poly(dA) region, or a combination thereof. The molecular beacon sequence may be 2 to 20 nucleotides in length. The universal primer may be 5 to 50 nucleotides in length. The universal primer may comprise an amplification primer, a sequencing primer, or a combination thereof.
[0033] In some embodiments, the cellular component target comprises a cell surface protein, a cell marker, a B cell receptor, a T cell receptor, a major histocompatibility complex, a tumor antigen, a receptor, or any combination thereof. The cellular component target can be selected from a group comprising 10 to 100 different cellular component targets. In some embodiments, at least one of the one or more protein targets or the one or more cellular component targets is on the cell surface. In some embodiments, the protein target or cellular component target comprises a carbohydrate, a lipid, a protein, an extracellular protein, a cell surface protein, a cell marker, a B cell receptor, a T cell receptor, a major histocompatibility complex, a tumor antigen, a receptor, an intracellular protein, or any combination thereof. The protein target or cellular component target can be selected from a group comprising 10 to 100 different protein targets or cellular component targets.
[0034] In some embodiments, a cellular component-binding aptamer is associated with two or more sample-indexing oligonucleotides having the same sequence. A cellular component-binding aptamer may also be associated with two or more sample-indexing oligonucleotides having different sample-indexing sequences. In some embodiments, the sample indexing composition comprises a second cellular component-binding aptamer, wherein the second cellular component-binding aptamer is capable of specifically binding to at least one of the one or more cellular component targets. The cellular component-binding aptamer and the second cellular component-binding aptamer may be capable of binding to the same cellular component target of one or more cellular component targets, and the second cellular component-binding aptamer may not be associated with a sample indexing oligonucleotide. The second cellular component-binding aptamer may be associated with a second sample indexing oligonucleotide comprising a second sample indexing sequence. The sample indexing sequence and the second sample indexing sequence may be identical. The sample indexing sequence and the second sample indexing sequence may be different.
[0035] In some embodiments, the cellular component-binding aptamer and the second cellular component-binding aptamer are at least 60%, 70%, 80%, 90%, or 95% identical. The cellular component-binding aptamer and the second cellular component-binding aptamer may be identical. The cellular component-binding aptamer and the second cellular component-binding aptamer may be capable of binding to different regions of the same protein target. The cellular component-binding aptamer and the second cellular component-binding aptamer may be capable of binding to different protein targets of one or more protein targets.
[0036] Disclosed herein are embodiments of a kit for identifying a sample. In some embodiments, the kit includes a plurality of sample indexing compositions. In some embodiments, the kit includes a plurality of barcodes. A barcode of the plurality of barcodes includes a target binding region. The target binding region may include a capture sequence configured to capture (or capable of capturing, hybridizing, or binding to) a sequence of the sample indexing oligonucleotide. The cellular component-binding aptamer may include a sample indexing oligonucleotide and a first cellular component-binding aptamer.
[0037] Disclosed herein are embodiments of a method for measuring protein expression in cells. In some embodiments, the method includes contacting a plurality of protein-binding aptamers with a plurality of cells comprising a plurality of protein targets, each of the plurality of protein-binding aptamers comprising an aptamer-specific oligonucleotide comprising a unique identifier sequence for the cellular component-binding aptamer, and the protein-binding aptamer being capable of specifically binding to at least one of the plurality of protein targets; distributing the plurality of cells associated with the plurality of protein-binding aptamers into a plurality of compartments, each of the plurality of compartments comprising a single cell derived from the plurality of cells associated with the protein-binding aptamers; and distributing the barcode in the compartment comprising the single cell. The method includes contacting bar-coded particles with aptamer-specific oligonucleotides, wherein the bar-coded particles comprise a plurality of oligonucleotide probes, each comprising a target-binding region and a barcode sequence selected from a diverse set of unique barcode sequences; extending the oligonucleotide probes hybridized to the aptamer-specific oligonucleotides to produce a plurality of labeled nucleic acids, wherein each of the labeled nucleic acids comprises a unique identifier sequence or its complementary sequence and a barcode sequence; and obtaining sequence information of the plurality of labeled nucleic acids or portions thereof to determine the amount of one or more of the plurality of protein targets in one or more of the plurality of cells.
[0038] The disclosure herein includes the embodiment of a method for measuring cellular component expression in cells.In some embodiments, this method includes: contacting a plurality of cellular component binding aptamers with a plurality of cells, comprising a plurality of cellular component targets, wherein each of the plurality of cellular component binding aptamers comprises an aptamer-specific oligonucleotide comprising a unique identifier sequence for the cellular component binding aptamer, and the cellular component binding aptamer can specifically bind to at least one of the plurality of cellular component targets; extending the oligonucleotide probe hybridized to the aptamer-specific oligonucleotide to produce a plurality of labeled nucleic acids, wherein each of the labeled nucleic acids comprises a unique identifier sequence or its complementary sequence and a barcode sequence; and obtaining the sequence information of the plurality of labeled nucleic acids or a portion thereof, and determining the amount of one or more of the plurality of cellular component targets in one or more of the plurality of cells. The method may include, prior to extending the oligonucleotide probe, distributing a plurality of cells having associated thereto a plurality of cellular component-binding aptamers into a plurality of compartments, wherein a compartment of the plurality of compartments contains a single cell derived from the plurality of cells having associated thereto a cellular component-binding aptamer; and contacting a bar-coding particle with an aptamer-specific oligonucleotide in the compartment containing the single cell, wherein the bar-coding particle includes a plurality of oligonucleotide probes, each of the bar-coding particles including a target-binding region and a barcode sequence selected from a diverse set of unique barcode sequences. The plurality of cellular component targets includes a plurality of protein targets, and the cellular component-binding aptamer is capable of specifically binding to at least one of the plurality of protein targets.
[0039] In some embodiments, the aptamer-specific oligonucleotide and the aptamer form a single polynucleotide. The aptamer may be located 5' of the aptamer-specific oligonucleotide in the single polynucleotide. The aptamer may be located 3' of the aptamer-specific oligonucleotide in the single polynucleotide. The aptamer-specific oligonucleotide may be associated with the aptamer. The aptamer-specific oligonucleotide may be attached to the aptamer. The aptamer-specific oligonucleotide may be covalently attached to the aptamer. The aptamer-specific oligonucleotide may be non-covalently attached to the aptamer. The aptamer-specific oligonucleotide may be conjugated to the aptamer. The aptamer-specific oligonucleotide may be conjugated to the aptamer via a chemical group selected from the group consisting of UV light-cleavable groups, streptavidin, biotin, amines, and combinations thereof. The aptamer-specific oligonucleotide may be covalently attached to the aptamer. In some embodiments, the method includes detaching the aptamer-specific oligonucleotide from the aptamer.
[0040] In some embodiments, the aptamer comprises a nucleotide aptamer. The nucleotide aptamer may comprise deoxyribonucleic acid (DNA), ribonucleic acid (RNA), xenonucleic acid (XNA), or a combination thereof. The nucleotide aptamer may comprise a base analog. The base analog may comprise a fluorescent base analog. The nucleotide aptamer may comprise a fluorophore. The aptamer may comprise a peptide aptamer. In some embodiments, the plurality of protein-binding aptamers includes a second protein-binding aptamer, or the plurality of cellular component-binding aptamers includes a second cellular component-binding aptamer. The aptamer and the second aptamer may be associated with a first carrier. The aptamer may be associated with a first carrier, and the second aptamer may be associated with a second carrier. The aptamer and the second aptamer may have at least 60%, 70%, 80%, 90%, or 95% sequence identity. The aptamer and the second protein-binding aptamer may be identical. The aptamer and the second aptamer may be different. In some embodiments, the protein target or cellular component target of the aptamer and the second aptamer are the same.The aptamer and the second aptamer may be capable of binding to different regions of the protein target or cellular component target.The protein target or cellular component target of the aptamer and the second aptamer may be different.
[0041] In some embodiments, the first carrier comprises a metal nanomaterial. The first carrier may comprise a metal nanostructure, a metal nanoparticle, or a combination thereof. The first carrier may comprise a gold nanomaterial. The first carrier may comprise a gold nanostructure, a gold nanoparticle, or a combination thereof. The first carrier may comprise a lysosome, a micelle, a vesicle, a lipid membrane, a lipid bilayer, a lipid monolayer, or a combination thereof.
[0042] In some embodiments, the aptamer is attached to the first carrier. The aptamer may be covalently attached to the first carrier. The aptamer may be conjugated to the first carrier. The aptamer may be conjugated to the first carrier via a chemical group selected from the group consisting of a UV light-cleavable group, streptavidin, biotin, an amine, and combinations thereof. The aptamer may be non-covalently linked to the first carrier. The aptamer may be associated with the first carrier via a linker. The aptamer may be immobilized on the first carrier, partially immobilized on the first carrier, immobilized within the first carrier, partially immobilized within the first carrier, encapsulated within the first carrier, partially encapsulated within the first carrier, embedded within the first carrier, partially embedded within the first carrier, or a combination thereof.
[0043] In some embodiments, the method includes dissociating the aptamer from the first carrier. Dissociation may occur after contact with the compartment containing the single cell. Dissociation may occur before contact with the compartment containing the single cell. In some embodiments, contacting the plurality of aptamers with the plurality of cells comprises contacting a first carrier containing the aptamers with a cell from the plurality of cells containing the plurality of protein targets. The first carrier may be internalized into the cell. The first carrier may be internalized into the cell by endocytosis, pinocytosis, nanopinocytosis, micropinocytosis, phagocytosis, membrane fusion, or a combination thereof. The first carrier may be internalized into the cell by clathrin-mediated internalization, caveolin-mediated internalization, receptor-dependent internalization, receptor-independent internalization, or a combination thereof.
[0044] In some embodiments, the aptamer-specific oligonucleotide comprises a sequence complementary to a capture sequence configured to capture (or capable of capturing, hybridizing, or binding to) the sequence of the aptamer-specific oligonucleotide. The barcode may comprise a target binding region comprising the capture sequence. The target binding region may comprise a poly(dT) region. The sequence of the aptamer-specific oligonucleotide complementary to the capture sequence may comprise a poly(dA) region. In some embodiments, the aptamer-specific oligonucleotide changes from a first conformation in which the poly(dA) region is inaccessible to a second conformation in which the poly(dA) region is accessible when the aptamer contacts at least one of the multiple protein targets or cellular component targets. The poly(dA) region of the aptamer-specific oligonucleotide in the first conformation may form a hairpin structure. The poly(dA) region with the hairpin structure may be accessible to the poly(dT) region of the target-binding region of the barcode. The oligonucleotide probe can be hybridized to the aptamer-specific oligonucleotide by hybridization between the poly(dA) region of the aptamer-specific oligonucleotide and the poly(dT) region of the oligonucleotide probe.
[0045] In some embodiments, the aptamer-specific oligonucleotide comprises an alignment sequence adjacent to the poly(dA) region. The alignment sequence may be one or more nucleotides in length. The alignment sequence may be two nucleotides in length. The alignment sequence may comprise guanine, cytosine, thymine, uracil, or a combination thereof. The alignment sequence may comprise a poly(dT) region, a poly(dG) region, a poly(dC) region, a poly(dU) region, or a combination thereof. In some embodiments, the aptamer-specific oligonucleotide comprises a molecular beacon sequence, a binding site for a universal primer, or both. The molecular beacon sequence may be 2 to 20 nucleotides in length. The universal primer may be 5 to 50 nucleotides in length. The universal primer may comprise an amplification primer, a sequencing primer, or a combination thereof.
[0046] In some embodiments, the method includes contacting the plurality of aptamers with the plurality of cells and then removing aptamers that are not in contact with the plurality of cells. Removing aptamers that are not in contact with the plurality of cells may include removing aptamers that are not in contact with at least one corresponding one of the plurality of protein targets. Removing aptamers that are not in contact with at least one corresponding one of the plurality of protein targets may include using an aptamer sponge to remove aptamers that are not in contact with at least one corresponding one of the plurality of protein targets. The aptamer sponge may include multiple sequences complementary to the sequences of the aptamers. Aptamers that are not in contact with at least one corresponding one of the plurality of protein targets can hybridize to the multiple sequences of the aptamer sponge. In some embodiments, the aptamer comprises a complementary sequence. Aptamers that are not in contact with at least one corresponding protein target of the plurality of protein targets may hybridize with each other to form aggregates. The aggregates may be targeted for degradation in cells.
[0047] In some embodiments, removing aptamers that are not in contact with at least one corresponding one of the multiple protein targets comprises using multiple antisense aptamers to remove aptamers that are not in contact with at least one corresponding one of the multiple protein targets, wherein each antisense aptamer of the multiple antisense aptamers comprises a second aptamer comprising an antisense aptamer-specific oligonucleotide that comprises a complementary sequence or a portion of an aptamer-specific oligonucleotide of one of the aptamers that is not in contact with at least one corresponding one of the multiple protein targets, and wherein the aptamer-specific oligonucleotide and the antisense aptamer-specific oligonucleotide form a double-stranded or partially double-stranded deoxyribonucleic acid (DNA) molecule. In some embodiments, distributing the plurality of cells comprises distributing a plurality of bar-coding particles, including a plurality of cells associated with a plurality of aptamers, into a plurality of compartments, wherein a compartment of the plurality of compartments comprises a single cell derived from the plurality of cells associated with the oligonucleotide-conjugated aptamers and bar-coding particles. In some embodiments, the compartment is a well or a droplet. In some embodiments, obtaining sequencing information of the plurality of labeled nucleic acids or a portion thereof comprises subjecting the labeled nucleic acids to one or more reactions to generate a set of nucleic acids for nucleic acid sequencing. In some embodiments, the plurality of protein targets comprises cell surface proteins, intracellular proteins, cell markers, B cell receptors, T cell receptors, antibodies, major histocompatibility complexes, tumor antigens, receptors, or combinations thereof. In some embodiments, the plurality of cells comprises T cells, B cells, tumor cells, myeloid cells, blood cells, normal cells, fetal cells, maternal cells, or a mixture thereof.
[0048] In some embodiments, each of the oligonucleotide probes comprises a cell label, a binding site for a universal primer, an amplification adaptor, a sequencing adaptor, or a combination thereof. In some embodiments, the bar-coding particles are sepharose beads, streptavidin beads, agarose beads, magnetic beads, silica beads, silica-like beads, hydrogel beads, anti-biotin microbeads, anti-fluorescent dye microbeads, or hydrogel beads.
[0049] Disclosed herein is a composition comprising a plurality of cellular component-binding aptamers, each of the plurality of cellular component-binding aptamers comprising an aptamer-specific oligonucleotide comprising a unique identifier sequence for the cellular component-binding aptamer, wherein the cellular component-binding aptamer is capable of specifically binding to at least one of the plurality of cellular component targets. In some embodiments, the aptamer-specific oligonucleotide and the aptamer form a single polynucleotide. The cellular component-binding aptamer may be located 5' to the aptamer-specific oligonucleotide in the single polynucleotide. The cellular component-binding aptamer may be located 3' to the aptamer-specific oligonucleotide in the single polynucleotide.
[0050] In some embodiments, the aptamer-specific oligonucleotide is associated with the cellular component-binding aptamer. The aptamer-specific oligonucleotide may be attached to the cellular component-binding aptamer. The aptamer-specific oligonucleotide may be covalently attached to the cellular component-binding aptamer. The aptamer-specific oligonucleotide may be non-covalently attached to the cellular component-binding aptamer. The aptamer-specific oligonucleotide may be conjugated to the cellular component-binding aptamer. The aptamer-specific oligonucleotide may be conjugated to the cellular component-binding aptamer via a chemical group selected from the group consisting of UV light-cleavable groups, streptavidin, biotin, amines, and combinations thereof. The aptamer-specific oligonucleotide may be covalently attached to the cellular component-binding aptamer.
[0051] In some embodiments, the cellular component-binding aptamer comprises a nucleotide aptamer. The nucleotide aptamer may comprise deoxyribonucleic acid (DNA), ribonucleic acid (RNA), xenonucleic acid (XNA), or a combination thereof. The nucleotide aptamer may comprise a base analog. The base analog may comprise a fluorescent base analog. The nucleotide aptamer may comprise a fluorophore. The cellular component-binding aptamer may comprise a peptide aptamer. In some embodiments, the plurality of cellular component-binding aptamers includes a second cellular component-binding aptamer. The cellular component-binding aptamer and the second cellular component-binding aptamer may be associated with a first carrier. The cellular component-binding aptamer may be associated with the first carrier, and the second cellular component-binding aptamer may be associated with a second carrier.
[0052] In some embodiments, the cellular component-binding aptamer and the second cellular component-binding aptamer may have at least 60%, 70%, 80%, 90%, or 95% sequence identity. The cellular component-binding aptamer and the second protein-binding aptamer may be the same. The cellular component-binding aptamer and the second cellular component-binding aptamer may be different. In some embodiments, the protein target or cellular component target of the cellular component-binding aptamer and the second cellular component-binding aptamer are the same. The cellular component-binding aptamer and the second cellular component-binding aptamer may be capable of binding to different regions of the protein target or cellular component target. The protein target or cellular component target of the cellular component-binding aptamer and the second cellular component-binding aptamer may be different.
[0053] In some embodiments, the first carrier comprises a metal nanomaterial. The first carrier may comprise a metal nanostructure, a metal nanoparticle, or a combination thereof. The first carrier may comprise a gold nanomaterial. The first carrier may comprise a gold nanostructure, a gold nanoparticle, or a combination thereof. The first carrier may comprise a lysosome, a micelle, a vesicle, a lipid membrane, a lipid bilayer, a lipid monolayer, or a combination thereof. In some embodiments, the cellular component-binding aptamer is attached to a first carrier. The cellular component-binding aptamer may be covalently attached to the first carrier. The cellular component-binding aptamer may be conjugated to the first carrier. The cellular component-binding aptamer may be conjugated to the first carrier via a chemical group selected from the group consisting of a UV light-cleavable group, streptavidin, biotin, an amine, and combinations thereof. The cellular component-binding aptamer may be non-covalently linked to the first carrier. The cellular component-binding aptamer may be associated with the first carrier via a linker. The cellular component-binding aptamer may be immobilized on the first carrier, partially immobilized on the first carrier, immobilized within the first carrier, partially immobilized within the first carrier, encapsulated within the first carrier, partially encapsulated within the first carrier, embedded within the first carrier, partially embedded within the first carrier, or a combination thereof.
[0054] In some embodiments, the first carrier is configured to be (or be capable of being) internalized into a cell. The first carrier may be configured to be (or be capable of being) internalized into a cell by endocytosis, pinocytosis, nanopinocytosis, micropinocytosis, phagocytosis, membrane fusion, or a combination thereof. The first carrier may be configured to be (or be capable of being) internalized into a cell by clathrin-mediated internalization, caveolin-mediated internalization, receptor-dependent internalization, receptor-independent internalization, or a combination thereof.
[0055] In some embodiments, the aptamer-specific oligonucleotide comprises a sequence complementary to a capture sequence configured to capture (or capable of capturing, binding to, or hybridizing to) the sequence of the aptamer-specific oligonucleotide. The barcode may comprise a target-binding region comprising the capture sequence. The target-binding region may comprise a poly(dT) region. The sequence of the aptamer-specific oligonucleotide complementary to the capture sequence may comprise a poly(dA) region. In some embodiments, the aptamer-specific oligonucleotide changes from a first conformation in which the poly(dA) region is inaccessible to a second conformation in which the poly(dA) region is accessible when the aptamer contacts at least one of the multiple protein targets or cellular component targets. The poly(dA) region of the aptamer-specific oligonucleotide in the first conformation may form a hairpin structure. The poly(dA) region with the hairpin structure may be accessible to the poly(dT) region of the target-binding region of the barcode. The oligonucleotide probe can be hybridized to the aptamer-specific oligonucleotide by hybridization between the poly(dA) region of the aptamer-specific oligonucleotide and the poly(dT) region of the oligonucleotide probe.
[0056] In some embodiments, the aptamer-specific oligonucleotide comprises an alignment sequence adjacent to the poly(dA) region. The alignment sequence may be one or more nucleotides in length. The alignment sequence may be two nucleotides in length. The alignment sequence may comprise guanine, cytosine, thymine, uracil, or a combination thereof. The alignment sequence may comprise a poly(dT) region, a poly(dG) region, a poly(dC) region, a poly(dU) region, or a combination thereof. In some embodiments, the aptamer-specific oligonucleotide may comprise a molecular beacon sequence, a binding site for a universal primer, or both. The molecular beacon sequence may be 2 to 20 nucleotides in length. The universal primer may be 5 to 50 nucleotides in length. The universal primer may comprise an amplification primer, a sequencing primer, or a combination thereof.
[0057] In some embodiments, the composition comprises an aptamer sponge. The aptamer sponge may comprise multiple sequences complementary to the sequence of an aptamer. In some embodiments, the aptamers comprise complementary sequences. The aptamers may be configured to hybridize (or be allowed to hybridize) with one another to form aggregates. The aggregates may be targeted for degradation in the cell.
[0058] In some embodiments, the composition comprises a plurality of antisense aptamers, each antisense aptamer of the plurality of antisense aptamers comprising a second aptamer comprising an antisense aptamer-specific oligonucleotide that comprises a complementary sequence or a portion of an aptamer-specific oligonucleotide of one of the aptamers, and the aptamer-specific oligonucleotide and the antisense aptamer-specific oligonucleotide form a double-stranded or partially double-stranded deoxyribonucleic acid (DNA) molecule. [Brief explanation of the drawings]
[0059] [Figure 1] FIG. 1 illustrates a non-limiting exemplary probabilistic barcode. [Figure 2] FIG. 1 illustrates a non-limiting exemplary workflow for probabilistic barcoding and electronic counting. [Figure 3] 1 is a schematic diagram showing a non-limiting exemplary process for creating an indexed library of stochastically barcoded targets from multiple targets. [Figure 4]1 shows a schematic diagram of an exemplary protein binding reagent (an antibody as shown in this figure) with an oligonucleotide associated therewith that comprises a unique identifier for the protein binding reagent. [Figure 5] 1 shows a schematic diagram of an exemplary binding reagent (antibody shown in this figure) associated with an oligonucleotide containing a unique identifier for sample indexing to determine cells from the same or different samples. [Figure 6] 1 shows a schematic diagram of an exemplary workflow using antibodies conjugated with oligonucleotides to simultaneously determine expression of cellular components (e.g., protein expression) and gene expression in a high-throughput manner. [Figure 7] 1 shows a schematic diagram of an exemplary workflow using antibodies conjugated with oligonucleotides for sample indexing. [Figure 8A] 1 shows a schematic diagram of an exemplary workflow using aptamers associated (e.g., conjugated) with oligonucleotides to determine expression of cellular components (e.g., protein expression). [Figure 8B] 1 shows a schematic diagram of an exemplary workflow using aptamers associated (e.g., conjugated) with oligonucleotides to determine expression of cellular components (e.g., protein expression). [Figure 8C] 1 shows a schematic diagram of an exemplary workflow using aptamers associated (e.g., conjugated) with oligonucleotides to determine expression of cellular components (e.g., protein expression). [Figure 8D] 1 shows a schematic diagram of an exemplary workflow using aptamers associated (e.g., conjugated) with oligonucleotides to determine expression of cellular components (e.g., protein expression). [Figure 8E] 1 shows a schematic diagram of an exemplary workflow using aptamers associated (e.g., conjugated) with oligonucleotides to determine expression of cellular components (e.g., protein expression). [Figure 9] FIG. 1 shows non-limiting exemplary aptamers and associated sequences. [Figure 10A] 1 shows a schematic diagram of an exemplary workflow using aptamers with attached oligonucleotides for sample indexing. [Figure 10B] 1 shows a schematic diagram of an exemplary workflow using aptamers with attached oligonucleotides for sample indexing. [Figure 11A] FIG. 1 shows non-limiting exemplary designs of oligonucleotides for simultaneous determination of protein and gene expression and for sample indexing. [Figure 11B] FIG. 1 shows non-limiting exemplary designs of oligonucleotides for simultaneous determination of protein and gene expression and for sample indexing. [Figure 11C] FIG. 1 shows non-limiting exemplary designs of oligonucleotides for simultaneous determination of protein and gene expression and for sample indexing. [Figure 11D] FIG. 1 shows non-limiting exemplary designs of oligonucleotides for simultaneous determination of protein and gene expression and for sample indexing. [Figure 12] 1 shows a schematic representation of non-limiting exemplary oligonucleotide sequences for simultaneously determining protein and gene expression and for sample indexing. DETAILED DESCRIPTION OF THE INVENTION
[0060] In the following detailed description, reference is made to the accompanying drawings, which form a part of this specification. In the drawings, like symbols typically indicate like components unless the context dictates otherwise. The illustrative embodiments described in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter provided herein. It will be readily understood that the aspects of the present disclosure, as generally described herein and illustrated in the figures, may be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are expressly contemplated herein and form part of the disclosure herein. All patents, published patent applications, other publications, and GenBank and other database sequences referenced herein are incorporated by reference in their entirety for relevant art.
[0061] Quantification of a small number of nucleic acids, such as messenger ribonucleotide acid (mRNA) molecules, is clinically important, for example, to determine the genes expressed in cells at different developmental stages or under different environmental conditions. However, determining the absolute number of nucleic acid molecules (e.g., mRNA molecules) can be very difficult, especially when the number of molecules is very small. One method for determining the absolute number of molecules in a sample is digital polymerase chain reaction (PCR). Ideally, PCR produces identical copies of molecules in each cycle. However, PCR has the disadvantage that each molecule is replicated with a stochastic probability, which varies depending on the PCR cycle and gene sequence, resulting in amplification bias and inaccurate gene expression measurement. Using stochastic barcodes with unique molecular labels (also called molecular indexes (MIs)), the number of molecules can be counted and amplification bias can be corrected. Stochastic barcoding, such as the Precise™ assay (Cellular Research, Inc., Palo Alto, CA), can correct for biases induced by PCR and library preparation steps by labeling mRNA using molecular beacons (MLs) during reverse transcription (RT).
[0062] The Precise™ assay uses a non-depleting pool of stochastic barcodes with multiple, e.g., 6561-65536, unique molecular labels on poly(T) oligonucleotides for hybridization to all poly(A) mRNAs in a sample during the RT step. The stochastic barcodes may contain universal PCR priming sites. During RT, target gene molecules react randomly with the stochastic barcodes. Each target molecule can hybridize to the stochastic barcode to generate a stochastically barcoded complementary ribonucleotide acid (cDNA) molecule. After labeling, the stochastically barcoded cDNA molecules from the microwells of a microwell plate can be pooled into a single tube for PCR amplification and sequencing. The raw sequence data can be analyzed to obtain the number of reads, the number of stochastic barcodes with unique molecular labels, and the number of mRNA molecules.
[0063] The method for determining the mRNA expression profile of a single cell can be performed in a massively parallel manner. For example, using the Precise™ assay, the mRNA expression profile of more than 10,000 cells can be determined simultaneously. The number of single cells analyzed per sample (e.g., hundreds or thousands of single cells) may be less than the capacity of current single-cell technologies. Pooling cells from different samples can improve the use of the capacity of current single-cell technologies, thereby reducing reagent waste and the cost of single-cell analysis. The present disclosure provides a sample indexing method for distinguishing cells from different samples to prepare cDNA libraries for cell analysis, such as single-cell analysis. Pooling cells from different samples can minimize variability in cDNA library preparation of cells from different samples, allowing for more accurate comparison of different samples.
[0064] The disclosure herein includes embodiments of a method for identifying samples. In some embodiments, the method includes: contacting each of a plurality of samples with a sample indexing composition from a plurality of sample indexing compositions, each of the plurality of samples comprising one or more cells, each comprising one or more protein targets; the sample indexing composition comprises an aptamer composition comprising an aptamer and a sample indexing oligonucleotide, wherein the aptamer is capable of specifically binding to at least one of the one or more protein targets, and the sample indexing oligonucleotide comprises a sample indexing sequence, and the sample indexing sequences of at least two of the plurality of sample indexing compositions comprise different sequences; barcoding the sample indexing oligonucleotides using a plurality of barcodes to generate a plurality of barcoded sample indexing oligonucleotides; obtaining sequencing data of the plurality of barcoded sample indexing oligonucleotides; and identifying the sample origin of at least one cell from the one or more cells based on the sample indexing sequence of at least one barcoded sample indexing oligonucleotide from the plurality of barcoded sample indexing oligonucleotides.
[0065] Disclosed herein are embodiments of a plurality of sample indexing compositions. In some embodiments, each of the plurality of sample indexing compositions comprises an aptamer composition comprising a first cellular component-binding aptamer and a sample indexing oligonucleotide, wherein the cellular component-binding aptamer is capable of specifically binding to at least one cellular component target, and the sample indexing oligonucleotide comprises a sample indexing sequence for identifying the sample origin of one or more cells of the sample, and the sample indexing sequences of at least two sample indexing compositions of the plurality of sample indexing compositions comprise different sequences.
[0066] Disclosed herein are embodiments of a method for measuring protein expression in cells. In some embodiments, the method includes contacting a plurality of protein-binding aptamers with a plurality of cells comprising a plurality of protein targets, each of the plurality of protein-binding aptamers comprising an aptamer-specific oligonucleotide comprising a unique identifier sequence for the cellular component-binding aptamer, and the protein-binding aptamer being capable of specifically binding to at least one of the plurality of protein targets; distributing the plurality of cells associated with the plurality of protein-binding aptamers into a plurality of compartments, each of the plurality of compartments comprising a single cell derived from the plurality of cells associated with the protein-binding aptamers; and distributing the barcode in the compartment comprising the single cell. The method includes contacting bar-coded particles with aptamer-specific oligonucleotides, wherein the bar-coded particles comprise a plurality of oligonucleotide probes, each comprising a target-binding region and a barcode sequence selected from a diverse set of unique barcode sequences; extending the oligonucleotide probes hybridized to the aptamer-specific oligonucleotides to produce a plurality of labeled nucleic acids, wherein each of the labeled nucleic acids comprises a unique identifier sequence or its complementary sequence and a barcode sequence; and obtaining sequence information of the plurality of labeled nucleic acids or portions thereof to determine the amount of one or more of the plurality of protein targets in one or more of the plurality of cells.
[0067] The disclosure herein includes the embodiment of a method for measuring cellular component expression in cells.In some embodiments, this method includes: contacting a plurality of cellular component binding aptamers with a plurality of cells, comprising a plurality of cellular component targets, wherein each of the plurality of cellular component binding aptamers comprises an aptamer-specific oligonucleotide comprising a unique identifier sequence for the cellular component binding aptamer, and the cellular component binding aptamer can specifically bind to at least one of the plurality of cellular component targets; extending the oligonucleotide probe hybridized to the aptamer-specific oligonucleotide to produce a plurality of labeled nucleic acids, wherein each of the labeled nucleic acids comprises a unique identifier sequence or its complementary sequence and a barcode sequence; and obtaining the sequence information of the plurality of labeled nucleic acids or a portion thereof, and determining the amount of one or more of the plurality of cellular component targets in one or more of the plurality of cells.
[0068] Disclosed herein is a composition comprising a plurality of cellular component-binding aptamers, each of the plurality of cellular component-binding aptamers comprising an aptamer-specific oligonucleotide comprising a unique identifier sequence for the cellular component-binding aptamer, wherein the cellular component-binding aptamer is capable of specifically binding to at least one of the plurality of cellular component targets.
[0069] definition Unless otherwise defined, technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this disclosure belongs.See, for example, Singleton et al., Dictionary of Microbiology and Molecular Biology 2nd ed., J. Wiley & Sons (New York, NY 1994); Sambrook et al., Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Press (Cold Spring Harbor, NY 1989).For the purposes of this disclosure, the following terms are defined below.
[0070] As used herein, the term "adapter" can refer to a sequence that facilitates amplification or sequencing of an associated nucleic acid. The associated nucleic acid may comprise a target nucleic acid. The associated nucleic acid may comprise one or more of a spatial label, a target label, a sample label, an indexing label, or a barcode sequence (e.g., a molecular label). The adapter may be linear. The adapter may be a pre-adenylated adapter. The adapter may be double-stranded or single-stranded. One or more adapters may be located at the 5' or 3' end of the nucleic acid. When the adapter comprises a known sequence at the 5' and 3' ends, the known sequences may be the same or different sequences. The adapter located at the 5' and / or 3' end of the polynucleotide may be capable of hybridizing to one or more oligonucleotides immobilized on a surface. In some embodiments, the adapter may comprise a universal sequence. The universal sequence may be a region of nucleotide sequence common to two or more nucleic acid molecules. Alternatively, the two or more nucleic acid molecules may have regions of different sequences. Thus, for example, the 5' adapters may contain identical and / or universal nucleic acid sequences, and the 3' adapters may contain identical and / or universal sequences. Universal sequences that may be present in different members of a plurality of nucleic acid molecules can enable the replication or amplification of multiple different sequences using a single universal primer complementary to the universal sequence. Similarly, at least one, two (e.g., a pair), or more universal sequences that may be present in different members of a collection of nucleic acid molecules can enable the replication or amplification of multiple different sequences using at least one, two (e.g., a pair), or more single universal primers complementary to the universal sequences. Thus, a universal primer comprises a sequence that can hybridize to such a universal sequence.The target nucleic acid sequence-carrying molecule can be modified to attach a universal adaptor (e.g., a non-target nucleic acid sequence) to one or both ends of different target nucleic acid sequences. One or more universal primers attached to the target nucleic acid can provide a site for hybridization of the universal primer. The one or more universal primers attached to the target nucleic acid can be the same or different from each other.
[0071] As used herein, an antibody may be a full-length (e.g., naturally occurring or formed by conventional immunoglobulin gene fragment recombination processes) immunoglobulin molecule (e.g., an IgG antibody) or an immunologically active (i.e., specifically binding) portion of an immunoglobulin molecule such as an antibody fragment.
[0072] In some embodiments, the antibody is a functional antibody fragment. For example, the antibody fragment may be a portion of an antibody, such as F(ab')2, Fab', Fab, Fv, and sFv. An antibody fragment can bind to the same antigen recognized by the full-length antibody. An antibody fragment can include isolated fragments consisting of the variable regions of an antibody, such as an "Fv" fragment consisting of the variable regions of the heavy and light chains, and a recombinant single-chain polypeptide molecule in which the variable regions of the light and heavy chains are connected by a peptide linker ("scFv protein"). Exemplary antibodies include, but are not limited to, antibodies against cancer cells, antibodies against viruses, antibodies that bind to cell surface receptors (e.g., CD8, CD34, and CD45), and therapeutic antibodies.
[0073] As used herein, the terms "associated" or "associated with" can mean that two or more species are identifiable as co-localized at a time. Association can mean that two or more species are or were present in the same container. Association can also be informatic association. For example, digital information about two or more species can be stored and used to determine that one or more of the species were co-localized at a time. Association can also be physical association. In some embodiments, two or more associated species are "tethered," "attached," or "immobilized" to each other or to a common solid or semi-solid surface. Association can refer to covalent or non-covalent means for attaching a label to a solid or semi-solid support, such as a bead. Association can also be a covalent bond between a target and a label. Association can include hybridization between two molecules (such as between a target molecule and a label).
[0074] As used herein, the term "complementary" may refer to the ability for precise pairing between two nucleotides. For example, if a nucleotide at a given position in a nucleic acid is capable of hydrogen bonding with a nucleotide in another nucleic acid, the two nucleic acids are considered complementary to each other at that position. Complementarity between two single-stranded nucleic acid molecules may be "partial," in which only a portion of the nucleotides bind, or may be complete, in which complete complementarity exists between the single-stranded molecules. A first nucleotide sequence can be said to be the "complement" of a second sequence if the first nucleotide sequence is complementary to the second nucleotide sequence. A first nucleotide sequence can be said to be the "reverse complement" of a second sequence if the first nucleotide sequence is complementary to a sequence that is the reverse of the second nucleotide sequence (i.e., the order of the nucleotides is reversed). As used herein, the terms "complement," "complementary," and "reverse complement" can be used interchangeably. In the present disclosure, it is understood that if a molecule is capable of hybridizing to another molecule, it may be the complement of the hybridizing molecule.
[0075] As used herein, the term "digital counting" may refer to a method for estimating the number of target molecules in a sample. Digital counting may include determining the number of unique labels associated with targets in a sample. This methodology may be probabilistic in nature, and converts the problem of molecule counting from the problem of locating and identifying identical molecules into a series of digital yes / no questions regarding the detection of a set of predefined labels.
[0076] As used herein, the term "label" or "labels" can refer to a nucleic acid code associated with a target in a sample. A label may be, for example, a nucleic acid label. A label may be a fully or partially amplifiable label. A label may be a fully or partially sequenceable label. A label may be a portion of a naturally occurring nucleic acid that is distinct and identifiable. A label may be a known sequence. A label may comprise a junction of a nucleic acid sequence, e.g., a junction of a naturally occurring sequence and a non-naturally occurring sequence. As used herein, the term "label" can be used interchangeably with the terms "index," "tag," or "label tag." A label can convey information. For example, in various embodiments, a label can be used to determine the identity of a sample, the source of a sample, the identity of a cell, and / or a target.
[0077] As used herein, the term "non-depletion reservoir" may refer to a pool of barcodes (e.g., stochastic barcodes) composed of many different labels. The non-depletion reservoir may contain many different barcodes, so that when the non-depletion reservoir is associated with a pool of targets, each target is likely to be associated with a unique barcode. The uniqueness of each labeled target molecule can be determined by random selection statistics and depends on the copy number of the same target molecule in the collection compared to the diversity of the labels. The size of the resulting set of labeled target molecules can be determined by the stochastic nature of the barcoding process, and analyzing the number of barcodes detected allows the number of target molecules present in the original collection or sample to be calculated. If the ratio of the copy number of target molecules present to the number of unique barcodes is low, the labeled target molecules are highly unique (i.e., it is highly unlikely that more than one target molecule will be labeled with a given label).
[0078] As used herein, the term "nucleic acid" refers to a polynucleotide sequence or a fragment thereof. A nucleic acid can comprise nucleotides. A nucleic acid can be exogenous or endogenous to a cell. A nucleic acid can be present in a cell-free environment. A nucleic acid can be a gene or a fragment thereof. A nucleic acid can be DNA. A nucleic acid can be RNA. A nucleic acid can comprise one or more analogs (e.g., modified backbones, sugars, or nucleobases). Some non-limiting examples of analogs include 5-bromouracil, peptide nucleic acid, xenonucleic acid, morpholino, locked nucleic acid, glycol nucleic acid, threose nucleic acid, dideoxynucleotide, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein linked to the sugar), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queusine, and wyosine. "Nucleic acid," "polynucleotide," "target polynucleotide," and "target nucleic acid" can be used interchangeably.
[0079] Nucleic acids may contain one or more modifications (e.g., base modifications, backbone modifications) to provide the nucleic acid with new or enhanced characteristics (e.g., improved stability). Nucleic acids may contain nucleic acid affinity tags. Nucleosides may be base-sugar combinations. The base portion of a nucleoside may be a heterocyclic base. The two most common types of such heterocyclic bases are purines and pyrimidines. A nucleotide may be a nucleoside further comprising a phosphate group covalently linked to the sugar portion of the nucleoside. In the case of nucleosides containing a pentofuranosyl sugar, the phosphate group may be linked to the 2', 3', or 5' hydroxyl moiety of the sugar. When forming nucleic acids, the phosphate groups can covalently link adjacent nucleosides to each other to form a linear polymeric compound. The respective ends of this linear polymeric compound may then be further joined to form a circular compound, although linear compounds are generally preferred. In addition, linear compounds may have internal nucleotide base complementarity and therefore may fold to produce fully or partially double-stranded compounds. Within nucleic acids, the phosphate groups are sometimes commonly referred to as forming the internucleoside backbone of the nucleic acid. The linkage or backbone may be a 3'-5' phosphodiester linkage.
[0080] Nucleic acids can include modified backbones and / or modified internucleoside linkages. Modified backbones can include those that retain a phosphorus atom in the backbone and those that do not have a phosphorus atom in the backbone. Suitable modified nucleic acid backbones containing phosphorus atoms include, for example, phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotriesters, aminoalkylphosphotriesters, 3'-alkylenephosphonates, 5'-alkylenephosphonates, methyl and other alkylphosphonates such as chiral phosphonates, phosphinates, phosphoramidates including 3'-aminophosphoramidates and aminoalkylphosphoramidates, phosphorodiamidates, thionophosphoramidates, thionoalkylphosphonates, thionoalkylphosphotriesters, selenophosphates, and boranophosphates with normal 3'-5' linkages, 2'-5' linkage analogs, and those with reverse polarity, where one or more internucleotide linkages are 3'-3', 5'-5', or 2'-2' linkages.
[0081] Nucleic acids can include polynucleotide backbones formed by short chain alkyl or cycloalkyl internucleoside linkages, mixed heteroatom and alkyl or cycloalkyl internucleoside linkages, or one or more short chain heteroatom or heterocyclic internucleoside linkages, including those with morpholino linkages (formed in part from the sugar portion of the nucleoside), siloxane backbones, sulfide, sulfoxide, and sulfone backbones, formacetyl and thioformacetyl backbones, methyleneformacetyl and thioformacetyl backbones, riboacetyl backbones, alkene-containing backbones, sulfamate backbones, methyleneimino and methylenehydrazino backbones, sulfonate and sulfonamide backbones, amide backbones, and others with mixed N, O, S, and CH moieties.
[0082] Nucleic acids can include nucleic acid mimetics. The term "mimetics" is intended to include polynucleotides in which only the furanose ring or both the furanose ring and the internucleotide linkage are replaced with non-furanose groups; replacement of only the furanose ring is sometimes referred to as a sugar surrogate. The heterocyclic base moiety or modified heterocyclic base moiety may be maintained for hybridization with an appropriate target nucleic acid. One such nucleic acid may be a peptide nucleic acid (PNA). In a PNA, the sugar backbone of a polynucleotide may be replaced with an amide-containing backbone, particularly an aminoethylglycine backbone. Nucleotides may be retained or linked directly or indirectly to the aza nitrogen atoms of the amide portion of the backbone. The backbone of a PNA compound may contain two or more linked aminoethylglycine units, giving the PNA an amide-containing backbone. The heterocyclic base moiety may be linked directly or indirectly to the aza nitrogen atoms of the amide portion of the backbone.
[0083] The nucleic acid may comprise a morpholino backbone structure. For example, the nucleic acid may comprise a six-membered morpholino ring instead of a ribose ring. In some of these embodiments, phosphorodiamidates or other non-phosphodiester internucleoside linkages may replace the phosphodiester linkages.
[0084] Nucleic acids may contain linked morpholino units (e.g., morpholino nucleic acids) with heterocyclic bases attached to a morpholino ring. Linking groups can link the morpholino monomer units of a morpholino nucleic acid. Nonionic morpholino-based oligomeric compounds may have fewer undesirable interactions with cellular proteins. Morpholino-based polynucleotides may be nonionic mimics of nucleic acids. Different linking groups can be used to join various compounds within the morpholino class. A further type of polynucleotide mimic may be called cyclohexenyl nucleic acid (CeNA). The furanose ring normally present in nucleic acid molecules can be replaced with a cyclohexenyl ring. CeNA DMT-protected phosphoramidite monomers can be prepared and used in oligomeric compound synthesis using phosphoramidite chemistry. Incorporation of CeNA monomers into nucleic acid chains can increase the stability of DNA / RNA hybrids. CeNA oligoadenylates can form complexes with nucleic acid complements with stability similar to that of the native complex. Further modifications include locked nucleic acids (LNAs), in which a 2'-hydroxyl group is linked to the 4' carbon atom of the sugar ring, thereby forming a 2'-C, 4'-C oxymethylene linkage, thereby forming a bicyclic sugar moiety. The linkage may be a methylene (-CH2) group (where n is 1 or 2) bridging the 2' oxygen atom and the 4' carbon atom. LNAs and LNA analogs can exhibit very high duplex thermal stability with complementary nucleic acids (Tm = +3 to +10°C), stability against 3'-exonuclease degradation, and good solubility properties.
[0085] Nucleic acids may also contain nucleobase (often simply referred to as "base") modifications or substitutions. As used herein, "unmodified" or "natural" nucleobases can include purine bases (e.g., adenine (A) and guanine (G)) and pyrimidine bases (e.g., thymine (T), cytosine (C), and uracil (U)). Modified nucleobases include 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyl (-C=C-CH3) uracil and cytosine, and other alkynyl derivatives of pyrimidine bases, 6-azouracil, cytosine and thymine, 5-uracil (pso- Other synthetic and natural nucleobases include uracil, 4-uracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl, and other 8-substituted adenines and guanines, 5-halo, particularly 5-bromo, 5-trifluoromethyl, and other 5-substituted uracils and cytosines, 7-methylguanine and 7-methyladenine, 2-F-adenine, 2-aminoadenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine, and 3-deazaguanine and 3-deazaadenine.Modified nucleobases include tricyclic pyrimidines such as phenoxazine cytidine (1H-pyrimido(5,4-b)(1,4)benzoxazin-2(3H)-one), phenothiazine cytidine (1H-pyrimido(5,4-b)(1,4)benzothiazin-2(3H)-one), substituted phenoxazine cytidines (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzoxazin-2(3H)-one), phenothiazine cytidine (1H-pyrimido(5, Examples of G-clamps include G-clamps such as 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzothiazin-2(3H)-one), substituted phenoxazine cytidines (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzoxazin-2(3H)-one), carbazole cytidine (2H-pyrimido(4,5-b)indol-2-one), and pyridoindole cytidine (H-pyrido(3',2':4,5)pyrrolo[2,3-d]pyrimidin-2-one).
[0086] As used herein, the term "sample" can refer to a composition that contains a target. Samples suitable for analysis by the methods, devices, and systems of the present disclosure include cells, tissues, organs, or organisms.
[0087] As used herein, the term "sample collection device" or "device" can refer to a device capable of collecting a section of a sample and / or depositing the section on a substrate. A sample device can refer to, for example, a fluorescence activated cell sorting (FACS) machine, a cell sorting machine, a biopsy needle, a biopsy device, a tissue sectioning device, a microfluidic device, a blade grid, and / or a microtome.
[0088] As used herein, the term "solid support" can refer to a discrete solid or semi-solid surface to which multiple barcodes (e.g., stochastic barcodes) can be attached. A solid support can encompass any type of solid, porous, or hollow sphere, ball, bearing, cylinder, or other similar configuration composed of plastic, ceramic, metal, or polymeric material (e.g., hydrogel) to which nucleic acids can be immobilized (e.g., covalently or non-covalently). A solid support can be spherical (e.g., microsphere) or can include discrete particles that can have non-spherical or irregular shapes, such as cubes, cuboids, pyramidal, cylindrical, conical, rectangular, or discoidal shapes. Beads can be non-spherical in shape. A plurality of solid supports spaced apart in an array may not constitute a substrate. A solid support may be used synonymously with the term "beads."
[0089] As used herein, the term "stochastic barcode" can refer to a polynucleotide sequence comprising a label of the present disclosure. A stochastic barcode may be a polynucleotide sequence that can be used for stochastic barcoding. A stochastic barcode can be used to quantify a target in a sample. A stochastic barcode can be used to control errors that may occur after a label is attached to a target. For example, a stochastic barcode can be used to evaluate amplification errors or sequencing errors. A stochastic barcode attached to a target may be referred to as a stochastic barcode-target or a stochastic barcode-tag-target. As used herein, the term "gene-specific stochastic barcode" can refer to a polynucleotide sequence that includes a label and a gene-specific target-binding region. A stochastic barcode can be a polynucleotide sequence that can be used for stochastic barcoding. A stochastic barcode can be used to quantify a target in a sample. A stochastic barcode can be used to control errors that may occur after a label is attached to a target. For example, a stochastic barcode can be used to evaluate amplification errors or sequencing errors. A stochastic barcode attached to a target may be referred to as a stochastic barcode-target or a stochastic barcode-tag-target.
[0090] As used herein, the term "probabilistic barcoding" can refer to random labeling (e.g., barcoding) of nucleic acids. Probabilistic barcoding can use a recursive Poisson strategy to associate and quantify the labels associated with targets. As used herein, the term "probabilistic barcoding" can be used synonymously with "probabilistic labeling." As used herein, the term "target" can refer to a composition that can be associated with a barcode (e.g., a stochastic barcode). Exemplary suitable targets for analysis by the methods, devices, and systems of the present disclosure include oligonucleotides, DNA, RNA, mRNA, microRNA, tRNA, and the like. Targets can be single-stranded or double-stranded. In some embodiments, targets can be proteins, peptides, or polypeptides. In some embodiments, targets are lipids. As used herein, "target" can be used synonymously with "species."
[0091] As used herein, the term "reverse transcriptase" can refer to a group of enzymes that have reverse transcriptase activity (i.e., catalyze the synthesis of DNA from an RNA template). Generally, such enzymes include, but are not limited to, retroviral reverse transcriptases, retrotransposon reverse transcriptases, retroplasmid reverse transcriptases, retron reverse transcriptases, bacterial reverse transcriptases, group II intron-derived reverse transcriptases, and mutants, variants, or derivatives thereof. Non-retroviral reverse transcriptases include non-LTR retrotransposon reverse transcriptases, retroplasmid reverse transcriptases, retron reverse transcriptases, and group II intron reverse transcriptases. Examples of group II intron reverse transcriptases include the Lactococcus lactis LI.LtrB intron reverse transcriptase, the Thermosynechococcus elongatus TeI4c intron reverse transcriptase, or the Geobacillus stearothermophilus GsI-IIC intron reverse transcriptase. Other classes of reverse transcriptases include many classes of non-retroviral reverse transcriptases (i.e., retrons, group II introns, and diversity-generating retroelements, among others).
[0092] The terms "universal adapter primer," "universal primer adapter," or "universal adapter sequence" are used interchangeably to refer to a nucleotide sequence that can hybridize with a barcode (e.g., a stochastic barcode) and be used to generate a gene-specific barcode. The universal adapter sequence may be, for example, a known sequence that is universal across all barcodes used in the methods of the present disclosure. For example, when multiple targets are labeled using the methods disclosed herein, each of the target-specific sequences may be linked to the same universal adapter sequence. In some embodiments, more than one universal adapter sequence can be used in the methods disclosed herein. For example, when multiple targets are labeled using the methods disclosed herein, at least two of the target-specific sequences are linked to different universal adapter sequences. The universal adapter primer and its complement may be contained in two oligonucleotides, one of which contains the target-specific sequence and the other of which contains the barcode. For example, the universal adapter sequence may be part of an oligonucleotide that contains a target-specific sequence to generate a nucleotide sequence complementary to a target nucleic acid. A second oligonucleotide comprising the barcode and the complementary sequence of the universal adapter sequence can hybridize with the nucleotide sequence to generate a target-specific barcode (e.g., a target-specific stochastic barcode). In some embodiments, the universal adapter primer has a different sequence than the universal PCR primer used in the disclosed methods.
[0093] Barcode Barcoding, such as probabilistic barcoding, is described, for example, in U.S. Patent Application Publication No. 20150299784, WO 2015031691, and Fu et al, Proc Natl Acad Sci USA 2011 May 31; 108(22):9026-31, the contents of which are incorporated herein in their entireties. In some embodiments, the barcodes disclosed herein may be probabilistic barcodes, which may be polynucleotide sequences that can be used to stochastically label (e.g., barcode, tag) targets. The barcodes are stochastic barcodes in which the ratio of the number of different barcode sequences in the stochastic barcode to the number of occurrences of any of the targets to be labeled is 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, or about 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, or a value or range between any two of these values, which may be referred to as a stochastic barcode. The target may be an mRNA species that includes mRNA molecules with identical or nearly identical sequences.The barcodes are stochastic barcodes in which the ratio of the number of distinct barcode sequences in the stochastic barcode to the number of occurrences of any of the targets to be labeled is at least 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, 110:1, 120:1, 130:1, 140:1, 150:1, 160:1, 170:1, 180:1, 190:1, 210:1, 220:1, 230:1, 240:1, 250:1, 260:1, 270:1, 280:1, 290:1, 300:1, 310:1, 320:1, 330:1, 340:1, 350:1, 360:1, 370:1, 380:1, 390:1, 410:1, 420:1, 430:1, 440:1, 450:1, 460:1, 470:1, 480:1, 490:1, 510:1, 520:1, 530:1, 540:1, 550:1, 560:1, A stochastic barcode can be referred to as a stochastic barcode if it is 0:1 or 100:1, or at most 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, or 100:1. The barcode sequence of a stochastic barcode may be referred to as a molecular beacon.
[0094] A barcode, e.g., a stochastic barcode, can include one or more labels. Exemplary labels can include a universal label, a cell label, a barcode sequence (e.g., a molecular label), a sample label, a plate label, a spatial label, and / or a pre-spatial label. FIG. 1 shows an exemplary barcode 104 having a spatial label. The barcode 104 can include a 5' amine that can link the barcode to a solid support 105. The barcode can include a universal label, a dimension label, a spatial label, a cell label, and / or a molecular label. The order of the various labels (including, but not limited to, the universal label, the dimension label, the spatial label, the cell label, and the molecular label) within the barcode can vary. For example, as shown in FIG. 1, the universal label can be the 5'-most label and the molecular label can be the 3'-most label. The spatial label, the dimension label, and the cell label can be in any order. In some embodiments, the universal label, the spatial label, the dimension label, the cell label, and the molecular label are in any order. The barcode can include a target binding region. The target binding region can interact with a target (e.g., target nucleic acid, RNA, mRNA, DNA) in a sample. For example, the target binding region can include an oligo(dT) sequence that can interact with the poly(A) tail of mRNA. In some cases, the labels of the barcode (e.g., universal label, dimension label, spatial label, cell label, and barcode sequence) can be separated by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more nucleotides.
[0095] A label, e.g., a cellular label, may comprise a unique set of nucleic acid subsequences, each of a defined length, e.g., 7 nucleotides (equivalent to the number of bits used in some Hamming error-correcting codes), which may be designed to provide error correction capabilities. The set of error-correcting subsequences may comprise 7 nucleotide sequences, and may be designed so that any pairwise combination of sequences within the set represents a defined "genetic distance" (or number of mismatched bases); for example, a set of error-correcting subsequences may be designed to represent a genetic distance of 3 nucleotides. In this case, by inspecting the error-correcting sequences (described more fully below) of a set of sequence data for a labeled target nucleic acid molecule, it may be possible to detect or correct amplification or sequencing errors. In some embodiments, the length of the nucleic acid subsequences used to create the error-correcting code can vary, for example, can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 31, 40, 50 nucleotides in length, or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 31, 40, 50 nucleotides in length, or a number or range of nucleotides between any two of these values. In some embodiments, nucleic acid subsequences of other lengths can be used to create the error-correcting code.
[0096] The barcode may include a target-binding region. The target-binding region can interact with a target in a sample. The target may be or include ribonucleic acid (RNA), messenger RNA (mRNA), microRNA, small interfering RNA (siRNA), RNA degradation products, RNAs each containing a poly(A) tail, or any combination thereof. In some embodiments, the multiple targets may include deoxyribonucleic acid (DNA).
[0097] In some embodiments, the target binding region may include an oligo(dT) sequence that can interact with the poly(A) tail of mRNA. One or more of the labels of the barcode (e.g., universal label, dimensional label, spatial label, cellular label, and barcode sequence (e.g., molecular label)) may be separated from one or two of the remaining labels of the barcode by a spacer. The spacer may be, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more nucleotides. In some embodiments, none of the labels of the barcode are separated by a spacer.
[0098] Universal Signage A barcode may include one or more universal labels. In some embodiments, the one or more universal labels may be the same for all barcodes in a set of barcodes attached to a given solid support. In some embodiments, the one or more universal labels may be the same for all barcodes attached to a plurality of beads. In some embodiments, the universal label may include a nucleic acid sequence capable of hybridizing to a sequencing primer. The sequencing primer may be used to sequence the barcode comprising the universal label. The sequencing primer (e.g., a universal sequencing primer) may include a sequencing primer associated with a high-throughput sequencing platform. In some embodiments, the universal label may include a nucleic acid sequence capable of hybridizing to a PCR primer. In some embodiments, the universal label may include a nucleic acid sequence capable of hybridizing to a sequencing primer and a PCR primer. The nucleic acid sequence of a universal label capable of hybridizing to a sequencing primer or a PCR primer may be referred to as a primer binding site. The universal label may include a sequence that can be used to initiate transcription of the barcode. The universal label may include a sequence that can be used to extend the barcode or a region within the barcode. A universal label may be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or a number or range of nucleotides between any two of these values. For example, a universal label may comprise at least about 10 nucleotides. A universal label may be at least 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length, or may be at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length.In some embodiments, the cleavable linker or modified nucleotide may be part of the universal label sequence to allow the barcode to be cleaved from the support.
[0099] Dimension Indicators A barcode may include one or more dimensional labels. In some embodiments, the dimensional labels may include nucleic acid sequences that provide information about the dimension in which labeling (e.g., stochastic labeling) occurred. For example, the dimensional labels can provide information about the time point at which a target was barcoded. The dimensional labels may be associated with the time point of barcoding (e.g., stochastic barcoding) within a sample. The dimensional labels can be activated at the time of labeling. Different dimensional labels can be activated at different time points. The dimensional labels provide information about the order in which targets, groups of targets, and / or samples were barcoded. For example, a population of cells can be barcoded at the G0 phase of the cell cycle. Cells can be re-pulsed with a barcode (e.g., a stochastic barcode) at the G1 phase of the cell cycle. Cells can be re-pulsed with a barcode at the S phase of the cell cycle, etc. The barcode in each pulse (e.g., each phase of the cell cycle) can include a different dimensional label. In this way, the dimensional labels provide information about which targets were labeled at which phase of the cell cycle. Dimensional labels allow for the investigation of multiple different biological time points. Exemplary biological time points can include, but are not limited to, cell cycle, transcription (e.g., transcription initiation), and transcript degradation. In another example, a sample (e.g., a cell, a population of cells) can be labeled before and / or after treatment with a drug and / or therapy. The change in copy number of distinct targets can indicate the response of the sample to the drug and / or therapy.
[0100] The dimension label may be activatable. An activatable dimension label can be activated at a specific time. An activatable label can, for example, be constitutively activated (e.g., not turned off). An activatable dimension label can, for example, be reversibly activated (e.g., an activatable dimension label can be turned on and off). For example, the dimension label may be reversibly activatable at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more times. The dimension label may be reversibly activatable at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more times. In some embodiments, the dimension label can be activated by fluorescence, light, a chemical event (e.g., cleavage, ligation of another molecule, addition of a modification (e.g., pegylation, sumoylation, acetylation, methylation, deacetylation, demethylation), a photochemical event (e.g., photocaging), and the introduction of a non-natural nucleotide.
[0101] The dimension labels, in some embodiments, may be the same for all barcodes (e.g., stochastic barcodes) attached to a given solid support (e.g., a bead), but may be different for different solid supports (e.g., beads). In some embodiments, at least 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100% of the barcodes on the same solid support may comprise the same dimension labels. In some embodiments, at least 60% of the barcodes on the same solid support may comprise the same dimension labels. In some embodiments, at least 95% of the barcodes on the same solid support may comprise the same dimension labels.
[0102] For multiple solid supports (e.g., beads), 10 6As many as one or more unique dimension label sequences may be provided. The dimension labels may be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or a number or range of nucleotides in length between any two of these values. The dimension labels may be at least 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length, or may be at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length. Dimension labels may comprise from about 5 to about 200 nucleotides. Dimension labels may comprise from about 10 to about 150 nucleotides. Dimension labels may comprise from about 20 to about 125 nucleotides in length.
[0103] spatial sign The barcode may include one or more spatial labels. In some embodiments, the spatial label may include a nucleic acid sequence that provides information about the spatial orientation of the target molecule to which the barcode is attached. The spatial label may be associated with coordinates of the sample. The coordinates may be fixed coordinates. For example, the coordinates may be fixed relative to the substrate. The spatial label may be referenced to a two-dimensional or three-dimensional grid. The coordinates may be fixed relative to a landmark. The landmark may be identifiable in space. The landmark may be a structure that can be imaged. The landmark may be a biological structure, e.g., an anatomical landmark. The landmark may be a cellular landmark, e.g., an organelle. The landmark may be a non-natural landmark, such as a color code, a barcode, a magnetic property, a fluorescent material, radioactivity, or a structure with an identifiable identifier, such as a unique size or shape. The spatial label may be associated with a physical compartment (e.g., a well, a container, or a droplet). In some embodiments, multiple spatial labels are used together to encode one or more locations in space.
[0104] The spatial labels may be the same for all barcodes attached to a given solid support (e.g., a bead), or may be different for different solid supports (e.g., beads). In some embodiments, the percentage of barcodes on the same solid support that contain the same spatial label may be at or about 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%, or a number or range between any two of these values. In some embodiments, the percentage of barcodes on the same solid support that contain the same spatial label may be at least 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100%, or may be up to 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100%. In some embodiments, at least 60% of the barcodes on the same solid support may contain the same spatial label. In some embodiments, at least 95% of the barcodes on the same solid support may contain the same spatial label.
[0105] For multiple solid supports (e.g., beads), 10 6As many as one or more unique spatial marker sequences may be provided. Spatial markers may be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or a number or range of nucleotides in length between any two of these values. Spatial markers may be at least 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length, or may be at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length. Spatial labels may contain from about 5 to about 200 nucleotides. Spatial labels may contain from about 10 to about 150 nucleotides. Spatial labels may contain from about 20 to about 125 nucleotides in length.
[0106] cell labeling A barcode (e.g., a stochastic barcode) may include one or more cellular labels. In some embodiments, the cellular label may include a nucleic acid sequence that provides information for determining which target nucleic acid originated from which cell. In some embodiments, the cellular label is the same for all barcodes attached to a given solid support (e.g., a bead), but is different for different solid supports (e.g., beads). In some embodiments, the percentage of barcodes on the same solid support that include the same cellular label may be at or about 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%, or a number or range between any two of these values. In some embodiments, the percentage of barcodes on the same solid support that contain the same cell label may be or may be about 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100%. For example, at least 60% of the barcodes on the same solid support may contain the same cell label. As another example, at least 95% of the barcodes on the same solid support may contain the same cell label.
[0107] For multiple solid supports (e.g., beads), 10 6As many as one or more unique cell marker sequences may be provided. The cell marker may be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or a number or range of nucleotides between any two of these values. The cell marker may be at least 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length, or may be at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length. For example, a cellular label may comprise from about 5 to about 200 nucleotides. As another example, a cellular label may comprise from about 10 to about 150 nucleotides. As yet another example, a cellular label may comprise from about 20 to about 125 nucleotides in length.
[0108] Barcode sequence The barcode may include one or more barcode sequences. In some embodiments, the barcode sequence may include a nucleic acid sequence that provides information for identifying a particular type of target nucleic acid species hybridized to the barcode. The barcode sequence may also include a nucleic acid sequence that provides a means for counting (e.g., provides a rough approximation of) the specific occurrence of a target nucleic acid species hybridized to the barcode (e.g., target binding region).
[0109] In some embodiments, a diverse set of barcode sequences is attached to a given solid support (e.g., a bead). 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 pieces, or about 10 2 , 10 3 , 10 4 , 10 5, 10 6 , 10 7 , 10 8 , 10 9 There may be at least 10 unique molecular label sequences, or a number or range between any two of these values. For example, the plurality of barcodes may include about 6561 barcode sequences having distinct sequences. As another example, the plurality of barcodes may include about 65536 barcode sequences having distinct sequences. In some embodiments, at least 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , or 10 9 of, or at most 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , or 10 9 There may be unique barcode sequences. The unique molecular label sequences may be attached to a given solid support (e.g., a bead).
[0110] The length of a barcode may vary in different implementations. For example, a barcode may be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or a number or range of nucleotides between any two of these values. As another example, a barcode may be at least 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length, or may be at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length.
[0111] molecular label A barcode (e.g., a stochastic barcode) may include one or more molecular labels. A molecular label may include a barcode sequence. In some embodiments, a molecular label may include a nucleic acid sequence that provides information for identifying a particular type of target nucleic acid species hybridized to the barcode. A molecular target may include a nucleic acid sequence that provides a means for counting specific occurrences of a target nucleic acid species hybridized to the barcode (e.g., a target binding region).
[0112] In some embodiments, a diverse set of molecular labels is attached to a given solid support (e.g., beads). 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 pieces, or about 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 There may be at least 10 unique molecular label sequences, or a number or range between any two of these values. For example, the plurality of barcodes may include about 6561 molecular labels with distinct sequences. As another example, the plurality of barcodes may include about 65536 molecular labels with distinct sequences. In some embodiments, at least 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , or 10 9 of, or at most 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , or 10 9There may be a number of unique molecular tag sequences. A barcode with a unique molecular tag sequence may be attached to a given solid support (e.g., a bead).
[0113] In the case of stochastic barcoding using multiple stochastic barcodes, the ratio of the number of different molecular label sequences and the number of occurrences of any of the targets is 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1 , 90:1, 100:1, or about 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, or a value or range between any two of these values. The target may be an mRNA species comprising mRNA molecules with identical or nearly identical sequences. In some embodiments, the ratio of the number of different molecular beacon sequences and the number of occurrences of any of the targets is at least 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1 , 90:1, or 100:1, or at most 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, or 100:1.
[0114] Molecular labels may be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or a number or range of nucleotides between any two of these values. Molecular labels may be at least 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length, or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length.
[0115] Target binding region The barcode may include one or more target-binding regions, such as a capture probe. In some embodiments, the target-binding region may hybridize to a target of interest. In some embodiments, the target-binding region may include a nucleic acid sequence that specifically hybridizes to a target (e.g., a target nucleic acid, a target molecule, e.g., a cellular nucleic acid to be analyzed), e.g., to a specific gene sequence. In some embodiments, the target-binding region may include a nucleic acid sequence that can attach (e.g., hybridize) to a specific position of a specific target nucleic acid. In some embodiments, the target-binding region may include a nucleic acid sequence that is capable of specific hybridization to a restriction enzyme site overhang (e.g., an EcoRI sticky end overhang). The barcode can then be ligated to any nucleic acid molecule that includes a sequence complementary to the restriction site overhang.
[0116] In some embodiments, the target binding region may include a non-specific target nucleic acid sequence. A non-specific target nucleic acid sequence may refer to a sequence that can bind to multiple target nucleic acids regardless of the specific sequence of the target nucleic acid. For example, the target binding region may include a random multimer sequence or an oligo(dT) sequence that hybridizes to a poly(A) tail on an mRNA molecule. The random multimer sequence may be, for example, a random dimer, trimer, tetramer, pentamer, hexamer, heptamer, octamer, nonamer, decamer, or higher-order multimer sequence of any length. In some embodiments, the target binding region is the same for all barcodes attached to a given bead. In some embodiments, the target binding regions of multiple barcodes attached to a given bead may include two or more different target binding sequences. A target binding region may be 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or a number or range of nucleotides between any two of these values. A target binding region may be up to about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or longer.
[0117] In some embodiments, the target binding region may comprise an oligo(dT) capable of hybridizing to an mRNA containing a polyadenylated end. The target binding region may be gene-specific. For example, the target binding region may be configured to hybridize to a specific region of the target. A target binding region can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26 27, 28, 29, 30 nucleotides in length, or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26 27, 28, 29, 30 nucleotides in length, or a number or range of nucleotides in length between any two of these values. The target binding region may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides in length, or may be at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides in length. The target binding region may be about 5-30 nucleotides in length. When a barcode includes a gene-specific target binding region, the barcode may be referred to herein as a gene-specific barcode.
[0118] Orientation Characteristics A stochastic barcode (e.g., a stochastic barcode) may include one or more orientation features that can be used to orient (e.g., align) the barcode. The barcode may include a moiety for isoelectric focusing. Different barcodes can constitute different isoelectric focusing points. When such barcodes are introduced into a sample, the sample can be subjected to isoelectric focusing to orient the barcodes in a known direction. In this manner, the orientation feature can be used to create a known map of the barcodes within the sample. Exemplary orientation features can include electrophoretic mobility (e.g., based on the size of the barcode), isoelectric point, spin, conductivity, and / or self-assembly. For example, a barcode with a self-assembly orientation feature can self-assemble into a particular direction (e.g., a nucleic acid nanostructure) when activated.
[0119] affinity properties A barcode (e.g., a stochastic barcode) may include one or more affinity features. For example, a spatial label may include an affinity feature. The affinity feature may include a chemical and / or biological moiety that can facilitate binding of the barcode to another entity (e.g., a cellular receptor). For example, the affinity feature may include an antibody, e.g., an antibody specific for a particular moiety (e.g., a receptor) on a sample. In some embodiments, the antibody can direct the barcode to a particular cell type or molecule. Targets on and / or near a particular cell type can be labeled (e.g., stochastically labeled). Because the antibody can direct the barcode to a specific location, in some embodiments, the affinity feature can provide spatial information in addition to the nucleotide sequence of the spatial label. The antibody may be a therapeutic antibody, e.g., a monoclonal or polyclonal antibody. The antibody may be humanized or chimerized. The antibody may be a naked antibody or a fusion antibody.
[0120] An antibody may be a full-length (i.e., naturally occurring or generated by conventional immunoglobulin gene fragment recombination processes) immunoglobulin molecule (e.g., an IgG antibody) or an immunologically active (i.e., specifically binding) portion of an immunoglobulin molecule, such as an antibody fragment. An antibody fragment may be a portion of an antibody, such as, for example, F(ab')2, Fab', Fab, Fv, and sFv. In some embodiments, an antibody fragment can bind to the same antigen recognized by the full-length antibody. Antibody fragments can include isolated fragments consisting of the variable regions of an antibody, such as an "Fv" fragment consisting of the variable regions of the heavy and light chains, and a recombinant single-chain polypeptide molecule in which the variable regions of the light and heavy chains are connected by a peptide linker ("scFv protein"). Exemplary antibodies include, but are not limited to, antibodies against cancer cells, antibodies against viruses, antibodies that bind to cell surface receptors (CD8, CD34, CD45), and therapeutic antibodies.
[0121] Universal Adapter Primer A barcode may include one or more universal adapter primers. For example, a gene-specific barcode, such as a gene-specific stochastic barcode, may include a universal adapter primer. A universal adapter primer can refer to a nucleotide sequence that is universal across all barcodes. A universal adapter primer can be used to construct a gene-specific barcode. Universal adapter primers can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26 27, 28, 29, 30 nucleotides in length, or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26 27, 28, 29, 30 nucleotides in length, or a number or range of nucleotides in length between any two of these. The universal adapter primer may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides in length, or may be at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides in length. The universal adapter primer may be 5 to 50 nucleotides in length.
[0122] Linker When a barcode contains more than one type of label (e.g., more than one cellular label or more than one barcode sequence, such as one molecular label), the labels may be dispersed with a linker label sequence. The linker label sequence may be at least about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, or more nucleotides in length. The linker label sequence may be at most about 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, or more nucleotides in length. In some cases, the linker label sequence is 12 nucleotides in length. The linker label sequence can be used to facilitate the synthesis of the barcode. The linker label may include an error-correcting (e.g., Hamming) code.
[0123] solid support In some embodiments, barcodes, such as the stochastic barcodes disclosed herein, may be associated with a solid support. The solid support may be, for example, a synthetic particle. In some embodiments, some or all of the barcode sequences, such as molecular labels of a stochastic barcode (e.g., a first barcode sequence) of a plurality of barcodes (e.g., a first plurality of barcodes) on a solid support, differ by at least one nucleotide. Cell labels of barcodes on the same solid support may be the same. Cell labels of barcodes on different solid supports may differ by at least one nucleotide. For example, a first cell label of a first plurality of barcodes on a first solid support may have the same sequence, and a second cell label of a second plurality of barcodes on a second solid support may have the same sequence. A first cell label of a first plurality of barcodes on a first solid support and a second cell label of a second plurality of barcodes on a second solid support may differ by at least one nucleotide. Cell labels may be, for example, about 5-20 nucleotides in length. The barcode sequence may be, for example, about 5 to 20 nucleotides in length. The synthetic particle may be, for example, a bead.
[0124] The beads may be, for example, silica gel beads, controlled pore glass beads, magnetic beads, Dynabeads, Sephadex / Sepharose beads, cellulose beads, polystyrene beads, or any combination thereof. The beads may comprise materials such as polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogels, paramagnetic materials, ceramics, plastics, glass, methylstyrene, acrylic polymers, titanium, latex, Sepharose, cellulose, nylon, silicone, or any combination thereof.
[0125] In some embodiments, the beads may be polymer beads, e.g., deformable beads or gel beads, functionalized with barcodes or stochastic barcodes (e.g., gel beads from 10X Genomics, San Francisco, CA). In some embodiments, the gel beads may comprise a polymer-based gel. Gel beads can be generated, for example, by encapsulating one or more polymer precursors within droplets. Gel beads can be generated by exposing the polymer precursors to an accelerator (e.g., tetramethylethylenediamine (TEMED)).
[0126] In some embodiments, the particles may be degradable. For example, polymer beads can be dissolved, melted, or decomposed under desired conditions. The desired conditions can include environmental conditions. The desired conditions can cause the polymer beads to dissolve, melt, or decompose in a controlled manner. Gel beads can be dissolved, melted, or decomposed by chemical, physical, biological, thermal, magnetic, electrical, or optical stimuli, or any combination thereof.
[0127] Analytes and / or reagents, such as oligonucleotide barcodes, may be coupled / immobilized, for example, to the interior surface of gel beads (e.g., the diffusion-accessible interior of the oligonucleotide barcodes and / or materials used to generate the oligonucleotide barcodes) and / or the exterior surface of gel beads or other microcapsules described herein. Coupling / immobilization may be via any form of chemical bond (e.g., covalent bond, ionic bond) or physical phenomenon (e.g., van der Waals forces, dipole-dipole interactions, etc.). In some embodiments, the coupling / immobilization of reagents to gel beads or any other microcapsules described herein may be reversible, such as, for example, via a labile moiety (e.g., via a chemical crosslinker, including those described herein). Application of a stimulus can cleave the labile moiety, releasing the immobilized reagent. In some embodiments, the labile moiety is a disulfide bond. For example, if an oligonucleotide barcode is immobilized to a gel bead via a disulfide bond, exposing the disulfide bond to a reducing agent can cleave the disulfide bond and release the oligonucleotide barcode from the bead. The labile moiety may be included as part of the gel bead or microcapsule, as part of a chemical linker connecting the reagent or analyte to the gel bead or microcapsule, and / or as part of the reagent or analyte. In some embodiments, at least one barcode of the plurality of barcodes may be immobilized on a particle, partially immobilized on a particle, encapsulated within a particle, partially encapsulated within a particle, or any combination thereof.
[0128] In some embodiments, the gel beads may comprise a wide variety of polymers, including, but not limited to, polymers, thermosensitive polymers, light-sensitive polymers, magnetic polymers, pH-sensitive polymers, salt-sensitive polymers, chemically sensitive polymers, polyelectrolytes, polysaccharides, peptides, proteins, and / or plastics. Polymers include, but are not limited to, poly(N-isopropylacrylamide) (PNIPAAm), poly(styrenesulfonate) (PSS), poly(allylamine) (PAAm), poly(acrylic acid) (PAA), poly(ethyleneimine) (PEI), poly(diallyldimethylammonium chloride) (PDADMAC), poly(pyrolle) (PPy), poly(vinylpyrrolidone) (PVPON ... Examples of materials that can be used include poly(vinyl acrylate) (PVP), poly(methacrylic acid) (PMAA), poly(methyl methacrylate) (PMMA), polystyrene (PS), poly(tetrahydrofuran) (PTHF), poly(phthaladehyde) (PTHF), poly(hexylviologen) (PHV), poly(L-lysine) (PLL), poly(L-arginine) (PARG), and poly(lactic-co-glycolic acid) (PLGA).
[0129] A number of chemical stimuli can be used to induce bead rupture, dissolution, or degradation. Examples of such chemical changes include, but are not limited to, pH-mediated changes to the bead wall, disintegration of the bead wall by chemical cleavage of cross-links, inducing bead wall depolymerization, and bead wall switching reactions. Bulk changes can also be used to induce bead rupture.
[0130] Additionally, bulk or physical changes to microcapsules upon various stimuli offer numerous advantages in the design of capsules for reagent release. Bulk or physical changes occur on a macroscopic scale, with bead rupture being the result of mechanical or physical forces induced by the stimulus. Such processes can include, but are not limited to, pressure-induced rupture, bead wall melting, or changes in bead wall porosity. Biological stimuli can also be used to induce disruption, dissolution, or degradation of beads. Generally, biological triggers are similar to chemical triggers, but many examples use biomolecules, or molecules commonly found in biological systems, such as enzymes, peptides, saccharides, fatty acids, and nucleic acids. For example, beads may contain polymers with peptide crosslinks that are susceptible to cleavage by specific proteases. More specifically, one example may include microcapsules containing GFLGK peptide crosslinks. Addition of a biological trigger, such as the protease cathepsin B, cleaves the peptide crosslinks in the shell wall, releasing the contents of the bead. In other cases, the protease can be heat-activated. In another example, beads contain a shell wall containing cellulose. Addition of the hydrolytic enzyme chitosan serves as a biological trigger for cleaving the cellulose bonds, depolymerizing the shell wall, and releasing its internal contents. Beads can also be induced to release their contents upon application of a thermal stimulus. Temperature changes can cause various changes in beads. Heat changes can cause beads to melt and collapse the bead walls. In other cases, heat can increase the internal pressure of the bead's internal components, causing the bead to rupture or explode. In still other cases, heat can deform the beads, causing them to shrink and dehydrate. Heat can also act on thermosensitive polymers within the bead walls, causing the bead to break down.
[0131] The inclusion of magnetic nanoparticles in the bead walls of the microcapsules can trigger bead bursting and allow the beads to be guided in an array. The devices of the present disclosure may include magnetic beads for any purpose. In one example, the incorporation of Fe3O4 nanoparticles into the polyelectrolyte containing beads triggers bursting in the presence of an oscillating magnetic field stimulus. Additionally, beads can be disrupted, dissolved, or decomposed as a result of electrical stimulation. Similar to the magnetic particles described in the previous section, electrically sensitive beads can both trigger bead rupture and perform other functions such as alignment in an electric field, electrical conductivity, or redox reactions. In one example, beads containing electrically sensitive materials can be aligned in an electric field to control the release of internal reagents. In another example, an electric field can induce redox reactions within the bead wall itself, thereby increasing porosity. Optical stimulation can also be used to disrupt the beads. Numerous photoinducers are possible, including systems using various molecules such as nanoparticles and chromophores that can absorb photons in specific wavelength ranges. For example, metal oxide coatings can be used as capsule inducers. UV irradiation of SiO2-coated polyelectrolyte capsules can induce the collapse of the bead wall. In yet another example, photoswitchable substances such as azobenzene groups can be incorporated into the bead wall. Upon application of UV or visible light, chemicals such as these undergo reversible cis-trans isomerization upon absorption of photons. In this embodiment, the incorporation of a photon switch results in a bead wall that can collapse or become more porous upon application of the photoinducer.
[0132] For example, in a non-limiting example of barcoding (e.g., stochastic barcoding) illustrated in FIG. 2, after cells, such as single cells, are introduced into a plurality of microwells of a microwell array in block 208, beads can be introduced into a plurality of microwells of the microwell array in block 212. Each microwell may contain one bead. The beads may contain multiple barcodes. The barcodes may comprise 5' amine regions attached to the beads. The barcodes may comprise a universal label, a barcode sequence (e.g., a molecular label), a target binding region, or any combination thereof.
[0133] The barcodes disclosed herein may be associated (e.g., attached) to a solid support (e.g., a bead). Each barcode associated with a solid support may comprise a barcode sequence selected from a group comprising at least 100 or 1000 barcode sequences having a unique sequence. In some embodiments, different barcodes associated with a solid support may comprise barcodes having different sequences. In some embodiments, a percentage of the barcodes associated with a solid support comprise the same cell marker. For example, the percentage may be 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%, or about 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%, or a number or range between any two of these values. As another example, the percentage may be at least 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100% or may be at most 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100%. In some embodiments, barcodes associated with a solid support may have the same cell marker. Barcodes associated with different solid supports may have different cell markers selected from a group including at least 100 or 1000 cell markers having unique sequences.
[0134] The barcodes disclosed herein may be associated (e.g., attached) to a solid support (e.g., a bead). In some embodiments, barcoding of multiple targets in a sample can be performed using a solid support comprising multiple synthetic particles having multiple barcodes associated therewith. In some embodiments, the solid support may comprise multiple synthetic particles having multiple barcodes associated therewith. The spatial labeling of multiple barcodes on different solid supports may differ by at least one nucleotide. The solid support may comprise multiple barcodes, for example, in two or three dimensions. The synthetic particles may be beads. The beads may be silica gel beads, controlled pore glass beads, magnetic beads, Dynabeads, Sephadex / Sepharose beads, cellulose beads, polystyrene beads, or any combination thereof. The solid support may comprise a polymer, a matrix, a hydrogel, a needle array device, an antibody, or any combination thereof. In some embodiments, the solid support may be buoyant. In some embodiments, the solid support may be embedded in a semi-solid or solid array. The barcodes may not be associated with the solid support. The barcodes may be individual nucleotides. The barcode may be associated with the substrate.
[0135] As used herein, the terms "tethered," "attached," and "immobilized" are used interchangeably and can refer to covalent or non-covalent means for attaching a barcode to a solid support. Any of a variety of different solid supports can be used as the solid support for attaching a pre-synthesized barcode or for in situ solid phase synthesis of a barcode.
[0136] In some embodiments, the solid support is a bead. The bead may comprise one or more types of solid, porous, or hollow spheres, balls, bearings, cylinders, or other similar structures capable of immobilizing nucleic acids (e.g., covalently or non-covalently). The bead may be composed of, for example, plastic, ceramic, metal, polymeric material, or any combination thereof. The bead may be spherical (e.g., microspheres) or may be or comprise individual particles having a non-spherical or irregular shape, such as a cube, cuboid, pyramidal, cylindrical, conical, rectangular, or discoid. In some embodiments, the bead may be non-spherical in shape.
[0137] The beads may comprise a variety of materials, including, but not limited to, paramagnetic materials (e.g., magnesium, molybdenum, lithium, and tantalum), superparamagnetic materials (e.g., ferrite (Fe3O4; magnetite) nanoparticles), ferromagnetic materials (e.g., iron, nickel, cobalt, some alloys thereof, and some rare earth metal compounds), ceramic, plastic, glass, polystyrene, silica, methylstyrene, acrylic polymers, titanium, latex, sepharose, agarose, hydrogels, polymers, cellulose, nylon, or any combination thereof. In some embodiments, the beads (e.g., the beads to which the labels are attached) are hydrogel beads. In some embodiments, the beads comprise a hydrogel.
[0138] Some embodiments disclosed herein include one or more particles (e.g., beads). Each of the particles may include a plurality of oligonucleotides (e.g., barcodes). Each of the plurality of oligonucleotides may include a barcode sequence (e.g., a molecular label sequence), a cell label, and a target binding region (e.g., an oligo(dT) sequence, a gene-specific sequence, a random multimer, or a combination thereof). The cell label sequence of each of the plurality of oligonucleotides may be the same. The cell label sequences of oligonucleotides on different particles may be different so that the oligonucleotides on different particles can be distinguished. The number of different cell label sequences may vary in different implementations. In some embodiments, the number of cell labeling sequences is 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 , 10 7 , 10 8 , 10 9 , or approximately 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 , 10 7 , 10 8 , 10 9 In some embodiments, the number of cell labeling sequences may be at least 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 , 10 7, 10 8 , or 10 9 or at most 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 , 10 7 , 10 8 , or 10 9 In some embodiments, up to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or more of the plurality of particles comprise oligonucleotides having the same cellular sequence. In some embodiments, up to 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, or more of the plurality of particles comprise oligonucleotides having the same cellular sequence. In some embodiments, none of the plurality of particles have the same cellular labeling sequence.
[0139] The multiple oligonucleotides on each particle may comprise different barcode sequences (e.g., molecular labels). In some embodiments, the number of barcode sequences is 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 , 10 7 , 10 8 , 10 9, or approximately 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 , 10 7 , 10 8 , 10 9 In some embodiments, the number of barcode sequences may be at least 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 , 10 7 , 10 8 , or 10 9 or at most 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 , 10 7 , 10 8 , or 10 9For example, at least 100 of the plurality of oligonucleotides may comprise different barcode sequences. As another example, in a single particle, at least 100, 500, 1000, 5000, 10000, 15000, 20000, 50000, a number or range between any two of these values, or more of the plurality of oligonucleotides may comprise different barcode sequences. Some embodiments provide a plurality of particles comprising barcodes. In some embodiments, the ratio of the occurrence (or copy or number) of the target to be labeled and the different barcode sequences may be at least 1:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:11, 1:12, 1:13, 1:14, 1:15, 1:16, 1:17, 1:18, 1:19, 1:20, 1:30, 1:40, 1:50, 1:60, 1:70, 1:80, 1:90, or more. In some embodiments, each of the plurality of oligonucleotides further comprises a sample label, a universal label, or both. The particle may be, for example, a nanoparticle or a microparticle.
[0140] The beads may vary in size. For example, the diameter of the beads may range from 0.1 micrometers to 50 micrometers. In some embodiments, the diameter of the beads may be 0.1, 0.5, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50 micrometers, or about 0.1, 0.5, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50 micrometers, or a value or range between any two of these values.
[0141] The diameter of the beads may be related to the diameter of the wells of the substrate. In some embodiments, the diameter of the beads may be 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100% or about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100% or a value or range between any two of these values longer or shorter than the diameter of the wells. The diameter of the beads may be related to the diameter of the cells (e.g., single cells confined in the wells of the substrate). In some embodiments, the diameter of the beads may be at least or up to 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% longer or shorter than the diameter of the well. The diameter of the beads may be related to the diameter of a cell (e.g., a single cell confined in a well of a substrate). In some embodiments, the diameter of the bead may be 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 150%, 200%, 250%, 300% longer or shorter than the diameter of the cell by at least about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 150%, 200%, 250%, 300%, or a number or range between any two of these values. In some embodiments, the diameter of the bead may be at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 150%, 200%, 250%, or 300% longer or shorter than the diameter of the cell by up to 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 150%, 200%, 250%, or 300%.
[0142] The beads may be attached to and / or embedded in a substrate. The beads may be attached to and / or embedded in a gel, hydrogel, polymer, and / or matrix. The spatial location of the beads within the substrate (e.g., gel, matrix, scaffold, or polymer) can be identified using spatial labels present in barcodes on the beads, which can serve as location addresses. Examples of beads include, but are not limited to, streptavidin beads, agarose beads, magnetic beads, Dynabeads®, MACS® microbeads, antibody-conjugated beads (e.g., anti-immunoglobulin microbeads), protein A-conjugated beads, protein G-conjugated beads, protein A / G-conjugated beads, protein L-conjugated beads, oligo(dT)-conjugated beads, silica beads, silica-like beads, anti-biotin microbeads, anti-fluorescent dye microbeads, and BcMag™ carboxyl-terminated magnetic beads.
[0143] The beads may be associated with (e.g., impregnated with) quantum dots or fluorescent dyes such that the beads fluoresce in one or more optical channels. The beads may be associated with iron oxide or chromium oxide to make them paramagnetic or ferromagnetic. The beads may be distinguishable. For example, the beads can be imaged using a camera. The beads may have a detectable code associated with them. For example, the beads may include a barcode. The beads may change size, for example, by swelling in an organic or inorganic solution. The beads may be hydrophobic. The beads may be hydrophilic. The beads may be biocompatible.
[0144] The solid support (e.g., beads) can be visualized. The solid support may include a visualization tag (e.g., a fluorescent dye). The solid support (e.g., beads) may be imprinted with an identifier (e.g., a number). The identifier can be visualized by imaging the beads. A solid support may comprise an insoluble, semi-soluble, or insoluble material. A solid support may be referred to as "functionalized" if it comprises a linker, scaffold, building block, or other reactive moiety attached thereto, and a solid support may be referred to as "non-functionalized" if it lacks such a reactive moiety attached thereto. A solid support may be used suspended in solution, such as in a microtiter well format, a flow-through format such as a column, or a dip stick. The solid support may comprise a membrane, paper, plastic, coated surface, flat surface, glass, slide, chip, or any combination thereof. The solid support may be in the form of a resin, gel, microsphere, or other geometric configuration. The solid support may comprise a flat support such as a silica chip, microparticle, nanoparticle, plate, array, capillary, glass fiber filter, glass surface, metal surface (steel, gold / silver, aluminum, silicon, and copper), a glass support, a plastic support, a silicon support, a chip, a filter, a membrane, a microwell plate, a slide, a multiwell plate, or a plastic material (e.g., made of polyethylene, polypropylene, polyamide, polyvinylidene difluoride), including a membrane, and / or a wafer, comb, pin, or needle (e.g., an array of pins suitable for combinatorial synthesis or analysis), or a wafer (e.g., a silicon wafer), a pitted wafer with or without a filter bottom, or a bead in an array of pits or nanoliter wells on a flat surface.
[0145] The solid support may comprise a polymer matrix (e.g., a gel, a hydrogel). The polymer matrix may be capable of penetrating intracellular spaces (e.g., around organelles). The polymer matrix may be capable of being pumped throughout the circulatory system.
[0146] Substrates and microwell arrays As used herein, a substrate can refer to a type of solid support. A substrate can refer to a solid support that may include a barcode or stochastic barcode of the present disclosure. A substrate can include, for example, a plurality of microwells. For example, a substrate can be a well array including two or more microwells. In some embodiments, a microwell can include a small reaction chamber of a defined volume. In some embodiments, a microwell can confine one or more cells. In some embodiments, a microwell can confine only one cell. In some embodiments, a microwell can confine one or more solid supports. In some embodiments, a microwell can confine only one solid support. In some embodiments, a microwell confines a single cell and a single solid support (e.g., a bead). A microwell can include a barcode reagent of the present disclosure.
[0147] Barcoding methods The present disclosure provides methods for estimating the number of distinct targets at distinct locations in a physical sample (e.g., a tissue, an organ, a tumor, a cell). The method may include placing a barcode (e.g., a stochastic barcode) near the sample, lysing the sample, associating the barcodes with distinct targets, amplifying the targets, and / or digitally counting the targets. The method may further include analyzing and / or visualizing information obtained from the spatial labels in the barcode. In some embodiments, the method includes visualizing a plurality of targets in the sample. Mapping the plurality of targets to a map of the sample may include generating a two-dimensional or three-dimensional map of the sample. The two-dimensional and three-dimensional maps can be generated before or after barcoding (e.g., stochastically barcoding) the plurality of targets in the sample. Visualizing a plurality of targets in the sample may include mapping the plurality of targets to a map of the sample. Mapping the plurality of targets to a map of the sample may include generating a two-dimensional or three-dimensional map of the sample. The two-dimensional map and the three-dimensional map can be generated before or after barcoding the multiple targets in the sample.In some embodiments, the two-dimensional map and the three-dimensional map can be generated before or after lysing the sample.The lysing of the sample before or after generating the two-dimensional map or the three-dimensional map can include heating the sample, contacting the sample with a detergent, changing the pH of the sample, or any combination thereof.
[0148] In some embodiments, barcoding the plurality of targets comprises hybridizing a plurality of barcodes to the plurality of targets to create barcoded targets (e.g., stochastically barcoded targets). Barcoding the plurality of targets may comprise generating an indexed library of barcoded targets. Generating an indexed library of barcoded targets can be performed using a solid support comprising a plurality of barcodes (e.g., stochastic barcodes).
[0149] Contact between sample and barcode The present disclosure provides a method for contacting a sample (e.g., cells) with a substrate of the present disclosure. For example, a sample including a thin section of cells, an organ, or a tissue can be contacted with a barcode (e.g., a stochastic barcode). For example, the cells can be contacted by gravity flow, where the cells can settle to create a monolayer. The sample can be a thin section of tissue. The thin section can be placed on a substrate. The sample can be one-dimensional (e.g., can form a planar surface). The sample (e.g., cells) can be spread across the substrate, for example, by growing / culturing the cells on the substrate. When the barcode is present in proximity to the target, the target can hybridize to the barcode. The barcodes can be contacted in a non-depleting ratio so that each distinct target can be associated with a distinct barcode of the present disclosure. To ensure efficient association of the target with the barcode, the target can be cross-linked to the barcode.
[0150] Cell lysis After the cells and barcodes are distributed, the cells can be lysed to release the target molecule.Cell lysis can be achieved by any of a variety of means, for example, by chemical or biochemical means, by osmotic shock, or by thermal lysis, mechanical lysis, or optical lysis.Cells can be lysed by adding a cell lysis buffer containing a detergent (e.g., SDS, Li dodecyl sulfate, Triton X-100, Tween-20, or NP-40), an organic solvent (e.g., methanol or acetone), or a digestive enzyme (e.g., proteinase K, pepsin, or trypsin), or any combination thereof.To increase the attachment of the barcode to the target, the diffusion rate of the target molecule can be altered, for example, by reducing the temperature and / or increasing the viscosity of the lysate.
[0151] In some embodiments, the sample may be lysed using filter paper. The filter paper can be impregnated with a lysis buffer on top of the filter paper. The sample can be applied to the filter paper with a pressure that can facilitate the lysis of the sample and the hybridization of the sample target with the substrate. In some embodiments, lysis can be performed by mechanical lysis, thermal lysis, optical lysis, and / or chemical lysis. Chemical lysis can include the use of digestive enzymes such as proteinase K, pepsin, and trypsin. Lysis can be performed by adding a lysis buffer to the substrate. The lysis buffer can include Tris HCl. The lysis buffer can include at least about 0.01, 0.05, 0.1, 0.5, or 1 M, or higher, of Tris HCl. The lysis buffer can include up to about 0.01, 0.05, 0.1, 0.5, or 1 M, or higher, of Tris HCl. The lysis buffer can include about 0.1 M Tris HCl. The pH of the lysis buffer can be at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or higher. The pH of the lysis buffer may be up to about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or higher. In some embodiments, the pH of the lysis buffer is about 7.5. The lysis buffer may contain a salt (e.g., LiCl). The salt concentration in the lysis buffer may be at least about 0.1, 0.5, or 1 M, or higher. The salt concentration in the lysis buffer may be up to about 0.1, 0.5, or 1 M, or higher. In some embodiments, the salt concentration in the lysis buffer is about 0.5 M. The lysis buffer may contain a detergent (e.g., SDS, Li dodecyl sulfate, triton X, tween, NP-40). The concentration of the detergent in the lysis buffer may be at least about 0.0001%, 0.0005%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, or 7%, or higher. The concentration of the detergent in the lysis buffer may be up to about 0.0001%, 0.0005%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, or 7%, or higher. In some embodiments, the concentration of the detergent in the lysis buffer is about 1% Li dodecyl sulfate. The time used in the lysis method may depend on the amount of detergent used.In some embodiments, the more detergent used, the shorter the time required for lysis. The lysis buffer may contain a chelating agent (e.g., EDTA, EGTA). The concentration of the chelating agent in the lysis buffer may be at least about 1, 5, 10, 15, 20, 25, or 30 mM, or higher. The concentration of the chelating agent in the lysis buffer may be up to about 1, 5, 10, 15, 20, 25, or 30 mM, or higher. In some embodiments, the concentration of the chelating agent in the lysis buffer is about 10 mM. The lysis buffer may contain a reducing agent (e.g., beta-mercaptoethanol, DTT). The concentration of the reducing agent in the lysis buffer may be at least about 1, 5, 10, 15, or 20 mM, or higher. The concentration of the reducing agent in the lysis buffer may be up to about 1, 5, 10, 15, or 20 mM, or higher. In some embodiments, the concentration of the reducing agent in the lysis buffer is about 5 mM. In some embodiments, the lysis buffer may include about 0.1 M Tris HCl, about pH 7.5, about 0.5 M LiCl, about 1% lithium dodecyl sulfate, about 10 mM EDTA, and about 5 mM DTT.
[0152] Lysing can be carried out at a temperature of about 4, 10, 15, 20, 25, or 30° C. Lysing can be carried out for about 1, 5, 10, 15, or 20 minutes, or longer. Lysed cells can contain at least about 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, or 700,000, or more target nucleic acid molecules. Lysed cells can contain up to about 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, or 700,000, or more target nucleic acid molecules.
[0153] Attaching a barcode to a target nucleic acid molecule After lysing the cells and releasing the nucleic acid molecules therefrom, the nucleic acid molecules can be randomly associated with the barcodes on the co-localized solid support. The association can involve hybridization of the target recognition region of the barcode with a complementary portion of the target nucleic acid molecule (e.g., the oligo(dT) of the barcode can interact with the poly(A) tail of the target). The assay conditions (e.g., buffer pH, ionic strength, temperature, etc.) used for hybridization can be selected to promote the formation of specific and stable hybrids. In some embodiments, the nucleic acid molecules released from the lysed cells can be associated with (e.g., hybridized to) multiple probes on a substrate. If the probes include oligo(dT), mRNA molecules can hybridize to the probes and be reverse transcribed. The oligo(dT) portion of the oligonucleotide can act as a primer for first-strand synthesis of cDNA molecules. For example, in the non-limiting example of barcoding illustrated in FIG. 2, mRNA molecules can be hybridized to barcodes on beads in block 216. For example, a single-stranded nucleotide fragment can hybridize to a target binding region of a barcode.
[0154] The attachment may further include ligating the target recognition region of the barcode with a portion of the target nucleic acid molecule. For example, the target binding region may include a nucleic acid sequence that may be capable of specific hybridization to a restriction site overhang (e.g., an EcoRI sticky end overhang). The assay procedure may further include treating the target nucleic acid with a restriction enzyme (e.g., EcoRI) to create a restriction site overhang. The barcode can then be ligated to any nucleic acid molecule that contains a sequence complementary to the restriction site overhang. A ligase (e.g., T4 DNA ligase) can be used to join the two fragments.
[0155] For example, in the non-limiting example of barcoding illustrated in Figure 2, the labeled targets (e.g., target-barcode molecules) from multiple cells (or multiple samples) may then be pooled, e.g., into tubes, in block 220. The labeled targets can be pooled, e.g., by collecting the barcodes and / or beads to which the target-barcode molecules are attached. The solid support-based collection of attached target-barcode molecules can be recovered using magnetic beads and an externally applied magnetic field. Once the target-barcode molecules are pooled, all further processing can proceed in a single reaction vessel. Further processing can include, for example, reverse transcription, amplification, cleavage, dissociation, and / or nucleic acid extension reactions. Further processing reactions can be performed within microwells, i.e., without first pooling labeled target nucleic acid molecules from multiple cells.
[0156] Reverse transcription The present disclosure provides a method for generating a target-barcode conjugate using reverse transcription (e.g., in block 224 of Figure 2). The target-barcode conjugate may include a barcode and a complementary sequence of all or part of a target nucleic acid (i.e., a barcoded cDNA molecule, such as a stochastically barcoded cDNA molecule). Reverse transcription of the associated RNA molecule can occur by adding a reverse transcription primer along with a reverse transcriptase. The reverse transcription primer may be an oligo(dT) primer, a random hexanucleotide primer, or a target-specific oligonucleotide primer. The oligo(dT) primer may be 12-18 nucleotides in length, or about 12-18 nucleotides in length, and can bind to the endogenous poly(A) tail at the 3' end of mammalian mRNA. The random hexanucleotide primer can bind to mRNA at various complementary sites. The target-specific oligonucleotide primer typically selectively primes from the mRNA of interest.
[0157] In some embodiments, reverse transcription of the labeled RNA molecule can occur by adding a reverse transcription primer. In some embodiments, the reverse transcription primer is an oligo(dT) primer, a random hexanucleotide primer, or a target-specific oligonucleotide primer. Generally, oligo(dT) primers are 12-18 nucleotides in length and bind to the endogenous poly(A) tail at the 3' end of mammalian mRNAs. Random hexanucleotide primers can bind to mRNAs at various complementary sites. Target-specific oligonucleotide primers typically selectively prime from the mRNA of interest. Reverse transcription can occur repeatedly to produce multiple labeled cDNA molecules. The methods disclosed herein can include performing at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 reverse transcription reactions. The methods can include performing at least about 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 reverse transcription reactions.
[0158] amplification One or more nucleic acid amplification reactions (e.g., in block 228 of FIG. 2) can be performed to generate multiple copies of the labeled target nucleic acid molecule. Amplification can be performed in a multiplexed format, in which multiple target nucleic acid sequences are simultaneously amplified. The amplification reaction can be used to add sequencing adapters to the nucleic acid molecule. The amplification reaction can include amplifying at least a portion of the sample label, if present. The amplification reaction can include amplifying at least a portion of the cell label and / or barcode sequence (e.g., molecular label). The amplification reaction can include amplifying at least a portion of the sample tag, cell label, spatial label, barcode sequence (e.g., molecular label), target nucleic acid, or a combination thereof. The amplification reaction may include amplifying 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 100%, or a range or value between any two of these values of the plurality of nucleic acids. The method may further include performing one or more cDNA synthesis reactions to produce one or more cDNA copies of the target-barcode molecule comprising the sample label, cell label, spatial label, and / or barcode sequence (e.g., molecular label).
[0159] In some embodiments, amplification can be carried out using polymerase chain reaction (PCR). As used herein, PCR can refer to a reaction for amplifying specific DNA sequences in vitro by simultaneous primer extension of complementary strands of DNA. As used herein, PCR can encompass derivative forms of this reaction, including, but not limited to, RT-PCR, real-time PCR, nested PCR, quantitative PCR, multiplex PCR, digital PCR, and assembly PCR.
[0160] Amplification of labeled nucleic acids may include non-PCR-based methods. Examples of non-PCR-based methods include, but are not limited to, multiple displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), real-time SDA, rolling circle amplification, or circle-to-circle amplification. Other non-PCR-based amplification methods include DNA-dependent RNA polymerase-driven RNA transcription amplification or multiple cycles of RNA-directed DNA synthesis and transcription to amplify DNA or RNA targets, ligase chain reaction (LCR), and Qβ replicase (Qβ) methods, the use of palindromic probes, strand displacement amplification, oligonucleotide-driven amplification using restriction endonucleases, amplification methods in which a primer hybridizes to a nucleic acid sequence and the resulting duplex is cleaved before extension and amplification, strand displacement amplification using a nucleic acid polymerase lacking 5' exonuclease activity, rolling circle amplification, and ramification extension amplification (RAM). In some embodiments, the amplification does not produce circularized transcripts.
[0161] In some embodiments, the methods disclosed herein further include performing a polymerase chain reaction on the labeled nucleic acid (e.g., labeled RNA, labeled DNA, labeled cDNA) to produce a labeled amplicon (e.g., a stochastically labeled amplicon). The labeled amplicon may be a double-stranded molecule. The double-stranded molecule may comprise a double-stranded RNA molecule, a double-stranded DNA molecule, or an RNA molecule hybridized to a DNA molecule. One or both strands of the double-stranded molecule may comprise a sample label, a spatial label, a cell label, and / or a barcode sequence (e.g., a molecular label). The labeled amplicon may be a single-stranded molecule. The single-stranded molecule may comprise DNA, RNA, or a combination thereof. The nucleic acids of the present disclosure may include synthetic or modified nucleic acids.
[0162] Amplification may include the use of one or more non-natural nucleotides. Non-natural nucleotides may include photolabile or inducible nucleotides. Examples of non-natural nucleotides include, but are not limited to, peptide nucleic acids (PNAs), morpholinos, and locked nucleic acids (LNAs), as well as glycol nucleic acids (GNAs) and threose nucleic acids (TNAs). Non-natural nucleotides may be added to one or more cycles of the amplification reaction. The addition of non-natural nucleotides can be used to distinguish products at specific cycles or time points of the amplification reaction.
[0163] Conducting one or more amplification reactions may include the use of one or more primers. The one or more primers may contain, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 or more nucleotides. The one or more primers may contain at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 or more nucleotides. The one or more primers may contain fewer than 12 to 15 nucleotides. The one or more primers may anneal to at least a portion of the multiple labeled targets (e.g., stochastically labeled targets). The one or more primers may anneal to the 3' or 5' ends of the multiple labeled targets. The one or more primers may anneal to an internal region of the multiple labeled targets. The internal region may be at least about 50, 100, 150, 200, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 650, 700, 750, 800, 850, 900, or 1000 nucleotides from the 3' end of the plurality of labeled targets. The one or more primers may comprise a panel of primers. The one or more primers may include at least one or more custom primers. The one or more primers may include at least one or more control primers. The one or more primers may include at least one or more gene-specific primers.
[0164] The one or more primers may comprise a universal primer. The universal primer may anneal to a universal primer binding site. The one or more custom primers may anneal to a first sample label, a second sample label, a spatial label, a cell label, a barcode sequence (e.g., a molecular label), a target, or any combination thereof. The one or more primers may comprise a universal primer and a custom primer. The custom primers may be designed to amplify one or more targets. The targets may comprise a subset of all nucleic acids in one or more samples. The targets may comprise a subset of all labeled targets in one or more samples. The one or more primers may comprise at least 96 or more custom primers. The one or more primers may comprise at least 960 or more custom primers. The one or more primers may comprise at least 9600 or more custom primers. The one or more custom primers may anneal to two or more different labeled nucleic acids. The two or more different labeled nucleic acids may correspond to one or more genes.
[0165] Any amplification scheme can be used in the disclosed method. For example, in one scheme, the first round of PCR can use a gene-specific primer and a primer for the universal Illumina sequencing primer 1 sequence to amplify the molecules attached to the beads. The second round of PCR can use a nested gene-specific primer flanked by the Illumina sequencing primer 2 sequence and a primer for the universal Illumina sequencing primer 1 sequence to amplify the first PCR product. The third round of PCR adds P5 and P7 and a sample index, converting the PCR product into an Illumina sequencing library. Sequencing using 150 bp x 2 sequencing can reveal cell markers and barcode sequences (e.g., molecular markers) in read 1, genes in read 2, and sample indexes in index 1 read.
[0166] In some embodiments, nucleic acids can be removed from a substrate using chemical cleavage. For example, chemical groups or modified bases present in the nucleic acid can be used to facilitate removal of the nucleic acid from a solid support. For example, enzymes can be used to remove nucleic acids from a substrate. For example, nucleic acids can be removed from a substrate by restriction endonuclease digestion. For example, nucleic acids containing dUTP or ddUTP can be removed from a substrate by treating the substrate with uracil-d-glycosylase (UDG). For example, nucleic acids can be removed from a substrate using enzymes that perform nucleotide excision, such as base excision repair enzymes such as apurinic / apyrimidinic (AP) endonucleases. In some embodiments, nucleic acids can be removed from a substrate using photocleavable groups and light. In some embodiments, nucleic acids can be removed from a substrate using a cleavable linker. For example, the cleavable linker can include at least one of biotin / avidin, biotin / streptavidin, biotin / neutravidin, Ig-Protein A, a photolabile linker, an acid- or base-labile linker group, or an aptamer.
[0167] If the probe is gene-specific, the molecule can be hybridized to the probe and reverse transcribed and / or amplified. In some embodiments, the nucleic acid can be amplified after it is synthesized (e.g., reverse transcribed). Amplification can be performed in a multiplex manner, in which multiple target nucleic acid sequences are amplified simultaneously. Amplification can add sequencing adapters to the nucleic acid.
[0168] In some embodiments, amplification can be performed on the substrate, for example, by bridge amplification. Homopolymer tails can be added to the cDNA to generate ends suitable for bridge amplification using oligo(dT) probes on the substrate. In bridge amplification, a primer complementary to the 3' end of the template nucleic acid can be the first primer of each pair covalently attached to a solid particle. A sample containing the template nucleic acid is contacted with the particle and a single thermal cycle is performed, allowing the template molecule to anneal to the first primer and extend the first primer in the forward direction by adding nucleotides to form a duplex molecule consisting of the template molecule and a newly formed DNA strand complementary to the template. The heating step of the next cycle can denature the duplex molecule, releasing the template molecule from the particle and leaving a complementary DNA strand attached to the particle via the first primer. During the annealing stage of the subsequent annealing and extension step, the complementary strand can hybridize to a second primer complementary to the segment of the complementary strand removed from the first primer. This hybridization can result in the formation of a bridge between the first and second primers, where the complementary strand is covalently attached to the first primer and hybridized to the second primer. In the extension step, the second primer can be extended in the reverse direction by adding nucleotides to the same reaction mixture, thereby converting the bridge into a double-stranded bridge. The next cycle then begins, denaturing the double-stranded bridge to yield two single-stranded nucleic acid molecules, each with one end attached to the particle surface via the first and second primers, respectively, and the other end unattached. In the annealing and extension step of this second cycle, each strand can hybridize to a complementary primer on the same particle that was not previously used, forming a new single-stranded bridge. At this point, the two hybridized, previously unused primers are extended, converting the two new bridges into double-stranded bridges.
[0169] The amplification reaction may amplify at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 100% of the plurality of nucleic acids. Amplification of labeled nucleic acids may include PCR-based or non-PCR-based methods. Amplification of labeled nucleic acids may include exponential amplification of labeled nucleic acids. Amplification of labeled nucleic acids may include linear amplification of labeled nucleic acids. Amplification may be performed by polymerase chain reaction (PCR). PCR may refer to a reaction for in vitro amplification of specific DNA sequences by simultaneous primer extension of complementary strands of DNA. PCR may encompass derivative forms of this reaction, including, but not limited to, RT-PCR, real-time PCR, nested PCR, quantitative PCR, multiplex PCR, digital PCR, suppression PCR, semi-suppressive PCR, and assembly PCR.
[0170] In some embodiments, the amplification of labeled nucleic acids includes non-PCR-based methods. Examples of non-PCR-based methods include, but are not limited to, multiple displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), real-time SDA, rolling circle amplification, or circle-to-circle amplification. Other non-PCR-based amplification methods include DNA-dependent RNA polymerase-driven RNA transcription amplification or multiple cycles of RNA-directed DNA synthesis and transcription to amplify DNA or RNA targets, ligase chain reaction (LCR), Qβ replicase (Qβ) method, the use of palindromic probes, strand displacement amplification, oligonucleotide-driven amplification using restriction endonucleases, amplification methods in which a primer hybridizes to a nucleic acid sequence and the resulting duplex is cleaved before extension and amplification, strand displacement amplification using a nucleic acid polymerase lacking 5' exonuclease activity, rolling circle amplification, and / or branched extension amplification (RAM).
[0171] In some embodiments, the methods disclosed herein further include performing a nested polymerase chain reaction on the amplified amplicon (e.g., target). The amplicon may be a double-stranded molecule. The double-stranded molecule may comprise a double-stranded RNA molecule, a double-stranded DNA molecule, or an RNA molecule hybridized to a DNA molecule. One or both strands of the double-stranded molecule may comprise a sample tag or molecular identifier label. Alternatively, the amplicon may be a single-stranded molecule. The single-stranded molecule may comprise DNA, RNA, or a combination thereof. The nucleic acid of the present disclosure may include a synthetic or modified nucleic acid.
[0172] In some embodiments, the methods include repeatedly amplifying labeled nucleic acids to produce multiple amplicons. The methods disclosed herein may include performing at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amplification reactions. Alternatively, the methods may include performing at least about 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 amplification reactions. The amplification may further include adding one or more control nucleic acids to one or more samples containing the plurality of nucleic acids. The amplification may further include adding one or more control nucleic acids to the plurality of nucleic acids. The control nucleic acids may include a control label.
[0173] Amplification may include the use of one or more non-natural nucleotides. The non-natural nucleotides may include photolabile and / or inducible nucleotides. Examples of non-natural nucleotides include, but are not limited to, peptide nucleic acids (PNAs), morpholinos, and locked nucleic acids (LNAs), as well as glycol nucleic acids (GNAs) and threose nucleic acids (TNAs). The non-natural nucleotides may be added to one or more cycles of the amplification reaction. The addition of non-natural nucleotides can be used to distinguish products at specific cycles or time points of the amplification reaction.
[0174] Conducting one or more amplification reactions may include the use of one or more primers. The one or more primers may include one or more oligonucleotides. The one or more oligonucleotides may include at least about 7-9 nucleotides. The one or more oligonucleotides may include less than 12-15 nucleotides. The one or more primers may anneal to at least a portion of the plurality of labeled nucleic acids. The one or more primers may anneal to the 3' and / or 5' ends of the plurality of labeled nucleic acids. The one or more primers may anneal to an internal region of the plurality of labeled nucleic acids. The internal region can be at least about 50, 100, 150, 200, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 650, 700, 750, 800, 850, 900, or 1000 nucleotides from the 3' end of the plurality of labeled nucleic acids. The one or more primers can comprise a panel of primers. The one or more primers may include at least one or more custom primers. The one or more primers may include at least one or more control primers. The one or more primers may include at least one or more housekeeping gene primers. The one or more primers may include a universal primer. The universal primer can anneal to a universal primer binding site. The one or more custom primers can anneal to a first sample tag, a second sample tag, a molecular identifier label, a nucleic acid, or a product thereof. The one or more primers may include a universal primer and a custom primer. The custom primer can be designed to amplify one or more target nucleic acids.The target nucleic acids may comprise a subset of the total nucleic acids in one or more samples. In some embodiments, the primers are probes attached to the arrays of the present disclosure.
[0175] In some embodiments, barcoding (e.g., stochastic barcoding) a plurality of targets in a sample further comprises generating an indexed library of barcoded targets (e.g., stochastically barcoded targets) or barcoded fragments of the targets. The barcode sequences of different barcodes (e.g., molecular labels of different stochastic barcodes) may be different from each other. Generating an indexed library of barcoded targets comprises generating a plurality of indexed polynucleotides from the plurality of targets in the sample. For example, in an indexed library of barcoded targets comprising a first indexed target and a second indexed target, the labeled region of the first indexed polynucleotide may differ from the labeled region of the second indexed polynucleotide by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, or a number or range between any two of these values. In some embodiments, generating an indexed library of barcoded targets includes contacting a plurality of targets, such as mRNA molecules, with a plurality of oligonucleotides comprising a poly(T) region and a label region; and performing first-strand synthesis using a reverse transcriptase to produce single-stranded, labeled cDNA molecules, each comprising a cDNA region and a label region, wherein the plurality of targets comprises at least two mRNA molecules of different sequences, and the plurality of oligonucleotides comprises at least two oligonucleotides of different sequences. Generating an indexed library of barcoded targets may further include amplifying the single-stranded, labeled cDNA molecules to produce double-stranded, labeled cDNA molecules; and performing nested PCR on the double-stranded, labeled cDNA molecules to produce labeled amplicons. In some embodiments, the method may include generating adapter-labeled amplicons.
[0176] Barcoding (e.g., stochastic barcoding) may involve labeling individual nucleic acid (e.g., DNA or RNA) molecules using nucleic acid barcodes or tags. In some embodiments, barcoding involves adding DNA barcodes or tags to cDNA molecules as they are generated from mRNA. Nested PCR can be performed to minimize PCR amplification bias. For example, adapters can be added for sequencing using next-generation sequencing (NGS). Sequencing results can be used to determine the sequences of nucleotide fragments of one or more copies of the cell label, molecular label, and target, for example, at block 232 of FIG. 2.
[0177] Figure 3 is a schematic diagram illustrating a non-limiting, exemplary process for generating an indexed library of barcoded targets (e.g., stochastically barcoded targets), such as barcoded mRNAs or fragments thereof. As shown in step 1, the reverse transcription process can encode each mRNA molecule with a unique molecular label, a cellular label, and a universal PCR site. In particular, RNA molecules 302 can be reverse transcribed by hybridizing (e.g., stochastic hybridization) a set of barcodes (e.g., stochastic barcodes) 310 to poly(A) tail regions 308 of the RNA molecules 302 to produce labeled cDNA molecules 304 containing cDNA regions 306. Each of the barcodes 310 can include a target-binding region, e.g., a poly(dT) region 312, a label region 314 (e.g., a barcode sequence or molecule), and a universal PCR region 316.
[0178] In some embodiments, the cellular label may comprise 3 to 20 nucleotides. In some embodiments, the molecular label may comprise 3 to 20 nucleotides. In some embodiments, each of the plurality of stochastic barcodes further comprises one or more of a universal label and a cellular label, wherein the universal label is the same for the plurality of stochastic barcodes on the solid support, and the cellular label is the same for the plurality of stochastic barcodes on the solid support. In some embodiments, the universal label may comprise 3 to 20 nucleotides. In some embodiments, the cellular label comprises 3 to 20 nucleotides.
[0179] In some embodiments, label region 314 may include a barcode sequence or molecular label 318 and a cellular label 320. In some embodiments, label region 314 may include one or more of a universal label, a dimensional label, and a cellular label. The barcode sequence or molecular label 318 may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides in length, may be about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides in length, may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides in length, or may be a number or range of nucleotides in length between any of these values. The cell label 320 may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides in length, may be about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides in length, may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides in length, or may be a number or range of nucleotides in length between any of these values.A universal label can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides in length, can be about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides in length, can be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides in length, or can be a number or range of nucleotides in length between any of these values. The universal label can be the same for multiple stochastic barcodes on the solid support, and the cell label is the same for multiple stochastic barcodes on the solid support. A dimension label may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides in length, may be about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides in length, may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides in length, or may be a number or range of nucleotides in length between any of these values.
[0180] In some embodiments, the label region 314 may include 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 different labels, such as barcode sequences or molecular labels 318 and cellular labels 320, It may include at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 different labels, or it may include at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 different labels, or any number or range between these values. Each label may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides in length, may be about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides in length, or may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides in length, or may be at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides in length, or any number or range of nucleotides in length. A set of barcodes or probabilistic barcodes 310 may include: 10, 20, 40, 50, 70, 80, 90, 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 1010 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 20 barcodes or stochastic barcodes 310, and may contain about 10, 20, 40, 50, 70, 80, 90, 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 20 barcodes or stochastic barcodes 310, and may contain at least 10, 20, 40, 50, 70, 80, 90, 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 20 barcodes or stochastic barcodes 310, or up to 10, 20, 40, 50, 70, 80, 90, 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 20A set of barcodes or stochastic barcodes 310 may contain a number or range of barcodes or stochastic barcodes 310 between any of these values. Also, a set of barcodes or stochastic barcodes 310 may, for example, each contain a unique labeled region 314. The labeled cDNA molecules 304 may be purified to remove excess barcodes or stochastic barcodes 310. Purification may include Ampure bead purification.
[0181] As shown in step 2, the products of the reverse transcription process in step 1 may be pooled into one tube and PCR amplified using a first pool of PCR primers and a first universal PCR primer. Pooling is possible due to the presence of uniquely labeled regions 314. In particular, labeled cDNA molecules 304 may be amplified to produce nested PCR labeled amplicons 322. The amplification may include multiplex PCR amplification. The amplification may include multiplex PCR amplification in a single reaction volume using 96 multiplex primers. In some embodiments, the multiplex PCR amplification may be performed in a single reaction volume using 10, 20, 40, 50, 70, 80, 90, 10, 25, 30, 45, 50, 60, 75, 80, 90, 100, 25, 30, 45, 50 ... 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 20 Multiplex primers may be used, and may be about 10, 20, 40, 50, 70, 80, 90, 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10, 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 20 Multiplex primers may be used, and may be at least 10, 20, 40, 50, 70, 80, 90, 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 20 Multiplex primers may be used, or up to 10, 20, 40, 50, 70, 80, 90, 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 20 multiplex primers may be used, or any number or range of multiplex primers between these values may be used. Amplification may include using a first PCR primer pool 324 that includes custom primers 326A-C that target specific genes and a universal primer 328. Custom primer 326 can hybridize to a region within cDNA portion 306' of labeled cDNA molecule 304. Universal primer 328 can hybridize to universal PCR region 316 of labeled cDNA molecule 304.
[0182] As shown in step 3 of Figure 3, the product of the PCR amplification in step 2 may be amplified using a nested PCR primer pool and a second universal PCR primer. Nested PCR can minimize PCR amplification bias. In particular, nested PCR-labeled amplicons 322 may be further amplified by nested PCR. Nested PCR may include multiplex PCR in a single reaction volume using a nested PCR primer pool 330 of nested PCR primers 332a-c and a second universal PCR primer 328'. The nested PCR primer pool 328 may contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 different nested PCR primers 330, and may contain about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 different nested PCR primers 330, and may contain at least 1, 2, 3, 4, 5, 6, 7, 8, 9 , 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 different nested PCR primers 330, or up to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 different nested PCR primers 330, or any number or range between any of these values. Nested PCR primer 332 contains an adaptor 334 and can hybridize to a region within cDNA portion 306'' of labeled amplicon 322. Universal primer 328' contains an adaptor 336 and can hybridize to universal PCR region 316 of labeled amplicon 322.Thus, step 3 produces adapter-labeled amplicon 338. In some embodiments, nested PCR primer 332 and second universal PCR primer 328' may not contain adapters 334 and 336. Instead, adapters 334 and 336 may be ligated to the product of the nested PCR to produce adapter-labeled amplicon 338.
[0183] As shown in step 4, the PCR products of step 3 may be PCR amplified for sequencing using library amplification primers. In particular, adapters 334 and 336 may be used to perform one or more additional assays on adapter-labeled amplicons 338. Adapters 334 and 336 may hybridize to primers 340 and 342. One or more primers 340 and 342 may be PCR amplification primers. One or more primers 340 and 342 may be sequencing primers. One or more adapters 334 and 336 may be used for further amplification of adapter-labeled amplicons 338. One or more adapters 334 and 336 may be used for sequencing of adapter-labeled amplicons 338. Primer 342 may contain a plate index 344 so that amplicons generated using the same set of barcodes or stochastic barcodes 310 can be sequenced in a single sequencing reaction using next-generation sequencing (NGS).
[0184] Compositions Comprising Cellular Component Binding Reagents Associated with Oligonucleotides - Patent application Some embodiments disclosed herein provide a plurality of compositions each comprising a cellular component binding reagent (e.g., a protein-binding reagent) conjugated to an oligonucleotide, the oligonucleotide comprising a unique identifier for the cellular component binding reagent to which it is conjugated. Cellular component binding reagents (e.g., barcoded antibodies) and their uses (e.g., for sample indexing of cells) are described in U.S. Patent Application Publication No. 2018 / 0088112 and U.S. Patent Application No. 15 / 937,713, the contents of each of which are incorporated by reference in their entirety.
[0185] In some embodiments, the cellular component-binding reagent is capable of specifically binding to a cellular component target. For example, the binding target of the cellular component-binding reagent may be or may include a carbohydrate, lipid, protein, extracellular protein, cell surface protein, cell marker, B cell receptor, T cell receptor, major histocompatibility complex, tumor antigen, receptor, integrin, intracellular protein, or any combination thereof. In some embodiments, the cellular component-binding reagent (e.g., a protein-binding reagent) is capable of specifically binding to an antigen target or a protein target. In some embodiments, each of the oligonucleotides may include a barcode, such as a stochastic barcode. The barcode may include a barcode sequence (e.g., a molecular label), a cell label, a sample label, or any combination thereof. In some embodiments, each of the oligonucleotides may include a linker. In some embodiments, each of the oligonucleotides may include a binding site for an oligonucleotide probe, such as a poly(A) tail. For example, the poly(A) tail may be untethered to a solid support or tethered to a solid support. The poly(A) tail may be about 10-50 nucleotides in length. In some embodiments, the poly(A) tail may be 18 nucleotides in length. The oligonucleotide may contain deoxyribonucleotides, ribonucleotides, or both.
[0186] The unique identifier may be, for example, a nucleotide sequence having any suitable length, for example, from about 4 nucleotides to about 200 nucleotides. In some embodiments, the unique identifier is a nucleotide sequence that is 25 nucleotides to about 45 nucleotides in length. In some embodiments, the unique identifier is a nucleotide sequence that is 4 nucleotides, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 15 nucleotides, 20 nucleotides, 25 nucleotides, 30 nucleotides, 35 nucleotides, 40 nucleotides, 45 nucleotides, 50 nucleotides, 55 nucleotides, 60 nucleotides, 70 nucleotides, 80 nucleotides, 90 nucleotides, 100 nucleotides, 200 nucleotides in length, or about ... The fragment may have a length of less than 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 70, 80, 90, 100, 200 nucleotides, or a length of more than 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 70, 80, 90, 100, 200 nucleotides, or a length that is a range between any two of the above values.
[0187] In some embodiments, the unique identifiers are selected from a diverse set of unique identifiers, which may include 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 5000 different unique identifiers, or about 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 5000 different unique identifiers, or a number or range between any two of these values. A diverse set of unique identifiers may include at least, or may include at most, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, or 5000 different unique identifiers. In some embodiments, a set of unique identifiers is designed to have minimal sequence homology to the DNA or RNA sequences of the sample to be analyzed. In some embodiments, the sequences of a set of unique identifiers differ from each other or their complements by 1 nucleotide, 2 nucleotides, 3 nucleotides, 4 nucleotides, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, or by about 1 nucleotide, 2 nucleotides, 3 nucleotides, 4 nucleotides, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, or a number or range between any two of these values.In some embodiments, the sequences of a set of unique identifiers differ from each other or their complements by at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 nucleotides, or by at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 nucleotides. In some embodiments, the sequences of a set of unique identifiers differ from each other or their complements by at least 3%, at least 5%, at least 8%, at least 10%, at least 15%, at least 20%, or more.
[0188] In some embodiments, the unique identifier may comprise a binding site for a primer, such as a universal primer. In some embodiments, the unique identifier may comprise at least two binding sites for a primer, such as a universal primer. In some embodiments, the unique identifier may comprise at least three binding sites for a primer, such as a universal primer. The primers may be used to amplify the unique identifier, for example, by PCR amplification. In some embodiments, the primers may be used in a nested PCR reaction.
[0189] The present disclosure contemplates any suitable cellular component binding reagent, such as a protein binding reagent, an antibody or fragment thereof, an aptamer, a small molecule, a ligand, a peptide, an oligonucleotide, etc., or any combination thereof. In some embodiments, the cellular component binding reagent may be a polyclonal antibody, a monoclonal antibody, a recombinant antibody, a single-chain antibody (sc-Ab), or a fragment thereof, such as a Fab or Fv. In some embodiments, the plurality of cellular component binding reagents may include 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 5000 different cellular component reagents, or about 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 5000 different cellular component reagents, or any number or range between any two of these values. In some embodiments, the plurality of cellular component binding reagents may include at least 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, or 5000 different cellular component reagents, or may include at most 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, or 5000 different cellular component reagents.
[0190] The oligonucleotide may be conjugated to the cellular component binding reagent by various mechanisms. In some embodiments, the oligonucleotide may be covalently conjugated to the cellular component binding reagent. In some embodiments, the oligonucleotide may be non-covalently conjugated to the cellular component binding reagent. In some embodiments, the oligonucleotide is conjugated to the cellular component binding reagent via a linker. The linker may be, for example, cleavable or detachable from the cellular component binding reagent and / or the oligonucleotide. In some embodiments, the linker may contain a chemical group that reversibly attaches the oligonucleotide to the cellular component binding reagent. The chemical group may be conjugated to the linker via, for example, an amine group. In some embodiments, the linker may contain a chemical group that forms a stable bond with another chemical group conjugated to the cellular component binding reagent. For example, the chemical group may be a UV light-cleavable group, a disulfide bond, streptavidin, biotin, an amine, or the like. In some embodiments, the chemical group may be conjugated to the cellular component-binding reagent via a primary amine or the N-terminus of an amino acid, such as lysine. Commercially available conjugation kits, such as the Protein-Oligo Conjugation Kit (Solulink, Inc., San Diego, CA) or the Thunder-Link® Oligo Conjugation System (Innova Biosciences, Cambridge, UK), can be used to conjugate the oligonucleotide to the cellular component-binding reagent.
[0191] The oligonucleotide can be conjugated to any suitable site on the cellular component binding reagent (e.g., a protein binding reagent) so long as it does not interfere with specific binding between the cellular component binding reagent and its cellular component target. In some embodiments, the cellular component binding reagent is a protein, such as an antibody. In some embodiments, the cellular component binding reagent is not an antibody. In some embodiments, the oligonucleotide is conjugated anywhere other than the antigen-binding site of an antibody, e.g., to the Fc region, C H 1 domain, CH 2 domains, C H 3 domains, C L The oligonucleotide can be conjugated to a domain or the like. Methods for conjugating oligonucleotides to cellular component-binding reagents (e.g., antibodies) have previously been disclosed, for example, in U.S. Pat. No. 6,531,283, the contents of which are expressly incorporated herein by reference in their entirety. The stoichiometry of the oligonucleotide relative to the cellular component-binding reagent may vary. To increase the sensitivity for detecting cellular component-binding reagent-specific oligonucleotides in sequencing, it may be advantageous to increase the ratio of oligonucleotide to cellular component-binding reagent during conjugation. In some embodiments, each cellular component-binding reagent may be conjugated to a single oligonucleotide molecule. In some embodiments, each cellular component-binding reagent may be conjugated to more than one oligonucleotide molecule, e.g., at least 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000 oligonucleotide molecules, or up to 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000 oligonucleotide molecules, or a number or range between any two of such values, each of the oligonucleotide molecules comprising the same or a different unique identifier. In some embodiments, each cellular component-binding reagent may be conjugated to more than one oligonucleotide molecule, e.g., at least 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000 oligonucleotide molecules, or up to 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000 oligonucleotide molecules, each of the oligonucleotide molecules comprising the same or a different unique identifier.
[0192] In some embodiments, the multiple cellular component binding reagents are capable of specifically binding to multiple cellular component targets in a sample, such as a single cell, multiple cells, a tissue sample, a tumor sample, or a blood sample. In some embodiments, the multiple cellular component targets include cell surface proteins, cell markers, B cell receptors, T cell receptors, antibodies, major histocompatibility complexes, tumor antigens, receptors, or any combination thereof. In some embodiments, the multiple cellular component targets may include intracellular cellular components. In some embodiments, the multiple cellular component targets may include intracellular cellular components. In some embodiments, the plurality of cellular constituents may be 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, 99% of the total cellular constituents (e.g., proteins) in a cell or organism, or about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, 99%, or a number or range between any two of these values. In some embodiments, the plurality of cellular components may represent at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, or 99% of the total cellular components (e.g., proteins) in a cell or organism. In some embodiments, the plurality of cellular component targets may include 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000, 10,000 different cellular component targets, or may include about 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000, 10,000 different cellular component targets, or a number or range between any two of these values.In some embodiments, the plurality of cellular component targets may include at least 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000, 10,000 different cellular component targets, or may include at most 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000, 10,000 different cellular component targets.
[0193] Figure 4 shows a schematic diagram of an exemplary cellular component binding reagent (e.g., an antibody) having attached thereto (e.g., conjugated thereto) an oligonucleotide comprising a unique identifier sequence of the antibody. An oligonucleotide conjugated to a cellular component binding reagent, an oligonucleotide for conjugation to a cellular component binding reagent, or an oligonucleotide previously conjugated to a cellular component binding reagent may be referred to herein as an antibody oligonucleotide (abbreviated as binding reagent oligonucleotide). An oligonucleotide conjugated to an antibody, an oligonucleotide for conjugation to an antibody, or an oligonucleotide previously conjugated to an antibody may be referred to herein as an antibody oligonucleotide (abbreviated as "AbOligo" or "AbO"). The oligonucleotide may also comprise additional components, including, but not limited to, one or more linkers, one or more unique identifiers of the antibody, optionally one or more barcode sequences (e.g., molecular beacons), and a poly(dA) tail. In some embodiments, the oligonucleotide may comprise, from 5' to 3', a linker, a unique identifier, a barcode sequence (e.g., molecular beacon), and a poly(dA) tail. The antibody oligonucleotide may be an mRNA mimic.
[0194] FIG. 5 shows a schematic diagram of an exemplary cellular component binding reagent (e.g., an antibody) associated with (e.g., conjugated to) an oligonucleotide comprising a unique identifier sequence of the antibody. The cellular component binding reagent may be capable of specifically binding to at least one cellular component target, such as an antigen target or a protein target. The binding reagent oligonucleotide (e.g., a sample indexing oligonucleotide or an antibody oligonucleotide) may comprise a sequence (e.g., a sample indexing sequence) for performing a method of the present disclosure. For example, the sample indexing oligonucleotide may comprise a sample indexing sequence for identifying the sample origin of one or more cells of the sample. The indexing sequences (e.g., sample indexing sequences) of at least two compositions (e.g., sample indexing compositions) comprising two cellular component binding reagents of a plurality of compositions comprising a cellular component binding reagent may comprise different sequences. In some embodiments, the binding reagent oligonucleotide is not homologous to a genomic sequence of a species. The binding reagent oligonucleotide may be configured to be detachable or non-detachable from the cellular component binding reagent (or may be detachable or non-detachable).
[0195] The oligonucleotide conjugated to the cellular component binding reagent may comprise, for example, a barcode sequence (e.g., a molecular label sequence), a poly(dA) tail, or a combination thereof. The oligonucleotide conjugated to the cellular component binding reagent may be an mRNA mimic. In some embodiments, the sample indexing oligonucleotide comprises a sequence complementary to the capture sequence of at least one barcode of the plurality of barcodes. The target binding region of the barcode may comprise a capture sequence. The target binding region may comprise, for example, a poly(dT) region. In some embodiments, the sequence of the sample indexing oligonucleotide complementary to the capture sequence of the barcode may comprise a poly(dA) tail. The sample indexing oligonucleotide may comprise a molecular label.
[0196] In some embodiments, the binding reagent oligonucleotides (e.g., sample oligonucleotides) are 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 128, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 40 0, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800, 810, 820, 830, 840, 850, 860, 870, 880, 890, 900, 910, 920, 930, 940, 950, 960, 970, 980, 990, 10 nucleotide sequences of about 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 128, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800, 810, 820, 830, 840, 850, 860, 870, 880, 890, 900, nucleotide sequences of 50, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800, 810, 820, 830, 840, 850, 860, 870, 880, 890, 900, 910, 920, 930, 940, 950, 960, 970, 980, 990, 1000 nucleotides in length,In some embodiments, the binding reagent oligonucleotide comprises a nucleotide sequence of at least 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 128, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800, 810, 820, 830, 840, 850, 860, 870, 880, 890, 900, 910, 920, 930, 940, 950, 960, 970, 980, 990, or 1000 Nu. of leotide length or up to 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 128, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 47 and a nucleotide sequence that is 0, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800, 810, 820, 830, 840, 850, 860, 870, 880, 890, 900, 910, 920, 930, 940, 950, 960, 970, 980, 990, or 1000 nucleotides in length.
[0197] In some embodiments, the cellular component binding reagent comprises an antibody, a tetramer, an aptamer, a protein scaffold, or a combination thereof. The binding reagent oligonucleotide may be conjugated to the cellular component binding reagent, for example, via a linker. The binding reagent oligonucleotide may comprise a linker. The linker may comprise a chemical group. The chemical group may be reversibly or irreversibly attached to the cellular component binding reagent molecule. The chemical group may be selected from the group consisting of a UV light-cleavable group, a disulfide bond, streptavidin, biotin, an amine, and any combination thereof.
[0198] In some embodiments, the cellular component binding reagent is capable of binding to ADAM10, CD156c, ANO6, ATP1B2, ATP1B3, BSG, CD147, CD109, CD230, CD29, CD298, ATP1B3, CD44, CD45, CD47, CD51, CD59, CD63, CD97, CD98, SLC3A2, CLDND1, HLA-ABC, ICAM1, ITFG3, MPZL1, NA K ATPase alpha 1, ATP1A1, NPTN, PMCA ATPase, ATP2B1, SLC1A5, SLC29A1, SLC2A1, SLC44A2, or any combination thereof.
[0199] In some embodiments, the protein target is or comprises an extracellular protein, an intracellular protein, or any combination thereof. In some embodiments, the antigen or protein target is or comprises a cell surface protein, a cell marker, a B cell receptor, a T cell receptor, a major histocompatibility complex, a tumor antigen, a receptor, an integrin, or any combination thereof. The antigen or protein target may be or comprise a lipid, a carbohydrate, or any combination thereof. The protein target can be selected from a group including at least a few protein targets. The number of antigen or protein targets may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, or about 1, 2, It may be 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, or a number or range between any two of these values. The number of protein targets was at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10,000. or may be at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10000.
[0200] A cellular component binding reagent (e.g., a protein binding reagent) may be associated with two or more binding reagent oligonucleotides (e.g., sample-indexing oligonucleotides) having the same sequence. A cellular component binding reagent may be associated with two or more binding reagent oligonucleotides having different sequences. The number of binding reagent oligonucleotides associated with a cellular component binding reagent may vary in different implementations. In some embodiments, the number of binding reagent oligonucleotides, whether having identical or different sequences, may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or a number or range between any two of these values. In some embodiments, the number of binding reagent oligonucleotides may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000, or may be at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000.
[0201] A plurality of compositions (e.g., a plurality of sample indexing compositions) comprising a cellular component binding reagent may include one or more additional cellular component binding reagents that are not conjugated to a binding reagent oligonucleotide (e.g., a sample indexing oligonucleotide), also referred to herein as cellular component binding reagents that do not comprise a binding reagent oligonucleotide (e.g., a cellular component binding reagent that does not comprise a sample indexing oligonucleotide). The number of additional cellular component binding reagents in a plurality of compositions may vary in different implementations. In some embodiments, the number of additional cellular component binding reagents may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or a number or range between any two of these values. In some embodiments, the number of additional cellular component binding reagents may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100, or may be at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100. In some embodiments, the cellular component binding reagent and any additional cellular component binding reagents may be the same.
[0202] In some embodiments, a mixture is provided that includes a cellular component binding reagent conjugated with one or more binding reagent oligonucleotides (e.g., sample indexing oligonucleotides) and a cellular component binding reagent not conjugated with binding reagent oligonucleotides. This mixture can be used in some embodiments of the methods disclosed herein, for example, for contacting a sample and / or cells. The ratio of (1) the number of cellular component binding reagents conjugated with binding reagent oligonucleotides to (2) the number of other cellular component binding reagents (e.g., the same cellular component binding reagents) not conjugated with binding reagent oligonucleotides (e.g., sample indexing oligonucleotides) or other binding reagent oligonucleotides in the mixture can vary in different implementations. In some embodiments, this ratio is 1:1, 1:1.1, 1:1.2, 1:1.3, 1:1.4, 1:1.5, 1:1.6, 1:1.7, 1:1.8, 1:1.9, 1:2, or 1:2.5, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:11, 1:12, 1:13, 1:14, 1:15, 1:16, 1:17, 1:18, 1:19, 1:20, 1:21, 1:22, 1:23, 1:24, 1:25, 1:26, 1:27, 1:28, 1:29, 1:30, 1:31, 1:32, 1:33, 1:34, 1:35, 1:36, 1:37 , 1:38, 1:39, 1:40, 1:41, 1:42, 1:43, 1:44, 1:45, 1:46, 1:47, 1:48, 1:49, 1:50, 1:51, 1:52, 1:53, 1:54, 1:55, 1:56, 1:57, 1:58, 1:59, 1:60, 1:61, 1:62, 1:63, 1:64, 1:65, 1:66, 1:67, 1:68, 1:69, 1:70, 1:71 1, 1:72, 1:73, 1:74, 1:75, 1:76, 1:77, 1:78, 1:79, 1:80, 1:81, 1:82, 1:83, 1:84, 1:85, 1:86, 1:87, 1:88, 1:89, 1:90, 1:91, 1:92, 1:93, 1:94, 1:95, 1:96, 1:97, 1:98, 1:99, 1:100, 1:200, 1:300, 1:400, 1:500 00, 1:600, 1:700, 1:800, 1:900, 1:1000, 1:2000, 1:3000, 1:4000, 1:5000, 1:6000, 1:7000, 1:8000, 1:9000, 1:10000 or approximately 1:1, 1:1.1, 1:1.2, 1:1.3, 1:1.4, 1:1.5, 1:1.6, 1:1.7, 1:1.8, 1:1.9, 1:2, 1:2.5, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:11, 1:12, 1:13, 1:14, 1:15, 1:16, 1:17, 1:18, 1:19, 1:20, 1:21, 1:22, 1:23, 1:24, 1:25, 1:26, 1:27, 1:28, 1:29, 1:30, 1:31, 1:32, 1:33, 1:34, 1:35 , 1:36, 1:37, 1:38, 1:39, 1:40, 1:41, 1:42, 1:43, 1:44, 1:45, 1:46, 1:47, 1:48, 1:49, 1:50, 1:51, 1:52, 1:53, 1:54, 1:55, 1:56, 1:57, 1:58, 1:59, 1:60, 1:61, 1:62, 1:63, 1:64, 1:65, 1:66, 1:67, 1:68, 1:69, 1:70, 1:71, 1:72, 1:73, 1:74, 1:75, 1:76, 1:77, 1:78, 1:79, 1:80, 1:81, 1:82, 1:83, 1:84, 1:85, 1:86, 1:87, 1:88, 1:89, 1:90, 1:91, 1:92, 1:93, 1:94, 1:95, 1:96, 1:97, 1:98, 1:99, 1:100, 1:101, 1:102, 1:103, 1:104, 1:105, 1:106, 1:107, 1:108, 1:109, 1:110, 1:111, 1:112, 1:113, 1:114, 1:115, 1:116, 1:117 7, 1:68, 1:69, 1:70, 1:71, 1:72, 1:73, 1:74, 1:75, 1:76, 1:77, 1:78, 1:79, 1:80, 1:81, 1:82, 1:83, 1:84, 1:85, 1:86, 1:87, 1:88, 1:89, 1:90, 1:91, 1:92, 1:93, 1:94, 1:95, 1:96, 1:97, 1:98, 1: 1:99, 1:100, 1:200, 1:300, 1:400, 1:500, 1:600, 1:700, 1:800, 1:900, 1:1000, 1:2000, 1:3000, 1:4000, 1:5000, 1:6000, 1:7000, 1:8000, 1:9000, 1:10000, or a number or range between any two of the above values. In some embodiments, the ratio is at least 1:1, 1:1.1, 1:1.2, 1:1.3, 1:1.4, 1:1.5, 1:1.6, 1:1.7, 1:1.8, 1:1.9, 1:2.5, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:11, 1:12, 1:13, 1:14, 1:15, 1:16, 1:17, 1:18, 1:19, 1:20, 1:21, 1:22, 1:23, 1:24, 1:25, 1:26, 1:27, 1:28, 1:29, 1:30, 1:31, 1:32, 1:33, 1:34, 1:35, 1:36, 1:37, 1: 38, 1:39, 1:40, 1:41, 1:42, 1:43, 1:44, 1:45, 1:46, 1:47, 1:48, 1:49, 1:50, 1:51, 1:52, 1:53, 1:54, 1:55, 1:56, 1:57, 1:58, 1:59, 1:60, 1:61, 1:62, 1:63, 1:64, 1:65, 1:66, 1:67, 1:68, 1:69, 1:70, 1:71, 1:72 , 1:73, 1:74, 1:75, 1:76, 1:77, 1:78, 1:79, 1:80, 1:81, 1:82, 1:83, 1:84, 1:85, 1:86, 1:87, 1:88, 1:89, 1:90, 1:91, 1:92, 1:93, 1:94, 1:95, 1:96, 1:97, 1:98, 1:99, 1:100, 1:200, 1:300, 1:400, 1:500, 1:600, It may be 1:700, 1:800, 1:900, 1:1000, 1:2000, 1:3000, 1:4000, 1:5000, 1:6000, 1:7000, 1:8000, 1:9000, or 1:10000, or up to 1:1, 1:1.1, 1:1.2, 1:1.3, 1:1.4, 1:1.5, 1:1.6, 1:1.7, 1:1.8, 1:1.9, 1:20, 1:21.5, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:11, 1:12, 1:13, 1:14, 1:15, 1:16, 1:17, 1:18, 1:19, 1:20, 1:21, 1:22, 1:23, 1:24, 1:25, 1:26, 1:27, 1:28, 1:29, 1:30, 1:31, 1:32, 1:33, 1:34 , 1:35, 1:36, 1:37, 1:38, 1:39, 1:40, 1:41, 1:42, 1:43, 1:44, 1:45, 1:46, 1:47, 1:48, 1:49, 1:50, 1:51, 1:52, 1:53, 1:54, 1:55, 1:56, 1:57, 1:58, 1:59, 1:60, 1:61, 1:62, 1:63, 1:64, 1:65 5, 1:66, 1:67, 1:68, 1:69, 1:70, 1:71, 1:72, 1:73, 1:74, 1:75, 1:76, 1:77, 1:78, 1:79, 1:80, 1:81, 1:82, 1:83, 1:84, 1:85, 1:86, 1:87, 1:88, 1:89, 1:90, 1:91, 1:92, 1:93, 1:94, 1:95, 1: The ratio may be 96, 1:97, 1:98, 1:99, 1:100, 1:200, 1:300, 1:400, 1:500, 1:600, 1:700, 1:800, 1:900, 1:1000, 1:2000, 1:3000, 1:4000, 1:5000, 1:6000, 1:7000, 1:8000, 1:9000, or 1:10000.
[0203] In some embodiments, the ratio is 1:1, 1.1:1, 1.2:1, 1.3:1, 1.4:1, 1.5:1, 1.6:1, 1.7:1, 1.8:1, 1.9:1, 2:1, 2.5:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 21:1, 22:1, 23:1, 24:1, 25:1, 26:1, 27:1, 28:1, 29:1, 30:1, 31:1, 32:1, 33:1, 34:1, 35:1, 36:1, 37:1, 38:1, 39:1, 40:1, 41:1, 42:1, 43:1, 44:1, 45:1, 46:1, 47:1, 48:1, 49:1, 50:1, 51:1, 52:1, 53:1, 54:1, 55:1, 56:1, 57:1, 58:1, 59:1, 60:1, 61:1, 62:1, 63:1, 64:1, 65:1, 66:1, 67:1, 68:1, 69:1, 70:1, 71:1, 72:1, 73:1, 74:1 :1, 26:1, 27:1, 28:1, 29:1, 30:1, 31:1, 32:1, 33:1, 34:1, 35:1, 36:1, 37:1, 38:1, 39:1, 40:1, 41:1, 42:1, 43:1, 44:1, 45:1, 46:1, 47:1, 48:1, 49:1, 50:1, 51:1, 52:1, 53:1, 54:1, 55:1, 56:1, 57:1, 58:1, 59:1, 60:1, 61:1, 62:1, 6 3:1, 64:1, 65:1, 66:1, 67:1, 68:1, 69:1, 70:1, 71:1, 72:1, 73:1, 74:1, 75:1, 76:1, 77:1, 78:1, 79:1, 80:1, 81:1, 82:1, 83:1, 84:1, 85:1, 86:1, 87:1, 88:1, 89:1, 90:1, 91:1, 92:1, 93:1, 94:1, 95:1, 96:1, 97:1, 98:1, 99:1, 100:1 , 200:1, 300:1, 400:1, 500:1, 600:1, 700:1, 800:1, 900:1, 1000:1, 2000:1, 3000:1, 4000:1, 5000:1, 6000:1, 7000:1, 8000:1, 9000:1, 10000:1, or approximately 1:1, 1.1:1, 1.2:1, 1.3:1, 1.4:1, 1.5:1, 1.6:1, 1.7:1, 1.8:1, 1.9:1, 2:1, 2.5:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 21:1, 22:1, 23:1, 24:1, 25:1, 26:1, 27:1, 28:1, 29:1, 30:1, 31:1, 32:1, 33:1, 34:1, 35 :1, 36:1, 37:1, 38:1, 39:1, 40:1, 41:1, 42:1, 43:1, 44:1, 45:1, 46:1, 47:1, 48:1, 49:1, 50:1, 51:1, 52:1, 53:1, 54:1, 55:1, 56:1, 57:1, 58:1, 59:1, 60:1, 61:1, 62:1, 63:1, 64:1, 65:1, 66:1, 67 :1, 68:1, 69:1, 70:1, 71:1, 72:1, 73:1, 74:1, 75:1, 76:1, 77:1, 78:1, 79:1, 80:1, 81:1, 82:1, 83:1, 84:1, 85:1, 86:1, 87:1, 88:1, 89:1, 90:1, 91:1, 92:1, 93:1, 94:1, 95:1, 96:1, 97:1, 98:1, 99 1.1:1, 1.2:1, 1.3:1, 1.4:1, 1.5:1, 1.6:1, 1.7:1, 1.8:1, 1.9:1, 2:1, 2:2, 3:1, 3:1, 4:1, 4:2, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 10:1, 100:1, 200:1, 300:1, 4000:1, 5000:1, 6000:1, 7000:1, 8:1, 9:1, 10:1, 10:2, 10:1 ...5:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 21:1, 22:1, 23:1, 24:1, 25:1, 26:1, 27:1, 28:1, 29:1, 30:1, 31:1, 32:1, 33:1, 34:1, 35:1, 36:1, 37:1, 38:1, 39:1, 40:1, 41:1, 42:1, 43:1, 44:1, 45:1, 46:1, 47:1, 48:1, 49:1, 50:1, 51:1, 52:1, 53:1, 54:1, 55:1, 56:1, 57:1, 58:1, 59:1, 60:1, 61:1, 62:1, 63:1, 64:1, 65:1, 66:1, 67:1, 68:1, 69:1, 70:1, 71:1, 72 :1, 73:1, 74:1, 75:1, 76:1, 77:1, 78:1, 79:1, 80:1, 81:1, 82:1, 83:1, 84:1, 85:1, 86:1, 87:1, 88:1, 89:1, 90:1, 91:1, 92:1, 93:1, 94:1, 95:1, 96:1, 97:1, 98:1, 99:1, 100:1, 200:1, 300:1, 400:1, 500:1, 600: It may be 1, 700:1, 800:1, 900:1, 1000:1, 2000:1, 3000:1, 4000:1, 5000:1, 6000:1, 7000:1, 8000:1, 9000:1, or 10000:1, or may be at most 1:1, 1.1:1, 1.2:1, 1.3:1, 1.4:1, 1.5:1, 1.6:1, 1.7:1, 1.8:1, 1.9:1, 2.0:1, 2.1:1.5:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 21:1, 22:1, 23:1, 24:1, 25:1, 26:1, 27:1, 28:1, 29:1, 30:1, 31:1, 32:1, 33:1, 34 :1, 35:1, 36:1, 37:1, 38:1, 39:1, 40:1, 41:1, 42:1, 43:1, 44:1, 45:1, 46:1, 47:1, 48:1, 49:1, 50:1, 51:1, 52:1, 53:1, 54:1, 55:1, 56:1, 57:1, 58:1, 59:1, 60:1, 61:1, 62:1, 63:1, 64:1, 65 :1, 66:1, 67:1, 68:1, 69:1, 70:1, 71:1, 72:1, 73:1, 74:1, 75:1, 76:1, 77:1, 78:1, 79:1, 80:1, 81:1, 82:1, 83:1, 84:1, 85:1, 86:1, 87:1, 88:1, 89:1, 90:1, 91:1, 92:1, 93:1, 94:1, 95:1, 9 It may be 6:1, 97:1, 98:1, 99:1, 100:1, 200:1, 300:1, 400:1, 500:1, 600:1, 700:1, 800:1, 900:1, 1000:1, 2000:1, 3000:1, 4000:1, 5000:1, 6000:1, 7000:1, 8000:1, 9000:1, or 10000:1.
[0204] The cellular component binding reagent may or may not be conjugated to a binding reagent oligonucleotide (e.g., a sample indexing oligonucleotide). In some embodiments, the percentage of cellular component binding reagents conjugated to a binding reagent oligonucleotide (e.g., a sample indexing oligonucleotide) in a mixture containing cellular component binding reagents conjugated to a binding reagent oligonucleotide (e.g., a sample indexing oligonucleotide) and cellular component binding reagents not conjugated to a binding reagent oligonucleotide is 0.000000001%, 0 .00000001%, 0.0000001%, 0.000001%, 0.00001%, 0.0001%, 0.001%, 0.01%, 0.1%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31% ,32%,33%,34%,35%,36%,37%,38%,39%,40%,41%,42%,43%,44%,45%,46%,47%,48%,49%,50%,51%,52%,53%,54%,55%,56%,57%,58%,59%,60%,61%,62%,63%,64%,65%,66%,67%,68%,69%,70%,71%,72%,73%,74%,75%,76% , 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or approximately 0.000000001%, 0.00000001%, 0.0000001%, 0.00001%, 0.0001%, 0.01%, 0.1%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, The percentage may be 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or range between any two of these values. In some embodiments, the percentage of cellular component binding reagents in the mixture to which sample-indexing oligonucleotides are conjugated is at least 0.000000001%, 0.00000001%, 0.0000001%, 0.000001%, 0.0001%, 0.001%, 0.01%, 0.1%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, 101%, 102%, 103 %, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70% , 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, or may be up to 0.000000001%, 0.00000001%, 0.0000001%, 0.000001%, 0.00001%, 0.0001%, 0.001%, 0.01%, 0.1%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43% , 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%.
[0205] In some embodiments, in a mixture comprising a cellular component binding reagent conjugated with a binding reagent oligonucleotide (e.g., a sample indexing oligonucleotide) and a cellular component binding reagent not conjugated with a sample indexing oligonucleotide, The percentages are: 0.000000001%, 0.00000001%, 0.0000001%, 0.000001%, 0.00001%, 0.0001%, 0.001%, 0.01%, 0.1%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27% ,28%,29%,30%,31%,32%,33%,34%,35%,36%,37%,38%,39%,40%,41%,42%,43%,44%,45%,46%,47%,48%,49%,50%,51%,52%,53%,54%,55%,56%,57%,58%,59%,60%,61%,62%,63%,64%,65%,66%,67%,68%,69%,70%,71%,72%,73%,74% , 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or approximately 0.000000001%, 0.00000001%, 0.0000001%, 0.00001%, 0.0001%, 0.01%, 0.1%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, The percentage may be 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or a number or range between any two of these values. In some embodiments, the percentage of cellular component binding reagent in the mixture that does not have a binding reagent oligonucleotide conjugated thereto is at least 0.000000001%, 0.00000001%, 0.0000001%, 0.000001%, 0.00001%, 0.0001%, 0.001%, 0.01% ,0.1%,1%,2%,3%,4%,5%,6%,7%,8%,9%,10%,11%,12%,13%,14%,15%,16%,17%,18%,19%,20%,21%,22%,23%,24%,25%,26%,27%,28%,29%,30%,31%,32%,33%,34%,35%,36%,37 %,38%,39%,40%,41%,42%,43%,44%,45%,46%,47%,48%,49%,50%,51%,52%,53%,54%,55%,56%,57%,58%,59%,60%,61%,62%,63%,64%,65%,66%,67%,68%,69%,70%,71%,72%,73%,74%,75%,76%,77%,78%,79 ...0%,71%,72%,73%,74%,75%,76%,77%,78%,79%,70%,71%,72%,73%,74%,75%,76%,77%,78%,79%,70%,71%,72%,73%,74%,75%,76%,77%,78%,79 It may be 3%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, or at most 0.000000001%, 0.00000001%, 0.0000001%, 0.000001%, 0.00001%, 0.0001%, 0.001%, 0.01%, 0.1%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 4 It can be 5%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%.
[0206] Cellular component cocktail In some embodiments, a cocktail of cellular component binding reagents (e.g., an antibody cocktail) can be used to increase the labeling sensitivity of the methods disclosed herein. Without being bound by any particular theory, this may be because cellular component or protein expression can vary between cell types and cell states, making it difficult to find a universal cellular component binding reagent or antibody that labels all cell types. For example, using a cocktail of cellular component binding reagents can enable more sensitive and efficient labeling of more sample types. A cocktail of cellular component binding reagents can include two or more different types of cellular component binding reagents, e.g., a broader range of cellular component binding reagents or antibodies. Cellular component binding reagents that label different cellular component targets can be pooled together to create a cocktail that sufficiently labels all cell types or one or more cell types of interest.
[0207] In some embodiments, each of the plurality of compositions (e.g., sample indexing compositions) comprises a cellular component binding reagent. In some embodiments, a composition of the plurality of compositions comprises two or more cellular component binding reagents, each of the two or more cellular component binding reagents being associated with a binding reagent oligonucleotide (e.g., a sample indexing oligonucleotide), and at least one of the two or more cellular component binding reagents being capable of specifically binding to at least one of one or more cellular component targets. The sequences of the binding reagent oligonucleotides associated with the two or more cellular component binding reagents may be identical. The sequences of the binding reagent oligonucleotides associated with the two or more cellular component binding reagents may comprise different sequences. Each of the plurality of compositions may comprise two or more cellular component binding reagents.
[0208] The number of different types of cellular component binding reagents (e.g., CD147 antibodies and CD47 antibodies) in a composition may vary in different implementations. A composition having two or more different types of cellular component binding reagents may be referred to herein as a cellular component binding reagent cocktail (e.g., a sample indexing composition cocktail). The number of different types of cellular component binding reagents in a cocktail may vary. In some embodiments, the number of different types of cellular component binding reagents in the cocktail can be at or about 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 10000, 100000, or a number or range between any two of these values. In some embodiments, the number of different types of cellular component binding reagents in the cocktail may be at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 10,000, or 100,000, or may be at most 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 10,000, or 100,000. Different types of cellular component binding reagents may be conjugated to binding reagent oligonucleotides having the same or different sequences (eg, sample indexing sequences).
[0209] Methods for quantitative analysis of cellular component targets In some embodiments, the methods disclosed herein can also be used for quantitative analysis of multiple cellular component targets (e.g., protein targets) in a sample using the compositions disclosed herein and oligonucleotide probes in which barcode sequences (e.g., molecular beacon sequences) can be attached to the oligonucleotides of a cellular component-binding reagent (e.g., a protein-binding reagent). The oligonucleotides of a cellular component-binding reagent can be or include antibody oligonucleotides, sample-indexing oligonucleotides, cell-identification oligonucleotides, control particle oligonucleotides, control oligonucleotides, interaction-determining oligonucleotides, etc. In some embodiments, the sample can be a single cell, multiple cells, a tissue sample, a tumor sample, a blood sample, etc. In some embodiments, the sample can include normal cells, tumor cells, blood cells, B cells, T cells, maternal cells, fetal cells, etc., or a mixture of cells from various subjects. In some embodiments, the sample may comprise a plurality of single cells separated into individual compartments, such as microwells in a microwell array.
[0210] In some embodiments, the binding targets of the multiple cellular component targets (i.e., cellular component targets) may be or include carbohydrates, lipids, proteins, extracellular proteins, cell surface proteins, cell markers, B cell receptors, T cell receptors, major histocompatibility complexes, tumor antigens, receptors, integrins, intracellular proteins, or any combination thereof. In some embodiments, the cellular component targets are protein targets. In some embodiments, the multiple cellular component targets include cell surface proteins, cell markers, B cell receptors, T cell receptors, antibodies, major histocompatibility complexes, tumor antigens, receptors, or any combination thereof. In some embodiments, the multiple cellular component targets may include intracellular cellular components. In some embodiments, the plurality of cellular constituents may represent at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 6%, at least 7%, at least 8%, at least 9%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or more of all cellular constituents encoded in the organism. In some embodiments, the plurality of cellular constituent targets may include at least 2, at least 3, at least 4, at least 5, at least 10, at least 20, at least 30, at least 40, at least 50, at least 100, at least 1,000, at least 10,000, or more different cellular constituent targets.
[0211] In some embodiments, multiple cellular component binding reagents are contacted with the sample for specific binding to multiple cellular component targets. Unbound cellular component binding reagents can be removed, for example, by washing. In embodiments in which the sample contains cells, any cellular component binding reagents that are not specifically bound to cells can be removed.
[0212] In some examples, cells from a cell population can be separated (e.g., isolated) into wells of a substrate of the present disclosure. The cell population can be diluted before separation. The cell population can be diluted so that at least 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% of the wells of the substrate receive single cells. The cell population can be diluted so that at most 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% of the wells of the substrate receive single cells. The cell population can be diluted so that the number of cells in the diluted population is or is at least 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% of the number of cells in the wells on the substrate. The cell population can be diluted so that the number of cells in the diluted population is 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% of the number of wells on the substrate. In some cases, the cell population is diluted so that the number of cells is about 10% of the number of wells on the substrate.
[0213] The distribution of single cells into the wells of the substrate may follow a Poisson distribution. For example, there may be at least a 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10% or higher probability that a well of the substrate has two or more cells. There may be at least a 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10% or higher probability that a well of the substrate has two or more cells. The distribution of single cells into the wells of the substrate may be random. The distribution of single cells into the wells of the substrate may not be random. Cells may be separated so that a well of the substrate receives only one cell.
[0214] In some embodiments, the cellular component binding reagents may be further conjugated to fluorescent molecules to enable flow sorting of cells into individual compartments. In some embodiments, the methods disclosed herein provide for contacting a sample with multiple compositions for specific binding to multiple cellular component targets. It will be appreciated that the conditions used may allow for specific binding of a cellular component binding reagent, e.g., an antibody, to the cellular component targets. After the contacting step, unbound compositions may be removed. For example, in embodiments where the sample contains cells and the compositions specifically bind to cellular component targets that are cell surface cellular components, such as cell surface proteins, the unbound compositions may be removed by washing the cells with a buffer solution, such that only compositions that specifically bind to the cellular component targets remain with the cells.
[0215] In some embodiments, the methods disclosed herein may include associating oligonucleotides (e.g., barcodes or stochastic barcodes) comprising barcode sequences (e.g., molecular labels), cell labels, sample labels, etc., or any combination thereof, with a plurality of oligonucleotides associated with the cellular component-binding reagent. For example, a plurality of oligonucleotide probes comprising barcodes can be used to hybridize to a plurality of oligonucleotides of the composition.
[0216] In some embodiments, the multiple oligonucleotide probes may be immobilized on a solid support. The solid support may be floating, e.g., beads in solution. The solid support may be encapsulated in a semi-solid or solid array. In some embodiments, the multiple oligonucleotide probes may not be immobilized on a solid support. When the multiple oligonucleotide probes are present in close proximity to the multiple oligonucleotides of the cellular component binding reagent associated therewith, the multiple oligonucleotides of the cellular component binding reagent may hybridize to the oligonucleotide probes. The oligonucleotide probes may be contacted in a non-depleting ratio so that each distinct oligonucleotide of the cellular component binding reagent may be associated with an oligonucleotide probe having a different barcode sequence (e.g., molecular label) of the present disclosure.
[0217] In some embodiments, the methods disclosed herein provide for desorbing oligonucleotides from cellular component-binding reagents specifically bound to cellular component targets. Desorption can be achieved by various methods that separate chemical groups from the cellular component-binding reagent, such as UV light cleavage, chemical treatment (e.g., dithiothreitol treatment), heating, enzymatic treatment, or a combination thereof. Desorption of oligonucleotides from the cellular component-binding reagent can be performed before, after, or during the step of hybridizing multiple oligonucleotide probes to multiple oligonucleotides of the composition.
[0218] Method for simultaneous quantitative analysis of cellular components and nucleic acid targets In some embodiments, the methods disclosed herein can also be used for the simultaneous quantitative analysis of multiple cellular component targets (e.g., protein targets) and multiple nucleic acid target molecules in a sample using oligonucleotide probes that can be associated with barcode sequences (e.g., molecular label sequences) on both the oligonucleotides and nucleic acid target molecules of the compositions and cellular component binding reagents disclosed herein. Other methods for the simultaneous quantitative analysis of multiple cellular component targets and multiple nucleic acid target molecules are described in U.S. Patent Application No. 15 / 715,028, filed September 25, 2017, the entire contents of which are incorporated herein by reference. In some embodiments, the sample can be a single cell, multiple cells, a tissue sample, a tumor sample, a blood sample, etc. In some embodiments, the sample can include a mixture of cell types, such as normal cells, tumor cells, blood cells, B cells, T cells, maternal cells, fetal cells, or a mixture of cells from different subjects.
[0219] In some embodiments, the sample may comprise a plurality of single cells separated into individual compartments, such as microwells in a microwell array.
[0220] In some embodiments, the multiple cellular component targets include cell surface proteins, cell markers, B cell receptors, T cell receptors, antibodies, major histocompatibility complexes, tumor antigens, receptors, or any combination thereof. In some embodiments, the multiple cellular component targets can include intracellular cellular components. In some embodiments, the plurality of cellular components may be present at or about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, 99%, or a number or range between any two of these values, of all cellular components, such as proteins, expressed in the organism or in one or more cells of the organism. In some embodiments, the plurality of cellular constituents may represent at least or up to 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, or 99% of all cellular constituents, such as proteins, that may be expressed in an organism or one or more cells of the organism. In some embodiments, the plurality of cellular constituent targets may include 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000, 10,000, or a number or range between any two of these values, or may include about 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000, 10,000, or a number or range between any two of these values, different cellular constituent targets. In some embodiments, the multiple cellular component targets may include at least 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000, or 10,000 different cellular component targets, or may include up to 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000, or 10,000 different cellular component targets. In some embodiments, multiple cellular component binding reagents are contacted with the sample to specifically bind to multiple cellular component targets. Unbound cellular component binding reagents can be removed, for example, by washing. In embodiments in which the sample contains cells, any cellular component binding reagents that are not specifically bound to cells can be removed.
[0221] In some examples, cells from a cell population can be separated (e.g., isolated) into wells of a substrate of the present disclosure. The cell population can be diluted before separation. The cell population can be diluted so that at least 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% of the wells of the substrate receive single cells. The cell population can be diluted so that at most 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% of the wells of the substrate receive single cells. The cell population can be diluted so that the number of cells in the diluted population is or is at least 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% of the number of cells in the wells on the substrate. The cell population can be diluted so that the number of cells in the diluted population is 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% of the number of wells on the substrate. In some cases, the cell population is diluted so that the number of cells is about 10% of the number of wells on the substrate.
[0222] The distribution of single cells into the wells of the substrate may follow a Poisson distribution. For example, there may be at least a 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10% or higher probability that a well of the substrate has two or more cells. There may be at least a 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10% or higher probability that a well of the substrate has two or more cells. The distribution of single cells into the wells of the substrate may be random. The distribution of single cells into the wells of the substrate may not be random. Cells may be separated so that a well of the substrate receives only one cell.
[0223] In some embodiments, the cellular component binding reagents may be further conjugated to fluorescent molecules to enable flow sorting of cells into individual compartments. In some embodiments, the methods disclosed herein provide for contacting a sample with multiple compositions for specific binding to multiple cellular component targets. It will be appreciated that the conditions used may allow for specific binding of a cellular component binding reagent, e.g., an antibody, to the cellular component targets. After the contacting step, unbound compositions may be removed. For example, in embodiments where the sample contains cells and the compositions specifically bind to cellular component targets on the cell surface, such as cell surface proteins, unbound compositions may be removed by washing the cells with a buffer solution, such that only compositions that specifically bind to the cellular component targets remain with the cells.
[0224] In some embodiments, the methods disclosed herein may provide for the release of multiple nucleic acid target molecules from a sample, for example, cells. For example, cells may be lysed to release multiple nucleic acid target molecules. Cell lysis may be achieved by any of a variety of means, such as chemical treatment, osmotic shock, heat treatment, mechanical treatment, optical treatment, or any combination thereof. Cells may be lysed by adding a cell lysis buffer containing a detergent (e.g., SDS, Li-dodecyl sulfate, Triton X-100, Tween-20, or NP-40), an organic solvent (e.g., methanol or acetone), or a digestive enzyme (e.g., proteinase K, pepsin, or trypsin), or any combination thereof.
[0225] It will be recognized by those skilled in the art that a plurality of nucleic acid molecules can comprise various nucleic acid molecules.In some embodiments, a plurality of nucleic acid molecules can comprise DNA molecules, RNA molecules, genomic DNA molecules, mRNA molecules, rRNA molecules, siRNA molecules, or a combination thereof, and can be double-stranded or single-stranded.In some embodiments, a plurality of nucleic acid molecules comprises 100, 1000, 10000, 20000, 30000, 40000, 50000, 100000, 1000000 species, or any two of these values between the number or range of species, or comprises about 100, 1000, 10000, 20000, 30000, 40000, 50000, 100000, 1000000 species, or any two of these values between the number or range of species. In some embodiments, the plurality of nucleic acid molecules comprises at least 100, 1000, 10,000, 20,000, 30,000, 40,000, 50,000, 100,000, or 1,000,000 species, or up to 100, 1000, 10,000, 20,000, 30,000, 40,000, 50,000, 100,000, or 1,000,000 species. In some embodiments, the plurality of nucleic acid molecules can be derived from a sample, for example, a single cell, or multiple cells. In some embodiments, the plurality of nucleic acid molecules can be pooled from multiple samples, for example, multiple single cells.
[0226] In some embodiments, the methods disclosed herein may include associating a plurality of nucleic acid target molecules and a plurality of oligonucleotides of a cellular component-binding reagent with barcodes (e.g., stochastic barcodes), which may include barcode sequences (e.g., molecular labels), cell labels, sample labels, etc., or any combination thereof. For example, a plurality of oligonucleotide probes including the stochastic barcodes may be used to hybridize a plurality of nucleic acid target molecules with a plurality of oligonucleotides of a composition.
[0227] In some embodiments, multiple oligonucleotide probes can be immobilized on a solid support. The solid support can be floating, for example, beads in solution. The solid support can be enclosed in a semi-solid or solid array. In some embodiments, multiple oligonucleotide probes can be unimmobilized on a solid support. When multiple oligonucleotide probes are present in proximity to multiple oligonucleotides of multiple nucleic acid target molecules and cellular component binding reagents, the multiple oligonucleotides of multiple nucleic acid target molecules and cellular component binding reagents can hybridize to the oligonucleotide probes. The oligonucleotide probes can be contacted in a non-depleting ratio so that each of the oligonucleotides of the distinct nucleic acid target molecules and cellular component binding reagents can be associated with an oligonucleotide probe (e.g., molecular label) having a different barcode sequence of the present disclosure.
[0228] In some embodiments, the methods disclosed herein provide for desorbing oligonucleotides from cellular component-binding reagents specifically bound to cellular component targets. Desorption can be achieved by various methods, such as UV light cleavage, chemical treatment (e.g., dithiothreitol treatment), heating, enzymatic treatment, or a combination thereof, to separate the chemical group from the cellular component-binding reagent. Desorption of the oligonucleotides from the cellular component-binding reagent can be performed before, after, or during the step of hybridizing multiple oligonucleotide probes to multiple oligonucleotides of multiple nucleic acid target molecules and compositions.
[0229] Simultaneous quantitative analysis of protein and nucleic acid targets In some embodiments, the methods disclosed herein can also be used for the simultaneous quantitative analysis of multiple target molecules, e.g., protein and nucleic acid targets. For example, the target molecules can be or include cellular components. Figure 6 shows a schematic diagram of an exemplary method for the simultaneous quantitative analysis of both nucleic acid targets and other cellular component targets (e.g., proteins) in a single cell. In some embodiments, multiple compositions 605, 605b, 605c, etc., each comprising a cellular component binding reagent, such as an antibody, are provided. Different cellular component binding reagents, such as antibodies, that bind to different cellular component targets are conjugated with different unique identifiers. The cellular component binding reagents can then be incubated with a sample containing multiple cells 610. The different cellular component binding reagents can specifically bind to cellular components on the cell surface, such as cell markers, B cell receptors, T cell receptors, antibodies, major histocompatibility complexes, tumor antigens, receptors, or any combination thereof. Unbound cellular component binding reagents can be removed, for example, by washing the cells with a buffer solution. The cells containing the cellular component binding reagents can then be separated into multiple compartments, such as a microwell array, where a single compartment 615 is sized to fit a single cell and a single bead 620. Each bead can contain multiple oligonucleotide probes, which can include a cell label common to all oligonucleotide probes on the bead and a barcode sequence (e.g., a molecular beacon sequence). In some embodiments, each oligonucleotide probe can include a target-binding region, e.g., a poly(dT) sequence. The oligonucleotides 625 conjugated to the cellular component binding reagents can be released from the cellular component binding reagents using chemical, optical, or other means. The cells can be lysed (635) to release intracellular nucleic acids, such as genomic DNA or cellular mRNA 630. The cellular mRNA 630, the oligonucleotides 625, or both can be captured by the oligonucleotide probes on the beads 620, for example, by hybridizing to the poly(dT) sequence.Using a reverse transcriptase, oligonucleotide probes hybridized to the cellular mRNA 630 and oligonucleotide 625 can be extended using the cellular mRNA 630 and oligonucleotide 625 as templates. The extension products generated by the reverse transcriptase can be subjected to amplification and sequencing. The sequencing reads can be subjected to demultiplexing of sequences or identifiers, such as cell labels, barcodes (e.g., molecular labels), genes, oligonucleotides specific for cellular component-binding reagents (e.g., antibody-specific oligonucleotides), etc., which can result in a digital representation of the cellular component and gene expression of each single cell in the sample.
[0230] Barcode attachment Oligonucleotides associated with cellular component-binding reagents (e.g., antigen-binding reagents or protein-binding reagents) and / or nucleic acid molecules can be randomly associated with oligonucleotide probes (e.g., barcodes, such as stochastic barcodes). Oligonucleotides associated with cellular component-binding reagents, referred to herein as binding reagent oligonucleotides, can be or include oligonucleotides of the present disclosure, such as antibody oligonucleotides, sample-indexing oligonucleotides, cell-identification oligonucleotides, control particle oligonucleotides, control oligonucleotides, interaction-determining oligonucleotides, etc. Association can involve, for example, hybridization of the target-binding region of the oligonucleotide probe to the complementary portion of the target nucleic acid molecule and / or protein-binding reagent oligonucleotide. For example, the oligo(dT) region of the barcode (e.g., stochastic barcode) can interact with the poly(A) tail of the target nucleic acid molecule and / or the poly(dA) tail of the protein-binding reagent oligonucleotide. Assay conditions used for hybridization (e.g., buffer pH, ionic strength, temperature, etc.) can be selected to promote the formation of specific, stable hybrids.
[0231] The present disclosure provides a method for using reverse transcription to attach a molecular label to an oligonucleotide attached to a target nucleic acid and / or a cellular component-binding reagent. Reverse transcription can use both RNA and DNA as templates. For example, the oligonucleotide originally conjugated to the cellular component-binding reagent can be either RNA or DNA bases, or both. The binding reagent oligonucleotide can be copied and linked (e.g., covalently linked) to the sequence of the binding reagent sequence, or a portion thereof, as well as a cellular label and a barcode sequence (e.g., a molecular label). As another example, an mRNA molecule can be copied and linked (e.g., covalently linked) to the sequence of the mRNA molecule, or a portion thereof, as well as a cellular label and a barcode sequence (e.g., a molecular label). In some embodiments, molecular labels can be added by ligation of the oligonucleotide probe target-binding region with an oligonucleotide that is (e.g., currently or previously associated with) a portion of the target nucleic acid molecule and / or a cellular component-binding reagent. For example, the target-binding region can include a nucleic acid sequence that can specifically hybridize to a restriction site overhang (e.g., an EcoRI sticky end overhang). The method can further include treating the oligonucleotide associated with the target nucleic acid and / or cellular component-binding reagent with a restriction enzyme (e.g., EcoRI) to generate the restriction site overhang. A ligase (e.g., T4 DNA ligase) can be used to join the two fragments.
[0232] Determining the number or presence of unique molecular signature sequences In some embodiments, the methods disclosed herein include determining the number or presence of unique molecular label sequences for each unique identifier, each nucleic acid target molecule, and / or each binding reagent oligonucleotide (e.g., antibody oligonucleotide). For example, sequencing reads can be used to determine the number of unique molecular label sequences for each unique identifier, each nucleic acid target molecule, and / or each binding reagent oligonucleotide. As another example, sequencing reads can be used to determine the presence or absence of molecular label sequences (e.g., molecular label sequences associated with targets in sequencing reads, binding reagent oligonucleotides, antibody oligonucleotides, sample-indexing oligonucleotides, cell-identification oligonucleotides, control particle oligonucleotides, control oligonucleotides, interaction-determining oligonucleotides, etc.).
[0233] In some embodiments, the number of unique molecular label sequences for each unique identifier, each nucleic acid target molecule, and / or each binding reagent oligonucleotide indicates the amount of each cellular component target (e.g., antigen target or protein target) and / or each nucleic acid target molecule in the sample. In some embodiments, the amount of a cellular component target and the amount of its corresponding nucleic acid target molecule, e.g., mRNA molecule, can be compared to each other. In some embodiments, the ratio of the amount of a cellular component target to the amount of its corresponding nucleic acid target molecule, e.g., mRNA molecule, can be calculated. The cellular component target can be, for example, a cell surface protein marker. In some embodiments, the ratio between the protein level of the cell surface protein marker and the mRNA level of the cell surface protein marker is low.
[0234] The methods disclosed herein can be used for a variety of applications. For example, the methods disclosed herein can be used for proteome and / or transcriptome analysis of a sample. In some embodiments, the methods disclosed herein can be used to identify cellular component targets and / or nucleic acid targets, i.e., biomarkers, in a sample. In some embodiments, the cellular component targets and nucleic acid targets correspond to each other, i.e., the nucleic acid target encodes the cellular component target. In some embodiments, the methods disclosed herein can be used to identify cellular component targets that have a desired ratio between the amount of the cellular component target and the amount of its corresponding nucleic acid target molecule, e.g., mRNA molecule, in a sample. In some embodiments, the ratio is 0.001, 0.01, 0.1, 1, 10, 100, 1000, or a number or range between any two of the above values, or is about 0.001, 0.01, 0.1, 1, 10, 100, 1000, or a number or range between any two of the above values. In some embodiments, the ratio is at least or at most 0.001, 0.01, 0.1, 1, 10, 100, or 1000. In some embodiments, the methods disclosed herein can be used to identify cellular component targets in a sample whose corresponding nucleic acid target molecule abundance in the sample is at or about 1000, 100, 10, 5, 2, 1, 0, or a number or range between any two of these values. In some embodiments, the methods disclosed herein can be used to identify cellular component targets in a sample whose corresponding nucleic acid target molecule abundance in the sample is greater than or less than 1000, 100, 10, 5, 2, 1, or 0.
[0235] Compositions and Kits Some embodiments disclosed herein provide kits and compositions for the simultaneous quantitative analysis of multiple cellular components (e.g., proteins) and / or multiple nucleic acid target molecules in a sample. In some embodiments, the kits and compositions can include multiple cellular component-binding reagents (e.g., multiple protein-binding reagents), each conjugated to an oligonucleotide, where the oligonucleotide comprises a unique identifier for the cellular component-binding reagent and multiple oligonucleotide probes, each of which comprises a target-binding region, a barcode sequence (e.g., a molecular label sequence), and the barcode sequence is derived from a diverse set of unique barcode sequences. In some embodiments, each of the oligonucleotides can comprise a molecular label, a cell label, a sample label, or any combination thereof. In some embodiments, each of the oligonucleotides can comprise a linker. In some embodiments, each of the oligonucleotides can comprise a binding site for the oligonucleotide probe, e.g., a poly(A) tail. For example, the poly(A) tail can be, for example, an oligodA 18 (not tethered to a solid support) or oligoA 18 V (anchored to a solid support). The oligonucleotide may comprise DNA residues, RNA residues, or both.
[0236] The present disclosure includes a plurality of sample indexing compositions. Each of the plurality of sample indexing compositions may include two or more cellular component binding reagents. Each of the two or more cellular component binding reagents may be associated with a sample indexing oligonucleotide. At least one of the two or more cellular component binding reagents may be capable of specifically binding to at least one cellular component target. The sample indexing oligonucleotide may include a sample indexing sequence for identifying the sample origin of one or more cells of the sample. The sample indexing sequences of at least two sample indexing compositions of the plurality of sample indexing compositions may include different sequences.
[0237] The present disclosure includes a kit including a sample indexing composition for identifying cells. In some embodiments, each of two sample indexing compositions includes a cellular component binding reagent (e.g., a protein binding reagent) associated with a sample indexing oligonucleotide, wherein the cellular component binding reagent is capable of specifically binding to at least one of one or more cellular component targets (e.g., one or more protein targets), and the sample indexing oligonucleotide includes a sample indexing sequence, and the sample indexing sequences of the two sample indexing compositions include different sequences. In some embodiments, the sample indexing oligonucleotide includes a molecular beacon sequence, a binding site for a universal primer, or a combination thereof.
[0238] The present disclosure includes a kit for identifying cells. In some embodiments, the kit includes two or more sample indexing compositions. Each of the two or more sample indexing compositions includes a cellular component binding reagent (e.g., an antigen-binding reagent) associated with a sample indexing oligonucleotide, where the cellular component binding reagent is capable of specifically binding to at least one of one or more cellular component targets, and the sample indexing oligonucleotide includes a sample indexing sequence, where the sample indexing sequences of the two sample indexing compositions include different sequences. In some embodiments, the sample indexing oligonucleotide includes a molecular beacon sequence, a binding site for a universal primer, or a combination thereof. The present disclosure includes a kit for multiplex identification. In some embodiments, the kit includes two sample indexing compositions. Each of the two sample indexing compositions comprises a cellular component binding reagent (e.g., an antigen binding reagent) associated with a sample indexing oligonucleotide, the antigen binding reagent being capable of specifically binding to at least one of one or more cellular component targets (e.g., antigen targets), the sample indexing oligonucleotide comprising a sample indexing sequence, and the sample indexing sequences of the two sample indexing compositions comprise different sequences.
[0239] The unique identifier (or an oligonucleotide associated with a cellular component binding reagent, such as a binding reagent oligonucleotide, antibody oligonucleotide, sample indexing oligonucleotide, cell discrimination oligonucleotide, control particle oligonucleotide, control oligonucleotide, or interaction determining oligonucleotide) can have any suitable length, for example, from about 25 nucleotides to about 45 nucleotides in length.In some embodiments, the unique identifier is 4 nucleotides, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 15 nucleotides, 20 nucleotides, 25 nucleotides, 30 nucleotides, 35 nucleotides, 40 nucleotides, 45 nucleotides, 50 nucleotides, 55 nucleotides, 60 nucleotides, 70 nucleotides, 80 nucleotides, 90 nucleotides, 100 nucleotides, 200 nucleotides, or a range between any two of the above values, about 4 nucleotides, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 15 nucleotides, 20 nucleotides, 25 nucleotides, 30 nucleotides, 35 nucleotides, 40 nucleotides, 45 nucleotides, 50 nucleotides, 55 nucleotides, 60 nucleotides, 70 nucleotides, 80 nucleotides, 90 nucleotides, 100 nucleotides, 200 nucleotides, or a range between any two of the above values. and may have a length of less than 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 70, 80, 90, 100, 200 nucleotides, or a range between any two of the above values, and more than 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 70, 80, 90, 100, 200 nucleotides, or a range between any two of the above values.
[0240] In some embodiments, the unique identifiers are selected from a diverse set of unique identifiers. The diverse set of unique identifiers may include 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 5000, or a number or range between any two of these values, or may include about 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 5000, or a number or range between any two of these values. A diverse set of unique identifiers may include at least 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, or 5000 different unique identifiers, or may include up to 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, or 5000 different unique identifiers. In some embodiments, the set of unique identifiers is designed to have minimal sequence homology to the DNA or RNA sequences of the sample being analyzed. In some embodiments, the sequences of a set of unique identifiers differ from each other by, or are complements of, 1 nucleotide, 2 nucleotides, 3 nucleotides, 4 nucleotides, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, or a number or range between any two of these values, or by about 1 nucleotide, 2 nucleotides, 3 nucleotides, 4 nucleotides, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, or a number or range between any two of these values. In some embodiments, the sequences of a set of unique identifiers differ from each other by, or are complements of, at least or by at most 1 nucleotide, 2 nucleotides, 3 nucleotides, 4 nucleotides, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, or 10 nucleotides.
[0241] In some embodiments, the unique identifier may comprise a binding site for a primer, for example, a universal primer. In some embodiments, the unique identifier may comprise at least two binding sites for a primer, for example, a universal primer. In some embodiments, the unique identifier may comprise at least three binding sites for a primer, for example, a universal primer. The primers may be used to amplify the unique identifier, for example, by PCR amplification. In some embodiments, the primers may be used for nested PCR reactions.
[0242] Any suitable cellular component-binding reagent is contemplated in the present disclosure, such as any protein-binding reagent (e.g., an antibody or fragment thereof, an aptamer, a small molecule, a ligand, a peptide, an oligonucleotide, etc., or any combination thereof). In some embodiments, the cellular component-binding reagent may be a polyclonal antibody, a monoclonal antibody, a recombinant antibody, a single-chain antibody (scAb), or a fragment thereof, such as a Fab, Fv, etc. In some embodiments, the plurality of protein binding reagents may include 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 5000 different protein binding reagents, or a number or range between any two of these values, or may include about 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 5000 different protein binding reagents, or a number or range between any two of these values. In some embodiments, the plurality of protein binding reagents may include at least 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, or 5000 different protein binding reagents, or may include up to 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, or 5000 different protein binding reagents.
[0243] In some embodiments, the oligonucleotide is conjugated to the cellular component-binding reagent via a linker. In some embodiments, the oligonucleotide may be covalently conjugated to the protein-binding reagent. In some embodiments, the oligonucleotide may be non-covalently conjugated to the protein-binding reagent. In some embodiments, the linker may contain a chemical group that reversibly or irreversibly attaches the oligonucleotide to the protein-binding reagent. The chemical group may be conjugated to the linker via, for example, an amine group. In some embodiments, the linker may contain a chemical group that forms a stable bond with another chemical group conjugated to the protein-binding reagent. For example, the chemical group may be a UV photocleavable group, a disulfide bond, streptavidin, biotin, an amine, or the like. In some embodiments, the chemical group may be conjugated to the protein-binding reagent via a primary amine on an amino acid such as lysine or the N-terminus. The oligonucleotide may be conjugated to any suitable site on the protein-binding reagent, as long as it does not interfere with specific binding between the protein-binding reagent and its protein target. In embodiments where the protein binding reagent is an antibody, the oligonucleotide may be an antibody that binds to the antigen binding site, e.g., the Fc region, C H 1 domain, C H 2 domains, C H 3 domains, C LThe protein-binding reagents may be conjugated to any site on an antibody other than the antibody domain, etc. In some embodiments, each protein-binding reagent may be conjugated to a single oligonucleotide molecule. In some embodiments, each protein-binding reagent may be conjugated to 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000, or a number or range between any two of these values, or to about 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000, or a number or range between any two of these values, each of the oligonucleotide molecules comprising the same unique identifier. In some embodiments, each protein-binding reagent may be conjugated to two or more oligonucleotide molecules, for example, at least or up to 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, or 1000 oligonucleotide molecules, each of the oligonucleotide molecules comprising the same unique identifier.
[0244] In some embodiments, the multiple cellular component binding reagents (e.g., protein binding reagents) are capable of specifically binding to multiple cellular component targets (e.g., protein targets) in a sample. The sample may be or may include a single cell, multiple cells, a tissue sample, a tumor sample, a blood sample, etc. In some embodiments, the multiple cellular component targets include cell surface proteins, cell markers, B cell receptors, T cell receptors, antibodies, major histocompatibility complexes, tumor antigens, receptors, or any combination thereof. In some embodiments, the multiple cellular component targets may include intracellular proteins. In some embodiments, the multiple cellular component targets may include intracellular proteins. In some embodiments, the multiple cellular component targets may represent 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, 99%, or a number or range between any two of these values, of all cellular component targets (e.g., expressed or expressible proteins) in the organism, or may represent about 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, 99%, or a number or range between any two of these values. In some embodiments, the plurality of cellular component targets may represent at least or up to 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, or 99% of all cellular component targets (e.g., expressed or expressible proteins) in the organism. In some embodiments, the plurality of cellular component targets may include 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000, 10,000, or a number or range between any two of these values, or may include about 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000, 10,000, or a number or range between any two of these values, of different cellular component targets.In some embodiments, the multiple cellular component targets may include at least 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000, or 10,000 different cellular component targets, or may include up to 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000, or 10,000 different cellular component targets.
[0245] Sample indexing using oligonucleotide-conjugated cellular component binding reagents The present disclosure includes a method for identifying a sample. In some embodiments, the method includes: contacting one or more cells from each of a plurality of samples with a sample indexing composition from a plurality of sample indexing compositions, wherein each of the one or more cells comprises one or more cellular component targets, and each of the plurality of sample indexing compositions comprises a cellular component binding reagent associated with a sample indexing oligonucleotide, wherein the cellular component binding reagent is capable of specifically binding to at least one of the one or more cellular component targets, the sample indexing oligonucleotide comprising a sample indexing sequence, and the sample indexing sequences of at least two of the plurality of sample indexing compositions comprise different sequences; barcoding (e.g., stochastically barcoding) the sample indexing oligonucleotides using a plurality of barcodes (e.g., stochastic barcodes) to generate a plurality of barcoded sample indexing oligonucleotides; obtaining sequencing data for the plurality of barcoded sample indexing oligonucleotides; and identifying the sample origin of at least one cell from the one or more cells based on the sample indexing sequence of at least one barcoded sample indexing oligonucleotide from the plurality of barcoded sample indexing oligonucleotides.
[0246] In some embodiments, barcoding the sample-indexing oligonucleotides using a plurality of barcodes includes contacting the plurality of barcodes with the sample-indexing oligonucleotides to generate barcodes hybridized to the sample-indexing oligonucleotides, and extending the barcodes hybridized to the sample-indexing oligonucleotides to generate a plurality of barcoded sample-indexing oligonucleotides. Extending the barcodes may include extending the barcodes using a DNA polymerase to generate a plurality of barcoded sample-indexing oligonucleotides. Extending the barcodes may include extending the barcodes using a reverse transcriptase to generate a plurality of barcoded sample-indexing oligonucleotides.
[0247] An oligonucleotide conjugated to an antibody, an oligonucleotide for conjugation to an antibody, or an oligonucleotide previously conjugated to an antibody is referred to herein as an antibody oligonucleotide ("AbOligo"). An antibody oligonucleotide in the context of sample indexing is referred to herein as a sample indexing oligonucleotide. An antibody conjugated to an antibody oligonucleotide is referred to herein as a hot antibody or oligonucleotide antibody. An antibody not conjugated to an antibody oligonucleotide is referred to herein as a cold antibody or oligonucleotide-free antibody. An oligonucleotide conjugated to a binding reagent (e.g., a protein binding reagent), an oligonucleotide for conjugation to a binding reagent, or an oligonucleotide previously conjugated to a binding reagent is referred to herein as a reagent oligonucleotide. A reagent oligonucleotide in the context of sample indexing is referred to herein as a sample indexing oligonucleotide. A binding reagent conjugated to an antibody oligonucleotide is referred to herein as a hot binding reagent or an oligonucleotide binding reagent. A binding reagent that is not conjugated to an antibody oligonucleotide is referred to herein as a cold binding reagent or an oligonucleotide-free binding reagent.
[0248] 7 shows a schematic diagram of an exemplary workflow using cellular component binding reagents associated with oligonucleotides for sample indexing. In some embodiments, multiple compositions 705a, 705b, etc. are provided, each comprising a binding reagent. The binding reagent may be a protein binding reagent, such as an antibody. The cellular component binding reagent may include an antibody, a tetramer, an aptamer, a protein scaffold, or a combination thereof. The binding reagents of the multiple compositions 705a, 705b may bind to the same cellular component target. For example, the binding reagents of the multiple compositions 705a, 705b may be identical (except for the sample-indexing oligonucleotide associated with the binding reagent).
[0249] The different compositions can comprise binding reagents conjugated to sample-indexing oligonucleotides having different sample-indexing sequences. The number of different compositions can vary in different implementations. In some embodiments, the number of different compositions is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, or a number or range between any two of these values. or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, or a number or range between any two of these values. In some embodiments, the number of different compositions may be at least or up to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10,000.
[0250] In some embodiments, the sample-indexing oligonucleotides of the binding reagents in a single composition may contain the same sample-indexing sequence. The sample-indexing oligonucleotides of the binding reagents in a single composition do not have to be identical. In some embodiments, the percentage of sample-indexing oligonucleotides of binding reagents in a composition that have identical sample-indexing sequences is 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.9%, or it may be a number or range between any two of these values, or may be about 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.9%, or a number or range between any two of these values. In some embodiments, the percentage of sample-indexing oligonucleotides of binding reagents in a composition that have identical sample-indexing sequences may be at least or up to 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.9%.
[0251] Compositions 705a and 705b can be used to label samples from various samples. For example, the sample indexing oligonucleotide of the cellular component binding reagent in composition 705a may have one sample indexing sequence and be used to label cell 710a, shown as a black circle, in sample 707a, such as a patient sample. The sample indexing oligonucleotide of the cellular component binding reagent in composition 705b may have another sample indexing sequence and be used to label cell 710b, shown as a shaded circle, in sample 707b, such as a sample from another patient or another sample from the same patient. The cellular component binding reagent may specifically bind to a cellular component target or protein on the cell surface, such as a cell marker, a B cell receptor, a T cell receptor, an antibody, a major histocompatibility complex, a tumor antigen, a receptor, or any combination thereof. Unbound cellular component binding reagent can be removed, for example, by washing the cells with a buffer solution.
[0252] The cells bearing the cellular component binding reagent can then be separated into multiple compartments, such as a microwell array, where each compartment 715a, 715b is sized to fit a single cell 710a and a single bead 720a or a single cell 710b and a single bead 720b. Each bead 720a, 720b can contain multiple oligonucleotide probes, which can include a cell label and a molecular label sequence common to all oligonucleotide probes on the bead. In some embodiments, each oligonucleotide probe can include a target binding region, e.g., a poly(dT) sequence. The sample indexing oligonucleotide 725a conjugated to the cellular component binding reagent of composition 705a can be configured to be (or can be) detachable or non-detachable from the cellular component binding reagent. The sample indexing oligonucleotide 725a conjugated to the cellular component binding reagent of composition 705a can be detached from the cellular component binding reagent using chemical, optical, or other means. The sample indexing oligonucleotide 725b conjugated to the cellular component binding reagent of composition 705b can be configured to be (or can be) detachable or non-detachable from the cellular component binding reagent. The sample indexing oligonucleotide 725b conjugated to the cellular component binding reagent of composition 705b can be detached from the cellular component binding reagent using chemical, optical, or other means.
[0253] Cell 710a can be lysed to release nucleic acids within cell 710a, such as genomic DNA or cellular mRNA 730a. Lysed cell 735a is shown as a dashed circle. Cellular mRNA 730a, sample-indexing oligonucleotide 725a, or both can be captured by oligonucleotide probes on beads 720a, for example, by hybridizing to poly(dT) sequences. Reverse transcriptase can be used to extend the oligonucleotide probes hybridized to cellular mRNA 730a and oligonucleotide 725a using cellular mRNA 730a and oligonucleotide 725a as templates. The extension products generated by reverse transcriptase can be subjected to amplification and sequencing.
[0254] Similarly, cells 710b can be lysed to release nucleic acids within the cells 710b, such as genomic DNA or cellular mRNA 730b. Lysed cells 735b are shown as dashed circles. The cellular mRNA 730b, sample-indexing oligonucleotides 725b, or both can be captured by oligonucleotide probes on beads 720b, for example, by hybridizing to poly(dT) sequences. Using reverse transcriptase, the oligonucleotide probes hybridized to the cellular mRNA 730b and oligonucleotides 725b can be extended using the cellular mRNA 730b and oligonucleotides 725b as templates. The extension products generated by the reverse transcriptase can be subjected to amplification and sequencing.
[0255] The sequencing reads can be subjected to demultiplexing of cell labels, molecular labels, gene identities, and sample identities (e.g., with respect to the sample-indexing sequences of sample-indexing oligonucleotides 725a and 725b). Demultiplexing of cell labels, molecular labels, and gene identities can provide a digital representation of the gene expression of each single cell in the sample. Demultiplexing of cell labels, molecular labels, and sample identities using the sample-indexing sequences of the sample-indexing oligonucleotides can be used to determine sample origin.
[0256] In some embodiments, cellular component-binding reagents for cell surface cellular component-binding reagents are conjugated to a library of unique sample indexing oligonucleotides, allowing cells to retain their sample identity. For example, antibodies to cell surface markers are conjugated to a library of unique sample indexing oligonucleotides, allowing cells to retain their sample identity. This allows multiple samples to be loaded onto the same Rhapsody™ cartridge, because information about the sample source is retained throughout library preparation and sequencing. Sample indexing can allow multiple samples to be run together in a single experiment, simplifying and shortening experiment time and eliminating batch effects.
[0257] The present disclosure includes a method for identifying a sample. In some embodiments, the method includes: contacting one or more cells from each of a plurality of samples with a sample indexing composition from a plurality of sample indexing compositions, wherein each of the one or more cells comprises one or more cellular component targets, and each of the plurality of sample indexing compositions comprises a cellular component binding reagent associated with a sample indexing oligonucleotide, wherein the cellular component binding reagent is capable of specifically binding to at least one of the one or more cellular component targets, the sample indexing oligonucleotide comprises a sample indexing sequence, and the sample indexing sequences of at least two of the plurality of sample indexing compositions comprise different sequences; and removing unbound sample indexing composition from the plurality of sample indexing compositions. The method may include barcoding (e.g., stochastically barcoding) sample indexing oligonucleotides using a plurality of barcodes (e.g., stochastic barcodes) to generate a plurality of barcoded sample indexing oligonucleotides; obtaining sequencing data of the plurality of barcoded sample indexing oligonucleotides; and identifying a sample origin of at least one cell of the one or more cells based on the sample indexing sequence of at least one barcoded sample indexing oligonucleotide of the plurality of barcoded sample indexing oligonucleotides.
[0258] In some embodiments, a method for identifying a sample includes contacting one or more cells from each of a plurality of samples with a sample indexing composition of a plurality of sample indexing compositions, wherein each of the one or more cells comprises one or more cellular component targets, and each of the plurality of sample indexing compositions comprises a cellular component binding reagent associated with a sample indexing oligonucleotide, wherein the cellular component binding reagent is capable of specifically binding to at least one of the one or more cellular component targets, and the sample indexing oligonucleotide comprises a sample indexing sequence, and the sample indexing sequences of at least two sample indexing compositions of the plurality of sample indexing compositions comprise different sequences; removing unbound sample indexing composition of the plurality of sample indexing compositions; and identifying the sample origin of at least one cell of the one or more cells based on the sample indexing sequence of at least one sample indexing oligonucleotide of the plurality of sample indexing compositions.
[0259] In some embodiments, identifying the sample origin of the at least one cell comprises barcoding (e.g., probabilistically barcoding) sample indexing oligonucleotides of a plurality of sample indexing compositions using a plurality of barcodes (e.g., probabilistic barcodes) to generate a plurality of barcoded sample indexing oligonucleotides; obtaining sequencing data of the plurality of barcoded sample indexing oligonucleotides; and identifying the sample origin of the cell based on the sample-indexing sequence of at least one barcoded sample indexing oligonucleotide of the plurality of barcoded sample indexing oligonucleotides. In some embodiments, barcoding the sample indexing oligonucleotides using a plurality of barcodes to generate a plurality of barcoded sample indexing oligonucleotides comprises probabilistically barcoding the sample indexing oligonucleotides using a plurality of stochastic barcodes to generate a plurality of stochastically barcoded sample indexing oligonucleotides.
[0260] In some embodiments, identifying the sample origin of the at least one cell may include identifying the presence or absence of a sample indexing sequence of at least one sample indexing oligonucleotide of the plurality of sample indexing compositions. Identifying the presence or absence of a sample indexing sequence includes replicating at least one sample indexing oligonucleotide to generate a plurality of replicate sample indexing oligonucleotides; obtaining sequencing data of the plurality of replicate sample indexing oligonucleotides; and identifying the sample origin of the cell based on the sample indexing sequence of the replicate sample indexing oligonucleotide of the plurality of sample indexing oligonucleotides that corresponds to the at least one barcoded sample indexing oligonucleotide in the sequencing data.
[0261] In some embodiments, replicating at least one sample indexing oligonucleotide to generate a plurality of replicate sample indexing oligonucleotides comprises ligating a replication adapter to the at least one barcoded sample indexing oligonucleotide prior to replicating the at least one barcoded sample indexing oligonucleotide. Replicating at least one barcoded sample indexing oligonucleotide may comprise replicating the at least one barcoded sample indexing oligonucleotide using a replication adapter ligated to the at least one barcoded sample indexing oligonucleotide to generate a plurality of replicate sample indexing oligonucleotides.
[0262] In some embodiments, replicating at least one sample indexing oligonucleotide to generate a plurality of replicate sample indexing oligonucleotides comprises contacting a capture probe with the at least one sample indexing oligonucleotide to generate a capture probe hybridized to the sample indexing oligonucleotide before replicating the at least one barcoded sample indexing oligonucleotide, and extending the capture probe hybridized to the sample indexing oligonucleotide to generate a sample indexing oligonucleotide associated with the capture probe. Replicating at least one sample indexing oligonucleotide may comprise replicating the sample indexing oligonucleotide associated with the capture probe to generate a plurality of replicate sample indexing oligonucleotides.
[0263] Cell overloading and multiplet discrimination The present disclosure also includes methods, kits, and systems for identifying cellular overloading and multiplets, which may be used in or in combination with any suitable methods, kits, and systems disclosed herein, such as, for example, methods, kits, and systems for measuring cellular component expression levels (e.g., protein expression levels) using oligonucleotide-associated cellular component-binding reagents.
[0264] Using current cell loading technology, when approximately 20,000 cells are loaded into a microwell cartridge or array with approximately 60,000 microwells, the number of microwells or droplets with more than one cell (referred to as doublets or multiplets) may be minimal. However, as the number of loaded cells increases, the number of microwells or droplets with multiple cells may increase significantly. For example, when approximately 50,000 cells are loaded into the approximately 60,000 microwells of a microwell cartridge or array, the percentage of microwells with multiple cells may be quite high, e.g., 11–14%. Loading such a large number of cells into microwells is sometimes referred to as cell overloading. However, if cells are divided into several groups (e.g., 5) and the cells in each group are labeled with sample-indexing oligonucleotides with distinct sample-indexing sequences, cell labels associated with more than one sample-indexing sequence (e.g., barcode cell labels, such as stochastic barcodes) can be identified in the sequencing data and removed from subsequent processing. In some embodiments, when cells are divided into multiple groups (e.g., 10,000), and cells in each group are labeled with sample-indexing oligonucleotides having distinct sample-indexing sequences, sample labels associated with two or more sample-indexing sequences can be identified in the sequencing data and removed from subsequent processing. In some embodiments, various cells are labeled with cell-identifying oligonucleotides having distinct cell-identifying sequences, and cell-identifying sequences associated with two or more cell-identifying oligonucleotides can be identified in the sequencing data and removed from subsequent processing. Such a large number of cells can be loaded into the microwells relative to the number of microwells in the microwell cartridge or array.
[0265] The disclosure herein includes a method for identifying a sample. In some embodiments, the method includes contacting a first plurality of cells and a second plurality of cells with two sample indexing compositions, respectively, where each of the first plurality of cells and each of the second plurality of cells comprises o...
Claims
1. 1. A method for identifying a sample, comprising: contacting each of a plurality of samples, respectively, with a sample indexing composition of a plurality of sample indexing compositions; each of the plurality of samples comprises one or more cells each comprising one or more cellular component targets, the sample indexing composition comprises an aptamer composition comprising an aptamer and a sample indexing oligonucleotide, the aptamer capable of specifically binding to at least one of the one or more cellular component targets; the sample indexing oligonucleotide comprises a sequence complementary to a capture sequence configured to capture a sequence of the sample indexing oligonucleotide, the sequence of the sample indexing oligonucleotide complementary to the capture sequence comprises a poly(dA) region; contacting, wherein the sample indexing oligonucleotide comprises a sample indexing sequence, and the sample indexing sequences of at least two sample indexing compositions of the plurality of sample indexing compositions comprise different sequences; and identifying a sample origin of at least one cell of said one or more cells based on a sample indexing sequence of at least one sample indexing oligonucleotide of said plurality of sample indexing compositions. The method includes:
2. (a) identifying the sample origin of the at least one cell by barcoding sample indexing oligonucleotides of said plurality of sample indexing compositions with a plurality of barcodes to generate a plurality of barcoded sample indexing oligonucleotides; obtaining sequencing data for said plurality of barcoded sample-indexing oligonucleotides; and identifying a sample origin of the cell based on a sample indexing sequence of at least one barcoded sample indexing oligonucleotide of the plurality of barcoded sample indexing oligonucleotides in the sequencing data. Including, (b) identifying the sample origin of the at least one cell comprises identifying the presence or absence of a sample indexing sequence of at least one sample indexing oligonucleotide of the plurality of sample indexing compositions; Optionally, identifying the presence or absence of said sample indexing sequence comprises: replicating said at least one sample indexing oligonucleotide to generate a plurality of replicate sample indexing oligonucleotides; obtaining sequencing data for said plurality of replicate sample indexing oligonucleotides; and identifying a sample origin of the cell based on a sample indexing sequence of a replicate sample indexing oligonucleotide of the plurality of sample indexing oligonucleotides that corresponds to the at least one barcoded sample indexing oligonucleotide in the sequencing data. may include Optionally, replicating said at least one sample indexing oligonucleotide to generate said plurality of replicate sample indexing oligonucleotides comprises: (i) ligating a replication adapter to the at least one barcoded sample indexing oligonucleotide prior to replicating the at least one barcoded sample indexing oligonucleotide, wherein replicating the at least one barcoded sample indexing oligonucleotide comprises replicating the at least one barcoded sample indexing oligonucleotide using the replication adapter ligated to the at least one barcoded sample indexing oligonucleotide to generate the plurality of replicate sample indexing oligonucleotides; or (ii) prior to replicating the at least one barcoded sample indexing oligonucleotide, contacting a capture probe with the at least one sample indexing oligonucleotide to generate a capture probe hybridized to the sample indexing oligonucleotide, and extending the capture probe hybridized to the sample indexing oligonucleotide to generate a sample indexing oligonucleotide associated with the capture probe, wherein replicating the at least one sample indexing oligonucleotide comprises replicating the sample indexing oligonucleotide associated with the capture probe to generate the plurality of replicate sample indexing oligonucleotides; may include, The method of claim 1.
3. (a) the cellular component target comprises a protein target; (b) the sample indexing sequence is between 6 and 60 nucleotides in length or between 50 and 500 nucleotides in length; (c) the sample indexing sequences of at least 10, 100, or 1000 sample indexing compositions of said plurality of sample indexing compositions comprise different sequences; (d) the single polynucleotide comprises said sample-indexing oligonucleotide and said aptamer, optionally wherein said aptamer comprises: (i) may be 5' to the sample-indexing oligonucleotide in the single polynucleotide; or (ii) may be located 3' to the sample-indexing oligonucleotide in the single polynucleotide; and / or (e) the sample indexing oligonucleotide is associated with the aptamer; (f) said sample indexing oligonucleotide comprises: (i) attached to said aptamer, and optionally said sample-indexing oligonucleotide is covalently attached to said aptamer; or (ii) conjugated to the aptamer, optionally the sample indexing oligonucleotide being conjugated to the aptamer via a chemical group selected from the group consisting of a UV light cleavable group, streptavidin, biotin, an amine, and combinations thereof; and / or (g) the sample indexing oligonucleotide is non-covalently attached to the aptamer; (h) the sample indexing oligonucleotide is attached to the aptamer via a linker; and / or (i) the aptamer (A) a nucleotide aptamer, optionally the nucleotide aptamer may comprise a deoxyribonucleic acid (DNA), a ribonucleic acid (RNA), a xenonucleic acid (XNA), a fluorophore, or a combination thereof, optionally the nucleotide aptamer may comprise a base analog, optionally the base analog may comprise a fluorescent base analog; or (B) comprising a peptide aptamer; The method according to any one of claims 1 to 2.
4. (a) the sample indexing composition is associated with a first carrier; (b) the aptamer composition is associated with a first carrier; and / or (c) the aptamer composition comprises a second aptamer capable of specifically binding to at least one of the one or more protein targets or the one or more cellular component targets, and a second sample indexing oligonucleotide comprising a second sample indexing sequence, and optionally (i) the aptamer and the second aptamer may be associated with a first carrier; or (ii) the aptamer is associated with a first support and the second aptamer is associated with a second support; and, optionally, (iii) the aptamer and the second aptamer may have at least 60%, 70%, 80%, 90%, or 95% sequence identity; and / or (iv) the aptamer and the second aptamer are the same, or the aptamer and the second aptamer may be different; and / or (v) the protein target or the cellular component target of the aptamer and the second aptamer may be the same, and optionally, the aptamer and the second aptamer may be capable of binding to different regions of a protein target or a cellular component target; The method according to any one of claims 1 to 3.
5. (a) the first carrier is (i) a metal nanomaterial, which may optionally include metal nanostructures, metal nanoparticles, or a combination thereof; and / or (ii) a gold nanomaterial, which may optionally be a gold nanostructure, a gold nanoparticle, or a combination thereof; and / or (iii) a lysosome, a micelle, a vesicle, a lipid membrane, a lipid bilayer, a lipid monolayer, or a combination thereof Including, (b) the aptamer is attached to the first support, and optionally the aptamer is covalently attached to the first support; (c) the aptamer is conjugated to the first carrier, optionally the aptamer is conjugated to the first carrier via a chemical group selected from the group consisting of a UV light cleavable group, streptavidin, biotin, an amine, and combinations thereof; (d) the aptamer is non-covalently linked to the first support; (e) the aptamer is attached to the first carrier via a linker; and / or (f) the aptamer is immobilized on the first support, partially immobilized on the first support, immobilized within the first support, partially immobilized within the first support, encapsulated within the first support, partially encapsulated within the first support, embedded within the first support, partially embedded within the first support, or a combination thereof. The method according to claim 3 or 4.
6. (a) dissociating the aptamer from the first support, optionally comprising: (i) may occur after barcoding the sample-indexing oligonucleotide; or (ii) may occur prior to barcoding the sample-indexing oligonucleotide; and / or (b) contacting each of the plurality of samples with the sample indexing composition comprises contacting cells of the one or more cells of the sample with the sample indexing composition, and optionally the first carrier may be internalized into the cells, and optionally the internalization may be by endocytosis, pinocytosis, nanopinocytosis, micropinocytosis, phagocytosis, membrane fusion, clathrin-mediated internalization, caveolin-mediated internalization, receptor-dependent internalization, receptor-independent internalization, or a combination thereof; 6. The method of claim 4(c) or 5.
7. (a) said at least one of said one or more protein targets or said one or more cellular component targets is on a cell surface; (b) the method includes removing unbound sample indexing compositions of the plurality of sample indexing compositions; (c) removing the unbound sample indexing composition comprises: (a) washing the one or more cells from each of the plurality of samples with a wash buffer; (b) selecting cells that are bound to at least one aptamer using flow cytometry; or both (a) and (b). (d) the method comprises lysing the one or more cells from each of the plurality of samples. (e) the sample indexing oligonucleotide comprises: (i) configured to be non-detachable from the aptamer; or (ii) configured to be detachable from the aptamer; and / or (f) the method comprises desorbing the sample indexing oligonucleotide from the aptamer, and optionally, desorbing the sample indexing oligonucleotide may comprise desorbing the sample indexing oligonucleotide from the aptamer by UV light cleavage, chemical treatment, heating, enzymatic treatment, or any combination thereof; The method according to any one of claims 1 to 6.
8. (a) the sample-indexing oligonucleotides are not homologous to any genomic sequence of the one or more cells, are homologous to a genomic sequence of a species, or a combination thereof, optionally the species may be a non-mammalian species; (b) a sample of the plurality of samples comprises a plurality of cells, a plurality of single cells, a tissue, a tumor sample, or any combination thereof; (c) the plurality of samples comprises mammalian cells, bacterial cells, viral cells, yeast cells, fungal cells, or any combination thereof; (d) the barcode comprises a target binding region that includes the capture sequence, and optionally the target binding region may include a poly(dT) region; and / or (e) the sample-indexing oligonucleotide comprises an alignment sequence adjacent to the poly(dA) region, and optionally: (a) the alignment sequence may be one or more nucleotides in length or two or more nucleotides in length; (b) the alignment sequence may comprise guanine, cytosine, thymine, uracil, or a combination thereof; and / or (c) the alignment sequence may comprise a poly(dT) region, a poly(dG) region, a poly(dC) region, a poly(dU) region, or a combination thereof; and / or (f) the sample indexing oligonucleotide comprises a molecular beacon sequence, a binding site for a universal primer, or both, and optionally the molecular beacon sequence may be 2-20 nucleotides in length or 5-50 nucleotides in length, and optionally the universal primer may comprise an amplification primer, a sequencing primer, or a combination thereof; (g) the protein target or the cellular component target comprises a carbohydrate, a lipid, a protein, an extracellular protein, a cell surface protein, a cell marker, a B cell receptor, a T cell receptor, a major histocompatibility complex, a tumor antigen, a receptor, an intracellular protein, or any combination thereof; (h) the protein target or the cellular component target is selected from a group comprising 10 to 100 different protein targets or cellular component targets; and / or (i) the aptamer comprises (A) the same sequence, or (B) Different sample indexing sequences and / or (j) the sample indexing composition of the plurality of sample indexing compositions comprises a second aptamer to which the sample indexing oligonucleotide is not associated, and optionally, the aptamer and the second aptamer may be identical; and / or (k) each of said plurality of sample indexing compositions comprises said aptamer; The method according to any one of claims 1 to 7.
9. A plurality of sample indexing compositions, comprising: each of said plurality of sample indexing compositions comprises an aptamer composition comprising a first cellular component-binding aptamer and a sample indexing oligonucleotide; the cellular component-binding aptamer is capable of specifically binding to at least one cellular component target; the sample indexing oligonucleotide comprises a sample indexing sequence for identifying the sample origin of one or more cells of the sample; the sample indexing oligonucleotide comprises a sequence complementary to a capture sequence configured to capture a sequence of the sample indexing oligonucleotide; the sequence of the sample-indexing oligonucleotide that is complementary to the capture sequence comprises a poly(dA) region; the sample indexing sequences of at least two sample indexing compositions of said plurality of sample indexing compositions comprise different sequences; A plurality of sample indexing compositions.
10. (a) the cellular component target comprises a protein target; (b) the sample indexing sequence is between 6 and 60 nucleotides in length or between 50 and 500 nucleotides in length; (c) the sample indexing sequences of at least 10, 100, or 1000 sample indexing compositions of said plurality of sample indexing compositions comprise different sequences; (d) the single polynucleotide comprises said sample-indexing oligonucleotide and said aptamer, optionally wherein said aptamer comprises: (i) may be 5' to the sample-indexing oligonucleotide in the single polynucleotide; or (ii) may be located 3' to the sample-indexing oligonucleotide in the single polynucleotide; and / or (e) the sample indexing oligonucleotide is associated with the aptamer; (f) said sample indexing oligonucleotide comprises: (i) attached to said aptamer, and optionally said sample-indexing oligonucleotide is covalently attached to said aptamer; or (ii) conjugated to the aptamer, optionally the sample indexing oligonucleotide being conjugated to the aptamer via a chemical group selected from the group consisting of a UV light cleavable group, streptavidin, biotin, an amine, and combinations thereof; and / or (g) the sample indexing oligonucleotide is non-covalently attached to the aptamer; (h) the sample indexing oligonucleotide is attached to the aptamer via a linker; and / or (i) the aptamer (A) a nucleotide aptamer, optionally the nucleotide aptamer may comprise a deoxyribonucleic acid (DNA), a ribonucleic acid (RNA), a xenonucleic acid (XNA), a fluorophore, or a combination thereof, optionally the nucleotide aptamer may comprise a base analog, optionally the base analog may comprise a fluorescent base analog; or (B) comprising a peptide aptamer; 10. A multiple sample indexing composition according to claim 9.
11. (a) the sample indexing composition is associated with a first carrier; (b) the aptamer composition is associated with a first carrier; and / or (c) the aptamer composition comprises a second aptamer capable of specifically binding to at least one of the one or more protein targets or the one or more cellular component targets, and a second sample indexing oligonucleotide comprising a second sample indexing sequence, and optionally (i) the aptamer and the second aptamer may be associated with a first carrier; or (ii) the aptamer is associated with a first support and the second aptamer is associated with a second support; and, optionally, (iii) the aptamer and the second aptamer may have at least 60%, 70%, 80%, 90%, or 95% sequence identity; and / or (iv) the aptamer and the second aptamer are the same, or the aptamer and the second aptamer may be different; and / or (v) the protein target or the cellular component target of the aptamer and the second aptamer may be the same, and optionally, the aptamer and the second aptamer may be capable of binding to different regions of a protein target or a cellular component target; 11. A multiple sample indexing composition according to claim 9 or 10.
12. (a) the first carrier is (i) a metal nanomaterial, which may optionally include metal nanostructures, metal nanoparticles, or a combination thereof; and / or (ii) a gold nanomaterial, which may optionally be a gold nanostructure, a gold nanoparticle, or a combination thereof; and / or (iii) a lysosome, a micelle, a vesicle, a lipid membrane, a lipid bilayer, a lipid monolayer, or a combination thereof Including, (b) the aptamer is attached to the first support, and optionally the aptamer is covalently attached to the first support; (c) the aptamer is conjugated to the first carrier, optionally the aptamer is conjugated to the first carrier via a chemical group selected from the group consisting of a UV light cleavable group, streptavidin, biotin, an amine, and combinations thereof; (d) the aptamer is non-covalently linked to the first support; (e) the aptamer is attached to the first carrier via a linker; and / or (f) the aptamer is immobilized on the first support, partially immobilized on the first support, immobilized within the first support, partially immobilized within the first support, encapsulated within the first support, partially encapsulated within the first support, embedded within the first support, partially embedded within the first support, or a combination thereof.
12. A multiple sample indexing composition according to claim 10 or 11.
13. (a) the sample-indexing oligonucleotides are not homologous to any genomic sequence of the one or more cells, are homologous to a genomic sequence of a species, or a combination thereof, optionally the species may be a non-mammalian species; (b) a sample of the plurality of samples comprises a plurality of cells, a plurality of single cells, a tissue, a tumor sample, or any combination thereof; (c) the plurality of samples comprises mammalian cells, bacterial cells, viral cells, yeast cells, fungal cells, or any combination thereof; (d) the barcode comprises a target binding region that includes the capture sequence, and optionally the target binding region may include a poly(dT) region; and / or (e) the sample-indexing oligonucleotide comprises an alignment sequence adjacent to the poly(dA) region, and optionally: (a) the alignment sequence may be one or more nucleotides in length or two or more nucleotides in length; (b) the alignment sequence may comprise guanine, cytosine, thymine, uracil, or a combination thereof; and / or (c) the alignment sequence may comprise a poly(dT) region, a poly(dG) region, a poly(dC) region, a poly(dU) region, or a combination thereof; and / or (f) the sample indexing oligonucleotide comprises a molecular beacon sequence, a binding site for a universal primer, or both, and optionally the molecular beacon sequence may be 2-20 nucleotides in length or 5-50 nucleotides in length, and optionally the universal primer may comprise an amplification primer, a sequencing primer, or a combination thereof; (g) the protein target or the cellular component target comprises a carbohydrate, a lipid, a protein, an extracellular protein, a cell surface protein, a cell marker, a B cell receptor, a T cell receptor, a major histocompatibility complex, a tumor antigen, a receptor, an intracellular protein, or any combination thereof; (h) the protein target or the cellular component target is selected from a group comprising 10 to 100 different protein targets or cellular component targets; and / or (i) the aptamer comprises (A) the same sequence, or (B) Different sample indexing sequences and / or (j) the sample indexing composition of the plurality of sample indexing compositions comprises a second aptamer to which the sample indexing oligonucleotide is not associated, and optionally, the aptamer and the second aptamer may be identical; and / or (k) each of said plurality of sample indexing compositions comprises said aptamer; A multiple sample indexing composition according to any one of claims 9 to 12.
14. 1. A method for measuring cellular component expression in a cell, comprising: contacting a plurality of cellular component-binding aptamers with a plurality of cells comprising a plurality of cellular component targets, each of the plurality of cellular component-binding aptamers comprising an aptamer-specific oligonucleotide comprising a unique identifier sequence for the cellular component-binding aptamer, the cellular component-binding aptamer capable of specifically binding to at least one of the plurality of cellular component targets, the aptamer-specific oligonucleotide comprising a sequence complementary to a capture sequence configured to capture a sequence of the aptamer-specific oligonucleotide, the sequence of the aptamer-specific oligonucleotide complementary to the capture sequence comprising a poly(dA) region; extending oligonucleotide probes hybridized to the aptamer-specific oligonucleotides to produce a plurality of labeled nucleic acids, each of the labeled nucleic acids comprising a unique identifier sequence or its complement and a barcode sequence; and obtaining sequence information of said plurality of labeled nucleic acids or portions thereof to determine the amount of one or more of said plurality of cellular component targets in one or more of said plurality of cells. The method includes:
15. Prior to extending the oligonucleotide probe, distributing the plurality of cells having associated therewith the plurality of cellular component-binding aptamers into a plurality of compartments, a compartment of the plurality of compartments comprising a single cell derived from the plurality of cells having associated therewith the cellular component-binding aptamers; contacting a bar-coded particle with the aptamer-specific oligonucleotide in a compartment containing the single cell, the bar-coded particle comprising a plurality of oligonucleotide probes each comprising a target binding region and a barcode sequence selected from a diverse set of unique barcode sequences.
15. The method of claim 14, comprising:
16. (a) the plurality of cellular component targets comprises a plurality of protein targets, and the cellular component-binding aptamer is capable of specifically binding to at least one of the plurality of protein targets; and / or (b) said aptamer-specific oligonucleotide and said aptamer form a single polynucleotide, and optionally said aptamer comprises: (i) may be located 5' to the aptamer-specific oligonucleotide in the single polynucleotide; or (ii) may be located 3' to the aptamer-specific oligonucleotide in the single polynucleotide; and / or (c) the aptamer-specific oligonucleotide is associated with the aptamer; (d) the aptamer-specific oligonucleotide is attached to the aptamer; and / or (e) the aptamer-specific oligonucleotide (i) is covalently attached to the aptamer; or (ii) is non-covalently attached to the aptamer; and / or (f) the aptamer-specific oligonucleotide is conjugated to the aptamer, and optionally the aptamer-specific oligonucleotide is conjugated to the aptamer via a chemical group selected from the group consisting of a UV light cleavable group, streptavidin, biotin, an amine, and combinations thereof; (g) the aptamer comprises a nucleotide aptamer, optionally the nucleotide aptamer may comprise a deoxyribonucleic acid (DNA), a ribonucleic acid (RNA), a xenonucleic acid (XNA), a fluorophore, or a combination thereof, optionally the nucleotide aptamer may comprise a base analog, optionally the base analog may comprise a fluorescent base analog, and / or (h) the aptamer comprises a peptide aptamer; The method according to any one of claims 14 to 15.
17. The plurality of protein-binding aptamers comprises a second protein-binding aptamer, or the plurality of cellular component-binding aptamers comprises a second cellular component-binding aptamer, and optionally (a)(i) the aptamer and the second aptamer may be associated with a first carrier; or (ii) the aptamer is associated with a first carrier and the second aptamer is optionally associated with a second carrier; and / or (b) the aptamer and the second aptamer may be at least 60%, 70%, 80%, 90%, or 95% identical in sequence; and / or (c) the aptamer and the second protein-binding aptamer (i) may be identical, or (ii) may be different, and optionally (A) the protein target or the cellular component target of the aptamer and the second aptamer may be the same; (B) the aptamer and the second aptamer may be capable of binding to different regions of a protein or cellular component target; or (C) the protein target or the cellular component target of the aptamer and the second aptamer may be different; and / or (d) the first carrier is (i) a metal nanomaterial, which may optionally be a metal nanostructure, a metal nanoparticle, or a combination thereof; and / or (ii) a gold nanomaterial, which may optionally be a gold nanostructure, a gold nanoparticle, or a combination thereof; and / or (iii) a lysosome, a micelle, a vesicle, a lipid membrane, a lipid bilayer, a lipid monolayer, or a combination thereof and / or (e) the aptamer may be attached to the first support, and optionally the aptamer may be covalently attached to the first support; (f) the aptamer may be conjugated to the first carrier, and optionally the aptamer may be conjugated to the first carrier via a chemical group selected from the group consisting of a UV light cleavable group, streptavidin, biotin, an amine, and combinations thereof; (g) the aptamer may be non-covalently linked to the first carrier; (h) the aptamer may be attached to the first carrier via a linker; and / or (i) the aptamer may be immobilized on the first carrier, may be partially immobilized on the first carrier, may be immobilized within the first carrier, may be partially immobilized within the first carrier, may be encapsulated within the first carrier, may be partially encapsulated within the first carrier, may be embedded within the first carrier, may be partially embedded within the first carrier, or may be a combination thereof; The method according to any one of claims 14 to 16.
18. (a) the barcode comprises a target binding region that includes the capture sequence, and optionally the target binding region may comprise a poly(dT) region; and / or (b) the aptamer-specific oligonucleotide changes from a first conformation in which the poly(dA) region is inaccessible to a second conformation in which the poly(dA) region is accessible upon contact of the aptamer with the at least one of the multiple protein or cellular component targets, and optionally the poly(dA) region of the aptamer-specific oligonucleotide in the first conformation may constitute a hairpin structure, and optionally the poly(dA) region with the hairpin structure may be accessible to the poly(dT) region of the target binding region of the barcode; (c) the oligonucleotide probe is hybridized to the aptamer-specific oligonucleotide by hybridization of the poly(dA) region of the aptamer-specific oligonucleotide with the poly(dT) region of the oligonucleotide probe; and / or (d) the aptamer-specific oligonucleotide comprises an alignment sequence adjacent to the poly(dA) region, and optionally: (a) the alignment sequence may be one or more nucleotides in length or two or more nucleotides in length; (b) the alignment sequence may comprise guanine, cytosine, thymine, uracil, or a combination thereof; and / or (c) the alignment sequence may comprise a poly(dT) region, a poly(dG) region, a poly(dC) region, a poly(dU) region, or a combination thereof; and / or (e) the aptamer-specific oligonucleotide comprises a molecular beacon sequence, a binding site for a universal primer, or both, and optionally the molecular beacon sequence may be 2-20 nucleotides in length, or the universal primer may be 5-50 nucleotides in length, or both, and optionally the universal primer may comprise an amplification primer, a sequencing primer, or a combination thereof; 20. The method of claim 17.
19. A composition comprising a plurality of cellular component-binding aptamers, each of the plurality of cellular component-binding aptamers comprising an aptamer-specific oligonucleotide comprising a unique identifier sequence for the cellular component-binding aptamer, the cellular component-binding aptamer capable of specifically binding to at least one of a plurality of cellular component targets, the aptamer-specific oligonucleotide comprising a sequence complementary to a capture sequence configured to capture a sequence of the aptamer-specific oligonucleotide, and the sequence of the aptamer-specific oligonucleotide complementary to the capture sequence comprises a poly(dA) region.
20. (a) the plurality of cellular component targets comprises a plurality of protein targets, and the cellular component-binding aptamer is capable of specifically binding to at least one of the plurality of protein targets; and / or (b) said aptamer-specific oligonucleotide and said aptamer form a single polynucleotide, and optionally said aptamer comprises: (i) may be located 5' to the aptamer-specific oligonucleotide in the single polynucleotide; or (ii) may be located 3' to the aptamer-specific oligonucleotide in the single polynucleotide; and / or (c) the aptamer-specific oligonucleotide is associated with the aptamer; (d) the aptamer-specific oligonucleotide is attached to the aptamer; and / or (e) the aptamer-specific oligonucleotide (i) is covalently attached to the aptamer; or (ii) is non-covalently attached to the aptamer; and / or (f) the aptamer-specific oligonucleotide is conjugated to the aptamer, and optionally the aptamer-specific oligonucleotide is conjugated to the aptamer via a chemical group selected from the group consisting of a UV light cleavable group, streptavidin, biotin, an amine, and combinations thereof; (g) the aptamer comprises a nucleotide aptamer, optionally the nucleotide aptamer may comprise a deoxyribonucleic acid (DNA), a ribonucleic acid (RNA), a xenonucleic acid (XNA), a fluorophore, or a combination thereof, optionally the nucleotide aptamer may comprise a base analog, optionally the base analog may comprise a fluorescent base analog, and / or (h) the aptamer comprises a peptide aptamer; 20. The composition of claim 19.
21. The plurality of protein-binding aptamers comprises a second protein-binding aptamer, or the plurality of cellular component-binding aptamers comprises a second cellular component-binding aptamer, and optionally (a)(i) the aptamer and the second aptamer may be associated with a first carrier; or (ii) the aptamer is associated with a first carrier and the second aptamer is optionally associated with a second carrier; and / or (b) the aptamer and the second aptamer may be at least 60%, 70%, 80%, 90%, or 95% identical in sequence; and / or (c) the aptamer and the second protein-binding aptamer (i) may be identical, or (ii) may be different, and optionally (A) the protein target or the cellular component target of the aptamer and the second aptamer may be the same; (B) the aptamer and the second aptamer may be capable of binding to different regions of a protein or cellular component target; or (C) the protein target or the cellular component target of the aptamer and the second aptamer may be different; and / or (d) the first carrier is (i) a metal nanomaterial, which may optionally be a metal nanostructure, a metal nanoparticle, or a combination thereof; and / or (ii) a gold nanomaterial, which may optionally be a gold nanostructure, a gold nanoparticle, or a combination thereof; and / or (iii) a lysosome, a micelle, a vesicle, a lipid membrane, a lipid bilayer, a lipid monolayer, or a combination thereof and / or (e) the aptamer may be attached to the first support, and optionally the aptamer may be covalently attached to the first support; (f) the aptamer may be conjugated to the first carrier, and optionally the aptamer may be conjugated to the first carrier via a chemical group selected from the group consisting of a UV light cleavable group, streptavidin, biotin, an amine, and combinations thereof; (g) the aptamer may be non-covalently linked to the first carrier; (h) the aptamer may be attached to the first carrier via a linker; and / or (i) the aptamer may be immobilized on the first carrier, may be partially immobilized on the first carrier, may be immobilized within the first carrier, may be partially immobilized within the first carrier, may be encapsulated within the first carrier, may be partially encapsulated within the first carrier, may be embedded within the first carrier, may be partially embedded within the first carrier, or may be a combination thereof; 21. The composition according to claim 19 or 20.
22. (a) the barcode comprises a target binding region that includes the capture sequence, and optionally the target binding region may comprise a poly(dT) region; and / or (b) the aptamer-specific oligonucleotide changes from a first conformation in which the poly(dA) region is inaccessible to a second conformation in which the poly(dA) region is accessible upon contact of the aptamer with the at least one of the multiple protein or cellular component targets, and optionally the poly(dA) region of the aptamer-specific oligonucleotide in the first conformation may constitute a hairpin structure, and optionally the poly(dA) region with the hairpin structure may be accessible to the poly(dT) region of the target binding region of the barcode; (c) the oligonucleotide probe is hybridized to the aptamer-specific oligonucleotide by hybridization of the poly(dA) region of the aptamer-specific oligonucleotide with the poly(dT) region of the oligonucleotide probe; and / or (d) the aptamer-specific oligonucleotide comprises an alignment sequence adjacent to the poly(dA) region, and optionally: (a) the alignment sequence may be one or more nucleotides in length or two or more nucleotides in length; (b) the alignment sequence may comprise guanine, cytosine, thymine, uracil, or a combination thereof; and / or (c) the alignment sequence may comprise a poly(dT) region, a poly(dG) region, a poly(dC) region, a poly(dU) region, or a combination thereof; and / or (e) the aptamer-specific oligonucleotide comprises a molecular beacon sequence, a binding site for a universal primer, or both, and optionally the molecular beacon sequence may be 2-20 nucleotides in length, or the universal primer may be 5-50 nucleotides in length, or both, and optionally the universal primer may comprise an amplification primer, a sequencing primer, or a combination thereof; 22. The composition of claim 21.