Aptomer barcoding

JP7842143B2Active Publication Date: 2026-04-07BECTON DICKINSON & CO
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-05-31
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Current technologies lack the capability to quantitatively analyze protein expression in cells while simultaneously measuring both protein and gene expression in a large-scale manner.

Method used

A method involving the use of aptamer compositions and sample indexing oligonucleotides to barcode cells, allowing for the identification of sample origin and quantification of protein and gene expression by sequencing data analysis.

Benefits of technology

Enables simultaneous and quantitative analysis of protein and gene expression in cells, providing detailed insights into cellular processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007842143000001
    Figure 0007842143000001
  • Figure 0007842143000002
    Figure 0007842143000002
  • Figure 0007842143000003
    Figure 0007842143000003
Patent Text Reader

Abstract

To provide systems, methods, compositions, and kits for sample identification and protein expression profiling.SOLUTION: A composition can comprise, for example, a protein binding aptamer associated with an oligonucleotide, such as a sample indexing oligonucleotide. Different oligonucleotides can have different sequences. Sample origin of cells, or protein expression profiles of cells, can be determined based on the sequences of the oligonucleotides by, for example, barcoding the oligonucleotides.SELECTED DRAWING: Figure 8A
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Related applications This application claims the interests of U.S. Provisional Patent Application No. 62 / 719,406, filed on 17 August 2018, pursuant to Section 119(e) of the U.S. Patent Act. The contents of this related application are incorporated herein by reference in their entirety for all purposes. Sequence listing reference This application is filed together with an electronic sequence listing. The sequence listing is provided as a file named SequenceListing, created on August 2, 2019, with a size of 4 kilobytes. The electronic information of the sequence listing is incorporated herein by reference in its entirety.

[0002] This disclosure generally relates to the field of molecular biology, for example, to identifying cells in different samples and determining protein expression profiles in cells using molecular barcoding. [Background technology]

[0003] Current technology allows for the measurement of single-cell gene expression in a large-scale parallel manner (e.g., >1000 cells) by co-localizing each cell within a compartment with barcoded reagent beads and attaching cell-specific oligonucleotide barcodes to poly(A)mRNA molecules derived from individual cells. Gene expression can influence protein expression. Protein-protein interactions can influence both gene expression and protein expression. There is a need for systems and methods that can quantitatively analyze protein expression in cells and simultaneously measure both protein and gene expression in cells. [Overview of the Initiative]

[0004] The disclosure herein includes embodiments of a method for identifying a sample. In some embodiments, the method includes contacting each of a plurality of samples with a sample indexing composition from a plurality of sample indexing compositions, each of the plurality of samples comprising one or more cells each comprising one or more protein targets, the sample indexing composition comprising an aptamer composition comprising an aptamer and a sample indexing oligonucleotide, wherein the aptamer is capable of specifically binding to at least one of the one or more protein targets, the sample indexing oligonucleotide comprising a sample indexing sequence, wherein the sample indexing sequences of at least two of the plurality of sample indexing compositions comprise different sequences; barcoding the sample indexing oligonucleotides using a plurality of barcodes to generate a plurality of barcoded sample indexing oligonucleotides; obtaining sequencing data for the plurality of barcoded sample indexing oligonucleotides; and identifying the sample origin of at least one of the one or more cells based on the sample indexing sequence of at least one of the barcoded sample indexing oligonucleotides. In some embodiments, the method involves contacting each of a plurality of samples with a sample indexing composition from a plurality of sample indexing compositions, each of the plurality of samples comprising one or more cells each comprising one or more cellular component targets, the sample indexing composition comprising an aptamer composition comprising an aptamer and a sample indexing oligonucleotide, the aptamer being capable of specifically binding to at least one of the one or more cellular component targets, the sample indexing oligonucleotide comprising a sample indexing sequence, the sample indexing sequences of at least two of the plurality of sample indexing compositions comprising different sequences; and identifying the sample origin of at least one of the one or more cells based on the sample indexing sequence of at least one sample indexing oligonucleotide of the plurality of sample indexing compositions. In some embodiments, the cellular component targets comprise protein targets. In some embodiments, the identification of the sample origin of at least one cell includes: barcoding sample indexing oligonucleotides of multiple sample indexing compositions using multiple barcodes to generate multiple barcoded sample indexing oligonucleotides; obtaining sequencing data for the multiple barcoded sample indexing oligonucleotides; and identifying the sample origin of the cell based on the sample indexing sequence of at least one of the multiple barcoded sample indexing oligonucleotides in the sequencing data.

[0005] In some embodiments, the identification of the sample origin of at least one cell includes identifying the presence or absence of a sample indexing sequence in at least one sample indexing oligonucleotide of a plurality of sample indexing compositions. The identification of the presence or absence of a sample indexing sequence includes replicating at least one sample indexing oligonucleotide to produce a plurality of replicated sample indexing oligonucleotides; obtaining sequencing data for the plurality of replicated sample indexing oligonucleotides; and identifying the sample origin of the cell based on the sample indexing sequences of the plurality of replicated sample indexing oligonucleotides corresponding to at least one barcoded sample indexing oligonucleotide in the sequencing data.

[0006] In some embodiments, duplicating at least one sample indexing oligonucleotide to produce multiple replicated sample indexing oligonucleotides includes ligating a replication adapter to at least one barcoded sample indexing oligonucleotide before duplicating at least one barcoded sample indexing oligonucleotide, and duplicating at least one barcoded sample indexing oligonucleotide includes duplicating at least one barcoded sample indexing oligonucleotide using a replication adapter ligated to at least one barcoded sample indexing oligonucleotide to produce multiple replicated sample indexing oligonucleotides. In some embodiments, duplicating at least one sample indexing oligonucleotide to produce multiple replicated sample indexing oligonucleotides includes contacting a capture probe to at least one sample indexing oligonucleotide before duplicating at least one barcoded sample indexing oligonucleotide to produce a capture probe hybridized to the sample indexing oligonucleotide; and extending the capture probe hybridized to the sample indexing oligonucleotide to produce a sample indexing oligonucleotide accompanied by the capture probe. The replication of at least one sample indexing oligonucleotide may include replicating a sample indexing oligonucleotide accompanied by a capture probe to generate multiple replicated sample indexing oligonucleotides.

[0007] In some embodiments, the sample indexing sequence is 6 to 60 nucleotides long. The sample indexing oligonucleotide may be 50 to 500 nucleotides long. The sample indexing sequences of at least 10, 100, or 1000 of the multiple sample indexing compositions may contain different sequences.

[0008] In some embodiments, a single polynucleotide comprises a sample index-granting oligonucleotide and an aptamer. The aptamer may be located at the 5' end of the sample index-granting oligonucleotide in the single polynucleotide. The aptamer may be located at the 3' end of the sample index-granting oligonucleotide in the single polynucleotide. The sample index-granting oligonucleotide may be attached to the aptamer. The sample index-granting oligonucleotide may be covalently attached to the aptamer. The sample index-granting oligonucleotide may be conjugated to the aptamer. The sample index-granting oligonucleotide may be conjugated to the aptamer via a chemical group selected from the group consisting of UV-cleavable groups, streptavidin, biotin, amines, and combinations thereof. In some embodiments, the sample index-granting oligonucleotide may be noncovalently attached to the aptamer. The sample index-granting oligonucleotide may be attached to the aptamer via a linker.

[0009] In some embodiments, the aptamer includes a nucleotide aptamer. The nucleotide aptamer may include deoxyribonucleic acid (DNA), ribonucleic acid (RNA), xenonucleic acid (XNA), or a combination thereof. The nucleotide aptamer may include a base analog. The base analog may include a fluorescent base analog. The nucleotide aptamer may include a fluorophore. The aptamer may include a peptide aptamer. In some embodiments, the sample indexing composition may be attached to the first carrier. Different sample indexing compositions among a plurality of sample indexing compositions may be attached to different first carriers.

[0010] In some embodiments, the aptamer composition may be attached to a first carrier. The aptamer composition may include a second aptamer capable of specifically binding to at least one of one or more protein targets or one or more cellular component targets, and a second sample indexing oligonucleotide containing a second sample indexing sequence. The aptamer and the second aptamer may be attached to a first carrier. The aptamer may be attached to the first carrier, and the second aptamer may be attached to the second carrier. The aptamer and the second aptamer may be identical in sequence by at least 60%, 70%, 80%, 90%, or 95%. The aptamer and the second aptamer may, for example, have identical sequences. The aptamer and the second aptamer may be different. The protein targets or cellular component targets of the aptamer and the second aptamer may be the same. The aptamer and the second aptamer may be capable of binding to different regions of the protein target or cellular component target. The protein target or cellular component target of the aptamer and the second aptamer may be different.

[0011] In some embodiments, the first support comprises a metal nanomaterial. The first support may comprise a metal nanostructure, metal nanoparticles, or a combination thereof. The first support may comprise a gold nanomaterial. The first support may comprise a gold nanostructure, gold nanoparticles, or a combination thereof. The first support may comprise lysosomes, micelles, vesicles, lipid membranes, lipid bilayers, lipid monolayers, or a combination thereof. In some embodiments, the aptamer may be attached to the first support. The aptamer may be covalently attached to the first support. The aptamer may be conjugated to the first support. The aptamer may be conjugated to the first support via a chemical group selected from the group consisting of UV-cleavable groups, streptavidin, biotin, amines, and combinations thereof. The aptamer may be noncovalently linked to the first support. The aptamer may be attached to the first support via a linker. The aptamer may be immobilized on the first support, partially immobilized on the first support, immobilized within the first support, partially immobilized within the first support, encapsulated within the first support, partially encapsulated within the first support, embedded within the first support, partially embedded within the first support, or a combination thereof. In some embodiments, the method includes dissociating the aptamer from the first support. Dissociation may occur after barcoding the sample index-granting oligonucleotide. Dissociation may occur before barcoding the sample index-granting oligonucleotide.

[0012] In some embodiments, contacting each of a plurality of samples with the sample indexing composition includes contacting one or more cells of the sample with the sample indexing composition. The first carrier may be internally transported into the cell. The first carrier may be internally transported into the cell by endocytosis, pinocytosis, nanopinocytosis, micropinocytosis, phagocytosis, membrane fusion, or a combination thereof. The first carrier may be internally transported into the cell by clathrin-mediated internal transport, caveolin-mediated internal transport, receptor-dependent internal transport, receptor-independent internal transport, or a combination thereof. In some embodiments, the method includes removing unbound sample indexing compositions from a plurality of sample indexing compositions. Removal of unbound sample indexing compositions may include washing one or more cells derived from each of the plurality of samples with a washing buffer. Removal of unbound sample indexing compositions may include selecting cells bound to at least one aptamer using flow cytometry. In some embodiments, the method includes lysing one or more cells derived from each of the plurality of samples.

[0013] In some embodiments, the sample index-granting oligonucleotide is configured to be indesorbable from the aptamer (or may be indesorbable). The sample index-granting oligonucleotide may be configured to be desorbable from the aptamer (or may be desorbable). The method may include desorbing the sample index-granting oligonucleotide from the aptamer. Desorbing the sample index-granting oligonucleotide may include desorbing the sample index-granting oligonucleotide from the aptamer by UV cutting, chemical treatment, heating, enzymatic treatment, or any combination thereof. In some embodiments, the sample indexing oligonucleotide is either not homologous to any of the genome sequences of one or more cells, homologous to the genome sequence of a species, or a combination thereof. The species may be a non-mammalian species.

[0014] In some embodiments, the samples among the multiple samples include multiple cells, multiple single cells, tissues, tumor samples, or any combination thereof. The multiple samples may include mammalian cells, bacterial cells, viral cells, yeast cells, fungal cells, or any combination thereof. In some embodiments, the sample indexing oligonucleotide includes a sequence complementary to a capture sequence configured to capture (or capable of capturing, hybridizing, or binding) the sequence of the sample indexing oligonucleotide. The barcode may include a target-binding region containing the capture sequence. The target-binding region may include a poly(dT) region. The sequence of the sample indexing oligonucleotide complementary to the capture sequence may include a poly(dA) region.

[0015] In some embodiments, the sample indexing oligonucleotide includes an alignment sequence adjacent to the poly(dA) region. The alignment sequence may have a nucleotide length of 1 or more. The alignment sequence may have a nucleotide length of 2 or more. The alignment sequence may include guanine, cytosine, thymine, uracil, or a combination thereof. The alignment sequence may include a poly(dT) region, a poly(dG) region, a poly(dC) region, a poly(dU) region, or a combination thereof. The sample indexing oligonucleotide may include a molecular labeling sequence, a universal primer binding site, or both. The molecular labeling sequence may be 2 to 20 nucleotides long. The universal primer may be 5 to 50 nucleotides long. The universal primer may include an amplification primer, a sequencing primer, or a combination thereof.

[0016] In some embodiments, the aptamer is accompanied by two or more sample indexing oligonucleotides having the same sequence. The aptamer may be accompanied by two or more sample indexing oligonucleotides having different sample indexing sequences. In some embodiments, one of the multiple sample indexing compositions includes a second aptamer that is not accompanied by a sample indexing oligonucleotide. The aptamer and the second aptamer may be identical. In some embodiments, each of the multiple sample indexing compositions includes an aptamer.

[0017] In some embodiments, one of a plurality of sample indexing compositions comprises a second aptamer composition containing a second protein-binding aptamer, the second aptamer capable of specifically binding to at least one of one or more protein targets or to at least one of one or more cellular component targets. The aptamer and the second aptamer may be capable of binding to the same protein target of one or more protein targets or one or more cellular component targets, and the second aptamer may not be accompanied by a sample indexing oligonucleotide. The second aptamer may be accompanied by a second sample indexing oligonucleotide containing a second sample indexing sequence. The aptamer and the second aptamer may be sequence-identical by at least 60%, 70%, 80%, 90%, or 95%. The aptamer and the second aptamer may, for example, have identical sequences. The aptamer and the second aptamer may be different. The aptamer and the second aptamer may bind to different regions of the same protein target or to different regions of the same cellular component target. The aptamer and the second aptamer may bind to different protein targets of one or more protein targets or to different cellular component targets of one or more cellular component targets. The sample indexing sequence and the second sample indexing sequence may be the same. The sample indexing sequence and the second sample indexing sequence may be different.

[0018] In some embodiments, the method includes pooling multiple samples that have been in contact with multiple sample indexing compositions before barcoding the sample indexing oligonucleotide. In some embodiments, some of the barcodes among the multiple barcodes include a target-binding region and a molecular labeling sequence, and the molecular labeling sequences of at least two of the multiple barcodes include different molecular labeling sequences. The barcodes may include a cell labeling sequence, a binding site for a universal primer, or any combination thereof. The target-binding region may include a poly(dT) region. In some embodiments, multiple barcodes are associated with the particles. At least one of the multiple barcodes may be immobilized on the particle, partially immobilized on the particle, encapsulated within the particle, partially encapsulated within the particle, or a combination thereof. The particles may be destructible. The particles may contain beads. The particles may contain Sepharose beads, streptavidin beads, agarose beads, magnetic beads, conjugate beads, protein A conjugate beads, protein G conjugate beads, protein A / G conjugate beads, protein L conjugate beads, oligo(dT) conjugate beads, silica beads, silica-like beads, hydrogel beads, antibiotin microbeads, antifluorescent dye microbeads, or any combination thereof, or the particles may contain materials selected from the group consisting of polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogel, paramagnetic materials, ceramics, plastics, glass, methylstyrene, acrylic polymers, titanium, latex, Sepharose, cellulose, nylon, silicone, and any combination thereof. The particles may also contain destructible hydrogel beads.

[0019] In some embodiments, the particle barcode includes molecular labeling sequences selected from at least 1,000, 10,000, or combinations thereof of different molecular labeling sequences. The molecular labeling sequences of the barcode may include random sequences. The particle may contain at least 10,000 barcodes. In some embodiments, barcoding a sample indexing oligonucleotide using multiple barcodes includes contacting the multiple barcodes with the sample indexing oligonucleotide to generate barcodes hybridized to the sample indexing oligonucleotide; and extending the barcodes hybridized to the sample indexing oligonucleotide to generate multiple barcoded sample indexing oligonucleotides.

[0020] In some embodiments, the method includes pooling barcodes hybridized to sample index-granting oligonucleotides before extending the barcodes hybridized to the sample index-granting oligonucleotides, and extending the barcodes hybridized to the sample index-granting oligonucleotides includes extending the pooled barcodes hybridized to the sample index-granting oligonucleotides to generate a pool of pooled barcoded sample index-granting oligonucleotides. Barcode extension may include extending the barcodes using DNA polymerase to generate a pool of barcoded sample index-granting oligonucleotides. Barcode extension may also include extending the barcodes using reverse transcriptase to generate a pool of barcoded sample index-granting oligonucleotides.

[0021] In some embodiments, the method includes amplifying multiple barcoded sample indexing oligonucleotides to produce multiple amplicons. The amplification of multiple barcoded sample indexing oligonucleotides may include amplifying at least a portion of the molecularly labeled sequence and at least a portion of the sample indexing oligonucleotides using polymerase chain reaction (PCR). Obtaining sequencing data for multiple barcoded sample indexing oligonucleotides may include obtaining sequencing data for multiple amplicons. Obtaining sequencing data may include sequencing at least a portion of the molecularly labeled sequence and at least a portion of the sample indexing oligonucleotides.

[0022] In some embodiments, generating multiple barcoded sample indexing oligonucleotides by barcoding them using multiple barcodes includes generating multiple probabilistic barcoded sample indexing oligonucleotides by probabilistically barcoding them using multiple probabilistic barcodes.

[0023] In some embodiments, the method comprises barcoding multiple targets of a cell using multiple barcodes to generate multiple barcoded targets, wherein each of the multiple barcodes contains a cell-labeling sequence, and at least two of the multiple barcodes contain the same cell-labeling sequence; and obtaining sequencing data for the barcoded targets. Barcoding multiple targets using multiple barcodes to generate multiple barcoded targets may include contacting a copy of the target with the target-binding region of the barcode; and reverse transcribing the multiple targets using the multiple barcodes to generate multiple reverse-transcribed targets. The method may also include amplifying the barcoded targets to generate multiple amplified barcoded targets before obtaining sequencing data for the multiple barcoded targets. Amplifying the barcoded targets to generate multiple amplified barcoded targets may include amplifying the barcoded targets by polymerase chain reaction (PCR). Generating multiple barcoded targets by barcoding multiple targets of a cell using multiple barcodes may also include generating multiple probabilistic barcoded targets by probabilistically barcoding multiple targets of a cell using multiple probabilistic barcodes.

[0024] The disclosure herein provides embodiments of a plurality of sample indexing compositions. In some embodiments, each of the plurality of sample indexing compositions comprises an aptamer composition comprising a first cell component-binding aptamer and a sample indexing oligonucleotide, wherein the cell component-binding aptamer is capable of specifically binding to at least one cell component target, and the sample indexing oligonucleotide comprises a sample indexing sequence for identifying the sample origin of one or more cells in the sample, wherein the sample indexing sequences of at least two of the plurality of sample indexing compositions comprise different sequences.

[0025] In some embodiments, the sample indexing sequence is 6 to 60 nucleotides long. The sample indexing oligonucleotide may be 50 to 500 nucleotides long. The sample indexing sequences of at least 10, 100, or 1000 of the multiple sample indexing compositions may contain different sequences. In some embodiments, a single polynucleotide comprises a sample indexing oligonucleotide and a cell component-binding aptamer. The aptamer may be located at the 5' end of the sample indexing oligonucleotide in the single polynucleotide. The aptamer may be located at the 3' end of the sample indexing oligonucleotide in the single polynucleotide.

[0026] In some embodiments, the sample index-granting oligonucleotide is attached to an aptamer. The sample index-granting oligonucleotide may be attached to a cell component-binding aptamer. The sample index-granting oligonucleotide may be covalently attached to a cell component-binding aptamer. The sample index-granting oligonucleotide may be conjugated to a cell component-binding aptamer. The sample index-granting oligonucleotide may be conjugated to a cell component-binding aptamer via a chemical group selected from the group consisting of UV-cleavable groups, streptavidin, biotin, amines, and combinations thereof. The sample index-granting oligonucleotide may be noncovalently attached to a cell component-binding aptamer. The sample index-granting oligonucleotide may be attached to a protein-binding aptamer or a cell component-binding aptamer via a linker.

[0027] In some embodiments, the protein-binding aptamer or cell component-binding aptamer includes a nucleotide aptamer. The nucleotide aptamer may include deoxyribonucleic acid (DNA), ribonucleic acid (RNA), xenonucleic acid (XNA), or a combination thereof. The nucleotide aptamer may include a base analog. The base analog may include a fluorescent base analog. The nucleotide aptamer may include a fluorophore. The cell component-binding aptamer may include a peptide aptamer. In some embodiments, the sample indexing composition is attached to the first carrier. Different sample indexing compositions among a plurality of sample indexing compositions may be attached to different first carriers. In some embodiments, the aptamer composition is attached to the first carrier. In some embodiments, the aptamer composition includes a second cell component-binding aptamer capable of specifically binding to at least one of one or more protein targets or one or more cell component targets, and a second sample indexing oligonucleotide containing a second sample indexing sequence. The cell component-binding aptamer and the second cell component-binding aptamer may be attached to the first carrier. The cell component-binding aptamer may be attached to the first carrier, and the second cell component-binding aptamer may be attached to the second carrier.

[0028] In some embodiments, the cell component-binding aptamer and the second cell component-binding aptamer have at least 60%, 70%, 80%, 90%, or 95% sequence identity. The cell component-binding aptamer and the second cell component-binding aptamer may, for example, have identical sequences. The cell component-binding aptamer and the second cell component-binding aptamer may be different. In some embodiments, the protein target or cellular component target of the cell component-binding aptamer and the second cell component-binding aptamer are the same. The cell component-binding aptamer and the second cell component-binding aptamer may be capable of binding to different regions of the protein target or cellular component target. The protein target or cellular component target of the cell component-binding aptamer and the second cell component-binding aptamer may be different.

[0029] In some embodiments, the first support comprises a metal nanomaterial. The first support may comprise a metal nanostructure, metal nanoparticles, or a combination thereof. The first support may comprise a gold nanomaterial. The first support may comprise a gold nanostructure, gold nanoparticles, or a combination thereof. The first support may comprise lysosomes, micelles, vesicles, lipid membranes, lipid bilayers, lipid monolayers, or a combination thereof. In some embodiments, the cell component-binding aptamer is attached to a first support. The cell component-binding aptamer may be covalently attached to the first support. The cell component-binding aptamer may be conjugated to the first support. The cell component-binding aptamer may be conjugated to the first support via a chemical group selected from the group consisting of UV-cleavable groups, streptavidin, biotin, amines, and combinations thereof. The cell component-binding aptamer may be noncovalently linked to the first support. The cell component-binding aptamer may be attached to the first support via a linker. In some embodiments, the cell component-binding aptamer is immobilized on the first carrier, partially immobilized on the first carrier, immobilized within the first carrier, partially immobilized within the first carrier, encapsulated within the first carrier, partially encapsulated within the first carrier, embedded within the first carrier, partially embedded within the first carrier, or a combination thereof.

[0030] In some embodiments, the first carrier is configured to be internally transported into (or can be internally transported into) a cell. The first carrier may be internally transported into a cell by endocytosis, pinocytosis, nanopinocytosis, micropinocytosis, phagocytosis, membrane fusion, or a combination thereof. The first carrier may be internally transported into a cell by clathrin-mediated internal transport, caveolin-mediated internal transport, receptor-dependent internal transport, receptor-independent internal transport, or a combination thereof.

[0031] In some embodiments, the sample indexing oligonucleotide is not homologous to any of the genomic sequences of one or more cells. At least one of the samples may include one or more single cells, multiple cells, tissue, tumor samples, or any combination thereof. The samples may include mammalian samples, bacterial samples, viral samples, yeast samples, fungal samples, or any combination thereof. In some embodiments, the sample indexing oligonucleotide includes a sequence complementary to the capture sequence, configured to capture (or capable of capturing, hybridizing, or binding) the sequence of the sample indexing oligonucleotide. The target-binding region may include a poly(dT) region. The sequence of the sample indexing oligonucleotide complementary to the capture sequence may include a poly(dA) region.

[0032] In some embodiments, the sample indexing oligonucleotide includes an alignment sequence adjacent to the poly(dA) region. The alignment sequence may have a nucleotide length of or longer than that. The alignment sequence may have a nucleotide length of 2 or longer. The alignment sequence may include guanine, cytosine, thymine, uracil, or a combination thereof. The alignment sequence may include a poly(dT) region, a poly(dG) region, a poly(dC) region, a poly(dU) region, or a combination thereof. In some embodiments, the sample indexing oligonucleotide includes a molecular labeling sequence, a poly(dA) region, or a combination thereof. The molecular labeling sequence may be 2 to 20 nucleotides long. The universal primer may be 5 to 50 nucleotides long. The universal primer may include an amplification primer, a sequencing primer, or a combination thereof.

[0033] In some embodiments, cellular component targets include cell surface proteins, cell markers, B cell receptors, T cell receptors, major histocompatibility complexes, tumor antigens, receptors, or any combination thereof. The cellular component targets can be selected from a group comprising 10 to 100 different cellular component targets. In some embodiments, at least one of the one or more protein targets or one or more cellular component targets is located on the cell surface. In some embodiments, the protein targets or cellular component targets include carbohydrates, lipids, proteins, extracellular proteins, cell surface proteins, cell markers, B cell receptors, T cell receptors, major histocompatibility complexes, tumor antigens, receptors, intracellular proteins, or any combination thereof. The protein targets or cellular component targets can be selected from a group comprising 10 to 100 different protein targets or cellular component targets.

[0034] In some embodiments, the cell component-binding aptamer is accompanied by two or more sample indexing oligonucleotides having identical sequences. The cell component-binding aptamer may also be accompanied by two or more sample indexing oligonucleotides having different sample indexing sequences. In some embodiments, the sample indexing composition comprises a second cell component-binding aptamer capable of specifically binding to at least one of one or more cell component targets. The cell component-binding aptamer and the second cell component-binding aptamer may be capable of binding to the same cell component target of one or more cell component targets, and the second cell component-binding aptamer may not be associated with a sample index-granting oligonucleotide. The second cell component-binding aptamer may be associated with a second sample index-granting oligonucleotide containing a second sample index-granting sequence. The sample index-granting sequence and the second sample index-granting sequence may be the same. The sample index-granting sequence and the second sample index-granting sequence may be different.

[0035] In some embodiments, the cell component-binding aptamer and the second cell component-binding aptamer are at least 60%, 70%, 80%, 90%, or 95% identical. The cell component-binding aptamer and the second cell component-binding aptamer may be identical. The cell component-binding aptamer and the second cell component-binding aptamer may be capable of binding to different regions of the same protein target. The cell component-binding aptamer and the second cell component-binding aptamer may be capable of binding to different protein targets of one or more protein targets.

[0036] The disclosure herein describes embodiments of a kit for identifying a sample. In some embodiments, the kit comprises a plurality of sample indexing compositions. In some embodiments, the kit comprises a plurality of barcodes. One of the barcodes comprises a target-binding region. The target-binding region may comprise a capture sequence configured to capture (or capable of capturing, hybridizing, or binding) the sequence of the sample indexing oligonucleotide. The cell component-binding aptamer may comprise the sample indexing oligonucleotide and a first cell component-binding aptamer.

[0037] The disclosure herein is an embodiment of a method for measuring protein expression in cells. In some embodiments, the method involves contacting a plurality of protein-binding aptamers with a plurality of cells containing a plurality of protein targets, each of the plurality of protein-binding aptamers comprising an aptamer-specific oligonucleotide containing a unique identifier sequence for a cell component-binding aptamer, and the protein-binding aptamer being capable of specifically binding to at least one of the plurality of protein targets; distributing the plurality of cells associated with the plurality of protein-binding aptamers into a plurality of compartments, each compartment containing a single cell derived from the plurality of cells associated with the protein-binding aptamers; and in the compartment containing the single cell, The method comprises contacting a barcoding particle with an aptamer-specific oligonucleotide, wherein each barcoding particle comprises a plurality of oligonucleotide probes, each comprising a target-binding region and a barcode sequence selected from a diverse set of unique barcode sequences; extending the oligonucleotide probes hybridized to the aptamer-specific oligonucleotide to produce a plurality of labeled nucleic acids, each of which comprises a unique identifier sequence or its complementary sequence and a barcode sequence; and obtaining sequence information of the plurality of labeled nucleic acids or a portion thereof to determine the amount of one or more of a plurality of protein targets in one or more of a plurality of cells.

[0038] The disclosures herein include embodiments of methods for measuring the expression of cellular components in cells. In some embodiments, the method comprises contacting a plurality of cellular component-binding aptamers with a plurality of cells containing a plurality of cellular component targets, each of the plurality of cellular component-binding aptamers comprising an aptamer-specific oligonucleotide comprising a unique identifier sequence for the cellular component-binding aptamer, and the cellular component-binding aptamer being capable of specifically binding to at least one of the plurality of cellular component targets; extending an oligonucleotide probe hybridized to the aptamer-specific oligonucleotide to produce a plurality of labeled nucleic acids, each of which comprises a unique identifier sequence or its complementary sequence and a barcode sequence; and obtaining sequence information of the plurality of labeled nucleic acids or a portion thereof to determine the amount of one or more of the plurality of cellular component targets in one or more of the plurality of cells. The method comprises distributing multiple cells associated with multiple cell component-binding aptamers into multiple compartments before extending an oligonucleotide probe, wherein each compartment contains a single cell derived from the multiple cells associated with the cell component-binding aptamers; and in the compartment containing the single cell, contacting a barcoding particle with an aptamer-specific oligonucleotide, wherein the barcoding particle contains multiple oligonucleotide probes, each containing a target-binding region and a barcode sequence selected from a diverse set of unique barcode sequences. The multiple cell component targets include multiple protein targets, and the cell component-binding aptamer is capable of specifically binding to at least one of the multiple protein targets.

[0039] In some embodiments, the aptamer-specific oligonucleotide and the aptamer form a single polynucleotide. The aptamer may be located at the 5' end of the aptamer-specific oligonucleotide in the single polynucleotide. The aptamer may be located at the 3' end of the aptamer-specific oligonucleotide in the single polynucleotide. The aptamer-specific oligonucleotide may be attached to the aptamer. The aptamer-specific oligonucleotide may be covalently attached to the aptamer. The aptamer-specific oligonucleotide may be noncovalently attached to the aptamer. The aptamer-specific oligonucleotide may be conjugated to the aptamer. The aptamer-specific oligonucleotide may be conjugated to the aptamer via a chemical group selected from the group consisting of UV-cleavable groups, streptavidin, biotin, amines, and combinations thereof. The aptamer-specific oligonucleotide may be covalently attached to the aptamer. In some embodiments, this method includes detaching an aptamer-specific oligonucleotide from the aptamer.

[0040] In some embodiments, the aptamer includes a nucleotide aptamer. The nucleotide aptamer may include deoxyribonucleic acid (DNA), ribonucleic acid (RNA), xenonucleic acid (XNA), or a combination thereof. The nucleotide aptamer may include a base analog. The base analog may include a fluorescent base analog. The nucleotide aptamer may include a fluorophore. The aptamer may include a peptide aptamer. In some embodiments, the multiple protein-binding aptamers include a second protein-binding aptamer, or the multiple cell component-binding aptamers include a second cell component-binding aptamer. The aptamers and the second aptamer may be attached to the first carrier. The aptamers may be attached to the first carrier, and the second aptamer may be attached to the second carrier. The aptamers and the second aptamer may be identical in sequence by at least 60%, 70%, 80%, 90%, or 95%. The aptamers and the second protein-binding aptamer may be identical. The aptamers and the second aptamer may be different. In some embodiments, the protein target or cellular component target of the aptamer and the second aptamer are identical. The aptamer and the second aptamer may be capable of binding to different regions of the protein target or cellular component target. The protein target or cellular component target of the aptamer and the second aptamer may be different.

[0041] In some embodiments, the first support comprises a metal nanomaterial. The first support may comprise a metal nanostructure, metal nanoparticles, or a combination thereof. The first support may comprise a gold nanomaterial. The first support may comprise a gold nanostructure, gold nanoparticles, or a combination thereof. The first support may comprise lysosomes, micelles, vesicles, lipid membranes, lipid bilayers, lipid monolayers, or a combination thereof.

[0042] In some embodiments, the aptamer is attached to a first support. The aptamer may be covalently attached to the first support. The aptamer may be conjugated to the first support. The aptamer may be conjugated to the first support via a chemical group selected from the group consisting of UV-cleavable groups, streptavidin, biotin, amines, and combinations thereof. The aptamer may be noncovalently linked to the first support. The aptamer may be attached to the first support via a linker. The aptamer may be immobilized on the first support, partially immobilized on the first support, immobilized within the first support, partially immobilized within the first support, encapsulated within the first support, partially encapsulated within the first support, embedded within the first support, partially embedded within the first support, or a combination thereof.

[0043] In some embodiments, the method includes dissociating the aptamer from the first carrier. Dissociation may occur after contact in a compartment containing a single cell. Dissociation may occur before contact in a compartment containing a single cell. In some embodiments, contacting multiple aptamers with multiple cells involves contacting a first carrier containing the aptamers with a cell among multiple cells containing multiple protein targets. The first carrier may be internally transported into the cell. The first carrier may be internally transported into the cell by endocytosis, pinocytosis, nanopinocytosis, micropinocytosis, phagocytosis, membrane fusion, or a combination thereof. The first carrier may be internally transported into the cell by clathrin-mediated internal transport, caveolin-mediated internal transport, receptor-dependent internal transport, receptor-independent internal transport, or a combination thereof.

[0044] In some embodiments, the aptamer-specific oligonucleotide includes a sequence complementary to the capture sequence, configured to capture (or capable of capturing, hybridizing, or binding) the sequence of the aptamer-specific oligonucleotide. The barcode may include a target-binding region containing the capture sequence. The target-binding region may include a poly(dT) region. The aptamer-specific oligonucleotide sequence complementary to the capture sequence may include a poly(dA) region. In some embodiments, an aptamer-specific oligonucleotide changes from a first conformation in which the poly(dA) region is inaccessible to a second conformation in which the poly(dA) region is accessible when the aptamer comes into contact with at least one of a plurality of protein targets or cellular component targets. The poly(dA) region of the aptamer-specific oligonucleotide in the first conformation may constitute a hairpin structure. The poly(dA) region with the hairpin structure may be accessible to the poly(dT) region of the target-binding region of the barcode. An oligonucleotide probe can be hybridized to an aptamer-specific oligonucleotide by hybridization of the poly(dA) region of the aptamer-specific oligonucleotide and the poly(dT) region of the oligonucleotide probe.

[0045] In some embodiments, the aptamer-specific oligonucleotide includes an alignment sequence adjacent to the poly(dA) region. The alignment sequence may have a nucleotide length of 1 or more. The alignment sequence may have a nucleotide length of 2. The alignment sequence may include guanine, cytosine, thymine, uracil, or a combination thereof. The alignment sequence may include a poly(dT) region, a poly(dG) region, a poly(dC) region, a poly(dU) region, or a combination thereof. In some embodiments, the aptamer-specific oligonucleotide comprises a molecular labeling sequence, a binding site for a universal primer, or both. The molecular labeling sequence may be 2 to 20 nucleotides long. The universal primer may be 5 to 50 nucleotides long. The universal primer may comprise an amplification primer, a sequencing primer, or a combination thereof.

[0046] In some embodiments, the method includes contacting multiple aptamers with multiple cells and then removing aptamers that have not come into contact with any of the cells. Removal of aptamers that have not come into contact with any of the cells may include removing aptamers that have not come into contact with at least one corresponding protein target of the multiple protein targets. Removal of aptamers that have not come into contact with at least one corresponding protein target of the multiple protein targets may include using an aptamer sponge to remove aptamers that have not come into contact with at least one corresponding protein target of the multiple protein targets. The aptamer sponge may contain multiple sequences complementary to the aptamer sequences. Aptamers that have not come into contact with at least one corresponding protein target of the multiple protein targets can hybridize to multiple sequences of the aptamer sponge. In some embodiments, the aptamers contain complementary sequences. Aptamers that are not in contact with at least one corresponding protein target may hybridize with each other to form aggregates. These aggregates may be targeted for degradation in cells.

[0047] In some embodiments, the removal of an aptamer that is not in contact with at least one corresponding protein target of a plurality of protein targets includes using a plurality of antisense aptamers to remove an aptamer that is not in contact with at least one corresponding protein target of a plurality of protein targets, each antisense aptamer comprising a second aptamer comprising an antisense aptamer-specific oligonucleotide comprising a complementary sequence or a portion thereof of one aptamer-specific oligonucleotide of the aptamer that is not in contact with at least one corresponding protein target of the plurality of protein targets, wherein the aptamer-specific oligonucleotide and the antisense aptamer-specific oligonucleotide form a double-stranded or partially double-stranded deoxyribonucleic acid (DNA) molecule. In some embodiments, the distribution of multiple cells involves distributing multiple cells associated with multiple aptamers and multiple barcoding particles, each containing a single cell derived from multiple cells associated with oligonucleotide conjugate aptamers and barcoding particles, into multiple compartments. In some embodiments, the compartments are wells or droplets. In some embodiments, obtaining sequencing information for multiple labeled nucleic acids or a portion thereof involves subjecting the labeled nucleic acids to one or more reactions to produce a set of nucleic acids for nucleic acid sequencing. In some embodiments, the multiple protein targets include cell surface proteins, intracellular proteins, cell markers, B cell receptors, T cell receptors, antibodies, major histocompatibility complexes, tumor antigens, receptors, or combinations thereof. In some embodiments, the multiple cells include T cells, B cells, tumor cells, myeloid cells, blood cells, normal cells, fetal cells, maternal cells, or mixtures thereof.

[0048] In some embodiments, each oligonucleotide probe includes a cell label, a binding site for a universal primer, an amplification adapter, a sequencing adapter, or a combination thereof. In some embodiments, the barcoding particles are Sepharose beads, streptavidin beads, agarose beads, magnetic beads, silica beads, silica-like beads, hydrogel beads, antibiotin microbeads, antifluorescent dye microbeads, or hydrogel beads.

[0049] Disclosure herein includes compositions comprising a plurality of cell component-binding aptamers, each of which comprises an aptamer-specific oligonucleotide containing a unique identifier sequence for the cell component-binding aptamer, and the cell component-binding aptamer is capable of specifically binding to at least one of a plurality of cell component targets. In some embodiments, the aptamer-specific oligonucleotide and the aptamer form a single polynucleotide. The cell component-binding aptamer may be located at the 5' end of the aptamer-specific oligonucleotide in the single polynucleotide. The cell component-binding aptamer may be located at the 3' end of the aptamer-specific oligonucleotide in the single polynucleotide.

[0050] In some embodiments, the aptamer-specific oligonucleotide is attached to a cell component-binding aptamer. The aptamer-specific oligonucleotide may be attached to the cell component-binding aptamer. The aptamer-specific oligonucleotide may be covalently attached to the cell component-binding aptamer. The aptamer-specific oligonucleotide may be noncovalently attached to the cell component-binding aptamer. The aptamer-specific oligonucleotide may be conjugated to the cell component-binding aptamer. The aptamer-specific oligonucleotide may be conjugated to the cell component-binding aptamer via a chemical group selected from the group consisting of UV-cleavable groups, streptavidin, biotin, amines, and combinations thereof. The aptamer-specific oligonucleotide may be covalently attached to the cell component-binding aptamer.

[0051] In some embodiments, the cell component-binding aptamer includes a nucleotide aptamer. The nucleotide aptamer may include deoxyribonucleic acid (DNA), ribonucleic acid (RNA), xenonucleic acid (XNA), or a combination thereof. The nucleotide aptamer may include a base analog. The base analog may include a fluorescent base analog. The nucleotide aptamer may include a fluorophore. The cell component-binding aptamer may include a peptide aptamer. In some embodiments, the multiple cell component-binding aptamers include a second cell component-binding aptamer. The cell component-binding aptamers and the second cell component-binding aptamer may be attached to the first carrier. The cell component-binding aptamers may be attached to the first carrier, and the second cell component-binding aptamer may be attached to the second carrier.

[0052] In some embodiments, the cell component-binding aptamer and the second cell component-binding aptamer may have at least 60%, 70%, 80%, 90%, or 95% sequence identity. The cell component-binding aptamer and the second protein-binding aptamer may be the same. The cell component-binding aptamer and the second cell component-binding aptamer may be different. In some embodiments, the protein target or cellular component target of the cell component-binding aptamer and the second cell component-binding aptamer are the same. The cell component-binding aptamer and the second cell component-binding aptamer may be capable of binding to different regions of the protein target or cellular component target. The protein target or cellular component target of the cell component-binding aptamer and the second cell component-binding aptamer may be different.

[0053] In some embodiments, the first support comprises a metal nanomaterial. The first support may comprise a metal nanostructure, metal nanoparticles, or a combination thereof. The first support may comprise a gold nanomaterial. The first support may comprise a gold nanostructure, gold nanoparticles, or a combination thereof. The first support may comprise lysosomes, micelles, vesicles, lipid membranes, lipid bilayers, lipid monolayers, or a combination thereof. In some embodiments, the cell component-binding aptamer is attached to a first support. The cell component-binding aptamer may be covalently attached to the first support. The cell component-binding aptamer may be conjugated to the first support. The cell component-binding aptamer may be conjugated to the first support via a chemical group selected from the group consisting of UV-cleavable groups, streptavidin, biotin, amines, and combinations thereof. The cell component-binding aptamer may be noncovalently linked to the first support. The cell component-binding aptamer may be attached to the first support via a linker. The cell component-binding aptamer may be immobilized on the first carrier, partially immobilized on the first carrier, immobilized within the first carrier, partially immobilized within the first carrier, encapsulated within the first carrier, partially encapsulated within the first carrier, embedded within the first carrier, partially embedded within the first carrier, or a combination thereof.

[0054] In some embodiments, the first carrier is configured to be internally transported into (or may be internally transported into) a cell. The first carrier may be configured to be internally transported into (or may be internally transported into) a cell by endocytosis, pinocytosis, nanopinocytosis, micropinocytosis, phagocytosis, membrane fusion, or a combination thereof. The first carrier may be configured to be internally transported into (or may be internally transported into) a cell by clathrin-mediated internal transport, caveolin-mediated internal transport, receptor-dependent internal transport, receptor-independent internal transport, or a combination thereof.

[0055] In some embodiments, the aptamer-specific oligonucleotide includes a sequence complementary to the capture sequence, configured to capture (or capable of capturing, binding to, or hybridizing) the sequence of the aptamer-specific oligonucleotide. The barcode may include a target-binding region containing the capture sequence. The target-binding region may include a poly(dT) region. The aptamer-specific oligonucleotide sequence complementary to the capture sequence may include a poly(dA) region. In some embodiments, an aptamer-specific oligonucleotide changes from a first conformation in which the poly(dA) region is inaccessible to a second conformation in which the poly(dA) region is accessible when the aptamer comes into contact with at least one of a plurality of protein targets or cellular component targets. The poly(dA) region of the aptamer-specific oligonucleotide in the first conformation may constitute a hairpin structure. The poly(dA) region with the hairpin structure may be accessible to the poly(dT) region of the target-binding region of the barcode. An oligonucleotide probe can be hybridized to an aptamer-specific oligonucleotide by hybridization of the poly(dA) region of the aptamer-specific oligonucleotide and the poly(dT) region of the oligonucleotide probe.

[0056] In some embodiments, the aptamer-specific oligonucleotide includes an alignment sequence adjacent to the poly(dA) region. The alignment sequence may have a nucleotide length of 1 or more. The alignment sequence may have a nucleotide length of 2. The alignment sequence may include guanine, cytosine, thymine, uracil, or a combination thereof. The alignment sequence may include a poly(dT) region, a poly(dG) region, a poly(dC) region, a poly(dU) region, or a combination thereof. In some embodiments, the aptamer-specific oligonucleotide may include a molecular labeling sequence, a binding site for a universal primer, or both. The molecular labeling sequence may be 2 to 20 nucleotides long. The universal primer may be 5 to 50 nucleotides long. The universal primer may include an amplification primer, a sequencing primer, or a combination thereof.

[0057] In some embodiments, the composition includes an aptamer sponge. The aptamer sponge may contain multiple sequences complementary to the aptamer sequence. In some embodiments, the aptamers include complementary sequences. The aptamers may be configured to hybridize with each other (or be enabled to hybridize) to form aggregates. These aggregates may be targeted for degradation in cells.

[0058] In some embodiments, the composition comprises a plurality of antisense aptamers, each of which comprises a second aptamer comprising an antisense aptamer-specific oligonucleotide comprising a complementary sequence or a portion thereof of one aptamer-specific oligonucleotide of the aptamer, wherein the aptamer-specific oligonucleotide and the antisense aptamer-specific oligonucleotide form a double-stranded or partially double-stranded deoxyribonucleic acid (DNA) molecule. [Brief explanation of the drawing]

[0059] [Figure 1] This figure shows an unrestricted, illustrative, probabilistic barcode. [Figure 2] This figure shows an unrestricted, illustrative workflow for probabilistic barcoding and electronic counting. [Figure 3] This is a schematic diagram illustrating a non-restrictive, illustrative process for creating an indexing library of probabilistically barcoded targets from multiple targets. [Figure 4]A schematic diagram of an exemplary protein-binding reagent (the antibody shown in this figure) in which an oligonucleotide containing a unique identifier is attached to the protein-binding reagent is shown. [Figure 5] A schematic diagram of an exemplary binding reagent (antibody shown in this figure) is shown, which is accompanied by an oligonucleotide containing a unique identifier for sample indexing to determine cells from the same or different samples. [Figure 6] This diagram illustrates an exemplary workflow using antibodies accompanied by oligonucleotides to simultaneously determine the expression of cellular components (e.g., protein expression) and gene expression in a high-throughput manner. [Figure 7] A schematic diagram illustrates an exemplary workflow using antibodies with oligonucleotides attached for sample indexing. [Figure 8A] A schematic diagram illustrates an exemplary workflow using aptamers with attached oligonucleotides (e.g., conjugated aptamers) to determine the expression of cellular components (e.g., protein expression). [Figure 8B] A schematic diagram illustrates an exemplary workflow using aptamers with attached oligonucleotides (e.g., conjugated aptamers) to determine the expression of cellular components (e.g., protein expression). [Figure 8C] A schematic diagram illustrates an exemplary workflow using aptamers with attached oligonucleotides (e.g., conjugated aptamers) to determine the expression of cellular components (e.g., protein expression). [Figure 8D] A schematic diagram illustrates an exemplary workflow using aptamers with attached oligonucleotides (e.g., conjugated aptamers) to determine the expression of cellular components (e.g., protein expression). [Figure 8E] A schematic diagram illustrates an exemplary workflow using aptamers with attached oligonucleotides (e.g., conjugated aptamers) to determine the expression of cellular components (e.g., protein expression). [Figure 9] This figure shows a non-restrictive, exemplary aptamer and its associated sequence. [Figure 10A] A schematic diagram illustrates an exemplary workflow using an oligonucleotide-associated aptamer for sample indexing. [Figure 10B] A schematic diagram illustrates an exemplary workflow using an oligonucleotide-associated aptamer for sample indexing. [Figure 11A] This figure shows a non-limiting, exemplary design of oligonucleotides for simultaneously determining protein and gene expression and for assigning sample indices. [Figure 11B] This figure shows a non-limiting, exemplary design of oligonucleotides for simultaneously determining protein and gene expression and for assigning sample indices. [Figure 11C] This figure shows a non-limiting, exemplary design of oligonucleotides for simultaneously determining protein and gene expression and for assigning sample indices. [Figure 11D] This figure shows a non-limiting, exemplary design of oligonucleotides for simultaneously determining protein and gene expression and for assigning sample indices. [Figure 12] A schematic diagram of non-restrictive, exemplary oligonucleotide sequences for simultaneously determining protein and gene expression and for assigning sample indices is shown. [Modes for carrying out the invention]

[0060] In the embodiments for carrying out the invention described herein, references are made to the accompanying drawings which form part of this specification. In the drawings, unless the context indicates otherwise, similar symbols typically indicate similar components. The exemplary embodiments described in the embodiments for carrying out the invention, the drawings, and the claims are not limiting. Other embodiments may be used and other modifications may be made without departing from the spirit or scope of the subject matter provided herein. The embodiments of this disclosure, as generally described herein and illustrated in the drawings, may be arranged in, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are expressly intended herein and will be readily understood to form part of this disclosure. All patents, published patent applications, other publications, and the array of GenBank and other databases referenced herein are incorporated by reference in their entirety with respect to the relevant technology.

[0061] Quantifying a small number of nucleic acids, such as messenger ribonucleotide (mRNA) molecules, is clinically important, for example, to determine which genes are expressed in cells at different developmental stages or under different environmental conditions. However, determining the absolute number of nucleic acid molecules (e.g., mRNA molecules) can be very difficult, especially when the number of molecules is very small. One method for determining the absolute number of molecules in a sample is digital polymerase chain reaction (PCR). Ideally, PCR produces identical copies of molecules in each cycle. However, PCR has the drawback that each molecule is replicated with a probabilistic probability, which varies depending on the PCR cycle and gene sequence, leading to amplification bias and inaccurate gene expression measurements. Using probabilistic barcodes with unique molecular labels (also called molecular indices (MIs)) allows for counting the number of molecules and correcting for amplification bias. Probabilistic barcoding methods such as the Precise® assay (Cellular Research, Inc., Palo Alto, California) can correct for biases induced by the PCR and library preparation steps by labeling mRNA using molecular labeling (ML) during reverse transcription (RT).

[0062] The Precise® assay allows the use of a non-exhaustion pool of numerous probabilistic barcodes, e.g., 6561–65536, each with unique molecular labels, for hybridizing to all poly(A)mRNA in the sample during the RT step. The probabilistic barcodes may contain universal PCR priming sites. During RT, target gene molecules react randomly with the probabilistic barcodes. Each target molecule can hybridize to a probabilistic barcode to generate a probabilistically barcoded complementary ribonucleotide (cDNA) molecule. After labeling, the probabilistically barcoded cDNA molecules from the microwells of a microwell plate can be pooled into a single tube for PCR amplification and sequencing. The raw sequence data can be analyzed to obtain the number of reads, the number of probabilistic barcodes with unique molecular labels, and the number of mRNA molecules.

[0063] Methods for determining single-cell mRNA expression profiles can be performed in a large-scale parallel manner. For example, the Precise® assay can simultaneously determine mRNA expression profiles of more than 10,000 cells. The number of single cells analyzed per sample (e.g., hundreds or thousands of single cells) may be less than the capacity of current single-cell technologies. By pooling cells from different samples, the use of the capacity of current single-cell technologies can be improved, thereby reducing reagent waste and the cost of single-cell analysis. This disclosure provides a sample indexing method for distinguishing cells from different samples in order to prepare cDNA libraries for cell analysis, such as single-cell analysis. By pooling cells from different samples, variability in cDNA library preparation from cells of different samples can be minimized, enabling more accurate comparisons of different samples.

[0064] The disclosure herein includes embodiments of a method for identifying a sample. In some embodiments, the method includes contacting each of a plurality of samples with a sample indexing composition from a plurality of sample indexing compositions, each of the plurality of samples comprising one or more cells each comprising one or more protein targets, the sample indexing composition comprising an aptamer composition comprising an aptamer and a sample indexing oligonucleotide, wherein the aptamer is capable of specifically binding to at least one of the one or more protein targets, the sample indexing oligonucleotide comprising a sample indexing sequence, wherein the sample indexing sequences of at least two of the plurality of sample indexing compositions comprise different sequences; barcoding the sample indexing oligonucleotides using a plurality of barcodes to generate a plurality of barcoded sample indexing oligonucleotides; obtaining sequencing data for the plurality of barcoded sample indexing oligonucleotides; and identifying the sample origin of at least one of the one or more cells based on the sample indexing sequence of at least one of the barcoded sample indexing oligonucleotides.

[0065] The disclosure herein provides embodiments of a plurality of sample indexing compositions. In some embodiments, each of the plurality of sample indexing compositions comprises an aptamer composition comprising a first cell component-binding aptamer and a sample indexing oligonucleotide, wherein the cell component-binding aptamer is capable of specifically binding to at least one cell component target, and the sample indexing oligonucleotide comprises a sample indexing sequence for identifying the sample origin of one or more cells in the sample, wherein the sample indexing sequences of at least two of the plurality of sample indexing compositions comprise different sequences.

[0066] The disclosure herein is an embodiment of a method for measuring protein expression in cells. In some embodiments, the method involves contacting a plurality of protein-binding aptamers with a plurality of cells containing a plurality of protein targets, each of the plurality of protein-binding aptamers comprising an aptamer-specific oligonucleotide containing a unique identifier sequence for a cell component-binding aptamer, and the protein-binding aptamer being capable of specifically binding to at least one of the plurality of protein targets; distributing the plurality of cells associated with the plurality of protein-binding aptamers into a plurality of compartments, each compartment containing a single cell derived from the plurality of cells associated with the protein-binding aptamers; and in the compartment containing the single cell, The method comprises contacting a barcoding particle with an aptamer-specific oligonucleotide, wherein each barcoding particle comprises a plurality of oligonucleotide probes, each comprising a target-binding region and a barcode sequence selected from a diverse set of unique barcode sequences; extending the oligonucleotide probes hybridized to the aptamer-specific oligonucleotide to produce a plurality of labeled nucleic acids, each of which comprises a unique identifier sequence or its complementary sequence and a barcode sequence; and obtaining sequence information of the plurality of labeled nucleic acids or a portion thereof to determine the amount of one or more of a plurality of protein targets in one or more of a plurality of cells.

[0067] The disclosures herein include embodiments of methods for measuring the expression of cellular components in cells. In some embodiments, the method comprises contacting a plurality of cellular component-binding aptamers with a plurality of cells containing a plurality of cellular component targets, each of the plurality of cellular component-binding aptamers comprising an aptamer-specific oligonucleotide comprising a unique identifier sequence for the cellular component-binding aptamer, and the cellular component-binding aptamer being capable of specifically binding to at least one of the plurality of cellular component targets; extending an oligonucleotide probe hybridized to the aptamer-specific oligonucleotide to produce a plurality of labeled nucleic acids, each of which comprises a unique identifier sequence or its complementary sequence and a barcode sequence; and obtaining sequence information of the plurality of labeled nucleic acids or a portion thereof to determine the amount of one or more of the plurality of cellular component targets in one or more of the plurality of cells.

[0068] Disclosure herein includes compositions comprising a plurality of cell component-binding aptamers, each of which comprises an aptamer-specific oligonucleotide containing a unique identifier sequence for the cell component-binding aptamer, and the cell component-binding aptamer is capable of specifically binding to at least one of a plurality of cell component targets.

[0069] definition Unless otherwise defined, technical and scientific terms used herein have the same meanings as those generally understood by those skilled in the art to which this disclosure belongs. See, for example, Singleton et al., Dictionary of Microbiology and Molecular Biology 2nd ed., J. Wiley & Sons (New York, NY 1994); Sambrook et al., Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Press (Cold Spring Harbor, NY 1989). For the purposes of this disclosure, the following terms are defined below.

[0070] As used herein, the term “adapter” may mean a sequence for facilitating the amplification or sequencing of an associated nucleic acid. The associated nucleic acid may include a target nucleic acid. The associated nucleic acid may include one or more spatial labels, target labels, sample labels, indexing labels, or barcode sequences (e.g., molecular labels). The adapter may be linear. The adapter may be a pre-adenylated adapter. The adapter may be double-stranded or single-stranded. One or more adapters may be located at the 5' or 3' end of the nucleic acid. If the adapter includes known sequences at the 5' and 3' ends, the known sequences may be the same or different sequences. Adapters located at the 5' and / or 3' ends of a polynucleotide may be hybridized to one or more oligonucleotides immobilized on a surface. In some embodiments, the adapter may include a universal sequence. The universal sequence may be a region of nucleotide sequences common to two or more nucleic acid molecules. Alternatively, two or more nucleic acid molecules may have regions of different sequences. Therefore, for example, a 5' adapter may contain the same and / or universal nucleic acid sequence, and a 3' adapter may contain the same and / or universal sequence. Universal sequences that may be present in different members of multiple nucleic acid molecules can be used to enable replication or amplification of multiple different sequences using a single universal primer complementary to the universal sequence. Similarly, at least one, two (e.g., a pair), or more universal sequences that may be present in different members of a collection of nucleic acid molecules can be used to enable replication or amplification of multiple different sequences using at least one, two (e.g., a pair), or more single universal primers complementary to the universal sequence. Therefore, a universal primer contains a sequence that can hybridize to such a universal sequence.A molecule containing a target nucleic acid sequence can be modified to allow a universal adapter (e.g., a non-target nucleic acid sequence) to be attached to one or both ends of a different target nucleic acid sequence. One or more universal primers attached to the target nucleic acid can provide sites for hybridization of the universal primers. The one or more universal primers attached to the target nucleic acid may be the same or different from each other.

[0071] As used herein, an antibody may be a full-length immunoglobulin molecule (e.g., one that is naturally occurring or formed by a normal immunoglobulin gene fragment recombination process) (e.g., an IgG antibody) or an immunologically active (i.e., specific binding) portion of an immunoglobulin molecule, such as an antibody fragment.

[0072] In some embodiments, the antibody is a functional antibody fragment. For example, the antibody fragment may be part of an antibody such as F(ab')2, Fab', Fab, Fv, and sFv. The antibody fragment can bind to the same antigen recognized by the full-length antibody. The antibody fragment may include isolated fragments of the variable regions of an antibody, such as an "Fv" fragment consisting of heavy and light chain variable regions, and a recombinant single-chain polypeptide molecule ("scFv protein") in which the light chain variable region and the heavy chain variable region are linked by a peptide linker. Examples of antibodies, but not limited to these, include antibodies against cancer cells, antibodies against viruses, antibodies that bind to cell surface receptors (e.g., CD8, CD34, and CD45), and therapeutic antibodies.

[0073] As used herein, the terms “associated” or “associated with” may mean that two or more species are identifiable as colocalized at a given time. Association may mean that two or more species are present in or have been present in a similar container. Association may be informatic; for example, digital information about two or more species can be stored and used to determine that one or more of those species were colocalized at a given time. Association may also be physical; in some embodiments, two or more associated species are “anchored,” “attached,” or “immobilized” to each other or to a common solid or semi-solid surface. Association may refer to covalent or non-covalent means for attaching a label to a solid or semi-solid support, such as beads. Association may be a covalent bond between a target and a label. Association may include 2-molecule hybridization (such as between a target molecule and a label).

[0074] As used herein, the term “complementary” may refer to the exact ability of two nucleotides to pair. For example, if a nucleotide at a given position in one nucleic acid can form a hydrogen bond with a nucleotide in another nucleic acid, then these two nucleic acids are considered complementary at that position. Complementarity between two single-stranded nucleic acid molecules may be “partial,” where only a portion of the nucleotides are bonded, or it may be complete, if complete complementarity exists between the single-stranded molecules. A first nucleotide sequence can be said to be the “complement” of a second nucleotide sequence if the first nucleotide sequence is complementary to the second nucleotide sequence. A first nucleotide sequence can be said to be the “reverse complement” of a second nucleotide sequence if the first nucleotide sequence is complementary to the second nucleotide sequence in the reverse order (i.e., the nucleotides are in the reverse order). As used herein, the terms “complement,” “complementary,” and “reverse complement” can be used synonymously. In this disclosure, it is understood that if a molecule can hybridize to another molecule, it may be a complement of the molecule it is hybridizing to.

[0075] As used herein, the term “digital counting” may refer to a method for estimating the number of target molecules in a sample. Digital counting may include a step of determining the number of unique labels associated with the target in the sample. This methodology may be probabilistic in nature and transforms the molecular counting problem from a problem of finding and identifying identical molecules into a series of yes / no digital questions regarding the detection of a set of predefined labels.

[0076] As used herein, the terms “label(s)” or “label(plural)” may refer to a nucleic acid code associated with a target in a sample. A label may, for example, be a nucleic acid label. A label may be a fully or partially amplified label. A label may be a fully or partially sequenceable label. A label may be a portion of a distinctly identifiable native nucleic acid. A label may be a known sequence. A label may include a junction of nucleic acid sequences, for example, a junction between a native and a non-native sequence. As used herein, the term “label” may be used synonymously with the terms “index,” “tag,” or “labeled tag.” A label can convey information. For example, in various embodiments, a label can be used to determine the identity of a sample, the source of a sample, the identity of cells, and / or targets.

[0077] As used herein, the term “non-exhausted reservoir” may refer to a pool of barcodes (e.g., probabilistic barcodes) consisting of a large number of different labels. A non-exhausted reservoir may contain a large number of different barcodes such that, when the non-exhausted reservoir is associated with a pool of targets, each target is highly likely to be associated with a unique barcode. The uniqueness of each labeled target molecule can be determined by the statistics of random selection and depends on the number of copies of the same target molecule in the collection compared to the diversity of labels. The size of the resulting set of labeled target molecules can be determined by the probabilistic nature of the barcoding process, and by analyzing the number of barcodes detected, it is possible to calculate the number of target molecules present in the original collection or sample. Labeled target molecules are highly unique (i.e., it is very unlikely that more than one target molecule will be labeled with a given label) if the ratio of the number of copies of the present target molecule to the number of unique barcodes is low.

[0078] As used herein, the term “nucleic acid” refers to a polynucleotide sequence or a fragment thereof. Nucleic acids may contain nucleotides. Nucleic acids may be exogenous or endogenous to cells. Nucleic acids may exist in a cell-free environment. Nucleic acids may be genes or fragments thereof. Nucleic acids may be DNA. Nucleic acids may be RNA. Nucleic acids may contain one or more analogs (e.g., modified skeletons, sugars, or nucleic acid bases). Some non-limiting examples of analogs include 5-bromouracil, peptide nucleic acids, xeno nucleic acids, morpholino, locked nucleic acids, glycol nucleic acids, threose nucleic acids, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein linked to sugars), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, cuosin, and waiosin. The terms "nucleic acid," "polynucleotide," "targeted polynucleotide," and "targeted nucleic acid" can be used synonymously.

[0079] Nucleic acids may contain one or more modifications (e.g., base modifications, skeletal modifications) to provide new or enhanced characteristics to the nucleic acid (e.g., improved stability). Nucleic acids may contain nucleic acid affinity tags. Nucleosides may be base-sugar combinations. The base portion of a nucleoside may be a heterocyclic base. Two of the most common types of such heterocyclic bases are purines and pyrimidines. A nucleotide may be a nucleoside further containing a phosphate group covalently linked to the sugar portion of the nucleoside. In the case of a nucleoside containing a pentofuranosyl sugar, the phosphate group may be linked to the 2', 3', or 5' hydroxyl portion of the sugar. When forming nucleic acids, the phosphate group can covalently link adjacent nucleosides to each other to form a linear polymer compound. Each end of this linear polymer compound may then be further joined to form a cyclic compound, but linear compounds are generally preferred. In addition, linear compounds may have internal nucleotide base complementarity and therefore may fold to produce fully or partially double-stranded compounds. Within nucleic acids, phosphate groups are generally referred to as forming the internucleoside skeleton of the nucleic acid. The linkage or skeleton may be a 3'-5' phosphodiester linkage.

[0080] Nucleic acids may contain a modified skeleton and / or modified nucleoside linkages. Modified skeletons include those that retain a phosphorus atom in the skeleton and those that do not. Suitable modified nucleic acid skeletons containing a phosphorus atom include, for example, phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotryesters, aminoalkyl phosphotryesters, methyl and other alkyl phosphonates such as 3'-alkylene phosphonates, 5'-alkylene phosphonates, chiral phosphonates, phosphoamides including 3'-aminophosphorumidates and aminoalkyl phosphoramidates, phosphorodiamidates, thionophosphorumidates, thionoalkyl phosphotryesters, selenophosphates, as well as boranophosphates having the usual 3'-5' linkage, 2'-5' linked analogs, and those having reverse polarity with one or more nucleotide linkages being 3'-3', 5'-5', or 2'-2' links.

[0081] Nucleic acids can include polynucleotide skeletons formed by short-chain alkyl or cycloalkyl nucleoside linkages, mixed heteroatoms and alkyl or cycloalkyl nucleoside linkages, or one or more short-chain heteroatoms or heterocyclic nucleoside linkages. Such include those having morpholino linkages (partially formed from the sugar moiety of the nucleoside); siloxane skeletons; sulfide, sulfoxide, and sulfone skeletons; formacetyl and thioformacetyl skeletons; methyleneformacetyl and thioformacetyl skeletons; riboacetyl skeletons; alkene-containing skeletons; sulfamate skeletons; methyleneimino and methylenehydrazino skeletons; sulfonate and sulfonamide skeletons; amide skeletons; and others having mixed N, O, S, and CH2 component moieties.

[0082] Nucleic acids may include nucleic acid mimes. The term “mimicking” may be intended to include polynucleotides in which only the furanose ring, or both the furanose ring and internucleotide linkages, are replaced with non-furanose groups; replacement of only the furanose ring may also be called a sugar substitute. A heterocyclic base moiety or a modified heterocyclic base moiety may be maintained for hybridization with a suitable target nucleic acid. One such nucleic acid may be a peptide nucleic acid (PNA). In PNAs, the sugar backbone of the polynucleotide may be replaced with an amide-containing backbone, particularly an aminoethylglycine backbone. The nucleotide may be held or bound directly or indirectly to the aza nitrogen atom of the amide moiety of the backbone. The backbone of a PNA compound may contain two or more linked aminoethylglycine units that confer the amide-containing backbone to the PNA. The heterocyclic base moiety may be bound directly or indirectly to the aza nitrogen atom of the amide moiety of the backbone.

[0083] The nucleic acid may contain a morpholino skeletal structure. For example, the nucleic acid may contain a six-membered morpholino ring instead of a ribose ring. In some of these embodiments, phosphorodiamidates or other non-phosphodiester nucleoside linkages may replace the phosphodiester linkages.

[0084] Nucleic acids may contain linked morpholino units (e.g., morpholino nucleic acids) having heterocyclic bases attached to a morpholino ring. Linking groups can link morpholino monomer units of morpholino nucleic acids. Nonionic morpholino-based oligomeric compounds may have fewer undesirable interactions with cellular proteins. Morpholino-based polynucleotides may be nonionic mimics of nucleic acids. Different linking groups can be used to link various compounds within the morpholino type. Further types of polynucleotide mimics are sometimes called cyclohexenyl nucleic acids (CeNA). The furanose ring normally present in nucleic acid molecules can be replaced with a cyclohexenyl ring. CeNA DMT-protected phosphoramidite monomers can be prepared and used in the synthesis of oligomeric compounds using phosphoramidite chemical reactions. Incorporation of CeNA monomers into nucleic acid chains can increase the stability of DNA / RNA hybrids. CeNA oligoadenylates can form complexes with nucleic acid complements with similar stability to natural complexes. Further modifications may include locked nucleic acids (LNAs) in which a 2'-hydroxyl group is linked to the 4' carbon atom of the sugar ring, thereby forming a 2'-C,4'-C oxymethylene linkage, and thereby forming a bicyclic sugar moiety. The linkage may also be a methylene (-CH2) group (wherein n is 1 or 2) bridging the 2' oxygen atom and the 4' carbon atom. LNAs and LNA analogs can exhibit very high double-chain thermal stability with complementary nucleic acids (Tm = +3 to +10°C), stability against 3'-exonuclease degradation, and good solubility properties.

[0085] Furthermore, nucleic acids may include modifications or substitutions of nucleic acid bases (often simply called "bases"). As used herein, "unmodified" or "natural" nucleic acid bases include purine bases (e.g., adenine (A) and guanine (G)) and pyrimidine bases (e.g., thymine (T), cytosine (C), and uracil (U)). Modified nucleic acid bases include 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyl(-C=C-CH3)uracil and cytosine, as well as other alkynyl derivatives of pyrimidine bases, 6-azouracil, cytosine and thymine, and 5-uracil (pso Other examples of synthetic and natural nucleic acid bases include ioduracil, 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl, and other 8-substituted adenines and guanines, 5-halo, especially 5-bromo, 5-trifluoromethyl, and other 5-substituted uracils and cytosines, 7-methylguanine and 7-methyladenine, 2-F-adenine, 2-aminoadenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine, as well as 3-deazaguanine and 3-deazaadenine.Modified nucleic acid bases include tricyclic pyrimidines such as phenoxazinecytidine (1H-pyrimido(5,4-b)(1,4)benzoxazine-2(3H)-one) and phenothiazinecytidine (1H-pyrimido(5,4-b)(1,4)benzothiadin-2(3H)-one), substituted phenoxazinecytidines (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzoxazine-2(3H)-one), and phenothiazinecytidine (1H-pyrimido(5, Examples of G-clamps include 4-b)(1,4)benzothiazine-2(3H)-one), substituted phenoxazine cytidines (e.g., 9-(2-aminoethoxy)-H-pyrimide(5,4-(b)(1,4)benzoxazine-2(3H)-one), carbazole cytidine (2H-pyrimide(4,5-b)indole-2-one), and pyridoindole cytidine (H-pyrimide(3',2':4,5)pyrrolo[2,3-d]pyrimidine-2-one).

[0086] As used herein, the term “sample” may refer to a composition containing a target. Suitable samples for analysis by the methods, devices, and systems of this disclosure include cells, tissues, organs, or organisms.

[0087] As used herein, the terms “sample collection device” or “device” may refer to a device capable of collecting sections of a sample and / or placing sections on a substrate. Sample devices may refer to, for example, fluorescence-activated cell sorting (FACS) machines, cell sorting machines, biopsy needles, biopsy devices, tissue sectioning devices, microfluidic devices, blade grids, and / or microtomes.

[0088] As used herein, the term “solid support” may refer to individual solid or semi-solid surfaces to which multiple barcodes (e.g., probabilistic barcodes) can be attached. A solid support may encompass any type of solid, porous, or hollow sphere, ball, bearing, cylinder, or other similar configuration made of plastic, ceramic, metal, or polymer material (e.g., hydrogel) to which nucleic acids can be immobilized (e.g., covalently or non-covalently). A solid support may include individual particles that are spherical (e.g., microspheres) or have non-spherical or irregular shapes, such as cubes, pyramidal, cylindrical, conical, rectangular, or disc-shaped. Beads may be non-spherical in shape. Multiple solid supports spaced apart in an array may not constitute a substrate. The term “solid support” may be used synonymously with “beads.”

[0089] As used herein, the term “probabilistic barcode” may refer to a polynucleotide sequence containing the label of this disclosure. A probabilistic barcode may also be a polynucleotide sequence that can be used for probabilistic barcoding. A probabilistic barcode can be used to quantify a target in a sample. A probabilistic barcode can be used to control for errors that may occur after the label has been attached to the target. For example, a probabilistic barcode can be used to evaluate amplification errors or sequencing errors. A probabilistic barcode attached to a target may be referred to as a probabilistic barcode-target or probabilistic barcode-tag-target. As used herein, the term “gene-specific probabilistic barcode” may refer to a polynucleotide sequence containing a label and a target-binding region that is gene-specific. A probabilistic barcode may also be a polynucleotide sequence that can be used for probabilistic barcoding. A probabilistic barcode can be used to quantify a target in a sample. A probabilistic barcode can be used to control for errors that may occur after the label has been attached to the target. For example, a probabilistic barcode can be used to evaluate amplification errors or sequencing errors. A probabilistic barcode attached to a target may be called a probabilistic barcode-target or probabilistic barcode-tag-target.

[0090] As used herein, the term “probabilistic barcoding” may refer to the random labeling of nucleic acids (e.g., barcoding). Probabilistic barcoding can use a recursive Poisson strategy to associate and quantify the labels associated with a target. As used herein, the term “probabilistic barcoding” may be used synonymously with “probabilistic labeling.” As used herein, the term “target” may refer to a composition to which a barcode (e.g., a probabilistic barcode) can be attached. Exemplary and preferred targets for analysis by the methods, devices, and systems of this disclosure include oligonucleotides, DNA, RNA, mRNA, microRNA, and tRNA. The target may be single-stranded or double-stranded. In some embodiments, the target may be a protein, peptide, or polypeptide. In some embodiments, the target may be a lipid. As used herein, “target” may be used synonymously with “species.”

[0091] As used herein, the term “reverse transcriptase” can refer to a group of enzymes that possess reverse transcriptase activity (i.e., enzymes that catalyze the synthesis of DNA from an RNA template). Generally, such enzymes include, but are not limited to, retroviral reverse transcriptases, retrotransposon reverse transcriptases, retroplasmid reverse transcriptases, retroron reverse transcriptases, bacterial reverse transcriptases, group II intron-derived reverse transcriptases, and their mutants, variants, or derivatives. Non-retroviral reverse transcriptases include non-LTR retrotransposon reverse transcriptases, retroplasmid reverse transcriptases, retroron reverse transcriptases, and group II intron reverse transcriptases. Examples of group II intron reverse transcriptases include Lactococcus lactis LI.LtrB intron reverse transcriptase, Thermosynechococcus elongatus TeI4c intron reverse transcriptase, and Geobacillus stearothermophilus GsI-IIC intron reverse transcriptase. Other classes of reverse transcriptases include many classes of non-retroviral reverse transcriptases (namely, retrons, group II introns, and diversity-generating retroelements).

[0092] The terms “universal adapter primer,” “universal primer adapter,” or “universal adapter sequence” are used synonymously to refer to nucleotide sequences that can be used to hybridize with a barcode (e.g., a probabilistic barcode) to generate a gene-specific barcode. A universal adapter sequence may be, for example, a known sequence that is universal across all barcodes used in the methods of this disclosure. For example, if multiple targets are labeled using the methods disclosed herein, each of the target-specific sequences may be ligated to the same universal adapter sequence. In some embodiments, more than one universal adapter sequence may be used in the methods disclosed herein. For example, if multiple targets are labeled using the methods disclosed herein, at least two of the target-specific sequences may be ligated to different universal adapter sequences. A universal adapter primer and its complement may be contained in two oligonucleotides, one containing a target-specific sequence and the other containing a barcode. For example, a universal adapter sequence may be part of an oligonucleotide containing a target-specific sequence for generating a nucleotide sequence complementary to a target nucleic acid. A second oligonucleotide containing complementary sequences to the barcode and universal adapter sequences can hybridize with the nucleotide sequence to generate a target-specific barcode (e.g., a target-specific probabilistic barcode). In some embodiments, the universal adapter primer has a different sequence from the universal PCR primer used in the methods of this disclosure.

[0093] barcode Barcoding, such as probabilistic barcoding, is described, for example, in U.S. Patent Application Publication No. 20150299784, International Publication No. 2015031691, and Fu et al, Proc Natl Acad Sci USA 2011 May 31; 108(22):9026-31, the contents of which are incorporated herein by reference in their entirety. In some embodiments, the barcodes disclosed herein may be probabilistic barcodes, which may be polynucleotide sequences, that can be used to probabilistically label (e.g., barcodes, tags) targets. The barcode is a probabilistic barcode where the ratio of the number of different barcode sequences to the number of occurrences of any of the targets to be labeled is 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, Alternatively, if the ratio is approximately 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, or any number or range between any two of these values, it can be called a probabilistic barcode. The target may be an mRNA species containing mRNA molecules with identical or nearly identical sequences.The barcode is a probabilistic barcode with a ratio of the number of different barcode sequences and the number of occurrences of any of the targets to be labeled, such that the ratio is at least 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 9 A barcode can be called a probabilistic barcode if its ratio is 0:1, 100:1, or at most 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, or 100:1. The barcode sequence of a probabilistic barcode is sometimes called a molecular label.

[0094] A barcode, such as a probabilistic barcode, may include one or more labels. Exemplary labels include universal labels, cellular labels, barcode sequences (e.g., molecular labels), sample labels, plate labels, spatial labels, and / or pre-spatial labels. Figure 1 shows an exemplary barcode 104 having a spatial label. Barcode 104 may include a 5' amine that can link the barcode to a solid support 105. The barcode may include a universal label, a dimension label, a spatial label, a cellular label, and / or a molecular label. The order of the various labels within the barcode (including, but not limited to, universal labels, dimension labels, spatial labels, cellular labels, and molecular labels) may vary. For example, as shown in Figure 1, the universal label may be the label furthest to the 5' end, and the molecular label may be the label furthest to the 3' end. The spatial label, dimension label, and cellular label may be in any order. In some embodiments, the universal label, spatial label, dimension label, cellular label, and molecular label are in any order. The barcode may include a target-binding region. The target-binding region can interact with a target in the sample (e.g., target nucleic acid, RNA, mRNA, DNA). For example, the target-binding region may include an oligo(dT) sequence that can interact with the poly(A) tail of mRNA. In some cases, the barcode labels (e.g., universal labels, dimensional labels, spatial labels, cellular labels, and barcode sequences) may be separated by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 or more nucleotides.

[0095] Labels, such as cell labels, may contain a specific set of nucleic acid subsequences of a predetermined length, for example, each of 7 nucleotides (equal to the number of bits used in some Hamming error correction codes), and may be designed to provide error correction capabilities. The set of error correction subsequences may contain 7 nucleotide sequences, and any combination of pairs of sequences in the set may be designed to represent a predetermined "genetic distance" (or number of mismatched bases), for example, the set of error correction subsequences may be designed to represent a genetic distance of 3 nucleotides. In this case, it may be possible to detect or correct amplification errors or sequencing errors by scrutinizing the error correction sequences of the set of sequence data of the labeled target nucleic acid molecule (as fully described below). In some embodiments, the length of the nucleic acid subsequence used to generate the error correction code may vary, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 31, 40, 50 nucleotides long, or approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 30, 31, 40, 50 nucleotides long, or a number or range of nucleotide lengths between any two of these values. In some embodiments, nucleic acid subsequences of other lengths may be used to generate the error correction code.

[0096] The barcode may include a target-binding region. The target-binding region can interact with a target in the sample. The target may be or may include ribonucleic acid (RNA), messenger RNA (mRNA), microRNA, small interfering RNA (siRNA), RNA degradation products, RNA each containing a poly(A) tail, or any combination thereof. In some embodiments, multiple targets may include deoxyribonucleic acid (DNA).

[0097] In some embodiments, the target-binding region may include an oligo(dT) sequence that can interact with the poly(A) tail of mRNA. One or more of the barcode labels (e.g., universal labels, dimensional labels, spatial labels, cellular labels, and barcode sequences (e.g., molecular labels)) may be separated by spacers from one or two other of the remaining barcode labels. The spacers may be, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 nucleotides, or more. In some embodiments, none of the barcode labels are separated by spacers.

[0098] Universal Sign A barcode may include one or more universal labels. In some embodiments, one or more universal labels may be the same for all barcodes in a set of barcodes attached to a given solid support. In some embodiments, one or more universal labels may be the same for all barcodes attached to multiple beads. In some embodiments, a universal label may include a nucleic acid sequence that can hybridize to a sequencing primer. The sequencing primer can be used to sequence the barcode containing the universal label. The sequencing primer (e.g., a universal sequencing primer) may include sequencing primers that accompany a high-throughput sequencing platform. In some embodiments, a universal label may include a nucleic acid sequence that can hybridize to a PCR primer. In some embodiments, a universal label may include a nucleic acid sequence that can hybridize to both a sequencing primer and a PCR primer. The nucleic acid sequence of a universal label that can hybridize to a sequencing primer or a PCR primer may be called a primer binding site. A universal label may include a sequence that can be used to initiate the transcription of a barcode. A universal label may include a sequence that can be used to extend a barcode or a region within a barcode. A universal label may be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 nucleotides long, or approximately 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 nucleotides long, or a number or range of nucleotides between any two of these values. For example, a universal label may contain at least approximately 10 nucleotides. A universal label may be at least 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides long, or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides long.In some embodiments, the cleavable linker or modified nucleotide may be part of a universal labeling sequence so that the barcode can be cleaved from the support.

[0099] Dimensional Sign A barcode may include one or more dimension labels. In some embodiments, the dimension label may include a nucleic acid sequence that provides information about the dimension in which the labeling (e.g., probabilistic labeling) occurred. For example, a dimension label may provide information about the time when a target was barcoded. A dimension label may be associated with the time of barcoding (e.g., probabilistic barcoding) within a sample. A dimension label may be activated at the time of labeling. Different dimension labels may be activated at different time points. Dimension labels provide information about a target, a group of targets, and / or the order in which a sample was barcoded. For example, a population of cells may be barcoded in the G0 phase of the cell cycle. Cells may be re-pulsed with barcodes (e.g., probabilistic barcodes) in the G1 phase of the cell cycle. Cells may be re-pulsed with barcodes in the S phase of the cell cycle, etc. Each pulse (e.g., each phase of the cell cycle) may include a different dimension label. Thus, dimension labels provide information about which target was labeled at which phase of the cell cycle. Dimension labels allow for the investigation of numerous different biological time points. Exemplary biological time points, though not limited to these, include the cell cycle, transcription (e.g., transcription initiation), and transcript degradation. In another example, a sample (e.g., cells, a population of cells) can be labeled before and / or after treatment with a drug and / or therapy. Changes in the copy number of distinct targets can indicate the sample's response to a drug and / or therapy.

[0100] Dimensional labels may be activatable. Activatable dimensional labels can be activated at specific points in time. Activatable labels can, for example, be constitutively activated (e.g., not turned off). Activatable dimensional labels can, for example, be reversibly activated (e.g., activatable dimensional labels can be turned on and off). For example, a dimensional label may be reversibly activated at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 times, or more. In some embodiments, dimensional labels can be activated by fluorescence, light, chemical events (e.g., cleavage, ligation of another molecule, addition of modifications (e.g., pegylation, SUMOylation, acetylation, methylation, deacetylation, demethylation)), photochemical events (e.g., photocaging), and introduction of non-native nucleotides.

[0101] In some embodiments, the dimension marker may be the same for all barcodes (e.g., probabilistic barcodes) attached to a given solid support (e.g., beads), but may be different for each different solid support (e.g., beads). In some embodiments, at least 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100% of the barcodes on the same solid support may contain the same dimension marker. In some embodiments, at least 60% of the barcodes on the same solid support may contain the same dimension marker. In some embodiments, at least 95% of the barcodes on the same solid support may contain the same dimension marker.

[0102] Multiple solid supports (e.g., beads) have 10 6One or more intrinsic dimension-labeling sequences may be provided. The dimension labels may be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides long, or approximately 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides long, or a number or range of nucleotides between any two of these values. The dimension labels may be at least 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides long, or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides long. The dimension label may contain approximately 5 to 200 nucleotides. The dimension label may contain approximately 10 to 150 nucleotides. The dimension label may have a length of approximately 20 to 125 nucleotides.

[0103] spatial sign A barcode may include one or more spatial markers. In some embodiments, the spatial marker may include a nucleic acid sequence that provides information about the spatial orientation of the target molecule to which the barcode is attached. The spatial marker may be associated with the coordinates of a sample. The coordinates may be fixed coordinates. For example, the coordinates may be fixed relative to the substrate. The spatial marker may be relative to a two-dimensional or three-dimensional grid. The coordinates may be fixed relative to a marker. The marker may be identifiable in space. The marker may be a structure that can be imaged. The marker may be a biological structure, such as an anatomical marker. The marker may be a cellular marker, such as an organelle. The marker may be a non-natural marker such as a color code, barcode, magnetic properties, fluorescent substance, radioactivity, or a structure with an identifiable identifier such as a unique size or shape. The spatial marker may be associated with a physical compartment (e.g., a well, container, or droplet). In some embodiments, multiple spatial markers are used together to encode one or more locations in space.

[0104] The spatial identifier may be the same for all barcodes attached to a given solid support (e.g., beads), but may be different for each different solid support (e.g., beads). In some embodiments, the percentage of barcodes on the same solid support that contain the same spatial identifier may be 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%, or approximately 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%, or any number or range between any two of these values. In some embodiments, the percentage of barcodes on the same solid support containing the same spatial marker may be at least 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100%, or at most 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100%. In some embodiments, at least 60% of the barcodes on the same solid support may contain the same spatial marker. In some embodiments, at least 95% of the barcodes on the same solid support may contain the same spatial marker.

[0105] Multiple solid supports (e.g., beads) have 10 6One or more unique spatial label sequences may be provided. The spatial label may be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides long, or approximately 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides long, or a number or range of nucleotides between any two of these values. The spatial label may be at least 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides long, or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides long. The spatial label may contain approximately 5 to 200 nucleotides. The spatial label may contain approximately 10 to 150 nucleotides. The spatial label may have a length of approximately 20 to 125 nucleotides.

[0106] cell labeling A barcode (e.g., a probabilistic barcode) may include one or more cell labels. In some embodiments, the cell labels may include nucleic acid sequences that provide information for determining which target nucleic acids originate from which cells. In some embodiments, the cell labels are identical for all barcodes attached to a given solid support (e.g., beads), but different for each different solid support (e.g., beads). In some embodiments, the percentage of barcodes on the same solid support that contain the same cell label may be 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%, or approximately 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%, or a number or range between any two of these values. In some embodiments, the percentage of barcodes on the same solid support containing the same cell label may be 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100%, or approximately 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100%. For example, at least 60% of the barcodes on the same solid support may contain the same cell label. In another example, at least 95% of the barcodes on the same solid support may contain the same cell label.

[0107] Multiple solid supports (e.g., beads) have 10 6One or more than one unique cell labeling sequences may be provided. The cell label may be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or about 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50 nucleotides in length, or a number or range of nucleotides between any two of these values. The cell label may be at least 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length, or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides in length. For example, the cell label may contain about 5 to about 200 nucleotides. As another example, the cell label may contain about 10 to about 150 nucleotides. As yet another example, the cell label may contain about 20 to about 125 nucleotides in length.

[0108] Barcode sequence The barcode may contain one or more barcode sequences. In some embodiments, the barcode sequence may contain a nucleic acid sequence that provides information for identifying a specific type of target nucleic acid species that hybridizes to the barcode. The barcode sequence may contain a nucleic acid sequence that provides a means (e.g., provides a rough approximation) for counting the specific occurrences of a target nucleic acid species that hybridizes to the barcode (e.g., the target-binding region).

[0109] In some embodiments, a diverse set of barcode sequences is attached to a given solid support (e.g., beads). In some embodiments, 10 2 、10 3 、10 4 、10 5 、10 6 、10 7 、10 8 、10 9 individual, or about 10 2 、10 3 、10 4 、10 5, 10 6 , 10 7 , 10 8 , 10 9 There may be a number or range of unique molecular label sequences, or any two of these values. For example, a group of barcodes may contain about 6561 barcode sequences, each having distinct sequences. In another example, a group of barcodes may contain about 65536 barcode sequences, each having distinct sequences. In some embodiments, at least 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , or 10 9 one, or up to 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , or 10 9 There may be individual unique barcode sequences. The unique molecular label sequences may be attached to a given solid support (e.g., beads).

[0110] The length of a barcode may vary depending on the implementation. For example, a barcode may be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 nucleotides long, or approximately 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 nucleotides long, or a number or range of nucleotides between any two of these values. As another example, a barcode may be at least 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides long, or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides long.

[0111] molecular label A barcode (e.g., a probabilistic barcode) may include one or more molecular labels. The molecular labels may include a barcode sequence. In some embodiments, the molecular labels may include a nucleic acid sequence that provides information for identifying a specific type of target nucleic acid species hybridized to the barcode. The molecular target may include a nucleic acid sequence that provides a means for counting the specific occurrences of the target nucleic acid species hybridized to the barcode (e.g., a target-binding region).

[0112] In some embodiments, a diverse set of molecular labels is attached to a given solid support (e.g., beads). In some embodiments, 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 10, or about 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 There may be a number or range of unique molecular label sequences, or any two of these values. For example, multiple barcodes may contain about 6561 molecular labels having distinct sequences. As another example, multiple barcodes may contain about 65536 molecular labels having distinct sequences. In some embodiments, at least 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , or 10 9 one, or up to 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , or 10 9There may be individual molecular labeling sequences. A barcode having a specific molecular labeling sequence may be attached to a given solid support (e.g., beads).

[0113] For probabilistic barcoding using multiple probabilistic barcodes, the ratio of the number of different molecular label sequences to the number of occurrences of any of the targets is 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1. The target may be 90:1, 100:1, or approximately 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, 100:1, or a number or range between any two of these values. The target may be an mRNA species containing mRNA molecules with identical or nearly identical sequences. In some embodiments, the ratio of the number of different molecularly labeled sequences to the number of occurrences of any of the targets is at least 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1 , 90:1, or 100:1, or at most 1:1, 2:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 30:1, 40:1, 50:1, 60:1, 70:1, 80:1, 90:1, or 100:1.

[0114] The molecular label may be 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 nucleotides long, or approximately 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 nucleotides long, or a number or range of nucleotides between any two of these values. The molecular label may be at least 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides long, or at most 1, 2, 3, 4, 5, 10, 15, 20, 25, 30, 35, 40, 45, 50, 100, 200, or 300 nucleotides long.

[0115] Target binding region The barcode may include one or more target-binding regions, such as capture probes. In some embodiments, the target-binding region can hybridize with a target of interest. In some embodiments, the target-binding region may include a nucleic acid sequence that specifically hybridizes to the target (e.g., a target nucleic acid, a target molecule, e.g., a cellular nucleic acid to be analyzed), for example, to a specific gene sequence. In some embodiments, the target-binding region may include a nucleic acid sequence that can attach (e.g., hybridize) to a specific location on a particular target nucleic acid. In some embodiments, the target-binding region may include a nucleic acid sequence that is capable of specific hybridization to restriction enzyme site overhangs (e.g., EcoRI sticky end overhangs). The barcode can then be ligated to any nucleic acid molecule containing a sequence complementary to the restriction site overhang.

[0116] In some embodiments, the target-binding region may include a non-specific target nucleic acid sequence. The non-specific target nucleic acid sequence may refer to a sequence capable of binding to multiple target nucleic acids independent of the specific sequence of the target nucleic acid. For example, the target-binding region may include a random multimer sequence or an oligo(dT) sequence that hybridizes to a poly(A) tail on an mRNA molecule. The random multimer sequence may be, for example, a random dimer, trimer, tetramer, pentamer, hexamer, heptamer, octamer, noumer, decamer, or a higher-order multimer sequence of any length. In some embodiments, the target-binding region is the same for all barcodes attached to a given bead. In some embodiments, the target-binding regions of multiple barcodes attached to a given bead may include two or more different target-binding sequences. The target-binding region may be 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 nucleotides long, or approximately 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 nucleotides long, or a number or range of nucleotides between any two of these values. The target-binding region may be up to approximately 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 nucleotides long, or longer.

[0117] In some embodiments, the target-binding region may include an oligo(dT) that can hybridize with mRNA containing a polyadenylated end. The target-binding region may also be gene-specific. For example, the target-binding region may be configured to hybridize to a specific region of the target. The target-binding region may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 nucleotides long, or approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 nucleotides long, or a number or range of nucleotides between any two of these values. The target-binding region may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides long, or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides long. The target-binding region may be approximately 5 to 30 nucleotides long. When a barcode contains a gene-specific target-binding region, the barcode may be referred to herein as a gene-specific barcode.

[0118] Directional characteristics A probabilistic barcode (e.g., a probabilistic barcode) may include one or more directional properties that can be used to orient (e.g., align) the barcode. The barcode may include portions for isoelectric focusing. Different barcodes may constitute different isoelectric focusing points. Once such barcodes are introduced into a sample, the sample can be subjected to isoelectric focusing to orient the barcodes in known directions. Thus, using directional properties, a known map of the barcodes within the sample can be created. Exemplary directional properties include electrophoretic mobility (e.g., based on the size of the barcode), isoelectric focus, spin, conductivity, and / or self-assembly. For example, a barcode with a self-assembly directional property may self-assemble in a specific direction (e.g., nucleic acid nanostructure) when activated.

[0119] affinity properties A barcode (e.g., a probabilistic barcode) may include one or more affinity properties. For example, a spatial label may include affinity properties. The affinity properties may include chemical and / or biological parts that facilitate the binding of the barcode to another entity (e.g., a cell receptor). For example, the affinity properties may include an antibody, e.g., an antibody specific to a particular part on a sample (e.g., a receptor). In some embodiments, the antibody can lead the barcode to a particular cell type or molecule. It can label (e.g., probabilistically) targets that are on and / or near a particular cell type. Because the antibody can lead the barcode to a specific location, in some embodiments, the affinity properties can provide spatial information in addition to the nucleotide sequence of the spatial label. The antibody may be a therapeutic antibody, e.g., a monoclonal antibody or a polyclonal antibody. The antibody may be humanized or chimeric. The antibody may be a naked antibody or a fusion antibody.

[0120] Antibodies may be full-length immunoglobulin molecules (i.e., naturally occurring or formed by normal immunoglobulin gene fragment recombination processes) (e.g., IgG antibodies) or immunologically active (i.e., specifically binding) portions of immunoglobulin molecules, such as antibody fragments. The antibody fragment may be part of an antibody such as F(ab')2, Fab', Fab, Fv, and sFv. In some embodiments, the antibody fragment can bind to the same antigen recognized by the full-length antibody. The antibody fragment may include isolated fragments of the variable regions of an antibody, such as an "Fv" fragment consisting of heavy and light chain variable regions, and a recombinant single-chain polypeptide molecule ("scFv protein") in which the light chain variable region and the heavy chain variable region are linked by a peptide linker. Examples of antibodies, but not limited to these, include antibodies against cancer cells, antibodies against viruses, antibodies that bind to cell surface receptors (CD8, CD34, CD45), and therapeutic antibodies.

[0121] Universal Adapter Primer A barcode may contain one or more universal adapter primers. For example, a gene-specific barcode, such as a gene-specific probabilistic barcode, may contain a universal adapter primer. The universal adapter primer can point to a nucleotide sequence that is universal across all barcodes. The universal adapter primer can be used to construct a gene-specific barcode. The universal adapter primer may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 nucleotides long, or approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 nucleotides long, or a number or range of nucleotides between any two of these. The universal adapter primer may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides long, or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides long. The universal adapter primer may also be 5 to 50 nucleotides long.

[0122] Linker If a barcode contains more than one type of label (e.g., more than one cellular label or more than one barcode sequence, such as one molecular label), the labels may be dispersed in a linker label sequence. The linker label sequence may be at least about 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 nucleotides long, or longer. The linker label sequence may be at most about 5, 10, 15, 20, 25, 30, 35, 40, 45, or 50 nucleotides long, or longer. In some cases, the linker label sequence is 12 nucleotides long. The linker label sequence can be used to facilitate the synthesis of barcodes. The linker label may include error correction (e.g., Hamming) codes.

[0123] solid support Barcodes, such as probabilistic barcodes, disclosed herein may, in some embodiments, be associated with a solid support. The solid support may be, for example, synthetic particles. In some embodiments, some or all of the barcode sequences, such as molecular labels of probabilistic barcodes (e.g., a first barcode sequence) of multiple barcodes (e.g., a first barcode sequence) on a solid support, differ by at least one nucleotide. Cellular labels of barcodes on the same solid support may be the same. Cellular labels of barcodes on different solid supports may differ by at least one nucleotide. For example, the first cellular labels of the first multiple barcodes on a first solid support may have the same sequence, and the second cellular labels of the second multiple barcodes on a second solid support may have the same sequence. The first cellular labels of the first multiple barcodes on a first solid support and the second cellular labels of the second multiple barcodes on a second solid support may differ by at least one nucleotide. The cellular labels may be, for example, about 5 to 20 nucleotides long. The barcode sequence may be, for example, about 5 to 20 nucleotides long. The synthetic particles may be, for example, beads.

[0124] The beads may be, for example, silica gel beads, pore-controlled glass beads, magnetic beads, DynaBeads, Sephadex / Sepharose beads, cellulose beads, polystyrene beads, or any combination thereof. The beads may also contain materials such as polydimethylsiloxane (PDMS), polystyrene, glass, polypropylene, agarose, gelatin, hydrogel, paramagnetic material, ceramic, plastic, glass, methylstyrene, acrylic polymer, titanium, latex, Sepharose, cellulose, nylon, silicone, or any combination thereof.

[0125] In some embodiments, the beads may be polymer beads functionalized with barcodes or probabilistic barcodes, such as deformable beads or gel beads (e.g., gel beads from 10X Genomics (San Francisco, California)). In some embodiments, the gel beads may comprise a polymer-based gel. Gel beads can be produced, for example, by encapsulating one or more polymer precursors in droplets. Exposure of the polymer precursors to an accelerator (e.g., tetramethylethylenediamine (TEMED)) can produce gel beads.

[0126] In some embodiments, the particles may be degradable. For example, polymer beads can be dissolved, melted, or degraded under desired conditions. Desired conditions may include environmental conditions. Desired conditions can result in the dissolution, melting, or degradation of polymer beads in a controlled manner. Gel beads can be dissolved, melted, or degraded by chemical, physical, biological, thermal, magnetic, electrical, or optical stimuli, or any combination thereof.

[0127] Analytes and / or reagents, such as oligonucleotide barcodes, may be coupled / immobilized, for example, to the inner surface of gel beads (e.g., the interior accessible by diffusion of the material used to generate the oligonucleotide barcodes and / or the oligonucleotide barcodes) and / or to the outer surface of gel beads or other microcapsules described herein. Coupling / immobilization may be by any form of chemical bonding (e.g., covalent bonding, ionic bonding) or physical phenomena (e.g., van der Waals forces, dipole interactions, etc.). In some embodiments, the coupling / immobilization of reagents with gel beads or any other microcapsules described herein may be reversible, for example, via an unstable portion (e.g., via a chemical crosslinker containing a chemical crosslinker described herein). Applying a stimulus can cleave the unstable portion and release the immobilized reagent. In some embodiments, the unstable portion is a disulfide bond. For example, if oligonucleotide barcodes are immobilized to gel beads via a disulfide bond, exposing the disulfide bond to a reducing agent can cleave the disulfide bond and release the oligonucleotide barcode from the beads. The unstable portion may be included as part of the gel beads or microcapsules, as part of a chemical linker that links the reagent or analyte to the gel beads or microcapsules, and / or as part of the reagent or analyte. In some embodiments, at least one of the multiple barcodes may be immobilized on the particle, partially immobilized on the particle, encapsulated within the particle, partially encapsulated within the particle, or any combination thereof.

[0128] In some embodiments, the gel beads may contain a wide variety of polymers, including, but are not limited to, polymers, heat-sensitive polymers, photosensitive polymers, magnetic polymers, pH-sensitive polymers, salt-sensitive polymers, chemical-sensitive polymers, polyelectrolytes, polysaccharides, peptides, proteins, and / or plastics. Examples of polymers include, but are not limited to, poly(N-isopropylacrylamide) (PNIPAAm), poly(styrene sulfonate) (PSS), poly(allylamine) (PAAm), poly(acrylic acid) (PAA), poly(ethyleneimine) (PEI), poly(diallyldimethylammonium chloride) (PDADMAC), poly(pyrrole) (poly(pyrolle)) (PPy), poly(vinylpyrrolidone) (PVPON), and poly(vinylpyrrolidone). Examples of materials include lysine (PVP), poly(methacrylic acid) (PMAA), poly(methyl methacrylate) (PMMA), polystyrene (PS), poly(tetrahydrofuran) (PTHF), poly(phthalaldehyde) (PTHF), poly(hexyl viologen) (PHV), poly(L-lysine) (PLL), poly(L-arginine) (PARG), and poly(lactic acid-co-glycolic acid) (PLGA).

[0129] Numerous chemical stimuli can be used to induce the breakdown, dissolution, or decomposition of beads. Examples of such chemical changes, though not limited to these, include pH-mediated changes to the bead wall, collapse of the bead wall due to chemical cleavage of crosslinking bonds, induction of depolymerization of the bead wall, and bead wall switching reactions. Additionally, bulk changes can be used to induce bead breakdown.

[0130] Furthermore, bulk or physical changes to microcapsules induced by various stimuli offer numerous advantages in the design of capsules for reagent release. These bulk or physical changes occur on a macroscopic scale, and bead rupture is a result of mechanophysical forces induced by the stimulus. Examples of such processes, though not limited to these, include pressure-induced rupture, bead wall melting, or changes in the porosity of the bead wall. Furthermore, biological stimuli can be used to induce the destruction, dissolution, or degradation of the beads. Generally, biological inducers are similar to chemical inducers, but in many cases, biomolecules or molecules commonly found in biological systems, such as enzymes, peptides, saccharides, fatty acids, and nucleic acids, are used. For example, the beads may contain polymers having peptide crosslinkages that are sensitive to cleavage by specific proteases. More specifically, one example may include microcapsules containing GFLGK peptide crosslinkages. Upon addition of a biological inducer such as the protease cathepsin B, the peptide crosslinkages of the shell wall are cleaved, releasing the contents of the beads. In other cases, the protease can be thermally activated. In another example, the beads contain a shell wall containing cellulose. The addition of the hydrolytic enzyme chitosan acts as a biological inducer for cleavage of the cellulose bonds, depolymerization of the shell wall, and release of its internal contents. Furthermore, beads can be induced to release their contents when heat is applied. Temperature changes can cause various changes in beads. Heat changes can cause the beads to melt and the bead walls to collapse. In other cases, heat can increase the internal pressure of the bead's internal components, causing the bead to burst or explode. In yet other cases, heat can deform the beads, causing them to shrink and become dehydrated. Heat can also act on heat-sensitive polymers within the bead walls, causing the beads to break down.

[0131] By incorporating magnetic nanoparticles into the bead walls of microcapsules, it may be possible to induce bead rupture and guide the beads in an array. The devices of this disclosure may include magnetic beads for any purpose. In one example, the incorporation of Fe3O4 nanoparticles into a polymer electrolyte containing beads induces rupture in the presence of an oscillating magnetic field stimulus. Furthermore, beads can be destroyed, dissolved, or decomposed as a result of electrical stimulation. Similar to the magnetic particles described in the previous section, electrosensitive beads can enable both bead rupture induction and other functions such as alignment in an electric field, electrical conductivity, or redox reactions. In one example, beads containing electrosensitive material can be aligned in an electric field to control the release of internal reagents. In another example, an electric field can induce redox reactions within the bead wall itself, thereby increasing porosity. Furthermore, beads can be destroyed using photostimulation. Numerous photocatalysts are conceivable, including systems using various molecules such as nanoparticles and chromophores capable of absorbing photons in specific wavelength ranges. For example, metal oxide coatings can be used as capsule catalysts. UV irradiation of SiO2-coated polymer electrolyte capsules can cause the bead walls to collapse. In yet another example, photoswitchable materials such as azobenzene groups can be incorporated into the bead walls. When UV or visible light is applied, these chemicals undergo reversible cis-trans isomerization upon photon absorption. In this embodiment, the incorporation of photon switches results in bead walls that may collapse or become more porous upon application of the photocatalyst.

[0132] For example, in a non-restrictive example of barcoding (e.g., probabilistic barcoding) illustrated in Figure 2, after introducing cells such as single cells into multiple microwells of a microwell array in block 208, beads can be introduced into multiple microwells of the microwell array in block 212. Each microwell may contain one bead. The beads may contain multiple barcodes. The barcodes may contain a 5' amine region attached to the bead. The barcodes may contain a universal label, a barcode sequence (e.g., a molecular label), a target-binding region, or any combination thereof.

[0133] The barcodes disclosed herein may be attached to (e.g., adhering to) a solid support (e.g., beads). Each barcode attached to a solid support may include a barcode sequence selected from the group including at least 100 or 1000 barcode sequences having unique sequences. In some embodiments, different barcodes attached to a solid support may include barcodes having different sequences. In some embodiments, a certain percentage of barcodes attached to a solid support may include the same cell label. For example, the percentage may be 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%, or approximately 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, 100%, or a number or range between any two of these values. As another example, the percentage may be at least 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100%, or at most 60%, 70%, 80%, 85%, 90%, 95%, 97%, 99%, or 100%. In some embodiments, barcodes associated with solid supports may have the same cell marker. Barcodes associated with different solid supports may have different cell markers selected from the group comprising at least 100 or 1000 cell markers having unique sequences.

[0134] The barcodes disclosed herein may be attached to (e.g., adhering to) a solid support (e.g., beads). In some embodiments, barcoding of multiple targets in a sample can be performed using a solid support comprising multiple synthetic particles to which multiple barcodes are attached. In some embodiments, the solid support may comprise multiple synthetic particles to which multiple barcodes are attached. Spatial labeling of multiple barcodes on different solid supports may differ by at least one nucleotide. The solid support may comprise, for example, two-dimensional or three-dimensional barcodes. The synthetic particles may be beads. The beads may be silica gel beads, pore-controlled glass beads, magnetic beads, DynaBeads, Sephadex / Sepharose beads, cellulose beads, polystyrene beads, or any combination thereof. The solid support may comprise a polymer, matrix, hydrogel, needle array device, antibody, or any combination thereof. In some embodiments, the solid support may be floating. In some embodiments, the solid support may be embedded in a semi-solid or solid array. The barcodes may not be attached to the solid support. The barcodes may comprise individual nucleotides. The barcode may be attached to the substrate.

[0135] As used herein, the terms “moored,” “attached,” and “immobilized” are used synonymously and may refer to covalent or non-covalent means for attaching a barcode to a solid support. Any of the various different solid supports can be used as a solid support for attaching a pre-synthesized barcode or for solid-phase synthesis of a barcode in situ.

[0136] In some embodiments, the solid support is a bead. The bead may include one or more types of solid, porous, or hollow spheres, balls, bearings, cylinders, or other similar configurations that can immobilize nucleic acids (e.g., covalently or non-covalently). The bead may be made of, for example, plastic, ceramic, metal, polymer material, or any combination thereof. The bead may be spherical (e.g., microspheres) or may be individual particles having a non-spherical or irregular shape, such as a cube, pyramidal, cylindrical, conical, rectangular, or disc. In some embodiments, the bead may have a non-spherical shape.

[0137] The beads may include, but are not limited to, a variety of materials including paramagnetic materials (e.g., magnesium, molybdenum, lithium, and tantalum), superparamagnetic materials (e.g., ferrite (Fe3O4; magnetite) nanoparticles), ferromagnetic materials (e.g., iron, nickel, cobalt, some alloys thereof, and some rare earth metal compounds), ceramics, plastics, glass, polystyrene, silica, methylstyrene, acrylic polymers, titanium, latex, Sepharose, agarose, hydrogels, polymers, cellulose, nylon, or any combination thereof. In some embodiments, the beads (e.g., beads to which a label is attached) are hydrogel beads. In some embodiments, the beads contain a hydrogel.

[0138] Some embodiments disclosed herein include one or more particles (e.g., beads). Each particle may contain a plurality of oligonucleotides (e.g., barcodes). Each of the plurality of oligonucleotides may contain a barcode sequence (e.g., a molecular label sequence), a cell label, and a target-binding region (e.g., an oligo(dT) sequence, a gene-specific sequence, a random multimer, or a combination thereof). The cell label sequences of each of the plurality of oligonucleotides may be the same. The cell label sequences of oligonucleotides on different particles may be different so as to be able to distinguish oligonucleotides on different particles. The number of different cell label sequences may vary depending on the implementation. In some embodiments, the number of cell-labeled sequences is 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 , 10 7 , 10 8 , 10 9 , or approximately 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 , 10 7 , 10 8 , 10 9 , or a number or range between any two of these values. In some embodiments, the number of cell-labeled sequences may be at least 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 , 10 7, 10 8 , or 10 9 It may be, or up to 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 , 10 7 , 10 8 , or 10 9 This may also be the case. In some embodiments, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 or less of the plurality of particles contain oligonucleotides having the same cell sequence. In some embodiments, the plurality of particles containing oligonucleotides having the same cell sequence may be up to 0.1%, 0.2%, 0.3%, 0.4%, 0.5%, 0.6%, 0.7%, 0.8%, 0.9%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, or more. In some embodiments, none of the plurality of particles have the same cell labeling sequence.

[0139] Multiple oligonucleotides on each particle may contain different barcode sequences (e.g., molecular labels). In some embodiments, the number of barcode sequences may be 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, or 10 6 , 10 7 , 10 8 , 10 9, or about 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 10 7 10 8 10 9 or a number or range between any two of these values. In some embodiments, the number of barcode arrays is at least 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 10 7 10 8 or 10 9 or up to 10, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 10 6 10 7 10 8 or 10 9This may also be the case. For example, at least 100 of the oligonucleotides may contain different barcode sequences. Another example is that in a single particle, at least 100, 500, 1000, 5000, 10000, 15000, 20000, 50000 of the oligonucleotides, a number or range between any two of these values, or more may contain different barcode sequences. Some embodiments provide multiple particles containing barcodes. In some embodiments, the ratio of the appearance (or copies or number) of the target to be labeled and the different barcode sequences may be at least 1:1, 1:2, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:11, 1:12, 1:13, 1:14, 1:15, 1:16, 1:17, 1:18, 1:19, 1:20, 1:30, 1:40, 1:50, 1:60, 1:70, 1:80, 1:90, or greater. In some embodiments, each of the plurality of oligonucleotides further includes sample labeling, universal labeling, or both. The particles may be, for example, nanoparticles or microparticles.

[0140] The size of the beads may vary. For example, the diameter of the beads may be in the range of 0.1 micrometers to 50 micrometers. In some embodiments, the diameter of the beads may be 0.1, 0.5, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50 micrometers, or approximately 0.1, 0.5, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50 micrometers, or any number or range between any two of these values.

[0141] The diameter of the beads may be related to the diameter of the wells in the substrate. In some embodiments, the diameter of the beads may be 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, or approximately 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, or a number or range between any two of these values, longer or shorter than the diameter of the wells. The diameter of the beads may also be related to the diameter of the cells (e.g., single cells confined in the wells of the substrate). In some embodiments, the diameter of the beads may be at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% longer or shorter than the diameter of the well, or up to 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% longer or shorter than the diameter of the well. The diameter of the beads may also be related to the diameter of the cells (e.g., single cells trapped in the wells of the substrate). In some embodiments, the diameter of the beads may be 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 150%, 200%, 250%, 300% longer or shorter than the diameter of the cell by approximately 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 150%, 200%, 250%, 300%, or by a number or range between any two of these values. In some embodiments, the diameter of the beads may be at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 150%, 200%, 250%, or 300% longer or shorter than the diameter of the cells, or up to 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 100%, 150%, 200%, 250%, or 300% longer or shorter.

[0142] The beads may be attached to and / or embedded in a substrate. The beads may be attached to and / or embedded in a gel, hydrogel, polymer, and / or matrix. The spatial location of the beads within the substrate (e.g., gel, matrix, scaffold, or polymer) can be identified using a spatial label present on the barcode on the bead, which can serve as a spatial address. Examples of beads include, but are not limited to, streptavidin beads, agarose beads, magnetic beads, Dynabeads®, MACS® microbeads, antibody conjugate beads (e.g., anti-immunoglobulin microbeads), protein A conjugate beads, protein G conjugate beads, protein A / G conjugate beads, protein L conjugate beads, oligo(dT) conjugate beads, silica beads, silica-like beads, anti-biotin microbeads, anti-fluorescent dye microbeads, and BcMag® carboxyl-terminated magnetic beads.

[0143] The beads may be accompanied by quantum dots or fluorescent dyes (e.g., impregnated) so that the beads fluoresce in one or more optical channels. The beads may be paramagnetic or ferromagnetic by being accompanied by iron oxide or chromium oxide. The beads may be identifiable. For example, the beads can be imaged using a camera. The beads may have a detectable code associated with them. For example, the beads may contain a barcode. The beads may change size, for example, by swelling in an organic or inorganic solution. The beads may be hydrophobic. The beads may be hydrophilic. The beads may be biocompatible.

[0144] A solid support (e.g., a bead) can be visualized. The solid support may include a visualization tag (e.g., a fluorescent dye). The solid support (e.g., a bead) may be engraved with an identifier (e.g., a number). The identifier can be visualized by imaging the bead. Solid supports may contain insoluble, semi-soluble, or insoluble materials. A solid support can be called "functionalized" if it contains linkers, scaffolds, structural blocks, or other reactive parts attached thereto, and a solid support can be called "unfunctionalized" if it lacks such reactive parts attached thereto. Solid supports can be used suspended in solution in microtiter wells, flow-through formats such as columns, or immersion sticks. The solid support may include films, paper, plastics, coated surfaces, flat surfaces, glass, slides, chips, or any combination thereof. The solid support may also take the form of resins, gels, microspheres, or other geometric configurations. The solid support may include flat supports such as silica chips, microparticles, nanoparticles, plates, arrays, capillaries, glass fiber filters, glass surfaces, metal surfaces (steel, gold, silver, aluminum, silicon, and copper), glass supports, plastic supports, silicon supports, chips, filters, films, microwell plates, slides, multiwell plates, or films made of plastic materials (e.g., formed from polyethylene, polypropylene, polyamide, polyvinylidene difluoride), and / or wafers, combs, pins, or needles (e.g., arrays of pins suitable for combinatorial synthesis or analysis), or wafers (e.g., silicon wafers), or beads in pits or arrays of nanoliter wells on flat surfaces such as wafers with or without filter bottoms.

[0145] The solid support may contain a polymer matrix (e.g., a gel, a hydrogel). The polymer matrix may be able to penetrate into the intracellular space (e.g., around organelles). The polymer matrix may be able to pump throughout the circulatory system.

[0146] Substrate and microwell array As used herein, a substrate can refer to a certain type of solid support. A substrate can refer to a solid support which may contain the barcode or probabilistic barcode of the Disclosure. A substrate may, for example, include a plurality of microwells. For example, a substrate may be a well array containing two or more microwells. In some embodiments, a microwell may include a small reaction chamber of a specified volume. In some embodiments, a microwell may contain one or more cells. In some embodiments, a microwell may contain only one cell. In some embodiments, a microwell may contain one or more solid supports. In some embodiments, a microwell may contain only one solid support. In some embodiments, a microwell contains a single cell and a single solid support (e.g., a bead). A microwell may contain the barcode reagent of the Disclosure.

[0147] Barcoding methods This disclosure provides a method for estimating the number of distinct targets located at distinct locations within a physical sample (e.g., tissue, organ, tumor, cell). The method may include placing barcodes (e.g., probabilistic barcodes) near the sample, dissolving the sample, associating barcodes with distinct targets, amplifying the targets, and / or digitally counting the targets. The method may further include analyzing and / or visualizing information derived from spatial labels on the barcodes. In some embodiments, the method includes visualizing multiple targets within the sample. Mapping multiple targets onto a map of the sample may include generating a two-dimensional or three-dimensional map of the sample. The two-dimensional and three-dimensional maps can be generated before or after barcoding (e.g., probabilistically barcoding) the multiple targets within the sample. Visualization of multiple targets within a sample may include mapping multiple targets onto a map of the sample. Mapping multiple targets onto a map of the sample may include generating a two-dimensional or three-dimensional map of the sample. Two-dimensional and three-dimensional maps can be generated before or after barcoding multiple targets in a sample. In some embodiments, two-dimensional and three-dimensional maps can be generated before or after dissolving the sample. Dissolving the sample before or after generating a two-dimensional or three-dimensional map may include heating the sample, contacting the sample with a surfactant, changing the pH of the sample, or any combination thereof.

[0148] In some embodiments, barcoding multiple targets involves hybridizing multiple barcodes with multiple targets to create barcoded targets (e.g., probabilistic barcoded targets). Barcoding multiple targets may also involve generating an indexing library of barcoded targets. The generation of an indexing library of barcoded targets can be carried out using individual supports containing multiple barcodes (e.g., probabilistic barcodes).

[0149] Contact between the sample and the barcode This disclosure provides a method for bringing a sample (e.g., cells) into contact with a substrate of this disclosure. For example, a sample comprising thin sections of cells, organs, or tissues can be brought into contact with a barcode (e.g., a probabilistic barcode). For example, cells can be brought into contact by gravity flow, in which case the cells can settle to form a monolayer. The sample may be thin sections of tissue. The thin sections can be placed on a substrate. The sample may be one-dimensional (e.g., it may form a planar surface). The sample (e.g., cells) can be spread across the substrate, for example, by growing / culturing cells on the substrate. When a barcode is present in the vicinity of a target, the target can hybridize with the barcode. The barcode can be brought into contact with each separate target in a non-depletion ratio so that each separate target may be associated with a separate barcode of this disclosure. To ensure efficient association between the target and the barcode, the target may be bridged with the barcode.

[0150] cell lysis After distributing the cells and barcodes, the cells may be lysed to release the target molecules. Cell lysis can be achieved by any of the following means, for example, by chemical or biochemical means, by osmotic shock, or by thermal lysis, mechanical lysis, or optical lysis. Cells can be lysed by adding a cell lysis buffer containing a surfactant (e.g., SDS, Lithium dodecyl sulfate, Triton X-100, Tween-20, or NP-40), an organic solvent (e.g., methanol or acetone), or a digestive enzyme (e.g., proteinase K, pepsin, or trypsin), or any combination thereof. To increase the attachment of the barcodes to the target, the diffusion rate of the target molecules may be altered, for example, by reducing the temperature and / or increasing the viscosity of the lysis solution.

[0151] In some embodiments, the sample may be dissolved using filter paper. The filter paper can be impregnated with a dissolution buffer on its upper surface. The sample can be applied to the filter paper at a pressure that facilitates the dissolution of the sample and the hybridization of the sample target with the substrate. In some embodiments, dissolution can be carried out by mechanical dissolution, thermal dissolution, optical dissolution, and / or chemical dissolution. Chemical dissolution may involve the use of digestive enzymes such as proteinase K, pepsin, and trypsin. Dissolution can be carried out by adding a dissolution buffer to the substrate. The dissolution buffer may contain Tris HCl. The dissolution buffer may contain at least about 0.01, 0.05, 0.1, 0.5, or 1 M, or higher concentrations of Tris HCl. The dissolution buffer may contain up to about 0.01, 0.05, 0.1, 0.5, or 1 M, or higher concentrations of Tris HCl. The dissolution buffer may contain about 0.1 M of Tris HCl. The pH of the dissolution buffer may be at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or higher. The pH of the lysis buffer may be up to about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or higher. In some embodiments, the pH of the lysis buffer is about 7.5. The lysis buffer may contain a salt (e.g., LiCl). The salt concentration in the lysis buffer may be at least about 0.1, 0.5, or 1 M, or higher. The salt concentration in the lysis buffer may be up to about 0.1, 0.5, or 1 M, or higher. In some embodiments, the salt concentration in the lysis buffer is about 0.5 M. The lysis buffer may contain a surfactant (e.g., SDS, Li dodecyl sulfate, Triton X, Tween, NP-40). The concentration of the surfactant in the dissolution buffer may be at least about 0.0001%, 0.0005%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, or 7%, or higher. The concentration of the surfactant in the dissolution buffer may be at most about 0.0001%, 0.0005%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, or 7%, or higher. In some embodiments, the concentration of the surfactant in the dissolution buffer is about 1% Li dodecyl sulfate. The time used in the dissolution method may depend on the amount of surfactant used.In some embodiments, the more surfactant used, the shorter the time required for dissolution. The dissolution buffer may contain a chelating agent (e.g., EDTA, EGTA). The concentration of the chelating agent in the dissolution buffer may be at least about 1, 5, 10, 15, 20, 25, or 30 mM, or higher. The concentration of the chelating agent in the dissolution buffer may be up to about 1, 5, 10, 15, 20, 25, or 30 mM, or higher. In some embodiments, the concentration of the chelating agent in the dissolution buffer is about 10 mM. The dissolution buffer may contain a reducing agent (e.g., beta-mercaptoethanol, DTT). The concentration of the reducing agent in the dissolution buffer may be at least about 1, 5, 10, 15, or 20 mM, or higher. The concentration of the reducing agent in the dissolution buffer may be up to about 1, 5, 10, 15, or 20 mM, or higher. In some embodiments, the concentration of the reducing agent in the lysis buffer is about 5 mM. In some embodiments, the lysis buffer may contain about 0.1 M Tris HCl, about pH 7.5, about 0.5 M LiCl, about 1% lithium dodecyl sulfate, about 10 mM EDTA, and about 5 mM DTT.

[0152] Lysis can be carried out at a temperature of approximately 4, 10, 15, 20, 25, or 30°C. Lysis can be carried out for approximately 1, 5, 10, 15, or 20 minutes, or longer. Lysified cells may contain at least approximately 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, or 700,000 target nucleic acid molecules, or more. Lysified cells may contain up to approximately 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, or 700,000 target nucleic acid molecules, or more.

[0153] Barcode attachment to target nucleic acid molecules After lysing cells and releasing nucleic acid molecules from them, barcodes on a co-localized solid support can be randomly attached to the nucleic acid molecules. The attachment may involve hybridization between the target recognition region of the barcode and a complementary portion of the target nucleic acid molecule (e.g., the oligo(dT) of the barcode can interact with the poly(A) tail of the target). Assay conditions used for hybridization (e.g., buffer pH, ionic strength, temperature) can be selected to promote the formation of specific and stable hybrids. In some embodiments, multiple probes on a substrate can be attached to the nucleic acid molecules released from lysed cells (e.g., hybridize with probes on the substrate). If the probes contain oligo(dT), the mRNA molecule can hybridize to the probe and be reverse transcribed. The oligo(dT) portion of the oligonucleotide can act as a primer for the first-strand synthesis of the cDNA molecule. For example, in a non-limiting example of barcoding illustrated in Figure 2, at block 216, the mRNA molecule can be hybridized to a barcode on a bead. For example, a single-stranded nucleotide fragment can be hybridized to the target-binding region of a barcode.

[0154] The attachment may further involve ligation between the target recognition region of the barcode and a portion of the target nucleic acid molecule. For example, the target-binding region may contain a nucleic acid sequence that may be capable of specific hybridization to restriction site overhangs (e.g., EcoRI sticky end overhangs). The assay procedure may further involve treating the target nucleic acid with a restriction enzyme (e.g., EcoRI) to create restriction site overhangs. The barcode can then be ligated to any nucleic acid molecule containing a sequence complementary to the restriction site overhangs. The two fragments can be joined using a ligase (e.g., T4 DNA ligase).

[0155] For example, in the non-limiting example of barcoding illustrated in Figure 2, labeled targets (e.g., target-barcode molecules) derived from multiple cells (or multiple samples) may then be pooled in block 220, for example, in a tube. Labelled targets can be pooled, for example, by recovering barcodes and / or beads to which target-barcode molecules are attached. The collection of the attached target-barcode molecules based on a solid support can be carried out using magnetic beads and an externally applied magnetic field. Once the target-barcode molecules are pooled, all further processing can be carried out in a single reaction vessel. Further processing may include, for example, reverse transcription reactions, amplification reactions, cleavage reactions, dissociation reactions, and / or nucleic acid elongation reactions. Further processing reactions may be carried out in microwells, i.e., without first pooling the labeled target nucleic acid molecules derived from multiple cells.

[0156] Reverse transcription This disclosure provides a method for producing a target-barcode conjugate using reverse transcription (for example, in block 224 of Figure 2). The target-barcode conjugate may contain a barcode and a complementary sequence of all or part of the target nucleic acid (i.e., a barcoded cDNA molecule, such as a probabilistic barcoded cDNA molecule). Reverse transcription of the accompanying RNA molecule can be achieved by adding a reverse transcription primer along with reverse transcriptase. The reverse transcription primer may be an oligo(dT) primer, a random hexanucleotide primer, or a target-specific oligonucleotide primer. Oligo(dT) primers may be 12–18 nucleotides long, or approximately 12–18 nucleotides long, and can bind to the endogenous poly(A) tail at the 3' end of mammalian mRNA. Random hexanucleotide primers can bind to mRNA at various complementary sites. Target-specific oligonucleotide primers typically produce selective priming from the mRNA of interest.

[0157] In some embodiments, reverse transcription of a labeled RNA molecule can be achieved by adding a reverse transcription primer. In some embodiments, the reverse transcription primer is an oligo(dT) primer, a random hexanucleotide primer, or a target-specific oligonucleotide primer. Generally, oligo(dT) primers are 12-18 nucleotides long and bind to the endogenous poly(A) tail at the 3' end of mammalian mRNA. Random hexanucleotide primers can bind to mRNA at various complementary sites. Target-specific oligonucleotide primers typically result in selective priming from the target mRNA. Reverse transcription can occur repeatedly to produce multiple labeled cDNA molecules. The methods disclosed herein may include performing at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 reverse transcription reactions. The methods may include performing at least about 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 reverse transcription reactions.

[0158] amplification One or more nucleic acid amplification reactions (e.g., in block 228 of Figure 2) can be performed to create multiple copies of the labeled target nucleic acid molecule. Amplification can be performed in a multiplex manner in which multiple target nucleic acid sequences are amplified simultaneously. The amplification reaction can be used to attach a sequencing adapter to the nucleic acid molecule. The amplification reaction may include amplifying at least a portion of the sample label, if present. The amplification reaction may include amplifying at least a portion of the cell label and / or barcode sequence (e.g., molecular label). The amplification reaction may include amplifying at least a portion of the sample tag, cell label, spatial label, barcode sequence (e.g., molecular label), target nucleic acid, or a combination thereof. The amplification reaction may include amplifying 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 100%, or a range or numerical value between any two of these values. The method may further include carrying out one or more cDNA synthesis reactions to produce one or more cDNA copies of a target-barcode molecule containing a sample label, cell label, spatial label, and / or barcode sequence (e.g., molecular label).

[0159] In some embodiments, amplification can be carried out using polymerase chain reaction (PCR). As used herein, PCR can refer to a reaction for in vitro amplification of a specific DNA sequence by simultaneous primer extension of the complementary strand of DNA. As used herein, PCR can encompass derivative forms of this reaction, including, but is not limited to, RT-PCR, real-time PCR, nested PCR, quantitative PCR, multiplex PCR, digital PCR, and assembly PCR.

[0160] Amplification of labeled nucleic acids may include non-PCR-based methods. Examples of non-PCR-based methods, but not limited to, include multiple substitution amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand substitution amplification (SDA), real-time SDA, rolling circle amplification, or circle-to-circle amplification. Other non-PCR-based amplification methods include DNA-dependent RNA polymerase-driven RNA transcription amplification or multiple cycles of RNA-directed DNA synthesis and transcription for amplifying DNA or RNA targets, ligase chain reaction (LCR), and Qβ replicase (Qβ) methods, the use of palindromic probes, strand substitution amplification, oligonucleotide-driven amplification using restriction endonucleases, amplification methods in which primers hybridize to nucleic acid sequences and the resulting double helix is ​​cleaved before extension and amplification, strand substitution amplification using nucleic acid polymerases lacking 5' exonuclease activity, rolling circle amplification, and ramification extension amplification (RAM). In some embodiments, amplification does not produce a cyclic transcript.

[0161] In some embodiments, the methods disclosed herein further include carrying out a polymerase chain reaction on a labeled nucleic acid (e.g., labeled RNA, labeled DNA, labeled cDNA) to produce a labeled amplicon (e.g., a probabilistic labeled amplicon). The labeled amplicon may be a double-stranded molecule. The double-stranded molecule may include a double-stranded RNA molecule, a double-stranded DNA molecule, or an RNA molecule hybridized to a DNA molecule. One or both strands of the double-stranded molecule may include a sample label, spatial label, cell label, and / or a barcode sequence (e.g., a molecular label). The labeled amplicon may be a single-stranded molecule. The single-stranded molecule may include DNA, RNA, or a combination thereof. The nucleic acids of this disclosure may include synthetic or modified nucleic acids.

[0162] Amplification may involve the use of one or more non-natural nucleotides. Non-natural nucleotides may include photounstable or photocatalytic nucleotides. Examples of non-natural nucleotides, but not limited to, include peptide nucleic acids (PNA), morpholino, and locked nucleic acids (LNA), as well as glycol nucleic acids (GNA) and threose nucleic acids (TNA). Non-natural nucleotides may be added to one or more cycles of the amplification reaction. The addition of non-natural nucleotides can be used to identify the product at a particular cycle or point in time of the amplification reaction.

[0163] One or more amplification reactions may involve the use of one or more primers. One or more primers may contain, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides, or more. One or more primers may contain at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 nucleotides, or more. One or more primers may contain 12 to less than 15 nucleotides. One or more primers can anneal to at least a portion of a plurality of labeled targets (e.g., probabilistic labeled targets). One or more primers can anneal to the 3' or 5' ends of a plurality of labeled targets. One or more primers can anneal to the internal regions of a plurality of labeled targets. The internal region may consist of at least approximately 50, 100, 150, 200, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 650, 700, 750, 800, 850, 900, or 1000 nucleotides from the 3' end of multiple labeled targets. One or more primers may include primers from a specific panel. One or more primers may include at least one custom primer. One or more primers may include at least one control primer. One or more primers may include at least one gene-specific primer.

[0164] One or more primers may include a universal primer. The universal primer can anneal to a universal primer binding site. One or more custom primers can anneal to a first sample label, a second sample label, a spatial label, a cell label, a barcode sequence (e.g., a molecular label), a target, or any combination thereof. One or more primers may include a universal primer and custom primers. Custom primers can be designed to amplify one or more targets. The targets may include a subset of all nucleic acids in one or more samples. The targets may include a subset of all labeled targets in one or more samples. One or more primers may include at least 96 or more custom primers. One or more primers may include at least 960 or more custom primers. One or more primers may include at least 9600 or more custom primers. One or more custom primers can anneal to two or more different labeled nucleic acids. Two or more different labeled nucleic acids may correspond to one or more genes.

[0165] Any amplification scheme can be used in the method of this disclosure. For example, in one scheme, the first round of PCR can amplify molecules attached to beads using gene-specific primers and primers for universal Illumina sequencing primer 1 sequence. In the second round of PCR, the first PCR product can be amplified using nested gene-specific primers flanked for Illumina sequencing primer 2 sequence and primers for universal Illumina sequencing primer 1 sequence. In the third round of PCR, P5 and P7 and the sample index are added, and the PCR product is converted into an Illumina sequencing library. Sequencing using 150 bp × 2 sequencing can reveal cell labels and barcode sequences (e.g., molecular labels) in read 1, genes in read 2, and the sample index in index 1 read.

[0166] In some embodiments, nucleic acids can be removed from the substrate using chemical cleavage. For example, chemical groups or modified bases present in the nucleic acid can be used to facilitate the removal of nucleic acids from the solid support. For example, enzymes can be used to remove nucleic acids from the substrate. For example, nucleic acids can be removed from the substrate by restriction endonuclease digestion. For example, nucleic acids can be removed from the substrate by treating dUTP or nucleic acids containing ddUTP with uracil-d-glycosylase (UDG). For example, nucleic acids can be removed from the substrate using enzymes that perform nucleotide excision, such as base excision repair enzymes, such as depurine / depyrimidine (AP) endonucleases. In some embodiments, nucleic acids can be removed from the substrate using photocleavable groups and light. In some embodiments, cleavable linkers can be used to remove nucleic acids from the substrate. For example, the cleavable linker may contain at least one of biotin / avidin, biotin / streptavidin, biotin / neutraavidin, Ig-protein A, a photo-unstable linker, an acid or base-unstable linker group, or an aptamer.

[0167] If the probe is gene-specific, the molecule can hybridize to the probe and be reverse transcribed and / or amplified. In some embodiments, the nucleic acid may be amplified after it has been synthesized (e.g., after reverse transcription). Amplification can be carried out in a multiplex manner in which multiple target nucleic acid sequences are amplified simultaneously. Amplification allows a sequencing adapter to be attached to the nucleic acid.

[0168] In some embodiments, amplification can be carried out on a substrate, for example, by bridge amplification. A homopolymer tail may be added to the cDNA to generate ends suitable for bridge amplification using an oligo(dT) probe on the substrate. In bridge amplification, the primer complementary to the 3' end of the template nucleic acid may be each pair of first primers covalently attached to a solid particle. When a sample containing the template nucleic acid is brought into contact with the particle and a single thermal cycle is performed, the template molecule can anneal to the first primer, and by adding nucleotides, the first primer can be extended in the forward direction to form a double-stranded molecule consisting of the template molecule and a newly formed DNA strand complementary to the template. In the heating step of the next cycle, the double-stranded molecule can be denatured, releasing the template molecule from the particle and leaving the complementary DNA strand attached to the particle via the first primer. In the annealing step of the subsequent annealing and extension steps, the complementary strand can hybridize to a second primer complementary to the segment of the complementary strand at the position removed from the first primer. This hybridization can cause the formation of a crosslink between the first and second primers, where the complementary strand is held to the first primer by covalent bonding and to the second primer by hybridization. In the extension step, the second primer can be extended in the reverse direction by adding nucleotides to the same reaction mixture, thereby converting the crosslink into a double-stranded crosslink. The next cycle is then initiated, and the double-stranded crosslink is denatured to obtain two single-stranded nucleic acid molecules, each with one end attached to the particle surface via the first and second primers, respectively, and the other end unattached. In the annealing and extension steps of this second cycle, each strand can hybridize to further complementary primers on the same particle that were not previously used, forming a new single-stranded crosslink. At this point, the two previously unused primers that are hybridizing are extended, and the two new crosslinks are converted into double-stranded crosslinks.

[0169] The amplification reaction can amplify at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 100% of multiple nucleic acids. Amplification of labeled nucleic acids may include PCR-based or non-PCR-based methods. Amplification of labeled nucleic acids may include exponential amplification of labeled nucleic acids. Amplification of labeled nucleic acids may include linear amplification of labeled nucleic acids. Amplification can be carried out by polymerase chain reaction (PCR). PCR can refer to a reaction for in vitro amplification of a specific DNA sequence by simultaneous primer extension of complementary strands of DNA. PCR can encompass, but is not limited to, derivative forms of this reaction, including RT-PCR, real-time PCR, nested PCR, quantitative PCR, multiplex PCR, digital PCR, suppression PCR, semi-suppressive PCR, and assembly PCR.

[0170] In some embodiments, amplification of labeled nucleic acids includes non-PCR-based methods. Examples of non-PCR-based methods, but not limited to, include multiple substitution amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand substitution amplification (SDA), real-time SDA, rolling circle amplification, or circle-to-circle amplification. Other non-PCR-based amplification methods include DNA-dependent RNA polymerase-driven RNA transcription amplification or multiple cycles of RNA-directed DNA synthesis and transcription for amplifying DNA or RNA targets, ligase chain reaction (LCR), Qβ replicase (Qβ) method, use of palindromic probes, strand substitution amplification, oligonucleotide-driven amplification using restriction endonucleases, amplification methods in which primers hybridize to nucleic acid sequences and the resulting double helix is ​​cleaved before extension and amplification, strand substitution amplification using nucleic acid polymerases lacking 5' exonuclease activity, rolling circle amplification, and / or branched extension amplification (RAM).

[0171] In some embodiments, the methods disclosed herein further include carrying out a nested polymerase chain reaction on an amplified amplicon (e.g., a target). The amplicon may be a double-stranded molecule. The double-stranded molecule may include a double-stranded RNA molecule, a double-stranded DNA molecule, or an RNA molecule hybridized to a DNA molecule. One or both strands of the double-stranded molecule may contain a sample tag or molecular identifier label. Alternatively, the amplicon may be a single-stranded molecule. The single-stranded molecule may include DNA, RNA, or a combination thereof. The nucleic acids of this disclosure may include synthetic or modified nucleic acids.

[0172] In some embodiments, the method includes repeatedly amplifying a labeled nucleic acid to produce multiple amplicons. The methods disclosed herein may include performing at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amplification reactions. Alternatively, the method may include performing at least about 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 amplification reactions. Amplification may further involve adding one or more control nucleic acids to one or more samples containing multiple nucleic acids. Amplification may further involve adding one or more control nucleic acids to multiple nucleic acids. The control nucleic acids may include a control label.

[0173] Amplification may involve the use of one or more non-natural nucleotides. Non-natural nucleotides may include photounstable and / or photocatalytic nucleotides. Examples of non-natural nucleotides, but not limited to, include peptide nucleic acids (PNA), morpholino, and locked nucleic acids (LNA), as well as glycol nucleic acids (GNA) and threose nucleic acids (TNA). Non-natural nucleotides may be added to one or more cycles of the amplification reaction. The addition of non-natural nucleotides can be used to identify the product at a particular cycle or point in time of the amplification reaction.

[0174] One or more amplification reactions may involve the use of one or more primers. One or more primers may contain one or more oligonucleotides. One or more oligonucleotides may contain at least about 7 to 9 nucleotides. One or more oligonucleotides may contain 12 to less than 15 nucleotides. One or more primers can anneal to at least a portion of multiple labeled nucleic acids. One or more primers can anneal to the 3' and / or 5' ends of multiple labeled nucleic acids. One or more primers can anneal to the internal regions of multiple labeled nucleic acids. The internal region may consist of at least about 50, 100, 150, 200, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 650, 700, 750, 800, 850, 900, or 1000 nucleotides from the 3' end of multiple labeled nucleic acids. One or more primers may include primers from a specific panel. One or more primers may include at least one or more custom primers. One or more primers may include at least one or more control primers. One or more primers may include at least one or more housekeeping gene primers. One or more primers may include a universal primer. The universal primer can anneal to a universal primer binding site. One or more custom primers can anneal to a first sample tag, a second sample tag, a molecular identifier label, a nucleic acid, or their products. One or more primers may include a universal primer and custom primers. Custom primers can be designed to amplify one or more target nucleic acids.The target nucleic acid may include a subset of all nucleic acids in one or more samples. In some embodiments, the primer is a probe attached to the array of this disclosure.

[0175] In some embodiments, barcoding of multiple targets in a sample (e.g., probabilistic barcoding) further includes generating an indexed library of barcoded targets (e.g., probabilistic barcoded targets) or barcoded fragments of targets. The barcode sequences of different barcodes (e.g., molecular labels of different probabilistic barcodes) may be different from one another. Generating an indexed library of barcoded targets includes generating multiple indexed polynucleotides from multiple targets in a sample. For example, in the case of an indexed library of barcoded targets including a first indexed target and a second indexed target, the label region of the first indexed polynucleotide may differ from the label region of the second indexed polynucleotide by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50 nucleotides, approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50 nucleotides, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50 nucleotides, or up to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50 nucleotides, or a number or range between any two of these values. In some embodiments, the generation of a barcoded target indexing library includes contacting multiple targets, such as mRNA molecules, with multiple oligonucleotides containing a poly(T) region and a labeling region; and performing first-strand synthesis using reverse transcriptase to produce single-stranded labeled cDNA molecules, each containing a cDNA region and a labeling region, wherein the multiple targets include at least two mRNA molecules with different sequences, and the multiple oligonucleotides include at least two oligonucleotides with different sequences. The generation of a barcoded target indexing library may further include amplifying the single-stranded labeled cDNA molecules to produce double-stranded labeled cDNA molecules; and performing nested PCR on the double-stranded labeled cDNA molecules to produce labeled amplicons. In some embodiments, the method may include generating adapter-labeled amplicons.

[0176] Barcoding (e.g., probabilistic barcoding) may involve labeling individual nucleic acid (e.g., DNA or RNA) molecules using nucleic acid barcodes or tags. In some embodiments, barcoding includes attaching a DNA barcode or tag to a cDNA molecule when the cDNA molecule is generated from mRNA. Nested PCR can be performed to minimize PCR amplification bias. For example, an adapter can be attached for sequencing using next-generation sequencing (NGS). Using the sequencing results, cell labels, molecular labels, and sequences of one or more copies of a target nucleotide fragment can be determined, for example, in block 232 of Figure 2.

[0177] Figure 3 is a schematic diagram illustrating a non-limiting, exemplary process for generating an indexed library of barcoded targets (e.g., probabilistic barcoded targets), such as barcoded mRNA or fragments thereof. As shown in Step 1, the reverse transcription process can encode each mRNA molecule having a unique molecular label, a cellular label, and a universal PCR site. In particular, a set of barcodes (e.g., probabilistic barcodes) 310 can be reverse transcribed by hybridizing (e.g., probabilistic hybridization) the poly(A) tail region 308 of the RNA molecule 302 to produce a labeled cDNA molecule 304 containing a cDNA region 306. Each of the barcodes 310 may include a target-binding region, e.g., a poly(dT) region 312, a labeling region 314 (e.g., a barcode sequence or molecule), and a universal PCR region 316.

[0178] In some embodiments, the cell label may contain 3 to 20 nucleotides. In some embodiments, the molecular label may contain 3 to 20 nucleotides. In some embodiments, each of a plurality of probabilistic barcodes further comprises one or more universal labels and cell labels, where the universal label is the same for a plurality of probabilistic barcodes on a solid support, and the cell label is the same for a plurality of probabilistic barcodes on a solid support. In some embodiments, the universal label may contain 3 to 20 nucleotides. In some embodiments, the cell label contains 3 to 20 nucleotides.

[0179] In some embodiments, the labeling region 314 may include a barcode sequence or molecular label 318 and a cell label 320. In some embodiments, the labeling region 314 may include one or more universal labels, dimensional labels, and cell labels. The barcode sequence or molecular label 318 may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides long, approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides long, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides long, or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides long, or a number or range of nucleotide lengths between any of these values. The cell label 320 may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides long, approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides long, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides long, or a number or range of nucleotides between any of these values.The universal label may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides long, approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides long, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides long, or a number or range of nucleotide lengths between any of these values. A universal label may be the same across multiple probabilistic barcodes on a solid support, and a cell label is the same across multiple probabilistic barcodes on a solid support. The dimension label may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides long, approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides long, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100 nucleotides long, or a number or range of nucleotide lengths between any of these values.

[0180] In some embodiments, the labeling region 314 may contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 different labels, such as barcode sequences or molecular labels 318 and cell labels 320, and may contain approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 different labels, It may contain at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 different markers, or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 different markers, or a number or range of different markers between any of these values. Each label may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides long, approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides long, at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100 nucleotides long, or any number or range of these values. One set of barcodes or probabilistic barcodes 310 consists of 10, 20, 40, 50, 70, 80, 90, and 10. 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 1010 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 20 It may contain 10 barcodes or probabilistic barcodes 310, and approximately 10, 20, 40, 50, 70, 80, 90, 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 20 It may contain 10 barcodes or probabilistic barcodes 310, and at least 10, 20, 40, 50, 70, 80, 90, 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 20 It may contain 1 barcode or probabilistic barcode 310, or up to 10, 20, 40, 50, 70, 80, 90, 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 20The labeled cDNA molecule 304 may contain 1 barcode or probabilistic barcode 310, or a number or range of barcodes or probabilistic barcodes 310 between any of these values. Furthermore, each set of barcodes or probabilistic barcodes 310 may, for example, contain a unique labeling region 314. To remove excess barcodes or probabilistic barcodes 310, the labeled cDNA molecule 304 may be purified. Purification may include Ampure bead purification.

[0181] As shown in Step 2, the product of the reverse transcription process in Step 1 may be pooled in one tube and PCR amplified using the first PCR primer pool and the first universal PCR primer. Pooling is possible because of the uniquely labeled region 314. In particular, the labeled cDNA molecule 304 may be amplified to produce a nested PCR-labeled amplicon 322. The amplification may include multiplex PCR amplification. The amplification may include multiplex PCR amplification in a single reaction volume using 96 multiplex primers. In some embodiments, the multiplex PCR amplification is performed in a single reaction volume with 10, 20, 40, 50, 70, 80, 90, 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 20 Multiplex primers may be used in quantities of approximately 10, 20, 40, 50, 70, 80, 90, and 10. 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10, 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 20 It may be possible to use 10, 20, 40, 50, 70, 80, 90, 10 2 , 10 3 , 10 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 20 It may be possible to use 10, 20, 40, 50, 70, 80, 90, 10 2 , 10 3 , 10<了 4 , 10 5 , 10 6 , 10 7 , 10 8 , 10 9 , 10 10 , 10 11 , 10 12 , 10 13 , 10 14 , 10 15 , 10 20 It may be possible to use 10, 20, 40, 50, 70, 80, 90, 10

[0182] As shown in step 3 of Figure 3, the product of the PCR amplification in step 2 may be amplified using a nested PCR primer pool and a second universal PCR primer. Nested PCR can minimize PCR amplification bias. In particular, the nested PCR-labeled amplicon 322 may be further amplified by nested PCR. Nested PCR may include multiplex PCR in a single reaction volume using a nested PCR primer pool 330 of nested PCR primers 332a-c and a second universal PCR primer 328'. Nested PCR primer pool 328 may contain 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000 different nested PCR primers 330, and may contain at least 1, 2, 3, 4, 5, 6, 7, 8, 9 It may contain 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 different nested PCR primers 330, or up to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 different nested PCR primers 330, or a number or range of nested PCR primers 330 between any of these values. Nested PCR primer 332 contains adapter 334 and can hybridize to a region within the cDNA portion 306'' of labeled amplicon 322. Universal primer 328' contains adapter 336 and can hybridize to the universal PCR region 316 of labeled amplicon 322.Thus, in step 3, adapter-labeled amplicon 338 is produced. In some embodiments, nested PCR primers 332 and the second universal PCR primer 328' may not contain adapters 334 and 336. Instead, adapters 334 and 336 may be ligated to the product of the nested PCR to produce adapter-labeled amplicon 338.

[0183] As shown in step 4, the PCR product of step 3 may be PCR amplified for sequencing using library amplification primers. In particular, one or more additional assays may be performed on adapter-labeled amplicon 338 using adapters 334 and 336. Adapters 334 and 336 can hybridize to primers 340 and 342. One or more of primers 340 and 342 may be PCR amplification primers. One or more of primers 340 and 342 may be sequencing primers. One or more of adapters 334 and 336 may be used for further amplification of adapter-labeled amplicon 338. One or more of adapters 334 and 336 may be used for sequencing of adapter-labeled amplicon 338. Primer 342 may contain plate index 344 so that the amplicon generated using the same set of barcodes or probabilistic barcodes 310 can be sequenced in one sequencing reaction using next-generation sequencing (NGS).

[0184] A composition comprising a cell component-binding reagent associated with an oligonucleotide Some embodiments disclosed herein provide a plurality of compositions each comprising a cell component-binding reagent (such as a protein-binding reagent) conjugated with an oligonucleotide, wherein the oligonucleotide comprises a unique identifier of the cell component-binding reagent to which it is conjugated. Cell component-binding reagents (such as barcoded antibodies) and their uses (such as assigning sample indexes to cells) are described in U.S. Patent Application Publication 2018 / 0088112 and U.S. Patent Application 15 / 937,713. The contents of each of these are incorporated in whole by reference.

[0185] In some embodiments, cell component binding reagents are capable of specifically binding to cell component targets. For example, the binding targets of cell component binding reagents may include carbohydrates, lipids, proteins, extracellular proteins, cell surface proteins, cell markers, B cell receptors, T cell receptors, major histocompatibility complexes, tumor antigens, receptors, integrins, intracellular proteins, or any combination thereof. In some embodiments, cell component binding reagents (e.g., protein binding reagents) are capable of specifically binding to antigen targets or protein targets. In some embodiments, each oligonucleotide may include a barcode, such as a probabilistic barcode. The barcode may include a barcode sequence (e.g., molecular label), a cell label, a sample label, or any combination thereof. In some embodiments, each oligonucleotide may include a linker. In some embodiments, each oligonucleotide may include a binding site for an oligonucleotide probe, such as a poly(A) tail. For example, the poly(A) tail may or may not be anchored to a solid support. The poly(A)tail may be about 10 to 50 nucleotides long. In some embodiments, the poly(A)tail may be 18 nucleotides long. The oligonucleotide may consist of deoxyribonucleotides, ribonucleotides, or both.

[0186] The unique identifier may be, for example, a nucleotide sequence having any preferred length, for example, about 4 nucleotides to about 200 nucleotides. In some embodiments, the unique identifier is a nucleotide sequence with a length of 25 nucleotides to about 45 nucleotides. In some embodiments, the unique identifier may have lengths of 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 70, 80, 90, 100, 200 nucleotides, lengths of about 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 70, 80, 90, 100, 200 nucleotides, lengths of 4, 5 nucleotides, 5 nucleos The length may be less than 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 70, 80, 90, 100, or 200 nucleotides, or greater than 4, 5, 6, 7, 8, 9, 100, or 200 nucleotides, or within a range of any two of the above values.

[0187] In some embodiments, the unique identifier is selected from a diverse set of unique identifiers. The diverse set of unique identifiers may include 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, or 5000 different unique identifiers, or may include about 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, or 5000 different unique identifiers, or may include a number or range of different unique identifiers between any two of these values. A diverse set of unique identifiers may contain at least 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, or 5000 different unique identifiers, or up to 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, or 5000 different unique identifiers. In some embodiments, a set of unique identifiers is designed to have minimal sequence homology to the DNA or RNA sequence of the sample to be analyzed. In some embodiments, a set of unique identifier sequences differ from each other or from their complements by only 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides, or by approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides, or by a number or range between any two of these values.In some embodiments, the sequences of a set of unique identifiers differ from each other or from their complements by at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides, or by a maximum of 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides. In some embodiments, the sequences of a set of unique identifiers differ from each other or from their complements by at least 3%, at least 5%, at least 8%, at least 10%, at least 15%, at least 20%, or more.

[0188] In some embodiments, the unique identifier may include binding sites to a primer, such as a universal primer. In some embodiments, the unique identifier may include at least two binding sites to a primer, such as a universal primer. In some embodiments, the unique identifier may include at least three binding sites to a primer, such as a universal primer. The primer can be used, for example, to amplify the unique identifier by PCR amplification. In some embodiments, the primer can be used in a nested PCR reaction.

[0189] This disclosure envisions any suitable cell component binding reagent, such as protein-binding reagents, antibodies or their fragments, aptamers, small molecules, ligands, peptides, oligonucleotides, or any combination thereof. In some embodiments, the cell component binding reagent may be a polyclonal antibody, a monoclonal antibody, a recombinant antibody, a single-chain antibody (sc-Ab), or a fragment thereof such as Fab, Fv. In some embodiments, the multiple cell component binding reagents may contain 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, and 5000 different cell component reagents, or about 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, and 5000 different cell component reagents, or a number or range of different cell component reagents between any two of these values. In some embodiments, the multiple cell component binding reagents may contain at least 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, or 5000 different cell component reagents, or up to 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, or 5000 different cell component reagents.

[0190] Oligonucleotides may be conjugated to cell component-binding reagents by various mechanisms. In some embodiments, oligonucleotides may be covalently conjugated to cell component-binding reagents. In some embodiments, oligonucleotides may be noncovalently conjugated to cell component-binding reagents. In some embodiments, oligonucleotides are conjugated to cell component-binding reagents via a linker. The linker may be cleavable or detachable from, for example, the cell component-binding reagent and / or oligonucleotide. In some embodiments, the linker may contain a chemical group that reversibly attaches the oligonucleotide to the cell component-binding reagent. The chemical group may be conjugated to the linker via, for example, an amine group. In some embodiments, the linker may contain a chemical group that forms a stable bond with another chemical group conjugated to the cell component-binding reagent. For example, the chemical group may be a UV-cleavable group, a disulfide bond, streptavidin, biotin, or an amine. In some embodiments, the chemical group may be conjugated to a cell component-binding reagent via a primary amine of an amino acid such as lysine or via its N-terminus. Oligonucleotides can be conjugated to cell component-binding reagents using commercially available conjugation kits, such as the Protein-Oligo Conjugation Kit (Solulink, Inc., San Diego, California) or the Thunder-Link® Oligo Conjugation System (Innova Biosciences, Cambridge, UK).

[0191] Oligonucleotides can be conjugated to any suitable site on a cell component-binding reagent (e.g., a protein-binding reagent), as long as they do not interfere with the specific binding between the cell component-binding reagent and its cell component target. In some embodiments, the cell component-binding reagent is a protein such as an antibody. In some embodiments, the cell component-binding reagent is not an antibody. In some embodiments, the oligonucleotide can be conjugated to any site on the antibody other than the antigen-binding site, for example, the Fc region, C H 1 domain, CH 2 domains, C H 3 domains, C L They can be conjugated to domains, etc. Methods for conjugating oligonucleotides to cell component-binding reagents (e.g., antibodies) have been previously disclosed, for example, in U.S. Patent No. 6,531,283. This is expressly incorporated herein by reference in its entirety. The stoichiometry of oligonucleotides to cell component-binding reagents may vary. To increase the sensitivity for detecting cell component-binding reagent-specific oligonucleotides in sequencing, it may be advantageous to increase the ratio of oligonucleotides to cell component-binding reagents in the conjugation. In some embodiments, each cell component-binding reagent may be conjugated with a single oligonucleotide molecule. In some embodiments, each cell component-binding reagent may be conjugated to more than one oligonucleotide molecule, for example, at least 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000 oligonucleotide molecules, or up to 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000 oligonucleotide molecules, or a number or range of oligonucleotide molecules between any two such values, each of which oligonucleotide molecules contains the same or different unique identifier.

[0192] In some embodiments, the multiple cell component binding reagent can specifically bind to multiple cell component targets in a sample such as a single cell, multiple cells, tissue sample, tumor sample, or blood sample. In some embodiments, the multiple cell component targets include cell surface proteins, cell markers, B cell receptors, T cell receptors, antibodies, major histocompatibility complexes, tumor antigens, receptors, or any combination thereof. In some embodiments, the multiple cell component targets may include intracellular cellular components. In some embodiments, the multiple cell component targets may include intracellular cellular components. In some embodiments, the number of cellular components may be 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, 99% of the total cellular components (e.g., proteins) in the cell or organism, or a number or range between any two of these values. In some embodiments, the multiple cellular components may represent at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, or 99% of the total cellular components (e.g., proteins) in the cell or organism, or up to 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, or 99%. In some embodiments, the multiple cellular component targets may include 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000, or 10000 different cellular component targets, or approximately 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000, or 10000 different cellular component targets, or may include a number or range of different cellular component targets between any two of these values.In some embodiments, the multiple cellular component targets may include at least 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000, or 10000 different cellular component targets, or up to 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000, or 10000 different cellular component targets.

[0193] Figure 4 shows a schematic diagram of an exemplary cell component-binding reagent (e.g., an antibody) accompanied (e.g., conjugated) with an oligonucleotide containing an antibody-specific identifier sequence. An oligonucleotide conjugated to a cell component-binding reagent, an oligonucleotide for conjugating to a cell component-binding reagent, or an oligonucleotide previously conjugated to a cell component-binding reagent may be referred to herein as an antibody oligonucleotide (abbreviated as "binding reagent oligonucleotide"). An oligonucleotide conjugated to an antibody, an oligonucleotide for conjugating to an antibody, or an oligonucleotide previously conjugated to an antibody may be referred to herein as an antibody oligonucleotide (abbreviated as "AbOligo" or "AbO"). The oligonucleotide may also contain additional components, but is not limited to, one or more linkers, one or more antibody-specific identifiers, optionally one or more barcode sequences (e.g., molecular labels), and a poly(dA) tail. In some embodiments, the oligonucleotide may contain, from 5' to 3', a linker, a specific identifier, a barcode sequence (e.g., molecular labels), and a poly(dA) tail. Antibody oligonucleotides may also be mRNA mimics.

[0194] Figure 5 shows a schematic diagram of an exemplary cell component-binding reagent (e.g., an antibody) accompanied (e.g., conjugated) by an oligonucleotide containing an antibody-specific identifier sequence. The cell component-binding reagent may be capable of specifically binding to at least one cell component target, such as an antigen target or a protein target. The binding reagent oligonucleotide (e.g., a sample indexing oligonucleotide, or an antibody oligonucleotide) may contain a sequence (e.g., a sample indexing sequence) for carrying out the methods of this disclosure. For example, a sample indexing oligonucleotide may contain a sample indexing sequence for identifying the sample origin of one or more cells in a sample. The indexing sequences (e.g., sample indexing sequences) of at least two compositions containing two cell component-binding reagents (e.g., sample indexing compositions) of a plurality of compositions containing cell component-binding reagents may contain different sequences. In some embodiments, the binding reagent oligonucleotide is not homologous to the genome sequence of a species. The binding reagent oligonucleotide may be configured to be detachable or indedetachable from the cell component-binding reagent (or may be detachable or indedetachable).

[0195] The oligonucleotide conjugated to the cell component binding reagent may include, for example, a barcode sequence (e.g., a molecular label sequence), a poly(dA) tail, or a combination thereof. The oligonucleotide conjugated to the cell component binding reagent may also be an mRNA mimeograph. In some embodiments, the sample indexing oligonucleotide includes a sequence complementary to the capture sequence of at least one of the multiple barcodes. The target binding region of the barcode may include a capture sequence. The target binding region may include, for example, a poly(dT) region. In some embodiments, the sequence of the sample indexing oligonucleotide complementary to the capture sequence of the barcode may include a poly(dA) tail. The sample indexing oligonucleotide may include a molecular label.

[0196] In some embodiments, the binding reagent oligonucleotide (e.g., sample oligonucleotide) is 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 128, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 40 0, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800, 810, 820, 830, 840, 850, 860, 870, 880, 890, 900, 910, 920, 930, 940, 950, 960, 970, 980, 990, 10 A nucleotide sequence of 00 nucleotides in length, or approximately 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 128, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 4 Nucleotide sequences of 50, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800, 810, 820, 830, 840, 850, 860, 870, 880, 890, 900, 910, 920, 930, 940, 950, 960, 970, 980, 990, and 1000 nucleotide lengths.Or a nucleotide sequence having a nucleotide length of a number or range between any two of these values. In some embodiments, the binding reagent oligonucleotide contains at least 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 128, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800, 810, 820, 830, 840, 850, 860, 870, 880, 890, 900, 910, 920, 930, 940, 950, 960, 970, 980, 990, or 1000 Nuk Rheotide length, or up to 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 80, 90, 100, 110, 120, 128, 130, 140, 150, 160, 170, 180, 190, 200, 210, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 47 Contains nucleotide sequences of 0, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 610, 620, 630, 640, 650, 660, 670, 680, 690, 700, 710, 720, 730, 740, 750, 760, 770, 780, 790, 800, 810, 820, 830, 840, 850, 860, 870, 880, 890, 900, 910, 920, 930, 940, 950, 960, 970, 980, 990, or nucleotide sequences of 1000 nucleotides in length.

[0197] In some embodiments, the cell component binding reagent includes an antibody, a tetramer, an aptamer, a protein scaffold, or a combination thereof. The binding reagent oligonucleotide may be conjugated to the cell component binding reagent, for example, via a linker. The binding reagent oligonucleotide may contain a linker. The linker may contain a chemical group. The chemical group may be reversibly or irreversibly attached to the molecule of the cell component binding reagent. The chemical group can be selected from the group consisting of UV-cleavable groups, disulfide bonds, streptavidin, biotin, amines, and any combination thereof.

[0198] In some embodiments, the cell component-binding reagent can bind to ADAM10, CD156c, ANO6, ATP1B2, ATP1B3, BSG, CD147, CD109, CD230, CD29, CD298, ATP1B3, CD44, CD45, CD47, CD51, CD59, CD63, CD97, CD98, SLC3A2, CLDND1, HLA-ABC, ICAM1, ITFG3, MPZL1, NA K ATPase alpha 1, ATP1A1, NPTN, PMCA ATPase, ATP2B1, SLC1A5, SLC29A1, SLC2A1, SLC44A2, or any combination thereof.

[0199] In some embodiments, the protein target is or includes an extracellular protein, an intracellular protein, or any combination thereof. In some embodiments, the antigen or protein target is or includes a cell surface protein, a cell marker, a B cell receptor, a T cell receptor, a major histocompatibility complex, a tumor antigen, a receptor, an integrin, or any combination thereof. The antigen or protein target may also be or include a lipid, a carbohydrate, or any combination thereof. The protein target can be selected from a group comprising a number of protein targets. The number of antigen targets or protein targets is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, or approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, or any number or range between any two of these values. The number of protein targets is at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10000. It may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10000.

[0200] A cell component binding reagent (e.g., a protein binding reagent) may be accompanied by two or more binding reagent oligonucleotides having the same sequence (e.g., sample indexing oligonucleotides). A cell component binding reagent may be accompanied by two or more binding reagent oligonucleotides having different sequences. The number of binding reagent oligonucleotides accompanied by a cell component binding reagent may differ depending on the implementation. In some embodiments, the number of binding reagent oligonucleotides may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 1000, or a number or range between any two of these values, regardless of whether they have the same sequence or different sequences. In some embodiments, the number of binding reagent oligonucleotides may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000, or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000.

[0201] Multiple compositions containing cell component binding reagents (e.g., multiple sample indexing compositions) may include one or more additional cell component binding reagents that are not conjugated to the binding reagent oligonucleotide (e.g., sample indexing oligonucleotides), also referred herein as cell component binding reagents that do not contain the binding reagent oligonucleotide (e.g., cell component binding reagents that do not contain the sample indexing oligonucleotide). The number of additional cell component binding reagents in the multiple compositions may vary depending on the implementation. In some embodiments, the number of additional cell component binding reagents may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or a number or range between any two of these values. In some embodiments, the number of additional cell component binding reagents may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100, or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100. In some embodiments, the cell component binding reagent and any of the additional cell component binding reagents may be identical.

[0202] In some embodiments, a mixture is provided comprising a cell component binding reagent conjugated with one or more binding reagent oligonucleotides (e.g., sample index-granting oligonucleotides) and a cell component binding reagent not conjugated with a binding reagent oligonucleotide. This mixture can be used, for example, in contact with a sample and / or cells in some embodiments of the methods disclosed herein. The ratio of (1) the number of cell component binding reagents conjugated with binding reagent oligonucleotides and (2) the number of other cell component binding reagents (e.g., the same cell component binding reagent) not conjugated with other binding reagent oligonucleotides in the mixture may vary depending on the implementation. In some embodiments, this ratio is 1:1, 1:1.1, 1:1.2, 1:1.3, 1:1.4, 1:1.5, 1:1.6, 1:1.7, 1:1.8, 1:1.9, 1:2, 1:2.5, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:11, 1:12, 1:13, 1:14, 1:15, 1:16, 1:17, 1:18, 1:19, 1:20, 1:21, 1:22, 1:23, 1:24, 1:25, 1:26, 1:27, 1:28, 1:29, 1:30, 1:31, 1:32, 1:33, 1:34, 1:35, 1:36, 1:37 , 1:38, 1:39, 1:40, 1:41, 1:42, 1:43, 1:44, 1:45, 1:46, 1:47, 1:48, 1:49, 1:50, 1:51, 1:52, 1:53, 1:54, 1:55, 1:56, 1:57, 1:58, 1:59, 1:60, 1:61, 1:62, 1:63, 1:64, 1:65, 1:66, 1:67, 1:68, 1:69, 1:70, 1:7 1, 1:72, 1:73, 1:74, 1:75, 1:76, 1:77, 1:78, 1:79, 1:80, 1:81, 1:82, 1:83, 1:84, 1:85, 1:86, 1:87, 1:88, 1:89, 1:90, 1:91, 1:92, 1:93, 1:94, 1:95, 1:96, 1:97, 1:98, 1:99, 1:100, 1:200, 1:300, 1:400, 1:5 00, 1:600, 1:700, 1:800, 1:900, 1:1000, 1:2000, 1:3000, 1:4000, 1:5000, 1:6000, 1:7000, 1:8000, 1:9000, 1:10000, or approximately 1:1, 1:1.1, 1:1.2, 1:1.3, 1:1.4, 1:1.5, 1:1.6, 1:1.7, 1:1.8, 1:1.9, 1:2, 1:2.5, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:11, 1:12, 1:13, 1:14, 1:15, 1:16, 1:17, 1:18, 1:19, 1:20, 1:21, 1:22, 1:23, 1:24, 1:25, 1:26, 1:27, 1:28, 1:29, 1:30, 1:31, 1:32, 1:33, 1:34, 1:35 , 1:36, 1:37, 1:38, 1:39, 1:40, 1:41, 1:42, 1:43, 1:44, 1:45, 1:46, 1:47, 1:48, 1:49, 1:50, 1:51, 1:52, 1:53, 1:54, 1:55, 1:56, 1:57, 1:58, 1:59, 1:60, 1:61, 1:62, 1:63, 1:64, 1:65, 1:66, 1:6 7, 1:68, 1:69, 1:70, 1:71, 1:72, 1:73, 1:74, 1:75, 1:76, 1:77, 1:78, 1:79, 1:80, 1:81, 1:82, 1:83, 1:84, 1:85, 1:86, 1:87, 1:88, 1:89, 1:90, 1:91, 1:92, 1:93, 1:94, 1:95, 1:96, 1:97, 1:98, 1: 99, 1:100, 1:200, 1:300, 1:400, 1:500, 1:600, 1:700, 1:800, 1:900, 1:1000, 1:2000, 1:3000, 1:4000, 1:5000, 1:6000, 1:7000, 1:8000, 1:9000, 1:10000, or any number or range between any two of the above values. In some embodiments, this ratio may be at least 1:1, 1:1.1, 1:1.2, 1:1.3, 1:1.4, 1:1.5, 1:1.6, 1:1.7, 1:1.8, 1:1.9, 1:2, 1:2.5, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:11, 1:12, 1:13, 1:14, 1:15, 1:16, 1:17, 1:18, 1:19, 1:20, 1:21, 1:22, 1:23, 1:24, 1:25, 1:26, 1:27, 1:28, 1:29, 1:30, 1:31, 1:32, 1:33, 1:34, 1:35, 1:36, 1:37, 1: 38, 1:39, 1:40, 1:41, 1:42, 1:43, 1:44, 1:45, 1:46, 1:47, 1:48, 1:49, 1:50, 1:51, 1:52, 1:53, 1:54, 1:55, 1:56, 1:57, 1:58, 1:59, 1:60, 1:61, 1:62, 1:63, 1:64, 1:65, 1:66, 1:67, 1:68, 1:69, 1:70, 1:71, 1:72 , 1:73, 1:74, 1:75, 1:76, 1:77, 1:78, 1:79, 1:80, 1:81, 1:82, 1:83, 1:84, 1:85, 1:86, 1:87, 1:88, 1:89, 1:90, 1:91, 1:92, 1:93, 1:94, 1:95, 1:96, 1:97, 1:98, 1:99, 1:100, 1:200, 1:300, 1:400, 1:500, 1:600, The ratios may be 1:700, 1:800, 1:900, 1:1000, 1:2000, 1:3000, 1:4000, 1:5000, 1:6000, 1:7000, 1:8000, 1:9000, or 1:10000, or at most 1:1, 1:1.1, 1:1.2, 1:1.3, 1:1.4, 1:1.5, 1:1.6, 1:1.7, 1:1.8, 1:1.9, 1:2, 1:2.5, 1:3, 1:4, 1:5, 1:6, 1:7, 1:8, 1:9, 1:10, 1:11, 1:12, 1:13, 1:14, 1:15, 1:16, 1:17, 1:18, 1:19, 1:20, 1:21, 1:22, 1:23, 1:24, 1:25, 1:26, 1:27, 1:28, 1:29, 1:30, 1:31, 1:32, 1:33, 1:34 , 1:35, 1:36, 1:37, 1:38, 1:39, 1:40, 1:41, 1:42, 1:43, 1:44, 1:45, 1:46, 1:47, 1:48, 1:49, 1:50, 1:51, 1:52, 1:53, 1:54, 1:55, 1:56, 1:57, 1:58, 1:59, 1:60, 1:61, 1:62, 1:63, 1:64, 1:6 5, 1:66, 1:67, 1:68, 1:69, 1:70, 1:71, 1:72, 1:73, 1:74, 1:75, 1:76, 1:77, 1:78, 1:79, 1:80, 1:81, 1:82, 1:83, 1:84, 1:85, 1:86, 1:87, 1:88, 1:89, 1:90, 1:91, 1:92, 1:93, 1:94, 1:95, 1: 96, 1:97, 1:98, 1:99, 1:100, 1:200, 1:300, 1:400, 1:500, 1:600, 1:700, 1:800, 1:900, 1:1000, 1:2000, 1:3000, 1:4000, 1:5000, 1:6000, 1:7000, 1:8000, 1:9000, or 1:10000 may also be used.

[0203] In some embodiments, this ratio is 1:1, 1.1:1, 1.2:1, 1.3:1, 1.4:1, 1.5:1, 1.6:1, 1.7:1, 1.8:1, 1.9:1, 2:1, 2.5:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 21:1, 22:1, 23:1, 24:1, 25 :1, 26:1, 27:1, 28:1, 29:1, 30:1, 31:1, 32:1, 33:1, 34:1, 35:1, 36:1, 37:1, 38:1, 39:1, 40:1, 41:1, 42:1, 43:1, 44:1, 45:1, 46:1, 47:1, 48:1, 49:1, 50:1, 51:1, 52:1, 53:1, 54:1, 55:1, 56:1, 57:1, 58:1, 59:1, 60:1, 61:1, 62:1, 6 3:1, 64:1, 65:1, 66:1, 67:1, 68:1, 69:1, 70:1, 71:1, 72:1, 73:1, 74:1, 75:1, 76:1, 77:1, 78:1, 79:1, 80:1, 81:1, 82:1, 83:1, 84:1, 85:1, 86:1, 87:1, 88:1, 89:1, 90:1, 91:1, 92:1, 93:1, 94:1, 95:1, 96:1, 97:1, 98:1, 99:1, 100:1 , 200:1, 300:1, 400:1, 500:1, 600:1, 700:1, 800:1, 900:1, 1000:1, 2000:1, 3000:1, 4000:1, 5000:1, 6000:1, 7000:1, 8000:1, 9000:1, 10000:1, or approximately 1:1, 1.1:1, 1.2:1, 1.3:1, 1.4:1, 1.5:1, 1.6:1, 1.7:1, 1.8:1, 1.9:1, 2:1, 2.5:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 21:1, 22:1, 23:1, 24:1, 25:1, 26:1, 27:1, 28:1, 29:1, 30:1, 31:1, 32:1, 33:1, 34:1, 35 :1, 36:1, 37:1, 38:1, 39:1, 40:1, 41:1, 42:1, 43:1, 44:1, 45:1, 46:1, 47:1, 48:1, 49:1, 50:1, 51:1, 52:1, 53:1, 54:1, 55:1, 56:1, 57:1, 58:1, 59:1, 60:1, 61:1, 62:1, 63:1, 64:1, 65:1, 66:1, 67 :1, 68:1, 69:1, 70:1, 71:1, 72:1, 73:1, 74:1, 75:1, 76:1, 77:1, 78:1, 79:1, 80:1, 81:1, 82:1, 83:1, 84:1, 85:1, 86:1, 87:1, 88:1, 89:1, 90:1, 91:1, 92:1, 93:1, 94:1, 95:1, 96:1, 97:1, 98:1, 99 :1, 100:1, 200:1, 300:1, 400:1, 500:1, 600:1, 700:1, 800:1, 900:1, 1000:1, 2000:1, 3000:1, 4000:1, 5000:1, 6000:1, 7000:1, 8000:1, 9000:1, 10000:1, or a number or range between any two of the above values. In some embodiments, this ratio may be at least 1:1, 1.1:1, 1.2:1, 1.3:1, 1.4:1, 1.5:1, 1.6:1, 1.7:1, 1.8:1, 1.9:1, 2:1, 2.5:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 21:1, 22:1, 23:1, 24:1, 25:1, 26:1, 27:1, 28:1, 29:1, 30:1, 31:1, 32:1, 33:1, 34:1, 35:1, 36:1, 37:1, 38:1, 39:1, 40:1, 41:1, 42:1, 43:1, 44:1, 45:1, 46:1, 47:1, 48:1, 49:1, 50:1, 51:1, 52:1, 53:1, 54:1, 55:1, 56:1, 57:1, 58:1, 59:1, 60:1, 61:1, 62:1, 63:1, 64:1, 65:1, 66:1, 67:1, 68:1, 69:1, 70:1, 71:1, 72 :1, 73:1, 74:1, 75:1, 76:1, 77:1, 78:1, 79:1, 80:1, 81:1, 82:1, 83:1, 84:1, 85:1, 86:1, 87:1, 88:1, 89:1, 90:1, 91:1, 92:1, 93:1, 94:1, 95:1, 96:1, 97:1, 98:1, 99:1, 100:1, 200:1, 300:1, 400:1, 500:1, 600: 1, 700:1, 800:1, 900:1, 1000:1, 2000:1, 3000:1, 4000:1, 5000:1, 6000:1, 7000:1, 8000:1, 9000:1, or 10000:1, or at most 1:1, 1.1:1, 1.2:1, 1.3:1, 1.4:1, 1.5:1, 1.6:1, 1.7:1, 1.8:1, 1.9:1, 2:1, 2.5:1, 3:1, 4:1, 5:1, 6:1, 7:1, 8:1, 9:1, 10:1, 11:1, 12:1, 13:1, 14:1, 15:1, 16:1, 17:1, 18:1, 19:1, 20:1, 21:1, 22:1, 23:1, 24:1, 25:1, 26:1, 27:1, 28:1, 29:1, 30:1, 31:1, 32:1, 33:1, 34 :1, 35:1, 36:1, 37:1, 38:1, 39:1, 40:1, 41:1, 42:1, 43:1, 44:1, 45:1, 46:1, 47:1, 48:1, 49:1, 50:1, 51:1, 52:1, 53:1, 54:1, 55:1, 56:1, 57:1, 58:1, 59:1, 60:1, 61:1, 62:1, 63:1, 64:1, 65 :1, 66:1, 67:1, 68:1, 69:1, 70:1, 71:1, 72:1, 73:1, 74:1, 75:1, 76:1, 77:1, 78:1, 79:1, 80:1, 81:1, 82:1, 83:1, 84:1, 85:1, 86:1, 87:1, 88:1, 89:1, 90:1, 91:1, 92:1, 93:1, 94:1, 95:1, 9 6:1, 97:1, 98:1, 99:1, 100:1, 200:1, 300:1, 400:1, 500:1, 600:1, 700:1, 800:1, 900:1, 1000:1, 2000:1, 3000:1, 4000:1, 5000:1, 6000:1, 7000:1, 8000:1, 9000:1, or 10000:1 are also acceptable.

[0204] The cell component binding reagent may or may not be conjugated with a binding reagent oligonucleotide (e.g., a sample index-granting oligonucleotide). In some embodiments, in a mixture containing a cell component binding reagent conjugated with a binding reagent oligonucleotide (e.g., a sample index-granting oligonucleotide) and a cell component binding reagent not conjugated with a binding reagent oligonucleotide, the percentage of the cell component binding reagent conjugated with a binding reagent oligonucleotide is 0.000000001%, 0 0.00000001%, 0.0000001%, 0.000001%, 0.00001%, 0.0001%, 0.001%, 0.01%, 0.1%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31% 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76% , 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or approximately 0.000000001%, 0.00000001%, 0.0000001%, 0.000001%, 0.00001%, 0.0001%, 0.001%, 0.01%, 0.01%.1%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or any number or range between any two of these values. In some embodiments, the percentage of the cell component-binding reagent in the mixture conjugated with a sample index-granting oligonucleotide is at least 0.000000001%, 0.00000001%, 0.0000001%, 0.000001%, 0.00001%, 0.0001%, 0.001%, 0.01%, 0.1%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35% %, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70% , 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, or at most 0.000000001%, 0.00000001%, 0.0000001%, 0.000001%, 0.00001%, 0.0001%, 0.001%, 0.01%, 0.1%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% may be used.

[0205] In some embodiments, a mixture containing a cell component binding reagent conjugated with a binding reagent oligonucleotide (e.g., a sample index-granting oligonucleotide) and a cell component binding reagent not conjugated with a sample index-granting oligonucleotide is used, where the cell component binding reagent not conjugated with a binding reagent oligonucleotide (e.g., a sample index-granting oligonucleotide) The percentages are 0.000000001%, 0.00000001%, 0.0000001%, 0.000001%, 0.00001%, 0.0001%, 0.001%, 0.01%, 0.1%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, and 27%. 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74% , 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or approximately 0.000000001%, 0.00000001%, 0.0000001%, 0.000001%, 0.00001%, 0.0001%, 0.001%, 0.01%, 0.01%.1%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 100%, or any number or range between any two of these values. In some embodiments, the percentage of the cell component-binding reagent in the mixture that is not conjugated with the binding reagent oligonucleotide is at least 0.000000001%, 0.00000001%, 0.0000001%, 0.000001%, 0.00001%, 0.0001%, 0.001%, 0.01% 0.1%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37 %, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 45%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 7 3%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%, or up to 0.000000001%, 0.00000001%, 0.0000001%, 0.000001%, 0.00001%, 0.0001%, 0.001%, 0.01%, 0.1%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 11%, 12%, 13%, 14%, 15%, 16%, 17%, 18%, 19%, 20%, 21%, 22%, 23%, 24%, 25%, 26%, 27%, 28%, 29%, 30%, 31%, 32%, 33%, 34%, 35%, 36%, 37%, 38%, 39%, 40%, 41%, 42%, 43%, 44%, 4 5%, 46%, 47%, 48%, 49%, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100%.

[0206] Cellular component cocktail In some embodiments, a cocktail of cell component-binding reagents (e.g., an antibody cocktail) can be used to increase the sensitivity of labeling in the methods disclosed herein. While not bound by any particular theory, this may be because it is difficult to find a universal cell component-binding reagent or antibody that labels all cell types, as cell component expression or protein expression can vary between cell types and cellular states. For example, using a cocktail of cell component-binding reagents can enable more sensitive and efficient labeling of a wider range of sample types. A cocktail of cell component-binding reagents may contain two or more different types of cell component-binding reagents, e.g., a broader range of cell component-binding reagents or antibodies. Cell component-binding reagents that label different cell component targets can be pooled together to create a cocktail that adequately labels all cell types of interest or one or more cell types.

[0207] In some embodiments, each of the multiple compositions (e.g., sample index-granting compositions) contains a cell component-binding reagent. In some embodiments, one of the multiple compositions contains two or more cell component-binding reagents, each of which is accompanied by a binding reagent oligonucleotide (e.g., a sample index-granting oligonucleotide), and at least one of the two or more cell component-binding reagents is capable of specifically binding to at least one of one or more cell component targets. The sequences of the binding reagent oligonucleotides accompanied by the two or more cell component-binding reagents may be identical. The sequences of the binding reagent oligonucleotides accompanied by the two or more cell component-binding reagents may include different sequences. Each of the multiple compositions may contain two or more cell component-binding reagents.

[0208] The number of different types of cell component binding reagents (e.g., CD147 antibody and CD47 antibody) in a composition may vary depending on the implementation. A composition having two or more different types of cell component binding reagents may be referred to herein as a cell component binding reagent cocktail (e.g., a sample indexing composition cocktail). The number of different types of cell component binding reagents in a cocktail may vary. In some embodiments, the number of different types of cell component-binding reagents in the cocktail may be 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 10000, 100000, or approximately 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 100000, 100000, or a number or range between any two of these values. In some embodiments, the number of different types of cell component-binding reagents in the cocktail may be at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 10000, or 100000, or at most 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 1000, 10000, or 100000. Different types of cell component-binding reagents may be conjugated with binding reagent oligonucleotides having the same or different sequences (e.g., sample indexing sequences).

[0209] Methods for quantitative analysis of cellular component targets In some embodiments, the methods disclosed herein can also be used for the quantitative analysis of multiple cellular component targets (e.g., protein targets) in a sample using oligonucleotide probes to which the compositions and barcode sequences (e.g., molecularly labeled sequences) disclosed herein can be attached to oligonucleotides of cell component binding reagents (e.g., protein binding reagents). The oligonucleotides of the cell component binding reagents may be or include antibody oligonucleotides, sample indexing oligonucleotides, cell-recognizing oligonucleotides, control particle oligonucleotides, control oligonucleotides, interaction-determining oligonucleotides, etc. In some embodiments, the sample may be a single cell, multiple cells, a tissue sample, a tumor sample, a blood sample, etc. In some embodiments, the sample may include normal cells, tumor cells, blood cells, B cells, T cells, maternal cells, fetal cells, etc., or a mixture of cells from various subjects. In some embodiments, the sample may consist of multiple single cells separated into individual compartments, such as microwells in a microwell array.

[0210] In some embodiments, the binding targets of the multiple cellular component targets (i.e., cellular component targets) may be carbohydrates, lipids, proteins, extracellular proteins, cell surface proteins, cell markers, B cell receptors, T cell receptors, major histocompatibility complexes, tumor antigens, receptors, integrins, intracellular proteins, or any combination thereof. In some embodiments, the cellular component targets are protein targets. In some embodiments, the multiple cellular component targets include cell surface proteins, cell markers, B cell receptors, T cell receptors, antibodies, major histocompatibility complexes, tumor antigens, receptors, or any combination thereof. In some embodiments, the multiple cellular component targets may include intracellular cellular components. In some embodiments, the multiple cellular components may represent at least 1%, at least 2%, at least 3%, at least 4%, at least 5%, at least 6%, at least 7%, at least 8%, at least 9%, at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, at least 99%, or a higher proportion of all cellular components encoded in the organism. In some embodiments, the multiple cellular component targets may comprise at least 2, at least 3, at least 4, at least 5, at least 10, at least 20, at least 30, at least 40, at least 50, at least 100, at least 1000, at least 10000, or more different cellular component targets.

[0211] In some embodiments, multiple cell component-binding reagents are brought into contact with the sample for specific binding to multiple cell component targets. Unbound cell component-binding reagents can be removed, for example, by washing. In embodiments where the sample contains cells, any cell component-binding reagents that are not specifically bound to cells can be removed.

[0212] In some examples, cells derived from a cell population can be separated (e.g., isolated) in wells of the substrate of this disclosure. The cell population can be diluted before separation. The cell population can be diluted so that at least 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% of the wells of the substrate are receptive to single cells. The cell population can be diluted so that up to 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% of the wells of the substrate are receptive to single cells. A cell population can be diluted such that the number of cells in the diluted population is 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% of the number of wells on the substrate, or at least 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100%. A cell population can be diluted such that the number of cells in the diluted population is 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% of the number of wells on the substrate, or at least 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100%. In some examples, the cell population is diluted so that the number of cells is approximately 10% of the number of wells on the substrate.

[0213] The distribution of single cells into the substrate wells may follow a Poisson distribution. For example, there may be at least 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10%, or higher, probabilities that a substrate well contains two or more cells. The distribution of single cells into the substrate wells may be random. The distribution of single cells into the substrate wells may not be random. Cells may be separated so that each substrate well receives only one cell.

[0214] In some embodiments, the cell component-binding reagent may be further conjugated with a fluorescent molecule to enable flow sorting of cells into individual compartments. In some embodiments, the methods disclosed herein provide for contacting a sample with multiple compositions for specific binding to multiple cellular component targets. The conditions used will be recognized as being able to enable the specific binding of cellular component-binding reagents, such as antibodies, to cellular component targets. After the contact step, unbound compositions may be removed. For example, in embodiments in which the sample contains cells and the compositions specifically bind to cellular component targets, which are cell surface cellular components such as cell surface proteins, the unbound compositions can be removed by washing the cells with a buffer so that only the compositions that specifically bind to the cellular component targets remain with the cells.

[0215] In some embodiments, the methods disclosed herein may involve associating oligonucleotides (e.g., barcodes or probabilistic barcodes), including barcode sequences (e.g., molecular labels), cell labels, sample labels, etc., or any combination thereof, with a plurality of oligonucleotides associated with a cell component-binding reagent. For example, a plurality of oligonucleotide probes including barcodes can be used to hybridize a plurality of oligonucleotides in a composition.

[0216] In some embodiments, the oligonucleotide probes may be immobilized on a solid support. The solid support may be floating, for example, beads in solution. The solid support may be encapsulated in a semi-solid or solid array. In some embodiments, the oligonucleotide probes may not be immobilized on a solid support. If the oligonucleotide probes are located in close proximity to multiple oligonucleotides of a cell component binding reagent, the oligonucleotides of the cell component binding reagent may hybridize to the oligonucleotide probes. The oligonucleotide probes may be contacted in a non-depleting ratio such that each of the distinct oligonucleotides of the cell component binding reagent may be accompanied by an oligonucleotide probe having a different barcode sequence (e.g., molecular label) of this disclosure.

[0217] In some embodiments, the methods disclosed herein provide for detaching oligonucleotides from cell component-binding reagents that are specifically bound to cell component targets. Desorption can be carried out by various methods for separating chemical groups from cell component-binding reagents, such as UV cleavage, chemical treatment (e.g., dithiothreitol treatment), heating, enzymatic treatment, or a combination thereof. Desorption of oligonucleotides from cell component-binding reagents can be carried out before, after, or during the step of hybridizing multiple oligonucleotide probes to multiple oligonucleotides in a composition.

[0218] Method for simultaneous quantitative analysis of cellular components and nucleic acid targets In some embodiments, the methods disclosed herein may also be used for the simultaneous quantitative analysis of multiple cellular component targets (e.g., protein targets) and multiple nucleic acid target molecules in a sample, using oligonucleotide probes that may be accompanied by barcode sequences (e.g., molecular labeling sequences) on both the oligonucleotides and nucleic acid target molecules of the compositions and cellular component-binding reagents disclosed herein. Other methods for the simultaneous quantitative analysis of multiple cellular component targets and multiple nucleic acid target molecules are described in U.S. Patent Application No. 15 / 715028, filed September 25, 2017, whose entire contents are incorporated herein by reference. In some embodiments, the sample may be a single cell, multiple cells, a tissue sample, a tumor sample, a blood sample, etc. In some embodiments, the sample may include a mixture of cell types such as normal cells, tumor cells, blood cells, B cells, T cells, maternal cells, fetal cells, or a mixture of cells from different subjects.

[0219] In some embodiments, the sample may consist of multiple single cells separated into individual compartments, such as microwells in a microwell array.

[0220] In some embodiments, the multiple cellular component targets include cell surface proteins, cell markers, B cell receptors, T cell receptors, antibodies, major histocompatibility complexes, tumor antigens, receptors, or any combination thereof. In some embodiments, the multiple cellular component targets may include intracellular cellular components. In some embodiments, the multiple cellular components may be present in a number or range between 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, 99%, or any two of these values, of all cellular components of the organism or proteins expressed in one or more cells of the organism, or in a number or range between approximately 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, 99%, or any two of these values. In some embodiments, the multiple cellular components may be present in at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, or 99% of all cellular components, such as proteins, that can be expressed in one or more cells of an organism. In some embodiments, the multiple cellular component targets may include 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000, 10000, or a number or range between any two of these values, or approximately 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000, 10000, or a number or range between any two of these values. In some embodiments, the multiple cellular component targets may include at least 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000, or 10,000 different cellular component targets, or up to 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000, or 10,000 different cellular component targets. In some embodiments, multiple cell component-binding reagents are brought into contact with the sample to specifically bind to multiple cell component targets. Unbound cell component-binding reagents can be removed, for example, by washing. In embodiments where the sample contains cells, any cell component-binding reagents that are not specifically bound to cells can be removed.

[0221] In some examples, cells derived from a cell population can be separated (e.g., isolated) in wells of the substrate of this disclosure. The cell population can be diluted before separation. The cell population can be diluted so that at least 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% of the wells of the substrate can accept single cells. The cell population can be diluted so that up to 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% of the wells of the substrate can accept single cells. A cell population can be diluted such that the number of cells in the diluted population is 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% of the number of wells on the substrate, or at least 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100%. A cell population can be diluted such that the number of cells in the diluted population is 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100% of the number of wells on the substrate, or at least 1%, 5%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, or 100%. In some examples, the cell population is diluted so that the number of cells is approximately 10% of the number of wells on the substrate.

[0222] The distribution of single cells into the substrate wells may follow a Poisson distribution. For example, there may be at least 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, or 10%, or higher, probabilities that a substrate well contains two or more cells. The distribution of single cells into the substrate wells may be random. The distribution of single cells into the substrate wells may not be random. Cells may be separated so that each substrate well receives only one cell.

[0223] In some embodiments, the cell component-binding reagent may be further conjugated with a fluorescent molecule to enable flow sorting of cells into individual compartments. In some embodiments, the methods disclosed herein provide contacting a sample with multiple compositions for specific binding to multiple cellular component targets. The conditions used will be recognized as being able to enable the specific binding of cellular component-binding reagents, such as antibodies, to cellular component targets. After the contact step, unbound compositions may be removed. For example, in embodiments where the sample contains cells and the compositions specifically bind to cellular component targets on the cell surface, such as cell surface proteins, the unbound compositions can be removed by washing the cells with a buffer so that only the compositions that specifically bind to the cellular component targets remain with the cells.

[0224] In some embodiments, the methods disclosed herein may provide for the release of multiple nucleic acid target molecules from a sample, for example, cells. For example, cells can be lysed to release multiple nucleic acid target molecules. Cell lysis can be achieved by any of the following means, for example, chemical treatment, osmotic shock, heat treatment, mechanical treatment, optical treatment, or any combination thereof. Cells can be lysed by adding a cell lysis buffer containing a surfactant (e.g., SDS, Li dodecyl sulfate, Triton X-100, Tween-20, or NP-40), an organic solvent (e.g., methanol or acetone), or a digestive enzyme (e.g., proteinase K, pepsin, or trypsin), or any combination thereof.

[0225] Those skilled in the art will recognize that multiple nucleic acid molecules may include a variety of nucleic acid molecules. In some embodiments, multiple nucleic acid molecules may include DNA molecules, RNA molecules, genomic DNA molecules, mRNA molecules, rRNA molecules, siRNA molecules, or combinations thereof, and may be double-stranded or single-stranded. In some embodiments, multiple nucleic acid molecules may include 100, 1000, 10000, 20000, 30000, 40000, 50000, 100000, 1000000 species, or a number or range of species between any two of these values, or approximately 100, 1000, 10000, 20000, 30000, 40000, 50000, 100000, 1000000 species, or a number or range of species between any two of these values. In some embodiments, the plurality of nucleic acid molecules may include at least 100, 1000, 10000, 20000, 30000, 40000, 50000, 100000, or 1,000000 species, or up to 100, 1000, 10000, 20000, 30000, 40000, 50000, 100000, or 1,000000 species. In some embodiments, the plurality of nucleic acid molecules may originate from a sample, e.g., a single cell or a group of cells. In some embodiments, the plurality of nucleic acid molecules may be pooled from a group of samples, e.g., a group of single cells.

[0226] In some embodiments, the methods disclosed herein may involve associating barcodes (e.g., probabilistic barcodes) that may include barcode sequences (e.g., molecular labels), cell labels, sample labels, or any combination thereof, with multiple nucleic acid target molecules and multiple oligonucleotides of cell component-binding reagents. For example, multiple oligonucleotide probes including probabilistic barcodes can be used to hybridize multiple nucleic acid target molecules with multiple oligonucleotides of a composition.

[0227] In some embodiments, multiple oligonucleotide probes may be immobilized on a solid support. The solid support may be floating, for example, beads in solution. The solid support may be semi-solid or encapsulated in a solid array. In some embodiments, the multiple oligonucleotide probes may not be immobilized on a solid support. When the multiple oligonucleotide probes are located in close proximity to multiple nucleic acid target molecules and multiple oligonucleotides of cell component binding reagents, the multiple oligonucleotides of the nucleic acid target molecules and cell component binding reagents may hybridize to the oligonucleotide probes. Oligonucleotide probes can be brought into contact in non-depletion ratios such that each of the oligonucleotides of separate nucleic acid target molecules and cell component binding reagents may be accompanied by an oligonucleotide probe (e.g., molecularly labeled) having a different barcode sequence of this disclosure.

[0228] In some embodiments, the methods disclosed herein provide for detaching oligonucleotides from cell component-binding reagents specifically bound to cell component targets. Desorption can be carried out by various methods such as UV cleavage, chemical treatment (e.g., dithiothreitol treatment), heating, enzymatic treatment, or a combination thereof, to separate chemical groups from the cell component-binding reagent. Desorption of oligonucleotides from cell component-binding reagents can be carried out before, after, or during the step of hybridizing multiple oligonucleotide probes to multiple oligonucleotides among multiple nucleic acid target molecules and compositions.

[0229] Simultaneous quantitative analysis of protein and nucleic acid targets In some embodiments, the methods disclosed herein can also be used for the simultaneous quantitative analysis of multiple target molecules, e.g., proteins and nucleic acid targets. For example, the target molecules may be cellular components or may contain cellular components. Figure 6 shows a schematic diagram of an exemplary method for the simultaneous quantitative analysis of both nucleic acid targets and other cellular component targets (e.g., proteins) in a single cell. In some embodiments, several compositions 605, 605b, 605c, etc., are provided, each containing a cellular component-binding reagent such as an antibody. Different cellular component-binding reagents, such as antibodies, which bind to different cellular component targets, are conjugated with different unique identifiers. The cellular component-binding reagents can then be incubated with a sample containing multiple cells 610. Different cellular component-binding reagents can specifically bind to cellular components on the cell surface, such as cell markers, B cell receptors, T cell receptors, antibodies, major histocompatibility complexes, tumor antigens, receptors, or any combination thereof. Unbound cellular component-binding reagents can be removed, for example, by washing the cells with a buffer. The cells containing the cell component binding reagent can then be separated into multiple compartments, such as a microwell array, where each compartment 615 is sized to accommodate a single cell and a single bead 620. Each bead may contain multiple oligonucleotide probes, which may include all oligonucleotide probes on the bead, a common cell label, and a barcode sequence (e.g., a molecular label sequence). In some embodiments, each oligonucleotide probe may include a target-binding region, e.g., a poly(dT) sequence. The oligonucleotide 625 conjugated to the cell component binding reagent can be detached from the cell component binding reagent by chemical, optical, or other means. The cells can be lysed (635) to release intracellular nucleic acids, such as genomic DNA or cellular mRNA 630. The cellular mRNA 630, oligonucleotide 625, or both, can be captured by the oligonucleotide probe on the bead 620, for example, by hybridizing to a poly(dT) sequence.Using reverse transcriptase, oligonucleotide probes hybridized to cellular mRNA 630 and oligonucleotide 625 can be extended using cellular mRNA 630 and oligonucleotide 625 as templates. The extension product produced by reverse transcriptase can be subjected to amplification and sequencing. Sequence reads can be subjected to demultiplexing of sequences or identifiers such as cell labels, barcodes (e.g., molecular labels), genes, and oligonucleotides specific to cellular component-binding reagents (e.g., antibody-specific oligonucleotides), which can result in a digital representation of the cellular components and gene expression of each single cell in the sample.

[0230] Barcode attached Oligonucleotides associated with cell component-binding reagents (e.g., antigen-binding reagents or protein-binding reagents) and / or nucleic acid molecules may be randomly associated with oligonucleotide probes (e.g., barcodes such as probabilistic barcodes). Oligonucleotides associated with cell component-binding reagents, as referred to herein as binding reagent oligonucleotides, may be, or include, the oligonucleotides of this disclosure, e.g., antibody oligonucleotides, sample indexing oligonucleotides, cell-identifying oligonucleotides, control particle oligonucleotides, control oligonucleotides, interaction-determining oligonucleotides, etc. The association includes, for example, hybridization of the target-binding region of the oligonucleotide probe to a complementary portion of the oligonucleotide of the target nucleic acid molecule and / or protein-binding reagent. For example, the oligo(dT) region of a barcode (e.g., probabilistic barcode) can interact with the poly(A) tail of the target nucleic acid molecule and / or the poly(dA) tail of the oligonucleotide of the protein-binding reagent. Assay conditions used for hybridization (e.g., pH, ionic strength, temperature of the buffer, etc.) can be selected to facilitate the formation of a particular stable hybrid.

[0231] This disclosure provides a method for conjugating a molecular label to an oligonucleotide conjugated to a target nucleic acid and / or a cell component-binding reagent using reverse transcription. Reverse transcription may use both RNA and DNA as templates. For example, the oligonucleotide originally conjugated to the cell component-binding reagent may be either RNA or DNA bases, or both. The binding reagent oligonucleotide may be copied and conjugated (e.g., covalently) to the sequence of the binding reagent sequence, or a portion thereof, as well as to the cell label and barcode sequence (e.g., molecular label). As another example, an mRNA molecule may be copied and conjugated (e.g., covalently) to the sequence of the mRNA molecule, or a portion thereof, as well as to the cell label and barcode sequence (e.g., molecular label). In some embodiments, molecular labeling may be added by ligation of an oligonucleotide probe target-binding region with an oligonucleotide (e.g., currently or previously associated) with a portion of the target nucleic acid molecule and / or a cell component-binding reagent. For example, the target-binding region may include a nucleic acid sequence that can be specifically hybridized to a restriction site overhang (e.g., an EcoRI adherent end overhang). The method may further include treating the target nucleic acid and / or the oligonucleotide associated with the cell component-binding reagent with a restriction enzyme (e.g., EcoRI) to generate a restriction site overhang. A ligase (e.g., T4 DNA ligase) can be used to ligate the two fragments.

[0232] Determination of the number or presence of specific molecular labeling sequences In some embodiments, the methods disclosed herein include determining the number or presence of unique molecular labeling sequences for each unique identifier, each nucleic acid target molecule, and / or each binding reagent oligonucleotide (e.g., antibody oligonucleotide). For example, sequencing reads can be used to determine the number of unique molecular labeling sequences for each unique identifier, each nucleic acid target molecule, and / or each binding reagent oligonucleotide. As another example, sequencing reads can be used to determine the presence or absence of molecular labeling sequences (e.g., molecular labeling sequences associated with targets in sequencing reads, binding reagent oligonucleotides, antibody oligonucleotides, sample indexing oligonucleotides, cell identification oligonucleotides, control particle oligonucleotides, control oligonucleotides, interaction determining oligonucleotides, etc.).

[0233] In some embodiments, the number of unique identifiers, each nucleic acid target molecule, and / or unique molecular labeling sequences for each binding reagent oligonucleotide indicates the amount of each cellular component target (e.g., antigen target or protein target) and / or each nucleic acid target molecule in the sample. In some embodiments, the amount of a cellular component target and the amount of its corresponding nucleic acid target molecule, e.g., mRNA molecule, can be compared with each other. In some embodiments, the ratio of the amount of a cellular component target and the amount of its corresponding nucleic acid target molecule, e.g., mRNA molecule, can be calculated. A cellular component target may be, for example, a cell surface protein marker. In some embodiments, the ratio of the protein level of a cell surface protein marker to the mRNA level of a cell surface protein marker is low.

[0234] The methods disclosed herein may be used for a variety of applications. For example, the methods disclosed herein may be used for proteome and / or transcriptome analysis of a sample. In some embodiments, the methods disclosed herein may be used to identify cellular component targets and / or nucleic acid targets, i.e., biomarkers, in a sample. In some embodiments, the cellular component targets and nucleic acid targets correspond to each other, i.e., the nucleic acid targets encode the cellular component targets. In some embodiments, the methods disclosed herein may be used to identify cellular component targets having a desired ratio between the amount of the cellular component target in a sample and the amount of its corresponding nucleic acid target molecule, e.g., mRNA molecule. In some embodiments, the ratio is 0.001, 0.01, 0.1, 1, 10, 100, 1000, or a number or range between any two of the above values, or approximately 0.001, 0.01, 0.1, 1, 10, 100, 1000, or a number or range between any two of the above values. In some embodiments, the ratio is at least, or at most, 0.001, 0.01, 0.1, 1, 10, 100, or 1000. In some embodiments, the methods disclosed herein can be used to identify cellular component targets in a sample where the amount of the corresponding nucleic acid target molecule in the sample is 1000, 100, 10, 5, 2¹, 0, or a number or range between any two of these values, or approximately 1000, 100, 10, 5, 2¹, 0, or a number or range between any two of these values. In some embodiments, the methods disclosed herein can be used to identify cellular component targets in a sample where the amount of the corresponding nucleic acid target molecule in the sample is greater than 1000, 100, 10, 5, 2¹, or 0, or less than 1000, 100, 10, 5, 2¹, or 0.

[0235] Compositions and kits Some embodiments disclosed herein provide kits and compositions for the simultaneous quantitative analysis of multiple cellular components (e.g., proteins) and / or multiple nucleic acid target molecules in a sample. In some embodiments, the kits and compositions may include multiple cellular component-binding reagents (e.g., multiple protein-binding reagents), each conjugated with an oligonucleotide, the oligonucleotide comprising a unique identifier for the cellular component-binding reagent and multiple oligonucleotide probes, each of which comprises a target-binding region, a barcode sequence (e.g., a molecular label sequence), and the barcode sequence derived from a diverse set of unique barcode sequences. In some embodiments, each oligonucleotide may include a molecular label, a cellular label, a sample label, or any combination thereof. In some embodiments, each oligonucleotide may include a linker. In some embodiments, each oligonucleotide may include a binding site for the oligonucleotide probe, e.g., a poly(A) tail. For example, the poly(A) tail may be, for example, oligodA 18 (Not anchored to a solid support) or oligoA 18 It may be V (anchored to a solid support). Oligonucleotides may contain DNA residues, RNA residues, or both.

[0236] The disclosure herein includes a plurality of sample indexing compositions. Each of the plurality of sample indexing compositions may include two or more cell component binding reagents. Each of the two or more cell component binding reagents may be accompanied by a sample indexing oligonucleotide. At least one of the two or more cell component binding reagents may be capable of specifically binding to at least one cell component target. The sample indexing oligonucleotide may include a sample indexing sequence for identifying the sample origin of one or more cells in the sample. The sample indexing sequences of at least two of the plurality of sample indexing compositions may include different sequences.

[0237] The disclosure herein includes a kit comprising sample indexing compositions for identifying cells. In some embodiments, each of the two sample indexing compositions comprises a cell component binding reagent (e.g., a protein binding reagent) accompanied by a sample indexing oligonucleotide, the cell component binding reagent being capable of specifically binding to at least one of one or more cell component targets (e.g., one or more protein targets), the sample indexing oligonucleotide comprising a sample indexing sequence, and the sample indexing sequences of the two sample indexing compositions comprising different sequences. In some embodiments, the sample indexing oligonucleotide comprises a molecular labeling sequence, a binding site for a universal primer, or a combination thereof.

[0238] The disclosure herein includes kits for identifying cells. In some embodiments, the kit includes two or more sample indexing compositions. Each of the two or more sample indexing compositions includes a cell component binding reagent (e.g., an antigen binding reagent) accompanied by a sample indexing oligonucleotide, the cell component binding reagent being capable of specifically binding to at least one of one or more cell component targets, the sample indexing oligonucleotide including a sample indexing sequence, and the sample indexing sequences of the two sample indexing compositions including different sequences. In some embodiments, the sample indexing oligonucleotide includes a molecular labeling sequence, a binding site for a universal primer, or a combination thereof. The disclosure herein also includes kits for multiple identification. In some embodiments, the kit includes two sample indexing compositions. Each of the two sample indexing compositions comprises a cell component binding reagent (e.g., an antigen binding reagent) accompanied by a sample indexing oligonucleotide, the antigen binding reagent being capable of specifically binding to at least one of one or more cell component targets (e.g., an antigen target), the sample indexing oligonucleotide comprising a sample indexing sequence, and the sample indexing sequences of the two sample indexing compositions comprising different sequences.

[0239] The unique identifier (or oligonucleotide associated with cell component-binding reagents such as binding reagent oligonucleotides, antibody oligonucleotides, sample indexing oligonucleotides, cell identification oligonucleotides, control particle oligonucleotides, control oligonucleotides, or interaction-determining oligonucleotides) may have any suitable length, for example, from about 25 nucleotides to about 45 nucleotides.In some embodiments, the unique identifier is in the range of 4 nucleotides, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 15 nucleotides, 20 nucleotides, 25 nucleotides, 30 nucleotides, 35 nucleotides, 40 nucleotides, 45 nucleotides, 50 nucleotides, 55 nucleotides, 60 nucleotides, 70 nucleotides, 80 nucleotides, 90 nucleotides, 100 nucleotides, 200 nucleotides, or any two of the above values. The length may be less than 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 70, 80, 90, 100, or 200 nucleotides, or within the range of any two of the above values, and may be greater than 4, 5, 6, 7, 8, 9, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 70, 80, 90, 100, or 200 nucleotides, or within the range of any two of the above values.

[0240] In some embodiments, the unique identifier is selected from a diverse set of unique identifiers. The diverse set of unique identifiers may include 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 5000 unique identifiers, or a different number or range between any two of these values; or it may include approximately 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 5000 unique identifiers, or a different number or range between any two of these values. A diverse set of unique identifiers may include at least 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, or 5000 different unique identifiers, or up to 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, or 5000 different unique identifiers. In some embodiments, the set of unique identifiers is designed to have minimal sequence homology to the DNA or RNA sequence of the sample being analyzed. In some embodiments, the sequence of the set of unique identifiers is distinct from each other or complementary to each other by a number or range between 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 nucleotides, or any two of these values, or by a number or range between approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 nucleotides, or any two of these values. In some embodiments, the sequence of the set of unique identifiers is distinct from each other or complementary to each other by at least, or up to 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nucleotides.

[0241] In some embodiments, the unique identifier may include binding sites to a primer, such as a universal primer. In some embodiments, the unique identifier may include at least two binding sites to a primer, such as a universal primer. In some embodiments, the unique identifier may include at least three binding sites to a primer, such as a universal primer. The primer may be used for amplification of the unique identifier, for example, by PCR amplification. In some embodiments, the primer may be used for a nested PCR reaction.

[0242] Any suitable cell component-binding reagent, such as any protein-binding reagent (e.g., antibodies or their fragments, aptamers, small molecules, ligands, peptides, oligonucleotides, etc., or any combination thereof), is intended in this disclosure. In some embodiments, the cell component-binding reagent may be a polyclonal antibody, a monoclonal antibody, a recombinant antibody, a single-chain antibody (scAb), or a fragment thereof, such as Fab, Fv, etc. In some embodiments, the multiple protein-binding reagents may include 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 5000 types, or a number or range between any two of these values ​​of different protein-binding reagents, or approximately 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 5000 types, or a number or range between any two of these values ​​of different protein-binding reagents. In some embodiments, the plurality of protein-binding reagents may include at least 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, or 5000 different protein-binding reagents, or up to 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, or 5000 different protein-binding reagents.

[0243] In some embodiments, oligonucleotides are conjugated with cell component-binding reagents via a linker. In some embodiments, oligonucleotides may be covalently conjugated with protein-binding reagents. In some embodiments, oligonucleotides may be noncovalently conjugated with protein-binding reagents. In some embodiments, the linker may include a chemical group that reversibly or irreversibly attaches the oligonucleotide to the protein-binding reagent. The chemical group may be conjugated to the linker via, for example, an amine group. In some embodiments, the linker may include a chemical group that forms a stable bond with another chemical group conjugated to the protein-binding reagent. For example, the chemical group may be a UV-cleavable group, a disulfide bond, streptavidin, biotin, or an amine. In some embodiments, the chemical group may be a primary amine on an amino acid such as lysine, or conjugated to the protein-binding reagent via the N-terminus. Oligonucleotides may be conjugated to any suitable site on the protein-binding reagent, as long as they do not interfere with the specific binding between the protein-binding reagent and its protein target. In embodiments where the protein-binding reagent is an antibody, the oligonucleotide has an antigen-binding site, for example, the Fc region, C H 1 domain, C H 2 domains, C H 3 domains, C LIt can be conjugated to any site of the antibody other than the domain. In some embodiments, each protein-binding reagent can be conjugated to a single oligonucleotide molecule. In some embodiments, each protein-binding reagent can be conjugated to 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000 oligonucleotide molecules, or a number or range between any two of these values, or to about 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000 oligonucleotide molecules, or a number or range between any two of these values, each oligonucleotide molecule containing the same unique identifier. In some embodiments, each protein-binding reagent can be conjugated to two or more oligonucleotide molecules, for example, at least, or up to 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, or 1000 oligonucleotide molecules, each oligonucleotide molecule containing the same unique identifier.

[0244] In some embodiments, a plurality of cell component-binding reagents (e.g., protein-binding reagents) can specifically bind to a plurality of cell component targets (e.g., protein targets) in the sample. The sample may be a single cell, a plurality of cells, a tissue sample, a tumor sample, a blood sample, or a combination thereof. In some embodiments, the plurality of cell component targets may include cell surface proteins, cell markers, B cell receptors, T cell receptors, antibodies, major histocompatibility complexes, tumor antigens, receptors, or any combination thereof. In some embodiments, the plurality of cell component targets may include intracellular proteins. In some embodiments, the multiple cellular component targets may represent 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, 99%, or any number or range between any two of these values ​​of all cellular component targets in the organism (e.g., proteins that are expressed or can be expressed), or approximately 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, 99%, or any number or range between any two of these values. In some embodiments, the multiple cellular component targets may represent at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 95%, 98%, or 99% of all cellular component targets in the organism (e.g., proteins that are expressed or can be expressed). In some embodiments, the multiple cellular component targets may include 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000, 10000, or a number or range between any two of these values, or approximately 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000, 10000, or a number or range between any two of these values.In some embodiments, the multiple cellular component targets may include at least 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000, or 10,000 different cellular component targets, or up to 2, 3, 4, 5, 10, 20, 30, 40, 50, 100, 1000, or 10,000 different cellular component targets.

[0245] Sample indexing using cell component-binding reagents conjugated with oligonucleotides. The disclosure herein includes methods for identifying samples. In some embodiments, the method includes contacting one or more cells derived from each of a plurality of samples with a sample indexing composition from a plurality of sample indexing compositions, each of the one or more cells comprising one or more cell component targets, each of the plurality of sample indexing compositions comprising a cell component binding reagent accompanied by a sample indexing oligonucleotide, the cell component binding reagent being capable of specifically binding to at least one of the one or more cell component targets, the sample indexing oligonucleotide comprising a sample indexing sequence, the sample indexing sequences of at least two of the plurality of sample indexing compositions comprising different sequences; barcoding (e.g., probabilistically barcoding) the sample indexing oligonucleotides using a plurality of barcodes (e.g., probabilistic barcodes) to generate a plurality of barcoded sample indexing oligonucleotides; obtaining sequencing data for the plurality of barcoded sample indexing oligonucleotides; and identifying the sample origin of at least one of the one or more cells based on the sample indexing sequence of at least one of the barcoded sample indexing oligonucleotides.

[0246] In some embodiments, barcoding a sample indexing oligonucleotide using multiple barcodes includes contacting the multiple barcodes with the sample indexing oligonucleotide to generate barcodes hybridized to the sample indexing oligonucleotide, and extending the barcodes hybridized to the sample indexing oligonucleotide to generate multiple barcoded sample indexing oligonucleotides. Extending the barcodes may include extending the barcodes using DNA polymerase to generate multiple barcoded sample indexing oligonucleotides. Extending the barcodes may also include extending the barcodes using reverse transcriptase to generate multiple barcoded sample indexing oligonucleotides.

[0247] Oligonucleotides conjugated to antibodies, oligonucleotides for conjugation with antibodies, or oligonucleotides previously conjugated to antibodies are referred to herein as antibody oligonucleotides ("AbOligo"). Antibody oligonucleotides in the context of sample indexing are referred to herein as sample indexing oligonucleotides. Antibodies conjugated to antibody oligonucleotides are referred to herein as hot antibodies or oligonucleotide antibodies. Antibodies not conjugated to antibody oligonucleotides are referred to herein as cold antibodies or oligonucleotide-free antibodies. Oligonucleotides conjugated to conjugating reagents (e.g., protein-binding reagents), oligonucleotides for conjugation with conjugating reagents, or oligonucleotides previously conjugated to conjugating reagents are referred to herein as reagent oligonucleotides. Reagent oligonucleotides in the context of sample indexing are referred to herein as sample indexing oligonucleotides. Conjugating reagents conjugated to antibody oligonucleotides are referred to herein as hot-binding reagents or oligonucleotide-binding reagents. Binding reagents that are not conjugated to antibody oligonucleotides are referred to herein as cold-binding reagents or oligonucleotide-free binding reagents.

[0248] Figure 7 shows a schematic diagram of an exemplary workflow using a cell component binding reagent accompanied by an oligonucleotide for sample indexing. In some embodiments, several compositions 705a, 705b, etc., are provided, each containing a binding reagent. The binding reagent may be a protein binding reagent, such as an antibody. The cell component binding reagent may include an antibody, tetramer, aptamer, protein scaffold, or a combination thereof. The binding reagents of several compositions 705a, 705b may bind to the same cell component target. For example, the binding reagents of several compositions 705, 705b may be identical (except for the sample indexing oligonucleotide accompanied by the binding reagent).

[0249] The various compositions may include binding reagents conjugated to sample indexing oligonucleotides having various sample indexing sequences. The number of various compositions may differ in different implementations. In some embodiments, the number of various compositions is 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, or a number or range between any two of these values. It may be approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, or a number or range between any two of these values. In some embodiments, the number of different compositions may be at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, or 10000.

[0250] In some embodiments, the sample index-granting oligonucleotides of the binding reagent in a single composition may include the same sample index-granting sequence. The sample index-granting oligonucleotides of the binding reagent in a single composition do not have to be identical. In some embodiments, the percentage of sample indexing oligonucleotides of a binding reagent in a single composition having the same sample indexing sequence is 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.9%, Or it may be a number or range between any two of these values, or approximately 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.9%, or a number or range between any two of these values. In some embodiments, the percentage of sample indexing oligonucleotides of a binding reagent in a single composition having the same sample indexing sequence may be at least, or up to, 50%, 51%, 52%, 53%, 54%, 55%, 56%, 57%, 58%, 59%, 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 99.9%.

[0251] Compositions 705a and 705b can be used to label samples from a variety of samples. For example, the sample indexing oligonucleotide of the cell component binding reagent in composition 705a may have one sample indexing sequence and can be used to label cells 710a, shown as black circles, in sample 707a, such as a patient's sample. The sample indexing oligonucleotide of the cell component binding reagent in composition 705b may have another sample indexing sequence and can be used to label cells 710b, shown as shaded circles, in sample 707b, such as a sample from a different patient or another sample from the same patient. The cell component binding reagent can specifically bind to cell component targets or proteins on the cell surface, such as cell markers, B cell receptors, T cell receptors, antibodies, major histocompatibility complexes, tumor antigens, receptors, or any combination thereof. Unbound cell component binding reagent can be removed, for example, by washing cells with a buffer.

[0252] The cells having the cell component binding reagent can then be separated into multiple compartments, such as a microwell array, where single compartments 715a, 715b are sized to accommodate a single cell 710a and a single bead 720a or a single cell 710b and a single bead 720b. Each bead 720a, 720b may contain multiple oligonucleotide probes, which may include a cell label and molecular label sequence common to all oligonucleotide probes on the bead. In some embodiments, each oligonucleotide probe may include a target-binding region, e.g., a poly(dT) sequence. The sample indexing oligonucleotide 725a conjugated to the cell component binding reagent of composition 705a may be configured to be detachable or indedetachable from the cell component binding reagent (or may be so). The sample indexing oligonucleotide 725a conjugated to the cell component binding reagent of composition 705a may be detached from the cell component binding reagent by chemical, optical or other means. The sample index-granting oligonucleotide 725b conjugated to the cell component-binding reagent of composition 705b may be configured to be detachable or indedetachable from the cell component-binding reagent (or may be detachable). The sample index-granting oligonucleotide 725b conjugated to the cell component-binding reagent of composition 705b may be detached from the cell component-binding reagent by chemical, optical, or other means.

[0253] Cell 710a can be lysed to release nucleic acids within cell 710a, such as genomic DNA or cell mRNA 730a. Lysed cell 735a is shown as a dashed circle. Cell mRNA 730a, sample indexing oligonucleotide 725a, or both can be captured by an oligonucleotide probe on bead 720a, for example, by hybridizing to a poly(dT) sequence. Using reverse transcriptase, the oligonucleotide probe hybridized to cell mRNA 730a and oligonucleotide 725a can be extended using cell mRNA 730a and oligonucleotide 725a as templates. The extension product produced by reverse transcriptase can be subjected to amplification and sequencing.

[0254] Similarly, cell 710b may be lysed to release nucleic acids within cell 710b, such as genomic DNA or cell mRNA 730b. Lysed cell 735b is shown as a dashed circle. Cell mRNA 730b, sample indexing oligonucleotide 725b, or both can be captured by an oligonucleotide probe on bead 720b, for example, by hybridizing to a poly(dT) sequence. Using reverse transcriptase, the oligonucleotide probe hybridized to cell mRNA 730b and oligonucleotide 725b can be extended using cell mRNA 730b and oligonucleotide 725b as templates. The extension product produced by reverse transcriptase can be subjected to amplification and sequencing.

[0255] Sequence reads can be subjected to demultiplexing for cell labeling, molecular labeling, gene identity, and sample identity (e.g., with respect to the sample indexing sequences of sample indexing oligonucleotides 725a and 725b). Demultiplexing for cell labeling, molecular labeling, and gene identity allows for a digital representation of gene expression in each single cell within the sample. Sample origin can be determined using demultiplexing for cell labeling, molecular labeling, and sample identity, with respect to the sample indexing sequences of sample indexing oligonucleotides.

[0256] In some embodiments, cell component-binding reagents on the cell surface can be conjugated to a library of unique sample-indexing oligonucleotides, allowing cells to retain sample identity. For example, antibodies against cell surface markers can be conjugated to a library of unique sample-indexing oligonucleotides, allowing cells to retain sample identity. This makes it possible to load multiple samples onto the same Rhapsody® cartridge, as information about the sample source is retained throughout library preparation and sequencing. Sample indexing can enable running multiple samples together in a single experiment, simplifying the process, reducing experimental time, and eliminating batch effects.

[0257] The disclosure herein includes methods for identifying samples. In some embodiments, the method includes contacting one or more cells derived from each of a plurality of samples with a sample indexing composition from a plurality of sample indexing compositions, each of the one or more cells comprising one or more cell component targets, each of the plurality of sample indexing compositions comprising a cell component binding reagent accompanied by a sample indexing oligonucleotide, the cell component binding reagent being capable of specifically binding to at least one of the one or more cell component targets, the sample indexing oligonucleotide comprising a sample indexing sequence, and the sample indexing sequences of at least two of the plurality of sample indexing compositions comprising different sequences; and removing an unbound sample indexing composition from the plurality of sample indexing compositions. This method may include: generating multiple barcoded sample indexing oligonucleotides by barcoding (e.g., probabilistically barcoding) sample indexing oligonucleotides using multiple barcodes (e.g., probabilistic barcodes); obtaining sequencing data for the multiple barcoded sample indexing oligonucleotides; and identifying the sample origin of at least one cell from one or more cells based on the sample indexing sequence of at least one barcoded sample indexing oligonucleotide from the multiple barcoded sample indexing oligonucleotides.

[0258] In some embodiments, a method for identifying a sample includes contacting one or more cells derived from each of a plurality of samples with a sample indexing composition from a plurality of sample indexing compositions, each of the one or more cells comprising one or more cell component targets, each of the plurality of sample indexing compositions comprising a cell component binding reagent accompanied by a sample indexing oligonucleotide, the cell component binding reagent being capable of specifically binding to at least one of the one or more cell component targets, the sample indexing oligonucleotide comprising a sample indexing sequence, the sample indexing sequences of at least two of the plurality of sample indexing compositions comprising different sequences; removing an unbound sample indexing composition from the plurality of sample indexing compositions; and identifying the sample origin of at least one of the one or more cells based on the sample indexing sequence of at least one sample indexing oligonucleotide from the plurality of sample indexing compositions.

[0259] In some embodiments, the identification of the sample origin of at least one cell includes barcoding (e.g., probabilistically barcoding) sample indexing oligonucleotides of multiple sample indexing compositions using multiple barcodes (e.g., probabilistic barcodes) to generate multiple barcoded sample indexing oligonucleotides; obtaining sequencing data for the multiple barcoded sample indexing oligonucleotides; and identifying the sample origin of the cell based on the sample indexing sequence of at least one of the multiple barcoded sample indexing oligonucleotides. In some embodiments, barcoding sample indexing oligonucleotides using multiple barcodes to generate multiple barcoded sample indexing oligonucleotides includes probabilistically barcoding sample indexing oligonucleotides using multiple probabilistic barcodes to generate multiple probabilistically barcoded sample indexing oligonucleotides.

[0260] In some embodiments, the identification of the sample origin of at least one cell may include identifying the presence or absence of a sample indexing sequence in at least one sample indexing oligonucleotide from a plurality of sample indexing compositions. The identification of the presence or absence of a sample indexing sequence includes replicating at least one sample indexing oligonucleotide to produce a plurality of replicated sample indexing oligonucleotides; obtaining sequencing data for the plurality of replicated sample indexing oligonucleotides; and identifying the sample origin of the cell based on the sample indexing sequences of the plurality of replicated sample indexing oligonucleotides corresponding to at least one barcoded sample indexing oligonucleotide in the sequencing data.

[0261] In some embodiments, replicating at least one sample-indexing oligonucleotide to produce multiple replica sample-indexing oligonucleotides includes ligating a replication adapter to at least one barcoded sample-indexing oligonucleotide before replicating the at least one barcoded sample-indexing oligonucleotide. Replicating at least one barcoded sample-indexing oligonucleotide may include replicating at least one barcoded sample-indexing oligonucleotide using a replication adapter ligated to at least one barcoded sample-indexing oligonucleotide to produce multiple replica sample-indexing oligonucleotides.

[0262] In some embodiments, duplicating at least one sample indexing oligonucleotide to produce multiple replicated sample indexing oligonucleotides includes contacting a capture probe with at least one sample indexing oligonucleotide before duplicating at least one barcoded sample indexing oligonucleotide to produce a capture probe hybridized to the sample indexing oligonucleotide, and extending the capture probe hybridized to the sample indexing oligonucleotide to produce a sample indexing oligonucleotide with the capture probe attached. Duplicating at least one sample indexing oligonucleotide may include duplicating a sample indexing oligonucleotide with the capture probe attached to produce multiple replicated sample indexing oligonucleotides.

[0263] Cell overloading and multiplet identification The disclosures herein also include methods, kits, and systems for identifying cellular overloading and multiplets. Such methods, kits, and systems may be used in or in combination with any suitable methods, kits, and systems disclosed herein, such as methods, kits, and systems for measuring the expression levels of cellular components (e.g., protein expression levels) using cell component-binding reagents accompanied by oligonucleotides.

[0264] With current cell loading techniques, when approximately 20,000 cells are loaded into a microwell cartridge or array with approximately 60,000 microwells, the number of microwells or droplets containing two or more cells (referred to as doublets or multiplets) can be minimized. However, as the number of loaded cells increases, the number of microwells or droplets containing multiple cells can increase significantly. For example, when approximately 50,000 cells are loaded into approximately 60,000 microwells in a microwell cartridge or array, the percentage of microwells containing multiple cells can be quite high, for example, 11-14%. Loading such a large number of cells into microwells is sometimes referred to as cell overloading. However, if the cells are divided into several groups (e.g., 5) and the cells in each group are labeled with sample-indexing oligonucleotides having distinct sample-indexing sequences, cell labels associated with two or more sample-indexing sequences (e.g., cell labels such as barcodes, such as probabilistic barcodes) can be identified in the sequencing data and removed from subsequent processing. In some embodiments, when cells are divided into a large number of groups (e.g., 10,000) and the cells in each group are labeled with a sample indexing oligonucleotide having a distinct sample indexing sequence, sample labels associated with two or more sample indexing sequences can be identified in the sequencing data and removed from subsequent processing. In some embodiments, various cells are labeled with cell-identifying oligonucleotides having distinct cell-identifying sequences, and cell-identifying sequences associated with two or more cell-identifying oligonucleotides can be identified in the sequencing data and removed from subsequent processing. Such a large number of cells can be loaded into microwells, depending on the number of microwells in a microwell cartridge or array.

[0265] The disclosure herein includes methods for identifying samples. In some embodiments, the method involves contacting a first plurality of cells and a second plurality of cells with two sample indexing compositions, each of which comprises one or more cellular components, and each of the two sample indexing compositions comprises a cellular component-binding reagent accompanied by a sample indexing oligonucleotide, the cellular component-binding reagent being specifically capable of binding to at least one of the one or more cellular components, and the sample indexing oligonucleotide comprises a sample indexing sequence, and the sample indexing sequences of the two sample indexing compositions comprise different sequences; or barcoding the sample indexing oligonucleotide using a plurality of barcodes to produce a plurality of barcoded sample indexing oligonucleotides, each of which comprises a cellular labeling sequence, a barcode sequence (e.g.) The barcode sequence of at least two of the barcodes includes a molecularly labeled sequence, a binding site for a universal primer, or a combination thereof.

[0266] For example, this method can be used to load more than 50,000 cells (compared to 10,000-20,000 cells) using sample indexing. Sample indexing allows cells from different samples to be labeled with a unique sample index using a cell component binding reagent (e.g., an antibody) conjugated with an oligonucleotide to a cell component (e.g., a universal protein marker) or a cell component binding reagent. When two or more cells from different samples, two or more cells from different cell populations of a sample, or two or more cells from a sample are captured in the same microwell or droplet, the combined "cells" (or contents of two or more cells) may be associated with sample indexing oligonucleotides having different sample indexing sequences (or cell-identifying oligonucleotides having different cell-identifying sequences). The number of different cell populations may vary in different implementations. In some embodiments, the number of different populations may be 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or a number or range between any two of these values, or approximately 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or a number or range between any two of these values. In some embodiments, the number of different populations may be at least, or at most, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100. The number of cells in each population, or the average number, may differ in different implementations. In some embodiments, the number of cells in each population, or the average number, may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or a number or range between any two of these values, or approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, or a number or range between any two of these values.In some embodiments, the number of cells in each population, or the average number, may be at least, or at most, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100. If the number of cells in each population, or the average number, is sufficiently small (e.g., equal to or less than 50, 25, 10, 5, 4, 3, 2, or 1 cell per population), the sample indexing composition for cell overloading and multiplet identification may be referred to as a cell identification composition.

[0267] The cells of a sample may be divided into multiple populations by dividing the cells of the sample into multiple equal populations. A “cell” associated with two or more sample indexing sequences in sequencing data may be identified as a “multiplet” based on two or more sample indexing sequences associated with one cell labeling sequence (e.g., a barcode, e.g., a cell labeling sequence for a probabilistic barcode) in sequencing data. The sequencing data of a combined “cell” is also referred to herein as a multiplet. A multiplet may be a doublet, triplet, quartet, quintet, sextet, septet, octet, nonet, or any combination thereof. A multiplet may also be any n-plette. In some embodiments, n is a range of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or any two of these values, or approximately a range of 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or any two of these values. In some embodiments, n is at least, or at most, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20.

[0268] When determining the expression profile of a single cell, two cells can be identified as one cell, and the expression profiles of two cells can be identified as the expression profile of one cell (referred to as a doublet expression profile). For example, when determining the expression profiles of two cells using barcoding (e.g., probabilistic barcoding), the mRNA molecules of the two cells may be associated with barcodes having the same cell marker. As another example, two cells may be associated with a single particle (e.g., a bead). The particle may contain a barcode having the same cell marker. After lysing the cells, the mRNA molecules of the two cells may be associated with the barcode of the particle, and therefore may have the same cell marker. Doublet expression profiles can distort the interpretation of expression profiles.

[0269] A doublet can refer to a combined "cell" accompanied by two sample-indexing oligonucleotides having different sample-indexing sequences. A doublet can also refer to a combined "cell" accompanied by two sample-indexing oligonucleotides having two different sample-indexing sequences. A doublet can occur when two cells accompanied by two sample-indexing oligonucleotides with different sequences (or two or more cells accompanied by sample-inde...

Claims

1. A method for identifying a sample, Each of the multiple samples is brought into contact with a sample indexing composition from among the multiple sample indexing compositions, Each of the plurality of samples comprises one or more cells, each containing one or more cellular component targets, and the sample index-granting composition comprises an aptamer composition comprising an aptamer and a sample index-granting oligonucleotide, wherein the aptamer is capable of specifically binding to at least one of the one or more cellular component targets. The sample index-granting oligonucleotide includes a sequence complementary to a capture sequence configured to capture the sequence of the sample index-granting oligonucleotide, and the sequence of the sample index-granting oligonucleotide complementary to the capture sequence includes a poly(dA) region. The sample indexing oligonucleotide comprises a sample indexing sequence, and the sample indexing sequences of at least two of the plurality of sample indexing compositions comprise different sequences, and Based on the sample indexing sequence of at least one sample indexing oligonucleotide of the plurality of sample indexing compositions, the sample origin of at least one of the one or more cells is identified. A method that includes this.

2. (a) Identification of the sample origin of the at least one cell is: To generate multiple barcoded sample indexing oligonucleotides of the multiple sample indexing compositions by barcoding them using multiple barcodes; Obtaining sequencing data for the plurality of barcoded sample indexing oligonucleotides; and The sample origin of the cells is identified based on the sample indexing sequence of at least one barcoded sample indexing oligonucleotide among the plurality of barcoded sample indexing oligonucleotides in the sequencing data. Includes, (b) Identification of the sample origin of the at least one cell includes identifying the presence or absence of a sample indexing sequence of at least one sample indexing oligonucleotide of the plurality of sample indexing compositions, and / or (c) Identification of the sample origin of at least one cell includes identifying the presence or absence of a sample indexing sequence of at least one sample indexing oligonucleotide of the plurality of sample indexing compositions, The identification of the presence or absence of the aforementioned sample indexing sequence is as follows: To generate multiple replicated sample indexing oligonucleotides by replicating at least one of the sample indexing oligonucleotides; Obtaining sequencing data for the plurality of replicated sample indexed oligonucleotides; and The sample origin of the cells is identified based on the sample indexing sequences of the replicated sample indexing oligonucleotides of the plurality of sample indexing oligonucleotides corresponding to the at least one barcoded sample indexing oligonucleotide in the sequencing data. including, The method according to claim 1.

3. (a) The cellular component target includes a protein target, (b) The sample indexing sequence is 6 to 60 nucleotides long or 50 to 500 nucleotides long. (c) The sample indexing sequences of at least 10, 100, or 1,000 of the sample indexing compositions include different sequences. (d) A single polynucleotide comprises the sample index-granting oligonucleotide and the aptamer, the aptamer is (i) In the single polynucleotide, the 5' side of the sample index-assigning oligonucleotide, (ii) The 3' side of the sample indexing oligonucleotide in the single polynucleotide, and / or (e) The sample index-assigning oligonucleotide is associated with the aptamer, (f) The sample index-assigned oligonucleotide is (i) attached to the aptamer, or (ii) Conjugated to the aptamer and / or (g) The sample index-assigning oligonucleotide is attached to the aptamer by non-covalent bond, (h) The sample index-granting oligonucleotide is attached to the aptamer via a linker, and / or (i) The aptamer comprises a nucleotide aptamer, the nucleotide aptamer comprises deoxyribonucleic acid (DNA), ribonucleic acid (RNA), xenonucleic acid (XNA), fluorophores, or combinations thereof. The method according to any one of claims 1 to 2.

4. (a) The sample indexing composition is attached to the first carrier, (b) The aptamer composition is attached to the first carrier and / or (c) The aptamer composition comprises a second aptamer capable of specifically binding to at least one of the one or more protein targets or the one or more cellular component targets, and a second sample indexing oligonucleotide containing a second sample indexing sequence, (i) The aptamer and the second aptamer are attached to the first carrier and / or (ii) The aptamer is attached to the first carrier, and the second aptamer is attached to the second carrier, and / or (iii) The aptamer and the second aptamer have at least 60%, 70%, 80%, 90%, or 95% sequence identity, and / or (iv) The aptamer and the second aptamer are the same, or the aptamer and the second aptamer are different, and / or (v) The protein target or cellular component target of the aptamer and the second aptamer are the same, The method according to any one of claims 1 to 3.

5. (a) The first carrier is (i) Metal nanomaterials, and / or (ii) Gold nanomaterials, and / or (iii) Lysosomes, micelles, vesicles, lipid membranes, lipid bilayers, lipid monolayers, or combinations thereof including, (b) The aptamer is attached to the first carrier, (c) The aptamer is conjugated to the first carrier, (d) The aptamer is non-covalently linked to the first carrier, (e) The aptamer is attached to the first carrier via a linker, and / or (f) The aptamer is immobilized on the first carrier, partially immobilized on the first carrier, immobilized within the first carrier, partially immobilized within the first carrier, encapsulated within the first carrier, partially encapsulated within the first carrier, embedded within the first carrier, partially embedded within the first carrier, or a combination thereof. The method according to claim 3 or 4.

6. (a) Dissociating the aptamer from the first carrier, wherein the dissociation is (i) after barcoding the sample indexing oligonucleotide, or (ii) before barcoding the sample indexing oligonucleotide, and / or (b) Contacting each of the plurality of samples with the sample indexing composition includes contacting one or more cells of the sample with the sample indexing composition, and / or (c) Contacting each of the plurality of samples with the sample indexing composition includes contacting one or more cells of the sample with the sample indexing composition, wherein the first carrier is internally transported into the cells, and the internal transport is by endocytosis, pinocytosis, nanopinocytosis, micropinocytosis, phagocytosis, membrane fusion, clathrin-mediated internal transport, caveolin-mediated internal transport, receptor-dependent internal transport, receptor-independent internal transport, or a combination thereof. The method according to any one of claims 4(c) or 5.

7. (a) At least one of the one or more protein targets or the one or more cellular component targets is located on the cell surface. (b) The method includes removing an unbound sample index-granting composition from among the plurality of sample index-granting compositions, (c) Removal of the unbound sample indexing composition comprises (a) washing one or more cells derived from each of the plurality of samples with a washing buffer, (b) selecting cells bound to at least one aptamer using flow cytometry, or both of (a) and (b). (d) The method comprises lysing one or more cells derived from each of the plurality of samples, (e) The sample index-assigning oligonucleotide is (i) configured to be inseparable from the aptamer, (ii) configured to be detachable from the aptamer, and / or (f) The method includes detaching the sample index-granting oligonucleotide from the aptamer, and / or (g) The method comprises detaching the sample index-granting oligonucleotide from the aptamer, wherein the detachment of the sample index-granting oligonucleotide comprises detaching the sample index-granting oligonucleotide from the aptamer by UV light cutting, chemical treatment, heating, enzymatic treatment, or any combination thereof. The method according to any one of claims 1 to 6.

8. (a) The sample index-assigned oligonucleotide is either not homologous to any of the genome sequences of the one or more cells, homologous to the genome sequence of a species, or a combination thereof, and the species is a non-mammalian species. (b) The sample among the plurality of samples includes a plurality of cells, a plurality of single cells, tissue, tumor sample, or any combination thereof. (c) The plurality of samples include mammalian cells, bacterial cells, viral cells, yeast cells, fungal cells, or any combination thereof. (d) The barcode includes a target binding region including the capture sequence, and / or (e) The sample indexing oligonucleotide comprises an alignment sequence adjacent to the poly(dA) region, wherein (a) the alignment sequence has a nucleotide length of 1 or more or a nucleotide length of 2 or more, (b) the alignment sequence comprises guanine, cytosine, thymine, uracil, or a combination thereof, and / or (c) the alignment sequence comprises a poly(dT) region, a poly(dG) region, a poly(dC) region, a poly(dU) region, or a combination thereof, and / or (f) The sample index-assigning oligonucleotide includes a molecular labeling sequence, a binding site for a universal primer, or both. (g) The protein target or the cellular component target includes carbohydrates, lipids, proteins, extracellular proteins, cell surface proteins, cell markers, B cell receptors, T cell receptors, major histocompatibility complexes, tumor antigens, receptors, intracellular proteins, or any combination thereof. (h) The protein target or cellular component target is selected from the group comprising 10 to 100 different protein targets or cellular component targets, and / or (i) The aptamer includes: (A) The same sequence, or (B) Different sample indexing sequences It is accompanied by two or more sample index-granting oligonucleotides having and / or (j) The sample index-granting composition among the plurality of sample index-granting compositions includes a second aptamer that is not accompanied by the sample index-granting oligonucleotide, and the aptamer and the second aptamer are the same, and / or (k) Each of the plurality of sample index-granting compositions comprises the aptamer, The method according to any one of claims 1 to 7.

9. A composition for assigning multiple sample indices, Each of the plurality of sample index-granting compositions comprises an aptamer composition containing a first cell component-binding aptamer and a sample index-granting oligonucleotide, The cellular component-binding aptamer is capable of specifically binding to at least one cellular component target. The sample indexing oligonucleotide comprises a sample indexing sequence for identifying the sample origin of one or more cells in the sample. The sample index-granting oligonucleotide includes a sequence complementary to a capture sequence configured to capture the sequence of the sample index-granting oligonucleotide. The sequence of the sample index-assigning oligonucleotide complementary to the capture sequence includes a poly(dA) region. The sample indexing sequences of at least two of the above-mentioned plurality of sample indexing compositions include different sequences. Multiple sample indexing compositions.

10. (a) The cellular component target includes a protein target, (b) The sample indexing sequence is 6 to 60 nucleotides long or 50 to 500 nucleotides long. (c) The sample indexing sequences of at least 10, 100, or 1,000 of the sample indexing compositions include different sequences. (d) A single polynucleotide comprises the sample index-granting oligonucleotide and the aptamer, the aptamer is (i) In the single polynucleotide, the 5' side of the sample index-assigning oligonucleotide, (ii) The 3' side of the sample indexing oligonucleotide in the single polynucleotide, and / or (e) The sample index-assigning oligonucleotide is associated with the aptamer, (f) The sample index-assigned oligonucleotide is (i) attached to the aptamer, or (ii) Conjugated to the aptamer and / or (g) The sample index-assigning oligonucleotide is attached to the aptamer by non-covalent bond, (h) The sample index-granting oligonucleotide is attached to the aptamer via a linker, and / or (i) The aptamer comprises a nucleotide aptamer, the nucleotide aptamer comprises deoxyribonucleic acid (DNA), ribonucleic acid (RNA), xenonucleic acid (XNA), fluorophores, or combinations thereof. A plurality of sample indexing compositions according to claim 9.

11. (a) The sample indexing composition is attached to the first carrier, (b) The aptamer composition is attached to the first carrier and / or (c) The aptamer composition comprises a second aptamer capable of specifically binding to at least one of the one or more protein targets or the one or more cellular component targets, and a second sample indexing oligonucleotide containing a second sample indexing sequence, (i) The aptamer and the second aptamer are attached to the first carrier and / or (ii) The aptamer is attached to the first carrier, and the second aptamer is attached to the second carrier, and / or (iii) The aptamer and the second aptamer have at least 60%, 70%, 80%, 90%, or 95% sequence identity, and / or (iv) The aptamer and the second aptamer are the same, or the aptamer and the second aptamer are different, and / or (v) The protein target or cellular component target of the aptamer and the second aptamer are the same, A plurality of sample indexing compositions according to claim 9 or 10.

12. (a) The first carrier is (i) Metal nanomaterials, and / or (ii) Gold nanomaterials, and / or (iii) Lysosomes, micelles, vesicles, lipid membranes, lipid bilayers, lipid monolayers, or combinations thereof including, (b) The aptamer is attached to the first carrier, (c) The aptamer is conjugated to the first carrier, (d) The aptamer is non-covalently linked to the first carrier, (e) The aptamer is attached to the first carrier via a linker, and / or (f) The aptamer is immobilized on the first carrier, partially immobilized on the first carrier, immobilized within the first carrier, partially immobilized within the first carrier, encapsulated within the first carrier, partially encapsulated within the first carrier, embedded within the first carrier, partially embedded within the first carrier, or a combination thereof. A plurality of sample indexing compositions according to claim 10 or 11.

13. (a) The sample index-assigning oligonucleotide is either not homologous to any of the genome sequences of the one or more cells, homologous to the genome sequence of a species, or a combination thereof, and the species is a non-mammalian species. (b) The sample among the plurality of samples includes a plurality of cells, a plurality of single cells, tissue, tumor sample, or any combination thereof. (c) The plurality of samples include mammalian cells, bacterial cells, viral cells, yeast cells, fungal cells, or any combination thereof. (d) The sample indexing oligonucleotide comprises an alignment sequence adjacent to the poly(dA) region, wherein (a) the alignment sequence has a nucleotide length of 1 or more or a nucleotide length of 2 or more, (b) the alignment sequence comprises guanine, cytosine, thymine, uracil, or a combination thereof, and / or (c) the alignment sequence comprises a poly(dT) region, a poly(dG) region, a poly(dC) region, a poly(dU) region, or a combination thereof, and / or (e) The sample index-assigning oligonucleotide includes a molecular labeling sequence, a binding site for a universal primer, or both. (f) The protein target or cellular component target includes carbohydrates, lipids, proteins, extracellular proteins, cell surface proteins, cell markers, B cell receptors, T cell receptors, major histocompatibility complexes, tumor antigens, receptors, intracellular proteins, or any combination thereof. (g) The protein target or cellular component target is selected from the group comprising 10 to 100 different protein targets or cellular component targets, and / or (h) The aptamer includes, (A) The same sequence, or (B) Different sample indexing sequences It is accompanied by two or more sample index-granting oligonucleotides having and / or (i) The sample index-granting composition among the plurality of sample index-granting compositions comprises a second aptamer that is not accompanied by the sample index-granting oligonucleotide, and / or (j) The sample index-granting composition among the plurality of sample index-granting compositions includes a second aptamer that is not accompanied by the sample index-granting oligonucleotide, and the aptamer and the second aptamer are the same, and / or (k) Each of the plurality of sample index-granting compositions comprises the aptamer, A plurality of sample index-granting compositions according to any one of claims 9 to 12.

14. A method for measuring the expression of cellular components in cells, A step of contacting a plurality of cell component-binding aptamers with a plurality of cells containing a plurality of cell component targets, wherein each of the plurality of cell component-binding aptamers comprises an aptamer-specific oligonucleotide comprising a unique identifier sequence for the cell component-binding aptamer, the cell component-binding aptamer is capable of specifically binding to at least one of the plurality of cell component targets, the aptamer-specific oligonucleotide comprises a sequence complementary to a capture sequence configured to capture the sequence of the aptamer-specific oligonucleotide, and the sequence of the aptamer-specific oligonucleotide complementary to the capture sequence comprises a poly(dA) region; A step of extending an oligonucleotide probe hybridized to the aptamer-specific oligonucleotide to produce a plurality of labeled nucleic acids, wherein each of the labeled nucleic acids includes a unique identifier sequence or its complementary sequence and a barcode sequence; and A step of obtaining sequence information of the plurality of labeled nucleic acids or a portion thereof, and determining the amount of one or more of the plurality of cellular component targets in one or more of the plurality of cells. A method that includes this.

15. Before extending the oligonucleotide probe, A step of distributing the plurality of cells, each associated with the plurality of cell component-binding aptamers, into a plurality of compartments, wherein each compartment contains a single cell derived from the plurality of cells associated with the cell component-binding aptamers; A step of contacting a barcoding particle with an aptamer-specific oligonucleotide in a compartment containing the single cell, wherein the barcoding particle comprises a plurality of oligonucleotide probes, each comprising a target-binding region and a barcode sequence selected from a diverse set of unique barcode sequences. The method according to claim 14, including the method described in claim 14.

16. (a) The plurality of cellular component targets comprises a plurality of protein targets, and the cellular component-binding aptamer is capable of specifically binding to at least one of the plurality of protein targets, and / or (b) The aptamer-specific oligonucleotide and the aptamer form a single polynucleotide, and the aptamer is (i) In the single polynucleotide, the 5' end of the aptamer-specific oligonucleotide, (ii) The 3' side of the aptamer-specific oligonucleotide in the single polynucleotide, and / or (c) The aptamer-specific oligonucleotide is associated with the aptamer, (d) The aptamer-specific oligonucleotide is attached to the aptamer and / or (e) The aptamer-specific oligonucleotide is (i) Covalently attached to the aptamer, (ii) Non-covalently attached to the aptamer, and / or (f) The aptamer-specific oligonucleotide is conjugated to the aptamer and / or (g) The aptamer comprises a nucleotide aptamer, the nucleotide aptamer comprises deoxyribonucleic acid (DNA), ribonucleic acid (RNA), xenonucleic acid (XNA), fluorophore, or a combination thereof. The method according to any one of claims 14 to 15.

17. The plurality of cell component-binding aptamers include a second cell component-binding aptamer, (a) (i) The aptamer and the second aptamer are attached to the first carrier, or (ii) The aptamer is attached to the first carrier, and the second aptamer is attached to the second carrier, and / or (b) The aptamer and the second aptamer are identical in sequence by at least 60%, 70%, 80%, 90%, or 95%, and / or (c) The aptamer and the second aptamer (i) They are the same, or (ii) Unlike, (A) The cellular component target of the aptamer and the second aptamer is the same. (B) The aptamer and the second aptamer are capable of binding to different regions of cellular component targets, or (C) The cellular component targets of the aptamer and the second aptamer are different, and / or (d) The first carrier is (i) Metal nanomaterials, and / or (ii) Gold nanomaterials, and / or (iii) Lysosomes, micelles, vesicles, lipid membranes, lipid bilayers, lipid monolayers, or combinations thereof including and / or (e) The aptamer is attached to the first carrier, (f) The aptamer is conjugated to the first carrier, (g) The aptamer is non-covalently linked to the first carrier, (h) The aptamer is attached to the first carrier via a linker, and / or (i) The aptamer is immobilized on the first carrier, partially immobilized on the first carrier, immobilized within the first carrier, partially immobilized within the first carrier, encapsulated within the first carrier, partially encapsulated within the first carrier, embedded within the first carrier, partially embedded within the first carrier, or a combination thereof. The method according to any one of claims 14 to 16.

18. (a) The barcode includes a target binding region including the capture sequence, and / or (b) The aptamer-specific oligonucleotide changes from a first conformation in which the poly(dA) region is inaccessible to a second conformation in which the poly(dA) region is accessible when the aptamer comes into contact with at least one of the plurality of cellular component targets. (c) The oligonucleotide probe is hybridized to the aptamer-specific oligonucleotide by hybridization between the poly(dA) region of the aptamer-specific oligonucleotide and the poly(dT) region of the oligonucleotide probe, and / or (d) The aptamer-specific oligonucleotide comprises an alignment sequence adjacent to the poly(dA) region, wherein (a) the alignment sequence has a nucleotide length of 1 or more or a nucleotide length of 2 or more, (b) the alignment sequence comprises guanine, cytosine, thymine, uracil, or a combination thereof, and / or (c) the alignment sequence comprises a poly(dT) region, a poly(dG) region, a poly(dC) region, a poly(dU) region, or a combination thereof, and / or (e) The aptamer-specific oligonucleotide comprises a molecular labeling sequence, a binding site for a universal primer, or both. The method according to claim 17.

19. A composition comprising a plurality of cell component-binding aptamers, each of the plurality of cell component-binding aptamers comprising an aptamer-specific oligonucleotide comprising a unique identifier sequence for the cell component-binding aptamer, the cell component-binding aptamer being capable of specifically binding to at least one of a plurality of cell component targets, the aptamer-specific oligonucleotide comprising a sequence complementary to a capture sequence configured to capture the sequence of the aptamer-specific oligonucleotide, the sequence of the aptamer-specific oligonucleotide complementary to the capture sequence comprising a poly(dA) region.

20. (a) The plurality of cellular component targets comprises a plurality of protein targets, and the cellular component-binding aptamer is capable of specifically binding to at least one of the plurality of protein targets, and / or (b) The aptamer-specific oligonucleotide and the aptamer form a single polynucleotide, and the aptamer is (i) In the single polynucleotide, the 5' end of the aptamer-specific oligonucleotide, (ii) The 3' side of the aptamer-specific oligonucleotide in the single polynucleotide, and / or (c) The aptamer-specific oligonucleotide is associated with the aptamer, (d) The aptamer-specific oligonucleotide is attached to the aptamer and / or (e) The aptamer-specific oligonucleotide is (i) Covalently attached to the aptamer, (ii) Non-covalently attached to the aptamer, and / or (f) The aptamer-specific oligonucleotide is conjugated to the aptamer and / or (g) The aptamer comprises a nucleotide aptamer, the nucleotide aptamer comprises deoxyribonucleic acid (DNA), ribonucleic acid (RNA), xenonucleic acid (XNA), fluorophore, or a combination thereof. The composition according to claim 19.

21. The plurality of cell component-binding aptamers include a second cell component-binding aptamer, (a) (i) The aptamer and the second aptamer are attached to the first carrier, or (ii) The aptamer is attached to the first carrier, and the second aptamer is attached to the second carrier, and / or (b) The aptamer and the second aptamer are identical in sequence by at least 60%, 70%, 80%, 90%, or 95%, and / or (c) The aptamer and the second aptamer (i) They are the same, or (ii) Unlike, (A) The cellular component target of the aptamer and the second aptamer is the same. (B) The aptamer and the second aptamer are capable of binding to different regions of cellular component targets, or (C) The cellular component targets of the aptamer and the second aptamer are different, and / or (d) The first carrier is (i) Metal nanomaterials, and / or (ii) Gold nanomaterials, and / or (iii) Lysosomes, micelles, vesicles, lipid membranes, lipid bilayers, lipid monolayers, or combinations thereof including and / or (e) The aptamer is attached to the first carrier, (f) The aptamer is conjugated to the first carrier, (g) The aptamer is non-covalently linked to the first carrier, (h) The aptamer is attached to the first carrier via a linker, and / or (i) The aptamer is immobilized on the first carrier, partially immobilized on the first carrier, immobilized within the first carrier, partially immobilized within the first carrier, encapsulated within the first carrier, partially encapsulated within the first carrier, embedded within the first carrier, partially embedded within the first carrier, or a combination thereof. The composition according to claim 19 or 20.

22. (a) The barcode includes a target binding region including the capture sequence, and / or (b) The aptamer-specific oligonucleotide changes from a first conformation in which the poly(dA) region is inaccessible to a second conformation in which the poly(dA) region is accessible when the aptamer comes into contact with at least one of the plurality of cellular component targets. (c) The oligonucleotide probe is hybridized to the aptamer-specific oligonucleotide by hybridization between the poly(dA) region of the aptamer-specific oligonucleotide and the poly(dT) region of the oligonucleotide probe, and / or (d) The aptamer-specific oligonucleotide comprises an alignment sequence adjacent to the poly(dA) region, wherein (a) the alignment sequence has a nucleotide length of 1 or more or a nucleotide length of 2 or more, (b) the alignment sequence comprises guanine, cytosine, thymine, uracil, or a combination thereof, and / or (c) the alignment sequence comprises a poly(dT) region, a poly(dG) region, a poly(dC) region, a poly(dU) region, or a combination thereof, and / or (e) The aptamer-specific oligonucleotide comprises a molecular labeling sequence, a binding site for a universal primer, or both. The composition according to claim 21.

Citation Information

Patent Citations

  • Assay for simultaneous genomic and proteomic analysis

    US20180208975A1

  • Measurement of protein expression using reagents with barcoded oligonucleotide sequences

    WO2018058073A2