Methods and compositions for obtaining linked functional and sequence data of single cells

The method integrates protein and gene expression data in single cells by assaying and indexing with nucleic acid-barcoded particles, improving cellular function understanding and translational research.

JP2026505689APending Publication Date: 2026-02-18BECTON DICKINSON & CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025534297
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-01-31
Filing Date
2024-01-26
Publication Date
2026-02-18

AI Technical Summary

Technical Problem

Current methods fail to effectively integrate protein expression data with transcriptome data in single cells, limiting understanding of cellular function and hindering translational research, diagnostic assays, and therapeutic development.

Method used

A method for linking functional and sequence data in single cells by functionally assaying divided cells, visually indexing with nucleic acid-barcoded identification particles, obtaining sequence data, and combining these data sets.

Benefits of technology

Enables comprehensive analysis of single cells by integrating protein and gene expression data, enhancing insights into cellular function and facilitating novel biomarker identification and therapeutic development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026505689000001_ABST
    Figure 2026505689000001_ABST
Patent Text Reader

Abstract

Methods for obtaining linked functional and sequence data for single cells (e.g., single cells from a cell sample) are provided. Aspects of the methods include functionally assaying the divided single cells; visually indexing the functionally assayed divided single cells using unique combinations of different nucleic acid-barcoded identification particles; obtaining sequence data for the visually indexed, functionally assayed, and divided single cells; and linking the functional and sequence data for the sequenced, visually indexed, functionally assayed, and divided single cells. Compositions for carrying out the methods of the invention are also provided.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] REFERENCE TO RELATED APPLICATIONS This application claims priority to the filing date of U.S. Provisional Patent Application No. 63 / 442,227, filed January 31, 2023, the disclosure of which is incorporated herein by reference. [Background technology]

[0002] Current technology allows for the measurement of gene expression in single cells in a massively parallel fashion (e.g., >10,000 cells) by attaching cell-specific oligonucleotide barcodes to poly(A) mRNA molecules from individual cells within a compartment where the individual cells colocalize with barcoding reagent beads. One platform capable of measuring single-cell gene expression in a massively parallel fashion is the BD Rhapsody™ Single-Cell Analysis System. The BD Rhapsody™ Single-Cell Analysis System enables high-throughput capture of nucleic acids from single cells using a simple cartridge workflow and a multitiered barcoding system. The resulting captured information can be used to generate various types of next-generation sequencing (NGS) libraries, including libraries suitable for whole-transcriptome analysis. Applications include discovery biology for sensitive transcript detection and targeted RNA analysis. Shum et al., “Quantitation of mRNA Transcripts and Proteins Using the BD Rhapsody(TM) Single-Cell Analysis System,” Adv Exp Med Biol. 2019;1129:63-79.

[0003] Gene expression can affect protein expression. Protein-protein interactions can affect gene and protein expression. Therefore, in recent years, systems and methods have been developed that can quantitatively analyze protein expression in cells and simultaneously measure protein and gene expression in cells. One such system is the BD Abseq platform. AbSeq is a method for profiling proteins in single cells. In AbSeq, typical fluorescent dye-conjugated antibodies are replaced with nucleic acid sequence tags that can be read at the single-cell level (e.g., through barcoding and NGS sequencing). The goal of Abseq is to characterize proteins in large numbers of single cells in a sensitive, accurate, and comprehensive manner. Cells are bound with antibodies against different target epitopes, similar to traditional immunostaining, except that the antibodies are labeled with unique sequence tags. When an antibody binds to its target, it carries along a DNA tag, allowing the presence of the target to be inferred based on the presence of the tag. In this method, by counting the number of tags, the number of different epitopes present in the cells detected by antibody binding can be estimated. Shahi et al., ``Abseq: Ultrahigh-throughput single cell protein profiling with droplet microfluidic barcoding. Sci Rep 7, 44447 (2017).'' Summary of the Invention [Problem to be solved by the invention]

[0004] summary

[0005] The present inventors have recognized that combining protein expression data with transcriptome data (e.g., as performed in AbSeq) provides important insights into single cells, while multiple genes, post-transcriptional and post-translational factors, and signaling pathways regulate cellular function. Understanding omics data in relation to cellular function and phenotype is valuable not only for further understanding single cells, but also for developing better strategies in translational research, including identifying novel biomarkers, developing diagnostic assays, and exploring novel therapeutics and clinical solutions. Therefore, the present inventors have recognized the need to provide a method for linking functional and sequence data (e.g., expression and / or transcriptome data) in single cells. [Means for solving the problem]

[0006] Embodiments of the present invention fulfill this need.

[0007] Methods for linking and obtaining functional and sequence data for single cells (e.g., cell samples) are provided. Aspects of the methods include functionally assaying divided single cells; visually indexing the functionally assayed divided single cells using unique combinations of different nucleic acid-barcoded identification particles; obtaining sequence data for the visually indexed, functionally assayed, and divided single cells; and linking the functional and sequence data for the sequenced, visually indexed, functionally assayed, and divided single cells. Compositions for carrying out the methods of the invention are also provided.

[0008] The invention will be best understood from the following detailed description when read in conjunction with the accompanying drawings, in which: [Brief explanation of the drawings]

[0009] [Figure 1A]1A and 1B show schematic diagrams of cell-binding beads according to embodiments of the present invention. [Figure 1B] 1A and 1B show schematic diagrams of cell-binding beads according to embodiments of the present invention. [Figure 2] FIG. 2 shows a schematic diagram of three different nucleic acid-barcoded identification particles (Pheno Seq particles) according to an embodiment of the present invention. [Figure 3A] 3A and 3B show views of partitions visually indexed with unique combinations of identification particles according to two different embodiments of the present invention. [Figure 3B] 3A and 3B show views of partitions visually indexed with unique combinations of identification particles according to two different embodiments of the present invention. [Figure 4] FIG. 4 shows how cell-binding bead nucleic acids are hybridized to both cell-capture bead nucleic acids and identification particle nucleic acids according to an embodiment of the present invention. [Figure 5A] 5A and 5B illustrate aspects of library preparation according to an embodiment of the present invention. [Figure 5B] 5A and 5B illustrate aspects of library preparation according to an embodiment of the present invention. [Figure 6] FIG. 6 illustrates how functional data is matched with sequence data according to an embodiment of the present invention. [Figure 7] FIG. 7 illustrates a workflow according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0010] definition

[0011] Unless otherwise defined, technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. See, e.g., Singleton et al., Dictionary of Microbiology and Molecular Biology 2nd ed., J. Wiley & Sons (New York, NY 1994); Sambrook et al., Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Press (Cold Spring Harbor, NY 1989). For purposes of this disclosure, the following terms are defined as follows:

[0012] As used herein, an antibody may be a full-length (e.g., naturally occurring or formed by the recombination process of normal immunoglobulin gene fragments) immunoglobulin molecule (e.g., an IgG antibody) or an immunologically active (i.e., specifically binding) portion of an immunoglobulin molecule (e.g., an antibody fragment). In some embodiments, an antibody is a functional antibody fragment. For example, an antibody fragment may be a portion of an antibody, such as F(ab')2, Fab', Fab, Fv, or sFv. An antibody fragment can bind to the same antigen recognized by a full-length antibody. Antibody fragments may include isolated fragments consisting of the variable regions of an antibody, such as an "Fv" fragment consisting of the variable regions of the heavy and light chains, and a recombinant single-chain polypeptide molecule ("scFv protein") in which the variable regions of the light and heavy chains are connected by a peptide linker. Exemplary antibodies may include, but are not limited to, antibodies against cancer cells, viruses, cell surface receptors (e.g., CD8, CD34, and CD45), and therapeutic antibodies.

[0013] As used herein, the terms "bound" or "associated with" may mean that two or more species are identifiable as being present in the same location at the same time. Association may also mean that two or more species are or were present in similar containers. Association may also mean an informatic association. For example, digital information about two or more species may be stored and used to determine that one or more of the species were present in the same location at a particular time. Association may also be a physical association. In some embodiments, two or more related species are "fixed," "bound," or "immobilized" to each other or to a common solid or semi-solid surface. Association may refer to covalent or non-covalent means for fixing a label to a solid or semi-solid support, such as a bead. Association may also be a covalent bond between a target and a label. Association can include hybridization between two molecules (eg, a target molecule and a label).

[0014] As used herein, the term "complementary" may refer to the ability for precise pairing between two nucleotides. For example, if a nucleotide at a given position in a nucleic acid can form a hydrogen bond with a nucleotide in another nucleic acid, the two nucleic acids are considered to be complementary to each other at that position. Complementarity between two single-stranded nucleic acid molecules may be "partial," in which only some nucleotides bind, or may be complete, in which case complete complementarity exists between the single-stranded molecules. If a first nucleotide sequence is complementary to a second nucleotide sequence, the first nucleotide sequence can be said to be "complementary" of the second sequence. If a first nucleotide sequence is complementary to the reverse (i.e., the order of nucleotides is reversed) sequence of the second sequence, the first nucleotide sequence can be said to be "reverse complementary" of the second sequence. As used herein, the terms "complement," "complementary," and "reverse complement" are used interchangeably. It is understood from the disclosure that if a molecule is capable of hybridizing to another molecule, then it is likely that the molecule is complementary to the hybridizing molecule.

[0015] As used herein, "nucleic acid" refers to a polynucleotide sequence or a fragment thereof. A nucleic acid can comprise nucleotides. A nucleic acid can be exogenous or endogenous to a cell. A nucleic acid can exist in a cell-free environment. A nucleic acid can be a gene or a fragment thereof. A nucleic acid can be DNA. A nucleic acid can be RNA. A nucleic acid can comprise one or more analogs (e.g., modified backbone, sugar, or nucleotide base). Non-limiting examples of analogs include 5-bromouracil, peptide nucleic acid, xenonucleic acid, morpholino, locked nucleic acid, glycol nucleic acid, threose nucleic acid, dideoxynucleotide, cordycepin, 7-deaza-GTP, fluorescent dyes (e.g., sugar-linked rhodamine or fluorescein), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queosine, and wyosine. "Nucleic acid," "polynucleotide," "target polynucleotide," and "target nucleic acid" can be used interchangeably.

[0016] Nucleic acids can contain one or more modifications (e.g., base modifications, backbone modifications) to confer new or enhanced properties (e.g., improved stability) to the nucleic acid. Nucleic acids can contain nucleic acid affinity tags. Nucleosides can be a base-sugar combination. The base portion of a nucleoside can be a heterocyclic base. The two most common classes of such heterocyclic bases are purines and pyrimidines. Nucleotides can also be nucleosides that further contain a phosphate group covalently attached to the sugar portion of the nucleoside. In nucleosides containing pentofluorosugars, the phosphate group is attached to the 2', 3', or 5' hydroxyl group of the sugar. In the formation of nucleic acids, the phosphate groups can covalently link adjacent nucleosides to each other to form linear polymeric compounds. The ends of this linear polymeric compound can further join to form a circular compound; however, linear compounds are generally preferred. Furthermore, linear compounds can have internal nucleotide base complementarity and thus can fold to form fully or partially double-stranded compounds. Within nucleic acids, the phosphate groups are commonly referred to as forming the internucleoside backbone of the nucleic acid. This linkage or backbone may be a 3' to 5' phosphodiester linkage.

[0017] Nucleic acids can contain modified backbones and / or modified internucleoside linkages. Modified backbones can include those that contain a phosphorus atom in the backbone and those that do not contain a phosphorus atom in the backbone. Suitable modified nucleic acid backbones containing a phosphorus atom can include, for example, phosphorothioates, chiral phosphorothioates, phosphorodithioates, phosphotriesters, aminoalkylphosphotriesters, methyl and other alkyl phosphonates, such as 3'-alkylene phosphonates, 5'-alkylene phosphonates, chiral phosphonates, phosphinates, phosphoramidates (including 3'-aminophosphoramidates and aminoalkylphosphoramidates), phosphorodiamidates, thiophosphoramidates, thiodialkylphosphonates, thiodialkylphosphotriesters, selenium phosphates, and boranophosphates with normal 3'-5' linkages, analogs with 2'-5' linkages, and those with reverse polarity, where one or more internucleoside linkages are 3'-to-3', 5'-to-5', or 2'-to-2'.

[0018] Nucleic acids can comprise polynucleotide backbones formed from short chain alkyl or cycloalkyl internucleoside linkages, mixed heteroatom and alkyl or cycloalkyl internucleoside linkages, or one or more short chain heteroatom or heterocyclic internucleoside linkages, including those with morpholino linkages (formed in part from the sugar portion of the nucleoside), siloxane backbones, sulfide, sulfoxide and sulfone backbones, formaacetyl and thioformacetyl backbones, methyleneformacetyl and thioformacetyl backbones, riboacetyl backbones, alkene-containing backbones, sulfamate backbones, methyleneimino and methylenehydrazino backbones, sulfonic acid and sulfonamide backbones, amide backbones, and others with mixed N, O, S and CH components.

[0019] Nucleic acids can include nucleic acid mimetics. The term "mimetic" is intended to include polynucleotides in which only the furanose ring or both the furanose ring and the internucleotide linkage are replaced with non-furanose groups; replacement of only the furanose ring may be referred to as a sugar substitute. The heterocyclic base group or modified heterocyclic base group may be maintained for hybridization with an appropriate target nucleic acid. One example of such a nucleic acid is peptide nucleic acid (PNA). In PNA, the sugar backbone of a polynucleotide may be replaced with an amide-containing backbone, particularly an aminoethylglycine backbone. The nucleotides may be retained and are bound directly or indirectly to aza nitrogen atoms in the amide portion of the backbone. The backbone in PNA compounds can contain two or more linked aminoethylglycine units, resulting in a PNA having an amide-containing backbone. The heterocyclic base group can be bound directly or indirectly to the aza nitrogen atoms in the amide portion of the backbone.

[0020] Nucleic acids can include morpholino backbone structures. For example, nucleic acids can include six-membered morpholino rings instead of ribose rings. In some of these embodiments, phosphorodiamidate or other non-phosphodiester internucleoside linkages can be substituted for phosphodiester linkages.

[0021] Nucleic acids can include structures consisting of linked morpholino units (e.g., morpholino nucleic acids), each consisting of a morpholino ring bound to a heterocyclic base. Linking groups can link the morpholino monomer units in morpholino nucleic acids. Nonionic morpholino-based oligomeric compounds may exhibit fewer undesired interactions with cellular proteins. Morpholino-based polynucleotides can also be nonionic nucleic acid mimics. Diverse compounds within the morpholino class can be conjugated using different linking groups. Yet another class of polynucleotide mimics is called cyclohexenyl nucleic acids (CeNA). The furanose ring normally present in nucleic acid molecules can be replaced with a cyclohexene ring. CeNA DMT-protected phosphoramidite monomers can be used to synthesize oligomeric compounds using phosphoramidite chemistry. Incorporation of CeNA monomers into nucleic acid chains can improve the stability of DNA / RNA hybrids. CeNA oligoadenylates can form complexes with nucleic acid complements with stabilities similar to those of native complexes. Further modifications can include locked nucleic acids (LNAs), in which a 2'-hydroxy group is attached to the 4' carbon atom of the sugar ring, forming a 2'-C,4'-C-oxymethylene bond to form a bicyclic sugar group. This bond is a methylene (-CH2) group bridging the 2' oxygen atom and the 4' carbon atom, where n can be 1 or 2. LNAs and LNA analogs can exhibit very high duplex thermal stability with complementary nucleic acids (Tm = +3 to +10°C), stability against 3'-exonucleolytic degradation, and good solubility.

[0022] Nucleic acids can also include modifications or substitutions of nucleobases (commonly abbreviated as "bases"). As used herein, "unmodified" or "natural" nucleobases can include purine bases (e.g., adenine (A) and guanine (G)) and pyrimidine bases (e.g., thymine (T), cytosine (C), and uracil (U)). Modified nucleobases may include other synthetic or natural nucleobases, such as 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine, 2-thiocytosine, 5-halouracil and cytosine, 5-propynyl (-C=C-CH3) uracil and other alkynyl derivatives of cytosine and pyrimidine bases, 6- Azo-uracil, cytosine and thymine, 5-uracil (pseudouracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxy and other 8-substituted adenines and guanines, 5-halo (especially 5-bromo), 5-trifluoromethyl and other 5-substituted uracils and cytosines, 7-methylguanine and 7-methyladenine, 2-F-adenine, 2-aminoadenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine, and 3-deazaguanine and 3-deazaadenine.Modified nucleobases may include tricyclic pyrimidines, e.g., phenoxazine cytidine (1H-pyrimido(5,4-b)(1,4)benzoxazin-2(3H)-one), phenothiazine cytidine (1H-pyrimido(5,4-b)(1,4)benzothiazin-2(3H)-one) ...-2(3H)-one), G-clamps (e.g., substituted phenoxazines), Phenoxazine cytidines (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzoxazin-2(3H)-one), phenothiazine cytidines (1H-pyrimido(5,4-b)(1,4)benzothiazin-2(3H)-one), G-clamps (e.g., substituted phenoxazine cytidines (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzoxazin-2(3H)-one), carbazole cytidines (2H-pyrimido(4,5-b)indol-2-one), pyridoindole cytidines (H-pyrido(3',2':4,5)pyrrolo[2,3-d]pyrimidin-2-one).

[0023] As used herein, the term "sample" may refer to a composition containing a target. Samples suitable for analysis by the disclosed methods, devices, and systems include cells, tissues, organs, or organisms. A cell sample is a composition of cells, e.g., a composition containing a plurality of different cells (e.g., an aqueous composition of a single cell), and the number of cells may vary.

[0024] As used herein, the terms "sampling device" or "device" may refer to a device capable of withdrawing a portion of a sample or depositing a portion of the sample on a substrate. A sample device may refer to, for example, a fluorescence activated cell sorter (FACS), a cell sorter, a biopsy needle, a biopsy device, a tissue sectioning device, a microfluidic device, a blade grid, and / or a microtome.

[0025] As used herein, the term "solid support" may refer to a discrete solid or semi-solid surface to which nucleic acids can be bound. A solid support may include a solid, porous, or hollow sphere, ball, bearing, cylinder, or other similar shape, including plastic, ceramic, metal, or polymeric materials (e.g., hydrogel), to which nucleic acids may be immobilized (e.g., covalently or non-covalently attached). A solid support may include discrete particles having a spherical shape (e.g., microsphere) or a non-spherical or irregular shape (e.g., cube, rectangular prism, pyramidal, cylindrical, conical, ellipsoid, or discoid). Beads may be non-spherical. A plurality of solid supports arranged in an array may not include a substrate. The terms "bead" and "particle" may be used interchangeably.

[0026] As used herein, the term "target" can refer to a composition that can be analyzed according to embodiments of the present invention. Examples of targets suitable for analysis by the disclosed methods, devices, and systems include oligonucleotides, DNA, RNA, mRNA, microRNA, tRNA, etc. Targets can be single-stranded or double-stranded. In some embodiments, targets can be proteins, peptides, or polypeptides. In some embodiments, targets are lipids. As used herein, "target" can be used interchangeably with "species."

[0027] As used herein, the term "reverse transcriptase" may refer to a group of enzymes possessing reverse transcriptase activity (i.e., the activity of catalyzing the synthesis of DNA from an RNA template). Generally, such enzymes include, but are not limited to, retroviral reverse transcriptases, retrotransposon reverse transcriptases, retroplasmid reverse transcriptases, retron reverse transcriptases, bacterial reverse transcriptases, group II intron-derived reverse transcriptases, and mutants, variants, or derivatives thereof. Non-retroviral reverse transcriptases include non-LTR retrotransposon reverse transcriptases, retroplasmid reverse transcriptases, retron reverse transcriptases, and group II intron reverse transcriptases. Examples of group II intron reverse transcriptases include Lactococcus lactis LI.LtrB intron reverse transcriptase, Thermosynechococcus elongatus TeI4c intron reverse transcriptase, or Geobacillus stearothermophilus GsI-IIC intron reverse transcriptase. Other classes of reverse transcriptases include many classes of non-retroviral reverse transcriptases (eg, retrons, group II introns, diversity-generating retroelements, etc.).

[0028] Detailed Description

[0029] Methods for obtaining linked functional and sequence data of single cells (e.g., of a cell sample) are provided. Aspects of the methods include functionally assaying the divided single cells; visually indexing the functionally assayed divided single cells using unique combinations of different nucleic acid-barcoded identification particles; obtaining sequence data of the visually indexed, functionally assayed, and divided single cells; and linking the functional and sequence data of the sequenced, visually indexed, functionally assayed, and divided single cells. Compositions for carrying out the methods of the invention are also provided.

[0030] Before describing the present invention in further detail, it is to be understood that the invention is not limited to particular embodiments described, as such may, of course, vary, and the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting; for this reason, it is to be understood that the scope of the present invention will be limited only by the appended claims.

[0031] Where a range of values ​​is provided, unless the context clearly dictates otherwise, it is understood that every intervening value (to the tenth of the unit of the lower limit) between the upper and lower limits of that range, and every other value or intervening value stated within that range, is included in the invention. The upper and lower limits of these smaller ranges may each independently be included in the smaller ranges and are also included in the invention unless there is an expressly excluded limit in that range. Where the stated range includes one or both limits, ranges excluding either or both of those included limits are also included in the invention.

[0032] In this specification, certain ranges are presented with the term "about" preceding the numerical values. In this specification, the term "about" not only literally supports the exact numerical value of the preceding numerical value, but also encompasses values ​​that are close to or approximately that numerical value. When determining whether a numerical value is close to or approximately a specifically stated numerical value, the unstated numerical value that is close or approximately may be a numerical value that, in the context in which it is presented, provides a value that is substantially equivalent to the specifically stated numerical value.

[0033] Unless otherwise defined, all technical and scientific terms used herein have the meaning commonly understood by one of ordinary skill in the art to which this invention belongs. Although methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present invention, representative exemplary methods and materials are described below.

[0034] All publications and patents cited herein are incorporated by reference as if each individual publication or patent was specifically and individually indicated to be incorporated by reference, and all publications and patents cited herein are incorporated by reference to disclose or describe the methods and / or materials for which the publications are cited. The citation of any publication is to indicate the publication date of the publication prior to the filing date of the present application and should not be construed as an admission that the present invention is not entitled to priority over such publication. Further, the stated publication dates may be different from the actual publication dates, which may require independent confirmation.

[0035] It should be noted that, in this specification and the appended claims, the singular forms "a," "an," and "the" include the plural forms unless the context clearly dictates otherwise. It should also be noted that the claims may be drafted to exclude optional elements. Accordingly, this statement is intended to serve as a predicate for using exclusive terminology, such as "only," "only," or "negative" limitations in the recitation of claim elements.

[0036] It will be apparent to those skilled in the art upon reading this disclosure that the individual embodiments described and illustrated herein have individual components and features which may be readily separated or combined with the features of any of the other embodiments without departing from the scope or spirit of the invention. Any method described can be carried out in the order of events described or in any other order which is logically possible.

[0037] While the systems and methods are described with functional descriptions for the sake of grammatical fluency, it is expressly understood that, unless expressly recited under 35 U.S.C. § 112, the claims are not to be construed as necessarily limited in any sense by the construction of "means" or "step" limitations, and that the claims are to be construed under the doctrine of judicial equivalents to encompass the meaning of the definitions provided by the claims and the full range of equivalents thereof; and, if a claim is expressly recited under 35 U.S.C. § 112, the claim shall be accorded the full range of statutory equivalents under 35 U.S.C. § 112.

[0038] method

[0039] As summarized above, methods are provided for obtaining linked functional and sequence data for a single cell (e.g., a single cell from an initial cell sample). "Linked functional and sequence data" refers to a combined dataset containing both functional data and nucleic acid sequence data that can be attributed to the same cell, such that the two types of data can be considered to originate from the same cell. In other words, linked functional and sequence data refers to a dataset containing both functional and nucleic acid sequence data obtained from the same cell. Functional data refers to data obtained from a cell using a functional assay (i.e., a phenotypic assay), where examples of functional data that may be obtained from a single cell in embodiments of the present invention and that may be linked to sequence data are described in further detail below. Nucleic acid sequence data refers to data obtained using nucleic acid sequencing techniques that identify the sequence of nucleotides within a nucleic acid molecule. Nucleic acid sequence data from a cell includes the sequence of one or more nucleic acid sequences (e.g., RNA molecules) present in the cell. Such data may include gene expression data. Such data may also include protein expression data (e.g., obtained using AbSeq). In some cases, the sequence data may be multi-omics data, which may be obtained using a variety of sequencing protocols, including next-generation sequencing (NGS) protocols.

[0040] As summarized above, aspects of the method include: (a) functionally assaying the divided single cells; (b) visually indexing the functionally assayed divided single cells using unique combinations of different nucleic acid-barcoded identification particles; (c) obtaining sequence data of the visually indexed, functionally assayed, divided single cells; and (d) linking the functional and sequence data of the sequenced, visually indexed, functionally assayed, and divided single cells. Embodiments of each of these steps are described in detail below.

[0041] Functional assays of divided single cells

[0042] In embodiments of the present invention, method aspects include functional assays of divided single cells. "Functional assays of divided single cells" refers to functionally assaying single cells that have been divided from one another, including, for example, assaying one or more phenotypic characteristics. In embodiments, single cells are provided in a compartment that is fluidly isolated from other compartments and functionally assayed within the compartment. Embodiments of functional assays of divided single cells include: (a) contacting cells of a cell sample with nucleic acid-barcoded cell-bound beads to generate bead-bound cells; (b) dividing the bead-bound cells to generate divided bead-bound single cells; and (c) functionally assaying the divided bead-bound single cells to obtain functional data for the divided bead-bound single cells.

[0043] Cell samples

[0044] In embodiments of the present invention, the single cell functionally assayed may be a cell originally present in a cell sample. The number of cells in a given cell sample may vary, and in some cases, the number of cells may range from 50 to 50,000,000, including, for example, 100 to 1,000,000 or 500 to 100,000. The cells present in a given cell sample may be of any type, including prokaryotic and eukaryotic cells. Suitable prokaryotic cells include, but are not limited to, E. coli, various Bacillus species, and extremophilic microorganisms such as thermophilic bacteria. Suitable eukaryotic cells include, but are not limited to, fungi such as yeast and filamentous fungi (including Aspergillus, Trichoderma, Neurospora, etc.), plant cells (including corn, sorghum, tobacco, canola, soybean, cotton, tomato, potato, alfalfa, sunflower, etc.), and animal cells (fish, bird, mammal, etc.). Suitable fish cells include, but are not limited to, cells from species such as salmon, trout, tilapia, tuna, carp, flounder, halibut, swordfish, cod, and zebrafish. Suitable bird cells include, but are not limited to, cells from chicken, duck, pheasant, pheasant, turkey, and other jungle and game birds. Suitable mammalian cells include, but are not limited to, cells from horse, cow, buffalo, deer, sheep, rabbit, mouse, rat, hamster, and other rodents such as guinea pig, goat, pig, primate, and marine mammals (including dolphins and whales) and cell lines (including cell lines of all human tissues and stem cell types), as well as stem cells (including pluripotent and non-pluripotent) and non-human fertilized eggs. Suitable cells also include cell types implicated in various disease symptoms, even when not in a disease state.Thus, suitable eukaryotic cell types include, but are not limited to, tumor cells of all types (e.g., melanoma, myeloid leukemia, lung, breast, ovarian, colon, kidney, prostate, pancreatic, and testicular carcinomas), cardiomyocytes, dendritic cells, endothelial cells, epithelial cells, lymphocytes (T cells and B cells), mast cells, eosinophils, vascular endothelial cells, macrophages, natural killer cells, erythrocytes, hepatocytes, leukocytes (including monocytes), stem cells (e.g., hematopoietic stem cells, neural stem cells, skin stem cells, lung stem cells, kidney stem cells, liver stem cells, and muscle stem cells) (for use in screening for differentiation and dedifferentiation factors), osteoclasts, chondrocytes and other connective tissue cells, keratinocytes, melanocytes, hepatocytes, renal cells, and adipocytes. In certain embodiments, the cells are primary disease-state cells, e.g., primary tumor cells. Suitable cells also include, but are not limited to, known research cells (Jurkat T cells, NIH3T3 cells, CHO, COS, etc.). See the ATCC cell line catalog, expressly incorporated herein by reference.

[0045] In certain embodiments, the cells used in the present invention are obtained from a subject. As used herein, "subject" refers to humans and other animals, such as laboratory animals, as well as other organisms. Thus, the methods and compositions described herein are applicable in both human and veterinary applications. In certain embodiments, the subject is a mammal, including embodiments in which the subject is a human patient with (or suspected of having) a disease or condition.

[0046] In certain embodiments, cells of interest are enriched (e.g., as described in more detail below) prior to indexing. For example, if the cells of interest are white blood cells from a human subject, the subject's whole blood may be subjected to density gradient centrifugation to enrich for peripheral blood mononuclear cells (PBMCs, or white blood cells). Cells may be enriched using any suitable method known in the art, including fluorescence-activated cell sorting (FACS), magnetic-activated cell sorting (MACS), density gradient centrifugation, etc. Parameters used to enrich for specific cells from a mixed population include, but are not limited to, physical parameters (e.g., size, shape, density, etc.), in vitro growth characteristics (e.g., response to specific nutrients in cell culture), and molecular expression (e.g., expression of cell surface proteins or carbohydrates, reporter molecules, e.g., green fluorescent protein, etc.).

[0047] In certain embodiments, the cells are viable cells that maintain viability over the course of the assay. "Maintaining viability" means that a certain percentage of the cells remain viable at the end of the assay, including from about 20% viability to about 100% viability. In other specific embodiments, the methods of the invention are performed in a manner that renders the cells nonviable over the course of the assay; for example, the cells may be fixed, permeabilized, or maintained in buffers or conditions that render the cells nonviable. Such parameters will generally be determined by the nature of the assay being performed and the reagents used.

[0048] In some cases, the cells may be treated with, for example, a stimulus. The stimuli with which the cells can be treated may vary, such as culture conditions, exposure to temperature changes (e.g., heat or cold), exposure to electromagnetic radiation (e.g., light), exposure to an activating agent, exposure to mechanical changes, etc. If desired, different samples of the plurality of cell samples can be treated with the same or different stimuli. Thus, in some cases, the method includes treating two or more of the plurality of cell samples differently, for example, contacting the two or more different samples with different active substances or contacting them with different concentrations of the same active substance.

[0049] Contacting the cells with nucleic acid-barcoded cell-binding beads

[0050] In embodiments, the method includes contacting cells of the cell sample with nucleic acid-barcoded cell-binding beads to generate bead-bound cells. The nucleic acid-barcoded cell-binding beads contacted with the cells of the cell sample can vary, including, for example, beads, specific binding moieties, and cell-binding bead nucleic acids comprising barcodes.

[0051] The bead component of nucleic acid-barcoded cell-binding beads can vary as needed and can be any solid support capable of binding nucleic acids, such as, for example, a discrete solid or semi-solid surface. Beads can be solid, porous, or hollow spheres, balls, bearings, cylinders, or other similar shapes, including plastic, ceramic, metal, or polymeric materials (e.g., hydrogels), to which nucleic acids can be immobilized (e.g., covalently or non-covalently). Beads can include discrete particles that can have a spherical shape (e.g., microspheres) or a non-spherical or irregular shape (e.g., cube, rectangular, pyramidal, cylindrical, conical, ellipsoid, or discoid). Beads can also be non-spherical. In some cases, beads can be magnetic. Further details regarding beads that can be used in embodiments of the present invention are described in U.S. Patent Application Publication No. US2018 / 0088112; U.S. Patent Application Publication No. 2018 / 0200710; U.S. Patent Application Publication No. US2018 / 0346970; U.S. Patent Application Publication No. 2019 / 0056415; U.S. Patent Application Publication No. US2020 / 0248263; U.S. Patent Application Publication No. 2020 / 0299672; and U.S. Patent Application Publication No. 2021 / 0171940, the disclosures of which are incorporated herein by reference.

[0052] The beads can display specific binding moieties on their surfaces. The specific binding moiety components of the nucleic acid-barcoded cell-binding beads employed in embodiments of the present invention may vary. The term "specific binding" refers to the direct binding of two molecules, for example, through covalent, electrostatic, hydrophobic, ionic, and / or hydrogen-bonding interactions (including interactions such as salt bridges and water bridges). "Specific binding moieties" refer to components of a molecular pair that have binding specificity for each other. Components of a specific binding pair may be naturally occurring or synthetically produced in whole or in part. One component of the molecular pair has a region or cavity on its surface that specifically binds to a particular spatial and polar organization of the other component of the molecular pair, and thus is complementary to each other. Thus, the components of the pair possess the property of specifically binding to each other. Examples of specific binding moiety pairs include antigen-antibody, biotin-avidin, hormone-hormone receptor, receptor-ligand, enzyme-substrate, etc. Specific binding moieties of a binding pair exhibit high affinity and binding specificity when binding to each other. Typically, the affinity between the specific binding moieties of a pair is greater than or equal to K d (dissociation constant) is 10 -6 It is expressed in M ​​or less, for example, 10 -7 M or less, e.g. 10 -8 M or less, 10 -9 M or less, 10 -10 M or less, 10 -11 M or less, 10 -12 M or less, 10 -13 M or less, 10 -14 M or less, 10 -15M or less. Affinity refers to the strength of binding, with higher binding affinity corresponding to a lower KD. In one embodiment, affinity is determined by surface plasmon resonance (SPR), for example, as used in the Biacore system. The affinity of one molecule for another molecule is determined by the binding kinetics of the interaction, measured, for example, at 25°C. Affinity refers to the strength of binding, with higher binding affinity corresponding to a lower KD. In one embodiment, affinity is determined by surface plasmon resonance (SPR), for example, as used in the Biacore system. The affinity of one molecule for another molecule is determined by the binding kinetics of the interaction, measured, for example, at 25°C. Specific binding moieties can vary; examples of specific binding moieties include, but are not limited to, polypeptides, nucleic acids, carbohydrates, lipids, peptoids, etc. In some cases, the specific binding moiety is proteinaceous. As used herein, the term "proteinaceous" refers to a group composed of amino acid residues. The proteinaceous group may be a polypeptide. In certain cases, the proteinaceous specific binding moiety is an antibody. In certain embodiments, the proteinaceous specific binding moiety is an antibody fragment, e.g., a binding fragment of an antibody that specifically binds to an epitope, such as a cell surface protein. As used herein, the terms "antibody" and "antibody molecule" are used interchangeably to refer to a protein consisting of one or more polypeptides substantially encoded by some or all of the recognized immunoglobulin genes. Recognized immunoglobulin genes include, for example, in humans, the kappa (k), lambda (l), and heavy chain loci, which together constitute a diverse group of variable region genes, and the mu (u), delta (d), gamma (g), sigma (e), and alpha (a) constant region genes, which encode the IgM, IgD, IgG, IgE, and IgA isotypes, respectively. The variable region of an immunoglobulin light or heavy chain is composed of a "framework" region (FR) interspersed with three hypervariable regions (also called complementarity-determining regions or CDRs).The sizes of framework and CDR regions are precisely defined (see "Sequences of Proteins of Immunological Interest," E. Kabat et al., US Department of Health and Human Services, (1991)). All antibody amino acid sequence numbering discussed herein conforms to the Kabat system. The sequences of framework regions of different light or heavy chains are relatively conserved within a species. The framework region of an antibody is the combined framework region of the constituent light and heavy chains and contributes to the positioning and arrangement of the CDRs. The CDRs are primarily responsible for binding to an epitope of an antigen. The term "antibody" refers to a full-length antibody and may refer to natural antibodies from any organism, modified antibodies, or antibodies engineered for experimental, therapeutic, or other purposes, as described in more detail below. Antibody fragments of interest include, but are not limited to, Fab, Fab', F(ab')2, Fv, scFv, or antigen-binding subsequences of antibodies, including those generated by modification of whole antibodies or synthesized de novo using recombinant DNA technology. Antibodies may be monoclonal or polyclonal and may have other specific activities against cells (e.g., antagonist, agonist, neutralizing, inhibitory, or stimulatory). It is understood that antibodies may have additional conservative amino acid substitutions that do not substantially affect antigen binding or other antibody functions. In certain embodiments, the specific binding component is a Fab fragment, F(ab')2 fragment, scFv, dibody, or triabody. In certain embodiments, the specific binding component is an antibody. In some cases, the specific binding component is a murine antibody or binding fragment thereof. In certain instances, the specific binding component is a recombinant antibody or binding fragment thereof.

[0053] The specific binding moiety component of the nucleic acid-barcoded cell-binding bead may specifically bind to any convenient cell marker. In some cases, the specific binding moiety binds to a cell surface marker, where the cell surface marker of interest includes, but is not limited to, a ubiquitous cell surface marker, i.e., a cell surface marker that is at least predicted to be present on all cells of a given cell sample processed in a given workflow according to the present invention. Examples of ubiquitous cell surface markers to which the specific binding moiety / oligonucleotide sub-barcode can specifically bind include, but are not limited to, CD44, CD45, beta-2 microglobulin, etc.

[0054] Nucleic acid-barcoded cell-binding beads employed in embodiments of the present invention also include cell-binding bead nucleic acids that include a barcode component. The barcode component may vary in length, for example, from 10 to 500 nt, e.g., from 15 to 100 nt. In some cases, the barcode component may be composed of ribonucleic acid or deoxyribonucleic acid, as needed. The barcode component of embodiments of the present invention may include a barcode region or domain as well as other domains that find use in embodiments of the present invention, where such domains may include a bead identifier domain, a capture sequence, a primer binding site, a domain complementary to a domain in particle nucleic acid identification, etc. The barcode component may be covalently attached to the bead directly or via an intermediate (e.g., a disulfide bond), as needed. In certain embodiments, the cell-binding bead nucleic acid may be attached to the bead via a cleavable linker that can be cleaved by cell lysis conditions, examples of which include, but are not limited to, disulfide linkers.

[0055] The barcode region (i.e., barcode domain) of a barcode component is a domain or subsequence, i.e., stretch, of the barcode component that serves as an identifier for the bead to which it is attached. The sequence of a given barcode region can be used as an identifier for the bead to which that barcode is bound, thereby allowing an amplicon bearing the barcode to be assigned as originating from a nucleic acid-barcoded cell-bound bead. Thus, the sequence of the barcode region corresponds to the bead to which it is attached. The barcode region may have any convenient sequence and may vary in length. The barcode may be, for example, a nucleotide sequence of any suitable length, e.g., from about 4 nucleotides to about 200 nucleotides in length. In some embodiments, the barcode region is a nucleotide sequence that is 25 nucleotides to about 45 nucleotides in length. In some embodiments, the length of the barcode region can be less than, more than, or equal to 4 nucleotides, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 15 nucleotides, 20 nucleotides, 25 nucleotides, 30 nucleotides, 35 nucleotides, 40 nucleotides, 45 nucleotides, 50 nucleotides, 55 nucleotides, 60 nucleotides, 70 nucleotides, 80 nucleotides, 90 nucleotides, 100 nucleotides, 200 nucleotides, or a range between any two of the above values.

[0056] The barcode component may include a first domain complementary to the target binding region of the nucleic acid capture bead, which may be referred to as a capture sequence. The capture sequence is a region or regions that function as a binding site for the target binding region (e.g., of the bead-bound nucleic acid of the cell capture bead), e.g., as described below. The intended capture sequence may vary as needed and may be specific, random, or semi-random. In some cases, the capture sequence is a sequence that hybridizes with the target binding region of the bead-bound nucleic acid of the cell capture bead (e.g., as described in detail below). In some cases, the capture sequence is a poly(A) sequence, which is configured to hybridize with an oligo-dT target binding region, e.g., as described in detail below. In such cases, the length of the poly(A) capture sequence may vary and in some cases may range from 3 to 50 nt, e.g., 5 to 25 nt. If present, the capture sequence may be located 3' to the barcode domain. In some cases, the capture sequence is located at the 3' end of the cell-binding bead nucleic acid.

[0057] In some cases, the barcode component further includes a second domain complementary to a sequence present in the particle identification nucleic acid in the nucleic acid-barcoded identification particle (e.g., as described in detail below). The sequence of this second domain is selected to hybridize with a sequence present in the identification particle nucleic acid and may have any convenient sequence. The length of the sequence of this second domain may vary and in some cases may range from 3 to 50, e.g., 5 to 25 nt. If present, the second domain may optionally be located 3' of the barcode domain. In some cases, this second domain may be located 3' of the capture sequence (i.e., the first domain), where in some cases, this second domain is present at the 3' end of the barcode component. In other cases, this second domain is located 3' of the barcode domain but 5' of the capture sequence (first domain).

[0058] The barcode component may further include a primer binding site. If present, the primer binding site may be configured to bind to a primer used, for example, in preparing a sequenceable nucleic acid. For example, the barcode component may include a primer binding site common to all nucleic acid-barcoded cell-bound beads used in a given workflow. This primer binding site may be different from other primer binding sites (e.g., a universal primer binding site, described in detail below) used, for example, in preparing a sequenceable library. In some cases, the primer binding site may be, or may be approximately, the following number: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or a number or range of lengths between any two of these nucleotides. The length of the primer binding site may vary and may be at least or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides in length. The primer binding site may be located at the 5' end of the barcode component. In some cases, the primer binding site may be the same as the primer binding site present in an oligonucleotide-labeled cellular component binding reagent, such as, for example, the primer binding site found in AbSeq oligonucleotide-labeled antibodies, as described in more detail below.

[0059] FIG. 1A shows a schematic diagram of a nucleic acid-barcoded cell-binding bead usable in embodiments of the present invention. As shown in FIG. 1, the nucleic acid-barcoded cell-binding bead 100 includes a magnetic bead 110 to which an antibody 120 specific for a cell surface marker (e.g., CD45) is bound. From the 5' end to the 3' end, the bead also includes a primer binding site 132, a barcode region 134, a first domain (i.e., a capture sequence) 136, and a second domain 138 that is complementary to a sequence present in the identification particle nucleic acid of the nucleic acid-barcoded identification particle (e.g., as described in detail below). As shown in the figure, the barcode component 130 can be covalently attached to the bead directly or via an intermediate (e.g., a disulfide bond), as desired. FIG. 1B shows a schematic diagram of an alternative embodiment of a nucleic acid-barcoded cell-binding bead usable in embodiments of the present invention. As shown in FIG. 1B, the barcode component has a capture sequence located at the 3' end and a second domain positioned between the capture sequence and the barcode domain.

[0060] Cells of a cell sample can be contacted with nucleic acid-barcoded cell-binding beads using any suitable protocol to generate bead-bound cells. The cells can be contacted with the beads under conditions sufficient to allow specific binding moieties (e.g., antibodies) displayed on the surface of the beads to specifically bind to corresponding epitopes present on cells of the cell sample, generating a composition of bead-bound cells. The number of beads bound to cells in a given bead-bound cell can, in some cases, range from 1 to 5, e.g., 1 to 2. Among the bead-bound cells, the barcode components can optionally share common domains, such as primer binding sites, first and / or second domains.

[0061] Splitting of bead-bound cells

[0062] After producing the bead-bound cells, e.g., as described above, method embodiments include splitting the bead-bound cells to generate split bead-bound single cells, where each split bead-bound cell is stably bound to one or more nucleic acid-barcoded cell-bound beads, e.g., via specific binding of a specific binding moiety (e.g., an antibody) on the bead to an epitope (e.g., an epitope of a cell surface protein of the cell). Thus, after generating the bead-bound cells, e.g., as described above, method embodiments include splitting the bead-bound cells to generate split bead-bound single cells. In some cases, splitting includes distributing the bead-bound cells into partitions or compartments such that the compartment contains a single bead-bound cell (i.e., the compartment contains only one bead-bound cell). By "splitting" is meant placing the bead-bound cells into small reaction chambers, where the reaction chambers may be fluidically isolated structures defined by a solid material (e.g., microwells) and configured to accommodate the bead-bound cells. Some embodiments of the disclosed methods, devices, and systems use a plurality of microwells randomly distributed on a substrate. In some embodiments, the plurality of microwells is distributed on the substrate in an ordered pattern, e.g., an ordered array. In some embodiments, the plurality of microwells is distributed on the substrate in a random pattern (e.g., a random array). The microwells may be configured in a variety of shapes and sizes. Suitable well shapes include, but are not limited to, cylindrical, elliptical, cubic, conical, hemispherical, rectangular, or polyhedral (e.g., a three-dimensional shape consisting of multiple planar faces), such as a rectangular prism, hexagonal prism, octagonal prism, inverted triangular pyramid, inverted square pyramid, inverted pentagonal pyramid, inverted hexagonal pyramid, or inverted truncated pyramid. In some embodiments, non-cylindrical microwells (e.g., wells having an elliptical or square footprint) may offer advantages in that they can accommodate larger cells.In some embodiments, the top and / or bottom edges of the well walls may be rounded to avoid sharp angles and reduce electrostatic forces that may arise due to electrostatic field concentration at sharp angles or sharp points. Thus, rounding the corners may improve the ability to retrieve beads from the microwells. Microwell dimensions may be characterized in absolute terms. In some cases, the average microwell diameter may range from about 5 μm to about 100 μm. In other embodiments, the average microwell diameter is at least 5 μm, at least 10 μm, at least 15 μm, at least 20 μm, at least 25 μm, at least 30 μm, at least 35 μm, at least 40 μm, at least 45 μm, at least 50 μm, at least 60 μm, at least 70 μm, at least 80 μm, at least 90 μm, or at least 100 μm. In still other embodiments, the average microwell diameter is at most 100 μm, at most 90 μm, at most 80 μm, at most 70 μm, at most 60 μm, at most 50 μm, at most 45 μm, at most 40 μm, at most 35 μm, at most 30 μm, at most 25 μm, at most 20 μm, at most 15 μm, at most 10 μm, or at most 5 μm. The microwell volume used in the methods of the invention is, in some cases, about 200 μm. 3 to approximately 800,000 μm 3 In some embodiments, the microwell volume varies in a range of at least 200 μm 3 , at least 500 μm 3 , at least 1,000 μm 3 , at least 10,000 μm 3 , at least 25,000 μm 3 , at least 50,000 μm 3 , at least 100,000 μm 3 , at least 200,000 μm 3 , at least 300,000 μm 3 , at least 400,000 μm 3 , at least 500,000 μm 3 , at least 600,000 μm 3 , at least 700,000 μm 3 , or at least 800,000 μm3 In other embodiments, the microwell volume is at most 800,000 μm 3 , up to 700,000 μm 3 , up to 600,000 μm 3 , 500,000 μm 3 , up to 400,000 μm 3 , up to 300,000 μm 3 , up to 200,000 μm 3 , up to 100,000 μm 3 , up to 50,000μm 3 , up to 25,000 μm 3 , up to 10,000 μm 3 , up to 1,000 μm 3 , up to 500 μm 3 , or a maximum of 200 μm 3The number of microwells in a given device used in embodiments of the invention may vary, and in some cases may include 100 or more, e.g., 250 or more, e.g., 500 or more, 1,000 or more, e.g., 5,000 or more, e.g., 10,000 or more, and in some cases, the number may be 15,000 or less, e.g., 12,500 or less. Suitable microwells that can be used in embodiments of the invention are described in further detail in PCT Application No. PCT / US2016 / 014612 (published as WO / 2016 / 118915), the disclosure of which is incorporated herein by reference. As used herein, a substrate can refer to a type of solid support. A substrate can, for example, include multiple microwells. For example, a substrate can be a well array including two or more microwells. In some embodiments, a microwell can include a small reaction chamber having a defined volume. In some embodiments, a microwell can capture one or more cells. In some embodiments, a microwell can capture only one cell. In some embodiments, the microwells can capture one or more solid supports. The number of wells (e.g., microwells) in a well plate (e.g., a microwell array) can vary in a given division step, and in some cases the number can be 100 or more, such as 250 or more, for example 500 or more, 1000 or more, including 5,000 or more, for example 10,000 or more, 100,000 or more, 250,000 or more, and in some cases the number can be 500,000 or less, for example 400,000 or less, and in some cases 15,000 or less, for example 12,500 or less.

[0063] When dividing the bead-bound cells, the bead-bound cells may be placed into compartments, such as, for example, microwells of a microwell array, using any suitable protocol. The present disclosure provides methods for compartmentalizing bead-bound cells into partitions for dividing the bead-bound cells. For example, a collection of bead-bound cells can be introduced into a structure, such as, for example, a microwell, to divide the bead-bound cells. The bead-bound cells can be contacted with the compartments, for example, by gravity flow, in which case the bead-bound cells can settle within the dividing structure. In some cases, an aqueous composition of bead-bound cells can be contacted with the array of microwells, for example, by flowing the composition through the array of microwells, such that the bead-bound cells settle into the microwells. The aqueous composition containing the bead-bound cells can be passed through a flow cell in fluid communication with the microwells. Suitable protocols and systems for dividing bead-bound cells into microwells are described in PCT Application No. PCT / US2016 / 014612 (published as WO / 2016 / 118915), the disclosure of which is incorporated herein by reference. Any suitable protocol can be employed to divide the bead-bound cells. For example, a fraction of the bead-bound cells may be distributed into compartments, such as by pipetting, or the sample may be flowed over the surface of a well plate.

[0064] Functional assay of split bead-bound single cells

[0065] After producing the split bead-bound single cells, for example as described above, method embodiments include functionally assaying the split bead-bound single cells to obtain functional data for the split bead-bound single cells. The functional assays that can be performed on the split bead-bound single cells can vary. In some cases, functional assays of split bead-bound single cells to obtain functional data for the split bead-bound single cells include evaluating the split bead-bound single cells over a time course. This time period can vary and in some cases can range from 1 second to 24 hours or more, including 1 second to 12 hours, 30 seconds to 6 hours, etc. In some cases, functional assays of split bead-bound single cells to obtain functional data for the split bead-bound single cells include evaluating the split bead-bound single cells in response to a stimulus. The stimulus used in such cases can vary, and examples of stimuli include, but are not limited to, chemical stimuli, mechanical stimuli, physical stimuli, or combinations thereof. The evaluation may detect a change arising from any number of sources, and may include, for example, the generation of a signal from a signal generating system (e.g., a reagent reporter system), a change in morphology, and the like.

[0066] Functional assays that can be performed on divided, bead-bound single cells include, but are not limited to: targeted cell lysis assays; cell chemotaxy assays; single-cell secretome assays (including real-time single-cell secretome assays), CAR-T cell evaluation and characterization assays, single-cell phenotypic characterization (organelle) assays (e.g., lysosomal activation, endocytosis, phagocytosis, autophagy, calcium signaling, Akt, NFkB translocation, and reporter assays); assays to evaluate the effects of small molecules (such as novel therapeutics) (e.g., evaluation in the tumor microenvironment); single-cell antibody production assays; dendritic cell maturation, macrophage differentiation, cell differentiation (morphological change) assays; neural differentiation assays; and virus production and regulation assays.

[0067] Functional assays performed on the partitioned, bead-bound single cells generate functional data for those cells, which may be recorded as originating from a particular partition and thus from the cells present within that partition, and may be linked to sequence data obtained for the cells within that partition using a visual index obtained for that partition, as described in more detail below.

[0068] Visual indexing of functionally assayed and divided single cells

[0069] After functional data for the partitioned, bead-bound single cells is obtained, e.g., as described above, method embodiments include visually indexing the functionally assayed, partitioned single cells. "Visually indexing" refers to acquiring indexing image data for each partition, where this indexing image data may be used to correlate functional data acquired from cells in a given partition with sequence data acquired from cells in that same partition (e.g., as described in more detail below). Visual indexing may include acquiring image data (i.e., indexing image data) of partitions containing functionally assayed signals, and may particularly include image data of unique combinations of identification particles present in the partition. In some cases, the indexing image data is acquired using unique combinations of different nucleic acid-barcoded identification particles. In such cases, the method includes introducing unique combinations of different nucleic acid-barcoded identification particles into partitions containing functionally assayed, bead-bound single cells to generate indexing partitions, where the indexing partitions contain unique combinations of functionally assayed, bead-bound single cells, and nucleic acid-barcoded identification particles. Image data of the indexed partitions is then acquired, and the unique combinations of nucleic acid-barcoded identification particles of the indexed partitions are identified.

[0070] Nucleic acid-barcoded identification particles

[0071] As can be seen, embodiments of the present invention include indexing partitions using unique combinations of nucleic acid-barcoded identification particles. Nucleic acid-barcoded identification particles (which may also be referred to as Pheno-Seq particles) are solid supports (e.g., beads) with a known size and color signature (e.g., color (i.e., hue) and brightness (e.g., due to fluorescence from one or more fluorescent dyes incorporated into the particle). Any type of solid support, such as those described above, may be used as a nucleic acid-barcoded identification particle. Thus, particles may include any type of solid, porous, or hollow sphere, ball, bearing, cylinder, or other similar construct, which may be composed of plastic, ceramic, metal, or polymeric material (e.g., hydrogel), to which a nucleic acid is immobilized (e.g., The solid support may be tethered to a solid support, e.g., covalently or non-covalently bound, and may incorporate a color-imparting agent (e.g., one or more fluorophores). The solid support may comprise discrete particles having a spherical shape (e.g., microspheres) or a non-spherical or irregular shape (e.g., cubes, cuboids, pyramidal, cylindrical, conical, ellipsoidal, or discoidal). The size of the nucleic acid-barcoded identification particles may vary, but in some cases, the size of a given particle is in the range of 1 to 50 micrometers, including 2 to 25 micrometers, 3 to 20 micrometers, etc. In some cases, the size of the particle is 3, 7, 10, or 16 micrometers.

[0072] As noted above, the identification particles used in embodiments of the present invention include a color signature (a collection of one or more hues and / or intensities thereof). In embodiments, the identification particles include one or more color-imparting agents (e.g., pigments, phosphors, etc.), where the color-imparting agent(s) and the amounts thereof constitute the color signature of the particle.

[0073] The color signature of a given identification particle in embodiments of the present invention is provided by one or more fluorescent dyes. When a given color signature is provided by multiple fluorescent dyes, the two or more fluorescent dyes collectively constitute the color signature of the identification particle. Thus, in embodiments of the present invention, a given color signature may be composed of a single fluorescent dye, or may be composed of two or more fluorescent dyes (e.g., 2 to 5, e.g., 2 to 4, including 2 to 3), which collectively constitute the color signature of the bead. Thus, the number of different fluorescent dyes constituting a given color signature may vary, and in some cases may range from 1 to 5, e.g., 1 to 3, including 1 to 2. Any two given distinguishable color signatures may be distinguishable from each other based on the type of fluorescent dyes and / or the brightness of the signals they provide. Thus, any two distinguishable color signatures of different identification particles may be distinguishable based on the fluorescent signals (e.g., emission wavelength maxima) and / or their intensities, or the total amount, of the fluorescent dyes collectively constituting the color signature. For example, two distinct color signatures of two different identification particles may be distinguishable from one another because they are composed of different combinations of fluorescent dyes (e.g., one contains fluorescent dye a and the other contains fluorescent dye b). Two distinct color signatures may also be distinguishable from one another because of different amounts of fluorescent dyes, e.g., one signature may be composed of fluorescent dye a present in a first amount in a given identification particle, while the other signature is composed of fluorescent dye a present in a second amount that differs from the first amount by a detectable amount (e.g., a difference in signal brightness). Any desired number of unique color signatures can be provided using combinations of fluorescent dye types and amounts.

[0074] The signature optionally comprises one or more fluorescent dyes. Thus, an identification particle may comprise a single type of fluorescent dye. Alternatively, a given identification particle may comprise two or more different types of fluorescent dyes. Examples of fluorescent dyes that may be present on an identification particle include, but are not limited to, acridine and its derivatives (e.g., acridine, acridine orange, acridine yellow, acridine red, acridine isothiocyanate, etc.); 5-(2'-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS); 4-amino-N-[3-vinylsulfonyl)phenyl]naphthalimide-3,5 disulfonic acid (luciferin yellow VS); N-(4-amino-1-naphthyl ) maleimide; anthranilamide; brilliant yellow; coumarin and its derivatives (e.g., coumarin, 7-amino-4-methylcoumarin (AMC, coumarin 120), 7-amino-4-trifluoromethylcoumarin (coumaran 151), etc.); cyanine glycol and its derivatives (e.g., cyanosine, Cy3, Cy5, Cy5.5, Cy7); 4',6-diaminodino-2-phenylindole (DAPI); 5',5''-dibromopyrogallolsulfonephetalein (bromo Pyrogallol Red; 7-Diethylamino-3-(4'-isothiocyanatophenyl)-4-methylcoumarin; Diethylaminocoumarin; Diethylenetriaminepentaacetate; 4,4'-Diisothiocyanatodihydro-stilbene-2,2'-disulfonic acid; 4,4'-Diisothiocyanatostilbene-2,2'-disulfonic acid; 5-[Dimethylamino]naphthalene-1-sulfonyl chloride (DNS, Dansul chloride); 4-(4'-Dimethylaminophenylazo)ammonium chloride benzoic acid (DABCYL); 4-dimethylaminophenylazophenyl-4'-isothiocyanate (DABITC); eosin and its derivatives (e.g., eosin and eosin isothiocyanate); erythrosin and its derivatives (e.g., erythrosin B and erythrosine isothiocyanate); ethidium; fluorescein and its derivatives (e.g., 5-carboxyfluorescein (FAM), 5-(4,6-dichlorotriazin-2-yl)aminofluorescein (DTAF));2'7'-Dimethoxy-4'5'-dichloro-6-carboxyfluorescein (JOE), fluorescein isothiocyanate (FITC), fluorescein chlorotriazinyl, naphthafluorescein, and QFITC ​​(XRITC); fluorescamine; IR144; IR1446; Lissamine™; Lissamine rhodamine, luciferin yellow; malachite green isothiocyanate; 4-methylumberferon; orthocresolphthalein; nitrotyrosine; pararosaniline; Nile red; Oregon green; phenol red; B-phycoerythrin; o-phthaldialdehyde; pyrene and its derivatives (pyrene, pyrene butyrate, sucrinimide-1-pyrene butyrate, etc.); Reactive red 4 (Cibacron™) Brilliant Red 3B-A; rhodamine and its derivatives (e.g., 6-carboxy-X-rhodamine (ROX), 6-carboxyrhodamine (R6G), 4,7-dichlororhodamine Lissamine, rhodamine B sulfonate, rhodamine (Rhod), rhodamine B, rhodamine 123, rhodamine X isothiocyanate, sulforhodamine B, sulforhodomin 101, sulfonate derivative of sulforhodomin 101 (Texas Red), N,N,N',N'-tetramethyl-6-carboxyrhodamine (TAMRA), tetramethylrhodomine, and tetramethylrhodomine isothiocyanate (TRITC); riboflavin; rosolic acid and terbium chelate derivatives; xanthenes; Alexa Fluor dyes (e.g., Alexa Fluor 350, Alexa Fluor 430, Alexa Fluor 488, Alexa Fluor 546, Alexa Fluor 600). 555, Alexa Fluor 568, Alexa Fluor 594, Alexa Fluor 633, Alexa Fluor 647, Alexa Fluor 660, Alexa Fluor 680, Alexa Fluor 700, Alexa Fluor 750), Pacific Blue, Pacific Orange, Cascade Blue, Cascade Yellow; Quantum Dot dyes (Quantum Dot Corporation);Dylight dyes (including Dylight 800, Dylight 680, Dylight 649, Dylight 633, Dylight 549, Dylight 488, Dylight 405) from Pierce (Rockford, Illinois) or combinations thereof. Other fluorescent dyes or combinations thereof known to those skilled in the art may also be used, including those available from Molecular Probes (Eugene, Oregon) and Exciton (Dayton, Ohio);

[0075] In some instances, the fluorescent dye is a polymeric dye (e.g., a fluorescent polymeric dye). The fluorescent polymeric dyes that find use in the present methods are diverse. In some instances of the methods, the polymeric dye comprises a conjugated polymer. Conjugated polymers (CPs) are characterized by a delocalized electronic structure, including a backbone with alternating unsaturated bonds (e.g., double and / or triple bonds) and saturated bonds (e.g., single bonds), allowing π electrons to move between bonds. Thus, the conjugated backbone may impart an extended linear structure to the polymeric dye, and the bond angles between the repeating units of the polymer may be restricted. For example, proteins and nucleic acids, although also polymers, in some cases do not form extended rod-like structures but fold into higher-dimensional three-dimensional shapes. Furthermore, CPs may form a "rigid rod" polymeric backbone, and the twist (e.g., torsion) angles between the repeating monomeric units along the polymer backbone chain may be restricted. In some instances, CPs with a rigid rod-like structure are included in the polymeric dye. The structural features of the polymer dye can affect the fluorescent properties of the molecule.

[0076] Any suitable polymer dye may be used. In some cases, the polymer dye is a multichromophore, a structure that imparts the ability to absorb light to amplify the fluorescent output of a fluorophore. In some cases, the polymer dye has the ability to absorb light and efficiently convert it to emitted light at a longer wavelength. In some cases, the polymer dye has a light-absorbing multichromophore that can efficiently transfer energy to a nearby fluorescent species (e.g., a "signal chromophore"). Energy transfer mechanisms include, for example, resonance energy transfer (e.g., Förster (or fluorescence) resonance energy transfer, FRET), quantum charge exchange (Dexter energy transfer), etc. In some cases, these energy transfer mechanisms are relatively short-range, i.e., the proximity of the light-absorbing multichromophore system and the signal chromophore allows for efficient energy transfer. Under conditions of efficient energy transfer, a large number of individual chromophores in a light-absorbing multi-chromophore system will result in amplification of the emission from the signal chromophore, i.e., when the incident light ("excitation light") is at a wavelength that is absorbed by the light-absorbing multi-chromophore system, the emission from the signal chromophore will be stronger than if the signal chromophore were directly excited by the pump light.

[0077] Multichromophores can also be conjugated polymers. Conjugated polymers (CPs) are characterized by delocalized electronic structures and can be used as highly sensitive optical reporters for chemical and biological targets. Because the effective conjugation length is significantly shorter than the length of the polymer chain, the backbone contains many closely spaced conjugated segments. Therefore, conjugated polymers are efficient at absorbing light, allowing optical amplification via Förster energy transfer.

[0078] Polymeric dyes of interest include, but are not limited to, those dyes described in U.S. Patent Nos. 7,270,956, 7,629,448, 8,158,944, and 8,227,187; 7,270,956; 7,629,448; 8,158,444; 8,227,187; 8,455,613; 8,575,303; 8,802,450; 8,969,509; 9,139,869; 9,371,559; 9,547,008; 10,094,838; 10,302,648; 10,458,989; Nos. 10,641,775 and 10,962,546, the disclosures of which are incorporated herein by reference in their entireties, and Gaylord et al., J. Am. Chem. Soc., 2001, 123 (26), pp. 6417-6418; Feng et al., Chem. Soc. Rev., 2010, 39, 2411-2419; and Traina et al., J. Am. Chem. Soc., 2011, 133 (32), pp. 12600-12607, the disclosures of which are incorporated herein by reference in their entireties. Specific polymer dyes that can be used include, but are not limited to, BD Horizon Brilliant™ Dyes (e.g., BD Horizon Brilliant™ Violet Dyes (e.g., BV421, BV510, BV605, BV650, BV711, BV786)), BD Horizon Brilliant™ Ultraviolet Dyes (e.g., BUV395, BUV496, BUV737, BUV805), and BD Horizon Brilliant™ Blue Dyes (e.g., BB515, BB550, BB790) (BD Biosciences, San Jose, Calif.). Fluorescent dyes known to those of skill in the art (including, but not limited to, those listed above) or to be discovered in the future may be used in the present methods.

[0079] In some cases, one or more fluorescent dyes that make up a given color signature can each be excited by a common light source (e.g., a common laser). In such cases, the fluorescent dyes that make up a given color signature each have a common excitation maximum wavelength, but may differ from one another in emission maximum wavelength.

[0080] As explained above, any two given distinguishable color signatures may be distinguishable from one another based on the type of fluorescent dyes comprising the barcode and / or the brightness of the signal provided thereby. Thus, any two different color signatures may be distinguishable based on the fluorescent signal obtained from the color signature and / or its intensity. For example, two distinguishable color signatures may be distinguishable from one another because they are composed of different types of fluorescent dyes (e.g., one includes fluorescent dye a and the other includes fluorescent dye b). Two distinguishable color signatures may be distinguishable from one another because they are composed of different amounts of fluorescent dyes (e.g., one signature includes a first amount of fluorescent dye a bound to an identification particle, while the other signature includes fluorescent dye a present in a second amount that differs from the first amount by a detectable amount (e.g., a difference in signal brightness)). The difference in brightness between the identification particles can be easily achieved by varying the amount of fluorescent dye(s) bound to the particles. Combinations of fluorescent dye types and amounts can be used to provide any desired number of unique color signatures.

[0081] In addition to the color signature, the nucleic acid-barcoded identification particles employed in embodiments of the present invention include an oligonucleotide barcode component, which may be referred to as a Pheno-Seq oligonucleotide barcode component. The oligonucleotide barcode component may vary and, in some cases, may range in length from 10 to 500 nt, e.g., from 15 to 100 nt. In some cases, the oligonucleotide barcode component may be composed of ribonucleic acid or deoxyribonucleic acid, as appropriate. The oligonucleotide barcode component of embodiments of the present invention may include an identification particle (i.e., Pheno-Seq particle) barcode domain and other domains that find use in embodiments of the present invention, including, but not limited to, an identification particle (i.e., Pheno-Seq particle) primer binding site, a second domain complementary to a sequence present in the cell-binding bead nucleic acid, etc. The barcode component may be covalently attached to the particle directly or via an intermediate (e.g., disulfide bond), as appropriate. In certain embodiments, the cell-binding bead nucleic acid may be attached to the bead by a cleavable linker that may be cleaved by cell lysis conditions, examples of such linkers include, but are not limited to, disulfide linkers.

[0082] The identification particle barcode domain is a unique identifier, a domain or region that may be used to identify the associated identification particle, e.g., by its sequence. The unique identifier may be, for example, a nucleotide sequence of any suitable length, for example, from about 4 nucleotides to about 200 nucleotides. In some embodiments, the unique identifier is a nucleotide sequence that is 25 nucleotides to about 45 nucleotides in length. In some embodiments, the unique identifier may be approximately, less than, or greater than the following values ​​in length: 4 nucleotides, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 15 nucleotides, 20 nucleotides, 25 nucleotides, 30 nucleotides, 35 nucleotides, 40 nucleotides, 45 nucleotides, 50 nucleotides, 55 nucleotides, 60 nucleotides, 70 nucleotides, 80 nucleotides, 90 nucleotides, 100 nucleotides, 200 nucleotides, or a range between any two of the above values.

[0083] The oligonucleotide barcode of the dual-indexed bead may include an identification particle primer binding site. If a primer binding site is present, the primer binding site may be configured to bind, for example, to a primer used to prepare a sequenceable nucleic acid. In an embodiment, the identification particle primer binding site is used in combination with a primer that is different from any universal primer that may be used in a given workflow. Thus, the identification particle primer binding site may bind to a primer configured to prime nucleic acid synthesis using only the identification particle nucleic acid as a template, and may not bind to other nucleic acids that may be present. In some cases, the primer binding site may be at or about the following numbers in length: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or a number or range of lengths between any two of these nucleotides. The length of the identifying particle primer binding site may vary and may be at least or at most the following numbers in length: 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30. Identification particles configured for use in combination with such primer binding sites may vary in length, and in some cases may range in length from 5 to 30 nucleotides. The primer binding site may be located at the 5' end of the oligonucleotide barcode component.

[0084] The oligonucleotide barcode may comprise a cell-binding bead-complementary sequence, e.g., a domain or region complementary to a sequence present in the cell-binding bead nucleic acid described above. The cell-binding bead-complementary sequence of interest may vary as desired and may be specific, random, or semi-random. In some cases, the cell-binding bead-complementary sequence may be a sequence that hybridizes with a domain or region of the cell-binding bead nucleic acid (e.g., as described in detail above). The length of this domain may vary, and in some cases, the length of this domain may range from 3 to 50 nt, e.g., 5 to 25 nt. If present, this domain may be located at the 3' end of the oligonucleotide barcode.

[0085] As described in detail below, partitions are indexed by unique combinations of identification particles. Among the different identification particles that make up a given unique combination, the oligonucleotide barcode components may share a common domain. For example, the oligonucleotide barcode components of the identification particles may have a common primer binding site, a domain complementary to the cell-binding bead nucleic acid, etc. In such cases, the common domain has the same sequence, and therefore, different particles of the unique combination may have the same common sequence (e.g., the same primer binding site, the same cell-binding bead nucleic acid complementary domain, etc.).

[0086] FIG. 2 shows examples of three different identification particles 210, 220, and 230 according to an embodiment of the present invention. As shown in FIG. 2, each different identification particle comprises a particle and an oligonucleotide barcode component. Identification particle 210 comprises a green, 3-micrometer particle and an oligonucleotide barcode component comprising an identification particle primer binding site at its 5' end, a unique barcode domain, and a cell-binding bead-complementary domain at its 3' end. Identification particle 220 comprises a red, 10-micrometer particle and an oligonucleotide barcode component comprising an identification particle primer binding site at its 5' end, a unique barcode domain, and a cell-binding bead-complementary domain at its 3' end. Identification particle 230 comprises a gray, 3-micrometer particle and an oligonucleotide barcode component comprising an identification particle primer binding site at its 5' end, a unique barcode domain, and a cell-binding bead-complementary domain at its 3' end. If desired, the barcode component may be covalently attached to the particle directly or via an intermediate (e.g., a disulfide bond, as described above). The particles can have any suitable color, examples of suitable colors include, but are not limited to, green, red, blue, gray, yellow, and black.

[0087] Introduction of unique combinations of different nucleic acid-barcoded identification particles into partitions

[0088] For visual indexing of partitions containing functionally assayed single cells, method embodiments include introducing unique combinations of different nucleic acid-barcoded identification particles, e.g., as described above, to partitions containing functionally assayed bead-bound single cells to generate indexed partitions, where the indexed partitions contain functionally assayed bead-bound single cells and unique combinations of nucleic acid-barcoded identification particles.

[0089] The unique combination of different nucleic acid-barcoded identification particles provided to a partition containing a functionally assayed single cell is composed of a plurality of different nucleic acid-barcoded identification particles that differ from one another in one or more of size, color, and brightness. The number of different nucleic acid-barcoded identification particles that make up a given unique combination that may be present in a given partition may vary, and in some cases may range from 1 to 15, e.g., 1 to 10, or 1 to 5, inclusive, e.g., 2 to 5. The size of the different identification particles in a given unique combination may vary, and in some cases may range from 1 to 50 micrometers, e.g., 2 to 25 micrometers, or 3 to 20 micrometers, inclusive. In some cases, the size of the particles that make up a given unique combination is 3, 7, 10, or 16 micrometers. The colors of the different identification particles that make up a given unique combination (e.g., colors provided by the fluorescent dyes of the identification particles) may also vary, and in some cases, the colors of the different particles that make up a unique combination may be green, red, blue, gray, black, or yellow. The brightness between the particles that make up a unique combination may also vary, for example, at least one particle may be bright and at least one particle may be dark.

[0090] Method embodiments include introducing a unique combination of identification particles into a partition containing functionally assayed single cells to provide a visual index of the partition, the visual index being provided by the unique combination of identification particles present in the partition. The unique combination of identification particles may be introduced into the partition using any suitable protocol. In some cases, introducing a unique combination of different nucleic acid-barcoded identification particles into a partition containing functionally assayed bead-bound single cells includes introducing a composition of different nucleic acid-barcoded identification particles into a flow cell having microwells at the bottom containing the functionally assayed bead-bound single cells. The nucleic acid-barcoded identification particle composition used in embodiments of the invention may be a composition of a plurality of different nucleic acid-barcoded identification particles present in a liquid (e.g., an aqueous solution). The composition may be contacted with the compartment by, for example, gravity flow, such that the identification particles can settle into the partition structure. In some cases, an aqueous composition of identification particles is contacted with an array of microwells, for example, by flowing the aqueous composition through the microwell array, such that the identification particles settle into the microwells. The aqueous composition containing the identification particles may be passed through a flow cell in fluid communication with the microwells. Suitable protocols and systems for partitioning bead-bound cells into microwells are described in PCT Application No. PCT / US2016 / 014612 (published as WO / 2016 / 118915), the disclosure of which is incorporated herein by reference. To ensure that the identification particles are deposited in sufficient numbers in the partitions (e.g., as described above), method embodiments may further include introducing a second composition of different nucleic acid-barcoded identification particles into the flow cell, allowing the particles to deposit in the partitions. The number of such repetitions may vary as needed, including, for example, from 2 to 5 times, for example, from 2 to 4 times, or from 2 to 3 times.

[0091] Retrieving image data for indexed partitions

[0092] For example, after introducing a unique combination of identification particles into a partition to generate an indexed partition as described above, an embodiment of the method of the present invention includes acquiring image data of the indexed partition to identify the unique combination of nucleic acid-barcoded identification particles therein, thereby acquiring index data for the partition. The indexed partition may be imaged using any suitable protocol, and image data of divided single cells may be acquired. The acquired image data may vary. Image data may be acquired from any divided cell of interest, and may be acquired from a partition containing a functionally assayed cell of interest. The type of image data acquired may vary and may include live-cell image data. Any suitable protocol may be used to acquire image data of cells in the partition. Examples of imaging protocols that can be used include, but are not limited to, microscopic imaging protocols such as phase contrast microscopy, fluorescence microscopy, quantitative phase contrast microscopy, holotomography, and the BD Rhapsody System (Becton, Dickinson and Company). Images may be generated, for example, by fluorescence imaging. Imaging may include microscopy techniques such as brightfield imaging, oblique illumination, darkfield imaging, dispersion staining, phase contrast, differential interference contrast, interference reflection microscopy, fluorescence, confocal, single plane illumination, or combinations thereof. Imaging may include imaging a portion of a sample (e.g., a slide / array). Imaging may include imaging at least 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% of the divided cells. In some cases, imaging may be performed in discrete steps (e.g., images need not be consecutive). Imaging may include taking at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more different images. Imaging may include taking at most 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more different images.

[0093] In embodiments of the invention, obtaining imaging data for a divided single cell includes obtaining a partition-specific visual index of the partition of interest, and therefore of the functionally assayed bead-bound cells present therein. Thus, the method may include obtaining, for each divided, functionally assayed, bead-bound cell of interest, a partition-specific identification of the unique combination of identification particles present therein, and therefore of the combination of identification particles present in the same partition as that cell. The visual index may be viewed as image data of the unique collection of identification particles present in a given partition.

[0094] Figure 3A shows cells containing partitions visually indexed with unique combinations of identification particles. As shown in Figure 3, different wells contain unique combinations of identification particles, the images of which can be used to visually index those wells. As shown, two different wells contain cells (white arrows). The first of these wells contains two 0.5 micron identification particles, providing a visual index for that well. The second of these wells contains 7.5 micron black beads, providing a visual index for that well. The visual indexes for the first and second wells are different. As described in more detail below, sequence reads obtained from the identification particle barcode component can be used in combination with the visual index to determine the wells from which those reads were obtained. This allows the reads to be matched with any functional data obtained for those wells. As described in more detail below, by matching the reads to cell target reads, the functional data can be linked to sequence data (e.g., omics data) for the wells. In some embodiments, a sequential imaging protocol is used to visually index the wells. Sequential imaging allows for visual indexing of more wells using a subset or fewer Pheno-Seq particles (e.g., compared to a non-sequencing protocol such as that shown in Figure 3A). In sequential imaging, two or more Pheno-Seq particle compositions are sequentially introduced into a partition, with an image acquired after each introduction. That is, two or more iterations of Pheno-Seq particle splitting and subsequent image acquisition are performed. Here, the number of iterations can vary, and in some cases, the number of iterations can range from 2 to 5, e.g., 2 to 4, e.g., 2 to 3. At each time point after splitting, the split Pheno-Seq beads have a unique indexed oligo. An embodiment of a sequential imaging protocol for visually indexing partitions is shown in Figure 3B.As shown in Figure 3B, at time T1, wells A, B, and C contain the same combination of black and green beads. However, at time T2, Pheno-Seq particles / beads are added, resulting in additional black particles in well "A," none in well "B," and red particles in well "C." Note that the Pheno-Seq particles added at T2 have different oligo sequences and are distinguishable by sequencing (e.g., as described in detail below). The difference between the T1 and T2 images is shown in ΔT1 and T2 images. As shown, even though the same black and red particles are present at T2, the oligos are different and distinguishable across time points, allowing for visual indexing of more wells / cells (e.g., compared to protocols using only a single segmentation / imaging step) (as shown in Figure 3A).

[0095] Obtaining sequence data

[0096] After generating visually indexed, functionally assayed, and divided single cells (e.g., as described above), method embodiments may include obtaining sequence data for the visually indexed, functionally assayed, and divided single cells. The sequence data may be obtained for the visually indexed, functionally assayed, and divided single cells using any suitable protocol. In some cases, the sequence data is obtained by using a protocol to generate a sequenceable library of nucleic acids for the visually indexed, functionally assayed, and divided single cells, and then sequencing the library, e.g., using a next-generation sequencing protocol.

[0097] Generation of sequence-ready libraries

[0098] Sequenable libraries of nucleic acids can be prepared from divided cells using any suitable protocol. Of interest are protocols that generate libraries containing cell-labeling domains, which can be used to identify and distinguish nucleic acids from a given divided single cell from nucleic acids from other divided single cells. In some cases, the protocol for preparing a sequenceable library is a cell-capture bead-mediated protocol. In such cases, functionally assayed bead-bound divided single cells may be spatially proximate to cell-capture beads bearing cell-labeling domain nucleic acids containing target-binding regions (e.g., as described in detail below). When the cell-labeling domain nucleic acids are proximate to targets in functionally assayed divided single cells, the targets can hybridize with the cell-labeling domain nucleic acids. The cell-labeling domains containing nucleic acids can be contacted in a non-depletable ratio so that each different target can bind to a different cell-labeling domain containing a nucleic acid with its own unique UMI.

[0099] Thus, in some embodiments, the method further includes providing cell capture beads comprising a cell label comprising a bead-bound nucleic acid to a partition containing the functionally assayed single cell, wherein the cell label comprising the nucleic acid is used in preparing a nucleic acid sequence-ready composition (e.g., a sequence-ready library) from the functionally assayed and partitioned single cell. The nucleic acid of the cell capture bead includes the cell label. The cell label is a unique identifier and is a domain or region that can be used to identify the bound cell capture bead and nearby cells, e.g., by its sequence. The unique identifier can be, for example, a nucleotide sequence of any suitable length, e.g., from about 4 nucleotides to about 200 nucleotides. In some embodiments, the unique identifier is a nucleotide sequence of 25 nucleotides to about 45 nucleotides in length. In some embodiments, a unique identifier may be, approximately, less than, or more than: 4 nucleotides, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 15 nucleotides, 20 nucleotides, 25 nucleotides, 30 nucleotides, 35 nucleotides, 40 nucleotides, 45 nucleotides, 50 nucleotides, 55 nucleotides, 60 nucleotides, 70 nucleotides, 80 nucleotides, 90 nucleotides, 100 nucleotides, 200 nucleotides, or a range between any two of the above values.

[0100] In some cases, the cell capture bead-bound nucleic acid comprises a target binding region that binds to a complementary sequence of a nucleic acid species of interest in a functionally assayed cell, for example, and a target binding region that binds to a capture sequence of a barcode component of a nucleic acid-barcoded cell-binding bead, for example, as described above. For example, if the target nucleic acid species is cellular mRNA and the barcode component comprises a poly(A) capture sequence, the cell capture bead-bound nucleic acid may comprise a poly(T) domain as the target binding region.

[0101] In addition to the target binding region, the bound nucleic acid may further comprise one or more additional domains, including, but not limited to, an additional barcode domain, a molecular index domain (e.g., a unique molecular identifier (UMI) domain), a universal primer binding domain, etc. Further details regarding particles with bound nucleic acids that may be provided within compartments can be found in U.S. Patent Application Publication Nos. US2018 / 0088112; US2018 / 0200710; US2018 / 0346970; US2019 / 0056415; US2020 / 0248263; US2020 / 0299672; and US2021 / 0171940, the disclosures of which are incorporated herein by reference. The nucleic acid-bound beads may be provided to the compartments using any suitable protocol, including but not limited to, the methods described above for dividing cells, and further including but not limited to, the methods described in PCT Application No. PCT / US2016 / 014612 (published as WO / 2016 / 118915). Particles (e.g., beads) can be provided in conjunction with, e.g., near, the single cells, as needed, before, after, or in some cases in combination with the single cells.

[0102] After producing visually indexed, functionally assayed, and divided single cells in the vicinity of the bead-containing cell labeling domain as described above, the cells may be lysed to release the target molecule, allowing the released target molecule (e.g., nucleic acid) to bind to the nucleic acid target-binding region of the cell labeling domain to generate captured nucleic acid. Cell lysis can be performed by any of a variety of means, including chemical or biochemical means, osmotic shock, heat lysis, mechanical lysis, or photolysis. Particle lysis can be achieved by adding a cell lysis buffer containing a detergent (e.g., SDS, lidocaine, Triton X-100, Tween-20, or NP-40), an organic solvent (e.g., methanol or acetone), or a digestive enzyme (e.g., proteinase K, pepsin, or trypsin), or a combination thereof. To increase target-barcode binding, the diffusion rate of the target molecule can be altered, for example, by lowering the temperature or increasing the viscosity of the lysis solution. In some embodiments, the sample can be lysed using filter paper. The filter paper can be soaked in a lysis buffer on top of the filter paper. The filter paper can be applied under pressure to the sample, which can promote lysis of the sample and hybridization of the sample's targets to the substrate. In some embodiments, lysis can be performed by mechanical lysis, heat lysis, photolysis, and / or chemical lysis. Chemical lysis can include the use of digestive enzymes such as proteinase K, pepsin, and trypsin. Lysis can be performed by adding a lysis buffer to the substrate. The lysis buffer can include TrisHCl. The lysis buffer can include at least about 0.01, 0.05, 0.1, 0.5, or 1 M or more TrisHCl. The lysis buffer can include up to about 0.01, 0.05, 0.1, 0.5, or 1 M or more TrisHCl. The lysis buffer can include about 0.1 M TrisHCl. The pH of the lysis buffer may be at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or higher.The pH of the lysis buffer may be up to about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more. In some embodiments, the pH of the lysis buffer is about 7.5. The lysis buffer may contain a salt (e.g., LiCl). The concentration of the salt in the lysis buffer may be at least about 0.1, 0.5, or 1 M, or more. The concentration of the salt in the lysis buffer may be up to about 0.1, 0.5, or 1 M, or more. In some embodiments, the concentration of the salt in the lysis buffer is about 0.5 M. The lysis buffer may contain a detergent (e.g., SDS, Li dodecyl sulfate, Triton X, Tween, NP-40). The concentration of the detergent in the lysis buffer may be at least about 0.0001%, 0.0005%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, or 7%, or more. The concentration of the detergent in the lysis buffer may be up to about 0.0001%, 0.0005%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, or 7%, or more. In some embodiments, the concentration of the detergent in the lysis buffer is about 1% Li-dodecyl sulfate. The time required for the lysis method may depend on the amount of detergent used. In some embodiments, the greater the amount of detergent used, the shorter the time required for lysis. The lysis buffer can include a chelating agent (e.g., EDTA, EGTA). The concentration of the chelating agent in the lysis buffer can be at least about 1, 5, 10, 15, 20, 25, or 30 mM, or more. The concentration of the chelating agent in the lysis buffer can be up to about 1, 5, 10, 15, 20, 25, or 30 mM, or more. In some embodiments, the concentration of the chelating agent in the lysis buffer is about 10 mM. The lysis buffer can include a reducing agent (e.g., β-mercaptoethanol, DTT). The concentration of the reducing agent in the lysis buffer can be at least about 1, 5, 10, 15, or 20 mM, or more. The concentration of the reducing agent in the lysis buffer can be up to about 1, 5, 10, 15, or 20 mM, or more.In some embodiments, the concentration of the reducing agent in the lysis buffer is about 5 mM. In some embodiments, the lysis buffer can comprise about 0.1 M TrisHCl, about pH 7.5, about 0.5 M LiCl, about 1% Li dodecyl sulfate, about 10 mM EDTA, and about 5 mM DTT. Lysis can be performed at a temperature of about 4, 10, 15, 20, 25, or 30° C. Lysis can be performed for about 1, 5, 10, 15, or 20 minutes or more. The lysed cells can comprise at least about 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, or 700,000 or more target nucleic acid molecules. The lysed cells can contain up to about 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, or 700,000 or more target nucleic acid molecules.

[0103] After cell lysis and release of nucleic acid molecules therefrom, the nucleic acid molecules can randomly bind to the cell-labeling domain nucleic acid of the co-localized cell capture bead. Binding can involve hybridization between the target recognition region of the cell-labeling domain nucleic acid and a complementary portion of the target nucleic acid molecule (e.g., an oligo(dT) in the barcode can interact with the poly(A) tail of the target). The assay conditions (e.g., buffer pH, ionic strength, temperature, etc.) used for hybridization can be selected to promote the formation of specific and stable hybrids. In some embodiments, nucleic acid molecules released from lysed cells can bind to (e.g., hybridize with) multiple probes on a substrate. If the probes contain oligo(dT), mRNA molecules can hybridize to the probes, and the mRNA molecules can be reverse-transcribed. The oligo(dT) portion of the oligonucleotide can serve as a primer for first-strand synthesis of cDNA molecules (e.g., when exposed to DNA synthesis reaction conditions to generate first-strand cDNA domains containing the capture nucleic acid).

[0104] The cell-label domain nucleic acid can also hybridize to the complementary capture sequence of the barcode component (e.g., poly(A) sequence) of the cell-binding bead stably bound to the cell. In this way, the cell-label domain nucleic acid can serve as a primer for reverse transcription using the barcode as a template (e.g., as described in detail below). In addition to the generation of the cell-label domain / cell-binding bead nucleic acid hybridization complex, a hybridization complex between the cell-binding bead nucleic acid and the identification particle nucleic acid is also generated via the complementary domains of the cell-binding bead nucleic acid and the identification particle nucleic acid. Figure 4 shows a schematic diagram of the oligo interactions between the cell capture bead, cell-binding beads (i.e., IMag beads bound to cells), and identification particles (Pheno-Seq particles) within a single well. As shown in Figure 4, the cell-capture bead nucleic acid, which includes a universal primer binding site, a cell-label domain (a unique index for each well), and a target-binding region (polyT), hybridizes to the capture sequence (polyA) of the cell-binding bead nucleic acid. Furthermore, the cell-binding bead nucleic acid hybridizes to the complementary domain in the identification particle nucleic acid. In this way, a reverse-transcribeable hybridization complex is generated between: (1) cell capture bead / cell binding bead nucleic acid; and (2) cell binding bead and identification particle nucleic acid.

[0105] In some examples, the method further includes using the oligonucleotide-labeled cellular component-binding reagent in applications where detection (e.g., quantification) of one or more cellular components, such as surface proteins, is desired, such as in AbSeq applications. The oligonucleotide-labeled cellular component-binding reagent used in such embodiments includes a cellular component-binding reagent (e.g., an antibody or binding fragment thereof) conjugated to a cellular component-binding reagent-specific oligonucleotide, the oligonucleotide including a cellular component-binding reagent-specific recognition sequence, and the cellular component-binding reagent has the cellular component-binding reagent-specific oligonucleotide conjugated to it. In such cases, the cell capture beads may include nucleic acids, which may be configured to capture, for example, a domain (e.g., a poly-T sequence) of the cellular component-binding reagent-specific oligonucleotide (e.g., as described above). In this manner, protein expression can be measured in combination with gene expression, for example, when multi-omic analysis (e.g., combined analysis of the transcriptome and proteome) is desired. In such cases, the method may include preparing a sample captured using the oligonucleotide-labeled cellular component-binding reagent and capturing the cellular component-binding reagent-specific oligonucleotide released from the captured and separated cells. Further details regarding the use of oligonucleotide-labeled cellular component binding reagents are described in U.S. Published Patent Application Nos. US20180267036 and US20200248263, the disclosures of which are incorporated herein by reference.

[0106] Optionally, a given workflow may include a pooling step, in which, for example, a product composition consisting of captured nucleic acid, synthesized first-strand cDNA, or synthesized double-strand cDNA may be combined or pooled with product compositions obtained from one or more additional samples (e.g., combined barcoded cells). In some cases, the pooling step is performed immediately after the hybridization step between the cell label domain nucleic acid and the target nucleic acid (e.g., as described above). In such embodiments, the number of different product compositions generated from different samples (e.g., cells) that are combined or pooled may vary, and in some cases may be in the range of 2 to 1,000,000, e.g., 3 to 200,000, 4 to 100,000, 5 to 50,000, and in some cases may be in the range of 100 to 10,000, e.g., 1,000 to 5,000. Before or after pooling, the product composition can be amplified (e.g., by polymerase chain reaction (PCR)) (e.g., as described in detail below). After the target cell domain-labeled complexes and cell-bound bead nucleic acid / identification particle nucleic acid complexes are pooled, all further processing can proceed within a single reaction vessel. Further processing may include, for example, reverse transcription reactions, amplification reactions, cleavage reactions, dissociation reactions, and / or nucleic acid extension reactions. Further processing reactions can be performed within the microwells, i.e., without first pooling the hybridized complexes from multiple cells.

[0107] The present disclosure provides methods for creating target-cell-label domain conjugates using any suitable protocol, such as reverse transcription or nucleotide extension. The target-cell-label domain conjugate can comprise a cell-label domain and a sequence complementary to all or part of a target nucleic acid. Reverse transcription of the bound RNA molecule can occur by adding a reverse transcription primer and a reverse transcriptase. The reverse transcription primer can be an oligo(dT) primer, a random hexanucleotide primer, or a target-specific oligonucleotide primer. The oligo(dT) primer can be 12 to 18 nucleotides in length, or approximately 12 to 18 nucleotides in length, and can bind to the endogenous poly(A) tail at the 3' end of mammalian mRNA. The random hexanucleotide primer can bind to multiple complementary sites on the mRNA. The target-specific oligonucleotide primer typically selectively primes the mRNA of interest. Reverse transcription can be repeated to generate multiple cDNA molecules. The methods of the disclosure may include performing at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 reverse transcription reactions. The methods can include performing at least about 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 reverse transcription reactions.

[0108] One or more nucleic acid amplification reactions can be performed to generate multiple copies of a target nucleic acid molecule. Amplification can be performed in a multiplexed format, simultaneously amplifying multiple target nucleic acid sequences. Amplification reactions can be used to add sequence adapters to nucleic acid molecules. The amplification reaction can include amplifying at least a portion of a sample label, if present. The amplification reaction can include amplifying at least a portion of a cell label and / or a barcode sequence (e.g., a molecular label). The amplification reaction can include amplifying at least a portion of a sample tag, a cell label, a spatial label, a barcode sequence (e.g., a molecular label), a target nucleic acid, or a combination thereof. The amplification reaction can include amplifying by the following percentages: 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 100%, or a range or value between any two of these values. Further, the method can include performing one or more cDNA synthesis reactions to produce one or more cDNA copies of a target barcode molecule comprising a sample label, a cell label, a spatial label, and / or a barcode sequence (e.g., a molecular label).

[0109] In some embodiments, amplification can be performed using polymerase chain reaction (PCR). As used herein, PCR can refer to a reaction in vitro that amplifies a specific DNA sequence by simultaneously performing primer extension of complementary strands of DNA. As used herein, PCR can include derivative forms of the reaction, including, but not limited to, RT-PCR, real-time PCR, nested PCR, quantitative PCR, multiplex PCR, digital PCR, and assembly PCR.

[0110] Nucleic acid amplification can include non-PCR-based methods. Examples of non-PCR-based methods include, but are not limited to, multiple displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), real-time SDA, rolling circle amplification, or circle-to-circle amplification. Other PCR-independent amplification methods include amplification of DNA or RNA targets by DNA-dependent RNA polymerase-driven RNA transcription amplification or RNA-guided DNA synthesis and transcription, ligase chain reaction (LCR), Qβ replicase (Qβ) method, the use of palindromic probes, strand displacement amplification, oligonucleotide-driven amplification (using restriction endonucleases), amplification methods in which a primer is hybridized to a nucleic acid sequence and the resulting duplex is cleaved before extension and amplification, strand displacement amplification using a nucleic acid polymerase lacking 5' exonuclease activity, rolling circle amplification, and branched extension amplification (RAM). In some embodiments, amplification does not produce circular transcripts.

[0111] In some embodiments, the methods disclosed herein further include performing a polymerase chain reaction on the nucleic acid (e.g., RNA, DNA, cDNA) to generate labeled amplicons (e.g., stochastically labeled amplicons). The labeled amplicons may be double-stranded molecules. The double-stranded molecules may include double-stranded RNA molecules, double-stranded DNA molecules, or RNA molecules hybridized with DNA molecules. One or both strands of the double-stranded molecules may include a sample label, a spatial label, a cellular label, and / or a barcode sequence (e.g., a molecular label). The labeled amplicons may be single-stranded molecules. The single-stranded molecules may include DNA, RNA, or a combination thereof. The nucleic acids of the present disclosure may include synthetic or modified nucleic acids. Accordingly, the methods may include producing an amplicon composition from a first-strand cDNA domain that includes a capture nucleic acid.

[0112] Amplification may include the use of one or more non-natural nucleotides. Non-natural nucleotides may include photocleavable or triggerable nucleotides. Examples of non-natural nucleotides may include, but are not limited to, peptide nucleotides (PNAs), morpholinos, locked nucleotides (LNAs), glycol nucleotides (GNAs), and threose nucleotides (TNAs). Non-natural nucleotides may be added to one or more cycles of the amplification reaction. The addition of non-natural nucleotides can be used to identify products at specific cycles or time points in the amplification reaction.

[0113] Performing one or more amplification reactions can include the use of one or more primers. The one or more primers can include, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 or more nucleotides. The one or more primers can include at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 or more nucleotides. The one or more primers can include 12 to fewer than 15 nucleotides. The one or more primers can anneal to at least a portion of the multiple labeled targets (e.g., stochastically labeled targets). The one or more primers can anneal to the 3' or 5' ends of the multiple labeled targets. The one or more primers can anneal to an internal region of the multiple labeled targets. The internal region can be at least about 50, 100, 150, 200, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 650, 700, 750, 800, 850, 900, or 1000 nucleotides from the 3' end of the plurality of labeled targets. The one or more primers can comprise a fixed panel of primers. The one or more primers may include at least one or more custom primers. The one or more primers may include at least one or more control primers. The one or more primers may include at least one or more gene-specific primers.

[0114] The one or more primers can include a universal primer. The universal primer can anneal to the universal primer binding site. The one or more custom primers can anneal to a first sample label, a second sample label, a spatial label, a cellular label, a barcode sequence (e.g., a molecular label), a target, or any combination thereof. The one or more primers can include a universal primer and a custom primer. The custom primers can be designed to amplify one or more targets. The targets can constitute a subset of the total nucleic acids in one or more samples. The targets can comprise a subset of the total labeled targets in one or more samples. The one or more primers can include at least 96 or more custom primers. The one or more primers can include at least 960 or more custom primers. The one or more primers can include at least 9600 or more custom primers. The one or more custom primers can anneal to two or more different labeled nucleic acids. The two or more different labeled nucleic acids may correspond to one or more genes.

[0115] Any amplification scheme can be used in the disclosed methods. For example, in one scheme, a first round of PCR can amplify bead-bound molecules using a gene-specific primer and a primer for the universal Illumina sequencing primer 1 sequence. A second round of PCR can amplify the first round PCR product using a nested gene-specific primer flanked by Illumina sequencing primer 2 sequences and a primer for the universal Illumina sequencing primer 1 sequence. A third round of PCR adds P5, P7, and a sample index, converting the PCR product into an Illumina sequencing library. Sequencing using 150 bp x 2 sequences reveals cell labels and barcode sequences (e.g., molecule labels) in read 1, genes in read 2, and sample indexes in index 1 read.

[0116] In some embodiments, nucleic acids can be removed from a substrate using chemical cleavage. For example, chemical groups or modified bases present in the nucleic acid can be used to facilitate removal from the solid support. For example, nucleic acids can be removed from a substrate using enzymes. For example, nucleic acids can be removed from a substrate by restriction enzyme digestion. For example, nucleic acids containing dUTP or ddUTP can be removed from a substrate by treating the substrate with uracil-D-glycosidase (UDG). For example, nucleic acids can be removed from a substrate using nucleotide excision enzymes, such as apurinic / apyrimidinic (AP) endonuclease. In some embodiments, nucleic acids can be removed from a substrate using photocleavable groups and light. In some embodiments, a cleavable linker can be used to remove nucleic acids from a substrate. For example, the cleavable linker can include at least one of biotin / avidin, biotin / streptavidin, biotin / neutravidin, Ig-Protein A, a photocleavable linker, an acid- or base-cleavable linker group, or an aptamer.

[0117] In some embodiments, amplification can be performed on the substrate. For example, bridge amplification can be used. The cDNA can be homopolymerically tailed to generate ends compatible with bridge amplification using oligo(dT) probes on the substrate. In bridge amplification, a primer complementary to the 3' end of the template nucleic acid can be the first primer of each pair covalently attached to solid particles. When a sample containing the template nucleic acid is contacted with the particles and a single thermal cycle is performed, the template molecule can anneal to the first primer, and the first primer can be extended in the forward direction by adding nucleotides, forming a duplex molecule consisting of the template molecule and a newly formed DNA strand complementary to the template. In the heating step of the next cycle, the duplex molecule can be denatured, the template molecule can be released from the particle, and the complementary DNA strand can remain attached to the particle via the first primer. In the annealing stage of the subsequent annealing and extension step, the complementary strand can hybridize with a second primer complementary to the segment of the complementary strand located where the first primer was removed. This hybridization can result in the complementary strand being covalently bound to the first primer and hybridized to the second primer, forming a bridge between the first and second primers. In the extension step, the second primer can be extended in the reverse direction by adding nucleotides to the same reaction mixture, thereby converting the bridge into a double-stranded bridge. The next cycle can then begin, with the double-stranded bridge denaturing to generate two single-stranded nucleic acid molecules, one end of which is bound to the particle surface via the first primer and the second primer, and the other end of which is unbound. In the annealing and extension step of this second cycle, each strand can hybridize to a previously unused complementary primer present on the same particle, forming a new single-stranded bridge. The two currently hybridized, unused primers are extended, converting the two new bridges into double-stranded bridges.The amplification reaction can include at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 100% amplification of a plurality of nucleic acids.

[0118] Amplification of labeled nucleic acids can include PCR-based methods or non-PCR-based methods. Amplification of labeled nucleic acids can include exponential amplification of labeled nucleic acids. Amplification of labeled nucleic acids can include linear amplification of labeled nucleic acids. Amplification can be performed by polymerase chain reaction (PCR). PCR can refer to a reaction that simultaneously performs primer extension of complementary strands of DNA to achieve in vitro amplification of specific DNA sequences. PCR can include derivative forms of the reaction, including, but not limited to, RT-PCR, real-time PCR, nested PCR, quantitative PCR, multiplex PCR, digital PCR, suppression PCR, semi-suppression PCR, and assembly PCR.

[0119] In some embodiments, amplification of labeled nucleic acids includes non-PCR-based methods. Examples of non-PCR-based methods include, but are not limited to, multiple displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), real-time SDA, rolling circle amplification, or circle-to-circle amplification. Other PCR-independent amplification methods include: amplification of DNA or RNA targets by DNA-dependent RNA polymerase-driven RNA transcription amplification or RNA-guided DNA synthesis and transcription, ligase chain reaction (LCR), Qβ replicase (Qβ), use of palindromic probes, strand displacement amplification, oligonucleotide-driven amplification using restriction endonucleases, amplification methods in which a primer is hybridized to a nucleic acid sequence and the resulting duplex is cleaved prior to extension and amplification, strand displacement amplification using a nucleic acid polymerase lacking 5' exonuclease activity, rolling circle amplification, and / or branched extension amplification (RAM).

[0120] In some embodiments, the methods disclosed herein further include performing a nested polymerase chain reaction on the amplified amplicon (e.g., target). The amplicon may be a double-stranded molecule. The double-stranded molecule may comprise a double-stranded RNA molecule, a double-stranded DNA molecule, or an RNA molecule hybridized with a DNA molecule. One or both strands of the double-stranded molecule may comprise a sample tag or molecular identification label. Alternatively, the amplicon may be a single-stranded molecule. The single-stranded molecule may comprise DNA, RNA, or a combination thereof. The nucleic acids of the present invention may include synthetic or modified nucleic acids.

[0121] In some embodiments, the methods involve iteratively amplifying labeled nucleic acids to generate multiple amplicons. The methods disclosed herein can include performing at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amplification reactions. Alternatively, the methods can include performing at least about 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 amplification reactions.

[0122] The amplification can further include adding one or more control nucleic acids to one or more samples containing the plurality of nucleic acids. The amplification can further include adding one or more control nucleic acids to the plurality of nucleic acids. The control nucleic acids can include a control label.

[0123] Amplification can include the use of one or more non-natural nucleotides. Non-natural nucleotides can include photocleavable and / or triggerable nucleotides. Examples of non-natural nucleotides include, but are not limited to, peptide nucleotides (PNAs), morpholinos, locked nucleotides (LNAs), glycol nucleotides (GNAs), and threose nucleotides (TNAs). Non-natural nucleotides can be added to one or more cycles of the amplification reaction. The addition of non-natural nucleotides can be used to identify products at specific cycles or time points in the amplification reaction.

[0124] Performing one or more amplification reactions may include using one or more primers. The one or more primers may include one or more oligonucleotides. The one or more oligonucleotides may include at least about 7-9 nucleotides. The one or more oligonucleotides may include 12-15 nucleotides or less. The one or more primers may anneal to at least a portion of the plurality of labeled nucleotides. The one or more primers may anneal to the 3' and / or 5' ends of the plurality of labeled nucleic acids. The one or more primers may anneal to an internal region of the plurality of labeled nucleic acids. The internal region can be at least about 50, 100, 150, 200, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 650, 700, 750, 800, 850, 900, or 1000 nucleotides from the 3' end of the plurality of labeled nucleic acids. The one or more primers can comprise a fixed panel of primers. The one or more primers may include at least one or more custom primers. The one or more primers may include at least one or more control primers. The one or more primers may include at least one or more housekeeping gene primers. The one or more primers may include a universal primer. The universal primer may anneal to a universal primer binding site. The one or more custom primers may anneal to a first sample tag, a second sample tag, a molecular identifier label, a nucleic acid, or a product thereof. The one or more primers may include a universal primer and a custom primer. The custom primer may be designed to amplify one or more target nucleic acids.The target nucleic acids can comprise a subset of the total nucleic acids in one or more samples. In some embodiments, the primers are probes bound to the disclosed arrays.

[0125] In some embodiments, barcoding (e.g., stochastic barcoding) a plurality of targets in a sample further includes generating an indexed library of barcoded targets (e.g., stochastically barcoded targets) or barcoded fragments of the targets. The barcode sequences of different barcodes (e.g., molecular labels of different stochastic barcodes) may differ from each other. Generating an indexed library of barcoded targets includes generating a plurality of indexed polynucleotides from the plurality of targets in the sample. For example, in an indexed library of barcoded targets including a first indexed target and a second indexed target, the label region of the first indexed polynucleotide differs from the label region of the second indexed polynucleotide by, approximately, at least, or at most 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, or any number or range between these values. In some embodiments, generating an indexed library of barcoded targets includes contacting a plurality of targets, e.g., mRNA molecules, with a plurality of oligonucleotides each comprising a poly(T) region and a label region; performing first-strand synthesis using a reverse transcriptase to generate single-stranded, labeled cDNA molecules, each comprising a cDNA region and a label region, wherein the plurality of targets includes at least two mRNA molecules of different sequences and the plurality of oligonucleotides includes at least two oligonucleotides of different sequences. Generating an indexed library of barcoded targets can further include amplifying the single-stranded, labeled cDNA molecules to generate double-stranded, labeled cDNA molecules; and performing nested PCR on the double-stranded, labeled cDNA molecules to generate labeled amplicons. In some embodiments, the method can include generating adapter-labeled amplicons.

[0126] Barcoding (e.g., probabilistic barcoding) can involve labeling individual nucleic acid (e.g., DNA or RNA) molecules using nucleic acid barcodes or tags. In some embodiments, this involves adding DNA barcodes or tags to cDNA molecules generated from mRNA. Nested PCR can be performed to minimize PCR amplification bias. Adapters may be added for sequencing, for example, using next-generation sequencing (NGS). Sequencing results can be used to determine the sequence of cellular labels, molecular labels, and nucleotide fragments of one or more target copies.

[0127] In certain embodiments, the provided methods further include subjecting the prepared expression library, e.g., the amplicon composition generated as described above, to a sequencing protocol, such as an NGS protocol. The protocol can be performed on any suitable NGS sequencing platform. NGS sequencing platforms of interest include, but are not limited to, sequencing platforms provided by Illumina® (e.g., HiSeq TM , MiSeq TM and / or NextSeq TM sequencing system), Ion Torrent TM (e.g., Ion PGM TM and / or Ion Proton TM sequencing systems); Pacific Biosciences (e.g., PACBIO RS II Sequel sequencing system); Life Technologies TM(e.g., SOLiD sequencing system), Oxford Nanopore (e.g., Minion), Roche (e.g., 454 GS FLX+ and / or GS Junior sequencing systems), or other target sequencing platforms. NGS protocols vary depending on the NGS sequencing system used. Detailed sequencing protocols (e.g., including amplicon sequencing, which may include further amplification (e.g., solid-phase amplification), and analysis of sequence data) are available from the manufacturer of the NGS sequencing system used.

[0128] Figures 5A and 5B show example amplification workflows that can be used to generate sequence-ready libraries according to embodiments of the present invention. Figure 5A illustrates how cell-bound bead primers (I-Mag AbSeq primers) are included in the amplification of signals from cell-bound beads (IMag AbSeq beads), and (C) how cell-bound bead primers (I-Mag AbSeq primers) and identification particle primers (PhenoSeq primers) are included in the amplification of signals from cell-bound beads (IMag AbSeq beads) and identification particles (PhenoSeq). Figure 5B illustrates a workflow in which library sequences for identifying the oligo products of the identification particles (PhenoSeq) are included in the final identification particle (PhenoSeq Index) PCR step. This step includes a modified PCR forward primer that combines the library forward primer and PhenoSeq primer sequences. This is used because the identification particle / cell-bound bead (e.g., PhenoSeq - IMag AbSeq) products amplified from the cell-bound bead nucleic acid (IMag bead nucleic acid) lack a universal primer sequence.

[0129] Details of methods for obtaining sequence data from single cells are provided in, for example, U.S. Patent Application Publication No. US2018 / 0088112; U.S. Patent Application Publication No. 2018 / 0200710; U.S. Patent Application Publication No. US2018 / 0346970; U.S. Patent Application Publication No. 2019 / 0056415; U.S. Patent Application Publication No. US 2020 / 0248263; U.S. Patent Application Publication No. 2020 / 0299672; and U.S. Patent Application Publication No. 2021 / 0171940, the disclosures of which are incorporated herein by reference.

[0130] This sequencing protocol generates sequence data for the combined barcoded cells. This sequence data can be easily linked to image data for the combined barcoded cells, allowing image data and sequence data obtained from the same combined barcoded cells to be paired. That is, a given image data set and a given sequence data set may be linked as being obtained from the same combined barcoded cells (e.g., as described in more detail below).

[0131] Linking single-cell functional and sequence data

[0132] After obtaining sequence data (e.g., as described above), embodiments of the method include linking functional and sequence data for single cells that have been sequenced, visually indexed, functionally assayed, and partitioned. In such cases, the functional and sequence data obtained from a given partition (and thus the cells present in that partition) are linked. By "linked," we mean that the functional and sequence data are paired as originating from the same partition, and therefore from the cells present in that partition when the functional data for that partition was obtained. Thus, functional and sequence data obtained from the same cell may be paired. In other words, a given set of functional data and sequence data (which may also be associated with a given set of omics data, such as transcriptome data and proteome data, e.g., as obtained with an AbSeq platform) may be identified as originating from the same cell and may then be paired or otherwise associated. This method allows for the acquisition of linked functional and sequence data for single cells of a cell sample.

[0133] Functional data and sequence data are linked by using cell-binding bead nucleic acid barcodes to identify sequence reads of nucleic acid amplicons obtained from the same partition. The obtained sequence data (e.g., as described above) includes sequence reads of cell targets, cell-binding bead nucleic acids, and identification particle nucleic acids. Reads are obtained from nucleic acids that contain both cell-binding bead nucleic acid barcodes and cell label barcodes, and this cell label barcode is also present in all reads of cell target nucleic acids within a given partition. Additionally, reads are obtained from nucleic acids that contain both the identification particle barcode provided by the identification particles from a given partition and the cell-binding bead nucleic acid barcode. The cell-binding bead nucleic acid barcodes may be used to identify the partitions from which the reads resulting from the cell targets and the reads resulting from the identification particles were obtained. That is, for each cell assayed in a given workflow, the sequence of the identification particle barcode associated with the partition containing that cell and the sequence of the target nucleic acid from that cell (e.g., mRNA from the cell) are obtained. For each cell, these resulting sequences are obtained using a protocol for generating a library from the original sequence (a next-generation sequencing protocol, as described above), where each member of a given library generated from the same partition shares a common cell label of the cell target and the identification particle nucleic acid, which can be linked by a common cell-binding bead nucleic acid barcode. In this way, sequence reads of the cell target nucleic acid and the identification particle obtained from the same cell can be identified as originating from the same partition by linking different reads using the cell-binding bead nucleic acid barcode. When linking functional data and sequence data, all reads that have the same cell label domain, i.e., share a common cell label from both: (a) the target nucleic acid read; and (b) the cell-binding bead nucleic acid read. This pairing or linking creates a set of reads including the target nucleic acid read and the cell-binding bead barcode nucleic acid read, and these reads can be identified as originating from the same cell.Furthermore, all reads that share the same cell-binding bead nucleic acid barcode, i.e., a common cell-binding bead nucleic acid barcode from both the identification particle nucleic acid read and the cell-binding bead nucleic acid read, may be paired or linked. The commonality of the cell-binding bead nucleic acid barcodes allows the reads from the cellular target and the reads from the identification particle to be paired or linked, identifying these reads as originating from the same partition and therefore from cells present therein.

[0134] The resulting sequence data, including the target nucleic acid and identification particle reads, may then be matched, i.e., paired or linked, with the functional data. As described above, the functional data of a cell may be assigned to a particular partition, which may be visually indexed by the unique combination of identification particles present within that partition. Reads from a unique combination of identification particles are obtained from a given partition (e.g., as described above), and the reads are matched or linked to the visual index obtained for that partition to determine their origin. Different partitions in a given workflow have their own unique visual index, provided by the unique combination of identification particles present therein. A given unique combination of identification particles that constitutes such a partition-specific visual index can be assigned to a given portion of the sequence reads because the sequence of the barcodes of the identification particles from which that visual index was obtained is known. Thus, each partition-specific visual index obtained for a given partition and the cells present within that partition can be used to determine the sequence of the different identification particle barcodes associated with that cell. Because the sequences of the identifying particle barcodes are present in the barcode reads, a given set of identifying barcodes can be determined to be associated with a given set of sequence data. When a partition-specific set of identifying particle barcode sequence reads is associated with a given set of sequence data, the sequence data can be determined to be obtained from the same cells that were present in the partition from which the set of identifying particle sequence reads was obtained. Thus, from the visual index obtained from a given partition (a composite of different signals obtained from unique combinations of identifying particles within a partition), a set of sequences of identifying particle barcode regions for the given partition can be obtained.This sequence series or collection of identifying particle barcode regions may be used to identify all sequence data obtained from that partition, for example, by linking the sequences using cell-binding bead nucleic acid barcodes, as described above. This identification may be performed by determining that sequence reads obtained from cells present in the partition: (a) have a common cell barcode and cell-binding bead nucleic acid barcode; and (b) have a partition-specific set of identifying particle barcode sequences that also include cell-binding bead nucleic acid barcodes. Once sequence data is assigned to a given partition, that sequence data can be easily linked to functional data obtained from that partition. This method provides linked functional and sequence data for a single cell in a cell sample.

[0135] Figure 6 illustrates how single-cell identity can be bioinformatically deconvolved and correlated with its function / phenotype. As shown, each well is associated with a unique index sequence derived from the cell capture beads (CCB#) present in that well, multiple unique index sequences derived from cell-binding beads (IMag AbSeq beads) (BBB, DDD, …, GGG), and a unique sequence present on the identification (Phenoseq) particles (15 micrometer red beads, 7 micrometer black beads, 7 micrometer green beads).

[0136] Typical Workflow

[0137] A representative workflow according to an embodiment of the present invention is shown in Figure 7. Figure 7 shows the BD Rhapsody TM The workflow executed on the system is shown. As shown in the figure, BD Rhapsody TMBD Rhapsody is being adopted as a platform for assessing the phenotype and function of single cells in real time and combining the resulting data with CITE-Seq (Cellular indexing of Transcriptome, Epitopes) information. TM The scanner and cartridges support single-cell functional assays, such as targeted cell lysis assays, chemotaxy assays, and assays evaluating response to small molecules. Images acquired in steps 3 and 6 are used to visually index cells using a combination of Pheno-Seq particles and match them with associated phenotypes and functions. (The index associated with the cell capture beads contains transcriptome and epitope information associated with each single cell, which is also matched to PCR products amplified from oligos tagged to Pheno-Seq particles. Together, these pieces of information are used to combine single-cell transcriptome and epitope data with cellular phenotype and function.)

[0138] The workflow of the proposed method is as follows:

[0139] 1. Cells obtained from in vitro culture or ex vivo culture are purified using iMag AbSeq beads (CD45 + The cells are mixed with iMag beads, which contain antibodies that target cells of interest, such as white blood cells.

[0140] 2. Cells bound to iMag AbSeq beads are dispensed into a Rhapsody cartridge to obtain single cells in the wells.

[0141] 3. Perform single cell functional assays in the wells of the Rhapsody cartridge.

[0142] 4. Read the output of the single-cell functional assay with the Rhapsody scanner to capture the cell phenotype.

[0143] 5. Optional step: While the cartridge is still on the magnet, wash the reagents from the functional assay (add target cells to mimic the cancer microenvironment). The IMag AbSeq beads will bind to the cells of interest and retain them in the well.

[0144] 6. Add Pheno-Seq particles to achieve an average of 6 particles per well. Image the wells to ensure that the majority of wells containing single cells are indexed with a unique combination of Pheno-Seq particles. You may want to include sequential Pheno-Seq particles until most wells containing single cells are indexed.

[0145] 7. Dispense cell capture beads and continue with cell lysis and subsequent steps according to your current Rhapsody workflow.

[0146] 8. Include an additional library to amplify PCR products from oligos bound to IMag Ab-seq beads.

[0147] 9. Correlate sequencing results with Rhapsody scanner images to identify relevant phenotypes and functions of single cells.

[0148] kit

[0149] Aspects of the invention further include kits and compositions for use in practicing various embodiments of the methods of the invention. Kits of the invention can include one or more of the following: cell-binding beads and / or groups of cell-binding bead nucleic acid primers; distinct identification particles and / or groups of identification particle nucleic acid primers; beads (such as those described above) containing bead-bound nucleic acids that include a cell label domain and a target binding region.

[0150] The kit may further include one or more additional components used to carry out the method embodiments. For example, the kit may include one or more components used to obtain sequence data, such as one or more of the following: primers, polymerase (e.g., a thermostable polymerase, a reverse transcriptase with hot-start properties, or the like), dsDNAse, exonuclease, dNTPs, metal cofactors, one or more nuclease inhibitors (e.g., an RNase inhibitor and / or a DNase inhibitor), one or more molecular crowding agents, polyethylene glycol, etc., one or more enzyme stabilizing components (e.g., DTT), a stimuli-responsive polymer, or any other kit components (e.g., a device, solid support, container, cartridge, tube, bead, plate, microfluidic chip, etc., as described above). The components of the kit may be present in separate containers, or multiple components may be contained in a single container.

[0151] In addition to the above components, the subject kits can further (in certain embodiments) include instructions for practicing the subject methods. These instructions may be present in a variety of forms within the subject kits, one or more of which may be present within the kit. One form in which these instructions may be present is as information printed on a suitable medium or substrate, such as one or more pieces of paper with the information printed on the kit packaging, package insert, etc. Another form in which these instructions may be present is as information recorded on a computer-readable medium (e.g., a floppy disk, compact disc (CD), portable flash drive, etc.). Yet another form in which these instructions may be present is a website address that can be used to access the information at a remote site via the internet.

[0152] The following are examples and not limitations.

[0153] experiment

[0154] I. Pheno Seq particles

[0155] Pheno-Seq particles (also referred to herein as "identification particles") are a combination of numerous micrometer-sized particles with unique phenotypes (size, color, shape, or fluorescence intensity) tagged with oligonucleotides (such as Ab-seq antibodies). Each unique Pheno-Seq particle is bound to a unique oligonucleotide sequence that can be matched to that particle. Pheno-Seq particle combinations are used to index different cells within the Rhapsody cartridge. By sequentially adding 60 Pheno-Seq particle combinations with a target of six or fewer per well, it is possible to index cells in over 300 million wells. The current Rhapsody workflow supports 5,000–10,000 single cells per cartridge. To index approximately 10,000 wells with single cells, 20 Pheno-Seq particles with a target of six or fewer per well are sufficient. The unique oligonucleotide sequences tagged to these particles, when combined with cell capture beads, help generate a set of unique indexing PCR products. Because the sequence of the oligo tagged to each Pheno-Seq particle is unique, the sequence data can be deconvoluted and matched to the combination of Pheno-Seq particles in the well.

[0156] II. Visual Indexing

[0157] Here, we describe a method to convert the location of each single cell within a cell-containing well into a unique barcode index that can be identified using sequence data acquired with the Rhapsody workflow. Particles, distinguishable based on their size, shape, and color, are tagged with unique sequences, such as Ab-seq oligos, with minor modifications to the design of the tagged oligos.

[0158] As shown in Table 1, particles are randomly distributed and follow a Poisson distribution. Starting with a pool of 20 unique phenotypic particles and aiming for an average of 6 particles per well, a median of 50% of wells will contain 4–7 particles, allowing approximately 39,000 wells to be indexed as containing cells. Most wells containing single cells (2.5–5%) likely contain a unique combination of Pheno-Seq particles. Note that the unique combination of beads in a single-cell well is important (though the same combination present in a cell-free well will not result in a contradictory result, as this product will not be amplified).

[0159] Table 1: Cumulative Poisson distribution of Pheno-Seq particles with a target average of 6 particles per well (containing 20 different Pheno-Seq particles in the pool).

[0160] [Table 1]

[0161] If the number of wells containing single cells is not clearly distinguishable from one another, additional particles can be sequentially included in multiple batches of 20 particles with unique index sequences, with the goal of a target average number of 8 that are distinguishable in images captured by the Rhapsody Imager. Results from the sequential inclusion of cells or beads of different sizes demonstrate that the Rhapsody cartridge supports the inclusion of cells and micrometer-sized particles in a sequential manner that can be imaged multiple times by the scanner.

[0162] To obtain 20 beads, we need a combination of sizes (3, 7, 10, and 16 micron spherical beads), fluorescence intensities (none, two different intensities of green and red (dark / bright), and red and green double positives, at least 5 clearly detectable by the scanner). Total = 4 (sizes) x 5 (fluorescence combinations) = 20 beads.

[0163] The oligonucleotides tagged onto the Pheno-Seq particles are specifically designed to be amplified only when cells are present in the wells, not when cells are absent. To achieve this, cells are tagged with beads, such as I-Mag beads, but also contain unique indexed oligonucleotides. These beads contain oligonucleotides with unique indexes, but attached to the beads with or without disulfide bonds. These oligos contain a common I-Mag-AbSeq primer region, followed by a unique index sequence, a poly(A), which base pairs with the poly(T) of the Cell Capture Bead oligo and has a consensus sequence complementary to the 3' end of the Pheno-Seq oligo.

[0164] III. Technical Effects

[0165] We will expand the current capabilities of the Rhapsody cartridge and scanner to function as a single-cell functional assay platform. We will identify genotypic, transcriptomic, and epitope information related to cell function and phenotype. Each Rhapsody cartridge contains approximately 200,000 wells, and the current Rhapsody workflow recommends including 5,000–10,000 single cells per cartridge, representing 2.5–5% of the total number of wells. We then convert the locations of these 5,000–10,000 single cells within the wells into unique barcode indexes that can be identified using sequence data acquired by the Rhapsody workflow. By combining these barcode indexes with CITE-seq data from each single cell, we will identify genotypic, transcriptomic, and associated epitope information related to cell function and phenotype at the single-cell level.

[0166] Combining protein expression data with transcriptome data contributes to the reliability of information obtained from single cells. Multiple genes, post-transcription factors, and signaling pathways also regulate cellular function. Understanding the relationship between omics data and cellular function and phenotype is valuable for deepening understanding and establishing strategies in translational research, including identifying novel biomarkers, developing diagnostic assays, and exploring novel therapeutics and clinical solutions.

[0167] Tumor-associated immune cells, such as CD8 T cells, are a good example. These cells share transcriptome and epitope information, but can differ functionally in terms of multi-functionality (degree and range of cytokine secretion), proliferative capacity, and target cell lytic efficacy. CD8 T cell subsets have very specific functions, such as cell killing and immunoregulation. CD8 + Polyfunctional T cells (secreting more than two cytokines) are more effective in immune control of cancer and infectious diseases and are associated with a better prognosis. TCR sequence information, along with cytokine profile and transcriptome data, is important for the development of novel immunotherapies. Real-time assessment of single cells in response to antigenic and other stimuli is valuable for assessing the effectiveness of immune responses. This also applies to CD4 T cells. Highly proliferative, polyfunctional, and long-term viable CAR-T cells are associated with a better prognosis in oncology management. Similarly, macrophages can share transcriptome and epitope information but differ in the secretion of certain cytokines, growth factors, and phagocytic activity.

[0168] On the other hand, characterizing cancer cells that are resistant to treatment (including cell therapy) is important for developing better strategies. This platform can characterize single cells derived from clinical specimens and can also determine prognosis and response to treatment.

[0169] This is achieved by extending the current capabilities of the Rhapsody cartridge and scanner to leverage it as a platform for performing single-cell functional assays and capturing cellular phenotypes, and then including Pheno-Seq particles to identify genomic, transcriptomic, and epitope information associated with cellular function and phenotype with minimal modifications to the Rhapsody workflow.

[0170] The method of the present invention finds use in a variety of different applications, including but not limited to:

[0171] 1. Characterizing cells based on their function is crucial for distinguishing polyfunctional antigen-specific CD4 and CD8 T cells and their associated T cell receptor (TCR) sequences and transcriptome profiles.

[0172] 2. This is also important for other applications, such as characterization of CAR-T cells, T regulatory cells, tumor-associated macrophages, and NK cells.

[0173] 3. Other applications include characterization of differentiated cells (based on their phenotype and function) from stem cell (cancer, regenerative medicine, etc.) populations, which will be extremely useful in understanding developmental pathways in health and disease, primarily cancer onset, prognosis, and outcome.

[0174] 4. The inclusion of functional and phenotypic information along with CITE-seq will broaden the scope of applications for biomarker development.

[0175] 5. Help characterize B cells, identify the antibody sequences responsible for virus neutralization, antibody-dependent cellular cytotoxicity (ADCC), and determine their efficacy.

[0176] 6. This method supports drug development efforts by allowing the evaluation of the efficacy of small molecules or therapeutic agents in multiple patient samples in a single experiment, helping to correlate the functional efficacy of novel therapeutics and small molecules across multiple donors and identifying genotypes, transcriptomes, and epitopes associated with efficacy. Each patient sample is identified by sample tag sequencing, and cells are dispensed into Rhapsody wells for single cell acquisition and treatment with the small molecule of interest. A single cartridge can accommodate 5,000–10,000 single cells, supporting efficacy evaluation of 50–100 cells from 100 patients. Furthermore, samples from multiple Rhapsody cartridges can be combined and included in a single RNA-seq run.

[0177] Notwithstanding the appended claims, the present disclosure is also defined by the following clauses.

[0178] 1. A method for obtaining linked functional and sequence data for a single cell of a cell sample, the method comprising: contacting cells of the cell sample with nucleic acid-barcoded cell-binding beads to produce bead-bound cells; dividing the bead-bound cells to generate divided bead-bound single cells, each bead-bound single cell being stably bound to a nucleic acid-barcoded cell-binding bead(s); functionally assaying the divided bead-bound single cells to obtain functional data for the divided bead-bound single cells; introducing unique combinations of different nucleic acid-barcoded identification particles into the partitions containing the functionally assayed bead-bound single cells to generate indexed partitions comprising: Functionally assayed bead-bound single cells; and Unique combination of nucleic acid-barcoded identification particles; obtaining image data of the indexed partitions and identifying the unique combinations of nucleic acid-barcoded identification particles contained therein; obtaining sequence data of a nucleic acid containing a barcode present in the partition; and Obtaining linked functional and sequence data for a single cell of the cell sample from the image data and the sequence data.

[0179] 2. The method of clause 1, wherein the nucleic acid-barcoded cell-binding beads comprise: beads; a specific binding component; and Cell-binding bead nucleic acid containing barcode.

[0180] 3. The method of claim 2, wherein the cell-bound bead nucleic acid further comprises: a first domain complementary to the target binding region of the nucleic acid capture bead; and A second domain complementary to a sequence present in the identification particle nucleic acid in said nucleic acid-barcoded identification particle.

[0181] 4. The method of any of paragraphs 2 and 3, wherein the beads are magnetic.

[0182] 5. The method of any preceding clause, wherein the specific binding moiety specifically binds to a cell surface marker.

[0183] 6. The method of paragraph 5, wherein the specific binding member comprises an antibody or a binding fragment thereof.

[0184] 7. The method of any preceding clause, wherein said dividing comprises distributing said bead-bound cells into partitions.

[0185] 8. The method of paragraph 7, wherein the distributing comprises introducing the bead-bound cells into a flow cell having microwells on the bottom surface.

[0186] 9. The method of any preceding clause, wherein functionally assaying the split, bead-bound single cells to obtain functional data for the split, bead-bound single cells comprises evaluating the split, bead-bound single cells over time.

[0187] 10. The method of any preceding clause, wherein functionally assaying the divided, bead-bound single cells to obtain functional data for the divided, bead-bound single cells comprises evaluating the divided, bead-bound single cells in response to a stimulus.

[0188] 11. The method of clause 10, wherein the stimulus is selected from the group consisting of a chemical stimulus, a mechanical stimulus, a physical stimulus, or a combination thereof.

[0189] 12. The method of any preceding clause, wherein the unique combination of different nucleic acid-barcoded identification particles comprises a plurality of different nucleic acid-barcoded identification particles that differ from one another in one or more of size, color, and brightness.

[0190] 13. The method according to any of the preceding paragraphs, wherein the number of different nucleic acid-barcoded identification particles constituting the combination in a partition ranges from 1 to 5.

[0191] 14. The method of any of the preceding clauses, wherein the different nucleic acid-barcoded identification particles range in size from 3 to 20 micrometers.

[0192] 15. The method of any of the preceding paragraphs, wherein the different nucleic acid-barcoded identification particles have a color selected from the group consisting of green, red, blue, gray, yellow, and black.

[0193] 16. The method of any preceding clause, wherein introducing unique combinations of different nucleic acid-barcoded identification particles to the partitions containing the functionally assayed bead-bound single cells comprises introducing the different compositions of nucleic acid-barcoded identification particles to a flow cell having microwells on its bottom surface, wherein the microwells contain the functionally assayed bead-bound single cells.

[0194] 17. The method of paragraph 16, wherein the composition of different nucleic acid-barcoded identification particles comprises between 2 and 15 different nucleic acid-barcoded identification particles.

[0195] 18. The method of clause 16 and clause 17, wherein the method further comprises introducing a second composition of different nucleic acid-barcoded identification particles into the flow cell.

[0196] 19. A method according to any preceding clause, wherein the sequencing comprises providing beads comprising bead-bound nucleic acids in the partition comprising bead-bound single cells, the bead-bound nucleic acids comprising a cell label domain and a target binding region.

[0197] 20. The method of paragraph 19, wherein the bead-bound nucleic acid further comprises one or more of a molecular index domain and a universal primer binding domain.

[0198] 21. The method of any of the preceding clauses, wherein obtaining sequence data of the split, combined barcoded single cells comprises employing a next-generation sequencing protocol.

[0199] 22. The method of claim 21, wherein the next generation sequencing protocol includes generating a sequence ready library.

[0200] 23. The method of claim 22, wherein generating the sequence-ready library comprises a reverse transcription step and an amplification step.

[0201] 24. The method of claim 23, wherein the amplifying step comprises using a first primer set and a second primer set, the first primer set for amplifying a nucleic acid comprising a cell label barcode, and the second primer set for amplifying a nucleic acid of a different nucleic acid-barcoded identification particle.

[0202] 25. The method of any preceding clause, wherein the sequence data comprises multiomic data.

[0203] 26. A composition of a plurality of different nucleic acid-barcoded identification particles, wherein the plurality of different nucleic acid-barcoded identification particles differ from one another in one or more of size, color, and brightness.

[0204] 27. The composition of paragraph 26, wherein the composition of different nucleic acid-barcoded identification particles comprises between 2 and 15 different nucleic acid-barcoded identification particles.

[0205] 28. The composition of clause 26 and clause 27, wherein the different nucleic acid-barcoded identifier particles range in size from 3 to 20 micrometers.

[0206] 29. The composition of any one of clauses 26 to 28, wherein the different nucleic acid-barcoded identification particles have a color selected from the group consisting of green, red, blue, gray, yellow, and black.

[0207] 30. A kit for obtaining linked functional and sequence data of single cells of a cell sample, comprising: a composition of a plurality of different nucleic acid-barcoded identification particles, wherein the plurality of different nucleic acid-barcoded identification particles differ from one another in one or more of size, color, and brightness; and Nucleic acid-barcoded cell-binding beads.

[0208] 31. The kit of paragraph 30, wherein the composition of different nucleic acid-barcoded identification particles comprises between 2 and 15 different nucleic acid-barcoded identification particles.

[0209] 32. The kit of paragraph 30 and paragraph 31, wherein the different nucleic acid-barcoded identifier particles are between 3 and 20 micrometers in size.

[0210] 33. The kit of clauses 30 to 32, wherein the different nucleic acid-barcoded identification particles have a color selected from the group consisting of green, red, blue, gray, yellow, and black.

[0211] 34. The kit of any of paragraphs 30 to 33, wherein the nucleic acid-barcoded cell-binding beads comprise the following components: beads; a specific binding component; and Cell-binding bead nucleic acid containing barcode.

[0212] 35. The kit of paragraph 34, wherein the cell-binding bead nucleic acid further comprises: a first domain complementary to the target binding region of the nucleic acid capture bead; and A second domain complementary to a sequence present in the identification particle nucleic acid in said nucleic acid-barcoded identification particle.

[0213] 36. The kit of either paragraph 34 or paragraph 35, wherein the beads are magnetic.

[0214] 37. The kit of any of the preceding paragraphs, wherein the specific binding moiety specifically binds to a cell surface marker.

[0215] 38. The kit of paragraph 37, wherein the specific binding moiety comprises an antibody or binding fragment thereof.

[0216] 39. The kit of any one of paragraphs 30 to 38, comprising: The kit further comprises beads comprising a bead-bound nucleic acid comprising a cell label domain and a target binding region.

[0217] 40. The kit of any of paragraphs 30 to 39, wherein the kit further comprises a flow cell having a microwell on its bottom surface.

[0218] Although the foregoing invention has been described in some detail by way of illustration and example, for purposes of clarity of understanding, it will be apparent to those skilled in the art that, in light of the teachings of the invention, certain changes and modifications can be made without departing from the spirit or scope of the appended claims.

[0219] Thus, the foregoing merely describes the principles of the present invention. Those skilled in the art will recognize that, even if not explicitly described or shown, they can devise various configurations that embody the principles of the present invention and are within its spirit and scope. Furthermore, all examples and conditional language described herein are primarily intended to aid the reader in understanding the concepts of the inventors who contributed to the development of the principles and technology of the present invention, and should not be construed as being limited to the specifically described examples and conditions. Furthermore, all statements herein that recite principles, aspects, and embodiments of the present invention, as well as specific examples thereof, are intended to encompass both structural and functional equivalents. Furthermore, such equivalents are intended to include both currently known equivalents and equivalents developed in the future, i.e., elements developed that perform the same function, regardless of structure. Furthermore, the subject matter disclosed herein is not intended to be appropriated to the public, regardless of whether such disclosure is expressly recited in the claims.

[0220] Accordingly, the scope of the present invention is not limited to the embodiments shown and described herein. Rather, the scope and spirit of the present invention is embodied by the appended claims. In the claims, 35 USC §112(f) or 35 USC §112(6) is expressly defined to apply to a claim limitation only if the precise phrase "means for" or "step" appears in the preface of that claim limitation. If that precise phrase is not used in a claim limitation, 35 USC §112(f) or 35 USC §112(6) does not apply.

Claims

1. 1. A method for obtaining linked functional and sequence data of a single cell of a cell sample, the method comprising: contacting cells of the cell sample with nucleic acid-barcoded cell-binding beads to produce bead-bound cells; dividing the bead-bound cells to generate divided bead-bound single cells, each bead-bound single cell being stably bound to a nucleic acid-barcoded cell-binding bead(s); functionally assaying the divided bead-bound single cells to obtain functional data for the divided bead-bound single cells; introducing unique combinations of different nucleic acid-barcoded identification particles into the partitions containing the functionally assayed bead-bound single cells to generate indexed partitions comprising: Functionally assayed bead-bound single cells; and unique combination of nucleic acid-barcoded identification particles; acquiring image data of the indexed partition and identifying the unique combination of nucleic acid-barcoded identification particles contained therein; obtaining sequence data of a nucleic acid containing a barcode present in the partition; and Obtaining linked functional and sequence data for a single cell of the cell sample from the image data and the sequence data.

2. 10. The method of claim 1, wherein the nucleic acid-barcoded cell-binding beads comprise: beads; a specific binding component; and Cell-binding bead nucleic acid containing barcode.

3. 3. The method of claim 2, wherein the cell-binding bead nucleic acid further comprises: a first domain complementary to the target binding region of the nucleic acid capture bead; and A second domain complementary to a sequence present in the identification particle nucleic acid in the nucleic acid-barcoded identification particle.

4. The method according to any one of claims 2 and 3, wherein the beads are magnetic.

5. 10. The method of any preceding claim, wherein the specific binding moiety specifically binds to a cell surface marker.

6. 6. The method of claim 5, wherein the specific binding component comprises an antibody or a binding fragment thereof.

7. 10. The method of claim 9, wherein the dividing step comprises distributing the bead-bound cells into partitions.

8. 8. The method of claim 7, wherein the dispensing comprises introducing the bead-bound cells into a flow cell having microwells on its bottom surface.

9. 10. The method of any preceding claim, wherein functionally assaying the split, bead-bound single cells to obtain functional data for the split, bead-bound single cells comprises evaluating the split, bead-bound single cells over time.

10. 10. A method according to any preceding claim, wherein functionally assaying the divided bead-bound single cells to obtain functional data for the divided bead-bound single cells comprises evaluating the divided bead-bound single cells in response to a stimulus.

11. 10. The method of any preceding claim, wherein the unique combination of different nucleic acid-barcoded identification particles comprises a plurality of different nucleic acid-barcoded identification particles that differ from each other in one or more of size, color, and brightness.

12. 10. The method of any preceding claim, wherein introducing unique combinations of different nucleic acid-barcoded identification particles into partitions containing functionally assayed bead-bound single cells comprises introducing compositions of different nucleic acid-barcoded identification particles into a flow cell having microwells on its bottom surface, wherein the microwells contain functionally assayed bead-bound single cells.

13. A method according to any preceding claim, wherein the sequencing comprises providing beads comprising bead-bound nucleic acids in the partition comprising bead-bound single cells, the bead-bound nucleic acids comprising a cell label domain and a target binding region.

14. 10. The method of any preceding claim, wherein obtaining sequence data of the split combined barcoded single cells comprises employing a next generation sequencing protocol.

15. 10. The method of any preceding claim, wherein the sequence data comprises multiomic data.

16. A composition of a plurality of different nucleic acid-barcoded identification particles, wherein the plurality of different nucleic acid-barcoded identification particles differ from one another in one or more of size, color, and brightness.

17. 1. A kit for obtaining linked functional and sequence data of a single cell of a cell sample, comprising: a composition of a plurality of different nucleic acid-barcoded identification particles, wherein the plurality of different nucleic acid-barcoded identification particles differ from one another in one or more of size, color, and brightness; and Nucleic acid-barcoded cell-binding beads.