Methods and compositions for obtaining linked image and sequence data for single cells

By using specific binding members in single cells/オリゴヌクレオチドブーード in single cells and obtaining the correlation between their image and sequence data, the problem of difficulty in linking single-cell imaging data and second-generation sequencing data in the prior art is solved, and high-throughput data acquisition in single-cell multiomics applications is achieved.

JP2025514783APending Publication Date: 2025-05-09BECTON DICKINSON & CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024561944
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-04-18
Filing Date
2023-04-12
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The prior art is difficult to effectively link single-cell imaging data with large-scale parallel secondary sequencing data, resulting in the inability to achieve high-throughput and efficient data acquisition in single-cell multiomics applications.

Method used

By using a specific binding member / オリゴヌクレオチドノーードドードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードードー� Then, the image and sequence data of the single cells are obtained by acquiring the image and sequence data of these single cells and correlating the two by sharing the same combinatorial barcodes.

Benefits of technology

It realizes effective correlation between single-cell image and sequence data, supports high-throughput data acquisition in single-cell multiomics applications, and improves data efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025514783000001_ABST
    Figure 2025514783000001_ABST
Patent Text Reader

Abstract

Aspects of the invention include methods of obtaining linked image and sequence data for, for example, single cells of a cell sample. An embodiment of the method includes combinatorially barcoding cells obtained, for example, from an initial cell sample, with specific binding members / oligonucleotide sub-barcodes to generate combinatorially barcoded cells. The resulting combinatorially barcoded cells are then partitioned to generate partitioned combinatorially barcoded single cells, each having a combinatorial barcode. Image data and sequence data are then obtained for the partitioned combinatorially barcoded single cells, and subsequently, image data and sequence data sharing a common combinatorial barcode are joined to obtain linked image and sequence data for the single cells of the cell sample. Also provided are compositions for carrying out the methods of the invention.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] Current technology allows gene expression of single cells to be measured in a massively parallel manner (e.g., >10,000 cells) by appending cell-specific oligonucleotide barcodes to poly(A) mRNA molecules from individual cells, as each of the cells colocalizes with a barcoded reagent bead within a compartment.

[0002] One platform that allows for measuring gene expression in a massively parallel manner in single cells is the BD Rhapsody™ Single-Cell Analysis System. The BD Rhapsody™ Single-Cell Analysis System is a platform that allows for high-throughput capture of nucleic acids from single cells using a simple cartridge workflow and a multi-layer barcode system. The resulting captured information can be used to generate various types of next-generation sequencing (NGS) libraries, including libraries suitable for whole-transcriptome analysis for, for example, discovery biology and targeted RNA analysis for sensitive transcript detection. Shum et al., “Quantitation of mRNA Transcripts and Proteins Using the BD Rhapsody™ Single-Cell Analysis System,” Adv Exp Med Biol. 2019;1129:63-79.

[0003] Gene expression can affect protein expression. Protein-protein interactions can affect gene expression and protein expression. Therefore, more recently, systems and methods have been developed that can quantitatively analyze protein expression in cells and simultaneously measure protein expression and gene expression in cells. One such platform is the BD Abseq platform. AbSeq is a method for profiling proteins in single cells. In Abseq, the usual fluorophore-labeled antibodies are replaced with nucleic acid sequence tags, which can be read out at the single cell level, for example, via barcoding and NGS sequencing. "The goal of Abseq is to enable sensitive, accurate and comprehensive characterization of proteins in large numbers of single cells. Cells are bound with antibodies against different target epitopes, similar to traditional immunostaining, except that the antibodies are labeled with unique sequence tags. If the antibody binds to its target, a DNA tag is carried along with it, allowing the presence of the target to be inferred based on the presence of the tag. In this way, counting the tags gives an estimate of the different epitopes present in the cell, detected via antibody binding." Shahi et al., "Abseq: Ultrahigh-throughput single cell protein profiling with droplet microfluidic barcoding. Sci Rep 7, 44447 (2017)". Summary of the Invention

[0004] The inventors have recognized that it is desirable to link image data to massively parallelized NGS data in single cell analysis, including single cell multi-omics applications. The inventors are not aware of any current protocols that exist for linking single cell imaging data from the same cells to single cell multi-omics data. One can first perform single cell sorting (FACS) of cells into macro-well plates (such as 96 wells) and then a plate-based single cell multi-omics workflow on the sorted cells, but this does not provide image data linked to NGS data for the cells. Plate-based workflows do not provide the same throughput or efficiency as massively parallel single cell multi-omics workflows. Furthermore, the indexed data is not microscope-based, and data from common flow cytometers currently used lacks two-dimensional (spatial) information. The present embodiments fulfill the need in the art for methods and compositions for easily obtaining linked image data and sequence data for single cells.

[0005] Aspects of the invention include methods of obtaining linked image and sequence data for, for example, single cells of a cell sample. An embodiment of the method includes combinatorially barcoding cells obtained, for example, from an initial cell sample, with specific binding members / oligonucleotide sub-barcodes to generate combinatorially barcoded cells. The resulting combinatorially barcoded cells are then partitioned to generate partitioned combinatorially barcoded single cells, each having a combinatorial barcode. Image data and sequence data are then obtained for the partitioned combinatorially barcoded single cells, followed by concatenating the image and sequence data sharing a common combinatorial barcode to obtain concatenated image and sequence data for the single cells of the cell sample. Also provided are compositions for carrying out the methods of the invention. [Brief description of the drawings]

[0006] The invention may be best understood from the following detailed description when read in conjunction with the accompanying drawing figures, which include:

[0007] [Figure 1-1] 1 illustrates a schematic of splitting / pooling individual cells and sample indexing according to an embodiment of the present invention. [Figure 1-2] Continuing with FIG. 1-1, individual cell splitting / pooling and sample indexing are shown diagrammatically in accordance with an embodiment of the present invention. [Diagram 2] FIG. 13 shows a schematic of split / pool ab-oligo labeling of cells to generate a single cell index that can be read out by imaging and downstream single cell multi-omics, according to an embodiment of the present invention. [Diagram 3] 1 provides an example of using cyclic immunofluorescence to decode ab oligo signatures of individual cells, according to an embodiment of the present invention.

[0008] definition Unless otherwise defined, technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this disclosure belongs.See, for example, Singleton et al., Dictionary of Microbiology and Molecular Biology 2nd ed., J. Wiley & Sons (New York, NY 1994); Sambrook et al., Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Press (Cold Spring Harbor, NY 1989).For purposes of this disclosure, the following terms are defined below.

[0009] As used herein, an antibody can be a full-length (e.g., naturally occurring or formed by recombinant processes of normal immunoglobulin gene fragments) immunoglobulin molecule (e.g., an IgG antibody) or an immunologically active (i.e., specifically binding) portion of an immunoglobulin molecule (such as an antibody fragment). In some embodiments, an antibody is a functional antibody fragment. For example, an antibody fragment can be a portion of an antibody, such as F(ab')2, Fab', Fab, Fv, sFv, etc. An antibody fragment can bind to the same antigen recognized by a full-length antibody. An antibody fragment can include isolated fragments consisting of the variable regions of an antibody, such as an "Fv" fragment consisting of the variable regions of a heavy and light chain, and recombinant single-chain polypeptide molecules in which the variable regions of a light and heavy chain are connected by a peptide linker ("scFv protein"). Exemplary antibodies include, but are not limited to, antibodies against cancer cells, antibodies against viruses, antibodies that bind to cell surface receptors (e.g., CD8, CD34, and CD45), and therapeutic antibodies.

[0010] As used herein, the term "associated" or "related to" may mean that two or more species are identifiable as coexisting at a time. Association may mean that two or more species are in or were in similar containers. Association may be informatic association. For example, digital information about two or more species can be stored and used to determine that one or more of the species were coexisting at a time. Association may be physical association. In some embodiments, two or more related species are "tethered," "bound," or "immobilized" to each other or to a common solid or semi-solid surface. Association may refer to a covalent or non-covalent means for attaching a label to a solid or semi-solid support such as a bead. Association may be a covalent bond between a target and a label. Association may include hybridization between two molecules (e.g., a target molecule and a label).

[0011] As used herein, the term "complementary" may refer to the ability for exact pairing between two nucleotides. For example, if a nucleotide at a given position of a nucleic acid can hydrogen bond with a nucleotide of another nucleic acid, the two nucleic acids are considered to be complementary to each other at that position. Complementarity between two single-stranded nucleic acid molecules may be "partial," where only some of the nucleotides bind, or may be complete, where there is complete complementarity between the single-stranded molecules. If a first nucleotide sequence is complementary to a second nucleotide sequence, the first nucleotide sequence may be said to be "complementary" to the second nucleotide sequence. If a first nucleotide sequence is complementary to a sequence that is the reverse of the second sequence (i.e., the order of the nucleotides is reversed), the first nucleotide sequence may be said to be the "reverse complement" of the second sequence. As used herein, the terms "complement," "complementary," and "reverse complement" may be used interchangeably. It is understood from this disclosure that when a molecule is capable of hybridizing to another molecule, it may be the complement of the hybridizing molecule.

[0012] As used herein, the term "nucleic acid" refers to a polynucleotide sequence or a fragment thereof. A nucleic acid may comprise nucleotides. A nucleic acid may be exogenous or endogenous to a cell. A nucleic acid may be present in a cell-free environment. A nucleic acid may be a gene or a fragment thereof. A nucleic acid may be DNA. A nucleic acid may be RNA. A nucleic acid may comprise one or more analogs (e.g., modified backbones, sugars, or nucleobases). Some non-limiting examples of analogs include 5-bromouracil, peptide nucleic acid, xenonucleic acid, morpholinos, locked nucleic acids, glycol nucleic acids, threose nucleic acids, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein linked to a sugar), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogs, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, queousine, and wyosine. "Nucleic acid," "polynucleotide," "target polynucleotide," and "target nucleic acid" may be used interchangeably.

[0013] Nucleic acids may include one or more modifications (e.g., base modifications, backbone modifications) to provide nucleic acids with new or enhanced characteristics (e.g., improved stability). Nucleic acids may include a nucleic acid affinity tag. Nucleosides may be base-sugar combinations. The base portion of the nucleoside may be a heterocyclic base. The two most common classes of such heterocyclic bases are purines and pyrimidines. Nucleotides may be nucleosides that further include a phosphate group covalently linked to the sugar portion of the nucleoside. For those nucleosides that include a pentofuranosyl sugar, the phosphate group may be linked to the 2', 3', or 5' hydroxyl moiety of the sugar. In forming nucleic acids, the phosphate group may covalently link adjacent nucleosides to each other to form a linear polymeric compound. The respective ends of this linear polymeric compound may then be further linked to form a circular compound, although linear compounds are generally preferred. In addition, linear compounds may have internal nucleotide base complementarity and therefore may fold to produce fully or partially double-stranded compounds. Within nucleic acids, the phosphate groups may be commonly referred to as forming the internucleoside backbone of the nucleic acid. The linkage or backbone may be a 3' to 5' phosphodiester bond.

[0014] The nucleic acids may contain modified backbones and / or modified internucleoside linkages. Modified backbones include those that retain a phosphorus atom in the backbone and those that do not have a phosphorus atom in the backbone. Suitable modified nucleic acid backbones containing a phosphorus atom therein include, for example, phosphorothioates, chiral phosphorothioates, phosphodithioates, phosphotriesters, aminoalkyl phosphotriesters, methyl and other alkyl phosphonates (e.g., 3'-alkylene phosphonates, 5'-alkylene phosphonates), chiral phosphonates, phosphinates, phosphoramidates (including 3'-amino phosphoramidates and aminoalkyl phosphoramidates), phosphorodiamidates, thionophosphoramidates, thionoalkylphosphonates, thionoalkylphosphotriesters, selenophosphates, and boranophosphates having normal 3'-5' linkages, 2'-5' linkage analogs, and those with inverted polarity in which one or more internucleotide linkages are 3' to 3', 5' to 5', or 2' to 2' linkages.

[0015] Nucleic acids may contain polynucleotide backbones formed by short chain alkyl or cycloalkyl internucleoside linkages, mixed heteroatom and alkyl or cycloalkyl internucleoside linkages, or one or more short chain heteroatom or heterocyclic internucleoside linkages. These may include those with morpholino linkages (formed in part from the sugar portion of the nucleoside), siloxane backbones, sulfide, sulfoxide, and sulfone backbones, formacetyl and thioformacetyl backbones, methyleneformacetyl and thioformacetyl backbones, riboacetyl backbones, alkene-containing backbones, sulfamate backbones, methyleneimino and methylenehydrazino backbones, sulfonate and sulfonamide backbones, amide backbones, and others with mixed N, O, S, and CH2 constituent moieties.

[0016] Nucleic acids may include nucleic acid mimetics. The term "mimetics" may be intended to include polynucleotides in which only the furanose ring or both the furanose ring and the internucleotide linkage are replaced with non-furanose groups, and replacement of only the furanose ring may also be referred to as sugar surrogate sugar. The heterocyclic base moiety or modified heterocyclic base moiety may be maintained for hybridization with an appropriate target nucleic acid. One such nucleic acid may be a peptide nucleic acid (PNA). In a PNA, the sugar backbone of the polynucleotide may be replaced with an amide-containing backbone, in particular an aminoethylglycine backbone. The nucleotides may be retained and are directly or indirectly bound to the aza nitrogen atoms of the amide portion of the backbone. The backbone in a PNA compound may contain two or more linked aminoethylglycine units, giving the PNA an amide-containing backbone. The heterocyclic base moiety may be directly or indirectly bound to the aza nitrogen atoms of the amide portion of the backbone.

[0017] The nucleic acid may include a morpholino backbone structure. For example, the nucleic acid may include a six-membered morpholino ring instead of a ribose ring. In some of these embodiments, phosphorodiamidate or other non-phosphodiester internucleoside linkages may replace the phosphodiester linkage.

[0018] Nucleic acids may include linked morpholino units (e.g., morpholino nucleic acids) having heterocyclic bases attached to the morpholino ring. Linking groups can link morpholino monomer units in morpholino nucleic acids. Nonionic morpholino-based oligomeric compounds can have less undesirable interactions with cellular proteins. Morpholino-based polynucleotides can be nonionic mimics of nucleic acids. Various compounds within the morpholino class can be linked using different linking groups. A further class of polynucleotide mimics can be referred to as cyclohexenyl nucleic acids (CeNAs). The furanose rings normally present in nucleic acid molecules can be replaced with cyclohexenyl rings. CeNA DMT-protected phosphoramidite monomers can be prepared and used to synthesize oligomeric compounds using phosphoramidite chemistry. Incorporation of CeNA monomers into nucleic acid strands can increase the stability of DNA / RNA hybrids. CeNA oligoadenylates can form complexes with nucleic acid complements with stability similar to the natural complexes. Further modifications may include Locked Nucleic Acids (LNAs) in which a 2'-hydroxyl group is attached to the 4' carbon atom of the sugar ring, thereby forming a 2'-C, 4'-C-oxymethylene linkage, thereby forming a bicyclic sugar moiety. The linkage may be a methylene (-CH2), a group bridging the 2' oxygen atom and the 4' carbon atom, where n is 1 or 2. LNAs and LNA analogs may exhibit very high duplex thermal stability with complementary nucleic acids (Tm=+3 to +10°C), stability against 3'-exonucleolytic degradation, and good solubility properties.

[0019] Nucleic acids may also include modifications or substitutions of nucleobases (often simply referred to as "bases"). As used herein, "unmodified" or "natural" nucleobases may include purine bases (e.g., adenine (A) and guanine (G)) and pyrimidine bases (e.g., thymine (T), cytosine (C), and uracil (U)). Modified nucleobases include other synthetic and natural nucleobases, such as 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyl (-C=C-CH3) uracil and other alkynyl derivatives of cytosine and pyrimidine bases, 6-azouracil, These may include cytosine and thymine, 5-uracil (pseudouracil), 4-thiouracil, 8-halo, 8-amino, 8-thioalkyl, 8-hydroxyl, and other 8-substituted adenines and guanines, 5-halo (especially 5-bromo), 5-trifluoromethyl, and other 5-substituted uracils and cytosines, 7-methylguanine and 7-methyladenine, 2-F-adenine, 2-aminoadenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine, and 3-deazaguanine and 3-deazaadenine.Modified nucleobases include tricyclic pyrimidines, such as phenoxazine cytidine (1H-pyrimido(5,4-b)(1,4)benzoxazin-2(3H)-one), phenothiazine cytidine (1H-pyrimido(5,4-b)(1,4)benzoxazin-2(3H)-one), G clamps, such as substituted phenoxazine cytidines (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzoxazin-2(3H)-one), phenothiazine cytidine and cytidines such as 1H-pyrimido(5,4-b)(1,4)benzothiazin-2(3H)-ones), G clamps, such as substituted phenoxazine cytidines (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzoxazin-2(3H)-ones), carbazole cytidines (2H-pyrimido(4,5-b)indol-2-ones), pyridoindole cytidines (H-pyrido(3',2':4,5)pyrrolo[2,3-d]pyrimidin-2-ones).

[0020] As used herein, the term "sample" may refer to a composition that contains a target. Samples suitable for analysis by the disclosed methods, devices, and systems include cells, tissues, organs, or organisms. A cell sample is a composition that is comprised of a plurality of cells (the number of cells may vary), such as a composition that includes a plurality of different cells, such as an aqueous composition of a single cell.

[0021] As used herein, the term "sampling device" or "device" may refer to a device that can take a portion of a sample and / or place the portion on a substrate. A sample device may refer to, for example, a fluorescence activated cell sorting (FACS) machine, a cell sorter, a biopsy needle, a biopsy device, a tissue sectioning device, a microfluidic device, a blade grid, and / or a microtome.

[0022] As used herein, the term "solid support" may refer to a discrete solid or semi-solid surface to which a nucleic acid may be attached. A solid support may include any type of solid, porous, or hollow sphere, ball, bearing, cylinder, or other similar configuration made of plastic, ceramic, metal, or polymeric material (e.g., hydrogel) to which a nucleic acid may be immobilized (e.g., covalently or non-covalently). A solid support may be spherical (e.g., microsphere) or may include discrete particles having a non-spherical or irregular shape (such as a cube, rectangular prism, pyramid, cylinder, cone, oval, or disk). The shape of a bead may be non-spherical. A plurality of solid supports spaced apart in an array may not include a substrate. A solid support may be used interchangeably with the term "bead."

[0023] As used herein, the term "target" may refer to a composition that may be analyzed according to embodiments of the present invention. Exemplary suitable targets for analysis by the disclosed methods, devices, and systems include oligonucleotides, DNA, RNA, mRNA, microRNA, tRNA, and the like. Targets may be single-stranded or double-stranded. In some embodiments, targets may be proteins, peptides, or polypeptides. In some embodiments, targets are lipids. As used herein, "target" may be used interchangeably with "species."

[0024] As used herein, the term "reverse transcriptase" may refer to a group of enzymes that have reverse transcriptase activity (i.e., catalyze the synthesis of DNA from an RNA template). In general, such enzymes include, but are not limited to, retroviral reverse transcriptases, retrotransposon reverse transcriptases, retroplasmid reverse transcriptases, retron reverse transcriptases, bacterial reverse transcriptases, group II intron-derived reverse transcriptases, and mutants, variants, or derivatives thereof. Non-retroviral reverse transcriptases include non-LTR retrotransposon reverse transcriptases, retroplasmid reverse transcriptases, retron reverse transcriptases, and group II intron reverse transcriptases. Examples of group II intron reverse transcriptases include Lactococcus lactis LI.LtrB intron reverse transcriptase, Thermosynechococcus elongatus TeI4c intron reverse transcriptase, or Geobacillus stearothermophilus GsI-IIC intron reverse transcriptase. Other classes of reverse transcriptases can include many classes of non-retroviral reverse transcriptases (ie, retrons, group II introns, and diversity-generating retroelements, among others). DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0025] Aspects of the invention include methods of obtaining linked image and sequence data for, for example, single cells of a cell sample. An embodiment of the method includes combinatorially barcoding cells obtained, for example, from an initial cell sample, with specific binding members / oligonucleotide sub-barcodes to generate combinatorially barcoded cells. The resulting combinatorially barcoded cells are then partitioned to generate partitioned combinatorially barcoded single cells, each having a combinatorial barcode. Image data and sequence data are then obtained for the partitioned combinatorially barcoded single cells, followed by concatenating the image and sequence data sharing a common combinatorial barcode to obtain concatenated image and sequence data for the single cells of the cell sample. Also provided are compositions for carrying out the methods of the invention.

[0026] Before describing the present invention in more detail, it is to be understood that this invention is not limited to particular embodiments described, as such may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting, since the scope of the present invention will be limited only by the appended claims.

[0027] Where a range of values ​​is provided, unless the context clearly indicates otherwise, it is understood that each intervening value is included, to the tenth of the unit of the lower limit, between the upper and lower limit of that range and any other stated or intervening value in that stated range. The upper and lower limits of these smaller ranges may independently be included in the smaller ranges and are also encompassed within the invention, subject to any specific excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included in the invention.

[0028] Certain ranges are presented herein with the term "about" preceding the numerical values. The term "about" is used herein to provide literal support for the exact number it precedes, as well as a number that is close to or approximately the number it precedes. In determining whether a number is close to or approximately a specifically recited number, the number that is close to or approximately the unrecited number may be a number that provides substantial equivalence to the specifically recited number in the context presented.

[0029] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. Although any methods and materials similar or equivalent to those described herein can also be used in the practice or testing of the present invention, representative exemplary methods and materials are described herein.

[0030] All publications and patents cited herein are incorporated by reference as if each individual publication or patent was specifically and individually indicated to be incorporated by reference, and are incorporated by reference herein to disclose and describe the methods and / or materials in connection with which the publications are cited. The citation of any publication is for its disclosure prior to the filing date and should not be construed as an admission that the present invention is not entitled to antedate such publication by virtue of prior invention. Further, the dates of publication provided may be different from the actual publication dates, which may need to be independently confirmed.

[0031] It should be noted that, as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. It should be further noted that the claims may be drafted to exclude any optional element. Thus, this statement is intended to serve as a predicate for the use of exclusive terminology such as "solely" and "only" in connection with the recitation of claim elements, or for the use of a "negative" limitation.

[0032] As will be apparent to those skilled in the art upon reading this disclosure, each of the separate embodiments described and illustrated herein has separate components and features which may be readily separated from or combined with the features of any of the other various embodiments without departing from the scope or spirit of the invention. Any recited method may be carried out in the order of events recited or in any other order which is logically possible.

[0033] Although the systems and methods have been described or will be described for grammatical fluidity with functional descriptions, it is expressly understood that the claims, unless expressly formulated under 35 U.S.C. ...

[0034] method As summarized above, methods are provided for obtaining concatenated image and sequence data of a single cell, e.g., an initial cell sample. Concatenated image and sequence data refers to a combined data set that includes both image data and nucleic acid sequence data, which may be attributed to the same cell, such that they may be considered to originate from the same cell. In other words, concatenated image and sequence data is a data set that includes both image data and nucleic acid sequence data obtained from the same cell. Image data is data obtained from a cell using imaging techniques. The term "image" is used in its conventional sense to refer to a representation of an object (e.g., a cell), e.g., generated by illumination via light. Image data can be data that collectively constitutes a representation and may be obtained using any convenient protocol. In some embodiments, the image data obtained by the methods of the present invention is microscopic image data. Microscopic image data refers to image data obtained using a microscope to observe objects and regions of objects (e.g., cells) that are not visible to the naked eye. Nucleic acid sequence data refers to data obtained using nucleic acid sequencing techniques that identify the sequence of nucleotides in a nucleic acid molecule. Nucleic acid sequencing data from a cell includes one or more nucleic acid sequences present in the cell, e.g., sequences of RNA molecules. Such data can be obtained using a variety of sequence protocols, including next generation sequencing (NGS) protocols.

[0035] As summarized above, aspects of the method include combinatorially barcoding cells of a cell sample with specific binding members / oligonucleotide sub-barcodes to generate combinatorially barcoded cells, partitioning the combinatorially barcoded cells to generate partitioned combinatorially barcoded single cells each having a combinatorial barcode, obtaining image data and sequence data for the partitioned combinatorially barcoded single cells, and concatenating image data and sequence data that share a common combinatorial barcode to obtain concatenated image and sequence data for a single cell of the cell sample. Embodiments of each of these steps are now described in more detail.

[0036] Combinatorial barcoding of cells in a cell sample using specific binding member / oligonucleotide sub-barcodes An embodiment of the method includes combinatorially barcoding cells of a cell sample with specific binding members / oligonucleotide sub-barcodes. By combinatorially barcoded cells, we mean that cells of an initial cell sample are modified to stably associate with a unique combination of sub-barcodes (provided by the specific binding member / oligonucleotide sub-barcode combination), which collectively constitutes the unique combinatorial barcode of that cell. By stably associated, we mean that the specific binding members / oligonucleotide sub-barcodes that constitute a given combinatorial barcode of a combinatorially barcoded cell are attached to the surface of that cell in such a manner that they do not dissociate from the cell during the conditions experienced by the cell under the method of the invention, e.g., as described in more detail below. In some cases, the stable association is provided by specific binding interactions, e.g., as described in more detail below. In combinatorial barcoded cells of embodiments of the invention, a unique combination of sub-barcodes is stably associated with a cell using a combinatorial protocol that associates a unique combination of specific binding member / oligonucleotide sub-barcodes with a given cell, the unique combination being obtained from an initial collection of specific binding member / oligonucleotide sub-barcodes. Combinatorial protocols used in embodiments of the invention include, for example, a split / pool protocol, as described in more detail below.

[0037] The sub-barcodes that collectively provide a unique combinatorial barcode to a cell are provided by specific binding members / oligonucleotide sub-barcodes. A specific binding member / oligonucleotide sub-barcode comprises a specific binding member component and an oligonucleotide sub-barcode component, which are stably associated with each other, for example, by a suitable bond or linking group (e.g., covalent bond). Thus, a specific binding member / oligonucleotide sub-barcode may be considered to have a specific binding member conjugated to an oligonucleotide sub-barcode component. Next, embodiments of each of these components are described in more detail.

[0038] The specific binding member component of the specific binding member / oligonucleotide sub-barcode used in the embodiments of the present invention may vary. The term "specific binding" refers to a direct association between two molecules by covalent, electrostatic, hydrophobic, and ionic and / or hydrogen bonding interactions, including interactions such as salt bridges and water bridges. The term "specific binding member" describes a member of a pair of molecules that has binding specificity for one another. Members of a specific binding pair may be naturally occurring or wholly or partially synthetically produced. One member of the pair of molecules has an area or cavity on its surface that specifically binds to, and is therefore complementary to, a particular spatial and polar configuration of the other member of the pair of molecules. Thus, the members of the pair have the property of specifically binding to one another. Examples of specific binding members of a pair are antigen-antibody, biotin-avidin, hormone-hormone receptor, receptor-ligand, enzyme-substrate. Specific binding members of a binding pair exhibit high affinity and binding specificity for binding to one another. Typically, the affinity between specific binding members of a pair is greater than 10 -6 M or less, e.g., 10 -7 M or less (10 -8 M or less), e.g., 10 -9 M or less, 10 -10 M or less, 10 -11 M or less, 10 -12 M or less, 10-13 M or less, 10 -14 M or less (10 -15 (including M and below) dA specific binding member is characterized by a dissociation constant (KD). "Affinity" refers to the strength of binding, and an increase in binding affinity correlates with a lower KD. In one embodiment, affinity is determined by surface plasmon resonance (SPR), for example, as used by Biacore systems. The affinity of one molecule to another is determined, for example, by measuring the on-rate of the interaction at 25°C. "Affinity" refers to the strength of binding, and an increase in binding affinity correlates with a lower KD. In one embodiment, affinity is determined by surface plasmon resonance (SPR), for example, as used by Biacore systems. The affinity of one molecule to another is determined, for example, by measuring the on-rate of the interaction at 25°C. The specific binding member can be a variety of. Examples of specific binding members include, but are not limited to, polypeptides, nucleic acids, carbohydrates, lipids, peptides, and the like. In some cases, the specific binding member is proteinaceous. As used herein, the term "proteinaceous" refers to a moiety composed of amino acid residues. The proteinaceous moiety can be a polypeptide. In certain cases, the proteinaceous specific binding member is an antibody. In certain embodiments, the proteinaceous specific binding member is an antibody fragment, e.g., a binding fragment of an antibody that specifically binds to a polymeric dye. As used herein, the terms "antibody" and "antibody molecule" are used interchangeably and refer to a protein consisting of one or more polypeptides substantially encoded by all or part of recognized immunoglobulin genes. Recognized immunoglobulin genes include, for example, in humans, the kappa (k), lambda (l), and heavy chain loci (which together contain a myriad of variable region genes), as well as the constant region genes mu (u), delta (d), gamma (g), sigma (e), and alpha (a) (which encode IgM, IgD, IgG, IgE, and IgA isotypes, respectively). An immunoglobulin light or heavy chain variable region consists of a framework region (FR) interrupted by three hypervariable regions, also called "complementarity determining regions" or "CDRs".The extent of the framework regions and CDRs have been precisely defined (see "Sequences of Proteins of Immunological Interest," E. Kabat et al., USDepartment of Health and Human Services, (1991)). The numbering of all antibody amino acid sequences described herein conforms to the Kabat system. The sequences of the framework regions of different light or heavy chains are relatively conserved within a species. The framework regions of an antibody, the combined framework regions of the constituent light and heavy chains, serve to position and align the CDRs. The CDRs primarily contribute to binding to an epitope of an antigen. The term "antibody" is meant to include full-length antibodies and may refer to natural antibodies from any organism, as further defined below, engineered antibodies, or antibodies recombinantly produced for experimental, therapeutic, or other purposes. Antibody fragments of interest include, but are not limited to, Fab, Fab', F(ab')2, Fv, scFv, or other antigen-binding subsequences of antibodies, either produced by modification of a whole antibody or synthesized de novo using recombinant DNA technology. The antibody may be monoclonal or polyclonal and may have other specific activity on cells (e.g., antagonist, agonist, neutralizing, inhibitory, or stimulatory antibody). It is understood that the antibody may have additional conservative amino acid substitutions that do not substantially affect antigen binding or other antibody functions. In certain embodiments, the specific binding member is a Fab fragment, a F(ab')2 fragment, an scFv, a diabody, or a triabody. In certain embodiments, the specific binding member is an antibody. In some cases, the specific binding member is a murine antibody or a binding fragment thereof. In certain cases, the specific binding member is a recombinant antibody or a binding fragment thereof.

[0039] The specific binding members / oligonucleotide sub-barcodes may specifically bind to any convenient cell marker. In some examples, the specific binding members / oligonucleotide sub-barcodes bind to cell surface markers. Cell surface markers of interest include, but are not limited to, ubiquitous cell surface markers, i.e., cell surface markers that are at least predicted to be present on all cells of a given cell sample processed in a given workflow according to the present invention. Examples of ubiquitous cell surface markers to which the specific binding members / oligonucleotide sub-barcodes may specifically bind include, but are not limited to, CD44, CD45, beta-2 microglobulin, and the like.

[0040] In addition to the specific binding member component, the specific binding member / oligonucleotide sub-barcode also includes an oligonucleotide sub-barcode component. The length of the oligonucleotide sub-barcode component can vary, in some cases ranging from 10-500 nt, e.g., 15-100 nt. In some cases, the oligonucleotide sub-barcode component can be comprised of ribonucleic acid or deoxyribonucleic acid, as appropriate. The oligonucleotide sub-barcodes of embodiments of the invention can include image labeling regions, as well as other domains used in embodiments of the invention, such domains can include unique identifiers for the specific binding member, primer binding sites, and the like.

[0041] The image labeling region of an oligonucleotide sub-barcode component is a domain or subsequence (i.e., stretch) of the oligonucleotide sub-barcode component that serves as a specific binding site for a labeled oligonucleotide used in the imaging step of an embodiment of the present invention, for example as described in greater detail below. The sequence of the image labeling region can be used as an identifier for the label (e.g., fluorescent label) of a labeled oligonucleotide that hybridizes to the image labeling region. Thus, the sequence of the image labeling region corresponds to the label of the labeled oligonucleotide that binds to that image labeling region. The image labeling region can have any convenient sequence and vary in length, in some cases ranging from 5 to 100 nt (e.g., 10 to 50 nt). A given oligonucleotide sub-barcode component can include a single image labeling region, or two or more image labeling regions (e.g., three or more image labeling regions), and in some cases the number of image labeling regions ranges from 1 to 5 (e.g., 2 to 3).

[0042] In addition to the image label region, the oligonucleotide sub-barcode component may include one or more of a unique identifier for a specific binding member, a capture sequence, a primer binding site, and the like. The unique identifier of a specific binding member is a domain or region that can be used (e.g., by its sequence) to identify the specific binding member. The unique identifier can be, for example, a nucleotide sequence having any suitable length, for example, from about 4 nucleotides to about 200 nucleotides. In some embodiments, the unique identifier is a nucleotide sequence that is 25 nucleotides to about 45 nucleotides in length. In some embodiments, the unique identifier can have a length of about 4 nucleotides, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 15 nucleotides, 20 nucleotides, 25 nucleotides, 30 nucleotides, 35 nucleotides, 40 nucleotides, 45 nucleotides, 50 nucleotides, 55 nucleotides, 60 nucleotides, 70 nucleotides, 80 nucleotides, 90 nucleotides, 100 nucleotides, 200 nucleotides, or a range between, less than, or greater than any two of the above values.

[0043] The oligonucleotide building blocks may include a capture sequence, e.g., a domain or region that serves as a binding site for a target binding region (e.g., of a bead-bound barcode nucleic acid), as described above. The capture sequence of interest may be varied and may be specific or random or semi-random, as desired. In some cases, the capture sequence hybridizes to the target binding region of a bead-bound nucleic acid, e.g., as described in more detail below. In some cases, the capture sequence is a poly(A) sequence, which is configured to hybridize to an oligo-dT target binding region, as described in more detail below. In such cases, the length of the poly(A) capture sequence may vary, in some cases ranging from 3 to 50 nt (e.g., 5 to 25 nt). The capture sequence, if present, may be located at the 5' end of the oligonucleotide building block.

[0044] The oligonucleotide component may include a primer binding site. The primer binding site, if present, may be configured to bind to a primer used (e.g., in preparing a sequenceable nucleic acid). For example, the oligonucleotide component may include a universal primer. A universal primer may refer to a nucleotide sequence that is universal or common across all specific binding members / oligonucleotide sub-barcodes used in a given workflow. In some examples, the primer binding site may be about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26 27, 28, 29, 30 nucleotides in length, or any two numbers or ranges thereof. The length of the primer binding site may vary and may be at least or up to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26 27, 28, 29, or 30 nucleotides in length. The length of the universal primer may vary and in some cases may range from 5 to 30 nucleotides in length. The primer binding site may be located at the 5' end of the oligonucleotide sub-barcode component.

[0045] As discussed above, in a specific binding member / oligonucleotide sub-barcode, the specific binding member is conjugated to the oligonucleotide sub-barcode component. The oligonucleotide component is conjugated to the specific binding member component through various mechanisms. In some embodiments, the oligonucleotide component may be covalently conjugated to the specific binding member component. In some embodiments, the oligonucleotide component may be non-covalently conjugated to the specific binding member component. In some embodiments, the oligonucleotide component is conjugated to the specific binding member component reagent via a linker. The linker may be, for example, cleavable or separable from the specific binding member and / or oligonucleotide component. In some embodiments, the linker may include a chemical group that reversibly bonds the oligonucleotide to the specific binding member. The chemical group may be conjugated to the linker, for example, via an amine group. In some embodiments, the linker may include a chemical group that forms a stable bond with another chemical group conjugated to the specific binding member component. For example, the chemical group may be a UV light cleavable group, a disulfide bond, streptavidin, biotin, an amine, etc. In some embodiments, the chemical group may be conjugated to the specific binding member component via an amino acid, e.g., a primary amine on a lysine or N-terminus. Commercially available conjugation kits, e.g., Protein-Oligo Conjugation Kit (Solulink, Inc., San Diego, California), Thunder-Link® Oligo Conjugation System (Innova Biosciences, Cambridge, United Kingdom), and the like, may be used to conjugate the oligonucleotide component to the specific binding member component. The oligonucleotide component may be conjugated to any suitable site on the specific binding member component (e.g., a protein binding reagent) so long as it does not interfere with specific binding between the specific binding member component and its cellular component target.Methods for conjugating oligonucleotides to specific binding members (e.g., antibodies) have been previously disclosed, for example, in U.S. Patent No. 6,531,283, the contents of which are incorporated herein by reference. The stoichiometry of the oligonucleotide to the specific binding member can vary.

[0046] Further details regarding specific binding member / oligonucleotide sub-barcode reagents and their components useful in embodiments of the present invention are provided in U.S. Patent Application Publication Nos. 2018 / 0088112, 2018 / 0200710, 2018 / 0346970, 2019 / 0056415, 2020 / 0248263, 2020 / 0299672, and 2021 / 0171940, the disclosures of which are incorporated herein by reference.

[0047] A given combinatorially barcoded cell may contain one or more specific binding members / oligonucleotide sub-barcodes stably associated therewith. In some cases, a given combinatorially barcoded cell contains multiple (i.e., two or more) different specific binding members / oligonucleotide sub-barcodes stably associated therewith, the different specific binding members / oligonucleotide sub-barcodes differing from one another at least with respect to the cell markers (e.g., cell surface proteins) to which they specifically bind. In some cases, the number of different specific binding members / oligonucleotide sub-barcodes stably associated with a combinatorially labeled cell ranges from 2 to 10 (e.g., 2 to 5, e.g., 3 to 4).

[0048] In embodiments of the methods of the invention, cells of a cell sample may be combinatorially barcoded using any convenient protocol. In some cases, the combinatorial barcode comprises one or more split / pool iterations in which cells of the cell sample are sequentially contacted with different specific binding members / oligonucleotide sub-barcodes. In some cases, each split / pool iteration comprises allocating cells of the cell sample into different compartments, introducing different (i.e., different) specific binding members / oligonucleotide sub-barcodes, where the oligonucleotide sub-barcode components differ from each other, into the different compartments to generate sub-barcoded cells, and pooling the sub-barcoded cells of the different compartments.

[0049] For a given split / pool iteration, the cells of the cell sample are allocated to different compartments such that they are separate from one another. The number of different compartments to which the cells are allocated can vary, and in some cases ranges from 5 to 1,000 (e.g., 5 to 500, including 5 to 100, e.g., 25 to 100). In some cases, the compartments are present in a substrate, e.g., wells of a well plate (e.g., wells of a macrowell plate). Examples of well plates to which the cell sample may be allocated include 36-well plates, 96-well plates, and 384-well plates, and in some embodiments, the well plate is a 36 or 96-well plate. Any convenient protocol can be used to allocate the cells of the cell sample to the different compartments, e.g., dispensing, pipetting, aliquoting the cell sample into the compartments, flowing the sample over the surface of the well plate, etc.

[0050] After cell allocation, different specific binding members / oligonucleotide sub-barcodes, whose oligonucleotide sub-barcode components differ from each other, are introduced into the different compartments to generate sub-barcoded cells. Different specific binding members / oligonucleotide sub-barcodes may be introduced into each compartment such that cells in the different compartments are stably associated with the specific binding members / oligonucleotide sub-barcodes introduced into those compartments. In this manner, cells in the different compartments are stably associated with different specific binding members / oligonucleotide sub-barcodes. In this step, the number of different specific binding members / oligonucleotide sub-barcodes introduced into the different compartments may vary in some cases in the range of 5-1,000 (e.g., 5-500), and in some cases the number is close to the number of compartments. Compartmentalized cells in stable association with specific binding members / oligonucleotide sub-barcodes may be referred to as sub-barcoded cells.

[0051] Following generation of the sub-barcoded cells, the sub-barcoded cells of the different compartments may be combined or pooled, for example to generate a pooled composition of sub-barcoded cells. The sub-barcoded cells may be combined or pooled using any convenient protocol. For example, the liquid compositions of the different compartments generated are collected from the compartments and combined, for example, in a suitable tube of sufficient volume.

[0052] Each split / sub-barcode / pool sequence in a given combinatorial labeling workflow may be referred to as an iteration. A given combinatorial labeling workflow may have any desired number of iterations, with more iterations providing more complex barcodes and a greater number of cells that may be processed in a given assay. In some instances, the number of split / pool iterations ranges from 2-10 (e.g., 2-5).

[0053] Sorting combinatorial barcoded cells to generate sorted combinatorial barcoded single cells, each having a combinatorial barcode After generation of the combinatorial barcoded cells, for example as described, embodiments of the method include partitioning the combinatorial barcoded cells to generate partitioned combinatorial barcoded single cells, each having a combinatorial barcode. In some cases, partitioning includes distributing the combinatorial barcoded cells into partitions or compartments, such that the compartments contain single combinatorial barcoded cells. By partitioning, it is meant placing the combinatorial barcoded cells into small reaction chambers, which can be fluidically isolated structures defined by solid materials, such as microwells configured to accommodate the combinatorial barcoded cells. In some embodiments of the disclosed methods, devices, and systems, a plurality of microwells randomly distributed across a substrate is used. In some embodiments, the plurality of microwells is distributed across a substrate in a regular pattern (e.g., a regular array). In some embodiments, the plurality of microwells is distributed across a substrate in a random pattern (e.g., a random array). The microwells can be manufactured in a variety of shapes and sizes. Suitable well geometries include, but are not limited to, three-dimensional geometries composed of several planes, such as cylinders, ellipses, cubes, cones, hemispheres, rectangles, or polyhedra, e.g., rectangular prisms, hexagonal prisms, octagonal prisms, inverted triangular pyramids, inverted square pyramids, inverted pentagonal pyramids, inverted hexagonal pyramids, or inverted truncated pyramids. In some embodiments, non-cylindrical microwells, e.g., wells having an elliptical or square footprint, may provide advantages in that they can accommodate more cells. In some embodiments, the top and / or bottom edges of the well walls are rounded to avoid sharp corners, thereby reducing electrostatic forces that may arise at sharp edges or points due to electrostatic field concentration. Thus, the use of rounded corners may improve the ability to retrieve beads from the microwells. The dimensions of the microwells may be characterized in terms of absolute dimensions. In some cases, the average diameter of the microwells may range from about 5 μm to about 100 μm.In other embodiments, the average microwell diameter is at least 5 μm, at least 10 μm, at least 15 μm, at least 20 μm, at least 25 μm, at least 30 μm, at least 35 μm, at least 40 μm, at least 45 μm, at least 50 μm, at least 60 μm, at least 70 μm, at least 80 μm, at least 90 μm, or at least 100 μm. In still other embodiments, the average microwell diameter is up to 100 μm, up to 90 μm, up to 80 μm, up to 70 μm, up to 60 μm, up to 50 μm, up to 45 μm, up to 40 μm, up to 35 μm, up to 30 μm, up to 25 μm, up to 20 μm, up to 15 μm, up to 10 μm, or up to 5 μm. The volume of the microwells used in the methods of the invention can vary, and in some cases is about 200 μm. 3 ~about 800,000μm 3 In some embodiments, the microwell volume is at least 200 μm 3 , at least 500 μm 3 , at least 1,000 μm 3 , at least 10,000 μm 3 , at least 25,000 μm 3 , at least 50,000 μm 3 , at least 100,000 μm 3 , at least 200,000 μm 3 , at least 300,000 μm 3 , at least 400,000 μm 3 , at least 500,000 μm 3 , at least 600,000 μm 3 , at least 700,000 μm 3 , or at least 800,000 μm 3 In other embodiments, the volume of the microwell is up to 800,000 μm 3 , up to 700,000μm 3 , up to 600,000μm 3 , 500,000μm 3 , up to 400,000μm 3 , up to 300,000μm 3 , up to 200,000μm 3, up to 100,000μm 3 , up to 50,000μm 3 , up to 25,000μm 3 , up to 10,000μm 3 , up to 1,000μm 3 , up to 500μm 3 , or up to 200 μm 3 The number of microwells in a given device used in embodiments of the invention can vary, and in some cases the number is 100 or more, e.g., 250 or more, e.g., 500 or more, 1000 or more, e.g., 5,000 or more, e.g., 10,000 or more, and in some cases the number is 15,000 or less, e.g., 12,500 or less. Microwells suitable for use in embodiments of the invention are further described in PCT Application No. PCT / US2016 / 014612 (published as WO / 2016 / 118915), the disclosure of which is incorporated herein by reference. As used herein, a substrate can refer to a type of solid support. A substrate can, for example, include a plurality of microwells. For example, a substrate can be a well array including two or more microwells. In some embodiments, a microwell can include a small reaction chamber of a defined volume. In some embodiments, a microwell can capture one or more cells. In some embodiments, a microwell can capture only one cell. In some embodiments, a microwell can capture one or more solid supports. In some embodiments, a microwell can capture only one solid support. In some embodiments, a microwell captures a single cell and a single solid support (e.g., a bead). The number of wells (e.g., microwells) in a well plate (e.g., a microwell array) can vary at a given allocation step, and in some cases the number ranged from 5 to 500 (e.g., 5 to 100).

[0054] When partitioning the combinatorial barcoded cells, the combinatorial barcoded cells can be placed into compartments (e.g., microwells of a microwell array) using any convenient protocol. The present disclosure provides methods for partitioning the combinatorial barcoded cells into compartments to partition the combinatorial barcoded cells. A collection of combinatorial barcoded cells can be introduced into a structure (e.g., a microwell) to partition the combinatorial barcoded cells, for example. The combinatorial barcoded cells can be contacted, for example, by gravity flow, and the combinatorial barcoded cells can settle into the partition structure. In some cases, an aqueous composition of the combinatorial barcoded cells is contacted with the array of microwells, for example, by flowing across it, such that the combinatorial barcoded cells are deposited in the microwells. The aqueous composition containing the combinatorial barcoded cells can flow through a flow cell in fluid communication with the microwells. Suitable protocols and systems for partitioning the captured particles into the microwells are described. Microwells suitable for use in embodiments of the present invention are further described in PCT Application No. PCT / US2016 / 014612 (published as WO / 2016 / 118915), the disclosure of which is incorporated herein by reference. Any convenient protocol can be used to compartmentalize the cells of the cell sample, such as dispensing, pipetting, aliquoting the cell sample into the compartment, flowing the sample over the surface of a well plate, etc.

[0055] In some embodiments, partitioning the plurality of combinatorial barcoded cells further includes providing particles (e.g., beads) containing particle (e.g., bead)-bound nucleic acids to the partition containing the single cells, and the bound nucleic acids are used to prepare a nucleic acid sequence composition (e.g., sequence library) from the combinatorial barcoded cells. In some cases, the particle (e.g., bead)-bound nucleic acids include a target binding region that binds to a complementary sequence in, for example, a nucleic acid species of interest in the combinatorial cells, as well as a capture sequence of an oligonucleotide sub-barcode component. For example, if the target nucleic acid species is cellular mRNA and the oligonucleotide sub-barcode component includes a poly(A) capture sequence, the bead-bound nucleic acid may include a poly(T) domain as the target binding region. In addition to the target binding region, many bound nucleic acids further include one or more additional domains, including, but not limited to, a cell labeling domain, a barcode domain, a molecular index domain (e.g., a unique molecular identifier (UMI) domain), a universal primer binding domain, and the like. Further details regarding particles with bound nucleic acids that can be provided to the compartments can be found in U.S. Patent Application Publication Nos. 2018 / 0088112, 2018 / 0200710, 2018 / 0346970, 2019 / 0056415, 2020 / 0248263, 2020 / 0299672, and 2021 / 0171940, the disclosures of which are incorporated herein by reference. Beads with bound nucleic acids can be provided in the compartments using any convenient protocol. Protocols include, but are not limited to, those described above for cell partitioning and are further described in PCT Application No. PCT / US2016 / 014612 (published as WO / 2016 / 118915), the disclosures of which are incorporated herein by reference. Particles (e.g., beads) can, in some cases, be partitioned into cells before or after the combinatorial barcoded cells, or, optionally, in combination with the combinatorial barcoded cells.

[0056] Acquiring image data and sequence data for sectioned combinatorial barcoded single cells As summarized above, after generation of the sorted combinatorial barcoded single cells, each having a combinatorial barcode, image data and sequence data are obtained for the sorted combinatorial barcoded single cells. In embodiments, the image data for the sorted barcoded single cells is obtained prior to obtaining the sequence data for the sorted combinatorial barcoded single cells.

[0057] Acquisition of image data The partitioned combinatorial barcoded single cells may be imaged using any convenient protocol to obtain image data of the partitioned single cells. The image data obtained may vary. Image data may be obtained for any combinatorial barcoded cell of interest, and from a partition containing the combinatorial barcoded cell of interest. The type of image data obtained may vary and may include image data of live cells. Any convenient protocol may be used to obtain image data for the combinatorial barcoded cells in the partition. Examples of imaging protocols that may be used include, but are not limited to, microscopic imaging protocols such as phase contrast microscopy, fluorescence microscopy, quantitative phase contrast microscopy, holotomography, BD Rhapsody System, and the like. The image may be generated, for example, by fluorescence imaging. Imaging may include microscopy such as bright field imaging, oblique illumination, dark field imaging, dispersion staining, phase contrast, differential interference contrast, interference reflection microscopy, fluorescence, confocal, and single plane illumination, or any combination thereof. Imaging may include imaging a portion of the sample (e.g., slide / array). The imaging may include imaging at least 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100% of the sectioned cells. In some cases, the imaging may be performed in separate steps (e.g., the images may not be consecutive). The imaging may include taking at least 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10, or more different images. The imaging may include taking up to 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10, or more different images. If desired, the image data may include images taken from two or more different imaging iterations, each imaging iteration including a labeling step and then an imaging step. In such cases, the acquisition of image data from the sectioned cells may be considered as a cyclic imaging step.

[0058] In an embodiment of the invention, acquiring image data for a partitioned combinatorial barcoded single cell includes acquiring a partition-specific fluorescent barcode for the partition of interest and thus the combinatorial barcoded cells present therein. Thus, the method may include acquiring for each combinatorial barcoded cell of interest the partition containing the cell and thus the partition-specific fluorescent barcode for that cell. A partition-specific fluorescent barcode is a collection of fluorescent signals acquired in a given partition that correspond to image-labeled regions of the combinatorial barcode of cells in that partition. In some cases, the collection of fluorescent signals that make up a given barcode has a specific sequence, such as a time sequence corresponding to time points acquired in a given workflow and / or a location of image-labeled regions on the oligonucleotide sub-barcode components that correspond to a given fluorescent signal of the barcode.

[0059] In some cases, the section-specific fluorescent barcodes are obtained in two or more imaging iterations (e.g., 2-20 iterations, including 2-10 iterations). Each imaging iteration includes contacting the sectioned combinatorial barcoded single cell with one or more labeled oligonucleotides that bind to the image label region of the oligonucleotide sub-barcode component of the specific binding member / oligonucleotide sub-barcode to generate a labeled sectioned combinatorial barcoded single cell, and capturing an image of the labeled sectioned combinatorial barcoded single cell to obtain a fluorescent signal from the label of the labeled oligonucleotide hybridized to the image label region. In such an embodiment, the labeled oligonucleotide is contacted with a labeled oligonucleotide that binds to the image label region of the oligonucleotide sub-barcode component of the combinatorial barcoded cell. The labeled oligonucleotide is an oligonucleotide that hybridizes to the image label region and includes a detectable label. In the labeled oligonucleotides used in the embodiments of the present invention, the detectable label can be part of a labeled nucleic acid that hybridizes to the image label region of the oligonucleotide sub-barcode unit. In such cases, the length of the labeled nucleic acid may vary, in some cases ranging from 5-100 nt in length, and may include one or more detectable moieties attached thereto. In some embodiments, the detectable moiety includes an optical moiety, a luminescent moiety, an electrochemically active moiety, a nanoparticle, or a combination thereof. In some embodiments, the luminescent moiety includes a chemiluminescent moiety, an electroluminescent moiety, a photoluminescent moiety, or a combination thereof. In some embodiments, the photoluminescent moiety includes a fluorescent moiety, a phosphorescent moiety, or a combination thereof. In some embodiments, the fluorescent moiety is a fluorescent dye. In some embodiments, the nanoparticle includes a quantum dot. In some embodiments, the method includes performing a reaction to convert a precursor of the detectable moiety to the detectable moiety.Detectable moieties that are useful in embodiments of the present invention include those described in U.S. Patent Application Publication Nos. 2018 / 0088112, 2018 / 0200710, 2018 / 0346970, 2019 / 0056415, 2020 / 0248263, 2020 / 0299672, and 2021 / 0171940, the disclosures of which are incorporated herein by reference.

[0060] For example, as described above, the compartmented combinatorial barcoded single cells are contacted with one or more labeled oligonucleotides that bind to the image label region of the oligonucleotide sub-barcode component of the specific binding member / oligonucleotide sub-barcode to generate labeled compartmented combinatorial barcoded single cells. In the set of labeled compartmented combinatorial barcoded single cells, the single cells in the compartment contain the image label region of the sub-barcode component hybridized to the labeled oligonucleotide. To facilitate imaging, the same labeled oligonucleotide with the same label (e.g., fluorescent dye) can be contacted with all of the combinatorial labeled cells in all compartments. These combinatorial labeled cells with image label regions complementary to the labeled oligonucleotides hybridize to the labeled oligonucleotides and are detectable in a subsequent imaging step.

[0061] Following generation of the labeled, sectioned, combinatorial barcoded single cells, the detectable labels of the labeled, sectioned, combinatorial barcoded single cells can be detected to obtain a fluorescent signal from the label. The detection of the fluorescent signal can be performed using any convenient protocol, which can include excitation of the cells with light at a suitable wavelength and detection of light from the label associated with the cells. As outlined above, the signal from the sectioned, combinatorial barcoded cells can be obtained in successive iterations, with each detection iteration including a labeling step followed by a detection step. In such a case, obtaining image data from the sectioned cells can be considered as a cyclic imaging step, whereby the sectioned specific fluorescent barcodes are obtained using the cyclic imaging step. In such an embodiment, for example, a set of sectioned, combinatorial barcoded single cells can be contacted with a first labeled oligonucleotide that specifically binds to a first image domain of a sub-barcode component that can be associated with a different cell of the sectioned cells. A first subset of image data can be obtained from the cells. The sectioned cells may then be contacted with a second labeled oligonucleotide that specifically binds to a second image domain of the sub-barcode component that may be associated with a different cell of the sectioned cells. A second subset of image data may then be acquired from the cells. This process may be repeated for any desired number of iterations. In some cases, the number of imaging iterations used to acquire image data ranges from 2 to 20 (e.g., 2 to 10). Between each iteration, the previous set of labeled oligonucleotides may be removed from the sectioned cells, for example, by washing the cells. Alternatively, the labels used in the previous iteration may be inactivated so that they are not detectable in the subsequence imaging iteration. In still other embodiments, the set of labels selected for use in a given imaging iteration protocol may be selected such that the labels are distinguishable in terms of excitation and / or emission maxima. In such cases, a single labeling step may be used in which two or more different labeled oligonucleotides are introduced to the section under hybridization conditions.After removing unbound labeled oligonucleotides, the labeled sectioned cells may then be cycled through two or more imaging steps, each imaging step differing in excitation and / or detection of the cells. Thus, a section-specific fluorescent barcode for a given section may be obtained by first contacting the given section with a plurality of different labeled oligonucleotides, a subset of the plurality of different labeled oligonucleotides binding to corresponding image-labeled regions of specific binding members / oligonucleotide sub-barcodes associated with cells present in the section. After removal of unbound labeled oligonucleotides, the remaining bound labeled oligonucleotides may be detected to obtain a fluorescent barcode for that section, and the detection protocol may be repetitive, such as a cyclic imaging protocol.

[0062] Acquiring sequence data For example, as described above, partitioning of combinatorial barcoded cells results in partitioned combinatorial barcoded cells that are spatially adjacent to particles (e.g., beads), which have bound cell labeling domain nucleic acids that include target binding regions, as described above. When the cell labeling domain nucleic acid is adjacent to a target of a combinatorial barcoded single cell, the target can hybridize to the cell labeling domain nucleic acid. The cell labeling domains that include nucleic acids can be contacted in a non-depleting ratio, if desired, such that each different target can associate with a different cell labeling domain that includes a nucleic acid with its own unique UMI.

[0063] After partitioning the combinatorial barcoded cells as described above, the combinatorial barcoded cells can be lysed to release the target molecules, whereby the released target molecules (e.g., nucleic acids) can bind to the target binding regions of the cell labeling domain nucleic acids to generate captured nucleic acids. Cell lysis can be achieved by any of a variety of means, for example, by chemical or biochemical means, by osmotic shock, or by either thermal, mechanical, or optical lysis. The particles can be lysed by adding a cell lysis buffer containing a detergent (e.g., SDS, Li dodecyl sulfate, TritonX-100, Tween-20, or NP-40), an organic solvent (e.g., methanol or acetone), or a digestive enzyme (e.g., proteinase K, pepsin, or trypsin), or any combination thereof. To increase the association of the target and the barcode, the diffusion rate of the target molecules can be altered, for example, by lowering the temperature and / or increasing the viscosity of the lysate. In some embodiments, the sample can be lysed using filter paper. The filter paper can be soaked in the lysis buffer on top of the filter paper. The filter paper can be applied to the sample with pressure to facilitate lysis of the sample and hybridization of the sample to the target substrate. In some embodiments, lysis can be performed by mechanical lysis, heat lysis, optical lysis, and / or chemical lysis. Chemical lysis can include the use of digestive enzymes such as proteinase K, pepsin, and trypsin. Lysis can be performed by adding a lysis buffer to the substrate. The lysis buffer can include Tris HCl. The lysis buffer can include at least about 0.01, 0.05, 0.1, 0.5, or 1 M Tris HCl, or more. The lysis buffer can include up to about 0.01, 0.05, 0.1, 0.5, or 1 M Tris HCL, or more. The lysis buffer can include about 0.1 M Tris HCl. The pH of the lysis buffer can be at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more. The pH of the lysis buffer can be up to about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more. In some embodiments, the pH of the lysis buffer is about 7.5.The lysis buffer may include a salt (e.g., LiCl). The concentration of the salt in the lysis buffer may be at least about 0.1, 0.5, or 1 M, or more. The concentration of the salt in the lysis buffer may be up to about 0.1, 0.5, or 1 M, or more. In some embodiments, the concentration of the salt in the lysis buffer is about 0.5 M. The lysis buffer may include a detergent (e.g., SDS, Li-dodecyl sulfate, Triton-X, tween, NP-40). The concentration of the detergent in the lysis buffer may be at least about 0.0001%, 0.0005%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, or 7%, or more. The concentration of the detergent in the lysis buffer can be up to about 0.0001%, 0.0005%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, or 7%, or more. In some embodiments, the concentration of the detergent in the lysis buffer is about 1% Li-dodecyl sulfate. The time used in the lysis method can depend on the amount of detergent used. In some embodiments, the more detergent used, the less time is required for lysis. The lysis buffer can include a chelating agent (e.g., EDTA, EGTA). The concentration of the chelating agent in the lysis buffer can be at least about 1, 5, 10, 15, 20, 25, or 30 mM, or more. The concentration of the chelating agent in the lysis buffer can be up to about 1, 5, 10, 15, 20, 25, or 30 mM, or more. In some embodiments, the concentration of the chelating agent in the lysis buffer is about 10 mM. The lysis buffer may include a reducing agent (e.g., beta-mercaptoethanol, DTT). The concentration of the reducing agent in the lysis buffer may be at least about 1, 5, 10, 15, or 20 mM, or more. The concentration of the reducing agent in the lysis buffer may be up to about 1, 5, 10, 15, or 20 mM, or more. In some embodiments, the concentration of the reducing agent in the lysis buffer is about 5 mM. In some embodiments, the lysis buffer may include about 0.1 M Tris HCl (about pH 7.5), about 0.5 M LiCl, about 1% lithium dodecyl sulfate, about 10 mM EDTA, and about 5 mM DTT.Lysing can be performed at a temperature of about 4, 10, 15, 20, 25, or 30° C. Lysing can be performed for about 1, 5, 10, 15, or 20 minutes or more. Lysed cells can contain at least about 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, or 700,000, or more target nucleic acid molecules. Lysed cells can contain up to 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, or 700,000, or more target nucleic acid molecules.

[0064] Following lysis of combinatorial barcoded cells and release of nucleic acid molecules therefrom, the nucleic acid molecules can randomly associate with cell labeling domain nucleic acids of a co-localized solid support (e.g., beads). The association can include hybridization of the target recognition region of the cell labeling domain nucleic acid to a complementary portion of the target nucleic acid molecule (e.g., the oligo(dT) of the barcode can interact with the poly(A) tail of the target). The assay conditions (e.g., buffer pH, ionic strength, temperature, etc.) used for hybridization can be selected to promote the formation of specific and stable hybrids. In some embodiments, the nucleic acid molecules released from the lysed cells can associate with (e.g., hybridize to) multiple probes on a substrate. If the probes include oligo(dT), the mRNA molecules can hybridize to the probes and be reverse transcribed. The oligo(dT) portion of the oligonucleotide can function as a primer for synthesis of the first strand of a cDNA molecule, for example, when subjected to DNA synthesis reaction conditions to generate a first strand cDNA domain that includes the capture nucleic acid. The cell labeling domain nucleic acid can also hybridize to a complementary capture sequence (e.g., a poly(A) sequence) of an oligonucleotide sub-barcode component of a specific binding member / oligonucleotide sub-barcode associated with a combinatorial barcoded cell. In this manner, the cell labeling domain nucleic acid can function as a primer for reverse transcription, e.g., using the oligonucleotide sub-barcode as a template, as described in more detail below.

[0065] If desired, a given workflow may include a pooling step, e.g., a product composition consisting of captured nucleic acids, synthesized first strand cDNA, or synthesized double strand cDNA is combined or pooled with product compositions obtained from one or more additional samples, e.g., combinatorially barcoded cells. In some cases, the pooling step is performed immediately after a hybridization step between the cell labeling domain nucleic acid and the target nucleic acid, e.g., as outlined above. The number of different product compositions generated from different samples (e.g., cells) that are combined or pooled in such embodiments may vary, in some cases the number is in the range of 2-1,000,000 (e.g., 3-200,000, including 4-100,000, such as 5-50,000), and in some cases the number is in the range of 100-10,000 (e.g., 1,000-5,000). Before or after pooling, the product composition may be amplified, for example, by polymerase chain reaction (PCR), as described in more detail below. Once the target cell domain labeling molecules are pooled, all further processing may proceed within a single reaction vessel. Further processing may include, for example, reverse transcription reactions, amplification reactions, cleavage reactions, dissociation reactions, and / or nucleic acid extension reactions. Further processing reactions may be performed within the microwells, i.e., without first pooling the labeled target nucleic acid molecules from multiple cells.

[0066] The present disclosure provides methods for making target cell labeling domain conjugates using any convenient protocol, such as reverse transcription or nucleotide extension. The target cell labeling domain conjugates may include a cell labeling domain and a complementary sequence of all or part of the target nucleic acid. Reverse transcription of the associated RNA molecule may occur by adding a reverse transcription primer along with a reverse transcriptase. The reverse transcription primer may be an oligo(dT) primer, a random hexanucleotide primer, or a target-specific oligonucleotide primer. The oligo(dT) primer may be 12-18 nucleotides in length or about 12-18 nucleotides in length and binds to the endogenous poly(A) tail at the 3' end of the mammalian mRNA. The random hexanucleotide primer may bind to the mRNA at various complementary sites. The target-specific oligonucleotide primer typically selectively primes the mRNA of interest. Reverse transcription may occur repeatedly to generate multiple cDNA molecules. The methods disclosed herein can include performing at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 reverse transcription reactions. The methods can include performing at least about 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 reverse transcription reactions.

[0067] One or more nucleic acid amplification reactions can be performed to generate multiple copies of a target nucleic acid molecule. Amplification can be performed in a multiplexed manner, where multiple target nucleic acid sequences are amplified simultaneously. The amplification reaction can be used to add sequencing adaptors to the nucleic acid molecule. The amplification reaction can include amplifying at least a portion of the sample label, if present. The amplification reaction can include amplifying at least a portion of the cell label and / or barcode sequence (e.g., molecular label). The amplification reaction can include amplifying at least a portion of the sample tag, cell label, spatial label, barcode sequence (e.g., molecular label), target nucleic acid, or a combination thereof. The amplification reaction may include amplifying 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 100%, or a range or value between any two of these values ​​of the plurality of nucleic acids. The method may further include performing one or more cDNA synthesis reactions to generate one or more cDNA copies of the target barcode molecule that includes the sample label, cell label, spatial label, and / or barcode sequence (e.g., molecular label).

[0068] In some embodiments, the amplification may be performed using polymerase chain reaction (PCR). As used herein, PCR may refer to a reaction for in vitro amplification of specific DNA sequences by simultaneous primer extension of complementary strands of DNA. As used herein, PCR may encompass derivative forms of the reaction, including, but not limited to, RT-PCR, real-time PCR, nested PCR, quantitative PCR, multiplexed PCR, digital PCR, and assembly PCR.

[0069] Amplification of nucleic acids may include non-PCR-based methods. Examples of non-PCR-based methods include, but are not limited to, multiple displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), real-time SDA, rolling circle amplification, or circle-to-circle amplification. Other non-PCR-based amplification methods include DNA-dependent RNA polymerase-driven RNA transcription amplification or multiple cycles of RNA-directed DNA synthesis and transcription to amplify DNA or RNA targets, ligase chain reaction (LCR), and Qβ replicase (Qβ) method, use of palindromic probes, strand displacement amplification, oligonucleotide-driven amplification using restriction endonucleases, amplification methods in which primers are hybridized to nucleic acid sequences and the resulting duplex is cleaved before extension reaction and amplification, strand displacement amplification using nucleic acid polymerases lacking 5' exonuclease activity, rolling circle amplification, and branched extension amplification (RAM). In some embodiments, the amplification does not generate circularized transcripts.

[0070] In some embodiments, the methods disclosed herein further include performing a polymerase chain reaction on the nucleic acid (e.g., RNA, DNA, cDNA) to generate a labeled amplicon (e.g., a stochastically labeled amplicon). The labeled amplicon can be a double-stranded molecule. The double-stranded molecule can include a double-stranded RNA molecule, a double-stranded DNA molecule, or an RNA molecule hybridized to a DNA molecule. One or both strands of the double-stranded molecule can include a sample label, a spatial label, a cell label, and / or a barcode sequence (e.g., a molecular label). The labeled amplicon can be a single-stranded molecule. The single-stranded molecule can include DNA, RNA, or a combination thereof. The nucleic acid of the present disclosure can include a synthetic or modified nucleic acid. Thus, the method can include generating an amplicon composition from a first strand cDNA domain that includes a capture nucleic acid.

[0071] Amplification may include the use of one or more non-natural nucleotides. Non-natural nucleotides may include photodegradable or inducible nucleotides. Examples of non-natural nucleotides may include, but are not limited to, peptide nucleic acid (PNA), morpholino and locked nucleic acid (LNA), as well as glycol nucleic acid (GNA) and threose nucleic acid (TNA). Non-natural nucleotides may be added to one or more cycles of the amplification reaction. The addition of non-natural nucleotides may be used to identify products as specific cycles or time points in the amplification reaction.

[0072] Conducting the one or more amplification reactions may include the use of one or more primers. The one or more primers may include, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 or more nucleotides. The one or more primers may include at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 or more nucleotides. The one or more primers may include less than 12-15 nucleotides. The one or more primers may anneal to at least a portion of the multiple labeled targets (e.g., stochastically labeled targets). The one or more primers may anneal to the 3' or 5' ends of the multiple labeled targets. The one or more primers may anneal to an internal region of the multiple labeled targets. The internal region can be at least about 50, 100, 150, 200, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 650, 700, 750, 800, 850, 900, or 1000 nucleotides from the 3' end of the plurality of labeled targets. The one or more primers can comprise a fixed panel of primers. The one or more primers may include at least one or more custom primers. The one or more primers may include at least one or more control primers. The one or more primers may include at least one or more gene-specific primers.

[0073] The one or more primers may include a universal primer. The universal primer may anneal to the universal primer binding site. The one or more custom primers may anneal to a first sample label, a second sample label, a spatial label, a cell label, a barcode sequence (e.g., a molecular label), a target, or any combination thereof. The one or more primers may include a universal primer and a custom primer. The custom primers may be designed to amplify one or more targets. The targets may include a subset of all nucleic acids in one or more samples. The targets may include a subset of all labeled targets in one or more samples. The one or more primers may include at least 96 or more custom primers. The one or more primers may include at least 960 or more custom primers. The one or more primers may include at least 9600 or more custom primers. The one or more custom primers may anneal to two or more different labeled nucleic acids. The two or more different labeled nucleic acids may correspond to one or more genes.

[0074] Any amplification scheme can be used in the disclosed method. For example, in one scheme, the first round of PCR can amplify the molecules bound to the beads using a gene-specific primer and a primer to the universal Illumina sequencing primer 1 sequence. The second round of PCR can amplify the first PCR product using a nested gene-specific primer adjacent to the Illumina sequencing primer 2 sequence and a primer to the universal Illumina sequencing primer 1 sequence. The third round of PCR adds P5 and P7 and a sample index to make the PCR products into an Illumina sequencing library. Sequencing using 150bp x 2 sequencing can reveal cell label and barcode sequences (e.g., molecular label) on read 1, genes on read 2, and sample index on index 1 read.

[0075] In some embodiments, the nucleic acid may be removed from the substrate using chemical cleavage. For example, chemical groups or modified bases present on the nucleic acid may be used to facilitate its removal from the solid support. For example, an enzyme may be used to remove the nucleic acid from the substrate. For example, the nucleic acid may be removed from the substrate through restriction endonuclease digestion. For example, the nucleic acid may be removed from the substrate using treatment of a nucleic acid containing dUTP or ddUTP with uracil-d-glycosylase (UDG). For example, the nucleic acid may be removed from the substrate using an enzyme that performs nucleotide removal, such as a base excision repair enzyme, such as an apurinic / apyrimidinic (AP) endonuclease. In some embodiments, the nucleic acid may be removed from the substrate using a photocleavable group and light. In some embodiments, a cleavable linker may be used to remove the nucleic acid from the substrate. For example, the cleavable linker may include at least one of biotin / avidin, biotin / streptavidin, biotin / neutravidin, Ig-protein A, a photocleavable linker, an acid or base degradable linker group, or an aptamer.

[0076] In some embodiments, amplification can be performed on the substrate, for example, using bridge amplification. The cDNA can be homopolymer tailed to generate ends compatible with bridge amplification using oligo(dT) probes on the substrate. In bridge amplification, a primer complementary to the 3' end of the template nucleic acid can be the first primer of each pair covalently attached to the solid particle. When a sample containing the template nucleic acid is contacted with the particle and a single thermal cycle is performed, the template molecule can be annealed to the first primer, and the first primer can be extended in the forward direction by the addition of nucleotides to form a duplex molecule consisting of the template molecule and a newly formed DNA strand complementary to the template. The heating step of the next cycle can denature the duplex molecule and release the template molecule from the particle, leaving the complementary DNA strand attached to the particle via the first primer. In the annealing phase of the subsequent annealing and extension step, the complementary strand can hybridize to a second primer that is complementary to the segment of the complementary strand at the position removed from the first primer. This hybridization forms a bridge between the first and second primers, which is covalently fixed to the first primer and hybridized to the second primer. In the extension step, the second primer can be extended in the reverse direction by adding nucleotides in the same reaction mixture, thereby converting the bridge into a double-stranded bridge. The next cycle then begins, denaturing the double-stranded bridge to obtain two single-stranded nucleic acid molecules. Each molecule is bound to the particle surface at one end via the first and second primers, respectively, and the other end is unbound. In the annealing and extension step of this second cycle, each strand can hybridize to a further complementary primer not previously used on the same particle to form a new single-stranded bridge. The now hybridized two previously unused primers are extended to convert the two new bridges into double-stranded bridges.The amplification reaction may include amplifying at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 100% of the plurality of nucleic acids.

[0077] The amplification of the labeled nucleic acid may include PCR-based or non-PCR-based methods. The amplification of the labeled nucleic acid may include exponential amplification of the labeled nucleic acid. The amplification of the labeled nucleic acid may include linear amplification of the labeled nucleic acid. The amplification may be performed by polymerase chain reaction (PCR). PCR may refer to a reaction for in vitro amplification of specific DNA sequences by simultaneous primer extension of complementary strands of DNA. PCR may include derivative types of reactions, including but not limited to RT-PCR, real-time PCR, nested PCR, quantitative PCR, multiplexed PCR, digital PCR, suppression PCR, semi-suppression PCR, and assembly PCR.

[0078] In some embodiments, the amplification of the labeled nucleic acid comprises a non-PCR-based method. Examples of non-PCR-based methods include, but are not limited to, multiple displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand displacement amplification (SDA), real-time SDA, rolling circle amplification, or circle-to-circle amplification. Other non-PCR-based amplification methods include DNA-dependent RNA polymerase-driven RNA transcription amplification or multiple cycles of RNA-directed DNA synthesis and transcription to amplify DNA or RNA targets, ligase chain reaction (LCR), Qβ replicase (Qβ), the use of palindromic probes, strand displacement amplification, oligonucleotide-driven amplification using restriction endonucleases, amplification methods in which a primer is hybridized to a nucleic acid sequence and the resulting duplex is cleaved before extension reaction and amplification, strand displacement amplification using a nucleic acid polymerase lacking 5' exonuclease activity, rolling circle amplification, and / or branched extension amplification (RAM).

[0079] In some embodiments, the methods disclosed herein further comprise performing a nested polymerase chain reaction on the amplified amplicon (e.g., target). The amplicon may be a double-stranded molecule. The double-stranded molecule may comprise a double-stranded RNA molecule, a double-stranded DNA molecule, or an RNA molecule hybridized to a DNA molecule. One or both of the strands of the double-stranded molecule may comprise a sample tag or molecular identifier label. Alternatively, the amplicon may be a single-stranded molecule. The single-stranded molecule may comprise DNA, RNA, or a combination thereof. The nucleic acids of the invention may include synthetic or modified nucleic acids.

[0080] In some embodiments, the methods include repeatedly amplifying the labeled nucleic acid to generate a plurality of amplicons. The methods disclosed herein may include performing at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amplification reactions. Alternatively, the methods include performing at least about 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 amplification reactions.

[0081] The amplification may further include adding one or more control nucleic acids to one or more samples comprising the plurality of nucleic acids. The amplification may further include adding one or more control nucleic acids to the plurality of nucleic acids. The control nucleic acids may include a control label.

[0082] Amplification may include the use of one or more non-natural nucleotides. Non-natural nucleotides may include photocleavable and / or inducible nucleotides. Examples of non-natural nucleotides include, but are not limited to, peptide nucleic acid (PNA), morpholino and locked nucleic acid (LNA), as well as glycol nucleic acid (GNA) and threose nucleic acid (TNA). Non-natural nucleotides may be added to one or more cycles of the amplification reaction. The addition of non-natural nucleotides may be used to identify products as specific cycles or time points in the amplification reaction.

[0083] Conducting the one or more amplification reactions may include the use of one or more primers. The one or more primers may include one or more oligonucleotides. The one or more oligonucleotides may include at least about 7-9 nucleotides. The one or more oligonucleotides may include less than 12-15 nucleotides. The one or more primers may anneal to at least a portion of the plurality of labeled nucleic acids. The one or more primers may anneal to the 3' and / or 5' ends of the plurality of labeled nucleic acids. The one or more primers may anneal to an internal region of the plurality of labeled nucleic acids. The internal region can be at least about 50, 100, 150, 200, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 650, 700, 750, 800, 850, 900, or 1000 nucleotides from the 3' end of the plurality of labeled nucleic acids. The one or more primers can comprise a fixed panel of primers. The one or more primers may include at least one or more custom primers. The one or more primers may include at least one or more control primers. The one or more primers may include at least one or more housekeeping gene primers. The one or more primers may include a universal primer. The universal primer may anneal to a universal primer binding site. The one or more custom primers may anneal to a first sample tag, a second sample tag, a molecular identifier label, a nucleic acid, or a product thereof. The one or more primers may include a universal primer and a custom primer. The custom primer may be designed to amplify one or more target nucleic acids. The target nucleic acid may include a subset of the total nucleic acid in one or more samples. In some embodiments, the primer is a probe attached to the array of the present disclosure.

[0084] In some embodiments, barcoding (e.g., stochastic barcoding) a plurality of targets in a sample further comprises generating an indexed library of barcoded targets (e.g., stochastic barcoded targets) or barcoded fragments of the targets. The barcode sequences of the different barcodes (e.g., molecular labels of the different stochastic barcodes) can differ from each other. Generating an indexed library of barcoded targets comprises generating a plurality of indexed polynucleotides from the plurality of targets in the sample. For example, in the case of an indexed library of barcoded targets including a first indexed target and a second indexed target, the labeled region of the first indexed polynucleotide can differ from the labeled region of the second indexed polynucleotide by about, at least, or up to 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, or a number or range between any two of these values. In some embodiments, generating an indexed library of barcoded targets includes contacting a plurality of targets (e.g., mRNA molecules) with a plurality of oligonucleotides comprising a poly(T) region and a label region, and performing first strand synthesis using a reverse transcriptase to generate single-stranded labeled cDNA molecules, each comprising a cDNA region and a label region, wherein the plurality of targets comprises at least two mRNA molecules of different sequences, and the plurality of oligonucleotides comprises at least two oligonucleotides of different sequences. Generating an indexed library of barcoded targets may further include amplifying the single-stranded labeled cDNA molecules to generate double-stranded labeled cDNA molecules, and performing nested PCR on the double-stranded labeled cDNA molecules to generate labeled amplicons. In some embodiments, the method may include generating adapter-labeled amplicons.

[0085] Barcoding (e.g., stochastic barcoding) may involve labeling individual nucleic acid (e.g., DNA or RNA) molecules with a nucleic acid barcode or tag. In some embodiments, it involves adding a DNA barcode or tag to a cDNA molecule generated from an mRNA. Nested PCR may be performed to minimize PCR amplification bias. For example, next generation sequencing (NGS) may be used to add adapters for sequencing. Sequencing results can be used to determine the sequence of cellular labels, molecular labels, and nucleotide fragments of one or more copies of a target.

[0086] In certain embodiments, the methods provided further include subjecting the prepared expression library, e.g., the amplicon composition generated as described above, to a sequencing protocol, such as an NGS protocol. The protocol may be performed on any suitable NGS sequencing platform. NGS sequencing platforms of interest include, but are not limited to, Illumina® (e.g., HiSeq™, MiSeq™, and / or NextSeq™ sequencing systems), Ion Torrent™ (e.g., Ion PGM™ and / or Ion Proton™ sequencing systems), Pacific Biosciences (e.g., PACBIO RSII Sequel sequencing systems), Life Technologies™ (e.g., SOLiD sequencing systems), Oxford Nanopore (e.g., Minion), Roche (e.g., 454 GS FLX+ and / or GS Junior sequencing systems), or any other sequencing platform of interest. The NGS protocol will vary depending on the particular NGS sequencing system used. Detailed protocols for sequencing, which may include, for example, further amplification (e.g., solid-phase amplification), sequencing of the amplicons, and analysis of the sequencing data, are available from the manufacturer of the NGS sequencing system used.

[0087] In some cases, the method further includes using an oligonucleotide-labeled cellular component binding reagent, for example, in applications where detection (e.g., quantification) of one or more cellular components (e.g., surface proteins) is desired. The oligonucleotide-labeled cellular component binding reagent used in such embodiments includes a cellular component binding reagent (e.g., an antibody or binding fragment thereof) bound to a cellular component binding reagent-specific oligonucleotide that includes an identifier sequence of the cellular component binding reagent with which the cellular component binding reagent-specific oligonucleotide is associated. In such cases, the magnetic capture bead may include a nucleic acid configured to capture (e.g., specifically bind) a domain of the cellular component binding reagent-specific oligonucleotide. In this manner, protein expression may be assayed in conjunction with gene expression, for example, when a multi-omics analysis (e.g., combined analysis of the transcriptome and proteome) is desired. In such cases, the method may include preparing a captured sample with the oligonucleotide-labeled cellular component binding reagent, and then providing capture of the cellular component binding reagent-specific oligonucleotide released from the captured and sorted cells. Further details regarding the use of oligonucleotide-labeled cellular component binding reagents can be found in U.S. Published Patent Applications Nos. 2018 / 0267036 and 2020 / 0248263, the disclosures of which are incorporated herein by reference.

[0088] For example, as noted above, further details regarding methods for obtaining sequence data from single cells are provided in U.S. Patent Application Publication Nos. 2018 / 0088112, 2018 / 0200710, 2018 / 0346970, 2019 / 0056415, 2020 / 0248263, 2020 / 0299672, and 2021 / 0171940 (the disclosures of which are incorporated by reference herein).

[0089] The sequence protocol generates sequence data for the combinatorial barcoded cells. This sequence data can then be readily linked to image data for the combinatorial barcoded cells such that image data and sequence data obtained from the same combinatorial barcoded cells can be paired. In other words, for example, as described in more detail below, a given set of image data and a given set of sequence data can be linked as obtained from the same combinatorial barcoded cells.

[0090] Linking image data and sequence data sharing a common combinatorial barcode For example, following acquisition of image data and sequence data as described above, the image and sequence data acquired from a given partition (and thus the cells present in that partition) are concatenated. By concatenation, it is meant that the image data and sequence data are paired as originating from the same partition, and thus the combinatorial barcoded cells present in that partition when the image data for that partition was acquired. Thus, image data and sequence data acquired from the same combinatorial barcoded cells can be paired. In other words, a given set of image data and a given set of sequence data can be identified as having been acquired from the same combinatorial barcoded cells, and then paired or otherwise associated with each other. In this way, concatenated image and sequence data can be acquired for a single cell of a cell sample.

[0091] The image data and sequence data are linked by using the combinatorial barcodes of the combinatorial barcoded cells from which the image data and sequence data were obtained. For example, as described above, in the obtained sequence data, sequence reads are obtained for both the cellular targets and oligonucleotide barcode subunits of the combinatorial barcoded cells. In other words, for each combinatorial barcoded cell assayed in a given workflow, the sequences of the oligonucleotide sub-barcodes associated with that cell and the sequences of the target nucleic acids (e.g., mRNA from the cell) from that cell are obtained. For each combinatorial barcoded cell, these obtained sequences are obtained using a protocol as described above (which may be a next-generation sequencing protocol), and a library is generated from the original sequences, and each member of a given library generated from the same partition shares a common cellular label. Thus, the sequence reads from the cellular target nucleic acids and oligonucleotide sub-barcodes obtained from the same combinatorial barcoded cells all share the same cellular label, i.e., they all have a common cellular label. When linking the cell and image data, all reads from both the target nucleic acid reads and the oligonucleotide sub-barcode reads that have the same cell label domain (i.e., share a common cell label) can be paired or linked, resulting in a set of reads that includes both the target nucleic acid and the oligonucleotide sub-barcode nucleic acid reads, which can be identified as originating from the same combinatorially barcoded cell.

[0092] The resulting sequence data, including both the target nucleic acid and oligonucleotide sub-barcode nucleic acid reads, can then be matched (i.e., paired or concatenated) with the image data. As outlined above, the image data for a combinatorially barcoded cell includes a set of fluorescent signals acquired from different labeled oligonucleotides detected from a given combinatorially labeled cell during the imaging step. This set or collection of fluorescent signals acquired from the same partition can be referred to as a partition-specific fluorescent barcode. Different partitions of a given workflow will have their own unique partition-specific fluorescent barcode. A given fluorescent signal that constitutes such a partition-specific fluorescent barcode can be assigned to a given portion of the read sequence since the sequence of the labeled oligonucleotide from which the fluorescent signal is acquired is known. Thus, each partition-specific fluorescent barcode acquired for a given combinatorially barcoded cell present in that partition can be used to determine the sequence of the different image-labeled regions associated with that combinatorially barcoded cell. If the sequence of the image labeling region is present in the read of the oligonucleotide sub-barcode, then a given section-specific fluorescent barcode may be determined to be associated with a given set of sequence data. Once a section-specific fluorescent barcode is associated with a given set of sequence data, then the sequence data may be determined to be obtained from the same combinatorial barcoded cells that were in the section from which the section-specific fluorescent barcode was obtained. In other words, from a set of fluorescent signals obtained from a given section, a set of sequences of image labeling regions for the given section may be obtained. This set or set of sequences of image labeling regions may then be used to identify all sequence data obtained from that section. This identification may be done by determining that sequence reads that have both (a) a common cell barcode and (b) a section that identifies a set of sequences of image labeling regions are obtained from combinatorial cells that were in the section from which the section-specific fluorescent barcode was obtained. Once the sequence data is assigned to a given section, then the sequence data may be easily linked with image data obtained from that section.In this way, linked image and sequence data can be obtained for a single cell of a cell sample.

[0093] kit Aspects of the invention further include kits and compositions that are useful in carrying out various embodiments of the methods of the invention. Kits of the invention may include, for example, as described above, a population of specific binding members / oligonucleotide sub-barcodes, a population of labeled oligonucleotides that bind to image labeling regions of the oligonucleotide sub-barcode components of the specific binding members / oligonucleotide sub-barcodes, and beads that include bead-bound nucleic acids that include a cell labeling domain and a target binding region. A population of specific binding members / oligonucleotide sub-barcodes may include a varying number of different specific binding members / oligonucleotide sub-barcodes that differ from one another with respect to the specific binding members and / or oligonucleotide sub-barcodes (e.g., differ from one another with respect to the image labeling regions present in the sub-barcode components). The number of different specific binding members / oligonucleotide sub-barcodes in a given population may vary, but in some cases the number is in the range of 5 to 1,000 (e.g., 10 to 500). The population of labeled oligonucleotides present in the kit can also vary, and in some cases the number of different labeled oligonucleotides that differ from each other in terms of their oligonucleotide sequence and / or label ranges from 5 to 1000 (e.g., 10 to 500, e.g., 10 to 100).

[0094] The kit may further include one or more additional components useful in carrying out the embodiments of the present method. For example, the kit may include components (e.g., macrowell plates, liquid containers, e.g., tubes, etc.) used to generate combinatorial barcoded cells. In addition, the kit may include one or more components (e.g., primers, polymerases (e.g., both thermostable polymerases, reverse transcriptases with hot start properties, etc.), dsDNAse, exonucleases, dNTPs, metal cofactors, one or more nuclease inhibitors (e.g., RNase inhibitors and / or DNase inhibitors), one or more molecular crowding agents (e.g., polyethylene glycol, etc.), one or more enzyme stabilizing components (e.g., DTT), stimuli-responsive polymers, or other desired kit components, devices, solid supports, containers, cartridges, e.g., tubes, beads, plates, microfluidic chips, etc., as described above. The components of the kit may be present in separate containers, or multiple components may be present in a single container.

[0095] In addition to the above components, the subject kits may further include (in certain embodiments) instructions for carrying out the subject methods. These instructions may be present in the subject kits in a variety of forms, one or more of which may be present in the kit. One form in which these instructions may be present is as printed information on a suitable medium or substrate, such as a sheet or sheets of paper on which the information is printed, in the kit packaging, in a package insert, etc. Yet another form in which these instructions may be present is a computer readable medium on which the information is recorded, such as a diskette, a compact disc (CD), a portable flash drive, etc. Yet another form in which these instructions may be present is a website address that may be used via the Internet to access the information at a remote site.

[0096] The following are offered by way of example and not by way of limitation.

[0097] Experimental Example 1-3 provide a workflow diagram according to an embodiment of the present invention.

[0098] Notwithstanding the scope of the appended claims, the present disclosure is also defined by the following clauses. 1. A method for obtaining linked image and sequence data for a single cell of a cell sample, the method comprising: combinatorially barcoding cells of said cell sample with specific binding members / oligonucleotide sub-barcodes to generate combinatorially barcoded cells; Sorting the combinatorial barcoded cells to generate sorted combinatorial barcoded single cells, each having a combinatorial barcode; obtaining image data and sequence data for the sorted combinatorially barcoded single cells; and linking the image data and sequence data that share a common combinatorial barcode to obtain linked image and sequence data for a single cell of the cell sample.

[0099] 2. The method of clause 1, wherein combinatorial barcoding comprises one or more split / pool iterations in which cells of the cell sample are sequentially contacted with different specific binding members / oligonucleotide sub-barcodes.

[0100] 3. Each split / pool iteration allocating cells of said cell sample into different compartments; introducing different specific binding members / oligonucleotide sub-barcodes, each of which differs in its oligonucleotide sub-barcode composition, into said different compartments to generate sub-barcoded cells; and pooling said sub-barcoded cells of said different compartments.

[0101] 4. The method according to clause 3, wherein the number of different compartments is in the range of 5 to 100.

[0102] 5. The method according to any one of clauses 3 and 4, wherein the compartment is a well of a well plate.

[0103] 6. The method of any one of clauses 2-5, wherein the number of split / pool iterations is in the range of 2 to 5.

[0104] 7. The method of any one of the preceding clauses, wherein said specific binding member / oligonucleotide sub-barcode comprises a specific binding member conjugated to an oligonucleotide sub-barcode component.

[0105] 8. The method of clause 7, wherein the specific binding member comprises an antibody or a binding fragment thereof.

[0106] 9. The method of any one of clauses 7 and 8, wherein said oligonucleotide sub-barcode components comprise an image label region.

[0107] 10. The method of claim 9, wherein the oligonucleotide sub-barcode component further comprises one or more of a unique identifier for the specific binding member, a capture sequence, and a primer binding site.

[0108] 11. The method of any one of the preceding clauses, wherein said partitioning comprises distributing said combinatorially barcoded cells into partitions containing single combinatorially barcoded cells.

[0109] 12. The method of clause 11, wherein said dispensing comprises introducing said combinatorial barcoded cells into a flow cell, said flow cell having microwells at its bottom surface.

[0110] 13. The method of claim 12, further comprising providing a section containing a single combinatorially barcoded cell with a bead comprising a bead-bound nucleic acid comprising a cell labeling domain and a target binding region.

[0111] 14. The method of claim 13, wherein the bead-bound nucleic acid further comprises one or more of a molecular index domain and a universal primer binding domain.

[0112] 15. Acquiring image data for the sorted combinatorial barcoded single cells includes one or more imaging iterations, each imaging iteration comprising: contacting said partitioned combinatorial barcoded single cell with one or more labeled oligonucleotides that bind to an image label region of an oligonucleotide sub-barcode component of a specific binding member / oligonucleotide sub-barcode to generate a labeled partitioned combinatorial barcoded single cell; and capturing an image of the labeled, sorted, combinatorially barcoded single cells.

[0113] 16. The method of claim 15, wherein the sorted combinatorial barcoded single cells are contacted with 2 to 5 differently labeled oligonucleotides that bind to different image label regions.

[0114] 17. The method of any one of clauses 15 and 16, wherein the one or more labeled oligonucleotides are fluorescently labeled.

[0115] 18. The method of any one of clauses 15 to 17, wherein the number of imaging repetitions is in the range of 2 to 20.

[0116] 19. The method of any one of the preceding clauses, wherein obtaining sequence data for the sorted combinatorially barcoded single cells comprises using a next-generation sequencing protocol.

[0117] 20. The method of any one of the preceding clauses, wherein the sequence data comprises multi-omics data.

[0118] 21. A kit for obtaining linked image and sequence data for a single cell of a cell sample, the kit comprising: a population of specific binding members / oligonucleotide sub-barcodes; a population of labeled oligonucleotides that bind to the image label region of the oligonucleotide sub-barcode component of said specific binding member / oligonucleotide sub-barcode; and a bead comprising a bead-bound nucleic acid comprising a cell labeling domain and a target binding region.

[0119] 22. The kit of clause 21, wherein said specific binding member / oligonucleotide sub-barcode comprises a specific binding member conjugated to an oligonucleotide sub-barcode component.

[0120] 23. A kit according to clause 22, wherein the specific binding member comprises an antibody or a binding fragment thereof.

[0121] 24. The kit of any one of clauses 22 and 23, wherein the oligonucleotide barcode component comprises an image label region.

[0122] 25. The kit of clause 24, wherein the oligonucleotide sub-barcode component further comprises one or more of a unique identifier for a specific binding member, a capture sequence, a primer binding site.

[0123] 26. The kit of any one of clauses 21-25, wherein said labeled oligonucleotide is fluorescently labeled.

[0124] 27. The kit of any one of clauses 21 to 26, wherein the bead-bound nucleic acid further comprises one or more of a molecular index domain and a universal primer binding domain.

[0125] 28. The kit of any one of clauses 21 to 26, wherein the kit further comprises a multi-well plate.

[0126] 29. The kit of clause 28, wherein the multi-well plate comprises a 36-96 well plate.

[0127] 30. The kit of any one of clauses 21 to 29, wherein the kit further comprises a flow cell, the flow cell having a microwell at its bottom surface.

[0128] Although the foregoing inventions have been described in some detail by way of illustration and example for purposes of clarity of understanding, it will be readily apparent to those skilled in the art that, in light of the teachings of the invention, certain changes and modifications may be made thereto without departing from the spirit or scope of the appended claims.

[0129] Thus, the above is merely illustrative of the principles of the invention. It will be appreciated that those skilled in the art can devise various arrangements that embody the principles of the invention and are within its spirit and scope, although not explicitly described or illustrated herein. Furthermore, all examples and conditional language described herein are intended primarily to aid the reader in understanding the principles of the invention and the concepts the inventors contribute to furthering the art, and should not be construed as being limited to such specifically described examples and conditions. Furthermore, all statements herein that describe the principles, aspects, and embodiments of the invention, as well as specific examples thereof, are intended to encompass both structural and functional equivalents thereof. Additionally, such equivalents are intended to include both currently known equivalents and equivalents developed in the future, regardless of structure, i.e., any elements developed to perform the same function, regardless of structure. Furthermore, nothing disclosed herein is intended to be dedicated to the public, regardless of whether such disclosure is expressly set forth in the claims.

[0130] Therefore, the scope of the present invention is not intended to be limited to the exemplary embodiments shown and described herein. Rather, the scope and spirit of the present invention is embodied in the appended claims. In the claims, 35 U.S.C. 112(f) or 35 U.S.C. 112(6) is expressly defined to be invoked for a claim limitation only if the precise phrase "means for" or the precise phrase "step for" is recited at the beginning of such limitation of the claim, and if such precise phrase is not used in the claim limitation, 35 U.S.C. 112(f) or 35 U.S.C. 112(6) is not invoked.

[0131] CROSS-REFERENCE TO RELATED APPLICATIONS Pursuant to 35 U.S.C. §119(e), this application claims priority to the filing date of U.S. Provisional Patent Application No. 63 / 332,087, filed April 18, 2022, the disclosure of which is incorporated herein by reference in its entirety.

Claims

1. 1. A method for obtaining linked image and sequence data for a single cell of a cell sample, the method comprising: combinatorially barcoding cells of said cell sample with specific binding members / oligonucleotide sub-barcodes to generate combinatorially barcoded cells; Sorting the combinatorial barcoded cells to generate sorted combinatorial barcoded single cells, each having a combinatorial barcode; obtaining image data and sequence data for the sorted combinatorially barcoded single cells; and linking the image data and sequence data that share a common combinatorial barcode to obtain linked image and sequence data for a single cell of the cell sample.

2. 2. The method of claim 1, wherein combinatorial barcoding comprises one or more split / pool iterations in which cells of the cell sample are sequentially contacted with different specific binding members / oligonucleotide sub-barcodes.

3. Each split / pool iteration allocating cells of said cell sample into different compartments; introducing different specific binding members / oligonucleotide sub-barcodes, each of which differs in its oligonucleotide sub-barcode composition, into said different compartments to generate sub-barcoded cells; and pooling the sub-barcoded cells of the different compartments.

4. The method of claim 3 , wherein the compartment is a well of a well plate.

5. 5. The method of any one of claims 1 to 4, wherein the specific binding member / oligonucleotide sub-barcode comprises a specific binding member conjugated to an oligonucleotide sub-barcode component.

6. The method of claim 5 , wherein the specific binding member comprises an antibody or a binding fragment thereof.

7. The method of claim 6 , wherein the oligonucleotide sub-barcode component comprises an image label region.

8. 8. The method of claim 7, wherein the oligonucleotide sub-barcode components further comprise one or more of a unique identifier for the specific binding member, a capture sequence, and a primer binding site.

9. 9. The method of any one of claims 1 to 8, wherein said partitioning comprises distributing the combinatorially barcoded cells into partitions containing single combinatorially barcoded cells.

10. 10. The method of claim 9, wherein the dispensing comprises introducing the combinatorial barcoded cells into a flow cell, the flow cell having microwells on its bottom surface.

11. 11. The method of claim 10, further comprising providing a section containing a single combinatorially barcoded cell with a bead comprising a bead-bound nucleic acid comprising a cell labeling domain and a target binding region.

12. Acquiring image data for the sorted combinatorial barcoded single cells comprises one or more imaging iterations, each imaging iteration comprising: contacting the partitioned combinatorial barcoded single cell with one or more labeled oligonucleotides that bind to an image label region of an oligonucleotide sub-barcode component of a specific binding member / oligonucleotide sub-barcode to generate a labeled partitioned combinatorial barcoded single cell; and capturing an image of the labeled, sorted, combinatorially barcoded single cells.

13. 13. The method of any one of claims 1 to 12, wherein obtaining sequence data for the sorted combinatorially barcoded single cells comprises using a next generation sequencing protocol.

14. The method of any one of claims 1 to 13, wherein the sequence data comprises multi-omics data.

15. 1. A kit for obtaining linked image and sequence data for a single cell of a cell sample, the kit comprising: a population of specific binding members / oligonucleotide sub-barcodes; a population of labeled oligonucleotides that bind to image label regions of the oligonucleotide sub-barcode components of said specific binding member / oligonucleotide sub-barcode; and a bead comprising a bead-bound nucleic acid comprising a cell labeling domain and a target binding region.