Dual-indexed specific binding members for obtaining linked single-cell cytometry and sequence data.
The use of dual-indexed specific binding members with fluorescent and oligonucleotide barcodes addresses the integration of flow cytometry data with sequence data, enabling comprehensive cell analysis by providing unique signatures for each cell, thereby enhancing data correlation and analysis efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-26
- Publication Date
- 2026-03-13
AI Technical Summary
Current technologies lack the ability to seamlessly integrate flow cytometry data with downstream sequence data, such as multi-omics data, for comprehensive cell analysis.
A method utilizing dual-indexed specific binding members with fluorescent and oligonucleotide barcodes to index individual cells, allowing for the correlation of fluorescence data with sequence data through flow cytometry and sequencing workflows.
Enables the linkage of flow cytometry data with sequence data, facilitating comprehensive analysis of individual cells by providing unique fluorescent and oligonucleotide signatures for each cell, enhancing data correlation and analysis efficiency.
Smart Images

Figure 2026508864000001_ABST
Abstract
Description
[Background technology]
[0001] Current technology enables large-scale, parallel (e.g., >10,000 cells) single-cell gene expression measurement by attaching cell-specific oligonucleotide barcodes to poly(A)mRNA molecules from individual cells as each cell colocalizes with barcoded reagent beads within its compartment. One platform that enables large-scale parallel single-cell gene expression measurement is the BD Rhapsody® Single-Cell Analysis System. The BD Rhapsody® Single-Cell Analysis System is a platform that enables high-throughput capture of nucleic acids from single cells using a simple cartridge workflow and a multi-layer barcoding system. The captured material can be used to generate various types of next-generation sequencing (NGS) libraries, including libraries suitable for whole-transcriptome analysis, such as discovery biology and targeted RNA analysis for highly sensitive transcript detection. Shum et al., "Quantitation of mRNA Transcripts and Proteins Using the BD Rhapsody" TM Single-Cell Analysis System”, Adv Exp Med Biol.2019;1129:63-79.
[0002] Gene expression can influence protein expression. Protein-protein interactions can influence both gene expression and protein expression. Thus, more recently, systems and methods have been developed that can quantitatively analyze protein expression in cells and simultaneously measure protein and gene expression in cells. One such platform is the BD Abseq platform. AbSeq is a method for profiling proteins within and on single cells. In Abseq, conventional fluorophore labeling on antibodies is replaced with nucleic acid sequence tags that can be read at the single-cell level, for example, via barcoding and NGS sequencing. "The objective of Abseq is to enable highly sensitive, accurate, and comprehensive characterization of proteins along with mRNA transcripts in a large number of single cells. Cells bind to antibodies against different target epitopes, similar to conventional immunohistochemistry, except that the antibodies are labeled with unique sequence tags. When an antibody binds to its target, the DNA tag is carried along with it, and the presence of the target can be inferred based on the presence of the tag. In this way, the counting tags provide an estimate of different epitopes present in the cell, so that they can be detected via antibody binding." Shahi et al., "Abseq: Ultrahigh-throughput single cell protein profiling with droplet microfluidic barcoding. Sci Rep 7, 44447 (2017)"
[0003] Flow cytometry is a technique used to characterize the physical and / or chemical properties of cell samples, for example, for detection, measurement, etc. In flow cytometry, a cell sample is suspended in a fluid and injected into a flow cytometer instrument. The sample is focused so that, ideally, one cell at a time flows through a laser beam, and the scattered light is characteristic of the cells and their components. Cells are often labeled with fluorescent markers (e.g., antibody-fluorophores) so that the light is absorbed and then emitted in a specific wavelength range. Tens of thousands of cells can be rapidly examined, and data can be collected from them. Flow cytometry is routinely used in basic research, clinical practice, and clinical trials. Applications in which flow cytometry is used include, but are not limited to, cell counting, cell sorting, determination of cell characterization and function, detection of microorganisms, biomarker detection, protein manipulation detection, diagnosis of health problems, and measurement of genome size. A flow cytometry analyzer is an instrument that provides quantifiable data from a sample. Other instruments that use flow cytometry include cell sorters that physically separate the target cells and thereby purify them based on their optical properties. [Overview of the project]
[0004] The inventors recognized the desirable ability to link flow cytometry data with downstream sequence data (e.g., multi-omics data) so that flow cytometry and sequence data, including image cytometry data, can be easily obtained for the same cell. A novel method for indexing individual cells to correlate fluorescence data (immunophenotyping, functional, or otherwise) or bright-field imaging data with downstream single-cell sequence data (e.g., multi-omics) is described herein. Embodiments of the present invention provide the ability to select desired live or preserved cells via imaging and / or fluorescence analysis (e.g., fluorescence-activated cell sorting) and then obtain sequence data for such selected cells.
[0005] For example, methods and compositions are provided for preparing indexed cell populations that can be used in protocols to obtain linked single-cell flow cytometry data and sequence (multi-omic, etc.) data. Embodiments of the method include dividing a cell sample into first plurality of parts; combining different parts of the first plurality of parts with a separate first dual-indexed specific binding member having separate fluorescent and oligonucleotide barcodes to stably associate the cells of the part with the first dual-indexed specific binding member; combining the first plurality of parts to produce a first pool; dividing the first pool into second plurality of parts; and combining different parts of the second plurality of parts with a separate second dual-indexed specific binding member having separate fluorescent and oligonucleotide barcodes to stably associate the cells of the part with the second dual-indexed specific binding member to produce an indexed cell population. Individual cell indexing is a result of the unique combinatorial labeling of the cells. In various embodiments, the resulting indexed cell population can then be subjected to flow cytometry and sequencing workflows, and the resulting flow cytometry and sequencing data can be linked. Compositions for use in carrying out this method are also provided.
[0006] Aspects of the present invention utilize dual-indexed specific binding members having associated oligo barcodes and fluorescent addresses (i.e., fluorescent barcodes). Aspects of the method utilize a large and diverse set of dual-indexed specific binding members. Within the set, unique oligo barcodes are directly associated with each unique fluorescent barcode. In various embodiments, dual-indexed specific binding members are randomly associated with cells via a combinatorial protocol, such as a split / pool protocol. The association of random combinations of dual-indexed specific binding members with individual cells provides unique combinations of barcodes associated with cells in the sample, resulting in a large, if not all, of the cells having a unique fluorescent signature consisting of one or more barcodes of the dual-indexed specific binding members associated with that cell. The fluorescent barcodes of the cell-associated specific binding members can then be analyzed using the same modality (such as flow cytometry) that collects fluorescence data from the cells. Subsequently, the oligo barcodes of dual-indexed specific binding members can be released and captured by complementary oligos in a single-cell multi-omics workflow, simultaneously with the capture of endogenous nucleic acids and other cell-associated barcodes (e.g., antibody-associated barcodes) during cell lysis. Oligo barcodes from dual-indexed specific binding members can be analogous to sample multiplexing oligos in routine cell hashing experiments. Then, using standard single-cell multi-omics library preparation, sequencing protocols, and data analysis, combinations of oligo barcodes unique to all cells in the experiment can be identified. Direct association between all combo-oligo barcodes and all combo-fluorescent barcodes allows NGS single-cell data to correlate one-to-one with flow cytometry single-cell data. [Brief explanation of the drawing]
[0007] The present invention can be best understood from the following detailed description when read in conjunction with the accompanying drawings. The drawings include the following figures.
[0008] [Figure 1-1] Provide a schematic diagram of a workflow according to an embodiment of the present invention. [Figure 1-2] Provide a schematic diagram of a workflow according to an embodiment of the present invention. [Figure 2-1] Provide a schematic diagram of a workflow according to another embodiment of the present invention. [Figure 2-2] Provide a schematic diagram of a workflow according to another embodiment of the present invention. [Figure 3-1] Provide details regarding dual-indexed reagents that can be used in embodiments of the present invention. [Figure 3-2] Provide details regarding dual-indexed reagents that can be used in embodiments of the present invention.
Modes for Carrying Out the Invention
[0009] Definitions Unless otherwise defined, technical terms and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this disclosure belongs. See, for example, Singleton et al., Dictionary of Microbiology and Molecular Biology 2nd ed., J. Wiley & Sons (New York, NY 1994); Sambrook et al., Molecular Cloning, A Laboratory Manual, Cold Spring Harbor Press (Cold Spring Harbor, NY 1989). For the purposes of this disclosure, the following terms are defined below.
[0010] As used herein, an antibody may be a full-length immunoglobulin molecule (e.g., an IgG antibody) or an immunologically active (i.e., specifically binding) portion of an immunoglobulin molecule, such as an antibody fragment, or an antibody fragment. In some embodiments, the antibody is a functional antibody fragment. For example, an antibody fragment may be a portion of an antibody such as F(ab')2, Fab', Fab, Fv, or sFv. The antibody fragment can bind to the same antigen recognized by a full-length antibody. The antibody fragment may include isolated fragments consisting of the variable region of an antibody, such as an “Fv” fragment consisting of the variable regions of the heavy and light chains, and recombinant single-chain polypeptide molecules ("scFv proteins") in which the light and heavy chain variable regions are linked by a peptide linker. Exemplary antibodies may include, but are not limited to, antibodies against cancer cells, antibodies against viruses, antibodies that bind to cell surface receptors (e.g., CD8, CD34, and CD45), and therapeutic antibodies.
[0011] As used herein, the terms “associated” or “associated with” can mean that two or more species are identifiable as being in the same location at a given time. An association can mean that two or more species are, or were, in similar containers. An association can be an association in information science. For example, digital information about two or more species can be stored and used to determine that one or more species were located in the same location at a given time. An association can also be a physical association. In some embodiments, two or more associated species are “attached,” “adhered,” or “fixed” to each other or to a common solid or semi-solid surface. An association can refer to a shared or non-shared means for attaching a label to a solid or semi-solid support, such as beads. An association can be a covalent bond between a target and a label. An association can include hybridization between two molecules (such as a target molecule and a label).
[0012] As used herein, the term "complementary" can refer to the ability of exact pairing between two nucleotides. For example, when the nucleotide at a given location of a nucleic acid can hydrogen bond with the nucleotide of another nucleic acid, the two nucleic acids are considered to be complementary to each other at that location. Complementarity between two single-stranded nucleic acid molecules may be "partial" where only a portion of the nucleotides bind, or may be complete if there is complete complementarity between the single-stranded molecules. A first nucleotide sequence can be said to be the "complement" of a second sequence if the first nucleotide sequence is complementary to the second nucleotide sequence. A first nucleotide sequence can be said to be the "reverse complement" of a second sequence if the first nucleotide sequence is complementary to a sequence that is the reverse of the second sequence (i.e., the order of the nucleotides is reversed). As used herein, the terms "complement", "complementary", and "reverse complement" can be used interchangeably from this disclosure, it is understood that if a molecule can hybridize to another molecule, it can be the complement of the molecule to which it is hybridizing.
[0013] As used herein, the term “nucleic acid” refers to a polynucleotide sequence or a fragment thereof. Nucleic acids may contain nucleotides. Nucleic acids may be exogenous or endogenous to cells. Nucleic acids may exist in a cell-free environment. Nucleic acids may be genes or fragments thereof. Nucleic acids may be DNA. Nucleic acids may be RNA. Nucleic acids may contain one or more analogues (e.g., modified skeletons, sugars, or nucleic acid bases). Some non-limiting examples of analogues include 5-bromouracil, peptide nucleic acids, heteronucleotides, morpholino, locked nucleic acids, glycol nucleic acids, threose nucleic acids, dideoxynucleotides, cordycepin, 7-deaza-GTP, fluorophores (e.g., rhodamine or fluorescein linked to sugars), thiol-containing nucleotides, biotin-linked nucleotides, fluorescent base analogues, CpG islands, methyl-7-guanosine, methylated nucleotides, inosine, thiouridine, pseudouridine, dihydrouridine, cuosin, and viosin. The terms "nucleic acid," "polynucleotide," "targeted polynucleotide," and "targeted nucleic acid" can be used interchangeably.
[0014] Nucleic acids can include one or more modifications (e.g., base modifications, skeletal modifications) to provide new or enhanced characteristics to the nucleic acid (e.g., improved stability). Nucleic acids can include nucleic acid affinity tags. A nucleoside can be a base-sugar combination. The base portion of a nucleoside can be a heterocyclic base. Two of the most common classes of such heterocyclic bases are purines and pyrimidines. A nucleotide can be a nucleoside that further includes a phosphate group covalently bonded to the sugar portion of the nucleoside. In the case of a nucleoside containing a pentofuranosyl sugar, the phosphate group can be linked to the 2', 3', or 5' hydroxyl portion of the sugar. When forming nucleic acids, the phosphate group can covalently bond adjacent nucleosides to each other to form a linear polymer compound. The respective ends of this linear polymer compound can then be further joined to form a cyclic compound. However, linear compounds are generally preferred. Furthermore, linear compounds may have internal nucleotide-base complementarity and thus may be folded to produce fully or partially double-stranded compounds. Within nucleic acids, phosphate groups can generally be said to form the internucleoside skeleton of the nucleic acid. The bond or skeleton may be a 3' to 5' phosphodiester bond.
[0015] Nucleic acids may contain modified skeletons and / or modified nucleoside bonds. Modified skeletons may include skeletons that retain phosphorus atoms and skeletons that do not contain phosphorus atoms. Suitable modified nucleic acid skeletons containing a phosphorus atom include, for example, phosphorothioates, chiral phosphorothioates, phosphodithioates, phosphotriesters, aminoalkyl phosphotriesters, methyl and other alkylphosphonates, such as 3'-alkylene phosphonates, 5'-alkylene phosphonates, chiral phosphonates, phosphinates, phosphoramides including 3'-aminophosphoramides and aminoalkylphosphoramides, phosphorodiamidates, thionophosphoramides, thionoalkyl phosphonates, thionoalkyl phosphotriesters, selenophosphates, as well as boranophosphates having the usual 3'-5' linkage and 2'-5' linkage analogues, and boranophosphates having reverse polarity where one or more internucleotide bonds are 3'-to-3', 5'-to-5', or 2'-to-2' links.
[0016] Nucleic acids can include polynucleotide skeletons formed by short-chain alkyl or cycloalkyl nucleoside bonds, mixed heteroatoms and alkyl or cycloalkyl nucleoside bonds, or one or more short-chain heteroatoms or heterocyclic nucleoside bonds. These can include morpholino bonds (partially formed from the sugar moieties of nucleosides); siloxane skeletons; sulfide, sulfoxide, and sulfone skeletons; formacetyl and thioformacetyl skeletons; methyleneformacetyl and thioformacetyl skeletons; riboacetyl skeletons; alkene-containing skeletons; sulfamic acid skeletons; methyleneimino and methylenehydrazino skeletons; sulfonic acid and sulfonamide skeletons; those having amide skeletons; and others having mixed N, O, S, and CH2 constituent parts.
[0017] Nucleic acids can include nucleic acid mimes. The term “mimicking” may refer to polynucleotides in which only the furanose ring, or both the furanose ring and the internucleotide bond, are substituted with a non-furanose group; substitution of only the furanose ring may also be called a sugar substitute. Heterocyclic base moieties or modified heterocyclic base moieties can be retained for hybridization with a suitable target nucleic acid. One such nucleic acid may be peptide nucleic acid (PNA). In PNA, the sugar backbone of a polynucleotide can be replaced with an amide-containing backbone, particularly an aminoethylglycine backbone. Nucleic acids can be retained and bond directly or indirectly to the aza nitrogen atom of the amide moiety of the backbone. The backbone in a PNA compound may contain two or more linked aminoethylglycine units that give PNA an amide-containing backbone. The heterocyclic base moiety bonds directly or indirectly to the aza nitrogen atom of the amide moiety of the backbone.
[0018] Nucleic acids can contain a morpholino backbone structure. For example, a nucleic acid can contain a six-membered morpholino ring instead of a ribose ring. In some of these embodiments, a phosphorodiamidate or other non-phosphodiester nucleoside bond can replace the phosphodiester bond.
[0019] Nucleic acids can contain linked morpholino units (e.g., morpholino nucleic acids) having heterocyclic bases attached to a morpholino ring. Linking groups can link morpholino monomer units in morpholino nucleic acids. Nonionic morpholino-based oligomeric compounds may have fewer undesirable interactions with cellular proteins. Morpholino-based polynucleotides can be nonionic mimics of nucleic acids. Various compounds within the morpholino class can be joined using different linking groups. A further class of polynucleotide mimics can be called cyclohexenyl nucleic acids (CeNA). The furanose ring normally present in nucleic acid molecules can be replaced with a cyclohexenyl ring. CeNA DMT-protected phosphoramidite monomers can be prepared and used in the synthesis of oligomeric compounds using phosphoramidite chemistry. Incorporation of CeNA monomers into nucleic acid chains can enhance the stability of DNA / RNA hybrids. CeNA oligoadenylates can form complexes with nucleic acid complements that have similar stability to natural complexes. Further modifications include locked nucleic acids (LNAs) in which a 2'-hydroxyl group is linked to the 4' carbon atom of the sugar ring, thereby forming a 2'-C,4'-C-oxymethylene bond and thus a bicyclic sugar moiety. The bond can be a methylene group (-CH2) or a group that bridges the 2' oxygen atom and the 4' carbon atom (wherein n is 1 or 2). LNAs and LNA analogs can exhibit very high double-chain thermal stability with complementary nucleic acids (Tm = +3 to +10°C), stability against 3'-exonuclease degradation, and good solubility.
[0020] Nucleic acids may also include modifications or substitutions of nucleic acid bases (often simply called “bases”). As used herein, “unmodified” or “natural” nucleic acid bases may include purine bases (e.g., adenine (A) and guanine (G)) and pyrimidine bases (e.g., thymine (T), cytosine (C) and uracil (U)). Modified nucleic acid bases may include other synthetic and natural nucleic acid bases, e.g., 5-methylcytosine (5-me-C), 5-hydroxymethylcytosine, xanthine, hypoxanthine, 2-aminoadenine, 6-methyl and other alkyl derivatives of adenine and guanine, 2-propyl and other alkyl derivatives of adenine and guanine, 2-thiouracil, 2-thiothymine and 2-thiocytosine, 5-halouracil and cytosine, 5-propynyl (-C=C-CH3)uracil and cytosine, as well as other alkynyl derivatives of pyrimidine bases, 6-azouracil, cytosine It may contain tosine and thymine, 5-uracil (pseudouracil), 4-thiouracil, 8-halo, 8-amino, 8-thiol, 8-thioalkyl, 8-hydroxyl and other 8-substituted adenines and guanines, 5-halo, especially 5-bromo, 5-trifluoromethyl and other 5-substituted uracils and cytosine, 7-methylguanine and 7-methyladenine, 2-F-adenine, 2-aminoadenine, 8-azaguanine and 8-azaadenine, 7-deazaguanine and 7-deazaadenine, as well as 3-deazaguanine and 3-deazaadenine.Modified nucleic acid bases include tricyclic pyrimidines, e.g., phenoxazinecytidine (1H-pyrimido(5,4-b)(1,4)benzoxazine-2(3H)-one), phenothiazinecytidine (1H-pyrimido(5,4-b)(1,4)benzothiadin-2(3H)-one), G-clamps, e.g., substituted phenoxazinecytidine (e.g., 9-(2-aminoethoxy)-H-pyrimido(5,4-(b)(1,4)benzoxazine-2(3H)-one), phenothiazinecytidine (1 This may include H-pyrimide(5,4-b)(1,4)benzothiadin-2(3H)-one), G-clamps, such as substituted phenoxazine cytidines (e.g., 9-(2-aminoethoxy)-H-pyrimide(5,4-(b)(1,4)benzoxazine-2(3H)-one), carbazole cytidine (2H-pyrimide(4,5-b)indole-2-one), and pyridoindole cytidine (H-pyrimide(3',2':4,5)pyrrolo[2,3-d]pyrimidine-2-one).
[0021] As used herein, the term “sample” may refer to a composition containing a target. Samples suitable for analysis by the methods, devices, and systems disclosed include cells, tissues, organs, or organisms. Cell samples are compositions composed of multiple cells, such as compositions containing multiple different cells, such as single-cell aqueous compositions, and the number of cells may vary.
[0022] As used herein, the terms “sampling device” or “device” may refer to a device capable of taking a portion of a sample and / or placing that portion onto a substrate. Sample devices may refer to, for example, fluorescence-activated cell sorting (FACS) machines, cell sorters, biopsy needles, biopsy devices, tissue sectioning devices, microfluidic devices, blade grids, and / or microtomes.
[0023] As used herein, the term “solid support” may refer to an individual solid or semi-solid surface to which nucleic acids can adhere. A solid support may encompass any type of solid, porous, or hollow sphere, ball, bearing, cylinder, or other similar configuration made of plastic, ceramic, metal, or polymer material (e.g., hydrogel) to which nucleic acids can be immobilized (e.g., covalently or non-covalently). A solid support may comprise individual particles that are spherical (e.g., microspheres) or have non-spherical or irregular shapes such as cubes, cuboids, pyramidal, cylindrical, conical, oval, or disc-shaped. Beads may be non-spherical. Multiple solid supports spaced apart in an array may not include a substrate. The term “solid support” may be used interchangeably with the term “beads.”
[0024] As used herein, the term “target” may refer to a composition that can be analyzed according to embodiments of the present invention. Exemplary suitable targets for analysis by the disclosed methods, devices, and systems include oligonucleotides, DNA, RNA, mRNA, microRNA, tRNA, and the like. Targets may be single-stranded or double-stranded. In some embodiments, targets may be proteins, peptides, or polypeptides. In some embodiments, targets may be lipids. As used herein, “target” may be used interchangeably with “species.”
[0025] As used herein, the term “reverse transcriptase” may refer to a group of enzymes that possess reverse transcriptase activity (i.e., those that catalyze the synthesis of DNA from an RNA template). Generally, such enzymes include, but are not limited to, retroviral reverse transcriptases, retrotransposon reverse transcriptases, retroplasmid reverse transcriptases, retron reverse transcriptases, bacterial reverse transcriptases, group II intron-derived reverse transcriptases, and their variants, variants, or derivatives. Non-retroviral reverse transcriptases include non-LTR retrotransposon reverse transcriptases, retroplasmid reverse transcriptases, retron reverse transcriptases, and group II intron reverse transcriptases. Examples of group II intron reverse transcriptases include Lactococcus lactis LI.LtrB intron reverse transcriptase, Thermosynechococcus elongatus TeI4c intron reverse transcriptase, or Geobacillus stearothermophilus GsI-IIC intron reverse transcriptase. Other classes of reverse transcriptases may include many classes of non-retroviral reverse transcriptases (i.e., retrons, group II introns, and diversity-generating retroelements, among others).
[0026] The term "specific binding" refers to the direct association between two molecules by covalent interaction, electrostatic interaction, hydrophobic interaction, and ionic and / or hydrogen bond interaction, including interactions such as salt bridge and water bridge. Specific binding members describe the members of a pair of molecules having binding specificity for each other. The members of a specific binding pair may be of natural origin or may be produced synthetically, either wholly or in part. One member of the pair of molecules has a region on its surface or cavity that specifically binds to the specific spatial and polar configuration of the other member of the pair of molecules and is thus complementary. Thus, the members of the pair have the property of specifically binding to each other. Examples of specific binding member pairs are antigen-antibody, biotin-avidin, hormone-hormone receptor, receptor-ligand, enzyme-substrate. The specific binding members of the binding pair exhibit high affinity and binding specificity for binding to each other. Typically, the affinity between the specific binding members of the pair is 10 -6 M or less, for example, 10 -7 M or less (including 10 -8 M or less), for example, 10 -9 M or less, 10 -10 M or less, 10 -11 M or less, 10 -12 M or less, 10 -13 M or less, 10 -14 M or less (including 10 -15 M or less) of K d (dissociation constant). "Affinity" refers to the strength of binding, and increased binding affinity correlates with a lower KD. In one embodiment, the affinity is determined by surface plasmon resonance (SPR) as used, for example, by a Biacore system. The affinity of one molecule for another molecule is determined, for example, by measuring the binding rate of the interaction at 25°C. "Affinity" refers to the strength of binding, and increased binding affinity correlates with a lower KD. In one embodiment, the affinity is determined by surface plasmon resonance (SPR) as used, for example, by a Biacore system. The affinity of one molecule for another molecule is determined, for example, by measuring the binding rate of the interaction at 25°C.
[0027] The specific binding member may be proteinaceous. As used herein, the term “proteinaceous” refers to a portion consisting of amino acid residues. The proteinaceous portion may be a polypeptide. In certain cases, the proteinaceous specific binding member is an antibody. In certain embodiments, the proteinaceous specific binding member is an antibody fragment, e.g., a binding fragment of an antibody that specifically binds to a target analyte. As used herein, the terms “antibody” and “antibody molecule” are used interchangeably and refer to a protein consisting of one or more polypeptides substantially encoded by all or part of a recognized immunoglobulin gene. Recognized immunoglobulin genes include, for example in humans, the kappa (k), lambda (l), and heavy chain loci (together comprising a multitude of variable region genes), as well as the constant region genes mu (u), delta (d), gamma (g), sigma (e), and alpha (a) (encoding IgM, IgD, IgG, IgE, and IgA isotypes, respectively). The immunoglobulin light or heavy chain variable region consists of a framework region (FR) interrupted by three hypervariable regions, also called “complementarity-determining regions” or “CDRs.” The extent of the framework region and CDRs is precisely defined (see “Sequences of Proteins of Immunological Interest,” E. Kabat et al., USD Department of Health and Human Services, (1991)). The numbering of all antibody amino acid sequences discussed herein follows the Kabat system. The sequences of different light or heavy chain framework regions are relatively conserved within species. The antibody framework region, i.e., the combined framework region of the constituent light and heavy chains, helps to position and align the CDRs. The CDRs are primarily responsible for binding to the antigen's epitope. The term “antibody” means including full-length antibodies and may refer to any naturally occurring antibody, engineered antibody, or recombinantly produced antibody of any biological origin, as further defined below, for experimental, therapeutic, or other purposes.
[0028] The antibody fragment of interest may include, but is not limited to, Fab, Fab', F(ab')2, Fv, scFv, or other antigen-binding subsequences of an antibody produced by modification of the whole antibody or novel synthesis using recombinant DNA technology. The antibody may be monoclonal or polyclonal and may have other specific activity against cells (e.g., antagonist, agonist, neutralizing, inhibitory, or stimulating antibody). It is understood that the antibody may have additional conserved amino acid substitutions that do not substantially affect antigen binding or other antibody functions.
[0029] In certain embodiments, the specific binding member is a Fab fragment, an F(ab')2 fragment, an scFv, a diabody, or a triabody. In certain embodiments, the specific binding member is an antibody. In some cases, the specific binding member is a mouse antibody or its binding fragment. In certain cases, the specific binding member is a recombinant antibody or its binding fragment.
[0030] For example, methods and compositions are provided for preparing an indexed cell population that can be used in protocols to obtain linked single-cell flow cytometry data and sequencing (multi-omic, etc.) data. Embodiments of the method include dividing a cell sample into a first plurality of parts, combining different parts of the first plurality of parts with a separate first dual-indexed specific binding member having separate fluorescent and oligonucleotide barcodes to stably associate the cells of the part with the first dual-indexed specific binding member, combining the first plurality of parts to produce a first pool, dividing the first pool into a second plurality of parts, and combining different parts of the second plurality of parts with a separate second dual-indexed specific binding member having separate fluorescent and oligonucleotide barcodes to stably associate the cells of the part with the second dual-indexed specific binding member to produce an indexed cell population. In various embodiments, the obtained indexed cell population can then be subjected to flow cytometry and sequencing workflows to link the obtained flow cytometry and sequencing data. Compositions for use when carrying out the method are also provided.
[0031] Before describing the present invention in more detail, it should be understood that the present invention is not limited to the specific embodiments described and is therefore naturally subject to change. It should also be understood that the scope of the present invention is limited only by the appended claims, and therefore the terms used herein are intended solely to describe specific embodiments and are not intended to limit them.
[0032] Where a range of values is provided, unless otherwise explicitly stated in the context, each value between the upper and lower limits of that range, up to one-tenth of the lower limit, and any other values or values in between within that stated range are understood to be included in the invention. The upper and lower limits of these smaller ranges may independently be included within smaller ranges and are also included in the invention, subject to any specifically excluded limits of the stated range. Where a stated range includes one or both limits, the range excluding one or both of those limits is also included in the invention.
[0033] In this specification, certain ranges are presented with the term “approximately” preceding the number. The term “approximately” is used herein to provide literal support for the exact number it precedes, as well as for any number that is close to or approximates the number it precedes. In determining whether a number is close to or approximates a specifically enumerated number, any close or approximate number that is not enumerated may be a number that, in the context in which the number is presented, presents a substantial equivalent of the specifically enumerated number.
[0034] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those generally understood by those skilled in the art to which the present invention pertains. Any methods and materials similar or equivalent to those described herein may also be used in carrying out or testing the present invention, but representative exemplary methods and materials are described herein.
[0035] All publications and patents cited herein are incorporated herein by reference as if each individual publication or patent were specifically and individually indicated to be incorporated by reference, and are incorporated herein by reference to disclose and describe the manner and / or materials by which the publications are cited. Any reference to a publication is for the purpose of its disclosure prior to the filing date and should not be construed as an acknowledgment that the present invention has no prior rights to such publication by prior art. Furthermore, the published dates presented may differ from the actual published dates and may need to be independently verified.
[0036] Where used herein and in the appended claims, the singular forms “a,” “an,” and “the” refer to multiple subjects unless otherwise explicitly stated in the context. Furthermore, it should be noted that claims may be drafted to exclude any one of their elements. Therefore, this statement is intended to serve as an antecedent for the use of exclusive terms such as “exclusively” and “only” in relation to the enumeration of elements of the claims or the use of “negative” limitations.
[0037] As will be obvious to those skilled in the art upon reading this disclosure, each of the individual embodiments described and illustrated herein has separate components and features that can be readily separated or combined with features of any of several other embodiments without departing from the scope or spirit of the invention. Any of the listed methods may be carried out in the order of the listed events, or in any other logically possible order.
[0038] While systems and methods are described for grammatical fluidity with functional descriptions, claims should not necessarily be interpreted as being limited by constructing a limitation of “means” or “steps” unless explicitly formulated under 35 U.S. SC § 112, and should be given the meaning of the definitions and the full scope of equivalents provided by the claims under the doctrine of equivalents, and if the claims are explicitly formulated under 35 U.S. SC § 112, it should be clearly understood that a full set of statutory equivalents should be given under 35 U.S. SC § 112.
[0039] method As summarized above, a method for preparing an indexed cell population is provided. In some cases, the indexed cell population consists of multiple distinctly indexed cells, each having a different or unique fluorescent signature associated with it. "Fluorescent signature" (i.e., fluorescent identifier) means a composite or aggregate spectrum provided by, for example, a cell-associated dual-indexed specific binding member (described in more detail below), consisting of one or more fluorescence emission signals obtained from one or more fluorophores stably associated with the indexed cell. Different cells in the indexed cell population produced by the method of the embodiments of the present invention have distinct fluorophores that constitute their associated fluorescent identifiers, and therefore provide different fluorescent signatures when assayed, for example, by a flow cytometry protocol. Thus, different indexed cells in a population are distinguishable from one another by the presence of their unique fluorescent signatures.
[0040] Dual-indexed specific binding members As described above, the fluorescent signature of a given indexed cell produced by an embodiment of the present invention is provided by a dual-indexed specific binding member stably associated with the cell. The dual-indexed specific binding member is a specific binding member comprising one or more fluorophores, the one or more fluorophores collectively constituting a fluorescent barcode and an oligonucleotide barcode of the dual-indexed specific binding member. The dual-indexed specific binding member may vary as desired and may include a specific binding member component, a fluorescent barcode component, and an oligonucleotide barcode component.
[0041] Specific binding members The specific binding member components of a dual-indexed specific binding member can vary, and suitable specific binding members include, but are not limited to, those listed above. In some cases, the specific binding member is an antibody or antibody fragment, e.g., a binding fragment of an antibody that specifically binds to a target analyte. As used herein, the terms “antibody” and “antibody molecule” are interchangeable and refer to a protein consisting of one or more polypeptides substantially encoded by all or part of a recognized immunoglobulin gene. The antibody fragment of interest may include, but is not limited to, Fab, Fab', F(ab')2, Fv, scFv, or other antigen-binding subsequences of an antibody produced by modification of the whole antibody or novelly synthesized using recombinant DNA technology. Antibodies may be monoclonal or polyclonal and may have other specific activities against cells (e.g., antagonist, agonist, neutralizing, inhibitory, or stimulating antibody). It is understood that antibodies may have additional conserved amino acid substitutions that do not substantially affect antigen binding or other antibody functions. In certain embodiments, the specific binding member is a Fab fragment, an F(ab')2 fragment, an scFv, a diabody, or a triabody. In certain embodiments, the specific binding member is an antibody. In some cases, the specific binding member is a mouse antibody or its binding fragment. In certain cases, the specific binding member is a recombinant antibody or its binding fragment.
[0042] fluorescent barcode A given dual-indexed specific binding member fluorescent barcode contains one or more fluorescent dyes. If a given fluorescent barcode contains more than one fluorescent dye, two or more fluorescent dyes collectively constitute the fluorescent barcode of the dual-indexed beads. Thus, a given fluorescent barcode may consist of a single fluorescent dye, or two or more fluorescent dyes that collectively constitute the fluorescent barcode of the beads, e.g., 2-20, e.g., 2-10, e.g., 2-5, e.g., 2-3 fluorescent dyes. Thus, the number of different fluorophores constituting a given fluorescent barcode can vary, ranging in some cases from 1-10, e.g., 1-5, and including 1-3. Any two given distinguishable fluorescent barcodes may be distinguishable from one another based on the type of fluorophores and / or the signal brightness provided thereby. Thus, any two distinguishable fluorescent barcodes may be distinguishable based on the fluorescence signal (e.g., maximum emission wavelength) and / or intensity and / or amount) of the fluorescent dyes that collectively constitute the fluorescent barcode. For example, two distinguishable fluorescent barcodes may be distinguishable from each other because they consist of combinations of different types of fluorophore dyes, e.g., one containing fluorophores a, b, and c, and the other containing fluorophores b, c, and d. Two distinguishable fluorescent barcodes may also be distinguishable from each other because they consist of different amounts of fluorescent dyes, e.g., one consisting of fluorophores a, b, and c present in a first amount in a given dual-indexed specific binding member, and the other consisting of fluorophores present in a second amount different from the first amount, e.g., a value that can be detected by the difference in signal brightness. Using combinations of fluorophore types and amounts, any desired number of unique fluorescent barcodes may be provided.
[0043] A fluorescent barcode may contain one or more fluorophores as desired. Therefore, a dual-indexed specific binding member may contain a single type of fluorophore, or a given dual-indexed specific binding member may contain two or more different types of fluorophores. Examples of fluorophores, but not limited to, include: acridine, and derivatives such as acridine, acridine orange, acridine yellow, acridine red, and acridine isothiocyanate; 5-(2'-aminoethyl)aminonaphthalene-1-sulfonic acid (EDANS); 4-amino-N-[3-(vinylsulfonyl)phenyl]naphthalimide-3,5 disulfonate (Lucifer Yellow VS); N-(4-anilino-1-naphthyl)maleimide; anthranilamide; brilliant yellow; coumarin, and derivatives such as coumarin, 7-amino-4-methylcoumarin (AMC, coumarin 120), 7-amino-4-trifluoromethylcruarin (coumaran 151); cyanine, and derivatives such as cyanosine, Cy3, Cy5, Cy5.5, and Cy7; 4',6-diaminidino-2-phenylindole (DA PI); 5',5''-dibromopyrogallol-sulfonphthalein (bromopyrogallol red); 7-diethylamino-3-(4'-isothiocyanatophenyl)-4-methylcoumarin; diethylaminocoumarin; diethylenetriaminepentaacetate; 4,4'-diisothiocyanatodihydrostilbene-2,2'-disulfonic acid; 4,4'-diisothiocyanatostilbene-2,2'-disulfonic acid; 5-[dimethylamino]naphthalene-1-sulfonyl chloride (DNS, dansyl chloride); 4-(4'-dimethylaminophenylazo)benzoic acid (DABCYL); 4-dimethylaminophenylazophenyl-4'-isothiocyanate (DABITC); eosin, and derivatives such as eosin and eosin isothiocyanate; erythrosine, and derivatives such as erythrosine B and erythrosine isothiocyanate; ethidium;Fluorescein, and derivatives such as 5-carboxyfluorescein (FAM), 5-(4,6-dichlorotriazine-2-yl)aminofluorescein (DTAF), 2'7'-dimethoxy-4'5'-dichloro-6-carboxyfluorescein (JOE), fluorescein isothiocyanate (FITC), fluorescein chlorotriazinyl, naphthofluorescein, and QFITC (XRITC); fluorescein; IR144; IR1446; Lisamin (trademark); Lisamin Rhodamine, Lucifer Yellow; Malachite Green Isothiocyanate; 4-methylumbelliferone; Orthocresolphthalein; Nitrotyrosine; Pararoseaniline; Nile Red; Oregon Green; Phenol Red; β-phycoerythrin; o-phthalidaldehyde; Pyrene, and derivatives such as pyrene, pyrenebutyrate, and succinimimidyl 1-pyrenebutyrate; Reactive Red4 (Cibacron® Brilliant Red 3B-A); rhodamine, and derivatives such as 6-carboxy-X-rhodamine (ROX), 6-carboxyrhodamine (R6G), 4,7-dichlororhodamine lysamine, rhodamine B sulfonyl chloride, rhodamine (Rhod), rhodamine B, rhodamine 123, rhodamine X isothiocyanate, sulforhodamine B, sulforhodamine 101, sulfonyl chloride derivatives of sulforhodamine 101 (Texas Red), N,N,N',N'-tetramethyl-6-carboxyrhodamine (TAMRA), tetramethylrhodamine, and tetramethylrhodamine isothiocyanate (TRITC); riboflavin; rosolic acid and terbium chelate derivatives; xanthenes;Alexa-Fluor dyes (e.g., Alexa Fluor 350, Alexa Fluor 430, Alexa Fluor 488, Alexa Fluor 546, Alexa Fluor 555, Alexa Fluor 568, Alexa Fluor 594, Alexa Fluor 633, Alexa Fluor 647, Alexa Fluor 660, Alexa Fluor 680, Alexa Fluor 700, Alexa Fluor 750), Pacific Blue, Pacific Orange, Cascade Blue, Cascade Yellow; Quantum Dot dyes (Quantum Dot Corporation); Dylight dyes from Pierce (Rockford, IL), including Dylight 800, Dylight 680, Dylight 649, Dylight 633, Dylight 549, Dylight 488, Dylight 405; or combinations thereof. Other fluorophores or combinations thereof known to those skilled in the art, such as those available from Molecular Probes (Eugene, Oreg.) and Exciton (Dayton, Ohio), may also be used.
[0044] In some cases, the fluorophore is a polymer dye (e.g., a fluorescent polymer dye). A variety of fluorescent polymer dyes are found to be used in the methods of interest. In some cases of this method, the polymer dyes include conjugated polymers. Conjugated polymers (CPs) are characterized by a delocalized electronic structure containing alternating backbones of unsaturated bonds (e.g., double and / or triple bonds) and saturated bonds (e.g., single bonds), where π electrons can move from one bond to the other. Thus, the conjugated backbone can confer an elongated linear structure on the polymer dye, where the bond angles between the repeating units of the polymer are restricted. For example, proteins and nucleic acids are polymers, but in some cases they do not form elongated rod structures, but rather fold into higher-order three-dimensional shapes. Furthermore, CPs can form a “rigid rod” polymer backbone, experiencing limited torsional (e.g., twist) angles between monomer repeating units along the polymer backbone chain. In some cases, the polymer dyes include CPs with rigid rod structures. The structural features of the polymer dye can affect the fluorescence properties of the molecule.
[0045] Any suitable polymer dye can be used in the device and method of interest. In some cases, the polymer dye is a multichromophore having a structure that can collect light to amplify the fluorescence output of a fluorophore. In some cases, the polymer dye can collect light and efficiently convert it into emission of longer wavelengths. In some cases, the polymer dye has a light-gathering multichromophore system that can efficiently transfer energy to a nearby emission species (e.g., a "signaling chromophore"). Mechanisms for energy transfer include, for example, resonance energy transfer (e.g., Forster (or fluorescence) resonance energy transfer, Fret), quantum charge exchange (Dexter energy transfer), etc. In some cases, these energy transfer mechanisms are over relatively short distances, i.e., proximity of the light-gathering multichromophore system to the signaling chromophore provides efficient energy transfer. Under conditions for efficient energy transfer, amplification of emission from the signaling chromophore occurs when there are a large number of individual chromophores in the light-gathering multichromophore system. In other words, the emission from a signal transduction chromophore is stronger when the incident light ("excitation light") is at a wavelength absorbed by a light-gathering multichromophore system than when the signal transduction chromophore is directly excited by pump light.
[0046] Multiple chromophores can be conjugated polymers. Conjugated polymers (CPs) are characterized by a delocalized electronic structure and can be used as highly responsive optical reporters for chemical and biological targets. Because the effective conjugation length is significantly shorter than the length of the polymer chain, the backbone contains numerous closely spaced conjugated segments. Thus, conjugated polymers are efficient at light collection and enable optical amplification via Forster energy transfer.
[0047] The polymer dyes of interest include, but are not limited to, U.S. Patents No. 7,270,956, 7,629,448, 8,158,444, 8,227,187, 8,455,613, 8,575,303, 8,802,450, 8,969,509, 9,139,869, 9,371,559, 9,547,008, 10,094,838, 10,302,648, 10,458,989, 10,641,775, and 10,962,546 (these disclosures are incorporated herein by reference in their entirety), and Gaylord et al. Examples of dyes include those described in al., J.Am.Chem.Soc., 2001, 123(26), pp 6417-6418, Feng et al., Chem.Soc.Rev., 2010, 39, 2411-2419, and Traina et al., J.Am.Chem.Soc., 2011, 133(32), pp 12600-12607 (these disclosures are incorporated herein by reference in their entirety). Specific polymer dyes that may be used include, but are not limited to, BD Horizon Brilliant® Dyes, e.g., BD Horizon Brilliant® Violet Dyes (e.g., BV421, BV510, BV605, BV650, BV711, BV786); BD Horizon Brilliant® Ultraviolet Dyes (e.g., BUV395, BUV496, BUV737, BUV805); and BD Horizon Brilliant® Blue Dyes (e.g., BB515, BB550, BB790) (BD Biosciences, San Jose, CA). Any fluorescent dyes known to those skilled in the art (including, but not limited to, those listed above) or fluorescent dyes not yet discovered may be used in the method in question.
[0048] In some cases, each fluorophore constituting a given barcode can be excited by a common light source, such as a common laser. In such cases, each of the multiple fluorophores constituting a given barcode may have a common excitation wavelength range (for example, they are excited by wavelengths that differ from each other by only up to 50 nm, e.g., up to 25 nm, including up to 10 nm, e.g., up to 5 nm), but they differ from each other in terms of maximum emission.
[0049] As outlined above, any given two distinguishable fluorescent barcodes may be distinguishable from one another based on the type of fluorophores constituting the barcode and / or the signal luminance provided thereby. Thus, any two different barcodes may be distinguishable based on the fluorescent signal and / or intensity of the fluorescent signal obtained from the barcode. For example, two distinguishable fluorescent barcodes may be distinguishable from one another because they are composed of combinations of different types of fluorophores, e.g., one includes fluorophores a, b, and c, and the other includes fluorophores b, c, and d. Two distinguishable fluorescent barcodes may also be distinguishable from one another because they are composed of different amounts of fluorophores, e.g., one consists of fluorophores a, b, and c present in a first amount associated with a dual-indexed bead, and the other consists of fluorophores present in a second amount different from the first amount, e.g., a value that can be detected by, for example, a difference in signal luminance. Different luminances can easily be provided by having different amounts of fluorophores associated with a dual-indexed bead. Using combinations of fluorophore types and amounts, any desired number of unique fluorescent barcodes can be provided.
[0050] Oligonucleotide barcodes In addition to fluorescent barcodes, the dual-indexed specific binding members used in embodiments of the present invention include oligonucleotide barcodes. The length of the oligonucleotide barcode can vary, and in some cases may be in the range of 10 to 500 nt, for example, 15 to 100 nt. In some cases, the oligonucleotide barcode may consist of ribonucleic acid or deoxyribonucleic acid, as desired. The oligonucleotide barcode of embodiments of the present invention may include a barcode domain of the dual-indexed specific binding member, as well as other domains for which applications are found in embodiments of the present invention, such domains may include capture sequences, primer binding sites, and the like.
[0051] An oligonucleotide barcode may include one or more of the barcode domain, capture sequence, primer binding site, etc., of a dual-indexed specific binding member. The barcode domain of a dual-indexed specific binding member is a unique identifier, a domain or region that can be used to identify the dual-indexed specific binding member to which it is associated, for example, by its sequence. The unique identifier may be, for example, a nucleotide sequence having any suitable length, e.g., about 4 nucleotides to about 200 nucleotides. In some embodiments, the unique identifier is a nucleotide sequence of 25 nucleotides to about 45 nucleotides in length. In some embodiments, the unique identifier may have a length of 4 nucleotides, 5 nucleotides, 6 nucleotides, 7 nucleotides, 8 nucleotides, 9 nucleotides, 10 nucleotides, 15 nucleotides, 20 nucleotides, 25 nucleotides, 30 nucleotides, 35 nucleotides, 40 nucleotides, 45 nucleotides, 50 nucleotides, 55 nucleotides, 60 nucleotides, 70 nucleotides, 80 nucleotides, 90 nucleotides, 100 nucleotides, 200 nucleotides, about that value, less than that value, greater than that value, or in a range between any two of the above values.
[0052] Oligonucleotide barcodes may include a capture sequence, which is a domain or region that acts as a binding site to the target binding region of a bead-bound barcode nucleic acid, such as those described above. The capture sequence of choice may vary as desired and may be specific, random, or semi-random. In some cases, the capture sequence is a sequence that hybridizes to the target binding region of the bead-bound nucleic acid, as described in more detail below. In some cases, the capture sequence is a poly(A) sequence, which is configured to hybridize to the oligo(dT) target binding region, as described in more detail below. In such cases, the length of the poly(A) capture sequence may vary, in some cases ranging from 3 to 50 nt, for example, from 5 to 25 nt. If present, the capture sequence may be located at the 5' end of the oligonucleotide component.
[0053] Dual-indexed oligonucleotide barcodes of specific binding members may include a primer binding site. The primer binding site, if present, may be configured to bind to a primer used, for example, when preparing a sequenceable nucleic acid. For example, an oligonucleotide component may include a universal primer. A universal primer can point to a nucleotide sequence that is universal or common across all specific binding member / oligonucleotide subbarcodes used in a given workflow. In some cases, the primer binding site may be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30 nucleotide lengths, approximately these nucleotide lengths, or a number or range between any two of these nucleotide lengths. The length of the primer binding site can vary, and may be at least, or at most, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, or 30 nucleotides long. The length of the universal primer can also vary, and in some cases may range from 5 to 30 nucleotides long. The primer binding site may be located at the 5' end of the oligonucleotide barcode component.
[0054] As outlined in more detail below, when indexing cells in a cell sample, the cell sample is indexed with multiple distinct dual-indexed specific binding members using a combinatorial (e.g., split / pool) indexing protocol. Among the multiple distinct dual-indexed specific binding members used in a given indexing protocol, oligonucleotide barcodes may share common domains. For example, the oligonucleotide barcodes of distinct dual-indexed specific binding members may have a common capture domain, primer binding site, etc. In such cases, the capture domain, primer binding site, and other common domains may have the same sequence, so that the multiple distinct dual-indexed specific binding members have the same common sequence, e.g., the same primer binding site, the same capture domain, etc.
[0055] Cell indexing of cell samples using multiple dual-indexed specific binding members As summarized above, the methods of embodiments of the present invention provide a plurality of distinctly indexed cells, each having a different fluorescent signature associated therewith, the given fluorescent signature consisting of a fluorescent barcode(s) provided by a dual-indexed specific binding member associated with the cell. The number of different dual-indexed specific binding members associated with a given indexed cell can vary, but in some cases the number ranges from 1 to 10, for example 1 to 5, including, for example 1 to 4, for example 1 to 3, for example 1 to 2, including, for example 1, 2, or 3. Each different dual-indexed specific binding member has its own unique fluorescent barcode, and the set of distinct fluorescent barcodes of the dual-indexed specific binding members associated with the cell collectively provides the fluorescent signature associated with the indexed cell.
[0056] In carrying out embodiments of this method, a cell sample containing multiple cells is provided. The number of cells in a given cell sample can vary, but in some cases the number of cells ranges from 50 to 50,000,000, for example, from 100 to 1,000,000, and includes 500 to 100,000. The cells present in a given cell sample may be any type of cell, including prokaryotic and eukaryotic cells. Suitable prokaryotic cells include, but are not limited to, bacteria such as Escherichia coli, various Bacillus species, and extreme bacteria such as thermophilic bacteria. Suitable eukaryotic cells include, but are not limited to, fungi such as yeast and filamentous fungi (including species of Aspergillus, Trichoderma, and Neurospora); plant cells including those of corn, sorghum, tobacco, canola, soybean, cotton, tomato, potato, alfalfa, and sunflower; and animal cells including fish, birds, and mammals. Suitable fish cells include, but are not limited to, those from salmon, trout, tilapia, tuna, carp, flounder, halibut, swordfish, cod, and zebrafish. Suitable bird cells include, but are not limited to, those from chickens, ducks, quail, pheasants, and turkeys, as well as other jungle or game birds. Suitable mammalian cells include, but are not limited to, those from horses, cattle, buffalo, deer, sheep, rabbits, rodents such as mice, rats, hamsters, and guinea pigs, goats, pigs, primates, marine mammals such as dolphins and whales, as well as cell lines, such as any tissue or stem cell type human cell lines, as well as stem cells including pluripotent and non-pluripotent cells, and cells from non-human zygotes. Suitable cells also include cell types involved in a wide variety of disease symptoms, even in non-disease states.Therefore, appropriate eukaryotic cell types include, but are not limited to, all types of tumor cells (e.g., melanoma, myeloid leukemia, lung, breast, ovarian, colon, kidney, prostate, pancreatic, and testicular carcinomas), cardiomyocytes, dendritic cells, endothelial cells, epithelial cells, lymphocytes (T cells and B cells), mast cells, eosinophils, vascular endometrial cells, macrophages, natural killer cells, erythrocytes, hepatocytes, leukocytes including mononuclear leukocytes, hematopoietic cells, stem cells such as nerve, skin, lung, kidney, liver, and muscle cell stem cells (for use in screening for differentiation and dedifferentiation factors), osteoclasts, chondrocytes, and other connective tissue cells, keratinocytes, melanocytes, hepatocytes, renal cells, and adipocytes. In certain embodiments, cells are primary disease state cells such as primary tumor cells. Appropriate cells also include, but are not limited to, known research cells such as Jurkat T cells, NIH3T3 cells, CHO, COS, etc. Please refer to the ATCC cell line catalog, which is explicitly incorporated herein by reference.
[0057] In certain embodiments, the cells used in the present invention are collected from a subject. As used herein, “subject” refers to both humans and other animals, as well as other living organisms such as laboratory animals. Accordingly, the methods and compositions described herein are applicable to both human and veterinary uses. In certain embodiments, the subject is a mammal, and embodiments include those in which the subject is a human patient having (or suspected to have) a disease or pathological condition.
[0058] In certain embodiments, the cells to be analyzed are enriched before indexing, for example, as described in more detail below. For example, if the cells of interest are leukocytes derived from a human subject, whole blood from the subject can be subjected to density gradient centrifugation to enrich peripheral blood mononuclear cells (PBMCs, or leukocytes). Cells may be enriched using any convenient method known in the art, including fluorescence-activated cell sorting (FACS), magnetically activated cell sorting (MACS), density gradient centrifugation, etc. Parameters used to enrich specific cells from a mixed population include, but are not limited to, physical parameters (e.g., size, shape, density, etc.), in vitro growth characteristics (e.g., in response to specific nutrients in cell culture), and molecular expression (e.g., expression of cell surface proteins or carbohydrates, reporter molecules, such as green fluorescent protein).
[0059] In certain embodiments, the cells are live cells that retain viability throughout the assay. “Retaining viability” means that a certain percentage of cells remain alive at the end of the assay, ranging from approximately 20% viable to approximately 100% viable. In certain other embodiments, the method of the present invention is carried out in such a manner that the cells are rendered inviolable during the course of the assay; for example, the cells may be maintained in a buffer or under conditions where the cells are inviolable by fixation, permeabilization, or other means. Such parameters are generally determined by the nature of the assay being performed and the reagents used.
[0060] In some cases, cells may be treated with stimuli, for example. Stimuli that can treat cells can vary widely, ranging from culture conditions, exposure to temperature changes (e.g., heat or cold), exposure to electromagnetic radiation (e.g., light), exposure to active agents, and exposure to mechanical changes. If necessary, multiple different cell samples may be treated with the same or different stimuli. Thus, in some cases, the method involves differentially treating two or more of multiple cell samples, for example, by contacting two or more different samples with different active agents, or with the same active agent at different concentrations.
[0061] In carrying out embodiments of the present invention, the method includes indexing cells in a cell sample with a plurality of distinct dual-indexed specific binding members, for example, as described above, using a combinatorial protocol. A combinatorial protocol means a protocol in which cells are indexed by a random combination of different dual-indexed specific binding members selected from a population or set of dual-indexed specific binding members. In some cases, the combinatorial protocol is a splitting / pooling protocol. In some cases, the splitting / pooling protocol used in embodiments of the present invention is a protocol in which an initial cell sample is subjected to two or more iterations of distribution into a plurality of parts, with different parts of the plurality (e.g., each of the plurality of parts) in contact with different distinct dual-indexed specific binding members, and then the parts that have come into contact with different dual-indexed specific binding members are pooled. The number of splitting / contacting / pooling iterations used in a given splitting / pooling protocol can vary, and in some cases the number of iterations is in the range of 2 to 5, for example 2 to 4, for example 2 to 3.
[0062] In some cases, a given splitting / pooling protocol for producing indexed cells from an initial cell sample involves first splitting the cell sample into a first plurality of parts, i.e., aliquots or volumes. The number of different parts in the first plurality of parts can vary as desired, and in some cases, the number ranges from 2 to 25,000 parts, e.g., 10 to 10,000 parts, and includes 20 to 1,000 parts, e.g., 50 to 500 parts, e.g., 96 to 384 parts. Each part consists of multiple cells, i.e., contains multiple cells. The number of cells constituting a given subpart of the first plurality of parts can vary, but in some cases, the number ranges from 1 to 10,000, e.g., 10 to 1,000, and includes 100 to 1,000.
[0063] The portion used in the method of the present invention can take various different formats. The portion can take the form of any suitable reaction vessel, including, but not limited to, tubes, wells of a multiwell plate, etc. In some cases, the portion is a well of a multiwell device (e.g., a multiwell plate or multiwell tip or droplet, etc.). The reaction vessel that can function as the portion to which the reaction mixture and its components may be added and as the portion to which the reaction of the method in question may take place varies. Useful reaction vessels include, but are not limited to, tubes (e.g., single tubes, multi-tube strips, etc.) and wells (multi-well plates (e.g., 96-well plates, 384-well plates, or plates having any number of wells, e.g., 2000, 4000, 6000, or 10,000 or more)). Multi-well plates may be standalone or as part of a chip and / or device, as will be described in more detail below, for example. For example, a 96-well plate, a 384-well plate, or a plate having any number of wells, e.g., 2000, 4000, 6000, or 10,000 or more. Multi-well plates may be part of a chip and / or device. This disclosure is not limited by the number of wells in a multi-well plate. No. In various embodiments, the total number of wells on the plate is 100 to 200,000, or 5,000 to 10,000. In other embodiments, the plate includes smaller chips, each containing 5,000 to 20,000 wells. For example, a square chip may contain 125 × 125 nanowells with a diameter of 0.1 mm. Wells (e.g., nanowells) in a multiwell plate can be manufactured in any convenient size, shape, or volume. Wells may have a length of 100 μm to 1 mm, a width of 100 μm to 1 mm, and a depth of 100 μm to 5 mm, or more. In some cases, wells may have a depth of 5 mm or less, but are not limited to, for example, 4 mm or less, 3 mm or less, 2 mm or less, or 1 mm or less. In various embodiments, each nanowell has an aspect ratio (depth-to-width ratio) of 1 to 6 or more.In one embodiment, each nanowell has an aspect ratio of 1:6. The cross-sectional area may be circular, elliptical, oval, conical, rectangular, triangular, polyhedron, or any other arbitrary shape. The cross-sectional area at any given depth of the well may vary in size and shape. In a particular embodiment, the well has a volume of 0.1 nL to 1 mL. The nanowell may have a volume of 1 μL or less, for example, 500 nL or less. The volume may be 200 nL or less, for example, 100 nL or less. In one embodiment, the volume of the nanowell is 100 nL. If necessary, the nanowells can be manufactured to increase the surface area-to-volume ratio, thereby promoting heat transfer through the unit and reducing the ramp time of the thermal cycle. The cavity of each well (e.g., nanowell) can take on various configurations. For example, the cavity within the well may be divided by straight or curved walls to form separate but adjacent compartments, or by circular walls to form inner and outer annular compartments.
[0064] After dividing a cell sample into a first plurality of parts, the method of the embodiment of the present invention includes, for example, as described above, combining different parts of the first plurality of parts with a separate first dual-indexed specific binding member having separate fluorescent and oligonucleotide barcodes to stably associate the cells of the part with the first dual-indexed specific binding member. In this step, the plurality of parts are brought into contact with the dual-indexed specific binding member, where the dual-indexed specific binding member having a given part is separate from the dual-indexed specific binding member in contact with another part, and therefore each part in contact with the dual-indexed specific binding member comes into contact with its own intrinsic dual-indexed specific binding member. The contact between the parts and the dual-indexed specific binding member results in a stable association between the dual-indexed specific binding member and the cellular components of the part. In various embodiments, the parts and the dual-indexed specific binding member are combined in a liquid, for example, an aqueous composition. The combination can be achieved under any suitable conditions that provide a stable association between the dual-indexed specific binding member and the cells of the cellular composition. A dual-indexed specific binding member can come into contact with cells in a cell sample by introducing the dual-indexed specific binding member into a container of the portion (e.g., a reaction vessel such as a well), for example, by manual or automated fluid dispensing. The portion and the dual-indexed specific binding member are combined such that the dual-indexed specific binding member is stably associated with the cells in the cell sample, resulting in indexed cells. The cells and the dual-indexed specific binding member may be mixed and combined as needed and incubated for a certain period of time at a temperature suitable for providing a stable association between the dual-indexed specific binding member and the cells.In some cases, the combination of cells and dual-indexed specific binding members is incubated for a period of 15–120 minutes, for example, 30–90 minutes (e.g., 60 minutes), at a temperature of 20–25°C, for example, 20–22°C.
[0065] Next, the first multiple parts are combined with each other to produce a first pool. Following the production of parts in contact with the first dual-indexed specific binding member, the resulting parts are combined or pooled to produce a first pool of cells containing the first dual-indexed specific binding member. Cells from different parts can be combined or pooled using any convenient protocol. The number of cells in the resulting first pool can vary, in some cases ranging from 2 to 10,000,000 cells, e.g., 10,000 to 1,000,000 cells or 10,000 to 100,000 cells.
[0066] Next, the method of the embodiment of the present invention includes dividing the first pool into a second plurality of parts. Thus, after pooling, the resulting first cell pool is allocated into second subsets, each containing a plurality of cells, each containing a first dual-indexed specific binding member. In other words, the first pool of cell sources is divided or separated into a plurality of subparts, which collectively constitute a second subset, and the different subparts contain multiple or multiple cell sources, and the plurality of cell sources constituting the different subparts contain, for example, the first identifier-tagged nucleic acid as described above. The number of parts constituting the second plurality of parts can vary, but in some cases the number ranges from 2 to 25,000 parts, e.g., 10 to 10,000 parts, and includes 20 to 1,000 parts, e.g., 50 to 500 parts, e.g., 96 to 384 parts. In some cases the number of second plurality of parts is the same as the number of first plurality of parts. As described above, each of the second plurality of parts consists of a plurality of cell sources, i.e., contains a plurality of cell sources. The number of cell sources constituting a given subpart of the second set can vary, but in some cases the number ranges from 1 to 10,000, for example, from 100 to 1,000, including 100 to 500.
[0067] Within a given portion of a second plurality of portions, the multiple cell sources constituting that portion are distinct from one another with respect to a first dual-indexed specific binding member associated with the cells of that portion. For the pooling / redistribution step, cells from different portions of the first plurality of portions are combined into the same portion of the second plurality of portions. Within this same portion, the first dual-indexed specific binding members of the cells are distinct from one another, at least with respect to the barcode domain of the dual-indexed specific binding member. Thus, a given portion of the second plurality of portions has several distinct dual-indexed specific binding members associated with it, each distinct dual-indexed specific binding member associated with its own cell source.
[0068] In various embodiments, after the production of a plurality of second parts, the method includes combining different parts of the second plurality of parts with a separate second dual-indexed specific binding member having separate fluorescent and oligonucleotide barcodes, thereby stably associating the cells of the part with the second dual-indexed specific binding member. This step may be carried out with respect to the contact between the first plurality of parts and the first dual-indexed specific binding member as described above.
[0069] As described above, a given combinatorial (e.g., split / pool), indexing protocol may include two or more split / pool repeats. Therefore, in some cases, the method involves splitting a second pool into a third plurality of parts, and combining different parts of the third plurality of parts with a separate third dual-indexed specific binding member having distinct fluorescent and oligonucleotide barcodes, thereby stably associating the cells of that part with the third dual-indexed specific binding member.
[0070] The method described above results in the production of an indexed cell population. An indexed cell population consists of multiple cells, each of which is uniquely indexed and stably associated with its own unique combination of dual-indexed specific binding members. Thus, different uniquely indexed cells have associated combinations of dual-indexed specific binding members that are different from any other combination of dual-indexed specific binding members associated with any other cell in the multiple.
[0071] Phenotypic label In certain embodiments of the present invention, the method may include the detection of one or more phenotypic features of cells. Detectable phenotypic features include, but are not limited to, the presence of analytes, e.g., cell surface or internal markers, physical features (e.g., size, shape, particle size, etc.), cell number (or frequency), etc. Substantially any detectable feature of interest can be assayed as a detectable phenotypic feature of interest. In certain embodiments, the method of the present invention is directed, either qualitatively or quantitatively, to the detection of the presence of analytes associated with (e.g., intracellular, on, or attached to) the cells being assayed, e.g., markers. In some cases, the markers used for phenotypic labeling are not markers to which the cell association member of a dual-indexed specific binding member binds, for example, as described above.
[0072] In certain embodiments of these embodiments, the method involves contacting an indexed cell sample with a detectable analyte-specific binder. “Analyte-specific binder” and its grammatical equivalent mean any molecule, e.g., nucleic acids, small organic molecules and proteins, nucleic acid-binding dyes (e.g., ethidium bromide), that can be associated with a particular analyte (or a particular isoform of an analyte) in the cell more than any other. The analyte of interest includes any molecule associated with or present in the cells being analyzed by the method in question. Thus, the analyte of interest includes, but is not limited to, proteins, carbohydrates, organelles, nucleic acids, infectious particles (e.g., viruses, bacteria, parasites), metabolites, and the like. In certain embodiments, the analyte-specific binder is a protein. In certain embodiments of these embodiments, the analyte-specific binder is, for example, an antibody or its binding fragment. Therefore, using the methods and compositions of the present invention, any particular element isoform in a sample that is antigenically detectable and antigenically distinguishable from other isoforms of the activatable element present in the sample can be detected.
[0073] In certain embodiments, multiple detectable analyte-specific binders are used in the method according to the present invention. “Multiple analyte-specific binders” means that at least two or more analyte-specific binders are used, including three or more, four or more, five or more, and so on. In certain embodiments, each of the different analyte-specific binders is labeled (again, directly or indirectly) with a separately detectable label (e.g., a fluorophore having an emission wavelength detectable by separate channels on a flow cytometer, with or without correction). Multiple analyte-specific binders can bind to the same analyte intracellularly or on a cell (e.g., two antibodies binding to different epitopes on the same protein), to different analytes intracellularly or on a cell, or to any combination (e.g., two drugs binding to the same analyte and a third drug binding to a separate analyte). The upper limit of the number of analyte-specific binders depends largely on the assay parameters and the detection capability of the detection system used. If the analyte of interest is intracellular, the indexed cells can be permeabilized, for example, using protocols known in the art.
[0074] Protein expression analysis In some cases, a given workflow can evaluate protein expression and evaluate it in combination with gene expression. In such embodiments, indexed cells may be contacted with a phenotypic biomarker label (e.g., a phenotypic biomarker-specific binding member, e.g., an antibody or its binding fragment), and the phenotypic biomarker label may be conjugated to a phenotypic biomarker-identifying oligonucleotide (e.g., an oligonucleotide comprising a specific binding member that specifically binds and a barcode domain that identifies a phenotypic biomarker (e.g., an antigen)). For example, embodiments of the method may involve contacting indexed cells with one or more commercially available examples of phenotypic biomarker-labeled oligonucleotide conjugates, such as AbSeq antibody-oligonucleotide conjugates (Becton, Dickinson and Company). Further details regarding such phenotypic biomarker-labeled oligonucleotide conjugations and their uses are provided in the following published PCT applications: International Publication No. 2022 / 109343, International Publication No. 2021 / 163374, International Publication No. 2021 / 146207, International Publication No. 2020 / 159757, International Publication No. 2018 / 058073, and in "Shahi et al., "Abseq: Ultrahigh-throughput single cell protein profiling with droplet microfluidic barcoding. Sci Rep 7, 44447 (2017)," the disclosures of which are incorporated herein by reference.
[0075] Acquisition of flow cytometry data After the production of the indexed composition, the method may include assaying the indexed composition by flow cytometry, as described above. "Assimilation by flow cytometry" means performing a flow cytometry assay on the composition, e.g., the assay composition, as described above. The flow cytometry assay may include characterizing a sample, e.g., a sample containing the assay composition, using a flow cytometer system. The flow cytometry assay may include introducing the assay composition into a flow cytometer. A flow cytometer typically includes a sample reservoir for receiving a fluid sample, such as a sample containing the assay composition, and a sheath reservoir for containing sheath fluid. The flow cytometer transports particles in the fluid sample (e.g., cells from the assay composition) into the flow cell as a cell stream, while directing the sheath fluid towards the flow cell. Light is shone into the flow stream to characterize the components of the flow stream. Variations in the material in the flow stream, such as the form or presence of fluorescent labeling, may cause variations in the observed light, and these variations allow for characterization and separation. For example, particles such as molecules in a fluid suspension, analyte-bound beads, or individual cells pass through a detection region where the particles are typically exposed to excitation light from one or more lasers, and the light scattering and fluorescence properties of the particles are measured. The particles or their components are typically labeled with fluorescent dyes to facilitate detection. By labeling different particles or components using spectrally distinct fluorescent dyes, a large number of different particles or components can be detected simultaneously. In some implementations, the analyzer includes multiple detectors, one for each of the scattering parameters being measured and one or more for each of the distinct dyes being detected. For example, some embodiments include spectral configurations in which two or more sensors or detectors are used per dye. The resulting data include the signal and fluorescence emission measured for each of the light scattering detectors. In certain embodiments, a flow cytometry assay may detect a signal indicating the presence of a labeled secondary antibody in the sample.If a signal is detected, the sample may contain one antibody against the antigenic determinant of the coronavirus antigen.
[0076] As summarized above, a sample (e.g., in the flow stream of a flow cytometer) may be illuminated with light from a light source. In some embodiments, the light source is a broadband light source that emits light with a wide range of wavelengths, for example, extending above 50 nm, for example above 100 nm, for example above 150 nm, for example above 200 nm, for example above 250 nm, for example above 300 nm, for example above 350 nm, for example above 400 nm (including those extending above 500 nm). For example, one suitable broadband light source emits light with wavelengths from 200 nm to 1500 nm. Another example of a suitable broadband light source includes a light source that emits light with wavelengths from 400 nm to 1000 nm. If the method involves irradiating with a broadband light source, the broadband light source protocol of interest may include, but is not limited to, halogen lamps, deuterium arc lamps, xenon arc lamps, stabilized fiber-coupled broadband light sources, broadband LEDs with continuous spectra, superluminescent light-emitting diodes, semiconductor light-emitting diodes, wide-spectrum LED white light sources, multi-LED integrated white light sources, or any combination thereof, among other broadband light sources.
[0077] In other embodiments, the method includes irradiating with a narrowband light source that emits a specific wavelength or a narrow range of wavelengths, for example, a light source that emits light with wavelengths such as 40 nm or less, 30 nm or less, 25 nm or less, 20 nm or less, 15 nm or less, 10 nm or less, 5 nm or less, 2 nm or less, and a light source that emits light with wavelengths such as 2 nm or less, and a light source that emits light with specific wavelengths (i.e., monochromatic light). If the method includes irradiating with a narrowband light source, the narrowband light source protocol of interest may include, but is not limited to, a narrow-wavelength LED, laser diode or broadband light source coupled with one or more optical bandpass filters, diffraction gratings, monochromators or any combination thereof.
[0078] In certain embodiments, the method includes irradiating a sample with one or more lasers. As described above, the type and number of lasers vary depending on the sample and the desired light to be collected, and include gas lasers such as helium-neon lasers, argon lasers, krypton lasers, xenon lasers, nitrogen lasers, CO2 lasers, CO lasers, argon-fluorine (ArF) excimer lasers, krypton-fluorine (KrF) excimer lasers, xenon-chlorine (XeCl) excimer lasers, or xenon-fluorine (XeF) excimer lasers, or combinations thereof. In other cases, the method includes irradiating a flow stream with a dye laser such as a stilbene, coumarin, or rhodamine laser. In further cases, the method involves irradiating a flowstream with a metallic vapor laser, such as a helium-cadmium (HeCd) laser, a helium-mercury (HeHg) laser, a helium-selenium (HeSe) laser, a helium-silver (HeAg) laser, a strontium laser, a neon-copper (NeCu) laser, a copper laser or a gold laser, or a combination thereof. In yet another case, the method involves irradiating a flowstream with a solid-state laser, such as a ruby laser, a Nd:YAG laser, a NdCrYAG laser, an Er:YAG laser, a Nd:YLF laser, a Nd:YVO4 laser, a Nd:YCa4O(BO3)3 laser, a Nd:YCOB laser, a titanium-sapphire laser, a turium YAG laser, a ytterbium YAG laser, a ytterbium-2O3 laser or a cerium-doped laser, or a combination thereof.
[0079] The sample may be irradiated with one or more of the above-mentioned light sources, for example, two or more light sources, for example, three or more light sources, for example, four or more light sources, for example, five or more light sources, and for example, ten or more light sources. The light sources may include any combination of light sources of any type. For example, in some embodiments, the method involves irradiating the sample in a flow stream with an array of lasers, such as an array having one or more gas lasers, one or more dye lasers and one or more solid-state lasers. If desired, at least one laser is used to excite the fluorescent barcode, and other lasers are used for other fluorophores associated with the cells.
[0080] In certain cases, flowstream is described in Diebold, et al. Nature Photonics Vol. 7(10); 806-810 (2013), as well as in U.S. Patent Nos. 9,423,353, 9,784,661, 9,983,132, 10,006,852, 10,078,045, 10,036,699, 10,222,316, 10,288,546, 10,324,019, 10,408,758, 10,451,538, 10,620,111, and U.S. Patent Application Publications. Cells in a flow stream are imaged by fluorescence imaging using high-frequency tagged emission (FIRE) by irradiating them with multiple beams of frequency-shifted light to generate frequency-coded images as described in Nos. 2017 / 0133857, 2017 / 0328826, 2017 / 0350803, 2018 / 0275042, 2019 / 0376895, and 2019 / 0376894, and these disclosures are incorporated herein by reference. In such cases, flow cytometry data may include image data of cells in the composition being assayed (see, for example, Schraivogel et al., Science Vol. 375(6578); 315-320(2022)).
[0081] Embodiments of this method include collecting fluorescence with a fluorescence detector. In some cases, the fluorescence detector may be configured to detect fluorescence emission from fluorescent molecules associated with particles in a flow cell, such as labeled specific binding members (e.g., labeled antibodies that specifically bind to a marker of interest). In certain embodiments, the method includes detecting fluorescence from a sample with one or more fluorescence detectors (e.g., two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, fifteen or more, and including 25 or more fluorescence detectors). In various embodiments, each fluorescence detector is configured to generate a fluorescence data signal. Fluorescence from the sample may be detected independently by each fluorescence detector over one or more wavelength ranges from 200 nm to 1200 nm. In some cases, the method involves detecting fluorescence from a sample over a wavelength range, for example, 200 nm to 1200 nm, 300 nm to 1100 nm, 400 nm to 1000 nm, and 500 nm to 900 nm (including 600 nm to 800 nm). In other cases, the method involves detecting fluorescence at one or more specific wavelengths with each fluorescence detector. For example, fluorescence may be detected at one or more of the following wavelengths, depending on the number of different fluorescence detectors in the photodetector system in question: 450 nm, 518 nm, 519 nm, 561 nm, 578 nm, 605 nm, 607 nm, 625 nm, 650 nm, 660 nm, 667 nm, 670 nm, 668 nm, 695 nm, 710 nm, 723 nm, 780 nm, 785 nm, 647 nm, 617 nm and any combination thereof. In a particular embodiment, the method involves detecting the wavelength of light corresponding to the fluorescence peak wavelength of a specific fluorophore present in the sample. In various embodiments, fluorescence flow cytometer data is received from one or more fluorescence detectors (e.g., one or more detection channels), for example, two or more, for example, three or more, for example, four or more, for example, five or more, for example, six or more, and including eight or more fluorescence detectors (e.g., eight or more detection channels).
[0082] The light from the sample may be measured at one or more wavelengths, for example, five or more different wavelengths, for example, ten or more different wavelengths, for example, twenty-five or more different wavelengths, for example, fifty or more different wavelengths, for example, one hundred or more different wavelengths, for example, two hundred or more different wavelengths, for example, three hundred or more different wavelengths, and the measurement of light collected at four hundred or more different wavelengths is included.
[0083] In certain embodiments, the method involves spectrally decomposing the light from each fluorophore in a fluorophore-biomolecular reagent pair in a sample. In some embodiments, the overlap between different fluorophores is determined, and the contribution of each fluorophore to the overlapping fluorescence is calculated. In some embodiments, spectrally decomposing the light from each fluorophore involves calculating a spectral separation matrix of the fluorescence spectra for each of the multiple fluorophores having overlapping fluorescence in the sample detected by a photodetector system. In certain cases, spectrally decomposing the light from each fluorophore and calculating the spectral separation matrix for each fluorophore can be used to estimate the abundance of each fluorophore, for example, to determine the abundance of target cells in a sample.
[0084] In certain embodiments, the method includes spectrally decomposing light detected by multiple photodetectors, such as those described in U.S. Patent No. 11,009,400, U.S. Patent Application Publication No. 20210247293, and U.S. Patent Application Publication No. 20210325292, the entirety of which is incorporated herein by reference. For example, spectrally decomposing light detected by multiple photodetectors of a second set of photodetectors may include solving the spectral separation matrix using one or more of the following: 1) a weighted least squares algorithm; 2) a Sherman-Morrison iterative inverse updater; 3) LU matrix decomposition, e.g., where the matrix is decomposed into a product of lower triangular (L) and upper triangular (U) matrices; 4) a modified Cholesky decomposition; 5) by QR factorization; and 6) by singular value decomposition to compute the weighted least squares algorithm. In certain embodiments, the method further includes characterizing the spillover diffusion of light detected by multiple photodetectors, for example, as described below: U.S. Patent Application Publication No. 20210349004, the disclosure of which is incorporated herein by reference.
[0085] In certain cases, the abundance of fluorophores associated with a target particle (e.g., chemically (i.e., covalently, ionically) or physically) is calculated from spectrally decomposed light from each fluorophore associated with the particle. For example, in one example, the relative abundance of each fluorophore associated with the target particle is calculated from spectrally decomposed light from each fluorophore. In another example, the absolute abundance of each fluorophore associated with the target particle is calculated from spectrally decomposed light from each fluorophore. In certain embodiments, particles can be identified or classified based on the relative abundance of each fluorophore determined to be associated with the particle. In these embodiments, particles can be identified or classified by any convenient protocol, such as comparing the relative or absolute abundance of each fluorophore associated with the particle to a control sample having particles of known identity, or performing spectroscopic analysis or other assay analysis of a population of particles (e.g., cells) having the calculated relative or absolute abundance of the associated fluorophores.
[0086] In certain embodiments, the method includes sorting one or more particles of a sample (e.g., cells) identified based on the estimated abundance of fluorophores associated with the particles. The term “sorting” is used herein in its conventional sense and refers to separating components of a sample (e.g., droplets containing cells, droplets containing non-cellular particles such as biological macromolecules) and, in some cases, delivering the separated components to one or more sample collection containers. For example, the method may include sorting two or more components of a sample, e.g., three or more components, e.g., four or more components, e.g., five or more components, e.g., ten or more components, e.g., fifteen or more components, or sorting twenty-five or more components of a sample.
[0087] In sorting particles identified based on the abundance of fluorophores associated with the particles, the method includes data acquisition, analysis, and recording using a computer or the like, with multiple data channels recording data from each detector used to obtain overlapping spectra of multiple fluorophore-biomolecular reagent pairs associated with the particles. In these embodiments, the analysis includes spectral decomposition of light from multiple fluorophores of a fluorophore-biomolecular reagent pair having overlapping spectra associated with the particles (e.g., by calculating a spectral separation matrix), and identification of the particles based on the estimated abundance of each fluorophore associated with the particles. This analysis can be transmitted to a sorting system configured to generate a set of digitized parameters based on the particle classification. In some embodiments, methods for sorting components of a sample include sorting particles (e.g., cells in a biological sample), as described below: U.S. Patents No. 3,960,449, No. 4,347,935, No. 4,667,830, No. 5,245,318, No. 5,464,581, No. 5,483,469, No. 5,602,039, No. 5,643,796, No. 5,700,692, No. 6,372,506, and No. 6,809,804, the disclosures of which are incorporated herein by reference. In some embodiments, the method includes sorting components of a sample using a particle sorting module such as those described below: U.S. Patent Nos. 9,551,643 and 10,324,019, U.S. Patent Application Publication No. 2017 / 0299493 and International Publication No. 2017 / 040151, the disclosures of which are incorporated herein by reference. In certain embodiments, cells of a sample are sorted using a sorting decision module having multiple sorting decision units such as those described below: U.S. Patent No. 11,085,868, the disclosures of which are incorporated herein by reference.
[0088] Flow cytometry assay procedures are well known in the art. For example, the disclosures incorporated herein by reference are: Ormerod (ed.), Flow Cytometry: A Practical Approach, Oxford Univ. Press (1997); Jaroszeski et al. (eds.), Flow Cytometry Protocols, Methods in Molecular Biology No. 91, Humana Press (1997); Practical Flow Cytometry, 3rd ed., Wiley-Liss (1995); Virgo, et al. (2012) Ann Clin Biochem. Jan; 49 (pt 1): 17-28; Linden, et al., Semin Throm Hemost. 2004 Oct; 30 (5): 502-11; Alison, et al. J Pathol, 2010 Dec; 222 (4): 335-344; and Herbig, et al. (2007) Crit Rev Ther Drug Carrier See Syst.24(3):203-255. In certain embodiments, assaying a composition by flow cytometry involves using a flow cytometer capable of simultaneous excitation and detection of multiple fluorophores, such as a BD Biosciences FACSCanto® flow cytometer, used substantially according to the manufacturer's instructions. The methods of this disclosure may include image cytometry, such as those described in Holden et al. (2005) Nature Methods 2:773 and Valet et al. 2004 Cytometry 59:167-171, whose disclosures are incorporated herein by reference.
[0089] Appropriate flow cytometry systems are disclosed herein by reference in the following publications: Ormerod (ed.), Flow Cytometry: A Practical Approach, Oxford Univ. Press (1997); Jaroszeski et al. (eds.), Flow Cytometry Protocols, Methods in Molecular Biology No. 91, Humana Press (1997); Practical Flow Cytometry, 3rd ed., Wiley-Liss (1995); Virgo, et al. (2012) Ann Clin Biochem. Jan; 49 (pt 1): 17-28; Linden, et al., Semin Throm Hemost. 2004 Oct; 30 (5): 502-11; Alison, et al. J Pathol, 2010 Dec; 222 (4): 335-344; and Herbig, et al. (2007) Crit Rev Ther Drug Carrier. This may include, but is not limited to, the items listed in Syst.24(3):203-255.In certain cases, the target flow cytometry system is the BD Biosciences FACSCanto® flow cytometer, BD Biosciences FACSCanto® II flow cytometer, BD Accuri® flow cytometer, BD Accuri® C6 Plus flow cytometer, BD Biosciences FACSCelesta® flow cytometer, BD Biosciences FACSLyric® flow cytometer, BD Biosciences FACSVerse® flow cytometer, BD Biosciences FACSymphony® flow cytometer, BD Biosciences LSRFortessa® flow cytometer, BD Biosciences LSRFortessa® X-20 flow cytometer, BD Biosciences FACSPresto® flow cytometer, BD Biosciences FACSVia® flow cytometer, and BD Biosciences FACSCalibur® cell sorter, BD Biosciences FACSCount® cell sorter, BD Biosciences This includes FACSLyric® cell sorters, BD Biosciences Via® cell sorters, BD Biosciences Influx® cell sorters, BD Biosciences Jazz® cell sorters, BD Biosciences Aria® cell sorters, BD Biosciences FACSAria® II cell sorters, BD Biosciences FACSAria® III cell sorters, BD Biosciences FACSAria® Fusion cell sorters, and BD Biosciences FACSMelody® cell sorters, BD Biosciences FACSymphony® S6 cell sorters, etc.
[0090] In some embodiments, the system under consideration is a flow cytometry system such as those listed below: U.S. Patent Nos. 10,663,476, 10,620,111, 10,613,017, 10,605,713, 10,585,031, 10,578,542, 10,578, No. 469, No. 10,481,074, No. 10,302,545, No. 10,145,793, No. 10,113,967, No. 10,006,852 , No. 9,952,076, No. 9,933,341, No. 9,726,527, No. 9,453,789, No. 9,200,334, No. 9,097, No. 640, No. 9,095,494, No. 9,092,034, No. 8,975,595, No. 8,753,573, No. 8,233,146, No. 8 ,140,300, No.7,544,326, No.7,201,875, No.7,129,505, No.6,821,740, No.6,813,017 The disclosures of the same Nos. 6,809,804, 6,372,506, 5,700,692, 5,643,796, 5,627,040, 5,620,842, 5,602,039, 4,987,086, and 4,498,766 are incorporated herein by reference in their entirety.
[0091] In some embodiments, the system in question is a particle sorting system configured to sort particles using a sealed particle sorting module such as that described below: U.S. Patent Application Publication 2017 / 0299493, the disclosure of which is incorporated herein by reference. In certain embodiments, particles of a sample (e.g., cells) are sorted using a sorting decision module having multiple sorting decision units such as that described below: U.S. Patent Application Publication 2020 / 0256781, the disclosure of which is incorporated herein by reference. In some embodiments, the system in question includes a particle sorting module having deflection plates such as that described below: U.S. Patent Application Publication 2017 / 0299493 (filed March 28, 2017), the disclosure of which is incorporated herein by reference.
[0092] In certain cases, the flow cytometry system of the present invention is configured to image particles in a flow stream by fluorescence imaging using high-frequency tagged emission (FIRE), as described below: Diebold, et al. Nature Photonics Vol. 7(10); 806-810 (2013) and U.S. Patents 9,423,353, 9,784,661, 9,983,132, 10,006,852, 10,078,045, 10,036,699, 10,222,316, 10,288,546, 10,324,019, and 10,40 U.S. Patent Nos. 8,758, 10,451,538, 10,620,111, and U.S. Patent Application Publications 2017 / 0133857, 2017 / 0328826, 2017 / 0350803, 2018 / 0275042, 2019 / 0376895, and 2019 / 0376894, the disclosures of which are incorporated herein by reference. Figure 4 provides a schematic diagram of the acquisition of images of labeled cells according to embodiments of the present invention, including flow cytometry data via the FIRE protocol using a FACSDiscover flow cytometer, for example, as described in Schraivogel et al., Science Vol. 375(6578); 315-320(2022). As illustrated, image data can be obtained from fluorescent barcodes provided by fluorophores that have little effect on changes in other detectors, such as barcodes provided by Horizon® conjugated polymer dyes BB515, BB550, and BB790 (BD Biosciences).
[0093] As described above, this method includes cytometric analysis, which may include sorting. The target cells identified in the sample may be sorted and subsequently analyzed by any convenient analytical technique. Subsequent analytical techniques of interest may include, but are not limited to, sequencing; assays using CellSearch (described in Food and Drug Administration (2004) Final rule. Fed Regist 69:26036-26038); assays using CTC Chips (described in Nagrath, et al. (2007) Nature 450:1235-1239); assays using MagSweeper (described in Talasaz, et al. (2009). Proc Natl Acad Sci USA 106:3970-3975); and assays using nanostructured substrates (described in Wang S, et al. (2011) Angew Chem Int Ed Engl 50:3084-3088); their disclosures are incorporated herein by reference. If desired, the sorting protocol may include distinguishing between viable and dead cells, and any convenient staining protocol for identifying such cells may be incorporated into the method. Of particular interest in certain embodiments are the flow cytometry data obtained using the BD FACSDiscover® S8 cell sorter and BD CellView® Image Technology (BD Biosciences).
[0094] The analysis of data obtained from the indexed sample of the present invention involves analyzing the cells for a desired detectable feature(s) (e.g., as described in more detail above). The analysis of detectable features can be performed at any convenient step in the data analysis phase, including before, during, or after deconvolution. In fact, since the obtained data can be freely analyzed and re-analyzed, there are no intended restrictions on the order of deconvolution and analysis of the detectable feature(s). For the cells of interest, the obtained data may include, for example, the cell's fluorescence signature, as well as other cellular features such as cell-associated markers, cell images, etc., as provided by a dual-indexed specific binding member associated with the cell (as described above). The data may be provided in any convenient format, for example, the Flow Cytometry Standard (FCS) file format.
[0095] Acquisition of indexed single-cell sequence data For example, after obtaining cytometry data of an indexed cell population as described above, the method of the embodiment of the present invention may include obtaining sequence data of the indexed cells of the sample (the sequence data may be obtained for all cells of the labeled sample or for a subset of cells of the labeled sample, such as selected cells obtained via the cytometry step). The sequence data may be obtained using any convenient protocol. In some cases, the sequence data may be obtained using a protocol that includes partitioning the indexed cells, then generating a sequenceable library of nucleic acids obtained from the partitioned cells, and then reading the sequenceable library.
[0096] Indexed cell segmentation Following the production of indexed cells (i.e., cells stably associated with a dual-indexed specific binding member, as described, for example), embodiments of the Method include partitioning the indexed cells to produce partitioned single indexed cells, each having a dual-indexed specific binding member associated with it. In some cases, partitioning includes distributing the indexed cells into partitions or compartments such that each compartment contains a single indexed cell. "Partitioning" means that the indexed cells are placed in a small reaction chamber, which may be a fluidly isolated structure defined by a solid material such as microwells configured to contain the indexed cells. In some embodiments of the disclosed Methods, Devices, and Systems, a plurality of microwells randomly distributed across a substrate are used. In some embodiments, the plurality of microwells are distributed across the substrate in an ordered pattern, e.g., an ordered array. In some embodiments, the plurality of microwells are distributed across the substrate in a random pattern, e.g., a random array. Microwells can be manufactured in a variety of shapes and sizes. Suitable well geometric shapes include, but are not limited to, cylindrical, elliptical, cubic, conical, hemispherical, rectangular, or polyhedron three-dimensional geometric shapes consisting of several planes, such as cuboids, hexagonal prisms, octagonal prisms, inverted triangular pyramids, inverted square pyramids, inverted pentagonal pyramids, inverted hexagonal pyramids, or inverted truncated pyramids. In some embodiments, non-cylindrical microwells, such as wells with elliptical or square footprints, may offer advantages in that they can accommodate larger cells. In some embodiments, the upper and / or lower edges of the well walls may be rounded to avoid sharp corners, thereby reducing electrostatic forces that may arise due to the concentration of electrostatic fields at sharp edges or points. Thus, the use of rounded corners may improve the ability to collect beads from the microwells. The dimensions of the microwells can be characterized with respect to absolute dimensions.In some cases, the average diameter of the microwells may range from approximately 5 μm to approximately 100 μm. In other embodiments, the average microwell diameter is at least 5 μm, at least 10 μm, at least 15 μm, at least 20 μm, at least 25 μm, at least 30 μm, at least 35 μm, at least 40 μm, at least 45 μm, at least 50 μm, at least 60 μm, at least 70 μm, at least 80 μm, at least 90 μm, or at least 100 μm. In yet another embodiment, the average microwell diameter is up to 100 μm, up to 90 μm, up to 80 μm, up to 70 μm, up to 60 μm, up to 50 μm, up to 45 μm, up to 40 μm, up to 35 μm, up to 30 μm, up to 25 μm, up to 20 μm, up to 15 μm, up to 10 μm, or up to 5 μm. In some cases, the volume of the microwells used in the method of the present invention is approximately 200 μm. 3 ~about 800,000μm 3 It can vary within the range. In some embodiments, the microwell volume is at least 200 μm 3 , at least 500 μm 3 , at least 1,000 μm 3 at least 10,000 μm 3 at least 25,000 μm 3 at least 50,000 μm 3 at least 100,000 μm 3 at least 200,000 μm 3 at least 300,000 μm 3 at least 400,000 μm 3 at least 500,000 μm 3 at least 600,000 μm 3 at least 700,000 μm 3 Or, at least 800,000 μm 3 In other embodiments, the microwell volume is up to 800,000 μm 3 up to 700,000 μm 3 up to 600,000 μm 3 , 500,000 μm 3 up to 400,000 μm 3up to 300,000 μm 3 up to 200,000 μm 3 up to 100,000 μm 3 up to 50,000 μm 3 up to 25,000 μm 3 up to 10,000 μm 3 up to 1,000 μm 3 up to 500 μm 3 , or up to 200 μm 3 The number of microwells in a given device used in embodiments of the present invention can vary, in some cases the number may be 100 or more, e.g., 250 or more, e.g., 500 or more, and 1,000 or more, e.g., 5,000 or more, e.g., 10,000 or more, and in some cases the number may be 15,000 or less, e.g., 12,500 or less. Microwells suitable for use in embodiments of the present invention are further described in PCT application serial number PCT / US2016 / 014612, published as International Publication No. 2016 / 118915, the disclosure of which is incorporated herein by reference. Where used herein, substrate can refer to a type of solid support. A substrate may include, for example, a plurality of microwells. For example, a substrate may be a well array containing two or more microwells. In some embodiments, a microwell may include a small reaction chamber of a defined volume. In some embodiments, a microwell may capture one or more cells. In some embodiments, a microwell may capture only one cell. In some embodiments, a microwell can capture one or more solid supports. In some embodiments, a microwell can capture only one solid support. In some embodiments, a microwell captures a single cell and a single solid support (e.g., a bead). The number of wells, e.g., microwells, in a well plate, e.g., a microwell array, can vary in a given distribution step, but in some cases, it is in the range of 5 to 500, e.g., 5 to 100.
[0097] When partitioning indexed cells, the indexed cells can be placed in partitions, for example, in the microwells of a microwell array, using any convenient protocol. This disclosure provides a method for partitioning indexed cells into partitions for partitioning indexed cells. An aggregate of indexed cells can be introduced into a structure, for example, a microwell, for partitioning the indexed cells. The indexed cells can be contacted, for example, by a gravity flow that allows the indexed cells to settle within the partitioning structure. In some cases, an aqueous composition of indexed cells is brought into contact with an array of microwells, for example, by flowing it across the array of microwells so that the indexed cells accumulate in the microwells. The aqueous composition containing indexed cells can flow through a flow cell that is in fluid communication with the microwells. Suitable protocols and systems for partitioning captured particles into microwells, and microwells suitable for use in embodiments of the present invention, are further described in PCT application serial number PCT / US2016 / 014612, published as International Publication No. 2016 / 118915, the disclosure of which is incorporated herein by reference. To compartmentalize cells in a cell sample, any convenient protocol may be used, such as dispensing by pipetting, aliquoting into compartments of the cell sample, or flowing the sample onto the surface of a well plate.
[0098] In some embodiments, partitioning multiple indexed cells further includes providing particles (e.g., beads) containing particle (e.g., beads)-conjugated nucleic acids to partitions containing single cells, and the conjugated nucleic acids are used in preparing nucleic acid sequence preparation compositions, e.g., sequence preparation libraries, from labeled cells. In some cases, the particle (e.g., beads)-conjugated nucleic acids include a target binding region that binds to a complementary sequence of a target nucleic acid species in a cell, for example, and also to a capture sequence of a dual-indexed bead. For example, if the target nucleic acid species is cellular mRNA and the oligonucleotide barcode of a dual-indexed specific binding member includes a poly(A) capture sequence, the bead-conjugated nucleic acid may include a poly(T) domain as the target binding region. In addition to the target binding region, the conjugated nucleic acid may further include one or more additional domains, e.g., but not limited to, a cell labeling domain, a barcode domain, a molecular index domain (e.g., a unique molecular identifier (UMI) domain), a universal primer binding domain, and the like. Further details relating to particles having bound nucleic acids that may be provided within a compartment can be found in U.S. Patent Publication Nos. 2018 / 0088112, 2018 / 0200710, 2018 / 0346970, 2019 / 0056415, 2020 / 0248263, 2020 / 0299672, and 2021 / 0171940, the disclosures thereof incorporated herein by reference. Beads containing bound nucleic acids may be provided within the compartment using any convenient protocol, including, but not limited to, those described above for cell compartmentalization and further described in PCT application serial number PCT / US2016 / 014612, published as International Publication No. 2016 / 118915, the disclosure of which is incorporated herein by reference. The particles, e.g., beads, may be compartmentalized into cells before or after indexed cells, or in some cases in combination with indexed cells, as desired.
[0099] Generating a sequenceable library For example, the segmentation of indexed cells as described above results in segmented labeled cells spatially adjacent to particles (e.g., beads) having a cell-labeling domain nucleic acid containing, for example, a target-binding region. If the cell-labeling domain nucleic acid is adjacent to the target of an indexed single cell and / or the oligonucleotide barcode of a dual-indexed specific binding member, the target / oligonucleotide barcode can hybridize to the cell-labeling domain nucleic acid. Cell-labeling domains containing nucleic acids can be contacted in an inexhaustible ratio, if desired, so that each distinct target can be associated with a distinct cell-labeling domain containing a nucleic acid having its own unique UMI.
[0100] Following the compartmentalization of indexed cells, the indexed cells can be lysed to release target molecules as described above, and as a result, the released target molecules, such as nucleic acids, can bind to the target binding region of the cell-labeled domain nucleic acid to produce captured nucleic acids. Cell lysis can be achieved by any of the following means, for example, by chemical or biochemical means, by osmotic shock, or by thermal lysis, mechanical lysis or optical lysis. Particles can be dissolved by adding a cell lysis buffer containing a detergent (e.g., SDS, Lithium dodecyl sulfate, Triton X-100, Tween-20, or NP-40), an organic solvent (e.g., methanol or acetone), or a digestive enzyme (e.g., proteinase K, pepsin or trypsin), or any combination thereof. To increase the association between the target and the barcode, the diffusion rate of the target molecules can be altered, for example, by lowering the temperature of the lysis and / or increasing its viscosity. In some embodiments, the sample can be dissolved using filter paper. The top of the filter paper can be immersed in the lysis buffer. The filter paper can be applied to the sample at a pressure that facilitates the dissolution of the sample and the hybridization of the sample target to the substrate. In some embodiments, dissolution can be carried out by mechanical dissolution, thermal dissolution, optical dissolution, and / or chemical dissolution. Chemical dissolution may include the use of digestive enzymes such as proteinase K, pepsin, and trypsin. Dissolution can be carried out by adding a lysis buffer to the substrate. The lysis buffer may contain Tris HCl. The lysis buffer may contain at least about 0.01, 0.05, 0.1, 0.5, or 1 M or more of Tris HCl. The lysis buffer may contain up to about 0.01, 0.05, 0.1, 0.5, or 1 M or more of Tris HCl. The lysis buffer may contain about 0.1 M of Tris HCl. The pH of the lysis buffer may be at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or more. The pH of the lysis buffer can be up to approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 or higher.In some embodiments, the pH of the lysis buffer is about 7.5. The lysis buffer may contain a salt (e.g., LiCl). The concentration of the salt in the lysis buffer can be at least about 0.1, 0.5, or 1 M or more. The concentration of the salt in the lysis buffer can be at most about 0.1, 0.5, or 1 M or more. In some embodiments, the concentration of the salt in the lysis buffer is about 0.5 M. The lysis buffer may contain a detergent (e.g., SDS, Li dodecyl sulfate, Triton X, Tween, NP-40). The concentration of the detergent in the lysis buffer can be at least about 0.0001%, 0.0005%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, or 7% or more. The concentration of detergent in the dissolution buffer can be up to approximately 0.0001%, 0.0005%, 0.001%, 0.005%, 0.01%, 0.05%, 0.1%, 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, or 7% or more. In some embodiments, the concentration of detergent in the dissolution buffer is approximately 1% Li dodecyl sulfate. The time used in the dissolution method may depend on the amount of detergent used. In some embodiments, the more detergent used, the shorter the time required for dissolution. The dissolution buffer may contain a chelating agent (e.g., EDTA, EGTA). The concentration of the chelating agent in the dissolution buffer can be at least approximately 1, 5, 10, 15, 20, 25, or 30 mM or more. In some embodiments, the concentration of the chelating agent in the lysis buffer is about 10 mM. The lysis buffer may contain a reducing agent (e.g., beta-mercaptoethanol, DTT). The concentration of the reducing agent in the lysis buffer can be at least about 1, 5, 10, 15, or 20 mM. The concentration of the reducing agent in the lysis buffer can be up to about 1, 5, 10, 15, or 20 mM, or greater. In some embodiments, the concentration of the reducing agent in the lysis buffer is about 5 mM.In some embodiments, the lysis buffer may contain approximately 0.1 M TrisHCl, approximately pH 7.5, approximately 0.5 M LiCl, approximately 1% lithium dodecyl sulfate, approximately 10 mM EDTA, and approximately 5 mM DTT. Lysis can be carried out at a temperature of approximately 4, 10, 15, 20, 25, or 30°C. Lysis can be carried out for approximately 1, 5, 10, 15, or 20 minutes or more. Lysed cells may contain at least approximately 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, or 700,000 or more target nucleic acid molecules. Lysed cells may contain up to approximately 100,000, 200,000, 300,000, 400,000, 500,000, 600,000, or 700,000 or more target nucleic acid molecules.
[0101] Following the lysis of indexed cells and the release of nucleic acid molecules therefrom, the nucleic acid molecules can be randomly associated with a colocalized solid support, such as cell-labeled domain nucleic acids on beads. The association may include hybridization between the target recognition region of the cell-labeled domain nucleic acid and a complementary portion of the target nucleic acid molecule (e.g., the oligo(dT) of a barcode may interact with the poly(A) tail of the target). The assay conditions used for hybridization (e.g., buffer pH, ionic strength, temperature) can be selected to facilitate the formation of specific and stable hybrids. In some embodiments, nucleic acid molecules released from lysed cells can be associated with multiple probes on a substrate (e.g., hybridize with probes on the substrate). If the probes contain oligo(dT), the mRNA molecule can hybridize to the probe and be reverse transcribed. The oligo(dT) portion of an oligonucleotide can act as a primer for the first-strand synthesis of a cDNA molecule, for example, when subjected to DNA synthesis reaction conditions to produce a first-strand cDNA domain containing captured nucleic acid. Cell-labeled domain nucleic acids can also hybridize to a complementary capture sequence, e.g., a poly(A) sequence, of a dual-indexed specific binding member oligonucleotide barcode associated with the labeled cell. In this way, the cell-labeled domain nucleic acid can act as a primer for reverse transcription using the dual-indexed specific binding member oligonucleotide barcode as a template, for example, as described in more detail below.
[0102] If desired, a given workflow may include a pooling step in which a product composition, consisting of, for example, captured nucleic acid, synthesized first-strand cDNA, or synthesized double-strand cDNA, is combined with or pooled with one or more additional samples, e.g., product compositions obtained from labeled cells. In some cases, the pooling step is performed immediately after the hybridization step between the cell-labeled domain nucleic acid and the target nucleic acid, e.g., as outlined above. In such embodiments, the number of different samples, e.g., different product compositions produced from cells, that are combined with or pooled can vary, ranging in some cases from 2 to 1,000,000, e.g., 3 to 200,000, and 4 to 100,000, e.g., 5 to 50,000, and in some cases from 100 to 10,000, e.g., 1,000 to 5,000. Before or after pooling, the product composition(s) can be amplified, e.g., by polymerase chain reaction (PCR), as described in more detail below. Once the target-cell-labeled domain molecules are pooled, all further processing can be carried out within a single reaction vessel. Further processing may include, for example, reverse transcription, amplification, cleavage, dissociation, and / or nucleic acid elongation. These further processing reactions can be carried out within microwells, i.e., without initially pooling labeled target nucleic acid molecules from multiple cells.
[0103] This disclosure provides a method for producing target-cell-labeled domain couples using any convenient protocol, such as reverse transcription or nucleotide elongation. Target-cell-labeled domain couples may include complementary sequences of the cell-labeled domain and all or part of the target nucleic acid. Reverse transcription of the associated RNA molecule can occur by adding a reverse transcription primer along with reverse transcriptase. The reverse transcription primer may be an oligo(dT) primer, a random hexanucleotide primer, or a target-specific oligonucleotide primer. Oligo(dT) primers may be 12–18 nucleotides long and bind to the endogenous poly(A) tail at the 3' end of mammalian mRNA. Random hexanucleotide primers can bind to mRNA at various complementary sites. Target-specific oligonucleotide primers typically selectively prime the mRNA of interest. Reverse transcription can occur repeatedly to produce multiple cDNA molecules. The methods disclosed herein may include performing at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 reverse transcription reactions. The methods may also include performing at least about 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 reverse transcription reactions.
[0104] One or more nucleic acid amplification reactions can be performed to produce multiple copies of a target nucleic acid molecule. Amplification can be carried out in a multiplexing manner in which multiple target nucleic acid sequences are amplified simultaneously. Amplification reactions can be used to add sequencing adapters to nucleic acid molecules. Amplification reactions may include amplifying at least a portion of sample labels, if present. Amplification reactions may include amplifying at least a portion of cell labels and / or barcode sequences (e.g., molecular labels). Amplification reactions may include amplifying at least a portion of sample tags, cell labels, spatial labels, barcode sequences (e.g., molecular labels), target nucleic acids, or combinations thereof. The amplification reaction may include amplifying multiple nucleic acids to 0.5%, 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, 100%, or a range or number between any two of these values. The method may further include performing one or more cDNA synthesis reactions to produce one or more cDNA copies of a target barcode molecule containing sample labels, cell labels, spatial labels, and / or barcode sequences (e.g., molecular labels).
[0105] In some embodiments, amplification can be performed using polymerase chain reaction (PCR). As used herein, PCR can refer to a reaction for in vitro amplification of a specific DNA sequence by simultaneous primer extension of complementary strands of DNA. As used herein, PCR can encompass derivative forms of reactions including, but not limited to, RT-PCR, real-time PCR, nested PCR, quantitative PCR, multiplexed PCR, digital PCR, and assembly PCR.
[0106] Nucleic acid amplification may include non-PCR-based methods. Examples of non-PCR-based methods include, but are not limited to, multi-displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand-displacement amplification (SDA), real-time SDA, rolling circle amplification, or inter-circle amplification. Other non-PCR-based amplification methods include multi-cycle DNA-dependent RNA polymerase-driven RNA transcription amplification or RNA-directed DNA synthesis and transcription for amplifying DNA or RNA targets, ligase chain reaction (LCR) and Qβ replicase (Qβ) methods, use of palindromic probes, strand-displacement amplification, oligonucleotide-driven amplification using restriction endonucleases, amplification methods in which primers are hybridized to nucleic acid sequences and the resulting double helix is cleaved before extension and amplification, strand-displacement amplification using nucleic acid polymerases lacking 5' exonuclease activity, rolling circle amplification, and branched extension amplification (RAM). In some embodiments, amplification does not produce a cyclic transcript.
[0107] In some embodiments, the methods disclosed herein further include carrying out a polymerase chain reaction on nucleic acids (e.g., RNA, DNA, cDNA) to produce labeled amplicons (e.g., stochastically labeled amplicons). The labeled amplicons may be double-stranded molecules. The double-stranded molecules may include double-stranded RNA molecules, double-stranded DNA molecules, or RNA molecules hybridized to DNA molecules. One or both strands of the double-stranded molecule may include sample labels, spatial labels, cell labels, and / or barcode sequences (e.g., molecular labels). The labeled amplicons may be single-stranded molecules. The single-stranded molecules may include DNA, RNA, or a combination thereof. The nucleic acids of this disclosure may include synthetic nucleic acids or modified nucleic acids. Accordingly, the methods may include producing an amplicon composition from a first-strand cDNA domain containing a captured nucleic acid.
[0108] Amplification may involve the use of one or more non-natural nucleotides. Non-natural nucleotides may include photounstable or triggerable nucleotides. Examples of non-natural nucleotides include, but are not limited to, peptide nucleic acids (PNA), morpholino and locked nucleic acids (LNA), as well as glycol nucleic acids (GNA) and threose nucleic acids (TNA). Non-natural nucleotides may be added to one or more cycles of the amplification reaction. The addition of non-natural nucleotides can be used to identify the product as a specific cycle or point in time in the amplification reaction.
[0109] Performing one or more amplification reactions may include the use of one or more primers. One or more primers may contain, for example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 or more nucleotides. One or more primers may contain at least 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, or 15 or more nucleotides. One or more primers may contain 12 to less than 15 nucleotides. One or more primers may anneal to at least a portion of multiple labeled targets (e.g., stochastically labeled targets). One or more primers may anneal to the 3' or 5' ends of multiple labeled targets. One or more primers may anneal to the internal regions of multiple labeled targets. The internal region may consist of at least approximately 50, 100, 150, 200, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 650, 700, 750, 800, 850, 900, or 1000 nucleotides from the 3' end of multiple labeled targets. One or more primers may include primers from a fixed panel. One or more primers may include at least one custom primer. One or more primers may include at least one control primer. One or more primers may include at least one gene-specific primer.
[0110] One or more primers may include universal primers. Universal primers can anneal to universal primer binding sites. One or more custom primers can anneal to a first sample label, a second sample label, a spatial label, a cellular label, a barcode sequence (e.g., a molecular label), a target, or any combination thereof. One or more primers may include universal primers and custom primers. Custom primers can be designed to amplify one or more targets. Targets may include subsets of all nucleic acids in one or more samples. Targets may include subsets of all labeled targets in one or more samples. One or more primers may include at least 96 custom primers. One or more primers may include at least 960 custom primers. One or more primers may include at least 9600 custom primers. One or more custom primers can anneal to two or more different labeled nucleic acids. Two or more different labeled nucleic acids may correspond to one or more genes.
[0111] The method of this disclosure can be used with any amplification scheme. For example, in one scheme, a first round of PCR can amplify molecules attached to beads using gene-specific primers and primers for universal Illumina sequencing primer 1 sequence. A second round of PCR can amplify the first PCR product using nested gene-specific primers adjacent to Illumina sequencing primer 2 sequence and primers for universal Illumina sequencing primer 1 sequence. A third round of PCR adds P5 and P7 and a sample index to the PCR product to form an Illumina sequencing library. Sequencing using 150 bp × 2 sequencing can reveal cell labels and barcode sequences (e.g., molecular labels) on read 1, genes on read 2, and a sample index on index 1 read.
[0112] In some embodiments, nucleic acids can be removed from a substrate using chemical cleavage. For example, chemical groups or modified bases present in the nucleic acid can be used to facilitate its removal from a solid support. For example, enzymes can be used to remove nucleic acids from a substrate. For example, nucleic acids can be removed from a substrate by restriction endonuclease digestion. For example, nucleic acids can be removed from a substrate by treatment of nucleic acids containing dUTP or ddUTP with uracil-d-glycosylase (UDG). For example, nucleic acids can be removed from a substrate using enzymes that perform nucleotide removal, such as base removal repair enzymes, e.g., aprin / apyrimidine (AP) endonuclease. In some embodiments, photocleavable groups and light can be used to remove nucleic acids from a substrate. In some embodiments, cleavable linkers can be used to remove nucleic acids from a substrate. For example, a cleavable linker may include at least one of biotin / avidin, biotin / streptavidin, biotin / neutraavidin, Ig-protein A, a photounstable linker, an acid or base-unstable linker group, or an aptamer.
[0113] In some embodiments, amplification can be performed on a substrate, for example, using bridge amplification. The cDNA can be tailed with a homopolymer to generate suitable ends for bridge amplification using an oligo(dT) probe on the substrate. In bridge amplification, the primer complementary to the 3' end of the template nucleic acid can be each pair of first primers covalently attached to a solid particle. When a sample containing the template nucleic acid is brought into contact with the particle and a single thermal cycle is performed, the template molecule can be annealed to the first primer, which is extended forward by the addition of nucleotides to form a double-stranded molecule consisting of the template molecule and a newly formed DNA strand complementary to the template. In the heating step of the next cycle, the double-stranded molecule can be denatured, releasing the template molecule from the particle while leaving the complementary DNA strand attached to the particle via the first primer. In the annealing step of the subsequent annealing and extension steps, the complementary strand can be hybridized to a second primer that is complementary to a segment of the complementary strand at a position away from the first primer. This hybridization allows the complementary strand to be fixed to the first primer by covalent bond, and a bridge to be formed between the first and second primers, which are then fixed to the second primer by hybridization. In the extension step, the second primer can be extended in the reverse direction by adding nucleotides in the same reaction mixture, thereby converting the crosslink into a double-stranded crosslink. The next cycle then begins, and the double-stranded crosslink can be denatured to obtain two single-stranded nucleic acid molecules, each with one end attached to the particle surface via the first and second primers, and the other end of each unattached. In the annealing and extension steps of this second cycle, each strand can hybridize to further previously unused complementary primers on the same particle to form new single-stranded crosslinks. The two previously unused primers that are hybridized here are extended, converting the two new crosslinks into double-stranded crosslinks.The amplification reaction may involve amplifying at least 1%, 2%, 3%, 4%, 5%, 6%, 7%, 8%, 9%, 10%, 15%, 20%, 25%, 30%, 35%, 40%, 45%, 50%, 55%, 60%, 65%, 70%, 75%, 80%, 85%, 90%, 95%, 97%, or 100% of multiple nucleic acids.
[0114] Amplification of labeled nucleic acids can include PCR-based or non-PCR-based methods. Amplification of labeled nucleic acids can include exponential amplification of labeled nucleic acids. Amplification of labeled nucleic acids can include linear amplification of labeled nucleic acids. Amplification can be carried out by polymerase chain reaction (PCR). PCR can refer to a reaction for in vitro amplification of a specific DNA sequence by simultaneous primer extension of the complementary strand of DNA. PCR can encompass derivative forms of reactions including, but not limited to, RT-PCR, real-time PCR, nested PCR, quantitative PCR, multiplexed PCR, digital PCR, repression PCR, semi-repression PCR, and assembly PCR.
[0115] In some embodiments, amplification of labeled nucleic acids includes non-PCR-based methods. Examples of non-PCR-based methods include, but are not limited to, multi-displacement amplification (MDA), transcription-mediated amplification (TMA), nucleic acid sequence-based amplification (NASBA), strand-displacement amplification (SDA), real-time SDA, rolling circle amplification, or inter-circle amplification. Other non-PCR-based amplification methods include multi-cycle DNA-dependent RNA polymerase-driven RNA transcription amplification or RNA-directed DNA synthesis and transcription for amplifying DNA or RNA targets, ligase chain reaction (LCR), Qβ replicase (Qβ), use of palindromic probes, strand-displacement amplification, oligonucleotide-driven amplification using restriction endonucleases, amplification methods in which primers are hybridized to nucleic acid sequences and the resulting double helix is cleaved before extension and amplification, strand-displacement amplification using nucleic acid polymerases lacking 5' exonuclease activity, rolling circle amplification and / or branched extension amplification (RAM).
[0116] In some embodiments, the methods disclosed herein further include performing a nested polymerase chain reaction on an amplified amplicon (e.g., a target). The amplicon may be a double-stranded molecule. The double-stranded molecule may include a double-stranded RNA molecule, a double-stranded DNA molecule, or an RNA molecule hybridized to a DNA molecule. One or both strands of the double-stranded molecule may contain a sample tag or molecular identifier label. Alternatively, the amplicon may be a single-stranded molecule. The single-stranded molecule may include DNA, RNA, or a combination thereof. The nucleic acids of the present invention may include synthetic nucleic acids or modified nucleic acids.
[0117] In some embodiments, the method involves repeatedly amplifying a labeled nucleic acid to produce multiple amplicons. The methods disclosed herein may include performing at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, or 20 amplification reactions. Alternatively, the method may include performing at least about 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100 amplification reactions.
[0118] Amplification may further involve adding one or more control nucleic acids to one or more samples containing multiple nucleic acids. The control nucleic acids may include a control label.
[0119] Amplification may involve the use of one or more non-natural nucleotides. Non-natural nucleotides may include photounstable and / or triggerable nucleotides. Examples of non-natural nucleotides include, but are not limited to, peptide nucleic acids (PNA), morpholino and locked nucleic acids (LNA), as well as glycol nucleic acids (GNA) and threose nucleic acids (TNA). Non-natural nucleotides may be added to one or more cycles of the amplification reaction. The addition of non-natural nucleotides can be used to identify the product as a specific cycle or point in time in the amplification reaction.
[0120] Performing one or more amplification reactions may include the use of one or more primers. One or more primers may contain one or more oligonucleotides. One or more oligonucleotides may contain at least approximately 7 to 9 nucleotides. One or more oligonucleotides may contain 12 to less than 15 nucleotides. One or more primers may anneal to at least a portion of multiple labeled nucleic acids. One or more primers may anneal to the 3' and / or 5' ends of multiple labeled nucleic acids. One or more primers may anneal to the internal regions of multiple labeled nucleic acids. The internal region may consist of at least approximately 50, 100, 150, 200, 220, 230, 240, 250, 260, 270, 280, 290, 300, 310, 320, 330, 340, 350, 360, 370, 380, 390, 400, 410, 420, 430, 440, 450, 460, 470, 480, 490, 500, 510, 520, 530, 540, 550, 560, 570, 580, 590, 600, 650, 700, 750, 800, 850, 900, or 1000 nucleotides from the 3' end of multiple labeled nucleic acids. One or more primers may include primers from a fixed panel. One or more primers may include at least one custom primer. One or more primers may include at least one control primer. One or more primers may include at least one housekeeping gene primer. One or more primers may include a universal primer. A universal primer can anneal to a universal primer binding site. One or more custom primers can anneal to a first sample tag, a second sample tag, a molecular identifier label, a nucleic acid or its product. One or more primers may include both a universal primer and a custom primer. A custom primer may be designed to amplify one or more target nucleic acids. The target nucleic acids may include a subset of all nucleic acids in one or more samples. In some embodiments, the primers are probes attached to the array of this disclosure.
[0121] In some embodiments, barcoding (e.g., probabilistically barcoding) multiple targets in a sample further includes generating an indexed library of barcoded targets (e.g., probabilistically barcoded targets) or barcoded target fragments. The barcode sequences of different barcodes (e.g., molecular labels of different probabilistic barcodes) may be different from one another. Generating an indexed library of barcoded targets includes generating multiple indexed polynucleotides from multiple targets in a sample. For example, in the case of an indexed library of barcoded targets containing a first indexed target and a second indexed target, the label region of the first indexed polynucleotide may differ from the label region of the second indexed polynucleotide by 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50 nucleotides, at least or at most these nucleotides, or a number or range of nucleotides between any two of these values. In some embodiments, generating an indexed library of barcoded targets involves contacting multiple targets, such as mRNA molecules, with multiple oligonucleotides containing a poly(T) region and a labeling region, and performing a first strand synthesis using reverse transcriptase to produce single-stranded labeled cDNA molecules, each containing a cDNA region and a labeling region, wherein the multiple targets include at least two mRNA molecules of different sequences, and the multiple oligonucleotides include at least two oligonucleotides of different sequences. Generating an indexed library of barcoded targets may further include amplifying the single-stranded labeled cDNA molecules to produce double-stranded labeled cDNA molecules, and performing nested PCR on the double-stranded labeled cDNA molecules to construct labeled amplicons. In some embodiments, the method may include generating adapter-labeled amplicons.
[0122] Barcoding (e.g., probabilistic barcoding) may involve using nucleic acid barcodes or tags to label individual nucleic acid (e.g., DNA or RNA) molecules. In some embodiments, this may involve adding DNA barcodes or tags to cDNA molecules as they are generated from mRNA. Nested PCR can be performed to minimize PCR amplification bias. Adapters may be added for sequencing, for example, using next-generation sequencing (NGS). Sequencing results can be used for cell labeling, molecular labeling, and determining the sequence of nucleotide fragments of one or more copies of a target.
[0123] Sample indexing In some cases, embodiments of the method may include sample indexing, for example, combining a given sample with other samples for downstream processing, such as flow cytometry analysis or sequencing. In such cases, a given indexed cell population prepared from a first cell sample may be contacted with the sample indexing reagent so that the sample indexing reagent is stably associated with the indexed cells of the indexed population. The sample indexing reagent may vary, but in some cases, the sample indexing reagent may be conjugated to a specific binding member sample-identifying oligonucleotide, and these components may be as described above. In embodiments in which protein expression is assayed, for example, if a phenotypic biomarker label containing a conjugated phenotypic biomarker-identifying oligonucleotide is used (as described above), the phenotypic biomarker label containing a conjugated phenotypic biomarker-identifying oligonucleotide may also be conjugated to the sample-identifying oligonucleotide. For example, an AbSeq antibody containing both an antibody-identifying barcode and a sample indexing barcode may be used. One or more additional indexed cell populations from one or more additional samples may be contacted in the same manner as additional sample indexing reagents. After the preparation of entirely different indexed samples, the indexed samples may be combined or pooled for further processing, such as flow cytometry analysis or sequencing, as described in more detail elsewhere in this application. Sample indexing protocols are further described in the following published PCT application publication numbers: International Publication 2020 / 046833, International Publication 2020 / 037065, and International Publication 2018 / 226293; their disclosures are incorporated herein by reference. In some cases, commercially available sample multiplexing systems, such as the BD(trademark) Single-Cell Multiplexing Kit (Becton, Dickinson and Company), may be used.
[0124] Sequencing In certain embodiments, the provided method further includes subjecting a prepared expression library, e.g., an amplicon composition produced as described above, to a sequencing protocol, e.g., a next-generation sequencing (NGS) protocol. The protocol may be implemented on any suitable NGS sequencing platform. The NGS sequencing platform of interest includes, but is not limited to, sequencing platforms provided by Illumina® (e.g., HiSeq®, MiSeq®, and / or NextSeq® sequencing systems); Ion Torrent® (e.g., Ion PGM®, and / or Ion Proton® sequencing systems); Pacific Biosciences (e.g., PACBIO RS II sequencing system); Life Technologies® (e.g., SOLiD sequencing system); Oxford Nanopore (e.g., Minion); Roche (e.g., 454 GS FLX+, and / or GS Junior sequencing systems); or any other sequencing platform of interest. The NGS protocol will vary depending on the specific NGS sequencing system used. Detailed protocols for sequencing, which may include, for example, further amplification (e.g., solid-phase amplification), amplicon sequencing, and analysis of sequencing data, are available from the manufacturer of the NGS sequencing system being used.
[0125] In some cases, the method further includes using an oligonucleotide-labeled cell component binding reagent in applications where, for example, the detection, e.g., quantification, of one or more cell components, e.g., surface proteins, is desired (e.g., via the BD AbSeq protocol). The oligonucleotide-labeled cell component binding reagent used in such embodiments includes a cell component binding reagent, e.g., an antibody or its binding fragment, coupled to a cell component binding reagent-specific oligonucleotide containing an identifier sequence for the cell component binding reagent to which the cell component binding reagent-specific oligonucleotide is associated. In such cases, the magnetic capture beads may contain nucleic acids configured to capture (e.g., specifically bind) the domain of the cell component binding reagent-specific oligonucleotide. Thus, protein expression can be assayed in conjunction with gene expression, for example, when multi-ohmic analysis (e.g., combined transcriptome and proteome analysis) is desired. In such cases, the method may include preparing a captured sample with an oligonucleotide-labeled cell component binding reagent, and then providing capture of the cell component binding reagent-specific oligonucleotide released from the captured segmented cells. Further details regarding the use of oligonucleotide-labeled cell component binding reagents can be found in U.S. Patent Application Publication No. 2018 / 0267036 and U.S. Patent Application Publication No. 2020 / 0248263, the disclosures of which are incorporated herein by reference.
[0126] For example, further details regarding methods for obtaining sequence data from single cells, as described above, are provided in U.S. Patent Publication Nos. 2018 / 0088112, 2018 / 0200710, 2018 / 0346970, 2019 / 0056415, 2020 / 0248263, 2020 / 0299672, and 2021 / 0171940, the disclosures thereof incorporated herein by reference.
[0127] Linking cytometry data and sequence data The sequencing protocol generates sequence data from labeled cells. This sequence data can then be easily linked to cytometry data from the labeled cells, and the cytometry data and sequence data obtained from the same cells can be paired. In other words, a given set of cytometry (e.g., images) data and a given set of sequence data can be linked as being obtained from the same cells, as will be explained in more detail below.
[0128] Following the acquisition of cytometry and sequence data, the acquired cytometry and sequence data obtained from a given cell are linked, for example, as described above. Linking means that the cytometry and sequence data are paired as originating from the same cell. In this way, cytometry and sequence data obtained from the same labeled cell can be paired. In other words, a given set of cytometry data and a given set of sequence data can be identified as originating from the same cell, and then paired or otherwise associated with each other. In this manner, linked cytometry and sequence data can be obtained for a single cell in a cell sample.
[0129] Cytometry data and sequence data are linked by using fluorescence signatures and oligonucleotide barcodes provided by dual-indexed specific binding members associated with the indexed cells from which cytometry data and sequence data are obtained. In the resulting sequence data, sequence reads are obtained for both the oligonucleotide barcodes of the dual-indexed specific binding members of the cell target and the indexed cells, for example, as described above. In other words, for each indexed cell assayed in a given workflow, the sequence of the oligonucleotide barcode of the dual-indexed specific binding member associated with that cell and the sequence of the target nucleic acid from that cell, e.g., mRNA from the cell, are obtained. For each indexed cell, these obtained sequences are obtained using a protocol such as the one described above (which may be a next-generation sequencing protocol), and a library is generated from the original sequence, with each member of a given library generated from the same segment sharing a common cell marker. Thus, the sequence reads from the cell target nucleic acid and the oligonucleotide barcodes of the dual-indexed specific binding members obtained from the cells all share the same cell marker; that is, they all have a common cell marker. When linking cells and image data, all reads that have the same cell-labeling domain (i.e., share a common cell label) can be paired or linked from both the target nucleic acid reads and the oligonucleotide barcode reads of the dual-indexed specific binding member. This pairing or linking results in a set of reads containing both the target nucleic acid and the oligonucleotide barcode nucleic acid reads of the dual-indexed specific binding member, and these reads can be identified as originating from the same cell.
[0130] Next, the resulting sequence data, which includes reads of both the target nucleic acid and the oligonucleotide barcode nucleic acid of the dual-indexed specific binding member, can be matched (i.e., paired or linked) with the cytometry data. As outlined above, cytometry data of indexed cells includes the fluorescence signature of those cells, which is provided by one or more dual-indexed specific binding members associated with those cells. When cells are cytologically analyzed to obtain cytometry data for that purpose, a series or set of fluorescence signals obtained from the dual-indexed specific binding members associated with those cells is obtained, and this series can be called the cell-specific fluorescence signature. Different cells in a given workflow have their own unique cell-specific fluorescence signature. A given fluorescence signal provided by the dual-indexed specific binding members constituting such a cell-specific fluorescence signature can be assigned to a given portion of the sequence read, since the sequence of the oligonucleotide barcode of the dual-indexed specific binding member from which that fluorescence signal is obtained is known. Therefore, using each cell-specific fluorescence signature obtained for a given labeled cell, the sequences of oligonucleotide barcodes of different dual-indexed specific binding members associated with that indexed cell can be determined. Since the sequences of oligonucleotide barcodes of dual-indexed specific binding members are present in the oligonucleotide barcode reads, a given cell-specific fluorescence signature can be determined when associated with a given sequence dataset. When a cell-specific fluorescence signature is associated with a given sequence dataset, it can be determined that the sequence data was obtained from the same indexed cell that was in the same segment from which the cell-specific fluorescence signature was obtained. In other words, a fluorescence signature for a given cell can be obtained from a series of fluorescence signals obtained from that cell during cytometry analysis.Since a given fluorescent signature can be matched with reads from oligonucleotide barcodes from a dual-indexed specific binding member, the fluorescent signature can be matched with sequence reads from the dual-indexed specific binding member that produced the fluorescent signature, and the matched sequence reads from the dual-indexed specific binding member can then be used to identify a given compartment and all sequence data obtained from the cells within that compartment. Once the sequence data is assigned to a given compartment, the sequence data can be easily linked with cytometry data obtained for the cells within that compartment. In this way, linked cytometry data and sequence data can be obtained for a single cell in a cell sample.
[0131] kit Aspects of the present invention further include kits and compositions used to carry out various embodiments of the methods of the present invention. A kit of the present invention may include: a collection of dual-indexed specific binding members; beads containing bead-bound nucleic acids, for example, cell-labeling domains and target-binding regions as described above; and / or other reagents as desired. The collection of dual-indexed specific binding members may include various numbers of distinct fluorescent barcodes and oligonucleotide barcodes that are different from each other. The number of distinct dual-indexed specific binding members in a given collection may vary, but in some cases the number is in the range of 5 to 1,000, for example, 10 to 500.
[0132] The kit may further include one or more additional components that find use in carrying out embodiments of the method. For example, the kit may include components used to produce labeled cells, such as macrowell plates, liquid containers, such as tubes. In some cases, the kit may include a sample indexing reagent, such as an SMK reagent. Such a reagent may be included in any convenient format if desired. Where provided, the sample indexing (e.g., SMK) reagent may be included in a multi-container format, such as a multi-well format. For example, Figure 3 provides a diagram of a multi-well plate containing an SMK reagent, such as a different combination of reagents in each well. The reagent may be available in a storage-stable format, such as a dry format, such as a lyophilized format. Furthermore, the kit may include one or more components used to obtain sequence data, such as primers, polymerases (e.g., thermally stable polymerases, reverse transcriptases (both with hot-start properties)), dsDNAse, exonucleases, dNTPs, metal cofactors, one or more nuclease inhibitors (e.g., RNase inhibitors and / or DNase inhibitors), one or more molecular crowding agents (e.g., polyethylene glycol), one or more enzyme stabilizing components (e.g., DTT), stimulus-responsive polymers, or any other desired kit components(s), such as devices, solid supports, containers, cartridges, e.g., tubes, beads, plates, microfluidic chips, etc. The kit components may reside in separate containers, or a large number of components may reside in a single container.
[0133] In addition to the components described above, the kit may further include (in certain embodiments) instructions for carrying out the method. These instructions may be present in the kit in various forms, and one or more of these forms may be present in the kit. One possible form of these instructions is information printed on a suitable medium or substrate, e.g., one or more sheets of paper on which the information is printed, the kit's packaging, or accompanying documents. Yet another form of these instructions is a computer-readable medium on which the information is recorded, e.g., a diskette, a compact disc (CD), or a portable flash drive. Yet another possible form of these instructions is a website address that can be used via the Internet to access the information at a remote site.
[0134] The following are provided as examples, not as limitations. [Examples]
[0135] Figures 1-1 and 1-2 provide schematic diagrams of a workflow according to one embodiment of the present invention. As shown in Figures 1-1 and 1-2, dual-labeled AbSeq is used in combination with dual-labeled SMK (Single-Cell Multiplexing Kits, Becton Dickinson) tags. In this case, combinatorial SMK labeling does not need to be very extensive (not all cells receive a unique cell index). However, the workflow shows that one unique index from each identifiable cell cluster (population) is bulk sorted via FACS. Identifiable cell populations are selected based on phenotypic clustering using Ab-fluor FACS data. Since Ab-fluor also contains covalently bound oligo barcodes, the same population can be identified in downstream single-cell multi-omics data. Within each phenotypic cluster, 100 cell barcodes can be identified (via combinatorial SMK) and remapped to the corresponding cell data in the FACS data file.
[0136] Figures 2-1 and 2-2 provide schematic diagrams of a workflow according to one embodiment of the present invention. The workflow shown in Figures 2-1 and 2-2 uses standard FACS ab-fluors (Becton Dickinson), and the workflow does not depend on the correlation between FACS population clusters and AbSeq population clusters. In this case, 100 cells from each target FACS cluster are sorted into individual vials. These sorted cells are then labeled with a standard cell hashing mechanism such as BD's SMK. From there, a single-cell multi-omics workflow follows (with or without AbSeq). Standard SMK tagging correlates back to the individual FACS populations, and from there, a 100-cell index utilizing dual-labeled combinatorial tags provides cells through cell correlation from FACS to scM.
[0137] Figures 3-1 and 3-2 provide details of combinatorial dual-labeling reagents that may be used in embodiments of the present invention and how they may be provided in plate format.
[0138] Notwithstanding the attached claims, this disclosure is also defined by the following clauses:
[0139] 1. A method for preparing an indexed cell population from a cell sample, Dividing a cell sample into several parts, To stably associate the cells of a portion with the first dual-indexed specific binding member by combining different portions of a first plurality of portions with a separate first dual-indexed specific binding member having separate fluorescent and oligonucleotide barcodes, Combining a first set of parts to produce a first pool, Dividing the first pool into a second set of parts, and To stably associate a cell of a second part with the second dual-indexed specific binding member by combining a different part of a second part with a separate second dual-indexed specific binding member having separate fluorescent and oligonucleotide barcodes, Includes, A method for producing an indexed population of cells.
[0140] 2. The method according to Clause 1, wherein a separate fluorescent barcode of a dual-indexed specific binding member includes a unique combination of one or more fluorophores with one or more signal levels.
[0141] 3. The method according to Clause 2, wherein one or more fluorophores are fluorophores in the range of 1 to 4.
[0142] 4. The method described in Clause 2 or 3, wherein one or more signal levels include 1 to 5 signal levels.
[0143] 5. The method according to any one of the claims 2 to 4, wherein one or more fluorophores comprise a conjugated polymer dye.
[0144] 6. The method according to any of the claims 1 to 5, wherein the separate oligonucleotide barcodes of the dual-indexed specific binding members are in the range of 10 to 500 nt in length.
[0145] 7. A separate oligonucleotide barcode is located in the 5' to 3' direction. Primer binding site, Dual-index bead barcode domain, and Capture Domain The method described in any of clauses 1 to 6, including the method described in any of clauses 1 to 6.
[0146] 8. The method according to Clause 7, wherein separate oligonucleotide barcodes have a common primer binding site and a capture domain.
[0147] 9. The method according to any one of the clauses 7 to 8, wherein the capture domain contains a polyA sequence.
[0148] 10. The method according to any one of the claims 1 to 9, wherein a dual-indexed specific binding member specifically binds to a cell marker.
[0149] 11. The method according to clause 10, wherein the cell marker is a surface marker or an internal marker.
[0150] 12. The method according to clauses 10 and 11, wherein a specific binding member specifically binds to a universal cell marker.
[0151] 13. The method according to Clause 12, wherein the universal cell marker is a non-phenotypic marker.
[0152] 14. The method according to clause 13, wherein the universal cell marker is selected from the group consisting of CD44, CD45, CD47, and β-2 microglobulin.
[0153] 15. The method according to any one of the claims 1 to 14, wherein the specific binding member comprises an antibody or a binding fragment thereof.
[0154] 16. The method according to any of the provisions 1 to 15, further comprising combining a second set of parts to produce a second pool.
[0155] 17. Dividing the second pool into a third or more parts, and To stably associate a cell of a third part with the third dual-indexed specific binding member by combining a different part of the third part with a separate third dual-indexed specific binding member having separate fluorescent and oligonucleotide barcodes. The method described in Article 16, further including the method described in Article 16.
[0156] 18. The method according to any of the provisions 1 to 17, wherein the cell sample contains 50 to 50,000,000 cells.
[0157] 19. The method according to any one of the claims 1 to 18, further comprising labeling cells with a phenotypic biomarker.
[0158] 20. The method according to Clause 19, wherein the phenotypic biomarker labeling includes a fluorescently labeled specific binding member.
[0159] 21. The method according to clause 20, wherein the fluorescently labeled specific binding member comprises a fluorescently labeled antibody.
[0160] 22. The method according to any one of the claims 19 to 21, wherein the phenotypic biomarker labeling comprises a conjugated phenotypic biomarker-identifying oligonucleotide.
[0161] 23. The method according to Clause 22, wherein the phenotypic biomarker label further comprises a conjugated sample identification oligonucleotide.
[0162] 24. The method according to any one of the claims 1 to 23, comprising combining an indexed cell population with a second indexed cell population produced from a second sample.
[0163] 25. The method according to Clause 24, comprising labeling cells of each combined indexed population with a sample-indexed oligonucleotide.
[0164] 26. The method according to any one of the claims 1 to 25, further comprising assaying an indexed cell population by flow cytometry in order to obtain flow cytometry data for the indexed cell population.
[0165] 27. The method according to Clause 24, wherein the flow cytometry data includes image data.
[0166] 28. The method according to any one of the clauses 1 to 27, further comprising obtaining sequence data for an indexed cell population.
[0167] 29. The method according to clause 28, wherein sequence data is obtained using a next-generation sequencing protocol.
[0168] 30. The method according to Clause 29, wherein the next-generation sequencing protocol includes producing a sequence-prepared library from an indexed cell population.
[0169] 31. The method according to clause 30, wherein the sequence preparation library is produced using a barcoded beads / partitioning protocol.
[0170] 32. The method according to any of the clauses 28-31, further comprising linking sequence data with flow cytometry data for one or more of the assayed cells.
[0171] 33. A group of distinct dual-indexed specific binding members, each possessing distinct fluorescence and oligonucleotide barcodes.
[0172] 34. The population according to Clause 33, wherein the distinct fluorescent barcodes of dual-indexed specific binding members include a unique combination of one or more fluorophores at one or more signal levels.
[0173] 35. The group described in Clause 34, wherein one or more fluorophores are in the range of 1 to 4 fluorophores.
[0174] 36. A group as described in Clause 34 or 35, in which one or more signal levels include one to five signal levels.
[0175] 37. A group according to any one of the clauses 34 to 36, comprising one or more fluorophores and a conjugated polymer dye.
[0176] 38. A population according to any of the clauses 33 to 37, wherein the separate oligonucleotide barcodes of the dual-indexed specific binding members are in the range of 10 to 500 nt in length.
[0177] 39. Separate oligonucleotide barcodes, in the direction from 5' to 3', Primer binding site, Dual-index bead barcode domain, and Capture Domain A group including any of the groups described in any of clauses 33 to 38.
[0178] 40. The group according to Clause 39, wherein separate oligonucleotide barcodes have a common primer binding site and capture domain.
[0179] 41. A population as described in any of clauses 39 to 40, wherein the capture domain contains a polyA sequence.
[0180] 42. The population according to any one of the clauses 33 to 41, further comprising a cell association member configured to provide a stable association with cells, wherein the dual-indexed beads are further configured to provide a stable association with cells.
[0181] 43. A population according to any of clauses 33 to 42, wherein a dual-indexed specific binding member specifically binds to a cell marker.
[0182] 44. The population described in Clause 43, wherein the cell marker is a surface marker or an internal marker.
[0183] 45. The populations described in clauses 43 and 44, wherein the specific binding member specifically binds to a universal cell marker.
[0184] 46. A population as described in Clause 45, for which the universal cell marker is a non-phenotypic marker.
[0185] 47. The population described in Clause 46, wherein the universal cell marker is selected from the group consisting of CD44, CD45, CD47, and β-2 microglobulin.
[0186] 48. A population according to any of clauses 33 to 47, wherein the specific binding member includes an antibody or a binding fragment thereof.
[0187] 49. Dual-indexed specific binding members of any of the populations described in any of clauses 31-48.
[0188] 50. A kit comprising a population of dual-indexed specific binding members as described in any of clauses 31 to 48.
[0189] While the aforementioned inventions have been described in some detail as examples and illustrations to clarify their understanding, it will be readily apparent to those skilled in the art that, in light of the teachings of the present invention, several changes and modifications can be made without departing from the spirit or scope of the appended claims.
[0190] Therefore, the above merely illustrates the principles of the present invention. Those skilled in the art will understand that various configurations embodying the principles of the present invention and falling within its spirit and scope can be devised, although not expressly described or shown herein. Furthermore, all examples and conditional statements enumerated herein are intended primarily to help the reader understand the principles of the present invention and the concepts to which the inventors have contributed to advancing the art, and should be construed as not being limited to such specifically enumerated examples and conditions. Furthermore, all descriptions herein enumerating the principles, aspects, and embodiments of the present invention, as well as specific examples thereof, are intended to encompass both their structural and functional equivalents. Moreover, such equivalents are intended to include both currently known equivalents and those to be developed in the future, i.e., any developed elements that perform the same function regardless of their structure. Furthermore, nothing disclosed herein is intended to be made available to the public, whether such disclosure is expressly described in the claims or not.
[0191] Accordingly, the scope of the present invention is not intended to be limited to the exemplary embodiments shown and described herein. Rather, the scope and spirit of the present invention are embodied in the appended claims. In the claims, 35 U.SC § 112(f) or 35 U.SC § 112(6) are explicitly defined as being invoked for limitation in the claims only if the exact phrase “means” or the exact phrase “step” is stated at the beginning of such limitation in the claims. If such exact phrase is not used in limitation in the claims, 35 U.SC § 112(f) or 35 U.SC § 112(6) is not invoked.
[0192] Cross-reference of related applications This application claims priority to the filing date of U.S. Provisional Patent Application 63 / 448,935, filed on 28 February 2023. The disclosure of that application is incorporated herein by reference.
Claims
1. A method for preparing an indexed cell population from a cell sample, Dividing the cell sample into a first number of parts, To stably associate the cells of the portion with the first dual-indexed specific binding member by combining different portions of the first plurality of portions with a separate first dual-indexed specific binding member having separate fluorescent and oligonucleotide barcodes, Combining the first plurality of parts to produce the first pool, Dividing the first pool into a second plurality of parts, and To stably associate the cells of the portion with the second dual-indexed specific binding member by combining different portions of the second plurality of portions with a separate second dual-indexed specific binding member having separate fluorescent and oligonucleotide barcodes, Includes, A method for producing an indexed population of cells.
2. The method according to claim 1, wherein the separate fluorescent barcode of the dual-indexed specific binding member comprises a unique combination of one or more fluorophores of one or more signal levels.
3. A separate oligonucleotide barcode is located in the direction from 5' to 3'. Primer binding site, Dual-index bead barcode domain, and Capture Domain The method according to any one of claims 1 to 2, including the method described above.
4. The method according to any one of claims 1 to 3, wherein the dual-indexed specific binding member specifically binds to a cell marker.
5. The method according to claim 4, wherein the cell marker is a surface marker or an internal marker.
6. The method according to claims 4 to 5, wherein the specific binding member specifically binds to a universal cell marker.
7. The method according to claim 6, wherein the universal cell marker is a non-phenotypic marker.
8. The method according to claim 7, wherein the universal cell marker is selected from the group consisting of CD44, CD45, CD47, and β-2 microglobulin.
9. The method according to any one of claims 1 to 8, wherein the specific binding member comprises an antibody or a binding fragment thereof.
10. The method according to any one of claims 1 to 9, further comprising labeling the cells with a phenotypic biomarker.
11. The method according to any one of claims 1 to 10, further comprising assaying the indexed cell population by flow cytometry in order to obtain flow cytometry data for the indexed cell population.
12. The method according to claim 11, wherein the flow cytometry data includes image data.
13. The method according to any one of claims 1 to 12, further comprising obtaining sequence data for the indexed cell population.
14. The method according to any one of claims 11 to 13, further comprising linking sequence data with flow cytometry data for one or more of the assayed cells.
15. A group of distinct, dual-indexed, specific binding members, each possessing a distinct fluorescence and oligonucleotide barcode.
16. A dual-indexed specific binding member of the population according to claim 15.
17. A kit comprising a group of dual-indexed specific binding members as described in claim 15.