Genotype and phenotype coupling

By using source-specific nucleic acid barcoding to label target molecules, the problem of linking genotype and phenotype in existing technologies has been solved, enabling high-throughput, low-cost genotype-phenotype pairing and intermolecular affinity determination, thus improving analytical efficiency and accuracy.

CN107614700BActive Publication Date: 2026-03-31THE BROAD INST INC +1
View PDF 23 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2016-03-11
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently link specific encoded molecular phenotypes to corresponding genotypes in highly complex gene variant ensembles, and effective analytical methods are lacking.

Method used

Source-specific nucleic acid barcodes are used to label target molecules. By identifying these barcodes, the origin and characteristics of the target molecules can be determined. Then, multiplex analysis is performed using next-generation sequencing technology to achieve genotype-phenotype coupling.

Benefits of technology

It enables high-throughput, low-cost pairing of genotype and phenotypic characteristics of cells or samples, tracks changes in target molecules, and correlates them with source samples and test conditions, thereby improving the speed and accuracy of intermolecular affinity determination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN107614700B_ABST
    Figure CN107614700B_ABST
Patent Text Reader

Abstract

The present disclosure provides methods and compositions useful for labeling target molecules with source-specific nucleic acid identifiers (e.g., barcodes) that can be subsequently used to identify, quantify, or otherwise characterize features or activities of target molecules originating from a particular discrete volume. Such target molecules can include polypeptides expressed by cells, wherein the nucleic acid molecules encoding the polypeptides are labeled with the same or matching source-specific nucleic acid identifiers.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This disclosure of the field

[0002] This disclosure relates to methods and compositions for specifically labeling molecules with indexable nucleic acid identifiers that can couple a genotype of a cellular or non-cellular system with at least one phenotype of said cellular or non-cellular system. The methods and compositions allow for multiplex analysis of diverse coupled genotypes and phenotypes while maintaining information about the origin of the sample or subsample.

[0003] background

[0004] Modern genetic engineering methods allow for the rapid and inexpensive preparation of nucleic acid constructs and variants. High-complexity ensembles of gene design changes can be constructed in high-throughput manner, offering enormous potential for exploring a given design space. However, the analysis of such large and complex ensembles of gene variants requires the ability to easily distinguish individual variants from one another after, for example, assessing the characteristics of the encoded molecules. Methods to achieve this have long been lacking. Therefore, there is a need for methods to associate specific encoded molecular phenotypes with corresponding genotypes in highly complex variant ensembles.

[0005] summary

[0006] This disclosure provides methods and compositions for high-throughput labeling of target molecules using specific nucleic acid barcodes (e.g., source-specific nucleic acid barcodes). The identity, quantity, and / or activity of target molecules and / or target nucleic acids derived from a particular sample or a portion thereof can be determined by recognizing the source-specific nucleic acid barcode, optionally in combination with sequences recognizing additional barcodes. Using this information, the properties of the molecules can be determined. For example, the disclosed methods and compositions can be used to determine the affinity and / or specificity of molecules such as antibodies or antigens. Other aspects of the disclosed methods and compositions allow for the pairing of genotypic characteristics of cells or samples with phenotypic characteristics, wherein changes observed in target molecules and / or target nucleic acids can be tracked and / or correlated with source samples and / or different test conditions and / or test reagents encountered.

[0007] A method is disclosed for assigning a set of target molecules to specific compartments while maintaining information about the source of the target molecules (e.g., the source from which the target molecules originate or are separated into a compartment or discrete volume), which may be linked to information such as, but not limited to, the conditions experienced by the molecules in the compartment. The method includes: providing a sample comprising cellular or non-cellular systems; separating a portion of a single cellular or non-cellular system from the sample into individual compartments or discrete volumes, wherein each compartment or discrete volume further includes a source-specific barcode, wherein the source-specific barcode contains a unique nucleic acid recognition sequence. The unique nucleic acid recognition sequence maintains or carries information about the cellular or non-cellular system in the sample, such as the origin (e.g., compartmental or volumetric origin); the target molecules in individual compartments are labeled with source-specific barcodes present in individual compartments to form source-labeled target molecules, wherein the source-labeled target molecules from each individual compartment contain at least one of the same or matching uniquely indexed nucleic acid recognition sequences; the target molecules are optionally processed individually or in a multi-system; and the nucleotide sequence of the source-specific barcode is detected, thereby assigning the set of target molecules to specific individual compartments while maintaining information about the compartmental origin of the target molecules.

[0008] A method for assigning a set of target molecules to a sample or a set of samples, while maintaining information about the origin of the target molecules and target nucleic acids, wherein the sample comprises a cellular or non-cellular system, the method comprising: providing a sample, the sample comprising a cellular or non-cellular system; separating a portion of a single cellular or non-cellular system from the sample into individual compartments, wherein each individual compartment further comprises a source-specific barcode containing a unique nucleic acid recognition sequence, the unique nucleic acid recognition sequence maintaining or carrying information about the origin (such as compartment origin) of the cellular or non-cellular system in the sample; labeling the target molecules and target nucleic acids in the individual compartments with the source-specific barcodes present in the individual compartments to form source-labeled target molecules and source-labeled target nucleic acids, wherein the source-labeled target molecules from each individual compartment contain the same or matching uniquely indexed nucleic acid recognition sequences; optionally processing the target molecules individually or in a multi-system; and detecting the nucleotide sequence of the source-specific barcode, thereby assigning the set of target molecules to the target nucleic acids while maintaining information about the sample origin (such as compartment origin) of the target molecules.

[0009] A method for determining the specificity of a test agent for a target molecule is further disclosed, the method comprising: assigning a target molecule aggregate to compartments; contacting cells expressing the target molecule with an aggregate of test agents labeled with a test agent-specific barcode prior to compartmentation; isolating the target molecule bound to the test agent; and determining the sequence of the test agent-specific barcode and the sequence of the source-specific barcode, thereby identifying the test agent bound to the target molecule.

[0010] A further disclosed method for determining the affinity and / or specificity of a test agent for a target molecule is disclosed, the method comprising: distributing a set of target molecules into compartments; contacting a labeled target molecule with a test agent that binds to a detectable label; isolating the labeled target molecule bound to the test agent using the detectable label; determining the sequence of a source-specific barcode on the isolated target molecule; and quantifying the source-specific barcode associated with the isolated target molecule, thereby determining the affinity of the test agent for the target molecule.

[0011] A method for determining the expression of a target molecule on the surface of a cell assembly is further disclosed, the method comprising: assigning the target molecule assembly to compartments; contacting the sample cells to be separated with an assembly of test agents, each test agent being labeled with a unique test agent barcode; and determining the sequences of the source-specific barcode and the test agent-specific barcode on the test agents bound to the cells, thereby determining the expression of the molecule on the surface of the cell assembly.

[0012] A method for identifying proteins with specific activities of interest from a cell population is further disclosed, the method comprising: assigning a set of target molecules to compartments; isolating target molecules with specific activities of interest; and identifying source-specific barcodes of the isolated target molecules with specific activities of interest.

[0013] A further disclosed barcode labeling complex comprises a solid or semi-solid substrate and a plurality of barcode elements reversibly coupled thereto, wherein each of the barcode elements comprises an index nucleic acid recognition sequence; and one or more of a nucleic acid capture sequence specifically binding to a target nucleic acid and a specific binder specifically binding to a target molecule.

[0014] As will become apparent, this disclosure offers several advantages over the prior art. For example, determining intermolecular affinity (e.g., between antibody and antigen) using established methods requires expensive and laborious assays (e.g., ELISA). Furthermore, the methods and compositions of this invention can be used to assess intermolecular interactions between a first protein and a second protein, a protein and a nucleic acid, and / or a first nucleic acid and a second nucleic acid. The methods of this disclosure facilitate intermolecular affinity determination by utilizing the speed, low cost, and capacity of multiplex assays, as well as compatibility with next-generation sequencing. This provides significant benefits in applications such as designing high-value drugs or affinity-binding reagents, and many others.

[0015] The foregoing and other objectives and features of this disclosure will become more apparent from the following detailed description, which is made with reference to the accompanying drawings. Brief description of the attached diagram

[0017] Figure 1 This is a schematic diagram illustrating an exemplary method for labeling an amplicon set from a cell set.

[0018] Figure 2 To show the inclusion of the method for... Figure 1 This is a schematic diagram illustrating an exemplary composition of beads that use source-specific barcodes to label amplicon sets. In this example, each hydrogel bead carries a randomly constructed multipart index sequence (e.g., D...). i / C i / B i / A i In addition to achieving gene-specific capture / amplification and The clone population of sequences constructed from the library, G = "universal" capture sequence; CL1 = cleavable adapter 1 (RNA or other RNA intended for release from the bead); P7 sequence; D i = Randomly selected index "D" (i = e.g., 1 to 192); C i = Randomly selected index "C" (i = e.g., 1 to 192); B i = Randomly selected index "B" (i = e.g., 1 to 192); A i = Randomly selected index "A" (i = for example, 1 to 192); Sequencing primers; V = target sequence for gene / amplicon-specific capture; CL2 = cleavable adapter 2 (intended to release from the captured Ab or protein); X = capture protein or antibody (e.g., protein G for antibody).

[0019] Figure 3This is a schematic diagram illustrating an exemplary composition of beads comprising source-specific barcodes for labeling both the target nucleic acid and the target molecule. In this example, each hydrogel bead carries (i) a randomly constructed multipart index sequence (D i / C i / B i / A i In addition to achieving gene-specific capture / amplification and (ii) a clonal population of sequences constructed from the library, and (iii) a second population of antibody or protein capture sequences, each linked to a clonal population of, for example, the same multipart index sequence. These multipart index sequences constitute source-specific barcodes. G = “universal” capture sequence; CL1 = cleavable adapter 1 (RNA or other designed to be released from the bead); P7 sequence; D i = Randomly selected index "D" (i = e.g., 1 to 192); C i = Randomly selected index "C" (i = e.g., 1 to 192); B i = Randomly selected index "B" (i = e.g., 1 to 192); A i = Randomly selected index "A" (i = for example, 1 to 192); Sequencing primers; V = target sequence for gene / amplicon-specific capture; CL2 = cleavable adapter 2 (intended to release from the captured Ab or protein); X = capture protein or antibody (e.g., protein G for antibody).

[0020] Figure 4 This is a schematic diagram illustrating an exemplary method for simultaneously barcoding amplicon and expressed antibody or protein using an indexed hydrogel. In simple terms, this exemplary method involves the following steps: A single cell and a single indexed hydrogel bead are encapsulated, for example, in an emulsion microdroplet. The cell expresses, for example, an antibody or protein. After cell lysis, the expressed antibody and protein, as well as nucleic acids (e.g., mRNA or cDNA), are tagged with source-specific barcodes. Multiple barcoded samples can then be pooled (e.g., by disrupting the emulsion). Individual constructs can be retrieved by PCR and / or the labeled antibody or protein can be analyzed.

[0021] Figure 5This diagram illustrates an exemplary affinity analysis of labeled antibodies. After disrupting the emulsion, the labeled antibody can be exposed to a bulk antigen and, for example, captured on a column. A source-specific barcode can then be lysed from the captured antibody and sequenced. The source-specific barcodes of unbound labeled antibodies can be lysed and sequenced separately. Concentrations can be normalized using sequencing and abundance quantification of antibodies that remain bound compared to unbound antibodies. Furthermore, this sequence information can be combined with RNA-Seq information from nucleic acids (e.g., mRNA or cDNA) labeled with source-specific barcodes, thereby coupling genotypic information with phenotypic information.

[0022] Figure 6 This is a schematic diagram illustrating an exemplary alternative method for simultaneously barcoding amplicon and expressed antibody or protein using indexed hydrogels. Each hydrogel bead carries (i) a randomly constructed multipart index sequence (D... i / C i / B i / A i In addition to achieving gene-specific capture / amplification and (ii) A clonal population of sequences constructed from the library, and (iii) a second population of antibody or protein capture sequences, each linked to a clonal population of, for example, the same multipart index sequence. These multipart index sequences constitute source-specific barcodes. In short, the method includes the following steps: Cells expressing a specific binding agent (e.g., B cells) are encapsulated in emulsion droplets along with a single indexed hydrogel bead and a target cell expressing a cell surface antigen. The cells are allowed to express a binding moiety (e.g., an antibody or protein), which can be secreted and subsequently bind to the cell surface antigen on the target cell. The cells are then lysed, releasing the nucleic acid, the cell surface antigen, and the bound binding moiety. The binding moiety and nucleic acid are then labeled with a source-specific barcode and aggregated (e.g., by disrupting the emulsion). Individual constructs can be retrieved by PCR and / or the labeled antibody or protein can be analyzed. G = "universal" capture sequence; CL1 = cleavable adapter 1 (RNA or other intended for release from the bead); P7 sequence; D i = Randomly selected index "D" (i = e.g., 1 to 192); C i = Randomly selected index "C" (i = e.g., 1 to 192); B i = Randomly selected index "B" (i = e.g., 1 to 192); A i = Randomly selected index "A" (i = for example, 1 to 192); Sequencing primers; T = target sequence for gene / amplicon-specific capture; CL2 = cleavable adapter 2 (intended to release from the captured Ab or protein); X = capture protein or antibody (e.g., protein G for antibody).

[0023] Figure 7 To illustrate, for example, according to Figure 3 The diagram illustrates an exemplary affinity analysis of the labeled binding fraction (e.g., antibody) in the protocol shown. After emulsion disruption, a cell wall-specific affinity column can be used to capture the labeled, specifically binding cell surface antigen complex. The source-specific barcodes from each of the bound and unbound fractions can be cleaved and sequenced separately. Concentrations can be normalized using sequencing and abundance quantification of antibodies that remain bound relative to unbound antibodies. Furthermore, this sequence information can be combined with RNA-Seq information from nucleic acids (e.g., mRNA or cDNA) labeled with source-specific barcodes, thereby coupling genotypic information with phenotypic information.

[0024] Figure 8 This is a schematic diagram of a tagged antigen structure. The structure includes streptavidin bound to a single biotinylated antigen and a biotinylated nucleic acid, the nucleic acid including nucleic acid barcodes, such as those described herein.

[0025] Figure 9 This is a schematic diagram of the DNA barcode structure of hydrogel beads.

[0026] Figure 10 To show, for example Figure 9 A schematic diagram illustrating an example of "separate-aggregate" DNA barcode labeling on hydrogel beads.

[0027] Figure 11 This is a schematic diagram illustrating the encapsulation of reverse transcriptase and other reagents, as well as hydrogel beads, into emulsion droplets using a microfluidic device.

[0028] Figure 12 This is a schematic diagram illustrating reverse transcription in an emulsion droplet.

[0029] Figure 13 This is a schematic diagram illustrating the results obtained from reverse transcription from emulsion droplets.

[0030] Figure 14 This diagram illustrates the encapsulation of cells and hydrogel beads carrying partially double-stranded DNA molecules within emulsion droplets. The emulsion is then injected into another microfluidic device, where the droplets fuse with other droplets containing a dissolution buffer, RT and its buffer, DNA polymerase, and BclI restriction enzyme. Fusion is achieved using an electric field generated by electrodes, causing the two droplets to merge when the electric field is applied.

[0031] Figure 15 This is a schematic diagram illustrating the secretion of antibodies within the droplet. The droplet is then combined with the RT reagent.

[0032] Figure 16 This is a schematic diagram illustrating the results obtained from reverse transcription from emulsion droplets.

[0033] Figure 17 This is a schematic diagram illustrating batch purification and amplification.

[0034] Figure 18 This is a schematic diagram illustrating the microfluidic encapsulation of labeled cells together with lysis reagents, DNA polymerase, BclI restriction enzyme, and hydrogel beads containing a partial double-stranded DNA molecule.

[0035] Figure 19 This is a schematic diagram illustrating the results obtained from reverse transcription from emulsion droplets.

[0036] Figure 20 This is a schematic diagram illustrating batch purification and amplification.

[0037] Figure 21 This is a schematic diagram of the labeled antibody structure.

[0038] Figure 22 This diagram illustrates the microfluidic encapsulation of a single cell with a DNA-tagged target-specific antibody, dissolution buffer, DNA polymerase (New England BioLabs Klenow Fragment (3'—>5'exo-)) and its buffer, BclI restriction enzyme, biotin-tagged antibody, and hydrogel beads. The hydrogel beads carry aggregates of partially double-stranded DNA molecules, all possessing the same 96-base-pair DNA barcode; this DNA barcode is different on each hydrogel bead.

[0039] Figure 23 This diagram illustrates the cleavage of barcoded oligonucleotides by the BclI restriction enzyme and their release into the entire volume of a droplet. Simultaneously, cell lysis and the released target protein are captured by a labeled antibody, which anneals to a complementary sequence on the single-stranded portion of the barcoded DNA bound to the hydrogel. A polymerase then elongates the barcoded DNA molecule, thereby copying the DNA tag. These antibody-target complexes are captured by a biotin-labeled antibody. Adding DNA barcoding to the antibody-bound DNA tag confers single-cell specificity to these sequences.

[0040] Figure 24This is a schematic diagram illustrating the recovery of phosphorylated target proteins and the separation of phosphorylated and unphosphorylated target proteins.

[0041] Figure 25 This is a schematic diagram of the labeled antibody structure.

[0042] Figure 26 This is a schematic diagram illustrating the microfluidic encapsulation of a single cell with a lysis buffer, a BclI restriction enzyme, a biotin-labeled antibody, and hydrogel beads carrying a DNA barcode linked to the DNA-tagged antibody.

[0043] Figure 27 This is a schematic diagram illustrating protein capture.

[0044] Figure 28 This is a schematic diagram illustrating the recovery of antibody-protein complexes.

[0045] Figure 29 This diagram illustrates the microfluidic encapsulation of a single cell with a DNA-tagged target-specific antibody, a dissolution buffer, reverse transcriptase and its buffer, DNA polymerase, BclI restriction enzyme, a biotin-tagged antibody as described above, and hydrogel beads, the hydrogel beads carrying aggregates of partially double-stranded DNA molecules, all possessing the same 96-base-pair DNA barcode; this DNA barcode is different on each hydrogel bead.

[0046] Figure 30 This is a schematic diagram illustrating protein capture, RT, and DNA polymerization within a microdroplet.

[0047] Figure 31 This is a schematic diagram illustrating the recovery of antibody-protein complexes and cDNA.

[0048] Figure 32 This diagram illustrates the cleavage and release of barcoded oligonucleotides into the entire volume of a droplet by the BclI restriction enzyme. Simultaneously, the cell lyses, and the released mRNA is captured by annealing to the single-stranded portion of the released barcoded DNA, while reverse transcriptase elongates the barcoded DNA molecule, thereby copying the mRNA sequence. The DNA barcode, linked to the cDNA sequence, confers single-cell specificity to these cDNA sequences.

[0049] Figure 33 This is a schematic diagram illustrating RT activity and the results of subsequent purification and amplification steps.

[0050] Figure 34 The agarose gel containing amplified RT products in droplets is shown, along with quality control analysis performed on an Agilent bioanalyzer prior to sequencing.

[0051] Figure 35 The results of the analysis of the sequencing data are shown. The top shows the expected heavy and light chain sequence pairings, while the bottom shows the number and percentage of correct pairings obtained in the sequencing data.

[0052] Figure 36 This is a schematic diagram showing the sequence and structure of the antibody tag (Ab-tag) before conjugation.

[0053] Figure 37 This is a schematic diagram illustrating RT activity and the results of subsequent purification and amplification steps.

[0054] Figure 38 The agarose gel containing amplified RT products in droplets is shown, along with quality control analysis performed on an Agilent bioanalyzer prior to sequencing.

[0055] Figure 39 This is a graph illustrating the results of the sequencing data analysis. Hollow black outline bars indicate the expected proportion of Ab-tag 1 (those in the input). Gray bars indicate the proportion of Ab-tag 1 in the reads.

[0056] Detailed description

[0057] I. Terminology

[0058] Unless otherwise specified, technical terms are used according to their usual usage. Definitions of commonly used terms in molecular biology can be found in Benjamin Lewin, Genes IX, by Jones and Bartlet, 2008 (ISBN 0763752223); Kendrew et al. (eds.), The Encyclopedia of Molecular Biology, by Blackwell Science Ltd., 1994 (ISBN 0632021829); and Robert A. Meyers (ed.), Molecular Biology and Biotechnology: a Comprehensive Desk Reference, by VCH Publishers, Inc., 1995 (ISBN 9780471185710); and other similar references.

[0059] As used herein, unless the context clearly indicates otherwise, the singular forms “a (kind)” and “the” refer to both the singular and plural. For example, the term “source-specific barcode” includes a singular or plural source-specific barcode and can be considered equivalent to the phrase “at least one source-specific barcode”.

[0060] As used herein, the term "includes" means "includes". Therefore, "source-specific barcode" means "includes source-specific barcode" without excluding other elements.

[0061] While many similar or equivalent methods and materials may be used, particularly suitable methods and materials are described below. In case of conflict, this specification, including the interpretation of terminology, shall prevail. Furthermore, the materials, methods, and examples are illustrative only and are not intended to be limiting.

[0062] To facilitate a summary of the various embodiments of this disclosure, the following explanations of terms are provided:

[0063] Amplification: The purpose of amplification is to increase the copy number of nucleic acid molecules, such as those containing indexable nucleic acid identifiers (such as source-specific barcodes as described herein). The resulting amplification product is typically called an "amplifier." Amplification of nucleic acid molecules (such as DNA or RNA molecules) refers to the use of techniques to increase the copy number of nucleic acid molecules (including fragments). In some instances, amplicons are nucleic acids from cellular or non-cellular systems, such as amplified mRNA or DNA.

[0064] One example of amplification is polymerase chain reaction (PCR), in which a sample is contacted with a pair of oligonucleotide primers under conditions that allow the primers to hybridize to a nucleic acid template in the sample. The primers are extended under suitable conditions, dissociated from the template, re-annealed, extended, and dissociated again to amplify the copy number of the nucleic acid. This cycle can be repeated. The amplified products can be characterized by techniques such as electrophoresis, restriction endonuclease cleavage patterns, oligonucleotide hybridization or ligation, and / or nucleic acid sequencing.

[0065] Other examples of in vitro amplification techniques include quantitative real-time PCR; reverse transcriptase PCR (RT-PCR); real-time PCR (rtPCR); real-time reverse transcriptase PCR (rt RT-PCR); nested PCR; strand displacement amplification (see US Patent No. 5,744,311); transcription-free isothermal amplification (see US Patent No. 6,033,881); repair strand reaction amplification (see WO 90 / 01069); ligase chain reaction amplification (see European Patent Publication EP-A-320 308); gap-filling ligase chain reaction amplification (see US Patent No. 5,427,930); coupled ligase detection and PCR (see US Patent No. 6,027,889); and NASBA. TM RNA transcription-free amplification (see US Patent No. 6,025,134), etc.

[0066] Antibody: A polypeptide ligand, such as a protein, or fragment thereof, comprising at least a light chain and / or a heavy chain immunoglobulin variable region (or fragment thereof) that specifically recognizes and binds to an epitope of an antigen. Antibodies may include heavy and light chains, each having a variable region, referred to as a variable heavy chain (VH) region and a variable light chain (VL) region. The term also includes recombinant forms, such as chimeric antibodies (e.g., humanized mouse antibodies) and heteroconjugated antibodies (e.g., bispecific antibodies). Antibodies or fragments thereof may be multispecific, such as bispecific. Antibodies include all known forms of antibodies and other protein backbones with antibody-like properties. For example, antibodies may be monoclonal antibodies, polyclonal antibodies, human antibodies, humanized antibodies, bispecific antibodies, monovalent antibodies, chimeric antibodies, immunoconjugations, or protein backbones with antibody-like properties, such as fibronectin or ankyrin repeat sequences. Antibodies may have any of the following isotypes: IgG (e.g., IgG1, IgG2, IgG3, and IgG4), IgM, IgA (e.g., IgA1, IgA2, and IgAsec), IgD, or IgE.

[0067] In most mammals, including humans, intact antibodies have at least two heavy (H) chains and two light (L) chains linked by disulfide bonds. Each heavy chain includes a heavy chain variable region (VL). H ) and heavy chain constant region (C H However, this also includes single-stranded V, such as those found in camels. HH Variants and their fragments. The heavy-chain constant region comprises three structural domains, namely C H 1. C H 2 and C H 3; and C H 1 and C H The hinge area between 2. Each light chain includes a light chain variable area (V). L The light chain constant region includes the structural domain C. L V H and V L The region can be further divided into highly variable regions called complementarity-determining regions (CDRs), interspersed with more conservative regions called framework regions (FRs). Each V H and V L It consists of three CDRs and four FRs, arranged in the following order from the amino terminus to the carboxyl terminus: FR1, CDR1, FR2, CDR2, FR3, CDR3, FR4. The variable regions of the heavy and light chains contain binding domains that interact with the antigen.

[0068] This includes complete immunoglobulins and their variants and portions well known in the art, such as Fab fragments, Fab' fragments, F(ab)'2 fragments, single-chain Fv proteins (“scFv”), and disulfide-bonded stable Fv proteins (“dsFv”) Fd, Feb, or SMIP. Antibody fragments can be, for example, bifunctional antibodies, trifunctional antibodies, affibody, nanobody, aptamer, domain antibody, linear antibody, single-chain antibody, or multispecific antibody formed from antibody fragments. Examples of antibody fragments include (i) Fab fragments: composed of V... L V H C L And C H (ii) F(ab')2 segment: a divalent segment consisting of two Fab segments bonded by disulfide bridges in the hinge region; (iii) Fd segment: a segment consisting of V H and C H (iv) Fv fragment: a fragment consisting of a single arm of the antibody. L and V H Fragments composed of structural domains; (v)dAb fragments: including V H and V L Fragments of structural domains; (vi)dAb fragment: composed of V H Domain or V HH Fragments composed of structural domains (such as nanobodies) TM (vii)dAb fragment: from V H or V L (viii) A fragment composed of structural domains; (ix) Separate complementarity-determining regions (CDRs); and (ix) A combination of two or more separate CDRs optionally connected by a synthesis joint. Furthermore, although the two structural domains of the Fv fragment (V... L and V H V is encoded by individual genes, but they can be linked together using recombination methods, for example, by synthetic linkers that enable them to be made into single protein chains, in which V L and V H Regions pair to form monovalent molecules (referred to as single-chain Fvs (scFvs)). Antibody fragments can be obtained using conventional techniques known to those skilled in the art, and in some cases, can be used in the same manner as intact antibodies. Antigen-binding fragments can be generated by recombinant DNA techniques or by enzymatic or chemical cleavage of intact immunoglobulins. Antibody fragments may further comprise any of the antibody fragments described above, plus additional C-terminal amino acids, N-terminal amino acids, or amino acids that separate the individual fragments.

[0069] An antibody is called chimeric if it comprises one or more variable or constant regions derived from a first species and one or more variable or constant regions derived from a second species. Chimeric antibodies can be constructed, for example, through genetic engineering. Chimeric antibodies may include immunoglobulin gene segments belonging to different species (e.g., from mice and humans).

[0070] Human antibodies are specific binding agents having variable regions, both the framework region and the CDR region, derived from the human immunoglobulin sequence. Furthermore, if the antibody contains a constant region, that constant region is also derived from the human immunoglobulin sequence. Human antibodies may include amino acid residues not recognized in the human immunoglobulin sequence, such as one or more sequence variations, such as mutations. Variations or additional amino acids may be introduced, for example, through human manipulation. The human antibodies disclosed herein are not chimeric.

[0071] Antibodies can be humanized, meaning that antibodies include one or more complementarity-determining regions (e.g., at least one CDR) substantially derived from non-human immunoglobulins, or antibodies are manipulated to include at least one immunoglobulin domain having a variable region including a variable framework region substantially derived from human immunoglobulins or antibodies.

[0072] Antigen or immunogen: A compound, composition, or substance that can stimulate the production of antibody or T cell responses in an animal, including compositions injected or absorbed into the animal. An antigen reacts with a product having specific humoral or cellular immunogenicity, including those induced by a heterologous antigen (such as the disclosed antigen). An "epitaph" or "antigenic determinant" refers to a region of an antigen to which B cells and / or T cells respond. In one embodiment, T cells respond to an epitope when it binds to an MHC molecule. Epitopes can be formed from consecutive amino acids or discontinuous amino acids juxtaposed by the ternary folding of a protein. Epitopes formed from consecutive amino acids are typically retained upon exposure to denaturing solvents, while epitopes formed by ternary folding are typically lost upon treatment with denaturing solvents. Epitopes typically comprise at least 3 amino acids, and more generally, at least 5, about 9, or about 8-10 amino acids in a unique spatial conformation. Methods for determining the spatial conformation of an epitope include, for example, X-ray lenticography and nuclear magnetic resonance.

[0073] Examples of antigens include, but are not limited to, peptides, lipids, polysaccharides, and nucleic acids containing antigenic determinants (such as those recognized by immune cells). In some instances, antigens include peptides derived from pathogens of interest. Exemplary pathogens include bacteria, fungi, viruses, and parasites. In specific instances, antigens are derived from HIV, such as peptides like gp120, gp140, and gp160, or antigenic fragments thereof, such as the gp120 outer domain.

[0074] A "target epitope" is a specific epitope on an antigen that specifically binds to an antibody of interest (such as a monoclonal antibody). In some instances, a target epitope includes amino acid residues that contact the antibody of interest, making the target epitope selectable by identifying the amino acid residues that contact the antibody of interest.

[0075] Biotin-16-UTP: A bioactive analog of uridine-5'-triphosphate readily incorporated into RNA by RNA polymerases such as T7, T3, or SP6 RNA polymerases during in vitro transcription reactions. In some instances, biotin-16-UTP is incorporated into a source-specific barcode (or any other barcode) during reverse transcription from a probe DNA template, such as during in vitro transcription using RNA polymerases such as T7, T3, or SP6 RNA polymerases.

[0076] Capture portion: When attached to another molecule, such as a nucleic acid barcode disclosed herein, it allows the capture of a target probe molecule or other substance via interaction with something bound to the capture portion, such as a specific surface and / or molecule, such as a specific binding molecule capable of specifically binding to the capture portion. In a specific instance, the capture portion is biotin, and the specific binding agent for the capture portion is avidin or streptavidin.

[0077] Contact: Placing in a direct physical association manner, including in solid or liquid form, such as bringing the sample into contact with a nucleic acid barcode.

[0078] Sufficient conditions for detection: allowing the activity required for detection, such as allowing the detection and / or quantification of nucleic acids, such as nucleic acid barcodes, transcription products and / or their amplified products, in any environment.

[0079] Control: Reference standard. A control can be a known value or range of values ​​indicating a baseline level or quantity, or present in a tissue or cell or its population (such as normal non-cancerous cells). A control can also be a cell or tissue control, such as tissue from a disease-free state and / or exposed to different environmental conditions. The difference between the test sample and the control sample can be an increase or conversely a decrease. This difference can be qualitative or quantitative, such as a statistically significant difference.

[0080] Covalent bonding refers to the covalent bonds formed between atoms, characterized by the sharing of electron pairs between them. In one example, a covalent bond is the bond between oxygen and phosphorus, such as the phosphodiester bond in the backbone of a nucleic acid chain. In another example, a covalent bond is the bond between a nucleic acid barcode and a solid or semi-solid substrate, such as beads, for example, hydrogel beads.

[0081] Detection: This aims to determine the presence or absence of a reagent (such as a signal or specific nucleic acid, such as a nucleic acid barcode, or a protein). In some instances, this may further include quantification performed on a sample or sample fraction (such as a specific cell or multiple cells).

[0082] Detectable markers: compounds or compositions directly or indirectly conjugated to another molecule to facilitate the detection of said molecule. Specific non-limiting examples of markers include fluorescent tags, enzyme-linked tags, and radioisotopes. In some instances, the marker is linked to an antibody or nucleic acid to facilitate the detection of molecules to which the antibody or nucleic acid specifically binds. In specific instances, detectable markers comprise nucleic acid barcodes, such as source-specific barcodes.

[0083] DNA sequencing: The process of determining the nucleotide sequence of a given DNA molecule. Common methods include automated Sanger sequencing (AB13730x1 genome analyzer), pyrosequencing on solid-state vectors (454 sequencing, Roche), and sequencing using reversibly terminated synthetic methods. Genome analyzer, ligation sequencing Or sequencing can be performed using the synthesis of virtual terminators. Sequencing is performed. In some implementations, nucleic acid identity is determined by DNA or RNA sequencing. Typically, automated Sanger sequencing (AB13730x1 genome analyzer), pyrosequencing on solid vectors (454 sequencing, Roche), and sequencing using reversibly terminated synthesis methods can be used. Genome analyzer, ligation sequencing Or sequencing can be performed using the synthesis of virtual terminators. Sequencing methods include: Moleculo sequencing (see Voskoboynik et al. eLife 2013 2:e00569 and U.S. Patent Application No. 13 / 608,778 filed September 10, 2012); DNA nanosphere sequencing; single-molecule real-time (SMRT) sequencing; nanopore DNA sequencing; hybridization sequencing; mass spectrometry sequencing; and microfluidic Sanger sequencing.

[0084] In some implementations, DNA sequencing is performed using a chain termination method developed by Frederick Sanger, and is therefore referred to as “Sanger-based sequencing” or “SBS.” This technique utilizes sequence-specific termination of a DNA synthesis reaction using a modified nucleotide substrate. Extension is initiated at specific sites on the template DNA using short oligonucleotide primers complementary to the template in the region. DNA polymerase is used to extend the oligonucleotide primers in the presence of four deoxynucleotide bases (DNA building blocks) and a low concentration of chain-terminating nucleotides (most commonly dideoxynucleotides). The limited merging of chain-terminating nucleotides by DNA polymerase produces a series of related DNA fragments that terminate only at the location of the specific nucleotide. The fragments are then size-separated by electrophoresis using a polyacrylamide gel or in a narrow glass tube (capillary) filled with a viscous polymer. As an alternative to using labeled primers, labeled terminators are used; this method is often referred to as “dye-terminator sequencing.”

[0085] Pyrosequencing is an array-based method that has been commercialized by 454Life Sciences. In some embodiments of the array-based method, single-stranded DNA is annealed to beads and then... Amplification. These DNA-bound beads are then placed in wells on a fiber optic chip along with an enzyme that generates light in the presence of ATP. When free nucleotides are eluted from the chip, light is generated as PCR amplification occurs and ATP is produced when the nucleotides bind to their complementary base pairs. Adding one (or more) nucleotides triggers a reaction that generates a light signal, which is recorded, for example, by a charge-coupled device (CCD) camera within the instrument. The signal intensity is proportional to the number of nucleotides (e.g., homopolymer extensions) incorporated into a single nucleotide stream.

[0086] Compartment: A discrete volume or space that may contain target molecules and indexable nucleic acid identifiers (e.g., nucleic acid barcodes), such as a container, reservoir, or other arbitrarily defined volume or space or any combination thereof, said defined volume or space being defined by properties that prevent and / or inhibit the migration of target molecules, for example by physical properties that may be impermeable or semi-permeable, such as walls (e.g., pore walls), tubes, or droplet surfaces; or by other means such as chemical, diffusion rate limiting, electromagnetic, or light-induced. "Diffusion rate limiting" (e.g., diffusion-defined volume) means a space into which only certain molecules or reactants can enter, because diffusion limiting effectively defines the space or volume as a case of two parallel thin-layer streams, where diffusion restricts the migration of target molecules from one stream to another. "Chemically" defined volume or space means a space into which only certain target molecules can be present due to their chemical or molecular properties (e.g., size), where, for example, gel beads may prevent certain substances from entering the beads but not others, such as based on the surface charge of the beads, matrix size, or other physical properties, thereby allowing selection of substances that can enter the interior of the beads. The “electromagnetic” defined volume or space refers to a region of space that can be defined using the electromagnetic properties (such as charge or magnetic properties) of the target molecule or its carrier, such as the space where magnetic particles are trapped in a magnetic field or directly on a magnet. The “optical” defined volume refers to any region of space that can be labeled by irradiating it with visible, ultraviolet, infrared, or other wavelengths of light, such that only the target molecule within the defined space or volume can be labeled. One advantage of using non-walled or semi-permeable materials is that some reagents (such as buffers, chemical activators, or other agents) can be delivered through the discrete volume, while other materials (such as the target molecule) can be maintained within the discrete volume or space. Typically, the discrete volume will include a fluid medium (e.g., aqueous solution, oil, buffer, and / or culture medium capable of supporting cell growth) suitable for labeling target molecules with indexable nucleic acid identifiers under labeling-permissible conditions. Exemplary discrete volumes or spaces applicable to the disclosed methods include droplets (e.g., microfluidic droplets and / or emulsion droplets), hydrogel beads or other polymer structures (e.g., polyethylene glycol diacrylate beads or agarose beads), tissue slides (e.g., fixed formalin-embedded tissue slides having specific regions, volumes, or spaces defined by chemical, optical, or physical means), microscope slides having regions defined by reagents deposited in ordered arrays or irregular patterns, tubes (such as centrifuge tubes, microcentrifuge tubes, test tubes, cuvettes, conical tubes, etc.), bottles (such as glass bottles, plastic bottles, ceramic bottles, Erlenmeyer flasks, scintillation bottles, etc.), wells (such as wells in plates), plates, pipettes, or pipette tips, etc. In some embodiments, the compartments are aqueous droplets in a water-in-oil emulsion.

[0087] Hybridization: Oligonucleotides and their analogues hybridize via hydrogen bonds between complementary bases, including Watson-Crick, Hoogsteen, or anti-Hoogsteen hydrogen bonds. Typically, nucleic acids consist of nitrogenous bases, either pyrimidines (cytosine (C), uracil (U), and thymine (T)) or purines (adenine (A) and guanine (G)). These nitrogenous bases form hydrogen bonds between pyrimidines and purines, and this bonding of pyrimidine to purine is called "base pairing." More specifically, A bonds hydrogen to T or U, while G bonds to C. "Complementarity" refers to base pairing that occurs between two different nucleic acid sequences or two different regions of the same nucleic acid sequence.

[0088] "Specific hybridization" and "specific complementarity" are terms indicating a degree of complementarity sufficient to allow stable and specific binding between an oligonucleotide (or its analogue) and a DNA or RNA target. An oligonucleotide or oligonucleotide analogue does not need to be 100% complementary to its target sequence for specific hybridization. Oligonucleotides or analogues exhibit specific hybridization when there is a degree of complementarity sufficient to prevent non-specific binding to non-target sequences under conditions requiring specific binding. This type of binding is called specific hybridization.

[0089] Isolated: "Isolated" biological components (such as nucleic acids) have been substantially isolated or purified from other biological components naturally present in the organism's cells, such as extrachromatin DNA and RNA, proteins, and organelles. The term also covers nucleic acids and proteins prepared through recombinant expression in host cells, as well as chemically synthesized nucleic acids. It should be understood that the term "isolated" does not imply that the biological component is free of trace contaminants and can include nucleic acid molecules that are at least 50% isolated, such as at least 75%, 80%, 90%, 95%, 98%, 99%, or even 100% isolated.

[0090] Expression level: can refer to RNA expression level, protein expression level, or both.

[0091] Nucleic acids (molecules or sequences): deoxyribonucleotides or ribonucleotide polymers, including but not limited to cDNA, mRNA, genomic DNA, and synthetic (such as chemically synthesized) DNA or RNA or hybrids thereof. Nucleic acids can be double-stranded (ds) or single-stranded (ss). In the case of single-stranded nucleic acids, they can be sense or antisense strands. Nucleic acids can include natural nucleotides (such as A, T / U, C, and G), and may also include analogues of natural nucleotides, such as labeled nucleotides. Some examples of nucleic acids include the probes disclosed herein.

[0092] The main building blocks of DNA polymerized nucleotides are deoxyadenosine 5'-triphosphate (dATP or A), deoxyguanosine 5'-triphosphate (dGTP or G), deoxycytidine 5'-triphosphate (dCTP or C), and deoxythymidine 5'-triphosphate (dTTP or T). The main building blocks of RNA polymerized nucleotides are adenosine 5'-triphosphate (ATP or A), guanosine 5'-triphosphate (GTP or G), cytidine 5'-triphosphate (CTP or C), and uridine 5'-triphosphate (UTP or U).

[0093] In some instances, nucleotides include those containing modified bases, modified sugar moieties, and modified phosphate backbones, such as those described in U.S. Patent No. 5,866,336 to Nazarenko et al. Examples of modified base moieties that can be used to modify the structure of nucleotides at any position include, but are not limited to: 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, acetylcytidine, 5-(carboxyhydroxymethyl)uracil, 5-carboxymethylaminomethyl-2-thiouracil, 5-carboxymethylaminomethyluracil, dihydrouracil, β-D-galactosylqueosine, inosine, N-6-isopentenyladenine, 1-methylguanine, 1-methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, and 5-methylcytosine. N6-adenine, 7-methylguanine, 5-methylaminomethyluracil, methoxyaminomethyl-2-thiouracil, β-D-mannosyl piracetamidine, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, 2-methylthio-N6-isopentenyladenine, uracil-5-oxyacetic acid, pseudouracil, piracetamidine, 2-thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, uracil-5-oxyacetic acid methyl ester, uracil-S-oxyacetic acid, 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, 2,6-diaminopurine, and biotinylated analogs, etc. Examples of modified sugar moieties that can be used to modify nucleotides at any position in the structure include, but are not limited to, arabinose, 2-fluoroarabinose, xylose, and hexose, or modified components of the phosphate backbone such as thiophosphates, dithiophosphates, aminothiophosphates, aminophosphates, diaminophosphates, methylphosphonates, alkyl phosphates, or formacetals or their analogues.

[0094] Nucleic acid barcodes, barcodes, unique molecular identifiers, or UMIs: Short sequences of nucleotides (e.g., DNA, RNA, or combinations thereof) used as identifiers for associated molecules (such as target molecules and / or target nucleic acids). Nucleic acid barcodes or UMIs may have a length of at least, for example, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 nucleotides and may be single-stranded or double-stranded. One or more nucleic acid barcodes and / or UMIs may be linked or "tag-linked" to target molecules and / or target nucleic acids. This connection can be direct (e.g., barcode covalent or non-covalent binding to the target molecule) or indirect (e.g., via additional molecules, such as specific binders, like antibodies (or other proteins), or barcode-receiving adaptors (or other nucleic acid molecules). Target molecules and / or target nucleic acids can be labeled with multiple nucleic acid barcodes (such as nucleic acid barcode multiplicands) in a combined manner. Typically, nucleic acid barcodes are used to identify target molecules and / or target nucleic acids as originating from a specific compartment (e.g., discrete volume), possessing specific physical properties (e.g., affinity, length, sequence, etc.), or having undergone certain treatment conditions. The target can be... Molecules and / or target nucleic acids associate with multiple nucleic acid barcodes to provide information about all of these characteristics (and more). On the other hand, members of a given population of UMIs typically associate with individual members of a specific set of nucleic acid barcodes of the same specificity (e.g., discrete volume specificity, physical property specificity, or therapeutic condition specificity) (e.g., components covalently bound to or identical to the same molecule). Thus, for example, members of a set of source-specific nucleic acid barcodes with the same or matching barcode sequences may associate with distinct or different UMIs (e.g., components covalently bound to or identical to the same molecule).

[0095] Nucleic acid capture sequences: Nucleic acid sequences that specifically bind to another nucleic acid (such as a target nucleic acid and / or a barcode, such as a source-specific barcode or a target molecule identification barcode). Nucleic acid capture sequences (e.g., DNA, RNA, or hybrid molecules) recognize target nucleic acid molecules (e.g., DNA or RNA molecules) through hybridization or base-pairing interactions. Such capture sequences can be single-stranded or have a dangling portion of a nucleic acid sequence that includes a sequence capable of hybridizing to the target nucleic acid molecule. In some cases, nucleic acid capture sequences can be attached to nucleic acid barcodes, for example, by ligation or by synthesizing a single continuous nucleic acid with a nucleic acid-specific binder and a barcode. The sequence length required for hybridization can vary depending on, for example, nucleotide content and the conditions used, but generally, the length can be at least 4, 8, 12, 16, 20, 25, 30, 40, 50, 75, or 100 nucleotides.

[0096] Primers: Short nucleic acid molecules, such as DNA oligonucleotides, for example, sequences of at least 15 nucleotides, which can anneal to complementary nucleic acid molecules through nucleic acid hybridization to form a hybrid between the primer and the nucleic acid chain. Primers can be extended along the nucleic acid molecule using polymerases. Therefore, primers can be used to amplify nucleic acid molecules, wherein the primer sequence is specific to the nucleic acid molecule, for example, such that the primer will hybridize to the nucleic acid molecule under extremely stringent hybridization conditions. The specificity of primers increases with their length. Thus, for example, a primer including 30 consecutive nucleotides will anneal to the sequence with higher specificity than a corresponding primer with only 15 nucleotides. Therefore, to obtain greater specificity, probes and primers including at least 15, 20, 25, 30, 35, 40, 45, 50 or more consecutive nucleotides can be selected.

[0097] In certain instances, the primer is at least 15 nucleotides long, such as at least 15 consecutive nucleotides complementary to the nucleic acid molecule. Specific lengths of primers that can be used to practice the methods of this disclosure include primers having at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 45, at least 50 or more consecutive nucleotides complementary to the target nucleic acid molecule to be amplified, such as primers of 15-60 nucleotides, 15-50 nucleotides, or 15-30 nucleotides.

[0098] Nucleic acid sequences can be amplified using primer pairs, such as by PCR, real-time PCR, or other nucleic acid amplification methods known in the art. The "upstream" or "forward" primer is the primer at the 5' end of the reference site on the nucleic acid sequence. The "downstream" or "reverse" primer is the primer at the 3' end of the reference site on the nucleic acid sequence. Generally, the amplification reaction includes at least one forward and one reverse primer. This can be achieved, for example, by using a computer program designed for this purpose, such as Primer (version 0.5). In 1991, the Whitehead Institute for Biomedical Research (Cambridge, MA) derived PCR primer pairs from known sequences.

[0099] For example, Sambrook et al. (1989) Molecular Cloning: A Laboratory Manual, ColdSpring Harbor, New York; and Ausubel et al. (1987) Current Protocols in Molecular Biology, Greene Publ. Assoc. & Wiley-Intersciences describe methods for preparing and using primers. In one instance, the primers include markers.

[0100] Probe: An isolated nucleic acid capable of hybridizing to a specific nucleic acid (such as a nucleic acid barcode or target nucleic acid). A detectable label or reporter molecule can be attached to the probe. Typical labels include radioactive isotopes, enzyme substrates, cofactors, ligands, chemiluminescent or fluorescent agents, haptens, and enzymes. In some instances, probes are used to isolate and / or detect specific nucleic acids.

[0101] For example, Sambrook et al., Molecular Cloning: A Laboratory Manual, Cold Spring Harbor Laboratory Press (1989) and Ausubel et al., Current Protocols in Molecular Biology, Greene Publishing Associates and Wiley-Intersciences (1987) discuss labeling methods and guidance on selecting suitable labels for various purposes.

[0102] Probes are typically between approximately 15 and 160 nucleotides in length, such as 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, and 57. 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 10 7, 108, 109, 110, 111, 112, 113, 114, 115, 116, 117, 118, 119, 120, 121, 122, 123, 124, 125, 126, 127, 128, 129, 130, 131, 132, 133, 134, 135, 136, 137, 138, 139, 140, 141, 142, 143, 144, 145, 146 147, 148, 149, 150, 151, 152, 153, 154, 155, 156, 157, 158, 159, 160 consecutive nucleotides complementary to a specific nucleic acid molecule, such as 50-140 nucleotides, 75-150 nucleotides, 60-70 nucleotides, 30-130 nucleotides, 20-60 nucleotides, 20-50 nucleotides, 20-40 nucleotides, or 20-30 nucleotides.

[0103] Sequence identity / similarity: The identity / similarity between two or more nucleic acid sequences or two or more amino acid sequences is expressed as the degree of identity or similarity between the sequences. Sequence identity can be measured as a percentage of identity; the higher the percentage, the more identical the sequences. Homologous or orthologous nucleic acid or amino acid sequences have a relatively high degree of sequence identity / similarity when aligned using standard methods. The sequence alignment methods used for comparison are well known in the art. Various procedures and alignment algorithms are described in the following literature: Smith and Waterman, Adv. Appl. Math. 2:482, 1981; Needleman and Wunsch, J. Mol. Biol. 48:443, 1970; Pearson and Lipman, Proc. Natl. Acad. Sci. USA 85:2444, 1988; Higgins and Sharp, Gene, 73:237-44, 1988; Higgins and Sharp, CABIOS 5:151-3, 1989; Corpet et al., Nuc. Acids Res. 16:10881-90, 1988; Huang et al., Computer Appls. in the Biosciences 8, 155-65, 1992; and Pearson et al., Meth. Mol. Bio. 24:307-31, 1994. Altschul et al., J.Mol.Biol.215:403-10, 1990, presented detailed considerations regarding sequence alignment methods and homology calculations. The NCBI Basic Local Alignment Search Tool (BLAST) (Altschul et al., J.Mol.Biol.215:403-10, 1990) is available from several sources, including the National Center for Biological Information (NCBI, National Library of Medicine, Building 38A, Room 8N805, Bethesda, MD 20894) and on the Internet, for use in conjunction with the sequence analysis programs blastp, blastn, blastx, tblastn, and tblastx. BLAST is used to compare nucleic acid sequences, while blastp is used to compare amino acid sequences. Additional information can be found on the NCBI website.

[0104] After alignment, the number of matches is determined by counting the number of positions where the two sequences contain the same nucleotide or amino acid residues. The sequence identity percentage is determined by dividing the number of matches by the length of the sequence shown in the identified sequence, or by the hinge length (such as 100 consecutive nucleotide or amino acid residues in the sequence shown in the identified sequence), and then multiplying the resulting value by 100. For example, a nucleic acid sequence with 1166 matches when aligned with a test sequence having 1554 nucleotides is 75.0% identical to the test sequence (1166 ÷ 1554 * 100 = 75.0). The sequence identity percentage value is rounded to the nearest decimal place. For example, 75.11, 75.12, 75.13, and 75.14 are rounded down to 75.1, while 75.15, 75.16, 75.17, 75.18, and 75.19 are rounded up to 75.2. Length values ​​will always be integers. In another example, the target sequence containing a 20-nucleotide region aligned with 20 consecutive nucleotides of the identified sequence below contains a region that has 75% sequence identity with the identified sequence (i.e., 15 ÷ 20 * 100 = 75).

[0105] Specific binders: Agents that bind essentially or preferentially only to a defined target (such as polypeptides, proteins, enzymes, polysaccharides, oligonucleotides, DNA, RNA, recombinant vectors, or small molecules). In one example, a "capture fraction specific binder" is capable of binding to a capture fraction that is linked to a nucleic acid, such as a nucleic acid barcode.

[0106] Nucleic acid-specific binders bind primarily to specific nucleic acids, such as RNA, or specific regions within nucleic acids. In some embodiments, the specific binder is a nucleic acid barcode that specifically binds to the target nucleic acid of interest.

[0107] Protein-specific binders bind essentially only to a defined protein, or a specific region within a protein. For example, “specific binders” include antibodies and other agents that bind essentially to a given polypeptide. Antibodies can be monoclonal or polyclonal antibodies specific to a polypeptide, along with their immunogenic portions (“fragments”). The determination that a specific agent binds essentially only to a specific polypeptide can be readily achieved by using or modifying standard procedures. A suitable in vitro assay utilizes the Western blotting procedure (described in many standard texts, including Harlow and Lane, Using Antibodies: A Laboratory Manual, CSHL, New York, 1999).

[0108] Carrier: A solid or semi-solid substrate to which something (such as a nucleic acid barcode, e.g., a source-specific barcode) may be attached. The attachment may be removable. Non-limiting examples of carriers suitable for the methods of this disclosure include hydrogels, cells, beads, columns, filters, the surface of a glass slide, or the inner wall or container of a compartment (such as a well in a microtiter plate). In some embodiments, the carrier is a hydrogel (such as hydrogel beads) coupled with one or more source-specific barcodes. The source-specific barcode reversibly coupled to the carrier may be detached from the carrier, for example, by enzyme cleavage at a cleavage site on the source-specific barcode. The carrier may be present in a compartment as described herein. In some embodiments, the carrier is a hydrogel bead present in an emulsion droplet.

[0109] Target molecule: This refers to a molecule whose source, expression, type, etc., are desired, or in some instances, a molecular complex. In some embodiments, the target molecule is labeled according to the methods disclosed herein. Examples of target molecules include, but are not limited to, peptides, polypeptides, proteins, antibodies, antibody fragments, amino acids, nucleic acids (such as RNA and DNA), nucleotides, carbohydrates, polysaccharides, lipids, small molecules, organic molecules, inorganic molecules, and their complexes. In some instances, multiple target molecules (e.g., the same target molecule or multiple copies of more than one different target molecule) may be present in a sample (such as a sample separated into compartments such as discrete volumes or spaces). In some embodiments, the target nucleic acid molecule (e.g., an RNA molecule) encodes a polypeptide target molecule (e.g., a protein) present in the compartment. In certain embodiments, the polypeptide target molecule and the nucleic acid target molecule are expressed by cells or a cell-free expression system present in the compartment.

[0110] Target molecules can be bound by a specific binder that associates with nucleic acid barcodes, such that the target molecules are labeled with nucleic acid barcodes (e.g., source-specific barcodes and / or target-specific barcodes). Multiple target molecules (e.g., multiple copies of the same target molecule or multiple copies of more than one different target molecule) may be present and labeled in compartments such as discrete volumes or spaces. In some embodiments, the target nucleic acid molecule (such as a DNA or RNA molecule) encodes a target molecule, such as a target protein present in the same compartment, and the target nucleic acid molecule and the peptide target molecule are labeled with the same barcode or matching barcodes (e.g., barcodes pre-identified as corresponding to each other) (such as source-specific nucleic acid barcodes). In a particular embodiment, the target molecules and target nucleic acid molecules are expressed by a cellular or cell-free expression system present in a specific compartment.

[0111] Target nucleic acid molecule: Any nucleic acid present or believed to be present in a sample from which information is desired. In some embodiments, the target nucleic acid of interest is RNA, such as mRNA, for example, mRNA encoding the target molecule. In some embodiments, the target nucleic acid of interest is DNA.

[0112] Test agent: Any agent used to test for effects (e.g., effects on cells of interest or target molecules). In some embodiments, the test agent is a compound such as a chemotherapeutic agent, antibiotic, or even an agent with unknown biological properties. In some instances, the test agent is a peptide or protein, such as an antibody, antigen, or immunogen.

[0113] Under permissible binding conditions: a phrase used to describe any environment that allows the desired activity, such as conditions that enable two or more molecules (such as nucleic acid molecules and / or protein molecules) to bind.

[0114] The following describes methods and materials suitable for practicing or testing this disclosure. These methods and materials are illustrative only and are not intended to be limiting. Other methods and materials similar to or equivalent to those described herein may be used. For example, various general and more specific references describe conventional methods well-known in the field to which this disclosure pertains, including, for instance, Sambrook et al., *Molecular Cloning: A Laboratory Manual*, 2nd ed., Cold Spring Harbor Laboratory Press, 1989; Sambrook et al., *Molecular Cloning: A Laboratory Manual*, 3rd ed., Cold Spring Harbor Press, 2001; Ausubel et al., *Current Protocols in Molecular Biology*, Greene Publishing Associates, 1992 (and supplements in 2000); Ausubel et al., *Short Protocols in Molecular Biology: A Compendium of Methods from Current Protocols in Molecular Biology*, 4th ed., Wiley & Sons, 1999; Harlow and Lane, *Antibodies: A Laboratory Manual*, Cold Spring Harbor Laboratory Press, 1990; and Harlow and Lane, *Using Antibodies: A Laboratory Manual*, Cold Spring Harbor Laboratory Press, 1990. Press, 1999. Furthermore, the materials, methods, and examples described are illustrative only and are not intended to be limiting.

[0115] II. Description of Several Implementation Schemes

[0116] A. Introduction

[0117] This disclosure provides methods and compositions for high-throughput labeling of target molecules using specific nucleic acid barcodes. The source-specific nucleic acid barcodes that specifically label target molecules can be used to multiplex the identification, quantity, and / or activity of target molecules, for example, by detecting the sequence of the nucleic acid barcode (such as by sequencing or other methods providing information about the nucleic acid sequence). Because source-specific barcodes can be traced back to the original sample or subsample (e.g., individual compartments), target molecules and / or target nucleic acids derived from the sample can be separated, processed differentially, and then pooled for multiplex analysis. Therefore, the methods and compositions of this disclosure enable highly complex combinatorial analyses by associating relevant information (such as phenotypic information) about different target molecules, their origin, physical properties, and / or different processing conditions with different nucleic acid barcodes, which can be deconvolved from the pooled sample, for example, using high-throughput sequencing technology.

[0118] This disclosure further features the ability to co-label target peptides of interest (such as antibodies, antigens, and immunogens) and DNA or RNA molecules (e.g., DNA (such as cDNA) and / or RNA encoding the co-labeled target peptide, or DNA and / or RNA of a reporter cell line) using the same or matching source-specific nucleic acid barcodes, thereby achieving, for example, the coupling of genotypes and phenotypes of interest. Thus, this disclosure enables rapid, large-scale multiplex screening of candidate target molecules. The methods of this disclosure can be further used to characterize gene expression systems in large-scale multiplex forms, for example, in the form of multipart DNA components.

[0119] In non-limiting applications of the disclosed methods to antibodies, the affinity of an antibody (or an antibody ensemble) for a specific target (such as a protein and / or a specific epitope visible on a protein) can be determined. In another instance, the methods disclosed herein can be used to potentially determine the expression, such as relative expression, of cell surface markers on cell types of interest. By utilizing arrays or ensembles of antibodies labeled with indexable nucleic acid identifiers, cell populations can be analyzed multiplexed to determine cell surface markers expressed on individual cells (e.g., individual cells separated into individual compartments).

[0120] In various instances of the methods disclosed herein, the phenotype of an individual cell or molecule (e.g., the binding characteristics of an antibody produced by a cell) can be associated with a corresponding genotype (i.e., a nucleic acid molecule, such as a nucleic acid molecule encoding an antibody). This can be performed after the aggregation and overall screening of a large number of cellular or target molecules (e.g., antibodies) produced by the cell (e.g., by a high-multiple affinity measurement). Therefore, this disclosure provides a high-throughput method for conjugating genotypes with phenotypes (e.g., antibody expression) in a large-scale, multi-contextual context, if desired.

[0121] B. Indexing methods

[0122] This paper discloses methods for genotype-phenotype coupling or indexing, for example, by assigning a set of target molecules to a sample or a set of target nucleic acids while maintaining information about the sample origin of the target molecules and target nucleic acids. In other words, using the methods disclosed herein, specific target molecules can be paired with specific target nucleic acids, or in some cases, with a set of target nucleic acids and target molecules. The disclosed methods include labeling target molecules and / or target nucleic acids with different nucleic acid barcodes associated with specific compartments, thus allowing the origin of individual compartments or labeled molecules within individual compartments to be determined at a later time (e.g., at the end of the experiment). Nucleic acid barcodes that are already present when introducing target molecules and / or target nucleic acids into a compartment or generating target molecules and / or target nucleic acids in a compartment can be used to label target molecules and / or target nucleic acids.

[0123] Some embodiments of the disclosed method include providing a sample comprising cells, or in some cases, a non-cellular system, and separating individual cells from the sample or discrete portions of the non-cellular system from the sample into individual compartments. Each compartment also includes a source-specific barcode comprising a unique nucleic acid recognition sequence (such as a unique nucleic acid recognition sequence comprising DNA, RNA, or combinations thereof), which is maintained, for example, after collection and analysis, or carries information about the source of the cells or non-cellular system in the sample, such as which compartment the labeled target molecule and / or target nucleic acid originates from. In this way, the compartments are effectively labeled with source-specific barcodes, so that the contents of the compartments can be tracked using the source-specific barcodes throughout the experiment and / or analysis (e.g., an experiment and / or analysis that exposes individual compartments to different conditions to measure the effect of said conditions on target molecules and / or target nucleic acids). Target molecules and / or target nucleic acids present in the compartments are labeled with source-specific barcodes present in the compartments to form source-labeled molecules and / or source-labeled nucleic acids. As discussed above, each compartment contains source-labeled molecules and / or source-labeled nucleic acids carrying the same or matching unique indexed nucleic acid recognition sequences or source-specific barcodes. Detection of the nucleotide sequence of the source-specific barcode assigns the target molecule set to the target nucleic acid in the sample or sample set, while maintaining information about the sample or compartment origin of the target molecules and target nucleic acids. The source-specific barcode sequence can be detected by any method known in the art, such as amplification, sequencing, hybridization, or any combination thereof, among other sequences (such as the sequence of the target nucleic acid and / or other barcodes).

[0124] In some implementations, the method is used to label a set of target molecules, which are nucleic acids from a sample, such as amplicon from cells, cell aggregates, or non-cellular systems. Figure 1 An example of this method is shown in the image. In such methods, source-specific barcodes are used (…). Figure 2 (Examples are shown in the text) Nucleic acids from individual compartments are labeled with different source-specific barcodes, while nucleic acids from another compartment or multiple other compartments are labeled with different source-specific barcodes, thereby allowing multiple analysis of nucleic acids, such as studying differences in gene expression in individual compartments when exposed to different conditions or belonging to different cell types.

[0125] As disclosed herein, unique nucleic acid identifiers (such as nucleic acid barcodes) are used to label target molecules and / or target nucleic acids, such as source-specific barcodes. Nucleic acid identifiers (such as nucleic acid barcodes) may include short nucleotide sequences that can be used as identifiers of associating molecules, locations, or conditions. In some embodiments, the nucleic acid identifier further includes one or more unique molecular identifiers and / or barcode-receiving adaptors. The length of the nucleic acid identifier may be, for example, about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 35, 40, 45, 50, 60, 70, 80, 90, or 100 base pairs (bp) or nucleotides (nt). In some implementations, nucleic acid identifiers can be constructed in a combinatorial manner by combining randomly selected indices (e.g., about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 indices). Each such index is a short sequence of nucleotides (e.g., DNA, RNA, or a combination thereof) with a different sequence. The length of the index can be, for example, about 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, or 25 bp or nt. Nucleic acid identifiers can be generated, for example, by a split-pool synthesis method (such as those described, for example, in International Patent Publications WO 2014 / 047556 and WO2014 / 143158, each of which is incorporated herein by reference in its entirety).

[0126] One or more nucleic acid identifiers (e.g., nucleic acid barcodes) can be linked or "tag-linked" to a target molecule. This link can be direct (e.g., the nucleic acid identifier covalently or non-covalently binds to the target molecule) or indirect (e.g., via an additional molecule). Such indirect links may, for example, include a barcode bound to a binding portion that recognizes the target molecule. In some embodiments, the barcode is linked to a protein G, and the target molecule is an antibody or antibody fragment. Linking the barcode to a target molecule (e.g., a protein or other biomolecule) can be performed using standard methods well known in the art. For example, the barcode can be linked via a cysteine ​​residue (e.g., a C-terminal cysteine ​​residue). In other instances, the barcode can be chemically introduced into a peptide (e.g., an antibody) via various functional groups on the peptide using appropriate group-specific reagents (see, for example, www.drmr.com / abcon). In some embodiments, as described herein, barcode tagging can be performed via a barcode receiving an adaptor associated (e.g., linked) with the target molecule.

[0127] Target molecules can optionally be labeled in combination with multiple barcodes (e.g., multiple barcodes using one or more specific binders that specifically identify the target molecule), thus greatly increasing the number of possible unique identifiers within a particular barcode collection. In some embodiments, barcodes are added to an increasing barcode multiplex linked to the target molecule, for example, one at a time. In other embodiments, multiple barcodes are assembled prior to being linked to the target molecule. Compositions and methods for multiplying multiple barcodes are described, for example, in International Patent Publication No. WO 2014 / 047561, which is incorporated herein by reference in its entirety.

[0128] Unique molecular identifiers are subtypes of nucleic acid barcodes that can be used, for example, to normalize samples for variable amplification efficiencies. For instance, in various embodiments characterized by solid or semi-solid carriers (e.g., hydrogel beads) linked to nucleic acid barcodes (e.g., multiple barcodes having the same sequence), each of the barcodes can be further coupled to a unique molecular identifier, such that each barcode on a particular solid or semi-solid carrier receives a different unique molecular identifier. The unique molecular identifier can then be transferred, for example, to a target molecule having the associated barcode, such that the target molecule receives not only the nucleic acid barcode but also a unique identifier among the identifiers derived from the solid or semi-solid carrier.

[0129] Nucleic acid identifiers may further include unique molecular identifiers and / or additional barcodes specific to, for example, a common carrier linked to one or more of said nucleic acid identifiers. Thus, an aggregate of target molecules may be added, for example, to a compartment containing multiple solid or semi-solid carriers (e.g., beads) representing different treatment conditions (and / or, for example, one or more additional solid or semi-solid carriers may be sequentially added to the compartment after the introduction of the target molecule aggregate), such that the precise combination of conditions exposed to a given target molecule can be determined subsequently by sequencing the unique molecular identifier associated with it.

[0130] Labeled target molecules and / or target nucleic acids associated with a source-specific nucleic acid barcode (optionally combined with other nucleic acid barcodes as described herein) can be amplified using methods known in the art, such as polymerase chain reaction (PCR). For example, the nucleic acid barcode may contain a universal primer recognition sequence that can be bound by PCR primers for PCR amplification and subsequent high-throughput sequencing. In some embodiments, the nucleic acid barcode includes or is linked to a sequencing adaptor (e.g., a universal primer recognition sequence) such that both the barcode and the sequencing adaptor element are coupled to the target molecule. In a particular instance, PCR is used, for example, to amplify the sequence of the source-specific barcode. In some embodiments, the source-specific barcode further includes a sequencing adaptor. In some embodiments, the source-specific barcode further includes a universal priming site. In some embodiments, a nucleic acid identifier (e.g., a nucleic acid barcode) can be linked to a sequence that allows amplification and sequencing (e.g., for use with...). The sequenced P7 (SEQ ID NO:11), SBS3 (SEQ ID NO:2), and P5 (SEQ ID NO:12) elements. In some embodiments, the nucleic acid barcode may further include a hybridization site for a primer (e.g., a single-stranded DNA primer) for attachment to the end of the barcode. For example, a source-specific barcode may be a nucleic acid comprising a barcode and a hybridization site for a specific primer. In a particular embodiment, the set of source-specific barcodes includes, for example, unique primer-specific barcodes prepared using the randomized oligonucleotide type NNNNNNNNNNNN (SEQ ID NO:1). Other lengths and components of barcodes contemplated for use in this invention are provided in this disclosure.

[0131] Optionally, nucleic acid barcodes (or multiples thereof), target nucleic acid molecules (e.g., DNA or RNA molecules), nucleic acids encoding target peptides or polypeptides, and / or nucleic acids encoding specific binding agents can be sequenced using any method known in the art (e.g., high-throughput sequencing methods, also known as next-generation sequencing or deep sequencing). Nucleic acid target molecules labeled with barcodes (e.g., source-specific barcodes) can be sequenced using said barcodes to produce single reads of both the target molecule and the barcode and / or contigs containing said sequences, or portions thereof. Exemplary next-generation sequencing technologies include, for example... Sequencing methods include Ion Torrent sequencing, 454 sequencing, SOLiD sequencing, and nanopore sequencing.

[0132] In some implementations, the sequence of the labeled target molecule is determined using methods other than sequencing. For example, variable-length probes or primers can be used to distinguish barcodes labeled with different target molecules based on, for example, the length of the barcode, the length of the target nucleic acid, or the length of the nucleic acid encoding the target polypeptide (e.g., source-specific barcodes). In other cases, the barcode may include a sequence identifying, for example, the molecular type of a specific target molecule (e.g., polypeptide, nucleic acid, small molecule, or lipid). For example, in an aggregate of labeled target molecules containing multiple types of target molecules, polypeptide target molecules may receive one recognition sequence, while target nucleic acid molecules may receive different recognition sequences. Such recognition sequences can be used to selectively amplify barcodes labeled with specific types of target molecules, for example, by using PCR primers specific to recognition sequences specific to specific types of target molecules. For example, barcodes labeled with polypeptide target molecules can be selectively amplified from an aggregate, thereby retrieving barcodes only from a subset of polypeptides in the target molecule aggregate.

[0133] In some embodiments of the disclosed method, determining the identity of nucleic acids (such as nucleic acid barcodes) involves detection via nucleic acid hybridization. Nucleic acid hybridization involves providing a probe and target nucleic acid under conditions where the probe and its complementary target can form a stable hybrid duplex through complementary base pairing. Typically, nucleic acids that do not form a hybrid duplex are then washed away by detecting the attached detectable tag, leaving the hybrid nucleic acid to be detected. It is generally accepted that nucleic acids are denatured by increasing the temperature or decreasing the salt concentration of the buffer containing the nucleic acid. Hybrid duplexes (e.g., DNA:DNA, RNA:RNA, or RNA:DNA) will form even under low-tightness conditions (e.g., low temperature and / or high salt), even if the annealed sequences are not completely complementary. Therefore, at lower tightness, the specificity of hybridization is reduced. Conversely, at higher tightness conditions (e.g., higher temperature or lower salt), successful hybridization requires fewer mismatches. Those skilled in the art will appreciate that hybridization conditions can be designed to provide different levels of tightness.

[0134] Generally, there is a trade-off between hybridization specificity (strictness) and signal strength. Therefore, in one implementation, washing is performed with the highest strictness, yielding consistent results and providing a signal strength greater than approximately 10% of the background strength. Thus, the hybridization array can be washed with solutions of progressively higher strictness, and readings are taken between each wash. Analysis of the resulting dataset will reveal a wash strictness beyond which the hybridization pattern does not change significantly, and which provides sufficient signal about the specific oligonucleotide probe of interest. In some instances, Northern blotting or in situ hybridization (Parker and Barnes, Methods in Molecular Biology 106:247-283, 1999); RNase protection assays (Hod, Biotechniques 13:852-4, 1992); and PCR-based methods, such as reverse transcription polymerase chain reaction (RT-PCR) (Weis et al., Trends in Genetics 8:263-4, 1992), are used to detect RNA.

[0135] In one embodiment, hybridized nucleic acids are detected by detecting one or more tags linked to sample nucleic acids. Tags can be incorporated by any of many methods. In one instance, tags are incorporated simultaneously during the amplification step in the preparation of sample nucleic acids. Thus, for example, polymerase chain reaction (PCR) using labeled primers or labeled nucleotides will provide labeled amplified products. In one embodiment, as described above, tags are incorporated into transcribed nucleic acids using transcriptional amplification of labeled nucleotides (such as fluorescein-labeled UTPs and / or CTPs).

[0136] Suitable detectable labels include any composition detectable by spectroscopic, photochemical, biochemical, immunochemical, electrical, optical, or chemical methods. Applicable labels include biotin for staining with labeled streptavidin conjugates, and magnetic beads (e.g., DYNABEADS). TM Fluorescent dyes (e.g., fluorescein, Texas red, rhodamine, green fluorescent protein, etc.), radioactive labels (e.g.) 3 H, 125 I, 35 S, 14 C or 32P), enzymes (such as horseradish peroxidase, alkaline phosphatase, and other enzymes commonly used in ELISA), and colorimetric labels, such as colloidal gold or colored glass or plastic (such as polystyrene, polypropylene, latex, etc.) beads. Patents teaching the use of such labels include U.S. Patent Nos. 3,817,837; 3,850,752; 3,939,350; 3,996,345; 4,277,437; 4,275,149; and 4,366,241.

[0137] Methods for detecting such labels are well known. For example, photographic film or a scintillation counter can be used to detect radioactive labels, while photodetectors can be used to detect emitted light to detect fluorescent labels. Enzyme labels are typically detected by providing a substrate to the enzyme and detecting the reaction products produced by the enzyme's action on the substrate, while colorimetric labels are detected simply by visually observing the colored label.

[0138] A label can be added to the target (sample) nucleic acid before or after hybridization. A so-called "direct label" is a detectable label that is directly linked to the target (sample) nucleic acid or incorporated into it before hybridization. In contrast, a so-called "indirect label" is linked to the hybrid duplex after hybridization. Typically, the indirect label is linked to a binding moiety that has already been linked to the target nucleic acid before hybridization. Thus, for example, the target nucleic acid can be biotinylated before hybridization. After hybridization, the avidin-conjugated fluorophore binds to the biotin-containing hybrid duplex, providing an easily detectable label (see Laboratory Techniques in Biochemistry and Molecular Biology, Vol. 24: Hybridization With Nucleic Acid Probes, ed. P. Tijssen, Elsevier, NY, 1993).

[0139] In some embodiments, labeling the target molecule includes directly linking a source-specific barcode to the target molecule. In some embodiments, labeling the target molecule includes indirectly linking a source-specific barcode to the target molecule. Indirect linking includes binding a target-specific binder to the target molecule, wherein the target-specific binder is indirectly or directly linked to the source-specific barcode, for example, through covalent linking. In some embodiments, the specific binder is an antibody, such as a complete antibody or an antibody fragment, such as an antigen-binding fragment. Alternatively, the specific binder may be a protein or peptide that is not an antibody. Therefore, specific binders can be, for example, kinases, phosphatases, proteasome proteins, protein chaperones, receptors (e.g., innate immune receptors or signal peptide receptors), synthetic antibodies, artificial antibodies, proteins having thioredoxin folds (e.g., disulfide isomerase, DsbA, glutathione, glutathione S-transferase, calcitonin, glutathione peroxidase, or glutathione peroxidase), proteins having folds derived from thioredoxin folds, repetitive proteins, proteins known to participate in protein complexes, proteins known in the art to be capable of participating in protein-protein interactions, or any variants thereof (e.g., variants with altered structure or binding properties). Specific binders can be any protein or polypeptide having a protein-binding domain known in the art, including any natural or synthetic protein containing a protein-binding domain. Specific binders can also be any protein or polypeptide having a polynucleotide-binding domain known in the art, including any natural or synthetic protein containing a polynucleotide-binding domain. In some cases, specific binders are recombinant specific binders.

[0140] Specific binding agents (e.g., antibodies) can be linked to nucleic acid barcodes (e.g., source-specific barcodes). For example, the specific binding agent may include cysteine ​​residues that can be linked to the nucleic acid barcode. In other cases, the binding portion may be a nucleic acid linked to the barcode. The specific binding agent linked to the nucleic acid barcode can recognize a target molecule of interest. The nucleic acid barcode can identify the specific binding agent as recognizing a specific target molecule of interest. The nucleic acid barcode may be, for example, cleavable from the specific binding agent after it has been bound to the target molecule.

[0141] Nucleic acid barcodes can be sequenced, for example, after cleavage, to determine the presence, quantity, or other characteristics of target molecules. In some embodiments, the nucleic acid barcode can be further linked to another nucleic acid barcode. For example, the nucleic acid barcode can be cleaved from the binding site (e.g., a peptide tag cleaved from the target molecule) after the binding site is attached to the target molecule or tag, and then the nucleic acid barcode can be linked to a source-specific barcode. The resulting nucleic acid barcode multiplicands can be pooled with other such multiplicands and sequenced. Sequencing reads can be used to identify which target molecules were initially present in which compartment.

[0142] In some embodiments, the target molecule comprises a target polypeptide, and the specific binder of the target molecule specifically bound to the sample comprises a polypeptide-specific binder specifically bound to the target polypeptide. In some embodiments, the polypeptide-specific binder comprises an antibody or a fragment thereof and / or a protein-binding domain or a fragment thereof, or a nucleic acid sequence specifically bound to the target polypeptide, for example, if the target polypeptide includes a nucleic acid-binding domain. In some embodiments, the target molecule-specific binder specifically binds to both the target molecule and the source-specific barcode. In some instances, the target molecule and the target molecule-specific binder binding both the target molecule and the source-specific barcode are incubated together, and the target molecule-specific binder not bound to the target molecule and / or the source-specific barcode is removed before being separated into individual compartments. In some instances, the target molecule-specific binder comprises a target molecule-specific binder barcode that encodes the identity of the target molecule-specific binder. In some instances, the target molecule-specific binder barcode may bind to the source-specific barcode via base-pairing interactions. In certain instances, the source-specific barcode is a primer for the complementary strand used to synthesize a target-specific binding barcode. In some instances, the sequence of the target-specific binding barcode is detected, among other sequences. The sequence of the target-specific binding barcode can be detected by any method known in the art, such as amplification, sequencing, hybridization, and any combination thereof. In addition to the source-specific nucleic acid barcode, the target molecule can be labeled with additional nucleic acid barcodes (optionally in the form of nucleic acid barcode polymorphs) in a specific manner based on any of a number of different properties of the target molecule and / or the conditions under which it is exposed, thereby facilitating characterization at other levels. In other embodiments, the target molecule is a polypeptide and is directly labeled with nucleic acid barcodes (such as target-specific binding and / or source-specific barcodes) via encoded cysteine ​​residues (e.g., C-terminal cysteine ​​residues).

[0143] In specific instances, such as PCR, the sequence of a target molecule-specific binding barcode may be amplified. The nucleic acid encoding the antigen or specific binding agent may be subcloned into an expression vector for producing the protein (e.g., an expression vector for producing the binding moiety in, for example, *E. coli*). The resulting binding moiety may be purified, for example, by affinity chromatography. In embodiments where the specific binding agent comprises one or more fragments substantially similar to an antibody or antibody fragment, the fragment may be incorporated into a known antibody framework for expression. For example, if the specific binding agent is scFv, then the heavy and light chain sequences of scFv may be cloned into a vector for expressing those chains within an IgG molecule.

[0144] In some embodiments, the target molecule comprises a target nucleic acid, such as DNA or RNA, and a specific binder that specifically binds to the target molecule in the sample comprises a nucleic acid sequence or nucleic acid-binding domain that specifically binds to and / or hybridizes to the target nucleic acid. The target nucleic acid includes RNA, such as mRNA; and DNA, such as cDNA. In some embodiments of the disclosed methods, cDNA is synthesized from the target nucleic acid, wherein the cDNA comprises a nucleic acid sequence or fragment thereof of the target nucleic acid and a source-specific barcode sequence. In some instances, the source-specific barcode is a primer for cDNA synthesis. In some embodiments, the target nucleic acid or its complement encodes a polypeptide of interest. In some embodiments, the target molecule comprises target DNA, and a specific binder that specifically binds to the target molecule in the sample comprises a nucleic acid sequence or DNA-binding domain that specifically binds to and / or hybridizes to the target DNA.

[0145] In some embodiments, the source-specific barcode is reversibly coupled to a solid or semi-solid substrate. In some embodiments, the source-specific barcode further includes a nucleic acid capture sequence that specifically binds to the target nucleic acid and / or a specific binder that specifically binds to the target molecule. In a particular embodiment, the source-specific barcode comprises two or more groups of source-specific barcodes, wherein a first group contains a nucleic acid capture sequence and a second group contains a specific binder that specifically binds to the target molecule. Figure 3 The diagram illustrates this situation. In some instances, the first source-specific barcode population further includes target nucleic acid barcodes, wherein the target nucleic acid barcodes identify the population as a population labeled with nucleic acids. In some instances, the second source-specific barcode population further includes target molecule barcodes, wherein the target molecule barcodes identify the population as a population labeled with target molecules.

[0146] Nucleic acid barcodes may be cleavable from a specific binder, for example, after the specific binder has bound to the target molecule. In some embodiments, the source-specific barcode further includes one or more cleavage sites. In some embodiments, at least one cleavage site is oriented such that cleavage at said site releases the source-specific barcode from a substrate (such as beads, e.g., hydrogel beads) to which it is coupled. In some embodiments, at least one cleavage site is oriented such that cleavage at said site releases the source-specific barcode from the target molecule specific binder. In some embodiments, the cleavage site is an enzyme cleavage site, such as a nuclease site present in a specific nucleic acid sequence. In other embodiments, the cleavage site is a peptide cleavage site, such that a specific enzyme can cleave the amino acid sequence. In other embodiments, the cleavage site is a chemical cleavage site.

[0147] In some implementations, each of the source-specific barcodes includes one or more indexes, one or more sequences for gene-specific capture and / or amplification, and / or one or more sequences for sequencing library construction.

[0148] In some embodiments, a target molecule is linked to a source-specific barcode receiver adaptor, such as a nucleic acid. In some embodiments, the source-specific barcode receiver adaptor includes a protrusion, and the source-specific barcode includes a sequence capable of hybridizing to the protrusion. A barcode receiver adaptor is a molecule configured to accept or receive a nucleic acid barcode (such as a source-specific nucleic acid barcode). For example, a barcode receiver adaptor may include a single-stranded nucleic acid sequence (e.g., a protrusion) capable of hybridizing to a given barcode (e.g., a source-specific barcode), for example, via a sequence complementary to a portion or all of the nucleic acid barcode. In some embodiments, this portion of the barcode is a standard sequence that remains constant across individual barcodes. Hybridization couples the barcode receiver adaptor to the barcode. In some embodiments, the barcode receiver adaptor may be associated (e.g., linked) with a target molecule. Thus, the barcode receiver adaptor may act as a means of linking a source-specific barcode to a target molecule. The barcode receiver adaptor may be linked to the target molecule according to methods known in the art. For example, a barcode receiver adaptor can be attached to a polypeptide target molecule at a cysteine ​​residue (e.g., a C-terminal cysteine ​​residue). The barcode receiver adaptor can be used to identify specific conditions associated with one or more target molecules, such as the source cell or source compartment. For example, the target molecule could be a cell surface protein expressed by a cell that receives a cell-specific barcode receiver adaptor. When a cell is exposed to one or more conditions, the barcode receiver adaptor can be conjugated to one or more barcodes, allowing the original source cell of the target molecule and the conditions to which the cell was exposed to to be subsequently determined by recognizing the sequence of the barcode receiver adaptor / barcode multiplier.

[0149] In some embodiments, more than one target-specific binder is linked to nucleic acid barcodes having the same sequence, such as source-specific barcodes. In some embodiments, more than one target-specific binder is linked to nucleic acid barcodes with different sequences. In some cases, multiple target-specific binders may be added to the compartment. Alternatively, multiple target-specific binders may be added separately. The different target-specific binders may optionally be related to, for example, experimental conditions.

[0150] One of the desirable features of the disclosed method is that it allows for the analysis of samples (such as the contents of multiple compartments) together in a single reaction (e.g., a pooling reaction). Therefore, in some instances, individual compartments are pooled to form a pooled sample. Target molecules and / or target nucleic acids labeled according to the disclosed method from multiple compartments can be combined to form a pool. For example, labeled target molecules and / or target nucleic acids in multiple emulsion droplets can be combined by disrupting the emulsion. Therefore, in some embodiments, the emulsion is disrupted. The aggregates may contain labeled target molecules and / or target nucleic acids from a large number of individual compartments or discrete volumes (e.g., at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 100, 500, 1,000, 2,500, 5,000, 10,000, 50,000, 100,000, 500,000, 1,000,000, 2,000,000 or more; in various instances, such as those utilizing plates, these numbers may be, for example, at least 6, 24, 96, 192, 384, 1,536, 3,456 or 9,600), thus facilitating the simultaneous processing of very large quantities of samples (e.g., by high multiple affinity measurements), resulting in high efficiency.

[0151] Labeled target molecules and / or target nucleic acids can be separated from the aggregate. Exemplary separation techniques include, but are not limited to, affinity capture, immunoprecipitation, chromatography (e.g., size exclusion chromatography, hydrophobic interaction chromatography, reversed-phase chromatography, ion exchange chromatography, affinity chromatography, metal-binding chromatography, immunoaffinity chromatography, high-performance liquid chromatography (HPLC), and liquid chromatography-mass spectrometry (LC-MS)), electrophoresis, hybridization to capture oligonucleotides, phenol-chloroform extraction, microcolumn purification, or precipitation with ethanol or isopropanol. Chromatography is described in detail, for example, by Hedhammar et al. (“Chromatographic methods for protein purification,” Royal Institute of Technology, Stockholm, Sweden), which is incorporated herein by reference. Such techniques may utilize capture molecules that recognize labeled target molecules or barcodes or binding moieties associated with said target molecules. For example, protein G can be used to separate target antibodies via affinity capture. The labeled target molecules can be further labeled with capture tags (such as biotin). Multiple target molecules can be separated simultaneously or separately (e.g., multiple identical target molecules, or a group of target molecules including multiple different target molecules).

[0152] In some embodiments, the source-specific barcode further includes a covalently or non-covalently linked capture portion. Thus, in some embodiments, the source-specific barcode and anything bound to or attached thereto, including the capture portion, are captured using a specific binder that specifically binds to the capture portion. In some embodiments, the capture portion is adsorbed or otherwise captured onto a surface. In certain embodiments, the targeting probe is labeled with biotin, for example, by incorporating biotin-16-UTP during in vitro transcription, thereby allowing subsequent capture by streptavidin. Other means for labeling, capturing, and detecting source-specific barcodes include: incorporating aminoallyl-labeled nucleotides, incorporating thiol-labeled nucleotides, incorporating nucleotides containing allyl or azide groups, and many other methods described in Bioconjugate Techniques (2nd edition), Greg T. Hermanson, Elsevier (2008), which is specifically incorporated herein by reference. In some embodiments, the targeting probe is covalently coupled to a solid support or other capture device prior to contact with the sample using methods such as incorporating an aminoally labeled nucleotide followed by coupling 1-ethyl-3-(3-dimethylaminopropyl)carbodiimide (EDC) to a carboxyl-activated solid support, or other methods described in Bioconjugate Techniques. In some embodiments, a specific binding agent has been immobilized on the solid support, thereby separating the source-specific barcode.

[0153] "Solid or semi-solid support / carrier" is intended to be any carrier capable of binding a source-specific barcode. Well-known carriers include hydrogels, glass, polystyrene, polypropylene, polyethylene, dextran, nylon, amylase, natural and modified cellulose, polyacrylamide, agarose, gabbro, and magnetite. For the purposes of this disclosure, the carrier may be soluble or insoluble to some extent. The carrier material may have virtually any possible structural configuration, as long as the coupled molecules can bind to the target probe. Thus, the carrier configuration may be spherical, such as in beads; or cylindrical, such as on the inner surface of a test tube or the outer surface of a rod. Alternatively, the surface may be flat, such as a sheet or test strip. Therefore, some embodiments of the method include, for example, selectively separating source-labeled molecules and source-labeled nucleic acids from a pooled sample. In some embodiments, cells are lysed.

[0154] Target molecules include any molecule present in a sample or subsample (such as an individual compartment or a collection of compartments) with desired information about it, such as expression, activity, etc. In some embodiments, target molecules that can be labeled and characterized according to the disclosed methods include polypeptides (such as, but not limited to, proteins, antibodies, antigens, immunogens, protein complexes, and peptides), said polypeptides being modified in some cases, such as post-translational modifications, e.g., glycosylation, acetylation, amidation, formylation, γ-carboxyglutamate hydroxylation, methylation, phosphorylation, sulfation, or modification with pyrrolidone carboxylic acid. The methods disclosed herein are particularly suitable for screening libraries of antibodies. In certain embodiments, the target molecules are antibodies, such as antibody sets, screened for activity (e.g., specificity and / or affinity). In some embodiments, the antibody is an anti-HIV antibody, such as an anti-gp41 or gp120 antibody. The methods disclosed herein are also particularly suitable for screening libraries of potential antigens or immunogens, for example, to determine their potential efficacy in evoking an antibody response against a specific pathogen (e.g., a neutralizing antibody response). In certain embodiments, the target molecule is an antigen or immunogen that is screened for activity (e.g., specificity and / or affinity), such as a collection of antigens or immunogens. In some embodiments, the antigen or immunogen is an HIV antigen or immunogen, such as gp41 or gp120 antigens or immunogens, such as immunogenic fragments of gp41 and / or gp120. Target molecules also include nucleic acids (such as DNA and RNA, e.g., mRNA and cDNA), carbohydrates, lipids, small molecules (such as potential or realized therapeutic agents), compounds, and inorganic compounds, as well as conjugates and complexes of the exemplary target molecule types described herein. Target molecules may be natural, recombinant, or synthetic. As disclosed herein, a given compartment may include one or more different target molecules, meaning that several targets may be present in the compartment. In some cases, the compartment may include a target polypeptide molecule and a target nucleic acid molecule encoding the target polypeptide molecule, making it easy to determine the nucleic acid sequence encoding the target polypeptide. Thus, in some cases, the polypeptide target molecule and the nucleic acid target molecule are produced by the same cell. In certain cases, the peptide target molecule and the target nucleic acid are labeled with the same source-specific barcode or a matching barcode. In some instances, the target molecule is not encoded by the target nucleic acid. The target molecule may be expressed in cells or extracts (such as cell-free extracts). In some embodiments, the target molecule comprises peptides, nucleic acids, polysaccharides, and / or small molecules. In certain embodiments, the peptide comprises an antibody, an antigen, or a fragment thereof. In certain instances, the target molecule represents a library of randomly mutated peptides. In certain embodiments, the target molecule is expressed on a cell surface, such as a cell surface protein, or a fragment thereof, such as a cell surface domain of a protein.

[0155] In some instances, the target molecule (optionally in associated form) and the target nucleic acid are produced by individual cells within a compartment. Thus, in some instances, the target molecule is a polypeptide and the target nucleic acid encodes the target molecule in its respective individual compartment. In some instances, the cell is a B cell and the target molecule is an antibody, and the target nucleic acid encodes the antibody. In some embodiments, the target molecule is present on the cell surface. Thus, in some embodiments, the target molecule (such as a protein or polypeptide) is a target molecule typically found on the cell surface. In other embodiments, the target molecule (such as a protein or polypeptide) is not typically found on the cell surface, but is expressed on the cell surface, for example, through recombinant means. In some cases, the target molecule is typically found intracellularly, such as in the cytoplasm or in organelles. In these embodiments, it may be necessary to lyse the cell in which the target molecule is produced for labeling. Protein or nucleic acid target molecules may be naturally produced by the cell or may be recombinantly produced, for example, based on the presence of synthetic constructs in the cell.

[0156] In some embodiments, the target molecule is a protein or peptide, or a fragment or variant thereof, that is available in a protein or peptide database (e.g., SWISS-PROT, TrEMBL, SBASE, PFAM, or other databases known in the art). The target molecule may be a protein or peptide, or a fragment or variant thereof, that can be derived (e.g., through transcription and / or translation) from nucleic acid sequences known in the art, such as nucleic acid sequences available in nucleic acid databases (e.g., GenBank, TIGR, or other databases known in the art).

[0157] Target molecules such as peptides may optionally be produced by one or more synthetic multigene constructs, which may be present in cells or in compartments in the presence of cell-free extracts as described herein. Such target molecules may include, for example, variants of antibody sequences (e.g., CDR sequences) and / or various combinations of antibody light and heavy chains.

[0158] The target molecules can be proteins or peptides that are endogenous to the organism, such as proteins or peptides selectively expressed or displayed by one or more cells of the organism. For example, the protein or peptide can be a cell surface marker expressed by one or more cells. The organism can be, for example, a eukaryote (e.g., a mammal, such as a human), a virus, bacteria, or a fungus. In some embodiments, multiple different target molecules are selected from a single organism. In alternative embodiments, multiple different target molecules are selected from multiple different organisms.

[0159] In some embodiments of the disclosed method, the cells are fixed. Methods for fixing cells are well known in the art and include, for example, fixation with acetaldehyde. In some embodiments, the cells are not fixed.

[0160] In various embodiments, target molecules are displayed or expressed on the cell surface. Thus, in the case of proteins, target molecules can be, for example, antibodies, cell surface receptors, signal transduction proteins, transport proteins, cell adhesion proteins, enzymes, or fragments thereof. Cell surface target molecules include known transmembrane proteins, i.e., proteins known to have one or more transmembrane domains, proteins previously identified as associating with the cell membrane, or proteins predicted to have one or more transmembrane domains by one or more domain prediction methods known in the art. Cell surface target molecules may also be fragments of such proteins. Cell surface target molecules may be intrinsic membrane proteins or peripheral membrane proteins, and / or may be proteins or peptides having sequences that are present in nature, substantially similar to sequences present in nature, engineered from or otherwise modified from sequences present in nature, or artificially generated, for example, by molecular biology techniques (e.g., fusion proteins).

[0161] Target molecules can also be nucleic acids, such as DNA or RNA molecules. For example, RNA molecules (or corresponding cDNA molecules) transcribed from a gene of interest can be labeled and subsequently isolated according to the methods of this disclosure, and sequenced using their barcode labels. Multiple such labeled nucleic acids can be labeled simultaneously, for example, each nucleic acid receiving a different barcode and / or one or more unique molecular identifiers and barcode-receiving adaptors. In some embodiments, nucleic acids are produced from cells, cell lysates, or cell-free extracts. In certain embodiments, nucleic acids are produced by microorganisms (such as prokaryotic cells). In some embodiments, nucleic acid target molecules encode polypeptide target molecules in the same compartment.

[0162] In some embodiments, the target molecule may be associated with diseased cells or disease conditions. For example, the target molecule may be associated with cancer cells, such as proteins, peptides, or nucleic acids selectively expressed or not expressed by cancer cells; or it may specifically bind to such proteins or peptides (e.g., antibodies or fragments thereof, as described herein). In some cases, the target molecule is a tumor marker, such as a substance produced by a tumor or produced by non-cancer cells (e.g., stromal cells) in response to the presence of a tumor. Many tumor markers are not expressed solely by cancer cells, but are expressed at altered (i.e., elevated or decreased) levels in cancer cells or at altered (i.e., elevated or decreased) levels in non-cancer cells in response to the presence of a tumor. In some embodiments, the target molecule may be a protein, peptide, or nucleic acid expressed in response to any disease or condition known in the art.

[0163] In some embodiments, the sample comprises one or more synthetic gene constructs containing one or more polypeptide coding sequences linked to a promoter, such as a collection of synthetic gene constructs that optionally include portions generated by combination.

[0164] As disclosed herein, a compartment (such as a discrete volume or space) means any type of region or volume, which can be defined as a region or volume in which barcoded molecules (such as labeled target molecules or labeled nucleic acids) cannot freely escape or move. Compartments include droplets, such as those from water-in-oil emulsions or droplets deposited on surfaces, such as microfluidic droplets deposited, for example, on glass slides. Other types of compartments include, but are not limited to, tubes, wells, plates, pipettes, pipette tips, and bottles. Other types of compartments include, for example, “practical” containers defined by areas exposed to light, diffusion limits, or electromagnetic means. Such compartments can also exist in the form of: diffusion-defined volumes or spaces that effectively define space due to diffusion limitations so that only certain molecules or reactants can enter, such as chemically defined volumes or spaces where only certain target molecules can be present due to their chemical or molecular properties (such as size); or electromagnetically defined volumes or spaces where certain regions within the space can be defined using the electromagnetic properties (such as charge or magnetic properties) of the target molecule or its carrier. Such discrete volumes can also be optically defined volumes or spaces, which can be defined by irradiating them with visible, ultraviolet, infrared, or other wavelengths of light so that only target molecules within the defined space can be labeled. These compartments can be made of, for example, plastic, metal, composite materials, and / or glass. Such compartments are suitable for placement in centrifuges (e.g., microcentrifuges, ultracentrifuges, benchtop centrifuges, refrigerated centrifuges, or clinical centrifuges). The discrete volumes can exist as a single entity or as part of an array of such discrete volumes, for example, in the form of strips, microplates, or microtiter plates. The volume of the compartment can be, for example, from at least about 1 femtoliter (fl) to about 1000 ml, such as about 1 fl, 10 fl, 100 fl, 250 fl, 500 fl, 750 fl, 1 picoliter (pl), 10 pl, 100 pl, 250 pl, 500 pl, 750 pl, 1 nl, 10 nl, 100 nl, 250 nl, 500 nl, 750 nl, 1 μl, 5 μl, 10 μl, 20 μl, 25 μl, 50 μl, 100 μl, 200μl, 250μl, 500μl, 750μl, 1ml, 1.25ml, 1.5ml, 2ml, 2.5ml, 5ml, 10ml, 15ml, 20ml, 25ml, 50ml, 100ml, 150m l, 200ml, 250ml, 300ml, 350ml, 400ml, 450ml, 500ml, 550ml, 600ml, 650ml, 700ml, 750ml, 800ml, 900ml or 1000ml.

[0165] In some embodiments, the compartments are droplets, such as droplets in an emulsion and / or microfluidic droplets. Emulsification can be used in the methods of this disclosure to separate or partition a sample or sample collection into a series of compartments, such as compartments having discrete portions of single cells or non-cell samples (such as cell-free extracts or cell-free transcripts and / or cell-free translation mixtures). Typically, when used in conjunction with the methods and compositions disclosed herein, an emulsion will comprise a plurality of droplets, each droplet comprising one or more target molecules and / or target nucleic acids and a source-specific barcode, such that each droplet includes a unique barcode that distinguishes it from other droplets. Emulsification can be used in the methods of this disclosure to compartmentalize one or more target molecules within emulsion droplets having one or more nucleic acid barcodes (such as source-specific barcodes). Emulsions as disclosed herein will typically comprise a plurality of droplets, each droplet comprising one or more target molecules, target nucleic acids, and one or more nucleic acid barcodes, such as source-specific barcodes. Droplets in an emulsion can be sorted and / or separated according to methods well known in the art. For example, a conventional fluorescence-activated cell sorting (FACS) machine can be used at a rate of >10 cells per second. 4The rate of droplet analysis and / or sorting of dual emulsion droplets containing fluorescent signals has been used to improve the activity of enzymes produced by in vitro translation of single cells or single genes (Aharoni et al., Chem Biol 12(12):1281-1289, 2005; Mastrobattista et al., Chem Biol 2(12):1291-1300, 2005). However, emulsions are highly polydisperse, which limits quantitative analysis and makes it difficult to add new reagents to pre-formed droplets (Griffiths et al., Trends Biotechnol 24(9):395-402, 2006). However, these limitations can be overcome by using schemes based on droplet-based microfluidic systems (see, for example, Teh et al., Lab on a chip 8(2):198-220, 2008; Theberge et al., Angew Chem Int Ed Engl 49(34):5846-5868, 2010; and Guo et al., Lab on a chip 12(12):2146, 2012), in which highly monodisperse droplets with picoliter volumes can be prepared (Anna et al., Appl Phys Lett 82(3):364-366, 2003), fused (Song et al., Angew Chem Int Edit 42(7):767-772, 2003; Chabert et al., Electrophoresis 26(19):3706-3715, 2005), or separated (Song et al., Angew Chem Int Edit 42(7):767-772, 2003; Chabert et al., Electrophoresis 26(19):3706-3715, 2005). 42(7):767-772,2003; Link et al., Phys Rev Lett 92(5):054503,2004), incubated (Song et al., Angew Chem Int Edit 42(7):767-772,2003; Frenz et al., Lab on a chip 9(10):1344-1348,2009) and sorted, triggered by fluorescence at kHz frequency (Baret et al., Lab on a chip 9(13):1850-1858,2009), such as those described in Mazutis et al. (Nat. Protoc. 8(5):870-891,2013), which are incorporated herein by reference. As disclosed herein, emulsions may include various compounds, enzymes or reagents in addition to target molecules, target nucleic acids and source-specific barcodes. These additives may be included in the emulsion solution prior to emulsification. Alternatively, the additive can be added to individual droplets after emulsification.

[0166] Emulsions can be achieved by a variety of methods known in the art (see, for example, US2006 / 0078888 A1, paragraphs

[0139] -

[0143] of which are incorporated herein by reference). In some embodiments, the emulsion is stable to denaturation temperatures, for example, up to 95°C or higher. An exemplary emulsion is a water-in-oil emulsion. In some embodiments, the continuous phase of the emulsion comprises a fluorinated oil. The emulsion may contain a surfactant or emulsifier (e.g., a detergent, anionic surfactant, cationic surfactant, or amphoteric surfactant) to stabilize the emulsion. Other oil / surfactant mixtures, such as silicone oils, may also be used in certain embodiments. The emulsion may be contained in one or more wells (such as a plate) for ease of handling. In some instances, one or more target molecules, target nucleic acids, and nucleic acid barcodes are compartmentalized. The emulsion may be a monodisperse or polydisperse emulsion. Each droplet in the emulsion may contain, or on average contain, 0-1,000 or more target molecules. For example, a given emulsion droplet may contain 0, 10, 20, 30, 40, 50, 100, 200, 300, 400, 500, or more target molecules. In a particular embodiment, a given droplet may contain 0, 1, 2, or 3 cells capable of expressing or secreting the target molecule (e.g., a clonal population of the target molecule). On average, when rounded to the nearest integer, the droplets of the emulsions of this disclosure may contain 0-3 cells capable of expressing or secreting the target molecule, such as 0, 1, 2, or 3 cells capable of expressing or secreting the target molecule. In some embodiments, the average number of cells capable of expressing or secreting the target molecule in each emulsion droplet will be 1, between 0 and 1, or between 1 and 2. In other embodiments, the droplets may contain non-cellular systems, such as cell-free extracts.

[0167] In some embodiments, the target molecule, target nucleic acid, and nucleic acid barcode can be compartmentalized into the well due to physical limitations related to the mass or size of the target molecule and nucleic acid barcode, the size of the well, or a combination thereof. The well can be an optical fiber panel in which a central core is etched with an acid (such as an acid to which the core coating is resistant). The well can be a molded well. The wells can be covered to prevent communication between wells, such that beads present in a particular well are retained within the well or prevented from moving to different wells. The covering can be a solid sheet or physical barrier, such as a neoprene gasket; or a liquid barrier, such as a fluorinated oil. Methods suitable for this disclosure are known in the art (e.g., Shukla et al., J. Drug Targeting 13:7-18, 2005; Koster et al., Lab on a Chip 8:1110-1115, 2008).

[0168] In some implementations, a single cell or a portion of a non-cellular system from a sample is encapsulated together with beads (such as hydrogel beads), the beads including a source-specific barcode reversibly coupled thereto. Figure 1 The diagram illustrates this type of packaging. (Reference) Figure 1 For example, PDMS chips can be used to form an ensemble of uniformly sized hydrogel beads (such as PEG-DA beads). In some embodiments, uniformly sized PEG-DA hydrogel beads are copolymerized with universal capture oligonucleotides, which can be used to construct nucleic acid recognition sequences unique to each bead. Unique nucleic acid barcodes can be added to each bead using automated techniques and separate-aggregate labeling (see, for example, International Patent Publication No. WO2014 / 047561, which is specifically incorporated herein by reference). Using microfluidics, individual beads can be placed into a single droplet and then a single cell can be added, such that each droplet in the emulsion contains a single cell and a single hydrogel containing a unique, source-specific barcode. Figure 1 As shown, this system can be used to label all amplicones derived from a cell with unique barcodes. If the emulsion is then disrupted, the result is a pooled sample of amplicones labeled with barcodes on microdroplets. All of these amplicones can be traced back to the single cell from which they originated. Figure 2 As illustrated, in some embodiments, the beads include exemplary beads and a source-specific barcode for labeling the target nucleic acid. In a particular embodiment, the source-specific barcode is delivered to the compartment by delivering individual beads to each compartment, wherein each bead carries multiple copies of a single source-specific barcode sequence.

[0169] In some embodiments of the method, cells are contacted with one or more test agents, such as small molecules, nucleic acids, peptides, or polysaccharides. In specific instances, the peptides contain antibodies or antibody fragments. In some embodiments, the test agents are also labeled with source-specific barcodes. In some embodiments, individual test agents are labeled with test agent-specific barcodes.

[0170] In some embodiments, the method further includes amplifying one or more source-specific barcodes, target molecule-specific binding agent barcodes, test agent-specific barcodes, target nucleic acid barcodes, and target molecule barcodes.

[0171] In some embodiments, the method further includes detecting one or more of a target molecule-specific binding barcode, a test drug-specific barcode, a target nucleic acid barcode, and a target molecule barcode, for example, using hybridization, sequencing, or a combination thereof to detect the sequence of the source-specific barcode, the target molecule-specific binding barcode, the test drug-specific barcode, the target nucleic acid barcode, and / or the target molecule barcode.

[0172] In some embodiments, the method further includes quantifying one or more of the following: source-specific barcodes, target molecule-specific binding agent barcodes, test agent-specific barcodes, target nucleic acid barcodes, and target molecule barcodes.

[0173] This disclosure relates to methods for determining the specificity and / or affinity of a test agent for one or more target molecules. The disclosed methods include contacting cells expressing the target molecule with an aggregate of the test agent labeled with a test agent-specific barcode prior to separation. The target molecule bound to the test agent is separated, and the sequences of the test agent-specific barcode and the source-specific barcode are determined, thereby identifying the test agent bound to the target molecule. In some embodiments, the target molecule comprises a cell surface protein, and the target nucleic acid encodes the cell surface protein. In some instances of the method, unbound test agent is washed from the cells. In some embodiments, the target nucleic acid is cDNA or mRNA. Some embodiments of the method further include labeling the test agent with a source-specific barcode.

[0174] Examples of test agents include small molecule compounds, nucleic acids, peptides (such as proteins, antibodies, antigens, and / or immunogens), or polysaccharides. In some embodiments, screening of test agents involves testing combinatorial libraries containing a large number of potential modulatory compounds. Combinatorial chemical libraries can be collections of various compounds produced by combining many chemical “building blocks” (such as reagents) through chemical or biosynthesis. For example, linear combinatorial chemical libraries, such as peptide libraries, are formed by combining a collection of chemical building blocks (amino acids) in every possible way for a given compound length (e.g., the number of amino acids in a peptide compound). Millions of compounds can be synthesized through such combinations of chemical building blocks.

[0175] Libraries may contain suitable reagents, such as synthetic or natural compounds in combinatorial libraries. Many libraries are commercially available or readily produced; methods for the random and direct synthesis of a wide variety of organic compounds and biomolecules, including the randomized expression of oligonucleotides (such as antisense oligonucleotides and oligopeptides), are also known. Alternatively, libraries of natural compounds in the form of bacterial, fungal, plant, and animal extracts are available or readily produced. Furthermore, naturally or synthetically produced libraries and compounds can be readily modified using conventional chemical, physical, and biochemical methods and can be used to generate combinatorial libraries. Such libraries are suitable for screening large numbers of different compounds.

[0176] Compounds identified using the methods disclosed herein may serve as conventional “lead compounds” or may be used on their own as potential or actual therapeutic agents. In some cases, assemblies of candidate agents may be identified and further screened to determine which individual agents or which agent sub-assemblies within the assemblies possess the desired activity.

[0177] In other embodiments, determining the affinity and / or specificity of the test agent for the target molecule includes: contacting a labeled target molecule with a test agent bound to a detectable label; isolating the labeled target molecule bound to the test agent using the detectable label; determining the sequence of a source-specific barcode on the isolated target molecule; and quantifying the source-specific barcode associated with the isolated target molecule, thereby determining the affinity of the test agent for the target molecule. In some instances, the source-specific barcode not isolated with the test agent is quantified to normalize concentration. In some embodiments, the source-specific barcode is cleaved after isolation. In some embodiments, the isolated barcode is associated with a target nucleic acid to determine the sequence of the target molecule bound to the test agent. In some instances, the target molecule comprises a peptide, such as an antibody, antigen, and / or immunogen. In specific instances, the antibody comprises an anti-HIV antibody, such as an anti-gp41 or anti-gp120 antibody, and the test agent comprises a collection of potential HIV immunogens. In certain instances, the test agent comprises an antibody, and the target molecule comprises a protein expressed on the cell surface; for example, the test agent comprises an anti-HIV antibody, such as an anti-gp41 or gp120 antibody, and the target molecule comprises a collection of HIV immunogens. In some embodiments, the method includes determining one or more of the following: a dissociation constant, an association rate, or a dissociation rate of the target molecule in relation to the test agent.

[0178] Aspects of the disclosed method also relate to determining the expression of target molecules on the surface of cells or cell collections. The method includes: assigning a set of target molecules to target nucleic acids in a sample or sample collection; contacting the sample cells to be separated with a collection of test agents, each individually labeled with a unique test agent barcode; and determining the sequence of the source-specific barcode, thereby determining the expression of the molecules on the surface of the cell collection. Some examples include associating the separated barcode with the target nucleic acid to determine the sequence of the target molecule bound to the test agent. In a particular embodiment, the target molecule comprises a polypeptide, such as an antibody, that is specific to a cell surface marker. In some embodiments, the molecules expressed on the cell surface are correlated with other measures of cell type, cell cycle, or cell state.

[0179] This disclosure further relates to a method for identifying proteins with a specific activity of interest from a cell population. The method includes: assigning a set of target molecules to a sample or a sample set containing target nucleic acids; isolating target molecules with the specific activity of interest; and identifying a source-specific barcode of the isolated target molecules with the specific activity of interest. In some embodiments, the method further includes identifying target nucleic acid molecules encoding proteins by matching the sequence of the source-specific nucleic acid barcode. In some instances, the activity is antigen binding, and the protein is an antibody. In some embodiments, the antibody of interest is prepared by expressing a nucleic acid molecule identified as encoding an antibody of interest.

[0180] This disclosure relates to methods for determining the levels of post-translational modification target molecules (such as target proteins) in cells or cell samples (such as individual cells). Determination of any post-translational modification is possible, such as glycosylation, acetylation, amidation, formylation, γ-carboxyglutamate hydroxylation, methylation, phosphorylation, sulfation, or modification with pyrrolidone carboxylic acid. In the disclosed method, a sample (such as a sample of cells of interest) is contacted with a first specific binder and a second specific binder, the first specific binder specifically binding the post-translational modified target molecule at the modification site, and the second specific binder specifically binding the non-post-translational modified target molecule at the modification site, wherein the first and second specific bindings are marked with specific binder-specific barcodes. Individual cells are separated into individual compartments, wherein each individual compartment further includes a source-specific barcode containing a unique nucleic acid recognition sequence that maintains or carries information about the source compartment of the separated cells in the sample. In some instances, cells are lysed to release the cell contents, allowing the specific binders to interact and thus bind the target molecules in the compartments. First and second specific binders in individual compartments are labeled with source-specific barcodes present in the individual compartments, and the target molecule is isolated, thereby isolating the first and second specific binders bound to the target molecule. Sequences of the source-specific barcode and the specific binder barcode are present on the isolated specific binders, thereby determining the presence of post-translational modified proteins in the sample. In some embodiments, the method includes quantifying the levels of post-translational modified and unmodified target proteins, for example, determining the ratio of modified to unmodified target molecules. In some instances, isolation includes contacting the target molecule with a third specific binder that does not bind to modification sites, wherein the third specific binder is detectably labeled. In some embodiments, the method further includes labeling the target nucleic acid in the individual compartment with a source-specific barcode present in the individual compartment to form a source-labeled target nucleic acid, wherein the source-labeled target nucleic acid from each individual compartment contains a unique indexed nucleic acid recognition sequence that is the same as or matches the source-labeled target molecule, and then detecting the nucleotide sequence of the source-specific barcode, thereby assigning the target protein set to the target nucleic acid in the sample or sample set, while maintaining information about the source compartment of the target protein and the target nucleic acid.

[0181] C. Exemplary Applications

[0182] 1. Genotype-phenotype coupling

[0183] This disclosure provides a method for co-detecting an expressed target molecule (e.g., a peptide, protein, protein complex, or any other gene product) and the nucleic acid (e.g., RNA or DNA) encoding the expressed target molecule. In some embodiments, each of a plurality of discrete volumes contains one or more cells permitted to express the target peptide molecule. After expression and optional cell lysis, the expressed peptide is barcoded with a source-specific nucleic acid by maintaining the discrete volumes under conditions that allow for barcoding. The nucleic acid molecule encoding the target peptide molecule (e.g., RNA or cDNA reverse transcribed therefrom) is also barcoded with a source-specific nucleic acid.

[0184] Source-specific nucleic acid barcodes used to label target polypeptide molecules and corresponding nucleic acid molecules within a specific discrete volume are typically matched nucleic acid barcodes. Matched nucleic acid barcodes may be, for example, at least 80% identical (e.g., 80%, 85%, 90%, 95%, 99%, or 100% identical). In some embodiments, matched nucleic acid barcodes are 100% identical. In other embodiments, matched barcodes may have different sequences but can be identified as members of a matching pair based on their sequences (e.g., by pre-defining the introduction of two specific barcodes into the same container such that molecules labeled with these two specific barcodes must originate from the same container). Optionally, source-specific nucleic acid barcodes may be introduced into discrete volumes incorporated into a single solid or semi-solid carrier (such as beads as disclosed herein). In such cases, a source-specific nucleic acid barcode intended for binding to a target polypeptide molecule may be included within a tag element comprising an affinity moiety (e.g., an antibody) specific to the target polypeptide molecule, while a source-specific nucleic acid barcode intended for binding to a corresponding nucleic acid molecule may be included within a tag element comprising an affinity moiety (e.g., a nucleic acid molecule) specific to the corresponding nucleic acid molecule.

[0185] After being labeled with source-specific nucleic acid barcodes, target polypeptide molecules and / or corresponding nucleic acids can be combined to form an aggregate, and the relevant barcodes (and optional target nucleic acid molecules) can be amplified. The source discrete volume of a given target polypeptide molecule or nucleic acid in the aggregate can be determined by recognizing the sequence of the relevant source-specific nucleic acid barcode.

[0186] In certain variations of these methods, specific portions of the aggregate may optionally be separated before the sequence is identified. For example, when the target polypeptide molecule is an antibody or a fragment thereof, chromatographic analysis may optionally be performed on a column comprising a fixed antigen binding to the antibody under specific or varying stringent conditions to allow the separation of antibodies, for example, those with particularly high binding affinity. In another example, when the target polypeptide molecule is an antigen containing an antigen or immunogen, chromatographic analysis may be performed using a column comprising a fixed antibody (or an antigen-binding fragment thereof). In other examples, activity or properties other than binding affinity may be evaluated. In any case, a specific target polypeptide molecule with desired characteristics may be separated, and the identity or other characteristics (e.g., activity and / or affinity) of the target polypeptide molecule may be determined by sequencing source-specific nucleic acid barcodes. The genotype corresponding to the phenotype of the selected target polypeptide molecule may be determined by using sequencing to identify matching source-specific nucleic acid barcodes in the aggregated nucleic acid sample, and then optionally by sequencing the nucleic acids linked to these barcodes.

[0187] In other variations of these methods, the separation steps described above are not required. More specifically, peptides or polypeptides containing epitopes of interest are labeled with epitope-specific nucleic acid barcodes. These peptides or polypeptides are mixed with cells that express antibodies on their surfaces (e.g., B cells such as those obtained through in vitro immunization or from an immune donor), hybridomas, or other cells expressing recombinant antibodies). Unbound peptides or polypeptides are washed from the cells and then separated into discrete volumes (e.g., one cell per discrete volume). Source-specific nucleic acid barcodes can then be used to label the antibody produced by the cell, the nucleic acid encoding the antibody, and the peptide or polypeptide (or epitope-specific nucleic acid barcode). Multiplex sequencing of the labeled molecules aggregated from multiple discrete volumes can be used to associate cells producing antibodies that bind specific epitopes with the corresponding coding sequences.

[0188] The methods also include methods for preparing antibodies using nucleic acid molecules identified by the methods of this disclosure as encoding antibodies of interest. In these methods, the identified nucleic acid molecules are expressed in cells using standard methods in the art.

[0189] 2. Proteomics

[0190] The methods disclosed herein can be used in proteomics applications to, for example, assess the expression levels and / or functions of various proteins expressed in cells or cell-free systems. Proteins may optionally be encoded by synthetic constructs (e.g., multi-gene constructs). In various instances, proteins are components of metabolic pathways, and proteomics analyses can be performed, for example, to assemble and optimize novel pathways and / or recognize rate-limiting steps. Regarding the latter, based on the analysis results, the expression of rate-limiting components of the pathway can be altered to optimize the pathway. These methods can be used, for example, in the context of optimizing metabolically engineered microorganisms in synthetic biology.

[0191] In other applications, the proteomics methods disclosed herein can be used for the identification and validation of biomarkers for diseases such as cancer, and these methods can optionally be used to assess proteomic changes in cells exposed to different conditions. Different proteins analyzed using the methods of this disclosure may be sequence variants of different pathway components and / or base sequences, in which case the methods can be used to identify variants with specific expression levels, stability, or other functional characteristics. The changes assessed may range from individual amino acid substitutions to domain or subunit substitutions or exchanges, and may include many combinations thereof. These changes may be random or may be generated, for example, by mixing and matching sequences derived from, for example, different species.

[0192] 3. Cell surface marker analysis

[0193] This disclosure provides a method for large-scale multiplexing of cell surface markers (e.g., cell surface proteins) using nucleic acid barcoding. The target molecule may be a cell surface marker associated with a barcode-receiving adaptor. For example, a cell surface marker may be linked to an oligonucleotide barcode-receiving adaptor, which includes protrusions capable of receiving a portion of a barcode. The barcode may, for example, include corresponding protrusions capable of hybridizing to the protrusions on the barcode-receiving adaptor. Thus, multiple cell surface markers can be labeled with barcodes. The barcode-receiving adaptor may also act as another identifier. For example, cells in a discrete volume may express a cell surface marker, which is then associated with a barcode-receiving adaptor. One or more barcodes can then be linked to individual barcode-receiving adaptors. For example, all barcode-receiving adaptors in a discrete volume and / or associated with a cell surface marker expressed by a particular cell may receive the same barcode. In some embodiments, the individual barcode-receiving adaptors are different for the individual cell surface marker. In an alternative implementation, all cell surface markers expressed by a particular cell receive the same barcode-receiving connective.

[0194] Cell surface markers and nucleic acids (e.g., mRNA) encoding cell surface markers can be associated with the same barcode-receiving adaptor and / or barcode, for example, according to the genotype-phenotype coupling method described herein. In some embodiments, if one or more barcodes associated with a cell are detected, but mRNA corresponding to the cell surface marker is not also detected, then cells that do not express the specific cell surface marker can be identified.

[0195] In examples of cell surface marker-related applications, multiple cells (e.g., B cells) each expressing different cell surface-specific binding agents (e.g., cell surface proteins, such as antibodies) are mixed with multiple target epitopes (e.g., epitopes of HIV proteins, such as Gag, Pol, Env, or Nef, gp120, gp41) labeled with oligonucleotide barcodes, for example, identifying each unique type of epitope. Optionally, a unique molecular identifier is associated with, for example, an epitope-specific nucleic acid barcode. Thus, each epitope-specific nucleic acid barcode of HIV epitope #1 would, for example, be associated with a different unique molecular identifier for quantification. The oligonucleotide barcode may further include, for example, a unique molecular identifier. Binding of the target epitope to the cell surface binding portion can be allowed, followed by washing away excess epitopes. Each cell, along with its expressed cell surface binding portion and the bound epitope, can then be encapsulated in a discrete volume (e.g., emulsion droplets). The cells can be dissolved, for example, in the discrete volume to release their internal contents. The discrete volume may further include a source-specific barcode, which can be used to label the cell surface binding moiety and / or target epitope. Thus, the cell surface binding moiety and / or target epitope may be the target molecules of this disclosure. In some cases, the source-specific barcode may also label, for example, a nucleic acid encoding the cell surface binding moiety (e.g., an mRNA or cDNA molecule encoding the cell surface binding moiety, such as RNA encoding the heavy and / or light chains of antibodies produced by B cells), thereby labeling the cell surface binding moiety and the nucleic acid encoding the cell surface binding moiety with the same barcode. In some cases, the target epitope is further labeled with the same source-specific barcode (e.g., by linking the source-specific barcode to an oligonucleotide barcode of the labeled epitope). In some cases, the cell surface binding moiety, the target epitope, and / or the nucleic acid encoding the cell surface binding moiety are labeled with the same or different barcodes (e.g., by linking to the same vector, pre-determined barcodes associated with each other). According to the methods of this disclosure, barcodes labeling cell surface binding sites, target epitopes, and / or nucleic acids encoding cell surface binding sites can be collected, isolated, amplified, and / or sequenced. By sequencing both source-specific and epitope-specific nucleic acid barcodes, antibodies binding to specific epitopes of interest can be identified. Furthermore, sequencing RNA associated with source-specific nucleic acid barcodes enables the identification of antibody nucleic acid sequences that can be used to prepare and / or further characterize antibodies of interest.

[0196] 4. Affinity Analysis

[0197] The methods disclosed herein can be used to determine the binding affinity between a target molecule and another molecule (e.g., a specific binder). For example, methods known in the art can be used to measure the equilibrium constant (Kd) and dissociation rate (k) of the binding interaction between the target molecule and another molecule (e.g., a specific binder).解离 For example, if a specific binding agent is attached to a solid support (e.g., a column, chip, surface, or bead), Kd can be measured by titrating the target conjugated to a barcode in various amounts. After incubation, the solid support can be washed, and the barcode can be cleaved and sequenced to determine the amount of target bound to the specific binding agent. This analysis can be performed in reverse, with the target molecule bound to the solid support and the specific binding agent conjugated to the barcode. Alternative methods involve a “sandwich” form of batch affinity purification or analysis, in which the target molecule is complexed with a specific binding agent (e.g., a specific binding agent labeled with a barcode such as a source-specific barcode) and also with another specific binding agent bound to the solid support. In this method, instead of directly binding the target molecule to the source-specific barcode, it is instead labeled by binding to a specific binding agent labeled with a source-specific barcode.

[0198] These analyses can also be performed in a competitive manner by adding a molecule that competes with the target molecule for binding to the specific binder. The competitor can be added simultaneously with the target molecule or after the target molecule has been added and subsequently complexed with the specific binder. Optionally, the competitor can be conjugated to a barcode. For example, the competitor can be the target molecule, such as a target molecule labeled with a source-specific barcode. In some cases, k can be determined by adding a non-barcoded competing molecule that can compete with the target molecule for binding to the specific binder. 解离 It can isolate and retain target molecules bound to a vector, and can sequence the barcode associated with the target molecule to determine the amount of target molecule retained bound to the vector.

[0199] An exemplary method for measuring binding affinity includes binding a labeled target-molecule-specific binding agent complex to a carrier (e.g., via a biotin moiety attached to a component of the complex). The carrier can then be exposed to one or more washing conditions. Barcodes associated with the target molecule that are removed in each wash can be sequenced separately to determine the abundance of the unbound fraction. The barcodes can also be cleaved from the complex that remains bound to the carrier after washing to determine the abundance of the binding fraction. In some embodiments, the binding fraction is allowed to remain bound to the carrier for approximately one to two days. The relative abundance of the unbound fraction and the bound fraction can be compared to determine the dissociation rate. Multiple washes can be performed in this manner to generate a Kd curve. In an alternative embodiment, a specific binding agent can be used to separate fractions of multiple target molecules (e.g., by immunoprecipitation using an antibody as a specific binding agent). This binding fraction or a portion thereof can then be washed from the specific binding agent. Barcodes associated with the target molecule can be separated from the bound and unbound fractions. The relative abundance of the target molecule in the bound and unbound fractions can then be calculated to determine an affinity measurement. Multiple washes can be performed, with each subsequent wash removing additional portions of the binding fraction, and the relevant nucleic acid barcodes are sequenced to generate affinity curves. In one embodiment, the target molecule is allowed to bind to a specific binding agent for approximately one to two days prior to washing.

[0200] 5. Reporter cells

[0201] The methods disclosed herein can be used to analyze the properties of target molecules (e.g., target peptide molecules). For example, target molecules can induce responses in cells. In some cases, target molecules can be soluble signals (e.g., secreted proteins, peptides, small molecules, or other specific binders) capable of interacting with cells (e.g., by binding to cell surface receptors), thereby inducing downstream effects in the cell. Exemplary downstream effects include, but are not limited to, changes in gene expression (e.g., increases or decreases) (e.g., changes in mRNA expression levels), changes in intracellular signaling pathways, and / or activation or inhibition of cell activity (e.g., cell proliferation, cell growth, cell death, changes in cell morphology, and / or changes in cell viability).

[0202] Therefore, in some embodiments, the discrete volumes of the methods disclosed herein may include one or more reporter cells in which target molecules may induce such downstream effects. For example, the expression of one or more mRNAs by reporter cells may be altered (e.g., increased or decreased) through direct or indirect interaction with target molecules. Such mRNAs may, for example, encode target molecules, or may not encode target molecules. In some cases, the mRNA may encode a peptide tag. In some embodiments, the mRNA is labeled with a source-specific barcode (e.g., the same source-specific barcode used to label target molecules). For example, the mRNA or corresponding cDNA may be ligated to a source-specific barcode. The labeled mRNA or cDNA can then be obtained (e.g., by lysing reporter cells) and sequenced to determine how the target molecule alters the expression of mRNAs in reporter cells (e.g., by measuring the amount of each different mRNA in a particular discrete volume). In some embodiments, reporter cells are lysed, and source-specific barcode-labeled mRNAs from multiple discrete volumes are pooled (e.g., together with labeled target molecules or labeled tags), amplified, and sequenced. The labeled mRNA can be converted into labeled cDNA, for example, prior to collection, amplification, and / or sequencing. In a particular embodiment, the entire transcriptome of the reporter cell is labeled with source-specific barcodes and sequenced according to RNA-seq methods known in the art.

[0203] 6. Protein modification detection

[0204] The methods disclosed herein can be used to determine the modification state, such as phosphorylation state, of proteins in samples (such as cells and / or non-cell systems). Proteins may undergo post-translational modifications such as phosphorylation, for example, tyrosine kinases. In some embodiments, antibodies specific to different test proteins are labeled with different single-stranded DNA tags. Cells of interest are incubated in an appropriate culture medium, with or without a test agent (such as a pharmaceutical agent or potential pharmaceutical agent). Cells are separated into compartments along with DNA-tagged protein target-specific antibodies, and in some cases, biotin-tagged antibodies specific to target domains lacking phosphorylation sites, and hydrogel beads carrying a mixture of barcoded primers complementary to the DNA tags of the DNA-tagged target-specific antibodies.

[0205] Target protein-specific antibodies may be commercially available, pre-conjugated with biotin, or may be biotin-modified, for example, by reacting with an NHS-ester of a lysine residue, or by partially conjugating an acylhydrazine to an oxidized antibody carbohydrate residue. DNA-tagged target-specific antibodies are specific to the phosphorylated regions of the target protein. An antibody targeting the phosphorylated protein and an antibody targeting the unphosphorylated protein are used to encode specific tags for each state. In some instances, a third antibody containing a detectable tag, specific to the non-phosphorylated domain of the protein, is used to label the binding complex. This labeling allows subsequent separation of the antibody-target complex from the unbound antibody. Alternatively, the complex can be purified by size exclusion chromatography without the use of a third antibody.

[0206] The compartment is incubated with a lysis agent that will lyse and release barcoded oligonucleotides into the entire volume. Simultaneously, cells are lysed, and the released target proteins are captured by labeled antibodies that anneal to complementary sequences on single-stranded portions of the barcoded DNA bound to the hydrogel. A polymerase then elongates the barcoded DNA molecule, thereby copying the DNA tag. These antibody-target complexes are captured by biotin-labeled antibodies. Adding DNA barcoding to antibody-bound DNA tags confers single-cell specificity to these sequences.

[0207] In an alternative implementation, antibodies specific to different target proteins are labeled with different double-stranded DNA tags. Target protein phosphorylation data are quantified at the single-cell level, and the number of target-specific DNA tag reads via each DNA barcode is compared to the total amount of target protein. Single-cell level mRNA expression and sequence information can also be obtained from sequencing.

[0208] 7. Combinatorial Chemistry

[0209] This disclosure provides a method for coupling a target molecule to multiple nucleic acid barcodes, for example, in the form of a barcode multiplex. Each barcode can be associated with specific conditions, such as compounds (e.g., small molecule compounds, nucleic acids, peptides, or polysaccharides), temperature, incubation time, atmospheric conditions, and pH, so that the set of conditions in which the target molecule was exposed can be determined by sequencing the barcodes. In some embodiments, each barcode is added to a growing barcode multiplex while exposing the target molecule to the conditions associated with the barcode, such that sequencing the barcode multiplex reveals the order in which the target molecule was exposed to the conditions. The barcodes can be associated with affinity portions that recognize the target molecule and thus facilitate the coupling of the barcodes to the target molecule. In some embodiments, the target molecule is exposed to multiple compounds (e.g., small molecules, nucleic acids, peptides, polysaccharides, or combinations thereof), and each of the compounds is associated with a different nucleic acid barcode, which can then be coupled to the target molecule (e.g., by adding each barcode to a growing barcode multiplex). For example, a target molecule can be exposed to a set of reaction conditions, each characterized by a specific compound, wherein each exposure results in the addition of a different barcode to the target molecule. Therefore, the compounds exposed to the target molecule and their order of exposure can then be determined by sequencing the barcodes associated with the target molecule.

[0210] D. Compositions and kits

[0211] This disclosure also relates to compositions and kits that can be used to perform the methods of this disclosure. In one example, the disclosed composition includes a barcode-tagged complex. The barcode-tagged complex includes a solid or semi-solid substrate (such as beads, e.g., hydrogel beads) and a plurality of barcode elements reversibly coupled thereto, wherein each of the barcode elements comprises an index nucleic acid recognition sequence (such as RNA, DNA, or a combination thereof), and one or more of a nucleic acid capture sequence specifically binding to a target nucleic acid and a specific binder specifically binding to a target molecule. In some examples, the source-specific barcode further includes one or more cleavage sites. In some examples, at least one cleavage site is oriented such that cleavage at said site releases the source-specific barcode from the substrate coupled thereto. In some examples, at least one cleavage site is oriented such that cleavage at said site releases the source-specific barcode from the target molecule specific binder. In some examples, the source-specific barcode further includes one or more capture portions covalently or non-covalently linked. In some examples, one or more capture portions comprise biotin, such as biotin-16-UTP. In some instances, the source-specific barcode further includes one or more sequencing adaptors or universal initiation sites. In some instances, the target-molecule-specific binder includes an antibody or a fragment thereof, a polypeptide or peptide containing an epitope recognized by the target molecule, or a nucleic acid. In some instances, each of the source-specific barcodes includes one or more indexes, one or more sequences for gene-specific capture and / or amplification, and / or one or more sequences for sequencing library construction. In some instances, each of the barcodes includes four indexes that can be combined and assembled. In some instances, the sequences for sequencing library construction include an Illumina P7 (SEQ ID NO:11) sequence and / or Illumina sequencing primers. In some instances, the barcode includes primers for DNA synthesis, such as primers suitable for DNA synthesis on a DNA template or RNA template. Kits that may include any and all of the compositions disclosed herein are also disclosed.

[0212] The invention is further defined with reference to the following numbered entries:

[0213] 1. A method for assigning coupled phenotype-genotype to a set of target molecules associated with a specific compartment, the method comprising:

[0214] Provide samples, including cellular or non-cellular systems, containing target molecules of interest and / or nucleic acids encoding target molecules of interest;

[0215] A subset of cells, a single cell, or a portion of the non-cellular system from the sample is separated into individual compartments, wherein each individual compartment further includes a source-specific barcode containing a unique nucleic acid recognition sequence that maintains or carries information about the source compartment of the separated cellular or non-cellular system in the sample;

[0216] The target molecules in the individual compartments are labeled with the source-specific barcodes present in the individual compartments to form source-labeled target molecules, wherein the source-labeled target molecules from each individual compartment contain the same unique indexed nucleic acid recognition sequence or a matching index sequence, and optionally the target molecules are further labeled or separated according to their physical and chemical properties to provide additional property or abundance information also related to the source-specific barcodes;

[0217] Associating the properties of the target molecule with the individual compartment and / or the cells within the compartment.

[0218] The nucleotide sequence of the source-specific barcode is detected, thereby assigning the set of target molecules to a specific compartment and the characteristics of the target molecules in the specific compartment.

[0219] 2. The method of Item 1, further comprising: allocating the set of target molecules to a target nucleic acid in the sample or a set of samples, wherein the sample contains, further comprises, the target nucleic acid of interest; and the method further comprises:

[0220] The target nucleic acid in the individual compartment is labeled with the source-specific barcode present in the individual compartment to form a source-labeled target nucleic acid, wherein the source-labeled target nucleic acid from each individual compartment contains a unique indexed nucleic acid recognition sequence that is the same as or matches the source-labeled target molecule;

[0221] The nucleotide sequence of the source-specific barcode is detected, thereby assigning the set of target molecules to the sample or a set of samples and maintaining information about the source compartments of the target molecules and the target nucleic acids.

[0222] 3. The method as described in item 1 or 2, wherein the source-specific barcode comprises RNA, DNA, or a combination thereof.

[0223] 4. The method of any one of items 1 to 3, wherein the source-specific barcode is reversibly coupled to a solid or semi-solid substrate.

[0224] 5. The method of any one of entries 1 to 4, further comprising encapsulating the single cell or portion of the non-cellular system from the sample together with beads containing the source-specific barcode reversibly coupled thereto.

[0225] 6. The method of any one of items 1 to 5, wherein the source-specific barcode further comprises a nucleic acid capture sequence specifically bound to the target nucleic acid and / or a specific binder specifically bound to the target molecule.

[0226] 7. The method of Item 6, wherein the source-specific barcode comprises two or more groups of source-specific barcodes, wherein the first group comprises the nucleic acid capture sequence and the second group comprises the specific binder specifically binding to the target molecule.

[0227] 8. The method of any one of items 2 to 7, wherein the target nucleic acid comprises RNA or DNA.

[0228] 9. The method as described in item 8, wherein the target nucleic acid includes mRNA, genomic DNA, or cDNA.

[0229] 10. The method of any one of items 2 to 9, further comprising synthesizing cDNA from the target nucleic acid, wherein the cDNA comprises a nucleic acid sequence of the target nucleic acid or a fragment thereof and a sequence of the source-specific barcode.

[0230] 11. The method of Item 10, wherein the source-specific barcode is a primer used for the synthesis of the cDNA.

[0231] 12. The method as described in item 11, wherein the target nucleic acid or its complement encodes a polypeptide of interest.

[0232] 13. The method of any one of items 7 to 12, wherein the target molecule comprises a target polypeptide, and the specific binder of the target molecule specifically bound to the sample comprises a polypeptide-specific binder specifically bound to the target polypeptide.

[0233] 14. The method of Item 13, wherein the polypeptide-specific binder comprises an antibody or a fragment thereof and / or a protein-binding domain or a fragment thereof, or a nucleic acid sequence specifically bound to the target polypeptide or expressing a cell surface marker specifically bound to the target polypeptide.

[0234] 15. The method of any one of items 1 to 14, wherein the target molecule comprises target DNA, and the specific binder of the target molecule specifically binding to the sample comprises a nucleic acid sequence or DNA-binding domain that specifically binds to and / or hybridizes to the target DNA.

[0235] 16. The method of any one of items 1 to 15, wherein the source-specific barcode further comprises a sequencing adaptor.

[0236] 17. The method of any one of items 1 to 16, wherein the source-specific barcode further comprises a universal initiation site.

[0237] 18. The method of any one of items 1 to 17, further comprising collecting the individual compartments to form a pooled sample.

[0238] 19. The method of Item 18, further comprising selectively separating the source-labeled molecule and the source-labeled nucleic acid from the aggregated sample.

[0239] 20. The method of any one of items 1 to 19, wherein the source-specific barcode further comprises one or more capture portions covalently or non-covalently bonded.

[0240] 21. The method of Item 20, wherein separating the source-labeled molecule and the source-labeled nucleic acid comprises capturing the source-specific barcode via the one or more capture portions.

[0241] 22. The method of any one of items 20 to 21, wherein the one or more capture portions are captured with a capture portion-specific binder specifically bound to the one or more capture portions.

[0242] 23. The method of any one of items 20 to 22, wherein the one or more capture portions are captured on a solid carrier.

[0243] 24. The method of any one of items 22 to 23, wherein the capture portion-specific binder is attached to the solid carrier.

[0244] 25. The method of any one of entries 20 to 24, wherein the one or more capturing portions contain biotin.

[0245] 26. The method of any one of items 22 to 25, wherein the capture-part-specific binder comprises streptavidin.

[0246] 27. The method of any one of items 1 to 26, wherein the source-specific barcode comprises biotin-16-UTP.

[0247] 28. The method of any one of items 1 to 27, wherein the labeling of the target molecule comprises directly linking the source-specific barcode to the target molecule.

[0248] 29. The method of any one of items 1 to 28, wherein the labeling of the target molecule comprises indirectly linking the source-specific barcode to the target molecule.

[0249] 30. The method of any one of Item 29, wherein indirect connection comprises binding a target molecule-specific binder to the target molecule, wherein the target molecule-specific binder is indirectly or directly connected to the source-specific barcode.

[0250] 31. The method of Item 30, wherein the target molecule specific binder comprises an antibody or a fragment thereof, a polypeptide or peptide specifically bound to the target molecule, or a nucleic acid.

[0251] 32. The method of any one of items 1 to 31, wherein the source-specific barcode further comprises a primer-specific region.

[0252] 33. The method of any one of items 1 to 32, wherein each of the source-specific barcodes further comprises a unique molecular identifier.

[0253] 34. The method of any one of items 1 to 33, wherein each of the source-specific barcodes comprises one or more indexes, one or more sequences for achieving gene-specific capture and / or amplification, and / or one or more sequences for achieving sequencing library construction.

[0254] 35. The method of any one of items 1 to 34, wherein the target molecule is connected to a source-specific barcode receiver connector.

[0255] 36. The method of 35, wherein the source-specific barcode receiving adaptor comprises a nucleic acid.

[0256] 37. The method of 36, wherein the source-specific barcode receiving connective includes a protrusion, and the source-specific barcode includes a sequence capable of hybridizing to the protrusion.

[0257] 38. The method of any one of items 30 to 37, wherein the target molecule-specific binder specifically binds both the target molecule and the source-specific barcode.

[0258] 39. The method of any one of items 30 to 38, wherein the target molecule is incubated together with a target molecule-specific binder that binds both the target molecule and the source-specific barcode, and the target molecule-specific binder that is not bound to the target molecule and / or the source-specific barcode is removed before separation into the individual compartments.

[0259] 40. The method of any one of items 30 to 39, wherein the target molecule-specific binder comprises a target molecule-specific binder barcode, the target molecule-specific binder barcode encoding the identity of the target molecule-specific binder.

[0260] 41. The method of Item 40, wherein the nucleic acid containing the target molecule-specific binding barcode can bind to the nucleic acid containing the source-specific barcode via base pairing interactions.

[0261] 42. The method of any one of items 40 to 41, wherein the source-specific barcode is a primer of the complementary strand for synthesizing the target molecule-specific binding barcode.

[0262] 43. The method of any one of items 40 to 42, further comprising detecting the sequence of the target molecule-specific binding barcode.

[0263] 44. The method of any one of items 1 to 43, wherein the source-specific barcode is delivered to the individual compartment by delivering individual beads to the individual compartments, wherein each bead carries multiple copies of a single source-specific barcode.

[0264] 45. The method of any one of items 1 to 44, wherein the compartment comprises aqueous droplets in an emulsion.

[0265] 46. ​​The method of Item 45, wherein the emulsion comprises one or more surfactants, thereby stabilizing the emulsion.

[0266] 47. The method of any one of items 45 to 46, wherein the emulsion comprises a continuous phase, and the continuous phase of the emulsion comprises a fluorinated oil.

[0267] 48. The method of any one of items 46 to 47, wherein the one or more surfactants comprises one or more fluorinated surfactants.

[0268] 49. The method of any one of items 45 to 48, further comprising breaking down the emulsion, thereby collecting the contents of the individual compartments.

[0269] 50. The method of any one of items 1 to 49, wherein the target molecule comprises a polypeptide, nucleic acid, polysaccharide, and / or small molecule.

[0270] 51. The method of Item 50, wherein the polypeptide comprises an antibody, an antigen, or a fragment thereof.

[0271] 52. The method of any one of entries 1 to 51, wherein the target molecule represents a library of randomly or systematically mutated peptides.

[0272] 53. The method of any one of items 1 to 52, wherein the target molecule is expressed on the surface of the cell.

[0273] 54. The method of Item 53, wherein the target molecule comprises a cell surface protein or a fragment thereof, such as a cell surface domain of a protein.

[0274] 55. The method of any one of items 1 to 54, wherein the sample comprises one or more cells.

[0275] 56. The method of any one of items 1 to 55, wherein the sample comprises the non-cellular system of the target molecule and the target nucleic acid.

[0276] 57. The method as described in Item 56, comprising cell-free extracts or cell-free transcripts and / or cell-free translation mixtures.

[0277] 58. The method of any one of items 1 to 57, wherein the sample comprises one or more synthetic gene constructs, the one or more synthetic gene constructs comprising one or more polypeptide coding sequences operatively linked to a promoter.

[0278] 59. The method of Item 58, wherein the one or more synthetic gene constructs comprise a collection of synthetic gene constructs, the collection optionally including portions generated by combination.

[0279] 60. The method of any one of entries 7 to 59, wherein the first source-specific barcode population further comprises a target nucleic acid barcode, wherein the target nucleic acid barcode identifies the population as a population of labeled nucleic acids.

[0280] 61. The method of any one of items 7 to 60, wherein the second source-specific barcode population further comprises a target molecule barcode, wherein the target molecule barcode identifies the population as a population that labels a target molecule.

[0281] 62. The method of any one of items 1 to 61, wherein the source-specific barcode further comprises one or more cleavage sites.

[0282] 63. The method of Item 62, wherein at least one cleavage site is oriented such that cleavage at said site releases the source-specific barcode from the substrate associated therewith.

[0283] 64. The method of Item 63, wherein the substrate comprises the beads.

[0284] 65. The method of any one of entries 62 to 64, wherein at least one cleavage site is oriented such that cleavage at said site releases the source-specific barcode from the target-molecule-specific binder.

[0285] 66. The method of any one of items 1 to 65, further comprising dissolving the cells.

[0286] 67. The method of any one of items 1 to 66, wherein the target molecule and the target nucleic acid, optionally in an associated form, are produced by the individual cells in the individual compartment.

[0287] 68. The method of any one of items 1 to 67, wherein the cell is a B cell, a plasmablast, or a plasma cell, and the target molecule is an antibody, and the target nucleic acid encodes the antibody.

[0288] 69. The method of any one of items 1 to 68, wherein the target molecule is a polypeptide and the target nucleic acid encodes the target molecule in its respective individual compartment.

[0289] 70. The method of any one of items 1 to 69, wherein the cells are contacted with one or more test agents.

[0290] 71. The method of Item 70, wherein the test reagent comprises a small molecule, nucleic acid, polypeptide, or polysaccharide.

[0291] 72. The method of Item 71, wherein the polypeptide comprises an antibody or an antibody fragment.

[0292] 73. The method of any one of items 70 to 72, further comprising labeling the test reagent with the source-specific barcode.

[0293] 74. The method of any one of items 70 to 73, wherein the individual test reagent is marked with a test reagent-specific barcode.

[0294] 75. The method of any one of items 1 to 74, further comprising amplifying one or more of the source-specific barcode, the target molecule-specific binding agent barcode, the test agent-specific barcode, the target nucleic acid barcode, and the target molecule barcode.

[0295] 76. The method of any one of items 1 to 75, further comprising detecting one or more of the target molecule-specific binding barcode, the test agent-specific barcode, the target nucleic acid barcode, and the target molecule barcode.

[0296] 77. The method of any one of items 1 to 76, wherein detecting the sequence of the source-specific barcode, the target molecule-specific binding agent barcode, the test agent-specific barcode, the target nucleic acid barcode, and / or the target molecule barcode comprises hybridization, sequencing, or a combination thereof.

[0297] 78. The method of any one of items 1 to 77, further comprising quantifying one or more of the source-specific barcode, the target molecule-specific binding agent barcode, the test-drug-specific barcode, the target nucleic acid barcode, and the target molecule barcode.

[0298] 79. The method of any one of items 1 to 78, further comprising, prior to separation, contacting the cells with a specific binder that specifically binds to target molecules on the surface of the cells.

[0299] 80. The method of entrant 79, wherein the specific binder comprises an antigen and the target molecule comprises an antibody.

[0300] 81. The method of any one of entries 79 to 80, further comprising, after separation, dissolving the cells, wherein the specific binder is bound by the source-specific barcode.

[0301] 82. The method of Item 81, wherein the antigen is further labeled with ssDNA or partially double-stranded DNA, and the source-specific barcode is used as a primer to synthesize a complementary strand.

[0302] 83. The method of any one of entries 81 to 82, further comprising synthesizing cDNA from mRNA encoding antibody heavy and / or light chains, wherein the source-specific barcode triggers the synthesis of the cDNA.

[0303] 84. The method of any one of entries 79 to 83, further comprising collecting the compartments and quantifying the amount of antigen bound to the antibody on each cell surface using the source-specific barcode on the antigen.

[0304] 85. The method of any one of items 79 to 84, further comprising determining the sequence of the heavy chain and / or light chain, and assigning the sequence to the antigen that binds the antibody.

[0305] 86. The method of any one of items 79 to 85, wherein the antigen comprises an HIV antigen, such as gp41 and / or gp120.

[0306] 87. A method for determining the specificity of a test agent for a target molecule, the method comprising assigning a set of target molecules to a compartment according to any one of items 1 to 78, wherein the method further comprises:

[0307] Prior to separation, cells expressing the target molecule are contacted with an aggregate of test drugs labeled with test drug-specific barcodes;

[0308] Isolate the target molecule bound to the test agent; and

[0309] The sequence of the test agent-specific barcode and the sequence of the source-specific barcode are determined to identify the test agent bound to the target molecule.

[0310] 88. The method of Item 87, wherein the target molecule comprises a cell surface protein, and the target nucleic acid encodes the cell surface protein.

[0311] 89. The method of any one of entries 87 to 88, further comprising washing unbound test reagent from said cells.

[0312] 90. The method of any one of entries 87 to 89, wherein the target nucleic acid is DNA or RNA.

[0313] 91. The method of any one of entries 87 to 90, wherein the target nucleic acid is cDNA or mRNA.

[0314] 92. The method of any one of entries 87 to 91, further comprising labeling the test reagent with the source-specific barcode.

[0315] 93. The method of any one of items 87 to 92, wherein the test reagent comprises a small molecule compound, nucleic acid, polypeptide or polysaccharide.

[0316] 94. A method for determining the affinity and / or specificity of a test agent for a target molecule, said method comprising assigning a set of target molecules to a compartment according to any one of items 1 to 78, said method further comprising:

[0317] The labeled target molecule is brought into contact with a test agent that binds to the same detectable label;

[0318] The detectable label is used to separate the labeled target molecule bound to the test agent;

[0319] Determine the sequence of the source-specific barcode on the isolated target molecule; and

[0320] The source-specific barcode associated with the isolated target molecule is quantified, thereby determining the affinity of the test agent for the target molecule.

[0321] 95. The method of Item 94, further comprising quantifying the source-specific barcode that was not separated from the test reagent to normalize the concentration.

[0322] 96. The method as described in any one of entries 94 to 95, further comprising splitting the source-specific barcode after separation.

[0323] 97. The method of any one of items 94 to 96, further comprising associating the isolated barcode with the target nucleic acid to determine the sequence of the target molecule bound to the test agent.

[0324] 98. The method of any one of entries 94 to 97, wherein the target molecule comprises a polypeptide.

[0325] 99. The method as described in entry 98, wherein the polypeptide comprises an antibody.

[0326] 100. The method of Item 99, wherein the antibody comprises an anti-HIV antibody, such as an anti-gp41 or anti-gp120 antibody, and the test agent comprises a collection of potential HIV antigens and / or immunogens.

[0327] 101. The method of Item 100, wherein the test agent comprises an antibody and the target molecule comprises a protein expressed on the surface of the cell.

[0328] 102. The method of 101, wherein the test agent comprises an anti-HIV antibody, such as an anti-gp41 or gp120 antibody, and the target molecule comprises a collection of HIV antigens and / or immunogens.

[0329] 103. The method of any one of items 94 to 102, further comprising determining one or more of the dissociation constant, association rate, or dissociation rate of the target molecule in relation to the test agent.

[0330] 104. A method for determining the expression of a target molecule on the surface of a cell aggregate, the method comprising assigning the target molecule aggregate to compartments according to any one of items 1 to 78, the method further comprising:

[0331] The sample cells to be separated are brought into contact with a collection of test reagents, each labeled with a unique test reagent barcode;

[0332] The sequence of a source-specific barcode is determined, thereby determining the expression of the molecule on the surface of the cell assembly.

[0333] 105. The method of Item 104, further comprising associating the isolated barcode with the target nucleic acid to determine the sequence of the target molecule bound to the test agent.

[0334] 106. The method of any one of items 104 to 105, wherein the test reagent comprises a polypeptide.

[0335] 107. The method of Item 106, wherein the polypeptide comprises an antibody specific to a cell surface marker.

[0336] 108. The method of any one of items 104 to 108, further comprising associating molecules expressed on the surface of the cell with other measures of cell type, cell cycle or cell state.

[0337] 109. A method for identifying proteins with specific activities of interest from a cell population, the method comprising assigning a set of target molecules to compartments according to any one of items 1 to 78, the method further comprising:

[0338] Isolate the target molecule having the specific activity of interest; and

[0339] Identify the source-specific barcode of the isolated target molecule with the specific activity of interest.

[0340] 110. The method of claim 110, further comprising identifying the target nucleic acid molecule encoding the protein by matching the sequence of a source-specific nucleic acid barcode.

[0341] 111. The method of claim 111, wherein the activity is antigen binding and the protein is an antibody.

[0342] 112. The method of claim 112, further comprising generating the antibody of interest by expressing a nucleic acid molecule identified as encoding an antibody of interest.

[0343] 113. A method for detecting and / or quantifying post-translational modified target molecules in a sample, the method comprising assigning a set of target molecules to a specific compartment according to any one of items 1-78, the method further comprising:

[0344] The labeled target molecule is brought into contact with a first specific binder and a second specific binder, wherein the first specific binder specifically binds to the post-translational modified target molecule at the modification site, and the second specific binder specifically binds to the untranslated modified target molecule at the modification site, wherein the first and second specific bindings are specifically marked with specific barcodes by the specific binders.

[0345] The target molecule is isolated, thereby separating the specific binder bound to the target molecule;

[0346] The presence of the post-translational modified protein in the sample is determined by measuring the source-specific barcode and the sequence of the specific binding agent barcode present on the isolated specific binding agent.

[0347] 114. The method as described in item 113, wherein the post-translational modification includes phosphorylation.

[0348] 115. The method of any one of entries 113 to 114, wherein the separation comprises contacting the target molecule with a third specific binder that does not bind to the modified site.

[0349] 116. The method of any one of items 113 to 115, wherein the specific binder comprises an antibody.

[0350] 117. A barcode marking complex comprising:

[0351] Solid or semi-solid substrates, and

[0352] The barcode includes a plurality of barcode elements reversibly coupled thereto, each of which comprises an index nucleic acid recognition sequence, a nucleic acid capture sequence specifically binding to a target nucleic acid, and one or more specific binders specifically binding to the target molecule.

[0353] 118. The complex as described in entry 117, wherein the substrate comprises beads.

[0354] 119. The composite as described in entry 117, wherein the substrate comprises a hydrogel.

[0355] 120. The complex as described in any one of entries 117 to 119, wherein the source-specific barcode further comprises one or more cleavage sites.

[0356] 121. The complex as described in entry 120, wherein at least one cleavage site is oriented such that cleavage at said site releases the source-specific barcode from the substrate to which it is coupled.

[0357] 122. The complex of any one of entries 117 to 121, wherein at least one cleavage site is oriented such that cleavage at said site releases the source-specific barcode from the target-molecule-specific binder.

[0358] 123. The complex of any one of entries 117 to 122, wherein the indexed nucleic acid recognition sequence comprises RNA, DNA, or a combination thereof.

[0359] 124. The complex of any one of entries 117 to 123, wherein the source-specific barcode comprises RNA, DNA, or a combination thereof.

[0360] 125. The complex of any one of entries 117 to 124, wherein the source-specific barcode further comprises one or more capture portions covalently or non-covalently bonded.

[0361] 126. The complex as described in entry 125, wherein one or more of the capturing portions contain biotin.

[0362] 127. The complex of any one of entries 117 to 126, wherein the source-specific barcode comprises biotin-16-UTP.

[0363] 128. The complex of any one of entries 117 to 127, wherein the source-specific barcode further comprises a sequencing adaptor.

[0364] 129. The complex of any one of entries 117 to 128, wherein the source-specific barcode further comprises a universal initiation site.

[0365] 130. The complex of any one of entries 117 to 129, wherein the target molecule specific binder comprises an antibody or a fragment thereof, a polypeptide or peptide containing an epitope recognized by the target molecule, or a nucleic acid.

[0366] 131. The complex of any one of entries 117 to 130, wherein each of the source-specific barcodes comprises one or more indexes, one or more sequences for achieving gene-specific capture and / or amplification, and / or one or more sequences for achieving sequencing library construction.

[0367] 132. The complex as described in entry 131, wherein each of the barcodes comprises four indices.

[0368] 133. The complex as described in any one of entries 131 to 132, wherein the one or more indices are assembled in combination.

[0369] 134. The complex of any one of entries 131 to 133, wherein each of the sequences that enables the construction of the sequencing library comprises an Illumina P7 sequence and / or an Illumina sequencing primer.

[0370] 135. The complex of any one of entries 117 to 133, wherein each of the barcodes contains a primer for DNA synthesis.

[0371] 136. The complex as described in entry 135, wherein the primer is suitable for DNA synthesis on a DNA template or an RNA template.

[0372] 137. A method for determining the level of a post-translational modified target protein in a sample, the method comprising:

[0373] Provide samples containing cells;

[0374] A single cell or a portion of the sample is co-separated into individual compartments along with a first specific binder and a second specific binder, wherein the first specific binder specifically binds to the post-translational modified target molecule at the modification site, and the second specific binder specifically binds to the untranslated post-modification target molecule at the modification site, wherein the first and second specific bindings are marked with specific binder-specific barcodes, and wherein each individual compartment further includes a source-specific barcode containing a unique nucleic acid recognition sequence that maintains or carries information about the source compartment of the segregated cell in the sample;

[0375] The first specific binder and the second specific binder in the individual compartments are marked with the source-specific barcode present in the individual compartments;

[0376] The target molecule is separated, thereby separating the first and second specific binders bound to the target molecule;

[0377] The presence of the post-translational modified protein in the sample is determined by detecting the source-specific barcode present on the separation-specific binder and the nucleotide sequence of the specific binder barcode.

[0378] 138. The method as described in entry 137, wherein the post-translational modification includes phosphorylation.

[0379] 139. The method of any one of entries 137 to 138, comprising quantifying the levels of post-translational modifications and unmodified target proteins.

[0380] 140. The method of any one of entries 137 to 139, wherein the separation comprises contacting the target molecule with a third specific binder that does not bind to the modified site, wherein the third specific binder is detectably labeled.

[0381] 141. The method of any one of items 137 to 140, wherein the specific binder comprises an antibody.

[0382] 142. The method of any one of items 137 to 141, further comprising labeling the target nucleic acid in the individual compartment with the source-specific barcode present in the individual compartment to form a source-labeled target nucleic acid, wherein the source-labeled target nucleic acid from each individual compartment contains a unique indexed nucleic acid recognition sequence that is the same as or matches the source-labeled target molecule;

[0383] The nucleotide sequence of the source-specific barcode is detected, thereby assigning the target protein set to the sample or target nucleic acid in the sample set, while maintaining information about the source compartment of the target protein and the target nucleic acid.

[0384] The following examples are intended to illustrate, but not limit, the invention. Example

[0385] Example 1

[0386] Compartment-specific labeling of target molecules

[0387] The following examples demonstrate the labeling of molecular assemblies using the methods described herein.

[0388] In this embodiment, the selected target molecule is an antigen, such as the natural isotype (wtGP120) or the N332A mutant isotype (gp120). N332AHIV envelope glycoprotein gp120 in the form of α. Two isotypes of the antigen were labeled with two different single-stranded DNA tags. B cells exhibiting affinity for either isotype of the antigen, immunoglobulin M (IgM), were then incubated with the mixture of labeled antigens. Cells were then washed and encapsulated in approximately 100 pL droplets with dissolution buffer, reverse transcription (RT) buffer, and enzymes (RT enzyme, DNA polymerase, and restriction enzyme BclI) and hydrogel beads using a microfluidic device (Kim et al. - Fabrication of monodisperse gel shells and functional microgels in microfluidic devices. Angew Chem Int Ed Engl. 2007; 46(11):1819-22; Abate, AR et al. - Beating Poisson encapsulation statistics using close-packed ordering. Lab Chip (2009). 9(18), 2628-31. doi:10.1039 / b909386a). The hydrogel beads carried a mixture of barcoded primers complementary to the DNA tag and the mRNA of heavy and light chain antibody genes.

[0389] Antigen labeling:

[0390] The attachment of an antigen to a DNA tag can be direct (through a covalent bond formed by chemical reactions such as the NHS-ester reaction) or indirect (non-covalent bonding, such as the biotin-streptavidin interaction). The second case is described below.

[0391] The labeled antigen, in its final form, is constructed as a biotinylated antigen bound to one pocket of a streptavidin molecule, which has three additional pockets bound by a biotinylated single-stranded DNA tag (see [link to documentation]). Figure 8 DNA tags can be obtained directly as biotinylated oligonucleotides or prepared by PCR and treated with λ exonuclease. The second strategy is described below.

[0392] DNA tags are formed by PCR using a pair of oligonucleotides with a 5' extension of 20 nucleotides (nt). The extension on one of the oligonucleotides is... The first product is amplified by a second PCR using a set of primers, namely SBS3 (ACACTCTTTCCCTACACGACGCTCTTCCGATCT (SEQ ID NO:2)) and an extension on another oligonucleotide, which is a randomly selected sequence (such as GGAGTTGTCCCAATTCTTGT (SEQ ID NO:3)). The tag-specific primers are common to all antigen tags. The amplified region is a 100-base-pair (bp) sequence of the pUC19 plasmid. A second PCR amplification of this first product is performed using a second set of primers, identical to the 5' extension of the first primer set, namely SBS3 (SEQ ID NO:2) and TSP. The SBS3 (SEQ ID NO:2) primer in the second set is biotinylated at its 5' end, and the TSP primer is phosphated at its 5' end. The final double-stranded product is then incubated with a λ exonuclease that targets the 5' phosphorylated strand for degradation and biotinylate-protects the other strand. The final product is a single-stranded 5' biotinylated DNA molecule. This procedure can generate different DNA tags by altering the amplification region of the plasmid in the first PCR step. The sequences will be identical for approximately 20 nt at both ends, but different in the center, thus providing different tags for different target molecules.

[0393] Here, by targeting two types of gp120 antigen (wtGP120 and gp120) N332A Two different tags were prepared by targeting two different regions of the pUC19 plasmid (of the same type).

[0394] The targeted molecule was a recombinant antigen displaying biotin. This antigen was first incubated with a 10-fold excess of free streptavidin, then washed and purified using a commercially available size exclusion column. The purified streptavidin-bound antigen was then incubated with a 10-fold excess of 5'-biotinylated single-stranded DNA molecules (the commercially available oligonucleotide or product of the λ exonuclease described above), washed again, and purified using a size exclusion column.

[0395] The final product is an antigen whose biotinylated tag is bound to one pocket of the streptavidin molecule, while the other three pockets are occupied by a 5' biotinylated single-stranded DNA tag. Here, wtGP120-tag 1 and gp120 are added before incubation with cells. N332A - The final products of label 2 are mixed in equal proportions (see Figure 8 ).

[0396] Cell markers:

[0397] Prior to compartmentalization, various cell populations were incubated with labeled antigens for 20 minutes, followed by three washes in large volumes of PBS. In the first control experiment, the populations consisted of cells with specific antigens that were distinct from wtGP120 or gp120. N332A The mixture consisted of four different cell lines with varying affinity. In subsequent experiments, the cell populations were derived from samples from HIV-infected patients, such as those showing broadly neutralizing HIV-1 antibodies in their serum, or those at different stages of infection.

[0398] Preparation of hydrogel beads and synthesis of barcoded oligonucleotides:

[0399] In some instances, hydrogel beads were prepared from PEG-DA oligomers in microfluidic chips, wherein an aqueous PEG-DA solution was dispersed by hydrodynamic flow focusing to form droplets in a fluorinated oil continuous phase (Anna, S., Bontoux, N., and Stone, H. (2003). Formation of dispersions using “flow focusing” in microchannels. Applied Physics Letters, 52(3), 364-366. doi:10.1063 / 1.1537519). The beads were then crosslinked via a UV-activated photoinitiator. 400 μM of a double-stranded DNA oligonucleotide (double strand) called RanA was added, which had a 5'acrydite modification at one end and a 4-nt 5' overhang on the other side (top chain: 5'Acrydite-TCTTCACGGAACGA (SEQ ID NO:4); bottom chain: 5' phosphate- CAGT TCGTTCCGTGAAGA (SEQ ID NO:5) was covalently cross-linked with the hydrogel matrix via the acrylate terminal groups of the PEG-DA oligomer. After polymerization and washing in Tris-HCl pH 7.4 20 mM; NaCl 50 mM; Tween 0.01%; EDTA 1 mM, a first duplex was ligated using T7 DNA ligase. The first duplex had a 4-nt 5' protrusion on one side compatible with the acrydite duplex, and another 4-nt 5' protrusion on the other side compatible with the downstream linker. In its duplex portion, this first duplex had a randomly selected 8-nt sequence called RanB (GACTAGAA (SEQ ID NO:6)) followed by a BclI restriction site. TGATCA (SEQ ID NO:7) ), followed by SBS12 Sequence (GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT (SEQ ID NO: 8)).

[0400] After ligation and washing, barcodes were synthesized through four consecutive ligations mediated by T7 DNA ligase using 20-nt DNA duplexes with 4-nt protrusions at both 5' ends. Different 4-nt protrusions were used in each ligation step to ensure that only four indices could be assembled in the correct order. To create a variety of barcodes, hydrogel bead batches were evenly distributed into the wells of a 96-well plate. Each well contained a duplex with a unique 20-nt sequence (index) designed to maintain a clear, at most three errors, along with ligation buffer and the enzyme. After ligation incubation, the entire reaction volume of the plate was pooled into a tube and washed. The next ligation step was performed in the same manner: the pooled batch was evenly distributed into a new plate containing another set of 96 different duplexes, along with ligation buffer and the enzyme. This separate-pool synthesis resulted in a combinatorial diversity of 96. 4 (Over 84 million). Finally, the last double-stranded strand ligated to the newly synthesized barcode is partially double-stranded to allow ligation via T7 DNA ligase and terminates at a long single-stranded 3' end (see example). Figure 9 and 10 The double-stranded region is a defined adapter sequence (TACGCTACGGAACGA (SEQ ID NO:9). The single-stranded region consists of a randomized 12-nt sequence (GNNNGNNGNNNG (SEQ ID NO:10)) and an antisense sequence for either the mRNA that initiates reverse transcription or the TSP that initiates DNA polymerization. The 12-nt random sequences act as unique molecular identifiers (UMIs): they allow the differentiation between sequences originating from different RT-initiated events (with different UMIs) and sequences amplified by PCR from the same cDNA (with the same UMI) (Shiroguchi, K., Jia, TZ, Sims, PA and Xie, XS (2012). Digital RNA sequencing minimizes sequence-dependent bias and amplification noise with optimized singlemolecule barcodes. Proceedings of the National Academy of Sciences, 109(4), 1347-1352. doi:10.1073 / pnas.1118018109).

[0401] The release of oligonucleotides via restriction enzyme cleavage can be achieved by replacing this sequence with any cleavable chemical group that can be linked to the nucleic acid (such as a photocleavable or pH-sensitive moiety).

[0402] Encoding phenotypes into DNA-DNA polymerization and RT within droplets:

[0403] Using a microfluidic chip, labeled cells were encapsulated in droplets along with a dissolution buffer, reverse transcriptase and its buffer, DNA polymerase, BclI restriction enzyme, and hydrogel beads. The hydrogel beads carried aggregates of partially double-stranded DNA molecules, all possessing the same DNA barcode; this DNA barcode was unique on each hydrogel bead. Droplets were generated using a Poisson distribution during cell encapsulation, and the cell concentration was selected such that the average number of cells per droplet was <1, ensuring that most droplets contained no more than one cell. Subsequently, deformable hydrogel beads were injected in a close-packed array (Abate, AR et al. (2009)) to ensure that most droplets contained a single bead. Figure 11 The single-stranded portion of the DNA molecule bound to the hydrogel beads consists, in sequence, of a UMI sequence (SEQ ID NO: 10) and an antisense sequence (three types of termination sequences in equal proportions on the beads) that is antisense to most of the 5' end of the constant (Fc) region of the heavy and light chain mRNA of the DNA tag or the 3' end of the antibody gene.

[0404] The emulsion was then incubated at 55°C for 1 hour and 30 minutes. During this incubation, BclI restriction enzymes cleaved the barcoded oligonucleotides and released them into the entire volume of the droplet. Simultaneously, labeled cells were lysed, and the released mRNA annealed to its complementary sequence on the single-stranded portion of the released barcoded oligonucleotide, while RT enzymes elongated the barcoded DNA, thereby copying the sequences of the antibody heavy and light chain mRNAs. Simultaneously, the antigen-binding DNA tag annealed to its complementary sequence on the single-stranded portion of the barcoded DNA bound to the hydrogel, and polymerases extended the barcoded DNA molecule, thereby copying the DNA tag. Adding DNA barcoding to both the mRNA-derived cDNA and the antigen-binding DNA tag confers single-cell specificity to these sequences (see, for example...). Figure 12 ).

[0405] Batch amplification and sequencing:

[0406] The emulsion was then placed at 70°C to inactivate the RT enzyme, disrupt the emulsion, and the aqueous phase was recovered. The DNA was then purified using a commercial kit (Agencourt RNAClean XP).

[0407] Then, for the heavy chain cDNA, light chain cDNA, and DNA tag, the cDNA and the barcoded tag were amplified in separate PCRs. The primers used matched the ends of the cDNA and DNA tag and had a 5' extension containing... The sequences required for sequencing are the anchoring sequences P7 (CAAGCAGAAGACGGCATACGAGAT (SEQ ID NO:11)) and P5 (AATGATACGGCGACCACCGAGATCT (SEQ ID NO:12)) (see example...). Figure 13 ).

[0408] Example 2

[0409] Phenotypic and sequence recovery of antibody-secreting cells

[0410] This embodiment describes a variation of Example 1, which is suitable for sequencing selected antibody-secreting cells, for example, to screen for antibodies that can bind to the HIV gp120 protein.

[0411] In this embodiment, the target molecule is an antigen, specifically the natural isotype (gp120). wt ) or present as the N332A mutant isotype (gp120) N332A HIV envelope glycoprotein gp120 in the form of [missing information]. Two isotypes of the antigen are tagged with two different single-stranded DNA tags. Plasma cells are encapsulated in approximately 100 μL droplets using a microfluidic device. Each droplet comprises a mixture of culture medium, the DNA-tagged (barcoded) antigen, and hydrogel beads carrying a mixture of barcoded primers complementary to the DNA tag and the mRNA of the heavy and light chain antibody genes. After incubation at 37°C to allow antibody secretion from the plasma cells, the droplets are fused using a microfluidic device with other droplets containing a dissolution buffer, an RT buffer, and enzymes (RT enzymes, DNA polymerases, and restriction enzymes such as BclI) and hydrogel beads carrying a mixture of barcoded primers complementary to the DNA tag and the mRNA of the heavy and light chain antibody genes. After incubation at 55°C (this facilitates enzymatic digestion via BclI, RT of mRNA, and DNA polymerization), the emulsion is chemically disrupted and the aqueous phase is recovered. Antibodies and associated DNA-tagged (barcode) antigens were purified using commercially available protein-A / G agarose resin. cDNA was purified from the protein-A / G agarose resin flow. The antigen tag and cDNA were separately amplified by PCR and then sequenced.

[0412] Antigen labeling is performed as in Example 1.

[0413] Preparation of hydrogel beads and synthesis of barcoded oligonucleotides:

[0414] As in Example 1, proceed as is.

[0415] Encoding phenotypes into DNA-DNA polymerization and RT within droplets:

[0416] As described in Example 1, using a microfluidic device, single cells were compartmentalized in microdroplets along with culture medium and single hydrogel beads, each hydrogel bead carrying a partially double-stranded DNA molecule containing a barcode, which was unique on each hydrogel bead. The collected emulsion was then incubated at 37°C for 30 minutes to 6 hours to allow the cells to secrete antibodies (see Example 1). Figure 14 The image above and Figure 15 ).

[0417] The emulsion was then injected into another microfluidic device, in which the droplet fused with other droplets containing a dissolution buffer, an RT buffer, and enzymes, DNA polymerase, and BclI restriction enzyme. Fusion was achieved using an electric field generated by electrodes, such that the two droplets merged when the electric field was applied (Chabert, M., Dorfman, K., and Viovy, J. (2005). Droplet fusion by alternating current (AC) field electrocoalescence in microchannels. Electrophoresis, 26(19), 3706-3715. doi:10.1002 / elps.200500109) (see also...) Figure 14 The image below and Figure 15 ).

[0418] As described in Example 1, the emulsion was then incubated at 55°C for 1 hour and 30 minutes for oligonucleotide release and RT. Simultaneously, the antigen-bound DNA tag annealed to its complementary sequence on the single-stranded portion of the released barcoded DNA, and polymerase extended the barcoded DNA molecule, thereby copying the DNA tag. Thus, barcoded cDNA of both heavy and light chains was generated via RT, and a barcoded antigen-DNA tag was generated via DNA polymerization (see example...). Figure 16 ).

[0419] Batch purification and amplification, as well as sequencing:

[0420] The emulsion was then placed at 70°C to inactivate the RT enzyme. After cooling on ice, a concentrated solution of unlabeled antigen was added to the top of the emulsion to prevent free labeled antigen (and now barcoded antigen) from one droplet from binding to unbound antibody from another droplet after the emulsion was disrupted. Figure 17).

[0421] The emulsion was then disrupted and the aqueous phase was recovered. Antibodies (bound to barcoded labeled antigens) were purified from this phase using commercially available protein-A / G agarose resin. cDNA was purified from the protein-A / G agarose resin flow using a commercial kit (Agencourt RNAClean XP).

[0422] Similar to Example 1, the barcoded cDNA and the barcoded tag were then amplified in separate PCRs for the heavy chain cDNA, light chain cDNA, and DNA tag. Figure 16 and 17 The primers used match the ends of the cDNA and DNA tag and have a 5' extension containing... The sequences required for sequencing are the anchor sequences P7 (SEQ ID NO: 11) and P5 (SEQ ID NO: 12).

[0423] Example 3

[0424] Blood cell counting via sequencing

[0425] In this embodiment, cells are incubated with DNA-tagged antibodies, washed, and encapsulated with DNA polymerase, restriction enzyme BclI, and hydrogel beads carrying partially single-stranded barcoded oligonucleotides (antisense material of the DNA tag). The emulsion is placed at 55°C to release the oligonucleotides and trigger DNA polymerization (extending the barcoded oligonucleotides to copy the antibody-bound DNA tag). Sequencing then reveals all antibodies bound to each cell, indicating the surface markers exhibited by the cells. In some embodiments, reverse transcription (RT) reagents are added to determine the mRNA sequence or transcription level; in this case, the cells are lysed after encapsulation.

[0426] Antibody labeling:

[0427] In some embodiments, target-specific antibodies are modified by coupling to biotin, for example via an NHS-ester reaction with lysine residues, or by partial conjugation of hydrazide to oxidized antibody carbohydrate residues.

[0428] The following biotinylated target-specific antibody was labeled (similar to Example 1, but using a biotinylated antibody instead of a biotinylated antigen and a modified purification kit). First, the antibody was incubated with a 10-fold excess of free streptavidin, then washed and purified using commercially available protein-A / G agarose resin. The purified streptavidin-bound antibody was then incubated with a 10-fold excess of 5' biotinylated single-stranded DNA molecules, then washed again and purified using commercially available protein-A / G agarose resin. The final product was an antibody whose biotinylate tag was bound to one pocket of the streptavidin molecule, while the other three pockets were occupied by a 5' biotinylated single-stranded DNA tag.

[0429] Cell markers:

[0430] Before compartmentalization, the cell population was incubated with labeled antibodies for 20 minutes and then washed three times in large amounts of PBS.

[0431] Preparation of hydrogel beads and synthesis of barcoded oligonucleotides:

[0432] As in Example 1, proceed as is.

[0433] Encode the phenotype into DNA-DNA polymerization and (optionally) RT in the droplet:

[0434] Using a microfluidic chip, labeled cells, along with DNA polymerase, BclI restriction enzyme, and hydrogel beads, are encapsulated in microdroplets. The hydrogel beads carry aggregates of partially double-stranded DNA molecules. Figure 18 If RT is to be performed, a cell lysis reagent is also added. The single-stranded portion of the DNA bound to the hydrogel beads consists, in sequence, of a UMI sequence (SEQ ID NO:10) and a sequence antisense to the 3' end of the DNA tag.

[0435] The emulsion was then incubated at 55°C for 1 hour and 30 minutes. During this incubation, the BclI restriction enzyme cleaved the barcoded oligonucleotides and released them into the entire volume of the droplet. The antibody-bound DNA tag annealed to its complementary sequence on the single-stranded portion of the barcoded DNA bound to the hydrogel, and the polymerase extended the barcoded DNA molecule, thereby copying the DNA tag. Figure 19 DNA barcodes linked to antibody DNA tags give these sequences single-cell specificity.

[0436] In RT, a portion of the oligonucleotide aggregate bound to the hydrogel beads is complementary to the mRNA, and the cells are lysed by incubation at 55°C to release the mRNA, which anneals to the complementary sequence on the single-stranded portion of the released barcoded oligonucleotide, and the RT enzyme extends the barcoded DNA to copy the sequence of the mRNA.

[0437] Batch purification and amplification, as well as sequencing:

[0438] The emulsion was then placed at 70°C to inactivate the RT enzyme. After cooling on ice, the emulsion was disrupted and the aqueous phase was recovered. The DNA was then purified using a commercial kit (Agencourt RNAClean XP).

[0439] The barcoded tag (and cDNA in the RT case) was then amplified by PCR. The primers used matched the ends of the tag and had a 5' extension containing... The sequences required for sequencing, namely the anchor sequences P7 (SEQ ID NO:11) and P5 (SEQ ID NO:12) Figure 20 ).

[0440] Example 4

[0441] Protein phosphorylation analysis

[0442] In this embodiment, the target molecule is a protein that has undergone post-translational phosphorylation modification, such as tyrosine kinases epidermal growth factor receptor (EGFR) and Janus kinase 2 (JAK2), and downstream kinases phosphorylating transcription factors, signal transducers, and transcription activators 3 (STAT3). This embodiment provides a method for measuring protein phosphorylation levels.

[0443] Antibodies specific to different target proteins are labeled with different single-stranded DNA tags. Figure 21 Cells of interest were incubated in an appropriate culture medium, with or without a test agent (such as a drug or potential drug). The cells were then washed and encapsulated in approximately 100 pL droplets using a microfluidic device, along with DNA-tagged target-specific antibodies, lysis buffer, PCR buffer, enzymes (DNA polymerase and restriction enzyme BclI), biotin-tagged antibodies specific to target domains lacking phosphorylation sites, and hydrogel beads (Abate, AR et al. (2009)). These hydrogel beads carried a mixture of barcoded primers complementary to the DNA tag of the DNA-tagged target-specific antibody.

[0444] Antibody labeling:

[0445] DNA tags are generally formed as in Example 1. However, any suitable plasmid can be used to generate tags, and the 5' extension is common to all antibody tags (SEQ ID NO:3), referred to as the tag-specific primer (TSP). Different tags are prepared by amplifying different regions of the plasmid, such as the three required in this example.

[0446] The target-specific antibody may be commercially available and pre-conjugated with biotin, or may be biotinylated via reaction with an NHS-ester of a lysine residue, or modified by partial conjugation of an acylhydrazine to an oxidized antibody carbohydrate residue. As in Example 1, the biotinylated target-specific antibody is typically labeled with a DNA tag.

[0447] The final product is an antibody whose biotin tag is bound to one pocket of the streptavidin molecule, while the other three pockets are occupied by a 5' biotinylated single-stranded DNA tag. The labeled target-specific antibodies are mixed in equal proportions before analysis.

[0448] Preparation of hydrogel beads and synthesis of barcoded oligonucleotides:

[0449] This is performed as in Example 1, where the UMI will distinguish between target-specific antibodies derived from different DNA-tagged sequences rather than sequences derived from different RT-induced events.

[0450] Encoding phenotypes into DNA-DNA polymerization within droplets:

[0451] The target-specific antibody, tagged with a DNA tag, is specific to the phosphorylated region of the target protein. An antibody targeting the phosphorylated protein and an antibody targeting the unphosphorylated protein are used to encode specific tags for each state. A third antibody, specific to the non-phosphorylated domain of the protein, is used to tag the binding complex. This antibody is biotin-tagged using the method described in the "Antibody Tagging" section of this embodiment. This antibody does not have a DNA tag and is encapsulated in all droplets. Biotin tagging allows for subsequent separation of the antibody-target complex from unbound antibody. Alternatively, the complex can be purified by size exclusion chromatography without using a third antibody.

[0452] Using a microfluidic chip, single cells were encapsulated in droplets along with a DNA-tagged target-specific antibody, a dissolution buffer, DNA polymerase (New England Biolabs Cleno fragment (3'→5'exo-)) and its buffer, BclI restriction enzyme, a biotin-tagged antibody as described above, and hydrogel beads carrying aggregates of partially double-stranded DNA molecules, all possessing the same DNA barcode; this DNA barcode was unique on each hydrogel bead. Droplets were generated using a Poisson distribution during cell encapsulation, and the cell concentration was chosen such that the average number of cells per droplet was <1, ensuring that most droplets contained no more than one cell. Deformable hydrogel beads were injected in a close-packed array (Abate, AR et al. (2009)) to ensure that most droplets contained a single bead. Figure 22 The single-stranded portion of the DNA molecule bound to the hydrogel beads consists of a UMI sequence and an antisense sequence relative to the 3' end of the DNA tag.

[0453] The emulsion was then incubated at 37°C for 1 hour and 30 minutes. During this incubation, BclI restriction enzymes cleaved the barcoded oligonucleotides and released them into the entire volume of the droplet. Simultaneously, cells lysed, and the released target proteins were captured by labeled antibodies, which annealed to complementary sequences on the single-stranded portions of the barcoded DNA bound to the hydrogel, and polymerases extended the barcoded DNA molecules, thereby copying the DNA tag. These antibody-target complexes were captured by biotin-labeled antibodies. Adding DNA barcoding to antibody-bound DNA tags confers single-cell specificity to these sequences. Figure 23 ).

[0454] Recovery of phosphorylated target proteins and separation of phosphorylated and unphosphorylated target proteins:

[0455] The emulsion was placed at 75°C to inactivate the polymerase, then the emulsion was disrupted and the aqueous phase was recovered. It was then processed via streptavidin agarose column chromatography. (Streptavidin agarose column) separates phosphorylated and unphosphorylated proteins, along with linked antibodies, DNA tags, and barcodes, from unbound DNA-tagged antibodies. Figure 24 Alternatively, the complex can be purified by size exclusion chromatography.

[0456] Batch amplification and sequencing:

[0457] The barcoded DNA tags were purified using a commercial kit (Agencourt AMPure XP). Then, for tags derived from phosphorylated and unphosphorylated proteins, the barcoded tags were amplified in separate PCRs. The primers used matched the ends of the DNA tags and had a 5' extension containing… The sequences required for sequencing are the anchor sequences P7 (SEQ ID NO:11) and P5 (SEQ ID NO:12).

[0458] Single-cell-level target protein phosphorylation data were quantified by the number of target-specific DNA tag reads or UMIs for each DNA barcode and compared with the total target protein mass.

[0459] Example 5

[0460] Alternative protein phosphorylation analysis

[0461] In this embodiment, the target molecule is a protein that has undergone post-translational phosphorylation modification, such as tyrosine kinases epidermal growth factor receptor (EGFR) and Janus kinase 2 (JAK2), and downstream kinases phosphorylating transcription factors, signal transducers, and transcription activators 3 (STAT3). This embodiment provides a method for measuring protein phosphorylation levels.

[0462] Antibodies specific to different target proteins are labeled with different double-stranded DNA tags. Figure 25 Cells of interest were incubated in an appropriate culture medium, with or without a test agent (such as a drug agent or potential drug agent). The cells were then washed and encapsulated in approximately 100 pL droplets using a microfluidic device, along with a dissolution buffer, a restriction enzyme BclI, a biotin-labeled antibody specific to a target domain lacking phosphorylation sites, and hydrogel beads (Abate, AR et al. (2009)) carrying barcoded primers linked to a DNA tag of a target-specific antibody labeled with a DNA tag.

[0463] Antibody labeling:

[0464] DNA tags are formed by annealing two commercially produced oligonucleotides having the following characteristics: a 4-nucleotide (nt) 5' extension with phosphate-capped ends, followed by a 30-nt random sequence, and then... The SBS3 sequence (SEQ ID NO:2) and a 10nt random sequence are present. This 4nt protrusion is common to all tags and allows the tag to attach to barcode-enabled hydrogel beads. The reverse sequence is terminated at the 5' end with an amino, aldehyde, or NHS-ester modifier, depending on the method used to conjugate the tag to the antibody, and has a 10nt 5' single-strand protrusion as a flexible linker. Different DNA tags can be generated through this same process by changing the 30nt random sequence. The sequences will be identical at both ends of the nucleotides, but different in the center, thus providing different tags for different target antibodies.

[0465] The annealed double-stranded DNA tag was then conjugated to the target-specific antibody as follows. Following Kozlov et al. (Kozlov, IA et al. - Efficient strategies for the conjugation of oligonucleotides to antibodies enabling highly sensitive protein detection. Biopolymers (2004). 73(5), 621-630. doi:10.1002 / bip.20009), a tag with a 5'4nt single-strand overhang and an aldehyde 5' modification (on the reverse sequence) was conjugated to the antibody via hydrazone bond formation. The antibody was first incubated in PBS with a 20-fold excess of 4-hydrazine succinimide acetone hydrazone (SANH; Solulink). It was then purified by size exclusion chromatography on an Illustra NAP-5 column (GE Healthcare) and resuspended in 100 mM citrate buffer at pH 6.0. The antibody was then incubated with a 10-fold excess of DNA tag and purified again using an Illustra (GE Healthcare) size exclusion column.

[0466] The final product is an antibody bound to a 5' aldehyde double-stranded DNA tag. The labeled target-specific antibodies are mixed in equal proportions before analysis.

[0467] Preparation of hydrogel beads and synthesis of barcoded oligonucleotides:

[0468] This is performed as in Example 1. However, the final double strand linked to the newly synthesized barcode is terminated with a 4nt 5' protrusion. This protrusion is complementary to the 5' 4nt protrusion common to all target-specific DNA tags, thus allowing it to be ligated to the hydrogel beads via T7 DNA ligase. Additionally, UMI is used to distinguish sequences derived from different target-specific antibodies with different DNA tags, rather than sequences from different RT-induced events.

[0469] Encoding phenotypes into DNA:

[0470] The hydrogel beads are aggregates of partially double-stranded DNA molecules, all possessing the same DNA barcode; this DNA barcode is unique on each hydrogel bead and terminates with a 4nt 5' protrusion. Before analysis, an antibody tagged with a DNA tag is ligated into the hydrogel beads via T7 ligase. This process links the target-specific tag sequence to the hydrogel bead barcode. This ligation of the DNA tag to the DNA barcode confers single-cell specificity to the target molecule.

[0471] These labeled antibodies are specific to the phosphorylated regions of the target protein. An antibody targeting the phosphorylated protein and an antibody targeting the unphosphorylated protein are used to encode specific tags for each state. A third antibody specific to the non-phosphorylated domains of the protein is used to label the binding complex. This antibody is biotin-labeled using the method described in the “Labeling of Target Antibody” section of Example 4. This antibody does not have a DNA tag and is encapsulated in all droplets. Biotin labeling allows for subsequent separation of the antibody-target complex from unbound antibody. Alternatively, the complex can be purified by size exclusion chromatography without the third antibody.

[0472] Using a microfluidic chip, single cells were encapsulated in droplets along with a dissolution buffer, BclI restriction enzyme, a biotinylated antibody as described above, and hydrogel beads carrying DNA barcodes linked to the DNA-tagged antibody. Droplets were generated using a Poisson distribution during cell encapsulation, and the cell concentration was chosen such that the average number of cells per droplet was <1, ensuring that most droplets contained no more than one cell. Deformable hydrogel beads were injected in a close-packed array (Abate, AR et al. (2009)) to ensure that most droplets contained a single bead. Figure 26 ).

[0473] The emulsion was then incubated at 37°C for 1 hour and 30 minutes. During this incubation, the BclI restriction enzyme lysed and released the barcoded and labeled antibodies into the entire volume of the droplet. Simultaneously, cells lysed, and the released target proteins were captured by these antibodies. These antibody-target complexes were captured by biotin-labeled antibodies. Figure 27 ).

[0474] Recovery of phosphorylated target proteins and separation of phosphorylated and unphosphorylated target proteins:

[0475] The emulsion was disrupted and the aqueous phase was recovered. Then it was passed through a streptavidin agarose column (…). Streptavidin agarose column) combines phosphorylated and unphosphorylated proteins with linked antibodies, DNA tags, and barcodes, along with unbound DNA-tagged antibodies. Figure 28 Separation. Alternatively, the complex can be purified by size exclusion chromatography.

[0476] Batch amplification and sequencing:

[0477] The barcoded DNA tags were purified using a commercial kit (Agencourt AMPure XP). Then, for tags derived from phosphorylated and unphosphorylated proteins, the barcoded tags were amplified in separate PCRs. The primers used matched the ends of the DNA tags and had a 5' extension containing… The sequences required for sequencing are the anchor sequences P7 (SEQ ID NO:11) and P5 (SEQ ID NO:12).

[0478] Single-cell-level target protein phosphorylation data were quantified by the number of target-specific DNA tag reads or UMIs for each DNA barcode and compared with the total target protein mass.

[0479] Example 6

[0480] Alternative protein phosphorylation analysis

[0481] In this embodiment, the target molecule is a protein that has undergone post-translational phosphorylation modification, such as tyrosine kinases epidermal growth factor receptor (EGFR) and Janus kinase 2 (JAK2), as well as downstream kinases that phosphorylate transcription factors, signal transducers, and activators of transcription 3 (STAT3). This embodiment provides a method for simultaneously measuring protein phosphorylation levels and performing targeted sequencing of mRNA to determine the target sequence and expression level.

[0482] Antibodies specific to different target proteins were labeled with different single-stranded DNA tags. Cells of interest were incubated in appropriate culture media, with or without test agents (such as pharmaceutical agents or potential pharmaceutical agents). Cells were then washed and encapsulated in approximately 100 pL droplets using a microfluidic device, along with DNA-tagged target-specific antibodies, reverse transcription buffer and enzymes (reverse transcriptase, DNA polymerase, and restriction enzyme BclI), biotin-labeled antibodies specific to target domains lacking phosphorylation sites, and hydrogel beads (Abate, AR et al. (2009)). The hydrogel beads carried a mixture of barcoded primers complementary to the DNA tag of the labeled target-specific antibody and the targeted mRNA.

[0483] Antibody labeling is performed as in Example 4, forming a DNA tag.

[0484] Preparation of hydrogel beads and synthesis of barcoded oligonucleotides:

[0485] This is done as in Example 1, where the UMI distinguishes between target-specific antibodies from different DNA-tagged sources and sequences from different RT-induced events.

[0486] Encoding phenotypes into DNA-DNA polymerization and RT within droplets:

[0487] Target-specific antibodies labeled with DNA tags are specific to the phosphorylated regions of target proteins. An antibody targeting the phosphorylated protein and an antibody targeting the unphosphorylated protein are used to encode specific tags for each state (see, for example...). Figure 29 The binding complex is labeled using a third antibody that is specific to the non-phosphorylated domains of the protein. This antibody is biotin-labeled using the method described in the “Labeling of Target Molecule” section of Example 4. This antibody does not have a DNA tag and is encapsulated in all droplets. Biotin labeling allows for subsequent separation of the antibody-target complex from unbound antibody. Alternatively, the complex can be purified by size exclusion chromatography without the use of a third antibody.

[0488] Using a microfluidic chip, single cells were encapsulated in droplets along with a DNA-tagged target-specific antibody, dissolution buffer, reverse transcriptase and its buffer, DNA polymerase, BclI restriction enzyme, a biotin-tagged antibody as described above, and hydrogel beads. The hydrogel beads carried aggregates of partially double-stranded DNA molecules, all possessing the same DNA barcode; this DNA barcode was unique on each hydrogel bead. The single-stranded portion of the DNA molecule bound to the hydrogel beads consisted, in sequence, of a DNA adapter sequence, a 3' antisense sequence to the DNA tag, or the targeting mRNA sequence JAK2. Droplets were generated using a Poisson distribution during cell encapsulation, and the cell concentration was selected such that the average number of cells per droplet was <1, ensuring that most droplets contained no more than one cell. Deformable hydrogel beads were injected in a close-packed array (Abate, AR et al. (2009)) to ensure that most droplets contained a single bead. Figure 29 , 30 ).

[0489] The emulsion was then incubated at 55°C for 1 hour and 30 minutes. During this incubation, BclI restriction enzymes cleaved the barcoded oligonucleotides and released them into the entire volume of the droplets. Simultaneously, cells lysed, and the released target proteins were captured by labeled antibodies, which annealed to complementary sequences on the single-stranded portions of the barcoded DNA bound to the hydrogel, and polymerases extended the barcoded DNA molecules, thereby copying the DNA tag. These antibody-target complexes were captured by biotin-labeled antibodies. Simultaneously, the released mRNA annealed to complementary sequences on the single-stranded portions of the released barcoded oligonucleotides, and RT enzymes extended the barcoded DNA, thereby copying the target mRNA sequence. Adding DNA barcoding to antibody-bound DNA tags confers single-cell specificity to these sequences. Figure 30 ).

[0490] Recovery of phosphorylated target proteins and separation of phosphorylated and unphosphorylated target proteins:

[0491] The emulsion was placed at 70°C to inactivate the polymerase, then the emulsion was disrupted and the aqueous phase was recovered. It was then passed through a streptavidin agarose column (…). (Streptavidin agarose column) separates phosphorylated and unphosphorylated proteins, along with linked antibodies, DNA tags, and barcodes, from unbound DNA-tagged antibodies and barcode-coded cDNA from RT. Figure 31 Alternatively, the complex can be purified by size exclusion chromatography.

[0492] Batch amplification and sequencing:

[0493] Barcoded DNA tags from the targeted protein were purified using a commercial kit (Agencourt AMPure XP). cDNA from RT DNA was purified using a commercial kit (Agencourt RNAClean XP). Then, as with cDNA from RT, the barcoded tags from both phosphorylated and unphosphorylated proteins were amplified in separate PCRs. Primers used matched the ends of the DNA tags and had a 5' extension containing… The sequences required for sequencing are the anchor sequences P7 (SEQ ID NO:11) and P5 (SEQ ID NO:12).

[0494] Single-cell level target protein phosphorylation data can be quantified by the number of target-specific DNA tag reads or UMIs for each DNA barcode, and compared with the total target protein amount. Single-cell level mRNA expression and sequence information can also be obtained from sequencing.

[0495] Example 7

[0496] Antibody profiling at the single-cell level:

[0497] This embodiment describes a method for sequencing two mRNAs encoding two peptides that constitute an antibody and provide it with specificity. In this embodiment, a mixture of antibody-producing cell lines (hybridomas) is encapsulated in approximately 100 μL droplets using a microfluidic device. The droplets comprise a mixture of culture medium, cell lysis and reverse transcription reagents, and hydrogel beads carrying a mixture of barcoded primers complementary to the heavy and light chain antibody gene mRNAs. After incubation at 55°C (which enables enzymatic digestion via BclI, RT of the mRNA, and DNA polymerization), the emulsion is chemically disrupted and the aqueous phase is recovered. The cDNA is purified using solid-phase reversible fixation (SPRI) beads and amplified by PCR prior to sequencing.

[0498] Preparation of hydrogel beads and functionalization for reverse transcription:

[0499] Hydrogel beads were prepared from PEG-DA oligomers in a microfluidic chip, wherein an aqueous PEG-DA solution was dispersed by hydrodynamic flow focusing to form microdroplets in a fluorinated oil continuous phase (Anna, S., Bontoux, N. and Stone, H. (2003). Formation of dispersions using “flow focusing” in microchannels. Applied Physics Letters, 52(3), 364-366. doi:10.1063 / 1.1537519). The beads were then crosslinked via a UV-activated photoinitiator. A double-stranded DNA oligonucleotide (double-stranded) called RanA, with a 5' acrydite modification at one end and a 4-nt 5' overhang on the other (top chain: 5'Acrydite-TCTTCACGGAACGA (SEQ ID NO:4); bottom chain: 5' phosphate-CAGTTCGTTCCGTGAAGA (SEQ ID NO:5)), was added and covalently cross-linked to the hydrogel matrix via the acrylate end groups of the PEG-DA oligomer. After polymerization and washing in Tris-HCl pH 7.4 20 mM; NaCl 50 mM; Tween 0.01%; EDTA 1 mM, a first double-stranded DNA ligase was used to ligate it. The first double-stranded DNA ligase had a 4-nt 5' overhang on one side compatible with the acrydite double-stranded DNA overhang, and another 4-nt 5' overhang on the other side compatible with the downstream linker. In its double-stranded portion, this first double-stranded form, called RanB, has a BclI restriction site (TGATCA (SEQ ID NO:13)) in the sequence, followed by a shorter form. Read 2 sequence (GTGTGCTCTTCCGATCT(SEQ ID NO:14)).

[0500] After ligation and washing, barcodes were synthesized by four consecutive ligations mediated by T7 DNA ligase using 20-nt DNA duplexes with 4-nt protrusions at both 5' ends. Different 4-nt protrusions were used in each ligation step to ensure that only four indices could be assembled in the correct order. To create a variety of barcodes, hydrogel bead batches were evenly distributed into the wells of a 96-well plate. Each well contained a duplex with a unique 20-nt sequence (index) designed to maintain a clear, at most three errors, along with ligation buffer and the enzyme. After ligation incubation, the entire reaction volume of the plate was pooled into a tube and washed. The next ligation step was performed in the same manner: the pooled batch was evenly distributed into a new plate containing another 96 different duplexes, along with ligation buffer and the enzyme. This separate-pool synthesis resulted in a combinatorial diversity of 96.4 (Over 84 million). Finally, the last double-stranded strand ligated to the newly synthesized barcode is partially double-stranded to allow ligation via T7 DNA ligase and terminates at a long single-stranded 3' end (see example). Figure 9 and 10 The double-stranded region is a defined adapter sequence (TACGCTACGGAACGA (SEQ ID NO:15)). The single-stranded region consists sequentially of a randomized 5-nt sequence (NNNNN (SEQ ID NO:16)) and the complementary sequence of the targeted gene. In this embodiment, we use two slightly degenerate sequences: HyLRT1 (TTGATTTCCAGCTTGGTCCC (SEQ ID NO:17)) designed to target the light chain mRNA of each hybridoma cell, and HyHRT1 (GGCCAGTGGATAGACYGATG (SEQ ID NO:17)) designed to target the heavy chain mRNA of each hybridoma cell. NO:18). Note that any published primer mix designed for maximum diversity targeting light and heavy chain mRNAs can be used. 5-nt random sequences act as unique molecular identifiers (UMIs): they allow the differentiation of sequences derived from different RT priming events (with different UMIs) from sequences derived from PCR amplification of the same cDNA (with the same UMI) (Shiroguchi, K., Jia, TZ, Sims, PA and Xie, XS (2012)). Digital RNA sequencing minimizes sequence-dependent bias and amplification noise with optimized single-molecule barcodes. Proceedings of the National Academy of Sciences, 709(4), 1347-1352. doi:10.1073 / pnas,1118018109).

[0501] The release of oligonucleotides via restriction enzyme cleavage can be achieved by replacing this sequence with any cleavable chemical group that can be linked to the nucleic acid (such as a photocleavable or pH-sensitive moiety).

[0502] Encoding the phenotype into the DNA within the droplet:

[0503] Using a microfluidic chip, 50,000 cells were encapsulated in microdroplets following a Poisson distribution at λ = 0.04 (0,008% double or more encapsulation events). The cell population consisted of a 1:1:1 mixture of hybridoma cell lines with known sequences.

[0504] The chip's triple-flow design brings together cells, RT enzymes, BclI restriction enzymes, and hydrogel beads at a nozzle that disperses them in 100pl droplets. These hydrogel beads carry aggregates of partially double-stranded DNA molecules. Figure 1 The single-stranded portion of the DNA bound to the hydrogel beads consists of a UMI sequence (SEQ ID NO:10) and an antisense sequence against the 3' end of the target mRNA.

[0505] The emulsion was then incubated at 55°C for 1 hour and 30 minutes. During this incubation, the BclI restriction enzyme releases the barcoded oligonucleotide into the entire volume of the droplet, where it encounters the mRNA released from the lysed cells. The two structures anneal to their complementary sequences, and the RT enzyme extends the barcoded DNA molecule, thereby copying the mRNA sequence. Figure 32 DNA barcodes are linked to cDNA, giving these sequences single-cell specificity.

[0506] Batch purification and amplification, as well as sequencing:

[0507] The emulsion was then placed at 70°C to inactivate the RT enzyme. After cooling on ice, the emulsion was disrupted and the aqueous phase was recovered, and the DNA was purified using Agencourt's RNAClean XP SPRI.

[0508] The barcoded cDNA was then amplified by PCR: a first 15-cycle PCR was performed using primer pairs from a 40% cDNA preparation, the primers matching both ends of the cDNA with a 5' extension, which introduced Illumina read 1 and completed read 2. The PCR was purified using Agencourt's AMPure XP SPRI, and the entire product was used for another 10-cycle PCR using primers P7-read 2 (CA AGC AGA AGAC GGC AT ACGAGA T GT GAC T GGAGTT C AGAC GT GT GCTCTTCCGATCT (SEQ ID NO:19)) and P5-read 1 (AATGATACGGCGACCACCGAGATCTACACTCTTTCCCTACACGACGCTCT TCCGATCT (SEQ ID NO:20)). Figure 33The results were purified again using Agencourt's AMPure XP SPRI, eluted in small volumes (10 to 20 μl), and deposited onto a 1.5% agarose gel. Bands of the expected size were purified from the gel (approximately 420 bp for light chain cDNA and approximately 500 bp for heavy chain cDNA) and mixed in equimolar ratios. The size purity of this final DNA preparation was examined on an Agilent Bioanalyzer and an Illumina MiSeq machine. Figure 34 ).

[0509] Next-generation sequencing (NGS) and bioinformatics analysis of data:

[0510] The sample run was performed with end-paired 2x300bp sequences: each processed molecule was sequenced from both ends and read 300 nucleotides inward. Given our product size, there is 50 to 100 nucleotide overlap between the reads at both ends.

[0511] Customized programs analyze data by clustering similar barcodes together, extracting associated cDNA sequences, and calculating the properties and frequencies of these sequences. Figure 35 With a strict threshold of 40 reads for each chain, 5102 pairs were obtained, of which 96.4% were correct (i.e., the heavy and light chains correspond to one of the three possible associations present in our three cell lines).

[0512] Example 8

[0513] Simultaneous reproduction of single-cell phenotype and transcriptional levels:

[0514] CytoSeq and RNAseq on B cells and T cells

[0515] The following examples demonstrate the reproduction of transcriptional and phenotypic information at the single-cell level using the methods described herein.

[0516] In this embodiment, cells are labeled using antibodies tagged with short RNA (or DNA) tags (Ab-tags). Each antibody's Ab-tag carries a unique sequence indicating its antigen specificity. After washing, the labeled cells are compartmentalized in a microfluidic system with hydrogel beads carrying barcoded primers, cell lysis reagent, reverse transcriptase (RT), and dNTPs. Most droplets will not contain more than one cell, and most droplets containing a single cell will also contain a single hydrogel bead carrying a primer with a unique barcode. The barcoded primers are released from the beads using a restriction enzyme (RE) and used to initiate cDNA synthesis using the Ab-tag as a template. The emulsion is then disrupted, the droplet contents are pooled, and the cDNA is purified to remove unpooled primers. The barcoded cDNA is then amplified by PCR to attach Illumina sequencing primer sites, and sequencing is performed using NextSeq or HiSeq 150nt operation to quantify the Ab-tags associated with each cell.

[0517] The tag can be RNA, DNA, or a hybrid substitution of RNA and DNA nucleotides; all of these structures can be used as templates by RT enzymes. Regardless of whether the template used is RNA, DNA, or a hybrid, the primer that elongates due to RT enzyme activity is consistently referred to as cDNA throughout the document. Note that for this example, any DNA polymerase can be used instead of an RT enzyme.

[0518] This description includes simultaneous reproducibility of phenotypic information (CytoSeq) and RNAseq methods used to analyze the transcriptional levels of 40 genes and a 50:50 mixture of surface markers from B lymphocytes and T lymphocytes derived from the human Ramos cell line (ATCC CRL-1596) and the human Jurkatt cell line (ATCC TIB-152). Four cell marker antibodies (CD1A, CD3D, CD72, and CD79B) were each conjugated to fluorescent antibody-specific oligonucleotides. T cells were preferentially labeled with CD1A and CD3D, while B cells were preferentially labeled with CD52 and CD79B. In this test, each barcode was fully correlated with either the T cell morphology (i.e., the combination of antibody-type-specific nucleotide sequences) or the B cell morphology; only two-cell events would result in a barcode with a mixed morphology.

[0519] Antibody-oligonucleotide conjugation:

[0520] Use a dedicated kit (e.g., Innova Bioscience) Oligonucleotide conjugation systems can covalently link antibodies to DNA or RNA oligonucleotides. Simply put, conjugation chemistry is based on both the modification of lysine residues in the antibody and the activation of the amine moiety on the oligonucleotide. The modified antibody and the activated oligonucleotide can react with each other to form a covalent bond (conjugation).

[0521] Double-stranded RNA oligonucleotides (duplexes) with several modifications specifically targeted to CytoSeq have been designed. Figure 36 The top strand sequence contains a BclI restriction site, followed by a short form. Read 1 sequence (ACACGACGCTCTTCCGATCT (SEQ ID NO:21)), followed by a 4nt UMI, then a different 6-nt sequence (antibody recognition sequence (AIS)) for each antibody, followed by an 18-nt sequence called TSP2 (TGAGTAAAGGAGAAGAAC (SEQ ID NO:22)) shared by each top chain. Four different top chains were synthesized, each with its own AIS: ACATCG; GATCTT; CTGAGC and TGCGAA to recognize CD1A, CD3D, CD72 and CD79B antibodies, respectively. The bottom chains have a complementary sequence to the top chain at the BclI site and a short... This is a portion of the read 1 sequence. It contains a 3' amine modification necessary for antibody conjugation, a fluorophore linked to its 5' end, and a BclI restriction site. Two different forms of this bottom chain were identified, one with the fluorophore 6-FAM (excited at 495 nm) and the other with the fluorophore TYE-655 (excited at 665 nm).

[0522] Prior to conjugation, the goal was to anneal the top chain of B cell-specific antibodies (CD72 and CD79B) to the bottom chain of 6-FAM, while annealing the T cell-specific antibodies (CD1A and CD3D) to the bottom chain of TYE-665. Conjugation efficiency was validated on SDS-PAGE, and cell marker efficacy was examined using a hematology counter.

[0523] Cell markers:

[0524] Before encapsulation into droplets, the mixture of B and T lymphocytes was labeled with a mixture of four labeled antibodies according to a standard labeling protocol: 10 6 Up to 10 7 The cells were washed in staining buffer (PBS, 0.5% bovine serum albumin (BSA), 2 mM EDTA), incubated in the dark and at 4°C in 100 μl of 4 μg labeled antibody for 10 min, and finally washed twice with 1 ml staining buffer.

[0525] Preparation of hydrogel beads and functionalization for reverse transcription:

[0526] Hydrogel beads were prepared from PEG-DA oligomers in a microfluidic chip, wherein an aqueous PEG-DA solution was dispersed by hydrodynamic flow focusing to form microdroplets in a fluorinated oil continuous phase (Anna, S., Bontoux, N. and Stone, H. (2003). Formation of dispersions using “flow focusing” in microchannels. Applied Physics Letters, 52(3), 364-366. doi:10.1063 / 1.1537519). The beads were then crosslinked via a UV-activated photoinitiator. A double-stranded DNA oligonucleotide (double-stranded) called RanA, with a 5' acrydite modification at one end and a 4-nt 5' overhang on the other side (top chain: 5'Acrydite-TCTTCACGGAACGA (SEQ ID NO:4); bottom chain: 5' phosphate-CAGTTCGTTCCGTGAAGA (SEQ ID NO:5)), was added and covalently cross-linked to the hydrogel matrix via the acrylate end groups of the PEG-DA oligomer. After polymerization and washing in Tris-HCl pH 7.4 20 mM; NaCl 50 mM; Tween 0.01%; EDTA 1 mM, a first double-stranded DNA oligomer was ligated using T7 DNA ligase. The first double-stranded DNA oligomer had a 4-nt 5' overhang on one side compatible with the acrydite double-stranded DNA overhang, and another 4-nt 5' overhang on the other side compatible with the downstream linker. In its double-stranded portion, this first double-stranded form, called RanB, has a BclI restriction site (TGATCA (SEQ ID NO:7)) in the sequence, followed by a shorter form. Read 2 sequence (GTGTGCTCTTCCGATCT(SEQ ID NO:23)).

[0527] After ligation and washing, barcodes were synthesized by four consecutive ligations mediated by T7 DNA ligase using 20-nt DNA duplexes with 4-nt protrusions at both 5' ends. Different 4-nt protrusions were used in each ligation step to ensure that only four indices could be assembled in the correct order. To create a variety of barcodes, hydrogel bead batches were evenly distributed into the wells of a 96-well plate. Each well contained a duplex with a unique 20-nt sequence (index) designed to maintain a clear, at most three errors, along with ligation buffer and the enzyme. After ligation incubation, the entire reaction volume of the plate was pooled into a tube and washed. The next ligation step was performed in the same manner: the pooled batch was evenly distributed into a new plate containing another 96 different duplexes, along with ligation buffer and the enzyme. This separate-pool synthesis resulted in a combinatorial diversity of 96. 4 (Over 84 million). Finally, the last double-stranded strand ligated to the newly synthesized barcode is partially double-stranded to allow ligation via T7 DNA ligase and terminates at a long single-stranded 3' end (see example). Figure 9 and 10 The double-stranded region is a defined adapter sequence (TACGCTACGGAACGA (SEQ ID NO:9)). The single-stranded region consists of a randomized 5-nt sequence (NNNNN (SEQ ID NO:24)) and an antisense sequence of ASP2 for initiating RT (using an Ab-tag as a template) (GTTCTTCTCCTTTACTCA (SEQ ID NO:25)). The 5-nt random sequences act as unique molecular identifiers (UMIs): they allow the differentiation between sequences originating from different RT initiation events (with different UMIs) and sequences amplified by PCR from the same cDNA (with the same UMI) (Shiroguchi, K., Jia, TZ, Sims, PA, and Xie, XS (2012). Digital RNA sequencing minimizes sequence-dependent bias and amplification noise with optimized single-molecule barcodes. Proceedings of the National Academy of Sciences of Sciences, 109(4),1347-1352.doi:10.1073 / pnas.1118018109).

[0528] The release of oligonucleotides via restriction enzyme cleavage can be achieved by replacing this sequence with any cleavable chemical group that can be linked to the nucleic acid (such as a photocleavable or pH-sensitive moiety).

[0529] Encoding the phenotype into the DNA within the droplet:

[0530] Using a microfluidic chip, labeled cells, along with RT enzymes, BclI restriction enzymes, and hydrogel beads, were encapsulated in microdroplets. The hydrogel beads carried aggregates of partially double-stranded DNA molecules. Figure 18 The single-stranded portion of the DNA bound to the hydrogel beads consists of a UMI sequence (SEQ ID NO:10) and an antisense sequence to the 3' end of the Ab-tag.

[0531] The emulsion was then incubated at 55°C for 1 hour and 30 minutes. During this incubation, the BclI restriction enzyme cleaved the barcoded oligonucleotide and Ab-tag and released them into the entire volume of the droplet. These two structures annealed at their complementary sequences (TSP2 sense and antisense), and the RT enzyme extended the barcoded DNA molecule, thereby copying the Ab-tag (…). Figure 19 DNA barcoding linked to Ab-tags gives these sequences single-cell specificity.

[0532] Batch purification and amplification, as well as sequencing:

[0533] The emulsion was then placed at 70°C to inactivate the RT enzyme. After cooling on ice, the emulsion was disrupted and the aqueous phase was recovered. The DNA was then purified using a commercial kit (Agencourt RNAClean XP).

[0534] The barcode-bearing tag was then amplified by PCR. The primers used matched the ends of the tag and had a 5' extension containing... The sequences required for sequencing, namely the anchor sequences P7 (SEQ ID NO:11) and P5 (SEQ ID NO:12) Figure 37 ).

[0535] Method sensitivity:

[0536] This method was performed using cells labeled with free, soluble Ab-tags instead of Ab-tags. The concentration of Ab-tags in the droplets was 100 pM, corresponding to approximately 6000 Ab-tag molecules per droplet. Ab-tag solutions introduced into the droplets at various ratios (1:1 (twice); 7:3; 9:1 (twice); or 99:1) were mixtures of two different Ab-tags (Ab-tag 1 and Ab-tag 2 with different AIS). The amplified PCR products were purified on 2% agarose gels, and size purity was examined on an Agilent Bioanalyzer. Figure 38The data was then fed into an Illumina NextSeq machine. The operation on the sample was a single 150bp read, covering the entire barcode and AIS. Results showed a good correlation between the Ab-tag ratio in the input (before encapsulation) and in the NGS data (read distribution). Figure 39 Notably, we were able to detect two types of Ab-tags in the data from the 99:1 ratio. This corresponds to a sensitivity of less than 70 molecules.

[0537] All publications, patents, and patent applications mentioned herein are incorporated herein by reference to the same extent that if specifically and individually indicated, each individual publication, patent, or patent application were incorporated by full reference. In the event of any discrepancy between the definitions set forth herein and those in the documents incorporated herein by reference, the definitions set forth herein shall prevail. Various modifications and variations to the methods, pharmaceutical compositions, and kits described in this disclosure will be apparent to those skilled in the art without departing from the scope and spirit of the invention. While the invention has been described in conjunction with specific embodiments, it should be understood that further modifications are possible and the invention, as claimed, should not be unduly limited to such specific embodiments. Indeed, the various modifications to the modes of carrying out the invention described, which will be apparent to those skilled in the art, are intended to be within the scope of the invention. This application is intended to cover any changes, uses, or modifications to the invention that generally follow the principles of the invention, and includes any deviations from known conventions in the field to which this disclosure pertains and which may apply to the essential features set forth above herein.

Claims

1. A method of assigning a binding phenotype of an expressed polypeptide to a potential genotype, expression level, or both, comprising: partitioning a single cell or a portion of a non-cellular system from a sample comprising a cell, a population of cells, or a non-cellular system into individual compartments; wherein the single cell or portion of a non-cellular system in the individual compartments comprises an expressed target polypeptide and an expressed target nucleic acid, and the target nucleic acid encodes the corresponding target polypeptide; wherein the individual compartments comprise a source-specific barcode and a polypeptide capture molecule, wherein the source-specific barcode comprises a unique nucleic acid sequence that identifies the individual compartment, and the polypeptide capture molecule comprises a capture molecule nucleic acid identifier that identifies the polypeptide capture molecule; allowing the expressed target polypeptide in each individual compartment to bind to the polypeptide capture molecule to produce a bound form target polypeptide-polypeptide capture molecule complex; labeling the expressed target nucleic acid and the capture molecule specific nucleic acid identifier of the bound form polypeptide capture molecule with the source-specific barcode to produce a barcoded expressed target nucleic acid and a barcoded polypeptide capture molecule; detecting the sequence of the barcoded expressed target nucleic acid and barcoded capture molecule nucleic acid identifier; grouping the expressed target nucleic acid and the expressed target polypeptide according to a common source-specific barcode, thereby identifying the type of polypeptide capture molecule bound by the expressed target polypeptide in an individual compartment and the expressed target nucleic acid in the same individual compartment; wherein the source-specific barcode comprises a first source-specific barcode type comprising a target nucleic acid binding sequence; and a second source-specific barcode type comprising a capture molecule nucleic acid identifier binding sequence, wherein for a given individual compartment, the first and second types of source-specific barcode comprise the same or matching unique nucleic acid sequence that identifies the individual compartment.

2. The method of claim 1, wherein each individual compartment comprises a single cell.

3. The method of claim 1 or 2, wherein the polypeptide capture molecule is allowed to bind to the expressed target polypeptide on the surface of a cell or population of cells prior to partitioning the sample or a portion thereof into individual compartments.

4. The method of any one of claims 1 to 3, further comprising lysing the cell or population of cells prior to allowing the expressed target polypeptide to bind to the polypeptide capture molecule.

5. The method of any one of claims 1 to 4, further comprising pooling all of the samples prior to detecting the sequence of the labeled expressed target nucleic acid and labeled capture molecule nucleic acid identifier.

6. The method of any one of claims 1 to 5, further comprising purifying the barcoded expressed target nucleic acid and barcoded capture molecule nucleic acid identifier prior to detecting the sequence of the barcoded target nucleic acid and barcoded capture molecule nucleic acid identifier.

7. The method of claim 1, wherein the origin-specific barcode further comprises one or more primer sequences, sequencing adaptors, one or more restriction sites, a capture moiety for facilitating enrichment of the origin-specific barcode from the sample, or a combination thereof.

8. The method of claim 7, wherein the primer sequence is a universal primer sequence.

9. The method of claim 1, wherein the origin-specific barcode comprises RNA, DNA, or a combination of RNA and DNA.

10. The method of any one of claims 1 to 6, wherein labeling the expressed target nucleic acids and the capture molecule nucleic acid identifiers comprises introducing to the individual compartments reagents sufficient to allow the first and second types of origin-specific barcodes to hybridize to the target nucleic acids and the capture molecule nucleic acid identifiers of the polypeptide capture molecules, respectively, and to serve as templates to generate cDNA copies of all or a portion of the target nucleic acids and capture molecule specific nucleic acid identifiers, such that sequences of oligonucleotide barcodes are incorporated in each target nucleic acid cDNA product and capture molecule nucleic acid identifier cDNA product.

11. The method of claim 10, comprising amplifying the cDNA products and detecting sequences of the amplified cDNA products.

12. The method of any one of claims 1 to 11, wherein the origin-specific barcodes are reversibly or irreversibly attached to a solid substrate.

13. The method of claim 12, wherein the origin-specific barcodes are attached to the solid substrate through a barcode receiving adaptor attached to the surface of the solid substrate.

14. The method of claim 13, wherein the barcode adaptor is a nucleic acid sequence complementary to an adaptor binding sequence on the origin-specific oligonucleotide.

15. The method of claim 12, wherein the solid substrate is a hydrogel bead.

16. The method of claim 12 or 15, wherein the oligonucleotide barcodes are released from the solid substrate prior to labeling the oligonucleotide tags of the expressed target nucleic acids and the polypeptide capture molecules.

17. The method of claim 16, wherein the oligonucleotide barcodes are released from the solid substrate by introducing to each individual compartment conditions sufficient to cause the solid substrate to dissolve or disintegrate or by chemically, photochemically, or enzymatically cleaving the origin-specific barcodes from the solid substrate.

18. The method of any one of claims 1 to 17, wherein the polypeptide capture molecules comprise small molecules, antigens, antibodies, protein binding domains, nucleic acids, or polysaccharides.

19. The method of claim 18, wherein the polypeptide capture molecules can discriminate between post-translational modifications of the target polypeptides.

20. The method of any one of claims 1 to 18, wherein the polypeptide capture molecules are antibodies specific for the corresponding target polypeptides.

21. The method of any one of claims 1 to 18, wherein the target polypeptide is displayed on the surface of the cell, and the polypeptide capture molecule is a binding partner of the target polypeptide.

22. The method of claim 21, wherein the target polypeptide is a cell surface receptor, and the target nucleic acid is an mRNA encoding the cell surface receptor.

23. The method of any one of claims 1 to 18, wherein the target polypeptide is an antibody expressed by the cell, and the polypeptide binding molecule is an antigen for the corresponding antibody.

24. The method of claim 23, wherein the target nucleic acid comprises an mRNA encoding a light chain, an mRNA encoding a heavy chain, an mRNA encoding a CDR, or a combination thereof.

25. The method of claim 23 or 24, wherein the cell is a B cell, a T cell, a plasmablast, or a plasma cell.

26. The method of any one of claims 1 to 25, wherein the polypeptide capture molecule identifier is directly or indirectly conjugated to the polypeptide capture molecule.

27. The method of claim 26, wherein the polypeptide capture molecule identifier is indirectly bound via a binding pair, wherein a first member of the binding pair is part of or linked to the polypeptide capture molecule and a second member of the binding pair is part of or linked to the capture molecule identifier.

28. The method of claim 27, wherein the binding pair is streptavidin-biotin.

29. The method of claim 26, wherein the polypeptide capture molecule and the capture molecule identifier are bound to a common substrate.

30. The method of claim 29, wherein the polypeptide capture molecule and the capture molecule identifier are biotinylated and bound to a common streptavidin substrate.

31. The method of claim 30, wherein multiple copies of the capture molecule identifier are bound to the common streptavidin substrate.

32. The method of any one of claims 1 to 31, wherein the individual compartments are single droplets generated on a microfluidic device.

33. The method of any one of claims 1 to 32, wherein the single droplet is formed by merging a first droplet comprising the cell, cell population, or non-cellular system with a second droplet comprising the origin-specific barcode.

34. The method of claim 33, wherein the origin-specific barcode in the second droplet is bound to a single solid substrate.

35. The method of claim 32, further comprising merging the single droplet with a third droplet comprising additional reagents.

36. The method of claim 35, wherein the third droplet comprises one or more of the following: a cell lysis reagent, a reverse transcription reagent, a restriction enzyme for releasing the origin-specific barcode from a solid substrate, dNTPs, and a DNA polymerase.

37. The method of claim 32, wherein additional reagents are injected into the single droplet.

38. The method of claim 37, wherein the additional reagents are one or more of: a cell lysis reagent, a reverse transcription reagent, a restriction enzyme for releasing the origin- specific barcodes from a solid substrate, dNTPs, and a DNA polymerase.

39. The method of any one of claims 1-38, further comprising introducing into each individual compartment a second type of protein capture molecule that has the same binding affinity for a target polypeptide as the original protein capture molecule and comprises a capture moiety, wherein the second protein capture molecule binds the same target polypeptide to form a sandwich complex with the original protein capture molecule, wherein the sandwich complex is purifiable from the individual compartment or pooled individual compartments by the capture moiety.

40. The method of any one of claims 1-39, wherein the expression levels of the target nucleic acids and target polypeptides are determined based at least in part on the detected origin- specific barcodes.

Citation Information

Patent Citations

  • Method for detecting a target nucleic acid sequence

    EP0320308A2

  • In vitro evolution in microfluidic systems

    US20060078888A1

  • Methods for obtaining a sequence

    US20130079231A1

  • Enzyme amplification assay

    US3817837A

  • Process for the demonstration and determination of low molecular compounds and of proteins capable of binding these compounds specifically

    US3850752A