Assessment of mRNA transcript levels using ddPCR for early pool and single cell clone assessment in cell line development

By using droplet digital PCR technology to measure mRNA transcript levels and selecting high-expression single-cell clones and cell pools, the problem of high-cost production of biopharmaceuticals in existing technologies has been solved, achieving efficient and low-cost production of biopharmaceuticals.

CN122055458APending Publication Date: 2026-05-15AMGEN INC
View PDF 45 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
AMGEN INC
Filing Date
2024-10-17
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently select single-cell clones and cell pools that produce large quantities of the target protein, resulting in high production costs for biopharmaceuticals, particularly posing challenges in terms of speed to market and cost-effectiveness of biosimilars.

Method used

mRNA transcript levels were measured using droplet digital PCR (ddPCR) technology. cDNA was generated through reverse transcription and quantitative analysis was performed. Single-cell clones or cell pools that did not meet the transcript level ratio were excluded, and cell pools with excellent expression characteristics were selected for fed-batch culture.

Benefits of technology

By assessing transcript levels early, titers and product quality during the production process can be predicted, saving time, costs, and resources, improving the efficiency of high-throughput cell line development, and reducing the cost of biopharmaceutical production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122055458A_ABST
    Figure CN122055458A_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of cell culture. It provides a method that facilitates the selection of single cell clones or cell pools for the manufacture of antigen binding proteins having two to four different antibody chains. The method utilizes ddPCR to quantify the transcript level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application generally relates to methods for selecting single-cell clones and cell pools that produce large quantities of a target protein. More specifically, this application relates to the use of digital PCR techniques (such as droplet digital PCR (ddPCR)) for measuring mRNA transcript levels to select high-expression single-cell clones and cell pools. Background Technology

[0002] Biologics are used worldwide for a variety of applications, such as therapy and diagnostics, due to their wide range of uses. Mammalian cell lines are the primary expression systems for these biologics, with Chinese hamster ovary (CHO) cells being the main cell factory. See Lalonde et al., 2017, J Biotechnol [Biotechnology Journal] 251:128-140. In particular, with the advent of biosimilars, speed to market and cost-effectiveness are now more important than ever.

[0003] The cost of manufacturing biologics is high due to the complexity of their production, which involves multi-step processes including selecting optimal cell lines, mass-culturing cells, and purifying the desired biologic from the cell harvest. While these costs are decreasing due to improvements in various aspects of manufacturing, they can still be prohibitive when widely adopted as first-line therapies.

[0004] To make biologics more accessible to patients, reducing the commodity cost of the manufacturing process is an attractive proposition. One area of ​​particular interest is the selection of optimal cell lines. Cell line development (CLD) plays a crucial role in generating stable, high-expression cell lines for the production of biologics. In a platform-based CLD workflow, single-cell clones and stable cell pools are generated from transfected CHO cells, followed by evaluation of protein production post-production and submission for titer and product quality analysis—a process that can take several weeks. With the increasing number of novel, complex molecules formed from difficult-to-express domains or the assembly of several polypeptide chains (such as antibody chains), it is necessary to investigate several variables—such as scaffold design, the nucleotide sequence of the target protein, carrier elements, and culture medium or process parameters—to increase productivity, thereby generating and screening large numbers of pools or single-cell clones.

[0005] One method for identifying highly expressed cell clones is automated titer measurement. See European Patent Application No. EP1901068 A1. Another method for identifying optimal cell clones is to apply next-generation sequencing (NGS) to genomic DNA to identify the genomic location of transgene integration. See Stadermann et al., 2021, Biotechnology & Bioengineering 119:868-880.

[0006] Methods for selecting single-cell clones and pools that produce large quantities of the target protein are still needed. Summary of the Invention

[0007] This disclosure relates to a method for selecting single-cell clones or cell pools for manufacturing antigen-binding proteins having two to four different antibody chains, the method comprising: a) passage a single-cell clone or cell pool expressing antibody chains such that the cells have at least 85% viability and a doubling time of less than 40 hours for at least two consecutive passages; b) extracting mRNA from the cells; c) reverse transcribing the mRNA to generate cDNA; d) performing ddPCR to quantify the transcript levels of the antibody chains; and e) excluding single-cell clones or cell pools having at least the lowest 30th percentile of the sum of transcript levels of all antibody chains.

[0008] In some embodiments, at least one of the antibody chains is a heavy chain or heavy chain-ScFv, and / or at least one of the antibody chains is a light chain.

[0009] In some embodiments, a single-cell clone or cell pool expresses an antigen-binding protein having two distinct antibody chains. In one aspect of this embodiment, the single-cell clone or cell pool expresses an antigen-binding protein having one heavy chain (HC) or heavy chain-scFv (HC-scFv) and one light chain (LC). In one sub-aspect, the method further includes excluding cell clones or cell pools expressing an antigen-binding protein having one heavy chain (or HC-scFv) and one light chain, wherein the transcript level ratio of HC:LC (or HC-scFv:LC) or LC-HC (or LC-HC-scFv) is less than 0.2 or greater than 1.5. This excludes cell clones or cell pools with either extreme ratio (i.e., too low or too high).

[0010] In some embodiments, single-cell clones or cell pools express antigen-binding proteins having three or four different antibody chains. In one aspect of this embodiment, cells express antigen-binding proteins having two heavy chains (e.g., HC1 and HC2, or HC1 and HC2-scFv) and one light chain (LC). In one sub-aspect, the method further includes excluding cell clones or cell pools where the transcript level ratio of HC1:HC2 or HC2:HC1 is less than 0.2 or greater than 1.5, or the ratio of HC1:HC2-scFv or HC2-scFv:HC1 is less than 0.2 or greater than 1.5. In one aspect of this embodiment, cells express antigen-binding proteins having two heavy chains (HC1 and HC2) and two light chains (LC1 and LC2). In one sub-aspect, the method further includes excluding cell clones or cell pools where the transcript level ratio of HC1:HC2 or HC2:HC1 is less than 0.2 or greater than 1.5.

[0011] In any of the above embodiments, excluding single-cell clones or cell pools excludes single-cell clones or cell pools having at least the lowest 50th percentile of the sum of all strand transcript levels.

[0012] In some embodiments, the method further includes f) subjecting the remaining single-cell clones or cell pools to fed batch culture for at least an additional 8 days while maintaining cell viability above 80%.

[0013] In some embodiments, ddPCR is a one-step ddPCR or a two-step ddPCR.

[0014] In some embodiments, the cells are selected from the group consisting of CHO, HEK293, NSO, or Sp2 / 0 cells. In one aspect of this embodiment, the cells are CHO cells. In one sub-aspect of this embodiment, the CHO cells are DHFR- (dihydrofolate reductase deficiency) or GSKO (glutamine synthase knockout) types.

[0015] In some embodiments, the cells are generated using single-cell printing.

[0016] This disclosure also relates to a method for selecting single-cell clones or cell pools expressing antigen-binding proteins having one or more antibody heavy chains and one or more antibody light chains to achieve stable growth, the method comprising: a) passage the cells until they have greater than 80% viability and a doubling time of less than 35 to 40 hours, for at least two passages; b) extracting mRNA from these cells; c) reverse transcribing the mRNA to generate cDNA; d) performing ddPCR to quantify the transcript levels of one or more heavy chains and one or more light chains; and e) excluding cell clones or cell pools having a total transcript level of all antibody chains at least at the lowest 30th percentile. In one embodiment, the doubling time has a deviation of less than 50% between at least two passages. Attached Figure Description

[0017] Figure 1A -C: Transcription levels and titers for various antibody forms (feedback batches). For (A) 4-chain heterologous IgG, (B) 3-chain asymmetric antibodies with only HC binding (AmAb-HCAb or [Fab]). VH]-heterologous Fc), and (C)3-chain asymmetric C1mAb, summation of fed-batch transcript expression is represented by gray bars, and normalized total titers are represented by black bars. Normalized titer values ​​are obtained by dividing each total titer value by the lowest titer value in the sample set. The best pool performers selected based on titer and product quality attributes are marked with a black asterisk.

[0018] Figure 2A -C: Transcriptions and titers for various antibody forms (normal passages). For (A) 4-chain heterologous IgG, (B) AmAb-HCAb, and (C) 3-chain asymmetric C1mAb, total normal passage transcript expression is represented by light gray bars, total fed-batch expression by gray bars, and normalized total fed-batch titers by black bars. Normalized titer values ​​are obtained by dividing each total titer value by the lowest titer value in the sample set. The best pool performers (marked with a black asterisk) selected based on titer and product quality attributes are ranked in the top two-thirds of the transcript expression range.

[0019] Figure 3A -B: Spearman correlation matrix between 3-strand AmAb-HCAb molecular transcript / chain ratio and product quality (where HC2 is the shorter heavy chain of 3-strand AmAb-HCAb). Spearman analysis was performed to identify the correlation between transcript expression and fed-batch titers and product quality attributes from (A) fed-batch culture time points and (B) normal subculture time points.

[0020] Figure 4Transcripts of bispecific molecules (B2HmAb, C2HmAb, B2mAb, and C2mAb) with different molecular structures but similar binding properties, evaluated from normal passage time points. The sum of heavy chain (HC) and light chain (LC) transcripts is represented by light gray bars, and the normalized total feed titer is represented by black bars. The selected best pool is marked with a black asterisk, and the alternative pools are marked with a light gray asterisk. Normalized titer values ​​are obtained by dividing each total titer value by the lowest titer value in the sample set. Each bar represents the average of four pools with the same vector configuration, and each pool has three replicate expression values; error bars represent standard deviations.

[0021] Figure 5A -D: Clonal transcript expression analysis of 3-strand AmAb-HCAb from normal passage time points using rapid RNA extraction methods. (A) Total clonal expression is shown in gray bars, and normalized total titers are shown in black bars, with the best pool performers (marked with a black asterisk). Error bars represent the standard deviation of three replicates. (B) Individual strand expression titer results; the clone with the lowest titer is marked with a light gray asterisk. (C) Spearman correlation matrix providing the HC1 / HC2 ratio (where HC2 is the shorter heavy strand of 3-strand AmAb-HCAb). (D) HC1:HC2 ratio with MP and LMW. Best clones are marked with black squares. Normalized titer values ​​are obtained by dividing each total titer value by the lowest titer value in the sample set. Detailed Implementation

[0022] This invention is partly based on the finding that droplet-based digital PCR (ddPCR) quantification of mRNA levels in specific cells is correlated with the titer of secreted proteins. To simplify this highly resource-intensive cloning process, a strategy using one-step or two-step RT-ddPCR for single-cell cloning and / or early pool assessment is employed. One-step RT-ddPCR is a sensitive and high-throughput assay performed in a single reaction vessel, which can determine the mRNA transcript levels of a target gene (e.g., each gene encodes both heavy and light chains) in a short time. Two-step ddPCR separates the RT and PCR steps in two different reaction vessels. To assess whether transcriptomic analysis of cultures derived from single-cell clones and / or stable pools at early time points can serve as a predictor of titers and product quality in the production process, mRNA levels were analyzed from samples from single-cell clones or normal passaged pools during the fed-batch process, and compared with protein titers and final product quality results (e.g., size exclusion chromatography (SEC) high molecular weight (HMW) % and / or low molecular weight (LMW) %; non-reducing capillary electrophoresis (nr-CE) pre-peak % and / or post-peak %). This assay can serve as a powerful tool for predicting best-performing single-cell clones and / or pools and for excluding poorly performing single-cell clones and / or pools before reaching the fed-batch production stage. Because this assay will further enable next-generation high-throughput cell line development (CLD) workflows before reaching the fed-batch production stage, it has the potential to save time, cost, and resources.

[0023] In the production of biological agents, this invention has discovered a specific utility in evaluating transcript levels of light and heavy chains (from products containing both heavy and light chains), or combinations of three or four heavy or light chains, in different vector configurations to find the optimal vector configuration.

[0024] Digital PCR is a method for the amplification and detection of nucleic acids based on diluting template DNA into independent, non-interacting partitions. See Sykes et al., 1992, BioTechniques 13: 444-449. Following Poisson statistics of highly diluted DNA templates, the presence of nucleic acids in each reaction is independently queried with single-molecule sensitivity. In recent years, several companies have introduced commercially available methods for automation and expanding partition range. This includes droplet-based digital PCR (ddPCR) systems (e.g., Bio-Rad QX200), which randomly disperse template DNA into equal-volume emulsion droplets. See Hindson et al., 2011, Anal. Chem. 83:8604-8610.

[0025] Recently, digital PCR has been more widely used as an analytical tool for research and clinical applications. For example, digital PCR can be used as a robust tool for analyzing copy number variations seen in specific gene amplifications or deletions, detecting mutations, and quantifying specific nucleic acid types.

[0026] Although the terminology used herein is standard in the art, definitions of certain terms are provided herein to ensure clarity and definiteness of the meaning of the claims. Units, prefixes, and symbols may be expressed in their SI (International System of Units) acceptable forms. The numerical ranges enumerated herein include the numbers defining the ranges and include and support every integer within the defined ranges. Unless otherwise indicated, the methods and techniques described herein may be performed according to conventional methods well known in the art and as described in the various general and more specific references cited and discussed throughout this specification. See, for example, Sambrook et al., Molecular Cloning: A Laboratory Manual, 3rd ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York (2001); Ausubel et al., Current Protocols in Molecular Biology, Greene Publishing Associates (1992); and Harlow and Lane, Antibodies: A Laboratory Manual, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York (1990).

[0027] Furthermore, unless the context otherwise requires, singular terms should include plural terms and plural terms should include singular terms. Generally, the nomenclature and techniques used in conjunction with the cell and tissue culture, molecular biology, immunology, microbiology, genetics, and protein and nucleic acid chemistry and hybridization described herein are those well-known and commonly used in the art.

[0028] All documents or portions thereof (including, but not limited to, patents, patent applications, articles, books, and monographs) referenced in this application are expressly incorporated herein by reference. The content described in the embodiments of the invention may be combined with other embodiments of the invention.

[0029] This disclosure provides methods for selecting single-cell clones and / or cell pools expressing high levels of a “target protein.” A “target protein” includes naturally occurring proteins, recombinant proteins, and engineered proteins (e.g., proteins that do not exist in nature and have been designed and / or produced by humans). The target protein may be, but does not necessarily have to be, a protein known or suspected of having therapeutic relevance.

[0030] As used herein, the term "heterologous" in conjunction with nucleic acids means having nucleic acids that are not naturally present in host cells. This can include mutated sequences, such as sequences different from those naturally occurring. This can include sequences from other species. This can also include having the sequence at a different location in the genome than the sequence naturally occurring in the host cell. This generally does not include natural mutations that may occur in the host cell. Cells that already contain heterologous nucleic acids encoding a target protein, for example through stable integration of an expression cassette, will be considered to contain heterologous nucleic acid sequences. For clarity, CHO cells or derivatives thereof (e.g., DHFR- or GS knockout types) containing nucleic acids encoding antigen-binding proteins will be considered to contain heterologous nucleic acids.

[0031] As used herein, “cell culture” or “culture” or “normal passage” refers to the growth and proliferation of cells outside a multicellular organism or tissue. Suitable culture conditions for mammalian cells are known in the art. See, for example, *Animal cell culture: A Practical Approach*, edited by D. Rickwood, Oxford University Press, New York (1992). Mammalian cells can be cultured in suspension or attached to a solid substrate. Fluidized bed bioreactors, hollow fiber bioreactors, roller flasks, shake flasks, or stirred tank bioreactors with or without microcarriers can be used.

[0032] The phrase “cell culture medium” (also known as “culture medium”, “cell culture media”, or “tissue culture medium”) refers to any nutrient solution used to grow cells (e.g., animal or mammalian cells) and typically provides at least one or more of the following components: energy (usually in the form of carbohydrates such as glucose); one or more of all essential amino acids, typically twenty basic amino acids plus cysteine; vitamins and / or other organic compounds typically required in low concentrations; lipids or free fatty acids; and trace elements, such as inorganic compounds or naturally occurring elements, typically required in very low concentrations (usually in the micromolar range).

[0033] "Feed-batch culture" refers to a form of suspension culture and means a method of culturing cells in which additional components are provided to the culture at one or more points after the start of the culture process. The provided components typically include nutrient supplements that have been depleted by the cells during the culture process. Alternatively, the additional components may include supplemental components (e.g., cell cycle inhibitory compounds). Fed-batch cultures typically stop at a certain point, and the cells and / or components in the culture medium are harvested and optionally purified.

[0034] "Cell density" refers to the number of cells in a given volume of culture medium. "Viable cell density" refers to the number of viable cells in a given volume of culture medium, as determined by standard viability assays (such as trypan blue exclusion assay).

[0035] "Cell viability" refers to the ability of cells in a culture to survive under a given set of culture conditions or experimental variations. The term also refers to the number of living cells at a given time, relative to the total number of cells (living and dead) in the culture at that point.

[0036] "Titer" refers to the total amount of a target polypeptide or protein (which may be naturally occurring or recombinant) produced by a cell culture in a given volume of culture medium. Titer can be expressed in milligrams or micrograms per milliliter of culture medium (or other volumetric measure). "Cumulative titer" is the titer produced by the cells during the culture process and can be determined, for example, by measuring the daily titer and using those values ​​to calculate the cumulative titer.

[0037] As used herein, the term “host cell” should be understood to include cells that have been genetically engineered to express a target polypeptide. Genetic engineering of cells involves transfecting, transforming, or transducing cells with a nucleic acid encoding a recombinant polynucleotide molecule (“target gene”), and / or otherwise altering (e.g., through homologous recombination and gene activation or fusion of recombinant and non-recombinant cells) to induce the host cell to express the desired recombinant polypeptide. Methods and vectors for genetically engineering cells and / or cell lines to express target peptides are well known to those skilled in the art; for example, various techniques are illustrated in Current Protocols in Molecular Biology, edited by Ausubel et al. (Wiley & Sons, New York, 1988, and quarterly updates); Sambrook et al., Molecular Cloning: A Laboratory Manual (Cold Spring Laboratory Press, 1989); and Kaufman, RJ., Large Scale Mammalian Cell Culture, 1990, pp. 15–69. The term includes the offspring of the parent cell, regardless of whether the offspring are morphologically or genetically identical to the original parent cell, provided the target gene is present. Cell cultures may contain one or more host cells.

[0038] As used in this article, the phrase “total transcript expression” refers to the sum of transcript expression of each antibody chain.

[0039] Workflow

[0040] In a typical workflow, cell samples are obtained from frozen samples, such as those from single-cell clones or stabilization pools, as well as frozen samples from normal passages or fed-batch processes. In some embodiments, cells are generated by single-cell printing, for example, using commercially available equipment from suppliers such as NamoCell (San Jose, California). The cell samples contain nucleic acids expressing the target protein. The cell samples can be cultured to obtain a sufficient number of cells for mRNA extraction. Generally, single-cell clones and pools are cultured until they achieve 85% viability (approximately 2-4 weeks) in at least two or three consecutive passages, with an average doubling time of less than 35-40 hours.

[0041] Cell samples are subjected to mRNA extraction. See, for example, Cheng et al., 2021, Anal. Methods 3:289-298; and Svec, 2013, Front. Oncol. 3:1-11. For example, samples can be treated to disrupt or lyse cells, such as by treating the samples with one or more detergents and / or denaturing agents (e.g., guanidine salt reagents). Nucleic acids can also be extracted from the samples, for example, after detergent treatment and / or denaturation. Total nucleic acid extraction can be performed using known techniques (e.g., by non-specific binding to a solid phase, such as silica gel). See, for example, U.S. Patent Nos. 5,234,809, 6,849,431; 6,838,243; 6,815,541; and 6,720,166. Commercially available kits are available from companies such as Qiagen, Inc. (Germantown, MD). Qiagen's kit uses silica-based extraction to select mRNAs based on size > 200 bp. First, the sample is treated with DNase to remove DNA and inactivate RNase. Then, mRNA is enriched using an RNAeasy silica column.

[0042] After mRNA extraction, the mRNA is reverse transcribed into cDNA using commonly known methods involving reverse transcriptase. Alternatively, cell samples can be subjected to a rapid lysis kit for cDNA transcription without RNA isolation. An example of such a kit is SingleShot from Bio-Rad (Hercules, California).

[0043] Finally, digital PCR was used to quantify cDNA expression (and mRNA expression).

[0044] Digital PCR

[0045] The method described herein employs digital PCR. For digital PCR, prior to PCR, a sample containing nucleic acids is aliquoted into numerous partitions. These partitions can be implemented in a variety of ways known in the art (e.g., using microplates, capillaries, emulsions, microchamber arrays, or nucleic acid-binding surfaces). Sample separation can involve allocating any suitable portion, including up to the entire sample, between the partitions. Each partition includes a fluid volume separated from the fluid volumes of other partitions. The partitions can be separated from each other by a fluid phase (such as a continuous phase of an emulsion), a solid phase (such as at least one wall of a container), or a combination thereof. In some embodiments, the partitions may contain droplets disposed within a continuous phase, such that the droplets and the continuous phase together form an emulsion.

[0046] Partitions can be formed by any suitable procedure in any suitable manner and have any suitable characteristics. For example, partitions can be formed using fluid dispensers (such as pipettes), droplet generators, or by agitating the sample (e.g., shaking, stirring, sonicating, etc.). Therefore, partitions can be formed continuously, in parallel, or in batches. Partitions can have one or more of any suitable volumes. Partitions can have substantially uniform volumes or can have different volumes. An exemplary partition with substantially uniform volumes is a monodisperse droplet. Example volumes of partitions include average volumes less than about 100, 10, or 1 μL, less than about 100, 10, or 1 nL, or less than about 100, 10, or 1 pL, etc.

[0047] After sample separation, PCR is performed in partitions. A partition is formed capable of carrying out one or more reactions within it. Alternatively, after partition formation, one or more reagents can be added to the partitions to enable them to perform the reactions. Reagents can be added using any suitable mechanism, such as a fluid dispenser, microdroplet fusion, etc.

[0048] Following PCR amplification, nucleic acids are quantified by counting partitions of PCR amplicons containing the target polynucleotide and / or reference polynucleotide. Partitioning the sample by assuming the molecular population follows a Poisson distribution allows for the quantification of the number of different molecules. For descriptions of digital PCR methods, see, for example, Hindson et al., 2011, Anal. Chem. 83:8604-8610; Pohl and Shih, 2004, Expert Rev. Mol. Diagn. 4:41-47; Pekin et al., 2011, Lab Chip 11: 2156-2166; Pinheiro et al., 2012, Anal. Chem. 84:1003-1011; Day et al., 2013, Methods 59:101-107; these references are incorporated herein by reference in their entirety.

[0049] Nucleic acid amplification of cDNA from mRNA transcripts involves digital PCR. Digital PCR can include any method, process, and / or protocol that uses instruments and / or kits associated with performing these methods, processes, and / or protocols to discretely amplify and quantify nucleic acids within individual partitions of a sample. Individual partitions for digital PCR can be generated via microfluidic processes (e.g., using microfluidic devices) and / or via droplet generation processes. Generating individual partitions via microfluidic processes (e.g., using microfluidic devices) and / or via droplet generation processes to provide multiple partitions in the form of droplets and to amplify nucleic acids thereon is more specifically described in the art as “droplet-based digital PCR.” Droplets generated for droplet-based digital PCR can be provided, for example, in the form of a water-in-oil emulsion. In some embodiments, methods, processes, and / or protocols, as well as instruments and / or kits, for performing nucleic acid amplification on partitions of droplets generated using microfluidic devices / processes and / or droplet generation processes are commercially available, such as, but not limited to, those provided by Bio-Rad Laboratories, 10X Genomics, Qiagen, and / or Thermo Fisher Scientific. In exemplary embodiments, nucleic acid amplification includes, but is not limited to, performing droplet digital PCR (ddPCR™) using Bio-Rad Laboratories' QX100™ or QX200™ droplet digital PCR system and analyzing the nucleic acid amplification products generated therefrom.

[0050] Reagents required for nucleic acid amplification (e.g., oligonucleotide primers, oligonucleotide probes, NTPs / dNTPs, polymerases, etc.) can be provided to each of multiple partitions, and nucleic acid amplification can be performed on each of the multiple partitions. In some embodiments, the reagents for nucleic acid amplification can be contained within the partitions, such as an aqueous phase of an emulsion, or microdroplets. A partition containing nucleic acid amplification reagents / components can be merged / fused with a partition containing a sample / nucleic acid extracted from a sample, and nucleic acid amplification can be performed on the merged / fused partitions.

[0051] Microfluidic devices can be used to merge partitions containing nucleic acid amplification reagents with partitions containing sample / nucleic acid extracts, such that each sample / nucleic acid partition includes nucleic acid amplification components / reagents. The number of partitions provided is not particularly limited. In some embodiments, for example in a sequencing reaction, the number of partitions containing different samples / polynucleotides can be about, greater than about, less than about, or at least about 1000, 5000, 10,000, 20,000, 30,000, 40,000, 50,000, 60,000, 70,000, 80,000, 90,000, 100,000, 200,000, 300,000, 40 0,000, 500,000, 600,000, 700,000, 800,000, 900,000, 1,000,000, 2,000,000, 3,000,000, 4,000,000, 5,000,000, 6,000,000, 7,000,000, 8,000,000, 9,000,000, or 10,000,000 partitions. In some embodiments, in the methods described herein, approximately 1,000 to approximately 10,000, approximately 10,000 to approximately 100,000, approximately 10,000 to approximately 500,000, approximately 100,000 to approximately 500,000, approximately 100,000 to approximately 1,000,000, approximately 500,000 to approximately 1,000,000, approximately 1,000,000 to approximately 5,000,000, or approximately 1,000,000 to approximately 10,000,000 partitions may be generated.

[0052] Methods for generating droplets are described, for example, in U.S. Patent Application Publication No. 2011 / 0053798, which is incorporated herein by reference in its entirety. In some embodiments, internal droplets (or partitions) can be fused with external droplets (or partitions) by heating / cooling to change temperature, applying pressure, changing composition (e.g., via chemical additives), applying acoustic energy (e.g., via ultrasonic treatment), exposing to light (e.g., to stimulate photochemical reactions), applying an electric field, or any combination thereof. In some cases, internal droplets may spontaneously fuse with external droplets. The treatment can be continuous or can vary over time (e.g., pulsed, impulsive, and / or repeated treatments). The treatment can provide gradual or rapid changes in emulsion parameters to achieve steady-state or transient initiation of droplet fusion. By selecting appropriate surfactant types, surfactant concentrations, critical micelle concentrations, ionic strengths, etc., for one or more phases of the internal / external partitions, the stability of the partitions and their responsiveness to treatments that induce droplet fusion can be determined during partition formation.

[0053] Fusion can occur spontaneously, requiring no treatment other than a sufficient time delay (or no delay) before processing the fused droplets. Alternatively, internal / external droplets can be treated to controllably induce droplet fusion, thereby forming the assay mixture.

[0054] The emulsion resulting from fusion can be processed. Processing may include subjecting the emulsion to any condition or set of conditions under which at least one desired reaction can occur (and / or cease), and for any suitable period of time. Thus, processing may include maintaining the temperature of the emulsion near a predetermined set point, varying the temperature of the emulsion between two or more predetermined set points (e.g., thermally cycling the emulsion), exposing the emulsion to light, changing the pressure applied to the emulsion, adding at least one chemical substance to the emulsion, applying an electric field to the emulsion, or any combination thereof.

[0055] Signals in the emulsion can be detected after and / or during processing. Signals can be detected optically, electrically, chemically, or in combination thereof. The detected signals may include test signals corresponding to at least one desired reaction occurring in the emulsion. Alternatively or additionally, the detected signals may include code signals corresponding to a code present in the emulsion. Test signals and code signals are generally distinguishable and can be detected using the same or different detectors. For example, both test signals and code signals can be detected as fluorescence signals, which can be distinguished based on excitation wavelength (or spectrum), emission wavelength (or spectrum), and / or unique location within the fused droplets (e.g., for fused droplets, the code signal may be detected as more localized than the test signal). As another example, test signals and code signals can be detected as different optical characteristics, such as the test signal being detected as fluorescence and the code signal being detected as optical reflection. As another example, test signals can be detected optically, while code signals can be detected electrically, and vice versa.

[0056] Partitions can be formed using any separation mode available for digital PCR. A partition can be a well in a microfluidic channel, nanofluidic or microfluidic device, or a microtiter plate, or a reaction chamber within a microfluidic device. A partition can be a region on an array surface. A partition can be the aqueous phase of an emulsion (e.g., microdroplets).

[0057] The droplets described herein may include emulsion compositions (or mixtures of two or more immiscible fluids) as described in U.S. Patent No. 7,622,280, and droplets generated by the apparatus described in International Patent Application Publication No. WO 2010 / 036352. As used herein, the term "emulsion" may refer to a mixture of immiscible fluids (e.g., oil and water). Oil phases and / or water-in-oil emulsions can allow the reaction mixture to be compartmentalized within aqueous droplets. In some embodiments, the emulsion may comprise aqueous droplets situated within a continuous oil phase. In other embodiments, the emulsions provided herein are oil-in-water emulsions, wherein the droplets are oil droplets situated within a continuous aqueous phase. The droplets provided herein can be used to prevent mixing between compartments, and each compartment can prevent its contents from evaporating and coalescing with the contents of other compartments. One or more enzymatic reactions may occur within the droplets.

[0058] Methods for generating partitions / droplets, additives, primers, probes, polymerases for PCR, etc., are described, for example, in U.S. Patent No. 9,347,059.

[0059] As described herein, dividing samples into smaller reaction volumes allows for the use of reduced reagent quantities, thereby lowering analytical material costs. Reducing sample complexity through partitioning also improves the dynamic range of detection because higher abundance molecules can be separated from lower abundance molecules in different compartments, allowing a greater proportion of lower abundance molecules to come into contact with the reaction reagents, which in turn enhances the detection of lower abundance molecules. Data can be stored and processed using a computer. Computer-executable logic can be employed for functions such as grouping and / or analyzing data. The computer can be used to display, store, retrieve, or compute diagnostic results from molecular profiling; display, store, retrieve, or compute raw data; or display, store, retrieve, or compute any sample or patient information useful in the methods described herein.

[0060] In some embodiments, a reference sequence, such as a housekeeping gene (e.g., a gene required to maintain basic cellular function), may be used, where each diploid genome exists in two copies. Dividing the concentration or amount of the target by the concentration or amount of the reference sequence yields an estimate of the number of target copies per genome.

[0061] Housekeeping genes that may be used as a reference in the methods described herein may include genes encoding the following: transcription factors, transcription repressors, RNA splicing genes, translation factors, tRNA synthetases, RNA-binding proteins, ribosomal proteins, RNA polymerases, protein processing proteins, heat shock proteins, histones, cell cycle regulators, apoptosis regulators, oncogenes, DNA repair / replication genes, carbohydrate metabolism regulators, citrate cycle regulators, lipid metabolism regulators, amino acid metabolism regulators, nucleotide synthesis regulators, NADH dehydrogenases, cytochrome C oxidases, ATPases, mitochondrial proteins, lysosomal proteins, proteasome proteins, ribonucleases, oxidases / reductases, cytoskeletal proteins, cell adhesion proteins, channel or transport proteins, receptors, kinases, growth factors, tissue necrosis factors, etc. Specific examples of housekeeping genes that may be used in the described methods include, for example, the human hydroxymethylcholine synthase (HMBS) or BRAF gene.

[0062] Methods for detecting nucleic acids are well known in the art and may include specific hybridization of a probe with a nucleic acid sequence, for example, thereby detecting fluorescence emission from a fluorescently labeled or labeled oligonucleotide probe that hybridizes with nucleic acids and / or nucleic acid amplification products and emits / releases fluorescence emission due to hybridization with nucleic acids and / or nucleic acid amplification products.

[0063] Any measurement tool known in the art can be used to analyze nucleic acids present in samples and / or subjects as described above, such as spectrophotometers for absorption or calorimetry measurements, fluorometers or flow cytometers for fluorescence measurements, scintillation counters or gamma counters for radioactivity measurements, and automated cell counters, automated plate counters, or manual plate counters for cell number measurements. As another example, microwell readers can be used for fluorescence, absorbance, or calorimetry measurements. In some embodiments, measurements for detecting nucleic acids / nucleic acid amplification products may include analysis via digital PCR. In some embodiments, measurements for detecting nucleic acids / nucleic acid amplification products may include analysis using droplet digital PCR (using, for example, but not limited to, ddPCR™) with Bio-Rad's QX100™ or QX200™ droplet digital PCR system. Measurement and analysis of amplification products may include analysis of one-dimensional and / or two-dimensional graphs of fluorescence amplitude exhibited in droplet digital PCR amplification.

[0064] In some embodiments, measurements for identifying nucleic acid / nucleic acid amplification products may include analysis by digital PCR. In some embodiments, measurements for identifying nucleic acid / nucleic acid amplification products may include analysis using droplet digital PCR (using, for example, but not limited to, ddPCR™) with Bio-Rad's QX100™ or QX200™ droplet digital PCR system. Measurement and analysis of amplification products may include analysis of one-dimensional and / or two-dimensional plots of fluorescence amplitude exhibited in droplet digital PCR amplification. In some embodiments, amplification, detection, analysis, and / or identification of nucleic acid / nucleic acid amplification products according to the concepts described herein may be performed as described in Maggi et al., 2020. J. Microbiol. Methods. [Journal of Microbiological Methods] 176, 106022.

[0065] In some embodiments, the capture and concentration of nucleic acids and / or nucleic acid sequences using highly porous hydrogel particles can be achieved using, for example, NANOTRAP® technology (Ceres Nanosciences, Inc. (Manassas, Virginia)), including NANOTRAP® particles and / or NANOTRAP® magnetic particles, as described in, for example, U.S. Patent Application Publications Nos. 2009 / 0087346, 2009 / 0148961, 2012 / 0164749, and 2014 / 0045274. Nanoparticles (e.g., NANOTRAP® particles and NANOTRAP® magnetic particles) functionalized with affinity decoys, affinity groups, and / or affinity ligands that have extremely high affinity for specific nucleic acids and / or nucleic acid sequences (as biomarkers for specific demanding microorganisms) can be used, for example, to capture specific nucleic acids and / or nucleic acid sequences onto nanoparticles under conditions of binding to / high affinity for the nanoparticles, followed by separation / isolation of nanoparticles including the captured nucleic acids and / or nucleic acid sequences bound to the nanoparticles from the sample, and elution of the captured nucleic acids and / or nucleic acid sequences from the nanoparticles, to provide an enriched / enriched sample comprising the specific nucleic acid and / or nucleic acid sequence. The enriched sample can be used for further analysis, such as detection and / or analysis of nucleic acids / nucleic acid sequences using the amplification, detection, and identification methods described herein.

[0066] The method described herein utilizes at least two sets of primers to amplify target regions within a target polynucleotide sequence and a reference polynucleotide sequence for comparison, respectively. The primers can be designed with 5′ tail sequences to allow for differentiation between the target and reference polynucleotide amplicon. A first set of primers with tails of a specific length can be used to amplify the target region within the target polynucleotide sequence, while a second set of primers with tails of a different length can be used to amplify the reference polynucleotide sequence. This results in amplification of nucleic acids using the two sets of primers producing amplicons with detectable target and reference polynucleotides of different lengths.

[0067] Fluorescent DNA dyes can be used to detect amplicones generated by PCR reactions. Any DNA dye that binds nonspecifically to DNA (i.e., binds to any sequence of DNA) can be used, allowing amplicon differentiation by nucleic acid length and quantitative measurements. Exemplary DNA dyes that can be used include EvaGreen (EG) dye, SYBR Green, SYBR Green II, SYBR Gold, oxazolium yellow (YO), YOYO, thiazole orange (TO), PicoGreen (PG), and SYTO dye.

[0068] Primers can be designed to have a region complementary to a portion of the template nucleic acid to be amplified, allowing polymerization to be initiated. Non-complementary nucleotides can be added to the 5′ end of the primer to create a 5′ tail. Typically, primer lengths range from 7 to 100 nucleotides (e.g., 15–60, 20–40, etc.), more typically between 20 and 40 nucleotides, and any length within the stated range.

[0069] It is desirable that the amplification efficiencies of the target and reference sequences be similar or approximately equal to allow for comparison in quantitative analysis. For this reason, primers and amplification conditions should be selected to achieve this result. For example, primer lengths can be extended or shortened at the 5′ or 3′ end to produce primers with the desired melting temperature. Although the complementary regions of primers designed for amplifying the target and reference sequences can be of the same length, their non-complementary tail sequences typically have different lengths to allow for differentiation of the amplicon of the target polynucleotide and the reference polynucleotide. Primers with shorter tails can be designed with higher GC content, while primers with longer tails can be designed with higher AT content to roughly match the melting temperatures of the primer-template complexes of the target and reference polynucleotides (preferably within 3°C of each other). The amplification efficiency of any primer pair can be easily determined using conventional techniques (see, for example, Furtado et al., “Application of real-time quantitative PCR in the analysis of gene expression.” DNA amplification: Current Technologies and Applications. Wymondham, Norfolk, UK: Horizon Bioscience, pp. 131-145 (2004)).

[0070] Primers can be readily synthesized using standard techniques, such as solid-phase synthesis via phosphoramide chemistry, as disclosed in U.S. Patent Nos. 4,458,066 and 4,415,732 (which are incorporated herein by reference); Beaucage et al., Tetrahedron (1992) 48:2223-2311; and Applied Biosystems User Bulletin No. 13 (April 1, 1987). Other chemical synthesis methods include, for example, the phosphotriester method described by Narang et al., Meth. Enzymol. (1979) 68:90, and the phosphodiester method disclosed by Brown et al., Meth. Enzymol. (1979) 68:109. Using these same methods, poly(A) or poly(C) or other non-complementary nucleotide extensions can be introduced into oligonucleotides. Ethylene hexaoxide derivatives can be coupled to oligonucleotides using methods known in the art. Cload et al., J. Am. Chem. Soc. (1991) 113:6324-6326; US Patent No. 4,914,210 to Levenson et al.; Durand et al., Nucleic Acids Res. (1990) 18:6353-6359; and Horn et al., Tet. Lett. (1986) 27:4705-4708.

[0071] In addition, one or more PCR additives or enhancers may be included to increase the yield of amplification reactions, for example, by reducing secondary structure or mis-priming events in nucleic acids. Such additives or enhancers include, but are not limited to, dimethyl sulfoxide (DMSO), N,N,N-trimethylglycine (betaine), formamide, glycerol, nonionic detergents (e.g., Triton X-100, Tween 20, and Nonidet P-40 (NP-40)), 7-dezo-2′-deoxyguanosine, bovine serum albumin, T4 phage gene 32 protein, polyethylene glycol, 1,2-propanediol, and tetramethylammonium chloride.

[0072] As explained above, primers with 5′ tails can be designed for use in polymerase chain reaction (PCR) techniques to distinguish amplicon generated from different polynucleotide sequences based on length. PCR can be used to amplify desired target nucleic acid sequences contained in nucleic acid molecules or mixtures of molecules. In PCR, an excess of a pair of primers is used to hybridize with the complementary strand of the target nucleic acid. Each primer is extended using the target nucleic acid as a template by polymerase. The extension product becomes the target sequence itself after dissociating from the original target strand. Then, new primers hybridize and are extended by polymerase, and this cycle is repeated to geometrically increase the number of target sequence molecules. PCR methods for amplifying target nucleic acid sequences in samples are well known in the art and are described, for example, in Innis et al. (ed.) PCR Protocols (Academic Press, New York, 1990); Taylor (1991) Polymerase chain reaction: basic principles and automation, PCR: A Practical Approach, McPherson et al. (ed.) IRL Press, Oxford; Saiki et al. (1986) Nature 324:163; and U.S. Patent Nos. 4,683,195, 4,683,202 and 4,889,818, all of which are incorporated herein by reference in their entirety.

[0073] Specifically, PCR uses relatively short oligonucleotide primers, with the target nucleotide sequence to be amplified on one side, oriented such that the 3′ ends of the primers and the target nucleotide sequence face each other, with each primer extending towards the other end. The polynucleotide sample is extracted and preferably denatured by heating, and then hybridized with an excess of the first and second primers. In the presence of four deoxyribonucleoside triphosphates (dNTPs—dATP, dGTP, dCTP, and dTTP), polymerization was catalyzed using primers and template-dependent polynucleotide polymerases such as any enzyme capable of producing primer extension products, for example, *E. coli* DNA polymerase I, the Klenow fragment of DNA polymerase I, T4 DNA polymerase, thermostable DNA polymerases isolated from *Taq* (available from multiple sources, e.g., PerkinElmer), *Thermophilic* (United States Biochemicals), *Bacillus stearothermophilus* (Bio-Rad Laboratories), or *Thermococcus litoralis* ("Vent" polymerase, New England Biolabs). This produced two "long products" containing corresponding primers covalently linked at the 5' end to a newly synthesized complementary sequence of the original strand. Then, the reaction mixture is restored to polymerization conditions, for example by lowering the temperature, inactivating the denaturant, or adding more polymerase, and a second cycle begins. The second cycle provides two original strands, two long products from the first cycle, two new long products replicated from the original strands, and two “short products” replicated from the long products. The short products contain the target sequence and have primers at both ends. In each additional cycle, two additional long products are produced, and the number of short products equals the number of long and short products remaining at the end of the previous cycle. Therefore, the number of short products containing the target sequence increases exponentially with each cycle. Preferably, PCR is performed using a commercially available thermal cycler (e.g., PerkinElmer). The amplified products can be detected in solution or using a solid support.

[0074] application

[0075] The methods described herein can be used to quantify the mRNA transcript levels of heterologous sequences encoding a target protein. Gene copy number and mRNA transcript levels have been shown to be associated with antibody production. See Jiang et al., 2006, Biotechnology Progress 22:313-318. In some embodiments, the target protein will be an antigen-binding protein having both heavy and light chains. In such embodiments, the mRNA transcript levels of both heavy and light chains can be measured. In such embodiments, a 1:1 heavy chain to light chain ratio may be desired for optimal antibody formation. In other embodiments, the target protein will be an antigen-binding protein having two heavy chains (HC1 and HC2) and one light chain (LC). In still other embodiments, the target protein will be an antigen-binding protein having two heavy chains (HC1 and HC2) and two light chains (LC1 and LC2). For triple- and quadruple-chain molecules, any combination of transcript level ratios (e.g., HC1:HC2) can be used to evaluate potential candidates.

[0076] Because of potential problems with individual strands (such as antibody strands) that may affect product quality (e.g., product-related impurities (HMW impurities) resulting from misfolded domains that may cause protein aggregation, incomplete molecules (half-molecules), or incorrectly assembled strands (incorrect pairing of heavy strands, incorrect pairing of heavy and light strands (LMW impurities)), transcription-to-translation efficiency, secretion efficiency, etc.), and the possibility of missing the optimal ratio of specific strand pairs, preferred workflows select and amplify those cells that produce products that meet process and product quality requirements. In one workflow, cells expressing the lowest levels of transcripts are excluded from further processing. For example, single-cell clones or cell pools with the lowest 80th, 70th, 60th, 50th, 40th, or 30th percentile transcript levels may be excluded from the workflow.

[0077] Because a wide range of chain ratios (such as antibody chains) indicates a low probability that a single-cell clone or cell pool is a high-expressor of the desired target protein and / or a high probability of producing high levels of product-related impurities, single-cell clones or cell pools in which the transcript level ratio of any two antibody chains is less than 0.2 or greater than 1.5 can be excluded. The ratio of 3 or 4 antibody chains is the same for any two chains. The chain ratio of the expressed peptide (such as antibody chains) can be measured using techniques known in the art. After excluding a certain percentage of poorly expressing single-cell clones or cell pools, the remaining single-cell clones or cell pools are cultured in batches for at least an additional 8, 9, or 10 days to evaluate the best-performing single-cell clones or cell pools. Preferably, the cell viability is higher than 80%.

[0078] As discussed in more detail below, expression vectors can be optimized by using different combinations of promoters, signal sequences, enhancer elements, etc., to achieve the desired heavy chain to light chain ratio.

[0079] Due to the design of the expression vector, the target protein is expected to be expressed at high transcript levels. Transcript expression is normalized relative to the transcript expression of the reference housekeeping gene, ensuring that significant differences in RNA levels across samples are not distorted by variations in sample volume. Expression of each strand can be summed within the same sample and compared between different samples to determine relative expression.

[0080] The methods described herein are adaptable to multiplex PCR, for example, to simultaneously detect and / or quantify multiple target polynucleotides. Therefore, multiple primer sets comprising forward and reverse primers can be used in each reaction mixture, each set targeting a different target polynucleotide sequence and containing detectable 5′ tails of varying lengths, so that amplicones generated from different target polynucleotides can be distinguished based on the intensity of the fluorescence signal after binding to a nonspecific fluorescent DNA dye. In some cases, multiplexing is performed to allow simultaneous detection and / or quantification of all types of various target polynucleotides from a single sample. In some embodiments, multiple primer sets are used to simultaneously detect and / or quantify different target alleles at the same locus or different target loci. In other embodiments, multiple primer sets are used to simultaneously quantify transcript-level expression of different target polynucleotide sequences (e.g., LC, HC).

[0081] A particular embodiment is a method in which at least one step is performed in a perforated plate, and a method in which at least step b) is performed in a perforated plate. The perforated plate may be a 96-well plate or a 384-well plate, preferably a 384-well plate.

[0082] Another embodiment is the following method, wherein single-cell clones or cell pools are monitored for a period of time sufficient to obtain batch sample titer curves, preferably 5-15 days, with samples taken every 2-3 days.

[0083] Another embodiment is the method in which steps a) to e) are performed in a 96-well plate and steps g) to j) are performed in a 384-well plate.

[0084] For ddPCR, sample tracking can be ensured using methods well-known in the art via barcode plates and barcode readers.

[0085] Generation of mammalian host cells expressing the target protein

[0086] The expression of target proteins in cells can be achieved transiently or stably using well-known methods (Davis et al., Basic Methods in Molecular Biology, 2nd ed., Appleton & Lange, Norwalk, Connecticut, 1994; Sambrook et al., Molecular Cloning: A Laboratory Manual, 3rd ed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York, 2001).

[0087] Stable integration methods are well known in the art. In short, stable integration is typically achieved by transiently introducing a heteropolynucleotide or a vector containing a heteropolynucleotide into a host cell, which facilitates the stable integration of the heteropolynucleotide into the cellular genome. Typically, the heteropolynucleotide is flanked by homologous arms (i.e., sequences homologous to upstream and downstream regions of the integration site). Before their introduction into mammalian host cells, circular vectors can be linearized to facilitate integration into the cellular genome. Methods for introducing vectors into cells are well known in the art and include transfection using biological methods (e.g., viral delivery), chemical methods (e.g., using cationic polymers, calcium phosphate, cationic lipids, or cationic amino acids), physical methods (e.g., electroporation or microinjection), or hybrid methods (e.g., protoplast fusion).

[0088] To stably transfect mammalian cells, it is known that only a small fraction of cells can integrate foreign DNA into their genome, depending on the expression vector and transfection technique used. To identify and select these integrators, genes encoding selection markers (e.g., those for antibiotic resistance) are typically introduced into the host cell along with the target gene. Preferred selection markers include those conferring resistance to drugs such as G418, hygromycin, and methotrexate. Cells stably transfected with the introduced nucleic acid can be identified by methods such as drug selection (e.g., cells with introduced selection marker genes will survive while other cells die).

[0089] Stable integration is achieved through specific methods using recombinase-mediated cassette exchange (RMCE; Bode and Baer, ​​2001, CurrOpin Biotechnol. [Current Biotechnical Perspective] 12:473-80, and Bode et al., 2000, Biol. Chem. [Biochemistry] 381:801-813) for site-specific integration into the genome (also known as “targeted integration”). Site-specific recombinases such as Flp and Cre mediate recombination between two copies of their target sequence, referred to as FRT and loxP, respectively. Using two incompatible target sequences, such as FRT combined with F3 (Schlake and Bode, 1994, Biochemistry, 33:12746-51) and an inverted recognition target site (Feng et al., 1999, J. Mol. Biol. 292:779-85), allows DNA fragments to be inserted into predetermined chromosomal loci carrying target sequences of similar conformation. See also European Patent No. EP 1781796 B1 and European Patent Application Publication No. EP 2789691 A1.

[0090] Insertion of RMCEs into specific sites in the genome can be mediated by nucleases (e.g., zinc finger proteins (ZFPs), transcription activator-like effector nucleases (TALENs), and clustered regularly spaced short palindromic repeats (CRISPR) / CRISPR-associated protein 9 (Cas9)). These nucleases can be engineered to generate single-strand and double-strand breaks (SSBs / DSBs) in the genome. There are two main and distinct pathways for DSB repair—homologous recombination and non-homologous end joining (NHEJ). Homologous recombination requires the presence of a homologous sequence as a template (e.g., a “donor” containing an RMCE) to guide the cellular repair process, and the repair outcome is error-free and predictable. In the absence of a template (or “donor”) sequence for homologous recombination, cells typically attempt to repair DSBs via the unpredictable and error-prone process of non-homologous end joining (NHEJ).

[0091] Vectors can be any molecule or entity suitable for transferring and / or transporting proteins encoding information to host cells and / or specific locations and / or compartments within host cells (e.g., nucleic acids, plasmids, bacteriophages, transposons, kinases, chromosomes, viruses, viral capsids, virions, naked DNA, complex DNA, etc.). Vectors can include viral and nonviral vectors, and non-attachment mammalian vectors. Vectors are commonly referred to as expression vectors, such as recombinant expression vectors and cloning vectors. Vectors can be introduced into host cells to allow replication of the vector itself, thereby amplifying copies of the polynucleotides contained therein. Cloning vectors may contain sequence components, which typically include, but are not limited to, origin of replication, promoter sequences, transcription initiation sequences, enhancer sequences, and selectability markers. These elements can be appropriately selected by those skilled in the art.

[0092] Vectors can be used to transform host cells and contain nucleic acid sequences that direct and / or control (alongside the host cell) the expression of one or more heterologous coding regions operatively linked to them. Expression constructs can include, but are not limited to, sequences that affect or control transcription, translation, and, in the presence of introns, influence RNA splicing of coding regions operatively linked to them. "Operably linked" means that the components to which this term applies are in a relationship that allows them to perform their inherent functions. For example, in a vector operatively linked to a protein-coding sequence, a control sequence (e.g., a promoter) is arranged such that normal activity of the control sequence leads to transcription of the protein-coding sequence, resulting in recombinant expression of the encoded protein.

[0093] Vectors that are functional in the specific host cell used can be selected (i.e., the vector is compatible with the host cell structure, thereby allowing gene amplification and / or expression). In some embodiments, the vector used employs protein fragment complementation assays using a reporter protein such as dihydrofolate reductase (see, for example, U.S. Patent No. 6,270,964). Suitable expression vectors are known in the art and are commercially available.

[0094] Typically, vectors used in host cells will contain sequences for plasmid maintenance and for cloning and expressing exogenous nucleotide sequences. Such sequences will typically include one or more of the following nucleotide sequences: promoter, one or more enhancer sequences, origin of replication, transcription and translation control sequences, transcription termination sequences, complete intron sequences containing donor and acceptor splicing sites, various pro- or pro-sequence sequences that improve glycosylation or yield, natural or heterologous signal sequences (lead sequences or signal peptides) for polypeptide secretion, ribosome binding sites, polyadenylated sequences, internal ribosome entry sites (IRES) sequences, expression enhancement sequence elements (EASE), triplet leader sequences (TPA) and VA gene RNA from adenovirus 2, multi-connector regions for inserting multinucleotides encoding the polypeptide to be expressed, and selective labeling elements. Vectors can be constructed from starter vectors (such as commercially available vectors), and additional elements can be obtained separately and ligated into the vector. Methods for obtaining the components are well known to those skilled in the art.

[0095] Vector components can be homologous (i.e., from the same species and / or strain as the host cell), heterologous (e.g., from a species or strain different from the host cell), heterozygous (i.e., a combination of side sequences from more than one source), synthetic, or natural. The sequences of useful components in these vectors can be obtained using methods well-known in the art, such as those previously identified by mapping and / or by restriction endonucleases. Furthermore, they can be obtained by polymerase chain reaction (PCR) and / or by screening genomic libraries with suitable probes.

[0096] Ribosome binding sites are typically required for the initiation of mRNA translation and are characterized by a Shine-Dalgarno sequence (prokaryotes) or a Kozak sequence (eukaryotes). This element is typically located at the 3' of the promoter and at the 5' of the coding sequence of the polypeptide to be expressed.

[0097] Origin of replication facilitates the amplification of vectors within host cells. These can be included as part of commercially available prokaryotic vectors or chemically synthesized based on known sequences and ligated into vectors. Various viral sources (e.g., SV40, polyomaviruses, adenoviruses, vesicular stomatitis virus (VSV), or papillomaviruses such as HPV or BPV) can be used to clone vectors in mammalian cells.

[0098] Transcriptional and translational control sequences for mammalian host cell expression vectors can be excised from the viral genome. Commonly used promoter and enhancer sequences are derived from polyomaviruses, adenovirus 2, simian virus 40 (SV40), and human cytomegalovirus (CMV). For example, the human CMV promoter / enhancer of the immediate early gene 1 can be used. See, for example, Patterson et al., 1994, Applied Microbiol. Biotechnol. [Applied Microbiology and Biotechnology] 40:691-98. DNA sequences derived from the SV40 viral genome (e.g., SV40 start site, early and late promoters, enhancers, splice sites, and polyadenylation sites) can be used to provide additional genetic elements for the expression of structural gene sequences in mammalian host cells. Early and late viral promoters are particularly useful because they are readily available as fragments from the viral genome and can also contain the origin of viral replication (Fiers et al., 1978, Nature 273:113; Kaufman, 1990, Meth. in Enzymol. 185:487-511). Smaller or larger SV40 fragments can also be used, provided they include approximately 250 bp of sequence extending from the Hind III site to the BglI site located at the origin of SV40 viral replication.

[0099] Transcription termination sequences are typically located at the 3′ end of the polypeptide coding region and are used to terminate transcription. In prokaryotic cells, the transcription termination sequence is usually a GC-rich fragment followed by a poly-T sequence. While the sequence can be readily cloned from libraries or even commercially available as part of a vector, it can also be readily synthesized using nucleic acid synthesis methods known to those skilled in the art.

[0100] Selective marker genes encode proteins essential for the survival and growth of host cells grown in selective media. Typical selective marker genes encode proteins that: (a) confer resistance to antibiotics or other toxins (e.g., ampicillin, tetracycline, or kanamycin for prokaryotic host cells); (b) compensate for cellular nutritional deficiencies; or (c) provide essential nutrients not available in complex or limited media. Specific selective markers are kanamycin resistance genes, ampicillin resistance genes, and tetracycline resistance genes. Advantageously, neomycin resistance genes can also be used for selection in both prokaryotic and eukaryotic host cells.

[0101] Other selective marker genes can be used to amplify genes to be expressed. Amplification is the process of tandemly repeating genes required for the production of proteins essential for growth or cell survival within the chromosomes of recombinant cells across successive generations. Examples of suitable selective markers for mammalian cells include glutamine synthase (GS), dihydrofolate reductase (DHFR), and promoter-free thymidine kinase genes. Transformed mammalian cells are subjected to selection pressure, where only the transformant is uniquely suited for survival due to the presence of the selective marker gene in the vector. This selection pressure is applied by culturing transformed cells under conditions of continuously increasing selectant concentrations in the culture medium, thereby amplifying the selective marker gene and the DNA encoding the target protein. Consequently, an increased amount of the target polypeptide is synthesized from the amplified DNA.

[0102] In some cases, such as when glycosylation is desired in eukaryotic host cell expression systems, various pre- or pro-sequences can be manipulated to improve glycosylation or yield. For example, altering the peptidase cleavage site of a specific signal peptide, or adding the pro-sequence, can also affect glycosylation. The final protein product may have one or more readily expressible additional amino acids at the -1 position (relative to the first amino acid of the mature protein), which may not have been completely removed. For example, the final protein product may have one or two amino acid residues attached to the amino terminus that are present at the peptidase cleavage site. Alternatively, if the enzyme cleaves in such a region within the mature polypeptide, using some enzyme cleavage site may produce a slightly truncated form of the desired polypeptide.

[0103] Expression and cloning typically involve promoters containing molecules that are recognized by the host organism and operatively linked to encode a target protein. A promoter is a non-transcribed sequence (typically within approximately 100 to 1000 bp) located upstream (i.e., 5') of the start codon of a structural gene and controls the transcription of that gene. Promoters are generally grouped into one of two categories: inducible promoters and constitutive promoters. Inducible promoters elicit an increased level of transcription in response to a change in culture conditions (such as the presence or absence of nutrients, or a change in temperature) under their control. Constitutive promoters, on the other hand, consistently transcribe the gene they are operatively linked to; that is, they have little or no control over gene expression. Many promoters recognized by a variety of potential host cells are well-known.

[0104] Suitable promoters for mammalian host cells are well known and include, but are not limited to, promoters derived from viral genomes, such as polyomaviruses, fowlpoxviruses, adenoviruses (e.g., adenovirus 2), bovine papillomaviruses, avian sarcomaviruses, cytomegaloviruses, retroviruses, hepatitis B viruses, and simian virus 40 (SV40). Other suitable mammalian promoters include heterologous mammalian promoters, such as heat shock promoters and actin promoters.

[0105] Additional targeted promoters include, but are not limited to: the SV40 early promoter (Benoist and Chambon, 1981, Nature 290:304-310); the CMV promoter (Thornsen et al., 1984, Proc. Natl. Acad. USA 81:659-663); promoters contained in the 3' long terminal repeat sequence of Rouss sarcoma virus (Yamamoto et al., 1980, Cell 22:787-797); the herpesvirus thymidine kinase promoter (Wagner et al., 1981, Proc. Natl. Acad. Sci. USA 78:1444-1445); glyceraldehyde-3-phosphate dehydrogenase (GAPDH); promoters and regulatory sequences from metallothionein genes (Prinster et al., 1982, Nature). 296:39-42); and prokaryotic promoters, such as β-lactamase promoters (Villa-Kamaroff et al., 1978, Proc. Natl. Acad. Sci. USA [Proceedings of the National Academy of Sciences] 75:3727-3731); or tac promoters (DeBoer et al., 1983, Proc. Natl. Acad. Sci. USA [Proceedings of the National Academy of Sciences] 80:21-25). Also of interest are the following animal transcriptional control regions, which exhibit tissue specificity and have been used in transgenic animals: the elastase I gene control region active in pancreatic acinar cells (Swift et al., 1984, Cell 38:639-646; Ornitz et al., 1986, Cold Spring Harbor Symp. Quant. Biol. 50:399-409; MacDonald, 1987, Hepatology 7:425-515); the insulin gene control region active in pancreatic β cells (Hanahan, 1985, Nature 315:115-122); and the immunoglobulin gene control region active in lymphoid cells (Grosschedl et al., 1984, Cell 38:647-658; Adames et al., 1985, Nature). 318:533-538; Alexander et al., 1987, Mol. Cell. Biol.[Molecular Cell Biology] 7:1436-1444); Active mouse mammary tumor virus control region in testes, mammary glands, lymphoid and mast cells (Leder et al., 1986, Cell [Cell] 45:485-495); Active albumin gene control region in the liver (Pinkert et al., 1987, Genes and Development. [Genes and Development] 1:268-276); Active α-fetoprotein gene control region in the liver (Krumlauf et al., 1985, Mol. Cell. Biol. [Molecular Cell Biology] 5:1639-1648; Hammer et al., 1987, Science [Science] 253:53-58); Active α1-antitrypsin gene control region in the liver (Kelsey et al., 1987, Genes and Devel. [Genes and Development] 1:161-171); The active β-globulin gene control region in myeloid cells (Mogram et al., 1985, Nature 315:338-340; Kollias et al., 1986, Cell 46:89-94); the active myelin basic protein gene control region in oligodendrocytes of the brain (Readhead et al., 1987, Cell 48:703-712); the active myosin light chain-2 gene control region in skeletal muscle (Sani, 1985, Nature 314:283-286); and the active gonadotropin-releasing hormone gene control region in the hypothalamus (Mason et al., 1986, Science 234:1372-1378).

[0106] Enhancer sequences can be inserted into this vector to increase transcription in higher eukaryotes. Enhancers are cis-acting elements of DNA, typically about 10–300 bp in length, that act on the promoter to increase transcription. Enhancers are relatively independent in orientation and location, residing at the 5' and 3' positions of the transcription unit. Several enhancer sequences are known from mammalian genes (e.g., globulins, elastases, albumins, alpha-fetoproteins, and insulin). However, enhancers derived from viruses are typically used. The SV40 enhancer, cytomegalovirus early promoter enhancer, polyomavirus enhancer, and adenovirus enhancer known in the art are exemplary enhancing elements for activating eukaryotic promoters. Although enhancers can be located at the 5' or 3' of the coding sequence in the vector, they are typically located at the 5' site of the promoter.

[0107] A sequence encoding an appropriate natural or heterologous signal sequence (lead sequence or signal peptide) can be introduced into an expression vector to promote the extracellular secretion of the target protein. The choice of signal peptide or leader sequence depends on the type of host cell from which the target protein is to be produced, and the heterologous signal sequence can replace the natural signal sequence. Examples of functional signal peptides in mammalian host cells include: the interleukin-7 signal sequence described in U.S. Patent No. 4,965,195; the interleukin-2 receptor signal sequence described in Cosman et al., 1984, Nature [Nature] 312:768; the interleukin-4 receptor signal peptide described in European Patent No. 0367566; the type I interleukin-1 receptor signal peptide described in U.S. Patent No. 4,968,607; and the type II interleukin-1 receptor signal peptide described in European Patent No. 0460846.

[0108] Additional control sequences that have been shown to improve the expression of heterologous genes from mammalian expression vectors include elements such as expression enhancement sequence elements (EASE) derived from CHO cells (Morris et al., in Animal Cell Technology, pp. 529-534 (1997); U.S. Patent Nos. 6,312,951 B1, 6,027,915 and 6,309,841 B1) and triplet leader sequences (TPL) and VA gene RNA derived from adenovirus 2 (Gingeras et al., 1982, J. Biol. Chem. 257:13475-13491). Virus-derived internal ribosome entry site (IRES) sequences enable efficient translation of bicistronic mRNAs (Oh and Sarnow, 1993, Current Opinion in Genetics and Development 3:295-300; Ramesh et al., 1996, Nucleic Acids Research 24:2697-2700).

[0109] After construction, one or more vectors can be inserted into suitable cells for amplification and / or peptide expression. Transformation of the expression vector into selected cells can be accomplished by well-known methods, including transfection, infection, calcium phosphate co-precipitation, electroporation, nuclear transfection, microinjection, DEAE-dextran-mediated transfection, cationic lipid-mediated delivery, liposome-mediated transfection, microbombardment, receptor-mediated gene delivery, and polylysine, histone, chitosan, and peptide-mediated delivery. The chosen method will vary in part depending on the type of host cells used. These methods, and other suitable methods, are well known to those skilled in the art and are described in manuals and other technical publications, such as Sambrook et al., *Molecular Cloning: A Laboratory Manual*, 3rd edition, Cold Spring Harbor Laboratory Press, Cold Spring Harbor, New York (2001).

[0110] The term "transformation" refers to a change in the genetic characteristics of a cell, and a cell is considered transformed when it is modified to contain new DNA or RNA. For example, a cell is transformed when it undergoes genetic modification from its native state by introducing new genetic material via transfection, transduction, or other techniques. After transfection or transduction, the transformed DNA can recombine with the cell's DNA through physical integration into the cell's chromosome, or it can be temporarily maintained as an addendum element without replication, or it can replicate independently as a plasmid. A cell is considered "stablely transformed" when the transformed DNA replicates with cell division.

[0111] The term “transfection” refers to the absorption of foreign or exogenous DNA by cells. Many transfection techniques are well known in the art and are disclosed herein. See, for example, Graham et al., 1973, Virology 52:456; Sambrook et al., 2001, Molecular Cloning: A Laboratory Manual, ibid.; Davis et al., 1986, Basic Methods in Molecular Biology, Elsevier; Chu et al., 1981, Gene 13:197.

[0112] The term “transduction” refers to the process by which foreign DNA is introduced into cells via viral vectors. See Jones et al., (1998). Genetics: principles and analysis. Boston: Jones & Bartlett Publ.

[0113] Target protein

[0114] "Target protein" includes naturally occurring proteins, recombinant proteins, engineered proteins (e.g., proteins that do not exist in nature and have been designed and / or produced by humans), or peptides. Target peptides and proteins may have scientific or commercial significance. Target proteins may be, but do not necessarily be, proteins known or suspected of having therapeutic relevance.

[0115] Target proteins include, in particular, secreted proteins, non-secreted proteins, intracellular proteins, or membrane-bound proteins. Target peptides and proteins can be produced using cell culture methods through recombinant animal cell lines and may be referred to as “recombinant proteins.” One or more expressed proteins can be produced intracellularly or secreted into a culture medium, from which they can be recovered and / or collected. The terms “isolated protein” or “isolated recombinant protein” refer to a target peptide or protein purified from proteins or peptides or other contaminants that would interfere with its therapeutic, diagnostic, preventative, research, or other uses. Target proteins include proteins that exert therapeutic effects by binding to targets, particularly those listed below (including targets derived from them, associated targets, and modifications thereof).

[0116] Target proteins include "antigen-binding proteins." Antigen-binding proteins are proteins or polypeptides containing an antigen-binding region or portion that has an affinity for another molecule (antigen) to which it binds. Antigen-binding proteins encompass antibodies, peptides, antibody fragments, antibody derivatives, antibody analogs, fusion proteins (including single-chain variable fragments (scFv), double-chain (bivalent) scFv, and IgG scFv) (see, for example, Orcutt et al., 2010, Protein EngDes Sel [Protein Engineering, Design & Selection] 23:221-228), heterologous IgG (see, for example, Liu et al., 2015, JBiol Chem [Journal of Biochemistry] 290:7535-7562), mutant proteins, and proteins derived from XmAb. ®Antibodies prepared by Xencor, Inc. (Monrovia, California). Examples of antigen-binding proteins include human antibodies, humanized antibodies; chimeric antibodies; recombinant antibodies; single-chain antibodies; biantibodies; triantibodies; tetraantibodies; Fab fragments; F(ab')2 fragments; IgD antibodies; IgE antibodies; IgM antibodies; IgG1 antibodies; IgG2a antibodies; IgG3 antibodies; or IgG4 antibodies, and fragments thereof. Also included are bispecific T-cell binder molecules (BiTE). ® Bispecific T-cell binder molecules with extensions (such as half-life extensions) (e.g., HLE BiTE) ® Molecules, HeteroIg BiTE ® And others), chimeric antigen receptors (CAR, CAR T) and T cell receptors (TCR). As used in this article, B1mAb refers to [Fab-scFv Fab]-heterogeneous Fc, B1HmAb refers to [Fab-VH Fab-heterogeneous Fc; B2mAb refers to Fab-scFv-Fc; B2HmAb refers to Fab-VH-Fc; C1mAb refers to Fab-heterogeneous Fc-[scFv] C2mAb refers to IgG-scFv; C2HmAb refers to IgG-VH.

[0117] As used herein, the term “antigen-binding protein” is used in its broadest sense and refers to a protein that contains a portion that binds to an antigen or target, and optionally contains a scaffold or framework portion that allows the antigen-binding portion to adopt a conformation that promotes the binding of the antigen-binding protein to the antigen. Antigen-binding proteins may contain, for example, artificial scaffolds that can replace the protein scaffold or have grafted CDRs or CDR derivatives. Such scaffolds include, but are not limited to, antibody-derived scaffolds containing mutations introduced to, for example, stabilize the three-dimensional structure of the antigen-binding protein; and fully synthetic scaffolds containing, for example, biocompatible polymers. See, for example, Korndorfer et al., 2003, Proteins: Structure, Function, and Bioinformatics, 53(1):121-129; Roque et al., 2004, Biotechnol. Prog. 20:639-654. Furthermore, peptide antibody mimics (“PAMs”) and scaffolds based on antibody mimics utilizing fibronectin components as scaffolds may be used.

[0118] Antigen-binding proteins may contain one or more antibody chains. A typical antibody has four antibody chains, produced by the expression of one heavy chain and one light chain. As used herein, “antibody chain” or “chain” refers to an antibody light chain, antibody heavy chain, antibody light chain fusion protein, and antibody heavy chain fusion protein, scFv-Fc fusion, VHH fusion, etc. The terms “antibody heavy chain” and “antibody light chain” have their standard meaning in the art and include, for example, the various antibody heavy and light chains described elsewhere herein (e.g., the heavy and light chains of lgG1, lgG2, lgG3, and lgG4 mAbs). The terms “antibody heavy chain” and “antibody light chain” include standard full-length antibody heavy and light chains, as well as their derivatives containing the corresponding variable regions (VL or VH). The terms “antibody heavy chain fusion” and “antibody heavy chain fusion protein” refer to a polypeptide containing an antibody heavy chain covalently linked to one or more additional proteins or peptides. For example, an “antibody heavy chain fusion protein” could be an antibody heavy chain covalently linked to a cytokine. The linkage can be direct or via a peptide linker (e.g., a glycine-serine linker). In antibody heavy chain fusion proteins, the antibody heavy chain can be linked to one or more additional proteins at the N-terminus or C-terminus (or both) of the heavy chain. The terms "antibody light chain fusion protein" and "antibody light chain fusion" have the same meaning as described immediately following "antibody heavy chain fusion protein" above, except for the antibody light chain. As used herein, "antibody fusion protein" refers to an antibody as provided herein, covalently linked (e.g., via the antibody's heavy or light chain) to one or more additional proteins or polypeptides. Thus, an antibody fusion protein contains at least one antibody heavy chain fusion protein or antibody light chain fusion protein as one of the polypeptides of an antibody fusion protein. Most commonly, an antibody fusion protein is a molecule containing two antibody light chains, one antibody heavy chain, and one antibody heavy chain fusion protein, such that the additional protein is linked to one of the antibody's heavy chains.

[0119] Antigen-binding proteins can have structures such as those of naturally occurring immunoglobulins. An immunoglobulin is a tetrameric molecule. In naturally occurring immunoglobulins, each tetramer consists of two pairs of identical polypeptide chains, each pair having a "light chain" (approximately 25 kDa) and a "heavy chain" (approximately 50-70 kDa). The amino-terminal portion of each chain includes a variable region of approximately 100 to 110 or more amino acids, which is primarily responsible for antigen recognition. The carboxyl-terminal portion of each chain defines a constant region primarily responsible for effector function. Human light chains are classified as κ light chains and λ light chains. Heavy chains are classified as μ, δ, γ, α, or ε, and antibody isotypes are defined as IgM, IgD, IgG, IgA, and IgE, respectively.

[0120] Naturally occurring immunoglobulin chains exhibit the same general structure of a relatively conserved framework region (FR) linked by three hypervariable regions (also known as complementarity-determining regions or CDRs). Both the light and heavy chains contain domains FR1, CDR1, FR2, CDR2, FR3, CDR3, and FR4 from the N-terminus to the C-terminus. Each domain can be assigned amino acids according to the definition in Sequences of Proteins of Immunological Interest, 5th Edition, US Dept. of Health and Human Services, PHS, NIH, NIH Publication No. 91-3242, (1991). If necessary, the CDR can also be redefined according to alternative naming schemes such as Chothia's (see Chothia and Lesk, 1987, J. Mol. Biol. [Journal of Molecular Biology] 196:901-917; Chothia et al., 1989, Nature [Nature] 342:878-883 or Honegger and Pluckthun, 2001, J . Mol. Biol. [Journal of Molecular Biology] 309:657-670).

[0121] In the context of this disclosure, when the dissociation constant (KD) ≤ 10 -8 When M, the antigen-binding protein is said to "specifically bind" or "selectively bind" to its target antigen. When KD ≤ 5 × 10⁻⁶ -9 At time M, the antibody binds to the antigen with "high affinity," when KD ≤ 5 × 10⁻⁶. -10 When M occurs, the antibody binds to the antigen with "extremely high affinity".

[0122] Unless otherwise stated, the term "antibody" includes any isotype or subtype of glycosylated and non-glycosylated immunoglobulin, or its antigen-binding region that competes with an intact antibody for specific binding. Additionally, unless otherwise stated, the term "antibody" refers to an intact immunoglobulin or its antigen-binding portion that competes with an intact antibody for specific binding. The antigen-binding portion can be produced by recombinant DNA technology or by enzymatic cleavage or chemical cutting of an intact antibody and can form elements of the target protein. Unless otherwise stated, antibodies include human, humanized, chimeric, multispecific, monoclonal, polyclonal, heterologous IgG, bispecific antibodies, and their oligomers or antigen-binding fragments. Antibodies include IgG1, IgG2, IgG3, or IgG4 types. It also includes proteins having antigen-binding fragments or antigen-binding regions, such as Fab, Fab', F(ab')2, Fv, biantibodies, Fd, dAb, maximal antibodies, single-chain antibody molecules, single-domain VHH, complementarity-determining region (CDR) fragments, scFv, biantibodies, triantibodies, tetraantibodies, and polypeptides that contain at least a portion of an immunoglobulin sufficient to bind a specific antigen to a target polypeptide.

[0123] Antigen-binding proteins may have one or more binding sites. If more than one binding site is present, these binding sites may be the same or different from each other. For example, naturally occurring human immunoglobulins typically have two identical binding sites, while "bispecific" or "bifunctional" antibodies have two different binding sites.

[0124] The Fab fragment is a monovalent fragment having VL, VH, CL, and CH1 domains; the F(ab')2 fragment is a divalent fragment having two Fab fragments connected by a disulfide bridge in the hinge region; the Fd fragment has VH and CH1 domains; the Fv fragment has VL and VH domains in the antibody arm; and the dAb fragment is an antigen-binding fragment having VH, VL, or either VH or VL domains (US Patent Nos. 6,846,634, 6,696,245, US Patent Application Publication Nos. 2005 / 0202512, 2004 / 0202995, 2004 / 0038291, 2004 / 0009507, 2003 / 0039958, Ward et al., 1989, Nature 341:544-546).

[0125] Single-chain antibodies (scFvs) are antibodies in which the VL and VH regions are linked via a linker (e.g., a synthetic sequence of amino acid residues) to form a continuous protein chain, wherein the linker is long enough to allow the protein chain to fold back and form a monovalent antigen-binding site (see, for example, Bird et al., 1988, Science 242:423-26 and Huston et al., 1988, Proc. Natl. Acad. Sci. USA 85:5879-83, U.S. Patents 7,741,465 and 6,319,494, and Eshhar et al., 1997, Cancer Immunol Immunotherapy 45:131-136). scFvs retain the ability of the parent antibody to specifically interact with the target antigen.

[0126] A biantibody is a divalent antibody consisting of two polypeptide chains, each containing VH and VL domains linked by a linker that is too short to allow pairing between the two domains on the same chain, thus allowing each domain to pair with a complementary domain on the other polypeptide chain (see, for example, Holliger et al., 1993, Proc. Natl. Acad. Sci. USA [Proceedings of the National Academy of Sciences] 90:6444-48; and Poljak et al., 1994, Structure [Structure] 2:1121-23). ​​If the two polypeptide chains of a biantibody are identical, the biantibody resulting from their pairing will have two identical antigen-binding sites. Polypeptide chains with different sequences can be used to prepare biantibodies with two different antigen-binding sites. Similarly, triantibodies and tetraantibodies are antibodies consisting of three and four polypeptide chains, respectively, forming three and four antigen-binding sites that may be identical or different.

[0127] For clarity, and as described herein, note that antigen-binding proteins may, but do not have to, be of human origin (e.g., human antibodies), and in some cases will contain non-human proteins, such as rat or mouse proteins, and in other cases antigen-binding proteins may contain hybrids of human and non-human proteins (e.g., humanized antibodies).

[0128] The target protein may include a human antibody. The term "human antibody" includes all antibodies having one or more variable and constant regions derived from a human immunoglobulin sequence. In one embodiment, all variable and constant domains are derived from a human immunoglobulin sequence (a fully human antibody). Such antibodies can be prepared in a variety of ways, including by immunizing mice genetically modified to express genes encoding human heavy and / or light chains, such as mice derived from the Xenomouse®, UltiMab™, or Velocimmune® systems, or rats derived from UniRat®, with the target antigen. Phage-based methods may also be used.

[0129] Alternatively, the target protein may include a humanized antibody. The sequence of a “humanized antibody” differs from that of an antibody derived from a non-human species in that one or more amino acid substitutions, deletions, and / or additions are made such that, when administered to a human subject, the humanized antibody is less likely to induce an immune response and / or induce a less severe immune response compared to a non-human species antibody. In one embodiment, certain amino acid mutations are made in the framework and constant domains of the heavy and / or light chains of a non-human species antibody to produce a humanized antibody. In another embodiment, one or more constant domains from a human antibody are fused to one or more variable domains from a non-human species. Examples of how humanized antibodies can be prepared can be found in U.S. Patent Nos. 6,054,297, 5,886,152, and 5,877,293.

[0130] It also includes modified proteins, such as those chemically modified by non-covalent, covalent, or both covalent and non-covalent bonds. It further includes proteins containing one or more post-translational modifications, which can be prepared by modification through cellular modification systems or by in vitro introduction or other means by enzymatic and / or chemical methods.

[0131] The target protein may also include recombinant fusion proteins, which include, for example, polymerized domains such as leucine zippers, coiled helices, and the Fc portion of immunoglobulins. It also includes proteins containing all or part of the amino acid sequence of the differentiating antigen (called CD proteins) or their ligands, or proteins substantially similar to any of these.

[0132] In some embodiments, the target protein may include a colony-stimulating factor, such as granulocyte colony-stimulating factor (G-CSF). Such G-CSF reagents include, but are not limited to, Neupogen® and Neulasta®. It also includes erythropoiesis stimulants (ESAs), such as Epogen® (epogenetin α), Aranesp® (dabepoetin α), Dynepo® (epogenetin δ), Mircera® (methoxy-polyethylene glycol-epogenetin β), Hematide®, MRK-2578, INS-22, Retacrit® (epogenetin ζ), Neorecormon® (epogenetin β), Silapo® (epogenetin ζ), Binocrit® (epogenetin α), epogenetin α Hexal, Abseamed® (epogenetin α), Ratioepo® (epogenetin θ), Eporatio® (epogenetin θ), Biopoin® (epogenetin θ), epogenetin α, epogenetin β, epogenetin ζ, epogenetin θ and epogenetin δ, epogenetin ω, epogenetin ι, tissue plasminogen activators, GLP-1 receptor agonists, and molecules of any of the foregoing substances or their variants or analogues and biosimilars.

[0133] In some embodiments, the target protein may include proteins that specifically bind to: one or more CD proteins, HER receptor family proteins, cell adhesion molecules, growth factors, nerve growth factor, fibroblast growth factor, transforming growth factor (TGF), insulin-like growth factor, bone-inducing factor, insulin and insulin-related proteins, coagulation and coagulation-related proteins, colony-stimulating factor (CSF), other blood and serum proteins, blood group antigens; receptors, receptor-related proteins, growth hormone, growth hormone receptor, T cell receptors; neurotrophic factors, neurotrophic proteins, relaxin, interferon, interleukin, viral antigens, lipoproteins, integrins, rheumatoid factor, immunotoxins, surface membrane proteins, transport proteins, homing receptors, addressins, regulatory proteins, and immunoadhesins.

[0134] In some embodiments, the target protein binds alone or in any combination to one or more of the following: CD proteins (including, but not limited to, CD3, CD4, CD5, CD7, CD8, CD19, CD20, CD22, CD25, CD30, CD33, CD34, CD38, CD40, CD70, CD123, CD133, CD138, CD171, and CD174), HER receptor family proteins (including, for example, HER2, HER3, HER4, and EGF receptors), EGFRvIII, cell adhesion molecules (e.g., LFA-1, Mol, p150,95, VLA-4, ICAM-1, VCAM, and αv / β3 integrin), growth factors (including, but not limited to, vascular endothelial growth factor (“VEGF”); VEGFR2, growth hormone, thyroid-stimulating hormone, follicle-stimulating hormone, luteinizing hormone, growth hormone-releasing factor, parathyroid hormone, and Müllerian-inhibiting substances). The following are listed: human macrophage inflammatory protein (MIP-1-α), erythropoietin (EPO), nerve growth factors (such as NGF-β), platelet-derived growth factor (PDGF), fibroblast growth factors (including, for example, aFGF and bFGF), epidermal growth factor (EGF), Cripto, transforming growth factor (TGF) (especially including TGF-α and TGF-β (including TGF-β1, TGF-β2, TGF-β3, TGF-β4, or TGF-β5)), insulin-like growth factor-I and insulin-like growth factor-II (IGF-I and IGF-II), des(1-3)-IGF-I (brain IGF-I) and bone-inducing factor, insulin and insulin-related proteins (including but not limited to insulin, insulin A chain, insulin B chain, proinsulin, and insulin-like growth factor binding protein); coagulation proteins and coagulation-related proteins (especially such as factor VIII, tissue factor, van Wilbond). Willebrand factor, protein C, α-1-antitrypsin, plasminogen activators (such as urokinase and tissue plasminogen activator (“t-PA”), bombazine, thrombin, thrombopoietin and thrombopoietin receptor), colony-stimulating factor (CSF) (especially including M-CSF, GM-CSF and G-CSF), other blood and serum proteins (including but not limited to albumin, IgE and blood group antigens), receptors and receptor-related proteins (including, for example, flk2 / flt3 receptors, obesity (OB) receptors, growth hormone receptors and T cell receptors); neurotrophic factors (including but not limited to bone-derived neurotrophic factor (BDNF) and neurotrophin-3, neurotrophin-4, neurotrophin-5 or neurotrophin-6 (NT-3, NT-4, NT-5 or NT-6));Relaxin A chain, relaxin B chain and pro-relaxin, interferons (including, for example, interferon α, interferon β and interferon γ), interleukins (ILs) (e.g. IL-1 to IL-10, IL-12, IL-15, IL-17, IL-23, IL-12 / IL-23, IL-2Ra, IL-1-R1, IL-6 receptor, IL-4 receptor and / or IL-13 receptor, IL-13RA2 or IL-17 receptor, IL-1RAP); viral antigens, including but not limited to AIDS envelope virus antigens, lipoproteins, calcitonin, glucagon, atrial natriuretic peptide, and pulmonary surfactant. Tumor necrosis factor-α and tumor necrosis factor-β, enkephalin, BCMA, IgKappa, ROR-1, ERBB2, mesothelin, RANTES (activated and regulated normal T cell expression and secretion factor), mouse gonadotropin-related peptide, DNase, FR-α, inhibin and activin, integrin, protein A or D, rheumatoid factor, immunotoxin, bone morphogenetic protein (BMP), superoxide dismutase, surface membrane protein, decay accelerator factor (DAF), AIDS envelope, transport protein, homing receptor, MIC (MIC-a, MIC-B), ULBP 1-6, EPCAM, addressin, regulatory protein, immunoadhesin, antigen-binding protein, growth hormone, CTGF, CTLA4, eotaxin-1, MUC1, CEA, c-MET, Claudin-18, GPC-3, EPHA2, FPA, LMP1, MG7, NY-ESO-1, PSCA, ganglioside GD2, ganglioside GM2, BAFF, OPGL (RANKL), myostatin, Dickkopf-1 (DKK-1), Ang2, NGF, IGF-1 receptor, hepatocyte growth factor (HGF), TRAIL-R2, c-Kit, B7RP-1, PSMA, NKG2D-1, programmed cell death protein 1 and ligand, PD1 and PDL1, mannose receptor / hCGβ, hepatitis C virus, mesothelin dsFv [PE38] conjugate, Legionella pneumophila (lly), IFN γ, interferon-gamma inducible protein 10 (IP10), IFNAR, TALL-1, thymic stromal lymphopoietin (TSLP), proprotein convertase subtilisin / Kexin type 9 (PCSK9), stem cell factor, Flt-3, calcitonin gene-related peptide (CGRP), OX40L, α4β7, platelet-specific (platelet glycoprotein IIb / IIIb (PAC-1), transforming growth factor β (TFGβ), zona pellucida sperm-binding protein 3 (ZP-3), TWEAK, platelet-derived growth factor receptor α (PDGFRα), sclerostin, and any bioactive fragments or variants of the foregoing.

[0135] In another embodiment, the target protein includes abciximab, adalimumab, adelimumab, aflibercept, alenmab, alicurumab, anakinase, asceticipeptide, baliximab, belimumab, bevacizumab, biosozumab, bonatumab, bentuximab, brodamarab, mocantozumab, konatumab, cetuximab, cetuximab, konatumab, dalizumab, denosumab, ikulimab, izollium, efalizumab, epazolizumab, etanercept, evokulumab, galiximab, genitalumab, gemutuzumab, golimumab, tiimozumab, infliximab, ipilimumab, and ixeximab. kizumab), lerdelimumab, ruximab, mapalimumab, motesanibdiphosphate, moromab-CD3, natelizumab, nesiritide, nimotuzumab, nivolumab, olizumab, oflamimumab, omalizumab, interleukin, pallizumab, panitumumab, pembrolizumab, pertuzumab, pectinumab, ranibizumab, rituximab, rituximab, romista, lomoxoluzumab, saxagsta, tocilizumab, tosimomab, trastuzumab, ustekinumab, vedozalimumab, vexizalimumab, voloximab, zalumab, zalumab, and any biosimilars of the foregoing substances.

[0136] The target protein according to the invention encompasses all the foregoing and further includes antibodies containing 1, 2, 3, 4, 5, or 6 complementarity-determining regions (CDRs) of any of the aforementioned antibodies. One or more CDRs can be covalently or non-covalently incorporated into the molecule to make it an antigen-binding protein. The antigen-binding protein can be incorporated into one or more CDRs as part of a larger polypeptide chain, can be covalently linked to one or more CDRs to another polypeptide chain, or can be non-covalently incorporated into one or more CDRs. CDRs allow the antigen-binding protein to bind specifically to a particular target antigen. Variations are also included that include regions identical in amino acid sequence to a reference amino acid sequence of the target protein at 70% or higher, particularly 80% or higher, more particularly 90% or higher, even more particularly 95% or higher, especially 97% or higher, even more particularly 98% or higher, even more particularly 99% or higher. This identity can be determined using a variety of well-known and readily available amino acid sequence analysis software. Preferred software includes those implementing the Smith-Waterman algorithm, which is considered a satisfactory solution to the problem of searching and aligning sequences. Other algorithms can also be used, especially when speed is a significant consideration. Commonly used programs for DNA, RNA, and peptide alignment and homology matching include FASTA, TFASTA, BLASTN, BLASTP, BLASTX, TBLASTN, PROSRCH, BLAZE, and MPSRCH, the latter being an implementation of the Smith-Waltman algorithm for execution on massively parallel processors manufactured by MasPar.

[0137] The “Fc” region, as used herein, comprises two heavy chain segments containing the CH2 and CH3 domains of an antibody. The two heavy chain segments are held together by two or more disulfide bonds and by hydrophobic interactions of the CH3 domains. Target proteins containing the Fc region (including antigen-binding proteins and Fc fusion proteins) form another aspect of this disclosure.

[0138] A "half-antibody" is an immunofunctional immunoglobulin construct comprising a complete heavy chain, a complete light chain, and a second heavy chain Fc region paired with the Fc region of the complete heavy chain. A linker may, but is not necessary, connect the heavy chain Fc region and the second heavy chain Fc region. In a particular embodiment, the half-antibody is a monovalent form of the antigen-binding protein disclosed herein. In other embodiments, one Fc region may be associated with a second Fc region using paired charged residues. In the context of this disclosure, the half-antibody may be the target protein.

[0139] As used herein, unless otherwise expressly stated, the term "a / an" means one or more. Furthermore, unless the context requires otherwise, singular terms shall include the plural and plural terms shall include the singular. Generally, the nomenclature and techniques used in conjunction with the cell and tissue culture, molecular biology, immunology, microbiology, genetics, and protein and nucleic acid chemistry and hybridization described herein are those well-known and commonly used in the art.

[0140] All documents or portions thereof (including, but not limited to, patents, patent applications, articles, books, and monographs) referenced in this application are expressly incorporated herein by reference. The content described in the embodiments of the invention may be combined with other embodiments of the invention.

[0141] This invention is not limited in scope to the specific embodiments described herein, which are intended as individual illustrations of various aspects of the invention, and functionally equivalent methods and components are also within the scope of the invention. In fact, various modifications to the invention will become apparent to those skilled in the art from the foregoing description and drawings, in addition to those shown and described herein. Such modifications are intended to fall within the scope of the appended claims. Example

[0142] Example 1

[0143] A predictive ddPCR assay was developed for screening pools in cell line development prior to fed-batch culture. mRNA transcript levels were analyzed from samples from both normal-pass and fed-batch (FB) cultures and compared with final product quality and titer results.

[0144] Materials and Methods

[0145] Transfection and cell culture

[0146] Serum-free suspension GSKO CHO cells were transfected with 25 µg of plasmid DNA (an expression vector representing the antibody heavy and light chains). For each transfection, 20 × 10⁻⁶ cells were subjected to treatment using a GenePulser Xcell (BioRad) at 3175 µF, 200 V, and 700 Ω. 6 Electroporation was performed on individual cells. Transfected cells were placed in 20 mL of cell culture medium in 50 mL conical tubes and stored in a Swiss Krones shaker at 36.5°C and 5% CO2. Three days after selection, cells were transferred to glutamine-free selection medium. The clone pools were cultured until they achieved >85% viability (approximately 2–4 weeks).

[0147] Single-cell cloning

[0148] Single-cell cloning of the pool was performed using a Berkeley Lights instrument (Emeryville, California), and the cells were exported into 96-well plates using a Biomek FX. p Beckman Coulter has carefully selected or chosen Gas Permeable Rapid Expansion 24-well plates (G-rex). ® In Wilson Wolf, St. Paul, Minnesota, the process was then scaled up and passaged in 24-hole deep-hole plates.

[0149] Feed generation and analysis

[0150] Once single-cell clones achieved >90% viability with a stable doubling time, fed-batch culture was performed to assess protein expression and transcript levels in early time-point FB single-cell clones and normal passage pools. The correlation between the HC1 / HC2 transcript ratio and PQA (product quality attribute) was analyzed.

[0151] Cells were passaged into FB medium, and monitoring was performed on days 3, 6, 8, and 10. Nutrients were added to the culture on days 3, 6, and 8. The supernatant was harvested on day 10, and the conditioned medium was analyzed for titer and product quality assessment on ATOLL purification material.

[0152] Cell sample preparation

[0153] CHO pools expressing three different types of molecules were evaluated: 4-chain heterologous IgG, 3-chain AmAb, and 3-chain C1mAb. 1 × 10⁶ cells were collected during normal pool passage and fed-batch production, meeting the standard of cell viability ≥ 80% at harvest. 6 - 5 × 10 6 Cells were centrifuged at 300 xg for 5 minutes and then washed in 1x DPBS (Life Technologies) to remove any residual culture medium. Cell samples were frozen in 1 mL of cryopreservation medium (cell culture medium with 10% DMSO) or 350 µL of RNeasy kit RLT buffer (Qaijie) and stored at -80°C until ready for cDNA generation.

[0154] cDNA generation

[0155] To extract mRNA, the RNeasy kit (Kiagen) was used according to the manufacturer's protocol. To remove any remaining DNA, a digestion step was performed using FastDigest enzyme (Thermo Scientific) according to the manufacturer's protocol. Alternatively, RNA for use in RT-PCR was prepared directly from cell culture samples using the RNA Rapid Extraction Kit (BioRad SingleShot Cell Lysis Kit). To generate cDNA, a reverse transcription step was performed using the iScript cDNA Synthesis Kit (BioRad), with incubation at 46°C for 20 minutes using a thermal cycler (Applied Biosystems). After mRNA isolation and cDNA synthesis, the quality of mRNA and cDNA was checked using a NanoDrop spectrophotometer at a 260 / 280 absorbance ratio.

[0156] Primer and probe set design

[0157] Design primer and probe sets specific to the unique sequences in each heavy and light chain of each molecule using PrimerExpress or the PrimerQuest tool from Integrated DNA Technologies (IDT) (ordered from IDT). The reference gene CHO β-actin (ChoBAct) is used to normalize the DNA loading in different samples.

[0158] ddPCR assay for quantifying mRNA transcript levels

[0159] Two-step RT-ddPCR was used to rapidly quantify mRNA transcript levels (e.g., each heavy and light chain of a multispecific antibody). Each reaction consisted of: 0.2 ng / µL cDNA, 18 μM of each primer, and 5 μM of probes for the target gene and the ChoBAct reference gene (Sigma-Aldrich) in combination with probes using Supermix (dUTP-free) (Bio-Rad Laboratories). Water was added to a final volume of 20 µL per well. Droplets were generated using an AutoDG system (Bio-Rad Laboratories), and PCR amplification was performed using a thermal cycler (Applied Biosystems) under the following thermal cycling conditions: 40 cycles of 95°C for 10 min, 94°C for 30 sec, annealing at 60°C, and extension for 1 min (heating cap: 105°C; sample volume: 40 µL), followed by a final cycle of 98°C for 10 min and a hold at 4°C. Droplets were measured using a Qx200 droplet reader (Bio-Rayet Corporation), and the quantities were analyzed using Quantasoft software (Bio-Rayet Corporation).

[0160] Statistical data analysis

[0161] The Spearman correlation matrix in R (available from the R Project for Statistical Computing) was used to analyze the correlation between transcript levels and functional fed-batch titers and product quality attributes.

[0162] result

[0163] The protein expression and transcript levels from the FB pool are shown in Figure 1, while those from the normal passage pool are shown in Figure 2. Positive correlations were found between protein expression and transcript levels and the titers of (A) 4-chain heterologous IgG, (B) 3-chain molecules, and (C) asymmetric C1 mAb.

[0164] Spearman correlation analysis of the 3-chain molecular results indicated a strong correlation between total expression and titer, with heavy chain (HC)1 showing the highest positive correlation with titer, where HC1 is the longer of the two HCs in the 3-chain molecule. The HC1 / HC2 ratio was directly proportional to the main peak and high molecular weight (HMW) in size exclusion chromatography (SEC) determination, but inversely proportional to the leading peak and low molecular weight (LMW). A consistent trend was observed at both FB and normal passage time points, such as... Figure 3A and Figure 3B As shown.

[0165] Transcripts of structurally different bispecific molecules (B2HmAb, C2HmAb, B2mAb, and C2mAb) with similar antigen-binding properties were evaluated. Figure 4 The results showed that the sum of heavy chain (HC) and light chain (LC) transcripts was positively correlated with the normalized total fed-batch titer. Despite structural differences, the selected optimal and alternative pools fell within the best half of the samples used to assess transcript expression.

[0166] Clonal transcript expression analysis using a rapid RNA extraction method showed a positive correlation between total clonal expression (grey bars) and normalized total titer (black bars), with the best pool of 3-strand molecules (marked with a black asterisk). See also Figure 5A The error bars represent the standard deviation of three replicates. Figure 5B The single-strand expression results showed that clones with the lowest titers (marked with light gray asterisks) produced little to no HC1, which explains the poor clonal titer performance and supports the predictive power of the assay. Figure 5C The Spearman correlation matrix indicates that the HC1 / HC2 ratio is proportional to the main peak but inversely proportional to the LMW (SEC) (where HC2 is the shorter of the two HCs in the 3-chain molecule). Figure 5D The results show a strong correlation between the HC1:HC2 ratio and MP and LMW. The optimal clones fall at approximately 0.6 HC1:HC2 ratio, marked with black squares.

[0167] This assay can be used not only to predict titers in pools and clones, but also to predict product quality attributes (PQA). For example, for 3-chain molecules, the correlation between the HC1 / HC2 transcript ratio and PQA was analyzed, and it was found that both were directly proportional to the percentage of main peak and high molecular weight (HMW) (as measured by size exclusion chromatography (SEC)) at both fed-batch and normal subculture time points, and inversely proportional to the percentage of leading peak (as measured by non-reducing capillary electrophoresis (nr-CE)) and low molecular weight (LMW, as measured by SEC). See also Figure 3A -B. Overall, this assay can serve as a powerful tool for predicting the best-performing pool prior to FB evaluation and also has potential applications in clone screening.

Claims

1. A method for selecting a single-cell clone or cell pool for manufacturing an antigen-binding protein having two to four different antibody chains, the method comprising: a) Passage single-cell clones or cell pools expressing the antigen-binding protein to a level where they exhibit at least 85% viability in at least two consecutive passages and a doubling time of less than 40 hours. b) Extract mRNA from the cell; c) Reverse transcribe the mRNA to generate cDNA; d) Perform ddPCR to quantify the transcript levels of this antibody chain, and e) Exclude single-cell clones or cell pools that are at least in the lowest 30th percentile of the total sum of the levels of all transcripts of that chain.

2. The method of claim 1, wherein at least one of the antibody chains is a heavy chain or heavy chain-ScFv.

3. The method of claim 1, wherein at least one of the antibody chains is a light chain.

4. The method of claim 1, wherein the single-cell clone or cell pool expresses an antigen-binding protein having two different antibody chains.

5. The method of claim 4, wherein the single-cell clone or cell pool expresses an antigen-binding protein having a heavy chain (HC) or heavy chain-ScFv and a light chain (LC).

6. The method of claim 1, wherein the single-cell clone or cell pool expresses an antigen-binding protein having three or four different antibody chains.

7. The method of any one of claims 1-6, wherein the exclusion of single-cell clones or cell pools is performed on single-cell clones or cell pools within the lowest 50th percentile of the sum of all such chain transcript levels.

8. The method of claim 1, further comprising f) subjecting the remaining single-cell clones or cell pools to fed batch culture for at least an additional 8 days while maintaining the cell viability above 80%.

9. The method of claim 5, further comprising excluding cell clones or cell pools expressing an antigen-binding protein having one heavy chain and one light chain, wherein the transcript level ratio of HC:LC or LC-HC is less than 0.2 or greater than 1.

5.

10. The method of claim 6, wherein the cell expresses an antigen-binding protein having two heavy chains and one light chain (LC), the heavy chain being selected from HC1 and HC2, and HC1 and HC2-scFv.

11. The method of claim 10, further comprising excluding single-cell clones or cell pools in which the transcript level ratio of HC1:HC2 or HC2:HC1 is less than 0.2 or greater than 1.5, or the ratio of HC1:HC2-scFv or HC2-scFv:HC1 is less than 0.2 or greater than 1.

5.

12. The method of claim 6, wherein the cell expresses an antigen-binding protein having two heavy chains (HC1 and HC2) and two light chains (LC1 and LC2).

13. The method of claim 12, further comprising excluding single-cell clones or cell pools in which the transcript level ratio of HC1:HC2 or HC2:HC1 is less than 0.2 or greater than 1.

5.

14. The method of claim 1, wherein the ddPCR is a one-step ddPCR.

15. The method of claim 1, wherein the ddPCR is a two-step ddPCR.

16. The method of claim 1, wherein the cell is selected from the group consisting of CHO, HEK293, NSO, or Sp2 / O cells.

17. The method of claim 16, wherein the cell is a CHO cell.

18. The method of claim 17, wherein the CHO cell is DHFR- (dihydrofolate reductase deficiency) or GSKO (glutamine synthase knockout) type.

19. The method of claim 1, wherein the cell is generated by single-cell printing.

20. A method for selecting single-cell clones or cell pools expressing antigen-binding proteins having one or more antibody heavy chains and one or more antibody light chains to achieve stable growth, the method comprising: a) Passage the cells until they exhibit greater than 80% viability and a doubling time of less than 35 to 40 hours in at least two passages. b) Extract mRNA from the cell; c) Reverse transcribe the mRNA to generate cDNA; d) Perform ddPCR to quantify the transcript levels of one or more heavy chains and one or more light chains; and e) Exclude single-cell clones or cell pools that are at least in the lowest 30th percentile of the total sum of all strand transcript levels.

21. The method of claim 20, wherein the doubling time has a deviation of less than 50% between successive generations.