A method for predicting the production stability of clonal cell lines
A modeling framework for predicting clonal cell line stability in biopharmaceuticals addresses the inefficiencies of current methods by accurately selecting stable cell lines early in the development process, enhancing production stability and reducing resource consumption.
Patent Information
- Application Number
- JP2025528531
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-11-16
- Filing Date
- 2023-11-14
- Publication Date
- 2025-11-18
AI Technical Summary
Current methods for predicting the production stability of clonal cell lines used in biopharmaceuticals are time-consuming and resource-intensive, with up to 63% of CHO cells evaluated as unstable, leading to significant impacts on manufacturing schedules and product distribution.
A method and system utilizing a modeling framework with multivariate latent variable modeling, multiway analysis, and evolutionary model structure to predict clonal cell line stability by analyzing product concentration profiles, enabling early selection of stable cell lines.
Facilitates early and robust prediction of clonal cell line production stability, reducing the time and resources required for cell line development by accurately identifying unstable cell lines, thereby improving production capacity and efficiency.
Smart Images

Figure 2025537576000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to the development of protein-producing cell lines for biopharmaceuticals, in particular to methods and systems for determining or predicting the production stability of clonal cell lines, and methods and systems for selecting clonal cell lines. [Background technology]
[0002] Fully human antibodies can be obtained in a variety of ways, for example, from yeast-based libraries or from animals such as mice that have been genetically engineered to produce a repertoire of human antibodies. Yeast displaying human antibodies on their surface that bind to the antigen of interest can be selected using a FACS (fluorescence-activated cell sorting)-based system or by bead capture with labeled antigen. Transgenic animals engineered to express human immunoglobulin genes can be immunized with the antigen of interest and antigen-specific human antibodies isolated using B cell selection techniques. Human antibodies produced by these techniques can be evaluated for desirable properties such as affinity, developability, and selectivity.
[0003] Mammalian cell lines are used to generate therapeutic protein-producing clonal cell lines. The nucleic acid sequence encoding the protein of interest is cloned into an expression vector and subsequently transfected into a host cell line. The transfected cell population is bulked and single-cell sorted, and these single-cell-sorted clonal cell lines are evaluated for production of the protein of interest by expansion. The clonal cell lines are ranked based on their product concentration, and through a series of selection events, approximately 50 clonal cell lines are selected for production stability evaluation.
[0004] Examples of such mammalian cell lines include mouse myeloma cells (NS0), baby hamster kidney cells (BHK), human embryonic kidney cells (HEK293), and Chinese hamster ovary cells (CHO). Currently, over 80% of approved recombinant proteins are produced in CHO cells (Butler & Spearman, 2014; Walsh, 2018). Reasons for the widespread use of CHO cells include the relative ease of introducing and expressing foreign DNA, rapid growth rate, reliable protein folding machinery, adaptability to serum-free suspension culture, and the ability to produce proteins with human-like post-translational modifications.
[0005] However, cell lines such as CHO cells are known for their high levels of chromosomal and genomic heterogeneity (Wurm & Wurm, 2017). As a result of this genetic plasticity, cell lines can exhibit production instability, whereby a decrease in the quantity and quality of recombinant proteins can be observed during long-term culture (Dahodwala & Lee, 2019).
[0006] Therefore, biopharmaceutical manufacturers must demonstrate to regulatory authorities that the cells used to manufacture their drug products maintain stable protein quality over time. Cell line production stability studies are typically designed with three or four evaluation points spanning 60–80 cell line generations. Cell lines tested in these stability studies are typically defined as stable when there is less than a 30% decrease in recombinant protein concentration (Dahodwala & Lee, 2019). Under this criterion, it is estimated that up to 63% of all CHO cells evaluated in production stability studies are classified as unstable (Dahodwala & Lee, 2019).
[0007] Because manufacturing schedules are typically planned a year in advance, the yield of one process can have a significant impact on the manufacturing schedule if product concentration is not maintained throughout the manufacturing period. Therefore, an unexpectedly low product yield can have a significant impact on the schedule through process repetition and can have cascading effects on product distribution. Therefore, these stability studies represent an important, yet time-consuming and resource-consuming, endeavor for manufacturers.
[0008] To date, modeling techniques have been used to monitor industrial processes, typically in the context of soft sensors (Ramaker et al., 2005; Camacho et al., 2008; Gunther et al., 2008). Soft sensors incorporate online measurements of process variables, such as temperature, pH, and dissolved oxygen, into models to predict output variables. For example, Gunther et al. used online process variable measurements to predict product titer. However, soft sensors have not been able to satisfactorily predict key qualities of cell line production stability. Therefore, cell line production stability testing remains a time-consuming and expensive undertaking.
[0009] Therefore, there is currently a need for methods to reduce the time and resources consumed in cell line production stability testing. Summary of the Invention
[0010] According to a first embodiment, there is provided a method for selecting a clonal cell line for use in producing a therapeutic protein, the method comprising: measuring a product concentration for each clonal cell line for a plurality of clonal cell lines; determining product concentration profile data for each clonal cell line based on the product concentrations; inputting the product concentration profile data into a learning model comprising a modeling framework, wherein the modeling framework comprises a modeling framework including multivariate latent variable modeling, a multifactorial analysis structure, and an evolutionary model structure; generating an output indicative of the production stability of each clonal cell line using the learning model; and selecting a clonal cell line for use in therapeutic protein production based on the output.
[0011] According to a second embodiment, there is provided a method for producing a therapeutic protein, the method comprising: measuring a product concentration for each clonal cell line for a plurality of clonal cell lines; determining product concentration profile data for each clonal cell line based on the product concentrations; inputting the product concentration profile data into a learning model comprising a modeling framework, wherein the modeling framework comprises multivariate latent variable modeling, a multiway analysis structure, and an evolutionary model structure; using the learning model to generate an output indicative of the production stability of each clonal cell line; and selecting a clonal cell line for use in therapeutic protein production based on the output.
[0012] According to a further embodiment, there is provided a system for determining clonal cell line production stability or selecting clonal cell lines, the system comprising: (a) an input for receiving product concentration profile data of the clonal cell lines; (b) a learning model for determining clonal cell line production stability or selecting clonal cell lines, the learning model comprising a modeling framework including multivariate latent variable modeling, a multifactorial analysis structure, and an evolutionary model structure; (c) one or more processors for processing the clonal cell line product concentration profile data using the learning model; and (d) an output indicative of the production stability of the clonal cell lines based on the product concentration profile data of the clonal cell lines according to the learning model.
[0013] According to a further embodiment, there is provided a computer program comprising instructions which, when executed by a computer or data processor, cause the computer to perform operations in accordance with the methods described herein.
[0014] According to a further embodiment, a computer readable medium is provided containing instructions that, when executed by a computer or data processor, cause the computer to perform operations in accordance with the methods described herein.
[0015] According to a further embodiment there is provided a computer readable data carrier having stored thereon a computer program as described herein.
[0016] The methods and systems of the present invention are advantageous in that they facilitate early and robust prediction of clonal cell line production stability compared to other methods. Specifically, incorporating clonal cell line product concentration profile data into the modeling framework described herein provides a technical advantage in that it predicts clonal cell line production stability with greater accuracy and efficiency than other methods, thereby aiding in clonal cell line selection during biopharmaceutical development. Applying this methodology early during cell line development (CLD) enables early selection of clonal cell lines predicted to exhibit unstable production, thereby improving CLD production capacity and shortening the overall chemistry, manufacturing, and control (CMC) process.
[0017] The details of one or more embodiments of the invention are set forth in the accompanying description below. Other features, objects, and advantages of the invention will be apparent from the description and claims. [Brief explanation of the drawings]
[0018] [Figure 1] 1 shows an exemplary flow chart for selecting a clonal cell line for use in the production of a therapeutic protein. [Figure 2] FIG. 1 shows a schematic diagram of batchwise expansion to address multi-way sequences in multivariate classification of clonal cell line production stability. [Figure 3] 1 shows a schematic diagram of the evolving multiway classification procedure. [Figure 4] 1 shows an exemplary process for classification of the production stability of clonal cell lines using a modeling framework trained incorporating a labeled calibration dataset. [Figure 5] Shown is a 3D score plot of the unsupervised evolutionary multi-way principal component analysis (EMPCA) model for Set A. Here, squares represent stable clonal cell lines and circles represent unstable clonal cell lines. The principal component space represents the reduced dimension resulting from multivariate analysis of titer features, but without formal class supervision (unsupervised). [Figure 6] Figure 1 shows an unsupervised evolutionary multi-way principal component analysis (EMPCA) model for Set B, where squares represent stable clonal cell lines and circles represent unstable clonal cell lines. The principal component space represents the dimensional reduction resulting from multivariate analysis of titer features, but without formal class supervision (unsupervised). [Figure 7] As described in Example 4, variables important for prediction (VIP) are shown for the PLS-DA model generated after production run 3. Variables with high VIP scores are potentially important for predicting clonal cell line production stability. Specific Description of the Invention
[0019] Before discussing specific embodiments with reference to the accompanying figures, the following description of the embodiments is provided.
[0020] According to a first embodiment, there is provided a method for selecting a clonal cell line for use in producing a therapeutic protein, the method including: measuring a product concentration for each clonal cell line for a plurality of clonal cell lines; determining product concentration profile data for each clonal cell line based on the product concentrations; inputting the product concentration profile data into a learning model including a modeling framework, wherein the modeling framework includes multivariate latent variable modeling, a multiway analysis structure, and an evolutionary model structure; using the learning model to generate an output indicative of the production stability of each clonal cell line; and selecting a clonal cell line for use in therapeutic protein production based on the output.
[0021] The multivariate latent variable modeling may include at least one of partial least squares regression (PLS) and principal component analysis (PCA), and the PLS may include at least one of PLS-discriminant analysis (PLS-DA), PLS-support vector machine (PLS-SVM), PLS-neural network (PLS-NN), PLS-logistic regression (PLS-LR), PLS-k-nearest neighbor (PLS-KNN), PLS-decision tree analysis (PLS-DT), PLS-Naive Bayes (PLS-NB), PLS-Random Forest (PLS-RF), or PLS-Gradient Boosting (PLS-GB).
[0022] The product concentration profile data may include at least one of the mean product concentration, standard deviation, skewness, kurtosis, difference, maximum product concentration, and maximum-minimum slope. The product concentration may include the concentration of the monoclonal antibody.
[0023] Measuring the product concentration of each clonal cell line for the plurality of clonal cell lines may include measuring the product concentration of each clonal cell line across multiple production runs. The product concentration of the clonal cell line can be measured in two, three, or four production runs. The product concentration of the clonal cell line can be measured in two production runs. The product concentration of the clonal cell line can be measured in three production runs. The product concentration of the clonal cell line can be measured in four production runs. The product concentration of the clonal cell line can be measured for up to 150 generations. The method may further include obtaining a generation number distribution for each production run.
[0024] The evolutionary model construction may include analyzing the final clonal cell line product concentration profile of each production run. The clonal cell line may be a mammalian cell line. The clonal cell line may be a CHO cell line.
[0025] The concentration of the clonal cell line product can be measured at multiple bioreactor scale. The concentration of the clonal cell line product can be measured at 15 ml bioreactor scale.
[0026] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly used in the art to which this invention belongs. All patents and publications mentioned in this specification are incorporated herein by reference in their entirety.
[0027] The term "comprising" encompasses "including" and "consisting," e.g., a composition "comprising" X may consist solely of X, or may include additional components such as X+Y.
[0028] The phrase "consisting essentially of" limits the scope of a technical feature to specific materials or steps, and those that do not materially affect the basic characteristics of the claimed feature.
[0029] The phrase "consisting of" excludes the presence of additional components.
[0030] The term "about" in relation to a numerical value x means, for example, x±10%, 5%, 2%, 1%.
[0031] The term "clonal cell line" or "clone," as used herein, refers to an isolated host cell containing a gene of interest. Here, isolation of host cells means separation from other host cells using techniques known in the art, such as FACS (fluorescence-activated cell sorting) or dilution cloning. The clonal cell line can be subjected to therapeutic protein production stability analysis while the isolated clonal cell line is grown in cell culture medium. The cells grown in the cell culture medium share a common ancestor with each clonal cell line.
[0032] As used herein, the term "clonal cell line production stability" or "production stability" refers to the stability of therapeutic protein production by a clonal cell line, i.e., the production of a consistent product concentration or titer of a therapeutic protein over 50-150 generations. A range of product concentrations or titers is defined as less than a 30% decrease in therapeutic protein production concentration or titer.
[0033] As used herein, the terms "clonal cell line product concentration," "product concentration," "titer," or "productivity titer" refer to the amount of therapeutic protein produced by a clonal cell line, i.e., the concentration of the therapeutic protein produced by the clonal cell line. The concentration of the therapeutic protein produced by the clonal cell line is measured using techniques such as ELISA, HPLC, Western blotting, immunoassay, detection of protein biological activity, FACS analysis, fluorescence microscopy, direct detection of fluorescent protein by FACS analysis, spectrophotometry, or other techniques known in the art. As used herein, the terms "clonal cell line product concentration profile," "clonal cell line product concentration profile data," "titer profile," "product concentration profile," or "product concentration profile data" refer to a set of variables (features) calculated from the product concentration data measured for each clonal cell line. For example, the clonal cell line product concentration profile may include any one or more of the mean product concentration, standard deviation, skewness, kurtosis, difference, maximum product concentration, and maximum-minimum slope.
[0034] As used herein, the term "production run" refers to the process of expressing a protein encoded by a gene of interest in a clonal cell line, in which one or more clonal cell lines in individual vessels undergo the steps of feeding, growth, and therapeutic protein production. A production stability study consists of multiple production runs.
[0035] A "batch" can be defined as a single container of a clonal cell line that goes through the feeding, growth, and therapeutic protein production steps of a production run. Thus, a production run consists of multiple batches of separate clonal cell lines that go through all of the feeding, growth, and therapeutic protein production steps for the production of a protein product. In production stability testing, the production stability of multiple clonal cell lines is evaluated, with each clonal cell line being present in a separate batch for each production run. In production stability testing, the production stability of multiple clonal cell lines is evaluated, with each clonal cell line being present in a separate batch for each production run.
[0036] The term "generation number" or "generation," as used herein, refers to the number of times a clonal cell line has doubled. For example, if the generation number is 80, it means that the clonal cell line has doubled 80 times.
[0037] As used herein, the term "product" refers to the protein produced by the clonal cell line. Thus, the terms "protein product," "cell line product," and "therapeutic protein product" are equivalent to the term "product" for purposes of the described invention. "Therapeutic protein production" refers to the process of production of a therapeutic protein product by a clonal cell line.
[0038] As used herein, the term "antibody" refers to a molecule having an immunoglobulin-like domain (e.g., IgG, IgM, IgA, IgD, or IgE), including monoclonal antibodies, recombinant antibodies, polyclonal antibodies, chimeric antibodies, human antibodies, humanized antibodies, multispecific antibodies (including bispecific antibodies), and heteroconjugate antibodies. It also includes modified forms thereof, such as single variable domains (e.g., domain antibodies (DABs)), antigen-binding antibody fragments, Fab, F(ab'), Fv, disulfide-linked Fv, single-chain Fv, disulfide-linked scFv, diabodies, and TANDABS.
[0039] The terms "complete antibody," "whole antibody," or "intact antibody" have the same meaning herein and refer to a heteromeric glycoprotein with a molecular weight of approximately 150,000 daltons. An intact antibody consists of two identical heavy chains (HC) and two identical light chains (LC) linked by covalent disulfide bonds. This H2L2 structure folds to form three functional domains: two antigen-binding fragments known as "Fab" fragments and an "Fc" crystallizable fragment. The Fab fragment consists of an amino-terminal variable domain, variable heavy (VH) or variable light (VL), and a carboxyl-terminal constant domain, CH1 (heavy) and CL (light). The Fc fragment consists of two domains formed by the dimerization of paired CH2 and CH3 regions. The Fc can elicit effector functions by binding to receptors on immune cells or by binding to C1q, the first component of the classical complement pathway. The five classes of antibodies, IgM, IgA, IgG, IgE, and IgD, are defined by different heavy chain amino acid sequences, designated μ, α, γ, ε, and δ, respectively, and each heavy chain may be paired with a K or λ light chain. The majority of antibodies in serum belong to the IgG class, and human IgG exists in four isotypes (IgG1, IgG2, IgG3, and IgG4), which differ primarily in the sequences of their hinge regions.
[0040] The present invention relates to a method and system for evaluating the production stability of clonal cell lines to select them for use in the production of therapeutic protein products. Clonal cell line production stability evaluation is essential. For a clonal cell line to proceed to the manufacturing stage, it must produce consistent amounts of therapeutic protein over the entire production period (typically 3-6 months). A standard production stability evaluation involves growing the clonal cell line and inoculating production vessels over a 3-6 month period to reflect the length of the production period. To calculate production stability, product concentration measurements are taken for each production run, and the rate of change in product concentration throughout the production run is calculated. Generally, a clonal cell line that can maintain protein expression within 30% of its original peak product concentration during stability evaluation is considered stable.
[0041] In industry, for each therapeutic protein, approximately 50 clonal cell lines are typically advanced to production stability evaluation, from which a single clonal cell line deemed viable for manufacturing is selected. Therefore, the cell line development (CLD) process for biopharmaceutical manufacturing requires significant investment of time and resources. Data analysis can be time-consuming and inconsistent, potentially leading to the selection of unstable clones or / and the selection of poorly performing clones.
[0042] The high-throughput automation platform for advanced microscale bioreactors (AMBR 15) with working volumes of 10 to 15 mL offers the opportunity to standardize experimental workflows while collecting and systematically storing vast amounts of data. Data availability and automation pave the way for the development of appropriate data modeling frameworks to exploit the full capabilities of acquired measurements. The ultimate aim is to improve stability study design, reduce preventable experimental effort, and improve the stability characterization of future clonal cell lines.
[0043] Variables such as temperature, pH, dissolved oxygen (DO), and air flow can be monitored online during the manufacturing process. Soft sensors can use these monitored variables as input data to models that predict output target measurements. Such soft sensors have been developed for situations where online analyzers are not available or economically unfeasible for the process variables of interest. For example, models using online monitoring of variables such as temperature, pH, and DO have been used to predict endpoint offline measurements of product titer (Gunther et al., 2008).
[0044] In this disclosure, applicants have surprisingly found that incorporating product titer measurements as input data into the modeling framework described herein can accurately predict the production stability of clonal cell lines. In contrast, incorporating other variables measured during online measurements of the production process into the modeling framework did not predict stability in later clonal cell line production cells. Using the methods described herein, the overall process for selecting productively stable clonal cell lines during cell line development can be shortened by accurately predicting clonal cell line production stability at an early stage.
[0045] Using data from production stability testing, we can develop a model that facilitates early and reliable prediction of clonal cell line production stability. By analyzing the product concentration profiles of successive production runs during stability testing, an early "fingerprint" can be identified for the clonal cell line, allowing for prediction of subsequent production stability. The methodology developed in this way utilizes advanced data-driven modeling approaches along with machine learning classification techniques to assess production stability. Applying this methodology to the early stages of the CLD period, 3-6 months, makes it possible to triage clonal cell lines predicted to be productively unstable. This increases CLD capacity and shortens chemistry, manufacturing, and control (CMC) steps. The system described here allows for the selection of stable clonal cell lines for subsequent recombinant therapeutic protein production processes if the clonal cell lines are predicted or determined to be stable or unstable.
[0046] The sensitivity of a method for predicting or determining production stability is related to its accuracy in predicting stable clonal cell lines. The specificity of a method for predicting or discriminating production stability is related to its accuracy in predicting unstable clonal cell lines. The purpose of this system is to predict, discriminate, or select stable clonal cell lines during CLD for therapeutic protein production. Therefore, high sensitivity is more desirable than specificity.
[0047] FIG. 1 shows an example of a method for selecting a clonal cell line for therapeutic protein production. In step 101, multiple clonal cell lines are generated. This step may include cloning a nucleic acid sequence encoding a protein of interest into an expression vector and transfecting the sequence into a host cell line. Host cells expressing the gene of interest can then be isolated via separation from other host cells using techniques known in the art, such as FACS (fluorescence-activated cell sorting) or dilution cloning. Those skilled in the art will understand that the process for generating a set of clonal cell lines can be accomplished through any technique known in the art and used for this purpose. The method for selecting a clonal cell line may include generating multiple clonal cell lines. The initial set of clonal cell lines can then be used for production runs. Each set of clonal cell lines may then be fed and propagated so that they express the therapeutic protein encoded by the gene of interest.
[0048] In step 102, the concentration of the therapeutic protein expressed by each clonal cell line of the set of clonal cell lines (also referred to as the productivity titer of each clonal cell line) can be measured. This product concentration can be measured using techniques such as ELISA, HPLC, Western blot, immunoassay, detection of protein biological activity, FACS analysis, fluorescence microscopy, direct detection of fluorescent proteins by FACS analysis, spectrophotometry, or other techniques known in the art. Multiple product concentration measurements can be performed at each production run of the stability test. Measurement of the product concentration of each clonal cell line can be automated. Further, calculations can be made based on the productivity titer, including, but not limited to, the mean product concentration, standard deviation (standard deviation of the product concentration profile), skewness (degree of asymmetry), kurtosis (sharpness), delta (difference between two consecutive measurements for all product concentration profile data points), maximum product concentration (maximum of the product concentration profile), and max-min slope (slope of the line between the maximum and minimum product concentration profiles). These variables, alone or in combination, may be referred to as "clonal cell line product concentration profiles," "clonal cell line product concentration profile data," "titer profiles," "product concentration profiles," or "product concentration profile data." Differential variable values are calculated from an array of n-1 cells, where n is the number of product concentration profile data points for the production run. In this way, multiple differences can be calculated from a single product concentration profile. For example, differential 6 is defined as the difference between the sixth and seventh product concentration data points measured in the product concentration profile. Stability studies to evaluate the production stability of clonal cell lines may be performed at multiple bioreactor scales. For example, product concentration profile data can be generated from production runs at the following scales or vessel sizes: 384-well plates, 96-well plates, 48-well plates, 24-well plates, 12-well plates, 6-well plates, T25, T75, T150, AMBR 15, or AMBR 250.Those skilled in the art will further understand that steps 101 and 102 are repeated multiple times in each iteration of the process shown in FIG. 1 , and that each iteration may include multiple production runs of the initial set of clonal cell lines. Clonal cell line production stability testing typically consists of at least three consecutive production runs, and often four or more consecutive runs over a four- to six-month period. The method may include determining product concentration profile data for two, three, four, five, six, seven, eight, nine, or ten production runs. That is, product concentration profile data may be determined after each run of two, three, four, five, six, seven, eight, nine, or ten production runs. The method may include determining product concentration profile data for two or more, three or more, or four or more production runs. The method may include determining product concentration profile data for two production runs. The method may include determining product concentration profile data for three production runs. The method may include determining product concentration profile data for four production runs. The stability testing can consider a generational framework of 0 to 150 clonal cell lines, whereby each successive production run is performed with an increasing number of generations for each clonal cell line. The method can include determining product concentration profile data for clonal cell lines having at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, or at least 50 generations. The method can also include determining product concentration profile data for clonal cell lines having up to 50, up to 60, up to 70, up to 80, up to 90, up to 100, up to 110, up to 120, up to 130, up to 140, or up to 150 generations.
[0049] In step 103, the product concentration profile data collected and / or calculated in step 102 may be provided to a trained modeling framework. The utilized modeling framework may include multivariate latent variable modeling, multiway analysis, and evolutionary modeling. This modeling framework can be defined as evolutionary multiway projection onto latent structures (EMPLS) or evolutionary multiway principal component analysis (EMPCA). Multivariate latent variable modeling can be used to develop classification models that enable the modeling framework to distinguish stable from unstable clonal cell lines and assign them according to the probability of stability classification. Examples of multivariate latent variable modeling include partial least squares regression (PLS) (Brereton & Lloyd, 2014) and principal component analysis (PCA). Multivariate latent variable modeling may include PLS, PCA, EMPLS, EMPCA, evolutionary multiway projection onto latent structures discriminant analysis (EMPLS-DA), or any combination thereof.
[0050] A method using EMPLS-DA may include the following steps. - Multivariate latent variable classification, i.e., Projection to Latent Structures-Discriminant Analysis (Brereton & Lloyd, 2014), is used to distinguish between stable and unstable clonal cell lines and to provide probabilistic assignment of stability categories, taking into account measurement correlations (cross-correlation, auto-correlation, and time correlation); - a multi-way methodology to explain measurement differences between batches within a single experimental production run (Nomikos & MacGregor, 1994); and - To study sequential stability test production runs conducted over increasing numbers of clonal generations through an evolutionary modeling framework (Ramaker et al., 2005).
[0051] Partial least squares regression (PLS) (Geladi & Kowalski, 1986; Would et al., 1983) is a multivariate regression technique. PLS deals with large amounts of correlated input data, where M variables have been measured and relate them to C corresponding response variables collected in a matrix Y [N × C], and stored in a matrix of regressors X [NM] for N experiments performed on different clonal cell lines. The data are typically "autoscaled," i.e., mean-centered and scaled to unit variance to avoid the effects of different measurement units. PLS reduces the dimensionality of the original X space by finding orthogonal (i.e., independent) latent variables (LVs) that explain the variance of the input Xs that are most predictive of the response Y. Thus, PLS decomposes X and Y as follows:
number
[0052] In the above formula, T[N×A] and U[N×A] are the score matrices of X and Y, respectively. P[M×A] and Q[C×A] are the loading matrices of X and Y, respectively, whose columns are p a and q a The superscript T denotes transposition. E[N×M] and F[N×P] are the residuals of X and Y, respectively, which are minimized by least squares to obtain a good fit to the calibration data. t a and u a are the columns of T and U, respectively, and the regression coefficients b a The relationship is linear through . The scores project the original data onto a reduced area of LVs and identify the relationship between N experiments on different clones. The loadings are the direction cosines of LVs relative to the original variables and identify the correlation between variables. The residual matrix contains the non-systematic part of the dataset (if the number of LVs A is appropriately chosen). Usually, a small number of LVAs << min(N,M,P) is sufficient to retain important information in the variation of both X and Y, no matter how large the dimensions of these matrices are. Cross-validation (Wold, 1978) is usually used to determine the optimal number of LVs.
[0053] New observation x i If measurements of M regressors are available in x i By projecting into the model space, we can estimate each response as follows:
number
[0054] In the above equation, y is the vector of estimated responses, and W[M×A] is the vector of x i is the weight matrix for projecting the
[0055] PLS can be effectively used for multivariate classification in the so-called Projection to Latent Structures-Discriminant Analysis (PLS-DA) (Brereton & Lloyd, 2014). In such cases, the matrix Y consists of categorical variables (rather than continuous variables). In particular, for binary classification problems, the Y variable is represented as a "dummy" variable, where 1 indicates that the sample belongs to that class and 0 indicates that the sample does not belong to that class.
[0056] In this case, stable clonal cell lines (y i =1) and unstable clonal cell lines (y i = 0), there are two classes of new samples x i The class attribute of can be obtained from Equation 4. However, the estimation error e i :
number
number
number
number
number
number
number
number
number
number
[0057] The definition of confidence in a prediction allows us to adequately explain the following situation:
number
[0058] This method of multivariate latent variable modeling can be supervised or unsupervised. Supervised modeling is performed using a labeled dataset, i.e., prior knowledge of data outputs, to train an algorithm to classify the data and predict outcomes. Unsupervised modeling is performed without labeled output variables, allowing the algorithm to identify hidden patterns within the dataset. Unsupervised modeling can be used to determine whether differences in CLD platforms and processes affect the predictive ability of a modeling framework trained on a calibration dataset from another platform or process. Unsupervised modeling can also be used to determine whether a modeling framework distinguishes between stable and unstable clonal cell lines, i.e., whether product concentration profile data from an initial production run indicates a natural fingerprint of the final production stability of a clonal cell line. Therefore, unsupervised modeling analysis of product concentration profile data from different designs, processes, or platforms may enable understanding of the identification of production stability, expressed as a product concentration profile fingerprint, across a series of production runs within a production stability study. Various processes and platforms that can be explored through unsupervised multivariate latent variable modeling include, but are not limited to, alternative cell line transfection systems, changes in media composition, seeding conditions, and / or process parameters (such as pH and temperature ranges).
[0059] In step 104, the modeling framework can output an index of productive stability for each input clonal cell line. The result of this application can be a classification of each input clonal cell line as productively stable or unstable. This classification can be achieved using one or more classification methods as part of the multivariate latent variable modeling of the methods described herein. Examples of classification methods include, but are not limited to, discriminant analysis (DA), support vector machines (SVM), neural networks (NN), logistic regression (LR), k-nearest neighbors (KNN), decision trees (DT), naive Bayes (NB), random forests (RF), and gradient boosting (GB). Thus, in one example, the multivariate latent variable modeling of the methods described herein can include PLS discriminant analysis (PLS-DA). In another example, PLS can include PLS-DA, PLS-SVM, PLS-NN, PLS-LR, PLS-KNN, PLS-DT, PLS-NB, PLS-RF, PLS-GB, or any combination thereof.
[0060] This step can be performed through the calculation of the probability density function (PDF) of the input product concentration profile data. The intersection of the probability density functions (PDF) of the stable and unstable classes provides a threshold value that distinguishes whether a clonal cell line belongs to the stable class. If the probability value of a particular clonal cell line is greater than the threshold value, the clonal cell line is assigned to the stable class. Assigning clonal cell lines to the stable or unstable class facilitates the prediction, determination, or selection of productively stable clonal cell lines. Multivariate latent variable modeling can assign one or more clonal cell lines to the stable class. Multivariate latent variable modeling can assign one or more clonal cell lines to the unstable class.
[0061] Optionally, a confidence value may be assigned to the probability of a clonal cell line belonging to the stable or unstable class. If the difference between the probability value of a clonal cell line being stable and the probability value of an unstable cell line is high, medium, or low, the high probability class is predicted with high, medium, or low confidence, respectively. A confidence value can be assigned to the production stability of a clonal cell line. If the confidence value is high, the high probability class can be predicted with high confidence. A high confidence value can be defined as 0.9 or higher, 0.8 or higher, 0.7 or higher, or 0.6 or higher. A medium confidence value can be defined as 0.5 or higher, 0.4 or higher, or 0.3 or higher. A low confidence value can be defined as 0.3 or lower, 0.2 or lower, or 0.1 or lower. The determined production stability of one or more clonal cell lines can be assigned a high, medium, or low confidence value.
[0062] In step 105, clonal cell lines shown to be productively stable can be selected. This step suggests to those skilled in the art that it may involve deselecting clonal cell lines that have been marked as unstable. The resulting clonal cell lines may be selected for use in subsequent recombinant therapeutic protein production processes and may be used to generate or produce a therapeutic protein product.
[0063] Figure 2 illustrates an example of the type of multi-way methodology that can be used in the present invention, as described in connection with step 103 of Figure 1. This methodology is particularly important when considering batch processes because the classification problem must consider a third dimension of data: time. To handle such cases, a modeling framework incorporating a multi-way analytical structure may structure input data into a multi-dimensional array 201. The multi-dimensional array is a batch data matrix of product concentration profile variables, where multi-dimensional array 201 is a three-way array with dimensions of the experimental batch (N), the product concentration profile variables (M), and the total number of time samples collected during the entire experimental production run (K). To process the three-dimensional array via PLS, array 201 is unfolded through process 202 to form a two-dimensional array. This can be formed by cutting array 201 into vertical slices 203, which are then juxtaposed in a so-called batchwise unfolding system 204 (Nomikos & MacGregor, 1994), resulting in a two-dimensional matrix 205. PLS-DA can be applied to the resulting data matrix 205, which may be defined by dimensions [N × MK]. Batchwise unfolding 204 requires that all batches have the same length K. Feature variables from the variable profile, e.g., potency features from a potency profile, can be intentionally defined instead of using raw data (Meneghetti et al., 2016). The placement of features follows a batch-like structure. PLS-DA applied to the unfolded data matrix 205 is referred to as multiway PLS-DA (MPLS-DA).
[0064] Figure 3 shows an example of an evolutionary model structure that can be used in the process of Figure 1. Evolutionary model structure (Ramaker et al., 2005), also referred to as evolutionary modeling or evolutionary methodology, is a method by which data from the latest production runs of a production stability study can be incorporated into a classification model. As production runs are completed, data collected from them can be fed into a horizontally linked classification model (Ramaker et al., 2005). In this way, incremental information from the latest data can be utilized by the model to capture cross-correlations between production runs.
[0065] The "generation number" of a cell indicates the number of times the cell has propagated. The model is also constructed under the assumption that clonal cell lines associated with a particular production run have similar generation numbers across batches and project molecules. This allows for flexibility in handling different numbers of production runs (e.g., a project-independent model with fewer than four production runs across stability analyses). Evaluating the generation number distribution for each production run may allow for alignment of production runs across different projects. The methods described herein may further include evaluating or obtaining a generation number distribution for each production run.
[0066] Thus, the evolutionary model structure allows for the incremental information from successive production runs to be utilized and for the correlation between production runs to be captured. Once a production run is completed, the values of the product concentration profile variables within the production run are expanded using multiway analysis and incorporated into the classification model generated by multivariate latent variable modeling in the modeling framework. Thus, the classification model generated after each production run in the stability study facilitates the classification of clonal cell lines as stable or unstable.
[0067] Use of the evolutionary model architecture may include analysis of clonal cell line product concentration profiles of consecutive production runs over the growth period of clonal cell line generation. The evolutionary model architecture may include analysis of clonal cell line product concentration profiles of two to four consecutive production runs. The evolutionary model architecture may include analysis of clonal cell line product concentration profiles of two or more consecutive production runs. The evolutionary model architecture may include analysis of clonal cell line product concentration profiles of three or more consecutive production runs. PLS and PLS-DA, when applied to multiway analysis architecture and evolutionary model architecture, are referred to as EMPLS and EMPLS-DA, respectively. PCA applied to multiway analysis architecture and evolutionary model architecture is referred to as EMPCA.
[0068] Important variables for prediction Analyzing product concentration variables using the methods described above may allow for the determination of variables important for prediction (VIP) within a predictive model, if desired. The VIP score represents the relative contribution of a particular variable to the classification of a clonal cell line as stable or unstable. Variables with higher VIP scores may contribute more to predicting production stability than variables with lower VIP scores. Because VIP evaluation is affected by the calibration dataset used, the identified VIP may vary across individual production runs or stability studies. Analyzing multiple production runs of operational stability study datasets may allow for the identification of a subset of variables with higher VIP scores across simultaneous production runs, if desired. Thus, in some examples, the VIP of a product concentration profile may include the maximum, maximum-minimum slope, standard deviation, difference, or any combination thereof.
[0069] Calibration Data Set The modeling framework of the present invention may use a calibration dataset to train an algorithm to predict the stability class of one or more clonal cell lines. Therefore, a previously acquired dataset is reused as a calibration dataset to develop a model that predicts the production stability of a clonal cell line. The previously acquired calibration dataset is used as an input dataset for the modeling framework to generate an output model that can classify a clonal cell line as stable or unstable. Analysis of the clonal cell line product concentration profiles may include the use of a calibration dataset. The calibration dataset may also evaluate the performance of the modeling framework by comparing predicted and observed results. The calibration dataset used may include product concentration profile data from a therapeutic protein-producing clonal cell line. The calibration dataset may include product concentration profile data including one or more of the mean product concentration, standard deviation, skewness, kurtosis, difference, maximum product concentration, and maximum-minimum slope. The calibration dataset and the prediction dataset may include product concentration profile data from clonal cell lines derived from the same or similar parent cell line. For example, a calibration dataset may include product concentration profile data from one or more production runs of a clonal CHO cell line expressing recombinant therapeutic protein X, and a prediction dataset may include product concentration profile data from one or more production runs of a clonal CHO cell line expressing recombinant therapeutic protein Y. A calibration dataset may include product concentration profile data from multiple clonal cell lines. A calibration dataset may include product concentration profile data from two or more different clonal cell lines from production runs using different platforms or processes.
[0070] FIG. 4 shows a diagram of the production stability classification of clonal cell lines using a trained modeling framework incorporating a labeled calibration dataset. As can be seen, the labeled calibration dataset may be provided to the modeling framework to train it. At this point, clonal cell line product concentration profile data obtained in the manner described in connection with FIG. 1 may be input into the trained modeling framework in the manner described in connection with FIGS. 1-3. A stability classification may then be output from the trained modeling framework, as described in connection with FIG. 1. While not shown, it is suggested to one skilled in the art that this output may be used to select clonal cell lines for use in the production of a therapeutic protein product.
[0071] The stability of protein production by any clonal cell line can be determined by the methods and systems described herein. The clonal cell line may be a mammalian cell line, an insect cell line, a plant cell line, a yeast cell line, a Xenopus cell line, or a zebrafish cell line. The clonal cell line may be a mammalian cell line. The mammalian cell line may be a CHO (Chinese Hamster Ovary) cell line, a BHK cell line, an NS0 cell line, a Jurkat cell line, a K562 cell line, a HeLa cell line, a HEK293 cell line, a HEK293T cell line, or a PerC6 cell line. The mammalian cell line may be a CHO cell line. The mammalian cell line may be a CHO cell line expressing a monoclonal antibody.
[0072] The clonal cell line may be transfected or transformed. The clonal cell line can be transfected or transformed with a nucleic acid encoding a protein of interest. For example, the clonal cell line may be transfected or transformed with a nucleic acid encoding a monoclonal antibody or a fragment thereof. The clonal cell line may be transfected or transformed with a nucleic acid encoding a monoclonal antibody, a hormone, an anticoagulant, a blood factor, an interferon, a cytokine, a modified protein scaffold, an Fc fusion protein, an enzyme, or any other suitable therapeutic protein.
[0073] The methods described herein may be implemented in whole or in part on a computer. The methods for selecting clonal cell lines may be implemented on a computer. The modeling frameworks, including multivariate latent variable modeling, multifactorial analysis structures, and evolutionary model structures, may be implemented on a computer.
[0074] The present disclosure also provides a method for predicting production stability, comprising: a. inputting the clonal cell line product concentration profile data into a learning model comprising a modeling framework including multivariate latent variable modeling, a multifactorial analysis structure, and an evolutionary model structure; b. outputting an output from the learning model that indicates the production stability of the clonal cell line; and c. discriminating the clonal cell line as a stable cell line or an unstable cell line based on the output.
[0075] The present disclosure also provides a method for predicting the production stability of a clonal cell line, comprising the steps of: a. Receiving clonal cell line product concentration profile data; and b. processing the clonal cell line product concentration profile data with a system configured to process the clonal cell line product concentration profile data to predict clonal cell line production stability. A method is provided that includes:
[0076] Also disclosed herein is a system for determining the stability of clonal cell line production or for selecting clonal cell lines, comprising: a. an input for receiving clonal cell line product concentration profile data; b. A learning model for determining the production stability of a clonal cell line or for selecting a clonal cell line, the learning model comprising a modeling framework including multivariate latent variable modeling, a multifactorial analysis structure, and an evolutionary model structure; c. one or more processors for processing the clonal cell line product concentration profile data with the learning model; and d. An output section that provides an indicator of clonal cell line production stability based on processing the clonal cell line product concentration profile data by the learning model. A system is provided comprising:
[0077] In addition, in this disclosure, a. one or more processors; and b. A non-transitory computer-readable medium containing one or more programs executable by one or more processors to perform the methods described herein. A system is provided comprising:
[0078] Those skilled in the art will appreciate that certain embodiments relating to the methods described herein are applicable to the systems described herein, and vice versa.
[0079] Also provided in the present disclosure is a computer program comprising instructions that, when the computer program is executed by a computer or data processor, cause the computer to perform operations in accordance with the methods described herein.
[0080] Also provided in this disclosure is a computer-readable medium containing instructions that, when executed by a computer or data processor, cause the computer to perform operations in accordance with the methods described herein.
[0081] The computer program or computer-readable medium may include instructions that, when executed, implement (a) a learning model and / or any associated functions, and / or (b) a system for predicting clonal cell line production stability. The computer program or computer-readable medium may include instructions that, when executed, implement (a) a learning model and / or any associated functions, and / or (b) a system for determining clonal cell line production stability. The computer program or computer-readable medium may include instructions that, when executed, implement (a) a learning model and / or any associated functions, and / or (b) a system for selecting stable clonal cell lines. The computer program or computer-readable medium may include instructions that, when executed, result in an output indicative of clonal cell line production stability.
[0082] Also provided is a computer readable data carrier having stored thereon a computer program product as described herein. [Example]
[0083] Example 1: Scientific validation framework for potency signature fingerprints The production stability of a cell line is expressed as a measure of productivity loss over multiple runs with increasing generation numbers. Clonal cell lines tested in these production stability studies are typically defined as unstable if their recombinant protein titer decreases by 30% or more over 60-80 generations. Initial analyses (not shown) hypothesized that modeling titer profile variables during CLD may be able to predict the production stability of clonal cell lines.
[0084] Therefore, productivity titer measurements of clonal cell lines and their evolution over time may represent an early fingerprint of the clonal cell line's stability. To determine whether a clonal cell line exhibits this early stability fingerprint, the entire titer profile was used as a set of multivariate inputs for the modeling framework described herein. The titer features used as regressors in the modeling approach are listed in Table 1.
[0085] [Table 1]
[0086] Data analysis approaches can be described as either "supervised" or "unsupervised." "Supervised" analysis uses labeled input and output training datasets to teach a model how to accurately predict outcomes, while "unsupervised" analysis models an unlabeled dataset to find naturally occurring patterns in the training dataset.
[0087] Supervised learning models work well when the set of input regressors describes the classes, i.e., the regressors have a fingerprint of the regressor variables. This can be confirmed by exploring the data in an unsupervised manner so that cluster-related influences can be identified in advance. The EMPLS-DA modeling approach can be transformed into an unsupervised method by replacing the PLS-DA method with PCA (EMPCA), where score diagnostics are ideal for visualizing the clusters of all evolving models in a low-dimensional multivariate space.
[0088] EMPCA modeling was performed on two data sets (Set A and Set B) generated from different process platforms. Traditionally, these production stability studies are performed over four production runs before final stability measurements are made on individual clonal cell lines.
[0089] Production stability testing was performed on Set A and Set B over four production runs, totaling 150 and 80 generations, respectively. The model constructed after the first two production runs did not demonstrate strong cluster separation, suggesting a possible weak correlation between the titer profile features used as regressors in the first two production runs and the stability of the clonal cell lines (Figures 5 and 6). However, the model was able to distinguish stable from unstable clonal cell lines in both datasets after the third production run. Thus, the multivariate evolutionary approach (EMPCA) begins to systematically accumulate stability-related information after the third production run, thereby displaying a final stability call fingerprint on the clonal cell lines earlier than the full set of endpoint titer data required by traditional methods. These considerations based on two different stability test designs demonstrate the potential discriminatory power of supervised approaches that could be developed when operating in these situations. Different designs and process platforms would require similar unsupervised analysis to understand which production runs are represented by the stability fingerprint of the titer features. Platform processes that result in differently aggressive behavior of titer profiles need to be considered in separate models, as the extracted features are explanatory of titer dynamics.
[0090] Example 2: Testing in a project The performance of the developed model was tested by simulating its use under conditions where production stability data for the final clonal cell line was available, allowing for evaluation of model performance after each production run. Under these simulated conditions, each project was run sequentially, allowing for the availability of more project data in order to expand the calibration dataset. This approach allowed for testing how differences in the size and composition of the calibration dataset affected modeling predictions.
[0091] Four hypothetical scenarios were considered, each using a different calibration dataset, prediction dataset, process platform, and number of generations completed in the stability test (Table 2). The model was evaluated for misclassification error rate, specificity, and sensitivity after each calibration and prediction production run under these hypothetical scenarios. Calibration results indicate the robustness of the model at the time of calibration, while prediction results indicate actual predictive performance. Model calibration results with misclassification error less than 30% and sensitivity greater than 70% were considered satisfactory based on a risk-based approach in which sensitivity, i.e., stable prediction accuracy, is prioritized over unstable prediction accuracy.
[0092] In addition to predicting the production stability of clonal cell lines, the developed methodology also estimates class probabilities that provide an overall reliability for each clonal cell line. This reliability coefficient is calculated using the PLS-DA method, as described in detail herein. PLS-DA provides a method for predicting the production stability of a particular clonal cell line (P i ) is stable (P i,stable ) or unstable (P i,unstable ) can be used to describe whether the scenario prediction results are highly or low in confidence, depending on the difference in probability. Therefore, this model can provide additional information for low-confidence predictions, allowing for possible uncertainty.
[0093] [Table 2]
[0094] The final stability of the clonal cell lines in each project used to evaluate model calibration and predict performance is shown in Table 3 .
[0095] [Table 3]
[0096] In scenario 1, the model was calibrated with 46 clonal cell lines from one dataset (A) and measured for up to 150 cell generations in a stability study using a process platform (X). After the third production run, calibration performance was satisfactory, with an error rate of 7% or less, sensitivity of 93% or more, and specificity of 94% or more. However, the first two runs had a margin of error of 20%, which was still reasonable (Table 4). Prediction results showed error rates of 46 to 54% in the first two production runs, but the final two runs showed good prediction performance, with misclassification errors of 15% or less. The low sensitivity of prediction performance may be partly explained by the small number of stable clonal cell lines (4) in prediction dataset B. A detailed examination of prediction reliability after production run 4 revealed that although 3 / 3 of the misclassified stable clonal cell lines were flagged as low reliability, the reliability of all misclassified stable clonal cell lines was approximately 60% after the third run.
[0097] [Table 4]
[0098] In Scenario 2, the model was calibrated with 70 clonal cell lines across datasets A and B. Calibration performance was consistent across production runs, with almost equal absolute values compared to Scenario 1 (Table 5), where only one dataset was used for calibration. Meanwhile, very high sensitivity (100%) and specificity (over 92%) were observed after consecutive production runs, significantly improving predictive performance compared to Scenario 1.
[0099] [Table 5]
[0100] In scenario 3, the model was calibrated with 48 clonal cell lines from one dataset (C) using a process platform (Y) measuring up to 80 generations in stability testing. Calibration performance showed a misclassification error rate of only 19% after the first production run, decreasing to 0% after the third and fourth runs (Table 6). Prediction performance indicated that the model had sensitivity issues after the first two runs, with a misclassification error of approximately 30%. However, in the third analysis, the misclassification error was 16% (7 out of 45 clonal cell lines), and prediction performance was satisfactory, as 6 / 7 misclassified clonal cell lines were identified as unreliable predictions.
[0101] [Table 6]
[0102] Finally, in Scenario 4, we used a process platform (Y) measuring up to 80 generations in the stability test and calibrated the model with a total of 93 clonal cell lines from two datasets (C and D). Although the calibration performance was slightly reduced compared to Scenario 3, the prediction performance improved, with Test 2 acting as an outlier due to a shift in the number of generations (Table 7).
[0103] [Table 7]
[0104] Therefore, this modeling approach uses a fingerprint of clonal cell line production stability in productivity titer data to predict production stability as early as possible, intermixing with the traditional approach of full stability testing. The model described herein significantly reduces the time and resources required to conduct stability testing by providing a means for selecting the most promising stable clonal cell lines early in the cell line development process.
[0105] Example 3: Model testing within an ongoing project Following the initial evaluation detailed in Example 3, the modeling approach was tested on current projects and projects using new platforms. Tables 8, 9, and 10 show the model's predictive performance for each project (numbered 1–5) and platform (A or B) tested. Sensitivity was consistently high (80–100%) in Study 3, consistent with previous results. Specificity was generally low, but sometimes manifested as being particularly low for the more stable Platform B due to the low number of available, unstable clonal cell lines (referring to one unstable clonal cell line in Project 3 and two unstable clonal cell lines in Project 5). The overall error rate in Study 3 was low (typically below 10%). These are all measurements that demonstrate satisfactory model performance in predicting the production stability of clonal cell lines.
[0106] Project 2 behaved somewhat differently from the other analyzed projects, exhibiting an unusual level of confidence in the actual final production stability calculation in Test 3. Consequently, specificity in Test 3 was very low because all unstable clonal cell lines were erroneously predicted with low confidence, but stable clonal cell lines still performed well. Therefore, full model diagnostics were performed on Project 2, which showed that it was the only project in which the majority of clonal cell lines exhibited titer profile characteristics that differed from the model structure of multiple variables important for prediction (VIPs), which were not correlated. VIP scores for each variable were calculated according to the formula described in Andersen & Bro, 2010. Such discrepancies were found to be statistically significant in terms of confidence in previous projects where the model was routinely tested and run. This analysis was performed by comparing model residuals (both exploratory PCA and PLS) and their contributions to the VIPs of the calibrated model in Test 3. Therefore, Project 2 behaved as an outlier identified by diagnostics and is likely due to atypical clonal cell lines. Four out of five stable clonal cell lines were correctly predicted after Test 3 of Project 2. Therefore, this model provided effective results for the selection of stable clonal cell lines.
[0107] The high sensitivity of the Platform B project (98–100%) is likely due to the high percentage of stable clonal cell lines generated by this particular platform and the efficient prediction of stable clonal cell lines by the model.
[0108] [Table 8]
[0109] [Table 9]
[0110] [Table 10]
[0111] Example 4: VIP To determine the VIP within the model for predicting production stability, we performed an analysis of the VIP of the model constructed by combining two training datasets (Datasets C and D in Scenario 4 in Table 2 above). Figure 7 shows a plot of the VIP scores after Test 3 for each variable tested in the PLS-DA model. Variables with a VIP score greater than 1 are defined as VIPs that may have a significant impact on predicting the production stability of the clonal cell line. This analysis revealed three groups of VIPs after each production run: those with a low VIP index (close to 1), those with a medium VIP index, and those with a high VIP index after production runs 2 through 4, respectively (Table 11). Table 11 shows that the maximum value of production run 1 (Max R1), the maximum-minimum slope of production run 1 (Slope R1), and the difference 6 of production run 1 (Diff 6 R1) were identified as variables with high VIP scores in production runs 2, 3, and 4, respectively. The standard deviation of the titer profile in production run 1 (Deviation R1) was also identified as a variable with high VIP scores after production runs 3 and 4. Thus, statistical features of the titer profile that influence the prediction of clonal cell line production stability may include maximum, maximum-minimum slope, difference, and / or standard deviation.
[0112] [Table 11]
[0113] [Table 12]
Claims
1. 1. A method for selecting a clonal cell line for use in the production of a therapeutic protein, comprising: Measuring the product concentration of each clonal cell line for multiple clonal cell lines; determining product concentration profile data for each clonal cell line based on said product concentrations; inputting the product concentration profile data into a learning model including a modeling framework, wherein the modeling framework includes multivariate latent variable modeling, a multiway analytical structure, and an evolutionary model structure; using the learned model to generate an output indicative of the production stability of each clonal cell line; and selecting a clonal cell line for use in producing a therapeutic protein based on said output; A method comprising:
2. 1. A method for producing a therapeutic protein, comprising: Measuring the product concentration of each clonal cell line for multiple clonal cell lines; determining product concentration profile data for each clonal cell line based on said product concentrations; inputting the product concentration profile data into a learning model including a modeling framework, wherein the modeling framework includes multivariate latent variable modeling, a multiway analytical structure, and an evolutionary model structure; using the learned model to generate an output indicative of the production stability of each clonal cell line; and Selecting a clonal cell line for use in producing a therapeutic protein based on the output: A method comprising:
3. 3. The method of claim 1 or 2, wherein the multivariate latent variable modeling comprises at least one of partial least squares regression (PLS) and principal component analysis (PCA).
4. 4. The method of claim 3, wherein the PLS comprises at least one of PLS-Discriminant Analysis (PLS-DA), PLS-Support Vector Machine (PLS-SVM), PLS-Neural Network (PLS-NN), PLS-Logistic Regression (PLS-LR), PLS-k-Nearest Neighbor (PLS-KNN), PLS-Decision Tree Analysis (PLS-DT), PLS-Naive Bayes (PLS-NB), PLS-Random Forest (PLS-RF), or PLS-Gradient Boosting (PLS-GB).
5. 5. The method of any one of claims 1 to 4, wherein the clonal cell line product concentration profile data comprises one or more of mean product concentration, standard deviation, skewness, kurtosis, difference, maximum product concentration, and max-min slope.
6. 6. The method of any one of claims 1-5, wherein measuring the product concentration of each clonal cell line for the plurality of clonal cell lines comprises measuring the product concentration of each clonal cell line across multiple production runs.
7. 7. The method of any one of claims 1 to 6, wherein measuring the product concentration of each clonal cell line comprises measuring the product concentration of each clonal cell line for up to 150 generations.
8. The method of claim 6 or 7, wherein the measurement method further comprises obtaining a generation number distribution for each production run.
9. The method of any one of claims 1 to 8, wherein the evolutionary model construction comprises analyzing product concentration profile data of sequential production runs.
10. The method of any one of claims 1 to 9, wherein the clonal cell line is a mammalian cell line.
11. 11. The method of claim 10, wherein the mammalian cell line is a CHO cell line.
12. 12. The method of any one of claims 1 to 11, wherein the product concentration is measured at multiple bioreactor scales.
13. The method of any one of claims 1 to 12, further comprising using the selected clonal cell line to produce the therapeutic protein.
14. The method of any one of claims 1 to 13, wherein the product concentration comprises the concentration of a monoclonal antibody.
15. 1. A system for determining the production stability of a clonal cell line or for selecting a clonal cell line, comprising: a. Input for receiving clonal cell line product concentration profile data; b. A learning model for determining the production stability of or selecting a clonal cell line, the learning model comprising a modeling framework including multivariate latent variable modeling, a multifactorial analysis architecture, and an evolutionary model architecture; c. one or more processors for processing clonal cell line product concentration profile data using the learning model; and d. an output providing an indication of clonal cell line production stability based on processing the clonal cell line product concentration profile data with said learning model; A system comprising:
16. A computer program comprising instructions which, when executed by a computer or data processor, cause a computer to perform operations according to the method of any one of claims 1 to 14.
17. A computer readable medium comprising instructions which, when executed by a computer or data processor, cause the computer to perform operations according to the method of any one of claims 1 to 14.
18. 18. The program of claim 16 or the medium of claim 17, wherein the instructions, when executed, implement (a) a learning model and / or any associated functions, and / or (b) a system for predicting the stability of clonal cell line production.
19. 19. The computer program or medium of claim 16, 17, or 18, wherein the instructions, when executed, produce an output indicative of the production stability of the clonal cell line.
20. 17. A computer-readable data carrier having stored thereon a computer program according to claim 16.