Method for predicting production stability of cloned cell lines
By constructing a learning model of multivariate latent variable variable modeling and multi-directional analysis structure, predicting the production stability of cloned cell lines, solving the problem of time-consuming and resource-intensive evaluation of cell lines production stability in the prior art, and achieving more efficient biopharmaceutical production.
Patent Information
- Application Number
- CN202380078904.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-16
- Filing Date
- 2023-11-14
- Publication Date
- 2025-07-01
AI Technical Summary
The prior art is time-consuming and resource-intensive in evaluating cell line production stability, making it difficult to predict the production stability of cloned cell lines early, resulting in waste of resources and delays in production during biopharmaceutical development.
By measuring product concentrations of multiple cloned cell lines, a learning model including multivariate latent variable modeling, multidirectional analysis of structures and evolutionary model structures is constructed to generate outputs indicating the production stability of each cloned cell line, thereby selecting suitable cloned cell lines for the production of therapeutic proteins.
This method can predict the production stability of cloned cell lines with higher accuracy and efficiency, reduce resource waste and time extension during cell line development, and improve the efficiency of biopharmaceutical production.
Smart Images

Figure CN120239743A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to methods for developing cell lines for biopharmaceutical protein production, in particular methods and systems for determining or predicting the production stability of clonal cell lines, and methods and systems for selecting clonal cell lines. Background Art
[0002] Fully human antibodies can be obtained using a variety of methods (e.g., using yeast-based libraries capable of generating human antibody repertoires or transgenic animals (e.g., mice)). Yeast presenting human antibodies that bind to an antigen of interest on their surface can be selected using FACS (fluorescence-activated cell sorting)-based methods or by capture on beads using a labeled antigen. Transgenic animals that have been modified to express human immunoglobulin genes can be immunized with the antigen of interest and antigen-specific human antibodies isolated using B cell sorting techniques. The desired properties of the human antibodies produced using these techniques, such as affinity, developability, and selectivity, can then be characterized.
[0003] Mammalian cell lines are used as cell factories to create clonal cell lines that produce therapeutic proteins. The nucleic acid sequence encoding the protein of interest is cloned into an expression vector and then transfected into a host cell line. The transfected pool is expanded, single cells are sorted, and these sorted single cells are grown, and then the clonal cell lines are evaluated for production of the protein of interest. The clonal cell lines are ranked according to their product concentration and subjected to a series of sorting events until approximately 50 clonal cell lines are selected for production stability assessment.
[0004] Examples of such mammalian cell lines include murine myeloma cells (NS0), baby hamster kidney cells (BHK), human embryonic kidney cells (HEK293), and Chinese hamster ovary cells (CHO). More than 80% of currently approved recombinant proteins are produced in CHO cells (Butler & Spearman, 2014; Walsh, 2018). The reasons for the widespread use of CHO cells include their relative ease in receiving and expressing foreign DNA, their fast growth rate, the reliability of their protein folding machinery, their adaptability to serum-free suspension culture, and their ability to produce proteins with human-like post-translational modifications.
[0005] However, known cell lines (e.g., CHO cells) have a high level of chromosomal and genomic heterogeneity (Wurm & Wurm, 2017). Due to this genetic plasticity, cell lines may exhibit production instability, whereby a decline in the quantity and quality of recombinant proteins is observed during long-term culture (Dahodwala & Lee, 2019).
[0006] Therefore, biopharmaceutical manufacturers must demonstrate to regulatory agencies that the cells used for the production of therapeutic products maintain a stable protein quality over time. Cell line production stability tests are typically designed with 3 or 4 evaluation points spanning 60 to 80 cell line generations. The cell lines tested in these stability tests are generally defined as stable when they exhibit a less than 30% decrease in recombinant protein concentration (Dahodwala & Lee, 2019). Under this criterion, it is estimated that up to 63% of all CHO cell lines evaluated in production stability tests are classified as unstable (Dahodwala & Lee, 2019).
[0007] In the case of not maintaining the product concentration throughout the manufacturing period, the process yield can have a significant impact on the timeline as manufacturing schedules are typically booked at least one year in advance. Thus, unexpectedly low product yields can lead to repeated manufacturing runs that have a huge impact on scheduling and a chain reaction on product distribution. Therefore, these stability tests represent important but time-consuming and resource-intensive work for manufacturers.
[0008] In the context of the soft sensor field, modeling techniques have generally been applied to industrial process monitoring (Ramaker et al., 2005; Camacho et al., 2008; Gunther et al., 2008). Soft sensors can incorporate online measurements of process variables such as temperature, pH, and dissolved oxygen (DO) into the model to predict output variables. For example, Gunther et al. used online process variable measurements to predict product titer. However, a satisfactory prediction of the important quality of cell line production stability has not been achieved in the soft sensor field. Therefore, cell line production stability tests remain a time-consuming and expensive task.
[0009] Therefore, there is a current need for methods to reduce the amount of time and resources spent during cell line production stability tests. Summary of the Invention
[0010] According to a first aspect, there is provided a method for selecting a clonal cell line for the production of a therapeutic protein, the method comprising: for a plurality of clonal cell lines, measuring the product concentration of each clonal cell line; based on the product concentration, determining the product concentration profile data of each clonal cell line; inputting the product concentration profile data into a learning model comprising a modeling framework; wherein the modeling framework comprises multivariate latent variable modeling, multi-way analysis structure, and an evolutionary model structure; using the learning model, generating an output indicating the production stability of each clonal cell line; and based on the output, selecting a clonal cell line for the production of a therapeutic protein product.
[0011] According to a second aspect, there is provided a method for producing a therapeutic protein, the method comprising: for a plurality of clonal cell lines, measuring the product concentration of each clonal cell line; based on the product concentration, determining the product concentration profile data of each clonal cell line; inputting the product concentration profile data into a learning model comprising a modeling framework; wherein the modeling framework comprises multivariate latent variable modeling, multi-way analysis structure and evolutionary model structure; using the learning model, generating an output indicative of the production stability of each clonal cell line; and based on the output, selecting a clonal cell line for producing a therapeutic protein product.
[0012] According to a further aspect, there is provided a system for determining the production stability of a clonal cell line or selecting a clonal cell line, the system comprising: (a) an input for receiving product concentration profile data of a clonal cell line; (b) a learning model for determining the production stability of a clonal cell line or selecting a clonal cell line, the learning model comprising a modeling framework, the modeling framework comprising multivariate latent variable modeling, multi-way analysis structure and evolutionary model structure; (c) one or more processors for processing the product concentration profile data of the clonal cell line using the learning model; and (d) an output for providing an indication of the production stability of the clonal cell line based on the processing of the product concentration profile data of the clonal cell line by the learning model.
[0013] According to a further aspect, there is provided a computer program comprising instructions which, when the program is executed by a computer or one or more data processors, cause the computer to perform the operations of the method described herein.
[0014] According to a further aspect, there is provided a computer-readable medium comprising instructions which, when executed by a computer or (one or more) data processors, cause the computer to perform the operations of the method described herein.
[0015] According to a further aspect, there is provided a computer-readable data carrier having stored thereon the computer program described herein.
[0016] The methods and systems of the present disclosure are advantageous because they facilitate early and robust prediction of the production stability of clonal cell lines. In particular, incorporating the measurement results of the product concentration profiles of clonal cell lines into the modeling framework described herein provides a technical advantage, since the production stability of clonal cell lines can be predicted with higher accuracy and efficiency compared to other methods, to assist in the selection of clonal cell lines during the biopharmaceutical development process. By applying this method at an early stage of the cell line development (CLD) phase, it is possible to classify and predict clonal cell lines with production instability earlier, thereby increasing the CLD capacity and reducing the chemistry, manufacturing and controls (CMC) timeline.
[0017] Details of one or more embodiments of the present invention are set forth in the following accompanying description. Other features, objects, and advantages of the invention will be apparent from the specification and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 An exemplary flowchart showing the selection of a clonal cell line for the production of a therapeutic protein is shown.
[0019] Figure 2 A schematic diagram showing the batch unfolding of a multi-way array in the multivariate classification of the production stability of a clonal cell line is shown.
[0020] Figure 3 A schematic diagram showing an evolving multi-way classification procedure is shown.
[0021] Figure 4 An exemplary process for classifying the production stability of a clonal cell line using a trained modeling framework with a calibration data set incorporating markers is shown.
[0022] Figure 5 A 3D score plot of an unsupervised evolving multi-way principal component analysis (EMPCA) model for Group A is shown, where squares are stable clonal cell lines and circles are unstable clonal cell lines. The principal component space represents the dimensionality reduction resulting from the multivariate analysis of the titer characteristics, but without formal class supervision (unsupervised).
[0023] Figure 6 A 3D score plot of an unsupervised evolving multi-way principal component analysis (EMPCA) model for Group B is shown, where squares are stable clonal cell lines and circles are unstable clonal cell lines. The principal component space represents the dimensionality reduction resulting from the multivariate analysis of the titer characteristics, but without formal class supervision (unsupervised).
[0024] Figure 7 Variables important in prediction (VIP) for the PLS-DA model generated after production run #3 as discussed in Example 4 are shown. Variables with high VIP scores may be important for predicting the production stability of a clonal cell line. DETAILED DESCRIPTION
[0025] Before discussing specific embodiments with reference to the accompanying drawings, the following description of the embodiments is provided.
[0026] In a first aspect, a method for selecting a clonal cell line for the production of a therapeutic protein is provided, the method comprising: for a plurality of clonal cell lines, measuring the product concentration of each clonal cell line; based on the product concentration, determining the product concentration profile data of each clonal cell line; inputting the product concentration profile data into a learning model comprising a modeling framework; wherein the modeling framework comprises multivariate latent variable modeling, multi-way analysis structure, and an evolutionary model structure; using the learning model to generate an output indicative of the production stability of each clonal cell line; and based on the output, selecting a clonal cell line for producing a therapeutic protein product.
[0027] The multivariate latent variable modeling may include at least one of projection to latent structures (PLS) and principal component analysis (PCA). The PLS may include at least one of PLS discriminant analysis (PLS-DA), PLS support vector machine (PLS-SVM), PLS neural network (PLS-NN), PLS logistic regression (PLS-LR), PLS-k-nearest neighbor (PLS-KNN), PLS decision tree (PLS-DT), PLS naive Bayes (PLS-NB), PLS random forest (PLS-RF), and PLS gradient boosting (PLS-GB).
[0028] The product concentration profile data may include one or more of average product concentration, standard deviation, skewness, kurtosis, differential, maximum product concentration, and maximum-minimum gradient. The product concentration may include the concentration of a monoclonal antibody.
[0029] For a plurality of clonal cell lines, measuring the product concentration of each clonal cell line may include measuring the product concentration of each clonal cell line across multiple production runs. The product concentration of the clonal cell line may be measured for 2, 3, or 4 production runs. The product concentration of the clonal cell line may be measured for 2 production runs. The product concentration of the clonal cell line may be measured for 3 production runs. The product concentration of the clonal cell line may be measured for 4 production runs. The product concentration of the clonal cell line may be measured up to 150 generations. The method may further include obtaining the generation number distribution for each production run.
[0030] The evolutionary model structure may include analyzing the product concentration profiles of clonal cell lines in consecutive production runs.
[0031] The clonal cell line may be a mammalian cell line. The clonal cell line may be a CHO cell line.
[0032] The product concentration of the clonal cell line may be measured at multiple bioreactor scales. The product concentration of the clonal cell line may be measured at a bioreactor scale of 15 mL.
[0033] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs. All patents and publications mentioned herein are incorporated by reference in their entirety.
[0034] The term "comprising" encompasses "including" or "consisting of", e.g., a composition "comprising" X can consist solely of X or can include additional things, e.g., X + Y.
[0035] The term "consisting essentially of" limits the scope of a feature to the specified materials or steps and those materials or steps that do not materially affect the basic properties of the claimed feature(s).
[0036] The term "consisting of" does not include the presence of any additional components.
[0037] The term "about" in relation to a numerical value x means, for example, x ± 10%, 5%, 2% or 1%.
[0038] As used herein, the term "clonal cell line" or "clone" refers to an isolated host cell containing a gene of interest. Herein, isolation of the host cell means separation from other host cells using techniques known in the art, such as FACS (fluorescence-activated cell sorting) or dilution cloning. The clonal cell line can undergo a therapeutic protein production stability analysis as described herein, during which the isolated clonal cell line will be grown in cell culture. The cells grown in the cell culture will share a common ancestor of the corresponding clonal cell line.
[0039] As used herein, the term "clonal cell line production stability" or "production stability" refers to the stability of therapeutic protein production by a clonal cell line, i.e., production of a therapeutic protein with a consistent product concentration or titer within 50 to 150 generations. A consistent product concentration or titer can be defined as a decrease in the therapeutic protein product concentration or titer of < 30%.
[0040] As used herein, the terms "clonal cell line product concentration", "product concentration", "titer", or "productivity titer" refer to the amount of therapeutic protein produced by a clonal cell line, that is, the concentration of therapeutic protein produced by a clonal cell line. The concentration of therapeutic protein produced by a cell line can be measured using techniques such as ELISA, HPLC, Western blot, immunoassay, detection of protein biological activity, FACS analysis, fluorescence microscopy, direct detection of fluorescent proteins by FACS analysis, spectrophotometry, or other techniques known in the art. As used herein, the terms "clonal cell line product concentration profile", "clonal cell line product concentration profile data", "titer profile", "product concentration profile", or "product concentration profile data" refer to a set of variables (features) calculated based on the measured product concentration data for each clonal cell line. For example, the product concentration profile of a clonal cell line can include one or more of the following: average product concentration, standard deviation, skewness, kurtosis, difference, maximum product concentration, and maximum-minimum gradient.
[0041] As used herein, the term "production run" refers to the process of expressing the protein encoded by the gene of interest produced by a clonal cell line, whereby one or more clonal cell lines in a single vessel undergo feeding, growth, and therapeutic protein production phases. Production stability testing includes multiple production runs.
[0042] A "batch" can be defined as a single vessel of clonal cell lines that has undergone the complete feeding, growth, and therapeutic production phases of a production run. Thus, a "production run" includes multiple batches of individual clonal cell lines undergoing the complete feeding, growth, and therapeutic protein production phases for the production of a protein product. During production stability testing, the production stability of multiple clonal cell lines is evaluated, and each clonal cell line is present in a separate batch of each production run.
[0043] As used herein, the terms "number of generations" or "generation" refer to the number of times a clonal cell line has doubled. For example, a number of generations of 80 means that a clonal cell line has doubled 80 times.
[0044] As used herein, the term "product" refers to the protein produced by a clonal cell line. Thus, for the purposes of the described invention, the terms "protein product", "cell line product", and "therapeutic protein product" are interchangeable with the term "product". "Therapeutic protein production" refers to the process of producing a therapeutic protein product by a clonal cell line.
[0045] As used herein, the term "antibody" refers to a molecule having an immunoglobulin-like domain (e.g., IgG, IgM, IgA, IgD, or IgE), and includes monoclonal, recombinant, polyclonal, chimeric, human, humanized, multispecific antibodies (including bispecific antibodies) and heteroconjugate antibodies; single variable domains (e.g., domain antibodies (DAB)), antigen-binding antibody fragments, Fab, F(ab’)2, Fv, disulfide-linked Fv, single-chain Fv, disulfide-linked scFv, diabodies, tandabs, etc. and any modified versions of any of the foregoing.
[0046] The terms full antibody, whole antibody or intact antibody, which are used interchangeably herein, refer to a heterotetrameric glycoprotein having a molecular weight of approximately 150,000 daltons. An intact antibody consists of two identical heavy chains (HC) and two identical light chains (LC) linked by covalent disulfide bonds. This H2L2 structure folds to form three functional domains, which include two antigen-binding fragments (referred to as "Fab" fragments) and one "Fc" crystallizable fragment. The Fab fragment consists of the variable domain at the amino terminus (heavy chain variable (VH) or light chain variable (VL)) and the constant domain at the carboxyl terminus (CH1 (heavy chain) and CL (light chain)). The Fc fragment consists of two domains formed by dimerization of paired CH2 and CH3 regions. Fc can initiate effector functions by binding to receptors on immune cells or by binding Clq (the first component of the classical complement pathway). The five classes of antibodies, IgM, IgA, IgG, IgE and IgD, are defined by different heavy chain amino acid sequences, which are designated μ, α, γ, ε and δ, respectively, and each heavy chain can pair with a K or λ light chain. Most antibodies in serum belong to the IgG class, and there are four subtypes of human IgG (IgG1, IgG2, IgG3 and IgG4), which differ mainly in their hinge regions.
[0047] This application relates to methods and systems for assessing the production stability of clonal cell lines to select clonal cell lines for the production of therapeutic protein products. Assessment of the production stability of clonal cell lines is essential. In order for a clonal cell line to progress to the manufacturing stage, it must produce a consistent amount of therapeutic protein during the manufacturing window (usually 3 to 6 months). Standard production stability assessment involves scaling up the clonal cell line and inoculating production vessels over a period of 3 to 6 months to reflect the length of the manufacturing window. To calculate production stability, product concentration measurements are made for each production run, and the percentage change in product concentration over a time series is calculated. Typically, a clonal cell line that is able to maintain its protein expression within 30% of its original peak product concentration during the stability assessment is considered stable.
[0048] In an industrial setting, for each therapeutic protein, typically around 50 clonal cell lines are progressed to production stability assessment, from which a single clonal cell line deemed manufacturable will be selected. Thus, the cell line development (CLD) process for manufacturing biopharmaceutical products requires a significant investment of time and resources. Data analysis can involve considerable time expenditure and inconsistencies in methods, which can lead to the selection of unstable clones and / or the rejection of the best-performing clones.
[0049] A high-throughput automation platform with advanced micro-bioreactors (AMBR 15) having a working volume of 10 - 15 mL represents an opportunity to standardize the experimental workflow while collecting and systematically storing large amounts of data. Data availability and automation pave the way for developing appropriate data modeling frameworks to leverage the full power of the measurements obtained, with the ultimate goal of improving stability test design, reducing preventable experimental work, and enhancing the stability characteristics of future clonal cell lines.
[0050] Variables such as temperature, pH, dissolved oxygen (DO), and airflow can be monitored online during the industrial process. Soft sensors can utilize these monitored variables as input data in models to continue predicting output target measurements. Such soft sensors have been developed when online analyzers are not available or economically infeasible for the process variables of interest. For example, models using online monitoring of variables including temperature, pH, and DO have been used to predict the off-line measurement results of product titer (Gunther et al., 2008).
[0051] In the present disclosure, the inventors have surprisingly found that product titer measurements can be incorporated as input data into the modeling framework described herein to accurately predict the production stability of clonal cell lines. In contrast, incorporating other variables measured during the online monitoring of the industrial process into the modeling framework does not predict the subsequent production stability of clonal cell lines. Using the method described herein, the timeline for selecting a production-stable clonal cell line during cell line development can be shortened by accurately predicting the production stability of clonal cell lines at an earlier stage.
[0052] Production stability test data can be used to develop models that help in the early robust prediction of the production stability of clonal cell lines. By analyzing the product concentration profiles of consecutive production runs during the stability test, early "fingerprints" can be identified in clonal cell lines, which are predictive of later production stability. The developed method utilizes an advanced data-driven modeling approach that combines multivariate analysis diagnostic capabilities with machine learning classification techniques to evaluate production stability. By applying this method in the early stage of the 3- to 6-month CLD period, it is possible to classify and predict earlier clonal cell lines that are production-unstable, thereby increasing CLD capacity and reducing the chemistry, manufacturing, and controls (CMC) timeline. When a clonal cell line is predicted or determined to be stable or unstable by the method described herein, stable clonal cell lines can be selected for implementation in subsequent recombinant therapeutic protein production processes.
[0053] The sensitivity of the method for predicting or determining production stability is related to the accuracy of the prediction of stable clonal cell lines. The specificity of the method for predicting or determining production stability is related to the accuracy of the prediction of unstable clonal cell lines. Since the purpose of the method is to predict, determine, or select stable clonal cell lines during the CLD to produce therapeutic proteins, high sensitivity is preferred over specificity.
[0054] Figure 1 An exemplary method of selecting a clonal cell line for the production of a therapeutic protein is shown. In step 101, a plurality of clonal cell lines are generated. This step 101 can include cloning a nucleic acid sequence encoding a protein of interest into an expression vector and transfecting the sequence into a host cell line. Then the host cells expressing the gene of interest can be isolated from other host cells by using techniques known in the art (such as FACS (fluorescence-activated cell sorting) or dilution cloning). Those skilled in the art will understand that the process of generating a set of clonal cell lines can be achieved by any technique known in the art and used for this purpose. The method for selecting a clonal cell line can include the step of generating a plurality of clonal cell lines. Then an initial group of clonal cell lines can be subjected to a production run, where each group of clonal cell lines can be fed and grown such that they express the therapeutic protein encoded by the gene of interest.
[0055] In step 102, the concentration of the therapeutic protein expressed by each clone cell line in the clone cell line group (also referred to as the productivity titer of each clone cell line) can be measured. This product concentration can be measured using techniques such as ELISA, HPLC, Western blotting, immunoassay, detection of protein bioactivity, FACS analysis, fluorescence microscopy, direct detection of fluorescent proteins by FACS analysis, spectrophotometry, or other techniques known in the art. Multiple product concentration measurements can be performed in each production run of the stability test. The measurement of the product concentration of each clone cell line can be automated. Other data can be calculated based on the productivity titer, including (but not limited to) average product concentration, standard deviation (standard deviation of the product concentration profile), skewness (degree of asymmetry), kurtosis (sharpness), difference (difference between two consecutive measurements of all product concentration profile data points), maximum product concentration (maximum value of the product concentration profile), and maximum-minimum gradient (gradient of the line between the maximum product concentration and the minimum value of the product concentration profile). Any one of these variables (alone or in combination) can be referred to as "clone cell line product concentration profile", "clone cell line product concentration profile data", "titer profile", "product concentration profile", or "product concentration profile data". The difference variable value is calculated from an array of n - 1 units, where n is the number of data points of the product concentration profile of the production run. Thus, multiple differences can be calculated from one product concentration profile. For example, difference 6 is defined as the difference between the sixth and seventh product concentration data points measured in the product concentration profile. The stability test for evaluating the production stability of clone cell lines can be performed at multiple bioreactor scales. For example, the product concentration profile data can be generated according to the production run at scales or vessel sizes such as 384-well plates, 96-well plates, 48-well plates, 24-well plates, 12-well plates, 6-well plates, T25, T75, T150, AMBR 15, or AMBR 250. Those skilled in the art will also understand that steps 101 and 102 can be performed in Figure 1This is repeated multiple times in each iteration of the process shown, and each iteration can include multiple production runs of the initial clone cell line group. The production stability test of the clone cell line typically includes at least 3 consecutive production runs, usually 4 or more consecutive production runs during a period of 4 to 6 months. The method can include determining product concentration profile data for 2, 3, 4, 5, 6, 7, 8, 9, or 10 production runs. That is, the product concentration profile data can be determined after each production run of 2, 3, 4, 5, 6, 7, 8, 9, or 10 production runs. The method can include determining product concentration profile data for 2 or more, 3 or more, or 4 or more production runs. The method can include determining product concentration profile data for 2 production runs. The method can include determining product concentration profile data for 3 production runs. The method can include determining product concentration profile data for 4 production runs. The stability test can consider a window of 0 to 150 clone cell line generations, whereby for each clone cell line each consecutive production run is performed with an increasing number of generations. The method can include determining product concentration profile data for clone cell lines having at least 15, at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, or at least 50 generations. The method can include determining product concentration profile data for clone cell lines having up to 50, up to 60, up to 70, up to 80, up to 90, up to 100, up to 110, up to 120, up to 130, up to 140, or up to 150 generations.
[0056] In step 103, the product concentration profile data collected and / or calculated in step 102 can be provided to a trained modeling framework. The modeling frameworks utilized can include multivariate latent variable modeling, multi-way analysis structures, and evolving model structures. Collectively, such modeling frameworks can be defined as an evolving multi-way projection of latent structure (EMPL) or evolving multi-way principal component analysis (EMPCA). The multivariate latent variable modeling can be used to develop a classification model that enables the modeling framework to distinguish stable from unstable clone cell lines and give a probabilistic assignment of the stability class. Examples of multivariate latent variable modeling include projection to latent structure (PLS) (Brereton & Lloyd, 2014), also known as partial least squares, and principal component analysis (PCA). The multivariate latent variable modeling can include PLS, PCA, EMPLS, EMPCA, evolving multi-way projection of discriminant analysis of latent structure (EMPLS-DA), or any combination thereof.
[0057] In a method that utilizes EMPLS-DA, the method can include:
[0058] · Multivariate latent variable classification method, i.e., projection-discriminant analysis of latent structure (Brereton & Lloyd, 2014), is used to distinguish whether the cloned cell line is stable or unstable, also considering measurement correlations (cross-correlation, auto-correlation, and temporal correlation) and giving a probabilistic attribution of the stability class;
[0059] · Multivariate approach (Nomikos & MacGregor, 1994), is used to explain the measurement differences between batches in a single experimental production run; and
[0060] · Evolutionary modeling structure (Ramaker et al., 2005), is used to consider the progress of consecutive stability pilot production runs carried out with increasing clone generations.
[0061] Projection to latent structure (PLS) (Geladi & Kowalski, 1986; Wold et al., 1983) is a multivariate regression technique. PLS processes a large amount of correlated input data stored in the regressor X [N×M] matrix for N experiments conducted on different cloned cell lines (where M variables are measured), and correlates them with C corresponding response variables collected in the matrix Y [N×C]. The data is usually "auto-scaled", i.e., centered at the mean and scaled to unit variance to avoid the influence of different measurement units. PLS reduces the dimension of the original X space by finding A orthogonal (i.e., independent) latent variables (LVs), which explain the input X variance that is most predictive of the response Y. Thus, PLS decomposes X and Y into:
[0062]
[0063] u a =b a t a (3)
[0064] where T [N×A] and U [N×A] are the score matrices of X and Y respectively; P [M×A] and Q [C×A] are the loading matrices of X and Y respectively, whose columns are p a and q a ; the superscript T represents transpose; E [N×M] and F [N×P] are the residuals of X and Y respectively, minimized in the least squares sense to properly fit the calibration data; t a and u a are the columns of T and U respectively, and are related by the regression coefficient b aLinear correlation. The scores are the projections of the original data into the reduced space of the LVs and identify the relationships in the N experiments of different clones; the loadings are the direction cosines of the LVs with respect to the original variables and identify the correlations between variables; the residual matrix includes the non-systematic part of the data set (when the number A of LVs is appropriately chosen). Usually, a small number of LVs A << min(N, M, P) is sufficient to retain the important information in the variability of X and Y, regardless of the dimensions of these matrices. Cross-validation (Wold, 1978) is usually used to determine the optimal number of LVs.
[0065] When measurements of M regressors are available for a new observation x i then the corresponding response can be estimated by projecting x i onto the model space according to the following equation:
[0066]
[0067] where is the vector of the estimated response and W[M×A] is the weight matrix of the projection of x i onto the latent space.
[0068] PLS can be effectively used for multivariate classification, namely the so-called projection-discriminant analysis of latent structures (PLS-DA) (Brereton & Lloyd, 2014). In this case, the matrix Y consists of categorical variables (instead of continuous variables). In particular, in a binary classification problem, the Y variables are represented as "dummy" variables, where 1 indicates that the sample belongs to the class and 0 indicates that the sample does not belong to the class.
[0069] In this case, there are two classes: stable clonal cell lines (y i = 1) and unstable clonal cell lines (y i = 0). The class membership of a new sample x i can be obtained from Equation 4. However, for class membership, there is always an estimation error e i :
[0070]
[0071] For this reason, the distribution of the estimates i around the true value y is calculated from all calibration data (whose classes are known a priori). The probability density function (PDF) is assumed for all i of the calibration data set around the true value y The distribution from 0 to N is calculated as a Gaussian distribution. Additionally, the intersection among the PDFs of the two categories (stable and unstable) provides a threshold value th, which differentiates whether the estimated clone cell line i belongs to the stable category. Further, the cumulative density function (CDF) is calculated based on the above PDFs of all categories, in such a way that the inversion of the CDF provides the probability P associated with each category. In summary, not only can each new clone cell line i be assigned to a category (i.e., stable or unstable) based on the fact of or or but also the confidence associated with this decision can be understood, that is, the probability (P i,稳定 / i,不稳定 ) that the clone cell line i actually belongs to either of the two categories. To simplify the concept of the two-category probabilities into a single estimated value, an arbitrary definition of "confidence" in the prediction is defined as:
[0072] V i = |P i, stable - P i, unstable| (6)
[0073] This definition of "confidence" in the prediction ensures that the following scenarios are appropriately described:
[0074] - If P i,稳定 >> P i,不稳定 , or vice versa, the difference between the two values is high, and thus the category with the higher probability is also predicted with high confidence;
[0075] - If P i,稳定 > P i,不稳定 , or vice versa, the difference between the two values is medium, and thus the category with the higher probability is also predicted with medium confidence;
[0076] - If P i,稳定 ≈ P i,不稳定 , the difference is low, and thus the confidence is low, being of little importance for the stable or unstable prediction.
[0077] The multivariate latent variable modeling of the method can be supervised or unsupervised. Supervised modeling is performed using a labeled data set, i.e., there is prior knowledge of the data output, and its training algorithm classifies the data and predicts the results. Unsupervised modeling is performed in the absence of a labeled output variable, enabling the algorithm to identify hidden patterns in the data set. The use of unsupervised modeling can be used to determine whether differences in the CLD platform and process affect the predictive ability of a modeling framework trained on a calibration data set from another platform or process. Unsupervised modeling can also be used to determine whether the modeling framework differentiates between stable and unstable clonal cell lines, i.e., the unsupervised modeling framework determines whether the product concentration profile data from an early production run exhibits the natural fingerprint of the final production stability of the clonal cell line. Thus, unsupervised modeling analysis of the product concentration profile data of different designs, processes, or platforms can reveal which production runs' production stability differentiates as the fingerprint of the product concentration profile over a series of production runs within a production stability test. Different processes and platforms that can be explored by unsupervised multivariate latent variable modeling include, but are not limited to, alternative cell line transfection systems, changes in media composition, inoculation conditions, and / or process parameters (such as pH or temperature range).
[0078] In step 104, the modeling framework can output an indication of the production stability of each clonal cell line for which input has been provided. This indication can classify each clonal cell line for which input has been provided as either production-stable or unstable. One or more classification methods can be used as part of the multivariate latent variable modeling of the method described herein to achieve this classification. Examples of classification methods include, but are not limited to, discriminant analysis (DA), support vector machine (SVM), neural network (NN), logistic regression (LR), k-nearest neighbor (KNN), decision tree (DT), naive Bayes (NB), random forest (RF), or gradient boosting (GB). Thus, in one example, the multivariate latent variable modeling of the method described herein can include PLS-discriminant analysis (PLS-DA). According to other examples, PLS can include PLS-DA, PLS-SVM, PLS-NN, PLS-LR, PLS-KNN, PLS-DT, PLS-NB, PLS-RF, PLS-GB, or any combination thereof.
[0079] This step can be performed by calculating the probability density function (PDF) of the input product concentration profile data. The intersection between the probability density functions (PDFs) of the stable and unstable classes provides a threshold for discriminating whether a cloned cell line belongs to the stable class. If the probability value of a specific cloned cell line is greater than the threshold, then the cloned cell line will be assigned to the stable class. Assigning a cloned cell line to a stable or unstable class helps to predict, identify, or select a production-stable cloned cell line. Multivariate latent variable modeling can assign one or more cloned cell lines to the stable class. Multivariate latent variable modeling can assign one or more cloned cell lines to the unstable class.
[0080] Optionally, a confidence value can be assigned to the probability of a cloned cell line belonging to the stable or unstable class. If the difference between the probability values for the stable and unstable cloned cell lines is high, medium, or low, then the class with the higher probability is predicted with high, medium, or low confidence, respectively. The production stability of a cloned cell line can be assigned a confidence value. The confidence value can be high, and thus the class with the higher probability can be predicted with high confidence. A high confidence value can be defined as ≥0.9, ≥0.8, ≥0.7, or ≥0.6. A medium confidence value can be defined as ≥0.5, ≥0.4, or ≥0.3. A low confidence value can be defined as ≤0.3, ≤0.2, or ≤0.1. The production stability of one or more identified cloned cell lines can be assigned a high confidence value, a medium confidence value, or a low confidence value.
[0081] In step 105, a cloned cell line that has been indicated as production-stable can be selected. Those skilled in the art will understand that this step can include deselecting a cloned cell line that has been marked as unstable. The resulting cloned cell line can be selected for subsequent recombinant therapeutic protein production processes and can be used to generate or produce a therapeutic protein product.
[0082] Figure 2 Illustrative examples of the types of multi-directional methods that can be used in the present invention are provided, and in particular as combined with Figure 1As described in step 103. This method is particularly important when considering batch processing, as the classification problem must take into account the third dimension of the data, i.e., time. To handle such a situation, a modeling framework including a multi-way analysis structure can structure the input data into a multi-dimensional array 201. The multi-dimensional array is a batch data matrix of product concentration profile variables, where the multi-dimensional array 201 is a three-way array with dimensions of experimental batches (N), product concentration profile variables (M), and the total number of time samples (K) collected during the entire experimental production run. To process the three-dimensional array by PLS, the array 201 is unfolded through process 202 to form a two-dimensional array. This can be done by cutting the array 201 into vertical slices 203, and then placing the vertical slices side by side in a so-called batch unfolding manner 204 (Nomikos & MacGregor, 1994), resulting in a two-dimensional matrix 205. PLS-DA can be applied to the resulting data matrix 205, which can be defined by dimensions [N×MK]. Batch unfolding 204 requires that all batches have the same length K. Characteristic variables can be specifically defined from the variable profiles (e.g., titer characteristics from the titer profile) instead of using the raw data (Meneghetti et al., 2016). The arrangement of the characteristics will still follow the batch-wise structure. PLS-DA applied to the unfolded data matrix 205 is called multi-way PLS-DA (MPLS-DA).
[0083] Figure 3 provides illustrative examples of evolution model structures that can be used in Figure 1 the process. The evolution model structure (Ramaker et al., 2005), also known as evolution modeling or the evolution method, is a method that allows data from the most recent production runs in a production stability test to be incorporated into a classification model. When a production run is completed, the data collected from it can be included in the classification model that horizontally concatenates the data of that run (Ramaker et al., 2005). In this way, the incremental information from the most recent data can be utilized by the model to capture cross-correlations across production runs.
[0084] The "number of cell generations" represents the number of times a cell has doubled. It is also assumed that the clonal cell lines associated with a certain production run should have similar numbers of cell generations across batches and project molecules to build the model. This allows for flexible handling of different numbers of production runs (e.g., independent models for projects where the number of production runs in the entire stability study < 4). Evaluation of the distribution of the number of cell generations for each production run can allow for the alignment of production runs from different projects. The methods described herein can also include evaluating or obtaining the distribution of the number of cell generations for each production run.
[0085] Thus, the evolving model structure utilizes the incremental information generated by sequential production runs and allows for the capture of cross-correlations across production runs. When a production run is completed, the values of the product concentration profile variables within the production run are unfolded by multi-way analysis and then included in the classification model generated by the multivariate latent variable modeling of the modeling framework. In this way, the classification model generated after each production run of the stability test facilitates the classification of clone cell lines as stable or unstable.
[0086] The use of the evolving model structure can include analyzing the clone cell line product concentration profiles of successive production runs in increasing clone cell line generations. The evolving model structure can include analyzing the clone cell line product concentration profiles of 2 to 4 successive production runs. The evolving model structure can include analyzing the clone cell line product concentration profiles of 2 or more successive production runs. The evolving model structure can include analyzing the clone cell line product concentration profiles of 3 or more successive production runs. PLS and PLS-DA applied to the multi-way analysis structure and the evolving model structure are referred to as EMPL and EMPLS-DA, respectively. PCA applied to the multi-way analysis structure and the evolving model structure is referred to as EMPCA.
[0087] Variables important for prediction
[0088] Analysis of the product concentration variables in the above method can optionally allow for the determination of variables important for prediction (VIP) in the prediction model. The VIP score represents the relative contribution that a given variable must have to classify a clone cell line as stable or unstable. Variables with high VIP scores may contribute relatively more to the prediction of production stability than variables with low VIP scores. Since the evaluation of VIP is affected by the calibration dataset used, the identified VIPs may vary between individual production runs or stability tests. By analyzing multiple production runs of the production stability test dataset, a subset of variables with high VIP scores across production runs can optionally be determined. Thus, in some examples, the VIP of the product concentration profile can include any one of a maximum value, a maximum-minimum gradient, a standard deviation, a difference, or any combination thereof.
[0089] Calibration dataset
[0090] The modeling framework of the present invention can use a calibration dataset to train an algorithm to predict the stability class of one or more clonal cell lines. Thus, a previously obtained dataset is reused as the calibration dataset to develop a model to predict the production stability of clonal cell lines. The previously obtained calibration dataset serves as the input dataset for the modeling framework to produce an output model that can classify clonal cell lines as stable or unstable. Analyzing the product concentration profiles of clonal cell lines can include using the calibration dataset. The calibration dataset also allows the performance of the modeling framework to be evaluated by comparing the predicted results with the observed results. The calibration dataset used can include product concentration profile data from clonal cell lines producing therapeutic proteins. The calibration dataset can include product concentration profile data, and the product concentration profile data includes one or more of average product concentration, standard deviation, skewness, kurtosis, difference, maximum product concentration, and maximum-minimum gradient. The calibration dataset and the prediction dataset can include product concentration profile data of clonal cell lines derived from the same or similar parental cell lines. For example, the calibration dataset can include product concentration profile data of one or more production runs of a clonal CHO cell line expressing recombinant therapeutic protein X, and the prediction dataset can include product concentration profile data of one or more production runs of a clonal CHO cell line expressing recombinant therapeutic protein Y. The calibration dataset can also include product concentration profile data of more than one clonal cell line. The calibration dataset can include product concentration profile data of two or more different clonal cell lines from production runs using different platforms or processes.
[0091] Figure 4 An illustration of the production stability classification of clonal cell lines using a trained modeling framework incorporating a labeled calibration dataset is provided. As can be seen, the labeled calibration dataset can be provided to the modeling framework to train the modeling framework. At this time, the clonal cell line product concentration profile data obtained as described in connection with Figure 1 can be input into the trained modeling framework as described in connection with Figures 1 to 3 . Then a stability classification can be output from the trained modeling framework as described in connection with Figure 1 . Those skilled in the art will understand that although not shown, this output can be used to select a clonal cell line for producing a therapeutic protein product.
[0092] The stability of protein production in any cloned cell line can be determined by the methods and systems described herein. The cloned cell line can be a mammalian cell line, an insect cell line, a plant cell line, a yeast cell line, an African clawed frog cell line, or a zebrafish cell line. The cloned cell line can be a mammalian cell line. The mammalian cell line can be a CHO (Chinese hamster ovary) cell line, a BHK cell line, an NS0 cell line, a Jurkat cell line, a K562 cell line, a HeLa cell line, a HEK293 cell line, a HEK293T cell line, or a PerC6 cell line. The mammalian cell line can be a CHO cell line. The mammalian cell line can be a CHO cell line expressing a monoclonal antibody.
[0093] The cloned cell line can be transfected or transformed. The cloned cell line can be transfected or transformed with a nucleic acid encoding a protein of interest. For example, the cloned cell line can be transfected or transformed with a nucleic acid encoding a monoclonal antibody or a fragment thereof. The cloned cell line can be transfected or transformed with a nucleic acid encoding a monoclonal antibody, a hormone, an anticoagulant, a blood factor, an interferon, a cytokine, an engineered protein scaffold, an Fc fusion protein, an enzyme, or any other suitable therapeutic protein.
[0094] The methods described herein can be wholly or partly computer-implemented. The method for selecting a cloned cell line can be a computer-implemented method. The modeling framework including multivariate latent variable modeling, multi-way analysis structure, and evolving model structure can be computer-implemented.
[0095] This article also provides a method for predicting production stability, the method including:
[0096] a. Inputting the cloned cell line product concentration profile data into a learning model including a modeling framework, the modeling framework including multivariate latent variable modeling, multi-way analysis structure, and evolving model structure;
[0097] b. Outputting from the learning model an output indicating the production stability of the cloned cell line; and
[0098] c. Identifying the cloned cell line as stable or unstable based on the output.
[0099] There is also provided a method for predicting the production stability of a cloned cell line, the method including the following steps:
[0100] a. Receiving the cloned cell line product concentration profile data; and
[0101] b. Processing the cloned cell line product concentration profile data with a system configured to process the cloned cell line product concentration profile data to predict the production stability of the cloned cell line.
[0102] Also provided is a system for determining the production stability of a cloned cell line or selecting a cloned cell line, the system comprising:
[0103] a. An input for receiving data on the concentration profile of the cloned cell line product;
[0104] b. A learning model for determining the production stability of a cloned cell line or selecting a cloned cell line, the learning model comprising a modeling framework, the modeling framework comprising multivariate latent variable modeling, multi-way analysis structure, and an evolutionary model structure;
[0105] c. One or more processors for processing the data on the concentration profile of the cloned cell line product using the learning model; and
[0106] d. An output for providing an indication of the production stability of the cloned cell line based on the processing of the data on the concentration profile of the cloned cell line product by the learning model.
[0107] Also provided is a system, the system comprising:
[0108] a. One or more processors; and
[0109] b. A non-transitory computer-readable storage medium comprising one or more programs executable by the one or more processors for performing the methods described herein.
[0110] Those skilled in the art will understand that certain embodiments related to the methods will be applicable to the systems described herein, and vice versa.
[0111] Also provided is a computer program comprising instructions that, when the program is executed by a computer or one or more data processors, cause the computer to perform the operations of the methods described herein.
[0112] Also provided is a computer-readable medium comprising instructions that, when executed by a computer or one or more data processors, cause the computer to perform the operations of the methods described herein.
[0113] The computer program or computer-readable medium may include instructions which, when executed, implement (a) a learning model and / or any associated functions and / or (b) a system to predict the production stability of a cloned cell line. The computer program or computer-readable medium may include instructions which, when executed, implement (a) a learning model and / or any associated functions and / or (b) a system to determine the production stability of a cloned cell line. The computer program or computer-readable medium may include instructions which, when executed, implement (a) a learning model and / or any associated functions and / or (b) a system to select a stable cloned cell line. The computer program or computer-readable medium may include instructions which, when executed, output an output indicative of the production stability of a cloned cell line.
[0114] There is also provided a computer-readable data carrier having stored thereon the computer program product described herein.
[0115] Examples
[0116] Example 1: Scientific validation framework for titer feature fingerprint
[0117] Cell line production stability is expressed as a measure of productivity loss over multiple runs at increasing generations. Cloned cell lines tested in these production stability trials are generally defined as unstable when they exhibit a recombinant protein titer drop of greater than 30% over 60 to 80 generations. Initial analysis (not shown) generated the hypothesis that modeling of titer profile variables during CLD could allow prediction of cloned cell line production stability.
[0118] Thus, the productivity titer measurements of a cloned cell line and its evolution over time can represent an early fingerprint of cloned cell line stability. To test whether a cloned cell line exhibits this early stability fingerprint, the complete titer profile is used as a set of multivariate inputs to the modeling framework described herein. The titer characteristics used as regressors for the modeling method are reported in Table 1.
[0119] Table 1. Titer characteristics (i.e., production concentration profile variables) used as regressors for the modeling method
[0120]
[0121] Data analysis methods can be described as "supervised" or "unsupervised". "Supervised" analysis uses labeled input and output training data sets such that the model learns how to accurately predict the outcome. On the other hand, "unsupervised" analysis models unlabeled data sets in order to find naturally occurring patterns in the training data set.
[0122] If the group of input regressors is a categorical description, i.e., the regressors have the fingerprint of the regressand, the supervised learning model is effective. This can be verified by exploring the data in an unsupervised manner such that cluster-related effects can be identified in advance. By replacing the PLS-DA method with PCA (EMPCA), the EMPLS-DA modeling method can be transformed into an unsupervised method, where score diagnostics are desirable in order to visualize the clusters of all evolving models into a lower-dimensional multivariate space.
[0123] EMPCA modeling was performed on two sets of data (Group A and Group B) generated from different process platforms. Traditionally, these production stability trials are conducted in 4 production runs before making a final stability call on a single clonal cell line.
[0124] The production stability trials were conducted in 4 production runs for a total of 150 generations and 80 generations in Group A and Group B, respectively. The models established after the first two production runs did not show strong cluster separation, indicating a potential weak correlation between the titer profile features used as regressors in the first two production runs and the stability of the clonal cell line ( Figure 5 and Figure 6 ). However, after the third production run, the model was able to distinguish between stable and unstable clonal cell lines in both datasets. Therefore, the multivariate evolution method (EMPCA) begins to accumulate stability-related information in a systematic manner after production run 3 and thus shows the fingerprint for the final stability call of the clonal cell line earlier than the full set of endpoint titer data required by traditional methods. These considerations based on two different stability trial designs demonstrate the potential discriminatory power of the supervised method that may be developed when operating in these contexts. Different designs or process platforms require similar unsupervised analyses to understand at which production run the stability discrimination is represented as the fingerprint of the titer characteristics. Platform processes that result in different dynamic behaviors of the titer profile should be considered in separate models as the feature extraction is a description of the titer dynamics.
[0125] Example 2: Project test
[0126] The performance of the developed model was tested by simulating its adoption in scenarios where the production stability data of the final clonal cell line is available so that the performance of the model can be evaluated after each production run. These simulation scenarios consider the sequential execution of each different project such that more project data becomes gradually available to expand the calibration dataset. This approach allows testing whether differences in the size and composition of the calibration dataset affect the modeling predictions.
[0127] Four scenarios (Table 2) were considered using different calibration datasets, prediction datasets, process platforms, and numbers of generations at the end of the stability test. The models were evaluated under these scenarios by assessing the misclassification error rate, specificity, and sensitivity after each production run during calibration as well as prediction. When calibrating, the calibration results give an indication of the model robustness, while the prediction results are an indication of the actual prediction performance. According to the risk-based approach, model calibration results with a misclassification error below 30% and a sensitivity greater than 70% are considered satisfactory, where sensitivity, i.e., stable prediction accuracy, is superior to unstable prediction accuracy.
[0128] In addition to predicting the production stability of clonal cell lines, the developed method also estimates the class probabilities of the overall confidence factors for generating each clonal cell line. This confidence factor is calculated using the PLS-DA method described in detail herein. PLS-DA is capable of describing the difference between the probabilities where a given clonal cell line (P i ) is stable (P i,稳定 ) or unstable (P i,不稳定 ) as high or low, and thus the cases predicted with high or low confidence. Therefore, the model can provide additional information where low-confidence predictions are flagged as potentially deceptive.
[0129] Table 2. Simulation scenarios considered for model testing
[0130]
[0131] The final stability calls for the clonal cell lines in each item used to evaluate model calibration and prediction performance are shown in Table 3.
[0132] Table 3. Process platforms, stability test end criteria design, and final stability calls for the unique clonal cell lines considered in each item
[0133]
[0134] In Scenario 1, during the stability test, the model was calibrated on 46 clonal cell lines from a single dataset (A) using a process platform (X) that measures up to 150 generations of cells. After the third production run, the calibration performance was satisfactory, with an error rate of <7%, sensitivity of >93%, and specificity of >94%, but there was still a reasonable 20% marginal error for the first two runs (Table 4). The prediction performance had an error rate of 46 - 54% in the first two production runs, while the last two runs showed good prediction performance with a misclassification error of approximately 15%. The low sensitivity of the prediction performance could be partially explained by a small number (4) of stable clonal cell lines in the prediction dataset B. A closer inspection of the prediction confidence after Production Run 4 revealed that 3 / 3 of the stably misclassified clonal cell lines were labeled as low confidence, but the confidence level for all stably misclassified clonal cell lines was approximately 60% after Run 3.
[0135] Table 4. Model performance for model production stability classification in calibration (left) and prediction (right) in Scenario 1. Results are presented as misclassification error, specificity, and sensitivity after each production run.
[0136]
[0137]
[0138] In Scenario 2, the model was calibrated on 70 clonal cell lines from dataset A + B. The calibration performance was consistent across production runs, but the absolute values remained similar compared to Scenario 1 where only one dataset was used for calibration (Table 5). On the other hand, the prediction performance was significantly improved compared to Scenario 1 due to very high sensitivity (=100%) and specificity (>92%) observed after all consecutive production runs.
[0139] Table 5. Model performance for model production stability classification in calibration (left) and prediction (right) in Scenario 2. Results are presented as misclassification error, specificity, and sensitivity after each production run.
[0140]
[0141] In Scenario 3, during the stability test, the model was calibrated on 48 clonal cell lines from a single dataset (C) using a process platform (Y) that measures up to 80 generations. After the first production run, the calibration performance had only a misclassification error rate of 19%. After the third and fourth runs, the misclassification error rate decreased to 0% (Table 6). The prediction performance showed that the model had a sensitivity issue after the first two runs, with a misclassification error of approximately 30%. However, the prediction performance was satisfactory after the third run, with a misclassification error of 16% (7 out of 45 clonal cell lines), and 6 / 7 of the misclassified clonal cell lines were identified as low-confidence predictions.
[0142] Table 6. Model performance for model production stability classification in calibration (left) and prediction (right) for Scenario 3. Results are presented as misclassification error, specificity, and sensitivity after each production run.
[0143]
[0144] Finally, in Scenario 4, in the stability test, the model was calibrated on 93 clonal cell lines from two datasets (C+D) using a process platform (Y) that measured up to 80 generations. Compared to Scenario 3, the calibration performance decreased slightly, but the prediction performance increased with Run 2 being an outlier, which can be attributed to the generation number shift (Table 7).
[0145] Table 7. Model performance for model production stability classification in calibration (left) and prediction (right) for Scenario 4. Results are presented as misclassification error, specificity, and sensitivity after each production run.
[0146]
[0147] Thus, the modeling approach uses the fingerprint of clonal cell line production stability in productivity titer data to predict production stability earlier than would be possible using traditional full stability test methods. The models described herein provide a tool for selecting the most promising stable clonal cell lines at the early stages of the cell line development process, thus significantly reducing the time and resources required to conduct stability tests.
[0148] Example 3: Model test for actual combat project
[0149] After the initial evaluation detailed in Example 3, it was decided to test the modeling approach on in - the - wild projects as well as projects using a new platform. Tables 8, 9, and 10 show the prediction performance of the models for each project (numbered 1 - 5) and platform (A or B) tested. The sensitivity was always high at Run #3 (80 to 100%), and was consistent with previous results. The specificity was generally low, and sometimes extremely low in the more stable Platform B, due to the low number of available unstable clonal cell lines (i.e., 1 unstable clonal cell line in Project 3, 2 unstable clonal cell lines in Project 5). The overall error rate at Run #3 was low (usually about 10% or less). These are all measures of satisfactory model performance in predicting clonal cell line production stability.
[0150] Project 2 behaved slightly differently from any of the other projects analyzed because it showed an unusual confidence level in terms of the actual final production stability calculation during Run #3. This led to a very low specificity for Run #3, as all unstable clonal cell lines were incorrectly predicted with low confidence while stable clonal cell lines were still predicted well. Therefore, a complete model diagnosis was performed on Project 2, which confirmed that it was the only project in which the majority of clonal cell lines showed different and irrelevant model structures with respect to the titer profile characteristics and multiple variables important for prediction (VIP). The VIP score for each variable was calculated according to the equation described in Andersen & Bro 2010. This difference was found to be statistically significant for past projects in which the model was regularly tested and implemented in terms of confidence level. The analysis was performed by comparing the model residuals (exploratory PCA and PLS) and their contributions with respect to the VIP of the calibration model during Run #3. Thus, Project 2 presented as an outlier identified by the diagnosis, believed to be generated by clonal cell lines with atypical expression profiles and very low productivity. After Run #3 of Project 2, four out of five stable clonal cell lines were correctly predicted. Therefore, the model still gave valuable results for the selection of stable clonal cell lines.
[0151] The high sensitivity (98 - 100%) of the Platform B project may be due to the higher fraction of stable clonal cell lines generated by this specific platform and the effective prediction of the model for stable clonal cell lines.
[0152] Table 8. Error rates of the models for each run of Projects 1 - 5.
[0153]
[0154] Table 9. Sensitivity rates of the models for each run of Projects 1 - 5.
[0155]
[0156]
[0157] Table 10. Specificity rates of the models for each run of Projects 1 - 5
[0158]
[0159] Example 4: VIP
[0160] To determine the VIP in the model for predicting production stability, the VIP of the model constructed using the two combined training datasets (Datasets C and D in Scenario 4 of Table 2 above) was analyzed. Figure 7Shows the VIP score plot for each variable tested in the PLS-DA model after Run #3. Variables with VIP scores greater than 1 are defined as VIPs having a greater impact on predicting the production stability of the clone cell line. This analysis revealed three groups of VIPs after each production run; those with low VIP indices (close to 1), those with medium VIP indices, and those with high VIP indices after each of Production Runs 2 - 4 (Table 11). Table 11 shows that the maximum of the titer profile for Production Run 1 (max R1), the maximum-minimum gradient for Production Run 1 (grad R1), and the difference 6 for Production Run 1 (diff6 R1) were identified as variables with high VIP scores after each of Production Runs 2, 3, and 4. The standard deviation of the titer profile for Production Run 1 (std R1) was also identified as a variable with high VIP scores after Production Runs 3 and 4. Thus, the statistical features of the titer profile that affect predicting the production stability of the clone cell line can include the maximum, the maximum-minimum gradient, the difference, and / or the standard deviation.
[0161] Table 11. Variables important for prediction (VIP)
[0162]
[0163] References
[0164] Andersen, C.M., & Bro, R. Variable selection in regression – a tutorial. Journal of Chemometrics 24(2010):728 - 737.
[0165] Brereton, Richard G., and Gavin R. Lloyd. "Partial least squares discriminant analysis: taking the magic away." Journal of Chemometrics 28.4(2014):213 - 225.
[0166] Butler, M. & Spearman, M. The choice of mammalian cell host and possibilities for glycosylation engineering. Curr Opin Biotechnol 30, 107 - 112, doi:10.1016 / j.copbio.2014.06.010(2014).
[0167] Dahodwala, H., & Lee K. H. The fickle CHO: a review of the causes, implications, and potential alleviation of the CHO cell line instability problem. Curr Opin Biotechnol 60, 128 - 137 https: / / doi.org / 10.1016 / j.copbio.2019.01.011 (2019).
[0168] García S., MacGregor, J. F., Kourti, T., 2005. Product transfer between sites using Joint - Y PLS. Chemom. Intell. Lab. Syst. 79, 101–114. doi:10.1016 / j.chemolab.2005.04.009
[0169] Geladi, P., & Kowalski, B. R. (1986). Partial least - squares regression: a tutorial. Analytica chimica acta, 185, 1 - 17.
[0170] Gower, J. C., 1975. Generalized procrustes analysis. Psychometrika 40, 33–51. doi:10.1007 / BF02291478
[0171] Gunther, J. C., Baclaski, J., Seborg, D. E., Conner, J. S., 2009. Pattern matching in batch bioprocesses—Comparisons across multiple products and operating conditions. Comput. Chem. Eng. 33, 88–96. doi:10.1016 / j.compchemeng.2008.07.001
[0172] Haykin,S.,2008.Neural Networks and Learning Machines,3edizione.ed.Prentice Hall,New York.
[0173] Jackson,J.E.,2003.A User’s Guide to Principal Components.Wiley-Interscience,Hoboken,N.J.
[0174] Meneghetti,N.,Facco,P.,Bezzo,F.,Himawan,C.,Zomer,S.,Barolo,M.,2016.Knowledge management in secondary pharmaceuticalmanufacturing by miningof data historians—A proof-of-concept study.Int.J.Pharm.505,394–408.https: / / doi.org / 10.1016 / j.ijpharm.2016.03.035
[0175] Nomikos,Paul,and John F.MacGregor."Monitoring batchprocesses usingmultiway principal component analysis."AIChE Journal40.8(1994):1361-1375.
[0176] Nomikos,P.,MacGregor,J.F.,1995.Multi-way partial least squaresinmonitoring batch processes.Chemom.Intell.Lab.Syst.,InCINC’94Selected papersfrom the First International Chemometrics InternetConference 30,97–108.https: / / doi.org / 10.1016 / 0169-7439(95)000437
[0177] Ramaker, Henk-Jan, et al. "Fault detection properties of global, local and time evolving models for batch process monitoring." Journal of Process control 15.7 (2005): 799 - 805.
[0178] Walczak, B., and D. L. Massart. "Dealing with missing data: Part I." Chemometrics and Intelligent Laboratory Systems 58.1 (2001): 15 - 27.
[0179] Walsh, G. Biopharmaceutical benchmarks 2018. Nat Biotechnol 36, 1136 - 1145, doi:10.1038 / nbt.4305 (2018).
[0180] Wold, S., Sjostrom, M., 1977. SIMCA: A Method for Analyzing Chemical Data in Terms of Similarity and Analogy, in: Chemometrics: Theory and Application, ACS Symposium Series. American Chemical Society, pp. 243–282. https: / / doi.org / 10.1021 / bk - 1977 - 0052.ch012
[0181] Wold, S., Martens, H., & Wold, H. (1983). The multivariate calibration problem in chemistry solved by the PLS method. In Matrix pencils (pp. 286 - 293). Springer, Berlin, Heidelberg.
[0182] Wold, S. (1978). Cross-validatory estimation of the number of components in factor and principal components models. Technometrics 20, 397-405.
[0183] Wurm, Florian M., and Maria Wurm. "Cloning of CHO cells, productivity and genetic stability—a discussion." Processes 5.2 (2017): 20
Claims
1. A method for selecting a clonal cell line for the production of a therapeutic protein, the method comprising: For a plurality of clonal cell lines, measuring the product concentration of each clonal cell line; Based on the product concentration, determining product concentration profile data for each clonal cell line; Inputting the product concentration profile data into a learning model comprising a modeling framework; wherein the modeling framework comprises multivariate latent variable modeling, multi-way analysis structure, and an evolutionary model structure; Using the learning model to generate an output indicative of the production stability of each clonal cell line; And Selecting a clonal cell line for the production of a therapeutic protein based on the output.
2. A method for the production of a therapeutic protein, the method comprising: For a plurality of clonal cell lines, measuring the product concentration of each clonal cell line; Based on the product concentration, determining product concentration profile data for each clonal cell line; Inputting the product concentration profile data into a learning model comprising a modeling framework; wherein the modeling framework comprises multivariate latent variable modeling, multi-way analysis structure, and an evolutionary model structure; Using the learning model to generate an output indicative of the production stability of each clonal cell line; And Selecting a clonal cell line for the production of a therapeutic protein based on the output.
3. The method according to claim 1 or 2, wherein the multivariate latent variable modeling comprises at least one of the following: projection to latent structure (PLS) and principal component analysis (PCA).
4. The method according to claim 3, wherein the PLS comprises at least one of the following: PLS discriminant analysis (PLS-DA), PLS support vector machine (PLS-SVM), PLS neural network (PLS-NN), PLS logistic regression (PLS-LR), PLS-k-nearest neighbor (PLS-KNN), PLS decision tree (PLS-DT), PLS naive Bayes (PLS-NB), PLS random forest (PLS-RF), and PLS gradient boosting (PLS-GB).
5. The method according to any one of the preceding claims, wherein the product concentration profile data of the clonal cell line comprises one or more of the following: average product concentration, standard deviation, skewness, kurtosis, difference, maximum product concentration, and maximum-minimum gradient.
6. The method according to any one of the preceding claims, wherein for a plurality of clonal cell lines, measuring the product concentration of each clonal cell line comprises measuring the product concentration of each clonal cell line across multiple production runs.
7. The method according to any one of the preceding claims, wherein measuring the product concentration of each clonal cell line comprises measuring the product concentration of each clonal cell line up to 150 generations of cell growth.
8. The method according to claim 6 or 7, further comprising obtaining the generation number distribution of each production run.
9. The method according to any one of the preceding claims, wherein the evolutionary model structure comprises analyzing the product concentration profile data of consecutive production runs.
10. The method according to any one of the preceding claims, wherein the clonal cell line is a mammalian cell line.
11. The method according to claim 10, wherein the mammalian cell line is a CHO cell line.
12. The method according to any one of the preceding claims, wherein the product concentration is measured at multiple bioreactor scales.
13. The method according to any one of the preceding claims, further comprising using the selected clonal cell line to produce the therapeutic protein.
14. The method according to any one of the preceding claims, wherein the product concentration comprises the concentration of a monoclonal antibody.
15. A system for determining the production stability of a clonal cell line or selecting a clonal cell line, the system comprising: a. an input for receiving clonal cell line product concentration profile data; b. a learning model for determining the production stability of a clonal cell line or selecting a clonal cell line, the learning model comprising a modeling framework, the modeling framework comprising multivariate latent variable modeling, multi-way analysis structure, and an evolutionary model structure; c. one or more processors for processing the clonal cell line product concentration profile data using the learning model; and d. an output for providing an indication of the production stability of a clonal cell line based on the processing of the clonal cell line product concentration profile data by the learning model.
16. A computer program comprising instructions that, when executed by a computer or a data processor, cause the computer to perform the operations of the method according to any one of claims 1-14.
17. A computer-readable medium comprising instructions that, when executed by a computer or a data processor, cause the computer to perform the operations of the method according to any one of claims 1-14.
18. The program according to claim 16 or the medium according to claim 17, wherein the instructions, when executed, implement (a) a learning model and / or any associated functions and / or (b) a system for predicting the production stability of a clonal cell line.
19. The program or medium according to claim 16, 17 or 18, wherein the instructions, when executed, output an output having an indication of the production stability of a clonal cell line.
20. A computer-readable data carrier having stored thereon the computer program according to claim 16.