Advanced Data-Driven Modeling for Biopharmaceutical Refinery Processes
Patent Information
- Application Number
- JP2024545897
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-02-04
- Filing Date
- 2023-01-24
- Publication Date
- 2026-01-30
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 306,971, filed February 4, 2022, the disclosure of which is incorporated by reference in its entirety herein.
[0002] The present disclosure relates generally to evaluating the performance of chemical processes, and more particularly to using machine learning and data modeling techniques to evaluate the performance of instances of chemical processes having a series of successive phases. [Background technology]
[0003] Purification is a key process in biopharmaceutical manufacturing that allows the separation of a therapeutic protein in its active form from other impurities. A typical purification process can include several chromatography-based unit operations, each of which can include multiple phases.
[0004] During the operation of each chromatography step, continuous (time series data for each parameter for each batch) can be generated by in-line / online sensors installed on the chromatography skids on the production floor, and batch data (e.g., one data point for each parameter for each batch) can be generated by at-line / offline in-process samples, respectively. These biomanufacturing process data can be leveraged for the development of advanced data-driven models that can generate their insights to support the decisions and actions of process experts.
[0005] Traditionally, each of the in-line / on-line / at-line / off-line analytical control charts for monitoring biomanufacturing processes tend to be univariate (e.g., one parameter per chart). This results in the need to consider multiple charts simultaneously to find any correlations between parameters. This makes real-time early fault detection and retrospective root cause analysis time-consuming and tedious. In addition, attempting to find relationships between multiple attributes by simply considering the individual charts of each parameter can be exceptionally difficult and limited in capturing all of the underlying correlations. Multivariate Data Analysis (MVDA) is a methodology that includes advanced statistical techniques that can be used to effectively analyze large, complex, heterogeneous data sets all at the same time. The development and deployment of such MVDA models enables more effective and efficient near real-time process monitoring, early fault detection and diagnosis. MVDA models can be used to monitor multiple process variables using only a few multivariate metrics while leveraging useful process information found in the correlation structures between process variables. MVDA is therefore a powerful methodology that can be used to assist process engineers and scientists with root cause identification of process excursions and provide numerous insights into manufacturing operations that can be leveraged to enhance overall process understanding and control. Summary of the Invention
[0006] Embodiments of the present disclosure include the application of advanced data-driven modeling to affinity chromatography columns in commercial biologics manufacturing. This involves developing a multivariate model that uses process parameters and in-process control parameters in the purification steps and visualizes the correlations between them. Specifically, embodiments of the present disclosure (a) present the application of hierarchical data-driven modeling methodology for effective monitoring of purification unit operations and their corresponding phases, (b) highlight the utility of such data-driven models, and (c) can be used to provide an overview of the key steps involved in developing advanced data-driven models in biologics manufacturing. Although the models are developed for affinity chromatography columns, similar modeling approaches can be adopted to monitor other types of chromatography columns, such as ion exchange, hydrogen ion concentration, etc. The basic concept of the data-driven modeling approach discussed herein is to find correlations and patterns in the data generated during the biomanufacturing process.
[0007] An exemplary method for evaluating performance of an instance of a chemical process having a series of consecutive phases includes obtaining data associated with the instance of the chemical process and evaluating the performance of the instance of the chemical process using a plurality of performance thresholds based on the data associated with the instance of the chemical process, where the plurality of performance thresholds are obtained by training a hierarchical model based on one or more historical instances of the chemical process, the hierarchical model including a plurality of batch evolutionary models (BEMs) at a first level of the hierarchy, where each BEM model corresponds to one phase of the series of consecutive phases, a plurality of batch level models (BLMs) at a second level above the first level of the hierarchy, where each BLM model corresponds to one phase of the series of consecutive phases, and a third level overall performance model at a third level above the second level of the hierarchy, where the overall performance model corresponds to all of the series of consecutive phases.
[0008] In some embodiments, the chemical process is a purification process to separate the recombinant protein from other proteins in the cell culture medium using one or more chromatography columns.
[0009] In some embodiments, the sequence of phases includes equilibrating, packing, washing, and eluting one or more chromatography columns.
[0010] In some embodiments, the chemical process comprises a cell culture development process, a cell separation process, a viral inactivation process, a pharmaceutical manufacturing process, or any combination thereof.
[0011] In some embodiments, each BEM of the plurality of BEMs is trained to obtain one or more performance thresholds for evaluating in-line data associated with a phase of a chemical process.
[0012] In some embodiments, the one or more performance thresholds include Hotelling's T2 method and one or more model residuals.
[0013] In some embodiments, multiple BEMs are trained using in-line data associated with one or more historical instances of a chemical process.
[0014] In some embodiments, the in-line data includes time series data obtained from one or more sensors.
[0015] In some embodiments, the in-line data is interpolated at a defined frequency.
[0016] In some embodiments, each BEM model of the plurality of BEMs is a partial least squares (PLS) model.
[0017] In some embodiments, each BLM of the plurality of BLMs is trained to obtain one or more performance thresholds for evaluating in-line, at-line, and offline data associated with a phase of a chemical process.
[0018] In some embodiments, the one or more performance thresholds include Hotelling's T2 method and one or more model residuals.
[0019] In some embodiments, multiple BLMs are trained using in-line, at-line, and offline data associated with one or more historical instances of a chemical process.
[0020] In some embodiments, the at-line and offline data include protein solution (bulk) attributes, bulk melt process attributes, column packing attributes, column attributes, elution attributes, sample measurements, or any combination thereof.
[0021] In some embodiments, each BLM model of the plurality of BLMs is a principal component analysis (PCA) model.
[0022] In some embodiments, the overall performance model is trained based on the second level trained BLM model.
[0023] In some embodiments, the method further includes displaying one or more results of the evaluated performance of the instance of the chemical process on a display.
[0024] In some embodiments, the method further comprises updating variables of the chemical process based on the evaluated performance of the instance of the chemical process.
[0025] An exemplary system for evaluating performance of an instance of a chemical process having a series of successive phases comprises one or more processors; a memory; and one or more programs, the one or more programs stored in the memory and comprising instructions, when executed by the one or more processors, for obtaining data associated with the instance of the chemical process and evaluating performance of the instance of the chemical process using a plurality of performance thresholds based on the data associated with the instance of the chemical process, the plurality of performance thresholds being obtained by training a hierarchical model based on one or more historical instances of the chemical process, the hierarchical model including: a plurality of batch evolutionary models (BEM) at a first level of the hierarchy, each BEM model corresponding to one phase of the series of successive phases; a plurality of batch level models (BLM) at a second level above the first level of the hierarchy, each BLM model corresponding to one phase of the series of successive phases; and a third level overall performance model at a third level above the second level of the hierarchy, the overall performance model corresponding to all of the series of successive phases.
[0026] An exemplary non-transitory computer-readable storage medium stores one or more programs for evaluating performance of an instance of a chemical process having a series of successive phases, the one or more programs including instructions that, when executed by one or more processors of an electronic device, cause the electronic device to obtain data associated with the instance of the chemical process and evaluate performance of the instance of the chemical process using a plurality of performance thresholds based on the data associated with the instance of the chemical process, the plurality of performance thresholds being obtained by training a hierarchical model based on one or more historical instances of the chemical process, the hierarchical model including: a plurality of batch evolutionary models (BEM) at a first level of the hierarchy, each BEM model corresponding to one phase of the series of successive phases; a plurality of batch level models (BLM) at a second level above the first level of the hierarchy, each BLM model corresponding to one phase of the series of successive phases; and a third level overall performance model at a third level above the second level of the hierarchy, the overall performance model corresponding to all of the series of successive phases. [Brief description of the drawings]
[0027] [Figure 1] 1 is a schematic diagram of an exemplary hierarchical model, according to some embodiments. Information from all base-level models is conveyed to the top level through their respective matrices (Tj) to a consensus matrix R. R is used to further summarize data from all base-level models by generating a top-level score TTL and a load PTL. [Diagram 2] FIG. 1 illustrates an exemplary cross-validation protocol, according to some embodiments. The dataset is split into training and test subsets. For each cross-validation round (with a different split of the dataset), a model is developed. The cross-validation results from all models are averaged to estimate the final predictive power of the models. [Diagram 3]1A and 1B are diagrams illustrating exemplary data structures according to some embodiments. (A) For the batch evolution model, datasets of different colors represent data from different batches. Each dataset includes "m" X variables (process-related variables), "n" observations (time points), and one Y variable (column volume). (B) For the batch level model, the dataset columns used for the batch evolution model are transposed to generate the batch level model dataset. Each row represents a different batch. X1 time1 indicates the value of the first X1 variable at time 1. Similarly, Xm timen indicates the value of the Xm variable at time n. [Figure 4] 1 is an exemplary schematic diagram of a purification column monitoring model structure, according to some embodiments. The bottom level of the hierarchy is a batch evolutionary model (BEM), which is a PLS model with in-line data only. The next lower level in the hierarchy is a batch level model (BLM) for each of the phases, which is a PCA model with in-line and at-line / off-line data. Finally, the top level model is a comprehensive PCA model that includes the in-line and at-line / off-line data of all phases combined. [Diagram 5] 1 is an exemplary BEM score plot showing the scores of the first principal component t[1] for X-space as a function of column volume (considered a maturity variable for a purification process) according to some embodiments. The green dashed line indicates the mean of the data, and the red dashed lines correspond to ±3 standard deviations from the mean. [Figure 6] FIG. 1 illustrates an exemplary model training, according to some embodiments. A batch evolution model is shown for different phases of the purification process. Each plot shows the progression of a batch summarized by the MVDA score t[1] in X-space versus the column volume of material passed through the purification column. [Figure 7]FIG. 1 illustrates an exemplary model training, according to some embodiments. Batch-level models of four phases in a refinery process are shown. Each plot represents a single phase, and the boundaries of the ellipses indicate the 95% confidence level of the data used to build the model for each of the phases. Each circle within the ellipses refers to a single batch, showing all the information for the batch summarized for that particular phase by the MVDA score. [Figure 8] FIG. 1 shows an exemplary model training according to some embodiments. The figure shows a top-level model of an affinity purification column that considers four phases: equilibration, loading, washing, and elution of the process. [Figure 9] FIG. 1 illustrates examples of model metrics used for excursion detection and diagnosis of a single BLM, according to some embodiments. Specifically shown are: (A) Hotelling's T2; (B) model residuals, both used to identify excursions during process monitoring; and (C) a contribution chart showing the contribution of variables to a batch excursion compared to the average across all batches. [Figure 10] FIG. 1 illustrates an example of model benchmarking, according to some embodiments. Also shown are model-detected excursions using MVDA metrics, and identified contributing process parameters using contribution charts. Also described are: (A) process excursions detected for a batch through the MVDA metrics "Hotelling's T2" and "Residual" in the packing phase batch level model (BLM); (B) excursions confirmed by the MVDA score plot of the packing phase batch evolution model (BEM); (C) contribution chart used to identify process parameters associated with the excursion, where the column pump flow rate was found to have the highest contribution to this excursion; and (D) univariate representation of pump flow rate versus column volume, where the column packing pump flow rate was dramatically reduced for some time during the packing phase of the batch due to a period of pump stoppage. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0028] The following description is presented to enable those skilled in the art to make and use the various embodiments. Descriptions of specific devices, techniques, and applications are provided only as examples. Various modifications to the examples described herein will be apparent to those skilled in the art, and the general principles defined herein may be applied to other examples and applications without departing from the spirit and scope of the various embodiments. Thus, the various embodiments are not intended to be limited to the examples described and illustrated herein, but should be accorded the scope consistent with the claims.
[0029] 1. Materials and Methods 1.1.Refining Process The purification process is downstream of the cell culture and separation steps in the manufacturing process of any recombinant therapeutic protein. During purification, the selected recombinant protein is separated from the pool of countless proteins, DNA, metabolites, etc., synthesized by mammalian host cells in cell culture as well as other process and product-related impurities. Depending on the type of protein being purified, different chromatography columns are used during the purification of a particular protein. Ion exchange, hydrophobic interaction, and affinity chromatography are among the most widely used separation techniques implemented for protein purification. The purification process is usually separated into several phases such as equilibration, loading, washing, elution, and finally production and storage of the purification column. A multivariate model developed for the online monitoring of therapeutic protein purification processes using affinity chromatography columns is discussed herein. The column contains peptide ligands (with target protein binding domains) on stationary phase beads that capture target protein molecules during the "loading" phase and release the protein molecules during the "elution" phase. Non-target proteins with no affinity for the column ligands flow through the column as waste.
[0030] 1.1.1.Equilibration During equilibration, the purification column is equilibrated with respect to its internal pH and conductivity before loading the target protein. This is accomplished by running a buffer through the column at the appropriate conditions for the selected protein.
[0031] Filling The column is first loaded with the target protein solution. During this phase, the therapeutic protein molecules that have an affinity for the packed beads in the column bind to the beads while the impurities flow through the column to waste, as they have no affinity for the peptide ligand.
[0032] Cleaning A wash buffer is passed through the column to remove only loosely bound impurities, leaving tightly held target protein molecules bound to the stationary phase beads.
[0033] 1.1.4.Elution An elution buffer is passed through the column, which disrupts the bond between the target protein and the peptide ligand and facilitates removal of the target protein molecules from the column. The column eluate containing the target protein is collected for further processing.
[0034] 1.2. Data and Data Sources The data is the basis of the modeling effort described herein. There are two categories of data used in the development of the MVDA model for affinity chromatography columns.
[0035] Inline / Online Data The in-line measurements used in the model are of the following types: (a) total volume of effluent from the chromatography column, (b) conductivity, (c) ultraviolet absorbance (UV), (d) temperature, (e) pressure, and (f) flow rate. Data from the process measurements are stored in a database called the PI process historian (OSIsoft). All time series data obtained from process sensors such as conductivity sensors are stored in the PI archive and their corresponding batch context (e.g., batch ID, start and end timestamps of individual process phases, etc.) is stored in the PI Asset Framework (AF) database.
[0036] At-line / off-line data Both at-line and offline data are accessible through Discoverant (BIOVIA), a relational database. Structured Query Language (SQL) is used to retrieve data from underlying data systems such as Manufacturing Execution Systems (MES), Laboratory Information Management Systems (LIMS), and Systems Applications and Products (SAP). The types of at-line / offline data used for model development include protein solution (bulk) attributes, bulk melt process attributes, column packing attributes, column attributes, elution attributes, and sample measurements.
[0037] 1.3. Software The following software was used in this case study:
[0038] Modeling: Simca 14.1 (Sartorius Stedim Biotech) and Matlab 2015b (MathWorks) Data acquisition, preprocessing, visualization, and model automation: Matlab, Python 3.6 (Python Software Foundation).
[0039] 1.4. Inline / Online Data Preprocessing In-line / online data obtained from the PI process historian must be processed in a standardized manner to remove obvious anomalies (e.g., chromatogram baseline offsets) while capturing the data and aligning the data to a standard format to facilitate batch-to-batch comparisons. Data pre-processing for the purification process includes the following steps: (a) interpolation, (b) segmentation, and (c) alignment.
[0040] In-line data captured by various sensors during the progression of a batch are stored in the PI historian with non-uniform sampling frequencies for different process parameters. In-line data is interpolated for all parameters at defined frequencies to monitor how each of the batches progresses through the purification process and compare their performance. Metadata including start and end timestamps were utilized to segment the in-line data into corresponding phases by extracting the time series data recorded continuously between the start and end time points for each of several phases involving affinity column batches. The time series data for each column sensor was pre-processed to ensure that all batches were aligned with respect to the start time for each sub-phase of the affinity purification process.
[0041] 1.5. Multivariate Data Analysis Multivariate Data Analysis (MVDA) refers to statistical techniques and algorithms used to jointly analyze data from three or more variables. Specifically, these algorithms can be used to detect patterns and relationships within the data. Some applications of these methods are clustering (detecting groupings), classification (determining group / class membership), and regression (determining the relationship between inputs and continuous numerical outputs). Some of the widely used MVDA techniques are Principal Component Analysis (PCA) and Partial Least Squares Projection onto Latent Structures (PLS - hereafter referred to as Partial Least Squares).
[0042] 1.5.1.Principal component analysis Principal Component Analysis (PCA) is an MVDA method that can be used to obtain an overview of the underlying data without a priori information and labeling or mapping of a priori information to target or output values. PCA is able to find structures and patterns in the data by reducing the dimensionality of a dataset where collinear relationships exist. The working principle of PCA is to summarize the original data by defining new orthogonal latent variables called principal components. These principal components (PCs) contain linear combinations of the original variables in the dataset. They are chosen such that the variance explained by a fixed number of PCs is maximized. The values of the original data in the new latent variable space are called scores. Given a dataset described by an nxm matrix with n observations and m variables, let T denote an nxk matrix containing k principal component values, called scores. Each individual variable X with i=1,...,n for the principal components is ij The coefficients p with J=1,...,m and q=1,...,k determine the contribution of jq is called the loadings. The m x k matrix P is called the loadings matrix, and the relationship between T, X, and P is given in matrix notation by equation (2.1).
[0043] X=TP T +E (2.1) where E denotes the residual nxm matrix. The residuals contain the variance not explained by the principal components l to k. A detailed introduction to PCA is available in Basilevsky, A., “Statistical factor analysis and related methods: theory and applications”, John Wiley & Sons, 2009, which is incorporated herein by reference. Model quality is assessed using cross-validation (see Section 2.5.4), as well as external datasets, if available. For this purpose, R 2 and Q 2 The statistics are evaluated. R 2 The statistic describes the proportion of the sum of squares explained by the model, Q 2Statistics convey information about the predictive ability of the model. A detailed derivation of both is given in Eriksson, L., Byrne, T., Johansson, E., Trygg, J., and Vikstrom, E., “Multi- and megavariate data analysis: Basic Principles and Applications” (2013): 425, which is incorporated herein by reference.
[0044] Partial Least Squares Partial Least Squares (PLS) regression is a MVDA method that aims to determine the functional relationship between inputs and outputs. This method is further described in "The Collinearity Problem in Linear Regression: The Partial Least Squares (PLS) Approach to Generalized Inverses" by Wold S et al., published in SIAM J.Sci.Stat.Comput.5(3)1984:735-743 and "PLS Regression: A Basic Tool of Chemometrics" by Wold S et al., published in Intell.Lab.Syst.58(2)2001:109-130, both of which are incorporated herein by reference. Briefly, a similar approach to PCA is taken in that the regression is not performed on the original variables available in the data set, but on a smaller number of orthogonal variables called latent variables. These are linear combinations of the original variables. In contrast to PCA, where the latent variables are selected to maximize the variance, in PLS the latent variables are determined to maximize the covariance between the dependent and independent variables. To obtain the solution to the regression problem, the following operations are performed in both X-space and Y-space: In X-space, a linear transformation is defined as follows:
[0045] T=XW * (2.2) and X=TP T +E (2.3) Here, T represents an X - score n x k matrix, P represents an X - loading m x k matrix, W * represents an X - weight m x k matrix, and E represents an X - residual n x m matrix with k < m. In the Y - space, the transformation is obtained as follows.
[0046] Y = UC T + G (2.4) Here, U represents a Y - score n x k matrix, C represents a Y - weight q x k matrix, and G represents a Y - residual n x q matrix. The X - scores are selected to minimize the X - residual E and be good predictors of Y, and the Y - scores are selected to minimize the Y - residual G. Similar to PCA, R 2 and Q 2 can be calculated for the PLS model.
[0047] 1.5.3. Hierarchical Modeling Hierarchical modeling facilitates the combination of data from different models, either PCA, PLS, or both. This is typically done to summarize information from different parts of processes that are not exactly similar but are interconnected. An application of this is combining different phases in affinity - chromatography - based purification processes such as equilibration, loading, washing, and elution, all of which are executed sequentially to achieve specific goals for each phase and ultimately output a purified product.
[0048] Hierarchical MVDA models include multiple levels. A detailed description of hierarchical models can be found in Wold, S., Kettaneh, N., Friden, H. and Holmberg, A., “Modelling and diagnostics of batch processes and analogous kinetic experiments”, Chemometrics and Intelligent Laboratory Systems 44(1998):331-340, which is incorporated herein by reference. An example of a two-level hierarchical model structure with a base-level (BL) model and a top-level (TL) model is shown in FIG. 1 for a process with two phases 1 and 2 with data X1 and X2, respectively. The base-level model (a) is a number of numbers, based on either PCA or PLS, and (b) has a set of latent variables (i.e., score matrix T j ) to summarize the input data, and (c) the P j and so on, where j denotes a different BL model. Information from both base-level models (corresponding to datasets X1 and X2) is fed to the top-level model through respective score matrices T1 and T2 with dimensions nxk1 and nxk2. The number of observations is denoted by n, and the number of latent variables in the BL models in phases 1 and 2 are k1 and k2, respectively. The TL model input is defined by an nx(k1+k2) matrix R that contains the scores from the two BL models. Specifically, the score matrices T j are combined to form the consensus matrix R (equation (2.5)) that is used to calculate the scores and loadings of the TL model. In the PCA TL model, the score matrix T TP , load P TP The relationship between, and the R matrix is given by equation (2.6).
[0049] R = [T1, T2] (2.5) T TP =RP TP (2.6) In general, k TL<(k1 + k2), which indicates that the MVDA hierarchical modeling structure facilitates the compression of all different BL models. An important advantage of the hierarchical model is that each of the data blocks such as X1 and X2 with different dimensions maintains a significant contribution to the TL model. Even when T1 contains fewer latent variables compared to T2 (k1 < k2), the hierarchical modeling processes the score matrices from both BL models with similar weighting.
[0050] 1.5.4. Cross - Validation Cross - validation is a model - testing technique used to evaluate whether the underlying statistical relationships in the data are general enough to predict a dataset that was not used in model training. In the cross - validation technique, a given dataset is split into training and test subsets. The model is developed using the training dataset and then evaluated against the test subset. Several rounds of cross - validation are performed (using different splits), resulting in multiple parallel models (see Figure 2). The results from all the parallel models are averaged to estimate the final predictive power of the model. The main purpose of cross - validation is to reduce the possibility of overfitting, where the model fits the training dataset very well but is not general enough to reasonably well predict an independent dataset.
[0051] Results and Discussion To make the MVDA purification monitoring model discussed in this specification a tool usable for end - users, the following factors were considered: (a) implementation of a meaningful modeling approach so that the model can detect process excursions, (b) benchmarking of new batches against historical batches. For this purpose, the modeling work was carried out in two stages: (a) model development and (b) benchmarking.
[0052] 1.6. Model Development The development of the MVDA monitoring model for an affinity chromatography column can include three steps: model selection, model training, and model testing.
[0053] Model Selection Evaluation of the batch trajectory of every single phase of affinity chromatography, e.g., equilibration, loading, washing, and elution (hereafter referred to as phases), requires the development of models that take into account the changes in the in-line data as a function of batch progression. Such models are called batch evolution models (BEMs).
[0054] Each phase can be further evaluated after the completion of a refinery batch by considering at-line and offline data. Therefore, there is a need for an MVDA model that can incorporate in-line time series data in addition to at-line / off-line discrete process parameters and attributes. In this regard, a batch-level model (BLM) can be used.
[0055] Finally, a comprehensive evaluation of an affinity chromatography unit operation requires the ability to evaluate all phases together. Such an objective can be achieved through a hierarchical modeling structure. Details of each of the levels of the hierarchical model are described in subsequent sections.
[0056] 2.1.1.1 Batch Evolution Model The batch evolution model is the first level in this hierarchical model structure. The batch evolution model provides an idea of how the batch is progressing by considering in-line data of various process parameters. The batch progression (either in terms of time of processing or volume of material processed) is represented as a function of all available in-line process parameters, summarized by a small number of latent variables. The BEM is a PLS model with the process parameters as X variables and the batch progression maturity as the Y variable. In some embodiments, 11 in-line process parameters include the X variables and column volume is used as the variable Y. The BEM focuses on maximizing the covariance between all process parameters X and the batch maturity Y. The dataset used to generate the BEM includes time series data of multiple batches. Each of the columns in the dataset corresponds to a different variable used for model development. Each of the rows corresponds to a different time point in the measurement of that batch (FIG. 3A).
[0057] 2.1.1.2 Batch-level model The batch-level model is the second level in the hierarchical model structure. It takes into account in-line and at-line / off-line data to provide an idea of how a batch will perform as phases of the refinery process are completed, compared to historical batches. The BLM here is essentially a PCA model that focuses on explaining the variation present in different process variables. All in-line time series data is transposed so that each row of the BLM dataset represents a single batch (see Figure 3B).
[0058] 2.1.1.3 Top-level model The top-level model is the third and highest level of the hierarchical model structure. The TL model combines the different levels in the multivariate modeling structure and provides a comprehensive view of the performance of a single batch through all phases of the refinery process (see Figure 4). The lowest level in the hierarchy is the batch evolutionary model (BEM), which is a PLS model that has only in-line data for each phase. The next lowest level in the hierarchy is the batch level model (BLM), which is a PCA model that combines the in-line and at-line / off-line data for each phase. Finally, the top-level model is a PCA model that includes the in-line and at-line / off-line data for all phases combined.
[0059] Model training After defining the model structure, in this case a hierarchical structure with batch evolution and batch-level models at the base level and a comprehensive top-level model, the next step is to train the model. Model training here refers to the process of using historical data to define multivariate control limits that are "acceptable operating ranges". In some embodiments, historical data including 60 drug substance (DS) batches were used for model training. All of these batches were considered for model training because they represent acceptable operating ranges. Specifically, none of the batches were excluded for model training because the quality of the final product produced by these DS batches was acceptable for publication.
[0060] Training the model with historical data (acceptable batches) makes it possible to define multivariate control limits that are in fact acceptable operating ranges. At the BEM level, the original time series data is described by only a few latent variables, which can be visualized as a function of column volume. In Figure 5, the score plot of the first principal component of a single batch from the BEM is shown as a function of column volume. The multivariate limits of the BEM are ±3 standard deviations (indicated by the red dashed line) of the historical data mean (indicated by the green dashed line). Figure 6 shows the BEM representation of all the refinery phases and batches considered for model training. Most of the batches used for model training fall within the multivariate limits. To increase the variability in the training dataset (to reduce the possibility of overfitting), we included some batches that are outside the multivariate limits but do not affect the process and product downstream of the refinery process. The score plots of the BLM and the top-level model are shown in Figures 7 and 8, respectively.
[0061] Process monitoring relies on two multivariate metrics: Hotelling’s T 2 and model residuals are used to facilitate this. 2 represents the distance of an observation from the historical mean. Residuals refer to the parts of the dataset that cannot be explained by the model, usually noise in the data, or occurrences not previously seen by the model. Hotelling's T for a batch 2 The tolerance of the residuals is defined by the 95% critical level. 2 If the deviations fall within the tolerances for the mean and residuals, no action is taken. However, if a batch is outside these tolerances for one or both metrics, further investigation of contributing factors is brought about. The contribution plot provides a quantitative comparison of the potential contribution of different process parameters to a particular excursion. It shows the difference for a selected batch or group of batches against the average of all batches.
[0062] Figure 9 shows the two excursion detection metrics (Hotelling's T 2 and model residuals) and one diagnostic metric (variable contribution), however these were calculated for all BLMs (shown in Fig. 7) and for the top level (shown in Fig. 8).
[0063] 2.1.3. Model testing The MVDA model is tested based on the following objectives: First, a test is performed to ensure that the model developed using the training dataset is general enough to describe an independent dataset. For this purpose, cross-validation is performed (see Section 1.5.4). Seven rounds of cross-validation were used for model testing purposes.
[0064] Further tests are performed to demonstrate the model's ability to detect excursions and determine the underlying contributing parameters. Eleven additional batches for the affinity chromatography process were used twice to detect process excursions and for model benchmarking.
[0065] 2.2 Model benchmarking Model benchmarking refers to the evaluation of a new batch (a batch not used for model training) against expectations from history that represent the acceptable operating range of the process, allowing evaluation of potential excursions and investigation of identified contributing factors, if any.
[0066] Eleven refined batches (not included in the training dataset) were used for model benchmarking. This served as a test of the model's ability to detect excursions (as described in Section 3.1.3 of Model Testing). To evaluate the batches, we used a multivariate metric, namely Hotelling's T 2 and model residuals were used. An example of model testing / benchmarking is shown in Figure 10. For affinity chromatography columns, Hotelling's T 2A process excursion was detected in one of the batches with both the model residual values (both outside the tolerance level) and the model error values (both outside the tolerance level). The excursion was identified in the MVDA score space of the filling phase BEM. In the contribution plot shown in Figure 10(C), it was found that the pump flow rate had the highest contribution to this excursion. A deeper look into the univariate plots revealed that the pump had stopped for some time during the filling phase in Figure 10(D). In addition, the subject matter experts from Manufacturing Sciences confirmed that the pump had indeed stopped due to some technical issue during the filling phase. Thus, through this monitoring procedure, an excursion can be detected whether or not it affects the product quality.
[0067] 3. Conclusion During commercial manufacturing of biopharmaceuticals, a wealth of process and product data is generated. These large and complex data sets are typically generated from in-line / online sensors for various unit operations, as well as from benchtop analyzers on the production floor and in quality control laboratories. This disclosure describes how large amounts of manufacturing data for purification processes can be utilized to develop advanced data-driven models that can be leveraged to generate process expert insights to support organizational decisions. Specifically, a case study of preparative affinity chromatography used in the production of recombinant therapeutic proteins is presented.
[0068] Multivariate models were developed for affinity chromatography columns for the purpose of effective and efficient in-line / on-line process monitoring using available in-line, online, at-line, and offline data. A multivariate hierarchical modeling approach was employed to consider several purification phases, including affinity chromatography unit operations, and to facilitate their comprehensive evaluation. This implies that the hierarchical model can monitor the trajectory of process parameters per process phase, in addition to the joint evaluation of process parameters with in-process controls. Specifically, individual batch evolution and batch-level models were developed for each phase, allowing the evaluation of the progress of new batches in terms of predictions from history. The available historical data was utilized for training these models, and additional data was used for model testing and benchmarking. The developed models describe the historically acceptable operating conditions that are used to evaluate new batches. Benchmarking can be performed with little, if any, intervention of multivariate diagnostics and contribution analysis to highlight factors (original variables) potentially contributing to excursions. The models presented herein have been tested and illustrated as being capable of detecting excursions.
[0069] This case study demonstrates how the development of advanced hierarchical data-driven models enables effective purification process monitoring through comprehensive evaluation of all phases, including unit operations, as well as the ability to detect patterns and relationships within each phase and across different phases. Multivariate modeling also ensures efficient process monitoring, as many process parameters can be evaluated through only a few multivariate metrics while maintaining the ability to investigate individual univariate analyses in detail. Also, the modeling approach discussed herein is not limited to purification alone, but can be applied to multiple unit operations during a biomanufacturing process. Developing multivariate models for cell culture, viral inactivation, and final product manufacturing (fill and finish) processes can also provide additional process understanding and efficient methods of holistic process monitoring and early fault detection.
[0070] Overall, advanced multivariate data-driven modeling can assist the overall orchestrated effort for process understanding and control of biopharmaceutical manufacturing processes while simultaneously enhancing process monitoring for early fault detection and fault diagnosis of purification unit operations.
[0071] Although the present disclosure and examples have been fully described with reference to the accompanying drawings, it should be noted that various changes and modifications will be apparent to those skilled in the art, and such changes and modifications should be understood as being included within the scope of the present disclosure and examples, as defined by the claims.
[0072] The above description has been described with reference to specific embodiments for the purpose of explanation. However, the above exemplary description is not intended to be exhaustive or to limit the invention to the precise form disclosed. Many modifications and variations are possible in light of the above teachings. The embodiments have been chosen and described to best explain the principles of the present teachings and their practical application. This will enable those skilled in the art to best utilize the present technology and various embodiments with various modifications suited to the particular use contemplated.
Claims
1. 1. A method for evaluating the performance of an instance of a chemical process having a series of successive phases, comprising: obtaining data associated with the instance of the chemical process; evaluating the performance of the instance of the chemical process using a plurality of performance thresholds based on the data associated with the instance of the chemical process; Including, the plurality of performance thresholds are obtained by training a hierarchical model based on one or more historical instances of the chemical process; The hierarchical model is a plurality of batch evolutionary models (BEMs) at a first level of a hierarchy, each BEM model corresponding to one phase of the series of successive phases; a plurality of batch level models (BLMs) at a second level above the first level of the hierarchy, each BLM model corresponding to one phase of the series of successive phases; a third-level overall performance model at a third level above the second level of the hierarchy, the third-level overall performance model corresponding to all of the series of successive phases; A method comprising:
2. 10. The method of claim 1, wherein the chemical process is a purification process for separating the recombinant protein from other proteins in the cell culture medium using one or more chromatography columns.
3. 3. The method of claim 2, wherein the series of phases includes equilibrating, packing, washing, and eluting the one or more chromatography columns.
4. The chemical process comprises: Refining process, cell culture development process, cell separation process, virus inactivation process, pharmaceutical manufacturing processes, or Any combination of these The method of claim 1 , comprising:
5. 5. The method of claim 1, wherein each BEM of the plurality of BEMs is trained to obtain one or more performance thresholds for evaluating in-line data related to a phase of the chemical process.
6. The method of claim 5 , wherein the one or more performance thresholds include Hotelling's T2 and one or more model residuals.
7. The method of claim 1 , wherein the plurality of BEMs are trained using in-line data associated with the one or more historical instances of the chemical process.
8. The method of claim 7 , wherein the in-line data comprises time-series data obtained from one or more sensors.
9. The method of claim 7 , wherein the inline data is interpolated at a defined frequency.
10. The method of claim 1 , wherein each BEM model of the plurality of BEMs is a partial least squares (PLS) model.
11. 10. The method of claim 1, wherein each BLM of the plurality of BLMs is trained to obtain one or more performance thresholds for evaluating in-line, at-line, and offline data associated with a phase of the chemical process.
12. The method of claim 11 , wherein the one or more performance thresholds include Hotelling's T2 and one or more model residuals.
13. 10. The method of claim 1, wherein the plurality of BLMs are trained using in-line, at-line, and offline data associated with the one or more historical instances of the chemical process.
14. 14. The method of claim 13, wherein the at-line and offline data include protein solution (bulk) attributes, bulk melt process attributes, column packing attributes, column attributes, elution attributes, sample measurements, or any combination thereof.
15. The method of claim 1 , wherein each BLM model of the plurality of BLMs is a principal component analysis (PCA) model.
16. The method of claim 1 , wherein the overall performance model is trained based on the trained BLM model of the second level.
17. The method of claim 1 , further comprising displaying one or more results of the evaluated performance of the instance of the chemical process on a display.
18. The method of claim 1 , further comprising updating variables of the chemical process based on the evaluated performance of the instance of the chemical process.
19. 1. A system for evaluating the performance of an instance of a chemical process having a series of successive phases, comprising: one or more processors; Memory and and one or more programs stored in the memory and configured to be executed by the one or more processors, the one or more programs comprising: obtaining data relating to the instance of the chemical process; evaluating the performance of the instance of the chemical process using a plurality of performance thresholds based on the data associated with the instance of the chemical process; including instructions for the plurality of performance thresholds are obtained by training a hierarchical model based on one or more historical instances of the chemical process; The hierarchical model is a plurality of batch evolutionary models (BEMs) at a first level of a hierarchy, each BEM model corresponding to one phase of the series of successive phases; a plurality of batch level models (BLMs) at a second level above the first level of the hierarchy, each BLM model corresponding to one phase of the series of successive phases; a third-level overall performance model at a third level above the second level of the hierarchy, the third-level overall performance model corresponding to all of the series of successive phases; Including, the system.
20. 1. A non-transitory computer-readable storage medium storing one or more programs for evaluating performance of an instance of a chemical process having a series of successive phases, the one or more programs, when executed by one or more processors of an electronic device, causing the electronic device to: obtaining data relating to the instance of the chemical process; evaluating the performance of the instance of the chemical process using a plurality of performance thresholds based on the data associated with the instance of the chemical process; Contains instructions, the plurality of performance thresholds are obtained by training a hierarchical model based on one or more historical instances of the chemical process; The hierarchical model is a plurality of batch evolutionary models (BEMs) at a first level of a hierarchy, each BEM model corresponding to one phase of the series of successive phases; a plurality of batch level models (BLMs) at a second level above the first level of the hierarchy, each BLM model corresponding to one phase of the series of successive phases; a third-level overall performance model at a third level above the second level of the hierarchy, the third-level overall performance model corresponding to all of the series of successive phases; 1. A non-transitory computer-readable storage medium comprising: