Pathogen Clearance Systems and Methods
A computer-implemented pathogen clearance model using historical data predicts the effectiveness of pathogen clearance processes, addressing contamination risks in recombinant protein production by reducing the need for extensive experimental validation and optimizing resource use.
Patent Information
- Application Number
- JP2023517970
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-09-21
- Filing Date
- 2021-08-23
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2041-08-23
AI Technical Summary
The production of recombinant therapeutic proteins from animal or human cell lines carries the risk of contamination with pathogens, and current pathogen clearance methods require significant resources and efforts in wet laboratory experiments for validation.
A computer-implemented method using a pathogen clearance model constructed from historical experimental data to predict the effectiveness and efficiency of pathogen clearance processes based on process parameters, reducing the need for extensive experimental validation.
Enables efficient prediction of pathogen clearance processes, allowing for prioritization and selection of effective processes without the need for numerous experimental trials, thereby optimizing resource utilization.
Smart Images

Figure 0007756155000005 
Figure 0007756155000006 
Figure 0007756155000007
Abstract
Description
[Background technology]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of U.S. Provisional Patent Application No. 63 / 080,856, filed September 21, 2020, the disclosure of which is incorporated herein by reference in its entirety.
[0002] background Today, the majority of recombinant therapeutic proteins are produced by large-scale fermentation of animal- or human-derived host cells that have been genetically engineered to express the gene of interest. Typical host cells include baby hamster kidney (BHK-21), Chinese hamster ovary (CHO), mouse myeloma (NS0), and several potential human cell lines, namely HEK293.
[0003] The production of recombinant therapeutic proteins from animal or human cell lines carries the risk of contamination with pathogens.
[0004] Host cells may contain endogenous retrovirus-like particle (ERVLP) coding sequences (provirus-like) inherently integrated into their chromosomes. ERVLPs are spontaneously produced in cell culture, resulting in viral contamination.
[0005] Adventitious pathogens can be introduced into a bioprocess or final product via the cell substrate, raw materials, and mechanical, environmental, personnel, and process-related factors.
[0006] Therefore, effective pathogen clearance, such as virus removal and / or inactivation, through the manufacturing process is critical to ensure the pathogen safety of biologics.
[0007] Pathogen clearance is typically achieved by dedicated unit operations such as low pH inactivation, viral filtration, chromatographic separation and / or other techniques.
[0008] The capability and robustness of pathogen clearance by a manufacturing process must be verified by virus clearance studies using scaled-down models. Validation of pathogen clearance studies should be performed in accordance with GLP (Good Laboratory Practices) guidance. Studies are performed using scaled-down test systems that represent manufacturing process conditions with test articles, parts, and process intermediates artificially spiked with model viruses to evaluate virus clearance performance. Scaled-down test systems are developed and qualified to represent the Current Good Manufacturing Practices (cGMP) processes used in the manufacturing facility. Typically, log 10 Virus reduction value (log 10 Expressed as a viral reduction value (LRV), the viral clearance obtained from these small-scale studies represents the viral clearance potential of the corresponding process steps in a cGMP manufacturing facility.
[0009] Therefore, process development of each pathogen clearance step for new therapeutic protein production requires significant effort and resources invested in wet laboratory experiments for process characterization studies. Summary of the Invention [Means for solving the problem]
[0010] overview To reduce such efforts and resources, the present embodiments provide a tool for predicting the performance of a (new) pathogen clearance process based on past experiments.
[0011] In a first aspect, the present invention provides a computer-implemented method comprising: receive a large number of training datasets, - wherein each training dataset of said multiple training datasets comprises values of at least two process parameters and at least one pathogen clearance score value, wherein the values of said process parameters characterize a pathogen clearance process and the value of said pathogen clearance score represents the effectiveness and / or efficiency of said pathogen clearance process; constructing a pathogen clearance model based on said multiple training data sets; - wherein the pathogen clearance model is configured to determine a pathogen clearance process, and a value of a pathogen clearance score from values of process parameters characterizes the pathogen clearance process; receiving an evaluation dataset; - wherein said evaluation dataset comprises values of at least two process parameters, said values of said at least two process parameters characterizing the pathogen clearance process being evaluated; inputting the evaluation dataset into the pathogen clearance model; receiving a resulting value of a pathogen clearance score as output from the pathogen clearance model; and Output the result value and / or one or more results associated with it The computer-implemented method includes:
[0012] In a second aspect, the present embodiment provides a computer system, comprising: Receiving unit a processing unit, and Output Unit Including, wherein the receiving unit is configured to receive an evaluation dataset, the evaluation dataset comprising values of at least two process parameters, the values of the at least two process parameters characterizing a pathogen clearance process being evaluated; wherein the processing unit is further configured to input the evaluation dataset into a pathogen clearance model and to receive a resulting pathogen clearance score value as output from the pathogen clearance model, the resulting pathogen clearance score value representing an effectiveness and / or efficiency of the pathogen clearance process being evaluated; wherein the pathogen clearance model is configured to predict a relationship between process parameters of a pathogen clearance process and a pathogen clearance score based on a number of training data sets, wherein the process parameters characterize the pathogen clearance process and the pathogen clearance score represents an effectiveness and / or efficiency of the pathogen clearance process; wherein the output unit is configured to output the resulting value of the pathogen clearance score and / or one or more results associated therewith. The computer system is provided.
[0013] In a third aspect, the present embodiments provide a non-transitory computer-readable storage medium comprising processor-executable instructions for performing operations for determining a result value of a pathogen clearance score for a pathogen clearance process based on an evaluation dataset, said operations comprising: receiving the evaluation dataset, wherein the evaluation dataset comprises at least two values of process parameters characterizing the pathogen clearance process; determining a value of the pathogen clearance score from the evaluation dataset by using a pathogen clearance model, wherein the pathogen clearance model is configured to predict a relationship between a process parameter of the pathogen clearance process and a pathogen clearance score based on a number of training datasets; and Output the result value and / or one or more results associated with it This includes: [Brief explanation of the drawings]
[0014] [Figure 1] FIG. 1 is a schematic diagram of an exemplary purification process. [Figure 2] Figure 2 is an example of a score plot obtained from a PCA analysis that provides an overview of the experiments performed in a reduced dimensional space. Each data point in the score plot represents a single experiment. The score plot also helps identify groupings, as indicated by the highlighted regions (Group-1 and Group-2). [Figure 3] FIG. 3 is an example of a contribution plot that allows identification of process parameters responsible for grouping experimental runs. [Figure 4] FIG. 4 is an example of a loading plot depicting the relationship between score vectors and process parameters in a reduced dimensional space. [Figure 5A] Figure 5 shows example confusion matrices that can be obtained for the accuracy of five different machine learning algorithms: (A) OPLS-DA, (B) LR, (C) SVM, (D) DT, and (E) RF machine learning algorithms. [Figure 5B] Figure 5 shows example confusion matrices that can be obtained for the accuracy of five different machine learning algorithms: (A) OPLS-DA, (B) LR, (C) SVM, (D) DT, and (E) RF machine learning algorithms. [Figure 5C] Figure 5 shows example confusion matrices that can be obtained for the accuracy of five different machine learning algorithms: (A) OPLS-DA, (B) LR, (C) SVM, (D) DT, and (E) RF machine learning algorithms. [Figure 5D] Figure 5 shows example confusion matrices that can be obtained for the accuracy of five different machine learning algorithms: (A) OPLS-DA, (B) LR, (C) SVM, (D) DT, and (E) RF machine learning algorithms. [Figure 5E] Figure 5 shows example confusion matrices that can be obtained for the accuracy of five different machine learning algorithms: (A) OPLS-DA, (B) LR, (C) SVM, (D) DT, and (E) RF machine learning algorithms. [Figure 6]FIG. 6 is an example of a loading plot for an exemplary OPLS-DA classification model. [Figure 7] FIG. 7 is an example of a decision tree from an exemplary random forest classification model. [Figure 8] FIG. 8 is a schematic diagram of a method for determining process parameters for which specific criteria for pathogen clearance scores are met. [Figure 9] FIG. 9 shows a graphical representation of the temperature and pH ranges for achieving the target LRV under the parameter bounds given in Table 2. [Figure 10] Figure 10 shows a comparison of measured vs. model-predicted overall LRV for the protein chromatography OPLS regression model. Each data point in the plot represents a single experiment. [Figure 11] Figure 11 shows a comparison of measured vs. model-predicted protein recovery for the Protein-A chromatography OPLS regression model. Each data point in the plot represents a single experiment. [Figure 12] FIG. 12 is a block diagram of an exemplary computer system of the present disclosure suitable for determining one or more pathogen clearance scores for a pathogen clearance process based on at least two process parameters. DETAILED DESCRIPTION OF THE INVENTION
[0015] drawing The drawings described herein are only for purposes of illustrating selected embodiments, not all possible implementations, and are not intended to limit the scope of the present disclosure.
[0016] Detailed Description The embodiments are described in more detail below without distinguishing between the subject matter of the embodiments (method, computer system, storage medium). Rather, the following descriptions are intended to apply equally to all subject matter of the embodiments, regardless of the context in which they occur.
[0017] This embodiment is useful for determining a pathogen clearance score for a pathogen clearance process.
[0018] A pathogen, in its broadest sense, is anything that can cause disease and / or harm an organism. In some embodiments of the present invention, a pathogen is also referred to as an infectious microorganism or pathogen, such as a virus, bacterium, protozoan, prion, viroid, or fungus. In some embodiments, the pathogen is a virus.
[0019] In some embodiments of the present invention, the term "viral inactivation" is used synonymously with the term "pathogen clearance."
[0020] A pathogen clearance process is any activity aimed at removing pathogens from a material, reducing the amount of pathogens in a material, and / or inactivating pathogens so that they no longer cause harm or are less likely to cause harm. In the case of inactivation, pathogens may remain in the material but in a non-infectious or less infective form.
[0021] Typical methods for removing or reducing the number of pathogens in materials include nanofiltration and chromatography.For more information about nanofiltration and / or chromatography, please refer to various related literature (for example, the chapter "Virus Removal by Nanofiltration" in "Therapeutic Proteins" edited by CMSmales and DCJames, Springer Protocols 2005, pp. 221-231; and the chapter "Downstream Industrial Biotechnology" edited by MCFlickinger, John Wiley & Sons, 2013, especially chapters 9.7-9.9).
[0022] Affinity chromatography, a commonly used pathogen clearance process for the production of therapeutic monoclonal antibodies, relies on the specific and reversible binding of antibodies to an immobilized ligand. Crude feedstock is passed through a column under conditions that promote binding of proteins (e.g., antibodies) in the feedstock to the ligand. After loading is complete, the column is washed under conditions that do not interfere with the specific interaction between the target protein and the ligand but disrupt nonspecific interactions between process impurities (e.g., host cell proteins) and the stationary phase. The bound protein is then eluted under mobile phase conditions that disrupt the target / ligand interaction. The most widely used affinity system for antibody purification is staphylococcal protein A (protein-A) and smaller ligands derived from it. The affinity between protein-A and IgG was one of the first natural interactions explored for the development of affinity systems for protein purification. In addition to Protein-A, other immunoglobulin-binding bacterial proteins such as Protein-G, Protein-A / G, and Protein-L are all commonly used to purify, immobilize, or detect immunoglobulins. For details on affinity chromatography, please refer to various textbooks and articles on affinity chromatography (e.g., "Affinity Chromatography - Methods and Protocols," edited by P. Bailon et al., Methods in Molecular Biology, Vol. 147, Humana Press Inc., 2000).
[0023] Many viruses contain lipid or protein coats that can be inactivated, for example, by chemical alteration or denaturation. Typical methods for such pathogen inactivation include solvent / detergent inactivation, pasteurization (heating), acidic pH inactivation (also known as low pH inactivation), and ultraviolet inactivation. For more information on pathogen inactivation, please refer to various related literature (e.g., "Filtration and Purification in the Biopharmaceutical Industry," edited by M.W. Jornitz, 3rd ed., CRC Press, 2020; "Continuous Biomanufacturing," edited by G. Subramanian, Wiley-VCH, 2017).
[0024] Different removal and / or inactivation steps are often combined. A typical downstream purification process for recombinant human monoclonal antibodies (rhumAb) is shown in Figure 1. First, harvested cell culture fluid (HCCF) is loaded onto a Protein-A column. The captured rhumAb is then eluted with a low pH solution after extensive equilibration and washing with a high-salt wash buffer. The eluate is then adjusted to a low pH in the range of 3.7–3.9 and held for a duration of 2 hours or more to inactivate enveloped viruses. The low-pH viral inactivation eluate is then neutralized and further polished by an anion or cation exchange column or membrane adsorption chromatography step to remove impurities and viruses. The product intermediate is then further filtered through a virus filter to remove potential viruses.
[0025] In many cases, the concentration of viruses in a given sample is extremely low. While low-level impurities can be ignored in other extraction processes, viruses are infectious impurities, so even a single virus particle can be sufficient to discard the entire process chain. Analytical limitations usually make it impossible to demonstrate absolute virus absence. Therefore, virus verification tests are performed both to demonstrate the clearance of viruses known to be associated with the product and to estimate the robustness of the process to clear potential adventitious viral contaminants (which may have gained access to the product) by characterizing the process's ability to clear nonspecific "model" viruses.
[0026] A "spiking study" is a study conducted to determine possible methods of virus removal or inactivation. For each process step being evaluated for its virus inactivation / removal ability, material is removed from the previous manufacturing process step, a known amount of virus is added (spiked), and the sample is processed through a scaled-down version of the manufacturing process step. The amount of infectious virus before and after the scaled-down process step is measured, for example, by infecting indicator cells in an endpoint dilution setup. The virus reduction ability of a process step can be calculated and presented, for example, as a log reduction value (LRV):
[0027] LRV=log 10 [(V1×T1) / (V2×T2)]
[0028] where V1 = volume of spiked feedstock before the clearance step; T1 = virus concentration of spiked feedstock before the clearance step; V2 = volume of material after the clearance step; and T2 = virus concentration of material after the clearance step.
[0029] The LRV is an example of a pathogen clearance score for a pathogen clearance process. A pathogen clearance score is typically a number that represents the effectiveness and / or efficiency of a pathogen clearance process.
[0030] According to this embodiment, the one or more pathogen clearance scores are determined using a pathogen clearance model that is constructed from a large amount of historical experimental data, also referred to herein as training data or a training data set.
[0031] The past experimental data is a number of data sets, each of which includes values of at least two process parameters and at least one value of a pathogen clearance score. The values of the at least two process parameters characterize a pathogen clearance process. The value of the pathogen clearance score represents the effectiveness and / or efficiency of the respective pathogen clearance process.
[0032] Process parameters typically relate to the conditions under which the pathogen clearance process is / was carried out.
[0033] For example, if the pathogen clearance process is a low pH inactivation process, process parameters characterizing the process include, for example, temperature, pH value, incubation time, initial virus titer, sample volume, virus load, spike ratio, protein type (e.g., antibody class and / or subclass), and / or pH-lowering agent. Additional and / or other process parameters are contemplated.
[0034] For example, if the pathogen clearance process is a Protein-A chromatography process (or any other affinity chromatography process), experimental features characterizing said process may include, for example: - Chromatography process settings: Chromatography column bed volume, load capacity, load density - For equilibrium steps: step volume, flow rate, conductivity, pH - For load steps: step volume, flow rate, conductivity, spike dilution, protein concentration, pH, capacity - For the first cleaning stage: step volume, flow rate, conductivity, pH - For the second wash step: step volume, flow rate, conductivity, pH - Elution steps: step volume, flow rate, eluate pH, conductivity, protein concentration - For regeneration steps: step volume, flow rate, conductivity, pH Includes:
[0035] Further and / or other experimental features are contemplated. The respective pathogen clearance score for a pathogen clearance process may be, for example, the LRV achieved in the pathogen clearance process and / or any other value related to the performance (effectiveness and / or efficiency) of the pathogen clearance process.
[0036] The pathogen clearance score can be, for example, the time required to achieve a predetermined concentration level of the pathogen in the material (inactivation time). Particularly for low-pH inactivation processes in which the virus is exposed to a low pH value during the incubation time, it can be interesting to know how long the incubation time should last to completely inactivate the virus. Therefore, the inactivation time is the appropriate pathogen clearance score for determining the performance of the pathogen clearance process.
[0037] In another preferred embodiment, the pathogen clearance score is the LRV achieved after a certain time (e.g., after 30 minutes, 60 minutes, or 90 minutes, or 120 minutes, or any other period of time) or any other value related to the (residual) amount of pathogen in the material. Such a pathogen clearance score is particularly suitable for evaluating the performance of a pathogen clearance process in which the material is subjected to a particular treatment for a predetermined period of time, such as low pH inactivation, heat treatment, UV irradiation, etc.
[0038] Classes can also be defined, with each class representing a pathogen clearance score value. Staying with the example of inactivation time, various pathogen clearance processes can be classified according to the time it takes for complete virus inactivation. For example, there can be two classes: a first class encompassing pathogen clearance processes (characterized by their respective process parameters) in which complete inactivation is achieved within a predetermined time limit, and a second class encompassing pathogen clearance processes in which the pathogens are not completely inactivated within the predetermined time limit. The pathogen clearance score identifies which class a particular pathogen clearance process belongs to. Of course, the number of classes is not limited to two. It is possible to have three or more classes, for example, three, four, five, or more classes, with each class representing a particular group of pathogen clearance processes with equivalent (similar) performance.
[0039] In particular, in the case of affinity chromatography purification processes, protein recovery (e.g., antibody recovery) is another example of a pathogen clearance score.
[0040] From multiple (historical, experimental) data sets, pathogen clearance models are constructed. Such pathogen clearance models correlate process parameters of the pathogen clearance process with respective pathogen clearance scores. There are many types of models and methods (algorithms) for creating these models, including, but not limited to, random forests, support vector machines, logistic regression, tree-based algorithms, naive Bayes, linear / logistic regression, artificial neural networks, nearest neighbor methods, Gaussian process regression, and / or various forms of recommender system algorithms (for details, see, for example, Kevin P. Murphy, "Machine learning: a probabilistic perspective," MIT Press, 2012). Scores generated by various methods can be combined using techniques such as, but not limited to, bagging, boosting, Brembling, ensembling, Bayesian model combination (BMC), simple averaging, and weighted averaging (see, e.g., Giovanni Seni and John Elder, "Ensemble Methods in Data Mining: Improving Accuracy Through Combining Predictions," 2010 (Morgan and Claypool Publishers); Opitz & Maclin, "Popular ensemble methods: An empirical study," 1999, Journal of Artificial Intelligence Research 11:169-98; and Rokach, "Ensemble-based classifiers," 2010, Artificial Intelligence Review 33(1-2):1-39).
[0041] Once a pathogen clearance model has been generated based on a large number of training data sets, the pathogen clearance model can be used to predict pathogen clearance for new pathogen clearance processes.
[0042] A new pathogen clearance process is typically the pathogen clearance process that is the subject of an evaluation. The purpose of the evaluation is to determine the performance of the new pathogen clearance process (the new pathogen clearance process being evaluated).
[0043] The pathogen clearance process being evaluated is characterized by an evaluation dataset, the evaluation dataset including values of at least two process parameters, the values of the at least two process parameters characterizing the pathogen clearance process being evaluated.
[0044] The evaluation data set is input into a pathogen clearance model, which outputs a resulting pathogen clearance score, which is a numerical value that represents the effectiveness and / or efficiency of the pathogen clearance process being evaluated.
[0045] The result value of the pathogen clearance score of the evaluated pathogen clearance process can be output, for example, to a monitor and / or a printer. Instead of or in addition to outputting the result value, one or more results related to the result value can be output. The result value can, for example, be compared to a predetermined reference value (e.g., a target pathogen clearance score). If the result value has a predetermined deviation from the reference value, a message can be output that the pathogen clearance process is deemed suitable for pathogen clearance or that the pathogen clearance process is deemed not suitable for pathogen clearance.
[0046] A number of new pathogen clearance processes can be evaluated by determining the respective result values of the pathogen clearance score based on the pathogen clearance model. The pathogen clearance processes of the number of new pathogen clearance processes can be ranked according to their result values. From the ranking list, portions of the pathogen clearance process (top performers) can be selected for further evaluation (e.g., experimental validation of predicted results).
[0047] Therefore, the embodiments described herein enable prioritization. It is not necessary to conduct numerous experiments to determine whether a pathogen clearance process can remove and / or inactivate pathogens in the process according to predetermined performance criteria. By calculating one or more pathogen clearance scores and comparing the one or more pathogen clearance scores with one or more predetermined reference values, it is possible to determine whether a new pathogen clearance process meets the performance criteria. Pathogen clearance processes whose respective pathogen clearance scores meet the predetermined criteria can be selected. The selected pathogen clearance processes can be further evaluated experimentally, and the non-selected pathogen clearance processes can be ignored.
[0048] Some examples of specific pathogen clearance models and pathogen clearance scores are given below, without any intention of limiting the present embodiments to said examples.
[0049] In one embodiment of the present invention, the pathogen clearance model is constructed by Principal Component Analysis.
[0050] Principal component analysis (PCA) is an unsupervised machine learning method used to reduce the dimensionality of collinear datasets. As shown in Table 1, a model with four principal components captures most (92%) of the variance present in the dataset.
[0051] Table 1: Overview of Protein-A Chromatography PCA Model
[0052] [Table 1]
[0053] PCA provides an effective and efficient way to contextualize experimental data. As shown in Figure 2, score plots can be used to identify groupings between experiments as well as atypical results. Additional patterns can emerge by color-coding observations based on available metadata (e.g., product name). Furthermore, PCA facilitates (a) the identification of process parameters, which facilitate grouping of experimental results in score plots, and (b) the identification of relationships between process parameters. For example, differences between experimental results for low-pH virus inactivation can be easily identified in Figure 2 and related to the original process parameters via the contribution plot in Figure 3. Additionally, the loading plot in Figure 4 can reveal relationships between underlying process parameters, helping scientists confirm known relationships or identify new ones. For example, inactivation time is observed to be positively correlated with pH.
[0054] In a preferred embodiment, the pathogen clearance model is trained using a supervised training method.
[0055] Supervised machine learning can be leveraged to model the relationship between process parameters and pathogen clearance performance. As an example, a classification modeling approach was used to build a pathogen clearance prediction model for low-pH virus inactivation. Two categories were defined based on inactivation time: 1. Fast inactivation (complete inactivation within a certain period of time) 2. Slow inactivation (incomplete inactivation within a certain period of time).
[0056] Inactivation time is determined as the first time point at which the viral titer falls below the assay limit of detection. This study evaluates the predictive potential and interpretability of multiple machine learning algorithms. Specifically, the following algorithms were evaluated: Orthogonal Partial Least Squares-Discriminant Analysis (OPLS-DA) Logistic Regression (LR) Support Vector Machine (SVM) · Decision tree (DT) Random Forest (RF)
[0057] The predictive ability of each classification model was evaluated using measures of overall model accuracy and individual class accuracy calculated using cross-validation (CV). CV overall accuracy and class accuracy were calculated as the average of their respective values over the total number of CV groups.
[0058] Table 2: Model performance summary for OPLS-DA, LR, SVM, DT, and RF
[0059] [Table 2]
[0060] Based on the CV overall accuracy results shown in Table 2, SVM and RF have the lowest (0.74) and highest (0.94) predictive abilities, respectively. OPLS-DA, with a CV overall accuracy of 0.89, outperforms LR, SVM, and DT, which have CV accuracies of 0.86, 0.74, and 0.83, respectively. The same conclusion can be drawn regarding the accuracy of the class "Slow." RF was the only machine learning algorithm evaluated in this study that performed better than OPLS-DA. Although OPLS-DA had a lower CV overall accuracy compared to RF, it still performed equally well to predict the class "Slow" with a CV class accuracy value of 0.92.
[0061] An overall accuracy of 1 clearly indicates the tendency of DT models to overfit, a condition in which the model memorizes specific patterns in the training dataset and fails to generalize them well to new data, as reflected in lower overall accuracy scores and low test accuracy scores in cross-validation. Overfitting was addressed by using an ensemble of DTs in a random forest algorithm, which yielded superior predictive performance.
[0062] Confusion matrices were generated for all five machine learning algorithms based on the entire training dataset, as shown in Figure 5. The confusion matrices were used to evaluate the performance of the classification models through the number of correctly predicted and inaccurately predicted observations per class. For example, there was only one incorrectly classified observation per class for the RF algorithm. However, for DT, there were no misclassified results for the "Slow" and "Fast" classes, which were previously discussed as indicating model overfitting. Therefore, the use of cross-validation in conjunction with confusion matrices can detect instances of model overfitting, allowing for effective evaluation of model performance.
[0063] Each machine learning method offers a different way to interpret modeling results. For linear methods, the model can be interpreted in terms of the direction and magnitude of the correlation between input and output. This is best demonstrated using the loading plot of the OPLS-DA model shown in Figure 6. For example, a positive loading for pH means that the higher the pH value, the slower the inactivation.
[0064] Nonlinear models may not be interpreted in the same way as linear models. Tree-based classification models maximize their accuracy by finding a split in the predictor variables that minimizes the Gini index. As shown in Figure 7, one of the decision trees from a random forest classification model identifies initial virus titer and pH as the two most important variables in achieving high model accuracy.
[0065] Once a pathogen clearance model is constructed, it can also be used to determine the process range within which a predetermined requirement for the pathogen clearance score is met. In the first step, a target pathogen clearance score, e.g., a threshold that should not be exceeded, is defined. Then, a combination of process parameters can be determined that results in a pathogen clearance score that does not exceed the predetermined threshold. A schematic diagram of this method is shown in Figure 8. To determine the process parameters that satisfy a given pathogen clearance score, an inverse problem must be solved that maps the pathogen clearance score to the process parameters. This inverse problem is formulated as a constrained optimization problem, the numerical solution of which results in a combination of process parameters that satisfies the target pathogen clearance score. For such an optimization problem, one or more of the process parameters can be fixed or subject to specific constraints. This is illustrated by the following example of a low-pH viral inactivation process. The configuration of this example is outlined in Table 3.
[0066] Table 3: Operating ranges of time series model parameters
[0067] [Table 3]
[0068] The target LRV was defined to be ≧5.00, and limits were defined for the process parameters pH and temperature. All other process parameters and conditions remained fixed.
[0069] The results are shown in Figure 9. Each plotted point in the figure corresponds to a valid solution that satisfies the target LRV and the specified constraints. The region where a process parameter combination meets the target LRV is indicated by data points colored by the achieved LRV. If the target LRV cannot be achieved for a particular combination, no point is plotted. For example, at a pH of 3.65, an LRV of at least 5 can be achieved for all temperatures. This is not true for higher pH values; for a pH of 3.8, a temperature above 19.5°C must be used to achieve the target LRV. This application can enable evaluation of how changing requirements for process parameters translate into achievable LRV.
[0070] Figures 10 and 11 show the predictive accuracy of the Protein-A chromatography OPLS (orthogonal partial least squares) regression model. This model was used to predict LRV and protein recovery throughout the Protein-A chromatography process. The process parameters are summarized in Table 4. Some of the process parameters were transformed nonlinearly to account for their nonlinear relationship to LRV and protein recovery.
[0071] Table 4: Input variables of the Protein-A chromatography model and their transformations
[0072] [Table 4]
[0073] Another example of a pathogen clearance model is an artificial neural network trained to determine one or more pathogen clearance scores from the values of the process parameters.
[0074] Such an artificial neural network comprises at least three layers of processing elements: a first layer with input neurons (nodes), an Nth layer with at least one output neuron (node), and N-2 inner layers, where N is a natural number greater than 2. In such a network, the output neurons function to predict at least one value of at least one pathogen clearance score. The input neurons function to receive values of process parameters. The processing elements of the layers are interconnected in a predetermined pattern with predetermined connection weights between them. Each network node represents a simple calculation of a weighted sum of inputs from previous nodes and a nonlinear output function. The combined calculations of the network nodes relate inputs to outputs. A separate network can be developed for each property measurement, or a group of properties can be included in a single network. Training estimates network weights that enable the network to calculate an output value(s) close to the measured output value. Supervised training methods can be used, in which output data is used to direct the training of the network weights. The network weights are initialized with small random values or with the weights of a previous, partially trained network. Training data inputs are applied to the network, and output values are calculated for each training sample. The network output values are compared to the measured output values. A backpropagation algorithm is applied to adjust the weight values in a direction that reduces the error between the measured and calculated outputs. This process is repeated until no further error reduction occurs or a predetermined prediction accuracy is reached. A cross-validation method can be used to split the data into training and validation data sets. The training data set is used for backpropagation training of the network weights. The validation data set is used to verify that the trained network generalizes and makes good predictions. The best set of network weights can be considered to best predict the output of the test data set.Similarly, varying the number of network hidden nodes and determining which network performs best on the dataset optimizes the number of hidden nodes.
[0075] Forward prediction uses a trained network to calculate one or more pathogen clearance scores for a (new) process based on its process parameters. Values of the process parameters are input into the trained network. A feedforward calculation through the network is performed to predict output property values. The predicted measurements can be compared to one or more property target values or tolerances. Because embodiment methods are based on historical data of property values, predictions of property values using such methods typically have errors approaching those of the empirical data; therefore, predictions are often as accurate as validation experiments.
[0076] Details on setting up artificial neural networks and training the networks can be found, for example, in CC Aggarwal: "Neural Networks and Deep Learning", Springer 2018, ISBN 978-3-319-94462-3.
[0077] The present embodiment is implemented using a computer system. Figure 12 illustrates an exemplary computer system 200. In connection therewith, computer system 200 may be configured to implement, via executable instructions, various algorithms and other operations described herein.
[0078] Exemplary computer system 200 may include, for example, one or more servers, workstations, personal computers, laptops, tablets, smartphones, other suitable computing devices, combinations thereof, etc. Additionally, computer system 200 may include a single computing device, or it may include multiple computing devices located in close proximity or distributed across a geographic region and coupled to each other via one or more networks. Such networks may include, without limitation, the Internet, an intranet, a private or public local area network (LAN), a wide area network (WAN), a mobile network, a telecommunications network, combinations thereof, or other suitable networks, etc.
[0079] In this case, the illustrated computer system 200 includes a processing unit 202 and a memory 204 coupled to (and in communication with) the processing unit 202. The processing unit 202 may include one or more processors (e.g., multi-core configurations, etc.), including without limitation a central processing unit (CPU), a microcontroller, a reduced instruction set computer (RISC) processor, an application specific integrated circuit (ASIC), a programmable logic device (PLD), a gate array, and / or any other circuit or processor capable of the functionality described herein. The above list is exemplary only, and thus is not intended to limit in any way the definition and / or meaning of the term processing unit.
[0080] Memory 204 is one or more devices that allow for the storage and retrieval of information, such as executable instructions and / or other data, as described herein. Memory 204 may include one or more computer-readable storage media, such as, without limitation, dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), solid-state devices, flash drives, CD-ROMs, thumb drives, tapes, hard disks, and / or any other type of volatile or non-volatile physical or tangible computer-readable medium. Memory 204 may be configured to store, without limitation, process parameters, pathogen clearance scores, pathogen clearance models, and / or other types of data (and / or data structures) suitable for use as described herein. In various embodiments, computer-executable instructions are stored in memory 204 for execution by processing unit 202 to cause processing unit 202 to perform one or more of the functions described herein, such that memory 204 is a physical, tangible, and non-transitory computer-readable storage medium. It should be appreciated that memory 204 may include a variety of different memories, each implemented in one or more of the functions or processes described herein.
[0081] In the exemplary embodiment, computer system 200 also includes an output unit 206 coupled to (and in communication with) processing unit 202. Output unit 206 outputs or presents information to a user of computer system 200, such as, for example, without limitation, by displaying and / or otherwise outputting information such as pathogen clearance scores, process parameters, and / or any other type of data. It should be further appreciated that in some embodiments, output unit 206 may comprise a display device such that various interfaces (e.g., applications (network-based or otherwise) etc.) may be displayed in computer system 200, particularly on the display device, to display such information, data, etc. Also, in some examples, computer system 200 may cause interfaces to be displayed on a display device of another computing device, including, for example, a server hosting a website having multiple web pages or a server interacting with web applications used on other computing devices. Output unit 206 may include, without limitation, a liquid crystal display (LCD), a light-emitting diode (LED) display, an organic LED (OLED) display, an "electronic ink" display, combinations thereof, etc. In some embodiments, the output unit 206 may include multiple units.
[0082] Computer system 200 further includes input device(s) 208 that receive input from a user. Input device(s) 208 are coupled to (and in communication with) processing unit 202 and may include, for example, a keyboard, a pointing device, a mouse, a stylus, a touch-sensitive panel (e.g., a touchpad or touchscreen), another computing device, and / or an audio input device. Additionally, in some exemplary embodiments, a touchscreen, such as one included in a tablet or similar device, can function as both output unit 206 and input device 208. Note that, in at least one exemplary embodiment, output unit 206 and input device 208 may be omitted.
[0083] Additionally, the illustrated computer system 200 includes a network interface 210 coupled to (and in communication with) the processing unit 202 (and, in some embodiments, also to the memory 204). The network interface 210 may include, without limitation, a wired network adapter, a wireless network adapter, a telecommunications adapter, or other device capable of communicating to one or more different networks.
[0084] The network interface (210) and / or the input device (208) are also referred to herein as receiving units.
Claims
1. 1. A computer-implemented method comprising: (a) receiving a plurality of training datasets, wherein each training dataset of the plurality of training datasets comprises values of at least two process parameters and at least one pathogen clearance score value, the process parameter values characterizing a pathogen clearance process and the pathogen clearance score value representing an effectiveness and efficiency of the pathogen clearance process; (b) constructing a pathogen clearance model based on the plurality of training datasets, wherein the pathogen clearance model is configured to determine, for a pathogen clearance process, a pathogen clearance score value from values of process parameters characterizing the pathogen clearance process; (c) receiving an evaluation dataset, wherein said evaluation dataset comprises values of at least two process parameters, said values of said at least two process parameters characterizing the pathogen clearance process being evaluated; (d) inputting the evaluation data set into the pathogen clearance model; and (e) receiving a pathogen clearance score result value as an output from the pathogen clearance model and outputting the result value and one or more outcomes associated therewith; The method comprising:
2. 10. The computer-implemented method of claim 1, wherein the pathogen clearance process is a low pH viral inactivation process.
3. 3. The computer-implemented method of claim 2, wherein the pathogen clearance score is a virus reduction capacity in the form of a log reduction value after a predetermined period of time and inactivation time required to achieve a predetermined concentration level of virus in a material.
4. 4. The computer-implemented method of claim 3, wherein the process parameters are two or more parameters selected from the group consisting of temperature, pH value, incubation time, initial virus titer, sample volume, virus amount, spike ratio, protein type, and degrading agent.
5. 10. The computer-implemented method of claim 1, wherein the pathogen clearance process is an affinity chromatography process used to purify antibodies.
6. 6. The computer-implemented method of claim 5, wherein the pathogen clearance score is a virus reduction capacity in the form of a log reduction value and the amount of antibody recovered.
7. 7. The computer-implemented method of claim 6, wherein the process parameters are two or more parameters selected from each of the group consisting of: for chromatography process settings: bed volume, load capacity, load density of the chromatography column; for equilibration phase: step volume, flow rate, conductivity, pH; for loading phase: step volume, flow rate, conductivity, spike dilution, protein concentration, pH, volume; for the first wash phase: step volume, flow rate, conductivity, pH; for the second wash phase: step volume, flow rate, conductivity, pH; for elution phase: step volume, flow rate, eluate pH, conductivity, protein concentration; and for regeneration phase: step volume, flow rate, conductivity, and pH.
8. 2. The computer-implemented method of claim 1, wherein the pathogen clearance model is selected from the group consisting of an orthogonal partial least squares discriminant analysis model, a logistic regression model, a support vector machine, a random forest, gradient boosting, and an artificial neural network.
9. 10. The computer-implemented method of claim 1, further comprising receiving a target pathogen clearance score, receiving information regarding process parameter constraints, determining a combination of resultant values of process parameters that satisfies the process parameter constraints and achieves the target pathogen clearance score, and outputting the values of the resultant values of process parameters.
10. 2. The computer-implemented method of claim 1, further comprising: determining a plurality of outcome values for a plurality of pathogen clearance processes to be evaluated; ranking the pathogen clearance processes of the plurality of pathogen clearance processes according to their outcome values to create a ranking list; selecting a portion of top performers from the ranking list; and conducting experiments on the selected top performers for validation purposes.
11. 1. A computer system comprising: (a) a receiving unit; (b) a processing unit; and (c) an output unit, wherein the receiving unit is configured to receive an evaluation dataset, wherein the evaluation dataset comprises values of at least two process parameters, the values of the at least two process parameters characterizing the pathogen clearance process being evaluated; wherein the processing unit is further configured to input the evaluation dataset into a pathogen clearance model and to receive a resultant value of a pathogen clearance score as output from the pathogen clearance model, the resultant value of the pathogen clearance score representing the effectiveness and efficiency of the pathogen clearance process being evaluated; wherein the pathogen clearance model is configured to predict, based on a plurality of training data sets, a relationship between process parameters of a pathogen clearance process and a pathogen clearance score, wherein the process parameters characterize the pathogen clearance process and the pathogen clearance score represents the effectiveness and efficiency of the pathogen clearance process; and wherein the output unit is configured to output the result value of the pathogen clearance score and one or more results associated therewith. The computer system.
12. 1. A non-transitory computer-readable storage medium comprising processor-executable instructions for performing operations of determining a result value of a pathogen clearance score for a pathogen clearance process based on an evaluation dataset, the operations comprising: (a) receiving the evaluation dataset, wherein the evaluation dataset comprises at least two values of process parameters characterizing the pathogen clearance process; (b) determining the value of the pathogen clearance score from the evaluation dataset by using a pathogen clearance model, wherein the pathogen clearance model is configured to predict a relationship between a process parameter of a pathogen clearance process and a pathogen clearance score based on a plurality of training datasets; and (c) outputting the result value and one or more results associated therewith. The non-transitory computer-readable storage medium.
Citation Information
Patent Citations
Method for producing high-titer, high-purity virus stocks and method for using them.
JP2013523175A
Determining conditions for purification of proteins
WO2019165148A1