Automatic data amplification, scaling and forwarding

JP2025527110A5Pending Publication Date: 2026-07-06AMGEN INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
AMGEN INC
Filing Date
2023-06-29
Publication Date
2026-07-06

AI Technical Summary

Technical Problem

Existing biopharmaceutical manufacturing processes face challenges in data scaling due to the 'low-N problem', where limited manufacturing history hinders effective real-time multivariate statistical process monitoring, particularly during scale-up studies, and similarity theory has limitations in deriving accurate scaling models.

Method used

A data scaling framework using machine learning methods to generate time-varying scaling relationships between processes, enabling data transfer and amplification, particularly in biopharmaceutical manufacturing, by employing a Bayesian framework for parameter estimation in stochastic state space models to address non-uniform scaling.

Benefits of technology

This approach reduces the time required to develop predictive models for larger-scale processes by reusing data from smaller-scale processes, ensuring accurate scaling and model calibration even with limited historical data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method for scaling data across different processes includes obtaining first time series data indicative of one or more input, state, and / or output parameters of a first process over time and second time series data indicative of one or more input, state, and / or output parameters of a second process over time. The method also includes generating a scaling model that specifies a time-varying scaling relationship of the input, state, and / or output parameters between the first and second processes, and using the scaling model to transfer source time series data associated with a source process to target time series data associated with a target process. The source time series data indicative of one or more input, state, and / or output parameters of the source process over time, and the target time series data indicative of the input, state, and / or output parameters of the target process over time. The method further includes storing the target time series data in a memory.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates generally to the application of machine learning methods to automate and streamline the process of transferring data between different processes, such as processes associated with different manufacturing sites, different products, and / or different scales. [Background technology]

[0002] Despite decades of research and advances in industrial process monitoring and control, existing monitoring methods are not particularly effective for use with batch processes, particularly in the biopharmaceutical industry. Unlike other batch processes, biopharmaceutical processes pose unique challenges from a process monitoring and control perspective, which can be referred to as the "low-N problem." The low-N problem describes situations where manufacturing history is limited, with "N" referring to the length of a drug product's manufacturing history or the number of previous campaigns. The low-N problem has its roots in the way modern biopharmaceutical manufacturing companies operate. In biopharmaceutical manufacturing, drugs with long manufacturing histories often have extensive repositories of past campaign data from which to build robust monitoring models. However, as newer drugs are discovered and introduced to the market, long manufacturing histories are often unavailable. In fact, it is common for manufacturing histories to contain only a few or no previous campaigns prior to the actual GMP (Good Manufacturing Practice) campaign. Real-time multivariate statistical process monitoring (RT-MSPM) for these manufacturing processes has traditionally required large, real-scale datasets to build representative models, limiting its usefulness for critical operations related to new product introduction (NPI).

[0003] The low-N problem manifests itself in, among other areas, scale-up studies. Scale-up studies typically involve attempting to replicate laboratory processes in incremental steps to develop the expected performance and set of best practices for the final industrial facility. The problem of scaling between variables is not new and has been widely studied (particularly in scale-up studies) using similarity theory. See Skoglund, 1967, Similitude: Theory and Applications, International Textbook Co. See also Coutinho et al., 2016, Engineering Structures, 119:81-94. Similarity theory is a branch of engineering concerned with establishing necessary and sufficient conditions for similarity between phenomena. See Coutinho et al., 2016, Engineering Structures, 119:81-94. A prototype model is said to have similarity to a real-world application if they share geometric, kinematic, and dynamic similarities. Similarity theory is the main theory behind many formulas in fluid mechanics and is closely related to dimensional analysis. Sonin, 2001, “The Physical Basis of Dimensional Analysis”2 nd ed., Massachusetts Institute of Technology. See also Yunus and Cimbala, 2006, "Fluid Mechanics: Fundamentals and Applications," International Edition, McGraw Hill Publication. Similarity theory is widely used in hydraulic engineering to design and test fluid flow conditions in practical experiments using prototype models.

[0004] For example, scale-up for microbial growth is based on maintaining a constant dissolved oxygen concentration in the liquid (broth), regardless of the size of the bioreactor. This is typically achieved by keeping the impeller end (tip) speed the same in both the pilot and commercial reactors. If the impeller speed is too fast, the impeller movement can lyse the bacteria. If the speed is too slow, the contents of the bioreactor will not be mixed well. Similarity theory can be used to calculate the impeller speed required in a commercial bioreactor, given the speed in the pilot bioreactor. Suppose:

number

number

[0005] [Non-Patent Document 1] Skoglund, 1967, Similarity: Theory and Applications, International Textbook Co. [Non-patent document 2] Coutinho et al.,2016,Engineering Structures,119:81-94 [Non-patent document 3] Sonin, 2001, “The Physical Basis of Dimensional analysis” 2nd ed., Massachusetts Institute of Technology. [Non-patent document 4] Yunus and Cimbala, 2006, “Fluid Mechanics: Fundamentals and Applications,” International Edition, McGraw Hill Publication [Non-Patent Document 5] Hubbard et al.,1988,Chemical Engineering Progress,84:55-61 Summary of the Invention [Problem to be solved by the invention]

[0006] Therefore, there is a need for a general framework for determining scaling between arbitrary variables while addressing some of the long-standing data scaling issues in biopharmaceutical manufacturing or other processes. [Means for solving the problem]

[0007] To address some of the limitations of current industry best practices, embodiments are described herein that relate to systems and methods that improve upon conventional techniques for data scaling, transfer, and / or amplification of biopharmaceutical or other processes. "Data scaling" generally refers to the process of discovering and / or applying a mathematical relationship between two datasets, which may be referred to as a "source" dataset and a "target" dataset. In data scaling, a linear model captures the scaling relationship between the source dataset and the target dataset using specific parameters (e.g., slope and intercept). Scaling models, and the process of developing such models, can provide specific insights and have a variety of use cases.

[0008] One such use case is “data transfer,” which generally refers to the process of transferring data from one process (“source” process) to another (“target” process). For example, the source and target processes may be biopharmaceutical processes associated with different sites, scales, and / or formulations. As a more specific example, extensive experimental data from a benchtop-scale (e.g., 2 liter) bioreactor may be scaled / transferred to a pilot-scale (e.g., 500 liter) or commercial-scale (e.g., 20,000 liter) bioreactor, the latter with very limited experimental data, to generate a predictive or inferential model (e.g., a regression model or a machine learning model such as a neural network) for a larger-scale target process. In some embodiments, the data transfer process is intentionally perturbed to ensure that the target dataset has certain desired characteristics (e.g., to control variability in the transferred data), which is generally referred to herein as “data amplification.” This may be done by manually modifying certain parameters of the data scaling model to achieve the desired characteristics.

[0009] The data scaling, transfer, and / or amplification process can effectively reuse or repurpose data available from a source process, thereby significantly reducing the time required to generate, calibrate, and / or maintain a model of a target process, particularly in situations such as pipeline drug development and / or production where there is little or no past manufacturing history. Many other use cases are also possible, some of which are described in more detail below.

[0010] In some embodiments, a method for scaling data across different processes includes obtaining first time series data indicative of one or more input, state, and / or output parameters of a first process over time and obtaining second time series data indicative of one or more input, state, and / or output parameters of a second process over time. The method also includes generating, by one or more processors, a scaling model that specifies a time-varying scaling relationship between the input, state, and / or output parameters of the first process and the input, state, and / or output parameters of the second process. The method also includes transferring, by the one or more processors and using the scaling model, source time series data associated with a source process to target time series data associated with a target process. The source time series data indicative of one or more input, state, and / or output parameters of the source process over time, and the target time series data indicative of one or more input, state, and / or output parameters of the target process over time. The method also includes storing, by the one or more processors, the target time series data in a memory.

[0011] In another embodiment, a system includes one or more processors and one or more computer-readable media storing instructions. When executed by the one or more processors, the instructions cause the one or more processors to obtain first time-series data indicative of one or more input, state, and / or output parameters of a first process over time and obtain second time-series data indicative of one or more input, state, and / or output parameters of a second process over time. The instructions also cause the one or more processors to generate a scaling model specifying a time-varying scaling relationship between the input, state, and / or output parameters of the first process and the input, state, and / or output parameters of the second process. The instructions also cause the one or more processors to transfer source time-series data associated with a source process to target time-series data associated with a target process using the scaling model. The source time-series data is indicative of one or more input, state, and / or output parameters of a source process over time, the target time-series data is indicative of one or more input, state, and / or output parameters of a target process over time, and the instructions cause the one or more processors to store the target time-series data in memory.

[0012] Those skilled in the art will understand that the figures described herein are included for illustrative purposes and are not intended to limit the present disclosure. The figures are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the present disclosure. It should be understood that in some instances, various aspects of the described implementations may be shown exaggerated or enlarged to facilitate understanding of the described implementations. In the figures, like reference numbers throughout the various views generally refer to functionally similar and / or structurally similar components. [Brief explanation of the drawings]

[0013] [Figure 1]FIG. 1 is a simplified block diagram of an example system capable of implementing one or more of the data scaling, forwarding and / or amplification techniques described herein. [Figure 2] 1 shows normalized oxygen flux profiles for an exemplary bioreactor process run at two different scales. [Figure 3] FIG. 1 is a flow diagram of an exemplary method for scaling data across different processes. [Figure 4A] Normalized oxygen flow profile, estimated scaling factors and uncertainties, and actual and estimated (normalized) target signals are shown for a use case where the source process is a 300 liter pilot-scale bioreactor process for biologics production and the target process is a 10,000 liter commercial-scale bioreactor process for biologics production. [Figure 4B] Normalized oxygen flow profile, estimated scaling factors and uncertainties, and actual and estimated (normalized) target signals are shown for a use case where the source process is a 300 liter pilot-scale bioreactor process for biologics production and the target process is a 10,000 liter commercial-scale bioreactor process for biologics production. [Figure 4C] Normalized oxygen flow profile, estimated scaling factors and uncertainties, and actual and estimated (normalized) target signals are shown for a use case where the source process is a 300 liter pilot-scale bioreactor process for biologics production and the target process is a 10,000 liter commercial-scale bioreactor process for biologics production. [Figure 4D]Normalized oxygen flow profile, estimated scaling factors and uncertainties, and actual and estimated (normalized) target signals are shown for a use case where the source process is a 300 liter pilot-scale bioreactor process for biologics production and the target process is a 10,000 liter commercial-scale bioreactor process for biologics production. [Figure 5A] Normalized viable cell density (VCD) profiles of biologics produced in a 2,000 liter commercial-scale bioreactor and a 2 liter small-scale bioreactor are shown, along with estimated scaling factors. [Figure 5B] Normalized viable cell density (VCD) profiles of biologics produced in a 2,000 liter commercial-scale bioreactor and a 2 liter small-scale bioreactor are shown, along with estimated scaling factors. [Figure 5C] Normalized viable cell density (VCD) profiles of biologics produced in a 2,000 liter commercial-scale bioreactor and a 2 liter small-scale bioreactor are shown, along with estimated scaling factors. [Figure 6A] Normalized oxygen flow profiles for producing six source products and one target product, a similarity measure for each source product relative to the target product, and a ranking of the source products relative to the target product are shown. [Figure 6B] Normalized oxygen flow profiles for producing six source products and one target product, a similarity measure for each source product relative to the target product, and a ranking of the source products relative to the target product are shown. [Figure 6C] Normalized oxygen flow profiles for producing six source products and one target product, a similarity measure for each source product relative to the target product, and a ranking of the source products relative to the target product are shown. [Figure 6D]Normalized oxygen flow profiles for producing six source products and one target product, a similarity measure for each source product relative to the target product, and a ranking of the source products relative to the target product are shown. [Figure 7A] Figure 1 shows the product sieving performance of hollow membrane fibers in a 50-liter perfusion bioreactor equipped with an alternating tangential flow (ATF) system, the corresponding (normalized) Raman spectral scans from the bioreactor and permeate, and the similarity measure between the Raman spectral scans as a function of time. [Figure 7B] Figure 1 shows the product sieving performance of hollow membrane fibers in a 50-liter perfusion bioreactor equipped with an alternating tangential flow (ATF) system, the corresponding (normalized) Raman spectral scans from the bioreactor and permeate, and the similarity measure between the Raman spectral scans as a function of time. [Figure 7C] Figure 1 shows the product sieving performance of hollow membrane fibers in a 50-liter perfusion bioreactor equipped with an alternating tangential flow (ATF) system, the corresponding (normalized) Raman spectral scans from the bioreactor and permeate, and the similarity measure between the Raman spectral scans as a function of time. [Figure 8] 1 shows exemplary scaling relationships between different products and scales. [Figure 9A] 1 shows the actual normalized oxygen flow rate profile for the source process at a 300 liter scale, and the actual vs. predicted normalized oxygen flow rate profile at a 15,000 liter scale. [Figure 9B] 1 shows the actual normalized oxygen flow rate profile for the source process at a 300 liter scale, and the actual vs. predicted normalized oxygen flow rate profile at a 15,000 liter scale. DETAILED DESCRIPTION OF THE INVENTION

[0014] The various concepts introduced above and discussed in more detail below can be implemented in any of many ways, and the concepts described are not limited to any particular implementation. Several example implementations are provided for illustrative purposes.

[0015] Exemplary System FIG. 1 is a simplified block diagram of an exemplary system 100 that can be used to amplify, scale, and transfer data from a first process ("Process A") to a second process ("Process B"). As used herein, the terms "scale" or "scaling" can be used to refer to the act of transferring or projecting data from one process to another (e.g., from Process A to Process B), or to refer to the relative physical size of equipment and / or materials associated with a process. To clarify which application is intended, the former meaning is primarily referred to herein in connection with "data," "parameters," or "variables" (e.g., "scaled data / parameters / variables" or "scaling of data / parameters / variables"), while the latter is primarily referred to herein with respect to a process involving one or more objects (e.g., "scaling up" a bioreactor process).

[0016] FIG. 1 illustrates an exemplary embodiment in which Process A and Process B are bioreactor processes (for producing / growing a biopharmaceutical product) that use different sized bioreactors and therefore have different content volumes. Each bioreactor discussed herein may be any suitable vessel, device, or system that supports a biologically active environment that may include living organisms in a medium and / or materials derived therefrom (e.g., cell culture). The bioreactor may contain recombinant proteins expressed by the cell culture, for example, for research purposes, clinical use, commercial sale or other distribution, etc. Depending on the biopharmaceutical process, the medium may include a specific fluid (e.g., "broth") and specific nutrients, and may have a target pH level or range, a target temperature or temperature range, etc. Collectively, the contents and parameters / characteristics of the medium are referred to herein as the "medium profile."

[0017] In an "upscaling" scenario or embodiment, process A uses a small-scale bioreactor and process B uses a large-scale bioreactor. For example, process A may use a 2 liter benchtop scale bioreactor and process B may use a 500 liter pilot scale bioreactor, or process A may use a 500 liter pilot scale bioreactor and process B may use a 20,000 liter commercial scale bioreactor, etc. "Downscaling" scenarios or embodiments are also possible, where process A uses a larger scale bioreactor than process B (e.g., for small-scale model qualification, as discussed below).

[0018] In other embodiments, Process A and Process B may differ from each other in other (or additional) ways. For example, Process A may be a bioreactor process for producing a particular biopharmaceutical product at a first site (e.g., a first manufacturing facility), and Process B may be a bioreactor process for producing the same biopharmaceutical product at a different second site (e.g., a second manufacturing facility). Additionally or alternatively, Process A may be a bioreactor process for producing / growing a first biopharmaceutical product, and Process B may be a bioreactor process for producing / growing a different second biopharmaceutical product. In still other embodiments, Process A and Process B may involve the use of equipment other than a bioreactor, such as, for example, purification or filtration systems of different sizes.

[0019] In yet other embodiments, Process A and Process B are not biopharmaceutical processes. For example, Process A and Process B can be processes for developing or manufacturing small molecule pharmaceuticals or products, or industrial processes completely unrelated to pharmaceutical development or production (e.g., an oil refining process using Processes A and B with different operating parameters and / or different types of refinery equipment).

[0020] System 100 includes computing system 102, which in this example includes processing hardware 120, a network interface 122, a display device 124, a user input device 126, and a memory 128. Processing hardware 120 includes one or more processors, each of which may be a programmable microprocessor that executes software instructions stored in memory 128 to perform some or all of the functions of computing system 102 described herein. Alternatively, one or more of the processors in processing hardware 120 may be other types of processors (e.g., application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), etc.). Memory 128 may include one or more physical memory devices or units, including volatile and / or non-volatile memory. Any suitable type of memory may be used, such as a read-only memory (ROM), a solid-state drive (SSD), a hard disk drive (HDD), etc. In some embodiments, a portion of memory 128 stores an operating system, another portion of memory 128 stores instructions for software applications, and another portion of memory 128 stores data used and / or generated by the software applications (e.g., any of the time series data or “signals” discussed herein).

[0021] Network interface 122 may include any suitable hardware (e.g., front-end transmitter and receiver hardware), firmware, and / or software configured to communicate over one or more networks using an appropriate communications protocol. For example, network interface 122 may be or include an Ethernet interface. In general, network interface 122 may enable computing system 102 to receive data related to process A (and possibly process B and / or other processes) from one or more local or remote sources (e.g., via one or more wired and / or wireless local area networks (LANs) and / or one or more wired and / or wireless wide area networks (WANs), such as the Internet or an intranet).

[0022] Display device 124 may use any suitable display technology (e.g., LED, OLED, LCD, etc.) to present information to a user, and user input device 126 may include a keyboard or other suitable input device (e.g., a microphone). In some embodiments, display device 124 and user input device 126 are integrated into a single device (e.g., a touchscreen display). Generally, display device 124 and user input device 126 may be combined to allow a user to interact with a user interface (e.g., a graphical user interface (GUI)) generated by processing hardware 120.

[0023] As mentioned above, memory 128 may store instructions for one or more software applications. One such application is an automatic data amplification, scaling, and transfer (AD ASTRA) application 130. When executed by processing hardware 120, AD ASTRA application 130 is generally configured to generate scaling models that specify time-varying scaling relationships between process data associated with different processes, such as Process A and Process B, and to project / transfer (and possibly amplify) data between the processes using such scaling models. The process data may include time-series data indicative of one or more process input parameters, one or more process state parameters, and / or one or more process output parameters (e.g., one value per day, one value per hour, etc.) over some time interval. The processes from which and to which AD ASTRA application 130 transfers data are referred to herein as the “source process” and the “target process,” respectively, and the data associated with these processes are referred to herein as the “source data” (or “source time-series data,” etc.) and the “target data” (or “target time-series data,” etc.), respectively. Thus, in the example of Figure 1, process A is the source process and process B is the target process.

[0024] The AD ASTRA application 130 includes a scaling model generation unit 140 configured to generate a scaling model based on at least one experimental data set from each of process A and process B. The AD ASTRA application 130 also includes a data conversion unit 142 configured to transfer / scale data from process A to process B using the generated scaling model. In some embodiments, the AD ASTRA application 130 is flexible enough to generate scaling models and transfer / scale data for a variety of source / target processes and / or use cases. The AD ASTRA application 130 also includes a user interface unit 144 configured to generate a user interface (which can be presented on the display device 124) that allows a user to interact with the scaling / conversion process. For example, the user interface may allow a user to manually select source and / or target processes / data sets, set parameters that alter (i.e., amplify) the variance of the source data, and / or view the source and / or target data (and / or metrics associated with that data).

[0025] The parameters manipulated by the AD ASTRA application 130 depend on the nature of process A and process B, as well as the use case. For example, one common use case is to develop a machine learning model that predicts or infers product quality attributes or other parameters (e.g., yield, titer, future glucose or other metabolite concentrations, etc.) of process B based on measurable medium profiles and / or other parameters (e.g., pH, temperature, current metabolite concentrations, etc.) of process B, in order to control particular inputs to process B (e.g., glucose feed rate) or for other purposes (e.g., to aid in the design of process B). If few experiments for process B have been performed, it may be difficult or impossible to create such a reliable predictive or predictive model using only experimental data from process B. Therefore, the scaling model generation unit 140 may generate a scaling model that transfers process A data reflecting parameters (e.g., pH, temperature, current metabolite concentrations, etc.) used as inputs to the predictive or predictive model to similar data for process B. Various exemplary use cases are discussed in more detail below.

[0026] It is understood that other configurations and / or components may be used instead of (or in addition to) those shown in system 100 of FIG. 1 . For example, a first other computing system may send process A and process B data to computing system 102, and / or a second other computing system may receive the scaled / transferred data from computing system 102 and, optionally, use (or facilitate the use of) the scaled data (e.g., to train and / or use a machine learning model, such as the predictive or inferential model described above, or any other suitable application). Alternatively, computing system 102 may itself include these other (possibly distributed) computing devices. System 100 may also include instruments for measuring parameters in process A and / or process B (e.g., a Raman spectroscopy system with probes, flow sensors, etc.) and / or instruments for controlling parameters in process A and / or process B (e.g., a glucose pump, a device with heating and / or cooling elements, etc.).

[0027] In some embodiments, the AD ASTRA application 130 can compare any two parameters given their time series data. The techniques applied by the AD ASTRA application 130 can be purely data-driven, requiring no prior knowledge of how or whether the parameters are related at all. This can provide flexibility in addressing certain long-term data scaling issues in biopharmaceutical manufacturing or other processes, some examples of which are discussed below.

[0028] Exemplary Scaling Model We now discuss in more detail the scaling models generated by the scaling model generation unit 140 of Figure 1, according to various embodiments. To address issues with similarity-based scaling models (discussed above in the background section), the scaling model generation unit 140 applies an improved data-based framework to calculate optimal (in some embodiments) scaling between arbitrary variables.

[0029] First, {X t} t∈N and {Y t} t∈N Let denote two general signals assumed to be related to the following model:

number

number

number

number

number

number

[0030] The model in Equation 2a is referred to herein as the scaling model because it establishes a scaling relationship between the signals, where:

number

number

number

number

number

number

number

[0031] As deterministic, {θ t} t∈N Unlike frequentist methods (e.g., OLS or ML) that assume {θ t} t∈Nis considered as a random variable, and has an initial density θ t We have ~p(θ0). The initial density captures the a priori information available about the parameters. For example, if the scaling parameters are assumed to be within some interval constraints, we can define a uniform or Gaussian density over a given interval.

[0032] Given p(θ0) and D, the Bayesian approach is to t} t∈N We seek to compute the posterior density for D. The posterior density can be constructed in both real-time (or "online") and non-real-time (or "offline") settings. To distinguish between these two settings, we refer to D 1:t ={x 1:t ,y 1:t}, where x 1:t =[x1,x2,…,x t ] T and y 1:t =[y1,y2,…,y t ] T Here, for real-time estimation in Equation 5b, the filtered posterior density {p(θ t |D 1:t )} t∈N is calculated recursively. The filtering density is D 1:t Given the unknown parameters θ t encapsulates all information about p(θ t |D 1:t ), only information up to time t is used. The filtering formulation is particularly useful in applications where real-time scaling relationships are required. For offline estimation, the Bayesian method is based on the smoothed posterior density {p(θ t |D 1:t )} t∈N Again, we try to calculate p(θ t |D 1:t ), all information up to time T is used. For ease of explanation, real-time learning is addressed here; however, it is understood that similar techniques and / or calculations can be used for offline learning.

[0033] To calculate the filtering density of the parameters, Equation 5b is expressed using a stochastic state space model (SSM) formulation as shown below:

number

number

number

number

number

number

number

number

number

[0034] In the SSM formulation of the scaling model in Eqs. 6a and 6b, {θ t} t∈N represents the state, and {Y t} t∈N is the measurement value, and {X t} t∈N is a parameter. In Equations 6a and 6b, {θ t} t∈N and {Y t} t∈N are respectively,

number

number

number

number

number

number

number

number

number

number

number

[0035] The Kalman filter propagates the mean and covariance functions (sufficient statistics of a Gaussian distribution) through the update (Equation 13a) and prediction (Equation 13b) steps to compute the posterior density in Equation 13c. This is outlined in Algorithm 1 below. The Kalman filter yields the minimum mean squared error for the state estimation problem in Equations 6a and 6b. In other words, Algorithm 1 is optimal in MSE for all t∈N. See Chen et al., 2003, Statistics, 182:1-69. Furthermore, conditioning on past measurements (Equation 11) reduces the impact of noisy measurements and parameters on the state estimate.

[0036] In some embodiments, Algorithm 1, which may be implemented by the AD ASTRA application 130, is as follows:

number

number

number

number

[0037] 3 is a flow diagram of an example method 300 for scaling data across different processes. Method 300 may be performed, in whole or in part, by, for example, computing system 102 of FIG. 1 (e.g., by processing hardware 120 in executing instructions of AD ASTRA application 130 stored in memory 128).

[0038] At block 302, first time series data indicative of one or more parameters of a first process is obtained. The first time series data indicative of one or more input parameters (e.g., feed rate), state parameters (e.g., metabolite concentrations), and / or output parameters (e.g., yield) of the first process. Block 302 may include retrieving the first time series data from a database in response to a user selecting a particular data set, for example, via user input device 126, display device 124, and user interface unit 144. The parameters represented by the first time series data may be, for example, parameters of any of the "source" data sets described above with reference to the various use cases.

[0039] At block 304, second time series data indicative of one or more parameters of the second process is obtained. The second time series data indicative of one or more input, state, and / or output parameters of the second process (e.g., the same types of parameters as those obtained at block 302 for the first process). Block 304 may include retrieving the second time series data from the database in response to a user selecting a particular data set, for example, via user input device 126, display device 124, and user interface unit 144. The parameters represented by the second time series data may be, for example, parameters of any of the “target” data sets described above with reference to the various use cases.

[0040] In block 306, a scaling model is generated that specifies the time-varying scaling relationship between the parameters of the first and second processes. The scaling model may be, for example, any of the models disclosed herein (with time-varying scaling) for any of the use cases discussed above, or another suitable scaling model built on similar principles. Preferably, the scaling model is a probabilistic estimator, such as the Kalman filter (or extended Kalman filter, etc.) described above.

[0041] In block 308, source time series data associated with the source process is transferred to target time series data associated with the target process using the scaling model generated in block 306. The source time series data represents one or more input, state, and / or output parameters of the source process over time, and the target time series data represents the input, state, and / or output parameters of the target process over time. Block 308 can occur in part or in its entirety, such as in substantially real time as the source time series data is acquired, or as a batch process.

[0042] At block 310, the target time series data is stored in memory (e.g., in a different unit, device, and / or portion of memory 128). For example, the target time series data may be stored in a local or remote training database (e.g., in an additional block of method 300) for use in training a machine learning (predictive or inferential) model for use with the target process (e.g., for monitoring, such as metabolite concentration or product screening monitoring, and / or control, such as glucose feed rate control).

[0043] As non-limiting examples, parameters represented by the first, second, source, and / or target time series data may include oxygen flow rate, pH, agitation, and / or dissolved oxygen. However, virtually any parameter is possible. In some embodiments and / or use cases, the parameters of the first / source time series data differ at least in part from the parameters of the second / target time series data, such that some source parameters are used to determine different target parameters.

[0044] In some embodiments and / or use cases, the source time series data and source process are first time series data and a first process, respectively, and / or the target time series data and target process are second time series data and a second process, respectively. However, in other embodiments and / or use cases, this is not the case. For example, the scaling model generated in block 306 may relate process A to process B, while block 308 projects / transfers a different process C to a different process D, as long as process A is sufficiently similar to process C and process B is sufficiently similar to process D (or more precisely, as long as the relationship between process A and process B is known or expected to be similar to the relationship between process C and process D). By way of example only, process A may be for a particular formulation, site, and scale, while process C may be for the same formulation and scale but for a different site. While this may reduce the accuracy of data scaling in some cases, it may nevertheless be acceptable as long as the different sites are sufficiently similar or the process is not overly sensitive to the process site.

[0045] In some embodiments and use cases, the first process and source process (which may be the same or different from one another) are associated with a first process site, and the second process and target process (which may be the same or different from one another) are associated with a second, different process site. For example, the first / source process site may be in one manufacturing facility, and the second / target process site may be in another manufacturing facility. Additionally or alternatively, the first process and source process may be associated with a first process scale (e.g., a smaller bioreactor size), and the second process and target process may be associated with a second, different process scale (e.g., a larger bioreactor size). Additionally or alternatively, the first process and source process may be a bioreactor process for growing a first biopharmaceutical product, and the second process and target process may be a bioreactor process for growing a second, different biopharmaceutical product.

[0046] In some embodiments, method 300 includes one or more other additional blocks not shown in Figure 3. For example, method 300 may include an additional block in which the target time series data is used to generate a machine learning model (e.g., a predictive or inferential neural network or a regression model, etc.) of the target process, and possibly another block in which the trained machine learning model is used to control one or more inputs to the target process (e.g., feed rate, etc.).

[0047] As another example, method 300 may include, at some point before block 308 occurs, a first additional block in which additional time-series data (indicating one or more input, state, and / or output parameters of the one or more additional processes over time) is obtained, a second additional block in which one or more additional scaling models (each specifying a time-varying scaling relationship between the input, state, and / or output parameters of the first process and the input, state, and / or output parameters of each of the one or more additional processes) are generated, and a third additional block in which, based on the scaling model from block 306 and the additional scaling models, it is determined that the parameters of the second process have the closest similarity measure to the input, state, and / or output parameters of the first process (i.e., closer than the additional processes). That determination may be made using, for example, the Kullback-Leibler Divergence (KLD) similarity measure or the Weitzman similarity measure described above.

[0048] As yet another example, method 300 may include a first additional block in which a user interface is provided to a user (e.g., by user interface unit 144 via display device 124) and a second additional block in which control settings are received from the user via the user interface. In such an embodiment, block 306 may include setting the covariance using the control settings when generating the scaling model.

[0049] 3 do not have to occur in the order shown. For example, block 304 may occur before or simultaneously with block 302, and / or block 306 may occur in real time as data is received at blocks 302 and 304.

[0050] Exemplary Use Cases While Algorithm 1 generally provides an optimal approach for extracting scaling information between target and source signals from their corresponding time-series data, the details of the approach are use-case specific. In this section, several problems in industrial biopharmaceutical manufacturing are presented, each of which can be formulated as a data scaling problem. We then demonstrate the effectiveness of Algorithm 1 on these reformulated problems. The applications / use cases discussed herein are non-limiting and can be broadly categorized into one of the following problem classes: (1) comparison of two signals, (2) comparison of multiple signals, (3) prediction of missing signals, and (4) generation of new signals. Each of these classes presents unique data scaling challenges and requires appropriate modifications of Algorithm 1.

[0051] Use Case 1: Comparing Two Signals The problem of comparing two parameters / variables (also called "signals") from time series data is one common use case. This is an important class of data scaling problems that has many practical applications in industrial biopharmaceutical manufacturing and other fields. Generally, there are different ways to compare two variables. However, here, signals are compared based on how they scale relative to each other using Algorithm 1. The effectiveness of Algorithm 1 is demonstrated below through a specific exemplary use case that can be implemented, for example, by system 100 of FIG. 1.

[0052] Use Case 1, Example A: Process Scale-Up Study A typical life cycle of commercial biopharmaceutical manufacturing involves cell culture operations at three different scales: benchtop, pilot, and commercial. The cell culture process is first developed in a benchtop bioreactor and then scaled up to a pilot-scale bioreactor, where the process design and parameters are further refined and the control strategy is improved / optimized. Finally, the cell culture process is scaled up to an industrial-scale bioreactor for commercial production. See Heath and Kiss, 2007, Biotechnology Progress, 23:46-51. At each stage of process scale-up (from benchtop to pilot scale and from pilot scale to commercial scale), the full-scale process performance of the bioreactor is continuously verified against a small-scale bioreactor. This is to ensure that the full-scale and small-scale cell cultures exhibit comparable productivity and comparable product quality attributes (PQAs). Successful scale-up operations typically result in comparable titer concentrations, viable cell densities (VCDs), metabolite profiles, and glycosylation isoform profiles in full-scale and small-scale bioreactors. This is primarily achieved by manipulating common process variables such as oxygen flow rate, pH, agitation, and dissolved oxygen. Examining how these manipulated parameters / variables compare across process scales is important for assessing full-scale equipment suitability and helps devise optimal full-scale control recipes. See Junker, 2004, Journal of Bioscience and Bio-engineering, 97:347-364; Xing et al., 2009, Biotechnology and Bioengineering, 103:733-746.

[0053] For illustrative purposes, oxygen flux profiles for biopharmaceuticals produced in pilot-scale and commercial-scale bioreactors are compared. Formally, {X t} t∈N and {Y t} t∈Nrepresent the oxygen flux profiles of a biologic produced in a pilot-scale and a commercial-scale bioreactor, respectively. Figures 4A-D show experimental results for one exemplary implementation in which the automated data amplification, scaling, and transfer techniques disclosed herein were used to estimate the target signal for a 10,000-liter commercial-scale bioreactor based on the oxygen flux profile of a biologic produced in a 300-liter pilot-scale bioreactor. In the plot of Figure 4A, the "source" signal represents the oxygen flux (normalized) measured for the 300-liter pilot-scale bioreactor, and the "target" signal represents the oxygen flux (normalized) measured for the 10,000-liter commercial-scale bioreactor. In biologics manufacturing, oxygen flux is a key manipulated variable for controlling dissolved oxygen concentration in cell cultures. As shown in Figure 4A, the oxygen flux through the commercial-scale (target) bioreactor is higher than that of the pilot-scale (source) bioreactor. This is primarily due to the larger volume and higher viable cell number in the commercial-scale bioreactor. Oxygen flow rate is a critical parameter that must be continuously monitored as the process scales. However, due to the lack of suitable mathematical tools for continuous monitoring of scale-up processes, it has traditionally been monitored only at discrete times using visual-based analysis. For example, peak oxygen values (i.e., the peak in Figure 4A, where the oxygen flow rate is greatest), which are also a critical parameter, are compared at different scales to evaluate mass transfer efficiency. Despite the availability of complete time series data in Figure 4A, comparative analysis, with the exception of this peak value analysis, is typically not performed.

[0054] To address the limitations of existing methods, the AD ASTRA application 130 uses Algorithm 1 to calculate {X t} t∈N and {Y t} t∈N can be compared continuously and in real time. First, t} t∈N and {Y t}t∈N is calculated according to the SSM of Equations 6a and 6b:

number

number

number

number

[0055] 4B and 4C show the state estimates

number

[0056] Another approach to assess the quality of the estimates obtained using Algorithm 1 is to compare the true (actual) target signal with the predicted target signal. t|t} t∈N’ is Y t|t =C t θ t|t This is calculated as

number

number

[0057] Figures 4B-D are specific to the scaling model defined in Equations 15a-15c. Changing the system parameters in Equations 15a-15c and / or 16a-b defines a new model and results in different state estimates. Because scaling is model-dependent, it is often difficult to attribute a meaningful physical interpretation to the results. For example, it is not always obvious to physically interpret the state estimates in Figures 4B and 4C so that they are consistent with the process behavior shown in Figure 4A. Nevertheless, it is often possible to attribute a mathematical interpretation to the results. In summary, Algorithm 1 applies when quantifying and analyzing the behavior of manipulated variables in scale-up operations. However, the developed tools can be generalized and used in other related applications, such as scale-down model qualification, process characterization studies (see, e.g., Tsang et al., 2014, Biotechnology Progress, 30:152-160; Li et al., 2006, Biotechnology Progress, 22:696-703), comparison of media formulations (see, e.g., Jerums et al., 2005, BioProcess Int., 3:38-44; Wurm, Nature Biotechnology, 2004, at 22:1393), and mixing efficiency in disposable and stainless steel bioreactors (see, e.g., Eibl et al., 2010, Applied Microbiology and Biotechnology, 86:41-49; Diekmann et al., 2011, BMC Proceedings, 5:103). These are important yet challenging problems in biopharmaceutical manufacturing, and the database scaling methods disclosed herein can complement existing knowledge-based solutions.

[0058] Use Case 1, Example B: Qualification of Small Models Process characterization (PC) studies are an important step in biopharmaceutical manufacturing to identify critical process parameters (CPPs), material attributes, control strategies, and design space. See Godavarti et al., 2005, Biotechnology and Bioprocessing Series, 29:69. PC studies typically involve running multiple experiments on a commercial process at various process conditions to identify the optimal design space. Because conducting many PC evaluations at a commercial scale is impractical, primarily due to economic considerations, PC studies are typically performed as bench-scale processes. To ensure that bench-scale process performance is representative of the commercial process, it is generally important to first build a qualified bench-scale process. This is because inaccurate models often lead to conclusions based on laboratory data that are not applicable at large scale, thus often leading to the failure of the validation campaign. See Varga et al., 2001, Biotechnology and Bioengineering, 74:96-107. A qualified scale-down model eliminates (or reduces) the need to conduct expensive experiments using full-scale equipment. As a result, small-scale models are widely used in PC studies, process suitability studies, manufacturing troubleshooting, viral clearance studies, raw material variability investigations, cell line selection studies, and process and media improvement studies. See FDA, 2011, "Guidance for industry process validation: General principles and practices," US Department of Health and Human Services, Rockville, MD, USA, vol. 1, pp. 1-22.

[0059] A typical cell culture process involves several operating scales, encompassing inoculum development and seed expansion through production. To ensure that small-scale and commercial processes meet the same operating window, it is important to establish equivalence between scales based on key performance parameters such as product quality, product titer, viable cell density (VCD), carbon dioxide profile, pH profile, osmolality profile, and metabolite profiles (e.g., glucose, lactate, glutamate, glutamine, ammonium). Fully qualified small-scale model and commercial processes are expected to exhibit similar profiles across all key performance parameters.

[0060] Qualification of cell culture processes is a challenging and time-consuming task that requires multiple experiments on small-scale bioreactors and careful design and tuning of control parameters. The FDA recently published a process validation guidance report (FDA, 2011, “Guidance for Industry Process Validation: General Principles and Practices,” US Department of Health and Human Services, Rockville, MD, USA, vol. 1, pp. 1–22) stating, “It is important to understand how well the model represents the commercial process, including any differences that may exist, because these differences may affect the relevance of information derived from the model.” Recently, several statistical methods, such as mean equivalence tests and multivariate statistics, have been proposed to assess the quality of small-scale models. The basic idea of an equivalence test is as follows: First, an a priori interval is defined within which the difference in the mean of some key performance parameter between two scales (small-scale and commercial) is assumed to be insignificant. Then, the difference in the mean between the two different scales is evaluated using a two-tailed t-test (TOST), which calculates a confidence interval for the difference in the means. Interscale equivalence (for selected performance parameters) is then established by comparing the confidence intervals obtained from TOST with predefined intervals. See Li et al., 2006, Biotechnology Progress, 22:696-703. Mean equivalence tests are commonly used to validate important parameters such as peak VCD, integrated VCD, final titer, and percentage of glycosylated isoforms. Most of the performance parameters validated with TOST assume a single value rather than a time series. For example, it is unclear how TOST can be used to compare time-varying metabolite concentrations at different scales.

[0061] Current methods for comparing time-varying parameters involve the use of qualitative methods. For example, metabolite profiles at different scales are often compared using visual-based methods or through simple statistics such as mean and variance. See Li et al., 2006, Biotechnology Progress, 22:696-703. While visual-based methods are simple, it is often difficult to quantify the degree of similarity (or dissimilarity) between time-varying parameters. Alternatively, multivariate statistical methods have been proposed for comparing time-varying parameters. See Tsang et al., 2014, Biotechnology Progress, 30:152-160. The key idea is as follows: First, a partial least squares (PLS) model is constructed using historical data for commercial process parameters (e.g., VCD, glucose, lactate, glutamine, glutamate, ammonium, carbon dioxide, cell viability, pH, etc.). Next, for a given PLS model, the parameters of the small-scale process are projected onto the model plane. If the small-scale model is fully qualified for the commercial process, the projected dataset can be explained by the PLS model; otherwise, discrepancies will occur. In other words, a PLS model built for a commercial process can explain the variability of the small-scale process only if the small-scale process is fully qualified. This observation is valid for volume-independent parameters such as pH, dissolved oxygen, and temperature. However, this is not necessarily true for volume-dependent parameters such as working volume, feed volume, agitation, and aeration. This is because volume-dependent parameters scale with the volume of the bioreactor. Furthermore, building a reliable PLS model for a commercial process requires access to a large amount of historical data (see Tulsyan et al., 2018, Biotechnology and Bioengineering, 115:1915-1924). This requirement contradicts the goal of building a qualified small-scale model, which is to reduce the number of experiments required for commercial processes.Finally, none of the existing methods quantify the degree of similarity, or lack thereof, in performance parameters. As stated in the 2011 FDA guidance, understanding the degree to which a small-scale model represents a commercial process can better understand the relevance of the information derived from the model.

[0062] The effectiveness of Algorithm 1 in comparing time-varying parameters arising in small-scale model qualification studies has been demonstrated. For illustrative purposes, only VCD profiles for biologics produced in small-scale and commercial-scale bioreactors were compared. However, it is understood that the proposed method can be extended to compare other performance parameters as well. Formally, {X t} t∈N and {Y t} t∈N Let σ represent the average VCD profiles in the commercial-scale and small-scale bioreactors, respectively. FIG. 5A shows normalized VCD profiles for a biologic produced in a 2000-liter commercial-scale bioreactor (herein referred to as the "source" process) and a 2-liter small-scale bioreactor (herein referred to as the "target" process). The average profile in FIG. 5A is calculated by averaging the VCD profiles across multiple small-scale and commercial-scale runs. Given the profiles in FIG. 5A, the goal is to quantify how similar (or dissimilar) the profiles are at each sample time. As mentioned above, traditional methods for comparing the time-varying performance parameters in FIG. 5A are based on either "visual" inspection or the use of rudimentary process knowledge, both of which are suboptimal methods and do not quantify the degree of similarity. To address the limitations of existing methods, the AD ASTRA application 130 (e.g., scaling model generation unit 140) may, in some embodiments, use Algorithm 1 to calculate {X t} t∈N and {Y t} t∈N can be compared continuously and in real time. First,t} t∈N and {Y t} t∈N is calculated according to the SSM of Equations 6a and 6b:

number

number

number

number

number

number

number

number

number

[0063] 5B and 5C show the state

number

number

number

number

[0064] In summary, the application of Algorithm 1 in small-scale model qualification of cell culture processes has been demonstrated. Again, the developed tool may be general and can be used in other related applications, such as process scale-up studies (see Junker, 2004, Journal of Bioscience and Bio-engineering, 97:347-364; Xing et al., 2009, Biotechnology and Bioengineering, 103:733-746), comparison of media formulations (see Jerums et al., 2005, BioProcess Int., 3:38-44; Wurm, 2004, Nature Biotechnology, 22:1393), and mixing efficiency in disposable and stainless steel bioreactors (see Eibl et al., 2010, Applied Microbiology and Biotechnology, 86:41-49; Diekmann et al., 2011, BMC Proceedings, 5:P103).

[0065] Use Case 2: Comparing Multiple Signals The developments in the previous section are now generalized to include multiple signals. Formally, a given target signal is compared to M source signals. Many problems in industrial biopharmaceutical manufacturing can be reformulated and transformed into problems requiring the comparison of multiple signals. Formally, this problem is defined as follows:

[0066] For all i = 1, 2, ..., M, {X i,t} t∈N teeth,

number

number

[0067] The next objective is to determine whether the signal is equal to the target y for all i = 1, 2, ..., M. 1:Tsource signal x based on how similar it is to i,1:T A simple approach to rank the source signals closest to the target is based on Euclidean distance. For example, (x i,1:T ,y 1:T ) is the Euclidean distance between

number

number

number

[0068] In the machine learning literature,

number

number

number

number

number

number

[0069] P and Q are respectively P~N(μ p ,Σ p ) and the KLD measure (

number

number

number

number

number

number

number

number

[0070]

number

number

number

number

number

number

number

number

number

number

[0071] In some embodiments, Algorithm 2, which may be implemented by the AD ASTRA application 130, is as follows:

number

[0072] Other similarity measures can be used instead of (or in addition to) the KLD measure to rank source signals, such as the Weitzman measure (Weitzman, 1970, US Bureau of the Census, vol. 22), the Matsushita measure (Matusita, 1955, Annals of Mathematical Statistics, pp. 631-640), and the Morishita measure (Morisita, 1959, Mem. Fac. Sci. Kyushu Univ. Series E, 3-65-80). For example, the Weitzman measure calculates the overlap between two PDFs, and the greater the overlap, the more similar the PDFs are. Mathematically, the Weitzman measure is

number

number

number

number

number

number

[0073] Similar to the KLD measure, source signals can be ranked based on the Weitzmann measure. This is done by replacing the KLD measure in Algorithm 2 with the Weitzmann measure in Equation 28. However, because Equation 22 and Equation 28 are two separate similarity measures, the ranking of source signals may differ. The framework described herein for comparing and ranking signals based on similarity is general and can be used to address several challenging problems in biopharmaceutical manufacturing amenable to reformulation that require signal comparison and ranking. For example, in Trunflo et al., the authors considered the problem of placing purchase orders for mammalian cell culture raw materials to meet biologics production requirements. See Trunflo et al., 2017, Biotechnology Progress, 33:1127-1138. The authors proposed a chemometric model that compares spectroscopic scans of raw materials obtained from multiple vendors with nominal material lots. The order is placed with the vendor whose raw material scan most closely resembles the nominal lot. While Trunflo et al. use a chemometric model to compare spectroscopic scans to nominal scans, the AD ASTRA application 130 can do the same using Algorithm 2. Indeed, an advantage of Algorithm 2 over chemometric methods such as those found in Trunflo et al. is that it does not require a model of the nominal lot. This reduces or eliminates the need to collect large numbers of historical scans of the nominal lot. As an example, consider the problem of ranking biotherapeutic proteins in a portfolio of products produced in a commercial bioreactor based on their oxygen uptake profiles. In general, this is an important class of problems in biopharmaceutical manufacturing because comparing key process variables across multiple products improves fundamental understanding of the process dynamics of different products and helps design strategies for controlling similar process parameters across different products.For example, if two biologics have similar oxygen uptake profiles, their cell growth profiles can be expected to be similar. Furthermore, knowledge of products with similar growth profiles allows engineers to develop similar strategies for controlling the process. Next, we discuss analyzing and ranking proteins using Algorithm 2.

[0074] Use Case 2, Example A: Comparison of Multiple Products Next, consider the problem of comparing oxygen flux profiles for various biologics produced in commercial bioreactors and ranking the biologics based on how their oxygen uptake profile compares to that of a reference biologic. For example, Figure 6A shows normalized oxygen flux profiles for seven biotherapeutic proteins produced in commercial bioreactors. From Figure 6A, it is clear that different biologics may have very different oxygen uptake requirements. Six of the seven profiles shown in Figure 6A (S1, S2, S3, S4, S5, S6) are for "source" biologics, and the other one (T1) is for the "target" biologic. Note that the distinction between source and target biologics is strictly mathematical and determined based on the problem formulation. In this example, consider the following: Given all the profiles in Figure 6A, the goal is to find the profile in the set (S1, S2, S3, S4, S5, S6) that is most similar to T1, or more generally, to rank the profiles in (S1, S2, S3, S4, S5, S6) based on their similarity to T1. This is an important problem because oxygen uptake is a key variable for controlling the level of dissolved oxygen in a bioreactor, and by comparing profiles between different products, process engineers can better understand and control cell growth profiles. To compare and rank the oxygen flux profiles in Figure 6A using Algorithm 2, the source and target profiles are, respectively, x i,1:T and y 1:Twhere i=1,...,M. Then, the dummy target signal

number

number

number

number

number

[0075] Figure 6B shows

number

number

number

Number

Number

Number

Number

Number

number

[0076] Weizmann measure

number

number

number

number

number

number

[0077] Use Case 2, Example B: Monitoring product screening In biopharmaceutical manufacturing, recombinant proteins are typically produced in batch or fed-batch bioreactors, where cells are cultured for 2–3 weeks to produce the protein of interest. As protein-based therapeutics continue to drive demand for cheaper, higher-volume production methods, continuous production options such as perfusion bioreactors are becoming popular choices in industry. See Wang et al., 2017, Journal of Biotechnology, 246:52–60 and Pollock et al., 2013, Biotechnology and Bioengineering, 110:206–219. Unlike batch or fed-batch bioreactors, perfusion bioreactors cultivate cells for much longer periods by continuously supplying the cells with fresh medium and removing spent medium while maintaining the cells in culture. In addition to continuously removing proteins before they are exposed to excess waste products that could cause degradation, perfusion bioreactors offer several advantages over traditional batch processes, including superior product quality, stability, scalability, and cost savings. See Wang et al., 2017, Journal of Biotechnology, 246:52-60.

[0078] Tangential flow filtration (TFF) and alternating tangential flow (ATF) systems are commonly used for product recovery in perfusion systems. TFF operations continuously pump feed from the bioreactor and back through filter channels, while withdrawing and collecting cell-free permeate. ATF systems use an alternating-flow diaphragm pump to pull feed from and push it into the bioreactor, while withdrawing cell-free permeate. See Hadpe et al., 2017, Journal of Chemical Technology and Biotechnology, 92:732-740. Cell retention devices are at the heart of any perfusion system, as they are often associated with scalability, reliability, cell viability, and efficiency in terms of cell clarification and product recovery at the desired cell density. See Wang et al., 2017, Journal of Biotechnology, 246:52-60. In industry, hollow fiber membranes are the most preferred technology for cell retention because they satisfy many of the aforementioned considerations. See Clincke et al., 2013, Biotechnology Progress, 29:754-767. Despite their widespread use, hollow fiber filtration systems are susceptible to product sieving and membrane fouling. See Mercille et al., 1994, Biotechnology and Bioengineering, 43,:833-846. Membrane fouling is a significant problem in any perfusion system, as it generally results in ineffective product recovery across the membrane and a gradual decline in permeate over time, potentially prematurely terminating its operation. See Wang et al., 2017, Journal of Biotechnology, 246:52-60.

[0079] In practice, product sieving through a hollow fiber is defined as the ratio of the protein concentration in the permeate line to the protein concentration in the bioreactor. A product sieving level of 100% indicates that all product has passed through the membrane, and a product sieving level of 0% indicates zero product recovery. Mathematically,

number

number

number

[0080] While the Equation 32 model is commonly used in practice to evaluate sieving performance, its resolution is limited. For example, much of the intraday product sieving information in Figure 7A is unavailable. This is because current technologies for real-time titer measurement or product sieving in Equation 32 are unreliable or too expensive. One approach to address the limited titer measurement is to use Raman-based chemometric models. Partial least squares (PLS) models have been used to correlate Raman spectra to protein concentrations in cell cultures. Andre et al., 2015, Analytica Chimica Acta, 892:148-152. Once PLS models are available, rapidly sampled spectral data can be used to predict protein concentrations inline. While chemometric models improve the resolution of sieving profiles, building PLS models is a tedious task that requires access to large historical datasets. Furthermore, the quality of the prediction depends on both the process and the medium concentration. While these efforts can be used for real-time applications such as closed-loop titer control in cell culture (see Matthews et al., 2016, Biotechnology and Bioengineering, 113:2416-2424), chemometric models may not be necessary for assessing membrane fouling. In this section, an alternative approach is provided for real-time monitoring of product sieving through hollow fibers, which can be implemented by the AD ASTRA application 130 by directly manipulating Raman spectra and using Algorithm 2.

[0081] First, in this example, two Raman spectroscopy probes were installed in a 50-liter perfusion bioreactor: one in the bioreactor and one in the permeate line. The Raman probes used were immersion-type probes constructed of stainless steel. The probes were connected to a RamanRXN3 (Kaiser Optical Systems, Inc.) Raman spectroscopy system / instrument. Optical excitation was performed using a laser at 785 nm, yielding approximately 200 mW of power at the output of each probe. Excitation in the far-red region of the visible spectrum yielded fluorescent signals from the culture and permeate components. Each Raman spectrum was collected with 10 accumulations and a 75-second exposure time. Dark spectrum subtraction and a cosmic ray filter were also employed. Raman spectra were measured every 15 minutes. Figure 7B shows the Raman spectra collected from the bioreactor and permeate at two different times, expressed as normalized relative intensity values. Note in Figure 7B that any differences (in the Euclidean sense) in the bioreactor and permeate spectra at a given time are due to differences in protein and metabolite concentrations across the hollow fiber membrane.

[0082] Next, instead of using a chemometric model to track changes in protein concentration, we tracked changes in the Raman spectrum. Since the Raman spectral signal implicitly contains / represents titer information, information about membrane fouling can be obtained by directly tracking the spectral signal (as it is a function of titer, see Equation 32). The AD ASTRA application 130 can do this using Algorithm 2, as follows: First,

number

number

number

number

number

number

number

number

number

number

[0083] Algorithm 2 is

number

number

number

number

number

number

number

number

[0084] Figure 7C shows the change in the saturation voltage as a function of time.

number

number

number

number

[0085] Figures 7A and 7C provide complementary illustrations of the product screening problem. For example, Figure 7A shows instantaneous product screening information, and Figure 7C shows product screening kinetics. This is because Figure 7C uses the initial membrane condition as the reference condition. If the initial membrane condition changes, the results in Figure 7C will change accordingly. Also, Figure 7A is based on the difference in titer concentration, while Figure 7C is based on the overall concentration difference, including titer and metabolite concentrations. This is because Figure 7C uses Raman spectra that encode both titer and metabolite information. If necessary, the effects of metabolite concentration and / or other media components can be mitigated by selecting a spectral region that is sensitive to titer alone.

[0086] In summary, this highlights how the problem of monitoring product sieving can be reformulated as one requiring the comparison of multiple signals, and how Algorithm 2 provides an effective practical solution to that problem. Again, in this example, Algorithm 2 was used to monitor product sieving, but the developed method is general and can be used in other applications requiring the comparison of multiple signals.

[0087] Use Case 3: Data Projection Here, we consider the data projection problem, where the goal is to project a signal (or dataset) generated at one scale (e.g., pilot scale) onto another scale (e.g., commercial scale) so that the projected signal represents the process at the new scale. Data projection is an important class of problem in biopharmaceutical manufacturing because a typical lifecycle of biologics production generates data across three different scales of operation: benchtop scale, pilot scale, and commercial scale. In such multiscale process operations, projecting datasets from one scale to another allows for early derivation of important process insights, data reuse and recycling across different scales, and a potential reduction in the experiments required at different scales. The data projection problem can be reformulated and viewed as a data scaling problem, where the goal is to rescale a signal generated at (size) scale 1 to represent the process behavior at (size) scale 2. The signals at scale 1 and scale 2 are referred to herein as source and target signals and are generally

number

[0088]

number

number

number

[0089] One approach to generating a scaling model between source and target spaces is to define the scaling model in terms of signals. For example, for any pair of source-target signals {x i,1:T ,y j,1:T}, where (i,j)∈{1,…,M}×{1,…,N}, these signals are assumed to be related according to the following scaling model: y j;t =c i,t θ i,j,t +ε i,j,t , (Equation 33) During the ceremony,

number

number

number

number

number

number

number

number

number

number

number

[0090] A Bayesian approach to projecting source signals onto target space under uncertainty is to use the posterior density

number

number

number

number

number

number

[0091] In Equation 37a, the scaling factor becomes insignificant.

number

number

number

number

number

number

[0092] Comparing Equation 40 with Equation 36, both the Bayesian and frequentist approaches compute x on the target space. m,t Note that while this yields the same point estimate for the projection of , it is also possible in a Bayesian approach to ascribe quality to the point estimate of Equation 40. This can be done using the variance of the posterior density of Equation 39.

[0093] Finally, Algorithm 3 outlines how a signal can be projected from source space to target space using the proposed method. Algorithm 3 is as follows:

number

[0094] Use Case 4: Predicting Missing Signals Here, we consider the problem of predicting the profile of a parameter / variable at a real scale by studying its behavior at other scales. Formally, this problem can be stated as follows: y 1,1:T and y 2,1:T Let y denote the variables (e.g., CO2 flow rate) of product A produced at scale 1 (e.g., pilot scale) and scale 2 (e.g., commercial scale), respectively. 1,1:T Assuming that only y is known, the goal is to 2,1:T In other words, given the dynamics of a variable at scale 1, the goal is to predict its dynamics at scale 2. Note that without a priori process knowledge or a clear understanding of the relationship between scale 1 and scale 2, solving this problem based on data alone can be non-trivial and very difficult and tedious.

[0095] y 2,1:T We discuss the application of scaling methods to predict x. First, an arbitrary product, product B, has been previously produced at scales 1 and 2. 1,1:T and x 2,1:T Let x denote the dynamics of product B at scales 1 and 2, respectively. 1,1:T and x 2,1:T Assume that the two signals y are measured and available. Then, the AD ASTRA application 130 uses information from product B to predict product A at scale 2. In some embodiments, this is done as follows: First, two signals y 2,t and x2,t are assumed to be related according to the following scaling model:

number

number

number

number

number

[0096]

number

number

number

number

number

number

number

[0097] First, regarding scale invariance, p(θ 1,t |y 1,1:t ,x 1,1:t ) and p(θ 2,t |y 2,1:t ,x 2,1:t ) denote the posterior density of the scaling coefficients between product A and product B at scale 1 and scale 2, respectively, where y 1,1:T and y 2,1:T indicates the data of product A at scales 1 and 2, respectively, and x 1,1:T and x 2,1:T denote the data of product B at scales 1 and 2, respectively. The system is scale invariant if the following relation holds for all t = 1, ..., T:

number

[0098] Next, referring to product invariance, p(θ 3,t |x 2,1:t ,x 1,1:t ) and p(θ 4,t |y 2,1:t ,y 1,1:t ) denote the posterior density of the scaling coefficients of product A and product B between scale 1 and scale 2, respectively, where y 1,1:T and y 2,1:T indicates the data of product A at scales 1 and 2, respectively, and x 1,1:T and x 2,1:T denote the data for product B at scales 1 and 2, respectively. The system is product invariant if the following relation holds for all t = 1, ..., T:

number

[0099] A schematic diagram illustrating the scaling relationship 800 between product A and product B produced at scales 1 and 2 is shown in Figure 8. The solid rectangles in Figure 8 represent the variables being measured (i.e., x 1:1:T , x 2,1:T and y1,1:T ), and the dashed rectangles represent missing variables (i.e., y 2,1:T ) The scaling between different products at different scales is indicated by arrows, which point towards the target signal. The corresponding scaling coefficients are shown next to the arrows.

[0100] In the current problem setting, it is assumed that product A and product B are scale invariant, i.e., the scaling between products is maintained across different scales. Theoretically, scale invariance is not a limiting assumption because any similarity or difference between product A and product B at scale 1 will continue to exist at scale 2 as long as products A and B are produced consistently (i.e., by maintaining the initial conditions of the process across scales). In certain scenarios, the system can be shown to be product invariant, i.e., different products scale similarly across different scales. The method proposed in this section is also valid under the product invariance assumption. Next, from the marginalization formula, p(θ) in Equation 44 can be calculated as 2,t |y 1,1:t ,x 1,1:t ) can be written as follows:

number

number

number

[0101] Comparing Equation 55 (below) with Equation 41, the true θ 2,t If there is no value, the value is θ 1,tAs will be discussed in this section, θ 1,t and θ 2,t The equivalence to is established under the assumption of scale invariance in Equation 45.

number

[0102] y 2,1:T is predicted as follows: First, the scaling θ 1,1:T x 1,1:T and y 1,1:T and then calculate between θ 1,1:T and x 2,1:T Using y 2,1:T Mathematically, θ 1,t |(y 1,1:t ,x 1,1:t )~N(θ 1,t|t ,P 1,t|t ) for all t = 1, ..., T, y 1,1:t and x 1,1:t If we assume that the a posteriori scaling between 1,t The estimate of is obtained as follows:

number

number

[0103] Equation 54 is y 2,t gives the best estimate of θ' 1,t|t However, under the assumption of scale invariance, θ' in Eq. 1,t|t is θ1,t|t and for all t = 1, ..., T,

number

[0104] An exemplary algorithm 4 that may be implemented by the AD ASTRA application 130 is as follows:

number

[0105] While the scale invariance assumption in Algorithm 4 may be theoretically unbounded, it rarely holds in practice due to inherent process and raw material variations, as well as other known and unknown process disturbances. In other words, θ' 1,t is θ 1,t This results in predictions from Equation 55 that may drift from the optimal prediction of Equation 54. To reduce such drift, a real-time implementation of Algorithm 4 allows feedback of information from Product 1 at Scale 2 to Equation 55. For k≦T, y 2,1:k Assuming that the objective is to extract the remaining y 2,k+1:T As mentioned above, the goal is to predict y 2,1:k Given θ', we use Algorithm 1 to find the scaling posterior θ' 1,t |(y2,1:t ,x 2,1:t )~N(θ' 1,t|t ,P' 1,t|t ) can be easily calculated. Here, for t≦k,

number

number

number

number

number

number

[0106] δ in Equation 56 kcompensates for the predicted drift observed in Equation 55, but the predictor in Equation 55 does not necessarily eliminate or prevent the drift in the first place. 1,t|t and θ' 1,t|t This is due to the inherent difference between θ 1,t|t To estimate, Algorithm 1 uses the dataset (x 1,1:T ,y 1,1:T ) and θ' 1,t|t To estimate, Algorithm 1 uses (x 2,1:t ,y 2,1:t ) where θ 1,t|t and θ' 1,t|t To reduce the difference between 1,t|t θ' 1,t|t This is the projection of the dataset (x 1,1:T ,y 1,1:T ) and (x 2,1:k ,y 2,1:k ) can be combined as follows: 1,t|t This is done by re-estimating

number

[0107] An exemplary algorithm 5 that may be implemented by the AD ASTRA application 130 is as follows:

number

[0108] Use Case 4, Example A: Scaling up monoclonal antibody production Here, we consider the scale-up of production of a monoclonal antibody (e.g., Product A) from a pilot-scale to a commercial-scale facility. During process scale-up, it is routine to evaluate the automation, equipment, and operational constraints at the receiving / target site to ensure the process can be successfully scaled up and the product produced to the required specifications. For example, several tasks, such as process characterization and gap analysis, are routinely performed to evaluate process parameters and constraints in consideration of increasing production volume at commercial scale. These studies typically lead to recommendations regarding process and equipment changeovers that need to be performed before the product can be transferred and produced at a commercial facility.

[0109] In this example, the pilot-scale facility was equipped with a 300-liter fed-batch stainless steel bioreactor, while the commercial facility operated a 15,000-liter fed-batch stainless steel bioreactor. To accommodate larger production volumes, the commercial bioreactor was operated under different aeration conditions than the pilot bioreactor. For example, the amount of oxygen required to maintain the target dissolved oxygen level was much higher in the commercial bioreactor than in the pilot bioreactor. To ensure that the pumps in the commercial facility can deliver the required oxygen at the required rate, it is important to first predict the oxygen demand required at commercial scale. Current methods only predict peak oxygen demand using simple volume scaling methods, instead of predicting oxygen demand at each sampling time, which could help better assess pump power ratings and other key attributes. Predictions based on volume scaling methods are, at best, approximate because these methods do not consider process disturbances or specific process configurations that may affect the actual oxygen demand in the commercial bioreactor. For example, if a commercial bioreactor is equipped with an inefficient impeller design, the actual oxygen required to maintain the target dissolved oxygen level will differ from that suggested by the volume scaling method.

[0110] Here, the oxygen demand in the commercial bioreactor at each sampling time was predicted using the scaling method described above to predict missing signals. Figure 9A shows the normalized oxygen demand for Product A in the pilot bioreactor. As shown in Figure 9A, as cells grow, the oxygen required to maintain the target dissolved oxygen level also increases. Next, an optional product, Product B, was introduced. Product B was previously produced in both pilot-scale and commercial-scale facilities. For simplicity, the oxygen demand profiles of Product B in the pilot and commercial bioreactors are not shown. Here, given the oxygen profile of Product A at pilot scale and the oxygen profiles of Product B at pilot and commercial scales, an application corresponding to an embodiment of the AD ASTRA application 130 used Algorithm 3 to predict the oxygen demand of Product A at commercial scale. Figure 9B compares the "offline" prediction from Algorithm 3 with the "actual" oxygen demand (both normalized). As shown in Figure 9B, Algorithm 3 predicts the oxygen demand at each sampling time, including the peak condition that also corresponds to the maximum VCD. Offline Prediction

number

number

number

number

[0111] The normalized values of peak oxygen demand (as normalized flow rate) predicted by the volume scaling method and Algorithm 3 are 0.691 and 0.813, respectively, compared to an actual peak demand of 0.918. Clearly, the predictions from Algorithm 3 are much more accurate compared to the predictions from the volume method. This further demonstrates the effectiveness of Algorithm 3 in predicting oxygen demand in bioreactors.

[0112] Next, Algorithm 4 was used to mitigate the offset observed in Figure 9B for the "offline" predictions of Algorithm 3. As discussed, Algorithm 4 is an "online" method that uses information from a full-scale bioreactor to improve future predictions. Figure 9B also compares the online predictions from Algorithm 4 with the actual demand. The results in Figure 9B are shown for t=300, which indicates that y 2,1:300 This means that y is assumed to be available. 2,1:300 is already known, so in Algorithm 4

number

[0113] Additional considerations regarding the present disclosure are now provided.

[0114] Some of the drawings described herein show example block diagrams having one or more functional components. It will be understood that such block diagrams are for illustrative purposes, and that the devices described and shown may have more, fewer, or alternative components than those shown. Also, in various embodiments, the components (and the functionality provided by each component) may be associated with or otherwise integrated as part of any suitable component.

[0115] Embodiments of the present disclosure relate to non-transitory computer-readable storage media having computer code thereon for performing various computer-implemented operations. The term “computer-readable storage medium” is used herein to encompass any medium capable of storing or encoding a sequence of instructions or computer code for performing the operations, methods, and techniques described herein. The medium and computer code may be those specially designed and constructed for the purposes of the embodiments of the present disclosure, or may be of the kind known and available to those skilled in the computer software arts. Examples of computer-readable storage media include, but are not limited to, magnetic media such as hard disks, floppy disks, and magnetic tape; optical media such as CD-ROMs and holographic devices; magneto-optical media such as optical disks; and hardware devices specially configured to store and execute program code, such as ASICs, programmable logic devices (“PLDs”), and ROM and RAM devices.

[0116] Examples of computer code include machine code, such as produced by a compiler, and files containing high-level code executed by a computer using an interpreter or compiler. For example, an embodiment of the present disclosure may be implemented using Java, C++, or other object-oriented programming languages and development tools. Further examples of computer code include encrypted code and compressed code. Moreover, an embodiment of the present disclosure may be downloaded as a computer program product, which may be transferred from a remote computer (e.g., a server computer) to a requesting computer (e.g., a client computer or another server computer) over a transmission channel. Other embodiments of the present disclosure may be implemented in hardwired circuitry in place of, or in combination with, machine-executable software instructions.

[0117] As used herein, the singular terms "a," "an," and "the" can include plural referents unless the context clearly dictates otherwise.

[0118] As used herein, the terms "nearly," "substantially," "substantial," and "about" are used to describe and account for small variations. When used in conjunction with an event or circumstance, these terms can refer not only to instances in which the event or circumstance occurs exactly, but also to instances in which the event or circumstance occurs approximately. For example, when used in conjunction with a numerical value, these terms can refer to a range of variation of ±10% or less of the numerical value, e.g., ±5% or less, ±4% or less, ±3% or less, ±2% or less, ±1% or less, ±0.5% or less, ±0.1% or less, or ±0.05% or less. For example, two numerical values can be considered "substantially" identical if the difference is ±10% or less of the mean of the numerical values, e.g., ±5% or less, ±4% or less, ±3% or less, ±2% or less, ±1% or less, ±0.5% or less, ±0.1% or less, or ±0.05% or less.

[0119] Additionally, amounts, ratios, and other numerical values may be presented herein in a range format. It should be understood that such range format is used for convenience and brevity and should be understood flexibly to include not only the numerical values explicitly stated as the limits of a range, but also to include all individual numerical values or subranges subsumed within that range as if each numerical value and subrange were expressly stated.

[0120] While the present disclosure has been described and illustrated with reference to specific embodiments, these descriptions and illustrations are not intended to limit the disclosure. Those skilled in the art will recognize that various changes may be made and equivalents may be substituted without departing from the true spirit and scope of the present disclosure, as defined by the appended claims. Illustrations may not necessarily be to scale. Differences between the artistic representations in this disclosure and the actual devices may occur due to manufacturing processes, tolerances, and / or other reasons. There may be other embodiments of the present disclosure not specifically illustrated. The specification (other than the claims) and drawings should be considered illustrative rather than restrictive. Modifications may be made to adapt particular situations, materials, compositions, techniques, or processes to the objective, spirit, and scope of the present disclosure. All such modifications are intended to be within the scope of the claims appended hereto. While the technology disclosed herein has been described with reference to particular operations performed in a particular order, it will be understood that these operations may be combined, subdivided, or reordered to form equivalent techniques without departing from the teachings of the present disclosure. Accordingly, unless specifically indicated herein, the order and grouping of the operations is not a limitation of the present disclosure.

Claims

1. A method for scaling data across different processes, One or more processors acquire first time-series data showing one or more input, state, and / or output parameters of a first process over time, The one or more processors acquire second time-series data indicating the input, state, and / or output parameters of one or more second processes over time, The one or more processors generate a scaling model that specifies the time-varying scaling relationship between the input, state, and / or output parameters of the first process and the input, state, and / or output parameters of the second process, Transferring source time series data related to a source process to target time series data related to a target process using one or more processors and the scaling model, wherein the source time series data represents one or more input, state and / or output parameters of the source process over time, and the target time series data represents one or more input, state and / or output parameters of the target process over time. The one or more processors store the target time-series data in memory. Methods that include...

2. The method according to claim 1, wherein the scaling model includes a probabilistic estimator.

3. The method according to claim 2, wherein the probabilistic estimator is a Kalman filter.

4. The method according to claim 1, wherein the first process is the source process and the second process is the target process.

5. The method according to claim 1, wherein the first process and the source process are associated with a first process site, and the second process and the target process are associated with a second process site different from the first process site.

6. The method according to claim 1, wherein the first process and the source process are associated with a first process scale, and the second process and the target process are associated with a second process scale different from the first process scale.

7. The method according to claim 6, wherein the first process and the source process are bioreactor processes using a first bioreactor size, the second process and the target process are bioreactor processes using a second bioreactor size, and the first bioreactor size is smaller than the second bioreactor size.

8. The one or more processors train a machine learning model of the target process using the target time series data. Using the trained machine learning model, control one or more inputs to the target process. The method according to claim 1, further comprising:

9. The first process, the second process, the source process, and the target process are bioreactor processes, The method according to claim 1, wherein the input, state, and / or output parameters of the first process, the second process, the source process, and the target process include oxygen flow rate, pH, stirring, and / or dissolved oxygen.

10. The method according to claim 1, wherein the first process and the source process are bioreactor processes for growing a first biopharmaceutical product, and the second process and the target process are bioreactor processes for growing a second biopharmaceutical product different from the first biopharmaceutical product.

11. Before transferring the aforementioned source time-series data, The aforementioned one or more processors acquire additional time-series data showing the input, state, and / or output parameters of one or more additional processes over time, The one or more processors each generate one or more additional scaling models that specify a time-varying scaling relationship between the input, state, and / or output parameters of the first process and the respective input, state, and / or output parameters of the one or more additional processes, The one or more processors determine, based on the scaling model and the one or more additional scaling models, that the input, state, and / or output parameters of the second process have the closest similarity measure to the input, state, and / or output parameters of the first process. The method according to claim 1, further comprising:

12. The method according to claim 11, wherein determining that the input, state and / or output parameters of the second process have the closest similarity measure to the input, state and / or output parameters of the first process includes using the Kullback-Leibler divergence (KLD) similarity measure or the Weizmann similarity measure.

13. The first process is a first bioreactor process at a first process scale, The second process is a second bioreactor process at a second process scale, The source process is the third bioreactor process at the first process scale, The target process is the fourth bioreactor process at the second process scale, The first, second, third, and fourth bioreactor processes are different processes. The method according to claim 1, wherein the first process scale is different from the second process scale.

14. The method according to claim 1, wherein at least part of transferring the source time series data to the associated target time series data occurs substantially in real time when the source time series data is acquired.

15. The provision of a user interface to the user by the aforementioned one or more processors and via a display device, The control settings are received by the one or more processors and from the user via the user interface. It further includes, The method according to claim 1, wherein generating the scaling model includes setting the covariance using the control settings when generating the scaling model.

16. One or more processors, One or more computer-readable media for storing instructions and A system comprising, wherein when the instruction is executed by the one or more processors, the system causes the one or more processors to perform the method according to any one of claims 1 to 15.